secp256k1 Sequential Scan — GPU Throughput (with working math)
Hello, I built an OpenCL implementation of secp256k1 that can process millions to billions per second, generate a public keys on a modern GPU.
https://github.com/ipsbruno3/secp256k1-gpu-accelerator
TL;DR
-
Per-GPU (RTX 5090): ~500–600M public keys/second minimum
-
12× rig: ~7.2–10.0B public keys/second observed
-
Per day: with 600M×12 → $$6.2208\times 10^{14}$$ keys/day; with 10B×12 → $$8.64\times 10^{14}$$ keys/day
-
wNAF: Larger windows improved scalar-mult throughput (trade VRAM for fewer additions)
To-do
- Used-address hash set: O(1) membership checks; increases detection, not raw throughput
Throughput math
Let:
-
$$R_{\mathrm{gpu}}$$ = per-GPU public keys per second
-
$$G$$ = number of GPUs
- $$T_{\mathrm{day}} = 86{,}400\ \mathrm{s}$$
- $$R_{\mathrm{rig}} = R_{\mathrm{gpu}}\cdot G$$
- $$K_{\mathrm{day}} = R_{\mathrm{rig}}\cdot T_{\mathrm{day}}$$
Example (600M keys/s per GPU, 12 GPUs):
$$
R_{\mathrm{gpu}} = 6.0\times 10^{8}\ \frac{\text{keys}}{\text{s}},\quad
G=12,\quad
T_{\mathrm{day}}=86{,}400\ \text{s}
$$
$$
R_{\mathrm{rig}} = 6.0\times 10^{8}\cdot 12 = 7.2\times 10^{9}\ \frac{\text{keys}}{\text{s}}
$$
$$
K_{\mathrm{day}} = 7.2\times 10^{9}\cdot 86{,}400
= 6.2208\times 10^{14}\ \text{keys/day}
$$
per day
And if you multiplier with hashtable used iaddress (53 millions)
per day validations (don’t multiply by set size)
Hit probability (order-of-magnitude)
With $$N \approx 5.3\times 10^{7}$$ known used addresses and address space $$M=2^{160}$$:
$$
\mathbb{E}[\text{hits/day}] \approx K_{\mathrm{day}}\cdot \frac{N}{M}.
$$
Using $$K_{\mathrm{day}}=6.2208\times 10^{14}$$:
$$
\mathbb{E}\approx \frac{6.2208\times 10^{14}\cdot 5.3\times 10^{7}}{2^{160}}
\approx 2.26\times 10^{-26}\ \text{hits/day}.
$$
Brute force on random keys stays infeasible. You only get traction with constrained keyspaces (partial seeds, weak RNGs, human patterns, etc.).
Implementation notes
-
wNAF window: Larger $$w$$ reduced additions and improved throughput on 5090s; optimal $$w$$ depends on VRAM vs. occupancy.
-
Used-address filter: Keep it as an in-memory hash set; serialize once and memory-map for fast cold starts.
-
I/O: Batch keys → compress/pack → single pass membership checks to avoid cache thrash.
-
Ethics: Test only against keys you own.
Other projects
https://github.com/ipsbrunoreserva/bitcoin_cracking
— High-performance PBKDF2-HMAC-SHA512 (OpenCL). WIP.
https://github.com/ipsbruno3/bitcoin_cracking_final
— Continuation of @ipsbrunoreserva/bitcoin_cracking_final. WIP (Final Version).
https://github.com/ipsbruno3/bitcoin_electrum_cracking
— Electrum seed verification trick + sequential BIP-39 scan + used-address membership checks (Works 100%)
secp256k1 Sequential Scan — GPU Throughput (with working math)
Hello, I built an OpenCL implementation of secp256k1 that can process millions to billions per second, generate a public keys on a modern GPU.
https://github.com/ipsbruno3/secp256k1-gpu-accelerator
TL;DR
To-do
Throughput math
Let:
Example (600M keys/s per GPU, 12 GPUs):
And if you multiplier with hashtable used iaddress (53 millions)
Hit probability (order-of-magnitude)
With$$N \approx 5.3\times 10^{7}$$ known used addresses and address space $$M=2^{160}$$ :
Using$$K_{\mathrm{day}}=6.2208\times 10^{14}$$ :
Brute force on random keys stays infeasible. You only get traction with constrained keyspaces (partial seeds, weak RNGs, human patterns, etc.).
Implementation notes
Other projects
https://github.com/ipsbrunoreserva/bitcoin_cracking
— High-performance PBKDF2-HMAC-SHA512 (OpenCL). WIP.
https://github.com/ipsbruno3/bitcoin_cracking_final
— Continuation of @ipsbrunoreserva/bitcoin_cracking_final. WIP (Final Version).
https://github.com/ipsbruno3/bitcoin_electrum_cracking
— Electrum seed verification trick + sequential BIP-39 scan + used-address membership checks (Works 100%)