Measurements
What has been measured for TNet v1, with the conditions of every figure. These results describe the listed devices and configurations only; they are not guarantees for other hardware, drivers or settings. Raw data is kept outside the repositories; the research notes record the source hashes and environments.
GPU attempts (research benchmark)
One attempt of 65,536 rows through the eight layers, measured with the research benchmark tnet_bench.cu using cuBLAS int8 GEMM, requantization as a separate pass and no fused epilogue.
| CMP 50HX (Turing, sm_75) | RTX 3090 (Ampere, sm_86) | |
|---|---|---|
| Attempt time | 611.6 ms | 334.5 ms |
| of which int8 GEMM | 539.6 ms | 290.1 ms |
| of which requantization | 41.2 ms | 23.9 ms |
| of which input expansion | 20.1 ms | 13.6 ms |
| of which ticket hashing | 10.7 ms | 6.8 ms |
| Share on tensor-core GEMM | 88.2% | 86.7% |
| Attempt throughput | 57.5 TMAC/s | 105.2 TMAC/s |
| Cost per ticket | 292 ns | 159.5 ns |
| Tickets per second (derived) | about 3.4 million | about 6.3 million |
| Conditions | driver 610.43.03, CUDA 13.3; 1 s warm-up, 3 attempts | rented machine, driver 570.211.01, CUDA 12.8; power-limited at a median of 328 W, SM about 1,628 MHz; median of 7 |
The RTX 3090 is 1.83× faster per ticket than the CMP 50HX. From coarse board-power sampling (1 s), the RTX 3090 used about 52 µJ per ticket.
Sources: tnet-v2, tnet-ampere-v1.
Mining software (CPPminer)
CPPminer replaces cuBLAS with built-in CUTLASS int8 tensor-core kernels, so it needs no extra libraries.
| GPU | Tickets per second | Conditions |
|---|---|---|
| CMP 50HX (Turing) | 3.6–3.7 million | CPPminer v0.5-fork.9 and v0.5-fork.10, Linux, CUDA 12.9 packages; in the test-network pool with every share accepted, and solo against a node |
| RTX 5070 (Blackwell) | 4.5–5.3 million | an earlier cuBLAS build, in the pool |
The RTX 3090 figure above comes from the research benchmark, not from CPPminer. CPPminer's kernel for Ampere and newer has not yet been benchmarked against cuBLAS; its correctness is checked against the CPU at every start.
Sources: CPPminer TNet guide, v0.5-fork.10 release notes.
Batching
Rows are independent, so a miner may use any batch size, but small batches waste tensor-core throughput.
- Mining one row at a time costs 42–45× more per ticket on the CMP 50HX and 51× more on the RTX 3090.
- On the RTX 3090, a batch of 256 rows costs 1.6× more per ticket than a full batch.
CPU verification
Verifying a block recomputes one row through the eight layers with the epoch weights already prepared. All figures on an AMD Ryzen 7 8745HS laptop processor, release builds, median of 7.
| Implementation | 1 thread | 8 threads |
|---|---|---|
| Requant node (portable build, AVX2 selected at run time) | 21.7 ms | 11.2 ms |
Abacus reference, target-cpu=native, transposed weights | 16.8 ms | 11.5 ms |
| Abacus reference, baseline x86-64, transposed weights | 81.8 ms | 17.6 ms |
Abacus reference, target-cpu=native, row-major weights | 46.1 ms | 15.5 ms |
| Abacus reference, baseline x86-64, row-major weights | 109.1 ms | 29.6 ms |
Once per epoch every node derives the 512 MiB of weights (5.0 s single-threaded, parallelisable) and transposes them (2.1 s). In the running node a whole block, including transactions, verifies in about 0.1 s.
Sources: SPEC.md §9, tnet-v2.
Bit-exact parity
- GPU-found tickets were recomputed byte for byte by the Rust reference on both architectures: on the CMP 50HX
(nonce 0, row 255, piece 23)with 15 leading zero bits; on the RTX 3090(0, 1751, 18)and(1, 6693, 12), also checked with the Requant node'stnet check. - Both architectures produced identical statistics of the last layer: 0.73% zeros, 2.08% saturated values, mean |x| of 43.45.
- CPPminer's kernels match the Rust reference byte for byte on Turing and Ampere.
- Frozen test vectors cover rows 0, 1, 255 and 65535 at the real parameters, and a small instance checked independently in Python.
Lottery fairness
Tickets are hashes, so the number of winning tickets should follow the expected count.
| Device | Threshold | Found | Expected |
|---|---|---|---|
| CMP 50HX | 14 leading zero bits | 388 | 384 |
| RTX 3090 | 14 leading zero bits | 902 | 896 |
| RTX 3090 | 20 leading zero bits | 4 | 3.75 |
Robustness: no cheap approximation
Measured with a NumPy probe on uniform int8 weights (8 rows), not the consensus derivation.
One error spreads. A single ±1 error after layer 1 changes this share of values in later layers:
| After layer | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|
| Values that differ | 1.0% | 7.4% | 21% | 35% | 44% | 50% | 55% |
Approximations fail.
- Dropping the 64 smallest terms of the last layer leaves only 25.8% of pieces exact; dropping 512 leaves none.
- Using int7 weights leaves no piece exact.
- Skipping exact zeros (0.7% of terms) is the only free saving, and tensor cores cannot exploit it: structured sparsity needs 50%.
Precomputation is bounded, not measured. Per-value lookup tables would need 64 GiB per layer at n = 8192 for no saving. Bit-plane lookup tables are 9–12× slower than tensor cores at group size 8 and 4–6× at group size 16 on the measured GPU, from arithmetic throughput alone; in silicon they trade an int8 multiplier for at least 128× the weight storage and bandwidth. No lookup-table kernel has been written yet (open questions).
Source: tnet-v2.
Earlier candidates, for comparison
| Measurement | Result | Source |
|---|---|---|
| Goldilocks matrix kernel, CMP 50HX | about 175 GMAC/s | gpu-suite-v1 |
| Exact int8 GEMM on tensor cores, CMP 50HX, n = 8192 | 77.7 TMAC/s, about 440× the Goldilocks kernel | int8-matmul-v1 |
| Per-attempt succinct commitment for int8 GEMM | 19–57× the cost of the product | a8-proof-cost-v1 |
| Freivalds block check at n = 256, 1 thread | 19.1 ms | verifier-throughput-v2 |
| Interactive attestation certified throughput, n ≥ 8192 | 40–49 TMAC/s | attest-v1 |