RequantTESTNET
Requant research

Measurements

What has been measured for TNet v1, with the conditions of every figure. These results describe the listed devices and configurations only; they are not guarantees for other hardware, drivers or settings. Raw data is kept outside the repositories; the research notes record the source hashes and environments.

GPU attempts (research benchmark)

One attempt of 65,536 rows through the eight layers, measured with the research benchmark tnet_bench.cu using cuBLAS int8 GEMM, requantization as a separate pass and no fused epilogue.

CMP 50HX (Turing, sm_75)RTX 3090 (Ampere, sm_86)
Attempt time611.6 ms334.5 ms
of which int8 GEMM539.6 ms290.1 ms
of which requantization41.2 ms23.9 ms
of which input expansion20.1 ms13.6 ms
of which ticket hashing10.7 ms6.8 ms
Share on tensor-core GEMM88.2%86.7%
Attempt throughput57.5 TMAC/s105.2 TMAC/s
Cost per ticket292 ns159.5 ns
Tickets per second (derived)about 3.4 millionabout 6.3 million
Conditionsdriver 610.43.03, CUDA 13.3; 1 s warm-up, 3 attemptsrented machine, driver 570.211.01, CUDA 12.8; power-limited at a median of 328 W, SM about 1,628 MHz; median of 7

The RTX 3090 is 1.83× faster per ticket than the CMP 50HX. From coarse board-power sampling (1 s), the RTX 3090 used about 52 µJ per ticket.

Sources: tnet-v2, tnet-ampere-v1.

Mining software (CPPminer)

CPPminer replaces cuBLAS with built-in CUTLASS int8 tensor-core kernels, so it needs no extra libraries.

GPUTickets per secondConditions
CMP 50HX (Turing)3.6–3.7 millionCPPminer v0.5-fork.9 and v0.5-fork.10, Linux, CUDA 12.9 packages; in the test-network pool with every share accepted, and solo against a node
RTX 5070 (Blackwell)4.5–5.3 millionan earlier cuBLAS build, in the pool

The RTX 3090 figure above comes from the research benchmark, not from CPPminer. CPPminer's kernel for Ampere and newer has not yet been benchmarked against cuBLAS; its correctness is checked against the CPU at every start.

Sources: CPPminer TNet guide, v0.5-fork.10 release notes.

Batching

Rows are independent, so a miner may use any batch size, but small batches waste tensor-core throughput.

  • Mining one row at a time costs 42–45× more per ticket on the CMP 50HX and 51× more on the RTX 3090.
  • On the RTX 3090, a batch of 256 rows costs 1.6× more per ticket than a full batch.

CPU verification

Verifying a block recomputes one row through the eight layers with the epoch weights already prepared. All figures on an AMD Ryzen 7 8745HS laptop processor, release builds, median of 7.

Implementation1 thread8 threads
Requant node (portable build, AVX2 selected at run time)21.7 ms11.2 ms
Abacus reference, target-cpu=native, transposed weights16.8 ms11.5 ms
Abacus reference, baseline x86-64, transposed weights81.8 ms17.6 ms
Abacus reference, target-cpu=native, row-major weights46.1 ms15.5 ms
Abacus reference, baseline x86-64, row-major weights109.1 ms29.6 ms

Once per epoch every node derives the 512 MiB of weights (5.0 s single-threaded, parallelisable) and transposes them (2.1 s). In the running node a whole block, including transactions, verifies in about 0.1 s.

Sources: SPEC.md §9, tnet-v2.

Bit-exact parity

  • GPU-found tickets were recomputed byte for byte by the Rust reference on both architectures: on the CMP 50HX (nonce 0, row 255, piece 23) with 15 leading zero bits; on the RTX 3090 (0, 1751, 18) and (1, 6693, 12), also checked with the Requant node's tnet check.
  • Both architectures produced identical statistics of the last layer: 0.73% zeros, 2.08% saturated values, mean |x| of 43.45.
  • CPPminer's kernels match the Rust reference byte for byte on Turing and Ampere.
  • Frozen test vectors cover rows 0, 1, 255 and 65535 at the real parameters, and a small instance checked independently in Python.

Lottery fairness

Tickets are hashes, so the number of winning tickets should follow the expected count.

DeviceThresholdFoundExpected
CMP 50HX14 leading zero bits388384
RTX 309014 leading zero bits902896
RTX 309020 leading zero bits43.75

Robustness: no cheap approximation

Measured with a NumPy probe on uniform int8 weights (8 rows), not the consensus derivation.

One error spreads. A single ±1 error after layer 1 changes this share of values in later layers:

After layer2345678
Values that differ1.0%7.4%21%35%44%50%55%

Approximations fail.

  • Dropping the 64 smallest terms of the last layer leaves only 25.8% of pieces exact; dropping 512 leaves none.
  • Using int7 weights leaves no piece exact.
  • Skipping exact zeros (0.7% of terms) is the only free saving, and tensor cores cannot exploit it: structured sparsity needs 50%.

Precomputation is bounded, not measured. Per-value lookup tables would need 64 GiB per layer at n = 8192 for no saving. Bit-plane lookup tables are 9–12× slower than tensor cores at group size 8 and 4–6× at group size 16 on the measured GPU, from arithmetic throughput alone; in silicon they trade an int8 multiplier for at least 128× the weight storage and bandwidth. No lookup-table kernel has been written yet (open questions).

Source: tnet-v2.

Earlier candidates, for comparison

MeasurementResultSource
Goldilocks matrix kernel, CMP 50HXabout 175 GMAC/sgpu-suite-v1
Exact int8 GEMM on tensor cores, CMP 50HX, n = 819277.7 TMAC/s, about 440× the Goldilocks kernelint8-matmul-v1
Per-attempt succinct commitment for int8 GEMM19–57× the cost of the producta8-proof-cost-v1
Freivalds block check at n = 256, 1 thread19.1 msverifier-throughput-v2
Interactive attestation certified throughput, n ≥ 819240–49 TMAC/sattest-v1