# Aere Network: Reproducible Benchmark Report (spec-18) **Status date:** 2026-07-19 **Chain:** Aere Network mainnet, chain ID 2800 (Hyperledger Besu QBFT). **Scope:** chain performance, zkVM proof-verification economics, and post-quantum cryptographic-operation gas, stated with a hard separation between what is freshly measured, what is cited from a committed repo artifact, and what still needs a measurement command run. ## The absolute rule this report obeys No performance number (TPS, latency, gas, proof size, proving time) appears here unless it was either (a) freshly measured by running a command this session, or (b) found in a committed repo artifact that is cited by path. Anything that cannot be measured or sourced is written as a measurement command and marked `[MEASURE]`. Anything sourceable only from memory and not yet confirmed against code is marked `[VERIFY]`. A fabricated benchmark would be a catastrophic failure of this document, so the labels below are load-bearing, not decoration. ### Label legend | Label | Meaning | |---|---| | `[MEASURED-FRESH]` | Run by this report on this machine, output shown or summarized. | | `[CITED: path]` | Value read from a committed repo artifact at the cited path. | | `[MEASURE: cmd]` | Not measured here. The command that would measure it is given. | | `[VERIFY: note]` | Sourceable but needs confirmation against live code or a live node. | ### Scope boundary that never moves Aere Network validators sign classical secp256k1 QBFT consensus messages. Nothing in this report makes Aere consensus post-quantum. The post-quantum work is an application and account layer capability. Any claim of post-quantum consensus would be false and is not made here. --- ## PART A: CHAIN PERFORMANCE ### A.0 The one honesty point, stated first The "273,000 TPS" figure is an ARCHITECTURAL CEILING derived from block-space arithmetic under ideal conditions. It is NOT the live mainnet transaction rate, it has never been a measured result, and it is not asserted as one. Three distinct quantities must never be conflated: | Quantity | Value | Label | |---|---|---| | Architectural ceiling (design maximum) | ~273,000 TPS | `[CITED]`, explicitly labeled a ceiling | | Testnet-measured L1 throughput | none exists | see A.1 | | Live-mainnet-observed throughput | far below the ceiling, demand-limited | see A.1 | `[CITED: aerenew/docs/WHITEPAPER-V2.md lines 73, 117, 633; aerenew/docs/wp2-sections/01-overview.md:59; 02-architecture.md:21; 09-governance-roadmap.md:294]` The whitepaper states it verbatim: the block parameters "admit a design ceiling on the order of 273,000 transactions per second; this is a theoretical maximum implied by the gas and block configuration, not a measured or sustained figure, and realized throughput on the live network is a small fraction of it." It is described as "a design ceiling derived from block-space arithmetic under ideal conditions" that "has never been a measured result." **On the derivation itself.** The repo labels 273,000 as a block-space-arithmetic ceiling but does not publish the explicit factors. The two facts that are sourceable: - Block period 0.5 s. `[CITED: aerenew/docs/wp2-sections/02-architecture.md:19; genesis-subsecond.json `config.transitions.qbft[0]` = `{ "block": 2137652, "xblockperiodmilliseconds": 500 }`]` The earlier form of this citation read `blockperiodmilliseconds: 500`, without the mandatory `x` prefix. Besu never reads that key, so the citation described a setting that would have had no effect. Corrected 2026-07-31 against the file on disk. See `aerenew/docs/TRANZITIA-QBFT-SETTLED-2026-07-20.md`. - Genesis block gas limit is ambiguous across genesis artifacts: `0x2625a00` (40,000,000) in `aerenew/genesis.json`, versus `0x1fffffffffffff` (9,007,199,254,740,991, effectively unbounded) in `aere-genesis-current.json` and `genesis-subsecond.json`. **RESOLVED by measurement 2026-08-01: the effectively-unbounded value is the live one.** `eth_getBlockByNumber("latest")` on `https://rpc.aere.network` at head block 11,810,041 returns header `gasLimit` = `0x1fffffffffffff` = 9,007,199,254,740,991. The 40,000,000 artifact is not what chain 2800 runs. Anyone can re-check this in one RPC call. Implied arithmetic (transparent, not a new claim): 273,000 TPS at a 0.5 s block is 136,500 tx per block, and at the 21,000-gas simple-transfer floor that is about 2.87 billion gas per block. That is far above the 40M limit and far below the effectively-unbounded limit. Since the live header limit is measured at 9,007,199,254,740,991 (see above), the block-space side of the ceiling is not what bounds it. The ceiling still presumes the cheapest possible transaction rather than a realistic workload, and that remains the honest caveat on the figure. The per-transaction EIP-7825 cap referenced here is **NOT MEASURED**: an `eth_call` above the stated 2^24 cap succeeded, `eth_estimateGas` does not discriminate, and the only conclusive probe is a transaction above the cap, which we do not send. `[VERIFY: publish the explicit block-space arithmetic behind 273,000 so the ceiling is reproducible]` ### A.1 Testnet-measured and live-observed throughput (the honest gap) **There is no testnet measurement that validates 273,000 TPS for L1.** The closest measured throughput number in the repo is a rollup-executor microbenchmark, and the artifact that produces it says plainly it neither reaches nor implies the ceiling. `[CITED: aere-research/parallel-executor/BLOCKSTM-ENDTOEND-THROUGHPUT-2026-07-13.md section 7]` The Block-STM end-to-end harness measured about 247,000 to 270,000 TPS at `gas=80` on a 16-vCPU box, but this is (a) the rollup executor with a light commit proxy, not L1, not a real Merkle Patricia Trie, and (b) entirely a function of the assumed per-transaction cost: the same box reports 115,000 to 500,000 TPS as the per-transaction weight changes. The `gas=80` proximity to 273,000 is called out in the artifact as a coincidence, not a validation. **Live mainnet 2800 runs far below the ceiling, and this is by demand, not by capacity.** `[CITED: aerenew/research/specs/spec-zk-stack.md section 0.2; aerenew/docs/wp2-sections/01-overview.md:59]` On-chain usage is thin (single-digit proof records for most verifiers; six genesis user wallets per project memory). Real throughput is demand-limited, not capacity-limited: the chain has ample block-space headroom and almost no load to fill it. `[MEASURE: sample a live window with eth_getBlockByNumber over N recent blocks against https://rpc.aere.network and divide total tx by elapsed time to publish an observed sustained TPS; expected result is near zero on an idle chain]` ### A.2 Block time and finality Deterministic single-slot QBFT finality at a 0.5 s block period. Because QBFT finalizes on commit rather than by accumulation of work, a transaction is irreversible in well under a second from inclusion, with no confirmation count and no reorg. `[CITED: aerenew/docs/wp2-sections/02-architecture.md:19]` The 0.5 s period was reached by a mid-chain QBFT parameter change at block 2,137,652, halved from the launch value of one second, with no re-genesis. `[CITED: aerenew/docs/wp2-sections/01-overview.md:59; aerenew/docs/WHITEPAPER-V2.md:73; genesis-subsecond.json `config.qbft.blockperiodseconds` = 1 at genesis, `config.transitions.qbft[0]` = `{ "block": 2137652, "xblockperiodmilliseconds": 500 }`]` The earlier form of this citation read `blockperiodseconds: 0, blockperiodmilliseconds: 500`. Neither appears in the file. `blockperiodseconds: 0` is rejected outright by Besu's `JsonUtil.getPositiveInt`, and the non-x key is silently ignored. Corrected 2026-07-31 by reading `genesis-subsecond.json` rather than by trusting the citation. Finality is a latency-to-irreversibility claim and is separate from throughput. The 0.5 s finality claim follows mechanically from QBFT plus the block period; it says nothing about TPS. ### A.3 Parallel execution (Block-STM): the honest result The honest result has three parts, and only the first is a mainnet-relevant guarantee. No end-to-end L1 throughput win is claimed, and Block-STM is NOT enabled on L1. **(1) Correctness proven bit-identical to sequential.** `[MEASURED-FRESH]` This session ran the executor's `bench` on this Windows dev machine. The `parallel==sequential` gate (committed state and keccak state root identical to the sequential oracle) passed on all four conflict profiles: ``` LOW-CONFLICT parallel==sequential check: PASS MEDIUM-CONFLICT parallel==sequential check: PASS HIGH-CONFLICT parallel==sequential check: PASS PATHOLOGICAL parallel==sequential check: PASS ``` This corroborates the cited larger runs: a 6,000-comparison correctness harness with 0 mismatches, and a re-import plus validation of 253 real exported blocks with zero state-root mismatch. `[CITED: aere-research/parallel-executor/BLOCKSTM-ENDTOEND-THROUGHPUT-2026-07-13.md section 3 and section 8]` **(2) Near-linear execution-phase scaling on parallelizable workloads.** `[MEASURED-FRESH]` on this machine (fewer physical cores than the reference 16-core Linux box, so absolute speedup is lower, which is expected and honest): | Workload (this box, 16 threads) | speedup vs sequential | abort% | |---|--:|--:| | Low-conflict (~2%) | 4.36x | 0.3% | | Medium-conflict (~15%) | 4.96x | 2.3% | | High-conflict (single hot counter) | 1.92x | 24.8% | | Pathological (all same slot) | 0.43x (correctly slower) | 11.5% | `[CITED: same artifact, section 4-5]` The reference 16-core Linux box measured about 8.3x to 9.4x execution-phase and about 8.3x to 9.1x end-to-end on realistic 0 to 50 percent conflict. The pathological all-same-slot profile is honestly slower than sequential (correct Block-STM behavior), on both boxes. **(3) No end-to-end throughput win claimed, and not on L1.** `[CITED: aerenew/research/specs/spec-parallel-execution.md section 9; BLOCKSTM-ENDTOEND section 7-8]` The Aere L1 mainnet executes sequentially on Besu. The parallel executor is a bounded four-kind VM (Transfer, Sweep, Increment, AmmSwap) over a balance and storage map, not a full EVM, and its measured speedups are single-machine microbenchmarks, not a network throughput or TPS claim. The real forked client's commit path was separately measured to anti-scale (parallel commit was slower than sequential), so the real-client end-to-end ceiling is tighter than the Rust harness, not looser. The invariant deliverable is the correctness gate plus the speedup curve and its Amdahl bound, not an absolute TPS. --- ## PART B: ZKVM COMPARATIVE (SP1 vs RISC Zero vs Halo2) Every cell is labeled. Gas and proof-size values that are cited come from repo deployment artifacts and are measured single-run values, not benchmarked averages (the source spec states this explicitly). `[CITED: aerenew/research/specs/spec-zk-stack.md footer]` ### B.1 Comparison table | Prover (version) | Proving time (off-chain) | On-chain gas to verify | Proof size | Verification latency | |---|---|---|---|---| | SP1 Groth16 (v6.1.0) | app-dependent, CITED single-runs: zkscreen ~72 to 75 s, over18 ~71 to 73 s, zkml-mnist ~200 to 205 s, aggregation Groth16 526.1 s `[CITED: contracts/deployments/proof-aggregator-scale.json]` | ~300k typical `[CITED: spec-zk-stack section 2]`; measured records: storage-proof 348,345, zkML 326,492, aggregation ~390k flat `[CITED: proof-aggregator-scale.json, zkml/storage-proof artifacts]` | ~260 to 356 bytes `[CITED: spec-zk-stack section 2]` | one staticcall, final at next block (~0.5 s) `[MEASURE: time eth_call wall-clock]` | | SP1 Plonk (v6.1.0) | `[MEASURE: prove a fixture with the Plonk prover]` | "somewhat higher than Groth16" `[CITED: spec-zk-stack section 2]`, exact `[MEASURE]` | `[MEASURE]` | one staticcall `[MEASURE]` | | RISC Zero Groth16 (5.0.0-rc.1) | `[MEASURE: r0vm prove the factor-guest]`, not recorded in repo | `[MEASURE: eth_call gas on RiscZeroVerifierRouter.verify]`, no gas number committed | seal 260 bytes `[CITED: contracts/deployments/risc0-verifier-fix.json sealLen]` | one staticcall `[MEASURE]` | | Halo2 (bn254 / KZG, SHPLONK) | `[MEASURE: halo2 prover on the cubic circuit]` | ~449,517 gas `[CITED: contracts/deployments/halo2-cubic.json; spec-zk-stack section 6]` (revm gas ~300,144) | 1,152 bytes `[CITED: same]` | one call `[MEASURE]` | | KZG / EIP-4844 point-eval precompile (0x0A) | n/a (precompile, no prover) | ~255,719 gas `[CITED: contracts/deployments/kzg-verifier.json; spec-zk-stack section 5]` | 192-byte precompile input `[CITED: same]` | one staticcall | Notes on the two recursion facts worth keeping, both `[CITED: proof-aggregator-scale.json; spec-zk-stack section 7]`: - A 10-proof recursive SP1 aggregation records on-chain in ~393,844 gas, roughly flat versus the 3-proof fold (~387,858 gas). Constant on-chain verification cost regardless of fold count is the point of recursion. - The proving side of that same 10-fold cost 526.1 s for the aggregation Groth16 plus 988 s of inner proving (measured, off-chain, single run). ### B.1.1 Every receipt-type figure in B.1 was fetched from chain 2800 and matches `[MEASURED-FRESH 2026-07-19]` The gas figures in B.1 that claim to be on-chain records are receipt-type claims, so each cited transaction was fetched from `https://rpc.aere.network` by `eth_getTransactionReceipt` and its `gasUsed` compared against the figure printed in this report. All five resolve, all five are status `0x1`, and all five match **exactly**, no roundings, no near-misses: | Figure in B.1 | Tx hash | Block | Receipt gasUsed | Report says | Result | |---|---|--:|--:|--:|---| | Halo2 cubic verify | `0x2be2449c...451c6c` | 8,926,951 | 449,517 | 449,517 | exact | | KZG point-eval (0x0A) | `0x28d4395c...8987b8` | 8,926,917 | 255,719 | 255,719 | exact | | SP1 aggregation n=10 | `0xd1fd4d60...44d978` | 8,930,005 | 393,844 | 393,844 | exact | | Storage-proof coprocessor | `0xe541c52b...936c54` | 8,921,209 | 348,345 | 348,345 | exact | | zkML MNIST verify | `0x4cb256ed...0b497a` | 8,741,874 | 326,492 | 326,492 | exact | This is the standard the Part C verify-and-record figures do not currently meet: these five name a transaction, the transaction exists on mainnet 2800, and its receipt says what the report says it says. A reader can re-run this check in five RPC calls. ### B.2 Routing logic (which prover for which job, and why) `[CITED: aerenew/research/specs/spec-zk-stack.md sections 1-3, 7-11]` Routing is by selector: the first 4 bytes of the proof pick the concrete verifier. - `SP1VerifierGateway` routes SP1 proofs by leading selector (Groth16 route `0x4388a21c`, Plonk `0x5a093a2f`). - `RiscZeroVerifierRouter` routes a RISC Zero seal by leading selector (`0xef6cb709` for 5.0.0-rc.1). Which prover is chosen: - **SP1 is the workhorse.** Every Aere application circuit routes through SP1: zk-KYC screen, over-18, zkML MNIST, compliance pool, storage-proof coprocessor, rollup validity anchor, and the recursive aggregator. The reason is concrete: smallest proofs (~260 to 356 bytes), lowest typical verify gas (~300k), one Rust guest toolchain, and native recursion via `verify_sp1_proof` inside the zkVM. - **RISC Zero is the diversity path.** A second independent RISC-V zkVM, kept so the verification surface is not single-prover; a real Groth16 factorization receipt is recorded on-chain. - **Halo2 / KZG is the direct-circuit path.** For circuits authored directly as a PLONK arithmetization (bn254 / KZG / SHPLONK), plus the raw EIP-4844 point-evaluation precompile for blob and commitment openings. The honest framing the source insists on: this is a multi-prover verification surface behind one gateway, unusual to have in one place, but not a claim that no other chain can verify any one of these. ### B.3 Why the on-chain-gas column motivates the STARK-verifier roadmap SP1 Groth16, RISC Zero Groth16, and Halo2 all terminate their on-chain check in a **BN254 pairing** (Groth16 or KZG). Two problems live in that column: the check costs ~300k to ~450k gas, and BN254 pairing security rests on discrete-log / pairing hardness, which is **Shor-breakable**. That single elliptic-curve link is the only quantum-vulnerable step in an otherwise hash-based SP1 proof chain. Roadmap item: replace the BN254 Groth16 wrap with a direct hash-based STARK verifier. `[CITED: aerenew/pqc-fork/pq-stark/README.md; aerenew/docs/PQ-STARK-VERIFIER-PRECOMPILE-2026-07-18.md]` The design is a native precompile at `0x0AE8`. The generic verifier skeleton built so far verifies a BabyBear / Plonky3 FRI STARK directly, with no BN254 wrap, so its security rests only on hash collision-resistance and Reed-Solomon proximity gaps. **Scope caveat (2026-07-19 research finding)** `[CITED: aerenew/docs/AERE-STARK-SP1-RECURSION-AIR-PORT-SPEC-SUMMARY.md]`: the six generic components confirmed so far verify **BabyBear + FRI STARKs** (Aere's OWN Plonky3 circuits), NOT SP1 6.1.0 proofs. The pinned SP1 6.1.0 is a **Hypercube** release (KoalaBear multilinear: BaseFold plus sumcheck-zerocheck plus LogUp-GKR), so its inner proof is hash-based but NOT a BabyBear FRI STARK, and FRI plus DEEP-ALI do not apply to it. Replacing the SP1 BN254 Groth16 wrap on real SP1 proofs is therefore a SEPARATE ~22 to 32 person-week retarget and a founder decision, not a completion of the skeleton below. **Honest status of that roadmap item, stated plainly** `[CITED: same README, "HONEST STATUS"]`: it is a REFERENCE SKELETON plus DESIGN. It is NOT audited, NOT a working verifier for SP1 proofs, and NOT activated on any Aere network. It fail-closes (returns "not verified" for every input, cannot emit a false accept). The SP1-recursion-specific crypto core is un-ported, and per the caveat above the SP1 target is a different (Hypercube) proof system entirely. Mainnet 2800 has exactly the five PQC precompiles `0x0AE1..0x0AE5`; `0x0AE8` is not live anywhere. --- ## PART C: CRYPTO-OP GAS (PQC vs ECDSA) ### C.1 Three different kinds of number, kept in three different columns The single most dangerous thing this table could do is print a schedule constant, an empirical cost, and a transaction receipt in one column as if they were the same kind of fact. They are not, so they are separated here: - **PRICE**, a fixed constant in the client's gas schedule. It is not an empirical result; it is a chosen number. What *can* be measured about a price is whether the live chain actually charges it. - **MEASURED-LIVE**, the price the live canonical client charged, isolated by a calibrated `eth_call` probe on mainnet 2800 (method in C.1.1). This is the only column that constitutes evidence about the running chain. - **RECEIPT**, a `gasUsed` from a real mined transaction, only meaningful when the tx hash and the chain it was mined on are both given. | Operation | Address | Schedule PRICE | MEASURED-LIVE on 2800 | Verdict | |---|---|--:|--:|---| | ECDSA (ecrecover) | 0x01 | 3,000 | **3,000** | price confirmed charged | | P-256 / passkey (**RIP-7951 / EIP-7951**) | 0x100 | **6,900** | **6,900** | price confirmed charged; see C.1.2 | | SHAKE256 (FIPS 202) | 0x0AE5 | 60 + 12/word, words counted **separately** (C.1.3) | **60 + 12/word, separate** (56/56 shapes exact) | price and word model confirmed | | Falcon-512 (NIST round-3) | 0x0AE1 | 40,000 | **40,000** | price confirmed charged | | Falcon-1024 (NIST round-3) | 0x0AE2 | 75,000 | **75,000** | price confirmed charged | | ML-DSA-44 (FIPS 204) | 0x0AE3 | 55,000 | **55,000** | price confirmed charged | | SLH-DSA-SHA2-128s (FIPS 205) | 0x0AE4 | 350,000 | **350,000** | price confirmed charged | | ML-KEM-768 (FIPS 203, encap) | 0x0AE6 | 60,000 `[CITED: MLKEM768PrecompiledContract.java line 73]` | **not live** (probe returns 0) | testnet-only, as documented | | Falcon HashToPoint | 0x0AE7 | 60 + 12/word `[CITED: HashToPointPrecompiledContract.java lines 51-52]` | **not live** (probe returns 0) | testnet-only, as documented | `[MEASURED-FRESH 2026-07-19, mainnet 2800, head block 10,487,300]` Every MEASURED-LIVE cell above was produced this session by the calibrated probe, not carried over from the scratch fork. The five signature and hash precompiles (`0x0AE1..0x0AE5`) are confirmed live and charging; `0x0AE6`, `0x0AE7` and `0x0AE8` are confirmed **absent** on mainnet (the probe measures exactly 0 marginal gas, i.e. a call to a codeless address), which independently corroborates the "KAT-validated on isolated testnet, founder-gated for mainnet" status. Activation of `0x0AE1..0x0AE5` is stated as block 9,189,161. `[VERIFY: this activation block could NOT be confirmed by measurement. The check that would prove it is a historical `eth_call` probe at blocks 9,189,160 and 9,189,161, showing the precompile charging 0 before and 40,000 after. The public endpoint retains state for only about the last 256 to 1,024 blocks (probes succeed at head-256 and return "Internal error" at head-1024), so state at block 9.19M is long pruned. Confirming this requires an archive node. The Nethermind second client cannot serve it either: it is mid-sync well below 9.19M. The activation block therefore remains a cited claim, not a measured one.]` **Verify-and-record transaction gasUsed, moved out of the table above, because these are NOT mainnet receipts.** | Operation | verify-and-record gasUsed | Chain it was mined on | Tx hash | |---|--:|---|---| | SHAKE256 | 21,470 | scratch fork **28099** | not recorded | | Falcon-512 | 86,336 | scratch fork **28099** | not recorded | | Falcon-1024 | 145,496 | scratch fork **28099** | not recorded | | ML-DSA-44 | 351,050 | scratch fork **28099** | not recorded | | SLH-DSA-SHA2-128s | 558,276 | scratch fork **28099** | not recorded | `[CITED: aere-research-repo/docs/PROVEN-RESULTS-2026-07-11.md line 93, which states plainly "Measured on the isolated scratch fork (chain 28099), full verify-and-record transactions"; aerenew/research/aip-draft-pqc-precompiles.md sections 3-7]` `[VERIFY: no transaction hash is recorded for any of these five, so none of them can be fetched and re-verified by a reader. They are single-run receipts from a disposable chain. Either publish the tx hashes and the chain-28099 export, or re-run all five as real transactions on mainnet 2800 (the precompiles are live, so this is now possible and would replace fork receipts with citable mainnet ones). Until then these five numbers are the weakest-provenance figures in this report and MUST NOT be cited as mainnet results.]` An independent data point showing how environment-dependent these end-to-end receipts are: the keyless statefork rehearsal on a *different* isolated network measured Falcon-1024 verify-and-record at **172,720** gas and ML-DSA-44 at **351,970** gas `[CITED: PROVEN-RESULTS-2026-07-11.md line 95]`, versus 145,496 and 351,050 on chain 28099. The marginal precompile price is a hard constant; the verify-and-record total is not, because it includes contract storage writes and calldata that vary with the harness. Scope reminder attached to this table: these precompiles are an application and account layer capability. They do not make Aere consensus post-quantum. #### C.1.1 How the MEASURED-LIVE column was produced, and how the instrument was calibrated `eth_estimateGas` is insufficient for this job: the EIP-7623 calldata floor can mask precompile price differences entirely. Instead a probe contract was installed at a scratch address via `eth_call` state override (read-only, nothing was sent to the chain). The probe does `CALLDATACOPY`, a **pre-warming** `STATICCALL` to the target, then `GAS / STATICCALL / GAS / SWAP1 / SUB` around the measured call, and returns the delta. The pre-warming call absorbs both the EIP-2929 cold-access charge and the memory expansion, so the second call's delta is the precompile's own `gasRequirement` plus a fixed probe overhead. That overhead is removed by subtracting the identical measurement against a codeless address. The instrument was calibrated **before** being trusted, against two operations whose price no scheduled EIP can move: - IDENTITY `0x04`, expected `15 + 3*ceil(len/32)`: **exact on all 12 input lengths tested** (0, 1, 31, 32, 33, 64, 96, 100, 128, 200, 256, 1000 bytes). - SHA-256 `0x02`, expected `60 + 12*ceil(len/32)`: **exact on all 4 lengths tested**. 16/16 exact, zero residual, including the `len=0` case returning exactly the bare base cost, which demonstrates the control subtraction leaves no residue. Only after that was any PQC or P-256 number recorded. #### C.1.2 The P-256 figure: a corrected 2x error, and exactly where it came from **A previous revision of this report stated P-256 at `~3,450` gas and attributed that to "RIP-7951, aka RIP-7212". Both halves of that were wrong.** The measured price on mainnet 2800 is **6,900 gas**, flat and independent of input length (measured at input lengths 0, 160 and 161 bytes; all 6,900). The resolution is not "spec price versus measured cost including overhead". The probe subtracts its own overhead, the raw probe delta was 7,029 and the codeless control was 129, leaving 6,900 as the precompile's own charge with zero call overhead included. So 6,900 is *both* the spec price and the measured price, and 3,450 is neither. 3,450 is the price in **RIP-7212**, the original Layer-2 rollup spec. **RIP-7951 / EIP-7951**, the mainnet successor scheduled in Fusaka, keeps RIP-7212's address (`0x100`), its 160-byte input and its output format, full bytecode compatibility , but **doubles the gas to 6,900**, having addressed security issues found in RIP-7212. Aere tracks Ethereum's ruleset (Pectra plus Fusaka), so Aere ships the RIP-7951 price of 6,900, correctly. The "aka RIP-7212" phrasing was the source of the error: the two specs are interface-compatible but are *not* price-compatible, and the older number was carried in under the newer name. `[VERIFY: this correction propagates. aerenew/research/aip-draft-pqc-precompiles.md section 4 anchors its Falcon-512 pricing rationale on "RIP-7951 prices a native P-256 verification at a fixed roughly 3,450 gas". That anchor is wrong. The argument's direction survives but its magnitude does not: Falcon-512 is 40,000/6,900 = 5.8x a P-256 verify, not 40,000/3,450 = 11.6x.]` #### C.1.3 SHAKE256 word counting: independently re-measured, separate not combined The SHAKE256 gas word count is `ceil(len(data)/32) + ceil(outLen/32)`, the two lengths rounded up **separately**, not `ceil((len(data)+outLen)/32)` combined. This was re-measured this session against live mainnet 2800 with well-formed input (32-byte `outLen` word followed by data) across **56 shapes** (8 output lengths x 7 data lengths), fitting both candidate models with no hypothesis assumed: ``` SEPARATE model 60 + 12*(ceil(d/32) + ceil(o/32)) : 56/56 exact, SSE = 0 COMBINED model 60 + 12*ceil((d+o)/32) : 47/56 exact, SSE = 1296 ``` The models diverge exactly when `len(data) mod 32` and `outLen mod 32` are both nonzero. Behaviour was confirmed alongside price: `eth_call` to `0x0AE5` with `outLen=32` and data `abc` returns `0x483366601360a8771c6863080cc4114d8db44530f8f1e1ee4f94ea37e78b5739`, which is `shake_256(b'abc').hexdigest(32)` per FIPS 202. `[VERIFY: aerenew/research/aip-draft-pqc-precompiles.md section 3 still specifies the combined form `words = ceil((len(data) + outLen) / 32)` normatively. That is the same defect already corrected in aerenew/eips/eip-draft-shake256-precompile.md and it contradicts the running chain. It is corrected in the AIP draft as of this report; flagged here because this report cited that document as its source.]` `[VERIFY: a previous revision of this report cited aerenew/pqc-fork/precompiles/HashToPointPrecompiledContract.java lines 51-52 as the source for the SHAKE256 (0x0AE5) gas constants. That file is the HashToPoint precompile, address 0x0AE7, which this session's probe confirms is NOT live on mainnet. Its BASE_GAS=60 / GAS_PER_WORD=12 happen to equal SHAKE256's constants but it is not their source. The 0x0AE5 constants are now sourced from live measurement above.]` Scope reminder attached to this table: these precompiles are an application and account layer capability. They do not make Aere consensus post-quantum. ### C.2 Two further cross-checks Note the register difference between these two and the C.1 measurements: the first is **arithmetic evaluated from source code**, not a measurement of a running chain, and the second is a **scratch-fork benchmark artifact**. Neither is a mainnet result, and neither is presented as one. `[COMPUTED-FROM-SOURCE, not measured]` The SHAKE256 gas schedule (60 base plus 12 per word) is confirmed present in source, and the HashToPoint precompile's own gas was computed from that exact source formula (`gasRequirement`, lines 78-93). This is evaluating a formula in a file, and HashToPoint (`0x0AE7`) is not live on mainnet 2800, so this figure could not have been measured on-chain even in principle: ``` Falcon-512 (n=512) 32-byte message, 73-byte input: words=38 -> gasRequirement = 516 Falcon-1024 (n=1024) 32-byte message, 73-byte input: words=73 -> gasRequirement = 936 ``` 516 gas matches the source javadoc's "one ~500-gas staticcall" claim, which is the mechanism by which HashToPoint cuts a full Solidity Falcon-512 verify. This is the precompile's own intrinsic gas; it is distinct from the end-to-end verify gas below, and the two must not be conflated. `[CITED: aerenew/pqc-fork/results/bench-results.json]` The HashToPoint-precompile optimization measured on the scratch fork (chain 28777, NOT mainnet), full Falcon-512 verify: 9,140,858 gas in pure EVM dropping to 7,186,367 gas with the precompile, a 21.4 percent reduction across 8 test vectors, all accept and tamper-reject checks passing. Re-checked against the artifact this session: the JSON records `avg_verify_gas_in_evm` 9140858, `avg_verify_gas_precompile` 7186367, `avg_verify_pct_drop` 21.4, `chainId` 28777, `vectors` 8, `integration_all_pass` true. The report's figures match the artifact exactly. ### C.3 The honest story the numbers tell Post-quantum verification is materially more expensive than ECDSA. Ratios below are fresh arithmetic over the **live-measured** prices of C.1, divided by the live-measured 3,000-gas ecrecover and by the live-measured 6,900-gas P-256. Both baselines are given because they answer different questions: ecrecover is what a plain EOA signature costs, P-256 is what a *passkey* signature costs, and a passkey is the realistic thing a post-quantum account would be replacing. | Scheme | Live-measured gas | vs ECDSA (3,000) | vs P-256 passkey (6,900) | |---|--:|--:|--:| | Falcon-512 | 40,000 | ~13.3x | ~5.8x | | ML-DSA-44 | 55,000 | ~18.3x | ~8.0x | | Falcon-1024 | 75,000 | ~25.0x | ~10.9x | | SLH-DSA-SHA2-128s | 350,000 | ~116.7x | ~50.7x | | ML-KEM-768 (not live on 2800) | 60,000 (price only) | ~20.0x | ~8.7x | So PQC verification runs roughly 13x (Falcon-512) to about 117x (SLH-DSA-SHA2-128s) the cost of an ECDSA recover, or about 6x to 51x the cost of a P-256 passkey verify. The ML-KEM-768 row is a schedule price only: that precompile is not live on mainnet 2800, so unlike the rows above it it has no live-measured value. That premium is exactly what motivates two roadmap directions: - **Aggregation.** Amortize one expensive PQC verify across many authentications rather than paying per signature. `[VERIFY: the specific "AerePQAggregate" contract name is not found in the repo; PQC aggregation is a documented roadmap direction, see aerenew/docs/FRONTIER-PROGRAM-2026-07-12.md]` - **Finality certificate.** Carry post-quantum finality as a certificate over a finalized block rather than re-verifying PQC signatures per transaction. This ties to the dual-quorum PQ-consensus design work. `[CITED: aerenew/docs/PQ-CONSENSUS-STEP2-2026-07-18.md]` Note the scope line: this is design and attestation alongside QBFT, it does not make consensus post-quantum. --- ## PART D: INDEPENDENT REPRODUCTION ### D.1 The strongest benchmarks should be third-party reproduced The credibility of every number above rises sharply when a party other than Aere reruns it. The honest reproduction candidates, in priority order: - **The Nethermind team (warmest angle).** Aere already runs a patched Nethermind as a second client that validates and follows live chain 2800, and it runs the PQC precompiles and can produce QBFT blocks Besu accepts. `[CITED: aerenew/docs/wp2-sections/09-governance-roadmap.md:199, 253; NETHERMIND-2ND-CLIENT-LIVE memory]` That existing relationship makes them the natural first ask for an independent client-side reproduction of the precompile gas and the consensus behavior. - **A university lab.** Cheap under a research grant and citable in a way a vendor benchmark is not; a good fit for the Block-STM correctness-and-scaling harness and the zkVM proving-time measurements, which are self-contained and need no Aere infrastructure. - **Explicitly NOT the expensive audit firms.** The audit-firm budget is reserved for the security audit, which is a different deliverable from a performance reproduction. `[CITED: aerenew/docs/mainnet_authority context; AUDIT-OUTREACH-2026-07-12.md]` The outreach kit is prepared by us. The actual third-party run and the relationship are founder and business gated, not something this report initiates. ### D.2 The run-it-yourself manifest (REPRODUCE) This report is the NUMBERS and how they were measured. The run-it-yourself commands live in the reproducibility bundle. Today that bundle is split across several committed files rather than one consolidated `REPRODUCE.md`; the honest recommendation is to consolidate them, but each is real and runnable now: | Domain | Where to run it | File | |---|---|---| | Block-STM correctness + scaling (Part A.3, freshly run here) | `cd aerenew/parallel-executor && cargo test --release && cargo run --release -- bench` and `-- endtoend` | `aere-research/parallel-executor/BLOCKSTM-ENDTOEND-THROUGHPUT-2026-07-13.md` section 10 | | PQC precompile KATs + gas (Part C) | `aerenew/pqc-fork/run-kats.sh`, `aere-research/bench/bench.py` | `aerenew/pqc-fork/` (results in `aere-research/results/bench-results.json`, `results/kat-results-*.json`) | | Reproducible Besu client build | `docker build ... -f aerenew/node/Dockerfile.reproducible` then `aerenew/scripts/verify-besu-image.sh` | `aerenew/node/REPRODUCIBLE.md` (note: public digest manifest publication is still roadmap, so the third-party comparison step is not yet live) | | PQ-consensus reproduction | committed harness | `aerenew/research-stage/formal-consensus/REPRODUCE.md`, `aerenew/audit-package-pq-consensus/REPRODUCE.md` | `[MEASURE: create a single top-level aerenew/REPRODUCE.md that indexes all four domains above with pinned commands, so a third party has one entry point]` --- ## Appendix: provenance summary of every number **Freshly measured this session `[MEASURED-FRESH]`:** - Block-STM `parallel==sequential` gate: PASS on all 4 conflict profiles (this Windows box). - Block-STM 16-thread execution-phase speedup this box: low 4.36x, medium 4.96x, high 1.92x, pathological 0.43x. **Freshly measured against live mainnet 2800 (2026-07-19, head block 10,487,300), calibrated probe, read-only `eth_call` only `[MEASURED-LIVE]`:** - Probe calibration: IDENTITY `0x04` exact on 12 lengths, SHA-256 `0x02` exact on 4 lengths, 16/16, zero residual. - ecrecover `0x01` = **3,000** (flat). Closes the prior `[VERIFY]` on this constant. - P-256 `0x100` = **6,900** (flat across 0/160/161-byte input). **Corrects the prior "~3,450".** - Falcon-512 `0x0AE1` = **40,000**; Falcon-1024 `0x0AE2` = **75,000**; ML-DSA-44 `0x0AE3` = **55,000**; SLH-DSA `0x0AE4` = **350,000**. All flat, all matching the published schedule prices. - SHAKE256 `0x0AE5` = **60 + 12 per word, words counted separately**: 56/56 shapes exact, SSE 0; combined-word model 47/56, SSE 1296. - SHAKE256 behaviour: `shake256("abc", 32)` returns the correct FIPS 202 bytes. - `0x0AE6`, `0x0AE7`, `0x0AE8` measure **0 marginal gas** = not live on mainnet, corroborating their documented testnet-only status. - **Re-measured and re-derived 2026-07-20** by `aerenew/pqc-fork/probe-precompile-liveness.js` at block 10,611,947, with a second, independent probe design (an in-EVM `STATICCALL` gas prober injected by `eth_call` state override, rather than transaction-level `eth_estimateGas`). Identity calibration exact at 6 lengths. Charges: `0x0AE1` 40,000, `0x0AE2` 75,000, `0x0AE3` 55,000, `0x0AE4` 350,000, `0x0AE5` 24,648, `0x0AE6` **0**, `0x0AE7` **0**. Both probe designs agree. - **Caveat on the `eth_estimateGas` form of this probe, discovered 2026-07-20.** Under EIP-7623 a transaction pays `max(21000 + standard calldata cost + execution, calldata floor cost)`, and the floor term dominates. At a 64-byte input the measured headroom is **1,873 gas of execution that is free inside the floor and therefore invisible**. Measured consequence: the estimateGas probe reports the **live** identity precompile at `0x04` as charging **0 gas**, at all six lengths tried. So "estimateGas equals the codeless control" is sound evidence of absence only for a precompile that would bill *more* than the floor headroom on that input. It is sound for `0x0AE6` (flat 60,000 for every input). It is **not** sound on its own for `0x0AE7`, which bills only about 84 gas on a malformed 64-byte input. `0x0AE7`'s absence rests on the in-EVM probe, which has no such blind spot. Raw numbers at that input for the record: codeless control 23,897; `0x0AE1` 62,853; `0x0AE2` 97,992; `0x0AE3` 77,913; `0x0AE4` 376,144; `0x0AE5` 47,055; `0x0AE6` 23,897; `0x0AE7` 23,897. - Part B receipts re-fetched: Halo2 449,517 / KZG 255,719 / SP1-agg 393,844 / storage-proof 348,345 / zkML 326,492, all five resolve on 2800, all status 1, all exact. **Computed from source this session, NOT measured on a chain `[COMPUTED-FROM-SOURCE]`:** - HashToPoint precompile intrinsic gas from source formula: 516 gas (Falcon-512, 32-byte message), 936 gas (Falcon-1024). HashToPoint is not live on mainnet, so this is formula evaluation, not measurement. **Cited from committed repo artifacts `[CITED]`:** - 273,000 TPS is a design ceiling, never measured (WHITEPAPER-V2, wp2-sections). - 0.5 s block period, halved at block 2,137,652 by `config.transitions.qbft[0].xblockperiodmilliseconds = 500`, single-slot QBFT finality (wp2-sections, genesis-subsecond.json). - Rollup end-to-end harness ~247k to 270k TPS at gas=80, swinging 115k to 500k, not the L1 ceiling (BLOCKSTM-ENDTOEND section 7). - Reference 16-core Block-STM ~8.3x to 9.4x exec, ~8.3x to 9.1x end-to-end (same artifact). - SP1 Groth16 ~300k gas, ~260 to 356 byte proofs; measured records storage-proof 348,345, zkML 326,492, aggregation ~390k (spec-zk-stack; deployment artifacts). - SP1 proving times: zkscreen ~72 to 75 s, over18 ~71 to 73 s, zkml-mnist ~200 to 205 s, aggregation Groth16 526.1 s (proof-aggregator-scale.json). - RISC Zero Groth16 seal 260 bytes, 5.0.0-rc.1 (risc0-verifier-fix.json). - Halo2 ~449,517 gas, 1,152-byte proof; KZG point-eval ~255,719 gas (deployment artifacts, spec-zk-stack). - PQC verify-and-record tx gasUsed: 86,336 / 145,496 / 351,050 / 558,276 / 21,470, **scratch fork chain 28099, no tx hashes recorded, NOT mainnet** (see C.1). - HashToPoint optimization 9,140,858 to 7,186,367 gas, -21.4%, **scratch fork chain 28777** (bench-results.json). **Superseded by live measurement this session (previously `[CITED]`, now `[MEASURED-LIVE]`):** the PQC marginal-gas constants and the ecrecover constant are no longer cited from the scratch fork or from memory; they were re-measured on mainnet 2800 and all matched. The P-256 figure did **not** match and was corrected. **Corrections applied to this report on 2026-07-19:** 1. **P-256 `0x100`: `~3,450` -> `6,900`.** The old value was RIP-7212's L2 price carried in under RIP-7951's name; RIP-7951/EIP-7951 doubles it to 6,900, and 6,900 is what chain 2800 measurably charges. A 2x understatement. See C.1.2. 2. **Verify-and-record gasUsed figures re-labelled.** They were presented as "on-chain receipts" in a report headed "Chain: mainnet 2800"; they are chain-28099 scratch-fork receipts with no recorded tx hashes. Moved to their own table with the chain named and the gap flagged. See C.1. 3. **SHAKE256 source citation corrected.** The report cited `HashToPointPrecompiledContract.java` (which is `0x0AE7`, not live on mainnet) as the source for `0x0AE5`'s constants. Now sourced from live measurement. 4. **SHAKE256 word model made explicit and verified separate, not combined** (56/56, SSE 0), and the surviving combined-form error in the cited AIP draft was flagged and fixed. See C.1.3. 5. **Register separation introduced in C.1**: schedule PRICE, MEASURED-LIVE, and RECEIPT are now three distinct columns rather than one "Marginal gas" column. 6. **C.3 ratio table** rebased on live-measured values and given a P-256 baseline alongside ecrecover. **To measure `[MEASURE]` or verify `[VERIFY]` (open items):** - `[MEASURE]` Observed live mainnet sustained TPS over a real block window. - `[MEASURE]` SP1 Plonk proving time, gas, proof size. - `[MEASURE]` RISC Zero on-chain verify gas and proving time. - `[MEASURE]` Halo2 proving time. - `[MEASURE]` Wall-clock verification latency per prover. - `[MEASURE]` Consolidated top-level aerenew/REPRODUCE.md. - `[MEASURE]` **Re-run the five PQC verify-and-record transactions on mainnet 2800 and publish their tx hashes**, replacing the chain-28099 fork receipts. The precompiles are live, so this is now possible and is the single highest-value provenance fix left in this report. - `[VERIFY]` The explicit block-space arithmetic behind 273,000 TPS. - ~~`[VERIFY]` Which genesis gas limit is live on 2800 (40M vs effectively unbounded).~~ **CLOSED, measured 2026-08-01:** header `gasLimit` at head block 11,810,041 is `0x1fffffffffffff` = 9,007,199,254,740,991, the effectively-unbounded value, identical to genesis. **Opened in its place:** the EIP-7825 per-transaction cap of 2^24 is **NOT MEASURED** and cannot be inferred from the block limit. - ~~`[VERIFY]` ECDSA ecrecover 3,000~~, **CLOSED**, measured 3,000 on live 2800. - `[VERIFY]` The "AerePQAggregate" contract name (not found in repo; aggregation is a documented roadmap direction).