# Block-STM end-to-end block-production throughput (measured) **Date:** 2026-07-13 **Module:** `parallel-executor/` (the from-scratch Rust Block-STM executor) **Scope:** ISOLATED benchmark on a throwaway cloud box. Nothing here touches live chain 2800, its validators, or any production node. This does **not** make Block-STM mainnet-active; that remains a founder-supervised fork activation. The site's honest framing ("testnet demonstration, not mainnet-active; execution-phase win") stays intact. This report only **adds a real, measured end-to-end number** to that framing. --- ## 0. Why this exists (the gap being closed) The existing `bench` command times **only** the Block-STM execution phase and reports the documented near-linear execution-phase scaling (~1.94x / 3.61x / 6.19x / 9.38x at 2/4/8/16 workers). That is an honest execution-phase number, but it is **not** an end-to-end block-throughput number. A block is not finished when execution finishes; it is finished when the state is **committed** and a **state root** is produced. Those tail phases are serial, so by Amdahl's law the end-to-end speedup must be **lower** than the execution-phase speedup. The site states exactly this gap: > "this is an execution-phase win; an end-to-end throughput win still needs real > execution-heavy load plus a parallel commit." This report closes that gap with a real measurement: the **full pipeline** (parallel execution + commit + state-root), parallel vs sequential, at 1/2/4/8/16 workers, across conflict ratios, with serial-equivalence checked on every run. New reproducible harness: `aere-block-stm endtoend [gas] [ntx] [reps]` (added to `src/main.rs`; phase timing exposed via the additive `execute_block_parallel_split` in `src/parallel.rs`). --- ## 1. Environment (honest) | | | |---|---| | Box | Throwaway cloud VM, 16 vCPU, 32 GB, Ubuntu 24.04, US region | | CPU | AMD EPYC-Rome, 16 cores, **Thread(s)/core = 1** (no in-guest SMT), 2.0 GHz | | Kernel / toolchain | Linux 6.8.0, rustc 1.97.0, release build (`lto=true`, `codegen-units=1`) | | Firewall | dedicated scratch firewall, SSH restricted to operator addresses | | Lifecycle | provisioned for this benchmark, **DELETED at end** (see §9) | **Caveat, stated up front:** this is a **shared-vCPU** cloud box, not the dedicated-core infra machine the original 9.38x figure was taken on. Shared vCPU adds run-to-run variance to **absolute** timings. Mitigations: 16 single-threaded vCPUs (so worker count maps to real cores with no hyperthread muddying), median of 7 runs, and the fact that the **headline deliverable is a ratio** (end-to-end vs execution-phase speedup, and the serial fraction) in which shared-vCPU noise affects numerator and denominator together and largely cancels. --- ## 2. Methodology **Pipeline, timed in three phases** (see `endtoend()` in `src/main.rs`): 1. **exec** - Block-STM parallel execution + validation + re-execution. PARALLEL; this is the part that scales with worker count. Timed inside `execute_block_parallel_split` between thread spawn and join. 2. **commit** - `MvMemory::snapshot_final`: materialize the multi-version memory into the flat committed `StateMap`. SERIAL (single-threaded, after join). 3. **root** - `keccak::state_root`: keccak256 over the sorted committed state, the on-chain rollup commitment. SERIAL. `end-to-end = exec + commit + root`. The **sequential baseline** is `execute_block_sequential` (which produces the live map directly, so it has no separate commit step) `+ state_root`. The state-root cost is identical serial work in both numerator and denominator, i.e. a pure Amdahl serial tail. **Workload (execution-heavy, not empty blocks).** The deterministic SplitMix64 generator emits a realistic EVM-style mix per batch: **60% ERC-20 transfers** (two balance touches), **15% sweeps** (value-dependent), **15% shared-counter increments** (the serializability stressor), **10% AMM swaps** (two-slot pool conflict). Each transaction also pays a deterministic simulated EVM cost of `gas` keccak-rounds (a stand-in for opcode/keccak/ecrecover interpreter time), so the executor is timed doing real per-tx work rather than pure scheduling overhead. **Conflict ratios.** `conflict_ppm` tunes how often a transaction touches the shared hot set. Four points span the spectrum: **0%** (disjoint accounts, 64 contracts), **10%**, **50%**, and a **pathological ~100%** single-hot-slot floor. The labelled % is the *designed* hot-touch rate; the honest *realized* conflict indicator is the **measured abort%** in each table. **Defaults:** `gas=80`, `ntx=20000`, median of 7 (chosen to match the regime the documented ~9.38x execution-phase figure was taken in, so the two are directly comparable). A gas sweep (§6) shows sensitivity. **Serial-equivalence, checked every run.** On every single parallel run the harness asserts the committed state equals the sequential oracle **and** the two keccak state roots are bit-identical. A single mismatch aborts the whole run. This is the property that makes the parallel root a valid commitment. --- ## 3. Serial-equivalence confirmation (the non-negotiable gate) | Check | Result | |---|---| | `cargo test --release` (7 tests incl. shared-counter serialization, 40-seed random equivalence, keccak KAT) | **7/7 PASS** | | Correctness harness: 6 profiles x 200 batches x {1,2,4,8,16} threads | **6000 comparisons, 0 mismatches** | | End-to-end harness inline check (state == oracle AND roots identical), every run | **140 checks, 0 mismatches** | `RESULT: PASS -- every parallel run committed the exact sequential state + root.` The `04-endtoend.stderr` log is empty (zero mismatch lines). Raw logs in `endtoend-logs-2026-07-13/`. --- ## 4. Execution-phase baseline on THIS box (for comparison) Running the existing `bench` (execution-phase only) on the same box reproduces the documented range, confirming the box is a faithful stand-in: | Workload | 2 | 4 | 8 | **16** (vs sequential) | |---|---|---|---|---| | Low-conflict (~2%) | 1.74x | 3.24x | 5.68x | **8.30x** | | Medium-conflict (~15%) | 1.82x | 3.48x | 6.18x | **9.62x** | | High-conflict (~80%) | 1.74x | 3.10x | 4.96x | **5.13x** | | Pathological (all one slot) | 0.93x | 0.86x | 0.82x | **0.76x** | (Medium-conflict vs-1-thread: 1.94x / 3.71x / 6.59x / 10.26x, matching the documented 1.94x/3.61x/6.19x/9.38x.) These are **execution-phase** numbers. --- ## 5. END-TO-END RESULTS (the deliverable) `gas=80`, `ntx=20000`, median of 7. `execSpd` = execution-phase speedup; `e2eSpd` = **end-to-end** speedup (exec+commit+root) vs the sequential end-to-end baseline. TPS is real transactions/second of the full pipeline. All rows PASS serial-equivalence. ### 0% designed conflict - 49,637 committed state keys (largest state) Sequential end-to-end: 672.84 ms (root-share 1.4%), 29,725 TPS. | W | exec ms | commit ms | root ms | e2e ms | execSpd | **e2eSpd** | **TPS** | abort% | (commit+root)/e2e | |--:|--:|--:|--:|--:|--:|--:|--:|--:|--:| | 1 | 705.51 | 9.01 | 12.52 | 727.04 | 0.94x | 0.93x | 27,509 | 0.0% | 3.0% | | 2 | 354.89 | 9.56 | 11.41 | 375.85 | 1.87x | 1.79x | 53,212 | 0.0% | 5.6% | | 4 | 184.02 | 9.72 | 12.37 | 206.11 | 3.60x | 3.26x | 97,037 | 0.1% | 10.7% | | 8 | 97.99 | 8.52 | 8.31 | 114.82 | 6.77x | 5.86x | 174,191 | 0.1% | 14.7% | | **16** | 64.23 | 8.08 | 8.51 | 80.83 | **10.32x** | **8.32x** | **247,444** | 0.3% | 20.5% | Amdahl serial tail (commit+root) = 15.33 ms; e2e ceiling 8.27x @16w, 43.88x @inf cores. ### 10% designed conflict - 18,342 keys Sequential end-to-end: 672.71 ms, 29,731 TPS. | W | exec ms | commit ms | root ms | e2e ms | execSpd | **e2eSpd** | **TPS** | abort% | tail% | |--:|--:|--:|--:|--:|--:|--:|--:|--:|--:| | 1 | 699.73 | 5.15 | 4.77 | 709.65 | 0.96x | 0.95x | 28,183 | 0.0% | 1.4% | | 2 | 366.73 | 5.31 | 4.38 | 376.42 | 1.83x | 1.79x | 53,133 | 0.1% | 2.6% | | 4 | 187.09 | 5.15 | 3.88 | 196.13 | 3.58x | 3.43x | 101,975 | 0.3% | 4.6% | | 8 | 96.64 | 5.06 | 4.29 | 105.99 | 6.93x | 6.35x | 188,699 | 0.7% | 8.8% | | **16** | 66.57 | 4.37 | 3.20 | 74.14 | **10.06x** | **9.07x** | **269,750** | 1.2% | 10.2% | Amdahl serial tail = 6.68 ms; e2e ceiling 9.64x @16w, 100.71x @inf. ### 50% designed conflict - 9,006 keys Sequential end-to-end: 699.97 ms, 28,573 TPS. | W | exec ms | commit ms | root ms | e2e ms | execSpd | **e2eSpd** | **TPS** | abort% | tail% | |--:|--:|--:|--:|--:|--:|--:|--:|--:|--:| | 1 | 712.46 | 2.63 | 1.96 | 717.06 | 0.98x | 0.98x | 27,892 | 0.0% | 0.6% | | 2 | 367.61 | 2.46 | 2.01 | 372.07 | 1.90x | 1.88x | 53,753 | 2.1% | 1.2% | | 4 | 195.00 | 2.58 | 1.98 | 199.56 | 3.58x | 3.51x | 100,222 | 5.7% | 2.3% | | 8 | 113.28 | 2.50 | 1.72 | 117.51 | 6.17x | 5.96x | 170,202 | 10.7% | 3.6% | | **16** | 74.60 | 2.35 | 1.59 | 78.53 | **9.36x** | **8.91x** | **254,686** | 16.6% | 5.0% | Amdahl serial tail = 4.15 ms; e2e ceiling 8.73x @16w, 168.85x @inf. ### Pathological ~100% conflict - 1,996 keys (honest floor) Sequential end-to-end: 664.65 ms, 30,091 TPS. | W | exec ms | commit ms | root ms | e2e ms | execSpd | **e2eSpd** | **TPS** | abort% | |--:|--:|--:|--:|--:|--:|--:|--:|--:| | 1 | 682.90 | 0.06 | 0.31 | 683.28 | 0.97x | 0.97x | 29,271 | 0.0% | | 2 | 378.02 | 0.04 | 0.37 | 378.42 | 1.76x | 1.76x | 52,851 | 6.5% | | 4 | 213.51 | 0.04 | 0.38 | 213.93 | 3.11x | 3.11x | 93,488 | 14.5% | | 8 | 133.61 | 0.04 | 0.30 | 133.95 | 4.96x | 4.96x | 149,314 | 21.2% | | **16** | 135.50 | 0.04 | 0.40 | 135.94 | **4.90x** | **4.89x** | **147,125** | 25.0% | (Here the state is tiny, so commit+root is negligible and e2e ~= exec; the limit is conflict, not the serial tail. 16 workers do not beat 8 under 25% aborts.) --- ## 6. The honest Amdahl boundary **End-to-end speedup is consistently lower than execution-phase speedup.** At 16 workers, on the realistic 0-50% conflict range: | Conflict | execSpd @16 | **e2eSpd @16** | gap | |---|--:|--:|--:| | 0% | 10.32x | **8.32x** | -1.9x | | 10% | 10.06x | **9.07x** | -1.0x | | 50% | 9.36x | **8.91x** | -0.45x | The gap is **largest at 0% conflict** and this is the important, non-obvious finding: the workload with the *most* execution parallelism (10.32x) has the *least* end-to-end benefit (8.32x), because it touches the *most* state (49,637 keys), which makes the serial commit+root the largest fraction of the pipeline (20.5% of end-to-end time at 16 workers). More parallelizable execution buys a bigger, not smaller, serial-tail penalty, exactly the Amdahl trade. **Serial fraction and ceiling.** With the commit+root tail fixed and only exec scaling, the end-to-end speedup saturates at `T_seq / (serial_tail + exec_floor)`. Measured 16-worker ceilings: **8.27x (0%), 9.64x (10%), 8.73x (50%)**. Even with infinite cores (exec -> 0) the flat-keccak tail alone caps end-to-end at 44x / 101x / 169x respectively; the real cap is far lower because exec has its own floor (residual serialization + thread overhead) and, in the real client, the commit is much heavier (§7). **Bottom line:** the honest end-to-end block-production speedup of this Rust Block-STM executor at 16 workers is roughly **8.3x to 9.1x** on realistic 0-50%-conflict execution-heavy batches, versus the ~9.4-10.3x execution-phase figure on the same box. The end-to-end win is real and large, but it is **lower** than the execution-phase headline, precisely as Amdahl predicts. --- ## 7. Real end-to-end TPS, and why it does NOT reach or imply the 273k ceiling The measured end-to-end TPS at `gas=80` (~247k-270k on realistic conflict) lands **coincidentally** near the documented 273k architectural ceiling. It must not be read as reaching or validating it. Two independent reasons: **(a) TPS is entirely a function of the assumed per-tx cost.** The gas sweep at 16 workers proves the absolute number is not a fixed capability: | per-tx cost | e2eSpd @16 (0%/10%/50%) | **end-to-end TPS @16 (0%/10%/50%)** | |---|---|---| | gas=20 (light) | 3.67x / 4.36x / 3.97x | **416k / 502k / 477k** | | gas=80 (default) | 8.32x / 9.07x / 8.91x | **247k / 270k / 255k** | | gas=200 (heavy) | 11.97x / 11.91x / 9.55x | **138k / 143k / 115k** | Change the simulated tx weight and the harness reports TPS from ~115k to ~500k on the same box. The `gas=80` proximity to 273k is a coincidence of one chosen per-tx cost, not a capability ceiling. (Note the honest inverse: heavier txs give *higher* speedup, because compute amortizes the serial tail and scheduler overhead, but *lower* TPS; lighter txs give lower speedup but higher TPS.) **(b) This is the rollup executor with a light commit proxy, not L1, not a real MPT.** The state model is balances + kv slots (not full EVM), and `state_root` is a flat keccak fold, a *lighter* commit than a real Merkle Patricia Trie. The 273k figure is a separate documented-conditions ceiling. These are different measurements under different conditions and this report neither reaches nor implies it. The invariant, honest deliverable is the **speedup curve and its Amdahl bound**, not the absolute TPS. --- ## 8. Corroboration: the real client's commit is heavier and anti-scales The Rust `state_root` is a light proxy for the commit. The real forked client's commit was measured separately (`aerenew/parallel/`, Besu Bonsai world-state, a throwaway box, DELETED). Findings there, which make the end-to-end picture in the *real* client even more conservative than this Rust harness: - Bit-identical state roots proven four independent ways (in-JVM harness, faithful storage-phase test, genesis-root equality, and re-import + validation of 253 real exported blocks with zero root mismatch). - The multi-threaded **parallel-commit** path **anti-scales**: storage-commit scaling vs w=1 was 1.00x / ~0.8x / ~0.6x / ~0.4x / ~0.4x at w=1/2/4/8/16. More commit threads made it **slower**. The shipped optimization is therefore a **sequential** storage trie (faster than the fork/join parallel one), not a parallel commit. So in the real client the commit is not merely a fixed serial tail; parallelizing it is counterproductive. The real-client end-to-end ceiling is thus **tighter** than the Rust flat-keccak numbers here, not looser. Honest direction of the error bar: down. --- ## 9. Box lifecycle - Provisioned: a dedicated single-purpose scratch VM behind its own scratch firewall, created solely for this benchmark. - Live chain 2800, its validators, and all production nodes were **not** touched. No benchmark compute ran on production infrastructure. - **Destroyed at end of run; destruction confirmed** (see `endtoend-logs-2026-07-13/`). The scratch firewall was also removed. --- ## 10. Reproduce ``` cd aerenew/parallel-executor cargo test --release # 7 serial-equivalence tests cargo run --release -- harness # 6000-comparison serializability gate cargo run --release -- bench # execution-phase speedup (the ~9.4x) cargo run --release -- endtoend # END-TO-END pipeline (this report), gas=80 ntx=20000 cargo run --release -- endtoend 200 20000 5 # heavier tx cargo run --release -- endtoend 20 20000 5 # lighter tx ``` Raw logs: `endtoend-logs-2026-07-13/` (env, cargo-test, harness, bench, endtoend at gas 80/200/20, empty stderr = zero mismatches). --- ## 11. One-paragraph honest summary The AERE Block-STM executor's ~9.4x execution-phase scaling was, until now, an execution-phase-only measurement. Measured end-to-end (parallel execution + serial commit + serial keccak state root) on a 16-vCPU box, under an execution-heavy realistic transaction mix, at 0/10/50% conflict, the real end-to-end block-production speedup at 16 workers is **~8.3x to 9.1x** - clearly lower than the execution-phase figure, exactly by the Amdahl serial tail (commit + state-root), which is 5-20% of pipeline time and grows with state footprint. Serial-equivalence (parallel state and keccak root identical to sequential) held on all 6,000 + 140 checks. The measured end-to-end TPS (~250k at gas=80) is a real rollup-batch number for this harness and this box, swings from ~115k to ~500k as the assumed per-tx cost changes, and is **not** the 273k documented-conditions architectural ceiling and does not reach or imply it. This remains a testnet-grade executor measurement; it does not make Block-STM mainnet-active on chain 2800.