aere-research/parallel-executor/BLOCKSTM-ENDTOEND-THROUGHPUT-2026-07-13.md
Aere Network 6cb0140fae Republished from a clean root: the compiled artifact is gone from history, and the local line of work joins the sanitized public line
The public history carried kat/__pycache__/mlkem768_reference.cpython-314.pyc,
a compiled Python artifact embedding the operator's absolute local path. Text
secret scanners do not read compiled binaries, which is exactly how it slipped
through, and removing it from the tip would have left it reachable through the
old root commits. So this repository is republished from a single clean root.

This root also carries, from the previously unpublished line of work:
- corrected LICENSE year, LICENSING.md, VERIFY-POLICY.md, and
  CITATIONS-UNRESOLVED.md remeasured 2026-08-11 (101 paths, README aligned)
- O-018: run_consensus_verification.py ran 19 of 29 models and reported PASS;
  it now runs all 29, and computemarket_smt.py gains resolveByTimeout /
  reclaimUnsettled cases plus a negative control
- O-006: the word 'audited' removed from next to Bouncy Castle, twice, after a
  concurrent edit resurrected it
- O-014: prior art named and dated - Algorand's native falcon_verify shipped
  about ten months before AERE's precompiles; the primacy claim is withdrawn
  where it was implied
- bench/ scripts parametrized so they actually run for an outsider (the
  earlier textual sanitization left $STAGING unexpanded inside Python strings)
- AIP-2/AIP-3 errata with measured figures, spec remeasurements at 2026-08-01,
  and the spec-zk-stack retractions (owner is an operational key, not the
  Foundation; 'maximally sound' withdrawn; aggregator V1 deprecated)
The redacted bench-host environment files from the sanitized line are kept
exactly as published; the unredacted local variants are not carried.
2026-08-15 13:52:14 +03:00

326 lines
16 KiB
Markdown

# Block-STM end-to-end block-production throughput (measured)
**Date:** 2026-07-13
**Module:** `parallel-executor/` (the from-scratch Rust Block-STM executor)
**Scope:** ISOLATED benchmark on a throwaway cloud box. Nothing here touches live
chain 2800, its validators, or any production node.
This does **not** make Block-STM mainnet-active; that remains a founder-supervised
fork activation. The site's honest framing ("testnet demonstration, not
mainnet-active; execution-phase win") stays intact. This report only **adds a
real, measured end-to-end number** to that framing.
---
## 0. Why this exists (the gap being closed)
The existing `bench` command times **only** the Block-STM execution phase and
reports the documented near-linear execution-phase scaling
(~1.94x / 3.61x / 6.19x / 9.38x at 2/4/8/16 workers). That is an honest
execution-phase number, but it is **not** an end-to-end block-throughput number.
A block is not finished when execution finishes; it is finished when the state is
**committed** and a **state root** is produced. Those tail phases are serial, so
by Amdahl's law the end-to-end speedup must be **lower** than the execution-phase
speedup. The site states exactly this gap:
> "this is an execution-phase win; an end-to-end throughput win still needs real
> execution-heavy load plus a parallel commit."
This report closes that gap with a real measurement: the **full pipeline**
(parallel execution + commit + state-root), parallel vs sequential, at 1/2/4/8/16
workers, across conflict ratios, with serial-equivalence checked on every run.
New reproducible harness: `aere-block-stm endtoend [gas] [ntx] [reps]`
(added to `src/main.rs`; phase timing exposed via the additive
`execute_block_parallel_split` in `src/parallel.rs`).
---
## 1. Environment (honest)
| | |
|---|---|
| Box | Throwaway cloud VM, 16 vCPU, 32 GB, Ubuntu 24.04, US region |
| CPU | 16 cores, **Thread(s)/core = 1** (no in-guest SMT), 2.0 GHz |
| Kernel / toolchain | Linux 6.8.0, rustc 1.97.0, release build (`lto=true`, `codegen-units=1`) |
| Firewall | dedicated scratch firewall, SSH restricted to operator addresses |
| Lifecycle | provisioned for this benchmark, **DELETED at end** (see §9) |
**Caveat, stated up front:** this is a **shared-vCPU** cloud box, not the
dedicated-core infra machine the original 9.38x figure was taken on. Shared vCPU
adds run-to-run variance to **absolute** timings. Mitigations: 16 single-threaded
vCPUs (so worker count maps to real cores with no hyperthread muddying), median of
7 runs, and the fact that the **headline deliverable is a ratio** (end-to-end vs
execution-phase speedup, and the serial fraction) in which shared-vCPU noise
affects numerator and denominator together and largely cancels.
---
## 2. Methodology
**Pipeline, timed in three phases** (see `endtoend()` in `src/main.rs`):
1. **exec** - Block-STM parallel execution + validation + re-execution. PARALLEL;
this is the part that scales with worker count. Timed inside
`execute_block_parallel_split` between thread spawn and join.
2. **commit** - `MvMemory::snapshot_final`: materialize the multi-version memory
into the flat committed `StateMap`. SERIAL (single-threaded, after join).
3. **root** - `keccak::state_root`: keccak256 over the sorted committed state, the
on-chain rollup commitment. SERIAL.
`end-to-end = exec + commit + root`. The **sequential baseline** is
`execute_block_sequential` (which produces the live map directly, so it has no
separate commit step) `+ state_root`. The state-root cost is identical serial work
in both numerator and denominator, i.e. a pure Amdahl serial tail.
**Workload (execution-heavy, not empty blocks).** The deterministic SplitMix64
generator emits a realistic EVM-style mix per batch: **60% ERC-20 transfers**
(two balance touches), **15% sweeps** (value-dependent), **15% shared-counter
increments** (the serializability stressor), **10% AMM swaps** (two-slot pool
conflict). Each transaction also pays a deterministic simulated EVM cost of `gas`
keccak-rounds (a stand-in for opcode/keccak/ecrecover interpreter time), so the
executor is timed doing real per-tx work rather than pure scheduling overhead.
**Conflict ratios.** `conflict_ppm` tunes how often a transaction touches the
shared hot set. Four points span the spectrum: **0%** (disjoint accounts, 64
contracts), **10%**, **50%**, and a **pathological ~100%** single-hot-slot floor.
The labelled % is the *designed* hot-touch rate; the honest *realized* conflict
indicator is the **measured abort%** in each table.
**Defaults:** `gas=80`, `ntx=20000`, median of 7 (chosen to match the regime the
documented ~9.38x execution-phase figure was taken in, so the two are directly
comparable). A gas sweep (§6) shows sensitivity.
**Serial-equivalence, checked every run.** On every single parallel run the harness
asserts the committed state equals the sequential oracle **and** the two keccak
state roots are bit-identical. A single mismatch aborts the whole run. This is the
property that makes the parallel root a valid commitment.
---
## 3. Serial-equivalence confirmation (the non-negotiable gate)
| Check | Result |
|---|---|
| `cargo test --release` (7 tests incl. shared-counter serialization, 40-seed random equivalence, keccak KAT) | **7/7 PASS** |
| Correctness harness: 6 profiles x 200 batches x {1,2,4,8,16} threads | **6000 comparisons, 0 mismatches** |
| End-to-end harness inline check (state == oracle AND roots identical), every run | **140 checks, 0 mismatches** |
`RESULT: PASS -- every parallel run committed the exact sequential state + root.`
The `04-endtoend.stderr` log is empty (zero mismatch lines). Raw logs in
`endtoend-logs-2026-07-13/`.
---
## 4. Execution-phase baseline on THIS box (for comparison)
Running the existing `bench` (execution-phase only) on the same box reproduces the
documented range, confirming the box is a faithful stand-in:
| Workload | 2 | 4 | 8 | **16** (vs sequential) |
|---|---|---|---|---|
| Low-conflict (~2%) | 1.74x | 3.24x | 5.68x | **8.30x** |
| Medium-conflict (~15%) | 1.82x | 3.48x | 6.18x | **9.62x** |
| High-conflict (~80%) | 1.74x | 3.10x | 4.96x | **5.13x** |
| Pathological (all one slot) | 0.93x | 0.86x | 0.82x | **0.76x** |
(Medium-conflict vs-1-thread: 1.94x / 3.71x / 6.59x / 10.26x, matching the
documented 1.94x/3.61x/6.19x/9.38x.) These are **execution-phase** numbers.
---
## 5. END-TO-END RESULTS (the deliverable)
`gas=80`, `ntx=20000`, median of 7. `execSpd` = execution-phase speedup;
`e2eSpd` = **end-to-end** speedup (exec+commit+root) vs the sequential end-to-end
baseline. TPS is real transactions/second of the full pipeline. All rows PASS
serial-equivalence.
### 0% designed conflict - 49,637 committed state keys (largest state)
Sequential end-to-end: 672.84 ms (root-share 1.4%), 29,725 TPS.
| W | exec ms | commit ms | root ms | e2e ms | execSpd | **e2eSpd** | **TPS** | abort% | (commit+root)/e2e |
|--:|--:|--:|--:|--:|--:|--:|--:|--:|--:|
| 1 | 705.51 | 9.01 | 12.52 | 727.04 | 0.94x | 0.93x | 27,509 | 0.0% | 3.0% |
| 2 | 354.89 | 9.56 | 11.41 | 375.85 | 1.87x | 1.79x | 53,212 | 0.0% | 5.6% |
| 4 | 184.02 | 9.72 | 12.37 | 206.11 | 3.60x | 3.26x | 97,037 | 0.1% | 10.7% |
| 8 | 97.99 | 8.52 | 8.31 | 114.82 | 6.77x | 5.86x | 174,191 | 0.1% | 14.7% |
| **16** | 64.23 | 8.08 | 8.51 | 80.83 | **10.32x** | **8.32x** | **247,444** | 0.3% | 20.5% |
Amdahl serial tail (commit+root) = 15.33 ms; e2e ceiling 8.27x @16w, 43.88x @inf cores.
### 10% designed conflict - 18,342 keys
Sequential end-to-end: 672.71 ms, 29,731 TPS.
| W | exec ms | commit ms | root ms | e2e ms | execSpd | **e2eSpd** | **TPS** | abort% | tail% |
|--:|--:|--:|--:|--:|--:|--:|--:|--:|--:|
| 1 | 699.73 | 5.15 | 4.77 | 709.65 | 0.96x | 0.95x | 28,183 | 0.0% | 1.4% |
| 2 | 366.73 | 5.31 | 4.38 | 376.42 | 1.83x | 1.79x | 53,133 | 0.1% | 2.6% |
| 4 | 187.09 | 5.15 | 3.88 | 196.13 | 3.58x | 3.43x | 101,975 | 0.3% | 4.6% |
| 8 | 96.64 | 5.06 | 4.29 | 105.99 | 6.93x | 6.35x | 188,699 | 0.7% | 8.8% |
| **16** | 66.57 | 4.37 | 3.20 | 74.14 | **10.06x** | **9.07x** | **269,750** | 1.2% | 10.2% |
Amdahl serial tail = 6.68 ms; e2e ceiling 9.64x @16w, 100.71x @inf.
### 50% designed conflict - 9,006 keys
Sequential end-to-end: 699.97 ms, 28,573 TPS.
| W | exec ms | commit ms | root ms | e2e ms | execSpd | **e2eSpd** | **TPS** | abort% | tail% |
|--:|--:|--:|--:|--:|--:|--:|--:|--:|--:|
| 1 | 712.46 | 2.63 | 1.96 | 717.06 | 0.98x | 0.98x | 27,892 | 0.0% | 0.6% |
| 2 | 367.61 | 2.46 | 2.01 | 372.07 | 1.90x | 1.88x | 53,753 | 2.1% | 1.2% |
| 4 | 195.00 | 2.58 | 1.98 | 199.56 | 3.58x | 3.51x | 100,222 | 5.7% | 2.3% |
| 8 | 113.28 | 2.50 | 1.72 | 117.51 | 6.17x | 5.96x | 170,202 | 10.7% | 3.6% |
| **16** | 74.60 | 2.35 | 1.59 | 78.53 | **9.36x** | **8.91x** | **254,686** | 16.6% | 5.0% |
Amdahl serial tail = 4.15 ms; e2e ceiling 8.73x @16w, 168.85x @inf.
### Pathological ~100% conflict - 1,996 keys (honest floor)
Sequential end-to-end: 664.65 ms, 30,091 TPS.
| W | exec ms | commit ms | root ms | e2e ms | execSpd | **e2eSpd** | **TPS** | abort% |
|--:|--:|--:|--:|--:|--:|--:|--:|--:|
| 1 | 682.90 | 0.06 | 0.31 | 683.28 | 0.97x | 0.97x | 29,271 | 0.0% |
| 2 | 378.02 | 0.04 | 0.37 | 378.42 | 1.76x | 1.76x | 52,851 | 6.5% |
| 4 | 213.51 | 0.04 | 0.38 | 213.93 | 3.11x | 3.11x | 93,488 | 14.5% |
| 8 | 133.61 | 0.04 | 0.30 | 133.95 | 4.96x | 4.96x | 149,314 | 21.2% |
| **16** | 135.50 | 0.04 | 0.40 | 135.94 | **4.90x** | **4.89x** | **147,125** | 25.0% |
(Here the state is tiny, so commit+root is negligible and e2e ~= exec; the limit is
conflict, not the serial tail. 16 workers do not beat 8 under 25% aborts.)
---
## 6. The honest Amdahl boundary
**End-to-end speedup is consistently lower than execution-phase speedup.** At 16
workers, on the realistic 0-50% conflict range:
| Conflict | execSpd @16 | **e2eSpd @16** | gap |
|---|--:|--:|--:|
| 0% | 10.32x | **8.32x** | -1.9x |
| 10% | 10.06x | **9.07x** | -1.0x |
| 50% | 9.36x | **8.91x** | -0.45x |
The gap is **largest at 0% conflict** and this is the important, non-obvious
finding: the workload with the *most* execution parallelism (10.32x) has the
*least* end-to-end benefit (8.32x), because it touches the *most* state (49,637
keys), which makes the serial commit+root the largest fraction of the pipeline
(20.5% of end-to-end time at 16 workers). More parallelizable execution buys a
bigger, not smaller, serial-tail penalty, exactly the Amdahl trade.
**Serial fraction and ceiling.** With the commit+root tail fixed and only exec
scaling, the end-to-end speedup saturates at `T_seq / (serial_tail + exec_floor)`.
Measured 16-worker ceilings: **8.27x (0%), 9.64x (10%), 8.73x (50%)**. Even with
infinite cores (exec -> 0) the flat-keccak tail alone caps end-to-end at
44x / 101x / 169x respectively; the real cap is far lower because exec has its own
floor (residual serialization + thread overhead) and, in the real client, the
commit is much heavier (§7).
**Bottom line:** the honest end-to-end block-production speedup of this Rust
Block-STM executor at 16 workers is roughly **8.3x to 9.1x** on realistic
0-50%-conflict execution-heavy batches, versus the ~9.4-10.3x execution-phase
figure on the same box. The end-to-end win is real and large, but it is
**lower** than the execution-phase headline, precisely as Amdahl predicts.
---
## 7. Real end-to-end TPS, and why it does NOT reach or imply the 273k ceiling
The measured end-to-end TPS at `gas=80` (~247k-270k on realistic conflict) lands
**coincidentally** near the documented 273k architectural ceiling. It must not be read
as reaching or validating it. Two independent reasons:
**(a) TPS is entirely a function of the assumed per-tx cost.** The gas sweep at 16
workers proves the absolute number is not a fixed capability:
| per-tx cost | e2eSpd @16 (0%/10%/50%) | **end-to-end TPS @16 (0%/10%/50%)** |
|---|---|---|
| gas=20 (light) | 3.67x / 4.36x / 3.97x | **416k / 502k / 477k** |
| gas=80 (default) | 8.32x / 9.07x / 8.91x | **247k / 270k / 255k** |
| gas=200 (heavy) | 11.97x / 11.91x / 9.55x | **138k / 143k / 115k** |
Change the simulated tx weight and the harness reports TPS from ~115k to ~500k on
the same box. The `gas=80` proximity to 273k is a coincidence of one chosen
per-tx cost, not a capability ceiling. (Note the honest inverse: heavier txs give
*higher* speedup, because compute amortizes the serial tail and scheduler
overhead, but *lower* TPS; lighter txs give lower speedup but higher TPS.)
**(b) This is the rollup executor with a light commit proxy, not L1, not a real
MPT.** The state model is balances + kv slots (not full EVM), and `state_root` is
a flat keccak fold, a *lighter* commit than a real Merkle Patricia Trie. The
273k figure is a separate documented-conditions ceiling. These are different
measurements under different conditions and this report neither reaches nor
implies it.
The invariant, honest deliverable is the **speedup curve and its Amdahl bound**,
not the absolute TPS.
---
## 8. Corroboration: the real client's commit is heavier and anti-scales
The Rust `state_root` is a light proxy for the commit. The real forked client's
commit was measured separately (`aerenew/parallel/`, Besu Bonsai world-state, a
throwaway box, DELETED). Findings there, which make the end-to-end picture in the
*real* client even more conservative than this Rust harness:
- Bit-identical state roots proven four independent ways (in-JVM harness, faithful
storage-phase test, genesis-root equality, and re-import + validation of 253 real
exported blocks with zero root mismatch).
- The multi-threaded **parallel-commit** path **anti-scales**: storage-commit
scaling vs w=1 was 1.00x / ~0.8x / ~0.6x / ~0.4x / ~0.4x at w=1/2/4/8/16. More
commit threads made it **slower**. The shipped optimization is therefore a
**sequential** storage trie (faster than the fork/join parallel one), not a
parallel commit.
So in the real client the commit is not merely a fixed serial tail; parallelizing
it is counterproductive. The real-client end-to-end ceiling is thus **tighter**
than the Rust flat-keccak numbers here, not looser. Honest direction of the error
bar: down.
---
## 9. Box lifecycle
- Provisioned: a dedicated single-purpose scratch VM behind its own scratch firewall,
created solely for this benchmark.
- Live chain 2800, its validators, and all production nodes were **not** touched.
No benchmark compute ran on production infrastructure.
- **Destroyed at end of run; destruction confirmed** (see
`endtoend-logs-2026-07-13/`). The scratch firewall was also removed.
---
## 10. Reproduce
```
cd aerenew/parallel-executor
cargo test --release # 7 serial-equivalence tests
cargo run --release -- harness # 6000-comparison serializability gate
cargo run --release -- bench # execution-phase speedup (the ~9.4x)
cargo run --release -- endtoend # END-TO-END pipeline (this report), gas=80 ntx=20000
cargo run --release -- endtoend 200 20000 5 # heavier tx
cargo run --release -- endtoend 20 20000 5 # lighter tx
```
Raw logs: `endtoend-logs-2026-07-13/` (env, cargo-test, harness, bench,
endtoend at gas 80/200/20, empty stderr = zero mismatches).
---
## 11. One-paragraph honest summary
The AERE Block-STM executor's ~9.4x execution-phase scaling was, until now, an
execution-phase-only measurement. Measured end-to-end (parallel execution +
serial commit + serial keccak state root) on a 16-vCPU box, under an
execution-heavy realistic transaction mix, at 0/10/50% conflict, the real
end-to-end block-production speedup at 16 workers is **~8.3x to 9.1x** - clearly
lower than the execution-phase figure, exactly by the Amdahl serial tail (commit +
state-root), which is 5-20% of pipeline time and grows with state footprint.
Serial-equivalence (parallel state and keccak root identical to sequential) held
on all 6,000 + 140 checks. The measured end-to-end TPS (~250k at gas=80) is a real
rollup-batch number for this harness and this box, swings from ~115k to ~500k as
the assumed per-tx cost changes, and is **not** the 273k documented-conditions
architectural ceiling and does not reach or imply it. This remains a testnet-grade
executor measurement; it does not make Block-STM mainnet-active on chain 2800.