The public history carried kat/__pycache__/mlkem768_reference.cpython-314.pyc, a compiled Python artifact embedding the operator's absolute local path. Text secret scanners do not read compiled binaries, which is exactly how it slipped through, and removing it from the tip would have left it reachable through the old root commits. So this repository is republished from a single clean root. This root also carries, from the previously unpublished line of work: - corrected LICENSE year, LICENSING.md, VERIFY-POLICY.md, and CITATIONS-UNRESOLVED.md remeasured 2026-08-11 (101 paths, README aligned) - O-018: run_consensus_verification.py ran 19 of 29 models and reported PASS; it now runs all 29, and computemarket_smt.py gains resolveByTimeout / reclaimUnsettled cases plus a negative control - O-006: the word 'audited' removed from next to Bouncy Castle, twice, after a concurrent edit resurrected it - O-014: prior art named and dated - Algorand's native falcon_verify shipped about ten months before AERE's precompiles; the primacy claim is withdrawn where it was implied - bench/ scripts parametrized so they actually run for an outsider (the earlier textual sanitization left $STAGING unexpanded inside Python strings) - AIP-2/AIP-3 errata with measured figures, spec remeasurements at 2026-08-01, and the spec-zk-stack retractions (owner is an operational key, not the Foundation; 'maximally sound' withdrawn; aggregator V1 deprecated) The redacted bench-host environment files from the sanitized line are kept exactly as published; the unredacted local variants are not carried.
8.0 KiB
aere-block-stm
A real Block-STM optimistic-concurrency parallel EVM-transaction executor for AERE, implemented from scratch in Rust with zero external dependencies.
It executes a batch of transactions across many CPU cores and commits a single state root that is provably identical to sequential execution (deterministic serializability), which is exactly what makes the resulting root a valid rollup commitment.
Reference algorithm: Gelashvili, Spiegelman, Xiang, Danezis, Li, Malkhi, Xia, Zhou, "Block-STM: Scaling Blockchain Execution by Turning Ordering Curse to a Performance Blessing" (arXiv:2203.06871, 2022) - the same design that powers Aptos, and the family Sui and Monad build on.
Honest scope (read this first)
What this is NOT: this does not make AERE's Layer-1 base execution parallel. AERE's L1 runs Hyperledger Besu, which executes transactions sequentially. Parallelizing Besu itself would be a fork of a Java execution client, a multi-month effort, and is explicitly out of scope here. Nothing in this module changes how the AERE L1 mainnet executes blocks. The L1 stays Besu-sequential.
What this IS: a standalone, correct, benchmarked parallel executor that
AERE's rollup layer can use. AERE already ships a Rollup-as-a-Service stack
(contracts/contracts/raas/AereRaaSFactory.sol +
AereRollupSettlement.sol). A rollup sequencer executes a batch of L2
transactions off-chain, computes the resulting state root, and commits it via
AereRollupSettlement.proposeStateRoot(epoch, root, l2BlockHash). That batch
execution is a pure, well-defined, embarrassingly-parallelizable workload, and
it is the concrete, real path to parallel execution on AERE. This module is that
executor, wired into the sequencer (see ../rollup-sequencer/).
So the honest one-line framing is:
Real Block-STM parallel executor, wired into the AERE rollup sequencer; the L1 base layer stays Besu-sequential.
What it implements (the real algorithm)
- Multi-version memory (MVMemory) -
src/mvmemory.rs. For every state key, an orderedtxn_index -> cellmap. A read by txnireturns the write of the highest txnj < i, giving each speculative execution a serial-equivalent view. Sharded across 256 mutexes so non-conflicting txns never contend. - Read-set / write-set tracking -
src/parallel.rs. Each execution records exactly what it read (and the version it read) and what it wrote. - Dynamic dependency detection - a read that hits an
ESTIMATEplaceholder (left by a lower txn that is currently being re-executed) aborts the reader and parks it on the blocking txn; it is re-armed when the blocker finishes. No static dependency analysis, no condvars. - Abort + re-execution -
src/scheduler.rs. Collaborative validation: validating txnire-checks its read set against current MVMemory; on mismatch it is aborted, its incarnation bumped, its writes turned into ESTIMATEs, and it is rescheduled. Cursors are decreased so the affected suffix is revalidated. - Serial-equivalent commit - the block is done only when every txn has been validated against the final memory. The committed state provably equals sequential execution.
- State model - account balances + a key/value contract store
(
src/types.rs), with genuine EVM-style transactions whose control flow and written values depend on reads:Transfer,Sweep(value-dependent),Increment(the canonical shared-counter serializability probe), andAmmSwap(two-slot DeFi pool conflict). - keccak256 state root -
src/keccak.rs, a from-scratch keccak256 (verified against known-answer vectors) folds the final state into a real, EVM-compatible 32-byte root for on-chain commitment.
Zero external crates: the scheduler, MVMemory, PRNG and keccak are all built on
std. Build is fully offline and reproducible ($0).
Correctness guarantee (the hard part, proven)
The whole point of Block-STM is that the parallel result must equal
sequential execution. Both paths call the exact same transaction-semantics
function (src/vm.rs::execute_txn), so any divergence can only come from a
concurrency-control bug. The harness proves there are none:
cargo run --release -- harness
runs 6 workload profiles x 200 randomized batches x {1,2,4,8,16} threads =
6000 comparisons (1200 distinct randomized batches), and asserts the
Block-STM committed state equals the sequential oracle for every one. Profiles
include adversarial 100%-conflict cases (all-hot, brutal-single-counter).
Result:
total (batch x threads) comparisons: 6000
mismatches : 0
RESULT: PASS -- parallel Block-STM == sequential on ALL cases
cargo test --release additionally checks:
shared_counter_is_serialized: 500 txns all incrementing one slot yield exactly 500 at 1/2/4/8/16 threads (naive parallelism would lose updates).parallel_equals_sequential_random: 40 seeds x 4 thread counts.keccak_known_answer: keccak256("") and keccak256("abc") match the standard vectors.
Measured speedup (honest numbers)
Measured on the AERE infra box: 16 physical cores, Linux 6.8, Rust 1.96.
20,000 transactions per batch, each carrying a realistic simulated EVM execution
cost (gas = 80 keccak-rounds, a stand-in for the microseconds a real EVM tx
spends in the interpreter). Speedup is versus the single-threaded sequential
oracle. Median of 7 runs.
| Workload | 2 cores | 4 cores | 8 cores | 16 cores | abort % @16 |
|---|---|---|---|---|---|
| Low-conflict (~2% hot-touch) | 1.78x | 3.22x | 5.71x | 8.59x | 0.3% |
| Medium-conflict (~15%) | 1.87x | 3.54x | 6.24x | 9.31x | 1.9% |
| High-conflict (single hot counter, 15% of txns) | 1.70x | 3.04x | 4.87x | 4.62x | 24% |
| Pathological (EVERY tx hits ONE slot) | 0.85x | 0.85x | 0.82x | 0.73x | 19% |
Honest reading of these numbers:
- Low/medium-conflict batches scale nearly linearly: ~8.6-9.3x on 16 cores.
- The high-conflict batch still gets ~4.6x because only ~15% of its transactions touch the single hot slot; the rest are independent transfers.
- The pathological batch (every transaction increments the same slot) has zero available parallelism. Block-STM correctly serializes it and is actually slower than sequential (0.73x) because it pays for aborted speculation. This is the expected, honest worst case: parallelism collapses under total conflict, and no correct parallel executor can beat sequential when there is nothing to parallelize.
There is also an honest floor case worth stating: if transactions do
near-zero compute (gas = 0), the concurrency-control overhead dominates and
the parallel executor is slower than sequential. That is why the benchmark
charges each transaction a realistic execution cost - real EVM transactions are
compute-heavy (opcode interpretation, keccak, ecrecover), which is precisely the
regime Block-STM is designed for. Run bench 0 to see the floor, bench 200
for heavier txs.
Reproduce:
cargo run --release -- bench # gas=80, 20k txns (default)
cargo run --release -- bench 200 # heavier txs
cargo run --release -- bench 0 # trivial txs: scheduler-overhead floor
Sequencer integration
The rollup sequencer calls the execute subcommand (a JSON-out CLI) with a batch
and gets back a self-checked keccak256 state root to commit on-chain. See
../rollup-sequencer/README.md and ../rollup-sequencer/src/blockstm/ for the
TypeScript wrapper and the exact proposeStateRoot flow. The library API
(execute_block_parallel, execute_block_sequential, state_root) is also
usable directly from Rust.
Every execute run computes the batch both sequentially and in parallel and
refuses to emit a root unless they match (parallelEqualsSequential: true), so a
concurrency bug can never silently produce a bad rollup commitment.
Build
source ~/.cargo/env # on the AERE infra box
cargo build --release
cargo test --release
cargo run --release -- harness
cargo run --release -- bench