Primary retrospective result
Does Predicted Complexity Track Time to Mainnet?
At fork level, the study compares the time from development underway in earnest—the first devnet with at least two independent execution-layer clients—to mainnet against three summaries of predicted complexity.
Fork-level results on this page use one evaluation, Opus 5.5 · v3: Claude Opus 5.5 applying complexity checklist revision 3 to every EIP. Earlier assessments made with other models or checklist revisions remain in the research repository and are not published here.
Predicted Complexity by Fork
The primary prediction is the score sum for EIPs included by each fork’s evaluation cutoff: its last initial scope-setting decision before the main implementation cycle, such as Fusaka’s April 2025 scope freeze. These EIPs form the initial fork scope; scoping continued afterward, so it is not the shipped scope. EIPs added afterward could not have informed that initial prediction, so the stacked chart reports their complexity as a separate increment. Both subtotals reflect the number and splitting of EIPs.
- Included by cutoff (primary prediction)
- Added after cutoff
Values: score at cutoff, added-later score, and final-scope score per fork
| Fork | EIPs by cutoff | Score at cutoff | EIPs added later | Added-later score | Final-scope score |
|---|---|---|---|---|---|
| Shanghai / Shapella | 4 | 58 | 1 | 3 | 61 |
| Cancun / Dencun | 5 | 109 | 1 | 5 | 114 |
| Prague / Pectra | 8 | 159 | 3 | 29 | 188 |
| Osaka / Fusaka | 8 | 72 | 4 | 24 | 96 |
| Amsterdam / Glamsterdam | 13 | 238 | 2 | 33 | 271 |
The at-cutoff subtotal is used in the shipping-time comparison; later additions remain visible as a separate segment. Hegotá is excluded because it is a forward-looking evaluation. Hover or focus a segment for its EIP count and share.
Which Kinds of Complexity Made Each Fork Heavy?
Each bar stacks the criterion scores of the EIPs included by the fork’s evaluation cutoff, so a segment is that criterion’s total contribution to the fork’s predicted complexity. Absolute mode keeps the shared scale used elsewhere on this page; composition mode normalizes each fork to 100% so the mix can be compared independently of size. Colours, names, and order match every per-EIP bar on the site.
Composition table: points per criterion and fork
| Criterion | Shanghai | Cancun | Prague | Osaka | Amsterdam |
|---|---|---|---|---|---|
| Added opcodes | 1 · 2% · 1 EIP | 6 · 6% · 4 EIPs | 0 | 0 | 4 · 2% · 2 EIPs |
| Modified opcodes | 3 · 5% · 1 EIP | 3 · 3% · 1 EIP | 3 · 2% · 1 EIP | 0 | 9 · 4% · 3 EIPs |
| Added precompiles | 0 | 1 · 1% · 1 EIP | 5 · 3% · 2 EIPs | 1 · 1% · 1 EIP | 0 |
| Modified precompiles | 0 | 0 | 0 | 4 · 6% · 2 EIPs | 0 |
| Added system contracts | 0 | 2 · 2% · 1 EIP | 2 · 1% · 1 EIP | 0 | 1 · 0% · 1 EIP |
| Modified system contracts | 0 | 0 | 1 · 1% · 1 EIP | 0 | 1 · 0% · 1 EIP |
| EVM Gas rule changes | 4 · 7% · 2 EIPs | 1 · 1% · 1 EIP | 5 · 3% · 2 EIPs | 2 · 3% · 2 EIPs | 10 · 4% · 6 EIPs |
| State-access ordering within opcode execution | 2 · 3% · 1 EIP | 2 · 2% · 1 EIP | 3 · 2% · 2 EIPs | 0 | 5 · 2% · 2 EIPs |
| Blob gas accounting changes | 0 | 2 · 2% · 1 EIP | 0 | 4 · 6% · 2 EIPs | 0 |
| State gas accounting changes | 0 | 0 | 0 | 0 | 6 · 3% · 3 EIPs |
| New transaction types | 0 | 3 · 3% · 1 EIP | 3 · 2% · 1 EIP | 0 | 0 |
| New or modified transaction validity mechanisms | 2 · 3% · 1 EIP | 3 · 3% · 1 EIP | 3 · 2% · 2 EIPs | 2 · 3% · 2 EIPs | 11 · 5% · 7 EIPs |
| New block / header fields | 3 · 5% · 1 EIP | 6 · 6% · 2 EIPs | 9 · 6% · 3 EIPs | 0 | 9 · 4% · 3 EIPs |
| Encoding changes (RLP/SSZ) | 3 · 5% · 1 EIP | 6 · 6% · 2 EIPs | 12 · 8% · 4 EIPs | 3 · 4% · 1 EIP | 9 · 4% · 3 EIPs |
| Block syncing changes | 3 · 5% · 1 EIP | 5 · 5% · 2 EIPs | 10 · 6% · 4 EIPs | 4 · 6% · 3 EIPs | 7 · 3% · 3 EIPs |
| New fork activation mechanism | 0 | 0 | 0 | 0 | 3 · 1% · 1 EIP |
| Engine API changes | 2 · 3% · 1 EIP | 3 · 3% · 2 EIPs | 5 · 3% · 3 EIPs | 0 | 4 · 2% · 3 EIPs |
| Transition-tool interface changes | 2 · 3% · 1 EIP | 3 · 3% · 2 EIPs | 8 · 5% · 5 EIPs | 1 · 1% · 1 EIP | 6 · 3% · 4 EIPs |
| Patterns affecting pre-existing tests | 6 · 10% · 4 EIPs | 8 · 7% · 5 EIPs | 7 · 4% · 6 EIPs | 7 · 10% · 5 EIPs | 26 · 11% · 13 EIPs |
| New invariant on pre-existing tests | 2 · 3% · 1 EIP | 4 · 4% · 2 EIPs | 8 · 5% · 4 EIPs | 0 | 8 · 3% · 4 EIPs |
| New test-framework primitives | 3 · 5% · 2 EIPs | 4 · 4% · 3 EIPs | 10 · 6% · 7 EIPs | 5 · 7% · 4 EIPs | 13 · 5% · 9 EIPs |
| Security risks | 4 · 7% · 3 EIPs | 9 · 8% · 5 EIPs | 12 · 8% · 7 EIPs | 7 · 10% · 7 EIPs | 18 · 8% · 13 EIPs |
| Performance risks | 2 · 3% · 2 EIPs | 7 · 6% · 4 EIPs | 9 · 6% · 6 EIPs | 5 · 7% · 5 EIPs | 16 · 7% · 8 EIPs |
| Edge/boundary conditions | 6 · 10% · 4 EIPs | 12 · 11% · 5 EIPs | 15 · 9% · 7 EIPs | 9 · 13% · 6 EIPs | 24 · 10% · 13 EIPs |
| Cryptography | 0 | 3 · 3% · 1 EIP | 4 · 3% · 2 EIPs | 1 · 1% · 1 EIP | 0 |
| Cross-EIP interactions | 5 · 9% · 3 EIPs | 7 · 6% · 4 EIPs | 10 · 6% · 8 EIPs | 9 · 13% · 7 EIPs | 27 · 11% · 13 EIPs |
| Unspecified behavior requiring cross-client consensus | 5 · 9% · 3 EIPs | 9 · 8% · 5 EIPs | 15 · 9% · 8 EIPs | 8 · 11% · 6 EIPs | 21 · 9% · 11 EIPs |
| Total | 58 | 109 | 159 | 72 | 238 |
Criterion legend for the fork composition bars
Every stacked bar, comparison matrix, and criterion table on this site uses the same criterion colours, abbreviations, and order. Colour marks the criterion group; the abbreviation and name identify the criterion. Scores are 0–3 per criterion (4 is exceptional; cross-EIP interactions is uncapped).
EVM surface
Opcodes, precompiles, and system contracts that are added or modified.
- Added opcodesIntroduces new opcodes
- Modified opcodesModifies pre-existing opcodes
- Added precompilesIntroduces new precompiles
- Modified precompilesModifies pre-existing precompiles logic or gas-accounting
- Added system contractsIntroduces new system contract, stateful or not
- Modified system contractsModifies pre-existing system contracts
Gas and accounting
Execution, blob, and state gas rules, refunds, and where charges happen inside opcodes.
- EVM Gas rule changesNew EVM gas accounting rules
- State-access ordering within opcode execution · not in checklist revision 1Changes *where inside an opcode's execution* state is accessed, or where gas is charged relative to that access. Because a state access is recorded in the block-level access list only if execution had enough gas to reach it, this ordering is consensus-critical: moving it changes the BAL at every gas boundary of every affected opcode.
- Blob gas accounting changesNew Blob gas accounting rules which potentially affect pre-existing tests
- State gas accounting changes · not in checklist revision 1New state gas accounting rules. State gas is the cost of *writing* state, as opposed to accessing or executing it: `StateGasCosts`, `COST_PER_STATE_BYTE`, the block-level state gas budget, and the spill path into execution gas.
- New EVM gas refundNew gas-refund mechanism
Blocks, transactions, and encoding
Transaction types and validity, block and header fields, encodings, syncing, and activation-time changes.
- New transaction typesIntroduces a new transaction type
- New or modified transaction validity mechanismsCreates new or modifies pre-existing transaction types' validation mechanisms
- New block / header fieldsIntroduces new block or block header fields
- Encoding changes (RLP/SSZ)Introduces encoding changes at the transaction/block/interfaces level
- Block syncing changesModifies block RLP validation mechanisms that require test client syncing.
- New fork activation mechanismModifies state, internal variables, or similar, at the fork activation block
Client interfaces
Engine API and transition-tool interface changes.
- Engine API changesIntroduces new fields to the Engine API directives
- Transition-tool interface changesModifies or adds new fields to the transition tool interface.
Testing impact
Rework, new invariants, and new primitives required in the test framework.
- Patterns affecting pre-existing testsImplements a new validation mechanism or rule that translates in reworking pre-existing tests
- New invariant on pre-existing tests · not in checklist revision 1Tests that are **not about this EIP** must nonetheless assert something this EIP produces. Their logic does not change; they gain a new thing to check.
- New test-framework primitives · not in checklist revision 1Requires new abstractions in the test framework itself — expectation types, modifiers, helpers — beyond writing test functions with what already exists.
Risk and validation
Security, performance, boundary conditions, and cryptography that need validation.
- Security risksIntroduces or modifies mechanisms that could compromise the security of the chain, users, validators, or other stakeholders, if not implemented properly.
- Performance risksIntroduces or modifies mechanisms and requires performance validation.
- Edge/boundary conditionsFeature contains edge/boundary conditions.
- CryptographyIntroduces new cryptography mechanisms or modifies existing functionality that involves cryptography
Coordination
Cross-EIP interactions and behavior that clients must agree on before tests exist.
- Cross-EIP interactionsIntroduces or modifies mechanisms that affect other EIPs in either the same or past forks.
- Unspecified behavior requiring cross-client consensus · not in checklist revision 1The EIP text does not determine the answer for cases a test can construct. Clients must agree on a previously unspecified detail before tests can be baselined. The cost here is coordination and re-baselining, not test writing.
Fork Shipping Time Versus Predicted Complexity, With Hegotá
Exploratory: scale complexity by frontier AI capability
More info
In short: each fork’s complexity is divided by a factor that grows with the AI coding capability available when its client development began. Shanghai and Cancun are the reference, so the slider never moves them; later forks shrink as α grows. AI can speed up only part of the work: coordinating Ethereum’s independent client teams sets a floor. This view is off by default, display-only, and never a study result.
What H measures
H is METR’s 50% task-completion time horizon: the length of software task, in human working minutes, that a model completes with 50% success on METR’s task suite. For each fork, H is the horizon of the best model released on or before the fork’s first devnet with at least two independent execution-layer clients, the point where implementation work begins in earnest. The values come from METR’s published results file, retrieved 2026-10-08; the table below the chart lists the model and horizon used for every fork.
Why Shanghai and Cancun do not move
H₀ is the horizon at the earliest fork’s devnet, 0.6 min, and every fork at that level has H ÷ H₀ = 1, so log₂(1) = 0 and its divisor is 1 for any α. Shanghai’s first two-client devnet launched on 2022-12-09; Cancun’s first two-client devnet launched on 2023-01-23. METR measured gpt-3.5-turbo-instruct (released March 2022) and next GPT-4 (released 14 March 2023), with nothing in between; ChatGPT, released on 30 November 2022, has no METR score. Both devnets fall in that gap, so both get the same model and the same horizon. Cancun’s tie is partly an artefact of that gap: most of its 415-day implementation window came after GPT-4.
What α means
α is the share of testing and implementation effort assumed to be saved for every doubling of the horizon. It is chosen by the reader, not estimated: five forks cannot separate it from everything else that changed between them. The fit line and its 50/80/95% bands are recomputed on the scaled values for display. Hegotá has no devnet yet, so it uses the latest measured model (Claude Mythos Preview (early), 1044.78 min); its eventual horizon will be at least that, so its scaled total is an upper bound.
Why AI help has a floor
Much of a fork’s implementation burden is human coordination rather than code. Ethereum deliberately runs several independent client implementations, so that a bug in any one of them cannot take down the network; client diversity is a security property, not an inefficiency. Its cost is that every fork is implemented separately by each client team, and the teams must agree on every observable behaviour, including edge cases. That agreement comes through specification work, shared test suites, multi-client devnets and coordination calls.
AI can shorten the work inside each team: reading the specification, writing code and tests, and finding bugs. It does much less for the coordination between teams, which moves at the pace of people reaching agreement and of devnets that have to run before a fork ships. That work sets a floor under how much AI can help protocol developers take on more complex forks. Dividing a fork’s whole complexity by one factor, as this view does, therefore overstates the effect; read α as applying only to the share of the work AI can actually accelerate.
Limitations
The horizon is measured on METR’s tasks, not Ethereum client work; the newest model is not necessarily the one client teams used; and teams adopted AI tools with a lag the measure ignores. Treat the view as a what-if, not an adjustment the study endorses.
Sources: METR, Time Horizon 1.1 · METR, Measuring AI Ability to Complete Long Tasks · METR results file
Loading interactive chart…
| Evaluation | Days per point | Intercept (days) | Residual SE (days) | R² |
|---|---|---|---|---|
| Opus 5.5 · v3 | 1.3743 | 123.19 | 99.31 | 0.578 |
| Scenario | Scored EIPs | Not applicable | Not yet assessed | EL-rubric total |
|---|---|---|---|---|
| SFI | 2 | 0 | 0 | 77 |
| SFI+CFI | 19 | 2 | 1 | 385 |
| SFI+CFI+PFI | 25 | 8 | 2 | 513 |
| Fork | First ≥2-EL devnet | Frontier model then | 50% horizon, min [METR CI] | METR measurement |
|---|---|---|---|---|
| Shanghai / Shapella | 2022-12-09 | gpt-3.5-turbo-instruct (2022-03-15) | 0.6 [0.26–1.12] | gpt_3_5_turbo_instruct |
| Cancun / Dencun | 2023-01-23 | gpt-3.5-turbo-instruct (2022-03-15) | 0.6 [0.26–1.12] | gpt_3_5_turbo_instruct |
| Prague / Pectra | 2024-05-17 | GPT-4o (2024-05-13) | 6.99 [4–12.91] | gpt_4o_inspect |
| Osaka / Fusaka | 2025-05-26 | o3 (2025-04-16) | 119.73 [74.62–190.94] | o3_inspect |
| Amsterdam / Glamsterdam | 2025-11-04 | GPT-5 (2025-08-07) | 203.01 [112.64–405.55] | gpt_5_2025_08_07_inspect |
| Hegotá | not yet | Claude Mythos Preview (early) (2026-04-07; latest measured, lower bound) | ≥ 1044.78 [508.88–3304.26] | claude_mythos_preview_early_inspect |
Other Summaries of Predicted Complexity
- To estimate when Execution Layer development began in earnest, we define the start date as the launch of the first devnet running at least two independent Execution Layer implementations.
- Filled points represent the four shipped forks; the open Amsterdam point uses the projected 15 December 2026 mainnet date.
- Each plot contains the same five forks and uses only the complexity of the initial fork scope, summarized differently (Opus 5.5 · v3).
Loading interactive chart…
Loading interactive chart…
| Shanghai / Shapella | withdrawals-devnet-0 2022-12-09 | 2023-04-12 | 124 | 58 | 26 | EIP-4895: 26 | 3 |
| Cancun / Dencun | dencun-devnet-4 2023-01-23 | 2024-03-13 | 415 | 109 | 73 | EIP-4844: 43 | 5 |
| Prague / Pectra | pectra-devnet-0 2024-05-17 | 2025-05-07 | 355 | 159 | 83 | EIP-7002: 31 | 29 |
| Osaka / Fusaka | fusaka-devnet-0 2025-05-26 | 2025-12-03 | 191 | 72 | 0 | EIP-7892: 14 | 24 |
| Amsterdam / Glamsterdam (projected) | bal-devnet-0 2025-11-04 | 2026-12-15 | 405 | 238 | 109 | EIP-8037: 43 | 33 |
Cancun is the only fork whose ≥2-EL start differs from its first EL devnet: devnets 1–3 ran a single patched geth fork, so this clock begins at dencun-devnet-4. Cancun also largely distinguishes the total-load and critical-path readings in this five-fork sample.
EIP-level observed proxies
As a separate view, each retrospective fork–EIP relationship is compared with two score-blind counts derived from public artifacts. Predicted complexity is the Opus 5.5 · v3 score:
- Specification rework counts substantive revisions committed after the assessment cutoff and before mainnet. Clarifications can count, and author commit practices can affect the total.
- New EIP interactions counts distinct
requiresorinteracts_withrelationships first evidenced after the assessment cutoff. It captures newly documented interaction surface, not necessarily a dependency that first arose at that moment.
Amsterdam is omitted from these plots because it had not reached mainnet when the data was frozen, so its counts were still incomplete.
Loading interactive chart…
Loading interactive chart…
| Metric | N | Pooled Spearman ρ |
|---|---|---|
| Specification rework | 34 | 0.62 |
| New EIP interactions | 34 | 0.51 |