Retrospective LLM-Based Complexity Evaluations

Primary retrospective result

Does Predicted Complexity Track Time to Mainnet?

At fork level, the study compares the time from development underway in earnest—the first devnet with at least two independent execution-layer clients—to mainnet against three summaries of predicted complexity.

Fork-level results on this page use one evaluation, Opus 5.5 · v3: Claude Opus 5.5 applying complexity checklist revision 3 to every EIP. Earlier assessments made with other models or checklist revisions remain in the research repository and are not published here.

Five forks, descriptive only. Amsterdam uses the projected 15 December 2026 mainnet date and is potentially in-sample. These plots do not establish that complexity caused delivery time; one fork can materially change the apparent pattern.

Predicted Complexity by Fork

The primary prediction is the score sum for EIPs included by each fork’s evaluation cutoff: its last initial scope-setting decision before the main implementation cycle, such as Fusaka’s April 2025 scope freeze. These EIPs form the initial fork scope; scoping continued afterward, so it is not the shipped scope. EIPs added afterward could not have informed that initial prediction, so the stacked chart reports their complexity as a separate increment. Both subtotals reflect the number and splitting of EIPs.

Shanghai4 by cutoff · 1 added later
Cancun5 by cutoff · 1 added later
Prague8 by cutoff · 3 added later
Osaka8 by cutoff · 4 added later
Amsterdam13 by cutoff · 2 added later
Values: score at cutoff, added-later score, and final-scope score per fork
Predicted Execution Layer complexity at and after each fork’s evaluation cutoff
ForkEIPs by cutoffScore at cutoffEIPs added laterAdded-later scoreFinal-scope score
Shanghai / Shapella4581361
Cancun / Dencun510915114
Prague / Pectra8159329188
Osaka / Fusaka87242496
Amsterdam / Glamsterdam13238233271

The at-cutoff subtotal is used in the shipping-time comparison; later additions remain visible as a separate segment. Hegotá is excluded because it is a forward-looking evaluation. Hover or focus a segment for its EIP count and share.

Which Kinds of Complexity Made Each Fork Heavy?

Each bar stacks the criterion scores of the EIPs included by the fork’s evaluation cutoff, so a segment is that criterion’s total contribution to the fork’s predicted complexity. Absolute mode keeps the shared scale used elsewhere on this page; composition mode normalizes each fork to 100% so the mix can be compared independently of size. Colours, names, and order match every per-EIP bar on the site.

View
Shanghai58 points · 4 EIPs
Cancun109 points · 5 EIPs
Prague159 points · 8 EIPs
Osaka72 points · 8 EIPs
Amsterdam238 points · 13 EIPs
Composition table: points per criterion and fork
Criterion points included by each fork’s evaluation cutoff
CriterionShanghaiCancunPragueOsakaAmsterdam
Added opcodes1 · 2% · 1 EIP6 · 6% · 4 EIPs004 · 2% · 2 EIPs
Modified opcodes3 · 5% · 1 EIP3 · 3% · 1 EIP3 · 2% · 1 EIP09 · 4% · 3 EIPs
Added precompiles01 · 1% · 1 EIP5 · 3% · 2 EIPs1 · 1% · 1 EIP0
Modified precompiles0004 · 6% · 2 EIPs0
Added system contracts02 · 2% · 1 EIP2 · 1% · 1 EIP01 · 0% · 1 EIP
Modified system contracts001 · 1% · 1 EIP01 · 0% · 1 EIP
EVM Gas rule changes4 · 7% · 2 EIPs1 · 1% · 1 EIP5 · 3% · 2 EIPs2 · 3% · 2 EIPs10 · 4% · 6 EIPs
State-access ordering within opcode execution2 · 3% · 1 EIP2 · 2% · 1 EIP3 · 2% · 2 EIPs05 · 2% · 2 EIPs
Blob gas accounting changes02 · 2% · 1 EIP04 · 6% · 2 EIPs0
State gas accounting changes00006 · 3% · 3 EIPs
New transaction types03 · 3% · 1 EIP3 · 2% · 1 EIP00
New or modified transaction validity mechanisms2 · 3% · 1 EIP3 · 3% · 1 EIP3 · 2% · 2 EIPs2 · 3% · 2 EIPs11 · 5% · 7 EIPs
New block / header fields3 · 5% · 1 EIP6 · 6% · 2 EIPs9 · 6% · 3 EIPs09 · 4% · 3 EIPs
Encoding changes (RLP/SSZ)3 · 5% · 1 EIP6 · 6% · 2 EIPs12 · 8% · 4 EIPs3 · 4% · 1 EIP9 · 4% · 3 EIPs
Block syncing changes3 · 5% · 1 EIP5 · 5% · 2 EIPs10 · 6% · 4 EIPs4 · 6% · 3 EIPs7 · 3% · 3 EIPs
New fork activation mechanism00003 · 1% · 1 EIP
Engine API changes2 · 3% · 1 EIP3 · 3% · 2 EIPs5 · 3% · 3 EIPs04 · 2% · 3 EIPs
Transition-tool interface changes2 · 3% · 1 EIP3 · 3% · 2 EIPs8 · 5% · 5 EIPs1 · 1% · 1 EIP6 · 3% · 4 EIPs
Patterns affecting pre-existing tests6 · 10% · 4 EIPs8 · 7% · 5 EIPs7 · 4% · 6 EIPs7 · 10% · 5 EIPs26 · 11% · 13 EIPs
New invariant on pre-existing tests2 · 3% · 1 EIP4 · 4% · 2 EIPs8 · 5% · 4 EIPs08 · 3% · 4 EIPs
New test-framework primitives3 · 5% · 2 EIPs4 · 4% · 3 EIPs10 · 6% · 7 EIPs5 · 7% · 4 EIPs13 · 5% · 9 EIPs
Security risks4 · 7% · 3 EIPs9 · 8% · 5 EIPs12 · 8% · 7 EIPs7 · 10% · 7 EIPs18 · 8% · 13 EIPs
Performance risks2 · 3% · 2 EIPs7 · 6% · 4 EIPs9 · 6% · 6 EIPs5 · 7% · 5 EIPs16 · 7% · 8 EIPs
Edge/boundary conditions6 · 10% · 4 EIPs12 · 11% · 5 EIPs15 · 9% · 7 EIPs9 · 13% · 6 EIPs24 · 10% · 13 EIPs
Cryptography03 · 3% · 1 EIP4 · 3% · 2 EIPs1 · 1% · 1 EIP0
Cross-EIP interactions5 · 9% · 3 EIPs7 · 6% · 4 EIPs10 · 6% · 8 EIPs9 · 13% · 7 EIPs27 · 11% · 13 EIPs
Unspecified behavior requiring cross-client consensus5 · 9% · 3 EIPs9 · 8% · 5 EIPs15 · 9% · 8 EIPs8 · 11% · 6 EIPs21 · 9% · 11 EIPs
Total5810915972238
Criterion legend for the fork composition bars

Every stacked bar, comparison matrix, and criterion table on this site uses the same criterion colours, abbreviations, and order. Colour marks the criterion group; the abbreviation and name identify the criterion. Scores are 0–3 per criterion (4 is exceptional; cross-EIP interactions is uncapped).

EVM surface

Opcodes, precompiles, and system contracts that are added or modified.

  • Added opcodes
    Introduces new opcodes
  • Modified opcodes
    Modifies pre-existing opcodes
  • Added precompiles
    Introduces new precompiles
  • Modified precompiles
    Modifies pre-existing precompiles logic or gas-accounting
  • Added system contracts
    Introduces new system contract, stateful or not
  • Modified system contracts
    Modifies pre-existing system contracts

Gas and accounting

Execution, blob, and state gas rules, refunds, and where charges happen inside opcodes.

  • EVM Gas rule changes
    New EVM gas accounting rules
  • State-access ordering within opcode execution · not in checklist revision 1
    Changes *where inside an opcode's execution* state is accessed, or where gas is charged relative to that access. Because a state access is recorded in the block-level access list only if execution had enough gas to reach it, this ordering is consensus-critical: moving it changes the BAL at every gas boundary of every affected opcode.
  • Blob gas accounting changes
    New Blob gas accounting rules which potentially affect pre-existing tests
  • State gas accounting changes · not in checklist revision 1
    New state gas accounting rules. State gas is the cost of *writing* state, as opposed to accessing or executing it: `StateGasCosts`, `COST_PER_STATE_BYTE`, the block-level state gas budget, and the spill path into execution gas.
  • New EVM gas refund
    New gas-refund mechanism

Blocks, transactions, and encoding

Transaction types and validity, block and header fields, encodings, syncing, and activation-time changes.

  • New transaction types
    Introduces a new transaction type
  • New or modified transaction validity mechanisms
    Creates new or modifies pre-existing transaction types' validation mechanisms
  • New block / header fields
    Introduces new block or block header fields
  • Encoding changes (RLP/SSZ)
    Introduces encoding changes at the transaction/block/interfaces level
  • Block syncing changes
    Modifies block RLP validation mechanisms that require test client syncing.
  • New fork activation mechanism
    Modifies state, internal variables, or similar, at the fork activation block

Client interfaces

Engine API and transition-tool interface changes.

  • Engine API changes
    Introduces new fields to the Engine API directives
  • Transition-tool interface changes
    Modifies or adds new fields to the transition tool interface.

Testing impact

Rework, new invariants, and new primitives required in the test framework.

  • Patterns affecting pre-existing tests
    Implements a new validation mechanism or rule that translates in reworking pre-existing tests
  • New invariant on pre-existing tests · not in checklist revision 1
    Tests that are **not about this EIP** must nonetheless assert something this EIP produces. Their logic does not change; they gain a new thing to check.
  • New test-framework primitives · not in checklist revision 1
    Requires new abstractions in the test framework itself — expectation types, modifiers, helpers — beyond writing test functions with what already exists.

Risk and validation

Security, performance, boundary conditions, and cryptography that need validation.

  • Security risks
    Introduces or modifies mechanisms that could compromise the security of the chain, users, validators, or other stakeholders, if not implemented properly.
  • Performance risks
    Introduces or modifies mechanisms and requires performance validation.
  • Edge/boundary conditions
    Feature contains edge/boundary conditions.
  • Cryptography
    Introduces new cryptography mechanisms or modifies existing functionality that involves cryptography

Coordination

Cross-EIP interactions and behavior that clients must agree on before tests exist.

  • Cross-EIP interactions
    Introduces or modifies mechanisms that affect other EIPs in either the same or past forks.
  • Unspecified behavior requiring cross-client consensus · not in checklist revision 1
    The EIP text does not determine the answer for cases a test can construct. Clients must agree on a previously unspecified detail before tests can be baselined. The cost here is coordination and re-baselining, not test writing.

Fork Shipping Time Versus Predicted Complexity, With Hegotá

Very uncertain: five forks, and the Hegotá scope is still changing. The line is a least-squares fit over five points, one of them projected (Amsterdam). The nested shaded bands are its central 50%, 80% and 95% prediction intervals for one new fork, the innermost most strongly shaded; they assume the straight-line model is right, which five points cannot confirm. Hegotá is shown only as a position on the complexity axis. Read where it lands against earlier forks, not a ship date.
Hegotá scope from EIP-8081 after ACDE 247

Lists as of EIPs f154af8 (2026-10-08); scores from the Opus 5.5 · v3 assessment of each EIP at 6dac5e7. EIPs added since are listed as not yet assessed and never count as zero.

EL-rubric total 385 across 19 scored EIPs

Choose individual EIPs
SFI: Scheduled for Inclusion
CFI: Considered for Inclusion
PFI: Proposed for Inclusion

Exploratory: scale complexity by frontier AI capability
More info

In short: each fork’s complexity is divided by a factor that grows with the AI coding capability available when its client development began. Shanghai and Cancun are the reference, so the slider never moves them; later forks shrink as α grows. AI can speed up only part of the work: coordinating Ethereum’s independent client teams sets a floor. This view is off by default, display-only, and never a study result.

What H measures

H is METR’s 50% task-completion time horizon: the length of software task, in human working minutes, that a model completes with 50% success on METR’s task suite. For each fork, H is the horizon of the best model released on or before the fork’s first devnet with at least two independent execution-layer clients, the point where implementation work begins in earnest. The values come from METR’s published results file, retrieved 2026-10-08; the table below the chart lists the model and horizon used for every fork.

Why Shanghai and Cancun do not move

H₀ is the horizon at the earliest fork’s devnet, 0.6 min, and every fork at that level has H ÷ H₀ = 1, so log₂(1) = 0 and its divisor is 1 for any α. Shanghai’s first two-client devnet launched on 2022-12-09; Cancun’s first two-client devnet launched on 2023-01-23. METR measured gpt-3.5-turbo-instruct (released March 2022) and next GPT-4 (released 14 March 2023), with nothing in between; ChatGPT, released on 30 November 2022, has no METR score. Both devnets fall in that gap, so both get the same model and the same horizon. Cancun’s tie is partly an artefact of that gap: most of its 415-day implementation window came after GPT-4.

What α means

α is the share of testing and implementation effort assumed to be saved for every doubling of the horizon. It is chosen by the reader, not estimated: five forks cannot separate it from everything else that changed between them. The fit line and its 50/80/95% bands are recomputed on the scaled values for display. Hegotá has no devnet yet, so it uses the latest measured model (Claude Mythos Preview (early), 1044.78 min); its eventual horizon will be at least that, so its scaled total is an upper bound.

Why AI help has a floor

Much of a fork’s implementation burden is human coordination rather than code. Ethereum deliberately runs several independent client implementations, so that a bug in any one of them cannot take down the network; client diversity is a security property, not an inefficiency. Its cost is that every fork is implemented separately by each client team, and the teams must agree on every observable behaviour, including edge cases. That agreement comes through specification work, shared test suites, multi-client devnets and coordination calls.

AI can shorten the work inside each team: reading the specification, writing code and tests, and finding bugs. It does much less for the coordination between teams, which moves at the pace of people reaching agreement and of devnets that have to run before a fork ships. That work sets a floor under how much AI can help protocol developers take on more complex forks. Dividing a fork’s whole complexity by one factor, as this view does, therefore overstates the effect; read α as applying only to the share of the work AI can actually accelerate.

Limitations

The horizon is measured on METR’s tasks, not Ethereum client work; the newest model is not necessarily the one client teams used; and teams adopted AI tools with a lag the measure ignores. Treat the view as a what-if, not an adjustment the study endorses.

Sources: METR, Time Horizon 1.1 · METR, Measuring AI Ability to Complete Long Tasks · METR results file

Calendar days from each fork’s first devnet with at least two independent EL clients to mainnet, against the complexity score sum of its initial scope: the EIPs in scope at the fork’s last initial scope-setting decision before its main implementation cycle (Task 04b, provisional). EIPs added later are excluded. Open point: Amsterdam, projected mainnet 15 December 2026. Shaded bands: central 50% (strongest shading), 80% and 95% prediction intervals around the dashed least-squares line. With the exploratory AI-capability scaling on, each fork’s complexity is divided by its scaling factor and the line and bands are refitted on the scaled values.
Least-squares fit over the five forks (Opus 5.5 · v3)
EvaluationDays per pointIntercept (days)Residual SE (days)R²
Opus 5.5 · v31.3743123.1999.310.578
Hegotá EL-rubric totals by EIP-8081 list as of ACDE 247 (Opus 5.5 · v3)
ScenarioScored EIPsNot applicableNot yet assessedEL-rubric total
SFI20077
SFI+CFI1921385
SFI+CFI+PFI2582513
Frontier AI capability at each fork’s first ≥2-client devnet: METR’s 50% time horizon, measured by METR (Time Horizon 1.1), values from METR’s published results file retrieved 2026-10-08
ForkFirst ≥2-EL devnetFrontier model then50% horizon, min [METR CI]METR measurement
Shanghai / Shapella2022-12-09gpt-3.5-turbo-instruct (2022-03-15)0.6 [0.26–1.12]gpt_3_5_turbo_instruct
Cancun / Dencun2023-01-23gpt-3.5-turbo-instruct (2022-03-15)0.6 [0.26–1.12]gpt_3_5_turbo_instruct
Prague / Pectra2024-05-17GPT-4o (2024-05-13)6.99 [4–12.91]gpt_4o_inspect
Osaka / Fusaka2025-05-26o3 (2025-04-16)119.73 [74.62–190.94]o3_inspect
Amsterdam / Glamsterdam2025-11-04GPT-5 (2025-08-07)203.01 [112.64–405.55]gpt_5_2025_08_07_inspect
Hegotánot yetClaude Mythos Preview (early) (2026-04-07; latest measured, lower bound)≥ 1044.78 [508.88–3304.26]claude_mythos_preview_early_inspect

Other Summaries of Predicted Complexity

Fork shipping time compared with the at-cutoff sum of scores from EIPs assessed as High complexity.
Fork shipping time compared with the highest-scoring Execution Layer EIP present by each fork’s evaluation cutoff.
Fork-level values plotted above (Opus 5.5 · v3)
Shanghai / Shapellawithdrawals-devnet-0
2022-12-09
2023-04-121245826EIP-4895: 263
Cancun / Dencundencun-devnet-4
2023-01-23
2024-03-1341510973EIP-4844: 435
Prague / Pectrapectra-devnet-0
2024-05-17
2025-05-0735515983EIP-7002: 3129
Osaka / Fusakafusaka-devnet-0
2025-05-26
2025-12-03191720EIP-7892: 1424
Amsterdam / Glamsterdam (projected)bal-devnet-0
2025-11-04
2026-12-15405238109EIP-8037: 4333

Cancun is the only fork whose ≥2-EL start differs from its first EL devnet: devnets 1–3 ran a single patched geth fork, so this clock begins at dencun-devnet-4. Cancun also largely distinguishes the total-load and critical-path readings in this five-fork sample.

EIP-level observed proxies

As a separate view, each retrospective fork–EIP relationship is compared with two score-blind counts derived from public artifacts. Predicted complexity is the Opus 5.5 · v3 score:

Amsterdam is omitted from these plots because it had not reached mainnet when the data was frozen, so its counts were still incomplete.

Specification rework across 34 fork–EIP relationships in the four shipped forks.
New EIP interactions across 34 fork–EIP relationships in the four shipped forks.
Descriptive rank associations across shipped forks
MetricNPooled Spearman ρ
Specification rework340.62
New EIP interactions340.51
Browse all EIP assessmentsCompare human and LLM evaluations