All Core Devs · Hegotá scoping
Could EIP complexity assessments help size Hegotá?
How much are we putting into Hegotá, compared with what previous forks actually took to ship?
Study site: Retrospective LLM-Based Complexity Evaluations · Hegotá evaluations
Background
Amsterdam saw the introduction of the “EIP Complexity Assessment Template (v1)”
The aim was to help gauge how difficult it would be to implement and ship each Execution Layer EIP to mainnet, and to help us size and scope devnets.
v1 · November 2025
Template v1: 24 criteria, nominal range 0–72. Used for the human Amsterdam assessments.
v2 · August 2026
Template v2 (ethspecs/pm#101): calibrated against the testing work each Amsterdam EIP actually produced. Five new criteria, cross-EIP interactions uncapped: 28 criteria, range 0–84.
Example criterion · v1 and v2
Added opcodes — Introduces new opcodes
- 0No new opcodes are introduced.
- 1A new simple opcode is introduced (no data portion, no complex stack mechanics, and a constant gas cost).
- 2Multiple new simple opcodes are introduced, or a single new complex opcode is introduced (has data portion, or complex stack mechanics, or a dynamic gas cost).
- 3Multiple new opcodes are introduced, and at least one of them is complex.
One answer per criterion; the scores add up to the EIP’s total.
The study
A retrospective study to gauge the reliability of the complexity assessments
- Score every EIP of previous forks retrospectively: does total fork complexity track how long the fork took to ship? Is there any signal?
- By hand that is a huge amount of work; with LLMs it is very feasible, and assessments are already AI-assisted anyway.
- Fully automated: same model (Claude Opus 5.5), template and sandbox for every EIP, each read at its fork’s scoping version; Hegotá at current EIPs master.
- It estimates testing and specification work, not engineering hours.
- Execution Layer only: the templates assess EL effort, so consensus-layer work such as ePBS (EIP-7732) is not counted.
More: Study design and limitations · Human vs LLM on Amsterdam · STEEL
Template v3 · ethspecs/pm#147 (open)
Humans read the template well; LLMs need less room to read between the lines
Template v3 keeps the same 28 criteria and scores, but pins down what each one means.
+OP Added opcodes
“Introduces new opcodes.”
Introduces previously undefined EVM instructions. Contract bytecode, transaction frames and new introspection selector values are not new opcodes. A prerequisite’s instruction already exists in the baseline.
INV New invariant on pre-existing tests
“Tests that are not about this EIP must nonetheless assert something this EIP produces.” Level 1: “a narrow, contrived category”; level 2: “a broad category”.
Baseline tests must check an output that did not exist before, such as a new log, header or receipt field, block commitment or protocol-mandated storage write. Changed gas costs or expected values are rework, not a new assertion.
Previous forks · EL side only
Forks vary a lot, and big forks are dominated by one or two headliners
EIP-8037: 43 (18%) ★
+ ePBS (EIP-7732) Amsterdam’s CL headliner, invisible to an EL-only rubric
EL complexity score sum of EIPs included by each fork’s evaluation cutoff. Blue: heaviest EIP (Amsterdam: the official EL headliner, BALs). Amber ★: EIP-8037, State Creation Gas Cost Increase, Amsterdam’s unofficial headliner.
Live demo
This is where the current scope sits (place your bets)
SFI: 77 points across 2 scored EIPs.
Very uncertain. Five data points, scope still changing, and the model knows how past forks turned out. The grey strip marks the range past forks cover; beyond it the dashed line is extrapolation. Lists as of ACDE 247, 8 October 2026 (EIPs f154af8), scores from EIPs 6dac5e7; every candidate: Hegotá evaluations.
Are we optimizing for the right outcome?
- Complexity scores provide signals — but imo these numbers aren’t the goal.
- There’s a bigger opportunity here…
- We should optimize this scoping process for hardening specs faster: giving authors concrete feedback.
- The question isn’t whether we use automated tooling and human judgment, but how we combine them most effectively.
Links: study site · these slides · template v3 (ethspecs/pm#147) · source · STEEL