Retrospective LLM-Based Complexity Evaluations

Amsterdam Comparison

Human and LLM Complexity Evaluations

A descriptive comparison of published human assessments with the study’s LLM assessments (Claude Opus 5.5, checklist revision 3) for 12 Amsterdam Execution Layer EIPs.

Why Amsterdam? It is currently the only fork with a completed human complexity evaluation; the separate manual evaluation for Hegotá is still underway (15 of 35 Hegotá candidates have a scored human checklist, and each Hegotá EIP page shows its Human status and a comparison where a scored checklist exists). This page compares the Amsterdam LLM-based and human assessments to examine how closely they align and whether their disagreements reflect systematic differences in how the complexity criteria were interpreted. It is a structured sanity check, not evidence that the LLM is calibrated to human judgment.
Different checklist revisions, same substance. The human reviewers used checklist revision 1 (24 criteria); the LLM used revision 3 (28 criteria), which phrases the same criteria more precisely. Per-criterion differences cover the 23 criteria both revisions share. Totals are compared as published, so the LLM total also includes five criteria revision 1 does not have: 4.4 points per EIP on average here. Both evaluated the same historical EIP revision.
EIPs compared
12
23 shared criteria per EIP
Mean Δ total (LLM − Human)
+6.58
mean |Δ| 7.58 · median |Δ| 6.5
Who scored higher
11 LLM · 1 Human
0 equal totals
Same tier
7 of 12
each evaluator’s own tier thresholds (revision 1: <10 · 10–19 · ≥20; revision 3: <12 · 12–22 · ≥23)

Complexity profiles per EIP

Human (revision 1) and LLM (revision 3) profiles for each EIP on one shared scale, ranked by the Human total. Δ is LLM minus Human. Open an EIP for its per-criterion comparison with the rationale from both evaluators.

  1. EIP-7928 Block-Level Access ListsΔ +12substantive drift
    HumanChecklist v1
    LLMChecklist v3
  2. HumanChecklist v1
    LLMChecklist v3
  3. HumanChecklist v1
    LLMChecklist v3
  4. HumanChecklist v1
    LLMChecklist v3
  5. HumanChecklist v1
    LLMChecklist v3
  6. HumanChecklist v1
    LLMChecklist v3
  7. HumanChecklist v1
    LLMChecklist v3
  8. EIP-7843 SLOTNUM opcodeΔ +12no substantive drift
    HumanChecklist v1
    LLMChecklist v3
  9. HumanChecklist v1
    LLMChecklist v3
  10. HumanChecklist v1
    LLMChecklist v3
  11. EIP-7976 Increase Calldata Floor CostΔ +3no substantive drift
    HumanChecklist v1
    LLMChecklist v3
  12. HumanChecklist v1
    LLMChecklist v3

Where the Evaluators Differ by Criterion

Mean per-criterion difference across the 12 EIPs, LLM minus Human, for the 23 criteria both revisions share. Positive values mean the LLM scored the criterion higher on average.

Per-criterion differences across 12 Amsterdam EIPs (shared criteria of revisions 1 and 3)
Cross-EIP interactions+1.421.581110
EVM Gas rule changes−1.081.08057
Encoding changes (RLP/SSZ)+0.750.75309
Security risks+0.671.17921
New or modified transaction validity mechanisms−0.420.75246
Block syncing changes+0.420.42309
Performance risks+0.330.67327
Edge/boundary conditions+0.330.50417
Modified system contracts−0.250.250210
New EVM gas refund−0.250.250210
New block / header fields+0.250.251011
Modified opcodes+0.170.67219
New fork activation mechanism+0.170.171011
Patterns affecting pre-existing tests−0.171.00525
Added system contracts−0.080.080111
Blob gas accounting changes−0.080.080111
Added opcodes00.000012
Added precompiles00.000012
Modified precompiles00.000012
New transaction types00.000012
Engine API changes00.171110
Transition-tool interface changes00.50228
Cryptography00.000012

Totals per EIP

Amsterdam totals per EIP: Human (checklist revision 1) and LLM (Opus 5.5, checklist revision 3)
EIP-7928Block-Level Access Lists2941+12High
EIP-8037State Creation Gas Cost Increase2843+15High
EIP-8038State-access gas cost update2014−6High vs Medium
EIP-2780Resource-based intrinsic transaction gas1325+12Medium vs High
EIP-7778Block Gas Accounting without Refunds1013+3Medium
EIP-7708ETH transfers emit a log917+8Low vs Medium
EIP-7610Revert creation in case of non-empty storage710+3Low
EIP-7843SLOTNUM opcode719+12Low vs Medium
EIP-7981Increase Access List Cost611+5Low
EIP-8024Backward compatible SWAPN, DUPN, EXCHANGE611+5Low
EIP-7976Increase Calldata Floor Cost58+3Low
EIP-7997Deterministic Factory Contract512+7Low vs Medium

What Differed Systematically?

These are descriptive differences across 12 EIPs. The comparison crosses checklist revisions and evaluation times, so it neither validates nor calibrates either evaluator.

Criterion legend for the revision-3 checklist

Every stacked bar, comparison matrix, and criterion table on this site uses the same criterion colours, abbreviations, and order. Colour marks the criterion group; the abbreviation and name identify the criterion. Scores are 0–3 per criterion (4 is exceptional; cross-EIP interactions is uncapped).

EVM surface

Opcodes, precompiles, and system contracts that are added or modified.

  • Added opcodes
    Introduces new opcodes
  • Modified opcodes
    Modifies pre-existing opcodes
  • Added precompiles
    Introduces new precompiles
  • Modified precompiles
    Modifies pre-existing precompiles logic or gas-accounting
  • Added system contracts
    Introduces new system contract, stateful or not
  • Modified system contracts
    Modifies pre-existing system contracts

Gas and accounting

Execution, blob, and state gas rules, refunds, and where charges happen inside opcodes.

  • EVM Gas rule changes
    New EVM gas accounting rules
  • State-access ordering within opcode execution · not in checklist revision 1
    Changes *where inside an opcode's execution* state is accessed, or where gas is charged relative to that access. Because a state access is recorded in the block-level access list only if execution had enough gas to reach it, this ordering is consensus-critical: moving it changes the BAL at every gas boundary of every affected opcode.
  • Blob gas accounting changes
    New Blob gas accounting rules which potentially affect pre-existing tests
  • State gas accounting changes · not in checklist revision 1
    New state gas accounting rules. State gas is the cost of *writing* state, as opposed to accessing or executing it: `StateGasCosts`, `COST_PER_STATE_BYTE`, the block-level state gas budget, and the spill path into execution gas.
  • New EVM gas refund
    New gas-refund mechanism

Blocks, transactions, and encoding

Transaction types and validity, block and header fields, encodings, syncing, and activation-time changes.

  • New transaction types
    Introduces a new transaction type
  • New or modified transaction validity mechanisms
    Creates new or modifies pre-existing transaction types' validation mechanisms
  • New block / header fields
    Introduces new block or block header fields
  • Encoding changes (RLP/SSZ)
    Introduces encoding changes at the transaction/block/interfaces level
  • Block syncing changes
    Modifies block RLP validation mechanisms that require test client syncing.
  • New fork activation mechanism
    Modifies state, internal variables, or similar, at the fork activation block

Client interfaces

Engine API and transition-tool interface changes.

  • Engine API changes
    Introduces new fields to the Engine API directives
  • Transition-tool interface changes
    Modifies or adds new fields to the transition tool interface.

Testing impact

Rework, new invariants, and new primitives required in the test framework.

  • Patterns affecting pre-existing tests
    Implements a new validation mechanism or rule that translates in reworking pre-existing tests
  • New invariant on pre-existing tests · not in checklist revision 1
    Tests that are **not about this EIP** must nonetheless assert something this EIP produces. Their logic does not change; they gain a new thing to check.
  • New test-framework primitives · not in checklist revision 1
    Requires new abstractions in the test framework itself — expectation types, modifiers, helpers — beyond writing test functions with what already exists.

Risk and validation

Security, performance, boundary conditions, and cryptography that need validation.

  • Security risks
    Introduces or modifies mechanisms that could compromise the security of the chain, users, validators, or other stakeholders, if not implemented properly.
  • Performance risks
    Introduces or modifies mechanisms and requires performance validation.
  • Edge/boundary conditions
    Feature contains edge/boundary conditions.
  • Cryptography
    Introduces new cryptography mechanisms or modifies existing functionality that involves cryptography

Coordination

Cross-EIP interactions and behavior that clients must agree on before tests exist.

  • Cross-EIP interactions
    Introduces or modifies mechanisms that affect other EIPs in either the same or past forks.
  • Unspecified behavior requiring cross-client consensus · not in checklist revision 1
    The EIP text does not determine the answer for cases a test can construct. Clients must agree on a previously unspecified detail before tests can be baselined. The cost here is coordination and re-baselining, not test writing.
Download criterion scores (CSV)Open a Human vs LLM pair in the comparison viewRead the methodology and limitations