Retrospective LLM-Based Complexity Evaluations

EIP complexity assessment

EIP-4758: Deactivate SELFDESTRUCT

Assessed in Hegotá. The score describes the EIP text available at the snapshot, not the EIP as it stands today.

ProspectiveHegotáSnapshot 2026-10-07EIP-8081: CFILayers: execution
LLM Completescore 11
Human Available in open PRscore 8 · Checklist revision 2· ethspecs/pm #105
Other checklist versions (1)

Evaluated on: · Spec revision: 2026-10-07 · 6dac5e7491 · EIP-8081 list: CFI

Scope at the cutoff. EIP-4758 renames SELFDESTRUCT to SENDALL and changes what it does. The opcode now only moves all ETH in the executing account to the target. It no longer deletes code or storage, and it no longer changes the nonce. The EIP also removes every refund related to SELFDESTRUCT. As a result, redeploying a contract at the same address with CREATE2 after destroying it no longer works. Applications that use SELFDESTRUCT only to pull out funds keep working.

11LowLow
Evaluator
LLMChecklist v3
Confidence
Medium
Under-specified at assessment cutoff
Yes — 4 criteria affected
Plausible range
11–15 (Low–Medium)
Snapshot
2026-10-07 · EIP revision 6dac5e7491 (2026-10-07)
Score bands · Checklist revision 3
  • Low <12
  • Medium 12–22
  • High ≥23

28 criteria scored 0–3 (4 in exceptional cases; cross-EIP interactions is uncapped); nominal maximum 84.

Complexity profile

Each segment is one criterion's contribution to the LLM total. Hover or focus a segment for its score and rationale.

Top complexity drivers

  1. Modified opcodes3
  2. Patterns affecting pre-existing tests2
  3. Cross-EIP interactions2
  4. Unspecified behavior requiring cross-client consensus2

Under-specified at assessment cutoff: Yes

The EIP text available at the assessment cutoff left material behavior unresolved. The affected criteria and the plausible total range record that uncertainty.

Why: Several details are left open. The Abstract says ETH goes to the 'caller' while the Specification says 'target'. The EIP does not say whether SENDALL halts execution. It does not say what happens to ETH when target==self. Its gas cost is not restated.

Plausible total

11–15
recorded score 11 · plausible tiers Low, Medium

Unresolved questions at the cutoff (4)
  • Does ETH go to the caller (Abstract) or to the stack-provided target (Specification)?
  • Does SENDALL halt execution as SELFDESTRUCT does, or does execution continue?
  • When target==self, is the balance kept or burned?
  • Do the SELFDESTRUCT gas schedule (base, cold access, new-account charges) and static-context restriction stay unchanged?
Notable ambiguities noted by the assessor (4)
  • The Abstract ('to the caller') contradicts the Specification ('to the target') on who receives the ETH.
  • The EIP does not say whether SENDALL halts execution.
  • The EIP does not say what happens to ETH when the target is the account itself. In the baseline, a contract created in the same transaction burns ETH this way.
  • The refund removal appears to be a no-op against the Amsterdam baseline. The EIP predates the baseline and does not address that.

Criterion breakdown

EIP-4758 Hegotá: LLM criterion scores and rationale
CriterionScoreWhy this scoreEvidence / uncertainty
Modified opcodes3The state effects of SELFDESTRUCT change: it no longer deletes code, storage or the nonce, even for contracts created in the same transaction. This is a semantic change beyond gas or ordering, which is level 3.
  • eip.md · Specification — "now only immediately moves all ETH in the account to the target; it no longer destroys code or storage or alters the nonce" The state effects of SELFDESTRUCT change.
Confidence: High
Patterns affecting pre-existing testsUnder-specified2Several families need their expected results reworked in specific cases. These include SELFDESTRUCT tests where the contract was created in the same transaction (including SELFDESTRUCT inside initcode and burning ETH by naming itself as beneficiary), CREATE/CREATE2 redeploy-after-destroy scenarios, and state/access-list expectations that involve destroyed accounts. Ordinary SELFDESTRUCT cases on pre-existing contracts that only send ETH are mostly unaffected in the baseline. This is localized rework in several families, which is level 2.
  • eip.md · Specification — "it no longer destroys code or storage or alters the nonce" In the baseline, SELFDESTRUCT still destroys a contract created in the same transaction. Those cases now keep their code, storage and nonce.
  • eip.md · Backwards Compatibility — "re-created at the same address using `CREATE2` (after a `SELFDESTRUCT`)" CREATE2 redeployment after a SELFDESTRUCT now fails, which changes the expected results of create/collision scenarios.
Confidence: Medium
Uncertainty: How much rework is needed depends on the unresolved halting semantics and on how target==self is handled. If SENDALL did not halt, rework would spread across ordinary SELFDESTRUCT tests (level 3).
Cross-EIP interactions2Coordinated cases with CREATE2 are needed: a contract created with CREATE2, then SENDALL, then redeployed with CREATE2 at the same address now hits an address collision because code and nonce persist. Cases are also needed for SELFDESTRUCT inside CREATE/CREATE2 initcode. EIP-20 is only cited as an application example, so it does not qualify as an interaction.
  • eip.md · Backwards Compatibility — "re-created at the same address using `CREATE2` (after a `SELFDESTRUCT`)" The EIP's change interacts with CREATE2 redeployment and address collision behavior.
  • supporting/eip-20.md · (whole file) — "This file was moved" EIP-20 is only an application-level example of token burning. It has no protocol interaction.
Confidence: Medium
Uncertainty: CREATE2 and the collision rule are not given as EIP numbers in the supplied documents. Interactions with the baseline SELFDESTRUCT restriction and block-level access lists are inferred from general knowledge of the baseline.
Unspecified behavior requiring cross-client consensus2Several local outcomes have competing readings. The recipient is the caller in the Abstract but the target in the Specification. The EIP does not say whether execution halts after SENDALL. It also does not say what happens when target==self, where the baseline same-transaction case burns ETH. Expected test results cannot be fixed until these are agreed, which is level 2.
  • eip.md · Abstract — "send all Ether in the account to the caller" The Abstract says ETH goes to the caller.
  • eip.md · Specification — "moves all ETH in the account to the target" The Specification says ETH goes to the target. This contradicts the Abstract on who receives the ETH.
  • eip.md · Specification The EIP does not say whether SENDALL still halts execution, or what happens to ETH when the target is the account itself.
Confidence: Medium
Uncertainty: Most readers would likely take the Specification ('target') and keep halting semantics, but the text does not say so.
Security risksUnder-specified1The changed security conditions can be checked locally. Tests need to show that a contract stays callable after SENDALL, that code and storage persist, and that redeploying at the same address collides. No cross-component trust invariant changes, so this is level 1.
  • eip.md · Security Considerations The EIP lists broken patterns: burning tokens via SELFDESTRUCT, CREATE2 redeployment, using destruction to block reuse, and upgradeability through redeployment.
Confidence: Medium
Uncertainty: If SENDALL does not halt, new re-entrancy and continuation behaviors would appear, which could raise this to level 2.
Edge/boundary conditionsUnder-specified1One changed mechanism, the SENDALL semantics, has boundary-sensitive outcomes: same-transaction creation versus a pre-existing account, beneficiary equal to self, zero balance, and initcode context. These are all parts of one rule, which is level 1.
  • eip.md · Specification — "no longer destroys code or storage or alters the nonce" Deletion no longer depends on whether the contract was created in the same transaction. Edge cases include target==self, a zero balance, SELFDESTRUCT inside initcode, and re-entry after SENDALL.
Confidence: Medium
Uncertainty: If halting semantics or the target==self outcome were settled as separate rules, this could count as multiple mechanisms (level 2).
Show 22 zero-score criteria
Zero-score criteria (Checklist revision 3)
CriterionScoreWhy this scoreEvidence / uncertainty
Added opcodes0Replacing an existing instruction's semantics is scored under Modified opcodes.
  • eip.md · Abstract — "renames the `SELFDESTRUCT` opcode to `SENDALL`" SENDALL replaces the semantics of an existing opcode. It is not a new instruction.
Added precompiles0No precompile is introduced.
  • eip.md · Specification The EIP adds no precompiles.
Modified precompiles0No precompile changes.
  • eip.md · Specification The EIP modifies no precompiles.
Added system contracts0No system contract is added.
  • eip.md · Specification The EIP introduces no system contract.
Modified system contracts0No system contract's rules change.
  • eip.md · Specification The EIP does not mention any system contract.
EVM Gas rule changesUnder-specified0In the Amsterdam baseline, SELFDESTRUCT already gives no gas refund (general protocol knowledge about the base fork), so removing the refunds changes nothing there. The EIP changes no opcode cost, charging rule or limit, which fits level 0.
  • eip.md · Specification — "All refunds related to `SELFDESTRUCT` are removed" The only gas-related statement removes SELFDESTRUCT refunds. It defines no new costs and no new accounting mechanism.
Uncertainty: The spec does not say whether SENDALL still halts execution. If execution continued after it, the gas results of baseline sequences would change (level 1). The spec also does not restate the gas schedule of the renamed opcode.
State-access ordering within opcode execution0The opcode touches the same accounts as before, and the EIP does not change the order of its state accesses or gas charges. This is level 0.
  • eip.md · Specification — "now only immediately moves all ETH in the account to the target" The opcode still reads and writes the same accounts (self and target). The EIP defines no new ordering of accesses or gas charges.
Uncertainty: Account deletion in same-transaction-created contracts no longer happens, which may change which state writes a block-level access list records. The EIP says nothing about this.
Blob gas accounting changes0The EIP does not change blob-gas accounting.
  • eip.md · Specification The EIP does not mention blobs or blob gas.
State gas accounting changes0The EIP adds no state-gas charging and changes no state-gas parameters.
  • eip.md · Specification The EIP defines no state-gas costs, budgets or spill rules.
New EVM gas refund0Removing a refund belongs under the gas criterion, not here. The EIP adds no refund mechanism.
  • eip.md · Specification — "All refunds related to `SELFDESTRUCT` are removed" The EIP removes refunds and adds none.
New transaction types0No new transaction envelope.
  • eip.md · Specification The EIP adds no transaction type.
New or modified transaction validity mechanisms0Transaction validity does not change.
  • eip.md · Specification The EIP changes no transaction-validity or intrinsic-gas rule.
New block / header fields0No header member is added.
  • eip.md · Specification The EIP adds no header fields.
Encoding changes (RLP/SSZ)0No codec or schema changes.
  • eip.md · Specification The EIP changes no serialized schema.
Block syncing changes0Only execution rules change.
  • eip.md · Specification Block decoding and structural validation do not change.
New fork activation mechanism0The fork only selects new rules. No one-time state transition is needed.
  • eip.md · Backwards Compatibility — "This EIP requires a hard fork" The fork only switches rules. The EIP specifies no state migration.
Engine API changes0The Engine API contract does not change.
  • eip.md · Specification The EIP mentions no Engine API changes.
Transition-tool interface changes0No change to the transition tool's interface is needed.
  • eip.md · Specification Only opcode execution semantics change. The EIP specifies no change to inputs or outputs.
New invariant on pre-existing tests0Baseline tests need no new kind of assertion.
  • eip.md · Specification The EIP introduces no new log, receipt, header or storage output.
New test-framework primitives0Existing abstractions are enough to express the tests.
  • eip.md · Specification Testing the changed behavior only needs existing primitives: contract deployment, opcode calls and post-state balance/code/storage checks.
Performance risks0No new or larger workload needs performance validation.
  • eip.md · Motivation The change removes large state deletions. It adds no new workload.
Cryptography0No cryptographic mechanism changes.
  • eip.md · Specification The EIP involves no cryptographic primitives.
Assessment provenance
Assessed EIP revision
ethereum/EIPs@6dac5e7491 EIPS/eip-4758.md committed 2026-10-07 · information cutoff 2026-10-07T22:23:55Z
Current master · File history · blob df7433a9a0 · sha256 8e4973a1ea21
Rubric
Checklist revision 3 · ethspecs/pm@fe2f793b03
Evaluator
Opus 5.5 (claude-opus-5-5) at high effort, one tool-less call per EIP · isolation bubblewrap_claude_p_no_tools_v1
Source record
Frozen research record research/tasks/10-opus-v3-reassessment/prospective/outputs/assessments/hegota-2026-10-08/eip-4758.yaml · sha256 d211e8062d90
Supporting documents supplied with the EIP
supporting/eip-20.md

Evaluated on: Not recorded

8LowLow
Evaluator
HumanChecklist v2
Confidence
Not recorded
Under-specified at assessment cutoff
Not recorded in the checklist
Checklist published
2026-08-17
Score bands · Checklist revision 2
  • Low <12
  • Medium 12–22
  • High ≥23

28 criteria scored 0–3 (4 in exceptional cases; cross-EIP interactions is uncapped); nominal maximum 84.

Complexity profile

Each segment is one criterion's contribution to the Human total. Hover or focus a segment for its score and rationale.

Top complexity drivers

  1. Modified opcodes3
  2. Patterns affecting pre-existing tests1
  3. Security risks1
  4. Edge/boundary conditions1

Criterion breakdown

EIP-4758 Hegotá: Human criterion scores and rationale
CriterionScoreWhy this scoreNotes
Modified opcodes3Pre-existing opcode behavior modified: the deletion path is removed and `SELFDESTRUCT` becomes `SENDALL`.—
Patterns affecting pre-existing tests1EIP-6780 already narrowed deletion to same-tx-created accounts, so only that contrived family changes expectations: `tests/cancun/eip6780_selfdestruct`, `tests/amsterdam/eip8246_selfdestruct_no_burn`, CREATE2-redeploy-onto-tombstone cases. A trial removal in EELS fails only a small number of tests; the ~200 ported-static SELFDESTRUCT tests almost all use pre-existing contracts, which stopped deleting at Cancun.—
Security risks1Chain-level it removes an invariant-complicating mechanism (resurrection via CREATE2 redeploy). Breakage is application-layer (upgrade/redeploy patterns) and testable in isolation.—
Edge/boundary conditions1One genuinely new edge-prone surface: created-then-destroyed accounts now persist, so CREATE2 onto such an account hits the collision path. The other boundary cases (self-sweep, zero-balance sweep to a dead beneficiary, SENDALL in initcode, repeated SENDALL in one tx) already exist under EIP-6780/8246 and lose branches rather than gain them.—
Cross-EIP interactions1Supersedes EIP-6780's deletion gate and EIP-8246's tombstone: their suites (and those tests' EIP-7928 BAL vectors) get updated expectations, but the rework is small, mechanical, and testable independently. Interacting EIPs: 6780, 8246, 7928, 7708, 8037; no +1 increment (under six).—
Unspecified behavior requiring cross-client consensus1Stagnant two-line 2022 text predates EIP-6780/7708/8037/8246 and needs modernizing, but each open detail has an obvious intended reading: halt semantics, state-gas charging sites, and transfer-log rules carry over unchanged; only the deletion goes.—
Show 22 zero-score criteria
Zero-score criteria (Checklist revision 2)
CriterionScoreWhy this scoreNotes
Added opcodes0Rename of `0xFF`, not a new opcode.—
Added precompiles0No rationale recorded.—
Modified precompiles0No rationale recorded.—
Added system contracts0No rationale recorded.—
Modified system contracts0No rationale recorded.—
EVM Gas rule changes0Opcode gas schedule unchanged (base + cold access + account write + state gas all remain). The EIP's refund-removal clause is a no-op: SELFDESTRUCT refunds were removed by EIP-3529.—
State-access ordering within opcode execution0No access or gas-charge site inside the opcode moves. The removed behavior (account clearing) happens at end of transaction, outside opcode execution.—
Blob gas accounting changes0No rationale recorded.—
State gas accounting changes0SENDALL keeps the existing NEW_ACCOUNT / account-write charging sites; end-of-tx clearing is not gas-charged today, so removing it changes no accounting.—
New EVM gas refund0Adds none; the refunds it removes have been gone since London.—
New transaction types0No rationale recorded.—
New or modified transaction validity mechanisms0No rationale recorded.—
New block / header fields0No rationale recorded.—
Encoding changes (RLP/SSZ)0No rationale recorded.—
Block syncing changes0No rationale recorded.—
New fork activation mechanism0Pre-fork state, including EIP-8246 tombstones, is untouched at activation.—
Engine API changes0No rationale recorded.—
Transition-tool interface changes0No rationale recorded.—
New invariant on pre-existing tests0No rationale recorded.—
New test-framework primitives0No rationale recorded.—
Performance risks0Removes work (end-of-tx account clearing); no new mechanism requires performance validation.—
Cryptography0No rationale recorded.—
Assessment provenance
Rubric
Checklist revision 2 · ethspecs/pm@3d8c0128c5
Evaluator
STEEL team · ethspecs/pm complexity_assessments
Source record
Open pull request #105: Add EIP-4758 complexity assessment · checklist at 87fbde51cd · updated 2026-08-17
blob 524925b920 · sha256 5a0bbd26d962
Research record
research/tasks/09-hegota-human-assessment-snapshot/outputs/assessments/eip-4758.yaml · sha256 b4717dd137d0

The LLM applied checklist revision 3 and the human reviewers revision 2 to EIP-4758 in Hegotá. Revision 3 phrases the same criteria more precisely; differences cover the 28 criteria both revisions share, and each total keeps its own revision. Δ is LLM minus Human.

Using the latest scored LLM evaluation for this checklist: 2026-10-08 · spec 2026-10-07 · 6dac5e7491. The Human and LLM assessments may use different spec revisions.

LLM11Low
Human8Low
Δ total+3Same tier
Criteria25/28agree exactly · 3 differ by 1 · 0 differ by 2+

Complexity profiles side by side

LLM
Human

Largest disagreements: Patterns affecting pre-existing tests (+1), Unspecified behavior requiring cross-client consensus (+1), Cross-EIP interactions (+1)

Per-criterion scores, Human versus LLM, ordered by the size of the difference
CriterionLLMHumanΔAgreementRationale from each source
Patterns affecting pre-existing tests21+1Differ by 1
Show rationale

LLM Several families need their expected results reworked in specific cases. These include SELFDESTRUCT tests where the contract was created in the same transaction (including SELFDESTRUCT inside initcode and burning ETH by naming itself as beneficiary), CREATE/CREATE2 redeploy-after-destroy scenarios, and state/access-list expectations that involve destroyed accounts. Ordinary SELFDESTRUCT cases on pre-existing contracts that only send ETH are mostly unaffected in the baseline. This is localized rework in several families, which is level 2.

Human EIP-6780 already narrowed deletion to same-tx-created accounts, so only that contrived family changes expectations: `tests/cancun/eip6780_selfdestruct`, `tests/amsterdam/eip8246_selfdestruct_no_burn`, CREATE2-redeploy-onto-tombstone cases. A trial removal in EELS fails only a small number of tests; the ~200 ported-static SELFDESTRUCT tests almost all use pre-existing contracts, which stopped deleting at Cancun.

Cross-EIP interactions21+1Differ by 1
Show rationale

LLM Coordinated cases with CREATE2 are needed: a contract created with CREATE2, then SENDALL, then redeployed with CREATE2 at the same address now hits an address collision because code and nonce persist. Cases are also needed for SELFDESTRUCT inside CREATE/CREATE2 initcode. EIP-20 is only cited as an application example, so it does not qualify as an interaction.

Human Supersedes EIP-6780's deletion gate and EIP-8246's tombstone: their suites (and those tests' EIP-7928 BAL vectors) get updated expectations, but the rework is small, mechanical, and testable independently. Interacting EIPs: 6780, 8246, 7928, 7708, 8037; no +1 increment (under six).

Unspecified behavior requiring cross-client consensus21+1Differ by 1
Show rationale

LLM Several local outcomes have competing readings. The recipient is the caller in the Abstract but the target in the Specification. The EIP does not say whether execution halts after SENDALL. It also does not say what happens when target==self, where the baseline same-transaction case burns ETH. Expected test results cannot be fixed until these are agreed, which is level 2.

Human Stagnant two-line 2022 text predates EIP-6780/7708/8037/8246 and needs modernizing, but each open detail has an obvious intended reading: halt semantics, state-gas charging sites, and transfer-log rules carry over unchanged; only the deletion goes.

Added opcodes000Agree
Show rationale

LLM Replacing an existing instruction's semantics is scored under Modified opcodes.

Human Rename of `0xFF`, not a new opcode.

Modified opcodes330Agree
Show rationale

LLM The state effects of SELFDESTRUCT change: it no longer deletes code, storage or the nonce, even for contracts created in the same transaction. This is a semantic change beyond gas or ordering, which is level 3.

Human Pre-existing opcode behavior modified: the deletion path is removed and `SELFDESTRUCT` becomes `SENDALL`.

Added precompiles000Agree
Show rationale

LLM No precompile is introduced.

Human No rationale recorded.

Modified precompiles000Agree
Show rationale

LLM No precompile changes.

Human No rationale recorded.

Added system contracts000Agree
Show rationale

LLM No system contract is added.

Human No rationale recorded.

Modified system contracts000Agree
Show rationale

LLM No system contract's rules change.

Human No rationale recorded.

EVM Gas rule changes000Agree
Show rationale

LLM In the Amsterdam baseline, SELFDESTRUCT already gives no gas refund (general protocol knowledge about the base fork), so removing the refunds changes nothing there. The EIP changes no opcode cost, charging rule or limit, which fits level 0.

Human Opcode gas schedule unchanged (base + cold access + account write + state gas all remain). The EIP's refund-removal clause is a no-op: SELFDESTRUCT refunds were removed by EIP-3529.

State-access ordering within opcode execution000Agree
Show rationale

LLM The opcode touches the same accounts as before, and the EIP does not change the order of its state accesses or gas charges. This is level 0.

Human No access or gas-charge site inside the opcode moves. The removed behavior (account clearing) happens at end of transaction, outside opcode execution.

Blob gas accounting changes000Agree
Show rationale

LLM The EIP does not change blob-gas accounting.

Human No rationale recorded.

State gas accounting changes000Agree
Show rationale

LLM The EIP adds no state-gas charging and changes no state-gas parameters.

Human SENDALL keeps the existing NEW_ACCOUNT / account-write charging sites; end-of-tx clearing is not gas-charged today, so removing it changes no accounting.

New EVM gas refund000Agree
Show rationale

LLM Removing a refund belongs under the gas criterion, not here. The EIP adds no refund mechanism.

Human Adds none; the refunds it removes have been gone since London.

New transaction types000Agree
Show rationale

LLM No new transaction envelope.

Human No rationale recorded.

New or modified transaction validity mechanisms000Agree
Show rationale

LLM Transaction validity does not change.

Human No rationale recorded.

New block / header fields000Agree
Show rationale

LLM No header member is added.

Human No rationale recorded.

Encoding changes (RLP/SSZ)000Agree
Show rationale

LLM No codec or schema changes.

Human No rationale recorded.

Block syncing changes000Agree
Show rationale

LLM Only execution rules change.

Human No rationale recorded.

New fork activation mechanism000Agree
Show rationale

LLM The fork only selects new rules. No one-time state transition is needed.

Human Pre-fork state, including EIP-8246 tombstones, is untouched at activation.

Engine API changes000Agree
Show rationale

LLM The Engine API contract does not change.

Human No rationale recorded.

Transition-tool interface changes000Agree
Show rationale

LLM No change to the transition tool's interface is needed.

Human No rationale recorded.

New invariant on pre-existing tests000Agree
Show rationale

LLM Baseline tests need no new kind of assertion.

Human No rationale recorded.

New test-framework primitives000Agree
Show rationale

LLM Existing abstractions are enough to express the tests.

Human No rationale recorded.

Security risks110Agree
Show rationale

LLM The changed security conditions can be checked locally. Tests need to show that a contract stays callable after SENDALL, that code and storage persist, and that redeploying at the same address collides. No cross-component trust invariant changes, so this is level 1.

Human Chain-level it removes an invariant-complicating mechanism (resurrection via CREATE2 redeploy). Breakage is application-layer (upgrade/redeploy patterns) and testable in isolation.

Performance risks000Agree
Show rationale

LLM No new or larger workload needs performance validation.

Human Removes work (end-of-tx account clearing); no new mechanism requires performance validation.

Edge/boundary conditions110Agree
Show rationale

LLM One changed mechanism, the SENDALL semantics, has boundary-sensitive outcomes: same-transaction creation versus a pre-existing account, beneficiary equal to self, zero balance, and initcode context. These are all parts of one rule, which is level 1.

Human One genuinely new edge-prone surface: created-then-destroyed accounts now persist, so CREATE2 onto such an account hits the collision path. The other boundary cases (self-sweep, zero-balance sweep to a dead beneficiary, SENDALL in initcode, repeated SENDALL in one tx) already exist under EIP-6780/8246 and lose branches rather than gain them.

Cryptography000Agree
Show rationale

LLM No cryptographic mechanism changes.

Human No rationale recorded.

Criterion legend and glossary

Every stacked bar, comparison matrix, and criterion table on this site uses the same criterion colours, abbreviations, and order. Colour marks the criterion group; the abbreviation and name identify the criterion. Scores are 0–3 per criterion (4 is exceptional; cross-EIP interactions is uncapped).

EVM surface

Opcodes, precompiles, and system contracts that are added or modified.

  • Added opcodes
    Introduces new opcodes
    Score anchors
    0
    No new opcodes are introduced.
    1
    A new simple opcode is introduced (no data portion, no complex stack mechanics, and a constant gas cost).
    2
    Multiple new simple opcodes are introduced, or a single new complex opcode is introduced (has data portion, or complex stack mechanics, or a dynamic gas cost).
    3
    Multiple new opcodes are introduced, and at least one of them is complex (has data portion, or complex stack mechanics, or a dynamic gas cost).
    • Cryptography opcodes are not considered complex by default. Refer to the "Cryptography" section for a separate assessment.
  • Modified opcodes
    Modifies pre-existing opcodes
    Score anchors
    0
    No pre-existing opcode modifications are introduced.
    3
    At least one pre-existing opcode's behavior is modified (not including gas changes) or a pre-existing opcode is deprecated.
  • Added precompiles
    Introduces new precompiles
    Score anchors
    0
    No new precompiles are introduced.
    1
    A new simple precompile is introduced (constant input length, constant gas cost).
    2
    Multiple new simple precompiles are introduced, or a single new complex precompile is introduced (dynamic input length or dynamic gas cost).
    3
    Multiple new precompiles are introduced, and at least one of them is complex (dynamic input length or dynamic gas cost).
    • Cryptography precompiles are not considered complex by default. Refer to the "Cryptography" for a separate assessment.
  • Modified precompiles
    Modifies pre-existing precompiles logic or gas-accounting
    Score anchors
    0
    No pre-existing precompiles are modified.
    1
    At least one pre-existing precompile has its gas schedule modified.
    2
    Multiple pre-existing precompiles have their gas schedule modified, or a single pre-existing precompile has its behavior modified.
    3
    The behavior of multiple pre-existing precompiles, or a single complex pre-existing precompile modified.
  • Added system contracts
    Introduces new system contract, stateful or not
    Score anchors
    0
    No new system contracts are introduced.
    1
    A new system contract is introduced that is not stateful nor does it trigger a new system action (e.g. requests to the consensus layer).
    2
    Multiple new system contracts are introduced or a single new system contract that is either stateful or triggers a new system action (e.g. requests to the consensus layer).
    3
    Multiple new system contracts are introduced and at least one of them is either stateful or triggers a new system action (e.g. requests to the consensus layer).
  • Modified system contracts
    Modifies pre-existing system contracts
    Score anchors
    0
    No modifications to pre-existing system contracts are introduced, directly or indirectly.
    1
    Does not directly modify any system contract, but its behavior has minor indirect effects on one or more system contracts.
    2
    Does not directly modify any system contract, but its behavior has major indirect effects on one or more system contracts.
    3
    At least one pre-existing system contract code or state is modified, which would involve irregular state transition or a similarly complex transition methodology.

Gas and accounting

Execution, blob, and state gas rules, refunds, and where charges happen inside opcodes.

  • EVM Gas rule changes
    New EVM gas accounting rules
    Score anchors
    0
    No gas accounting changes.
    1
    Existing gas accounting mechanism is updated.
    2
    A new gas accounting mechanism is introduced but it does not affect existing mechanisms nor does it affect existing tests.
    3
    A new gas accounting mechanism is introduced and affects existing mechanisms which in turn affect existing tests.
  • State-access ordering within opcode execution · not in checklist revision 1
    Changes *where inside an opcode's execution* state is accessed, or where gas is charged relative to that access. Because a state access is recorded in the block-level access list only if execution had enough gas to reach it, this ordering is consensus-critical: moving it changes the BAL at every gas boundary of every affected opcode.
    Score anchors
    0
    No change to where state is accessed, or to where gas is charged relative to a state access, within any opcode.
    1
    A single opcode's state-access or gas-charge ordering changes.
    2
    Multiple opcodes' ordering changes, or a new state-accessing operation is introduced whose position in the order must be settled.
    3
    The ordering rule changes for a whole class of state-accessing opcodes at once, or what counts as a recordable state access is redefined — requiring existing BAL vectors to be re-derived across opcodes and forks.
    • Distinct from "Modified opcodes", which asks whether an opcode's **result** changed. This row asks about the **path to the result**, which is observable even when the result is identical. An EIP can be 0 on that row and 3 on this one.
    • Score changes **to** the ordering. Do not score the fact that state accesses are observable — they always are.
    • Each boundary must be re-tested against every other dimension that can change the answer (cold/warm, static/non-static, delegated/direct, revert/success), so the case count grows multiplicatively rather than additively. Note this explicitly under Special Considerations.
  • Blob gas accounting changes
    New Blob gas accounting rules which potentially affect pre-existing tests
    Score anchors
    0
    No blob gas accounting changes.
    1
    Existing blob gas accounting mechanism is updated.
    2
    A new blob gas accounting mechanism is introduced but it does not affect existing mechanisms nor does it affect existing tests.
    3
    A new blob gas accounting mechanism is introduced and affects existing mechanisms which in turn affect existing tests.
  • State gas accounting changes · not in checklist revision 1
    New state gas accounting rules. State gas is the cost of *writing* state, as opposed to accessing or executing it: `StateGasCosts`, `COST_PER_STATE_BYTE`, the block-level state gas budget, and the spill path into execution gas.
    Score anchors
    0
    No state gas accounting changes.
    1
    An existing state gas cost or `STATE_BYTES_PER_*` rate is adjusted.
    2
    A new state-gas-charging site is introduced, or the block-level state gas budget or reservoir allocation is modified.
    3
    A new state gas charging mechanism is introduced, or the spill interaction between state gas and execution gas is modified, affecting existing gas tests.
    • Harder to test than blob gas: the spill path means state gas cannot be metered independently of execution gas, and some costs (e.g. `NEW_ACCOUNT`) are state-dependent.
  • New EVM gas refund
    New gas-refund mechanism
    Score anchors
    0
    No new gas-refund mechanisms are introduced.
    1
    A new simple gas-refund mechanism is introduced that does not affect either existing tests or existing gas-refund mechanisms.
    2
    A new complex gas-refund mechanism is introduced or a simple mechanism that affects existing tests or existing gas-refund mechanisms.
    3
    A new complex gas-refund mechanism is introduced that affects existing tests or existing gas-refund mechanisms.

Blocks, transactions, and encoding

Transaction types and validity, block and header fields, encodings, syncing, and activation-time changes.

  • New transaction types
    Introduces a new transaction type
    Score anchors
    0
    No new transaction types are introduced.
    3
    A new transaction type is introduced.
  • New or modified transaction validity mechanisms
    Creates new or modifies pre-existing transaction types' validation mechanisms
    Score anchors
    0
    No changes are introduced to the validity rules of existing transaction types or to their intrinsic gas cost calculation.
    1
    Minor adjustments are introduced to validity rules or intrinsic gas cost calculation, but they do not significantly affect existing tests.
    2
    Changes to validity rules or intrinsic gas cost calculation affect existing tests, but require only limited updates to test cases and no redesign of the testing infrastructure.
    3
    Changes to validity rules or intrinsic gas cost calculation require extensive rework or redesign of the tests or testing infrastructure.
  • New block / header fields
    Introduces new block or block header fields
    Score anchors
    0
    No new block or header fields are introduced.
    3
    A new block or header field is introduced.
  • Encoding changes (RLP/SSZ)
    Introduces encoding changes at the transaction/block/interfaces level
    Score anchors
    0
    No encoding changes are introduced at the transaction, block, or interfaces levels.
    3
    An encoding change is introduced at transaction, block or interfaces level (e.g. RLP -> SSZ).
    • "Interfaces level" includes the Engine API. Score an Engine API encoding change (e.g. JSON -> SSZ) here.
  • Block syncing changes
    Modifies block RLP validation mechanisms that require test client syncing.
    Score anchors
    0
    No new RLP validation mechanism is introduced.
    1
    A single simple RLP validation mechanism is introduced.
    2
    Multiple simple RLP validation mechanisms are introduced or a single complex one.
    3
    Multiple RLP validation mechanisms are introduced and at least one of them is deemed complex.
  • New fork activation mechanism
    Modifies state, internal variables, or similar, at the fork activation block
    Score anchors
    0
    No state modifications, internal variables or similar are modified at the fork activation block.
    3
    Either a state modification or internal variables are modified at the fork activation block.
    • Initialization of new internal variable is not considered a modification.

Client interfaces

Engine API and transition-tool interface changes.

  • Engine API changes
    Introduces new fields to the Engine API directives
    Score anchors
    0
    No new fields or communication mechanisms are introduced to the Engine API.
    1
    A single new field is introduced in one of the Engine API endpoints.
    2
    Multiple fields are introduced to one or multiple Engine API end points, or a new Engine API end-point is introduced.
    3
    Multiple fields are introduced to one or multiple Engine API end points and a new Engine API end-point is introduced.
  • Engine API encoding changes · Checklist revision 1 only
    Engine API encoding changes (the revision-1 template defines no anchor text for this row).
  • Transition-tool interface changes
    Modifies or adds new fields to the transition tool interface.
    Score anchors
    0
    No modifications to the transition tool interface are required.
    1
    A single new field needs to be introduced to the transition tool interface.
    2
    Multiple new fields or a new mechanism has to be introduced to the transition tool interface.
    3
    Multiple new fields and a new mechanism has to be introduced to the transition tool interface.
    • Special consideration must be paid to this section if the EIP introduces a mechanism that requires the state transition tool to be aware whether the block it is processing is the fork-activation block.

Testing impact

Rework, new invariants, and new primitives required in the test framework.

  • Patterns affecting pre-existing tests
    Implements a new validation mechanism or rule that translates in reworking pre-existing tests
    Score anchors
    0
    No pre-existing tests are affected by this change.
    1
    Minor subset of existing tests are affected by this change.
    2
    Considerable subset of existing tests are affected by this change but involves only a contrived category of tests.
    3
    Major subset of existing tests are affected, including diverse category of tests (benchmarks, static, multiple forks, etc.).
  • New invariant on pre-existing tests · not in checklist revision 1
    Tests that are **not about this EIP** must nonetheless assert something this EIP produces. Their logic does not change; they gain a new thing to check.
    Score anchors
    0
    Pre-existing tests assert nothing new.
    1
    A narrow, contrived category of pre-existing tests gains a new assertion.
    2
    A broad category gains a new assertion, applied mechanically.
    3
    Every test in the fork gains the assertion regardless of what it tests, and pre-fork vectors must be re-derived to satisfy it.
    • Paired with the row above, and easy to confuse with it. "Patterns affecting pre-existing tests" asks whether existing tests must be **reworked**; this row asks whether they must **additionally assert something new**. Score both — an EIP can be low on one and high on the other.
  • New test-framework primitives · not in checklist revision 1
    Requires new abstractions in the test framework itself — expectation types, modifiers, helpers — beyond writing test functions with what already exists.
    Score anchors
    0
    Existing test primitives suffice.
    1
    Existing primitives need minor extension.
    2
    New expectation or modifier primitives are required, reusable within this EIP's own test suite.
    3
    New framework-level primitives are required that become a permanent part of the framework and are used by other EIPs' tests.

Risk and validation

Security, performance, boundary conditions, and cryptography that need validation.

  • Security risks
    Introduces or modifies mechanisms that could compromise the security of the chain, users, validators, or other stakeholders, if not implemented properly.
    Score anchors
    0
    No new mechanisms are introduced that could pose a security risk.
    1
    The introduced mechanisms are self-contained, can be validated in isolation, and do not alter existing invariants that could pose a security risk for any stakeholders.
    2
    The introduced mechanisms interact with a limited number of existing components, slightly altering their security assumptions and requiring a targeted security review or fuzzing.
    3
    The introduced mechanisms interact with multiple existing components, including critical ones, substantially altering their security assumptions and requiring an extensive security review and fuzzing.
  • Performance risks
    Introduces or modifies mechanisms and requires performance validation.
    Score anchors
    0
    No new mechanisms are introduced that require performance validation.
    1
    The introduced mechanisms can be benchmarked in isolation and do not affect existing performance behavior.
    2
    The introduced mechanisms cannot be fully benchmarked in isolation, but they only have a limited impact on the existing performance benchmarks.
    3
    The introduced mechanisms cannot be benchmarked in isolation and have a substantial impact on existing performance benchmarks or have complex interactions with existing mechanisms.
  • Edge/boundary conditions
    Feature contains edge/boundary conditions.
    Score anchors
    0
    No discernible edge cases or boundary conditions are introduced.
    1
    A single edge-case or boundary-condition prone mechanism is introduced.
    2
    Multiple edge-case or boundary-condition prone mechanisms are introduced, but none of them requires an elevated number of cases to test.
    3
    Multiple edge-case or boundary-condition prone mechanisms are introduced and at least one of them requires an elevated number of cases to test.
  • Cryptography
    Introduces new cryptography mechanisms or modifies existing functionality that involves cryptography
    Score anchors
    0
    No cryptography mechanisms are introduced.
    1
    A new cryptography mechanism is introduced but it is a well known mechanism that is known to have vast resources to aid on its testing.
    2
    Multiple new cryptography mechanisms are introduced that are well-known or a single but novel mechanism is introduced that is either untested or has limited resources.
    3
    Multiple new cryptography mechanisms are introduced and at least one of them is a novel mechanism.

Coordination

Cross-EIP interactions and behavior that clients must agree on before tests exist.

  • Cross-EIP interactions
    Introduces or modifies mechanisms that affect other EIPs in either the same or past forks.
    Score anchors
    0
    Fully self-contained EIP that does not depend on, modify, or conflict with any other EIP.
    1
    The EIP interacts with one or more other EIPs in a non-critical and limited way but can be tested independently for the most part.
    2
    The EIP depends on or modifies one or more other EIPs such that coordinated testing and consideration is required, but interactions are limited in scope and not complex.
    3
    The EIP has strong interdependencies with multiple EIPs, requiring extensive coordinated cross-EIP testing as well as potential re-design of existing test vectors.
    • +1 for every 3 additional interacting EIPs beyond the first 3, each of which requires its own coordinated test cases. List the EIPs in the rationale.
    • This row is intentionally uncapped, unlike every other anchor: each interacting EIP is another axis of the test matrix, so a ceiling would make a 12-EIP product indistinguishable from a 3-EIP one.
  • Unspecified behavior requiring cross-client consensus · not in checklist revision 1
    The EIP text does not determine the answer for cases a test can construct. Clients must agree on a previously unspecified detail before tests can be baselined. The cost here is coordination and re-baselining, not test writing.
    Score anchors
    0
    The EIP text determines the answer for every case a test could construct.
    1
    A few details are unspecified but have an obvious intended reading.
    2
    Details require client agreement before tests can be written, but they are localized.
    3
    A previously unspecified *and previously unobservable* behavior becomes consensus-critical; expect tests to be re-baselined on each round of EIP amendment.
    • Score this from the EIP's state at assessment time: whether it has client implementations, whether it has been through a devnet, and how many open questions remain on its discussion thread.