Retrospective LLM-Based Complexity Evaluations

EIP complexity assessment

EIP-7708: ETH transfers emit a log

Assessed in Amsterdam / Glamsterdam. The score describes the EIP text available at the assessment cutoff, not the EIP as it stands today.

RetrospectiveAmsterdam / GlamsterdamAssessment cutoff 2025-10-22Included by cutoffLayers: execution
LLM Completescore 17
Human Completescore 9 · Checklist revision 1· merged checklist

Evaluated on: · Spec revision: 2025-04-13 · 7466599174

Scope at the cutoff. At this revision, EIP-7708 makes every nonzero-value ETH transfer emit a log. That covers three cases: a nonzero-value CALL, a SELFDESTRUCT that transfers nonzero value, and a nonzero-value transaction. Each log is "identical to a LOG3". Its three topics are MAGIC (still TBD), the sender address and the recipient address. Its data is the 32-byte big-endian transfer value. The transaction-level log comes before any logs from EVM execution. CALL and SELFDESTRUCT logs are emitted when the transfer executes. Whether withdrawals and fee payments should also emit logs is left as an open question. No gas cost, emitter address or revert handling is specified for the logs.

17MediumMedium
Evaluator
LLMChecklist v3
Confidence
Medium
Under-specified at assessment cutoff
Yes — 8 criteria affected
Plausible range
14–23 (Medium–High)
Assessment cutoff
2025-10-22 · EIP revision 7466599174 (2025-04-13)
Score bands · Checklist revision 3
  • Low <12
  • Medium 12–22
  • High ≥23

28 criteria scored 0–3 (4 in exceptional cases; cross-EIP interactions is uncapped); nominal maximum 84.

Complexity profile

Each segment is one criterion's contribution to the LLM total. Hover or focus a segment for its score and rationale.

Top complexity drivers

  1. Modified opcodes3
  2. Patterns affecting pre-existing tests2
  3. New invariant on pre-existing tests2
  4. New test-framework primitives2

Under-specified at assessment cutoff: Yes

The EIP text available at the assessment cutoff left material behavior unresolved. The affected criteria and the plausible total range record that uncertainty.

Why: MAGIC is TBD. The log's emitter address is unspecified, as are gas charging, revert handling, and coverage of CALLCODE, CREATE/CREATE2 endowments, contract-creation transactions, SELFDESTRUCT-to-self and failed value calls. Withdrawals and fees are explicit open questions. There is no test-case section.

Unresolved questions at the cutoff (8)
  • What is the MAGIC topic value?
  • What is the log's address field, both for in-EVM transfers and for the transaction-level log?
  • Is any gas charged for the emitted log?
  • Are transfer logs from reverted frames discarded like ordinary LOG3 logs?
  • Do CALLCODE, CREATE/CREATE2 with value, and contract-creation transactions emit logs?
  • Does a value CALL that fails (insufficient balance, depth limit) emit a log?
  • Does SELFDESTRUCT to itself, or a burn, emit a log, and with what recipient?
  • Should withdrawals and fee payments emit logs?
Notable ambiguities noted by the assessor (5)
  • The abstract says all ETH transfers emit a log, but the specification lists only CALL, SELFDESTRUCT and transactions.
  • "identical to a LOG3" does not say which address emits the log, especially for the transaction-level log.
  • MAGIC is TBD, so the expected log contents cannot be fixed.
  • "nonzero-value CALL" could mean the requested value or an executed transfer, which matters for calls that fail before transferring.
  • Gas charging for the implicit log is never stated.

Criterion breakdown

EIP-7708 Amsterdam / Glamsterdam: LLM criterion scores and rationale
CriterionScoreWhy this scoreEvidence / uncertainty
Modified opcodes3Emitted logs are part of instruction semantics, and the logs emitted by CALL and SELFDESTRUCT change.
  • eip.md · Specification / Functionality - "a nonzero-value CALL, (ii) a nonzero-value-transferring SELFDESTRUCT" CALL and SELFDESTRUCT now emit logs when they transfer value.
Confidence: High
Uncertainty: It is unclear whether CALLCODE and CREATE/CREATE2 with an endowment are also covered.
Patterns affecting pre-existing testsUnder-specified2Baseline tests that assert log lists or log indices for value-transferring transactions or calls need revised expected logs. Examples are LOG-opcode tests in transactions with value, and CALL or SELFDESTRUCT tests with value. Their logs hash, bloom and receipt root also change. This rework covers ordinary cases in several log-checking families. It mostly means regenerating expectations rather than a common restructuring, so level 2.
  • eip.md · Specification / Functionality - "The LOG of a value-transferring transaction should be placed before any logs created by EVM execution" Existing logs in value-transferring transactions move to later positions, and the logs, bloom and receipt commitments change.
  • eip.md · Abstract All ETH transfers, including transactions, CALL and SELFDESTRUCT, emit a log.
Confidence: Medium
Uncertainty: Logs-hash and receipt-root changes affect nearly every value-transferring fixture. Whether that counts as a common cross-family rewrite (level 3) or as regeneration depends on how the tests are written.
New invariant on pre-existing tests2Baseline tests across many families (plain transfers, value calls, SELFDESTRUCT, contract interactions) must now check the new transfer log. Zero-value tests and pre-fork vectors are unaffected, so the requirement is neither universal nor re-derived for earlier forks.
  • eip.md · Specification / Functionality Every nonzero-value transaction, CALL or SELFDESTRUCT produces a new log.
Confidence: High
New test-framework primitivesUnder-specified2The tests need a new expectation helper that builds the expected transfer log from a transfer and places it correctly among the other logs (transaction log first; call logs inline). That is a new expectation abstraction within the target's suite.
  • eip.md · Specification / Functionality The log's topics and data are derived from the transfer (MAGIC, sender, recipient, value).
Confidence: Medium
Uncertainty: If the framework had to inject expected transfer logs into all existing log-checking families automatically, this could reach level 3. A simple log builder could be level 1.
Edge/boundary conditionsUnder-specified2Several boundary-sensitive mechanisms are introduced. These are the nonzero-value triggers for transactions, CALL and SELFDESTRUCT; whether the transfer actually executes (insufficient balance, call depth, OOG); log ordering; and survival of logs in reverted frames. They can mostly be tested independently.
  • eip.md · Specification / Functionality - "nonzero-value" The log is emitted only when value is nonzero (the 0 vs 1 wei boundary) in three separate trigger sources.
  • eip.md · Specification / Functionality - "placed before any logs created by EVM execution" Ordering rules for the transaction log and in-execution logs.
Confidence: Medium
Uncertainty: An elevated matrix is possible: value × call success × enclosing-frame revert × SELFDESTRUCT beneficiary self/other. That could justify level 3.
Cross-EIP interactionsUnder-specified2The SELFDESTRUCT trigger needs coordinated cases against the baseline (Osaka) SELFDESTRUCT rules: same-transaction-created vs pre-existing contracts, beneficiary equal to self (burn), and balance-only transfer. ERC-20 needs only a local check that the log layout is compatible.
  • eip.md · Rationale - "The log type is compatible with the ERC-20 token standard" The log layout is intended to match the shape of the ERC-20 Transfer event.
  • supporting/eip-20.md · front matter Only a moved-file stub, so the ERC-20 event definition is not supplied.
  • eip.md · Specification / Functionality - "nonzero-value-transferring SELFDESTRUCT" Interacts with the baseline SELFDESTRUCT semantics, defined elsewhere.
Confidence: Medium
Uncertainty: The SELFDESTRUCT rules come from an EIP that is not supplied or named. The ERC-20 definition is not supplied either.
Interacting EIPs: EIP-20
Unspecified behavior requiring cross-client consensusUnder-specified2Several consensus-visible outcomes (receipts and bloom) have competing interpretations: the MAGIC value, the log address field, gas charging, coverage of CALLCODE and CREATE, SELFDESTRUCT to self or burn, failed value calls, revert handling, and withdrawals and fees. Clients and the spec must agree on these before expected results can be fixed.
  • eip.md · Specification / Parameters - "MAGIC: TBD" The first topic of every transfer log is undefined.
  • eip.md · Specification / Functionality - "identical to a LOG3" The log's emitter address field is not specified, especially for the transaction-level log.
  • eip.md · Abstract - "All ETH-transfers" The abstract claims all transfers, but the functionality section lists only CALL, SELFDESTRUCT and transactions. CALLCODE, CREATE/CREATE2 endowments and creation transactions are ambiguous.
  • eip.md · Rationale / Open questions Logging for withdrawals and fee payments is unresolved.
Confidence: Medium
Uncertainty: The emitter address and MAGIC apply to every transfer log. Depending on interpretation, that could justify level 3.
Security risksUnder-specified1The new security conditions are local: logs must be emitted exactly when a transfer happens, logs from reverted frames must not survive, and genuine transfer logs should be distinguishable from contract-emitted LOG3 forgeries. No other EL component's assumptions change.
  • eip.md · Motivation Exchanges and indexers would rely on these logs to track ETH deposits.
  • eip.md · Specification / Parameters - "MAGIC: TBD" The topic is unspecified, so it is undefined whether contracts could forge identical logs with LOG3.
Confidence: Medium
Uncertainty: Without a defined emitter address, forgeability affects off-chain consumers. That could warrant targeted review (level 2).
Performance risksUnder-specified1Receipt and bloom generation and log storage grow on average. Component benchmarks of log-heavy value-transfer blocks are enough, and end-to-end assumptions stay bounded according to the spec.
  • eip.md · Security Considerations - "does not increase the worst-case number of logs" / "will somewhat increase the average number of logs" The worst case stays bounded by the 6700-gas minimum transfer cost. Average log volume rises.
Confidence: Low
Uncertainty: If the logs are free and the 6700 figure does not apply in some paths, worst-case log volume per gas could need validation.
Show 19 zero-score criteria
Zero-score criteria (Checklist revision 3)
CriterionScoreWhy this scoreEvidence / uncertainty
Added opcodes0No new opcodes.
  • eip.md · Specification No new instruction is introduced. The log is emitted implicitly.
Added precompiles0None added.
  • eip.md · Specification No precompile is defined.
Modified precompiles0No precompile semantics or gas change.
  • eip.md · Specification Precompile semantics are unchanged. A value CALL to a precompile emits a log through the CALL rule.
Added system contracts0None added.
  • eip.md · Specification No system contract is defined.
Modified system contracts0No existing system contract's rules or surrounding behavior change.
  • eip.md · Rationale / Open questions - "Should withdrawals also trigger a log?" Withdrawals and fees are left open, and no system contract behavior is changed.
Uncertainty: It is unspecified whether system calls carrying value would emit logs, though baseline system calls are normally zero-value.
EVM Gas rule changesUnder-specified0No execution-gas accounting rule is specified as changing. The log appears to be emitted without its own gas charge.
  • eip.md · Specification / Functionality Says to issue a log identical to a LOG3, but gives no gas charge for it.
  • eip.md · Security Considerations - "much more expensive than the LOG3 opcode (1500 gas)" Compares transfer cost with LOG3 cost to bound the number of logs. This implies the log is not separately charged, but the text does not say so.
Uncertainty: The text never says whether LOG3-equivalent gas is charged. If it were, the gas results of value-transferring CALL, SELFDESTRUCT and transactions would change (level 1).
State-access ordering within opcode execution0Emitting a log does not access state. The order of state access and gas charging in CALL and SELFDESTRUCT is unchanged, and no access-list rule is introduced.
  • eip.md · Specification / Functionality - "placed at the time that the value transfer executes" Sets where the log is placed, not the order of state access or gas charging.
Blob gas accounting changes0Blob gas is unaffected.
  • eip.md · Specification Nothing about blobs or blob gas.
State gas accounting changes0There are no state-gas accounting changes.
  • eip.md · Specification Logs are not state writes, and no state-gas accounting is defined.
New EVM gas refund0No refund is introduced.
  • eip.md · Specification No refund mechanism is mentioned.
New transaction types0None.
  • eip.md · Specification No new transaction type.
New or modified transaction validity mechanisms0Transaction validity is unchanged.
  • eip.md · Specification No validity or intrinsic-gas change.
New block / header fields0Changed values in existing fields do not count.
  • eip.md · Specification No header field is added. The existing receiptsRoot and logsBloom values change.
Encoding changes (RLP/SSZ)0New values within an unchanged receipt and log schema do not count.
  • eip.md · Specification / Functionality - "identical to a LOG3" Reuses the existing log and receipt schema.
Block syncing changes0Only receipt contents change, through ordinary execution. Structural validation is unchanged.
  • eip.md · Specification No change to block structure or decoding.
New fork activation mechanism0No activation-specific transition.
  • eip.md · Specification Only a rule change, with no state migration.
Engine API changes0No Engine API fields or endpoints change.
  • eip.md · Specification No Engine API change.
Transition-tool interface changes0Logs are already a t8n output. No new input, output field or mechanism is needed.
  • eip.md · Specification / Functionality - "identical to a LOG3" Uses the existing log format, which is already reported in receipts.
Uncertainty: No transition-tool evidence was supplied.
Cryptography0No cryptographic change.
  • eip.md · Specification No cryptographic rules are involved. Bloom and receipt hashing are unchanged.
Assessment provenance
Assessed EIP revision
ethereum/EIPs@7466599174 EIPS/eip-7708.md committed 2025-04-13 · information cutoff 2025-10-22T20:52:23Z
Current master · File history · blob a0758eccab · sha256 87520fb55cfb
Rubric
Checklist revision 3 · ethspecs/pm@fe2f793b03
Evaluator
Opus 5.5 (claude-opus-5-5) at high effort, one tool-less call per EIP · isolation bubblewrap_claude_p_no_tools_v1
Source record
Frozen research record research/tasks/10-opus-v3-reassessment/retrospective/outputs/assessments/amsterdam/eip-7708.yaml · sha256 6e6720ac90af
Supporting documents supplied with the EIP
supporting/eip-20.md

Evaluated on: Not recorded · Spec revision: 2025-04-13 · 7466599174

9LowLow
Evaluator
HumanChecklist v1
Confidence
Not recorded
Under-specified at assessment cutoff
Not recorded in the checklist
Checklist published
2025-11-05 · EIP at 7466599174
Score bands · Checklist revision 1
  • Low <10
  • Medium 10–19
  • High ≥20

24 criteria scored 0–3 (4 in exceptional cases; cross-EIP interactions is uncapped); nominal maximum 72.

Complexity profile

Each segment is one criterion's contribution to the Human total. Hover or focus a segment for its score and rationale.

Top complexity drivers

  1. Modified system contracts2
  2. Performance risks2
  3. Edge/boundary conditions2
  4. Modified opcodes1

Criterion breakdown

EIP-7708 Amsterdam / Glamsterdam: Human criterion scores and rationale
CriterionScoreWhy this scoreNotes
Modified system contracts2Some system contracts do receive/transfer Eth, and while existing tests might be sufficient, we need to double check that this change does not break current behavior (particularly the beacon deposit contract).—
Performance risks2Test worst-case scenario of a transaction/block transfering Eth in multiple ways.—
Edge/boundary conditions2All 0->1 boundaries have to be tested in all possible scenarios where Eth is transfered.—
Modified opcodes1*CALL opcodes and SELFDESTRUCT/SENDALL behavior is slightly modified.—
Transition-tool interface changes1Rather than a change to the T8N interface, we have to enhance execution-specs testing framework to verify the logs returned from T8N, because at the moment we just pass along the log list that was received from T8N to the test to be compared against the clients.—
Patterns affecting pre-existing tests1All tests will now issue Eth-transfer logs, which might be a non-issue but does touch basically every single test.—
Show 18 zero-score criteria
Zero-score criteria (Checklist revision 1)
CriterionScoreWhy this scoreNotes
Added opcodes0No rationale recorded.—
Added precompiles0No rationale recorded.—
Modified precompiles0No rationale recorded.—
Added system contracts0No rationale recorded.—
EVM Gas rule changes0No rationale recorded.—
Blob gas accounting changes0No rationale recorded.—
New EVM gas refund0No rationale recorded.—
New transaction types0No rationale recorded.—
New or modified transaction validity mechanisms0No rationale recorded.—
New block / header fields0No rationale recorded.—
Encoding changes (RLP/SSZ)0No rationale recorded.—
Block syncing changes0No rationale recorded.—
New fork activation mechanism0No rationale recorded.—
Engine API changes0No rationale recorded.—
Engine API encoding changes0No rationale recorded.—
Security risks0No rationale recorded.—
Cryptography0No rationale recorded.—
Cross-EIP interactions0No rationale recorded.—
Assessment provenance
Assessed EIP revision
ethereum/EIPs@7466599174 EIPS/eip-7708.md committed 2025-04-13
EIP master revision at the time of the human checklist commit; the checklist itself does not pin an EIP revision.
Current master · File history · blob a0758eccab · sha256 87520fb55cfb
Rubric
Checklist revision 1 · ethspecs/pm@d936bcb349
Evaluator
STEEL team · ethspecs/pm complexity_assessments
Source record
Merged checklist ethspecs/pm@72a1fc653d complexity_assessments/EIPs/EIP-7708.md · committed 2025-11-05
blob 092e5bab2a · sha256 a2652b18c2cf
Research record
research/tasks/05c-amsterdam-human-assessment-alignment/inputs/human/eip-7708.yaml · sha256 4e334dac152a

The LLM applied checklist revision 3 and the human reviewers revision 1 to EIP-7708 in Amsterdam / Glamsterdam. Revision 3 phrases the same criteria more precisely; differences cover the 23 criteria both revisions share, and each total keeps its own revision. Δ is LLM minus Human.

LLM17Medium
Human9Low
Δ total+8Tiers differ: Medium vs Low
Criteria16/23agree exactly · 4 differ by 1 · 3 differ by 2+
Confounded comparison. The human checklist evaluated the EIP as it stood when the checklist was written; the LLM evaluated the sealed historical revision. Input alignment: exact blob; human hindsight exposure: high exposure.
Exposure evidence

The rationale describes the then-current execution-specs test-framework handling of transition-tool logs and a required enhancement.

The human-time and approved Task 04 EIP commits resolve to the same Git blob.

Complexity profiles side by side

LLM
Human

Largest disagreements: Modified system contracts (−2), Modified opcodes (+2), Cross-EIP interactions (+2), Patterns affecting pre-existing tests (+1), Transition-tool interface changes (−1)

Per-criterion scores, Human versus LLM, ordered by the size of the difference
CriterionLLMHumanΔAgreementRationale from each source
Modified opcodes31+2Differ by 2+
Show rationale

LLM Emitted logs are part of instruction semantics, and the logs emitted by CALL and SELFDESTRUCT change.

Human *CALL opcodes and SELFDESTRUCT/SENDALL behavior is slightly modified.

Modified system contracts02−2Differ by 2+
Show rationale

LLM No existing system contract's rules or surrounding behavior change.

Human Some system contracts do receive/transfer Eth, and while existing tests might be sufficient, we need to double check that this change does not break current behavior (particularly the beacon deposit contract).

Cross-EIP interactions20+2Differ by 2+
Show rationale

LLM The SELFDESTRUCT trigger needs coordinated cases against the baseline (Osaka) SELFDESTRUCT rules: same-transaction-created vs pre-existing contracts, beneficiary equal to self (burn), and balance-only transfer. ERC-20 needs only a local check that the log layout is compatible.

Human No rationale recorded.

Transition-tool interface changes01−1Differ by 1
Show rationale

LLM Logs are already a t8n output. No new input, output field or mechanism is needed.

Human Rather than a change to the T8N interface, we have to enhance execution-specs testing framework to verify the logs returned from T8N, because at the moment we just pass along the log list that was received from T8N to the test to be compared against the clients.

Patterns affecting pre-existing tests21+1Differ by 1
Show rationale

LLM Baseline tests that assert log lists or log indices for value-transferring transactions or calls need revised expected logs. Examples are LOG-opcode tests in transactions with value, and CALL or SELFDESTRUCT tests with value. Their logs hash, bloom and receipt root also change. This rework covers ordinary cases in several log-checking families. It mostly means regenerating expectations rather than a common restructuring, so level 2.

Human All tests will now issue Eth-transfer logs, which might be a non-issue but does touch basically every single test.

Security risks10+1Differ by 1
Show rationale

LLM The new security conditions are local: logs must be emitted exactly when a transfer happens, logs from reverted frames must not survive, and genuine transfer logs should be distinguishable from contract-emitted LOG3 forgeries. No other EL component's assumptions change.

Human No rationale recorded.

Performance risks12−1Differ by 1
Show rationale

LLM Receipt and bloom generation and log storage grow on average. Component benchmarks of log-heavy value-transfer blocks are enough, and end-to-end assumptions stay bounded according to the spec.

Human Test worst-case scenario of a transaction/block transfering Eth in multiple ways.

Added opcodes000Agree
Show rationale

LLM No new opcodes.

Human No rationale recorded.

Added precompiles000Agree
Show rationale

LLM None added.

Human No rationale recorded.

Modified precompiles000Agree
Show rationale

LLM No precompile semantics or gas change.

Human No rationale recorded.

Added system contracts000Agree
Show rationale

LLM None added.

Human No rationale recorded.

EVM Gas rule changes000Agree
Show rationale

LLM No execution-gas accounting rule is specified as changing. The log appears to be emitted without its own gas charge.

Human No rationale recorded.

Blob gas accounting changes000Agree
Show rationale

LLM Blob gas is unaffected.

Human No rationale recorded.

New EVM gas refund000Agree
Show rationale

LLM No refund is introduced.

Human No rationale recorded.

New transaction types000Agree
Show rationale

LLM None.

Human No rationale recorded.

New or modified transaction validity mechanisms000Agree
Show rationale

LLM Transaction validity is unchanged.

Human No rationale recorded.

New block / header fields000Agree
Show rationale

LLM Changed values in existing fields do not count.

Human No rationale recorded.

Encoding changes (RLP/SSZ)000Agree
Show rationale

LLM New values within an unchanged receipt and log schema do not count.

Human No rationale recorded.

Block syncing changes000Agree
Show rationale

LLM Only receipt contents change, through ordinary execution. Structural validation is unchanged.

Human No rationale recorded.

New fork activation mechanism000Agree
Show rationale

LLM No activation-specific transition.

Human No rationale recorded.

Engine API changes000Agree
Show rationale

LLM No Engine API fields or endpoints change.

Human No rationale recorded.

Edge/boundary conditions220Agree
Show rationale

LLM Several boundary-sensitive mechanisms are introduced. These are the nonzero-value triggers for transactions, CALL and SELFDESTRUCT; whether the transfer actually executes (insufficient balance, call depth, OOG); log ordering; and survival of logs in reverted frames. They can mostly be tested independently.

Human All 0->1 boundaries have to be tested in all possible scenarios where Eth is transfered.

Cryptography000Agree
Show rationale

LLM No cryptographic change.

Human No rationale recorded.

State-access ordering within opcode execution0n/a—Only in revision 3
Show rationale

LLM Emitting a log does not access state. The order of state access and gas charging in CALL and SELFDESTRUCT is unchanged, and no access-list rule is introduced.

Human No rationale recorded.

State gas accounting changes0n/a—Only in revision 3
Show rationale

LLM There are no state-gas accounting changes.

Human No rationale recorded.

Engine API encoding changesn/a0—Only in revision 1—
New invariant on pre-existing tests2n/a—Only in revision 3
Show rationale

LLM Baseline tests across many families (plain transfers, value calls, SELFDESTRUCT, contract interactions) must now check the new transfer log. Zero-value tests and pre-fork vectors are unaffected, so the requirement is neither universal nor re-derived for earlier forks.

Human No rationale recorded.

New test-framework primitives2n/a—Only in revision 3
Show rationale

LLM The tests need a new expectation helper that builds the expected transfer log from a transfer and places it correctly among the other logs (transaction log first; call logs inline). That is a new expectation abstraction within the target's suite.

Human No rationale recorded.

Unspecified behavior requiring cross-client consensus2n/a—Only in revision 3
Show rationale

LLM Several consensus-visible outcomes (receipts and bloom) have competing interpretations: the MAGIC value, the log address field, gas charging, coverage of CALLCODE and CREATE, SELFDESTRUCT to self or burn, failed value calls, revert handling, and withdrawals and fees. Clients and the spec must agree on these before expected results can be fixed.

Human No rationale recorded.

Criterion legend and glossary

Every stacked bar, comparison matrix, and criterion table on this site uses the same criterion colours, abbreviations, and order. Colour marks the criterion group; the abbreviation and name identify the criterion. Scores are 0–3 per criterion (4 is exceptional; cross-EIP interactions is uncapped).

EVM surface

Opcodes, precompiles, and system contracts that are added or modified.

  • Added opcodes
    Introduces new opcodes
    Score anchors
    0
    No new opcodes are introduced.
    1
    A new simple opcode is introduced (no data portion, no complex stack mechanics, and a constant gas cost).
    2
    Multiple new simple opcodes are introduced, or a single new complex opcode is introduced (has data portion, or complex stack mechanics, or a dynamic gas cost).
    3
    Multiple new opcodes are introduced, and at least one of them is complex (has data portion, or complex stack mechanics, or a dynamic gas cost).
    • Cryptography opcodes are not considered complex by default. Refer to the "Cryptography" section for a separate assessment.
  • Modified opcodes
    Modifies pre-existing opcodes
    Score anchors
    0
    No pre-existing opcode modifications are introduced.
    3
    At least one pre-existing opcode's behavior is modified (not including gas changes) or a pre-existing opcode is deprecated.
  • Added precompiles
    Introduces new precompiles
    Score anchors
    0
    No new precompiles are introduced.
    1
    A new simple precompile is introduced (constant input length, constant gas cost).
    2
    Multiple new simple precompiles are introduced, or a single new complex precompile is introduced (dynamic input length or dynamic gas cost).
    3
    Multiple new precompiles are introduced, and at least one of them is complex (dynamic input length or dynamic gas cost).
    • Cryptography precompiles are not considered complex by default. Refer to the "Cryptography" for a separate assessment.
  • Modified precompiles
    Modifies pre-existing precompiles logic or gas-accounting
    Score anchors
    0
    No pre-existing precompiles are modified.
    1
    At least one pre-existing precompile has its gas schedule modified.
    2
    Multiple pre-existing precompiles have their gas schedule modified, or a single pre-existing precompile has its behavior modified.
    3
    The behavior of multiple pre-existing precompiles, or a single complex pre-existing precompile modified.
  • Added system contracts
    Introduces new system contract, stateful or not
    Score anchors
    0
    No new system contracts are introduced.
    1
    A new system contract is introduced that is not stateful nor does it trigger a new system action (e.g. requests to the consensus layer).
    2
    Multiple new system contracts are introduced or a single new system contract that is either stateful or triggers a new system action (e.g. requests to the consensus layer).
    3
    Multiple new system contracts are introduced and at least one of them is either stateful or triggers a new system action (e.g. requests to the consensus layer).
  • Modified system contracts
    Modifies pre-existing system contracts
    Score anchors
    0
    No modifications to pre-existing system contracts are introduced, directly or indirectly.
    1
    Does not directly modify any system contract, but its behavior has minor indirect effects on one or more system contracts.
    2
    Does not directly modify any system contract, but its behavior has major indirect effects on one or more system contracts.
    3
    At least one pre-existing system contract code or state is modified, which would involve irregular state transition or a similarly complex transition methodology.

Gas and accounting

Execution, blob, and state gas rules, refunds, and where charges happen inside opcodes.

  • EVM Gas rule changes
    New EVM gas accounting rules
    Score anchors
    0
    No gas accounting changes.
    1
    Existing gas accounting mechanism is updated.
    2
    A new gas accounting mechanism is introduced but it does not affect existing mechanisms nor does it affect existing tests.
    3
    A new gas accounting mechanism is introduced and affects existing mechanisms which in turn affect existing tests.
  • State-access ordering within opcode execution · not in checklist revision 1
    Changes *where inside an opcode's execution* state is accessed, or where gas is charged relative to that access. Because a state access is recorded in the block-level access list only if execution had enough gas to reach it, this ordering is consensus-critical: moving it changes the BAL at every gas boundary of every affected opcode.
    Score anchors
    0
    No change to where state is accessed, or to where gas is charged relative to a state access, within any opcode.
    1
    A single opcode's state-access or gas-charge ordering changes.
    2
    Multiple opcodes' ordering changes, or a new state-accessing operation is introduced whose position in the order must be settled.
    3
    The ordering rule changes for a whole class of state-accessing opcodes at once, or what counts as a recordable state access is redefined — requiring existing BAL vectors to be re-derived across opcodes and forks.
    • Distinct from "Modified opcodes", which asks whether an opcode's **result** changed. This row asks about the **path to the result**, which is observable even when the result is identical. An EIP can be 0 on that row and 3 on this one.
    • Score changes **to** the ordering. Do not score the fact that state accesses are observable — they always are.
    • Each boundary must be re-tested against every other dimension that can change the answer (cold/warm, static/non-static, delegated/direct, revert/success), so the case count grows multiplicatively rather than additively. Note this explicitly under Special Considerations.
  • Blob gas accounting changes
    New Blob gas accounting rules which potentially affect pre-existing tests
    Score anchors
    0
    No blob gas accounting changes.
    1
    Existing blob gas accounting mechanism is updated.
    2
    A new blob gas accounting mechanism is introduced but it does not affect existing mechanisms nor does it affect existing tests.
    3
    A new blob gas accounting mechanism is introduced and affects existing mechanisms which in turn affect existing tests.
  • State gas accounting changes · not in checklist revision 1
    New state gas accounting rules. State gas is the cost of *writing* state, as opposed to accessing or executing it: `StateGasCosts`, `COST_PER_STATE_BYTE`, the block-level state gas budget, and the spill path into execution gas.
    Score anchors
    0
    No state gas accounting changes.
    1
    An existing state gas cost or `STATE_BYTES_PER_*` rate is adjusted.
    2
    A new state-gas-charging site is introduced, or the block-level state gas budget or reservoir allocation is modified.
    3
    A new state gas charging mechanism is introduced, or the spill interaction between state gas and execution gas is modified, affecting existing gas tests.
    • Harder to test than blob gas: the spill path means state gas cannot be metered independently of execution gas, and some costs (e.g. `NEW_ACCOUNT`) are state-dependent.
  • New EVM gas refund
    New gas-refund mechanism
    Score anchors
    0
    No new gas-refund mechanisms are introduced.
    1
    A new simple gas-refund mechanism is introduced that does not affect either existing tests or existing gas-refund mechanisms.
    2
    A new complex gas-refund mechanism is introduced or a simple mechanism that affects existing tests or existing gas-refund mechanisms.
    3
    A new complex gas-refund mechanism is introduced that affects existing tests or existing gas-refund mechanisms.

Blocks, transactions, and encoding

Transaction types and validity, block and header fields, encodings, syncing, and activation-time changes.

  • New transaction types
    Introduces a new transaction type
    Score anchors
    0
    No new transaction types are introduced.
    3
    A new transaction type is introduced.
  • New or modified transaction validity mechanisms
    Creates new or modifies pre-existing transaction types' validation mechanisms
    Score anchors
    0
    No changes are introduced to the validity rules of existing transaction types or to their intrinsic gas cost calculation.
    1
    Minor adjustments are introduced to validity rules or intrinsic gas cost calculation, but they do not significantly affect existing tests.
    2
    Changes to validity rules or intrinsic gas cost calculation affect existing tests, but require only limited updates to test cases and no redesign of the testing infrastructure.
    3
    Changes to validity rules or intrinsic gas cost calculation require extensive rework or redesign of the tests or testing infrastructure.
  • New block / header fields
    Introduces new block or block header fields
    Score anchors
    0
    No new block or header fields are introduced.
    3
    A new block or header field is introduced.
  • Encoding changes (RLP/SSZ)
    Introduces encoding changes at the transaction/block/interfaces level
    Score anchors
    0
    No encoding changes are introduced at the transaction, block, or interfaces levels.
    3
    An encoding change is introduced at transaction, block or interfaces level (e.g. RLP -> SSZ).
    • "Interfaces level" includes the Engine API. Score an Engine API encoding change (e.g. JSON -> SSZ) here.
  • Block syncing changes
    Modifies block RLP validation mechanisms that require test client syncing.
    Score anchors
    0
    No new RLP validation mechanism is introduced.
    1
    A single simple RLP validation mechanism is introduced.
    2
    Multiple simple RLP validation mechanisms are introduced or a single complex one.
    3
    Multiple RLP validation mechanisms are introduced and at least one of them is deemed complex.
  • New fork activation mechanism
    Modifies state, internal variables, or similar, at the fork activation block
    Score anchors
    0
    No state modifications, internal variables or similar are modified at the fork activation block.
    3
    Either a state modification or internal variables are modified at the fork activation block.
    • Initialization of new internal variable is not considered a modification.

Client interfaces

Engine API and transition-tool interface changes.

  • Engine API changes
    Introduces new fields to the Engine API directives
    Score anchors
    0
    No new fields or communication mechanisms are introduced to the Engine API.
    1
    A single new field is introduced in one of the Engine API endpoints.
    2
    Multiple fields are introduced to one or multiple Engine API end points, or a new Engine API end-point is introduced.
    3
    Multiple fields are introduced to one or multiple Engine API end points and a new Engine API end-point is introduced.
  • Engine API encoding changes · Checklist revision 1 only
    Engine API encoding changes (the revision-1 template defines no anchor text for this row).
  • Transition-tool interface changes
    Modifies or adds new fields to the transition tool interface.
    Score anchors
    0
    No modifications to the transition tool interface are required.
    1
    A single new field needs to be introduced to the transition tool interface.
    2
    Multiple new fields or a new mechanism has to be introduced to the transition tool interface.
    3
    Multiple new fields and a new mechanism has to be introduced to the transition tool interface.
    • Special consideration must be paid to this section if the EIP introduces a mechanism that requires the state transition tool to be aware whether the block it is processing is the fork-activation block.

Testing impact

Rework, new invariants, and new primitives required in the test framework.

  • Patterns affecting pre-existing tests
    Implements a new validation mechanism or rule that translates in reworking pre-existing tests
    Score anchors
    0
    No pre-existing tests are affected by this change.
    1
    Minor subset of existing tests are affected by this change.
    2
    Considerable subset of existing tests are affected by this change but involves only a contrived category of tests.
    3
    Major subset of existing tests are affected, including diverse category of tests (benchmarks, static, multiple forks, etc.).
  • New invariant on pre-existing tests · not in checklist revision 1
    Tests that are **not about this EIP** must nonetheless assert something this EIP produces. Their logic does not change; they gain a new thing to check.
    Score anchors
    0
    Pre-existing tests assert nothing new.
    1
    A narrow, contrived category of pre-existing tests gains a new assertion.
    2
    A broad category gains a new assertion, applied mechanically.
    3
    Every test in the fork gains the assertion regardless of what it tests, and pre-fork vectors must be re-derived to satisfy it.
    • Paired with the row above, and easy to confuse with it. "Patterns affecting pre-existing tests" asks whether existing tests must be **reworked**; this row asks whether they must **additionally assert something new**. Score both — an EIP can be low on one and high on the other.
  • New test-framework primitives · not in checklist revision 1
    Requires new abstractions in the test framework itself — expectation types, modifiers, helpers — beyond writing test functions with what already exists.
    Score anchors
    0
    Existing test primitives suffice.
    1
    Existing primitives need minor extension.
    2
    New expectation or modifier primitives are required, reusable within this EIP's own test suite.
    3
    New framework-level primitives are required that become a permanent part of the framework and are used by other EIPs' tests.

Risk and validation

Security, performance, boundary conditions, and cryptography that need validation.

  • Security risks
    Introduces or modifies mechanisms that could compromise the security of the chain, users, validators, or other stakeholders, if not implemented properly.
    Score anchors
    0
    No new mechanisms are introduced that could pose a security risk.
    1
    The introduced mechanisms are self-contained, can be validated in isolation, and do not alter existing invariants that could pose a security risk for any stakeholders.
    2
    The introduced mechanisms interact with a limited number of existing components, slightly altering their security assumptions and requiring a targeted security review or fuzzing.
    3
    The introduced mechanisms interact with multiple existing components, including critical ones, substantially altering their security assumptions and requiring an extensive security review and fuzzing.
  • Performance risks
    Introduces or modifies mechanisms and requires performance validation.
    Score anchors
    0
    No new mechanisms are introduced that require performance validation.
    1
    The introduced mechanisms can be benchmarked in isolation and do not affect existing performance behavior.
    2
    The introduced mechanisms cannot be fully benchmarked in isolation, but they only have a limited impact on the existing performance benchmarks.
    3
    The introduced mechanisms cannot be benchmarked in isolation and have a substantial impact on existing performance benchmarks or have complex interactions with existing mechanisms.
  • Edge/boundary conditions
    Feature contains edge/boundary conditions.
    Score anchors
    0
    No discernible edge cases or boundary conditions are introduced.
    1
    A single edge-case or boundary-condition prone mechanism is introduced.
    2
    Multiple edge-case or boundary-condition prone mechanisms are introduced, but none of them requires an elevated number of cases to test.
    3
    Multiple edge-case or boundary-condition prone mechanisms are introduced and at least one of them requires an elevated number of cases to test.
  • Cryptography
    Introduces new cryptography mechanisms or modifies existing functionality that involves cryptography
    Score anchors
    0
    No cryptography mechanisms are introduced.
    1
    A new cryptography mechanism is introduced but it is a well known mechanism that is known to have vast resources to aid on its testing.
    2
    Multiple new cryptography mechanisms are introduced that are well-known or a single but novel mechanism is introduced that is either untested or has limited resources.
    3
    Multiple new cryptography mechanisms are introduced and at least one of them is a novel mechanism.

Coordination

Cross-EIP interactions and behavior that clients must agree on before tests exist.

  • Cross-EIP interactions
    Introduces or modifies mechanisms that affect other EIPs in either the same or past forks.
    Score anchors
    0
    Fully self-contained EIP that does not depend on, modify, or conflict with any other EIP.
    1
    The EIP interacts with one or more other EIPs in a non-critical and limited way but can be tested independently for the most part.
    2
    The EIP depends on or modifies one or more other EIPs such that coordinated testing and consideration is required, but interactions are limited in scope and not complex.
    3
    The EIP has strong interdependencies with multiple EIPs, requiring extensive coordinated cross-EIP testing as well as potential re-design of existing test vectors.
    • +1 for every 3 additional interacting EIPs beyond the first 3, each of which requires its own coordinated test cases. List the EIPs in the rationale.
    • This row is intentionally uncapped, unlike every other anchor: each interacting EIP is another axis of the test matrix, so a ceiling would make a 12-EIP product indistinguishable from a 3-EIP one.
  • Unspecified behavior requiring cross-client consensus · not in checklist revision 1
    The EIP text does not determine the answer for cases a test can construct. Clients must agree on a previously unspecified detail before tests can be baselined. The cost here is coordination and re-baselining, not test writing.
    Score anchors
    0
    The EIP text determines the answer for every case a test could construct.
    1
    A few details are unspecified but have an obvious intended reading.
    2
    Details require client agreement before tests can be written, but they are localized.
    3
    A previously unspecified *and previously unobservable* behavior becomes consensus-critical; expect tests to be re-baselined on each round of EIP amendment.
    • Score this from the EIP's state at assessment time: whether it has client implementations, whether it has been through a devnet, and how many open questions remain on its discussion thread.