Retrospective LLM-Based Complexity Evaluations

EIP complexity assessment

EIP-7979: Call and Return Opcodes for the EVM

Assessed in Hegotá. The score describes the EIP text available at the snapshot, not the EIP as it stands today.

ProspectiveHegotáSnapshot 2026-10-07EIP-8081: CFILayers: execution
LLM Completescore 13
Human Pending· No STEEL checklist existed on the ethspecs/pm default branch or in any open pull request at the snapshot.

LLM assessment

Evaluated on: · Spec revision: 2026-10-07 · 6dac5e7491 · EIP-8081 list: CFI

Scope at the cutoff. EIP-7979 adds three EVM instructions: CALLSUB (pop a destination, push PC+1 to a new return stack, jump), CALLDEST (a no-op subroutine-entry label costing 1 gas), and RETURNSUB (pop the return stack into the PC). The return stack is a new part of machine state that EVM code cannot read or write directly. It holds at most 1024 entries, and both overflow and underflow cause an exceptional halt. Jump-destination analysis is extended so CALLDEST is a valid target for CALLSUB, JUMP and JUMPI, which changes the set of valid destinations for the existing jump instructions. Opcode byte values are still to be decided. EIP-8173 is an informational background document and defines no protocol rules.

13MediumMedium
Evaluator
LLMChecklist v3
Confidence
Medium
Under-specified at assessment cutoff
Yes — 3 criteria affected
Plausible range
12–15 (Medium)
Snapshot
2026-10-07 · EIP revision 6dac5e7491 (2026-10-07)
Score bands · Checklist revision 3
  • Low <12
  • Medium 12–22
  • High ≥23

28 criteria scored 0–3 (4 in exceptional cases; cross-EIP interactions is uncapped); nominal maximum 84.

Complexity profile

Each segment is one criterion's contribution to the LLM total. Hover or focus a segment for its score and rationale.

Top complexity drivers

  1. Modified opcodes3
  2. Added opcodes2
  3. Security risks2
  4. Edge/boundary conditions2

Under-specified at assessment cutoff: Yes

The EIP text available at the assessment cutoff left material behavior unresolved. The affected criteria and the plausible total range record that uncertainty.

Why: The opcode byte values are not yet assigned. The return stack's scope (per execution frame vs shared across nested calls) is implied by calling it machine state, not stated explicitly. The order of CALLSUB's checks is unspecified but not observable.

Unresolved questions at the cutoff (3)
  • Which opcode byte values will CALLSUB, CALLDEST and RETURNSUB use?
  • Is the return stack explicitly fresh and empty for each new message-call or create frame, and discarded on return?
  • Does the 1024 return-stack limit apply per frame or across the whole call depth?
Notable ambiguities noted by the assessor (3)
  • The EIP says JUMP and JUMPI 'need no change', yet their valid-destination set now includes CALLDEST, which is an observable semantic change.
  • The placeholder opcodes 0xB0–0xB2 in the test cases are explicitly unconfirmed.
  • Per-frame return-stack isolation is inferred from the 'machine state' wording and the EELS Evm dataclass, not stated normatively.

Criterion breakdown

EIP-7979 Hegotá: LLM criterion scores and rationale
CriterionScoreWhy this scoreEvidence / uncertainty
Modified opcodes3JUMP and JUMPI change control-flow and exceptional-halt semantics: jumping to a CALLDEST byte now succeeds where it previously halted. This is a semantic change beyond gas, so the level is 3.
  • eip.md · CALLDEST - "A CALLDEST is also a valid JUMP and JUMPI destination" The set of valid destinations for the existing JUMP and JUMPI instructions grows, changing their exceptional-halt behavior.
  • eip.md · Reference Implementation - "CALLDEST positions ... lands in both sets" The shared jump-destination analysis is extended.
Confidence: High
Uncertainty: The EIP describes JUMP as needing 'no change' in code, but the observable semantics do change.
Added opcodes2Multiple instructions are added and all meet the rubric's definition of simple (no immediates, fixed stack effects, constant gas), so the level is 2. The new return-stack state adds test content but does not change the classification under the rubric definition.
  • eip.md · Abstract Three new instructions: CALLSUB, CALLDEST, RETURNSUB.
  • eip.md · Specification / Costs Each instruction has a constant gas cost (8, 1, 5), no immediate data, and a fixed data-stack effect (CALLSUB pops 1; the others pop 0).
Confidence: High
Security risksUnder-specified2Beyond local opcode checks, the change alters the jump-destination validity that existing JUMP/JUMPI and client code-analysis caches rely on. The return stack must also stay isolated per execution frame across CALL/CREATE boundaries. These bounded interactions call for targeted differential fuzzing and review.
  • eip.md · Security Considerations Runtime checks: CALLSUB destination must be a CALLDEST, empty-stack halt, and 1024-entry limit. Return addresses must stay inaccessible.
  • eip.md · CALLDEST - "also a valid JUMP and JUMPI destination" The jump-destination analysis shared with existing JUMP/JUMPI (including PUSH-data skipping) changes.
Confidence: Medium
Uncertainty: If these are treated as purely interpreter-local checks, level 1 applies.
Edge/boundary conditions2Several independent boundary-sensitive rules are introduced: the return-stack depth limit (1023/1024/1025), underflow on an empty return stack, destination validity (CALLDEST vs JUMPDEST vs PUSH data vs out of range vs large 256-bit values), and return past end of code. The instruction × destination-type table is small and can be enumerated, so it is not treated as an elevated matrix.
  • eip.md · CALLSUB - "or the return stack already holds 1024 items" There is a return-stack depth limit of 1024 (overflow halt).
  • eip.md · RETURNSUB - "If the return stack is empty" There is an underflow halt, including when a subroutine is entered by JUMP with an empty return stack.
  • eip.md · Reference Implementation - get_valid_destinations Destination validity depends on the byte being CALLDEST and not lying inside PUSH data. Out-of-range destinations and JUMPDEST targets of CALLSUB are invalid.
  • eip.md · Subroutine at end of code A return to PC past the end of code acts as an implicit STOP.
Confidence: Medium
Uncertainty: The {CALLSUB, JUMP, JUMPI} × {CALLDEST, JUMPDEST, PUSH-data, out-of-range} combinations could be argued to form an elevated matrix, which would give level 3.
Patterns affecting pre-existing testsUnder-specified1Baseline rework is limited to boundary cases: invalid/undefined-opcode tests that cover the three newly assigned bytes, and invalid-jump-destination cases that happen to target those bytes. Ordinary cases do not change, so this is level 1.
  • eip.md · Specification - "Opcode values are still to be determined" Three byte values that are currently undefined will gain meaning.
  • eip.md · CALLDEST - "A CALLDEST is also a valid JUMP and JUMPI destination" Bytecode that jumps to the chosen CALLDEST byte would no longer halt with an invalid jump destination.
  • eip.md · Backwards Compatibility The EIP states that the semantics of existing code do not change, apart from code that already had unspecified behavior.
Confidence: Medium
Uncertainty: Since the opcode values are not assigned, it is not yet known which baseline vectors use those bytes.
New test-framework primitives1Adding entries to the existing opcode definitions is a local extension of an existing primitive. The return stack is unobservable, so behavior is checked through ordinary storage and gas results, and no new expectation abstraction is needed.
  • eip.md · Test Cases - "placeholder opcode values 0xB0=CALLSUB, 0xB1=CALLDEST, 0xB2=RETURNSUB" The framework's opcode/bytecode definitions must add three instructions.
  • eip.md · CALLSUB - "return stack already holds 1024 items" Overflow tests need recursive bytecode generators, which can be built from ordinary bytecode.
Confidence: Medium
Performance risks1Component benchmarks of the new opcodes and of worst-case return-stack growth (including deep recursion) and the extended destination analysis cover the changed workload. Baseline end-to-end assumptions are not changed.
  • eip.md · Costs - "Benchmarking will be needed to tell if the costs are well-balanced" The cost of each new opcode needs benchmarking.
  • eip.md · CALLSUB Each frame's return stack can reach 1024 entries filled at 8 gas per push, which bounds worst-case memory per frame.
Confidence: Medium
Uncertainty: If return stacks are allocated per frame across nested call depth, a targeted integrated benchmark might be warranted (level 2).
Unspecified behavior requiring cross-client consensusUnder-specified1Some details are left open: the opcode values, and per-frame scope of the return stack is stated only implicitly. The surrounding text (machine state, the EELS Evm dataclass) supports one intended outcome. In CALLSUB, the order of the destination check and the overflow check is not observable because both are exceptional halts.
  • eip.md · RETURNSUB Notes - "Opcode values are still to be determined" The opcode byte assignments are not fixed.
  • eip.md · Specification - "The EVM's machine state includes ... This EIP adds a return stack" Per-frame scope follows only implicitly from placing the return stack in machine state (and the per-frame Evm dataclass).
Confidence: Medium
Uncertainty: Opcode assignment must be agreed before vectors are final, which could be read as level 2.
Show 20 zero-score criteria
Zero-score criteria (Checklist revision 3)
CriterionScoreWhy this scoreEvidence / uncertainty
Added precompiles0None.
  • eip.md · Specification No precompiles are added.
Modified precompiles0None.
  • eip.md · Specification No precompiles are changed.
Added system contracts0None.
  • eip.md · Specification No system contracts are added.
Modified system contracts0None.
  • eip.md · Specification No system contracts are referenced.
EVM Gas rule changes0The only gas items are constant costs for new instructions, using existing tiers. No metering, limit or settlement rule changes, and no existing operation's gas result changes. Testing these costs belongs under added opcodes.
  • eip.md · Costs The new opcodes use the existing constant tiers: mid (8), jumpdest (1) and low (5).
  • eip.md · Reference Implementation - "GAS_MID (8), GAS_LOW (5), and GAS_JUMPDEST (1) are EELS's existing constants" No new gas-accounting mechanism is added and no existing gas rule changes.
Uncertainty: The EIP notes that benchmarking may change the costs, but these would still be constant per-opcode costs.
State-access ordering within opcode execution0No state access or access-list behavior is added or reordered.
  • eip.md · Specification The three instructions touch only the data stack, the return stack and the PC. None of them accesses state.
Blob gas accounting changes0Blob gas is not affected.
  • eip.md · Specification No blob-related content.
State gas accounting changes0State-gas accounting is not affected.
  • eip.md · Specification No state writes or state-gas rules are involved.
New EVM gas refund0No refund mechanism is added.
  • eip.md · Costs Only constant charges are defined; there are no refunds.
New transaction types0None.
  • eip.md · Specification No transaction types are added.
New or modified transaction validity mechanisms0Transaction validity and intrinsic gas are unchanged.
  • eip.md · Rationale - Why no immediate arguments or code sections? No deploy-time validation or transaction validity rules are added; everything is checked at run time.
New block / header fields0None.
  • eip.md · Specification No header fields are added.
Encoding changes (RLP/SSZ)0No schema or codec changes.
  • eip.md · Specification No serialized objects change.
Block syncing changes0No block-level validation changes.
  • eip.md · Specification Only EVM execution changes; block decoding and structure are unchanged.
New fork activation mechanism0No activation-specific state transition is needed.
  • eip.md · Backwards Compatibility No state migration or code installation is required; only the rules change.
Engine API changes0The Engine API is unchanged.
  • eip.md · Specification No Engine API content.
Transition-tool interface changes0The transition tool's inputs and outputs are unchanged.
  • eip.md · RETURNSUB Notes The return stack is not consensus-critical. No new block, transaction or environment inputs are added.
Uncertainty: Optional trace output for the return stack would be tooling only, not a consensus interface change.
New invariant on pre-existing tests0Baseline tests need no new assertion.
  • eip.md · RETURNSUB Notes - "its actual state is not observable by EVM code, nor consensus-critical" The return stack is not an observable output, and no new receipt, header or log output is added.
Cryptography0No cryptographic mechanism is added or changed.
  • eip.md · Specification No cryptographic content.
Cross-EIP interactions0The only candidate EIP, EIP-8173, has no protocol rules to interact with. The supplied documents name no other EIP-defined behavior needing coordinated cases, so the level is 0.
  • supporting/eip-8173.md · Security Considerations - "specifies no changes to the protocol" EIP-8173 is informational and defines no testable behavior.
  • eip.md · Motivation EIP-8173 is cited only as background.
Uncertainty: Return-stack isolation across call frames (e.g., DELEGATECALL or CREATE-family frames) may warrant compatibility checks, but the supplied text does not name these interactions.
Assessment provenance
Assessed EIP revision
ethereum/EIPs@6dac5e7491 EIPS/eip-7979.md committed 2026-10-07 · information cutoff 2026-10-07T22:23:55Z
Current master · File history · blob 32de0d1ffe · sha256 fae355ffb3ee
Rubric
Checklist revision 3 · ethspecs/pm@fe2f793b03
Evaluator
Opus 5.5 (claude-opus-5-5) at high effort, one tool-less call per EIP · isolation bubblewrap_claude_p_no_tools_v1
Source record
Frozen research record research/tasks/10-opus-v3-reassessment/prospective/outputs/assessments/hegota-2026-10-08/eip-7979.yaml · sha256 ae4524a7ce48
Supporting documents supplied with the EIP
supporting/eip-8173.md
Criterion legend and glossary

Every stacked bar, comparison matrix, and criterion table on this site uses the same criterion colours, abbreviations, and order. Colour marks the criterion group; the abbreviation and name identify the criterion. Scores are 0–3 per criterion (4 is exceptional; cross-EIP interactions is uncapped).

EVM surface

Opcodes, precompiles, and system contracts that are added or modified.

  • Added opcodes
    Introduces new opcodes
    Score anchors
    0
    No new opcodes are introduced.
    1
    A new simple opcode is introduced (no data portion, no complex stack mechanics, and a constant gas cost).
    2
    Multiple new simple opcodes are introduced, or a single new complex opcode is introduced (has data portion, or complex stack mechanics, or a dynamic gas cost).
    3
    Multiple new opcodes are introduced, and at least one of them is complex (has data portion, or complex stack mechanics, or a dynamic gas cost).
    • Cryptography opcodes are not considered complex by default. Refer to the "Cryptography" section for a separate assessment.
  • Modified opcodes
    Modifies pre-existing opcodes
    Score anchors
    0
    No pre-existing opcode modifications are introduced.
    3
    At least one pre-existing opcode's behavior is modified (not including gas changes) or a pre-existing opcode is deprecated.
  • Added precompiles
    Introduces new precompiles
    Score anchors
    0
    No new precompiles are introduced.
    1
    A new simple precompile is introduced (constant input length, constant gas cost).
    2
    Multiple new simple precompiles are introduced, or a single new complex precompile is introduced (dynamic input length or dynamic gas cost).
    3
    Multiple new precompiles are introduced, and at least one of them is complex (dynamic input length or dynamic gas cost).
    • Cryptography precompiles are not considered complex by default. Refer to the "Cryptography" for a separate assessment.
  • Modified precompiles
    Modifies pre-existing precompiles logic or gas-accounting
    Score anchors
    0
    No pre-existing precompiles are modified.
    1
    At least one pre-existing precompile has its gas schedule modified.
    2
    Multiple pre-existing precompiles have their gas schedule modified, or a single pre-existing precompile has its behavior modified.
    3
    The behavior of multiple pre-existing precompiles, or a single complex pre-existing precompile modified.
  • Added system contracts
    Introduces new system contract, stateful or not
    Score anchors
    0
    No new system contracts are introduced.
    1
    A new system contract is introduced that is not stateful nor does it trigger a new system action (e.g. requests to the consensus layer).
    2
    Multiple new system contracts are introduced or a single new system contract that is either stateful or triggers a new system action (e.g. requests to the consensus layer).
    3
    Multiple new system contracts are introduced and at least one of them is either stateful or triggers a new system action (e.g. requests to the consensus layer).
  • Modified system contracts
    Modifies pre-existing system contracts
    Score anchors
    0
    No modifications to pre-existing system contracts are introduced, directly or indirectly.
    1
    Does not directly modify any system contract, but its behavior has minor indirect effects on one or more system contracts.
    2
    Does not directly modify any system contract, but its behavior has major indirect effects on one or more system contracts.
    3
    At least one pre-existing system contract code or state is modified, which would involve irregular state transition or a similarly complex transition methodology.

Gas and accounting

Execution, blob, and state gas rules, refunds, and where charges happen inside opcodes.

  • EVM Gas rule changes
    New EVM gas accounting rules
    Score anchors
    0
    No gas accounting changes.
    1
    Existing gas accounting mechanism is updated.
    2
    A new gas accounting mechanism is introduced but it does not affect existing mechanisms nor does it affect existing tests.
    3
    A new gas accounting mechanism is introduced and affects existing mechanisms which in turn affect existing tests.
  • State-access ordering within opcode execution · not in checklist revision 1
    Changes *where inside an opcode's execution* state is accessed, or where gas is charged relative to that access. Because a state access is recorded in the block-level access list only if execution had enough gas to reach it, this ordering is consensus-critical: moving it changes the BAL at every gas boundary of every affected opcode.
    Score anchors
    0
    No change to where state is accessed, or to where gas is charged relative to a state access, within any opcode.
    1
    A single opcode's state-access or gas-charge ordering changes.
    2
    Multiple opcodes' ordering changes, or a new state-accessing operation is introduced whose position in the order must be settled.
    3
    The ordering rule changes for a whole class of state-accessing opcodes at once, or what counts as a recordable state access is redefined — requiring existing BAL vectors to be re-derived across opcodes and forks.
    • Distinct from "Modified opcodes", which asks whether an opcode's **result** changed. This row asks about the **path to the result**, which is observable even when the result is identical. An EIP can be 0 on that row and 3 on this one.
    • Score changes **to** the ordering. Do not score the fact that state accesses are observable — they always are.
    • Each boundary must be re-tested against every other dimension that can change the answer (cold/warm, static/non-static, delegated/direct, revert/success), so the case count grows multiplicatively rather than additively. Note this explicitly under Special Considerations.
  • Blob gas accounting changes
    New Blob gas accounting rules which potentially affect pre-existing tests
    Score anchors
    0
    No blob gas accounting changes.
    1
    Existing blob gas accounting mechanism is updated.
    2
    A new blob gas accounting mechanism is introduced but it does not affect existing mechanisms nor does it affect existing tests.
    3
    A new blob gas accounting mechanism is introduced and affects existing mechanisms which in turn affect existing tests.
  • State gas accounting changes · not in checklist revision 1
    New state gas accounting rules. State gas is the cost of *writing* state, as opposed to accessing or executing it: `StateGasCosts`, `COST_PER_STATE_BYTE`, the block-level state gas budget, and the spill path into execution gas.
    Score anchors
    0
    No state gas accounting changes.
    1
    An existing state gas cost or `STATE_BYTES_PER_*` rate is adjusted.
    2
    A new state-gas-charging site is introduced, or the block-level state gas budget or reservoir allocation is modified.
    3
    A new state gas charging mechanism is introduced, or the spill interaction between state gas and execution gas is modified, affecting existing gas tests.
    • Harder to test than blob gas: the spill path means state gas cannot be metered independently of execution gas, and some costs (e.g. `NEW_ACCOUNT`) are state-dependent.
  • New EVM gas refund
    New gas-refund mechanism
    Score anchors
    0
    No new gas-refund mechanisms are introduced.
    1
    A new simple gas-refund mechanism is introduced that does not affect either existing tests or existing gas-refund mechanisms.
    2
    A new complex gas-refund mechanism is introduced or a simple mechanism that affects existing tests or existing gas-refund mechanisms.
    3
    A new complex gas-refund mechanism is introduced that affects existing tests or existing gas-refund mechanisms.

Blocks, transactions, and encoding

Transaction types and validity, block and header fields, encodings, syncing, and activation-time changes.

  • New transaction types
    Introduces a new transaction type
    Score anchors
    0
    No new transaction types are introduced.
    3
    A new transaction type is introduced.
  • New or modified transaction validity mechanisms
    Creates new or modifies pre-existing transaction types' validation mechanisms
    Score anchors
    0
    No changes are introduced to the validity rules of existing transaction types or to their intrinsic gas cost calculation.
    1
    Minor adjustments are introduced to validity rules or intrinsic gas cost calculation, but they do not significantly affect existing tests.
    2
    Changes to validity rules or intrinsic gas cost calculation affect existing tests, but require only limited updates to test cases and no redesign of the testing infrastructure.
    3
    Changes to validity rules or intrinsic gas cost calculation require extensive rework or redesign of the tests or testing infrastructure.
  • New block / header fields
    Introduces new block or block header fields
    Score anchors
    0
    No new block or header fields are introduced.
    3
    A new block or header field is introduced.
  • Encoding changes (RLP/SSZ)
    Introduces encoding changes at the transaction/block/interfaces level
    Score anchors
    0
    No encoding changes are introduced at the transaction, block, or interfaces levels.
    3
    An encoding change is introduced at transaction, block or interfaces level (e.g. RLP -> SSZ).
    • "Interfaces level" includes the Engine API. Score an Engine API encoding change (e.g. JSON -> SSZ) here.
  • Block syncing changes
    Modifies block RLP validation mechanisms that require test client syncing.
    Score anchors
    0
    No new RLP validation mechanism is introduced.
    1
    A single simple RLP validation mechanism is introduced.
    2
    Multiple simple RLP validation mechanisms are introduced or a single complex one.
    3
    Multiple RLP validation mechanisms are introduced and at least one of them is deemed complex.
  • New fork activation mechanism
    Modifies state, internal variables, or similar, at the fork activation block
    Score anchors
    0
    No state modifications, internal variables or similar are modified at the fork activation block.
    3
    Either a state modification or internal variables are modified at the fork activation block.
    • Initialization of new internal variable is not considered a modification.

Client interfaces

Engine API and transition-tool interface changes.

  • Engine API changes
    Introduces new fields to the Engine API directives
    Score anchors
    0
    No new fields or communication mechanisms are introduced to the Engine API.
    1
    A single new field is introduced in one of the Engine API endpoints.
    2
    Multiple fields are introduced to one or multiple Engine API end points, or a new Engine API end-point is introduced.
    3
    Multiple fields are introduced to one or multiple Engine API end points and a new Engine API end-point is introduced.
  • Engine API encoding changes · Checklist revision 1 only
    Engine API encoding changes (the revision-1 template defines no anchor text for this row).
  • Transition-tool interface changes
    Modifies or adds new fields to the transition tool interface.
    Score anchors
    0
    No modifications to the transition tool interface are required.
    1
    A single new field needs to be introduced to the transition tool interface.
    2
    Multiple new fields or a new mechanism has to be introduced to the transition tool interface.
    3
    Multiple new fields and a new mechanism has to be introduced to the transition tool interface.
    • Special consideration must be paid to this section if the EIP introduces a mechanism that requires the state transition tool to be aware whether the block it is processing is the fork-activation block.

Testing impact

Rework, new invariants, and new primitives required in the test framework.

  • Patterns affecting pre-existing tests
    Implements a new validation mechanism or rule that translates in reworking pre-existing tests
    Score anchors
    0
    No pre-existing tests are affected by this change.
    1
    Minor subset of existing tests are affected by this change.
    2
    Considerable subset of existing tests are affected by this change but involves only a contrived category of tests.
    3
    Major subset of existing tests are affected, including diverse category of tests (benchmarks, static, multiple forks, etc.).
  • New invariant on pre-existing tests · not in checklist revision 1
    Tests that are **not about this EIP** must nonetheless assert something this EIP produces. Their logic does not change; they gain a new thing to check.
    Score anchors
    0
    Pre-existing tests assert nothing new.
    1
    A narrow, contrived category of pre-existing tests gains a new assertion.
    2
    A broad category gains a new assertion, applied mechanically.
    3
    Every test in the fork gains the assertion regardless of what it tests, and pre-fork vectors must be re-derived to satisfy it.
    • Paired with the row above, and easy to confuse with it. "Patterns affecting pre-existing tests" asks whether existing tests must be **reworked**; this row asks whether they must **additionally assert something new**. Score both — an EIP can be low on one and high on the other.
  • New test-framework primitives · not in checklist revision 1
    Requires new abstractions in the test framework itself — expectation types, modifiers, helpers — beyond writing test functions with what already exists.
    Score anchors
    0
    Existing test primitives suffice.
    1
    Existing primitives need minor extension.
    2
    New expectation or modifier primitives are required, reusable within this EIP's own test suite.
    3
    New framework-level primitives are required that become a permanent part of the framework and are used by other EIPs' tests.

Risk and validation

Security, performance, boundary conditions, and cryptography that need validation.

  • Security risks
    Introduces or modifies mechanisms that could compromise the security of the chain, users, validators, or other stakeholders, if not implemented properly.
    Score anchors
    0
    No new mechanisms are introduced that could pose a security risk.
    1
    The introduced mechanisms are self-contained, can be validated in isolation, and do not alter existing invariants that could pose a security risk for any stakeholders.
    2
    The introduced mechanisms interact with a limited number of existing components, slightly altering their security assumptions and requiring a targeted security review or fuzzing.
    3
    The introduced mechanisms interact with multiple existing components, including critical ones, substantially altering their security assumptions and requiring an extensive security review and fuzzing.
  • Performance risks
    Introduces or modifies mechanisms and requires performance validation.
    Score anchors
    0
    No new mechanisms are introduced that require performance validation.
    1
    The introduced mechanisms can be benchmarked in isolation and do not affect existing performance behavior.
    2
    The introduced mechanisms cannot be fully benchmarked in isolation, but they only have a limited impact on the existing performance benchmarks.
    3
    The introduced mechanisms cannot be benchmarked in isolation and have a substantial impact on existing performance benchmarks or have complex interactions with existing mechanisms.
  • Edge/boundary conditions
    Feature contains edge/boundary conditions.
    Score anchors
    0
    No discernible edge cases or boundary conditions are introduced.
    1
    A single edge-case or boundary-condition prone mechanism is introduced.
    2
    Multiple edge-case or boundary-condition prone mechanisms are introduced, but none of them requires an elevated number of cases to test.
    3
    Multiple edge-case or boundary-condition prone mechanisms are introduced and at least one of them requires an elevated number of cases to test.
  • Cryptography
    Introduces new cryptography mechanisms or modifies existing functionality that involves cryptography
    Score anchors
    0
    No cryptography mechanisms are introduced.
    1
    A new cryptography mechanism is introduced but it is a well known mechanism that is known to have vast resources to aid on its testing.
    2
    Multiple new cryptography mechanisms are introduced that are well-known or a single but novel mechanism is introduced that is either untested or has limited resources.
    3
    Multiple new cryptography mechanisms are introduced and at least one of them is a novel mechanism.

Coordination

Cross-EIP interactions and behavior that clients must agree on before tests exist.

  • Cross-EIP interactions
    Introduces or modifies mechanisms that affect other EIPs in either the same or past forks.
    Score anchors
    0
    Fully self-contained EIP that does not depend on, modify, or conflict with any other EIP.
    1
    The EIP interacts with one or more other EIPs in a non-critical and limited way but can be tested independently for the most part.
    2
    The EIP depends on or modifies one or more other EIPs such that coordinated testing and consideration is required, but interactions are limited in scope and not complex.
    3
    The EIP has strong interdependencies with multiple EIPs, requiring extensive coordinated cross-EIP testing as well as potential re-design of existing test vectors.
    • +1 for every 3 additional interacting EIPs beyond the first 3, each of which requires its own coordinated test cases. List the EIPs in the rationale.
    • This row is intentionally uncapped, unlike every other anchor: each interacting EIP is another axis of the test matrix, so a ceiling would make a 12-EIP product indistinguishable from a 3-EIP one.
  • Unspecified behavior requiring cross-client consensus · not in checklist revision 1
    The EIP text does not determine the answer for cases a test can construct. Clients must agree on a previously unspecified detail before tests can be baselined. The cost here is coordination and re-baselining, not test writing.
    Score anchors
    0
    The EIP text determines the answer for every case a test could construct.
    1
    A few details are unspecified but have an obvious intended reading.
    2
    Details require client agreement before tests can be written, but they are localized.
    3
    A previously unspecified *and previously unobservable* behavior becomes consensus-critical; expect tests to be re-baselined on each round of EIP amendment.
    • Score this from the EIP's state at assessment time: whether it has client implementations, whether it has been through a devnet, and how many open questions remain on its discussion thread.