Retrospective LLM-Based Complexity Evaluations

EIP complexity assessment

EIP-7934: RLP Execution Block Size Limit

Assessed in Osaka / Fusaka. The score describes the EIP text available at the assessment cutoff, not the EIP as it stands today.

RetrospectiveOsaka / FusakaAssessment cutoff 2025-05-09Added after cutoffLayers: execution
LLM Completescore 7
Human Not available· Human complexity assessments were not produced for this fork; only the LLM assessment exists.

LLM assessment

Evaluated on: · Spec revision: 2025-05-06 · 040800a325

Scope at the cutoff. EIP-7934 (revision 040800a, Draft) adds a consensus rule that an execution block is invalid if its RLP encoding is larger than MAX_RLP_BLOCK_SIZE. That limit is defined as MAX_BLOCK_SIZE (10 MiB = 10,485,760 bytes) minus a 512 KiB margin (524,288 bytes), giving 9,961,472 bytes. Block producers must not build blocks over the limit, and nodes must reject such blocks during validation and propagation. The check does not depend on gas metrics. The EIP adds no header fields, transaction types, opcodes, gas rules or Engine API changes.

7LowLow
Evaluator
LLMChecklist v3
Confidence
Medium
Under-specified at assessment cutoff
Yes — 5 criteria affected
Plausible range
6–10 (Low)
Assessment cutoff
2025-05-09 · EIP revision 040800a325 (2025-05-06)
Score bands · Checklist revision 3
  • Low <12
  • Medium 12–22
  • High ≥23

28 criteria scored 0–3 (4 in exceptional cases; cross-EIP interactions is uncapped); nominal maximum 84.

Complexity profile

Each segment is one criterion's contribution to the LLM total. Hover or focus a segment for its score and rationale.

Top complexity drivers

  1. Unspecified behavior requiring cross-client consensus2
  2. Block syncing changes1
  3. New test-framework primitives1
  4. Security risks1

Under-specified at assessment cutoff: Yes

The EIP text available at the assessment cutoff left material behavior unresolved. The affected criteria and the plausible total range record that uncertainty.

Why: The exact limit is internally inconsistent. The abstract says the cap is 10 MiB; the bullets define MAX_RLP_BLOCK_SIZE = 10 MiB − 512 KiB; the pseudocode uses an undefined GOSSIP_UPPER_LIMIT and a differently named SAFETY_MARGIN. The EIP also does not define which 'block' encoding is measured beyond rlp.encode(block), how the rule maps to payloads received through the Engine API, or the expected builder behavior when the next transaction would exceed the cap.

Unresolved questions at the cutoff (5)
  • Is the enforced limit 10,485,760 bytes (as the abstract implies) or 9,961,472 bytes (MAX_BLOCK_SIZE − MARGIN)? What is GOSSIP_UPPER_LIMIT?
  • Exactly which structure is RLP-encoded for the check: the full EL block (header, transactions, ommers, withdrawals) or another representation?
  • How does the check apply to payloads received through engine_newPayload? Is a specific status or error expected?
  • Does the rule apply from the first block at or after the fork activation timestamp?
  • Is there any normative builder rule for transaction selection near the cap, beyond 'must ensure ... does not exceed'?
Notable ambiguities noted by the assessor (5)
  • The abstract says the cap is 10 MiB, but the specification computes 10,485,760 − 524,288 = 9,961,472 bytes.
  • The pseudocode references GOSSIP_UPPER_LIMIT, which is never defined, and SAFETY_MARGIN, while the bullets use MARGIN.
  • Which object is encoded ('block') is not specified precisely: the full EL block [header, transactions, ommers, withdrawals] versus some other representation such as an Engine API payload.
  • Builder behavior is only stated as 'must ensure ... does not exceed'. There is no rule for transaction selection near the cap and no Engine API error semantics.
  • Fork activation timing is not stated explicitly. Presumably the check applies to blocks from the fork onward.

Criterion breakdown

EIP-7934 Osaka / Fusaka: LLM criterion scores and rationale
CriterionScoreWhy this scoreEvidence / uncertainty
Unspecified behavior requiring cross-client consensusUnder-specified2The text contradicts itself on the exact limit. The abstract says 10 MiB; the bullets give 10 MiB minus 512 KiB = 9,961,472 bytes; the code uses an undefined GOSSIP_UPPER_LIMIT. The validity outcome for blocks between 9,961,473 and 10,485,760 bytes therefore needs spec or client agreement before fixtures can be fixed. These are localized competing outcomes (level 2).
  • eip.md · Abstract — "cap on the maximum RLP-encoded execution block size to 10 megabytes (MiB), which includes a margin of 512 KiB" The abstract frames the cap as 10 MiB.
  • eip.md · Block Size Cap — "MAX_RLP_BLOCK_SIZE = GOSSIP_UPPER_LIMIT - SAFETY_MARGIN" The pseudocode uses an undefined constant (GOSSIP_UPPER_LIMIT) and names the margin SAFETY_MARGIN, while the bullets name it MARGIN and define the limit as MAX_BLOCK_SIZE - MARGIN.
  • eip.md · Block Size Cap — "Any RLP-encoded block" Which block encoding counts is not stated beyond 'block'. It does not address the Engine API payload versus the reconstructed RLP block.
Confidence: Medium
Uncertainty: Most readers would probably resolve GOSSIP_UPPER_LIMIT as MAX_BLOCK_SIZE, giving 9,961,472 bytes. If that reading is treated as unambiguous, this would be level 1.
Block syncing changesUnder-specified1One new structural validation rule is added. It is computed from the block's own encoding and does not depend on other blocks or state, so I treat it as simple (level 1).
  • eip.md · Changes to Protocol Behavior — "Block Validation: Nodes must reject blocks whose RLP-encoded size exceeds MAX_RLP_BLOCK_SIZE" Adds one structural block-validity rule that must be exercised through block import.
  • eip.md · Protocol Adjustment — "integrate this size check as part of block validation and propagation" The check applies during validation and propagation.
Confidence: Medium
Uncertainty: The check depends on the encoding of every block component, not a single field. A strict reading of "simple validates one field locally" could push it to complex (level 2).
New test-framework primitives1Tests need a local extension: a helper that pads block contents (transaction calldata, transaction count, withdrawals) to hit an exact encoded block size, plus a new block-exception value. This extends existing block-construction primitives. It does not need a new shared abstraction.
  • eip.md · Block Size Cap — "len(rlp.encode(block)) > MAX_RLP_BLOCK_SIZE" Boundary tests need blocks whose full RLP encoding is exactly at, or one byte above, 9,961,472 bytes.
Confidence: Medium
Uncertainty: Hitting the size boundary may require unusually high gas limits in genesis or block environments. RLP length-prefix discontinuities make exact sizing fiddly, but this is still local.
Security risks1The new rejection condition can be checked locally: clients must reject at the boundary consistently, or there is a consensus-split risk. The CL gossip alignment is the motivation, not a changed EL security assumption that needs integration fuzzing.
  • eip.md · Security Considerations — "protection against deliberate oversized-block attacks" Adds a new validation boundary against oversized blocks.
  • eip.md · Rationale — "aligns with the gossip protocol constraint in Ethereum's consensus layer" The limit is chosen to match the CL gossip limit.
Confidence: Medium
Uncertainty: Whether the EL limit stays consistent with the CL gossip limit and beacon block overhead could be treated as a bounded cross-layer interaction (level 2).
Edge/boundary conditionsUnder-specified1There is exactly one boundary-sensitive mechanism, the encoded block size limit. Size can come from different block parts (header, transactions, ommers list, withdrawals), but they all feed the same single limit.
  • eip.md · Block Size Cap — "return len(rlp.encode(block)) > MAX_RLP_BLOCK_SIZE" A single strict-inequality byte-size boundary: a block of exactly MAX_RLP_BLOCK_SIZE bytes is valid and one byte more is invalid.
Confidence: High
Uncertainty: There is a question of which value is the boundary: 10 MiB (abstract) or 10 MiB minus 512 KiB (specification). It is recorded under UNSP and does not change the count.
Cross-EIP interactions1Interactions are local compatibility checks only. Boundary cases should make sure every body component (transactions of every type, withdrawals, header) counts toward the size, and that high-gas-limit blocks hit the cap. No coordinated multi-EIP scenario restructuring is required. The candidate list is empty, so no numbered EIPs are recorded.
  • eip.md · Block Size Cap — "len(rlp.encode(block))" The size counts every block component, including body lists such as withdrawals that other EIPs define in the baseline.
  • eip.md · Protocol Adjustment — "This limit applies independently of gas-related metrics." Whether the cap can be reached depends on the gas limit and calldata pricing defined elsewhere.
Confidence: Medium
Uncertainty: The EIP numbers for withdrawals, typed transactions and calldata floor pricing are not in the candidate list or the supplied text. The interactions are described generically.
Show 22 zero-score criteria
Zero-score criteria (Checklist revision 3)
CriterionScoreWhy this scoreEvidence / uncertainty
Added opcodes0No new opcode.
  • eip.md · Specification No new instructions.
Modified opcodes0No opcode semantics change.
  • eip.md · Specification No instruction semantics change.
Added precompiles0No new precompile.
  • eip.md · Specification No precompile is introduced.
Modified precompiles0No precompile modified.
  • eip.md · Specification No precompile changes.
Added system contracts0No system contract added.
  • eip.md · Specification No contracts are introduced.
Modified system contracts0No system contract modified.
  • eip.md · Specification No system contract is referenced.
EVM Gas rule changes0No execution-gas charging, metering or settlement rule changes. The new limit is a byte-size validity check.
  • eip.md · Protocol Adjustment — "This limit applies independently of gas-related metrics." The size cap is separate from gas accounting. Gas charging and gas limits are unchanged.
State-access ordering within opcode execution0No opcode's state-access or gas-charge ordering changes.
  • eip.md · Specification — exceed_max_rlp_block_size The only change is a block-level check on encoded size. No instruction behavior is touched.
Blob gas accounting changes0No blob-gas accounting change.
  • eip.md · Block Size Cap The limit applies to the RLP-encoded execution block. Blob gas is not mentioned.
State gas accounting changes0No state-gas accounting change.
  • eip.md · Specification No state-write accounting is introduced or changed.
New EVM gas refund0No new refund mechanism.
  • eip.md · Specification No refund mechanism is mentioned.
New transaction types0No new transaction envelope.
  • eip.md · Specification No transaction type is defined.
New or modified transaction validity mechanisms0Consensus transaction-validity and intrinsic-gas rules are unchanged. Builders leaving out transactions to stay under the cap is block-construction policy, not a transaction-validity rule.
  • eip.md · Changes to Protocol Behavior The rule makes a block invalid. It does not change per-transaction eligibility or intrinsic gas.
Uncertainty: Each transaction's inclusion now depends on the block's cumulative encoded size. One could read that as a new block-level eligibility condition on transactions (level 1–2).
New block / header fields0No new header or block field.
  • eip.md · Specification No header or block member is added.
Encoding changes (RLP/SSZ)0No serialized schema or codec changes. Only the length of the existing encoding is bounded.
  • eip.md · Block Size Cap — "len(rlp.encode(block))" Uses the existing block RLP encoding to measure size. The schema is unchanged.
New fork activation mechanism0Activating the fork only switches on a rule. No activation-specific state transition.
  • eip.md · Specification Only a validity rule is added. There is no state migration or code installation.
Engine API changesUnder-specified0No Engine API field or endpoint change is specified. An oversized payload would be handled by existing INVALID-status semantics, and payload building must respect the cap. Both are behaviors under existing contracts.
  • eip.md · Changes to Protocol Behavior — "Block Creation: Validators must ensure ... does not exceed" Builders must respect the limit, but no Engine API field or method is changed.
Uncertainty: The EIP does not say how the cap applies to the payload as it arrives through the Engine API (reconstructed block RLP), or whether getPayload behavior needs explicit testing. No API change is specified.
Transition-tool interface changesUnder-specified0The block RLP is assembled and checked outside the transaction-level state transition, so no t8n field or mechanism is required. A tool that builds blocks could optionally report RLP size or a block exception, but the EIP does not require this.
  • eip.md · Changes to Protocol Behavior — "Nodes must reject blocks whose RLP-encoded size exceeds" The rule applies to the whole encoded block. This is naturally checked at block import, not inside the state transition over transactions.
Uncertainty: If the framework relies on t8n to signal block-level exceptions or to limit which transactions it includes, one output field may be needed (level 1).
Patterns affecting pre-existing tests0Ordinary baseline tests use small blocks and are unaffected. Only tests that build blocks over about 9.5 MiB (for example, stress or high-gas-limit vectors) would need rework. No such family is established by the supplied evidence.
  • eip.md · Backwards Compatibility — "not backward-compatible with any blocks larger than the newly specified size limit" Only blocks over 9,961,472 bytes become invalid. Ordinary baseline test blocks are far smaller.
Uncertainty: A baseline stress or benchmark family that uses very large blocks could need adjusting (level 1). No supplied test suite lets me confirm this.
New invariant on pre-existing tests0Baseline tests need no additional output assertion. The size check only affects validity.
  • eip.md · Changes to Protocol Behavior No new header, receipt, log or commitment output is produced.
Performance risks0The cap reduces the worst-case workload and adds no new resource-consuming workload. Computing the encoded size is cheap. No additional performance validation is established.
  • eip.md · Motivation — "Extremely large blocks slow down propagation" The change only tightens the upper bound on block size. It adds no workload.
Uncertainty: Builders may need incremental size tracking while assembling blocks, which might justify a component benchmark (level 1).
Cryptography0No cryptographic mechanism changes.
  • eip.md · Specification Only the length of the RLP encoding is checked. No hashing or signature rule changes.
Assessment provenance
Assessed EIP revision
ethereum/EIPs@040800a325 EIPS/eip-7934.md committed 2025-05-06 · information cutoff 2025-05-09T21:56:48Z
Current master · File history · blob 028e8657ab · sha256 29c5f346e4a9
Rubric
Checklist revision 3 · ethspecs/pm@fe2f793b03
Evaluator
Opus 5.5 (claude-opus-5-5) at high effort, one tool-less call per EIP · isolation bubblewrap_claude_p_no_tools_v1
Source record
Frozen research record research/tasks/10-opus-v3-reassessment/retrospective/outputs/assessments/osaka/eip-7934.yaml · sha256 918300a8dadd
Criterion legend and glossary

Every stacked bar, comparison matrix, and criterion table on this site uses the same criterion colours, abbreviations, and order. Colour marks the criterion group; the abbreviation and name identify the criterion. Scores are 0–3 per criterion (4 is exceptional; cross-EIP interactions is uncapped).

EVM surface

Opcodes, precompiles, and system contracts that are added or modified.

  • Added opcodes
    Introduces new opcodes
    Score anchors
    0
    No new opcodes are introduced.
    1
    A new simple opcode is introduced (no data portion, no complex stack mechanics, and a constant gas cost).
    2
    Multiple new simple opcodes are introduced, or a single new complex opcode is introduced (has data portion, or complex stack mechanics, or a dynamic gas cost).
    3
    Multiple new opcodes are introduced, and at least one of them is complex (has data portion, or complex stack mechanics, or a dynamic gas cost).
    • Cryptography opcodes are not considered complex by default. Refer to the "Cryptography" section for a separate assessment.
  • Modified opcodes
    Modifies pre-existing opcodes
    Score anchors
    0
    No pre-existing opcode modifications are introduced.
    3
    At least one pre-existing opcode's behavior is modified (not including gas changes) or a pre-existing opcode is deprecated.
  • Added precompiles
    Introduces new precompiles
    Score anchors
    0
    No new precompiles are introduced.
    1
    A new simple precompile is introduced (constant input length, constant gas cost).
    2
    Multiple new simple precompiles are introduced, or a single new complex precompile is introduced (dynamic input length or dynamic gas cost).
    3
    Multiple new precompiles are introduced, and at least one of them is complex (dynamic input length or dynamic gas cost).
    • Cryptography precompiles are not considered complex by default. Refer to the "Cryptography" for a separate assessment.
  • Modified precompiles
    Modifies pre-existing precompiles logic or gas-accounting
    Score anchors
    0
    No pre-existing precompiles are modified.
    1
    At least one pre-existing precompile has its gas schedule modified.
    2
    Multiple pre-existing precompiles have their gas schedule modified, or a single pre-existing precompile has its behavior modified.
    3
    The behavior of multiple pre-existing precompiles, or a single complex pre-existing precompile modified.
  • Added system contracts
    Introduces new system contract, stateful or not
    Score anchors
    0
    No new system contracts are introduced.
    1
    A new system contract is introduced that is not stateful nor does it trigger a new system action (e.g. requests to the consensus layer).
    2
    Multiple new system contracts are introduced or a single new system contract that is either stateful or triggers a new system action (e.g. requests to the consensus layer).
    3
    Multiple new system contracts are introduced and at least one of them is either stateful or triggers a new system action (e.g. requests to the consensus layer).
  • Modified system contracts
    Modifies pre-existing system contracts
    Score anchors
    0
    No modifications to pre-existing system contracts are introduced, directly or indirectly.
    1
    Does not directly modify any system contract, but its behavior has minor indirect effects on one or more system contracts.
    2
    Does not directly modify any system contract, but its behavior has major indirect effects on one or more system contracts.
    3
    At least one pre-existing system contract code or state is modified, which would involve irregular state transition or a similarly complex transition methodology.

Gas and accounting

Execution, blob, and state gas rules, refunds, and where charges happen inside opcodes.

  • EVM Gas rule changes
    New EVM gas accounting rules
    Score anchors
    0
    No gas accounting changes.
    1
    Existing gas accounting mechanism is updated.
    2
    A new gas accounting mechanism is introduced but it does not affect existing mechanisms nor does it affect existing tests.
    3
    A new gas accounting mechanism is introduced and affects existing mechanisms which in turn affect existing tests.
  • State-access ordering within opcode execution · not in checklist revision 1
    Changes *where inside an opcode's execution* state is accessed, or where gas is charged relative to that access. Because a state access is recorded in the block-level access list only if execution had enough gas to reach it, this ordering is consensus-critical: moving it changes the BAL at every gas boundary of every affected opcode.
    Score anchors
    0
    No change to where state is accessed, or to where gas is charged relative to a state access, within any opcode.
    1
    A single opcode's state-access or gas-charge ordering changes.
    2
    Multiple opcodes' ordering changes, or a new state-accessing operation is introduced whose position in the order must be settled.
    3
    The ordering rule changes for a whole class of state-accessing opcodes at once, or what counts as a recordable state access is redefined — requiring existing BAL vectors to be re-derived across opcodes and forks.
    • Distinct from "Modified opcodes", which asks whether an opcode's **result** changed. This row asks about the **path to the result**, which is observable even when the result is identical. An EIP can be 0 on that row and 3 on this one.
    • Score changes **to** the ordering. Do not score the fact that state accesses are observable — they always are.
    • Each boundary must be re-tested against every other dimension that can change the answer (cold/warm, static/non-static, delegated/direct, revert/success), so the case count grows multiplicatively rather than additively. Note this explicitly under Special Considerations.
  • Blob gas accounting changes
    New Blob gas accounting rules which potentially affect pre-existing tests
    Score anchors
    0
    No blob gas accounting changes.
    1
    Existing blob gas accounting mechanism is updated.
    2
    A new blob gas accounting mechanism is introduced but it does not affect existing mechanisms nor does it affect existing tests.
    3
    A new blob gas accounting mechanism is introduced and affects existing mechanisms which in turn affect existing tests.
  • State gas accounting changes · not in checklist revision 1
    New state gas accounting rules. State gas is the cost of *writing* state, as opposed to accessing or executing it: `StateGasCosts`, `COST_PER_STATE_BYTE`, the block-level state gas budget, and the spill path into execution gas.
    Score anchors
    0
    No state gas accounting changes.
    1
    An existing state gas cost or `STATE_BYTES_PER_*` rate is adjusted.
    2
    A new state-gas-charging site is introduced, or the block-level state gas budget or reservoir allocation is modified.
    3
    A new state gas charging mechanism is introduced, or the spill interaction between state gas and execution gas is modified, affecting existing gas tests.
    • Harder to test than blob gas: the spill path means state gas cannot be metered independently of execution gas, and some costs (e.g. `NEW_ACCOUNT`) are state-dependent.
  • New EVM gas refund
    New gas-refund mechanism
    Score anchors
    0
    No new gas-refund mechanisms are introduced.
    1
    A new simple gas-refund mechanism is introduced that does not affect either existing tests or existing gas-refund mechanisms.
    2
    A new complex gas-refund mechanism is introduced or a simple mechanism that affects existing tests or existing gas-refund mechanisms.
    3
    A new complex gas-refund mechanism is introduced that affects existing tests or existing gas-refund mechanisms.

Blocks, transactions, and encoding

Transaction types and validity, block and header fields, encodings, syncing, and activation-time changes.

  • New transaction types
    Introduces a new transaction type
    Score anchors
    0
    No new transaction types are introduced.
    3
    A new transaction type is introduced.
  • New or modified transaction validity mechanisms
    Creates new or modifies pre-existing transaction types' validation mechanisms
    Score anchors
    0
    No changes are introduced to the validity rules of existing transaction types or to their intrinsic gas cost calculation.
    1
    Minor adjustments are introduced to validity rules or intrinsic gas cost calculation, but they do not significantly affect existing tests.
    2
    Changes to validity rules or intrinsic gas cost calculation affect existing tests, but require only limited updates to test cases and no redesign of the testing infrastructure.
    3
    Changes to validity rules or intrinsic gas cost calculation require extensive rework or redesign of the tests or testing infrastructure.
  • New block / header fields
    Introduces new block or block header fields
    Score anchors
    0
    No new block or header fields are introduced.
    3
    A new block or header field is introduced.
  • Encoding changes (RLP/SSZ)
    Introduces encoding changes at the transaction/block/interfaces level
    Score anchors
    0
    No encoding changes are introduced at the transaction, block, or interfaces levels.
    3
    An encoding change is introduced at transaction, block or interfaces level (e.g. RLP -> SSZ).
    • "Interfaces level" includes the Engine API. Score an Engine API encoding change (e.g. JSON -> SSZ) here.
  • Block syncing changes
    Modifies block RLP validation mechanisms that require test client syncing.
    Score anchors
    0
    No new RLP validation mechanism is introduced.
    1
    A single simple RLP validation mechanism is introduced.
    2
    Multiple simple RLP validation mechanisms are introduced or a single complex one.
    3
    Multiple RLP validation mechanisms are introduced and at least one of them is deemed complex.
  • New fork activation mechanism
    Modifies state, internal variables, or similar, at the fork activation block
    Score anchors
    0
    No state modifications, internal variables or similar are modified at the fork activation block.
    3
    Either a state modification or internal variables are modified at the fork activation block.
    • Initialization of new internal variable is not considered a modification.

Client interfaces

Engine API and transition-tool interface changes.

  • Engine API changes
    Introduces new fields to the Engine API directives
    Score anchors
    0
    No new fields or communication mechanisms are introduced to the Engine API.
    1
    A single new field is introduced in one of the Engine API endpoints.
    2
    Multiple fields are introduced to one or multiple Engine API end points, or a new Engine API end-point is introduced.
    3
    Multiple fields are introduced to one or multiple Engine API end points and a new Engine API end-point is introduced.
  • Engine API encoding changes · Checklist revision 1 only
    Engine API encoding changes (the revision-1 template defines no anchor text for this row).
  • Transition-tool interface changes
    Modifies or adds new fields to the transition tool interface.
    Score anchors
    0
    No modifications to the transition tool interface are required.
    1
    A single new field needs to be introduced to the transition tool interface.
    2
    Multiple new fields or a new mechanism has to be introduced to the transition tool interface.
    3
    Multiple new fields and a new mechanism has to be introduced to the transition tool interface.
    • Special consideration must be paid to this section if the EIP introduces a mechanism that requires the state transition tool to be aware whether the block it is processing is the fork-activation block.

Testing impact

Rework, new invariants, and new primitives required in the test framework.

  • Patterns affecting pre-existing tests
    Implements a new validation mechanism or rule that translates in reworking pre-existing tests
    Score anchors
    0
    No pre-existing tests are affected by this change.
    1
    Minor subset of existing tests are affected by this change.
    2
    Considerable subset of existing tests are affected by this change but involves only a contrived category of tests.
    3
    Major subset of existing tests are affected, including diverse category of tests (benchmarks, static, multiple forks, etc.).
  • New invariant on pre-existing tests · not in checklist revision 1
    Tests that are **not about this EIP** must nonetheless assert something this EIP produces. Their logic does not change; they gain a new thing to check.
    Score anchors
    0
    Pre-existing tests assert nothing new.
    1
    A narrow, contrived category of pre-existing tests gains a new assertion.
    2
    A broad category gains a new assertion, applied mechanically.
    3
    Every test in the fork gains the assertion regardless of what it tests, and pre-fork vectors must be re-derived to satisfy it.
    • Paired with the row above, and easy to confuse with it. "Patterns affecting pre-existing tests" asks whether existing tests must be **reworked**; this row asks whether they must **additionally assert something new**. Score both — an EIP can be low on one and high on the other.
  • New test-framework primitives · not in checklist revision 1
    Requires new abstractions in the test framework itself — expectation types, modifiers, helpers — beyond writing test functions with what already exists.
    Score anchors
    0
    Existing test primitives suffice.
    1
    Existing primitives need minor extension.
    2
    New expectation or modifier primitives are required, reusable within this EIP's own test suite.
    3
    New framework-level primitives are required that become a permanent part of the framework and are used by other EIPs' tests.

Risk and validation

Security, performance, boundary conditions, and cryptography that need validation.

  • Security risks
    Introduces or modifies mechanisms that could compromise the security of the chain, users, validators, or other stakeholders, if not implemented properly.
    Score anchors
    0
    No new mechanisms are introduced that could pose a security risk.
    1
    The introduced mechanisms are self-contained, can be validated in isolation, and do not alter existing invariants that could pose a security risk for any stakeholders.
    2
    The introduced mechanisms interact with a limited number of existing components, slightly altering their security assumptions and requiring a targeted security review or fuzzing.
    3
    The introduced mechanisms interact with multiple existing components, including critical ones, substantially altering their security assumptions and requiring an extensive security review and fuzzing.
  • Performance risks
    Introduces or modifies mechanisms and requires performance validation.
    Score anchors
    0
    No new mechanisms are introduced that require performance validation.
    1
    The introduced mechanisms can be benchmarked in isolation and do not affect existing performance behavior.
    2
    The introduced mechanisms cannot be fully benchmarked in isolation, but they only have a limited impact on the existing performance benchmarks.
    3
    The introduced mechanisms cannot be benchmarked in isolation and have a substantial impact on existing performance benchmarks or have complex interactions with existing mechanisms.
  • Edge/boundary conditions
    Feature contains edge/boundary conditions.
    Score anchors
    0
    No discernible edge cases or boundary conditions are introduced.
    1
    A single edge-case or boundary-condition prone mechanism is introduced.
    2
    Multiple edge-case or boundary-condition prone mechanisms are introduced, but none of them requires an elevated number of cases to test.
    3
    Multiple edge-case or boundary-condition prone mechanisms are introduced and at least one of them requires an elevated number of cases to test.
  • Cryptography
    Introduces new cryptography mechanisms or modifies existing functionality that involves cryptography
    Score anchors
    0
    No cryptography mechanisms are introduced.
    1
    A new cryptography mechanism is introduced but it is a well known mechanism that is known to have vast resources to aid on its testing.
    2
    Multiple new cryptography mechanisms are introduced that are well-known or a single but novel mechanism is introduced that is either untested or has limited resources.
    3
    Multiple new cryptography mechanisms are introduced and at least one of them is a novel mechanism.

Coordination

Cross-EIP interactions and behavior that clients must agree on before tests exist.

  • Cross-EIP interactions
    Introduces or modifies mechanisms that affect other EIPs in either the same or past forks.
    Score anchors
    0
    Fully self-contained EIP that does not depend on, modify, or conflict with any other EIP.
    1
    The EIP interacts with one or more other EIPs in a non-critical and limited way but can be tested independently for the most part.
    2
    The EIP depends on or modifies one or more other EIPs such that coordinated testing and consideration is required, but interactions are limited in scope and not complex.
    3
    The EIP has strong interdependencies with multiple EIPs, requiring extensive coordinated cross-EIP testing as well as potential re-design of existing test vectors.
    • +1 for every 3 additional interacting EIPs beyond the first 3, each of which requires its own coordinated test cases. List the EIPs in the rationale.
    • This row is intentionally uncapped, unlike every other anchor: each interacting EIP is another axis of the test matrix, so a ceiling would make a 12-EIP product indistinguishable from a 3-EIP one.
  • Unspecified behavior requiring cross-client consensus · not in checklist revision 1
    The EIP text does not determine the answer for cases a test can construct. Clients must agree on a previously unspecified detail before tests can be baselined. The cost here is coordination and re-baselining, not test writing.
    Score anchors
    0
    The EIP text determines the answer for every case a test could construct.
    1
    A few details are unspecified but have an obvious intended reading.
    2
    Details require client agreement before tests can be written, but they are localized.
    3
    A previously unspecified *and previously unobservable* behavior becomes consensus-critical; expect tests to be re-baselined on each round of EIP amendment.
    • Score this from the EIP's state at assessment time: whether it has client implementations, whether it has been through a devnet, and how many open questions remain on its discussion thread.