Retrospective LLM-Based Complexity Evaluations

EIP complexity assessment

EIP-7805: Fork-choice enforced Inclusion Lists (FOCIL)

Assessed in Hegotá. The score describes the EIP text available at the snapshot, not the EIP as it stands today.

ProspectiveHegotáSnapshot 2026-10-07EIP-8081: SFILayers: execution, consensus
LLM Completescore 25
Human Completescore 15 · Checklist revision 1· merged checklist
Other checklist versions (1)

Evaluated on: · Spec revision: 2026-10-07 · 6dac5e7491 · EIP-8081 list: SFI

Scope at the cutoff. EIP-7805 (FOCIL) adds a per-slot committee of 16 validators. Each member gossips a signed inclusion list (IL) of up to 8 KiB of RLP-encoded transactions taken from its mempool view. On the execution side, `engine_newPayload` gains a post-execution check. For every IL transaction not already in the block, the EL checks whether `T.gas <= gas_left` and whether the transaction passes nonce and balance checks against the post-block state. If any such transaction would be valid, the EL returns a new `INCLUSION_LIST_UNSATISFIED` status. The block stays valid, but attesters will not vote for it. The Engine API also gains `engine_getInclusionListV1`, an IL field in `forkchoiceUpdated` `payloadAttributes` for payload building, and an IL-transactions parameter in `newPayload`. There are no EVM, gas, header, transaction-type or system-contract changes; the remaining changes (committee, gossip, equivocation handling and fork choice) are consensus-layer only.

25HighHigh
Evaluator
LLMChecklist v3
Confidence
Medium
Under-specified at assessment cutoff
Yes — 7 criteria affected
Plausible range
19–29 (Medium–High)
Snapshot
2026-10-07 · EIP revision 6dac5e7491 (2026-10-07)
Score bands · Checklist revision 3
  • Low <12
  • Medium 12–22
  • High ≥23

28 criteria scored 0–3 (4 in exceptional cases; cross-EIP interactions is uncapped); nominal maximum 84.

Complexity profile

Each segment is one criterion's contribution to the LLM total. Hover or focus a segment for its score and rationale.

Top complexity drivers

  1. Encoding changes (RLP/SSZ)3
  2. Engine API changes3
  3. Security risks3
  4. Edge/boundary conditions3

Under-specified at assessment cutoff: Yes

The EIP text available at the assessment cutoff left material behavior unresolved. The affected criteria and the plausible total range record that uncertainty.

Why: The EL IL-satisfaction check is only described as 'nonce and balance checks'. The supplied text does not specify: the full validity predicate, how the balance requirement is computed, how presence in the block is determined, how undecodable or duplicate entries are handled, or the exact Engine API schemas and versions. The transition-tool and framework requirements are inferred.

Unresolved questions at the cutoff (6)
  • Does 'balance check' mean value + gas_limit × max_fee_per_gas (plus blob fees), and against which base fee?
  • Do signature, chain ID, max-fee-vs-base-fee, intrinsic-gas and type-specific rules apply when deciding whether a missing IL transaction is 'valid'?
  • How are blob transactions (no sidecars) and undecodable or malformed entries in an IL treated?
  • Is presence in B determined by transaction hash, and how are duplicates across ILs handled?
  • Is the gas-fit check against execution gas only, or also blob gas?
  • What are the exact parameters, field names and versions of engine_getInclusionListV1, the payloadAttributes IL field and the newPayload IL parameter?
Notable ambiguities noted by the assessor (4)
  • IL satisfaction is not a block-validity rule; it is a separate newPayload status. This blurs whether it counts as transaction validity or syncing.
  • The order in which IL transactions are evaluated, and the early-termination rule, could matter if implementations evaluate IL transactions with interdependent effects differently.
  • EIP-7547 is a stagnant alternative design used only for comparison and is not part of the baseline.
  • Builder-side payload update timing ('exact timings will be defined after running some tests/benchmarks') is left open.

Criterion breakdown

EIP-7805 Hegotá: LLM criterion scores and rationale
CriterionScoreWhy this scoreEvidence / uncertainty
Encoding changes (RLP/SSZ)3Engine API serialized schemas change: a new IL field in payloadAttributes, a new IL-transactions parameter in newPayload, and a new response schema for getInclusionListV1.
  • eip.md · Engine API Changes The payloadAttributes schema gains an IL member, and the newPayload parameter list gains IL transactions.
  • eip.md · IL Building — "for all of the RLP encoded transactions" ILs are lists of RLP-encoded transactions exchanged over the Engine API (getInclusionList response).
Confidence: High
Uncertainty: Exact JSON field names and versions are not given in the supplied text.
Engine API changes3Endpoint-level changes (a new method and a new newPayload response status) come together with multiple field changes (the payloadAttributes IL and the newPayload IL parameter).
  • eip.md · Engine API Changes — "Add `engine_getInclusionListV1` endpoint" New endpoint.
  • eip.md · Engine API Changes — "`payloadAttributes` is extended to include the IL" New field in forkchoiceUpdated payloadAttributes.
  • eip.md · Engine API Changes — "Modify `engine_newPayload` ... include a parameter for transactions in ILs ... `INCLUSION_LIST_UNSATISFIED` status" New newPayload parameter and new response-status rule.
Confidence: High
Security risks3A shared validation invariant now spans the EL newPayload check, the EL payload builder and CL fork choice. Clients must agree exactly on IL satisfaction for adversarially crafted ILs (malformed or undecodable transactions, edge nonce/balance cases, dependency chains, equivocation-filtered sets). Otherwise blocks are wrongly rejected or censorship is missed. This needs coordinated adversarial scenarios across components.
  • eip.md · CL P2P Validation Rules — "ILs are allowed to contain any transactions—valid or invalid" The EL receives untrusted, unvalidated transaction bytes through the Engine API.
  • eip.md · Core Properties — Fork-choice enforced The EL's satisfaction result decides whether attesters vote, so divergent EL results across clients split fork choice.
  • eip.md · Consensus Liveness Builders unable to satisfy ILs cannot produce canonical blocks.
Confidence: Medium
Uncertainty: Could be argued as level 2 if the check is treated as a bounded EL/CL interaction.
Edge/boundary conditionsUnder-specified3Several boundary-sensitive mechanisms are introduced: gas-fit against `gas_left`, the post-state nonce check, the post-state balance check, and the 8 KiB IL size bound. The satisfaction outcome depends on interacting dimensions that cannot be tested independently. These are: whether the transaction is already in the block, remaining gas, sender nonce and balance as changed by in-block transactions, multiple IL transactions from the same sender, duplicates across ILs, and multiple ILs.
  • eip.md · Execution Layer — "If `T.gas` > `gas_left`, then jump to the next transaction" Gas-fit boundary (T.gas equal to vs exceeding gas_left).
  • eip.md · Execution Layer — "Validate `T` against `S` by checking the nonce and balance of `T.origin`" Nonce boundary (exact vs off by one) and balance boundary (exactly sufficient vs 1 wei short), both evaluated against the post-block state.
  • eip.md · IL Building — "maximum size of `MAX_BYTES_PER_INCLUSION_LIST = 8 KiB`" Byte-size bound on ILs produced by the EL.
  • eip.md · Payload Construction Dependencies between IL transactions (balance transfers, nonce chains) and in-block changes to sender nonce/balance alter the outcome.
Confidence: Medium
Uncertainty: If gas-fit, nonce and balance are treated as separable conjunctive checks, level 2 would apply.
New or modified transaction validity mechanismsUnder-specified2Block inclusion validity is unchanged, but FOCIL adds a new fork-choice-relevant transaction-validity evaluation: nonce, balance and gas-fit of IL transactions against the post-block state. It needs dedicated cases built from block contents plus IL transactions. Existing validity sequencing can be reused for the nonce and balance checks themselves.
  • eip.md · Execution Layer — "Validate `T` against `S` by checking the nonce and balance of `T.origin`" A new post-block validity evaluation of non-included transactions against the post-execution state; its result is consensus-relevant via fork choice.
  • eip.md · Attesters — "determining if any missing transactions are invalid when appended to the end of the payload" Validity is judged as if the transaction were appended to the end of the payload.
Confidence: Low
Uncertainty: A strict reading (block validity unchanged) gives 0. If coordinated block/IL/state scenarios are judged to need restructured shared construction, it could be 3.
Transition-tool interface changesUnder-specified2To fill IL-satisfaction expectations, the transition tool must accept IL transactions and run a new post-block check: is each transaction present, does `T.gas` fit in `gas_left`, does it pass nonce/balance checks. It must also report the result. That is a new mechanism with at least one new input field. Level 2 is the best-supported score, but adding the output status as a second field would reach level 3.
  • eip.md · Execution Layer — "After all of the transactions in the payload have been executed, we check whether any transaction from ILs..." A post-execution evaluation phase over externally supplied IL transactions against the post-state, needed to derive expected IL-satisfaction outcomes.
  • eip.md · Engine API Changes IL transactions are a new input to payload validation and building.
Confidence: Low
Uncertainty: No transition-tool evidence was supplied. Expected status could instead be hand-specified (lower), or both an IL input and a satisfaction output could be required (level 3).
New test-framework primitivesUnder-specified2Testing needs new abstractions: an IL attachment modifier for blocks/newPayload, a new 'valid but IL-unsatisfied' expectation, and payload-building or getInclusionList tests that depend on mempool state. These are new expectation and construction abstractions within the target's suite (level 2). They do not clearly change how unrelated families are built or checked.
  • eip.md · Engine API Changes — "If the IL is not satisfied an `INCLUSION_LIST_UNSATISFIED` status must be returned." A third newPayload outcome, distinct from VALID and INVALID, must be expressible and checked.
  • eip.md · Engine API Changes — "`payloadAttributes` is extended to include the IL" Payload-building tests must supply ILs and check that the built payload includes valid IL transactions.
  • eip.md · IL Building engine_getInclusionListV1 draws on the mempool, which needs a mempool-driven test setup.
Confidence: Medium
Uncertainty: If the engine fixture format must change for all target-fork fixtures, or a shared mempool/payload-building facility is required, this could reach level 3.
Performance risksUnder-specified2Targeted integrated benchmarks are needed for two bounded interactions. The first is newPayload latency with maximum-size adversarial ILs (many distinct senders forcing cold state reads, invalid signatures). The second is payload-update time in forkchoiceUpdated with ILs under the slot deadline.
  • eip.md · Execution Layer newPayload must decode and recover the sender of each IL transaction and read sender nonce and balance from the post-state, adding to validation latency.
  • eip.md · Preset — IL_COMMITTEE_SIZE 16, MAX_BYTES_PER_INCLUSION_LIST 8192 Bounds the workload: up to 16 × 8 KiB of IL transactions per slot.
  • eip.md · Payload Construction — "`O(n^2)`" A naive payload update is quadratic; builders must update payloads under tight timing (t=11s).
  • eip.md · Consensus Liveness Liveness depends on enough time to update the payload after the IL freeze.
Confidence: Medium
Uncertainty: Coupling with attestation deadlines and mempool-driven IL building could justify level 3.
Cross-EIP interactionsUnder-specified2Coordinated cases are needed where IL-transaction validity depends on behavior defined elsewhere. Examples: in-block account-abstraction or delegated-code transactions changing an IL sender's balance or nonce without the sender's own transaction, and IL transactions of other typed formats such as blob-carrying transactions. The supplied text does not name these EIPs. EIP-7547 is only a comparison design and needs no tests.
  • eip.md · Payload Construction — "when the `balance` changes without a `nonce` increment (e.g., after an Account Abstraction (AA) transaction has interacted with that EOA)" IL validity depends on account-abstraction-style balance changes from other transactions in the block.
  • eip.md · CL P2P Validation Rules — "ILs are allowed to contain any transactions" Any baseline transaction type can appear in an IL, including types defined by other EIPs.
  • supporting/eip-7547.md · Abstract EIP-7547 is a superseded forward-IL design that is only referenced for comparison; it is not part of the baseline.
Confidence: Low
Uncertainty: Interacting EIPs are inferred from unnumbered descriptions. If treated as local compatibility checks, the score would be 1.
Unspecified behavior requiring cross-client consensus2These are localized competing outcomes within the IL-satisfaction check. Clients could disagree on whether a transaction with insufficient max fee or too-low intrinsic gas, a blob transaction, or an undecodable entry makes the IL unsatisfied. Agreement is needed before expected statuses can be fixed.
  • eip.md · Execution Layer — "(i.e. nonce and balance checks pass)" Only nonce and balance are named. The text does not say whether other validity rules apply (signature/chain ID, max fee vs base fee, intrinsic gas, blob transactions without sidecars, transaction-type rules) or how the balance requirement is computed.
  • eip.md · Execution Layer — "Check whether `T` is present in `B`" The presence criterion (hash equality?) and the handling of undecodable or duplicate IL entries are unspecified.
  • eip.md · Engine API Changes Parameters, schema and versioning of the endpoints are not given.
Confidence: Medium
Uncertainty: Linked consensus-spec or execution-API documents might resolve some of these but were not supplied.
Patterns affecting pre-existing testsUnder-specified1Baseline expected results do not change. Only the engine-API form of baseline blockchain tests needs a changed input: an empty or absent IL passed to the new newPayload variant. That is confined to one family (engine newPayload invocation) and is mechanical.
  • eip.md · Engine API Changes — "Modify `engine_newPayload` endpoint to include a parameter for transactions in ILs" Every newPayload call in the target fork carries a new IL parameter.
  • eip.md · Execution Layer — "Although the block is valid" Block validity and execution results are unchanged, so expected state and validity outcomes of baseline tests stay the same.
Confidence: Medium
Uncertainty: If the framework injects an empty IL automatically, no per-test rework is needed (level 0).
Show 17 zero-score criteria
Zero-score criteria (Checklist revision 3)
CriterionScoreWhy this scoreEvidence / uncertainty
Added opcodes0No new opcode.
  • eip.md · Specification No EVM instruction is defined.
Modified opcodes0No opcode semantics or availability change.
  • eip.md · Specification No instruction semantics change.
Added precompiles0No new precompile.
  • eip.md · Specification No precompile is defined.
Modified precompiles0No precompile change.
  • eip.md · Specification No precompile changes.
Added system contracts0No system contract is introduced.
  • eip.md · Specification No protocol-designated contract is defined.
Modified system contracts0No system-contract change.
  • eip.md · Specification No existing system contract is referenced or changed.
EVM Gas rule changes0Remaining block gas is used only as an input to the IL-satisfaction check. No execution-gas accounting rule or parameter changes.
  • eip.md · Execution Layer — "Let `gas_left` be the gas remaining after execution of B." gas_left is only read to decide IL satisfaction; no charging, metering or settlement rule changes.
State-access ordering within opcode execution0No instruction's state-access or gas-charge ordering changes, and no new state-accessing operation is introduced in the EVM.
  • eip.md · Execution Layer The new check runs after all payload transactions have executed and does not touch opcode execution.
Blob gas accounting changes0There are no blob-gas charging, pricing or limit changes. How blob transactions in ILs are treated is a validity or specification question, not accounting.
  • eip.md · Execution Layer — "checking the nonce and balance of `T.origin`" Only nonce and balance checks are applied to IL transactions; blob gas pricing and limits are untouched.
Uncertainty: The text does not say whether the balance check includes blob fees or whether the gas check covers blob gas. That is recorded under UNSP, not here.
State gas accounting changes0State-gas accounting does not change.
  • eip.md · Specification No state-write cost or state-gas rule is defined.
New EVM gas refund0No new refund mechanism.
  • eip.md · Core Properties — "No incentive mechanism" No rewards or refunds are introduced; no refund rule appears in the specification.
New transaction types0No new transaction type.
  • eip.md · Specification No new EIP-2718 transaction envelope is defined.
New block / header fields0No EL block-level or header member is added. The IL is an API-only input.
  • eip.md · Execution Layer ILs are passed via the Engine API and are not committed in the execution block or header.
Block syncing changes0There is no execution-block decoding or structural validation change. The IL check affects fork-choice attestation, not block import.
  • eip.md · Execution Layer — "Although the block is valid, the CL will not attest to it." IL satisfaction is not a block validity rule. The block's RLP and structural validation are unchanged.
New fork activation mechanism0No activation-specific EL state transition.
  • eip.md · Backwards Compatibility A hard fork is needed for CL validation changes. No EL state migration or code installation is described.
New invariant on pre-existing tests0Baseline tests gain no new output to assert. IL-satisfaction status matters only for tests that supply ILs, which are new feature tests.
  • eip.md · Execution Layer With no IL supplied, the only new output (INCLUSION_LIST_UNSATISFIED) cannot occur. No header, receipt or log output is added.
Cryptography0The EL does not execute any new or changed cryptographic verification, signing or hashing rule. IL signature verification is confined to the CL.
  • eip.md · CL P2P Validation Rules — "the only nontrivial check performed on IL propagation is signature verification" The new IL signatures are BLS signatures checked by the CL. On the EL, sender recovery of transactions is unchanged.
  • eip.md · New containers — SignedInclusionList The signed IL container is a CL object.
Assessment provenance
Assessed EIP revision
ethereum/EIPs@6dac5e7491 EIPS/eip-7805.md committed 2026-10-07 · information cutoff 2026-10-07T22:23:55Z
Current master · File history · blob 0a3955d9e8 · sha256 81096dd37cab
Rubric
Checklist revision 3 · ethspecs/pm@fe2f793b03
Evaluator
Opus 5.5 (claude-opus-5-5) at high effort, one tool-less call per EIP · isolation bubblewrap_claude_p_no_tools_v1
Source record
Frozen research record research/tasks/10-opus-v3-reassessment/prospective/outputs/assessments/hegota-2026-10-08/eip-7805.yaml · sha256 919a70bea4ef
Supporting documents supplied with the EIP
supporting/eip-7547.md

Evaluated on: Not recorded

15MediumMedium
Evaluator
HumanChecklist v1
Confidence
Not recorded
Under-specified at assessment cutoff
Not recorded in the checklist
Checklist published
2026-01-12
Score bands · Checklist revision 1
  • Low <10
  • Medium 10–19
  • High ≥20

24 criteria scored 0–3 (4 in exceptional cases; cross-EIP interactions is uncapped); nominal maximum 72.

Complexity profile

Each segment is one criterion's contribution to the Human total. Hover or focus a segment for its score and rationale.

Top complexity drivers

  1. Engine API changes3
  2. New or modified transaction validity mechanisms2
  3. Transition-tool interface changes2
  4. Security risks2

Criterion breakdown

EIP-7805 Hegotá: Human criterion scores and rationale
CriterionScoreWhy this scoreNotes
Engine API changes32 new engine apis included, 1 modified—
New or modified transaction validity mechanisms2Need to validate each transation in the IL, but limited in scope and straightforward checks—
Transition-tool interface changes2new field, new mechanism for tx validation, new error—
Security risks2Consensus-critical validations, new economic incentives—
Performance risks2Each transaction not in block requires state access, adding latency to block validation time—
Edge/boundary conditions2Many boundary/edge cases to test, but standard checks—
Patterns affecting pre-existing tests1"After all of the transactions in the payload have been executed, we check whether any transaction from ILs, that is not already present in the payload, could be validly included...". Additionally, there is further validation for each Transaction in the IL—
Cross-EIP interactions1Interacts with AA—
Show 16 zero-score criteria
Zero-score criteria (Checklist revision 1)
CriterionScoreWhy this scoreNotes
Added opcodes0No rationale recorded.—
Modified opcodes0No rationale recorded.—
Added precompiles0No rationale recorded.—
Modified precompiles0No rationale recorded.—
Added system contracts0No rationale recorded.—
Modified system contracts0No rationale recorded.—
EVM Gas rule changes0No rationale recorded.—
Blob gas accounting changes0No rationale recorded.—
New EVM gas refund0No rationale recorded.—
New transaction types0No rationale recorded.—
New block / header fields0No rationale recorded.—
Encoding changes (RLP/SSZ)0No rationale recorded.—
Block syncing changes0No rationale recorded.—
New fork activation mechanism0No rationale recorded.—
Engine API encoding changes0No rationale recorded.—
Cryptography0No rationale recorded.—
Assessment provenance
Rubric
Checklist revision 1 · ethspecs/pm@d936bcb349
Evaluator
STEEL team · ethspecs/pm complexity_assessments
Source record
Merged checklist ethspecs/pm@3d8c0128c5 complexity_assessments/EIPs/EIP-7805.md · committed 2026-01-12
blob b452b3c506 · sha256 ed018ce5a7fc
Research record
research/tasks/09-hegota-human-assessment-snapshot/outputs/assessments/eip-7805.yaml · sha256 0a29be7f88ae

The LLM applied checklist revision 3 and the human reviewers revision 1 to EIP-7805 in Hegotá. Revision 3 phrases the same criteria more precisely; differences cover the 23 criteria both revisions share, and each total keeps its own revision. Δ is LLM minus Human.

Using the latest scored LLM evaluation for this checklist: 2026-10-08 · spec 2026-10-07 · 6dac5e7491. The Human and LLM assessments may use different spec revisions.

LLM25High
Human15Medium
Δ total+10Tiers differ: High vs Medium
Criteria19/23agree exactly · 3 differ by 1 · 1 differ by 2+

Complexity profiles side by side

LLM
Human

Largest disagreements: Encoding changes (RLP/SSZ) (+3), Edge/boundary conditions (+1), Security risks (+1), Cross-EIP interactions (+1)

Per-criterion scores, Human versus LLM, ordered by the size of the difference
CriterionLLMHumanΔAgreementRationale from each source
Encoding changes (RLP/SSZ)30+3Differ by 2+
Show rationale

LLM Engine API serialized schemas change: a new IL field in payloadAttributes, a new IL-transactions parameter in newPayload, and a new response schema for getInclusionListV1.

Human No rationale recorded.

Security risks32+1Differ by 1
Show rationale

LLM A shared validation invariant now spans the EL newPayload check, the EL payload builder and CL fork choice. Clients must agree exactly on IL satisfaction for adversarially crafted ILs (malformed or undecodable transactions, edge nonce/balance cases, dependency chains, equivocation-filtered sets). Otherwise blocks are wrongly rejected or censorship is missed. This needs coordinated adversarial scenarios across components.

Human Consensus-critical validations, new economic incentives

Edge/boundary conditions32+1Differ by 1
Show rationale

LLM Several boundary-sensitive mechanisms are introduced: gas-fit against `gas_left`, the post-state nonce check, the post-state balance check, and the 8 KiB IL size bound. The satisfaction outcome depends on interacting dimensions that cannot be tested independently. These are: whether the transaction is already in the block, remaining gas, sender nonce and balance as changed by in-block transactions, multiple IL transactions from the same sender, duplicates across ILs, and multiple ILs.

Human Many boundary/edge cases to test, but standard checks

Cross-EIP interactions21+1Differ by 1
Show rationale

LLM Coordinated cases are needed where IL-transaction validity depends on behavior defined elsewhere. Examples: in-block account-abstraction or delegated-code transactions changing an IL sender's balance or nonce without the sender's own transaction, and IL transactions of other typed formats such as blob-carrying transactions. The supplied text does not name these EIPs. EIP-7547 is only a comparison design and needs no tests.

Human Interacts with AA

Added opcodes000Agree
Show rationale

LLM No new opcode.

Human No rationale recorded.

Modified opcodes000Agree
Show rationale

LLM No opcode semantics or availability change.

Human No rationale recorded.

Added precompiles000Agree
Show rationale

LLM No new precompile.

Human No rationale recorded.

Modified precompiles000Agree
Show rationale

LLM No precompile change.

Human No rationale recorded.

Added system contracts000Agree
Show rationale

LLM No system contract is introduced.

Human No rationale recorded.

Modified system contracts000Agree
Show rationale

LLM No system-contract change.

Human No rationale recorded.

EVM Gas rule changes000Agree
Show rationale

LLM Remaining block gas is used only as an input to the IL-satisfaction check. No execution-gas accounting rule or parameter changes.

Human No rationale recorded.

Blob gas accounting changes000Agree
Show rationale

LLM There are no blob-gas charging, pricing or limit changes. How blob transactions in ILs are treated is a validity or specification question, not accounting.

Human No rationale recorded.

New EVM gas refund000Agree
Show rationale

LLM No new refund mechanism.

Human No rationale recorded.

New transaction types000Agree
Show rationale

LLM No new transaction type.

Human No rationale recorded.

New or modified transaction validity mechanisms220Agree
Show rationale

LLM Block inclusion validity is unchanged, but FOCIL adds a new fork-choice-relevant transaction-validity evaluation: nonce, balance and gas-fit of IL transactions against the post-block state. It needs dedicated cases built from block contents plus IL transactions. Existing validity sequencing can be reused for the nonce and balance checks themselves.

Human Need to validate each transation in the IL, but limited in scope and straightforward checks

New block / header fields000Agree
Show rationale

LLM No EL block-level or header member is added. The IL is an API-only input.

Human No rationale recorded.

Block syncing changes000Agree
Show rationale

LLM There is no execution-block decoding or structural validation change. The IL check affects fork-choice attestation, not block import.

Human No rationale recorded.

New fork activation mechanism000Agree
Show rationale

LLM No activation-specific EL state transition.

Human No rationale recorded.

Engine API changes330Agree
Show rationale

LLM Endpoint-level changes (a new method and a new newPayload response status) come together with multiple field changes (the payloadAttributes IL and the newPayload IL parameter).

Human 2 new engine apis included, 1 modified

Transition-tool interface changes220Agree
Show rationale

LLM To fill IL-satisfaction expectations, the transition tool must accept IL transactions and run a new post-block check: is each transaction present, does `T.gas` fit in `gas_left`, does it pass nonce/balance checks. It must also report the result. That is a new mechanism with at least one new input field. Level 2 is the best-supported score, but adding the output status as a second field would reach level 3.

Human new field, new mechanism for tx validation, new error

Patterns affecting pre-existing tests110Agree
Show rationale

LLM Baseline expected results do not change. Only the engine-API form of baseline blockchain tests needs a changed input: an empty or absent IL passed to the new newPayload variant. That is confined to one family (engine newPayload invocation) and is mechanical.

Human "After all of the transactions in the payload have been executed, we check whether any transaction from ILs, that is not already present in the payload, could be validly included...". Additionally, there is further validation for each Transaction in the IL

Performance risks220Agree
Show rationale

LLM Targeted integrated benchmarks are needed for two bounded interactions. The first is newPayload latency with maximum-size adversarial ILs (many distinct senders forcing cold state reads, invalid signatures). The second is payload-update time in forkchoiceUpdated with ILs under the slot deadline.

Human Each transaction not in block requires state access, adding latency to block validation time

Cryptography000Agree
Show rationale

LLM The EL does not execute any new or changed cryptographic verification, signing or hashing rule. IL signature verification is confined to the CL.

Human No rationale recorded.

State-access ordering within opcode execution0n/a—Only in revision 3
Show rationale

LLM No instruction's state-access or gas-charge ordering changes, and no new state-accessing operation is introduced in the EVM.

Human No rationale recorded.

State gas accounting changes0n/a—Only in revision 3
Show rationale

LLM State-gas accounting does not change.

Human No rationale recorded.

Engine API encoding changesn/a0—Only in revision 1—
New invariant on pre-existing tests0n/a—Only in revision 3
Show rationale

LLM Baseline tests gain no new output to assert. IL-satisfaction status matters only for tests that supply ILs, which are new feature tests.

Human No rationale recorded.

New test-framework primitives2n/a—Only in revision 3
Show rationale

LLM Testing needs new abstractions: an IL attachment modifier for blocks/newPayload, a new 'valid but IL-unsatisfied' expectation, and payload-building or getInclusionList tests that depend on mempool state. These are new expectation and construction abstractions within the target's suite (level 2). They do not clearly change how unrelated families are built or checked.

Human No rationale recorded.

Unspecified behavior requiring cross-client consensus2n/a—Only in revision 3
Show rationale

LLM These are localized competing outcomes within the IL-satisfaction check. Clients could disagree on whether a transaction with insufficient max fee or too-low intrinsic gas, a blob transaction, or an undecodable entry makes the IL unsatisfied. Agreement is needed before expected statuses can be fixed.

Human No rationale recorded.

Criterion legend and glossary

Every stacked bar, comparison matrix, and criterion table on this site uses the same criterion colours, abbreviations, and order. Colour marks the criterion group; the abbreviation and name identify the criterion. Scores are 0–3 per criterion (4 is exceptional; cross-EIP interactions is uncapped).

EVM surface

Opcodes, precompiles, and system contracts that are added or modified.

  • Added opcodes
    Introduces new opcodes
    Score anchors
    0
    No new opcodes are introduced.
    1
    A new simple opcode is introduced (no data portion, no complex stack mechanics, and a constant gas cost).
    2
    Multiple new simple opcodes are introduced, or a single new complex opcode is introduced (has data portion, or complex stack mechanics, or a dynamic gas cost).
    3
    Multiple new opcodes are introduced, and at least one of them is complex (has data portion, or complex stack mechanics, or a dynamic gas cost).
    • Cryptography opcodes are not considered complex by default. Refer to the "Cryptography" section for a separate assessment.
  • Modified opcodes
    Modifies pre-existing opcodes
    Score anchors
    0
    No pre-existing opcode modifications are introduced.
    3
    At least one pre-existing opcode's behavior is modified (not including gas changes) or a pre-existing opcode is deprecated.
  • Added precompiles
    Introduces new precompiles
    Score anchors
    0
    No new precompiles are introduced.
    1
    A new simple precompile is introduced (constant input length, constant gas cost).
    2
    Multiple new simple precompiles are introduced, or a single new complex precompile is introduced (dynamic input length or dynamic gas cost).
    3
    Multiple new precompiles are introduced, and at least one of them is complex (dynamic input length or dynamic gas cost).
    • Cryptography precompiles are not considered complex by default. Refer to the "Cryptography" for a separate assessment.
  • Modified precompiles
    Modifies pre-existing precompiles logic or gas-accounting
    Score anchors
    0
    No pre-existing precompiles are modified.
    1
    At least one pre-existing precompile has its gas schedule modified.
    2
    Multiple pre-existing precompiles have their gas schedule modified, or a single pre-existing precompile has its behavior modified.
    3
    The behavior of multiple pre-existing precompiles, or a single complex pre-existing precompile modified.
  • Added system contracts
    Introduces new system contract, stateful or not
    Score anchors
    0
    No new system contracts are introduced.
    1
    A new system contract is introduced that is not stateful nor does it trigger a new system action (e.g. requests to the consensus layer).
    2
    Multiple new system contracts are introduced or a single new system contract that is either stateful or triggers a new system action (e.g. requests to the consensus layer).
    3
    Multiple new system contracts are introduced and at least one of them is either stateful or triggers a new system action (e.g. requests to the consensus layer).
  • Modified system contracts
    Modifies pre-existing system contracts
    Score anchors
    0
    No modifications to pre-existing system contracts are introduced, directly or indirectly.
    1
    Does not directly modify any system contract, but its behavior has minor indirect effects on one or more system contracts.
    2
    Does not directly modify any system contract, but its behavior has major indirect effects on one or more system contracts.
    3
    At least one pre-existing system contract code or state is modified, which would involve irregular state transition or a similarly complex transition methodology.

Gas and accounting

Execution, blob, and state gas rules, refunds, and where charges happen inside opcodes.

  • EVM Gas rule changes
    New EVM gas accounting rules
    Score anchors
    0
    No gas accounting changes.
    1
    Existing gas accounting mechanism is updated.
    2
    A new gas accounting mechanism is introduced but it does not affect existing mechanisms nor does it affect existing tests.
    3
    A new gas accounting mechanism is introduced and affects existing mechanisms which in turn affect existing tests.
  • State-access ordering within opcode execution · not in checklist revision 1
    Changes *where inside an opcode's execution* state is accessed, or where gas is charged relative to that access. Because a state access is recorded in the block-level access list only if execution had enough gas to reach it, this ordering is consensus-critical: moving it changes the BAL at every gas boundary of every affected opcode.
    Score anchors
    0
    No change to where state is accessed, or to where gas is charged relative to a state access, within any opcode.
    1
    A single opcode's state-access or gas-charge ordering changes.
    2
    Multiple opcodes' ordering changes, or a new state-accessing operation is introduced whose position in the order must be settled.
    3
    The ordering rule changes for a whole class of state-accessing opcodes at once, or what counts as a recordable state access is redefined — requiring existing BAL vectors to be re-derived across opcodes and forks.
    • Distinct from "Modified opcodes", which asks whether an opcode's **result** changed. This row asks about the **path to the result**, which is observable even when the result is identical. An EIP can be 0 on that row and 3 on this one.
    • Score changes **to** the ordering. Do not score the fact that state accesses are observable — they always are.
    • Each boundary must be re-tested against every other dimension that can change the answer (cold/warm, static/non-static, delegated/direct, revert/success), so the case count grows multiplicatively rather than additively. Note this explicitly under Special Considerations.
  • Blob gas accounting changes
    New Blob gas accounting rules which potentially affect pre-existing tests
    Score anchors
    0
    No blob gas accounting changes.
    1
    Existing blob gas accounting mechanism is updated.
    2
    A new blob gas accounting mechanism is introduced but it does not affect existing mechanisms nor does it affect existing tests.
    3
    A new blob gas accounting mechanism is introduced and affects existing mechanisms which in turn affect existing tests.
  • State gas accounting changes · not in checklist revision 1
    New state gas accounting rules. State gas is the cost of *writing* state, as opposed to accessing or executing it: `StateGasCosts`, `COST_PER_STATE_BYTE`, the block-level state gas budget, and the spill path into execution gas.
    Score anchors
    0
    No state gas accounting changes.
    1
    An existing state gas cost or `STATE_BYTES_PER_*` rate is adjusted.
    2
    A new state-gas-charging site is introduced, or the block-level state gas budget or reservoir allocation is modified.
    3
    A new state gas charging mechanism is introduced, or the spill interaction between state gas and execution gas is modified, affecting existing gas tests.
    • Harder to test than blob gas: the spill path means state gas cannot be metered independently of execution gas, and some costs (e.g. `NEW_ACCOUNT`) are state-dependent.
  • New EVM gas refund
    New gas-refund mechanism
    Score anchors
    0
    No new gas-refund mechanisms are introduced.
    1
    A new simple gas-refund mechanism is introduced that does not affect either existing tests or existing gas-refund mechanisms.
    2
    A new complex gas-refund mechanism is introduced or a simple mechanism that affects existing tests or existing gas-refund mechanisms.
    3
    A new complex gas-refund mechanism is introduced that affects existing tests or existing gas-refund mechanisms.

Blocks, transactions, and encoding

Transaction types and validity, block and header fields, encodings, syncing, and activation-time changes.

  • New transaction types
    Introduces a new transaction type
    Score anchors
    0
    No new transaction types are introduced.
    3
    A new transaction type is introduced.
  • New or modified transaction validity mechanisms
    Creates new or modifies pre-existing transaction types' validation mechanisms
    Score anchors
    0
    No changes are introduced to the validity rules of existing transaction types or to their intrinsic gas cost calculation.
    1
    Minor adjustments are introduced to validity rules or intrinsic gas cost calculation, but they do not significantly affect existing tests.
    2
    Changes to validity rules or intrinsic gas cost calculation affect existing tests, but require only limited updates to test cases and no redesign of the testing infrastructure.
    3
    Changes to validity rules or intrinsic gas cost calculation require extensive rework or redesign of the tests or testing infrastructure.
  • New block / header fields
    Introduces new block or block header fields
    Score anchors
    0
    No new block or header fields are introduced.
    3
    A new block or header field is introduced.
  • Encoding changes (RLP/SSZ)
    Introduces encoding changes at the transaction/block/interfaces level
    Score anchors
    0
    No encoding changes are introduced at the transaction, block, or interfaces levels.
    3
    An encoding change is introduced at transaction, block or interfaces level (e.g. RLP -> SSZ).
    • "Interfaces level" includes the Engine API. Score an Engine API encoding change (e.g. JSON -> SSZ) here.
  • Block syncing changes
    Modifies block RLP validation mechanisms that require test client syncing.
    Score anchors
    0
    No new RLP validation mechanism is introduced.
    1
    A single simple RLP validation mechanism is introduced.
    2
    Multiple simple RLP validation mechanisms are introduced or a single complex one.
    3
    Multiple RLP validation mechanisms are introduced and at least one of them is deemed complex.
  • New fork activation mechanism
    Modifies state, internal variables, or similar, at the fork activation block
    Score anchors
    0
    No state modifications, internal variables or similar are modified at the fork activation block.
    3
    Either a state modification or internal variables are modified at the fork activation block.
    • Initialization of new internal variable is not considered a modification.

Client interfaces

Engine API and transition-tool interface changes.

  • Engine API changes
    Introduces new fields to the Engine API directives
    Score anchors
    0
    No new fields or communication mechanisms are introduced to the Engine API.
    1
    A single new field is introduced in one of the Engine API endpoints.
    2
    Multiple fields are introduced to one or multiple Engine API end points, or a new Engine API end-point is introduced.
    3
    Multiple fields are introduced to one or multiple Engine API end points and a new Engine API end-point is introduced.
  • Engine API encoding changes · Checklist revision 1 only
    Engine API encoding changes (the revision-1 template defines no anchor text for this row).
  • Transition-tool interface changes
    Modifies or adds new fields to the transition tool interface.
    Score anchors
    0
    No modifications to the transition tool interface are required.
    1
    A single new field needs to be introduced to the transition tool interface.
    2
    Multiple new fields or a new mechanism has to be introduced to the transition tool interface.
    3
    Multiple new fields and a new mechanism has to be introduced to the transition tool interface.
    • Special consideration must be paid to this section if the EIP introduces a mechanism that requires the state transition tool to be aware whether the block it is processing is the fork-activation block.

Testing impact

Rework, new invariants, and new primitives required in the test framework.

  • Patterns affecting pre-existing tests
    Implements a new validation mechanism or rule that translates in reworking pre-existing tests
    Score anchors
    0
    No pre-existing tests are affected by this change.
    1
    Minor subset of existing tests are affected by this change.
    2
    Considerable subset of existing tests are affected by this change but involves only a contrived category of tests.
    3
    Major subset of existing tests are affected, including diverse category of tests (benchmarks, static, multiple forks, etc.).
  • New invariant on pre-existing tests · not in checklist revision 1
    Tests that are **not about this EIP** must nonetheless assert something this EIP produces. Their logic does not change; they gain a new thing to check.
    Score anchors
    0
    Pre-existing tests assert nothing new.
    1
    A narrow, contrived category of pre-existing tests gains a new assertion.
    2
    A broad category gains a new assertion, applied mechanically.
    3
    Every test in the fork gains the assertion regardless of what it tests, and pre-fork vectors must be re-derived to satisfy it.
    • Paired with the row above, and easy to confuse with it. "Patterns affecting pre-existing tests" asks whether existing tests must be **reworked**; this row asks whether they must **additionally assert something new**. Score both — an EIP can be low on one and high on the other.
  • New test-framework primitives · not in checklist revision 1
    Requires new abstractions in the test framework itself — expectation types, modifiers, helpers — beyond writing test functions with what already exists.
    Score anchors
    0
    Existing test primitives suffice.
    1
    Existing primitives need minor extension.
    2
    New expectation or modifier primitives are required, reusable within this EIP's own test suite.
    3
    New framework-level primitives are required that become a permanent part of the framework and are used by other EIPs' tests.

Risk and validation

Security, performance, boundary conditions, and cryptography that need validation.

  • Security risks
    Introduces or modifies mechanisms that could compromise the security of the chain, users, validators, or other stakeholders, if not implemented properly.
    Score anchors
    0
    No new mechanisms are introduced that could pose a security risk.
    1
    The introduced mechanisms are self-contained, can be validated in isolation, and do not alter existing invariants that could pose a security risk for any stakeholders.
    2
    The introduced mechanisms interact with a limited number of existing components, slightly altering their security assumptions and requiring a targeted security review or fuzzing.
    3
    The introduced mechanisms interact with multiple existing components, including critical ones, substantially altering their security assumptions and requiring an extensive security review and fuzzing.
  • Performance risks
    Introduces or modifies mechanisms and requires performance validation.
    Score anchors
    0
    No new mechanisms are introduced that require performance validation.
    1
    The introduced mechanisms can be benchmarked in isolation and do not affect existing performance behavior.
    2
    The introduced mechanisms cannot be fully benchmarked in isolation, but they only have a limited impact on the existing performance benchmarks.
    3
    The introduced mechanisms cannot be benchmarked in isolation and have a substantial impact on existing performance benchmarks or have complex interactions with existing mechanisms.
  • Edge/boundary conditions
    Feature contains edge/boundary conditions.
    Score anchors
    0
    No discernible edge cases or boundary conditions are introduced.
    1
    A single edge-case or boundary-condition prone mechanism is introduced.
    2
    Multiple edge-case or boundary-condition prone mechanisms are introduced, but none of them requires an elevated number of cases to test.
    3
    Multiple edge-case or boundary-condition prone mechanisms are introduced and at least one of them requires an elevated number of cases to test.
  • Cryptography
    Introduces new cryptography mechanisms or modifies existing functionality that involves cryptography
    Score anchors
    0
    No cryptography mechanisms are introduced.
    1
    A new cryptography mechanism is introduced but it is a well known mechanism that is known to have vast resources to aid on its testing.
    2
    Multiple new cryptography mechanisms are introduced that are well-known or a single but novel mechanism is introduced that is either untested or has limited resources.
    3
    Multiple new cryptography mechanisms are introduced and at least one of them is a novel mechanism.

Coordination

Cross-EIP interactions and behavior that clients must agree on before tests exist.

  • Cross-EIP interactions
    Introduces or modifies mechanisms that affect other EIPs in either the same or past forks.
    Score anchors
    0
    Fully self-contained EIP that does not depend on, modify, or conflict with any other EIP.
    1
    The EIP interacts with one or more other EIPs in a non-critical and limited way but can be tested independently for the most part.
    2
    The EIP depends on or modifies one or more other EIPs such that coordinated testing and consideration is required, but interactions are limited in scope and not complex.
    3
    The EIP has strong interdependencies with multiple EIPs, requiring extensive coordinated cross-EIP testing as well as potential re-design of existing test vectors.
    • +1 for every 3 additional interacting EIPs beyond the first 3, each of which requires its own coordinated test cases. List the EIPs in the rationale.
    • This row is intentionally uncapped, unlike every other anchor: each interacting EIP is another axis of the test matrix, so a ceiling would make a 12-EIP product indistinguishable from a 3-EIP one.
  • Unspecified behavior requiring cross-client consensus · not in checklist revision 1
    The EIP text does not determine the answer for cases a test can construct. Clients must agree on a previously unspecified detail before tests can be baselined. The cost here is coordination and re-baselining, not test writing.
    Score anchors
    0
    The EIP text determines the answer for every case a test could construct.
    1
    A few details are unspecified but have an obvious intended reading.
    2
    Details require client agreement before tests can be written, but they are localized.
    3
    A previously unspecified *and previously unobservable* behavior becomes consensus-critical; expect tests to be re-baselined on each round of EIP amendment.
    • Score this from the EIP's state at assessment time: whether it has client implementations, whether it has been through a devnet, and how many open questions remain on its discussion thread.