Empirical Validation & Benchmark Metrics
Source-linked structural validation for four Base fixtures, hermetic parser/proxy tests, and a separate live provider-agreement, freshness, failure, and RPC-path latency benchmark.
NOT_MEASURED.
EVM Disassembly Mechanics: Linear Opcode Walking
Why Substring / Regex Matching Fails on EVM Bytecode
Standard scanners often use naive string searching (indexOf("f4")) to detect DELEGATECALL. This causes catastrophic false positives because data pushed onto the stack via PUSH1..PUSH32 (such as addresses, hashes, or numeric constants) frequently contains 0xF4 as literal data bytes.
M2M Sentinel implements an instruction-by-instruction EVM walker (lib/dissect.js) that records non-executable PUSH operand offsets before scanning for selected opcodes:
Multi-RPC Quorum & Trust Grading Specification
Deterministic On-Chain State Verification on Base Mainnet
To eliminate RPC spoofing or single-node divergence, M2M Sentinel categorizes every analysis response with a cryptographic trust grading:
| Trust Level | Quorum Condition | Evidence Grade | Machine Policy Action |
|---|---|---|---|
| HIGH_TRUST_PRIMARY | TLS 1.3 credentialed connection with verified Chain ID 8453 handshake. | TRUE | Caller policy may use the evidence grade and provenance in its own decision. |
| QUORUM_PUBLIC | Agreement across ≥ 2 independent public Base RPC providers. | TRUE | Consensus validated across disparate nodes. |
| DEGRADED_LOW_TRUST | Lone fallback node or split quorum response. | FALSE | Bot policy can pause or require secondary validation. |
Capability Validation Scope
The published engine currently observes five selected capability types: DELEGATECALL, SELFDESTRUCT, MINT_SELECTOR, PAUSE_SELECTOR, and FREEZE_SELECTOR. It also resolves supported common-proxy structures.
Parser unit fixtures cover exact byte offsets, PUSH-operand masking, CBOR metadata stripping, and selected proxy branches. The small source-linked structural validation report compares selected ABI selectors and common-proxy observations for its stated fixtures.
Recall, false-negative rate, false-positive rate on live contracts, runtime reachability, exploitability, and transaction outcomes are NOT_MEASURED. No BLACKLIST or CREATE2 detector is currently advertised by the API contract.
Hermetic 1,000-Cycle Engine Microbenchmark
Local parser and proxy-resolver execution
This mocked-RPC harness runs 1,000 local engine cycles across 20 labelled fixtures. Real Base addresses are labels only and receive neutral fixture bytecode; synthetic vectors exercise selected branches. These timings exclude HTTP, live RPC, facilitator, and edge-gateway latency.
| Metric | Measured Value | Scope |
|---|---|---|
| Fixture runs | 1,000 |
Mocked RPC, local process |
| Completed fixture assertions | 1,000 |
Agreement with declared mock expectations |
| Mean local execution time | 0.08 ms |
In-process engine only |
| p95 local execution time | 0.23 ms |
In-process engine only |
| Live latency / chain accuracy / safety accuracy | NOT_MEASURED |
Outside hermetic harness scope |
Reproduce the hermetic microbenchmark, then run the separate live provider-agreement and freshness benchmark:
Latest live run (2026-08-20T19:16:37Z): gate failed because only one of three providers returned a complete USDC observation. Across 19 successful and 11 failed measurements, full RPC-path p50 was 410.79 ms and p95 was 2051.58 ms; successful observations agreed structurally and maximum block drift was one. See benchmark.json.
⛓️ Optimism EAS Verified Project Metadata Attestation
Optimism EAS VerifiedProject identity, open-source repository lineage, and payout configurations are verified on-chain via the Ethereum Attestation Service on the Optimism Superchain:
Methodology & Limitations
- The source-linked structural report contains four well-known Base token fixtures; it is not a representative population sample.
- Verified Base Blockscout ABI metadata is used only for three exact selector signatures, and Blockscout proxy metadata is compared for the four fixtures.
- Separate hermetic fixtures cover selected parser and proxy-resolver branches; mock expectations are not live-chain ground truth.
- The Contracts Hub is a 50-address reference catalog, not a 50-case accuracy study.
- Static preflight observations do not measure dynamic runtime state, call reachability, or economic security.
Machine-readable evidence: validation.json and benchmark.json. Separately, explore the 50-address reference catalog in the Curated Reference Contracts Hub →