Empirical Validation & Benchmark Metrics

Source-linked structural validation for four Base fixtures, hermetic parser/proxy tests, and a separate live provider-agreement, freshness, failure, and RPC-path latency benchmark.

What this benchmark does and does not measure: This benchmark measures selected opcode capability observations, selector extraction, and proxy pattern agreement. It does not label contracts safe, malicious, or financially exploitable. Safety classification accuracy is NOT_MEASURED.
4Source-Linked Structural Fixtures
3Exact ABI Selectors Compared
NOT MEASUREDSafety Classification Accuracy
331.91 msLive Full-Path p50 · 2026-08-24
937.90 msLive Full-Path p95 · 2026-08-24
PASSEDLive Provider Gate · 2026-08-24 · 30/30 measurements

EVM Disassembly Mechanics: Linear Opcode Walking

Why Substring / Regex Matching Fails on EVM Bytecode

Standard scanners often use naive string searching (indexOf("f4")) to detect DELEGATECALL. This causes catastrophic false positives because data pushed onto the stack via PUSH1..PUSH32 (such as addresses, hashes, or numeric constants) frequently contains 0xF4 as literal data bytes.

M2M Sentinel implements an instruction-by-instruction EVM walker (lib/dissect.js) that records non-executable PUSH operand offsets before scanning for selected opcodes:

[Raw Bytecode Hex] ──► [buildPushDataOffsets] ──► Set<Non-Executable Byte Offsets> │ ├── Opcode Scan (0xF4 DELEGATECALL, 0xFF SELFDESTRUCT) │ └── If offset in PushData Set: Ignored (Literal Data) │ └── If offset outside PushData: Confirmed Executable Opcode └── Selector Extraction (Function Dispatch Table Jump Destinations)

Multi-RPC Quorum & Trust Grading Specification

Deterministic On-Chain State Verification on Base Mainnet

To eliminate RPC spoofing or single-node divergence, M2M Sentinel categorizes every analysis response with a cryptographic trust grading:

Trust Level Quorum Condition Evidence Grade Machine Policy Action
HIGH_TRUST_PRIMARY TLS 1.3 credentialed connection with verified Chain ID 8453 handshake. TRUE Caller policy may use the evidence grade and provenance in its own decision.
QUORUM_PUBLIC Agreement across ≥ 2 independent public Base RPC providers. TRUE Consensus validated across disparate nodes.
DEGRADED_LOW_TRUST Lone fallback node or split quorum response. FALSE Bot policy can pause or require secondary validation.

Capability Validation Scope

The published engine currently observes five selected capability types: DELEGATECALL, SELFDESTRUCT, MINT_SELECTOR, PAUSE_SELECTOR, and FREEZE_SELECTOR. It also resolves supported common-proxy structures.

Parser unit fixtures cover exact byte offsets, PUSH-operand masking, CBOR metadata stripping, and selected proxy branches. The small source-linked structural validation report compares selected ABI selectors and common-proxy observations for its stated fixtures.

Recall, false-negative rate, false-positive rate on live contracts, runtime reachability, exploitability, and transaction outcomes are NOT_MEASURED. No BLACKLIST or CREATE2 detector is currently advertised by the API contract.

Hermetic 1,000-Cycle Engine Microbenchmark

Local parser and proxy-resolver execution

This mocked-RPC harness runs 1,000 local engine cycles across 20 labelled fixtures. Real Base addresses are labels only and receive neutral fixture bytecode; synthetic vectors exercise selected branches. These timings exclude HTTP, live RPC, facilitator, and edge-gateway latency.

Metric Measured Value Scope
Fixture runs 1,000 Mocked RPC, local process
Completed fixture assertions 1,000 Agreement with declared mock expectations
Mean local execution time 0.08 ms In-process engine only
p95 local execution time 0.23 ms In-process engine only
Live latency / chain accuracy / safety accuracy NOT_MEASURED Outside hermetic harness scope

Reproduce the hermetic microbenchmark, then run the separate live provider-agreement and freshness benchmark:

git clone https://github.com/M2M-Sentinel/m2m-sentinel-sdk.git npm install node scripts/benchmark_agent_decisions.js npm run benchmark:latency

Latest live run (2026-08-20T19:16:37Z): gate failed because only one of three providers returned a complete USDC observation. Across 19 successful and 11 failed measurements, full RPC-path p50 was 410.79 ms and p95 was 2051.58 ms; successful observations agreed structurally and maximum block drift was one. See benchmark.json.

⛓️ Optimism EAS Verified Project Metadata Attestation

Optimism EAS Verified

Project identity, open-source repository lineage, and payout configurations are verified on-chain via the Ethereum Attestation Service on the Optimism Superchain:

EAS Registry: 0x4200000000000000000000000000000000000021 ↗
Attestation UID: 0x3fcc74a0...24b05bf2 ↗
Schema: Schema #472 (Retro Funding Project Metadata Snapshot)
Project / Payout Recipient: 0x6d6c398390cfb88f1cd42715b84906a0bd6652aa
Note: This EAS record contains project-declared metadata and repository/payout references. It is not an independent technical audit. Bytecode observations are computed from sampled RPC evidence.

Methodology & Limitations

Machine-readable evidence: validation.json and benchmark.json. Separately, explore the 50-address reference catalog in the Curated Reference Contracts Hub →