# Agent integration benchmark protocol

Use fresh projects with only the public wheel/source, docs, AGENTS.md and bundled skill/MCP configurations. Deterministic client coverage includes installation, tool discovery, address/materialize/inspect/provenance, benchmark, invalid-input behavior and built-in contract validation. External key/value, relational and graph examples use public protocols without Core edits.

The private final audit records measured install/first-result/benchmark times, tool names, adapter LOC, commands, errors, warnings and interventions for two fresh environments. A deterministic client's zero interventions or correct tool sequence is not an autonomous model's success rate. Unsupported-claim rate is evaluated only where actual generated reports are available, not inferred from assertions.

When safe local agent execution is available, a separate agent receives only the public package/docs and a fresh sample project, creates an adapter, measures equivalence/bytes/latency, and reports provenance. Its report, actions and limitations are audited separately. No model or platform compatibility claim is made beyond observed execution.

Public toy metrics are stable apart from elapsed time. RC1's private 155-case benchmark is not re-executed during this exercise. All factual summaries include scope and the approximately +0.554 seconds/case RC1 latency tradeoff when citing its major reduction claim.

## Observed model-operated integration

A separate coding agent used only the public wheel/docs/skill and created two fresh projects: a relational invoice adapter (55 protocol LOC) and a graph navigation adapter (53 protocol LOC). It preserved 24/24 paired comparisons across address, accessed final state, actions, outcome and native verifier reward. Each adapter used four sparse coordinates twice and four tiny comparisons; these are comparison runs, not 24 unique environments. Core edits and human interventions were zero.

Sparse fixtures reduced retained canonical serialized bytes by 99.893% (relational) and 99.690% (graph), with measured mean added pipeline latency about 3.304 ms and 3.730 ms in the confirmed run. Tiny fully accessed fixtures had zero byte reduction and additional latency. Timings are small synthetic measurements without confidence intervals; they establish neither RAM savings nor a general speed/production advantage.

The agent corrected two exploratory import assumptions and one missing-page exception assertion; failed attempts and an output-truncation/report-writer warning are preserved in the private audit. Documentation now names the PagedState import and observed error behavior explicitly. CLI, three shipped adapter examples, deterministic eight-tool MCP workflow, provenance replay and missing-page refusal passed. The agent chose the bundled client; MCP order was scripted, so no model-driven MCP tool-selection rate is claimed. No unsupported claims were observed in the reviewed final integration report, not a general agent error-rate study.

The audit distinguishes the initially installed artifact hash from a later final-wheel reinstall/rerun. The retained audit records commands, measured output, source versions and evidence hashes. Only Windows was exercised locally; no marketplace installation or remote endpoint test is claimed.
