M04: Planning Measurement Disposition
Result
NOT ESTABLISHED. No production token, cost, or latency improvement is established for
0.7.12. This is a measurement-gap report, not a completed paid campaign, a campaign lock, or
provider qualification. Passing local replay and installed-package checks does not change this
verdict. The report belongs to task 202609261720-KKE9ZN, WorkItem PL-11.
Exact Evidence Boundary
- Candidate source commit:
bb6cfaf0da9981370fb03d0ebcd574b88202fa8e. - Candidate tree:
c62e8980574b219114948142fe86244e8b9f3372. - Previous-release source, resolved from
v0.7.11:65b1c24e83576b7daca306e6dc6b7085c522cf58. - Candidate package version at this boundary:
0.7.12-beta.1. - Measurement contract: PL-11 and execution charter.
The following SHA-256 values pin source bytes at the candidate commit. They are not hashes of published packages or evidence of a provider run.
| Input | SHA-256 |
|---|---|
scripts/bench/paired-production-driver.mjs | 1789091a57c67ead5b00924bd07d2eb83268674f677725cc71378855ee261f84 |
scripts/bench/paired-result-report.mjs | 8010931dba4f039c19d0965efe5ee34931532a0e316ab00f065b7a21700a1bf4 |
scripts/lib/installed-planning-matrix.mjs | da1430971559329e6359b279a0e0b086e7eca08457b9e0cd1183a21fa91bd7ad |
agentplane-roadmap-r2/tasks/PL-11.md | b9fe7e7d85207a97d608acd64baf447b10984e67673e42cdef6ee23da79a68c5 |
agentplane-roadmap-r2/EXECUTION-CHARTER.md | b110063e261636cdcf0303955513ecaf8b049ba7a0052a74bcb3b5a344cd5cf5 |
There is no M04 campaign.lock.json, assigned-run ledger, minimal-agent artifact, pinned
provider/model configuration, or full-host telemetry for this candidate. Those campaign inputs
are unknown, not inferred from historical M01/M03 fixtures. No paid M04 attempts were assigned
or launched by this WorkItem. The number and cost of upstream development-host calls are unknown;
zero locally assigned benchmark attempts must not be interpreted as zero upstream cost.
Separate Strata
| Stratum | Available evidence | Host construction and retries | First-mutation latency | Verified-result latency | Efficiency verdict |
|---|---|---|---|---|---|
| Caller-supplied Plan; no separate PLANNER | Installed CLI reaches the real approval boundary and then EXECUTOR/EVALUATOR without a PLANNER WorkOrder | Unknown | Unknown | Unknown | NOT ESTABLISHED |
| Required PLANNER; managed planning bridge | Contract tests cover supported capability declarations, typed result admission and replay; installed unsupported adapter fails before launch | Unknown | Unknown | Unknown | NOT ESTABLISHED |
These strata are not pooled with each other or with managed/external transport labels. Input provenance and transport are separate dimensions. The second row does not certify a live provider sandbox: contract test doubles cannot replace observed runtime containment receipts. An unverified receipt still stops admission.
PL-10 exercised six installed planning scenarios and eight migration scenarios. Its native release-critical check passed 64 tests in 11 test files. These observations establish local behavior, not elapsed provider time or token savings. Test-run durations are not substituted for either latency metric. Review in this development session used the same host and is not an independent second-agent measurement review.
Conditions For A Measured Claim
A subsequent campaign must pin exact artifacts and source SHAs for candidate, previous release, and minimal-agent arms. Its immutable lock must also pin target repository commit/tree, intent, oracle and final review requirements, adapter, model, effort, sandbox, network, cache/session policy, randomized assignment, retry limits, and budget before assigning attempts. Every arm in each stratum must receive the same final oracle and review requirements. Missing inputs prohibit execution as a locked campaign; this report must not be used as a lock with placeholder values.
Retain every assigned attempt, including blocked, failed, cancelled, interrupted, and retried attempts. Attribute upstream intent/Plan construction, revisions, host orchestration, managed calls, verification and review to the same assignment. Record unknown usage as unknown, never zero. A retry is a cost-bearing observation, not permission to discard an earlier failure. Total cost per verified success includes all assigned costs in the numerator. With zero verified successes or incomplete accounting, no finite full-cost estimate is established. Cached input and reasoning tokens remain subsets and must not be added a second time.
Measure first mutation and independently verified completion from the same assignment start, including upstream construction. Retain absent endpoints for blocked or failed attempts rather than removing those attempts from the population. Report both latency distributions separately for each provenance stratum and transport. A reduction in dispatch count is not a measured percentage reduction in end-to-end work.
The existing paired driver and reporter provide artifact pinning, paired oracle checks, and attempt-level usage handling. Their transport-only summaries and recorded agent-stage timing do not, by themselves, establish complete upstream accounting or either M04 latency endpoint. Do not relabel their existing historical output as an M04 full-host result.
Regression Checks
The required commands are bun run bench:agent-efficiency:check and
bun run bench:agent-efficiency:replay:check. Their results are recorded by the Task controller.
They protect frozen structural and replay evidence. Neither command conducts a live M04
campaign, certifies provider access, or upgrades this report's efficiency verdict.