Skip to main content

M04: Planning Measurement Disposition

Result​

NOT ESTABLISHED. No production token, cost, or latency improvement is established for 0.7.12. This is a measurement-gap report, not a completed paid campaign, a campaign lock, or provider qualification. Passing local replay and installed-package checks does not change this verdict. The report belongs to task 202609261720-KKE9ZN, WorkItem PL-11.

Exact Evidence Boundary​

  • Candidate source commit: bb6cfaf0da9981370fb03d0ebcd574b88202fa8e.
  • Candidate tree: c62e8980574b219114948142fe86244e8b9f3372.
  • Previous-release source, resolved from v0.7.11: 65b1c24e83576b7daca306e6dc6b7085c522cf58.
  • Candidate package version at this boundary: 0.7.12-beta.1.
  • Measurement contract: PL-11 and execution charter.

The following SHA-256 values pin source bytes at the candidate commit. They are not hashes of published packages or evidence of a provider run.

InputSHA-256
scripts/bench/paired-production-driver.mjs1789091a57c67ead5b00924bd07d2eb83268674f677725cc71378855ee261f84
scripts/bench/paired-result-report.mjs8010931dba4f039c19d0965efe5ee34931532a0e316ab00f065b7a21700a1bf4
scripts/lib/installed-planning-matrix.mjsda1430971559329e6359b279a0e0b086e7eca08457b9e0cd1183a21fa91bd7ad
agentplane-roadmap-r2/tasks/PL-11.mdb9fe7e7d85207a97d608acd64baf447b10984e67673e42cdef6ee23da79a68c5
agentplane-roadmap-r2/EXECUTION-CHARTER.mdb110063e261636cdcf0303955513ecaf8b049ba7a0052a74bcb3b5a344cd5cf5

There is no M04 campaign.lock.json, assigned-run ledger, minimal-agent artifact, pinned provider/model configuration, or full-host telemetry for this candidate. Those campaign inputs are unknown, not inferred from historical M01/M03 fixtures. No paid M04 attempts were assigned or launched by this WorkItem. The number and cost of upstream development-host calls are unknown; zero locally assigned benchmark attempts must not be interpreted as zero upstream cost.

Separate Strata​

StratumAvailable evidenceHost construction and retriesFirst-mutation latencyVerified-result latencyEfficiency verdict
Caller-supplied Plan; no separate PLANNERInstalled CLI reaches the real approval boundary and then EXECUTOR/EVALUATOR without a PLANNER WorkOrderUnknownUnknownUnknownNOT ESTABLISHED
Required PLANNER; managed planning bridgeContract tests cover supported capability declarations, typed result admission and replay; installed unsupported adapter fails before launchUnknownUnknownUnknownNOT ESTABLISHED

These strata are not pooled with each other or with managed/external transport labels. Input provenance and transport are separate dimensions. The second row does not certify a live provider sandbox: contract test doubles cannot replace observed runtime containment receipts. An unverified receipt still stops admission.

PL-10 exercised six installed planning scenarios and eight migration scenarios. Its native release-critical check passed 64 tests in 11 test files. These observations establish local behavior, not elapsed provider time or token savings. Test-run durations are not substituted for either latency metric. Review in this development session used the same host and is not an independent second-agent measurement review.

Conditions For A Measured Claim​

A subsequent campaign must pin exact artifacts and source SHAs for candidate, previous release, and minimal-agent arms. Its immutable lock must also pin target repository commit/tree, intent, oracle and final review requirements, adapter, model, effort, sandbox, network, cache/session policy, randomized assignment, retry limits, and budget before assigning attempts. Every arm in each stratum must receive the same final oracle and review requirements. Missing inputs prohibit execution as a locked campaign; this report must not be used as a lock with placeholder values.

Retain every assigned attempt, including blocked, failed, cancelled, interrupted, and retried attempts. Attribute upstream intent/Plan construction, revisions, host orchestration, managed calls, verification and review to the same assignment. Record unknown usage as unknown, never zero. A retry is a cost-bearing observation, not permission to discard an earlier failure. Total cost per verified success includes all assigned costs in the numerator. With zero verified successes or incomplete accounting, no finite full-cost estimate is established. Cached input and reasoning tokens remain subsets and must not be added a second time.

Measure first mutation and independently verified completion from the same assignment start, including upstream construction. Retain absent endpoints for blocked or failed attempts rather than removing those attempts from the population. Report both latency distributions separately for each provenance stratum and transport. A reduction in dispatch count is not a measured percentage reduction in end-to-end work.

The existing paired driver and reporter provide artifact pinning, paired oracle checks, and attempt-level usage handling. Their transport-only summaries and recorded agent-stage timing do not, by themselves, establish complete upstream accounting or either M04 latency endpoint. Do not relabel their existing historical output as an M04 full-host result.

Regression Checks​

The required commands are bun run bench:agent-efficiency:check and bun run bench:agent-efficiency:replay:check. Their results are recorded by the Task controller. They protect frozen structural and replay evidence. Neither command conducts a live M04 campaign, certifies provider access, or upgrades this report's efficiency verdict.