Token totals tell you how much a session consumed. They do not tell you whether its code was accepted, how much review it required, or what a subscription actually charged. Use one completed change as the unit of comparison.
Define accepted before starting
Write two or three checks for the change: a behavior, a relevant automated check, and a review condition. For example, a settings form might save a name after refresh, reject an empty name, and leave no unresolved high-severity PR findings. If one run meets the checks and another does not, comparing their token spend alone is misleading.
Capture the whole attempt
Record the implementation session, reviewer session, retries, and human review time under one task ID. Note the CLI, model, branch, starting commit, acceptance result, and exact commands run. When a model fails and a stronger model completes the job, count both sessions. Excluding failed first attempts makes a routing strategy appear cheaper than it was.
Keep usage categories separate
If the CLI reports them, record uncached input, cached input, cache writes, and output separately for each model. Provider pricing treats these categories differently and can change. Our session-cost calculator lets you enter current rates and calculate an API-equivalent estimate. A plan quota or percentage is a different signal, and subscriptions, included usage, discounts, and reporting gaps can make the actual payment diverge from the estimate.
Compare outcomes on similar tasks
For each strategy, calculate accepted changes divided by total attempts, estimated spend per accepted change, and human minutes per accepted change. If a task never reaches acceptance, record it as incomplete rather than assigning it zero cost. Compare several tasks of similar scope before concluding that a small-first or premium-first route helps. Different repositories and acceptance bars invalidate a simple ranking.
Use the result to change one decision
Look for the expensive boundary: broad prompts, repeated discovery, context resets, review failures, or unnecessary premium-model turns. Change one part of the workflow and repeat the measurement. Keep quality visible by recording regressions or later fixes; a fast merge that needs rework next week was not a cheap completion.
Copyable resources
Copyable accepted-change worksheet
One row per session, then summarize all sessions for the task.
Task ID and repository: [ ]
Starting commit and branch: [ ]
Acceptance checks: [ ]
Session 1: CLI/model [ ]; role [implement/review/fix]; outcome [ ]; tokens by category [ ]; estimated cost [ ]
Session 2 or retry: [same fields]
Human review minutes: [ ]
Provider charge attributable to task, if knowable: [unknown or amount with method]
Final state: [accepted / incomplete / reverted]
Later rework: [none / issue and time]
Compare only with tasks of similar scope and the same acceptance bar. Frequently asked questions
Can I divide my subscription fee by the number of PRs?
You can use that as a rough monthly allocation if you label the method, but it is not the provider's per-PR charge. Keep subscription payments separate from API-equivalent token estimates.
What if a CLI does not report cache tokens?
Mark the category unknown. Do not silently treat it as zero; the estimate may be incomplete.