The lowest per-token price does not guarantee the lowest cost to finish a feature. A model that needs repeated investigation, corrections, and review can use more total resources than a stronger model that reaches an accepted change once. The reverse is also true for a simple task. Measure the whole job.
Start with one accepted outcome
Define the same repository state, task, acceptance checks, and review standard for both model choices. For example: fix a login redirect, pass the existing auth tests, and show the final diff. A first-pass response is not an outcome if the fix fails in the browser. Count every attempt until the change is accepted or explicitly abandoned. Keep model settings, tools, and permissions in the record; they can change both token use and quality.
A worked comparison, not a price claim
Imagine two hypothetical paths priced in arbitrary cost units. Path A spends 2 units on each of four attempts, then 6 units on a separate review: 14 units and one accepted patch. Path B spends 7 units on one implementation and 3 units on review: 10 units and one accepted patch. The cheaper attempt cost more per accepted patch. If A succeeds first try, its 8-unit total wins. The numbers are illustrative, not current provider rates or measured Canopy results; the point is to count all attempts and review under the same acceptance rule.
| Path | Implementation | Review | Total | Result |
|---|---|---|---|---|
| A: lower cost per attempt | 4 × 2 = 8 | 6 | 14 | Accepted after four tries |
| B: higher cost per attempt | 1 × 7 = 7 | 3 | 10 | Accepted after one try |
| A if it works first try | 1 × 2 = 2 | 6 | 8 | Accepted after one try |
Separate token categories and billing arrangements
Input, cached input, cache writes where applicable, and output can have different API rates. OpenAI's API usage reference exposes input, cached input, output, and request counts; its prompt-caching guide says cache behavior depends on reusable prefixes and reports cache categories separately. Anthropic also documents distinct cache read and write accounting. If you use a subscription CLI, an API-equivalent token estimate is a comparison aid, not a charge or a direct measure of remaining plan allowance. Use the provider's billing and plan views for the actual money and limits.
Add the cost the token chart cannot see
Record the human time spent re-explaining the task, resolving conflicts, reproducing failures, and reading the final diff. Also note whether the implementation passed checks, survived review, and needed a follow-up fix. You need not invent a dollar rate for a colleague's time; a separate minutes column makes the tradeoff visible. A ten-minute extra review may be worthwhile when it prevents a bad merge. A low-token response that leaves no inspectable result should not be counted as a successful cheap run.
Run a small local comparison
Pick a bounded task with a deterministic check. Use separate worktrees from the same commit, one model path per worktree, and identical acceptance criteria. Keep each session's usage, model, timestamps, retries, final diff, test result, and review findings. In Canopy, the usage screen can help group supported CLI signals by session and model, while the agent workspace helps connect work to branches and review. Check missing or stale usage data before calculating totals. With one or two trials, report only what happened in those trials; do not declare a universal winning model.
Copyable resources
Cost per accepted change trial
Copy once for each model path and compare outcomes before prices.
Task, repository, starting commit: [ ]
Acceptance checks and reviewer: [ ]
CLI, model, account/plan, settings: [ ]
Attempts and reason for each retry: [ ]
Input / cached input / cache writes / output: [ ]
Estimated API-equivalent cost and rate date: [ ]
Actual provider charge or plan usage, if known: [ ]
Human review and correction minutes: [ ]
Final tests, diff, and accepted/rejected outcome: [ ] Frequently asked questions
Is the least expensive model always cheaper for simple coding tasks?
It may be. If it passes the same checks with little rework, the lower rate can win. Measure a bounded task rather than assuming either model is cheaper overall.
Should cached tokens be priced like new input?
No. API providers publish separate cache rates and usage categories. Use the applicable model, date, and billing path; a CLI subscription can follow different rules.
Can Canopy tell me which model produced the best code?
No. Canopy can show supported usage and project evidence, but a person must evaluate the running behavior, tests, diff, and review findings.
Can one trial prove a model is better?
No. It shows one outcome under recorded conditions. Repeat across representative tasks before making a team policy.