Blog / 2026-09-28

Why cheaper AI tokens can make a coding task more expensive

A worked task-cost comparison that counts retries, cache categories, human review, and accepted code instead of ranking models by input-token price.

Canopy usage view with token and estimated-cost summaries
Canopy usage view with token and estimated-cost summaries · Illustrative data; displayed costs are estimates, not provider bills.

The lowest per-token price does not guarantee the lowest cost to finish a feature. A model that needs repeated investigation, corrections, and review can use more total resources than a stronger model that reaches an accepted change once. The reverse is also true for a simple task. Measure the whole job.

Start with one accepted outcome

Define the same repository state, task, acceptance checks, and review standard for both model choices. For example: fix a login redirect, pass the existing auth tests, and show the final diff. A first-pass response is not an outcome if the fix fails in the browser. Count every attempt until the change is accepted or explicitly abandoned. Keep model settings, tools, and permissions in the record; they can change both token use and quality.

A worked comparison, not a price claim

Imagine two hypothetical paths priced in arbitrary cost units. Path A spends 2 units on each of four attempts, then 6 units on a separate review: 14 units and one accepted patch. Path B spends 7 units on one implementation and 3 units on review: 10 units and one accepted patch. The cheaper attempt cost more per accepted patch. If A succeeds first try, its 8-unit total wins. The numbers are illustrative, not current provider rates or measured Canopy results; the point is to count all attempts and review under the same acceptance rule.

Illustrative units only; include all attempts through acceptance.
PathImplementationReviewTotalResult
A: lower cost per attempt4 × 2 = 8614Accepted after four tries
B: higher cost per attempt1 × 7 = 7310Accepted after one try
A if it works first try1 × 2 = 268Accepted after one try

Separate token categories and billing arrangements

Input, cached input, cache writes where applicable, and output can have different API rates. OpenAI's API usage reference exposes input, cached input, output, and request counts; its prompt-caching guide says cache behavior depends on reusable prefixes and reports cache categories separately. Anthropic also documents distinct cache read and write accounting. If you use a subscription CLI, an API-equivalent token estimate is a comparison aid, not a charge or a direct measure of remaining plan allowance. Use the provider's billing and plan views for the actual money and limits.

Add the cost the token chart cannot see

Record the human time spent re-explaining the task, resolving conflicts, reproducing failures, and reading the final diff. Also note whether the implementation passed checks, survived review, and needed a follow-up fix. You need not invent a dollar rate for a colleague's time; a separate minutes column makes the tradeoff visible. A ten-minute extra review may be worthwhile when it prevents a bad merge. A low-token response that leaves no inspectable result should not be counted as a successful cheap run.

Run a small local comparison

Pick a bounded task with a deterministic check. Use separate worktrees from the same commit, one model path per worktree, and identical acceptance criteria. Keep each session's usage, model, timestamps, retries, final diff, test result, and review findings. In Canopy, the usage screen can help group supported CLI signals by session and model, while the agent workspace helps connect work to branches and review. Check missing or stale usage data before calculating totals. With one or two trials, report only what happened in those trials; do not declare a universal winning model.

Copyable resources

Cost per accepted change trial

Copy once for each model path and compare outcomes before prices.

Task, repository, starting commit: [ ]
Acceptance checks and reviewer: [ ]
CLI, model, account/plan, settings: [ ]
Attempts and reason for each retry: [ ]
Input / cached input / cache writes / output: [ ]
Estimated API-equivalent cost and rate date: [ ]
Actual provider charge or plan usage, if known: [ ]
Human review and correction minutes: [ ]
Final tests, diff, and accepted/rejected outcome: [ ]

Frequently asked questions

Is the least expensive model always cheaper for simple coding tasks?

It may be. If it passes the same checks with little rework, the lower rate can win. Measure a bounded task rather than assuming either model is cheaper overall.

Should cached tokens be priced like new input?

No. API providers publish separate cache rates and usage categories. Use the applicable model, date, and billing path; a CLI subscription can follow different rules.

Can Canopy tell me which model produced the best code?

No. Canopy can show supported usage and project evidence, but a person must evaluate the running behavior, tests, diff, and review findings.

Can one trial prove a model is better?

No. It shows one outcome under recorded conditions. Repeat across representative tasks before making a team policy.

Browse more Canopy questions →

Sources and further reading