Use case / 2026-09-28

Should I use a cheaper AI model first, then a stronger model to review?

Choose current Claude or OpenAI models for triage, implementation, and review; copy the handoff prompts and compare complete task costs.

Canopy Agents rail showing separate coding-agent sessions
Canopy Agents rail showing separate coding-agent sessions

Can I give Haiku or Luna the cheap first pass and ask Opus, Fable, Sol, or Astra to review? Yes, if the first agent produces a bounded, verifiable finding and the second agent can inspect the final change. The split can also waste money when both agents repeat the same investigation. Use this workflow to decide from the accepted result, not a model's price alone.

Pick a model for the job you can actually test

As of 28 September 2026, OpenAI describes Luna for scoped tasks and triage, Sol for everyday coding and careful review, and Astra for ambiguous or demanding work. Anthropic describes Haiku 4.5 as an efficiency starting point, Sonnet 5 for everyday agent and coding work, Opus 5.5 for complex agentic coding, and Fable 5.1 for its highest-capability long-running work. These are provider starting points, not a ranking of which CLI will fix your bug. Model availability, tool support, reasoning or effort controls, and limits differ by product, plan, version, and account. Check the model selector in your installed CLI before planning a route. In Canopy, each CLI owns its model and account; separate tabs do not pool subscriptions.

Current provider guidance checked 28 September 2026. Choose from models available in your installed CLI and compare complete outcomes.
WorkCandidate starting pointEvidence to collect
Read-only, narrow triageHaiku 4.5 or Luna; try lower effort on a capable model tooFirst failing command, relevant files, uncertainty
Routine implementationSonnet 5 or Sol when availableFocused diff, passing and failing tests
Complex implementationOpus 5.5 or Astra; evaluate Fable 5.1 for the hardest long-running workAccepted behavior across the affected system
Independent reviewA capable model with the final commit and acceptance checksReproducible finding or explicit no-finding report

Give the first model a bounded question

Suppose a settings-form test fails only after an API timeout. Give a lightweight model the failing command, branch, and expected retry behavior. Ask it to identify the first relevant error, trace the code path, and report a probable cause with file and line references. Keep this pass read-only. Its stop point is a concise finding, not a speculative rewrite of authentication or a PR merge. If it cannot reproduce or cite the path within the agreed time or usage limit, stop and escalate with the uncertainty intact.

Hand off facts the second model can verify

Open a second session in the intended checkout and provide the original acceptance checks, exact failing command and output, first model's cited finding, current branch or PR, and what was not tested. Ask the stronger model to inspect those files before editing, because the first diagnosis may be wrong. Canopy can show supported agent messages and project context, but the CLI conversations remain separate and integration depth varies by agent. A six-line handoff is more useful than pasting the entire first transcript into a new expensive session.

Use capability where uncertainty is highest

A strong model can be worthwhile when the fix crosses API and UI state, changes a shared contract, or needs an independent reviewer to challenge a plausible patch. A mechanical typo with a focused failing test may not justify an extra premium pass. The CLI selects its own model; Canopy does not automatically route prompts or pool provider subscriptions. Anthropic's cost guidance treats model choice, caching, and token hygiene as separate levers, and its examples do not prove a universal winner for your repository. Keep the task and acceptance criteria fixed when comparing routes.

Possible roles for one bounded defect, not model rankings.
StepUseful outputEscalate or stop when
Smaller-model triageReproduction, likely path, cited filesNo evidence, repeated guesses, or sensitive cross-system change
Stronger-model implementationFocused fix and exact testsScope expands beyond the agreed behavior
Independent reviewReproduced finding or justified no-finding reportReviewer cannot inspect the final commit or running path
Human acceptanceObserved retry behavior and PR decisionTests, diff, or production risk remain unresolved

Count every attempt in the experiment

Record the triage session, implementation session, any retries, and review as one route. For a baseline, use a comparably scoped issue with the same CLI account, repository, acceptance checks, and review standard in a single capable-model session. Note elapsed time, tokens by available category, estimated API-equivalent cost, human review time, and whether the final patch was accepted. A failed low-cost first pass plus a full second investigation can cost more than starting with the stronger model. A cheap first pass that narrows the risk and prevents rework can save time. Both are hypotheses until measured on complete tasks.

End with one change owner

The implementer owns the final branch and latest diff. The reviewer reports evidence-backed findings without silently editing the same checkout. A person decides whether the retry works, the tests cover both success and failure, and the PR is ready. Canopy's usage view can help compare supported session estimates, while the provider account remains authoritative for actual charges or plan limits. If the handoff becomes longer than the fix, simplify the next experiment rather than adding another agent.

Copyable resources

First-pass triage prompt

Send to a separate read-only CLI session with your real command and branch.

Find the cause of [failing behavior] on branch [branch] at commit [SHA]. Expected behavior: [ ]. Failing command and first error: [ ]. Do not edit files. Reproduce if possible, trace only the relevant path, and cite file paths and lines. Return: first verified failure; likely cause with evidence; what you could not verify; one next test. Stop after [time or usage limit] or if the investigation crosses into [sensitive area].

Independent review prompt

Give the reviewer a stable commit or PR, not an uncommitted moving target.

Review commit [SHA] / PR [URL] against these acceptance checks: [ ]. The implementation claims: [ ]. Inspect the final diff and affected call sites independently. Run or describe the relevant tests. Do not edit, approve, or merge. Report each finding with file, line, reproduction, and impact; distinguish untested concerns from verified failures. If you find none, state exactly what you checked and what remains unverified.

Two-model experiment card

Fill one card per accepted task; use a comparable baseline before claiming savings.

Task and acceptance checks: [ ]
Starting commit/branch: [ ]
Triage model/CLI: [ ]; read-only finding and cited files: [ ]; stop time: [ ]
Implementation model/CLI: [ ]; changed files: [ ]; tests and results: [ ]
Independent review: [finding or no finding with evidence]
Retries and rework: [ ]
Usage by session and category: [ ]; estimated API equivalent: [ ]
Human review minutes: [ ]; elapsed time: [ ]
Final outcome: accepted / rejected / further work
Comparison route with same review standard: [ ]

Frequently asked questions

Can I run two different agent CLIs in one Canopy project?

Yes, for supported installed CLIs. Each session uses the CLI's own model and account settings.

Does a smaller first-pass model always reduce cost?

No. Repeated retries, poor handoffs, and a second full investigation can increase total usage. Compare the accepted result and complete workflow.

Can a Haiku tab hand work to an Astra or Opus tab automatically?

Do not assume that. Sessions have separate conversations. Pass the task, branch or commit, evidence, acceptance checks, and open questions explicitly; inspect any agent messaging supported by the installed Canopy and CLI versions.

Should I use Fable or Astra for every PR review?

No. Start with the smallest capable review process that finds real defects on your work. Compare verified findings, missed defects, latency, usage, and human review; a difficult cross-system change may justify more capability.

Browse more Canopy questions →

Sources and further reading