Guide / 2026-09-28

How to evaluate an AI coding workspace on a real project

A repeatable trial for terminal, IDE, and multi-agent workspaces that follows one change from setup through running behavior, review, resume, and accepted result.

Canopy Agent Workspace connecting a coding session to files, checks, and PR review
Canopy Agent Workspace connecting a coding session to files, checks, and PR review

A feature checklist tells you what a product claims to expose. It does not show whether your team can finish a change with it. Evaluate an AI coding workspace using a task from your own repository, a visible acceptance bar, and the entire path from a fresh checkout to an accepted diff. Keep the agent and model fixed if you want to isolate the workspace; change them in a separate experiment.

Choose a task with a failure you can observe

Use a bounded change that resembles the work you actually ship. For example: a settings form saves a display name, shows a useful empty-name error, survives refresh, and works at mobile width. Supply the same starting commit, issue, test data, and success and failure checks in each trial. Avoid a task whose answer is already in a public benchmark or its repository history. A benchmark patch score tests one narrow outcome; it does not measure how readily a person finds the right checkout, runs the app, reviews the final change, or resumes tomorrow. SWE-bench's own evaluation protocol records patches and test outcomes under a fixed dataset, and its submission checklist calls out solution lookup and attempt counts as controls. Your workspace trial needs equally explicit controls for its different question.

Hold the agent constant when comparing workspaces

Record each product's installed version and plan, operating system, repository commit, CLI and model, account, prompt, permissions, and time window. If one product cannot run the same CLI, state that limitation and treat the result as a comparison of complete tool stacks. Do not call it a workspace-only result. Use separate branches or Git worktrees so a first attempt cannot quietly supply files to the second. Reset test data and local services before each run. Git's worktree documentation explains how linked working trees have distinct working files while sharing repository history; a worktree alone does not isolate your database or third-party services.

Observe the whole task, including the human queue

Start the project and note missing dependencies, secrets, and run commands. Ask the agent for the change, then find the correct session, checkout, and running URL. Inspect normal and error states, one narrow visual correction, the final diff, and the PR checks on the latest commit. Ask a separate reviewer to identify one concrete risk or explain why it found none; verify the finding yourself. GitHub documents that required checks must pass on the latest commit, so a green result from before the agent's last fix is insufficient. Pause and resume once with only an artifact handoff: goal, branch, changed files, observed behavior, checks, decisions, and next action.

Record observations; do not award points for a feature you did not use.
StageEvidenceFailure to notice
SetupVersion, checkout, prerequisites, commands, time to first ready URLThe wrong branch or an unrepeatable environment
ImplementationPrompt, interventions, changed files, observed success and error statesA plausible summary with a broken app
ReviewLatest diff, specific finding, CI commit, owner decisionAn old green check or unexamined generated code
ResumeSame task and checkout recovered from a short handoffA new session repeating work or losing the decision
CostAttempts, human minutes, provider usage source, accepted resultComparing token prices without rework

Compare accepted outcomes before speed or tokens

For each attempt, mark accepted, rejected, or still uncertain against the same checks. Record elapsed time, human setup and review minutes, retries, rework, and any provider-reported usage separately. A quick first draft is not a quick accepted change. A Canopy usage estimate is not the provider's invoice. If one trial had a network outage, a hidden credential, or an extra reviewer, annotate it rather than hiding it in a winner score. METR's work on experienced developers is a reminder that perceived speed and measured task completion can diverge; its study population and time period do not establish how a particular 2026 workspace performs on your project.

Run the Canopy trial as a complete workflow

Open the project in Canopy and confirm the installed CLI and checkout. Save the real website, API, or worker commands in Servers, then inspect their ready lines and the URL in Preview. Give one agent the bounded implementation; if another agent reviews, keep its request tied to the current branch and diff. Use Agent Workspace and Git or PR context to inspect the resulting files and checks. Pause and locate the session again before the final decision. Repeat the same task in a terminal, IDE, or alternative workspace with the same controls. If your app uses a single process, do not add artificial services just to make a product look busy. The useful verdict is which setup made your team's accepted change easier to produce and verify, with its remaining limits written down.

Copyable resources

Copyable workspace trial record

Use one record per attempt; duplicate it for the other workspace and compare accepted outcomes.

Task, starting commit, and issue: [ ]
Success check and failure check: [ ]
Workspace/version/plan/platform: [ ]
CLI/model/account/permissions: [ ]
Branch or worktree, clean starting state: [ ]
Project commands, test data, and first ready URL: [ ]
Prompt and human interventions: [ ]
Normal/error behavior and visual correction: [ ]
Reviewer finding and response: [ ]
Latest diff, PR commit, checks, and human decision: [ ]
Pause/resume handoff and what was recovered: [ ]
Elapsed time, human minutes, retries, usage source: [ ]
Accepted, rejected, or uncertain; reason and limitation: [ ]

Frequently asked questions

Can a coding-agent benchmark tell me which workspace to use?

It can inform expectations about a model and agent under the benchmark's protocol. Test setup, running behavior, review, resume, and accepted work on your own repository to choose a workspace.

Should I use the same model in both products?

Yes if you want to isolate the workspace. If the products require different CLIs or models, label the trial as a comparison of complete stacks.

Is a faster first draft the winner?

Only if it also passes the same acceptance and review checks. Record rework, human time, and the final decision separately.

How many tasks make the result reliable?

One task can expose a deal-breaker or guide a trial, but it cannot establish a general speed ranking. Repeat representative tasks before making a broad team decision.

Browse more Canopy questions →

Sources and further reading