Guide / 2026-09-28

Did your coding agent weaken a test to make CI pass?

Review changed assertions, skips, mocks, and browser setup against the original behavior before accepting an agent's green check.

Canopy code diff and pull request review for an agent-written change
Canopy code diff and pull request review for an agent-written change

A green check says a command passed on a commit. It does not say the check still guards the behavior you asked for. If an agent changed both application code and tests, inspect how the test changed before accepting the result. This guide uses a Save button that should recover from an API error. It distinguishes a legitimate test update from a weaker assertion, skipped case, or test that repairs the app at runtime.

Keep the original behavior visible

Write down the acceptance check before reading the new test: after a failed save, the form shows an error, keeps the entered value, and lets the user retry successfully. Identify the last known failing reproduction or test run and the commit it used. If requirements changed during implementation, have the product owner record the new behavior; a test can properly change with a changed requirement. Without that decision, a passing assertion against a different outcome is not evidence that the original request was delivered.

  • Name one success path and one failure path in terms a user can observe.
  • Record the exact PR head and check run you are reviewing.
  • Ask the agent which behavior each changed assertion now protects.

Review the tests and app code in the same diff

Open the PR's changed files in Canopy or GitHub and inspect test edits alongside production edits. Watch for a removed assertion, a snapshot updated without visual review, a broad mock that replaces the failing service, a new skip or only marker, a catch block that swallows failure, or a test that changes page code before asserting. None of these strings alone proves misconduct: mocks, snapshots, and skips can have legitimate uses. The decision comes from comparing the new test's observable claim with the approved behavior and its real execution path. GitHub's PR Files changed view exposes both sides of the patch; its green status checks report results for configured commands, not the meaning of every assertion.

In the Save-button example, inspect what each test change proves.
Test changeQuestion to askEvidence needed
Expect error then retry succeedsDoes it exercise the real failed-save state?Visible error, second request, persisted result
Expect spinner to disappear onlyCould save still fail silently?Check response, message, and saved value
Mock every save response as successIs the API failure path still covered elsewhere?Separate error-path test and integration check
Skip the failing testWho approved removing this gate?Reason, owner, replacement check, follow-up

Prove the test can fail for the relevant defect

Run the changed test on the PR head and record its exact name and result. In a disposable checkout, use a safe, temporary reversal of the intended fix or the known failing base commit to see whether the test catches the broken behavior. Restore the checkout afterward and inspect its status. A test that stays green with the defect present needs investigation. A test that fails on the old state and passes on the fixed state is stronger evidence, though it still may miss another path. For browser tests, Playwright recommends testing user-visible behavior with web-first assertions and isolated state; inspect any setup script that injects code or replaces requests before the page is used.

  • Keep the negative-control trial out of the PR branch; do not commit a deliberately broken app.
  • Compare the same test command and fixture on both states.
  • If a test is flaky, diagnose the nondeterminism instead of weakening its assertion until it stops failing.

Use the running product as an independent check

Open the app on the PR's intended checkout and trigger a save failure with an approved test service or fixture. Observe the error, retry, and final saved value without the test runner's injected setup. Browser and server logs can confirm the second request and persisted result. If the app behaves correctly but the old test was wrong, document why the test update was necessary and review the new assertion. If the app still fails while tests pass, treat the PR as incomplete and send the agent the reproduction, relevant test diff, and first differing observation. A screenshot of a single success state cannot establish the retry path.

  • Use disposable accounts and test data; keep private records out of public PR comments.
  • Check the latest diff after any correction, because the agent may modify tests again.
  • Have a person approve a changed acceptance rule, especially for access, payments, or data loss.

Close the PR with a defensible record

Ask the implementer to report the reason for each changed test, the command and exact result, and the live behavior it exercised. A separate reviewer can inspect the diff and run the failure path without editing it. Confirm the CI check is attached to the latest head, then record which original checks passed, which were replaced, and why. A passing suite is useful evidence when its checks still correspond to the intended product; it is not a substitute for that correspondence. If the test only passes because it patches the app at runtime, repair the app and redesign the test before merging.

  • Do not accuse an agent of intent from the diff; focus on the observable test boundary.
  • Retain a regression check for the defect that motivated the task.
  • Escalate disagreements about what the product should do to the owner of that requirement.

Copyable resources

Test-integrity review card

Run negative controls only in a disposable checkout, then restore and verify its status.

Task and approved behavior: [ ]
PR head and latest CI run: [ ]
Original failing reproduction or test: [ ]
Production files changed: [ ]
Test files, assertions, mocks, skips, snapshots, and setup changed: [ ]
What each changed test now claims: [ ]
New test on PR head: [command, named case, result]
Negative control on known broken state: [command, result]
Live success path: [observed]
Live failure and retry path: [observed]
Changed requirement approved by: [owner or none]
Reviewer decision and remaining gap: [ ]

Frequently asked questions

Is it always wrong for an agent to change a test while fixing code?

No. A changed requirement, corrected fixture, or better behavioral assertion may require a test edit. Review why it changed and whether it still catches the relevant defect.

How can I tell whether a new test actually protects the fix?

Run it on the fixed commit and, in a disposable checkout, against a known broken state. It should fail for the relevant defect and pass after the fix, while the running product should show the expected behavior.

Does a green GitHub Actions check prove the test is good?

No. It proves the configured check completed successfully on that commit. Inspect changed assertions and exercise the behavior the test is meant to protect.

What if a browser test injects code into the page?

Inspect whether that setup supplies harmless fixtures or changes the application behavior being tested. A test that repairs the broken UI before asserting does not demonstrate that users receive a working app.

Browse more Canopy questions →

Sources and further reading