Guide / 2026-09-28

Why do my agent tests pass locally but fail in GitHub Actions?

Compare the exact PR commit, failing CI step, runner, install, environment, and test command before changing code or rerunning a job.

Canopy pull request and running project beside the agent workspace
Canopy pull request and running project beside the agent workspace

An agent reports passing tests in its local checkout, while the PR has a red GitHub Actions check. Neither observation explains the other until they refer to the same commit and command. This guide uses a Node app whose agent ran a local test against existing dependencies while CI failed during clean installation. The same evidence order applies to other languages, but use the project’s actual workflow and package manager.

Pin the failed run to the PR's current commit

Open the PR Checks or Actions run and note its run ID, triggering event, commit SHA, job, and attempt number. Compare that SHA with the PR's latest head and the agent's local checkout. An earlier failure can remain visible after a newer commit, while an old green local log may predate the last agent edit. GitHub's workflow logs show the step that failed and the runner image used. A rerun repeats the original event's SHA and ref; it does not test a newer PR commit. If you need evidence for the latest code, wait for or trigger the workflow appropriate to that commit.

  • Record the last successful run and its commit if you need a comparison.
  • Distinguish a failed run, a cancelled run, and a job that never started.
  • Do not declare a fix from a rerun of the old SHA.

Find the first failing step, not the last red summary

Read the failing job from setup through its first actionable error. A dependency install failure needs a different owner from a type error, test assertion, browser timeout, or deployment credential failure. Capture the command, short sanitized error, runner OS and tool versions, working directory, and any matrix entry. GitHub-hosted jobs usually start on fresh runner instances, so an agent's existing local dependency folder or untracked file is not present there. Compare the workflow YAML and lockfile in the commit that actually ran; the current working tree may already have changed.

A compact local-versus-CI comparison for one failed PR run.
BoundaryAgent's local evidenceGitHub Actions evidence
CodeCheckout branch and exact commitRun SHA and PR head
InstallPackage manager and local dependency stateInstall command and committed lockfile
RuntimeOS, language version, environmentRunner image, matrix entry, setup step
CheckCommand and result after final editFirst failing step and error

Reproduce the relevant CI boundary locally

For the Node example, the workflow uses npm ci but the agent only ran npm test with an old node_modules directory. npm documents that npm ci performs a clean install and fails if package.json and package-lock.json disagree. After preserving any work, run the repository's documented clean-install and test commands in a disposable checkout or container that matches the runner as closely as practical. If the lockfile is inconsistent, update it with the project's package manager and review the dependency diff; do not disable the install check. If the test itself fails only on CI, compare language version, operating system, case-sensitive paths, timezone, and external services one at a time. The observed log determines which of those is worth investigating.

  • Use the exact workflow command, not a convenient local substitute.
  • Check untracked and ignored files when local success depends on fixtures or generated output.
  • Keep the reproduction isolated when it would replace dependencies or generate files in a teammate's checkout.

Treat missing configuration as an access boundary

If the first failure says a variable or credential is missing, identify whether the workflow should have it for this event. GitHub does not pass most Actions secrets to workflows triggered by PRs from forks, and reusable workflows need explicit secret passing. Do not paste a production secret into a PR, agent prompt, or log to force a green check. Ask the repository owner whether the test should use a safe fixture, a restricted test service, or a separate trusted workflow. Changing to a more privileged event solely to expose secrets can create a serious security problem. The right fix may be to make an untrusted PR test independent of credentials.

  • Record the secret name and required scope, never its value.
  • Distinguish an absent secret from an app error after a valid connection.
  • Have the owner approve any change to CI permissions or production service access.

Give the agent one failure and verify the new head

Send the agent the failing run URL, SHA, first failing step, sanitized error, and local comparison. Ask it to explain the difference before editing, make the smallest relevant change, and run the equivalent check. In Canopy, keep the agent session, branch, PR, and local services visible while inspecting the diff; confirm the controls in the installed release. After a new commit reaches the PR, inspect the latest diff and the checks attached to that new head. A green CI result means the configured checks passed; repeat the feature's user journey if the workflow does not cover it.

  • If a rerun passes without a code change, investigate flakiness or an external service rather than claiming an application fix.
  • If a fix changes test expectations, review whether the test still catches the original failure.
  • Record the final commit and which CI attempt passed before resolving the PR conversation.

Copyable resources

Local-pass, CI-fail handoff

Share only sanitized errors and configuration names; do not paste credentials or private fixture data.

PR URL and latest head commit: [ ]
Failed Actions run ID / attempt / SHA: [ ]
Runner OS, image, language version, matrix entry: [ ]
First failing job, step, command, and short error: [ ]
Agent local checkout/commit and command/result after final edit: [ ]
Workflow install command and committed lockfile: [ ]
Difference to test first: [ ]
Disposable local reproduction and result: [ ]
Fix owner and changed files: [ ]
New PR head and its check result: [ ]
Feature journey still passes: [ ]

Frequently asked questions

Why did npm test pass locally while npm ci failed in Actions?

A local test may use packages already installed in node_modules. npm ci starts a clean install and rejects a mismatch between package.json and the lockfile. Compare the failing install log and committed files before changing the test.

Will rerunning the failed Actions job test my latest PR commit?

No. GitHub documents that a rerun uses the original event's SHA and ref. Check the run attached to the latest PR head after a new commit.

Should I add my API key to Actions to make a fork PR pass?

Do not expose a production key to an untrusted PR. GitHub withholds most secrets from fork-triggered workflows. Ask the owner to design a fixture, restricted test service, or separate trusted check.

Can the agent fix CI by skipping the failing test?

Only if the team decides the test is invalid and documents why. Inspect what behavior it guarded, then add or retain an equivalent check before considering the PR ready.

Browse more Canopy questions →

Sources and further reading