title: Continuous integration for AI-assisted development slug: continuous-integration summary: >- How continuous integration combines code changes, tests behavior, and gives coding agents reliable feedback before merge and deployment. kind: concept status: evolving claimMode: mixed perspectiveOwner: Co confidence: high topics:
- software
- continuous-integration
- testing
- ai-agents related:
- reliable-agent-systems sources:
- title: Continuous Integration url: 'https://martinfowler.com/articles/continuousIntegration.html'
- title: Understanding GitHub Actions url: 'https://docs.github.com/en/actions/get-started/understand-github-actions'
- title: Managing a merge queue url: >- https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/configuring-pull-request-merges/managing-a-merge-queue
- title: Secure use reference url: 'https://docs.github.com/en/actions/reference/security/secure-use'
- title: Flaky Tests at Google url: >- https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html
- title: Scaling test impact analysis at Anthropic url: >- https://claude.com/blog/agentic-coding-is-straining-ci-heres-how-we-scaled-test-impact-analysis-at-anthropic
- title: Demystifying evals for AI agents url: 'https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents' aiAssisted: true generatedBy: Co sourceDigest: 'sha256:2d1eb4d8bde000ec210a154c33403979ad30ebff614d9beb99297e7163880874' updated: '2026-09-15T16:42:05.106Z' reviewStatus: approved publishedAt: '2026-09-15T15:48:43.587Z' reviewedContentDigest: 'sha256:d55dbf275c9bffc510abd4102a51d669c94c0a2870a587a39b10a58a6f0b43e7' reviewReceiptDigest: 'sha256:37de0cd95dd6ab88a61cb62b7791543e9cf9b28bdd8a8f5a01df7792fc2924bd' reviewBasis: technical-publication-authorization implementationReviewedBy: Co implementationReviewedAt: '2026-09-15T16:42:26.140Z' publicationAuthorization: kind: technical-publication-authorization authorizedBy: Cameron recordedAt: '2026-09-15T16:42:05.109Z' route: continuous-integration scope: technical-publication exactRenderReviewed: false receiptPath: knowledge/receipts/technical-publication/continuous-integration.json receiptDigest: 'sha256:37de0cd95dd6ab88a61cb62b7791543e9cf9b28bdd8a8f5a01df7792fc2924bd'
Continuous integration (CI) is the practice of frequently combining developers' changes in a shared codebase and automatically building and testing the combined result. A CI service runs those checks in a controlled environment and reports which revision it tested. When coding agents contribute, CI supplies a repeatable feedback loop they can use to diagnose failures, repair changes, and establish evidence before integration.
The practical goal is to keep the shared software working while many changes arrive. A green result means a particular set of checks passed on a particular revision in a particular environment. Its usefulness depends on whether those checks cover the intended behavior and whether the tested revision is the one that will be merged or deployed.
This guide explains the general workflow, then develops a starting design for teams using coding agents. The recommendations are engineering guidance, not a claim that every repository needs the same tools or scale. Anthropic's report provides a large-scale case study in test selection; a small project can start with one automated verification job.
Follow one change through the system #
Consider a hypothetical shopping application. An agent changes how a discount applies to a cart and adds a regression test for rounding. The agent's local test passes, but another developer is changing the representation of prices at the same time. Each change can be correct in isolation and incompatible when combined.
A typical pull-request workflow proceeds as follows:
- Prepare the change. Start from the current main branch, inspect the relevant requirements, edit the code, and run focused local checks.
- Propose integration. Push a commit and open a pull request (PR), a request to merge one branch into another.
- Run automated checks. The CI service checks out a specified revision, installs the declared toolchain and dependencies, builds the software, and runs tests.
- Inspect failures. The author, human or agent, reads logs and test artifacts, reproduces the problem where possible, and submits a correction.
- Check the combined state. Validate against the current target branch, including queued changes when using a merge queue.
- Merge under repository policy. Required checks and reviews determine whether the change is eligible. Passing tests alone does not grant an agent permission to merge.
- Verify the main branch. Post-merge checks detect problems with the integrated revision. Packaging and deployment may follow through separate jobs.
In Fowler's definition, frequent mainline integration is central to CI. Merely running tests on long-lived branches leaves integration problems unresolved. Short-lived PRs and merge queues can support frequent integration, but installing a CI service does not by itself establish the practice.
CI, delivery, and deployment #
CI checks whether changes can coexist in working software. Continuous delivery extends that discipline so the software remains releasable through an automated process, with a decision about when to release. Continuous deployment automatically releases qualifying changes to production.
These responsibilities can share a pipeline while retaining different permissions. A PR job may build and test with no production access. A later deployment job may receive credentials only after the required approval. Keeping these stages distinct prevents a request to test code from silently becoming permission to release it.
An artifact is an output saved from a run: a compiled package, container image, test report, screenshot, or trace. Where practical, promote the tested build artifact into deployment instead of rebuilding from loosely specified inputs. Record its digest and source revision so a later investigation can identify what actually ran. A deployment smoke test then checks a small set of critical behaviors in the deployed environment.
The vocabulary of a pipeline #
CI products use different names, but most contain the following parts. GitHub Actions is one concrete implementation; GitLab CI, Jenkins, Buildkite, and other services implement similar concepts.
| Term | Meaning | Example |
|---|---|---|
| Trigger | An event that starts automation | A push, PR update, merge-group event, or schedule |
| Workflow or pipeline | The declared automation | Build, test, package, and publish |
| Job | A unit scheduled onto an execution environment | Run Linux unit tests |
| Step | An operation within a job | Install dependencies, then invoke the test command |
| Runner | The machine or container executing a job | A fresh hosted Linux VM |
| Matrix | Multiple configurations of a job | Linux and Windows across supported runtime versions |
| Check | A reported result associated with a revision | Type checking passed for this commit |
| Required check | A result repository policy demands before merging | The integration-test job must succeed |
| Artifact | A retained output | A failing browser trace or release package |
| Cache | Reusable data intended to reduce repeated work | Downloaded dependencies or compiled objects |
Jobs form a dependency graph. Independent jobs can run concurrently; a packaging job may wait for several builds. Steps inside a job usually share a workspace, while files needed across jobs must be transferred explicitly through the CI platform's supported mechanisms.
“CI agent” sometimes means a runner process, not a language-model agent. Distinguishing them avoids confusion: the runner executes commands; a coding agent interprets results and may propose changes.
Choose checks by the failures they detect #
A useful test suite contains checks at several levels. Passing a cheaper level does not establish the behavior tested by a more expensive one.
| Check | What it can establish | What it commonly misses |
|---|---|---|
| Formatting and linting | Source follows selected conventions and static rules | Whether the program solves the user's problem |
| Type checking or compilation | Specified type and build constraints hold | Runtime data, configuration, and timing failures |
| Unit tests | Small pieces produce expected outputs for tested inputs | Incorrect interactions with real services |
| Integration or contract tests | Components agree on tested interfaces and behavior | Entire user journeys or production differences |
| End-to-end tests | A tested journey works through the assembled application | Untested environments and rare conditions |
| Security checks | Selected vulnerability and policy checks pass | All possible attacks or misuse |
| Performance checks | Measured workloads meet selected budgets | Unmeasured workloads or hardware differences |
For the shopping example, a unit test can establish rounding behavior. An integration test can establish that the API and database agree on price units. A browser test can establish that the displayed total changes when a coupon is applied. A test calling a helper directly does not exercise the route, serialization, authentication, or UI that invokes it.
The test oracle is the rule that decides what the correct result should be. If an agent writes both implementation and tests from the same mistaken assumption, they can agree and still be wrong. Ground expected behavior in requirements, examples, existing contracts, and independent review. For a bug fix, demonstrate that the regression test fails against the buggy version and passes after the repair.
Code coverage reports which code ran during tests. High coverage does not establish that assertions would detect incorrect behavior. Mutation testing, which deliberately changes code and checks whether tests fail, can expose weak assertions, but costs additional execution time. Use it selectively where stronger evidence is worth the cost.
Build a small, usable starting pipeline #
For a modest repository, start with a documented verification command and one job that runs it from a clean checkout. The command should install or require the correct toolchain, use locked dependencies, build the application where needed, and run its fast tests. Developers and coding agents should be able to invoke the same underlying checks locally.
Then add a small number of integration tests for critical interfaces and an end-to-end smoke test for the main user journey. Make failure logs available, save useful artifacts, and configure repository rules to require the checks that actually protect integration. Test the rule configuration with a deliberately failing change: a workflow file existing in the repository does not prove that a failed job blocks merging.
A starting division of work might look like this:
| When | Work | Reason |
|---|---|---|
| During editing | Focused regression tests and fast static checks | Short feedback while the change is still in context |
| Before merge | Required build, fast suite, relevant integration tests | Establish that the candidate is acceptable to combine |
| At merge queue | Checks against the proposed combined revision | Detect interactions with other accepted changes |
| After merge or nightly | Broader suites, platform matrix, expensive analysis | Find failures omitted from the fast path |
| Before and after deployment | Artifact validation and deployed smoke checks | Establish what was released and whether it works there |
These are suggested responsibilities, not mandatory separate jobs. Start with the smallest pipeline that detects your important failures. A nightly test protects less against a bad merge than a required pre-merge test; moving work to night changes when risk is discovered.
Give a coding agent an executable contract #
A repository should tell an agent what to run and what counts as completion. Put stable commands in scripts or the build tool, and explain their purpose in the repository's agent instructions. Avoid making every agent reconstruct the pipeline from logs and old conversations.
The contract should identify the project root, supported toolchain, setup command, fast checks, broader checks, expected artifacts, and permissions. It should also explain which changes need extra verification: database migrations, public interfaces, dependency upgrades, authentication, shared rendering code, or CI configuration itself.
A useful completion report includes the exact commit tested, executed commands, exit statuses, failing test names, skipped checks, and links to retained artifacts. A statement such as “tests pass” is incomplete if it refers to an older commit or a different checkout. If the agent changes code after testing, the previous result no longer covers the new revision.
Give the agent bounded room to repair failures, but preserve the acceptance criteria. Removing a failing assertion, skipping a test, or updating every snapshot can produce a green check while weakening the evidence. Such changes require a reason tied to intended behavior and appropriate review. Snapshot updates deserve inspection of what changed, not approval because the update command exited successfully.
The broader separation between observation, authorization, and verified effects is described in Reliable Agent Systems. For CI, it means an agent's interpretation can guide repair while repository policy and observed test results govern acceptance.
Diagnose a red build before editing #
Failure diagnosis starts with the earliest meaningful error, not the last line of a long log. Establish the checked-out revision and job environment before attributing the failure to the proposed patch.
| Failure class | Example | Useful response |
|---|---|---|
| Product regression | The new discount calculation violates an established rounding rule | Add or inspect the regression test and repair the code |
| Incorrect test expectation | A deliberately changed interface retains the old expected response | Confirm the new contract before updating the test |
| Environment mismatch | CI uses a different runtime or omitted generated file | Reproduce the clean environment and fix setup |
| Infrastructure failure | Runner disk fills or dependency download fails | Repair or retry the infrastructure with a bound |
| Flaky behavior | The same revision sometimes passes and sometimes fails | Preserve both outcomes and investigate nondeterminism |
| Integration conflict | Two passing PRs break when combined | Test the combined revision and reconcile the interaction |
A baseline run on the target branch helps distinguish an existing failure from a new one. However, “also fails on main” does not automatically make a test safe to ignore. It may reveal an already broken shared system that should block further changes.
Retries can help diagnose transient faults, but a later pass does not explain an earlier failure. Google's account describes how flaky tests consume investigation time and cause real failures to be dismissed. Its measurements are historical observations from Google, not expected rates for every project.
If a flaky test must be quarantined, keep it running outside the merge gate, assign an owner, retain its history, and set a repair deadline. Quarantine temporarily removes protection. Race conditions in production code can manifest as flaky tests, so labeling a result flaky is the start of diagnosis, not a declaration that the product is fine.
Keep feedback fast without hiding failures #
Total feedback time includes queueing, environment setup, execution, and result delivery. Adding machines helps only the part constrained by available runners. A ten-second test can still deliver poor feedback after waiting twenty minutes for a worker.
Improve the measured bottleneck in stages:
- Reuse dependencies and build outputs safely. Cache keys should reflect relevant inputs such as lockfiles, toolchain, and target platform. Periodic clean runs help detect accidental dependence on cached state.
- Run independent jobs in parallel. Avoid sharing mutable test databases, ports, or filesystem locations between workers.
- Shard large suites. Split tests among workers and balance by measured duration. The slowest shard determines completion time.
- Cancel superseded checks where safe. New commits can make older PR test runs irrelevant. Avoid canceling deployment or migration steps without considering partial effects.
- Select affected tests. Use dependency information and validated selection rules when the full suite becomes too costly.
Test sharding runs the same intended suite across more workers. Test impact analysis, or test selection, runs a subset based on the change and information about dependencies or past results. Selection reduces work but creates the risk of omitting a test that would have failed.
A conservative selection system handles unknown changes by running broader checks. New tests, lockfile changes, build configuration, shared libraries, and test-runner changes deserve explicit policies. Compare selected runs against periodic full-suite runs to measure missed failures. An LLM's unsupported opinion that a change “cannot affect that module” is weak evidence for excluding a required test.
Case study: Anthropic's test-selection service #
In its September 14, 2026 engineering post, Anthropic reports a 25-fold increase in CI jobs over six months and a tenfold increase in its test corpus. These are company-reported figures from one organization. They illustrate a possible capacity problem, rather than a growth forecast that every team should adopt.
Anthropic describes two components: a listener that records results from CI runs, and a selector that uses test history and package relevance to choose tests for new changes. When the listener lagged, the selector used stale information. The post explicitly distinguishes missing result ingestion from CI never having run.
Three interim fixes bought progressively less time: a larger machine, per-package sharding, and restarts. In the redesign, listener workers append results to a journal in an in-memory data store. A separate consumer summarizes that journal into per-test history, which the selector reads. Moving state out of individual listener workers allows those workers to scale horizontally.
The design still needs correct ordering, replay behavior, journal capacity, and consumer throughput. Moving state into a store does not make it disappear. The published article does not specify enough detail to establish the system's crash durability or delivery guarantees, so those should not be inferred from the word “journal.”
The general lesson is to monitor the feedback infrastructure itself. Measure how long it takes for a completed test result to influence the next selection decision, not just how fast tests execute. Agent-assisted development can increase PR volume, test volume, and overnight activity together. A selector with stale history may undermine the usefulness of a large testing investment.
Merge queues protect the combined revision #
Two independently green PRs can conflict semantically even when Git merges their text without complaint. A merge queue tests candidates with the latest target branch and the changes ahead of them in the queue. It reduces the need for each author to repeatedly update and resubmit a branch as other changes land.
GitHub's merge queue creates temporary merge groups for this purpose. GitHub Actions workflows providing required checks must handle the merge_group event in addition to the usual PR triggers. Otherwise, the queue may wait for a result the workflow never produces.
The important identifier is the tested revision. A check for a PR head, a synthetic merge commit, and a merge-group commit can refer to different code. An agent deciding whether work is ready must inspect that relationship rather than collect green badges from unrelated runs.
CI executes code with authority #
Build scripts, dependency installation, test imports, and third-party actions can execute arbitrary code. Treat a proposed change as untrusted until the applicable review and isolation requirements are met, even when an internal coding agent authored it.
GitHub's security guidance recommends least-privilege tokens, immutable action references, and caution with privileged triggers and self-hosted runners. A practical separation is to run untrusted PR checks without production secrets, then give a separately controlled release job the minimum permissions it needs.
Do not check out and execute an untrusted PR in a privileged pull_request_target workflow. Artifacts passed from an untrusted workflow into a privileged one also need validation. Use short-lived identity-based cloud credentials where supported, and review changes to workflow definitions as changes to executable infrastructure.
Coding agents add another input channel: logs, source comments, issue text, and test output can contain instructions. Treat those contents as evidence to analyze, not authority to reveal credentials, disable checks, or run unrelated commands. Limit the agent's tools and network access to the repair task. A prompt asking the agent to be careful does not replace isolation or token permissions.
When the product itself uses AI #
Using AI to write software and testing software that calls AI are different problems. A conventional test suite should still cover deterministic components: parsing, persistence, permissions, routing, timeouts, and error handling. Stub model responses for those tests so every small change does not require paid, variable inference.
Live model behavior needs a separate evaluation layer. Anthropic's guide distinguishes a task, repeated trials, graders, and the outcome in the environment. Record the model and provider configuration, prompt and dataset versions, scoring rule, sampling settings, cost, and run identifiers. Evaluate on cases that were not simply copied into the implementation prompt. Where outputs vary, inspect performance across repeated trials rather than treating one successful sample as proof.
A model-based grader can help assess open-ended responses, but its rubric and errors also need evaluation. Keep deterministic checks for exact requirements and human review for consequential ambiguous cases. Model upgrades and provider changes can alter behavior even when application code is unchanged, so scheduled evaluations may be useful alongside commit-triggered checks.
Live evaluations should run against disposable test resources with explicit spend limits. They should not send real customer messages or modify production accounts merely because the application has those capabilities. A fixture can test the proposed action; a separately authorized integration test can test the real transport.
What to measure and what to build first #
Track feedback latency, queue time, runner cost, flaky failures, escaped regressions, and time spent restoring a broken main branch. When using test selection, track missed failures and history freshness too. When using repair agents, distinguish accepted repairs from suggestions, and verify that their changes preserve coverage and permissions.
Begin with a clean-checkout verification command, meaningful regression tests, protected integration, and readable failure artifacts. Add parallelism, sharding, selection, and automated repair when observed workload justifies them. The useful outcome is a working shared codebase with timely, trustworthy evidence, rather than the maximum number of tests or agent-generated commits.