Running the suite in CI on every change
A test suite that only runs when someone remembers to run it is barely a suite at all. The power of automation comes when the tests run automatically, on every change, so a bug is caught the moment it is introduced — not days later when someone finally runs the suite. That is continuous integration (CI): your tests run on a server every time code is pushed. This lesson is getting your suite running in CI, using AgentPay's real GitHub Actions workflow as the model.
What CI is, and why it matters for testing
Continuous integration means: every time someone pushes code (or opens a pull request), an automated system checks out the code, builds it, and runs the tests — automatically, on a clean server. For a QA this is where automation delivers its value:
- Every change is tested — no reliance on a human remembering. A regression is caught on the pull request that caused it, before it merges.
- A red build blocks the merge — you can require the tests to pass before code can be merged, so broken code never reaches the main branch. This is the mechanism that actually protects the product.
- It runs on a clean, known environment — not "works on my machine", but a fresh server every time, which catches environment-dependent bugs and keeps results honest.
- Fast feedback — the developer learns within minutes that they broke something, while the change is fresh in their mind and cheap to fix (the cost-of-bugs lesson).
Automated tests that do not run in CI are a fraction as valuable as ones that do. CI is what turns a suite from "something we could run" into "the thing that guards every change".
A real CI workflow
CI is configured with a file in your repo. GitHub Actions (what AgentPay uses) reads .github/workflows/*.yml.
Here is AgentPay's actual workflow, and it is worth reading line by line because it is exactly the shape you
will write:
name: CI
on:
push:
branches: [main]
pull_request: # run on every PR and every push to main
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4 # get the code
- uses: actions/setup-node@v4 # install Node
with: { node-version: 20, cache: npm }
- run: npm ci # install dependencies (clean, from the lockfile)
- name: Typecheck
run: npm run typecheck
- name: API tests
run: npm run test:api # the 21 Vitest API tests
- name: Install Playwright browser
run: npx --workspace=@agentpay/web playwright install --with-deps chromium
- name: End-to-end tests
run: npm run test:e2e # the 12 Playwright E2E tests (starts the servers itself)
- uses: actions/upload-artifact@v4 # keep the report even on failure
if: ${{ !cancelled() }}
with: { name: playwright-report, path: apps/web/playwright-report/ }
Read what it does: on every push to main and every pull request, a fresh Ubuntu machine checks out the
code, installs Node and dependencies, typechecks, runs the API tests, installs the browser, runs the E2E
tests, and uploads the Playwright report. Every one of those steps must pass for the build to be green. This
is a complete, honest CI pipeline for a real test suite.
The pieces that make it work
A few things in that file are worth understanding, because they are what make CI reliable:
npm ci, notnpm install— installs exactly what the lockfile specifies, so CI runs the same dependency versions every time. Reproducibility is the whole point of CI.- The E2E suite starts its own servers — AgentPay's Playwright config has
webServer, so CI does not need a separate step to start the app; the suite is self-contained (the setup lesson). This is why it "just works" on a clean machine. - Installing the browser — Playwright's browsers are not in the repo, so CI installs them
(
playwright install --with-deps). A common first-CI-run failure is forgetting this. - Uploading the report even on failure (
if: !cancelled()) — so when a test fails in CI, you can download the Playwright report and trace and see why (the debugging lesson). Without this, a CI failure is a dead end.
That last point matters most for a QA: CI is only useful if a failure is actionable. Uploading reports, traces and screenshots is what lets someone diagnose a failure they cannot see on the CI server.
Making the tests a required check
Configuring the workflow makes the tests run; the final step is making them matter. In the repository
settings you mark the CI check as required for merging — now a pull request cannot merge while the tests
are red. This is the mechanism that actually protects main: the suite is not advisory, it is a gate. A QA
often owns advocating for this, because a suite that can be ignored eventually is — and required checks are
how automation becomes a real safeguard rather than a suggestion.
Check your work
CI runs your tests automatically on every push/PR, on a clean server. Value: every change is tested (no human memory needed), a red build blocks the merge (broken code never reaches main), a known clean environment (no "works on my machine"), and fast feedback while the change is fresh. Tests not in CI are a fraction as valuable.
A workflow (.github/workflows/*.yml) checks out code, installs Node + deps, and runs the checks.
AgentPay's runs on push-to-main and every PR: checkout → setup-node → npm ci → typecheck → API tests →
install browser → E2E tests → upload report.
Reliability pieces: npm ci (exact lockfile versions, reproducible), the E2E suite starting its own
servers (self-contained on a clean machine), installing Playwright browsers (a common miss), and uploading
the report/trace even on failure so a CI failure is actionable.
Make it a gate: mark the CI check required so a PR cannot merge while red — this is what actually
protects main. A suite that can be ignored will be; required checks make automation a real safeguard.
Practice
- Read AgentPay's
.github/workflows/ci.ymland describe what happens on each push and PR, step by step. - Explain why
npm ciis used instead ofnpm installin CI. - Identify the step that installs the Playwright browser and explain what fails without it.
- Explain why the report is uploaded with
if: !cancelled()and what a QA does with it after a failure. - Write a minimal GitHub Actions workflow that installs deps and runs a test command on every PR.
- Explain what "required status check" means and why a QA would advocate for it.
Official documentation
- GitHub Actions — Quickstart — Writing a workflow file.
- Playwright — Continuous Integration — Running Playwright in CI, including browser install.
- GitHub — Required status checks — Making the tests a merge gate.
Next: reporting so a failure is actionable.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship