Test environments and data, with Docker
Tests need somewhere to run and something to run against — an environment and data. Get these wrong and even well-written tests become unreliable: they pass locally and fail in CI, or fail because someone changed the shared test database, or because the environment drifted from production. Controlling the environment and the data is a large part of what makes a suite dependable. This lesson is test environments and test data, and how Docker helps make them consistent.
The environments a change moves through
Software usually flows through a series of environments, and a QA works across them:
- Local / development — the developer's machine, where they build and run tests as they work.
- CI — the clean server that runs the suite on every change (the running-in-ci lesson).
- Test / staging — a deployed environment that mirrors production, for end-to-end and manual testing before release.
- Production — the real thing, used by real users. You test against production carefully and mostly read-only.
The recurring danger is drift: when these differ (a different database version, different config, different data), a test passes in one and fails in another, and you waste hours on a "bug" that is really an environment mismatch. The goal is to make the environments as consistent as possible, so a test result means the same thing everywhere. This is the "works on my machine" problem, and controlling the environment is how you kill it.
Docker: the same environment everywhere
Docker packages an application and everything it needs — runtime, libraries, config — into a container that runs identically on any machine. For testing, this is transformative: instead of "install Postgres 15 and configure it just so", you run a container that is Postgres 15, configured, the same on your laptop and in CI.
Why it matters for a QA:
- Consistent dependencies — the database, the message queue, the cache your app needs run as containers, identical everywhere. No "it failed because CI had a different Postgres".
- Disposable and clean — spin a fresh container up for a test run, throw it away after. Every run starts from a known-clean state, which is exactly what reliable tests need.
- Reproducible — a
docker-compose.ymldescribes the whole set of services, so anyone (and CI) brings up the identical stack with one command.
AgentPay deliberately needs no database — it keeps data in memory and reseeds — precisely so the course can run without any infrastructure. But a real app has a database, and there Docker is how you give every environment the same one. A typical setup runs the app's dependencies as containers in CI before the tests, so the suite runs against a known stack.
Test data, again — the environment's other half
The environments lesson is incomplete without data, because a controlled environment with uncontrolled data is still flaky. The principles from the fixtures lesson scale up here:
- Each environment owns its data. Never test against a shared database that humans also use — someone will change a record your test depends on. Give tests their own data.
- Reset to a known state before a run (AgentPay's reseed; for a real DB, migrate-and-seed a fresh container), so results are deterministic.
- Seed realistic and sufficient data — enough rows to exercise pagination and performance (the dashboards lesson), not three demo records.
- Isolate parallel runs — a database per CI job or per worker, or unique data per test, so concurrent runs do not collide (the fixtures lesson's trade-off).
The combination — a containerised, consistent environment and controlled, reset, isolated data — is what makes a suite give the same answer every time. Flaky "environment" failures almost always trace back to one of these being uncontrolled.
Pointing tests at an environment
Finally, a test suite should be able to run against any environment without code changes — local, CI, staging — by configuration:
- A base URL / connection from an environment variable. AgentPay's Playwright
baseURLand the API'sAPI_URLare set by config, so the same suite can point at localhost or a staging deployment by changing one value (the setup lesson). Never hard-code the environment into tests. - Environment-specific config, not environment-specific tests. The tests stay the same; only the configuration (URLs, credentials, which data) changes. This is what lets you run smoke tests against staging with the very suite you run locally.
Keeping the environment in configuration, and the environment itself consistent via Docker, means a test is a statement about behaviour, not about a particular machine — which is the whole point.
Check your work
Environments: local/dev, CI, test/staging, production (test against carefully, mostly read-only). The danger is drift — differences that make a test pass in one and fail in another ("works on my machine"). Goal: make environments consistent so a result means the same everywhere.
Docker packages an app + its dependencies into a container that runs identically anywhere: consistent
dependencies (same DB version everywhere), disposable/clean (fresh state per run), reproducible (one
docker-compose brings up the stack). AgentPay needs no DB by design (in-memory reseed) so the course runs
infra-free; a real app runs its dependencies as containers.
Data (the other half): each environment owns its data (never a shared human-used DB); reset to a known state; seed realistic and sufficient data (enough for pagination/perf); isolate parallel runs. Uncontrolled data = flaky "environment" failures.
Point tests by config: base URL / connection from env vars (AgentPay's baseURL/API_URL); the same
suite runs against local/CI/staging by changing configuration, never code. Environment in config + consistent
via Docker = a test is about behaviour, not a machine.
Practice
- List the environments a change passes through and what a QA does in each.
- Explain "environment drift" with an example, and how Docker reduces it.
- Describe how you would use a
docker-compose.ymlto give CI the same database as your laptop. - Explain why AgentPay needs no database, and what a real app would do instead.
- List four rules for test data across environments and why each prevents flakiness.
- Show how AgentPay's
baseURL/API_URLconfig lets the same suite run against local and staging.
Official documentation
- Docker — Get started — Containers and images, and why they make environments consistent.
- Docker — Compose overview — Describing a multi-service stack for tests.
- Playwright — Test configuration (
baseURL) — Pointing the suite at an environment by config.
Next: maintaining a suite people trust.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship