Docker Compose Healthcheck: Postgres Ready Before App Start

Production deployments still fail in 2026 because developers trust depends_on without healthchecks. That’s the whole problem โ and it’s more widespread than most engineering teams want to admit.
Key Takeaways
- Docker’s
depends_onwithoutcondition: service_healthyonly waits for a container to start, not for Postgres to accept connections โ creating race conditions that surface unpredictably in production.- The
pg_isreadyhealthcheck pattern eliminates the most common class of container startup failures across multi-service stacks.- According to Docker documentation (2025),
service_healthyconditions require an explicithealthcheckblock on the dependency service โ omitting either half silently breaks the mechanism.- Teams running multi-container stacks on platforms like Railway, Render, or self-hosted VMs report that startup race conditions account for a disproportionate share of “works locally, fails in CI” bugs.
- A properly configured healthcheck adds 3โ10 seconds to cold start time โ a worthwhile trade against app crashes and manual restarts.
1. The Problem Nobody Reads the Docs About
Postgres takes time to be ready. Not just running โ actually ready to accept TCP connections, initialize the data directory, and process queries. That window is typically 2โ8 seconds on a cold container start. On slower VPS hardware or network-attached storage, it can stretch past 20 seconds.
Most Docker Compose tutorials show depends_on and stop there. That’s the trap.
depends_on in its default form tells Docker: “start this container after that one.” It says nothing about whether Postgres is actually accepting connections. So your Node, Python, or Go app boots up, tries to connect on port 5432, gets ECONNREFUSED, and either crashes or enters a broken state that’s annoying to diagnose and embarrassing to explain.
This matters more now because container-native deployments are the default across most mid-size engineering teams. According to the 2025 Stack Overflow Developer Survey, over 78% of professional developers use containers in their workflow โ up from 69% in 2023. More teams running containers means more teams hitting this exact startup race condition. And more of those teams are deploying to ephemeral environments โ CI pipelines, preview deployments, autoscaling groups โ where cold starts happen constantly.
The pg_isready healthcheck pattern exists specifically to close this gap. It’s not new. But it’s still misconfigured far more often than it should be.
What this analysis covers:
- How the race condition actually manifests and why
depends_onalone fails - The correct healthcheck configuration using
pg_isready - A comparison of common approaches โ restart policies, healthchecks, wait scripts
- Practical recommendations for production and CI environments
The short version: depends_on without a health condition is a silent failure waiting to happen. The fix is a two-part pattern โ a healthcheck block on the Postgres service and condition: service_healthy on the app โ and it takes under 10 lines of YAML.
Three things to understand before diving in:
- Postgres containers report as “started” while still initializing their data directory, making them unavailable for connections for several seconds.
- The
pg_isreadyutility (bundled with every official Postgres image) is the most reliable way to check actual connection readiness from inside a healthcheck. - Production configurations should tune
interval,timeout,retries, andstart_periodbased on the environment’s cold start characteristics.
2. Background: How Docker Compose Dependencies Actually Work
Docker Compose introduced depends_on back in version 2 of the Compose spec, and the intent was clear: control startup order. But the implementation had a gap. The original depends_on only checked container state โ started, healthy, or completed โ and the default was just started.
That default persisted for years. Countless tutorials copied it. Stack Overflow answers cemented it. The result: a generation of docker-compose.yml files that look correct but fail under load or on slow hardware.
Compose v2 (which became the canonical format around 2022โ2023) introduced condition as a proper dependency qualifier. Three values matter:
condition: service_startedโ the old default. Container is running. No readiness check.condition: service_healthyโ container has passed itshealthcheck. This is what you want.condition: service_completed_successfullyโ for one-shot init containers.
The service_healthy condition only works if the target service has a healthcheck block defined. Write condition: service_healthy without a corresponding healthcheck on Postgres, and Docker treats the dependency as immediately satisfied. Silent failure, again. This behavior is documented in the Docker Compose specification, but easy to miss when copying snippets from a tutorial at midnight.
The pg_isready utility โ shipped with every official Postgres Docker image โ is purpose-built for this check. It probes the Postgres socket and returns exit code 0 only when the server is ready to accept connections. That’s exactly what a healthcheck test command needs.
3. Main Analysis
The Race Condition in Detail
When you run docker compose up without healthchecks, the sequence looks like this:
- Postgres container starts.
- Docker marks it
started. - App container starts (because
depends_onsays “after Postgres starts”). - App tries to connect โ Postgres is still running
initdbor loading WAL. - Connection fails. App crashes or logs errors.
The timing window is narrow but consistent. On a typical 2-core VPS with a fresh Postgres 16 image, data directory initialization takes roughly 3โ6 seconds. Your app’s connection attempt often lands squarely inside that window.
The fix is making Docker wait for a real readiness signal before starting the app. That’s the core of this pattern.
The Correct YAML Configuration
This is the configuration that works:
services:
db:
image: postgres:16
environment:
POSTGRES_USER: appuser
POSTGRES_PASSWORD: secret
POSTGRES_DB: appdb
healthcheck:
test: ["CMD-SHELL", "pg_isready -U appuser -d appdb"]
interval: 5s
timeout: 5s
retries: 5
start_period: 10s
app:
image: myapp:latest
depends_on:
db:
condition: service_healthy
environment:
DATABASE_URL: postgres://appuser:secret@db:5432/appdb
A few things to understand here. The start_period tells Docker not to count failed healthchecks against the retry limit during the first 10 seconds โ giving Postgres time to cold-start without triggering a premature unhealthy state. With interval: 5s and retries: 5, the total grace period before marking unhealthy is about 35 seconds (start_period + retries ร interval).
The -U appuser -d appdb flags on pg_isready are critical in production. Without them, pg_isready only checks if Postgres is listening โ not if your specific database and user are accessible. Migrations running as appuser can still fail if the user isn’t provisioned yet. This is the detail most tutorials skip.
What Breaks When You Get This Wrong
Three failure modes show up repeatedly in production:
Missing healthcheck block. You write condition: service_healthy on the app but forget to add healthcheck to the Postgres service. Docker skips the wait entirely. App crashes on startup. This is the most common misconfiguration โ easy to introduce when someone copies half a pattern from different sources.
Using restart: unless-stopped as a workaround. Some teams skip healthchecks and add restart: unless-stopped to the app service. The app crashes, Docker restarts it, Postgres eventually becomes ready, and the third or fourth restart succeeds. It works, but it’s fragile. You’re adding 10โ30 seconds of crash-loop delay, polluting logs, and masking the root cause. In CI, it can trigger false failures before the eventual successful retry โ which creates the kind of flaky pipeline behavior that erodes trust in the entire test suite.
Overly short timeout. A timeout: 1s on a slow network or NFS-backed volume means pg_isready itself times out before Postgres responds, even when Postgres is perfectly healthy. Set timeout to at least 5s in production environments.
Comparison: Three Approaches to the Startup Race Condition
| Approach | Reliability | Complexity | CI Suitability | Cold Start Overhead |
|---|---|---|---|---|
depends_on (default, no healthcheck) | Low โ race condition always possible | Minimal | Poor | None |
restart: unless-stopped on app | Medium โ works eventually | Low | Poor โ noisy logs, slow | 10โ30s crash-loop |
pg_isready healthcheck + service_healthy | High โ deterministic | Low-Medium | Excellent | 3โ10s (tunable) |
Custom wait script (wait-for-it.sh) | High | Medium-High | Good | 3โ15s |
| Application-level retry logic | High (with circuit breaker) | High | Good | None extra |
The healthcheck pattern wins on reliability-to-complexity ratio. Custom wait scripts like wait-for-it.sh or dockerize were popular before Compose’s service_healthy condition matured โ they work, but they require injecting an extra binary into your app image or wrapping your entrypoint. That’s extra maintenance surface with no real advantage over the native approach.
Application-level retry logic with exponential backoff is worth adding alongside the healthcheck, not instead of it. Defense in depth. Your app shouldn’t assume the database is permanently reachable just because startup succeeded.
4. Practical Implications
For teams running local dev environments: The healthcheck pattern works identically in development and production. Add it to your base docker-compose.yml, not just a docker-compose.prod.yml. If it’s only in production, your dev environment becomes a different beast โ and you’ll debug production-only startup issues without being able to reproduce them locally. Consistency matters more than it seems.
For CI/CD pipelines: GitHub Actions, GitLab CI, and CircleCI all run Docker Compose natively. The service_healthy condition means your test runner waits for Postgres to be genuinely ready before running migrations or tests. Without it, you’re adding sleep 10 hacks to your pipeline YAML โ fragile, embarrassing to maintain, and invisible to anyone reviewing the repository for the first time.
Autoscaling and container restarts: Consider an app container on a platform like Render or Fly.io that gets OOM-killed and restarts. If the Postgres container is already running and healthy, the healthcheck passes almost immediately on the app’s restart โ typically within one interval cycle, so about 5 seconds. Clean, predictable recovery. The crash-loop approach can’t guarantee that timing, and in high-traffic scenarios, those extra seconds matter.
Docker Compose on a VPS: Teams running docker compose up -d on a DigitalOcean Droplet or Hetzner server after a reboot need deterministic startup order. Without healthchecks: reboot โ services start in parallel โ app crashes โ manual intervention at 2am. With the healthcheck pattern: the app simply waits. No intervention needed.
This approach isn’t always sufficient. If your Postgres container is healthy but your database migrations haven’t run yet, pg_isready will pass while your app still fails. For migration-dependent apps, combine the healthcheck pattern with an init container or a migration job that runs with condition: service_completed_successfully. The healthcheck solves the connection readiness problem โ not the schema readiness problem. Know the difference.
What to watch: Docker’s Compose v2.x roadmap continues to evolve healthcheck semantics. Better observability in docker compose ps output and improvements to start_period behavior are both flagged in 2025 Compose specification discussions. Worth tracking if you’re standardizing this pattern across a larger team.
5. Conclusion & Future Outlook
The pattern isn’t complicated. Two additions to your YAML: a healthcheck block on Postgres and condition: service_healthy on your app. Ten lines. Deterministic startup. No more crash loops at 2am.
Key takeaways:
depends_onalone doesn’t wait for Postgres readiness โ it only waits for container start, which is a meaningfully different thing.pg_isready -U <user> -d <dbname>is the correct healthcheck command โ the flags matter in production.start_periodprevents prematureunhealthystates during cold starts; don’t omit it.- Restart-policy workarounds eventually succeed but are fragile, slow, and obscure the underlying problem.
- Healthchecks solve connection readiness โ pair them with migration handling for full coverage.
As more teams move to declarative infrastructure, expect healthcheck patterns to become part of standard linting toolchains for Compose files. Tools like Hadolint already flag Dockerfile issues โ similar static analysis for docker-compose.yml is a logical next step. By late 2026, CI platforms may warn on missing healthchecks by default, the same way they now flag exposed secrets or unpinned base images.
The one action worth taking today: audit your docker-compose.yml files. If you see depends_on without condition: service_healthy, that’s a crash waiting for slower hardware to trigger it. Fix it before production does it for you.
What does your current Compose setup look like โ healthchecks in place, or still relying on restart policies? Drop your setup in the comments.
References
- How to Create Docker Compose Health Checks
- Docker Compose Service Dependencies: Solving Database Startup Sequence with Healthchecks ยท BetterLin
Photo by Rubaitul Azad on Unsplash


