Scale with multiple replicas

Split a self-hosted Quackback instance into web and worker replicas with QUACKBACK_ROLE, run migrations once, and wire health checks per role. No Redis or sticky sessions needed.

JM
James Morton
Written By James MortonLast updated about 2 hours ago

Run a self-hosted Quackback instance across several replicas by splitting HTTP serving from background processing. One container handles everything at first; split roles as traffic grows so a slow AI job never competes with a page load.

How Quackback processes work

Every Quackback process reads QUACKBACK_ROLE to decide what it does. The default, all, serves HTTP and runs background workers in the same process. That's what the Docker and Railway guides deploy, and it's the right choice until you have a specific reason to split.

Role

HTTP

Queue workers

Use for

all (default)

Yes

Yes

Single-container deployments

web

Yes

No (enqueues only)

Front-end replicas behind a load balancer

worker

Health probes only

Yes

Dedicated background processing

Note: web replicas serve HTTP but never claim queued jobs. They enqueue jobs and leave them for a worker. A deployment that only runs web replicas queues work that nothing ever processes.

Split web and worker replicas

Run the same image with different QUACKBACK_ROLE values:

# Web replicas: scale these behind your load balancer
docker run -d -e QUACKBACK_ROLE=web -e SKIP_MIGRATIONS=true -e DATABASE_URL=... -e SECRET_KEY=... -e BASE_URL=... ghcr.io/quackbackio/quackback:latest

# Worker replicas: one is enough for most workloads
docker run -d -e QUACKBACK_ROLE=worker -e SKIP_MIGRATIONS=true -e DATABASE_URL=... -e SECRET_KEY=... -e BASE_URL=... ghcr.io/quackbackio/quackback:latest

Tip: Route load-balancer traffic to web replicas only. worker replicas serve HTTP so their health probes work, but they shouldn't receive user traffic.

Background workers handle event delivery (webhooks, notifications, workflows), email sending and inbound polling, AI jobs (summaries, duplicate detection, extraction), analytics refreshes, workflow timers, and imports. Jobs live in a queue table in PostgreSQL, so running more than one worker replica is safe: each job and each scheduled sweep runs once. There is no Redis or separate queue service.

Warning: If you scale web replicas without a worker (or all) replica running somewhere, webhooks, notifications, and workflows stop firing. Jobs queue up but nothing consumes them.

With Docker Compose

The bundled docker-compose.prod.yml is single-node by design (the app publishes port 3000 directly). To split roles with Compose, keep the same postgres and minio services, replace the app service with a web service (with QUACKBACK_ROLE=web and several replicas) and a worker service (with QUACKBACK_ROLE=worker), set SKIP_MIGRATIONS=true on both, and put a reverse proxy or load balancer in front of the web service. Web replicas can't each publish port 3000 on the host.

Requirements

Every replica, whatever its role, needs:

  • The same DATABASE_URL (shared PostgreSQL, which also holds the job queue)
  • The same SECRET_KEY (sessions are stored in the database, so any replica can serve any request)
  • The same S3-compatible storage and email settings

Sticky sessions aren't required. Realtime updates travel through PostgreSQL LISTEN/NOTIFY, so DATABASE_URL must be a direct or session-mode connection.

Each process opens its own connection pool (DB_POOL_MAX, default 10 for web and all, 20 for worker). Make sure PostgreSQL's max_connections covers all replicas.

Run migrations once

With one replica, migrations run on startup. With several, set SKIP_MIGRATIONS=true on every replica so they don't race, and run migrations as a one-off container before you roll out a new image:

docker run --rm \
  -e DATABASE_URL="postgresql://user:pass@host:5432/quackback" \
  --entrypoint bun \
  ghcr.io/quackbackio/quackback:latest \
  /app/migrate.mjs

Then roll the new image out to the web and worker replicas.

Health checks per role

Point your orchestrator's liveness probe at /api/health/live (process is up, no I/O) and its readiness probe at /api/health/ready (database, migrations, and workers). A web replica doesn't run workers, so its readiness check reports workers.expected: false and still passes. See Health endpoints.

Was this helpful?

Your feedback shapes what we write next.