A Practical Checklist for Zero-Downtime Migrations
Published: August 12, 2026
Moving production workloads without an outage window is a discipline, not a lucky break. Here is the exact sequence our engineers follow.
Start with traffic, not servers
The most common migration failure is treating the move as a server problem when it is actually a traffic problem. Before a single workload is replicated, we map every path a request can take: DNS records, load balancers, CDNs, internal service discovery, and long-lived connections such as WebSockets and database pools.
Each of these paths becomes a line item on the migration board. If a path cannot be drained or redirected, it defines the shape of the whole cutover.
Replicate, then verify, then cut over
We never migrate state by copying it once. Instead, the target environment runs in parallel with continuous data replication, and we shift traffic gradually — usually 1%, then 10%, then 50% — comparing error rates and latency between environments at every step.
The old environment stays warm and authoritative until the new one has served a full daily traffic cycle without regression. Rollback is a routing change, not a restoration.
The pre-cutover checklist
- —Data replication lag measured in seconds, not minutes, under production load
- —Observability parity: dashboards, alerts, and log pipelines live on the target before traffic moves
- —Backups restored on the target and verified, not assumed
- —A written rollback plan with a decision owner and a time limit
- —Internal stakeholders notified with a runbook, not a vague maintenance window
Decommission deliberately
The migration is not finished when the new environment is live — it is finished when the old one is off the bill. We keep the source environment in a reduced, read-only state for at least two billing cycles, then archive its snapshots and decommission it. That final step is where most of the savings live, and it deserves the same rigor as the cutover itself.