Blue Green Deployment: Switch Systems With No Downtime
Blue green deployment explained by an engineer who led a live platform migration: traffic switching, database changes, rollback plans and when to use canary instead.

On this page · 7 sections
- What is blue green deployment?
- How does traffic move from the old environment to the new one?
- How does blue green deployment help when replacing a live system?
- What happens to the database during a blue green switch?
- How do you roll back if the new version fails?
- Blue green vs canary: which one should you use?
- Want a second opinion on your switch?
Key takeaways
- Blue green means two production-grade environments and a single switch in front of them.
- The switch is easy. Shared data and background jobs are where releases go wrong.
- Rollback is only fast if the database changes are backward compatible.
- Keep blue alive and untouched until you trust green.
- Blue green moves everyone at once. Canary moves a few users first. Many teams use both.
Blue green deployment is a release technique where you run two full production environments side by side: blue, which serves users today, and green, which holds the new version. Once green is tested, you move traffic from blue to green in one controlled switch.
Done well, users should not notice downtime, because blue keeps running the whole time. If green misbehaves, you point traffic back. It is one of the simplest ways to replace a live system, with one hard part: the database. Below I cover how the switch works, where it fits in a migration, and where it does not.
What is blue green deployment?
You maintain two environments that are as identical as you can make them. Blue is live. Green is idle, or receives only internal test traffic. You deploy the new version to green, run smoke tests, then flip the router so green becomes live. Blue becomes your standby. On the next release, the roles swap.
What you need before you try it
- A switch layer: a load balancer, reverse proxy, API gateway, DNS record, or a platform feature that decides where traffic goes.
- Matching environments: same infrastructure, same configuration shape, same secrets handling. Drift between blue and green is how surprises happen.
- Health checks: automated tests that prove green works before it takes real users.
- Monitoring: error rates, latency and at least one business metric, so you can tell whether the switch hurt. Having worked on the cart and checkout team at Total Wine & More, I always watch a money path during a switch, not just technical metrics.
If you cannot build the second environment repeatably, fix that first. Manual setup makes blue green slow and risky.
How does traffic move from the old environment to the new one?
Traffic moves wherever your routing layer lives. The common options differ mainly in how fast and how cleanly they switch.
| Method | How it switches | Watch out for |
|---|---|---|
| Load balancer target swap | Point the listener at the green group | Connection draining on blue |
| Reverse proxy or gateway | Change the upstream, reload config | Config errors in the proxy itself |
| DNS change | Update the record to the green address | Caching means some users stay on blue for a while |
| Kubernetes service selector | Change the label the service points at | Readiness probes must be accurate |
Whatever you choose, think about requests already in flight. Let blue finish them (connection draining) before you stop sending it traffic. Also check sticky sessions, in-memory caches and long-lived connections such as websockets. A user logged in on blue should not be logged out by the switch.
How does blue green deployment help when replacing a live system?
It gives you a moment where the old and new systems both exist and you choose which one users see. That is the safest shape for a migration, because the decision is reversible.
I led the full migration of a platform at avanzzada that applies to job openings automatically for its users, moving its whole architecture from a legacy stack to a new one while people kept using it every day. The lesson I took from that kind of work: put the switch in front of everything before you write new code. Once you can route traffic, every later step becomes a small, reversible decision instead of a launch night.
Where it fits, and where it does not
- Good fit: replacing a whole service or application when the old and new can run in parallel.
- Weaker fit: a large system where the new stack cannot yet do everything. There, move piece by piece, and use blue green per route or per service.
- Cost: you pay to run two environments during the overlap. In my experience that is usually cheaper than a big-bang rewrite, which has to keep the old system running in parallel anyway.
One trap is background work. If both environments run the same scheduled jobs or queue consumers, users can get duplicate emails, charges or automated actions. Decide which environment owns the workers, and turn them off in the other. When I take over a codebase, my order is access first, then the build, then the money paths, then the risks, and an inventory of jobs and consumers is one of the first risks I list.
What happens to the database during a blue green switch?
This is the real difficulty. Two application versions usually share one database, so the schema has to work for both at the same time.
Option 1: shared database, backward-compatible changes
The most common approach, and the one I default to, is expand and contract:
- Expand: add new columns or tables without removing anything. Blue still works.
- Deploy green and switch traffic. Green uses the new structure and tolerates the old.
- Contract: after you are sure you will not roll back, remove the old columns in a later release.
Never combine a destructive schema change with the traffic switch. That removes your rollback.
Option 2: separate databases
When the new system has a different data model or database engine, green gets its own database. You then need to keep both in sync until the switch, using replication or change data capture, and often a short write freeze at cutover. It is more work and carries more risk, so I treat it as its own project and rehearse it on a copy of production data first.
How do you roll back if the new version fails?
Rolling back means sending traffic to blue again. If you prepared for it, it usually takes seconds, and it is the main reason to use this technique.
- Define triggers in advance. For example: error rate above a threshold you agreed on, failed logins, or a drop in a key business action such as checkouts or submissions. Do not debate this during an incident.
- Keep blue running and patched for a set observation window. Shutting it down right after the switch removes your safety net.
- Protect the data. Anything green wrote must still be readable by blue. That is why backward-compatible schema changes matter. With separate databases, you may need to sync data back.
- Rehearse. Practice the rollback in staging. An untested rollback is a guess.
Also note the limits. A rollback cannot undo side effects already sent to the outside world, such as emails or payments. Idempotent operations and good logs help you clean those up.
Blue green vs canary: which one should you use?
Both reduce release risk, but they trade different things.
- Blue green: all users move at once, rollback is a single switch, and the logic is easy to reason about. It costs a second full environment, and a hidden bug hits everyone before you notice.
- Canary: a small share of users gets the new version first, so a bug affects fewer people. It needs weighted routing and solid metrics to judge the result, and it means two versions serve users for longer.
My rule of thumb, as an estimate and not a standard: use blue green when you want a clean, reversible cutover and can run both environments cheaply. Use canary when the change is risky for real user behavior and you have the monitoring to read the results. Combining them works well: verify green with test traffic, then shift users in steps through weighted routing, keeping blue ready until the end.
Want a second opinion on your switch?
If you are planning a cutover and are unsure about the database or rollback plan, email me@filipeeduardo.dev with the details. I am glad to compare notes and tell you what I would check first.
Frequently asked questions
Does blue green deployment really mean zero downtime?
It aims for no visible downtime, but it is not a guarantee. Dropped connections, cache differences, schema mismatches or a failed health check can still cause errors. Draining, testing and monitoring reduce those risks.
Is blue green deployment expensive?
You pay for a second environment while both exist. On cloud infrastructure you can create green for the release and tear it down afterward, which limits the cost. Compare it with the cost of a failed release or a long rewrite, not with zero.
Can I use blue green with a monolith?
Yes. It works at the application level, so a monolith is fine as long as you can run two copies and handle the shared database carefully. Monoliths with heavy local state or file storage need extra thought.
What is the difference between blue green and rolling deployment?
A rolling deployment replaces instances one by one inside a single environment. Blue green keeps a complete second environment and switches at once, which makes rollback faster and cleaner.

