# Canary Deployment: Release to a Few Users First

> Canary deployment explained: how to pick canary users, which metrics to compare, when to promote or roll back, and how canaries de-risk a live platform migration.

- Author: Filipe Eduardo, Senior Software Engineer (https://filipeeduardo.dev/)
- Published: October 6, 2026 · Updated: October 7, 2026
- Topic: Legacy modernization · Tags: canary deployment
- Canonical URL: https://filipeeduardo.dev/blog/canary-deployment

![Illustration of canary deployment: a few users routed to a new server topped by an orange canary bird](https://cms.filipeeduardo.dev/wp-content/uploads/2026/10/canary-deployment-1.webp)

## Key takeaways

- A canary sends a small share of real traffic to the new version and compares it with the old one.
- Choose the canary group deliberately: internal users first, then low-risk customers, then everyone else.
- Decide your health metrics and rollback thresholds before the release, not during it.
- Watch errors, latency, and business outcomes such as completed checkouts or sign-ups, not just server health.
- You don’t need a fancy platform to start. A load balancer rule or a feature flag can be enough.
- Canaries work best inside a larger incremental migration, not as a replacement for testing. Pair it with refactoring without changing behavior for safer changes.

A **canary deployment** releases a new version of your software to a small slice of real traffic while everyone else stays on the current version. If the canary behaves well, you widen the slice step by step. If it doesn’t, you send traffic back and only a few users were affected.

The name comes from the canary in the coal mine: a small early warning that something is wrong before the whole group is exposed. I use it most when changing a system that real people rely on every day, because it turns a risky release into a series of small, reversible ones.

## What is a canary deployment?

In a normal release, every user moves to the new version at the same moment. If there’s a bug, everyone meets it. In a canary release, two versions run side by side: the stable one and the canary. A router, load balancer, or application-level switch decides which requests go where.

You start with a small share, observe, and increase it in stages. At each stage you ask one question: is the canary at least as healthy as the stable version? The comparison matters. A raw error rate means little on its own, but an error rate that is higher on the canary than on the stable version tells you something specific.

### How is it different from blue-green and feature flags?

[Blue-green keeps two full environments](https://filipeeduardo.dev/blog/blue-green-deployment) and switches all traffic from one to the other in one move. A canary shifts traffic gradually. Feature flags control behavior inside the code, often per user, and are frequently used to implement a canary. They overlap, and many teams combine them: deploy the code dark, then expose it to a growing group.

## How do you choose which users or traffic go to the canary?

This is the decision that makes or breaks the approach. A canary that only sees traffic nobody cares about proves nothing. A canary that sees your biggest customer first is reckless.

A sequence that has worked for me:

1.  **Internal users and staff.** They find obvious problems and will forgive them.
2.  **A small random sample of traffic.** Random sampling gives a fair comparison with the stable version.
3.  **Low-risk segments.** For example, free accounts, one region, or one tenant who agreed to go early.
4.  **Wider groups,** in stages, ending with high-value customers last.

### Random or sticky assignment?

Make assignment sticky. If a user bounces between versions on each request, they may see inconsistent screens, lose sessions, or hit data written in a format the other version can’t read. Assign by user ID, account ID, or a cookie so the same person stays on the same version during the test.

### Be careful with money and data paths

Anything that charges cards, sends emails, or writes data that the old version must also read deserves extra caution. Having worked on the cart and checkout team at Total Wine & More, I treat the money path as the last place to get adventurous. Keep the canary small for longer on those paths, and check that both versions can live with the data the other one writes. A canary can’t undo a corrupted record, so data compatibility has to be solved before traffic moves.

## Which metrics tell you the canary is healthy?

Pick the metrics and the thresholds before you start. If you decide what counts as healthy while you’re staring at a dashboard under pressure, you’ll rationalize. Compare the canary against the stable version over the same time window.

Category

What to compare

Why it matters

Errors

HTTP 5xx rate, unhandled exceptions, failed jobs

The fastest signal that something broke

Latency

Median and slow-tail response times

Averages hide the slow requests users actually feel

Saturation

CPU, memory, database connections, queue depth

Shows problems that appear only as load grows

Business outcomes

Completed checkouts, sign-ups, successful submissions

Catches bugs that return a happy 200 but break the flow

User signals

Support tickets, client-side errors, retries

Reveals what server metrics miss

### Business metrics are the ones people skip

A release can pass every technical check and still quietly stop people from finishing a purchase or an application. Server dashboards won’t show that. If the product has a core action, track its success rate for both versions. In a migration, this is often the most honest health signal you have.

### Mind the sample size

With a very small canary, a handful of requests can swing a percentage widely. Don’t promote because the first ten minutes looked clean. Let the canary run long enough, and across enough varied traffic, that the comparison means something. How long depends on your volume, so treat any specific duration as something to tune for your system.

## When do you promote or roll back a canary?

I write the rules down as a short checklist before release. A simple version:

-   **Promote** when the canary stays within your agreed margin of the stable version on errors, latency, and the business metric for the full observation window.
-   **Hold** when results are ambiguous. Stay at the current percentage and gather more data rather than pushing forward.
-   **Roll back** when any hard limit is crossed, such as a clear rise in errors, a drop in the core business action, or any sign of data being written incorrectly.

Rollback should be boring. Ideally it means changing a routing weight or flipping a flag back, not redeploying under stress. Test that path before you need it. If rolling back takes a long manual procedure, your canary is protecting you less than you think.

Automated promotion and rollback are possible, and some tools do it based on metric thresholds. I’d still run the first few releases with a human watching, so the thresholds are tuned against reality before you trust them to act alone.

## How do canary releases reduce risk in a migration?

[Migrating a live platform is where canaries pay off most](https://filipeeduardo.dev/blog/migrating-a-live-platform-to-a-new-stack). When I led the migration of a platform that applies to job openings automatically for its users, the whole architecture moved from a legacy stack to a new one while people kept using it every day. A single cutover would have meant betting everything on one moment. Moving gradually meant each piece could be proven with real traffic first.

The pattern is straightforward:

1.  Put a switch in front of everything before writing new code, so you control which version handles each request. This is always my first step.
2.  Move one small piece of functionality to the new stack.
3.  Send a small share of traffic to it as a canary and compare.
4.  Widen the share, then retire the old piece when the new one has proven itself.

This is why I prefer [incremental migration over a big-bang rewrite](https://filipeeduardo.dev/blog/legacy-modernization). A rewrite usually costs more and takes longer, partly because the old system has to keep running in parallel anyway. A canary makes that parallel period useful: you learn from real behavior instead of waiting for a launch day to find out.

One limit to keep in mind: a canary tells you how the new version behaves for the traffic it receives. It won’t catch a rare edge case that none of that traffic triggered, so it complements tests and code review rather than replacing them. For older systems, see [how to change old code safely](https://filipeeduardo.dev/blog/legacy-code-refactoring).

## Do you need special infrastructure for canary deployments?

No. Specialized tooling makes it easier at scale, but the idea works with modest setups. Options, from simplest to most involved:

-   **Feature flags in the application.** Expose the new code path to a percentage or list of users. This needs no routing changes.
-   **Weighted routing at a load balancer or reverse proxy.** Many common ones support sending a percentage of requests to a second set of servers.
-   **Platform-level traffic splitting.** Container orchestration and cloud platforms often offer weighted rollouts, and some hosting platforms offer traffic splitting. Check your provider’s current documentation for what’s supported.
-   **Dedicated progressive delivery tools** that automate analysis and rollback.

What you do need, whatever the tooling, is observability. If you can’t see errors, latency, and a business metric split by version, you’re not running a canary, you’re just releasing slowly. Tagging logs and metrics with the version is the first thing I set up; the routing can be crude.

You also need to handle shared state. [Both versions will hit the same database](https://filipeeduardo.dev/blog/change-data-capture), so schema changes must be backward compatible while the two coexist. Add columns before using them, and remove old ones only after the old version is gone.

## Planning a canary for your own release?

If you’re weighing how to roll out a migration or a risky release, email me@filipeeduardo.dev with a short description of your setup. I’m glad to take a look and tell you what I’d check first.

## Frequently asked questions

### Is a canary deployment the same as a staged rollout?

They’re close, and people use the terms loosely. A staged rollout increases exposure in steps. A canary adds the explicit comparison between the new and stable versions, with criteria to promote or roll back at each step.

### What percentage of traffic should the canary get first?

There’s no universal number. Start small enough that a failure is tolerable, but large enough to produce meaningful data. Low-traffic products may need to run the canary longer rather than increase the share.

### Can small teams or early-stage startups use canaries?

Yes. A feature flag and a version-tagged dashboard are enough to begin. The discipline of defining healthy metrics and a rollback rule matters more than the tooling.

### Do canaries work for database changes?

Only partly. You can canary the application code, but the schema must work with both versions during the overlap. Risky data migrations need their own plan, such as backward-compatible steps and verification, beyond traffic splitting.

### What’s the biggest mistake teams make?

Not deciding the success criteria in advance, and watching only technical metrics. Define what healthy means, including a business metric, before the first user reaches the canary.
