# Migrating a Live Platform to a New Stack Without Downtime

> How to migrate a live platform to a new stack with zero downtime: run both systems side by side, move one journey at a time and keep every step reversible.

- Author: Filipe Eduardo, Senior Software Engineer (https://filipeeduardo.dev/)
- Published: September 30, 2026
- Topic: Rescue & modernization · Tags: Architecture, Feature flags, Migration, Tech leadership
- Canonical URL: https://filipeeduardo.dev/blog/migrating-a-live-platform-to-a-new-stack

## Key takeaways

- Never have a day when only the new system exists: route every request through a switch that can send it to the old or the new stack.
- Migrate user journeys, not layers, so each step ships, is measured and can be rolled back on its own.
- Agree on the numbers that mean “safe” before the first real user moves.

At avanzzada I led the full migration of a platform that applies to job openings automatically on behalf of its users, moving its whole architecture from a legacy stack to a new one while people kept using it every day. Stopping the product for a weekend cutover was never an option: for its users, a missed day means missed applications.

This post is the approach that made it possible. It is not specific to that platform. The same pattern works for moving a monolith to services, a PHP app to Node or Go, or an old frontend to Next.js, as long as the product has to keep running while you work.

## Why migrate at all?

A rewrite is rarely the right answer, so the case has to be concrete before you start. The honest reasons are almost always one of these:

-   **Every change is slow.** Small features take weeks because nobody can predict what they will break.
-   **Releases are risky.** Deploys happen rarely, at night, with someone on standby.
-   **The platform limits the product.** The next thing the business needs (scale, new integrations, a mobile app) is hard or impossible on the current stack.
-   **Nobody wants to work on it.** Hiring and keeping engineers for an outdated stack gets harder every year.

“The code is ugly” is not on that list. If the only problem is aesthetics, refactor in place. A migration is only worth it when the old system is blocking the business, and you should be able to say how in one sentence to the people paying for it.

## How do you run two systems side by side?

The first decision is the one that makes everything else possible: never have a day when only the new system exists. Every request goes through a switch that decides, per user, which stack handles it. This is the [strangler fig pattern](https://martinfowler.com/bliki/StranglerFigApplication.html), with a [feature flag](https://martinfowler.com/articles/feature-toggles.html) as the switch.

```ts
// one switch, per user, reversible without a deploy
const useNew = await flags.isEnabled('new-apply-flow', { userId });
return useNew ? newApi.apply(job, user) : legacyApi.apply(job, user);
```

That switch gives you three things you cannot get any other way:

-   Internal users on the new stack first, then a small share of real users, then everyone.
-   The same kind of request handled by both systems, so results can be compared.
-   A rollback that is a flag change, not an emergency deploy.

*Figure: Every request passes through one switch; the share on the new stack grows week by week.*

The switch has a cost. For a while you run and pay for two systems, and some data has to be readable by both. Plan for that period explicitly: decide where the source of truth lives for each piece of data, and which system writes it, before any traffic moves.

## Why move one flow at a time?

Instead of migrating layers (first the database, then the backend, then the frontend), migrate user journeys: sign-up, then the core action, then billing, and so on. Each journey moves end to end, ships on its own, and has its own metrics and its own rollback.

Layer-by-layer migrations look tidy on a plan and deliver nothing until the very end. By the time the last layer lands, months of work meet real users at once, and if something is wrong you cannot tell which layer caused it.

Journey-by-journey migrations deliver value from the first flow, and they teach you. The first flow you move will expose problems in your plan: data you did not know existed, integrations that behave differently, edge cases nobody documented. Better to learn that on one flow than on all of them.

Choose the first flow carefully. It should be real enough to prove the new stack under production traffic, but not the one that earns the most money. A secondary journey with steady traffic is ideal.

> A migration is a product launch that nobody is supposed to notice.

### Measuring that nothing broke

For each flow, agree on the numbers before any user moves:

-   **Success rate** of the key action (for an application platform: applications submitted successfully).
-   **Latency** at the 95th percentile, compared between old and new for the same kind of request.
-   **Error rate** in the new path, with an alert that fires before users complain.

Then move traffic in steps (internal users, 5%, 25%, 50%, 100%) and only advance when the numbers hold for long enough to trust them. If they do not, flip the switch back, fix, and try again. Nobody outside the team needs to know.

## How long does a migration like this take?

It depends on the number of journeys, not the size of the codebase. A small flow can move in a few weeks once the switch and the monitoring exist; a whole platform usually takes months. The first flow is always the slowest, because it pays for the infrastructure every later flow reuses: the switch, the dashboards, the data sync.

Two things shorten the timeline more than anything else: freezing new features on the old stack (new work goes to the new stack only), and deleting old code as soon as a flow reaches 100%. Every week both versions of a flow stay alive, you pay to maintain both.

## What would I do again?

> Make every step reversible, and the migration stops being scary for the team and invisible to the users.
> 
> Key takeaway

If I had to reduce the approach to three rules:

1.  Put a switch in front of everything before writing new code.
2.  Migrate journeys, not layers.
3.  Agree on the numbers that mean “safe” before the first user moves.

The rest is discipline: moving in small steps, watching the numbers, and resisting the temptation to “just switch everyone over” when the first flow looks good.

> Planning a migration, or stuck in the middle of one? Email me at [me@filipeeduardo.dev](mailto:me@filipeeduardo.dev) with the current stack and what is blocking you. I’ll tell you where I would start.
> 
> Work with me

## Frequently asked questions

### How do you migrate a live system without downtime?

Run the old and new systems side by side behind a switch, usually a feature flag, and move traffic to the new stack gradually. Compare success rate, latency and errors for the same kind of request, and roll back by flipping the flag instead of deploying. Users never see a cutover, because there isn’t one.

### Should you migrate by layer or by feature?

By feature, or user journey. Migrating one journey end to end lets you ship, measure and roll back each step independently, and you learn from the first flow before moving the rest. Layer-by-layer migrations only deliver value at the very end, when every change meets real users at once.

### What should you measure during a migration?

The numbers users feel: the success rate of the key action, latency at the 95th percentile and the error rate, compared between the old and new stack for the same traffic. Agree on the thresholds that mean “safe” before moving anyone, and only advance to the next traffic step when they hold.

## Sources

- [Martin Fowler — Strangler Fig Application](https://martinfowler.com/bliki/StranglerFigApplication.html)
- [Martin Fowler — Feature Toggles (aka Feature Flags)](https://martinfowler.com/articles/feature-toggles.html)
