Legacy Code Refactoring: How to Change Old Code Safely
How to refactor legacy code safely: characterization tests, picking targets by risk, small reversible steps, and when to switch to an incremental migration.

On this page · 6 sections
Key takeaways
- Refactoring means changing structure, not behavior. If behavior changes, it is a feature or a bug fix, and it should be a separate commit.
- Write characterization tests first. They record what the code does today, even when that behavior looks wrong.
- Start where change is frequent and risky, not where the code is ugliest.
- Work in small steps that can be merged and reverted on their own. Long-lived refactoring branches are where projects die.
- If the stack itself is the problem (unsupported runtime, no hiring pool, no way to test), stop refactoring in place and move to an incremental migration with a switch in front of everything.
Legacy code refactoring is safe when you change the structure of code without changing what it does, and you can prove that with tests before you touch anything. The order is: lock in current behavior, pick the parts that hurt most, change them in small reversible steps, and keep shipping features while you do it.
It goes wrong when teams tidy code they do not understand, with no safety net, in one giant branch.
What makes legacy code hard to refactor?
Legacy code is hard to change because nobody can predict what a change will break; ugliness is secondary. In the codebases I have taken over, the same things keep showing up:
- No tests, or tests that do not test anything. You cannot tell whether a change preserved behavior.
- Hidden coupling. A function that looks self-contained reads a global, writes to a shared table, or depends on call order.
- Behavior nobody documented. Odd-looking conditions often encode a real customer case or a past incident. Deleting them as dead code is a classic mistake.
- Missing context. The original developer or agency is gone, and the commit history says things like “fix” and “update”.
- Money paths mixed with everything else. Billing, checkout, or permissions logic is tangled into UI or data code, so a small edit has a large blast radius.
The first goal is not to clean anything but to make change cheap to verify: if you cannot run the project locally, build it reproducibly, and see a test result in minutes, fix that first. When I take over an inherited codebase, I follow the same order every time: access first, then the build, then the money paths, then the risks. Refactoring starts once the first two are solid.
How do you add tests before touching anything?
You add characterization tests: tests that capture what the code actually does right now, not what it should do. Michael Feathers popularized this approach in Working Effectively with Legacy Code.
Start at the outside, not the inside
Unit tests on tangled internals are slow to write and brittle. Begin at a boundary you can observe:
- An HTTP endpoint: send a request, record the response.
- A database state: run an operation, assert on the rows that changed.
- A rendered page or an end-to-end flow: sign up, check out, apply, pay.
These tests are coarse, but they catch changes users can see. I always cover the money paths first: checkout, billing, signup, anything that moves funds or locks people out.
Record, then assert
The process is simple:
- Call the code with realistic inputs.
- Write down the output, even if it looks wrong.
- Turn that output into an assertion.
- Repeat for the edge cases you can find in logs, support tickets, and old bug reports.
If you find a real bug while doing this, note it, keep the test asserting current behavior, and fix it later in its own commit. Mixing bug fixes into a refactor makes it impossible to tell what changed.
Use production as a source of test data
Sanitized samples of production requests show the cases the code really handles. Strip personal data before it reaches a fixture.
Break dependencies only as much as needed
When code is impossible to test because it calls a payment API or the system clock, introduce a thin seam: pass the dependency in, or wrap it in a small function you can replace in tests. Keep these seam changes tiny and mechanical, since they are the one kind of edit you make before you have tests.
Where should you start refactoring first?
Start where the code changes often and mistakes are expensive. Ugly code nobody touches can stay ugly.
When I rank candidates, these are the signals I look at, in roughly this order:
- Change frequency. Version control history shows which files are edited every sprint; messy structure there taxes you repeatedly.
- Bug clusters. Support tickets and incident notes point to modules that keep breaking.
- Risk. Money paths and authentication, where failures cost money or trust.
- Roadmap blockers. If the next three roadmap items touch the same module, start there.
My rule of thumb: if a module is both frequently changed and frightening to change, it goes first.
What not to start with
- Renaming and reformatting the whole codebase. It creates noisy diffs and merge conflicts and fixes nothing.
- The oldest code, simply because it is oldest.
- Code you are planning to delete soon. Do not polish what is about to go away.
Refactoring should be driven by evidence about cost and risk, not by taste.
How do you refactor in small steps while shipping features?
You make each refactoring step small enough to merge to the main branch on its own, and you interleave those steps with feature work instead of pausing the roadmap. In my experience, teams that stop everything for a “refactoring quarter” often lose stakeholder trust and end up with a half-finished branch.
Use the “leave it better” rule at the point of work
When a feature needs you to touch a messy module, spend a bounded amount of effort making that area easier to change first, then build the feature. Kent Beck’s phrasing is “make the change easy, then make the easy change.” The feature justifies the refactor, and the business sees both.
Keep steps reversible
Good small steps look like this:
- Extract a function, run the tests, merge.
- Move a function to a better module, run the tests, merge.
- Introduce a new implementation behind a flag, route a little traffic to it, compare results, then switch over and delete the old one.
That last pattern is the safest way I know to change risky logic: if the new path misbehaves, you can usually flip back without a deploy.
Separate commits by intent
Keep refactor commits and behavior-change commits apart. Reviewers can then read a refactor commit knowing nothing should change, and a failing test immediately tells you something did.
What this looked like on my team
At Inovaula I built v2 of the web platform and then led the team, and we reached the steadiest delivery flow the team had had. Looking back, I credit that less to any big cleanup than to the habit described above: small changes, reviewed carefully, merged often, with structure improving as part of normal work.
When should refactoring give way to a migration?
Refactoring improves code inside a stack, but cannot fix a stack that is itself the obstacle. Consider a migration when:
- The runtime, framework, or key libraries no longer receive security updates.
- You cannot hire, or onboarding takes far too long, because of the technology.
- The architecture prevents testing at all, for example logic living in the database or in templates with no seams.
- Each refactor needs so much scaffolding that you are effectively rebuilding anyway.
Even then, I prefer incremental migration over a big-bang rewrite. The rule I follow: put a switch in front of everything before writing new code. A routing layer, proxy, or feature flags lets you send one slice of traffic to the new system while the old one keeps serving everyone else. A rewrite usually costs more and takes longer, because the old system has to keep running in parallel the whole time.
At avanzzada I led the full migration of a platform that applies to job openings automatically for its users. The whole architecture moved from a legacy stack to a new one while people kept using it every day. What made it workable were the habits in this guide: knowing what the current system does, putting the switch in first, and moving one piece at a time.
Want a second opinion on your codebase?
If you are staring at an old codebase and unsure whether to refactor or migrate, email me@filipeeduardo.dev with a short description or a link, and I’ll tell you what I would check first.
Frequently asked questions
What is the difference between refactoring and rewriting?
Refactoring changes structure in small steps while behavior stays the same and the system keeps running. A rewrite replaces the code wholesale, often in parallel, and typically takes longer and carries more risk because you have to rediscover every behavior the old code handled.
Can I refactor legacy code without tests?
Only for the smallest mechanical changes, such as renaming through a tool you trust or adding a seam. For anything else, write characterization tests first, starting with the highest-risk paths. Refactoring blind means you will learn about breakage from your users.
How much time should a team spend on refactoring?
There is no universal number. Treat it as part of feature work, plus targeted work on modules that cause repeated incidents, and avoid separate refactoring projects with no visible business outcome.
Should I use AI tools to refactor legacy code?
They can help with mechanical edits and with explaining unfamiliar code, but they do not know your hidden business rules. Treat their output like a junior’s pull request: run it against your characterization tests and review every diff before merging.
How do I know the refactor did not break anything?
Tests that passed before still pass after, plus production monitoring. For risky logic, compare old and new paths behind a flag before cutting over. No process eliminates risk, but this keeps it small and visible.

