Software Evolution: When the Right Design Stops Being Right

Your Software Needs to Evolve

Have you ever worked on an application where every new feature seems to require another exception to how the system was designed? The code still runs, the tests pass, and it does what it was originally built to do. Yet making changes keeps getting harder.

I think one of the difficult parts of software engineering is accepting that software can be right when we write it and wrong a few years later without a single line changing. The requirements change, our understanding of the problem grows, and the assumptions we started with may no longer hold true.

When I think about this, I find myself comparing it to the Stone Age, Bronze Age, and Iron Age. It is a loose comparison, but each represents a period in which people developed and used tools with the knowledge and materials available to them. Those tools served a purpose, and over time we developed new ways to solve problems. We can appreciate what an earlier tool made possible while recognizing why we moved beyond it.

Software should have room to go through that same evolution. However, we sometimes treat changing an existing design as an admission that someone made a mistake.

A lot of projects begin with a straightforward problem. Maybe we write a small Go program that calls an API and does something with the response. The requirements are simple, and the code reflects that. At this point, there may be very little reason to build anything more complicated.

As people use the software, we learn more about what they need. Workflows become more involved, and the software takes on responsibilities we had not considered at the beginning. These changes can affect the basic assumptions behind the code, even if the original implementation still works exactly as intended.

Take something like provisioning a node. The first version might allocate a machine, configure networking, install the operating system, and mark it ready. We write the code around that sequence because getting the machine into service is the problem we need to solve.

Later, we need to manage more of the node’s lifecycle. Firmware needs to be updated. Hardware needs maintenance. A node may need to be drained, reconfigured, and returned to service. The software now needs to manage machines that already have workloads running on them, and the order in which we make changes matters.

We could revisit how we represent the node and its lifecycle, but parts of the provisioning code already do what we need. So we add a flag to skip the OS installation. Then another to preserve the existing network configuration. A firmware update needs a reboot, so we add a separate path for that, along with checks to avoid rebooting a node that is still running workloads.

Each addition serves a purpose, but we are gradually turning a process built to prepare a new machine into something responsible for managing that machine throughout its life. A node being ready is no longer the end of the process. We need to understand whether it is in service, undergoing maintenance, waiting for an update, or safe to return to use.

At that point, the original provisioning sequence no longer gives us enough information to manage the node safely. We need clear lifecycle states and rules for moving between them. Continuing to add flags leaves each workflow responsible for figuring out those rules on its own.

The original code solved the problem we had. The responsibility of the system grew, and now the design needs to grow with it.

That can be an uncomfortable conversation, especially on larger projects. Changing the underlying design may involve migrations, compatibility concerns, and coordination across several teams. People have spent time building and maintaining the existing system, and there is a reasonable concern about breaking something that customers already depend on.

There is also the roadmap. It can be difficult to justify revisiting an existing design when there are features waiting to be delivered. The next workaround may look smaller and easier to estimate, even when everyone understands that it adds to a problem we will eventually have to address.

So we decide to fit one more feature into the current design. We will come back to it later.

The trouble is that the software keeps growing while we wait. The next feature depends on the workaround, another team starts relying on its behavior, and what was supposed to be temporary becomes something we have to support. By the time we revisit the original problem, there is considerably more to untangle.

This is how we can end up adding technical debt under the label of an improvement. The feature itself is real, and it may solve an immediate customer need. But we should also be honest about whether we have improved the design or added another way around its limitations.

Eventually, that cost becomes part of everyday development. A small change requires updates in several places. A new engineer asks why a rule has so many exceptions, and the explanation involves years of decisions that no longer make much sense together. A fix in one workflow breaks another because they were built around different assumptions.

We keep paying for the old assumption with every new change.

What bothers me is the idea that we need to keep defending a decision because it was the right one when we made it. “This was the right design” and “we need to change this design” can both be true. The first version solved the problem we understood at the time. We should expect to learn something from building and using it.

I think it is important that teams can have this conversation without turning it into a judgment of the engineers who wrote the original code. Otherwise, we make it easier to propose another workaround than to explain why the design needs to change. We end up protecting a past decision at the expense of the work we need to do now.

Of course, every new requirement does not mean we need a rewrite. Sometimes extending the existing implementation is the right approach. There is value in working software, and changing it has a cost. But when we repeatedly run into the same limitation, we should question whether the original approach still makes sense.

This evolution can happen incrementally. We can introduce a new approach, move one workflow at a time, and remove the old behavior once it is no longer needed. It takes planning and collaboration, and it needs to be treated as part of the work rather than something we hope to get to between features.

To me, this is part of the lifecycle of software. We build it, learn from it, and reconsider parts of it as the problem changes. Some code will continue to serve us well. Other parts will need to be replaced or removed. We need to be comfortable making those decisions.

The longer we avoid that change, the more of the system we build around a decision that no longer fits. Eventually, the hardest part of building the next feature is accommodating everything we were afraid to change.

The original design earned its place by solving a problem. We need to be willing to question it when it becomes the problem.