Sebastian MercierChief Technology Officer · Essays & perspective
← All essays

Responsible AI3 min read

Reliability belongs in product prioritisation

Reliability work competes poorly when it is presented only as engineering housekeeping.

Its value becomes clearer when linked to customer experience, incident exposure and the organisation's ability to deliver change safely.

I want product and engineering leaders to share that view. Technical debt should be described through the consequences it creates, with a reasoned sequence for addressing it. That supports a business decision instead of a recurring negotiation over an undefined engineering allowance.

Reliability loses prioritisation debates for predictable reasons. Features have advocates, visible launch dates and customers waiting for them. Reliability work is most valuable when nothing happens, which makes its benefit hard to see and easy to defer. Incidents produce a burst of attention and a list of remedial actions, but the attention fades as the next product commitment arrives. Engineers learn to describe the work in technical terms that product leaders cannot easily weigh, and the work is postponed until a more serious incident forces it back onto the agenda.

I hold a regular reliability review jointly with product leaders. We agree service objectives in customer terms, describing the experience a customer should be able to rely on rather than abstract system measures. Each service then has an agreed tolerance for disruption. While a service stays within its tolerance, teams prioritise feature work freely. When the tolerance is consumed, reliability work takes precedence by prior agreement, without a fresh negotiation each time. That arrangement turns a recurring argument into a rule both sides accepted while they were calm.

Imagine a platform where releases require manual steps and occasionally cause brief disruption. Product leaders see occasional incidents that seem tolerable. Engineers see releases that are slow and risky, so they batch changes together, which makes each release riskier still. Described in terms of consequences, the picture changes: customers in an important segment experience disruption at renewal time, and the product team cannot respond quickly to competitive moves. Investing in automated, safer releases addresses both problems. Framed that way, it becomes a product decision that product leaders can weigh against other priorities.

Some leaders worry that reliability can absorb unlimited effort, and that engineers will pursue perfection at the expense of progress. The concern is fair. Not every service needs the same standard, and a back-office reporting tool can tolerate far more disruption than a payment flow. Reliability targets should follow customer need, not engineering pride. Setting them explicitly is what prevents gold-plating, because a service that already meets its agreed objective has no claim on further investment until the objective itself changes. The discipline protects feature work as much as it protects reliability.

This asks product leaders to co-own reliability outcomes rather than treating them as engineering's private concern. It asks engineering to explain technical debt through its effects on customers and delivery, and to propose a sensible sequence for addressing it. It also asks the wider leadership team to accept that some roadmap items will move when a service exceeds its tolerance, and to explain those moves honestly to the people who were expecting them. That honesty inside the organisation is good preparation for the conversations that matter most outside it.

Making reliability visible in the roadmap also changes the conversation with customers. A clear account of where the platform is strong, where it is fragile and what is being done about it builds more trust than an unbroken promise of availability. Engineering credibility is earned in exactly those conversations.