Resilience has a shopping-list problem. Timeouts, retries with exponential backoff, circuit breakers, bulkheads, backpressure, fallbacks, dead-letter queues — the patterns are well documented, the libraries are good, and adding them is easy enough that it feels like progress.
It’s entirely possible to add every one of them and end up less reliable than you started. Not because the patterns are wrong, but because each one is an answer to a specific failure, and applied to a failure it doesn’t fit it does something worse than nothing: it converts a fast, visible error into a slow, silent one.
So this post isn’t a tour of the patterns. It’s about the step that has to come first, and the reason most retry code I’ve reviewed is decoration.
Classify the failure, or the pattern is decoration
Nearly every resilience decision follows from one question: what kind of failure is this? There are three answers and they want opposite treatments.
Transient. The downstream is briefly unavailable — a 503, a dropped connection, a node recycling, a database failing over. The operation would succeed if you tried it again in a moment. This is the only class retries were invented for. Let it throw and let the retry policy replay it.
Business or permanent. The record doesn’t exist, the state transition isn’t legal, validation failed. Trying again cannot help, because nothing about the world will be different next time. Retrying these is worse than useless: it burns the retry budget that the transient failures needed, and it delays a definitive answer somebody is waiting for. These want an explicit decision — handled, logged, or dead-lettered on purpose.
Unexpected or data-integrity. Something is wrong in a way you didn’t model. Here retries are actively harmful: you don’t want a hundred replays of an operation that might be corrupting something. You want it out of the pipeline fast and loudly, with enough context to diagnose.
The reason I lead with this rather than with the patterns: on one platform I worked on, the single
most recurring cause of data loss over a period wasn’t a missing pattern. It was handlers with a
blanket catch (Exception) that treated all three classes identically. A transient upstream 5xx
and an illegal state transition took exactly the same path, which meant a momentary blip became a
permanent silent drop — the message was gone, the operation hadn’t happened, and nothing had
technically failed. Four separate incidents in a single week traced back to that one shape.
Adding more retries on top of a handler like that doesn’t fix it. It just moves the silent drop further into the future.
A timeout is a product decision wearing technical clothes
Timeouts are the most under-thought item on the list, because they look like configuration.
The right question isn’t “how long might this take?” — it’s “how long is it acceptable for a customer to sit and wait?” Those are completely different questions, and only one of them is answerable by an engineer looking at a latency chart.
A worked example I’m fond of. A third-party lookup on a customer-facing journey had a six-second client timeout. Not a default, and not a guess: in production that provider returned in under two seconds, so six was a deliberate ceiling — comfortably above normal, low enough that a hanging provider couldn’t hold a customer indefinitely. The number encoded a decision about the customer’s experience, not about the provider’s performance.
The interesting part came in the lower environment, where the same provider routinely took thirty seconds. So the six-second timeout fired constantly, filled the logs with cancellations, and looked exactly like a bug. It wasn’t. The right response was to leave it alone and know why, because the number was correct for the environment it was chosen for, and “fixing” it would have quietly removed a customer-experience guarantee from production to tidy up test noise.
Two things fall out of that. A timeout without a stated reason will eventually be widened by someone who thinks it’s a bug — so record what it protects. And a timeout is one of the clearest cases of a non-functional requirement being the architecture: “customers wait at most six seconds” shapes the system far more than most feature decisions do.
Partial success is a design choice — make it on purpose
Here’s a pattern that doesn’t appear on the shopping list but is doing resilience work in most systems: carrying on after one item fails.
A nightly job generating a batch file for many independent records wraps each record and, on a parse failure, records the problem and moves to the next. That’s deliberate. One malformed record must not prevent the file being produced for the hundreds of healthy ones, because the alternative punishes every other record for the sins of one.
I initially read that catch as a defect and proposed removing it, which was wrong, and the correction is the useful bit: keep the resilience, add the signal. A swallowed failure with no telemetry is indistinguishable from no failure at all — you have built something that degrades invisibly, which is the failure mode you least want in anything with a reconciliation obligation. What it needed was a structured warning per dropped record, a counter you can alert on and trend, a span you can query, and a dropped-count surfaced in the response so the file could be reconciled before anyone relied on it.
Which is the general rule, and probably the most useful sentence here: resilience without observability is just silence. Every pattern on the list — retries, breakers, fallbacks — works by hiding a failure from the caller. That’s the point. But if it also hides the failure from you, the system now has a category of problem that never surfaces until it surfaces as a very large one. This is the argument observability as a first-class concern makes, arriving from the other direction: you can’t operate a system whose recoveries are as invisible as its failures.
Where the boundary sits matters more than the pattern
One structural point worth more than any individual pattern. If a handler updates local state and also acknowledges a message, the order of those two things decides what happens in a disaster.
Roll back the local transaction and let the message be lost, and you have the worst outcome available: the operation didn’t happen and nothing will retry it. That isn’t a missing resilience pattern; it’s two consistency boundaries that were never reasoned about together. Either the message is safely settled before you attempt the external work, or the rollback compensates explicitly — but “it’ll be fine, we’ve got retries” is not a design.
This is the dual-write problem wearing different clothes, and it’s why retries and idempotency are the same conversation: a retry only helps if repeating the work is safe, which is the whole subject of exactly-once being a lie.
When each one is overkill
The other half of “earn their keep” — the cases where a pattern costs more than it returns:
- A circuit breaker on something you call once an hour. Breakers need traffic to make a statistical judgement. With low volume it will trip on a coincidence and stay open through a window where everything was fine, and you’ve built an outage generator.
- Retries on anything you haven’t made idempotent. An ambiguous failure — the work committed, the acknowledgement was lost — retried against a non-idempotent operation applies it twice. You have traded a visible error for a duplicate, which is a much worse bug to find.
- Backpressure on a queue that never fills. Real, but not your problem yet, and the machinery isn’t free to reason about.
- A fallback that returns something plausible. The most dangerous item on the list. A cached or default value returned during an outage means the customer gets an answer that looks fine and isn’t. Sometimes right, but it’s a business decision about whether a stale answer beats an honest error, and it should be made by someone who owns that trade-off rather than by whoever was writing the method.
The uncomfortable general form: every resilience pattern adds a failure mode of its own. Retries add load exactly when a system is struggling. Breakers add a state machine that can be wrong. Timeouts add the possibility of abandoning work that would have succeeded. You’re not eliminating failure, you’re choosing which failures you’d rather have — which is the same how-much-is-enough shape that security has, and the same it depends underneath.
None of which is an argument against the patterns. It’s an argument for remembering what you bought them with: partial failure is the thing you took on the day you went distributed, and this whole list is instalments on that tax.
The honest limits
- Resilience competes for the same budget as everything else. Every pattern is code to write, understand, test and operate, and it shows up on the observability bill too. The question isn’t “is this more resilient” but “is this the most reliability I can buy with this week”.
- The patterns interact, mostly badly. Retries behind a breaker, breakers behind a load balancer, timeouts nested inside timeouts — combinations misbehave in ways none of them do individually, and the interactions are what you’ll actually be paged for.
- You cannot tune what you cannot see. Choosing sensible retry counts, timeouts and thresholds requires knowing your real latency distribution and real failure rates. Without that, every number in your resilience configuration is someone’s guess, including mine.
The one-line version
Resilience isn’t a set of patterns you add, it’s a classification you do first: decide which failures are transient, which are final and which mean something is broken, and the patterns follow almost mechanically. Then make sure every recovery leaves a trace — because a system that heals silently is a system that is lying to you about how well it works.