A service starts throwing BadImageFormatException on startup. Sometimes an ArgumentException
about a metadata token instead — different landing spot, same root cause. It’s intermittent, it
only happens around deploys, and the assembly it’s complaining about is one you didn’t change.
The instinct is to go and read the diff, and the diff is innocent. This post is about what actually turned out to be happening, but mostly it’s about the three safety nets that all failed to catch it, because each one failed for a reason that generalises well beyond this bug.
The correlation that named the mechanism
The useful diagnostic move came early and it was a boring one: work out what the recurrence actually tracks.
It didn’t track the version. The same build would be fine, then not fine, then fine. What it tracked was deploy frequency — the more often we shipped, the more often it happened. That alone rules out a code fault and points at the act of deploying.
The second clue was instance identity. The apps were configured single-instance, yet over
twenty-four hours the telemetry showed around fifteen distinct instance IDs. That isn’t
scale-out — it’s one worker recycling fifteen times. And on Linux App Service, a Zip Deploy
extracts the package into /home/site/wwwroot, storage shared by every instance, and each instance
then syncs those files into its own local container. Every recycle re-runs that sync. Fifteen
recycles, fifteen fresh opportunities for a file to land badly.
So the mechanism was legible without ever reading a stack trace properly: the corruption isn’t in the package, it’s in the repeated materialisation of the package. Which is why the diff was always innocent.
That’s a habit in its own right. When something intermittent won’t line up with your code, stop asking “what changed in the release” and start asking “what does the frequency correlate with”. The answer often names the mechanism before you understand it.
Slots don’t fix this, and the reason matters
The first safety net was deployment slots, and everyone’s intuition — mine included, repeatedly — is that they should help. Deploy to staging, warm it up, swap. Surely a corrupt instance never reaches production.
It doesn’t work, and the reason applies to every slot-based deployment: a swap is a routing change over already-warmed instances plus a warm-up request. It is not an integrity check.
Specifically, three things go wrong in sequence. The package extracts into the shared filesystem belonging to the staging slot and synchronises to the instance before the swap — so if corruption happens during that sync, the corrupt instance is precisely what gets promoted. The swap itself then triggers a restart, which re-runs the shared-to-local sync, which is another chance to corrupt something after you’ve finished checking. And the warm-up that’s supposed to catch a broken instance is a single HTTP request, which brings us to the second safety net.
The general form: slots protect you against a bad deploy of good code. They do nothing about a bad materialisation of a good package, because nothing in the swap ever verifies that what landed on disk is what you shipped.
The probe was blind by design
The warm-up ping hit the liveness endpoint. That endpoint had been deliberately written to be cheap — no database, no ORM, no dependency checks — because that is exactly what a liveness probe is supposed to be. It answers “is this process up”, quickly, without side effects.
A corrupt instance passes it easily. The assemblies that fail to load are the ones behind the data layer, and liveness never goes near them. The failure only appears on the readiness endpoint, which walks the ORM’s compiled model and blows up.
So the probe gating promotion was structurally incapable of detecting the fault, and it was incapable of it because it had been designed well. That’s the uncomfortable part. This wasn’t a badly written health check; it was a correctly written liveness check being used as a qualification gate, which is a different job. The lesson isn’t “make your liveness check deeper” — it’s that liveness and readiness answer different questions, and if you use one as a proxy for the other you will get a confident answer to a question you didn’t ask.
Surfacing a fault is not fixing it
This is the one I got wrong, and it’s the most transferable thing here.
We added a deeper readiness check that forces the compiled model to load, so a corrupt instance fails readiness loudly instead of quietly serving errors. Good. I then assumed — and it had already made its way into a triage runbook as “self-heals, no action needed” — that the platform would see the readiness failure and replace the instance.
It doesn’t. Checking the configuration settled it: always-on was enabled so there was no idle recycle, autoheal was disabled, and App Service’s Health check feature — which replaces an instance that stays unhealthy for an hour — was pointed at the shallow liveness path, which a corrupt instance passes. Nothing in that arrangement evicts anything.
Which means a single corrupt worker can sit there failing readiness indefinitely, until a deploy or a manual restart. And on a single-instance app, one corrupt worker is not a degraded service, it’s all of it.
Two things to take from that. A health check surfaces a fault; it does not remediate one. Whether anything acts on it depends entirely on what you wired it into — and the platform acts on the probe you configured, not the one you find most informative. And more generally: “it recovered on its own” is a claim, not an observation. It needs the same evidence as any other diagnosis, which is the whole argument of proving it was Azure’s fault, not ours turned back on yourself.
I still can’t tell you what cleared it on the day I looked hardest. Probably a recycle from the next deploy. Possibly a platform event. Not idle timeout, not autoheal, not the probe — those are ruled out. Saying “it self-healed” would be exactly the mistake this post is about.
Remove the window, don’t detect it faster
If the corruption happens during extraction and synchronisation, the durable fix isn’t a better
check — it’s to stop extracting. Run From Package (WEBSITE_RUN_FROM_PACKAGE=1) mounts the
deployment zip read-only as the site root and serves directly from it: no extraction, no
per-recycle sync, no window. The artefact becomes immutable, which is what you
wanted from a deployment artefact in the first place.
There’s a sub-lesson attached that I liked less at the time. This had been tried before and
abandoned, because it caused a ten-minute outage — which was true, and which had killed the idea for
about six weeks. Except the outage wasn’t caused by the setting. It was caused by the deployment task: the Azure
DevOps AzureWebApp@1 task’s runFromPackage method is a Functions-runtime mount, unsupported for
Linux web apps. The same setting deployed with a plain az webapp deploy --type zip is the
combination Microsoft documents.
A remedy that fails for an incidental reason takes a correct hypothesis down with it, and it stays down until somebody re-reads why it failed rather than that it failed. Worth checking your own list of things-we-tried-that-didn’t-work for entries whose failure was never actually diagnosed.
The honest limits
- Single-instance is doing a lot of the damage. Running more than one instance wouldn’t prevent corruption, but it would stop one bad worker being a hundred percent of capacity. That’s a cost conversation, not a bug.
- The deep readiness check is still worth having — it just buys observability, not recovery. Knowing which of those you’ve bought is the entire point; keeping it while believing it fixes something is worse than not having it.
- Immutable artefacts trade one failure class for a different operating model. No in-place patching, no editing a file on the host to get out of trouble at 2am. Mostly a good trade, and still a trade.
- This isn’t a general indictment of the platform. It’s a specific interaction between shared storage, per-recycle synchronisation and single-instance hosting. Different topology, different exposure — which is why the resilience question is always “what failure am I actually treating” before it’s “which pattern do I add”.
The short version
A platform that rebuilds your application directory on every recycle will eventually rebuild it wrong, and none of the obvious protections notice: slots promote the corrupt instance, the warm-up probe is too shallow to see it, and the deep probe that does see it isn’t wired to anything that acts. The fix is to stop materialising the artefact at all — but the lesson is that a check which tells you something is broken and a mechanism that does something about it are two different pieces of engineering, and it is remarkably easy to buy the first while believing you own the second.