Pillar 03 · Residency

Most resilience strategies are an act of faith with a storage invoice stapled to the back. Let's do the arithmetic instead, since apparently someone has to.

The default move is triple replication. Keep three full copies of everything. It feels safe, because three is a bigger number than one, and that is roughly where the analysis stops in most rooms.

So let's not stop there.

Three copies means you store 300% of your data to protect it. Two hundred percent of that is pure overhead — bytes you pay to store, power, cool, and replicate, that exist only to be redundant. In exchange, you survive losing any two copies. Triple the cost, two-failure tolerance. Write it down.

Now Reed-Solomon erasure coding. RS(5,2): split each object into five data shards, compute two parity shards, seven pieces total, and any five of the seven rebuild the original exactly. Storage overhead: 40%. Fault tolerance: the loss of any two of the seven shards.

Put them side by side and the comparison is almost rude. Identical two-failure tolerance. One-fifth the storage overhead. The erasure-coded version isn't a refinement of replication — it's the same guarantee for roughly a fifth of the bill. You were quoted five times the price for a tie.

And it gets worse for the replica, because so far I've been generous to it. Those three copies almost always live inside one provider — three disks, maybe three zones, one control plane, one company, one bad afternoon. So your 200% premium buys robust protection against a disk dying and approximately nothing against the provider itself going down, which, per every outage post-mortem ever written, is the failure that actually takes you offline.

Spread the seven shards across independent providers and your failure events stop being correlated. A whole vendor can evaporate — outage, billing dispute, regional event, doesn't matter — and you've lost one or two shards out of seven and rebuilt from parity before anyone noticed. The replica strategy can't say that at any price, because all three copies were standing in the same blast radius the entire time. Redundancy that fails together isn't redundancy. It's expensive synchronised regret.

So tally it. Triple replication costs five times more and defends against less. That isn't a trade-off. It's the wrong answer, sold confidently.

This is the slide nobody builds, because it ends the meeting. Resilience isn't a vibe and it isn't a vendor logo. It's a calculation, and the popular answer is off by a factor of five in the wrong direction.

Run the numbers on your own storage tonight. Then ask, out loud, what exactly you've been paying triple for.