The Decisions That Carried the Outage
My AI went dark for two days. The mornings shipped anyway, on rules written three weeks earlier.
At 1:19 in the morning on Saturday, my primary AI assistant hit its usage ceiling and went dark for two days.
I found out later that morning, the way you find out about good news: nothing was wrong. The Saturday content batch was drafted, staged, and waiting for my review, exactly where it always is. Sunday ran the same way. A backup had taken over both mornings, followed the same rules, and left the same paper trail.
Nothing about that weekend was lucky. Every decision that carried it was made three weeks earlier, in calm conditions, when nothing was on fire. That gap, between when protection gets designed and when it gets tested, is what this edition is about.
The three decisions
A takeover contract, written down where any tool can read it. There’s a small file on disk for each day. It says who owns the day’s run, when they claimed it, and what state the work is in. The rule lives in the file, in plain text: if the owner goes quiet for forty-five minutes without delivering, or marks the day failed, the backup may take over. A rule agreed on in advance and enforced by whoever shows up, instead of a meeting or a 7 a.m. judgment call.
When the primary went dark, the backup didn’t have to decide anything. It read the file, saw the conditions were met, and started working. The decision had been made weeks before, by me, once.
A deliberately underpowered backup. The backup can draft everything and publish nothing. It works in a sandbox with no network access: it can’t publish or email anything, and it can’t touch a live system. A separate delivery step, with its own checks, moves finished work to where I review it.
I weakened the backup on purpose, and trust had nothing to do with it. Permissions are the honest form of governance. Anyone can write a policy document that says “the backup will be careful.” The sandbox makes carefulness structural. When it took over, the worst it could possibly do was already decided, and it was small.
A gate that doesn’t move. Nothing publishes under my name without my word. That was true when my usual assistant was working, it was true while the backup ran the mornings, and it will be true for whatever model I’m running next year. The gate outranks the tools on both sides of it. Swap everything else out and the one thing a reader can rely on is unchanged: a human decided this was worth their time.
What broke anyway, and why that was fine
Honesty requires the other half. The backup couldn’t do everything. The parts of my morning routine that need live reads of the outside world, engagement data, other people’s posts, simply didn’t happen for two days. Those are exactly the powers the backup doesn’t have.
That’s isn’t a hole in the design. That is the design. The blast radius of the outage was decided in advance: drafting continues, publishing waits for me, and anything requiring live access pauses cleanly instead of running on stale assumptions. I’d rather lose two days of engagement plans than have a backup improvise with yesterday’s picture of the world.
A colleague-in-craft put the principle better than I had: Charlotte Ledoux wrote this week that AI agent governance is a design decision, not a clean-up job, and that most teams do it in the wrong order. They deploy, then discover what the agent shouldn’t have been able to do. The outage was my proof of the right order. By the time a system is being tested for real, the only decisions that count are the ones you already made.
One detail I keep coming back to
On Monday morning, with the primary back online, the backup still claimed the day first. It woke at 7:37. The primary showed up at 8:00 and found the work already staged, with a note in the contract file saying so. No collision, no duplicate, no confusion, because the same file that handles disasters also handles ordinary mornings.
Systems that only exist for emergencies rot before the emergency arrives. The takeover contract works because it isn’t emergency equipment. It runs every single day, which means it was rehearsed hundreds of times before it mattered.
The ten-minute version for your business
Pick one process that would hurt if you were unreachable for two days. Invoicing, client replies, publishing, payroll approval, whatever stings to imagine dropped.
Write three lines about it, somewhere your backup person or tool can actually see:
Who owns this, and how would anyone know they’ve gone quiet?
Who takes over, and after how long?
What is the taker-over explicitly NOT allowed to do without you?
That third line is the one people skip, and it’s the one that matters most. A backup with unlimited permissions isn’t a safety net. It’s a second way to have an accident.
The full system can come later. Something better already happened: the next outage now has rules waiting for it instead of improvisation.
If you’d like systems like this running your business, calm on their worst weekend, that’s what we build together, step by step, in coaching. The strategy call is thirty minutes, and you’ll leave with at least one takeover rule worth writing down. Book whenever suits you.



