The Checklist I Run Before I Trust Any AI Output
Last week's three checks, turned into the actual checklist I run against AI output. Steal it.
Last week I gave you the three checks I run against anything an AI hands me before I trust it. Reversible or not. Consequence versus checking effort. And the big one, did I verify it myself or just believe the tool’s report.
I promised I’d hand over the actual checklist I run those three against. Here it is. Steal it, keep it next to your desk, and run your next AI output through it.
First, the thirty-second version
Before you ship or send anything AI touched, ask three questions in order:
If this is wrong, can I undo it easily?
Does the cost of being wrong here justify how hard I’m checking?
Do I actually know this worked, or am I trusting a “done” message?
If all three come back clean, ship it. If any one of them snags, that’s exactly where you slow down. Most AI output passes all three in a few seconds. The checklist isn’t there to slow you down everywhere. It’s there to tell you the one place to slow down.
Now the full version, one check at a time.
Check 1: Reversibility
The question: if the AI got this wrong, what does it cost me to fix?
Run down this list:
Can I undo it in one click or one command? Editing a paragraph, re-running a draft. Green light, barely look.
Does this write to something permanent? A database, a billing total, a live post, a sent email. Slow down.
Could running it twice by accident do damage? If yes, it needs a guard before it needs your trust.
The rule I build into everything: a re-run or a script I fire twice can’t quietly cost me an afternoon. If an action isn’t safely repeatable, I make it safely repeatable before I let AI near it. Check-before-write. A dry run first.
Your ten-second version: is this edit-and-forget, or is this permanent? Permanent gets read.
Check 2: Consequence versus checking effort
The question: does the downside here justify the scrutiny, or am I wasting my attention?
There are two ways to lose here, and most people only guard against one.
Over-checking. You read every throwaway draft with the same intensity as production code, you burn out, and then, exhausted, you stop checking the things that actually mattered. Over-checking everything is how people end up checking nothing.
Under-checking. The first ten AI outputs were great, so you stopped looking. The eleventh quietly wasn’t, and it was the one that went to your biggest client.
The fix is to match effort to stakes, deliberately. A rough draft gets a glance. Client-facing copy gets a real read. Anything that touches money or someone’s data gets read like your name is on it, because it is.
And remember the thing a better model will never fix: context. A smarter AI still doesn’t know your biggest client hates being called a “partner,” or that one number in that report is the one your boss will circle. That’s not the model’s job. It’s yours.
Check 3: Did I verify it, or trust the report?
This is the one almost nobody runs, and it’s the one that saves you.
A tool telling you it succeeded and the thing actually succeeding are two separate facts.
I learned this watching automations report “done” while the actual thing hadn’t happened. So now, for anything that matters, I check the reality, not the log:
The tool says the post published. Did it actually appear on the profile?
The tool says the file sent. Is it in the sent folder, with the attachment?
The tool says the record saved. Is the row actually there when I query it?
The tool says “success.” Success at what, and can I see it with my own eyes?
The gap between “reported done” and “actually done” is small when models are bad, because they fail loudly. It gets more dangerous as models get good, because they fail quietly and confidently. The better your tools get, the more this check matters.
Your ten-second version: don’t read that it worked. Look at the thing.
How to actually use this
You’re not running a three-page audit on every prompt. That would be its own kind of failure.
Here’s the real rhythm. The reversibility and consequence checks happen in your head, in a second or two, and they tell you whether to bother. Ninety percent of the time the answer is “reversible, low stakes, move on.” The verify-it-yourself check is the one you actually stop and do, only on the small handful of things where being wrong is expensive and permanent.
That’s the whole system. Two fast filters that tell you where to look, and one real look where it counts.
Your ten minutes this week: take the last thing an AI did for you that actually mattered, something that sent or shipped. Don’t trust the confirmation. Go look at the real thing and confirm it happened the way you think it did. That habit, more than any model upgrade, is what separates people who build on AI from people who get surprised by it.
PS. I build these checks into systems so I don’t have to remember them: idempotency, and verify-then-report with a human in the loop. If you want that discipline wired into your own AI setup instead of living in your head, that’s what I do with people one on one: coaching.g8n.ai



