Proof It Worked
Six checks that tell you whether an AI tool did the job, or only said it did.
Want help applying this to your business? Get in touch and tell me what you are relying on.
Nothing In Here Crashed. That Was The Problem.
Every failure behind this checklist came with a green tick on it. No errors, no warnings, no red text anywhere. A tool told me the work was finished and the work was not finished, and the only way I ever found out was by going and looking at the thing itself.
That is the shape of the problem. A broken job and a finished job look identical from the outside, because the part reporting on the work is usually the same part that did it. Nobody grades their own homework honestly, and software does not even know it is doing it.
These six checks take a few minutes each. None of them require you to understand how the tool works. They only require you to stop reading the report and start reading the result.
1. Ask What It Checked, Never Whether It Is Done
"Is that everything?" has exactly one answer available to it, and you will get that answer every time regardless of the truth.
An AI told me a job was finished five separate times over two days. Four of those were honest. Each was true when it was said, and then a deeper look turned up something genuinely new. The fifth was a flat reversal: the same list of thirty-five open items, two opposite recommendations about an hour apart, delivered with the same confidence both times.
The question that works is different in shape. What did you check, and what have you not checked yet? That one cannot be answered enthusiastically, which is precisely why it is useful.
Ask: what did you look at, and what is still unexamined? Then decide for yourself whether the gap matters.
2. Read The Result, Not The Report
A report is written by the thing being reported on. That is a conflict of interest in every other part of life and we accept it from software without blinking.
I once ran five small jobs in a single session. Every one printed a success message. One claimed twenty-four changes and had made thirteen. Another looked at twenty-five files, got eighteen of them wrong, and called them fine. A third destroyed work that was already correct and printed a cheerful summary on the way out.
Not one of them failed loudly. Zero errors, five times. I had been reading that as good news, which is roughly like deciding your smoke alarm works because the house has never burned down.
Open the actual file, page or spreadsheet and count something. If the count and the report disagree, believe the count.
3. Ask What Your Score Cannot See
Any number that grades quality is measuring something narrower than quality. Find out what, before you trust a green result.
Twenty-five pages I had written all passed an automated gate and scored as easy reading on the standard formula. Then the client read one and asked whether it was actually plain English. The formula counts syllables and sentence length. It would happily approve a sentence like "cure a default on your principal residence", because every word in it is short.
A second check built for what the formula cannot see found a hundred and seventy-three problems. After fixing every one, the famous score moved by one tenth of a grade. It had been blind to the whole thing while printing a reassuring number.
Write down what your score actually measures. Then name one way the work could be bad without the score noticing.
4. Test Every Rule In Both Directions
A safety rule has two jobs. Catch the bad thing, and leave the good thing alone. Almost everybody tests the first and almost nobody tests the second.
I banned a particular word from a client's website because their regulator forbids promising outcomes. Sensible. Then the rule failed one page thirteen times over that single word, and I made two wrong fixes before working out my own rule was the fault. The word appears in the disclaimer their regulator requires them to display. My safety check was blocking the thing the law told them to say.
Now every rule gets two planted violations that must be caught and two legitimate uses that must be allowed through. It takes ten minutes and it has saved me twice.
Feed your rule two things it should block and two it should permit. A rule that has only ever been tested one way is half tested.
5. Look At The Whole Answer Before Acting On Part Of It
An incomplete report is harder to catch than a wrong one. A wrong number eventually contradicts something. A missing number never contradicts anything at all.
I spent an afternoon planning to fix an expensive-looking job. It sent the same large block of instructions on every single run, which looked like obvious waste. Before redesigning it, I printed one extra piece of information out of the response, mostly so the plan would have a figure in it.
Almost all of it was already free and had been the whole time. The report had never been wrong. It had simply stopped one field short of useful, every run, for months, and nobody had thought to look past it.
Before optimising anything, look at the complete response once. The field nobody prints is often the answer.
6. Ask What This Check Would Show If The Thing Were Broken
This is the one that makes the other five work, and it is the only question here that applies to absolutely everything.
I once confirmed a website update had gone live by checking whether the page loaded. It loaded. The update had failed completely, and the old page was loading exactly as it always had. My check could not tell the difference between success and failure, so it was never a check at all. It was a habit that felt like one.
Run the question over anything you currently rely on. If the answer is that your check would look identical either way, you do not have a check. You have a ritual.
For each check you trust: what would this show me if the work had gone wrong? If the answer is the same as when it goes right, replace it.
The Six, In Order
| Check | The question | What it catches |
|---|---|---|
| 1 | What did you check, and what have you not checked? | Confident claims of being finished |
| 2 | Does the actual file agree with the summary? | A report written by the thing it reports on |
| 3 | What does my score not measure? | A green number on a bad result |
| 4 | Does my rule allow the things it should? | Safety rules that block legitimate work |
| 5 | What is in the rest of the response? | Decisions made on partial information |
| 6 | What would this check show if it were broken? | Habits that feel like checks and are not |
None of these require technical knowledge. Every one of them is a question you can ask out loud, and the awkward pause that follows is usually the answer.
About Andrew Voskov
Andrew Voskov is the founder of Cherry Pi AI. He has spent twenty years building online businesses and now helps business owners work out which AI tools earn their place and which ones only demo well. Every failure in this checklist is one of his own, on real work, caught late enough to be embarrassing and early enough to fix.
Relying on something you cannot fully check? Get in touch and tell me what it is.
Found this useful? Follow along on LinkedIn — I post free systems and breakdowns every week.
Follow Andrew on LinkedIn