The AI Research Verification Checklist
Four passes that catch what a confident AI report hides. Run them before you act on any research.
Want help applying this to your business? Get in touch and tell me what you are working on.
Why the Most Confident Document You Own Needs Checking
An AI research report is the most confident-looking document you will ever be handed. Structure, tables, citations, a clear recommendation. Everything about the presentation says this has been checked.
I asked for one before a build decision and was fairly sold on its recommendation. Then I ran four verification passes on it. Each pass reversed something the previous pass had established. By the end, the report's core recommendation was gone, and so were several decisions I had made on top of it. The four passes cost me hours. The architecture they ruled out would have cost months.
This checklist is those four passes, in order. You do not need special tools. You need the discipline to run them before acting, not after.
- Check off each item as you complete it.
- The first two passes take minutes. Run those on everything.
- Run all four before anything you'll spend weeks or client money on.
Pass 1: Check Whether It Can Be Checked (~5 minutes)
Before reading the recommendation, try to open one citation. In my report, every single reference came through as a broken token. The document looked thoroughly sourced and was, as delivered, unverifiable.
This one observation reframes the whole document. A report with working sources is research. A report without them is a hypothesis, and you treat a hypothesis differently.
- ☐ Open 2-3 citations before reading the conclusion
- ☐ If links are broken or missing, relabel the report a hypothesis in your head
- ☐ Check whether the sources are independent or all tracing back to one origin. Fifteen of my sources agreed with each other. It was one idea repeated fifteen times.
Pass 2: Read What the Report Is Summarizing (~30 minutes)
Pull the actual sources the report leans on and read them yourself. My report characterized an influential approach, and when I read the original, it was substantially simpler than the report's description. The summary had layered its own assumptions on top and presented the result as the source's position.
This pass overturned the report's central recommendation. The engineering team behind the cited approach had explicitly rejected the architecture the report was recommending, calling it overkill for exactly my use case. That contradiction only shows up in the source, never in the summary of it.
- ☐ Identify the 2-3 sources the recommendation actually rests on
- ☐ Read them directly, not the report's description of them
- ☐ Note every place the original is simpler or different than the summary claimed
Pass 3: Gather Practitioner Accounts, and Ask for Deltas (~1 hour)
Find people actually doing the thing: videos, write-ups, demos. And when you feed them to an AI to digest, brief it on what you have already decided, then ask what changes about the plan. "Summarize this" produces summaries you still have to read. "Tell me what changes about my plan" produces decisions.
Know this pass's blind spot. My ten practitioner videos were all polished tutorials. Ten builds, zero failures shown. Useful for mechanics, useless for risk. That is what Pass 4 is for.
- ☐ Collect 5-10 practitioner accounts of the thing being done for real
- ☐ Brief the AI on your current plan before it reads them
- ☐ Ask for changes to the plan, not summaries
Pass 4: Ask for the Postmortems, Not the Tutorials (~1 hour)
This is the pass that matters most, and the one nobody runs. Same research tool, same topic, opposite brief: find the postmortems, the abandonment stories, the forum complaints, the people who tried it and quit. Explicitly exclude tutorials and product launches.
What came back for me was categorically different material, and it reversed most of what was left standing, including a scale threshold that was off by an order of magnitude. On any hyped topic, the published material is mostly enthusiasm, so a neutral prompt returns enthusiasm. You have to ask for the other side by name.
The single highest-value question I asked all session: "Make the strongest evidence-based case that I should NOT do this." Ask your AI to argue against the thing you want. The answer is worth more than another round of confirmation.
- ☐ Run one research pass that only hunts failures, postmortems and complaints
- ☐ Ask the AI to argue against your preferred conclusion, explicitly
- ☐ If a pass confirms everything you believed, redesign the pass. It wasn't built to find anything.
Where to Start
| Pass | Time | When to run it |
|---|---|---|
| 1: Open the citations | ~5 min | Every report, always |
| 2: Read primary sources | ~30 min | Anything you'll act on |
| 3: Practitioner deltas | ~1 hour | Before you build |
| 4: Hunt the failures | ~1 hour | Before you spend weeks or client money |
The minimum viable version, if you only do two things: open one citation, and run one adversarial pass asking for failures. Those two steps caught almost everything in my session.
About Andrew Voskov
Andrew Voskov is the founder of Cherry Pi AI. He runs this checklist on his own research before client recommendations, including the session it was built from, where it reversed a decision he was about to spend months on.
Want help applying this to your business? Get in touch and tell me what you are working on.
Found this useful? Follow along on LinkedIn — I post free systems and breakdowns every week.
Follow Andrew on LinkedIn