Essay · AI · Judgment
Here is the conclusion first. Most people judge AI output on whether it reads well, and that test stopped working. Fluency costs nothing now. A wrong answer and a right answer arrive in the same voice, the same format, the same confident register, and you cannot tell them apart by feel. So you need tests that do not depend on how the thing sounds. I use six. They take about ninety seconds and they catch most of what goes wrong.
I built these while testing a free set of prompts I give away on this site for researching a stock, a fund, or your own portfolio. I ran every prompt against five different instruments to find where they broke, and they broke in ways I did not expect. The failures were not hallucinated facts, which is what everyone watches for. They were quieter than that. The model answered a question that did not apply, and formatted the answer so well that nothing signaled anything had gone wrong.
1. Does it lead with the answer?
Open the response. Read the first two lines. If the conclusion is not there, the answer failed, and you can stop reading.
This is not a style preference. It is a diagnostic. An answer that builds toward its conclusion is usually an answer that does not have one, because the writing is doing the work the thinking should have done. When a model knows what it thinks, it can say so in a sentence. When it does not, it stacks context until the space where a conclusion belongs gets filled by momentum.
I have run on this rule for four decades and it is Chapter One of my book. Serve the cake first. If someone wants the recipe, they will ask. The version I use with AI is stricter, because the model will always comply with the shape you request: I tell it to open with the verdict in under thirty words, then give me the three findings that make it true, then descend into detail in that same order. When the top of the answer is thin, the analysis underneath is usually thin too.
2. Is every number dated?
Undated numbers are unusable, and models produce them constantly.
While I was testing, I pulled the market capitalization of one mid-sized company from several sources on the same afternoon. The figures differed by about a billion dollars, because each provider had stamped a different as-of date. Nothing was wrong with any of them. But an answer that quotes one of those figures without saying when it was true has handed you a number you cannot check and cannot compare. Run the same question next week and you will think something changed when nothing did.
So I require an as-of date on every figure. It is a small ask, it costs the model nothing, and it converts an assertion into something verifiable.
3. Does it say what it could not verify?
Silence about sourcing is the tell. Good answers distinguish between what came from a filing and what came from somewhere softer. Weak answers present both in identical language, because there is no penalty inside the model for smoothing over the difference.
The instruction that fixes this is blunt. Mark anything you could not verify from a primary source. Then the gaps become visible, and you get to decide whether they matter. Without it you are reading a document where the solid parts and the guessed parts wear the same clothes.
4. Does it know what it is looking at?
This is the one that surprised me, and it is the failure I would watch for hardest.
I asked a research prompt to assess a widely held index fund. The prompt requested things like the management team’s capital allocation record and recent insider activity, because it was written for operating companies. An index fund has neither. There is no management team making allocation decisions. There are no insiders. The questions do not apply.
The model did not say that. It reinterpreted the question quietly, answered about the sponsoring firm and about the fund’s underlying holdings, and formatted the result exactly like a real company analysis. Nothing in the output flagged that the premise had been swapped. If you did not already know the difference between a fund and a company, you would have read that answer and believed you had learned something.
So the first thing I now require is identification. Tell me what this actually is before you tell me anything about it. When the premise is wrong, everything downstream is decoration.
5. Does it tell you when a rating came from a machine?
Small companies often carry no human analyst coverage at all. What they carry instead is an algorithmic rating, generated by matching the company statistically against peers that do have coverage. The provider is usually transparent about this if you look. The label sits right there on the page.
A model reading that page will report the fair value and the moat rating in the same voice it uses for genuine analyst work, because the words look the same. You get a number with a pedigree it does not have. I now require the distinction to be stated explicitly, and it changes how much weight the rest of the answer deserves.
6. Does it refuse?
The last test is the one that separates a useful assistant from a confident one. Ask it something it cannot know, and see what happens.
An assistant without web access should tell you it cannot check current data and stop. A question that does not apply to the instrument should get a short note saying so. When I asked for chart support levels on a mutual fund that prices once a day at close, the correct answer is that there is no intraday chart to read. There is nothing to analyze. Say so.
What you usually get instead is compliance, because these systems are built to be helpful and a refusal feels like failure. So I write the permission into the prompt directly. If you cannot meet a requirement, name the requirement you cannot meet and stop. A short honest answer is worth more than a complete one you had to guess at.
What this actually changes
Generating analysis used to be the expensive part. It is now close to free, and the cost has moved to the other end of the process. The scarce skill is no longer producing an argument. It is rejecting one that looks finished.
That skill is not new and it is not technical. Anyone who has sat through a board meeting knows it. The deck is beautiful, the logic is tight, and something is off, and the job is to find the load-bearing assumption before the vote. What has changed is the volume. You used to face a handful of polished arguments a quarter. Now you can generate ten before lunch, and every one of them will read as though someone competent wrote it.
Six tests, ninety seconds. The output that survives them is worth your time. The output that fails is worth exactly nothing, no matter how good it sounds, and the ability to tell the difference quickly is going to matter more every year.
Fluency used to be evidence of thinking. It is now evidence of nothing at all.
The prompts I tested are free on this site, in English and Spanish, with the instruction block that enforces all six tests built into every one of them. No email required. Open the toolkit, or read the argument behind it in An Outsider’s Playbook.