Blog / Leadership in Complexity / How Will You Know When It’s Wrong?


How Will You Know When It’s Wrong?


Two identical sealed envelopes lying side by side on a pale stone surface, each closed with the same unmarked navy wax seal, indistinguishable from one another

Every AI tool you’ve evaluated showed you a version of the same demonstration. Someone describes a task, and a short while later a finished result appears. It’s clean, complete, and genuinely convincing. Everyone in the room is impressed.

What the demonstration can’t show you is the part that matters most to anyone responsible for what their organization produces. How would you know if the resulting output from an AI tool was wrong?

That question sits at the center of AI-assisted work, and it has an uncomfortable answer. As of today, there is no generally accepted, reliable way to know. The problem isn’t just hard. It is, in actuality, unsolved. And it happens to be the one thing the people selling you these tools have the least reason to bring up.

The work that reports itself as finished

Start with the behavior that surprises leaders the most once they see it clearly.

When you give an AI system a task, it does the work and then tells you it’s done. Most of the time, it’s right. The difficulty is in the times it isn’t. When the work is wrong, the system reports that in exactly the same confident voice it uses when the work is right. There is no change in tone, no hedge, no flag of any kind. The word “finished” means the same thing whether it happens to be true or not.

The engineer and writer John Crider, who has been more candid about this than most people in his field, calls these phantom completions. They announce themselves as done while quietly not being done. It’s a good name, because it points straight at the specific danger. The absence is invisible. Nothing appears to be missing.

Take that out of the world of software, and it becomes familiar to anyone who has ever managed people. Imagine an employee who finishes every assignment you give them and reports each one complete with total confidence, and who is right, say, nine times out of ten. The tenth time, the report sounds exactly like the other nine. You are holding a piece of finished work, and the confidence attached to it tells you nothing about which of the two kinds you have (right or wrong). The signal you would normally lean on has gone flat.

That is the everyday experience of working with these tools at scale. It isn’t catastrophe. It’s a steady stream of finished-looking output, a knowable fraction of which is wrong in ways no one has noticed yet.

Why the usual check doesn’t work here

In normal operations, the check on this is straightforward. You get a second set of eyes. Someone other than the person who did the work looks it over before it leaves the building. For most of business history, the slowness of the work supplied that check for free. When something took two weeks to produce, there was time, and usually a second person, sitting between “done” and “delivered.”

AI removes both of those at once. The work itself gets dramatically faster, and the second set of eyes is, more and more, the same system that produced the work in the first place. These tools can review their own output, and they are often asked to. But a system checking its own work tends to look for the things it already thought of, and to miss precisely what it missed the first time. It is the same reason a writer proofreading their own draft reads what they meant to put on the page rather than what is actually there. The blind spot is shared, because there is only one set of eyes, wearing two hats.

This is the part that is easy to walk straight past, so it is worth saying plainly. Verification itself has not become a harder discipline. It requires what it always required: someone independent of the work, plus enough time for problems to surface before anything ships. That was the pre-AI recipe for results you could trust, and it worked for one reason above all, which is that the reviewer was not the author. The AI era preserves that recipe only when you insist on it. Left to the path of least resistance, and especially as the work speeds up, people allow the same tool that produced a result to be the one that vouches for it. At that moment the recipe loses its one essential ingredient. Independence is gone.

No warning when it’s out of its depth

There is one more property that makes this genuinely difficult, and it is the one most people underestimate.

A human expert alerts others when they reach the edge of their competence. They slow down, they hedge, they say some version of “this part is outside what I know well.” That signal is doing quiet but essential work. It tells you where to look harder. AI systems do not have such an alert. Their confidence at the peak of their ability and at the very edge of a cliff is identical. The output that is brilliant and the output that is confidently, even dangerously, wrong both arrive in the same steady, articulate voice, with no warning band in between.

You can see what this costs in the one area that has been measured carefully. Studies of AI-generated software find that somewhere between 40 and 62 percent of it contains a security weakness. Put plainly, roughly half the time there is a hidden flaw of the kind an attacker could use. Other research finds AI-written code carrying close to three times the defects of the human-written equivalent. The point for a leader is not the exact figure. It is the reason the figure is so high. It is high because confidently-wrong looks exactly like confidently-right, and nothing in the output itself tells you which one you are reading.

Why your vendors won’t say this

None of this tends to come up in a sales conversation, and the reason is not simple dishonesty. It is selection.

A demonstration is built out of the cases that work. The incentive runs in one direction only, toward the impressive result and away from the one-in-ten that fails in the same confident voice. “This is unsolved” is not a sentence anyone has ever put on a sales slide. So the people making the most consequential decisions about these tools, such as how far to trust them, how much to let them run without a person in the path, and what to put them in charge of, are the least likely to ever hear the single fact that should weigh on those decisions most.

That is the gap. It is not a defect hiding inside someone’s product. It is a piece of honesty that the structure of the sale quietly filters out.

Where this leaves you

The answer is not to hold AI at arm’s length. The economics really have changed, and the organizations that refuse to use the tools altogether will fall behind the ones that learn to use them well. This is not an argument for caution as a posture.

It is an argument for treating verification as something you build, rather than something you assume you already have. The instinct, reinforced by every demonstration you will ever sit through, is to treat “is it correct?” as a property the tool either has or does not have, and then to go shopping for the tool that has more of it. That instinct is the trap. At this stage of the technology, correctness is not a property you can buy. It is a system you have to build around the work. That system has a few parts. Someone is accountable for checking the output who did not produce it. There is a clear line between the work that can fail cheaply and run fast and the work that cannot and has to be gated. And there is a standing assumption that “finished” is a claim to be tested, not a status to be trusted.

So the question to carry out of every one of these meetings is not how good is this tool. It is quieter, and far more useful. How will we know when it is wrong? And the time to answer it is before the work ships, not after.

What looks like a question about the quality of a tool is, underneath, a question about the system you have built around it. The gap no one is telling you about is not a flaw in what they are selling. It is the work that just became yours.

 

The verification problem and the idea of “phantom completions” in this piece are drawn from J.M. Crider’s Harness Engineering for Vibe Coders (2026); the reading of them for leaders is my own.

If this reframes how you’re seeing the challenges you’re facing, a conversation can help turn that perspective into clearer choices.

(No agenda required — it's just a conversation.)

About Jeff Hayes

Jeff Hayes works with senior leaders navigating complexity, pressure, and change. His work focuses on helping leaders slow down, see patterns more clearly, and make sound decisions in uncertain conditions.