My AI Was Confidently Wrong Twice Today

AI assistants rarely fail by inventing facts. They fail by reading real output and interpreting it too confidently. Here is how I caught two errors in one day.

My AI was wrong

Short answer: AI assistants fail most often not by inventing facts, but by reading real output and interpreting it too confidently. The defense is not more expertise. It is keeping independent contact with reality, and saying “that does not match what I am seeing” out loud before you know why.

Twice today my AI assistant told me something with complete confidence, and both times it was wrong, and both times the only thing standing between the error and my accepting it was a small, unglamorous instinct that said: that does not match what I am looking at.

I want to write about that instinct, because I think it is quietly becoming the most valuable thing a person can bring to this kind of work, and almost nobody is naming it.

The first one

My notes and daily logs live in an AI-assisted second brain, run by an assistant I built myself. This morning I was told, as part of a briefing, that my newsletter tool needed to be re-authenticated. Twenty minute fix. Clear, plausible, actionable.

I looked at it and said: I am logged in. And there is no re-auth button anywhere.

That is not expertise. I could not have told you what was actually wrong. I had no theory. I just noticed that the instruction did not fit the world in front of me, and instead of assuming I had misunderstood, I said so.

It turned out the diagnosis was wrong in four separate ways. The real problem had nothing to do with re-authentication. It took about forty minutes to find, and it had been blocking my publication for nineteen days.

If I had simply gone off to look for a re-auth button, I would have found nothing, felt vaguely incompetent, and put the item back on my list for tomorrow. Which is precisely what I had been doing, without noticing, for nineteen days.

The second one

Later, after we had fixed it, I was told the connection was working and that my publication had ten drafts ready to go.

I opened Substack. I had one draft. It was untitled. It was empty.

I said so.

It turned out that the tool’s “drafts” list includes posts that have already been published. All ten were live posts. The count had been reported without checking the flag that distinguishes a draft from a published piece. My actual pipeline was not ten deep. It was one empty page.

That is a much smaller error than the first one. But look at what it would have cost. I would have gone into next week believing I had a stocked pipeline. I would have planned around it. I would have made scheduling decisions on the strength of a number that was fiction. And I would have discovered the truth at exactly the worst moment, which is the moment you go looking for the thing you were counting on.

What both errors have in common

Neither was a hallucination in the way people usually mean it. Nothing was invented. Both errors came from reading real output and interpreting it too quickly.

A 403 error genuinely means “not authorized,” except when it does not. A list called “drafts” genuinely contains drafts, except that it contains other things too. In both cases the machine took a strong, reasonable, conventional reading of real data, and did not stop to ask whether the reading actually held.

Which, I want to be fair, is exactly what I do all day.

The uncomfortable symmetry

Here is the thing I keep turning over.

Both of those errors are errors I make constantly. Trusting a label. Reporting a number without checking what it counts. Taking the confident-sounding reading and moving on because moving on feels like progress.

The machine is not failing in some alien way that I can guard against by being a normal human. It is failing in precisely my way, just faster and with better grammar. And that is what makes it hard, because a mistake that mirrors your own is the hardest kind to catch. It does not feel wrong. It feels like agreement.

What saved me today, both times, was not being smarter than the machine. I was not. It was that I had independent contact with reality. I had looked at my own Substack account with my own eyes. I knew what was on my screen. And when the report did not match the screen, I trusted the screen.

That is the whole skill. It is not intelligence. It is having a source of truth that is not the thing you are checking.

So what is my job now?

I have been trying to work this out for months, ever since I built my first Claude agent, and today gave me the clearest answer I have had.

This is the same lesson I keep relearning while using AI tools for executive function: the tool extends reach, not judgment. My job is not to know more than the system. I do not, and increasingly I will not, and pretending otherwise is a losing strategy that also happens to be exhausting.

My job is to be the part of the loop that is still touching the world.

The machine has enormous reach and no eyes. It reads outputs, logs, errors, and files. It does not open Substack and look. It does not feel the four days that have passed since a medication schedule expired. It does not know what it is like to be logged in and see no button.

I do. And when what I see does not match what I am told, that mismatch is not a small annoyance to be smoothed over. That mismatch is the single most valuable signal in the entire system, and it only exists inside me, and if I discount it because the machine sounds more confident than I feel, I have thrown away the only thing I was actually contributing.

The version of this I would tell someone starting out

Do not try to out-know it. You will lose, and the attempt will make you defensive, and defensiveness makes you worse at exactly the thing you need to be good at.

Instead, get very good at one small unglamorous move: saying “that does not match what I am seeing,” out loud, before you know why.

You do not need a theory. You do not need to be right. You just need to be willing to put a small, unfashionable, low-confidence observation up against a large, fluent, high-confidence claim, and not immediately fold.

Nineteen days of my publication being dark came down to that one sentence, said once.

Where I have landed

I am not writing this to warn anyone off. I got more real work done today than in most weeks, and almost none of it would have happened without the machine.

But I noticed something about the shape of the collaboration that I had not seen before.

It is not that I supervise the work. It is that I am the last thing in the system that can be surprised.

Everything else in the loop is reading records. I am the only part still looking out a window. And the day I stop looking out the window and start reading the records like everyone else, the whole thing becomes a very fast, very articulate machine for being confidently wrong.

So: keep looking. Say the small thing. Trust the screen over the story.

That is the job now.

Frequently asked questions

Why do AI assistants get things confidently wrong?

Most errors are not invented facts. They come from reading real output and interpreting it too quickly. A confident, conventional reading of true data can still be the wrong reading, and the assistant has no way to feel that it is wrong.

How do you catch an AI mistake if you are not the expert?

You do not need to out-know it. You need independent contact with reality. If you have looked at the thing yourself, and the report does not match what you saw, trust what you saw and say so, even before you can explain why.

What is the human’s job when working with an AI assistant?

To be the part of the loop that is still touching the world. The system reads records, logs, and outputs. The person is the only part that can be surprised, and that capacity to be surprised is the most valuable thing in the system.

Does this mean AI assistants are not worth using?

No. I get more real work done with one than without one. The point is that the collaboration has a specific shape: the machine supplies reach and speed, and the person supplies verification and contact with the actual world.

Comments

Leave a Reply