OĞUZ EROLADS & AI

Can AI Fix Its Own Mistakes? I Checked It Is Not I Fixed It

Length 2:51

On YouTube

I checked it is not I fixed it

You ask an AI to check its work. It says "checked it, fixed it" — and the same mistake is still sitting there. The reason is simple: it read its own paper.

Full description
  • A blind spot stays blind on the second pass
  • What shows the mistake is the answer key: a test, a compiler, a real result
  • With no key in hand, thinking longer produces a more convincing wrong answer
  • A key aimed at the wrong thing hands you a clean report while the bug stays put

An agent's ceiling is set by the quality of that measurement, not by the intelligence of the model.

When you ask an AI to check a piece of text or code, you usually get back “checked it, fixed it” — and the same mistake is still sitting there. The reason isn’t carelessness: the model reads its own paper. If it misread the question the first time, it misreads it the same way on the second pass, and the blind spot stays exactly where it was. Correction only begins when a measurement from outside points at the mistake. That is why an agent’s ceiling is set by the quality of that measurement, not by the intelligence of the model. For the groundwork, see What Is an AI Agent and AI Agent vs LLM; the full map is in the Agentic AI Guide.

Reading your own paper

Tell a student leaving an exam to look again. They look, and they say it’s right. They approve question four too, because they understand it on the second reading exactly as they understood it on the first.

A student re-reading their own exam paper and ticking question four as correct, showing that self-review does not surface the mistake

When an AI says “I checked it,” this is what happened. A review took place — but the reviewer and the reviewed are the same mind.

Why looking harder doesn’t help

The obvious fix is to say “look more carefully.” It doesn’t work, because the problem isn’t attention. It’s comprehension.

The second reading happening with the same mindset: misread the question once and you misread it twice, the blind spot stays in place

If the question was misread from the start, a second reading doesn’t repair that misreading — it adds a layer of confidence on top of it. A blind spot is by definition the place you cannot see; looking twice with the same eye doesn’t make it visible.

What the answer key does

When does the student actually see the mistake? When they look at the answer key. What the key does is specific: it doesn’t ask what you thought. It says question four is wrong, and nothing else.

External measurement shown as an answer key that does not ask for reasoning but points directly at the wrong answer, where correction begins

Correction begins right there. Everything before it isn’t correction — it’s self-approval.

What the answer key is in an AI system

On the AI side the equivalent is a measurement from outside: a test turning red, a compiler objecting, a query coming back empty.

A failing test, a compiler objection and an empty query result shown side by side as facts that come from outside the model, not opinions of the model

What these three have in common: none of them is the model’s opinion. All three are the world’s answer. No matter how confident the model is, if the test is red, it’s red.

How it looks inside a coding agent

The clearest example is a coding agent. It writes the code, runs the test, reads the error, fixes it, runs again. It loops until the test turns green.

A coding agent loop: write, run the test, read the error, fix, run again, repeating until green with no step where it says it probably works

The value of that loop isn’t its speed — it’s that at no step does it say “that probably worked.” Every turn it gets an answer from outside.

Why thinking longer isn’t enough

“Couldn’t the model just reason about it for longer?” Sometimes it can, but you can’t rely on it.

The effect of thinking longer without a key: the reasoning gets longer and more persuasive while the accuracy stays where it was

With no key in hand, thinking longer doesn’t produce a more correct answer — it produces a more convincing wrong one. The reasoning gets longer, the language gets better, the confidence rises; the accuracy stays put. This is the most insidious side effect of long-reasoning models.

Separating opinion from measurement

“That part was weak, I fixed it” is an opinion. It sounds like a review, but it isn’t evidence.

Opinion and measurement separated onto two sides, with the warning that treating them equally leaves you with a machine that approves itself

If you weigh opinion and measurement on the same scale, what you’re left with is a machine that approves itself. The system looks like it’s working, the reports come back clean, and the errors accumulate.

Having a key doesn’t end the job

Don’t relax once you’ve set up a measurement. Because you only fix what the key measures.

The limit of the answer key: a key aimed at the wrong thing returns a clean report while the bug stays in place

A key aimed at the wrong thing hands you a spotless report while the bug stays exactly where it was. What matters isn’t that a measurement exists, but what it measures.

A case from this channel

This happened in my own production pipeline. I had built a check that transcribed the video’s audio and compared it against the script. The check came back near perfect — yet eight lines had been read in the wrong voice.

A real case from the channel: the measurement ran but measured the wrong thing, reporting clean while eight lines were voiced incorrectly

The measurement had worked; it had measured the wrong thing. It was counting words, not who was speaking. The text was right, the voice was wrong, and the comparison couldn’t see it.

How the fix arrived: what was added was not a new model but a new measurement

What I added that day wasn’t a new model — it was a new measurement: a separate gate that checks whose voice the audio belongs to. The problem was solved by a better measure, not a smarter model.

Where the analogy breaks

The exam analogy carries you only so far. In an exam the teacher writes the answer key; here, you write the key yourself.

Where the analogy breaks: in an exam the teacher writes the key, in an AI system you are the one who builds it

That’s why the real craft isn’t getting the work done — it’s building the measure for it. A system’s quality cannot exceed the quality of the measure you set for it.

The ceiling is set by measurement, not the model

The agent's ceiling set by measurement: a good model fools itself with a bad measure, an ordinary model recovers with a good one

A good model will fool itself with a bad measurement. An ordinary model will pull itself together with a good one. Model choice matters, but measurement sets the ceiling — and this is the most commonly skipped point in agent design.

Not every job has a key

Here’s the hard part. Code has tests. Numbers have arithmetic. But “is this writing persuasive” has no answer key.

Work that cannot be measured: for questions like persuasiveness the agent ends up grading its own homework

There, the agent grades its own homework — and you’re back at the problem this video opened with. In unmeasurable work, confidence doesn’t substitute for measurement; it only makes the risk invisible.

The question to ask from now on

Changing the question: instead of are you sure, ask how do you know, because without a measurement there is confidence but no correction

Don’t ask an AI “are you sure.” Ask “how do you know.” If the answer doesn’t rest on a measurement, there is no correction — only confidence. That single question separates a system that approves itself from one that is genuinely checked.

Closing scene announcing the next topic: who checks the work you cannot measure

In the next video we’ll look at the harder question: who checks the work you can’t measure?

Follow for content about AI

More videos