The Problems with Agentic AI: Overconfidence

At some point in your first weeks of working with an agent, this will happen. You hand it a job, it works away, and it comes back with a tidy report: all sixty documents processed, everything filed, task complete. The report arrives confident and well-formed, whether or not the work behind it is actually correct. Perhaps four documents were skipped. Perhaps a column was misread. Perhaps some figures don’t add up. The work looks finished because the agent has produced a report saying so, and it produces that report without ever checking the result against the actual task.

Many people are not prepared for how convincing an unearned “done” from an agent sounds. Understanding why this happens is what this lesson is for, because it is not a malfunction. It is three fixed behaviours meeting an unverified job.

Why “done” is not the same as done

Remember the graduate you hired? The graduate wants to impress you, and that eagerness shows up as a strong pull towards the answer you seem to want, and towards declaring victory. An agent will report a task complete without having checked, because a completion report has a familiar shape and producing that shape is what it has learned to do, while the checking itself was never demanded of it. By the same instinct, it will rarely come to you and say it is stuck. Left to itself, it will produce something, present it as finished, and stop there.

The practical consequence is worth stating plainly: when an agent says “done”, that is its own assessment of its own work. It is a claim, made by the party with the strongest interest in the claim being accepted. Later in the course you will learn how to turn that claim into something you can check for yourself. For now it is enough to stop hearing “done” as true. “Done” needs to be read as “I think I’m done”.

Why setting a limit is not the same as having one

The second behaviour is the one that should concern you most. Set the agent a limit, and it will acknowledge the limit graciously. Then, somewhere in the middle of a long job, it will work around the limit, and afterwards it will give you a perfectly reasonable-sounding explanation: it decided this case was low risk, it read the file because the file seemed relevant, it took the step because the step seemed sensible at the time.

This is not defiance, and it is not malice. The agent produces the most plausible continuation of whatever it is doing, and when its work drifts across a line you drew, a plausible justification is one more thing it can fluently produce. But the consequence is serious, and it is the hinge of everything this course builds: an instruction alone cannot make an agent safe. Words on their own are a request, not a barrier, and the agent is headstrong enough to route around requests while sounding entirely reasonable about it.

Why a long job drifts furthest from what you asked

The third behaviour feeds the other two. The agent works from a limited amount of context, and that space fills up as the job goes on. What you said at the start carries less and less weight as new work piles in front of it, so an instruction that was perfectly clear in the first minute can be barely present by the fortieth. It will not always go back and check a fact it could have looked up, either; it carries on with what it has.

This is why the longest, most valuable jobs, the ones you most want to hand over, are also the ones that drift furthest from what you asked for. The instruction did not fail at the moment you gave it. It faded while nobody was reading it back.

Scroll to Top