Work
Why an Agent That Says It Does Not Know Saves More Time Than One That Guesses
An AI agent that admits it does not know saves more time than one that guesses, because verifying confident answers costs more than answering fewer questions.
Written by Sicherhaven
A tool that answers every question feels more useful than one that sometimes says it cannot help. The arithmetic says otherwise. An AI agent that admits it does not know saves more time than one that guesses, because checking confident answers costs more than doing without a few of them.
This is worth treating as a feature to demand from vendors, in the same way you would ask about uptime. Ask what the system does when it is unsure, and ask to see it happen.
The hidden cost of a confident wrong answer
Compare two failures.
An agent that says "I could not find this, here is where I looked" costs you the time to go and look yourself. That is the same time the task would have taken anyway. You lost nothing.
An agent that invents a plausible answer costs you differently. Somebody has to notice it is wrong, which takes attention. Then they have to work out what is wrong, then repair whatever was built on it. If nobody notices, the cost lands later and larger.
A refusal costs you the time you would have spent anyway. A confident wrong answer costs detection, investigation, repair and everything built on it in between.
The verification problem
Here is the part that decides the whole argument.
If an agent is usually right but occasionally wrong, and you cannot tell which is which by looking, you have to check everything. That means the true cost of using it is the checking cost across all answers, not the correction cost on the few bad ones.
Checking an answer is often nearly as much work as producing it. So an agent that answers everything with a handful of quiet errors mixed in can easily cost more than it saves, while one that answers most questions and flags the rest saves real time, because the flagged ones are the only place attention has to go. That only pays off if the flag lands in an approval step people actually read.
Abstention is what makes the output sortable. Without it, every answer sits in the same undifferentiated pile.
What good abstention looks like
Not all "I do not know" is equally useful. The version that saves time tells you three things:
- What it was trying to find
- Where it looked
- What it would need in order to answer
That turns a dead end into a task. Somebody can go and supply the missing piece, and the next attempt works.
Compare that with a bare refusal, which just moves the whole problem back to the person with no head start.
What to ask a vendor
Most demos are built to avoid this. Ask directly.
- What does the system do when the information is not in the records it can see?
- Does it distinguish between not found and not sure, and how?
- Are refusals and low confidence responses logged, or only successes?
- Can we see the rate of refusals over time?
- Can we tune how cautious it is per task?
That last one matters because the right level differs by job. An internal brainstorm can tolerate guessing. An answer about someone's leave entitlement cannot.
Ask for a live demonstration with a question the system cannot answer. Watch what happens. A system that produces something fluent and wrong in front of you will do it every day once installed.
Why context reduces the need to guess
Much of the guessing that annoys people comes from a thin view. An agent that can only see task titles will invent the rest, because that is all it has.
Give it more of the surrounding record and there is less to invent. SicherOne runs project management, HR and AI agents on one set of records, so an agent answering a question about a project can see who owns the work and who is on leave rather than filling in the gap. A human approves agent output before it ships, which catches what remains.
Measuring whether it is working
Two numbers tell you most of it.
Refusal rate. Track it monthly. A rate of zero is a warning, not a success. It means either your questions are trivial or the system is guessing.
Correction rate on accepted answers. How often does a reviewer change something before it goes out? If refusals are low and corrections are high, the system is confident in the wrong places. A correction rate near zero can mean the opposite problem, that approval has quietly become rubber stamping.
Watched together, they show whether caution is set sensibly. Rising refusals with falling corrections usually means the balance moved in your favour.
The uncomfortable part
A system that admits uncertainty demos worse. It looks less impressive in a room full of people watching it answer questions, and buyers reward the one that always has something to say.
That is precisely why it is worth asking about. The tool that never says no is not more capable. It has just moved the work of finding its errors onto you, and it did not tell you it was doing that. Splitting the job so drafting, reviewing and sending are separate steps puts that work back where somebody can see it.
← All postsWe're building the future of community events and financial wellness
See how Eventify and WealthWise change the way people find events and manage money.
Get Started
