Skip to main content

Work

Six Numbers to Track Before You Switch On an AI Agent

Six baseline numbers to capture in the two weeks before an AI agent goes live, because none of them can be reconstructed once the old process has changed.

Written by Sicherhaven

Six months after switching on an AI agent, somebody will ask whether it helped. If you did not measure the old way of working first, the honest answer is that nobody knows. Capture six numbers in the two weeks before rollout, because none of them can be reconstructed afterwards.

They are all cheap to collect. The reason teams skip them is that the fortnight before launch is busy, and the cost of skipping only shows up much later.

The six numbers, one by one

One: how long the task takes now

Time the existing process end to end, for the same job the agent will do. Not the ideal version, the real one including the interruptions.

Ask three or four people to note start and stop times for a week. Take the spread as well as the average, because the slow cases are usually where the value is.

Once the agent is in place, the old timings are gone. Nobody remembers how long invoicing took before, and estimates made after the fact drift towards whatever supports the decision already made.

Two: how often it happens

Volume decides whether any of this matters. A task that takes an hour and happens twice a year is not worth automating no matter how tedious it is.

Count occurrences over a representative period, and note whether they cluster. Ten invoices spread across a month is a different problem from ten on the last Friday.

Three: the current error rate

This is the number people most regret not having. When an agent makes its first visible mistake, the conversation becomes "it is unreliable", and without a baseline nobody can say whether it is worse than the humans it replaced.

Count how often the existing process produces something that needed correcting, and what kind of correction. You do not need perfect data. A tally sheet for two weeks is enough to say whether the current rate is roughly one in five or roughly one in fifty.

The most valuable pre launch number is your current error rate. Without it, the agent's first mistake becomes an argument about trust rather than a comparison.

Four: how long corrections take

For each error you counted, note how long it took to spot and fix. This is the number that turns risk into something you can weigh against savings, because it is in the same unit as the time saved. That is the whole idea behind pricing a wrong agent action in hours.

Include the detection delay. An error found immediately and an error found three weeks later cost very different amounts even when the repair is identical.

Five: how much of the job is waiting

Split the elapsed time into work and waiting. A task that takes four days but only forty minutes of effort has a queue problem, not an effort problem, and an agent that drafts faster will not fix it.

Note where the waiting happens: an approval, a colleague's reply, a system that only runs overnight. If the biggest block is somebody's inbox, an agent may just produce drafts faster and add to the pile.

Six: who does it, and what else they were doing

Record which roles carry the work and roughly what share of their week it takes. Two reasons.

First, the savings only exist if the freed time goes somewhere useful. If a job takes twenty minutes a day across six people, that time will quietly evaporate into other work and no one will feel a difference.

Second, adoption depends on the people involved. If the task belongs to one person who has done it for years, that person's cooperation is the whole project.

Collecting all six without a project

The lightweight version takes two weeks and one shared sheet:

  • A row per occurrence of the task
  • Columns for who, start, finish, waiting time, and whether anything needed correcting
  • A separate note when a correction happens: how it was spotted and how long it took
  • A one line summary at the end of each week

That is it. No tooling, no dashboard, no baseline study.

If your records already sit in one place, a lot of this is easier to pull. SicherOne keeps project management, HR and AI agents on the same set of records, which means task history and who was available sit together rather than in separate systems that have to be joined by hand.

What to do with them after launch

Track the same six for the first quarter after the agent goes live. Compare like for like, and expect the first month to look worse while people learn the new shape of the work. Add a column for time spent checking agent output, because that is the cost nobody budgets for.

The comparison worth watching is not time saved on its own, and it helps to agree in advance on a way of counting minutes that survives scrutiny. It is time saved minus correction time, against the old baseline. That single line settles most arguments about whether to expand, adjust or switch it off.

← All posts

We're building the future of community events and financial wellness

See how Eventify and WealthWise change the way people find events and manage money.

Get Started