Skip to main content

Work

A Quarterly AI Audit You Can Run Without Hiring a Consultant

A half day agenda for reviewing what your AI agents did last quarter: the records to pull, how much to sample, and the four findings that demand a change.

Written by Sicherhaven

Your team turned on AI tools months ago and nobody has checked what came out of them since. You do not need an outside firm to fix that. A quarterly AI audit is a half day of work: pull a set of records, sample them, and answer four questions honestly.

The point of the audit is not to grade the technology. It is to find out where agent output is being accepted without anyone reading it, and to change one or two things before the next quarter starts.

What the audit is actually asking

Keep the scope small enough to finish in one sitting. Three questions cover most of it.

  • Where did agents produce work, and who signed off on it?
  • Where did a person accept output without reading it?
  • Where did the agent work from stale or missing context?

Anything else is interesting, but it will not change what you do on Monday.

Records to pull before you sit down

Book the room, then spend an hour gathering. You want the raw trail, not a summary someone prepared for the meeting. The hour only stays an hour if you have been keeping a record of who approved what all along.

  • Every task or ticket in the quarter where an agent drafted, summarised or generated something.
  • The approval record for each of those: who approved, and how long after the draft appeared.
  • The final version that shipped, next to the draft the agent produced.
  • Any complaint, correction or rework item that traces back to one of those tasks.
  • The list of records the agent could read, so you know what it was working from.

If your tools keep project work, HR records and agent activity in separate systems, this hour becomes three. That gap is the reason SicherOne puts project management, HR and AI agents on one set of records: the audit trail is a query rather than an archaeology project. A human still approves agent output before it ships, which is what makes the approval column meaningful at all.

The half day agenda

Ninety minutes of review, then decisions. Do not let it run longer.

First 30 minutes: read samples together. Put the draft and the shipped version side by side. Read them out loud if the group is small. You are looking for edits, and for their absence.

Next 30 minutes: check the approvals. For each sampled item, ask who approved it and what they changed. An approval with no edits is not automatically bad. An approval that landed two minutes after a long draft appeared is worth a question. It is the same reading skill as a manager's Monday morning pass over an agent's work log, done in bulk.

Next 30 minutes: check the inputs. For a handful of items, list what the agent could see. Then ask whether the right answer needed something it could not see.

Final hour: pick changes. Two at most. Write them down with an owner and a date.

How much to sample

You are not doing statistics. You are looking for patterns, and patterns show up fast.

  • Under 50 agent assisted items in the quarter: read all of them.
  • 50 to 200: sample 20, chosen at random rather than by whoever volunteers them.
  • Over 200: sample 20 at random, plus every item that produced a complaint or a rework.

Random matters more than volume. If someone hand picks the sample, you will audit the best work in the building.

Four findings that should trigger a change

Most audits turn up small annoyances. These four are different, because each one means the current setup will keep producing the same problem next quarter.

Rubber stamp approvals. A reviewer approving long drafts within seconds, repeatedly. The fix is usually fewer approval steps with more weight, not more steps.

No named owner. Output that shipped with an approval from a shared inbox or a rotating queue. If nobody owns it, nobody reads it.

Blind spots in context. The agent gave a confident answer that a person would have caught, because the agent could not see a record the person had. Widen the access or narrow the task.

Drift between draft and shipped. Heavy editing on every sampled item means the agent is doing work in the wrong place. That is a scope problem, not a quality problem.

Write it down, then stop

One page. What you sampled, what you found, the two changes and who owns them. Date it and file it where next quarter's audit will find it. It doubles as evidence, alongside what an audit trail should contain if a regulator ever asks.

The value compounds when you can compare quarters. A finding that shows up twice is no longer a finding. It is a decision you have been avoiding.

← All posts

We're building the future of community events and financial wellness

See how Eventify and WealthWise change the way people find events and manage money.

Get Started