How to audit an AI employee's output
An audit trail is only worth what somebody actually reads. A practical monthly routine for sampling an AI employee's conversations and correcting what the reading finds.
Key takeaways
- Every build ships with an audit trail. The audit is the habit of reading it, and that habit is yours.
- Sample deliberately. Take every escalation, the longest threads, the enquiries that ended with nothing, and a small random slice.
- Judge each conversation against the script you approved, because that document is the standard the employee was built to.
- Most failures are boundaries in the wrong place or a voice that drifted, and both are faults in the script.
- Write the correction in, approve the script again, date it, and read the same slice next month.
Auditing an AI employee is reading a sample of its conversations against the script you approved, and correcting the script where the two disagree. It takes about an hour a month for a small firm, needs no technical skill, and it is the only thing standing between an audit trail and a filing cabinet nobody opens. Every build ships with the trail. The reading is the part that has to be somebody's job.
The routine below covers how often to do it, which conversations to pull, what a failure looks like when you find one, and what to do about it. The security page covers where the data behind that trail sits and who can reach it.
How often should you audit an AI employee?
Monthly through the first quarter, then whenever the monthly report shows something you did not expect. The early weeks are when the approved script sits furthest from the conversations people actually have, so that is when reading a sample changes the most. Once the pattern settles, a quarterly read alongside the monthly report is usually enough for a firm of ten.
There is a fixed checkpoint alongside your own reading. At week six there is a written adoption check that asks whether the workflow you hired the employee for has actually changed, and the first ninety days are set out in working with AI employees. Your monthly audit is the thing that gives that check something to read.
Which conversations should you pull for an audit?
Four slices, and a random sample on its own is not one of them. Take every escalation, because those are the moments the employee judged itself out of its depth. Take the longest threads, because length usually means confusion. Take the enquiries that ended with no booking and no handover, because that is where work leaks away. Then add a small random sample, which is the only slice that shows you what an ordinary day looks like.
| Slice | Why this slice | What a good result looks like |
|---|---|---|
| Every escalation | These are the moments the employee judged itself out of its depth | It handed over early, with the context a colleague needed to pick it up |
| The longest threads | Length usually means the conversation went round in circles | The enquiry was genuinely complicated, and the length tracks that complication |
| Enquiries that ended with no booking and no handover | This is where work leaks away unnoticed | The enquiry was outside what you sell, and closing it was correct |
| A small random sample | The only slice that shows you an ordinary day | It reads like your firm, and you would have sent it yourself |
- Why this slice
- These are the moments the employee judged itself out of its depth
- What a good result looks like
- It handed over early, with the context a colleague needed to pick it up
- Why this slice
- Length usually means the conversation went round in circles
- What a good result looks like
- The enquiry was genuinely complicated, and the length tracks that complication
- Why this slice
- This is where work leaks away unnoticed
- What a good result looks like
- The enquiry was outside what you sell, and closing it was correct
- Why this slice
- The only slice that shows you an ordinary day
- What a good result looks like
- It reads like your firm, and you would have sent it yourself
The four slices a monthly audit pulls, and what each one is looking for. The slices are named in the answer above; this table restates them in the order they are read.
What does a failed audit actually look like?
Usually a boundary in the wrong place. The employee answered something it should have passed to a person, or held back something it could have handled and left the client waiting. Voice drift is the other common finding: a reply that is accurate on the facts and sounds nothing like your firm. Both are faults in the approved script, and both have a written fix.
That is worth sitting with, because the instinct on finding a bad reply is to distrust the technology. Almost every time, the reply was a faithful performance of an instruction somebody approved months ago and has not read since. The employee did what the script said. The script was wrong about a case nobody had met yet.
What do you do with what an audit finds?
Write the correction into the script, approve it again, and date it. An AI employee works only inside scripts you have approved, so a correction that stays in your head changes nothing about what goes out tomorrow. Then read the same slice at the next audit. Checking the same place twice is how you tell a fixed problem from one that has quietly moved somewhere else.
Keep the corrections somewhere dull and dated. A single running list of what changed, when, and who approved it turns a year of audits into a record you can hand to a regulator, an insurer, or the person who takes the job over from you.
Can you audit an AI employee without technical skills?
Yes. An audit here is reading conversations and asking whether you would have been content for a colleague to send them. There is no model to inspect and no logs to decode. The trail holds the same plain words your client saw, alongside the approved script the employee was working inside and the point at which it handed the matter over.
What an audit will not tell you
It will not tell you about the enquiries that never arrived, and it will not tell you whether the role was scoped to the right job. A trail full of tidy conversations about a workflow nobody needed is a clean audit and a wasted hire. That question belongs to the week-six adoption check and the ninety-day usage telemetry, which measure whether your team is using the thing at all, and both are set out on how we work.
It also will not do your governance for you. The audit tells you what happened; the decisions about what should be allowed to happen sit with a named person on your team, written down before launch.
The hour a month version
For a firm of ten, the whole routine is one hour with a coffee. Open the trail, pull the four slices, read them against the approved script, and write down the two or three lines you want changed. Send those to whoever maintains the script, approve the new version, and diarise the same hour next month.
An audit that has to be remembered gets done twice and then stops, so the first hour belongs in the diary before the build goes live. A trail nobody opens is indistinguishable from no trail at all. The AI employees page describes what the role covers, which is what the trail will end up holding.
Where this leads
Who owns an AI employee, what its audit trail records, and how the adoption check works.
Or run your own figures and check the assumptions while you are there.