Degradation is gradual and silent
Nothing announces that an AI employee has got worse. Knowledge ages, the model shifts, the business changes, and the first signal is usually a customer complaint months later — unless somebody looks on purpose.
1. What did it escalate, and was that right?
Escalations are self-reporting and quick to scan. Rising volume means scope is wrong; falling volume may mean it has started answering things it shouldn’t.
2. What did it resolve that it shouldn’t have?
The expensive failures never escalate, so this needs a random sample of resolved conversations rather than a look at the exception queue.
Thirty minutes a week is the difference between a system that improves and one that quietly rots.
3. What couldn’t it answer?
The unanswered questions are a ranked list of what your documentation is missing, and the cheapest content research available — the exhaust worth reading.
4. What changed in the business?
New prices, new policies, a product change, a seasonal cut-off. If nobody connects those to the knowledge base, the system keeps confidently stating the old version.
5. What goes into the evaluation set?
Anything found this week becomes a permanent test, so the same problem cannot recur unnoticed. That loop is what makes review compounding rather than repetitive.
One named person, every week
A review that belongs to everyone happens to nobody. It is the concrete form of ownership, and it should survive holidays and busy periods — those are when drift accumulates.