Degradation is gradual and silent

Nothing announces that an AI employee has got worse. Knowledge ages, the model shifts, the business changes, and the first signal is usually a customer complaint months later — unless somebody looks on purpose.

1. What did it escalate, and was that right?

Escalations are self-reporting and quick to scan. Rising volume means scope is wrong; falling volume may mean it has started answering things it shouldn’t.

2. What did it resolve that it shouldn’t have?

The expensive failures never escalate, so this needs a random sample of resolved conversations rather than a look at the exception queue.

Thirty minutes a week is the difference between a system that improves and one that quietly rots.

3. What couldn’t it answer?

The unanswered questions are a ranked list of what your documentation is missing, and the cheapest content research available — the exhaust worth reading.

4. What changed in the business?

New prices, new policies, a product change, a seasonal cut-off. If nobody connects those to the knowledge base, the system keeps confidently stating the old version.

5. What goes into the evaluation set?

Anything found this week becomes a permanent test, so the same problem cannot recur unnoticed. That loop is what makes review compounding rather than repetitive.

One named person, every week

A review that belongs to everyone happens to nobody. It is the concrete form of ownership, and it should survive holidays and busy periods — those are when drift accumulates.