1. Scope that was never narrowed

“Handle customer questions” is not a job. Without explicit exclusions the system attempts everything and is unreliable at all of it — which is what a job description exists to prevent.

2. Knowledge that contradicts itself

Two documents disagreeing produces a confident wrong answer, because retrieval picks by similarity rather than correctness. Resolving conflicts beats writing anything new — the first pass of the knowledge work.

Almost every AI employee failure is an organisational decision nobody made, not a model that wasn’t clever enough.

3. Escalation with nowhere to land

A hand-off into an unwatched inbox moves the failure rather than removing it. The person on the other end has to exist, during the hours you claim to cover.

4. Nobody reviewing the resolved cases

Teams review escalations, which are self-reporting, and never sample what the system settled alone. The expensive failures live there — hence random sampling of resolved work.

5. No owner

Deployed by one team, used by another, maintained by nobody. It works at launch and degrades quietly as the business moves — the ownership gap is the most common root cause of all.

6. Scaling before it was stable

A second and third workflow added while the first still generates daily corrections multiplies review load and hides which system is failing. Expand when corrections flatten, not on a date.