The label stopped carrying information

“AI-powered” now describes everything from a genuine reasoning feature to a keyword matcher with a new badge. The word tells you nothing, so the evaluation has to be built from questions about behaviour instead.

1. What does it do when it doesn’t know?

The most revealing question you can ask. A vendor whose product has a supported way to decline and escalate has thought about production; one whose demo always answers has not.

2. Can I see it fail?

Ask them to run it on your messiest real inputs, not their examples. Reluctance here is the single strongest signal in the whole process, because every product looks good on curated data.

Ask to see it handle your worst data. How they respond tells you more than the answer would.

3. What happens to my data?

Retention, region, training use, sub-processors, and deletion. These are contractual answers with the same weight as any dependency on someone else’s model, and vague responses should be treated as no.

4. How do I measure whether it works?

A vendor who can tell you which numbers to watch and how to get them is confident. One who offers only usage statistics is selling activity — the same distinction as deflection rate versus confirmed resolution.

5. Who reviews the output, and how?

Ask what the review workflow looks like in practice and how much staff time it needs. That is the hidden cost that determines whether the tool pays back, and it rarely appears in pricing.

6. What if we leave?

Whether your data, corrections and configuration come with you. The corrections especially — they are the accumulated asset, and a product that keeps them has locked you in more effectively than any contract term.