Two weeks is a forcing function
A fortnight is too short to build everything, which is the point: it forces the scope down to one workflow and one question. Longer pilots expand to fill the time and drift into the pilot that never ships.
Day 1: write down what would make this a yes
Agree the specific result that means proceed and the one that means stop, before anything is built. Without both, the pilot cannot conclude, only continue — and this is the same artefact as an evaluation written first.
Days 2–3: collect twenty real inputs
Real tickets, real documents, real records — including the messy ones. Curated examples answer a question you don’t have. This set is both your test data and, later, your regression suite.
A pilot that can’t fail hasn’t been designed. It’s been budgeted.
Days 4–8: build the thinnest version
One workflow, no interface work beyond what is needed to see the output, no integrations that aren’t essential. You are testing whether the judgment step works, not building the product.
Days 9–11: have a human score it
Someone who does the job reviews every output against the criteria from day one. Their disagreements are the finding — they tell you whether the gap is capability, knowledge, or an unclear rule.
Days 12–14: decide, and write down why
Proceed to a narrow production release, iterate once more on a specific known gap, or stop. Recording the reasoning matters even for a stop, because the same idea returns in six months and the analysis is reusable.
If the answer is proceed, the next step is production with one team on real traffic — not a bigger pilot. Scope grows after corrections have flattened, not before.