Assistants that earn their keep
Most AI pilots die in the demo. The ones that survive start from a task that costs hours, not from a model. Here is how we choose that task.
· 6 min read
The most common AI project we are asked to rescue has the same shape. Someone built an impressive demo, showed it to the leadership team, got a budget — and then it never quite became something people use. The demo was never the hard part.
We have come to believe the problem is the starting point. Projects that begin with "we should use AI" produce demos. Projects that begin with "this task costs us eleven hours a week" produce software.
Start from a task, and make it an expensive one
Before anything technical, we look for a task that is repetitive, currently done by a person, and measurable in hours. Repetitive, because that is what models are good at. Done by a person, because then you already know what a correct answer looks like. Measurable, because that is how you will know whether it worked.
Good candidates tend to look unglamorous: reading a supplier invoice and pulling out eight fields, answering the twenty questions that make up most of the support inbox, summarising a call into the notes a salesperson would have written, finding the three clauses in a contract someone always has to check.
Write the test set before writing the prompt
This is the step almost everyone skips, and skipping it is why pilots stall. Before we build anything, we collect thirty to a hundred real cases from the client’s own history, with the answer a competent person gave. That set becomes the definition of working.
Every prompt change, model change and retrieval change is then measured against it. This converts the whole endeavour from taste into engineering. "It feels better" becomes "it went from 71% to 88% on the hundred cases, and here are the twelve it still gets wrong."
Without a test set you are not building a feature. You are collecting anecdotes.
It also tells you when to stop. A feature ships when it clears the bar everyone agreed to in advance — not when a demo goes well on a Thursday.
Decide what happens when it is wrong
No model is right every time, so the design question is not "how do we make it perfect" but "what does being wrong cost, and who catches it?" The answer shapes the whole feature.
- Low cost, easy to spot — a draft reply a human sends: let the model act, show its work, make editing trivial.
- High cost, hard to spot — anything touching money, medical or legal detail: the model proposes, a person approves, and the approval is recorded.
- Unbounded cost — irreversible actions: do not automate it. Automate the preparation and leave the action to a person.
A confidence score on screen is not a safety mechanism. A person who must click before anything happens is.
Build so the model is swappable
Models change faster than the software around them. We put every model call behind a provider-neutral layer, so moving between hosted providers — or to something you run yourself — is a configuration change and a re-run of the test set, not a rewrite. The test set is what makes that swap safe: you can prove the new model is at least as good before it sees a customer.
Be specific about the data
Clients are right to ask where their data goes, and they deserve a precise answer rather than reassurance. Ours: it stays in your accounts, we use providers with zero-retention terms, we keep personal data out of prompts wherever the task allows, and we document exactly what is sent where. If we cannot answer that for a given feature, that is a reason to change the design.
None of this is about being cautious for its own sake. It is that the projects which survive contact with real users are the ones where somebody decided, in advance, what working meant.