We are asked to build a great deal of AI that should not be built. The projects that pay off share a small number of characteristics, and the ones that fail share the opposite.
The test
- Volume: the task happens often enough that saved minutes accumulate into something real
- Rules: a competent new employee could be taught the task from a written document
- Tolerance: an occasional error is recoverable and detectable, not catastrophic and silent
- Data: the information required already exists somewhere machine-readable
- Baseline: you can measure what the task costs today, so you can prove the change
What passes
Enquiry classification and routing. Extracting line items from supplier invoices. Drafting first-pass responses to repeated questions. Summarising call notes into a CRM. Checking documents against a compliance checklist. None of it exciting, all of it valuable.
What fails
Anything requiring judgement your senior people would disagree about. Anything where a confident wrong answer reaches a customer without review. Anything where the underlying data is scattered across inboxes and nobody's spreadsheet. And anything justified primarily by the desire to be seen doing AI.
Start narrow
One workflow, real data, an accuracy target agreed in advance, and a measured baseline to compare against. Prove the return on something small before the scope grows. Projects that begin as a platform strategy rather than a single workflow tend to end as expensive infrastructure nobody uses.