Alice LawsonChief Data Officer · Essays & perspective
← All essays

Decision intelligence2 min read

AI evaluation starts with purpose and exposure

An AI use case needs to be assessed in the context of the decision or task it supports.

The consequences of error, the information involved and the human responsibilities shape what evaluation and monitoring are appropriate.

I want teams to state those conditions before moving from an experiment into business use. Technical performance remains important, but it is one part of an operating arrangement that must be explainable, reviewable and accountable.

AI evaluation tends to drift towards technical benchmarks because they are measurable and familiar to the teams building the system. Accuracy on a test set, ratings of response quality and speed are useful, but they say little about what happens when an output is wrong in a particular business context. A drafting assistant used for internal notes and the same assistant used for customer correspondence present very different exposures. When evaluation is designed around the technology rather than the use, organisations can approve a capable system for a purpose it does not suit.

My practice is to ask the sponsoring team for a short use statement before any pilot moves towards business use. It describes the task, the people affected by the output, the information the system will access, the consequence of a plausible error and the person accountable for the result. It also sets out how humans will review outputs and what monitoring will continue after launch. The statement is reviewed with risk and legal colleagues, and its level of detail scales with the exposure it describes.

Consider a team proposing to use a language model to summarise customer complaints for managers. The experiment produces good summaries. The use statement prompts further thought: some complaints contain personal information, a missed detail could hide a serious issue and managers might act on the summary without reading the original. Evaluation then focuses on whether serious issues are reliably flagged, how personal information is handled and whether managers still consult the source. The technical result matters, but the operating safeguards determine whether the use is acceptable.

A common objection is that this level of scrutiny will slow innovation and drive experimentation into unofficial channels. That risk is real. My answer is to make the path proportionate and quick for low-exposure uses, so that teams have little reason to avoid it. An internal drafting aid with no personal data might need only a brief record and a named owner. Uses that affect customers, employees or significant decisions receive more attention. A process that treats every experiment as high risk will be bypassed, and then it protects nobody.

This asks the leadership team to accept that accountability for an AI-assisted outcome remains with a named person in the business, not with the technology or the data function. It asks risk, legal and data colleagues to work from a shared understanding of exposure rather than separate checklists. And it asks executives to fund monitoring after deployment, which is less visible than the original experiment but often more important. Without that ongoing evidence, the organisation cannot say with confidence whether a use is still behaving as intended.

Boards increasingly ask how AI is governed. The most useful answer describes purpose, exposure, human accountability and the evidence that would trigger a change in use. That answer is only credible if the organisation can show it working on a real system, not only in a policy document.