Oxford Intelligence / Home

OPERATING NOTES · 17 SEPTEMBER 2026

An evidence checklist for an applied AI pilot

Decide what you need to learn before deciding how far to scale.

At Oxford Intelligence, our operating principles distinguish demonstrated capability, active development and future ambition. This checklist turns that distinction into a practical exercise for teams considering an AI-assisted workflow. It is an editorial planning tool, not a report of measured customer results.

1. Define one useful task

Write down who performs the work today, what starts it, what information is needed and what a usable result looks like. Avoid treating a broad ambition such as ‘improve productivity’ as a testable task. A pilot might help a team assemble evidence for a review; that does not mean it is authorised to make the final decision.

2. Record the starting point

Observe a representative sample of the existing workflow before changing it. Record completion time, corrections, unresolved cases and the effort required from reviewers. Keep the sample and its selection method visible. A comparison is only useful when the tasks and quality standards are sufficiently similar; a demonstration on unusually easy cases is not a baseline.

3. Name the human decision owner

Separate preparing information, proposing an action and authorising that action. Identify who can accept, correct or stop the result. Give the reviewer the relevant source evidence and a clear indication of uncertainty. Responsibility should be part of the working process, rather than an assumption that somebody will check later.

4. Test an exception

Choose a difficult case that the team already encounters: missing information, conflicting records or an unusual request. Record whether the system asks for clarification, routes the case to a person or stops. A confident-looking answer is not a successful result when the necessary evidence is absent.

5. Include the cost of checking

Measure the work needed to verify and correct an output alongside any time saved in producing it. Note whether mistakes are easy to spot and reverse. Review sample quality as well as speed; faster completion is not an improvement if the team must repair hidden errors later.

6. Decide what earns the next stage

Before reviewing results, agree what would justify continuing, narrowing the scope or stopping. Use thresholds appropriate to the task and its consequences, not a universal accuracy figure. A small pilot supports a decision about the next test; it does not establish that every customer, document or operating condition is covered.

Your one-page pilot record

  • Task and intended user
  • Existing workflow and dated baseline sample
  • Information the system may use
  • Actions it may take and actions reserved for a person
  • Exception and correction procedure
  • Quality, reviewer effort and completion-time measures
  • Evidence collected, remaining uncertainty and next decision

Keep sensitive customer information out of public case studies. Share only evidence you are authorised to disclose, and clearly distinguish observed results from intended benefits.

For the principles behind this worksheet, see Oxford Intelligence’s operating principles. To discuss a possible technology partnership, visit our partnership section.

Oxford Intelligence Limited is an independent company and is not affiliated with or endorsed by the University of Oxford.