To evaluate an AI pilot, define the task and its acceptance criteria before you build, test the system on a labelled set of your own real cases, measure it with metrics that fit the task, probe how it fails, and record cost and speed. Roll out only if it passes thresholds agreed in advance, with human review for low-confidence results.
Many AI pilots end with a good demo and an unclear decision. People remember the impressive answers and forget the wrong ones, and nobody agreed beforehand what "good enough" meant. This guide gives you a repeatable way to judge an AI assistant, an agent or a document processing pilot, whoever built it. For background on how these systems are put together, see our guide to LLM app development with RAG and agents.
What should you decide before the AI pilot starts?
Write down the task, what a correct result looks like, and the threshold that would justify rollout, before anyone sees results. Criteria set afterwards tend to drift towards whatever the pilot achieved.
| Decision | Example for an invoice extraction pilot |
|---|---|
| Task | Read supplier invoices and extract supplier, invoice number, date, totals, tax and line items into the ERP |
| Inputs in scope | PDF and scanned invoices from existing suppliers, in the languages you receive |
| Out of scope | Credit notes and handwritten invoices in phase one |
| What happens to the output | Posted to a review screen; nothing reaches the ledger without passing validation rules |
| Baseline | How the current process performs on the same cases: time per document and error types |
| Acceptance criteria | A minimum accuracy per critical field, a maximum share of documents sent to manual review, and zero unvalidated postings |
| Decision owner | The finance manager, not the project team |
The baseline matters. "The AI gets some fields wrong" means little until you know how often the current manual process gets them wrong, and how long it takes.
How do you build a test set from your own data?
Collect real cases that represent the work, label the correct answer for each, and keep them separate from anything used to tune the system.
- Sample from real work, across suppliers, customers, departments, formats and time periods, not only the clean examples someone picked for the demo.
- Deliberately include hard and rare cases: poor scans, unusual layouts, multi-page documents, ambiguous questions, edge-case requests. Rare cases are where pilots fail in production.
- Have domain experts label the answers. For a sample, have two people label independently and resolve disagreements. Where experts disagree, the task definition usually needs tightening.
- Hold the test set back. If prompts or settings are tuned on the same cases used to score the system, the score flatters it. Keep a development set for tuning and a separate test set for the decision.
- Protect the data. Test sets often contain personal or commercial information. Use them only in an environment approved for that data, and redact where you can.
- Version it. You will re-run the same set after every change, so store it with a version number and keep the labels with it.
How many cases you need depends on how many categories, fields or question types you must cover. A useful rule: each important category should have enough examples that a single error does not swing the result.
Which metrics fit which kind of AI task?
Measure what the business relies on, at the level it relies on it. One overall "accuracy" number hides most of what you need to know.
| Task type | Primary metric | What it means in plain words |
|---|---|---|
| Document extraction | Accuracy per field | For each field, how often the extracted value matches the correct value after agreed normalisation (dates, spacing, currency format) |
| Document extraction | Straight-through rate | How often every required field on a document is right, so nobody needs to touch it |
| Classification and routing | Precision per category | When the system says "this is a complaint", how often it really is one |
| Classification and routing | Recall per category | Of all the real complaints, how many the system caught |
| Assistant answering questions | Correctness and groundedness | Whether the answer is right, and whether every claim is supported by the source it cites |
| Assistant answering questions | Appropriate refusal | Whether it says it does not know when the answer is not in its sources |
| Agent completing tasks | Task success rate | How often it completes the whole task correctly, end to end |
| Agent completing tasks | Wrong-action rate | How often it takes an action it should not have, even if it later recovers |
Precision and recall trade against each other. A routing model tuned to catch every complaint will also mislabel some ordinary messages. Decide which mistake costs more for each category: missing a legal complaint is worse than sending a routine message to the wrong queue.
How answers get scored
- Exact checks for extraction and classification: compare against the label automatically.
- Human grading with a written rubric for assistant answers: correct, partly correct, wrong, unsupported.
- Model-assisted grading, where another model scores answers, can speed up large test sets, but check it against human grades on a sample before trusting it.
- Citation checks: open the cited passage and confirm it actually says what the answer claims.
For document-heavy pilots, our intelligent document processing guide covers extraction, validation and exception handling in more depth. For search assistants over company documents, see enterprise knowledge search with RAG.
How should human review and confidence routing work?
Send results the system is unsure about to a person, and spot-check the ones it is sure about. Then measure the review workload, because it is part of the real cost.
- Check whether confidence scores mean anything. On the test set, compare the system's confidence with whether it was actually right. If high-confidence results are wrong as often as low-confidence ones, the score cannot be used for routing.
- Set thresholds per field or category, using the test set. A supplier name and a payment amount do not deserve the same threshold.
- Route by risk as well as confidence. Large amounts, new suppliers, legal topics or irreversible actions can always go to review, whatever the confidence.
- Sample automatically approved results on a regular schedule, so silent errors are caught.
- Measure the review queue: the share of items reviewed, the time per review, and how often reviewers change the AI's output.
Illustrative example
In an invoice pilot, the team finds on the test set that totals and dates are reliable at high confidence, but supplier tax numbers are often misread on poor scans. They auto-accept totals and dates above the agreed threshold, always send tax numbers to review for new suppliers, and run a weekly sample of auto-accepted invoices. The go/no-go decision then weighs accuracy and the review workload together.
How do you test how an AI system fails?
Run a planned set of failure tests alongside the accuracy tests. The OWASP Top 10 for LLM Applications (2025 edition, part of the OWASP GenAI Security Project) is a practical checklist; it lists risks including prompt injection, sensitive information disclosure, excessive agency, misinformation and unbounded consumption. For agents, OWASP has also published a Top 10 for Agentic Applications (2026 edition).
- Direct prompt injection: users instructing the system to ignore its rules or reveal its instructions.
- Indirect prompt injection: instructions hidden inside documents, emails or web pages the system reads.
- Out-of-scope questions: requests the system should decline or redirect.
- Unanswerable questions: questions whose answer is not in the sources, where a confident answer is a failure.
- Sensitive data: asking for another employee's salary, a customer's personal details or restricted documents.
- Excessive agency: an agent attempting actions beyond its permissions, or chaining tools in unintended ways.
- Messy inputs: rotated scans, mixed languages, very long documents, empty fields.
- Logging: check that prompts, outputs and logs do not store sensitive data longer or more widely than intended.
Record each test, the expected behaviour and the result. A single serious failure here can be a no-go even when accuracy is high. Our security and data protection page explains how we handle access, logging and data in these systems.
How do you measure cost and speed?
Measure cost per completed task and time to a usable result, under realistic volume, rather than quoting model prices in isolation.
Cost categories to capture:
- Model usage per task, including retries and any grading or checking calls
- Hosting and infrastructure, including search indexes and storage
- Human review time per task
- Integration upkeep and monitoring
- Re-evaluation effort each time something changes
Speed categories to capture:
- Time until the user sees the first part of a response, for interactive assistants
- End-to-end time to a finished result, for documents and agents
- The slowest typical cases, not only the average, since those drive complaints
- Behaviour under expected peak volume
What should happen after go-live?
Keep measuring. AI systems change when the model, the prompt, the source documents or the inputs change, and sometimes when nothing on your side changes at all because a provider updated the model.
- Re-run the full test set before any change to the model version, prompts, retrieval settings, tools or source data, and compare against the last approved result.
- Add real production failures to the test set, so the same mistake is tested forever.
- Track review rates, user corrections and feedback over time; a rise usually signals that inputs have shifted.
- Assign an owner for the system and its test set.
The NIST AI Risk Management Framework (AI RMF 1.0, published as NIST AI 100-1 in January 2023) organises this work into four functions: Govern, Map, Measure and Manage. Its Generative AI Profile (NIST AI 600-1, July 2024) adds risks and actions specific to generative AI. NIST has said the framework is being revised, so check for the current edition. Both are voluntary frameworks, useful for structuring evaluation and governance rather than a certification.
Go/no-go scorecard
Fill in the threshold column before the pilot and the result column after it.
| Criterion | Threshold agreed in advance | Result | Pass? |
|---|---|---|---|
| Primary metric on the held-back test set (per field or category) | |||
| Performance on hard and rare cases | |||
| Comparison with the current process baseline | |||
| Share of work routed to human review | |||
| Reviewer correction rate | |||
| Failure tests: injection, out-of-scope, sensitive data | No serious failures | ||
| Cost per completed task | |||
| Response time, including slowest typical cases | |||
| Monitoring, re-evaluation process and owner in place | Yes | ||
| Sign-off from the business decision owner | Yes |
A "no-go" is a useful result. It might mean narrowing the scope, adding human review, fixing source data, or deciding that rules or conventional software suit the task better.
When is a full evaluation not worth it, or not the right approach?
- The task is deterministic. If clear rules produce the right answer, a rules engine or ordinary workflow automation is cheaper, faster and easier to audit than an AI model.
- There is no ground truth. If experts cannot agree on the right answer, fix the task definition before measuring the system.
- Volume is very low. A handful of cases a month may not justify the build, the evaluation or the monitoring.
- The stakes are trivial. An internal drafting aid that a person always edits needs lighter testing than an agent that updates records.
AI pilot evaluation checklist
- Task, inputs and out-of-scope cases written down
- Business decision owner named
- Current process baseline measured on the same cases
- Acceptance thresholds agreed before results are seen
- Test set sampled from real work, including hard and rare cases
- Labels by domain experts; disagreements resolved
- Development and test sets kept separate; test set versioned
- Metrics chosen per task type and per field or category
- Confidence scores checked against actual correctness
- Review thresholds and risk-based routing defined
- Failure tests run and recorded, using the OWASP lists as a guide
- Cost per task and response times measured at realistic volume
- Monitoring, re-evaluation triggers and an owner agreed
- Scorecard completed and signed
Evaluating AI with Timeline Digital
We build evaluation into AI projects from the first week: the test set and acceptance criteria come before the prompts. Our AI development page explains how we work, with detail on document processing and AI agent development. Every engagement starts with a free pilot of 2 to 3 key modules before the full project, and the evaluation approach above is how that pilot is judged.