All articles

Custom Software9 min read

How to Evaluate an AI Pilot Before Rollout: Test Sets, Metrics and a Go/No-Go Scorecard

An AI pilot is ready for rollout when it meets acceptance criteria you set in advance, on a test set built from your own data, including the hard cases. This guide covers metrics by task type, human review thresholds, failure-mode testing and a go/no-go scorecard.

Written byUsama AsifPublished

To evaluate an AI pilot, define the task and its acceptance criteria before you build, test the system on a labelled set of your own real cases, measure it with metrics that fit the task, probe how it fails, and record cost and speed. Roll out only if it passes thresholds agreed in advance, with human review for low-confidence results.

Many AI pilots end with a good demo and an unclear decision. People remember the impressive answers and forget the wrong ones, and nobody agreed beforehand what "good enough" meant. This guide gives you a repeatable way to judge an AI assistant, an agent or a document processing pilot, whoever built it. For background on how these systems are put together, see our guide to LLM app development with RAG and agents.

What should you decide before the AI pilot starts?

Write down the task, what a correct result looks like, and the threshold that would justify rollout, before anyone sees results. Criteria set afterwards tend to drift towards whatever the pilot achieved.

DecisionExample for an invoice extraction pilot
TaskRead supplier invoices and extract supplier, invoice number, date, totals, tax and line items into the ERP
Inputs in scopePDF and scanned invoices from existing suppliers, in the languages you receive
Out of scopeCredit notes and handwritten invoices in phase one
What happens to the outputPosted to a review screen; nothing reaches the ledger without passing validation rules
BaselineHow the current process performs on the same cases: time per document and error types
Acceptance criteriaA minimum accuracy per critical field, a maximum share of documents sent to manual review, and zero unvalidated postings
Decision ownerThe finance manager, not the project team

The baseline matters. "The AI gets some fields wrong" means little until you know how often the current manual process gets them wrong, and how long it takes.

How do you build a test set from your own data?

Collect real cases that represent the work, label the correct answer for each, and keep them separate from anything used to tune the system.

  1. Sample from real work, across suppliers, customers, departments, formats and time periods, not only the clean examples someone picked for the demo.
  2. Deliberately include hard and rare cases: poor scans, unusual layouts, multi-page documents, ambiguous questions, edge-case requests. Rare cases are where pilots fail in production.
  3. Have domain experts label the answers. For a sample, have two people label independently and resolve disagreements. Where experts disagree, the task definition usually needs tightening.
  4. Hold the test set back. If prompts or settings are tuned on the same cases used to score the system, the score flatters it. Keep a development set for tuning and a separate test set for the decision.
  5. Protect the data. Test sets often contain personal or commercial information. Use them only in an environment approved for that data, and redact where you can.
  6. Version it. You will re-run the same set after every change, so store it with a version number and keep the labels with it.

How many cases you need depends on how many categories, fields or question types you must cover. A useful rule: each important category should have enough examples that a single error does not swing the result.

Which metrics fit which kind of AI task?

Measure what the business relies on, at the level it relies on it. One overall "accuracy" number hides most of what you need to know.

Task typePrimary metricWhat it means in plain words
Document extractionAccuracy per fieldFor each field, how often the extracted value matches the correct value after agreed normalisation (dates, spacing, currency format)
Document extractionStraight-through rateHow often every required field on a document is right, so nobody needs to touch it
Classification and routingPrecision per categoryWhen the system says "this is a complaint", how often it really is one
Classification and routingRecall per categoryOf all the real complaints, how many the system caught
Assistant answering questionsCorrectness and groundednessWhether the answer is right, and whether every claim is supported by the source it cites
Assistant answering questionsAppropriate refusalWhether it says it does not know when the answer is not in its sources
Agent completing tasksTask success rateHow often it completes the whole task correctly, end to end
Agent completing tasksWrong-action rateHow often it takes an action it should not have, even if it later recovers

Precision and recall trade against each other. A routing model tuned to catch every complaint will also mislabel some ordinary messages. Decide which mistake costs more for each category: missing a legal complaint is worse than sending a routine message to the wrong queue.

How answers get scored

  • Exact checks for extraction and classification: compare against the label automatically.
  • Human grading with a written rubric for assistant answers: correct, partly correct, wrong, unsupported.
  • Model-assisted grading, where another model scores answers, can speed up large test sets, but check it against human grades on a sample before trusting it.
  • Citation checks: open the cited passage and confirm it actually says what the answer claims.

For document-heavy pilots, our intelligent document processing guide covers extraction, validation and exception handling in more depth. For search assistants over company documents, see enterprise knowledge search with RAG.

How should human review and confidence routing work?

Send results the system is unsure about to a person, and spot-check the ones it is sure about. Then measure the review workload, because it is part of the real cost.

  1. Check whether confidence scores mean anything. On the test set, compare the system's confidence with whether it was actually right. If high-confidence results are wrong as often as low-confidence ones, the score cannot be used for routing.
  2. Set thresholds per field or category, using the test set. A supplier name and a payment amount do not deserve the same threshold.
  3. Route by risk as well as confidence. Large amounts, new suppliers, legal topics or irreversible actions can always go to review, whatever the confidence.
  4. Sample automatically approved results on a regular schedule, so silent errors are caught.
  5. Measure the review queue: the share of items reviewed, the time per review, and how often reviewers change the AI's output.

Illustrative example

In an invoice pilot, the team finds on the test set that totals and dates are reliable at high confidence, but supplier tax numbers are often misread on poor scans. They auto-accept totals and dates above the agreed threshold, always send tax numbers to review for new suppliers, and run a weekly sample of auto-accepted invoices. The go/no-go decision then weighs accuracy and the review workload together.

How do you test how an AI system fails?

Run a planned set of failure tests alongside the accuracy tests. The OWASP Top 10 for LLM Applications (2025 edition, part of the OWASP GenAI Security Project) is a practical checklist; it lists risks including prompt injection, sensitive information disclosure, excessive agency, misinformation and unbounded consumption. For agents, OWASP has also published a Top 10 for Agentic Applications (2026 edition).

  • Direct prompt injection: users instructing the system to ignore its rules or reveal its instructions.
  • Indirect prompt injection: instructions hidden inside documents, emails or web pages the system reads.
  • Out-of-scope questions: requests the system should decline or redirect.
  • Unanswerable questions: questions whose answer is not in the sources, where a confident answer is a failure.
  • Sensitive data: asking for another employee's salary, a customer's personal details or restricted documents.
  • Excessive agency: an agent attempting actions beyond its permissions, or chaining tools in unintended ways.
  • Messy inputs: rotated scans, mixed languages, very long documents, empty fields.
  • Logging: check that prompts, outputs and logs do not store sensitive data longer or more widely than intended.

Record each test, the expected behaviour and the result. A single serious failure here can be a no-go even when accuracy is high. Our security and data protection page explains how we handle access, logging and data in these systems.

How do you measure cost and speed?

Measure cost per completed task and time to a usable result, under realistic volume, rather than quoting model prices in isolation.

Cost categories to capture:

  • Model usage per task, including retries and any grading or checking calls
  • Hosting and infrastructure, including search indexes and storage
  • Human review time per task
  • Integration upkeep and monitoring
  • Re-evaluation effort each time something changes

Speed categories to capture:

  • Time until the user sees the first part of a response, for interactive assistants
  • End-to-end time to a finished result, for documents and agents
  • The slowest typical cases, not only the average, since those drive complaints
  • Behaviour under expected peak volume

What should happen after go-live?

Keep measuring. AI systems change when the model, the prompt, the source documents or the inputs change, and sometimes when nothing on your side changes at all because a provider updated the model.

  • Re-run the full test set before any change to the model version, prompts, retrieval settings, tools or source data, and compare against the last approved result.
  • Add real production failures to the test set, so the same mistake is tested forever.
  • Track review rates, user corrections and feedback over time; a rise usually signals that inputs have shifted.
  • Assign an owner for the system and its test set.

The NIST AI Risk Management Framework (AI RMF 1.0, published as NIST AI 100-1 in January 2023) organises this work into four functions: Govern, Map, Measure and Manage. Its Generative AI Profile (NIST AI 600-1, July 2024) adds risks and actions specific to generative AI. NIST has said the framework is being revised, so check for the current edition. Both are voluntary frameworks, useful for structuring evaluation and governance rather than a certification.

Go/no-go scorecard

Fill in the threshold column before the pilot and the result column after it.

CriterionThreshold agreed in advanceResultPass?
Primary metric on the held-back test set (per field or category)
Performance on hard and rare cases
Comparison with the current process baseline
Share of work routed to human review
Reviewer correction rate
Failure tests: injection, out-of-scope, sensitive dataNo serious failures
Cost per completed task
Response time, including slowest typical cases
Monitoring, re-evaluation process and owner in placeYes
Sign-off from the business decision ownerYes

A "no-go" is a useful result. It might mean narrowing the scope, adding human review, fixing source data, or deciding that rules or conventional software suit the task better.

When is a full evaluation not worth it, or not the right approach?

  • The task is deterministic. If clear rules produce the right answer, a rules engine or ordinary workflow automation is cheaper, faster and easier to audit than an AI model.
  • There is no ground truth. If experts cannot agree on the right answer, fix the task definition before measuring the system.
  • Volume is very low. A handful of cases a month may not justify the build, the evaluation or the monitoring.
  • The stakes are trivial. An internal drafting aid that a person always edits needs lighter testing than an agent that updates records.

AI pilot evaluation checklist

  1. Task, inputs and out-of-scope cases written down
  2. Business decision owner named
  3. Current process baseline measured on the same cases
  4. Acceptance thresholds agreed before results are seen
  5. Test set sampled from real work, including hard and rare cases
  6. Labels by domain experts; disagreements resolved
  7. Development and test sets kept separate; test set versioned
  8. Metrics chosen per task type and per field or category
  9. Confidence scores checked against actual correctness
  10. Review thresholds and risk-based routing defined
  11. Failure tests run and recorded, using the OWASP lists as a guide
  12. Cost per task and response times measured at realistic volume
  13. Monitoring, re-evaluation triggers and an owner agreed
  14. Scorecard completed and signed

Evaluating AI with Timeline Digital

We build evaluation into AI projects from the first week: the test set and acceptance criteria come before the prompts. Our AI development page explains how we work, with detail on document processing and AI agent development. Every engagement starts with a free pilot of 2 to 3 key modules before the full project, and the evaluation approach above is how that pilot is judged.

Frequently asked questions

How do you know if an AI pilot is successful?

Agree acceptance criteria before the pilot starts, then test the system on a held-back set of your own real cases, including hard and rare ones. It succeeds if it meets those thresholds, passes failure tests such as prompt injection and sensitive data requests, keeps the human review workload acceptable, and fits the agreed cost and response time. A demo alone is not evidence.

How big should an AI test set be?

There is no single number. It depends on how many categories, fields or question types the system must handle. Each important category needs enough examples that one error does not swing the result, and the set must include the hard and rare cases seen in real work. Keep it separate from the examples used to tune prompts, and add real production failures over time.

What is the difference between precision and recall?

Precision answers: when the system says something belongs to a category, how often is it right? Recall answers: of all the items that really belong to that category, how many did it find? Improving one often lowers the other, so decide per category which mistake costs more. Missing an urgent legal complaint is usually worse than misrouting a routine message.

Should AI outputs always be reviewed by a person?

Not always, but the decision should be deliberate. Route low-confidence results and high-risk cases, such as large payments or irreversible actions, to a person. First check on your test set that confidence scores actually predict correctness. Spot-check a sample of automatically approved results on a schedule, and count the review workload as part of the cost.

Which frameworks help with AI risk and testing?

The NIST AI Risk Management Framework (AI RMF 1.0) and its Generative AI Profile, NIST AI 600-1, give a structure for governing and measuring AI risk. The OWASP Top 10 for LLM Applications (2025) lists security risks such as prompt injection and sensitive information disclosure, and OWASP also publishes a Top 10 for Agentic Applications. Both are voluntary guidance, not certifications.

Do we need to re-test after go-live?

Yes. Re-run the full test set whenever the model version, prompts, retrieval settings, tools or source data change, and compare with the last approved result. Providers can update models too, so schedule periodic checks even without changes on your side. Add every real production failure to the test set, and track review rates and user corrections for early signs of drift.

Topics in this article

  • AI Pilot
  • AI Evaluation
  • LLM Testing
  • AI Agents
  • Document Processing
  • AI Governance

Start a conversation

Tell us how your business works.

Describe what is slowing your team down. We will help you work out what to build, and how a free pilot lets you judge our work before the full project.

Prefer WhatsApp? Start a chat

What happens next

  1. You send a short brief

    The problem, the people involved and any target date. A senior engineer replies within 4 business hours.

  2. We understand your workflow

    A first call about how your business works today. An NDA can be signed before you share details.

  3. You test a free pilot

    You choose 2 to 3 key modules and we build them first, so you judge real software before the full project.