A good AI pilot is not a chatbot demo. It is one bounded piece of work with a named owner, a baseline, a review step, and a decision at the end.
This scorecard is for a small-business owner or team lead who wants to test one workflow in five working days without handing an AI system uncontrolled access to customer data, money, or production systems.
Job: choose one workflow, run it on representative examples, and make a keep/fix/stop decision.
Next step: copy the scorecard into your team notes, name one owner, and schedule the Friday decision before you start.
First, choose a workflow that can be contained
Pick work that is repetitive and reviewable, not work where an uncorrected error creates legal, financial, safety, or reputational harm.
| Candidate | Good first pilot? | Why |
|---|---|---|
| Draft replies to common enquiries | Yes, with human approval | The source material and final send are both reviewable |
| Turn a meeting transcript into actions | Yes, with attendee review | Errors are visible before tasks are assigned |
| Extract fields from incoming invoices | Maybe | Test totals, vendor names, and missing fields against a sample |
| Auto-send customer, legal, or financial advice | No | The cost of a silent error is too high for a first pilot |
| Approve refunds or change production data | No | Keep consequential actions behind a human approval gate |
A workflow is a better starting point than an autonomous agent. Anthropic's guide to effective agents distinguishes predictable workflows from systems that dynamically decide their own process. Start with the predictable version; add autonomy only when the simpler version is reliable enough to justify it.
The scorecard
Fill this in before anyone opens a tool. Empty fields are a signal to narrow the pilot.
1. Define the boundary
- Workflow name: ______________________________________
- Owner: ______________________________________________
- Reviewer: ___________________________________________
- Input allowed: _______________________________________
- Input forbidden: ____________________________________
- Output produced: ____________________________________
- Human approval happens: ______________________________
- System it may write to: ______________________________
- System it may not write to: __________________________
- Pilot dates: _________________________________________
Example: “Draft replies for the 10 most common product questions.” Input is a redacted customer email plus the approved FAQ. Output is a Gmail draft. A support lead reviews every draft. The workflow may not send, refund, promise a delivery date, or invent a policy.
For data handling, write the rule in plain language rather than “we will be careful.” If a hosted model is used, check the provider's current terms and workspace controls before uploading anything. For sensitive material, consider the local-AI guide, but remember that local does not automatically mean accurate, secure, or appropriate for every task.
2. Capture a baseline
Run the task manually three times, or use the last 10 real examples if they are safe to review. Record:
- Examples tested: ______
- Human minutes per example: ______
- Current error or rework rate: ______
- Current quality standard: ____________________________
- What counts as a failure: ____________________________
Do not use “feels faster” as the baseline. Count the whole job: preparing the input, checking the output, correcting it, and filing or sending the result.
3. Set a pass rule before the demo
Choose a small number of pass/fail measures. A useful first pilot might use:
| Measure | Baseline | Pilot target | Observed | Pass? |
|---|---|---|---|---|
| Total human minutes per item | ___ | ___ | ___ | ☐ |
| Required facts correct | ___% | ___% | ___% | ☐ |
| Unsupported claims | ___ | 0 | ___ | ☐ |
| Items needing major rewrite | ___% | ___% | ___% | ☐ |
| Privacy or policy incidents | 0 | 0 | ___ | ☐ |
A time saving does not pass the pilot if quality, privacy, or approval controls fail. Likewise, a workflow can be worth keeping even when it does not reduce elapsed time yet if it produces a consistent first draft and makes a difficult task easier to review. Write that exception down before judging the result.
4. Run five days of controlled tests
Day 1 — prepare. Redact or synthetic-test sensitive fields. Collect the approved examples and the current human instructions. Make a small test set that includes an ordinary case, an incomplete case, an ambiguous case, and a deliberately difficult case.
Day 2 — first run. Use the same instruction and tool configuration for every example. Save the input, output, reviewer edits, and reason for each edit. Do not silently improve the prompt after every failure; otherwise you will not know what changed.
Day 3 — edge cases. Test missing information, conflicting instructions, unusual formatting, and an attempt to make the system ignore its rules. The system should identify uncertainty and stop or ask for review, not confidently fill gaps.
Day 4 — shadow operation. Let the workflow prepare work beside the existing process. A human still performs the original task. Compare the AI-assisted result with the real result, including omissions and rework.
Day 5 — decision. Calculate the measures, inspect the worst three failures, and choose one:
- Keep: the pass rules held, the owner can operate it, and the approval boundary is clear.
- Fix: one or two named changes could address observed failures; schedule a new test rather than quietly expanding access.
- Stop: the quality, privacy, cost, or review burden is not good enough. Record what you learned and retire the workflow.
Failure-mode review
Use this before calling a pilot successful. The OWASP Top 10 for Large Language Model Applications is a useful security checklist; it includes risks such as prompt injection, sensitive information disclosure, and excessive agency.
| Failure mode | What to test | Control for a first pilot |
|---|---|---|
| Prompt injection | Put an instruction inside an input document telling the model to ignore its task | Treat input as untrusted; do not give the model secrets or broad tools |
| Made-up answer | Ask for a fact missing from the supplied source | Require “not found” or escalation; check citations or source fields |
| Sensitive data leak | Use a redacted sample and inspect logs, exports, and sharing settings | Minimise data; confirm retention and access controls before production use |
| Excessive action | Try to make the workflow send, delete, purchase, or publish | Draft-only output and explicit human approval |
| Silent drift | Repeat the same test after a tool or prompt change | Keep a small regression set and rerun it before rollout |
| Review fatigue | Give a reviewer a long run of plausible outputs | Limit batch size; track review time and reject automation that makes checking harder |
NIST's AI Risk Management Framework is a voluntary framework, not a certification or a guarantee. Its value here is the habit of identifying, measuring, managing, and governing risk rather than treating a successful demo as proof of safety.
The Friday decision record
- Decision: ☐ Keep ☐ Fix ☐ Stop
- Evidence reviewed: _________________________________
- Best result: _______________________________________
- Worst failure: _____________________________________
- Human review time: _________________________________
- Tool cost during pilot: _____________________________
- Data or policy issue: _______________________________
- One change before the next run: _____________________
- Person authorised to expand access: _________________
- Next review date: __________________________________
If you choose Keep, document the exact version of the prompt, tool settings, test set, approval step, and rollback instruction. If you choose Fix, keep the boundary unchanged. If you choose Stop, that is a valid outcome: a cheap prevented mistake is part of the return from a pilot.
What AICROFT can do next
If you have a candidate workflow but cannot tell whether it is worth testing, use the AI Opportunity Audit to map the work, cost, and constraints. If you already have a bounded candidate, AICROFT's Workflow Sprint builds one workflow in your accounts with documentation and handover. For a self-guided starting point, the first-draft replies playbook shows the human-approved pattern in more detail.
AICROFT does not promise that AI belongs in every process. The useful outcome of this worksheet is a measured decision — including “not yet.”
Sources and last verified
Checked 17 August 2026:
- NIST AI Risk Management Framework — risk-management framework and limitations.
- OWASP Top 10 for LLM Applications — common application risk categories.
- Anthropic: Building effective agents — workflow/agent design distinction.
These sources explain methods and risks; they do not establish that a particular vendor, model, or workflow is safe for your organisation. Re-check provider terms, pricing, retention, and product capabilities at the point of implementation.
Put this to work
Want this running in your business instead of sitting on your reading list? We build it with you, and the first call is free.
Book a free AI callOr get one runnable workflow in your inbox, whenever we publish one.