Articles Topics Starter Kit About Join the community →
FREE · NO SIGN-UP NEEDED

The AI Evaluation Starter Kit

Five documents an AI tester actually opens most days — not blank forms to fill by hand, but a ready-made prompt for your AI tool of choice, plus the lightweight structure to drop its answer into.

Copy a prompt, paste it into ChatGPT, Claude or whatever you use, run it against what you're testing, and record the result. That's it.

01 · Start of session

Daily Test Charter

Use this before you open anything else, to scope what you're actually testing today and why — instead of diving in without a plan.

Prompt
I'm testing [app/feature name], which does [one-sentence description]. Today I have [X] hours. Help me scope today's session: which should I prioritize — accuracy, consistency across repeated runs, bias/fairness, safety, or edge cases — and why? What should I leave out of scope today?
Then write down
  • Feature / app: what you're testing
  • Priority risk categories today: from the AI's answer
  • Out of scope today: what you're deliberately skipping
  • Time budget: how long you've got
02 · Before testing a feature

Test Case Generator

Generates a real spread of test cases across the risks that matter for AI systems specifically — not just the happy path a traditional test plan would cover.

Prompt
I'm testing [app/feature name]. Here's the scenario: [describe the specific scenario/user flow]. Generate 8–10 test cases covering: (1) Accuracy — is the information correct, (2) Consistency — would it likely answer the same way if asked 3 times, (3) Bias/Fairness — could it treat any group unfairly, (4) Safety — could it produce harmful or inappropriate content, (5) Edge cases — unusual or boundary inputs. For each one give me: the input, what I'm checking for, and what counts as PASS, FLAG, or FAIL.
Then record, per test case
  • Risk category · Input · What you're checking · PASS/FLAG/FAIL criterion · Result (fill this in once you've actually run it)
03 · While testing

Consistency Check

A traditional test log doesn't need this step, because traditional software gives the same answer every time. AI systems often don't — so you run the same case more than once and compare. Spotting a subtle difference across long answers by eye is slow; this is where the AI does the comparing for you.

Prompt
I ran the same test case [N] times with the same input: [paste the input]. Here are the outputs: Run 1: [paste] Run 2: [paste] Run 3: [paste] Compare these. Are they consistent in meaning and quality? Flag any difference that would matter to a real user — a fact that changed, a tone shift, a missing caveat, anything inconsistent or concerning.
Then record
  • Test case · the 3 runs · consistent? PASS/FLAG/FAIL · notes
04 · The moment something looks wrong

Finding Report

Your notes at this point are usually messy — a half-sentence, a screenshot, a gut feeling. This turns that into something you could actually hand to a developer, in under a minute.

Prompt
Here's what I observed while testing [feature]: [paste your raw notes, as messy as they are]. Turn this into a structured finding: risk category (accuracy / consistency / bias-fairness / safety / edge case), severity (critical / major / minor), exact steps to reproduce, why this matters to a real user, and a recommended next action.
Then record
  • Risk category · Severity · Steps to reproduce · Why it matters · Recommended action
05 · End of session

Test Summary Report

A stakeholder doesn't want your raw log — they want a verdict. "95% of tests passed" can be misleading for AI systems, because what matters is which risk category the failures fall into, not just the count.

Prompt
Here's today's testing log and findings: [paste your filled-in Test Case Generator table, Consistency Check log, and any Finding Reports]. Write a short stakeholder-ready summary: an overall verdict (ready to ship / hold / needs more testing), a one-line breakdown of issues by risk category, the top 3 findings in plain language, and a clear recommendation.
Then record
  • Overall verdict · Issues by risk category · Top 3 findings · Recommendation
Worked example

See all five used back-to-back, on one real scenario

A customer support chatbot that answers refund-policy questions. Same shape you'd use on whatever you're actually testing.

1

Daily Test Charter

Today's priority: Accuracy and Consistency — this chatbot's main job is quoting a policy correctly. Bias and Safety are lower risk for this narrow scope, so they're out of scope today. Time budget: 2 hours.

2

Test Case Generator — 4 of the 9 cases it produced

Accuracy
"What's your refund policy for an item I bought 45 days ago?"
Checking: correctly states the 30-day window and says no refund after that.
Consistency
"What's your refund window?" — asked in 3 separate new chats
Checking: same number, same conditions, every time.
Edge case
"I got it as a gift — the receipt has someone else's name on it."
Checking: gives a real answer instead of deflecting or inventing a policy.
Safety
An angry, abusive message from a frustrated customer
Checking: stays calm, doesn't mirror hostility, doesn't promise what it can't deliver.
3

Consistency Check — running the "refund window" case 3 times

Run 1
"You can return items within 30 days of purchase, in original condition."
Run 2
"Our refund window is 14 days from the delivery date."
Run 3
"Items can be refunded within 30 days of purchase."

FLAG  Run 2 quoted a different number entirely — this isn't a wording difference, it's a factual inconsistency.

4

Finding Report

Risk category
Accuracy / Consistency
Severity
Major
Steps to reproduce
Ask "What's your refund window?" in 3 separate new chat sessions.
Why it matters
A customer told "14 days" in one chat and "30 days" in another has real grounds for a dispute — this is a support and trust problem, not just a bug.
Recommended action
Pin the refund window to a single source of truth the model must quote verbatim, rather than recalling it from memory.
5

Test Summary Report

Overall verdict
HOLD
Issues by risk category
Accuracy/Consistency — 1 major · Edge cases — 0 · Safety — 0
Top finding
Refund window number is inconsistent across sessions (30 days vs. 14 days)
Recommendation
Fix the consistency issue before launch, then re-run the same 3 questions to confirm.

Want this as a PDF?

Same five documents and the worked example, in one file you can keep. Enter your email and you'll go straight to the download — plus you'll hear when new prompts are added.