Five documents an AI tester actually opens most days — not blank forms to fill by hand, but a ready-made prompt for your AI tool of choice, plus the lightweight structure to drop its answer into.
Copy a prompt, paste it into ChatGPT, Claude or whatever you use, run it against what you're testing, and record the result. That's it.
Use this before you open anything else, to scope what you're actually testing today and why — instead of diving in without a plan.
Generates a real spread of test cases across the risks that matter for AI systems specifically — not just the happy path a traditional test plan would cover.
A traditional test log doesn't need this step, because traditional software gives the same answer every time. AI systems often don't — so you run the same case more than once and compare. Spotting a subtle difference across long answers by eye is slow; this is where the AI does the comparing for you.
Your notes at this point are usually messy — a half-sentence, a screenshot, a gut feeling. This turns that into something you could actually hand to a developer, in under a minute.
A stakeholder doesn't want your raw log — they want a verdict. "95% of tests passed" can be misleading for AI systems, because what matters is which risk category the failures fall into, not just the count.
A customer support chatbot that answers refund-policy questions. Same shape you'd use on whatever you're actually testing.
Today's priority: Accuracy and Consistency — this chatbot's main job is quoting a policy correctly. Bias and Safety are lower risk for this narrow scope, so they're out of scope today. Time budget: 2 hours.
FLAG Run 2 quoted a different number entirely — this isn't a wording difference, it's a factual inconsistency.
Same five documents and the worked example, in one file you can keep. Enter your email and you'll go straight to the download — plus you'll hear when new prompts are added.