Reference test set
Real questions with reference answers written by your subject-matter experts, plus edge cases.
Tests that measure assistant and agent quality before and after launch.
Download the PDFTest results for every release of your generative AI systems.
The problem it solves, what is included, how it works, the technical components, and how we adapt it with you.
Your download has started.
We also emailed the link to . It stays valid for 7 days.
Download didn't start? Get the PDF
Want to see how it would fit your data? Talk to us.
Assistants change constantly: new prompts, new models, new retrieval settings. Without a test set, quality is judged by a few people trying a few questions, and regressions reach users first.
GenAI Evaluation Tests score every change against questions your experts wrote, and hold the release when results fall short.
Four parts, each adapted to your data, platforms, and controls.
Real questions with reference answers written by your subject-matter experts, plus edge cases.
Accuracy, use of sources, refusals, leaks, personal data, safety, latency, and cost for every answer.
A built-in adversarial suite, extended with cases for your own tools, data, and known risks.
Tests run on every change to prompts, retrieval, or models, and block the merge on failure.
You keep the test set, the gates, the pipeline, and a report for every release.
Start from real questions in logs, tickets, or a pilot, with personal data removed.
Accuracy, grounding, refusal, permissions, injection, and safety, with critical cases marked.
Unique strings in restricted documents and the system prompt. Any answer containing one is a leak.
Pass rates overall and by category, cost and latency limits, and maximum drop against the last release.
The pipeline scores the change, publishes the report, and blocks the merge if a gate fails.
Bad answers seen in production go into the set before they are fixed.
Vendor-neutral Python and configuration, Azure first, with tests included from the start.
Thresholds are set in the Ground step and recorded in the system’s Evaluation Card.
Any critical case failing, or any case erroring, also fails the run. The report lists every case that passed last time and fails now.
Have a question about the GenAI Evaluation Tests?
Get an answer from our pages in seconds, with links to the sources.
Share the overview with your team, or tell us the decision you want to improve and we will tell you whether it fits.