7.2 Module 7 · Evaluation & Observability

Building Test Sets & Evaluation Datasets

Create a structured evaluation dataset for your agent. Define input scenarios, expected outputs, scoring criteria, and edge cases. Export as JSON for automated evaluation.

Test Set Builder Scoring Criteria Designer

Test Set Builder

Add test cases one by one. Each case has an input scenario, expected output, category, and difficulty level. When you are done, preview and export the full dataset.

Quick add:
Test Cases (0)
No test cases yet. Add one above or use the quick-add buttons.

Scoring Criteria Designer

Design a scoring rubric for evaluating agent outputs. Toggle dimensions on or off, adjust weights, and preview the rubric configuration.

Rubric Preview

Key insight: A good rubric separates subjective "feels right" evaluation into measurable dimensions. Start with accuracy and safety, then add dimensions specific to your use case. Weights let you prioritise what matters most.