Building Test Sets & Evaluation Datasets
Create a structured evaluation dataset for your agent. Define input scenarios, expected outputs, scoring criteria, and edge cases. Export as JSON for automated evaluation.
Test Set Builder
Add test cases one by one. Each case has an input scenario, expected output, category, and difficulty level. When you are done, preview and export the full dataset.
Scoring Criteria Designer
Design a scoring rubric for evaluating agent outputs. Toggle dimensions on or off, adjust weights, and preview the rubric configuration.
Key insight: A good rubric separates subjective "feels right" evaluation into measurable dimensions. Start with accuracy and safety, then add dimensions specific to your use case. Weights let you prioritise what matters most.