Evaluation Playground

Test one example before running a full dataset

01 INPUT

What should be evaluated?

Provide the input, candidate response, and reference evidence as JSON.

Leave blank to use the deployed workspace key. A pasted key stays in this browser tab and is never exported.

02 Questions

Atomic evaluators

01Noul
inspect: candidateOutput, reference
02Choice
inspect: candidateOutput, reference
03Score
inspect: candidateOutput, reference
04LLM judge
inspect: candidateOutput, reference

03 Results

Normalized evidence

Run the suite to compare calibrated decisions and structured critique.