Keep arithmetic in code. Give judgment to the right model.
Exact bounds, parsing, latency, token, and cost checks stay deterministic. Jev handles atomic semantic decisions. LLMs add critique only when uncertainty warrants the extra cost.
Evaluation infrastructure, made inspectable
Combine exact checks, Jev's calibrated decisions, structured LLM judges, and human review in one local-first workbench.
Candidate: “Refunds are available within 30 days.”
Exact bounds, parsing, latency, token, and cost checks stay deterministic. Jev handles atomic semantic decisions. LLMs add critique only when uncertainty warrants the extra cost.
Accept clear outcomes, escalate uncertain ones, and always route release-critical cases to a human.
Datasets, rubrics, and candidates remain reproducible.
Inspect probability, disagreement, stability, cost, and regressions.
Start with evidence