
Judges
Judges in Evals: Flip Your Intuition
First-principles responses to common objections about using LLMs to judge LLMs.
If you are like most developers, your first instinct may be to reject the idea of using non-deterministic approaches in settings where reliability counts. This is especially true in AI reliability itself: using a model to judge the results of another model feels like fighting fire with fire.
This typically comes from a handful of credible doubts. Let us combat these concerns from first principles.

SutroTurn expert judgment into production-grade AI evals.Sutro provides infrastructure for expert annotation, optimization, and measurement.See how Sutro works
