OpenRouter's Ori Eval is the most credible self-serve agent eval we have seen ship this year, and it valid…
LLM-as-a-judge is reliable, but only when you control for its known failure modes. Used as a raw scorer it inherits three documented biases…