01 25 or 45 min · talk
Evals as CI: making stochastic systems boring
Why offline benchmarks lie, how to resample golden sets from production, and how to gate deploys on evals so model changes ship weekly instead of quarterly. Drawn from running LLM features on a platform handling 1M+ calls a day.