AI engineering
Evals before launch: how we take an LLM feature to production
A demo that works once is not a feature. How we build eval sets, regression gates and monitoring so an AI capability behaves the same way on day 90 as it did on day one.
6 min read