Insights
Field notes from systems that are actually running
Written by the engineers who built and operate them, including the parts that did not work the first time.
- 2 articles
- Written by practitioners
Latest
Recent articles
-
How to Build an LLM Evaluation Set That Is Worth Trusting
The practical method — where the cases come from, how to adjudicate them, and how to validate a model-as-judge before you trust its numbers.
EvaluationJuly 16, 2026 -
Why AI Agents Fail in Production
Six failure modes that account for almost every agent project we have been called in to rescue, and what actually prevents each one.
AI agentsJuly 2, 2026
Next Step
Working on something these notes touch on?
A 30-minute technical call with an engineer who has shipped this before — not a sales qualification round. You leave with a feasibility read, a rough shape for the build, and an honest answer about whether it is worth doing at all.