Getting AI agents from demo to dependable.
A working demo takes an afternoon. Something people can rely on every day is a different problem. I write about that gap, from systems I have actually shipped.
What I write about
- Agent design that survives real inputs, not just the happy path.
- Evaluating AI output when there is no single correct answer.
- Cost and latency decisions, and what they trade away.
- The failure modes that only appear once something is in production.