Field notes · published irregularly, measured always
Writing
Notes from running AI systems in production: agent evals, context engineering, and what actually survives contact with real users.
01Posts
First posts are in the works. This is what they will be about:
Agent evals
How to know whether an agent actually works — task-level evaluation, not vibes.
Context engineering
Shaping what a model sees: context files, skills, and the tooling around them.
AI in production
What survives contact with real users, and what quietly breaks.
RSS will be live at /rss.xml from the first post.