Field notes · published irregularly, measured always

Writing

Notes from running AI systems in production: agent evals, context engineering, and what actually survives contact with real users.

01Posts

First posts are in the works. This is what they will be about:

Agent evals

How to know whether an agent actually works — task-level evaluation, not vibes.

Context engineering

Shaping what a model sees: context files, skills, and the tooling around them.

AI in production

What survives contact with real users, and what quietly breaks.

RSS will be live at /rss.xml from the first post.