Notes
What I'm learning about agents, evals, and the platforms they run on.
-
Can I Build an AI Agent That Doesn't Write Slop?
I tested prompt engineering and model fine-tuning to see if an AI could act as a faithful copy editor. It's a step in the right direction, but nothing replaces good human judgment.
-
Fine-Tuning Was the Easy Part
Fine-tuning a small model for one narrow job worked in an afternoon. The hard developer-platform problem is distribution: getting that fix past a single adapter, across every job your developers do, and into the models they pick next.
-
A Model Router Needs a Scoreboard
A routing policy is a hypothesis until the same held-out tasks preserve quality and measure retries, latency, tokens, and cost across capability profiles.
-
Builder Platforms Grow by Owning the Agent Loop
Give coding agents a tested path into your platform, measure what works, and use the results to improve activation and retention.
-
Loop Engineering Coding Agent
A public operating contract with four role overlays, a structural check, and 17 specified scenarios. It is not a behavioral benchmark yet.
-
DevX Is a Growth Function
The growth loop starts when DevX owns repeated developer friction, ships the fix, distributes the better path, and measures whether behavior changed.
Elsewhere
-
Curated context beats raw model knowledge for building maps ↗
I built maps with AI to test where current models fail. Here's what worked, what broke, and why curated context beats raw model knowledge.
-
Native GeoJSON in BigQuery means no more transformation pipelines ↗
Query massive map datasets immediately. Here's how to skip complex transformation pipelines and process map data directly using BigQuery's native GeoJSON support.
Email list
Get new field notes by email
One email when something ships. One-click unsubscribe.
The address is stored only to deliver these updates. See Privacy.