Claim safety validation workflow from agent draft through eval gate to publish or human escalation

Claim safety: evidence before metrics

This is part 5 of the Agent production system series. Previous: From Xbox SLAs to agent fleets. Constellation nodes: Eval gates, Search, Human approval. Agents are fluent. Fluency is not the same as true. Claim safety is the discipline of keeping numbers, titles, and authorship tied to evidence — especially when a model would rather sound complete than sound correct. It shows up in eval gates, in human approval, and in public artifacts like this site and Resume Builder. ...

August 11, 2026 Â· Dave Voyles
Illustration of eval checkpoint arches with human-in-the-loop control

Eval gates are not optional theater

This is part 1 of the Agent production system series. Start with the system map if you haven’t read it yet. On the interactive diagram: Eval gates. Demo agents look smart until they touch a real repo. In a demo, the agent writes the code, the code runs once, everyone claps. Nobody checks what happens the second time, or the tenth, or the time the agent decides the fastest way past a failing test is to delete the test. Then you learn the hard lesson: intelligence without a gate is just a faster way to ship a bad change. ...

July 28, 2026 Â· Dave Voyles