Working Surfacetopics
Topic · 7 links · since 17 September 2026

Evals before the build

Before an AI feature is built, the team writes down what a good output looks like, so that the result can be tested against it.

BLOG@CACM
Nobody Did TDD for 25 Years. Now the Machine Requires It
Abtin Aghagolian · 28 September 2026
Abtin Aghagolian, chief technology officer of Pikd, a London augmented-reality company, argues that tests written before an AI agent writes code are now how engineers control its output.
How-to · agents · critique & review
28 Sep · 10
The Pragmatic Engineer
Design Engineering with Maggie Appleton
Gergely Orosz with Maggie Appleton · 23 September 2026
Maggie Appleton, a research engineer at GitHub, no longer reads the code AI agents write for her prototypes and writes a spec saying how the agent will check its work.
Company story · prototyping
26 Sep · 05
Anthropic
How to prepare for AI-driven code modernization projects
Jonah Ezekiel and Lexie Tonelli · 23 September 2026
Two Anthropic engineers advise companies modernising old code with AI agents to agree, before the work starts, what each change must prove and how much human review it gets.
Company story · critique & review · artifacts
24 Sep · 01
Lenny's Newsletter
Advanced evals: How to find (and fix) hidden AI failures in your product
Hamel Husain and Shreya Shankar · 22 September 2026
Two specialists in testing AI products argue that people should read real AI sessions and note the failures themselves before an agent looks for errors or anyone writes a metric.
How-to · critique & review · rituals
23 Sep · 03
Warp
Using LLM-as-a-judge scoring to measure your software factory
Zach Lloyd · 18 September 2026
Warp's chief executive explains how AI agents can grade a sample of coding agents' past work against written criteria, so that people review the failures rather than every change.
How-to · critique & review · artifacts
22 Sep · 03
The Pragmatic Engineer
AI Skills with Matt Pocock
Gergely Orosz with Matt Pocock · 17 September 2026
Engineer and educator Matt Pocock uses short instruction files that make AI coding agents question him closely about a plan, which leaves him more time for strategic work.
Company story · agents · decisions
20 Sep · 02
Product Talk
4 New Evals and 16 Experiment Variants to Fix 1 Customer Complaint
Teresa Torres · 16 September 2026
Teresa Torres spent three weeks and sixteen experiments fixing one customer complaint about an AI product, starting by measuring how often the error occurred.
How-to · critique & review
17 Sep · 02