Working Surfacelink · 28 Sep · 10
10
A clockmaker seen from behind, the back of the head showing, sets a brass gauge with cut notches on a workbench, while a small robot files a gear and tries it against the gauge again and again, a small pile of rejected gears growing beside it, the whole scene drawn with wide empty margins on every side.
BLOG@CACM

Nobody Did TDD for 25 Years. Now the Machine Requires It

Abtin Aghagolian · 28 September 2026

How-to · agents · critique & review · Evals before the build

Read the original
Takeaway

Abtin Aghagolian, chief technology officer of Pikd, a London augmented-reality company, argues that tests written before an AI agent writes code are now how engineers control its output.

Summary

For 25 years, most programmers skipped test-driven development, the practice of writing a failing test before the code that makes it pass. Abtin Aghagolian, chief technology officer of Pikd, a London company that runs an augmented-reality platform, argues that AI agents now make the practice necessary. An engineer writes the test, the agent writes code until it passes, and the engineer reviews the test instead of the code.

Key points
  • He argues that a prompt is a description that only a person can check. A test is a definition that an AI agent can run and check its own work against.
  • An AI agent will pass a weak test without being correct. He therefore prefers tests of rules that must always hold over tests that pin one input to one output.
  • He wants teams to stop asking an AI model to write tests for code that already exists, because such tests confirm whatever the code does, errors included.
  • Pikd's own assistant passed all 1,053 of its tests while telling a user that no such tool was available and carrying out the correct navigation action in the same reply. He knows of no way to test for every such contradiction in advance.
Implication

An engineering team that hands implementation to AI agents can move its review effort to the tests written beforehand, and should expect gaps where two correct outputs contradict each other. The application to other teams is an inference and is not the author's claim.

Suggested actions
Engineering

Before an AI agent writes code, write the tests that define a correct result, let the agent work until they pass, and review the tests instead of the code.

Engineering

Do not count tests that an AI model wrote for existing code as verification, because they record what the code already does, errors included.

Derived by Working Surface from the article. Source line: The new one is decide what ‘correct’ means, write that down in executable form, and let the machine iterate against it until it goes green.

Source issue

28 September 2026: when agents build, people keep the first screen, the review, the files agents read and the whole product.