Working Surfacelink · 22 Sep · 03
03
A small robot at the end of a production line holds each finished part against a cut metal template and drops the misfits into a separate bin, while a figure seen from behind, the back of the head showing, sits beside the bin and turns one rejected part over in both hands, the whole scene drawn with wide empty margins on every side.
Warp

Using LLM-as-a-judge scoring to measure your software factory

Zach Lloyd · 18 September 2026

How-to · critique & review · artifacts · Evals before the build

Read the original
Takeaway

Warp's chief executive explains how AI agents can grade a sample of coding agents' past work against written criteria, so that people review the failures rather than every change.

Summary

Warp sells a product for running many AI coding agents in the cloud, which it calls a software factory. Its chief executive, Zach Lloyd, argues that teams using coding agents should stop guessing how well the agents perform and grade their work directly. He describes how to set up grading agents, using Warp's own internal factory as the example.

Key points
  • Each grading agent reads the full record of a past coding session and returns a pass or a fail. A written prompt defines what it checks, such as whether the task was done, whether the work was efficient and whether the code was good.
  • Teams add criteria of their own. Warp checks whether its agents write redundant tests, a failure it saw often.
  • Only a sample of sessions is graded, because grading costs money. At Warp it takes about 3 percent of what the company spends on AI models.
  • People open the failed sessions, read what the coding agent did and why the grader failed it, and adjust the agents' instructions. Lloyd describes a further step in which another agent proposes those changes automatically.
Implication

A team that writes down what good agent work looks like can have a sample checked automatically and spend its review time on the failures; this is an inference. The post is Warp's account of its own product, and it does not say who in a team should write the criteria.

Source issue

22 September 2026: when agents write code faster than people can check it, teams move the checking into written rules and narrower reviews.