Using LLM-as-a-judge scoring to measure your software factory
Zach Lloyd · 18 September 2026
Read the originalWarp's chief executive explains how AI agents can grade a sample of coding agents' past work against written criteria, so that people review the failures rather than every change.
Warp sells a product for running many AI coding agents in the cloud, which it calls a software factory. Its chief executive, Zach Lloyd, argues that teams using coding agents should stop guessing how well the agents perform and grade their work directly. He describes how to set up grading agents, using Warp's own internal factory as the example.
- Each grading agent reads the full record of a past coding session and returns a pass or a fail. A written prompt defines what it checks, such as whether the task was done, whether the work was efficient and whether the code was good.
- Teams add criteria of their own. Warp checks whether its agents write redundant tests, a failure it saw often.
- Only a sample of sessions is graded, because grading costs money. At Warp it takes about 3 percent of what the company spends on AI models.
- People open the failed sessions, read what the coding agent did and why the grader failed it, and adjust the agents' instructions. Lloyd describes a further step in which another agent proposes those changes automatically.
A team that writes down what good agent work looks like can have a sample checked automatically and spend its review time on the failures; this is an inference. The post is Warp's account of its own product, and it does not say who in a team should write the criteria.
22 September 2026: when agents write code faster than people can check it, teams move the checking into written rules and narrower reviews.