AI coding agents now produce code changes faster than people and automated systems can check them. Four accounts from Warp, Linear and Bolt show where each company moved the checking.
Warp, a company that makes tools for software teams, runs an internal system of AI agents that turns a request in Slack into a tested code change. Its chief executive, Zach Lloyd, described the system in a podcast interview that Lenny Rachitsky recapped (22 Sep · 01). The system reaches a proposed change in 35 minutes on average, but the first human review arrives 3.5 hours later. Warp now lets the person who asked for the change review it, instead of waiting for a second reviewer. It also counts how often people have to step in on each change, and uses that count to measure how well the system works.
Lloyd also explains, in a post on Warp's blog, how the company grades its agents' work (22 Sep · 03). Grading agents read the records of a sample of past sessions and pass or fail each one against criteria written as a prompt. One of Warp's criteria is whether the agent wrote redundant tests, a failure the company saw often. People then read the failed sessions and adjust the agents' instructions. The grading costs Warp about 3 percent of what it spends on AI models.
Linear makes software for planning and tracking product work, and AI agents now write most of its tests. As the test suite almost quadrupled this year, the automated checks that every code change must pass grew slow and costly. The chief technology officer assigned the problem to an engineer, Mufeez Amjad (22 Sep · 02), whose team roughly halved the machine time per test. One speed-up risked tests affecting each other, so the team made it opt-in and changed its coding agents' instructions in the same piece of work.
Bolt is a mobility company whose website serves up to 60 million users a year in more than 50 markets. Its designers used to build working components with an AI agent in a separate workspace, and engineers then built each one again for the production library. Burak Özdemir, a tech lead at Bolt (22 Sep · 04), describes how designers now build in the production library itself. Designers approve the visual details, two engineers review the code, and a file of written rules tells the agent what the library expects. For routine components, the time from a working component to a merged one fell from a week to hours.
In each account, people stop checking every line of agent-written work and check a narrower part of it. At Warp, the person who asked for a change reviews it and grading agents check a sample; at Bolt, designers and engineers each review the part they know best. Linear and Bolt also write their rules into files the agents read, so the agents follow the rules before anyone reviews the work.
These are software companies describing their own engineering, and Warp sells the kind of system it describes. When the person who asked for a change also reviews it, nobody else is left to notice what that person did not think to ask for.
Warp, which makes tools for software teams, lets the person who asked its AI agent for a code change review that change, instead of waiting for a separate reviewer.
Linear rebuilt its automated code checks after AI agents made writing code faster than checking it, and taught its agents a new testing rule in the same change.
Warp's chief executive explains how AI agents can grade a sample of coding agents' past work against written criteria, so that people review the failures rather than every change.
At Bolt, designers now build interface components with AI agents directly in the library that ships, and engineers review those components instead of rebuilding them.



