Four pieces published on 15 and 16 September show where people's judgment goes when AI agents do the building. It moves to stating the problem and measuring the output before the build, and to reviewing the running result before it reaches production.
Gergely Orosz, who writes the engineering newsletter The Pragmatic Engineer, visited OpenAI and described how it builds software now that its coding agent, Codex, underpins almost all its work (17 Sep · 01). An engineer or product manager states the problem and the outcome, and Codex writes, tests and fixes the code. Several review agents, each set up as a specialist in an area such as security, then check each change. High-risk changes can require a person's review, and a person approves a change before an agent takes it into production.
Teresa Torres, who teaches product discovery, helps the software company Vistaly build an AI feature that groups customer needs from interviews into a map (17 Sep · 02). A beta customer found one part of her map with many needs listed side by side and no grouping. Before changing anything, Torres wrote checks to measure how often the error occurred, and tested her AI grader against labels she made by hand. After sixteen experiments over three weeks, a code check added to the feature's self-review step fixed the problem. Her advice on AI graders, which she calls judges, is this: "You have to calibrate the judge against your own judgment first."
Torres and Petra Wille, hosts of the podcast All Things Product, question the claim that AI agents have made software delivery free (17 Sep · 03). They say the first 60 to 70 percent of a product, a good-looking prototype, is now fast to build, but the last 30 percent still takes months to years. The full transcript is for paid subscribers, and this account comes from the open show notes. Teams that treat delivery as free build more features, and their data models grow tangled. As the episode puts it, "By the time you're getting to feature 15, your data model looks like a Frankenstein strategy."
Vercel, a cloud platform, published a customer story on Delphi, a company with ten engineers that turns people's writing and recordings into digital minds others can talk to (17 Sep · 04). Delphi's engineers hand most problems to AI agents, and every change gets a live preview link that anyone can open from a phone. Delphi ships to production more than 100 times a day, and its chief product officer and growth teams build and release their own experiments. In the story's words, "The preview deployment is where they first see what the agent built."
In each piece, agents do more of the building, and people's work moves to its two ends. At the start, someone states the problem or writes checks for good output; at the end, someone reviews the running result. Torres and Wille add that the last 30 percent of a product still needs skilled engineers.
Two of the four pieces come from people with a commercial stake: Torres works with Vistaly, and Vercel sells the platform Delphi uses. Most teams have neither OpenAI's pipeline nor three weeks to spend on one complaint.
Gergely Orosz describes how OpenAI builds software with its coding agent Codex: people define the problem, agents write and review the code, and high-risk changes can require a person's review.
Teresa Torres spent three weeks and sixteen experiments fixing one customer complaint about an AI product, starting by measuring how often the error occurred.
Teresa Torres and Petra Wille argue that AI coding agents made building a feature cheaper but did not make delivering a reliable product free.
Delphi, a company with ten engineers, ships to production over 100 times a day and first sees its AI agents' work on live preview links, according to Vercel.



