Working Surfacelink · 17 Sep · 01
01
Pull requests on a conveyor pass three robot inspectors. A human controls the gate at the end.
The Pragmatic Engineer

Inside OpenAI's agentic software factory

Gergely Orosz · 15 September 2026

Company story · agents · Human approval

Read the original
Takeaway

Gergely Orosz describes how OpenAI builds software with its coding agent Codex: people define the problem, agents write and review the code, and high-risk changes can require a person's review.

Summary

Gergely Orosz, who writes the engineering newsletter The Pragmatic Engineer, visited OpenAI and interviewed seven of its engineering leaders and engineers. Codex, OpenAI's coding agent, now underpins almost all work at the company, and code changes per engineer have grown so fast that build and release systems struggle. The article describes the pipeline OpenAI built around its agents, which it calls a software factory.

Key points
  • An engineer or product manager states the problem and the desired outcome. Codex then gathers context, writes the code, runs the tests and fixes failures until the checks pass.
  • Several review agents check each change, each set up as a specialist in one area. Low-risk changes can be approved automatically, and after a person approves a change for production, an agent watches its rollout.
  • Venkat Venkataramani, OpenAI's vice president of engineering for applied infrastructure, told Orosz that its engineers are becoming more like product managers than traditional systems engineers. Judgment, prioritisation and taste now matter most in the step where a person defines the problem.
  • Finance, recruitment and legal teams went from almost no use of Codex to 90 percent in four months. Experts from such fields now join engineering teams to say what good output looks like.
Implication

Where agents write most of the code, people's work moves to stating the problem and deciding which changes need a person's review. Orosz reports this shift at OpenAI, and its application to other teams is an inference.

Suggested actions
Engineering

Sort each code change written by an AI agent by risk, and require a person to review the high-risk ones.

Design

Put a designer on the team that sets up the company's AI agents, so that the agents' output is judged by a designer's standard of good work.

Derived by Working Surface from the article. Source line: A software engineer or product manager specifies the problem and desired outcome. Judgment, prioritization, and taste are becoming more important for this phase.

Source issue

17 September 2026: when agents build, people's judgment moves to stating the problem and checking the result.