When AI agents do the building, design and engineering leads still keep some of the work for themselves. Five pieces published between 24 and 27 September show which work: the first screen and the rules, the review, the files agents read, and responsibility for the whole product. Five engineering pieces from 28 September cover reviewing AI-written code, agents that change live systems, agents that build agents, new AI models and tests written first.
Peter Yang, who runs the interview newsletter Behind the Craft, spoke with the two leads who build Grok Bot, the AI assistant made by SpaceXAI (28 Sep · 01). Peng Zheng, the lead designer, builds the design system and one finished screen by hand. His design bot then extends that screen into every other screen of the flow, working directly in Figma. Lauren Tan, the lead engineer, hands large projects to a bot called Matcha, which splits the work among other coding bots. The full episode is for paid subscribers, and this account comes from its open show notes.
Leonardo De La Rocha leads design for software used by clinicians, a product with twelve years of interface patterns behind it (28 Sep · 02). Each new pattern is one more thing a clinician has to learn. He and one of his design leaders therefore wrote two rules: reuse existing patterns before inventing new ones, and judge every design within the whole screen around it. Both rules came out of design reviews. In one review, a designer showed three layouts made by AI for a deadline tracker. A senior designer saw that one layout looked like the product's clickable tabs, although the tracker cannot be clicked.
Clinicians at mental-health practices that use this software kept receiving patients' insurance-card photos sideways, and fixed each one by hand: download, rotate, upload again. A designer on the company's billing team fixed the problem himself in one afternoon, using an AI coding agent. De La Rocha then reviewed the fix with him (28 Sep · 03). His view is that the fix took the feature from not great to good in an afternoon. His review is for the step after, from good to great: "When it takes an afternoon, good is closer to a first draft."
Brian Houck, a scientist at the developer-research firm DX, wrote on Atlassian's blog about the context that AI agents read (28 Sep · 04). Context means the instructions, specifications and documents an agent is given before it works. Houck and four co-authors define five qualities good context needs, from clarity to security. They also say who should look after it: files that many agents read should be reviewed like code, and each should belong to a team. In their words, "Ownership should follow teams rather than individuals, so context survives reorgs."
David Heinemeier Hansson created the web framework Ruby on Rails and co-founded 37signals, the company behind Basecamp. At the Rails World conference he said that 37signals now rarely writes code by hand. Jared Smith, an engineering leader, wrote up the talk (28 Sep · 05). This spring, 37signals designers built the last features of Basecamp 5 with AI agents. Each change looked reasonable, but together they left the code, in Hansson's words, "a little like a Swiss cheese." Hansson expects a better AI model to prevent this. Smith disagrees, because nobody was responsible for how the changes fit together, and a better model does not change that.
Addy Osmani writes the Elevate newsletter and says he helps build an AI tool that writes code. In "The Code Nobody Reads" (28 Sep · 06), he withdraws his own advice of two years ago to review every line of AI-generated code. He was prompted in part by an engineer at a large company whose post about approving code nobody had read reached more than eight million people. Instead, Osmani proposes that AI agents review every pull request, that people review risky changes carefully, and that a person approve every merge. He also asks leaders to give engineers a say in the pace they are blamed for, because "Accountability needs authority."
Sriram Madapusi Vasudevan is a senior software engineer at Amazon Web Services (AWS), where he works on AWS DevOps Agent, an AI agent that helps on-call engineers during production incidents. In LeadDev (28 Sep · 07), he describes what changes once such an agent can change cloud resources instead of only recommending changes. An approval has to be stored with its exact target and expiry, and every change has to pass through one gateway. A task is finished only when the incident has cleared. He proposes that a platform team own this shared machinery, while each domain team decides what counts as done and operators keep the power to stop work. In his words, "Approval proves authority. Execution evidence proves what changed."
Salesforce's engineering blog interviewed Sohini Arya, who leads AI delivery for Marketing Cloud, the company's marketing software business (28 Sep · 08). Her engineers spent two to four hours building and testing each team of cooperating AI agents by hand. Her team built Agent Designer, in which a coordinating agent and seven specialist agents design, build and test the agent team, aiming to finish in 15 to 30 minutes. An engineer approves the design before any file is written, and agents that check the work must answer pass, warn or fail. "The team limited the self-healing process to two fix attempts and a $3 spending ceiling," after which an engineer takes over.
Databricks, a company that sells a data and AI platform, gives more than 10,000 employees each new AI model on the day it is released (28 Sep · 09). Its AI product and engineering team explained why care is needed: one new AI model cost more and scored lower with the company's engineers than the version before. Each new AI model now arrives in employees' coding tools marked as experimental, paid from a separate share of each person's budget. Within about three days the team keeps or drops it, using private tests, users' reports and cost per session. The account is Databricks' own and promotes Unity Gateway, a product it sells. To measure cost, the team compares the same early users before and after, because "early adopters tend to be power users of AI rather than average users."
Abtin Aghagolian, co-founder and chief technology officer of the London augmented-reality company Pikd, argued on the Communications of the ACM blog that AI agents have made test-driven development necessary (28 Sep · 10). For 25 years most programmers skipped the practice of writing a failing test before the code. Now an engineer can write the test, let the agent write code until it passes, and review the test instead of the code. Pikd's own assistant passed 1,053 tests while contradicting itself in a single reply, and he knows of no way to test for every such contradiction in advance. He warns that an agent will pass a weak test without being correct: "If the machine optimizes against your check, your check is the product."
In each piece, agents do more of the making, and people keep the parts that decide whether the result is right. The design lead keeps the first screen, the reviewer keeps the step from good to great, and a team keeps the files agents read. Smith adds one more: someone has to answer for how all the pieces fit together.
Most of these accounts come from small teams, or from one leader describing their own work, and the Grok Bot interview is mostly behind a paywall. In a larger organisation, a platform team may control the design system and the agents' files, so a product team cannot keep them.
Two leads at SpaceXAI show how they share work with AI agents: the designer builds the first screen by hand, and the engineer's agent divides projects among other agents.
A design team for clinical software drafted two rules for keeping its product simple, reuse existing patterns and judge each design in context, from reviews that included AI-generated layouts.
A designer fixed a billing problem for clinicians in an afternoon with an AI coding agent, and his design lead used the review to take it from good to great.
Brian Houck of the developer-research firm DX defines five qualities of the instructions AI agents read, and argues that each shared instruction file needs a team that owns it.
Engineering leader Jared Smith argues that when AI agents write the code, a person must still own how the changes fit together, which a better model will not fix.
Addy Osmani withdraws his advice to read every line of AI-written code and proposes automated review of every change, human review matched to risk, and a person approving every merge.
An Amazon Web Services engineer argues that once an AI agent can change live systems, teams must prove separately who approved a change, what ran and whether it worked.
A Salesforce team built AI agents that design and test teams of other agents, aiming to cut hours to minutes, with an engineer approving every design.
Databricks, a data and AI software company, gives staff each new AI model on release day within a capped budget, then keeps or drops it on measured cost and quality.
Abtin Aghagolian, chief technology officer of Pikd, a London augmented-reality company, argues that tests written before an AI agent writes code are now how engineers control its output.









