Four pieces argue that people can keep up with work made by AI agents only when it reaches them in small pieces that have already been checked. They also say that the judgment people bring has to be built first, by reading real cases and by working beside experienced colleagues.
Shopify is rebuilding its largest mobile app, which has more than 300 screens, in separate native code for iPhone and Android, using AI models to do the rewriting. Talha Naqvi of Shopify Engineering (23 Sep · 01) describes Helix, the internal tool that splits each screen into small steps, each described in a few words. Every step must pass tests of its behaviour, an AI comparison of screenshots with the old app, and two separate AI code reviewers before an engineer sees it. The agent may retry a failed check but cannot override it. In Naqvi's words, "An attempt is allowed to be wrong. It is not allowed to ship until it isn't."
Nicole Selig, an engineer at the software consultancy SEP (23 Sep · 02), writes that code review broke down once agents produced changes larger and faster than people could read them. Her remedies are established practice. She cuts work into thin slices that each run end to end through a feature, and the team reviews the agent's plan together before any code is written. Feature flags and gradual releases then catch problems after a change is merged.
Hamel Husain and Shreya Shankar have worked with more than 50 AI companies on testing their products (23 Sep · 03). They write that most teams measure quality with automated metrics before they know which failures matter. In their example, an AI leasing assistant said goodbye to a prospective tenant who found an apartment too expensive, and most AI agents would count that as a success. Their method has a person note what is wrong in at least 10 real sessions before an agent suggests anything. The agent then proposes further problems, and the person accepts or rejects each one.
Patrick Neeman, who has designed for the web since 1995 (23 Sep · 04), writes that design leaders are hiring senior designers and few juniors. He cites research showing that AI tools help novices most, and argues that this makes juniors a better hire than before. He pairs them with seniors for the sake of judgment. In his words, "A junior left alone with the tool learns to accept output, while a junior with a senior beside them learns to question it."
None of the four asks reviewers to work harder. Shopify and Selig control the size of what reaches a reviewer and check it automatically first. Husain, Shankar and Neeman describe how people learn to judge AI output: by reading real cases themselves, and by working beside someone who already questions it.
Shopify is rebuilding an app that already exists, so it has a reference to compare each screen against. Its post gives no example of a new feature built this way. Neeman argues for a hiring practice rather than reporting on a team that tried it.
Shopify rebuilds its mobile app with AI models in small steps, each checked by tests, a visual comparison and two reviewing agents before an engineer approves it.
An engineer at a software consultancy argues that when AI-written code outgrows review, teams should cut work into thin slices and review the agent's plan before any code exists.
Two specialists in testing AI products argue that people should read real AI sessions and note the failures themselves before an agent looks for errors or anyone writes a metric.
A designer argues that design leaders should hire juniors again, because AI tools make newcomers useful sooner and questioning AI output is learned beside a senior.



