Working Surfaceissue · 23 Sep
Issue · 4 links
Issue

Wednesday
23 September 2026

Four pieces argue that people can keep up with work made by AI agents only when it reaches them in small pieces that have already been checked. They also say that the judgment people bring has to be built first, by reading real cases and by working beside experienced colleagues.

Swipe for 02 to 04
23 Sep
The note

Four pieces argue that people can keep up with work made by AI agents only when it reaches them in small pieces that have already been checked. They also say that the judgment people bring has to be built first, by reading real cases and by working beside experienced colleagues.

Shopify is rebuilding its largest mobile app, which has more than 300 screens, in separate native code for iPhone and Android, using AI models to do the rewriting. Talha Naqvi of Shopify Engineering (23 Sep · 01) describes Helix, the internal tool that splits each screen into small steps, each described in a few words. Every step must pass tests of its behaviour, an AI comparison of screenshots with the old app, and two separate AI code reviewers before an engineer sees it. The agent may retry a failed check but cannot override it. In Naqvi's words, "An attempt is allowed to be wrong. It is not allowed to ship until it isn't."

Nicole Selig, an engineer at the software consultancy SEP (23 Sep · 02), writes that code review broke down once agents produced changes larger and faster than people could read them. Her remedies are established practice. She cuts work into thin slices that each run end to end through a feature, and the team reviews the agent's plan together before any code is written. Feature flags and gradual releases then catch problems after a change is merged.

Hamel Husain and Shreya Shankar have worked with more than 50 AI companies on testing their products (23 Sep · 03). They write that most teams measure quality with automated metrics before they know which failures matter. In their example, an AI leasing assistant said goodbye to a prospective tenant who found an apartment too expensive, and most AI agents would count that as a success. Their method has a person note what is wrong in at least 10 real sessions before an agent suggests anything. The agent then proposes further problems, and the person accepts or rejects each one.

Patrick Neeman, who has designed for the web since 1995 (23 Sep · 04), writes that design leaders are hiring senior designers and few juniors. He cites research showing that AI tools help novices most, and argues that this makes juniors a better hire than before. He pairs them with seniors for the sake of judgment. In his words, "A junior left alone with the tool learns to accept output, while a junior with a senior beside them learns to question it."

What they add up to

None of the four asks reviewers to work harder. Shopify and Selig control the size of what reaches a reviewer and check it automatically first. Husain, Shankar and Neeman describe how people learn to judge AI output: by reading real cases themselves, and by working beside someone who already questions it.

The case against

Shopify is rebuilding an app that already exists, so it has a reference to compare each screen against. Its post gives no example of a new feature built this way. Neeman argues for a hiring practice rather than reporting on a team that tried it.

Receipts
Shopify Engineering
Helix: The internal tool powering our Shopify app's native migrationCompany story · 23 Sep · 01
Talha Naqvi · 21 September 2026

Shopify rebuilds its mobile app with AI models in small steps, each checked by tests, a visual comparison and two reviewing agents before an engineer approves it.

SEP
So, AI Killed Your Code Review Process. Now What?Company story · 23 Sep · 02
Nicole Selig · 22 September 2026

An engineer at a software consultancy argues that when AI-written code outgrows review, teams should cut work into thin slices and review the agent's plan before any code exists.

Lenny's Newsletter
Advanced evals: How to find (and fix) hidden AI failures in your productHow-to · 23 Sep · 03
Hamel Husain and Shreya Shankar · 22 September 2026

Two specialists in testing AI products argue that people should read real AI sessions and note the failures themselves before an agent looks for errors or anyone writes a metric.

UX Collective
In the age of AI, the UX field survives on leaders who cultivate juniorsCompany story · 23 Sep · 04
Patrick Neeman · 22 September 2026

A designer argues that design leaders should hire juniors again, because AI tools make newcomers useful sooner and questioning AI output is learned beside a senior.