Working Surfaceplaybook

Plays for engineering teams working with AI coding agents

What engineers and their managers can try, from reviewing AI-written code to the instruction files agents read, taken from what the teams in these articles did. Jev, an AI model from TypeSafe, checks that each one makes sense without its article. How the check works

57 plays from 52 links

Human approval

A person other than the one who built a piece of AI-assisted work checks it and approves it before it reaches customers, often called keeping a human in the loop.

14 plays · Topic pageClearest first, as Jev scored them
CHECKED BY JEV · WORKING SURFACE ·7 OCT2026
Engineering

When a pull request was approved without every part being read, say in its description which parts nobody read, so that later reviewers and incident reviews know what was checked.

Elevate · The Code Nobody Reads · 28 Sep 2026
Engineering

Before an AI agent's code change goes live, have an engineer check the agent's plan for watching it in use, and agree in advance what happens if a problem appears.

Cursor · Bots for the last mile: Rollouts, Security Review · 24 Sep 2026
Engineering

Assign the review of each agent-built code change to the tech lead or engineering manager of the team that owns that code, not to whoever is free.

Design Outcomes · The Bottleneck Moved to Review · 21 Sep 2026
Engineering

Before reading an agent-written change line by line, rank its files by risk and spend close review on the riskiest ones.

JetBrains Research · Our Framework for Reviewing AI-Generated Code · 7 Oct 2026
Engineering

Keep a fixed list of changes that always need a person's approval, including changes to the approval rules, and check it with a script the AI agent cannot edit.

Medium · The Approve Button Is Not a Control · 29 Sep 2026
Engineering

Have a coding agent attach a note to each change listing what was asked, what changed, what has no test and what was left out, for the reviewer to check.

Medium · The Approve Button Is Not a Control · 29 Sep 2026
Engineering

Log the median time to approve and the approval rate for each step where a person approves an AI agent's action. Redesign any step with a median under two seconds or approval above 95 percent.

UX Tigers · User Vigilance Fails Twice in the AI Age: Watch Duty & Verdict Duty · 1 Oct 2026
Engineering

Send every change that an AI agent makes through the same validation and saving code that people's changes use, so that the product's rules are kept in one place.

Dropbox.Tech · Evolving our calendar assistant Reclaim to be AI-native without starting over · 30 Sep 2026
Engineering

Report AI agent spending by team per piece of finished, reviewed work, so the cost of reviewing agent output counts, not only the AI model's bill.

Agoda Engineering & Design · Key Takeaways from Agoda’s AI Developer Report 2026 · 29 Sep 2026
Engineering

Let the person who asked an AI agent for a code change review that change, and track how often people have to correct or redirect the agent.

Lenny's Newsletter · How I AI: Meta's Muse review + How Warp ships 2,000 PRs a month with AI factories · 22 Sep 2026
Engineering

Give every change an AI agent makes its own live preview link that people outside engineering can open.

Vercel · How Delphi ships 100 times a day with its Python backend on Vercel · 17 Sep 2026
Engineering

When an AI agent drafts text for users, let the AI model write the wording, and let ordinary code route requests, copy record numbers and check the output format.

Salesforce Engineering · How Deterministic Controls Turn AI Output Into Reliable Prompt Templates · 2 Oct 2026
Engineering

Review an AI agent's code by risk: a person for the risky changes, a lighter check for routine ones.

From 3 articles · OpenAI, GitHub and 1 moreShow all 3Hide
OpenAI · in The Pragmatic Engineer · 17 Sep 2026Each agent-written change is sorted by risk, and a person reviews the high-risk ones.
GitHub · in GitHub Blog · 21 Sep 2026The reviewer of risky code keeps going until they can explain the change and take responsibility for it.
Elevate · 28 Sep 2026Agents review every pull request, careful human review is kept for core and sensitive code, and a person approves every merge.

Agent context files

The files an AI agent reads before it works, such as AGENTS.md, CLAUDE.md, design-system rules and other written instructions, and the person or team who keeps them current.

9 plays · Topic pageClearest first, as Jev scored them
WORKING SURFACECHECKEDBY JEV · 7 OCT 2026
Engineering

Write down how senior engineers investigate a problem, including their checks, trusted signals and order of steps, and turn it into instructions an AI agent follows.

LeadDev · Meta turned engineers’ judgment into agent skills · 6 Oct 2026
Engineering

Test AI agents on the same tasks with and without their added instruction files, because extra context can confuse an agent and make its results worse.

Design Systems Collective · “You don’t fix the output, you fix the system” — Cristian Morales Achiardi on agentic design systems · 29 Sep 2026
Engineering

Store the instructions every AI agent task reads, the files for one task and incoming streams such as email in three separate folders.

Nielsen Norman Group · The New Big Ball of Mud: Why Agentic AI Systems Turn Fragile · 3 Oct 2026
Engineering

Keep every rule about who may see what in one central permissions table, so that adding a type of user or granting a permission is a one-line change.

Accidentally in Code · Good Architecture Is a Function of Time · 29 Sep 2026
Engineering

Split agent work into steps that each map to one pull request, and run the coordination between steps in ordinary code rather than in an AI agent.

Stripe Dot Dev Blog · Stripe’s Payment Method Factory: Orchestrating agents for repeated, custom integrations · 3 Oct 2026
Engineering

After each AI agent run, save what the agent had to learn, the decisions it made alone and the engineers' review comments, and update the reusable prompts from them.

Stripe Dot Dev Blog · Stripe’s Payment Method Factory: Orchestrating agents for repeated, custom integrations · 3 Oct 2026
Engineering

When AI agents write much of the code, invest in shared structure for the parts that change most, such as a design system or generated database queries.

Accidentally in Code · Good Architecture Is a Function of Time · 29 Sep 2026

Changing roles

How the jobs of product managers, designers, engineers and their managers change when AI agents do more of the building.

7 plays · Topic pageClearest first, as Jev scored them
CHECKEDJEV7 OCT 2026
Engineering

Report how long pull requests wait for review and how many are reworked afterwards, next to the amount of code produced.

LeadDev · Faster code isn’t faster delivery · 7 Oct 2026
Engineering

Require the author of an AI-assisted pull request to explain the change, run the relevant tests and check security concerns before requesting a review.

LeadDev · Faster code isn’t faster delivery · 7 Oct 2026
Engineering

State the team's high-level technical values to the coding agent explicitly before it starts work.

Sean Goedecke · Human-AI partnerships are for alignment, not capability · 3 Oct 2026
Engineering

Set stricter automated checks for test coverage and code complexity on AI-written code than on hand-written code, and run them on every change before merging.

Codacy · How AI Is Changing the Engineering Manager Role: More Context, More Capacity, and the New Job of Protecting Focus · 25 Sep 2026
Engineering

Read the code an AI model writes before it ships.

seangoedecke.com · Shipping is the foundation · 5 Oct 2026
Engineering

Cap how many topics each engineer and each squad may have open at once, and hold the cap when stakeholders or engineers push to start more AI-assisted work.

Codacy · How AI Is Changing the Engineering Manager Role: More Context, More Capacity, and the New Job of Protecting Focus · 25 Sep 2026
Engineering

Compare every proposed ticket, design document or delegated task against the cost of making the change directly.

seangoedecke.com · Shipping is the foundation · 5 Oct 2026

Agent ownership

A person or team answers for what an AI agent does, and decides its instructions, its default settings and when it must hand a decision to a person.

6 plays · Topic pageClearest first, as Jev scored them
CHECKED BY JEV · CHECKED BY JEV ·07.10.26
Engineering

Count an AI agent's incident task as finished only when a separate check shows the problem has gone, such as the alarm clearing, not when the agent's action reports success.

LeadDev · Your agent loop is not a production system · 28 Sep 2026
Engineering

Start using AI agents in one low-risk process that the rest of development does not depend on, and name one person in that team to set the agents up.

arXiv · Developing a Roadmap to an AI-first Organization: A Case Study in Embedded Software Development · 25 Sep 2026
Engineering

When AI agents repair their own failing work, cap the repair attempts and the spending, and hand the problem to an engineer once either limit is reached.

Salesforce Engineering · Engineering Multi-Agent AI Teams That Build and Test Themselves · 28 Sep 2026
Engineering

Give each AI coding agent one whole feature, from interface to storage, instead of one layer such as the frontend, so that agents seldom need to brief each other.

Medium · Durable Agentic Teams: Maximizing Your Software Development Workflow · 29 Sep 2026
Engineering

Have AI agents that check other agents' work answer pass, warn or fail in a fixed format, and compare several checkers' answers instead of trusting a single one.

Salesforce Engineering · Engineering Multi-Agent AI Teams That Build and Test Themselves · 28 Sep 2026
Engineering

When an AI agent takes over one low-risk development step, record the review and testing time of the steps before and after it. Compare that with the time saved on the step itself.

arXiv · Developing a Roadmap to an AI-first Organization: A Case Study in Embedded Software Development · 25 Sep 2026

Evals before the build

Before an AI feature is built, the team writes down what a good output looks like, so that the result can be tested against it.

4 plays · Topic pageClearest first, as Jev scored them
CHECKEDJEV7 OCT 2026
Engineering

Before trusting one AI model to grade another's output, label a set of examples by hand and measure how often the grader agrees.

Product Talk · 4 New Evals and 16 Experiment Variants to Fix 1 Customer Complaint · 17 Sep 2026
Engineering

Do not count tests that an AI model wrote for existing code as verification, because they record what the code already does, errors included.

BLOG@CACM · Nobody Did TDD for 25 Years. Now the Machine Requires It · 28 Sep 2026
ProductEngineering

Before an AI agent starts building, agree with the people who approve its changes which automatic tests every change must pass.

From 3 articles · Anthropic, GitHub Next, PikdShow all 3Hide
Anthropic · in Anthropic · 24 Sep 2026The automatic checks are agreed with the people who approve changes, and tests are added until those engineers would merge on the results alone.
GitHub Next · in The Pragmatic Engineer · 26 Sep 2026The spec for a prototype lists how the agent will verify each part of its work.
Pikd · in BLOG@CACM · 28 Sep 2026The tests that define a correct result are written first, the agent works until they pass, and people review the tests instead of the code.

Training juniors

How people new to a craft build judgement when AI tools do the work that used to teach it.

4 plays · Topic pageClearest first, as Jev scored them
CHECKEDJ7 OCT 2026
Engineering

Never assign a junior engineer a project alone, and have a senior engineer review the junior's pull requests, including code written with AI.

LeadDev · AI-coding tools could be breaking the junior engineer pipeline · 6 Oct 2026
Engineering

Run incident reviews that ask what the team believed that turned out to be untrue, so that junior engineers practise judging their own assumptions.

LeadDev · AI makes critical thinking harder to build · 29 Sep 2026
Engineering

Teach junior engineers to review what an AI tool produces and to explain why some of its choices may cause problems.

LeadDev · AI-coding tools could be breaking the junior engineer pipeline · 6 Oct 2026
Engineering

Write critical thinking into career frameworks as specific behaviours, such as validating assumptions before implementation, so that managers can see a junior engineer's growth that shipped tickets no longer show.

LeadDev · AI makes critical thinking harder to build · 29 Sep 2026

Understanding the work

A team checks on purpose that its people can still explain work that AI agents produced.

3 plays · Topic pageClearest first, as Jev scored them
7 OCT2026CHECKED BY JEV

Work cut to size

Before an AI agent starts, the work is split into pieces small enough for a person to review one at a time.

2 plays · Topic pageClearest first, as Jev scored them
CHECKED07.10.26JEV

After the prototype

AI makes a working prototype quick, so the effort moves to choosing a direction and getting from prototype to production with a reliable product.

1 play · Topic page
CHECKEDJEV07.10.26

Other plays

Plays from links that are not filed under a topic.

7 playsClearest first, as Jev scored them
CHECKED BY JEV · WORKING SURFACE ·7 OCT2026
Engineering

When AI agents build the same feature separately for iOS and Android, make both versions pass one shared test suite for the business logic before either ships.

The Pragmatic Engineer · Why has Shopify dropped React Native? · 30 Sep 2026
Engineering

Require a test coverage check to pass before any pull request merges, because an AI coding agent may ignore a testing rule written in its instruction file.

Accidentally in Code · Is the software factory working? · 7 Oct 2026
Engineering

Offer each new AI model to staff on release day as an experiment with a capped budget, and keep it only if benchmarks, user reports and cost data support it.

Databricks · How Databricks rolls out frontier models to 12,000 employees on Day 1 · 28 Sep 2026
Engineering

Compare a new AI model's cost per session on the same group of early users before and after, because early users are heavier users of AI than most staff.

Databricks · How Databricks rolls out frontier models to 12,000 employees on Day 1 · 28 Sep 2026
Engineering

Have the same developer build each mobile feature on both iOS and Android with AI coding agents, instead of keeping separate iOS and Android teams.

The Pragmatic Engineer · Why has Shopify dropped React Native? · 30 Sep 2026
Engineering

Sort pull requests by type each month and track features, automated checks and same-day emergency fixes side by side.

Accidentally in Code · Is the software factory working? · 7 Oct 2026
Engineering

Before an AI model grades an AI agent's output, run a free check in ordinary code first, such as a rule that flags a list with too many items.

Product Talk · Generating Opportunity Solution Trees with AI: How Vistaly Rebuilt Its Product Around Interview Synthesis, Evals, and Repair Loops · 5 Oct 2026