Working Surfaceissue · 24 Sep
Issue · 4 links
Issue

Thursday
24 September 2026

Four pieces from Anthropic, Stripe, Cursor and Evil Martians describe setting the rules for AI agents' work before the work starts. The rules go into automatic checks, review policies, monitoring plans and the components themselves, and people judge what falls outside them.

Swipe for 02 to 04
24 Sep
The note

Four pieces published on 22 and 23 September describe how teams prepare for work made by AI agents before the work arrives. They agree the rules in advance, build them into checks, tools or components, and leave people to judge what falls outside the rules.

Jonah Ezekiel and Lexie Tonelli, engineers who work with Anthropic's customers (24 Sep · 01), wrote a guide to rewriting old systems with AI agents in banks and other regulated companies. There, every change to a critical system is reviewed and approved, but agents produce changes far faster than anyone can review them one by one. The authors advise agreeing two things before the work starts. The first is a set of automatic checks that every change must pass, written with the people who will approve the changes. The second is a review policy, set from the top of the organisation, that gives risky changes more human review. Agreeing it in advance means, in their words, "responsibility for a bug that reaches production is shared, not pinned on whoever approved the change."

AI shopping agents have to pay on checkout pages built for people. On Stripe's ready-made payment page, a purchase cost an agent 39 actions and 1.8 million tokens on average. Tokens are the units of text an AI model reads and is billed by. Cara Mecozzi and Steve Kaliski of Stripe (24 Sep · 02) describe how the page now offers agents its actions as named tools, using an emerging browser standard called WebMCP. The page offers only the actions valid at each step, so the payment button appears only once the form is complete. The tools run on the same code as the checkout people use, and in Stripe's tests agents used 42 percent fewer tokens.

Cursor, which makes AI tools for writing code, launched two bots for the work after a code change is proposed. Rustam Lalkaka of Cursor (24 Sep · 03) writes that this work, checking security and watching the release, has not sped up as writing code has. Before a change is merged, the Rollouts bot writes a monitoring plan that a person can correct. After release, it compares live measurements with those from before, and it can propose reversing the change for a person to approve.

Anton Lovchikov and Yuri Mandrikov of the consultancy Evil Martians (24 Sep · 04) write about AI agents that build screens from bare components. The agents guess how to use them, override their styles and create duplicates. Their framework keeps an agent that builds a screen from changing the design system; it records the missing component and works around it. Each component carries a written contract saying what it is for and when to use it. Any local exception must state its reason and goes into a log that the team reviews.

What they add up to

In all four pieces, people decide the rules before the agents start. Anthropic puts them into checks and a review policy, Stripe and Evil Martians put them into the page or the component itself, and Cursor puts them into a monitoring plan. People then spend their time on the cases the rules do not settle: the riskiest changes, the proposed reversals and the logged exceptions.

The case against

Three of the four pieces come from companies selling the tools they describe, and only Stripe reports measured results, from its own tests. Anthropic's authors say their checks are hardest to write for new behaviour, where no old system exists to compare against, which is where much product work begins.

Receipts
Anthropic
How to prepare for AI-driven code modernization projectsCompany story · 24 Sep · 01
Jonah Ezekiel and Lexie Tonelli · 23 September 2026

Two Anthropic engineers advise companies modernising old code with AI agents to agree, before the work starts, what each change must prove and how much human review it gets.

Stripe
How Stripe is designing Checkout for AI agentsHow-to · 24 Sep · 02
Cara Mecozzi and Steve Kaliski · 22 September 2026

Stripe made its Checkout page cheaper for AI shopping agents to use by offering them only the actions valid at each step, built on the human checkout's own code.

Cursor
Bots for the last mile: Rollouts, Security ReviewCompany story · 24 Sep · 03
Rustam Lalkaka · 23 September 2026

Cursor launched two AI bots for the work after a code change is proposed: one plans and watches the release, and one reviews every change for security flaws.

Evil Martians
AI makes design system guardrails mandatory; this framework delivers themCompany story · 24 Sep · 04
Anton Lovchikov and Yuri Mandrikov · 23 September 2026

Evil Martians, a product consultancy, stops AI agents that build screens from changing the design system, and records every exception for a person to review.