Working Surfaceissue · 2 Oct
Issue · 4 links
Issue

Friday
2 October 2026

Three pieces describe where an AI agent stops and a person takes over, whether in product code, in a product's promises or in its tools for agents. Each one treats that limit as something to design and write down before anything fails. A fourth, a podcast on innovation, argues that a new solution is worth building only when it serves the customer better.

Swipe for 02 to 04
2 Oct
The note

Teams that build with AI agents are now writing down what the agent may not do, and what a person still has to check. Three pieces this week show that limit set in code, in the promises a product makes, and in the tools a product exposes to agents.

Salesforce, the business software company, sells Prompt Builder, a tool in which administrators write reusable instructions for its AI features. Vaibhav Raizada, a senior software engineer, and his colleagues Kumar Kasimala and Ashish Gite (2 Oct · 01) built an AI agent that drafts these templates. They judged the most dangerous failure to be a template that looked correct but broke later, because the AI model had invented a data field or altered a record number. They let the AI model interpret the request and write the text, and let ordinary code control routing, record numbers and the output format. The agent can propose a draft, but it cannot save, publish or run a template, so an administrator reviews each change and saves it separately.

Jakob Nielsen, the usability researcher who writes UX Tigers (2 Oct · 02), argues that users judge a product by comparing it with what they expected beforehand. He warns that fluent AI writing leads users to assume the facts are as reliable as the prose. He tells teams to show what the system checked, what it inferred and what still needs review, and to promise less than they deliver. He writes that "Conversational polish can therefore invite more delegation than the system's reliability warrants."

PostHog, which makes product analytics software, studied 63 million calls that AI agents made to its tools over 90 days. Natalia Amorim (2 Oct · 03) reports that only about 7 percent came from a chat window; most came from agents working in code editors and terminals. The two most common actions were running a database query and reading what data exists. The company built a way to replay an agent's session call by call, so that its team can see what these new users do.

The fourth piece is not about AI. Teresa Torres, who writes Product Talk, and Petra Wille host the podcast All Things Product (2 Oct · 04). In a short episode they argue that innovation is a means of solving a customer's problem and not a goal. They trace how logging in changed from passwords to emailed links to passkeys, and they note that each new method asks users to learn a new system. This account comes from the open show notes, because the transcript is for paying subscribers.

What they add up to

In each case the team decided in advance where the AI agent stops and a person takes over. Salesforce put the stop in the product's code, and Nielsen puts it in what the product tells its users. PostHog makes the agent's own behaviour visible, so that a team can decide where the stop belongs. Writing that limit down early appears to cost less than discovering it after a failure, which is an inference from these three pieces.

The case against

Two of the three pieces are vendors describing their own products, and Nielsen offers principles rather than a team's measured result. None of them shows that a written limit reduced failures in practice. The podcast does not mention AI agents, so its bearing on them is an inference.

Receipts
Salesforce Engineering
How Deterministic Controls Turn AI Output Into Reliable Prompt TemplatesHow-to · 2 Oct 2026
Vaibhav Raizada, Kumar Kasimala and Ashish Gite · 30 September 2026

Salesforce engineers let an AI agent draft prompt templates for administrators but gave it no power to save, publish or run them, so a person approves every change.

UX Tigers
User Satisfaction Depends on Delivered Quality Relative to ExpectationsHow-to · 2 Oct 2026
Jakob Nielsen · 1 October 2026

Usability researcher Jakob Nielsen argues that fluent AI writing makes users over-trust its facts, so teams must design user expectations as carefully as screens.

PostHog
How AI agents behave: lessons from 63M MCP tool callsCompany story · 2 Oct 2026
Natalia Amorim · 30 September 2026

PostHog studied 63 million AI agent calls to its tools, found most came from code editors, and built a way to replay them.

Product Talk
Unpacking Innovation - All Things Product Podcast with Teresa Torres & Petra WilleHow-to · 2 Oct 2026
Teresa Torres and Petra Wille · 29 September 2026

Teresa Torres and Petra Wille argue that innovation is a means of solving a customer's problem, and that teams which chase novelty build complex products customers then have to learn.