Teams that build with AI agents are now writing down what the agent may not do, and what a person still has to check. Three pieces this week show that limit set in code, in the promises a product makes, and in the tools a product exposes to agents.
Salesforce, the business software company, sells Prompt Builder, a tool in which administrators write reusable instructions for its AI features. Vaibhav Raizada, a senior software engineer, and his colleagues Kumar Kasimala and Ashish Gite (2 Oct · 01) built an AI agent that drafts these templates. They judged the most dangerous failure to be a template that looked correct but broke later, because the AI model had invented a data field or altered a record number. They let the AI model interpret the request and write the text, and let ordinary code control routing, record numbers and the output format. The agent can propose a draft, but it cannot save, publish or run a template, so an administrator reviews each change and saves it separately.
Jakob Nielsen, the usability researcher who writes UX Tigers (2 Oct · 02), argues that users judge a product by comparing it with what they expected beforehand. He warns that fluent AI writing leads users to assume the facts are as reliable as the prose. He tells teams to show what the system checked, what it inferred and what still needs review, and to promise less than they deliver. He writes that "Conversational polish can therefore invite more delegation than the system's reliability warrants."
PostHog, which makes product analytics software, studied 63 million calls that AI agents made to its tools over 90 days. Natalia Amorim (2 Oct · 03) reports that only about 7 percent came from a chat window; most came from agents working in code editors and terminals. The two most common actions were running a database query and reading what data exists. The company built a way to replay an agent's session call by call, so that its team can see what these new users do.
The fourth piece is not about AI. Teresa Torres, who writes Product Talk, and Petra Wille host the podcast All Things Product (2 Oct · 04). In a short episode they argue that innovation is a means of solving a customer's problem and not a goal. They trace how logging in changed from passwords to emailed links to passkeys, and they note that each new method asks users to learn a new system. This account comes from the open show notes, because the transcript is for paying subscribers.
In each case the team decided in advance where the AI agent stops and a person takes over. Salesforce put the stop in the product's code, and Nielsen puts it in what the product tells its users. PostHog makes the agent's own behaviour visible, so that a team can decide where the stop belongs. Writing that limit down early appears to cost less than discovering it after a failure, which is an inference from these three pieces.
Two of the three pieces are vendors describing their own products, and Nielsen offers principles rather than a team's measured result. None of them shows that a written limit reduced failures in practice. The podcast does not mention AI agents, so its bearing on them is an inference.
Salesforce engineers let an AI agent draft prompt templates for administrators but gave it no power to save, publish or run them, so a person approves every change.
Usability researcher Jakob Nielsen argues that fluent AI writing makes users over-trust its facts, so teams must design user expectations as carefully as screens.
PostHog studied 63 million AI agent calls to its tools, found most came from code editors, and built a way to replay them.
Teresa Torres and Petra Wille argue that innovation is a means of solving a customer's problem, and that teams which chase novelty build complex products customers then have to learn.



