Working Surfacelink · 28 Sep · 07
07
At night in an old railway signal box, an operator seen from behind, the back of the head showing, hands a small signed slip to a small robot standing at one lever in a long row of levers, while a large dial on the wall beside them still points high, the whole scene drawn with wide empty margins on every side.
LeadDev

Your agent loop is not a production system

Sriram Madapusi Vasudevan · 28 September 2026

How-to · agents · roles · Agent ownership

Read the original
Takeaway

An Amazon Web Services engineer argues that once an AI agent can change live systems, teams must prove separately who approved a change, what ran and whether it worked.

Summary

Sriram Madapusi Vasudevan, a senior software engineer at Amazon Web Services (AWS), works on AWS DevOps Agent, an AI agent that joins on-call engineers during production incidents. He describes what the system around such an agent must record and check once it can change cloud resources. He argues that engineering leaders should set these rules for every team before any agent may change production.

Key points
  • An approval should name the exact operation, resource, approver and expiry. It should survive a restart, be used once, and be asked for again if the target changes.
  • Every change an agent makes should pass through one gateway, whether the agent calls a tool directly, writes code, or hands the task to another agent.
  • A rollback can succeed while the errors continue, so the system should confirm that the incident has cleared before it counts the task as done.
  • He proposes that a platform team own the shared machinery and that domain teams own their tools and what counts as done. Operators own approval limits and the power to stop work.
Implication

An engineering team that lets an AI agent act on production can store each approval with its scope and expiry, and check the outcome before counting a task done. The author draws on his own work on an AWS product, and the application to other teams is an inference.

Suggested actions
Engineering

Before an AI agent changes a live system, require a person to approve that exact change, let each approval be used once, and ask again if the change is altered.

Engineering

Count an AI agent's incident task as finished only when a separate check shows the problem has gone, such as the alarm clearing, not when the agent's action reports success.

Derived by Working Surface from the article. Source line: Approval proves authority. Execution evidence proves what changed. Verification proves whether the goal was reached.

Source issue

28 September 2026: when agents build, people keep the first screen, the review, the files agents read and the whole product.