Working Surfacelink · 1 Oct · 02
02
A person seen directly from behind, only the back of the head and the shoulders showing, sits at a desk before a laptop whose screen shows a long column of identical small boxes, one hand resting on the trackpad, the whole scene drawn with wide empty margins on every side.
UX Tigers

User Vigilance Fails Twice in the AI Age: Watch Duty & Verdict Duty

Jakob Nielsen · 29 September 2026

How-to · agents · critique & review · Human approval

Read the original
Takeaway

Jakob Nielsen, a usability researcher, argues that people approving an AI agent's actions stop paying attention, and that products should ask less often and offer undo.

Summary

Products built on AI agents often ask a person to approve each action or to watch a stream of the agent's work. Jakob Nielsen, who writes about usability at UX Tigers, argues that both forms of supervision fail because human attention decays, during long watches and with each repeated approval. He cites a 2026 study by Ting Yan in which people who wrote permission rules in advance blocked less unwanted agent behaviour than people who approved each action by hand.

Key points
  • In Ting Yan's study, 113 US adults without a software background supervised a simulated AI assistant that attempted seven unrequested actions, including a $12 insurance purchase and reading private messages.
  • People who wrote standing rules blocked 40 percent of those actions, against 60 percent for people who approved every action by hand.
  • For long agent runs, Nielsen recommends asking rarely and putting context in each request. He also recommends tiering approvals by consequence, such as spending, publishing or deleting, and offering undo for reversible actions.
  • He recommends logging the median time to approve and the approval rate for each approval step. A median under two seconds or approval above 95 percent is a reason to redesign the step.
Implication

A design team could measure its agent approval steps and review the numbers alongside other usability data; this is an inference and is not the author's claim. The study used a simulated assistant and participants without technical training.

Suggested actions
Engineering

Log the median time to approve and the approval rate for each agent approval step, and redesign steps below two seconds or above 95 percent approval.

Design

Tier an AI agent's approval requests by consequence, such as spending, publishing, deleting and reading private data, rather than by which tool the agent uses.

Derived by Working Surface from the article. Source line: A reviewer who approves agent actions for 3 hours straight stops being a safeguard and becomes an expensive rubber stamp.

Source issue

1 October 2026: as AI agents make changes cheap to produce, teams need to check on purpose whether the people who own the work can still explain and judge it.