Working Surfacelink · 23 Sep · 03
03
A reader seen from behind, the back of the head showing, sits at a long table before a stack of bound logbooks and pencils a note in the margin of each open page, while two small robots at the far end of the table sort the finished notes into separate trays, the whole scene drawn with wide empty margins on every side.
Lenny's Newsletter

Advanced evals: How to find (and fix) hidden AI failures in your product

Hamel Husain and Shreya Shankar · 22 September 2026

How-to · critique & review · rituals · Evals before the build

Read the original
Takeaway

Two specialists in testing AI products argue that people should read real AI sessions and note the failures themselves before an agent looks for errors or anyone writes a metric.

Summary

Teams building AI products often measure quality with automated metrics before they know which failures matter. Hamel Husain and Shreya Shankar have worked with more than 50 AI companies on testing their products. In a guest post in Lenny's Newsletter, they describe the step most teams skip: reading records of real user sessions to find what goes wrong.

Key points
  • In their example, an AI leasing assistant said goodbye to a prospective tenant who found an apartment too expensive. Most AI agents would count that as a success, but the product's aim was to offer cheaper options.
  • A person reads at least 10 session records and notes anything wrong before an AI agent suggests problems. The authors call this a safeguard against trusting the agent too readily, and set 100 records as the target.
  • The agent then proposes further problems, which the person accepts or rejects. In a study of 100 sessions, agents missed failures that needed knowledge of the product and flagged some good answers as failures.
  • The post is behind a paywall from its third step, and only the open part is cited here.
Implication

The failures a team chooses to measure depend on its own idea of a good experience, which an agent cannot supply at the start. Treating this reading as design work, done by people who know the product, is an inference and is not the authors' claim.

Suggested actions
Product

Have a person read 10 real user sessions with an AI feature and note what went wrong, then 100, before anyone writes a measure of its quality.

Design

Check each failure that an agent proposes against the team's own notes on real sessions, and accept or reject it before the team starts to measure it.

Derived by Working Surface from the article. Source line: Most teams skip the first stage of error discovery and jump straight to writing metrics.

Source issue

23 September 2026: people keep up with agent-built work when it arrives small and already checked, and when they have learned to judge it.