Advanced evals: How to find (and fix) hidden AI failures in your product
Hamel Husain and Shreya Shankar · 22 September 2026
Read the originalTwo specialists in testing AI products argue that people should read real AI sessions and note the failures themselves before an agent looks for errors or anyone writes a metric.
Teams building AI products often measure quality with automated metrics before they know which failures matter. Hamel Husain and Shreya Shankar have worked with more than 50 AI companies on testing their products. In a guest post in Lenny's Newsletter, they describe the step most teams skip: reading records of real user sessions to find what goes wrong.
- In their example, an AI leasing assistant said goodbye to a prospective tenant who found an apartment too expensive. Most AI agents would count that as a success, but the product's aim was to offer cheaper options.
- A person reads at least 10 session records and notes anything wrong before an AI agent suggests problems. The authors call this a safeguard against trusting the agent too readily, and set 100 records as the target.
- The agent then proposes further problems, which the person accepts or rejects. In a study of 100 sessions, agents missed failures that needed knowledge of the product and flagged some good answers as failures.
- The post is behind a paywall from its third step, and only the open part is cited here.
The failures a team chooses to measure depend on its own idea of a good experience, which an agent cannot supply at the start. Treating this reading as design work, done by people who know the product, is an inference and is not the authors' claim.
Have a person read 10 real user sessions with an AI feature and note what went wrong, then 100, before anyone writes a measure of its quality.
Check each failure that an agent proposes against the team's own notes on real sessions, and accept or reject it before the team starts to measure it.
Derived by Working Surface from the article. Source line: Most teams skip the first stage of error discovery and jump straight to writing metrics.
23 September 2026: people keep up with agent-built work when it arrives small and already checked, and when they have learned to judge it.