4 New Evals and 16 Experiment Variants to Fix 1 Customer Complaint
Teresa Torres · 16 September 2026
Read the originalTeresa Torres spent three weeks and sixteen experiments fixing one customer complaint about an AI product, starting by measuring how often the error occurred.
Teresa Torres, who teaches product discovery and writes Product Talk, helps the software company Vistaly build an AI feature that turns customer interviews into an opportunity solution tree. The tree is a map of customer needs grouped under a business goal. A beta customer found one branch with many needs listed side by side and no grouping, and Torres spent three weeks fixing the cause.
- She first wrote evals, checks that measure how often an error occurs: a code check of the tree's shape and a second AI model that graded the groupings.
- The AI grader wrongly flagged many correct groupings. She measured it against her own hand labels, and wrote two more graders for earlier errors that were confusing it.
- Sixteen experiments with prompts, models and workflow followed. The fix was a code check in the agent's existing self-review step, which sends any overcrowded group back to the agent to split.
- The final version cut entries that merely restate the heading above them by 78 percent, and produced 65 percent more subgroups.
A team fixing an AI feature can measure the error first, and check any AI grader against its own judgment before trusting it. Torres works with Vistaly, so this is a partner's account of the product.
Before trusting one AI model to grade another's output, label a set of examples by hand and measure how often the grader agrees.
Measure how often a known error appears in AI output before changing the prompt, and measure it again after every change.
Derived by Working Surface from the article. Source line: You have to calibrate the judge against your own judgment first.
17 September 2026: when agents build, people's judgment moves to stating the problem and checking the result.