Working Surfaceissue · 18 Sep
Issue · 3 links
Issue

Friday
18 September 2026

Three pieces show that a person's review of an AI agent's work is only as good as the reference that person checks against. Without a reliable reference, a reviewer can add errors instead of catching them.

Swipe for 02 to 03
18 Sep
The note

Three pieces from DataHub and Airbnb ask what a person checks an AI agent's work against. Each finds that a review is only as good as the reference behind it, whether that is a set of checked answers, approved data definitions or a written method.

Ananya Das, a summer intern at DataHub, a maker of data-management software, tested Otto, an AI agent that answers technical questions in DataHub's community Slack of over 16,000 people (18 Sep · 01). She graded its answers to 21 real questions over four rounds. In the last round, another AI agent wrote reference answers and a human reviewer added corrections, and Otto's passes fell from eleven to seven. Re-checking against the code showed that five of the changed grades were wrong, because the reviewer had not checked the other agent's corrections closely. Her lesson is to "check whoever’s checking your agent, too."

In a second DataHub post, Lakshay Nasa argues that AI agents answering business questions fail mainly because the company data they read is wrong, out of date or contradictory (18 Sep · 02). Miro, maker of an online whiteboard, connected an agent to more than 20,000 datasets, and it first answered under 40 percent of 900 expert-written questions correctly. Checked descriptions of the data and curated documentation took it above 90 percent. In DataHub's product, experts approve the data definitions before any agent uses them, so the post is also a vendor's case for its own product.

Wren Dougherty describes Insight Miner, a tool Airbnb built to analyse large volumes of text such as customer-support conversations (18 Sep · 03). Before launching an AI customer service assistant, Airbnb needed to know what situations it would meet, and each investigation took months by hand. People now spend their time on the most ambiguous cases and on testing hypotheses, and most users work outside technical roles. Insight Miner builds the analysis method into software around an AI agent, so that, in Dougherty's words, "results can be reproduced, audited, and challenged."

What they add up to

In each piece, a person's approval is useful only when that person has something reliable to check against. Das's reviewer had no such reference and added errors. DataHub and Airbnb both put the reference in a shared record, approved data definitions or a written method, instead of in one person's confidence.

The case against

Two of the three pieces come from DataHub, and Das tested a support agent on questions whose correct answers were in the code. Much product work has no answer key, so a checked reference may not be possible to write.

Receipts
DataHub
What Shipping an AI Agent to 16,000 People Taught Me About Reviewing OneCompany story · 18 Sep · 01
Ananya Das · 17 September 2026

Ananya Das, an intern at DataHub, tested a support agent before its release and found that an unchecked human review had made the agent look worse than it was.

DataHub
Context Engineering for AI Agents: Why the Hard Part Isn't the Context WindowHow-to · 18 Sep · 02
Lakshay Nasa · 16 September 2026

DataHub, which sells data-management software, argues that AI agents give confident wrong answers because the company data they read was never checked, and that experts should approve it first.

The Airbnb Tech Blog
Beyond the model: Engineering AI infra with scientific judgementCompany story · 18 Sep · 03
Wren Dougherty · 15 September 2026

Airbnb built Insight Miner, a tool that writes its analysis method into software around an AI agent, so that findings from customer-support data can be reproduced and checked.