What Shipping an AI Agent to 16,000 People Taught Me About Reviewing One
Ananya Das · 17 September 2026
Read the originalAnanya Das, an intern at DataHub, tested a support agent before its release and found that an unchecked human review had made the agent look worse than it was.
DataHub makes software that helps companies find and govern their data, and runs a community Slack workspace of more than 16,000 data practitioners. Ananya Das, a summer intern who is not an engineer, tested Otto, a new AI agent that answers the community's technical questions, before it went live. She graded its answers to 21 real questions over four rounds, and the last round taught her to check the reviewers as well as the agent.
- Over the first three rounds, Otto's pass rate rose from 33 to 52 percent as the team fixed problems.
- In the fourth round, another internal AI agent wrote reference answers and a human reviewer added corrections. Graded against these, Otto's passes fell from eleven to seven.
- Das re-checked every changed grade against the code itself. Five were wrong because the reviewer had not checked the other agent's corrections closely, and reverting them restored the third round's result.
- Three real problems remained: Otto compared DataHub with competitors, lacked access to parts of the code, and invented answers when it lacked information. Each was fixed before Otto went live.
A human review adds error when the reviewer has no reliable reference to check against. DataHub's own conclusion is that the review must be structured and done by an expert in the subject, and the post is on a vendor's blog.
Before a person reviews an AI agent's answers, give that person a verified reference to check them against, and re-check any corrections the reviewer makes.
Ask an expert in the subject, not a generalist, to review an AI agent's output, with a structured way to check each answer.
Derived by Working Surface from the article. Source line: A human sign-off only makes AI-generated content trustworthy if the review itself is structured and checked by the right domain expert.
18 September 2026: a review of an agent's work is only as good as the reference behind it.