Our Framework for Reviewing AI-Generated Code
Katie Fraser, Agnia Sergeyuk and Ilya Zakharov · 6 October 2026
Read the originalJetBrains researchers propose that tools for reviewing agent-written code should show an overview, rank files by risk and only then open code, based on workshops with 17 practitioners.
Katie Fraser, Agnia Sergeyuk and Ilya Zakharov work in the Human-AI Experience team at JetBrains, which makes programming tools. With researchers at Lund University, they asked developers how a tool for reviewing AI-generated code should work. Their paper, to be presented at Empirical Software Engineering International Week in October, proposes a three-level review workflow.
- The design came from four workshops with 17 practitioners, followed by a survey of 43 software professionals.
- An AI model presents every line with the same apparent confidence, and a reviewer cannot ask it about its reasoning, so the cues a human author gives are missing.
- The proposed tool first gives an overview, then ranks files by risk before the reviewer reads code, and only then opens small pieces of code for close reading.
- The authors warn that tools which only make changes easier to understand may still lead reviewers to spend effort on low-risk code and miss high-risk code.
A team reviewing large agent-written changes can decide which files deserve close reading before anyone reads line by line. This is an inference and is not the authors' claim; the piece is a research proposal from a company that sells developer tools.
Before reading an agent-written change line by line, rank its files by risk and spend close review on the riskiest ones.
Derived by Working Surface from the article; more in the Playbook. Source line: The diff viewer was the right tool for reviewing what your colleague wrote.
7 October 2026: once AI agents do most of the making, checking and choosing become the scarce steps, and teams protect them with author responsibility, automated checks and tools that point to risk.