Reviewing AI code is a trust problem, not a diff problem, says a JetBrains and Lund study
JetBrains Research and researchers from Lund University have written up a study of how tools for reviewing AI-generated code should work. The paper will be presented at the Empirical Software Engineering International Week (ESEIW) 2026 in October. Its central claim is worth taking seriously because it describes a problem every team using coding agents already has.
Why the usual review instincts fail. When a colleague's code is reviewed, invisible scaffolding helps: the reviewer knows how senior the author is, which parts they rush, and can ask them in Slack. None of that exists for a model. Worse, a language model presents every line with the same apparent confidence, whatever its real uncertainty — the authentication logic and the shaky database migration read the same. The rational response is to read every line, which does not scale as agents produce larger change sets.
The reframe. The authors argue that reviewing large multi-file changes from a model is not a diffing problem but a trust-calibration problem: the skill of spending review effort in proportion to the risk of each segment when the author cannot be asked how sure it was. The diff viewer assumed the job was to understand what changed; that, they write, may not be the right tool for what an agent wrote.

What they propose. A three-level workflow following Shneiderman's information-visualisation principle — overview first, then zoom and filter, then details on demand: the reviewer forms high-level hypotheses and drills down selectively to test them. The design came from four workshops with 17 practitioners and a follow-up survey of 43 software professionals.
Where tools are today. The authors note partial counterparts: prose walkthroughs of pull requests in CodeRabbit, multiple reviewer agents tagging findings by severity in Claude Code, and Graphite's stacked pull requests as the nearest analogue to splitting a change into reviewable chunks. Their point is that no single tool yet exposes risk and confidence at the level where reviewers actually put their attention.
Why it matters. For teams, the practical lesson is cheap to apply even without new tools: ask the agent to split its change and say where it was unsure, and review those parts first.