Search Results

Documents authored by Ghaleb, Taher A.


Document
Emerging Results, Vision & Reflection Track Paper
AI-to-AI Code Reviews of GitHub Pull Requests

Authors: Niruthiha Selvanayagam and Taher A. Ghaleb

Published in: LIPIcs, Volume 394, 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)


Abstract
AI coding agents are increasingly integrated into software development workflows, operating on both sides of the pull request (PR) process: AI authoring agents, which create or modify PRs, and AI reviewers, which evaluate them. This creates a closed loop where one AI coding agent reviews contributions of another AI coding agent. In this paper, we construct an AI-to-AI code review dataset by linking AI-authored pull requests with AI-attributed review events from CodAGE, a public dataset of coding agent–generated GitHub events. Our dataset contains νm{248641} PRs across 12 AI coding agents with at least one AI review, including νm{45269} reviewed by a different AI product and νm{208145} by the same product. We observe that cross-product AI-to-AI code review occurs in only about 1.6% of identified agent-authored PRs but is substantial in absolute terms: 45k PRs written by one identifiable AI product and reviewed by another. This activity grows by more than two orders of magnitude from 2025-Q1 to 2025-Q3. We measure reviewer behavior using CodeRabbit comment categories, per-PR comment volume, and time to first review, and find that it varies across author–reviewer pairs. For example, Claude-Code PRs receive more refactor comments from CodeRabbit than Copilot PRs (35.0% vs. 10.5%), a difference that may stem from PRs themselves rather than the reviewer. For three of four dual-role reviewers, mean comments per PR were 58-65% higher in the same-product group, though effect sizes were small or negligible and the difference was concentrated in the upper tail. Median time from PR creation to first AI review was 1.2 minutes for cross-product pairs and 4.7 minutes for same-product pairs, reflecting which reviewer bots are in each group rather than the product pairing. Overall, our large-scale characterization shows that closed-loop AI-to-AI code review is on the rise but remains a minority phenomenon, with review output varying across authoring-agent groups and author-reviewer configurations.

Cite as

Niruthiha Selvanayagam and Taher A. Ghaleb. AI-to-AI Code Reviews of GitHub Pull Requests. In 20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026). Leibniz International Proceedings in Informatics (LIPIcs), Volume 394, pp. 75:1-75:15, Schloss Dagstuhl – Leibniz-Zentrum für Informatik (2026)


Copy BibTex To Clipboard

@InProceedings{selvanayagam_et_al:LIPIcs.ESEM.2026.75,
  author =	{Selvanayagam, Niruthiha and Ghaleb, Taher A.},
  title =	{{AI-to-AI Code Reviews of GitHub Pull Requests}},
  booktitle =	{20th International Symposium on Empirical Software Engineering and Measurement (ESEM 2026)},
  pages =	{75:1--75:15},
  series =	{Leibniz International Proceedings in Informatics (LIPIcs)},
  ISBN =	{978-3-95977-450-5},
  ISSN =	{1868-8969},
  year =	{2026},
  volume =	{394},
  editor =	{Feldt, Robert and Paasivaara, Maria and Mendez, Daniel and Wagner, Stefan and Bar\'{o}n, Marvin Mu\~{n}oz},
  publisher =	{Schloss Dagstuhl -- Leibniz-Zentrum f{\"u}r Informatik},
  address =	{Dagstuhl, Germany},
  URL =		{https://drops.dagstuhl.de/entities/document/10.4230/LIPIcs.ESEM.2026.75},
  URN =		{urn:nbn:de:0030-drops-280430},
  doi =		{10.4230/LIPIcs.ESEM.2026.75},
  annote =	{Keywords: AI coding agents, AI code review, closed-loop AI, Pull requests, Mining software repositories, GitHub events}
}

Any Issues?
X

Feedback on the Current Page

CAPTCHA

Thanks for your feedback!

Feedback submitted to Dagstuhl Publishing

Could not send message

Please try again later or send an E-mail