SELECTED PAPERS

Selected papers on benchmarks, datasets and RL for software engineering agents. SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents NeurIPS 2025 Datasets and Benchmarks · first author (145 citations) Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning arXiv 2025 (38 citations) SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale ICML 2026 · first author (18 citations) Scaling Data Collection for Training Software Engineering Agents Nebius blog 2024 · first author (13 citations) Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents ICML 2025 (12 citations) All papers → Google Scholar (citations as of Sep 2026)