PROJECTS
Benchmarks, datasets and RL environments for code agents. SWE-rebench Leaderboard Live benchmark for coding agents on fresh real-world GitHub tasks, refreshed every month. Issue dates are tracked against model release dates, so possibly contaminated results are flagged. Reports resolved rate, pass@5, cost and tokens per problem. about · paper SWE-rebench 21,000+ executable SWE tasks from 3,400+ Python repositories, collected by a fully automated pipeline and validated by environment setup and test execution. Built for RL training and evaluation of code agents. dataset · paper · NeurIPS 2025 SWE-rebench V2 32,000+ executable SWE tasks in 20 programming languages from 3,600+ repositories, with pre-built images. An interactive setup agent installs each repository; an ensemble of LLM judges filters out unsound tasks. Plus 126,000+ additional tasks built from pull requests. dataset · PR tasks · code · paper · ICML 2026