Key Features

Tasks come from licensed, private company codebases with real engineering requirements.
Examples cover billing, tax calculations, customer migrations, metering, and analytics.
Evaluations measure the model together with its native coding harness.
Resolution rates average pass-at-one outcomes across eight independent task runs.
Leaderboard comparisons include 95% confidence intervals.
Task environments expose relevant databases, infrastructure, and business tools.
Concise requests require discovering existing business rules and implementation conventions.
Published task analyses show model results and representative engineering challenges.

The benchmark runs model-and-harness combinations in task environments containing the services and tools needed for each workflow. Agents investigate concise requirements and modify existing systems. Resolution rates use pass-at-one results averaged across independent runs, with confidence intervals to make comparisons more informative.


Real-SWE is useful for studying repository navigation, cross-service changes, and failures that simpler coding exercises can miss. Its private origins reduce exposure to publicly available solutions. The published leaderboard and task analyses are accessible, while the sample tasks require an access request.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!