The benchmark runs model-and-harness combinations in task environments containing the services and tools needed for each workflow. Agents investigate concise requirements and modify existing systems. Resolution rates use pass-at-one results averaged across independent runs, with confidence intervals to make comparisons more informative.
Real-SWE is useful for studying repository navigation, cross-service changes, and failures that simpler coding exercises can miss. Its private origins reduce exposure to publicly available solutions. The published leaderboard and task analyses are accessible, while the sample tasks require an access request.

