Key Features

It evaluates improvements to executable methods or harnesses around a frozen model.
Agents inherit a functioning but improvable executable baseline.
The final deliverable is executable code rather than a reported score or prediction file.
The submitted artifact is rerun unchanged on a hidden evaluation set.
Visible experimentation and sealed hidden evaluation use isolated environments.
RSI-Exam 0.1 contains 88 tasks: 35 public and 53 private.
Fixed scoring definitions retain domain-native metrics while normalizing measured improvement.
Saved versions, experiment descriptions, resource use, and outcomes support analysis.

Each task provides an executable baseline and a visible development environment. The agent saves and tests candidate versions, then submits code that is rerun unchanged on a hidden set in a sealed environment. Domain-native metrics are mapped through fixed scoring definitions to make progress comparable across tasks.


The first release contains 88 tasks across six broad domains, including 35 public and 53 private tasks. RSI-Exam is useful for studying research-agent reliability, transferable improvements, and long-horizon experimentation. Saved trajectories, artifact versions, and resource use make the development process inspectable.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!