Each task provides an executable baseline and a visible development environment. The agent saves and tests candidate versions, then submits code that is rerun unchanged on a hidden set in a sealed environment. Domain-native metrics are mapped through fixed scoring definitions to make progress comparable across tasks.
The first release contains 88 tasks across six broad domains, including 35 public and 53 private tasks. RSI-Exam is useful for studying research-agent reliability, transferable improvements, and long-horizon experimentation. Saved trajectories, artifact versions, and resource use make the development process inspectable.

