RSI-Exam tests whether an AI agent can improve an existing frontier method in a real-world production or scientific setting. Every task is contributed by someone who works on that problem. We are looking for more experts to bring problems from their own fields.
Start with a short proposal and a scoping call. If you are authoring a task, you then build the references and review the resulting rollout. See the task-author process.
Your problem, a weak starter that runs, and a stronger method you can run yourself. We do the packaging.
≈ 18 hRead someone else's task and check the metric, the seal and the anchors. A substantive review earns four contribution credits.
≈ 4 h eachRead one agent trajectory end to end and say whether the score moved for a real reason.
≈ 2 h eachScoring, sealing, difficulty gates, the harness. This is the core-contributor track.
OpenThe agent gets a weak method that runs, visible data, and 12 hours. It rewrites the method. A sealed verifier then runs its code on hidden data, and that score is the result. See the 88 tasks in RSI-Exam 0.1.
Paper authorship is based on substantial benchmark work. One accepted task qualifies; so do four substantive reviews, or an equivalent mix.
Listed first, ordered by contribution. Three things count:
| what you deliver | credits |
|---|---|
| A task you wrote, accepted into the bank | 15 |
| A task we implemented from your problem, metric and references | 10 |
| A task that fails a gate after real work | Up to 10 |
| A review with a written verdict | 4 |
| A rollout audit | 2 |
| A pipeline or evaluation change we adopt | 2–8 |
| A referral | Acknowledged, 0 |
Each completed benchmark contribution receives credit under the schedule above. We share the contribution record with everyone on it before submission; if a line is wrong, tell us and we fix it.
For planning, the tables below estimate the work that needs your domain judgment. A first task usually fits into a few evenings a week over three or four weeks; we take care of the engineering, probes and compute.
| step | typical time (h) |
|---|---|
| Scope the problem, check the data licence | 2–4 |
| Design the visible / hidden split | 2–5 |
| Get the weak starter running | 2–4 |
| Run the two reference methods | 4–10 |
| Set the metric and the guards | 1–3 |
| Write the instruction | 1–2 |
| Read the rollout, revise once | 2–5 |
| Total | 14–33 · usually ≈ 18 |
The reference runs set the 0.30 and 0.60 marks on the score scale. A number from a paper does not set them; the method has to run on your data. A second task usually takes about two thirds as long as the first.
| step | typical time (h) |
|---|---|
| Read the task as if you had to solve it | 0.5–1 |
| Judge the metric and the realism | 0.5–1 |
| Check the leakage report | 0.5–1 |
| Check the anchors and the scale between them | 0.5–1.5 |
| Read one agent trajectory | 0.5–1 |
| Write the verdict | 0.5 |
| Total | 3–6 · usually ≈ 4 |
An automated pass runs the mechanical checks first and hands you the evidence: recomputed anchors, leakage candidates, the trivial-output baseline. You make the calls it cannot.
Your field, what you want to do, and a rough sketch if you already have a task in mind.
We settle the metric and the hidden split together. You leave with a written scope.
If the gap between them turns out to be noise, we stop here and the work is still credited. This is the honest failure mode and we put it early on purpose.
We run the probes and a 12-hour agent rollout. You check that the task is hard for the reason you meant, and revise once. A peer reviews it. If it passes review, it ships with your name on it.
Yes, and that is the easiest case. You need to be able to run the code today, and the data has to be redistributable or rebuildable offline.
Most tasks in the bank are CPU-only. GPU tasks are accepted with a reason, so raise it on the call. All agent rollouts run on our compute.
Then it does not ship as it stands, and the work is still credited. Calibration catches this before the expensive part. Often the fix is a harder hidden split rather than a new task.
Send them the form and ask them to mention your name when they submit it. We name referrers in the acknowledgements. A referral on its own does not earn credit, since authorship tracks work on the benchmark itself.
To join the paper accompanying RSI-Exam 1.0, complete your contribution and have its credits confirmed by September 15, 2026. The RSI-Exam 2.0 deadline is December 15, 2026. A first task usually takes three to four weeks at a few hours a week; we confirm the schedule on the scoping call.
If you already have a task in mind, bring one sentence about the problem and one about the metric. The call does the rest.