Call for contributors

Write a task from your own field.
Or review one.

RSI-Exam tests whether an AI agent can improve an existing frontier method in a real-world production or scientific setting. Every task is contributed by someone who works on that problem. We are looking for more experts to bring problems from their own fields.

Start with a short proposal and a scoping call. If you are authoring a task, you then build the references and review the resulting rollout. See the task-author process.

Four ways in

Choose how you want to contribute

01 · Task author

Write a task

Your problem, a weak starter that runs, and a stronger method you can run yourself. We do the packaging.

≈ 18 h
02 · Reviewer

Review a task

Read someone else's task and check the metric, the seal and the anchors. A substantive review earns four contribution credits.

≈ 4 h each
03 · Auditor

Audit a rollout

Read one agent trajectory end to end and say whether the score moved for a real reason.

≈ 2 h each
04 · Advisor

Shape the evaluation

Scoring, sealing, difficulty gates, the harness. This is the core-contributor track.

Open
What a task is

A research environment the agent works in

The agent gets a weak method that runs, visible data, and 12 hours. It rewrites the method. A sealed verifier then runs its code on hidden data, and that score is the result. See the 88 tasks in RSI-Exam 0.1.

What we need from youfive things
  • A problem from your own work, with a clear definition of better.
  • A weak starting point that runs end to end today, with real room above it.
  • A stronger method you can run yourself. It sets the bar the agent is measured against.
  • Data that is public, redistributable, or rebuildable offline.
  • A hidden set that differs from the visible one in a way that matters: a distribution shift, a regime change, a tighter budget.
Contribution & authorship

How authorship works

Paper authorship is based on substantial benchmark work. One accepted task qualifies; so do four substantive reviews, or an equivalent mix.

Paper author · 15 credits

  • Named on the paper, on the site, and on the page of every task you wrote.
  • One accepted task, four reviews, or any mix that reaches 15.
  • Authors outside the core group are listed alphabetically.

Core contributor

Listed first, ordered by contribution. Three things count:

  • Contribution. The breadth and depth of work in your contribution record.
  • Independence. How much of the work started with you and ran without us in the loop.
  • Reach. One adopted change to the pipeline or the evaluation, or steady participation in the working discussion.
How contributions are recordedcredit schedule
what you delivercredits
A task you wrote, accepted into the bank15
A task we implemented from your problem, metric and references10
A task that fails a gate after real workUp to 10
A review with a written verdict4
A rollout audit2
A pipeline or evaluation change we adopt2–8
A referralAcknowledged, 0

Each completed benchmark contribution receives credit under the schedule above. We share the contribution record with everyone on it before submission; if a line is wrong, tell us and we fix it.

Typical commitment

For planning, the tables below estimate the work that needs your domain judgment. A first task usually fits into a few evenings a week over three or four weeks; we take care of the engineering, probes and compute.

Writing one task
steptypical time (h)
Scope the problem, check the data licence2–4
Design the visible / hidden split2–5
Get the weak starter running2–4
Run the two reference methods4–10
Set the metric and the guards1–3
Write the instruction1–2
Read the rollout, revise once2–5
Total14–33 · usually ≈ 18

The reference runs set the 0.30 and 0.60 marks on the score scale. A number from a paper does not set them; the method has to run on your data. A second task usually takes about two thirds as long as the first.

Reviewing one task
steptypical time (h)
Read the task as if you had to solve it0.5–1
Judge the metric and the realism0.5–1
Check the leakage report0.5–1
Check the anchors and the scale between them0.5–1.5
Read one agent trajectory0.5–1
Write the verdict0.5
Total3–6 · usually ≈ 4

An automated pass runs the mechanical checks first and hands you the evidence: recomputed anchors, leakage candidates, the trivial-output baseline. You make the calls it cannot.

What we doso you don't have to
  • Packaging. Two isolated images, the sealed verifier, resource budgets, offline build.
  • Hardening. We try to break your grader with trivial and degenerate submissions.
  • Probes. We measure whether a frontier model can write your reference method from memory, and whether a two-parameter grid search reaches it.
  • Rollouts. 12-hour runs of frontier agents on your task, on our compute, with the full trajectories handed back to you.

What you get

  • Paper authorship when your recorded contributions reach the threshold described above.
  • If you author a task, your name appears on its page and the task is cited as yours.
  • Task authors receive frontier-model results on their problem, with full trajectories.
  • Task authors receive a sealed evaluation environment they can reuse in their own work.

What we won't ask for

  • Unpublished or embargoed data.
  • Docker work. You can do it if you want to, it is never the price of entry.
  • More than about 20 hours without checking in with you first.
  • Exclusivity. Your science stays yours and you keep every right to publish it.
Task-author process

From proposal to a published task

  1. You send the form10 min

    Your field, what you want to do, and a rough sketch if you already have a task in mind.

  2. Scoping call45 min

    We settle the metric and the hidden split together. You leave with a written scope.

  3. You build the baseline and the referencesmain commitment

    If the gap between them turns out to be noise, we stop here and the work is still credited. This is the honest failure mode and we put it early on purpose.

  4. We package it, you read the rollout2–5 h

    We run the probes and a 12-hour agent rollout. You check that the task is hard for the reason you meant, and revise once. A peer reviews it. If it passes review, it ships with your name on it.

Questions we get

Before you fill in the form

Can I build a task out of my own paper?

Yes, and that is the easiest case. You need to be able to run the code today, and the data has to be redistributable or rebuildable offline.

Do I need a GPU?

Most tasks in the bank are CPU-only. GPU tasks are accepted with a reason, so raise it on the call. All agent rollouts run on our compute.

What if a frontier agent just solves my task?

Then it does not ship as it stands, and the work is still credited. Calibration catches this before the expensive part. Often the fix is a harder hidden split rather than a new task.

I know someone better suited than me.

Send them the form and ask them to mention your name when they submit it. We name referrers in the acknowledgements. A referral on its own does not earn credit, since authorship tracks work on the benchmark itself.

How long do I have?

To join the paper accompanying RSI-Exam 1.0, complete your contribution and have its credits confirmed by September 15, 2026. The RSI-Exam 2.0 deadline is December 15, 2026. A first task usually takes three to four weeks at a few hours a week; we confirm the schedule on the scoping call.

Next step

Ten minutes on the form

If you already have a task in mind, bring one sentence about the problem and one about the metric. The call does the rest.