Sergei Makarov

Sergei Makarov

Project leader · applied AI

Research notes · ARC Prize 2026 · ARC-AGI-3

The server was the ceiling

An open-weights agent on a single GPU

Interim version. Scores, configurations and code are withheld until the competition's code deadline in November; the full write-up with all measurements will be published here then.

The task

ARC-AGI-3 is a benchmark of interactive puzzle games that an agent has never seen. There are no instructions, no stated goal and no reward signal: the agent sees a small coloured grid, chooses moves, and must work out both the rules and the objective as it plays. A level is scored by how close the agent's move count comes to a human's, so wasted moves are punished hard.

The competition runs on fixed hardware — one GPU, no internet, a time limit per run — which makes it a test of engineering judgement as much as of model ability.

What we did

We ran the project as a measurement programme rather than a sequence of ideas. Every change to the agent was compared against repeated runs of an unchanged baseline, with its success criterion written down before the run. A deterministic local replica of the game engine let us replay recorded games and test "what if" questions without spending any compute.

Over the season we tested a long list of improvements to the agent's instructions and tooling: goal checks, scripted exploration, injected rules and facts, voting, search, and several forms of an internal "world model". We also measured where points can come from at all, before building anything, by recomputing the score of recorded games under counterfactual assumptions.

What we learned

1. Measure the noise first

Identical runs of the same agent differ a lot — single games flip between failure and full success. Until that spread is known, every comparison is a guess. Most claimed gains in this competition are smaller than the spread of one build, ours included.

2. Compute the score budget before choosing a lever

A large share of the levels an agent completes is already scored at the maximum, so making it more efficient on levels it can already solve is capped. Reaching one level deeper in each game is worth more than everything else combined. That told us which ideas could matter before we built them.

3. Adding instructions did not help; giving the agent its memory back did

Text added to the prompt changed almost nothing: the model would repeat an injected rule and still not act on it. What the strongest open solutions had in common was not a better prompt but more of the game's own history kept in front of the model. Under a fixed context budget, every added instruction displaces a piece of the agent's own experience.

4. Audit the whole stack before the top layer

For weeks we tuned the agent while the inference server underneath it was starved — most requests were waiting in a queue, and we misread that as slow generation. The order of work we now recommend for any fixed-hardware AI system: model packaging, serving engine and version, how the memory cache is sized and spent, how time is shared between tasks — and only then the agent's prompt and logic. We did it in reverse.

5. Throughput is the score

Once the serving stack was right, our experiments on it pointed the same way: configurations that reduced the total amount of reasoning the server could produce lost, and nothing we varied made each individual move smarter. Building tools to measure the server directly — without playing games — turned out to be the cheapest way to make decisions.

6. Read a comparison for what it actually differs in

A stronger model we compared against played no better per move than ours; its whole advantage was speed. We first attributed that to its infrastructure instead of asking what infrastructure we could run ourselves. The answer to that question turned out to be most of the gap.

How the project was run

Coming in November

The full paper with all measurements, the score budget, the serving experiments and the code.