Project leader · applied AI
An open-weights agent on a single GPU
Interim version. Scores, configurations and code are withheld until the competition's code deadline in November; the full write-up with all measurements will be published here then.
ARC-AGI-3 is a benchmark of interactive puzzle games that an agent has never seen. There are no instructions, no stated goal and no reward signal: the agent sees a small coloured grid, chooses moves, and must work out both the rules and the objective as it plays. A level is scored by how close the agent's move count comes to a human's, so wasted moves are punished hard.
The competition runs on fixed hardware — one GPU, no internet, a time limit per run — which makes it a test of engineering judgement as much as of model ability.
We ran the project as a measurement programme rather than a sequence of ideas. Every change to the agent was compared against repeated runs of an unchanged baseline, with its success criterion written down before the run. A deterministic local replica of the game engine let us replay recorded games and test "what if" questions without spending any compute.
Over the season we tested a long list of improvements to the agent's instructions and tooling: goal checks, scripted exploration, injected rules and facts, voting, search, and several forms of an internal "world model". We also measured where points can come from at all, before building anything, by recomputing the score of recorded games under counterfactual assumptions.
Identical runs of the same agent differ a lot — single games flip between failure and full success. Until that spread is known, every comparison is a guess. Most claimed gains in this competition are smaller than the spread of one build, ours included.
A large share of the levels an agent completes is already scored at the maximum, so making it more efficient on levels it can already solve is capped. Reaching one level deeper in each game is worth more than everything else combined. That told us which ideas could matter before we built them.
Text added to the prompt changed almost nothing: the model would repeat an injected rule and still not act on it. What the strongest open solutions had in common was not a better prompt but more of the game's own history kept in front of the model. Under a fixed context budget, every added instruction displaces a piece of the agent's own experience.
For weeks we tuned the agent while the inference server underneath it was starved — most requests were waiting in a queue, and we misread that as slow generation. The order of work we now recommend for any fixed-hardware AI system: model packaging, serving engine and version, how the memory cache is sized and spent, how time is shared between tasks — and only then the agent's prompt and logic. We did it in reverse.
Once the serving stack was right, our experiments on it pointed the same way: configurations that reduced the total amount of reasoning the server could produce lost, and nothing we varied made each individual move smarter. Building tools to measure the server directly — without playing games — turned out to be the cheapest way to make decisions.
A stronger model we compared against played no better per move than ours; its whole advantage was speed. We first attributed that to its infrastructure instead of asking what infrastructure we could run ourselves. The answer to that question turned out to be most of the gap.
The full paper with all measurements, the score budget, the serving experiments and the code.