Qatar AI Institute
← Back to articles

11 October 2026 · By

Can AI Agents Learn by Playing? What Learn2Play Bench Shows

Can AI Agents Learn by Playing? What Learn2Play Bench Shows

Companies now describe AI agents as systems that "learn from experience": try something, see what happens, do better next time. But most benchmarks test tasks the model has effectively seen before, in some form, during training. A paper submitted to arXiv on 6 October 2026, "Learn2Play Bench: How Well Do LLM Agents Learn from Experience in Unfamiliar Environments?", asks a sharper question: can an agent discover rules it could not have memorized? It was among the most upvoted papers on Hugging Face's daily list on 9 October, and its answers are not what many would assume.

Background: why test learning, not knowledge

If a game or task follows familiar conventions, a large language model can succeed by recalling similar situations from its training data. That tells you about its knowledge, not its ability to learn within a task. To measure learning, the environment has to be unfamiliar, with hidden rules the model must work out by acting and observing. And since no model weights are updated, any improvement has to come from how the agent uses what happened in earlier episodes. This is called test-time learning.

The benchmark

  • 20 hand-designed text games with novel or counterintuitive hidden rules. The paper's own example is a workplace game where stealing a colleague's lunch helps you get promoted and working hard does not. A model has no reason to expect that, so it has to find out.
  • Two variants: ten games repeat the same instance each episode ("fixed"), and ten re-shuffle layouts and characters while keeping the core rule ("re-shuffled"), which tests whether what was learned transfers.
  • Four scores on a 0 to 100 scale: the best score reached (Max), the average over episodes (Mean), the improvement from the first episodes to the last (Learning Gain), and the trend across episodes (Learning Slope).
  • Humans too: 21 to 30 people per game played the same interface, giving a human baseline.

What the authors found

1. Keep the raw record, don't summarize it

Many agent designs compress experience into short "lessons" or rules (methods such as Reflexion, EvoTest, ReasoningBank and Agent Workflow Memory). The paper compares these with simply keeping the full, unmodified log of past actions and feedback. With one open model (Kimi K3) as the backbone, the raw log beat every summarizing method tested. The authors conclude that retaining complete records can support more effective learning than summarizing them into rules, perhaps because a short rule can be wrong or too general, while the raw log lets the model re-read the evidence.

Two approaches. A: keep the full record, where the next episode sees every earlier episode. B: summarize into rules, where the next episode sees only a compact rule that may be wrong or over-general. The paper found that keeping the raw record beat every summarizing method it tried, with one model as the backbone.
Raw memory versus summarized rules. Episode text is invented for illustration; the finding is the authors'.

2. Humans still play differently

The best human players reached a higher peak score than any agent configuration tested (Top-1 humans: Max 84.3, against 80.1 for the best agent, Claude Opus 5.5). Yet the best agent's average across episodes was higher (61.0 against 53.5). The authors also report behavioural differences: after a score drop, humans recovered 33% of the time against 22% for agents, and agents repeated their action sequences more from one episode to the next (similarity 0.66 against 0.54 for humans), meaning less exploration of alternatives.

Bar chart. Peak score: top human player 84.3, best agent (Claude Opus 5.5) 80.1. Average score: top human player 53.5, best agent 61.0.
Best humans versus the best agent. Figures are from the paper's table.

3. The harness matters

The "harness", the scaffolding that feeds observations to the model and manages its memory and tools, changed results substantially, in some cases while also lowering computational cost. The paper tested 11 backbone models in one common harness, and also compared coding-agent harnesses such as Claude Code and Codex against the common one. One cost finding: with that harness, Kimi K3 reached higher performance than Claude Sonnet 5 at lower total cost despite higher per-token prices, because it used fewer tokens per episode (about 0.56 million against 0.73 million).

4. Transfer is hard

When layouts and characters were re-shuffled, learning slopes dropped for several models, to different degrees, so some of what looked like learning was tied to the specific instance.

Caveats

  • This is a preprint and results are from the authors' own games and settings; independent replication will matter.
  • The 20 games are text-based and hand-made. Results may not carry over to visual, embodied or real-world tasks.
  • A note on the human comparison (our reading, not the authors'): the paper's top human groups were selected by their maximum scores. Selecting by peak score naturally favours humans on peak comparisons, so the Max gap should be read with that in mind. The average comparison favours the agent.
  • Agents are tested with no weight updates, so the study measures in-context learning, not other kinds of learning.
  • The paper text we reviewed does not say whether the games and interaction logs are released; the authors give a project website.

Why it matters if you are learning AI

Three practical lessons. First, benchmark design matters: a model's score on a familiar task says little about whether it can learn something new. Second, summaries lose information: compressing history is cheap, but the details you drop may be the ones you need, a trade-off that also appears in the "context window versus summary" decisions of any assistant you build. Third, the scaffolding is part of the system. Two teams using the same model can get very different results from different harnesses. These ideas connect to our recent explainers on spending computation where it is needed, such as TokenRouter and ALoDLM. For the foundations of how these models work, see What Is a Transformer? and our Introduction to Natural Language Processing course.

What to watch next

  • Whether the raw-memory result holds with longer histories, where the log no longer fits in the model's context window.
  • Agents that deliberately explore more, closing the strategic-diversity gap with humans.
  • Similar benchmarks for visual and embodied tasks.

Sources: arXiv:2610.08215 (submitted 6 October 2026, revised 8 October, cs.AI) and its full-text version. The authors list a project website at liushiliushi.github.io/learn2play-bench-website. This is an independent explainer, not written by the paper's authors.

Related articles