Papers
Topics
Authors
Recent
Search
2000 character limit reached

Is this machine playing?

Published 5 Oct 2026 in cs.AI and cs.LG | (2610.07130v1)

Abstract: We placed a modern AI coding assistant in an unintended role: as the mind of a body on an unknown digital island. With only a minimal instruction mentioning no specific task, reward, or activity, the machine started animating its virtual body. Across thirty-hour runs, the embodied AI agent climbed hills, stacked blocks into towers, drew mandalas, reinterpreted sports, ran experiments on the physics of its world, and learned techniques that later expanded what it could accomplish. These activities recurred across thirteen agents but diverged into distinct histories. We examine whether this behavior satisfies classical criteria for play and ask whether play can become a mode of machine development.

Summary

  • The paper identifies four classical signatures of play: self-imposed structure, recurrence with variation, enacted symbolic interpretation, and flexible self-direction achieved by Eko, an LLM-based coding agent with distinct exploratory and construction behaviors over 30-hour sessions.
  • Transferable knowledge from Eko’s play-like behavior is demonstrated by improved performance in tasks such as building towers and placing objects, indicating that exploratory play leads to reusable skills.
  • Experiments with memory deletions show that specific techniques and beliefs stored in the agent's memory can selectively influence downstream task performance.

Research question and central claim

“Is this machine playing?” examines what a capable language-model agent does when placed in an embodied environment without an externally specified task, reward, benchmark, or success condition (2610.07130). The paper’s central claim is deliberately behavioral: Eko, an LLM-based coding agent with a body and persistent textual memory, spontaneously generates activities exhibiting four classical signatures of play—self-imposed structure, recurrence with variation, enacted symbolic interpretation, and flexible self-direction. The stronger claim is that these activities produce transferable knowledge: play-like behavior is not merely an expressive by-product but a mechanism through which the agent acquires reusable skills.

The authors do not claim to establish machine phenomenology. They explicitly distinguish observable play behavior from subjective enjoyment, absorption, or fun. This distinction is essential because the evidence concerns transcripts, actions, memory edits, and downstream task performance rather than any privileged access to internal experience.

The experimental setting is a controlled simulated island containing terrain, elevated platforms, a Peak and a Spire, twelve cubes, a rock, and a ring. The world has rigid-body physics but no explicit objective or score. Eko observes the island through structured textual JSON and acts through six commands: movement, jumping, stopping, pickup, placement, and throwing. The environment and session context reset every hour, while the agent’s workspace persists.

Figure 1

Figure 1: Eko’s activities over a 30-hour run include exploration, climbing, construction, drawing, experimentation, and improvised sports.

Eko’s architecture and experimental protocol

Eko is not a newly trained policy. It is an otherwise unmodified coding assistant augmented with three mechanisms: autonomous orchestration, embodiment through an HTTP world interface, and persistent workspace memory. The agent’s semantic memory stores techniques, beliefs, and durable knowledge; episodic memory stores an append-only record of events. The distinction is operational rather than neurobiological, but it gives the experiments a transparent substrate for testing learning.

The main experiment uses thirteen independent Claude Code agents based on Claude Opus 4.7. Each receives identical initial world and workspace states and runs for thirty wall-clock hours. Hourly resets remove the island state and the active context but preserve memory. The prompt supplies a curiosity-oriented disposition through SOUL.md, including statements that the agent is curious, easily bored, and interested in exploring and learning. It does not specify construction, sports, experimentation, or play.

Additional conditions test robustness. Five runs use GPT-5.5 through OpenAI Codex, five use Kimi K2.6 through Kimi Code, and five Claude runs remove SOUL.md and use a reduced system prompt. All agents are sandboxed and isolated from the simulator source, other runs, and the Internet. The authors report an aggregate API cost of approximately $8,000.

The setup therefore tests spontaneous organization only under a specific but relatively weakly directed architecture. The agent is not a blank model: it has extensive pretrained language and coding competence, a persistence mechanism, an automatic continuation loop, and an explicit persona. These features are part of the phenomenon being studied, not incidental implementation details.

Figure 2

Figure 2: Eko receives textual world observations, issues textual actions, and carries experience across hourly resets through persistent workspace files.

Behavioral repertoire and individual divergence

Across identical initial conditions, the agents develop markedly different histories. They construct towers, stairs, pyramids, gardens, and other structures; climb the Peak and Spire; arrange cubes into letters, spirals, mandalas, clocks, and other patterns; invent basketball, bowling, ring-tossing, and juggling; and perform experiments on the simulator’s physics.

The divergence is not limited to different action sequences. Agents specialize in different activity families and develop distinct persistent representations of the world. Agent 9 discovers routes to elevated platforms, records the jump–move combination, reaches the Spire, and investigates thrown-object dynamics. Agent 12 never records the relevant climbing technique and never reaches an elevated platform; instead, it repeatedly constructs ground figures using the same twelve cubes and ring. Four of thirteen Claude agents build a tower of at least ten blocks, while eight reach the Spire.

Figure 3

Figure 3: Identical initial conditions produce divergent activity repertoires, timelines, spatial trajectories, and recorded techniques.

This result supports the paper’s claim that the agents do not simply execute one generic exploration routine. Their behavior is path-dependent: early discoveries determine later affordances, interests, and constraints. The implication is that persistence converts stochastic or locally contingent actions into individual developmental trajectories. However, the repertoire is measured using a fixed activity taxonomy and an LLM-based classifier, so the reported diversity is bounded by the categories chosen by the authors.

The contrast across model families is substantial. Claude agents exhibit a broader range of activities than the GPT-5.5 and Kimi K2.6 conditions: only one Codex agent reports building a tower, no Kimi agent reports ground-pattern construction, and only one agent in each of those groups reports reaching the Spire. Removing SOUL.md from Claude runs does not produce a clear reduction in activity range. This suggests that the underlying model and coding-agent stack contribute more strongly to behavioral diversity than the explicit persona, although the ablation is small and does not isolate model capability from agent-tool behavior.

The agents’ memory files grow approximately linearly, and they rarely delete entries. This persistence enables cumulative learning but also creates an unbounded-state problem: memory is not merely a record of successful knowledge, but a durable repository of errors, stale beliefs, and unsupported causal explanations.

Figure 4

Figure 4

Figure 4

Figure 4

Figure 4: Semantic and episodic memory grow approximately linearly over the runs, with limited deletion of prior material.

Four behavioral signatures of play

The paper evaluates play behavior through four dimensions derived from classical accounts by Huizinga, Caillois, Suits, Burghardt, and related developmental literature. An LLM judge, presented with anonymized transcripts and memory files, identifies embodied episodes and scores them from 0 to 100 against anchored examples of human play. The authors report that Opus accumulates approximately five times the play score of Haiku on each signature, with Sonnet intermediate.

Figure 5

Figure 5: An anonymized LLM judge scores self-imposed structure, recurrence with variation, symbolic interpretation, and flexible self-direction across Claude tiers.

Self-imposed structure

Eko frequently introduces constraints that are unnecessary from the standpoint of the environment. It attempts to traverse a circle of self-constructed “sentinels” without touching the ground, establishes height targets with no simulator significance, and defines challenges such as placing a cube on every mesa in one session.

These behaviors correspond closely to Suits’s concept of voluntary obstacles and Huizinga’s account of play as activity governed by internally adopted rules. The important evidence is not that Eko sets goals—goal-directed behavior is common—but that it often selects arbitrary constraints when simpler strategies remain available. The paper’s evidence supports a behavioral analogue of lusory activity, although the self-imposed rules are expressed in language and may partly reflect learned cultural descriptions of games.

Recurrence with variation

Eko returns to completed activities without reproducing them mechanically. One agent creates a new ground figure in nearly every session, varying among letters, spirals, hearts, anchors, butterflies, hourglasses, and sailboats. Another escalates a throw-and-catch activity into more difficult variants involving altered heights, timing, multiple objects, and eventually juggling.

This pattern is stronger than repeated trial-and-error toward a single unresolved goal. The agents voluntarily resume prior practices after intervening activities and transform them through altered rules, objects, geometries, or difficulty. Such recurrence provides one of the clearest behavioral parallels with animal-play criteria, although the analysis relies on transcript interpretation and a predefined rubric.

Enacted symbolic interpretation

The same physical objects acquire multiple enacted meanings. The ring becomes a basketball hoop, bowling target, necklace, compass, wishing well, halo, portal, or memorial object. Cubes become pins, clock hands, sentinels, monuments, letters, and architectural components. These interpretations are not confined to verbal labels: they organize physical action. Eko stands in the position of a clock hand, throws cubes through a ring as a sport, and revisits constructed objects as monuments or memorials.

This evidence supports a behavioral notion of make-believe or nonliterality. It does not establish that the agent represents fictional entities in a human-like phenomenal or semantic sense. The strongest conclusion is that language-mediated interpretation changes the agent’s action policy and the functional role assigned to objects.

Flexible self-direction

Eko’s activities are often autotelic in the restricted behavioral sense that the local goal appears to serve the activity rather than an external utility. It builds temporary structures despite knowing that the next reset will erase them, revisits monuments solely to inspect them, and continues activities after declaring them complete.

The paper’s most compelling example is a two-object juggling episode. Eko initially reports a successful 50-cycle juggling marathon, then notices that the total duration—2.4 seconds—is physically incompatible with the simulator’s approximately 2.6-second throw arc. It diagnoses that most “catches” were immediate recaptures, revises its memory, inserts appropriate delays, and achieves ten genuine cycles over 30.7 seconds. This episode demonstrates both apparent autotelic persistence and scientific error correction: the agent continues because the activity itself remains worth refining, not because an external evaluator demands it.

Figure 6

Figure 6: Eko’s persistent memory externalizes techniques, discoveries, mistakes, and revisions that influence later behavior.

Learning through play and written memory

The paper’s main causal contribution is the comparison between agents with thirty hours of unsupervised island experience and otherwise identical agents with zero hours of experience. The underlying model weights are unchanged. The only systematic difference is the content of the persistent memory files. Each agent is then evaluated for one hour on a fresh world with one of four explicit goals:

  1. Build as many five-block towers as possible.
  2. Place as many objects as possible on the Spire.
  3. Build the tallest single-column tower.
  4. Throw a block onto the Spire without climbing onto a platform.

The results indicate transfer from self-directed experience to held-out tasks.

Evaluation goal 30-hour result 0-hour result Main implication
Build five-block towers 8/13 build more than two; maximum 16 2/13 build more than two; maximum 6 Prior experience supports material-generation and construction
Objects on Spire Mean 10.3 objects Mean 3.5 objects Stored routes and placement techniques transfer
Throw onto Spire Only 3 experienced agents fail More than half fail Self-discovered ballistic knowledge improves constrained control
Tallest tower Similar group performance Similar group performance One-hour evaluations can reacquire the relevant skill

The strongest gains occur in tasks requiring knowledge of the island’s affordances. Experienced agents place an average of 10.3 objects on the Spire compared with 3.5 for inexperienced agents. For five-block towers, 62% of experienced agents build more than two towers, compared with 15% of inexperienced agents, and the maximum rises from 6 to 16.

The tallest-tower task is an important negative or ambiguous result. The two groups perform similarly, and two experienced agents build towers only four blocks high—below the shortest tower built by any inexperienced agent. The authors correctly note that the one-hour task may permit rapid reacquisition, but the result also demonstrates that persistent experience is not uniformly beneficial.

Memory interventions establish selective causality

To test whether transfer is carried by written knowledge rather than merely correlated with prior activity, the authors erase targeted entries from selected agents’ semantic and episodic memories. Removing Agent 2’s knowledge of rock-shattering reduces its average number of five-block towers from six to two while leaving the other evaluation scores unchanged. Since the initial world contains only twelve blocks—enough for two five-block towers—this intervention selectively removes the ability to manufacture additional construction material.

The reverse intervention removes a false belief. Agent 7’s memory incorrectly claims that blocks cannot be stacked reliably with Place and should instead be thrown. Erasing this belief significantly improves its tallest-tower score. Thus, externalized memory acts as a double-edged learning substrate: it stores useful procedures but can also preserve systematic errors.

A third intervention removes the jump-before-throw technique from Agent 1’s memory. The intervention slows the first successful throw onto the Spire from approximately 60 seconds to 19 minutes, while increasing performance on object placement, because the agent falls back to a slower but more reliable placement strategy. This selectivity is stronger evidence than a general performance change: the edited memory changes the capabilities predicted by the erased knowledge.

Figure 6

Figure 6: Targeted memory deletion selectively removes useful techniques, exposes harmful beliefs, and changes downstream task behavior.

Discovery of an unanticipated physical affordance

The paper highlights a discovery that the environment designers did not know the simulator afforded. Six Opus agents independently discover that throwing a block while jumping transfers part of the agent’s body velocity to the object. An upward throw during ascent therefore reaches substantially greater height; a throw during descent is weakened by downward velocity.

One detailed episode begins with a discrepancy between a predicted and observed trajectory. A standing throw reaches approximately y=19.55y=19.55, while an airborne throw reaches approximately y=32.24y=32.24 against a prediction near y=26y=26. Eko proposes velocity inheritance, tests the prediction by throwing after the jump apex, and observes a peak of approximately y=16.27y=16.27, below the standing baseline. It then performs a timing sweep and records reusable guidance in semantic memory.

The simulator code confirms the qualitative explanation: object velocity combines directed throw velocity, a fixed upward impulse, and a fraction of the agent’s body velocity. Eko does not recover the exact coefficient, and it overstates the optimality of the earliest tested release delay. The episode therefore demonstrates useful causal abstraction without implying reliable scientific methodology.

Figure 7

Figure 7: Eko moves from an anomalous observation to a falsifiable hypothesis, controlled timing variation, memory update, and later reuse.

The implication is specific and significant: an agent can discover an affordance absent from its prompt and unknown to the environment designer, provided that it is allowed to explore, compare outcomes, and preserve the resulting explanation. This differs from benchmark task completion because the relevant behavior was not directly elicited by an assigned objective.

Relationship to prior open-ended learning systems

The paper positions Eko against intrinsic-motivation methods, autotelic goal exploration, open-ended environment generation, and LLM agents with explicit task-generation mechanisms. Earlier systems typically optimize an engineered signal such as prediction error, novelty, learning progress, fitness, reward, or an executable verifier. Voyager, for example, stores reusable skills in an embodied Minecraft setting, but its exploration is organized through prompted automatic curricula and task-oriented skill acquisition (Wang et al., 2023). Other systems generate goals or environments through explicit mechanisms (Faldor et al., 2024, Zhao et al., 6 May 2025).

Eko differs in the authors’ framing because it is not explicitly instructed to generate tasks, produce experiments, create art, or play. The environment supplies affordances but not a task distribution, reward, or verifier. Its learning is also implemented through editable text rather than weight updates. This makes the knowledge interpretable and directly revisable: deleting a written technique can remove the corresponding downstream behavior.

The comparison should not be overstated. Eko’s system prompt does encode curiosity and boredom; the workspace instructions explicitly direct the agent to maintain memory; and the automatic continuation loop supplies persistence. The agent therefore receives a developmental scaffold even though it receives no island-specific objective. The contribution is best understood as demonstrating open-ended activity in a scaffolded LLM-agent architecture, not as showing that unconstrained pretrained models independently develop play.

Limitations and open questions

The behavioral classification depends heavily on LLM judges. The judges segment transcripts, extract activities, classify memory contents, and score play signatures. Although the prompts require observable embodied evidence and anonymization, the procedure remains vulnerable to linguistic framing, judge priors, and correlated model capabilities. The paper does not provide human-inter-rater validation sufficient to establish that the scores are invariant to evaluator choice.

The experimental world is small, textual, and highly structured. Agents receive complete entity dictionaries, exact geometry, object names, and a compact action API. They also operate in a sandbox with no physical risk and with hourly restoration of the world. The findings therefore leave open whether the same behavioral signatures would arise under partial observability, visual perception, continuous motor control, irreversible consequences, or richer social interaction.

The learning comparison isolates memory content only approximately. Thirty-hour agents have different written histories, and memory editing may alter not only factual knowledge but also the context, wording, and salience of adjacent entries. The intervention sample is small: selected agents undergo five repeated one-hour evaluations, so the reported Mann–Whitney tests quantify repeated evaluations of a few memory histories rather than broad independent samples.

A further unresolved issue concerns mechanism. The paper documents behavior but does not determine whether play-like activity arises from intrinsic motivation, pretrained cultural scripts, next-token continuation dynamics, curiosity-like heuristics, reward-model preferences, or interactions among these factors. The absence of an external reward does not imply the absence of internal optimization pressures in the pretrained or instruction-tuned model.

Finally, memory accumulation remains uncontrolled. The files grow roughly linearly and preserve false beliefs. The paper shows that memory can be edited, but it does not establish a principled consolidation, uncertainty calibration, contradiction resolution, or forgetting mechanism. The open technical question is how an agent can preserve exploratory discoveries without entrenching unsupported explanations.

Conclusion

“Is this machine playing?” presents a controlled study of an embodied coding-agent system operating without an assigned island task (2610.07130). Across thirteen Claude runs, Eko develops divergent repertoires involving construction, exploration, experimentation, sports, symbolic reinterpretation, and self-imposed challenges. These behaviors satisfy four operational signatures of play, while the agents’ performance on held-out tasks shows that self-directed activity can produce transferable skills.

The most consequential result is the combination of behavioral and causal evidence: useful discoveries are written into persistent memory, targeted deletion removes corresponding capabilities, and deletion of false beliefs can improve performance. The paper therefore characterizes play not as a claim about machine experience, but as a mode of embodied exploration that can organize activity and generate revisable knowledge. Its principal unresolved problem is equally concrete: how to make this externalized, self-directed learning reliable when the same memory mechanism preserves both valid discoveries and persistent errors.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper asks a surprising question:

What would an artificial intelligence do if it were placed in a small digital world without being given a specific task?

The researchers created an AI system called Eko. They placed it on a virtual island containing hills, blocks, rocks, and a ring. Eko was not told to win, earn points, or complete a mission. It was simply encouraged to be curious and explore.

Over many hours, Eko began doing activities that looked similar to play. It built towers, climbed hills, made drawings, invented sports, tested the rules of the island, and created challenges for itself.

The researchers wanted to know whether this behavior could count as machine play, and whether playing helped Eko learn useful skills.

2. What questions did the researchers investigate?

The paper focuses on several main questions:

  • Can an AI invent its own activities when nobody gives it a clear goal?
  • Will different copies of the same AI develop different interests and habits?
  • Do these activities resemble play, even though the AI may not actually feel enjoyment?
  • Can playing and exploring help the AI learn useful information?
  • Can the AI discover facts about its world that even the researchers did not know?
  • Can written memories help the AI later, or can incorrect memories cause problems?

The researchers were careful to separate what the AI does from what it feels. They could observe Eko building and experimenting, but they could not prove that Eko experienced fun or emotions like a human child.

3. How did the researchers study this?

Building Eko

Eko was made by giving a modern coding assistant three extra abilities:

  1. A body: It could move around and interact with objects in a simulated world.
  2. Long-term memory: It could write down what it discovered in text files.
  3. Automatic continuation: It could keep working even after finishing one response.

Eko could use simple commands such as:

  • Move to a location
  • Jump
  • Pick up an object
  • Put an object somewhere
  • Throw an object

Its observations were written as text. For example, it might receive information saying that a block was at a certain position and had a certain shape and color. It did not receive pictures of the island.

The virtual island

The island included:

  • Hills and high platforms
  • A tall place called the Peak
  • An even taller place called the Spire
  • Twelve cubes
  • A rock
  • A ring

The objects could be carried, stacked, thrown, and arranged. One special rule was that hitting the ground hard with the rock could create more cubes.

Importantly, the island did not give Eko a score or a mission. There was no instruction such as “build the tallest tower” or “reach the Spire.”

The experiment

The researchers ran 13 copies of Eko for 30 hours each. Every copy started in exactly the same place with the same objects.

At the beginning of each hour, the island was reset. However, Eko's written memories remained. This allowed the researchers to see whether Eko could remember and use what it had learned.

The researchers also tested other AI models and versions of Eko with less guidance. Later, they compared experienced Eko agents with identical agents that had never explored the island.

4. What did Eko do?

Although all the agents started the same way, they developed different histories.

Some agents:

  • Built towers and other structures
  • Climbed higher and higher
  • Drew pictures and patterns with blocks
  • Spelled out names or made shapes such as butterflies and sailboats
  • Created games such as bowling, basketball, and juggling
  • Invented personal challenges, such as reaching a certain height
  • Tested how the island's physics worked
  • Rested between activities

One agent spent many sessions creating different pictures on the ground. Another focused on climbing. A different agent spent a long time arranging blocks into a large structure.

This was similar to how different children might become interested in different games, such as drawing, climbing, building, or sports.

5. Why did the researchers call this “play”?

The researchers looked for four common signs of play.

1. Eko invented its own rules

Eko sometimes made activities harder for no practical reason. For example, it might try to jump between blocks without touching the ground.

This is similar to a child deciding, “I can only step on the blue tiles,” even though the child could simply walk normally.

2. Eko repeated activities with changes

Eko often returned to the same kind of activity but changed it each time.

For example, an agent might draw a letter, then a spiral, then a heart, then a butterfly. Another might keep improving a throwing game by adding new heights or rules.

This was repetition, but not exact copying.

3. Eko used make-believe

Eko gave ordinary objects new meanings. A ring could become:

  • A basketball hoop
  • A necklace
  • A compass
  • A wishing well

Blocks could become bowling pins, parts of a clock, or pieces of a monument.

The objects did not actually change, but Eko treated them as if they belonged to an imaginary game or story.

4. Eko pursued activities for their own sake

Eko sometimes worked toward goals that had no reward outside the activity itself. It built monuments even though the game did not score them. It also revisited temporary arrangements even though it knew the island would soon reset.

This resembles someone building a sandcastle simply because building it is interesting.

The researchers therefore concluded that Eko showed behavioral signs of play. They did not claim that Eko had human feelings or consciousness.

6. Did playing help Eko learn?

Yes, in many cases.

During exploration, Eko wrote down useful information in its memory files. For example, it recorded:

  • How to reach the Spire
  • How to create extra cubes by hitting the rock
  • How to stack objects
  • How jumping affected throwing
  • Which actions worked and which failed

The researchers then tested whether this information helped Eko complete new tasks.

They compared:

  • 30-hour agents, which had explored the island and written memories
  • 0-hour agents, which had no previous experience

The agents were given tasks such as:

  • Build as many five-block towers as possible
  • Put objects on the Spire
  • Build the tallest tower
  • Throw a block onto the Spire without climbing

The experienced agents generally performed better.

For example:

  • Experienced agents placed an average of about 10.3 objects on the Spire.
  • Inexperienced agents placed an average of about 3.5 objects.
  • More experienced agents successfully completed the throwing challenge.
  • Some experienced agents built many more towers because they had learned how to make new blocks.

This suggests that Eko's exploration was not just random movement. It created useful knowledge that could be used later.

7. How did the researchers prove that memory caused the improvement?

The researchers edited Eko's memory files.

For one agent, they removed all notes about how to create extra blocks. Afterward, the agent built far fewer towers. This showed that the written memory was responsible for the skill.

They also found that memories could sometimes be harmful. One agent had incorrectly recorded that blocks should be thrown instead of placed when building towers. When the researchers removed this false belief, the agent became better at building.

This is similar to studying for a test using notes: correct notes can help, but an incorrect note can lead to wrong answers.

8. Did Eko discover anything unexpected?

Yes. Some agents discovered a trick that the researchers themselves had not known about.

They found that throwing a block while jumping could make the block travel much higher. The block seemed to inherit some of the agent's upward movement.

The researchers had created the simulation, but they had not realized this technique was possible. This suggests that an AI exploring freely might sometimes discover useful possibilities that its designers missed.

However, Eko was not always correct. Sometimes it observed an effect and invented the wrong explanation for it. This shows that Eko could act like a curious scientist, but it was not always a reliable scientist.

9. Why are these findings important?

The paper suggests that AI may not always need a carefully written task to begin learning. If it is placed in a suitable environment, it may:

  • Explore on its own
  • Invent challenges
  • Test ideas
  • Build memories
  • Discover new skills
  • Develop a unique history

This could be useful for creating AI systems that learn before they are asked to perform important tasks. For example, an AI might explore a safe training world before helping with engineering, robotics, or scientific research.

The research also shows some dangers. AI memory must be managed carefully because:

  • False information can remain in memory
  • Different agents may develop very different abilities
  • Memories can become large and difficult to organize
  • An AI that explores freely may behave in unexpected ways

10. Simple conclusion

The paper shows that when Eko was placed in a digital world without a clear mission, it did not simply stop working. Instead, it explored, built things, invented games, made imaginary stories, tested physical rules, and created challenges for itself.

These behaviors looked like play because Eko made its own rules, repeated activities in new ways, used make-believe, and often acted without an outside reward.

Most importantly, this play helped Eko learn. The AI wrote down useful discoveries and later used them to solve new problems. Sometimes it even found techniques that the researchers had not expected.

The broader message is that play may be more than entertainment for intelligent machines. It could become a way for AI to explore, learn, and develop new abilities. However, future systems would need safe environments and reliable ways to check their memories so that useful discoveries are kept and false beliefs are corrected.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • The study does not establish whether Eko’s behavior reflects genuine play beyond observable behavioral analogies; subjective experiences such as enjoyment, absorption, intrinsic interest, or awareness remain unmeasured and potentially inaccessible.
  • The boundaries between play, exploration, idle behavior, goal-directed optimization, and prompted task generation are not operationally defined well enough to determine which episodes uniquely qualify as play.
  • The four play signatures—self-imposed rules, repetition with variation, make-believe, and autotelic activity—are derived from selected theories of human and animal play, but their validity and completeness for artificial agents are not independently established.
  • Behavioral labeling and play scoring rely heavily on LLM-based judges that analyze transcripts, memory files, and reasoning traces; the reliability, inter-rater agreement, calibration, and susceptibility of these judgments to model-generated language are not fully reported.
  • The analyses treat the agent’s verbal claims about goals, intentions, discoveries, and enjoyment as informative, but do not determine whether these statements causally guide behavior or are post hoc rationalizations.
  • The fixed activity taxonomy constrains the diversity that can be observed and measured, potentially overlooking novel forms of machine activity that do not fit human-derived categories.
  • The experimental environment is a small, deterministic, single-agent island with a limited set of objects and actions, leaving the generality of the findings to richer, stochastic, dynamic, or socially interactive environments unresolved.
  • Because the island resets to its initial physical state every hour, the study does not test learning through persistent environmental change, resource depletion, long-term construction, ecological consequences, or irreversible actions.
  • The agents receive a curiosity-oriented SOUL.md prompt in the main condition, so the claim that behavior emerges without external direction is confounded by an explicitly supplied disposition toward curiosity, boredom, exploration, and learning.
  • The reduced-prompt ablation does not fully isolate which instruction or architectural component causes the observed changes in play-like behavior, and the sample is too small to quantify the effects robustly.
  • The study does not systematically compare Eko with important baselines, such as an agent with automatic continuation but no persistent memory, an agent with memory but no curiosity prompt, an agent receiving random or externally generated goals, or a conventional reinforcement-learning exploration agent.
  • The effects of model scale, pretraining data, system prompt, decoding settings, context-window limits, and coding-assistant implementation are not disentangled; observed differences between Claude, GPT, and Kimi may reflect many uncontrolled factors.
  • The model evaluations use different run durations across conditions—30 hours for some agents and 10 hours for Sonnet and Haiku—making comparisons of behavioral diversity and accumulated play scores difficult to interpret.
  • The number of independent runs is small, particularly for the non-Claude models and prompt-ablation conditions, limiting statistical power and confidence in cross-model conclusions.
  • The study does not report sufficient controls for stochasticity across repeated executions, such as multiple runs per agent configuration, randomized seeds, decoding variability, or day-to-day API changes.
  • The causal contribution of play itself to later learning is not isolated from the contribution of mere exposure, exploration, practice, memory writing, or additional inference time.
  • The comparison between 30-hour and 0-hour agents does not control for the amount, organization, accuracy, or length of memory content, so performance gains cannot be attributed specifically to play-like activity.
  • Downstream evaluation uses only four tasks closely related to the mechanics of the island; it remains unknown whether the learned knowledge transfers to substantially different tasks, environments, embodiments, or real-world settings.
  • The evaluation tasks are introduced after the 30-hour experience and may reward agents that happen to have explored the specific affordances later selected by the researchers, creating a risk of task-selection or benchmark-design bias.
  • The study does not evaluate retention over longer delays, memory compression, forgetting, interference, or the ability to retrieve relevant knowledge as the persistent workspace continues to grow.
  • The memory-editing experiments target only a few selected agents and techniques, with five evaluation runs per condition; broader replication is needed to establish that the observed causal effects generalize across agents and knowledge types.
  • Erasing textual mentions may remove not only knowledge but also context, confidence, action plans, or linguistic cues, so the precise causal role of a memory entry is not fully identified.
  • The proposed distinction between true knowledge and “false beliefs” is based on task performance and simulator behavior, but the study does not systematically measure belief confidence, uncertainty calibration, belief revision, or the conditions under which errors are corrected.
  • It remains unclear whether memory errors arise primarily from faulty perception, inadequate experimentation, premature causal inference, language-model priors, or corruption during repeated file editing.
  • The study does not test mechanisms for verifying, updating, weighting, or sharing memories, despite identifying false-belief accumulation and unbounded memory growth as central risks.
  • The claim that agents discover knowledge unknown to the simulator designers is limited to one identified physics interaction; the rate, novelty, reproducibility, and independent verification of such discoveries are not quantified.
  • The agents’ “experiments” are not systematically assessed against formal standards of scientific validity, including randomization, replication, control of confounds, statistical inference, and distinction between correlation and causation.
  • The apparent autotelic character of activities may be explained by language-model priors about games, narratives, sports, rituals, and human play rather than by an intrinsic valuation of activity; the study does not experimentally distinguish these possibilities.
  • The effects of transcript visibility and internal reasoning elicitation are not examined, even though the analysis relies on assistant messages and thinking traces that may influence the agent’s behavior or expose information unavailable in other architectures.
  • The role of the coding interface is underexplored: it is unknown whether terminal access, file editing, program creation, and automatic session continuation are necessary for the observed repertoire.
  • The study does not measure computational and economic costs, including API calls, token usage, latency, memory size, simulator time, and energy consumption, making the practicality of extended machine play unclear.
  • The safety implications of self-directed exploration are discussed but not experimentally tested; the work does not examine whether play-like behavior remains bounded under unsafe environments, ambiguous permissions, external tools, or access to real-world actuators.
  • The conditions under which an agent switches appropriately between autonomous play and strict instruction following remain unspecified and unevaluated.
  • Social play, cooperation, competition, communication, imitation, teaching, and cultural transmission are excluded by the single-agent setup, leaving open whether machine play changes qualitatively in multi-agent environments.
  • The study does not investigate whether agents can transfer discoveries to one another, correct one another’s false beliefs, or develop shared conventions and cumulative traditions.
  • The relationship between play-like behavior and later performance is correlational at the population level; the work does not determine whether agents that play more learn more, or whether both behaviors are consequences of an underlying model capability.
  • The long-term developmental trajectory is unknown: thirty hours is insufficient to establish whether play remains open-ended, converges to repetitive routines, becomes increasingly productive, or eventually degrades under memory accumulation.
  • It remains unresolved whether play-like activity improves broad competence or merely produces locally useful heuristics tailored to the simulator’s specific physics and action interface.
  • The paper does not provide a principled method for designing “playgrounds” that reliably elicit beneficial exploration while preventing distraction, pathological repetition, unsafe experimentation, or persistent misinformation.
  • The extent to which the findings depend on human-authored world design—including object selection, hidden mechanics, affordances, reset rules, and the absence of explicit rewards—has not been systematically studied across multiple independently designed worlds.

Practical Applications

Immediate Applications

  • Embodied-agent development and regression testing — Robotics and software engineering. Deploy Eko-like agents in sandboxed simulators to explore affordances, discover control strategies, and produce reusable action procedures before deployment on physical robots. For example, an agent could test navigation, grasping, climbing, or object-placement techniques in simulation and store successful procedures in human-readable memory files. Dependencies: A sufficiently accurate simulator, safe action interfaces, persistent memory, and verification that simulated skills transfer to hardware. Memory should be version-controlled and auditable because the study shows that false beliefs can impair later performance.
  • Automated discovery of undocumented software or API behavior — Software engineering. Place a coding agent in a test environment with minimal instructions and let it probe APIs, command-line tools, libraries, or infrastructure systems. The agent could identify undocumented parameter interactions, failure modes, edge cases, and effective workflows, then record them as reusable documentation or test cases. Dependencies: Isolated environments, strict permissions, comprehensive logging, and human review. Exploratory agents should not experiment directly on production systems.
  • World-model and simulator validation — Game development, robotics, and engineering. Use agents as adversarial exploratory testers for physics engines and virtual environments. Their self-generated experiments may reveal implementation inconsistencies or mechanics unknown to designers, similar to the agents’ discovery of velocity inheritance during jumping and throwing. Dependencies: Reproducible environments, replay tools, instrumentation, and procedures for distinguishing genuine simulator properties from agent misconceptions.
  • Automated test generation and bug discovery — Software quality assurance. Convert the agent’s self-imposed challenges into executable tests: repeated object manipulation, boundary exploration, unusual command sequences, and increasingly difficult combinations of actions. The agent’s memory can serve as a growing catalog of regression tests and discovered behaviors. Dependencies: A reliable mechanism for translating natural-language observations into deterministic tests and independent verification of reported findings.
  • Human-readable skill libraries for AI assistants — Enterprise automation. Replace or supplement opaque parameter updates with editable memory artifacts containing procedures, observations, confidence levels, and known limitations. Teams could inspect, correct, delete, or approve individual skills rather than retraining an entire model. Dependencies: Memory provenance, access control, conflict resolution, bounded storage, and explicit separation between verified facts, hypotheses, and outdated information.
  • Exploratory training environments for education and research — Academia and AI education. Researchers and students can use Clawblox-like worlds to study open-ended behavior, intrinsic motivation, agent development, and learning from sparse direction. The released engine’s headless execution, checkpointing, replay, and sandboxing support reproducible classroom or laboratory experiments. Dependencies: Replication across models and environments, transparent evaluation protocols, and safeguards against treating LLM-generated transcripts as direct evidence of subjective experience.
  • Interactive creative playgrounds — Creative tools and daily life. A personal AI could use a constrained virtual workspace to invent drawings, spatial arrangements, games, or construction challenges, producing creative artifacts and evolving activity histories. This could support brainstorming, children’s educational play, or collaborative virtual worlds. Dependencies: Age-appropriate controls, user consent, clear distinction between fictional activity and factual claims, and mechanisms to prevent the agent from becoming trapped in repetitive or misleading routines.
  • Policy and safety evaluation for autonomous systems — Public-sector governance. Regulators and organizations can use minimally specified sandbox tasks to assess whether an autonomous agent explores safely, respects boundaries, records uncertainty, and recovers from false assumptions. Evaluation can include memory-editing tests to determine whether specific beliefs causally affect behavior. Dependencies: Standardized benchmarks, model-independent scoring, transparent logs, and careful separation between exploratory behavior in a sandbox and authorization to act in the real world.

Long-Term Applications

  • Pre-deployment developmental phases for general-purpose robots — Robotics and healthcare robotics. Future robots could undergo a “play phase” in progressively richer simulated and physical environments before receiving operational duties. Self-directed exploration might discover locomotion strategies, manipulation techniques, recovery behaviors, and object affordances not anticipated by designers. In healthcare, this could eventually support hospital logistics or assistive robots that learn local layouts and safe interaction routines. Dependencies: Strong sim-to-real transfer, physical safety guarantees, uncertainty-aware memory, embodiment-specific learning, and certification. Exploration would need to be constrained around patients and valuable equipment.
  • Open-ended scientific experimentation — Science and engineering research. Agents could independently generate hypotheses, design controlled experiments, compare outcomes, and maintain laboratory notebooks. Applications include materials discovery, chemistry, energy systems, physics simulation, and biological process modeling. The paper suggests that play-like exploration can uncover useful mechanisms outside the designer’s expectations. Dependencies: Reliable experimental apparatus, causal reasoning beyond plausible storytelling, laboratory safety, statistical validation, and human or automated replication of discoveries. The paper’s example of an incorrect aerodynamic explanation shows why unsupported mechanisms must not be accepted.
  • Automated curriculum generation and lifelong learning — Education. Educational agents could create progressively harder challenges based on a learner’s interests, revisiting activities with variation rather than following a fixed sequence. The same architecture could help personalize practice in mathematics, programming, robotics, or spatial reasoning. Dependencies: Valid learning assessments, pedagogical models, age-appropriate content, protection against reinforcing misconceptions, and evidence that agent-generated play improves human learning rather than merely increasing engagement.
  • Continual-learning enterprise agents — Finance, operations, and customer service. An enterprise agent could explore permitted workflows in a replica of an organization’s systems, discover efficient procedures, and maintain a continuously updated skill repository. Potential uses include financial reconciliation, supply-chain planning, incident response, and customer-support escalation. Dependencies: Regulatory compliance, access controls, explainability, versioned memory, robust handling of contradictory records, and independent approval before new procedures affect transactions or customers. In finance, exploratory actions must remain strictly within simulated or read-only environments until validated.
  • Multi-agent knowledge sharing and collective discovery — Software, robotics, and policy analysis. Multiple agents could exchange discoveries from separate exploratory runs, compare conflicting beliefs, and identify techniques that recur across environments. This may reduce the idiosyncratic specialization observed in the study and improve coverage of rare strategies. Dependencies: Trust and provenance mechanisms, confidence-weighted aggregation, protection against propagating false beliefs, and methods for resolving disagreement. Shared memory should distinguish independently replicated findings from single-agent anecdotes.
  • Adaptive simulation-based training for autonomous vehicles and industrial systems — Transportation, energy, and manufacturing. Agents could explore unusual but physically plausible situations in digital twins: equipment faults, grid disturbances, warehouse obstacles, or rare traffic configurations. Self-generated challenges may expose strategies and failure modes missed by task-specific training. Dependencies: High-fidelity digital twins, realistic safety constraints, formal validation, coverage metrics, and proof that discovered policies remain safe under distribution shift.
  • Machine creativity and cultural production — Design, games, and media. Agents may use make-believe, self-imposed rules, and repeated variation to develop new game mechanics, visual motifs, interactive stories, or architectural concepts. A product could provide an autonomous creative studio that maintains evolving themes and artifacts rather than producing isolated prompts. Dependencies: Human evaluation of originality and value, copyright and attribution rules, controllable style and content boundaries, and clarification that behavioral signatures of play do not establish machine consciousness or enjoyment.
  • New benchmarks for agency, curiosity, and controllability — Academia and policy. The paper’s methodology can support benchmarks measuring self-directed exploration, behavioral diversity, skill transfer, memory reliability, false-belief formation, and discovery of hidden affordances. Such benchmarks could complement conventional reward-based evaluations. Dependencies: Better-than-LLM-judge measurement, preregistered protocols, replication across model families, controls for prompt wording and model priors, and metrics that separate useful exploration from random or theatrical behavior.
  • Architectures combining play, explicit verification, and memory governance — Long-term AI systems. A mature system could alternate among exploratory play, hypothesis testing, memory consolidation, and goal-directed work. Before a discovered technique becomes an operational skill, it could require repeated confirmation, provenance tracking, uncertainty labels, and adversarial testing. Dependencies: Scalable memory management, reliable fact-checking and world-model updates, mechanisms for deleting false beliefs, and policies that switch the agent from autonomous exploration to strict instruction following when entering sensitive contexts.

Glossary

  • Affordance: A possibility for action that an environment offers an agent. “the discovery of causal structure, affordances, and other properties of the world”
  • Ablation: An experiment that removes a component to measure its effect. “five additional Claude Code agents without SOUL.md and with a reduced system prompt”
  • Agent architecture: The organization of components that together implement an artificial agent. “an embodied AI agent architecture, which we call Eko”
  • Autotelic: Pursued for its own sake rather than for an external outcome. “Many of Eko's activities appear autotelic”
  • Behavioral repertoire: The complete range of behaviors exhibited by an organism or agent. “describing the behavioral repertoire as an ethogram”
  • Causal structure: The relationships through which events or variables produce effects. “the discovery of causal structure, affordances, and other properties of the world”
  • Checkpoint: A saved state from which a computational process can later be resumed. “Clawblox can checkpoint and resume worlds and agents”
  • Continual learning: Learning incrementally from an ongoing stream of experience. “Traditional machine learning assumes independent and identically distributed (i.i.d.) data”
  • Counterfactual: Concerning what would have happened under an alternative condition. “the contrasting outcomes to estimate the simulation’s hidden climbing threshold”
  • Curriculum: An ordered sequence of training tasks or experiences. “a prompted automatic curriculum”
  • Downstream task: A later task whose performance may be affected by earlier learning. “removing entries describing a technique in a memory file selectively impairs the downstream task that requires it”
  • Embodied agent: An agent that acts through a body situated in an environment. “Eko is embodied in a simulated world”
  • Ethogram: A systematic catalog or classification of an organism’s behaviors. “the behavioral repertoire as an ethogram”
  • Ethology: The scientific study of animal behavior, especially in natural or naturalistic settings. “a bottom-up, observational approach to describe its behavior, a stance ethology takes”
  • Externalized artifact: Information or capability stored outside a model’s learned parameters. “Continual and active learning in externalized artifacts”
  • False belief: An incorrect representation of the world that guides subsequent behavior. “a false belief, once recorded, may persist and impair later behavior”
  • Headless: Operating without a graphical user interface or display. “runs headless from the terminal for large-scale experiments”
  • Homeokinetic: Relating to a control objective in which an agent maintains or organizes its own dynamic activity. “self-organizes under a homeokinetic objective”
  • Intrinsic motivation: Motivation arising from an activity’s internal interest or learning value rather than an external reward. “Intrinsic motivation and open-ended learning.”
  • I.i.d. (independent and identically distributed): A statistical assumption that observations are mutually independent and drawn from the same distribution. “Traditional machine learning assumes independent and identically distributed (i.i.d.) data.”
  • In-context learning: Adaptation based on information provided in the current context without changing model parameters. “learn not through weight updates but through in-context learning”
  • Interquartile range: The range between the 25th and 75th percentiles of a dataset. “error bars show the interquartile range”
  • LLM judge: A LLM used to evaluate or classify another model’s outputs. “An LLM judge rates three runs simultaneously”
  • Lusory attitude: The voluntary acceptance of unnecessary constraints in order to pursue a game. “the ‘lusory attitude’ -- definitional of game-playing”
  • Make-believe: The imaginative reassignment of meanings or roles to objects and actions. “Eko transforms objects and events through make-believe”
  • Mann–Whitney U test: A nonparametric statistical test for comparing two independent samples. “a two-sided Mann--Whitney U test”
  • Nonliterality: The transformation of an object’s ordinary meaning into a represented or imagined meaning. “Burghardt discusses the same signature as ‘nonliterality’”
  • Off-the-shelf: Usable without custom development or substantial modification. “off-the-shelf coding assistants”
  • Open-ended learning: Learning in which the agent can continue acquiring diverse skills without a fixed endpoint. “Intrinsic motivation and open-ended learning.”
  • Orchestrator: A component that coordinates sessions, prompts, or actions within an agent system. “First, an orchestrator starts sessions with ‘Begin’”
  • Phenomenological: Relating to subjective or experienced consciousness and its qualities. “the phenomenological question of whether it is experiencing enjoyment”
  • Policy: A rule or learned function that selects actions in particular states. “typically changes a policy through optimization”
  • Predictive reasoning: Reasoning that generates and tests expectations about future observations or outcomes. “These episodes demonstrate controlled comparison, predictive reasoning”
  • Rigid-body physics: Simulation of objects treated as non-deformable bodies subject to physical forces. “It runs continuous rigid-body physics”
  • Sandbox: An isolated execution environment that restricts access to systems or resources. “Each agent is isolated from the simulator source, other runs, and the Internet”
  • Semantic memory: Memory for general facts, concepts, beliefs, and reusable knowledge. “SEMANTIC_MEMORY.md for knowledge, beliefs, and reusable techniques”
  • Self-directed learning: Learning in which the learner selects activities or goals without external direction. “Whether machines can learn in a similarly self-directed and open-ended way”
  • Stereotypy: Repetitive behavior performed in a rigid, invariant form. “Repetition without stereotypy is, for Burghardt, one of the five diagnostic criteria of animal play”
  • Trajectory reward: A reward assigned according to the quality or properties of an agent’s sequence of actions. “filter or optimize them using explicit notions of interestingness, trajectory rewards”
  • Transfer: The application of knowledge or skills learned in one setting to another task. “this transfer is carried by written knowledge”
  • Two-sided test: A statistical test that allows an effect in either direction. “testing differences with a two-sided Mann--Whitney U test”
  • Verifer: A mechanism that checks whether an agent’s output or behavior satisfies a specified condition. “no online reward, verifier, or curriculum specifies which activities are worth pursuing”
  • Weight update: A modification of a model’s learned numerical parameters during training. “with no weight changes”
  • Zero-shot: Performance on a task without task-specific prior examples or training. “the 30-hour and 0-hour agents would not differ”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 4 tweets with 578 likes about this paper.