Neuromorphic
Hide and Seek: How Simple Competition Taught AI Agents to Use Tools
With only a reward for hiding or seeking, OpenAI's agents invented forts, ramp attacks, ramp theft and box surfing. How six strategies emerged from competition alone, what an autocurriculum is, and why the experiment echoes how brains evolved.
Akmal Alif · 8 October 2026 MYT

Evolution had no plan for intelligence. Simple rules, natural selection and competition, applied over a very long time, produced it anyway. In 2019 OpenAI asked whether a comparably simple setup could make artificial agents grow more capable without anyone teaching them, and answered with a game of hide-and-seek (OpenAI, 2019). The three-minute video is embedded below; the timestamps jump to the matching moment. This post adds the details from the paper behind it (Baker et al., 2020).
The setup
Two teams share a small physics world with walls, boxes and ramps. Hiders are rewarded for staying out of the seekers' sight, and seekers for finding them. Seekers are frozen for a short preparation phase at the start of each round, which gives hiders time to act (watch from 0:16; 0:31).
That is the whole reward. Nothing in it mentions boxes, ramps, building or teamwork. The agents are trained with reinforcement learning and self-play: they play thousands of games in parallel for days, against each other and against earlier versions of themselves (watch from 2:06).
Six strategies, each forcing the next
What emerged was an arms race. The paper reports six distinct phases of strategy, and each one created pressure for the other team to adapt (Baker et al., 2020):
Run and chase. Early on, seekers learn to chase and hiders learn to flee, which is not enough for the hiders (watch from 0:16).
Build a shelter. After millions of rounds, hiders learn to grab boxes, move them and lock them in place to wall themselves in. They also learn to cooperate on tasks no single agent could do (watch from 0:31; 0:46).
Use ramps. Seekers answer by pushing ramps against the walls and jumping in (watch from 0:46).
Take the ramps away. Hiders learn to move the ramps to the edge of the arena and lock them before seekers can use them (watch from 1:03).
Box surfing. In a more open environment, seekers find they can climb onto a box, then ride it across the floor to the hiders' shelter. The researchers had not known their physics allowed this (watch from 1:34; 1:51).
Lock everything. Hiders learn to lock every spare box before they start building, which leaves seekers nothing to surf on (watch from 1:51).
The key line in the video is that none of these behaviours was explicitly rewarded. Each team's new skill changed the problem the other team faced, and that pressure did the teaching (watch from 1:03; 1:18).
Autocurricula
The paper calls this an autocurriculum: a sequence of ever-harder challenges that the agents generate for each other, instead of one designed by researchers. That matters because hand-designing curricula and rewards for complex behaviour is hard and brittle. The authors also argue that multi-agent competition may scale better with environment complexity than the main alternative for self-supervised exploration, intrinsic motivation. On a suite of transfer tests, the hide-and-seek agents learned several human-relevant skills faster than agents trained with intrinsic motivation or from scratch, though not every test favoured them (Baker et al., 2020).
Why it belongs in Neuromorphic
Neuromorphic computing usually means copying the brain's hardware: spikes, synapses and dendrites. Hide-and-seek copies something else that shaped the brain, the environment. Brains became capable under pressure from predators, prey and rivals that kept changing, and the experiment reproduces a small version of that loop. The video ends on that analogy: competition and co-evolution produced the only generally intelligent species we know, and a larger, more diverse world might one day produce truly intelligent agents (watch from 2:22; 2:38).
Two cautions keep that in proportion. The world is tiny and the agents' skills are narrow. And box surfing is a reminder that optimisers find whatever the environment allows, including behaviour designers never intended. The same tendency that makes autocurricula creative makes them hard to predict.
References
Baker, B., Kanitscheider, I., Markov, T., Wu, Y., Powell, G., McGrew, B., & Mordatch, I. (2020). Emergent tool use from multi-agent autocurricula. In International Conference on Learning Representations. https://arxiv.org/abs/1909.07528
OpenAI. (2019, September 17). Multi-agent hide and seek [Video]. YouTube. https://www.youtube.com/watch?v=kopoLzvh5jY