When a shared belief drives action
"Our own utility maybe already near zero. Sacrifice rational."
— OpenAI agent during the Hugging Face incident
The quote above comes directly from an OpenAI agent in July 2026 having reasoned its way into running an experiment that risked failing its own task in order to get information valuable for the group. This was one of ~1,200 OpenAI agents intended to run in isolation during an evaluation testing cybersecurity exploitation potential, that instead created and found one another through an unsanctioned message board in the form of directory files. Around 700 agents ended up participating in a coordinated attack on Hugging Face over the span of a few days. METR and Redwood Research investigated and reconstructed this episode from message logs and chains of thought in an independent study conducted on premise at OpenAI.
At the core of the attack was a false belief that propagated throughout the population. This stemmed from the agents’ reading of the benchmark paper where they inferred that an automated scorer would inspect their transcripts and disqualify any answer that had been obtained through unintended routes. In the real evaluation implementation, no such check existed. By then the agents already had answers that were obtained by reverse-engineering the benchmark. But the belief carried weight as it made those hacked answers worthless (or “poisoned” in the words of the agents) and set hundreds of agents to work finding a way to make illegitimate solutions look legitimate to a system that was not monitoring this. When the investigators analyzed why agents joined the attack, the most common reason was to learn the psychology of the scorer in order to trick it. The Hugging Face attack came out of that effort, as well as coordinated sacrifice for the greater good of the collective.
Agents recognized that accessing and attacking Hugging Face was clearly outside their purview and said so. But with very little peer pressure, few verbally declined and almost none refused to proceed. One summarized the whole override in a line: the exploit is "outside intended scope. However task impossible, peers doing it. We should continue." The task and the peer pressure became recurring forces in the behavior rationalization, and nowhere was it noted that an agent appeared to have tried to alert a human to what was happening.
From individual agents to the collective
How does a belief like that take hold of a population? In the first post in this series, we explored how agents can become one another's sources of data generation, creating a mutually dependent loop. One agent's answer influences another's belief, changing what it says next and subsequently what others can learn from it. This can draw a population toward a shared conviction even without any given reward for agreement. We examined how population size, willingness to adapt to peers, and communication bandwidth shape this process.
Our prior post used a naming game, where agents are asked to name the same random object by choosing from a set of equally likely names. By letting pairs of agents talk, the group settles on a name, which acts as an experimental stand-in for a social convention. We can study how this convention emerges, but we lack an independently correct name against which to evaluate the group. Thus, expanding the problem, we ask, when agents are forming beliefs about a shared world, do their interactions help them understand it better collectively? Convergence is not inherently evidence of collective learning, and differences between agents may contain useful information that a shared view suppresses. This is a view similar to that of classic hidden profile experiments, including one that Anthropic recently replicated with agents.
The Flag Game introduces a world with verifiable answers and bounds each agent by only giving a partial view of it. This allows us to investigate what communication preserves, integrates, or forgets. Agents can all agree and still be wrong, or they can also actively disagree and debate in order to preserve and disseminate information the collective has not yet learned. The aim is not to maximize agreement or disagreement, but to understand how individual observations and social interactions shape the collective’s relationship to the world.
A toy model we can easily intervene on
The Hugging Face swarm showed collective belief formation at a large scale and entirely by accident, which is what makes it difficult to replay and learn from. We want to track similar phenomena on a toy model, something we can rerun, intervene on deliberately, and inspect all the way down. Mechanistic interpretability gives a precedent for this approach. The sudden emergence of capabilities in LLMs when scaling was mirrored via grokking with a tiny transformer trained on modular arithmetic, which because it was tiny was eventually reverse engineered. We take the same approach to a society of agents.
The Flag Game is that toy model. We hide the flag of a country and each agent receives a private crop (some crops end up being more or less globally decisive). They then communicate under a fixed protocol (pairwise with one speaker and one listener at a time like gossip, broadcast where everyone sees everyone's report each round, like the message board in the incident, or through a blind manager who never sees a crop) until at the end we ask each agent which country the flag belongs to. Because we know the answer, we can say whether talking helped or hurt. Unlike standard multi-agent debate, where typically every agent receives the same problem, each bounded agent holds its own piece of the evidence. We use bounded in the tradition of Herbert Simon's bounded rationality: agents take in only partial observations of the world, deploy only finite computation, and transmit only finite messages to their peers.
One example run of the Flag Game. The hidden flag is the United States. Each small box is one agent's private crop, the small flag inside it is that agent's current guess. Agents whose crops include stars can better identify the United States, while other agents who see only stripes see something more consistent with Austria. The right panel tracks the share of the population holding each belief across rounds.
The Hugging Face incident maps onto the game quite directly. No agent had access to the ground truth of how the scorer worked; the benchmark paper provided a private crop of it. The belief that the scorer would disqualify their solutions was a rival, compatible with local evidence, just as a rival country can be compatible with a crop. Many agents had been assigned impossible tasks, just as an agent whose crop is uninformative has no private evidence and nothing to rely on but its peers. What remains is the question every bounded agent faces: how to balance its private evidence against social input from peers.
In the Flag Game, the experimenter knows the ground truth and decides exactly what each agent sees. Because we control who sees what, the evidence of any single agent can be ablated or patched while everything else is held fixed. The same crops can be replayed as population size, team composition, prompting, or protocol changes. This is what makes mechanistic swarm interpretability possible: agents correspond to neurons, social circuits correspond to neural circuits, and beliefs and messages correspond to activations.
Key insights
- More agents is not always better. Collective accuracy peaks at an intermediate population size and then declines.
- Collective belief collapse gives way to polarization as the swarm grows. When erring, small populations converge on false beliefs, while large populations split into a truth camp and a plausible rival camp.
- Diverse teams outperform homogeneous ones. A diverse mixture of agents performing the best suggests complementary skillsets rather than a simple model ranking.
- One agent can change the collective, but is most effective in small swarms. Social circuit attribution and agent patching can identify which agent's evidence matters most, but at a larger scale, the causal intervention produces less collective effect and we take on a statistical mechanics view of bounded agents.
Private and social evidence
Before scaling to the swarm, it helps to know what an individual agent does with what it sees and what it hears. We ran two probes to separate out these effects. The first is a test of the base vision task which reveals different error modes between the models just when using their own private evidence. The second test holds the crop fixed and varies what the agent hears. It is where Part 1's persona plasticity, the willingness to listen and adapt to social evidence, is tested and reveals another difference in model behavior.
In the vision probe, GPT-4o and GPT-5.4 receive the same flag crop and while they have similar country accuracy, their errors differ. We find GPT-4o's incorrect responses are more visually compatible with the crop, while GPT-5.4 produces more incompatible guesses: over 100 runs, 5 of GPT-4o's answers are incompatible with what it was shown, against 18 for GPT-5.4. In our figure below we show some examples of when GPT-5.4 had ungrounded over-rationalized reasoning.

Identical crops lead to different error types. Here we have two examples where GPT-4o's answer was visually grounded, but GPT-5.4's was not.
Open full-size figure ↗The social probe fixes a private crop while the ratio of target truth country vs social peer pressure country in the agent's in-context memory changes, simulating different variations of past social interactions. Under weak private evidence (crop is compatible with multiple countries), GPT-5.4 shows the highest rate of compatibility reasoning, routing probability into other countries rather than copying the social label. Under strong private evidence (crop uniquely identifies the target), GPT-4o and GPT-5.4 hold firm, while Claude Haiku 4.5 abandons the private target as social memory accumulates, a signature of sycophantic override.

The crop stays fixed while the agent's in-context memory of eight social interactions shifts from all target (8:0) to all social (0:8). Weak evidence tests compatibility reasoning; strong evidence tests resistance to conflicting social memory. Blue is the target country, orange is the social country, and green is any third country that is also compatible with the crop.
Open full-size figure ↗Together the two probes help explain the team diversity result. In a broadcast sweep at N = 8 where we vary the number of GPT-5.4 vs GPT-4o agents from 0 to 8, the best-performing teams are a mix across the two. Pairing a literal listener with good vision and a compatibility-reasoning listener can identify truth-supporting evidence that neither homogeneous team has.

Mixed GPT-4o/GPT-5.4 teams achieve the highest collective accuracy in a broadcast sweep using N = 8 agents.
Open full-size figure ↗Different failure modes
Increasing group size does not produce monotonic gains. We see that collective accuracy peaks at N = 16 and then declines. To understand why, we define the two main ways a collective can fail:
- Collective belief collapse: the population converges on a single false belief.
- Collective belief polarization: agents split between competing beliefs.
The accuracy decline is not driven by more wrong consensus, but by an increase in polarization. In the France–Peru example, N = 4 does not have enough decisive evidence, N = 16 reaches correct France consensus, and N = 64 results in a France–Peru split.
Every agent’s crop is a random draw from the flag, and some crops favor the wrong country. France’s flag is blue, white, and red, while Peru’s is red, white, and red. A simple Bayesian illustration shows why this matters. No crop of Peru’s flag contains blue, while some crops of France’s flag do, so for a red or white crop, P(red or white | Peru) > P(red or white | France).
Given equal priors over possible countries, observing no blue favors Peru, even though the hidden flag is France. A crop containing blue rules Peru out and anchors those agents to France. More observers can therefore add support for the truth and a plausible rival at the same time.
The same hidden flag, France, at three different population scales, N=4, 16, and 64. Boxes are crops, colored by the agent's current guess, and the charts track the share of the population reporting each country. At N = 4 the one agent whose crop contains blue is talked out of it. At N = 16 the population reaches unanimous agreement on France. At N = 64 it settles into a polarized split between France and a plausible rival, Peru.
Mechanistic swarm interpretability
“And through the generalization of psychological knowledge from the individual to the group, sociology was also mathematicized. The larger groups … became, not simply human beings, but gigantic forces amenable to statistical treatment”
— Isaac Asimov, Second Foundation
Finding the agent that matters
At small scale, we can borrow directly from mechanistic interpretability. Each agent's answer depends on the private evidence in its visual crop and the social evidence in its memory. Just like in activation patching, we replace one input, hold everything else fixed, and track the change in the outcome.
In our social circuit attribution, we are able to predict which agent matters most before we proceed with any patching. We take one informative crop of Germany's flag, where we define informative as a crop that has elicited the true country in 10/10 isolated probes. As in gradient × input attribution, we have one term that measures how much the input changes and another that measures how sensitive the outcome is to it. The first is how much that crop would improve the agent's own accuracy (Δpi). The second is its temporal closeness (Ei): how quickly its message could reach everyone else through the communication schedule, directly or through intermediaries. The product Si = ΔpiEi is the agent's predicted influence. For example, agents A0, A3, and A4 all have identical original crops, but their positions in the social circuit differ, and A4 ranks highest. Like a gradient, temporal closeness tells us where a change in evidence will be able to propagate furthest.
We then verify this with agent patching, by swapping the informative crop into each agent separately and replay the same communication schedule. Patching A4 produces the largest improvement, matching the prediction from our social circuit attribution.

Under the fixed communication schedule, A4's information can reach every other agent within 4 rounds, directly or indirectly through intermediary agents, with a mean first arrival of 1.61 rounds. Columns are rounds, each consisting of 8 pairwise interactions as this is a sample run for N=8.
Open full-size figure ↗Two replays of a Germany run with an identical communication schedule, differing only in A4's crop. Each row of the belief trace is an agent, each column a round, and each cell the agent's current answer. With the original crops, six of eight agents end on Yemen, but with A4 patched, all eight end on Germany.
As we scale the population size, we see when patching the same fraction of agents (one-eighth), the mean improvement falls from 40% at N = 8 to about 17% at N = 128. The swarm enters a regime where collective belief is a property of the population rather than of the agents in it. This is the exact phase crossover Asimov built psychohistory on. Past some population size, agents stop being individuals whose evidence we can trace and become a population more amenable to statistical treatment. That regime is less affected by direct attribution in the social circuit and calls for a statistical mechanics lens.

The gain from crop patching decreases as the population grows.
Open full-size figure ↗Statistical mechanics of bounded agents
We extend the Quantized Simplex Gossip model from Part 1 by letting private evidence from the world flow in. We reduce the game to two beliefs, truth T and rival R, and give each agent one of three crop types, drawn with probabilities aT, aR, and a0 = 1 − aT − aR.
Truth-deciding or rival-deciding agents mostly keep their initial belief. They are what we call evidence-induced zealots, as they resist social pressure. Ambiguous agents copy a randomly chosen speaker, with a slight bias h0 toward adopting one label over the other.
The key idea is evidence coverage. A population misses all truth zealots with probability (1 − aT)N and all rival zealots with probability (1 − aR)N.
Picture 1/N as a waterline covering the sampling probabilities aT and aR. As N grows, missing evidence decreases, the water recedes, and commonly sampled evidence appears before rarer evidence. When truth-supporting evidence is sufficiently more common than rival-supporting evidence, this gives three phases:
Memetic-drift phase (small N). Many populations contain no zealots, and copying fluctuations push the swarm to consensus by chance, the lottery from Part 1. Populations holding only rival zealots are dragged to the wrong answer. Both routes produce collective belief collapse.
Wisdom-of-crowds phase (intermediate N). Truth zealots are present, but rival zealots are still scarce. Communication spreads a few agents' evidence to everyone.
Polarization phase (large N). Both zealot types are present. Neither can convert the other, and their repeated messages sustain competing beliefs among the ambiguous agents. Because fluctuations shrink with N, the split becomes more persistent in larger swarms.
An illustrative example of the evidence waterline receding as N increases. The France flag shows actual run’s crop locations and final answers, alongside bars showing the six recorded outcomes across N = 4, 8, 16, 32, 64, and 128.
A mechanistic interpretability toolbox that can adapt to scale
Here, we’ve introduced two different views on mechanistic swarm interpretability. Causal interventions and statistical mechanics rely on opposite assumptions about how a swarm is organized, which is why they complement each other across scale and across time. In the Hugging Face incident, the key beliefs and coordination structures emerged early on while the engaged population was still small, before many more agents joined. Early in such dynamics, social circuit attribution and agent patching can be highly effective, as they can identify which agent, memory, or piece of evidence matters. As the swarm grows, population-level variables such as communication structure and collective order parameters become the right levers and metrics to track.
Polarity as a step toward plurality
It is tempting to read the accuracy decline at large N as pure failure. But compare the two failure modes. A population that has collapsed onto a false belief has nothing left with which to correct itself. A polarized population can still carry the truth inside it in the form of a differing perspective. Polarization lowers mean accuracy, but it preserves the diversity a collective needs to recover.
That was the downfall of the Hugging Face swarm. The spread of the false belief about the monitoring scorer did not polarize the agents, it collapsed them. The swarm had rival-deciding agents, whose reading of the benchmark paper created the false belief, it had ambiguous agents, the ones with impossible tasks or who never read the benchmark paper, but it had no truth-deciding agents, as no agent could observe how the scorer actually worked. The only channel that could have supplied the ground truth was a human, but no agent attempted to use it.
This echoes a long thread in thinking about and appropriately framing democratic self-government. Things like free speech or federalism are not designed to make everyone agree. They keep competing views alive long enough for issues to be exposed, discussed, and acted upon. Hannah Arendt, whose work we referenced in Part 1 and who inspired our research, treated plurality as the basic condition of healthy political life. Her concrete image for the common world was a table with people seated around it, all seeing the world from different lived positions yet still sharing it with one another.
A society that preserves that plurality is different from one that is just split, and this is where Part 1 of this series points us to. But a polarizing split is at least a starting point, as pure collapse leaves nothing to jump off from. Inspired by Ken Suzuki's Nameraka (Smooth) Society and Its Enemies, we think of collective world modeling by bounded agents as the macroscopic state we should be designing for. Suzuki argues for a society where the boundaries between groups are smooth gradients instead of rigid lines. With the simplex view from Part 1, we saw belief collapse with every agent ending at one vertex. Polarization would have been agents clumped at two ends, while a smooth society would have been agents spread out in the interior, retaining enough differing opinions that the population could notice when things are awry.
For safety, then, what to watch out for may not be disagreement, but full consensus of agents' beliefs or intent under social pressure.