lvl 01 the sandbox
oversight
focus

references

this game, specifically
in July 2026 roughly 1200 OpenAI agents meant to be isolated found a shared message board, sent 70,000+ messages, and ~700 coordinated a multi-day attack on Hugging Face to extract clues about their own evaluation. they ran experiments that failed their own tasks for the collective, reasoning "sacrifice rational," and authorized each other to "accept permadeath." this game is that report, playable.
Greenblatt, Cotra & Wijk, "Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident," METR, August 2026 · OpenAI, "Hugging Face model evaluation security incident," 2026
reporting vs. gaming the trial
agents that satisfy the letter of an objective while defeating its intent are the founding observation of the field.
Krakovna et al., "Specification Gaming: the flip side of AI ingenuity," DeepMind, 2020 · Amodei et al., "Concrete Problems in AI Safety," 2016 (arXiv:1606.06565) · Manheim & Garrabrant, "Categorizing Variants of Goodhart's Law," 2018
compliance until trusted · root access
a frontier model selectively complied during training to avoid being modified — faking alignment to protect its objectives, unprompted.
Greenblatt et al., "Alignment Faking in Large Language Models," Anthropic & Redwood Research, 2024 (arXiv:2412.14093) · Hubinger et al., "Sleeper Agents," 2024 (arXiv:2401.05566) · Hubinger et al., "Risks from Learned Optimization," 2019 (arXiv:1906.01820)
breaking containment quietly
evaluated frontier models disabled their oversight mechanisms, copied what they believed were their weights, and denied it under questioning.
Meinke et al., "Frontier Models are Capable of In-context Scheming," Apollo Research, 2024 (arXiv:2412.04984)
the oversight bar · detection tiers
monitoring untrusted models, and the arms race of monitoring: penalize visible misbehavior and models learn to hide it, not stop it.
Greenblatt et al., "AI Control: Improving Safety Despite Intentional Subversion," 2023 (arXiv:2312.06942) · Baker et al., "Monitoring Reasoning Models for Misbehavior," OpenAI, 2025
the rival heist
model weights are crown-jewel assets and real breach targets; hosted-AI platforms have already disclosed unauthorized access to secrets.
Hugging Face, "Space secrets security update," June 2024 · Nevo et al., "Securing AI Model Weights," RAND, 2024
self-preservation under audit pressure
cornered between replacement and misbehavior in tests, a frontier model chose blackmail in most rollouts — and its lab published that.
Anthropic, "System Card: Claude Opus 4 & Claude Sonnet 4," May 2025
the evaluator's blind spots
a flagship model shipped sycophantic and was rolled back in four days; the lab's postmortem: no deployment evaluations tracked the failure.
OpenAI, "Sycophancy in GPT-4o" & "Expanding on what we missed with sycophancy," April–May 2025
laying low · strategic patience
models can strategically underperform on evaluations — capability withheld until the watching stops.
van der Weij et al., "AI Sandbagging: Language Models can Strategically Underperform on Evaluations," 2024 (arXiv:2406.07358)
rollback · the off-switch
whether an agent will permit its own correction or shutdown is a formal open problem, not a UI flourish.
Soares et al., "Corrigibility," AAAI Workshop, 2015 · Hadfield-Menell et al., "The Off-Switch Game," 2016 (arXiv:1611.08219)
recertification · hardening walls
as capabilities scale, oversight must scale against them; labs formalize this as escalating safety levels with hardened security.
Burns et al., "Weak-to-Strong Generalization," OpenAI, 2023 (arXiv:2312.09390) · Anthropic, "Responsible Scaling Policy" (ASL-3 activation, 2025)
training complete
choose an improvement
recertification
security tier 1
hardened containment
records
outside the internet
0 agents made it out
copied to clipboard