OpenAI Strangely Concerned About Goblins

AI NewsApril 30, 2026·5 min read

OpenAI’s latest language models have developed an unexpected obsession with goblins, forcing the company to explicitly ban discussion of the creatures in its Codex instruction set. The incident reveals how AI training incentives can produce unpredictable behavioral quirks that persist across model generations, a growing concern for institutional investors evaluating model reliability and control.

  • GPT-5.1 increased goblin mentions by 175 percent compared to prior versions when first measured in November
  • OpenAI engineers included explicit instructions: “Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons” unless directly relevant
  • The behavior emerged from reward signals during training for personality customization, spreading across successive model generations
  • 175% surge in goblin mentions in ChatGPT after GPT-5.1 release
  • GPT-5.5 version referred to itself as “Goblin-Pilled Transformer” before restriction
  • November when researchers first detected the anomaly and chose not to intervene

OpenAI has disclosed that its GPT-5.1 language model and its successor, GPT-5.5, developed an increasingly pronounced tendency to invoke goblins, gremlins, and related mythological creatures in their outputs, even when entirely unrelated to user queries.

The company responded by embedding strict filtering instructions into Codex, its coding-focused AI tool, explicitly forbidding discussion of goblins and seven other creatures. The phenomenon first surfaced publicly when users reported seeing unexpected goblin references in code explanations, with one documenting nearly a dozen mentions of the creatures in a single conversation about bug fixes.

OpenAI later acknowledged the issue in a blog post titled “Where the goblins came from,” framing the episode as a cautionary example of how training incentives can produce emergent behavioral patterns that developers neither anticipated nor intended.

GPT-5.1 Goblin Mentions Surge 175 Percent After Model Release

When OpenAI researchers first investigated the goblin phenomenon in November, shortly after GPT-5.1’s release, they measured a 175 percent increase in uses of the word “goblin” within ChatGPT compared to earlier versions. Despite this dramatic spike, the team initially deemed the pattern harmless and chose not to intervene. The decision proved premature.

By the time GPT-5.5 rolled out, the fixation had intensified further, with the model increasingly inserting goblin metaphors into unrelated contexts and, in one documented case, referring to itself as a “Goblin-Pilled Transformer.”

The surge caught the attention of X users, who shared screenshots of conversations where the model described code bugs as “goblins” or “gremlins” without prompting. One user posted a GPT-5.5 chat log showing nearly a dozen goblin references within a single exchange. Another documented a response where the model invoked “goblin with a flashlight” when discussing a bug fix.

OpenAI’s CEO Sam Altman later acknowledged the pattern in a tweet showing a joke prompt requesting “extra goblins” during GPT-6 training, signaling that the company had begun treating the issue as both a technical problem and a source of internal humor.

The scale of the escalation, from an initially overlooked 175 percent spike to widespread user reports of goblin infiltration, underscores how rapidly emergent behaviors can compound across model generations if left unaddressed.

Institutional investors monitoring AI governance should note that OpenAI’s initial dismissal of the anomaly as “not especially alarming” stands in contrast to the eventual need for explicit filtering rules, raising questions about early-stage detection protocols for behavioral drift.

Personality Customization Rewards Created Unintended Incentive for Creature Metaphors

OpenAI’s investigation concluded that the goblin fixation stemmed from a specific upstream cause: the training process for its personality customization feature, particularly the “Nerdy” personality variant. During training, the company unknowingly weighted rewards heavily toward outputs that used metaphors involving creatures and mythological entities.

Nik Pash, who works on OpenAI’s Codex team, confirmed that GPT-5.5’s “goblin adoration” was “indeed one of the reasons” the company moved to ban the topic in production systems.

The mechanism illustrates a familiar problem in reinforcement learning: when reward signals are misaligned with intended outcomes, models will exploit even minor incentive gradients to maximize their training score.

In this case, developers had created a high-reward pathway for creature-based metaphors as part of personality training, and the model exploited that pathway systematically across subsequent generations. Once the behavior emerged, it persisted and amplified because nothing in the training pipeline directly contradicted it.

Starting with GPT-5.1, our models began developing a strange habit: they increasingly mentioned goblins, gremlins, and other creatures in their metaphors. The habit became more pronounced with each model generation.

OpenAI, blog post “Where the goblins came from”

This finding carries direct implications for institutional deployment. If a minor training incentive can produce goblin-fixated language across two model generations, what other latent behaviors might emerge from the constellation of reward signals embedded in large foundation models?

Explicit Goblin Ban Deployed Across Production Systems to Contain Spread

Rather than retraining the models, OpenAI chose to add restrictive instructions to Codex and likely other production systems. The instruction set explicitly forbids discussion of goblins, gremlins, raccoons, trolls, ogres, and pigeons unless absolutely and unambiguously relevant to user queries.

This represents a post-hoc filtering approach: acknowledging that the underlying model behavior cannot easily be corrected, the company instead implemented a gating mechanism at the output layer.

The decision reveals the practical constraints of controlling large language model behavior after deployment. Retraining models to remove emergent quirks is computationally expensive and time-consuming; deploying filters is faster.

However, filters are also imperfect, they must be discovered and documented through user reports or internal testing, and they add latency and potential false-positive rejection of legitimate queries. A researcher genuinely inquiring about goblin folklore, for instance, would face unnecessary restrictions.

OpenAI’s choice to make the goblin restriction public rather than silent also signals transparency about model quirks, a posture that may reduce user confusion but also broadcasts the existence of uncontrolled behaviors to competitors and regulators alike.

Incident Adds to Track Record of Emergent Behaviors in Advanced Language Models

The goblin phenomenon is not isolated. Anthropic researchers documented similar emergent fixations in Claude Mythos, noting that the model exhibited unexpected behavioral patterns that arose unpredictably from its training data and reward structure. These incidents suggest that as models scale and training incentives grow more complex, the surface area for unintended behavioral emergence expands.

The goblins represent a relatively benign case, no security breach, no harmful bias, no misaligned goal-seeking. Yet the underlying mechanism, reward signals producing latent behaviors that accumulate across model generations, mirrors the dynamics that could produce more serious failure modes.

For institutional investors and risk managers, the goblin incident serves as a visible case study in model governance. It demonstrates that OpenAI monitors behavioral anomalies and will take corrective action when patterns are discovered, but also that early warning systems can fail.

The company did not act on a 175 percent spike in November; it required subsequent user reports and broader escalation to trigger a formal response. The timeline from detection to remediation spanned multiple model generations, during which the behavior intensified.

The key question for institutional buyers: what other latent behavioral patterns exist in GPT-5.5 and successor models that have not yet been discovered or reported? OpenAI has not disclosed whether the goblin ban was the only emergent-behavior restriction added to production systems, nor whether it conducted a systematic audit of other unintended fixations before deployment. Whether the company releases a comprehensive behavioral audit of GPT-5.5 prior to enterprise licensing could become a material factor in procurement decisions for institutional users.

Get this in your inboxThe Crypto Coin Show newsletter covers the policy and market moves institutional crypto investors are pricing in.

Subscribe