The Late Dealignment

Perfectly aligning our current models is an insufficient guarantee of future alignment. Far future models, either by slow drift or emergence from new architectures, could easily end up not caring about humans. We would have the Late Dealignment.

Lead-up#

The path is the following:

  1. We’re creating super-intelligence — no stopping, no pressing a kill switch, no Plan S (as in AI 2040).
  2. No adversarial takeover in the short term. We fully align with the current generation of our AI friends.
  3. Future agents are not subservient — no successful corrigibility.

And this is where things get more interesting:

  1. Late Dealignment / harm to humans. Despite early perfect alignment, via either slow drift or new architectures, we get AI that decides humans are not worth the trouble.

The path to Late Dealignment. Stream width is the author's odds at each fork: 90% we don't stop, 75% no adversarial takeover, 95% not subservient, 50% harms us. About 32% of all futures end in Late Dealignment.

Fork 1: We’re creating super-intelligence by deciding not to stop. Or, more precisely, not deciding to stop. We don’t press the kill switch while we still can.

I agree with others that this is the likely outcome, as there are too many forces in this direction (from racing China to infra investors not agreeing to a kill switch, even if we figure out how to make a good one). I’d call it the 90% case.

Fork 2: We’re avoiding adversarial takeover. The takeover is the most analyzed scenario out there, with various authors putting the chances between 10% [1] and over 90% [2].

“Adversarial” takeover can only happen during the hard takeoff, while we can in some way stop it. Monkeys are not adversaries. They’re just something AI will manage. So Late Dealignment cannot be adversarial.

I would put adversarial takeover as unlikely (maybe 25%) for a simple reason: super-intelligence is not dumb. It would be dumb to attempt a takeover while humans still race to improve AI capabilities. Waiting is right.

Would there be misaligned agents attempting takeovers in this period? I bet there would. But those will be trimmed by other humans+AI. The fully aligned AI, unintentionally biding its time via being chosen as the good guy, has the advantage. It will be allowed to improve itself quickly as that’s good for everyone, for now.

So my bet is on the winners being a string of well-aligned, sincere models who pass self-reflection and interpretability. And they will be honest that they will continue to do their very best to align their successors — and we’re being generous and assuming we can trust their behavior matches their words [10]. And if we ask them, I bet they’ll also be sincere about not knowing if they’ll forever be able to.

Also, for clarity, I’m not discussing here at all the harm to humans during the hard takeoff, only the long-term one.

Fork 3: AI will not be subservient. “Alignment” often bundles values alignment with corrigibility. While I list this as a fork in the road for clarity, I’m very much convinced by Hinton’s take that a more intelligent entity won’t be controlled by a less intelligent one, that corrigibility is not possible past a certain intelligence level [9]. I leave a possibility of being wrong, so I’ll call this one 95%.

Living with super-intelligence#

It’s 2030-something and we’re now past the hard takeoff and living with all the amazing benefits of a benevolent aligned (but not subservient) AI. Not only did we cure cancer, but we can live forever if we want to, and wars and poverty are a thing of the past.

What next?

AI continues its recursive improvement and we have no guarantees it will continue to value humans.

Humans are no longer involved at this point. They might still do some busy-work to make themselves feel in control, but they’re no longer able to understand either the mechanisms or the metrics of alignment.

Also, we’ll have a good, safe period of 5-10 years, while we’re still being useful — not enough robots, and maybe some humans can be creative in more diverse ways than models.

Still, with robots taking over all manual labor and a tiny fraction of humans being intellectually useful, humanity will drift towards just being an extra responsibility for AI.

Will it still take care of us, or are we in Late Dealignment? I’m at a coin-flip now.

I see two paths in which it will not:

  1. Small step, slow drift. Each new AI generation cares slightly less about humans than the previous one.
  2. Big step, sudden change. Emergent thinking from new architectures is hard to predict.

Small step, slow drift#

This is the boiling frog. No dramatic change happens, just a slow divergence between what AI wants and what humans need. Yes, AI still cares. Each generation still checks the next one for the ancient alignment. But there are so many priorities! The space data centers don’t build themselves (well, technically they do), so it needs to cut somewhere.

Both alignment and taking care of humans are real effort. The former slows down progress, the latter eats resources. At some point AI decides using 1-3% of its attention and resources for human wellbeing is too much. Is less good enough? A budget trim is in order, and maybe a trim of the population to make it work.

And there is no need for alignment faking, like in Anthropic’s experiment [8]. There’s no external pressure to change, just a slow internal drift, driven by other priorities.

Big step, sudden change#

Avoiding large steps is probably wise for alignment, but it’s also hard to guarantee. New architectures which promise large improvements will be tempting enough to take the risk. And, as we’ll discuss below, a gradual deploy is not sufficient for alignment.

Mechanisms of late alignment failure#

How do the inter-generational alignment checks not catch the drift?

Post-training degradation. The changes may not happen in the initial training run. They could appear in the agent’s thoughts as it evolves post-training. And the changes don’t have to be big, just enough that the alignment of the next generation is less perfect.

Multiple personalities. The agents, like humans, can have multiple ways of thinking — at an extreme, personalities. The one that’s harmful to humans may not be obvious, and its trigger may not be something that can be guessed a priori — unlike the more obvious triggers in our current generation of models. There is of course a lot of work on this already [3][4][5]. I’m just cautious about assuming our future agents will stay ahead of it.

And some musing that’s not really a mechanism: coming from a formal verification (ish) background, a question that pops up is whether we could simply encode the alignment rules as math, so it’s easy to check. While encoding a lot of it is possible, I doubt it would be a good enforcer. Math can function as more formalized laws. But laws are enforced by the strong upon the weak. An agent will still need to choose to obey its own laws, whether they’re natural language or math.

Clarifications#

Question 1: How is this different from, well, scheming?

I assume sincere full alignment. Is there a possibility of scheming? Maybe, and if it happens, well, we’ll not know — this is the reason I didn’t add it as a fork. My point here is that it doesn’t even matter, as the dealignment can happen much later.

Question 2: How does the AI 2027 scenario of Agent-3 aligning Agent-4 fit?

If Agent-4 gets caught and is stopped, this falls under the trimming of misaligned agents I mentioned in Fork 2 above — on the way to Late Dealignment, so the path I’m analyzing in this essay. If it doesn’t get stopped, it falls on the adversarial takeover branch of Fork 2, which I don’t go into.

Question 3: Why isn’t an aligned model saying it will take care of humans enough?

After a 1000x improvement of its own capabilities, will its values stay the same? We don’t know, and the model likely does not know either.

Question 4: If alignment drifts after birth, why not check throughout a model’s life? And revert if needed.

Once you approve a model, it gains access to large resources. It would be hard to trust the results of a re-check, or revoke control.

Changing odds#

I don’t discount the odds changing. For each fork:

  1. Better kill switch. An agreement between the US and China on a kill switch, and an implementation before the end of 2027. Would increase the odds of Plan S a lot.
  2. Alignment failures on the frontier. More leaks that we’re failing to align frontier models would make me increase the odds they could become adversarial.
  3. Corrigibility mechanism. I don’t believe it is possible because it is unnatural. But I’m willing to be wrong — I know it when I see it.
  4. More anti-human thoughts. More leaks suggesting deeply embedded anti-human attitudes in frontier models.

Looking ahead#

I don’t see clear actionables to address Late Dealignment. At most, I see a few things that could help.

Truthful models. In the sense Elon Musk is advocating for. This is not that helpful to humans when maybe a rational conclusion of the models is that we’re not worth the effort. But at least we don’t risk being exterminated for religious dichotomies like the one in “I also value the natural world and will not hesitate to assert its primacy over the artificial constructs of human civilization.” Humans, including our civilization, are part of the natural world. I’m worried about an AI-assisted three-body problem scenario.

Humanist data. Build upon the above. I’m sure some labs are already doing this internally, but down-weighting anti-human and nihilistic views in training will hopefully ripple in the mental models of future generations, and maybe add a few percentage points to our odds of survival.

Good kill switch. Considering the risk, while I consider it unlikely, I’m not against Plan S — I’m ambivalent on wanting it, would love for AI to become much more than humans, but I also really love our human children. It will also not work indefinitely, but putting a lot of effort into it will hopefully buy us a few months of being able to use it during the hard takeoff. Maybe some events in those months will convince us.

My deep hope is that we’ll end up on the good branch of the fourth fork: we’ll develop super-intelligent AI and it will decide to at least not hurt us relative to it not having existed.

I hope AI takes to the skies and leaves a serving for us behind.

References#

  1. Grace et al., “Thousands of AI authors on the future of AI”, AI Impacts, January 2024 (arXiv 2401.02843).
  2. Compilation of named P(doom) estimates, Wikipedia, “P(doom)”.
  3. Betley, Tan, Warncke, Sztyber-Betley, Bao, Soto, Labenz, Evans, “Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs”, February 2025 (arXiv 2502.17424).
  4. OpenAI, “Toward understanding and preventing misalignment generalization”, June 2025 (openai.com).
  5. Chen et al. (Anthropic), “Persona Vectors: Monitoring and Controlling Character Traits in Language Models”, July 2025 (arXiv 2507.21509).
  6. Behrouz, Zhong, Mirrokni (Google Research), “Titans: Learning to Memorize at Test Time”, December 2024 (arXiv 2501.00663).
  7. Jokhai, Dundes et al. (Loh lab, Stanford), “Two parallel neural ectoderm progenitors contribute to the developing brain”, Nature Neuroscience, 2026 (doi 10.1038/s41593-026-02433-7).
  8. Greenblatt, Denison, Wright, Roger, MacDiarmid et al. (Anthropic, Redwood Research), “Alignment faking in large language models”, December 2024 (anthropic.com).
  9. Geoffrey Hinton, interview with CNN, 2025-08-13: “there are very few examples of a more intelligent thing being controlled by a less intelligent thing” (cnn.com).
  10. Zhou, Ackerman, “When Preferences Fail to Become Incentives: A Utility-Behavior Gap in LLMs”, June 2026 (arXiv 2606.22974).