Alignment faking in large language models
A post on anthropic.com, published 18 December 2024, is cited by 26 distinct blogs to date (corpus.blog, measured 8 October 2026).
A record is what was published and who pointed at it. The text of the post is not held here: read it at the source, or at its Wayback capture.
26blogs citing
| Citing blogs | 26 |
|---|---|
| Links | 36 |
| First cited | 16 Dec 2024 |
| Latest citation | 11 Sept 2026 |
| Anchor phrases | 5 |
| External links | 5 |
| Words captured | 2,181 |
How it is described
The words blogs link this post with, by how many different blogs use them.
- “alignment faking” 3 blogs
- “faking alignment” 3 blogs
- “here” 2 blogs
- “identifying which transcripts in an experiment contain alignment-faking behaviors” 2 blogs
- “3” 1 blog
Cited by 26 blogs in 36 posts
One row per blog, the most recently citing first. Each shows the blog's newest citing post; open a row for the others.
linked as “8”
Read at the source (opens in a new tab)linked as “behave deceptively during training”
Read at the source (opens in a new tab)linked as “Alignment Faking Study”
Read at the source (opens in a new tab)linked as “subsequent discussion”
Read at the source (opens in a new tab)linked as “paper published in December 2024”
Read at the source (opens in a new tab)linked as “lying”
Read at the source (opens in a new tab)1 other post
linked as “tricked into not complying”
Read at the source (opens in a new tab)linked as “hide signs that their actual goals are not exactly what their creators intended”
Read at the source (opens in a new tab)linked as “hide some of its behavior from Anthropic”
Read at the source (opens in a new tab)linked as “pretending to be aligned”
Read at the source (opens in a new tab)linked as “resists attempts to make it harmful”
Read at the source (opens in a new tab)linked as “here”
Read at the source (opens in a new tab)3 other posts
linked as “alignment faking”
Read at the source (opens in a new tab)linked as “identifying which transcripts in an experiment contain alignment-faking behaviors”
Read at the source (opens in a new tab)linked as “here”
Read at the source (opens in a new tab)4 other posts
linked as “alignment faking”
Read at the source (opens in a new tab)linked as “identifying which transcripts in an experiment contain alignment-faking behaviors”
Read at the source (opens in a new tab)linked as “paper”
Read at the source (opens in a new tab)linked as “strategically fake alignment to preserve their goals”
Read at the source (opens in a new tab)linked as “3”
Read at the source (opens in a new tab)linked as “look for ways to protect themselves”
Read at the source (opens in a new tab)linked as “actively subvert the assessment itself”
Read at the source (opens in a new tab)1 other post
linked as “Alignment faking in large language models”
Read at the source (opens in a new tab)linked as “Alignment Faking by Anthropic AI”
Read at the source (opens in a new tab)linked as “Anthropic recently published”
Read at the source (opens in a new tab)linked as “a lot of effort”
Read at the source (opens in a new tab)Links from this post
Elsewhere on anthropic.com (8)
Something wrong here? Report a problem