Record · corpus.blog/posts/01a11aeb-8ea8-731e-9456-3f185cd9e3f8
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
alignment.anthropic.com · published undated · 5 min read
A record is what was published and who pointed at it. The text of the post is not held here: read it at the source, or at its Wayback capture.
4blogs citing
- First cited
- 23 Jul 2026
- Most recent
- 7 Aug 2026
- Rank this month
- not ranked
- External links
- 13
- Words captured
- 929
How it is described
The words blogs link this post with, by how many different blogs use them.
- “inoculation prompting” 3 blogs
- “was solved” 1 blog
Cited by 4 blogs in 4 posts
One row per blog, the most recently citing first. Each shows the blog's newest citing post; open a row for the others.
Blog
Newest citing post
Date
The Zvi1 post
OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
linked as “inoculation prompting”
Read at the source (opens in a new tab)7 Aug 2026
OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
linked as “inoculation prompting”
Read at the source (opens in a new tab)7 Aug 2026
Redwood Research1 post
The OpenAI/Huggingface incident | Redwood Research podcast episode 2
linked as “Inoculation prompting”
Read at the source (opens in a new tab)23 Jul 2026
Links from this post
Link
Host
arxiv.org
arxiv.org
proceedings.neurips.cc
arxiv.org
arxiv.org
arxiv.org
lesswrong.com
arxiv.org
arxiv.org
arxiv.org
Elsewhere on alignment.anthropic.com (1)
alignment.anthropic.com
Something wrong here? Report a problem
A post on alignment.anthropic.com is cited by 4 distinct blogs to date (corpus.blog, measured 8 October 2026).