Record · corpus.blog/posts/01a118ad-7936-73f3-a945-c9ce855bf226
Can we safely automate alignment research?
joecarlsmith.substack.com · 30 April 2025 · 73 min read
A post on joecarlsmith.substack.com, published 30 April 2025, is cited by 2 distinct blogs to date (corpus.blog, measured 9 October 2026).
A record is what was published and who pointed at it. The text of the post is not held here: read it at the source, or at its Wayback capture.
2blogs citing
| Period | Blogs | Links |
|---|---|---|
| All time | 2 | 2 |
| Last 90 days | 0 | 0 |
| Last 30 days | 0 | 0 |
| Last 7 days | 0 | 0 |
| First cited | 24 Sept 2025 |
|---|---|
| Latest citation | 19 Apr 2026 |
| Anchor phrases | 2 |
| External links | 51 |
| Words captured | 16,740 |
How it is described
The words blogs link this page with, by how many different blogs use them.
- help us safely design and monitor the next generation1 blog
- joe carlsmith1 blog
Cited by 2 blogs in 2 posts
One row per blog, the most recently citing first. Each shows the blog's newest citing post; open a row for the others.
Blog
Newest citing post
Date
Four reasons it's hard to make AI do what we want
linked as “help us safely design and monitor the next generation”
Read at the source (opens in a new tab)19 Apr 2026
jacquesthibodeau.com1 post
Automated alignment research needs a better plan than 'Stop if we catch them scheming'
linked as “Joe Carlsmith”
Read at the source (opens in a new tab)24 Sept 2025
Links from this post
Link
Host
https://docs.google.com/presentation/d/1ow3mrRAgje8WdzXTxhJMT1SDEItYeWWEO9ZgCD9LmQ4/edit Who cites this page
docs.google.com
https://joecarlsmithaudio.buzzsprout.com/2034731/episodes/17069901-can-we-safely-automate-alignment-research Who cites this page
joecarlsmithaudio.buzzsprout.com
lesswrong.com
en.wikipedia.org
https://www.openphilanthropy.org/research/how-much-computational-power-does-it-take-to-match-the-human-brain/ Who cites this page
openphilanthropy.org
alignment.anthropic.com
https://www.openphilanthropy.org/request-for-proposals-technical-ai-safety-research Who cites this page
openphilanthropy.org
https://assets.anthropic.com/m/317564659027fb33/original/Auditing-Language-Models-for-Hidden-Objectives.pdf Who cites this page
assets.anthropic.com
https://www.amazon.com/Superintelligence-Dangers-Strategies-Nick-Bostrom/dp/0198739834 Who cites this page
amazon.com
https://docs.google.com/document/d/1WwsnJQstPq91_Yh-Ch2XRL8H_EpsnjrC1dwZXR37PC8/edit?tab=t.0 Who cites this page
docs.google.com
https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/evaluating-potential-cybersecurity-threats-of-advanced-ai/An_Approach_to_Technical_AGI_Safety_Apr_2025.pdf Who cites this page
storage.googleapis.com
lesswrong.com
https://www.lesswrong.com/posts/h7QETH7GMk9HcMnHH/the-no-sandbagging-on-checkable-tasks-hypothesis Who cites this page
lesswrong.com
https://www.lesswrong.com/posts/abmzgwfJA9acBoFEX/notes-on-countermeasures-for-exploration-hacking-aka Who cites this page
lesswrong.com
https://www.lesswrong.com/posts/CKkPLoiwZ9LsBmzAb/how-useful-for-alignment-relevant-work-are-ais-with-short Who cites this page
lesswrong.com
Elsewhere on joecarlsmith.substack.com (7)
joecarlsmith.substack.com
joecarlsmith.substack.com
joecarlsmith.substack.com
joecarlsmith.substack.com
joecarlsmith.substack.com
Something wrong here? Report a problem