corpus.blog
Most cited
Talked about
Blogs
41554 blogs · [ { "id": "01a08c56-b9e4-730a-b8f3-5fe3a9da2880", "title": "How to Train an LLM to do proofs: Beyond Verifiable Rewards", "url": "https://tobysimonds.com/research/2025/09/29/Proofs.html/research/2025/09/29/Proofs.html", "published_at": "2025-09-28T14:00:00+00:00" }, { "id": "01a08c56-b9e4-730a-b8f3-5fe3a9dd7baf", "title": "All In on Winning: The dangers of RL without moral reward", "url": "https://tobysimonds.com/research/2025/09/29/Proofs.html/research/2025/08/22/PokerRL.html", "published_at": "2025-08-21T14:00:00+00:00" }, { "id": "01a08c56-b9e4-730a-b8f3-5fe3a9ed6899", "title": "RLSR: Reinforcement Learning from Self Reward", "url": "https://tobysimonds.com/research/2025/09/29/Proofs.html/research/2025/08/06/LLM-Self-Rewarding.html", "published_at": "2025-08-05T14:00:00+00:00" }, { "id": "01a08c56-b9e4-730a-b8f3-5fe3aab537ca", "title": "AlphaWrite: Inference time compute Scaling for Writing", "url": "https://tobysimonds.com/research/2025/09/29/Proofs.html/research/2025/06/06/AlphaWrite.html", "published_at": "2025-06-05T14:00:00+00:00" }, { "id": "01a08c56-b9e4-730a-b8f3-5fe3aae7ceba", "title": "LLMs for Engineering: Teaching Models to Design High Powered Rockets", "url": "https://tobysimonds.com/research/2025/09/29/Proofs.html/research/2025/04/27/LLM-Rorckets.html", "published_at": "2025-04-26T14:00:00+00:00" }, { "id": "01a08c56-b9e4-730a-b8f3-5fe3ab1f4e41", "title": "LADDER+TTRL", "url": "https://tobysimonds.com/research/2025/09/29/Proofs.html/research/2025/03/01/Ladder-TTRL.html", "published_at": "2025-02-28T13:00:00+00:00" }, { "id": "01a08c56-b9e4-730a-b8f3-5fe3ab359055", "title": "Turning the web into RL Questions", "url": "https://tobysimonds.com/research/2025/09/29/Proofs.html/research/2025/02/01/Web_to_RL.html", "published_at": "2025-01-31T13:00:00+00:00" }, { "id": "01a08c56-b9e4-730a-b8f3-5fe3abecd9c7", "title": "REL: Working out is all you need", "url": "https://tobysimonds.com/research/2025/09/29/Proofs.html/research/2024/09/01/REL.html", "published_at": "2024-08-31T14:00:00+00:00" }, { "id": "01a08c56-b9e4-730a-b8f3-5fe3ac150794", "title": "MoDEM: Mixture of Domain Expert Models", "url": "https://tobysimonds.com/research/2025/09/29/Proofs.html/research/2024/09/01/MoEA.html", "published_at": "2024-08-31T14:00:00+00:00" }, { "id": "01a08c56-b9e4-730a-b8f3-5fe3ac83bb6c", "title": "EAD: Entropy Adpative Decoding", "url": "https://tobysimonds.com/research/2025/09/29/Proofs.html/research/2024/09/01/EAD.html", "published_at": "2024-08-31T14:00:00+00:00" } ] posts
Claim your blog
Back to tobysimonds.com
Blog · corpus.blog/blogs/tobysimonds.com/posts
tobysimonds.com
tobysimonds.com
2025
How to Train an LLM to do proofs: Beyond Verifiable Rewards
original ↗
28 Sept 2025
All In on Winning: The dangers of RL without moral reward
original ↗
21 Aug 2025
RLSR: Reinforcement Learning from Self Reward
original ↗
5 Aug 2025
AlphaWrite: Inference time compute Scaling for Writing
original ↗
5 Jun 2025
LLMs for Engineering: Teaching Models to Design High Powered Rockets
original ↗
26 Apr 2025
LADDER+TTRL
original ↗
28 Feb 2025
Turning the web into RL Questions
original ↗
31 Jan 2025
2024
REL: Working out is all you need
original ↗
31 Aug 2024
MoDEM: Mixture of Domain Expert Models
original ↗
31 Aug 2024
EAD: Entropy Adpative Decoding
original ↗
31 Aug 2024