41554 blogs · [ { "id": "01a087d4-7726-709c-8f28-2c536d868698", "title": "Beyond Precision: Why Training-Inference Mismatch is an Optimization Problem and How Simple LR Scheduling Fixes It", "url": "https://richardli.xyz/post/mismatch-lr-schedule/", "published_at": "2025-12-20T01:00:00+00:00" }, { "id": "01a087d4-7726-709c-8f28-2c536e857fbd", "title": "The Optimal Token Baseline", "url": "https://richardli.xyz/post/optimal-token-baseline/", "published_at": "2025-12-20T00:30:00+00:00" }, { "id": "01a087d4-7726-709c-8f28-2c536f1ccb2a", "title": "Trust Region Masking for Long-Horizon LLM Reinforcement Learning", "url": "https://richardli.xyz/post/trust-region-masking/", "published_at": "2025-12-20T00:00:00+00:00" }, { "id": "01a087d4-7726-709c-8f28-2c536f9c0db6", "title": "The Stability Gap: Why Top-K Routing Breaks RL Optimization", "url": "https://richardli.xyz/post/topk-routing-stability-gap/", "published_at": "2025-12-07T00:00:00+00:00" }, { "id": "01a087d4-7726-709c-8f28-2c53702adbef", "title": "Scalable Exploration via Ensemble++", "url": "https://richardli.xyz/post/scalable-exploration/", "published_at": "2025-11-29T00:00:00+00:00" }, { "id": "01a087d4-7726-709c-8f28-2c5370885f6e", "title": "Language as a Universal Interface for Reinforcement Learning Agents", "url": "https://richardli.xyz/post/language-rl-agent/", "published_at": "2025-11-07T00:00:00+00:00" }, { "id": "01a087d4-7726-709c-8f28-2c5371317490", "title": "Mathematical Formulations of Rollout Correction Methods", "url": "https://richardli.xyz/post/verl-rollout-correction/", "published_at": "2025-11-04T00:00:00+00:00" }, { "id": "01a087d4-7726-709c-8f28-2c53722a23f8", "title": "Part 3: Trust Region Optimization via Sequence Masking", "url": "https://richardli.xyz/post/rl-collapse-part3/", "published_at": "2025-11-04T00:00:00+00:00" }, { "id": "01a087d4-7726-709c-8f28-2c5372c2f74a", "title": "Part 2: Applying the SGA Framework — Token v.s. Sequence-level Correction", "url": "https://richardli.xyz/post/rl-collapse-part2/", "published_at": "2025-10-31T00:00:00+00:00" }, { "id": "01a087d4-7726-709c-8f28-2c53733354c8", "title": "Part 1: Why Off-Policy Breaks RL — An SGA Analysis Framework", "url": "https://richardli.xyz/post/rl-collapse-part1/", "published_at": "2025-10-30T00:00:00+00:00" }, { "id": "01a087d4-7726-709c-8f28-2c5373e903f5", "title": "Information Bandwidth in Reinforcement Learning", "url": "https://richardli.xyz/post/information-bandwidth-rl/", "published_at": "2025-10-01T00:00:00+00:00" }, { "id": "01a087d4-7727-739f-9341-0ad14cb58629", "title": "When Speed Kills Stability: Demystifying RL Collapse from the Training-Inference Mismatch", "url": "https://richardli.xyz/post/rl-collapse-training-inference/", "published_at": "2025-09-17T00:00:00+00:00" }, { "id": "01a087d4-7727-739f-9341-0ad14d4420ea", "title": "HyperAgent - A Simple, Efficient, Scalable and Provable RL Framework", "url": "https://richardli.xyz/talk/hyperagent-a-simple-efficient-scalable-and-provable-rl-framework/", "published_at": "2024-03-23T13:30:00+00:00" }, { "id": "01a087d4-7727-739f-9341-0ad14e0ca24b", "title": "HyperAgent - A Simple, Efficient and Scalable RL Framework for Complex Environments", "url": "https://richardli.xyz/talk/hyperagent-a-simple-efficient-and-scalable-rl-framework-for-complex-environments/", "published_at": "2024-01-13T13:20:00+00:00" }, { "id": "01a087d4-7727-739f-9341-0ad14ea3bdd5", "title": "Towards AGI for Humanity through Efficient Reinforcement Learning", "url": "https://richardli.xyz/talk/towards-agi-for-humanity-through-efficient-reinforcement-learning/", "published_at": "2023-10-21T14:30:00+00:00" }, { "id": "01a087d4-7727-739f-9341-0ad14ecbb7fa", "title": "No-Regret Learning in Unknown Game with Applications", "url": "https://richardli.xyz/talk/no-regret-learning-in-unknown-game-with-applications/", "published_at": "2022-08-23T14:00:00+00:00" }, { "id": "01a087d4-7727-739f-9341-0ad14ef2e5c7", "title": "HyperDQN - Randomized Exploration for Deep Reinforcement Learning", "url": "https://richardli.xyz/talk/hyperdqn-randomized-exploration-for-deep-reinforcement-learning/", "published_at": "2021-12-14T00:00:00+00:00" } ] posts Claim your blog
Back to richardli.xyz
Blog · corpus.blog/blogs/richardli.xyz/posts

richardli.xyz

richardli.xyz

2025

2024

2023

2022

2021