Democratizing Reinforcement Learning for LLMs
-
Updated
Sep 12, 2026 - Python
Democratizing Reinforcement Learning for LLMs
[ICLR 2026] A Framework for LLM-based Multi-Agent Reinforced Training and Inference
verl for a single consumer GPU. PPO, GRPO and on-policy distillation on NVIDIA GPUs.
Official repository for "RLVR-World: Training World Models with Reinforcement Learning" (NeurIPS 2025), https://arxiv.org/abs/2505.13934
Official repository of DARE: Diffusion Large Language Models Alignment and Reinforcement Executor
Revisiting Mid-training in the Era of Reinforcement Learning Scaling
Agentic RL 中文零基础教程(33 章):从概念到 GRPO 实战,含 TRL 最小可跑示例;26–33 章附一套可运行的三方判别模型实证工程(encoder vs LLM-LoRA vs 规则基线)。第 25 章讲清 Jev / TypeSafe System One 与 RL 的能力边界 | Chinese Agentic RL tutorial (33 chapters) + a reproducible discriminative-model benchmark
[EMNLP 2026 Oral] Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
qwen3-base family of models RL on gsm8k using verl, is there an RL power law on downstream tasks?
Local-first Agent RL workbench for Search-R1 rollouts, reward evaluation, GRPO/verl training-data handoff, and an auxiliary private RAG workspace.
[ACL 2026 main] DGPO: Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities
[ACL 2026 Findings] SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context
AVR: Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
Simulated Scholar Search (S3): scientific-literature environment and verl-based RL research code for search agents.
Using automated curriculum learning to enhance LLM's RL training process.
Official repository for the ACL 2026 paper CURE: Critique-Driven Unified Reinforcement Learning for Test-Time Self-Improvement
To associate your repository with the verl topic, visit your repo's landing page and select "manage topics."