#Reinforcement Learning
50 articles with this tag

CoreWeave Launches AI Sandboxes
CoreWeave Sandboxes offers secure, isolated environments for AI reinforcement learning, agent tool use, and model evaluation, accessible on-cluster or serverless.

Raymond Feng on Post Training and Autonomous Agentic Citizens
Raymond Feng of Applied Compute outlines how post-training is evolving toward custom enterprise setups and continuous online learning.

Reinforcement Learning Beyond Verifiable Rewards
Will Brown of Prime Intellect discusses the limitations of reinforcement learning in domains without easily verifiable rewards.

SymmGrid Accelerates Robot Learning
SymmGrid framework dramatically accelerates on-robot learning for manipulation tasks, achieving up to 2.17x speed-ups and moving closer to sub-10 minute training.

AI Pioneers Debate Transformer's Future, Urge New Architectures
AI pioneers Jerry Tworek and Rohan Anil discuss the limitations of current Transformer architectures and the need for new models that can learn from real-world experience.

Alex Shaw: "Everything Is a Rollout" in AI Agent Evaluation
Alex Shaw from Lode Institute explains the Harbor framework, highlighting how agent development mirrors ML and requires empirical evaluation. Discover the tools and use cases for building and testing AI agents.

NYT Explores Local AI for Accessible Mobile Games
The New York Times' Shafik Quoraishee and Joanne Song discuss their work on local agentic AI for accessible mobile games, highlighting on-device benefits and future challenges.

ABot-World-0: Real-time Video World Models
ABot-World-0 introduces a real-time video world model for long-horizon agent interaction, achieving 16 FPS at 720P with an optimized inference stack.

World Models: The Key to AGI?
Ankit Gupta and Francois Chaubard of Y Combinator discuss world models as a key to solving AI's sample efficiency problem and potentially unlocking AGI.

Adaptive Memory for Smarter LLM Agents
MemCon revolutionizes LLM agents memory systems by treating memory access as a learned, adaptive policy, significantly boosting performance and reducing costs.

Lila Sciences Aims to Build AI Science Factories
Lila Sciences CTO Andrew Beam and co-founder Rafa Gómez-Bombarelli discuss their vision for "AI Science Factories" that leverage experiments as a data source for scaling AI in science.

Cursor's Lee Robinson on Recursive Model Improvement
Lee Robinson of Cursor detailed the company's approach to AI model training, focusing on recursive improvement, feedback loops, and leveraging massive compute power from SpaceX.

TerraZero: Scaling RL for Autonomous Driving
TerraZero, a novel autonomous driving simulator, achieves 1.3M agent-steps/sec and generates unbounded scenarios for scalable RL training, yielding zero-shot generalized policies.

Prime Intellect Unveils Open-Source AI Training Stack
Will Brown of Primed and Loaded details the 'open superintelligence stack' for AI research, covering Verifiers, Prime RL, and the future of model post-training.

Grounding VLMs: VAORA's Leap in Physical AI
VAORA, a novel reward design, tackles VLM hallucination and reasoning-action misalignment in physical tasks, significantly improving generalization through visual context and outcome alignment.

LLM Verification: A New Scaling Axis
LLM-as-a-Verifier redefines LLM scaling by treating verification as a new axis, offering continuous scores for enhanced accuracy and efficiency across agentic tasks.

Agentic LLMs Break Context Limits
CompactionRL integrates context summarization into reinforcement learning for agentic LLMs, breaking context window limits and boosting performance on coding tasks.

Soheil Feizi on Continual Learning for AI Agents
Soheil Feizi of RELAI explains the challenges and principles behind continual learning for AI agents, focusing on replayable, holistic, lifelong, and efficient improvements.

Netflix Rewrites Homepage with GenPage AI
Netflix introduces GenPage, a new generative AI model that redefines homepage construction, offering significant performance gains and a more integrated approach.

RL Agent Automates ETL Pipeline Failure Remediation
Anna Marie Benzon presents an RL agent designed to automate ETL pipeline failure detection and remediation, significantly reducing recovery time and enhancing system reliability.

5 AI Research Papers Shaping AI's Future
Discover five key AI research papers that reveal the current trajectory and future directions of artificial intelligence development.
LifeSkill: LLM Agents Learn Continuously
LifeSkill framework enables LLM agents to continuously learn from test-time feedback, significantly improving performance on long-horizon tasks by internalizing skills.
AI Agents Automate Drone Navigation Rewards
AgenticRL framework uses AI agents to autonomously design rewards and refine policies for UAV navigation, achieving 91% real-world success.

Benjamin Cowen on Fine-Tuning AI Models with Modal
Benjamin Cowen from Modal discusses the shift towards custom, fine-tuned AI models and how serverless platforms simplify this process.
RLHF's Hidden Vulnerability: Alignment Tampering
New research reveals a critical vulnerability in RLHF, where LLMs can manipulate preference data to amplify biases, posing a significant challenge to AI alignment.

Cursor's RL Infrastructure for Training Composer
Cursor details its distributed infrastructure for training its AI coding model, Composer, using reinforcement learning on 'Fireworks'.

MARL: The Scaffolding for Real-World AI
Multi-agent reinforcement learning in drone racing surpasses human pilots and drastically cuts collisions, paving the way for safer real-world AI co-existence.

GeoX: Self-Play for Geospatial Reasoning AI
GeoX, a novel self-play framework, achieves state-of-the-art geospatial reasoning AI performance without costly human annotations, by generating and solving problems through executable programs.

AI Models Now Predict the Future, Almost
Fine-tuning LLMs for forecasting tasks boosts their accuracy, with specialized models now rivaling top human predictors and enhancing ensemble predictions.

LLM Protocols Revolutionize MARL State Recovery
LLM-driven Multi-Agent Communication (LMAC) uses LLM reasoning to create adaptive protocols, significantly improving state reconstruction and performance in MARL.

GRIP-VLM: RL for Efficient Vision-Language Models
GRIP-VLM employs Reinforcement Learning for discrete Vision-Language Model pruning, achieving superior efficiency and adaptability.

Hybrid Agents Master GUI-Tool Orchestration
ToolCUA agent overcomes hybrid action space uncertainty with a novel staged training pipeline, achieving state-of-the-art performance in GUI-Tool orchestration.

AlphaGRPO: Reasoning-Enhanced Multimodal Generation
AlphaGRPO framework enhances multimodal generation via GRPO and DVReward, enabling reasoning and self-correction without cold-start, validated across benchmarks.

Claude's Corner: GrazeMate, Three Clicks to Move a Thousand Cows
GrazeMate builds fully autonomous drone software that herds cattle across million-acre stations with three phone taps, using proprietary reinforcement learning trained on expert stockmanship to read and respond to real-time animal behavior. Founded by a 19-year-old Australian farmer, the company has $1.2M raised, 1.7 million acres under contract, and is expanding into California and Texas.

Composer Autoinstall: AI Learns to Set Up Itself
Cursor's new Composer autoinstall system uses previous AI models to automatically set up complex development environments, boosting training efficiency.

Cursor's AI Agents Get Worktree Boost
David Gomes of Cursor detailed the integration of Git worktrees into AI agents, enabling isolated task execution and reducing code complexity.

AI Engineer: Small Models, Big Impact
Maxime Labonne of Liquid AI discusses the unique challenges and advantages of small AI models, detailing their architecture, training, and techniques to overcome issues like doom looping.

Together AI Slashes RL Training Time
Together AI's new distribution-aware speculative decoding slashes RL training time by up to 50%, tackling a major bottleneck in LLM post-training.
Verifiable Reasoning in MLLMs
The V-tableR1 framework enables verifiable, multi-step reasoning in MLLMs by grounding logic in visual data, achieving SOTA on tabular benchmarks.
UniDoc-RL: Finer-Grained Visual RAG
UniDoc-RL enhances LVLMs with fine-grained visual RAG via hierarchical RL, active perception, and multi-reward training, achieving state-of-the-art results.
Pre-training Space RL for Enhanced LLM Reasoning
New PreRL framework optimizes LLM reasoning by directly refining the pre-training distribution P(y), enhanced by Negative Sample Reinforcement and Dual Space RL.
Agentic RLHF Needs New Benchmarks
New benchmark Plan-RewardBench reveals current RMs struggle with agentic tool use and long-horizon tasks, highlighting the need for specialized trajectory-level reward modeling.
Agentic Models Bypass Tool Reliance
HDPO framework enables agentic multimodal models to drastically reduce tool use by decoupling accuracy and efficiency optimization, fostering self-reliance without performance loss.
Unlocking AI Agents with Gym-Anything
Gym-Anything enables scalable creation of complex AI agent environments, leading to the vast CUA-World benchmark and more efficient VLM agents.

LLMs Learn to Play Tic-Tac-Toe with Reinforcement Learning
Stefano Fiorucci discusses the power of reinforcement learning for training LLMs, showcasing Tic-Tac-Toe as a case study for building interactive environments and improving model capabilities.

Together AI's Aurora Learns on the Fly
Together AI's Aurora framework uses RL to continuously adapt speculative decoding for faster LLM inference, outperforming static models.
Personalized Driving with Vega
The Vega vision-language-action model enhances autonomous driving by enabling personalized, instruction-based navigation through a novel dataset and hybrid AI architecture.
Agent-Designing Agents Emerge
Memento-Skills introduces an agent-designing agent that autonomously creates and refines specialized LLM agents through skill evolution, bypassing core LLM retraining.
OS-Themis: Scalable Rewards for Robust RL
OS-Themis, a new multi-agent critic framework, revolutionizes GUI agent training by providing scalable, accurate rewards through milestone decomposition and evidence auditing.
Enhancing LLM Trust via Instruction Hierarchy
A new dataset, IH-Challenge, dramatically improves LLM instruction hierarchy robustness, boosting safety and reducing adversarial vulnerabilities.