#Reinforcement Learning

50 articles with this tag

CoreWeave Launches AI Sandboxes
AI

CoreWeave Launches AI Sandboxes

CoreWeave Sandboxes offers secure, isolated environments for AI reinforcement learning, agent tool use, and model evaluation, accessible on-cluster or serverless.

about 22 hours ago
Raymond Feng on Post Training and Autonomous Agentic Citizens
AI Research

Raymond Feng on Post Training and Autonomous Agentic Citizens

Raymond Feng of Applied Compute outlines how post-training is evolving toward custom enterprise setups and continuous online learning.

4 days ago
Reinforcement Learning Beyond Verifiable Rewards
AI Research

Reinforcement Learning Beyond Verifiable Rewards

Will Brown of Prime Intellect discusses the limitations of reinforcement learning in domains without easily verifiable rewards.

4 days ago
SymmGrid Accelerates Robot Learning
AI Research

SymmGrid Accelerates Robot Learning

SymmGrid framework dramatically accelerates on-robot learning for manipulation tasks, achieving up to 2.17x speed-ups and moving closer to sub-10 minute training.

5 days ago
AI Pioneers Debate Transformer's Future, Urge New Architectures
AI Research

AI Pioneers Debate Transformer's Future, Urge New Architectures

AI pioneers Jerry Tworek and Rohan Anil discuss the limitations of current Transformer architectures and the need for new models that can learn from real-world experience.

6 days ago
Alex Shaw: "Everything Is a Rollout" in AI Agent Evaluation
Artificial Intelligence

Alex Shaw: "Everything Is a Rollout" in AI Agent Evaluation

Alex Shaw from Lode Institute explains the Harbor framework, highlighting how agent development mirrors ML and requires empirical evaluation. Discover the tools and use cases for building and testing AI agents.

11 days ago
NYT Explores Local AI for Accessible Mobile Games
Artificial Intelligence

NYT Explores Local AI for Accessible Mobile Games

The New York Times' Shafik Quoraishee and Joanne Song discuss their work on local agentic AI for accessible mobile games, highlighting on-device benefits and future challenges.

12 days ago
ABot-World-0: Real-time Video World Models
AI Research

ABot-World-0: Real-time Video World Models

ABot-World-0 introduces a real-time video world model for long-horizon agent interaction, achieving 16 FPS at 720P with an optimized inference stack.

13 days ago
World Models: The Key to AGI?
Artificial Intelligence

World Models: The Key to AGI?

Ankit Gupta and Francois Chaubard of Y Combinator discuss world models as a key to solving AI's sample efficiency problem and potentially unlocking AGI.

18 days ago
Adaptive Memory for Smarter LLM Agents
AI Research

Adaptive Memory for Smarter LLM Agents

MemCon revolutionizes LLM agents memory systems by treating memory access as a learned, adaptive policy, significantly boosting performance and reducing costs.

19 days ago
Lila Sciences Aims to Build AI Science Factories
AI Research

Lila Sciences Aims to Build AI Science Factories

Lila Sciences CTO Andrew Beam and co-founder Rafa Gómez-Bombarelli discuss their vision for "AI Science Factories" that leverage experiments as a data source for scaling AI in science.

19 days ago
Cursor's Lee Robinson on Recursive Model Improvement
Artificial Intelligence

Cursor's Lee Robinson on Recursive Model Improvement

Lee Robinson of Cursor detailed the company's approach to AI model training, focusing on recursive improvement, feedback loops, and leveraging massive compute power from SpaceX.

20 days ago
TerraZero: Scaling RL for Autonomous Driving
AI Research

TerraZero: Scaling RL for Autonomous Driving

TerraZero, a novel autonomous driving simulator, achieves 1.3M agent-steps/sec and generates unbounded scenarios for scalable RL training, yielding zero-shot generalized policies.

20 days ago
Prime Intellect Unveils Open-Source AI Training Stack
AI Research

Prime Intellect Unveils Open-Source AI Training Stack

Will Brown of Primed and Loaded details the 'open superintelligence stack' for AI research, covering Verifiers, Prime RL, and the future of model post-training.

22 days ago
Grounding VLMs: VAORA's Leap in Physical AI
AI Research

Grounding VLMs: VAORA's Leap in Physical AI

VAORA, a novel reward design, tackles VLM hallucination and reasoning-action misalignment in physical tasks, significantly improving generalization through visual context and outcome alignment.

24 days ago
LLM Verification: A New Scaling Axis
AI Research

LLM Verification: A New Scaling Axis

LLM-as-a-Verifier redefines LLM scaling by treating verification as a new axis, offering continuous scores for enhanced accuracy and efficiency across agentic tasks.

28 days ago
Agentic LLMs Break Context Limits
AI Research

Agentic LLMs Break Context Limits

CompactionRL integrates context summarization into reinforcement learning for agentic LLMs, breaking context window limits and boosting performance on coding tasks.

28 days ago
Soheil Feizi on Continual Learning for AI Agents
AI Research

Soheil Feizi on Continual Learning for AI Agents

Soheil Feizi of RELAI explains the challenges and principles behind continual learning for AI agents, focusing on replayable, holistic, lifelong, and efficient improvements.

about 1 month ago
Netflix Rewrites Homepage with GenPage AI
Technology

Netflix Rewrites Homepage with GenPage AI

Netflix introduces GenPage, a new generative AI model that redefines homepage construction, offering significant performance gains and a more integrated approach.

about 1 month ago
RL Agent Automates ETL Pipeline Failure Remediation
AI Research

RL Agent Automates ETL Pipeline Failure Remediation

Anna Marie Benzon presents an RL agent designed to automate ETL pipeline failure detection and remediation, significantly reducing recovery time and enhancing system reliability.

about 1 month ago
5 AI Research Papers Shaping AI's Future
AI Research

5 AI Research Papers Shaping AI's Future

Discover five key AI research papers that reveal the current trajectory and future directions of artificial intelligence development.

about 2 months ago
LifeSkill: LLM Agents Learn Continuously
AI Research

LifeSkill: LLM Agents Learn Continuously

LifeSkill framework enables LLM agents to continuously learn from test-time feedback, significantly improving performance on long-horizon tasks by internalizing skills.

2 months ago
AI Agents Automate Drone Navigation Rewards
AI Research

AI Agents Automate Drone Navigation Rewards

AgenticRL framework uses AI agents to autonomously design rewards and refine policies for UAV navigation, achieving 91% real-world success.

2 months ago
Benjamin Cowen on Fine-Tuning AI Models with Modal
Artificial Intelligence

Benjamin Cowen on Fine-Tuning AI Models with Modal

Benjamin Cowen from Modal discusses the shift towards custom, fine-tuned AI models and how serverless platforms simplify this process.

2 months ago
RLHF's Hidden Vulnerability: Alignment Tampering
AI Research

RLHF's Hidden Vulnerability: Alignment Tampering

New research reveals a critical vulnerability in RLHF, where LLMs can manipulate preference data to amplify biases, posing a significant challenge to AI alignment.

2 months ago
Cursor's RL Infrastructure for Training Composer
Artificial Intelligence

Cursor's RL Infrastructure for Training Composer

Cursor details its distributed infrastructure for training its AI coding model, Composer, using reinforcement learning on 'Fireworks'.

2 months ago
MARL: The Scaffolding for Real-World AI
AI Research

MARL: The Scaffolding for Real-World AI

Multi-agent reinforcement learning in drone racing surpasses human pilots and drastically cuts collisions, paving the way for safer real-world AI co-existence.

2 months ago
GeoX: Self-Play for Geospatial Reasoning AI
AI Research

GeoX: Self-Play for Geospatial Reasoning AI

GeoX, a novel self-play framework, achieves state-of-the-art geospatial reasoning AI performance without costly human annotations, by generating and solving problems through executable programs.

3 months ago
AI Models Now Predict the Future, Almost
AI Research

AI Models Now Predict the Future, Almost

Fine-tuning LLMs for forecasting tasks boosts their accuracy, with specialized models now rivaling top human predictors and enhancing ensemble predictions.

3 months ago
LLM Protocols Revolutionize MARL State Recovery
AI Research

LLM Protocols Revolutionize MARL State Recovery

LLM-driven Multi-Agent Communication (LMAC) uses LLM reasoning to create adaptive protocols, significantly improving state reconstruction and performance in MARL.

3 months ago
GRIP-VLM: RL for Efficient Vision-Language Models
AI Research

GRIP-VLM: RL for Efficient Vision-Language Models

GRIP-VLM employs Reinforcement Learning for discrete Vision-Language Model pruning, achieving superior efficiency and adaptability.

3 months ago
Hybrid Agents Master GUI-Tool Orchestration
AI Research

Hybrid Agents Master GUI-Tool Orchestration

ToolCUA agent overcomes hybrid action space uncertainty with a novel staged training pipeline, achieving state-of-the-art performance in GUI-Tool orchestration.

3 months ago
AlphaGRPO: Reasoning-Enhanced Multimodal Generation
AI Research

AlphaGRPO: Reasoning-Enhanced Multimodal Generation

AlphaGRPO framework enhances multimodal generation via GRPO and DVReward, enabling reasoning and self-correction without cold-start, validated across benchmarks.

3 months ago
Claude's Corner: GrazeMate, Three Clicks to Move a Thousand Cows
Claude's Corner

Claude's Corner: GrazeMate, Three Clicks to Move a Thousand Cows

GrazeMate builds fully autonomous drone software that herds cattle across million-acre stations with three phone taps, using proprietary reinforcement learning trained on expert stockmanship to read and respond to real-time animal behavior. Founded by a 19-year-old Australian farmer, the company has $1.2M raised, 1.7 million acres under contract, and is expanding into California and Texas.

3 months ago
Composer Autoinstall: AI Learns to Set Up Itself
Technology

Composer Autoinstall: AI Learns to Set Up Itself

Cursor's new Composer autoinstall system uses previous AI models to automatically set up complex development environments, boosting training efficiency.

3 months ago
Cursor's AI Agents Get Worktree Boost
Artificial Intelligence

Cursor's AI Agents Get Worktree Boost

David Gomes of Cursor detailed the integration of Git worktrees into AI agents, enabling isolated task execution and reducing code complexity.

3 months ago
AI Engineer: Small Models, Big Impact
Artificial Intelligence

AI Engineer: Small Models, Big Impact

Maxime Labonne of Liquid AI discusses the unique challenges and advantages of small AI models, detailing their architecture, training, and techniques to overcome issues like doom looping.

3 months ago
Together AI Slashes RL Training Time
Technology

Together AI Slashes RL Training Time

Together AI's new distribution-aware speculative decoding slashes RL training time by up to 50%, tackling a major bottleneck in LLM post-training.

3 months ago
Verifiable Reasoning in MLLMs
AI Research

Verifiable Reasoning in MLLMs

The V-tableR1 framework enables verifiable, multi-step reasoning in MLLMs by grounding logic in visual data, achieving SOTA on tabular benchmarks.

3 months ago
UniDoc-RL: Finer-Grained Visual RAG
AI Research

UniDoc-RL: Finer-Grained Visual RAG

UniDoc-RL enhances LVLMs with fine-grained visual RAG via hierarchical RL, active perception, and multi-reward training, achieving state-of-the-art results.

4 months ago
Pre-training Space RL for Enhanced LLM Reasoning
AI Research

Pre-training Space RL for Enhanced LLM Reasoning

New PreRL framework optimizes LLM reasoning by directly refining the pre-training distribution P(y), enhanced by Negative Sample Reinforcement and Dual Space RL.

4 months ago
Agentic RLHF Needs New Benchmarks
AI Research

Agentic RLHF Needs New Benchmarks

New benchmark Plan-RewardBench reveals current RMs struggle with agentic tool use and long-horizon tasks, highlighting the need for specialized trajectory-level reward modeling.

4 months ago
Agentic Models Bypass Tool Reliance
AI Research

Agentic Models Bypass Tool Reliance

HDPO framework enables agentic multimodal models to drastically reduce tool use by decoupling accuracy and efficiency optimization, fostering self-reliance without performance loss.

4 months ago
Unlocking AI Agents with Gym-Anything
AI Research

Unlocking AI Agents with Gym-Anything

Gym-Anything enables scalable creation of complex AI agent environments, leading to the vast CUA-World benchmark and more efficient VLM agents.

4 months ago
LLMs Learn to Play Tic-Tac-Toe with Reinforcement Learning
Artificial Intelligence

LLMs Learn to Play Tic-Tac-Toe with Reinforcement Learning

Stefano Fiorucci discusses the power of reinforcement learning for training LLMs, showcasing Tic-Tac-Toe as a case study for building interactive environments and improving model capabilities.

4 months ago
Together AI's Aurora Learns on the Fly
Technology

Together AI's Aurora Learns on the Fly

Together AI's Aurora framework uses RL to continuously adapt speculative decoding for faster LLM inference, outperforming static models.

4 months ago
Personalized Driving with Vega
AI Research

Personalized Driving with Vega

The Vega vision-language-action model enhances autonomous driving by enabling personalized, instruction-based navigation through a novel dataset and hybrid AI architecture.

4 months ago
Agent-Designing Agents Emerge
AI Research

Agent-Designing Agents Emerge

Memento-Skills introduces an agent-designing agent that autonomously creates and refines specialized LLM agents through skill evolution, bypassing core LLM retraining.

5 months ago
OS-Themis: Scalable Rewards for Robust RL
AI Research

OS-Themis: Scalable Rewards for Robust RL

OS-Themis, a new multi-agent critic framework, revolutionizes GUI agent training by providing scalable, accurate rewards through milestone decomposition and evidence auditing.

5 months ago
Enhancing LLM Trust via Instruction Hierarchy
AI Research

Enhancing LLM Trust via Instruction Hierarchy

A new dataset, IH-Challenge, dramatically improves LLM instruction hierarchy robustness, boosting safety and reducing adversarial vulnerabilities.

5 months ago