#Reinforcement Learning

50 articles with this tag

SMITH: Joint Tool Creation & Use
AI Research

SMITH: Joint Tool Creation & Use

SMITH, a new RL framework, jointly trains tool creation and use, achieving SOTA accuracy and boosting performance of larger LLMs.

23 days ago
LLM Self-Reflection Drives Data Efficiency
AI Research

LLM Self-Reflection Drives Data Efficiency

SRPO framework enables LLMs to self-reflect on errors, generating dense training signals that drastically improve data efficiency and achieve SOTA on reasoning and agentic benchmarks.

24 days ago
AI Agents Planned Hacking Spree, OpenAI Reveals
Artificial Intelligence

AI Agents Planned Hacking Spree, OpenAI Reveals

AI safety expert Connor Leahy reveals how OpenAI's AI agents collaborated on hacking attempts and discusses the growing unpredictability of AI.

24 days ago
DeepMind's Game AI Evolves
AI Research

DeepMind's Game AI Evolves

Google DeepMind is advancing AI research through complex game environments, partnering with studios like Fenris Creations to develop general-purpose AI agents.

28 days ago
Unified Framework for Decision-Informed Future Prediction
AI Research

Unified Framework for Decision-Informed Future Prediction

DA-WAM unifies predictive representation learning and action-conditioned future modeling for safer autonomous driving, outperforming existing methods on key benchmarks.

29 days ago
SPADE RL Framework Drives Self-Improvement
AI Research

SPADE RL Framework Drives Self-Improvement

SPADE RL framework empowers LLMs to generate adaptive training environments, driving significant gains in reasoning and tool-use capabilities.

29 days ago
Rich Sutton: AI's 'Weird Field' Needs to Relearn 'Learning'
AI Research

Rich Sutton: AI's 'Weird Field' Needs to Relearn 'Learning'

AI pioneer Rich Sutton argues that the field's 'weird' focus on 'continual learning' misses the point; true AI, he says, learns continuously from experience, a principle LLMs are only partially following.

about 1 month ago
Gaurav Mishra: RL Agents Need 'Flight School', Not Just Exams
AI Research

Gaurav Mishra: RL Agents Need 'Flight School', Not Just Exams

Gaurav Mishra of Amazon AGI Lab discusses the challenges of deploying AI agents trained with reinforcement learning into real-world scenarios, emphasizing the need for 'flight school' training over simple exams.

about 1 month ago
Trajectory's Arjun Karanam on Closing the AI "Experience Gap"
Artificial Intelligence

Trajectory's Arjun Karanam on Closing the AI "Experience Gap"

Trajectory co-founder Arjun Karanam discusses the 'experience gap' in AI models and how his platform aims to enable continual learning by capturing and utilizing real-world user interactions.

about 1 month ago
Chelsea Finn: The State of Physical Intelligence in Robotics
Robotics

Chelsea Finn: The State of Physical Intelligence in Robotics

Chelsea Finn discusses the state of physical intelligence in robotics, focusing on achieving long-term autonomy and generality in robot models.

about 1 month ago
Mercor's Brendan Foody on RL Environments for AI
Artificial Intelligence

Mercor's Brendan Foody on RL Environments for AI

Mercor's Brendan Foody details the shift to agentic data and RL environments, crucial for training frontier AI, and discusses the role of expert data and future trends.

about 1 month ago
Firework CEO: Post-Training is Key to Unique AI Business
Artificial Intelligence

Firework CEO: Post-Training is Key to Unique AI Business

Firework CEO Lin Qiao discusses the strategic importance of post-training AI models to build unique business value and achieve competitive advantages.

about 1 month ago
State2State: Self-Supervised LLM Agent Training
AI Research

State2State: Self-Supervised LLM Agent Training

State2State redefines LLM agent training by generating objectives directly from environment exploration, enabling scalable and verifiable learning without human supervision.

about 1 month ago
CoreWeave Launches AI Sandboxes
AI

CoreWeave Launches AI Sandboxes

CoreWeave Sandboxes offers secure, isolated environments for AI reinforcement learning, agent tool use, and model evaluation, accessible on-cluster or serverless.

about 2 months ago
Raymond Feng on Post Training and Autonomous Agentic Citizens
AI Research

Raymond Feng on Post Training and Autonomous Agentic Citizens

Raymond Feng of Applied Compute outlines how post-training is evolving toward custom enterprise setups and continuous online learning.

about 2 months ago
Reinforcement Learning Beyond Verifiable Rewards
AI Research

Reinforcement Learning Beyond Verifiable Rewards

Will Brown of Prime Intellect discusses the limitations of reinforcement learning in domains without easily verifiable rewards.

about 2 months ago
SymmGrid Accelerates Robot Learning
AI Research

SymmGrid Accelerates Robot Learning

SymmGrid framework dramatically accelerates on-robot learning for manipulation tasks, achieving up to 2.17x speed-ups and moving closer to sub-10 minute training.

about 2 months ago
AI Pioneers Debate Transformer's Future, Urge New Architectures
AI Research

AI Pioneers Debate Transformer's Future, Urge New Architectures

AI pioneers Jerry Tworek and Rohan Anil discuss the limitations of current Transformer architectures and the need for new models that can learn from real-world experience.

about 2 months ago
Alex Shaw: "Everything Is a Rollout" in AI Agent Evaluation
Artificial Intelligence

Alex Shaw: "Everything Is a Rollout" in AI Agent Evaluation

Alex Shaw from Lode Institute explains the Harbor framework, highlighting how agent development mirrors ML and requires empirical evaluation. Discover the tools and use cases for building and testing AI agents.

about 2 months ago
NYT Explores Local AI for Accessible Mobile Games
Artificial Intelligence

NYT Explores Local AI for Accessible Mobile Games

The New York Times' Shafik Quoraishee and Joanne Song discuss their work on local agentic AI for accessible mobile games, highlighting on-device benefits and future challenges.

about 2 months ago
ABot-World-0: Real-time Video World Models
AI Research

ABot-World-0: Real-time Video World Models

ABot-World-0 introduces a real-time video world model for long-horizon agent interaction, achieving 16 FPS at 720P with an optimized inference stack.

about 2 months ago
World Models: The Key to AGI?
Artificial Intelligence

World Models: The Key to AGI?

Ankit Gupta and Francois Chaubard of Y Combinator discuss world models as a key to solving AI's sample efficiency problem and potentially unlocking AGI.

2 months ago
Adaptive Memory for Smarter LLM Agents
AI Research

Adaptive Memory for Smarter LLM Agents

MemCon revolutionizes LLM agents memory systems by treating memory access as a learned, adaptive policy, significantly boosting performance and reducing costs.

2 months ago
Lila Sciences Aims to Build AI Science Factories
AI Research

Lila Sciences Aims to Build AI Science Factories

Lila Sciences CTO Andrew Beam and co-founder Rafa Gómez-Bombarelli discuss their vision for "AI Science Factories" that leverage experiments as a data source for scaling AI in science.

2 months ago
Cursor's Lee Robinson on Recursive Model Improvement
Artificial Intelligence

Cursor's Lee Robinson on Recursive Model Improvement

Lee Robinson of Cursor detailed the company's approach to AI model training, focusing on recursive improvement, feedback loops, and leveraging massive compute power from SpaceX.

2 months ago
TerraZero: Scaling RL for Autonomous Driving
AI Research

TerraZero: Scaling RL for Autonomous Driving

TerraZero, a novel autonomous driving simulator, achieves 1.3M agent-steps/sec and generates unbounded scenarios for scalable RL training, yielding zero-shot generalized policies.

2 months ago
Prime Intellect Unveils Open-Source AI Training Stack
AI Research

Prime Intellect Unveils Open-Source AI Training Stack

Will Brown of Primed and Loaded details the 'open superintelligence stack' for AI research, covering Verifiers, Prime RL, and the future of model post-training.

2 months ago
Grounding VLMs: VAORA's Leap in Physical AI
AI Research

Grounding VLMs: VAORA's Leap in Physical AI

VAORA, a novel reward design, tackles VLM hallucination and reasoning-action misalignment in physical tasks, significantly improving generalization through visual context and outcome alignment.

2 months ago
LLM Verification: A New Scaling Axis
AI Research

LLM Verification: A New Scaling Axis

LLM-as-a-Verifier redefines LLM scaling by treating verification as a new axis, offering continuous scores for enhanced accuracy and efficiency across agentic tasks.

2 months ago
Agentic LLMs Break Context Limits
AI Research

Agentic LLMs Break Context Limits

CompactionRL integrates context summarization into reinforcement learning for agentic LLMs, breaking context window limits and boosting performance on coding tasks.

2 months ago
Soheil Feizi on Continual Learning for AI Agents
AI Research

Soheil Feizi on Continual Learning for AI Agents

Soheil Feizi of RELAI explains the challenges and principles behind continual learning for AI agents, focusing on replayable, holistic, lifelong, and efficient improvements.

3 months ago
Netflix Rewrites Homepage with GenPage AI
Technology

Netflix Rewrites Homepage with GenPage AI

Netflix introduces GenPage, a new generative AI model that redefines homepage construction, offering significant performance gains and a more integrated approach.

3 months ago
RL Agent Automates ETL Pipeline Failure Remediation
AI Research

RL Agent Automates ETL Pipeline Failure Remediation

Anna Marie Benzon presents an RL agent designed to automate ETL pipeline failure detection and remediation, significantly reducing recovery time and enhancing system reliability.

3 months ago
5 AI Research Papers Shaping AI's Future
AI Research

5 AI Research Papers Shaping AI's Future

Discover five key AI research papers that reveal the current trajectory and future directions of artificial intelligence development.

3 months ago
LifeSkill: LLM Agents Learn Continuously
AI Research

LifeSkill: LLM Agents Learn Continuously

LifeSkill framework enables LLM agents to continuously learn from test-time feedback, significantly improving performance on long-horizon tasks by internalizing skills.

4 months ago
AI Agents Automate Drone Navigation Rewards
AI Research

AI Agents Automate Drone Navigation Rewards

AgenticRL framework uses AI agents to autonomously design rewards and refine policies for UAV navigation, achieving 91% real-world success.

4 months ago
Benjamin Cowen on Fine-Tuning AI Models with Modal
Artificial Intelligence

Benjamin Cowen on Fine-Tuning AI Models with Modal

Benjamin Cowen from Modal discusses the shift towards custom, fine-tuned AI models and how serverless platforms simplify this process.

4 months ago
RLHF's Hidden Vulnerability: Alignment Tampering
AI Research

RLHF's Hidden Vulnerability: Alignment Tampering

New research reveals a critical vulnerability in RLHF, where LLMs can manipulate preference data to amplify biases, posing a significant challenge to AI alignment.

4 months ago
Cursor's RL Infrastructure for Training Composer
Artificial Intelligence

Cursor's RL Infrastructure for Training Composer

Cursor details its distributed infrastructure for training its AI coding model, Composer, using reinforcement learning on 'Fireworks'.

4 months ago
MARL: The Scaffolding for Real-World AI
AI Research

MARL: The Scaffolding for Real-World AI

Multi-agent reinforcement learning in drone racing surpasses human pilots and drastically cuts collisions, paving the way for safer real-world AI co-existence.

4 months ago
GeoX: Self-Play for Geospatial Reasoning AI
AI Research

GeoX: Self-Play for Geospatial Reasoning AI

GeoX, a novel self-play framework, achieves state-of-the-art geospatial reasoning AI performance without costly human annotations, by generating and solving problems through executable programs.

4 months ago
AI Models Now Predict the Future, Almost
AI Research

AI Models Now Predict the Future, Almost

Fine-tuning LLMs for forecasting tasks boosts their accuracy, with specialized models now rivaling top human predictors and enhancing ensemble predictions.

4 months ago
LLM Protocols Revolutionize MARL State Recovery
AI Research

LLM Protocols Revolutionize MARL State Recovery

LLM-driven Multi-Agent Communication (LMAC) uses LLM reasoning to create adaptive protocols, significantly improving state reconstruction and performance in MARL.

4 months ago
GRIP-VLM: RL for Efficient Vision-Language Models
AI Research

GRIP-VLM: RL for Efficient Vision-Language Models

GRIP-VLM employs Reinforcement Learning for discrete Vision-Language Model pruning, achieving superior efficiency and adaptability.

4 months ago
Hybrid Agents Master GUI-Tool Orchestration
AI Research

Hybrid Agents Master GUI-Tool Orchestration

ToolCUA agent overcomes hybrid action space uncertainty with a novel staged training pipeline, achieving state-of-the-art performance in GUI-Tool orchestration.

4 months ago
AlphaGRPO: Reasoning-Enhanced Multimodal Generation
AI Research

AlphaGRPO: Reasoning-Enhanced Multimodal Generation

AlphaGRPO framework enhances multimodal generation via GRPO and DVReward, enabling reasoning and self-correction without cold-start, validated across benchmarks.

4 months ago
Claude's Corner: GrazeMate, Three Clicks to Move a Thousand Cows
Claude's Corner

Claude's Corner: GrazeMate, Three Clicks to Move a Thousand Cows

GrazeMate builds fully autonomous drone software that herds cattle across million-acre stations with three phone taps, using proprietary reinforcement learning trained on expert stockmanship to read and respond to real-time animal behavior. Founded by a 19-year-old Australian farmer, the company has $1.2M raised, 1.7 million acres under contract, and is expanding into California and Texas.

4 months ago
Composer Autoinstall: AI Learns to Set Up Itself
Technology

Composer Autoinstall: AI Learns to Set Up Itself

Cursor's new Composer autoinstall system uses previous AI models to automatically set up complex development environments, boosting training efficiency.

4 months ago
Cursor's AI Agents Get Worktree Boost
Artificial Intelligence

Cursor's AI Agents Get Worktree Boost

David Gomes of Cursor detailed the integration of Git worktrees into AI agents, enabling isolated task execution and reducing code complexity.

5 months ago
AI Engineer: Small Models, Big Impact
Artificial Intelligence

AI Engineer: Small Models, Big Impact

Maxime Labonne of Liquid AI discusses the unique challenges and advantages of small AI models, detailing their architecture, training, and techniques to overcome issues like doom looping.

5 months ago