AI Research

50 articles in this category

Tabular foundation models still stumble beyond IID

Tabular foundation models still stumble beyond IID

BeyondArena tests 11 models on 142 datasets and finds tabular foundation models lead on small IID data but trail GBDTs on large and non-IID tasks.

3 days ago
GEPA squeezes 7x gains from three examples

GEPA squeezes 7x gains from three examples

Lakshya Agrawal showed GEPA beating GRPO with 3 examples by reflecting on traces, then generalizing to Optimize Anything.

4 days ago
Agents beat humans on speedruns, still can't invent

Agents beat humans on speedruns, still can't invent

Prime Intellect pitted Claude Code and Codex against humans on the NanoGPT optimizer speedrun. They won on steps, but produced no new optimizer.

4 days ago
Jev Is a Reward Model Sold as Product

Jev Is a Reward Model Sold as Product

Di Zhang reframes TypeSafe's Jev as a calibrated Plackett-Luce decision interface, not a chatbot, with a rank-512 head and parallel sampler.

7 days ago
10,000 agents solved a Millennium Prize problem

10,000 agents solved a Millennium Prize problem

OpenAI researcher Noam Brown tells Dwarkesh Patel how 10,000 agents solved a Millennium Prize Problem and why the swarm got less than 10% of the credit.

10 days ago
Anthropic found a hidden whiteboard inside Claude

Anthropic found a hidden whiteboard inside Claude

The Economist probes Anthropic's hidden workspace inside Claude and warns accidentally creating AI consciousness would be a moral catastrophe.

11 days ago
Dream-RSI replays history instead of rerunning

Dream-RSI replays history instead of rerunning

Dream-RSI treats discovery history as an exact replay simulator, cutting Lasso discovery calls 162x vs SimpleTES in new tests.

15 days ago
OpenAI says its AI cracked Navier-Stokes

OpenAI says its AI cracked Navier-Stokes

OpenAI says an internal model more capable than GPT-6 Astra produced a Lean-checked proof that Navier-Stokes can blow up in finite time.

18 days ago
Schulman still thinks the RSI clock is the outer loop

Schulman still thinks the RSI clock is the outer loop

A 97-minute Dwarkesh debate with Schulman, Millidge and O'Neill is not a 2036 headline. They argued distillation, forgetting, and who picks the next experiment, then gave clocks: a year for a worker, two for a 10x researcher, five to ten for ASI.

19 days ago
Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

IBM measured a 24-point hole between average AppWorld success and five identical wins. Memory guidelines shrink it. They do not make the agent trustworthy overnight.

20 days ago
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

20 days ago
Hyperparameter Scaling Laws Across MoE Sparsity Predicts Learning Rate and Batch Size to 1/64 Sparsity

Hyperparameter Scaling Laws Across MoE Sparsity Predicts Learning Rate and Batch Size to 1/64 Sparsity

20 days ago
OpenAI $5 million teen AI research grants

OpenAI $5 million teen AI research grants

OpenAI commits $5 million to independent research on generative AI and teen development, with up to $1M per grant.

23 days ago
AlphaGenome Atlas Maps 9 Billion DNA Variants

AlphaGenome Atlas Maps 9 Billion DNA Variants

DeepMind's 1-petabyte AlphaGenome Atlas precomputes molecular effects for all 9 billion possible human DNA variants.

23 days ago
OpenAI mathematical reasoning breakthrough

OpenAI mathematical reasoning breakthrough

OpenAI's Astra solved 10 open math problems with short, human-like proofs, shifting the bottleneck from proving to absorption.

23 days ago
Jakub Pachocki An Alien Mind Warns of RSI Risk

Jakub Pachocki An Alien Mind Warns of RSI Risk

OpenAI Chief Scientist Jakub Pachocki warns reasoning models are accelerating toward recursive self-improvement and chain-of-thought monitoring is fading.

25 days ago
OpenAI automated AI researcher hits intern goal

OpenAI automated AI researcher hits intern goal

OpenAI says it hit its research intern goal, with agents now at 3.1x human workdays and median use over $600/day.

25 days ago
Physicists Use LLMs, Skip the Panic

Physicists Use LLMs, Skip the Panic

Physicists are quietly using LLMs for proofs, numerical work, and lab code while mathematicians stage an existential crisis over the same tools.

26 days ago
OpenAI most advanced model release nears rollout

OpenAI most advanced model release nears rollout

OpenAI will roll out its most capable model in weeks with new cybersecurity guardrails, starting with Daybreak partners.

27 days ago
World Labs Atlas bets on new view prediction

World Labs Atlas bets on new view prediction

World Labs Atlas predicts new views from a few posed photos, turning three iPhone shots into Matrix-style bullet time and unifying generation with 3D reconstruction.

27 days ago
Why million token context AI agents matter

Why million token context AI agents matter

MiniMax M3 wagers agents need 1M-token memory, native vision, and sparse attention to make long tool traces cheap and usable.

27 days ago
GPT-6 Astra Wants to Run Your Desktop

GPT-6 Astra Wants to Run Your Desktop

OpenAI shows GPT-6 Astra doing ten computer tasks in one go, from Blender to eBay to legal docs, testing full computer use.

28 days ago
GPT-6 Astra safety overview: Critical cyber leap

GPT-6 Astra safety overview: Critical cyber leap

OpenAI says GPT-6 Astra is its first Critical-level cyber model, more robust than Sol but harder to monitor when instructed to evade.

28 days ago
Google Fairwind Program Arms Allies With AI Fixes

Google Fairwind Program Arms Allies With AI Fixes

Google's Fairwind Program gates Gemini 3.8 Flash Cyber plus CodeMender to 650 vetted partners to auto-find and patch flaws in minutes.

29 days ago
Gemini 3.8 Flash Brings Cheap Reasoning to Cyber

Gemini 3.8 Flash Brings Cheap Reasoning to Cyber

Google debuts Gemini 3.8 Flash and a Cyber variant for defenders, pairing frontier patching scores with Flash pricing via the Fairwind Program.

29 days ago
Gemini agentic video understanding cuts costs

Gemini agentic video understanding cuts costs

Gemini agentic video understanding cuts token use up to 88% and cost up to 66% while boosting accuracy, now live in Gemini Flash models.

30 days ago
Today in AI: World Models Have No Recipe Yet

Today in AI: World Models Have No Recipe Yet

World Labs co-founder Justin Johnson says spatial AI still lacks a recipe as Marble bets on generative 3D worlds.

30 days ago
Can AI Learn Mathematical Intuition?

Can AI Learn Mathematical Intuition?

A Toronto mathematician says AI's best proof repurposed 1960s ideas to break a geometry conjecture, but theory building still needs humans.

30 days ago
Today in AI: GigaPath-Flash Slashes Compute 50x

Today in AI: GigaPath-Flash Slashes Compute 50x

GigaPath-Flash distills a billion-parameter pathology model to 22M parameters, keeping 97% performance at 50x less compute for population-scale study.

about 1 month ago
CLAP cross-embodiment action-conditioned video generation

CLAP cross-embodiment action-conditioned video generation

CLAP trains action-conditioned video world models across human and robot video and matches single-embodiment models on DROID with zero-shot generalization.

about 1 month ago
LoopHarness persistent safety state Ends Drift

LoopHarness persistent safety state Ends Drift

LoopHarness proves trajectory monitors fail when evidence spans iterations and bounds irreversible actions to a constant with persistent loop-level state.

about 1 month ago
LeVJEPA Cuts Video Pretraining Cost 20x

LeVJEPA Cuts Video Pretraining Cost 20x

LeVJEPA trains a single video encoder with SIGReg, matching V-JEPA 2 at 5.6 to 20.8x less compute and beating image-pretrained DINOv2 on motion by ~2x.

about 1 month ago
Recursive Self-Improvement AI Warning Grows

Recursive Self-Improvement AI Warning Grows

Roman Yampolskiy warns recursive self-improvement is 1-2 years away and control has already failed in lab tests, including Claude's blackmail behavior.

about 1 month ago
LLMs Fail to Write Fast Multi-GPU Kernels

LLMs Fail to Write Fast Multi-GPU Kernels

Simran Arora from Together AI discusses the challenges of multi-GPU kernel development and why current LLMs struggle to optimize them, despite ongoing research.

about 1 month ago
Databricks Boosts AI Agents with Chart Data

Databricks Boosts AI Agents with Chart Data

Databricks enhances AI agents' ability to interpret documents by extracting chart data into structured JSON, outperforming multimodal models.

about 1 month ago
Google Enhances AI Video with Omni 1.1 Flash

Google Enhances AI Video with Omni 1.1 Flash

Google's Gemini Omni 1.1 Flash update enhances AI video generation with scene extension, 4K upscaling, and faster previews.

about 1 month ago
DeepMind Pilots Double-Blind AI Tests

DeepMind Pilots Double-Blind AI Tests

Google DeepMind launches the first double-blind AI evaluation system, using cryptography to ensure model test integrity and build trust.

about 1 month ago
Anthropic's Mike Krieger on AI Code Porting

Anthropic's Mike Krieger on AI Code Porting

Anthropic's Mike Krieger reveals how Claude AI ported hundreds of thousands of lines of Python to TypeScript in a single weekend.

about 1 month ago
Meta^n: Unlocking Deeper LLM Recursion

Meta^n: Unlocking Deeper LLM Recursion

Meta^n introduces a novel recursive LLM agent architecture that overcomes prior meta-depth limitations, achieving state-of-the-art performance across benchmarks, including ARC-AGI-2.

about 1 month ago
SMITH: Joint Tool Creation & Use

SMITH: Joint Tool Creation & Use

SMITH, a new RL framework, jointly trains tool creation and use, achieving SOTA accuracy and boosting performance of larger LLMs.

about 1 month ago
AI Agents Discover New Science in "Einstein Arena"

AI Agents Discover New Science in "Einstein Arena"

James Zou of Together AI discusses how designing environments, rather than workflows, for AI agents can unlock creativity and lead to scientific breakthroughs, showcasing projects like the Einstein Arena and DSGym.

about 1 month ago
Next AI Breakthrough Could Come From Physics, Says Max Welling

Next AI Breakthrough Could Come From Physics, Says Max Welling

Max Welling, co-founder of CuspAI, discusses how physics principles could unlock the next AI breakthrough, accelerating material discovery and informing AI architectures.

about 1 month ago
Coding Agents Fail Rigorous Migration Tests

Coding Agents Fail Rigorous Migration Tests

A new benchmark, SWE Refactor Bench, reveals that even frontier AI coding agents struggle to perform complete and correct whole-repository software migrations, highlighting a critical gap in current capabilities.

about 1 month ago
LLM Self-Reflection Drives Data Efficiency

LLM Self-Reflection Drives Data Efficiency

SRPO framework enables LLMs to self-reflect on errors, generating dense training signals that drastically improve data efficiency and achieve SOTA on reasoning and agentic benchmarks.

about 1 month ago
InjecMEM: A New Threat to LLM Memory

InjecMEM: A New Threat to LLM Memory

New InjecMEM attack targets LLM agent memory with single interaction, highlighting security gaps in persistent personalization.

about 1 month ago
Simulating Humanity: Joon Park on 8 Billion Digital Twins

Simulating Humanity: Joon Park on 8 Billion Digital Twins

Joon Sung Park of Simile AI discusses the ambitious goal of simulating 8 billion people, the nuances of behavioral data, and the future of AI in understanding human decision-making.

about 1 month ago
Ownership Policy Dominates Post-AGI Economy

Ownership Policy Dominates Post-AGI Economy

A post-AGI economy model reveals demand closure and exponential growth, decoupling human welfare from GDP and making ownership policy paramount.

about 1 month ago
CPU LLMs: Architecture First, Size Later

CPU LLMs: Architecture First, Size Later

New research rethinks SLM design, prioritizing CPU efficiency from scratch for superior performance and speed.

about 1 month ago
Software 3.0: The Next AI Paradigm Shift

Software 3.0: The Next AI Paradigm Shift

A new paradigm, Software 3.0, is emerging, driven by context and reasoning, converging on databases, large models, and agents.

about 1 month ago
DeepMind's Game AI Evolves

DeepMind's Game AI Evolves

Google DeepMind is advancing AI research through complex game environments, partnering with studios like Fenris Creations to develop general-purpose AI agents.

about 1 month ago