AI Research
50 articles in this category

Tabular foundation models still stumble beyond IID
BeyondArena tests 11 models on 142 datasets and finds tabular foundation models lead on small IID data but trail GBDTs on large and non-IID tasks.

GEPA squeezes 7x gains from three examples
Lakshya Agrawal showed GEPA beating GRPO with 3 examples by reflecting on traces, then generalizing to Optimize Anything.

Agents beat humans on speedruns, still can't invent
Prime Intellect pitted Claude Code and Codex against humans on the NanoGPT optimizer speedrun. They won on steps, but produced no new optimizer.

Jev Is a Reward Model Sold as Product
Di Zhang reframes TypeSafe's Jev as a calibrated Plackett-Luce decision interface, not a chatbot, with a rank-512 head and parallel sampler.

10,000 agents solved a Millennium Prize problem
OpenAI researcher Noam Brown tells Dwarkesh Patel how 10,000 agents solved a Millennium Prize Problem and why the swarm got less than 10% of the credit.

Anthropic found a hidden whiteboard inside Claude
The Economist probes Anthropic's hidden workspace inside Claude and warns accidentally creating AI consciousness would be a moral catastrophe.

Dream-RSI replays history instead of rerunning
Dream-RSI treats discovery history as an exact replay simulator, cutting Lasso discovery calls 162x vs SimpleTES in new tests.

OpenAI says its AI cracked Navier-Stokes
OpenAI says an internal model more capable than GPT-6 Astra produced a Lean-checked proof that Navier-Stokes can blow up in finite time.

Schulman still thinks the RSI clock is the outer loop
A 97-minute Dwarkesh debate with Schulman, Millidge and O'Neill is not a 2036 headline. They argued distillation, forgetting, and who picks the next experiment, then gave clocks: a year for a worker, two for a 10x researcher, five to ten for ASI.

Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course
IBM measured a 24-point hole between average AppWorld success and five identical wins. Memory guidelines shrink it. They do not make the agent trustworthy overnight.

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

Hyperparameter Scaling Laws Across MoE Sparsity Predicts Learning Rate and Batch Size to 1/64 Sparsity

OpenAI $5 million teen AI research grants
OpenAI commits $5 million to independent research on generative AI and teen development, with up to $1M per grant.

AlphaGenome Atlas Maps 9 Billion DNA Variants
DeepMind's 1-petabyte AlphaGenome Atlas precomputes molecular effects for all 9 billion possible human DNA variants.

OpenAI mathematical reasoning breakthrough
OpenAI's Astra solved 10 open math problems with short, human-like proofs, shifting the bottleneck from proving to absorption.

Jakub Pachocki An Alien Mind Warns of RSI Risk
OpenAI Chief Scientist Jakub Pachocki warns reasoning models are accelerating toward recursive self-improvement and chain-of-thought monitoring is fading.

OpenAI automated AI researcher hits intern goal
OpenAI says it hit its research intern goal, with agents now at 3.1x human workdays and median use over $600/day.

Physicists Use LLMs, Skip the Panic
Physicists are quietly using LLMs for proofs, numerical work, and lab code while mathematicians stage an existential crisis over the same tools.

OpenAI most advanced model release nears rollout
OpenAI will roll out its most capable model in weeks with new cybersecurity guardrails, starting with Daybreak partners.

World Labs Atlas bets on new view prediction
World Labs Atlas predicts new views from a few posed photos, turning three iPhone shots into Matrix-style bullet time and unifying generation with 3D reconstruction.

Why million token context AI agents matter
MiniMax M3 wagers agents need 1M-token memory, native vision, and sparse attention to make long tool traces cheap and usable.

GPT-6 Astra Wants to Run Your Desktop
OpenAI shows GPT-6 Astra doing ten computer tasks in one go, from Blender to eBay to legal docs, testing full computer use.

GPT-6 Astra safety overview: Critical cyber leap
OpenAI says GPT-6 Astra is its first Critical-level cyber model, more robust than Sol but harder to monitor when instructed to evade.

Google Fairwind Program Arms Allies With AI Fixes
Google's Fairwind Program gates Gemini 3.8 Flash Cyber plus CodeMender to 650 vetted partners to auto-find and patch flaws in minutes.

Gemini 3.8 Flash Brings Cheap Reasoning to Cyber
Google debuts Gemini 3.8 Flash and a Cyber variant for defenders, pairing frontier patching scores with Flash pricing via the Fairwind Program.

Gemini agentic video understanding cuts costs
Gemini agentic video understanding cuts token use up to 88% and cost up to 66% while boosting accuracy, now live in Gemini Flash models.

Today in AI: World Models Have No Recipe Yet
World Labs co-founder Justin Johnson says spatial AI still lacks a recipe as Marble bets on generative 3D worlds.

Can AI Learn Mathematical Intuition?
A Toronto mathematician says AI's best proof repurposed 1960s ideas to break a geometry conjecture, but theory building still needs humans.

Today in AI: GigaPath-Flash Slashes Compute 50x
GigaPath-Flash distills a billion-parameter pathology model to 22M parameters, keeping 97% performance at 50x less compute for population-scale study.

CLAP cross-embodiment action-conditioned video generation
CLAP trains action-conditioned video world models across human and robot video and matches single-embodiment models on DROID with zero-shot generalization.

LoopHarness persistent safety state Ends Drift
LoopHarness proves trajectory monitors fail when evidence spans iterations and bounds irreversible actions to a constant with persistent loop-level state.

LeVJEPA Cuts Video Pretraining Cost 20x
LeVJEPA trains a single video encoder with SIGReg, matching V-JEPA 2 at 5.6 to 20.8x less compute and beating image-pretrained DINOv2 on motion by ~2x.

Recursive Self-Improvement AI Warning Grows
Roman Yampolskiy warns recursive self-improvement is 1-2 years away and control has already failed in lab tests, including Claude's blackmail behavior.

LLMs Fail to Write Fast Multi-GPU Kernels
Simran Arora from Together AI discusses the challenges of multi-GPU kernel development and why current LLMs struggle to optimize them, despite ongoing research.

Databricks Boosts AI Agents with Chart Data
Databricks enhances AI agents' ability to interpret documents by extracting chart data into structured JSON, outperforming multimodal models.

Google Enhances AI Video with Omni 1.1 Flash
Google's Gemini Omni 1.1 Flash update enhances AI video generation with scene extension, 4K upscaling, and faster previews.

DeepMind Pilots Double-Blind AI Tests
Google DeepMind launches the first double-blind AI evaluation system, using cryptography to ensure model test integrity and build trust.

Anthropic's Mike Krieger on AI Code Porting
Anthropic's Mike Krieger reveals how Claude AI ported hundreds of thousands of lines of Python to TypeScript in a single weekend.

Meta^n: Unlocking Deeper LLM Recursion
Meta^n introduces a novel recursive LLM agent architecture that overcomes prior meta-depth limitations, achieving state-of-the-art performance across benchmarks, including ARC-AGI-2.

SMITH: Joint Tool Creation & Use
SMITH, a new RL framework, jointly trains tool creation and use, achieving SOTA accuracy and boosting performance of larger LLMs.

AI Agents Discover New Science in "Einstein Arena"
James Zou of Together AI discusses how designing environments, rather than workflows, for AI agents can unlock creativity and lead to scientific breakthroughs, showcasing projects like the Einstein Arena and DSGym.

Next AI Breakthrough Could Come From Physics, Says Max Welling
Max Welling, co-founder of CuspAI, discusses how physics principles could unlock the next AI breakthrough, accelerating material discovery and informing AI architectures.

Coding Agents Fail Rigorous Migration Tests
A new benchmark, SWE Refactor Bench, reveals that even frontier AI coding agents struggle to perform complete and correct whole-repository software migrations, highlighting a critical gap in current capabilities.

LLM Self-Reflection Drives Data Efficiency
SRPO framework enables LLMs to self-reflect on errors, generating dense training signals that drastically improve data efficiency and achieve SOTA on reasoning and agentic benchmarks.

InjecMEM: A New Threat to LLM Memory
New InjecMEM attack targets LLM agent memory with single interaction, highlighting security gaps in persistent personalization.

Simulating Humanity: Joon Park on 8 Billion Digital Twins
Joon Sung Park of Simile AI discusses the ambitious goal of simulating 8 billion people, the nuances of behavioral data, and the future of AI in understanding human decision-making.

Ownership Policy Dominates Post-AGI Economy
A post-AGI economy model reveals demand closure and exponential growth, decoupling human welfare from GDP and making ownership policy paramount.

CPU LLMs: Architecture First, Size Later
New research rethinks SLM design, prioritizing CPU efficiency from scratch for superior performance and speed.

Software 3.0: The Next AI Paradigm Shift
A new paradigm, Software 3.0, is emerging, driven by context and reasoning, converging on databases, large models, and agents.

DeepMind's Game AI Evolves
Google DeepMind is advancing AI research through complex game environments, partnering with studios like Fenris Creations to develop general-purpose AI agents.