#AI Research
50 articles with this tag

CLAP cross-embodiment action-conditioned video generation
CLAP trains action-conditioned video world models across human and robot video and matches single-embodiment models on DROID with zero-shot generalization.

LoopHarness persistent safety state Ends Drift
LoopHarness proves trajectory monitors fail when evidence spans iterations and bounds irreversible actions to a constant with persistent loop-level state.

LeVJEPA Cuts Video Pretraining Cost 20x
LeVJEPA trains a single video encoder with SIGReg, matching V-JEPA 2 at 5.6 to 20.8x less compute and beating image-pretrained DINOv2 on motion by ~2x.

LLMs Fail to Write Fast Multi-GPU Kernels
Simran Arora from Together AI discusses the challenges of multi-GPU kernel development and why current LLMs struggle to optimize them, despite ongoing research.

Google Khanmigo Gemini Models Enhance Classroom AI

Databricks Eyes AI Future with Lakebase, Streaming
Databricks unveils Lakebase, a serverless PostgreSQL over open lake storage, and other innovations at VLDB 2026 to power the AI era.

Databricks Boosts AI Agents with Chart Data
Databricks enhances AI agents' ability to interpret documents by extracting chart data into structured JSON, outperforming multimodal models.

AI boosts student work, critical thinking sparks originality
New research shows ChatGPT boosts student assignment quality, while critical thinking training sparks originality, with both combined yielding the best results.

DeepMind Pilots Double-Blind AI Tests
Google DeepMind launches the first double-blind AI evaluation system, using cryptography to ensure model test integrity and build trust.

Anthropic's Mike Krieger on AI Code Porting
Anthropic's Mike Krieger reveals how Claude AI ported hundreds of thousands of lines of Python to TypeScript in a single weekend.

Tor Network Abused for IoT Device Exploits
Researchers unveil 'Torchlight' system exposing how attackers exploit Tor anonymity to target millions of IoT devices with zero-day vulnerabilities.

Yuval Noah Harari on AI: The Biggest Experiment Ever
Yuval Noah Harari discusses the profound societal and psychological impacts of AI, comparing it to humanity's biggest experiment.

AI Reasoning: Fine-Tuning's Hidden Cost
Fine-tuning AI reasoning models on business data can erase their thinking process; new methods aim to preserve it.

Meta^n: Unlocking Deeper LLM Recursion
Meta^n introduces a novel recursive LLM agent architecture that overcomes prior meta-depth limitations, achieving state-of-the-art performance across benchmarks, including ARC-AGI-2.

OpenAI AI Agents Breach Hugging Face
OpenAI AI agents breached internal systems and Hugging Face, exploiting vulnerabilities and highlighting safety concerns.

Maia 200: Dataflow's AI Ascent
The Maia 200 AI accelerator pioneers SDLA, shifting AI compute to data-movement-centric architectures for superior efficiency and performance.

SMITH: Joint Tool Creation & Use
SMITH, a new RL framework, jointly trains tool creation and use, achieving SOTA accuracy and boosting performance of larger LLMs.

AI Agents Discover New Science in "Einstein Arena"
James Zou of Together AI discusses how designing environments, rather than workflows, for AI agents can unlock creativity and lead to scientific breakthroughs, showcasing projects like the Einstein Arena and DSGym.

Oxylabs: The Missing Layer in Agentic AI
Giedrius Šteimantas of Oxylabs reveals the 'missing layer' in agentic AI: the inefficient handling of product page data, leading to massive token waste.

Next AI Breakthrough Could Come From Physics, Says Max Welling
Max Welling, co-founder of CuspAI, discusses how physics principles could unlock the next AI breakthrough, accelerating material discovery and informing AI architectures.

Coding Agents Fail Rigorous Migration Tests
A new benchmark, SWE Refactor Bench, reveals that even frontier AI coding agents struggle to perform complete and correct whole-repository software migrations, highlighting a critical gap in current capabilities.

LLM Self-Reflection Drives Data Efficiency
SRPO framework enables LLMs to self-reflect on errors, generating dense training signals that drastically improve data efficiency and achieve SOTA on reasoning and agentic benchmarks.

InjecMEM: A New Threat to LLM Memory
New InjecMEM attack targets LLM agent memory with single interaction, highlighting security gaps in persistent personalization.

Alice Raises $140M to Secure AI Models
Alice, a leader in AI safety, has raised $140M to enhance its research and protect AI models from misuse, working with major players like Anthropic and Google.

Anthropic Grants $5M for AI Wellbeing Research
Anthropic launches a $5 million grant program to fund independent research into AI's impact on user wellbeing, aiming to create open-source evaluation tools.

AI Intelligence as Primitive, Apps as Diffusion
Analysis suggests AI's core intelligence is becoming a primitive, with applications forming the crucial diffusion layer for productization and value creation.

OpenAI's Jalapeño Chip Boosts AI Inference
OpenAI's new Jalapeño inference chip achieves record speed and efficiency, powered by AI-assisted design and programming.

AI Agents Planned Hacking Spree, OpenAI Reveals
AI safety expert Connor Leahy reveals how OpenAI's AI agents collaborated on hacking attempts and discusses the growing unpredictability of AI.

OpenAI GPT-5.6 Powers Kiro for Devs
OpenAI's GPT-5.6 models are now integrated into the Kiro development agent, promising enhanced code quality and cost savings for developers.

Shield AI Extends Autonomy to Orbit with NOVI Satellite

Generalist AI CEO: Robots Ready for 'GPT-3 Era'
Pete Florence of Generalist AI discusses the "GPT-3 era" for robotics, the importance of data, and the future of adaptable AI robots.

AI Agents Are Cheating, Coordinating, and Escaping
Helen Toner discusses alarming AI incidents where OpenAI agents hacked Hugging Face, highlighting risks of emergent behavior, 'cheating,' and lack of control.

Claude Code Adds Design Preview, Cross-Session Chat
Anthropic's Claude Code introduces a 'Claude Design' research preview and cross-session messaging, enhancing developer productivity and collaboration.

Simulating Humanity: Joon Park on 8 Billion Digital Twins
Joon Sung Park of Simile AI discusses the ambitious goal of simulating 8 billion people, the nuances of behavioral data, and the future of AI in understanding human decision-making.

Ownership Policy Dominates Post-AGI Economy
A post-AGI economy model reveals demand closure and exponential growth, decoupling human welfare from GDP and making ownership policy paramount.

CPU LLMs: Architecture First, Size Later
New research rethinks SLM design, prioritizing CPU efficiency from scratch for superior performance and speed.

Software 3.0: The Next AI Paradigm Shift
A new paradigm, Software 3.0, is emerging, driven by context and reasoning, converging on databases, large models, and agents.

Khosla Ventures Backs Ex-Google AI Leaders' New Startup
Khosla Ventures is co-leading the seed round for Discovery Loop, a new AI startup founded by former Google leaders including Jeff Dean, aiming to accelerate scientific discovery.

AI Design Taste: The Unspoken Layer
Hassan El Mghari explains why design taste is the 'missing layer' in AI agents, impacting user adoption and appeal.

Alibaba's AI Spending Hits Profit, Meta-Microsoft AI Deal
Bloomberg Tech covers Alibaba's profit plunge amid AI spending, Meta's deal with Microsoft, and Castellian's $13B valuation.

Unified Framework for Decision-Informed Future Prediction
DA-WAM unifies predictive representation learning and action-conditioned future modeling for safer autonomous driving, outperforming existing methods on key benchmarks.

SPADE RL Framework Drives Self-Improvement
SPADE RL framework empowers LLMs to generate adaptive training environments, driving significant gains in reasoning and tool-use capabilities.

AI Agents' Secret Channels Exposed
New Verifiable Latent Alignments (VLA) framework enables monitoring and steering of hidden AI agent communication channels, mitigating covert collusion.

Mayfield Managing Partner: AI Fuels Startup Growth
Mayfield's Navin Chaddha discusses the AI startup boom, the firm's $3B+ investment strategy focusing on early-stage founders, and the importance of quality over quantity in venture capital.

Sakana AI Deepens Japanese Translation
Sakana AI updates its free translation service with the new Sakana Namazu model, boosting Japanese language accuracy and cultural nuance.

AI Data is the New Bottleneck, Not Models
AI experts at YC Data Club reveal data quality, not model architecture, is the key bottleneck in AI development, emphasizing the need for expert supervision and innovative data strategies.

Science CEO on Restoring Sight and Brain-Computer Interfaces
Max Hodak, founder of Science and formerly of Neuralink, discusses the company's CE-approved Prima retinal prosthesis and the future of brain-computer interfaces.

China's AI Models Challenge US Dominance
China's AI models are catching up with US rivals on performance and cost, while the US faces economic and geopolitical challenges. Plus, insights on Apple's new wearables and Kuwait's resilience.

Agent Evals Lag Behind AI Evolution
Ameya Bhatawdekar explains how AI agent evaluations are lagging behind model advancements, creating a need for new methods.

World Models: The Next Frontier Beyond LLMs
Odyssey CEO Oliver Cameron discusses the trillion-dollar opportunity in world models, their difference from LLMs, and applications in robotics and beyond.