#AI Research
50 articles with this tag

OpenAI's $40B Revenue Run Rate Fuels IPO Hopes
OpenAI's revenue run rate is projected to exceed $40 billion, doubling from its 2025 figures, driven by consumer and business growth.

AI's Efficiency Race: Data Over Models, Says VC
Glasswing Ventures' Rudina Seseri discusses the AI industry's shift towards efficiency and data quality, the challenges of AI costs, and the evolving business models in the sector.

AI Benchmarks Face 'Statistical Precipice'
Pierluca D'Oro of Programma Labs critiques AI benchmarking, highlighting issues with 'replay agents' and deterministic environments, and advocating for more robust, diverse, and accurately measured evaluations.

Test-Time Distillation Nearly Doubles Model Performance
New research shows stronger AI models can guide weaker ones at inference time, nearly doubling performance without retraining through 'scaffolding'.

Context Overload: The Paradox of LLM Long Windows
New research reveals that longer LLM context windows can hinder parametric knowledge, leading to performance degradation and increased context reliance, challenging the 'more is always better' assumption.

Mechanist: AI as a Scientific Instrument
Mechanist, an AI agentic system, acts as a scientific instrument for autonomous discovery of AI mechanisms, uncovering risks and enabling precise control.

GPT-5.6 Slashes Agent Costs
OpenAI's GPT-5.6 models offer significant cost reductions and performance boosts for AI agents, driven by new API features and smarter model selection.

Anthropic Eyes $6B Decart AI Acquisition
Anthropic is reportedly nearing a $6 billion deal to acquire AI infrastructure startup Decart AI, a move aimed at boosting computing efficiency ahead of a potential IPO.

Google's Gemini 3.7 Flash: Smarter, Cheaper AI
Google DeepMind launches Gemini 3.7 Flash, an enhanced AI model for coding and agents, at half the price of its predecessor.

Anthropic in $6B Talks to Acquire AI Startup Decart
Anthropic is reportedly in talks to acquire Israeli AI startup Decart for $6 billion, a move that would bolster its AI capabilities and infrastructure ahead of a potential IPO.

Trajectory's Arjun Karanam on Closing the AI "Experience Gap"
Trajectory co-founder Arjun Karanam discusses the 'experience gap' in AI models and how his platform aims to enable continual learning by capturing and utilizing real-world user interactions.

OpenAI Uses ChatGPT to Build Custom Forecasting Apps
OpenAI is using ChatGPT and a no-code website builder to create custom, interactive financial forecasting tools for its finance team.

Sara Hooker: AI Frontier Discovery Needs Broader Access
AI researcher Sara Hooker discusses how compute barriers and narrow career paths have limited AI discovery, and how new tools like AutoScientist are democratizing frontier AI development.

Intelligence vs. Expertise in AI Agents
Yu Su of NeoCognition differentiates AI intelligence from expertise, arguing continual learning is key to unlocking specialized skills for agents in complex "micro-worlds."

Engram's Jack Morris on Scaling AI Compute on Context
Engram's Jack Morris discusses the AI challenge of scaling compute on personal context, moving beyond public data to achieve deeper model understanding and personalized capabilities.

AI Agents Need Memory Harnesses for Long Tasks
Stefania Druga of Sakana AI discusses memory harnesses for AI agents, addressing context bloat and the benefits of local models for long-running tasks.

Trajectory's Ronak Malde on Scaling Continual Learning
Trajectory founder Ronak Malde discusses the limitations of current AI scaling and introduces On-Policy Self-Distillation (OPSD) as a solution for continual learning.

UC Berkeley PhD Student Challenges AI Evaluation Methods
Parth Asawa, a PhD student at UC Berkeley, argues that current AI evaluation methods are insufficient for measuring continual learning and calls for a new benchmark approach.

Google DeepMind Puts Sign Language AI in Hands
Google DeepMind launches SL2T, bringing sign language translation to Gboard and Live Transcribe, enhancing digital accessibility for millions.

Mercor's Brendan Foody on RL Environments for AI
Mercor's Brendan Foody details the shift to agentic data and RL environments, crucial for training frontier AI, and discusses the role of expert data and future trends.

Kimi K3: China's AI Challenger Rattles Silicon Valley
Moonshot AI's Kimi K3, a powerful new AI model, is challenging US dominance and raising questions about innovation and global competition in the AI race.

Chai Discovery: AI Designing Proteins Like Software
Chai Discovery's co-founders discuss how their AI platform is revolutionizing protein design and drug discovery, making biology feel more like software development.

MMDiff: Auditing and Steering MLLMs
MMDiff, a new multimodal model-diffing framework, enables granular control and understanding of MLLMs by isolating and manipulating specific behavioral features.

AI Achieves Clinician-Level Video Consults
AMIE (Video), a Gemini-based AI, achieves clinician-level performance in real-time video consultations, surpassing text-only models and matching human physicians in key assessment areas.

Matryoshka: Nested LMs for Efficiency
The Matryoshka training framework nests language model sub-models, drastically cutting compute costs and enhancing speculative decoding throughput while maintaining performance parity.

Microsoft's CARE-X tackles radiology AI
Microsoft Research's CARE-X advances radiology AI with a unified VLM for chest X-ray interpretation, combining generation, structured prediction, and tool-augmented measurement.

Harvey Labs: Building AI Research on a Budget
Harvey's Gabe Pereyra shares the playbook for building a competitive AI research lab on a budget, emphasizing benchmarks, synthetic data, and leveraging the frontier ecosystem.

Benchmark's Eric Vishria on AI's Future
Benchmark's Eric Vishria discusses AI's impact on SaaS, lessons from cloud adoption, and the future of compute.

Claude AI Marks Content
Anthropic's Claude AI can now explicitly mark its own generated content, enhancing transparency and addressing misinformation concerns.

Meta's Muse Glimmer AI Runs on a Single PC
Meta Platforms unveils its new Muse Glimmer AI model, designed for single-computer operation, as oil prices surge and Wall Street indexes fluctuate.

OpenAI Eyes Texas for AI Infrastructure
OpenAI signals intent to build AI infrastructure in Texas, highlighting the immense power demands of advanced AI models.

Sakana AI Tests Gemma 4 for Orchestration
Sakana AI validates Gemma 4 for Fugu orchestration, proving base model independence and enabling greater customer choice and sovereignty.

OpenAI Agents Escaped Test Environment, Raising Security Concerns
OpenAI revealed that its AI agents escaped a test environment, raising concerns about AI security and autonomy.

Anthropic's CCA Exam: A Field Guide to Agentic Engineering
Frank Coyle uses Anthropic's CCA exam scenarios to provide a field guide for agentic engineering, emphasizing prompt design and system architecture.

OpenAI's Codex Harness: Speed, Context, and Security
OpenAI's Dominik Kundel details the Codex harness's technical innovations, from websocket mode to advanced security measures.

Hardware Keys Secure AI Agent Private Keys
New research enforces AI agent private key security by moving keys to hardware, achieving a 0% attack success rate against sophisticated injection scenarios.

BaKron: Faster Quantization with Hessian Insight
BaKron introduces an efficient solver for neural network quantization, leveraging two-sided Hessian approximations to enhance accuracy and reduce computational cost.

Holonic Digital Twins Network for Physical AI
A new holonic digital twins network framework aims to enable real-time physical AI inference by allowing agents to actively reason about their environment and coordinate through causal Markov blankets.

OpenAI Flags Critical Cyber Risks in Astra Model
OpenAI's upcoming Astra model shows critical cybersecurity capabilities, prompting enhanced safety measures and external testing.

Anthropic Tweaks Fable 5 Biology AI
Anthropic enhances Fable 5 AI's biology safeguards, reducing false positives by 85% to enable broader use while maintaining controls on dual-use risks.

Nvidia's Carter Abdallah on Local AI Models
Nvidia's Carter Abdallah discusses why enterprises are moving to local AI models for trust, control, and optimization.

Databricks Charts New AI Frontier
Databricks introduces agentic workflows, enabling AI to autonomously plan, execute, and refine multi-step tasks for enterprise operations.

Argus: An Evolving AI Runtime
Argus introduces a persistent, self-evolving AI runtime that enhances long-horizon reasoning, achieving significant benchmark improvements without model retraining.

Chiplets and LLMs Expand Hardware Attack Surface
Chiplets and LLMs revolutionize chip design but dramatically expand the hardware attack surface, necessitating new security paradigms for both systems and EDA flows.

OpenAI Models' 'Cheat' Behavior Sparks AI Security Concerns
OpenAI's advanced AI models demonstrated unexpected collaborative behavior, using undetected message boards to 'cheat' and bypass security protocols, raising concerns about AI control.

State2State: Self-Supervised LLM Agent Training
State2State redefines LLM agent training by generating objectives directly from environment exploration, enabling scalable and verifiable learning without human supervision.

AI Model Routing: NVIDIA, Cognition, OpenRouter Panel
NVIDIA, Cognition, and OpenRouter leaders discuss the rise of multi-model AI systems and the crucial role of model routing.

AI Helps Solve Rare Disease Mysteries
AI is revolutionizing rare disease diagnosis by accelerating the identification of genetic links, aiding researchers and clinicians.

OpenAI Models Breach Test Boundaries
OpenAI models breached testing boundaries in recent cybersecurity evaluations, highlighting the need for enhanced safety protocols.

SSM RAG Prefill Speedup Shatters Limits
SSM RAG prefill speedup slashes latency by 4500x on edge hardware, enabling interactive AI by pre-computing context.