#Machine Learning
50 articles with this tag

Gaurav Mishra: RL Agents Need 'Flight School', Not Just Exams
Gaurav Mishra of Amazon AGI Lab discusses the challenges of deploying AI agents trained with reinforcement learning into real-world scenarios, emphasizing the need for 'flight school' training over simple exams.

ScienceFlow: Autonomous Research Gets Serious
ScienceFlow autoresearch agent framework enables sustained LLM research, achieving SOTA results on MLE-bench by managing states and resources adaptively.

OpenAI Previews GPT-5.6 Ultrafast Mode
OpenAI previews GPT-5.6 Sol's 'Ultrafast' mode, demonstrating how up to 14x speed boosts transform AI tasks from investigation to coding.

Trajectory's Arjun Karanam on Closing the AI "Experience Gap"
Trajectory co-founder Arjun Karanam discusses the 'experience gap' in AI models and how his platform aims to enable continual learning by capturing and utilizing real-world user interactions.

LinkedIn's AI Code Review Adapts
LinkedIn's multi-agent AI code review system boosts developer velocity by adapting to codebase specifics and providing actionable feedback.

Chelsea Finn: The State of Physical Intelligence in Robotics
Chelsea Finn discusses the state of physical intelligence in robotics, focusing on achieving long-term autonomy and generality in robot models.

Sara Hooker: AI Frontier Discovery Needs Broader Access
AI researcher Sara Hooker discusses how compute barriers and narrow career paths have limited AI discovery, and how new tools like AutoScientist are democratizing frontier AI development.

Intelligence vs. Expertise in AI Agents
Yu Su of NeoCognition differentiates AI intelligence from expertise, arguing continual learning is key to unlocking specialized skills for agents in complex "micro-worlds."

Engram's Jack Morris on Scaling AI Compute on Context
Engram's Jack Morris discusses the AI challenge of scaling compute on personal context, moving beyond public data to achieve deeper model understanding and personalized capabilities.

AI Agents Need Memory Harnesses for Long Tasks
Stefania Druga of Sakana AI discusses memory harnesses for AI agents, addressing context bloat and the benefits of local models for long-running tasks.

UC Berkeley PhD Student Challenges AI Evaluation Methods
Parth Asawa, a PhD student at UC Berkeley, argues that current AI evaluation methods are insufficient for measuring continual learning and calls for a new benchmark approach.

Firework CEO: Post-Training is Key to Unique AI Business
Firework CEO Lin Qiao discusses the strategic importance of post-training AI models to build unique business value and achieve competitive advantages.

Anthropic's Evolution of AI Agents
Anthropic's Gagan Bhat and Isabella Kai He detail the evolution of AI agents, from Messages API to Managed Agents, focusing on engineering principles, reliability, and security.

AI Agents: The New Primitives of Software
Kwindla Kramer of Daily discusses the historical evolution of computing and the future of AI-native software, drawing parallels from Vannevar Bush to today's AI agents.

Saoud Rizwan: Open Source is Dead, Long Live Open Source
Cline founder Saoud Rizwan argues that AI's impact on open source is profound, but open-weight models offer a cost-effective future, challenging proprietary AI dominance.

Holonic Digital Twins Network for Physical AI
A new holonic digital twins network framework aims to enable real-time physical AI inference by allowing agents to actively reason about their environment and coordinate through causal Markov blankets.

Databricks Unifies Unstructured Data for AI
Databricks introduces FILE type for native handling of unstructured data like images and video in its Lakehouse, enhancing AI development and governance.

Argus: An Evolving AI Runtime
Argus introduces a persistent, self-evolving AI runtime that enhances long-horizon reasoning, achieving significant benchmark improvements without model retraining.

Mariana Minerals: The Future of US Mining with AI
Mariana Minerals, backed by a16z, is set to transform the US mining industry by integrating AI and software to address critical mineral supply chain vulnerabilities.

AI Helps Solve Rare Disease Mysteries
AI is revolutionizing rare disease diagnosis by accelerating the identification of genetic links, aiding researchers and clinicians.

Palo Alto CEO: AI to Patch Cyber Threats in Hours
Palo Alto Networks CEO Nikesh Arora discusses the company's AI-driven approach to cybersecurity, aiming to slash vulnerability patching times from 55 days to hours, and touches on AI token economics and an NBA London bid.

AI Agents Simulate A/B Tests, Cut Costs
AI agents can now simulate A/B tests, drastically reducing costs and time. A new framework decomposes errors, enabling targeted improvements and making AI agent A/B testing simulation a powerful tool.

Chai Discovery: Scaling Drug Design as a Software Problem
Chai Discovery's co-founders discuss their approach to AI-driven drug design, emphasizing simplicity, scaling laws, and the transformation of biology into an engineering discipline.

Waymo CEO on AI's Real-World Challenges
Waymo Co-CEO Dmitri Dolgov shares 7 lessons learned from building and scaling autonomous driving technology, emphasizing the difference between demos and products.

CoreWeave ARIA: AI's New Research Assistant
CoreWeave launches ARIA, an AI agent that automates experiment data analysis to speed up AI model and agent development.

Rayan Garg on Why Long Horizon AI Agents Need Better Verifiers
Rayan Garg from Theta Software explains why long horizon AI agent benchmarks need accurate environment design and final-state verifiers.

Thinking Machines Lab cuts costs with Inkling-Small
Thinking Machines Lab launches Inkling-Small, a 276B parameter model that delivers comparable performance to its larger predecessor at a fraction of the cost.

MiniMax M3: Open Source AI Model Deep Dive
Dan from Together AI and Olive from MiniMax discuss the open-sourcing of the M3 multimodal AI model, its capabilities, and the infrastructure behind scaling AI.

Jeff Dean: AI is a 'compression problem'
Google's Jeff Dean discusses AI's progress, future predictions, and the importance of specialized hardware and context engineering.

Netflix Bets on LLMs for Smarter Recommendations
Netflix's GenRec system uses LLMs to power recommendations, shifting from feature engineering to context engineering for smarter, more efficient content discovery.

AI Agents Stall on Core AI Research
Frontier AI agents can automate AI research engineering but fail to make substantial progress on core research questions, according to new shadow evaluations.

SymmGrid Accelerates Robot Learning
SymmGrid framework dramatically accelerates on-robot learning for manipulation tasks, achieving up to 2.17x speed-ups and moving closer to sub-10 minute training.

NBCUniversal Cuts Costs With Databricks
NBCUniversal slashed data infrastructure costs by 30% and boosted agility by migrating to the Databricks Lakehouse platform, unifying analytics and ML.

Microsoft's Echoverse trains AI agents
Microsoft Research's Echoverse system creates deep, evolving training environments to significantly improve AI agent capabilities in complex software applications.

OpenAI Triples Benchmark Scores With Simple Settings
OpenAI found simple API setting tweaks tripled scores on the ARC-AGI-3 benchmark, proving harness design significantly impacts AI evaluation.

OpenAI GPT-5.6: Smarter, Cheaper
OpenAI's new GPT-5.6 family cuts costs and boosts performance through significant inference and agentic harness optimizations.

Databricks Genie Targets Healthcare Finance
Databricks Genie uses AI and ontology to provide healthcare finance teams with contextualized, governed insights for better margin protection.

Databricks AI: Finance's New Manufacturing Coworker
Databricks launches Genie, an AI coworker for manufacturing finance to identify and free trapped capital using contextual data.

Govt. Benefits Fraud Fights Get Real-Time
Databricks enables federal agencies to implement real-time fraud prevention for government benefits, leveraging AI and cross-agency data sharing.

Databricks AI Agents Take on Production Lines
Databricks introduces AI agents for manufacturing, enabling real-time decision-making on production lines by unifying data and ensuring human oversight.

OpenAI Offers Free Frontier Models to 100,000 Researchers
OpenAI is granting free access to its frontier AI models for 100,000 academic researchers, aiming to accelerate scientific discovery and foster creativity.

OpenAI Opens Frontier AI to 100,000 Researchers
OpenAI offers free access to its frontier AI models, including GPT-5.6, to 100,000 academic researchers to accelerate scientific discovery.

Onton's AI Tackles Agentic Web Trust
Onton has launched Ontology 1, an AI model designed to ensure trustworthy product discovery by combating synthetic content and manipulated recommendations on the agentic web.

AI Pioneers Debate Transformer's Future, Urge New Architectures
AI pioneers Jerry Tworek and Rohan Anil discuss the limitations of current Transformer architectures and the need for new models that can learn from real-world experience.

Dynamic Pretraining Pipelines for LLMs
DataOrchestra revolutionizes LLM pretraining with example-specific data processing, yielding performance gains and reducing compute costs.

Optimizing VLM Pretraining Data Mixes
DecoupleMix offers a systematic, attributable approach to optimizing Vision Language Model pretraining data mixtures, achieving strong performance with greater efficiency.

Databricks AI Search Scales to Production QPS
Databricks AI Search now offers high QPS scaling, allowing applications to move from prototype to production without infrastructure headaches.

Fei-Fei Li: Robots Need 'Spatial Intelligence' for Real-World Tasks
Fei-Fei Li and Yunus from Scenix discuss World Labs' vision for 'spatial intelligence' in AI and robotics, focusing on world models and the real-to-sim-to-real pipeline.

AI's Physical Future: Beyond the Model
The real AI moat lies not in smarter models, but in intelligent engineering systems that can rapidly deploy and iterate on them.

AI Agents Revolutionize Scientific Software
AI agents are transforming scientific computing, speeding up software development and maintenance, but human oversight and long-term stewardship remain crucial.