#LLM
50 articles with this tag

AI Inference: 10x Faster Models & Self-Optimization
Philip Kiely and Ali Taha of Baseten discuss AI inference, LLM optimization, speculative decoding, and the engineering behind cutting-edge AI models.

Crusoe Cloud's fastokens v2 boosts AI speed
Crusoe Cloud's fastokens v2 offers native tiktoken support and major speed boosts for AI model inference and training.

AI Chatbot Race: Who's Really Winning?
New data reveals ChatGPT leads in reach and retention, but Claude AI is exploding in growth, challenging established dominance.

Netflix Bets on LLMs for Smarter Recommendations
Netflix's GenRec system uses LLMs to power recommendations, shifting from feature engineering to context engineering for smarter, more efficient content discovery.

Together AI partners with Moonshot AI
Together AI partners with Moonshot AI to offer Kimi K3 and future models, providing developers with day-zero access to large-scale open-source AI.

Ditch LLM Chasing, Build Once
Otari's unified gateway simplifies LLM integration, allowing teams to access multiple models without rebuilding infrastructure for each new provider.

Uber Eats Uses AI Agents to Enhance Food Photos at Scale
Uber Eats' computer vision team details their AI agent system for enhancing food photos, focusing on closed-loop feedback and continuous learning.

Google Experts Share AI Agent Evaluation Best Practices
Google's Preetika Bhateja & Daniel Bump share essential strategies for building effective AI agent evaluation systems, from initial 'vibing' to scaling with LLM judges.

OpenAI Explains Value Maximization with GPT-5.6
OpenAI's Build Hour explains 'value maxing' with GPT-5.6, detailing efficiency tips and Ploy's migration strategy.

Anthropic Launches Claude Opus 5
Anthropic releases Claude Opus 5, delivering near-frontier AI intelligence at a reduced cost with state-of-the-art performance in coding and knowledge work.

Kimi K3 Challenges Claude Fable 5 on Code Quality, Slashes Cost
Kimi K3 challenges Claude Fable 5 on coding benchmarks, offering similar quality at a third of the cost and the benefits of an open-weight model.

OpenCode CEO on 20x Growth & AI Agent Market
OpenCode CEO Jay V reveals how the platform achieved 20x growth, reaching 4.6M users by supporting any AI model and becoming a key player in the global coding agent market.

OpenWorker AI: Your Desktop Co-Worker
Andrew Ng and Rohit Prasad launch OpenWorker AI, an open-source desktop agent delivering finished work and prioritizing data privacy.

AI Models vs. Hackers: The Cybersecurity Arms Race
Thomas Wolf (Hugging Face) and Uri Rolls (Arithmetic) discuss training AI models to out-think cyber attackers using novel benchmarks and the potential of open-source AI.

Notion's AI Lead on Token Costs & Strategy
Notion's Sarah Sachs discusses the economics of AI tokens, the importance of product strategy over model choice, and navigating the 'wild west' of AI development.

Blackstone Sees AI Pivot Paying Off
Blackstone's CEO discusses the company's successful AI pivot, strong earnings, and future growth strategies, while addressing market concerns and the firm's diverse investments.

LLM Provenance: Tracking Data Origins with Graffiti
Daniel Chalef of Zep AI discusses the critical challenge of provenance in LLM-generated data and how the Graffiti framework addresses it through temporal graph modeling.

TwelveLabs Builds Video Memory Layer
James Le of TwelveLabs discusses the critical need for memory in video AI, detailing the limitations of current systems and introducing their 'Jockey' platform as a solution.

Agentic AI Needs Ontologies for Guardrails, Says UC Berkeley Expert
Frank Coyle of UC Berkeley explains how ontologies act as essential guardrails for AI agents, combining probabilistic reasoning with formal logic for safer, more coherent systems.

AI Assistants Need Graph Memory, Not Just More Tokens
Stephen Chin of Neo4j discusses how graph databases offer superior memory solutions for AI assistants compared to traditional file storage or vector databases.

Copilot vs. API: Where your AI dollars go
GitHub Copilot offers a workflow-integrated AI experience, distinct from raw API access which is suited for building custom systems. The choice depends on the scope of 'work you need to own'.

AI Guardrails Need Smarter Tools
AI guardrails require the same scrutiny as models, with agentic tools like web search proving vital for reliable deployment.

AI Agents Evolve: From Harnesses to Autonomous Claws
Mastra CEO Sam Bhagwat discusses the evolution of AI agents from LLMs to autonomous 'Claws,' the shift to cloud-based systems, and the inevitable market shakeout.

AI Agents as Supply Chain Actors: Patch Pilot's Security Model
Moritz Johner of Form3 discusses the limitations of automated dependency patching tools and the security considerations of using AI agents with production code access, introducing their 'Patch Pilot' system.

Nvidia Dev: ML Security Flaws Are 'Boring' Mistakes
Nvidia's Lavina D'Mello argues that ML security failures stem from 'boring' infrastructure misconfigurations, not exotic AI attacks, urging a return to foundational security practices.

Anthropic's Claude Code Creator on AI's Future
Boris Cherny, Head of Claude Code at Anthropic, discusses the evolution of AI models, the future of software engineering, and Anthropic's focus on AI safety.

Diane Lin on AI Agent Inconsistency
Diane Lin from Datadog explains how AI agents' inconsistent outputs often stem from ambiguous data and how to improve them using memory augmentation techniques.

Ace: Why Voice AI Needs Smarter Scaffolding, Not Bigger Models
Ace creators Ornella Bahidika and Joel Allou explain why smaller AI models coupled with smart 'scaffolding' are superior for responsive voice applications.

Bala Ramdoss on Generative UI for Agentic CX
Bala Ramdoss of Amazon discusses generative UI, the critical layer between LLM output and product experience, emphasizing rendering contracts, streaming, and BFF patterns for agentic CX.

AWS Experts Detail Turn-Taking in Voice Agents
AWS experts Chintan Agrawal and Daniel Wirjo discuss turn-taking in voice agents, covering the 200ms human constraint, pipeline components, and three levels of solutions.

Skills Are the New SDKs: Rethinking AI Agents
Elvin Aghammadzada of DataRobot argues that 'skills' are the new SDKs for AI agents, addressing context engineering challenges and the shift from 'friction' to 'fluency' moats.

AI Automates Oncology Workflows, Minimizing Human Touch
Anant Shankar from Trisca discusses how AI agents are automating oncology workflows, from eligibility checks to submission, aiming for 'no-touch' processing of prior authorizations.

Langfuse: Domain Expertise Crucial for AI Self-Improvement
Langfuse's Annabelle Schäfer explains why domain expertise is crucial for AI self-improvement, advocating for high-signal target functions and expert-driven data.

Netflix's LLM Engine Revealed
Netflix reveals its custom LLM serving infrastructure built on vLLM and NVIDIA Triton, enabling flexibility and performance within its production environment.

AI Stocks: Fundamentals vs. Fear, Says Mandeep Singh
Mandeep Singh of Bloomberg Intelligence discusses AI market volatility, the rise of open-weight models, and the impact of Alphabet's Gemini launch delay.

AI Race: China's Moonshot and the Shifting Tech Landscape
Bloomberg Intelligence discusses Moonshot's AI model, Asian markets' influence, Apple's AI strategy, and the evolving LLM landscape.

Databricks Builds AI Soccer Coach
Databricks' new 'La Pizarra' app turns massive soccer match data into real-time tactical insights, leveraging AI for scouting and opponent analysis.

Kimi K2.6 Open Sources Advanced Coding AI
Moonshot AI open-sources Kimi K2.6, a powerful AI model for coding and agentic workflows, boasting state-of-the-art long-horizon execution and agent swarm capabilities.

Databricks Tackles Agentic AI Skills Gap
Databricks launches industry-first Context Engineer certification and AI-assisted training to address the growing agentic AI skills gap.

Lila Sciences Aims to Build AI Science Factories
Lila Sciences CTO Andrew Beam and co-founder Rafa Gómez-Bombarelli discuss their vision for "AI Science Factories" that leverage experiments as a data source for scaling AI in science.

Databricks Scales Higher Ed Support with AI
Databricks leverages GenAI to transform higher education advisory services, improving call quality monitoring and student support through scalable AI solutions.

Anthropic's Cat Wu & Thariq Shihipar on AI in Software Dev
Anthropic's Cat Wu and Thariq Shihipar discuss the evolution of Claude Code and the impact of Claude Tag on software development, emphasizing increased efficiency and collaboration.

Open AI Models Ready, APIs Lag
Open source AI models are ready for agents, but their surrounding platforms lag behind frontier APIs, creating critical infrastructure gaps.

Anthropic Platform: Building Blocks for AI
Anthropic's Katelyn Lesse and Angela Jiang discuss building a flexible AI platform with layers of knowledge, execution, and coordination.

AI in Law: More Work or More Efficiency?
Gary Wingens of Lowenstein Sandler discusses AI's impact on law, from efficiency gains to the evolving role of junior lawyers, and the uncertain economics of AI pricing.

Mozilla.ai Launches Otari LLM Control Plane
Mozilla.ai has launched Otari, an open-source LLM control plane to simplify managing diverse language model infrastructure for AI application developers.

TNG AI Chess: AI Explains Chess Like a Human Trainer
Stephan Steinfurt from TNG Technology Consulting discusses their AI agent that explains chess games like a human trainer, combining LLMs with specialized tools to create automated video analyses.

Flint: AI Agents Get Smarter Charts
Microsoft Research's Flint visualization language empowers AI agents to generate polished, expressive charts from simple specifications, bypassing complex low-level details.

AI's Next Frontier: Cost, Control, and Compute Race
Perplexity's Aravind Srinivas discusses the evolving AI race, emphasizing cost, control, and compute as key drivers for enterprise adoption of open-source models and orchestration.

Meta Engineers on AI Game Development
Meta's Danielle An and David Hoe discuss the complexities of AI in game development, emphasizing 'taste' and dynamic gameplay.