#Large Language Models
50 articles with this tag

Context Overload: The Paradox of LLM Long Windows
New research reveals that longer LLM context windows can hinder parametric knowledge, leading to performance degradation and increased context reliance, challenging the 'more is always better' assumption.

Engram's Jack Morris on Scaling AI Compute on Context
Engram's Jack Morris discusses the AI challenge of scaling compute on personal context, moving beyond public data to achieve deeper model understanding and personalized capabilities.

Amazon Hits $3T Valuation, Alibaba's AI Challenger, Apple's Subscription Push
Amazon tops $3T, Alibaba challenges AI with new model, Apple bets on subscriptions, and Valar Atomics raises $1B for nuclear AI power.

AI Synthetic Personas: Promise and Pitfalls
Ishan Anand of Insight Sciences explains the rise of AI synthetic personas, their potential, and critical failure modes, drawing parallels to weather forecasting.

Snowflake Integrates Claude Opus 5
Snowflake Cortex AI now integrates Anthropic's Claude Opus 5, bringing advanced AI reasoning to enterprise data clouds for complex analysis and content generation.

Agentic AI Needs Ontologies for Guardrails, Says UC Berkeley Expert
Frank Coyle of UC Berkeley explains how ontologies act as essential guardrails for AI agents, combining probabilistic reasoning with formal logic for safer, more coherent systems.

Simple LLM Merging Surprises
Direct weighted averaging of LLMs with dimensional adaptation proves surprisingly effective, but ratio control is paramount to avoid capability collapse.

Databricks AI Classify Beats LLMs on Cost
Databricks unveils an AI Classify workflow combining vector search and AI functions, outperforming LLMs in accuracy and cost for massive document classification tasks.

Pinterest's Medic AI Tames Spark Failures
Pinterest's Drasko Profirovic details Medic, an AI agent designed to diagnose and fix Apache Spark job failures, highlighting the evolution from prototype to a sophisticated multi-agent system.

Meta's Muse 1.1 Now on Databricks
Meta's Spark Muse 1.1 is now available on Databricks via the Unity AI Gateway, simplifying access and governance for developers.

Kimi K2 Unleashes Open Agentic AI
Moonshot AI open-sources Kimi K2, a powerful LLM optimized for agentic tasks, democratizing advanced AI capabilities for developers and users.

Inkling Model Lands on Databricks
Thinking Machines Lab's Inkling model is now accessible on Databricks via Unity AI Gateway, enhancing enterprise AI development for coding and agentic tasks.

OpenAI Launches GPT-5.6, Sol Leads Charge
OpenAI launched its GPT-5.6 model family on July 9, 2026, introducing flagship Sol, balanced Terra, and cost-efficient Luna models.

GPT-5.6 Models Launched: Sol, Terra, and Luna Arrive
OpenAI announces the release of its new GPT-5.6 family of models: Sol, Terra, and Luna, already being tested for diverse applications.

OpenAI GPT 5.6 Lands on Snowflake
Snowflake integrates OpenAI's advanced GPT 5.6 models into its Cortex AI platform, enabling enterprises to deploy sophisticated AI applications securely.

Teaching AI Agents to Master Spreadsheets
Nuno Campos of Witan Labs discusses teaching AI agents to master spreadsheets using a REPL approach, improving accuracy and efficiency.

LLM Deception Monitor: Training Data Holds the Key
Sachin Kumar explains why LLM deception monitors fail and how analyzing activation 'deltas' from training data is the key to detecting hidden backdoors.

Mixedbread AI on Teaching Agents Better Retrieval
Mixedbread AI's Hanna Lichtenberg explains how their new search agent harness bridges the gap between LLM reasoning and effective information retrieval.

LLM Verification: A New Scaling Axis
LLM-as-a-Verifier redefines LLM scaling by treating verification as a new axis, offering continuous scores for enhanced accuracy and efficiency across agentic tasks.

Agentic LLMs Break Context Limits
CompactionRL integrates context summarization into reinforcement learning for agentic LLMs, breaking context window limits and boosting performance on coding tasks.

Anthropic's Thariq Shihipar on Fable: A Field Guide
Anthropic's Thariq Shihipar presents a 'Field Guide to Fable,' detailing its 'grown' nature, the evolution of AI agents, and how to leverage it to discover 'unknown unknowns.'

Memora: Microsoft's AI Memory Upgrade
Microsoft's Memora AI memory system revolutionizes long-term AI interactions by balancing detailed recall with efficient retrieval, outperforming existing solutions.

AI Coding Token Reduction: Rajkumar Sakthivel on Local Code Index
Rajkumar Sakthivel from Tesco discusses how a local code index reduced AI coding tokens by 94%, optimizing costs and performance by focusing on context over model improvements.

Erik Hanchett: Cut AI Agent Token Costs
AWS Developer Advocate Erik Hanchett shares five essential strategies to cut AI agent token costs, including caching prompts, routing by difficulty, and managing conversation history.

Isadora Martin-Dye: Layering AI Tone Instructions
Isadora Martin-Dye explains why simple tone instructions for AI are insufficient, advocating for a four-layered approach to prompt engineering.
AI Agents Need More Than Just Brains
AI agents require more than just powerful LLMs; they need a robust harness infrastructure for reliable real-world task execution.

Uber Eats Tries Agentic Shopping
Uber Eats' Cart Assistant uses AI to translate natural language requests into draft grocery carts, simplifying the shopping process.

Simulating Humans at Scale: Simile's Joon Sung Park
Simile's Joon Sung Park discusses simulating human behavior at scale using LLMs to understand societal dynamics and emergent phenomena.

TCS Taps Anthropic's Claude for Regulated Industries
TCS partners with Anthropic to bring Claude AI to regulated industries like finance and healthcare, integrating it into its own operations and client solutions.

DeepMind's Kilpatrick on AI Models Eating Harnesses
Google DeepMind's Logan Kilpatrick delves into the AI concept of models "eating the harness," explaining how over-specialization hinders generalization and what can be done to prevent it.
Steering LRMs Beyond Output Degradation
A new probe-based method, FPCG, distinguishes prediction from detection features to enable precise large reasoning models steering with minimal output quality degradation.

Google DeepMind Discusses Open Models & AI Ownership
Google DeepMind's Gus Martins and Ian Ballantyne discuss the benefits of open AI models like Gemma for ownership, control, and custom applications.
Claude on Databricks: Governed AI Access for Enterprise Data Teams
Anthropic's Claude Fable 5 is now available on Databricks, offering advanced AI capabilities with enterprise-grade governance and cost controls.

Alex Bowcut on RAG: Accuracy Over Obsolescence
Alex Bowcut of Sphere discusses why Retrieval Augmented Generation (RAG) remains vital for AI applications demanding accuracy, especially in specialized fields like tax compliance.
Images as the New Reasoning Medium
This paper introduces optical reasoning, enabling images to serve as the primary medium for LLM and MLLM reasoning, achieving higher token efficiency and competitive performance.

Google's Gemma 4 12B: AI on Your Laptop
Google's Gemma 4 12B model brings efficient, multimodal AI directly to laptops with a novel unified architecture.

AI Agents Get Dumber With More Context, Expert Warns
Nupur Sharma of Qodo explains how too much context can hinder AI agents, leading to the 'lost in the middle' problem, and discusses solutions like context engines and hybrid orchestration.
NF-CoT: High-Bandwidth Latent Reasoning
NF-CoT framework enables high-bandwidth latent reasoning using normalizing flows, boosting LLM performance and efficiency while preserving autoregressive strengths.

Brendon Dillon on Text Diffusion at Google DeepMind
Brendon Dillon from Google DeepMind discusses the advancements and potential of text diffusion models in language generation, highlighting advantages over autoregressive models.
Databricks Search Gets 3x Faster
Databricks' Instructed-Retriever-1 model uses parallel test-time scaling to boost Knowledge Assistant search speed by over 3x.

Together AI Masters MiniMax M3 Inference
Together AI details engineering feats enabling efficient MiniMax M3 inference, unlocking 1M-token context and multimodality.
Claude Code Embraces Opus 4.8
Claude Code updates its default model to Opus 4.8, introduces dynamic workflows, security plugins, and faster, cheaper Opus options.

Snowflake Adds Claude Opus 4.8
Snowflake integrates Anthropic's Claude Opus 4.8 into its Cortex AI platform, boosting agentic workflows, data analysis, and code generation for enterprises.

Anthropic Debuts Claude Opus 4.8
Anthropic unveils Claude Opus 4.8, boosting AI performance with new features like 'effort control' and 'dynamic workflows' for complex coding.

CAG vs. Long Context: AI's Memory Explained
IBM's Martin Keen explains how AI models use Long Context and Cache Augmented Generation (CAG) to process information, highlighting the trade-offs and efficiency gains of each approach.

Angus McLean on Bounded Autonomy in AI
Angus J. McLean of Oliver discusses 'Bounded Autonomy' in AI, exploring the shift to agentic processes in advertising and offering practical advice for building AI agents.

Unlocking LLM Recall: Data Composition is Key
New research reveals a sigmoid scaling law for LLM factual recall, driven by model size and training data composition, explaining up to 94% of performance variance.

Google Launches Gemini 3.5 Flash
Google unveils Gemini 3.5 Flash, a fast and intelligent AI model optimized for agentic tasks, now powering consumer and developer tools.

Spotify's Shivam Verma on LLMs and Personalization
Shivam Verma from Spotify discusses how LLMs are transforming personalization in recommendation systems, moving towards steerable and context-aware content discovery.

AI Delegation: Reliability Concerns Emerge
New Microsoft Research highlights how AI can degrade document fidelity in long, delegated tasks, stressing the need for better verification and orchestration.