#Large Language Models
50 articles with this tag

Meta^n: Unlocking Deeper LLM Recursion
Meta^n introduces a novel recursive LLM agent architecture that overcomes prior meta-depth limitations, achieving state-of-the-art performance across benchmarks, including ARC-AGI-2.

AI's Future: Cheaper Tokens, Longer Tasks
Sal research CEO Neil Movva discusses the future of AI, focusing on cheaper tokens, long-running agents, and the shift from low-latency to proactive intelligence.

Software 3.0: The Next AI Paradigm Shift
A new paradigm, Software 3.0, is emerging, driven by context and reasoning, converging on databases, large models, and agents.

SPADE RL Framework Drives Self-Improvement
SPADE RL framework empowers LLMs to generate adaptive training environments, driving significant gains in reasoning and tool-use capabilities.

Runway Research Unveils Autoregressive-to-Diffusion VLMs for Faster Visual AI

Context Overload: The Paradox of LLM Long Windows
New research reveals that longer LLM context windows can hinder parametric knowledge, leading to performance degradation and increased context reliance, challenging the 'more is always better' assumption.

Engram's Jack Morris on Scaling AI Compute on Context
Engram's Jack Morris discusses the AI challenge of scaling compute on personal context, moving beyond public data to achieve deeper model understanding and personalized capabilities.

Amazon Hits $3T Valuation, Alibaba's AI Challenger, Apple's Subscription Push
Amazon tops $3T, Alibaba challenges AI with new model, Apple bets on subscriptions, and Valar Atomics raises $1B for nuclear AI power.

AI Synthetic Personas: Promise and Pitfalls
Ishan Anand of Insight Sciences explains the rise of AI synthetic personas, their potential, and critical failure modes, drawing parallels to weather forecasting.

Snowflake Cortex AI Adds Claude Opus 5 for Enterprise
Snowflake Cortex AI now integrates Anthropic's Claude Opus 5, bringing advanced AI reasoning to enterprise data clouds for complex analysis and content generation.

Agentic AI Needs Ontologies for Guardrails, Says UC Berkeley Expert
Frank Coyle of UC Berkeley explains how ontologies act as essential guardrails for AI agents, combining probabilistic reasoning with formal logic for safer, more coherent systems.

Simple LLM Merging Surprises
Direct weighted averaging of LLMs with dimensional adaptation proves surprisingly effective, but ratio control is paramount to avoid capability collapse.

Databricks AI Classify Beats LLMs on Cost
Databricks unveils an AI Classify workflow combining vector search and AI functions, outperforming LLMs in accuracy and cost for massive document classification tasks.

Pinterest's Medic AI Tames Spark Failures
Pinterest's Drasko Profirovic details Medic, an AI agent designed to diagnose and fix Apache Spark job failures, highlighting the evolution from prototype to a sophisticated multi-agent system.

Meta's Muse 1.1 Now on Databricks
Meta's Spark Muse 1.1 is now available on Databricks via the Unity AI Gateway, simplifying access and governance for developers.

Kimi K2 Unleashes Open Agentic AI
Moonshot AI open-sources Kimi K2, a powerful LLM optimized for agentic tasks, democratizing advanced AI capabilities for developers and users.

Inkling Model Lands on Databricks
Thinking Machines Lab's Inkling model is now accessible on Databricks via Unity AI Gateway, enhancing enterprise AI development for coding and agentic tasks.

OpenAI Launches GPT-5.6, Sol Leads Charge
OpenAI launched its GPT-5.6 model family on July 9, 2026, introducing flagship Sol, balanced Terra, and cost-efficient Luna models.

GPT-5.6 Models Launched: Sol, Terra, and Luna Arrive
OpenAI announces the release of its new GPT-5.6 family of models: Sol, Terra, and Luna, already being tested for diverse applications.

OpenAI GPT 5.6 Lands on Snowflake
Snowflake integrates OpenAI's advanced GPT 5.6 models into its Cortex AI platform, enabling enterprises to deploy sophisticated AI applications securely.

Teaching AI Agents to Master Spreadsheets
Nuno Campos of Witan Labs discusses teaching AI agents to master spreadsheets using a REPL approach, improving accuracy and efficiency.

LLM Deception Monitor: Training Data Holds the Key
Sachin Kumar explains why LLM deception monitors fail and how analyzing activation 'deltas' from training data is the key to detecting hidden backdoors.

Mixedbread AI on Teaching Agents Better Retrieval
Mixedbread AI's Hanna Lichtenberg explains how their new search agent harness bridges the gap between LLM reasoning and effective information retrieval.

LLM Verification: A New Scaling Axis
LLM-as-a-Verifier redefines LLM scaling by treating verification as a new axis, offering continuous scores for enhanced accuracy and efficiency across agentic tasks.

Agentic LLMs Break Context Limits
CompactionRL integrates context summarization into reinforcement learning for agentic LLMs, breaking context window limits and boosting performance on coding tasks.

Anthropic's Thariq Shihipar on Fable: A Field Guide
Anthropic's Thariq Shihipar presents a 'Field Guide to Fable,' detailing its 'grown' nature, the evolution of AI agents, and how to leverage it to discover 'unknown unknowns.'

Memora: Microsoft's AI Memory Upgrade
Microsoft's Memora AI memory system revolutionizes long-term AI interactions by balancing detailed recall with efficient retrieval, outperforming existing solutions.

AI Coding Token Reduction: Rajkumar Sakthivel on Local Code Index
Rajkumar Sakthivel from Tesco discusses how a local code index reduced AI coding tokens by 94%, optimizing costs and performance by focusing on context over model improvements.

Erik Hanchett: Cut AI Agent Token Costs
AWS Developer Advocate Erik Hanchett shares five essential strategies to cut AI agent token costs, including caching prompts, routing by difficulty, and managing conversation history.

Isadora Martin-Dye: Layering AI Tone Instructions
Isadora Martin-Dye explains why simple tone instructions for AI are insufficient, advocating for a four-layered approach to prompt engineering.
AI Agents Need More Than Just Brains
AI agents require more than just powerful LLMs; they need a robust harness infrastructure for reliable real-world task execution.

Uber Eats Tries Agentic Shopping
Uber Eats' Cart Assistant uses AI to translate natural language requests into draft grocery carts, simplifying the shopping process.

Simulating Humans at Scale: Simile's Joon Sung Park
Simile's Joon Sung Park discusses simulating human behavior at scale using LLMs to understand societal dynamics and emergent phenomena.

TCS Taps Anthropic's Claude for Regulated Industries
TCS partners with Anthropic to bring Claude AI to regulated industries like finance and healthcare, integrating it into its own operations and client solutions.

DeepMind's Kilpatrick on AI Models Eating Harnesses
Google DeepMind's Logan Kilpatrick delves into the AI concept of models "eating the harness," explaining how over-specialization hinders generalization and what can be done to prevent it.
Steering LRMs Beyond Output Degradation
A new probe-based method, FPCG, distinguishes prediction from detection features to enable precise large reasoning models steering with minimal output quality degradation.

Google DeepMind Discusses Open Models & AI Ownership
Google DeepMind's Gus Martins and Ian Ballantyne discuss the benefits of open AI models like Gemma for ownership, control, and custom applications.
Claude Fable 5 on Databricks: Governed Access for Enterprise Data Teams
Anthropic's Claude Fable 5 is now available on Databricks, offering advanced AI capabilities with enterprise-grade governance and cost controls.

Alex Bowcut on RAG: Accuracy Over Obsolescence
Alex Bowcut of Sphere discusses why Retrieval Augmented Generation (RAG) remains vital for AI applications demanding accuracy, especially in specialized fields like tax compliance.
Images as the New Reasoning Medium
This paper introduces optical reasoning, enabling images to serve as the primary medium for LLM and MLLM reasoning, achieving higher token efficiency and competitive performance.

Google's Gemma 4 12B: AI on Your Laptop
Google's Gemma 4 12B model brings efficient, multimodal AI directly to laptops with a novel unified architecture.

AI Agents Get Dumber With More Context, Expert Warns
Nupur Sharma of Qodo explains how too much context can hinder AI agents, leading to the 'lost in the middle' problem, and discusses solutions like context engines and hybrid orchestration.
NF-CoT: High-Bandwidth Latent Reasoning
NF-CoT framework enables high-bandwidth latent reasoning using normalizing flows, boosting LLM performance and efficiency while preserving autoregressive strengths.

Brendon Dillon on Text Diffusion at Google DeepMind
Brendon Dillon from Google DeepMind discusses the advancements and potential of text diffusion models in language generation, highlighting advantages over autoregressive models.
Databricks Search Gets 3x Faster
Databricks' Instructed-Retriever-1 model uses parallel test-time scaling to boost Knowledge Assistant search speed by over 3x.

Together AI Masters MiniMax M3 Inference
Together AI details engineering feats enabling efficient MiniMax M3 inference, unlocking 1M-token context and multimodality.
Claude Code Sets Opus 4.8 as Default Model for Developers
Claude Code updates its default model to Opus 4.8, introduces dynamic workflows, security plugins, and faster, cheaper Opus options.

Snowflake Adds Claude Opus 4.8
Snowflake integrates Anthropic's Claude Opus 4.8 into its Cortex AI platform, boosting agentic workflows, data analysis, and code generation for enterprises.

Anthropic Debuts Claude Opus 4.8
Anthropic unveils Claude Opus 4.8, boosting AI performance with new features like 'effort control' and 'dynamic workflows' for complex coding.

CAG vs. Long Context: AI's Memory Explained
IBM's Martin Keen explains how AI models use Long Context and Cache Augmented Generation (CAG) to process information, highlighting the trade-offs and efficiency gains of each approach.