#Large Language Models

50 articles with this tag

Context Overload: The Paradox of LLM Long Windows
AI Research

Context Overload: The Paradox of LLM Long Windows

New research reveals that longer LLM context windows can hinder parametric knowledge, leading to performance degradation and increased context reliance, challenging the 'more is always better' assumption.

3 days ago
Engram's Jack Morris on Scaling AI Compute on Context
AI Research

Engram's Jack Morris on Scaling AI Compute on Context

Engram's Jack Morris discusses the AI challenge of scaling compute on personal context, moving beyond public data to achieve deeper model understanding and personalized capabilities.

4 days ago
Amazon Hits $3T Valuation, Alibaba's AI Challenger, Apple's Subscription Push
Artificial Intelligence

Amazon Hits $3T Valuation, Alibaba's AI Challenger, Apple's Subscription Push

Amazon tops $3T, Alibaba challenges AI with new model, Apple bets on subscriptions, and Valar Atomics raises $1B for nuclear AI power.

13 days ago
AI Synthetic Personas: Promise and Pitfalls
Artificial Intelligence

AI Synthetic Personas: Promise and Pitfalls

Ishan Anand of Insight Sciences explains the rise of AI synthetic personas, their potential, and critical failure modes, drawing parallels to weather forecasting.

18 days ago
Snowflake Integrates Claude Opus 5
Technology

Snowflake Integrates Claude Opus 5

Snowflake Cortex AI now integrates Anthropic's Claude Opus 5, bringing advanced AI reasoning to enterprise data clouds for complex analysis and content generation.

23 days ago
Agentic AI Needs Ontologies for Guardrails, Says UC Berkeley Expert
Artificial Intelligence

Agentic AI Needs Ontologies for Guardrails, Says UC Berkeley Expert

Frank Coyle of UC Berkeley explains how ontologies act as essential guardrails for AI agents, combining probabilistic reasoning with formal logic for safer, more coherent systems.

25 days ago
Simple LLM Merging Surprises
AI Research

Simple LLM Merging Surprises

Direct weighted averaging of LLMs with dimensional adaptation proves surprisingly effective, but ratio control is paramount to avoid capability collapse.

26 days ago
Databricks AI Classify Beats LLMs on Cost
Technology

Databricks AI Classify Beats LLMs on Cost

Databricks unveils an AI Classify workflow combining vector search and AI functions, outperforming LLMs in accuracy and cost for massive document classification tasks.

27 days ago
Pinterest's Medic AI Tames Spark Failures
Artificial Intelligence

Pinterest's Medic AI Tames Spark Failures

Pinterest's Drasko Profirovic details Medic, an AI agent designed to diagnose and fix Apache Spark job failures, highlighting the evolution from prototype to a sophisticated multi-agent system.

27 days ago
Meta's Muse 1.1 Now on Databricks
Technology

Meta's Muse 1.1 Now on Databricks

Meta's Spark Muse 1.1 is now available on Databricks via the Unity AI Gateway, simplifying access and governance for developers.

30 days ago
Kimi K2 Unleashes Open Agentic AI
AI

Kimi K2 Unleashes Open Agentic AI

Moonshot AI open-sources Kimi K2, a powerful LLM optimized for agentic tasks, democratizing advanced AI capabilities for developers and users.

about 1 month ago
Inkling Model Lands on Databricks
Technology

Inkling Model Lands on Databricks

Thinking Machines Lab's Inkling model is now accessible on Databricks via Unity AI Gateway, enhancing enterprise AI development for coding and agentic tasks.

about 1 month ago
OpenAI Launches GPT-5.6, Sol Leads Charge
Artificial Intelligence

OpenAI Launches GPT-5.6, Sol Leads Charge

OpenAI launched its GPT-5.6 model family on July 9, 2026, introducing flagship Sol, balanced Terra, and cost-efficient Luna models.

about 1 month ago
GPT-5.6 Models Launched: Sol, Terra, and Luna Arrive
AI Research

GPT-5.6 Models Launched: Sol, Terra, and Luna Arrive

OpenAI announces the release of its new GPT-5.6 family of models: Sol, Terra, and Luna, already being tested for diverse applications.

about 1 month ago
OpenAI GPT 5.6 Lands on Snowflake
Technology

OpenAI GPT 5.6 Lands on Snowflake

Snowflake integrates OpenAI's advanced GPT 5.6 models into its Cortex AI platform, enabling enterprises to deploy sophisticated AI applications securely.

about 1 month ago
Teaching AI Agents to Master Spreadsheets
Artificial Intelligence

Teaching AI Agents to Master Spreadsheets

Nuno Campos of Witan Labs discusses teaching AI agents to master spreadsheets using a REPL approach, improving accuracy and efficiency.

about 1 month ago
LLM Deception Monitor: Training Data Holds the Key
Artificial Intelligence

LLM Deception Monitor: Training Data Holds the Key

Sachin Kumar explains why LLM deception monitors fail and how analyzing activation 'deltas' from training data is the key to detecting hidden backdoors.

about 1 month ago
Mixedbread AI on Teaching Agents Better Retrieval
AI Research

Mixedbread AI on Teaching Agents Better Retrieval

Mixedbread AI's Hanna Lichtenberg explains how their new search agent harness bridges the gap between LLM reasoning and effective information retrieval.

about 1 month ago
LLM Verification: A New Scaling Axis
AI Research

LLM Verification: A New Scaling Axis

LLM-as-a-Verifier redefines LLM scaling by treating verification as a new axis, offering continuous scores for enhanced accuracy and efficiency across agentic tasks.

about 1 month ago
Agentic LLMs Break Context Limits
AI Research

Agentic LLMs Break Context Limits

CompactionRL integrates context summarization into reinforcement learning for agentic LLMs, breaking context window limits and boosting performance on coding tasks.

about 1 month ago
Anthropic's Thariq Shihipar on Fable: A Field Guide
AI Research

Anthropic's Thariq Shihipar on Fable: A Field Guide

Anthropic's Thariq Shihipar presents a 'Field Guide to Fable,' detailing its 'grown' nature, the evolution of AI agents, and how to leverage it to discover 'unknown unknowns.'

about 1 month ago
Memora: Microsoft's AI Memory Upgrade
AI Research

Memora: Microsoft's AI Memory Upgrade

Microsoft's Memora AI memory system revolutionizes long-term AI interactions by balancing detailed recall with efficient retrieval, outperforming existing solutions.

about 2 months ago
AI Coding Token Reduction: Rajkumar Sakthivel on Local Code Index
Artificial Intelligence

AI Coding Token Reduction: Rajkumar Sakthivel on Local Code Index

Rajkumar Sakthivel from Tesco discusses how a local code index reduced AI coding tokens by 94%, optimizing costs and performance by focusing on context over model improvements.

about 2 months ago
Erik Hanchett: Cut AI Agent Token Costs
Artificial Intelligence

Erik Hanchett: Cut AI Agent Token Costs

AWS Developer Advocate Erik Hanchett shares five essential strategies to cut AI agent token costs, including caching prompts, routing by difficulty, and managing conversation history.

about 2 months ago
Isadora Martin-Dye: Layering AI Tone Instructions
Artificial Intelligence

Isadora Martin-Dye: Layering AI Tone Instructions

Isadora Martin-Dye explains why simple tone instructions for AI are insufficient, advocating for a four-layered approach to prompt engineering.

about 2 months ago
AI Agents Need More Than Just Brains
Technology

AI Agents Need More Than Just Brains

AI agents require more than just powerful LLMs; they need a robust harness infrastructure for reliable real-world task execution.

about 2 months ago
Uber Eats Tries Agentic Shopping
tech

Uber Eats Tries Agentic Shopping

Uber Eats' Cart Assistant uses AI to translate natural language requests into draft grocery carts, simplifying the shopping process.

2 months ago
Simulating Humans at Scale: Simile's Joon Sung Park
Artificial Intelligence

Simulating Humans at Scale: Simile's Joon Sung Park

Simile's Joon Sung Park discusses simulating human behavior at scale using LLMs to understand societal dynamics and emergent phenomena.

2 months ago
TCS Taps Anthropic's Claude for Regulated Industries
Artificial Intelligence

TCS Taps Anthropic's Claude for Regulated Industries

TCS partners with Anthropic to bring Claude AI to regulated industries like finance and healthcare, integrating it into its own operations and client solutions.

2 months ago
DeepMind's Kilpatrick on AI Models Eating Harnesses
AI Research

DeepMind's Kilpatrick on AI Models Eating Harnesses

Google DeepMind's Logan Kilpatrick delves into the AI concept of models "eating the harness," explaining how over-specialization hinders generalization and what can be done to prevent it.

2 months ago
Steering LRMs Beyond Output Degradation
AI Research

Steering LRMs Beyond Output Degradation

A new probe-based method, FPCG, distinguishes prediction from detection features to enable precise large reasoning models steering with minimal output quality degradation.

2 months ago
Google DeepMind Discusses Open Models & AI Ownership
AI Research

Google DeepMind Discusses Open Models & AI Ownership

Google DeepMind's Gus Martins and Ian Ballantyne discuss the benefits of open AI models like Gemma for ownership, control, and custom applications.

2 months ago
Claude on Databricks: Governed AI Access for Enterprise Data Teams
Technology

Claude on Databricks: Governed AI Access for Enterprise Data Teams

Anthropic's Claude Fable 5 is now available on Databricks, offering advanced AI capabilities with enterprise-grade governance and cost controls.

2 months ago
Alex Bowcut on RAG: Accuracy Over Obsolescence
Artificial Intelligence

Alex Bowcut on RAG: Accuracy Over Obsolescence

Alex Bowcut of Sphere discusses why Retrieval Augmented Generation (RAG) remains vital for AI applications demanding accuracy, especially in specialized fields like tax compliance.

2 months ago
Images as the New Reasoning Medium
AI Research

Images as the New Reasoning Medium

This paper introduces optical reasoning, enabling images to serve as the primary medium for LLM and MLLM reasoning, achieving higher token efficiency and competitive performance.

2 months ago
Google's Gemma 4 12B: AI on Your Laptop
AI Research

Google's Gemma 4 12B: AI on Your Laptop

Google's Gemma 4 12B model brings efficient, multimodal AI directly to laptops with a novel unified architecture.

2 months ago
AI Agents Get Dumber With More Context, Expert Warns
Artificial Intelligence

AI Agents Get Dumber With More Context, Expert Warns

Nupur Sharma of Qodo explains how too much context can hinder AI agents, leading to the 'lost in the middle' problem, and discusses solutions like context engines and hybrid orchestration.

2 months ago
NF-CoT: High-Bandwidth Latent Reasoning
AI Research

NF-CoT: High-Bandwidth Latent Reasoning

NF-CoT framework enables high-bandwidth latent reasoning using normalizing flows, boosting LLM performance and efficiency while preserving autoregressive strengths.

2 months ago
Brendon Dillon on Text Diffusion at Google DeepMind
AI Research

Brendon Dillon on Text Diffusion at Google DeepMind

Brendon Dillon from Google DeepMind discusses the advancements and potential of text diffusion models in language generation, highlighting advantages over autoregressive models.

2 months ago
Databricks Search Gets 3x Faster
Technology

Databricks Search Gets 3x Faster

Databricks' Instructed-Retriever-1 model uses parallel test-time scaling to boost Knowledge Assistant search speed by over 3x.

2 months ago
Together AI Masters MiniMax M3 Inference
Technology

Together AI Masters MiniMax M3 Inference

Together AI details engineering feats enabling efficient MiniMax M3 inference, unlocking 1M-token context and multimodality.

2 months ago
Claude Code Embraces Opus 4.8
Technology

Claude Code Embraces Opus 4.8

Claude Code updates its default model to Opus 4.8, introduces dynamic workflows, security plugins, and faster, cheaper Opus options.

3 months ago
Snowflake Adds Claude Opus 4.8
Technology

Snowflake Adds Claude Opus 4.8

Snowflake integrates Anthropic's Claude Opus 4.8 into its Cortex AI platform, boosting agentic workflows, data analysis, and code generation for enterprises.

3 months ago
Anthropic Debuts Claude Opus 4.8
Artificial Intelligence

Anthropic Debuts Claude Opus 4.8

Anthropic unveils Claude Opus 4.8, boosting AI performance with new features like 'effort control' and 'dynamic workflows' for complex coding.

3 months ago
CAG vs. Long Context: AI's Memory Explained
Artificial Intelligence

CAG vs. Long Context: AI's Memory Explained

IBM's Martin Keen explains how AI models use Long Context and Cache Augmented Generation (CAG) to process information, highlighting the trade-offs and efficiency gains of each approach.

3 months ago
Angus McLean on Bounded Autonomy in AI
Artificial Intelligence

Angus McLean on Bounded Autonomy in AI

Angus J. McLean of Oliver discusses 'Bounded Autonomy' in AI, exploring the shift to agentic processes in advertising and offering practical advice for building AI agents.

3 months ago
Unlocking LLM Recall: Data Composition is Key
AI Research

Unlocking LLM Recall: Data Composition is Key

New research reveals a sigmoid scaling law for LLM factual recall, driven by model size and training data composition, explaining up to 94% of performance variance.

3 months ago
Google Launches Gemini 3.5 Flash
Artificial Intelligence

Google Launches Gemini 3.5 Flash

Google unveils Gemini 3.5 Flash, a fast and intelligent AI model optimized for agentic tasks, now powering consumer and developer tools.

3 months ago
Spotify's Shivam Verma on LLMs and Personalization
Artificial Intelligence

Spotify's Shivam Verma on LLMs and Personalization

Shivam Verma from Spotify discusses how LLMs are transforming personalization in recommendation systems, moving towards steerable and context-aware content discovery.

3 months ago
AI Delegation: Reliability Concerns Emerge
AI Research

AI Delegation: Reliability Concerns Emerge

New Microsoft Research highlights how AI can degrade document fidelity in long, delegated tasks, stressing the need for better verification and orchestration.

3 months ago