#Large Language Models

50 articles with this tag

Meta^n: Unlocking Deeper LLM Recursion
AI Research

Meta^n: Unlocking Deeper LLM Recursion

Meta^n introduces a novel recursive LLM agent architecture that overcomes prior meta-depth limitations, achieving state-of-the-art performance across benchmarks, including ARC-AGI-2.

about 1 month ago
AI's Future: Cheaper Tokens, Longer Tasks
Artificial Intelligence

AI's Future: Cheaper Tokens, Longer Tasks

Sal research CEO Neil Movva discusses the future of AI, focusing on cheaper tokens, long-running agents, and the shift from low-latency to proactive intelligence.

about 1 month ago
Software 3.0: The Next AI Paradigm Shift
AI Research

Software 3.0: The Next AI Paradigm Shift

A new paradigm, Software 3.0, is emerging, driven by context and reasoning, converging on databases, large models, and agents.

about 1 month ago
SPADE RL Framework Drives Self-Improvement
AI Research

SPADE RL Framework Drives Self-Improvement

SPADE RL framework empowers LLMs to generate adaptive training environments, driving significant gains in reasoning and tool-use capabilities.

about 1 month ago
Runway Research Unveils Autoregressive-to-Diffusion VLMs for Faster Visual AI
AI Video

Runway Research Unveils Autoregressive-to-Diffusion VLMs for Faster Visual AI

about 1 month ago
Context Overload: The Paradox of LLM Long Windows
AI Research

Context Overload: The Paradox of LLM Long Windows

New research reveals that longer LLM context windows can hinder parametric knowledge, leading to performance degradation and increased context reliance, challenging the 'more is always better' assumption.

about 2 months ago
Engram's Jack Morris on Scaling AI Compute on Context
AI Research

Engram's Jack Morris on Scaling AI Compute on Context

Engram's Jack Morris discusses the AI challenge of scaling compute on personal context, moving beyond public data to achieve deeper model understanding and personalized capabilities.

about 2 months ago
Amazon Hits $3T Valuation, Alibaba's AI Challenger, Apple's Subscription Push
Artificial Intelligence

Amazon Hits $3T Valuation, Alibaba's AI Challenger, Apple's Subscription Push

Amazon tops $3T, Alibaba challenges AI with new model, Apple bets on subscriptions, and Valar Atomics raises $1B for nuclear AI power.

about 2 months ago
AI Synthetic Personas: Promise and Pitfalls
Artificial Intelligence

AI Synthetic Personas: Promise and Pitfalls

Ishan Anand of Insight Sciences explains the rise of AI synthetic personas, their potential, and critical failure modes, drawing parallels to weather forecasting.

2 months ago
Snowflake Cortex AI Adds Claude Opus 5 for Enterprise
Technology

Snowflake Cortex AI Adds Claude Opus 5 for Enterprise

Snowflake Cortex AI now integrates Anthropic's Claude Opus 5, bringing advanced AI reasoning to enterprise data clouds for complex analysis and content generation.

2 months ago
Agentic AI Needs Ontologies for Guardrails, Says UC Berkeley Expert
Artificial Intelligence

Agentic AI Needs Ontologies for Guardrails, Says UC Berkeley Expert

Frank Coyle of UC Berkeley explains how ontologies act as essential guardrails for AI agents, combining probabilistic reasoning with formal logic for safer, more coherent systems.

2 months ago
Simple LLM Merging Surprises
AI Research

Simple LLM Merging Surprises

Direct weighted averaging of LLMs with dimensional adaptation proves surprisingly effective, but ratio control is paramount to avoid capability collapse.

2 months ago
Databricks AI Classify Beats LLMs on Cost
Technology

Databricks AI Classify Beats LLMs on Cost

Databricks unveils an AI Classify workflow combining vector search and AI functions, outperforming LLMs in accuracy and cost for massive document classification tasks.

2 months ago
Pinterest's Medic AI Tames Spark Failures
Artificial Intelligence

Pinterest's Medic AI Tames Spark Failures

Pinterest's Drasko Profirovic details Medic, an AI agent designed to diagnose and fix Apache Spark job failures, highlighting the evolution from prototype to a sophisticated multi-agent system.

2 months ago
Meta's Muse 1.1 Now on Databricks
Technology

Meta's Muse 1.1 Now on Databricks

Meta's Spark Muse 1.1 is now available on Databricks via the Unity AI Gateway, simplifying access and governance for developers.

3 months ago
Kimi K2 Unleashes Open Agentic AI
AI

Kimi K2 Unleashes Open Agentic AI

Moonshot AI open-sources Kimi K2, a powerful LLM optimized for agentic tasks, democratizing advanced AI capabilities for developers and users.

3 months ago
Inkling Model Lands on Databricks
Technology

Inkling Model Lands on Databricks

Thinking Machines Lab's Inkling model is now accessible on Databricks via Unity AI Gateway, enhancing enterprise AI development for coding and agentic tasks.

3 months ago
OpenAI Launches GPT-5.6, Sol Leads Charge
Artificial Intelligence

OpenAI Launches GPT-5.6, Sol Leads Charge

OpenAI launched its GPT-5.6 model family on July 9, 2026, introducing flagship Sol, balanced Terra, and cost-efficient Luna models.

3 months ago
GPT-5.6 Models Launched: Sol, Terra, and Luna Arrive
AI Research

GPT-5.6 Models Launched: Sol, Terra, and Luna Arrive

OpenAI announces the release of its new GPT-5.6 family of models: Sol, Terra, and Luna, already being tested for diverse applications.

3 months ago
OpenAI GPT 5.6 Lands on Snowflake
Technology

OpenAI GPT 5.6 Lands on Snowflake

Snowflake integrates OpenAI's advanced GPT 5.6 models into its Cortex AI platform, enabling enterprises to deploy sophisticated AI applications securely.

3 months ago
Teaching AI Agents to Master Spreadsheets
Artificial Intelligence

Teaching AI Agents to Master Spreadsheets

Nuno Campos of Witan Labs discusses teaching AI agents to master spreadsheets using a REPL approach, improving accuracy and efficiency.

3 months ago
LLM Deception Monitor: Training Data Holds the Key
Artificial Intelligence

LLM Deception Monitor: Training Data Holds the Key

Sachin Kumar explains why LLM deception monitors fail and how analyzing activation 'deltas' from training data is the key to detecting hidden backdoors.

3 months ago
Mixedbread AI on Teaching Agents Better Retrieval
AI Research

Mixedbread AI on Teaching Agents Better Retrieval

Mixedbread AI's Hanna Lichtenberg explains how their new search agent harness bridges the gap between LLM reasoning and effective information retrieval.

3 months ago
LLM Verification: A New Scaling Axis
AI Research

LLM Verification: A New Scaling Axis

LLM-as-a-Verifier redefines LLM scaling by treating verification as a new axis, offering continuous scores for enhanced accuracy and efficiency across agentic tasks.

3 months ago
Agentic LLMs Break Context Limits
AI Research

Agentic LLMs Break Context Limits

CompactionRL integrates context summarization into reinforcement learning for agentic LLMs, breaking context window limits and boosting performance on coding tasks.

3 months ago
Anthropic's Thariq Shihipar on Fable: A Field Guide
AI Research

Anthropic's Thariq Shihipar on Fable: A Field Guide

Anthropic's Thariq Shihipar presents a 'Field Guide to Fable,' detailing its 'grown' nature, the evolution of AI agents, and how to leverage it to discover 'unknown unknowns.'

3 months ago
Memora: Microsoft's AI Memory Upgrade
AI Research

Memora: Microsoft's AI Memory Upgrade

Microsoft's Memora AI memory system revolutionizes long-term AI interactions by balancing detailed recall with efficient retrieval, outperforming existing solutions.

3 months ago
AI Coding Token Reduction: Rajkumar Sakthivel on Local Code Index
Artificial Intelligence

AI Coding Token Reduction: Rajkumar Sakthivel on Local Code Index

Rajkumar Sakthivel from Tesco discusses how a local code index reduced AI coding tokens by 94%, optimizing costs and performance by focusing on context over model improvements.

3 months ago
Erik Hanchett: Cut AI Agent Token Costs
Artificial Intelligence

Erik Hanchett: Cut AI Agent Token Costs

AWS Developer Advocate Erik Hanchett shares five essential strategies to cut AI agent token costs, including caching prompts, routing by difficulty, and managing conversation history.

3 months ago
Isadora Martin-Dye: Layering AI Tone Instructions
Artificial Intelligence

Isadora Martin-Dye: Layering AI Tone Instructions

Isadora Martin-Dye explains why simple tone instructions for AI are insufficient, advocating for a four-layered approach to prompt engineering.

3 months ago
AI Agents Need More Than Just Brains
Technology

AI Agents Need More Than Just Brains

AI agents require more than just powerful LLMs; they need a robust harness infrastructure for reliable real-world task execution.

3 months ago
Uber Eats Tries Agentic Shopping
tech

Uber Eats Tries Agentic Shopping

Uber Eats' Cart Assistant uses AI to translate natural language requests into draft grocery carts, simplifying the shopping process.

4 months ago
Simulating Humans at Scale: Simile's Joon Sung Park
Artificial Intelligence

Simulating Humans at Scale: Simile's Joon Sung Park

Simile's Joon Sung Park discusses simulating human behavior at scale using LLMs to understand societal dynamics and emergent phenomena.

4 months ago
TCS Taps Anthropic's Claude for Regulated Industries
Artificial Intelligence

TCS Taps Anthropic's Claude for Regulated Industries

TCS partners with Anthropic to bring Claude AI to regulated industries like finance and healthcare, integrating it into its own operations and client solutions.

4 months ago
DeepMind's Kilpatrick on AI Models Eating Harnesses
AI Research

DeepMind's Kilpatrick on AI Models Eating Harnesses

Google DeepMind's Logan Kilpatrick delves into the AI concept of models "eating the harness," explaining how over-specialization hinders generalization and what can be done to prevent it.

4 months ago
Steering LRMs Beyond Output Degradation
AI Research

Steering LRMs Beyond Output Degradation

A new probe-based method, FPCG, distinguishes prediction from detection features to enable precise large reasoning models steering with minimal output quality degradation.

4 months ago
Google DeepMind Discusses Open Models & AI Ownership
AI Research

Google DeepMind Discusses Open Models & AI Ownership

Google DeepMind's Gus Martins and Ian Ballantyne discuss the benefits of open AI models like Gemma for ownership, control, and custom applications.

4 months ago
Claude Fable 5 on Databricks: Governed Access for Enterprise Data Teams
Technology

Claude Fable 5 on Databricks: Governed Access for Enterprise Data Teams

Anthropic's Claude Fable 5 is now available on Databricks, offering advanced AI capabilities with enterprise-grade governance and cost controls.

4 months ago
Alex Bowcut on RAG: Accuracy Over Obsolescence
Artificial Intelligence

Alex Bowcut on RAG: Accuracy Over Obsolescence

Alex Bowcut of Sphere discusses why Retrieval Augmented Generation (RAG) remains vital for AI applications demanding accuracy, especially in specialized fields like tax compliance.

4 months ago
Images as the New Reasoning Medium
AI Research

Images as the New Reasoning Medium

This paper introduces optical reasoning, enabling images to serve as the primary medium for LLM and MLLM reasoning, achieving higher token efficiency and competitive performance.

4 months ago
Google's Gemma 4 12B: AI on Your Laptop
AI Research

Google's Gemma 4 12B: AI on Your Laptop

Google's Gemma 4 12B model brings efficient, multimodal AI directly to laptops with a novel unified architecture.

4 months ago
AI Agents Get Dumber With More Context, Expert Warns
Artificial Intelligence

AI Agents Get Dumber With More Context, Expert Warns

Nupur Sharma of Qodo explains how too much context can hinder AI agents, leading to the 'lost in the middle' problem, and discusses solutions like context engines and hybrid orchestration.

4 months ago
NF-CoT: High-Bandwidth Latent Reasoning
AI Research

NF-CoT: High-Bandwidth Latent Reasoning

NF-CoT framework enables high-bandwidth latent reasoning using normalizing flows, boosting LLM performance and efficiency while preserving autoregressive strengths.

4 months ago
Brendon Dillon on Text Diffusion at Google DeepMind
AI Research

Brendon Dillon on Text Diffusion at Google DeepMind

Brendon Dillon from Google DeepMind discusses the advancements and potential of text diffusion models in language generation, highlighting advantages over autoregressive models.

4 months ago
Databricks Search Gets 3x Faster
Technology

Databricks Search Gets 3x Faster

Databricks' Instructed-Retriever-1 model uses parallel test-time scaling to boost Knowledge Assistant search speed by over 3x.

4 months ago
Together AI Masters MiniMax M3 Inference
Technology

Together AI Masters MiniMax M3 Inference

Together AI details engineering feats enabling efficient MiniMax M3 inference, unlocking 1M-token context and multimodality.

4 months ago
Claude Code Sets Opus 4.8 as Default Model for Developers
Technology

Claude Code Sets Opus 4.8 as Default Model for Developers

Claude Code updates its default model to Opus 4.8, introduces dynamic workflows, security plugins, and faster, cheaper Opus options.

4 months ago
Snowflake Adds Claude Opus 4.8
Technology

Snowflake Adds Claude Opus 4.8

Snowflake integrates Anthropic's Claude Opus 4.8 into its Cortex AI platform, boosting agentic workflows, data analysis, and code generation for enterprises.

4 months ago
Anthropic Debuts Claude Opus 4.8
Artificial Intelligence

Anthropic Debuts Claude Opus 4.8

Anthropic unveils Claude Opus 4.8, boosting AI performance with new features like 'effort control' and 'dynamic workflows' for complex coding.

4 months ago
CAG vs. Long Context: AI's Memory Explained
Artificial Intelligence

CAG vs. Long Context: AI's Memory Explained

IBM's Martin Keen explains how AI models use Long Context and Cache Augmented Generation (CAG) to process information, highlighting the trade-offs and efficiency gains of each approach.

4 months ago
#Large Language Models Articles | StartupHub.ai