#LLM
50 articles with this tag

LLMs Fail to Write Fast Multi-GPU Kernels
Simran Arora from Together AI discusses the challenges of multi-GPU kernel development and why current LLMs struggle to optimize them, despite ongoing research.

Tor Network Abused for IoT Device Exploits
Researchers unveil 'Torchlight' system exposing how attackers exploit Tor anonymity to target millions of IoT devices with zero-day vulnerabilities.

LinkedIn's AI Powers Smarter Follows
LinkedIn leverages LLMs to build a new recommendation engine, matching users with creators based on deep semantic understanding rather than just popularity.

AI Reasoning: Fine-Tuning's Hidden Cost
Fine-tuning AI reasoning models on business data can erase their thinking process; new methods aim to preserve it.

LLM Evaluation: Beyond Benchmarks
GitHub shares critical lessons on evaluating LLMs for production, emphasizing product decisions and rigorous testing over benchmarks.

Oxylabs: The Missing Layer in Agentic AI
Giedrius Å teimantas of Oxylabs reveals the 'missing layer' in agentic AI: the inefficient handling of product page data, leading to massive token waste.

LLM Self-Reflection Drives Data Efficiency
SRPO framework enables LLMs to self-reflect on errors, generating dense training signals that drastically improve data efficiency and achieve SOTA on reasoning and agentic benchmarks.

ChatGPT Text Chat Limits? 2 Unlimited Says 'No Limits'
A music video uses 2 Unlimited's 'No Limit' to declare that text chats with AI like ChatGPT have 'no limits.'

LLM Query Types Drive GPU Demand
LLM query types dramatically affect GPU demand, memory, and power, highlighting the need for optimized AI infrastructure planning.

Legora's YC Rejection to $100M ARR: A Startup Story
Legora's co-founder & CEO, Max Junestrand, shares the startup's incredible journey from YC rejection to $100M ARR, fueled by AI and a winning culture.

AI Agents Planned Hacking Spree, OpenAI Reveals
AI safety expert Connor Leahy reveals how OpenAI's AI agents collaborated on hacking attempts and discusses the growing unpredictability of AI.

Parag Agrawal on AI Search & the Future Web
Parag Agrawal discusses Parallel's mission to build a new web infrastructure optimized for AI agents, moving beyond human-centric search.

Databricks AI Tackles Incident Nightmares
Databricks deploys AI SRE to automate incident investigation, speeding up root cause analysis and enhancing transparency across its microservices.

DigitalOcean: Model Routing Beats Benchmarks
DigitalOcean's Archana Kamath and Tyler Gillam discuss model routing, arguing that preferences like cost and latency should dictate LLM choices over benchmarks, and showcase their open-source inference router.

Microsoft Unveils TokenOps for AI Agent Cost Control
Microsoft introduces TokenOps, a run-aware governance system for AI agents to control token spending and move from 'token maxing' to 'value maxing'.

Hugging Face Engineer Automates Job with AI Agents
Niels Rogge from Hugging Face shares how he uses AI agents to automate his job, from outreach to researchers to improving model discoverability on the Hugging Face Hub.

AI Transforms Keyword Research
AI is revolutionizing keyword research, offering speed, coverage, and structure beyond traditional methods. Discover how tools are evolving.

AI Search Reshapes Competitor Analysis
AI is reshaping how businesses understand their competition, moving beyond search rankings to generative AI visibility and audience behavior.

Open Source AI: Beyond Virtue to Ownership
Open source AI infrastructure is an ownership strategy, not charity. Controlling the AI model orchestration layer is key for long-term stability and auditability.

Anterior's Anuj Iravane on Synthetic Healthcare Data
Anuj Iravane of Anterior discusses how the company overcomes PHI challenges in healthcare AI by generating synthetic data, reversing inference workflows, and empowering clinicians.

Ufonia's AI: Shipping Healthcare Safely to a Million Patients
Jared Joselowitz of Ufonia explains how to safely deploy healthcare AI to millions of patients using simulation, automated prompt optimization, and rigorous evaluation, bypassing traditional A/B testing.

Dynamic Memory Activation Enhances LLMs
A new mechanism, Proteus, introduces incremental memory activation to LLMs, enhancing performance on long contexts without added computational cost.

Reactor's Ahmed Ahres on Real-Time Interactive Video
Ahmed Ahres of Reactor discusses the transformative potential of real-time interactive video, moving beyond static generative models to dynamic, programmable content.

Palmyra x6: Enterprise LLMs Get Focused
Palmyra x6 LLM redefines enterprise agentic tasks with a conservative, high-performance fine-tuning approach, excelling in benchmarks and safety.

Retail AI Needs a Control Plane
Retailers are moving beyond AI experimentation to enterprise-wide adoption, demanding a 'control plane for context' to manage governance, data, and costs.

Rich Sutton: AI's 'Weird Field' Needs to Relearn 'Learning'
AI pioneer Rich Sutton argues that the field's 'weird' focus on 'continual learning' misses the point; true AI, he says, learns continuously from experience, a principle LLMs are only partially following.

Context Engineering: The Key to Better AI Agents
AI experts discuss context engineering strategies, detailing how compaction, caching, and memory management improve AI agent performance and reduce costs.

AI Writing: Authorship's New Frontier
Philosophical debates on authorship from the 1960s illuminate our modern reactions to AI-generated text, revealing a social need for 'author-function' in content.

Vals: The AI Scorekeeper We Need
Vals is building the essential trust layer for AI, evaluating models on real-world tasks, not just academic benchmarks.

Similarweb Unlocks AI Ad Visibility
Similarweb launches AI Ads, offering marketers visibility into advertising on ChatGPT and Google AI.

LangChain CEO on Building Better AI Agents
LangChain CEO Harrison Chase discusses the essential components of AI agents, the importance of custom harnesses, and the power of evals and observability for continuous improvement.

UC Berkeley PhD Student Challenges AI Evaluation Methods
Parth Asawa, a PhD student at UC Berkeley, argues that current AI evaluation methods are insufficient for measuring continual learning and calls for a new benchmark approach.

Sonar: AI Coding Needs Verification, Not Just Generation
Anirban Chatterjee of Sonar discusses the challenges and solutions for AI-driven software development, emphasizing the need for verification and governance.

AI Agents: The New Primitives of Software
Kwindla Kramer of Daily discusses the historical evolution of computing and the future of AI-native software, drawing parallels from Vannevar Bush to today's AI agents.

Saoud Rizwan: Open Source is Dead, Long Live Open Source
Cline founder Saoud Rizwan argues that AI's impact on open source is profound, but open-weight models offer a cost-effective future, challenging proprietary AI dominance.

Databricks: The Data-Native AI Assistant
Databricks unveils its AI Assistant, emphasizing deep data integration and governance for enterprise use.

Cloudflare Adds AI to Internet Data Explorer
Cloudflare launches Radar Researcher, an AI tool that lets users query Internet data using plain language, simplifying access to complex insights.

Databricks Charts New AI Frontier
Databricks introduces agentic workflows, enabling AI to autonomously plan, execute, and refine multi-step tasks for enterprise operations.

AI Inference: 10x Faster Models & Self-Optimization
Philip Kiely and Ali Taha of Baseten discuss AI inference, LLM optimization, speculative decoding, and the engineering behind cutting-edge AI models.

Crusoe Cloud's fastokens v2 boosts AI speed
Crusoe Cloud's fastokens v2 offers native tiktoken support and major speed boosts for AI model inference and training.

AI Chatbot Race: Who's Really Winning?
New data reveals ChatGPT leads in reach and retention, but Claude AI is exploding in growth, challenging established dominance.

Netflix Bets on LLMs for Smarter Recommendations
Netflix's GenRec system uses LLMs to power recommendations, shifting from feature engineering to context engineering for smarter, more efficient content discovery.

Together AI partners with Moonshot AI
Together AI partners with Moonshot AI to offer Kimi K3 and future models, providing developers with day-zero access to large-scale open-source AI.

Ditch LLM Chasing, Build Once
Otari's unified gateway simplifies LLM integration, allowing teams to access multiple models without rebuilding infrastructure for each new provider.

Uber Eats Uses AI Agents to Enhance Food Photos at Scale
Uber Eats' computer vision team details their AI agent system for enhancing food photos, focusing on closed-loop feedback and continuous learning.

Google Experts Share AI Agent Evaluation Best Practices
Google's Preetika Bhateja & Daniel Bump share essential strategies for building effective AI agent evaluation systems, from initial 'vibing' to scaling with LLM judges.

OpenAI Explains Value Maximization with GPT-5.6
OpenAI's Build Hour explains 'value maxing' with GPT-5.6, detailing efficiency tips and Ploy's migration strategy.

Anthropic Launches Claude Opus 5
Anthropic releases Claude Opus 5, delivering near-frontier AI intelligence at a reduced cost with state-of-the-art performance in coding and knowledge work.

Kimi K3 Challenges Claude Fable 5 on Code Quality, Slashes Cost
Kimi K3 challenges Claude Fable 5 on coding benchmarks, offering similar quality at a third of the cost and the benefits of an open-weight model.

OpenCode CEO on 20x Growth & AI Agent Market
OpenCode CEO Jay V reveals how the platform achieved 20x growth, reaching 4.6M users by supporting any AI model and becoming a key player in the global coding agent market.