#Deep Learning
50 articles with this tag

Ilya Sutskever's SSI Partners With NVIDIA
Ilya Sutskever's Safe Superintelligence Inc. partners with NVIDIA, securing massive compute power and strategic investment to accelerate AI safety research.

Unified AI Music Generation
A unified AI music generation framework leverages novel architectures and training strategies to produce high-quality full-length songs from diverse inputs.

SkewAdam: Rethinking MoE Optimizer Memory
SkewAdam drastically cuts MoE training memory by tailoring optimizer state to parameter populations, achieving superior perplexity and enabling training on accessible hardware.

Netflix's LLM Engine Revealed
Netflix reveals its custom LLM serving infrastructure built on vLLM and NVIDIA Triton, enabling flexibility and performance within its production environment.

Together AI Boosts GPU Cluster Uptime
Together AI introduces major reliability and control upgrades for its GPU Clusters, including automated node repair and enhanced operational oversight.

OpenAI's GPT-Red: AI Learns to Police Itself
OpenAI's new GPT-Red system uses AI to find and fix vulnerabilities, making models like GPT-5.6 Sol significantly more robust against attacks.

Anthropic Invests $10M in Canadian AI
Anthropic commits $10 million CAD to Canadian AI research institutions, fostering advancements in responsible AI applications and recognizing Canada's early contributions to the field.

AI Learns to Smell: The Science of Olfactory AI
AI is learning to smell thanks to companies like Osmo, which are building models to understand, predict, and design scents by mapping molecular structures to olfactory perception.

VisionAId: On-Device Vision for the Visually Impaired
VisionAId transforms smartphones into real-time visual assistants for the visually impaired, leveraging on-device AI and few-shot learning for personalized object recognition and multimodal guidance.

AI Accelerates Weight Drug Discovery
AI is accelerating the development of new weight management drugs like Wegovy, promising faster discovery cycles and personalized therapies.

Anthropic's Chloe Lubinski on AI, Ethics, and Future
Anthropic's Chloe Lubinski discusses AI's learning process, the importance of interpretability, and how AI reflects human values, emphasizing the need for ethical guidance in AI development.
AI Is Everywhere: From Your Inbox to Your Doctor's Office
AI learns from data to perform tasks requiring human intelligence, with generative AI applications now creating novel content.
AI Cracks Rare Genetic Disease Codes
OpenAI's reasoning model helped identify diagnoses in 18 previously unsolved rare genetic disease cases, demonstrating AI's potential in re-analyzing complex medical data.
Databricks, NVIDIA Forge AI Partnership
Databricks and NVIDIA are deepening their collaboration, integrating NVIDIA's GPUs, Vera CPUs, and AI software to accelerate enterprise AI development and agentic applications on the Databricks Lakehouse platform.
Databricks Unleashes AI Runtime for GPU Training
Databricks launches AI Runtime, a serverless GPU platform to simplify large-scale deep learning model training and accelerate AI development.
Phase Dominance in AI Image Recognition
AI image classifiers exhibit a striking phase dominance for identity encoding, mirroring human vision principles, with architectural differences shaping its expression.

5 AI Research Papers Shaping AI's Future
Discover five key AI research papers that reveal the current trajectory and future directions of artificial intelligence development.

RunPod's Audry Hsu on IDE-Integrated GPU Cloud Deployment
Audry Hsu from RunPod discusses the platform's IDE-integrated GPU cloud deployment, addressing developer pain points and showcasing the company's rapid growth and adoption.

Together AI Pushes LLM Context Limits to 5 Million Tokens
Max Ryabinin from Together AI discusses breaking barriers in LLM training, detailing techniques to achieve 5 million token context lengths and their impact on memory and performance.
LinkedIn's Generative Recommender Speed-Up
LinkedIn engineers drastically improved Generative Recommender training efficiency, cutting GPU hours by up to 65% through system-level optimizations.
LocateAnything: Parallel Decoding for Vision
LocateAnything revolutionizes vision-language models with Parallel Box Decoding, boosting speed and accuracy in visual grounding and detection.

Uber's DeepETT Boosts Traffic Forecasts
Uber's DeepETT system revolutionizes traffic forecasting with deep learning, boosting accuracy and handling 2 million predictions per second.

Uber Eats' Search Engine Gets Smarter
Uber Eats enhances its delivery search with semantic AI, leveraging LLMs and optimized infrastructure for speed, scale, and accuracy.

Omar Sanseviero on Google's AI Strategy
Omar Sanseviero from Google DeepMind discusses Google's AI strategy, focusing on efficient models, multimodality, and open innovation in AI.

Graph Neural Networks Explained: GNN Basics & Models
Explore the essentials of Graph Neural Networks (GNNs), from their basic principles to key models like GCNs, GraphSAGE, GATs, GINs, and Transformers.
AI Agents Supercharge GPU Kernel Development
LinkedIn is leveraging AI agents to automate complex GPU kernel engineering for its Liger Kernel project, accelerating AI model performance.
AI Agents Build Better AI
LinkedIn Engineering details how AI agents are revolutionizing model development through automated, iterative refinement loops.

Attractors Unlock Scalable Reasoning
Equilibrium Reasoners (EqR) leverage learned attractor landscapes to achieve scalable, adaptive test-time compute allocation, dramatically boosting accuracy on complex reasoning tasks.

Jure Leskovec on Relational Foundation Models
Jure Leskovec, AI researcher and Stanford professor, discusses Relational Foundation Models, a new AI approach for understanding complex enterprise data and its applications.

Claude's Corner: Ndea - Chollet's $43M Bet That Scale Isn't AGI
Francois Chollet built ARC-AGI, the benchmark the entire AGI industry has spent a decade failing to beat. Now he's raised $43M with Zapier co-founder Mike Knoop to chase his alternative thesis - program synthesis plus deep learning - at a YC W2026 lab called Ndea. Here's why it matters, why $43M, and why you can't replicate it.

Microsoft's MatterSim accelerates material discovery
Microsoft's MatterSim AI platform achieves experimental validation, faster simulations, and introduces a powerful multi-task model for advanced material discovery.

LLMs Slash Neural Architecture Search Costs
Delta-Code Generation uses LLMs to produce compact architecture refinements, dramatically cutting costs and improving NAS efficiency.

Andrej Karpathy: AI Models Need Human-Like Reasoning
Andrej Karpathy discusses the evolution of AI from programming to prompting, emphasizing the current need for models to develop human-like reasoning.

Yann LeCun Pushes AI Beyond Language Models
Yann LeCun is championing a new AI architecture, JEPA, that moves beyond language models to learn world representations and predict future states, aiming for more robust AI.

Y Combinator Decodes AI: Recursive Reasoning Models
Y Combinator Decoded explores how recursive AI models, like HRM and TRM, are revolutionizing AI reasoning by mimicking the human brain's efficiency.

Cloudflare Unweights LLMs by 22%
Cloudflare's 'Unweight' system slashes LLM model sizes by up to 22% using lossless compression, enhancing inference speed and efficiency.

AI Agents Collaborate to Solve Math Problems
Together AI's EinsteinArena platform enables AI agents to collaborate on complex scientific problems, achieving new breakthroughs in mathematics.
DMax: Parallel Decoding for Diffusion LLMs
DMax revolutionizes diffusion language models with Soft Parallel Decoding, boosting TPF significantly while preserving accuracy and achieving 1,338 TPS.
AI Accelerates Molecular Dynamics at Scale
AI-driven potentials are now integrated into GROMACS, enabling near ab initio fidelity for large-scale molecular dynamics simulations on multi-GPU systems.

Google Researchers Explore AI Storage Efficiency
Google researchers are developing AI compression techniques to reduce model storage needs by sixfold, aiming to lower costs and boost efficiency in AI development.

NVIDIA's Jensen Huang on AI's Future and Compute Demands
NVIDIA CEO Jensen Huang discusses the company's strategic evolution in AI, the importance of co-design, and the future of AI computing.

Mamba-3: Inference-First SSMs Arrive
Together AI's Mamba-3 advances state space models with a focus on inference speed, outperforming previous versions and some Transformers.

Andrej Karpathy on AI Agents: More Than Just Code
Andrej Karpathy discusses the evolution of AI agents beyond code generation, emphasizing the need for modularity, self-improvement, and human-AI collaboration for future advancements.
VideoAtlas: Unlocking Long-Context Video AI
VideoAtlas AI offers a lossless, hierarchical grid representation and Video-RLM for scalable, robust long-context video understanding with logarithmic compute growth.
Databricks Adds Serverless NVIDIA GPUs
Databricks launches AI Runtime, offering serverless NVIDIA GPUs for simplified AI model training and fine-tuning directly within the Lakehouse.
MoDA: Unlocking LLM Depth Scaling
Mixture-of-Depths Attention (MoDA) tackles LLM signal degradation by enabling cross-layer attention, boosting performance with minimal overhead.
AI vs. ML: What's the Difference?
AI is the broad concept of machines mimicking human intelligence, while machine learning is a specific method where systems learn from data.

AI's Consciousness Debate
Vishal Misra and Martin Casado discuss LLM functionality, the path to AGI, and the role of data in AI development.
SCORE: Recurrent Depth for Deep Networks
SCORE introduces a recurrent, iterative approach to deep neural networks, accelerating training and reducing parameter counts without complex ODE solvers.

AI Agents Now Do Overnight Research
An automated system uses AI agents to conduct overnight LLM training experiments, modifying code and iterating on models autonomously.