AI Kernel Optimization & Local AI Efficiency
Experts discuss multi-GPU kernel optimization, the rise of 'Intelligence per Watt', and the growing viability of local AI.

Visual TL;DR
From the article 6 mentionsThe discussion highlighted the accelerating trend of specialization in AI hardware, driven by the vastly different requirements for training and inference workloads.
From the articleSul pointed out that while significant progress has been made in optimizing single-GPU performance, GPU networking remains a critical bottleneck.
From the articleJon Saad-Falcon, a Stanford PhD student, then introduced the concept of 'Intelligence per Watt' (IPW) and 'Intelligence per Joule' (IPJ) as key metrics for evaluating AI efficiency.
framework simplifying creation of efficient multi-GPU AI kernels for better performance
From the articleThe session kicked off with Stuart Sul, a researcher at Cursor and Stanford PhD student, who presented his work on ParallelKittens, a framework designed to simplify the creation of efficient multi-GPU AI kernels.
From the article 6 mentionsSpeakers emphasized that while cloud-based AI has dominated, the rise of more powerful local accelerators and capable smaller language models is paving the way for a significant shift towards distributed, on-device AI.
automating AI research and kernel development for faster innovation cycles
From the article 2 mentionsMark Saroufim, a former PyTorch maintainer and co-founder of Core Automation, discussed the challenges and opportunities in automating AI research, particularly in the realm of kernel development.
From the article 2 mentionsSpeakers emphasized that while cloud-based AI has dominated, the rise of more powerful local accelerators and capable smaller language models is paving the way for a significant shift towards distributed, on-device AI.
Contents(4)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.