Generative Video: Efficiency and Interaction Take Center Stage

Keegan McCallum of uRun discusses the rapid evolution of generative video, focusing on efficiency and real-time interaction, and how this tech is reshaping content creation and HCI.

Keegan McCallum speaking at a podium with a screen showing generative video examples.
AI Engineer
Visual TL;DR
Generative Video EvolutionContext
early 2023 examples like Will Smith spaghetti meme to Sora and SeeDance
From the article 4 mentionsKeegan McCallum, founder of uRun, a new inference provider focused on interactive media, recently presented a compelling overview of the rapid advancements in generative video technology.
Beyond Visual FidelityDriver
industry focus shifting from just video quality to efficiency and long-horizon generation
From the articleHowever, McCallum emphasized that the true frontier is now shifting beyond just visual fidelity to encompass efficiency and extended generation capabilities.
Efficiency & InteractionCore
McCallum highlights efficiency and real-time interaction as the new frontiers
From the article 6 mentionsHe also stressed the importance of accessibility, noting that generative video can provide a more intuitive and less text-heavy interaction method for individuals who think visually or find text-based AI interactions challenging.
uRun's RoleCore
Keegan McCallum's company, uRun, provides inference for interactive media
From the article 4 mentionsHe showcased a demo of uRun's model, Helios, a distilled version of Yuan 2.1 14B, demonstrating its ability to produce long, continuous generations and clips that are generated faster than they can be consumed.
Reshaping Content CreationEffect
new capabilities are fundamentally changing how content is produced and consumed
From the articleMcCallum also touched upon the shift in content creation, moving away from a 'slot machine' approach where users spend significant resources trying to achieve a desired shot.
New HCI ParadigmsEffect
generative video tech is also transforming human-computer interaction methods
Future of ContentOutcome
driving innovation in content creation and interactive experiences for users
From the article 2 mentionsWith real-time steerability, creators can now iteratively refine their generated content, piloting agents and directing the output with granular control.
Contents(3)

Keegan McCallum, founder of uRun, a new inference provider focused on interactive media, recently presented a compelling overview of the rapid advancements in generative video technology. Speaking at the AI Engineer World's Fair, McCallum highlighted that while the industry has been captivated by improvements in video quality, a parallel and equally significant leap is occurring in efficiency and the capability for long-horizon video generation.

Generative Video: Efficiency and Interaction Take Center Stage - AI Engineer
Generative Video: Efficiency and Interaction Take Center Stage, AI Engineer

The Evolution of Generative Video

McCallum contrasted early generative video examples from 2023, like the infamous 'Will Smith eating spaghetti' meme, with the more refined outputs seen in 2024 with models like Sora. He noted that while Sora represented a significant step forward, models like SeeDance are now achieving photorealism that is almost indistinguishable from reality. However, McCallum emphasized that the true frontier is now shifting beyond just visual fidelity to encompass efficiency and extended generation capabilities.

He showcased a demo of uRun's model, Helios, a distilled version of Yuan 2.1 14B, demonstrating its ability to produce long, continuous generations and clips that are generated faster than they can be consumed. These models, he explained, maintain a quality comparable to the frontier models of the previous year but operate at a significantly lower cost. "The generation in the bottom right corner," McCallum pointed out, referring to a video of a yellow sports car on a scenic road, "is a long continuous generation. And the other video is a bunch of clips that have been generated faster than you can consume them." He further elaborated that Helios is capable of running at 19.5 frames per second on a single NVIDIA H100 GPU and supports minute-scale generation.

Efficiency and Accessibility Driving Innovation

McCallum presented data illustrating the rapid improvement in efficiency, with models like Helios achieving state-of-the-art quality for both long and short video generation on a single GPU. He highlighted the dramatic cost reduction, stating that what once cost $10 per minute can now be achieved for approximately a hundredth of that cost. "$10 can get you 3 hours worth of generated video continuously with most of these models," he noted, underscoring the democratization of this technology.

This increased efficiency and speed are unlocking new use cases. McCallum envisioned applications like a 'magic mirror' where users can see themselves in different outfits or with different hairstyles in real-time. He also stressed the importance of accessibility, noting that generative video can provide a more intuitive and less text-heavy interaction method for individuals who think visually or find text-based AI interactions challenging. "There are more opportunities to have companions or visual mediums that are going to allow more people to experience the things a lot of us have with coding models," he said.

The Future of Content Creation and Interaction

McCallum also touched upon the shift in content creation, moving away from a 'slot machine' approach where users spend significant resources trying to achieve a desired shot. With real-time steerability, creators can now iteratively refine their generated content, piloting agents and directing the output with granular control. He cited modern models like Google's Gemini Omni as examples of advancements allowing for more full-fidelity clip rendering.

The technical challenges in building applications that leverage these advanced models are significant. McCallum outlined the need for distributed GPUs, robust networking infrastructure (WebRTC, ICE, TURN), and the ability to wire multiple models together in continuous streaming workflows. To address this, uRun has developed a React component that simplifies the integration of interactive video into applications, backed by a programmable Python runtime for building complex, asynchronous pipelines.

Looking ahead, McCallum predicted that by 2026, the industry will require not just platforms but 'software factories' to manage and deploy these AI capabilities. He concluded by inviting developers to collaborate as design partners, emphasizing that uRun is actively hiring to push the boundaries of human-computer interaction through generative video.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.