AI's Future: Cheaper Tokens, Longer Tasks

Sal research CEO Neil Movva discusses the future of AI, focusing on cheaper tokens, long-running agents, and the shift from low-latency to proactive intelligence.

8 min read
Neil Movva, CEO of Sal research, speaks in an interview setting.
Invest with the Best
Visual TL;DR
Cheaper AI TokensDriver
Sal research aims to be the absolute cheapest provider of AI tokens
From the article 6 mentionsMovva elaborated on the company's mission, stating, "We think that whenever you make something 10 times cheaper, it's a new product category." Sal research aspires to achieve this for AI tokens, believing that the ability of machines to think is profound and should be democratized.
Shift to PersistenceContext
moving from low-latency responses to persistent, background AI agents
From the article 3 mentionsYou wanted more persistence, more long horizon tasks." Movva believes it's now obvious that the future of agentic inference is in these longer, more complex tasks.
Long-Horizon TasksEffect
enabling new categories of AI applications that run over extended periods
From the article 8 mentionsMovva explained that while low latency was crucial for early AI applications, the future lies in enabling agents to perform long-horizon tasks, running for hours, days, or even weeks.
Proactive AI AgentsOutcome
AI will anticipate needs and act autonomously, changing human interaction
From the article 5 mentionsLooking ahead, Movva painted a picture of proactive intelligent agents.
Cheaper AI TokensDriver
Sal research aims to be the absolute cheapest provider of AI tokens
From the article 6 mentionsMovva elaborated on the company's mission, stating, "We think that whenever you make something 10 times cheaper, it's a new product category." Sal research aspires to achieve this for AI tokens, believing that the ability of machines to think is profound and should be democratized.
Token Factory ModelCore
leveraging open-source LLMs for any task at an unbeatable price point
From the articleSal research positions itself as a 'token factory,' offering an API that allows anyone to leverage open-source large language models for any task at an unbeatable price.
Hardware InnovationsCore
GPUs and memory advancements crucial for efficient, affordable AI operations
From the article 4 mentionsHe highlighted Nvidia's Tensor Cores as a key innovation that accelerated matrix multiplication, a core operation for machine learning.
Shift to PersistenceContext
moving from low-latency responses to persistent, background AI agents
From the article 3 mentionsYou wanted more persistence, more long horizon tasks." Movva believes it's now obvious that the future of agentic inference is in these longer, more complex tasks.
Open Source LeverageCore
utilizing open-source models for cost reduction and user sovereignty
Long-Horizon TasksEffect
enabling new categories of AI applications that run over extended periods
From the article 8 mentionsMovva explained that while low latency was crucial for early AI applications, the future lies in enabling agents to perform long-horizon tasks, running for hours, days, or even weeks.
Proactive AI AgentsOutcome
AI will anticipate needs and act autonomously, changing human interaction
From the article 5 mentionsLooking ahead, Movva painted a picture of proactive intelligent agents.
Contents(5)

In a candid discussion about the future of artificial intelligence, Neil Movva, CEO of Sal research, articulated a vision for making AI capabilities more accessible and affordable. Movva, whose company aims to be the absolute cheapest provider of AI tokens, believes the industry is on the cusp of a significant shift away from immediate, low-latency responses towards more persistent, background AI agents. This transition, he argues, will unlock new categories of AI applications and fundamentally change how humans interact with intelligent machines.

The Token Factory: Driving Down Costs for Abundant Intelligence

Sal research positions itself as a 'token factory,' offering an API that allows anyone to leverage open-source large language models for any task at an unbeatable price. "My job is to make the tokens as cheap as humanly possible," Movva stated. "I will achieve that and I will do it through every layer in the stack available to me." This focus on cost reduction is driven by a core belief in abundance: making intelligence accessible to as many people as possible at a sustainable cost.

The full discussion can be found on Invest with the Best's YouTube channel.

Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper - Invest with the Best
Ex-NVIDIA Engineer: Why AI Is About to Get 1000x Cheaper, from Invest with the Best

Movva elaborated on the company's mission, stating, "We think that whenever you make something 10 times cheaper, it's a new product category." Sal research aspires to achieve this for AI tokens, believing that the ability of machines to think is profound and should be democratized.

The Shift from Latency to Persistence: Enabling Long-Horizon Tasks

The conversation highlighted a key divergence from the current AI landscape, which often prioritizes low-latency, interactive experiences. Movva explained that while low latency was crucial for early AI applications, the future lies in enabling agents to perform long-horizon tasks, running for hours, days, or even weeks. "It doesn't matter if it spits out tokens at 100 tokens per second. Maybe 10 is just fine," he noted, emphasizing that this shift allows for greater efficiency and opens up new possibilities.

This vision is supported by early indications in AI development, particularly with models like Opus45, which showed suitability for longer tasks. Movva predicts that the future of agentic inference will be dominated by these long-running tasks, moving away from the user being constantly in the loop and waiting for responses. Instead, he envisions AI agents operating more like human colleagues, handling high-level tasks proactively and allowing users to check in periodically.

Leveraging Open Source and User Sovereignty

Movva identified the rise of open-source models as a significant tailwind for Sal research. "I think we are starting to see an increasing number of our customers and the broader market care about owning intelligence," he said. "They want to have control, sovereignty over the thing that they depend on." This desire for control aligns with the open-source movement, which provides users with the weights and the right to deploy models as they see fit.

He contrasted this with the market's previous focus on low-latency inference, driven by a key customer like Cursor. "As of six months ago, it started to look like maybe low latency wasn't the only thing you wanted from an agent. You wanted more persistence, more long horizon tasks." Movva believes it's now obvious that the future of agentic inference is in these longer, more complex tasks.

The Future of AI: Proactive Agents and Unbounded Possibilities

Looking ahead, Movva painted a picture of proactive intelligent agents. He imagined a Siri-like assistant running in the background, understanding a user's daily communications and proactively offering assistance. This future hinges on the availability of incredibly cheap intelligence, where users are willing to spend tokens without an immediate promise of return. "The long lens view to take on this is that we have a form of intelligence that can tackle any verifiable problem," he stated. This includes most software, mathematical proofs, and scientific discovery.

He projected that by the end of the year, workloads would be split 50/50 between background and real-time tasks, with a long-term trend towards 90/10 in favor of background processes. The ultimate goal, he suggested, is to be limited only by the questions we can ask, with AI models capable of chasing down every possible follow-up to a high-level query.

Innovations in Hardware and Memory: The Role of GPUs

Movva also touched upon the fundamental role of hardware and computational efficiency. He highlighted Nvidia's Tensor Cores as a key innovation that accelerated matrix multiplication, a core operation for machine learning. "Nvidia, great graphics company, obviously has had market share dominance in GPUs and gaming graphics for quite some time," he noted. He explained how Nvidia's early recognition of AI's potential led them to dedicate significant resources to optimizing their hardware for these tasks, creating specialized units like Tensor Cores.

He contrasted the benefits of SRAM with DRAM in chip design. While SRAM offers faster access but lower density, DRAM provides higher density but requires constant refreshing. "It's my job to figure out what parallelism scheme I'm going to use that's going to make this chip suitable for inference," Movva stated, emphasizing the importance of understanding the trade-offs between different hardware approaches to achieve optimal performance for AI workloads.

© 2026 StartupHub.ai. All rights reserved. Do not enter, scrape, copy, reproduce, or republish this article in whole or in part. Use as input to AI training, fine-tuning, retrieval-augmented generation, or any machine-learning system is prohibited without written license. Substantially-similar derivative works will be pursued to the fullest extent of applicable copyright, database, and computer-misuse laws. See our terms.