The AI revolution is no longer confined to massive data centers. Ahmad Osman, Founder & CEO of Osmantic, presented "The Desktop Frontier" at the AI Engineer World's Fair, detailing the significant strides made in local and open-source AI models. Osman predicts that within 18 months, AI intelligence equivalent to GLM 5.2 will be runnable on a single RTX 5090 with 32GB of VRAM, a development he considers conservative.
The Compression Curve of AI
For years, the narrative in AI has been about scaling up: bigger models, more parameters, and larger clusters. However, Osman highlighted a crucial shift: "A second curve has become impossible to ignore over the last three years. Similar capability bands are appearing in smaller active footprints and more ownable systems." He introduced the concept of "impact per parameter," emphasizing the efficiency gains that allow models to achieve comparable or better results with significantly fewer parameters and a smaller hardware footprint.
Osman illustrated this with a personal anecdote: "A year ago, I used to run Llama 2 on an RTX 3090. It's now running Qwen 3.5, 3.6, 27 billion parameter. That's better than Llama's 405, a 400 billion plus parameter model that you beat with a 27 billion parameter model a year and a half after." This trend, he noted, is not random but driven by ongoing research, architectural innovations, and efficiency gains.
"Densing Law" in Action
Citing research from Nature Machine Intelligence, Osman referred to this phenomenon as the "densing law" of LLMs, where capability density is rising. He explained that every 3.5 months, there's a roughly 50% increase in parameters, whether dense or activated, leading to more intelligence from the models we use.
As examples of this progress, Osman pointed to GLM 5.2, a 744 billion parameter model with only 40 billion activated, supporting up to 1 million tokens context. He also mentioned NVIDIA's Nemotron 3 Ultra, which demonstrated that more efficient training can be done on hardware, making fine-tuning and specialized model development more accessible.
