Filip Makraduli, representing Superlinked, took the stage to discuss a critical gap in the AI ecosystem: the lack of robust infrastructure for small model inference. He highlighted that while many teams focus on building powerful models, the systems to deploy and run them efficiently are often an afterthought. This oversight, Makraduli explained, leads to wasted resources and suboptimal performance, especially for AI agents and complex workflows.
Bridging the Infrastructure Gap
Makraduli began by framing the problem: the AI community's tendency to overlook the practicalities of inference, particularly for smaller models. He noted that his own previous work focused heavily on model training and evaluation, but he soon realized the crucial importance of the underlying infrastructure. This realization led him to join Superlinked, a company focused on building this essential layer for AI search and document processing.
The 'Yin' and 'Yang' of Inference
He introduced a helpful analogy of the 'yin' and 'yang' of model inference. The 'yin' represents the models themselves, their performance, accuracy, and the advancements in areas like open-source model development. The 'yang,' conversely, is the infrastructure required to make these models work effectively in production, encompassing aspects like routing, autoscaling, monitoring, and deployment.
Makraduli presented data showing the rapid growth of open-source models, highlighting that this ecosystem is not slowing down. However, he emphasized that the performance of these models is heavily reliant on the quality of the inference infrastructure. He pointed to a graph illustrating how model performance can degrade significantly with increased context length, a phenomenon known as 'context rot.' This underscores the need for sophisticated infrastructure to manage context effectively.
From Theory to Production: Superlinked's Solution
Superlinked's approach to this problem is to build a comprehensive inference engine that supports a wide array of models and architectures. Makraduli showcased a diagram illustrating their production cluster, which is designed for efficiency and scalability. This cluster features multiple pools of GPUs (L4, A100, H100) to accommodate different model requirements. Key components include:
