We’re not at the foothills. That would imply there’s a direction up the mountain. We’re in the inchoate state of the Agentic AI industry. 99% daily fodder among the furtive star-dust that will come together to represent the pillars of the industry.
Everything is up in the air, orbiting the central need for autonomous, personalized Large Language Models (LLMs) that are affordable and capable of rapidly learning, or even evolving based on their own internal decisions. Infrastructure tools, AI Agents, LLMs releasing newer models, newer modalities, AI Agent builders platform, rapid gamut expansion. I can enumerate the list for you but it grows by the hour.
My two cents; nobody has luxurious stability, including the LLM developers who are the first principal foundation of the industry. The structural makeup of the Agentic AI industry will likely confound us all.
What is clear is the importance of data. Indeed data is the lifeblood of pre-training. Venerated Ilya Sustkever just echoed that at NeurIPs 2024, hinting at his SSI1 foundation model that’s being trained on data beyond what the open internet has to offer. But data is the lifeblood of fine-tuning and outputting results, where the technique of Retrieval Augmented Generation (RAG) gives them task-specific utility. It’s where LLMs and applications thereof become useful to us, more than their representations during pre-training.
Access to premium and fresh data is now the fulcrum at which LLMs ride or die. It’s one of the last miles of LLMs’ value. It’s become the difference between a trial and a production deployment
