Sohail Shaikh and Ankush Rastogi of Prosodica present a compelling argument against the common LLM agent design pattern of statically loading all available tool definitions into every prompt. In their talk, "The 100-Tool Agent Is a Trap," they highlight the significant drawbacks of this approach, which they term the 'fat agent trap.' This method, while functional for small-scale demonstrations, quickly becomes inefficient and unreliable in production environments as the number of tools scales.
The core issue, as explained by Shaikh and Rastogi, is that the naive approach leads to several critical problems: 'token bloat,' where the prompt becomes excessively large due to the inclusion of all tool schemas; 'accuracy crashes,' as the model struggles to select the correct tool from an overwhelming list; 'cost explosions,' driven by the high token count per request; and 'context crowding,' leaving insufficient space for actual reasoning.
The 'Fat Agent Trap' Detailed
Shaikh and Rastogi illustrate these problems with concrete data. They show that with just 10 tools, an agent's accuracy is around 78%, but this plummets to 40% with 100 tools, and further degrades to a mere 13% accuracy when handling 741 tools. This decline is attributed to the model being forced to process an unnecessarily large amount of information for every single request. The sheer volume of tool schemas within the prompt overwhelms the LLM's ability to accurately identify and utilize the correct tool.
The financial and performance implications are equally stark. Loading 741 tools requires approximately 127,000 tokens per request. This not only drives up costs but also significantly increases latency, making the agent slow and unresponsive. The presentation contrasts this with a 'Just-In-Time' (JIT) approach, where only a handful of relevant tool schemas (typically 3-5) are injected into the prompt at runtime. This strategy maintains high accuracy (above 83% even with 700+ tools) and drastically reduces token usage to around 1,000 tokens per request, a 99% reduction, while keeping latency near-flat.
Semantic Routing as the Solution
The key to achieving this efficiency and accuracy lies in semantic routing. Shaikh and Rastogi describe this as akin to Retrieval Augmented Generation (RAG), but for tools instead of documents. The process involves:
