Visual TL;DR. Slow Search Latency addressed by Instructed-Retriever-1 Model. Instructed-Retriever-1 Model uses Parallel Test-Time Scaling. Parallel Test-Time Scaling enables Broader Evidence Retrieval. Parallel Test-Time Scaling enables Precise Context Selection. Broader Evidence Retrieval leads to 3x Faster Search. Precise Context Selection leads to 3x Faster Search. 3x Faster Search and Halved Answer Generation. Halved Answer Generation results in 2-Second Time to First Token.
- Slow Search Latency: Traditional agentic search systems process results sequentially, leading to higher latency
- Instructed-Retriever-1 Model: New Databricks model powering the performance upgrade for Knowledge Assistant
- Parallel Test-Time Scaling: Flipping sequential computation to fan out tasks in parallel during initial search
- Broader Evidence Retrieval: Allows for wider retrieval of relevant information upfront in the search process
- Precise Context Selection: Enables more accurate selection of context for better answer generation
- 3x Faster Search: Significant boost in Knowledge Assistant search speed, over three times faster
- Halved Answer Generation: Answer generation time is also reduced by approximately half
- 2-Second Time to First Token: Achieving approximately two seconds for the initial response to be delivered
Visual TL;DR