Databricks says its Proteus harness generated Qwen 3.5 122B kernels that ran 1.8 to 5.2 times faster than the best vLLM implementations.
Companies working on this
Profiles of the companies named in this story, with funding and a one-liner from our database.
The gain comes from abandoning generic kernels and specializing for the exact shapes seen at runtime.
