#LLM Efficiency
4 articles with this tag

AI Research
LOCKS Unlocks LLM Long Context Efficiency
LOCKS revolutionizes long-context LLMs by approximating KV cache with spectral summaries, achieving near full-quality inference while drastically cutting latency and computation.
6 days ago

AI Research
From LLM APIs to Local Neural Artifacts
Fuzzy-function programming enables compiling LLM-powered functions locally, matching large model performance with minimal resources.
about 1 month ago
AI Research
AdaCodec: Efficient Video MLLM Encoding
AdaCodec revolutionizes video MLLMs by using predictive visual coding to drastically cut tokenization costs and latency, achieving superior performance at a fraction of the budget.
2 months ago
AI Research
DMax: Parallel Decoding for Diffusion LLMs
DMax revolutionizes diffusion language models with Soft Parallel Decoding, boosting TPF significantly while preserving accuracy and achieving 1,338 TPS.
4 months ago