1 articles with this tag
The Matryoshka training framework nests language model sub-models, drastically cutting compute costs and enhancing speculative decoding throughput while maintaining performance parity.