The latest release of OLMo 2 positions it as the leading fully open language model to date, delivering performance that rivals and in some cases exceeds open-weight models, of comparable size.
With the introduction of their 7B and 13B parameter models trained on up to 5 trillion tokens, Ai2 says OLMo 2 achieves state-of-the-art efficiency and performance across a range of benchmarks.
In evaluations, OLMo-2-7B outperforms Llama-3.1-8B, while OLMo-2-13B beats Qwen-2.5-7B, despite requiring fewer computational resources for training. These results place OLMo 2 on the Pareto frontier for open models, balancing computational efficiency with high benchmark scores.
Key Results and Comparisons of OLMO 2
The OLMo 2 models excel across both familiar development benchmarks, such as ARC Challenge and HellaSwag, and unseen evaluation metrics, including AGIEval and GSM8k. Notably:
- OLMo-2-7B matches or surpasses larger models, proving highly efficient relative to training FLOPs.
- OLMo-2-13B, tuned for instruction tasks, outperforms competitors like Qwen-2.5-14B in instruction-following and reasoning tasks.
By combining high performance with complete transparency, releasing weights, datasets, training code, and recipes, OLMo 2 continues the trend of narrowing the gap between open and proprietary models.
