Mistral AI, a Paris-based competitor to OpenAI, has launched Magistral, a new series of reasoning-optimized large language models (LLMs). The series includes two models: Magistral Small (open-source) and Magistral Medium (proprietary), both boasting advanced features like chain-of-thought reasoning and multilingual capabilities. Magistral Medium, in particular, demonstrates impressive performance on complex tasks, outperforming some competitors in speed and accuracy.
The company's innovative training methodology, detailed in a research paper, utilizes reinforcement learning without a critic model, leading to significant improvements in LLM response quality. Mistral AI's work challenges existing assumptions in reinforcement learning and highlights the potential of this approach for developing high-performing LLMs.
