The quest for large language models tailored for specific enterprise agentic tasks is intensifying.
Controlled Fine-Tuning for Agentic Prowess
The Palmyra x6 LLM emerges as a significant advancement, meticulously optimized for enterprise-oriented agentic workloads. Its development eschews brute force, opting instead for a deliberate and conservative post-training strategy. The model is built upon a Mixture-of-Experts base, refined with Anchored Supervised Fine-Tuning using a compact, verified corpus of synthetic tool-use trajectories. This recipe is characterized by its restraint: a mere 626 trajectories, a single training epoch, a low learning rate, and a KL anchor to the frozen base model, all optimized with a Muon + Adam hybrid optimizer.
Benchmark Dominance and Safety Credentials
The impact of this controlled approach is evident in Palmyra x6's performance. It shows substantial gains over previous default models for Writer Agent tasks and stacks up favorably against numerous recent models on public benchmarks. Notably, it achieved the highest score of $0.785$ on BFCL Core and posted the highest six-benchmark mean within its cohort. Beyond raw performance, the model also distinguished itself in bias and safety evaluations, demonstrating competitive or leading results relative to comparators.
