A team led by researchers at Prior Labs and the University of Freiburg has put tabular foundation models to a broader test with BeyondArena, a unified benchmark that stresses IID, temporal and grouped tasks at once.
It is a reality check.
BeyondArena evaluates 11 models on 142 datasets, curated from 1,128 candidates under rigorous protocols, to cover classification and regression from 100 to 1 million rows, up to high dimensionality and with messy features like text and high-cardinality categories. To make that curation reproducible the authors publish DataFoundry, a Python framework and metadata schema for tabular data, and integrate the whole suite into TabArena's open-source ecosystem at tabarena.ai.
The headline result splits cleanly. In the paper's Elo comparison of the best model per family, three open-source tabular foundation models tested in in-context learning dominate tiny to medium IID slices, while traditional gradient-boosted decision trees and tuned MLPs, with default, tuned and post-hoc ensembled variants, still dominate non-IID, large and high-dimensional slices. The authors report the same pattern across grouped, temporal, wide and text-rich subsets. TabPFN-3.5's own report notes that tuned and ensembled MLPs still lead on BeyondArena's grouped, temporal and large-data slices, even as the model claims a 150 Elo lead overall.
That contrast matters because prior benchmarks narrowed the field to where foundation models already shine. The most quoted predecessor, TabArena, is described as the industry standard benchmark which contains datasets with up to 100,000 training data points, and much of its leaderboard has been fought over IID tables. Chatbot Arena, the language-model counterpart, ranks models by human head-to-head votes, producing an Elo-style rating, and has collected over 1.5M human votes from side-by-side LLM battles. BeyondArena borrows the Elo idea for supervised tabular error but applies it to temporal and grouped splits where leakage and distribution shift, not prompt preference, decide the winner. Co-authors include Lennart Purucker, Andrej Tschalzev, Nick Erickson of Prior Labs and AutoGluon, and Frank Hutter, a long-time AutoML lead.
