Tabular foundation models still stumble beyond IID

BeyondArena tests 11 models on 142 datasets and finds tabular foundation models lead on small IID data but trail GBDTs on large and non-IID tasks.

Tabular foundation models still stumble beyond IID
Image credit: StartupHub.ai

A team led by researchers at Prior Labs and the University of Freiburg has put tabular foundation models to a broader test with BeyondArena, a unified benchmark that stresses IID, temporal and grouped tasks at once.

It is a reality check.

BeyondArena evaluates 11 models on 142 datasets, curated from 1,128 candidates under rigorous protocols, to cover classification and regression from 100 to 1 million rows, up to high dimensionality and with messy features like text and high-cardinality categories. To make that curation reproducible the authors publish DataFoundry, a Python framework and metadata schema for tabular data, and integrate the whole suite into TabArena's open-source ecosystem at tabarena.ai.

The headline result splits cleanly. In the paper's Elo comparison of the best model per family, three open-source tabular foundation models tested in in-context learning dominate tiny to medium IID slices, while traditional gradient-boosted decision trees and tuned MLPs, with default, tuned and post-hoc ensembled variants, still dominate non-IID, large and high-dimensional slices. The authors report the same pattern across grouped, temporal, wide and text-rich subsets. TabPFN-3.5's own report notes that tuned and ensembled MLPs still lead on BeyondArena's grouped, temporal and large-data slices, even as the model claims a 150 Elo lead overall.

That contrast matters because prior benchmarks narrowed the field to where foundation models already shine. The most quoted predecessor, TabArena, is described as the industry standard benchmark which contains datasets with up to 100,000 training data points, and much of its leaderboard has been fought over IID tables. Chatbot Arena, the language-model counterpart, ranks models by human head-to-head votes, producing an Elo-style rating, and has collected over 1.5M human votes from side-by-side LLM battles. BeyondArena borrows the Elo idea for supervised tabular error but applies it to temporal and grouped splits where leakage and distribution shift, not prompt preference, decide the winner. Co-authors include Lennart Purucker, Andrej Tschalzev, Nick Erickson of Prior Labs and AutoGluon, and Frank Hutter, a long-time AutoML lead.

The design choices are deliberate. The authors define IID versus non-IID by the test split that mirrors deployment, random for random holdouts, time-ordered for future forecasting, grouped for unseen entities, and they separate temporal tabular tasks from time-series forecasting. Ablations in the paper probe what that means for validity, showing sensitivity to outer test splits for grouped data, inner validation splits for tiny or non-IID data, preprocessing for grouped and text features, and probability calibration for log loss.

For practitioners the guidance is pragmatic, not triumphal. If your data is small, clean and IID, a foundation model in a single forward pass can match a four-hour AutoGluon ensemble. If it is large, wide, or must generalize forward in time or across groups, classic GBDTs and deep tabular nets remain the safer default. The gap is not closed by scaling alone and the benchmark's next value will be whether future foundation models learn the inductive biases that temporal and grouped tasks require.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.