# Tabular foundation models still stumble beyond IID _BeyondArena tests 11 models on 142 datasets and finds tabular foundation models lead on small IID data but trail GBDTs on large and non-IID tasks._ **Published:** 2026-09-28 **Source:** https://www.startuphub.ai/ai-news/ai-research/2026/tabular-foundation-models-still-stumble-beyond-iid --- A team led by researchers at [Prior Labs](https://arxiv.org/html/2606.30410v1) and the University of Freiburg has put tabular foundation models to a broader test with [BeyondArena](https://arxiv.org/html/2606.30410v1), a unified benchmark that stresses IID, temporal and grouped tasks at once. It is a reality check. BeyondArena evaluates 11 models on 142 datasets, curated from 1,128 candidates under rigorous protocols, to cover classification and regression from 100 to 1 million rows, up to high dimensionality and with messy features like text and high-cardinality categories. To make that curation reproducible the authors publish DataFoundry, a Python framework and metadata schema for tabular data, and integrate the whole suite into TabArena's open-source ecosystem at tabarena.ai. The headline result splits cleanly. In the paper's Elo comparison of the best model per family, three open-source tabular foundation models tested in in-context learning dominate tiny to medium IID slices, while traditional gradient-boosted decision trees and tuned MLPs, with default, tuned and post-hoc ensembled variants, still dominate non-IID, large and high-dimensional slices. The authors report the same pattern across grouped, temporal, wide and text-rich subsets. TabPFN-3.5's own report notes that tuned and ensembled MLPs still lead on BeyondArena's grouped, temporal and large-data slices, even as the model claims a 150 Elo lead overall. That contrast matters because prior benchmarks narrowed the field to where foundation models already shine. The most quoted predecessor, TabArena, is described as the industry standard benchmark which contains datasets with up to 100,000 training data points, and much of its leaderboard has been fought over IID tables. Chatbot Arena, the language-model counterpart, ranks models by human head-to-head votes, producing an Elo-style rating, and has collected over 1.5M human votes from side-by-side LLM battles. BeyondArena borrows the Elo idea for supervised tabular error but applies it to temporal and grouped splits where leakage and distribution shift, not prompt preference, decide the winner. Co-authors include Lennart Purucker, Andrej Tschalzev, Nick Erickson of [Prior Labs](https://www.startuphub.ai/startups/prior-labs) and AutoGluon, and Frank Hutter, a long-time AutoML lead. The design choices are deliberate. The authors define IID versus non-IID by the test split that mirrors deployment, random for random holdouts, time-ordered for future forecasting, grouped for unseen entities, and they separate temporal tabular tasks from time-series forecasting. Ablations in the paper probe what that means for validity, showing sensitivity to outer test splits for grouped data, inner validation splits for tiny or non-IID data, preprocessing for grouped and text features, and probability calibration for log loss. For practitioners the guidance is pragmatic, not triumphal. If your data is small, clean and IID, a foundation model in a single forward pass can match a four-hour AutoGluon ensemble. If it is large, wide, or must generalize forward in time or across groups, classic GBDTs and deep tabular nets remain the safer default. The gap is not closed by scaling alone and the benchmark's next value will be whether future foundation models learn the inductive biases that temporal and grouped tasks require. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.