An independent benchmark of tabular machine-learning systems has found that pretrained foundation models can compete with, and in many cases outperform, a carefully tuned XGBoost baseline without retraining their weights for each new dataset. The test compared TabPFN and TabICL with tuned and untuned boosting across 14 datasets drawn from the Grinsztajn tabular benchmark.
The experiment used the same train-test split and timing approach for each contender on an Nvidia RTX 4070 Ti. Each dataset was capped at 3,000 rows and split 70-30 with stratification. Unlike conventional boosted-tree training, the foundation models consumed the training examples as context and generated predictions in a forward pass. Their application still exposes a fit step, but the author says that operation largely transfers data rather than running gradient descent.
That architecture changes where computing time is spent. Fitting was almost immediate for the foundation models, while inference carried more of the cost. XGBoost showed the opposite pattern because it spends time constructing an ensemble before serving predictions. The benchmark author reported that hyperparameter search for XGBoost took as long as about 50 seconds on an individual test dataset, whereas a foundation model could provide a quick baseline without that search.
Results varied by problem. On the credit, bank-marketing and hospital-readmission examples discussed in detail, the foundation models performed strongly against tuned boosting. The hospital dataset produced a larger gap in area under the curve than in raw accuracy, with the reported AUC reaching 0.6484 compared with 0.6264. On an electricity dataset, TabICL matched tuned boosting to four decimal places while TabPFN trailed.
The wide Bioresponse table exposed a clearer limitation. With 419 input columns, TabPFN recorded a score of 0.7567, fell behind even untuned XGBoost and required 18.1 seconds. TabICL remained the top performer in that case. The mixed results suggest that table width, not only row count, can materially affect whether this approach is attractive.
The test is a practitioner-run comparison rather than an independent peer-reviewed evaluation, and its small-row cap, hardware and chosen datasets limit how broadly the outcome can be generalized. It nevertheless illustrates a practical role for tabular foundation models: rapidly establishing a serious baseline before teams commit time and computing resources to model selection and tuning. Production use would still require evaluation on the organization’s actual data, latency constraints and accuracy metric. The comparison also reinforces that accuracy alone may not capture operational value: inference speed, tuning expense and metric selection can change which system is preferable for a particular table and decision process.



