
How Deep Learning Finally Cracked Messy Tables - Frank Hutter
Frank Hutter, co-founder of Prior Labs, talks about TabPFN, a tabular foundation model that makes predictions in a single forward pass, and the research behind it.
TabPFN is pre-trained on synthetic datasets drawn from a prior over structural causal models, rather than on real data. At prediction time it takes the whole training table as context and outputs an approximation of the Bayesian posterior predictive distribution, without per-dataset training or hyperparameter search. Frank explains how this grew out of his earlier work on AutoML and neural architecture search, how the priors are built and revised, and why tabular data was hard for deep learning for so long.
The conversation also covers the TabArena benchmark, how the architecture changed from TabPFN v1 to v3, scaling to larger tables, using the model with coding agents, test-time compute, Google's TabFM, causal inference and interventions, and relational data. At the end, a short update Frank recorded after the interview covers the TabPFN-3.5 release.
Prior Labs:
TabPFN-3.5: https://priorlabs.ai/tabpfn-3-5
https://priorlabs.ai/careers
TOC:
00:00 Introduction
00:44 Welcome and Frank's background
02:05 Why tabular data was hard for deep learning
10:17 Pre-training on synthetic data
12:52 The TabArena benchmark
19:28 From AutoML to neural architecture search
26:34 TabPFN as a learned algorithm
30:50 Bayesian prediction in one forward pass
39:37 Scaling to larger tables
47:48 Using TabPFN with coding agents
57:47 Output heads and architecture from v1 to v3
1:05:29 Test-time compute and adaptation
1:13:32 Google's TabFM
1:16:53 How the priors are designed
1:18:40 Correlation, causation and interventions
1:35:22 Relational and multimodal data
1:38:31 Use in organisations
1:46:38 The open research arm
1:50:21 Update: TabPFN-3.5
REFS:
TabPFN v2, Nature (Hollmann et al., 2025)
https://www.nature.com/articles/s41586-024-08328-6
Transformers Can Do Bayesian Inference (Müller et al.)
https://arxiv.org/abs/2112.10510
TabArena (Erickson et al.)
https://arxiv.org/abs/2506.16791
AutoGluon-Tabular (Erickson et al.)
https://arxiv.org/abs/2003.06505
Beyond IID: How General Are Tabular Foundation Models, Really?
https://arxiv.org/abs/2606.30410
Neural Architecture Search: A Survey (Elsken, Metzen & Hutter)
https://arxiv.org/abs/1808.05377
Auto-WEKA (Thornton et al.)
https://www.cs.ubc.ca/~hutter/papers/AutoWEKA-KDD2013.pdf
TabPFN v1 (Hollmann et al., 2022)
https://arxiv.org/abs/2207.01848
TabPFN-3 technical report
https://arxiv.org/abs/2605.13986
TabPFN-2.5 report
https://arxiv.org/abs/2511.08667
CAAFE (Hollmann et al.)
https://arxiv.org/abs/2305.03403
TabICL (Qu et al.)
https://arxiv.org/abs/2502.05564
TabICLv2 (Qu et al.)
https://arxiv.org/abs/2602.11139
Google TabFM
https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/
TALENT benchmark (Ye et al.)
https://arxiv.org/abs/2407.00956
Do-PFN (Robertson et al.)
https://arxiv.org/abs/2506.06039
CausalPFN (Balazadeh et al.)
https://arxiv.org/abs/2506.07918
Causal Foundation Models with Partial Graphs (Reuter et al.)
https://arxiv.org/abs/2602.14972
RelBench (Robinson et al.)
https://arxiv.org/abs/2407.20060
RelArena-α, TabPFN-Rel and RPI
https://arxiv.org/abs/2608.16319
TabPFN on GitHub
https://github.com/PriorLabs/TabPFN
TabPFN-3.5 technical report
https://arxiv.org/abs/2609.17895
Otto Group Product Classification Challenge (Kaggle, 2015)
https://www.kaggle.com/competitions/otto-group-product-classification-challenge
---RESCRIPT:https://app.rescript.info/share/e99676c25ee6189fbf54c9be07eb623e
D'autres épisodes de "Machine Learning Street Talk (MLST)"



Ne ratez aucun épisode de “Machine Learning Street Talk (MLST)” et abonnez-vous gratuitement à ce podcast dans l'application GetPodcast.








