Machine Learning Street Talk (MLST) podcast

How Deep Learning Finally Cracked Messy Tables - Frank Hutter

0:00
1:53:12
Reculer de 15 secondes
Avancer de 15 secondes

Frank Hutter, co-founder of Prior Labs, talks about TabPFN, a tabular foundation model that makes predictions in a single forward pass, and the research behind it.


TabPFN is pre-trained on synthetic datasets drawn from a prior over structural causal models, rather than on real data. At prediction time it takes the whole training table as context and outputs an approximation of the Bayesian posterior predictive distribution, without per-dataset training or hyperparameter search. Frank explains how this grew out of his earlier work on AutoML and neural architecture search, how the priors are built and revised, and why tabular data was hard for deep learning for so long.


The conversation also covers the TabArena benchmark, how the architecture changed from TabPFN v1 to v3, scaling to larger tables, using the model with coding agents, test-time compute, Google's TabFM, causal inference and interventions, and relational data. At the end, a short update Frank recorded after the interview covers the TabPFN-3.5 release.


Prior Labs:

TabPFN-3.5: https://priorlabs.ai/tabpfn-3-5

https://priorlabs.ai/careers


TOC:

00:00 Introduction

00:44 Welcome and Frank's background

02:05 Why tabular data was hard for deep learning

10:17 Pre-training on synthetic data

12:52 The TabArena benchmark

19:28 From AutoML to neural architecture search

26:34 TabPFN as a learned algorithm

30:50 Bayesian prediction in one forward pass

39:37 Scaling to larger tables

47:48 Using TabPFN with coding agents

57:47 Output heads and architecture from v1 to v3

1:05:29 Test-time compute and adaptation

1:13:32 Google's TabFM

1:16:53 How the priors are designed

1:18:40 Correlation, causation and interventions

1:35:22 Relational and multimodal data

1:38:31 Use in organisations

1:46:38 The open research arm

1:50:21 Update: TabPFN-3.5


REFS:

TabPFN v2, Nature (Hollmann et al., 2025)

https://www.nature.com/articles/s41586-024-08328-6

Transformers Can Do Bayesian Inference (Müller et al.)

https://arxiv.org/abs/2112.10510

TabArena (Erickson et al.)

https://arxiv.org/abs/2506.16791

AutoGluon-Tabular (Erickson et al.)

https://arxiv.org/abs/2003.06505

Beyond IID: How General Are Tabular Foundation Models, Really?

https://arxiv.org/abs/2606.30410

Neural Architecture Search: A Survey (Elsken, Metzen & Hutter)

https://arxiv.org/abs/1808.05377

Auto-WEKA (Thornton et al.)

https://www.cs.ubc.ca/~hutter/papers/AutoWEKA-KDD2013.pdf

TabPFN v1 (Hollmann et al., 2022)

https://arxiv.org/abs/2207.01848

TabPFN-3 technical report

https://arxiv.org/abs/2605.13986

TabPFN-2.5 report

https://arxiv.org/abs/2511.08667

CAAFE (Hollmann et al.)

https://arxiv.org/abs/2305.03403

TabICL (Qu et al.)

https://arxiv.org/abs/2502.05564

TabICLv2 (Qu et al.)

https://arxiv.org/abs/2602.11139

Google TabFM

https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/

TALENT benchmark (Ye et al.)

https://arxiv.org/abs/2407.00956

Do-PFN (Robertson et al.)

https://arxiv.org/abs/2506.06039

CausalPFN (Balazadeh et al.)

https://arxiv.org/abs/2506.07918

Causal Foundation Models with Partial Graphs (Reuter et al.)

https://arxiv.org/abs/2602.14972

RelBench (Robinson et al.)

https://arxiv.org/abs/2407.20060

RelArena-α, TabPFN-Rel and RPI

https://arxiv.org/abs/2608.16319

TabPFN on GitHub

https://github.com/PriorLabs/TabPFN

TabPFN-3.5 technical report

https://arxiv.org/abs/2609.17895

Otto Group Product Classification Challenge (Kaggle, 2015)

https://www.kaggle.com/competitions/otto-group-product-classification-challenge


---RESCRIPT:https://app.rescript.info/share/e99676c25ee6189fbf54c9be07eb623e

D'autres épisodes de "Machine Learning Street Talk (MLST)"