# TabFM TabFM (Tabular Foundation Model) is a pretrained model from Google Research for classification and regression on tables: spreadsheets, database exports, anything with rows and columns. You give it your labeled rows plus the rows you want predictions for, and it answers in a single forward pass. No training on your data, no hyperparameter search, no feature engineering pipeline. Google announced it on June 30, 2026 (blog post by Weihao Kong and Abhimanyu Das; Google Research also made TimesFM, the same zero-shot idea for time series), and the technical report landed on arXiv on September 29, 2026. Why does this matter? Tabular data sits behind a big share of predictive work in companies (churn, fraud, credit risk, pricing), and it's an area where deep learning kept losing to gradient-boosted trees like XGBoost. The tree approach works, but every new dataset means fitting from scratch, tuning, and hand-crafting features. TabFM bets that a single [[AI Foundation Models|foundation model]] can skip all of that. ## How it works The trick is in-context learning, the same mechanism that lets [[Large Language Models (LLMs)]] pick up a task from a few examples in the prompt. With TabFM, the "prompt" is your table. Calling `fit()` updates no weights at all; your training rows become the context, and the test rows attend to them at prediction time. Tables are awkward for a [[Transformers|Transformer]], though. Text is a 1D ordered sequence; a table is 2D, and shuffling its rows changes nothing. So TabFM borrows from two earlier models, TabPFN and TabICL: - **Cell embeddings**: each cell is encoded with learned Fourier features, with separate encodings for numerical and categorical columns - **Alternating column and row attention**: attention across rows (per column) learns each column's distribution; attention across features (per row) learns how columns interact. The column side uses inducing points (a Set Transformer trick) so cost grows linearly with the number of rows - **Row compression**: 8 learned CLS tokens squeeze each row into one 1,024-dimensional vector - **In-context predictor**: a 24-layer Transformer runs over those compressed rows. Test rows can only attend to labeled rows, so each prediction is independent of the other test rows Most of the weight sits in that last block: about 403M of the 409M parameters (the paper calls it a 400M-parameter model). Training used no real tables at all. Google generated hundreds of millions of [[Synthetic Data|synthetic datasets]] from structural causal models (random cause-and-effect graphs that spit out mixed numerical and categorical columns, missing values, label noise and class imbalance). The paper gives two reasons: good, diverse public tables are scarce, and industrial tables are proprietary or sensitive. Pretraining went through a four-stage curriculum, growing the context from 2,048 to 16,384 rows, with tables capped at 100 columns. ## Results Everything is measured on TabArena, a public benchmark of 51 datasets (38 classification, 13 regression, from 700 to 150,000 rows) that ranks methods with Elo scores from head-to-head wins. The paper compares 67 method configurations: | Method | Classification Elo | Regression Elo | |---|---|---| | TabFM-Auto (TabFM + Gemini feature engineering) | 1940.7 | 2392.4 | | TabFM+ (ensemble of 32 views, calibrated) | 1838.0 | 2189.2 | | **TabFM (zero-shot, single pass)** | **1768.6** | **2055.2** | | EXAONE-Tabular | 1768.0 | 1973.1 | | TabPFN-3 | 1641.9 | 1866.6 | | AutoGluon 1.5 (extreme preset) | 1669.7 | 1851.2 | Three things stand out to me: - **Zero-shot TabFM beats 4-hour AutoML runs**, and its regression lead is large. On classification it's basically tied with EXAONE-Tabular (0.6 Elo apart); the paper argues TabFM wins there on consistency across datasets (lowest error gap to a per-dataset oracle: 6.10% vs 9.47%) - **Head to head, across both tracks**, TabFM wins 75.9% of dataset folds against TabPFN-3 and 72.1% against AutoGluon 1.5 extreme - **The top spot goes to TabFM-Auto**, where Gemini 3.8 Flash writes and refines a Python feature engineering pipeline around the frozen model (up to 96 tries or 6 hours per dataset). It improves 41 of 51 datasets, mostly by building domain features (e.g. Reynolds numbers on an airfoil dataset). That's the same idea as [[LLM-Generated Features for Classical ML]]: let a language model describe the data, let a statistical model do the predicting An independent reproduction by Yash Raj Pandey (13 TabArena datasets, fold-matched) mostly backs the claims: zero-shot TabFM beat an XGBoost tuned with 100 Optuna trials on all 10 datasets he could compare, with margins up to +5.5 accuracy points. Two thin "wins" against TabPFN were within noise and he called them ties. He also found a bug that crashed `predict` on any multi-GPU machine; his fix was merged upstream. ## Availability and license - **Code**: [[Apache 2.0 License|Apache 2.0]], on GitHub. scikit-learn style `TabFMClassifier` / `TabFMRegressor`, with JAX and PyTorch backends - **Weights**: on [[HuggingFace]] (v1.0.0 in PyTorch and JAX, v1.1.0 in PyTorch) under the "TabFM Non-Commercial License v1.0". Research, testing and internal benchmarking only. No production use, no commercial use, no client deliverables, no redistribution. You need a separate commercial license from Google for any of that - **BigQuery**: TabFM runs inside Google Cloud BigQuery (in preview) through two SQL functions, `AI.PREDICT` and `AI.EVALUATE`. You pass a training table, a prediction table and the label column. BigQuery samples the training data and distributes inference over millions of rows So the open part is the code. The free weights are for experimenting; Google's paid route to production is BigQuery. ## Limits - **Max 10 classes** for classification, a hard architectural limit - **Width**: pretraining saw at most 100 columns; the library caps inference at 500 features. In the independent test, Bioresponse (1,777 features) failed outright - **Size**: pretraining saw tables of up to 16,384 rows. Bigger tables get subsampled. Pandey couldn't finish the 78k and 150k-row datasets in reasonable time - **Hardware**: the JAX backend needed about 17 GB of GPU memory regardless of table size (32-member ensemble by default). The later PyTorch backend with bfloat16 and activation chunking used 3 to 7 GB and fit 40,000 context rows on an RTX 4090. On CPU, one prediction with 5,000 context rows took about 21 minutes - **Synthetic-only training**: the model card says performance on specific real-world domains, minority groups and edge distributions isn't fully characterized. Free text isn't encoded semantically - **Explainability**: Google's own BigQuery post says to keep using XGBoost when you need feature importance, very large datasets, more features than TabFM supports, or full control over tuning ## Alternatives - **TabPFN** (Prior Labs): the model that started this whole family (prior-data fitted networks, published in Nature). TabFM's architecture builds on it directly. SAP announced in May 2026 that it would acquire Prior Labs - **TabICL / TabICLv2**: the source of the compressed-row idea TabFM reuses - **EXAONE-Tabular, Mitra-v2, Xiaomi-TabLDM**: other synthetic-trained tabular foundation models published in 2026 - **AutoGluon**: AutoML that stacks many models (trees, neural nets) under a time budget - **Gradient-boosted trees** (XGBoost, LightGBM, CatBoost): still the default when you need explainability, huge tables, or production use without licensing questions ## Community reaction The [[Hacker News]] thread (97 points, 14 comments) was impressed but skeptical about the reporting. One commenter called TabPFN state of the art and TabFM impressive, but criticized showing only Elo, since Elo hides how BIG the improvement is, and called the GitHub results section (a folder of undocumented parquet files) a mess. Another read that as hiding weak spots and flagged the missing tuned-XGBoost comparison; a reply noted that TabArena's AutoGluon baselines include tuned XGBoost. Someone pointed out that the plot leaves out the strongest TabPFN variants (thinking and ensembled), so the comparison isn't apples to apples. Two commenters remarked on the timing right after SAP's Prior Labs deal. And the "up to 150,000 rows" line got the expected jokes about Big Data, plus a side debate on whether more data actually makes a better tabular model. To be fair, the September paper answers part of that: it reports geometric-mean error, wins and oracle improvability next to Elo. ## My take Zero-shot prediction on a table you just dropped in is a big deal for everyone who isn't a data scientist. In BigQuery it's one SQL query. But the license is the real story for anyone outside Google Cloud: you can't ship the free weights in a product. If you need a tabular model in production on your own infrastructure, look at TabPFN or the tree models. And TabFM-Auto's results suggest the next gains come from LLMs engineering features around the model, more than from the model itself. ## References - [Introducing TabFM: A zero-shot foundation model for tabular data](https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/) (Weihao Kong, Abhimanyu Das; Google Research blog, 2026-06-30) - [TabFM: A Zero-Shot Foundation Model for Tabular Data](https://arxiv.org/abs/2609.37959) (Kong et al., arXiv 2609.37959, 2026-09-29) - [TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models](https://arxiv.org/abs/2609.37989) (Fu et al., arXiv 2609.37989, 2026-09-29; abstract only) - [google-research/tabfm on GitHub](https://github.com/google-research/tabfm) (README, CHANGELOG, `classifier_and_regressor.py`, `results/`) - [google/tabfm-1.0.0-pytorch model card](https://huggingface.co/google/tabfm-1.0.0-pytorch) and [TabFM Non-Commercial License v1.0](https://huggingface.co/google/tabfm-1.0.0-pytorch/blob/main/LICENSE) (Hugging Face) - [google/tabfm-1.1.0-pytorch](https://huggingface.co/google/tabfm-1.1.0-pytorch) (Hugging Face) - [Introducing TabFM in BigQuery: Predictive analytics reimagined](https://cloud.google.com/blog/products/data-analytics/tabfm-adds-predictive-ml-to-bigquery) (Vaibhav Sethi, Xi Cheng; Google Cloud blog) - [I Tried to Break Google's New Tabular Foundation Model. Then I Fixed It.](https://yashrajpandey.com/writing/breaking-google-tabfm/) (Yash Raj Pandey, 2026-07-01, updated 2026-07-07) - [SAP to Acquire Prior Labs](https://news.sap.com/2026/05/sap-to-acquire-prior-labs-establish-frontier-ai-lab-europe/) (SAP News Center, 2026-05-04) - [TabArena leaderboard](https://huggingface.co/spaces/TabArena/leaderboard) - [PriorLabs/TabPFN on GitHub](https://github.com/PriorLabs/TabPFN) - [Hacker News discussion](https://news.ycombinator.com/item?id=48739919) (via the HN Algolia API) and [HN thread on the independent evaluation](https://news.ycombinator.com/item?id=48805514) - arXiv listings for EXAONE Tabular 1.0 (2608.25774), Mitra-v2 (2609.04540) and Xiaomi-TabLDM (2609.03880), abstracts only ## Related - [[AI Foundation Models]] - [[Machine Learning (ML)]] - [[LLM-Generated Features for Classical ML]] (TabFM-Auto applies the same idea) - [[Zero-Shot Classification]] - [[Synthetic Data]] - [[Transformers]] - [[Data Science]] - [[Google]] - [[Gemini]]