The TabFM model
This document describes BigQuery's built-in TabFM tabular regression and classification model.
The built-in TabFM model is an implementation of Google Research's open source TabFM model. The Google Research TabFM model is a foundation model for tabular data that enables zero-shot regression and classification on structured data through in-context learning. Because the TabFM model is pre-trained on hundreds of millions of synthetic datasets generated using structural causal models, it captures complex feature interactions and generalizes well to unseen real-world tables across many domains.
You can use the TabFM model with the
AI.PREDICT function
to perform regression and classification on structured data in a single
forward pass without having to train a model, optimize hyperparameters, or
engineer features.
The prediction results are comparable to conventional supervised tree-based
algorithms such as XGBoost and random
forests. If you want more model tuning options than the TabFM model offers,
you can train a supervised model such as a
boosted tree
or
random forest
model and use it with the
ML.PREDICT function
instead.
To generate predictions with the TabFM model on tabular data, use the
AI.PREDICT function.
To evaluate predicted values from the TabFM model against the actual values,
use the
AI.EVALUATE function.
To learn more about the Google Research TabFM model, use the following resources:
When you use TabFM through BigQuery, your usage is governed by
the Google Cloud Terms of Service and allows
for commercial uses. The non-commercial license associated with the publicly
downloadable TabFM weights on GitHub and Hugging Face applies only to
self-hosted downloads and does not restrict usage within
BigQuery.