Not that long ago, building a machine learning system meant carefully designing features, choosing a model architecture, training it from scratch, tuning hyperparameters, and repeating this process for every new problem. Then foundation models changed how we think about text, images, and increasingly other modalities: pretrain once, then adapt to new tasks through context.
The same transition is now happening for structured data.
With Kumo Tabular, you provide a table with labeled examples and rows you want predictions for. The model produces predictions in a single forward pass --— with no task-specific training, no fine-tuning, and no feature engineering.
What makes this especially exciting to me is that this is not just a new model, but part of a rapidly growing research ecosystem around tabular foundation models. There is tremendous innovation happening across academia and industry in architectures, synthetic pretraining, in-context learning, evaluation, and efficient inference.
NVIDIA wants to be an active part of that ecosystem —-- contributing research, releasing models openly, and building infrastructure that helps the community push the field forward.
There are several aspects of the work I find particularly interesting.
** The models are pretrained entirely on synthetic tables generated from structural causal models, allowing us to expose them to enormous diversity without training on customer data or benchmark datasets. The largest model sees more than 100 million synthetic tables during pretraining.
** The resulting models are both accurate and efficient. Across major tabular benchmarks, Kumo Tabular improves upon strong existing approaches while requiring no per-dataset training. On TabArena, for example, Kumo Tabular Large sits on the accuracy–speed Pareto frontier and is 17× faster at prediction than LimiX-2.
To me, the bigger story is the direction ML is moving: from hand-built models for individual tasks to pretrained models that learn broad representations of a domain and can solve new problems from context. We have seen this transformation in language and vision. It is exciting to see it now reaching the enormous world of structured data.
Huge congratulations to the team — and to the broader tabular foundation model research community whose ideas and work are making this new paradigm possible.
We’re excited to contribute, learn, and help build this ecosystem together.