<script data-pm-proxy="intercept"></script><?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Mindful Modeler]]></title><description><![CDATA[Tabular foundation models, ML interpretability, and beyond by a statistician turned machine learner.]]></description><link>https://mindfulmodeler.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!WtQm!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png</url><title>Mindful Modeler</title><link>https://mindfulmodeler.substack.com</link></image><generator>Substack</generator><lastBuildDate>Tue, 01 Sep 2026 05:02:31 GMT</lastBuildDate><atom:link href="/__u/mindfulmodeler.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Christoph Molnar]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[mindfulmodeler@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[mindfulmodeler@substack.com]]></itunes:email><itunes:name><![CDATA[Christoph Molnar]]></itunes:name></itunes:owner><itunes:author><![CDATA[Christoph Molnar]]></itunes:author><googleplay:owner><![CDATA[mindfulmodeler@substack.com]]></googleplay:owner><googleplay:email><![CDATA[mindfulmodeler@substack.com]]></googleplay:email><googleplay:author><![CDATA[Christoph Molnar]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Trends in tabular foundation research]]></title><description><![CDATA[Based on 151 papers from ICML workshop "Foundation Models for Structured Data"]]></description><link>https://mindfulmodeler.substack.com/p/trends-in-tabular-foundation-research</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/trends-in-tabular-foundation-research</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Wed, 22 Jul 2026 11:33:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WtQm!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>151 papers on tabular and time series foundation models were presented at &#8220;Foundation Models for Structured Data&#8221;, a workshop that took place for the second time at ICML this month in Seoul.</p><p>I didn&#8217;t attend, but still marked the event in my calendar to check on current trends in tabular foundation research. Going through the titles and some abstracts was very informative, and I want to share my observations about the state of tabular foundation research.</p><p>If you are new to tabular foundation models, start with this series: </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;d9a17eef-1a30-4f76-b3fb-7cbec41faa9e&quot;,&quot;caption&quot;:&quot;Tree-based boosting algorithms have sat on the tabular throne for many years now. Many times, the deep learners have attempted to dethrone XGBoost, CatBoost, and other tree-based algorithms, but without success.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The rise of tabular foundation models&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:8489879,&quot;name&quot;:&quot;Christoph Molnar&quot;,&quot;bio&quot;:&quot;Writer and machine learning expert&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5dffee84-bdf9-4188-9e5d-b63c519aba2e_2998x2998.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-01-13T13:14:41.426Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!Sv6x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89dd7432-4067-4499-b283-e1a13335d5b3_1797x920.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:184297088,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:174,&quot;comment_count&quot;:26,&quot;publication_id&quot;:1078760,&quot;publication_name&quot;:&quot;Mindful Modeler&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WtQm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>Here are the trends and observations, based on this <a href="https://icml-structured-fm-workshop.github.io/accepted-papers/">list of papers</a>:</p><ul><li><p>Tabular foundation models are expanding beyond just regression and classification, <a href="/__u/mindfulmodeler.substack.com/p/im-betting-on-tabular-foundation">as I expected</a>. There are papers on survival analysis, Bayesian inference, causal identification, and more.</p></li><li><p>Time series was the biggest ML task/modality (besides tabular) with around 50 out of 151 papers. I attribute it to the workshop framing (&#8220;tabular and time-series&#8221;). By the way, <a href="https://tabularfoundationmodels.com/forecasting">I just published a book chapter on time series forecasting</a> with tabular foundation models.</p></li><li><p>Tabular foundation models are moving away from being just models toward becoming more of a system, with research on external infrastructure such as context management, input-space adapters, distillation, and integration into agents.</p></li><li><p>Benchmarking and evaluation were a huge topic. Also, some critique about the current, narrow way of benchmarking. I personally think this is improving thanks to new benchmarks like ScoringBench and BeyondArena.</p></li><li><p>Some researchers examined prior designs such as causal DAGs and graph priors.</p></li><li><p>An emerging topic is mechanistic interpretability for foundation models.</p></li><li><p>If you are curious about the application of tabular foundation models, you&#8217;ll find some applied papers in the list as well, in domains like healthcare.</p></li><li><p>Some papers explored multi-modal foundation models, like tabular+text and tabular+image.</p></li></ul><p>Some topics were surprisingly underrepresented:</p><ul><li><p>The focus was not as much on scaling foundation models, at least compared to  industry labs, where scaling tabular foundation models to larger data and speeding up inference is top priority.</p></li><li><p>There was a bit on relational foundation modeling (extending foundation models to a relational multi-table setup), but I expected more here.</p></li></ul><p>Papers that caught my interest: <a href="https://openreview.net/forum?id=r3RAi8Kqzl">Beyond Accuracy: Toward Trustworthy Tabular Foundation Models in Industrial Applications</a> and <a href="https://openreview.net/forum?id=PXSBtjo3Gd">Exploring Differences Between Tabular Enterprise Data and Public Benchmarks</a>.</p>]]></content:encoded></item><item><title><![CDATA[Time for a change]]></title><description><![CDATA[After 4 years of writing, I'm exploring what's next for me in ML / data science / AI (or whatever you want to call it)]]></description><link>https://mindfulmodeler.substack.com/p/time-for-a-change</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/time-for-a-change</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Thu, 16 Jul 2026 13:23:12 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/e917e8f4-41be-4861-b507-0d730a2978aa_4000x3000.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Sometimes you have to make a big change in life. You&#8217;ve probably been there yourself.</p><p>For me, this time is now.</p><p>In 2022, I made the scary decision to become self-employed as a writer for technical books. With the Interpretable ML book already under my belt, some financial runway, and a support network, I gave it a try. I didn&#8217;t want to look back at my life and regret never trying the author life. There was a real risk, since most authors can&#8217;t live from writing books. Fortunately, I managed to make a living from just books, and I am super grateful to everyone buying my books, making it possible for me to become a full-time author.</p><p>At some point, though, I felt my world shrinking due to the self-inflicted constraint of saying no to everything that is not about writing. So I added more projects to the mix: I gave workshops on ML topics like interpretability and uncertainty quantification, and I consulted clients. At one point, <a href="/__u/mindfulmodeler.substack.com/p/how-to-win-an-ml-competition-beyond">I even (successfully!) invested 3 months into a machine learning competition</a> to prove to myself that I&#8217;m still in the data science game and not just an impostor. Adding projects next to writing temporarily helped.</p><p>But now it&#8217;s time for a bigger change. I miss being part of a team, working together on a shared goal and projects. I miss ongoing exchanges with colleagues, the deadlines (yes, really), and this certain randomness (in the positive sense) you only get when being part of something bigger.</p><h2>What&#8217;s next</h2><p>I don&#8217;t have a concrete job lined up right now, but I started applying and having first conversations. I have no idea where I will end up yet, and that&#8217;s quite exciting.</p><p><strong>If you would like to work with me or know of a job opening in the ML/data science space, I&#8217;d be very happy to hear about it! Just reply to this email, or send an email to chris at christophmolnar.com.</strong></p><p>What it means for Mindful Modeler: I will continue posting here, but maybe on a more irregular schedule. Long-term, my writing for Mindful Modeler depends on what job I will have, whether it&#8217;s part-time or full-time. For some jobs, it might even make sense to post as part of the job itself.</p><p>I&#8217;ll continue writing the Tabular Foundation Models book during the job search. It&#8217;s probably going to be a bit shorter, focusing on the essentials. I already published a few chapters; <a href="https://tabularfoundationmodels.com/">check the tabular foundation book out here.</a></p><p>I&#8217;ll forever be grateful for having been able to write full-time. Thanks again to everyone who bought my books and followed along on my journey, here and elsewhere!</p>]]></content:encoded></item><item><title><![CDATA[TabFM minus the hype]]></title><description><![CDATA[top performance on TabArena; large model; slow inference; non-commercial license]]></description><link>https://mindfulmodeler.substack.com/p/tabfm-minus-the-hype</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/tabfm-minus-the-hype</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 07 Jul 2026 12:29:39 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WtQm!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Google Research published TabFM, a new tabular foundation model.</p><p>The AI influencers found out, and social media is now awash in TabFM posts.</p><p>Let&#8217;s talk about what TabFM is and isn&#8217;t, without the hype.</p><h2>TabFM is a tabular foundation model with in-context learning</h2><p>TabFM is a tabular foundation model like TabPFN, TabICL, and TabDPT.</p><p>More technically, TabFM is a transformer-based neural network pre-trained on hundreds of millions of synthetic datasets. TabFM makes predictions via in-context learning: there is no classic training step, but the training data is provided during inference time and serves as the context to predict the test data. It can do both regression and classification (up to 10 classes).</p><p>If you want to learn more about tabular foundation models in general, read my TFM series:</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;90713fc0-a536-43a2-9ca2-86ebec9a0247&quot;,&quot;caption&quot;:&quot;Tree-based boosting algorithms have sat on the tabular throne for many years now. Many times, the deep learners have attempted to dethrone XGBoost, CatBoost, and other tree-based algorithms, but without success.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The rise of tabular foundation models&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:8489879,&quot;name&quot;:&quot;Christoph Molnar&quot;,&quot;bio&quot;:&quot;Writer and machine learning expert&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5dffee84-bdf9-4188-9e5d-b63c519aba2e_2998x2998.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-01-13T13:14:41.426Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!Sv6x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89dd7432-4067-4499-b283-e1a13335d5b3_1797x920.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:184297088,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:169,&quot;comment_count&quot;:26,&quot;publication_id&quot;:1078760,&quot;publication_name&quot;:&quot;Mindful Modeler&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WtQm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>TabFM follows in the footsteps of especially TabPFN and TabICL, which the authors note in their <a href="https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/">research announcement post</a>. It re-uses architecture elements of both:</p><ul><li><p>For example, TabFM uses alternating row and column attention, as e.g.  TabPFN 2 did.</p></li><li><p>TabFM uses row compression and performs ICL over the compressed rows, as TabICL does.</p></li></ul><p>Also, TabFM is pre-trained on hundreds of millions of synthetic datasets generated with structural causal models, just like TabICL and TabPFN. However, not much is known about pretraining and prior, because the published code is just the inference code, and there is no white paper or paper published, just an inference code repo, the research announcement post, and the Hugging Face model (weight) release.</p><p>Being a tabular foundation model, TabFM inherits the &#8220;standard&#8221; pros and cons of modern foundation models: No tuning needed; highly performant; slow inference; need to provide training data at inference time; and so on.</p><p>So, TabFM is not the first tabular foundation model, nor the last. What is all the hype about?</p><p>One of the reasons for the hype: TabFM climbs to the top of TabArena, a benchmark for (primarily) tabular foundation models. If you want to learn more about TabArena, I've got you covered: </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;f0439b0b-41c4-45e0-9c93-703277464aa9&quot;,&quot;caption&quot;:&quot;Machine learning progresses through benchmarks. While I&#8217;ve been critical before of ML&#8217;s benchmark obsession and danger of getting stuck on benchmarks, benchmarks are essential to guide researchers and practitioners in the right direction. ImageNet, under the leadership of Fei-Fei Li, for example, ushered in the deep learning era. Also today, LLMs are la&#8230;&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;TabArena explained&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:8489879,&quot;name&quot;:&quot;Christoph Molnar&quot;,&quot;bio&quot;:&quot;Writer and machine learning expert&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5dffee84-bdf9-4188-9e5d-b63c519aba2e_2998x2998.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-03-31T12:09:21.288Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!hjDv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://mindfulmodeler.substack.com/p/tabarena-explained&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:192596694,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:18,&quot;comment_count&quot;:4,&quot;publication_id&quot;:1078760,&quot;publication_name&quot;:&quot;Mindful Modeler&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WtQm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>TabFM&#8217;s entry is not yet reflected in the live leaderboard at the time of writing, but results are waiting to be included in a <a href="https://github.com/autogluon/tabarena/pull/434">pull request</a>. These results are certainly impressive. </p><p>So, should we all be using TabFM now? I won&#8217;t be, for now, for two reasons.</p><h2>Slow inference and non-commercial license</h2><p>TabFM is larger than the other tabular foundation models. For example, it defaults to 32 estimators, when most other TFMs have a default of 8.  Also, many aspects of the architecture are scaled up: TabICL contains 4 CLS tokens; TabFM contains 8. TabICL has 12 ICL transformer blocks; TabFM has 24. TabICL has an embedding size of 128; TabFM has 256. In many ways, TabFM is larger than the current 2nd-generation (TabICL v2.0; TabPFN-3.0).</p><p>This scaling up and the alternating row and column attention come at a price: Inference is slower than for the other TFMs. The maximum number of features is 500.</p><p>The bigger issue no one talks about is the license. If you go to <a href="https://github.com/google-research/tabfm">their GitHub repo</a>, you see an <a href="https://github.com/google-research/tabfm?tab=Apache-2.0-1-ov-file">Apache-2.0 license</a>. Great news, right? But that repository only contains the inference code. The weights are <a href="https://huggingface.co/google/tabfm-1.0.0-pytorch">downloaded from Hugging Face</a> and come with the <a href="https://huggingface.co/google/tabfm-1.0.0-pytorch/blob/main/LICENSE">tabfm-non-commercial-v1.0 license</a>, which prohibits commercial use. In my understanding, if you want to use it commercially, you would have to contact Google and ask for permission.</p><p>For new TabFM users, I don&#8217;t think it&#8217;s obvious: At the time of writing, the non-commercial license is not mentioned in the README, and the weights are downloaded silently from Hugging Face. See also <a href="https://github.com/google-research/tabfm/issues/29">this GitHub issue</a>. I hope Google addresses this; otherwise, users might unknowingly violate the non-commercial license.</p><p>All in all, TabFM's performance seems impressive, and I am looking forward to seeing how it holds up in other benchmarks. The model seems to be geared towards performance, at the cost of inference speed. The non-commercial license makes it less attractive for me.</p>]]></content:encoded></item><item><title><![CDATA[Which ML models produce the best quantile estimates?]]></title><description><![CDATA[Results from ScoringBench]]></description><link>https://mindfulmodeler.substack.com/p/which-regressor-produces-the-best</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/which-regressor-produces-the-best</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 30 Jun 2026 08:10:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!YaZ6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>It&#8217;s not a secret that most ML models for quantile regression tend to be too biased towards the mean. You want a 10% quantile? The model&#8217;s quantile estimate will tend to be too large. And for large quantiles, the estimate is often too low. For example, I used two xgboost models to estimate 10% and 90% intervals, but ended up with an average coverage way below 80%  (more like 65%).</p><p>So, what is the best model (class) to use?</p><p>The short answer: Second-generation tabular foundation models like TabICL v2 and TabPFN-3.0.</p><p>These pre-trained models are capable of <a href="/__u/mindfulmodeler.substack.com/p/regression-should-predict-full-distributions">outputting the full predictive distribution</a>. Meaning you get the quantiles &#8220;for free&#8221;. Ignoring that TFM inference is kind of expensive, there is at least no additional cost.</p><h2>Evidence from ScoringBench</h2><p>The evidence for TFMs being good at predicting beyond the mean comes from the benchmark ScoringBench (<a href="http://scoringbench.com">website</a>|<a href="https://github.com/jonaslandsgesell/ScoringBench">code</a>|<a href="https://arxiv.org/html/2603.29928">paper</a>):</p><ul><li><p>This benchmark compares models based on proper scoring rules and other metrics that evaluate the entire predictive distribution.</p></li><li><p>An example of a simpler metric of the benchmark is the coverage of the 5%&#8211;95% prediction interval (which should be 90%).</p></li><li><p>An example of a proper scoring rule in the benchmark is the Continuous Ranked Probability Score (CRPS).</p></li></ul><p>The following figure shows the ranks of various ML models/algorithms:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!YaZ6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!YaZ6!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png 424w, /__u/substackcdn.com/image/fetch/$s_!YaZ6!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png 848w, /__u/substackcdn.com/image/fetch/$s_!YaZ6!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YaZ6!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!YaZ6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png" width="1456" height="1229" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1229,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:478803,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/195967381?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!YaZ6!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png 424w, /__u/substackcdn.com/image/fetch/$s_!YaZ6!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png 848w, /__u/substackcdn.com/image/fetch/$s_!YaZ6!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png 1272w, /__u/substackcdn.com/image/fetch/$s_!YaZ6!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9f34ec3a-89c7-470f-bdce-9f7dad26d3fe_1704x1438.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Comparing various models on various distribution metrics. Figure by <a href="https://arxiv.org/html/2603.29928">Landsgesell et al. (2026)</a>, CC BY-SA 4.0</figcaption></figure></div><p>But what exactly are the fine-tuned models that lead ScoringBench? Since these PFN-based tabular foundation models like TabICL and TabPFN are &#8220;just&#8221; neural networks, they can be fine-tuned. And if you want them to become better at predicting such scores, you can specifically fine-tune them for scoring-rule objectives. However, even there, it matters which metrics you fine-tune. From the paper: &#8220;TabICLv2 was fine-tuned with the CRPS objective, so it improves on CRPS but not on the log score, consistent with the expectation that fine-tuning shifts a model&#8217;s inductive bias toward the optimized scoring rule.&#8221;</p><h2>Two caveats</h2><p>The differences between tabular foundation models and the other ML algorithms seem large in the chart. But they are only large in terms of median ranks. The actual effect sizes are only negligible to small. However, paired with their often stronger performance, tabular foundation models are a great deal, if you can swallow the higher inference cost.</p><p>Even fine-tuning doesn&#8217;t <em>guarantee</em> coverage; it at best approximates it. Although it seems that TabPFN and TabICL are already pretty good at non-mean predictions. If you need (marginal) coverage guarantees, you would have to go with something like conformal prediction. If you are interested in that, I have a book for you:  <a href="https://christophmolnar.com/books/conformal-prediction">Introduction to Conformal Prediction With Python.</a></p><p>Anyways, if you need quantiles or anything beyond mean prediction, give tabular foundation models a try.</p>]]></content:encoded></item><item><title><![CDATA[When trees still beat tabular foundation models]]></title><description><![CDATA[As you might have noticed, I&#8217;m rather optimistic about tabular foundation models.]]></description><link>https://mindfulmodeler.substack.com/p/when-trees-still-beat-tabular-foundation</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/when-trees-still-beat-tabular-foundation</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 23 Jun 2026 13:53:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!oYhD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>As you might have noticed, I&#8217;m rather optimistic about tabular foundation models. And while TFMs have been gaining momentum, beating benchmark after benchmark, they don&#8217;t outperform every time. And it&#8217;s not just random tasks here and there, but there is a pattern to it: there are certain characteristics of datasets where gradient-boosted trees outperform tabular foundation models, and there is even a small cluster where tuned neural networks take the trophy.</p><h2>TabBench for classification tasks</h2><p>We take a look at the <a href="https://huggingface.co/spaces/Neuralk-AI/tabbench">TabBench V2</a> benchmark, which evaluates ML algorithms on 189 classification tasks. These classification tasks range from one hundred up to 150k rows, from binary to up to 100 classes, and from 2 to 1777 features. All datasets are IID classification datasets from OpenML. The company behind TabBench V2 is <a href="https://www.neuralk.ai/">NeuralkAI</a>, based in Paris, France.</p><p>TabBench compares three &#8220;clusters&#8221; of ML algorithms: tabular foundation models (TFMs), gradient boosted decision trees (GBDTs) like CatBoost, and tuned neural networks (NNs) like TabM.</p><p>Overall win rate of state-of-the-art (=2nd generation) Tabular Foundation Models: ~83% (caveat: based on accuracy; but TFMs are highly performant under the other metrics like F1 as well). The list of 2nd-generation TFMs includes TabICLv2, TabPFN-3.0, and Seldon.</p><p>Further down on the benchmarks website, there is a quite interesting figure, which drills down performance based on dataset characteristics.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!oYhD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!oYhD!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png 424w, /__u/substackcdn.com/image/fetch/$s_!oYhD!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png 848w, /__u/substackcdn.com/image/fetch/$s_!oYhD!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oYhD!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!oYhD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png" width="958" height="600" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e50e172c-79da-4794-887b-985db0bf8975_958x600.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:600,&quot;width&quot;:958,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:88780,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/203049531?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!oYhD!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png 424w, /__u/substackcdn.com/image/fetch/$s_!oYhD!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png 848w, /__u/substackcdn.com/image/fetch/$s_!oYhD!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png 1272w, /__u/substackcdn.com/image/fetch/$s_!oYhD!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe50e172c-79da-4794-887b-985db0bf8975_958x600.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Win &#8220;clusters&#8221;. Each dot represents a dataset. Image by NeuralkAI: https://huggingface.co/spaces/Neuralk-AI/tabbench</figcaption></figure></div><p>As far as I understand, the axes are the first two PCA components of the datasets &#215; model-metric rank matrix. The x-axis goes from small tasks with mostly numerical features to categorical-feature-heavy, large tasks; the y-axis goes from binary, balanced tasks to multi-class, imbalanced tasks.</p><p>Three distinct clusters emerge, which give us a rough idea of when to expect which ML algorithm to work well.</p><h2>When does which ML algorithm perform well?</h2><p>Based on that same data, the creators of TabBench also provide a decision tree (I assume they trained it on the results), which gives us a good rule of thumb when to expect which ML algorithm to perform the best classifications: </p><ul><li><p>Small to midsize (&lt;44k rows) classification dataset with few categorical features (&lt;15%) &#8594; Use <strong>Tabular Foundation Models </strong>(2nd-gen) </p></li><li><p>Smallish datasets with many categorical features &#8594; Use <strong>boosted trees</strong></p></li><li><p>Larger binary datasets (&gt;44k rows) &#8594; Use <strong>boosted trees</strong></p></li><li><p>Large multi-class or unbalanced dataset &#8594; Use <strong>tuned neural networks</strong></p></li></ul><p>These are, of course, just tendencies and not physical laws. Your task may be balanced, binary classification, and your best model might be a neural network, or even a logistic regression model. By the way, logistic regression was not part of the benchmark, which would have been interesting to see as a baseline. Current baseline is ensembled xgboost. Another caveat: The benchmark only goes up to 150k rows, so beyond that, we can only guess. My guess, though, is that beyond 150k rows, tabular foundation models will not (yet?) take home too many prizes.</p><p>Where does that leave tabular foundation models? This snapshot supports the view that tabular foundation models are state-of-the-art models that stand side-by-side with boosted trees and tuned neural networks. However, the development of TFMs still has a lot of momentum. Go back to December 2025, and the recommendation would have been something like: use tabular foundation models only for very small datasets. Since then, they have been eating into the boosted tree territory improvement by improvement. The question is: how far will the TFM territory expand? I&#8217;m very curious to see how the TFM field will further develop, and, of course, will share with you what I learn along the way.</p><p>In the meantime, <a href="https://huggingface.co/spaces/Neuralk-AI/tabbench">I encourage you to take a look at the TabBench V2 results</a> and play around with the interactive figures; the site contains many more insights than I covered here.</p><p>If you want to learn more about Tabular Foundation Models, have a look at the open book I am working on: <a href="https://tabularfoundationmodels.com/">tabularfoundationmodels.com</a>.</p>]]></content:encoded></item><item><title><![CDATA[What is TabPFN's Thinking mode?]]></title><description><![CDATA[When I read the TabPFN-3 Technical Report, the benchmark for &#8220;TabPFN-3-Thinking&#8221; stood out: it appeared at the top of the TabArena benchmark for large datasets.]]></description><link>https://mindfulmodeler.substack.com/p/what-is-tabpfns-thinking-mode</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/what-is-tabpfns-thinking-mode</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 16 Jun 2026 09:48:37 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!gjiu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When I read the TabPFN-3 Technical Report, the benchmark for &#8220;TabPFN-3-Thinking&#8221; stood out: it appeared at the top of the <a href="/__u/mindfulmodeler.substack.com/p/tabarena-explained">TabArena benchmark</a> for large datasets. TabPFN-3-Thinking also surpassed AutoGluon 1.5 extreme,  an open-source AutoML framework that trains and ensembles many models &#8212; including foundation models like TabPFN, TabICL, and TabDPT. As we know, ensembles are typically the way to go in machine learning if you want to squeeze out the last bits of performance. And yet, the Thinking mode beat it:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!gjiu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!gjiu!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png 424w, /__u/substackcdn.com/image/fetch/$s_!gjiu!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png 848w, /__u/substackcdn.com/image/fetch/$s_!gjiu!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gjiu!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!gjiu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png" width="1456" height="582" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:582,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:157512,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/197350531?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!gjiu!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png 424w, /__u/substackcdn.com/image/fetch/$s_!gjiu!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png 848w, /__u/substackcdn.com/image/fetch/$s_!gjiu!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gjiu!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7023d57a-ef28-4666-84b8-4ccedca2e31f_1600x640.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Image source: TabPFN-3 Technical Report: https://arxiv.org/abs/2605.13986</figcaption></figure></div><p>So, what is this mysterious Thinking mode?</p><h2>Thinking mode is not open source</h2><p>Not much is publicly known about this mysterious Thinking mode. It&#8217;s bundled into &#8220;TabPFN3-Plus,&#8221; the closed-source/commercial part of TabPFN. The technical report has only one paragraph on the Thinking mode and keeps things high-level. In general, Prior Labs has chosen a dual path: an open core with their model architecture being open source plus closed source parts, like the prior and the Thinking mode. This is a similar playbook to what open-core infrastructure companies like Red Hat and GitLab are doing. Give away the core for free, monetize some of the features, especially for enterprise.</p><h2>What we know about Thinking mode</h2><p>We have two sources available to go full detective. Let&#8217;s start with the <a href="https://arxiv.org/pdf/2605.13986">TabPFN-3 Technical Report</a>, which gives us the following hints:</p><ul><li><p>Thinking is &#8220;test-time compute scaling&#8221; of TabPFN.</p></li><li><p>The authors also refer to the mode as applying &#8220;additional inference-time computation.&#8221;</p></li><li><p>It outperforms AutGluon 1.5 extreme in less than a tenth of the runtime.</p></li><li><p>No use of LLMS, real data, internet, or any other model besides TabPFN.</p></li><li><p>The performance gains seem to be mostly for large datasets.</p></li></ul><p>A bit more information is provided in the <a href="https://docs.priorlabs.ai/capabilities/thinking-mode">TabPFN documentation</a>:</p><ul><li><p>The Thinking mode flag is set in <code>.fit(),</code> not in <code>.predict().</code> The docs also state that the Thinking mode &#8220;fuels recurring predictions&#8221;.</p></li></ul><ul><li><p>There is a <code>thinking_effort</code> parameter, which can be set to 'medium' or 'high'. This parameter steers the &#8220;effort &amp; compute&#8221; of the Thinking mode.</p></li><li><p>The <code>thinking_timeout_s</code> parameter sets a time budget.</p></li><li><p>A &#8220;<code>thinking_metric</code>&#8221; parameter controls which target metric to optimize during Thinking mode. Currently supported metrics are: Accuracy, LogLoss, ROC AUC, RMSE, and MAE.</p></li><li><p>There is a separate monthly quota for Thinking mode of 20 fits per month.</p></li></ul><p>Alright, this is all I found. Everything that follows is speculation based solely on the limited publicly available information.</p><h2>What Thinking mode is probably not</h2><p>Even though not much is known, I would exclude some possibilities for Thinking mode based on the available information.</p><ul><li><p><strong>I don&#8217;t think it&#8217;s an ensemble with other models</strong>, since the report states that no other models were used. However, it might be an ensemble of TabPFN models.</p></li><li><p>It&#8217;s probably <strong>not some novel pretraining procedure</strong>, as TabPFN-3 with and without thinking seem to share the same pretraining.</p></li><li><p><strong>No LLM chain-of-thought</strong> involved, as the report states no LLMs involved.</p></li><li><p><strong>I don&#8217;t believe Thinking mode involves gradient-based fine-tuning of TabPFN</strong>, since the metrics the user can pick are not differentiable.</p></li></ul><h2>Let&#8217;s speculate</h2><p>The TabArena results for plain tabular foundation models are quite strong. And yet, they &#8220;only&#8221; do in-context learning, which does not even change the model weights. Besides that, no particular optimization happens in vanilla TFMs. For example, TabPFN is already an ensemble, but the ensembling is agnostic of the underlying task: The variety between ensemble members comes from changing the feature order and the pre-processing pipeline. What I want to say: there is an opportunity to optimize beyond in-context learning.</p><p>In my opinion, there are strong hints that Thinking mode might involve some validation-driven search: there is some computational expense (up to 1/10th of the AutoGluon 1.5 extreme compute time); the budget parameters would be coherent with some type of search with a budget (<code>thinking_effort </code>and<code> thinking_timeout_s</code>); also, the user can pick a metric, which might be the objective of some optimization; It would also make sense given the team&#8217;s AutoML background.</p><p>If we assume that Thinking mode is search with a budget, the big question would be: What is the search space? Does it include feature engineering? Is it about smart context selection? Is it maybe a more task-specific way of constructing a TabPFN ensemble? The gain in performance (in Elo) is quite large, so I&#8217;m assuming that it&#8217;s optimizing multiple things, not just a bit of feature engineering or so. </p><p>Another possibility is that Thinking mode does some type of context distillation or optimization, which might explain the strong gains on large datasets.</p><p>I am also wondering whether the Thinking procedure is model-agnostic. And if it is, would it give the same gains to xgboost and co, or is there something about tabular foundation models that makes the Thinking mode more effective?</p><p>These are my thoughts so far. I know, I am leaving you with more questions than answers &#128517;. What are your guesses?</p>]]></content:encoded></item><item><title><![CDATA[How TabICL and TabPFN handle missing values]]></title><description><![CDATA[You can give training or test data with missing values to the tabular foundation models TabPFN and TabICL, and the prediction will &#8220;just work&#8221;.]]></description><link>https://mindfulmodeler.substack.com/p/how-tabicl-and-tabpfn-handle-missing</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/how-tabicl-and-tabpfn-handle-missing</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 09 Jun 2026 06:40:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WtQm!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>You can give training or test data with missing values to the <a href="/__u/mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird">tabular foundation models</a> TabPFN and TabICL, and the prediction will &#8220;just work&#8221;. But what happens in the background?</p><p>Let&#8217;s find out.</p><h2>How TabICL handles missing values</h2><p>For TabICLv2, we can find that information in the <a href="https://github.com/soda-inria/tabicl">README</a>, section &#8220;Preprocessing&#8221;:</p><blockquote><ul><li><p>Create a separate category for missing values in categorical features</p></li><li><p>Perform mean imputation for missing numerical values (encoded as NaN)</p></li></ul></blockquote><p>This looks like a pre-processing layer that you could attach to any model. Creating a missing value category is pretty standard. Mean imputation for numerical values is something I see often, too, unfortunately (mean imputation loses a lot of information).</p><p>A look into the <a href="https://arxiv.org/html/2602.11139">TabICLv2 paper</a> confirms that missing values handling is indeed just a pre-processing layer:</p><blockquote><p>Adding missing indicators (Le Morvan &amp; Varoquaux, <a href="https://arxiv.org/html/2602.11139v1#bib.bib42">2025</a>) or introducing missingness during pretraining may improve the handling of missing values, which are currently imputed by the mean, but remain unexplored.</p></blockquote><p>My recommendation for TabICL, in this current version: Deal with missing data imputation yourself. Or at least confirm that adding a missing value category for categorical features and mean imputation for numerical features is what you want.</p><h2>How TabPFN handles missing values</h2><p>Let&#8217;s have a look at TabPFNs <a href="https://github.com/PriorLabs/TabPFN#performance--limitations">README</a> to see how it handles missing values:</p><blockquote><p><strong>Q: Can TabPFN handle missing values?</strong> </p><p><strong>Yes!</strong></p></blockquote><p>Ok &#8230; cool, I guess. Not very informative though.</p><p>Let&#8217;s go deeper.  The <a href="https://arxiv.org/abs/2605.13986">technical report for TabPFN-3.0</a> says:</p><blockquote><p>Native missing-value handling. For each cell that is NaN, TabPFN-3 computes a binary indicator and concatenates it with the cell value before embedding. The model therefore receives an explicit signal about missing data and can condition its predictions accordingly, rather than relying on upstream imputation.</p></blockquote><p>This sounds promising. Let&#8217;s go even deeper.</p><p>In the <a href="https://github.com/PriorLabs/TabPFN/blob/main/src/tabpfn/architectures/encoders/steps/nan_handling_encoder_step.py">TabPFN code</a>, we can see that missing values get encoded and passed into the model: The <code>NanHandlingEncoderStep</code> class encodes NaNs as -2.0, Inf as 2.0, and -Inf as 4.0. The features themselves are then mean imputed, but the additional &#8220;missingness&#8221; channel  is concatenated with the original features before the linear embedding. It gets a bit more complicated because features are grouped, but that&#8217;s not important right now. We only need to know that the TabPFN-3.0 model has the information on missingness for each feature available during inference time. </p><p>Having these missingness indicators in the architecture is only half of the story. The other part is pre-training with missing values. Unfortunately, TabPFN&#8217;s pre-training code is not open source, but I guess that TabPFN was pre-trained on datasets with missing values.</p><p>Based on the way that TabPFN&#8217;s architecture encodes missingness, it is well equipped to deal with missing values. <a href="https://arxiv.org/abs/1902.06931">According to this research paper</a>, TabPFN can, in theory, even handle the most difficult type of missingness: missing not at random (MNAR). However, I haven&#8217;t tried it out myself so far and have not seen any benchmarks yet.</p>]]></content:encoded></item><item><title><![CDATA[TabPFN and TabICL are not everything]]></title><description><![CDATA[An overview of other families for tabular foundation models]]></description><link>https://mindfulmodeler.substack.com/p/tabpfn-and-tabicl-are-not-everything</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/tabpfn-and-tabicl-are-not-everything</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 02 Jun 2026 10:17:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!6fms!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>How do you pre-train a tabular model?</p><p>One answer is tabular foundation models like TabPFN and TabICL. They are <a href="/__u/mindfulmodeler.substack.com/p/how-pfns-make-tabular-foundation">prior-data fitted networks</a>, an approach to tabular foundation models. I covered the PFN family extensively in this blog (<a href="/__u/mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird">start here</a>). In short, TabPFN and TabICL are transformer-based networks, pre-trained on millions of synthetic datasets, allowing for in-context learning.</p><p>But what about other approaches to learning from some tables and transferring them to another table?</p><p>Besides PFNs, there are hypernetworks, cross-table transfers, and LLM-based approaches.</p><p>This post zooms out a bit, looking at these other tabular foundation model families.</p><h2>Hypernetworks</h2><p>Hypernetworks are pre-trained to &#8220;predict&#8221; or emit weights of a smaller model, typically a multi-layer perceptron (MLP), which is then used to make the actual prediction. The &#8220;prediction&#8221; step of the hypernetwork, therefore, replaces the training step of the MLP. And the actual prediction for your current data is then just regular inference with an MLP. Hypernetworks are often based on transformers (like <a href="https://arxiv.org/html/2312.08598">MotherNet</a>), but don&#8217;t have to be (for example, <a href="https://arxiv.org/html/2511.15941">iLTM</a>)</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!6fms!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!6fms!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png 424w, /__u/substackcdn.com/image/fetch/$s_!6fms!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png 848w, /__u/substackcdn.com/image/fetch/$s_!6fms!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6fms!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!6fms!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png" width="628" height="437.78846153846155" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1015,&quot;width&quot;:1456,&quot;resizeWidth&quot;:628,&quot;bytes&quot;:190019,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/200083727?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!6fms!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png 424w, /__u/substackcdn.com/image/fetch/$s_!6fms!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png 848w, /__u/substackcdn.com/image/fetch/$s_!6fms!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6fms!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe281ed90-32d8-4349-ac3c-4a64c4b11bbc_1788x1246.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Making predictions with hypernetwork-based TFMs. A) The pre-trained hypernetwork &#8220;predicts&#8221; weights to B) parameterize a multi-layer perceptron, which C) then can predict the data.</figcaption></figure></div><h2>Cross-table transfer plus fine-tuning</h2><p>This family follows the more classic foundation model recipe with self-supervised pre-training across many tables, then fine-tuning on your current data. The pre-training task is usually self-supervised, for example, a masked reconstruction task, where some cells are hidden, and the model predicts them. Approaches in this family have in common that they learn to encode diverse tables into a schema-agnostic representation, and they are often transformer-based.</p><p>During pre-training, each table might get its own input/output adapters (featurizer and head) that are discarded for inference. Instead, to apply the model to new data, you attach a fresh set of adapters and fine-tune.</p><p>The family holds a diverse set of models, differing mostly in how to encode tables into a schema-agnostic representation. For example, some approaches also encode semantic information such as column names (e.g., <a href="https://arxiv.org/html/2402.16785">CARTE</a>), while others don&#8217;t (e.g., <a href="https://arxiv.org/abs/2305.06090">XTab</a>).</p><h2>LLM-based approaches</h2><p>The basic idea: take a table, turn it into text, then fine-tune an LLM in a supervised fashion. This reframes table prediction as a language problem to force it into the LLM scheme. I find that the weirdest family, as LLMs have the wrong inductive biases for tabular data, especially when the tables have lots of numerical values. On top of that, inference cost is relatively hefty. The LLM-based family may be most interesting for  small tables with lots of semantic features, like free text features, categories, and telling feature names. This TFM family is the only one that allows zero-shot predictions, but can also do few-shot learning (aka In-Context Learning). An example is <a href="https://arxiv.org/abs/2210.10723">TabLLM</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!slyq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!slyq!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png 424w, /__u/substackcdn.com/image/fetch/$s_!slyq!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png 848w, /__u/substackcdn.com/image/fetch/$s_!slyq!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png 1272w, /__u/substackcdn.com/image/fetch/$s_!slyq!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!slyq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png" width="1083" height="277" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:277,&quot;width&quot;:1083,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:57952,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/200083727?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!slyq!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png 424w, /__u/substackcdn.com/image/fetch/$s_!slyq!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png 848w, /__u/substackcdn.com/image/fetch/$s_!slyq!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png 1272w, /__u/substackcdn.com/image/fetch/$s_!slyq!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94ffc5f5-2ea3-493b-acae-44cf1c50f51d_1083x277.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Turning a table row into a string that can be used for fine-tuning an LLM.</figcaption></figure></div><h2>Why I focus on PFNs in my book</h2><p>Part of what brought me to write this post is thinking about the scope of my latest book project, <a href="https://tabularfoundationmodels.com/">Tabular Foundation Models</a>.</p><p>The book is focused on just one of the families, the PFNs. So why not cover these other families in the book? Here is my reasoning for why I&#8217;m betting specifically on PFN-based models like TabPFN and TabICL:</p><ul><li><p>PFNs are currently state-of-the-art, even beating boosted tree models.</p></li><li><p>Great open source availability and ecosystem.</p></li><li><p>There is a lot of development happening in the PFN space.</p></li><li><p>Many startups and companies are betting on PFN-based models.</p></li><li><p>PFNs allow for in-context learning, no further training required.</p></li><li><p>No fine-tuning required. But you can fine-tune if you want to.</p></li><li><p>It&#8217;s easy to instill the right inductive biases, as you can include synthetic data in the pre-training.</p></li></ul><p>So, to me, PFN-based approaches like TabPFN and TabICL hold the greatest promise for foundation modeling. That&#8217;s why I&#8217;m betting that they&#8217;ll end up as a new state-of-the-art paradigm in tabular modeling.</p>]]></content:encoded></item><item><title><![CDATA[I’m writing a book on Tabular Foundation Models]]></title><description><![CDATA[tl;dr: In-progress book here: tabularfoundationmodels.com]]></description><link>https://mindfulmodeler.substack.com/p/im-writing-a-book-on-tabular-foundation</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/im-writing-a-book-on-tabular-foundation</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 26 May 2026 14:03:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!y3mi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>tl;dr: In-progress book here: <a href="https://tabularfoundationmodels.com">tabularfoundationmodels.com</a></p><p>In terms of book writing, my last year has been horrible. <a href="/__u/mindfulmodeler.substack.com/p/why-i-quit-writing-building-blocks">I quit writing Building Blocks of ML</a> after months of work, and I put Machine Learning 4 Remote Sensing on ice. My 2025 experience has made me wary of starting new book projects. So I focused on Mindful Modeler, especially the <a href="/__u/mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird">series on tabular foundation models</a>.</p><p>I enjoyed writing the TFM series a lot, and it resonated with many of you. This truly built up my motivation and confidence, and I started writing a book again:</p><p><strong><a href="https://tabularfoundationmodels.com/">Tabular Foundation Models &#8212; A Hands-On Guide to TabPFN, TabICL, and the Tabular Revolution</a></strong></p><p>There is such a huge gap between frontier labs pushing tabular foundation models and everyday modeling practice. Many data scientists I&#8217;ve talked to have barely heard about tabular foundation models, struggle to learn about them, or have had strong concerns with explainability and scaling. There is just so little coverage of foundation models: barely any tutorials, videos, posts, or books.</p><p>So, here is my plan.</p><h2>My plans for the Tabular Foundation Book</h2><p>I want the book to be the best introduction out there on tabular foundation models, covering the three parts: <strong>intuition, theory, and practice</strong>.</p><p>The<strong> book will be open</strong>, meaning a web version for anyone to read in the world, free forever. Just like I did with my book <a href="https://christophm.github.io/interpretable-ml-book/">Interpretable ML</a> and <a href="https://ml-science-book.com/">Supervised ML for Science</a>.</p><p>Model-wise, I will likely <strong>focus on TabICL and TabPFN</strong>. While there are many more models and alternative approaches to the foundation model paradigm, the book&#8217;s priority will be practicality and performance, and not academic coverage. So more of a book for cooks, less for botanists.</p><p>It will be a challenge to keep up with the fast pace of the tabular foundation field. That&#8217;s why I&#8217;ll put extra effort into <strong>making the book as &#8220;timeless&#8221; as possibl</strong>e. I want  the book be worth your time investment.</p><p>The good news is that I am fairly confident that code examples should age well, since TabICL and TabPFN adopted the scikit-learn interface. I&#8217;m planning to put <strong>many practical Python code examples</strong> into the book. From simple classification and regression examples, to time series forecasting, quantile regression, speeding up inference, imbalanced classification, &#8230; you name it.</p><p>I plan to move fast with the book, get an early version v1.0 ready by autumn (fingers crossed), and then keep updating it from there on.</p><p>That&#8217;s my plan. Only one piece of information is missing &#8230;</p><h2>Which cover mascot?</h2><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!y3mi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!y3mi!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!y3mi!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!y3mi!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!y3mi!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!y3mi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png" width="358" height="238.74862637362637" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:358,&quot;bytes&quot;:907072,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/199288553?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!y3mi!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png 424w, /__u/substackcdn.com/image/fetch/$s_!y3mi!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png 848w, /__u/substackcdn.com/image/fetch/$s_!y3mi!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png 1272w, /__u/substackcdn.com/image/fetch/$s_!y3mi!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fec1d5cb1-4679-4ecf-9439-c9cc6815d4b2_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">Which mascot should I pick? Image made by AI.</figcaption></figure></div><p>All my books have had a mascot on the cover: A mole, an octopus, a beaver, meerkats, and a raven. For the current book, I am looking for a cover mascot that captures the idea of foundation models. Some candidates:</p><ul><li><p>A chameleon would represent &#8220;one-shotting&#8221; patterns.</p></li><li><p>A mouse that can live in any environment.</p></li><li><p>Honeybees working on a grid of cells, building a foundation.</p></li></ul><p>If you have any ideas, feel free to leave a comment. The book is still in progress, which means I&#8217;d love to hear your feedback and requests for chapters.</p>]]></content:encoded></item><item><title><![CDATA[Tabular ML is entering a new benchmark era]]></title><description><![CDATA[From static and narrow benchmarks to live, capability-driven evaluation]]></description><link>https://mindfulmodeler.substack.com/p/tabular-ml-is-entering-a-new-benchmark</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/tabular-ml-is-entering-a-new-benchmark</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 19 May 2026 08:18:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WtQm!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>In tabular machine learning, benchmarks are old news. When ML researchers develop a new machine learning algorithm, they pick a set of datasets like <a href="https://openml.github.io/openml-python/main/examples/20_basic/simple_suites_tutorial.html">OpenML-CC18</a> and compare the performance of the new algorithm against the state-of-the-art algorithms.</p><p>But the tabular benchmark situation is changing; a change that goes hand-in-hand with <a href="/__u/mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird">the rise of tabular foundation models</a>.</p><p>Changes are two-fold: Benchmarks are becoming more &#8220;live&#8221; and focused on &#8220;capabilities&#8221; at least from the lens of tabular foundation models.</p><p>Live benchmarks, like <a href="https://huggingface.co/spaces/TabArena/leaderboard">TabArena</a>, are typically very rigorous with a strict protocol for standardized pre-processing and evaluation. But what makes them &#8220;live&#8221; is that they come with a website and active maintenance, reflecting the current state-of-the-art. This is in contrast to static benchmarks, which may be a table in a PDF paper, without updates. If you want to learn more about TabArena, I have a full blog post:</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;eb08afd7-90e6-44dc-8d9d-11263fd67bdf&quot;,&quot;caption&quot;:&quot;Machine learning progresses through benchmarks. While I&#8217;ve been critical before of ML&#8217;s benchmark obsession and danger of getting stuck on benchmarks, benchmarks are essential to guide researchers and practitioners in the right direction. ImageNet, under the leadership of Fei-Fei Li, for example, ushered in the deep learning era. Also today, LLMs are la&#8230;&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;TabArena explained&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:8489879,&quot;name&quot;:&quot;Christoph Molnar&quot;,&quot;bio&quot;:&quot;Writer and machine learning expert&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5dffee84-bdf9-4188-9e5d-b63c519aba2e_2998x2998.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-03-31T12:09:21.288Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!hjDv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://mindfulmodeler.substack.com/p/tabarena-explained&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:192596694,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:15,&quot;comment_count&quot;:4,&quot;publication_id&quot;:1078760,&quot;publication_name&quot;:&quot;Mindful Modeler&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WtQm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>Live benchmarks are quite common in LLM development, where we have <a href="https://www.swebench.com/">SWE-bench</a> for assessing coding, <a href="https://github.com/THUDM/LongBench">LongBench</a> for testing LLMs with long contexts, and <a href="https://lastexam.ai/">Humanity&#8217;s Last Exam</a> with a list of expert-level questions. Each benchmark addresses different &#8220;capabilities&#8221; of large language models.</p><h2>Benchmarking capabilities in tabular ML</h2><p>Tabular is, by nature, a narrow modality compared to language, which is more general-purpose (translation, coding, question-answering, &#8230;). However, even the tabular modality contains many tasks: classification, regression, quantile regression, missing data imputation, time series forecasting, and many, many more. If someone designs a new ML algorithm, they can benchmark it against any of these tasks (if the algorithm is flexible enough).</p><p>The novel appeal with benchmarking tabular foundation models is their even greater flexibility, and that we are not testing an algorithm, but a fixed model.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> Whether you use TabICL for regression, quantile regression, or time series forecasting, it&#8217;s always the same pre-trained model, and due to in-context learning, the weights don&#8217;t change. This parallels LLMs, where we have pre-trained models with in-context learning.  This invites us to reframe tasks as &#8220;capabilities.&#8221;</p><p>I&#8217;m excited about seeing a proliferation of benchmarks to test &#8220;capabilities&#8221; beyond just classification and regression. For example, <a href="https://scoringbench.com/">ScoringBench</a> evaluates ML algorithms and tabular foundation models based on their capability to predict the full predictive distribution.</p><p>While benchmarks have always been a catalyst in machine learning, it feels like the benchmark landscape for tabular ML is changing, due to tabular foundation models. For example, just recently, the <a href="https://arxiv.org/abs/2605.10616">MulTaBench</a> paper was put on arxiv. The benchmark contains 40 multimodal datasets, 20 of which are tabular plus text, and the other 20 are tabular plus image. Exactly the catalyst we need to move forward on multi-modal tabular foundation models.</p><p>A piece of evidence pointing toward the new benchmark situation is the <a href="https://storage.googleapis.com/prior-labs-tabpfn-public/reports/TabPFN_3_model_report.pdf">TabPFN-3.0 model report</a>. Out of the 20 main pages, 9 are &#8220;Experimental Results&#8221;, mostly benchmarks. For example, they test classification, regression, and quantile regression capabilities on <a href="https://huggingface.co/spaces/TabArena/leaderboard">TabArena</a>, and prediction with text columns on <a href="https://arxiv.org/abs/2505.18125">TabStar</a> data, and relational data with <a href="https://relbench.stanford.edu/">RelBenchV1</a>. Not all are &#8220;live&#8221; benchmarks, but they test different capabilities.</p><p>I remain excited about tabular foundation models. The tabular foundation models field already has a strong momentum, and having diverse benchmarks may serve as catalysts. However, there is the risk of overly focusing on benchmarks. Think overfitting and benchmarks-as-marketing, as we are seeing with LLMs. That&#8217;s why it&#8217;s  important that we have many diverse benchmarks from various parties.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Ignoring that classification and regression are usually separate models, e.g., in TabICL and TabPFN</p></div></div>]]></content:encoded></item><item><title><![CDATA[How to make Tabular Foundation Model inference faster]]></title><description><![CDATA[Practical strategies and tradeoffs for reducing prediction time in tabular foundation models]]></description><link>https://mindfulmodeler.substack.com/p/making-tabular-foundation-models</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/making-tabular-foundation-models</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 12 May 2026 09:49:54 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/1dfcc84c-a949-498f-af28-5073b11bca45_1491x1108.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>The greatest bottleneck with tabular foundation models: Prediction, aka inference, is slow.</p><p>This post is a collection of tips and tricks to make tabular foundation models much faster. BUT! There is always a price to pay. And you must decide on that bargain. Some improvements are cheaper, some are more expensive.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!22Nd!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!22Nd!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png 424w, /__u/substackcdn.com/image/fetch/$s_!22Nd!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png 848w, /__u/substackcdn.com/image/fetch/$s_!22Nd!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png 1272w, /__u/substackcdn.com/image/fetch/$s_!22Nd!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!22Nd!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png" width="564" height="461.3489010989011" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1191,&quot;width&quot;:1456,&quot;resizeWidth&quot;:564,&quot;bytes&quot;:221346,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/196631466?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!22Nd!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png 424w, /__u/substackcdn.com/image/fetch/$s_!22Nd!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png 848w, /__u/substackcdn.com/image/fetch/$s_!22Nd!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png 1272w, /__u/substackcdn.com/image/fetch/$s_!22Nd!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8d5cc38b-9b92-487b-8fe9-b9fd843968b9_1491x1220.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Quick overview of TFM inference optimization strategies</figcaption></figure></div><p>If you haven&#8217;t heard about tabular foundation models, check out my series on TFMs:</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;eabb09f3-592a-42e1-a778-d662e508dc4c&quot;,&quot;caption&quot;:&quot;Tree-based boosting algorithms have sat on the tabular throne for many years now. Many times, the deep learners have attempted to dethrone XGBoost, CatBoost, and other tree-based algorithms, but without success.&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The rise of tabular foundation models&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:8489879,&quot;name&quot;:&quot;Christoph Molnar&quot;,&quot;bio&quot;:&quot;Writer and machine learning expert&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5dffee84-bdf9-4188-9e5d-b63c519aba2e_2998x2998.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-01-13T13:14:41.426Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!Sv6x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F89dd7432-4067-4499-b283-e1a13335d5b3_1797x920.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:184297088,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:163,&quot;comment_count&quot;:27,&quot;publication_id&quot;:1078760,&quot;publication_name&quot;:&quot;Mindful Modeler&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WtQm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>Let&#8217;s start with a &#8220;cheap&#8221; option for improving the prediction time of tabular foundation models. </p><h2>Pick a faster foundation model (currently TabICL)</h2><p>There has been a flurry of new models in the TFM space. Not all have the same inference time. It may be worth comparing the speed of multiple TFMs.</p><p>Which one is the fastest TFM? This is subject to rapid changes, but here is a quick first idea, based on TabArena. If you don&#8217;t know TabArena, check out my post.</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;8bb9a07b-dc76-4e78-8527-55fd5d953966&quot;,&quot;caption&quot;:&quot;Machine learning progresses through benchmarks. While I&#8217;ve been critical before of ML&#8217;s benchmark obsession and danger of getting stuck on benchmarks, benchmarks are essential to guide researchers and practitioners in the right direction. ImageNet, under the leadership of Fei-Fei Li, for example, ushered in the deep learning era. Also today, LLMs are la&#8230;&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;TabArena explained&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:8489879,&quot;name&quot;:&quot;Christoph Molnar&quot;,&quot;bio&quot;:&quot;Writer and machine learning expert&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5dffee84-bdf9-4188-9e5d-b63c519aba2e_2998x2998.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-03-31T12:09:21.288Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!hjDv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://mindfulmodeler.substack.com/p/tabarena-explained&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:192596694,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:15,&quot;comment_count&quot;:4,&quot;publication_id&quot;:1078760,&quot;publication_name&quot;:&quot;Mindful Modeler&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WtQm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>The following leaderboard excerpt shows that TabICLv2 is, on average, the fastest TFM, but not the most performant. The best-performing model is TabPFN-2.6, with a slightly better Elo than TabICLv2.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!mizJ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!mizJ!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png 424w, /__u/substackcdn.com/image/fetch/$s_!mizJ!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png 848w, /__u/substackcdn.com/image/fetch/$s_!mizJ!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mizJ!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!mizJ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png" width="680" height="365.3191489361702" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:505,&quot;width&quot;:940,&quot;resizeWidth&quot;:680,&quot;bytes&quot;:100679,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/196631466?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!mizJ!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png 424w, /__u/substackcdn.com/image/fetch/$s_!mizJ!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png 848w, /__u/substackcdn.com/image/fetch/$s_!mizJ!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png 1272w, /__u/substackcdn.com/image/fetch/$s_!mizJ!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1159d8b7-1374-411c-ac55-72aa41e44ae0_940x505.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Screenshot from TabArena on May 5th, 2026. Source: <a href="https://huggingface.co/spaces/TabArena/leaderboard">https://huggingface.co/spaces/TabArena/leaderboard</a></figcaption></figure></div><p>However, results from TabArena are averages (or medians) and therefore just rough pointers. You won&#8217;t know which TFM performs best on your task until you try.</p><p>I compared the prediction time of TabPFN and TabICL on the <a href="https://archive.ics.uci.edu/dataset/14/breast+cancer">UCI breast cancer dataset</a>: Switching from TabPFN to TabICL reduces inference time by ~14% (TabICL: 0.42 seconds; TabPFN: 0.49 seconds). TabICL (acc: 0.98, logloss: 0.067) even slightly outperforms TabPFN (acc: 0.97, logloss: 0.087) on this data.</p><h2>Use a GPU</h2><p>Tabular foundation models are neural-network-based and are much faster on a GPU than on a CPU. If you are GPU-poor like me, this sucks, as it requires getting access to a GPU. With boosted trees and other classic machine learning algorithms, I was able to run most projects on my good old MacBook M1. With tabular foundation models, a GPU is much, much more performant. Switching from CPU to GPU is one of the best levers to improve prediction time.</p><p>To give you an idea: Classifying the breast cancer dataset on a CPU with TabICL (on Google Colab) took ~15 seconds. Switching to GPU reduced that time to ~0.5 seconds, which is a 30x improvement.</p><h2>Cache training representations</h2><p>For each prediction call, tabular foundation models push the entire training data through the network. Let&#8217;s say you have two test sets for which you make predictions in two separate .predict() calls. For each of these calls, most computations are identical. That&#8217;s because the most compute-intensive parts are the transformer modules in the neural network. Here, training data can attend to training data, and test data can attend to training data. That first part, training-attends-training, is the same for each .predict() call, regardless of what the test data looks like.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!gFm7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 424w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 848w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!gFm7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png" width="326" height="243.1119221411192" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:613,&quot;width&quot;:822,&quot;resizeWidth&quot;:326,&quot;bytes&quot;:90159,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:&quot;&quot;,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/185159883?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 424w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 848w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 1456w" sizes="100vw" loading="lazy" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Training data attends training data. Test data also attends training, but never other test data.</figcaption></figure></div><p>A clear case for caching, right?</p><p>Super easy to do, here with TabICL:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:null}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">
clf = TabICLClassifier(kv_cache=True)
clf.fit(X_train, y_train)  
clf.predict(X_test)</code></pre></div><p>With caching, the .fit() call stores the training data computations in a key-value store. The .predict() step can now run faster since it only needs to make computations for the test data and can look up all the attention values for the training data.</p><p>For the breast cancer data, this reduces the .predict() time from 0.5 seconds to 0.09 seconds for TabICL. </p><p>But there is a catch. Two, actually.</p><p>First, caching costs memory. A lot of memory. According to the <a href="https://arxiv.org/html/2511.08667v2">TabPFN 2.5 paper</a>, 6.1 KB of GPU memory and 48.8 KB of CPU memory per cell of the training data. So, in our case, for a small dataset of 455 rows and 31 columns, which means 14105 cells, we already have 86 MB of GPU memory and 688 MB (~0.67 GB) of CPU memory. And that&#8217;s for a small dataset. That makes caching more realistic for small data or for reduced contexts. For TabICL, I haven&#8217;t found the memory numbers, but I suspect they will be in the same ballpark.</p><p>The second catch: While .predict() is faster, .fit() is accordingly slower. So if you do a single train/test split and call .predict() exactly once with the same context data, the KV caching will be useless. Use caching only with repeated .predict() calls and when you actually do have the memory.</p><h2>Context Optimization</h2><p>Models like TabICL scale with <code>O(n^2m + nm^2)</code>. Assuming the number of features <code>m</code> is much smaller than the total number of rows <code>n</code>, and context data dominating <code>n</code>, then halving the context data size may almost mean a 4x reduction in compute time. Any strategy that reduces context size substantially improves speed.</p><p>Reducing the context size usually means lower predictive performance. The following chart shows how simple subsampling with different sample sizes of the context data affects predictive performance, in this case for the <a href="https://archive.ics.uci.edu/dataset/186/wine+quality">UCI wine quality dataset</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ulhx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ulhx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png" width="478" height="316.5135135135135" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:784,&quot;width&quot;:1184,&quot;resizeWidth&quot;:478,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>There are also smarter ways to reduce context size, such as a k-nearest neighbor-based approach and data/distribution distillation approaches.</p><p>More details on context optimization in this post:</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;258d35d7-3457-4cf6-ba9a-b31006e7c562&quot;,&quot;caption&quot;:&quot;Tabular foundation models such as TabPFN and TabICL don&#8217;t need to be trained to perform regression or classification. What they do is called in-context learning. What used to be the training data now becomes the context data at prediction time.&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Context is the new training&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:8489879,&quot;name&quot;:&quot;Christoph Molnar&quot;,&quot;bio&quot;:&quot;Writer and machine learning expert&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5dffee84-bdf9-4188-9e5d-b63c519aba2e_2998x2998.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-04-21T11:50:48.409Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!Ulhx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://mindfulmodeler.substack.com/p/context-is-the-new-training&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:194499979,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:35,&quot;comment_count&quot;:1,&quot;publication_id&quot;:1078760,&quot;publication_name&quot;:&quot;Mindful Modeler&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WtQm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>Since TabICL&#8217;s runtime of <code>O(n^2m + nm^2) </code>is also quadratic in terms of the features, reducing the number of features is also an option, especially if you start out with lots of features. To reduce the number of features, we have the entire toolbox available, from dimensionality reduction to feature selection.</p><h2>Reduce the ensemble size</h2><p>When you call tfm.predict(), you get the result from multiple TFM calls with slightly different parameters. This has to do with the dependence of TFMs on the column order. </p><p>The default ensemble size is 8 (coming down from a whopping 32 in the first version of TabICL). Hiding in there is therefore an up to 8x speedup. The speedup is usually at the cost of reduced performance, but not always. In the case of the breast cancer dataset, we get with TabICL:</p><ul><li><p>8 estimators: 0.42s and a logloss of 0.0671</p></li><li><p>4 estimators: 0.22s and a  logloss of 0.0659</p></li><li><p>2 estimators: 0.11s and a logloss of 0.0760</p></li><li><p>1 estimator: 0.06s and a logloss of 0.0717</p></li></ul><p>In this case, accuracy even stays the same between 8 and 1 estimators. Here, reducing the ensemble to 1 estimator is justified and gives us a huge speed improvement. For your own data, I recommend verifying whether an ensemble reduction brings speed improvements.</p><h2>Train a surrogate model</h2><p>You may distill a TFM into a surrogate model, like a tree ensemble or a multi-layer perceptron. The distillation approach can make sense if you do a lot of inference with fixed context/training data or run it on constrained devices. But you lose all the benefits of using TFMs in the first place, especially the flexibility in context used, embeddings, and so on. Especially if TFMs move into a multi-modal and relational direction, which I strongly suspect, this line will be much less relevant.</p><h2>Key speed improvements for the impatient</h2><ul><li><p>Use a GPU.</p></li><li><p>Switch to a faster TFM.</p></li><li><p>If memory allows, use KV caching for repeated predictions with the same context.</p></li><li><p>Let optimizations compound: For example, GPU inference plus KV caching can improve prediction time by well over 100x.</p></li></ul><p>Also: Scaling is a high priority for most TFM labs. I expect further improvements in architecture and inference tricks, which will bring down the .predict() time.</p>]]></content:encoded></item><item><title><![CDATA[Time series forecasting with tabular foundation models]]></title><description><![CDATA[This post is a quick primer on using tabular foundation models for time series forecasting.]]></description><link>https://mindfulmodeler.substack.com/p/time-series-forecasting-with-tabular</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/time-series-forecasting-with-tabular</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 05 May 2026 09:04:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!V5E8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This post is a quick primer on using tabular foundation models for time series forecasting.</p><p>In my post titled <a href="/__u/mindfulmodeler.substack.com/p/im-betting-on-tabular-foundation">I&#8217;m betting on tabular foundation models</a>, I argued that TFMs will work with all kinds of supervised ML tasks because you can pre-train for any task for which you can construct a data generator. This includes the task of time series forecasting.</p><p>Besides specific pre-training, there is another path: Reframe the forecasting task as a regression task and throw a tabular foundation model at it.</p><p>This post is about the reframing + TFM approach. This approach is called TabPFN-TS in <a href="https://arxiv.org/html/2501.02945v4">this paper for TabPFN</a>, but you can replace TabPFN with other TFMs, such as TabICL or TabDPT.</p><p>The idea is simple: Take a time series, automatically enhance it with temporal features based on the timestamp, then one-shot it with a TFM regressor.</p><p>The input can be as simple as a target plus a timestamp:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!6ypO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!6ypO!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png 424w, /__u/substackcdn.com/image/fetch/$s_!6ypO!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png 848w, /__u/substackcdn.com/image/fetch/$s_!6ypO!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6ypO!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!6ypO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png" width="264" height="248.69565217391303" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17ef567c-36ba-4100-8615-af5b4c181675_414x390.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:390,&quot;width&quot;:414,&quot;resizeWidth&quot;:264,&quot;bytes&quot;:33947,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/195961023?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!6ypO!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png 424w, /__u/substackcdn.com/image/fetch/$s_!6ypO!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png 848w, /__u/substackcdn.com/image/fetch/$s_!6ypO!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png 1272w, /__u/substackcdn.com/image/fetch/$s_!6ypO!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17ef567c-36ba-4100-8615-af5b4c181675_414x390.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Head of table with timestamp and target column. Source: <a href="https://tabicl.readthedocs.io/en/latest/tutorials/time_series_forecasting.html">TabICL Tutorial.</a></figcaption></figure></div><p>For this simple, univariate time series with equidistant time steps, we only need to provide the context data and the number of time steps to get a prediction, here with TabICL:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;4463fe15-0fa5-4544-aae1-939e537c89de&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">model = TabICLForecaster()
pred = model.predict_df(context_df, prediction_length=10)</code></pre></div><p>Time-based features are automatically added. These include an index that simply counts through the timestamps, cyclical (sin/cos) calendar features for day of year, day of week, and a few more features.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!V5E8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!V5E8!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png 424w, /__u/substackcdn.com/image/fetch/$s_!V5E8!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png 848w, /__u/substackcdn.com/image/fetch/$s_!V5E8!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png 1272w, /__u/substackcdn.com/image/fetch/$s_!V5E8!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!V5E8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png" width="1000" height="300" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b0c3443c-2245-4268-baf5-616381849210_1000x300.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:300,&quot;width&quot;:1000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Item ID: 0&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Item ID: 0" title="Item ID: 0" srcset="/__u/substackcdn.com/image/fetch/$s_!V5E8!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png 424w, /__u/substackcdn.com/image/fetch/$s_!V5E8!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png 848w, /__u/substackcdn.com/image/fetch/$s_!V5E8!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png 1272w, /__u/substackcdn.com/image/fetch/$s_!V5E8!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb0c3443c-2245-4268-baf5-616381849210_1000x300.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Forecasting a univariate time series with tabicl. Image source: <a href="https://tabicl.readthedocs.io/en/latest/tutorials/time_series_forecasting.html">TabICL tutorial.</a></figcaption></figure></div><p>Instead of using a vanilla tabular foundation model, you can also use a specific time series foundation model, such as TiRex, Toto, or Moirai-2.0. Based on the <a href="https://arxiv.org/html/2501.02945v4">paper</a>, the time series foundation models outperform tabular foundation models on univariate time series tasks. However, the leaderboard flips when introducing other features: TabPFN-TS outperforms the time series-specific models for covariate-informed forecasts. </p><h2>So what?</h2><p>Let&#8217;s disentangle this:</p><ul><li><p>Even if not pre-trained for forecasting, tabular foundation models like TabPFN work well as a backbone for time series forecasting.</p></li><li><p>I found it surprising that tabular foundation models outperformed time series foundation models in multivariate settings.</p></li><li><p>The wrapper that turns TabPFN into TabPFN-TS is not unique to tabular foundation models. At least to my understanding, it should also work with other &#8220;traditional&#8221; ML models.</p></li><li><p>However, providing such a wrapper for tabular foundation models is very on-brand, as it further pushes TFM down the path of becoming a batteries-included, does-a-lot-of-things-for-you-by-default type of model.</p></li></ul>]]></content:encoded></item><item><title><![CDATA[Train or Test for Feature Effect Estimation? We Finally Have an Answer]]></title><description><![CDATA[A guest post by XAI researcher Timo Hei&#223;]]></description><link>https://mindfulmodeler.substack.com/p/train-or-test-for-feature-effect</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/train-or-test-for-feature-effect</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 28 Apr 2026 09:11:43 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NrCF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is a guest post by the Explainable AI researcher <a href="https://www.linkedin.com/in/timo-heiss/">Timo Hei&#223;</a> (LMU Munich, Germany).</em></p><p>We use XAI methods to trust our models - but can we trust XAI itself? Ironically, XAI comes with its own trust problem: these methods are approximations, with their own errors and pitfalls that are easy to overlook. Take the choice of dataset: <em>should you compute your explanations on training or holdout (test) data?</em> For loss-based feature importance methods like permutation feature importance, the answer is clear: use holdout data. Using training data will give you optimistically biased importance scores if the model overfits &#8212; a common pitfall.</p><p>I have asked myself the same question for years when it comes to feature effects like partial dependence plots (PDP) or accumulated local effects (ALE). Unfortunately, there are no clear recommendations on which dataset to use. Practitioners disagree, software defaults vary, and many researchers quietly sidestep the question. In this post, I want to shed light on the errors lurking in feature effects, and finally (!) give concrete recommendations on which dataset to use when computing them.</p><h2>Feature effects: a quick refresher</h2><p>Feature effects describe how changes in an input variable influence a model&#8217;s predictions. Take a concrete example: <em>how is the predicted hourly power consumption in a city affected by the hour of the day?</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!NrCF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!NrCF!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png 424w, /__u/substackcdn.com/image/fetch/$s_!NrCF!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png 848w, /__u/substackcdn.com/image/fetch/$s_!NrCF!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png 1272w, /__u/substackcdn.com/image/fetch/$s_!NrCF!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!NrCF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png" width="1456" height="398" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:398,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!NrCF!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png 424w, /__u/substackcdn.com/image/fetch/$s_!NrCF!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png 848w, /__u/substackcdn.com/image/fetch/$s_!NrCF!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png 1272w, /__u/substackcdn.com/image/fetch/$s_!NrCF!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F267210ea-da1c-4986-914e-30f62e0d229b_2048x560.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em>Figure 1: Effects of the hour of the day on the power consumption prediction.</em></p><p>PDP and ALE give you exactly that: in our example (Figure 1), power consumption is lowest during night hours and peaks in the evening when people return from work.</p><p>PDPs work by marginalizing over all other features: for a given hour, you average predictions across all observed combinations of the remaining features. While simple and intuitive, this includes unrealistic combinations when features are correlated. ALE plots address this by working locally: they partition the feature into intervals, compute how predictions change within each interval, and accumulate those local effects. However, both are estimates, and estimates have errors.</p><h2>Feature effects also have bias &amp; variance</h2><p>You&#8217;re probably familiar with the bias-variance tradeoff for models: a model&#8217;s mean-squared error (MSE) decomposes into a systematic bias and a variance component.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!_Lih!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!_Lih!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png 424w, /__u/substackcdn.com/image/fetch/$s_!_Lih!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png 848w, /__u/substackcdn.com/image/fetch/$s_!_Lih!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_Lih!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!_Lih!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png" width="669" height="380.4478021978022" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:828,&quot;width&quot;:1456,&quot;resizeWidth&quot;:669,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!_Lih!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png 424w, /__u/substackcdn.com/image/fetch/$s_!_Lih!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png 848w, /__u/substackcdn.com/image/fetch/$s_!_Lih!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png 1272w, /__u/substackcdn.com/image/fetch/$s_!_Lih!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F00fa1b4b-71ad-4c55-b461-1887cb7b808c_2048x1165.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: center;"><em>Figure 2: Bias (systematic deviation from the true value) vs. variance (random variability around it)</em></p><p>The same logic applies to feature effects: we can decompose a feature effect&#8217;s pointwise deviation from the &#8220;true underlying effect&#8221; in the data into now 4 components: bias and variance inherited from the model, and bias and variance introduced by the feature effect computation itself.</p><h2>Model errors propagate to feature effects</h2><p>Naturally, the errors of your machine learning model flow into your feature effects. If your model is biased (e.g., <em>the complexity of your chosen model does not suffice to represent the highly nonlinear relationships in power consumption</em>), this leads to a bias in your feature effect (e.g., <em>your PDP of the feature &#8220;Hour&#8221; is systematically off</em>).</p><p>Similarly, if your model has high variance, as can happen when you overfit to the training data, it affects the variance of your feature effects (e.g., <em>you might get a very different PDP depending on the random seed used to train the model</em>). We&#8217;ll come back to what you can do about this in a minute.</p><h2>An additional bias on training data?</h2><p>There is another bias component that now depends on how you compute the feature effects - and here is where the training-vs-holdout question becomes relevant. In theory, this bias component vanishes only when we compute feature effects on holdout data. However, we conducted a large simulation study, and the results indicated that the bias arising from training data is practically negligible. So, using training data seems to be safe, even when models overfit!</p><h2>Size matters&#8230;</h2><p>Similarly, there is another variance component, stemming from the fact that we compute feature effects on finite samples of data. Interestingly, this variance is connected to feature interactions learned by the model. For models without interactions (standard linear models, GAMs), it is zero. However, most machine learning models learn complex interaction patterns. In these cases, the variance decreases with increasing dataset size.</p><p>For PDPs, the scaling factor is 1/n. For ALE, it is K/n, where K is the number of intervals. This leads to two conclusions: (1) Dataset size matters, and since the training set is usually larger, we should prefer it for feature effect computation. (2) ALE is particularly sensitive to the dataset size, especially with fine-grained intervals.</p><h2>Overfitting? Cross-validation helps</h2><p>Earlier, I flagged that model variance from overfitting reflects in feature effect variance. So, what can you do if you suspect your model is overfitting?</p><p>The fix is cross-validation: fit multiple models on different folds, average the resulting feature effects, and you reduce model variance! At the same time, it preserves the effective sample size of training-data estimation (at the cost of extra computation).</p><h2>Takeaways</h2><p><strong>1&#65039;&#8419;</strong> <strong>XAI methods have errors</strong> - for feature effects, we can decompose them into bias and variance, stemming from the model error and feature effect computation itself.</p><p>2&#65039;&#8419; <strong>Training data bias in feature effects appears negligible in practice</strong> - your feature effects won&#8217;t be meaningfully biased by using training instead of holdout data, even if your model overfits.</p><p>3&#65039;&#8419; <strong>The training set is often preferable for feature effect computation</strong> - larger sample sizes considerably improve feature effects, particularly for ALE!</p><p>4&#65039;&#8419; <strong>Overfitting models? Use cross-validation for feature effects </strong>- averaging feature effects across CV folds reduces model variance without sacrificing sample size.</p><p>We started with an uncomfortable truth: the tools we use to trust our models are themselves imperfect. Knowing where those imperfections lie doesn&#8217;t undermine the value of feature effects; it makes you a more honest user of them. And in this case, it comes with a practical finding: use your training data to estimate your feature effects - or cross-validation if your model overfits!</p><p><em>If you are interested in the details and theory: <a href="https://arxiv.org/pdf/2603.15057">https://arxiv.org/pdf/2603.15057</a></em></p>]]></content:encoded></item><item><title><![CDATA[Context is the new training]]></title><description><![CDATA[Tabular foundation models replace retraining with editable inference-time data]]></description><link>https://mindfulmodeler.substack.com/p/context-is-the-new-training</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/context-is-the-new-training</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 21 Apr 2026 11:50:48 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Ulhx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Tabular foundation models such as TabPFN and TabICL don&#8217;t need to be trained to perform regression or classification. What they do is called in-context learning. What used to be the training data now becomes the context data at prediction time.</p><p>This post explores the idea of context data and contrasts it with &#8220;classic&#8221; training data. Does moving from training data to context data change how we model? Does it enable something new?</p><p>Let&#8217;s dive in.</p><h1>Training versus context</h1><p>For traditional machine learning (linear regression, XGBoost, SVM), the training data shapes the model. Especially for trees, this is vivid: change the training data, and you might get out a differently shaped tree. Fit a linear regression model, and the weights (coefficients) become a function of the training data.</p><p>Not with tabular foundation models. Pre-trained. No classic training step. Prediction happens via in-context learning. When predicting with a TFM, you have to provide both the &#8220;training&#8221; (aka context) and the test data. Through multiple steps, the table cells are embedded, and through attention mechanisms and fully connected layers, the TFM enriches the cell representations by attending to other cells and computing stuff. All based on what the model learned in pre-training. Training and test data points can attend to the training data points, but not to other test data. The cell embeddings for y_test are used to predict or classify. If you want to learn more about how tabular foundation models work, check out my T<a href="/__u/mindfulmodeler.substack.com/p/the-architecture-behind-tabpfn">abPFN architecture post</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!gFm7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 424w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 848w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!gFm7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png" width="518" height="386.29440389294405" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:613,&quot;width&quot;:822,&quot;resizeWidth&quot;:518,&quot;bytes&quot;:90159,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/185159883?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="/__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 424w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 848w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 1272w, /__u/substackcdn.com/image/fetch/$s_!gFm7!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6cc90a13-2116-4f37-bec7-2f315fbf588b_822x613.png 1456w" sizes="100vw" loading="lazy" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Datapoint attention. Cells may attend to other cells in the same column. However, they may only attend to cells from the training data, not from the test data.</figcaption></figure></div><p>So what happens when we change the context data? The model weights don&#8217;t change. The only things that change are the predictions that come out of the model: because the context changed, the embeddings will change, and the set of points that can be attended to will change.</p><p>But besides not affecting the model, is it really worth thinking about the &#8220;training&#8221; data as context data?</p><p>What I tell you in the next two sections is not new: You could also do traditional ML and call it context instead of training. Wrap it all in a function where the .predict() actually does fit+predict. Except, there are things that make it worthwhile to think of context data instead of training data: the ability and necessity to work with smaller datasets (which TFMs work well on) and the inverted cost in which training cost effectively disappears (ignoring pre-training here) and prediction becomes the expensive step.</p><h1>Smaller and smarter contexts</h1><p>Making predictions with tabular foundation models is relatively expensive. Especially if you always use all your training data as context. The inference runtime complexity of TabICL is O(n^2 + nm^2), where n is the number of rows of context+test and m is the number of columns. Since it&#8217;s quadratic in the number of rows, it will absolutely explode when we move up the orders of magnitude: a 10x increase in training data rows is a 100x increase in runtime (assuming that training data size is much larger than test data size and much larger than the number of features). The quadratic scaling strongly incentivizes using a smaller context dataset.</p><p>There is a more positive outlook on small contexts: Tabular foundation models work especially well for smaller data. That&#8217;s where they typically outperform most other approaches. Perhaps it&#8217;s because, for smaller data, the inductive biases learned in pre-training can shine.</p><p>Anyways, we can think of downsizing our data. When I experimented with tabular foundation models I sometimes downsampled the data for improved runtime, but with the drawback of reduced predictive performance.</p><p>For example, the following figure shows the mean absolute error for predicting wine quality, comparing TabICL and a random forest. The TabICL model has the same performance with 700 context data points as the random forest with 1200 training data points. However, increasing the context data for TabICL to the full 1200 data points gives us a big boost in MAE. 1200 data points is easily handled by a TFM nowadays. The question is whether, for example, downsampling from 1 million to 10k data points shows a relevant loss in predictive performance or not.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Ulhx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Ulhx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png" width="582" height="385.3783783783784" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:784,&quot;width&quot;:1184,&quot;resizeWidth&quot;:582,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 424w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 848w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Ulhx!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1211af11-578b-48a9-bd81-8b452fb79103_1184x784.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We can also be a bit smarter about downsizing the context:</p><ul><li><p><a href="http://ttps://proceedings.neurips.cc/paper_files/paper/2024/file/c40daf14d7a6469e65116507c21faeb7-Paper-Conference.pdf">This paper</a> suggests using a kNN model for each test data point to decide on the context. An interesting quote from the paper: &#8220;We thus believe that using nearby points as context is a good inductive bias for tabular data classification.&#8221;</p></li><li><p><a href="https://arxiv.org/abs/2405.16156">Another paper</a> suggests clustering the training data. For a data point in the test data, check which cluster it belongs to, and then use the respective training cluster as context. </p></li><li><p>You can reduce the training data to representative data points.</p></li></ul><p>The following figure shows the cluster approach: Using 2 or 3 clusters is actually slightly more performant than using all data.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Nugk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Nugk!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png 424w, /__u/substackcdn.com/image/fetch/$s_!Nugk!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png 848w, /__u/substackcdn.com/image/fetch/$s_!Nugk!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Nugk!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Nugk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png" width="588" height="348.9756097560976" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:584,&quot;width&quot;:984,&quot;resizeWidth&quot;:588,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Nugk!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png 424w, /__u/substackcdn.com/image/fetch/$s_!Nugk!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png 848w, /__u/substackcdn.com/image/fetch/$s_!Nugk!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Nugk!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3692a6a8-d893-470b-ac90-9e36143d47dc_984x584.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Again, all of these approaches are possible with the classic training-test scheme, but with tabular foundation models, we are more incentivized to reduce the context size while at the same time the TFM models work quite well on smaller data.</p><h1>Context change is the new re-training</h1><p>With in-context learning, the cost inverts between training and inference: training costs close to nothing while inference becomes more expensive.</p><p>This has implications for topics such as interpretability, feature selection, and robustness analysis. For example, in model-agnostic model interpretability, we have methods that rely on shuffling a feature column, but there is also a version of that which relies on removing that column and re-training the model. I&#8217;ve covered this in the following post:</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;772521a0-1057-4a80-8468-9787ee77cb03&quot;,&quot;caption&quot;:&quot;Before &#8220;betting&#8221; on tabular foundation models, I made a similar bet on model-agnostic interpretable machine learning: I wrote two books and did a PhD. Model-agnostic interpretability fascinated me because of its potential &#8220;timelessness&#8221;. Since they work by studying the model predictions when manipulating the inputs, these methods should have a longer sh&#8230;&quot;,&quot;cta&quot;:&quot;Read full story&quot;,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;The interpretability tax on tabular foundation models&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:8489879,&quot;name&quot;:&quot;Christoph Molnar&quot;,&quot;bio&quot;:&quot;Writer and machine learning expert&quot;,&quot;photo_url&quot;:&quot;https://substackcdn.com/image/fetch/f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fbucketeer-e05bbc84-baa3-437e-9518-adb32be77984.s3.amazonaws.com%2Fpublic%2Fimages%2F95da3743-656a-41a3-9626-cdcbf28a88e6_2268x3180.jpeg&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-03-24T13:20:09.736Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!0FrG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://mindfulmodeler.substack.com/p/tabular-foundation-models-break-the&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:191965788,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:20,&quot;comment_count&quot;:9,&quot;publication_id&quot;:1078760,&quot;publication_name&quot;:&quot;Mindful Modeler&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!WtQm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p>In general, methods that require re-training are now much cheaper, relative to methods that require predicting twice with the same model. Usually, re-training was expensive. But for TFMs, re-training boils down to adapting the context data. The cost of re-training is close to zero, we only pay for making the predictions with TFMs.</p><h1>Context engineering for TFMs?</h1><p>So far, we have covered needs of smaller contexts for better scaling and changed costs for some post-hoc methods. Doesn&#8217;t feel like a big shift away from classic training+predict paradigm. I&#8217;m still trying to wrap my head around the context-data-framing. We gain a lot of flexibility without having to think about hyperparameter tuning, model selection, cross-validation, and so on. Just some random thoughts on what a context-paradigm might enable or simplify:</p><ul><li><p>Imagine a user requests their data to be fully deleted. Data your model was trained on. Removing them from a traditional ML model would require re-training, but for TFMs, it&#8217;s just a matter of removing their data from the context.</p></li><li><p>Imagine a forecasting task for which only the last 30 days are relevant, because data distribution is drifting. No problem with TFMs, we can just have a sliding context window.</p></li><li><p>Context data makes it easier to have per-request contexts. For example, contexts only with data from a certain region or customer segment, or even per user.</p></li><li><p>If you have to remove poisoned/wrong/problematic data, just delete the rows from the context.</p></li><li><p>Specialization is much easier. Imagine you have a global model that classifies transactions in your banking app. Then the user classifies some items on their own. These can be simply added to the context.</p></li><li><p>Customization is also simplified. Continuing with the transactions example, the user could say: Don&#8217;t learn from old transactions (= remove rows from context), or don&#8217;t use the merchant name as a feature (= remove column from context).</p></li></ul><p>None of this is truly new. But still, I found it very refreshing to think of context data rather than training data. And I feel like I have yet to fully grasp the context-mindset.</p>]]></content:encoded></item><item><title><![CDATA[Regression should predict full distributions]]></title><description><![CDATA[A "hidden" feature of tabular foundation models]]></description><link>https://mindfulmodeler.substack.com/p/regression-should-predict-full-distributions</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/regression-should-predict-full-distributions</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 14 Apr 2026 15:12:41 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Gx8X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When I took on a <a href="/__u/mindfulmodeler.substack.com/p/how-to-win-an-ml-competition-beyond?utm_source=publication-search">side quest to forecast water supply</a>, I had to predict the 10%, 50%, and 90% quantiles. I solved this by training three separate quantile models (ensembles of xgboost, actually). I would have preferred to only train a single model that could predict all three quantiles at once. While approaches like linear regression can output full predictive distributions, these often come with (too) strong distributional assumptions.</p><p>What if we always worked with machine learning models that produce the full predictive distribution? With classification, we are already at this point: Modern machine learning approaches output not just the majority class, but a probability for each class. Whether this probability is calibrated is another question. </p><p>With regression, we are a bit stuck with a point-based mindset. However, this could change with tabular foundation models. At least in theory: While these models produce the full predictive distribution (or at least a discretized approximation over a fixed support) it&#8217;s not the default and the output is a bit hidden. </p><h1>Tabular foundation models are pre-trained to output the full predictive distribution</h1><p>TabICL and TabPFN predict the full predictive distribution (discretized approximation). But by default, they only give you the conditional mean. Consider the following code snippet:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;c27fb2b7-8526-40f0-bcb7-ec7d61c209bb&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">reg = TabPFNRegressor()
reg.fit(X=X_train, y=y_train)
predictions = reg.predict(X_test)</code></pre></div><p>The <code>predictions</code> object in the code above contains just the means of the full predictive distributions. But in the background, the foundation model actually produces the full distribution (or at least a discretized version over a fixed support) and later aggregates it.</p><p>If you want the predictive distribution, you need to change a paramter in .predict(): </p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;06196755-5cb0-4d20-a4f2-b119c554c083&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">quantiles_range = np.arange(0.05, 0.95, 0.05)
quantiles = reg.predict(X_test, output_type='quantiles', quantiles=quantiles_range)</code></pre></div><p>Both output type options (mean, quantile) take the same amount of compute time, since the tabular foundation models predict the distribution anyways.</p><h2>Why prefer distributions over points?</h2><p>Starting from P(Y|X = x) may seem more tedious, but it gives us more modeling freedom. There are many situations where the mean is not appropriate, or where we want more informative predictions:</p><ul><li><p>You might prefer the predictive median over the mean to get more robust predictions.</p></li><li><p>If  you are interested in the tails of the distribution, you can extract the 10% and the 90% quantiles (or any other quantiles for that matter).</p></li><li><p>Computing quantile intervals or the variance of the predictive distribution can help quantify uncertainty.</p></li><li><p>Maybe you need the probability that the prediction exceeds a certain threshold, P(Y&gt;threshold).</p></li><li><p>Maybe you want to extract the modalities.</p></li><li><p>Or you can also visualize the entire distribution.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!Gx8X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!Gx8X!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png 424w, /__u/substackcdn.com/image/fetch/$s_!Gx8X!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png 848w, /__u/substackcdn.com/image/fetch/$s_!Gx8X!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Gx8X!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!Gx8X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png" width="602" height="327.85631517960604" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:470,&quot;width&quot;:863,&quot;resizeWidth&quot;:602,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!Gx8X!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png 424w, /__u/substackcdn.com/image/fetch/$s_!Gx8X!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png 848w, /__u/substackcdn.com/image/fetch/$s_!Gx8X!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png 1272w, /__u/substackcdn.com/image/fetch/$s_!Gx8X!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2beb351b-2bac-4b18-8147-8d01d4517dbe_863x470.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We have so many more options when we work with the predictive distribution. And we get it without any additional computations when using tabular foundation models.</p><h1>Are the predictive distributions calibrated?</h1><p>While TFMs are directly pre-trained to predict the full distribution, there is no guarantee that this distribution is calibrated, meaning, e.g., that the predicted 10% quantile matches the actual 10% quantile. What distinguishes TFMs from approaches like linear regression or Bayesian regression models is the lack of explicit distributional assumptions.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> But to get a better idea whether to trust the full distributions, we need benchmarks. Unfortunately,  popular benchmarks like  <a href="/__u/mindfulmodeler.substack.com/p/tabarena-explained">TabArena</a> focus on point estimates, not the entire distribution. <a href="https://arxiv.org/pdf/2603.08206">This paper</a> criticizes the field&#8217;s focus on the conditional mean and proposes reporting metrics in benchmarks that reflect calibration. Keep in mind that even when a model scores well in a benchmark, you have no guarantees for your own project. It&#8217;s just a rough pointer.</p><p>In the end, you need to validate calibration yourself, using proper scoring rules, or even calibrate using conformal prediction. But still, I find it exciting that tabular foundation models give us the option of predicting other aspects of the predictive distribution.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>TFMs do have some implicitly encoded assumptions through their pre-training on synthetic data.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[TabArena explained]]></title><description><![CDATA[What chess ratings tell us about ML models]]></description><link>https://mindfulmodeler.substack.com/p/tabarena-explained</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/tabarena-explained</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 31 Mar 2026 12:09:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!hjDv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Machine learning progresses through benchmarks. While I&#8217;ve been critical before of ML&#8217;s <a href="/__u/mindfulmodeler.substack.com/p/we-are-obsessed-with-benchmarks?utm_source=publication-search">benchmark obsession</a> and <a href="/__u/mindfulmodeler.substack.com/p/stuck-on-benchmark-island?utm_source=publication-search">danger of getting stuck on benchmarks</a>, benchmarks are essential to guide researchers and practitioners in the right direction. ImageNet, under the leadership of Fei-Fei Li, for example, ushered in the deep learning era. Also today, LLMs are largely driven by benchmarks, like <a href="https://lastexam.ai">Humanity&#8217;s Last Exam</a> or <a href="https://livecodebench.github.io">LiveCodeBench</a> (coding).</p><p>With the <a href="/__u/mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird">rise of tabular foundation models</a>, a benchmark that often comes up is TabArena. Tabular, of course, is already way more mature than deep learning was in 2012. However, TabArena still introduces a new dynamic and is a change from other tabular benchmarks. For example, two days ago, <a href="https://www.linkedin.com/posts/prior-labs_priorlabs-tabpfn-tabularfoundationmodels-activity-7442923489168232448-vBBb?utm_source=share&amp;utm_medium=member_desktop&amp;rcm=ACoAABLM5UsB_MMk8Jy1LghrgnOGSF3acMDe3c0">Prior Labs announced TabPFN v2.6</a>. A huge part of that announcement was the TabArena placement of the new model (first spot). TabArena has become a central element in the paradigm shift to tabular foundation models.</p><p>So what is TabArena? Let&#8217;s find out.</p><h2>TabArena &#8220;lives&#8221; on HuggingFace</h2><p>TabArena is a &#8220;living&#8221; benchmark, and you can find it in its <a href="https://huggingface.co/spaces/TabArena/leaderboard">natural habitat on HuggingFace</a>. It&#8217;s by far not the only tabular benchmark; there have been many before it, like <a href="https://arxiv.org/abs/1708.03731">OpenML-CC18</a> or <a href="https://github.com/naszilla/tabzilla">Tabzilla</a>.</p><p>These benchmarks have been crucial in advancing machine learning. Tabular benchmarks, at their minimum, define a collection of ML tasks connected to datasets. What sets them apart from a pure dataset repository, like UCI, is stronger curation of datasets, and sometimes a benchmarking protocol, e.g., for how to evaluate models.</p><p>TabArena is probably the benchmark with the strictest protocol and thoroughness:</p><ul><li><p>TabArena standardized the pre-processing and evaluation procedures.</p></li><li><p>It also includes ensembles for each model (except TabICL and TabDPT).</p></li><li><p>Has the highest limit for tuning and training the model, meaning more confidence that the search is maxed out.</p></li><li><p>Not only metric results, but also predictions are available.</p></li></ul><p>What makes it &#8220;live&#8221; is that it&#8217;s an actively maintained benchmark. Not just a static table in a paper, but a space on HuggingFace that the TabArena maintainers update with new ML algorithms and tabular foundation models. TabArena was created by researchers from Amazon Web Services, the University of Freiburg, the University of Mannheim, INRIA Paris, Ecole Normale Sup&#233;rieure, the ELLIS Institute T&#252;bingen, and Prior Labs. The core maintainers are Nick Erickson, Lennart Purucker, Andrej Tschalzev, and David Holzm&#252;ller.</p><h2>TabArena pits ML algorithms against each other</h2><p>TabArena rates ML algorithms and foundation models using an Elo rating system, which you might know from chess or other competitive games. Elo is for pairwise comparisons, and the rating of an ML algorithm / TFM can be interpreted as the expected win probability for a task. Winning, as in producing a better model than another algorithm on a given task. The following figure shows the data going into the Elo ranking. For example, TabICLv2 was beating CatBoost in 76% of tasks, or, inversely, CatBoost was beating TabICLv2 in 24% of the TabArena tasks.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!hjDv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!hjDv!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png 424w, /__u/substackcdn.com/image/fetch/$s_!hjDv!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png 848w, /__u/substackcdn.com/image/fetch/$s_!hjDv!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hjDv!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!hjDv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png" width="1446" height="1156" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1156,&quot;width&quot;:1446,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!hjDv!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png 424w, /__u/substackcdn.com/image/fetch/$s_!hjDv!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png 848w, /__u/substackcdn.com/image/fetch/$s_!hjDv!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png 1272w, /__u/substackcdn.com/image/fetch/$s_!hjDv!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa71d3775-e773-4dca-886f-318b69b3f6a7_1446x1156.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Win rate matrix of TabArena algorithms. Source: <a href="https://huggingface.co/spaces/TabArena/leaderboard">tabarena.ai</a></figcaption></figure></div><p>Elo reflects well how we select models in practice. When we compare models for a project, we want one that beats the others in terms of performance. There may be other constraints, such as interpretability and computational performance, but predictive performance ranks high. Elo only counts wins versus losses (or ties), but not by which margin. If two algorithms have the same win rate, their Elo will be the same, even when one algorithm fails catastrophically on the losses, and the other always comes in second by a small margin. This, however, is covered by other metrics on TabArena, such as improvability.</p><h2>TabArena contains 51 tasks</h2><p>TabArena benchmarks ML approaches based on 13 regression and 38 classification datasets.</p><p>Sounds small.</p><p>Especially small, given the amount of tabular data flying all around. To be fair, most of them are locked up in companies. You won&#8217;t find Aldi sales data, OpenAI churn tables, or spring coil test datasets flying around the internet with an open license. But still, OpenML hosts <a href="https://www.openml.org/search?type=data&amp;sort=runs&amp;status=active">over 6k datasets</a>! The only problem is: once you start being just a little selective, the number drops fast.</p><p>The TabArena team started with 1053 datasets and sequentially filtered them down:</p><ul><li><p>They removed 491 duplicate datasets.</p></li><li><p>135 datasets were tabular only in disguise: they were actually derived from other data modalities like images.</p></li><li><p>The team removed another 123 for lack of a real predictive task (e.g., deterministic outcomes).</p></li><li><p>Tiny data, quality issues, incompatible license &#8230; meant another 181 were excluded.</p></li><li><p>Of the remaining data, they reduced 70 non-IID datasets.</p></li></ul><p>Leaving only 51 datasets after this manual process, see also the <a href="https://arxiv.org/html/2506.16791v4">TabArena paper</a>.</p><h2>TabArena snapshot</h2><p>Let&#8217;s finally have a look at the current leaderboard of the overall benchmark:</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!DlXp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!DlXp!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png 424w, /__u/substackcdn.com/image/fetch/$s_!DlXp!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png 848w, /__u/substackcdn.com/image/fetch/$s_!DlXp!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png 1272w, /__u/substackcdn.com/image/fetch/$s_!DlXp!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!DlXp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png" width="1456" height="342" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:342,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!DlXp!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png 424w, /__u/substackcdn.com/image/fetch/$s_!DlXp!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png 848w, /__u/substackcdn.com/image/fetch/$s_!DlXp!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png 1272w, /__u/substackcdn.com/image/fetch/$s_!DlXp!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2e3b56c-89ad-4077-b568-d5401c228bf1_3450x810.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption">TabArena benchmark results. Source: <a href="https://tabarena.ai/">tabarena.ai</a></figcaption></figure></div><p>Current ceiling: AutoGluon 4h extreme. AutoGluon is an open-source AutoML framework by AWS, and &#8220;4h extreme&#8221; is a preset that allows up to 4 hours of training time with the most intensive ensemble and hyperparameter search settings for maximum predictive performance. The top 3 positions are all tabular foundation models: TabPFN 2.6, RealTabPFN-2.5, and TabICLv2. All the way down at position 7 appears the first boosted tree algorithm with LightGBM. Reminder:  Elo rankings don&#8217;t tell us about the absolute differences in performance. Also, the number of datasets is a caveat, and keep in mind that TabArena represents small to mid-sized IID data. Nonetheless, still impressive how far tabular foundation models have come.</p><p>If you want to dive deeper, there is a <a href="https://www.youtube.com/watch?v=mcPRMcJHW2Y">TabArena presentation on YouTube</a>.</p>]]></content:encoded></item><item><title><![CDATA[The interpretability tax on tabular foundation models]]></title><description><![CDATA[Before &#8220;betting&#8221; on tabular foundation models, I made a similar bet on model-agnostic interpretable machine learning: I wrote two books and did a PhD.]]></description><link>https://mindfulmodeler.substack.com/p/tabular-foundation-models-break-the</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/tabular-foundation-models-break-the</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 24 Mar 2026 13:20:09 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!0FrG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Before <a href="/__u/mindfulmodeler.substack.com/p/im-betting-on-tabular-foundation">&#8220;betting&#8221; on tabular foundation models</a>, I made a similar bet on model-agnostic interpretable machine learning: I wrote <a href="https://christophmolnar.com/books/interpretable-machine-learning">two</a> <a href="https://christophmolnar.com/books/shap">books</a> and did a PhD. Model-agnostic interpretability fascinated me because of its potential &#8220;timelessness&#8221;. Since they work by studying the model predictions when manipulating the inputs, these methods should have a longer shelf-life than model-specific methods.</p><p><a href="/__u/mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird">Tabular foundation models</a> are a paradigm shift from training+predict  to pretraining+ICL (in-context learning), making it a great test for model-agnostic interpretability tools.</p><p>Let&#8217;s dive in.</p><p>The remainder of this post is based on research from the paper <a href="https://arxiv.org/abs/2403.10923">Interpretable Machine Learning for TabPFN</a>.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></p><h2>Model-agnostic interpretation works out of the box for tabular foundation models</h2><p>Consider this minimal workflow: You train a model, and then study its feature importance using permutation feature importance (PFI). PFI is a classic, post-hoc model-agnostic interpretability method. Meaning it works regardless of the underlying model structure. It only needs access to the .predict() function. Behind &#8220;.predict()&#8221; could even secretly be Bob from marketing filling that Excel based on gut feeling. If PFI works for Bob, it will work for TFMs, right? Indeed, it does work and looks and feels very familiar.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;9313c54f-1d37-471f-b83d-a0153f84c736&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">reg = TabICLRegressor()
reg.fit(X_train, y_train)
pfi = sklearn.inspection.permutation_importance(reg, X_test, y_test)</code></pre></div><p>The sklearn implementation of PFI computes R-squared for the test data, then shuffles each feature and checks R-squared again. The larger the drop, the more important that feature was. This is repeated for all features, typically averaged over a couple of such permutations, five by default.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!wQxc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!wQxc!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png 424w, /__u/substackcdn.com/image/fetch/$s_!wQxc!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png 848w, /__u/substackcdn.com/image/fetch/$s_!wQxc!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png 1272w, /__u/substackcdn.com/image/fetch/$s_!wQxc!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!wQxc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png" width="470" height="326" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:326,&quot;width&quot;:470,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:11314,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/191965788?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!wQxc!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png 424w, /__u/substackcdn.com/image/fetch/$s_!wQxc!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png 848w, /__u/substackcdn.com/image/fetch/$s_!wQxc!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png 1272w, /__u/substackcdn.com/image/fetch/$s_!wQxc!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff33c7ced-7e04-4887-b3d9-931c1d648398_470x326.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Feature importance results based on PFI.</figcaption></figure></div><p>So far, so familiar.</p><p>If it weren&#8217;t for the changed costs.</p><h2>Inference-cost inversion changes the economics of interpretability</h2><p>On my MacBook Air M1, computing PFI took 104 seconds for this dataset of 1k rows (700 &#8220;training&#8221;, and 300 test data). Way more than a random forest would need.</p><p>The reason is that tabular foundation models are still slow. Not for training, but at prediction time. While traditional machine learning is training-expensive and inference-cheap, this relationship inverts for tabular foundation models: &#8220;Training&#8221; costs almost nothing, but predictions are expensive. Note that I&#8217;m excluding pre-training here, because I&#8217;m viewing it from an application perspective. Side note: While not a classic real training step, <code>reg.fit()</code> loads the TFM&#8217;s weights and pre-processes the &#8220;training&#8221; data, aka the context for in-context learning.</p><p>We could just ignore this cost inversion and throw more compute at our problem. Absence of training is neat anyway, and maybe we can just swallow the increased cost of inference. Certainly a possibility for the GPU-rich. But even then, the economics of interpretability are not so favorable. Permutation feature importance in the example above already requires 26x (1 + number of repetitions x number of features) times as many predictions as just predicting the test data. For an interpretability metric, you may only compute once. But if you start computing things like Shapley values along with every prediction, costs get out of hand quickly.</p><p>So what to do?</p><h2>Better scaling for inference-heavy interpretability</h2><p>We can speed up inference-heavy interpretability methods through TFM-friendly implementations. For traditionally trained models, it makes sense to chunk calls to the model: For PFI, you might call <code>model.predict() </code>with every new shuffling of the data. For TFMs, this no longer makes much sense. According to the paper &#8220;Interpretable Machine Learning for TabPFN&#8221;, the cost of inference per test data scales with <code>O(ntrain^2/ntest)</code>, where <code>ntrain</code> is the size of the training data and <code>ntest</code> the size of the test data (in rows). So it&#8217;s better to bundle more test data into one call (if memory permits), because &#8212; relatively &#8212; it gets cheaper. I tested this out for computing PFI for one feature (computation time was averaged over 100x doing this). The naive version is two calls to the model, once with the original test data, once with the permuted. The batched version concatenates both before sending the data to the TFM for prediction.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!0FrG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!0FrG!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png 424w, /__u/substackcdn.com/image/fetch/$s_!0FrG!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png 848w, /__u/substackcdn.com/image/fetch/$s_!0FrG!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0FrG!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!0FrG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png" width="470" height="326" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:326,&quot;width&quot;:470,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:10721,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/191965788?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!0FrG!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png 424w, /__u/substackcdn.com/image/fetch/$s_!0FrG!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png 848w, /__u/substackcdn.com/image/fetch/$s_!0FrG!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png 1272w, /__u/substackcdn.com/image/fetch/$s_!0FrG!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F150a56bd-72a6-48e7-9489-4f74e15d4a0f_470x326.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Comparing chunked versus batched computations of PFI for TabICL.</figcaption></figure></div><p>Efficiency for other methods like ICE, ALE, and PDP can also benefit from such implementation changes.</p><h2>Adapting interpretability tools for a training-cheap, inference-expensive world</h2><p>LOCO importance is a model-agnostic interpretability method that can also compute feature importances. It stands for &#8220;leave-one-covariate-out&#8221; and relies on re-training: Remove a feature from the training data, re-train the model, and compare performances on test data.</p><p>With traditional machine learning, LOCO is expensive when compared to PFI. To compute LOCO importance for an xgboost model with 100 features, you have to re-train the model 100 times.</p><p>This relation changes with TFMs as they are training-cheap and interpretability methods based on re-training become more attractive. Because under TFMs, LOCO costs roughly the same as PFI, due to the lack of a training step. &#8220;Re-training&#8221; under the TFM paradigm just becomes another forward pass of the data with one column removed.</p><p>For the following figure, I computed PFI and LOCO for one feature for both TabICL and a random forest, repeated 100 times, and measured how long it takes:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!c1uF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!c1uF!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png 424w, /__u/substackcdn.com/image/fetch/$s_!c1uF!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png 848w, /__u/substackcdn.com/image/fetch/$s_!c1uF!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c1uF!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!c1uF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png" width="562" height="374" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:374,&quot;width&quot;:562,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:15610,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/191965788?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!c1uF!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png 424w, /__u/substackcdn.com/image/fetch/$s_!c1uF!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png 848w, /__u/substackcdn.com/image/fetch/$s_!c1uF!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png 1272w, /__u/substackcdn.com/image/fetch/$s_!c1uF!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3f44125b-b3f4-422f-8465-bdfdff2f8ce3_562x374.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Comparing PFI and LOCO for TabICL and a Random Forest.</figcaption></figure></div><p>The random forest is, in total, much cheaper; however, note that no hyperparameter tuning is involved, so it&#8217;s not a 100% fair comparison. Look at the relation between PFI computation time versus LOCO computation time. For TabICL, it&#8217;s roughly the same, while for the random forest, LOCO is 10x more expensive compared to PFI.</p><p>Note that LOCO and PFI define feature importance differently. With PFI, we stick to a fixed model and see how it adapts to shuffling data. With LOCO, we ultimately compare two different models. When features are correlated, LOCO and PFI importances diverge and have different interpretations. These differences boil down to the question of whether our interpretation is <a href="/__u/mindfulmodeler.substack.com/p/audit-or-insight-know-your-interpretation?utm_source=publication-search">true to the data (LOCO) or true to the model (PFI)</a>. Or, in a more cynical view, yet another case of <a href="/__u/mindfulmodeler.substack.com/p/correlation-can-ruin-interpretability">correlation ruining interpretability</a>.</p><p>Other interpretability methods, like Shapley values, also have re-training-based versions, which become more attractive in conjunction with TFMs. But again, with a changed interpretation.</p><h2>So what?</h2><p>All model-agnostic tools are still available for tabular foundation models, but the costs have shifted. We can adapt interpretability methods to some degree, e.g., by batch calling, or we might shift to more training-focused versions.</p><p>And, who knows, we might also see different types of interpretability arise:</p><ul><li><p>Mechanistic interpretability for tabular foundation models.</p></li><li><p>Model-specific implementations. Think of TreeSHAP for tree-based models ( PFNshap?)</p></li><li><p>TFMs might enable new types of interpretability in the form of directly estimating properties of the underlying data.</p></li></ul><p>I&#8217;m looking forward to seeing how interpretability for TFMs will evolve.</p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Rundel, David, et al. &#8220;Interpretable machine learning for TabPFN.&#8221; World Conference on Explainable Artificial Intelligence. Cham: Springer Nature Switzerland, 2024.</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[I’m betting on tabular foundation models]]></title><description><![CDATA[The shipping container moment of data science]]></description><link>https://mindfulmodeler.substack.com/p/im-betting-on-tabular-foundation</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/im-betting-on-tabular-foundation</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 17 Mar 2026 09:24:10 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!-ksB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Before the 1950s, we shipped goods in boxes, crates, and sacks. Dock workers would load and unload each item individually.</p><p>Then came the shipping container.</p><p>Containerization came with huge upfront costs: ships had to be refitted, ports needed new cranes, and processes had to adapt. On the upside, containers drastically reduced loading times and labor. It also standardized transportation of goods: The same container could be loaded onto a truck or train without unpacking. Containerization standardized how goods move. Once everything fit into the same container, the logistics became simpler and faster.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!-ksB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!-ksB!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png 424w, /__u/substackcdn.com/image/fetch/$s_!-ksB!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png 848w, /__u/substackcdn.com/image/fetch/$s_!-ksB!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-ksB!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!-ksB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png" width="1456" height="469" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:469,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:34208,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/191133165?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!-ksB!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png 424w, /__u/substackcdn.com/image/fetch/$s_!-ksB!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png 848w, /__u/substackcdn.com/image/fetch/$s_!-ksB!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png 1272w, /__u/substackcdn.com/image/fetch/$s_!-ksB!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0be3299c-0807-4bef-840b-9f4c129de249_1800x580.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Before and after containerization</figcaption></figure></div><p>Machine learning has seen a similar kind of standardization. If you wanted to train an image recognition model for dogs in 2005, you&#8217;d maybe have applied edge detectors, thresholded intensities, and hand-crafted a feature vector. Then train something like a support vector machine on top. Every problem required its individual solution.</p><p>Today, you can start from the same place regardless of the task: a pre-trained foundation model. Take a Vision Transformer, fine-tune it on your data, or train a logistic classifier on the embeddings. And it will likely just work. Wherever you look, you&#8217;ll find pre-trained, often transformer-based, foundation models: for translation, image recognition, image segmentation, speech-to-text, text-to-speech, video generation, &#8230;</p><p>All data modalities and tasks are occupied by foundation models.</p><p>All? No! One small modality still holds out against them: tabular data.</p><p>But this resistance is crumbling.</p><h2>TFMs are a fundamental shift, not just a performance trade-off</h2><p>TabPFN opened the era of foundation models for tabular data. For small and mid-sized data, tabular foundation models now outperform other ML algorithms (see <a href="https://huggingface.co/spaces/TabArena/leaderboard">TabArena</a>). I feel like the discussion in the ML community has since switched from &#8220;performance not there&#8221; to:</p><ul><li><p>&#8220;Is a slight increase in performance worth the increased compute?&#8221;</p></li><li><p>&#8220;But it doesn&#8217;t work for large data!&#8221;</p></li><li><p>&#8220;Wow, so much wasted compute when you could just use linear regression / random forest.&#8221;</p></li></ul><p>I agree. For many applications, the increased compute time is an obstacle. But for how long will this be the case? Ultimately, I expect advances in hardware, architecture, and inference tricks to fix these concerns.</p><p>As a machine learning community, we are very much focused on benchmarks and computational performance. That&#8217;s both the strength and the <a href="/__u/mindfulmodeler.substack.com/p/we-are-obsessed-with-benchmarks">weakness</a> of machine learning. The focus on benchmarks was certainly difficult for me when I started publishing on ML interpretability, but that&#8217;s another story.</p><p>Purely judging tabular foundation models in terms of performance and compute distracts from the larger shift I believe tabular ML is undergoing. During <a href="/__u/mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird">my deep dive into tabular foundation model</a>s, my intuition grew that this is something big. It took me a while to figure out why exactly I think TFMs are a fundamental shift in data science. Why TFMs are not just slightly better-performing but more expensive models.</p><p>This post explains why I now believe TFMs are a really big deal.</p><h2>PFNs are a universal engine for tabular ML</h2><p>CatBoost, logistic regression, SVMs &#8211; the boxes, crates, and sacks of tabular ML &#8211; get the job done. Scikit-learn has done a great job of containerizing them at least API-wise with .fit, .predict, &#8230;. But underneath, they are very different algorithms, producing very different models. Some thrive on CPU, some on GPU. Very different optimization, failure modes, task-readiness, and so on. You always have to train from scratch, which is especially difficult with smaller datasets. Integrating prior domain knowledge is also inconsistent between ML algorithms. For example, if you suspect your data should have a maximum interaction depth of two, for linear models, you would have to actually add interaction terms, while for tree-based models, you would have to restrict the depth of learned trees.</p><p>All of this changes with tabular foundation models, specifically with prior-data fitted networks (PFNs). Pre-trained on millions of synthetic data examples, these transformer-based models are trained to perform in-context learning. Prediction essentially becomes a single forward pass of training and test data at once, and &#8211; without changing the weights &#8211; the model predicts the test data by attending to the training data.</p><p>To keep with the analogy of shipping, we finally have a general-purpose container for tabular machine learning, or even for data science more generally. The general purpose technology enabling this containerization is the prior-data fitted networks. Read <a href="/__u/mindfulmodeler.substack.com/p/how-pfns-make-tabular-foundation">my post about PFNs here</a>. PFNs are a general method of essentially self-supervised learning that consists of creating synthetic (tabular) tasks, on which a (transformer-based) model is pre-trained to predict new data via in-context learning. Thanks to PFNs, we now have tabular foundation models that can do classification and regression without needing to be trained on the particular task at hand.</p><p>If classification and regression were the only applications, I probably wouldn&#8217;t call it a revolution. While these two tasks already make up a substantial portion of business and research tasks, tabular ML and data science are much larger. For TFMs, classification and regression were the first milestones to be taken seriously by the ML community.</p><p>But TFMs are not stopping there. Any<strong> </strong>data science task that you can express as a self-supervised learning problem (including a data generator) can be learned by the PFN &#8220;engine&#8221; to pre-train a TFM that then one-shots this task.</p><p>Let&#8217;s have a look at the implications.</p><h2>If you can simulate it, you can pre-train for it</h2><p>Off-the-shelf classification and regression TFMs <a href="https://arxiv.org/abs/2601.22259">already work for survival analysis</a> and <a href="https://arxiv.org/html/2501.02945v3">time series forecasting</a>. That&#8217;s impressive and relevant, but not the paradigm shift I&#8217;m talking about. The big shift comes from being able to pre-train tabular foundation models for ANY data science problem for which we are able to simulate data with a ground truth. Researchers have already begun pre-training foundation models for other tasks:</p><ul><li><p><strong>Time-series forecasting:</strong>  <a href="https://www.emergentmind.com/topics/timepfn">TimePFN</a></p></li><li><p><strong>Anomaly Detection:</strong> <a href="https://arxiv.org/abs/2409.05672">FoMo-0D</a></p></li><li><p><strong>Classification and regression for graphs:</strong>  <a href="https://arxiv.org/html/2509.21489v2">GraphPFN</a>.</p></li><li><p><strong>p&gt;n prediction problems (like genomics):</strong> <a href="https://arxiv.org/abs/2510.06162">TabPFN-Wide</a> </p></li><li><p><strong>Missing data imputation:</strong>  <a href="https://arxiv.org/abs/2510.02625">TabImpute</a></p></li><li><p><strong>Clustering:</strong>  <a href="https://arxiv.org/abs/2601.21656">TabClustPFN</a>.</p></li><li><p><strong>Counterfactual fairness:</strong>  <a href="https://arxiv.org/abs/2407.05732">FairPFN</a></p></li><li><p><strong>Estimate Shapley values</strong>: <a href="https://github.com/joaopfonseca/ExplainerPFN">ExplainerPFN</a></p></li><li><p><strong>Causal inference:</strong> <a href="https://openreview.net/forum?id=Jb9yNhfsEM">Do-PFN</a> and <a href="https://github.com/yccm/CausalFM-toolkit">CausalFM</a></p></li><li><p><strong>Bayesian Optimization surrogate model:</strong>  <a href="https://arxiv.org/abs/2305.17535">PFNs4BO</a></p></li></ul><p>This non-exhaustive list shows how versatile the PFN-engine is to pre-train tabular foundation models beyond classification and regression. Pre-training a TFM is not a small feat, especially since designing the prior, aka the data distribution, is not always easy. But once pre-trained, you can use it on any table for that same task. I expect this list to become even larger while also niching down into more specific use cases.</p><p>The PFN-engine also allows you to instill any inductive bias into the model that you can define via the prior. For example, you can <a href="https://arxiv.org/abs/2512.03307">robustify TFMs through adversarial training</a>. Or, to absorb useful inductive biases from tree-based models, you can <a href="https://arxiv.org/abs/2502.05564">add tree-based relationships into the prior</a> (<a href="https://arxiv.org/abs/2405.13396">see also this paper</a>). This all comes on top of PFNs already being able to adapt to inductive biases through in-context learning alone, as <a href="https://arxiv.org/abs/2511.18278">research suggests</a>.</p><p>And since they are so versatile and predict the entire posterior predictive distribution, not just a point estimate, TFMs can also be used for <strong>density estimation</strong> and <strong>synthetic data generation</strong>. Both are already built into <a href="https://github.com/PriorLabs/TabPFN">TabPFN</a> and would also be doable with other TFMs.</p><p>All these facts and developments show how much of a general-purpose technology PFNs are: There will be TFMs for every suitable data science problem soon.</p><h2>Tabular ML now rides the deep learning wave</h2><p>The non-tabular machine learning world has already converged on deep neural networks, and many even on transformer-based foundation models. Tabular is, or rather, was the last bastion to hold onto its boxes, crates, and sacks. The only modality not using the same &#8220;container&#8221; that the other modalities are using.</p><p>By switching over to tabular foundation models, we are essentially streamlining tabular with the other modalities. I liken technical developments to a river. The huge investments and talent in deep learning/transformers make this a very strong river. Tabular is now dipping into this flow, and the transformer river is carrying tabular with it. Suddenly, tabular ML also benefits from the vast ecosystem:</p><ul><li><p>Investments in GPU infrastructure now benefit tabular.</p></li><li><p>Any further, more specific hardware for neural network training/inference makes tabular ML more efficient.</p></li><li><p>Race to make LLM inference cheap also helps TFMs.</p></li><li><p>Architectures (e.g., attention mechanisms) are getting better.</p></li><li><p>All the optimizations and learnings for training and deploying LLMs, like LoRA, quantization, and Muon, may be transferred.</p></li><li><p>We can build with very mature and user-friendly deep learning software like PyTorch (I still remember what a pain TensorFlow was in 2018).</p></li><li><p>Tabular data now becomes more interesting to a large pool of AI investors.</p></li></ul><p>All these developments carry tabular foundation models, but not SVMs or boosted trees. TFMs make tabular data science snap into the broader ecosystem.</p><p>Besides being carried, there&#8217;s improved interoperability with other modalities: TFMs give us embeddings for tabular data, making it straightforward to represent and reuse tables. For example, it has now become easier to train multi-modal models with tabular data as a modality.</p><h2>I&#8217;m betting on foundation models</h2><p>It&#8217;s been some time since I&#8217;ve been this excited about tabular ML. Tabular foundation models feel like the first refreshing change in a long time. I just don&#8217;t care about yet another algorithm that produces a tree ensemble.</p><p>Given my reasoning above (containerization), I&#8217;m personally betting on TFMs. (&#8220;Betting&#8221; as in investing my time and attention, not as in participating in prediction markets.) I am planning to learn more about them, to push them forward, and to educate. For example, I started working on <a href="https://tabular-foundation.christophmolnar.com">TFM overview</a>, a website that tracks libraries and labs.</p>]]></content:encoded></item><item><title><![CDATA[The Random Forest of the 2030s?]]></title><description><![CDATA[Three scenarios for the future of tabular ML, plus one that makes them all irrelevant]]></description><link>https://mindfulmodeler.substack.com/p/the-random-forest-of-the-2030s</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/the-random-forest-of-the-2030s</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 03 Mar 2026 13:32:36 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!WtQm!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc3dc1d90-fb9f-4e3c-80d5-756a5d6e8495_500x500.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>This is the final post in the series on tabular foundation models (see <a href="/__u/mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird">#1</a>, <a href="/__u/mindfulmodeler.substack.com/p/how-pfns-make-tabular-foundation">#2</a>, <a href="/__u/mindfulmodeler.substack.com/p/the-architecture-behind-tabpfn">#3</a>, <a href="/__u/mindfulmodeler.substack.com/p/how-tabular-foundation-models-are">#4</a>, <a href="/__u/mindfulmodeler.substack.com/p/the-state-of-tabular-foundation-models">#5</a>, <a href="/__u/mindfulmodeler.substack.com/p/a-regression-example-with-tabiclv2">#6</a>).</p><p>Tabular foundation models (TFMs) are a paradigm shift from traditional tabular ML: They are transformer-based architectures pre-trained on synthetic data. There is no classic training step. Instead, TFMs predict the test data in a single forward pass of combined training and test data without any parameter updates (in-context learning).</p><p>These last few weeks of deep-dive have reshaped how I think about TFMs and tabular ML as a whole. I won&#8217;t claim I can predict the future. I&#8217;ve been completely wrong before, like about how good AI would become at coding. Instead of predictions, here are a few scenarios of increasing impact of TFMs (levels) on everyday tabular ML work.</p><p>Let&#8217;s dive in.</p><h2>Level 1: TFMs enter the pantheon of ML algorithms</h2><p>In most projects, you already try a couple of ML algorithms. Who cares that TabICL and other TFMs are pre-trained? Just one more algorithm to throw at your data. An expensive one, so maybe you just use it for smaller datasets. In this scenario, TFMs would become an established member of the ML algorithms that you find in benchmarks and that people try out in model selection.</p><h2>Level 2: TFMs become the new Random Forest, aka the quick-and-dirty model</h2><p>Whenever I encounter a new tabular prediction problem, I throw the random forest at it. Why the random forest? It&#8217;s reasonably fast and works well without hyperparameter tuning. For you, this may be a different algorithm that gives you that feeling of comfort and convenience. Maybe it&#8217;s linear/logistic regression or the support vector machine. In this scenario, TFMs would become the quick-and-dirty baseline for most supervised ML tasks. It would make sense: Great performance without hyperparameter tuning, implementations for both regression and classification, can handle missing data, can also predict quantiles, and so on.</p><h2>Level 3: TFMs become synonymous with tabular ML</h2><p>In this scenario, TFMs become the go-to approach for all supervised machine learning tasks, from classification and regression to survival analysis and time series forecasting. Kind of what LLMs did in the NLP space. TFMs would dominate all other ML algorithms in performance, at least for most cases, but also in convenience. TFMs would become the Swiss knife of tabular ML: they would work for all types of supervised ML tasks, handle missing data with ease, quantify uncertainty, generate new data if you want, and much more. Developers would build an entire ecosystem of tools around TFMs, from easy deployment to interpretability and monitoring. The ecosystem grows and becomes so powerful and convenient that it&#8217;s hard to switch back to any other ML algorithm.</p><h1>Where do we stand? My impression</h1><p>For <strong>Level 1,</strong> the barriers are cleared: The software is there, it&#8217;s usable, and there are enough clever tricks to make TFMs (somewhat) work even  for larger datasets. I expect costs to keep dropping. However, it hasn&#8217;t permeated yet: Not everyone knows about TFMs, and it will take time for these models to reach all the toolboxes.</p><p>For me personally, <strong>Level 2</strong> is cleared, at least for small datasets. TFMs will be my go-to model for small to mid-sized data. Too convenient not to. (Sorry, Random Forest.) Whether it&#8217;s true for others modelers will remain to be seen and depends on how computational costs evolve.</p><p><strong>Level 3</strong> is clearly not reached, and I am not sure if it ever will be. Will TFMs ever become the default for supervised ML? This strongly depends on how much improvements in TFM architectures and inference tricks can bring down computational costs. And this probably won&#8217;t be a global binary thing, but can also be quite different between communities. For example, the random forest has dominated ecology for years, while other communities haven&#8217;t embraced the Random Forest as much. Currently, there&#8217;s definitely an ecosystem growing around these TFMs. The TFM ecosystem could become a little modeling universe, covering all the supervised ML needs, so no reason to use anything outside this universe.</p><h2>TFMs vs. traditional ML may become an irrelevant question</h2><p>Developments are happening in parallel with agentic AI, or, more specifically, coding agents that may make this entire TFM discussion irrelevant. I&#8217;ve jumped on the vibe-coding bandwagon, coding an Android app for personal use (basically a recording app that dispatches my voice notes to the right places). I&#8217;ve been critical about vibe-coding, but my recent experience with Claude Code has drastically changed my view on agentic AI. I believe the world of software engineering has shifted quite a bit. Similar automation may be coming for data science and supervised machine learning (or kind of already is already here). This may make the entire TFM versus &#8220;traditional&#8221; ML discussion irrelevant, because these choices are abstracted away. At least from a user perspecte, as the data scientist / ML engineer (or whatever the job may be called then) operates on a higher level of abstraction.</p><p>This concludes my series on tabular foundation models. Which doesn&#8217;t mean that I won&#8217;t keep posting on it, it&#8217;s just giving me the freedom to write about other things as well.</p>]]></content:encoded></item><item><title><![CDATA[A regression example with TabICLv2]]></title><description><![CDATA[Strong performance at a steep interpretability cost (at least on CPU)]]></description><link>https://mindfulmodeler.substack.com/p/a-regression-example-with-tabiclv2</link><guid isPermaLink="false">https://mindfulmodeler.substack.com/p/a-regression-example-with-tabiclv2</guid><dc:creator><![CDATA[Christoph Molnar]]></dc:creator><pubDate>Tue, 24 Feb 2026 11:34:51 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/0ad63a8d-79be-4b78-8664-ea200a7e7e20_742x441.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><em>This is post #6 of the Tabular Foundation Model (TFM) series (see <a href="/__u/mindfulmodeler.substack.com/p/tabular-ml-is-about-to-get-weird">#1</a>, <a href="/__u/mindfulmodeler.substack.com/p/how-pfns-make-tabular-foundation">#2</a>, <a href="/__u/mindfulmodeler.substack.com/p/the-architecture-behind-tabpfn">#3</a>, <a href="/__u/mindfulmodeler.substack.com/p/how-tabular-foundation-models-are">#4</a>, and <a href="/__u/mindfulmodeler.substack.com/p/the-state-of-tabular-foundation-models">#5</a>).</em></p><p>This post walks through a TabICLv2 regression workflow: loading data, &#8220;fitting&#8221;, predicting, quantifying uncertainty, and interpreting the predictions with SHAP.</p><p>Let&#8217;s get started.</p><h2>Preparing the dataset</h2><p>We&#8217;ll use the <a href="https://archive.ics.uci.edu/dataset/242/energy+efficiency">Energy Efficiency dataset</a> from UCI with 768 rows and 8 building features. The task is to predict heating and cooling loads (regression). First, let&#8217;s load the data and prepare it for modeling:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;d6085535-fe94-494c-9bbf-d6457c553987&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">from ucimlrepo import fetch_ucirepo
import pandas as pd
from sklearn.model_selection import train_test_split

energy = fetch_ucirepo(id=242)
X = energy.data.features
y = energy.data.targets["Y1"]  # Heating Load

X.columns = ["Relative Compactness", "Surface Area", "Wall Area", "Roof Area",
    "Overall Height", "Orientation", "Glazing Area", "Glazing Area Distribution"]

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

print(f"Train: {X_train.shape[0]} rows | Test: {X_test.shape[0]} rows")
X_train.head()</code></pre></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!KZzX!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!KZzX!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png 424w, /__u/substackcdn.com/image/fetch/$s_!KZzX!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png 848w, /__u/substackcdn.com/image/fetch/$s_!KZzX!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KZzX!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!KZzX!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png" width="716" height="259.1565934065934" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/61e27d59-1b35-4704-9651-3860d427c125_1624x588.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:527,&quot;width&quot;:1456,&quot;resizeWidth&quot;:716,&quot;bytes&quot;:112003,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/188903001?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!KZzX!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png 424w, /__u/substackcdn.com/image/fetch/$s_!KZzX!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png 848w, /__u/substackcdn.com/image/fetch/$s_!KZzX!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png 1272w, /__u/substackcdn.com/image/fetch/$s_!KZzX!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61e27d59-1b35-4704-9651-3860d427c125_1624x588.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h2>Fitting TabICLv2</h2><p>Well, fitting isn&#8217;t the right word, but most tabular foundation models emulate the sklearn API. &#8220;Fitting&#8221; the TabICLv2 regression model means downloading ~110MB of model weights from Hugging Face on first use. After that, it&#8217;s cached locally. The fit step also loads the model weights and pre-processes the data, like standardizing the features.</p><p>Let&#8217;s have a look at how this works and how long it takes. For comparison, I also trained a random forest. By the way, I am doing everything on my MacBook Air M1, no GPU. Not ideal for tabular foundation models, a GPU would be faster.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;a560de66-6491-4227-a0f2-9a4262cc7018&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import time
from tabicl import TabICLRegressor
from sklearn.ensemble import RandomForestRegressor

tabicl = TabICLRegressor(n_estimators=8, device="cpu", random_state=42)
rf = RandomForestRegressor(random_state=42)

for name, model in [("TabICL", tabicl), ("RandomForest", rf)]:
    t0 = time.time()
    model.fit(X_train, y_train)
    print(f"{name:&lt;12} fit: {time.time() - t0:.2f}s")</code></pre></div><p><em>TabICL       fit: 0.28s</em></p><p><em>RandomForest fit: 0.08s</em></p><p>Both are quite fast. While the random forest was faster, the .fit() step is where traditional tabular ML is costly for larger datasets, while TFMs should still be fast since no training happens here.</p><h2>Predicting with TabICLv2</h2><p>Let&#8217;s make predictions with TabICLv2 and evaluate performance.</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;33713b4d-ea03-4bb3-b8e5-82383bace2c7&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import numpy as np
from sklearn.metrics import root_mean_squared_error, mean_absolute_error
import matplotlib.pyplot as plt

results = {}
for name, model, kwargs in [
    ("TabICL",      tabicl, {"output_type": "mean"}),
    ("RandomForest", rf, {}),
]:
    t0 = time.time()
    y_pred = model.predict(X_test, **kwargs)
    results[name] = y_pred
    rmse = root_mean_squared_error(y_test, y_pred)
    mae = mean_absolute_error(y_test, y_pred)
    print(f"{name:&lt;12}  RMSE={rmse:.3f}  MAE={mae:.3f}  predict={time.time()-t0:.2f}s")</code></pre></div><p><em>TabICL             RMSE=0.424  MAE=0.304  predict=2.84s</em></p><p><em>RandomForest  RMSE=0.490  MAE=0.354  predict=0.01s</em></p><p>The Random Forest shows the usual pattern: Per data row, prediction is much cheaper, aka faster than training. For tabular foundation models, the predict step is the expensive one. And we can see it&#8217;s comparably slow. However, the out-of-the-box performance is much better than that of the random forest. This is by no means a real benchmark (e.g., no hyperparameter tuning). It shows, however, something I observed a few times now:  You can just throw a TFM at these tabular tasks, and it just works.</p><h2>Quantifying uncertainty with prediction intervals</h2><p>Predicting the heating load is a regression task.</p><p>I find it interesting how the <a href="https://team.inria.fr/soda/">Soda lab at INRIA</a> architected and pre-trained TabICLv2 for regression:</p><ul><li><p>TabICLv2 for regression is pre-trained separately from the classification model.</p></li><li><p>TabICLv2 is pre-trained to predict 999 quantiles using the (aggregated) pinball loss, while TabPFN treats regression as bin classification.</p></li><li><p>To predict the mean, the 999 quantile predictions are averaged.</p></li><li><p>We can also extract any quantile: 5%, median, 95%, &#8230;</p></li></ul><p>Since we get the entire predictive distribution, we can quantify uncertainty by outputting prediction intervals instead of point predictions:</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;e962378b-28af-4db8-a775-abd519eddb75&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">quantiles = tabicl.predict(X_test, output_type="quantiles", alphas=[0.05, 0.95])
coverage = np.mean((y_test.values &gt;= quantiles[:, 0]) &amp; (y_test.values &lt;= quantiles[:, 1]))
print(f"Empirical 90% coverage: {coverage:.1%}")</code></pre></div><p><em>Empirical 90% coverage: 90.9%</em></p><p>A well-calibrated 90% interval should contain roughly 90% of true values. An empirical coverage of 90.9% is pretty good and well-calibrated. However, this is just one anecdote. I also tried TabICL on the w<a href="/__u/mindfulmodeler.substack.com/p/how-to-win-an-ml-competition-beyond?utm_source=publication-search">atersupply competition I won</a>: The performance, out-of-the-box, was slightly worse than my XGBoost ensemble (but still very good). However, the coverage of the prediction interval, which should have been 90%, was something like 60%. To be fair, I had the same problem with the XGBoost ensemble, which I fixed using <a href="/__u/mindfulmodeler.substack.com/p/week-3-conformal-prediction-for-regression?utm_source=publication-search">conformal prediction</a>. The same fix is available for TFMs.</p><h2>Interpretability with SHAP</h2><p>Tabular foundation models are not inherently interpretable models, given that they are deep transformer-based neural networks, meaning we have to use additional methods to interpret these models, such as Shapley values (ShAP):</p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;python&quot;,&quot;nodeId&quot;:&quot;49f9aabb-f73b-4e88-98af-ff3f1596b9c6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-python">import shap

X_test_sample = X_test.iloc[:30]

def predict_fn(x):
    return tabicl.predict(pd.DataFrame(x, columns=X_train.columns), output_type="mean")
 background=/__u/mindfulmodeler.substack.com/shap.sample(X_train, 50, random_state=42)
explainer = shap.PermutationExplainer(predict_fn, background)

# TabICL PermutationExplainer
t0 = time.time()
shap_values = explainer(X_test_sample)
print(f"SHAP completed in {time.time() - t0:.1f}s")

# RF TreeExplainer
t0 = time.time()
tree_explainer = shap.TreeExplainer(rf)
tree_explainer(X_test_sample)
print(f"TreeExplainer    (RF): {time.time() - t0:.3f}s")

# RF PermutationExplainer
t0 = time.time()
perm_explainer_rf = shap.PermutationExplainer(rf.predict, background)
perm_explainer_rf(X_test_sample)
print(f"PermutationExplainer (RF): {time.time() - t0:.1f}s")

shap.summary_plot(shap_values, X_test_sample, show=False)</code></pre></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="/__u/substackcdn.com/image/fetch/$s_!G6_z!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="/__u/substackcdn.com/image/fetch/$s_!G6_z!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png 424w, /__u/substackcdn.com/image/fetch/$s_!G6_z!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png 848w, /__u/substackcdn.com/image/fetch/$s_!G6_z!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png 1272w, /__u/substackcdn.com/image/fetch/$s_!G6_z!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_webp, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png 1456w" sizes="100vw"><img src="/__u/substackcdn.com/image/fetch/$s_!G6_z!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png" width="576" height="342.33962264150944" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:441,&quot;width&quot;:742,&quot;resizeWidth&quot;:576,&quot;bytes&quot;:48840,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://mindfulmodeler.substack.com/i/188903001?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="/__u/substackcdn.com/image/fetch/$s_!G6_z!, /__u/mindfulmodeler.substack.com/w_424, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png 424w, /__u/substackcdn.com/image/fetch/$s_!G6_z!, /__u/mindfulmodeler.substack.com/w_848, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png 848w, /__u/substackcdn.com/image/fetch/$s_!G6_z!, /__u/mindfulmodeler.substack.com/w_1272, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png 1272w, /__u/substackcdn.com/image/fetch/$s_!G6_z!, /__u/mindfulmodeler.substack.com/w_1456, /__u/mindfulmodeler.substack.com/c_limit, /__u/mindfulmodeler.substack.com/f_auto, /__u/mindfulmodeler.substack.com/q_auto:good, /__u/mindfulmodeler.substack.com/fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fc026a7-189a-45e2-a617-0cc6fec10444_742x441.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em>SHAP completed in 2496.5s</em></p><p><em>TreeExplainer (RF): 0.139s</em></p><p><em>PermutationExplainer (RF): 7.8s</em></p><p><strong>Ouch. Big ouch.</strong></p><p>Especially since I&#8217;ve already severely dialed down the computational burden: I&#8217;ve reduced the background dataset to just 50 data points, and we are only computing the Shapley values for 30 of the 154 test data points. Again, it&#8217;s on a CPU, so a bit of a caveat for my anecdote here. But practically, it&#8217;s my current reality, as I&#8217;m used to running smaller data science projects on my MacBook.</p><p>The huge SHAP computation time gap between TabICL and the Random Forest exists because of two reasons:</p><ul><li><p>For TFMs, prediction is expensive compared to traditional tabular ML.</p></li><li><p>Model-agnostic methods, especially Shapley values, require many model.predict() calls. The only reason why Shapley values are reasonable for traditional tabular ML is that we have a fast model-specific estimator for tree-based models, which dominate(d) tabular ML.</p></li></ul><p>But not all is lost: In theory, with some adaptations, it should be possible to use something like the GradientExplainer for TFMs, a model-specific version for gradient-based models. I&#8217;ve also seen that TabPFN uses a different approach to estimating Shapley values by leveraging the in-context learning, as described <a href="https://arxiv.org/pdf/2403.10923">in this paper</a> (yet another paper in my ever-growing paper stack &#128579;).</p><p>So I guess we will have to see some optimization taking place to use the good old model-agnostic interpretability methods, but hopefully they shall be usable again. At the same time, I expect tabular foundation models to become more efficient. Also, maybe I&#8217;ll have to start using a GPU &#129335;&#8205;&#9794;&#65039;.</p>]]></content:encoded></item></channel></rss>