hn.today

TabPFN and TabICL vs. tuned XGBoost: the model that doesn't train won 14/14

efraingaray.com23 points12 comments
Screenshot of TabPFN and TabICL vs. tuned XGBoost: the model that doesn't train won 14/14

Commenters debated a blog post that claimed TabPFN/TabICL (tabular foundation models) beat tuned XGBoost across 14 datasets. rdedev argued tabular foundation models can be surprisingly effective, citing drug-property prediction where they and molecular foundation models approach state-of-the-art. Others echoed enthusiasm for trying foundation models but criticized the post’s breathless, AI-generated prose (icfly2, bronlund, NordStreamYacht), with bronlund calling it “AI slop” and noting the author’s defensiveness. opensandwich and not_a_feature urged broader, more rigorous comparisons and pointed to alternative benchmarks (TabBench-Bio, TabArena) as better evaluation grounds.

Skeptics focused on methodology and generalizability. lyelibi and 3abiton said transformers rarely beat XGBoost/CatBoost in industrial settings with very large datasets and that feature engineering and LightGBM remain powerful for time series. 3eb7988a1663 questioned fairness: XGBoost was optimized for accuracy rather than AUC, training-time accounting looked odd, and the tabular model’s pretraining on public benchmarks plus lack of released scripts reduced credibility. icfly2 and others noted that big-production datasets are often trimmed and split, so results on small benchmarks may not reflect real-world workflows. The divide centers on whether foundation models truly generalize beyond curated benchmarks and whether the evaluation and presentation in the post were rigorous.

Read on efraingaray.com12 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.