AI Colloquium

ScoringBench: A Benchmark for Evaluating Tabular Foundation Models with Proper Scoring Rules

by Dr Jonas Landsgesell (University of Stuttgart), Pascall Knoll (University of Stuttgart), Dr Tizian Wenzel (Ludwig Maximilian University of Munich, Munich Center for Machine Learning)

Europe/Berlin
JvF25/3-303 - Conference Room (Lamarr/RC Trust Dortmund)

JvF25/3-303 - Conference Room

Lamarr/RC Trust Dortmund

30
Show room on map
Description

Abstract:

While select tabular foundation models (TFMs), such as TabPFN and TabICL, are capable of producing full predictive distributions, prevailing benchmarks often rely exclusively on point-estimate metrics like RMSE. This focus rewards conditional mean accuracy while potentially ignoring the valuable distributional information these models are intended to capture.
To address this evaluation gap, we introduce ScoringBench, an open and extensible benchmark designed to evaluate tabular regression models using a comprehensive suite of proper scoring rules alongside standard point metrics.
Key Findings and Implications:

  • Metric Sensitivity: Across diverse model families, we observe that rankings shift substantially depending on the chosen metric. Models that perform well on point metrics often rank poorly when evaluated through probabilistic lenses.
  • Inductive Biases: Different proper scoring rules impose distinct inductive biases during training. Even when theoretically minimized by the true distribution, the choice of loss function directly influences the model’s behavior.
  • Fine-tuning Potential: We demonstrate that fine-tuning with scoring rules that were unseen during the pretraining phase can lead to improvements in the corresponding metrics.


These results suggest that the choice of metric is a critical modeling decision. We argue that benchmarks should prioritize the reporting of distributional metrics, and that consideration should be given to how a model's training objective aligns with the specific requirements of the downstream decision problem. Ultimately, choosing the loss function is an integral part of choosing the model itself.

Organised by

Matthias Feurer