Selecting Probability Calibrators for Tabular Classification

Authors

  • Haojie Li Southwest University, Chongqing, 400000, China

Keywords:

Probability Calibration, Tabular Classification, Post-Hoc Calibration, Model Selection, Openml

Abstract

Reliable probabilities matter in risk-sensitive tabular decision making. In credit scoring, fraud detection, medical triage, and review prioritization, confidence quality often matters as much as label accuracy. Yet post-hoc calibrators such as temperature scaling, sigmoid scaling, isotonic regression, Dirichlet calibration, and beta calibration do not behave uniformly across model families, class structures, calibration-set sizes, or evaluation protocols. Rather than searching for a universally best calibrator, this work treats tabular probability calibration as a selection problem. Under a strictly separated Train / Cal / Select / Test protocol, uncalibrated outputs and candidate calibrators are compared on 20 OpenML datasets, three baseline model families, and three random seeds. Selection on the held-out Select split uses a guarded policy, the Phase-8 Guarded Calibration Selector (GCS-8), with fallback conditions and isotonic-specific constraints. Across the multiseed benchmark, the frozen selector reduces mean test log-loss from 0.332157 to 0.298072; On held-out validation datasets, mean test log-loss reaches 0.189940. Gains persist under class imbalance and small calibration splits, while isotonic emerges as the least stable option in low-data calibration regimes and beta provides a useful binary-only extension. Overall, tabular probability calibration is better framed as choosing among candidate post-hoc calibrators under a strict protocol than as seeking a single globally best method.

Downloads

Published

2026-07-12

How to Cite

Li, H. (2026). Selecting Probability Calibrators for Tabular Classification. CPS Digital Library - Series of Conferences, 2, 229–236. Retrieved from https://seriesofconference.com/index.php/SCJ/article/view/295