When Does Cost-Sensitive Weighting Matter? A Classifier-Capacity Anal-ysis for Imbalanced Classification

Authors

  • Swati Satpute Research Scholar, Dept. of Computer Science, Yashwantrao Mohite College of Arts, Science and Commerce, Bharati Vidyapeeth (Deemed to be University), Pune, Maharashtra, India Author
  • Ajit More Professor, Institute of Management and Entrepreneurship Development, Bharati Vidyapeeth (Deemed to be University), Pune, Maharashtra, India Author

DOI:

https://doi.org/10.47392/IRJAEH.2026.0677

Keywords:

Imbalanced classification, cost-sensitive learning, class weighting, classifier capacity, ensemble learning, XGBoost, Friedman test, Matthews correlation coefficient

Abstract

Class weighting is the most widely used cost-sensitive remedy for class imbalance, and inverse-frequency “balanced” weights are often applied as an unquestioned default. This paper asks whether the choice among competing weighting mechanisms actually matters, and for which classifiers. Seven mechanisms, namely Balanced, Cost-Matrix, Sqrt-InverseFreq, Log-Damped, Effective-Number, a Search-Tuned power scheme, and focal loss, were evaluated on twenty-four benchmark datasets whose imbalance ratios reach 85:1, under three base classifiers of increasing representational capacity: logistic regression, a soft-voting ensemble of logistic regression, random forest, and a support vector machine, and XGBoost. Using leakage-free nested cross-validation and the Friedman–Nemenyi procedure applied separately to each classifier, mechanism choice produced large, highly significant differences under logistic regression (p < 0.0001), differences that shrank to non-significance under the voting ensemble (p = 0.179) and became indistinguishable from a random ordering under XGBoost (p = 0.955, average-rank spread of 0.46). At the same time, the two high-capacity classifiers outperformed logistic regression under every mechanism tested. The evidence supports a simple two-tier policy: tune the weighting scheme carefully for linear learners, and keep the plain balanced default for ensemble and boosting learners, where additional tuning yields no benefit that survives a significance test.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-01

How to Cite

When Does Cost-Sensitive Weighting Matter? A Classifier-Capacity Anal-ysis for Imbalanced Classification. (2026). International Research Journal on Advanced Engineering Hub (IRJAEH), 4(08), 5174-5186. https://doi.org/10.47392/IRJAEH.2026.0677