User Guide#

Introduction to Imbalanced Learning#

Class imbalance occurs when, for a classification problem, one target variable class is much less frequent than the other(s). This phenomenon creates challenges when training classifiers, as models tend to learn more from the majority class, leading to poor performance on the minority class. With smaller datasets and/or very high class imbalance, there can also simply be too few samples in the minority class to learn from. Typical solutions are either data-based or algorithmic.

Data-based techniques involve balancing the distribution of the target variable with undersampling, oversampling, or a mixture of the two. The imbalanced-learn library implements many of these methods whilst integrating with the sklearn API. See the imbalanced-learn user guide for further information on the implemented resampling techniques.

Algorithmic techniques involve training models with respect to a modified loss function, typically increasing the loss associated with misclassifying the minority class. In particular, the weighted cross-entropy loss is implemented through the class_weight parameter in sklearn classifiers, and scale_pos_weight parameter in the xgboost and lightgbm packages.

The Calibration Problem#

The main drawback to these methods is that they introduce bias. For both data-based and algorithmic methods, models are not calibrated to the original data distribution, and instead the new distribution induced by resampling or weighting the loss function. Resulting probability estimates thus cannot be interpreted directly, and probability-based evaluation metrics become inaccurate.

Fortunately, in the case of random undersampling, random oversampling, or the weighted cross-entropy loss, an analytical correction exists which transforms model output back to the true data distribution. Although this correction is well-documented in the literature, confusion can arise due to differing notation and parameterisations. Manually applying the correction reduces reproducibility and can create confusion, for example when saving or sharing a model object. Furthermore, for some parameterisations (such as sklearn’s class_weight), the correction is a function of both the training data and parameter value, creating an additional potential source of error.

To address this issue, imbalanced-calibrate provides the PriorCalibratedClassifier class, which acts as a calibration wrapper (similar to sklearn.calibration.CalibratedClassifierCV) around any sklearn-compatible binary classifier. In particular, the implementation automatically calculates the correct correction by inspecting the sub-estimator’s (or imbalanced-learn resampler’s) parameters.

imbalanced-calibrate API#

PriorCalibratedClassifier is fully integrated within the sklearn API, acting as as Meta-Estimator and Classifier object. As with all sklearn estimators, it implements the fit, predict_proba and predict methods.

The PriorCalibratedClassifier takes an (untrained) classifier and optional weight parameter at instantiation, for example:

>>> from imbcalibrate import PriorCalibratedClassifier
>>> from sklearn.linear_model import LogisticRegression

>>> clf = PriorCalibratedClassifier(LogisticRegression(class_weight='balanced'))

Calling clf.fit(X, y) fits the sub-estimator to the data, and infers the weight from the sub-estimator’s parameters (and training data when applicable) when possible. If weight is provided at instatiation, it overrides any inferred weight. Currently weight inference is implemented for the following estimator objects (in order of priority):

  • imblearn.pipeline.Pipeline: weight is inferred from the first instance of RandomOverSampler or RandomUnderSampler encountered in the pipeline steps.

  • sklearn.pipeline.Pipeline: weight is inferred from the last step of the pipeline, provided it is as classifier (also applies to imblearn.pipeline.Pipeline if no sampler is found).

  • Any sklearn classifier implementing the class_weight parameter.

  • xgboost.XGBClassifier, lightgbm.LGBClassifier, or any similar estimator which implements the scale_pos_weight parameter.

Once fitted, calls to clf.predict_proba(X) and clf.predict(X) use the calibration-corrected probability estimates. Crucially, this behaviour persists with model object saving/loading, as the weight is set and saved at fit time. The fitted sub-estimator can be accessed through the estimator_ attribute.

Mathematical Formulation#

Weighted Cross-Entropy Loss#

Let \(Y \sim \mathbb{P}\) be a binary random variable taking values 0 and 1, with \(\pi_0 = \mathbb{P}(Y=0)\) and \(\pi_1 = \mathbb{P}(Y=1)\). Recall the binary cross-entropy loss is given by

\[L(y, \hat{p}) = -y\log{\hat{p}} - (1-y)\log{(1-\hat{p})},\]

for a probability estimate \(\hat{p}\). W.l.o.g., we assume that \(Y=1\) is the minority class, in which case the weighted cross-entropy loss can be defined as

\[L^w(y, \hat{p}) = -y w \log{\hat{p}} - (1-y)\log{(1-\hat{p})},\]

for some weight \(w > 1\). Now, we have

\[\begin{split}\mathbb{E}_\mathbb{P}[L^w(y, \hat{p})] &= \mathbb{P}(Y=1)\mathbb{E}[L^w(1, \hat{p})|Y=1] + \mathbb{P}(Y=0)\mathbb{E}[L^w(0, \hat{p})|Y=0] \\ &= w \pi_1 \mathbb{E}[L(1, \hat{p})|Y=1] + \pi_0 \mathbb{E}[L(0, \hat{p})|Y=0] \\ &\propto \mathbb{P}_w(Y=1)\mathbb{E}[L(1, \hat{p})|Y=1] + \mathbb{P}_w(Y=0)\mathbb{E}[L(0, \hat{p})|Y=0] \\ &= \mathbb{E}_{\mathbb{P}_w}[L(y,\hat{p})],\end{split}\]

where \(\mathbb{P}_w\) is an implicit probability measure induced by \(w\), with \(\mathbb{P}_w(Y=0)=\pi_0/z\) and \(\mathbb{P}_w(Y=1)=(w\pi_1)/z\), where \(z = \pi_0 + w\pi_1\) is a normalising constant. We assume that \(\mathbb{P}(\cdot|Y) = \mathbb{P}_w(\cdot|Y)\), since the weight is entirely determined by the label \(Y\).

Notice that the prior odds induced by the weighting are given by

\[O_w = \frac{w \pi_1 / z}{\pi_0 / z} = w \frac{\pi_1}{\pi_0} = w \cdot O.\]

That is, minimising a weighted loss (with weight \(w\)) under the true distribution \(\mathbb{P}\) is proportional to minimising an unweighted loss under the artificial data distribution \(\mathbb{P}_w\), for which the odds of the positive class are multiplied by \(w\).

Resampling Methods#

In fact, there is a one-to-one relationship between the changes in prior odds induced by the weighted loss and resampling methods. The methods below rely on the assumption that the conditional distributions are not affected by resampling, i.e. \(P(\cdot|Y, S) = P(\cdot|Y)\), where \(S\) is a random variable indicating whether an observation is included in the resampled dataset. In practice, this assumption only holds for random oversampling and random undersampling .

Oversampling#

Let \(N_0\) and \(N_1\) be the number of samples of each class in the dataset, so that the empirical probability distribution is given by \(\hat{\mathbb{P}}(Y=1) = N_1/(N_0+N_1)\), with empirical odds ratio \(\hat{O}=N_1/N_0\). Suppose we oversample with rate \(w > 1\), so that the oversampled dataset has \(wN_1\) samples of the positive/minority class. The new empirical odds ratio is given by

\[\hat{O}_w = \frac{w N_1}{N_0} = w \hat{O}.\]

That is, oversampling with rate \(w\) results in the same shift in prior odds ratio as using the weight \(w\) in a weighted cross-entropy loss.

Undersampling#

Similarly, suppose we undersample with rate \(1/w\), so that the undersampled dataset has \(N_0/w\) samples of the negative/majority class. We obtain the same shift in odds ratio:

\[\hat{O}_w = \frac{N_1}{N_0/w} = w \hat{O}.\]

Imbalanced-Learn Parameterisation#

In imbalanced-learn samplers, the sampling rates are specified by the sampling_strategy parameter, which is defined as “the desired ratio of the number of samples in the minority class over the number of samples in the majority class after resampling.” That is, the sampling_strategy parameter corresponds to the desired odds ratio \(\hat{O}_w\). To recover the weight, we simply have

\[w = \hat{O}_w / \hat{O} = \frac{N_1^s}{N_0^s} \cdot \frac{N_0}{N_1},\]

where \(N_0^s\) and \(N_1^s\) are the number of samples in the negative and positive classes respectively in the resampled dataset.

Applying the Correction#

We have established the equivalence in odds ratio shift between the weighted loss, over- and under-sampling in terms of a weight w. It remains to transform the estimated (uncalibrated) probability estimates \(p_w\) to the true data distribution.

\[\begin{split}O_w &= w \cdot O \\ \frac{p_w}{1-p_w} &= w \cdot \frac{p}{1-p} \\ &\vdots \\ p &= \frac{\frac{p_w}{w(1-p_w)}}{1 + \frac{p_w}{w(1-p_w)}} \\ &= \frac{p_w}{w(1-p_w) + p_w}\end{split}\]

References & Further Reading#

Caplin, A., Martin, D., and Marx, P. (2022). Calibrating for Class Weights by Modeling Machine Learning. arXiv:2205.04613 [cs.LG].

Chawla, N. V., Japkowicz, N., and Kotcz, A. (2004). Editorial: special issue on learning from imbalanced data sets. SIGKDD Explor. Newsl., 6(1):1–6.

Elkan, C. (2001). The foundations of cost-sensitive learning. In Proceedings of the 17th international joint conference on Artificial intelligence - Volume 2, IJCAI’01, pages 973–978, San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.

He, H. and Ma, Y. (2013). Imbalanced learning : foundations, algorithms, and applications. Wiley, Hoboken, NJ.