Stabilization Uplift
Metrics for evaluating how stable a classifier stays under a sudden distribution shift, such as a macroeconomic shock.
- Stabilization Score (SS): how much one model's ROC AUC changed between the base (pre-shock) and shock periods, relative to the severity of the shift. In [0.5, 1]; 1 means no change.
- Stabilization Uplift (SU): whether model B (e.g. trained with synthetic outliers) is more stable and not worse than model A (a baseline) under the shock. In [0, 1); 0 means no uplift.
Calculator
Model A (baseline)
Model B (e.g. with synthetic outliers)
Compute it from data with
stabilization_uplift.distribution_shift (pip install "stabilization-uplift[shift]").SS of model A
–
SS of model B
–
SU of B over A
–
Use in code
Python package (PyPI), with an end-to-end example on the dataset and the API reference on its page:
pip install stabilization-uplift # SS and SU
pip install "stabilization-uplift[shift]" # + distribution_shift computed from data
from stabilization_uplift import stabilization_score, stabilization_uplift
stabilization_score(auc_base=0.80, auc_shock=0.70, dist_shift=0.2) # 0.915
stabilization_uplift(auc_base_A=0.80, auc_shock_A=0.70,
auc_base_B=0.80, auc_shock_B=0.81, dist_shift=0.2) # 0.725
Hugging Face evaluate (this Space is also an evaluate module):
pip install evaluate stabilization-uplift
import evaluate
metric = evaluate.load("zyplai/stabilization-uplift")
metric.compute(auc_base_A=[0.80], auc_shock_A=[0.70],
auc_base_B=[0.80], auc_shock_B=[0.81], dist_shift=[0.2])
# {'stabilization_score_A': [0.915...], 'stabilization_score_B': [0.992...], 'stabilization_uplift': [0.725...]}
Each row is one comparison of model B with model A, so several comparisons can be computed in one call.