DeLong test
Also known as: AUC comparison · DeLong et al. 1988 · roc.test
A statistical test that tells whether the difference between two AUCs measured on the same base is real or chance. It is the standard way to compare the power of two scores without assuming a distribution.
Legal basis
DeLong, DeLong & Clarke-Pearson (1988), Biometrics
When two scores are evaluated on the same population, their AUCs are not independent — the same cases enter both. The DeLong test handles that dependence and returns a p-value and a confidence interval for the difference between the two areas, without assuming the scores follow any particular distribution.
It is the right test for the incremental lift question: does the model with the new score order better than the model without it, or does the difference fit inside the noise? Proposed by DeLong, DeLong and Clarke-Pearson in 1988, it remains the reference for paired comparison of AUCs.
Frequently asked questions
What is the DeLong test for?
To compare the ordering power (AUC) of two scores evaluated on the same base and tell whether the difference between them is statistically significant, while handling the correlation between the two measurements.