Use Bayesian Paired Tests with a ROPE to Improve the Comparison of Machine Learning Models

Chris Williams

Video

Paper PDF

Thumbnail of paper pages

Abstract

This tutorial paper argues that model comparison in machine learning can be much improved by using \emph{paired testing}, i.e.\ comparing the predictions of methods A and B on each (common) test example. Due to the limitations of null hypothesis significance testing, a Bayesian approach is recommended, including the use of the region of practical equivalence (ROPE; Kruschke 2015a; Kruschke and Liddell 2018; Benavoli, Corani, Dem\v{s}ar, and Zaffalon 2017). We discuss a Bayesian $t$-test and a Bayesian McNemar test for comparisons on a single task, and Bayesian hierarchical models for comparisons over multiple tasks. Three worked examples are presented to illustrate the methods, and the use of reporting guidelines is discussed as a potential means of changing current practice.