Multiple Testing for IR and Recommendation System Experiments


Ngozi Ihemelandu and Michael D. Ekstrand. 2024. Multiple Testing for IR and Recommendation System Experiments. To appear as a short paper in Proceedings of the 46th European Conference on Information Retrieval (ECIR '24). Proc. ECIR '24. Acceptance rate: 24.3%.

This paper was led by my Ph.D student Ngozi Ihemelandu.


While there has been significant research on statistical techniques for comparing two information retrieval (IR) systems, many IR experiments test more than two systems. This can lead to inflated false discoveries due to the multiple-comparison problem (MCP). A few IR studies have investigated multiple comparison procedures; these studies mostly use TREC data and control the familywise error rate. In this study, we extend their investigation to include recommendation system evaluation data as well as multiple comparison procedures that controls for False Discovery Rate (FDR).

