Latest Findings · ASI Lab

Which test-selection metric actually helps? 1,640 scenarios say: it depends on your objective

The largest controlled comparison of 15 test-selection metrics across 3 testing objectives, 5 out-of-distribution shifts, 3 data modalities and 13 DNNs reveals surprising winners and losers.

Jingyu Zhang, Fan Wang, Jacky W. Keung, Yihan Liao, Yan Xiao, Lei Ma
Agentic Software Intelligence Research Lab · Department of Computer Science, City University of Hong Kong
FSE 2026 · PACMSE

This empirical study benchmarks 15 test-selection metrics on 1,640 scenarios spanning three testing objectives, five out-of-distribution shift types, three data modalities (Android, image, text) and 13 Deep Neural Networks. The winners are often not the metrics designed for the job — diversity-based metrics dominate performance estimation.

Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts
Fig. 1. Overview of our study: 1. Test Input Preparation ... 2. Test Input Selection ... 3. Evaluation: fault detection (#Mis. and #Clu.), performance estimation (AE%), retraining guidance (Acc.%).

Why it matters

Metrics that encourage test input diversity are effective in performance estimation, outperforming those specifically designed for this objective, and metrics designed for retraining guidance are also outperformed by other metrics.

Key findings

Empirical test-selection benchmark

What it means for practitioners

Don't default to the metric named for your goal. For accurate performance estimation, use diversity-based metrics regardless of budget; for fault detection pick high-surprise inputs for classifiers and broader neuron coverage for regressors; and reserve diverse inputs for retraining when the budget is loose.

Get the paper & cite it

Download full PDF ↓ View at DOI

Official citation: Jingyu Zhang, Fan Wang, Jacky W. Keung, Yihan Liao, Yan Xiao, Lei Ma (2026). Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts. Proceedings of the ACM on Software Engineering (FSE 2026). DOI: 10.1145/3797086.

Test selection Deep learning testing Out-of-distribution Neuron coverage Surprise adequacy Fault detection Performance estimation Retraining guidance