Which test-selection metric actually helps? 1,640 scenarios say: it depends on your objective
The largest controlled comparison of 15 test-selection metrics across 3 testing objectives, 5 out-of-distribution shifts, 3 data modalities and 13 DNNs reveals surprising winners and losers.
This empirical study benchmarks 15 test-selection metrics on 1,640 scenarios spanning three testing objectives, five out-of-distribution shift types, three data modalities (Android, image, text) and 13 Deep Neural Networks. The winners are often not the metrics designed for the job — diversity-based metrics dominate performance estimation.
Why it matters
Metrics that encourage test input diversity are effective in performance estimation, outperforming those specifically designed for this objective, and metrics designed for retraining guidance are also outperformed by other metrics.
Key findings
- Diversity beats intent Metrics that encourage test diversity (Rand, GD, STD) outperform metrics (CES, PACE, EST, DR) specifically designed for performance estimation; EST (22.68%) and CES (10.03%) give very inaccurate estimates.
- Fault detection picks Recommend DSA (high-surprise inputs) for classification and NC (broader neuron coverage) for regression.
- Retraining guidance KMNC and PACE give the highest retraining improvement for classification and regression respectively, outpacing metrics (MCP, DAT) built for this objective.
- OOD-stable, budget-sensitive Metric performance is largely stable across OOD scenarios, and across budgets for objectives 1 and 2 — but fluctuates considerably for retraining guidance (objective 3).
- Efficiency trade-off A clear performance-speed trade-off exists for fault-detection metrics, but for performance estimation and retraining guidance the best metrics are also fast.
Empirical test-selection benchmark
- 15 metrics (2014–2023) across six families — uncertainty, diversity, surprise, sampling, clustering, hybrid — on 1,640 controlled scenarios.
- Three data modalities (Android via AndroZoo, images via MNIST/Udacity, text via IMDb), five OOD shift types, 13 DNNs.
- Three testing objectives: fault detection (#Mis., #Clu.), performance estimation (AE%), retraining guidance (Acc.%).
- Four selection budgets per scenario; results reported with rigorous validity-threat mitigation (3 runs, optimal hyperparameters).
What it means for practitioners
Don't default to the metric named for your goal. For accurate performance estimation, use diversity-based metrics regardless of budget; for fault detection pick high-surprise inputs for classifiers and broader neuron coverage for regressors; and reserve diverse inputs for retraining when the budget is loose.
Get the paper & cite it
Official citation: Jingyu Zhang, Fan Wang, Jacky W. Keung, Yihan Liao, Yan Xiao, Lei Ma (2026). Empirical Insights of Test Selection Metrics under Multiple Testing Objectives and Distribution Shifts. Proceedings of the ACM on Software Engineering (FSE 2026). DOI: 10.1145/3797086.