Benchmark wins don't transfer: a task-oriented re-evaluation of log parsing and anomaly detection on real industrial logs
Across four datasets and five detectors, parser accuracy only weakly predicts downstream anomaly-detection gains — the authors argue for treating log analysis as an end-to-end task.
This study re-evaluates seven log parsers and five anomaly-detection models on real industrial logs (factory assembly, process control, SCADA) versus a curated benchmark. Nearly everything degrades on industrial data, LLM-based parsers win on accuracy, and improving parsing rarely flows through to better detection — a call to measure the pipeline end-to-end.
Why it matters
Log parsing and anomaly detection are widely studied, but whether improved parsing quality translates into downstream anomaly-detection gains in real industrial settings remains unclear.
Key findings
- Benchmark ≠ industrial All parsers and detectors degrade substantially on FALL/PIMS/SCADA-X versus the curated LHSB subset; benchmark results do not reliably transfer to industrial logs.
- LLM parsers win on accuracy LLM-based LUNAR best overall — FTA 0.828 (LHSB), 0.797 (FALL), 0.719 (SCADA-X); worst was rule-based IPLoM (0.329 SCADA-X, 0.362 PIMS). Rule-based Drain/IPLoM are far faster.
- Best detector SemiRALD outperformed all baselines (F1: LHSB 0.985, FALL 0.920, PIMS 0.921, SCADA-X 0.912); others dropped sharply on industrials (e.g. PLELog 0.615, LogBERT 0.755 on SCADA-X).
- Annotation scarcity hurts At 10–50% labels performance degrades sharply; semi-supervised models consistently beat the supervised baseline under scarce labels.
- Parsing quality ≠ detection gain Correcting parser errors only helps when the correction preserves anomaly-critical cues (Error Code Removed Δ=86.9%, Sensor Value Lost Δ=72.8%); global average ΔF1 gains were small (max +0.023).
Task-oriented industrial log analysis
- Four datasets: LHSB benchmark subset plus three industrial collections — FALL (factory assembly line), PIMS (process industry), SCADA-X (energy/SCADA).
- Seven parsers (4 rule-based + 3 LLM-based) and five detectors (4 semi-supervised + 1 supervised) evaluated under standardized conditions.
- Five research questions spanning generalization, label efficiency, dominant error patterns, and whether improved parsing transfers to detection.
- Expert annotation (4 domain experts, two-round cross-validation) and interviews with seven industrial maintenance engineers.
What it means for practitioners
Prioritize semantic fidelity over maximal abstraction when choosing a parser; allocate effort to high-impact error correction; treat annotation as an ongoing operational process; design for explainability and human-in-the-loop monitoring.
Get the paper & cite it
Official citation: Yicheng Sun, Jacky W. Keung, Xiaoxue Ma, Yihan Liao, Zhenyu Mao, Hi Kuen Yu (2026). Industrial log analysis revisited: A task-oriented evaluation of parsing and anomaly detection under real-world constraints. Information and Software Technology. DOI: 10.1016/j.infsof.2026.108205.