Latest Findings · ASI Lab

Benchmark wins don't transfer: a task-oriented re-evaluation of log parsing and anomaly detection on real industrial logs

Across four datasets and five detectors, parser accuracy only weakly predicts downstream anomaly-detection gains — the authors argue for treating log analysis as an end-to-end task.

Yicheng Sun, Jacky W. Keung, Xiaoxue Ma, Yihan Liao, Zhenyu Mao, Hi Kuen Yu
Agentic Software Intelligence Research Lab · Department of Computer Science, City University of Hong Kong
Information & Software Technology · 2026

This study re-evaluates seven log parsers and five anomaly-detection models on real industrial logs (factory assembly, process control, SCADA) versus a curated benchmark. Nearly everything degrades on industrial data, LLM-based parsers win on accuracy, and improving parsing rarely flows through to better detection — a call to measure the pipeline end-to-end.

Industrial log analysis revisited: A task-oriented evaluation of parsing and anomaly detection under real-world constraints
Fig. 2. Time expenditure of seven parsers on four full datasets.

Why it matters

Log parsing and anomaly detection are widely studied, but whether improved parsing quality translates into downstream anomaly-detection gains in real industrial settings remains unclear.

Key findings

Task-oriented industrial log analysis

What it means for practitioners

Prioritize semantic fidelity over maximal abstraction when choosing a parser; allocate effort to high-impact error correction; treat annotation as an ongoing operational process; design for explainability and human-in-the-loop monitoring.

Get the paper & cite it

Download full PDF ↓ View at DOI

Official citation: Yicheng Sun, Jacky W. Keung, Xiaoxue Ma, Yihan Liao, Zhenyu Mao, Hi Kuen Yu (2026). Industrial log analysis revisited: A task-oriented evaluation of parsing and anomaly detection under real-world constraints. Information and Software Technology. DOI: 10.1016/j.infsof.2026.108205.

Log parsing Log anomaly detection Industrial logs Semantic fidelity Label efficiency Empirical software engineering Semi-supervised learning LLM-based parsing