LogRoBERTa: Hybrid Language Model Detects Log Anomalies Without a Parser
RoBERTa plus an attention-based BiLSTM and a DPP-selected training set beats state-of-the-art baselines on four benchmark datasets using only 0.42% of logs for training.
Anomaly detection in software logs usually leans on time-consuming log parsers. LogRoBERTa discards the parser entirely, pairing RoBERTa's contextual embeddings with an attention-based BiLSTM. A Determinantal Point Process (DPP) sampler builds a small, diverse labeled set — just 20,000 logs on BGL (0.42% of the dataset). Across HDFS, BGL, Thunderbird, and Spirit it outperforms state-of-the-art baselines including three fully supervised models, stays strong on low-resource data, and cuts runtime by roughly a third.
Why it matters
We propose LogRoBERTa, an innovative anomaly detection model that eliminates the need for a parser. LogRoBERTa creates a stable and diverse labeled training set using the Determinantal Point Process (DPP) method, needing only a small amount of labeled data. The hybrid language model is based on RoBERTa's architecture, combined with an attention-based BiLSTM... Experiments on four widely used datasets demonstrate that LogRoBERTa outperforms state-of-the-art benchmark models—including three fully supervised approaches—without relying on a dedicated log parser.
Key findings
- Parser-free and still state of the art Running with no log parser at all, LogRoBERTa reaches F1 scores of 0.992, 0.991, 0.993, and 0.988 on HDFS, BGL, Thunderbird, and Spirit, outperforming three fully supervised baselines (LogRobust, NeuralLog, HitAnomaly).
- Tiny labeled set, large gains With only 20,000 labeled logs (just 0.42% of BGL), LogRoBERTa beats LogRobust, NeuralLog, and HitAnomaly by 2.2%, 1.2%, and 1.2% on HDFS; 16.1%, 1.4%, and 7.1% on BGL; 30.8%, 3.8%, and 10.8% on Thunderbird; and 4.1%, 7.6%, and 15.1% on Spirit.
- Parsers barely help — and sometimes hurt LogRoBERTa scores almost identically with and without a parser, while some models (PLELog with Drain or Spell, LogBERT with Spell) actually score lower with a parser than without one.
- Fastest of the compared models LogRoBERTa has the lowest runtime of the five models evaluated, averaging 32.9% lower training time and 24.1% lower testing time; omitting the parsing step alone saves 28.3% to 44.1% of runtime.
- Robust in low-resource scenarios Across six datasets its average F1 drops only 8.5% at the smallest scale; on the industrial ILMS and FALL datasets it beats DeepLog, PLELog, and LogRobust by 42.3% and 36.5% at the 10% data scale.
- Diverse training data drives accuracy Template coverage (RT) of the labeled set directly correlates with F1: DPP and k-Center Greedy give the highest coverage and best performance, while stratified random sampling performs worst.
LogRoBERTa
- Uses a Determinantal Point Process (DPP) to select a small, diverse, labeled training set by maximizing the determinant of a cosine-similarity submatrix, minimizing manual labeling effort.
- Encodes raw log sequences with RoBERTa, using dynamic masking and treating each log as a sentence token sequence; the CLS token output becomes the sequence semantic vector.
- Feeds those representations into an attention-based BiLSTM (256 hidden size, 8 attention heads) to capture forward and backward sequential dependencies.
- Focuses the model on key log information with a multi-head (8-head) attention layer before a fully connected output layer with a softmax classifier.
- Operates directly on raw, unparsed, variable-length logs, bypassing the need for a dedicated log parser.
What it means for practitioners
For teams maintaining log-based monitoring, a dedicated log parser is not a hard requirement for accurate anomaly detection. LogRoBERTa shows that a hybrid pre-trained language model plus an attention-based BiLSTM, trained on a small, deliberately diverse subset of logs, can match or beat parser-driven approaches while cutting preprocessing time substantially. The key practical lever is the diversity of the labeled subset (template coverage), not its raw size — so invest annotator effort in a diverse sample (via DPP or k-center greedy) rather than labeling the entire log stream.
Get the paper & cite it
Official citation: Yicheng Sun, Jacky Keung, Zhen Yang, Shuo Liu, Hi Kuen Yu (2026). Improving anomaly detection in software logs through hybrid language modeling and reduced reliance on parser. Automated Software Engineering. DOI: 10.1007/s10515-025-00548-y.