Latest Findings

What's new from the lab

Short, readable summaries of our newest publications — the methods, the findings, and why they matter. Ready to download, read, and cite.

When Does Parameter-Efficient Fine-Tuning Beat Full Fine-Tuning in Code-Change Learning?
IEEE TSE · 2026

When Does Parameter-Efficient Fine-Tuning Beat Full Fine-Tuning in Code-Change Learning?

The first systematic comparison of PEFT (LoRA, adapters, prompt tuning) against full-model fine-tuning across seven code LMs on real code-change tasks — plus PastaK, a probing-guided framework that trains only a fraction of the parameters and still beats full fine-tuning.

Shuo Liu, Jacky W. Keung, Zhi Jin, Zhen Yang, Fang Liu, Hao Zhang Read the summary →
Federated fine-tuning lets companies fix bugs together without sharing code
ACM TOSEM · 2026

Federated fine-tuning lets companies fix bugs together without sharing code

LLM-based automated program repair depends on high-quality code — but most real-world code is proprietary. This study applies federated learning to fine-tune LLMs on a private industrial dataset (TutorCode) for program repair, evaluated on the EvalRepair-Java benchmark, and finds federated fine-tuning significantly improves repair while preserving data privacy.

Wenqiang Luo, Jacky W. Keung, Boyang Yang, He Ye, Claire Le Goues, Tegawendé F. Bissyandé, Haoye Tian, Bach Le Read the summary →
Top models score under 37% on HumanEval-V — visual reasoning, not coding skill, is the bottleneck
ACM TOSEM · 2026

Top models score under 37% on HumanEval-V — visual reasoning, not coding skill, is the bottleneck

HumanEval-V is a benchmark of 253 human-annotated tasks that push large multimodal models to read an algorithm's logic straight from a diagram and emit working Python. Across 27 state-of-the-art LMMs, the best result is only 36.8% pass@1. The authors show the real failure point is high-level visual reasoning, not coding ability.

Fengji Zhang, Linquan Wu, Huiyu Bai, Guancheng Lin, Xiao Li, Xiao Yu, Yue Wang, Bei Chen, Jacky W. Keung Read the summary →
LogRoBERTa: Hybrid Language Model Detects Log Anomalies Without a Parser
Automated Software Engineering · 2026

LogRoBERTa: Hybrid Language Model Detects Log Anomalies Without a Parser

Anomaly detection in software logs usually leans on time-consuming log parsers. LogRoBERTa discards the parser entirely, pairing RoBERTa's contextual embeddings with an attention-based BiLSTM. A Determinantal Point Process (DPP) sampler builds a small, diverse labeled set — just 20,000 logs on BGL (0.42% of the dataset). Across HDFS, BGL, Thunderbird, and Spirit it outperforms state-of-the-art baselines including three fully supervised models, stays strong on low-resource data, and cuts runtime by roughly a third.

Yicheng Sun, Jacky W. Keung, Zhen Yang, Shuo Liu, Hi Kuen Yu Read the summary →
R2ComSync harnesses LLMs to automatically keep code comments in sync
Empirical Software Engineering · 2026

R2ComSync harnesses LLMs to automatically keep code comments in sync

R2ComSync is a Large Language Model (LLM) framework for Code-Comment Synchronization (CCS): automatically rewriting an outdated comment when its code changes. A pilot study showed vanilla LLMs fall short of state-of-the-art CCS methods. R2ComSync closes two gaps: Ensemble Hybrid Retrieval (EHR), which mixes code-comment semantic similarity with code-change-pattern similarity to build instructive in-context examples, and a Multi-turn Re-ranking (MR) strategy that prioritizes the most correct-prone candidates.

Zhen Yang, Hongyi Lin, Xiao Yu, Jacky W. Keung, Shuo Liu, Pak Yuen Patrick Chan, Yicheng Sun, Fengji Zhang Read the summary →
LogMeta: few-shot meta-learning for log anomaly detection that adapts to new systems
Journal of Systems and Software · 2026

LogMeta: few-shot meta-learning for log anomaly detection that adapts to new systems

Deep log-anomaly detectors struggle with heterogeneous log formats and scarce labeled anomalies. LogMeta combines Model-Agnostic Meta-Learning (MAML) with a hybrid language model — RoBERTa for semantics plus Bi-LSTM and attention for sequence dependencies — so it adapts to unseen systems from a few samples.

Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Wenqiang Luo Read the summary →
Benchmark wins don't transfer: a task-oriented re-evaluation of log parsing and anomaly detection on real industrial logs
Information & Software Technology · 2026

Benchmark wins don't transfer: a task-oriented re-evaluation of log parsing and anomaly detection on real industrial logs

This study re-evaluates seven log parsers and five anomaly-detection models on real industrial logs (factory assembly, process control, SCADA) versus a curated benchmark. Nearly everything degrades on industrial data, LLM-based parsers win on accuracy, and improving parsing rarely flows through to better detection — a call to measure the pipeline end-to-end.

Yicheng Sun, Jacky W. Keung, Xiaoxue Ma, Yihan Liao, Zhenyu Mao, Hi Kuen Yu Read the summary →
Bridging the responsibility gap: deriving human-oversight requirements for GenAI-enabled software
Information & Software Technology · 2026

Bridging the responsibility gap: deriving human-oversight requirements for GenAI-enabled software

As generative AI becomes an active component of reasoning and decision support, accountability becomes diffuse: developers can't see how models reach conclusions, users can't fully control behavior, organizations struggle to trace decisions. The authors argue human oversight should be formalized as an explicit category in requirements engineering, and propose a design methodology that identifies responsibility gaps by aligning system-side patterns (control, transparency) with human-side roles (authority, interaction), then derives oversight requirements and documents the reasoning in a reusable Deductive Backbone Table.

Read the summary →
FedDC: dual-layer protection that keeps federated learning accurate against black-box and white-box attacks
Journal of Information Security & Applications · 2026

FedDC: dual-layer protection that keeps federated learning accurate against black-box and white-box attacks

FedDC pairs differential privacy with a chaos-logistic-map scrambler, protecting only sensitive layers (not the whole model). Against black-box theft it drags accuracy to 9–12%; against white-box model inversion it raises reconstruction error dramatically — all with under 2.3–5% per-round overhead.

Yihan Liao, Jacky W. Keung, Jingyu Zhang, Yurou Dai Read the summary →
Computers in Human Behavior · 2026

AI decision support nearly triples fraud-detection odds — but vulnerability is age- and scam-specific

Most fraud research describes WHO is vulnerable via static risk profiles or single-scam studies; this paper tests whether real-time AI support reduces fraud risk at the moment of exposure, and how its effects differ across age and scam type. Grounded in the Person–Task Fit framework, a 2×2 online lab-in-the-field experiment compared AI-assisted vs. unaided judgment across younger (18–30) and older (50+) groups over 10 days using ecologically realistic simulated fraud scenarios across six scam categories. AI-assisted support improved detection accuracy and reduced risky intention, but vulnerability was patterned rather than uniform — and AI produced the largest reductions exactly in each group's most-vulnerable scam categories.

Yicheng Sun, Jacky W. Keung, Hi Kuen Yu, Yuchen Cao Read the summary →
Which test-selection metric actually helps? 1,640 scenarios say: it depends on your objective
FSE 2026 · PACMSE

Which test-selection metric actually helps? 1,640 scenarios say: it depends on your objective

This empirical study benchmarks 15 test-selection metrics on 1,640 scenarios spanning three testing objectives, five out-of-distribution shift types, three data modalities (Android, image, text) and 13 Deep Neural Networks. The winners are often not the metrics designed for the job — diversity-based metrics dominate performance estimation.

Jingyu Zhang, Fan Wang, Jacky W. Keung, Yihan Liao, Yan Xiao, Lei Ma Read the summary →