R2ComSync harnesses LLMs to automatically keep code comments in sync
An in-context-learning approach that retrieves the right demonstrations and re-ranks candidates — lifting Accuracy by up to 694% over state-of-the-art baselines.
R2ComSync is a Large Language Model (LLM) framework for Code-Comment Synchronization (CCS): automatically rewriting an outdated comment when its code changes. A pilot study showed vanilla LLMs fall short of state-of-the-art CCS methods. R2ComSync closes two gaps: Ensemble Hybrid Retrieval (EHR), which mixes code-comment semantic similarity with code-change-pattern similarity to build instructive in-context examples, and a Multi-turn Re-ranking (MR) strategy that prioritizes the most correct-prone candidates.
Why it matters
Code-Comment Synchronization aims to synchronize comments with code changes in an automated fashion, reducing developer workload during software maintenance; prior approaches lack generalization or need extensive resources, motivating an LLM-based solution.
Key findings
- LLMs underperform unless guided In zero- and two-shot settings, state-of-the-art baselines beat LLMs by wide margins (49.00%–463.67% in Accuracy on Liu's dataset); only meaningful demonstrations plus candidate re-ranking close the gap.
- Big gains across languages and models R2ComSync raises Accuracy by 188.98%–694.08%, Recall@5 by 83.65%–337.92%, and ESS Ratio by 40.50%–331.00% across three datasets (two Java, one Python); Llama3-70B is best overall, beating SOTA by 142.10% (Panth's), 166.98% (Pai's), and 31.64% (Liu's) on Accuracy.
- Both modules contribute Ablations show EHR beats CodeBERT-only, expert-only, and random retrieval, while the MR rules add 15.71%–46.65% Accuracy — though Rule 1 mainly helps on Liu's dataset (Rules 2 and 3 generalize better).
- Extends to small models R2ComSync lifts Qwen2.5-1.5B by 76.77% in Accuracy and Qwen2.5-Coder-1.5B by 213.73%, showing the approach also works on resource-limited LLMs.
R2ComSync
- Ensemble Hybrid Retrieval (EHR): retrieves top-(P/2) demonstrations from a semantic pool (CodeBERT-encoded old code, old comment, new code) and top-(P/2) from a change-pattern pool (eleven expert features), concatenated into balanced, instructive ICL prompts.
- Exclusive hard-prompt template: system message + instruction + placeholders filled with retrieved demonstrations (END_OF_DEMO delimiter) plus the CCS target, fed to the LLM to generate candidate comments.
- Multi-turn Re-ranking (MR): three heuristics from 80,591 CCS samples — Rule 1 (weak) updates a mentioned function name; Rule 2 (medium) limits new-sub-token ratio; Rule 3 (severe) limits comment edit-distance ratio — applied weak-to-severe to push violators back and prioritize correct-prone candidates.
What it means for practitioners
R2ComSync is a zero-training LLM pipeline for keeping comments in sync with code changes. Re-tune its published thresholds (σ=0.35; ϵ=0.25 on Liu's, 0.55 on Panth's, 0.2 on Pai's) per dataset; weight expert-based change-pattern retrieval more on short, high-variability samples and CodeBERT semantic retrieval on longer ones; pick the shot number per model/dataset since performance peaks then declines. Cost is acceptable — most models finish a synchronization in under a second with zero training cost.
Get the paper & cite it
Official citation: Zhen Yang, Hongyi Lin, Xiao Yu, Jacky W. Keung, Shuo Liu, Pak Yuen Patrick Chan, Yicheng Sun, Fengji Zhang (2026). R2ComSync: improving code-comment synchronization with in-context learning and reranking. Empirical Software Engineering. DOI: 10.1007/s10664-025-10800-4.