Latest Findings · ASI Lab

When Does Parameter-Efficient Fine-Tuning Beat Full Fine-Tuning in Code-Change Learning?

A systematic study across four PEFT methods and seven code-language models — and PastaK, a probing-guided framework that trains a fraction of the parameters and still wins.

Shuo Liu · Jacky W. Keung · Zhi Jin · Zhen Yang · Fang Liu · Hao Zhang
Agentic Software Intelligence Research Lab · Department of Computer Science, City University of Hong Kong
IEEE Transactions on Software Engineering · vol. 52, no. 1, pp. 3–21

Full-model fine-tuning (FMFT) is the gold standard for adapting a pre-trained code model to a downstream task — but it is expensive. Parameter-Efficient Fine-Tuning (PEFT) methods, which update only 0.01–5% of the weights, have often beaten FMFT on static code tasks such as code search and clone detection. Yet almost every prior result was on tasks where the input is a single unchanged code snippet. What about code-change tasks, where a model must read a diff — old code and new code side by side? This work is the first to test that question at scale.

Three code intelligence tasks: code summarization (static), commit message generation (dynamic), and just-in-time defect prediction (dynamic). The study moves from static comprehension to tasks that read code diffs.
Code comprehension tasks come in two flavours. Static tasks (a) read a single snippet; code-change tasks (b, c) read a diff. This study is the first large-scale PEFT evaluation on the dynamic ones.

What the researchers measured

The team compared four mainstream PEFT methods — Adapter Tuning (AT), LoRA, Prompt Tuning (PT), and Prefix Tuning (PreT) — against FMFT across seven popular pre-trained code models (CodeBERT, GraphCodeBERT, PLBART, UniXcoder, CodeT5, CodeT5+, Qwen2.5-Coder) on two widely used code-change tasks: Just-In-Time Defect Prediction (JIT-DP, is this commit buggy?) and Commit Message Generation (CMG, write a message for this diff). Datasets include JIT-Defects4J and the multi-language MCMD corpus across Java, C#, C++, Python and JavaScript.

Key findings

Why probing explains the gap

To explain why the methods differ, the authors designed two new probing tasks (Code Change Match and Line Type Prediction) alongside a static type-detection probe, and read the encoded semantics layer by layer. The result is a clear picture: PEFT better preserves a model's pre-trained dynamic code semantics in the layers where FMFT overfits, while FMFT is stronger in the bottom layers. The two are complementary — not rivals.

PastaK: probe once, tune better

Elite findings from the probing experiments, the team proposes PastaK (self-adaptive efficient layer-specific tuning). Instead of all-or-nothing, PastaK looks at the layer-wise probing results, then assigns the bottom K layers to full fine-tuning and the rest to a PEFT method. Because this allocation is read directly from the probe data, the framework adapts itself to each task and dataset.

On commit-message generation, PastaK surpasses every PEFT method and beats full FMFT by 1.48% (BLEU), 3.21% (METEOR) and 1.87% (ROUGE-L) — while saving 26.26% of training time and 20.65% of memory.

Why this matters to practitioners

The takeaway is not "PEFT is always better" — it is when. For classification-style code-change tasks (defect prediction, and transfer/new-language scenarios with scarce data), PEFT gives better accuracy at a fraction of the cost and should replace FMFT. For demanding generative tasks, use AT or LoRA when compute is limited, and consider PastaK when you want full-model quality without the full-model bill. The authors also open-source the replication package so results are reproducible.

Get the paper & cite it

Download full PDF ↓ View at DOI

Official citation: Liu, S., Keung, J. W., Jin, Z., Yang, Z., Liu, F., & Zhang, H. (2026). An Empirical Study of Parameter-Efficient Fine-Tuning in Code Change Learning and Beyond. IEEE Transactions on Software Engineering, 52(1), 3–21. DOI: 10.1109/TSE.2025.3637335.

@article{liu2026peft, title = {An Empirical Study of Parameter-Efficient Fine-Tuning in Code Change Learning and Beyond}, author = {Liu, Shuo and Keung, Jacky W. and Jin, Zhi and Yang, Zhen and Liu, Fang and Zhang, Hao}, journal = {IEEE Transactions on Software Engineering}, volume = {52}, number = {1}, pages = {3--21}, year = {2026}, doi = {10.1109/TSE.2025.3637335} }
LoRAAdapter TuningPrompt TuningJust-in-Time Defect PredictionCommit Message GenerationPastaKParameter-Efficient Fine-TuningCode Language Models