Latest Findings · ASI Lab

Federated fine-tuning lets companies fix bugs together without sharing code

An empirical study shows LLMs fine-tuned with federated learning on private industrial code rival centralized training — lifting program repair by up to 16.67% Top@10 and 18.44% Pass@10 while keeping the data private.

Wenqiang Luo, Jacky W. Keung, Boyang Yang, He Ye, Claire Le Goues, Tegawendé F. Bissyandé, Haoye Tian, Bach Le
Agentic Software Intelligence Research Lab · Department of Computer Science, City University of Hong Kong
ACM TOSEM · 2026

LLM-based automated program repair depends on high-quality code — but most real-world code is proprietary. This study applies federated learning to fine-tune LLMs on a private industrial dataset (TutorCode) for program repair, evaluated on the EvalRepair-Java benchmark, and finds federated fine-tuning significantly improves repair while preserving data privacy.

When Fine-Tuning LLMs Meets Data Privacy: An Empirical Study of Federated Learning in LLM-Based Program Repair
Fig. 1. The workflow of federated learning of parameter-efficient federated fine-tuning framework.

Why it matters

We investigate federated learning as a privacy-preserving method for fine-tuning LLMs on proprietary and decentralized data, using the private industrial dataset TutorCode and the EvalRepair-Java benchmark.

Key findings

ParamEfficient Federated Fine-Tuning

What it means for practitioners

If you hold proprietary code but want LLM repair that competes with central training, federated learning is a practical path — you contribute adapters, not code. Code heterogeneity is a minor concern, and picking the right federated algorithm for your model matters more than the privacy setup.

Get the paper & cite it

Download full PDF ↓ View at DOI

Official citation: Wenqiang Luo, Jacky W. Keung, Boyang Yang, He Ye, Claire Le Goues, Tegawendé F. Bissyandé, Haoye Tian, Bach Le (2026). When Fine-Tuning LLMs Meets Data Privacy: An Empirical Study of Federated Learning in LLM-Based Program Repair. ACM Transactions on Software Engineering and Methodology. DOI: 10.1145/3733599.

Program repair Federated learning Large language models Fine-tuning Data privacy LoRA Automated program repair