Federated fine-tuning lets companies fix bugs together without sharing code
An empirical study shows LLMs fine-tuned with federated learning on private industrial code rival centralized training — lifting program repair by up to 16.67% Top@10 and 18.44% Pass@10 while keeping the data private.
LLM-based automated program repair depends on high-quality code — but most real-world code is proprietary. This study applies federated learning to fine-tune LLMs on a private industrial dataset (TutorCode) for program repair, evaluated on the EvalRepair-Java benchmark, and finds federated fine-tuning significantly improves repair while preserving data privacy.
Why it matters
We investigate federated learning as a privacy-preserving method for fine-tuning LLMs on proprietary and decentralized data, using the private industrial dataset TutorCode and the EvalRepair-Java benchmark.
Key findings
- Big repair gains Federated fine-tuning improves program repair by up to +16.67% for Top@10 and +18.44% for Pass@10, rivaling — and in some cases exceeding — centralized learning.
- Code heterogeneity barely matters The negligible impact of code heterogeneity implies industries with very different codebases can still collaborate effectively under federated learning.
- Algorithm choice is model-dependent Different federated algorithms show unique strengths across LLMs; FedAvg performs best overall and personalization aids some models.
- Privacy-preserving collaboration Only lightweight adapters (LoRA/QLoRA) are exchanged between clients and server — not full models or raw data — dramatically cutting communication cost.
ParamEfficient Federated Fine-Tuning
- Fine-tunes LLMs on the private industrial TutorCode dataset (1,239 buggy C++ programs, 427 developers) with federated learning, evaluated on EvalRepair-Java (163 bugs) benchmark.
- Uses parameter-efficient fine-tuning (LoRA / QLoRA with 4-bit NF4 quantization) so clients upload only small adapters, not full models.
- Server aggregates adapter parameters (e.g. FedAvg) and broadcasts the global model back to clients — preserving raw-data privacy.
- Studies the impact of heterogeneous code (style, complexity, embedding) and multiple federated algorithms across several open LLMs.
What it means for practitioners
If you hold proprietary code but want LLM repair that competes with central training, federated learning is a practical path — you contribute adapters, not code. Code heterogeneity is a minor concern, and picking the right federated algorithm for your model matters more than the privacy setup.
Get the paper & cite it
Official citation: Wenqiang Luo, Jacky W. Keung, Boyang Yang, He Ye, Claire Le Goues, Tegawendé F. Bissyandé, Haoye Tian, Bach Le (2026). When Fine-Tuning LLMs Meets Data Privacy: An Empirical Study of Federated Learning in LLM-Based Program Repair. ACM Transactions on Software Engineering and Methodology. DOI: 10.1145/3733599.