Reviewer #2 Questions 1. Brief summary of the paper The paper proposes a survey of Federated learning for LLMs (FedLLM), focusing on two key aspects: fine-tuning and prompt learning in a federated setting. It also discusses potential directions for federated LLMs, including pre-training, federated agents, and LLMs for federated learning. 2. List three, or more, strong aspects of this paper. Please number each point. Extensive and recent coverage of the FedLLM literature. Clear categorization of methods and research topics. Timely discussion of emerging directions such as federated pre-training and LLM-based agents. 3. List three, or more, weak aspects of this paper. Please number each point. Limited critical analysis Lack of a unifying framework The novelty with respect to prior surveys should be more sharply articulated 4. Detailed comments to the authors. The paper proposes a survey of Federated learning for LLMs (FedLLM), focusing on two key aspects: fine-tuning and prompt learning in a federated setting. It also discusses potential directions for federated LLMs, including pre-training, federated agents, and LLMs for federated learning. The topic is timely and relevant, and the manuscript includes an extensive and up-to-date set of references, as well as a clear categorization of methods and research themes. However, the survey is largely descriptive and tends to enumerate existing works rather than offering a deeper critical analysis or synthesis. The lack of a unifying conceptual framework makes the paper closer to a structured bibliography than an analytical survey, and important aspects such as evaluation practices, reproducibility issues, and the comparability of experimental results across studies are only marginally discussed. In addition, the heavy reliance on arXiv-only references raises concerns regarding the maturity of some of the cited results. Further, the claimed novelty with respect to existing surveys should be articulated more clearly. Overall, while the paper provides a useful overview of the FedLLM landscape, it would benefit from stronger critical insight, clearer positioning with respect to prior surveys, and a more explicit discussion of open challenges and research trade-offs. 5. Overall Recommendation Weak Accept: Borderline paper, tending to accept Reviewer #3 Questions 1. Brief summary of the paper This paper offers a comprehensive survey of Federated Learning for Large Language Models, tackling the key challenge of balancing LLMs’ data-intensive training with privacy protection. Federated Learning enables decentralized collaboration without raw data sharing, reducing risks in sensitive fields but bringing unique issues like data and model heterogeneity, high communication and computational costs, and convergence instability. The survey addresses gaps in existing literature by covering security and privacy, communication and computation efficiency, prompt tuning, real-world applications, and emerging directions. It first focuses on federated fine-tuning, exploring heterogeneity, privacy and security, efficiency optimization, diverse frameworks, and evaluation methods. Next, the paper reviews prompt learning in Federated Learning, a cost-effective approach that fine-tunes soft prompts instead of full models. Key themes include prompt generation, few-shot scenarios, chain-of-thought reasoning, personalization, multi-domain adaptation, efficiency improvements, and black-box settings, along with applications across multiple domains. Finally, it outlines future directions: real-world deployment of personalized models on confidential data, multimodality model co-optimization, efficient federated pre-training, federated AI agents for local inference and cross-client collaboration, and leveraging LLMs to enhance Federated Learning while addressing ethical and legal concerns. Overall, the survey synthesizes latest FedLLM methodologies, highlights the role of fine-tuning and prompt learning, and provides a research roadmap, serving as a valuable resource for interdisciplinary researchers and practitioners. 2. List three, or more, strong aspects of this paper. Please number each point. 1.This paper systematically covers the key aspects of federated large language models, including federated fine-tuning, prompt learning, practical applications, and emerging directions, providing comprehensive coverage of critical research areas and filling gaps in the existing literature. 2.The paper is clearly structured and highly logical, with each chapter containing coherent subtopics, making it easier to understand and grasp the complexities of interdisciplinary fields. 3.The paper not only integrates the latest methodologies but also focuses on showcasing diverse practical applications in areas such as multilingual processing, recommendation systems, medical visual question answering, and weather forecasting, while also outlining specific directions for future research, making it highly relevant and informative in practice. 4. The paper comprehensively explores the core challenges of integrating federated learning with large language models, such as data and model heterogeneity, high communication costs, and privacy risks, and provides a detailed overview of various effective solutions, enhancing a deep understanding of the current progress in this field. 3. List three, or more, weak aspects of this paper. Please number each point. 1.Although the paper proposes numerous solutions to core challenges such as heterogeneity and efficiency, it fails to provide a detailed comparative evaluation of specific performance, including the advantages and disadvantages and applicable scenarios, and lacks quantitative data to support it. 2. Although the paper outlines potential areas for future research, it does not delve into the specific technical bottlenecks, implementation challenges, or key issues that need to be addressed for each direction. It merely provides a general description without offering actionable research paths or an in-depth analysis of potential obstacles. 3. The paper rarely addresses specific practical challenges, such as hardware resource limitations, cross-platform compatibility, lack of industry standards, and compliance issues, providing limited guidance for implementing federated large language models in real-world scenarios. 4. Lack of targeted analysis for different application domains although the paper highlights applications across multiple fields it does not distinguish the unique characteristics challenges and differentiated requirements of federated large language models in specific domains leading to overly generalized discussions that fail to provide precise and actionable insights for domain specific research and practice. 4. Detailed comments to the authors. 1.To address the lack of in-depth comparative analysis in the paper, it is recommended to supplement it with a quantitative comparative evaluation of the proposed solutions. The comparison dimensions can include training efficiency, communication overhead, level of privacy protection, convergence speed, and adaptability to heterogeneous data. For each core challenge, such as heterogeneity or efficiency, relevant methods should be compared in a structured manner, including specific performance metrics from experiments. Additionally, the paper can also include discussions on the trade-offs between various solutions, such as some methods improving efficiency at the expense of partial performance, thereby providing a more comprehensive reference for decision-making. 2.For papers that only superficially discuss emerging directions, it is recommended to focus on specific technical bottlenecks, implementation difficulties, and key issues, and to conduct an in-depth analysis of each potential research area. For example, when discussing federated pretraining, the core challenges of data exchange protocols should be elaborated, such as how to balance communication efficiency and data utility under limited bandwidth. In the area of federated AI agents, the technical obstacles among multiple client agents in designing coordination mechanisms, credit allocation, and communication efficiency should be analyzed. In addition, preliminary feasible research directions should be provided, such as proposing possible technical frameworks or key algorithm improvement directions for each emerging direction. 3.Regarding the issue in the paper about insufficient coverage of practical deployment obstacles, it is recommended to expand the discussion on specific real-world challenges and corresponding coping strategies. For hardware resource constraints, the paper should provide a detailed explanation of the computing and memory limitations of edge devices in cross-device scenarios, and analyze how to optimize model structures or training strategies to accommodate these limitations. For cross-platform compatibility, the paper should explore the technical difficulties of adapting federated large language model frameworks to different deep learning platforms and device systems, and propose potential standardization suggestions. Regarding regulatory compliance, the paper should discuss the specific requirements of data privacy laws in different regions for federated learning, and how to design models and training processes to meet these requirements. 4. To address the issue of the paper's lack of targeted analysis for different application domains, it is recommended to conduct differentiated discussions based on the characteristics of specific fields. For medical applications, the emphasis should be on the extremely high privacy requirements and the heterogeneity of medical data, and an analysis should be made on how to design more targeted privacy protection and data adaptation strategies for these characteristics. For recommendation systems, attention should be paid to the needs for personalization and real-time performance, and discussion should focus on how to optimize federated prompt learning or fine-tuning methods to balance personalization with model generalization. For multilingual processing, the challenges of low-resource languages and cross-lingual data heterogeneity should be explored, and targeted model adaptation solutions proposed. By organizing discussions by domain, the paper can provide more precise and actionable guidance, avoiding content that is overly generalized. 5. Overall Recommendation Weak Reject: Borderline paper, tending to reject Reviewer #4 Questions 1. Brief summary of the paper This paper presents a survey on the intersection of Federated Learning and Large Language Models, termed FedLLM. The authors aim to address the privacy and communication challenges inherent in training LLMs by leveraging FL. The survey is structured around two primary methodologies: Federated Fine-Tuning (including parameter-efficient methods like LoRA) and Prompt Learning. It categorizes existing literature based on heterogeneity, privacy, efficiency, and frameworks. Furthermore, the paper outlines emerging directions, and the use of LLMs to assist the FL process itself (e.g., synthetic data generation). The authors contrast their work with existing surveys by highlighting their focus on real-world applications and emerging agent-based paradigms. 2. List three, or more, strong aspects of this paper. Please number each point. 1. The paper addresses a rapidly evolving field and includes a significant number of recent references (2023-2024). 2. The separation of the survey into "Federated Fine-Tuning" and "Prompt Learning" is a logical and helpful distinction. The detailed tables (Table 2 and Table 3) provide a concise and useful overview of the literature, mapping specific papers to sub-problems like heterogeneity, privacy, and efficiency. 3. Unlike some standard surveys that stop at current methods, Section 4 explores forward-looking topics such as "Federated AI Agents" and "LLMs for FL". This adds value by pointing researchers toward under-explored avenues. 3. List three, or more, weak aspects of this paper. Please number each point. 1. The text frequently adopts a "listing" approach (e.g., "Paper A does X. Paper B does Y.") without providing a deep comparative analysis. There is little discussion on the trade-offs between these methods, contradictory results in the literature, or specific guidelines on which method to choose for a given constraint set. 2. The paper relies entirely on text and tables. A survey on complex architectures and frameworks significantly benefits from diagrams. For example, figures illustrating the difference between Federated Prompt Tuning and Federated Fine-Tuning, or the architecture of split learning in LLMs, are notably missing. 3. The survey's coverage of 2025 literature is sparse (only ~3 references noted). 4. Detailed comments to the authors. 1. The current reference list is heavily dominated by 2023 and 2024 papers. The lack of representation for 2025 research is a major gap. Please conduct a thorough literature search for 2025 to include the latest advancements. 2. Many sections read like an annotated bibliography. Please synthesize the information. 3. There is little discussion on how these models are evaluated. A deeper discussion on evaluation challenges in a federated setting might be needed. 5. Overall Recommendation Weak Accept: Borderline paper, tending to accept Reviewer #5 Questions 1. Brief summary of the paper This paper presents a survey on the integration of Federated Learning (FL) with Large Language Models (FedLLM). The authors identify that while LLMs are powerful, their application is often hindered by privacy concerns and data availability, roblems which FL aims to solve. 2. List three, or more, strong aspects of this paper. Please number each point. The paper reviews several techniques, such as LoRa and prompt learning, to address data heterogeneity, privacy issues and communication efficiency. 3. List three, or more, weak aspects of this paper. Please number each point. 1. Section 2.1 lists various methods for handling heterogeneity (FedDAT, FedLoRA, etc.). However, the paper describes what these methods do without analyzing why one should be preferred over the other in specific scenarios. 2. As authors stated, prompt learning is often a solution to the heterogeneity issues discussed in Section 2. The authors should explicitly discuss how Prompt Learning specifically mitigates the "Data Heterogeneity" challenges mentioned in Section 2.1. 3. The authors should elaborate on the architectural differences between standard FedLLM (training a model) and Federated Agents (collaborating agents). Specifically, how does the communication protocol differ? 4. Section 4.5 "LLMs for Federated Learning" discusses using LLMs to generate synthetic data. This seems like a technique to improve current FL, rather than a "future direction." The authors might consider moving this to the "Efficiency" or "Data Heterogeneity" sections in the main body, as synthetic data generation is already an active area of research. 4. Detailed comments to the authors. See above 5. Overall Recommendation Weak Accept: Borderline paper, tending to accept