面向隐私保护及数据共享的多医院共建危重症预测联邦学习模型

    A Federated Learning Model for Critical Illness Prediction Built by Multiple Hospitals for Privacy Protection and Data Sharing

    • 摘要: 本研究聚焦于危重症患者临床预测与资源调度中的核心矛盾:单一医院数据有限,难以训练出泛化能力强的可靠模型;而直接共享原始医疗数据又面临严格的隐私与合规限制。针对这一问题,构建基于联邦学习的双任务预测系统,在隐私保护的前提下,实现多中心协同建模;在死亡风险预测任务中,对比多种数据增强方法与机器学习模型,筛选出适用于医疗不平衡数据分类任务的最优组合;在住院天数预测任务中,依托联邦学习框架聚合多家医院分布式医疗知识,构建全局预测模型,并通过多维度指标与各医院本地模型、集中式模型进行对比验证。实验结果表明,新方法有效实现了数据隐私保护与多中心协作建模的兼顾,预测精度与集中式训练接近。在死亡风险预测中,ADASYN-MLP组合展现出显著优势,其构建的联邦学习全局模型在F1分数(0.30)和召回率(0.76)上均优于单一医院本地模型,有效解决了危重症“稀有事件”带来的样本不平衡问题;在住院天数预测方面,联邦学习全局模型在平均绝对误差(MAE = 7.13)和决定系数(R2 = 0.69)指标上显著优于各医院本地模型,展现出更优的预测精度与泛化能力。

       

      Abstract: This study focuses on the core contradiction in clinical prediction and resource allocation for critically ill patients: data from a single hospital is limited, making it difficult to train reliable models with strong generalization ability; meanwhile, directly sharing raw medical data faces strict privacy and compliance restrictions. To address this issue, we first constructed a federated learning-based dual-task prediction system to achieve multi-center collaborative modeling while protecting privacy. In the mortality risk prediction task, we compared various data augmentation methods and machine learning models to select the optimal combination suitable for imbalanced medical data classification tasks. In the hospital length-of-stay prediction task, we leveraged the federated learning framework to aggregate distributed medical knowledge from multiple hospitals, building a global prediction model, and validated it against local models at each hospital and centralized models using multidimensional metrics. Experimental results show that the new method effectively balances data privacy protection with multi-center collaborative modeling, achieving predictive accuracy close to centralized training. For mortality risk prediction, the ADASYN-MLP combination demonstrated significant advantages, with the federated learning global model achieving higher F1 score (0.30) and recall rate (0.76) than single-hospital local models, effectively addressing the sample imbalance problem caused by the 'rare event' nature of critical illness. For hospital length-of-stay prediction, the federated learning global model significantly outperformed local hospital models in mean absolute error (MAE = 7.13) and coefficient of determination (R2 = 0.69), demonstrating superior predictive accuracy and generalization ability.

       

    /

    返回文章
    返回