Scalable Distributed Deep Learning for Lung Disease Diagnosis: A Spark–Elephas Cross-Modality Framework with GPU Cluster Scalability Analysis

Authors

  • Zeineb Ben Messaoud Digital Research Center of Sfax, Laboratory of Signals, SysteMs, Artificial Intelligence and Networks (SM@RTS), Sfax, Tunisia
  • Maher Baccar Higher Institute of Computer Science and Multimedia Gabes, Gabes, Tunisia
  • Heni Bouhamed Advanced Technologies for Image and Signal Processing Unit (ATISP), Sfax University, Sfax. Tunisia Research and Development Department, Zetta-Spark, Sfax, Tunisia
  • Monia Hamdi Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia

DOI:

https://doi.org/10.15837/ijccc.2026.5.7566

Keywords:

Distributed transfer learning, Multi-class lung disease diagnosis, Apache Spark– Elephas, Cross-modality medical imaging, Parameter-server architecture, Synchronous vs. asynchronous training, Scalability characterization

Abstract

The rapid growth of large-scale medical imaging datasets has exposed critical limitations of single-node deep learning workflows for clinical decision support in lung disease diagnosis. We present a unified, Spark-based distributed framework for scalable multi-class lung disease classification from both chest X-ray (CXR) and computed tomography (CT) images. Our approach integrates Apache Spark with Elephas to enable synchronous and asynchronous data-parallel training of Keras models via a parameter-server architecture on a GPU-enabled cluster, supporting binary and multi-class classification across multiple pulmonary pathologies including COVID-19 pneumonia, viral pneumonia, lung opacity, and normal cases. We implement transfer learning with DenseNet169, ResNet50, and MobileNetV2 and evaluate three classification tasks: binary lung disease detection and three-class diagnosis on CXR images, and three-class diagnosis on CT scans. Experiments on large public datasets demonstrate that the proposed system achieves high diagnostic performance, reaching 97.63% accuracy for three-class CXR and 95.49% for binary CT classification, while providing substantial runtime gains over single-node baselines. On a multi-node Spark cluster, synchronous training attains near-linear to super-linear scaling with speedups up to 6.26× and parallel efficiency exceeding 156%, attributed to improved cache utilization and reduced I/O contention across distributed workers. Comparative analysis consistently demonstrates that synchronous parameter updates outperform asynchronous training in both convergence stability and final diagnostic accuracy across all tasks and modalities. This work delivers an end-to-end, cross-modality, distributed deep learning pipeline with explicit scalability characterization. The proposed framework is generalizable beyond pulmonary imaging and demonstrates how cluster computing infrastructures can be leveraged to build scalable, high-throughput experimental pipelines for multi-class lung disease classification in research settings.

Author Biography

Zeineb Ben Messaoud, Digital Research Center of Sfax, Laboratory of Signals, SysteMs, Artificial Intelligence and Networks (SM@RTS), Sfax, Tunisia



References

Alistarh, D.; Grubic, D.; Li, J.; Tomioka, R.; Vojnovic, M. (2017). QSGD: Communicationefficient SGD via gradient quantization and encoding, Advances in Neural Information Processing Systems, 30, 1709-1720.

Awan, M.J.; Bilal, M.H.; Yasin, A.; Nobanee, H.; Khan, N.S.; Zain, A.M. (2021). Detection of COVID-19 in chest X-ray images: A big data enabled deep learning approach, International Journal of Environmental Research and Public Health, 18(19), 10147. https://doi.org/10.3390/ijerph181910147

Benbrahim, H.; Hachimi, H.; Amine, A. (2020). Deep transfer learning with Apache Spark to detect COVID-19 in chest X-ray images, Romanian Journal of Information Science and Technology, 23, 117-129.

Benbrahim, H.; Hachimi, H.; Amine, A. (2021). Deep transfer learning pipelines with Apache Spark and Keras TensorFlow combined with logistic regression to detect COVID-19 in chest CT images, Walailak Journal of Science and Technology, 18(11), 13109. https://doi.org/10.48048/wjst.2021.13109

Buda, M.; Maki, A.; Mazurowski, M.A. (2018). A systematic study of the class imbalance problem in convolutional neural networks, Neural Networks, 106, 249-259. https://doi.org/10.1016/j.neunet.2018.07.011

Dean, J.; Corrado, G.; Monga, R.; et al. (2012). Large scale distributed deep networks, Advances in Neural Information Processing Systems, 25, 1223-1231.

Elephas: Distributed Deep Learning with Keras & Spark. https://github.com/maxpumperla/ elephas. Accessed: 2025-03-20.

Fawcett, T. (2006). An introduction to ROC analysis, Pattern Recognition Letters, 27(8), 861-874. https://doi.org/10.1016/j.patrec.2005.10.010

He, K.; Zhang, X.; Ren, S.; Sun, J. (2016). Deep residual learning for image recognition, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770-778. https://doi.org/10.1109/CVPR.2016.90

Huang, G.; Liu, Z.; Van Der Maaten, L.; Weinberger, K.Q. (2017). Densely connected convolutional networks, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4700-4708. https://doi.org/10.1109/CVPR.2017.243

Ioffe, S.; Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift, Proceedings of the 32nd International Conference on Machine Learning (ICML), 448-456.

Kingma, D.P.; Ba, J. (2014). Adam: A method for stochastic optimization, arXiv preprint arXiv:1412.6980.

Li, M.; Andersen, D.G.; Park, J.W.; Smola, A.J.; et al. (2014). Scaling distributed machine learning with the parameter server, Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation (OSDI), 583-598.

Lin, Y.; Han, S.; Mao, H.; Wang, Y.; Dally, W.J. (2017). Deep gradient compression: Reducing the communication bandwidth for distributed training, arXiv preprint arXiv:1712.01887.

Mezzoudj, S.; Belkessa, I.; Bouras, F.; Khelifa, M. (2023). A novel distributed deep learning approach for large-scale chest X-ray COVID-19 images detection, Research Square. https://doi. org/10.21203/rs.3.rs-2534755/v1 https://doi.org/10.21203/rs.3.rs-2534755/v1

Mezzoudj, S.; Khelifa, M.; Saadna, Y. (2025). Leveraging Spark-TensorFlow Distributor for distributed deep convolutional neural networks: Accelerating large-scale COVID-19 detection, Informatica, 49, 173-188. https://doi.org/10.31449/inf.v49i17.7505

Micikevicius, P.; Narang, S.; Alben, J.; et al. (2017). Mixed precision training, arXiv preprint arXiv:1710.03740.

Rahman, T.; et al. (2021). COVID-19 radiography database, Kaggle. https://www.kaggle.com/ datasets/tawsifurrahman/covid19-radiography-database

Ravikumar, A.; Sriraman, H. (2023). Real-time pneumonia prediction using pipelined Spark and high-performance computing, PeerJ Computer Science, 9, e1258. https://doi.org/10.7717/peerj-cs.1258

Roldán, A.; Sánchez-Solís, J.; Jiménez, V.; Juárez, R.; Zárate, G. (2020). Convolutional neural network in a pseudo-distributed environment for classification of chest X-ray images of patients with pneumonia, Research in Computing Science, 149, 101-110.

Rubin, G.D.; Ryerson, C.J.; Haramati, L.B.; et al. (2020). The role of chest imaging in patient management during the COVID-19 pandemic: A multinational consensus statement from the Fleischner Society, Radiology, 296(1), 172-180. https://doi.org/10.1148/radiol.2020201365

Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. (2018). MobileNetV2: Inverted residuals and linear bottlenecks, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4510-4520. https://doi.org/10.1109/CVPR.2018.00474

Shvachko, K.; Kuang, H.; Radia, S.; Chansler, R. (2010). The Hadoop distributed file system, Proceedings of the IEEE 26th Symposium on Mass Storage Systems and Technologies (MSST), 1-10. https://doi.org/10.1109/MSST.2010.5496972

Simpson, S.; et al. (2020). Radiological Society of North America expert consensus statement on reporting chest CT findings related to COVID-19, Radiology: Cardiothoracic Imaging, 2(2), e200152. https://doi.org/10.1148/ryct.2020200152

Soares, E.; Angelov, P.; Biaso, S.; Froes, M.H.; Abe, D.K. (2020). SARS-CoV-2 CT-scan dataset: A large dataset of real patients CT scans for SARS-CoV-2 identification, medRxiv.https://www. kaggle.com/datasets/plameneduardo/sarscov2-ctscan-dataset

Sokolova, M.; Japkowicz, N.; Szpakowicz, S. (2009). A systematic analysis of performance measures for classification tasks, Information Processing & Management, 45(4), 427-437. https://doi.org/10.1016/j.ipm.2009.03.002

Vavilapalli, V.K.; Murthy, A.C.; Douglas, C.; et al. (2013). Apache Hadoop YARN: Yet another resource negotiator, Proceedings of the 4th Annual Symposium on Cloud Computing (SOCC), 1-16. https://doi.org/10.1145/2523616.2523633

Wong, H.Y.F.; Lam, H.Y.S.; Fong, A.H.T.; et al. (2020). Frequency and distribution of chest radiographic findings in patients positive for COVID-19, Radiology, 296(2), E72-E78. https://doi.org/10.1148/radiol.2020201160

Zaharia, M.; Chowdhury, M.; Franklin, M.J.; Shenker, S.; Stoica, I. (2010). Spark: Cluster computing with working sets, Proceedings of the 2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud'10).

Zaharia, M.; Chowdhury, M.; Das, T.; Dave, A.; Ma, J.; McCauly, M.; Franklin, M.J.; Shenker, S.; Stoica, I. (2012). Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing, Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation (NSDI'12).

Additional Files

Published

2026-09-01

Most read articles by the same author(s)

Obs.: This plugin requires at least one statistics/report plugin to be enabled. If your statistics plugins provide more than one metric then please also select a main metric on the admin's site settings page and/or on the journal manager's settings pages.