Hierarchical Multimodal Fusion of Multi-Sequence MRI and Clinical Metadata for the Classification of Rotator Cuff Tears


AŞIK S., YAZICI A., AŞCI M., Okumuşer İ.

Journal of Clinical Medicine, cilt.15, sa.14, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 15 Sayı: 14
  • Basım Tarihi: 2026
  • Doi Numarası: 10.3390/jcm15145525
  • Dergi Adı: Journal of Clinical Medicine
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Chemical Abstracts Core, EMBASE, Academic Search Ultimate (EBSCO), Health Research Premium Collection (ProQuest)
  • Anahtar Kelimeler: clinical metadata, convolutional neural network, deep learning, multi-sequence MRI, multimodal fusion, rotator cuff tear, sequence selection, shoulder MRI, vision transformer
  • Eskişehir Osmangazi Üniversitesi Adresli: Evet

Özet

Background/Objectives: Rotator cuff tears are a leading cause of shoulder disability. While multi-sequence MRI is standard, the optimal deep learning integration of heterogeneous image series and clinical metadata remains unresolved. This study evaluated a hierarchical, sequence-aware multimodal framework for patient-level binary rotator cuff tear classification. Methods: A single-center cohort of 199 patients (100 tears, 99 controls) was analyzed across four MRI sequences (T1 coronal, T2 fat-suppressed sagittal, and proton density [PD] fat-suppressed coronal and transverse/axial) and nine demographic features. Under a patient-level stratified three-fold cross-validation scheme preventing data leakage, we evaluated ResNet50 and Vision Transformer baselines (Study 0), full-protocol fusion topologies (Study 1), and systematically mapped sequence-subset combinations with or without metadata (Study 2). Results: In Study 0, the PD coronal ResNet50 model was the top baseline (AUC = 0.9834, F1 = 0.9515). In Study 1, late decision fusion yielded the highest AUC (0.9909), while feature concatenation optimized threshold balance (F1 = 0.9502). In Study 2, a streamlined three-sequence subset with metadata (C14M: T2 + PDc + PDt) achieved peak performance (AUC = 0.9961, 95% CI: 0.9823–0.9987, F1 = 0.9618, MCC = 0.9238), outperforming the full protocol (AUC = 0.9909, F1 = 0.9355). Metadata utility was configuration-dependent, assisting only fluid-sensitive combinations. Conclusions: Rather than indiscriminately aggregating entire clinical protocols, multimodal fusion is optimized by selecting complementary imaging series. For binary classification, excluding non-fat-suppressed T1 images in favor of a streamlined T2 and PD set stabilized by clinical demographics maximized classification performance in this internally validated, single-center cohort.