Back to 2026 Abstracts
Seeing the Whole Study: Sequence-Aware Deep Learning for VUR Grading from Complete VCUG Series
HsinHsiao Scott Wang, MD, MPH, MBAn1, Michael Li, MBAn, PhD2, Raymond Bahng, MSc3, Carlos Estrada, MD, MBA1.
1Boston Children's Hospital, Boston, MA, USA, 2Harvard Business School, Boston, MA, USA, 3Massachusetts Institute of Technology, Cambridge, MA, USA.
Objective: Vesicoureteral reflux (VUR) grading on voiding cystourethrogram (VCUG) is subjective and associated with interrater variability. Accurate grading requires interpretation of the full image series across bladder filling and voiding, yet existing automated approaches generally do not incorporate the entire VCUG sequence. We aimed to develop and evaluate a deep learning model for bilateral VUR grade assignment from complete VCUG image series, incorporating study progression and visual explanation of model predictions.
Methods: We performed a retrospective single-center study of VCUG studies obtained from January 2002 through January 2025. A random sample of patients aged 5 years or younger was included. Patients with posterior urethral valves, neurogenic bladder, bladder exstrophy, or prior reconstructive genitourinary surgery were excluded. Bilateral left and right VUR grades were extracted from radiology reports and used as ground truth. The task was formulated as a 10-class bilateral grade assignment problem, including grades 0, 1, 1.5 (grade 1-2), 2, 2.5(grade2-3), 3, 3.5 (grade 3-4), 4, 4.5 (grade 4-5_, and 5. Images were resized to 224 × 224 pixels, with up to 16 images used per study. The model used a ResNet-50 backbone with multihead attention and separate left and right prediction heads. Performance was assessed on a held-out test set using exact accuracy, accuracy within ±0.5 grade, accuracy within ±1.0 grade, mean absolute error, and R². Signed Grad-CAM saliency maps were generated for interpretability.
Results: The final cohort included 6,076 VCUG studies from 4,943 patients. Studies were split at the patient level into training, validation, and test sets, including 4,253, 907, and 916 studies, respectively. Median age was 0.8 years (IQR, 0.3-2.0). The cohort included 3,934 female patients (64.7%). Each study contained a mean of 14.3 images.On the held-out test set, the model achieved exact side-specific grade accuracy of 53.7% for the left kidney and 59.2% for the right kidney. Accuracy within ±0.5 grade was 77.6% and 80.2%, respectively, while accuracy within ±1.0 grade was 92.7% and 94.1% (Figure 1). Across both sides, exact accuracy was 56.4%, accuracy within ±0.5 grade was 78.9%, and accuracy within ±1.0 grade was 93.4%. Row-normalized confusion matrices demonstrated that most prediction errors clustered near the diagonal, indicating that misclassifications were typically within one VUR grade rather than widely discordant. Saliency maps and frame-level attention weights suggested that predictions were driven by a limited subset of influential frames rather than uniformly across the full VCUG sequence.
Conclusions: An attention-based computer vision model that incorporates the entire VCUG image series with explicit image-order encoding demonstrated promising performance for VUR grade identification. By evaluating the full study sequence rather than isolated images, this approach more closely reflects clinical VCUG interpretation across bladder filling and voiding. Sequence-aware modeling and saliency-based visual explanation may help support more consistent, interpretable assessment of VUR severity and could serve as a foundation for future clinical decision-support tools after external validation.
Back to 2026 Abstracts