Back to 2026 Abstracts
Accuracy and Inter-Rater Reliability of Clinician Decision-Making Before and After Exposure to an AI Hydronephrosis Severity Index
Samer Maher, MHSc1, Adree Khondker, MD1, Jin Kyu Kim, MD2, Lauren Erdman, PhD3, Hadel Alsubaie, MD4, Laura Betcherman, MD4, Fabio Botelho, MD4, Valentina Bruno, MD4, Martina Bruneira, MD5, Joana Dos Santos, MD4, Usman Kahloon, MPH4, Jethro Kwong, MD1, Mirriam Mikhail, MD4, Asmaa A. Milyani, MD4, Beverly Miranda, MN4, David-Dan Nguyen, MD1, Innocent Nzeyimana, MD4, Mawuenyo Oyortey, MD4, Priya Saini, MD4, Nithiakishna Selvathesan, MD4, Chia W. Teoh, MD4, Marilyn Wong, MD4, Michael Chua, MD4, Armando Lorenzo, MD4, Mandy Rickard, MN4.
1University of Toronto, Toronto, ON, Canada, 2Riley Hospital for Children, Indianapolis, IN, USA, 3Cincinnati Children's Hospital, Cincinnati, OH, USA, 4The Hospital for Sick Children, Toronto, ON, Canada, 5University of Padova, Padova, Italy.
Background The Hydronephrosis Severity Index (HSI) is a convolutional neural network that assigns a continuous severity score to paediatric hydronephrosis from ultrasound images and has undergone multi-centre external validation. Validated discrimination alone does not establish clinical influence; a silent trial is the standard pre-deployment step to characterize how clinicians use a model and whether it changes management.
Methods Twenty-three clinicians from paediatric urology and nephrology (residents, fellows, advanced practice providers [APPs], staff physicians) independently reviewed 293 retrospective, de-identified hydronephrosis cases (age 0-24 months), yielding 6,739 clinician-case observations. For each case, participants received standardized focused clinical information (e.g., postnatal age, laterality, prior imaging) and two ultrasound images, then selected a plan: discharge, repeat ultrasound, diuretic renogram, or surgery. They were then shown the HSI output (numerical risk score, predicted probability of surgery, traffic-light category) and recorded a revised or unchanged plan. The primary outcome, accuracy, was defined as concordance between a surgical (vs. non-surgical) recommendation and whether surgery was performed during follow-up (reference standard: real-world surgical decision). Pre- and post-model comparisons used McNemar’s test and a generalized linear mixed-effects logistic regression with random effects for clinician and case. Inter-rater reliability (ICC[2,1]) was compared via bootstrap.
ResultsICC(2,1) rose from 0.63 (95% CI 0.59-0.67) to 0.71 (0.68-0.75); Δ=0.08, p=0.0035. Decisions changed in 19.6% of clinician-case pairings (95% CI 18.7-20.6%); of these, the majority moved toward the HSI recommendation. Overall accuracy improved from 73.8% to 76.8% (Δ=+3.0 percentage points, McNemar p<0.001). In the mixed-effects model, model exposure was associated with increased odds of a correct decision (OR 1.44, 95% CI 1.28-1.62; p<0.001). Improvements were significant across all designations, but the magnitude differed by experience (interaction p=0.04): novice clinicians gained roughly twice the accuracy improvement of more experienced staff (Table 1).
ConclusionThe HSI model influenced clinician decision-making in one in five cases and produced significant improvements in accuracy and inter-rater reliability, with the largest gains among less-experienced clinicians. Because the reference standard is the real-world surgical decision rather than adjudicated need, and trainees in this trial were already within paediatric urology and nephrology, prospective workflow integration studies linked to patient outcomes are needed before clinical deployment.
Table 1. Decision Change and Accuracy Before and After Model Exposure by Clinician Designation.| Designation | Decision change rate (%) | Pre-model accuracy (%) | Post-model accuracy (%) | Δ accuracy (pp) | p-value |
| Resident/novice APP (n=7) | 23.2 | 73.2 | 77.8 | 4.6 | <0.001 |
| Fellows (n=6) | 19.9 | 73.4 | 76.0 | 2.6 | <0.001 |
| Staff physician/expert APP (n=10) | 16.4 | 74.5 | 76.7 | 2.2 | <0.001 |
| Overall (n=23) | 19.6 | 73.8 | 76.8 | 3.0 | <0.001 |
Back to 2026 Abstracts