Societies for Pediatric Urology

SPU Home SPU Home Past & Future Meetings Past & Future Meetings

Back to 2026 Abstracts


Large Language Model-Assisted Rule-Based UTD Reclassification of a Prenatal Hydronephrosis Database: A Validation Study
Mandy Rickard, MN, NP, Samer Maher, MSc, Usman Kahloon, MSc, Adree Khondker, MD, Michael Chua, MD, Innocent Nzeyimana, MD, Rodrigo Romao, MD, Joao Pippi Salle, MD, Joana Dos Santos, MD, Chia Wei Teoh, MD, Nithiakishna Selvathesan, MD, Armando J. Lorenzo, MD.
SickKids, Toronto, ON, Canada.


BACKGROUND: The Urinary Tract Dilation (UTD) classification system has replaced the SFU grading system as the preferred framework for staging postnatal hydronephrosis, incorporating anteroposterior pelvic diameter, calyceal dilation, parenchymal status, and ureteral dilation. Large prospective databases collected under SFU grading lack UTD classifications, limiting modern comparative analyses. Manual reclassification at scale is resource-intensive. We evaluated whether a large language model (LLM) could reliably apply UTD rules to an existing SFU-coded database, using expert-assigned UTD grades as the validation standard.
METHODS:
A prospectively collected prenatal hydronephrosis database (1,512 patients, REDCap) containing SFU grade, anteroposterior diameter, and ureteral dilation at up to 11 scheduled ultrasound timepoints, plus one most-recent visit (12 total), was used. An LLM (Claude, Anthropic) was provided with the UTD classification decision framework and tasked with generating a rule-based classification script to be applied across all visits. No patient-identifying data were entered into the LLM; only the UTD decision framework and aggregate variable descriptions were provided. The generated script was reviewed by a human investigator before application. UTD grades (Normal, P1, P2, P3) were derived from: SFU grade (4 → P3; 3 → P2; 2 → P1), APD (≥15 mm → P2; 10-14 mm with SFU 1 → P1), and ureteral dilation (present → minimum P2). Validation was performed against expert-assigned UTD grades already present in the database at the 11 scheduled visits, using exact agreement, within-one-grade agreement, and Cohen’s κ. Three patients lacked complete data at US1, yielding 1,509 pairs for baseline validation.
RESULTS:
The LLM generated 7,690 UTD classifications for 1,512 patients in under 10 minutes. At US1 and US2, where expert UTD grades were prospectively assigned, agreement was strong: exact 80%, within-one-grade 99%, and κ = 0.74 (Table 1, Figure 1). At subsequent visits (US3-US11), where existing grades were assigned less systematically, exact agreement was 57% (κ = 0.40), though within-one-grade agreement remained 92% (Figure 1). At US3-US11, the LLM assigned a higher grade in 29% of discordant cases, consistent with a more systematic application of APD and ureteral dilation criteria. The US1 confusion matrix shows discordance concentrated in adjacent categories, with no cases misclassified by more than one grade in 99% of comparisons (Figure 2).
CONCLUSIONS:

LLM-assisted rule-based classification reliably replicated expert UTD grading at scale, achieving κ = 0.74 relative to prospectively assigned grades. Systematic discordance at later visits reflects inconsistency in retrospective human grading rather than LLM error, as evidenced by consistently high within-one-grade agreement and directional bias. This approach enables rapid, reproducible, and auditable modernization of historical hydronephrosis datasets without manual chart review



Back to 2026 Abstracts