Societies for Pediatric Urology

SPU Home SPU Home Past & Future Meetings Past & Future Meetings

Back to 2026 Display Posters


Charting the Future: Evaluating the Accuracy of AI-Assisted Chart Abstraction in Pediatric Urology
Raphael R. Pasala, BS, Caleb Q. Ashbrook, MD, Corina B. Do, BA, Vijay K. Rings, MD, Douglass B. Clayton, MD, Abby S. Taylor, MD, MPH, Lauren E. Corona, MD.
Vanderbilt University Medical Center, Nashville, TN, USA.


BackgroundManual chart abstraction is the standard for clinical research but is labor-intensive and subject to inter-reviewer variability and inaccuracy. Large language model (LLM)-based tools offer a promising mechanism for scalable, rapid data extraction from electronic health records. Undescended testis (UDT) represents an ideal use case given its volume and longitudinal documentation. Our objective was to evaluate the accuracy and time burden of an LLM-based chart abstraction tool (Brim; Brim Analytics) against manual (human) extraction for clinical variables in a large UDT cohort. MethodsDocumentation of patients seen at our institution between January 2000 and October 2025 with a diagnosis of undescended or retractile testicles was uploaded into Brim. Two independent reviewers (human A and human B) collected demographics and clinical variables (family history of cryptorchidism, use of ultrasound before referral, testicle location at time of referral, operative approach and concomitant surgery) for 100 randomly selected patients. Human A to human B consensus gold standard was developed and iterations of Brim-derived variables were compared and refined (three attempts on 100 patients). Inter-reviewer agreement (human-human and Brim-gold standard) was assessed using Cohen’s kappa, and accuracy was defined as the percent agreement between Brim-derived and gold standard variables. Time to review was recorded for both human and Brim abstraction.ResultsHuman A and human B demonstrated substantial inter-reviewer agreement across most variables (Table). Mean time to review each patient was 1.58 minutes (SD 0.735) for human A and 1.87 minutes (SD 0.675) for human B. Training and variable development in Brim performed by a physician without prior experience with LLM-based tools was 5 hours, followed by 3 hours of alternate cohort testing of variables (25 subjects) and 10 hours of variable testing and adjustment on the population of interest. Brim review of all notes for the 100 patients in attempt 1 was 4.93 minutes, 5.17 minutes in attempt 2 and 3.32 minutes in attempt 3. Marked improvement in accuracy across attempts was noted with near perfect inter-reviewer kappa with discrete, time-independent variables, especially in operative note contained parameters (figure 2). Variables requiring assessment of temporal relationships in notes (physical exam findings in a return versus initial visit) improved but remained lower than time-independent variables. Notably, Brim identified one pre-referral ultrasound that both human reviewers missed.Conclusion LLM-based chart abstraction demonstrated strong agreement to a human-derived gold standard for most discrete, note-contained clinical variables. Despite upfront development and testing time, LLM-based chart abstraction completes large cohort analysis rapidly. Acknowledging the human error of chart review and considering the scalable capacity of AI-guided chart abstraction, our findings suggest these tools may accelerate pediatric urology research with continued validation.


Back to 2026 Display Posters