Back to 2026 Abstracts
Retrieval-Augmented Conversational AI for Rare Pediatric Urologic Disease Education: A Pilot Study in Bladder Exstrophy
Scott Durham, MD1, Siddharth Ramanathan, MD1, Eloise Kropp, BSN2, Alican Kilicarslan, MS2, Brad Kropp, MD2.
1University of Toledo Medical Center, Toledo, OH, USA, 2OKC Kids Urology, Oklahoma City, OK, USA.
BACKGROUND: Rare congenital urologic conditions such as bladder exstrophy require highly specialized longitudinal care and patient education. Families frequently seek medical information online, where content quality and accuracy are highly variable. Although large language models (LLMs) and conversational artificial intelligence (AI) platforms have emerged as accessible tools for answering health-related questions, concerns remain regarding hallucinated or unsupported responses.
Objective: We developed a source-grounded conversational educational platform integrated within the “We the BE” application, a patient-centered educational ecosystem focused on bladder exstrophy and related conditions. The platform uses a retrieval-augmented generation (RAG) architecture to ground LLM responses in expert-curated educational videos. Videos were transcribed, segmented into timestamp-aligned text chunks, and indexed within a vector database for semantic retrieval. When a query is submitted, the system retrieves only the most relevant transcript segments to provide controlled context for the LLM. If no relevant content is identified, a fail-safe pathway is activated that avoids unsupported medical advice and directs users to clinical care or external educational resources.
METHODS: Expert-curated educational videos focused on bladder exstrophy and related topics were indexed within the “We the BE” conversational AI platform. Source-grounded lay-language questions directly answerable from video content were generated, resulting in 100 grounded questions. Fifty clinically relevant but non-indexed (“ungrounded”) questions were generated to evaluate fail-safe behavior and hallucination mitigation. For each query, the platform generated a patient-friendly summary response and exact video timestamps corresponding to relevant educational content. Two blinded pediatric urology reviewers independently graded summary accuracy (“accurate” vs “inaccurate”) and timestamp appropriateness (“appropriate” vs “inappropriate”). Ungrounded responses were evaluated for fail-safe activation, unsupported generated content, and hallucinated medical information. Retrieval fidelity was defined as the proportion of responses containing both an accurate summary and appropriate timestamp. Inter-rater reliability was assessed using Cohen’s κ statistic.
RESULTS: A total of 150 prompts were evaluated, including 100 grounded and 50 ungrounded questions. Among grounded queries, the platform demonstrated a retrieval fidelity of 90%, with accurate summary generation in 97% and appropriate timestamp retrieval in 92% of responses. Overall reviewer agreement was 94.5% (κ = 0.69). Among ungrounded questions, appropriate fail-safe activation occurred in 94% of cases. In the remaining 6%, the platform identified relevant indexed educational material and successfully returned accurate timestamp-linked responses. No hallucinated or unsupported medical responses were identified.
CONCLUSIONS:This pilot study demonstrates the feasibility of a timestamp-grounded RAG platform for bladder exstrophy education. By linking AI-generated summaries directly to expert-curated video timestamps, the system achieved high retrieval fidelity while maintaining source transparency and minimizing unsupported responses. This approach may represent a scalable model for improving patient and caregiver education in rare pediatric diseases.
Back to 2026 Abstracts