ai-chatbots-answer-thousands-of-rheumatology-patient-questions-in-landmark-real-world-test
AI Chatbots Answer Thousands of Rheumatology Patient Questions in Landmark Real-World Test

AI Chatbots Answer Thousands of Rheumatology Patient Questions in Landmark Real-World Test

For millions of people living with rheumatic diseases, the questions never stop coming. Between appointments, patients wonder whether their medication is working, whether a new symptom signals a flare, or whether it is safe to plan a pregnancy while on immunosuppressants. Routine consultations, typically squeezed into fifteen minutes, cannot absorb this steady stream of uncertainty. Now, one of the largest real-world evaluations of medical artificial intelligence to date suggests that carefully engineered chatbots may help fill the gap. In a nationwide German study published in the Journal of Medical Systems, researchers deployed ten disease-specific, guideline-grounded large language model chatbots across thirteen rheumatology centres and six patient organisations, and watched what happened when real patients started asking real questions.

The scale of uptake was striking. Between September 2025 and January 2026, the chatbots recorded 6,291 question-and-answer exchanges across 1,534 individual sessions. Rheumatoid arthritis chatbots saw the heaviest use, followed by those for axial spondyloarthritis and ANCA-associated vasculitis. Usage peaked during the day but extended well into the evening, and engagement dropped off at weekends, a pattern consistent with evidence that patients reach for digital health tools precisely when their clinics are closed. The median conversation lasted just three turns, suggesting that most users were hunting for a targeted answer rather than settling in for an extended dialogue, and longer sessions attracted proportionally more feedback ratings.

What patients actually asked about is revealing. Using a rheumatologist-guided, human-in-the-loop automated coding pipeline, the team identified thirteen question categories. Half of all questions, 50.2 percent, were disease-specific, covering topics such as how to prepare for a doctor’s appointment to discuss vasculitis. Medication and monitoring questions accounted for 37.3 percent, with patients asking things like when their newly started drugs should begin to work. Diagnostics came next at 28.2 percent, followed by non-pharmacological interventions and lifestyle at 27.8 percent and prognosis and complications at 20.1 percent. Because coding was multi-label, a single question could touch several themes at once. These are not abstract definitional queries; they are the practical, often anxious questions that arise between consultations and that static leaflets were never designed to answer.

The technical architecture behind the chatbots is central to the story. Rather than letting a general-purpose model answer from its internal knowledge, the researchers used retrieval-augmented generation, or RAG, a technique that anchors every response in a predefined corpus of documents. Here, each chatbot’s knowledge base was restricted to the latest valid German clinical guideline for its disease. Guideline text was hierarchically segmented and embedded in a vector database using a domain-specific embedding model, and candidate passages were retrieved through a hybrid of lexical and dense-vector search, then ranked by a statistical reranker weighting vector and token similarity at seventy to thirty. Passages falling below a twenty percent similarity threshold were discarded, and up to eight were inserted into each prompt alongside source identifiers that were displayed to users. The underlying GPT-4o model, accessed through an Azure endpoint hosted in Germany, was never fine-tuned; behaviour was controlled entirely by retrieval and the system prompt, with generation settings held deliberately conservative at a temperature of 0.10 and top-p of 0.30 to favour reproducible output.

Patients noticed the difference. Of the 2,671 responses that received a rating, 92.9 percent earned a thumbs-up. Among the 190 negative ratings, the dominant complaint, at 65.8 percent, was not error but insufficient detail, a telling signal that users wanted more depth than a single guideline could supply. A total of 602 users completed the optional evaluation questionnaire, and their verdicts were largely favourable: 84.6 percent agreed the chatbot was easy to use, 84.1 percent found the answers easy to understand, and 80.2 percent considered it a useful addition to existing patient education materials. Some 77.4 percent said it would save them time searching for information, and 56 percent preferred it to general internet searches such as Google. Trust, however, was more qualified, with only 67.3 percent describing the answers as trustworthy, a reminder that credibility in AI health tools rests on more than accuracy alone.

On the critical question of safety, the automated evaluation was broadly reassuring. Applying a six-dimensional LLM-as-a-judge framework adapted from prior work, the researchers rated all 6,291 exchanges on correctness, question difficulty, completeness, patient-readiness, safety and patient-centredness. Fully 95.3 percent of answers were rated completely safe, 79.1 percent completely correct, and 70.4 percent highly patient-centred. Only 0.2 percent of answers, eleven in total, were flagged as potentially unsafe, and when two board-certified rheumatologists with more than a decade of experience each independently reviewed those flagged cases, just one was confirmed as genuinely problematic, involving a suggestion that cortisone could be administered subcutaneously during disease flares. For comparison, a recent red-teaming study of publicly available, non-RAG chatbots reported unsafe response rates of five to thirteen percent, although the authors caution that naturalistic and adversarial testing conditions differ too much to declare their system categorically safer.

Yet the study is equally notable for what it exposes. Only 45 percent of responses were judged fully adherent to the underlying guidelines, with 44.2 percent partially adherent and 10.9 percent not adherent at all, and agreement between the automated adherence ratings and physician assessment was weak. In some cases, strict fidelity to the uploaded guidelines produced answers that were outdated, such as a response stating that no Janus kinase inhibitors were approved for ankylosing spondylitis, reflecting the guideline text rather than current regulatory reality. In others, the model drifted beyond its curated corpus despite instructions and conservative settings. This exposes a fundamental design tension: confining a chatbot to a single guideline preserves transparency but leaves legitimate questions unanswered, while loosening the leash undermines the very source-grounding that makes the system trustworthy. The researchers argue the way forward lies in refined retrieval configuration and the controlled integration of additional verified content.

The methodology itself may prove as influential as the results. Evaluating thousands of exchanges by hand is impractical, so the team refined a rheumatologist-guided LLM-as-a-judge approach, using a locally hosted, data-secure open model to perform qualitative coding and structured quality assessment, with physician-coded samples used to calibrate the prompts. Agreement between automated and human coding was strong for question and feedback categories, with Gwet’s prevalence-robust AC1 coefficients reaching 0.91 and 0.96 respectively, and substantial for safety, patient-centredness, correctness and completeness. But the approach faltered on patient-readiness and question difficulty, where the model rated far fewer responses at the highest level than the rheumatologists did, a pattern consistent with known central tendency bias in LLM-based ordinal scoring. The authors are blunt about the implication: automated evaluation should be treated as a scalable screening tool, not a substitute for expert validation.

The study also has honest limits. Use was voluntary and may have attracted digitally confident patients; only 42.5 percent of responses were rated; no user accounts meant repeat sessions could not be excluded; and diagnoses were self-reported. Crucially, no patient outcomes were measured, so whether chatbot use actually improves knowledge, self-efficacy or adherence remains an open question. There was no real-time clinical safety monitoring or escalation pathway, only a system-prompt instruction to seek professional advice for medication changes and Azure’s general content filters. Still, the overall picture is one of cautious optimism. Guideline-grounded chatbots, co-designed with patient research partners from Germany’s largest rheumatology patient organisation, delivered predominantly safe, correct and well-received answers to thousands of real questions asked at all hours. Realising that promise at scale, the authors conclude, will demand robust source governance, faster guideline updates, transparent answer boundaries, continuous safety monitoring and, above all, sustained human oversight.

Subject of Research: Guideline-grounded large language model chatbots for patient education and self-management in rheumatology

Article Title: Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology

Article References: Wilhelmi, T., Bartsch, V., Platt, M., Rashid, A., Hornig, J., Krusche, M., Hueber, A. J., Fink, D., Pfeil, A., Dischereit, G., Holzer, M.-T., Drott, U., Aries, P., Müller, M., Böhm, P., Morf, H., Labinsky, H., Mühlensiepen, F., Benavent, D., … Knitza, J. (2026). Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology. Journal of Medical Systems, 50(1), Article 143. https://doi.org/10.1007/s10916-026-02470-6

Image Credits: AI Generated

DOI: 10.1007/s10916-026-02470-6

Keywords: rheumatology, large language models, chatbots, retrieval-augmented generation, patient education, self-management, artificial intelligence, clinical guidelines, rheumatoid arthritis, patient safety, LLM-as-a-judge, digital health