Patient Portal Study Finds Claude and GPT Summaries Improved Understanding and Reduced Follow-Up Care Demand

A controlled, three-arm study in a primary-care practice found that patient-facing summaries generated by Claude Sonnet 4.5 and GPT-5.1 Thinking improved patients’ understanding of their care and ability to self-manage compared with standard portal access. The two LLM-supported groups also recorded fewer return visits and fewer requests for antibiotics, while clinician audits found no clinically significant hallucinations.[1]

The significance is not simply that generative AI made clinical notes easier to read. The reported intervention changed patient behavior in a live care setting: people appeared better able to act on medical information without seeking an unnecessary additional appointment or pressing for medication. That is a more consequential test for healthcare AI than satisfaction scores alone, though the study’s single-practice, primary-care setting puts clear limits on what can be inferred about broader costs, safety, or health outcomes.

By the numbers

  • 3 study arms: standard portal access and two LLM-summary groups.
  • 2 LLM systems: Claude Sonnet 4.5 and GPT-5.1 Thinking.
  • 0 clinically significant hallucinations reported: in clinician audits of the generated summaries.
Reported study design and safety audit result3study arms2LLM systems evaluated0clinically significanthallucinations
Data: International Journal of Medical Informatics

What the intervention appears to have changed

Patient portals have become a routine distribution channel for visit notes, test results, instructions, medication lists, and clinician messages. But access is not the same as comprehension. Medical documentation is written primarily to support care delivery, professional communication, coding, and legal recordkeeping. It may contain unfamiliar terminology, unresolved differentials, shorthand, copied-forward information, and instructions scattered across a lengthy record.

The study, published in the International Journal of Medical Informatics, compared conventional portal access with portal summaries generated by two large language model systems. Both AI groups performed better on comprehension and self-management measures and showed lower demand for return visits and antibiotic requests than the standard-access group.[1]

Those outcomes matter because they connect language simplification to operational behavior. A patient who understands what a diagnosis means, what self-care steps to take, what symptoms warrant escalation, and why an antibiotic is not indicated is less likely to create avoidable demand through a repeat appointment or message. In primary care, that can affect appointment availability, inbox workload, continuity, and prescribing quality.

The results are particularly relevant to low-acuity encounters, where uncertainty rather than clinical deterioration may drive a substantial share of follow-up requests. Better explanations can reduce uncertainty without withholding access to care. That distinction is important: a useful patient-facing system should make appropriate follow-up clearer, not merely suppress it.

Why clinician auditing was central to the design

LLM-generated medical text creates a safety problem that conventional readability tools do not: the model can produce plausible but unsupported content. In a clinical setting, a polished sentence is not enough. A summary must preserve the meaning of the underlying record, accurately distinguish instructions from possibilities, avoid inventing diagnoses or treatments, and communicate when a patient should contact a clinician.

The reported clinician audits found no clinically significant hallucinations in the summaries.[1] That is an encouraging finding, but it should be interpreted narrowly. It indicates that the reviewed output did not contain errors judged clinically meaningful in this deployment; it does not establish that either model is intrinsically hallucination-free, nor does it prove safety across specialties, patient populations, languages, or high-acuity cases.

Still, auditing is one of the study’s most consequential elements. A patient portal is a high-trust environment. Text presented alongside a clinician’s records can readily be interpreted as medical advice, even when it is generated automatically. A deployable system therefore needs more than a capable model. It needs source-grounded generation, clear handling of uncertainty, scope controls, escalation language, and a clinical review process that tests whether summaries stay faithful to the chart.

The study’s comparison of two separate models also matters. If both groups improved over standard portal access, the useful mechanism may be the care-workflow design—turning scattered clinical information into an actionable patient explanation—rather than a unique capability of one model. That is potentially good news for health systems, because it suggests the product and governance layer may matter at least as much as a model leaderboard.

A meaningful utilization signal, not a total-cost finding

The reduction in return visits and antibiotic requests is the strongest practical signal in the study, but it should not be overstated. The reported outcome concerns potentially unnecessary follow-up demand in a narrow, low-acuity primary-care setting. It does not by itself measure total healthcare consumption, emergency-department use, hospitalizations, clinical outcomes, delayed diagnoses, or long-term spending.

Nor does fewer follow-up encounters automatically mean better care. Some follow-up is clinically appropriate and can prevent harm. The relevant question is whether summaries reduce avoidable utilization while preserving or improving timely escalation for patients whose condition needs reassessment. The reported comprehension and self-management gains point in the right direction, but larger studies will need to test that trade-off explicitly.

Important design questions remain for any follow-on work: how patients were selected; whether groups were randomized; which conditions and visit types were included; how comprehension and self-management were measured; how frequently summaries required correction; and whether results varied by health literacy, age, language, disability, digital access, or complexity of illness. These details determine whether an apparently promising workflow can be safely generalized.

Implications for primary care, payers and antimicrobial stewardship

For accountable-care organizations and insurers, the study points to a potential route to reducing avoidable utilization without erecting new barriers to care. A patient who receives a comprehensible explanation may be less likely to seek a repeat visit for reassurance. That could be valuable in value-based payment arrangements, where organizations are financially accountable for avoidable demand and care coordination.

Independent practices may see a more immediate operational case. High portal-message volumes and constrained appointment capacity have turned patient communication into a major clinical workload. If chart-grounded summaries reduce confusion before it becomes a message, phone call, medication request, or booking, they could free capacity for patients who need direct evaluation. The economic value would depend on integration costs, clinician review time, liability policies, and whether reductions in demand persist outside a study setting.

The decline in antibiotic requests is also notable. Antibiotic stewardship often depends on communicating why an infection is likely viral, what symptom progression is expected, which supportive measures are appropriate, and which warning signs require review. LLM summaries may be useful when they help patients understand that plan in plain language. But they should not become automated refusal tools: prescribing decisions remain clinical judgments, and summaries must preserve exceptions and safety-net instructions.

What a credible next phase would test

The next evidence step is replication across multiple practices, health systems and patient groups, with pre-specified clinical and operational endpoints. Stronger trials would separately measure unnecessary and necessary follow-up, track urgent care and emergency use, examine medication adherence and symptom outcomes, and monitor whether patients with lower digital literacy benefit equally.

They should also evaluate the full system rather than the base model alone. That includes the exact source material provided to the model, prompt and template design, retrieval and grounding controls, multilingual performance, clinician override mechanisms, audit sampling, incident reporting, and the way summaries are labeled in the portal. A model version change can alter output behavior; healthcare organizations will need continuing validation rather than a one-time approval.

For now, the study offers evidence that patient-facing generative AI can be more than a documentation convenience. In a controlled primary-care deployment, it appears to have improved patient understanding and altered downstream demand in a direction that could benefit both patients and overstretched care teams.[1] The broader claim—that it reliably lowers costs or improves outcomes across healthcare—still requires substantially more evidence.

Editor’s Take

I find the behavioral result more compelling than the model comparison. Healthcare does not need another tool that makes notes sound friendlier; it needs systems that help patients make the right next decision. If a grounded portal summary prevents a needless return visit while making warning signs and escalation paths clearer, that is a concrete product outcome with real clinical and economic value.

I would watch the audit process as closely as the utilization data. “No clinically significant hallucinations” is encouraging, but it is not a blanket safety certification. The durable opportunity is a tightly governed workflow: summaries tied to the chart, explicit safety-netting, continuous clinical review, and measurement of missed necessary care as rigorously as avoided unnecessary care. Hype will outrun the facts if vendors sell this as autonomous medical guidance; the evidence here supports a more practical role as a carefully monitored communication layer.

References

  1. International Journal of Medical Informatics – https://www.sciencedirect.com/science/article/pii/S138650562600287X

Leave a Reply

Your email address will not be published. Required fields are marked *