New Benchmark Highlights Evidence Gaps in AI Responses to Clinical Questions

By HospiMedica International staff writers
Posted on 28 Sep 2026

Many clinical decisions still rely on limited evidence, while most artificial intelligence benchmarks overlook the patient context that shapes real-world care. General large language models can also struggle to retrieve evidence matched to an individual patient’s history. With an estimated 86% of medical decisions lacking high-quality evidence, hospitals need tools that can surface relevant, patient-specific data at the point of care. Addressing this gap, a new precision medicine benchmark has launched to evaluate and improve clinical LLM responses using patient-contextualized, real-world evidence.

Atropos Health (Palo Alto, CA, USA) has introduced Precision Evidence Bench, described as a first-of-its-kind precision medicine benchmark for evaluating large language models against clinical questions grounded in individual patient context and history. The benchmark measures how well models generate evidence-based answers tailored to the specific patient being assessed. Performance improved by more than 300% when models were provided Scalar Evidence Content from the company’s Alexandria Evidence Library, which contains hundreds of millions of precision Evidence-Based Findings (pEBFs), or study equivalents.


Image: Clinical LLM Performance Improves by >300% When Provided with High-Quality Real-World Evidence in New Precision Medicine Benchmark (Photo courtesy of Atropos Health)

The benchmark incorporates matched patient-history context into both the clinical questions and the evaluation rubric, highlighting limitations of models that rely primarily on published literature and guidelines. Leading general large language models were tested, including GPT 5.6 Sol and Astra 6 (OpenAI), Claude Opus 5 (Anthropic), and 3.8 Flash (Google). The evaluation covered hundreds of real-world precision medicine questions spanning complex patient populations, treatments, comparators, and outcomes using the Population, Intervention, Control, Outcome (PICO) framework.

Without access to the Alexandria Evidence Library, standard models produced a complete, fully cited answer addressing all PICO criteria in only 15% of queries. Performance improved as access to precision evidence increased, with models tested again after receiving access to 100 million and then 500 million pEBFs. According to the findings, this improvement highlights how limited access to large-scale patient data can constrain commercial models when generating accurate, evidence-based recommendations for questions in which individual patient history is essential.

As part of the release, Atropos Health has made 209 benchmark questions and the accompanying scoring rubric freely available on Hugging Face, allowing researchers, healthcare organizations, and model developers to evaluate new models and track progress. The company also plans additional public updates on benchmark performance and intends to extend precision context to other public benchmarks. The Alexandria Evidence Library currently contains 500 million pEBFs and is expected to reach two billion by the end of 2026, further expanding the evidence available to support individualized clinical decision-making.

“Large language models are already very capable of interpreting and communicating medical knowledge. But many of the most important questions for clinicians and patients have never been answered by traditional research. If we want AI to support more precise healthcare decisions, we need to give them evidence to inform those decisions,” said Saurabh Gombar, MD, PhD, Chief Medical Officer of Atropos Health.

Related Links
Atropos Health


Latest AI News