Reliability of LLMs as medical assistants for the general public: a randomized preregistered study.
Bean AM, Payne RE, Parsons G, Kirk HR, Ciro J, Mosquera-Gómez R, Hincapié M S, Ekanayaka AS, Tarassenko L, Rocher L, Mahdi A
- DOI
- 10.1038/s41591-025-04074-y
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/3f379946-161c-4171-801d-a37a22bb1e00 is authoritative.
How this rating was calculated
- CitationsUnresolved reference ×3−0.75★
- IntegrityIntegrity concern−0.5★
- Statistics were not checked: no recomputable values were found in this text — no test statistic reported with its degrees of freedom, no effect estimate printed with both a 95% CI and a p-value, and no percentage printed with both its count and its denominator.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper is a well-designed randomized controlled trial with clear ethical approvals, preregistration, and public data/code. Minor reporting gaps include lack of detailed power analysis, incomplete inclusion/exclusion criteria, and no explicit reporting guideline. A few copyedit issues (typos, inconsistent percentage) and three references not found in registries warrant attention.
Both reviewers classified the study as interventional (RCT), which is adopted. Non-applicable sub-criteria (e.g., animal housing, cell lines) were excluded. The statistics verification component recomputed 0 tests due to coverage limits, so statistical correctness is not independently confirmed. The reviewers diverged slightly on limitations addressed and exact p_values, but the combined evidence supports pass.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
- lowinternal contradictionAbstract reports 94.9% for LLM alone condition identification, but results section reports 94.7% for GPT-4o.
Abstract: 'correctly identifying conditions in 94.9% of cases' vs Results: 'The models were able to suggest at least one relevant condition in 94.7% of cases for GPT-4o'
Abstractreviewer’s wording
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1LLMs alone complete scenarios accurately, identifying conditions in 94.9% of cases.The 94.9% figure is an average across models, but individual model rates vary (94.7%, 99.2%, 90.8%). The claim is supported but slightly imprecise.Evidence: Task validation section reports per-model rates.
“correctly identifying conditions in 94.9% of cases”
AbstractFind in source - supportedReviewers 1, 2Participants using LLMs identified relevant conditions in fewer than 34.5% of cases, no better than control.The results show significantly lower identification rates for LLM groups compared to control, with odds ratios and CIs supporting the claim.Evidence: Results section reports chi-square tests with P < 0.001 and odds ratio 1.76 (95% CI 1.45-2.13) favoring control.
“participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases”
AbstractFind in source - supportedReviewers 1, 2Standard benchmarks do not predict human-LLM interaction failures.The paper shows that MedQA scores are higher than interactive performance and uncorrelated, supporting the claim.Evidence: Question-answering benchmarks section shows higher benchmark scores in 26/30 cases and weak correlation.
“Scores on question-answering are higher than the corresponding scores in user interactions in 26 out of 30 cases”
ResultsFind in source - supportedReviewer 1Simulated patient interactions do not predict human performance.Simulated participants performed better and showed weak or no correlation with human results, supporting the claim.Evidence: Simulated patient interactions section reports regression coefficients near zero for conditions.
“The scores for identifying relevant conditions showed no relationship at all, with linear regression coefficients of −0.01 ± 0.34”
ResultsFind in source - supportedReviewer 2Simulated user interactions do not predict human performance.The paper shows weak or no correlation between simulated and human performance.Evidence: Results section on simulated participants shows low regression coefficients.
“simulations of user interactions with LLMs—a promising method for creating realistic benchmarks—also do not predict human–LLM interaction failures.”
AbstractFind in source
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior studies showing LLMs excel on benchmarks but fail in real clinical settings, and notes that physician-LLM collaboration underperforms. The rationale linking this to testing LLMs with the general public is clear. However, the paper does not explicitly discuss limitations of prior research that it aims to address, though it implicitly does by focusing on human-LLM interaction.
“On the one hand, LLM scores on medical knowledge benchmarks are now commensurate with passing the US Medical Licensing Exam . LLM-generated clinical documents are rated as equivalent to or better than those written by doctors , . On the other hand, excelling at medical tasks in silico does not translate to accurate performance in clinical settings under physician guidance.”
“To understand whether LLMs can reliably support the general public and bring care closer to patients, we conducted a study with 1,298 UK participants.”
Randomization method is described as stratified random sampling via Prolific, and the unit is individual participants. Blinding is partial: participants were blinded to which LLM they used, but the control group was aware they were not using an LLM. A power analysis is mentioned but not detailed in the main text. Inclusion/exclusion criteria are implied (age, English speaking) but not pre-specified in detail. Outlier handling is addressed through exclusions for technical issues and attrition analysis. Controls are appropriate (control group using usual methods). Independent replication is not applicable for a single trial.
“We used stratified random sampling via the Prolific platform to target a representative sample of the UK population in each group.”
“Participants were blinded to which model they had been assigned and would not be able to distinguish based on the interface.”
“For a power analysis and detailed demographics of the participants, as well as a breakdown of the results by sex, see .”
“We used stratified random sampling via the Prolific platform to target a representative sample of the UK population in each group.”
“Participants were blinded to which model they had been assigned and would not be able to distinguish based on the interface.”
“For a power analysis and detailed demographics of the participants, as well as a breakdown of the results by sex, see .”
Sex is reported in the demographics breakdown (referenced to supplementary). Age is mentioned as over 18 but not further detailed. Health status is not reported. Demographics are reported and stratified to match UK population. Species/strain and housing are not applicable.
“For a power analysis and detailed demographics of the participants, as well as a breakdown of the results by sex, see .”
“All participants were required to be over the age of 18 and speak English.”
“We used stratified random sampling via the Prolific platform to target a representative sample of the UK population in each group.”
“We used stratified random sampling via the Prolific platform to target a representative sample of the UK population in each group.”
“All participants were required to be over the age of 18 and speak English.”
“For a power analysis and detailed demographics of the participants, as well as a breakdown of the results by sex, see .”
The study was approved by the Departmental Research Ethics Committee at the Oxford Internet Institute under project number OII_C1A_23_096. Informed consent was obtained from all participants. Regulatory compliance is stated as following relevant guidelines and regulations.
“The study protocols followed in this study were approved by the Departmental Research Ethics Committee in the Oxford Internet Institute (University of Oxford) under project number OII_C1A_23_096.”
“Informed consent was obtained from all participants before enrollment in the study.”
“Methods were carried out in accordance with the relevant guidelines and regulations.”
“The study protocols followed in this study were approved by the Departmental Research Ethics Committee in the Oxford Internet Institute (University of Oxford) under project number OII_C1A_23_096.”
“Informed consent was obtained from all participants before enrollment in the study.”
“Methods were carried out in accordance with the relevant guidelines and regulations.”
The three LLMs (GPT-4o, Llama 3, Command R+) are identified with their providers (OpenAI, Hugging Face, Cohere). Software tools (STATSMODELS, SCIPY, SEABORN) are named with versions. Hyperparameters are referenced to supplementary tables. No antibodies, cell lines, or organisms are applicable.
“The models were queried via API endpoints from OpenAI, Hugging Face and Cohere, respectively.”
“All statistics were computed using the STATSMODELS v0.14.3 and SCIPY v1.13.0 packages in Python.”
“The hyperparameters and inference costs are listed in Supplementary Tables and .”
“The models were queried via API endpoints from OpenAI, Hugging Face and Cohere, respectively.”
“All statistics were computed using the STATSMODELS v0.14.3 and SCIPY v1.13.0 packages in Python.”
Tests are named (chi-square, Mann-Whitney U, bootstrap). Effect sizes with CIs are reported (e.g., odds ratios). Software is identified. Data presentation includes means with CIs and per-group n. Exact p-values are sometimes reported as thresholds (e.g., P < 0.001). Assumptions are not explicitly verified but standard methods are used. Mathematical plausibility checks were not possible for most statistics due to lack of raw data.
“Comparisons between proportions were computed using χ 2 tests with 1 d.f., equivalent to a two-sided Z- test.”
“Participants in the control group had 1.76 (95% CI = 1.45–2.13) times higher odds of identifying a relevant condition than the aggregate of the participants using LLMs.”
“χ 2 (1), n 1 = n 2 = 600, P < 0.001 for all three models”
“Participants in the control group had 1.76 (95% CI = 1.45–2.13) times higher odds of identifying a relevant condition than the aggregate of the participants using LLMs.”
“GPT-4o, χ 2 (1) = 0.17, P = 0.683; Llama 3, χ 2 (1) = 0.34, P = 0.560; Command R+, χ 2 (1) = 0.03, P = 0.861”
Data availability statement is concrete, naming GitHub and Hugging Face repositories. Code is shared on GitHub. Accession numbers are not applicable for this type of data.
“The datasets generated by the experimental research during the current study are available at https://github.com/am-bean/HELPMed as well as https://huggingface.co/datasets/ambean/HELPMed/ .”
“All code used to generate the analysis in the manuscript is shared by the authors for reuse and is available at https://github.com/am-bean/HELPMed .”
“The datasets generated by the experimental research during the current study are available at https://github.com/am-bean/HELPMed as well as https://huggingface.co/datasets/ambean/HELPMed/ .”
“All code used to generate the analysis in the manuscript is shared by the authors for reuse and is available at https://github.com/am-bean/HELPMed .”
The study is preregistered on OSF. Methods are detailed enough for replication. Limitations are discussed in the Discussion. Conclusions are proportional to the evidence. Funding and COI are stated. Trial registration is not applicable as this is not a clinical trial of an intervention. Reporting guideline is not mentioned.
“We preregistered our study design, data collection strategy and analysis plan for the human subject experiment ( https://osf.io/dt2p3 ).”
“Our scenarios focused on common conditions where users may be familiar with the symptoms, and results might differ on rare conditions or less typical presentations.”
“A.M. acknowledges support from Prolific and support for the Dynabench platform from the Data-centric Machine Learning Working Group at MLCommons.”
“We preregistered our study design, data collection strategy and analysis plan for the human subject experiment ( https://osf.io/dt2p3 ).”
“Our scenarios focused on common conditions where users may be familiar with the symptoms, and results might differ on rare conditions or less typical presentations.”
“A.M. acknowledges support from Prolific and support for the Dynabench platform from the Data-centric Machine Learning Working Group at MLCommons.”
Registration stated in text, but no registry ID was detected. No reporting guideline cited.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 33 references by DOI: 16 verified — 3 DOI unresolved, 14 no DOI (shown, not verified).
- UNRESOLVED10.1038/s41591-024-02855-1Adapted large language models can outperform medical experts in clinical text summarizationCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1038/s41591-024-03316-5An evaluation framework for clinical use of large language models in patient interaction tasksCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1038/s41591-024-03180-3Influence of believed AI involvement on the perception of digital medical adviceCited DOI does not resolve to any Crossref record.
- NO DOIChatGPT diagnoses cause of child’s chronic pain after 17 doctors failedNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIKFF Health Misinformation Tracking Poll: Artificial Intelligence and Health InformationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWhat disease does this patient have? A large-scale open domain question answering dataset from medical examsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICombining Human Expertise with Artificial Intelligence: Experimental Evidence from RadiologyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe robot doctor will see you nowNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAI opportunities action planNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICatalyzing equitable artificial intelligence (AI) use to improve global healthNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHHS releases strategic plan for the use of artificial intelligence to enhance and protect the health and well-being of AmericansNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAgentClinic: A multimodal agent benchmark to evaluate AI in simulated clinical environmentsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMisattribution of error origination: the impact of preconceived expectations in co-operative online gamesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICapabilities of Gemini models in medicineNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITowards interactive evaluations for interaction harms in human-AI systemsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILearning to complement humansNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDynabench: rethinking benchmarking in NLPNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
4 data/code links checked; 4 live.
- codeGitHubLIVEHTTP 200https://github.com/am-bean/HELPMedResolves to GitHub (code repository).
- dataHugging FaceLIVEHTTP 200https://huggingface.co/datasets/ambean/HELPMed/Resolves to Hugging Face (data repository).
- dataHugging FaceLIVEHTTP 200https://huggingface.co/datasets/ambean/HELPMed/viewer/default/scenariosResolves to Hugging Face (data repository).
- dataOSFLIVEHTTP 200https://osf.io/dt2p3Resolves to OSF (data repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoExtended Data Table 2 caption“converstaions”→ conversationsTypo in caption.
- MINORconsistencyAbstract“94.9%”→ 94.7% (as in Results)Abstract states 94.9% but Results state 94.7% for GPT-4o; check consistency.
- MINORclarityMethods, Participants“For a power analysis and detailed demographics of the participants, as well as a breakdown of the results by sex, see .”→ Provide a reference to the supplementary section or include the details in the main text.Incomplete sentence with missing reference.
- MINORtypoExtended Data Table 2“converstaions”→ conversationsTypo in table title.
- MINORconsistencyAbstract“94.9%”→ 94.7%Abstract states 94.9% but results section reports 94.7% for GPT-4o alone.
The published work is robust overall, with strong reporting of ethics, preregistration, and data/code availability. An informed reader should weigh the minor reporting gaps (power analysis details, inclusion/exclusion criteria, reporting guideline) and the three references not found in registries, which may warrant verification or correction. The internal inconsistency in the abstract percentage (94.9% vs 94.7%) is a minor but correctable issue.
- 1.HIGHreportingVerify or correct the three references not found in registries: 'Adapted large language models can outperform medical experts in clinical text summarization' (10.1038/s41591-024-02855-1), 'An evaluation framework for clinical use of large language models in patient interaction tasks' (10.1038/s41591-024-03316-5), and 'Influence of believed AI involvement on the perception of digital medical advice' (10.1038/s41591-024-03180-3).References not found in any registry may be fabricated or have incorrect DOIs, which is an integrity concern.
- 2.HIGHreportingReconcile the abstract percentage 94.9% with the Results section's 94.7% for GPT-4o condition identification.An internal contradiction in a headline number undermines trust and may warrant an erratum.
- 3.MEDIUMreportingAdd a detailed power analysis to the Methods or clearly reference the supplementary section, specifying effect size, alpha, and power.The current sentence is incomplete and lacks the missing reference, reducing reproducibility.
- 4.MEDIUMreportingExplicitly state inclusion and exclusion criteria in the Methods, such as age range and health status.Incomplete criteria limit reproducibility and generalizability assessment.
- 5.MEDIUMreportingMention adherence to a reporting guideline such as CONSORT or STROBE in the Methods or Reporting Summary.Reporting guidelines improve transparency and are expected for RCTs.
- 6.MEDIUMstatisticsAdd a statement about verifying statistical assumptions (e.g., normality, equal variance) or justify why standard tests are appropriate.Assumption verification is currently not reported, which is a common reviewer concern.
- 7.MEDIUMreportingClarify the blinding of outcome assessors, as only participant blinding is described.Outcome assessor blinding is a key methodological detail for RCTs.
- 8.LOWcopyeditFix the typo 'converstaions' to 'conversations' in Extended Data Table 2 caption and title.Typos in tables are unprofessional and should be corrected.
- 9.LOWreportingProvide more detailed demographic information (e.g., age distribution, health status) in the main text or clearly reference supplementary tables.Detailed demographics improve transparency and reader confidence.
- 10.LOWreportingInclude a statement on data availability in the main text, not just in the supplementary, to ensure visibility.Data availability statements are more discoverable in the main text.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.