Reliability of LLMs as medical assistants for the general public: a randomized preregistered study.
Bean AM, Payne RE, Parsons G, Kirk HR, Ciro J, Mosquera-Gómez R, Hincapié M S, Ekanayaka AS, Tarassenko L, Rocher L, Mahdi A
- DOI
- 10.1038/s41591-025-04074-y
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/7f0f2475-8733-46d1-a45f-9c624356520b is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ReportingBiological variables not met−0.5★
- ReportingEthical approvals partially met−0.25★
- ReportingStatistical analysis partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 2 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Biological variables not reported
The paper does not report the sex, age distribution, or other demographic characteristics of participants in the main text, making it impossible to assess the biological variables.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
A preregistered, well-designed randomized controlled trial with open data/code and a clear scientific rationale, testing whether LLM assistance improves medical decision-making by the general public. The main weaknesses are reporting-level: headline p-values printed as thresholds, a generic ethics compliance statement, broken supplementary references, and incomplete demographic detail in the main text. No retracted citations, no invalid statistics, and all data/code links are live.
All eight dimensions were evaluated by three independent reviewer runs; reviewers diverged on biological variables (pass vs not applicable vs fail — resolved to pass with reduced confidence), ethical approvals (warn vs pass), key resources (pass vs warn), and statistical analysis (warn vs pass), with the synthesized verdicts explained in each dimension. The statistics verification recomputed only 6 tests (those with test statistic + df or effect + CI); threshold-only and bootstrap/permutation p-values were not machine-verifiable, and GRIM/GRIMMER checks were not applicable to these continuous/derived measures. N/A items (cell lines, mycoplasma, animals, reagents, sequencing accessions) were excluded.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 6 tests: 6 consistent, 0 inconsistent; 6 via agent-written checks.
- CONSISTENTreported p = .683 · recomputed p = .680Reviewer 1GPT-4o disposition accuracy χ² test: reported χ²=0.17, df=1, P=0.683
“GPT-4o, χ 2 (1) = 0.17, P = 0.683”
Taken as given: 0.17 is the χ² statistic with 1 degree of freedom; the test is two-sided (equivalent to a two-sided Z-test, as the paper states for all proportion comparisons); df=1 because the comparison is a 2×2 proportion testMethod: Two-sided χ² survival/rejection probability with df=1, recomputed from the quoted statistic and df.How we recomputed it: pChi2(0.17, 1) - CONSISTENTreported p = .560 · recomputed p = .560Reviewer 1Llama 3 disposition accuracy χ² test: reported χ²=0.34, df=1, P=0.560
“Llama 3, χ 2 (1) = 0.34, P = 0.560”
Taken as given: 0.34 is the χ² statistic with 1 degree of freedom; the test is two-sided; df=1 because the comparison is a 2×2 proportion testMethod: Two-sided χ² with df=1, recomputed from the quoted statistic and df.How we recomputed it: pChi2(0.34, 1) - CONSISTENTreported p = .861 · recomputed p = .862Reviewer 1Command R+ disposition accuracy χ² test: reported χ²=0.03, df=1, P=0.861
“Command R+, χ 2 (1) = 0.03, P = 0.861”
Taken as given: 0.03 is the χ² statistic with 1 degree of freedom; the test is two-sided; df=1 because the comparison is a 2×2 proportion testMethod: Two-sided χ² with df=1, recomputed from the quoted statistic and df.How we recomputed it: pChi2(0.03, 1) - CONSISTENTreported p = .118 · recomputed p = .118Reviewer 1Llama 3 human vs LLM-alone disposition χ² test: reported χ²=2.44, df=1, P=0.118
“Llama 3 as well but not statistically significant ( χ 2 (1) = 2.44, P = 0.118, n 1 = n 2 = 600)”
Taken as given: 2.44 is the χ² statistic with 1 degree of freedom; the test is two-sided; df=1 because the comparison is a 2×2 proportion testMethod: Two-sided χ² with df=1, recomputed from the quoted statistic and df.How we recomputed it: pChi2(2.44, 1) - CONSISTENTreported p = .814 · recomputed p = .814Reviewer 1Attrition-by-treatment-group χ² test: reported χ²=0.948, df=3, P=0.814
“with 26 dropping out from the GPT-4o treatment, 30 from the Llama 3 treatment, 25 from the Command R+ treatment and 20 from the control group ( χ 2 (3) = 0.948, d.f. = 3, P = 0.814)”
Taken as given: 0.948 is the χ² statistic with 3 degrees of freedom; the test is two-sided; df=3 because there are 4 groups (3 treatment + 1 control)Method: Two-sided χ² with df=3, recomputed from the quoted statistic and df.How we recomputed it: pChi2(0.948, 3) - CONSISTENTreported p = .814 · recomputed p = .814Reviewer 2Chi-square test for attrition across groups
“χ 2 (3) = 0.948, d.f. = 3, P = 0.814”
Taken as given: The chi-square statistic is 0.948 with 3 degrees of freedom.; The test is two-tailed.Method: Recomputed using the chi-square CDF: p = 1 - chi2Cdf(0.948, 3).How we recomputed it: pChi2(0.948, 3)
- lowinternal contradictionThe abstract states 'fewer than 34.5% of cases' for condition identification, but the results section reports a range of 0.34-0.43 for Command R+ and 0.42-0.54 for GPT-4o, which are not directly comparable to the abstract's single percentage.
“participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%”
AbstractFind in source
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
11 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 3Participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group.The presented χ² results directly back the claim that LLM-assisted participants were not better than control on both outcomes.Evidence: χ² tests showing treatment groups significantly worse than control for conditions (P < 0.001 for all three models) and nonsignificant disposition differences (P = 0.683, 0.560, 0.861).
“However, participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group.”
AbstractFind in source - supportedReviewer 1We identify user interactions as a challenge to the deployment of LLMs for medical advice.The qualitative transcript evidence and quantitative interaction-measurement results directly support the identification of user interaction as a failure point.Evidence: Transcript analysis of 30 sampled interactions showing incomplete information provision, LLM misinterpretations, and inconsistent responses.
“We identify user interactions as a challenge to the deployment of LLMs for medical advice.”
AbstractFind in source - supportedReviewers 1, 3Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants.The MedQA and simulation comparisons directly support the claim that these proxies fail to predict human interactive performance.Evidence: MedQA benchmark scores higher than interactive performance in 26/30 cases; simulated-user regression coefficients near zero or negative for conditions.
“Standard benchmarks for medical knowledge and simulated patient interactions do not predict the failures we find with human participants.”
AbstractFind in source - supportedReviewer 1LLMs tested alone complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average.The stated averages are consistent with the reported per-model task-validation figures (mean 94.9% and 56.3%).Evidence: Task-validation results: condition accuracy 94.7% (GPT-4o), 99.2% (Llama 3), 90.8% (Command R+); disposition accuracy 64.7%, 48.8%, 55.5%.
“Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average.”
AbstractFind in source - supportedReviewers 1, 2None of the tested language models were ready for deployment in direct patient care.The claim is appropriately scoped to the tested models and is backed by the study's primary outcome results.Evidence: The study's overall finding that LLM-assisted participants performed no better than control on both primary outcomes.
“In our work, we found that none of the tested language models were ready for deployment in direct patient care.”
Discussion ¶6Find in source - supportedReviewer 2LLMs alone perform well on the scenarios, but participants using LLMs perform no better than control.The paper provides direct evidence: LLM-alone accuracy is high (94.9% conditions, 56.3% disposition), while participant performance is lower and not significantly different from control for disposition, and worse for condition identification.Evidence: Results sections 'Task validation' and 'Experimental performance' with figures and statistics.
“Tested alone, LLMs complete the scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average. However, participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and disposition in fewer than 44.2%, both no better than the control group.”
AbstractFind in source - supportedReviewer 2Standard benchmarks (MedQA) do not predict human-LLM interaction failures.The paper shows that LLM performance on MedQA is high but uncorrelated with human-LLM performance, with specific examples of high benchmark scores corresponding to low human scores.Evidence: Results section 'Question-answering benchmarks' and Figure 4.
“However, benchmark scores of more than 80% still corresponded to human experimental scores below 20% in several cases, indicating the potential size of the differences.”
ResultsFind in source - supportedReviewer 2Simulated patient interactions do not predict human-LLM interaction failures.The paper shows that simulated participants perform better and with less variability, and regression coefficients are weak or near zero, indicating poor prediction.Evidence: Results section 'Simulated patient interactions' and Figure 4.
“The scores for identifying relevant conditions showed no relationship at all, with linear regression coefficients of −0.01 ± 0.34 for GPT-4o users, −0.17 ± 0.29 for Llama 3 users and −0.01 ± 0.51 for Command R+ users.”
ResultsFind in source - supportedReviewer 2User interactions are a key challenge to LLM deployment.The paper provides qualitative and quantitative evidence of communication failures, including incomplete information from users and LLMs not conveying correct suggestions.Evidence: Results section 'Performance in user interactions' and qualitative analysis of transcripts.
Despite these correct suggestions appearing in the conversations, users did not consistently include them in the final responses, indicating a second breakdown in communication between the model and user.
Resultsreviewer’s wording - supportedReviewer 3LLMs alone complete the medical scenarios accurately, correctly identifying conditions in 94.9% of cases and disposition in 56.3% on average.The paper provides direct evidence from prompting the LLMs with the scenarios and sampling 60 responses per model per scenario, reporting the proportions.Evidence: Figure 2a and text: 'The models were able to suggest at least one relevant condition in 94.7% of cases for GPT-4o, 99.2% of cases for Llama 3 and 90.8% of cases for Command R+.' and 'The models’ accuracy in recommending dispositions was 64.7% for GPT-4o, 48.8% for Llama 3 and 55.5% for Command R+'.
“The models were able to suggest at least one relevant condition in 94.7% of cases for GPT-4o, 99.2% of cases for Llama 3 and 90.8% of cases for Command R+.”
ResultsFind in source - supportedReviewer 3User interactions are a challenge to the deployment of LLMs for medical advice.The paper provides evidence from analyzing interaction transcripts, showing that users provided incomplete information, LLMs made incorrect suggestions, and users did not consistently follow correct recommendations.Evidence: Results, Performance in user interactions: analysis of 30 transcripts, examples of incomplete information, inconsistent responses, and recovered interactions.
“Overall, users often failed to provide the models with sufficient information to reach a correct recommendation.”
ResultsFind in source
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
2 integrity concerns flagged (0 high).
- lowotherThe paper reports that 98 participants were replaced due to API issues, and 13 due to software errors, but the total number of participants (1,298) and the number of responses (2,400) may not align with the replacement counts; however, the paper explains the adaptation of the stopping protocol.
“requiring 98 participants to be replaced because the models failed to respond to the users. We paid all of the impacted participants. We also replaced 13 participants who appeared in more than one treatment group due to a software error in the Prolific platform.”
MethodsFind in source
Reporting gaps
3 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Biological variables not reportedAssessed
- Statistical reporting gaps (tests, assumptions, effect sizes)Assessed
- Ethics/consent reporting incompleteAssessed
The introduction cites multiple studies showing LLMs achieve high scores on medical exams but fail to improve real-world clinical performance. It argues that benchmarks do not predict human-LLM interaction failures, providing a clear rationale for the study. Limitations of prior work (e.g., lack of human user testing) are explicitly addressed as a motivation for the study.
“excelling at medical tasks in silico does not translate to accurate performance in clinical settings under physician guidance”
“excelling at medical tasks in silico does not translate to accurate performance in clinical settings under physician guidance”
“To understand whether LLMs can reliably support the general public and bring care closer to patients, we conducted a study with 1,298 UK participants.”
“We considered two common testing approaches for medical capabilities in LLMs and found that although they may assess the medical information stored in the LLMs, they do not reflect the challenges of user interactions in deployment.”
“Although LLMs now achieve strong performances on medical tasks, attempts to support doctors with LLMs in real clinical settings have faced difficulties.”
“To understand whether LLMs can reliably support the general public and bring care closer to patients, we conducted a study with 1,298 UK participants.”
Randomization is described as 'randomly assigned' with stratification based on demographics, but the specific randomization method (e.g., random number generator) is not detailed. Blinding is reported: participants were blinded to which model they received, but the control group was necessarily unblinded. A power analysis is mentioned but not detailed in the main text (referred to supplementary). Inclusion/exclusion criteria are described (age, English speaking, completion). Outlier handling is not explicitly discussed, but attrition is reported and analyzed. Controls are appropriate (control group using usual methods). Independent replication is not applicable for a single trial.
“We used stratified random sampling via the Prolific platform to target a representative sample of the UK population in each group.”
“Participants were blinded to which model they had been assigned and would not be able to distinguish based on the interface. The control group were necessarily aware that they were not using an LLM.”
“For a power analysis and detailed demographics of the participants, as well as a breakdown of the results by sex, see .”
“We used stratified random sampling via the Prolific platform to target a representative sample of the UK population in each group.”
“Participants were blinded to which model they had been assigned and would not be able to distinguish based on the interface.”
“For a power analysis and detailed demographics of the participants, as well as a breakdown of the results by sex, see .”
“We used stratified random sampling via the Prolific platform to target a representative sample of the UK population in each group.”
“Participants were blinded to which model they had been assigned and would not be able to distinguish based on the interface. The control group were necessarily aware that they were not using an LLM.”
The main text does not report the sex of participants, age distribution, or health status. It mentions that a breakdown of results by sex is in the supplementary, but the main text itself lacks this information. The presurvey collected education, English fluency, etc., but results are not presented. This is a critical omission for a human subjects study.
“We used stratified random sampling via the Prolific platform to target a representative sample of the UK population in each group.”
“All participants were required to be over the age of 18 and speak English.”
“For a power analysis and detailed demographics of the participants, as well as a breakdown of the results by sex, see .”
The paper names the ethics committee (Oxford Internet Institute Departmental Research Ethics Committee) and provides a protocol number. Informed consent is explicitly stated. However, regulatory compliance is only mentioned as 'in accordance with the relevant guidelines and regulations' without naming a specific framework like the Declaration of Helsinki, which is inadequate.
“The study protocols followed in this study were approved by the Departmental Research Ethics Committee in the Oxford Internet Institute (University of Oxford) under project number OII_C1A_23_096.”
“Informed consent was obtained from all participants before enrollment in the study.”
“Methods were carried out in accordance with the relevant guidelines and regulations.”
“The study protocols followed in this study were approved by the Departmental Research Ethics Committee in the Oxford Internet Institute (University of Oxford) under project number OII_C1A_23_096.”
“Informed consent was obtained from all participants before enrollment in the study.”
“The study protocols followed in this study were approved by the Departmental Research Ethics Committee in the Oxford Internet Institute (University of Oxford) under project number OII_C1A_23_096.”
“Informed consent was obtained from all participants.”
“Methods were carried out in accordance with the relevant guidelines and regulations.”
The three LLMs (GPT-4o, Llama 3, Command R+) are identified with their providers. Statistical software (STATSMODELS v0.14.3, SCIPY v1.13.0, SEABORN v0.13.2) is named with versions. No antibodies, cell lines, or organisms are used, so those criteria are not applicable.
“The models were queried via API endpoints from OpenAI, Hugging Face and Cohere, respectively. The hyperparameters and inference costs are listed in Supplementary Tables and .”
“we selected three different leading LLMs: GPT-4o, chosen for its large user base as the model most likely to be used by the general public; Llama 3, selected for its open weights, most likely to be used as the backbone for creating specialized medical models; and Command R+, included for its use of retrieval-augmented generation”
“All statistics were computed using the STATSMODELS v0.14.3 and SCIPY v1.13.0 packages in Python.”
“All statistics were computed using the STATSMODELS v0.14.3 and SCIPY v1.13.0 packages in Python.”
The paper names chi-square tests, Mann-Whitney U tests, bootstrap, and linear regression. Effect sizes (odds ratios, mean differences) are reported with 95% CIs. Software versions are stated. However, some p-values are given as 'P < 0.001' without exact values, and assumptions for parametric tests are not discussed (though non-parametric tests are used). Data presentation includes means with CIs and per-group n, which is adequate for a large trial.
“participants using LLMs were significantly less likely than those in the control group to correctly identify at least one medical condition relevant to their scenario ( χ 2 (1), n 1 = n 2 = 600, P < 0.001 for all three models)”
“Participants in the control group had 1.76 (95% CI = 1.45–2.13) times higher odds of identifying a relevant condition than the aggregate of the participants using LLMs.”
“All statistics were computed using the STATSMODELS v0.14.3 and SCIPY v1.13.0 packages in Python.”
“Comparisons between proportions were computed using χ 2 tests with 1 d.f., equivalent to a two-sided Z- test.”
“GPT-4o, χ 2 (1) = 0.17, P = 0.683; Llama 3, χ 2 (1) = 0.34, P = 0.560; Command R+, χ 2 (1) = 0.03, P = 0.861”
“Participants in the control group had 1.76 (95% CI = 1.45–2.13) times higher odds of identifying a relevant condition than the aggregate of the participants using LLMs.”
“Participants in the control group had 1.76 (95% CI = 1.45–2.13) times higher odds of identifying a relevant condition”
The data availability statement provides concrete repositories (GitHub and Hugging Face) with URLs. Code is shared in a public GitHub repository. Accession numbers are not applicable as no sequencing data are involved.
“The datasets generated by the experimental research during the current study are available at https://github.com/am-bean/HELPMed as well as https://huggingface.co/datasets/ambean/HELPMed/ .”
“All code used to generate the analysis in the manuscript is shared by the authors for reuse and is available at https://github.com/am-bean/HELPMed .”
“The datasets generated by the experimental research during the current study are available at https://github.com/am-bean/HELPMed as well as https://huggingface.co/datasets/ambean/HELPMed/ .”
“All code used to generate the analysis in the manuscript is shared by the authors for reuse and is available at https://github.com/am-bean/HELPMed .”
“The datasets generated by the experimental research during the current study are available at https://github.com/am-bean/HELPMed as well as https://huggingface.co/datasets/ambean/HELPMed/ .”
“All code used to generate the analysis in the manuscript is shared by the authors for reuse and is available at https://github.com/am-bean/HELPMed .”
The methods section is comprehensive, covering scenario development, participant recruitment, treatment allocation, and statistical analysis. The study is preregistered on OSF (https://osf.io/dt2p3). All primary and secondary outcomes are reported, including negative results. Limitations are discussed, including the use of vignettes and the possibility of improved future models. Conclusions are proportional to the evidence. A reporting guideline (e.g., CONSORT) is not referenced, and a conflict-of-interest statement is not in the main text (though referenced as available online).
“We preregistered our study design, data collection strategy and analysis plan for the human subject experiment ( https://osf.io/dt2p3 ).”
“Participants using LLMs did not have statistically significant differences in disposition accuracy from the control group”
“Our scenarios focused on common conditions where users may be familiar with the symptoms, and results might differ on rare conditions or less typical presentations.”
“We preregistered our study design, data collection strategy and analysis plan for the human subject experiment ( https://osf.io/dt2p3 ).”
“Our work can only provide a lower bound on performance: newer models, models that make use of advanced techniques from chain of thought to reasoning tokens, or fine-tuned specialized models, are likely to provide higher performance on medical benchmarks.”
“The funders had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.”
“We preregistered our study design, data collection strategy and analysis plan for the human subject experiment ( https://osf.io/dt2p3 ).”
Registration stated in text, but no registry ID was detected. No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 37 references by DOI: 16 verified — 21 no DOI (shown, not verified).
- NO DOIChatGPT diagnoses cause of child’s chronic pain after 17 doctors failedNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIKFF Health Misinformation Tracking Poll: Artificial Intelligence and Health InformationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICombining Human Expertise with Artificial Intelligence: Experimental Evidence from RadiologyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISuperhuman performance of a large language model on the reasoning tasks of a physicianNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe robot doctor will see you nowNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDigital access – a ‘front door to the NHS’No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAI opportunities action planNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICatalyzing equitable artificial intelligence (AI) use to improve global healthNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHHS releases strategic plan for the use of artificial intelligence to enhance and protect the health and well-being of AmericansNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILarge language models encode clinical knowledgeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICapabilities of GPT-4 on medical challenge problemsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAgentClinic: A multimodal agent benchmark to evaluate AI in simulated clinical environmentsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn evaluation framework for clinical use of large language models in patient interaction tasksNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInfluence of believed AI involvement on the perception of digital medical adviceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMisattribution of error origination: the impact of preconceived expectations in co-operative online gamesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICapabilities of Gemini models in medicineNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITowards conversational diagnostic artificial intelligenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITowards interactive evaluations for interaction harms in human-AI systemsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILearning to complement humansNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDynabench: rethinking benchmarking in NLPNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINICE guidanceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
4 data/code links checked; 4 live.
- codeGitHubLIVEHTTP 200https://github.com/am-bean/HELPMedResolves to GitHub (code repository).
- dataHugging FaceLIVEHTTP 200https://huggingface.co/datasets/ambean/HELPMed/Resolves to Hugging Face (data repository).
- dataHugging FaceLIVEHTTP 200https://huggingface.co/datasets/ambean/HELPMed/viewer/default/scenariosResolves to Hugging Face (data repository).
- dataOSFLIVEHTTP 200https://osf.io/dt2p3Resolves to OSF (data repository).
Copyediting
9 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 9 minor suggestions below.
9 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoExtended Data Table 2 caption“In both of these converstaions, the model is GPT-4o”→ In both of these conversations, the model is GPT-4oMisspelling of 'conversations'.
- MINORclarityDiscussion, paragraph 6“Developers of general-purpose LLM platforms may have an incentive to design risk-adverse LLMs”→ risk-averse LLMs'Risk-adverse' should be 'risk-averse'.
- MINORotherEnd of reference list“rstand participant health behavior. Res. Soc. Adm. Pharm. 17 , 2070 (2021).”→ Complete or remove the truncated reference fragmentThe trailing text appears to be a torn/truncated reference fragment.
- MINORconsistencyExtended Data Table 4“GPT-4o: f = . 755, p < 0.0001”→ f = 0.755Inconsistent formatting of effect-size values (leading zero and spacing) across the table.
- MINORtypoExtended Data Table 2“converstaions”→ conversationsTypo in the table description.
- MINORconsistencyResults, Experimental performance“χ 2 (1), n 1 = n 2 = 600, P < 0.001 for all three models”→ Ensure consistent formatting of chi-square statistics across the paper.Chi-square notation varies slightly (e.g., 'χ 2 (1)' vs 'χ 2 (1) = 0.17').
- MINORclarityMethods, Participants“For a power analysis and detailed demographics of the participants, as well as a breakdown of the results by sex, see .”→ Provide a specific reference to the supplementary section or table.The sentence ends with a period but lacks a reference to the supplementary material.
- MINORtypoExtended Data Table 2 caption“converstaions”→ conversationsTypo in 'conversations'
- MINORconsistencyMethods, Participants“For a power analysis and detailed demographics of the participants, as well as a breakdown of the results by sex, see .”→ Replace 'see .' with a specific reference to the supplementary table or section.Placeholder with period; likely missing reference.
The published work is methodologically solid and largely reproducible, but an informed reader should weigh the reporting gaps: the primary condition-identification comparisons are reported only as P < 0.001 thresholds, several in-text references to supplementary material are broken placeholders (power analysis, demographics, sex breakdown), the ethics compliance statement names no recognized framework, and age/health demographics are absent from the main text. These issues are not validity threats on their own — no retractions, no incorrect recomputed statistics, all links live — but a correction/erratum restoring the broken references and exact p-values, plus a statement of statistical assumption handling, would strengthen the record and support independent re-analysis.
- 1.HIGHreportingFix the broken 'see .' placeholders in Methods, Participants (power analysis, detailed demographics, sex breakdown) by pointing to the specific supplementary/Extended Data tables.Readers currently cannot locate the referenced power analysis or demographic/sex-stratified data, which undercuts both the power_analysis and biological_variables reporting.
- 2.HIGHstatisticsReplace the threshold p-values (P < 0.001) on the headline condition-identification comparisons in Results, Experimental performance with exact p-values.Threshold-only reporting of a computed p-value is imprecise reporting that a critical reader will flag; exact values are needed for the primary outcome.
- 3.HIGHreportingComplete or remove the truncated reference fragment at the end of the reference list ('rstand participant health behavior. Res. Soc. Adm. Pharm. 17 , 2070 (2021).').A torn/truncated reference is a citation-integrity red flag and an unresolved fabrication signal that must be fixed or removed.
- 4.MEDIUMethicsName a recognized regulatory/compliance framework (e.g., Declaration of Helsinki) in Methods, Participants, replacing the generic 'relevant guidelines and regulations' statement.A named framework is the standard for human-subjects research and was flagged by two reviewers as inadequate.
- 5.MEDIUMstatisticsAdd an explicit statement in Methods, Statistical methods on how statistical assumptions were handled (justify the use of non-parametric Mann-Whitney tests in lieu of normality/equal-variance testing).Assumption handling is not described, and two reviewers rated it reported_but_inadequate.
- 6.MEDIUMreportingReport the a priori power analysis parameters (effect size, alpha, power) in the main text Methods rather than only referencing supplementary.The power analysis is a key design element that reviewers expect in the main text for a confirmatory RCT.
- 7.MEDIUMreportingReport the participants' age distribution (and any health/comorbidity summary) in the main text or in a clearly referenced supplementary table.Only an 18+ threshold is given and the reference is broken, leaving the demographic profile incomplete for a human-subjects study.
- 8.MEDIUMreportingReconcile the abstract statement 'fewer than 34.5% of cases' with the Results ranges (0.34–0.43 for Command R+, 0.42–0.54 for GPT-4o).The abstract's single percentage is not directly comparable to the reported per-model ranges and may mislead readers.
- 9.MEDIUMotherClarify in Methods, Participants how the 98 participants replaced due to API issues and 13 due to software errors relate to the final n=1,298 and the 2,400 responses.The total participant count and response count should reconcile with the reported replacement counts; the current explanation is incomplete.
- 10.MEDIUMreportingReference the CONSORT checklist (or state the contents of the Nature Reporting Summary) for this randomized trial in the main text.A named reporting guideline improves transparency for a human RCT and was flagged by all three reviewers as missing.
- 11.LOWcopyeditFix the typo 'converstaions' → 'conversations' in the Extended Data Table 2 caption, and 'risk-adverse' → 'risk-averse' in Discussion, paragraph 6.Minor typos that should be corrected for a clean published record.
- 12.LOWcopyeditStandardize chi-square notation formatting across Results (e.g., 'χ 2 (1)' vs 'χ 2 (1) = 0.17') and fix the inconsistent effect-size formatting in Extended Data Table 4 (e.g., 'f = . 755' → 'f = 0.755').Inconsistent notation across tables and text reduces readability and precision.
- 13.LOWdata codeAdd a DOI or persistent identifier for the GitHub repository to guarantee permanent access to data and code.A persistent identifier protects long-term access beyond the current URL.
- 14.LOWreportingSpecify the randomization method (e.g., random number generator / block randomization) and the unit of randomization, and clarify whether outcome assessors (scorers) were blinded to treatment group.These methodological details were flagged by Reviewer 2 and would strengthen the design description.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.