An LLM chatbot to facilitate primary-to-specialist care transitions: a randomized controlled trial.
Tao X, Zhou S, Ding K, Li S, Li Y, Wu B, Huang Q, Chen W, Shen M, Meng E, Chen X, Hu H, Zhang J, Zhou J, Zou L, Ma L, Han S
- DOI
- 10.1038/s41591-025-04176-7
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/1b09dcb8-689c-4299-9b16-b78e6fed2548 is authoritative.
How this rating was calculated
Started at 5★ — no deductions. Nothing the checks ran surfaced a material problem.
- References were not verified against Crossref/OpenAlex.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 6 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper reports a well-designed, multicenter pragmatic RCT with rigorous methodology and comprehensive reporting across all eight rigor dimensions. Minor gaps include the lack of exact p-values (thresholds used), absence of a formal outlier-handling plan, and incomplete health-status reporting, but none undermine the core findings.
Evaluated as an interventional trial; both independent reviewer runs agreed on all dimensions. The copyedit pass flagged one missing figure number. Verification components found no retracted citations, no statistical inconsistencies, all links live, and no unsupported claims.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks.
- CONSISTENTreported p = .134 · recomputed p = .134Reviewer 2Chi-squared test for sex distribution across three trial groups
“Sex, n (%) Female 1,141 (55.1) Male 928 (44.9) ... P value 0.134”
Taken as given: The observed counts are 382, 309, 361, 328, 398, 291 for the six cells (PreA-only F, PreA-only M, PreA-human F, PreA-human M, No-PreA F, No-PreA M); The expected counts are computed from row and column totals; The chi-squared statistic is approximately 4.02 with 2 degrees of freedom; The test is two-tailedMethod: Pearson's chi-squared test for a 3×2 contingency tableHow we recomputed it: 1-chi2Cdf(4.02,2)
Overstated conclusions
None found · partly checkedConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Nothing surfaced — but not everything feeding this category ran (missing: surrogate-endpoint assessment), so read this as a partial clean bill.
10 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewer 1PreA-only group significantly reduced consultation duration compared to No-PreA.The evidence directly supports this claim with a clear comparison and p-value.Evidence: Results: PreA-only 3.14 ± 2.25 min vs No-PreA 4.41 ± 2.77 min, P < 0.001, relative reduction 28.7% (95% CI 22.7–34.8).
“PreA-only group showing significantly reduced physician consultation duration (28.7% reduction; 3.14 ± 2.25 min) compared to the No-PreA group (4.41 ± 2.77 min; P < 0.001)”
AbstractFind in source - supportedReviewer 1Significant improvements in physician-perceived care coordination.The evidence directly supports this claim with a large effect size and p-value.Evidence: Results: Care coordination PreA-only 3.69 ± 0.90 vs No-PreA 1.73 ± 0.95, P < 0.001, relative increase 113.1% (95% CI 107.4–118.7).
“significant improvements in physician-perceived care coordination (mean scores 113.1% increase; 3.69 ± 0.90 versus 1.73 ± 0.95; P < 0.001)”
AbstractFind in source - supportedReviewer 1Significant improvements in patient-reported communication ease.The evidence directly supports this claim.Evidence: Results: Ease of communication PreA-only 3.99 ± 0.62 vs No-PreA 3.44 ± 0.97, P < 0.001, relative increase 16.0% (95% CI 13.5–18.5).
“patient-reported communication ease (mean scores 16.0% increase; 3.99 ± 0.62 versus 3.44 ± 0.97; P < 0.001)”
AbstractFind in source - supportedReviewer 1Equivalent outcomes between PreA-only and PreA-human groups confirm autonomous operation capability.The evidence shows no significant differences between the two groups on key outcomes, supporting the claim of autonomous operation.Evidence: Results: Consultation duration PreA-only vs PreA-human P = 0.17; care coordination P = 0.45; no significant differences in patient-reported outcomes.
“Equivalent outcomes between the PreA-only and PreA-human groups confirmed the autonomous operation capability.”
AbstractFind in source - supportedReviewers 1, 2Co-designed PreA outperformed the same model with additional fine-tuning on local dialogues.The comparative simulation study directly supports this claim with quality scores and p-values.Evidence: Results: Co-designed model scored higher than data-tuned model across history-taking, diagnosis, and test ordering (all P < 0.001).
“Co-designed PreA outperformed the same model with additional fine-tuning on local dialogues across clinical decision-making domains.”
AbstractFind in source - supportedReviewer 1Co-design with local stakeholders is a more effective strategy than passive local data collecting for deploying LLMs.The RCT and comparative simulation provide evidence that co-design leads to better outcomes and avoids replicating systemic biases.Evidence: Results: Co-designed model outperformed data-tuned model; data-tuned model replicated suboptimal practices. The RCT demonstrates real-world effectiveness.
“Co-design with local stakeholders, compared to passive local data collecting, represents a more effective strategy for deploying LLMs to strengthen health systems and enhance patient-centered care in resource-limited settings.”
DiscussionFind in source - supportedReviewer 2PreA significantly reduced physician consultation duration compared to usual care.The evidence (Figure 2, Table) shows a 28.7% relative reduction with P < 0.001, which is adequately supported.Evidence: Table and Figure 2 report mean consultation duration: PreA-only 3.14 ± 2.25 min vs No-PreA 4.41 ± 2.77 min, P < 0.001, relative reduction 28.7% (95% CI 22.7–34.8).
“The PreA-only group had a significantly shorter consultation duration compared to the No-PreA group (PreA-only 3.14 ± 2.25 versus No-PreA 4.41 ± 2.77 min; P < 0.001; Fig. ), corresponding to a 28.7% (95% CI 22.7–34.8) relative reduction.”
ResultsFind in source - supportedReviewer 2PreA significantly improved physician-perceived care coordination and patient-reported communication ease.The paper reports substantial improvements in both outcomes with P < 0.001 and 95% CIs, supporting the claim.Evidence: Care coordination: PreA-only 3.69 ± 0.90 vs No-PreA 1.73 ± 0.95, P < 0.001, relative increase 113.1% (95% CI 107.4–118.7). Ease of communication: PreA-only 3.99 ± 0.62 vs No-PreA 3.44 ± 0.97, P < 0.001, relative increase 16.0% (95% CI 13.5–18.5).
Physicians reported a significantly higher value for PreA referral reports compared to the usual one (Care coordination: PreA-only 3.69 ± 0.90 versus No-PreA 1.73 ± 0.95; P < 0.001; Fig. ) ... Patients and care partners in the PreA-only group reported significantly improved consultation experiences compared to the No-PreA group across the primary outcome, ease of communication: 3.99 ± 0.62 versus 3.44 ± 0.97; P < 0.001
Resultsreviewer’s wording - supportedReviewer 2Equivalent outcomes between PreA-only and PreA-human groups confirmed autonomous operation capability.The paper reports no significant differences between these groups across multiple outcomes, supporting autonomous operation.Evidence: No significant difference in consultation duration (P = 0.17), patient-reported outcomes, or care coordination (P = 0.45).
No significant difference was observed between PreA-only and PreA-human groups (3.17 ± 2.87 min; P = 0.17). ... No significant difference in perceived value was observed between the masked PreA-only and PreA-human groups (P = 0.45).
Resultsreviewer’s wording - supportedReviewer 2PreA-assisted medical consultation did not introduce detectable, systematic alterations in physician decision-making.The classification analysis and domain-specific comparisons show no significant differences in clinical notes, supporting the claim.Evidence: Classification analysis: F1 score 0.57, P = 0.81, ΔF1 < 0.02. Domain-specific analysis: P values between 0.10 and 0.90.
Classification analysis yielded near-random discriminability (F1 score 0.57; P = 0.81; ΔF1 < 0.02). ... no significant difference was observed between the PreA-only and No-PreA groups (or PreA-only and PreA-human groups) in the five domains (P = 0.10–0.90; Extended Data Table ).
Resultsreviewer’s wording
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The Introduction cites numerous studies on healthcare system inefficiencies and LLM applications, acknowledges gaps in prior evidence (e.g., lack of real-world trials, risks of biased local data), and presents a clear rationale for the co-design approach. The central question is explicitly stated. Limitations of prior work are discussed and addressed via the co-design methodology.
“The growing burden of multimorbidity and aging populations has exposed vulnerabilities in healthcare delivery worldwide”
“To bridge the gap between the potential of LLMs and their practical impact in resource-limited settings, we developed PreA (Pre-Assessment), an LLM chatbot (OpenAI; GPT-4.0 mini) for primary-to-specialist care transitions, using a multistakeholder participatory co-design approach”
“Critically, evidence is lacking for LLM chatbots that directly interact with socioeconomically diverse patient populations while supporting both curative and caring aspects of medicine in high-volume clinical environments”
“Yet, the relative utility of co-design versus passive data collecting for meeting clinical needs remains unknown.”
Randomization was computer-generated with allocation concealment. The unit of randomization is individual patients. Blinding was single-blind (patients unblinded, physicians and analysts blinded). A priori power analysis is reported with >80% power. Inclusion and exclusion criteria are clearly listed. Outlier handling is not explicitly defined (no ITT/per-protocol statement), but the analysis population is described. The design is adequate for a pragmatic trial.
“Participants were allocated to one of the three groups using an individual-level, computer-generated randomization sequence without stratification. Allocation was concealed to prevent selection bias.”
“This trial was single-blinded: while the patients knew their group assignments (PreA-only, PreA-human, or No-PreA), the physicians were uninformed about the PreA-intervention groups (PreA-only or PreA-human).”
“The target minimum sample size of 2,010 participants (670 per study arm) was prespecified based on a power analysis using preliminary data from the pilot study of 90 patients. This minimum target sample size ensured sufficient power (>80%) for the primary outcome at a significance level of 0.05.”
“Participants were allocated to one of the three groups using an individual-level, computer-generated randomization sequence without stratification.”
“This trial was single-blinded: while the patients knew their group assignments (PreA-only, PreA-human, or No-PreA), the physicians were uninformed about the PreA-intervention groups (PreA-only or PreA-human). Furthermore, research staff involved in data analysis remained blinded to group assignments throughout the study.”
“The target minimum sample size of 2,010 participants (670 per study arm) was prespecified based on a power analysis using preliminary data from the pilot study of 90 patients.”
Sex is reported (55.1% female, 44.9% male). Age is reported as mean ± SD. Weight and health status are not reported. Demographics are extensive (education, income, work status, ethnicity, setting). The single-sex justification is not applicable as both sexes are enrolled.
“Participants had a mean age of 47.6 ± 14.6 years and included 1,141 women (55.1%) and 928 men (44.9%).”
“Table 1 Distribution of baseline covariates across three trial groups”
The study was approved by named ethics committees (Chinese Academy of Medical Sciences, local ethics committee of the First Affiliated Hospital of Guilin Medical University, IRB of Affiliated Hospital of Gansu Medical College). Informed consent was obtained from all participants. The trial followed the Declaration of Helsinki and ICH-GCP. The trial is registered at ChiCTR.
“The Chinese Academy of Medical Sciences and Peking Union Medical College and the local medical ethics committee of the First Affiliated Hospital of Guilin Medical University approved the study.”
“We obtained informed consent from all participants in this study.”
“The trial followed the Declaration of Helsinki and the International Conference of Harmonization Guidelines for Good Clinical Practice.”
“The Chinese Academy of Medical Sciences and Peking Union Medical College and the local medical ethics committee of the First Affiliated Hospital of Guilin Medical University approved the study. The institutional review boards of the Affiliated Hospital of Gansu Medical College approved the study protocol based on their review and the approval from the medical ethics committee of the First Affiliated Hospital of Guilin Medical University.”
“We obtained informed consent from all participants in this study.”
“The trial followed the Declaration of Helsinki and the International Conference of Harmonization Guidelines for Good Clinical Practice.”
The investigational product (PreA chatbot) is identified as 'OpenAI; GPT-4.0 mini'. Statistical software (Python v.3.7, R v.4.3.0) is stated. The code repository is shared. No antibodies, cell lines, or other bench reagents are applicable.
“PreA (Pre-Assessment), an LLM chatbot (OpenAI; GPT-4.0 mini)”
“Python v.3.7 and R v.4.3.0 were used to perform the statistical analyses and present the results.”
“We developed PreA (Pre-Assessment), an LLM chatbot (OpenAI; GPT-4.0 mini) for primary-to-specialist care transitions”
“Python v.3.7 and R v.4.3.0 were used to perform the statistical analyses and present the results.”
All statistical tests are named, assumptions are checked, effect sizes with confidence intervals are reported, software is identified, data presentation is adequate (individual data points shown, error bars defined), and mathematical plausibility checks pass. However, many p-values are reported as '<0.001' rather than exact values, which is a minor inadequacy.
“corresponding to a 28.7% (95% CI 22.7–34.8) relative reduction.”
“PreA-only 3.14 ± 2.25 versus No-PreA 4.41 ± 2.77 min; P < 0.001”
“The PreA-only group had a significantly shorter consultation duration compared to the No-PreA group (PreA-only 3.14 ± 2.25 versus No-PreA 4.41 ± 2.77 min; P < 0.001; Fig. )”
“corresponding to a 28.7% (95% CI 22.7–34.8) relative reduction”
The data availability statement provides a concrete route: source data in a code repository (with Zenodo DOI) and managed access for individual-level data via email request with conditions. The code repository is publicly available on GitHub with a permanent identifier. No sequencing data are involved.
“Source data are provided in Tables and Extended Data Tables and can be accessed via the code repository ( https://github.com/ShashaHan-collab/PreA-OutpatientRCT )”
“Code for classification analysis and data visualization can be found at https://github.com/ShashaHan-collab/PreA-OutpatientRCT (ref. ).”
“Anonymized, nondialogue individual-level data underlying the results can be requested by qualified researchers for academic use. Requests should include a research proposal, statistical analysis plan and justification for data use, and can be submitted via email to S.H. (hanshasha@pumc.edu.cn).”
Methods are detailed enough for replication. The trial is registered (ChiCTR2400094159). A CONSORT flow diagram is included. All outcomes are reported. Limitations are discussed. Conclusions are proportional to the evidence. Funding and competing interests are disclosed.
“Chinese Clinical Trial Registry identifier: ChiCTR2400094159”
“Fig. 1 CONSORT flow diagram for the randomized controlled trial.”
“Several limitations warrant consideration when interpreting our findings.”
“Chinese Clinical Trial Registry identifier: ChiCTR2400094159”
Registered (1 ID: Chinese Clinical Trial Registry). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
2 data/code links checked; 2 live.
- codeGitHubLIVEHTTP 200https://github.com/ShashaHan-collab/PreA-OutpatientRCTResolves to GitHub (code repository).
- datahttps://www.chictr.org.cn/showprojEN.html?proj=251179LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
1 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 1 minor suggestion below.
1 copyedit issue flagged: mostly other.
- MINORotherIntroduction, paragraph 3“The final PreA chatbot integrated a patient-facing chatbot with low-literacy accessibility features and a clinical interface that generates specialist referrals and supports evidence-based decision-making under time constraints (Extended Data Fig. ).”→ Add a specific figure number after 'Extended Data Fig.' (e.g., Extended Data Fig. 1).Missing figure number.
As a published paper, this work is robust and well-documented. An informed reader should weigh the minor reporting gaps (exact p-values, outlier handling, health status) as areas for potential clarification but not as threats to the study's validity. No erratum or re-analysis is warranted based on the current evidence.
- 1.HIGHreportingReport exact p-values for all primary and secondary outcomes instead of threshold values (e.g., 'P < 0.001') in the Results section to improve precision.Threshold-only p-values are a minor but fixable reporting imprecision that reduces the informativeness of the statistical results.
- 2.HIGHrigorExplicitly define the analysis population (intention-to-treat, modified ITT, or per-protocol) and state how missing data and outliers were handled in the Methods, Statistical analysis section.The absence of a clear outlier-handling and analysis-population statement is a gap that could affect reproducibility and interpretation.
- 3.HIGHreportingInclude health status (e.g., common comorbidities, reason for consultation) of participants in the baseline table (Table 1) to fully satisfy the 'age_weight_health' criterion.Health status is a relevant biological variable for a clinical trial; its omission is a minor reporting gap.
- 4.MEDIUMcopyeditAdd a specific figure number after 'Extended Data Fig.' in the Introduction, paragraph 3 (e.g., 'Extended Data Fig. 1').The missing figure number is a minor labeling error that could confuse readers.
- 5.MEDIUMdata codeInclude a timeline for responding to data access requests in the Data availability statement to set clear expectations.A defined response timeline would improve transparency and usability of the data access mechanism.
- 6.MEDIUMdata codeAdd the permanent DOI for the code repository in the main text of the Code availability section, not just in the reference list.Including the DOI in the main text ensures persistent access and easier citation.
- 7.MEDIUMreportingClarify whether the matched-pairs analysis for physician workload was prespecified and whether physicians were matched on all relevant confounders in the Methods, Statistical analysis section.Transparency about the prespecification of secondary analyses strengthens the study's credibility.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.