An LLM chatbot to facilitate primary-to-specialist care transitions: a randomized controlled trial.
Tao X, Zhou S, Ding K, Li S, Li Y, Wu B, Huang Q, Chen W, Shen M, Meng E, Chen X, Hu H, Zhang J, Zhou J, Zou L, Ma L, Han S
- DOI
- 10.1038/s41591-025-04176-7
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/75a6bc70-86c5-48a2-b062-749083cfda62 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- CitationsUnresolved reference−0.25★
- Statistics were not checked: no recomputable values were found in this text — no test statistic reported with its degrees of freedom, no effect estimate printed with both a 95% CI and a p-value, and no percentage printed with both its count and its denominator.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 35 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcomes are consultation duration, physician-rated care coordination, and patient-rated ease of communication. These are process and patient-reported measures, not hard clinical outcomes. The paper does not establish a validated link between these surrogates and long-term clinical outcomes such as mortality or morbidity. Target engagement is not demonstrated in terms of dose-exposure or PK/PD.
“The primary outcomes were consultation duration, physician-rated care coordination, and patient-rated ease of communication. These metrics were selected based on co-design feedback, which identified time efficiency and care coordination as critical for…”
- 02Treatment effect not shown to be clinically meaningful
The reported effects are relative reductions/increases in consultation duration and patient-reported scores. For example, consultation duration reduced by 28.7% (from 4.41 to 3.14 minutes), and care coordination score increased by 113.1% (from 1.73 to 3.69 on a 5-point scale). However, these are not anchored to a minimal clinically important difference or a clear biological/clinical meaningfulness threshold. The absolute changes are small (e.g., ~1.3 minutes reduction in consultation time) and may not be clinically meaningful.
“The PreA-only group had a significantly shorter consultation duration compared to the No-PreA group (PreA-only 3.14 ± 2.25 versus No-PreA 4.41 ± 2.77 min; P < 0.001; Fig. ), corresponding to a 28.7% (95% CI 22.7–34.8) relative reduction.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported pragmatic RCT evaluating an LLM chatbot for care transitions. The study design is rigorous with clear randomization, blinding, and power analysis, and the manuscript thoroughly documents ethics, data/code availability, and limitations. Minor reporting gaps (e.g., missing CONSORT flow diagram, unclear handling of missing physician notes) and a single unresolved reference do not undermine the overall integrity.
Both reviewers independently scored all eight dimensions and agreed on every status; no divergence required reconciliation. The statistics verification component found no recomputable tests (coverage limited to tests with test statistic + df or effect + CI), so reported statistics remain unverified. The citation check flagged one reference not found in any registry (a GitHub/DOI entry), which is a potential fabrication signal but not confirmed.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
- lowinternal contradictionThe abstract states 2,069 patients (1,141 women; 928 men) but the text in Results says '1,141 women (55.1%) and 928 men (44.9%)' which sums to 2,069, consistent. However, the CONSORT flow diagram is not shown in the text, so the exact numbers of excluded patients are not verifiable.
2,069 patients (1,141 women; 928 men) ... included 1,141 women (55.1%) and 928 men (44.9%)
Abstractreviewer’s wording
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
6 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2Co-design with local stakeholders represents a more effective strategy for deploying LLMs to strengthen health systems.The claim is supported by the comparative simulation and the RCT results, but the RCT did not directly compare co-design vs. passive data collection; the simulation did. The claim is somewhat broader than the evidence.Evidence: The comparative simulation study showed co-designed model outperformed data-tuned model. The RCT demonstrated effectiveness of the co-designed PreA.
“Co-design with local stakeholders, compared to passive local data collecting, represents a more effective strategy for deploying LLMs to strengthen health systems and enhance patient-centered care in resource-limited settings.”
DiscussionFind in source - supportedReviewers 1, 2PreA-only group showed significantly reduced physician consultation duration compared to No-PreA.The reported difference in means and p-value support this claim.Evidence: Consultation duration: PreA-only 3.14 ± 2.25 vs No-PreA 4.41 ± 2.77 min; P < 0.001
The PreA-only group had a significantly shorter consultation duration compared to the No-PreA group (PreA-only 3.14 ± 2.25 versus No-PreA 4.41 ± 2.77 min; P < 0.001)
Resultsreviewer’s wording - supportedReviewers 1, 2PreA improved physician-perceived care coordination.The reported increase in mean scores and p-value support this claim.Evidence: Care coordination: PreA-only 3.69 ± 0.90 vs No-PreA 1.73 ± 0.95; P < 0.001
“Care coordination: PreA-only 3.69 ± 0.90 versus No-PreA 1.73 ± 0.95; P < 0.001”
ResultsFind in source - supportedReviewers 1, 2PreA improved patient-reported communication ease.The reported increase in mean scores and p-value support this claim.Evidence: Ease of communication: 3.99 ± 0.62 vs 3.44 ± 0.97; P < 0.001
“ease of communication: 3.99 ± 0.62 versus 3.44 ± 0.97; P < 0.001”
ResultsFind in source - supportedReviewers 1, 2Equivalent outcomes between PreA-only and PreA-human groups confirmed autonomous operation capability.The paper reports no significant differences between these groups on primary outcomes, supporting the claim.Evidence: No significant difference was observed between PreA-only and PreA-human groups (3.17 ± 2.87 min; P = 0.17) for consultation duration; similar for other outcomes.
“No significant difference was observed between PreA-only and PreA-human groups (3.17 ± 2.87 min; P = 0.17)”
ResultsFind in source - supportedReviewers 1, 2Co-designed PreA outperformed the same model with additional fine-tuning on local dialogues across clinical decision-making domains.The comparative simulation study shows significantly higher quality scores for the co-designed model across all domains.Evidence: Co-designed model achieved significantly higher-quality rating scores than the data-tuned counterpart across all domains: history-taking (4.56 ± 0.65 vs 3.86 ± 0.81; P < 0.001), diagnosis (4.67 ± 0.55 vs 2.47 ± 1.44; P < 0.001), testing order (4.23 ± 1.09 vs 2.21 ± 1.12; P < 0.001).
The co-designed model achieved significantly higher-quality rating scores than the data-tuned counterpart ... across all domains
Resultsreviewer’s wording
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcomes are consultation duration, physician-rated care coordination, and patient-rated ease of communication. These are process and patient-reported measures, not hard clinical outcomes. The paper does not establish a validated link between these surrogates and long-term clinical outcomes such as mortality or morbidity. Target engagement is not demonstrated in terms of dose-exposure or PK/PD.
“The primary outcomes were consultation duration, physician-rated care coordination, and patient-rated ease of communication. These metrics were selected based on co-design feedback, which identified time efficiency and care coordination as critical for adoption in high-workload settings, and are established proxies for clinical effectiveness and patient-centered care.”
- INADEQUATEEffect sizeThe reported effects are relative reductions/increases in consultation duration and patient-reported scores. For example, consultation duration reduced by 28.7% (from 4.41 to 3.14 minutes), and care coordination score increased by 113.1% (from 1.73 to 3.69 on a 5-point scale). However, these are not anchored to a minimal clinically important difference or a clear biological/clinical meaningfulness threshold. The absolute changes are small (e.g., ~1.3 minutes reduction in consultation time) and may not be clinically meaningful.
“The PreA-only group had a significantly shorter consultation duration compared to the No-PreA group (PreA-only 3.14 ± 2.25 versus No-PreA 4.41 ± 2.77 min; P < 0.001; Fig. ), corresponding to a 28.7% (95% CI 22.7–34.8) relative reduction.”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Data look implausibly cleanAssessed
3 integrity concerns flagged (0 high).
- lowdata too cleanBaseline characteristics are remarkably well balanced across groups with p-values ranging from 0.134 to 0.956, which is plausible for a large randomized trial but could be scrutinized.
“These baseline covariates were well balanced across the three trial arms, with no significant differences in distribution (Table ).”
Table 1Find in source - lowdata too cleanThe baseline characteristics in Table 1 show very similar means and SDs across groups, which is expected in a large RCT but could be seen as unusually balanced.
Age, years, mean ± s.d. | 47.6 ± 14.6 | 47.2 ± 14.4 | 47.7 ± 14.9 | 47.8 ± 14.6
Table 1reviewer’s wording
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites numerous studies on healthcare system inefficiencies, LLM applications, and co-design approaches, acknowledging both strengths and limitations of prior work. The rationale linking the premise to the study objectives is explicit, and the paper addresses limitations of prior research by proposing co-design as a novel strategy.
“Yet, the relative utility of co-design versus passive data collecting for meeting clinical needs remains unknown.”
Randomization used a computer-generated sequence without stratification, allocation was concealed, and the trial was single-blinded with physicians and data analysts blinded. A priori power analysis determined a target sample size of 2,010 participants. Inclusion/exclusion criteria were pre-specified. Outlier handling is addressed through the analysis population and missing-data approach.
“Participants were allocated to one of the three groups using an individual-level, computer-generated randomization sequence without stratification.”
“The target minimum sample size of 2,010 participants (670 per study arm) was prespecified based on a power analysis using preliminary data from the pilot study of 90 patients. This minimum target sample size ensured sufficient power (>80%) for the primary outcome at a significance level of 0.05.”
“Participants were allocated to one of the three groups using an individual-level, computer-generated randomization sequence without stratification.”
“This trial was single-blinded: while the patients knew their group assignments (PreA-only, PreA-human, or No-PreA), the physicians were uninformed about the PreA-intervention groups (PreA-only or PreA-human).”
“The target minimum sample size of 2,010 participants (670 per study arm) was prespecified based on a power analysis using preliminary data from the pilot study of 90 patients.”
The paper reports age, sex, ethnicity, education, work status, income, and medical discipline for all participants. Since both sexes are enrolled, sex justification is not applicable. Age and health status are reported via inclusion criteria and baseline characteristics.
“Participants had a mean age of 47.6 ± 14.6 years and included 1,141 women (55.1%) and 928 men (44.9%).”
“Participants had a mean age of 47.6 ± 14.6 years and included 1,141 women (55.1%) and 928 men (44.9%).”
“Characteristic | Total | PreA-only | Pre-human | No-PreA | P value”
The paper names the ethics committees that approved the study, states that informed consent was obtained from all participants, and declares adherence to the Declaration of Helsinki and ICH-GCP guidelines.
“The Chinese Academy of Medical Sciences and Peking Union Medical College and the local medical ethics committee of the First Affiliated Hospital of Guilin Medical University approved the study.”
“We obtained informed consent from all participants in this study.”
“The trial followed the Declaration of Helsinki and the International Conference of Harmonization Guidelines for Good Clinical Practice.”
“The Chinese Academy of Medical Sciences and Peking Union Medical College and the local medical ethics committee of the First Affiliated Hospital of Guilin Medical University approved the study.”
“We obtained informed consent from all participants in this study.”
“The trial followed the Declaration of Helsinki and the International Conference of Harmonization Guidelines for Good Clinical Practice.”
The PreA chatbot is described as based on OpenAI's GPT-4.0 mini, and the software used for analysis (Python v.3.7, R v.4.3.0) is identified. Since this is a clinical trial of a software intervention, bench resources like antibodies and cell lines are not applicable.
“we developed PreA (Pre-Assessment), an LLM chatbot (OpenAI; GPT-4.0 mini) for primary-to-specialist care transitions”
“Python v.3.7 and R v.4.3.0 were used to perform the statistical analyses and present the results.”
“we developed PreA (Pre-Assessment), an LLM chatbot (OpenAI; GPT-4.0 mini)”
“Python v.3.7 and R v.4.3.0 were used to perform the statistical analyses and present the results.”
The paper names the statistical tests used (t-tests, Mann-Whitney U, chi-squared, ANOVA, Wilcoxon signed-rank), reports exact p-values, and provides effect sizes with 95% CIs. Software is identified. Data presentation includes box plots and per-group n. Mathematical plausibility checks were not possible for all values due to continuous data, but no obvious errors were found.
“corresponding to a 28.7% (95% CI 22.7–34.8) relative reduction.”
“The PreA-only group had a significantly shorter consultation duration compared to the No-PreA group (PreA-only 3.14 ± 2.25 versus No-PreA 4.41 ± 2.77 min; P < 0.001; Fig. ), corresponding to a 28.7% (95% CI 22.7–34.8) relative reduction.”
“corresponding to a 28.7% (95% CI 22.7–34.8) relative reduction.”
The paper provides a data availability statement with a clear mechanism for requesting anonymized data, and code is available in a public GitHub repository. Raw conversation data are not shared due to privacy, which is appropriate.
“Anonymized, nondialogue individual-level data underlying the results can be requested by qualified researchers for academic use. Requests should include a research proposal, statistical analysis plan and justification for data use, and can be submitted via email to S.H.”
“Code for classification analysis and data visualization can be found at https://github.com/ShashaHan-collab/PreA-OutpatientRCT”
“Source data are provided in Tables and Extended Data Tables and can be accessed via the code repository ( https://github.com/ShashaHan-collab/PreA-OutpatientRCT ) .”
“Code for classification analysis and data visualization can be found at https://github.com/ShashaHan-collab/PreA-OutpatientRCT (ref. ).”
The trial is registered (ChiCTR2400094159), methods are detailed, limitations are discussed, and conclusions are proportional. Funding and competing interests are declared. A reporting summary is mentioned.
“Chinese Clinical Trial Registry identifier: ChiCTR2400094159”
“Several limitations warrant consideration when interpreting our findings.”
Registered (1 ID: Chinese Clinical Trial Registry). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 57 references by DOI: 50 verified — 1 DOI unresolved, 6 no DOI (shown, not verified).
- UNRESOLVED10.5281/zenodo.17330614ShashaHan-colab/PreA-Outpatient RCTCited DOI does not resolve to any Crossref record.
- NO DOIMultimorbidity: A Priority for Global Health ResearchNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPhysician-patient communication in managed careNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffective physician-patient communication and health outcomes: a reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPhysician-patient communication: psychosocial care, emotional well-being, and health outcomesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPre-train, prompt, and predict: a systematic survey of prompting methods in natural language processingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISimCSE: Simple contrastive learning of sentence embeddingsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- codeGitHubLIVEHTTP 200https://github.com/ShashaHan-collab/PreA-OutpatientRCTResolves to GitHub (code repository).
- datahttps://www.chictr.org.cn/showprojEN.html?proj=251179LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly clarity, grammar, consistency.
- MINORconsistencyAbstract“2,069 patients (1,141 women; 928 men)”→ Ensure consistency with Table 1 where sex is reported as 1,141 (55.1%) female and 928 (44.9%) male.The abstract uses a semicolon instead of parentheses for sex counts, which is a minor style inconsistency.
- MINORclarityResults, Patient flow and baseline data“This left 2,138 patients who were randomly assigned in 1:1:1 ratio to use PreA independently (PreA-only, n = 712), use it with staff support (PreA-human, n = 713) or not use it (No-PreA, n = 713). Subsequently, 69 patients opted out or were removed for various reasons.”→ Clarify that the 69 patients were excluded after randomization, leading to the final analysis set of 2,069.The flow is clear but could be more explicit about the timing of exclusions.
- MINORgrammarResults, Outpatient workflow“The PreA-only group had a significantly shorter consultation duration compared to the No-PreA group”→ Change 'compared to' to 'compared with' for formal academic style.Minor grammatical preference.
- MINORclarityResults, Outpatient workflow“corresponding to a 28.7% (95% CI 22.7–34.8) relative reduction.”→ Clarify that the relative reduction is in consultation duration.The sentence is clear but could be more explicit.
- MINORgrammarDiscussion, paragraph 4“The data-tuned model replication of suboptimal practices mirrors broader concerns”→ Change 'replication' to 'replication of' or rephrase to 'The data-tuned model's replication of suboptimal practices'.Missing possessive apostrophe.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (missing CONSORT flow diagram, unclear handling of missing physician notes) and the single unresolved reference. No erratum is warranted for the main findings, but the authors should verify the flagged reference and consider adding the missing flow diagram in any revision.
- 1.HIGHreportingVerify the reference 'ShashaHan-colab/PreA-Outpatient RCT' (DOI 10.5281/zenodo.17330614) that was not found in any registry; correct or remove it if it is erroneous or fabricated.An unresolved reference is a potential fabrication signal that must be resolved before the paper can be fully trusted.
- 2.HIGHreportingAdd a CONSORT-style flow diagram to the main text to clearly show participant flow from screening to analysis, including the 69 patients excluded after randomization.A flow diagram is a CONSORT requirement and would clarify the timing and reasons for exclusions, improving transparency.
- 3.MEDIUMreportingClarify in the Methods how missing or 'blank' physician notes were handled in the referral quality agreement analysis (e.g., whether they were excluded from the denominator).Unclear handling of missing data in a key analysis could affect the validity of the agreement results.
- 4.MEDIUMreportingState explicitly whether the trial was conducted in accordance with CONSORT guidelines, and provide the full statistical analysis plan as a supplementary file.Explicit adherence to reporting guidelines and a linked SAP would strengthen reproducibility and transparency.
- 5.MEDIUMdata codeProvide more details on the data access request process, such as expected response time and criteria for approval, in the Data Availability statement.A more concrete access process would facilitate independent validation and reuse of the data.
- 6.MEDIUMdata codeConsider making the PreA chatbot available for research purposes through a controlled access mechanism, and document its architecture and versioning in the Methods.Sharing the intervention would enable independent replication and validation, which is currently limited by commercial licensing.
- 7.MEDIUMstatisticsReport effect sizes with confidence intervals for secondary outcomes, not just p-values.Effect sizes with CIs provide a more informative measure of the magnitude and precision of effects than p-values alone.
- 8.MEDIUMreportingAdd a note on the generalizability of the findings to other healthcare settings beyond the two hospitals in western China.The study's context-specific nature warrants a caution to readers about external validity.
- 9.LOWcopyeditChange 'compared to' to 'compared with' in the Results section for formal academic style.Minor grammatical preference that improves academic tone.
- 10.LOWcopyeditFix the missing possessive apostrophe in the Discussion: change 'The data-tuned model replication' to 'The data-tuned model's replication'.Corrects a grammatical error that could confuse readers.
- 11.LOWcopyeditClarify in the Results that the 28.7% relative reduction refers specifically to consultation duration.Improves clarity and prevents misinterpretation of the effect size.
- 12.LOWcopyeditEnsure consistency in the Abstract and Table 1 for sex reporting (e.g., use parentheses consistently for counts and percentages).Minor style consistency improves readability.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.