A vaccine chatbot intervention for parents to improve HPV vaccination uptake among middle school girls: a cluster randomized trial.
Hou Z, Wu Z, Qu Z, Gong L, Peng H, Jit M, Larson HJ, Wu JT, Lin L
- DOI
- 10.1038/s41591-025-03618-6
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/e62cbd71-b512-4634-b8f4-855123287a12 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- CitationsUnresolved reference ×2−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 18 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported cluster randomized trial. The design is rigorous, ethics approvals are documented, and the statistical analysis is appropriate with effect sizes and CIs reported. Minor reporting gaps (outlier handling, assumption verification) and several copyedit issues (typos, percentage inconsistencies) do not undermine the core findings.
Both reviewers classified the study as interventional (cluster RCT), and this was adopted. The evaluation covered the full text, with N/A for non-applicable sub-criteria (e.g., animal/housing, cell lines). The statistics verification covered only 4 tests with test statistics/CIs; other p-values were not machine-verified. The citation check found 2 references not found in registries (arXiv preprints), which are flagged for verification.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 4 tests: 4 consistent, 0 inconsistent; 4 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary outcome: 92/1294 vs 25/1377, adjusted RR 3.85 (95% CI 2.48-5.97), p<0.001
“7.1% (92 out of 1,294) of parents in the chatbot group had scheduled or received an HPV vaccination for their daughters, compared with only 1.8% (25 out of 1,377) in the usual care group”
Taken as given: The counts 92 and 25 are the event counts in each group.; The denominators are 1294 and 1377 respectively.; The p-value is from a chi-square test on the 2x2 table.Method: Pearson chi-square test on the 2x2 table of events and non-events.How we recomputed it: pChi2x2(92, 1294-92, 25, 1377-25) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Secondary outcome: consultation 635/1294 vs 242/1377, adjusted RR 2.73 (95% CI 2.41-3.09), p<0.001
“49.1% of parents in the chatbot group consulting health professionals compared with 17.6% in the usual care group”
Taken as given: The counts 635 and 242 are the event counts in each group.; The denominators are 1294 and 1377 respectively.; The p-value is from a chi-square test on the 2x2 table.Method: Pearson chi-square test on the 2x2 table.How we recomputed it: pChi2x2(635, 1294-635, 242, 1377-242) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Primary outcome: HPV vaccine receipt or scheduled appointment, chatbot vs usual care (ITT).
“7.1% (92 out of 1,294) of parents in the chatbot group had scheduled or received an HPV vaccination for their daughters, compared with only 1.8% (25 out of 1,377) in the usual care group (Table ). The intention-to-treat (ITT) analysis indicated a statistically significant increase... ( P < 0.001).”
Taken as given: The 92 and 25 are the event counts in the chatbot and usual care groups, respectively.; The 1294 and 1377 are the total participants in each group.; The test used is a chi-square test (or GEE, but a simple chi-square is a reasonable approximation for the raw proportions).Method: Pearson chi-square test from 2x2 contingency table (events and non-events per group).How we recomputed it: pChi2x2(92, 1294-92, 25, 1377-25) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Secondary outcome: HPV vaccination-specific consultation, chatbot vs usual care.
“49.1% of parents in the chatbot group consulting health professionals compared with 17.6% in the usual care group... ( P < 0.001).”
Taken as given: The 635 and 242 are the event counts in the chatbot and usual care groups, respectively.; The 1294 and 1377 are the total participants in each group.; The test used is a chi-square test (or GEE approximation).Method: Pearson chi-square test from 2x2 contingency table.How we recomputed it: pChi2x2(635, 1294-635, 242, 1377-242)
- lowinternal contradictionTable 1 reports rural usual care percentage as 38.9% for 26 classes out of 90, which should be 28.9%.
“Rural | 51 (28.3) | 25 (27.8) | 26 (38.9)”
Table 1Find in source - lowinternal contradictionTable 2 reports a stray parenthesis in the age row for the chatbot group.
“Age, years (mean ± s.d.) | 40.4 ± 4.6 | 40.3 ± 4.4) | 40.5 ± 4.8”
Table 2Find in source
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
10 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewer 1The chatbot intervention significantly increased HPV vaccination receipt or scheduled appointment among daughters.The primary outcome analysis shows a statistically significant increase with adjusted RR 3.85 (95% CI 2.48-5.97), p<0.001.Evidence: Table 3: 7.1% vs 1.8%, adjusted RR 3.85 (95% CI 2.48-5.97), p<0.001.
“In intention-to-treat analyses, 7.1% of the intervention group met this outcome versus 1.8% of the control group ( P < 0.001)”
AbstractFind in source - supportedReviewers 1, 2The chatbot significantly increased HPV vaccination-specific consultations with health professionals.Secondary outcome analysis shows a significant increase with adjusted RR 2.73 (95% CI 2.41-3.09), p<0.001.Evidence: Table 3: 49.1% vs 17.6%, adjusted RR 2.73 (95% CI 2.41-3.09), p<0.001.
“there was a statistically significant increase in HPV vaccination-specific consultations with health professionals (49.1% versus 17.6%, P < 0.001)”
AbstractFind in source - supportedReviewer 1The chatbot enhanced vaccine literacy and rumor discernment among participants.HPV-related literacy scores improved significantly in the chatbot group compared to control, with p<0.001.Evidence: Table 3: HPV-related literacy mean change 0.7 vs <0.1, coefficient 0.70 (95% CI 0.52-0.88), p<0.001.
“along with enhanced vaccine literacy ( P < 0.001) and rumor discernment ( P < 0.001)”
AbstractFind in source - supportedReviewer 1The chatbot effectively increased vaccination and improved parental vaccine literacy.The primary and secondary outcomes support this claim, though the effect on vaccination is modest in absolute terms.Evidence: Primary outcome 7.1% vs 1.8%; literacy improvements significant.
“These findings indicate that the chatbot effectively increased vaccination and improved parental vaccine literacy”
AbstractFind in source - supportedReviewer 1The chatbot intervention was effective across diverse socioeconomic settings, with stronger increases in rural areas.Subgroup analysis shows significant effects in rural areas with RR 8.81 (95% CI 2.74-28.35).Evidence: Subgroup analysis: rural RR 8.81 (95% CI 2.74-28.35).
“with stronger increases in rural areas”
AbstractFind in source - supportedReviewer 2The chatbot intervention significantly increased HPV vaccine receipt or scheduled appointment among female middle school students.The primary outcome shows a statistically significant increase (7.1% vs 1.8%, adjusted RR 3.85, 95% CI 2.48-5.97, P<0.001) in the ITT analysis, supported by per-protocol analysis.Evidence: Table 3, Results, Primary outcome
“In intention-to-treat analyses, 7.1% of the intervention group met this outcome versus 1.8% of the control group ( P < 0.001) over a two-week intervention period.”
AbstractFind in source - supportedReviewer 2The chatbot enhanced vaccine literacy and rumor discernment.HPV-related literacy scores improved significantly (mean increase 0.7 points, 95% CI 0.52-0.88, P<0.001), with significant improvements in both knowledge and rumor screening subscales.Evidence: Table 3, Results, Secondary outcomes
“enhanced vaccine literacy ( P < 0.001) and rumor discernment ( P < 0.001) among participants using the chatbot.”
AbstractFind in source - supportedReviewer 2The chatbot was effective across diverse socioeconomic settings, with stronger effects in rural areas.Subgroup analysis shows significant effects in nearly all subgroups, with a notably higher RR in rural areas (8.81, 95% CI 2.74-28.35).Evidence: Figure 2, Results, Subgroup analysis
“In rural areas, vaccine receipt or scheduled appointment in the chatbot group was 8.81 times higher (95% CI 2.74–28.35) than that in the usual care group.”
ResultsFind in source - supportedReviewer 2Higher engagement with the chatbot was associated with greater vaccination uptake.Table 4 shows that high overall engagement (RR 2.16, 95% CI 1.35-3.44) and high interaction frequency (RR 2.06, 95% CI 1.28-3.31) were significantly associated with vaccination.Evidence: Table 4, Results, Chatbot engagement
“parents with high engagement levels in the chatbot intervention were 2.16 times (95% CI: 1.35–3.44) more likely to initiate HPV vaccination for their daughters than those with low engagement.”
ResultsFind in source - supportedReviewer 2The nurse persona was more effective than the expert persona in promoting vaccination.Table 4 shows that users who engaged with the nurse persona were 2.08 times (95% CI 1.25-3.45) more likely to initiate vaccination than those who only used the expert persona.Evidence: Table 4, Results, Chatbot engagement
“users who engaged with the nurse persona were 2.08 times (95% CI: 1.25–3.45) more likely to initiate HPV vaccination than those who exclusively interacted with the vaccine expert persona.”
ResultsFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary outcome is the receipt or scheduled appointment of the HPV vaccine, which is a hard clinical outcome (vaccination uptake). Although the trial uses a two-week intervention period and includes scheduled appointments as part of the primary outcome, the receipt of the vaccine is verified using official vaccination records. The primary outcome is not a surrogate biomarker but a direct measure of vaccination behavior.
“The primary outcome was the receipt or scheduled appointment of the HPV vaccine, determined by whether the participants’ daughters were vaccinated or had actively scheduled a vaccination in the two-week intervention period. Receipt of an HPV vaccine was verified using official vaccination records.”
- ADEQUATEEffect sizeThe primary outcome showed a statistically significant increase from 1.8% in the control group to 7.1% in the intervention group, with an adjusted relative risk of 3.85 (95% CI 2.48–5.97). The absolute increase of 5.3 percentage points is clinically meaningful in the context of HPV vaccination, where baseline uptake is low. The effect size is anchored to a meaningful clinical outcome (vaccination receipt or scheduled appointment) and is statistically supported.
“In intention-to-treat analyses, 7.1% of the intervention group met this outcome versus 1.8% of the control group (P < 0.001) ... parents in the chatbot group were 3.85 times (adjusted relative risk: 3.85, 95% confidence interval (CI) 2.48–5.97) more likely to initiate HPV vaccination (either by scheduling or receiving the vaccine) than those in the usual care group (P < 0.001).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on HPV burden, low vaccination coverage in China, parental hesitancy, and the potential of chatbots, while acknowledging limitations of prior work (e.g., lack of studies in non-routine immunization contexts with out-of-pocket costs). The rationale logically connects these gaps to the study's aim of evaluating a chatbot in China's context. Limitations of prior research are addressed by the study design (e.g., cluster RCT, diverse socioeconomic settings).
“Cervical cancer remains a major global health challenge, with 662,301 new cases and 348,874 deaths reported worldwide in 2022”
“Our study addresses these gaps by examining how a chatbot intervention might influence HPV vaccination in China, where substantial cost barriers exist”
“Our study addresses these gaps by examining how a chatbot intervention might influence HPV vaccination in China, where substantial cost barriers exist, while also exploring the pathways through which such digital intervention can impact vaccine-related decision-making.”
Randomization was computer-generated, stratified by region, school, and grade, with classes as the unit. Blinding of participants and implementers was not possible due to the intervention nature, but randomization was blinded to schools, teachers, and participants. A priori power analysis is reported with assumptions (ICC, design effect, dropout). Inclusion/exclusion criteria are pre-specified. Outlier handling is not explicitly discussed but the analysis uses robust methods (GEE, mixed models). Controls are appropriate (usual care). Independent replication is not reported (n/a for a single trial).
“Using class lists in each grade in a school, computer-generated randomization was used to determine whether class 1 and class 2 were assigned to intervention or control group”
“The nature of the intervention did not allow for masking of the intervention to class teachers, participants or study implementers.”
“Assuming a power of 80%, a two-sided significance level of 0.05 and incorporating this cluster design effect, we determined that a sample size of 648 participants per arm would be sufficient”
“The nature of the intervention did not allow for masking of the intervention to class teachers, participants or study implementers.”
Sex of participants (mother/father) and daughters (all female) is reported. The study focuses on female students, so single-sex is justified by the research question. Age, grade, region, education, income, and other demographics are reported in Table 2. Species/strain/housing are not applicable (human study).
“Mothers constituted 87.7% of participants”
“Age, years (mean ± s.d.) | 13.1 ± 1.1”
“Age, years (mean ± s.d.) | 13.1 ± 1.1”
“A vaccine chatbot intervention for parents to improve HPV vaccination uptake among middle school girls”
The study received approval from the IRB of Fudan University School of Public Health and the Human Research Ethics Committee of the University of Hong Kong, with an approval number. Informed consent is described in detail, including a 14-day opt-out period. Regulatory compliance is implied through adherence to guidelines.
“This study was a cluster randomized trial approved by both the Institutional Review Board (IRB) of Fudan University School of Public Health and the Human Research Ethics Committee of the University of Hong Kong.”
“All participants provided informed consent to participate this trial.”
“This study was conducted in alignment with evidence-based clinical practice guidelines”
“This study was conducted in alignment with evidence-based clinical practice guidelines and received approval from the IRB of Fudan University School of Public Health and Human Research Ethics Committee of the University of Hong Kong.”
The chatbot is described in detail (knowledge base, roles, GPT-4, RAG, prompt engineering) with a URL. Statistical software (STATA v.15.1, R v.4.4.1) is identified. No antibodies, cell lines, or organisms are used (n/a). Reagents are not applicable (no wet-lab components).
“In this study, we developed an AI-powered chatbot tailored for HPV vaccine consultation, designed specifically for the Chinese context.”
“Statistical analyses were performed using STATA v.15.1 and R v.4.4.1 statistical software.”
“The AI-powered chatbot is accessible via WeChat and through web browsers at https://hpvchatbot.social-insight.ai .”
“Statistical analyses were performed using STATA v.15.1 and R v.4.4.1 statistical software.”
“our chatbot uses advanced linguistic technologies through GPT-4 (ref. ) coupled with retrieval-augmented generation and prompt engineering”
Tests are named (GEE for categorical, mixed-effects for continuous, log-binomial for engagement). Assumptions are not explicitly verified but the methods (GEE, mixed models) are robust to typical violations. Exact p-values are reported (e.g., P < 0.001). Effect sizes with 95% CIs are reported for all outcomes. Software is identified. Data presentation includes per-group n, percentages, means, SDs, and CIs. Mathematical plausibility checks: the primary outcome counts (92/1294 = 7.1%, 25/1377 = 1.8%) are consistent; subgroup counts in Table 2 sum to totals (e.g., 322+956+788+605 = 2671).
“Mixed-effects models and generalized estimating equations (GEE) were used to evaluate the effectiveness of the chatbot intervention”
The data availability statement explains that data are not publicly available due to privacy/consent but provides a mechanism for access (request to corresponding authors, one-month timeline, data use agreement). This is adequate for patient-level data. Code is shared on GitHub (https://github.com/wu-zhengdong/HPV-vaccine-chatbot.git). Repository deposit and accession numbers are not applicable (no sequencing or depositable non-identifiable data).
“Access to data will be provided upon application, with a timeline of one month determined in accordance with the request.”
“All codes are freely available on GitHub at https://github.com/wu-zhengdong/HPV-vaccine-chatbot.git”
“All codes are freely available on GitHub at https://github.com/wu-zhengdong/HPV-vaccine-chatbot.git .”
Methods include trial design, setting, participants, intervention, outcomes, sample size, and statistical analysis in detail. Trial registration number is provided. CONSORT checklist is mentioned in supplementary materials. All pre-specified outcomes (primary and secondary) are reported with results. Limitations are discussed in detail (out-of-pocket costs, short duration, cross-contamination, etc.). Conclusions are proportional (e.g., 'further research is necessary to scale and sustain these gains'). Funding sources and COI are stated.
“Clinical trial registration: NCT06227689”
“Supplementary Document 4. CONSORT checklist.”
“The study has several limitations.”
“Clinical trial registration: NCT06227689 (https://clinicaltrials.gov/ct2/show/NCT06227689) .”
“Supplementary Document 4. CONSORT checklist.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 55 references by DOI: 2 verified — 2 DOI unresolved, 51 no DOI (shown, not verified).
- UNRESOLVED10.48550/arxiv.2303.08774GPT-4 technical reportCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.48550/arxiv.2311.05112A survey of large language models in medicine: progress, application, and challengeCited DOI does not resolve to any Crossref record.
- NO DOICervix uteri. Fact sheetNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIChina. Fact sheetNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITrend in cervical cancer incidence and mortality rates in China, 2006–2030: a Bayesian age-period-cohort modeling studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITemporal trends and projection of cancer attributable to human papillomavirus infection in China, 2007–2030No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHuman papillomavirus and cervical cancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe effects of the national HPV vaccination programme in England, UK, on cervical cancer and grade 3 cervical intraepithelial neoplasia incidence: a register-based observational studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDomestic HPV vaccine price and economic returns for cervical cancer prevention in China: a cost-effectiveness analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEstimated human papillomavirus vaccine coverage among females 9–45 years of age — China, 2017–2022No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBetween now and later: a mixed methods study of HPV vaccination delay among Chinese caregivers in urban Chengdu, ChinaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAccelerating the elimination of cervical cancer as a public health problem: Towards achieving 90–70–90 targets by 2030No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe status and challenges of HPV vaccine programme in China: an exploration of the related policy obstaclesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAwareness and knowledge about human papillomavirus vaccination and its acceptance in China: a meta-analysis of 58 observational studiesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPreference for human papillomavirus vaccine type and vaccination strategy among parents of school-age girls in Guangdong province, ChinaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIParental willingness of HPV vaccination in Mainland China: a meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWillingness and hesitancy towards the governmental free human papillomavirus vaccination among parents of eligible adolescent girls in Shenzhen, Southern ChinaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWillingness to pay for HPV vaccine among female health care workers in a Chinese nationwide surveyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAwareness of HPV and HPV vaccines, acceptance to vaccination and its influence factors among parents of adolescents 9 to 18 years of age in China: a cross-sectional studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA case-control study on factors of HPV vaccination for mother and daughter in ChinaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIKnowledge, attitude, and uptake of human papillomavirus (HPV) vaccination among Chinese female adults: a national cross-sectional web-based survey based on a large e-commerce platformNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIKnowledge about human papillomavirus (HPV) and HPV vaccine and willingness for their children′s vaccination among parents of 9–14 years old girls in Hangzhou cityNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA nationwide post-marketing survey of knowledge, attitude and practice toward human papillomavirus vaccine in general population: Implications for vaccine roll-out in mainland ChinaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIChinese mothers’ intention to vaccinate daughters against human papillomavirus (HPV), and their vaccine preferences: a study in Fujian ProvinceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDiscrepancy of human papillomavirus vaccine uptake and intent between girls 9–14 and their mothers in a pilot region of Shanghai, ChinaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILarge language models in medicineNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIConversational AI and vaccine communication: systematic review of the evidenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGenerative artificial intelligence can have a role in combating vaccine hesitancyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAI in healthcare: navigating opportunities and challenges in digital communicationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIChatGPT and vaccines: can AI chatbots boost awareness and uptake?No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIChatbot-delivered COVID-19 vaccine communication message preferences of young adults and public health workers in urban American communities: qualitative studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEarly usability assessment of a conversational agent for HPV vaccinationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExamining potential usability and health beliefs among young adults using a conversational agent for HPV vaccine counselingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffectiveness of chatbots in increasing uptake, intention, and attitudes related to any type of vaccination: a systematic review and meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFeasibility and acceptability of Saheli, a WhatsApp chatbot, on COVID-19 vaccination among pregnant and breastfeeding women in rural North IndiaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIChatbot-delivered online intervention to promote seasonal influenza vaccination during theCOVID-19 pandemic: a randomized clinical trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEfficacy, usability, and acceptability of a chatbot for promoting COVID-19 vaccination in unvaccinated or booster-hesitant young adults: pre–post pilot studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInformation delivered by a chatbot has a positive impact on COVID-19 vaccines attitudes and intentionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffectiveness of chatbots on COVID vaccine confidence and acceptance in Thailand, Hong Kong, and SingaporeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDigital health interventions to improve adolescent HPV vaccination: a systematic reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence-based chatbots for promoting health behavioral changes: systematic reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMeasuring the impact of COVID-19 vaccine misinformation on vaccination intent in the UK and USANo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOI‘You don’t know if it’s the truth or a lie’: exploring human papillomavirus (HPV) vaccine hesitancy among communities with low HPV vaccine uptake in Northern CaliforniaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITackling barriers to scale up human papillomavirus vaccination in China: progress and the way forwardNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWhen knowledge is not enough: changing behavior to change vaccination resultsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAttitudes and personal beliefs about the COVID-19 vaccine among people with COVID-19: a mixed-methods analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIndividual and social determinants of COVID-19 vaccine uptakeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPerceived benefits and barriers to Chinese COVID-19 vaccine uptake among young adults in ChinaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIConsiderations for addressing bias in artificial intelligence for health equityNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence in health care: will the value match the hype?No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMapping global trends in vaccine confidence and investigating barriers to vaccine uptake: a large-scale retrospective temporal modelling studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHPV.edu study protocol: a cluster randomised controlled evaluation of education, decisional support and logistical strategies in school-based human papillomavirus (HPV) vaccination of adolescentsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntracluster correlation coefficients from school-based cluster randomized trials of interventions for improving health outcomes in pupilsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- codeGitHubLIVEHTTP 200https://github.com/wu-zhengdong/HPV-vaccine-chatbot.gitResolves to GitHub (code repository).
Copyediting
10 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 10 minor suggestions below.
10 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORtypoIntroduction, paragraph 1“heavey”→ heavyTypographical error.
- MINORconsistencyTable 1“Rural | 51 (28.3) | 25 (27.8) | 26 (38.9)”→ Check percentage for usual care rural (should be 28.9% if 26/90)Percentage appears inconsistent with the count.
- MINORconsistencyTable 2“Age, years (mean ± s.d.) | 40.4 ± 4.6 | 40.3 ± 4.4) | 40.5 ± 4.8”→ Remove stray parenthesis after 4.4Formatting error.
- MINORconsistencyTable 2“1000,000–200,000 CNY”→ 100,000–200,000 CNYTypographical error in income range.
- MINORclarityMethods, Statistical analysis“with stepwise reduction of covariates used to ensure model convergence when needed”→ Clarify the stepwise procedure and criteria for covariate reduction.Could be more specific for reproducibility.
- MINORtypoAbstract, line 2“heavey”→ heavyTypo in 'heavey disease burden'.
- MINORtypoTable 2, income category“1000,000–200,000 CNY”→ 100,000–200,000 CNYExtra zero in '1000,000'.
- MINORconsistencyTable 1, region row“Rural | 51 (28.3) | 25 (27.8) | 26 (38.9)”→ 26 (28.9) for usual carePercentage for rural usual care is 26/90 = 28.9%, not 38.9% as printed. This appears to be a typo.
- MINORgrammarMethods, Study procedures“Participants provided informed consent to participate this trial”→ Participants provided informed consent to participate in this trialMissing preposition 'in'.
- MINORclarityResults, Subgroup analysis“vaccine receipt or scheduled appointment in the chatbot group was 8.81 times higher (95% CI 2.74–28.35) than that in the usual care group”→ vaccine receipt or scheduled appointment was 8.81 times higher (95% CI 2.74–28.35) in the chatbot group than in the usual care groupSlight rephrase for clarity.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (outlier handling, assumption verification) and the copyedit inconsistencies, but none threaten the validity of the conclusions. No erratum is warranted for the core findings, though the authors should correct the typographical and percentage errors in the tables.
- 1.HIGHcopyeditCorrect the percentage for rural usual care in Table 1 from 38.9% to 28.9% (26/90).The printed percentage is internally inconsistent with the count and could be misread as a data error.
- 2.HIGHcopyeditFix the income range typo in Table 2: change '1000,000–200,000 CNY' to '100,000–200,000 CNY'.The extra zero is a clear typographical error that could confuse readers.
- 3.HIGHcopyeditRemove the stray parenthesis in Table 2 age row: '40.3 ± 4.4)' should be '40.3 ± 4.4'.Formatting error in a key demographic table.
- 4.HIGHcopyeditFix the typo 'heavey' to 'heavy' in the Abstract and Introduction.Typographical error in the opening lines.
- 5.HIGHreportingAdd a statement in the Methods (Statistical analysis) on how outliers were handled (e.g., whether any data points were excluded and why).One reviewer flagged outlier handling as not explicitly reported; this is a transparency gap for a cluster RCT.
- 6.HIGHstatisticsExplicitly verify and report test assumptions (normality, equal variance) for continuous outcomes (e.g., HPV literacy scores) or justify the use of mixed-effects models without such verification.Assumption verification was rated as reported_but_inadequate; adding this strengthens statistical transparency.
- 7.HIGHreportingVerify the two references not found in registries (GPT-4 technical report, arXiv:2303.08774; and 'A survey of large language models in medicine', arXiv:2311.05112) and correct or replace them if they are not accessible.These references could not be located in Crossref/OpenAlex and may be fabricated or incorrectly cited.
- 8.MEDIUMreportingReport the intracluster correlation coefficient (ICC) for the primary outcome in the Results section.The ICC was used in the sample size calculation but not reported in the results, which is a useful detail for readers.
- 9.MEDIUMreportingClarify the stepwise covariate reduction procedure in the Methods (Statistical analysis), specifying the criteria used.The current description is vague and not fully reproducible.
- 10.MEDIUMreportingAdd a statement on how missing data were handled (e.g., complete-case analysis, imputation) for primary and secondary outcomes.Missing data handling is not described, which is a common reviewer concern.
- 11.MEDIUMreportingReport exact p-values for all secondary outcomes, including non-significant ones, rather than only thresholds.Exact p-values improve transparency and allow readers to assess borderline results.
- 12.MEDIUMreportingProvide a CONSORT-style flow diagram in the main text (currently only referenced as Fig. 1) to improve transparency.A flow diagram is a key CONSORT element and aids reader comprehension of participant flow.
- 13.MEDIUMdata codeConsider providing a more detailed data availability statement that includes a managed-access platform (e.g., Vivli, YODA) for patient-level data.A managed-access platform would enhance data sharing credibility beyond author requests.
- 14.MEDIUMreportingClarify the exact version of GPT-4 used (e.g., GPT-4-turbo) and whether any fine-tuning was performed on the chatbot's knowledge base.This improves reproducibility of the intervention.
- 15.MEDIUMreportingDescribe the randomization sequence generation in more detail (e.g., software used, seed) to enhance reproducibility.Details on randomization sequence generation are important for reproducibility.
- 16.LOWcopyeditFix grammar in Methods, Study procedures: 'participate this trial' → 'participate in this trial'.Missing preposition is a minor grammar error.
- 17.LOWcopyeditRephrase the subgroup analysis sentence in Results for clarity: 'vaccine receipt or scheduled appointment was 8.81 times higher (95% CI 2.74–28.35) in the chatbot group than in the usual care group'.The current phrasing is slightly awkward and could be clearer.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.