Model of integrated mental health video consultations for people with depression or anxiety in primary care (PROVIDE-C): assessor masked, multicentre, randomised controlled trial.
Haun MW, Tönnies J, Hartmann M, Wildenauer A, Wensing M, Szecsenyi J, Feißt M, Pohl M, Vomhof M, Icks A, Friederich HC
- DOI
- 10.1136/bmj-2024-079921
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/b00c65ba-24c6-4314-8824-00305329f12c is authoritative.
How this rating was calculated
- CitationsUnresolved reference ×3−0.75★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 83 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is the PHQ-ADS score, a patient-reported symptom scale, which is a surrogate for clinical benefit. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure) nor cite validated evidence linking changes in PHQ-ADS to hard clinical outcomes. The minimal clinically important difference is mentioned but not validated as a surrogate for long-term outcomes.
“The primary outcome was the absolute change in the mean severity of depressive and anxiety symptoms measured using the patient health questionnaire anxiety and depression scale (PHQ-ADS) at six months”
- 02Treatment effect not shown to be clinically meaningful
The reported effect size is small (Cohen's d = 0.21 at 6 months) and the adjusted mean difference (-2.4 points) is below the minimal clinically important difference of 3-5 points. The paper acknowledges this but argues population-level impact, yet the effect is not anchored to a clinically meaningful threshold for individual patients.
“The difference in the change of 2.3 points was lower than that of the minimal clinically important difference in the PHQ-ADS score (3-5 point change).”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomised controlled trial with rigorous design, clear ethical approvals, and appropriate statistical methods. Minor reporting gaps include lack of explicit Declaration of Helsinki mention and a few copyedit issues, but these do not undermine the scientific integrity.
Both reviewers classified the study as interventional; no divergence. The evaluation covered the full text, with statistics verification limited to 6 recomputable tests (all consistent) and citation check finding 3 references not found in registry (potential fabrication signals).
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 6 tests: 6 consistent, 0 inconsistent; 2 recomputed directly from the reported test statistics, 4 via agent-written checks.
- CONSISTENTreported p = .480 · recomputed p = .485Recomputed t(153)=0.70, P=0.48
“t(153)=0.70, P=0.48”
Taken as given: the printed df is 153, and it is the df of this statistic rather than of another test in the same sentence; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed t statistic and its df and compare it against the printed pHow we recomputed it: pT(0.7, 153) - CONSISTENTreported p = .070 · recomputed p = .074Recomputed t (153)=1.80, P=0.07
“t (153)=1.80, P=0.07”
Taken as given: the printed df is 153, and it is the df of this statistic rather than of another test in the same sentence; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed t statistic and its df and compare it against the printed pHow we recomputed it: pT(1.8, 153) - CONSISTENTreported p = .020 · recomputed p = .022Reviewer 1Primary outcome adjusted mean change difference p-value from CI
“adjusted mean change difference −2.4 points (−4.5 to −0.4), P=0.02”
Taken as given: The estimate is -2.4 and the 95% CI is -4.5 to -0.4.; The CI is two-sided at 95%.; The estimate is a difference in means, not a ratio.Method: Recomputed two-tailed p from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-2.4, -4.5, -0.4, 0) - CONSISTENTreported p = .007 · recomputed p = .008Reviewer 112-month primary outcome p-value from CI
“mean change difference −2.9 (−5.0 to −0.7), P=0.007”
Taken as given: The estimate is -2.9 and the 95% CI is -5.0 to -0.7.; The CI is two-sided at 95%.; The estimate is a difference in means, not a ratio.Method: Recomputed two-tailed p from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-2.9, -5.0, -0.7, 0) - CONSISTENTreported p = .020 · recomputed p = .022Reviewer 2Primary outcome adjusted mean change difference at 6 months
“adjusted mean change difference −2.4 points (95% confidence interval −4.5 to −0.4), P=0.02”
Taken as given: The estimate is -2.4 and the 95% CI is -4.5 to -0.4.; The CI is two-sided at 95%.; The estimate is on a linear scale (not log).Method: Recomputed two-sided p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(-2.4, -4.5, -0.4, 0) - CONSISTENTreported p = .007 · recomputed p = .008Reviewer 2Primary outcome adjusted mean change difference at 12 months
“mean change difference −2.9 (−5.0 to −0.7), P=0.007”
Taken as given: The estimate is -2.9 and the 95% CI is -5.0 to -0.7.; The CI is two-sided at 95%.; The estimate is on a linear scale.Method: Recomputed two-sided p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(-2.9, -5.0, -0.7, 0)
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2The small effect might cumulatively impact on population health in this population.The claim is speculative and not directly tested; it is an inference from the small effect size and high prevalence, but the paper does not provide direct evidence of population-level impact.Evidence: Discussion: 'Small effects that accumulate over time and at scale are highly likely to be consequential from a public health perspective.'
“Depression and anxiety disorders are prevalent and therefore the small effect might cumulatively impact on population health in this population.”
ConclusionFind in source - supportedReviewers 1, 2The PROVIDE intervention led to improvements in severity of depressive and anxiety symptoms at six months compared with usual care.The primary outcome analysis shows a statistically significant adjusted mean change difference of -2.4 points (95% CI -4.5 to -0.4, P=0.02), supporting the claim.Evidence: Primary outcome result: adjusted mean change difference −2.4 points (−4.5 to −0.4), P=0.02
“Compared with usual care, the PROVIDE intervention led to improvements in severity of depressive and anxiety symptom (adjusted mean change difference in the PHQ-ADS score −2.4 points (95% confidence interval −4.5 to −0.4), P=0.02) at six months.”
AbstractFind in source - supportedReviewers 1, 2The effects were sustained at 12 months.The 12-month analysis shows a significant difference of -2.9 points (95% CI -5.0 to -0.7, P<0.01), supporting the claim.Evidence: 12-month result: mean change difference −2.9 (−5.0 to −0.7), P<0.01
“The effects were sustained at 12 months (−2.9 (−5.0 to −0.7), P<0.01).”
AbstractFind in source - supportedReviewers 1, 2No serious adverse events were reported in either group.The paper states no serious adverse events were reported, and the harms section confirms this.Evidence: Harms section: 'No serious adverse events attributable to trial participation were reported in either group during the study.'
“No serious adverse events were reported in either group.”
AbstractFind in source - supportedReviewers 1, 2The PROVIDE model led to a decrease in depressive and anxiety symptoms with small effects in the short and long term.The effect sizes are small (Cohen's d 0.21 at 6 months, 0.30 at 12 months) and statistically significant, supporting the claim.Evidence: Effect sizes: Cohen's d 0.21 (95% CI 0.03 to 0.39) at 6 months; 0.30 (95% CI 0.08 to 0.52) at 12 months
“Through relatively low intensity treatment, the PROVIDE model led to a decrease in depressive and anxiety symptoms with small effects in the short and long term.”
ConclusionFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is the PHQ-ADS score, a patient-reported symptom scale, which is a surrogate for clinical benefit. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure) nor cite validated evidence linking changes in PHQ-ADS to hard clinical outcomes. The minimal clinically important difference is mentioned but not validated as a surrogate for long-term outcomes.
“The primary outcome was the absolute change in the mean severity of depressive and anxiety symptoms measured using the patient health questionnaire anxiety and depression scale (PHQ-ADS) at six months”
- INADEQUATEEffect sizeThe reported effect size is small (Cohen's d = 0.21 at 6 months) and the adjusted mean difference (-2.4 points) is below the minimal clinically important difference of 3-5 points. The paper acknowledges this but argues population-level impact, yet the effect is not anchored to a clinically meaningful threshold for individual patients.
“The difference in the change of 2.3 points was lower than that of the minimal clinically important difference in the PHQ-ADS score (3-5 point change).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites a systematic review of co-located care and notes the need for more rigorous RCTs, and mentions prior pilot work (PROVIDE-B). The rationale links the gap in scalable integrated mental health video consultations to the study objective. Limitations of prior research (e.g., conducted in highly regulated environments) are explicitly acknowledged.
“In 2020, the results from a systematic review of 15 studies showed that co-located specialty care was associated with mental health benefits, and concluded that more rigorous randomised controlled trials are needed.”
“The aim of this assessor masked randomised controlled trial was to investigate the effectiveness of this new mental health service model for treating people with depression or anxiety, or both, in primary care settings.”
“The limited number of published randomised controlled trials to date were conducted in highly regulated environments, such as the US Veterans Health Care Administration, or involved patients from inpatient facilities.”
“In 2020, the results from a systematic review of 15 studies showed that co-located specialty care was associated with mental health benefits, and concluded that more rigorous randomised controlled trials are needed.”
“The aim of this assessor masked randomised controlled trial was to investigate the effectiveness of this new mental health service model for treating people with depression or anxiety, or both, in primary care settings.”
“The potential scalability of these models to primary care settings, particularly in countries where smaller, single handed, or rural and remote practices dominate, remains uncertain.”
Randomization used a secure web-based system with computer-generated sequence, stratified by centre and severity, with random permuted blocks. Blinding of outcome assessors and data analysts is described, with a note on unintentional unmasking. Sample size calculation is provided with effect size, alpha, power, and dropout adjustment. Inclusion/exclusion criteria are detailed. Outlier handling is addressed through graphical evaluation of assumptions and multiple imputation for missing data. Controls are appropriate (usual care). Independent replication is not applicable for a single pivotal trial.
“Eligible participants were then randomly assigned (1:1) to the intervention or control group via a secure web based randomisation system (Randomiser V.2.0.2) operated by a data manager who was not involved in patient recruitment, centrally at the Institute of Medical Biometry, Heidelberg University.”
“While the patients, GPs, and mental health specialists were aware of the intervention assignment after allocation, the data analysts were masked to the allocation.”
“To detect the minimal clinically important difference in the PHQ-ADS score of 3 points (SD 9 points) with a two sided 5% significance level and a power of 80%, a sample size of 160 patients per group was necessary.”
“Eligible participants were then randomly assigned (1:1) to the intervention or control group via a secure web based randomisation system (Randomiser V.2.0.2) operated by a data manager who was not involved in patient recruitment, centrally at the Institute of Medical Biometry, Heidelberg University.”
“While the patients, GPs, and mental health specialists were aware of the intervention assignment after allocation, the data analysts were masked to the allocation.”
“To detect the minimal clinically important difference in the PHQ-ADS score of 3 points (SD 9 points) with a two sided 5% significance level and a power of 80%, a sample size of 160 patients per group was necessary.”
Sex is reported (63% female), age (mean 45, SD 14), and health status (chronic physical disease, symptom severity). Demographics include marital status, education, employment, income, and psychiatric history. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“Mean age was 45 years (standard deviation (SD) 14), 63% of the participants were female”
“A total of 220 (59%) participants had at least one chronic physical disease.”
“Of the 376 participants, 238 (63%) participants were female. The mean age was 45 years (SD 14; range 18-81).”
“A total of 220 (59%) participants had at least one chronic physical disease.”
“Table 1 Baseline characteristics of the intention-to-treat population”
The trial was approved by the Medical Faculty of the University of Heidelberg Ethics Committee (S-923/2019) on 7 January 2020. All participants provided written informed consent. Regulatory compliance is implied through adherence to CONSORT and ethical standards, though not explicitly naming a framework like the Declaration of Helsinki.
“The trial protocol was approved by the Medical Faculty of the University of Heidelberg Ethics Committee (S-923/2019) on 7 January 2020.”
“All participants provided written informed consent to participate in this trial.”
“We reported the PROVIDE-C trial in accordance with the CONSORT 2010 statement.”
“The trial protocol was approved by the Medical Faculty of the University of Heidelberg Ethics Committee (S-923/2019) on 7 January 2020.”
“All participants provided written informed consent to participate in this trial.”
“We reported the PROVIDE-C trial in accordance with the CONSORT 2010 statement.”
The intervention is described in detail, including the videoconferencing platform (arztkonsultation ak GmbH) and the intervention manual. The statistical software (R 4.4.0) is identified. No antibodies, cell lines, or organisms are used, so those criteria are not applicable.
“The intervention was delivered through individual, synchronous one-to-one video consultations conducted via an encrypted, web based videoconferencing platform on a subscription basis (arztkonsultation ak GmbH, Schwerin, Germany, https://arztkonsultation.de ).”
“The analyses were performed using R 4.4.0 or higher.”
“The intervention was delivered through individual, synchronous one-to-one video consultations conducted via an encrypted, web based videoconferencing platform on a subscription basis (arztkonsultation ak GmbH, Schwerin, Germany, https://arztkonsultation.de ).”
“The analyses were performed using R 4.4.0 or higher.”
“via a secure web based randomisation system (Randomiser V.2.0.2)”
The primary analysis uses a mixed linear model with multiple imputation, and assumptions are graphically evaluated. Exact p-values are reported (e.g., P=0.02). Effect sizes with 95% CIs are provided. Statistical software is identified. Data presentation includes per-group n and SDs. Mathematical plausibility checks were not possible for all values due to model-based estimates, but no obvious errors were found.
“We analysed the primary outcome with a mixed linear model, where a random intercept accounted for the primary care practice to which the patient belonged.”
“adjusted mean change difference −2.4 points (−4.5 to −0.4), P=0.02”
“The effect size (Cohen’s d) was 0.21 (95% CI 0.03 to 0.39).”
“We analysed the primary outcome with a mixed linear model, where a random intercept accounted for the primary care practice to which the patient belonged.”
“adjusted mean change difference −2.4 points (−4.5 to −0.4), P=0.02”
“The effect size (Cohen’s d) was 0.21 (95% CI 0.03 to 0.39).”
The data availability statement describes a managed-access process: deidentified participant data will be available nine months after publication to investigators whose proposed use is approved by an independent review committee. This is adequate for patient-level data. No public repository deposit or accession numbers are applicable for identifiable patient data. Code sharing is not applicable as no bespoke code is mentioned.
“For individual participant data meta-analysis, beginning nine months following article publication, deidentified participant data that underlie the results reported in this article (text, tables, figures, and appendices) and a data dictionary will be available to investigators whose proposed use of the data has been approved by an independent review committee identified for this purpose.”
“For individual participant data meta-analysis, beginning nine months following article publication, deidentified participant data that underlie the results reported in this article (text, tables, figures, and appendices) and a data dictionary will be available to investigators whose proposed use of the data has been approved by an independent review committee identified for this purpose.”
Trial registration is provided (NCT04316572). CONSORT 2010 is referenced. All pre-specified outcomes are reported, including non-significant ones. Limitations are thoroughly discussed. Conclusions are proportional to the evidence, acknowledging the small effect size. Funding and competing interests are declared.
“Trial registration ClinicalTrials.gov NCT04316572 (https://clinicaltrials.gov/ct2/show/NCT04316572) .”
“We reported the PROVIDE-C trial in accordance with the CONSORT 2010 statement.”
“This study has several limitations. Selection bias is a pressing issue in practice based clinical research.”
“Trial registration ClinicalTrials.gov NCT04316572 (https://clinicaltrials.gov/ct2/show/NCT04316572) .”
“We reported the PROVIDE-C trial in accordance with the CONSORT 2010 statement.”
“Funding: This trial was funded by a grant from the German Federal Ministry of Education and Research (BMBF) (grant no. 01GY16129).”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 72 references by DOI: 65 verified — 3 DOI unresolved, 4 no DOI (shown, not verified).
- UNRESOLVED10.1016/s0140-6736(18World mental health report: transforming mental health for allCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1016/s0140-6736(07Primary Care Mental HealthCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1002/14651858.cd000532.pub2/abstractOn‐site mental health workers delivering psychological therapy and psychosocial interventions to patients in primary care: effects on the professional practice of primary care providersCited DOI does not resolve to any Crossref record.
- NO DOIThe somatising effect of clinical consultation: What patients and doctors say and do not say when patients present medically unexplained physical symptomsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAchieving evidence-based psychotherapy practice: a psychodynamic perspective on the general acceptance of treatment manualsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe PRECIS-2 tool: designing trials that are fit for purposeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPatient preferences for specialist outpatient video consultations: A discrete choice experimentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT04316572LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
7 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 7 minor suggestions below.
7 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAbstract, Results“The effects were sustained at 12 months (−2.9 (−5.0 to −0.7), P<0.01).”→ Ensure consistent use of en-dash or hyphen in ranges.Minor formatting inconsistency.
- MINORconsistencyTable 2, footnote“Effect size measured with Cohen’s d.”→ Clarify whether effect sizes are for the minimally adjusted or adjusted model.The table has two effect size columns; the footnote is ambiguous.
- MINORclarityMethods, Statistical analysis“We decided to perform a responder analysis comparing the proportion of participants within each study arm who had a change at least as large as the minimal clinically important difference as part of the secondary analyses.”→ Rephrase for clarity: 'We decided to perform a responder analysis, comparing the proportion of participants within each study arm who had a change at least as large as the minimal clinically important difference, as part of the secondary analyses.'The sentence is long and could be clearer.
- MINORtypoAbstract, Results“The effects were sustained at 12 months (−2.9 (−5.0 to −0.7), P<0.01).”→ Consider adding 'points' for clarity: '−2.9 points (−5.0 to −0.7)'.Minor clarity issue.
- MINORconsistencyTable 2, row 'Intention to treat'“−2.43 (−4.48 to −0.38)”→ Ensure the primary outcome estimate in the text (−2.4) matches the table (−2.43) consistently.Rounding difference; consider standardizing to one decimal place.
- MINORtypoTable 3, row '12 item short form survey'“1.13 (−0.83 to 3.09)‡”→ Check the footnote symbol placement; it may be misplaced.Formatting issue.
- MINORgrammarDiscussion, Strengths and weaknesses“The trial has substantial strengths. The sample predominantly included female participants and middle class individuals with free access to health insurance.”→ Consider rephrasing to avoid a sentence fragment.Minor style issue.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (explicit regulatory framework, rounding inconsistencies) and the three unresolved references as potential concerns, but none warrant immediate erratum. The statistics are consistent for the subset checked, and the trial is registered and follows CONSORT.
- 1.HIGHreportingVerify or correct the three references not found in any registry: 'World mental health report: transforming mental health for all' (DOI 10.1016/s0140-6736(18), 'Primary Care Mental Health' (DOI 10.1016/s0140-6736(07), and 'On‐site mental health workers delivering psychological therapy and psychosocial interventions to patients in primary care' (DOI 10.1002/14651858.cd000532.pub2/abstract).References that cannot be located in Crossref/OpenAlex may be fabricated or contain incorrect DOIs, which is a serious integrity concern.
- 2.HIGHethicsIn the Ethics statements, explicitly name the regulatory framework (e.g., Declaration of Helsinki) and whether the trial followed ICH-GCP or equivalent.One reviewer flagged regulatory compliance as inadequate because the framework is not explicitly named; adding this strengthens the ethics reporting.
- 3.MEDIUMreportingStandardize the primary outcome estimate to one decimal place across text and Table 2 (text: −2.4; table: −2.43) and clarify the footnote on effect sizes in Table 2.Inconsistent rounding and ambiguous footnotes can confuse readers and are easily fixed.
- 4.MEDIUMcopyeditFix the sentence fragment in the Discussion, Strengths and weaknesses: 'The trial has substantial strengths. The sample predominantly included female participants...' by merging or rephrasing.Grammar issues detract from the paper's professionalism.
- 5.MEDIUMcopyeditRephrase the responder analysis sentence in Methods, Statistical analysis for clarity: 'We decided to perform a responder analysis, comparing the proportion of participants within each study arm who had a change at least as large as the minimal clinically important difference, as part of the secondary analyses.'The original sentence is long and could be misread.
- 6.MEDIUMcopyeditAdd 'points' to the 12-month effect in the Abstract: '−2.9 points (−5.0 to −0.7)' and ensure consistent dash usage.Clarifies the unit and improves formatting consistency.
- 7.MEDIUMcopyeditCheck the footnote symbol placement in Table 3, row '12 item short form survey' (1.13 (−0.83 to 3.09)‡).Misplaced footnote symbols can mislead readers.
- 8.MEDIUMdata codeConsider providing the statistical analysis code in a public repository or as a supplementary file, even though data are controlled access.Sharing code enhances reproducibility and is a common reviewer request.
- 9.MEDIUMreportingReport the number of participants with missing data for each secondary outcome at each time point in the tables.Transparency about missing data strengthens the reporting of secondary analyses.
- 10.MEDIUMreportingDiscuss the potential impact of the high rate of unintentional unmasking (20%) on the primary outcome in the limitations section.Unmasking can bias outcome assessment; addressing it directly is important for interpretation.
- 11.LOWreportingProvide a link to the published protocol or statistical analysis plan in the data availability statement.Enhances transparency and allows independent verification.
- 12.LOWreportingConsider reporting the intraclass correlation coefficient for the primary outcome in the abstract.Provides additional context for the cluster-randomized design.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.