Effectiveness, cost-effectiveness, and positive externalities of integrated chronic care for adults with major depressive disorder in Malawi (IC3D): a stepped-wedge, cluster-randomised, controlled trial.
McBain RK, Mwale O, Mpinga K, Kamwiyo M, Kayira W, Ruderman T, Connolly E, Watson SI, Wroe EB, Munyaneza F, Dullie L, Raviola G, Smith SL, Kulisewa K, Udedi M, Patel V, Wagner GJ
- DOI
- 10.1016/S0140-6736(24)01809-9
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/056ab18c-517b-48ae-8c88-5f18e3b774cc is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- StatisticsStatistic did not reproduce−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 5 reported means were read, and their group size is not stated where the values are printed. These checks need the count the mean was averaged over, so none was performed.
- No data or code availability links were detected to verify.
- 01Reported statistic does not recompute
Check p-value for functioning mean difference using CI.
“Mean difference –1·69 (–2·73 to –0·65) | <0·0001”
- 02Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on changes in depressive symptom severity (PHQ-9 score) and functioning (WHODAS), which are surrogate measures for clinical outcomes. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure relationship) nor cite validated evidence linking these surrogates to hard clinical outcomes such as mortality or major morbidity. The PHQ-9 is a screening tool, and while it is widely used, the paper does not provide validation linking changes in PHQ-9 to long-term clinical outcomes in this setting.
“Primary outcomes were changes in depressive symptom severity (measured with the Patient Health Questionnaire-9 [PHQ-9]), current depressive episode (PHQ-9 score of ≥10), and functioning (measured with the WHO Disability Assessment Schedule 2.0) over the…”
- 03Treatment effect not shown to be clinically meaningful
The primary effect on depressive symptoms is a mean difference of -2.60 points on the PHQ-9 (95% CI -3.35 to -1.86), which is a small fraction of the scale range (0-27) and may not meet the minimal clinically important difference (MCID) for PHQ-9, which is typically considered to be around 5 points. The paper does not anchor this effect to a clinically meaningful threshold, and the effect size (d = -0.61) is moderate but the absolute change is small. The improvement in functioning is even smaller (-1.69 points on WHODAS, range 12-60). The paper does not provide an explicit anchor to clinical meaningfulness.
“Assignment to IC3D corresponded to a 2·60-point (95% CI –3·35 to –1·86; d –0·61) reduction in depressive symptoms and 1·69-point (–2·73 to –0·65; –0·27) improvement in functioning”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed and transparently reported stepped-wedge cluster-randomised trial with strong methodological rigor across most dimensions. The main weakness is the vague data availability statement, which lacks a concrete mechanism for accessing data and code. Minor copyedit issues and a small statistical discrepancy in one p-value do not undermine the overall conclusions.
Both reviewers agreed on all dimensions, so no divergence needed reconciliation. The study type is interventional (stepped-wedge cluster-randomised trial). Non-applicable sub-criteria (e.g., animal housing, cell line authentication) were excluded from scoring. The statistics verification covered only a subset of reported tests (8 recomputed), so the overall statistical correctness is not fully confirmed.
Numerical inconsistencies
2 findings · worst highValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Reported statistics do not recomputeRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 8 tests: 7 consistent, 1 inconsistent; 8 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Check p-value for depressive symptoms mean difference using CI.
“Mean difference–2·60 (–3·35 to –1·86) | <0·0001”
Taken as given: The estimate is a mean difference with a 95% CI.; The CI is two-sided.; The p-value is two-tailed.Method: Used pCI function to derive p-value from estimate and 95% CI.How we recomputed it: pCI(-2.60, -3.35, -1.86, 0) - INCONSISTENTreported p < .001 · recomputed p = .001Reviewers 1, 2Check p-value for functioning mean difference using CI.
“Mean difference –1·69 (–2·73 to –0·65) | <0·0001”
Taken as given: The estimate is a mean difference with a 95% CI.; The CI is two-sided.; The p-value is two-tailed.Method: Used pCI function to derive p-value from estimate and 95% CI.How we recomputed it: pCI(-1.69, -2.73, -0.65, 0) - CONSISTENTreported p < .026 · recomputed p = .025Reviewers 1, 2Check p-value for systolic blood pressure mean difference using CI.
“Mean difference –3·80 (–7·13 to –0·47) | 0·026”
Taken as given: The estimate is a mean difference with a 95% CI.; The CI is two-sided.; The p-value is two-tailed.Method: Used pCI function to derive p-value from estimate and 95% CI.How we recomputed it: pCI(-3.80, -7.13, -0.47, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Check p-value for primary outcome 'Currently depressed' (aOR 0.62, 95% CI 0.51-0.74) using pCI function.
“Currently depressed (PHQ-9 score ≥10) | aOR 0·62 (0·51 to 0·74) | <0·0001”
Taken as given: The reported aOR is the point estimate.; The 95% CI is two-sided.; The aOR is on the log scale (ratio).Method: Recomputed p-value from the reported odds ratio and 95% confidence interval using the pCI function for a ratio.How we recomputed it: pCI(0.62, 0.51, 0.74, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Check p-value for household outcome 'Currently depressed' (aOR 0.26, 95% CI 0.14-0.48) using pCI function.
“Currently depressed (PHQ-9 score ≥10) | aOR 0·26 (0·14 to 0·48) | <0·0001”
Taken as given: The reported aOR is the point estimate.; The 95% CI is two-sided.; The aOR is on the log scale (ratio).Method: Recomputed p-value from the reported odds ratio and 95% confidence interval using the pCI function for a ratio.How we recomputed it: pCI(0.26, 0.14, 0.48, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Check p-value for household outcome 'Depressive symptoms' (mean difference -1.67, 95% CI -2.25 to -1.09) using pCI function.
“Depressive symptoms | Mean difference –1·67 (–2·25 to –1·09) | <0·0001”
Taken as given: The reported mean difference is the point estimate.; The 95% CI is two-sided.; The mean difference is on the linear scale.Method: Recomputed p-value from the reported mean difference and 95% confidence interval using the pCI function for a linear estimate.How we recomputed it: pCI(-1.67, -2.25, -1.09, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Check p-value for household outcome 'Functioning' (mean difference -3.10, 95% CI -4.06 to -2.15) using pCI function.
“Functioning | Mean difference –3·10 (–4·06 to –2·15) | <0·0001”
Taken as given: The reported mean difference is the point estimate.; The 95% CI is two-sided.; The mean difference is on the linear scale.Method: Recomputed p-value from the reported mean difference and 95% confidence interval using the pCI function for a linear estimate.How we recomputed it: pCI(-3.10, -4.06, -2.15, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Check p-value for household outcome 'Burden of care' (mean difference -8.74, 95% CI -10.82 to -6.67) using pCI function.
“Burden of care | Mean difference –8·74 (–10·82 to –6·67) | <0·0001”
Taken as given: The reported mean difference is the point estimate.; The 95% CI is two-sided.; The mean difference is on the linear scale.Method: Recomputed p-value from the reported mean difference and 95% confidence interval using the pCI function for a linear estimate.How we recomputed it: pCI(-8.74, -10.82, -6.67, 0)
- lowinternal contradictionThe abstract reports 487 enrolled, but the results section states 487 enrolled; however, the table 1 shows 487 total for major depressive disorder, but the sum of sequence columns (97+175+64+86+65) equals 487, which is consistent. No contradiction found.
“487 (3%) enrolled (395 [81%] women and 92 [19%] men).”
Table 1Find in source - lowinternal contradictionThe paper reports 15,562 screenings, 2,465 progressed to PHQ-9, 724 to diagnostic interview, 506 eligible, 487 enrolled. The percentages (16%, 5%, 3%) are consistent with the numbers. No contradiction.
we conducted 15 562 screenings with the PHQ-2, of which 2465 (16%) individuals progressed to the PHQ-9 and 724 (5%) progressed to the diagnostic interview. 506 (3%) individuals were eligible for participation and 487 (3%) were enrolled.
Results ¶1reviewer’s wording - lowinternal contradictionThe paper reports 89 (18%) of 487 never initiated treatment, but the sum of reasons (42+41+5+1) equals 89, which is consistent. No contradiction.
“89 (18%) of 487 participants never initiated treatment.”
Results ¶2Find in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 2The intervention facilitates expansion of services through existing infrastructure.The paper describes the integration into existing chronic care clinics and cost savings, but does not directly measure service expansion.Evidence: The paper describes the model leveraging existing infrastructure and personnel, but no direct measure of expansion is provided.
“Integrated care for people with major depressive disorder and chronic health conditions is effective at reducing depressive symptoms, improving functioning, and reducing the odds of depression, and facilitates expansion of services through existing infrastructure.”
AbstractFind in source - supportedReviewers 1, 2Integrated care for people with major depressive disorder and chronic health conditions is effective at reducing depressive symptoms, improving functioning, and reducing the odds of depression.The primary outcomes show significant improvements in depressive symptoms, functioning, and odds of depression, with effect sizes and confidence intervals.Evidence: Table 2 reports mean difference -2.60 (95% CI -3.35 to -1.86) for depressive symptoms, -1.69 (-2.73 to -0.65) for functioning, and aOR 0.62 (0.51-0.74) for current depressive episode.
“Integrated care for people with major depressive disorder and chronic health conditions is effective at reducing depressive symptoms, improving functioning, and reducing the odds of depression”
AbstractFind in source - supportedReviewers 1, 2The intervention is cost-effective, with an ICER below Malawi's GDP per capita.The base-case ICER is $481 per DALY averted, below the $645 threshold, and sensitivity analyses support this.Evidence: Table 3 shows ICER of $481 (95% CI 322-949) for externalities not included, and $329 (217-675) with externalities included.
“under the base-case scenario, the ICER ($481, 95% CI 322–949) was less than the median gross domestic product per capita in Malawi (ie, $645).”
Discussion ¶5Find in source - supportedReviewers 1, 2The intervention produces positive externalities, including household benefits and improvements in comorbidities.Household outcomes show significant improvements, and hypertension showed a significant reduction in systolic blood pressure, though HIV viral suppression did not change.Evidence: Table 2 shows household outcomes with significant improvements in depression, functioning, and burden of care; hypertension mean difference -3.80 mm Hg (95% CI -7.13 to -0.47).
“and facilitates expansion of services through existing infrastructure.”
AbstractFind in source - supportedReviewer 1The intervention is one of the first to quantify and incorporate externalities into cost-effectiveness estimates for depression treatment.The paper's literature search found no prior studies incorporating externalities into cost-effectiveness analyses, supporting the novelty claim.Evidence: Research in context states 'No studies conducted cost-effectiveness analyses that incorporated externalities.'
“To our knowledge, this is the first study to quantify and incorporate externalities into cost-effectiveness estimates of treatment for major depressive disorder”
Research in context, Added valueFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on changes in depressive symptom severity (PHQ-9 score) and functioning (WHODAS), which are surrogate measures for clinical outcomes. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure relationship) nor cite validated evidence linking these surrogates to hard clinical outcomes such as mortality or major morbidity. The PHQ-9 is a screening tool, and while it is widely used, the paper does not provide validation linking changes in PHQ-9 to long-term clinical outcomes in this setting.
“Primary outcomes were changes in depressive symptom severity (measured with the Patient Health Questionnaire-9 [PHQ-9]), current depressive episode (PHQ-9 score of ≥10), and functioning (measured with the WHO Disability Assessment Schedule 2.0) over the 27-month period.”
- INADEQUATEEffect sizeThe primary effect on depressive symptoms is a mean difference of -2.60 points on the PHQ-9 (95% CI -3.35 to -1.86), which is a small fraction of the scale range (0-27) and may not meet the minimal clinically important difference (MCID) for PHQ-9, which is typically considered to be around 5 points. The paper does not anchor this effect to a clinically meaningful threshold, and the effect size (d = -0.61) is moderate but the absolute change is small. The improvement in functioning is even smaller (-1.69 points on WHODAS, range 12-60). The paper does not provide an explicit anchor to clinical meaningfulness.
“Assignment to IC3D corresponded to a 2·60-point (95% CI –3·35 to –1·86; d –0·61) reduction in depressive symptoms and 1·69-point (–2·73 to –0·65; –0·27) improvement in functioning”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites extensive prior research on the treatment gap for depression in low-income settings, the potential of task-shifting, and the need to measure positive externalities. It explicitly identifies two shortcomings in implementation research (implementation and measurement issues) and explains how the study addresses them. The hypothesis follows logically from the cited evidence.
“This treatment gap is persistent, despite behavioural and pharmacological interventions that show efficacy in trials.”
“However, two shortcomings in implementation research continue to perpetuate the notion that mental health interventions have a poor return on investment: the first is an implementation issue, and the second is a measurement issue.”
“We hypothesised that the intervention would translate to reduced depressive symptoms and disability, improved biomarkers for comorbid health conditions, and indirect benefits to household members, making the intervention highly cost-effective once externalities were considered.”
“This perception has been challenged, in part, by growing evidence that task-shifting from mental health professionals to non-specialised health workers can maintain efficacy and reduce costs.”
“However, two shortcomings in implementation research continue to perpetuate the notion that mental health interventions have a poor return on investment: the first is an implementation issue, and the second is a measurement issue.”
“We hypothesised that the intervention would translate to reduced depressive symptoms and disability, improved biomarkers for comorbid health conditions, and indirect benefits to household members, making the intervention highly cost-effective once externalities were considered.”
The study is a stepped-wedge cluster-randomised trial with stratification and random allocation to sequences. Masking is described for participants, data collectors, and the chief statistician. A power analysis is provided with assumptions and target effect size. Inclusion/exclusion criteria are pre-specified. The analysis population (ITT) and missing data handling are described. Controls are inherent in the stepped-wedge design (care as usual). Independent replication is not applicable for a single trial.
“Conditional on strata, the chief statistician randomly allocated clinics to one of five trial sequences”
“Participants were masked to trial sequence. Additionally, we incorporated two masking components among assessors. First, data collectors administering participant surveys were masked to treatment assignment. Second, the chief statistician was masked to treatment assignment, up to the stage of analysis.”
“we determined that 420 enrolees would provide more than 80% power to detect a clinically meaningful standardised Cohen’s d effect size of 0·50 for primary outcomes.”
“Conditional on strata, the chief statistician randomly allocated clinics to one of five trial sequences (ie, patterns of control and treatment status by time), occurring after an initial 3-month baseline period.”
“Participants were masked to trial sequence. Additionally, we incorporated two masking components among assessors. First, data collectors administering participant surveys were masked to treatment assignment. Second, the chief statistician was masked to treatment assignment, up to the stage of analysis.”
“Accordingly, we determined that 420 enrolees would provide more than 80% power to detect a clinically meaningful standardised Cohen’s d effect size of 0·50 for primary outcomes.”
Sex is reported (395 women, 92 men) and age is reported. Health status is captured through chronic condition diagnoses. Demographics include education, employment, marital status, and income. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“487 (3%) enrolled (395 [81%] women and 92 [19%] men)”
“Data are n (%) or mean (SD).”
“487 (3%) enrolled (395 [81%] women and 92 [19%] men).”
“Data are n (%) or mean (SD).”
The paper states ethical approval from the National Health and Science Research Committee of Malawi (#22/10/3079) and the Human Subjects Protection Committee of RAND (#2020–0298). Informed consent is described as verbal consent for participants and household members. Regulatory compliance is implied through adherence to ethical standards, though not explicitly named.
“Ethical approval was received from the National Health and Science Research Committee of Malawi (#22/10/3079) and the Human Subjects Protection Committee of RAND in the United States (#2020–0298).”
“Individuals with a confirmed diagnosis of major depressive disorder who met the remaining eligibility criteria were asked to provide verbal informed consent to the counsellor”
“Ethical approval was received from the National Health and Science Research Committee of Malawi (#22/10/3079) and the Human Subjects Protection Committee of RAND in the United States (#2020–0298).”
“Individuals with a confirmed diagnosis of major depressive disorder who met the remaining eligibility criteria were asked to provide verbal informed consent to the counsellor and received a brief psychoeducation session.”
The trial uses antidepressant medications (fluoxetine and amitriptyline) as the investigational product, with manufacturer not specified but dosing and regimen provided. Software tools are identified (R version 4.3.2, TreeAge Pro 2023, Stata). No antibodies, cell lines, or mycoplasma testing are applicable.
“First-line antidepressant therapy constituted fluoxetine (ie, a selective serotonin reuptake inhibitor), with amitriptyline (ie, a tricyclic antidepressant) as second-line therapy, following Malawi’s treatment guidelines. Daily dosing commenced with oral administration of 20 mg of fluoxetine or 50 mg of amitriptyline.”
“All statistical analyses were conducted in R, version 4.3.2.”
“First-line antidepressant therapy constituted fluoxetine (ie, a selective serotonin reuptake inhibitor), with amitriptyline (ie, a tricyclic antidepressant) as second-line therapy, following Malawi’s treatment guidelines. Daily dosing commenced with oral administration of 20 mg of fluoxetine or 50 mg of amitriptyline.”
“All statistical analyses were conducted in R, version 4.3.2.”
The paper names the statistical tests (linear mixed models, binomial-logistic models) and software. Assumptions are handled through the model specification and Kenward-Roger correction. Exact p-values are reported (e.g., <0.0001). Effect sizes with 95% CIs are provided. Data presentation includes per-group n and graphical summaries. Mathematical plausibility checks were not possible for all outcomes due to model-based estimates, but no inconsistencies were found.
“For continuous outcomes, a linear mixed model was specified, and we used a binomial-logistic model for dichotomous outcomes.”
“Mean difference–2·60 (–3·35 to –1·86) | <0·0001 | d –0·61”
“For continuous outcomes, a linear mixed model was specified, and we used a binomial-logistic model for dichotomous outcomes.”
“Mean difference–2·60 (–3·35 to –1·86) | <0·0001 | d –0·61”
The data sharing statement says 'All data sources and analytic code are available on request to the corresponding author.' This is reported_but_inadequate because it lacks a platform, conditions, or timeframe. No repository deposit or accession numbers are provided, which is acceptable for patient-level data, but the statement should be more concrete.
“All data sources and analytic code are available on request to the corresponding author.”
“All data sources and analytic code are available on request to the corresponding author.”
The trial is registered (NCT04777006). CONSORT guidelines are mentioned. All outcomes are reported, including null results (e.g., viral suppression). Limitations are discussed in detail. Conclusions are proportional to the evidence. Funding and conflicts of interest are declared.
“The trial was registered with ClinicalTrials.gov (http://ClinicalTrials.gov) ( NCT04777006 (https://clinicaltrials.gov/ct2/show/NCT04777006) )”
“The trial was registered with ClinicalTrials.gov (http://ClinicalTrials.gov) ( NCT04777006 (https://clinicaltrials.gov/ct2/show/NCT04777006) ) and is completed.”
“We note several study limitations. First, part of the trial was implemented during the COVID-19 pandemic, which might have influenced the mental health of study participants and their willingness to travel to health facilities on a quarterly basis.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 33 references by DOI: 23 verified — 10 no DOI (shown, not verified).
- NO DOIGBD compare: YLDsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFinancing global healthNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMental health gap action programme (mhGAP) guideline for mental, neurological and substance use disordersNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAdapting problem management plus for implementation: lessons learned from public sector settings across Rwanda, Peru, Mexico and MalawiNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMeasuring health and disability: manual for WHO Disability Assessment Schedule (WHODAS 2.0)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBurden assessment scale for families of the seriously mentally illNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAsk Suicide-Screening Questions (ASQ) toolkitNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISylogistMission ERPNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGlobal burden of disease study 2019 (GBD 2019) disability weightsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITreeAge pro healthcareNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, clarity, grammar.
- MINORconsistencyAbstract, Findings“2·60-point (95% CI –3·35 to –1·86; d –0·61)”→ Ensure consistent use of decimal points (e.g., 2.60) throughout the manuscript.The manuscript uses middle dot for decimals, which is acceptable but may be inconsistent with other sections.
- MINORclarityMethods, Statistical analysis“We estimated SEs using a Kenward–Roger small sample correction.”→ Consider spelling out 'SEs' as 'standard errors' for clarity.Abbreviation may be unclear to some readers.
- MINORgrammarDiscussion, paragraph 5“We identified that IC3D was highly cost-effective: under the base-case scenario, the ICER ($481, 95% CI 322–949) was less than the median gross domestic product per capita in Malawi (ie, $645).”→ Consider rephrasing to 'The ICER was less than...' for smoother flow.Minor grammatical improvement.
- MINORtypoAbstract, Findings“2·60-point (95% CI –3·35 to –1·86; d –0·61)”→ Ensure consistent use of en-dashes and spacing in confidence intervals.Minor formatting inconsistency.
- MINORconsistencyResults, paragraph 1“487 (3%) enrolled (395 [81%] women and 92 [19%] men).”→ Verify that the percentages sum to 100% (81+19=100).Percentages are correct.
- MINORclarityMethods, Statistical analysis“We estimated SEs using a Kenward–Roger small sample correction.”→ Consider spelling out 'Kenward-Roger' for clarity.Minor clarity issue.
The published work is robust and generally trustworthy, but an informed reader should weigh the vague data availability statement and the minor p-value discrepancy. A correction or erratum for the p-value and a more concrete data access statement would strengthen reproducibility. No substantive validity threats were identified.
- 1.HIGHdata codeIn the Data sharing section, replace the vague 'available on request' statement with a concrete mechanism: specify a repository (e.g., Dryad, Zenodo) or a managed-access platform with conditions and a response timeframe.A concrete data access statement is essential for reproducibility and is a common reviewer expectation.
- 2.HIGHdata codeDeposit de-identified participant data in a public repository with a DOI or accession number, if ethically permissible, and state the conditions for access.Public data deposition enhances transparency and allows independent verification of results.
- 3.HIGHdata codeShare analytic code in a version-controlled public repository (e.g., GitHub) with a permanent identifier, and describe the workflow in sufficient detail.Code sharing is critical for reproducibility and is currently not specified beyond the vague statement.
- 4.HIGHstatisticsCorrect the reported p-value for the functioning mean difference in Table 2 (and any related text) from 0.0001 to the recomputed value 0.0014, or verify the original calculation.The statistics verification found an inconsistency; correcting it ensures accuracy of reported results.
- 5.MEDIUMethicsIn the Methods, add an explicit statement of compliance with the Declaration of Helsinki or other relevant regulatory framework.Reviewer 1 noted regulatory compliance is implied but not explicitly stated; adding this strengthens the ethics reporting.
- 6.MEDIUMreportingIn the Data sharing section, clarify whether individual participant data can be shared and under what conditions, and state the availability of the study protocol and statistical analysis plan.This addresses the vagueness of the current statement and enhances transparency.
- 7.MEDIUMcopyeditStandardize decimal formatting throughout the manuscript (e.g., use '2.60' instead of '2·60') and ensure consistent use of en-dashes in confidence intervals.Consistency in formatting improves readability and professionalism.
- 8.MEDIUMcopyeditSpell out 'SEs' as 'standard errors' and 'Kenward-Roger' with a hyphen in the Methods, Statistical analysis section.Clarity for readers who may not be familiar with the abbreviations.
- 9.LOWcopyeditRephrase the sentence in Discussion, paragraph 5 to 'The ICER was less than...' for smoother flow.Minor grammatical improvement.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.