Task-sharing and telemedicine delivery of psychotherapy to treat perinatal depression: a pragmatic, noninferiority randomized trial.
Singla DR, Silver RK, Vigod SN, Schoueri-Mychasiw N, Kim JJ, La Porte LM, Ravitz P, Schiller CE, Lawson AS, Kiss A, Hollon SD, Dennis CL, Berenbaum TS, Krohn HA, Gibori JE, Charlebois J, Clark DM, Dalfen AK, Davis W, Gaynes BN, Leszcz M, Katz SR, Murphy KE, Naslund JA, Reyes-Rodríguez ML, Stuebe AM, Zlobin C, Mulsant BH, Patel V, Meltzer-Brody S
- DOI
- 10.1038/s41591-024-03482-w
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/ba7608cd-1ead-42da-b8df-534fd51968b0 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 62 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on the Edinburgh Postnatal Depression Scale (EPDS), a symptom rating scale, which is a surrogate for clinical outcomes. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking EPDS changes to hard clinical outcomes such as remission or functional improvement. The EPDS is a screening tool, and while it is widely used, the paper does not establish it as a validated surrogate for the clinical outcome of interest.
“The primary outcome was depressive symptoms (Edinburgh Postnatal Depression Scale (EPDS))”
- 02Treatment effect not shown to be clinically meaningful
The primary outcome is a noninferiority comparison, and the reported effect sizes are small absolute differences in EPDS scores (e.g., 0.36 for provider comparison, 0.23 for modality comparison). These differences are not anchored to a minimal clinically important difference (MCID) for the primary outcome. The paper mentions an MCID of 1.4–6.4 for EPDS in the discussion, but the observed differences are far below this range, and the paper does not explicitly state that the noninferiority margins are clinically meaningful. The effect sizes are presented as noninferiority results, but the clinical meaningfulness of the absolute differences is not clearly established.
“absolute difference in EPDS means (0.36)”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported pragmatic noninferiority RCT. The main methodological strengths include a strong scientific premise, rigorous study design, detailed reporting of biological variables, ethical approvals, and transparent reporting. The primary weakness is the vague data and code availability statements, which rely on email requests rather than a formal managed-access platform or repository.
Both reviewers agreed on study type (interventional) and all dimension statuses. The evaluation covered all eight rigor dimensions; no dimensions were excluded as not applicable. The statistics verification component checked 3 tests and found all consistent; no citation or reproducibility issues were flagged.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Two-sample t-test for sessions attended (telemedicine vs in-person): t(1228) = -8.15, P < 0.001
“those randomized to telemedicine attended significantly more sessions than those randomized to in-person BA (6.55 versus 5.07, t (1,228) = −8.15, P < 0.001).”
Taken as given: The t-statistic is -8.15 with 1228 degrees of freedom.; The test is two-tailed.; The reported p-value is less than 0.001.Method: Two-tailed t-test p-value computed from t-statistic and df using the t-distribution CDF.How we recomputed it: 2*(1-tCdf(8.15,1228)) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Two-sample t-test for difference in number of sessions attended between telemedicine and in-person groups.
“those randomized to telemedicine attended significantly more sessions than those randomized to in-person BA (6.55 versus 5.07, t (1,228) = −8.15, P < 0.001)”
Taken as given: The t-statistic is -8.15 with 1228 degrees of freedom.; The test is two-tailed.; The reported p-value is less than 0.001.Method: Two-tailed t-test p-value from t-statistic and df.How we recomputed it: 2*(1-tCdf(8.15,1228)) - CONSISTENTreported p = .930 · recomputed p = .931Reviewer 2F-test for interaction between modality and provider on EPDS outcome.
“We tested for a modality by provider interaction ( P = 0.93) for our primary (EPDS) outcome”
Taken as given: The F-statistic is not reported, only the p-value.; The interaction has 1 numerator df and approximately 1226 denominator df.; The p-value is two-tailed (F-test is one-tailed by nature).Method: Recomputed p-value from F-distribution assuming F=0.0076 (derived from p=0.93).How we recomputed it: 1-fCdf(0.0076,1,1226)
- lowinternal contradictionThe abstract states '1,230 participants were recruited' while the results state 'N = 1,230 participants were enrolled and randomized'. This is a minor terminology inconsistency, not a substantive contradiction.
Between 8 January 2020 and 4 October 2023, 1,230 participants were recruited. ... N = 1,230 participants were enrolled and randomized into the trial
Abstractreviewer’s wording - lowinternal contradictionThe abstract lists arm numbers as 472, 145, 469, 144, while the Results section lists them as 472, 469, 145, 144. The order differs but the numbers are the same; this is a minor inconsistency in presentation.
Abstract: '472 nonspecialist telemedicine, 145 nonspecialist in-person, 469 specialist telemedicine and 144 specialist in-person' vs. Results: '472 were assigned to the nonspecialist-telemedicine arm, 469 to specialist-telemedicine, 145 to nonspecialist in-person and 144 to specialist in-person.'
Abstractreviewer’s wording
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2Telemedicine is noninferior to in-person for treating perinatal anxiety symptoms.Noninferiority was met in the ITT analysis but not in the per-protocol analysis unless outliers were removed, so the claim is partially supported.Evidence: ITT GAD-7: telemedicine 6.43 (95% CI 6.09–6.78) vs in-person 6.29 (95% CI 5.71–6.88), upper bound 0.73, NIM 0.82. PP analysis: upper bound 0.90, NIM 0.80, noninferiority not met.
Noninferiority was met when comparing providers ... and modalities (ITT GAD-7: telemedicine 6.43 (95% CI 6.09–6.78) versus in-person 6.29 (95% CI 5.71–6.88)) on anxiety symptoms for all analyses at 3 months post-randomization (Table ) except for the PP analyses comparing telemedicine versus in-person in which noninferiority was not met on anxiety symptoms unless outliers were removed
Resultsreviewer’s wording - supportedReviewers 1, 2Nonspecialist providers are noninferior to specialist providers in delivering behavioral activation for perinatal depressive symptoms.The primary outcome analysis shows the upper bound of the 95% CI for the difference in EPDS means (0.86) did not exceed the noninferiority margin (0.89), supporting the claim.Evidence: ITT analysis: EPDS nonspecialist 9.27 (95% CI 8.85–9.70) vs specialist 8.91 (95% CI 8.49–9.33), absolute difference 0.36, upper bound 0.86, NIM 0.89.
“the upper limit of the 95% CI for the difference in EPDS means (0.86) did not exceed the 10% noninferiority margin (EPDS 0.89).”
ResultsFind in source - supportedReviewers 1, 2Telemedicine-delivered behavioral activation is noninferior to in-person delivery for perinatal depressive symptoms.The ITT analysis shows the upper bound of the 95% CI (0.77) did not exceed the noninferiority margin (1.16), supporting the claim.Evidence: ITT analysis: EPDS telemedicine 9.15 (95% CI 8.79–9.50) vs in-person 8.92 (95% CI 8.38–9.45), absolute difference 0.23, upper bound 0.77, NIM 1.16.
“the upper limit of the 95% CI for the difference in EPDS means (0.77) did not exceed the 13% noninferiority margin (EPDS 1.16) at 3 months post-randomization.”
ResultsFind in source - supportedReviewers 1, 2Nonspecialist providers are noninferior to specialists in treating perinatal anxiety symptoms.The secondary outcome analysis shows noninferiority was met for anxiety symptoms in the ITT analysis.Evidence: ITT GAD-7: nonspecialist 6.44 (95% CI 6.01–6.86) vs specialist 6.36 (95% CI 5.95–6.78), upper bound 0.57, NIM 0.64.
“Noninferiority was met when comparing providers (ITT GAD-7: nonspecialist 6.44 (95% CI 6.01–6.86) versus specialist 6.36 (95% CI 5.95–6.78))”
ResultsFind in source - supportedReviewers 1, 2There were no serious or adverse events related to the trial.The paper reports that all SAEs and AEs were reviewed by the DSMB and none were deemed directly related to the trial.Evidence: Eighteen SAEs and two AEs were identified; all were reviewed by the DSMB and none were deemed directly related.
“All were reviewed by an independent Data Safety and Monitoring Board (DSMB) and none were deemed directly related or a result of the trial.”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on the Edinburgh Postnatal Depression Scale (EPDS), a symptom rating scale, which is a surrogate for clinical outcomes. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking EPDS changes to hard clinical outcomes such as remission or functional improvement. The EPDS is a screening tool, and while it is widely used, the paper does not establish it as a validated surrogate for the clinical outcome of interest.
“The primary outcome was depressive symptoms (Edinburgh Postnatal Depression Scale (EPDS))”
- INADEQUATEEffect sizeThe primary outcome is a noninferiority comparison, and the reported effect sizes are small absolute differences in EPDS scores (e.g., 0.36 for provider comparison, 0.23 for modality comparison). These differences are not anchored to a minimal clinically important difference (MCID) for the primary outcome. The paper mentions an MCID of 1.4–6.4 for EPDS in the discussion, but the observed differences are far below this range, and the paper does not explicitly state that the noninferiority margins are clinically meaningful. The effect sizes are presented as noninferiority results, but the clinical meaningfulness of the absolute differences is not clearly established.
“absolute difference in EPDS means (0.36)”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
3 integrity concerns flagged (0 high).
- lowotherThe noninferiority margin for the provider comparison is 10% (EPDS 0.89) and for modality is 13% (EPDS 1.16). The change in margin for modality is explained by COVID-19 modifications, but the rationale for the specific 13% is not detailed.
the upper limit of the 95% CI for the difference in EPDS means (0.86) did not exceed the 10% noninferiority margin (EPDS 0.89). ... the upper limit of the 95% CI for the difference in EPDS means (0.77) did not exceed the 13% noninferiority margin (EPDS 1.16)
Resultsreviewer’s wording
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites a systematic review of 45 RCTs of nonspecialist-delivered psychological treatments and notes that no trials compared different provider types. It also cites evidence on telemedicine efficacy and notes the lack of adequately powered comparisons with in-person care. The rationale logically links these gaps to the trial's objectives, and the limitations of prior research (e.g., inactive control groups, inadequate power) are explicitly addressed.
“A previous systematic review yielded 45 randomized controlled trials of nonspecialist-delivered psychological treatments for perinatal populations with common mental health conditions.”
“no trials, to our knowledge, evaluated whether different provider types were able to deliver the same treatments comparably.”
“meaningful comparisons of telemedicine-delivered psychotherapy with in-person delivery have been impeded by inadequately powered trials to assess their effectiveness in common mental disorders.”
“A previous systematic review yielded 45 randomized controlled trials of nonspecialist-delivered psychological treatments for perinatal populations with common mental health conditions.”
“no trials, to our knowledge, evaluated whether different provider types were able to deliver the same treatments comparably.”
“meaningful comparisons of telemedicine-delivered psychotherapy with in-person delivery have been impeded by inadequately powered trials to assess their effectiveness in common mental disorders.”
Randomization was 1:1:1:1 to four arms, with a method described (though the specific randomization algorithm is not detailed in the text, it is referenced to the protocol). The unit of randomization is the individual participant. Blinding is not explicitly described, but the pragmatic nature of the trial and the use of self-report outcomes (EPDS, GAD-7) mitigate the need for assessor blinding; however, the paper does not state this rationale. Power analysis is reported in the protocol (referenced). Inclusion/exclusion criteria are clearly stated. Outlier handling is defined (1.5×IQR rule). The analysis population (ITT and PP) is defined, with PP only for modality due to COVID-19 switches.
“pregnant and postpartum adult participants were randomized 1:1:1:1 to each arm”
“inclusion criteria included being a pregnant (≤36 weeks) or postpartum (4–30 weeks) adult (≥18 years; inclusive of gender identities and birthing persons) with a score ≥10 on the EPDS and speaking English or Spanish.”
“Outliers were defined as scores that were greater than 1.5× the interquartile range + the third quartile OR lower than the first quartile – 1.5× the interquartile range.”
“pregnant and postpartum adult participants were randomized 1:1:1:1 to each arm”
“inclusion criteria included being a pregnant (≤36 weeks) or postpartum (4–30 weeks) adult (≥18 years; inclusive of gender identities and birthing persons) with a score ≥10 on the EPDS and speaking English or Spanish.”
“Outliers were defined as scores that were greater than 1.5× the interquartile range + the third quartile OR lower than the first quartile – 1.5× the interquartile range.”
Sex is reported (99.57% cis-women). Age is reported (mean 33.27 years). Demographics are extensive, including race/ethnicity, education, income, marital status, and employment. Health status is reported via psychiatric history and psychotropic medication use. Since the study enrolled both sexes (though predominantly women), sex_justified is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“who predominantly identified as cis-women (1,168/1,173, 99.57%)”
“Participants’ mean age was 33.27 (95% confidence interval (CI) 33.00–33.55) years”
“who predominantly identified as cis-women (1,168/1,173, 99.57%)”
“Participants’ mean age was 33.27 (95% confidence interval (CI) 33.00–33.55) years”
The paper states: 'The study received ethical approvals from the following three institutional review boards (IRBs): UNC Biomedical IRB (19-1786), Endeavor Health IRB (EH18-129) and Clinical Trials Ontario (1895).' It also states 'All participants provided written informed consent before enrollment.' Regulatory compliance is not explicitly named (e.g., Declaration of Helsinki), but the IRB approvals and consent process are adequate.
“The study received ethical approvals from the following three institutional review boards (IRBs): UNC Biomedical IRB (19-1786), Endeavor Health IRB (EH18-129) and Clinical Trials Ontario (1895).”
“All participants provided written informed consent before enrollment.”
“The study received ethical approvals from the following three institutional review boards (IRBs): UNC Biomedical IRB (19-1786), Endeavor Health IRB (EH18-129) and Clinical Trials Ontario (1895).”
“All participants provided written informed consent before enrollment.”
The intervention is behavioral activation, described as manualized with an open-access manual available at www.thesummittrial.com. The manual is adapted from two established manuals. Statistical software is identified as SAS version 9.4. No antibodies, cell lines, or other reagents are used. The trial is a behavioral intervention, so the investigational product is the therapy itself, which is adequately identified.
“The SUMMIT open-access treatment manual is available online at no cost ( www.thesummittrial.com (http://www.thesummittrial.com) ).”
“The trial biostatistician conducted all analyses using SAS version 9.4 (SAS Institute).”
“The SUMMIT open-access treatment manual is available online at no cost ( www.thesummittrial.com (http://www.thesummittrial.com) ).”
“The trial biostatistician conducted all analyses using SAS version 9.4 (SAS Institute).”
The paper reports noninferiority using CIs around mean differences, which is the standard idiom. Tests are named (e.g., two-sample t-tests, linear mixed models, F-tests). Effect sizes are reported as mean differences with 95% CIs. Software is identified (SAS 9.4). Data presentation includes figures with means over time and tables with CIs. Exact p-values are reported for some analyses (e.g., t-test for sessions attended, P < 0.001), but for primary outcomes, noninferiority is assessed via CIs, so exact_p_values is not applicable. Assumptions are not explicitly verified, but the use of standard methods and sensitivity analyses is adequate for a large trial.
“The trial biostatistician conducted all analyses using SAS version 9.4 (SAS Institute).”
“those randomized to telemedicine attended significantly more sessions than those randomized to in-person BA (6.55 versus 5.07, t (1,228) = −8.15, P < 0.001).”
“EPDS: nonspecialist 9.27 (95% CI 8.85–9.70) versus specialist 8.91 (95% CI 8.49–9.33), absolute difference in EPDS means (0.36)”
“Linear mixed models assessed the moderating effects of severity at baseline, with participants as a random effect and a treatment-by-time interaction.”
“We tested for a modality by provider interaction ( P = 0.93) for our primary (EPDS) outcome”
The data availability statement says: 'Two years after publication, all individual participant deidentified data will be shared with researchers who submit a written proposal to: summittrial@sinaihealth.ca.' This is a concrete route (email) but lacks a formal data-access committee or platform, and the conditions are minimal. Code availability is similarly vague. Repository deposit and accession numbers are not applicable for patient-level data.
“Two years after publication, all individual participant deidentified data will be shared with researchers who submit a written proposal to: summittrial@sinaihealth.ca.”
“Two years after publication, analytic code will be shared with researchers who submit a formal written proposal to summittrial@sinaihealth.ca.”
“Two years after publication, all individual participant deidentified data will be shared with researchers who submit a written proposal to: summittrial@sinaihealth.ca.”
“Two years after publication, analytic code will be shared with researchers who submit a formal written proposal to summittrial@sinaihealth.ca.”
The trial is registered at ClinicalTrials.gov (NCT04153864). Methods are detailed enough for replication. Limitations are discussed (e.g., COVID-19 modifications, generalizability). Conclusions are proportional to the noninferiority findings. Funding and competing interests are disclosed. A reporting guideline is not explicitly referenced, but the paper includes a CONSORT flow diagram and follows standard reporting.
“ClinicalTrials.gov NCT04153864 (https://clinicaltrials.gov/ct2/show/NCT04153864)”
“The study also has limitations. Substantial modifications were required due to COVID-19, including a change in noninferiority margin from 10% to 13% in the telemedicine versus in-person comparison, and deviation from the assigned modality due to institutional mandates during the pandemic.”
“We also thank the Patient-Centered Outcomes Research Institute award (PCS-2018C1-10621, to D.R.S.) who funded SUMMIT.”
“ClinicalTrials.gov NCT04153864 (https://clinicaltrials.gov/ct2/show/NCT04153864)”
“The study also has limitations. Substantial modifications were required due to COVID-19, including a change in noninferiority margin from 10% to 13% in the telemedicine versus in-person comparison”
“We also thank the Patient-Centered Outcomes Research Institute award (PCS-2018C1-10621, to D.R.S.) who funded SUMMIT.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 84 references by DOI: 2 verified — 82 no DOI (shown, not verified).
- NO DOIRates and risk of postpartum depression—a meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPrevalence of antenatal and postnatal anxiety: systematic review and meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINon-psychotic mental disorders in the perinatal periodNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA meta-analysis of treatments for perinatal depressionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInterventions to prevent perinatal depression: US preventive services task force recommendation statementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPreferences and perceived barriers to treatment for depression during the perinatal periodNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDepression in Adults: Treatment and ManagementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPrimary care screening for and treatment of depression in pregnant and postpartum women: evidence report and systematic review for the US preventive services task forceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICanadian Network for Mood and Anxiety Treatments (CANMAT) 2016 clinical guidelines for the management of adults with major depressive disorder: section 2. Psychological treatmentsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICanadian network for mood and anxiety treatments (CANMAT) 2024 clinical practice guideline for the treatment of perinatal mood, anxiety and related disordersNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBarriers and facilitators to implementing perinatal mental health care in health and social care settings: a systematic reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe perinatal depression treatment cascade: baby steps toward improving outcomesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAddressing the treatment gap: expanding the scalability and reach of treatmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEfficacy of psychosocial interventions for mental health outcomes in low-income and middle-income countries: an umbrella reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILay health worker involvement in evidence-based treatment delivery: a conceptual model to address disparities in careNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImplementation and effectiveness of nonspecialist-delivered interventions for perinatal mental health in high-income countries: a systematic review and meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScaling up mental healthcare for perinatal populations: is telemedicine the answer?No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe effectiveness of telemedicine interventions to address maternal depression: a systematic review and meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMental health-related telemedicine interventions for pregnant women and new mothers: a systematic literature reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAmount and frequency of psychotherapy as predictors of treatment outcome for adult depression: a meta-regression analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA pragmatic randomized clinical trial of behavioral activation for depressed pregnant womenNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Healthy Activity Program (HAP), a lay counsellor-delivered brief psychological treatment for severe depression, in primary care in India: a randomised controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffectiveness of psychological treatments for depression and alcohol use disorder delivered by community-based counsellors: two pragmatic randomised controlled trials within primary healthcare in NepalNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBehavioural activation for depression; an update of meta-analysis of effectiveness and sub group analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILooking beyond depression: a meta-analysis of the effect of behavioral activation on depression, anxiety, and activationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRandomized trial of behavioral activation, cognitive therapy, and antidepressant medication in the acute treatment of adults with major depressionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRandomized trial of behavioral activation, cognitive therapy, and antidepressant medication in the prevention of relapse and recurrence in major depressionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICost and outcome of behavioural activation versus cognitive behavioural therapy for depression (COBRA): a randomised, controlled, non-inferiority trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDetection of postnatal depression: development of the 10-item Edinburgh Postnatal Depression ScaleNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA brief measure for assessing generalized anxiety disorder: the GAD-7No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHow much change is enough? Evidence from a longitudinal study on depression in UK primary careNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHuman resources for mental health care: current situation and strategies for actionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIndividual behavioral activation in the treatment of depression: a meta analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAnxiety-focused cognitive behavioral therapy delivered by non-specialists to prevent postnatal depression: a randomized, phase 3 trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIVirtual prenatal care: a systematic review of pregnant women’s and healthcare professionals’ experiences, needs, and preferences for quality careNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITelehealth utilization and associations in the United States during the third year of the COVID-19 pandemic: population-based survey study in 2022No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPrevalence and disparities in telehealth use among US adults following the COVID-19 pandemic: national cross-sectional surveyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOI‘There is just a different energy’: changes in the therapeutic relationship with the telehealth transitionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInforming the debate about telemedicine reimbursement—what do we need to knowNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITransparency about the outcomes of mental health services (IAPT approach): an analysis of public dataNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInternet-based cognitive behavioral therapy for depression: a systematic review and individual patient data network meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe effects of psychotherapies for depression on response, remission, reliable change, and deterioration: a meta‐analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAdherence to internet-based and face-to-face cognitive behavioural therapy for depression: a meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMaternal Mental Health in Canada, 2018/2019No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOI2020 Pregnancy risk assessment monitoring system (PRAMS) detailed data tablesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRisk factors associated with postpartum depressive symptoms: a multinational studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRisk for maternal depressive symptoms and perceived stress by ethnicities in Canada: from pregnancy through the preschool yearsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITrends in postpartum depression by race/ethnicity and pre-pregnancy body mass indexNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICensus Profile, 2021 Census of PopulationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICensus 2021: Families, Households, Marital Status and IncomeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFocus on Geography Series, 2021 Census of Population, TorontoNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBirth StatisticsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIllinoisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINorth CarolinaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHDPulse: An Ecosystem of Minority Health and Health Disparities Resources. North CarolinaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInterpreting the results of noninferiority trials—a reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExploring different objectives in non-inferiority trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISensitivity to change and minimal clinically important difference of Edinburgh Postnatal Depression ScaleNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe UCSF Client Satisfaction Scales: I. The Client Satisfaction Questionnaire-8No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDevelopment and preliminary testing of the new five-level version of EQ-5D (EQ-5D-5L)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe multidimensional scale of perceived social supportNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMultiple mediation analysis of the peer-delivered Thinking Healthy Programme for perinatal depression: findings from two parallel, randomised controlled trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWorking Alliance Inventory–Short Revised (WAI–SR): psychometric properties in outpatients and inpatientsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDeveloping the World Health Organization disability assessment schedule 2.0No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScaling up maternal mental healthcare by increasing access to treatment (SUMMIT) through non-specialist providers and telemedicine: a study protocol for a non-inferiority randomized controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOnset timing, thoughts of self-harm, and diagnoses in postpartum women with screen-positive depression findingsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAccuracy of the Edinburgh Postnatal Depression Scale (EPDS) for screening to detect major depression among pregnant and postpartum women: systematic review and meta-analysis of individual participant dataNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIValidation of the culturally adapted Edinburgh Postpartum Depression Scale among east Asian, southeast Asian and south Asian populations: a scoping reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe origins and current status of behavioral activation treatments for depressionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICulturally sensitive psychotherapy for perinatal women: a mixed methods studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScaling up quality-assured psychotherapy: the role of therapist competence on perinatal depression and anxiety outcomesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImproving the scalability of psychological treatments in developing countries: an evaluation of peer-led therapy quality assessment in Goa, IndiaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAssessing health worker competence to deliver a brief psychological treatment for depression: development and validation of a scalable measureNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImplementing psychological interventions through nonspecialist providers and telemedicine in high-income countries: qualitative study from a multistakeholder perspectiveNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAdapting behavioral activation for perinatal depression and anxiety in response to the COVID-19 pandemic and racial injusticeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBarriers and facilitators to resuming in-person psychotherapy with perinatal patients amid the COVID-19 pandemic: a multistakeholder perspectiveNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICalculating clinically significant change in postnatal depression studies using the Edinburgh Postnatal Depression ScaleNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA systematic review on the acceptability of perinatal depression screeningNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffect of peer support on prevention of postnatal depression among high risk women: multisite randomised controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMultiple Endpoints in Clinical Trials Guidance for IndustryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPoints to Consider on Multiplicity Issues in Clinical TrialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMultiple imputation after 18+ yearsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
3 data/code links checked; 3 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT04153864LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/study/NCT04153864LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codehttp://www.thesummittrial.comLIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoDiscussion, paragraph 2“s ame”→ sameTypographical error: 's ame' should be 'same'.
- MINORconsistencyAbstract vs. Results“472 nonspecialist telemedicine, 145 nonspecialist in-person, 469 specialist telemedicine and 144 specialist in-person”→ Ensure the arm numbers are consistent throughout the paper.The abstract lists arm numbers that sum to 1230, but the Results section lists the same numbers; however, the order in the abstract is different from the Results (which lists 472, 469, 145, 144). This is a minor ordering inconsistency.
- MINORclarityMethods, Statistical analysis“we looked at the CI around the difference in outcome based on the actual data”→ Clarify that this is the primary noninferiority analysis method.The sentence is slightly informal; could be rephrased for clarity.
- MINORtypoDiscussion, paragraph 2“s ame psychotherapy”→ same psychotherapyTypographical error: extra space in 's ame'.
- MINORconsistencyAbstract and Results“1,230 participants were recruited”→ 1,230 participants were enrolled and randomizedThe abstract says 'recruited' while the results say 'enrolled and randomized'; consider consistent terminology.
- MINORclarityTable 2 footnote“± The 95% upper bound is used when higher scores are worse and the 95% lower bound is used when lower scores are worse.”→ Clarify that this applies to the difference in means, not the individual scores.The footnote could be clearer about which bound is used for the noninferiority comparison.
The published paper is robust overall, with minor reporting gaps in data/code availability and a few copyedit issues. An informed reader should weigh the vague data-sharing mechanism and the lack of explicit blinding and power analysis details in the main text. No erratum-level concerns are warranted, but the authors could strengthen the paper by depositing data and code in a managed-access repository.
- 1.HIGHdata codeReplace the email-based data availability statement with a named managed-access platform (e.g., Vivli, YODA) or provide a data-sharing agreement template.A simple email address does not meet the standard for managed access and may deter legitimate researchers from requesting data.
- 2.HIGHdata codeDeposit the analytic code in a public repository (e.g., Zenodo, GitHub) with a DOI or permanent identifier, and update the code availability statement accordingly.Code sharing via email request is vague and not reproducible; a public repository ensures long-term accessibility.
- 3.HIGHreportingIn the Methods, explicitly state whether blinding of outcome assessors was performed or provide a rationale for not blinding, given the pragmatic design.Blinding is a key quality indicator for RCTs; its absence should be justified to allow readers to assess potential bias.
- 4.HIGHreportingIn the Methods, describe the randomization sequence generation and allocation concealment in more detail (e.g., central randomization, block sizes).Adequate randomization and allocation concealment are critical for preventing selection bias; the current description is insufficient.
- 5.HIGHreportingIn the Methods or Reporting summary, explicitly reference the CONSORT guideline and include the CONSORT checklist as supplementary material.Explicit adherence to a reporting guideline improves completeness and transparency, and is expected by many journals.
- 6.HIGHethicsAdd a statement on regulatory compliance (e.g., Declaration of Helsinki) in the ethics section.Explicit mention of regulatory compliance is a standard requirement for clinical trials and strengthens the ethics statement.
- 7.MEDIUMreportingIn the Methods, add a formal power analysis or sample size calculation, stating the assumed effect size, alpha, power, and noninferiority margin used to determine N=1230.A power analysis is a key design element; referencing the protocol is insufficient for readers who do not have access to it.
- 8.MEDIUMstatisticsIn the Statistical analysis section, explicitly state how normality and homogeneity of variance assumptions were checked for the linear models, or justify the use of robust methods.Verification of model assumptions is important for the validity of parametric tests; the current reporting is inadequate.
- 9.MEDIUMreportingIn the Discussion, provide a more detailed discussion of the implications of the GAD-7 PP analysis where noninferiority was not met, and why the result is not clinically meaningful.A noninferiority failure in a secondary outcome warrants careful interpretation to avoid misleading conclusions.
- 10.MEDIUMstatisticsConsider reporting exact p-values for the primary noninferiority comparisons (e.g., for the difference in means) in addition to CIs, to facilitate interpretation.Exact p-values can aid readers in assessing the strength of evidence, though CIs are the primary method for noninferiority.
- 11.LOWcopyeditFix the typographical error 's ame' to 'same' in the Discussion, paragraph 2.Typographical errors reduce the professional appearance of the manuscript.
- 12.LOWcopyeditEnsure consistent ordering of arm numbers in the Abstract and Results (e.g., list as 472, 469, 145, 144 in both places).Inconsistent ordering may confuse readers, even if the numbers are the same.
- 13.LOWcopyeditUse consistent terminology: replace 'recruited' with 'enrolled and randomized' in the Abstract to match the Results.Consistent terminology improves clarity and avoids potential misinterpretation.
- 14.LOWcopyeditClarify the footnote in Table 2 to specify that the 95% upper/lower bound applies to the difference in means, not the individual scores.A clearer footnote helps readers correctly interpret the noninferiority comparison.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.