Anxiety-focused cognitive behavioral therapy delivered by non-specialists to prevent postnatal depression: a randomized, phase 3 trial.
Surkan PJ, Malik A, Perin J, Atif N, Rowther A, Zaidi A, Rahman A
- DOI
- 10.1038/s41591-024-02809-x
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/f4dd28e1-e605-4d1c-8370-c87bde49783a is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×4−2★
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported phase 3 RCT with a strong scientific premise, rigorous design, and transparent reporting. The main weaknesses are minor reporting gaps: no explicit CONSORT reference, no explicit statistical assumption checks, and no analysis code sharing. The copyedit pass found only minor typos and rounding inconsistencies.
Both reviewers agreed on all dimensions and study type (interventional). The statistics verification recomputed 5 tests, all consistent; coverage is limited to tests with test statistics/df or effect estimates with CIs. The citation check found no retracted or unresolved references. The integrity check flagged only low-severity rounding differences, which are not validity threats.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 5 tests: 5 consistent, 0 inconsistent; 5 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Recompute p-value for the primary composite outcome (MDE or moderate-to-severe anxiety) using the reported adjusted odds ratio and 95% CI.
“Adjusted Odds Ratio | 0.19 (0.132, 0.268) | p | 1.6 ×10^E-20”
Taken as given: The adjusted odds ratio is 0.19.; The 95% confidence interval is (0.132, 0.268).; The p-value is two-sided.Method: Used pCI function to derive p-value from the odds ratio and its 95% confidence interval on the log scale.How we recomputed it: pCI(0.19, 0.132, 0.268, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Recompute p-value for the primary outcome of MDE alone using the reported adjusted odds ratio and 95% CI.
“Adjusted Odds Ratio | 0.19 (0.129, 0.275) | p | 5.5 ×10^E-18”
Taken as given: The adjusted odds ratio is 0.19.; The 95% confidence interval is (0.129, 0.275).; The p-value is two-sided.Method: Used pCI function to derive p-value from the odds ratio and its 95% confidence interval on the log scale.How we recomputed it: pCI(0.19, 0.129, 0.275, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Recompute p-value for the primary outcome of moderate-to-severe anxiety using the reported adjusted odds ratio and 95% CI.
“Adjusted Odds Ratio | 0.26 (0.167, 0.393) | p | 4.7 ×10^E-10”
Taken as given: The adjusted odds ratio is 0.26.; The 95% confidence interval is (0.167, 0.393).; The p-value is two-sided.Method: Used pCI function to derive p-value from the odds ratio and its 95% confidence interval on the log scale.How we recomputed it: pCI(0.26, 0.167, 0.393, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Recompute p-value for the difference in HADS anxiety scores using the reported adjusted mean difference and 95% CI.
“Adjusted Difference (reference control): Estimate (95% CI) | −3.80 (−4.414, −3.181) | p | 9.3×10^E-31”
Taken as given: The adjusted mean difference is -3.80.; The 95% confidence interval is (-4.414, -3.181).; The p-value is two-sided.Method: Used pCI function to derive p-value from the mean difference and its 95% confidence interval.How we recomputed it: pCI(-3.80, -4.414, -3.181, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Recompute p-value for the difference in PHQ-9 scores using the reported adjusted mean difference and 95% CI.
“Adjusted Difference (reference control): Estimate (95% CI) | −5.09 (−5.880, −4.282) | p | 1.6 ×10^E-32”
Taken as given: The adjusted mean difference is -5.09.; The 95% confidence interval is (-5.880, -4.282).; The p-value is two-sided.Method: Used pCI function to derive p-value from the mean difference and its 95% confidence interval.How we recomputed it: pCI(-5.09, -5.880, -4.282, 0)
- lowinternal contradictionThe abstract states '12% of women in the intervention group developed MDE' while Table 4 reports 11.6% (44/380). This is a rounding difference, not a true contradiction.
“12% of women in the intervention group developed MDE at six-weeks postpartum, versus 41% in the control group.”
Table 4Find in source - lowinternal contradictionThe text states that 755 women completed postnatal assessments, but the flow diagram (not shown) may have different numbers; however, the text is consistent.
“Of 1,200 women who were randomized, 755 (63%) completed the six-week postnatal interview”
ResultsFind in source - lowinternal contradictionThe text reports 'average HADS anxiety score ... 3.4 versus 7.2' while Table 5 reports 3.40 and 7.19. This is a rounding difference.
“the average postnatal HADS anxiety score among women in the intervention arm was significantly lower than among those in the control arm (3.4 versus 7.2, p < 0.001)”
Table 5Find in source - lowinternal contradictionThe abstract states '41% in the control group' while Table 4 reports 40.5% (152/375). This is a rounding difference.
“12% of women in the intervention group developed MDE at six-weeks postpartum, versus 41% in the control group.”
Table 4Find in source
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
5 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2The HMHB intervention reduced the odds of postnatal MDE or moderate-to-severe anxiety by 81%.The primary outcome analysis shows aOR=0.19 (95% CI 0.14-0.28), which is consistent with an 81% reduction in odds.Evidence: Table 4 reports adjusted OR 0.19 (95% CI 0.132-0.268) for the composite outcome.
“we found 81% reduced odds of having either a Major Depressive Episode (MDE) or moderate-to-severe anxiety for women randomized to the intervention (aOR=0.19, 95% CI: 0.14-0.28)”
AbstractFind in source - supportedReviewers 1, 2The intervention reduced the odds of postnatal MDE by 81%.The analysis for MDE alone shows aOR=0.19 (95% CI 0.13-0.28), consistent with an 81% reduction.Evidence: Table 4 reports adjusted OR 0.19 (95% CI 0.129-0.275) for MDE.
“We found reductions of 81% and 74% in the odds of postnatal MDE (aOR=0.19, 95% CI: 0.13-0.28)”
AbstractFind in source - supportedReviewers 1, 2The intervention reduced the odds of moderate-to-severe anxiety by 74%.The analysis for moderate-to-severe anxiety shows aOR=0.26 (95% CI 0.17-0.40), consistent with a 74% reduction.Evidence: Table 4 reports adjusted OR 0.26 (95% CI 0.167-0.393) for anxiety.
“and of moderate-to-severe anxiety (aOR=0.26, 95% CI: 0.17-0.40), respectively”
AbstractFind in source - supportedReviewers 1, 2The intervention was safe, with no adverse events related to participation.The paper reports that adverse events were unrelated to the intervention and no unexpected events occurred.Evidence: Results, Safety section states 'The study did not have any unexpected events or adverse events related to the intervention'.
“The study did not have any unexpected events or adverse events related to the intervention, thus the adverse events that occurred were unrelated to the trial.”
ResultsFind in source - supportedReviewers 1, 2The intervention is effective in preventing postnatal depression and anxiety in a low-resource setting.The trial was conducted in Pakistan, a low-resource setting, and showed significant reductions in outcomes.Evidence: The trial was conducted in Pakistan and the primary outcomes were met.
“This study was carried out in a low-resource setting where specialized mental health care is not readily available.”
DiscussionFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary outcome is a clinical diagnosis of Major Depressive Episode (MDE) via SCID and moderate-to-severe anxiety via HADS threshold, which are clinical outcomes, not surrogate biomarkers. The intervention shows target engagement through reduced symptom scores and the outcomes are validated clinical measures.
“The primary outcome was major depression, generalized anxiety disorder, or both at six-weeks after delivery.”
- ADEQUATEEffect sizeThe effect sizes are large (81% reduction in odds of MDE or moderate-to-severe anxiety, with absolute rates 12% vs 41% for MDE) and are explicitly anchored as clinically significant, with reference to established minimal clinically important differences for PHQ-9 and HADS.
“This is considered clinically significant, corresponding to approximately a five-point decrease in the PHQ-9 for symptoms of depression and to a four-point decrease on the HADS anxiety scale compared to those not receiving the intervention.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites multiple studies on prevalence, co-occurrence, and consequences of prenatal anxiety, and notes the lack of preventive approaches. It also references the authors' own formative qualitative work that informed the intervention. The hypothesis follows logically from the cited evidence. Limitations of prior research (e.g., focus on depression, not anxiety) are explicitly addressed.
“Prenatal anxiety predicts anxiety, depression and suicide risk in the postnatal period”
“Effective treatments exist, but preventive approaches that could reduce the prevalence of severe postnatal depression are lacking”
“much less attention has been paid to maternal anxiety and stress than depression”
“Prenatal anxiety predicts anxiety, depression and suicide risk in the postnatal period”
“According to a systematic review, much less attention has been paid to maternal anxiety and stress than depression, and even fewer studies have tested the effects of prenatal interventions on postnatal anxiety”
Randomization used a pseudo random-number generator with permuted blocks of sizes 4, 8, 12, 16, and allocation concealment via opaque envelopes. The unit of randomization is the individual woman. Outcome assessors were blinded to allocation. A power calculation assumed 30% CMD prevalence, 30% reduction, 85% power, and 30% attrition, targeting 1200 women. Inclusion/exclusion criteria are detailed. The primary analysis is ITT; missing data were handled by complete-case analysis, and Little's MCAR test was used to assess missingness. No separate replication cohort is reported, but this is not expected for a single pivotal trial.
“Arm assignment was generated using a pseudo random-number generator by a trial statistician in the US based on randomly permuted blocks of size 4, 8, 12 and 16.”
“Outcome assessors were blinded to the trial arm allocation.”
“we calculated needing 840 pregnant women (420 in each arm) to achieve 85% power to detectable a 30% reduction in CMDs”
“Arm assignment was generated using a pseudo random-number generator by a trial statistician in the US based on randomly permuted blocks of size 4, 8, 12 and 16.”
“Outcome assessors were blinded to the trial arm allocation.”
“Assuming this prevalence, we calculated needing 840 pregnant women (420 in each arm) to achieve 85% power to detectable a 30% reduction in CMDs (30% in the control arm compared to 21% in the intervention arm).”
Sex is inherently female (pregnant women), and both arms include women, so sex_justified is not applicable. Age, gestational age, and health status (anxiety/depression scores) are reported in Table 1. Demographics such as education, income, and family structure are also reported. Species/strain and housing conditions are not applicable for a human trial.
“Participants had an average age of 25 (standard deviation (SD) 4.7) years at enrollment, with an average gestational age of 16 (SD=5) weeks.”
“Participants had at least mild anxiety at enrollment by design, with an average Hospital Anxiety and Depression Scale (HADS) anxiety score of 11.0 (SD=2.0).”
“Participants had an average age of 25 (standard deviation (SD) 4.7) years at enrollment, with an average gestational age of 16 (SD=5) weeks.”
“Table 1. Description of 1200 pregnant women enrolled in the HMHB trial by arm.”
Ethical approval was obtained from three named IRBs (Rawalpindi Medical University, Human Development Research Foundation, Johns Hopkins Bloomberg School of Public Health) with protocol numbers, and an NIMH-appointed DSMB. Written informed consent was obtained from all participants. The trial was registered at ClinicalTrials.gov. Regulatory compliance is implied through IRB approvals and DSMB oversight.
“Ethical approval for this research was obtained from the Institutional Review Boards of Rawalpindi Medical University, Human Development Research Foundation, and the Johns Hopkins Bloomberg School of Public Health and an NIMH-appointed Data Safety Monitoring Board.”
“All participants provided written informed consent prior to screening and to data collection.”
“Ethical approval for this research was obtained from the Institutional Review Boards of Rawalpindi Medical University, Human Development Research Foundation, and the Johns Hopkins Bloomberg School of Public Health and an NIMH-appointed Data Safety Monitoring Board.”
“All participants provided written informed consent prior to screening and to data collection.”
“clinicaltrials.gov (http://clinicaltrials.gov) #: NCT03880032”
The HMHB intervention is described in detail (six core sessions, booster sessions, content based on CBT and THP). The control condition (enhanced routine care) is also described. The paper identifies the software used for data collection (Open Data Kit) and statistical analysis (R 4.2.0). No antibodies, cell lines, or other bench reagents are used, so those criteria are not applicable.
“Statistical analyses were performed with R 4.2.0.”
“The Open Data Kit (ODK) was used to store data in a de-identified format.”
“The HMHB intervention While the intervention content was informed by formative qualitative research – , it uses the same core principles and strategies of the Thinking Healthy Program (THP), an evidence-based psychosocial intervention for mothers experiencing perinatal depression .”
“Statistical analyses were performed with R 4.2.0.”
The paper names logistic and linear regression, Student's t-test, and Chi-square test. Effect sizes (ORs and mean differences) are reported with 95% CIs. Exact p-values are given for primary outcomes (e.g., p = 7.0 ×10^-11). Assumptions are not explicitly tested, but the use of standard methods for large samples is acceptable. Data presentation includes tables with per-group Ns and dispersion. Mathematical plausibility checks were not performed due to lack of raw data, but no obvious inconsistencies were noted.
“we performed standard statistical comparisons such as the Student’s t-test and the Chi-square test”
“aOR=0.19, 95% CI: 0.14-0.28”
“p | 7.0 ×10^E-11”
“we performed standard statistical comparisons such as the Student’s t-test and the Chi-square test, depending on the type of variable.”
“aOR=0.19, 95% CI: 0.14-0.28”
“p | 7.0 ×10^E-11 | 2.8 ×10^E-20 | 1.7 ×10^E-22”
The data availability statement names the NIMH Data Archive (nda.nih.gov) as the repository, which is a concrete access route. The paper also mentions that data structures and codebooks are uploaded there. No custom code is mentioned, so code_sharing is not applicable. Repository deposit and accession numbers are not applicable for patient-level data, but the NIMH Data Archive provides a managed access mechanism.
“All our data and relevant codebooks have been submitted to the US National Institute of Mental Health’s data archive for public access, which can be accessed at https://nda.nih.gov/ .”
“Data structures, variables and variables names and relevant codebooks used for the statistical analyses have been uploaded to the NIMH Data Archives and are freely available to the public for use https://nda.nih.gov/ .”
“All our data and relevant codebooks have been submitted to the US National Institute of Mental Health’s data archive for public access, which can be accessed at https://nda.nih.gov/ .”
Methods are detailed enough for replication. The trial is registered (NCT03880032). Limitations are discussed extensively, including loss to follow-up and COVID-19 impact. Conclusions are proportional to the evidence. Funding (NIMH grant) and COI (none declared) are stated. No reporting guideline (e.g., CONSORT) is explicitly mentioned, but the paper follows standard trial reporting.
“One limitation to our intervention was loss to follow-up.”
“This study was supported by the National Institute of Mental Health at the US National Institutes of Health Grant # RO1 MH111859”
“The trial was registered at the US National Library of Medicine ( clinicaltrials.gov (http://clinicaltrials.gov) : NCT03880032”
“One limitation to our intervention was loss to follow-up.”
“This study was supported by the National Institute of Mental Health at the US National Institutes of Health Grant # RO1 MH111859”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 72 references by DOI: 62 verified — 10 no DOI (shown, not verified).
- NO DOIAnxiety and depression in pregnant women presenting in the OPD of a teaching hospitalNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRethinking mental health care: bridging the credibility gapNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWorld Mental Health: Problems and Priorities in Low-income CountriesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe conditions and consequences of choice: reflections on the measurement of women’s empowermentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIResources, Agency, Achievements: Reflections on the Measurement of Women’s EmpowermentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Power to Choose: Bangladeshi Women and Labour Market Decisions in London and DhakaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWomen’s empowerment and the question of choiceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITranslation and cultural adaptation of health questionnairesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICenter-based prevalence of anxiety and depression in women of the northern areas of PakistanNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIComparison of the Personal Health Questionnaire and the Self Reporting Questionnaire in rural PakistanNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://nda.nih.gov/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
7 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 7 minor suggestions below.
7 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAbstract, Results“The primary outcomes were met.”→ Consider rephrasing to 'The primary outcome was met.' for grammatical consistency.Minor grammatical issue.
- MINORconsistencyResults, Primary outcomes“12% of women in the intervention group developed MDE at six-weeks postpartum, versus 41% in the control group”→ Ensure percentages are consistent with Table 4 (11.6% and 40.5%).Percentages are rounded, which is acceptable but could be clarified.
- MINORclarityMethods, Power calculation“we calculated needing 840 pregnant women (420 in each arm) to achieve 85% power to detectable a 30% reduction in CMDs”→ Change 'to detectable' to 'to detect'.Typo.
- MINORconsistencyTable 4“7.0 ×10^E-11”→ Use standard scientific notation (e.g., 7.0 × 10^-11) without the 'E'.Formatting issue in table.
- MINORotherMethods, Statistical analyses“P-values were two-sided.”→ Consider specifying the significance level (e.g., alpha = 0.05).Minor omission.
- MINORtypoAbstract“NGiven many women are required by their families to be escorted to the hospital”→ Remove stray 'N' at the beginning of the sentence.Typographical error in the Discussion section.
- MINORclarityMethods, Power calculation“to detectable a 30% reduction”→ Change to 'to detect a 30% reduction'.Grammatical error.
The published work is robust and well-reported. An informed reader should weigh the minor reporting gaps (no explicit CONSORT reference, no explicit assumption checks, no code sharing) as transparency limitations, but none undermine the validity of the findings. No erratum or correction is warranted based on the identified issues.
- 1.MEDIUMreportingAdd an explicit statement referencing the CONSORT reporting guideline in the Methods or a separate section.Enhances transparency and aligns with standard trial reporting expectations.
- 2.MEDIUMstatisticsAdd a statement about verifying statistical assumptions (e.g., normality, homoscedasticity) for linear regression models, or justify why they are not needed, in the Statistical analyses section.Reviewer 2 flagged this as a reporting gap; explicit assumption checks strengthen the statistical analysis.
- 3.MEDIUMdata codeConsider sharing the analysis code in a public repository (e.g., GitHub) to facilitate reproducibility.Even though data are available, code sharing would improve reproducibility and is a common reviewer request.
- 4.MEDIUMreportingClarify the handling of missing data for secondary outcomes and sensitivity analyses in the Statistical analyses section.The paper uses complete-case analysis; specifying sensitivity analyses would address potential bias from missing data.
- 5.MEDIUMreportingProvide a more detailed description of the randomization sequence generation and allocation concealment in the Methods to fully meet CONSORT requirements.Reviewer 2 suggested this to fully meet CONSORT requirements.
- 6.MEDIUMreportingInclude a statement on whether any deviations from the pre-specified statistical analysis plan occurred, beyond those already mentioned.Transparency about protocol deviations is important for trial reporting.
- 7.MEDIUMreportingConsider reporting the number of participants who received each number of booster sessions to better describe intervention dose.Provides a clearer picture of intervention adherence and dose-response.
- 8.MEDIUMreportingAdd a note on the generalizability of the findings given the high attrition rate and the COVID-19 pandemic context.Contextualizes the results and helps readers interpret external validity.
- 9.LOWcopyeditFix the typo in the Methods, Power calculation: change 'to detectable' to 'to detect'.Grammatical error that should be corrected.
- 10.LOWcopyeditRemove the stray 'N' at the beginning of the sentence in the Abstract: 'NGiven many women are required...'.Typographical error that should be corrected.
- 11.LOWcopyeditUse standard scientific notation in Table 4 (e.g., 7.0 × 10^-11) instead of '7.0 ×10^E-11'.Formatting issue that could confuse readers.
- 12.LOWcopyeditSpecify the significance level (e.g., alpha = 0.05) in the Statistical analyses section where it says 'P-values were two-sided'.Minor omission that clarifies the threshold for significance.
- 13.LOWcopyeditEnsure percentages in the text (e.g., '12%' and '41%') are consistent with Table 4 (11.6% and 40.5%) or clarify that they are rounded.Rounding differences are acceptable but should be noted for consistency.
- 14.LOWreportingConsider reporting the number of participants who received the intervention as randomized (per-protocol) in addition to ITT.Addresses the 25% who never received any session and provides a sensitivity analysis.
- 15.LOWreportingClarify the definition of 'enhanced routine care' to specify what constitutes 'enhanced' beyond the listed components.Improves clarity of the control condition.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.