Technology-assisted cognitive-behavioral therapy for perinatal depression delivered by lived-experience peers: a cluster-randomized noninferiority trial.
Rahman A, Malik A, Nazir H, Zaidi A, Nisar A, Waqas A, Atif N, Gibbs NK, Luo Y, Sikander S, Wang D
- DOI
- 10.1038/s41591-025-03655-1
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/95aee62c-4d30-4d67-9e4b-991e80481b1f is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsOverstated claim−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ReportingData & code availability partially met−0.25★
- LinksDead data/code link−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is remission from major depressive episode (MDE) diagnosed with SCID, which is a clinical diagnostic endpoint, not a surrogate. However, the efficacy claim also relies on secondary outcomes like PHQ-9 scores, which are symptom scales. The primary outcome is a hard clinical outcome, so the surrogate verdict is not applicable. But since the primary outcome is a clinical diagnosis, it is adequate.
“The primary outcome was defined as remission from MDE at 3 months postnatal, evaluated with the SCID MDE module.”
- 02Conclusion reaches beyond the evidence
Peer involvement may surpass the effectiveness of trained community health workers.
“The findings also suggest that peer involvement may have the potential to surpass the effectiveness of trained community health workers.”
Discussion ¶4Find in source - 03Declared data/code link does not resolve
Dead link — nothing to verify.
“https://elements.liverpool.ac.uk/”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported cluster-randomized noninferiority trial with a strong scientific premise, rigorous design, and transparent reporting. The main weakness is the data/code availability: the repository link is dead and the access route is vague, which undermines reproducibility. Minor reporting gaps include incomplete identification of the THP-TAP app and lack of explicit outlier handling.
Both reviewers classified the study as interventional (cluster-randomized trial), and I adopt that classification. The evaluation covered all eight dimensions; species/strain, housing, antibodies, cell lines, and mycoplasma testing were not applicable as this is a human trial. The statistics verification recomputed only 3 tests (all consistent); the remaining statistics are unverified. The citation check found no retracted or unresolved references.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Primary outcome risk difference in ITT crude analysis
“The crude risk difference was 8.91%. The lower boundary of the one-sided 97.5% confidence interval (CI) was 4.25%, larger than −10% of the noninferiority margin, thus meeting the prespecified criterion for noninferiority ( P noninferiority < 0.0001).”
Taken as given: The risk difference is 8.91% with a two-sided 95% CI of (4.25, 13.56).; The p-value is for a one-sided noninferiority test, but the CI-based p is two-sided; the one-sided p is half of the two-sided p.; The CI is symmetric on the risk difference scale.Method: Used pCI to compute two-sided p from estimate and CI, then halved for one-sided.How we recomputed it: pCI(8.91, 4.25, 13.56, 0) - CONSISTENTreported p = .007 · recomputed p = .007Reviewers 1, 2Secondary outcome PHQ-9 at 3 months crude mean difference
“crude mean difference = −1.13, 95% CI = −1.96 to −0.31, P = 0.0071”
Taken as given: The mean difference is -1.13 with a 95% CI of (-1.96, -0.31).; The p-value is two-sided.; The CI is symmetric on the mean difference scale.Method: Used pCI to compute two-sided p from estimate and CI.How we recomputed it: pCI(-1.13, -1.96, -0.31, 0) - CONSISTENTreported p = .002 · recomputed p = .001Reviewers 1, 2Secondary outcome PHQ-9 >=10 at 3 months crude odds ratio
“crude odds ratio = 0.45, 95% CI = 0.28–0.74, P = 0.0016”
Taken as given: The odds ratio is 0.45 with a 95% CI of (0.28, 0.74).; The p-value is two-sided.; The CI is on the log scale.Method: Used pCI with log=1 for odds ratio.How we recomputed it: pCI(0.45, 0.28, 0.74, 1)
- lowinternal contradictionThe abstract states '846/980 (86.3%) participants at 3 months postnatal' were assessed, but the results section reports 428+418=846 assessed for the primary outcome, which is consistent. However, the number of participants lost to follow-up at 3 months is 134 (59+75), which would imply 980-134=846 assessed, consistent. No contradiction.
“On assessment of 846/980 (86.3%) participants at 3 months postnatal”
AbstractFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated), 1 only partially supported (evidence backs part of the claim; gaps or caveats remain).
- overstatedReviewer 2Peer involvement may surpass the effectiveness of trained community health workers.The trial shows noninferiority and a statistically significant difference in PHQ-9 at 3 months, but the primary outcome is noninferiority, not superiority. The claim of 'may surpass' is speculative and not directly supported by the primary outcome.Evidence: The primary outcome shows noninferiority, not superiority. The secondary PHQ-9 difference is significant, but the claim of surpassing is not a prespecified hypothesis.
“The findings also suggest that peer involvement may have the potential to surpass the effectiveness of trained community health workers.”
Discussion ¶4Find in source - partialReviewers 1, 2THP-TAP is a potentially scalable alternative to WHO-THP in low-resource settings.The trial demonstrates noninferiority and lower delivery costs in the optimized scenario, but scalability is inferred from cost modeling and process evaluation, not directly tested.Evidence: Costs section: per-participant cost US$69 for THP-TAP vs US$59 for WHO-THP, but optimized costs US$24 vs US$44. Process evaluation shows high acceptability.
“Estimating the optimized, real-world intervention delivery costs in which equipment and training are used to full capacity, we estimate the per-patient cost to fall to US $24 for THP-TAP and US $44 for WHO-THP.”
DiscussionFind in source - supportedReviewers 1, 2THP-TAP is noninferior to WHO-THP for remission from perinatal depression at 3 months postnatal.The primary outcome analysis shows a risk difference of 8.91% with a lower CI bound of 4.25%, exceeding the -10% margin, and the noninferiority p-value is <0.0001.Evidence: Table 2: ITT crude analysis: 395/428 (92.3%) vs 349/418 (83.5%), risk difference 8.91 (4.25, 13.56), P<0.0001.
“The crude risk difference was 8.91%. The lower boundary of the one-sided 97.5% confidence interval (CI) was 4.25%, larger than −10% of the noninferiority margin, thus meeting the prespecified criterion for noninferiority ( P noninferiority < 0.0001).”
ResultsFind in source - supportedReviewers 1, 2THP-TAP reduces depressive symptoms more than WHO-THP at 3 months postnatal.The PHQ-9 mean difference at 3 months is -1.13 (95% CI -1.96 to -0.31, P=0.0071), indicating a statistically significant reduction.Evidence: Table 3: PHQ-9 at 3 months: mean difference -1.13 (95% CI -1.96 to -0.31), P=0.0071.
“THP-TAP group showed significantly lower mean depressive symptom scores measured with the 9-item Personal Health Questionnaire (PHQ-9) at 3 months, compared to the WHO-THP group, in the ITT analysis (crude mean difference = −1.13, 95% CI = −1.96 to −0.31, P = 0.0071”
ResultsFind in source - supportedReviewers 1, 2The intervention is safe with no adverse events linked to the intervention.The safety section reports no adverse events linked to the intervention, with two cases of deterioration due to domestic violence referred for care.Evidence: Safety section: 'there were no adverse events linked to intervention reported in THP-TAP or WHO-THP groups.'
“During the trial, there were no adverse events linked to intervention reported in THP-TAP or WHO-THP groups.”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary outcome is remission from major depressive episode (MDE) diagnosed with SCID, which is a clinical diagnostic endpoint, not a surrogate. However, the efficacy claim also relies on secondary outcomes like PHQ-9 scores, which are symptom scales. The primary outcome is a hard clinical outcome, so the surrogate verdict is not applicable. But since the primary outcome is a clinical diagnosis, it is adequate.
“The primary outcome was defined as remission from MDE at 3 months postnatal, evaluated with the SCID MDE module.”
- ADEQUATEEffect sizeThe primary outcome shows a remission rate of 92.3% in THP-TAP vs 83.5% in WHO-THP, with a risk difference of 8.91% and lower CI bound 4.25%, exceeding the noninferiority margin of -10%. This is a clinically meaningful difference, and the effect size is anchored to a prespecified margin and clinical remission.
“The difference in the remission rate was 8.91% with the lower boundary of the one-sided 97.5% confidence interval being 4.25%, larger than the prespecified −10% noninferiority margin (P noninferiority < 0.0001).”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
3 integrity concerns flagged (0 high).
- lowotherThe paper reports a noninferiority margin of -10% and a risk difference of 8.91% favoring THP-TAP, which is unusual for a noninferiority trial (typically the new treatment is expected to be slightly worse). This is not a validity threat but worth noting.
“The crude risk difference was 8.91%.”
ResultsFind in source - lowotherThe paper reports that the trial was single-blind, but the assessors were blinded. This is standard for psychotherapy trials, but the term 'single-blind' might be ambiguous; however, it is explained.
Our trial was a single-blind study, as participants cannot be blinded to 'talking therapies'.
Discussionreviewer’s wording
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites a systematic review and IPD meta-analysis of task-shared psychological interventions, the original THP trial, and WHO adoption, acknowledging both strengths and limitations (e.g., 'voltage-drop', 'program drift', critiques of reductionism). The rationale for the noninferiority design and the choice of margin are explicitly justified. The paper also addresses limitations of prior work by describing how peer delivery and digital assistance aim to overcome quality and fidelity issues.
“A systematic literature review and individual patient data meta-analysis of 11 trials of 4,145 participants from LMICs showed that task-shared psychological interventions for depression were associated with a greater decrease in symptom severity than control conditions”
“Furthermore, we believed depriving depressed women of an efficacious treatment would be unethical, which could be avoided by using the noninferiority trial design.”
Randomization used permuted blocks stratified by union council, with allocation by an independent statistician. Assessors were blinded, and participants/delivery agents could not be blinded due to the nature of the intervention, which is acknowledged. Power analysis is detailed with assumptions (75% remission, -10% margin, ICC 0.005, 20% loss, 87.2% power). Inclusion/exclusion criteria are pre-specified. Outlier handling is not explicitly described, but the analysis population (ITT and per-protocol) and missing data imputation are defined. Controls are inherent in the active comparator design.
“Randomization codes were generated via a permuted-block randomization method (stratified by UC). Block sizes varied at two, four and six. Allocation of clusters was carried out by an independent statistician based in Liverpool using the SAS PROC Plan.”
“Due to the nature of the intervention, it was not possible to mask participants and delivery agents to treatment allocation. However, outcome assessors, who were nonresidents of the study area and independent of the intervention delivery procedures, were masked to treatment allocation.”
“Allowing for 70 village clusters randomized with a 1:1 allocation ratio and 14 depressed participants per village cluster and 20% loss to follow-up, a total of 980 participants were required to detect noninferiority for the primary outcome at 3 months postnatal with a power of 87.2%.”
“Randomization codes were generated via a permuted-block randomization method (stratified by UC). Block sizes varied at two, four and six. Allocation of clusters was carried out by an independent statistician based in Liverpool using the SAS PROC Plan.”
“However, outcome assessors, who were nonresidents of the study area and independent of the intervention delivery procedures, were masked to treatment allocation.”
“Allowing for 70 village clusters randomized with a 1:1 allocation ratio and 14 depressed participants per village cluster and 20% loss to follow-up, a total of 980 participants were required to detect noninferiority for the primary outcome at 3 months postnatal with a power of 87.2%.”
Sex is reported (all female) and is justified by the condition (perinatal depression). Age, education, occupational status, and baseline PHQ-9/GAD-7/WHO-DAS scores are reported in Table 1. Species/strain and housing conditions are not applicable as this is a human trial. Demographics are adequately reported.
“We included women in their second or third pregnancy trimester, aged 18 years and above”
“Age (mean, s.d.) | 27.20 (5.15) | 27.29 (4.94)”
“PHQ-9 scores (mean, s.d.) | 16.73 (4.57) | 16.16 (4.51)”
“Age (mean, s.d.) | 27.20 (5.15) | 27.29 (4.94)”
“PHQ-9 scores (mean, s.d.) | 16.73 (4.57) | 16.16 (4.51)”
Ethical approval was obtained from the Ethics Review Committee at the University of Liverpool, the HDRF Pakistan, and the National Bioethics Committee, Pakistan. Informed consent is mentioned in the eligibility criteria ('Women with current MDE who provided informed consent were recruited'). Regulatory compliance is implied through adherence to ethical standards and the inclusion of an ethics statement.
“Ethical approval was obtained from the Ethics Review Committee at the University of Liverpool, the HDRF, Pakistan (implementing organization), and the National Bioethics Committee, Pakistan.”
“Women with current MDE who provided informed consent were recruited.”
“Safeguards were implemented to ensure the well-being and safety of all participants involved.”
“Ethical approval was obtained from the Ethics Review Committee at the University of Liverpool, the HDRF, Pakistan (implementing organization), and the National Bioethics Committee, Pakistan.”
“Women with current MDE who provided informed consent were recruited.”
The THP-TAP app is described as an Android tablet/smartphone-based app with in-built training, supervision, and monitoring features. The WHO-THP manual is referenced as downloadable from the WHO website. The peers and LHWs are described in terms of training and background. The app is a key resource, and its development is referenced to another publication. Software used for analysis (SAS 9.4) is identified.
“The Android tablet or smartphone-based App had in-built features of training, supervision and monitoring.”
“All statistical analyses were performed using SAS 9.4.”
“The Android tablet or smartphone-based App had in-built features of training, supervision and monitoring.”
“All statistical analyses were performed using SAS 9.4.”
Tests are named (GEE, GLMM). Assumptions are handled by design (e.g., GEE for clustered data). Exact p-values are reported (e.g., P=0.0071). Effect sizes with CIs are reported throughout. Software is identified. Data presentation includes per-group n and CIs. Mathematical plausibility checks: the primary outcome percentages (92.3% and 83.5%) are plausible given the numerators and denominators. No arithmetic errors detected.
“Primary outcome was analyzed using a generalized estimating equation (GEE) model.”
“crude mean difference = −1.13, 95% CI = −1.96 to −0.31, P = 0.0071”
“8.91 (4.25, 13.56)”
“Primary outcome was analyzed using a generalized estimating equation (GEE) model.”
“crude mean difference = −1.13, 95% CI = −1.96 to −0.31, P = 0.0071”
The data availability statement mentions deposition at the University of Liverpool repository, but the access route is 'by submitting a request to the corresponding author' with conditions to be assessed, which is less concrete than a named managed-access platform. Code availability is also 'on request' with no repository. Repository deposit is planned but not yet available.
“Data from our trial will be deposited at the University of Liverpool data repository ( https://elements.liverpool.ac.uk/ ). This will include deidentified individual participant data and the data dictionary. Dataset will be made available by submitting a request to the corresponding author.”
“After publication, analytic code will be shared with researchers who submit a formal proposal to the corresponding author (A.R.).”
“Data from our trial will be deposited at the University of Liverpool data repository ( https://elements.liverpool.ac.uk/ ). This will include deidentified individual participant data and the data dictionary. Dataset will be made available by submitting a request to the corresponding author.”
“After publication, analytic code will be shared with researchers who submit a formal proposal to the corresponding author (A.R.).”
The paper states it follows CONSORT guidelines. Trial registration number (NCT05353491) is provided. All primary and secondary outcomes are reported, though some prespecified analyses (EQ-5D, mediation) are deferred to future publications, which is disclosed. Limitations are discussed, including single-blind design and potential Hawthorne effects. Conclusions are proportional to the evidence, with appropriate caveats about subgroup analyses. Funding (NIHR) and competing interests are declared.
“ClinicalTrials.gov registration: NCT05353491”
“The trial results are reported following the CONSORT guidelines for cluster-randomized trials.”
“A potential limitation is that the site has been used in previous THP trials”
“ClinicalTrials.gov registration: NCT05353491”
“The trial results are reported following the CONSORT guidelines for cluster-randomized trials.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst mediumReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- Dead data/code linksRecomputed
Checked 42 references by DOI: 34 verified — 8 no DOI (shown, not verified).
- NO DOIMaternal depression: a global threat to children’s health, development, and behavior and to human rightsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThinking Healthy: A Manual for Psychosocial Management of Perinatal DepressionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITask-shifting or problem-shifting? How lay counselling is redefining mental healthcareNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStories and metaphors in cognitive-behavior therapyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBridging the Gender Digital Divide: Challenges and an Urgent Call for Action for Equitable Digital Skills DevelopmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILandscape analysis of the family planning situation in Pakistan—District profile: RawalpindiNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStructured Clinical Interview for DSM-5—Research Version (SCID-5 for DSM-5, Research Version; SCID-5-RV)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFramework analysis: a method for analysing qualitative dataNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 0 live, 1 dead.
- datahttps://elements.liverpool.ac.uk/DEADHTTP 400Dead link — nothing to verify.
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly typo, consistency, grammar.
- MINORtypoAbstract“technology-assisted peer-delivered THP (THP-TAP)”→ Ensure consistent use of 'THP-TAP' throughout.The acronym is defined but used inconsistently in the abstract.
- MINORconsistencyTable 1“Occupational status | Housewife | 473 (97.1%) | 479 (97.2)”→ Add '%' to the second percentage for consistency.Minor formatting inconsistency in table.
- MINORgrammarDiscussion“This opens up important avenues for discussion regarding the unique advantages peers might bring, such as reduced workloads, heightened motivation and deeper familiarity with the community, which could enhance intervention delivery and outcomes.”→ Consider splitting the long sentence for clarity.Long sentence but grammatically correct.
- MINORtypoAbstract“technology-assisted peer-delivered THP (THP-TAP)”→ Consider defining the abbreviation 'THP-TAP' at first use in the abstract, as it is used before the full term is given.The abbreviation is defined in the introduction, but the abstract uses it without definition.
- MINORconsistencyResults, 'Primary outcome'“P noninferiority < 0.0001”→ Consider reporting the exact p-value if available, or use 'P < 0.0001' consistently.The paper uses 'P noninferiority' in some places and 'P' in others; consistency would improve clarity.
The published work is robust overall, but an informed reader should weigh the dead data repository link and the vague data/code access route as a reproducibility concern. The over-claim about peers 'surpassing' CHWs should be tempered. No erratum is warranted for the core findings, but the authors should correct the data availability link and clarify the access mechanism.
- 1.HIGHdata codeFix the dead data repository link (https://elements.liverpool.ac.uk/) in the Data availability section, or replace it with a working URL or DOI.The reproducibility check found the link is dead, which directly undermines the data availability statement and prevents readers from accessing the data.
- 2.HIGHdata codeSpecify a concrete data access mechanism in the Data availability section, such as a named data access committee or a managed-access platform (e.g., Vivli), with conditions and a timeframe for data release.The current 'request to the corresponding author' route is vague and does not meet transparency standards for patient-level data.
- 3.HIGHdata codeDeposit the analytic code in a public repository (e.g., Zenodo or GitHub) with a DOI, and provide the link in the Code availability section.Code sharing 'on request' is not reproducible; a public repository with a persistent identifier is the expected standard.
- 4.HIGHreportingTemper the claim that 'peer involvement may surpass the effectiveness of trained community health workers' in the Discussion to reflect that the trial demonstrated noninferiority, not superiority.The primary outcome is noninferiority, so a claim of 'surpassing' is not supported by the evidence and could mislead readers.
- 5.MEDIUMrigorIn the Methods, THP-TAP section, provide the app version or a reference to its repository or documentation to fully identify the investigational product.The app is a key resource, and without a version or repository, the intervention cannot be precisely reproduced.
- 6.MEDIUMrigorIn the Methods, Statistical analysis, explicitly state how outliers were handled, or note that no outliers were excluded.One reviewer flagged outlier handling as inadequate; explicit reporting would resolve this ambiguity.
- 7.MEDIUMreportingIn the Abstract, define the abbreviation 'THP-TAP' at first use, as it is used before the full term is given.The copyedit pass noted the abbreviation is used without definition in the abstract, which could confuse readers.
- 8.MEDIUMcopyeditIn Table 1, add '%' to the second percentage for occupational status (e.g., '479 (97.2%)') for consistency.The copyedit pass flagged a minor formatting inconsistency in the table.
- 9.MEDIUMreportingIn the Results, report the exact p-value for the noninferiority test in the abstract instead of 'P noninferiority < 0.0001'.Exact p-values enhance precision and consistency with the rest of the paper.
- 10.LOWreportingIn the Methods, specify the exact version of the SCID used (e.g., SCID-5) to ensure reproducibility.The version of the diagnostic instrument is not stated, which is a minor reproducibility gap.
- 11.LOWreportingIn the Statistical Analysis section, describe how the intracluster correlation coefficient was estimated and whether it was prespecified.This would clarify the sample size calculation and analysis assumptions.
- 12.LOWreportingIn the Discussion, discuss the generalizability of the findings to other LMIC settings, given the specific cultural and health system context of Pakistan.The trial was conducted in a single district, so scalability to other contexts is uncertain.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.