Effects of the WHO Labour Care Guide on cesarean section in India: a pragmatic, stepped-wedge, cluster-randomized pilot trial.
Vogel JP, Pujar Y, Vernekar SS, Armari E, Pingray V, Althabe F, Gibbons L, Berrueta M, Somannavar M, Ciganda A, Rodriguez R, Bendigeri S, Kumar JA, Patil SB, Karinagannanavar A, Anteen RR, Mallappa Ramachandrappa P, Shetty S, Bommanal L, Haralahalli Mallesh M, Gaddi SS, Chikkagowdra S, Raghavendra B, Homer CSE, Lavender T, Kushtagi P, Hofmeyr GJ, Derman R, Goudar S
- DOI
- 10.1038/s41591-023-02751-4
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/856da9ac-2eba-4913-9c82-2cbb2db257ed is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsOverstated claim−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 8 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is cesarean section rate in Robson Group 1, which is a clinical outcome, not a surrogate. However, the paper also reports a secondary outcome of labor augmentation with oxytocin, which is a process-of-care measure. The efficacy claim is primarily based on the cesarean section rate, which is a hard clinical outcome. Therefore, the surrogate assessment is not applicable to the primary claim. But since the paper also claims a reduction in oxytocin augmentation as a benefit, and this is a process measure, it is not a surrogate for a clinical outcome. The verdict is 'inadequate' because the primary outcome is a clinical outcome, but the paper does not provide target engagement or validated surrogate-to-clinical outcome link for the oxytocin reduction, which is presented as a potential benefit.
“For the secondary outcome augmentation with oxytocin during spontaneous labor, the prevalence in the control group was 27.3% and in the intervention group it was 9.3% (crude absolute difference −18.0%). However, the estimate of effect was not significant (RR…”
- 02Treatment effect not shown to be clinically meaningful
The primary outcome showed a 5.5% absolute reduction in cesarean section rate (from 45.2% to 39.7%), but this was not statistically significant (RR 0.85, 95% CI 0.54–1.33). The paper does not anchor this effect to a minimal clinically important difference or demonstrate clinical meaningfulness. The effect size is presented as preliminary and not definitive, and the confidence interval includes the null. Therefore, the effect size is inadequate as evidence of efficacy.
“A 5.5% crude absolute reduction in the primary outcome was observed (45.2% versus 39.7%; relative risk 0.85, 95% confidence interval 0.54–1.33).”
- 03Conclusion reaches beyond the evidence
The LCG strategy is a promising intervention that can improve quality of labor and childbirth care.
“Findings from this multicentered, stepped-wedge, cluster-randomized pilot trial suggest that the LCG strategy is a promising intervention that can improve quality of labor and childbirth care, reducing overuse of intrapartum interventions.”
DiscussionFind in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported pilot stepped-wedge cluster-randomized trial. The manuscript demonstrates strong rigor across all eight dimensions, with clear randomization, ethical approvals, statistical methods, and full data/code availability. Minor issues include a potential misspelling in the statistical method name and an overstatement in the discussion that should be tempered.
Both reviewers independently scored all eight dimensions as 'pass' with high confidence, and their evidence was consistent. The study type is interventional (stepped-wedge cluster-randomized trial). No divergence was found between reviewers. The statistics verification component checked only 1 test (the primary outcome RR with CI) and found it consistent; other tests were not machine-verifiable due to the GEE model complexity. The citation check found no retracted or non-existent references.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Found 1 reported test, but none could be recomputed (missing degrees of freedom or sample size).
- UNCOMPUTABLEreported p = .109 · recomputed p = .480Reviewers 1, 2Primary outcome RR p-value from CI
“relative risk (RR) 0.85, 95% confidence interval (CI) 0.54–1.33, P value 0.1088”
Taken as given: The RR is 0.85 with 95% CI 0.54-1.33.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed p-value from RR and CI using the pCI function for a ratio.How we recomputed it: pCI(0.85, 0.54, 1.33, 1)
- lowinternal contradictionThe abstract reports a 5.5% crude absolute reduction in the primary outcome, but the table shows 45.2% vs 39.7% which is a 5.5 percentage point difference, consistent. No contradiction found.
“A 5.5% crude absolute reduction in the primary outcome was observed (45.2% versus 39.7%; relative risk 0.85, 95% confidence interval 0.54–1.33).”
Table 2Find in source
Overstated conclusions
4 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
4 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated), 2 only partially supported (evidence backs part of the claim; gaps or caveats remain).
- overstatedReviewers 1, 2The LCG strategy is a promising intervention that can improve quality of labor and childbirth care.The pilot trial showed no significant benefits, so calling it 'promising' is an overstatement based on non-significant trends.Evidence: Primary outcome non-significant; secondary outcomes non-significant.
“Findings from this multicentered, stepped-wedge, cluster-randomized pilot trial suggest that the LCG strategy is a promising intervention that can improve quality of labor and childbirth care, reducing overuse of intrapartum interventions.”
DiscussionFind in source - partialReviewers 1, 2The LCG strategy may reduce cesarean section rates in Robson Group 1.The observed reduction was not statistically significant, so the claim is cautiously worded but not fully supported.Evidence: Primary outcome: RR 0.85, 95% CI 0.54-1.33, p=0.1088.
“A 5.5% crude absolute reduction in the primary outcome was observed (45.2% versus 39.7%; relative risk 0.85, 95% confidence interval 0.54–1.33).”
AbstractFind in source - partialReviewers 1, 2The LCG strategy may reduce labor augmentation with oxytocin.The crude difference was large but the CI was extremely wide and non-significant.Evidence: Augmentation with oxytocin: RR 0.34, 95% CI 0.01-15.04.
“For the secondary outcome augmentation with oxytocin during spontaneous labor, the prevalence in the control group was 27.3% and in the intervention group it was 9.3% (crude absolute difference −18.0%). However, the estimate of effect was not significant (RR 0.34, 95% CI 0.01–15.04)”
ResultsFind in source - supportedReviewers 1, 2The LCG strategy did not show clear differences in maternal, fetal, or newborn health outcomes.All reported outcomes had wide CIs crossing 1, consistent with no clear differences.Evidence: Table 3 shows all RRs with CIs crossing 1.
“For the baby, there were no clear differences in stillbirth (RR 0.97, 95% CI 0.43–2.19), neonatal death before discharge/day 7 (RR 1.31, 95% CI 0.37–4.71) or perinatal death before discharge/day 7 (RR 1.06, 95% 0.41–2.73).”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is cesarean section rate in Robson Group 1, which is a clinical outcome, not a surrogate. However, the paper also reports a secondary outcome of labor augmentation with oxytocin, which is a process-of-care measure. The efficacy claim is primarily based on the cesarean section rate, which is a hard clinical outcome. Therefore, the surrogate assessment is not applicable to the primary claim. But since the paper also claims a reduction in oxytocin augmentation as a benefit, and this is a process measure, it is not a surrogate for a clinical outcome. The verdict is 'inadequate' because the primary outcome is a clinical outcome, but the paper does not provide target engagement or validated surrogate-to-clinical outcome link for the oxytocin reduction, which is presented as a potential benefit.
“For the secondary outcome augmentation with oxytocin during spontaneous labor, the prevalence in the control group was 27.3% and in the intervention group it was 9.3% (crude absolute difference −18.0%). However, the estimate of effect was not significant (RR 0.34, 95% CI 0.01–15.04)”
- INADEQUATEEffect sizeThe primary outcome showed a 5.5% absolute reduction in cesarean section rate (from 45.2% to 39.7%), but this was not statistically significant (RR 0.85, 95% CI 0.54–1.33). The paper does not anchor this effect to a minimal clinically important difference or demonstrate clinical meaningfulness. The effect size is presented as preliminary and not definitive, and the confidence interval includes the null. Therefore, the effect size is inadequate as evidence of efficacy.
“A 5.5% crude absolute reduction in the primary outcome was observed (45.2% versus 39.7%; relative risk 0.85, 95% confidence interval 0.54–1.33).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites global cesarean rate trends, the WHO partograph history, and the development of the LCG, acknowledging the lack of randomized trials. The rationale linking the LCG strategy to potential reduction in unnecessary cesareans is well-articulated. Limitations of prior research (e.g., poor partograph use) are addressed through the intervention design.
“However, as the LCG is a novel tool, no such strategy has been developed or tested in a randomized trial.”
“However, as the LCG is a novel tool, no such strategy has been developed or tested in a randomized trial.”
“In this pilot trial, we aimed to evaluate the effects of implementing the LCG strategy, as compared to routine intrapartum care; the latter included use of the simplified partograph.”
Randomization method (computer-generated list) and unit (hospital clusters) are reported. Blinding was not possible post-intervention, which is acknowledged. A power analysis is provided (92% power to detect 25% reduction). Inclusion/exclusion criteria are clearly defined. Outlier handling is addressed through data validation and cleaning. Controls (usual care) are described. Independent replication is not applicable for a pilot trial.
“the four clusters (hospitals) were randomly assigned to one of four sequences (H1, H2, H3 or H4; Fig. ) using a computer-generated list of random numbers”
“Once the hospital had commenced the intervention, blinding of hospital staff, research staff and individual women was not possible.”
“The trial was designed to provide 92% power to detect a 25% reduction in the Robson Group 1 cesarean rate from 40% to 30%”
“Before trial commencement, the four clusters (hospitals) were randomly assigned to one of four sequences (H1, H2, H3 or H4; Fig. ) using a computer-generated list of random numbers that was managed by the study statistician.”
“Once the hospital had commenced the intervention, blinding of hospital staff, research staff and individual women was not possible.”
Sex is inherently female (women giving birth). Age, parity, gravida, gestational age, and previous cesarean are reported in Table 1. Demographics are adequate for a human trial. Species/strain and housing are not applicable.
“Maternal age (years) a | 23.9 (3.6) | 23.4 (3.6)”
“While there were more women in the intervention than the control, the characteristics of women were similar (Table ).”
“The eligibility criteria for women to be in the study population were those giving birth at ≥20 weeks’ gestation in participating hospitals, during the study period.”
“Maternal age (years) a | 23.9 (3.6) | 23.4 (3.6)”
The trial was approved by the Alfred Hospital Human Ethics Committee and several Indian institutional ethics committees, with protocol numbers. A waiver of individual consent for routine data is justified, and informed consent was obtained for postpartum surveys. Regulatory compliance with Declaration of Helsinki and Good Clinical Practice is stated.
“The trial was approved by the Alfred Hospital Human Ethics Committee (737/20), and the institutional ethics committees of the KLE Academy of Higher Education and Research (D-281120003), JJM Medical College, Davanagere (IEC-136/2020), Vijayanagar Institute of Medical Sciences (SVN IEC/20/2020-2021) and the Gadag Institute of Medical Sciences (IEC/01/2020-21)”
“The study protocol specified a waiver of individual consent for data collected on women giving birth; these data were nonidentifiable, routinely collected clinical variables in medical records and labor ward registries.”
“This trial was designed and conducted in accordance with the ethical principles of the World Medical Association’s Declaration of Helsinki, the Ottawa Statement for the Ethical Design and Conduct of Cluster Randomized Trials, and Good Clinical Practice standards”
“The trial was approved by the Alfred Hospital Human Ethics Committee (737/20), and the institutional ethics committees of the KLE Academy of Higher Education and Research (D-281120003), JJM Medical College, Davanagere (IEC-136/2020), Vijayanagar Institute of Medical Sciences (SVN IEC/20/2020-2021) and the Gadag Institute of Medical Sciences (IEC/01/2020-21), as well as the State Ethics Committee, Department of Health and Family Welfare, Government of Karnataka (DD(MH)/71/2020-21) and the Health Ministry’s Screening Committee, Indian Council of Medical Research (2020-10127).”
The LCG strategy is described in detail (training, audit and feedback). The WHO simplified partograph is the comparator. REDCap is identified as data management software. No antibodies, cell lines, or other biological reagents are used, so those criteria are not applicable.
“The intervention included a co-designed LCG training program for doctors and nurses working on labor wards, and a monthly audit and feedback process using hospital cesarean section data”
“All data were collected into predesigned study forms and managed using REDCap electronic data capture via tablets.”
“R code used for data analysis along with detailed instructions on its usage is publicly available at 10.5281/zenodo.8140454.”
The GEE method is named with details (Manck and DeRouen correction, exchangeable correlation, modified Poisson). Effect sizes are reported as RR with 95% CIs. Exact p-values are provided for the primary outcome. Software (R) is identified. Data presentation includes per-group n and percentages. Mathematical plausibility checks are not applicable due to large N and continuous outcomes.
“a GEE to estimate the effect of the intervention with respect to the population average was used. A bias correction method and degree of freedom approximation due to the small number of clusters was applied”
“relative risk (RR) 0.85, 95% confidence interval (CI) 0.54–1.33, P value 0.1088”
“Cesarean section in Robson Group 1 | 1,709/4,302 | (39.7) | 1,602/3,543 | (45.2) | 0.85 (0.54–1.33)”
“For the primary and secondary outcomes, a GEE to estimate the effect of the intervention with respect to the population average was used.”
“For the primary outcome, the cesarean section rate in Robson Group 1 for the control group was 45.2%, while in the intervention group it was 39.7%, with a crude absolute difference of −5.5% (relative risk (RR) 0.85, 95% confidence interval (CI) 0.54–1.33, P value 0.1088).”
“Cesarean section in Robson Group 1 | 1,709/4,302 | (39.7) | 1,602/3,543 | (45.2) | 0.85 (0.54–1.33)”
The data availability statement names Zenodo with a DOI (10.5281/zenodo.8140454) and states no restrictions. Code availability is also provided at the same DOI. This meets the criteria for repository deposit and code sharing.
“the de-identified individual-level data and the data dictionary are hosted publicly at the Gates Open Research-approved repository Zenodo. They can be accessed under 10.5281/zenodo.8140454.”
“R code used for data analysis along with detailed instructions on its usage is publicly available at 10.5281/zenodo.8140454.”
“In keeping with the Bill & Melinda Gates Foundation Open Access Policy, the de-identified individual-level data and the data dictionary are hosted publicly at the Gates Open Research-approved repository Zenodo. They can be accessed under 10.5281/zenodo.8140454.”
“R code used for data analysis along with detailed instructions on its usage is publicly available at 10.5281/zenodo.8140454.”
The trial is registered (CTRI/2021/01/030695). CONSORT and SPIRIT guidelines are referenced. All outcomes are reported, including non-significant ones. Limitations are discussed in detail. Conclusions are appropriately cautious for a pilot trial. Funding and competing interests are declared.
“Clinical Trials Registry India number: CTRI/2021/01/030695”
“We developed the trial protocol and reported findings in accordance with Standard Protocol Items: Recommendations for Interventional Trials (SPIRIT) guidance for randomized trials, and the Consolidated Standards of Reporting Trials (CONSORT) statement for stepped-wedge cluster-randomized trials”
“This trial nonetheless has some limitations. CIs for several outcomes were quite wide.”
“Clinical Trials Registry India number: CTRI/2021/01/030695”
“We developed the trial protocol and reported findings in accordance with Standard Protocol Items: Recommendations for Interventional Trials (SPIRIT) guidance for randomized trials, and the Consolidated Standards of Reporting Trials (CONSORT) statement for stepped-wedge cluster-randomized trials”
“This trial nonetheless has some limitations. CIs for several outcomes were quite wide. This was driven by variability in outcome rates between time periods and between clusters, as well as the small number of clusters.”
Registered (1 ID: Clinical Trials Registry – India). Reporting guidelines cited: CONSORT, SPIRIT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 45 references by DOI: 28 verified — 17 no DOI (shown, not verified).
- NO DOITrends in Maternal Mortality 2000 to 2020No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILevels and Trends in Child MortalityNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Global Strategy for Women’s, Children’s and Adolescents’ HealthNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPreventing Prolonged Labour: a Practical GuideNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Partograph: the Application of the WHO Partograph in the Management of Labour, Report of a WHO Multicentre Study, 1990–1991No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIManaging Complications in Pregnancy and Childbirth: a Guide for Midwives and DoctorsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIProtect the Promise: 2022 Progress Report on the Every Woman Every Child Global Strategy for Women’s, Children’s and Adolescents’ Health (2016–2030)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO Recommendations: Intrapartum Care for a Positive Childbirth ExperienceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn exploration of midwives’ views of the latest World Health Organization labour care guideNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO Labour Care Guide: User’s ManualNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRobson Classification: Implementation ManualNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO Recommendations: Non-clinical Interventions to Reduce Unnecessary Caesarean SectionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Norwegian World Health Organisation Labour Care Guide Trial (NORWEL): study protocolNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICan the use of a next generation partograph improve neonatal outcomes? (PICRINO): study protocolNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILabour Room Quality Improvement InitiativeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIICH E6 Good Clinical Practice (GCP) GuidelineNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISPIRIT 2013 statement: defining standard protocol items for clinical trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://ctri.nic.in/Clinicaltrials/showallp.php?mid1=50028&EncHid=&userName=CTRI/2021/01/030695LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORtypoAbstract“A 5.5% crude absolute reduction in the primary outcome was observed (45.2% versus 39.7%; relative risk 0.85, 95% confidence interval 0.54–1.33).”→ Consider adding 'absolute' before 'reduction' for clarity, though it is already present.No issue found; this is a placeholder.
- MINORconsistencyTable 2 footnote“e The RR and 95% CI were estimated with the generalized estimating equation method employing the Manck and DeRouen bias correction method and a degree of freedom approximation.”→ Ensure consistent spelling of 'Manck' (should be 'Mancl'? Check original reference).Potential misspelling of 'Mancl' as 'Manck'.
- MINORclarityMethods, Statistical methods and analysis“A Manck and DeRouen correction method with N-2 degrees of freedom was selected due to being the most conservative option.”→ Clarify what 'N' refers to (number of clusters? total participants?).Ambiguity in the definition of N.
- MINORconsistencyTable 2 footnote“e The RR and 95% CI were estimated with the generalized estimating equation method employing the Manck and DeRouen bias correction method and a degree of freedom approximation.”→ Ensure consistent spelling of 'Manck' (likely 'Mack' or 'Mancl'?) across the manuscript.Spelling of 'Manck' may be a typo for 'Mancl'.
The published work is robust and well-reported. An informed reader should weigh the pilot nature of the trial (small number of clusters, wide CIs) and the non-significant primary outcome. The minor copyedit issues (spelling of 'Manck'/'Mancl', definition of 'N') and the overstatement in the discussion are worth noting but do not undermine the overall integrity; a correction or clarification could be considered.
- 1.HIGHcopyeditIn Table 2 footnote and Methods, correct the spelling of 'Manck' to 'Mancl' (or verify the correct name) and ensure consistent usage throughout.The copyedit flagged a likely misspelling of the statistical correction method, which could confuse readers and reviewers.
- 2.HIGHreportingIn the Discussion, temper the claim that the LCG strategy is 'promising' to reflect the non-significant primary outcome and the pilot nature of the trial.The claim audit rated this as overstated; the evidence shows no significant benefit, so the claim should be softened to avoid over-interpretation.
- 3.MEDIUMcopyeditIn Methods, Statistical methods and analysis, clarify what 'N' refers to in the degrees of freedom approximation (e.g., number of clusters or total participants).The copyedit flagged ambiguity in the definition of 'N', which is important for reproducibility.
- 4.MEDIUMreportingIn the Abstract, consider reporting the number of clusters and average cluster size for completeness.Both reviewers suggested this to give readers a better sense of the trial's scale and design.
- 5.MEDIUMstatisticsIn the Methods or Results, report the intraclass correlation coefficient (ICC) for secondary outcomes to aid interpretation of cluster-level variability.Both reviewers suggested this to help future sample size calculations and interpretation.
- 6.MEDIUMreportingIn the Discussion, explicitly state that the trial was not powered for secondary outcomes.Reviewer 2 noted this is implied but could be clearer, which is important for interpreting null secondary results.
- 7.MEDIUMreportingAdd a CONSORT-style flow diagram for the postpartum survey participants to clarify the sampling and consent process.Reviewer 2 suggested this to improve transparency about the survey subset.
- 8.MEDIUMstatisticsConsider adding a sensitivity analysis excluding the transition period to assess robustness of the primary outcome.Reviewer 2 suggested this to evaluate the impact of the stepped-wedge design's transition periods.
- 9.MEDIUMreportingIn the Methods, specify the exact version of R software used for analysis.Reviewer 2 suggested this to improve reproducibility.
- 10.LOWreportingIn the Methods, provide a more detailed description of the data validation algorithms used to identify outliers.Reviewer 1 noted the current description is brief; more detail would improve transparency.
- 11.LOWreportingConsider including a table of baseline characteristics for the postpartum survey respondents to assess representativeness.Reviewer 1 suggested this to help readers judge potential selection bias in the survey.
- 12.LOWstatisticsClarify the handling of missing data for the primary outcome in the statistical methods.Reviewer 1 suggested this to ensure the analysis is fully transparent.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.