A cost-effectiveness analysis of early detection and bundled treatment of postpartum hemorrhage alongside the E-MOTIVE trial.
Williams EV, Goranitis I, Oppong R, Perry SJ, Devall AJ, Martin JT, Mammoliti KM, Beeson LE, Sindhu KN, Galadanci H, Alwy Al-Beity F, Qureshi Z, Hofmeyr GJ, Moran N, Fawcus S, Mandondo S, Middleton L, Hemming K, Oladapo OT, Gallos ID, Coomarasamy A, Roberts TE
- DOI
- 10.1038/s41591-024-03069-5
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/64aace4e-0be3-45dd-8c66-ff11d1483a5d is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- LinksDead data/code link−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on the surrogate outcome of severe PPH (blood loss ≥1,000 ml) and DALYs averted. While DALYs are a composite measure, the trial's primary outcome is a surrogate (blood loss) and the paper does not provide evidence linking this surrogate to long-term clinical outcomes beyond the trial's short-term follow-up. Target engagement is not explicitly demonstrated for the intervention components.
“Severe PPH occurred in 786 of 48,678 patients (1.6%) in the E-MOTIVE group and in 2129 of 50,043 (4.3%) in the usual-care group (adjusted risk difference −2.6%, 95% confidence interval (CI) −3.1% to −2.1%; Table ).”
- 02Treatment effect not shown to be clinically meaningful
The effect size is reported as a risk difference of -2.6 percentage points for severe PPH, which is a reduction from 4.3% to 1.6%. While statistically significant, the clinical meaningfulness is not explicitly anchored to a minimal clinically important difference or other benchmarks. The DALY difference is small and not statistically significant (adjusted difference -0.00266, 95% CI -0.00814 to 0.00287).
“The adjusted DALY difference between E-MOTIVE and usual care per patient was −0.00266 (95% CI −0.00814 to 0.00287; Table ).”
- 03Declared data/code link does not resolve
Dead link — nothing to verify.
“https://github.com/ewbham/E-MOTIVE”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted economic evaluation alongside a cluster-randomized trial, with strong reporting of methods, ethics, and data/code availability. Minor reporting gaps include lack of explicit blinding statement, power analysis details, and a small inconsistency in the usual-care group N between text and table.
The study type is classified as observational (economic evaluation) rather than interventional, as it analyzes trial data without conducting a new intervention. The reviewers diverged on study type (one said observational, one said interventional); the economic evaluation is a secondary analysis of trial data, so observational is more accurate. Biological variables and several sub-criteria (e.g., blinding, power analysis) are not applicable or not fully reported but do not affect the overall pass.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks.
- CONSISTENTreported p < .050 · recomputed p = <.001Reviewer 1Check the p-value for the adjusted risk difference of severe PPH using the reported 95% CI.
“Severe PPH occurred in 786 of 48,678 patients (1.6%) in the E-MOTIVE group and in 2129 of 50,043 (4.3%) in the usual-care group (adjusted risk difference −2.6%, 95% confidence interval (CI) −3.1% to −2.1%; Table ).”
Taken as given: The adjusted risk difference is -2.6 percentage points.; The 95% CI is -3.1 to -2.1 percentage points.; The CI is two-sided at 95%.Method: Used pCI function to compute p-value from estimate and CI, assuming normal approximation.How we recomputed it: pCI(-2.6, -3.1, -2.1, 0) - CONSISTENTreported p > .050 · recomputed p = .344Reviewer 1Check the p-value for the adjusted DALY difference using the reported 95% CI.
“The adjusted DALY difference between E-MOTIVE and usual care per patient was −0.00266 (95% CI −0.00814 to 0.00287; Table ).”
Taken as given: The adjusted DALY difference is -0.00266.; The 95% CI is -0.00814 to 0.00287.; The CI is two-sided at 95%.Method: Used pCI function to compute p-value from estimate and CI, assuming normal approximation.How we recomputed it: pCI(-0.00266, -0.00814, 0.00287, 0)
- lowinternal contradictionThe usual-care group N is reported as 50,043 in the text but 50,044 in Table 1.
“Severe PPH occurred in 786 of 48,678 patients (1.6%) in the E-MOTIVE group and in 2129 of 50,043 (4.3%) in the usual-care group (adjusted risk difference −2.6%, 95% confidence interval (CI) −3.1% to −2.1%; Table ).”
Table 1Find in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2The E-MOTIVE intervention is cost-effective in each participating country.Country-level analyses are based on pooled clinical data with country-specific costs, which may not fully reflect country-specific effectiveness.Evidence: Country-level ICERs compared against country-specific thresholds in Extended Data Table 2.
“Briefly, the E-MOTIVE intervention was judged to be cost-effective for each participating country when the ICERs were compared against both country-specific GDP-based WTP thresholds and opportunity-cost-based WTP thresholds (Extended Data Table ).”
ResultsFind in source - supportedReviewers 1, 2The E-MOTIVE intervention is cost-effective compared with usual care.The paper provides ICERs below willingness-to-pay thresholds and sensitivity analyses supporting cost-effectiveness.Evidence: ICERs of 11.83 USD per severe PPH averted and 113.91 USD per DALY averted, both below the GDP-based and opportunity-cost thresholds.
“The estimated incremental cost-effectiveness ratios (ICERs) (Table ) are therefore 11.83 USD per case of severe PPH averted and 113.91 USD per DALY averted.”
Results ¶3Find in source - supportedReviewers 1, 2The E-MOTIVE intervention reduces severe PPH compared with usual care.The adjusted risk difference of -2.6% with a 95% CI excluding zero supports a significant reduction.Evidence: Adjusted risk difference -2.6% (95% CI -3.1% to -2.1%).
“Severe PPH occurred in 786 of 48,678 patients (1.6%) in the E-MOTIVE group and in 2129 of 50,043 (4.3%) in the usual-care group (adjusted risk difference −2.6%, 95% confidence interval (CI) −3.1% to −2.1%; Table ).”
Results ¶2Find in source - supportedReviewer 1The E-MOTIVE intervention is a worthwhile use of healthcare budgets.The cost-effectiveness results and sensitivity analyses support this conclusion.Evidence: ICERs below thresholds and sensitivity analyses showing robustness.
Therefore, provision of calibrated blood-collection drapes and use of bundled first-response treatment can be considered a worthwhile use of constrained healthcare budgets.
Discussionreviewer’s wording - supportedReviewer 2Reducing the cost of the calibrated drape could lead to cost savings.Sensitivity analyses show that at drape costs of $1, $0.75, and $0.50, the intervention becomes dominant (less costly and more effective), supporting the claim.Evidence: Table 2 and sensitivity analysis section.
If the device cost of the calibrated drape is reduced to 1 USD (2023 prices), the E-MOTIVE intervention becomes comparable in cost to usual care, while being more effective (Table 2). Further reductions in the cost of the calibrated drape could potentially result in cost savings.
Resultsreviewer’s wording
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on the surrogate outcome of severe PPH (blood loss ≥1,000 ml) and DALYs averted. While DALYs are a composite measure, the trial's primary outcome is a surrogate (blood loss) and the paper does not provide evidence linking this surrogate to long-term clinical outcomes beyond the trial's short-term follow-up. Target engagement is not explicitly demonstrated for the intervention components.
“Severe PPH occurred in 786 of 48,678 patients (1.6%) in the E-MOTIVE group and in 2129 of 50,043 (4.3%) in the usual-care group (adjusted risk difference −2.6%, 95% confidence interval (CI) −3.1% to −2.1%; Table ).”
- INADEQUATEEffect sizeThe effect size is reported as a risk difference of -2.6 percentage points for severe PPH, which is a reduction from 4.3% to 1.6%. While statistically significant, the clinical meaningfulness is not explicitly anchored to a minimal clinically important difference or other benchmarks. The DALY difference is small and not statistically significant (adjusted difference -0.00266, 95% CI -0.00814 to 0.00287).
“The adjusted DALY difference between E-MOTIVE and usual care per patient was −0.00266 (95% CI −0.00814 to 0.00287; Table ).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on PPH as a leading cause of maternal death, the inaccuracy of visual blood loss estimation, and the delayed/inconsistent use of effective interventions. It acknowledges limitations of prior work, such as the late administration of tranexamic acid and low uptake of recommendations. The rationale logically links these challenges to the study's objective of evaluating cost-effectiveness of early detection and bundled treatment.
“In this Article, we report the economic evaluation conducted alongside the E-MOTIVE trial, an integral component of the E-MOTIVE project, which aimed to assess the cost-effectiveness of the E-MOTIVE intervention compared with usual care.”
“To address these challenges, the cluster-randomized E-MOTIVE trial was designed to assess a multicomponent intervention for detection and treatment of PPH in patients having vaginal delivery.”
The trial is a parallel cluster-randomized trial with a baseline control phase. Randomization was performed using a minimization algorithm by an independent statistician, with the unit of randomization being the hospital. Blinding is not described, which is typical for a cluster trial of a complex intervention where blinding is infeasible; the paper does not mention blinding, but this is not a critical flaw given the pragmatic design. A power analysis was conducted to determine sample size, though the specific parameters are not detailed in this paper (referenced to the clinical outcomes paper). Inclusion/exclusion criteria for hospitals are clearly stated. The analysis accounts for clustering using multilevel models. The paper does not report a separate replication cohort, which is not applicable for a single pivotal trial.
“Hospitals were eligible for inclusion if they were geographically and administratively distinct from each other, had between 1,000 and 5,000 vaginal births per year, and were able to provide comprehensive obstetrical care with the ability to perform surgery for PPH.”
“Both were carried out on an intention-to-treat basis and relied on complete case analysis wherein cases without source-verified blood loss data were excluded.”
“Hospitals were eligible for inclusion if they were geographically and administratively distinct from each other, had between 1,000 and 5,000 vaginal births per year, and were able to provide comprehensive obstetrical care with the ability to perform surgery for PPH. Hospitals were excluded if they had already implemented a treatment bundle for PPH.”
The study involves human participants (women giving birth), but the analysis is at the population level for cost-effectiveness. The paper does not report individual-level biological variables like age or comorbidities, which is typical for a trial-level economic evaluation. The study focuses on healthcare system costs and outcomes, not biological mechanisms.
The paper lists specific ethics committees that approved the study, including the University of Birmingham STEM ethics committee, WHO-HRP, and national committees in Kenya, Nigeria, South Africa, and Tanzania, with protocol numbers. It also states that written informed consent was obtained from participants. Regulatory compliance is implied through adherence to national regulations and data protection.
“All participants provided written informed consent before participation in intervention training.”
“All participants provided written informed consent before participation in intervention training.”
The calibrated blood-collection drape is identified by manufacturer (Excellent Fixable Drapes in India) and cost. Drugs (oxytocic drugs, TXA, IV fluids) are identified with cost sources. Software (Stata, REDCap) is identified with version numbers. The paper does not report lot numbers or RRIDs, which is acceptable for a health economic evaluation.
“Calibrated blood-collection drape costs were obtained from Excellent Fixable Drapes in India, the manufacturer and supplier of the drapes used in the E-MOTIVE trial.”
“All analyses were carried out using Stata, version 17.1 (StataCorp).”
“Calibrated blood-collection drape costs were obtained from Excellent Fixable Drapes in India, the manufacturer and supplier of the drapes used in the E-MOTIVE trial.”
“All analyses were carried out using Stata, version 17.1 (StataCorp).”
“Resource use information was collected prospectively via electronic case report forms and recorded in REDCap (version 10.9.0–13.3.2).”
The paper names the statistical tests used (multilevel modeling with binomial/logit for severe PPH, Gaussian/identity for costs and DALYs, nonparametric permutation tests). Assumptions are addressed through the use of robust standard errors and permutation tests. Exact p-values are not reported; instead, 95% confidence intervals are provided, which is appropriate for an economic evaluation. Effect sizes (risk differences, cost differences) are reported with confidence intervals. Software is identified. Data presentation includes tables with per-group Ns, means, and SDs. Mathematical plausibility checks are not applicable due to the large sample size and continuous cost data.
“For severe PPH, we used the binomial family and logit link, in addition to robust standard errors, followed by marginal standardization to estimate risk difference.”
“All analyses were carried out using Stata, version 17.1 (StataCorp).”
“All analyses were carried out using Stata, version 17.1 (StataCorp).”
The data availability statement explains that patient data cannot be made public due to privacy, but provides a managed access route via the E-MOTIVE Trial Data Analysis Sub-Committee with a contact and review timeframe. Code is shared on GitHub. Repository deposit and accession numbers are not applicable for patient-level data.
“The complete de-identified patient data that support the findings of this study can be obtained from the Chief Investigator of the E-MOTIVE trial, on approval from the E-MOTIVE Trial Data Analysis Sub-Committee.”
“Stata codes are available via GitHub at https://github.com/ewbham/E-MOTIVE .”
“Patient data cannot be made publicly available due to privacy concerns. The complete de-identified patient data that support the findings of this study can be obtained from the Chief Investigator of the E-MOTIVE trial, on approval from the E-MOTIVE Trial Data Analysis Sub-Committee.”
“Stata codes are available via GitHub at https://github.com/ewbham/E-MOTIVE .”
The trial is registered on ClinicalTrials.gov (NCT04341662) and the Pan African Clinical Trials Registry. Methods are detailed enough for replication. A reporting guideline (Nature Portfolio reporting summary) is referenced. All outcomes are reported, including negative/null results (e.g., no significant difference in costs). Limitations are discussed, including the lack of bottom-up costing, societal perspective, and equity analysis. Conclusions are proportional to the evidence, stating cost-effectiveness without overclaiming. Funding and competing interests are declared.
“ClinicalTrials.gov identifier: NCT04341662”
“ClinicalTrials.gov identifier: NCT04341662 (https://clinicaltrials.gov/ct2/show/NCT04341662?term=NCT04341662) .”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
1 finding · worst mediumReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- Dead data/code linksRecomputed
Checked 39 references by DOI: 28 verified — 11 no DOI (shown, not verified).
- NO DOIWHO Recommendations for the Prevention and Treatment of Postpartum HaemorrhageNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGlobal Burden of Disease Study 2019 (GBD 2019) Disability WeightsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGlobal Burden of Disease Study 2019 (GBD 2019) Life Tables 1950–2019No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIUses of Medicines for Prevention and Treatment of Post-partum Hemorrhage and Other Obstetric PurposesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInternational Medical Product Price GuideNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIUnit costs of health care inputs in low and middle income regionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO-CHOICE Estimates of Cost for Inpatient and Outpatient Health Service DeliveryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISalary Scales, with Translation Keys, for Employees on Salary Levels 1 to 12 and Those Employees Covered by Occupation Specific Dispensions (OSDs)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIUNICEF Supply Catalogue Vol. 2023No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIUniform Patient Fee Schedule 2022 Vol. 2023No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGDP per Capita (Current US$)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 1 live, 1 dead.
- datahttps://clinicaltrials.gov/ct2/show/NCT04341662?term=NCT04341662LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubDEADHTTP 404https://github.com/ewbham/E-MOTIVEDead link — nothing to verify.
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, clarity, other.
- MINORconsistencyTable 1“Usual care ( N = 50,044)”→ Ensure the N matches the text (50,043) or clarify the discrepancy.The table lists N=50,044 for usual care, but the text states 50,043.
- MINORclarityResults, paragraph 2“The adjusted DALY difference between E-MOTIVE and usual care per patient was −0.00266 (95% CI −0.00814 to 0.00287; Table ).”→ Clarify that the negative value indicates a reduction in DALYs (benefit).The sign of the difference could be misinterpreted.
- MINORotherMethods, Statistical analysis“We included fixed effects for allocated exposure to E-MOTIVE, time period, country and covariates used in the randomization method (number of vaginal births per hospital, the proportion of patients with a clinical primary-outcome event at each hospital, and the quality of oxytocin at each hospital during the baseline phase).”→ Consider breaking this long sentence into shorter ones for readability.Long sentence with multiple clauses.
- MINORtypoAbstract“disability-adjusted life-year averted”→ Consider hyphenating 'disability-adjusted life-year' consistently; elsewhere it is 'DALY'.Minor inconsistency in hyphenation.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (blinding, power analysis details, N discrepancy) and the dead link in the data availability statement. No erratum is warranted for the N discrepancy, but a correction to the table or text would improve consistency.
- 1.HIGHreportingReconcile the usual-care group N in Table 1 (50,044) with the text (50,043) and correct the discrepancy.An internal contradiction in a key denominator undermines trust in the reported results.
- 2.HIGHreportingAdd a statement on blinding or rationale for lack thereof in the Methods section.Blinding is not described, and a brief note would address a common reviewer concern for cluster trials.
- 3.HIGHreportingInclude a brief note on the power analysis parameters (effect size, alpha, power) or reference the clinical outcomes paper for details.Power analysis is referenced but not detailed, which is a minor reporting gap for a trial-based analysis.
- 4.MEDIUMdata codeFix the dead link in the data availability statement (likely the ClinicalTrials.gov URL) to ensure all links are live.A broken link undermines the accessibility of the data and code sharing.
- 5.MEDIUMreportingClarify in the Results that a negative DALY difference indicates a reduction in DALYs (benefit).The sign of the difference could be misinterpreted by readers.
- 6.MEDIUMreportingExplicitly state adherence to a recognized ethical framework (e.g., Declaration of Helsinki) in the Ethical approval section.Reviewer 2 noted regulatory compliance is implied but not explicitly stated; adding this would strengthen the ethics reporting.
- 7.MEDIUMreportingConsider reporting exact p-values for the adjusted differences, even though the estimation-based approach is acceptable.Exact p-values would facilitate meta-analyses and align with some reporting standards.
- 8.MEDIUMreportingAdd a statement clarifying that the economic evaluation is a secondary analysis and that power was determined by the clinical trial.This avoids confusion about sample size justification.
- 9.MEDIUMdata codeSpecify the exact data sharing agreement terms and any conditions for access beyond the review period.The current statement is adequate but could be more detailed for transparency.
- 10.MEDIUMstatisticsProvide a more detailed description of the multiple imputation model, including the number of imputations and variables used.Enhances reproducibility of the sensitivity analysis.
- 11.MEDIUMreportingIn the limitations, explicitly mention the lack of formal implementation cost quantification as a potential source of underestimation of total costs.Addresses a potential source of bias in the cost estimates.
- 12.LOWreportingConsider adding a supplementary table with the full list of unit costs and their sources for transparency.Improves transparency of the cost inputs.
- 13.LOWstatisticsIn the methods, clarify the handling of missing data for the DALY calculations, as the complete case analysis may introduce bias.Clarifies a potential source of bias in the DALY estimates.
- 14.LOWreportingConsider reporting the results of the sensitivity analyses in the main text, not just in supplementary materials.Highlights robustness of the findings.
- 15.LOWreportingAdd a statement about the generalizability of the findings to other settings, given the specific countries and hospital types included.Helps readers interpret the applicability of the results.
- 16.LOWstatisticsProvide a more detailed description of the permutation test procedure, including the exact algorithm and number of replications.Aids replication of the statistical analysis.
- 17.LOWcopyeditBreak the long sentence in Methods, Statistical analysis into shorter sentences for readability.Improves clarity of the statistical methods description.
- 18.LOWcopyeditEnsure consistent hyphenation of 'disability-adjusted life-year' in the Abstract.Minor consistency issue in terminology.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.