Effectiveness of kangaroo mother care before clinical stabilisation versus standard care among neonates at five hospitals in Uganda (OMWaNA): a parallel-group, individually randomised controlled trial and economic evaluation.
Tumukunde V, Medvedev MM, Tann CJ, Mambule I, Pitt C, Opondo C, Kakande A, Canter R, Haroon Y, Kirabo-Nagemi C, Abaasa A, Okot W, Katongole F, Ssenyonga R, Niombi N, Nanyunja C, Elbourne D, Greco G, Ekirapa-Kiracho E, Nyirenda M, Allen E, Waiswa P, Lawn JE, OMWaNA Collaborative Authorship Group
- DOI
- 10.1016/S0140-6736(24)00064-3
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/10d39c70-617b-4252-80fb-cbf232e5fde6 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ReportingEthical approvals partially met−0.25★
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed and transparently reported randomised controlled trial: randomisation, concealment, power calculation, participant/maternal baseline data, ethics approvals, trial registration, and reporting guidelines are all adequately documented, and the verification components (6/6 statistics consistent, no retracted/unfound citations, all links live) found no substantive problems. The main weaknesses are reporting-level: threshold-only p-values in a few outcome tables, a Table 2 footnote that labels model-derived standard errors as 'mean (SD)', and the absence of an explicit Declaration of Helsinki/ICH-GCP compliance statement.
All three independent reviewer runs classified the paper as interventional and converged on passes for six of seven applicable dimensions; the two divergences (statistical analysis warn vs pass; data code availability warn vs pass) were resolved against the explicit scoring rules, with both positions preserved in the dimension details. key resources was excluded as not applicable (procedural trial with no investigational product). The statistics verification recomputed only 6 of the reported analyses (those with a test statistic + df or an effect estimate + CI); the remaining analyses are unverified, not confirmed.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 6 tests: 6 consistent, 0 inconsistent; 3 recomputed directly from the reported test statistics, 3 via agent-written checks.
- CONSISTENTreported p = .850 · recomputed p = .828Recomputed RR 0.97 (95% CI 0.74–1.28), reported p=0.85
“RR 0.97 [95% CI 0.74–1.28]; p=0.85”
Taken as given: 0.74–1.28 is a two-sided 95% confidence interval for the RR of 0.97, not a range, an IQR, or a different interval level; the RR is a RATIO measure, so the interval is symmetric on the log scale; p=0.85 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.97, 0.74, 1.28, 1) - CONSISTENTreported p = .072 · recomputed p = .068Recomputed RR 0.60 (95% CI 0.35–1.05), reported p=0.072
“RR 0.60 [95% CI 0.35–1.05]; p=0.072”
Taken as given: 0.35–1.05 is a two-sided 95% confidence interval for the RR of 0.60, not a range, an IQR, or a different interval level; the RR is a RATIO measure, so the interval is symmetric on the log scale; p=0.072 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.6, 0.35, 1.05, 1) - CONSISTENTreported p = .043 · recomputed p = .050Recomputed RR 0.86 (95% CI 0.74–1.00), reported p=0.043
“RR 0.86 [95% CI 0.74–1.00]; p=0.043”
Taken as given: 0.74–1.00 is a two-sided 95% confidence interval for the RR of 0.86, not a range, an IQR, or a different interval level; the RR is a RATIO measure, so the interval is symmetric on the log scale; p=0.043 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.86, 0.74, 1, 1) - CONSISTENTreported p = .960 · recomputed p = .963Reviewers 1, 2Crude 7-day mortality comparison
“81 (7·5%) of 1083 neonates in the intervention group and 83 (7·5%) of 1102 in the control group died (adjusted RR 0·97 [95% CI 0·74–1·28]; p=0·85).”
Taken as given: The crude 2x2 table counts are 81 events/1002 non-events for intervention and 83 events/1019 non-events for control.; The crude p-value reported is 0.96, which is consistent with the chi-square test for independence.Method: Pearson chi-square test from 2x2 cell countsHow we recomputed it: pChi2x2(81,1002,83,1019) - CONSISTENTreported p = .307 · recomputed p = .307Reviewers 1, 2Crude 28-day mortality comparison
“119 (11·3%) of 1051 neonates in the intervention group and 134 (12·8%) of 1049 in the control group died (RR 0·88 [0·71–1·09]; p=0·229).”
Taken as given: The crude 2x2 table counts are 119 events/932 non-events for intervention and 134 events/915 non-events for control.; The crude p-value reported is 0.307, which is consistent with the chi-square test for independence.Method: Pearson chi-square test from 2x2 cell countsHow we recomputed it: pChi2x2(119,932,134,915) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Crude chi-square test for hypothermia at 24 h (secondary outcome).
“The proportion of neonates with hypothermia at 24 h was 40·9% (448 of 1096) in the intervention group and 53·1% (585 of 1101) in the control group (adjusted RR 0·76 [95% CI 0·70–0·83]; p<0·0001).”
Taken as given: Crude chi-square test approximates the p-value; the adjusted p-value is <0.0001.; Cell counts are 448, 648, 585, 516.Method: pChi2x2 from 2×2 table.How we recomputed it: pChi2x2(448, 1096-448, 585, 1101-585)
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
10 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewer 3KMC initiated before stabilisation decreased the economic cost of neonatal care to society and providers.The mean cost differences were small and not statistically significant (societal –$7·2, p=0·33; provider –$9·4, p=0·058), so the claim that costs 'decreased' is stronger than the presented evidence.Evidence: Mean economic cost to society: US$359·1 vs $365·9 (adjusted mean difference –$7·2 [–21·5 to 7·2]; p=0·33); provider costs –$9·4 [–19·0 to 0·3]; p=0·058.
“Our economic evaluation found that, compared with standard care, KMC initiated before stabilisation decreased the economic cost of neonatal care to society and providers.”
Discussion ¶1Find in source - supportedReviewers 1, 2KMC initiated before stabilisation did not reduce early neonatal mortality (7 days).The primary outcome shows no statistically significant difference between groups (adjusted RR 0.97, 95% CI 0.74–1.28, p=0.85).Evidence: Primary outcome in Table 2: adjusted RR 0.97 (95% CI 0.74–1.28), p=0.85.
“From randomisation to 7 days of age, 81 (7·5%) of 1083 neonates in the intervention group and 83 (7·5%) of 1102 in the control group died (adjusted RR 0·97 [95% CI 0·74–1·28]; p=0·85).”
Table 2Find in source - supportedReviewers 1, 2KMC was cost-effective from the societal and provider perspectives.The economic evaluation shows that the intervention is cost-effective, with a high probability of being cost-saving even when no value is placed on averted deaths.Evidence: Economic results: adjusted mean societal cost difference –$7.2 (95% CI –21.5 to 7.2), 97% probability of cost-effectiveness from provider perspective.
“Even if policy makers place no value on averting neonatal deaths, the intervention would have 97% probability from the provider perspective and 84% probability from the societal perspective of being more cost-effective than standard care.”
ResultsFind in source - supportedReviewers 1, 2Secondary outcomes including hypothermia and daily weight gain were significantly improved in the intervention group.Hypothermia at 24h was significantly reduced (RR 0.76, p<0.0001) and daily weight gain at 28 days was significantly higher (mean difference 0.75 g, p=0.047).Evidence: Table 2: adjusted RR for hypothermia 0.76 (95% CI 0.70–0.83), p<0.0001; adjusted mean difference for daily weight gain 0.75 (95% CI 0.01–1.49), p=0.047.
The proportion of neonates with hypothermia at 24 h was 40·9% (448 of 1096) in the intervention group and 53·1% (585 of 1101) in the control group (adjusted RR 0·76 [95% CI 0·70–0·83]; p<0·0001). ... Mean weight gain at 28 days was 7·8 g per day (SD 0·3) in the intervention group and 7·1 g per day (0·3) in the control group (adjusted mean difference 0·75 [0·01–1·49]; p=0·047).
Table 2reviewer’s wording - supportedReviewers 1, 2A pooled meta-analysis showed a significant relative reduction in 28-day mortality of 19% overall.The meta-analysis of three trials (OMWaNA, iKMC, eKMC) shows a significant 19% reduction (95% CI 0.71–0.93, p=0.0019).Evidence: Results, meta-analysis: 'A pooled meta-analysis... showed a relative reduction of 19% in 28-day mortality (95% CI 0·71–0·93; p=0·0019).'
A pooled meta-analysis of the OMWaNA trial, WHO’s iKMC trial, and the eKMC trial in The Gambia showed a relative reduction of 19% in 28-day mortality (95% CI 0·71–0·93; p=0·0019).
Resultsreviewer’s wording - supportedReviewer 2KMC before stabilisation decreased the economic cost of neonatal care to society and providers.The economic results show lower mean costs in the intervention group, though not statistically significant at conventional levels.Evidence: Results: provider economic costs –$9.4 (–19.0 to 0.3, p=0.058); societal economic costs –$7.2 (–21.5 to 7.2, p=0.33).
Across all five hospitals in the OMWaNA trial, the mean economic cost to society per neonate (n=2221) was similar between the intervention group (US$359·1) and the control group ($365·9; adjusted mean difference –$7·2 [95% CI –21·5 to 7·2]; p=0·33).
Results ¶6reviewer’s wording - supportedReviewers 2, 3KMC initiated before stabilisation did not reduce early neonatal mortality.The primary outcome was null (RR 0·97, 95% CI 0·74–1·28, p=0·85), directly supporting the claim.Evidence: Primary outcome: 81/1083 (7.5%) vs 83/1102 (7.5%) deaths at 7 days, adjusted RR 0·97 (0·74–1·28), p=0·85.
“KMC initiated before stabilisation did not reduce early neonatal mortality; however, it was cost-effective from the societal and provider perspectives compared with standard care.”
AbstractFind in source - supportedReviewers 2, 3KMC initiated before stabilisation was cost-effective from the societal and provider perspectives compared with standard care.The cost-effectiveness analysis reports high probabilities of cost-effectiveness (97% provider, 84% societal) even at zero value per death averted, supporting the claim.Evidence: Incremental net benefit analysis: 97% probability (provider) and 84% probability (societal) of being more cost-effective than standard care.
“Even if policy makers place no value on averting neonatal deaths, the intervention would have 97% probability from the provider perspective and 84% probability from the societal perspective of being more cost-effective than standard care.”
ResultsFind in source - supportedReviewer 3A pooled meta-analysis of the three trials showed a significant relative reduction in 28-day mortality of 19% overall and 14% across African sites.The paper reports its own pooled analyses with CIs and p-values (overall RR 0·81 equivalent, 19% reduction; African-only 14% reduction, p=0·043), supporting the claim.Evidence: Pooled meta-analysis: 19% relative reduction in 28-day mortality (95% CI 0·71–0·93; p=0·0019); African sites only 14% reduction (0·74–1·00; p=0·043).
A pooled meta-analysis of the OMWaNA trial, WHO's iKMC trial, and the eKMC trial in The Gambia showed a relative reduction of 19% in 28-day mortality (95% CI 0·71–0·93; p=0·0019)... A meta-analysis including African sites only across the three trials showed a relative reduction of 14% in 28-day mortality (0·74–1·00; p=0·043).
Results ¶4reviewer’s wording - supportedReviewer 3Secondary outcomes, including hypothermia at 24 h and daily weight gain at 28 days, were significantly improved in the intervention group.Hypothermia at 24 h (RR 0·76, p<0·0001) and daily weight gain (mean difference 0·75 g/day, p=0·047) are significantly improved, supporting the claim.Evidence: Hypothermia at 24 h: 40·9% vs 53·1% (adjusted RR 0·76 [0·70–0·83]; p<0·0001); daily weight gain: 7·8 vs 7·1 g/day (adjusted mean difference 0·75 [0·01–1·49]; p=0·047).
“Secondary outcomes, including hypothermia at 24 h and daily weight gain at 28 days, were significantly improved among neonates in the intervention group.”
AbstractFind in source
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Methods and results do not matchAssessed
1 integrity concern flagged (0 high).
- lowmethod result mismatchTable 2 labels parenthetical values for model-derived continuous outcomes as 'mean (SD)', but values such as daily weight gain 7·8 (0·3) across 761 neonates and duration of hospital admission 7·3 (0·2) are implausibly small as standard deviations of raw data and are most likely standard errors from regression models.
Duration of hospital admission, days | 7·3 (0·2); 1083 (97·6%) | 6·1 (0·1); 1102 (99·2%) ... Daily weight gain at 28 days, g per day | 7·8 (0·3); 761 (68·6%) | 7·1 (0·3); 731 (65·8%) | 0·78 (0·02 to 1·55) | p=0·045 | 0·75 (0·01 to 1·49) | p=0·047
Table 2reviewer’s wording
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Ethics/consent reporting incompleteAssessed
The introduction cites the iKMC trial, eKMC trial, and WHO guidelines, discusses their limitations (e.g., underpowered in The Gambia, lack of economic data), and clearly states the rationale for the study in sub-Saharan Africa level 2 facilities. The limitations of prior work are addressed in terms of context and economic evaluation.
“We aimed to compare the effectiveness and safety of KMC initiated before stabilisation versus standard care among neonates ≤2000 g in sub-Saharan Africa. Additionally, we aimed to assess the incremental costs and cost-effectiveness of this intervention from the societal perspective.”
“The findings of these two trials left an evidence gap regarding the effects of immediate KMC at facilities without neonatal intensive care, where most births in LMICs occur, and questions regarding possibly lower impact in sub-Saharan Africa.”
“However, resource requirements were not well defined, including the incremental costs of initiating KMC before stabilisation and necessary facility infrastructure. Indeed, there are few economic evaluations of KMC overall.”
“The findings of these two trials left an evidence gap regarding the effects of immediate KMC at facilities without neonatal intensive care, where most births in LMICs occur, and questions regarding possibly lower impact in sub-Saharan Africa.”
Randomization was computer-generated with permuted blocks of varying sizes, stratified by birthweight and site. The independent statistician was masked. A power analysis was conducted assuming 25% baseline mortality and 80% power to detect a 5.6% absolute reduction. Inclusion/exclusion criteria are detailed. Missing data were handled by complete-case analysis with sensitivity analyses.
“A random allocation sequence was computer-generated with permuted blocks of varying sizes, stratified by birthweight (700–1000 g, 1000–1499 g, or 1500–2000 g) and recruitment site.”
“A sample size of 2188 neonates (1094 per group) was estimated to be required to detect an absolute reduction in 7-day mortality of 5·6% (22·4% relative reduction) at a 5% significance level (two-sided) and 80% power, allowing for 20% attrition.”
“A random allocation sequence was computer-generated with permuted blocks of varying sizes, stratified by birthweight (700–1000 g, 1000–1499 g, or 1500–2000 g) and recruitment site. Allocation concealment was done by programming the allocation sequence into the screening database and revealing treatment group only when screening of eligible neonates was complete.”
“A sample size of 2188 neonates (1094 per group) was estimated to be required to detect an absolute reduction in 7-day mortality of 5·6% (22·4% relative reduction) at a 5% significance level (two-sided) and 80% power, allowing for 20% attrition.”
“Masking of parents, caregivers, or health-care workers was not possible due to the nature of the KMC intervention; however, the independent statistician who conducted the analyses was masked to treatment allocation.”
“A random allocation sequence was computer-generated with permuted blocks of varying sizes, stratified by birthweight (700–1000 g, 1000–1499 g, or 1500–2000 g) and recruitment site.”
“Masking of parents, caregivers, or health-care workers was not possible due to the nature of the KMC intervention; however, the independent statistician who conducted the analyses was masked to treatment allocation.”
“A sample size of 2188 neonates (1094 per group) was estimated to be required to detect an absolute reduction in 7-day mortality of 5·6% (22·4% relative reduction) at a 5% significance level (two-sided) and 80% power, allowing for 20% attrition.”
Table 1 provides baseline characteristics by group, including sex (approximately 50% female), mean age at screening, gestational age, birthweight, mode of delivery, and maternal age, education, and income. Both sexes are enrolled, so sex justification is not required.
“Female | 558 (50·3%) | 561 (50·5%) | | Male | 552 (49·7%) | 550 (49·5%)”
The paper reports approval by the Uganda Virus Research Institute, Uganda National Council of Science and Technology, and LSHTM ethics committees with protocol numbers. Written informed consent was obtained. However, no statement of adherence to the Declaration of Helsinki or ICH-GCP is provided, which is expected for a clinical trial.
“The study was approved by the Research Ethics Committees of Uganda Virus Research Institute (GC/127/19/06/717), Uganda National Council of Science and Technology (HS 2645), and London School of Hygiene & Tropical Medicine (16972).”
“Written informed parental consent was obtained for all participants.”
“The study was approved by the Research Ethics Committees of Uganda Virus Research Institute (GC/127/19/06/717), Uganda National Council of Science and Technology (HS 2645), and London School of Hygiene & Tropical Medicine (16972).”
“Written informed parental consent was obtained for all participants.”
“The study was approved by the Research Ethics Committees of Uganda Virus Research Institute (GC/127/19/06/717), Uganda National Council of Science and Technology (HS 2645), and London School of Hygiene & Tropical Medicine (16972).”
“Written informed parental consent was obtained for all participants.”
The intervention is kangaroo mother care (skin-to-skin contact, wrap) and standard care (incubator/radiant heater). The paper does not identify specific brands or manufacturers of these devices. Since this is a behavioral/procedural trial, the dimension is not applicable per the scoring guidelines.
“Neonates were naked, except for hat and diaper; placed prone and skin-to-skin on the caregiver’s chest; and secured using a KMC wrap.”
“All statistical analyses were performed with Stata (version 18.1).”
All statistical tests are named (modified Poisson, Cox, linear regression). Effect sizes are reported with 95% CIs. Statistical software (Stata 18.1) is identified. Data presentation includes tables with n, percentages, means, SD, IQR, and a CONSORT flow diagram. Assumptions are handled by robust standard errors and standard methods. Some p-values are reported as '<0.0001' which is not exact, but this is a minor issue.
“Risk ratios (RRs) with 95% CI were estimated for 7-day and 28-day mortality and were compared between the intervention group and the control group with modified Poisson regression models. Modified Poisson regression was preferred over binomial regression for obtaining RRs since binomial regression is prone to non-convergence.”
“From randomisation to 7 days of age, 81 (7·5%) of 1083 neonates in the intervention group and 83 (7·5%) of 1102 in the control group died (adjusted RR 0·97 [95% CI 0·74–1·28]; p=0·85).”
“All analyses were performed with Microsoft Excel and Stata (version 18.1).”
The paper states: 'Data from the study will be deposited online at LSHTM Data Compass (https://datacompass.lshtm.ac.uk/), with access subject to approval.' This is a concrete route with a named repository and conditions (managed access). For a clinical trial with patient-level data, repository deposit with a persistent identifier is not expected due to privacy; controlled access satisfies the field norm. Code sharing is not required because the analysis used standard Stata commands. Thus the single applicable criterion is adequate.
“Data from the study will be deposited online at LSHTM Data Compass ( https://datacompass.lshtm.ac.uk/ ), with access subject to approval.”
Methods are complete for replication. Trial registration number is provided. CONSORT guideline is referenced. All pre-specified outcomes are reported in results. Limitations are discussed in detail. Conclusions are proportional to the evidence. Funding sources and conflicts of interest are declared.
“Our sample size would have been sufficient to detect an absolute reduction of 4·8% (relative reduction of 26·7%), even with an 18% mortality rate. However, we observed lower reductions in mortality rate than were expected from baseline data in both the control and intervention groups, which reduced our power to detect the prespecified relative reduction of 5·6%.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 44 references by DOI: 37 verified — 7 no DOI (shown, not verified).
- NO DOILevels & trends in child mortality: report 2023No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBorn too soon: decade of action on preterm birthNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO recommendations for care of the preterm or low-birth-weight infantNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe cost-savings of implementing kangaroo mother care in NicaraguaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIKangaroo mother care: a practical guideNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOfficial exchange rate (LCU per US$, period average)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIKangaroo mother care: implementation strategy for scale-up adaptable to different country contextsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
3 data/code links checked; 3 live.
- datahttps://datacompass.lshtm.ac.uk/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttp://Clinicaltrials.govLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT02811432LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
3 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 3 minor suggestions below.
3 copyedit issues flagged: mostly consistency, punctuation.
- MINORpunctuationSummary, Findings“adjusted relative risk [RR] 0·97 [95% CI 0·74–1·28]; p=0·85.”→ Add closing parenthesis: 'p=0·85).'Unbalanced parenthesis around the adjusted RR in the abstract.
- MINORconsistencyTable 2 footnote and Results text“Data are n/N (%), n (%), median (IQR), or mean (SD).”→ Clarify whether the parenthetical values for model-derived outcomes (e.g., daily weight gain 7·8 (0·3); duration of admission 7·3 (0·2)) are SDs or standard errors.Values such as 7·8 (0·3) and 7·3 (0·2) are implausibly small as raw-data SDs and are almost certainly standard errors from regression models; labelling them 'mean (SD)' is misleading.
- MINORconsistencySummary vs Table 2“(RR 0·88 [0·71–1·09]; p=0·229)”→ Reconcile p=0·229 (abstract) with p=0·23 (Table 2) and the CI style ([...] vs (...)).Same estimate is reported with slightly different rounding and bracket style in abstract versus table.
The published trial is methodologically robust and its headline conclusions (no early-mortality benefit; cost-effectiveness) are supported by the reported analyses and the verification checks performed. An informed reader should weigh three minor reporting gaps — the Table 2 'mean (SD)' labelling of model-derived standard errors (which warrants a correction/clarification), the threshold-only p-values, and the absent compliance statement — none of which overturn the findings or, on the evidence available, necessitate independent re-analysis.
- 1.HIGHstatisticsCorrect the Table 2 footnote so that parenthetical values for model-derived continuous outcomes (e.g., daily weight gain 7·8 [0·3]; duration of admission 7·3 [0·2]) are identified as standard errors from the regression models, not as 'mean (SD)' of raw data.The integrity check flagged these values as implausibly small to be raw-data SDs — a method–result mismatch that could mislead readers and warrants a published correction.
- 2.MEDIUMcopyeditReconcile the 28-day mortality estimate across the abstract (RR 0·88 [0·71–1·09]; p=0·229) and Table 2 (p=0·23) and standardise the bracket style ([...] vs (...)).The same estimate is reported with inconsistent rounding and bracket formatting in the abstract versus the table.
- 3.MEDIUMcopyeditFix the unbalanced parenthesis in the abstract findings by adding the closing ')' after 'p=0·85'.The copyedit pass flagged an unbalanced parenthesis around the adjusted RR in the abstract.
- 4.MEDIUMethicsAdd an explicit statement of compliance with the Declaration of Helsinki (or ICH-GCP) to the Methods/ethics section.All three reviewers flagged the absent regulatory-compliance statement as the only ethics reporting gap; the named approvals and consent documentation otherwise satisfy the criterion.
- 5.MEDIUMstatisticsReport exact p-values for outcomes currently printed only as thresholds (e.g., replace 'p<0·0001' with the actual value in Table 2).Threshold-only p-values constitute imprecise reporting of statistics the authors did compute, and exact values improve reproducibility.
- 6.LOWdata codeAdd a persistent identifier (DOI) for the LSHTM Data Compass deposit once available, and state when the dataset will be accessible.The data-availability statement is concrete but prospective ('will be deposited'); a DOI strengthens the commitment to reproducibility.
- 7.LOWdata codeAdd a brief code-availability note for the Stata do-files used in the analysis.Even where the code uses standard Stata commands, sharing or explicitly declining to share the analysis code enhances transparency.
- 8.LOWstatisticsState whether the proportional-hazards assumption was tested for the Cox time-to-event analyses (e.g., test for non-proportionality).Reviewer 2 noted the assumption is not explicitly verified for the Cox models reporting hazard ratios.
- 9.LOWstatisticsClarify the missing-data handling for the primary analysis (e.g., whether complete-case analysis was used for the primary outcome and multiple imputation only for costs).Reviewer 1 asked for explicit specification of how missing data were handled for the primary endpoint versus the economic analysis.
- 10.LOWotherProvide the rationale for the 5·6% absolute-reduction target chosen for the sample-size calculation as clinically meaningful.Reviewer 1 suggested justifying why the specified effect size was clinically meaningful, which would aid readers interpreting the power analysis.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.