Effects of Cooking with Liquefied Petroleum Gas or Biomass on Stunting in Infants.
Checkley W, Thompson LM, Sinharoy SS, Hossen S, Moulton LH, Chang HH, Waller L, Steenland K, Rosa G, Mukeshimana A, Ndagijimana F, McCracken JP, Díaz-Artiga A, Balakrishnan K, Garg SS, Thangavel G, Aravindalochanan V, Hartinger SM, Chiang M, Kirby MA, Papageorghiou AT, Ramakrishnan U, Williams KN, Nicolaou L, Johnson M, Pillarisetti A, Rosenthal J, Underhill LJ, Wang J, Jabbarzadeh S, Chen Y, Dávila-Román VG, Naeher LP, McCollum ED, Peel JL, Clasen TF, HAPIN Investigators
- DOI
- 10.1056/NEJMoa2302687
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/229e5dde-45f7-4129-8f19-cb36991c601a is authoritative.
How this rating was calculated
- StatisticsImpossible or misreported statistic ×3−3★
- ReportingData & code availability partially met−0.25★
- CitationsUnresolved reference−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) covered 23 of 26 reported means. The other 3 do not state their group size where the value is printed, and these checks need the count the mean was averaged over, so they were not checked.
- No data or code availability links were detected to verify.
- 01No integer data can produce this mean and SDdemonstrable
SD 1.1 exceeds the maximum possible (0.19) for n=1186 values on a [-0.22, 0.21] scale with mean -0.1
“−0.1±1.1”
Table 1 - 02No integer data can produce this mean and SDdemonstrable
SD 1 exceeds the maximum possible (0.22) for n=365 values on a [-0.22, 0.21] scale with mean 0
“0.0±1.0”
Table 1 - 03No integer data can produce this mean and SDdemonstrable
SD 1 exceeds the maximum possible (0.19) for n=339 values on a [-0.22, 0.21] scale with mean -0.1
“−0.1±1.0”
Table 1
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted RCT with a clear scientific premise, rigorous design, and transparent reporting. The main weaknesses are the lack of detailed ethics committee names and informed consent description in the main text, and the vague data availability statement. Minor reporting gaps include unspecified LPG stove model and lack of explicit CONSORT adherence.
Both reviewers agreed on study type (interventional). The synthesis weighted specific evidence; where reviewers diverged (e.g., randomization method detail, ethics reporting), the more conservative rating was adopted when the evidence supported a reporting gap. The statistics verification covered only a subset of tests; the 3 inconsistent recomputations were not detailed and may be due to rounding, so they do not change the pass status but warrant caution.
Numerical inconsistencies
1 finding · worst criticalValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks. 3 reported summary statistics mathematically impossible for the stated N (SPRITE).
- SPRITESD 1.1 exceeds the maximum possible (0.19) for n=1186 values on a [-0.22, 0.21] scale with mean -0.1
“−0.1±1.1”
Table 1 - SPRITESD 1 exceeds the maximum possible (0.22) for n=365 values on a [-0.22, 0.21] scale with mean 0
“0.0±1.0”
Table 1 - SPRITESD 1 exceeds the maximum possible (0.19) for n=339 values on a [-0.22, 0.21] scale with mean -0.1
“−0.1±1.0”
Table 1
- CONSISTENTreported p = .120 · recomputed p = .238Reviewers 1, 2Primary outcome relative risk p-value from CI
“relative risk, 1.10; 98.75% confidence interval, 0.94 to 1.29; P = 0.12”
Taken as given: The CI is a 98.75% confidence interval for the relative risk.; The p-value is two-sided and corresponds to the test of RR=1.Method: Recomputed p-value from the reported RR and CI using the pCI function with log=1.How we recomputed it: pCI(1.10, 0.94, 1.29, 1)
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
4 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2The LPG cookstove intervention did not reduce the risk of stunting in infants.The primary outcome analysis shows no significant difference between groups, supporting the claim.Evidence: Primary outcome: RR 1.10, 98.75% CI 0.94-1.29, P=0.12.
“An intervention strategy starting in pregnancy and aimed at mitigating household air pollution by replacing biomass fuel with LPG for cooking did not reduce the risk of stunting in infants.”
ConclusionFind in source - supportedReviewers 1, 2The intervention reduced personal exposures to fine particulate matter.The reported exposure measurements show substantial reductions in the intervention group.Evidence: Mean prenatal exposure 35.0 vs 103.3 μg/m3; postnatal 37.9 vs 109.2 μg/m3.
the intervention resulted in lower prenatal and postnatal 24-hour personal exposures to fine particulate matter than the control (mean prenatal exposure, 35.0 μg per cubic meter vs. 103.3 μg per cubic meter; mean postnatal exposure, 37.9 μg per cubic meter vs. 109.2 μg per cubic meter).
Abstractreviewer’s wording - supportedReviewers 1, 2The intervention did not reduce severe stunting at 12 months.The secondary outcome of severe stunting showed a relative risk of 1.36 with 95% CI 1.02-1.82, which is not statistically significant at the 5% level (though the CI excludes 1, the p-value is not reported; the authors note it was not adjusted for multiplicity).Evidence: Severe stunting: 8.1% vs 6.1%, RR 1.36 (95% CI 1.02-1.82).
The percentage of infants with severe stunting at 12 months of age was 8.1% in the intervention group and 6.1% in the control group (relative risk, 1.36; 95% CI, 1.02 to 1.82).
Resultsreviewer’s wording - supportedReviewer 2Adherence to the intervention was high.The median percentage of days using biomass stove was 0.4%, indicating high adherence.Evidence: Median percentage of monitored days using biomass stove was 0.4% (IQR 0-2.3).
“the median percentage of monitored days that intervention households used their biomass cookstove rather than the LPG cookstove during the trial period was 0.4% (interquartile range, 0 to 2.3).”
ResultsFind in source
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites multiple studies and meta-analyses linking household air pollution to stunting, and the rationale for the trial is well-articulated. The paper also discusses limitations of prior observational studies and the need for a randomized trial.
“A meta-analysis of 11 studies showed a 19% higher risk of stunting among children younger than 5 years of age who were exposed to household air pollution than among those who were not exposed.”
“The Household Air Pollution Intervention Network (HAPIN) trial was designed to assess the effects of replacing biomass cookstoves with liquefied petroleum gas (LPG) cookstoves on four primary outcomes, including stunting in infants.”
“Associations between exposures to household air pollution and stunting have been reported in observational studies.”
“Associations between exposures to household air pollution and stunting have been reported in observational studies. For example, a review of two studies showed a 27% higher risk of stunting among children younger than 5 years of age who were exposed to household air pollution than among those who were not exposed.”
“The Household Air Pollution Intervention Network (HAPIN) trial was designed to assess the effects of replacing biomass cookstoves with liquefied petroleum gas (LPG) cookstoves on four primary outcomes, including stunting in infants.”
“Associations between exposures to household air pollution and stunting have been reported in observational studies.”
Randomization was stratified by geographic region, and the method is described. The trial is open-label, but the outcome (length measurement) is objective and measured by trained personnel. Power analysis is provided. Inclusion/exclusion criteria are clear. Missing data handling is described.
“Randomization was stratified according to geographic region, of which there were 10 in the trial”
“We estimated that 1440 participants per trial group would be needed to provide the trial with 80% power to detect a relative risk of stunting of 0.81 favoring the intervention, with a baseline incidence of stunting of 30% among participants in the control group, at an alpha level of 0.0125”
“They measured the infants in the homes where the infants resided and thus were aware of the trial-group assignments.”
“Randomization was stratified according to geographic region, of which there were 10 in the trial”
“We estimated that 1440 participants per trial group would be needed to provide the trial with 80% power to detect a relative risk of stunting of 0.81 favoring the intervention, with a baseline incidence of stunting of 30% among participants in the control group, at an alpha level of 0.0125”
“They measured the infants in the homes where the infants resided and thus were aware of the trial-group assignments.”
The paper reports infant sex, maternal age, gestational age, and other health-related variables in Table 1. Demographics are well described. Since this is a human trial, species/strain and housing conditions are not applicable.
“Male sex — no. (%) | 608 (51.9) | 605 (51.0) | 192 (52.6) | 182 (53.7)”
“Maternal age — yr | 25.4±4.4 | 25.5±4.5 | 25.2±4.4 | 25.3±4.7”
“Male sex — no. (%) | 608 (51.9) | 605 (51.0) | 192 (52.6) | 182 (53.7)”
“Maternal age — yr | 25.4±4.4 | 25.5±4.5 | 25.2±4.4 | 25.3±4.7”
“Maternal education — yr | 8.2±3.8 | 7.9±3.6 | 8.3±3.7 | 8.3±3.5”
The paper states that the trial was approved by applicable ethics committees and mentions informed consent in the supplementary appendix. Regulatory compliance is implied through adherence to ethical standards.
“The trial was approved by the applicable ethics committees; details are provided in the , available with the full text of this article at NEJM.org”
The LPG stove is not explicitly named with a manufacturer, but the intervention is described. Monitoring devices are identified (Enhanced Children's MicroPEM, Lascar EL-USB-CO300). Statistical software is identified (R and Stata). No antibodies, cell lines, or mycoplasma testing are applicable.
“We performed the statistical analyses using R software, version 4.1.2 (R Project for Statistical Computing), and Stata SE software, version 15.1 (StataCorp).”
“using the Enhanced Children’s MicroPEM monitor (RTI International) and to carbon monoxide using the Lascar EL-USB-CO300 (Lascar Electronics)”
“We performed the statistical analyses using R software, version 4.1.2 (R Project for Statistical Computing), and Stata SE software, version 15.1 (StataCorp).”
The primary analysis uses log-binomial regression with intention-to-treat. Exact p-values are reported for the primary outcome. Effect sizes with confidence intervals are provided. Software is identified. Data presentation includes per-group n and dispersion. Mathematical plausibility checks are not applicable due to large sample sizes and continuous outcomes.
“We used a log-binomial regression model, with stunting at 12 months of age as the outcome and the trial-group assignment (with the control group as the reference) as the main covariate, adjusted for randomization strata.”
“relative risk, 1.10; 98.75% confidence interval, 0.94 to 1.29; P = 0.12”
“We used a log-binomial regression model, with stunting at 12 months of age as the outcome and the trial-group assignment (with the control group as the reference) as the main covariate, adjusted for randomization strata.”
“relative risk, 1.10; 98.75% confidence interval, 0.94 to 1.29; P = 0.12”
“1.10 (0.94 to 1.29)”
The paper states that a data sharing statement is available with the full text, but the specific mechanism is not described in the manuscript. No data repository or accession numbers are given. Code sharing is not mentioned.
“A data sharing statement provided by the authors is available with the full text of this article at NEJM.org”
“A data sharing statement provided by the authors is available with the full text of this article at NEJM.org”
The trial is registered at ClinicalTrials.gov (NCT02944682). Methods are comprehensive. Limitations are explicitly discussed, including missing data and exposure measurement. Conclusions are appropriately cautious. Funding and COI are disclosed.
“First, we were unable to measure linear growth in approximately 20% of the infants because of challenges during the Covid-19 pandemic.”
“First, we were unable to measure linear growth in approximately 20% of the infants because of challenges during the Covid-19 pandemic.”
“Supported by the National Institutes of Health (cooperative agreement 1UM1HL134590) in collaboration with the Bill and Melinda Gates Foundation (OPP1131279).”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 42 references by DOI: 32 verified — 1 DOI unresolved, 9 no DOI (shown, not verified).
- UNRESOLVED10.1101/2023.07.04.23292226v1Post-birth exposure contrasts for children during the Household Air Pollution Intervention Network randomized controlled trialCited DOI does not resolve to any Crossref record.
- NO DOIEstimating disease burden attributable to household air pollution: new methods within the Global Burden of Disease studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGBD compareNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO child growth standards: length/height-for-age, weight-for-age, weight-for-length, weight-for-height and body mass index-for-age: methods and developmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIChild malnutrition estimates: key findings of the 2020 Joint Child Malnutrition Estimates: UNICEF regionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAnthropometric standardization reference manualNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIR: a language and environment for statistical computingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMinimum dietary diversity for women: a guide for measurementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO ambient air quality database, 2021 updateNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO indoor air quality guidelines: household fuel combustionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
1 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 1 minor suggestion below.
1 copyedit issue flagged: mostly consistency.
- MINORconsistencyAbstract“1171 (76.2%) of the 1536 infants born to women in the intervention group and 1186 (77.8%) of the 1525 infants born to women in the control group had a valid length measurement at 12 months of age.”→ Consider rephrasing for clarity: '1171 of 1536 infants (76.2%) in the intervention group and 1186 of 1525 (77.8%) in the control group had a valid length measurement.'The percentage placement is slightly awkward.
The published work is robust overall, but an informed reader should weigh the incomplete ethics details and vague data availability statement. These are reporting gaps that could warrant a correction or clarification from the authors, but they do not undermine the core findings.
- 1.CRITICALstatisticsCorrect or explain the statistically impossible value: SPRITE: SD 1.1 exceeds the maximum possible (0.19) for n=1186 values on a [-0.22, 0.21] scale with mean -0.1Demonstrable critical failure — blocks the verdict from passing.
- 2.CRITICALstatisticsCorrect or explain the statistically impossible value: SPRITE: SD 1 exceeds the maximum possible (0.22) for n=365 values on a [-0.22, 0.21] scale with mean 0Demonstrable critical failure — blocks the verdict from passing.
- 3.CRITICALstatisticsCorrect or explain the statistically impossible value: SPRITE: SD 1 exceeds the maximum possible (0.19) for n=339 values on a [-0.22, 0.21] scale with mean -0.1Demonstrable critical failure — blocks the verdict from passing.
- 4.HIGHethicsIn the Methods section, name the specific ethics committees that approved the trial and provide protocol numbers.The current statement is vague and does not allow verification of ethical oversight.
- 5.HIGHethicsDescribe the informed consent process (written, oral, or waiver) in the Methods section.Informed consent is not explicitly described in the main text, which is a reporting gap.
- 6.HIGHdata codeProvide a concrete data availability statement in the main text, naming a repository or managed-access platform with conditions.The current statement refers to a separate statement without specifying how to access data, which is insufficient for reproducibility.
- 7.HIGHotherVerify or correct the reference 'Post-birth exposure contrasts for children during the Household Air Pollution Intervention Network randomized controlled trial' (DOI 10.1101/2023.07.04.23292226v1), which was not found in any registry.A reference that cannot be located may be fabricated or have an incorrect DOI; it should be verified or replaced.
- 8.MEDIUMreportingExplicitly mention adherence to CONSORT reporting guidelines in the Methods or acknowledgments.The paper follows CONSORT-like structure but does not state it, which is a minor transparency gap.
- 9.MEDIUMotherIdentify the LPG stove model and manufacturer in the Methods section.The intervention product is not fully specified, which limits reproducibility.
- 10.MEDIUMdata codeConsider depositing de-identified data in a public repository with a DOI, or provide a clear data access mechanism.Enhances transparency and allows independent verification of results.
- 11.MEDIUMdata codeShare analysis code in a public repository (e.g., GitHub) with a versioned DOI.Facilitates reproducibility of the statistical analyses.
- 12.LOWcopyeditRephrase the abstract sentence about valid length measurements for clarity: '1171 of 1536 infants (76.2%) in the intervention group and 1186 of 1525 (77.8%) in the control group had a valid length measurement.'The current percentage placement is slightly awkward and could be clearer.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.