Non-invasive high frequency oscillatory ventilation for primary respiratory support in extremely preterm infants: multicentre randomised controlled trial.
Li Y, Zhu X, Li LJ, Chen L, Yang Q, Xu L, Liang W, Lin X, Li C, Xue J, Liu L, Pan X, Ju R, Peng X, Tang W, Shi Y, NHFOV study group
- DOI
- 10.1136/bmj-2025-085569
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/73cec639-26d5-4b04-ab52-691df015cd19 is authoritative.
How this rating was calculated
Started at 5★ — no deductions. Nothing the checks ran surfaced a material problem.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported multicentre RCT comparing NHFOV with NCPAP in extremely preterm infants. The paper demonstrates strong methodological rigor across all eight dimensions, with minor reporting gaps in outlier handling and statistical assumption verification.
Both reviewers classified the study as interventional, which is adopted. The evaluation covers all eight dimensions; no dimensions were excluded as not applicable. The reviewers converged closely, with minor disagreements on outlier handling and demographics reporting, which were resolved by weighing the evidence.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks.
- CONSISTENTreported p = .007 · recomputed p = .007Reviewers 1, 2Primary outcome risk difference p-value
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007)”
Taken as given: The 27 and 48 are the event counts in the NHFOV and NCPAP groups, respectively.; The 170 and 172 are the total numbers in each group.; The test used is a chi-square test for 2x2 table.Method: Pearson chi-square test on the 2x2 table.How we recomputed it: pChi2x2(27, 170-27, 48, 172-48) - CONSISTENTreported p = .008 · recomputed p = .009Reviewers 1, 2Secondary outcome: treatment failure within seven days p-value
“Treatment failure within seven days after birth was reported in 36 of 170 infants (21.2%) in the NHFOV group and 58 of 172 (33.7%) in the NCPAP group (risk difference −12.5 percentage points, 95% confidence interval −21.9 to −3.2; P=0.008)”
Taken as given: The 36 and 58 are the event counts in the NHFOV and NCPAP groups, respectively.; The 170 and 172 are the total numbers in each group.; The test used is a chi-square test for 2x2 table.Method: Pearson chi-square test on the 2x2 table.How we recomputed it: pChi2x2(36, 170-36, 58, 172-58)
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
7 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1NHFOV may mitigate the need for invasive mechanical ventilation, particularly in infants of lower gestational age or with more severe respiratory failure.The trial shows benefit in the overall population, but subgroup analyses by gestational age or severity are not reported, so this claim extends beyond the presented evidence.Evidence: Discussion comparison with other studies.
“Taken together, these data suggest that NHFOV may mitigate the need for invasive mechanical ventilation, particularly in infants of lower gestational age or with more severe respiratory failure.”
DiscussionFind in source - supportedReviewer 1NHFOV is more efficacious than NCPAP in reducing invasive mechanical ventilation within 72 hours after birth.The primary outcome shows a statistically significant reduction in treatment failure within 72 hours, with a risk difference of -12.0 percentage points (95% CI -20.7 to -3.4, P=0.007).Evidence: Primary outcome result in Results section.
“Treatment failure within 72 hours occurred in 27 of` 170 infants (15.9%) in the NHFOV group and 48 of 172 infants (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007).”
AbstractFind in source - supportedReviewer 1NHFOV is superior to NCPAP in reducing the need for intubation when used as a primary respiratory support strategy in extremely preterm infants.The primary and secondary outcomes (treatment failure within 72 hours and 7 days) both show significant reductions, supporting the conclusion.Evidence: Primary and secondary outcome results.
“NHFOV appeared superior to NCPAP in reducing the need for intubation when used as a primary respiratory support strategy in extremely preterm infants.”
ConclusionFind in source - supportedReviewers 1, 2Both techniques did not show significant differences in neonatal adverse events.Safety outcomes showed no significant differences between groups, as reported in Table 3.Evidence: Safety outcomes in Table 3.
“Both techniques did not show significant differences in neonatal adverse events.”
ConclusionFind in source - supportedReviewer 2NHFOV is more efficacious than NCPAP in reducing invasive mechanical ventilation within 72 hours.The primary outcome shows a statistically significant reduction in treatment failure within 72 hours, with a risk difference of -12.0 percentage points (95% CI -20.7 to -3.4, P=0.007).Evidence: Primary outcome result in Results section.
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007; ).”
AbstractFind in source - supportedReviewer 2NHFOV is superior to NCPAP in reducing the need for intubation within seven days.The secondary outcome of treatment failure within seven days also shows a significant reduction (risk difference -12.5 percentage points, 95% CI -21.9 to -3.2, P=0.008).Evidence: Secondary outcome result in Results section.
“Treatment failure within seven days after birth was reported in 36 of 170 infants (21.2%) in the NHFOV group and 58 of 172 (33.7%) in the NCPAP group (risk difference −12.5 percentage points, 95% confidence interval −21.9 to −3.2; P=0.008; ).”
ResultsFind in source - supportedReviewer 2NHFOV may provide dual physiological advantages over NCPAP by maintaining lung recruitment while actively clearing carbon dioxide.This is a mechanistic explanation supported by physiological reasoning and cited literature, not directly tested in this trial.Evidence: Introduction paragraph 2.
“NHFOV may provide dual physiological advantages over NCPAP by maintaining lung recruitment while actively clearing carbon dioxide through oscillatory gas mixing—a mechanism particularly beneficial for non-uniformly diseased lungs.”
IntroductionFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary outcome is treatment failure, defined as the need for invasive mechanical ventilation within 72 hours after birth. This is a hard clinical outcome (need for intubation), not a surrogate biomarker. The trial directly measures a clinically meaningful event.
“The primary outcome was respiratory support failure, defined by the need for invasive mechanical ventilation within 72 hours after birth.”
- ADEQUATEEffect sizeThe primary outcome shows a 12.0 percentage point absolute reduction in treatment failure (from 27.9% to 15.9%), which is statistically significant (P=0.007) and clinically meaningful as it represents a substantial reduction in the need for invasive ventilation in extremely preterm infants.
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The paper cites prior work on NCPAP failure rates and NHFOV's physiological advantages, and notes inconsistent findings from previous RCTs. The hypothesis follows logically from the cited evidence. Limitations of prior research (e.g., inconsistent findings, underpowered studies) are acknowledged and addressed by the trial's design.
“However, cumulative evidence has suggested that up to 40% of extremely preterm infants are prone to a high risk of NCPAP failure because of chest wall collapse and poor diaphragmatic strength.”
“We hypothesised that NHFOV would be more effective than NCPAP in reducing treatment failure within 72 hours after birth in a homogeneous population of Chinese extremely preterm infants.”
“Several randomised controlled trials comparing NHFOV and NCPAP have reported inconsistent findings about their effectiveness in managing respiratory distress syndrome.”
“Several randomised controlled trials comparing NHFOV and NCPAP have reported inconsistent findings about their effectiveness in managing respiratory distress syndrome.”
“We hypothesised that NHFOV would be more effective than NCPAP in reducing treatment failure within 72 hours after birth in a homogeneous population of Chinese extremely preterm infants.”
Randomization used a computer-generated sequence with central allocation concealment and stratification by centre. Outcome assessors were masked, though caregivers were not due to the nature of the intervention. A power analysis was performed with a stated effect size, alpha, and power. Inclusion/exclusion criteria were pre-specified. The analysis was intention-to-treat. Outlier handling is not explicitly described, but the analysis population is defined.
“Simple randomisation was performed using a computer generated random number sequence, which was securely posted on a dedicated, password protected website accessible 24/7.”
“However, outcome assessors were masked to the treatment allocation to minimise ascertainment bias.”
“With a significance level (α) of 0.05 and 90% power, a sample size of 170 newborns per group was required, leading to a total enrolment target of at least 340.”
“Simple randomisation was performed using a computer generated random number sequence, which was securely posted on a dedicated, password protected website accessible 24/7.”
“However, outcome assessors were masked to the treatment allocation to minimise ascertainment bias.”
“With a significance level (α) of 0.05 and 90% power, a sample size of 170 newborns per group was required, leading to a total enrolment target of at least 340.”
The paper reports sex, gestational age, birth weight, and multiple other baseline characteristics. Since both sexes are enrolled, sex justification is not applicable. Demographics are reported in detail. Species/strain and housing conditions are not applicable for a human trial.
“Female | 74 (43.5) | 68 (39.5)”
“Gestational age (weeks), median (IQR) | 27.0 (26.0-28.0) | 27.0 (26.0-28.0)”
“Female | 74 (43.5) | 68 (39.5)”
“Gestational age (weeks), median (IQR) | 27.0 (26.0-28.0) | 27.0 (26.0-28.0)”
The paper states approval by the Ethics Committee of the Children's Hospital of Chongqing Medical University (No 2019.161) and local IRBs. Written informed consent was obtained. Compliance with CONSORT and relevant regulations is mentioned.
“The study was approved by the ethics committee of the Children's Hospital of Chongqing Medical University (No 2019.161)”
“Informed consent was obtained from parents or guardians antenatally or upon admission to the neonatal intensive care unit”
“The study was approved by the ethics committee of the Children's Hospital of Chongqing Medical University (No 2019.161)”
“Informed consent was obtained from parents or guardians antenatally or upon admission to the neonatal intensive care unit”
The trial uses specific devices (Comen NV8, Mindray NB350, Fabian-III, SLE 5000, Leoni+) and drugs (caffeine citrate, Curosurf) with manufacturers and doses. Software (R version 4.4.2) is identified. Antibodies, cell lines, and mycoplasma testing are not applicable.
“Continuous flow devices (Comen NV8, China; Mindray NB350, China; supplementary eTable 2) were used to provide NCPAP.”
“Surfactant treatment (Curosurf, Chiesi Pharmaceuticals) was administered at a dose of 200 mg/kg”
“All statistical analyses were conducted in R software (version 4.4.2).”
“Continuous flow devices (Comen NV8, China; Mindray NB350, China; supplementary eTable 2) were used to provide NCPAP.”
“Surfactant treatment (Curosurf, Chiesi Pharmaceuticals) was administered at a dose of 200 mg/kg”
“All statistical analyses were conducted in R software (version 4.4.2).”
The paper names tests (Student's t, Mann-Whitney U, χ²) and reports risk differences with 95% CIs. Exact p-values are given for primary and secondary outcomes. Data presentation includes per-group n and percentages. Mathematical plausibility checks are not applicable due to large N and continuous outcomes.
“Comparisons were performed using Student’s t test for parametric continuous variables, Mann-Whitney U test for non-parametric continuous variables, and χ 2 test for dichotomous variables, respectively.”
“risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007”
“risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4”
“Comparisons were performed using Student’s t test for parametric continuous variables, Mann-Whitney U test for non-parametric continuous variables, and χ 2 test for dichotomous variables, respectively.”
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007; ).”
The data availability statement provides a concrete route: a Mendeley Data repository with a DOI. Code is stated to be in supplemental files. Repository deposit and accession numbers are satisfied by the Mendeley DOI.
“The data underlying the findings in this paper are openly and publicly available and can be found at https://data.mendeley.com/datasets/66gc4zb37c/1 or Zhu, Xingwang (2025) “NHFOV as Primary Support in Very Preterm Infants With RDS,” Mendeley Data, V1, doi: 10.17632/66gc4zb37c.1”
“The code used to analyse the data in the paper can be found in the supplemental files.”
“The data underlying the findings in this paper are openly and publicly available and can be found at https://data.mendeley.com/datasets/66gc4zb37c/1 or Zhu, Xingwang (2025) “NHFOV as Primary Support in Very Preterm Infants With RDS,” Mendeley Data, V1, doi: 10.17632/66gc4zb37c.1 .”
“The code used to analyse the data in the paper can be found in the supplemental files.”
Trial registration number is provided. CONSORT compliance is stated. Limitations are discussed in detail. Conclusions are proportional. Funding sources and competing interests are declared.
“ClinicalTrials.gov NCT05141435”
“The trial was conducted in compliance with CONSORT (consolidated standards of reporting trials) guidelines.”
“Clinician masking was unfeasible owing to the inherently distinct interventions, which may have introduced selection, performance, or detection bias.”
“Trial registration ClinicalTrials.gov NCT05141435”
“The trial was conducted in compliance with CONSORT (consolidated standards of reporting trials) guidelines.”
“Clinician masking was unfeasible owing to the inherently distinct interventions, which may have introduced selection, performance, or detection bias.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 37 references by DOI: 36 verified — 1 no DOI (shown, not verified).
- NO DOIPractice of NeonatologyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- dataMendeley DataLIVEHTTP 200https://data.mendeley.com/datasets/66gc4zb37c/1Resolves to Mendeley Data (data repository).
- datahttps://clinicaltrials.gov/ct2/show/NCT05141435LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAuthor affiliations“Resrach”→ ResearchTypo in affiliation 7.
- MINORconsistencyTable 3“80/172(46.5)”→ 80/172 (46.5)Missing space before parenthesis.
- MINORclarityMethods, Statistical analysis“A two tailed P value <0.05 was considered statistically significant.”→ A two-tailed P value <0.05 was considered statistically significant.Hyphenation of 'two-tailed'.
- MINORtypoAuthor affiliations“Yunan”→ YunnanMisspelling of province name.
- MINORconsistencyAbstract and Results“27 of` 170”→ 27 of 170Stray backtick in abstract.
- MINORclarityMethods, Statistical analysis“For continuous outcomes, the mean difference with a 95% confidence interval or the Hodges-Lehmann median difference with interquartile range was reported, as appropriate.”→ Consider clarifying when each method is used.Ambiguity in reporting method selection.
The published paper is robust and well-reported. An informed reader should note the minor gaps in explicit outlier handling and statistical assumption verification, but these do not undermine the overall validity. No erratum or correction is warranted based on this audit.
- 1.HIGHreportingAdd an explicit statement in the Methods (Statistical analysis) on how outliers were handled, e.g., whether any data points were excluded and on what basis.Both reviewers noted that outlier handling is not explicitly described, which is a standard reporting expectation for clinical trials.
- 2.HIGHstatisticsClarify in the Methods (Statistical analysis) how normality and equal variance assumptions were verified for parametric tests, or state that non-parametric tests were used due to non-normality.Reviewer 2 flagged assumptions_verified as inadequate; explicit verification strengthens statistical rigor.
- 3.MEDIUMdata codeDeposit the analysis code in a public repository (e.g., GitHub, Zenodo) with a DOI, rather than only in supplemental files.Both reviewers suggested this to enhance long-term reproducibility and discoverability.
- 4.MEDIUMreportingProvide a link to the trial protocol publication or registration record for full methodological details.Reviewer 2 suggested this to allow readers to access the pre-specified analysis plan.
- 5.MEDIUMreportingClarify the exact role of the local institutional review boards in the ethics approval statement, as only the central committee's approval number is given.Reviewer 1 noted this as a minor reporting gap for transparency.
- 6.MEDIUMreportingConsider reporting additional demographic details such as race/ethnicity if available, to enhance generalizability.Reviewer 2 suggested this as a potential improvement for demographic reporting.
- 7.LOWcopyeditFix typo in Author affiliations: change 'Resrach' to 'Research'.Copyedit pass flagged this minor typo.
- 8.LOWcopyeditFix typo in Author affiliations: change 'Yunan' to 'Yunnan'.Copyedit pass flagged this misspelling of the province name.
- 9.LOWcopyeditAdd a space before parenthesis in Table 3: change '80/172(46.5)' to '80/172 (46.5)'.Copyedit pass flagged this consistency issue.
- 10.LOWcopyeditFix hyphenation in Methods (Statistical analysis): change 'two tailed' to 'two-tailed'.Copyedit pass flagged this clarity issue.
- 11.LOWcopyeditRemove stray backtick in Abstract: change '27 of` 170' to '27 of 170'.Copyedit pass flagged this typo.
- 12.LOWcopyeditClarify in Methods (Statistical analysis) when mean difference vs. Hodges-Lehmann median difference is used for continuous outcomes.Copyedit pass flagged this ambiguity in reporting method selection.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.