Non-invasive high frequency oscillatory ventilation for primary respiratory support in extremely preterm infants: multicentre randomised controlled trial
Li Y, Zhu X, Li LJ, Chen L, Yang Q, Xu L, Liang W, Lin X, Li C, Xue J, Liu L, Pan X, Ju R, Peng X, Tang W, Shi Y, NHFOV study group.
- DOI
- 10.1136/bmj-2025-085569
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/73e3dad1-4ce3-4c93-a092-1a4f3d5f7529 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ReportingEthical approvals partially met−0.25★
- ReportingData & code availability partially met−0.25★
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted, well-reported multicentre RCT comparing NHFOV vs. NCPAP in extremely preterm infants. The study demonstrates strong methodological rigor in design, statistical analysis, and reporting transparency. Two minor reporting gaps—lack of a named ethics framework and analysis code only in supplemental files—result in 'warn' ratings for ethical_approvals and data_code_availability, but these do not undermine the paper's core validity.
This is a post-publication audit of an already published RCT. The evaluation covered all eight dimensions of scientific rigor. The study is a human interventional trial, so criteria such as animal housing, cell line authentication, and independent replication were scored as not applicable. The three independent reviewer runs converged on most dimensions; the main disagreements were on ethical approvals and data code availability, which were resolved by applying the scoring rules to the checklist sub-criteria.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 11 tests: 11 consistent, 0 inconsistent; 11 via agent-written checks.
- CONSISTENTreported p = .007 · recomputed p = .007Reviewers 1, 2Primary outcome: Treatment failure within 72 hours (NHFOV vs NCPAP)
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007; ).”
Taken as given: The counts 27/170 and 48/172 represent event counts and total N for two independent groups.; The test used was a chi-squared test for independence.Method: Pearson's Chi-squared test, two-tailedHow we recomputed it: pChi2x2(27, 170-27, 48, 172-48) - CONSISTENTreported p = .008 · recomputed p = .009Reviewer 1Secondary outcome: Treatment failure within seven days (NHFOV vs NCPAP)
“Treatment failure within seven days after birth was reported in 36 of 170 infants (21.2%) in the NHFOV group and 58 of 172 (33.7%) in the NCPAP group (risk difference −12.5 percentage points, 95% confidence interval −21.9 to −3.2; P=0.008; ).”
Taken as given: The counts 36/170 and 58/172 represent event counts and total N for two independent groups.; The test used was a chi-squared test for independence.Method: Pearson's Chi-squared test, two-tailedHow we recomputed it: pChi2x2(36, 170-36, 58, 172-58) - CONSISTENTreported p = .050 · recomputed p = .054Reviewer 1Reason for treatment failure: Refractory hypoxia (NHFOV vs NCPAP)
“Refractory hypoxia | 19 (11.2) | 32 (18.6) | −7.4 (−14.9 to 0.1) | 0.05 |”
Taken as given: The counts 19/170 and 32/172 represent event counts and total N for two independent groups.; The test used was a chi-squared test for independence.Method: Pearson's Chi-squared test, two-tailedHow we recomputed it: pChi2x2(19, 170-19, 32, 172-32) - CONSISTENTreported p = .007 · recomputed p = .007Reviewer 2Primary outcome: Pearson chi-square on the 2x2 table (27/170 vs 48/172)
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007”
Taken as given: 27 and 48 are the event counts in the NHFOV and NCPAP arms respectively; The group totals are 170 and 172, giving non-events 143 and 124; df = 1 because the table is 2x2; the test is two-sided Pearson chi-squareMethod: Pearson chi-square test on the 2x2 cell counts.How we recomputed it: pChi2x2(27, 143, 48, 124) - CONSISTENTreported p = .008 · recomputed p = .009Reviewer 2Secondary outcome: risk difference in treatment failure within seven days
“Treatment failure within seven days was also lower in the NHFOV group (−12.5 percentage points, 95% confidence interval −21.9 to −3.2; P=0.008”
Taken as given: The −12.5 is the difference in percentage points between the two groups; The CI is two-sided at 95%; The test is a two-sided test of a difference of two proportionsMethod: Back-derived p-value from the estimate and 95% CI using a normal approximation.How we recomputed it: pCI(-12.5, -21.9, -3.2, 0) - CONSISTENTreported p = .008 · recomputed p = .009Reviewer 2Secondary outcome: Pearson chi-square on the seven-day failure 2x2 table (36/170 vs 58/172)
“Treatment failure within seven days after birth was reported in 36 of 170 infants (21.2%) in the NHFOV group and 58 of 172 (33.7%) in the NCPAP group (risk difference −12.5 percentage points, 95% confidence interval −21.9 to −3.2; P=0.008”
Taken as given: 36 and 58 are the event counts in the NHFOV and NCPAP arms respectively; The group totals are 170 and 172, giving non-events 134 and 114; df = 1 because the table is 2x2; the test is two-sided Pearson chi-squareMethod: Pearson chi-square test on the 2x2 cell counts.How we recomputed it: pChi2x2(36, 134, 58, 114) - CONSISTENTreported p = .050 · recomputed p = .053Reviewer 2Secondary outcome: risk difference in refractory hypoxia (treatment failure reason)
“those with reported hypoxia seemed to have borderline reduction in risk in the NHFOV group compared with the NCPAP group (−7.4 percentage points, 95% confidence interval −14.9 to 0.1; P=0.05;”
Taken as given: The −7.4 is the difference in percentage points between the two groups; The CI is two-sided at 95%; The test is a two-sided test of a difference of two proportionsMethod: Back-derived p-value from the estimate and 95% CI using a normal approximation.How we recomputed it: pCI(-7.4, -14.9, 0.1, 0) - CONSISTENTreported p = .007 · recomputed p = .007Reviewer 3Primary outcome: χ² test on 2×2 table (treatment failure within 72h, NHFOV 27/170 vs NCPAP 48/172).
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007)”
Taken as given: The 27 and 48 are the event counts in the NHFOV and NCPAP groups respectively; The 170 and 172 are the group totals; The non-event counts are 143 (=170−27) and 124 (=172−48); df = 1 because the table is 2×2; The paper used a χ² test for dichotomous outcomesMethod: Pearson χ² on the 2×2 table from cell counts, two-tailed, df=1; p≈0.0072, consistent with reported P=0.007.How we recomputed it: pChi2x2(27,143,48,124) - CONSISTENTreported p = .007 · recomputed p = .007Reviewer 3Primary outcome risk difference, p derived from estimate and 95% CI.
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007)”
Taken as given: −12.0 is the risk difference in percentage points; −20.7 and −3.4 are the lower and upper bounds of a two-sided 95% CI; log=0 because the estimate is a difference, not a ratioMethod: Two-tailed normal-approximation p from the estimate and 95% CI width; p≈0.0065, consistent with reported P=0.007.How we recomputed it: pCI(-12.0, -20.7, -3.4, 0) - CONSISTENTreported p = .008 · recomputed p = .009Reviewer 3Secondary outcome: treatment failure within 7 days, 2×2 χ² test (36/170 vs 58/172).
“Treatment failure within seven days after birth was reported in 36 of 170 infants (21.2%) in the NHFOV group and 58 of 172 (33.7%) in the NCPAP group (risk difference −12.5 percentage points, 95% confidence interval −21.9 to −3.2; P=0.008)”
Taken as given: The 36 and 58 are the event counts in the NHFOV and NCPAP groups respectively; The 170 and 172 are the group totals; The non-event counts are 134 (=170−36) and 114 (=172−58); df = 1 because the table is 2×2; The paper used a χ² test for dichotomous outcomesMethod: Pearson χ² on the 2×2 table, two-tailed, df=1; p≈0.0093, close to reported P=0.008, within rounding.How we recomputed it: pChi2x2(36,134,58,114) - CONSISTENTreported p = .008 · recomputed p = .009Reviewer 3Secondary outcome risk difference at 7 days, p derived from estimate and 95% CI.
“Treatment failure within seven days after birth was reported in 36 of 170 infants (21.2%) in the NHFOV group and 58 of 172 (33.7%) in the NCPAP group (risk difference −12.5 percentage points, 95% confidence interval −21.9 to −3.2; P=0.008)”
Taken as given: −12.5 is the risk difference in percentage points; −21.9 and −3.2 are the lower and upper bounds of a two-sided 95% CI; log=0 because the estimate is a difference, not a ratioMethod: Two-tailed normal-approximation p from the estimate and 95% CI width; p≈0.0088, consistent with reported P=0.008.How we recomputed it: pCI(-12.5, -21.9, -3.2, 0)
- lowinternal contradictionBaseline characteristics were not perfectly balanced: CRIB II score (P=0.05) and PaO2 before enrolment (P=0.02) were significantly higher in the NHFOV group; the authors disclose this and ran sensitivity analyses adjusting for them, so it is reported transparently rather than concealed.
“Baseline characteristics were well balanced between the two groups except for clinical risk index for babies II score (P=0.05) and PaO 2 before enrolment (P=0.02), which is significantly higher in the NHFOV group”
ResultsFind in source
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
12 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewer 1Our findings corroborate and extend the conclusions of two recent meta-analyses, and provide new insights into NHFOV’s potential use in managing respiratory distress syndrome in this vulnerable population.The study's findings align with some prior research and meta-analyses, but the claim of 'new insights' is a bit strong given the inconsistent findings in the field and the study's own limitations regarding generalizability to diverse populations.Evidence: Discussion, Comparison with other studies, paragraph 1
“Our findings corroborate and extend the conclusions of two recent meta-analyses, and provide new insights into NHFOV’s potential use in managing respiratory distress syndrome in this vulnerable population.”
Discussion ¶1Find in source - partialReviewer 3The null bronchopulmonary dysplasia result is attributable at least in part to the trial being underpowered for this secondary outcome.The paper states the trial was underpowered for BPD as a secondary outcome and notes a non-significant reduction; this is a reasonable interpretation but is an inference, not a demonstrated cause.Evidence: BPD incidence 65/170 (38.2%) vs 77/172 (44.8%), risk difference −6.5 pp (95% CI −17.0 to 3.9; P=0.22); author states 'our trial was underpowered for bronchopulmonary dysplasia as a secondary outcome'.
“Secondly, our trial was underpowered for bronchopulmonary dysplasia as a secondary outcome.”
DiscussionFind in source - supportedReviewer 1NHFOV is more efficacious than NCPAP in reducing invasive mechanical ventilation as primary respiratory support for extremely preterm infants with respiratory distress syndrome.The primary outcome results directly support this claim, showing a statistically significant reduction in treatment failure (need for invasive mechanical ventilation) in the NHFOV group.Evidence: Results, Primary outcome, paragraph 1; Table 2
“To test the hypothesis that non-invasive high frequency oscillatory ventilation (NHFOV) is more efficacious than nasal continuous positive airway pressure (NCPAP) in reducing invasive mechanical ventilation as primary respiratory support for extremely preterm infants with respiratory distress syndrome.”
AbstractFind in source - supportedReviewer 1Treatment failure within 72 hours occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 infants (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007).This is a direct reporting of the primary outcome, which is clearly presented in the results section and Table 2.Evidence: Results, Primary outcome, paragraph 1; Table 2
“Treatment failure within 72 hours occurred in 27 of` 170 infants (15.9%) in the NHFOV group and 48 of 172 infants (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007).”
AbstractFind in source - supportedReviewers 1, 3Both techniques did not show significant differences in neonatal adverse events.The safety outcomes section and Table 3 show no significant differences in various adverse events between the two groups, directly supporting this claim.Evidence: Results, Safety outcomes, paragraph 1; Table 3
“Both techniques did not show significant differences in neonatal adverse events.”
ConclusionFind in source - supportedReviewers 1, 2, 3NHFOV appeared superior to NCPAP in reducing the need for intubation when used as a primary respiratory support strategy in extremely preterm infants.This conclusion is directly supported by the statistically significant reduction in treatment failure (defined as need for invasive mechanical ventilation) in the NHFOV group.Evidence: Results, Primary outcome, paragraph 1; Table 2
“NHFOV appeared superior to NCPAP in reducing the need for intubation when used as a primary respiratory support strategy in extremely preterm infants.”
ConclusionFind in source - supportedReviewer 1Our trial was underpowered for bronchopulmonary dysplasia as a secondary outcome.The discussion explicitly states that the trial was underpowered for this specific secondary outcome, which is a transparent and supported limitation.Evidence: Discussion, Comparison with other studies, paragraph 2
“Secondly, our trial was underpowered for bronchopulmonary dysplasia as a secondary outcome.”
Discussion ¶2Find in source - supportedReviewer 2Treatment failure within seven days was also lower with NHFOV.The secondary outcome provides a significant result consistent with the claim.Evidence: 36/170 (21.2%) vs 58/172 (33.7%), risk difference −12.5 (−21.9 to −3.2), P=0.008.
“Treatment failure within seven days after birth was reported in 36 of 170 infants (21.2%) in the NHFOV group and 58 of 172 (33.7%) in the NCPAP group”
ResultsFind in source - supportedReviewers 2, 3Both respiratory support techniques were equally safe (no significant differences in neonatal adverse events).No significant differences were found in the reported safety outcomes, supporting the claim.Evidence: Table 3 safety outcomes: air leaks, thick secretions, death in hospital, and nasal injury all non-significant (P=0.18 to 0.64).
“Both techniques did not show significant differences in neonatal adverse events.”
AbstractFind in source - supportedReviewer 2NHFOV is feasible and potentially advantageous as a primary respiratory treatment for extremely preterm infants.The primary-result reduction in intubation and comparable safety profile back this claim, which is appropriately hedged.Evidence: Primary outcome reduction and comparable adverse-event rates across the trial.
“The results advance our current understanding by showing that NHFOV is feasible and potentially advantageous as a primary respiratory treatment for extremely preterm infants”
DiscussionFind in source - supportedReviewer 2The findings corroborate and extend two recent meta-analyses.The Discussion contextualises the results against prior RCTs and meta-analyses consistently, supporting the claim.Evidence: Discussion comparison with Malakian, Iranpour, Zhu, Mukerji, Rüegger, Klotz trials and two meta-analyses.
“Our findings corroborate and extend the conclusions of two recent meta-analyses, and provide new insights into NHFOV’s potential use in managing respiratory distress syndrome in this vulnerable population.”
DiscussionFind in source - supportedReviewer 3NHFOV reduced treatment failure within seven days compared with NCPAP.The secondary outcome is significant and reported with CI.Evidence: 36/170 (21.2%) vs 58/172 (33.7%), risk difference −12.5 pp (95% CI −21.9 to −3.2; P=0.008).
Treatment failure within seven days after birth was reported in 36 of 170 infants (21.2%) in the NHFOV group and 58 of 172 (33.7%) in the NCPAP group (risk difference −12.5 percentage points, 95% confidence interval −21.9 to −3.2; P=0.008)
Resultsreviewer’s wording
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- N/ASurrogate endpointThe primary efficacy outcome is a hard clinical outcome: the need for invasive mechanical ventilation within 72 hours after birth. No surrogate or biomarker is used as the primary basis for the efficacy claim.
“The primary outcome was respiratory support failure, defined by the need for invasive mechanical ventilation within 72 hours after birth.”
- ADEQUATEEffect sizeThe primary effect is a 12.0 percentage point absolute reduction in treatment failure (from 27.9% to 15.9%), which is statistically significant and clinically meaningful because it reflects avoidance of intubation in extremely preterm infants. The magnitude is substantial and not a small fraction of any reference value.
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007).”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
3 integrity concerns flagged (0 high).
- lowotherThe data availability statement contains a stray, unrelated DOI (10.1001/jamanetworkopen.2021.18904) embedded mid-sentence, which appears to be a copy/paste artifact rather than a validity threat.
“can be found at https://data.mendeley.com/datasets/66gc4zb37c/1 (10.1001/jamanetworkopen.2021.18904)”
Data availabilityFind in source - lowotherThe baseline table labels FiO2 with the unit 'mm Hg', but FiO2 is a fraction (values are 0.3), indicating a mislabeled unit in Table 1.
“FiO 2 before enrolment (mm Hg), median (IQR) | 0.3 (0.2-0.3)”
Table 1Find in source
Reporting gaps
2 findings · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
- Ethics/consent reporting incompleteAssessed
The introduction discusses the limitations of NCPAP, the potential advantages of NHFOV, and the inconsistent findings from previous trials, leading to a clear hypothesis for the current study. It explicitly states the gap in knowledge regarding NHFOV as primary respiratory support.
“Several randomised controlled trials comparing NHFOV and NCPAP have reported inconsistent findings about their effectiveness in managing respiratory distress syndrome.”
“NHFOV may provide dual physiological advantages over NCPAP by maintaining lung recruitment while actively clearing carbon dioxide through oscillatory gas mixing—a mechanism particularly beneficial for non-uniformly diseased lungs.”
“However, its efficacy as primary respiratory support for respiratory distress syndrome in preterm infants remains unproven.”
“Several randomised controlled trials comparing NHFOV and NCPAP have reported inconsistent findings about their effectiveness in managing respiratory distress syndrome.”
“We hypothesised that NHFOV would be more effective than NCPAP in reducing treatment failure within 72 hours after birth in a homogeneous population of Chinese extremely preterm infants.”
“Importantly, our study design addressed a key historical limitation in comparing NHFOV and NCPAP: variability in applied airway pressures.”
“European consensus guidelines recommend nasal continuous positive airway pressure (NCPAP) as the preferred first line respiratory support strategy”
“However, its efficacy as primary respiratory support for respiratory distress syndrome in preterm infants remains unproven.”
“We hypothesised that NHFOV would be more effective than NCPAP in reducing treatment failure within 72 hours after birth”
The paper describes simple randomisation using a computer-generated sequence and central management for allocation concealment. Outcome assessors were masked, and a power analysis determined the sample size. Inclusion and exclusion criteria were clearly stated, and the study used appropriate controls (NCPAP group).
“Simple randomisation was performed using a computer generated random number sequence, which was securely posted on a dedicated, password protected website accessible 24/7.”
“However, outcome assessors were masked to the treatment allocation to minimise ascertainment bias.”
“With a significance level (α) of 0.05 and 90% power, a sample size of 170 newborns per group was required, leading to a total enrolment target of at least 340.”
“Simple randomisation was performed using a computer generated random number sequence, which was securely posted on a dedicated, password protected website accessible 24/7.”
“However, outcome assessors were masked to the treatment allocation to minimise ascertainment bias.”
“With a significance level (α) of 0.05 and 90% power, a sample size of 170 newborns per group was required, leading to a total enrolment target of at least 340.”
“Simple randomisation was performed using a computer generated random number sequence, which was securely posted on a dedicated, password protected website accessible 24/7.”
“With a significance level (α) of 0.05 and 90% power, a sample size of 170 newborns per group was required”
The paper provides detailed demographics for the enrolled infants in Table 1, including gestational age, birth weight, sex (female percentage), and other relevant health statuses and clinical characteristics.
“Female | 74 (43.5) | 68 (39.5)”
“Female | 74 (43.5) | 68 (39.5)”
“Gestational age (weeks), median (IQR) | 27.0 (26.0-28.0) | 27.0 (26.0-28.0)”
“Female | 74 (43.5) | 68 (39.5)”
“Gestational age (weeks), median (IQR) | 27.0 (26.0-28.0) | 27.0 (26.0-28.0)”
The protocol was approved by a named ethics committee with a protocol number (No 2019.161) and local IRBs, and written informed consent was obtained from parents/guardians. However, regulatory compliance is only stated vaguely as 'data were anonymised in accordance with relevant local regulations' — no recognised framework (e.g., Declaration of Helsinki, ICH-GCP) is named. Under the scoring rule, any applicable inadequate/not-reported item yields a warn.
“The study was approved by the ethics committee of the Children's Hospital of Chongqing Medical University (No 2019.161)”
“Informed consent was obtained from parents or guardians antenatally or upon admission to the neonatal intensive care unit, and data were anonymised in accordance with relevant local regulations.”
“The trial was conducted in compliance with CONSORT (consolidated standards of reporting trials) guidelines.”
“The protocol was approved by the Ethics Committee of the Children’s Hospital of Chongqing Medical University (No 2019.161) and local institutional review boards of the participating centres.”
“All parents or guardians provided written informed consent.”
“data were anonymised in accordance with relevant local regulations.”
“The study was approved by the ethics committee of the Children's Hospital of Chongqing Medical University (No 2019.161)”
“All parents or guardians provided written informed consent.”
“data were anonymised in accordance with relevant local regulations. The trial was conducted in compliance with CONSORT”
The paper identifies the continuous flow devices used for NCPAP (Comen NV8, Mindray NB350) and the piston or membrane oscillators for NHFOV (Fabian-III, SLE 5000, Leoni+), along with their manufacturers. Specific medications like caffeine citrate and Curosurf are also named with their manufacturers. The statistical software and version are reported.
“Continuous flow devices (Comen NV8, China; Mindray NB350, China; supplementary eTable 2) were used to provide NCPAP.”
“NHFOV was delivered using piston or membrane oscillators capable of active expiration, including the Fabian-III (Acutronic, Switzerland), SLE 5000 (SLE (UK)) and Leoni+ (Löwenstein Medical, Germany; supplementary eTable 2).”
“All statistical analyses were conducted in R software (version 4.4.2).”
“NHFOV was delivered using piston or membrane oscillators capable of active expiration, including the Fabian-III (Acutronic, Switzerland), SLE 5000 (SLE (UK)) and Leoni+ (Löwenstein Medical, Germany”
“All statistical analyses were conducted in R software (version 4.4.2).”
“Surfactant treatment (Curosurf, Chiesi Pharmaceuticals) was administered at a dose of 200 mg/kg”
“All statistical analyses were conducted in R software (version 4.4.2).”
The paper explicitly names the statistical tests used (Student’s t test, Mann-Whitney U test, χ² test). It reports exact p-values and 95% confidence intervals for risk differences. The statistical software (R version 4.4.2) is identified. Data presentation in tables includes per-group Ns, percentages, medians with IQRs, and risk differences with 95% CIs.
“Comparisons were performed using Student’s t test for parametric continuous variables, Mann-Whitney U test for non-parametric continuous variables, and χ 2 test for dichotomous variables, respectively.”
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007; ).”
“All statistical analyses were conducted in R software (version 4.4.2).”
“Comparisons were performed using Student’s t test for parametric continuous variables, Mann-Whitney U test for non-parametric continuous variables, and χ 2 test for dichotomous variables, respectively.”
“Treatment failure within 72 hours after birth occurred in 27 of 170 infants (15.9%) in the NHFOV group and 48 of 172 (27.9%) in the NCPAP group (risk difference −12.0 percentage points, 95% confidence interval −20.7 to −3.4; P=0.007”
“Comparisons were performed using Student’s t test for parametric continuous variables, Mann-Whitney U test for non-parametric continuous variables, and χ 2 test for dichotomous variables”
“Female | 74 (43.5) | 68 (39.5)”
The data availability statement names a concrete route (Mendeley Data, doi: 10.17632/66gc4zb37c.1) and the data are deposited there, so data_availability_statement and repository_deposit are adequate. accession_numbers is n/a for patient data. code_sharing is reported_but_inadequate: 'The code used to analyse the data in the paper can be found in the supplemental files' — supplemental files are not a version-controlled public repo with a persistent identifier. Two of three applicable criteria adequate → warn.
“Data availability statement The code used to analyse the data in the paper can be found in the supplemental files. The data underlying the findings in this paper are openly and publicly available and can be found at https://data.mendeley.com/datasets/66gc4zb37c/1 or Zhu, Xingwang (2025) “NHFOV as Primary Support in Very Preterm Infants With RDS,” Mendeley Data, V1, doi: 10.17632/66gc4zb37c.1 .”
“The data underlying the findings in this paper are openly and publicly available and can be found at https://data.mendeley.com/datasets/66gc4zb37c/1 or Zhu, Xingwang (2025) “NHFOV as Primary Support in Very Preterm Infants With RDS,” Mendeley Data, V1, doi: 10.17632/66gc4zb37c.1 .”
“The code used to analyse the data in the paper can be found in the supplemental files.”
“The data underlying the findings in this paper are openly and publicly available and can be found at https://data.mendeley.com/datasets/66gc4zb37c/1”
“The code used to analyse the data in the paper can be found in the supplemental files.”
“The data underlying the findings in this paper are openly and publicly available and can be found at https://data.mendeley.com/datasets/66gc4zb37c/1”
“The code used to analyse the data in the paper can be found in the supplemental files.”
The methods section is comprehensive, detailing interventions, procedures, and outcome definitions. The trial is registered on ClinicalTrials.gov, and adherence to CONSORT guidelines is stated. Limitations are explicitly discussed, and funding sources and conflicts of interest are provided.
“Trial registration ClinicalTrials.gov NCT05141435 (https://clinicaltrials.gov/ct2/show/NCT05141435)”
“The trial was conducted in compliance with CONSORT (consolidated standards of reporting trials) guidelines.”
“Nevertheless, several limitations should be acknowledged. Clinician masking was unfeasible owing to the inherently distinct interventions, which may have introduced selection, performance, or detection bias.”
“Funding: The study was partially funded by National Key Research and Development Program of China (No 2022YFC2704803); Natural Science Foundation of Chongqing (CSTB2024NSCQ-MSX0158); Hunan Provincial Clinical Research Center for Newborn Diseases of Maternal Origins (2023SK4057); the Chongqing Maternal and Child Disease Prevention and Control and Public Health Research Project (CQFYJB01008); Key Research and Development Program of Jiangxi (20243BBI91020); and Clinical Research Project for the Summit Program of Children’s Hospital of Chongqing Medical University (CHCMU-2024-XKDF-1002).”
“Trial registration ClinicalTrials.gov NCT05141435 (https://clinicaltrials.gov/ct2/show/NCT05141435)”
“The trial was conducted in compliance with CONSORT (consolidated standards of reporting trials) guidelines.”
“The study was partially funded by National Key Research and Development Program of China (No 2022YFC2704803)”
“Trial registration ClinicalTrials.gov NCT05141435”
“The trial was conducted in compliance with CONSORT (consolidated standards of reporting trials) guidelines.”
“Clinician masking was unfeasible owing to the inherently distinct interventions, which may have introduced selection, performance, or detection bias.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 37 references by DOI: 36 verified — 1 no DOI (shown, not verified).
- NO DOIPractice of NeonatologyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT05141435LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- dataMendeley DataLIVEHTTP 200https://data.mendeley.com/datasets/66gc4zb37c/1Resolves to Mendeley Data (data repository).
Copyediting
19 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 19 minor suggestions below.
19 copyedit issues flagged: mostly punctuation, consistency, typo.
- MINORpunctuationAbstract, Results“27 of` 170 infants”→ 27 of 170 infantsTypo: backtick character before '170'.
- MINORconsistencyAbstract, Trial registration“Accepted 2025 Aug 15; Collection date 2025.”→ Accepted [Date]; Collection date [Date].The accepted date and collection date are in the future (2025), which is likely a placeholder or error in the preprint.
- MINORgrammarMethods, Study design, paragraph 1“All authors reviewed the protocol and ensured adherence throughout the trial.”→ All authors reviewed the protocol and ensured adherence to it throughout the trial.Slightly awkward phrasing, 'to it' improves clarity.
- MINORpunctuationMethods, Study participants, paragraph 1“diagnosis of respiratory distress syndrome and prerandomisation support with NCPAP set at 6 cm H 2 O of positive end expiratory pressure (after delivery room stabilisation at 6-8 cm H 2 O per local guidelines) with a fraction of inspired oxygen (FiO 2 ) >0.25 to maintain a target peripheral oxygen saturation (SpO 2 ) of 89-94%; less than two hours after birth at the time of enrolment; and informed parental consent obtained before randomisation.”→ diagnosis of respiratory distress syndrome and prerandomisation support with NCPAP set at 6 cm H 2 O of positive end expiratory pressure (after delivery room stabilisation at 6-8 cm H 2 O per local guidelines) with a fraction of inspired oxygen (FiO 2 ) >0.25 to maintain a target peripheral oxygen saturation (SpO 2 ) of 89-94%; less than two hours after birth at the time of enrolment; and informed parental consent obtained before randomisation.The semicolon after '89-94%' should probably be a comma or removed for better flow, as the next clause is part of the same list of criteria.
- MINORgrammarMethods, Study procedure, paragraph 4“Caffeine administration Prophylactic caffeine treatment (caffeine citrate injection, Chiesi Pharmaceuticals, Parma, Italy) was administered within the first 24 hours after birth.”→ Caffeine administration: Prophylactic caffeine treatment (caffeine citrate injection, Chiesi Pharmaceuticals, Parma, Italy) was administered within the first 24 hours after birth.Add a colon after the heading 'Caffeine administration'.
- MINORgrammarMethods, Other treatments, paragraph 1“Antibiotics are frequently started in newborns with respiratory distress syndrome until sepsis is excluded.”→ Antibiotics are frequently started in newborns with respiratory distress syndrome until sepsis is excluded.Consider 'Antibiotics are frequently started for newborns...' or 'Antibiotics are frequently initiated in newborns...' for slightly improved phrasing.
- MINORpunctuationResults, Study population, paragraph 1“Consequently, data from 342 newborns were included in the final analysis ().”→ Consequently, data from 342 newborns were included in the final analysis.Empty parenthesis at the end of the sentence.
- MINORpunctuationResults, Primary outcome, paragraph 1“P=0.007; ).”→ P=0.007).Extra closing parenthesis.
- MINORpunctuationResults, Secondary outcomes, paragraph 1“P=0.008; ).”→ P=0.008).Extra closing parenthesis.
- MINORconsistencyTable 3, footnote“* Risk difference reported as median (interquartile range).”→ * Median difference reported as median (interquartile range).The row 'Time to surfactant administration (hours)' and 'Duration of supplemental oxygen (days)' report median (IQR) and the footnote refers to 'Risk difference reported as median (interquartile range)', which is inconsistent. It should refer to median difference for continuous variables.
- MINORtypoAffiliation 7“Agency for Science, Technology & Resrach (A*STAR)”→ Agency for Science, Technology & Research (A*STAR)Misspelling of 'Research'.
- MINORtypoAffiliation 12“Women and Children's Health Hospital of Qujing, Yunan, China”→ Yunnan, ChinaMisspelling of the province name 'Yunnan'.
- MINORconsistencyTable 1“FiO 2 before enrolment (mm Hg), median (IQR) | 0.3 (0.2-0.3)”→ Change the unit to a fraction (e.g., 'FiO 2 before enrolment, median (IQR)')FiO2 is a fraction, not mm Hg; the unit label is incorrect.
- MINORconsistencyAssociated Data section“Data Availability Statement ... publicly available and can be found at https://data.mendeley.com/datasets/66gc4zb37c/1”→ Remove the duplicated Data Availability Statement blockThe Data Availability Statement text appears twice (main text and Associated Data section).
- MINORpunctuationAbstract, Results“Treatment failure within 72 hours occurred in 27 of` 170 infants”→ Remove the stray backtick: '27 of 170 infants'A stray backtick appears after 'of'.
- MINORotherAssociated Data / Data Availability Statement“can be found at https://data.mendeley.com/datasets/66gc4zb37c/1 (10.1001/jamanetworkopen.2021.18904)”→ Remove the stray unrelated DOI '(10.1001/jamanetworkopen.2021.18904)' from the data statement.An apparently unrelated JAMANetworkOpen DOI appears mid-sentence in the data availability statement; likely a template/citation leftover.
- MINORtypoAuthor affiliations, item 3“Women and Children's Health Hospital of Qujing, Yunan, China”→ Change 'Yunan' to 'Yunnan'.Misspelled province name.
- MINORtypoAuthor affiliations, item 7“Institute for Human Development & Potential (iHDP), Agency for Science, Technology & Resrach (A*STAR)”→ Change 'Resrach' to 'Research'.Misspelling in agency name.
- MINORconsistencyData availability statement (body vs. Associated Data)“can be found at https://data.mendeley.com/datasets/66gc4zb37c/1 ... or Zhu, Xingwang (2025)”→ Keep a single, consistent data availability statement.The data availability statement is duplicated verbatim in the body and in the Associated Data section.
As a post-publication audit, this paper is methodologically robust. An informed reader should weigh the two minor reporting gaps: the ethics statement lacks a named framework (Declaration of Helsinki/ICH-GCP), and the analysis code is only in supplemental files rather than a version-controlled repository. These are not validity threats but suggest that a correction or erratum could improve transparency. The primary findings appear reliable based on the statistical recomputation checks and the overall study design.
- 1.HIGHethicsAdd an explicit statement of compliance with a named regulatory framework (e.g., Declaration of Helsinki and/or ICH-GCP) to the ethics section.The current statement only references 'relevant local regulations' and CONSORT, which is a reporting guideline, not an ethics framework; this is a common expectation for human trials.
- 2.HIGHdata codeDeposit the analysis code in a version-controlled public repository (e.g., GitHub, Zenodo) with a permanent DOI and link to the data, and update the data availability statement accordingly.Providing code only in supplemental files lacks versioning and a persistent identifier, reducing reproducibility and discoverability.
- 3.HIGHreportingCorrect the unit label for FiO2 in Table 1 from 'mm Hg' to a fraction (or remove the unit), as FiO2 is a proportion, not a pressure.This mislabeling could confuse readers about the nature of the variable.
- 4.HIGHcopyeditRemove the stray unrelated DOI (10.1001/jamanetworkopen.2021.18904) from the data availability statement, and ensure only the correct Mendeley DOI appears.This appears to be a copy-paste artifact and could be misinterpreted as a related publication.
- 5.HIGHcopyeditRemove the duplicate Data Availability Statement block in the Associated Data section to avoid redundancy.The statement appears twice verbatim, which is untidy and may confuse readers.
- 6.MEDIUMcopyeditFix the stray backtick in the abstract: change '27 of` 170 infants' to '27 of 170 infants'.Typographical error.
- 7.MEDIUMcopyeditRemove the extra closing parenthesis in the primary outcome sentence: change 'P=0.007; )' to 'P=0.007).' (and similarly for the secondary outcome 'P=0.008; )').Punctuation error.
- 8.MEDIUMreportingAdd an explicit statement of race/ethnicity in the demographics table (e.g., 'Chinese, n (%)') to improve generalizability reporting.Although the population is implicitly Chinese, explicit reporting is standard practice for human studies.
- 9.MEDIUMcopyeditFix the footnote in Table 3: change '* Risk difference reported as median (interquartile range)' to '* Median difference reported as median (interquartile range)' for continuous variable rows.The current footnote is inconsistent with the data presented (median differences for continuous outcomes).
- 10.MEDIUMcopyeditFix the misspelling of 'Yunnan' in affiliation 12 (currently 'Yunan') and 'Research' in affiliation 7 (currently 'Resrach').Geographic and institutional name errors.
- 11.LOWcopyeditAdd a colon after the heading 'Caffeine administration' in the Methods section: change 'Caffeine administration Prophylactic...' to 'Caffeine administration: Prophylactic...'.Improves readability.
- 12.LOWcopyeditRemove the empty parenthesis at the end of the sentence in Results, Study population: 'included in the final analysis ().'Typographical artifact.
- 13.LOWcopyeditConsider replacing the semicolon after '89-94%' with a comma or period for better flow in the inclusion criteria list.Minor punctuation improvement.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.