Natural ovulation versus programmed regimens before frozen embryo transfer in ovulatory women: multicentre, randomised clinical trial.
Wei D, Qin Y, Sun Y, Yan J, Zhao H, Guan Y, Tan J, Guo T, Wang Z, Gong F, Hao C, Ma X, Zhang C, Zhang A, Geng L, Sun M, Li X, Ling X, Lu Q, Bao H, Chao L, Huang W, Shi Q, Zhao J, Lu Y, Wu S, Zhang S, Wang J, Guo M, Sun X, Ma Y, Wu Q, Li Y, Ou X, Fang Z, Chen J, Hao G, Zhang H, Legro RS, Chen ZJ, PnROVE Study Group
- DOI
- 10.1136/bmj-2025-087045
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/c6a093ec-be73-4aec-bc96-cbd11b392ee5 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ReportingData & code availability partially met−0.25★
- No data or code availability links were detected to verify.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed and transparently reported multicentre randomised trial with rigorous methodology, clear ethical approvals, and appropriate statistical analysis. The main weakness is the incomplete data availability statement with a truncated URL and lack of a persistent repository identifier for data and code.
Both reviewers independently scored all eight dimensions and agreed on every status; no divergence required reconciliation. The statistics verification recomputed only 3 of the reported tests (those with test statistics/df or effect estimates with CIs); the remaining statistics were not machine-verified and should not be assumed correct. The citation check found no retracted or unresolvable references.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks.
- CONSISTENTreported p = .490 · recomputed p = .490Reviewers 1, 2Healthy live birth primary outcome comparison
“910 (41.6%) of 2185 patients in the natural ovulation regimen group and 890 (40.6%) of 2191 in the programmed regimen group had a healthy live birth (difference between groups 1.0% (95% CI −1.9% to 3.9%); relative ratio 1.03 (95% CI 0.96 to 1.10); P=0.49)”
Taken as given: The numbers 910 and 890 are the event counts in each group.; The numbers 2185 and 2191 are the total patients in each group.; The test used is a chi-square test for 2x2 table.Method: Pearson chi-square test on the 2x2 table.How we recomputed it: pChi2x2(910, 2185-910, 890, 2191-890) - CONSISTENTreported p = .020 · recomputed p = .024Reviewers 1, 2Pre-eclampsia among clinical pregnancies comparison
“the natural ovulation regimen group had a lower risk of pre-eclampsia among those who achieved clinical pregnancy (2.9% (38 of 1302) v 4.6% (61 of 1326), P=0.02)”
Taken as given: The numbers 38 and 61 are the event counts in each group.; The numbers 1302 and 1326 are the clinical pregnancy totals in each group.; The test used is a chi-square test for 2x2 table.Method: Pearson chi-square test on the 2x2 table.How we recomputed it: pChi2x2(38, 1302-38, 61, 1326-61) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Postpartum haemorrhage comparison
“postpartum haemorrhage (2.0% (22 of 1117) v 6.1% (67 of 1100); 0.32 (0.20 to 0.52); P<0.001)”
Taken as given: The numbers 22 and 67 are the event counts in each group.; The numbers 1117 and 1100 are the total deliveries in each group.; The test used is a chi-square test for 2x2 table.Method: Pearson chi-square test on the 2x2 table.How we recomputed it: pChi2x2(22, 1117-22, 67, 1100-67)
- lowinternal contradictionThe data availability statement URL is truncated, but this is a reporting issue rather than a validity threat.
“The data underlying the findings in this paper are openly and publicly available at https://githu”
Data availabilityFind in source
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
7 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2The findings support recommending a natural ovulation regimen based on its safety profile.The recommendation is based on secondary outcomes and may be considered a clinical opinion; the primary outcome showed equivalence, not superiority.Evidence: Lower risks of maternal complications, but the primary outcome was not superior.
“Therefore, we recommend a natural ovulation regimen based on its safety profile, especially for patients at high risk of hypertensive disorders of pregnancy, such as patients aged 38 years or older and those with obesity.”
DiscussionFind in source - supportedReviewers 1, 2A natural ovulation regimen is as effective as a programmed regimen for achieving a healthy live birth.The primary outcome shows no significant difference between groups, supporting the claim of equivalence.Evidence: Healthy live birth rates: 41.6% vs 40.6%, RR 1.03 (95% CI 0.96-1.10), P=0.49.
“In the intention-to-treat analyses, 910 (41.6%) of 2185 patients in the natural ovulation regimen group and 890 (40.6%) of 2191 in the programmed regimen group achieved a healthy live birth (relative ratio 1.03 (95% confidence interval (CI) 0.96 to 1.10); P=0.49).”
AbstractFind in source - supportedReviewers 1, 2A natural ovulation regimen reduces the risk of pre-eclampsia compared with a programmed regimen.The secondary outcome shows a statistically significant reduction in pre-eclampsia among clinical pregnancies.Evidence: Pre-eclampsia among clinical pregnancies: 2.9% vs 4.6%, RR 0.63 (95% CI 0.43-0.94), P=0.02.
“The risk of pre-eclampsia was lower in the natural ovulation regimen group among patients who achieved clinical pregnancy than in the programmed regimen group (2.9% (38 of 1302) v 4.6% (61 of 1326); 0.63 (0.43 to 0.94); P=0.02).”
AbstractFind in source - supportedReviewer 1A natural ovulation regimen reduces the risk of several other maternal complications.Multiple secondary outcomes show significant reductions, though some were not pre-specified as primary.Evidence: Lower risks of early pregnancy loss, placental accreta spectrum, caesarean section, and postpartum haemorrhage.
“The incidences of early pregnancy loss (12.1% (158 of 1302) v 15.2% (201 of 1326); 0.80 (0.66 to 0.97)), placental accreta spectrum (1.8% (24 of 1302) v 3.6% (48 of 1326); 0.51 (0.31 to 0.83)), caesarean section (69.5% (776 of 1117) v 75.6% (831 of 1100); 0.92 (0.87 to 0.97)), and postpartum haemorrhage (2.0% (22 of 1117) v 6.1% (67 of 1100); 0.32 (0.20 to 0.52)) were lower in the natural ovulation regimen group.”
AbstractFind in source - supportedReviewer 1The natural ovulation regimen has a higher cycle cancellation rate.The result is clearly reported and statistically significant.Evidence: Cycle cancellation: 16.2% vs 11.5%, P<0.001.
“The rate of cycle cancellation was higher in the natural ovulation regimen (16.2% (354 of 2185) v 11.5% (251 of 2191), P<0.001).”
AbstractFind in source - supportedReviewer 2A natural ovulation regimen reduces the risk of several maternal complications including early pregnancy loss, placental accreta spectrum, caesarean section, and postpartum haemorrhage.These secondary outcomes show significant reductions, supporting the claim.Evidence: Secondary outcomes in abstract and Table 4.
“The incidences of early pregnancy loss (12.1% (158 of 1302) v 15.2% (201 of 1326); 0.80 (0.66 to 0.97)), placental accreta spectrum (1.8% (24 of 1302) v 3.6% (48 of 1326); 0.51 (0.31 to 0.83)), caesarean section (69.5% (776 of 1117) v 75.6% (831 of 1100); 0.92 (0.87 to 0.97)), and postpartum haemorrhage (2.0% (22 of 1117) v 6.1% (67 of 1100); 0.32 (0.20 to 0.52)) were lower in the natural ovulation regimen group.”
AbstractFind in source - supportedReviewer 2The rate of cycle cancellation is higher with a natural ovulation regimen.The result shows a significantly higher cancellation rate, supporting the claim.Evidence: Cycle cancellation result in abstract and Table 2.
“The rate of cycle cancellation was higher in the natural ovulation regimen (16.2% (354 of 2185) v 11.5% (251 of 2191), P<0.001).”
AbstractFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary efficacy outcome is a healthy live birth, which is a hard clinical outcome. The coprimary safety outcome is pre-eclampsia, also a clinical outcome. No surrogate biomarkers are used as primary endpoints.
“Primary outcomes were a healthy live birth and pre-eclampsia or eclampsia after a frozen embryo transfer.”
- ADEQUATEEffect sizeThe primary efficacy outcome (healthy live birth) showed no significant difference (RR 1.03, 95% CI 0.96-1.10), indicating non-inferiority. The safety outcome (pre-eclampsia) showed a relative reduction of 37% (RR 0.63, 95% CI 0.43-0.94), which is clinically meaningful and anchored to established clinical significance.
“The risk of pre-eclampsia was lower in the natural ovulation regimen group among patients who achieved clinical pregnancy than in the programmed regimen group (2.9% (38 of 1302) v 4.6% (61 of 1326); 0.63 (0.43 to 0.94); P=0.02).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites previous randomised trials and observational studies, noting that prior trials were underpowered for obstetric complications and that observational studies suggested increased risks with programmed regimens. The rationale links this to the need for a large randomised trial. The paper explicitly states the hypothesis and addresses limitations of prior work by designing a trial powered for pre-eclampsia.
“Previous randomised trials comparing the regimens for endometrial preparation focused on pregnancy and live birth rate, but they were not sufficiently powered to capture differences in the risks of obstetric and neonatal complications.”
“Observational studies suggested that programmed regimens were associated with increased risks of pre-eclampsia, hypertensive disorders of pregnancy, preterm birth, postpartum haemorrhage, macrosomia, and large-for-gestational-age babies compared with natural ovulation regimens.”
“In this multicentre randomised trial, we tested the hypothesis that a natural ovulation regimen would improve the chance of a healthy live birth and lower the risk of pre-eclampsia or eclampsia after frozen embryo transfer compared with a programmed regimen.”
“Previous randomised trials comparing the regimens for endometrial preparation focused on pregnancy and live birth rate, but they were not sufficiently powered to capture differences in the risks of obstetric and neonatal complications.”
“Observational studies suggested that programmed regimens were associated with increased risks of pre-eclampsia, hypertensive disorders of pregnancy, preterm birth, postpartum haemorrhage, macrosomia, and large-for-gestational-age babies compared with natural ovulation regimens.”
“In this multicentre randomised trial, we tested the hypothesis that a natural ovulation regimen would improve the chance of a healthy live birth and lower the risk of pre-eclampsia or eclampsia after frozen embryo transfer compared with a programmed regimen.”
Randomization used block randomization with variable block sizes, stratified by site, with sequence generated by a data coordinating centre. Blinding of outcome assessors is described. Power analysis is detailed with assumptions and sample size inflation. Inclusion/exclusion criteria are pre-specified. Outlier handling is addressed through ITT and per-protocol analyses. Controls are inherent in the comparator arm. Independent replication is not applicable for a single pivotal trial.
“We stratified randomisation by study site using block randomisation with variable block sizes of two, four, and six. The data coordinating centre at Shanghai Jiao Tong University generated the randomisation sequence.”
“Experienced obstetricians and paediatricians blinded to treatment allocation assessed all outcomes.”
“Assuming a 5% increase in the healthy live birth rate in the natural ovulation regimen group was clinically significant, the minimum sample size was 1969 patients in each group (3938 patients in total) providing 90% power at a two sided significance level of 0.05 and 84% power at a 0.025 significance level.”
“We stratified randomisation by study site using block randomisation with variable block sizes of two, four, and six.”
“Experienced obstetricians and paediatricians blinded to treatment allocation assessed all outcomes.”
“Assuming a 5% increase in the healthy live birth rate in the natural ovulation regimen group was clinically significant, the minimum sample size was 1969 patients in each group (3938 patients in total) providing 90% power at a two sided significance level of 0.05 and 84% power at a 0.025 significance level.”
The study enrolled only women, which is appropriate for the condition. Age, BMI, and health status are reported in baseline characteristics. Demographics are detailed. Species/strain and housing are not applicable for a human trial.
“Mean (SD) age at randomisation (years) | 32.68 (3.92) | 32.67 (3.89)”
“Diagnosis of chronic hypertension | 20 (0.9) | 19 (0.9)”
“The trial included women aged 20-40 years with regular menstrual cycles”
“Mean (SD) age at randomisation (years) | 32.68 (3.92) | 32.67 (3.89)”
“Data are number (%) or number/total number (%) unless otherwise specified”
The trial was approved by a named ethics committee with an approval ID, and all patients provided written informed consent. Regulatory compliance is implied through adherence to standard ethical principles, though not explicitly named.
“This trial was approved by the ethics committee of Hospital for Reproductive Medicine Affiliated to Shandong University (ID:2022-89) and all participating centres.”
“All patients provided written informed consent.”
“This trial was approved by the ethics committee of Hospital for Reproductive Medicine Affiliated to Shandong University (ID:2022-89) and all participating centres.”
“All patients provided written informed consent.”
The trial uses named drugs (dydrogesterone, vaginal progesterone, oestrogen) with manufacturers. The statistical software SAS version 9.4 is identified. No antibodies, cell lines, or mycoplasma testing are applicable.
“including oral dydrogesterone (Duphaston; Abbott) alone or combined with vaginal progesterone (Crinone; Merck Serono, or Utrogestan; Besins Healthcare)”
“All analyses were performed using SAS software (version 9.4).”
“including oral dydrogesterone (Duphaston; Abbott) alone or combined with vaginal progesterone (Crinone; Merck Serono, or Utrogestan; Besins Healthcare)”
“All analyses were performed using SAS software (version 9.4).”
Tests are named (chi-square, Fisher's, t-test, Wilcoxon). Assumptions are addressed via histogram and Kolmogorov-Smirnov test. Exact p-values are reported. Effect sizes with CIs are provided. Software is identified. Data presentation follows clinical trial conventions. Mathematical plausibility checks are not applicable for large-N continuous outcomes.
“Categorical variables were reported as frequency (percentage) and compared using χ 2 test or Fisher’s test as appropriate.”
“910 (41.6%) of 2185 patients in the natural ovulation regimen group and 890 (40.6%) of 2191 in the programmed regimen group had a healthy live birth (difference between groups 1.0% (95% CI −1.9% to 3.9%); relative ratio 1.03 (95% CI 0.96 to 1.10); P=0.49)”
“Categorical variables were reported as frequency (percentage) and compared using χ 2 test or Fisher’s test as appropriate.”
“relative ratio 1.03 (95% confidence interval (CI) 0.96 to 1.10)”
The data availability statement says data are 'openly and publicly available' but the URL is truncated and no repository or accession number is given. Code is mentioned as available in supplemental files, but no repository link is provided.
“The data underlying the findings in this paper are openly and publicly available at https://githu”
“The code used to analyse the data in the paper can be found in the supplemental files.”
“The data underlying the findings in this paper are openly and publicly available at https://githu”
“The code used to analyse the data in the paper can be found in the supplemental files.”
The trial is registered (ChiCTR2200057990). Methods are detailed. All outcomes are reported. Limitations are discussed. Conclusions are proportional. Funding and COI are stated. Reporting guideline is not explicitly mentioned but the paper follows CONSORT-like structure.
“Trial registration Chinese Clinical Trial Registry ChiCTR2200057990.”
“Funding: This trial was funded by grants from the National Natural Science Foundation of China (32588201, 82421004, and 82495194) and National Key Research and Development Programme of China (2022YFC2703502 and 2023YFC2705502).”
“Trial registration Chinese Clinical Trial Registry ChiCTR2200057990.”
“Funding: This trial was funded by grants from the National Natural Science Foundation of China (32588201, 82421004, and 82495194) and National Key Research and Development Programme of China (2022YFC2703502 and 2023YFC2705502).”
“Several limitations also need to be mentioned. Firstly, the proportion of protocol deviations (nearly 15% in each group), similar to the findings of our previous trials with a pragmatic design, may bias the results.”
Registered (7 IDs: ClinicalTrials.gov, Chinese Clinical Trial Registry). Reporting guideline cited: CONSORT.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 35 references by DOI: 35 verified.
Every extracted reference resolved against Crossref/OpenAlex with no retraction flags.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAbstract, Results“relative ratio 1.03 (95% confidence interval (CI) 0.96 to 1.10); P=0.49”→ Consider using 'risk ratio' or 'relative risk' consistently instead of 'relative ratio'.Terminology inconsistency.
- MINORconsistencyData availability statement“https://githu”→ Complete the URL and provide a full repository link.Truncated URL.
- MINORclarityMethods, Statistical analysis“A two sided P value <0.025 was considered statistically significant for the primary outcomes and <0.05 for the secondary outcomes.”→ Clarify that the 0.025 threshold is for the two primary outcomes to control family-wise error.Could be clearer.
- MINORtypoData availability statement“https://githu”→ Complete the URL to the actual repository.The URL is truncated.
- MINORconsistencyAbstract“relative ratio”→ Use 'relative risk' or 'risk ratio' consistently.The term 'relative ratio' is unusual; standard is 'relative risk'.
- MINORgrammarMethods, Statistical analysis“student’s t test”→ Capitalize 'Student's t test'.Proper noun should be capitalized.
The published work is methodologically robust and the findings are credible, but the incomplete data availability statement (truncated URL, no repository/DOI for data or code) is a transparency gap that an informed reader should weigh; it warrants a correction or clarification from the authors. No statistical or citation integrity concerns were identified.
- 1.HIGHdata codeComplete the truncated URL in the Data availability statement and provide a full, working repository link (e.g., Zenodo, Dryad, or GitHub) with a persistent identifier such as a DOI.The current statement is incomplete and does not allow readers to actually access the data, undermining reproducibility.
- 2.HIGHdata codeDeposit the analysis code in a version-controlled public repository (e.g., GitHub, Zenodo) with a DOI and provide the link in the Data availability statement.Code is currently only mentioned as available in supplemental files without a stable link, which is inadequate for a clinical trial.
- 3.MEDIUMreportingExplicitly state adherence to the CONSORT reporting guideline in the Methods or a separate section, and mention that the CONSORT checklist is available as supplementary material.The paper follows CONSORT-like structure but does not explicitly declare it, which is a transparency expectation for randomised trials.
- 4.MEDIUMethicsClarify the regulatory compliance framework (e.g., Declaration of Helsinki) in the Ethics statements section.The ethics statement names the committee and consent but does not explicitly reference the ethical principles followed.
- 5.MEDIUMcopyeditReplace the non-standard term 'relative ratio' with 'relative risk' or 'risk ratio' consistently throughout the Abstract and Results.The term 'relative ratio' is unusual and inconsistent with standard epidemiological terminology.
- 6.MEDIUMcopyeditCapitalize 'Student's t test' in the Methods, Statistical analysis section.Proper noun should be capitalized for correctness.
- 7.LOWstatisticsClarify in the Methods, Statistical analysis section that the 0.025 significance threshold applies to the two primary outcomes to control family-wise error.The current wording could be misinterpreted; explicit clarification improves precision.
- 8.LOWreportingConsider adding a statement about the availability of the full protocol and statistical analysis plan in a public repository.This would further enhance transparency and reproducibility for readers.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.