Preimplantation genetic testing for aneuploidy versus no genetic testing in couples undergoing intracytoplasmic sperm injection for severe male infertility: multicentre, open label, randomised controlled trial.
Lin X, Wu D, Zhang C, Wang L, Lu Y, Zhou P, Zhou C, Jin L, Wang L, Zhu H, Pan J, Xu C, Chen S, Gao L, Li L, Zhang S, Wu Y, Sun Y, Mol BW, Huang H
- DOI
- 10.1136/bmj-2025-084050
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/e81366b3-db3a-4c4d-b485-48dab3478ed4 is authoritative.
How this rating was calculated
Started at 5★ — no deductions. Nothing the checks ran surfaced a material problem.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomised controlled trial. The design is rigorous with adequate randomization, sample size justification, and pre-specified analyses. Reporting is exemplary with trial registration, CONSORT flow, and data/code sharing.
Both reviewers classified the study as interventional and agreed on all dimensions. The statistics verification covered only a subset of tests (8 of many); the rest are unverified but no errors were found. The citation check found no retracted or non-existent references.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 8 tests: 8 consistent, 0 inconsistent; 6 recomputed directly from the reported test statistics, 2 via agent-written checks.
- CONSISTENTreported p = .640 · recomputed p = .644Recomputed odds ratio 1.09 (95% CI 0.76–1.58), reported p=0.64
“odds ratio 1.09 (95% confidence interval (CI) 0.76 to 1.58), P=0.64”
Taken as given: 0.76–1.58 is a two-sided 95% confidence interval for the odds ratio of 1.09, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.64 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.09, 0.76, 1.58, 1) - CONSISTENTreported p = .860 · recomputed p = .844Recomputed odds ratio 1.04 (95% CI 0.70–1.53), reported p=0.86
“odds ratio 1.04 (95% CI 0.70 to 1.53); P=0.86”
Taken as given: 0.70–1.53 is a two-sided 95% confidence interval for the odds ratio of 1.04, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.86 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.04, 0.7, 1.53, 1) - CONSISTENTreported p = .920 · recomputed p = .917Recomputed odds ratio 0.98 (95% CI 0.67–1.43), reported p=0.92
“odds ratio 0.98 (95% CI 0.67 to 1.43); P=0.92”
Taken as given: 0.67–1.43 is a two-sided 95% confidence interval for the odds ratio of 0.98, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.92 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.98, 0.67, 1.43, 1) - CONSISTENTreported p = .990 · recomputed p = 1.000Recomputed hazard ratio 1.00 (95% CI 0.79–1.27), reported p=0.99
“hazard ratio 1.00 (95% CI 0.79 to 1.27); P=0.99”
Taken as given: 0.79–1.27 is a two-sided 95% confidence interval for the hazard ratio of 1.00, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.99 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1, 0.79, 1.27, 1) - CONSISTENTreported p = .030 · recomputed p = .028Recomputed odds ratio 3.08 (95% CI 1.13–8.41), reported p=0.03
“odds ratio 3.08 (95% CI 1.13 to 8.41); P=0.03”
Taken as given: 1.13–8.41 is a two-sided 95% confidence interval for the odds ratio of 3.08, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.03 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(3.08, 1.13, 8.41, 1) - CONSISTENTreported p = .290 · recomputed p = .281Recomputed odds ratio 0.72 (95% CI 0.40–1.32), reported p=0.29
“odds ratio 0.72 (95% CI 0.40 to 1.32); P=0.29”
Taken as given: 0.40–1.32 is a two-sided 95% confidence interval for the odds ratio of 0.72, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.29 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.72, 0.4, 1.32, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Pregnancy loss after first transfer (OR 0.26, 95% CI 0.14-0.50, P<0.001)
“0.26 (0.14 to 0.50), P<0.001”
Taken as given: The odds ratio is 0.26.; The 95% CI is 0.14 to 0.50.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed p-value from the reported OR and 95% CI using the normal approximation for the log odds ratio.How we recomputed it: pCI(0.26, 0.14, 0.50, 1) - CONSISTENTreported p = .001 · recomputed p = .002Reviewer 2Cumulative pregnancy loss (OR 0.43, 95% CI 0.25-0.72, P=0.001)
“0.43 (0.25 to 0.72), P=0.001”
Taken as given: The odds ratio is 0.43.; The 95% confidence interval is 0.25 to 0.72.; The p-value is two-sided.Method: Recomputed p-value from the odds ratio and its 95% confidence interval using the normal approximation for the log odds ratio.How we recomputed it: pCI(0.43, 0.25, 0.72, 1)
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2PGT-A could possibly benefit women with BMI ≥25.The subgroup analysis shows a significant benefit, but the authors acknowledge it is exploratory and post hoc.Evidence: Post hoc analysis: live birth rate 22/35 (63%) vs 11/31 (36%), OR 3.08 (1.13-8.41), P=0.03.
“higher live birth rates and cumulative live birth rates observed in women with a body mass index of 25 or more in the PGT-A group suggested that PGT-A could possibly benefit this group of women, although the result was exploratory”
DiscussionFind in source - supportedReviewers 1, 2PGT-A did not improve live birth rates in ICSI for severe male infertility compared to ICSI alone.The primary outcomes show no significant difference in live birth rates, supporting the claim.Evidence: Live birth after first transfer: 109 (48.4%) vs 104 (46.2%), OR 1.09 (0.76-1.58), P=0.64; cumulative live birth: 136 (60.4%) vs 137 (60.9%), OR 0.98 (0.67-1.43), P=0.92.
“PGT-A did not improve live birth rates in ICSI for severe male infertility compared to ICSI alone”
ConclusionFind in source - supportedReviewers 1, 2PGT-A reduced rates of pregnancy loss.The secondary outcomes show significantly lower pregnancy loss rates in the PGT-A group.Evidence: Pregnancy loss after first transfer: 13 (5.8%) vs 43 (19.1%), OR 0.26 (0.14-0.50), P<0.001; cumulative pregnancy loss: 25 (11.1%) vs 51 (22.7%), OR 0.43 (0.25-0.72), P=0.001.
“but reduced rates of pregnancy loss.”
ConclusionFind in source - supportedReviewers 1, 2The 42.6% aneuploidy rate in severe male infertility patients provides a rationale for evaluating PGT-A.The aneuploidy rate is reported and supports the rationale.Evidence: Among 664 blastocysts tested, 42.6% (283) were aneuploid.
“The 42.6% aneuploidy rate identified in our study of severe male infertility patients provides a rationale for evaluating the potential benefits of PGT-A for this specific population.”
DiscussionFind in source - supportedReviewer 2PGT-A is not recommended as a priority strategy for couples with severe male factor infertility.The lack of benefit in live birth rates supports this recommendation, though the reduced pregnancy loss is noted.Evidence: Primary outcomes show no significant difference; secondary outcomes show reduced pregnancy loss.
“These results further suggest that PGT-A is not recommended as a priority strategy for couples with severe male factor infertility”
DiscussionFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary outcomes are live birth after first embryo transfer and cumulative live birth, which are hard clinical outcomes. The secondary outcome of pregnancy loss is also a clinical outcome. No surrogate endpoints are used as the primary basis for efficacy claims.
“Primary outcomes were live birth after the first embryo transfer and cumulative live birth (up to three transfer cycles) within 12 months after randomisation.”
- ADEQUATEEffect sizeThe primary outcomes show no significant difference between groups (live birth rate 48.4% vs 46.2%, OR 1.09, 95% CI 0.76-1.58; cumulative 60.4% vs 60.9%, OR 0.98, 95% CI 0.67-1.43). The secondary outcome of pregnancy loss shows a significant reduction (5.8% vs 19.1%, OR 0.26, 95% CI 0.14-0.50). These are clinically meaningful outcomes with effect sizes reported and statistically supported.
“PGT-A did not improve live birth rates in ICSI for severe male infertility compared to ICSI alone, but reduced rates of pregnancy loss.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites retrospective studies suggesting PGT-A benefits in severe male infertility and notes the lack of RCTs. The rationale for the trial is clearly linked to the need for prospective evidence. Limitations of prior research are implicitly addressed by the trial's design.
“Recent retrospective studies showed that PGT-A was associated with a lower rate of early miscarriage and an increased live birth rate per transfer in couples with severe male factor infertility. However, there are no randomised controlled trials assessing the effect of PGT-A in couples with male factor infertility problems undergoing ICSI.”
“We performed a multicentre randomised controlled trial in couples with severe male factor infertility undergoing ICSI, and compared the live birth rate after ICSI in couples who had or had not also undergone PGT-A”
“Recent retrospective studies showed that PGT-A was associated with a lower rate of early miscarriage and an increased live birth rate per transfer in couples with severe male factor infertility. However, there are no randomised controlled trials assessing the effect of PGT-A in couples with male factor infertility problems undergoing ICSI.”
“We performed a multicentre randomised controlled trial in couples with severe male factor infertility undergoing ICSI, and compared the live birth rate after ICSI in couples who had or had not also undergone PGT-A through a combination of morphological assessments and next generation sequencing.”
Randomization was computer-generated, stratified by maternal age and BMI, with a 1:1 allocation. The trial is open-label with a clear rationale (nature of interventions). A priori power analysis is provided with effect size, alpha, and power. Inclusion/exclusion criteria are detailed. Outlier handling is addressed through ITT and per-protocol analyses and imputation for missing data. Controls are inherent in the comparator arm. Independent replication is not applicable for a single pivotal trial.
“independent staff performed randomisation stratified by maternal age (20-29.9, 30-34.9, ≥35 years) and body mass index (<18.5, 18.5-24.9, ≥25) using an online system with a computer generated randomisation list.”
“Because of the nature of the interventions, the participants, clinicians, embryologists, and investigators assessing the outcomes were not masked to the group allocation.”
“We determined that a sample size of 436 women (218 per group) would provide a power of 80% to demonstrate or refute this 14% difference at a two sided α level of 0.05, with an estimated dropout rate of 20%.”
“independent staff performed randomisation stratified by maternal age (20-29.9, 30-34.9, ≥35 years) and body mass index (<18.5, 18.5-24.9, ≥25) using an online system with a computer generated randomisation list.”
“Because of the nature of the interventions, the participants, clinicians, embryologists, and investigators assessing the outcomes were not masked to the group allocation.”
“We determined that a sample size of 436 women (218 per group) would provide a power of 80% to demonstrate or refute this 14% difference at a two sided α level of 0.05, with an estimated dropout rate of 20%.”
Sex is reported (both male and female partners). Age and BMI are reported for both partners. Health status is implied through inclusion/exclusion criteria. Demographics are detailed in Table 1. Species/strain and housing are not applicable for a human trial.
“Primary infertility | 172 (76.4) | 164 (72.9)”
“Primary infertility | 172 (76.4) | 164 (72.9)”
The study was approved by a named ethics committee with a protocol number. Written informed consent was obtained from all participants. Compliance with Good Clinical Practice and Declaration of Helsinki is stated.
“This study was approved by the ethics committee of International Peace Maternity and Child Health Hospital of Shanghai Jiao Tong University in China (GKLW2016-16 on 19 October 2016).”
“All participants provided written informed consent.”
“The study was performed in accordance with Good Clinical Practice and Declaration of Helsinki principles”
“This study was approved by the ethics committee of International Peace Maternity and Child Health Hospital of Shanghai Jiao Tong University in China (GKLW2016-16 on 19 October 2016).”
“All participants provided written informed consent.”
“The study was performed in accordance with Good Clinical Practice and Declaration of Helsinki principles”
The investigational product is ICSI with PGT-A, and the paper identifies the sequencing platforms and software used. Reagents such as dydrogesterone and progesterone are named with manufacturers. Antibodies, cell lines, and mycoplasma testing are not applicable.
“The next generation sequencing platforms used in our study included the Illumina NextSeq 550 or Ion PGM/Proton (Thermo Fisher Scientific, Waltham, MA, USA).”
“All analyses were performed with R software, version 4.0.”
“The next generation sequencing platforms used in our study included the Illumina NextSeq 550 or Ion PGM/Proton (Thermo Fisher Scientific, Waltham, MA, USA).”
“All analyses were performed with R software, version 4.0.”
“Patients using an artificial regimen were given vaginal progesterone gel at a dose of 90 mg once daily (Crinone, Merck Serono, Germany), or at a dose of 200 mg three times daily (Utrogestan, Belsins, Belgium)”
Statistical tests are named (Wilcoxon rank sum, chi-square, binomial regression, Cox proportional hazards). Assumptions are addressed through non-parametric tests and Schoenfeld residuals. Exact p-values are reported. Effect sizes with 95% CIs are provided. Software is identified. Data presentation includes per-group n and appropriate figures. Mathematical plausibility checks were not performed due to large N and continuous outcomes.
“comparisons between groups were analysed using the Wilcoxon rank sum test because of the non-normality of the variables.”
“odds ratio 1.09 (95% confidence interval (CI) 0.76 to 1.58)”
“comparisons between groups were analysed using the Wilcoxon rank sum test because of the non-normality of the variables.”
“odds ratio 1.09 (95% confidence interval (CI) 0.76 to 1.58)”
The data availability statement provides a concrete route: the dataset and code are on Open Science Framework with a URL. Access is governed by a data use agreement, which is appropriate for patient data. Repository deposit and accession numbers are not applicable for identifiable patient data, but the OSF link serves as the repository.
“The dataset and code for analysis can be found on Open Science Framework website ( https://osf.io/g5z2q/?view_only=1f5a8947b50d4fbe9b534433a6a51f22 ).”
“The dataset and code for analysis can be found on Open Science Framework website ( https://osf.io/g5z2q/?view_only=1f5a8947b50d4fbe9b534433a6a51f22 ).”
“Access is governed by a data use agreement to ensure appropriate use and privacy protection.”
The trial is registered (NCT02941965). A CONSORT diagram is included. All pre-specified outcomes are reported, including negative results. Limitations are discussed. Conclusions are proportional to the evidence. Funding and COI statements are provided.
“ClinicalTrials.gov NCT02941965”
“Fig 1 Trial profile. CONSORT (consolidated standards of reporting trials) diagram”
“However, limitations remain: firstly, male infertility should be diagnosed based on a comprehensive evaluation”
“Trial registration ClinicalTrials.gov NCT02941965”
“Fig 1 Trial profile. CONSORT (consolidated standards of reporting trials) diagram”
“However, limitations remain: firstly, male infertility should be diagnosed based on a comprehensive evaluation”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 33 references by DOI: 30 verified — 3 no DOI (shown, not verified).
- NO DOIWHO laboratory manual for the examination and processing of human semenNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDeveloping a core outcome set for future infertility research: an international consensus development studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPreimplantation genetic testing for aneuploidies (abnormal number of chromosomes) in in vitro fertilisationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- dataOSFLIVEHTTP 200https://osf.io/g5z2q/?view_only=1f5a8947b50d4fbe9b534433a6a51f22Resolves to OSF (data repository).
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, grammar, typo.
- MINORconsistencyAbstract, Results“ISCI=intracytoplasmic sperm injection”→ Change 'ISCI' to 'ICSI' for consistency.Typo in figure legend.
- MINORgrammarMethods, Randomisation and masking“After we obtained written informed from the eligible couples”→ Change to 'After we obtained written informed consent from the eligible couples'.Missing word 'consent'.
- MINORconsistencyResults, paragraph 2“The median numbers of retrieved oocytes were 14 (IQR 9-20) in the PGT-A group and 12 (9-16) in the no PGT-A group”→ Ensure consistent use of 'median number' (singular) or 'median numbers' (plural) throughout.Minor grammatical inconsistency.
- MINORtypoAbstract, Results“ISCI”→ ICSITypo in figure legend.
- MINORconsistencyMethods, Statistical analysis“We ultimately recruited 450 women (225 per group) according to our study protocol, and evaluated the sample size (n=450) for power because of a co-primary endpoint design”→ Clarify that the sample size was re-evaluated for the co-primary endpoint design.The sentence is slightly confusing; consider rephrasing.
- MINORclarityResults, paragraph 3“Eighty nine couples (39/191 (20.4%), PGT-A; 50/211 (23.7%), no PGT-A) underwent two transfer cycles”→ Consider rephrasing to 'Of the couples who underwent transfer, 39/191 (20.4%) in the PGT-A group and 50/211 (23.7%) in the no PGT-A group underwent two transfer cycles.'The sentence is a bit awkward.
The published work is robust and well-reported. An informed reader should weigh the open-label design and the exploratory subgroup finding in overweight women, but these are adequately discussed. No erratum or re-analysis is warranted based on this audit.
- 1.MEDIUMcopyeditIn the Abstract and figure legend, correct the typo 'ISCI' to 'ICSI'.Consistency in terminology is essential for professional presentation.
- 2.MEDIUMcopyeditIn Methods, Randomisation and masking, add the missing word 'consent' to 'written informed'.Grammar error that could confuse readers.
- 3.MEDIUMcopyeditIn Results, paragraph 2, ensure consistent use of 'median number' (singular) or 'median numbers' (plural).Minor grammatical inconsistency.
- 4.MEDIUMcopyeditIn Methods, Statistical analysis, clarify the sentence about sample size re-evaluation for the co-primary endpoint design.The current phrasing is confusing and could be misinterpreted.
- 5.MEDIUMcopyeditIn Results, paragraph 3, rephrase the sentence about couples undergoing two transfer cycles for clarity.The current sentence is awkward and could be clearer.
- 6.LOWreportingIn the Introduction, explicitly acknowledge limitations of prior retrospective studies (e.g., selection bias, lack of control).Strengthens the scientific premise by addressing limitations of prior work.
- 7.LOWreportingIn Methods, provide a brief rationale for the open-label design beyond 'nature of the interventions'.Addresses potential bias concerns and improves transparency.
- 8.LOWreportingIn the Discussion, discuss the generalizability of findings to other populations and settings.Helps readers interpret the applicability of the results.
- 9.LOWdata codeIn the Data Availability Statement, clarify the process for requesting data access (e.g., timeline, criteria).Enhances transparency and usability of the shared data.
- 10.LOWreportingIn the Results, report the number of participants with missing data for each outcome.Improves completeness and allows readers to assess potential bias.
- 11.LOWstatisticsIn the Statistical Analysis, specify the exact version of R and any packages used.Improves reproducibility.
- 12.LOWreportingIn the Discussion, avoid overemphasizing the subgroup finding in overweight women as it was exploratory and not pre-specified.Prevents overinterpretation of exploratory analyses.
- 13.LOWreportingIn the Methods, provide more detail on the embryo biopsy procedure and quality control measures.Enhances methodological transparency.
- 14.LOWreportingIn the Results, consider presenting the per-protocol analysis results in the main text.Transparency about sensitivity analyses.
- 15.LOWreportingIn the Discussion, discuss the potential impact of the funding limitation (only covering 1-3 embryos) on the results.Contextualizes a potential source of bias.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.