Deep learning versus manual morphology-based embryo selection in IVF: a randomized, double-blind noninferiority trial.
Illingworth PJ, Venetis C, Gardner DK, Nelson SM, Berntsen J, Larman MG, Agresta F, Ahitan S, Ahlström A, Cattrall F, Cooke S, Demmers K, Gabrielsen A, Hindkjær J, Kelley RL, Knight C, Lee L, Lahoud R, Mangat M, Park H, Price A, Trew G, Troest B, Vincent A, Wennerström S, Zujovic L, Hardarson T
- DOI
- 10.1038/s41591-024-03166-5
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/c11545a4-77ad-484b-8849-a43fc65a0416 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 17 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is clinical pregnancy, which is a surrogate for the ultimate outcome of live birth. The paper does not provide evidence that clinical pregnancy is a validated surrogate for live birth in this context, nor does it demonstrate target engagement for the deep learning algorithm at the tested dose. The algorithm's output (iDAScore) is a predictive score, and the trial does not establish a validated link between this surrogate and the clinical outcome beyond the trial's own results.
“The primary outcome for this study was the achievement of clinical pregnancy after the first embryo transfer... We defined clinical pregnancy as an intrauterine gestation with a fetal heartbeat observed after 7–9 weeks gestation.”
- 02Treatment effect not shown to be clinically meaningful
The primary outcome shows a risk difference of -1.7 percentage points (95% CI -7.7 to 4.3) for clinical pregnancy, which is not statistically significant and does not meet the noninferiority margin. The effect size is small and not anchored to a clinically meaningful difference, as the trial failed to demonstrate noninferiority. The paper does not provide an anchor to clinical meaningfulness for the observed difference.
“Noninferiority of embryo selection using the deep learning algorithm was not shown, with an absolute risk difference of −1.7 percentage points (95% confidence interval (CI), −7.7, 4.3) and a rate ratio of 0.96 (95% CI, 0.85, 1.10).”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted, rigorously reported multicenter RCT comparing deep learning versus standard morphology for embryo selection in IVF. The main methodological strengths are the randomized double-blind design, pre-specified analysis populations, and transparent reporting. The primary weakness is the vague data availability statement and proprietary code, which limit reproducibility.
Both reviewers classified the study as interventional, and this was adopted. The evaluation covered all eight dimensions; several sub-criteria were marked not applicable (e.g., animal-related items, cell line authentication) due to the human clinical trial context. The statistics verification recomputed 7 tests, all consistent; the citation check found no retracted or missing references.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 7 tests: 7 consistent, 0 inconsistent; 7 via agent-written checks.
- CONSISTENTreported p = .620 · recomputed p = .624Reviewer 1Primary outcome: clinical pregnancy rate comparison between study and control groups using Fisher's exact test.
“The primary outcome, clinical pregnancy, occurred in 248 of 533 patients (46.5%) in the study group and in 257 of 533 patients (48.2%) in the control group.”
Taken as given: The numbers 248 and 257 are the event counts in the study and control groups, respectively.; The numbers 533 and 533 are the total numbers in each group.; The test used is Fisher's exact test (two-sided) as stated in Table 2 footnote.Method: Two-sided Fisher's exact test on the 2x2 table (248, 285, 257, 276).How we recomputed it: pFisher2x2(248, 285, 257, 276, 0) - CONSISTENTreported p = .240 · recomputed p = .238Reviewer 1Live birth rate comparison between study and control groups using Fisher's exact test.
“Live birth rate was similar between the two groups where the live birth rate for the iDAScore group was 39.8% (212 of 533 patients) and 43.5% (232 of 533 patients) in the standard morphology criteria group (risk difference −3.9%; 95% CI, −9.9, 2.2; P = 0.24)”
Taken as given: The numbers 212 and 232 are the live birth counts in the study and control groups, respectively.; The numbers 533 and 533 are the total numbers in each group.; The test used is Fisher's exact test (two-sided) as implied by the primary analysis.Method: Two-sided Fisher's exact test on the 2x2 table (212, 321, 232, 301).How we recomputed it: pFisher2x2(212, 321, 232, 301, 0) - CONSISTENTreported p = .620 · recomputed p = .772Reviewer 1Subgroup analysis: clinical pregnancy in women >35 years, comparison between groups.
“In the prespecified subgroup analysis where the results in only women >35 years of age were compared, clinical pregnancy occurred in 85 of 228 (37.3%) in the study group and in 89 of 228 (39.0%) in the control group, with a risk difference between the groups of −1.8 (95% CI, −11.1, 7.6).”
Taken as given: The numbers 85 and 89 are the event counts in the study and control groups, respectively.; The numbers 228 and 228 are the total numbers in each group.; The test used is Fisher's exact test (two-sided) as implied by the primary analysis.Method: Two-sided Fisher's exact test on the 2x2 table (85, 143, 89, 139).How we recomputed it: pFisher2x2(85, 143, 89, 139, 0) - CONSISTENTreported p = .620 · recomputed p = .624Reviewer 2Primary outcome: clinical pregnancy rate comparison (ITT) using Fisher's exact test.
“Intention to treat | 248 of 533 (46.5%) (42.2%–50.9%) | 257 of 533 (48.2%) (43.9%–52.6%) | 0.62”
Taken as given: The numbers 248 and 257 are the event counts in the study and control groups, respectively.; The numbers 533 and 533 are the total group sizes.; The test is two-sided Fisher's exact test as stated in the table footnote.Method: Two-sided Fisher's exact test on the 2x2 table (248, 285, 257, 276).How we recomputed it: pFisher2x2(248, 285, 257, 276, 0) - CONSISTENTreported p = .240 · recomputed p = .238Reviewer 2Secondary outcome: live birth rate comparison using Fisher's exact test.
“Live birth rate was similar between the two groups where the live birth rate for the iDAScore group was 39.8% (212 of 533 patients) and 43.5% (232 of 533 patients) in the standard morphology criteria group (risk difference −3.9%; 95% CI, −9.9, 2.2; P = 0.24)”
Taken as given: The numbers 212 and 232 are the live birth counts in the study and control groups, respectively.; The numbers 533 and 533 are the total group sizes.; The test is two-sided Fisher's exact test.Method: Two-sided Fisher's exact test on the 2x2 table (212, 321, 232, 301).How we recomputed it: pFisher2x2(212, 321, 232, 301, 0) - CONSISTENTreported p = .350 · recomputed p = .354Reviewer 2Subgroup analysis: fresh transfer clinical pregnancy comparison.
“Clinical pregnancy rates were comparable in fresh cycles, with the study and control groups resulting in 48.1% (154 of 320) and 44.5% (157 of 353) ( P = 0.35).”
Taken as given: The numbers 154 and 157 are the clinical pregnancy counts in the study and control groups, respectively.; The numbers 320 and 353 are the total group sizes.; The test is two-sided Fisher's exact test.Method: Two-sided Fisher's exact test on the 2x2 table (154, 166, 157, 196).How we recomputed it: pFisher2x2(154, 166, 157, 196, 0) - CONSISTENTreported p = .032 · recomputed p = .032Reviewer 2Subgroup analysis: freeze-all transfer clinical pregnancy comparison.
“By contrast, the freeze-all cycles showed a significant difference ( P = 0.032) with the study and control groups resulting in 49.5% (94 of 190) and 61.3% (100 of 163) (Supplementary Table ).”
Taken as given: The numbers 94 and 100 are the clinical pregnancy counts in the study and control groups, respectively.; The numbers 190 and 163 are the total group sizes.; The test is two-sided Fisher's exact test.Method: Two-sided Fisher's exact test on the 2x2 table (94, 96, 100, 63).How we recomputed it: pFisher2x2(94, 96, 100, 63, 0)
- lowinternal contradictionThe abstract states 1,066 patients were included, but the results section mentions 1,002 in the per-protocol analysis after excluding 64. This is consistent, but the patient flow diagram is referenced but not shown in the text.
“Of these, 1,002 participants were included in the per-protocol (PP) analysis after the exclusion of 64 participants, owing to protocol violations (Patient flow diagram; Fig. ).”
Results ¶1Find in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2The deep learning approach provided a consistent user-independent approach.The paper shows reduced evaluation time and concordance in embryo selection, but does not directly measure consistency or user-independence beyond time.Evidence: Concordance analysis: same embryo selected in 65.8% of cases; time reduction.
“However, the deep learning approach studied did provide a consistent user-independent approach with a 10-fold reduction in assessment time.”
DiscussionFind in source - partialReviewers 1, 2The higher pregnancy rates observed in both groups may be due to the Hawthorne effect or rigorous assessment protocol.This is speculative; the paper offers no direct evidence for the Hawthorne effect, but it is a plausible explanation.Evidence: Discussion of higher-than-expected pregnancy rates compared to national datasets.
“The higher pregnancy rates observed in both groups, surpassing typical rates reported in US, European and Australian national datasets , may be a result of the participation in an RCT environment (the Hawthorne effect ).”
Discussion ¶3Find in source - supportedReviewers 1, 2The study was not able to demonstrate noninferiority of deep learning for clinical pregnancy rate compared to standard morphology.The primary outcome analysis shows the lower bound of the 95% CI for the risk difference (-7.7%) exceeds the noninferiority margin (-5%), so noninferiority is not demonstrated.Evidence: Primary outcome: risk difference -1.7% (95% CI -7.7, 4.3), with noninferiority margin -5%.
“This study was not able to demonstrate noninferiority of deep learning for clinical pregnancy rate when compared to standard morphology and a predefined prioritization scheme.”
AbstractFind in source - supportedReviewers 1, 2Deep learning significantly accelerates evaluation times compared to standard morphology.The time substudy shows a large and statistically significant reduction in evaluation time (21.3 vs 208.3 seconds, P<0.001).Evidence: Time substudy: mean time 21.3 ± 18.1 seconds vs 208.3 ± 144.7 seconds, P<0.001.
“The study group was found to have an almost 10-fold reduction in the time required for evaluation, with a mean standard deviation time of 21.3 ± 18.1 seconds in comparison to the control group, which took 208.3 ± 144.7 seconds ( P < 0.001)”
ResultsFind in source - supportedReviewer 2The deep learning algorithm underperformed in frozen-embryo transfers compared to fresh transfers.The subgroup analysis shows a significant difference in freeze-all cycles, but the paper notes this is a post hoc observation and may be due to chance.Evidence: Freeze-all: 49.5% vs 61.3%, P=0.032; fresh: 48.1% vs 44.5%, P=0.35.
“By contrast, the freeze-all cycles showed a significant difference ( P = 0.032) with the study and control groups resulting in 49.5% (94 of 190) and 61.3% (100 of 163)”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is clinical pregnancy, which is a surrogate for the ultimate outcome of live birth. The paper does not provide evidence that clinical pregnancy is a validated surrogate for live birth in this context, nor does it demonstrate target engagement for the deep learning algorithm at the tested dose. The algorithm's output (iDAScore) is a predictive score, and the trial does not establish a validated link between this surrogate and the clinical outcome beyond the trial's own results.
“The primary outcome for this study was the achievement of clinical pregnancy after the first embryo transfer... We defined clinical pregnancy as an intrauterine gestation with a fetal heartbeat observed after 7–9 weeks gestation.”
- INADEQUATEEffect sizeThe primary outcome shows a risk difference of -1.7 percentage points (95% CI -7.7 to 4.3) for clinical pregnancy, which is not statistically significant and does not meet the noninferiority margin. The effect size is small and not anchored to a clinically meaningful difference, as the trial failed to demonstrate noninferiority. The paper does not provide an anchor to clinical meaningfulness for the observed difference.
“Noninferiority of embryo selection using the deep learning algorithm was not shown, with an absolute risk difference of −1.7 percentage points (95% confidence interval (CI), −7.7, 4.3) and a rate ratio of 0.96 (95% CI, 0.85, 1.10).”
Data authenticity concerns
2 findings · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Methods and results do not matchAssessed
- Other integrity concernAssessed
3 integrity concerns flagged (0 high).
- lowmethod result mismatchThe paper reports a significant interaction between transfer type (fresh vs freeze-all) and treatment effect (P=0.022), but this is a post hoc analysis and the authors acknowledge it could be due to chance.
“In a further prespecified interaction analysis, we observed differing performance between the study and control groups in fresh-embryo transfer and freeze-all cycles (RR, 1.08 versus RR, 0.81; P = 0.022).”
ResultsFind in source - lowotherThe paper reports a significant interaction between treatment and transfer type (fresh vs frozen) with P=0.022, but this is a post hoc analysis and should be interpreted cautiously.
“In a further prespecified interaction analysis, we observed differing performance between the study and control groups in fresh-embryo transfer and freeze-all cycles (RR, 1.08 versus RR, 0.81; P = 0.022).”
ResultsFind in source
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior studies on AI in embryo selection, acknowledges the lack of RCTs for deep learning, and justifies the need for a noninferiority trial. The rationale is well-linked to the study objectives, and limitations of prior research (retrospective designs, small samples) are addressed by the prospective RCT design.
“A number of studies of AI have suggested significant improvements in consistency of embryo scoring , , but these evaluations have been retrospective. Two previous studies , have prospectively examined the efficacy of traditional machine learning algorithms based on previous data and human input for embryo selection.”
“Despite the potential utility of deep learning, there is a need to demonstrate noninferiority to standard embryologist assessment of the embryos.”
“Before this study, the performance of AI algorithms for blastocyst transfer and their impact on clinical pregnancy outcomes had not been directly compared to standard morphological criteria used by embryologists in a prospective RCT setting.”
“A number of studies of AI have suggested significant improvements in consistency of embryo scoring , , but these evaluations have been retrospective. Two previous studies , have prospectively examined the efficacy of traditional machine learning algorithms based on previous data and human input for embryo selection.”
“We therefore undertook a prospective randomized noninferiority trial to determine whether selection of a single blastocyst for transfer by deep learning results in a clinical pregnancy rate comparable with a 5% inferiority margin to that achieved by trained embryologists using standard morphology criteria.”
“No RCTs of deep learning in embryo selection have previously been reported.”
Randomization was performed via web-based interactive response technology, ensuring allocation concealment. Blinding was double-blind for patients and clinicians, though embryologists were unblinded by necessity. A priori power analysis was conducted for both the primary outcome and the time substudy. Inclusion/exclusion criteria were pre-specified. The analysis populations (ITT, PP, FAS) were defined. Outlier handling is addressed through the pre-specified analysis populations and sensitivity analyses.
“The randomization was performed within the eCRF using web-based interactive response technology.”
“The trial-group assignment was performed in an unblinded manner to the embryologist, but both the treating clinician and the patient remained blinded to the randomization outcome until after the first embryo transfer.”
“To demonstrate with 90% power ( α = 0.05 and β = 0.10) that the lower limit of the two-sided 95% CI for the difference between the iDAScore and the standard morphology criteria group would not be less than −5%, with an expected increase in clinical pregnancy of 5% or more in the iDAScore group, we required 494 women per group.”
“The randomization was performed within the eCRF using web-based interactive response technology.”
“The trial-group assignment was performed in an unblinded manner to the embryologist, but both the treating clinician and the patient remained blinded to the randomization outcome until after the first embryo transfer.”
“To demonstrate with 90% power ( α = 0.05 and β = 0.10) that the lower limit of the two-sided 95% CI for the difference between the iDAScore and the standard morphology criteria group would not be less than −5%, with an expected increase in clinical pregnancy of 5% or more in the iDAScore group, we required 494 women per group.”
Sex is reported (all female participants), age and BMI are reported in Table 1, and demographics are detailed. Since the study includes only women, sex_justified is not applicable as it is a single-sex study by design (IVF). Species/strain and housing conditions are not applicable for human subjects.
“Maternal age | 33.7 (3.7) 34 (22; 43) | 33.8 (3.8) 34 (20; 42) |”
“Reason for infertility (couple) | | No clinical subfertility | 44 (8.3%) | 42 (7.9%) |”
“Maternal age | 33.7 (3.7) 34 (22; 43) | 33.8 (3.8) 34 (20; 42)”
“The demographics and clinical characteristics of the patients at the time of randomization were well balanced between the trial groups, except from the type of transfer, with an extra 7.5% of patients undergoing a frozen–thawed embryo transfer in the study group compared to the control group (Table ).”
The paper states that each participating clinic's Human Research Ethics Committee approved the trial, and written informed consent was obtained from all participants. Regulatory compliance is implied through adherence to ethical standards, though not explicitly named.
“Each participating clinic’s respective Human Research Ethics Committee reviewed and approved the trial protocol”
“Written informed consent was obtained from all participants.”
“Each participating clinic’s respective Human Research Ethics Committee reviewed and approved the trial protocol”
“Written informed consent was obtained from all participants.”
“Australian New Zealand Clinical Trials Registry (ANZCTR) registration: 379161”
The iDAScore algorithm is described with version (v.1.2) and manufacturer (Vitrolife). The EmbryoScope incubator and EmbryoGlue are named. Statistical software (SAS v.9.4) is identified. Since this is a device/software trial, antibodies, cell lines, and mycoplasma testing are not applicable.
“iDAScore v.1.2 was performed to generate a score for each embryo”
“All the statistical analyses were executed using SAS v.9.4”
“The iDAScore v.1 algorithm was trained and evaluated based on a large dataset from 18 IVF centers consisting of 115,832 embryos, of which 14,644 were transferred embryos with known outcome .”
“All the statistical analyses were executed using SAS v.9.4”
“All blastocysts were transferred in EmbryoGlue (Vitrolife).”
Tests are named (Fisher's exact test, Farrington-Manning test, logistic regression). Assumptions are handled by design (non-parametric tests, adjusted analyses). Exact p-values are reported (e.g., P = 0.62). Effect sizes with 95% CIs are provided. Software is identified. Data presentation includes per-group n and confidence intervals. Mathematical plausibility checks were not performed due to large N and continuous outcomes.
“For comparison between groups, Fisher’s exact test (2-sided) was used.”
“absolute risk difference of −1.7 percentage points (95% confidence interval (CI), −7.7, 4.3)”
“For comparison between groups, Fisher’s exact test (2-sided) was used.”
“absolute risk difference of −1.7 percentage points (95% confidence interval (CI), −7.7, 4.3)”
The data availability statement says data will be made available on request to academic researchers after review by a steering committee, but does not specify a repository or platform. The code is proprietary and not shared. Since this is a clinical trial with patient data, repository deposit and accession numbers are not applicable, but the data access mechanism is not fully concrete.
“The data and documents will be made available on request to academic researchers following review by the study’s steering committee (P.J.I., C.V., D.K.G., S.M.N., J.B., M.G.L. and T.H.) and completion of a data sharing agreement.”
“The data and documents will be made available on request to academic researchers following review by the study’s steering committee (P.J.I., C.V., D.K.G., S.M.N., J.B., M.G.L. and T.H.) and completion of a data sharing agreement.”
The trial is registered with ANZCTR (379161). Methods are detailed enough for replication. Limitations are discussed. Conclusions are proportional to the evidence. Funding and competing interests are disclosed. Reporting guideline adherence is implied through Nature Portfolio reporting summary.
“Australian New Zealand Clinical Trials Registry (ANZCTR) registration: 379161”
“It is important to acknowledge several limitations in our trial.”
“The project was funded by a grant from Vitrolife.”
“Australian New Zealand Clinical Trials Registry (ANZCTR) registration: 379161”
“It is important to acknowledge several limitations in our trial. First, iDAScore was derived and tested solely within the context of the EmbryoScope incubator, limiting its generalizability to other time-lapse incubator systems.”
“The project was funded by a grant from Vitrolife.”
Registration stated in text, but no registry ID was detected. No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 34 references by DOI: 33 verified — 1 no DOI (shown, not verified).
- NO DOITowards Reproductive Certainty: Infertility and Genetics BeyondNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- codehttp://www.vitrolife.comLIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, clarity, grammar.
- MINORconsistencyAbstract“1,066 patients (533 in the iDAScore group and 533 in the morphology group)”→ Ensure consistent terminology: use 'study group' and 'control group' consistently.The abstract uses 'iDAScore group' and 'morphology group' while the rest of the paper uses 'study group' and 'control group'.
- MINORclarityResults, Secondary outcomes“The study group was found to have an almost 10-fold reduction in the time required for evaluation, with a mean standard deviation time of 21.3 ± 18.1 seconds”→ Rephrase to 'mean (standard deviation) time' for clarity.The phrase 'mean standard deviation time' is awkward; should be 'mean (SD) time'.
- MINORgrammarDiscussion, paragraph 3“The higher pregnancy rates observed could also be an outcome of the rigorous morphological assessment protocol employed.”→ Consider 'could also be a result of' for smoother phrasing.Minor grammatical improvement.
- MINORconsistencyAbstract“Australian New Zealand Clinical Trials Registry (ANZCTR) registration: 379161”→ Consider adding a space before the registration number for readability.Minor formatting.
- MINORclarityResults, Secondary outcomes“The study group was found to have an almost 10-fold reduction in the time required for evaluation, with a mean standard deviation time of 21.3 ± 18.1 seconds”→ Change 'mean standard deviation time' to 'mean time (standard deviation)' for clarity.Awkward phrasing.
- MINORgrammarDiscussion, paragraph 5“The higher pregnancy rates observed could also be an outcome of the rigorous morphological assessment protocol employed.”→ Consider rephrasing to 'could also be a result of' for smoother flow.Minor grammar.
The published work is methodologically robust and the findings are credible, but an informed reader should weigh the limited data/code availability and the post hoc nature of the fresh vs frozen subgroup analysis. No erratum appears necessary based on the checks performed, though the authors could improve transparency by providing a more concrete data access route.
- 1.HIGHdata codeIn the Data availability section, specify a concrete data access mechanism, such as a named repository (e.g., Zenodo) or a managed-access platform (e.g., Vivli), with conditions and timeframe, to move from 'on request' to a more concrete statement.The current statement is vague and does not provide a clear route for researchers to access the data, which is a reproducibility concern.
- 2.HIGHdata codeIn the Code availability section, consider providing a detailed description of the iDAScore algorithm's architecture and training procedure, or releasing a non-commercial version for academic use, to improve reproducibility.The proprietary code limits independent verification of the deep learning algorithm, which is central to the study.
- 3.HIGHreportingIn the Methods, explicitly state adherence to a reporting guideline such as CONSORT, and include the completed checklist as supplementary material.Explicit reporting guideline adherence enhances transparency and is expected for RCTs.
- 4.MEDIUMreportingIn the Discussion, clarify the post hoc nature of the fresh vs frozen subgroup analysis and avoid over-interpreting the interaction, as it was not pre-specified.The significant interaction (P=0.022) is post hoc and could be due to chance; readers should be cautioned against over-interpretation.
- 5.MEDIUMreportingIn the Results, consider reporting the exact p-value for the time evaluation comparison (currently reported as P < 0.001) to improve precision.Exact p-values are more informative than threshold-only values and align with the paper's overall reporting style.
- 6.MEDIUMcopyeditIn the Abstract, ensure consistent terminology: use 'study group' and 'control group' consistently instead of switching between 'iDAScore group' and 'morphology group'.Inconsistent terminology can confuse readers and detracts from the paper's professionalism.
- 7.MEDIUMcopyeditIn the Results, Secondary outcomes, rephrase 'mean standard deviation time' to 'mean time (standard deviation)' for clarity.The current phrasing is awkward and could be misinterpreted.
- 8.LOWcopyeditIn the Discussion, consider rephrasing 'could also be an outcome of' to 'could also be a result of' for smoother flow.Minor grammatical improvement for readability.
- 9.LOWcopyeditIn the Abstract, consider adding a space before the registration number for readability (e.g., 'registration: 379161').Minor formatting improvement.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.