AI-based triage and decision support in mammography and digital tomosynthesis for breast cancer screening: a paired, noninferiority trial.
Elías-Cabot E, Romero-Martín S, Raya-Povedano JL, Rodríguez-Ruiz A, Álvarez-Benito M
- DOI
- 10.1038/s41591-026-04277-x
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/9c55c09a-a45b-4e02-9c63-ca627e323559 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×6−3★
- ReportingStudy design partially met−0.25★
- CitationsUnresolved reference−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run on this paper: the pass that reads its reported means did not complete. No reported mean was checked for arithmetic impossibility.
- 01Internal contradictions in the reported numbers
The trial registration number is inconsistent across the manuscript: NCT04849776 (Abstract/Introduction), NCT04949776 (Results/Methods), and NTC04949776 (Methods typo). These cannot all be the correct identifier.
ClinicalTrials.gov: NCT04849776 ... ClinicalTrials.gov ID NCT04949776 ... (ClinicalTrials.gov ID: NTC04949776)
Abstractreviewer’s wording - 02Mathematically impossible statistic
The reported PPV p-value of 0.500 appears inconsistent with the data. Using a two-proportion z-test, the p-value is approximately 0.973, far from 0.500.
PPV of recall: 13.23 (11.63, 14.83) vs 13.19 (11.50, 14.90); P = 0.500
Table 2reviewer’s wording - 03Internal contradictions in the reported numbers
The clinical trial registration number is inconsistent: 'NCT04849776' in the abstract and 'NCT04949776' in the methods section.
Abstract: 'ClinicalTrials.gov: NCT04849776'; Methods: 'ClinicalTrials.gov ID: NCT04949776' and 'NTC04949776'
Abstractreviewer’s wording - 04Other integrity concern
The paper reports trial registration NCT04849776, but ClinicalTrials.gov has no record with that identifier. A registration that cannot be resolved does not support the claim that the trial was registered.
NCT04849776
reviewer’s wording - 05Methods and results do not match
The reported PPV P = 0.500 is inconsistent with a two-sided z-test on the given proportions (228/1,723 vs 198/1,501), which yields p ≈ 0.97. The authors should verify which test produced 0.500.
“PPV of recall | 228/1,723 | 13.23 | 198/1,501 | 13.19 | 0.04 | 0.31 | P = 0.500”
Table 2Find in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper reports a well-conducted prospective clinical trial with a strong scientific premise and adequate ethical approvals. However, several reporting gaps reduce its robustness: reader blinding is not addressed, missing data handling is unspecified, race/ethnicity and weight/health status are unreported, statistical analysis code is not shared, and the trial registration number is inconsistent and unresolvable. A potential discrepancy in the PPV p-value (0.500 vs ~0.97 from recomputation) requires investigation.
This assessment synthesizes three independent reviewer runs (all from the same model, AI) and a copyedit pass. The reviewers agreed on most dimensions, but there were disagreements on study design, biological variables, ethical approvals, statistical analysis, and data code availability. The synthesized statuses reflect the checklist-based scoring rules. The statistics verification component covered only 3 tests; the PPV p-value discrepancy was not verified by the component and is flagged by a reviewer's recomputation.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
- Mathematically impossible statisticAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Check the p-value for the CDR superiority test (McNemar) using the paired counts of cancers detected by each strategy.
“Table 2: CDR ... AI strategy: 228/31,301 ... Standard strategy: 198/31,301 ... Absolute difference: 1 ... P < 0.001 bc”
Taken as given: The 30 is the number of cancers detected by both strategies (from Extended Data Table 2: 198 detected by standard, 228 by AI, 252 total, so both = 198+228-252 = 174? Wait, recalc: 198+228-252 = 174, not 30. The paper states 24 cancers detected only by standard, 54 only by AI, so both = 252-24-54 = 174. The McNemar test uses discordant pairs: 24 (standard only) vs 54 (AI only). The p-value from McNemar test on discordant pairs (24 vs 54) is two-tailed.; The McNemar test is appropriate for paired binary data.; The p-value reported is from the superiority test (two-sided).Method: Two-tailed McNemar test using the discordant pair counts (24 vs 54) via pChi2x2 with the formula (b-c)^2/(b+c) = (54-24)^2/(54+24) = 900/78 ≈ 11.538, df=1, p = 1 - chi2Cdf(11.538, 1) ≈ 0.00068. The reported p < 0.001 is consistent.How we recomputed it: pFisher2x2(30, 198-30, 54-30, 31301-198-54+30, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 3CDR superiority, paired McNemar chi-square from discordant pair counts (54 cancers only by AI, 24 only by standard).
“Twenty-four cancers were detected only by the standard strategy ... there were 54 screen-detected cancers only detected by the AI strategy”
Taken as given: 54 and 24 are the discordant pair counts (detected only by AI and only by standard, respectively); McNemar chi-square without continuity correction approximates the paired test; df = 1 for a 2x2 paired comparisonMethod: Paired McNemar chi-square on discordant pairs, chi-square CDF with df=1; recomputed p ≈ 0.0007, consistent with the reported P < 0.001.How we recomputed it: pChi2(Math.pow(54-24,2)/(54+24), 1) - CONSISTENTreported p = .500 · recomputed p = .972Reviewer 3PPV of recall, two-sided z-test for equality of proportions (228/1,723 vs 198/1,501).
“PPV of recall | 228/1,723 | 13.23 | 198/1,501 | 13.19 | 0.04 | 0.31 | P = 0.500”
Taken as given: 228/1,723 and 198/1,501 are the PPV event counts and denominators for the AI and standard strategies; pooled proportion is (228+198)/(1723+1501); the paper's stated z-test is a two-sided test on the difference of proportions; the printed P is the two-tailed p-valueMethod: Two-proportion z-test with pooled SE, two-tailed normal CDF; recomputed p ≈ 0.97, which does not match the reported P = 0.500.How we recomputed it: pZ( (228/1723 - 198/1501) / Math.sqrt( (426/3224) * (1 - 426/3224) * (1/1723 + 1/1501) ) )
- mediuminternal contradictionThe trial registration number is inconsistent across the manuscript: NCT04849776 (Abstract/Introduction), NCT04949776 (Results/Methods), and NTC04949776 (Methods typo). These cannot all be the correct identifier.
ClinicalTrials.gov: NCT04849776 ... ClinicalTrials.gov ID NCT04949776 ... (ClinicalTrials.gov ID: NTC04949776)
Abstractreviewer’s wording - mediumimpossible statisticThe reported PPV p-value of 0.500 appears inconsistent with the data. Using a two-proportion z-test, the p-value is approximately 0.973, far from 0.500.
PPV of recall: 13.23 (11.63, 14.83) vs 13.19 (11.50, 14.90); P = 0.500
Table 2reviewer’s wording - mediuminternal contradictionThe clinical trial registration number is inconsistent: 'NCT04849776' in the abstract and 'NCT04949776' in the methods section.
Abstract: 'ClinicalTrials.gov: NCT04849776'; Methods: 'ClinicalTrials.gov ID: NCT04949776' and 'NTC04949776'
Abstractreviewer’s wording - lowinternal contradictionTable 3 reports the DBT workload as 4,820/13,986, but the DBT denominator throughout the paper (and the same table's other rows) is 13,968; the digits are transposed.
“Screening readings (workload) | 4,820/13,986 | 34.5”
Table 3Find in source
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
12 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewer 1AI triage and AI-supported screening could be a safe and effective screening strategy.The study demonstrates feasibility and noninferior CDR with increased RR, but the single-site design, single AI system, and lack of interval cancer data limit generalizability. The claim is supported for the specific setting but overstated as a general recommendation.Evidence: The study shows a 63.6% workload reduction, noninferior CDR, but increased RR. Limitations include single site, single vendor, single AI system, and no interval cancer data.
“AI triage and AI-supported screening that excludes low-risk mammograms from radiologist reading could be a safe and effective screening strategy that would allow for a substantial reduction in reading workload in breast cancer screening programs without negatively affecting cancer detection or the PPV of the recalled studies.”
DiscussionFind in source - partialReviewer 3It is safe to use AI to identify screening exams that can be automatically labeled as normal.CDR was noninferior, but the recall rate was higher and not noninferior (an additional 222 recalls), so the blanket 'safe' wording is only partially supported; the paper itself acknowledges this tension.Evidence: CDR noninferior/superior but RR +14.8% and not noninferior; 11 cancers missed by the AI strategy.
“demonstrate that in a prospective setting it is safe to use AI to identify screening exams that can be automatically labeled as normal and avoid radiologist reading”
Discussion ¶1Find in source - supportedReviewer 1AI-based triage and decision support reduces radiologist workload by 63.6%.The claim is directly supported by the primary outcome data: workload readings reduced from 31,301 to 11,384, a 63.6% reduction (95% CI 64.2% to 63.1%).Evidence: Table 2: Screening readings (workload) AI strategy: 11,384/31,301 (36.4%) vs Standard strategy: 31,301/31,301 (100%), absolute difference −19,917, relative difference −63.6%.
“Screening readings (workload) | 11,384/31,301 | 36.4 | 31,301/31,301 | 100 | −19,917 | −63.6”
Table 2Find in source - supportedReviewer 1The AI strategy is noninferior and superior to standard double reading in cancer detection rate.The CDR was 7.3 vs 6.3 per 1,000, a 15.2% relative increase. Noninferiority was demonstrated (lower bound of one-sided 97.5% CI above the −5% margin), and superiority was confirmed (P < 0.001).Evidence: Table 2: CDR AI strategy: 7.3 (95% CI 6.3, 8.2) vs Standard: 6.3 (95% CI 5.5, 7.2), absolute difference 1.0 (95% CI 0.4, 1.5), relative difference 15.2% (95% CI 6.6%, 24.4%), P < 0.001.
“the CDR in the AI strategy was noninferior and statistically superior to that of the standard strategy”
Table 2Find in source - supportedReviewer 1The recall rate in the AI strategy is not noninferior to the standard strategy.The RR was 5.5% vs 4.8%, a 14.8% relative increase. Noninferiority was not demonstrated (P = 0.997 for noninferiority test), as stated.Evidence: Table 2: RR AI strategy: 5.5% (95% CI 5.3%, 5.8%) vs Standard: 4.8% (95% CI 4.6%, 5.0%), absolute difference 0.7% (95% CI 0.4%, 1.0%), relative difference 14.8% (95% CI 9.0%, 20.6%), P = 0.997 for noninferiority.
“RR was not noninferior to the standard strategy”
Table 2Find in source - supportedReviewer 1The AI strategy detects more invasive carcinomas and carcinomas in situ than the standard strategy.The data show 174 vs 158 invasive carcinomas (10.1% increase, 95% CI 1.8%, 18.7%) and 54 vs 40 in situ carcinomas (35% increase, 95% CI 7.6%, 60.3%).Evidence: Table 4: Invasive carcinoma: AI 174, Standard 158, absolute difference 16, relative difference 10.1% (95% CI 1.8%, 18.7%). In situ carcinoma: AI 54, Standard 40, absolute difference 14, relative difference 35.0% (95% CI 7.6%, 60.3%).
“the AI strategy detected 10.1% more invasive carcinomas (95% CI 1.8%, 18.7%); 35% more carcinomas in situ (95% CI 7.6%, 60.3%)”
Table 4Find in source - supportedReviewer 1The AI strategy detects a higher proportion of grade I, T1, and N0 invasive carcinomas than the standard strategy.The data show relative increases of 30.2% for grade I, 13.5% for T1, and 15.6% for N0, all with 95% CIs excluding zero.Evidence: Table 4: Grade I: AI 56, Standard 43, relative difference 30.2% (95% CI 13.0%, 48.3%). T1: AI 126, Standard 111, relative difference 13.5% (95% CI 3.2%, 24.1%). N0: AI 141, Standard 122, relative difference 15.6% (95% CI 6.1%, 25.4%).
“a higher proportion of grade I invasive carcinomas (30.2% (95% CI 13.0%, 48.3%)), T1 (13.5% (95% CI 3.2%, 24.1%)) and N0 invasive carcinomas (15.6% (95% CI 6.1%, 25.4%)) than the standard strategy”
Table 4Find in source - supportedReviewer 1The AI strategy is safe, with no adverse events reported.The paper explicitly states 'No adverse events were reported in this study.' This is a direct statement of safety.Evidence: Results, Safety section: 'No adverse events were reported in this study.'
“No adverse events were reported in this study.”
ResultsFind in source - supportedReviewer 2AI could safely reduce workload by excluding low-risk exams from radiologist reading.The paper provides direct evidence of a 63.6% reduction in radiologist readings, with no reported adverse events.Evidence: Workload reduction from 62,602 to 22,768 readings; no adverse events reported.
“In the AI strategy, radiologist workload was 63.6% lower”
AbstractFind in source - supportedReviewer 2The cancer detection rate was 15.2% higher (95% CI 6.6%, 24.4%) in the AI strategy compared to the standard strategy, and was noninferior and statistically superior.The paper reports the CDR increase and shows noninferiority and superiority through statistical testing.Evidence: CDR: 7.3/1000 vs 6.3/1000; absolute difference 1.0/1000; P < 0.001 for superiority after noninferiority established.
“The cancer detection rate was 15.2% higher (95% confidence interval 6.6%, 24.4%), increasing from 6.3 of 1,000 to 7.3 of 1,000, P < 0.001”
AbstractFind in source - supportedReviewers 2, 3The recall rate was not noninferior and was 14.8% higher (95% CI 9.0%, 20.6%) in the AI strategy.The paper reports the RR increase and states that noninferiority was not demonstrated.Evidence: RR: 5.5% vs 4.8%; absolute difference 0.7%; P = 0.997 for noninferiority.
“The recall rate was not noninferior and was 14.8% higher (95% confidence interval 9.0%, 20.6%)”
AbstractFind in source - supportedReviewer 2Subanalyses by modality showed similar workload reduction in DM and DBT, but CDR and RR increased in DM while remaining stable in DBT.The paper provides subgroup results in Table 3, showing the differences between DM and DBT.Evidence: Workload reduction: -62.1% for DM, -65.5% for DBT; CDR increase 33.7% in DM, 0.9% in DBT; RR increase 28.2% in DM, -2.4% in DBT.
In the AI strategy with DM, the CDR was 33.7% (95% CI 19.8%, 50.5%) higher than in the standard strategy... In the AI strategy with DBT, the CDR and RR were similar to the standard strategy.
Table 3reviewer’s wording
Data authenticity concerns
2 findings · worst mediumAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Methods and results do not matchAssessed
- Other integrity concernAssessed
6 integrity concerns flagged (1 high).
- highotherThe paper reports trial registration NCT04849776, but ClinicalTrials.gov has no record with that identifier. A registration that cannot be resolved does not support the claim that the trial was registered.
NCT04849776
reviewer’s wording - mediummethod result mismatchThe reported PPV P = 0.500 is inconsistent with a two-sided z-test on the given proportions (228/1,723 vs 198/1,501), which yields p ≈ 0.97. The authors should verify which test produced 0.500.
“PPV of recall | 228/1,723 | 13.23 | 198/1,501 | 13.19 | 0.04 | 0.31 | P = 0.500”
Table 2Find in source
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Study-design details incomplete (controls, blinding, power)Assessed
The introduction cites numerous prior studies (refs. 1–22) establishing the background, including retrospective simulations and prospective trials. It acknowledges limitations of prior work (e.g., retrospective nature, need for prospective confirmation) and states how the current study addresses them (prospective, paired design, includes DBT). The hypothesis follows logically from the cited evidence.
“Over the last few years, several retrospective studies have shown that radiologists improved their cancer detection accuracy when using an AI system as concurrent reading support”
“Several prospective studies have been designed to overcome the limitations of retrospective studies and to address and quantify the real-life impact of AI in breast cancer screening”
“A common trend in the outcomes of many retrospective and prospective studies is the high performance accuracy of AI in identifying large numbers of low-risk breast screening exams with a very high negative predictive value.”
“Several prospective studies have been designed to overcome the limitations of retrospective studies and to address and quantify the real-life impact of AI in breast cancer screening”
“aims to confirm prospectively the results obtained in these studies by evaluating whether an AI system can be used to omit human reading completely in a large proportion of the screening exams classified as low risk for cancer”
“Several prospective studies have been designed to overcome the limitations of retrospective studies”
The study is a paired noninferiority trial with a prespecified sample size calculation and clear inclusion/exclusion criteria. However, the paper does not describe blinding of radiologists or provide a rationale for an open-label design. Additionally, there is no discussion of how outliers or missing data were handled, nor a statement on intention-to-treat vs per-protocol analysis.
“After signing informed consent, women were excluded from the clinical trial if they had symptoms or signs of suspected breast cancer; if they had breast prostheses; or if their images could not be processed by the AI system”
“Cases classified by the AI system with a score of 1 to 7 were considered low risk (approximately 70% of the studies) and automatically classified as normal.”
“After signing informed consent, women were excluded from the clinical trial if they had symptoms or signs of suspected breast cancer; if they had breast prostheses; or if their images could not be processed by the AI system (for example, because of the presence of visible breast implants or because of erroneous PACS image transfer).”
“the sample size needed was defined as 27,000 women in order to demonstrate the noninferiority of the AI strategy over the standard strategy globally”
“women were excluded from the clinical trial if they had symptoms or signs of suspected breast cancer; if they had breast prostheses; or if their images could not be processed by the AI system”
The study reports sex (100% women), age (median 59, IQR 54–64), breast density (BI-RADS categories A–D), and screening round (first vs. follow-up). Race/ethnicity is not individually recorded but the paper states 'the majority of the target population was Caucasian.' This is adequate for a screening trial in a defined geographic region.
“Women | 31,301 (100.0%)”
“The median age was 59 years old (interquartile range 64 to 54 years).”
“Race and ethnicity were not individually recorded, but the majority of the target population was Caucasian.”
“Women | 31,301 (100.0%)”
“The median age was 59 years old (interquartile range 64 to 54 years).”
“Race and ethnicity were not individually recorded, but the majority of the target population was Caucasian.”
The paper states that the study protocol and informed consent received a favorable ruling from the Institutional Review Board at Reina Sofía University Hospital of Córdoba Research Ethics Committee (IRB No 4932). All patients provided written informed consent. Compliance with HIPAA is stated. The trial is registered at ClinicalTrials.gov.
“In March 2021, the study protocol and the informed consent received a favorable ruling from the Institutional Review Board (IRB) at Reina Sofía University Hospital of Córdoba Research Ethics Committee (IRB No 4932).”
“All patients provided written informed consent before enrollment.”
“This prospective clinical trial was compliant with the Health Insurance Portability and Accountability Act”
“In March 2021, the study protocol and the informed consent received a favorable ruling from the Institutional Review Board (IRB) at Reina Sofía University Hospital of Córdoba Research Ethics Committee (IRB No 4932).”
“All patients provided written informed consent before enrollment.”
“This prospective clinical trial was compliant with the Health Insurance Portability and Accountability Act”
“the study protocol and the informed consent received a favorable ruling from the Institutional Review Board (IRB) at Reina Sofía University Hospital of Córdoba Research Ethics Committee (IRB No 4932)”
“All patients provided written informed consent before enrollment.”
“This prospective clinical trial was compliant with the Health Insurance Portability and Accountability Act”
The paper identifies the AI system by name, version, and vendor. The statistical software (R 4.4.2) and key libraries are listed. No antibodies, cell lines, or animal resources are used, so those sub-criteria are not applicable.
“This clinical trial used a commercially available AI system for breast cancer detection, Transpara (version 1.7 ScreenPoint Medical)”
“Images were acquired using four devices: three DM devices (Lorad Selenia, Hologic) and one DBT device (3Dimensions, Hologic).”
“The software used to perform the statistical analyses in this study was R, version 4.4.2.”
“The software used to perform the statistical analyses in this study was R, version 4.4.2.”
“This clinical trial used a commercially available AI system for breast cancer detection, Transpara (version 1.7 ScreenPoint Medical)”
“The software used to perform the statistical analyses in this study was R, version 4.4.2.”
The paper names statistical tests (McNemar test for paired comparisons, z-test for PPV, Wald statistics). Exact p-values are reported for CDR (P < 0.001) and RR (P = 0.997). Effect sizes with 95% CIs are provided for all primary and secondary outcomes. Statistical software (R 4.4.2) is identified. Data presentation includes proportions with 95% CIs and per-group n. Assumptions (normality, equal variance) are not explicitly verified, but the methods (McNemar, Wald) are standard for paired binary data and do not require those assumptions. Mathematical plausibility checks are not applicable for these aggregate statistics.
“P < 0.001 bc”
“For the CDR and RR, initially, one-tailed McNemar paired tests were applied for the noninferiority analyses.”
“increase in the CDR of 15.2% (95% CI 6.6%, 24.4%)”
The paper states that the individual deidentified participant dataset, data dictionary, protocol, and informed consent form are publicly accessible via Zenodo with a DOI. The proprietary AI algorithm code is not shared, but a technical description is provided. The statistical analysis code is not shared, but the use of standard R libraries is noted.
“Individual deidentified participant dataset, data dictionary defining each field, protocol and informed consent form-patients information sheet are publicly accessible via Zenodo at 10.5281/zenodo.17625633 (ref. ). Access ends 10 years following article publication.”
“The software used to perform the statistical analyses in this study was R, version 4.4.2. The main open-source libraries used for the analyses were: ggplot2 (for graphic purposes) and PropCIs (for confidence intervals estimation).”
“The code for training and developing the evaluated AI algorithm (Transpara version 1.7, ScreenPoint Medical) is part of a proprietary system.”
“Individual deidentified participant dataset, data dictionary defining each field, protocol and informed consent form-patients information sheet are publicly accessible via Zenodo at 10.5281/zenodo.17625633 (ref. ).”
“The code for training and developing the evaluated AI algorithm (Transpara version 1.7, ScreenPoint Medical) is part of a proprietary system.”
“Individual deidentified participant dataset, data dictionary defining each field, protocol and informed consent form-patients information sheet are publicly accessible via Zenodo at 10.5281/zenodo.17625633”
“The code for training and developing the evaluated AI algorithm (Transpara version 1.7, ScreenPoint Medical) is part of a proprietary system.”
The paper provides a comprehensive description of methods, including the AI system, study design, and statistical analysis. The trial is registered at ClinicalTrials.gov (NCT04849776). Limitations such as single-site, single-vendor, and the paired design are discussed. The conclusions are appropriately cautious, noting the need for further validation. Funding sources and competing interests are disclosed.
“ClinicalTrials.gov: NCT04849776 (http://clinicaltrials.gov/ct2/show/NCT04849776)”
“Our study has the limitation of being a single-site investigation with a screening workflow of double reading without arbitration by expert radiologists in breast imaging and breast screening who have several years of experience in the use of AI.”
“ClinicalTrials.gov: NCT04849776 (http://clinicaltrials.gov/ct2/show/NCT04849776)”
“Further information on research design is available in the linked to this article.”
“Our study has the limitation of being a single-site investigation with a screening workflow of double reading without arbitration”
“A.R.R. is an employee of ScreenPoint Medical. The other authors declare no competing interests.”
Registered (2 IDs: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 33 references by DOI: 28 verified — 1 DOI unresolved, 4 no DOI (shown, not verified).
- UNRESOLVED10.5281/zenodo.17625633Artificial Intelligence in Breast Cancer Screening Program in Cordoba (AITIC)Cited DOI does not resolve to any Crossref record.
- NO DOIEuropean Guidelines for Breast Cancer ScreeningNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITNM Classification of Malignant TumorsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBreast TumoursNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIACR BI-RADS Atlas, Breast Imaging Reporting and Data SystemNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
4 data/code links checked; 4 live.
- datahttp://clinicaltrials.gov/ct2/show/NCT04849776LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://classic.clinicaltrials.gov/ct2/show/NCT04949776LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttp://clinicaltrials.gov/ct2/show/NTC04949776LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT04949776LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
1 finding · worst lowWording, consistency and formatting errors that need correcting before submission.
- Wording or formatting errors that need correctingAssessed
13 copyedit issues flagged (2 major): mostly consistency, clarity, typo.
- MAJORconsistencyAbstract vs Methods“Abstract: 'ClinicalTrials.gov: NCT04849776'; Methods: 'ClinicalTrials.gov ID: NCT04949776' and 'NTC04949776'”→ Verify the correct trial registration number and use it consistently throughout the manuscript.The discrepancy between NCT04849776 and NCT04949776 could lead to confusion and affects data traceability.
- MAJORconsistencyAbstract / Results / Methods“ClinicalTrials.gov: NCT04849776 ... ClinicalTrials.gov ID NCT04949776 ... (ClinicalTrials.gov ID: NTC04949776)”→ Use a single, correct registration number throughout.NCT04849776 and NCT04949776 (and NTC04949776) are all used; one is a typo.
- MINORconsistencyAbstract“ClinicalTrials.gov: NCT04849776 (http://clinicaltrials.gov/ct2/show/NCT04849776)”→ Ensure the URL is consistent with the one in Methods (NCT04949776 vs NCT04849776).The abstract uses NCT04849776 while the Methods section uses NCT04949776; this appears to be a typo.
- MINORconsistencyMethods, Clinical trial design and patients“ClinicalTrials.gov ID: NTC04949776”→ Correct 'NTC' to 'NCT'.Typo in the ClinicalTrials.gov identifier prefix.
- MINORclarityResults, paragraph 1“The median age was 59 years old (interquartile range 64 to 54 years).”→ Rephrase to 'median age 59 years (interquartile range 54–64 years)' for standard presentation.The interquartile range is presented in descending order, which is unconventional.
- MINORclarityTable 2 footnote“For the CDR and RR, initially, one-tailed McNemar paired tests were applied for the noninferiority analyses. However, when noninferiority was achieved, the P value represented in the table is the one corresponding to the superiority analyses.”→ Clarify that the p-value for CDR (P < 0.001) is from the superiority test, not the noninferiority test.The footnote is clear but could be more explicit about which p-value is reported.
- MINORconsistencyTable 3“DBT | Screening readings (workload) | 4,820/13,986 | 34.5”→ Correct the denominator to 13,968 to match the total DBT participants.The denominator for DBT workload is 13,986 instead of 13,968; likely a typo.
- MINORclarityMethods, Statistical analyses“One-sided 97.5% CIs were established for the difference in proportions (AI standard) for the CDR and RR. These were derived from two-sided 95% CIs.”→ Clarify that the one-sided 97.5% CI corresponds to a one-sided alpha of 0.025, consistent with noninferiority testing.The explanation is technically correct but could be more accessible.
- MINORconsistencyExtended Data Fig. 1 and 2“Relative differences in CDR between the AI strategy and standard strategy for the whole population (N = 31301): 15.2 (6.6, 24.4)”→ Use commas in large numbers (31,301) for consistency with the main text.The figure captions omit the comma in the number 31,301.
- MINORconsistencyTable 3, DBT workload row“4,820/13,986”→ Change to 4,820/13,968.Denominator does not match the stated DBT total of 13,968.
- MINORclarityResults, paragraph 1“The median age was 59 years old (interquartile range 64 to 54 years).”→ Change to 'interquartile range 54 to 64 years'.IQR is written in descending order.
- MINORtypoMethods, Outcomes“The FPV was calculated as the number of women recalled who were not finally diagnosed with cancer”→ Change 'FPV' to 'FPR'.False-positive rate is abbreviated elsewhere as FPR.
- MINORconsistencyResults, Secondary outcomes“absolute difference of 0.04% (95% CI 2.3%, 2.4%)”→ Change to (95% CI −2.3%, 2.4%) to match Table 2.Negative sign missing; table shows (−2.3, 2.4).
The published paper is generally robust but has several issues that an informed reader should weigh: the trial registration number is inconsistent and one identifier is unresolvable, the PPV p-value appears inconsistent with the reported data, and several reporting details (blinding, missing data, code sharing) are missing. These issues do not necessarily invalidate the main conclusions, but they warrant a correction or independent verification, particularly for the PPV p-value and the registration number.
- 1.HIGHstatisticsInvestigate the discrepancy in the PPV p-value: the reported P=0.500 is inconsistent with a two-proportion z-test on the given counts (228/1723 vs 198/1501, which yields p≈0.97). Verify the test used, correct the p-value, or provide a detailed explanation for the discrepancy.A demonstrable error in a reported p-value undermines the credibility of the secondary analysis and may require a correction.
- 2.HIGHreportingCorrect the trial registration number: standardize to NCT04849776 or NCT04949776 (whichever is correct) throughout the manuscript, and ensure the identifier is resolvable on ClinicalTrials.gov. The current inconsistency (NCT04849776, NCT04949776, NTC04949776) is a serious reporting gap.An unresolvable or inconsistent trial registration number makes it impossible to verify pre-registration and undermines trust in the study's transparency.
- 3.HIGHreportingAdd a statement on reader blinding or explain why it was not applicable/feasible in the AI strategy.The absence of blinding information is a methodological gap that readers should be aware of.
- 4.HIGHreportingDescribe the handling of missing data and protocol deviations, including whether intention-to-treat or per-protocol analysis was used.Missing data handling is essential for assessing the robustness of the results.
- 5.HIGHdata codeDeposit the R analysis scripts used for statistical analyses (e.g., code for generating tables and figures) in a public repository (e.g., Zenodo or GitHub) to improve reproducibility.The current code availability is inadequate; sharing the analysis code allows independent verification.
- 6.MEDIUMreportingReport exact p-values instead of thresholds (e.g., replace 'P < 0.001' with the exact value, such as 'P = 0.00068') for all significance tests.Threshold p-values are less informative and deviate from best practices for reporting.
- 7.MEDIUMreportingInclude a statement on compliance with the Declaration of Helsinki or other relevant ethical framework in the Ethics approval section.The current statement only mentions HIPAA, which is a data privacy law, not an ethical guideline for human subjects research.
- 8.MEDIUMreportingReport the weight and health status (e.g., BMI, comorbidities) of participants, if available, or note that these data were not collected.Biological variables beyond age and sex are important for characterizing the study population.
- 9.MEDIUMreportingCorrect the typo in Table 3: change the DBT workload denominator from 13,986 to 13,968 to match the stated DBT total.The transposed digits introduce an inconsistency in the data presentation.
- 10.MEDIUMcopyeditReorder the interquartile range from '64 to 54 years' to '54 to 64 years' in the Results section.The IQR is presented in descending order, which is unconventional.
- 11.LOWcopyeditCorrect the abbreviation 'FPV' to 'FPR' in the Methods, Outcomes section.Inconsistent abbreviation may cause confusion.
- 12.LOWcopyeditAdd the missing negative sign in the CI for the absolute difference in the secondary outcome: change (95% CI 2.3%, 2.4%) to (95% CI -2.3%, 2.4%) to match Table 2.The missing sign creates a discrepancy with the table.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.