AI-based triage and decision support in mammography and digital tomosynthesis for breast cancer screening: a paired, noninferiority trial.
Elías-Cabot E, Romero-Martín S, Raya-Povedano JL, Rodríguez-Ruiz A, Álvarez-Benito M
- DOI
- 10.1038/s41591-026-04277-x
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/4c2d7704-1f83-4edb-ba55-91bffb50ffaf is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×4−2★
- ClaimsUnsupported claim (uncorroborated)−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- CitationsUnresolved reference−0.25★
- 01Conclusion not supported by the paper’s own evidence
The recall rate in the AI strategy is noninferior to the standard strategy.
“the recall rate was not noninferior and was 14.8% higher (95% confidence interval 9.0%, 20.6%)”
AbstractFind in source - 02Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is that the AI strategy improves cancer detection rate (CDR) and reduces workload. CDR is a surrogate for the clinical outcome of breast cancer mortality. The paper does not demonstrate target engagement at the tested dose (AI system) linking CDR to mortality, nor does it cite validated evidence that an increase in CDR translates to reduced mortality. The increase in CDR is presented as a benefit without establishing a validated surrogate-to-clinical-outcome link.
“the cancer detection rate was 15.2% higher (95% confidence interval 6.6%, 24.4%), increasing from 6.3 of 1,000 to 7.3 of 1,000, P < 0.001”
- 03Treatment effect not shown to be clinically meaningful
The primary reported effect is an increase in CDR from 6.3 to 7.3 per 1,000 (absolute difference of 1.0 per 1,000). This is a small absolute increase and is not anchored to a minimal clinically important difference or to a meaningful clinical outcome such as mortality reduction. The paper does not provide evidence that this magnitude of increase in CDR is clinically meaningful.
“an increase in the CDR of 15.2% (95% CI 6.6%, 24.4%), which is an absolute difference of 1.0 of 1,000 (95% CI 0.4 of 1,000, 1.5 of 1,000; P < 0.001)”
- 04Other integrity concern
The paper reports trial registration NCT04849776, but ClinicalTrials.gov has no record with that identifier. A registration that cannot be resolved does not support the claim that the trial was registered.
NCT04849776
reviewer’s wording
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper is a well-designed prospective, paired, noninferiority clinical trial evaluating an AI-based triage strategy in breast cancer screening. It demonstrates strong scientific premise, rigorous design, adequate reporting of biological variables, ethics, key resources, statistics, data availability, and transparency. However, there are minor reporting gaps (blinding details, exact p-values, reporting guideline) and a critical issue with the trial registration number that could not be resolved.
Both reviewers classified the study as interventional, which is adopted. The evaluation covered all eight dimensions, with several sub-criteria marked as not applicable (e.g., species/strain, housing, antibodies). The statistics verification checked only 1 test (consistent), and the citation check found 1 reference not found in registry (the Zenodo DOI). The integrity check flagged the trial registration number as unresolvable.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Check the p-value for the superiority test of CDR (AI vs standard) using McNemar test on the paired data.
“CDR | 228/31,301 | 7.3 | 198/31,301 | 6.3 | 1 | 15.2 | P < 0.001 bc”
Taken as given: The 228 and 198 are the number of cancers detected by AI and standard strategies, respectively.; The total N is 31,301.; The discordant pairs are 30 (cancers detected by AI but not standard) and 198 (cancers detected by standard but not AI) — this is inferred from the table showing 228 AI-detected and 198 standard-detected, with 252 total unique cancers, implying 30 cancers detected only by AI and 198 only by standard? Actually, the numbers are: AI detected 228, standard detected 198, total unique 252. So discordant pairs: AI+/standard- = 228 - (252-198) = 174? This is complex; the McNemar test requires the number of discordant pairs. The paper does not provide the exact 2x2 table of paired outcomes, so this check is approximate.; The test is two-sided for superiority.Method: McNemar test using the chi-square approximation with continuity correction, assuming the discordant pairs are (30, 198) as a rough approximation. This is not exact because the paper does not provide the full paired contingency table.How we recomputed it: pChi2x2(30, 198, 228, 30845)
- lowinternal contradictionThe abstract states 'Received 2025 Jan 12; Accepted 2026 Feb 4; Issue date 2026.' The acceptance date is after the study end date (January 2024), which is unusual but possible for a journal with a long review process.
“Received 2025 Jan 12; Accepted 2026 Feb 4; Issue date 2026.”
AbstractFind in source - lowinternal contradictionThe abstract states 31,301 women were included, but the Methods mention 31,856 agreed to participate and 555 were excluded, which sums to 31,301. This is consistent.
“A total of 31,856 women agreed to participate in this trial (AITIC, ClinicalTrials.gov ID NCT04949776 (https://clinicaltrials.gov/ct2/show/NCT04949776) ) and signed the informed consent form (the patient information sheet and informed consent form are included in Zenodo . Of these, 555 women (1.7%) were excluded because of the different reasons explained in Fig. , resulting in 31,301 women being included (17,333 DM and 13,968 DBT).”
AbstractFind in source - lowinternal contradictionIn Table 3, the denominator for DBT AI strategy screening readings is 13,986, but the total DBT N is 13,968. This is a minor inconsistency.
“Screening readings (workload) | 4,820/13,986 | 34.5 | 13,968/13,968 | 100 | −9,148 | −65.5”
Table 3Find in source
Overstated conclusions
4 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions not supported by the paper’s own evidenceAssessed
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated), 2 only partially supported (evidence backs part of the claim; gaps or caveats remain).
- unsupportedReviewers 1, 2The recall rate in the AI strategy is noninferior to the standard strategy.The paper explicitly states that noninferiority for RR was not demonstrated (P = 0.997). The claim in the abstract is that the recall rate was not noninferior.Evidence: Table 2 shows RR of 5.5% vs 4.8%, with noninferiority not demonstrated.
“the recall rate was not noninferior and was 14.8% higher (95% confidence interval 9.0%, 20.6%)”
AbstractFind in source - partialReviewer 1The AI strategy is safe and effective for partially automated screening.The claim of safety is supported by no adverse events and increased cancer detection, but the increased recall rate and missed cancers (11) warrant caution.Evidence: The paper reports no adverse events and increased CDR, but also notes 11 cancers missed by AI and increased RR.
“These results demonstrate the feasibility of a partially automated AI workflow in breast cancer screening, avoiding human reading of studies classified as low risk.”
DiscussionFind in source - partialReviewer 2The AI strategy is effective for both DM and DBT, with similar workload reduction.Workload reduction is similar (-62.1% for DM, -65.5% for DBT), but the impact on CDR and RR differs: CDR increased for DM but not for DBT, and RR increased for DM but not for DBT. The claim is partially supported as the workload reduction is consistent, but the effects on detection and recall differ by modality.Evidence: Table 3 shows modality-specific results.
“Subanalyses by modality highlighted a similar workload reduction in digital mammography (−62.1%) and digital breast tomosynthesis (−65.5%). However, in digital mammography, the cancer detection rate increased by 1.6 of 1,000 and the recall rate by 1.3%, while both remained stable in digital breast tomosynthesis.”
ResultsFind in source - supportedReviewers 1, 2AI-based triage reduces radiologist workload by 63.6%.The workload reduction is directly measured and reported with confidence intervals.Evidence: Table 2 reports workload readings of 11,384 vs 31,301, a relative reduction of -63.6%.
“In the AI strategy, radiologist workload was 63.6% lower”
AbstractFind in source - supportedReviewer 1AI strategy increases cancer detection rate by 15.2% compared to standard double reading.The increase is statistically significant and reported with confidence intervals.Evidence: Table 2 reports CDR of 7.3 vs 6.3 per 1,000, relative difference 15.2% (95% CI 6.6-24.4), P<0.001.
“the cancer detection rate was 15.2% higher (95% confidence interval 6.6%, 24.4%), increasing from 6.3 of 1,000 to 7.3 of 1,000, P < 0.001”
AbstractFind in source - supportedReviewer 1The AI strategy detects more invasive and in situ cancers than standard strategy.The paper reports higher detection of invasive and in situ cancers in the AI strategy, with statistical significance for some subgroups.Evidence: Table 4 shows AI detected 174 invasive vs 158, and 54 in situ vs 40, with relative differences of 10.1% and 35% respectively.
“Thus, the AI strategy detected 10.1% more invasive carcinomas (95% CI 1.8%, 18.7%); 35% more carcinomas in situ (95% CI 7.6%, 60.3%)”
ResultsFind in source - supportedReviewer 2AI-based triage and decision support in mammography and digital tomosynthesis for breast cancer screening is noninferior in cancer detection rate compared to standard double reading.The paper provides evidence that the CDR in the AI strategy was noninferior and statistically superior to the standard strategy (relative difference 15.2%, 95% CI 6.6% to 24.4%, P < 0.001).Evidence: Table 2 shows CDR of 7.3/1000 vs 6.3/1000, with noninferiority demonstrated.
“the cancer detection rate was 15.2% higher (95% confidence interval 6.6%, 24.4%), increasing from 6.3 of 1,000 to 7.3 of 1,000, P < 0.001”
AbstractFind in source - supportedReviewer 2The AI strategy is safe, with no adverse events reported.The paper reports that no adverse events were reported, and the safety monitoring plan is described.Evidence: Results, Safety section: 'No adverse events were reported in this study.'
“No adverse events were reported in this study.”
ResultsFind in source - supportedReviewer 2The AI strategy detects more invasive carcinomas and carcinomas in situ than the standard strategy.The paper reports that the AI strategy detected 10.1% more invasive carcinomas (95% CI 1.8%, 18.7%) and 35% more carcinomas in situ (95% CI 7.6%, 60.3%).Evidence: Table 4 shows these differences.
“the AI strategy detected 10.1% more invasive carcinomas (95% CI 1.8%, 18.7%); 35% more carcinomas in situ (95% CI 7.6%, 60.3%)”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is that the AI strategy improves cancer detection rate (CDR) and reduces workload. CDR is a surrogate for the clinical outcome of breast cancer mortality. The paper does not demonstrate target engagement at the tested dose (AI system) linking CDR to mortality, nor does it cite validated evidence that an increase in CDR translates to reduced mortality. The increase in CDR is presented as a benefit without establishing a validated surrogate-to-clinical-outcome link.
“the cancer detection rate was 15.2% higher (95% confidence interval 6.6%, 24.4%), increasing from 6.3 of 1,000 to 7.3 of 1,000, P < 0.001”
- INADEQUATEEffect sizeThe primary reported effect is an increase in CDR from 6.3 to 7.3 per 1,000 (absolute difference of 1.0 per 1,000). This is a small absolute increase and is not anchored to a minimal clinically important difference or to a meaningful clinical outcome such as mortality reduction. The paper does not provide evidence that this magnitude of increase in CDR is clinically meaningful.
“an increase in the CDR of 15.2% (95% CI 6.6%, 24.4%), which is an absolute difference of 1.0 of 1,000 (95% CI 0.4 of 1,000, 1.5 of 1,000; P < 0.001)”
Data authenticity concerns
1 finding · worst mediumAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
4 integrity concerns flagged (1 high).
- highotherThe paper reports trial registration NCT04849776, but ClinicalTrials.gov has no record with that identifier. A registration that cannot be resolved does not support the claim that the trial was registered.
NCT04849776
reviewer’s wording
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites multiple prior studies (retrospective and prospective) showing AI's potential to improve screening accuracy and reduce workload, and explicitly notes limitations of retrospective studies (e.g., 'simulated retrospective studies'). The premise is logically linked to the study objectives: to prospectively confirm that AI can safely omit human reading for low-risk exams. The paper also addresses how the current study overcomes prior limitations by being prospective and including DBT.
“Over the last few years, several retrospective studies have shown that radiologists improved their cancer detection accuracy when using an AI system as concurrent reading support – . Furthermore, AI systems implemented as a stand-alone solution have been found to achieve human-like performance in simulated retrospective studies – , classifying exams according to cancer probability and resulting in simulations of AI-based breast cancer strategies that suggested improved clinical outcomes ,– .”
“Several prospective studies have been designed to overcome the limitations of retrospective studies and to address and quantify the real-life impact of AI in breast cancer screening in which radiologists interact with AI.”
“The Artificial Intelligence in Breast Cancer Screening Program in Córdoba (AITIC) clinical trial is a prospective, paired, noninferiority, accuracy study ( NCT04849776 (http://clinicaltrials.gov/ct2/show/NCT04849776) ) that aims to confirm prospectively the results obtained in these studies by evaluating whether an AI system can be used to omit human reading completely in a large proportion of the screening exams classified as low risk for cancer, with noninferior results in the cancer detection rate (CDR) and recall rate (RR) compared to the standard strategy.”
“Several prospective studies have been designed to overcome the limitations of retrospective studies and to address and quantify the real-life impact of AI in breast cancer screening in which radiologists interact with AI.”
The study is a prospective paired noninferiority trial. Inclusion/exclusion criteria are clearly stated (e.g., women aged 50-71, exclusion for symptoms, prostheses, or AI-incompatible images). A power analysis is reported with parameters (alpha=0.05, beta=0.20, assumed cancer incidence 6/1000, noninferiority margin 5%). Randomization is not applicable as it is a paired design (each woman serves as her own control). Blinding is not applicable in the same sense as a parallel trial, but the standard strategy uses double reading without AI support, and the AI strategy uses AI-supported double reading; the paper does not explicitly state blinding of radiologists to the strategy, but the paired design inherently means each exam is read under both strategies by different readers. The paper does not discuss biological vs. technical replicates (n/a for human trial). Outlier handling is not explicitly discussed, but the analysis population is defined (all included women). Controls are inherent in the paired design (standard strategy as control). Independent replication is not reported (n/a for a single pivotal trial).
“Screening readings were randomly assigned according to availability of the nine dedicated breast radiologists in the hospital’s radiology department (3 to 21 years of experience in breast cancer screening), regardless of whether they were readings with or without AI assistance and regardless of the reader’s experience.”
“The standard of care: double human reading without AI support. The AI strategy: double human reading with AI support (by two additional radiologists) only for cases classified by the AI system with a score of 8 to 10 (approximately 30% of the studies most likely to have cancer).”
“For the sample size calculation, the method by Connor for paired studies was used, taking into account the sample size necessary to demonstrate noninferiority (5% margin) in terms of sensitivity for cancer detection via McNemar’s z -test (one-sided). An alpha (type I error estimate) value of 5% and a beta (type II error estimate) of 20% (representing a power for the study of 80%, 1 − beta) was used.”
“After signing informed consent, women were excluded from the clinical trial if they had symptoms or signs of suspected breast cancer; if they had breast prostheses; or if their images could not be processed by the AI system (for example, because of the presence of visible breast implants or because of erroneous PACS image transfer).”
Sex is reported (all women). Age is reported as median and interquartile range, and by age groups. Breast density (BI-RADS) is reported. Demographics include screening round (first vs. follow-up) and modality (DM vs. DBT). Race/ethnicity is not individually recorded but noted as predominantly Caucasian. Species/strain/source and housing conditions are not applicable for a human study.
“Women | 31,301 (100.0%)”
“The median age was 59 years old (interquartile range 64 to 54 years).”
“In terms of breast density, the distribution was as follows: A: 20.6%; B: 46.5%; C: 27.6%; and D: 5.3%.”
“Women | 31,301 (100.0%)”
“Age at screening, years, median (interquartile range) | 59 (54–64)”
The study reports written informed consent from all participants, a favorable ruling from a named IRB with protocol number, and compliance with HIPAA. The ethics statement is adequate.
“In March 2021, the study protocol and the informed consent received a favorable ruling from the Institutional Review Board (IRB) at Reina Sofía University Hospital of Córdoba Research Ethics Committee (IRB No 4932).”
“All patients provided written informed consent before enrollment.”
“This prospective clinical trial was compliant with the Health Insurance Portability and Accountability Act and the design was preregistered at https://classic.clinicaltrials.gov/ct2/show/NCT04949776 .”
“In March 2021, the study protocol and the informed consent received a favorable ruling from the Institutional Review Board (IRB) at Reina Sofía University Hospital of Córdoba Research Ethics Committee (IRB No 4932).”
“All patients provided written informed consent before enrollment.”
“This prospective clinical trial was compliant with the Health Insurance Portability and Accountability Act”
The AI system (Transpara version 1.7, ScreenPoint Medical) is identified with version and manufacturer. The mammography devices are identified (Lorad Selenia, Hologic; 3Dimensions, Hologic). The statistical software is identified (R version 4.4.2) with key libraries. Antibodies, cell lines, mycoplasma testing, and organisms are not applicable for this clinical trial. Reagents are not applicable beyond the imaging equipment and AI software.
“This clinical trial used a commercially available AI system for breast cancer detection, Transpara (version 1.7 ScreenPoint Medical), which had been used for previous research studies by the same group of radiologists involved in this study.”
“Images were acquired using four devices: three DM devices (Lorad Selenia, Hologic) and one DBT device (3Dimensions, Hologic).”
“The software used to perform the statistical analyses in this study was R, version 4.4.2.”
“This clinical trial used a commercially available AI system for breast cancer detection, Transpara (version 1.7 ScreenPoint Medical)”
“The software used to perform the statistical analyses in this study was R, version 4.4.2.”
“Images were acquired using four devices: three DM devices (Lorad Selenia, Hologic) and one DBT device (3Dimensions, Hologic).”
Statistical tests are named (McNemar test, z-test, Wald statistics). Effect sizes with 95% confidence intervals are reported for primary and secondary outcomes. Software is identified (R v4.4.2). Data presentation includes per-group n and proportions with CIs. Assumptions verification (e.g., normality) is not explicitly discussed, but for a large paired trial with McNemar tests, this is less critical. Exact p-values are reported for some comparisons (e.g., P < 0.001 for CDR superiority) but not all (e.g., P = 0.997 for RR noninferiority is a threshold). Mathematical plausibility checks: the reported numbers appear internally consistent (e.g., percentages sum to ~100% in Table 1).
“For PPV, a z -test was used for testing the two-sided hypothesis of equality of proportions, while for FPR a paired McNemar test was used.”
“CDR | 228/31,301 | 7.3 | 198/31,301 | 6.3 | 1 | 15.2 | P < 0.001 bc”
“an increase in the CDR of 15.2% (95% CI 6.6%, 24.4%), which is an absolute difference of 1.0 of 1,000 (95% CI 0.4 of 1,000, 1.5 of 1,000; P < 0.001)”
“CDR | 228/31,301 | 7.3 | 198/31,301 | 6.3 | 1 | 15.2 | P < 0.001 bc | | (6.3, 8.2) | (5.5, 7.2) | (0.4, 1.5) | (6.6, 24.4)”
“RR | 1,723/31,301 | 5.5 | 1,501/31,301 | 4.8 | 0.7 | 14.8 | P = 0.997 d”
The individual deidentified participant dataset is publicly accessible via Zenodo with a DOI. The AI algorithm code is proprietary but access conditions are stated. Statistical software is identified. The data availability statement is concrete.
“Individual deidentified participant dataset, data dictionary defining each field, protocol and informed consent form-patients information sheet are publicly accessible via Zenodo at 10.5281/zenodo.17625633 (ref. ).”
“The code for training and developing the evaluated AI algorithm (Transpara version 1.7, ScreenPoint Medical) is part of a proprietary system. The commercial AI system is available for external research evaluation collaborations for researchers who provide a relevant and methodologically sound proposal.”
“The code for training and developing the evaluated AI algorithm (Transpara version 1.7, ScreenPoint Medical) is part of a proprietary system.”
The trial is registered at ClinicalTrials.gov (NCT04849776). Methods are detailed enough for replication (procedures, AI system, statistical analysis). Limitations are discussed (single-site, single vendor, single AI system, sample size not powered for subgroups). Conclusions are proportional to the evidence (e.g., 'could be a safe and effective screening strategy'). Funding sources and conflicts of interest are disclosed. A reporting guideline is not explicitly mentioned in the main text, but the paper states 'Reporting summary Further information on research design is available in the linked to this article', which may refer to a Nature Portfolio reporting summary.
“ClinicalTrials.gov: NCT04849776 (http://clinicaltrials.gov/ct2/show/NCT04849776) .”
“Our study has the limitation of being a single-site investigation with a screening workflow of double reading without arbitration by expert radiologists in breast imaging and breast screening who have several years of experience in the use of AI.”
“In March 2023, EEC received for this study a grant (20,000 euros) from the SEDIM Foundation. The funder had no role in study design, data collection and analysis, decision to publish or preparation of the manuscript.”
“ClinicalTrials.gov: NCT04849776 (http://clinicaltrials.gov/ct2/show/NCT04849776)”
“Our study has the limitation of being a single-site investigation with a screening workflow of double reading without arbitration by expert radiologists in breast imaging and breast screening who have several years of experience in the use of AI.”
“The study was not funded by Industry. A.R.R. is an employee of ScreenPoint Medical. The other authors declare no competing interests.”
Registered (2 IDs: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 32 references by DOI: 1 verified — 1 DOI unresolved, 30 no DOI (shown, not verified).
- UNRESOLVED10.5281/zenodo.17625633Artificial Intelligence in Breast Cancer Screening Program in Cordoba (AITIC)Cited DOI does not resolve to any Crossref record.
- NO DOINational performance benchmarks for screening digital breast tomosynthesis: update from the Breast Cancer Surveillance ConsortiumNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIProspective study aiming to compare 2D mammography and tomosynthesis + synthesized mammography in terms of cancer detection and recall. From double reading of 2D mammography to single reading of tomosynthesisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAI-based strategies to reduce workload in breast cancer screening with mammography and tomosynthesis: a retrospective evaluationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImpact of real-life use of artificial intelligence as support for human reading in a population-based breast cancer screening program with mammography and tomosynthesisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImpact of artificial intelligence support on accuracy and reading time in breast tomosynthesis image interpretation: a multi-reader multi-case studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDetection of breast cancer with mammography: effect of an artificial intelligence support systemNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImproving breast cancer detection accuracy of mammography with the concurrent use of an artificial intelligence toolNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStand-alone use of artificial intelligence for digital mammography and digital breast tomosynthesis screening: a retrospective evaluationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStand-alone artificial intelligence for breast cancer detection in mammography: comparison with 101 radiologistsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExternal evaluation of 3 commercial artificial intelligence algorithms for independent assessment of screening mammogramsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInternational evaluation of an AI system for breast cancer screeningNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPossible strategies for use of artificial intelligence in screen-reading of mammograms, based on retrospective data from 122,969 screening examinationsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence for reducing workload in breast cancer screening with digital breast tomosynthesisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIdentifying normal mammograms in a large screening population using artificial intelligenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICan we reduce the workload of mammographic screening by automatic identification of normal exams with artificial intelligence? A feasibility studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBreast cancer screening with digital breast tomosynthesis: comparison of different reading strategies implementing artificial intelligenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence—supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): a clinical safety analysis of a randomised, controlled, non-inferiority, single-blinded, screening accuracy studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScreening performance and characteristics of breast cancer detected in the Mammography Screening with Artificial Intelligence trial (MASAI): a randomised, controlled, parallel-group, non-inferiority, single-blinded, screening accuracy studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence for breast cancer detection in screening mammography in Sweden: a prospective, population-based, paired-reader, non-inferiority studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINationwide real-world implementation of AI for cancer detection in population-based mammography screeningNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIProspective implementation of AI-assisted screen reading to improve early detection of breast cancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn artificial intelligence-based mammography screening protocol for breast cancer: outcome and radiologist workloadNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffect of artificial intelligence-based triaging of breast cancer screening mammograms on cancer detection and radiologist workload: a retrospective simulation studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICombining the strengths of radiologists and AI for breast cancer screening: a retrospective analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEarly indicators of the impact of using AI in mammography screening for breast cancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITNM Classification of Malignant Tumors International Union Against CancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO Classification of Tumours. Breast Tumours Vol. 2No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIACR BI-RADS Atlas, Breast Imaging Reporting and Data SystemNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISample size for testing differences in proportions for the paired-sample designNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAdjusted Wald confidence interval for a difference of binomial proportions based on paired dataNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
3 data/code links checked; 3 live.
- datahttps://classic.clinicaltrials.gov/ct2/show/NCT04949776LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttp://clinicaltrials.gov/ct2/show/NCT04849776LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttp://clinicaltrials.gov/ct2/show/NTC04949776LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
8 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 8 minor suggestions below.
8 copyedit issues flagged: mostly consistency, clarity, grammar.
- MINORconsistencyAbstract“NCT04849776”→ Ensure consistent use of ClinicalTrials.gov identifier format (NCT04849776 vs NTC04949776).The identifier appears as NCT04849776 in the abstract and NTC04949776 in the Methods, which is a typographical inconsistency.
- MINORclarityResults, Primary outcomes“The CDR and RR were 6.3 of 1,000 (95% confidence interval (CI) 5.5 of 1,000, 7.2 of 1,000) and 4.8% (95% CI 4.6%, 5.0%), respectively.”→ Clarify that CDR is per 1,000 and RR is per 100 to avoid confusion.The units for CDR and RR are different (per 1,000 vs per 100) and could be explicitly stated in the text.
- MINORgrammarDiscussion, paragraph 4“The benefits of double reading of high-AI-risk exams with AI support and the replacement of double reading by single reading in low-risk exams are also supported by the results of the observational study after AI implementation in the Capital Region of Denmark (increase in the CDR of 17%, reduction in false positives of 32% and reduction in workload of 33%) using the same AI system as MASAI and AITIC.”→ Consider splitting this long sentence for clarity.The sentence is long and could be broken into two for readability.
- MINORtypoAbstract“Received 2025 Jan 12; Accepted 2026 Feb 4; Issue date 2026.”→ Check dates: acceptance date (2026) is after the study end date (2024).This may be a placeholder or error.
- MINORconsistencyMethods, Clinical trial design and patients“ClinicalTrials.gov ID: NTC04949776”→ Should be NCT04949776 (NCT, not NTC).Typo in the trial registration number.
- MINORclarityResults, Table 2 footnote“For the CDR and RR, initially, one-tailed McNemar paired tests were applied for the noninferiority analyses. However, when noninferiority was achieved, the P value represented in the table is the one corresponding to the superiority analyses.”→ Clarify that the p-values in the table for CDR and RR are from superiority tests, not noninferiority tests.This is clear but could be misinterpreted.
- MINORconsistencyResults, Table 3“Screening readings (workload) | 4,820/13,986 | 34.5 | 13,968/13,968 | 100 | −9,148 | −65.5”→ The denominator for DBT AI strategy readings is 13,986, but the total DBT N is 13,968. Check if this is a typo.Minor inconsistency: 13,986 vs 13,968.
- MINORgrammarDiscussion, paragraph 1“The results from this prospective, paired, noninferiority clinical trial demonstrate that in a prospective setting it is safe to use AI to identify screening exams that can be automatically labeled as normal and avoid radiologist reading in a population screening program that uses both DM and DBT exams.”→ Consider rephrasing for clarity: '...demonstrate that it is safe to use AI to identify screening exams that can be automatically labeled as normal, thereby avoiding radiologist reading...'Minor grammatical issue.
The published work is largely robust, but an informed reader should weigh the unresolved trial registration number (NCT04849776 not found in ClinicalTrials.gov) as a potential validity concern. The minor reporting gaps (blinding, exact p-values, reporting guideline) are not fatal but warrant attention. A correction or clarification regarding the trial registration is recommended.
- 1.HIGHreportingVerify and correct the ClinicalTrials.gov registration number (NCT04849776 vs NTC04949776) in the Abstract and Methods; ensure the registration is actually resolvable.The integrity check found that the registration number could not be resolved in ClinicalTrials.gov, which undermines the claim of preregistration.
- 2.HIGHreportingAdd a statement clarifying that the recall rate (RR) noninferiority was NOT demonstrated (P = 0.997) in the abstract and results, and temper any claim that the AI strategy is noninferior on RR.The claim audit found the abstract claim of RR noninferiority unsupported; the paper itself states noninferiority was not achieved.
- 3.HIGHreportingExplicitly state whether radiologists were blinded to the AI scores in the standard strategy and whether they were aware of the study hypothesis in the Methods, Procedures section.Reviewer 2 flagged that blinding is not explicitly described for the AI strategy readers, which is a key methodological detail for a clinical trial.
- 4.HIGHstatisticsReport exact p-values for all comparisons, especially for the RR noninferiority test (currently reported as P = 0.997) in Table 2 and the text.Reviewer 2 noted that exact p-values are not always reported, which limits the reader's ability to assess the evidence.
- 5.HIGHreportingExplicitly reference a reporting guideline (e.g., CONSORT) in the main text, not just in a linked reporting summary.Both reviewers noted that no formal reporting guideline is explicitly mentioned, which is a transparency gap.
- 6.HIGHdata codeProvide the statistical analysis code in a public repository (e.g., GitHub) to enhance reproducibility, even if the AI algorithm code is proprietary.Reviewer 2 suggested that sharing the analysis code would improve reproducibility; the current code_sharing is marked as inadequate.
- 7.MEDIUMstatisticsAdd a statement on verification of statistical assumptions (e.g., appropriateness of McNemar test for paired binary data) in the Methods, Statistical analyses section.Reviewer 2 noted that assumptions verification is not explicitly discussed, which is a minor reporting gap.
- 8.MEDIUMstatisticsDiscuss how outliers were handled in the analysis (e.g., any extreme values in continuous variables) in the Methods, Statistical analyses section.Reviewer 2 flagged that outlier handling is not reported, which is a minor methodological detail.
- 9.MEDIUMreportingClarify the randomization process for assigning readings to radiologists: was it truly random or based on availability?Reviewer 2 raised a question about the randomization process, which could affect the interpretation of the paired design.
- 10.MEDIUMreportingAdd a statement on the handling of missing data (e.g., if any women were lost to follow-up or had incomplete data) in the Methods, Statistical analyses section.Reviewer 2 suggested that missing data handling is not addressed, which is a minor transparency issue.
- 11.MEDIUMcopyeditFix the typo in the ClinicalTrials.gov identifier: change 'NTC04949776' to 'NCT04849776' in the Methods, Clinical trial design and patients section.The copyedit pass flagged this inconsistency, which could confuse readers and affect the registration lookup.
- 12.MEDIUMcopyeditClarify the units for CDR (per 1,000) and RR (per 100) in the Results, Primary outcomes section to avoid confusion.The copyedit pass noted that the units differ and could be misinterpreted.
- 13.MEDIUMcopyeditCheck the denominator in Table 3 for DBT AI strategy screening readings (13,986 vs 13,968) and correct if it is a typo.The copyedit pass and integrity check flagged this minor inconsistency, which could affect the workload calculation.
- 14.LOWcopyeditSplit the long sentence in Discussion, paragraph 4 for clarity.The copyedit pass suggested breaking up a long sentence to improve readability.
- 15.LOWcopyeditRephrase the first sentence of Discussion, paragraph 1 for clarity.The copyedit pass noted a minor grammatical issue that could be improved.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.