Nationwide real-world implementation of AI for cancer detection in population-based mammography screening.
Eisemann N, Bunk S, Mukama T, Baltus H, Elsner SA, Gomille T, Hecht G, Heywang-Köbrunner S, Rathmann R, Siegmann-Luz K, Töllner T, Vomweg TW, Leibig C, Katalinic A
- DOI
- 10.1038/s41591-024-03408-6
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/da6620ca-ebd9-46bf-ad96-a902a7daa1bc is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- CitationsUnresolved reference ×2−0.5★
- ReportingStudy design partially met−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on breast cancer detection rate (BCDR), a surrogate for the ultimate clinical outcome of reduced breast cancer mortality. The study does not demonstrate target engagement at the tested dose (the AI system's predictions are not linked to a dose-response or PK/PD measure) and does not cite validated evidence linking increased BCDR to reduced mortality. The paper itself acknowledges that downstream effects on interval cancer rate and stage distribution are unknown and require follow-up.
“The important downstream effects of AI-supported screening on overall program performance metrics, including interval cancer rate and stage-at-diagnosis distribution at subsequent screening rounds, are subject to follow-up investigations.”
- 02Treatment effect not shown to be clinically meaningful
The primary effect is a 17.6% relative increase in breast cancer detection rate (from 5.7 to 6.7 per 1,000), which corresponds to one additional cancer per 1,000 women. This effect is not anchored to a minimal clinically important difference or to a demonstrated improvement in patient outcomes such as reduced mortality or morbidity. The clinical meaningfulness of detecting additional cancers, especially DCIS (which increased by 67.6%), is uncertain and may represent overdiagnosis.
“The BCDR in the AI group was considered noninferior and even statistically superior to that in the control group.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a large, well-conducted observational implementation study with strong reporting of methods, ethics, data/code availability, and statistical analysis. The main weakness is the non-randomized design with potential self-selection bias, which is acknowledged and partially mitigated by propensity score weighting and sensitivity analyses. Minor reporting inconsistencies (abstract vs. results numbers) and lack of a formal reporting guideline are the only copyedit-level concerns.
Both reviewers classified the study as observational, which is adopted. The evaluation covered all eight dimensions; randomization, blinding, and animal-related criteria were marked not applicable. The statistics verification checked only one reported effect estimate (consistent); other statistics were not machine-verified. The citation check flagged two references (Dryad and Zenodo DOIs) as not found in Crossref/OpenAlex, but these are data/code repositories, not typical citations, and the links are live; this is noted but not treated as a fabrication signal.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks.
- CONSISTENTreported p < .050 · recomputed p = <.001Reviewers 1, 2Check the p-value for the relative difference in BCDR (17.6% with 95% CI 5.7% to 30.8%)
“This represents a model-based absolute difference of one additional cancer per 1,000 screened women and a relative increase of 17.6% (95% confidence interval (CI): +5.7%, +30.8%).”
Taken as given: The CI is a 95% confidence interval for the relative difference (ratio).; The estimate is the relative difference (17.6%) and the CI bounds are 5.7% and 30.8%.; The p-value is for the test that the relative difference is zero (i.e., ratio = 1).Method: Used pCI function with log=1 for ratio, assuming the CI is for the relative difference.How we recomputed it: pCI(0.176, 0.057, 0.308, 1)
- lowinternal contradictionThe abstract reports 463,094 women screened, while the Results section reports 461,818 women participated. The difference is explained by exclusions, but the abstract could be clearer.
From July 2021 to February 2023, a total of 463,094 women were screened (260,739 with AI support) by 119 radiologists. ... Overall, 461,818 women who attended mammography screening at the 12 screening sites participated in the study.
Abstractreviewer’s wording
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
6 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1AI-supported screening can reduce reading workload.The paper shows reduced reading times for normal-tagged exams and a hypothetical workload reduction, but this is not a direct measure of actual workload reduction in practice.Evidence: Reading time analysis and fictitious automation scenario.
“Overall, radiologists spent 43% less time interpreting examinations tagged as normal, with a mean reading time of 39 s for normal examinations compared to 67 s for examinations not tagged as normal (Extended Data Fig. ).”
ResultsFind in source - partialReviewer 2AI-supported screening is feasible and safe and can reduce workload.Feasibility and safety are supported by the study, but workload reduction is based on a post hoc analysis and reading time measurements in the AI group only.Evidence: Reading times were lower for normal-tagged exams; a fictitious automation scenario showed potential workload reduction.
“In conclusion, our findings substantially add to the growing body of evidence suggesting that AI-supported mammography screening is feasible and safe and can reduce workload.”
DiscussionFind in source - supportedReviewers 1, 2AI-supported double reading was associated with a higher breast cancer detection rate without negatively affecting the recall rate.The paper provides model-based estimates showing a 17.6% higher BCDR with a CI excluding zero, and a recall rate difference with CI including zero, supporting the claim.Evidence: Table 3 shows BCDR 6.7 vs 5.7 per 1,000, relative difference 17.6% (5.7%, 30.8%); recall rate 37.4 vs 38.3, relative difference -2.5% (-6.5%, 1.7%).
Compared to standard double reading, AI-supported double reading was associated with a higher breast cancer detection rate without negatively affecting the recall rate.
Abstractreviewer’s wording - supportedReviewers 1, 2AI-supported screening increased breast cancer detection rates by a significant margin without affecting the recall rate.The claim is supported by the same evidence as above; the margin is statistically significant for BCDR and non-significant for recall.Evidence: Table 3 and sensitivity analyses.
“In a large-scale prospective study run across 12 sites in Germany and involving 463,094 women and 119 radiologists, AI-supported screening increased breast cancer detection rates by a significant margin without affecting the recall rate.”
AbstractFind in source - supportedReviewers 1, 2The AI system can improve mammography screening metrics.The claim is supported by the primary outcomes and subgroup analyses, though the long-term impact on interval cancers is not yet known.Evidence: Primary outcomes and subgroup analyses.
“Compared to standard double reading, AI-supported double reading was associated with a higher breast cancer detection rate without negatively affecting the recall rate, strongly indicating that AI can improve mammography screening metrics.”
AbstractFind in source - supportedReviewer 1The increased BCDR is not due to confounding by reading behavior.The paper uses propensity score weighting and extensive sensitivity analyses to address confounding, supporting the claim.Evidence: Sensitivity analyses including placebo intervention, bootstrapping, and alternative adjustment methods.
“We conducted a placebo intervention analysis to check whether the AI effect observed in the main analysis would vanish (as it should) when there is only a placebo intervention while all assumptions of the model are kept (that is, in the presence of residual confounding due to the reading behavior). As expected, the average model-based difference was minimal (0.8% (−9.9%, 11.6%)), indicating no residual confounding.”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on breast cancer detection rate (BCDR), a surrogate for the ultimate clinical outcome of reduced breast cancer mortality. The study does not demonstrate target engagement at the tested dose (the AI system's predictions are not linked to a dose-response or PK/PD measure) and does not cite validated evidence linking increased BCDR to reduced mortality. The paper itself acknowledges that downstream effects on interval cancer rate and stage distribution are unknown and require follow-up.
“The important downstream effects of AI-supported screening on overall program performance metrics, including interval cancer rate and stage-at-diagnosis distribution at subsequent screening rounds, are subject to follow-up investigations.”
- INADEQUATEEffect sizeThe primary effect is a 17.6% relative increase in breast cancer detection rate (from 5.7 to 6.7 per 1,000), which corresponds to one additional cancer per 1,000 women. This effect is not anchored to a minimal clinically important difference or to a demonstrated improvement in patient outcomes such as reduced mortality or morbidity. The clinical meaningfulness of detecting additional cancers, especially DCIS (which increased by 67.6%), is uncertain and may represent overdiagnosis.
“The BCDR in the AI group was considered noninferior and even statistically superior to that in the control group.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Study-design details incomplete (controls, blinding, power)Assessed
The introduction cites multiple retrospective and prospective studies on AI in mammography screening, acknowledges their limitations (small sample sizes, lack of heterogeneity), and explains how PRAIM addresses these gaps. The rationale for the study is logically linked to the premise, and the hypothesis (noninferiority of AI-supported double reading) follows from the cited evidence.
“However, these studies are limited by small sample sizes (which restrict the analysis of subgroups) and by the lack of heterogeneity in terms of screening sites, mammography equipment vendors and the radiologists involved, thereby reducing their generalizability to real-world settings.”
“In the PRAIM (PRospective multicenter observational study of an integrated AI system with live Monitoring) implementation study embedded in the German mammography screening program, we investigated whether the performance metrics achieved by double reading using an AI-supported CE (Conformité Européenne)-certified medical device with a decision referral approach were noninferior to those achieved by double reading without AI support in a real-world setting.”
“However, these studies are limited by small sample sizes (which restrict the analysis of subgroups) and by the lack of heterogeneity in terms of screening sites, mammography equipment vendors and the radiologists involved, thereby reducing their generalizability to real-world settings.”
“In the PRAIM (PRospective multicenter observational study of an integrated AI system with live Monitoring) implementation study embedded in the German mammography screening program, we investigated whether the performance metrics achieved by double reading using an AI-supported CE (Conformité Européenne)-certified medical device with a decision referral approach were noninferior to those achieved by double reading without AI support in a real-world setting.”
The study is observational, so randomization_method, randomization_unit, and blinding_levels are not applicable. Power_analysis is reported (target of 200,000 per group) but the actual analysis was adapted due to bias. Inclusion_exclusion criteria are described (women with missing outcome data or cancellations excluded). Outlier_handling is addressed through sensitivity analyses and robust statistical methods. Controls are not applicable in the traditional sense, but the control group serves as the comparator. Independent_replication is not applicable as this is a single study.
“A sample size of 200,000 women per study group was targeted for assessing the noninferiority of AI in terms of the BCDR with the originally planned analysis.”
“For analysis, women with missing outcome data (666 women, 0.1%; Fig. ) or cancellations (610 women, 0.1%) were excluded.”
“PRAIM is an observational study with no random assignment of screening examinations to the AI-supported and standard-of-care groups.”
“A sample size of 200,000 women per study group was targeted for assessing the noninferiority of AI in terms of the BCDR with the originally planned analysis.”
“For analysis, women with missing outcome data (666 women, 0.1%; Fig. ) or cancellations (610 women, 0.1%) were excluded.”
Sex is reported (all female), age is reported (median and IQR), and demographics are provided in Table 1. Since the study includes only women, sex_justified is not applicable (it is a screening program for women). Species/strain and housing conditions are not applicable as this is a human study.
“Median (IQR) | 58 (54–63) | 58 (54–63) | 58 (54–63)”
The study was approved by the ethics committee of the University of Lübeck (22-043), which waived the need for informed consent. The study protocol was registered in the German Clinical Trials Register (DRKS00027322). Compliance with GDPR is stated.
“The study protocol was registered in the German Clinical Trials Register (DRKS00027322) and was approved by the ethics committee of the University of Lübeck (22-043), which waived the need for informed consent.”
“Women were informed about the use of AI software at each screening site, and their data were processed according to applicable data privacy rules, including the General Data Protection Regulation.”
“The study protocol was registered in the German Clinical Trials Register (DRKS00027322) and was approved by the ethics committee of the University of Lübeck (22-043), which waived the need for informed consent.”
“their data were processed according to applicable data privacy rules, including the General Data Protection Regulation.”
The investigational product is the AI system (Vara MG), which is named with manufacturer and CE certification. The statistical software (R and Python) and packages are identified with versions. Antibodies, cell lines, mycoplasma, and organisms are not applicable.
“The AI system used was Vara MG (from the German company Vara), a CE-certified medical device designed to display mammograms (viewer software) and preclassify screening examinations to assist radiologists in their reporting routine.”
“All analyses were conducted with R (version 4.1.3) using the packages PSweight (version 1.2.0) , dagitty (version 0.3.1) and marginaleffects (version 0.18) , as well as Python (version 3.10) using the package dowhy (version 0.11.1) .”
“The AI system used was Vara MG (from the German company Vara), a CE-certified medical device designed to display mammograms (viewer software) and preclassify screening examinations to assist radiologists in their reporting routine.”
“During the study, the medical device being examined underwent a series of updates, transitioning from version 1.0.5 to 2.6.2 (a total of ten updates).”
The paper uses propensity score weighting with overlap weighting, and reports effect sizes with 95% CIs. Exact p-values are not reported, but the paper uses estimation with CIs, which is acceptable. Assumptions are addressed through sensitivity analyses and causal graph analysis. Software is identified. Data presentation includes tables with model-based predictions and CIs. Mathematical plausibility is not applicable as the data are model-based and large-N.
“This represents a model-based absolute difference of one additional cancer per 1,000 screened women and a relative increase of 17.6% (95% confidence interval (CI): +5.7%, +30.8%).”
“This represents a model-based absolute difference of one additional cancer per 1,000 screened women and a relative increase of 17.6% (95% confidence interval (CI): +5.7%, +30.8%).”
The data availability statement names a concrete repository (Dryad) with a DOI. Code is available via Zenodo with a DOI. The AI system code is not shared due to commercial reasons, but a process for external research evaluation is described. This meets the criteria for adequate data and code sharing.
“The anonymized analysis dataset, including individual participant data and a data dictionary defining each field, is available via Dryad at 10.5061/dryad.zs7h44jgn (ref. ).”
“The code and supporting information necessary to reproduce the results are available via Zenodo at 10.5281/zenodo.10822135 (ref. ).”
“The anonymized analysis dataset, including individual participant data and a data dictionary defining each field, is available via Dryad at 10.5061/dryad.zs7h44jgn”
“The code and supporting information necessary to reproduce the results are available via Zenodo at 10.5281/zenodo.10822135”
The study is registered (DRKS00027322). Methods are detailed. Limitations are discussed extensively. Conclusions are proportional to the evidence. Funding and competing interests are disclosed. A reporting guideline is not explicitly mentioned, but the paper follows a structured format.
“The study protocol was registered in the German Clinical Trials Register (DRKS00027322)”
“Our study has some limitations. PRAIM is an observational study with no random assignment of screening examinations to the AI-supported and standard-of-care groups.”
“The study protocol was registered in the German Clinical Trials Register (DRKS00027322)”
“PRAIM is an observational study with no random assignment of screening examinations to the AI-supported and standard-of-care groups. Thus, there was a risk that confounding factors influenced radiologists’ decision to use AI to interpret examinations, which could bias the findings.”
“The study was funded by Vara. Vara was involved in the study design, collection and interpretation of data, and writing of the report.”
Registered (1 ID: German Clinical Trials Register (DRKS)). No reporting guideline cited.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 42 references by DOI: 33 verified — 2 DOI unresolved, 7 no DOI (shown, not verified).
- UNRESOLVED10.5061/dryad.zs7h44jgnArtificial intelligence in breast cancer screening: results of a nationwide real-world prospective cohort study (PRAIM)Cited DOI does not resolve to any Crossref record.
- UNRESOLVED10.5281/zenodo.10822135Code & supporting documents for ‘Artificial intelligence for cancer detection in population-based mammography screening: results of a nationwide real-world prospective cohort study (PRAIM)’Cited DOI does not resolve to any Crossref record.
- NO DOIEuropean Guidelines for Quality Assurance in Breast Cancer Screening and DiagnosisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEuropean guidelines on breast cancer screening and diagnosisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScreening for breast cancer: US Preventive Services Task Force recommendation statementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Python Language Reference Manual: For Python Version 3.2No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDoWhy: an end-to-end library for causal inferenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIJahresbericht Evaluation 2021. Deutsches Mammographie-Screening-ProgrammNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA kernel statistical test of independenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://research.uni-luebeck.de/en/projects/prospective-multicenter-observational-study-of-an-integrated-ai-sLIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
3 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 3 minor suggestions below.
3 copyedit issues flagged: mostly consistency, clarity.
- MINORconsistencyAbstract“463,094 women were screened (260,739 with AI support)”→ Ensure the total number of women screened is consistent with the 461,818 reported in the results.The abstract states 463,094, while the results state 461,818. This discrepancy should be reconciled.
- MINORclarityMethods, Statistical methods“A sample size of 200,000 women per study group was targeted”→ Clarify whether this was the final sample size or if it was adjusted due to the bias.The sentence could be clearer about the final sample size achieved.
- MINORconsistencyAbstract vs. Results“Abstract states 463,094 women screened, Results state 461,818 women participated.”→ Clarify the difference between screened and participated (exclusions).The difference is explained by exclusions, but the abstract could be more precise.
The published work is robust and well-reported; an informed reader should weigh the observational design and potential self-selection bias when interpreting the noninferiority conclusion. The minor abstract/results number discrepancy and the lack of a formal reporting guideline are worth noting but do not undermine the study's validity. No erratum is warranted for the number discrepancy, but the authors could clarify it in a correction if desired.
- 1.HIGHreportingReconcile the discrepancy between the abstract's 463,094 women screened and the Results' 461,818 women participated; clarify that the difference is due to exclusions (666 missing outcome data + 610 cancellations) in the abstract.The copyedit pass flagged this internal inconsistency, which could confuse readers and reviewers about the exact cohort size.
- 2.HIGHreportingAdd an explicit statement in the Methods or Reporting Summary that the study follows the STROBE reporting guideline for observational studies.Both reviewers noted the absence of a formal reporting guideline, which is a standard expectation for observational research and would enhance transparency.
- 3.HIGHstatisticsClarify in the Methods, Statistical methods section whether the targeted sample size of 200,000 per group was achieved or how the final sample size differed after the analysis plan was adapted due to self-selection bias.The power analysis is reported but the adaptation is not fully explained, leaving uncertainty about the study's statistical power.
- 4.MEDIUMstatisticsExplicitly describe how outliers were handled in the statistical analysis, even if only to state that no outliers were excluded.One reviewer rated outlier handling as 'reported_but_inadequate' because it is not explicitly described, which is a minor reporting gap.
- 5.MEDIUMreportingAdd a brief statement on the generalizability of findings to other populations, given the lack of racial/ethnic diversity data in the cohort.One reviewer suggested this to address potential limitations in external validity.
- 6.MEDIUMreportingClarify the role of the funding source (Vara) in the analysis and interpretation, beyond the current statement that they were involved in study design and writing.The competing interests statement is adequate but could be more explicit about the extent of funder involvement in data analysis to address potential bias concerns.
- 7.LOWdata codeConsider making the AI system's training code available under a more permissive license, or provide a detailed description of the algorithms to facilitate independent replication.The AI code is proprietary, but a more detailed description would improve reproducibility, as suggested by both reviewers.
- 8.LOWreportingConsider reporting exact p-values for key comparisons in addition to confidence intervals to facilitate meta-analyses.One reviewer suggested this as a nice-to-have for future meta-analyses, though the current use of CIs is appropriate.
- 9.LOWreportingAdd a discussion of the potential for overdiagnosis due to increased DCIS detection and the balance between benefits and harms.One reviewer suggested this to provide a more balanced interpretation of the increased cancer detection rate.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.