AI-based selection of individuals for supplemental MRI in population-based breast cancer screening: the randomized ScreenTrustMRI trial.
Salim M, Liu Y, Sorkhei M, Ntoula D, Foukakis T, Fredriksson I, Wang Y, Eklund M, Azizpour H, Smith K, Strand F
- DOI
- 10.1038/s41591-024-03093-5
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/8f592a56-8349-4e51-afc7-3e0c037c3b54 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on cancer detection rate (a surrogate for clinical benefit such as reduced mortality or morbidity). The paper does not demonstrate target engagement at the tested dose (no PK/PD or dose-exposure data) and does not cite validated evidence linking the surrogate (cancer detection by MRI) to a hard clinical outcome (e.g., reduced mortality or interval cancer rate) in this context. The primary endpoint of the trial (advanced cancer at 27 months) is not yet reported.
“The primary endpoint of ScreenTrustMRI is advanced breast cancer defined as either interval cancer, invasive component larger than 15 mm or lymph node positive cancer, based on a 27-month follow-up time from the initial screening. Secondary endpoints,…”
- 02Treatment effect not shown to be clinically meaningful
The reported effect is a cancer detection rate of 64.4 per 1,000 MRI examinations. This is presented as a rate without anchoring to a minimal clinically important difference or to a meaningful clinical outcome (e.g., reduction in mortality or advanced cancer). The comparison to the DENSE trial's 16.5 per 1,000 is a relative efficiency claim, but the clinical meaningfulness of detecting additional cancers by MRI (without evidence of improved hard outcomes) is not established.
“We observed a cancer detection rate of 64.4 cancers per 1,000 MRI examinations... The cancer detection rate of our trial at 64 (95% CI 46.8–88.1) cancers per 1,000 MRIs corresponds to about 3.8 times higher supplemental cancer detection rate compared with the…”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported randomized controlled trial evaluating an AI-based selection tool for supplemental MRI in breast cancer screening. The paper has strong scientific premise, rigorous design, clear ethics reporting, and detailed methods. The main weaknesses are a vague data availability statement and minor reporting gaps (no explicit power analysis, no explicit reporting guideline, some p-values as thresholds).
Both reviewers classified the study as interventional (RCT), and this was adopted. The evaluation covered all eight dimensions; several sub-criteria were marked not applicable (e.g., blinding, power analysis for secondary endpoint, species/housing, antibodies/cell lines). The statistics verification recomputed only 3 tests (all consistent); the rest of the statistical results were not machine-verified and should not be assumed correct. The citation check found no retracted or non-existent references.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks.
- CONSISTENTreported p = .046 · recomputed p = .071Reviewers 1, 2Check p-value for BI-RADS 3 comparison (no breast cancer vs diagnosed breast cancer) in Table 2.
“BI-RADS 3: No breast cancer (n=47), Diagnosed breast cancer (n=7)”
Taken as given: The counts 47 and 7 are from the 'No breast cancer' and 'Diagnosed breast cancer' columns for BI-RADS 3.; The total for 'No breast cancer' is 523, so the complement for BI-RADS 3 is 523-47=476.; The total for 'Diagnosed breast cancer' is 36, so the complement is 36-7=29.; The test used is Fisher's exact test (two-sided) as is common for small cell counts.Method: Fisher's exact test (two-sided) from cell counts (47, 476, 7, 29).How we recomputed it: pFisher2x2(47, 476, 7, 29, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Check p-value for BI-RADS 4 comparison in Table 2.
“BI-RADS 4: No breast cancer (n=10), Diagnosed breast cancer (n=17)”
Taken as given: The counts 10 and 17 are from the 'No breast cancer' and 'Diagnosed breast cancer' columns for BI-RADS 4.; The total for 'No breast cancer' is 523, so the complement is 523-10=513.; The total for 'Diagnosed breast cancer' is 36, so the complement is 36-17=19.; The test used is Fisher's exact test (two-sided).Method: Fisher's exact test (two-sided) from cell counts (10, 513, 17, 19).How we recomputed it: pFisher2x2(10, 513, 17, 19, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Check p-value for BI-RADS 5 comparison in Table 2.
“BI-RADS 5: No breast cancer (n=2), Diagnosed breast cancer (n=12)”
Taken as given: The counts 2 and 12 are from the 'No breast cancer' and 'Diagnosed breast cancer' columns for BI-RADS 5.; The total for 'No breast cancer' is 523, so the complement is 523-2=521.; The total for 'Diagnosed breast cancer' is 36, so the complement is 36-12=24.; The test used is Fisher's exact test (two-sided).Method: Fisher's exact test (two-sided) from cell counts (2, 521, 12, 24).How we recomputed it: pFisher2x2(2, 521, 12, 24, 0)
- lowinternal contradictionIn Table 2, the column 'No breast cancer (n=523)' has a footnote 'Includes women that were biopsied with a benign finding and women that are undergoing follow-up.' However, the total number of women with BI-RADS 3-5 is 95, and 24 of those are undergoing follow-up. The sum of women with BI-RADS 3-5 in the 'No breast cancer' column is 47+10+2 = 59, which is less than 95-36=59, consistent. But the footnote mentions 'women that are undergoing follow-up' which are 24, but these are not explicitly accounted for in the 'No breast cancer' column. This is a minor inconsistency in labeling.
“a Includes women that were biopsied with a benign finding and women that are undergoing follow-up.”
Table 2Find in source - lowinternal contradictionTable 2 fibroglandular tissue counts sum to 557, not 559, but the footnote explains missing assessments.
“Amount of fibrograndular tissue | | Almost entirely fat | 4 | (0.7%) | 4 | (0.77%) | 0 | (0%) | N/A | | Scattered fibroglandular tissue | 206 | (36.9%) | 188 | (36%) | 18 | (50%) | 0.098 | | Heterogenous fibroglandular tissue | 288 | (52%) | 275 | (53%) | 13 | (36%) | 0.057 | | Extreme fibroglandular tissue | 59 | (10.6%) | 54 | (10.3%) | 5 | (13.9%) | 0.508”
Table 2Find in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1The cost per cancer detected using AISmartDensity is comparable with screening mammography.The claim is supported by a high-level cost-comparison analysis, but the analysis is based on assumptions (e.g., MRI is ten times more expensive) and not a formal cost-effectiveness study. The paper acknowledges this as an ad hoc analysis.Evidence: Methods, Cost-comparison: 'A high-level cost-comparison analysis was carried out... MRI is ten times more expensive per examination compared with mammography.' Results: 'Using the AISmartDensity method would make the detection cost per cancer similar to the cost in population-wide screening mammography.'
“Using the AISmartDensity method would make the detection cost per cancer similar to the cost in population-wide screening mammography and contribute to earlier detection of invasive cancer.”
Discussion ¶1Find in source - partialReviewer 2Using AISmartDensity makes the cost per cancer detected comparable with screening mammography.The cost-comparison is a high-level analysis with assumptions, not a full cost-effectiveness study.Evidence: A cost-comparison analysis is described in Methods, but no detailed results are presented in this paper.
“Altogether, our results show that using an AI-based score to select a small proportion (6.9%) of individuals for supplemental MRI after negative mammography detects many missed cancers, making the cost per cancer detected comparable with screening mammography.”
DiscussionFind in source - supportedReviewer 1The AI method was nearly four times more efficient in terms of cancers detected per 1,000 MRI examinations compared to traditional breast density measures.The claim is supported by the reported cancer detection rate of 64.4 per 1,000 MRI exams in this study versus 16.5 per 1,000 in the DENSE trial, which used traditional density.Evidence: Results section: 'The cancer detection rate of our trial at 64 (95% CI 46.8–88.1) cancers per 1,000 MRIs corresponds to about 3.8 times higher supplemental cancer detection rate compared with the traditional density method used in the DENSE trial at 16.5 cancers per 1,000 MRIs.'
“The cancer detection rate of our trial at 64 (95% CI 46.8–88.1) cancers per 1,000 MRIs corresponds to about 3.8 times higher supplemental cancer detection rate compared with the traditional density method used in the DENSE trial at 16.5 cancers per 1,000 MRIs.”
Discussion ¶2Find in source - supportedReviewer 1Using an AI-based score to select a small proportion (6.9%) of individuals for supplemental MRI after negative mammography detects many missed cancers.The claim is supported by the detection of 36 cancers in 559 women (6.9% of the screened population), with a cancer detection rate of 64.4 per 1,000.Evidence: Results section: 'Cancerous lesions were detected in 36 participants, corresponding to 64.4 (95% confidence interval (CI) 46.8–88.1) cancer detection rate per 1,000 MRI examinations.'
Cancerous lesions were detected in 36 participants, corresponding to 64.4 (95% confidence interval (CI) 46.8–88.1) cancer detection rate per 1,000 MRI examinations.
Resultsreviewer’s wording - supportedReviewer 1Most additional cancers detected were invasive and several were multifocal, suggesting that their detection was timely.The claim is supported by the cancer characteristics in Table 4, which show that 75% of cancers were invasive only, and 11% were multifocal on histopathology.Evidence: Table 4: 'Invasiveness, n (%): Invasive only: 27 (75%)' and 'Multiple lesions, n (%): 2 lesions: 2 (6%), ≥3 lesions: 2 (6%)'.
“In the histopathological analysis of the surgical specimens, most (22 or 61%) were a combination of invasive and ductal cancer in situ, with 5 (14% of 36) being in situ only.”
ResultsFind in source - supportedReviewer 1The AI tool was trained only on mammography images from Hologic equipment and will need to be validated for other equipment.This is stated as a limitation in the Discussion, and the paper acknowledges the need for external validation.Evidence: Discussion, paragraph 8: 'the AI tool was trained only on mammography images from Hologic equipment, using images of high quality assessed by highly experienced radiologists will need to be validated for other equipment for the method to be generalized.'
“the AI tool was trained only on mammography images from Hologic equipment, using images of high quality assessed by highly experienced radiologists will need to be validated for other equipment for the method to be generalized.”
Discussion ¶8Find in source - supportedReviewer 2The AI-based selection method detects cancers at a rate of 64.4 per 1,000 MRI examinations.The claim is directly supported by the reported cancer detection rate with 95% CI.Evidence: Results section reports 36 cancers in 559 MRI examinations, yielding 64.4 per 1,000 (95% CI 46.8–88.1).
Cancerous lesions were detected in 36 participants, corresponding to 64.4 (95% confidence interval (CI) 46.8–88.1) cancer detection rate per 1,000 MRI examinations.
Resultsreviewer’s wording - supportedReviewer 2The AI method is nearly four times more efficient than traditional breast density measures.The comparison to the DENSE trial's 16.5 per 1,000 is valid, though it is a cross-study comparison.Evidence: Discussion compares the 64.4 rate to the DENSE trial's 16.5 per 1,000, yielding ~3.8 times higher.
“The cancer detection rate of our trial at 64 (95% CI 46.8–88.1) cancers per 1,000 MRIs corresponds to about 3.8 times higher supplemental cancer detection rate compared with the traditional density method used in the DENSE trial at 16.5 cancers per 1,000 MRIs.”
Discussion ¶2Find in source - supportedReviewer 2Most additional cancers detected were invasive and several were multifocal, suggesting timely detection.The claim is supported by the cancer characteristics table showing 75% invasive and 11% multifocal.Evidence: Table 4 shows 27/36 (75%) invasive only, and 4/36 (11%) multifocal on histopathology.
“Among all the diagnosed cancers, 7 (19% of 36) presented with multiple mass lesions on MRI whilst histopathological analysis confirmed multifocality for 4 (11% of 36).”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on cancer detection rate (a surrogate for clinical benefit such as reduced mortality or morbidity). The paper does not demonstrate target engagement at the tested dose (no PK/PD or dose-exposure data) and does not cite validated evidence linking the surrogate (cancer detection by MRI) to a hard clinical outcome (e.g., reduced mortality or interval cancer rate) in this context. The primary endpoint of the trial (advanced cancer at 27 months) is not yet reported.
“The primary endpoint of ScreenTrustMRI is advanced breast cancer defined as either interval cancer, invasive component larger than 15 mm or lymph node positive cancer, based on a 27-month follow-up time from the initial screening. Secondary endpoints, prespecified in the study protocol to be reported before the primary outcome, include cancer detected by supplemental MRI, which is the focus of the current paper.”
- INADEQUATEEffect sizeThe reported effect is a cancer detection rate of 64.4 per 1,000 MRI examinations. This is presented as a rate without anchoring to a minimal clinically important difference or to a meaningful clinical outcome (e.g., reduction in mortality or advanced cancer). The comparison to the DENSE trial's 16.5 per 1,000 is a relative efficiency claim, but the clinical meaningfulness of detecting additional cancers by MRI (without evidence of improved hard outcomes) is not established.
“We observed a cancer detection rate of 64.4 cancers per 1,000 MRI examinations... The cancer detection rate of our trial at 64 (95% CI 46.8–88.1) cancers per 1,000 MRIs corresponds to about 3.8 times higher supplemental cancer detection rate compared with the traditional density method used in the DENSE trial at 16.5 cancers per 1,000 MRIs.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites multiple studies on interval cancers, MRI sensitivity in dense breasts, and the DENSE trial, establishing the problem of missed cancers and the potential of MRI. It acknowledges limitations of prior work (e.g., high cost of MRI, lack of qualified staff) and presents the AI tool as a solution to improve cost-effectiveness. The hypothesis that AI-based selection will be more effective than traditional density measures is clearly stated.
“However, as qualified MRI staff are lacking, the equipment is expensive to purchase and cost-effectiveness for screening may not be convincing, the utilization of MRI is currently limited.”
“Our hypothesis was that AI-based image analysis will provide a more effective selection tool than traditional density in terms of the proportion of MRI examinations leading to a cancer diagnosis.”
“Our hypothesis was that AI-based image analysis will provide a more effective selection tool than traditional density in terms of the proportion of MRI examinations leading to a cancer diagnosis.”
“Traditional mammographic density and risk models thus seem to capture a markedly lower amount of relevant image information compared with the AISmartDensity tool.”
Randomization was performed using a random number generator in Excel by a study administrator not involved in assessments. Inclusion/exclusion criteria are clearly defined (e.g., high AI score, no high-risk surveillance, no implants). The trial is registered (NCT04832594) with a prespecified protocol. Blinding is not applicable because the control group does not receive MRI, and the outcome (cancer detection) is objective. Power analysis is not reported for this secondary endpoint analysis, but the primary endpoint is powered. Replicate distinction and controls are not applicable for this human RCT design.
“Randomization was performed by having a random number generated in Microsoft Excel by a study administrator not involved in the radiological assessments.”
“ClinicalTrials.gov registration: NCT04832594”
“Randomization was performed by having a random number generated in Microsoft Excel by a study administrator not involved in the radiological assessments.”
“Individuals attending screening as part of a high-risk surveillance program (for example, personal history, family history or genetic mutations) undergo supplemental imaging and were therefore excluded from the trial.”
The study population is described in Table 1 with median age (56 years), weight, height, and various reproductive and breast cancer history variables. Sex is implicitly reported as all female (the screening population is women). Age is reported in categories. Demographics such as family history and previous breast cancer are included. Species/strain and housing conditions are not applicable for this human study.
“Age, years (0 missing), median (IQR) | 56 (50–65)”
“Under the Swedish national breast screening program, women between 40 and 74 years old are invited for mammographic screening every 2 years.”
“The median age of the cohort was 56 years (interquartile range (IQR) 50–65 years) (Table ).”
“Weight, kg (24 missing) (median, IQR) | 69 (63–77)”
The trial was approved by the ethics review board in Stockholm County. Written informed consent was obtained from participants. The study complies with the Declaration of Helsinki. An independent Data Safety and Monitoring Committee is mentioned. These meet the criteria for adequate reporting.
“The trial was approved by the ethics review board in Stockholm County and was monitored by an independent Data Safety and Monitoring Committee.”
“The study complies with all local and national regulations regarding the use of human study participants and was conducted in accordance to the criteria set by the Declaration of Helsinki.”
“The trial was approved by the ethics review board in Stockholm County and was monitored by an independent Data Safety and Monitoring Committee.”
“Those providing written informed consent were enrolled and randomized either to supplemental breast MRI or as part of an observational control arm, which is not included in the current analysis.”
“The study complies with all local and national regulations regarding the use of human study participants and was conducted in accordance to the criteria set by the Declaration of Helsinki.”
The AI tool AISmartDensity is described in detail, including its three component models and training data. The MRI scanner (GE Signa Premier 3 T) and protocol are specified. Statistical software (Stata v.15.1) is identified. Antibodies, cell lines, and organisms are not applicable for this clinical trial. The AI software is not CE-marked or FDA-approved, which is noted as a limitation.
“AISmartDensity uses an average of standardized scores from the three models.”
“All MRI images were obtained on a GE Signa Premier 3 T MRI scanner.”
“Stata statistical software v.15.1 was used for all statistical analyses.”
“All MRI images were obtained on a GE Signa Premier 3 T MRI scanner.”
“Stata statistical software v.15.1 was used for all statistical analyses.”
The paper reports cancer detection rates with 95% CIs, PPVs with CIs, and p-values for comparisons between cancer and no-cancer groups (e.g., Table 2). Statistical tests are named (e.g., two-sided tests, Stata 'proportion' command). Exact p-values are given (e.g., p=0.046, p<0.001). Assumptions verification is not explicitly stated, but for a clinical trial with standard methods, this is acceptable. Data presentation includes per-group n and CIs. Mathematical plausibility checks are not applicable for continuous outcomes.
“64.4 (95% confidence interval (CI) 46.8–88.1) cancer detection rate per 1,000 MRI examinations”
“Stata statistical software v.15.1 was used for all statistical analyses.”
“Stata statistical software v.15.1 was used for all statistical analyses. All statistical tests were two-sided.”
The data availability statement says de-identified prediction scores and diagnostic outcomes will be shared upon request to the corresponding author, with no platform or timeframe specified. This is 'reported_but_inadequate' per the guidelines. Code is shared via a GitHub repository (https://github.com/radiology2023/RetrospectiveAISmartDensity), which is adequate. No accession numbers are provided for depositable data.
“De-identified patient-level prediction scores and diagnostic outcomes will be shared upon request to the corresponding author who will respond within 4 weeks.”
“Source code is available at https://github.com/radiology2023/RetrospectiveAISmartDensity”
“De-identified patient-level prediction scores and diagnostic outcomes will be shared upon request to the corresponding author who will respond within 4 weeks.”
“Source code is available at https://github.com/radiology2023/RetrospectiveAISmartDensity .”
Methods are detailed enough for replication (AI tool, MRI protocol, statistical analysis). Trial registration is provided (NCT04832594). Limitations are discussed (e.g., comparison with other studies, small sample, generalizability). Conclusions are proportional to the evidence (e.g., noting the need for primary endpoint follow-up). Funding sources and competing interests are declared. A reporting guideline is not explicitly mentioned, but the paper includes a CONSORT diagram.
“ClinicalTrials.gov registration: NCT04832594”
“A key limitation of the present report is that the cancer detection rate using this method can be compared only with results using traditional density in other studies, not within this study.”
“ClinicalTrials.gov registration: NCT04832594 (https://classic.clinicaltrials.gov/ct2/show/NCT04832594) .”
“A key limitation of the present report is that the cancer detection rate using this method can be compared only with results using traditional density in other studies, not within this study.”
“Funding was provided by Region Stockholm (M. Salim, F.S.), Medtechlabs (Y.L., K.S., H.A., F.S.), the Swedish Breast Cancer Association (F.S.), the Swedish Cancer Society grant no 22-2010-Pj-01-H (F.S.), the Swedish Research Council grant no 2022-01465 (F.S.) and 2020-00692 (M.E.).”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 33 references by DOI: 30 verified — 3 no DOI (shown, not verified).
- NO DOIDifferences in Ki67 and c-erbB2 expression between screen-detected and true interval breast cancersNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMammographic density and breast cancer in three ethnic groupsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIACR BI-RADS® Atlas, Breast Imaging Reporting and Data SystemNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
3 data/code links checked; 3 live.
- datahttps://classic.clinicaltrials.gov/ct2/show/NCT04832594LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT04832594LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/radiology2023/RetrospectiveAISmartDensityResolves to GitHub (code repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORconsistencyAbstract“64 versus 16.5”→ Consider adding units (cancers per 1,000 MRI examinations) for clarity.The units are clear from context but could be explicit.
- MINORclarityDiscussion, paragraph 5“The potential to pre-emptively detect most cancers by offering MRI to a small proportion of individuals represents an important healthcare value proposition.”→ Consider rephrasing for clarity: 'The potential to pre-emptively detect most cancers by offering MRI to a small proportion of individuals represents an important healthcare value.'Minor stylistic suggestion.
- MINORtypoDiscussion, paragraph 4“As a alternative approach”→ As an alternative approachArticle usage error.
- MINORconsistencyDiscussion, paragraph 11“shifter towards lower stages”→ shifted towards lower stagesVerb form inconsistency.
- MINORclarityMethods, Statistical analysis“The level for statistical significance was set at alpha = 0.05.”→ The level for statistical significance was set at α = 0.05.Use Greek alpha for consistency.
The published work is methodologically robust and generally well-reported, but an informed reader should weigh the vague data availability statement, the lack of an explicit power analysis for the primary endpoint, and the absence of an explicitly named reporting guideline. These are reporting gaps rather than validity threats; no erratum is warranted based on the checks performed, though the authors could strengthen reproducibility by depositing data in a managed-access repository and versioning the code with a DOI.
- 1.HIGHdata codeReplace the vague 'available on request' data availability statement with a concrete managed-access plan (e.g., a named data access committee or platform such as Vivli/YODA) and specify conditions and a response timeframe.A vague data statement is inadequate for a clinical trial and undermines reproducibility claims.
- 2.HIGHdata codeDeposit the analysis code in a versioned repository with a persistent DOI (e.g., Zenodo) and cite that DOI in the Code availability section.A GitHub URL without a versioned DOI is not a stable, citable artifact for reproducibility.
- 3.HIGHreportingExplicitly state adherence to the CONSORT reporting guideline in the Methods or a Reporting Summary section.The paper includes a CONSORT diagram but never names the guideline, which is a transparency gap reviewers will note.
- 4.HIGHstatisticsReport exact p-values instead of thresholds (e.g., p=0.0004 rather than p<0.001) in Table 2 and elsewhere.Threshold-only p-values are imprecise reporting and prevent readers from assessing the strength of evidence.
- 5.MEDIUMreportingAdd a power analysis or sample size justification for the primary endpoint in the Methods, even though this is an interim secondary-endpoint report.The absence of any power analysis is a gap for a clinical trial and limits interpretation of the secondary findings.
- 6.MEDIUMreportingExplicitly state that the trial is open-label and justify why blinding was not feasible, or describe any blinding of outcome assessors.Blinding status is not explicitly stated, which is a minor reporting gap for an RCT.
- 7.MEDIUMstatisticsAdd a statement on verification of statistical assumptions (e.g., normality, equal variance) and describe handling of missing data and outliers in the Statistical analysis section.Assumptions and missing-data handling are not reported, which weakens statistical transparency.
- 8.MEDIUMreportingProvide the full study protocol as a supplementary file to allow verification of prespecified outcomes.A full protocol would strengthen confidence that all prespecified outcomes are reported.
- 9.MEDIUMreportingClarify the CONSORT diagram to report the number of individuals screened for eligibility and reasons for exclusion.The current diagram may not fully document the flow of participants, which is a CONSORT requirement.
- 10.LOWcopyeditFix the typo 'As a alternative approach' to 'As an alternative approach' in Discussion, paragraph 4.Article usage error is a minor copyedit issue.
- 11.LOWcopyeditChange 'shifter towards lower stages' to 'shifted towards lower stages' in Discussion, paragraph 11.Verb form inconsistency is a minor copyedit issue.
- 12.LOWcopyeditUse Greek alpha (α) instead of 'alpha' in the Statistical analysis section for consistency.Minor stylistic consistency improvement.
- 13.LOWcopyeditAdd units (cancers per 1,000 MRI examinations) to the abstract's '64 versus 16.5' for clarity.Units are clear from context but explicit units improve clarity.
- 14.LOWcopyeditRephrase 'represents an important healthcare value proposition' to 'represents an important healthcare value' in Discussion, paragraph 5.Minor stylistic clarity improvement.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.