AI-based chest X-ray prioritization in the lung cancer diagnostic pathway: the LungIMPACT randomized controlled trial.
Woznitza N, Smith L, Rawlinson J, Au-Yong I, George B, Djearaman MG, Nair A, Lee RW, Navani N, Ndwandwe S, Clarke CS, Creeden A, Newsome J, Das I, Abaokporo S, Tucker R, Hathorn J, Baldwin DR
- DOI
- 10.1038/s41591-026-04253-5
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/e5faf2f9-b656-4dc4-8de1-924a6a6b17a3 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×5−2.5★
- StatisticsStatistic did not reproduce ×2−1★
- StatisticsPrinted percentage does not match its own count (capped) ×3−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 3 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Printed percentage does not match its own count
76.1% does not match the reported count 21511/26505
“21,511 (76.1%)”
Table 4Find in source - 02Printed percentage does not match its own count
10.9% does not match the reported count 3077/26505
“3,077 (10.9%)”
Table 4Find in source - 03Printed percentage does not match its own count
2.4% does not match the reported count 672/26505
“672 (2.4%)”
Table 4 - 04Printed percentage does not match its own count
2% does not match the reported count 559/26505
“559 (2.0%)”
Table 4 - 05Printed percentage does not match its own count
1.7% does not match the reported count 488/26505
“488 (1.7%)”
Table 4
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted, prospectively registered multicenter RCT with clear randomization, power analysis, and pre-specified outcomes. The paper is transparent about data and code availability, and the conclusions are appropriately cautious. Minor reporting gaps include lack of explicit blinding description and incomplete inclusion/exclusion criteria.
Both reviewers classified the study as interventional (RCT), and no disagreement was present. The evaluation covered all eight dimensions; several sub-criteria were marked not applicable (e.g., species/strain, housing, antibodies) due to the human clinical trial nature. The statistics verification component checked only a subset of reported tests (those with test statistics/df or effect estimates with CIs); the 5 'inconsistent' recomputations were not detailed and may reflect rounding, and no decision errors were found.
Numerical inconsistencies
3 findings · worst highValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks. 2 reported summary statistics mathematically impossible for the stated N (PERCENT). 3 printed percentages that do not match their own count.
- PERCENT76.1% does not match the reported count 21511/26505
“21,511 (76.1%)”
Table 4Find in source - PERCENT2.4% does not match the reported count 672/26505
“672 (2.4%)”
Table 4 - PERCENT2% does not match the reported count 559/26505
“559 (2.0%)”
Table 4 - PERCENT1.7% does not match the reported count 488/26505
“488 (1.7%)”
Table 4 - PERCENT10.9% does not match the reported count 3077/26505
“3,077 (10.9%)”
Table 4Find in source
- CONSISTENTreported p = .310 · recomputed p = .196Reviewers 1, 2Primary outcome: time to CT scan, ratio of geometric means
“There was no significant difference in the time from CXR acquisition to CT scan according to AI prioritization, with a ratio of geometric means of 0.97 (95% confidence interval (CI) = 0.93–1.02, P = 0.31).”
Taken as given: The ratio of geometric means is 0.97.; The 95% CI is 0.93 to 1.02.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed p-value from the reported ratio and 95% CI using the pCI function with log=1 for a ratio.How we recomputed it: pCI(0.97, 0.93, 1.02, 1) - CONSISTENTreported p = .840 · recomputed p = .813Reviewers 1, 2Primary outcome: time to lung cancer diagnosis, ratio of geometric means
“There was no significant difference in the time from CXR acquisition to lung cancer diagnosis between AI prioritization days, with a ratio of geometric means of 0.98 (95% CI = 0.83–1.16, P = 0.84; Table ).”
Taken as given: The ratio of geometric means is 0.98.; The 95% CI is 0.83 to 1.16.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed p-value from the reported ratio and 95% CI using the pCI function with log=1 for a ratio.How we recomputed it: pCI(0.98, 0.83, 1.16, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Secondary outcome: time to CXR report, ratio of geometric means
“Time from CXR to CXR report (h) | 44,078 | 47.0 | (15.8–99) | 42,814 | 34.1 | (6.6–93.1) | 0.85 | (0.83–0.87 | <0.001”
Taken as given: The ratio of geometric means is 0.85.; The 95% CI is 0.83 to 0.87.; The CI is two-sided at 95%.; The p-value is two-tailed and reported as <0.001.Method: Recomputed p-value from the reported ratio and 95% CI using the pCI function with log=1 for a ratio.How we recomputed it: pCI(0.85, 0.83, 0.87, 1)
- lowinternal contradictionThe paper reports 13,347 CTs identified, but the sum of CTs in the two arms (6,674 + 6,673) equals 13,347, which is consistent. However, the number of CTs within 14 days (1,314 + 1,452 = 2,766) matches the abstract, so no contradiction.
“Among all patients, 13,347 had a valid CT scan and were included in the analysis, with a time window from CXR acquisition to CT scan analysis—6,674 with AI prioritization and 6,673 with no AI prioritization.”
ResultsFind in source - lowinternal contradictionThe abstract states 4,405 CXRs were excluded due to data compliance issues or failure of randomization, but the Methods mention exclusion of invalid dates. The specific breakdown of exclusions is not provided.
“Of 97,731 participant CXRs, 4,405 were excluded due to data compliance issues or failure of randomization, resulting in 93,326 CXRs analyzed”
AbstractFind in source - lowinternal contradictionThe abstract states 4,405 CXRs were excluded, but the CONSORT diagram may show a different number; however, the text is consistent.
“Of 97,731 participant CXRs, 4,405 were excluded due to data compliance issues or failure of randomization, resulting in 93,326 CXRs analyzed”
AbstractFind in source - lowinternal contradictionThe paper reports 558 lung cancers, but the sum of cancers in the two arms (269 + 289 = 558) is consistent. No contradiction.
“A total of 558 patients were diagnosed with lung cancer and were included in the analysis of time to cancer diagnosis, 269 with AI prioritization and 289 with no AI prioritization.”
ResultsFind in source
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
7 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2AI prioritization of chest X-rays did not significantly shorten time to CT or lung cancer diagnosis.The primary outcomes show no significant difference, with p-values of 0.31 and 0.84, and the confidence intervals include 1.Evidence: Primary outcomes: ratio of geometric means 0.97 (95% CI 0.93-1.02, P=0.31) for time to CT; 0.98 (95% CI 0.83-1.16, P=0.84) for time to diagnosis.
“AI prioritization of CXR requested by UK primary care has no significant impact on the lung cancer pathway.”
AbstractFind in source - supportedReviewer 1AI prioritization did not improve time to urgent referral, time to treatment, or stage at diagnosis.Secondary outcomes show no significant differences, with p-values of 0.13, 0.99, and 0.34 respectively.Evidence: Secondary outcomes: time to 2WW referral P=0.13, time to treatment P=0.99, stage at diagnosis P=0.34.
“No significant differences were observed in time to lung cancer referral (14 versus 15 days; P = 0.13), time to treatment (76 versus 72.5 days; P = 0.99) or stage at diagnosis ( P = 0.34).”
AbstractFind in source - supportedReviewers 1, 2AI prioritization reduced time to CXR report.The secondary outcome shows a significant reduction in time to report, from 47.0 to 34.1 hours, with p<0.001.Evidence: Secondary outcome: time from CXR to report median 47.0 vs 34.1 hours, ratio 0.85 (95% CI 0.83-0.87, P<0.001).
“A significant reduction in the median time from CXR acquisition to report was observed, from 47 h to 34.1 h; however, no significant differences were detected in any of the timings measured as primary or secondary outcomes (Table ).”
ResultsFind in source - supportedReviewer 1Discordance between AI and radiology reports occurred in 30.3% of CXRs, and expert review identified actionable findings in 23.9% of discordant cases.The numbers are reported consistently in the abstract and results.Evidence: Discordance reviews: 28,261 discordant CXRs (30.3%), 6,750 actionable findings (23.9%).
“Discordance between AI and radiology reports occurred in 28,261 CXRs (30.3%) and expert radiology review identified actionable findings in 6,750 cases (23.9%).”
AbstractFind in source - supportedReviewer 1The study provides the largest reliable dataset on discordance reviews.The paper states this as a strength, and the number of reviews (26,505) is large, supporting the claim.Evidence: Discussion: 'This has provided perhaps the largest reliable dataset on this aspect.'
“This has provided perhaps the largest reliable dataset on this aspect.”
Discussion ¶5Find in source - supportedReviewers 1, 2AI prioritization alone is unlikely to accelerate the lung cancer diagnostic pathway.The primary outcomes show no effect, and the discussion appropriately concludes that prioritization alone is not sufficient.Evidence: Primary outcomes show no significant difference; discussion states 'AI prioritization of CXR requested by UK primary care has no significant impact on the lung cancer pathway.'
“AI prioritization of CXR requested by UK primary care has no significant impact on the lung cancer pathway.”
AbstractFind in source - supportedReviewer 2CXR AI deployments should not include worklist prioritization in this context.The conclusion follows from the null primary outcomes and the added complexity/cost of prioritization.Evidence: Primary outcomes show no benefit; discussion argues that prioritization adds complexity and cost without improving outcomes.
“Therefore, CXR AI deployments should not include worklist prioritization in this context.”
AbstractFind in source
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
5 integrity concerns flagged (0 high).
- lowotherThe paper reports a large number of discordance reviews (26,505) but the total number of discordant CXRs is 28,261; the difference is explained by the prespecified cap on reviews for certain categories.
“Discordance reviews were completed for 26,505 of 28,261 discordant CXR reports (30.3%) of all CXRs. In 1,756 cases, discordance review was not performed because the category was not considered potentially serious and the prespecified limit of 1,000 reviews per category had been reached.”
ResultsFind in source
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on diagnostic delays, the National Optimal Lung Cancer Pathway, and previous work by the same team showing that immediate radiographer reporting reduced time to diagnosis. It also acknowledges the lack of evidence for AI prioritization, citing a NICE review that identified LungIMPACT as potentially capable of addressing key questions. The rationale linking the premise to the study objectives is explicit, and the study directly addresses the gap in evidence.
“The primary aim of the trial was to measure the impact of immediate AI-driven prioritization of abnormal CXRs for reporting on the time to CT and diagnosis of lung cancer.”
“However, a recent review of the clinical utility of this technology by NICE concluded that there was insufficient evidence to make any recommendations other than to ensure products were carefully evaluated . NICE cited the LungIMPACT study as the only investigation considered potentially capable of addressing important clinical questions .”
“The primary aim of the trial was to measure the impact of immediate AI-driven prioritization of abnormal CXRs for reporting on the time to CT and diagnosis of lung cancer.”
The study is a prospective, multicenter RCT with block randomization by day and site. The randomization method is described (random sampling via 1:1 randomization for whole-day sessions). Blinding is not explicitly described, but the design (AI available in both arms, prioritization randomized) implies that reporters were not blinded to AI availability, which is appropriate for a pragmatic trial. A power analysis is provided for both primary outcomes. Inclusion/exclusion criteria are implicit (primary care-requested CXRs, age ≥18). Outlier handling is addressed through data cleaning (exclusion of invalid dates). Controls are inherent in the comparator arm. Independent replication is not applicable for a single pivotal trial.
“Participants aged 18 years or older, attending for a primary care-requested CXR, were block-randomized by day and site to either immediate AI prioritization of reporting or no AI prioritization.”
“Based on data from previous work , the median time to lung cancer diagnosis was 63 days in the standard reporting group and using a conservative reduction of 10 days, we calculated that 265 cases per group would be needed to detect a difference with 95% power.”
“The AI was applied to both arms at the time of image acquisition, so that the reporter had access to the AI-marked images at the time of reporting in both arms of the study.”
“Participants aged 18 years or older, attending for a primary care-requested CXR, were block-randomized by day and site to either immediate AI prioritization of reporting or no AI prioritization.”
“Based on data from previous work , the median time to lung cancer diagnosis was 63 days in the standard reporting group and using a conservative reduction of 10 days, we calculated that 265 cases per group would be needed to detect a difference with 95% power.”
The study reports age and sex for the study population, stratified by AI prioritization day. Health status is not directly reported, but the population is primary care patients undergoing CXR, which is a relevant context. Demographics are adequately reported for a human trial. Species/strain and housing conditions are not applicable.
“Age (years) | 59.1 (17.4) | 59.1 (17.5) | 59.1 (17.4) | | Sex | | Male | 21,987 (46.4%) | 21,278 (46.3%) | 43,265 (46.4%)”
“Participants aged 18 years or older, attending for a primary care-requested CXR”
“Age (years) | 59.1 (17.4) | 59.1 (17.5) | 59.1 (17.4) | | Sex | | Male | 21,987 (46.4%) | 21,278 (46.3%) | 43,265 (46.4%) | | Female | 25,352 (53.6%) | 24,709 (53.7%) | 50,061 (53.6%)”
The paper states favorable ethical approval from the East of England—Cambridge East Research Ethics Committee with a protocol number (23/EE/0014). Informed consent was not obtained, but the study was non-consenting by design, with opt-out information provided, as agreed by the Ethics Committee. Regulatory compliance is stated (Good Clinical Practice).
“Favorable ethical approval was obtained from the East of England—Cambridge East Research Ethics Committee (23/EE/0014, 21 February 2023).”
“Consent was not obtained from patients, but in each department, clear messaging indicated that AI was being used as part of a research study, and details on how to opt out of the study were provided, as agreed by the Ethics Committee.”
“The study was undertaken with strict adherence to recommended CONSORT guidelines and Good Clinical Practice .”
“Favorable ethical approval was obtained from the East of England—Cambridge East Research Ethics Committee (23/EE/0014, 21 February 2023).”
“Consent was not obtained from patients, but in each department, clear messaging indicated that AI was being used as part of a research study, and details on how to opt out of the study were provided, as agreed by the Ethics Committee.”
“The study was undertaken with strict adherence to recommended CONSORT guidelines and Good Clinical Practice .”
The AI algorithm qXR (v 4.0, Qure.ai Technologies) is named with version and manufacturer. The statistical software Stata/MP 19.5 is identified. No antibodies, cell lines, or other wet-lab reagents are used, so those criteria are not applicable. The AI algorithm is the key resource and is adequately identified.
“AI algorithm qXR (v 4.0, Qure.ai Technologies, India) is a class IIb CE-certified deep learning algorithm already in routine clinical use in some NHS Hospitals.”
“Statistical analyses were performed using Stata/MP 19.5 (StataCorp).”
“AI algorithm qXR (v 4.0, Qure.ai Technologies, India) is a class IIb CE-certified deep learning algorithm already in routine clinical use in some NHS Hospitals.”
“Statistical analyses were performed using Stata/MP 19.5 (StataCorp).”
The primary analysis uses t-tests on log-transformed outcomes, reported as ratios of geometric means with 95% CIs. Secondary outcomes use chi-squared tests and kappa statistics. Exact p-values are reported for primary outcomes. Effect sizes with CIs are reported. Statistical software is identified. Data presentation includes medians and IQRs, and per-group n are stated. Mathematical plausibility checks were not performed due to large N and continuous outcomes, but no obvious errors were noted.
“A t test on the log transformation of these data was used to test the null hypothesis that mean values on the log-transformed scale are equal.”
“There was no significant difference in the time from CXR acquisition to CT scan according to AI prioritization, with a ratio of geometric means of 0.97 (95% confidence interval (CI) = 0.93–1.02, P = 0.31).”
“Statistical analyses were performed using a two-sided t- test on log-transformed outcomes and presented as the ratio of geometric means and 95% CIs.”
“There was no significant difference in the time from CXR acquisition to CT scan according to AI prioritization, with a ratio of geometric means of 0.97 (95% confidence interval (CI) = 0.93–1.02, P = 0.31).”
“There was no significant difference in the time from CXR acquisition to lung cancer diagnosis between AI prioritization days, with a ratio of geometric means of 0.98 (95% CI = 0.83–1.16, P = 0.84; Table ).”
The data availability statement provides a concrete route: data can be requested from the sponsor with a timeframe (within two months). The statistical analysis code is available on GitHub. The AI algorithm code is not available as it is commercial, which is acceptable. Repository deposit and accession numbers are not applicable for patient-level data.
“Access to fully anonymised data can be requested from the sponsor at: Research & Innovation, Nottingham University Hospitals NHS Trust, Queen’s Medical Centre, B Floor, Medical School, Derby Road, Nottingham NG7 2UH, UK. Email: nuhnt.researchsponsor@nhs.net. Data would normally be available within two months of the request.”
“The code for the statistical analysis is available from L.S. and has been uploaded to GitHub; see https://github.com/LFairleySmith/LungIMPACT .”
“All data for the study is stored by the study sponsor. Access to fully anonymised data can be requested from the sponsor at: Research & Innovation, Nottingham University Hospitals NHS Trust, Queen’s Medical Centre, B Floor, Medical School, Derby Road, Nottingham NG7 2UH, UK. Email: nuhnt.researchsponsor@nhs.net. Data would normally be available within two months of the request.”
“The code for the statistical analysis is available from L.S. and has been uploaded to GitHub; see https://github.com/LFairleySmith/LungIMPACT .”
The trial is registered with ISRCTN (78987039). The paper adheres to CONSORT guidelines. All pre-specified outcomes are reported, including negative results. Limitations are extensively discussed. Conclusions are proportional to the evidence, with appropriate caveats. Funding and competing interests are declared.
“ISRCTN registration: 78987039 (https://www.isrctn.com/ISRCTN78987039) .”
“The study was undertaken with strict adherence to recommended CONSORT guidelines and Good Clinical Practice .”
“A limitation of the study is that the primary outcomes evaluated only the impact of AI prioritization, not the presence or absence of AI in the clinical workflow.”
“ISRCTN registration: 78987039 (https://www.isrctn.com/ISRCTN78987039) .”
“The study was undertaken with strict adherence to recommended CONSORT guidelines and Good Clinical Practice .”
“A limitation of the study is that the primary outcomes evaluated only the impact of AI prioritization, not the presence or absence of AI in the clinical workflow.”
Registered (1 ID: ISRCTN). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 39 references by DOI: 5 verified — 34 no DOI (shown, not verified).
- NO DOIInternational differences in lung cancer survival by sex, histological type and stage at diagnosis: an ICBP SURVMARK-2 StudyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIProgress in cancer survival, mortality, and incidence in seven high-income countries 1995-2014 (ICBP SURVMARK-2): a population-based studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILung cancer survival and stage at diagnosis in Australia, Canada, Denmark, Norway, Sweden and the UK: a population-based study, 2004–2007No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILung cancer diagnosis and staging with endobronchial ultrasound-guided transbronchial needle aspiration compared with conventional approaches: an open-label, pragmatic, randomised controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImpact of timing of lobectomy on survival for clinical stage IA lung squamous cell carcinomaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITime to initial cancer treatment in the United States and association with survival over time: an observational studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICancer patients’ concerns regarding access to cancer care: perceived impact of waiting times along the diagnosis and treatment journeyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHow often do patients ask for the results of their radiological studies?No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReviewing imaging examination results with a radiologist immediately after study completion: patient preferences and assessment of feasibility in an academic departmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOptimal Care Pathway for People with Lung CancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINational Optimal Lung Cancer Pathway (Version 4.0)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImpact of radiographer immediate reporting of X-rays of the chest from general practice on the lung cancer pathway (radioX): a randomised controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAchieving earlier diagnosis of symptomatic lung cancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILung cancer stage-shift following a symptom awareness campaignNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGeneral practice chest X-ray rate is associated with earlier lung cancer diagnosis and reduced all-cause mortality: a retrospective observational studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEarly-stage lung cancer associated with higher frequency of chest x-ray up to three years prior to diagnosisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial Intelligence for Analysing Chest X-ray ImagesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDiagnostic Imaging Dataset Annual Statistical Release 2023/24No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILife Sciences Competitiveness Indicators 2024: SummaryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICan incorrect artificial intelligence (AI) results impact radiologists, and if so, what can we do about it? A multi-reader pilot study of lung cancer detection with chest radiographyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEstimating lung cancer risk from chest X-ray and symptoms: a prospective cohort studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIQuality assurance in radiology: peer review and peer feedbackNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICollaborative learning in radiology: from peer review to peer learning and peer coachingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence in neuroradiology: a smart prospective peer reviewerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIUtility of artificial intelligence tool as a prospective radiology peer reviewer—detection of unreported intracranial hemorrhageNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReporting radiographer peer review systems: a cross-sectional survey of London NHS trustsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRadiograph accelerated detection and identification of cancer in the lung (RADICAL): a mixed methods study to assess the clinical effectiveness and acceptability of Qure.ai artificial intelligence software to prioritise chest X-ray (CXR) interpretationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRAIQC supports evaluation of Lunit INSIGHT CXR AI tool for enhancing clinician chest X-ray interpretationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDiagnostic effect of artificial intelligence solution for referable thoracic abnormalities on chest radiography: a multicenter respiratory outpatient diagnostic cohort studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffect of a comprehensive deep-learning model on the accuracy of chest x-ray interpretation by radiologists: a retrospective, multireader multicase studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence in chest radiography reporting accuracy: added clinical value in the emergency unit setting without 24/7 radiology coverageNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEarly clinical evaluation of AI triage of chest radiographs: time to diagnosis for suspected cancer and number of urgent CT referralsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStatistical Power Analysis for the Behavioural Science, Second EditionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extensionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- codeGitHubLIVEHTTP 200https://github.com/LFairleySmith/LungIMPACTResolves to GitHub (code repository).
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, clarity.
- MINORconsistencyAbstract“AI prioritization of CXR requested by UK primary care has no significant impact on the lung cancer pathway.”→ Consider rephrasing to 'AI prioritization of CXRs requested by UK primary care' for grammatical consistency.Minor grammatical issue.
- MINORconsistencyResults, Secondary outcomes“The median time to initiation of cancer treatment was 76 days (IQR = 38–114 days) for AI prioritization days and 73 days (IQR = 43–121 days) for those with no AI prioritization (Table ).”→ Ensure the order of groups is consistent throughout the paper (e.g., 'AI prioritization' vs 'no AI prioritization').Minor inconsistency in group order.
- MINORclarityDiscussion, paragraph 5“Some of these will be from the same CXR (proportion of FP plus FN of all reviewed CXRs was 33,910/28,261; ratio = 1.2:1).”→ Clarify that the ratio is not a proportion but a ratio of counts, and consider rephrasing for clarity.Potentially confusing phrasing.
- MINORconsistencyAbstract“AI prioritization of CXR requested by UK primary care has no significant impact on the lung cancer pathway.”→ Consider rephrasing for clarity: 'AI prioritization of CXRs requested by UK primary care had no significant impact on the lung cancer pathway.'Minor grammatical issue.
- MINORconsistencyTable 2“Time to cancer treatment starting (days) | 200 | 72.5 | (43–120.5) | 200 | 76 | (37–114) | 1.00 | (0.84–1.19) | 0.99”→ Ensure consistent formatting of decimal places (e.g., 72.5 vs 76).Minor formatting inconsistency.
- MINORclarityDiscussion“This is one reason why NICE has not recommended any AI products for CXR interpretation in England .”→ Consider adding a reference to the NICE guidance for completeness.Reference missing.
The published work is robust and well-reported. An informed reader should weigh the minor reporting gaps (blinding description, inclusion/exclusion criteria, exact p-values for secondary outcomes) but these do not undermine the main conclusions. No erratum or re-analysis is warranted based on the available evidence.
- 1.HIGHrigorIn the Methods, explicitly state the blinding status of radiologists and outcome assessors, and describe any measures taken to minimize bias in outcome assessment.Blinding is not described; while the pragmatic design and objective outcomes reduce bias risk, explicit reporting is expected for RCTs.
- 2.HIGHrigorIn the Methods, provide explicit inclusion and exclusion criteria for CXRs (e.g., age, view type, primary care request) and for patients (e.g., multiple CXRs handling).Incomplete inclusion/exclusion criteria reduce reproducibility and transparency.
- 3.MEDIUMstatisticsIn the Results, report exact p-values for secondary outcomes where only thresholds are given (e.g., P < 0.001).Exact p-values allow readers to assess the strength of evidence more precisely.
- 4.MEDIUMstatisticsIn the Methods, state the statistical assumptions for the t-test on log-transformed data (e.g., normality of log-transformed outcomes) and how they were verified.Explicitly verifying assumptions strengthens the statistical analysis.
- 5.MEDIUMreportingIn the Results, report the number of participants who opted out and any baseline characteristics of excluded CXRs to enhance transparency.Transparency about exclusions and opt-outs improves the reader's ability to assess generalizability.
- 6.MEDIUMreportingIn the Methods, provide more detail on the data cleaning process, including the specific criteria for excluding CXRs due to 'data compliance issues'.Clarifying exclusion criteria for data cleaning improves reproducibility.
- 7.MEDIUMreportingIn the Results, report the number of CXRs excluded due to invalid dates and the reasons for these exclusions.Detailed exclusion counts help readers understand the flow of participants.
- 8.MEDIUMreportingIn the Discussion, consider discussing the potential impact of the lack of blinding on the results and any steps taken to mitigate it.Addressing blinding limitations proactively strengthens the discussion.
- 9.MEDIUMreportingIn the Discussion, consider discussing the generalizability of the findings to other AI algorithms and healthcare settings beyond the English NHS.Generalizability discussion is important for clinical adoption.
- 10.MEDIUMreportingIn the Methods, provide a reference for the CONSORT-AI extension and how it was applied.Referencing the specific reporting guideline enhances transparency.
- 11.MEDIUMdata codeIn the Data availability statement, consider providing a data access request form or a link to a data access committee for clarity.A clear data access process facilitates data sharing.
- 12.MEDIUMdata codeIn the Code availability, consider providing a versioned release or DOI for the GitHub repository to ensure permanence.A DOI or versioned release ensures the code remains accessible and citable.
- 13.LOWcopyeditIn the Abstract, rephrase 'AI prioritization of CXR requested by UK primary care' to 'AI prioritization of CXRs requested by UK primary care' for grammatical consistency.Minor grammatical fix improves clarity.
- 14.LOWcopyeditIn the Results, ensure the order of groups is consistent throughout the paper (e.g., 'AI prioritization' vs 'no AI prioritization').Consistent group ordering improves readability.
- 15.LOWcopyeditIn the Discussion, clarify that the ratio 33,910/28,261 is a ratio of counts, not a proportion, and consider rephrasing for clarity.Clarifying the ratio avoids confusion.
- 16.LOWcopyeditIn Table 2, ensure consistent formatting of decimal places (e.g., 72.5 vs 76).Consistent formatting improves professionalism.
- 17.LOWcopyeditIn the Discussion, add a reference to the NICE guidance for completeness.Missing reference reduces completeness.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.