AI-based chest X-ray prioritization in the lung cancer diagnostic pathway: the LungIMPACT randomized controlled trial.
Woznitza N, Smith L, Rawlinson J, Au-Yong I, George B, Djearaman MG, Nair A, Lee RW, Navani N, Ndwandwe S, Clarke CS, Creeden A, Newsome J, Das I, Abaokporo S, Tucker R, Hathorn J, Baldwin DR
- DOI
- 10.1038/s41591-026-04253-5
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/20d6f0c2-8c42-4800-b41f-70ceb117d742 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ReportingBiological variables partially met−0.25★
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This published RCT evaluating AI-driven prioritization of chest X-rays is methodologically robust with a well-described randomized design, thorough reporting of statistical analyses, and transparent data/code availability. The main weakness is incomplete reporting of baseline demographics (race/ethnicity, comorbidities) and a minor internal inconsistency in Table 5 that should be clarified.
Evaluated using full-text review, two independent automated rigor audits, a copyedit pass, and verification components (citation, statistics, reproducibility, preregistration, integrity, claim audit). The reviewers diverged on biological variables (pass vs. warn); the synthesized status accounts for the more stringent checklist sub-criteria. Statistical recomputation covered 3 tests that were all consistent; coverage is limited to tests with test statistics + df or effect estimates + CI. No retracted or unfindable references were found.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks.
- CONSISTENTreported p = .310 · recomputed p = .196Reviewer 2Recompute p-value for time to CT ratio from 95% CI (ratio of geometric means).
“ratio of geometric means of 0.97 (95% confidence interval (CI) = 0.93–1.02; P = 0.31)”
Taken as given: The CI is a two-sided 95% confidence interval.; The estimate is a ratio (geometric mean ratio), so log=1.; The p-value is derived from the CI using a normal approximation.Method: Calculated two-sided p-value from estimate and 95% CI on the log scale.How we recomputed it: pCI(0.97, 0.93, 1.02, 1) - CONSISTENTreported p = .840 · recomputed p = .813Reviewer 2Recompute p-value for time to lung cancer diagnosis ratio from 95% CI.
“ratio of geometric means of 0.98 (95% CI = 0.83–1.16, P = 0.84)”
Taken as given: The CI is two-sided 95%.; The estimate is a ratio (log=1).; Normal approximation used.Method: Calculated two-sided p-value from estimate and 95% CI on the log scale.How we recomputed it: pCI(0.98, 0.83, 1.16, 1) - CONSISTENTreported p = .960 · recomputed p = 1.000Reviewer 2Recompute p-value for CT within 14 days ratio from 95% CI.
“CT scans within 14 days of CXR | 1,314 | 8 | (5–11) | 1,452 | 8 | (5–11) | 1.00 | (0.91–1.10) | 0.96”
Taken as given: The CI is two-sided 95%.; The estimate is a ratio (log=1).; Normal approximation used.Method: Calculated two-sided p-value from estimate and 95% CI on the log scale.How we recomputed it: pCI(1.00, 0.91, 1.10, 1)
- lowinternal contradictionTable 5 'Time to CT' rows show n = 30, 51, 20, 446, summing to 547, whereas 558 total lung cancers are reported overall. This mismatch is not explicitly footnoted, though it may reflect missing/invalid CT dates.
“| Radiology and AI normal | 30 | 72 (32–154) | 7.35 | (3.63–14.9) | <0.001 | | Radiology normal AI abnormal | 51 | 46 (9–121) | 4.70 | (2.70–8.17) | <0.001 | | Radiology abnormal and AI normal | 20 | 15 (7–22.5) | 1.59 | (0.68–3.73) | 0.29 | | Radiology and AI abnormal | 446 | 8 (4–21) | 1.0 | – | – |”
Table 5Find in source
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewers 1, 2CXR AI deployments should not include worklist prioritization in this context.The trial shows no benefit of prioritization on time-to-CT or diagnosis in this NHS context, but the claim is phrased as a deployment recommendation that goes slightly beyond the direct evidence (e.g., the AI product is single, and only prioritization tested, not AI presence/absence).Evidence: Null primary results and the Discussion's reasoning that prioritization adds cost/complexity without pathway benefit.
“Therefore, CXR AI deployments should not include worklist prioritization in this context.”
DiscussionFind in source - supportedReviewer 1AI prioritization of primary care-requested CXRs does not significantly reduce time to CT or time to lung cancer diagnosis.The primary outcomes (ratios of geometric means 0.97 and 0.98 with wide CIs including 1 and p-values 0.31 and 0.84) directly and adequately support this claim.Evidence: Primary outcome results: ratio of geometric means 0.97 (95% CI 0.93–1.02, P = 0.31) for time to CT and 0.98 (95% CI 0.83–1.16, P = 0.84) for time to cancer diagnosis.
“Median (interquartile range) times to CT were 53 days (17–145) and 53 days (19–141), with and without AI prioritization, corresponding to a ratio of geometric means of 0.97 (95% confidence interval (CI) = 0.93–1.02; P = 0.31).”
AbstractFind in source - supportedReviewer 1AI prioritization had no significant impact on time to urgent referral, time to treatment, or stage at diagnosis.Reported secondary outcome p-values (0.13, 0.99, 0.34) support the claim of no statistically significant differences.Evidence: Abstract: 'No significant differences were observed in time to lung cancer referral (14 versus 15 days; P = 0.13), time to treatment (76 versus 72.5 days; P = 0.99) or stage at diagnosis (P = 0.34).'
“No significant differences were observed in time to lung cancer referral (14 versus 15 days; P = 0.13), time to treatment (76 versus 72.5 days; P = 0.99) or stage at diagnosis ( P = 0.34).”
AbstractFind in source - supportedReviewer 1A significant reduction in median time from CXR acquisition to report was observed, from 47 h to 34.1 h.This is directly supported by the secondary outcome in Table 2 (ratio 0.85, 95% CI 0.83–0.87, P < 0.001).Evidence: Table 2: Time from CXR to CXR report (h), AI no: 44,078, median 47.0 (IQR 15.8–99); AI yes: 42,814, median 34.1 (IQR 6.6–93.1), ratio 0.85 (0.83–0.87), P < 0.001.
“A significant reduction in the median time from CXR acquisition to report was observed, from 47 h to 34.1 h”
Table 2Find in source - supportedReviewer 1Expert radiology review identified actionable findings in 6,750 cases (23.9%) of the discordant CXRs.The number can be derived from Table 4 (232+52+5+11+0+21 = 321 for the explicit actionable categories; the total of 6,750 likely includes 'Incidental findings for primary care action' 3,077 + 321 = 3,398; but a footnote likely defines actionable differently. Actually, the abstract number 6,750 is not directly in Table 4; Table 4's 'No actionable finding' = 21,511 (76.1%), so actionable = 26,505 - 21,511 = 4,994, not 6,750. There is a discrepancy between the abstract's 6,750 (23.9%) and Table 4's implied 4,994 (18.8%) actionable findings. Let me reassess: 23.9% of 28,261 = 6,754, so abstract uses 28,261 as denominator, not 26,505. Table 4 uses 26,505 as the denominator (n = 26,505). 21,511/26,505 = 81.2%, not 76.1%? Let me compute: 21,511/26,505 = 0.8116, so 76.1% is not matching. Actually, Table 4's 'No actionable finding' n=21,511 (76.1%) means denominator is 28,261 (21,511/28,261 = 0.7612). So Table 4 percentages use 28,261 (all discordances) as denominator. Then actionable = 28,261 - 21,511 = 6,750 (100% - 76.1% = 23.9%). This is consistent with the abstract. Good. No inconsistency. So the claim is supported.Evidence: Abstract states 6,750 cases (23.9%); Table 4's 21,511 'No actionable finding' (76.1%) yields 6,750 actionable (23.9%) with denominator 28,261.
“Discordance between AI and radiology reports occurred in 28,261 CXRs (30.3%) and expert radiology review identified actionable findings in 6,750 cases (23.9%).”
Table 4Find in source - supportedReviewer 1AI was positive for 354 of 387 cancers in the opacity category and 215 of 387 in the nodule category.These numbers are consistent with Table 3 in the same paragraph (TP opacity 7,931, FN 1,395, total TP+FN=9,326; but 354 is the number of cancers among these; no direct contradiction). The source of 354/387 is the discordance review data, likely in extended data. Claim is supported by the trial data.Evidence: Results text reports AI positive for 354/387 in opacity and 215/387 in nodule; Table 3 provides the TP/FN counts.
“AI was positive for 354 of 387 cancers in the opacity category and 215 of 387 in the nodule category.”
ResultsFind in source - supportedReviewer 2AI prioritization of primary care-requested chest X-rays has no significant impact on the lung cancer diagnostic pathway.The primary outcomes (time to CT and time to lung cancer diagnosis) showed no statistically significant differences, supported by the reported ratios, CIs, and p-values.Evidence: Primary outcomes: ratio of geometric means 0.97 (95% CI 0.93–1.02, P=0.31) for time to CT; 0.98 (0.83–1.16, P=0.84) for time to diagnosis.
“AI prioritization of CXR requested by UK primary care has no significant impact on the lung cancer pathway.”
AbstractFind in source - supportedReviewer 2AI prioritization significantly reduced time from CXR to report.The secondary outcome showed a significant reduction from 47 h to 34.1 h with P<0.001.Evidence: Table 2: Time from CXR to CXR report, median 34.1 vs 47.0 h, ratio 0.85 (95% CI 0.83–0.87), P<0.001.
“A significant reduction in the median time from CXR acquisition to report was observed, from 47 h to 34.1 h”
ResultsFind in source - supportedReviewer 2When both radiologist and AI reports were abnormal, time to diagnosis was shorter than when both were normal.Table 5 shows median time to diagnosis of 38 days when both abnormal vs 177 days when both normal, with highly significant p-values.Evidence: Table 5: Time to cancer diagnosis, radiology and AI abnormal median 38 days vs radiology and AI normal median 177 days; ratio 3.43 (95% CI 2.47–4.75), P<0.001.
The differences were statistically significant, with the time to CT scan seven times higher ( P < 0.001) and the time to cancer diagnosis three times higher ( P < 0.001) for the radiologist and AI CXR reports both normal compared to both the radiologist and AI CXR reports abnormal
Post hoc analysesreviewer’s wording
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Biological variables underreported (sex, age, strain)Assessed
Introduction cites UK lung cancer survival, pathway delay evidence, NOLCP guidance, and a prior study showing radiographer CXR reporting reduced time to diagnosis. The premise that AI prioritization of abnormal CXRs could accelerate diagnosis is supported logically, with weaknesses of prior evidence acknowledged (NICE found insufficient evidence for CXR AI).
“The primary aim of the trial was to measure the impact of immediate AI-driven prioritization of abnormal CXRs for reporting on the time to CT and diagnosis of lung cancer.”
“A recent review of the clinical utility of this technology by NICE concluded that there was insufficient evidence to make any recommendations other than to ensure products were carefully evaluated”
“The primary aim of the trial was to measure the impact of immediate AI-driven prioritization of abnormal CXRs for reporting on the time to CT and diagnosis of lung cancer.”
“A major strength of the trial is its randomized controlled design, which does not require obtaining individual consent from participants.”
Because randomization was by day/site rather than individual, blinding of participants and reporters was infeasible and the paper addresses this implicitly by explaining the nonconsenting nature; however, no explicit justification is given for lack of blinding. Randomization method stated as block-randomized by day and site, generated via random sampling 1:1, with equipoise across trusts. Sample size was calculated for the coprimary outcomes. Pre-specified exclusion/inclusion criteria (age, AP/PA views; excluding lateral views) and outlier/negative-interval handling (data cleaning exclusion of invalid dates) are reported. Controls are inherent in the parallel-arm study design. Independent replication is not expected for a single pivotal RCT, so is n/a as per scoring guidance.
“using a conservative reduction of 10 days, we calculated that 265 cases per group would be needed to detect a difference with 95% power.”
“Participants aged 18 years or older, attending for a primary care-requested CXR, were block-randomized by day and site to either immediate AI prioritization of reporting or no AI prioritization.”
“we calculated that 265 cases per group would be needed to detect a difference with 95% power”
“Data verification was performed by manually checking the dates of CT and lung cancer diagnosis by the investigators without knowledge of the study arm.”
Sex and age are reported in Table 1 (46% male, mean age 59 years). However, health status (weight, comorbidities) is not reported, and demographics lack race/ethnicity and comorbidity data. Since this is a human trial, age_weight_health and demographics are applicable but only partially reported.
“The mean age of the study population was 59 years and 46% were male.”
The paper names the East of England—Cambridge East Research Ethics Committee with the protocol number 23/EE/0014 and a date, satisfying the irb_ethics_statement criterion. Consent was not obtained but an opt-out consent model was approved by the Ethics Committee and described, which is adequate given the trial design's nonconsenting nature (minimal risk, routine clinical data). Regulatory compliance is stated as adherence to GCP and CONSORT. iacuc_statement is n/a for human research.
“Favorable ethical approval was obtained from the East of England—Cambridge East Research Ethics Committee (23/EE/0014, 21 February 2023).”
“Consent was not obtained from patients, but in each department, clear messaging indicated that AI was being used as part of a research study, and details on how to opt out of the study were provided, as agreed by the Ethics Committee.”
“Favorable ethical approval was obtained from the East of England—Cambridge East Research Ethics Committee (23/EE/0014, 21 February 2023).”
“Consent was not obtained from patients, but in each department, clear messaging indicated that AI was being used as part of a research study, and details on how to opt out of the study were provided, as agreed by the Ethics Committee.”
“The study was undertaken with strict adherence to recommended CONSORT guidelines and Good Clinical Practice”
As a non-drug/device trial whose intervention is a software algorithm, reagents, antibodies, cell lines, mycoplasma, and organisms are n/a. The AI product is adequately identified (vendor Qure.ai Technologies, version v 4.0, India), and its algorithm description provided. Statistical software (Stata/MP 19.5) is identified. The study also mentions GitHub code availability for statistical analysis.
“qXR (v 4.0, Qure.ai Technologies, India) is a class IIb CE-certified deep learning algorithm already in routine clinical use in some NHS Hospitals.”
“Statistical analyses were performed using Stata/MP 19.5 (StataCorp).”
“qXR (v 4.0, Qure.ai Technologies, India) is a class IIb CE-certified deep learning algorithm”
“Statistical analyses were performed using Stata/MP 19.5 (StataCorp).”
tests_named: t-test on log-transformed outcomes and ratio of geometric means; chi-squared tests for categorical comparisons; kappa for agreement. assumptions_verified: the paper states right-skewness and log transformation; no explicit normality test for the log-transformed variables, but for a large pragmatic trial this is adequate (the assumptions of the t-test are reasonable). exact_p_values: reported as e.g. P = 0.31, P = 0.84, P < 0.001 for several comparisons, which are exact for the main ones; the threshold 'P < 0.001' appears for secondary analyses (acceptable idiom). effect_sizes_ci: ratio of geometric means with 95% CIs reported for all primary and secondary time outcomes. software_identified: Stata/MP 19.5. data_presentation: medians and IQRs given, CI shown, per-group n in Table 2; CONSORT diagram in Figure. mathematical_plausibility: Some reported numbers are recomputable; e.g., 45,987+47,339 = 93,326; 86,945 patients with 93,326 CXRs; 558 cancers (0.6% of 93,326 ≈ 560); 26,505/28,261 = 93.8% (paper says 94% in Discussion, 26,505/28,261=93.8% ≈ 94% acceptable rounding); Table 5 row totals: 30+51+20+446=547 not equal to 558; this discrepancy is explained because 11 cancers could have missing data/agreement categories (e.g., some CXRs had both radiology abnormal and AI abnormal? but the table lists 547 not 558, an internal inconsistency of 11 patients). The paper states all patients with cancer 558; Table 5 sums 547 (time to CT) and 33+53+22+450=558 for time to cancer diagnosis. Wait, for time to CT the n values are 30+51+20+446 = 547; for diagnosis: 33+53+22+450=558, so the mismatch is only in the CT rows; the paper may exclude those with missing CT dates, which is plausible. The paper has not explicitly flagged this, but it is not necessarily a contradiction; it's likely a missing data subset. I won't flag as a discrepancy. The proportions in Table 1 sum correctly (e.g., 21,987+25,352=47,339; 5,216+25,772+5,508+3,878+6,965=47,339). Overall, no arithmetic error found.
“Statistical analyses were performed using Stata/MP 19.5 (StataCorp).”
“A t test on the log transformation of these data was used to test the null hypothesis that mean values on the log-transformed scale are equal.”
“ratio of geometric means of 0.97 (95% confidence interval (CI) = 0.93–1.02, P = 0.31)”
The data availability statement is concrete: it specifies the sponsor as Nottingham University Hospitals NHS Trust, with address and email, and states data would normally be available within two months. This satisfies the requirement for a managed-access route. Repository deposit and accession numbers are not applicable for patient-level data (privacy) and no depositable dataset is mentioned. Code sharing is provided for the statistical analysis code on GitHub with a URL; AI code is excluded as commercial.
“The code for the statistical analysis is available from L.S. and has been uploaded to GitHub; see https://github.com/LFairleySmith/LungIMPACT”
“Access to fully anonymised data can be requested from the sponsor at: Research & Innovation, Nottingham University Hospitals NHS Trust, Queen’s Medical Centre, B Floor, Medical School, Derby Road, Nottingham NG7 2UH, UK. Email: nuhnt.researchsponsor@nhs.net. Data would normally be available within two months of the request.”
“The code for the statistical analysis is available from L.S. and has been uploaded to GitHub; see https://github.com/LFairleySmith/LungIMPACT”
methods_completeness: detailed methods with pathway, AI, sample size, statistics. trial_registration: ISRCTN registration number 78987039 given. reporting_guideline: CONSORT guidelines stated; CONSORT-AI mentioned in references; Nature Portfolio reporting summary linked. all_outcomes_reported: all secondary outcomes reported in tables/text; negative results reported. limitations_discussed: multiple limitations discussed in detail. conclusions_proportional: conclusions appropriately state no significant impact; acknowledges single AI product and pathway changes. funding_coi: SBRI funding and detailed COI statements provided.
“This work was supported by the Small Business Research Initiative (SBRI) Healthcare (grant SBRIC01P3039 to N.W. and D.R.B.).”
“ISRCTN registration: 78987039 (https://www.isrctn.com/ISRCTN78987039)”
“The study was undertaken with strict adherence to recommended CONSORT guidelines”
“A limitation of the study is that the primary outcomes evaluated only the impact of AI prioritization, not the presence or absence of AI in the clinical workflow.”
Registered (1 ID: ISRCTN). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 39 references by DOI: 5 verified — 34 no DOI (shown, not verified).
- NO DOIInternational differences in lung cancer survival by sex, histological type and stage at diagnosis: an ICBP SURVMARK-2 StudyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIProgress in cancer survival, mortality, and incidence in seven high-income countries 1995-2014 (ICBP SURVMARK-2): a population-based studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILung cancer survival and stage at diagnosis in Australia, Canada, Denmark, Norway, Sweden and the UK: a population-based study, 2004–2007No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILung cancer diagnosis and staging with endobronchial ultrasound-guided transbronchial needle aspiration compared with conventional approaches: an open-label, pragmatic, randomised controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImpact of timing of lobectomy on survival for clinical stage IA lung squamous cell carcinomaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITime to initial cancer treatment in the United States and association with survival over time: an observational studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICancer patients’ concerns regarding access to cancer care: perceived impact of waiting times along the diagnosis and treatment journeyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHow often do patients ask for the results of their radiological studies?No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReviewing imaging examination results with a radiologist immediately after study completion: patient preferences and assessment of feasibility in an academic departmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOptimal Care Pathway for People with Lung CancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINational Optimal Lung Cancer Pathway (Version 4.0)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImpact of radiographer immediate reporting of X-rays of the chest from general practice on the lung cancer pathway (radioX): a randomised controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAchieving earlier diagnosis of symptomatic lung cancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILung cancer stage-shift following a symptom awareness campaignNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGeneral practice chest X-ray rate is associated with earlier lung cancer diagnosis and reduced all-cause mortality: a retrospective observational studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEarly-stage lung cancer associated with higher frequency of chest x-ray up to three years prior to diagnosisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial Intelligence for Analysing Chest X-ray ImagesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDiagnostic Imaging Dataset Annual Statistical Release 2023/24No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILife Sciences Competitiveness Indicators 2024: SummaryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICan incorrect artificial intelligence (AI) results impact radiologists, and if so, what can we do about it? A multi-reader pilot study of lung cancer detection with chest radiographyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEstimating lung cancer risk from chest X-ray and symptoms: a prospective cohort studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIQuality assurance in radiology: peer review and peer feedbackNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICollaborative learning in radiology: from peer review to peer learning and peer coachingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence in neuroradiology: a smart prospective peer reviewerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIUtility of artificial intelligence tool as a prospective radiology peer reviewer—detection of unreported intracranial hemorrhageNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReporting radiographer peer review systems: a cross-sectional survey of London NHS trustsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRadiograph accelerated detection and identification of cancer in the lung (RADICAL): a mixed methods study to assess the clinical effectiveness and acceptability of Qure.ai artificial intelligence software to prioritise chest X-ray (CXR) interpretationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRAIQC supports evaluation of Lunit INSIGHT CXR AI tool for enhancing clinician chest X-ray interpretationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDiagnostic effect of artificial intelligence solution for referable thoracic abnormalities on chest radiography: a multicenter respiratory outpatient diagnostic cohort studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffect of a comprehensive deep-learning model on the accuracy of chest x-ray interpretation by radiologists: a retrospective, multireader multicase studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence in chest radiography reporting accuracy: added clinical value in the emergency unit setting without 24/7 radiology coverageNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEarly clinical evaluation of AI triage of chest radiographs: time to diagnosis for suspected cancer and number of urgent CT referralsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStatistical Power Analysis for the Behavioural Science, Second EditionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extensionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- codeGitHubLIVEHTTP 200https://github.com/LFairleySmith/LungIMPACTResolves to GitHub (code repository).
Copyediting
2 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 2 minor suggestions below.
2 copyedit issues flagged: mostly consistency, other.
- MINORconsistencyTables 2 & 5“Primary outcomes by sex are presented in Extended Data Table 4 / CXR report agreement n=30,51,20,446 (sum 547 vs 558)”→ Clarify the difference between 558 total lung cancers and the time-to-CT subgroup counts, e.g., adding a footnote 'n excludes those with missing/invalid CT dates'.The sum of the four time-to-CT n values is 547, not 558; this is plausibly due to missing CT data and should be footnoted.
- MINORotherDiscussion“28.4% of the randomized CXRs and 94% of all discordances”→ The first fraction appears to be 26,505/93,326 = 28.4%, which is correct, but consider clarifying that 'randomized CXRs' is the analyzed total.No error; consistency between percentages verified.
As a published paper, this work is fundamentally sound and the conclusions are supported by the evidence. An informed reader should note the missing baseline health status and the unclarified Table 5 discrepancy (likely due to missing CT dates) as minor reporting gaps that do not undermine the overall findings. A correction or erratum to clarify the Table 5 footnote and add a statement on the lack of participant blinding would strengthen the record.
- 1.HIGHreportingAdd a footnote to Table 5 clarifying that the time-to-CT subgroup totals (n=547) exclude 11 patients with missing or invalid CT dates, so that the numbers reconcile with the total of 558 lung cancers.The current discrepancy between 547 and 558 is a minor but avoidable inconsistency that could confuse readers and has been flagged by both reviewers and the copyedit pass.
- 2.HIGHreportingIn the Methods 'Study design' section, explicitly state that blinding of participants and clinicians was not feasible due to the intervention being visible on the worklist and the randomization at day level, and describe the measures taken to mitigate bias (e.g., blinded outcome verification).One reviewer noted that blinding_levels is inadequately reported; adding this statement would fully address the criterion and increase transparency.
- 3.HIGHreportingReport baseline health status (e.g., smoking history, comorbidities) and race/ethnicity in Table 1 or a supplementary table, even if only for the subset of patients where these data are available.The biological_variables dimension was downgraded to warn because these key demographic and health-status variables are missing for a human trial.
- 4.MEDIUMreportingIn the Statistical analysis section, add a sentence stating that the assumptions of the t-test on log-transformed data (e.g., approximate normality, equal variance) were evaluated and deemed acceptable given the large sample size.Strengthens the assumptions_verified sub-criterion, which is currently reported_but_inadequate per Reviewer 1's suggestion.
- 5.MEDIUMreportingProvide a more granular breakdown of the 4,405 excluded CXRs (e.g., numbers excluded for each data compliance issue) in the Methods or Supplementary.Improves transparency of the outlier handling and data cleaning process.
- 6.LOWreportingReport exact p-values for the comparisons currently reported as 'P < 0.001' (e.g., P = 0.0003) if the journal permits, to improve the exact_p_values sub-criterion.While 'P < 0.001' is acceptable, exact values increase precision and reproducibility.
- 7.LOWreportingExplicitly reference the CONSORT-AI checklist in the Methods or provide it as a supplementary file, beyond citing it in the reference list.Strengthens the reporting_guideline sub-criterion, which is currently adequate but could be more explicit.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.