AI-based chest X-ray prioritization in the lung cancer diagnostic pathway: the LungIMPACT randomized controlled trial.
Woznitza N, Smith L, Rawlinson J, Au-Yong I, George B, Djearaman MG, Nair A, Lee RW, Navani N, Ndwandwe S, Clarke CS, Creeden A, Newsome J, Das I, Abaokporo S, Tucker R, Hathorn J, Baldwin DR
Paper source
AI-based chest X-ray prioritization in the lung cancer diagnostic pathway: the LungIMPACT randomized controlled trial.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This published RCT evaluating AI-driven prioritization of chest X-rays is methodologically robust with a well-described randomized design, thorough reporting of statistical analyses, and transparent data/code availability. The main weakness is incomplete reporting of baseline demographics (race/ethnicity, comorbidities) and a minor internal inconsistency in Table 5 that should be clarified.
Evaluated using full-text review, two independent automated rigor audits, a copyedit pass, and verification components (citation, statistics, reproducibility, preregistration, integrity, claim audit). The reviewers diverged on biological variables (pass vs. warn); the synthesized status accounts for the more stringent checklist sub-criteria. Statistical recomputation covered 3 tests that were all consistent; coverage is limited to tests with test statistics + df or effect estimates + CI. No retracted or unfindable references were found.
01
Numerical inconsistencies
1 finding · worst low
Values that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Internal contradictions in the reported numbersAssessed
Conclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: all adequately supported.
03
Data authenticity concerns
None found
An adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
04
Reporting gaps
1 finding · worst medium
Required detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Introduction cites UK lung cancer survival, pathway delay evidence, NOLCP guidance, and a prior study showing radiographer CXR reporting reduced time to diagnosis. The premise that AI prioritization of abnormal CXRs could accelerate diagnosis is supported logically, with weaknesses of prior evidence acknowledged (NICE found insufficient evidence for CXR AI).
Sub-criteria
prior work citedADEQUATE
premise rationaleADEQUATE
limitations addressedADEQUATE
Per reviewer (2) — the votes behind the verdict
PASSReviewer 1· 95% conf
The paper establishes a clear premise based on prior literature, including the prior work by the study team, and links the rationale to the trial objectives and hypothesis.
Evidence
paraphrase[Introduction, paragraph 4]
“However, a recent review of the clinical utility of this technology by NICE concluded that there was insufficient evidence to make any recommendations other than to ensure products were carefully evaluated.”
“The primary aim of the trial was to measure the impact of immediate AI-driven prioritization of abnormal CXRs for reporting on the time to CT and diagnosis of lung cancer.”
paraphrase[Conclusion of Introduction / Methods study design]
“A randomized design avoids the confounding in previous before-after service evaluations, as later noted when discussing Storey et al. and the Scottish service evaluation.”
PASSReviewer 2· 90% conf
The study is well grounded in prior research, with a clear rationale linking the premise to the trial objectives and explicit acknowledgement of gaps in prior work.
Evidence
“A recent review of the clinical utility of this technology by NICE concluded that there was insufficient evidence to make any recommendations other than to ensure products were carefully evaluated”
“The primary aim of the trial was to measure the impact of immediate AI-driven prioritization of abnormal CXRs for reporting on the time to CT and diagnosis of lung cancer.”
Because randomization was by day/site rather than individual, blinding of participants and reporters was infeasible and the paper addresses this implicitly by explaining the nonconsenting nature; however, no explicit justification is given for lack of blinding. Randomization method stated as block-randomized by day and site, generated via random sampling 1:1, with equipoise across trusts. Sample size was calculated for the coprimary outcomes. Pre-specified exclusion/inclusion criteria (age, AP/PA views; excluding lateral views) and outlier/negative-interval handling (data cleaning exclusion of invalid dates) are reported. Controls are inherent in the parallel-arm study design. Independent replication is not expected for a single pivotal RCT, so is n/a as per scoring guidance.
Sub-criteria
controlsADEQUATE
power analysisADEQUATE
blinding levelsINADEQUATE
outlier handlingADEQUATE
randomization unitADEQUATE
inclusion exclusionADEQUATE
randomization methodADEQUATE
replicate distinctionN/A
independent replicationN/A
Per reviewer (2) — the votes behind the verdict
PASSReviewer 1· 95% conf
The trial employs a well-described randomized design with pragmatic randomization by day/site, pre-specified sample size calculations, clear inclusion criteria, prespecified analysis of the first CXR, and explicit CONSORT reporting.
Evidence
paraphrase[Methods, Study design]
“block-randomized by day and site to either immediate AI prioritization of reporting or no AI prioritization. ... Pre-allocation to the intervention or routine care was performed using random sampling via a 1:1 randomization method for Monday-to-Friday routine imaging whole-day sessions.”
absence[Methods, Study design]
Blinding not addressed explicitly
“using a conservative reduction of 10 days, we calculated that 265 cases per group would be needed to detect a difference with 95% power.”
Randomization, blinding, power analysis, inclusion/exclusion, and outlier handling are all adequately reported for this human RCT.
Evidence
“Participants aged 18 years or older, attending for a primary care-requested CXR, were block-randomized by day and site to either immediate AI prioritization of reporting or no AI prioritization.”
“Data verification was performed by manually checking the dates of CT and lung cancer diagnosis by the investigators without knowledge of the study arm.”
Sex and age are reported in Table 1 (46% male, mean age 59 years). However, health status (weight, comorbidities) is not reported, and demographics lack race/ethnicity and comorbidity data. Since this is a human trial, age_weight_health and demographics are applicable but only partially reported.
Sub-criteria
demographicsINADEQUATE
sex reportedADEQUATE
sex justifiedN/A
age weight healthINADEQUATE
housing conditionsN/A
species strain sourceN/A
Per reviewer (2) — the votes behind the verdict
PASSReviewer 1· 95% conf
Sex and age are reported for the study population; age, sex, health status (comorbidity) are reported in Table 1; demographic details are reported; sex-specific primary outcomes are given in Extended Data Table 4.
Evidence
paraphrase[Results, Table 1]
“Age (years) | 59.1 (17.4) | 59.1 (17.5) | 59.1 (17.4) | Sex | Male 43,265 (46.4%)”
paraphrase[Results, Primary outcomes]
“Primary outcomes by sex are presented in Extended Data Table 4.”
WARNReviewer 2· 80% conf
Age and sex are reported, but key demographic and health-status variables (race/ethnicity, comorbidities, weight) are not.
Evidence
“The mean age of the study population was 59 years and 46% were male.”
“Paraphrase: Table 1 lists age, sex, hospital trust, and quarter, but no race/ethnicity or comorbidity data.”
The paper names the East of England—Cambridge East Research Ethics Committee with the protocol number 23/EE/0014 and a date, satisfying the irb_ethics_statement criterion. Consent was not obtained but an opt-out consent model was approved by the Ethics Committee and described, which is adequate given the trial design's nonconsenting nature (minimal risk, routine clinical data). Regulatory compliance is stated as adherence to GCP and CONSORT. iacuc_statement is n/a for human research.
Sub-criteria
iacuc statementN/A
informed consentADEQUATE
irb ethics statementADEQUATE
regulatory complianceADEQUATE
Per reviewer (2) — the votes behind the verdict
PASSReviewer 1· 95% conf
Named ethics approval, opt-out consent procedure approved by the Ethics Committee, and stated compliance with CONSORT and Good Clinical Practice are reported.
Evidence
“Favorable ethical approval was obtained from the East of England—Cambridge East Research Ethics Committee (23/EE/0014, 21 February 2023).”
“Consent was not obtained from patients, but in each department, clear messaging indicated that AI was being used as part of a research study, and details on how to opt out of the study were provided, as agreed by the Ethics Committee.”
“Consent was not obtained from patients, but in each department, clear messaging indicated that AI was being used as part of a research study, and details on how to opt out of the study were provided, as agreed by the Ethics Committee.”
As a non-drug/device trial whose intervention is a software algorithm, reagents, antibodies, cell lines, mycoplasma, and organisms are n/a. The AI product is adequately identified (vendor Qure.ai Technologies, version v 4.0, India), and its algorithm description provided. Statistical software (Stata/MP 19.5) is identified. The study also mentions GitHub code availability for statistical analysis.
Sub-criteria
mycoplasma testingN/A
reagents identifiedADEQUATE
organisms identifiedN/A
antibodies identifiedN/A
cell line authenticationN/A
software tools identifiedADEQUATE
Per reviewer (2) — the votes behind the verdict
PASSReviewer 1· 95% conf
The investigational AI product qXR is identified by name (qXR, Qure.ai Technologies), version (v 4.0), regulatory status (class IIb CE-certified), and usage. Statistical software identified.
Evidence
“qXR (v 4.0, Qure.ai Technologies, India) is a class IIb CE-certified deep learning algorithm already in routine clinical use in some NHS Hospitals.”
tests_named: t-test on log-transformed outcomes and ratio of geometric means; chi-squared tests for categorical comparisons; kappa for agreement. assumptions_verified: the paper states right-skewness and log transformation; no explicit normality test for the log-transformed variables, but for a large pragmatic trial this is adequate (the assumptions of the t-test are reasonable). exact_p_values: reported as e.g. P = 0.31, P = 0.84, P < 0.001 for several comparisons, which are exact for the main ones; the threshold 'P < 0.001' appears for secondary analyses (acceptable idiom). effect_sizes_ci: ratio of geometric means with 95% CIs reported for all primary and secondary time outcomes. software_identified: Stata/MP 19.5. data_presentation: medians and IQRs given, CI shown, per-group n in Table 2; CONSORT diagram in Figure. mathematical_plausibility: Some reported numbers are recomputable; e.g., 45,987+47,339 = 93,326; 86,945 patients with 93,326 CXRs; 558 cancers (0.6% of 93,326 ≈ 560); 26,505/28,261 = 93.8% (paper says 94% in Discussion, 26,505/28,261=93.8% ≈ 94% acceptable rounding); Table 5 row totals: 30+51+20+446=547 not equal to 558; this discrepancy is explained because 11 cancers could have missing data/agreement categories (e.g., some CXRs had both radiology abnormal and AI abnormal? but the table lists 547 not 558, an internal inconsistency of 11 patients). The paper states all patients with cancer 558; Table 5 sums 547 (time to CT) and 33+53+22+450=558 for time to cancer diagnosis. Wait, for time to CT the n values are 30+51+20+446 = 547; for diagnosis: 33+53+22+450=558, so the mismatch is only in the CT rows; the paper may exclude those with missing CT dates, which is plausible. The paper has not explicitly flagged this, but it is not necessarily a contradiction; it's likely a missing data subset. I won't flag as a discrepancy. The proportions in Table 1 sum correctly (e.g., 21,987+25,352=47,339; 5,216+25,772+5,508+3,878+6,965=47,339). Overall, no arithmetic error found.
Sub-criteria
tests namedADEQUATE
exact p valuesADEQUATE
effect sizes ciADEQUATE
data presentationADEQUATE
software identifiedADEQUATE
assumptions verifiedADEQUATE
mathematical plausibilityN/A
Per reviewer (2) — the votes behind the verdict
PASSReviewer 1· 90% conf
Statistical tests are named (t-test on log-transformed outcomes, chi-squared, kappa, linear regression); exact p-values and 95% CIs reported; software identified; the analysis plan is described. Mathematical plausibility checks are not flagged by the reviewer, though some numbers are independently verifiable.
Evidence
paraphrase[Methods, Statistical analysis]
“A t test on the log transformation of these data was used ... The number of urgent referrals, incidence of lung cancer and stage of lung cancer at diagnosis were compared between the two groups using chi-squared tests.”
paraphrase[Results, Primary outcomes]
“with a ratio of geometric means of 0.97 (95% confidence interval (CI) = 0.93–1.02; P = 0.31).”
“Statistical analyses were performed using Stata/MP 19.5 (StataCorp).”
Statistical tests are named, assumptions handled via log transformation, exact p-values and CIs reported, software identified, and data presentation is thorough.
Evidence
“A t test on the log transformation of these data was used to test the null hypothesis that mean values on the log-transformed scale are equal.”
The data availability statement is concrete: it specifies the sponsor as Nottingham University Hospitals NHS Trust, with address and email, and states data would normally be available within two months. This satisfies the requirement for a managed-access route. Repository deposit and accession numbers are not applicable for patient-level data (privacy) and no depositable dataset is mentioned. Code sharing is provided for the statistical analysis code on GitHub with a URL; AI code is excluded as commercial.
Sub-criteria
code sharingADEQUATE
accession numbersN/A
repository depositN/A
data availability statementADEQUATE
Per reviewer (2) — the votes behind the verdict
PASSReviewer 1· 95% conf
Data availability statement names the sponsor with address, email, and a timeframe (within two months) for access to fully anonymised data; statistical analysis code is shared on GitHub. The AI algorithm code is not available as it is commercial.
Evidence
paraphrase[Data availability]
“Access to fully anonymised data can be requested from the sponsor at: Research & Innovation, Nottingham University Hospitals NHS Trust ... Email: nuhnt.researchsponsor@nhs.net. Data would normally be available within two months of the request.”
“The code for the statistical analysis is available from L.S. and has been uploaded to GitHub; see https://github.com/LFairleySmith/LungIMPACT”
Data availability statement provides a concrete access route, and statistical analysis code is shared on GitHub.
Evidence
“Access to fully anonymised data can be requested from the sponsor at: Research & Innovation, Nottingham University Hospitals NHS Trust, Queen’s Medical Centre, B Floor, Medical School, Derby Road, Nottingham NG7 2UH, UK. Email: nuhnt.researchsponsor@nhs.net. Data would normally be available within two months of the request.”
methods_completeness: detailed methods with pathway, AI, sample size, statistics. trial_registration: ISRCTN registration number 78987039 given. reporting_guideline: CONSORT guidelines stated; CONSORT-AI mentioned in references; Nature Portfolio reporting summary linked. all_outcomes_reported: all secondary outcomes reported in tables/text; negative results reported. limitations_discussed: multiple limitations discussed in detail. conclusions_proportional: conclusions appropriately state no significant impact; acknowledges single AI product and pathway changes. funding_coi: SBRI funding and detailed COI statements provided.
Sub-criteria
funding coiADEQUATE
trial registrationADEQUATE
reporting guidelineADEQUATE
methods completenessADEQUATE
all outcomes reportedADEQUATE
limitations discussedADEQUATE
conclusions proportionalADEQUATE
Per reviewer (2) — the votes behind the verdict
PASSReviewer 1· 95% conf
The paper references CONSORT and CONSORT-AI guidelines, reports the ISRCTN registration, includes a CONSORT diagram, discusses limitations and conflicts, and states funding. All outcomes appear reported.
Methods are complete, trial is registered, CONSORT is followed, all outcomes reported, limitations and COI are disclosed, and conclusions are proportional.
“A limitation of the study is that the primary outcomes evaluated only the impact of AI prioritization, not the presence or absence of AI in the clinical workflow.”
References checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 39 references by DOI: 5 verified — 34 no DOI (shown, not verified).
1 data/code link checked; 1 live.
06
Copyediting
2 minor
Wording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 2 minor suggestions below.
Kaimen Rigor Action ItemsFinal pass before journal submission
As a published paper, this work is fundamentally sound and the conclusions are supported by the evidence. An informed reader should note the missing baseline health status and the unclarified Table 5 discrepancy (likely due to missing CT dates) as minor reporting gaps that do not undermine the overall findings. A correction or erratum to clarify the Table 5 footnote and add a statement on the lack of participant blinding would strengthen the record.
1.
HIGHreporting
Add a footnote to Table 5 clarifying that the time-to-CT subgroup totals (n=547) exclude 11 patients with missing or invalid CT dates, so that the numbers reconcile with the total of 558 lung cancers.
The current discrepancy between 547 and 558 is a minor but avoidable inconsistency that could confuse readers and has been flagged by both reviewers and the copyedit pass.
2.
HIGHreporting
In the Methods 'Study design' section, explicitly state that blinding of participants and clinicians was not feasible due to the intervention being visible on the worklist and the randomization at day level, and describe the measures taken to mitigate bias (e.g., blinded outcome verification).
One reviewer noted that blinding_levels is inadequately reported; adding this statement would fully address the criterion and increase transparency.
3.
HIGHreporting
Report baseline health status (e.g., smoking history, comorbidities) and race/ethnicity in Table 1 or a supplementary table, even if only for the subset of patients where these data are available.
The biological_variables dimension was downgraded to warn because these key demographic and health-status variables are missing for a human trial.
4.
MEDIUMreporting
In the Statistical analysis section, add a sentence stating that the assumptions of the t-test on log-transformed data (e.g., approximate normality, equal variance) were evaluated and deemed acceptable given the large sample size.
Strengthens the assumptions_verified sub-criterion, which is currently reported_but_inadequate per Reviewer 1's suggestion.
5.
MEDIUMreporting
Provide a more granular breakdown of the 4,405 excluded CXRs (e.g., numbers excluded for each data compliance issue) in the Methods or Supplementary.
Improves transparency of the outlier handling and data cleaning process.
6.
LOWreporting
Report exact p-values for the comparisons currently reported as 'P < 0.001' (e.g., P = 0.0003) if the journal permits, to improve the exact_p_values sub-criterion.
While 'P < 0.001' is acceptable, exact values increase precision and reproducibility.
7.
LOWreporting
Explicitly reference the CONSORT-AI checklist in the Methods or provide it as a supplementary file, beyond citing it in the reference list.
Strengthens the reporting_guideline sub-criterion, which is currently adequate but could be more explicit.
Full text (56,371 chars)Analyzed: Aug 10, 2026
Powered by Kaimen Rigor — Alpha1, alpha1science.com Model-assisted review. Not a substitute for expert review.