Evaluation of a Clinical Decision Support System for Imaging Requests: A Cluster Randomized Clinical Trial.
Dijk SW, Wollny C, Barkhausen J, Jansen O, Mildenberger P, Halfmann MC, Stroeder J, Rizopoulos D, Hunink MGM, Kroencke T
- DOI
- 10.1001/jama.2024.27853
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/35bac816-2df1-4d67-ae36-0a56f811541d is authoritative.
How this rating was calculated
- ReportingData & code availability partially met−0.25★
- CitationsUnresolved reference−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 1 reported mean was read, and its group size is not stated where the value is printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- No data or code availability links were detected to verify.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper is a well-conducted cluster randomized trial with strong methodological reporting, including clear randomization, blinding, prespecified outcomes, and appropriate statistical methods. The main weakness is the vague data sharing statement, which lacks a concrete repository or access route. Minor copyedit issues and one unresolved reference are also noted.
Both reviewers agreed on all dimensions, so no divergence needed reconciliation. The statistics verification covered only one test (the primary outcome) and found it consistent; other statistical results were not machine-verified. The citation check flagged one reference as not found in registry, which is a potential fabrication signal.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Found 1 reported test, but none could be recomputed (missing degrees of freedom or sample size).
- UNCOMPUTABLEreported p = .690 · recomputed p = .180Reviewers 1, 2Primary outcome difference-in-differences p-value
“difference-in-differences value of 1.3 percentage points (99% CI, −2.0 to 1.8 percentage points; P = .69)”
Taken as given: The estimate is 1.3 percentage points.; The 99% CI is (-2.0, 1.8).; The CI is two-sided.Method: Recomputed p-value from the estimate and 99% CI assuming a normal approximation.How we recomputed it: pCI(1.3, -2.0, 1.8, 0)
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
4 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2The CDSS did not reduce the number of inappropriate imaging requests ordered by physicians in academic hospital settings.The primary outcome analysis shows a non-significant difference-in-differences, supporting the claim.Evidence: Difference-in-differences of 1.3 percentage points (99% CI, -2.0 to 1.8; P = .69).
“The CDSS did not reduce the number of inappropriate imaging requests ordered by physicians in academic hospital settings.”
Conclusion - supportedReviewer 1The intervention clusters showed a similar reduction in inappropriate imaging requests compared with control clusters.The reported mean differences and difference-in-differences support this claim.Evidence: Intervention mean difference -0.5% (99% CI, -2.4% to 0.4%); control mean difference -1.8% (99% CI, -4.3% to -0.4%); difference-in-differences 1.3 percentage points.
“The intervention clusters showed a similar reduction (mean difference, −0.5% [99% CI, −2.4% to 0.4%]) in inappropriate imaging requests compared with the control clusters (mean difference, −1.8% [99% CI, −4.3% to −0.4%]) and there was a difference-in-differences value of 1.3 percentage points (99% CI, −2.0 to 1.8 percentage points; P = .69), which was not statistically significant.”
Results - supportedReviewers 1, 2The study was not underpowered.The paper states the study achieved the required sample size and was not underpowered, which is consistent with the prespecified power analysis.Evidence: The paper states 'the study was not underpowered' and 'the study achieved the required sample size'.
The current trial did not demonstrate benefit in either the primary or secondary outcomes and the study was not underpowered.
Discussion ¶5reviewer’s wording - supportedReviewer 2Few physicians changed their imaging requests after CDSS feedback.The reported change rate of 1.0% supports the claim.Evidence: Only 95 of 9532 requests (1.0%) were changed.
There were 9532 imaging requests after implementation of the CDSS among the intervention clusters; only 95 (1.0%) of which were changed prior to the final imaging diagnostic test.
Resultsreviewer’s wording
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior studies showing mixed results and limitations, and the rationale for the trial is clearly linked to the need for real-world evidence. The limitations of prior research are acknowledged, and the study design addresses them.
Randomization was stratified by hospital and specialty type, with 13 clusters per arm. Blinding of physicians to appropriateness scores during the preimplementation period is described. A power analysis was prespecified with a clinically relevant difference of 2.5%. Inclusion criteria for departments and imaging requests are defined. The analysis population is intention-to-treat, and missing data are addressed.
“Physicians were blind to the appropriateness scores.”
“Physicians were blind to the appropriateness scores.”
Sex and age are reported for the imaging requests. The study includes both sexes, so sex justification is not applicable. Demographics are adequately reported in Table 2.
“50.1% of imaging requests were for female patients and the mean patient age was 64 years (SD, 17.1 years).”
“50.1% of imaging requests were for female patients and the mean patient age was 64 years (SD, 17.1 years)”
The paper reports approval from medical research ethics committees with protocol numbers for each site. Oral consent was obtained through department chairs. The study followed the Declaration of Helsinki.
“Oral consent was obtained through the chairs of the participating departments.”
“The study was conducted according to the principles of the Declaration of Helsinki.”
“Oral consent was obtained through the chairs of the participating departments.”
“The study was conducted according to the principles of the Declaration of Helsinki.”
The CDSS is described as the European Society of Radiology iGuide, and its basis on ACR criteria is stated. The statistical software R version 4.0.0 is identified. No other biological or chemical resources are applicable.
“the European Society of Radiology iGuide”
“the European Society of Radiology iGuide”
The primary analysis uses difference-in-differences with mixed-effects models, and the Mann-Whitney test for cluster-level comparisons. Exact p-values are reported (e.g., P = .69). Effect sizes are reported with 99% CIs. Software is identified. Data presentation includes per-group Ns and appropriate figures. Mathematical plausibility checks are not applicable due to large N and continuous outcomes.
“difference-in-differences value of 1.3 percentage points (99% CI, −2.0 to 1.8 percentage points; P = .69)”
The paper states 'Data Sharing Statement: See Supplement 4.' This is vague and does not specify a concrete access route. No repository deposit or accession numbers are provided. No custom code is mentioned.
“Data Sharing Statement: See Supplement 4.”
“Data Sharing Statement: See Supplement 4.”
The trial is registered at ClinicalTrials.gov (NCT05490290). CONSORT is followed. All prespecified outcomes are reported. Limitations are discussed in detail. Conclusions are proportional. Funding and conflicts of interest are disclosed.
“ClinicalTrials.gov Identifier: NCT05490290”
“Conflict of Interest Disclosures: Dr Dijk reported receiving research funding from the Gordon and Betty Moore Foundation.”
“ClinicalTrials.gov Identifier: NCT05490290”
“Our study was subject to several limitations.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 49 references by DOI: 41 verified — 1 DOI unresolved, 7 no DOI (shown, not verified).
- UNRESOLVED10.1515/dx-2023-008Features and functions of decision support systems for appropriate diagnostic imagingCited DOI does not resolve to any Crossref record.
- NO DOISources, effects and risks of ionizing radiationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImaging utilization trends and reimbursementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMethodology for ESR iGuide contentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAppropriate use criteria programNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMedicare and Medicaid Programs; CY 2024 Payment Policies Under the Physician Fee Schedule and Other Changes to Part B Payment and Coverage Policies; Medicare Shared Savings Program Requirements; Medicare Advantage; Medicare and Medicaid Provider and Supplier Enrollment Policies; and Basic Health ProgramNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffective radiation dose in adultsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIR: A language and environment for statistical computingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
1 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 1 minor suggestion below.
1 copyedit issue flagged: mostly typo.
- MINORtypoAbstract, Results“difference-in-differences value of 1.3 percentage points (99% CI, −2.0 to 1.8 percentage points; P = .69)”→ Consider rephrasing for clarity: 'the difference-in-differences was 1.3 percentage points'Minor wording issue.
The published work is methodologically robust and transparent, with only minor reporting gaps. An informed reader should weigh the vague data sharing statement and the unresolved reference as minor concerns; no erratum is warranted for the main findings, but the data sharing statement could be clarified in a correction.
- 1.HIGHdata codeIn the Data Sharing Statement section, replace 'See Supplement 4' with a concrete statement specifying a repository (e.g., Dryad, Zenodo) or a managed-access procedure with conditions and a contact mechanism.The current statement is vague and does not meet the standard for data availability, which is a common reviewer concern.
- 2.HIGHotherVerify the reference 'Features and functions of decision support systems for appropriate diagnostic imaging' (DOI: 10.1515/dx-2023-008) as it was not found in any registry; correct or remove it if it is fabricated.An unresolved reference is a potential fabrication signal that must be addressed.
- 3.MEDIUMdata codeIf ethically permissible, deposit de-identified aggregate data in a public repository and provide an accession number in the Data Sharing Statement.Providing a direct data access route enhances transparency and reproducibility.
- 4.MEDIUMdata codeIf custom analysis code was used, share it in a public repository (e.g., GitHub) with a DOI and reference it in the Methods section.Code sharing allows independent verification of the statistical analyses.
- 5.LOWcopyeditIn the Abstract Results, rephrase 'difference-in-differences value of 1.3 percentage points' to 'the difference-in-differences was 1.3 percentage points' for clarity.Minor wording improvement to enhance readability.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.