Neoadjuvant stereotactic body radiation therapy with durvalumab and oleclumab in ER(+)HER2(-) breast cancer: a randomized phase 2 trial.
De Caluwé A, Desmoulins I, Cao K, Remouchamps V, Baten A, Longton E, Peignaux K, Joaquin Garcia A, Venet D, Arecco L, Agostinetto E, Nader-Marta G, Denis Z, Dhont J, Kristanto P, Catteau X, Larsimont D, Salgado R, Poortmans P, Stagg J, Sotiriou C, Piccart M, Ignatiadis M, Romano E, Buisseret L
- DOI
- 10.1038/s41591-026-04453-z
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/694e2a97-f055-47ac-9d27-c74ff97245e4 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×10−0.25★
- CitationsUnresolved reference−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on pathological complete response (pCR) and residual cancer burden (RCB) 0/1, which are surrogate endpoints for long-term clinical outcomes such as event-free survival and overall survival. The paper does not provide evidence linking pCR or RCB to improved survival in this specific context, nor does it demonstrate target engagement at the tested dose that would validate these surrogates as reliable predictors of clinical benefit. The discussion acknowledges that EFS remains the clinically decisive endpoint and longer follow-up is needed.
“while RCB and pCR provide an early signal of activity, EFS remains the clinically decisive endpoint; therefore, longer follow-up and the conduct of future phase III trials will be essential to determine whether the increase in pCR observed with iSBRT + ICI…”
- 02Treatment effect not shown to be clinically meaningful
The reported effect sizes, such as pCR rates of 16.7% to 33.3% in the ITT population, are modest and not anchored to a minimal clinically important difference. The primary endpoint (RCB 0/1) did not reach statistical significance in the ITT population, and the significant pCR increase in the per-protocol population (P=0.04) is a secondary endpoint without predefined alpha control. The absolute increases, while statistically significant in some subgroups, are not explicitly compared to a clinically meaningful threshold.
“In the intention-to-treat population, the primary endpoint, residual cancer burden 0/1 rate, was 35.4% with No_ICI, 45.1% with Single_ICI and 47.9% with Double_ICI, without statistically significant differences. pCR rates were 16.7%, 29.4% and 33.3%,…”
- 03Printed percentage does not match its own count
35.6% is unattainable for n=42 (nearest: 33.3, 35.7%)
“35.6% in Double_ICI”
Per-protocol population, Double_ICI armFind in source - 04Printed percentage does not match its own count
20.88% is unattainable for n=48 (nearest: 20.83, 22.92%)
“20.88% (Double_ICI)”
Safety, Double_ICI armFind in source - 05Printed percentage does not match its own count
37.3% is unattainable for n=48 (nearest: 35.4, 37.5%)
“37.3% (Double_ICI)”
Safety, Double_ICI armFind in source - 06Printed percentage does not match its own count
1% is unattainable for n=48 (nearest: 0, 2%)
“1%”
Safety, Double_ICI arm
6 further findings of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomized phase 2 trial. The paper excels in ethical approvals, data/code availability, and biological variable reporting, with minor gaps in power analysis reporting, exact p-values for exploratory analyses, and explicit reporting guideline adherence.
Both reviewers classified the study as interventional and agreed on all dimension statuses; minor divergences in checklist ratings (power analysis, sex justified, assumptions verified, exact p_values, reporting guideline) were resolved conservatively. The statistics verification covered only 12 tests with test statistics/CI; threshold-only p-values and resampling-based tests were not machine-verifiable, so the absence of detected errors does not confirm overall statistical correctness.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks. 10 printed percentages that do not match their own count.
- PERCENT35.6% is unattainable for n=42 (nearest: 33.3, 35.7%)
“35.6% in Double_ICI”
Per-protocol population, Double_ICI armFind in source - PERCENT20.88% is unattainable for n=48 (nearest: 20.83, 22.92%)
“20.88% (Double_ICI)”
Safety, Double_ICI armFind in source - PERCENT37.3% is unattainable for n=48 (nearest: 35.4, 37.5%)
“37.3% (Double_ICI)”
Safety, Double_ICI armFind in source - PERCENT1% is unattainable for n=48 (nearest: 0, 2%)
“1%”
Safety, Double_ICI arm - PERCENT73.3% is unattainable for n=147 (nearest: 72.8, 73.5%)
“Breast-conserving surgery was performed in 73.3%”
ITT populationFind in source - PERCENT68.9% is unattainable for n=48 (nearest: 68.8, 70.8%)
“68.9% in No_ICI”
No_ICI armFind in source - PERCENT80% is unattainable for n=51 (nearest: 78.4, 80.4%)
“80.0% in Single_ICI”
Single_ICI armFind in source - PERCENT71.1% is unattainable for n=48 (nearest: 70.8, 72.9%)
“71.1% in Double_ICI”
Double_ICI armFind in source - PERCENT60.9% is unattainable for n=147 (nearest: 60.5, 61.2%)
“axillary lymph node dissection was performed in 60.9%”
ITT populationFind in source - PERCENT56.3% is unattainable for n=147 (nearest: 55.8, 56.5%)
“56.3% of the trial”
ITT populationFind in source
- CONSISTENTreported p = .059 · recomputed p = .059Reviewers 1, 2pCR rate comparison No_ICI vs Double_ICI in ITT (chi-squared test)
“pCR rates were 16.7%, 29.4% and 33.3%, respectively ( P = 0.059)”
Taken as given: The pCR counts are 8 (16.7% of 48) for No_ICI and 16 (33.3% of 48) for Double_ICI.; The non-events are 40 and 32 respectively.; The test is a two-sided chi-squared test without continuity correction.Method: Pearson chi-squared test on 2x2 table of pCR counts.How we recomputed it: pChi2x2(8,40,16,32) - CONSISTENTreported p = .040 · recomputed p = .033Reviewers 1, 2pCR rate comparison No_ICI vs Double_ICI in per-protocol (chi-squared test)
“16.3% in No_ICI (95% CI, 5.2–27.3), 32.6% in Single_ICI (95% CI, 18.6–46.6) and 35.6% in Double_ICI (95% CI, 21.6–49.5) (No_ICI versus Double_ICI: P = 0.04)”
Taken as given: The per-protocol N for No_ICI is 49 (16.3% of 49 ≈ 8) and for Double_ICI is 45 (35.6% of 45 ≈ 16).; The non-events are 41 and 29 respectively.; The test is a two-sided chi-squared test without continuity correction.Method: Pearson chi-squared test on 2x2 table of pCR counts.How we recomputed it: pChi2x2(8,41,16,29)
- lowinternal contradictionThe abstract reports pCR rates of 16.7%, 29.4%, and 33.3% for the ITT population, but the per-protocol population pCR rates are 16.3%, 32.6%, and 35.6%. The abstract states 'In the per-protocol population (MammaPrint High Risk, n = 131), pCR rates were 16.3%, 32.6% and 35.6%, respectively ( P = 0.040).' This is consistent with the results section.
“In the per-protocol population (MammaPrint High Risk, n = 131), pCR rates were 16.3%, 32.6% and 35.6%, respectively ( P = 0.040).”
AbstractFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2iSBRT + anti-PD-L1 may convert immune-cold ER+ HER2- BC into more inflamed tumors and improve response, particularly in PD-L1-negative disease.The claim is supported by the observed pCR increases and TME reprogramming in PD-L1-negative tumors, but the primary endpoint was not met and the study is phase 2, so the evidence is suggestive but not definitive.Evidence: pCR rates in PD-L1-negative tumors: 3.4% (No_ICI) vs 28.1% (Single_ICI) and 30.0% (Double_ICI); transcriptomic and IHC changes showing increased immune activation in iSBRT+ICI arms.
“These findings suggest that iSBRT + anti-PD-L1 may convert immune-cold ER + HER2 − BC into more inflamed tumors and improve response, particularly in PD-L1-negative disease.”
AbstractFind in source - supportedReviewers 1, 2The addition of ICI to iSBRT + NACT significantly increased pCR rate in the ITT population.The exploratory analysis shows a significant 14.6% pCR increase (95% CI, 0.7–28.6) when combining ICI arms vs No_ICI, which is supported by the data.Evidence: Exploratory analysis: overall pCR increase of 14.6% (95% CI, 0.7–28.6) in iSBRT+ICI vs iSBRT_only.
“Overall, the addition of ICI in the ITT population led to a significant 14.6% pCR increase (95% CI, 0.7–28.6)”
ResultsFind in source - supportedReviewers 1, 2The combination of CD73 blockade with NACT and anti-PD-L1 failed to improve outcome at surgery in comparison with anti-PD-L1 with NACT.The pCR rates in the Double_ICI arm were not significantly different from Single_ICI, supporting the claim of no added benefit.Evidence: pCR rates: 29.4% (Single_ICI) vs 33.3% (Double_ICI) in ITT; no significant difference reported.
“In our trial, however, the combination of CD73 blockade with NACT and anti-PD-L1 failed to improve outcome at surgery in comparison with anti-PD-L1 with NACT.”
DiscussionFind in source - supportedReviewer 1iSBRT + ICI reprograms the TME toward an inflamed phenotype, particularly in PD-L1-negative tumors.Paired biopsies show increased immune activation (IFN signatures, effector T cells, MHC-I, PD-L1) in iSBRT+ICI arms, especially in PD-L1-negative tumors.Evidence: RNA-seq and IHC analyses of paired baseline and week-6 biopsies show increased inflammatory signatures and MHC-I/PD-L1 upregulation in iSBRT+ICI arms.
“Paired RNA sequencing revealed early, dynamic reprogramming of the TME in PD-L1-negative tumors after iSBRT + ICI, with coordinated activation of IFN signaling, effector T cell programs and immune checkpoint pathways.”
ResultsFind in source - supportedReviewer 2iSBRT + ICI is associated with dynamic modulation of the TME toward an inflamed phenotype.The claim is supported by paired biopsy analyses showing increased immune activation and PD-L1 expression in iSBRT+ICI arms.Evidence: Paired RNA sequencing and IHC analyses showed increased inflammatory signatures, MHC-I, and PD-L1 in iSBRT+ICI arms.
“Paired RNA sequencing revealed early, dynamic reprogramming of the TME in PD-L1-negative tumors after iSBRT + ICI, with coordinated activation of IFN signaling, effector T cell programs and immune checkpoint pathways.”
DiscussionFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on pathological complete response (pCR) and residual cancer burden (RCB) 0/1, which are surrogate endpoints for long-term clinical outcomes such as event-free survival and overall survival. The paper does not provide evidence linking pCR or RCB to improved survival in this specific context, nor does it demonstrate target engagement at the tested dose that would validate these surrogates as reliable predictors of clinical benefit. The discussion acknowledges that EFS remains the clinically decisive endpoint and longer follow-up is needed.
“while RCB and pCR provide an early signal of activity, EFS remains the clinically decisive endpoint; therefore, longer follow-up and the conduct of future phase III trials will be essential to determine whether the increase in pCR observed with iSBRT + ICI translates to an improved EFS.”
- INADEQUATEEffect sizeThe reported effect sizes, such as pCR rates of 16.7% to 33.3% in the ITT population, are modest and not anchored to a minimal clinically important difference. The primary endpoint (RCB 0/1) did not reach statistical significance in the ITT population, and the significant pCR increase in the per-protocol population (P=0.04) is a secondary endpoint without predefined alpha control. The absolute increases, while statistically significant in some subgroups, are not explicitly compared to a clinically meaningful threshold.
“In the intention-to-treat population, the primary endpoint, residual cancer burden 0/1 rate, was 35.4% with No_ICI, 45.1% with Single_ICI and 47.9% with Double_ICI, without statistically significant differences. pCR rates were 16.7%, 29.4% and 33.3%, respectively (P = 0.059).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior studies on ICI in breast cancer, the immune-cold TME of ER+ HER2- tumors, and the immunomodulatory effects of RT and CD73 blockade. The rationale for combining iSBRT with ICI and anti-CD73 is logically developed. Limitations of prior work (e.g., low pCR in ER+ HER2-, limited ICI benefit in PD-L1-negative tumors) are explicitly addressed as the motivation for the trial.
“In triple-negative BC, the addition of anti-PD-1 to NACT increased pCR by 13.6% and improved overall survival by 5.0% at 5 years”
“RT has emerged as a promising strategy to enhance the efficacy of ICI by transforming an immune-cold TME into a more immune-responsive environment”
“Unlike triple-negative BC, the tumor microenvironment (TME) in ER + HER2 − BC is generally less inflamed with a lower presence of tumor-infiltrating lymphocytes (TILs) and lower PD-L1 expression, which may contribute to its lowered sensitivity to ICI”
“In triple-negative BC, the addition of anti-PD-1 to NACT increased pCR by 13.6% and improved overall survival by 5.0% at 5 years”
“RT has emerged as a promising strategy to enhance the efficacy of ICI by transforming an immune-cold TME into a more immune-responsive environment”
“Unlike triple-negative BC, the tumor microenvironment (TME) in ER + HER2 − BC is generally less inflamed with a lower presence of tumor-infiltrating lymphocytes (TILs) and lower PD-L1 expression, which may contribute to its lowered sensitivity to ICI”
Randomization method is described (1:1:1 ratio, web-response system) with stratification factors. The unit is the patient. Blinding is not applicable as the trial is open-label, which is stated. Power analysis is not explicitly reported in the text, but the trial is a phase 2 with prespecified endpoints. Inclusion/exclusion criteria are detailed. Outlier handling is addressed through ITT and per-protocol analyses. Controls are inherent in the No_ICI arm. Independent replication is not applicable for a single trial.
“Patients were randomly assigned (in a 1:1:1 ratio) to the three arms, with stratification based on centrally assessed baseline PD-L1 using the IC score (<1% versus ≥1%).”
“The Neo-CheckRay clinical trial is a prospective, randomized, multicenter, open-label phase 2 trial”
“Patients were randomly assigned (in a 1:1:1 ratio) to the three arms, with stratification based on centrally assessed baseline PD-L1 using the IC score (<1% versus ≥1%).”
“The Neo-CheckRay clinical trial is a prospective, randomized, multicenter, open-label phase 2 trial”
Sex is reported (all female). Age, menopausal status, BMI, and ECOG PS are reported in Table 1. Tumor characteristics (T-stage, N-stage, PD-L1, MammaPrint, histologic subtype, grade, ER/PR, Ki67, sTILs) are extensively reported. Species/strain and housing conditions are not applicable for a human trial. Demographics are reported in Table 1.
“all patients were female based on biological sex.”
“Age, years | Median (IQR) | 50 (41.75; 57) | 48 (42; 54) | 48.5 (41; 57.5)”
“T-stage, n (%) a,b | T1–T2 T3 | 39 (81.25%) 9 (18.75%) | 36 (70.59%) 15 (29.41%) | 38 (79.17%) 10 (20.83%)”
“all patients were female based on biological sex.”
“Age, years | Median (IQR) | 50 (41.75; 57) | 48 (42; 54) | 48.5 (41; 57.5)”
“Menopausal status, n (%) | Premenopausal Postmenopausal | 30 (62.5%) 18 (37.5%) | 37 (72.55%) 14 (27.45%) | 28 (58.33%) 20 (41.67%)”
The trial protocol was approved by named ethics committees in Belgium and France with reference numbers. All patients provided written informed consent. The trial was conducted in accordance with Good Clinical Practice and the Declaration of Helsinki. This is a human interventional trial, so IACUC is not applicable.
“The trial protocol was approved by the ethics committees in Belgium and in France. For Belgium: Commissie Medische Ethiek UZ Brussels/VUB, Laarbeeklaan 101, 1090 Brussels, Reference number: 2019/P/02. For France: CHU de Grenoble, Comité de Protection des personnes, CS 10217 38043, Grenoble Cedex 9, Reference number: 20-JUBO-01.”
“All the patients provided written, informed consent before enrollment.”
“All the authors attest that the trial was conducted in accordance with the protocol, its amendments and the standards of Good Clinical Practice. The protection of all clinical trial subjects was consistent with the principles of the Declaration of Helsinki.”
“The trial protocol was approved by the ethics committees in Belgium and in France. For Belgium: Commissie Medische Ethiek UZ Brussels/VUB, Laarbeeklaan 101, 1090 Brussels, Reference number: 2019/P/02. For France: CHU de Grenoble, Comité de Protection des personnes, CS 10217 38043, Grenoble Cedex 9, Reference number: 20-JUBO-01.”
“All the patients provided written, informed consent before enrollment.”
“All the authors attest that the trial was conducted in accordance with the protocol, its amendments and the standards of Good Clinical Practice. The protection of all clinical trial subjects was consistent with the principles of the Declaration of Helsinki.”
The investigational drugs (durvalumab, oleclumab) are named with doses and regimens. The PD-L1 assay (VENTANA SP263) is identified. Software for RNA-seq analysis is named (Trimmomatic, STAR, Salmon, DESeq2, etc.). Antibodies, cell lines, and mycoplasma testing are not applicable as this is a clinical trial without wet-lab assays.
“Durvalumab was 1,500 mg intravenously once every 4 weeks for four administrations and oleclumab 3,000 mg intravenously once every 2 weeks for four administrations followed by once every 4 weeks for three administrations.”
“The IC score is defined as the percentage of the tumor area occupied by PD-L1-positive immune cells and was assessed using the VENTANA SP263 IHC assay.”
“Trimmomatic: a flexible trimmer for Illumina sequence data”
“Durvalumab was 1,500 mg intravenously once every 4 weeks for four administrations and oleclumab 3,000 mg intravenously once every 2 weeks for four administrations followed by once every 4 weeks for three administrations.”
“The IC score is defined as the percentage of the tumor area occupied by PD-L1-positive immune cells and was assessed using the VENTANA SP263 IHC assay.”
“Trimmomatic: a flexible trimmer for Illumina sequence data”
Statistical tests are named (chi-squared tests, logistic regression, Wilcoxon rank-sum, Wilcoxon signed-rank). Assumptions are handled by design (e.g., chi-squared for proportions). Exact p-values are reported (e.g., P = 0.059, P = 0.040). Effect sizes with confidence intervals are reported. Software is identified (R). Data presentation includes forest plots and bar plots with CIs. Mathematical plausibility is not applicable for large-N continuous outcomes.
“The P values are based on chi-squared tests.”
“the primary endpoint RCB 0/1 rate was 35.4% in No_ICI (95% confidence interval (95% CI), 21.9–48.9)”
“pCR rates were 16.7%, 29.4% and 33.3%, respectively ( P = 0.059)”
“The P values are based on chi-squared tests.”
“pCR rates were 16.7%, 29.4% and 33.3%, respectively ( P = 0.059).”
“the primary endpoint RCB 0/1 rate was 35.4% in No_ICI (95% confidence interval (95% CI), 21.9–48.9)”
A data availability statement is present with concrete access routes: figshare for source data, EGA for raw sequencing data, and a managed-access procedure for additional patient-level data. Code is available via figshare. Repository deposits and accession numbers are provided.
“Data used for the analyses are available via figshare at 10.6084/m9.figshare.29489597”
“Raw sequencing data have been deposited in the European Genome-phenome Archive (EGA) under accession number EGAD50000002552”
“The R code used for the analyses reported in this study is available via figshare at 10.6084/m9.figshare.29489597”
“Data used for the analyses are available via figshare at 10.6084/m9.figshare.29489597”
“Raw sequencing data have been deposited in the European Genome-phenome Archive (EGA) under accession number EGAD50000002552”
“The R code used for the analyses reported in this study is available via figshare at 10.6084/m9.figshare.29489597”
The trial is registered (NCT03875573). Methods are detailed. Limitations are discussed. Conclusions are proportional. Funding and COI are disclosed. Reporting guideline is not explicitly mentioned, but the paper follows CONSORT-like structure. All outcomes are not fully reported as some secondary endpoints are deferred, but this is stated transparently.
“ClinicalTrials.gov identifier: NCT03875573”
“Limitations of our trial include the small sample size, the short follow-up time at this data cut-off and the use of iSBRT in all three arms”
“This trial is an investigator-initiated academic trial sponsored by the Institut Jules Bordet with support by AstraZeneca”
“ClinicalTrials.gov registration: NCT03875573”
“Limitations of our trial include the small sample size, the short follow-up time at this data cut-off and the use of iSBRT in all three arms”
“Additional prespecified secondary endpoints, including EFS and other efficacy outcomes, are not reported in this paper and will be analyzed at a later timepoint when follow-up is sufficiently mature.”
Registered (1 ID: ClinicalTrials.gov). Reporting guidelines cited: CONSORT, REMARK.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 65 references by DOI: 62 verified — 1 DOI unresolved, 2 no DOI (shown, not verified).
- UNRESOLVED10.6084/m9.figshare.29489597Source files and R code for Neoadjuvant stereotactic body radiation therapy with durvalumab and oleclumab in ER + /HER2 − breast cancer: a randomised phase 2 trialCited DOI does not resolve to any Crossref record.
- NO DOIAJCC Cancer Staging ManualNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICommon Terminology Criteria for Adverse Events (CTCAE)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
3 data/code links checked; 3 live.
- datahttps://clinicaltrials.gov/study/NCT03875573LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT03875573LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- dataEGALIVEHTTP 200https://ega-archive.org/datasets/EGAD50000002552Resolves to EGA (data repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity.
- MINORconsistencyAbstract“Single_IC versus No_ICI”→ Single_ICI versus No_ICITypo in figure legend: 'Single_IC' should be 'Single_ICI'.
- MINORconsistencyTable 2“10(20.88)”→ 10 (20.8)Inconsistent decimal places in percentage.
- MINORclarityDiscussion“the use of iSBRT in all three arms, which precludes disentangling the individual contributions of RT and concurrent systemic therapies—including paclitaxel and ICI—to TME modulation and pCR”→ Consider rephrasing for clarity.Long sentence with multiple clauses.
- MINORconsistencyAbstract“Single_IC versus No_ICI”→ Change to 'Single_ICI versus No_ICI' for consistency.Typo in figure legend.
- MINORconsistencyTable 2“10(20.88)”→ Change to '10 (20.8)' for consistency with other percentages.Inconsistent decimal places.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (missing power analysis, threshold-only p-values in exploratory analyses, no explicit CONSORT checklist) as limitations, but none warrant an erratum or independent re-analysis. The copyedit issues are minor and do not affect scientific integrity.
- 1.HIGHreportingAdd the a priori power/sample-size calculation to the Methods section, including assumed effect size, alpha, and power, to justify the 147-patient sample size.The absence of a power analysis is a notable reporting gap for a randomized trial and a common reviewer concern.
- 2.HIGHreportingExplicitly state adherence to a reporting guideline (e.g., CONSORT) in the Methods and provide the completed checklist as supplementary material.The paper follows CONSORT-like structure but does not explicitly reference the guideline, which is a transparency gap.
- 3.HIGHreportingReport all prespecified secondary endpoints (e.g., EFS) in a preliminary form or clearly state the planned timeline for their analysis in the paper.Deferring secondary endpoints without a clear timeline reduces completeness and transparency.
- 4.MEDIUMstatisticsProvide exact p-values for all exploratory analyses instead of threshold values like P < 0.05.Threshold-only p-values are imprecise and limit the reader's ability to assess the strength of evidence.
- 5.MEDIUMstatisticsAdd a statement on verification of statistical assumptions (e.g., normality, equal variance) for the tests used.Assumption verification is not explicitly reported, which is a minor methodological transparency gap.
- 6.MEDIUMotherAdd a brief justification for enrolling only female patients in the Methods or Discussion.The trial enrolls only females, and a justification would address the sex_justified criterion.
- 7.MEDIUMreportingClarify the role of the funder (AstraZeneca) in study design, data collection, analysis, and manuscript preparation in the competing interests section.The current funding statement does not specify the funder's role, which is important for transparency.
- 8.MEDIUMdata codeAdd a data availability statement for the clinical trial protocol and statistical analysis plan in a public repository.The protocol and SAP are not explicitly deposited, which would enhance reproducibility.
- 9.MEDIUMreportingProvide the full list of inclusion/exclusion criteria in the main text or as a supplement for reproducibility.The current criteria are summarized; full criteria would improve reproducibility.
- 10.MEDIUMstatisticsReport the number of patients with missing data for each analysis and the imputation methods used.Missing data handling is not fully described, which is important for interpreting results.
- 11.LOWcopyeditFix the typo 'Single_IC' to 'Single_ICI' in the Abstract and figure legend.Consistency in terminology is important for clarity.
- 12.LOWcopyeditCorrect the percentage '10(20.88)' to '10 (20.8)' in Table 2 for consistent decimal places.Inconsistent decimal places are a minor formatting issue.
- 13.LOWcopyeditRephrase the long sentence in the Discussion about iSBRT and systemic therapies for clarity.The sentence is complex and could be clearer.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.