Phase 3 Trial of Stereotactic Body Radiotherapy in Localized Prostate Cancer.
van As N, Griffin C, Tree A, Patel J, Ostler P, van der Voet H, Loblaw A, Chu W, Ford D, Tolan S, Jain S, Camilleri P, Kancherla K, Frew J, Chan A, Naismith O, Armstrong J, Staffurth J, Martin A, Dayes I, Wells P, Price D, Williamson E, Pugh J, Manning G, Brown S, Burnett S, Hall E
- DOI
- 10.1056/NEJMoa2403365
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/b57ecb64-8fd8-4d54-847a-c3f4fee2a9c5 is authoritative.
How this rating was calculated
- StatisticsImpossible or misreported statistic−1★
- IntegrityIntegrity concern ×2−1★
- StatisticsStatistic did not reproduce−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- No data or code availability links were detected to verify.
- 01Significance claim does not survive recomputationdemonstrable
Check p-value for non-inferiority from HR and 90% CI.
“SBRT was non-inferior to CRT with an unadjusted HR 0.73 (90%CI 0.48, 1.12), p-value for non-inferiority=0.004”
- 02Reported statistic does not recompute
Recomputed HR 1.59 (95% CI 1.18–2.12), reported p<0.001
“HR 1.59 (95%CI 1.18, 2.12), p<0.001”
- 03Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is biochemical/clinical failure, which is a surrogate for clinically meaningful outcomes like metastasis or survival. The paper does not provide evidence linking biochemical failure to patient-important outcomes in this context, nor does it demonstrate target engagement for the surrogate at the tested dose.
“The primary endpoint was freedom from biochemical/clinical failure with a critical hazard ratio for non-inferiority of 1.45.”
- 04Treatment effect not shown to be clinically meaningful
The reported effect is a small absolute difference of 1.43% in biochemical failure-free rates at 5 years, with confidence intervals crossing zero. No minimal clinically important difference is provided, and the effect is not anchored to clinical meaningfulness.
“The estimated absolute difference in the proportion of participants event free in the SBRT group compared with that in the CRT group at 5 years was: 1.43% (90% CI: -0.60, 2.78).”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted phase III randomized controlled trial with rigorous design, clear ethical approvals, and thorough reporting of biological variables and key resources. The main weaknesses are the lack of a specific data availability statement and the absence of a reporting guideline statement. A critical statistical verification finding shows that the reported non-inferiority p-value (0.004) is inconsistent with the recomputed value (0.145), which could affect the interpretation of the primary endpoint.
Both reviewers agreed on study type (interventional) and on all dimensions except statistical analysis, where the deterministic verification component overrode their pass. The statistics verification covered only 8 tests with test statistics or CIs; other p-values (e.g., threshold-only) were not machine-verified. The citation check found no retracted or unresolved references. The integrity check flagged only minor terminology inconsistencies.
Numerical inconsistencies
2 findings · worst highValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Reported statistics do not recomputeRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 8 tests: 6 consistent, 2 inconsistent (1 change significance at p<.05); 1 recomputed directly from the reported test statistics, 7 via agent-written checks.
- INCONSISTENTreported p < .001 · recomputed p = .002Recomputed HR 1.59 (95% CI 1.18–2.12), reported p<0.001
“HR 1.59 (95%CI 1.18, 2.12), p<0.001”
Taken as given: 1.18–2.12 is a two-sided 95% confidence interval for the HR of 1.59, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p<0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.59, 1.18, 2.12, 1) - CONSISTENTreported p = .220 · recomputed p = .223Reviewer 1Check p-value for superiority test from HR and 95% CI.
“A test for superiority was not significant (HR 0.73; 95% CI: 0.44 1.21; p=0.22).”
Taken as given: The HR is 0.73.; The 95% CI is (0.44, 1.21).; The p-value is two-sided.Method: Used pCI to compute two-sided p from HR and 95% CI.How we recomputed it: pCI(0.73, 0.44, 1.21, 1) - CONSISTENTreported p = .110 · recomputed p = .112Reviewers 1, 2Check p-value for RTOG GU toxicity comparison at 5 years.
“At 5 years, RTOG grade ≥2 GU toxicity was seen in 16/355(4.5%) participants who received CRT and 26/355 (7.3%) who received SBRT (p=0.11).”
Taken as given: The numbers 16 and 26 are the event counts.; The denominators are 355 for both groups.; The test is a chi-square test for 2x2 table.Method: Used pChi2x2 with cell counts: CRT events=16, CRT non-events=339, SBRT events=26, SBRT non-events=329.How we recomputed it: pChi2x2(16, 339, 26, 329) - CONSISTENTreported p = .320 · recomputed p = .315Reviewer 1Check p-value for CTCAE GU toxicity comparison at 5 years.
“CTCAE grade ≥2 GU toxicity was reported in 24/357 (6.7%) and 31/355 (8.7%) in the CRT and SBRT groups respectively at 5 years (p=0.32)”
Taken as given: The numbers 24 and 31 are the event counts.; The denominators are 357 and 355 respectively.; The test is a chi-square test for 2x2 table.Method: Used pChi2x2 with cell counts: CRT events=24, CRT non-events=333, SBRT events=31, SBRT non-events=324.How we recomputed it: pChi2x2(24, 333, 31, 324) - CONSISTENTreported p = .370 · recomputed p = .315Reviewers 1, 2Check p-value for RTOG GI toxicity comparison at 5 years.
“At 5 years, RTOG grade ≥2 GI toxicity was seen in 1/355(0.3%) receiving CRT and 3/354(0.8%) receiving SBRT (p=0.37)”
Taken as given: The numbers 1 and 3 are the event counts.; The denominators are 355 and 354 respectively.; The test is a chi-square test for 2x2 table.Method: Used pChi2x2 with cell counts: CRT events=1, CRT non-events=354, SBRT events=3, SBRT non-events=351.How we recomputed it: pChi2x2(1, 354, 3, 351) - CONSISTENTreported p = .430 · recomputed p = .427Reviewer 1Check p-value for CTCAE GI toxicity comparison at 5 years.
“No difference in CTCAE GI grade ≥2 events was noted at 5 years: 6/357 (1.7%) CRT vs 9/355 (2.5%) SBRT (p=0.43)”
Taken as given: The numbers 6 and 9 are the event counts.; The denominators are 357 and 355 respectively.; The test is a chi-square test for 2x2 table.Method: Used pChi2x2 with cell counts: CRT events=6, CRT non-events=351, SBRT events=9, SBRT non-events=346.How we recomputed it: pChi2x2(6, 351, 9, 346) - CONSISTENTreported p = .460 · recomputed p = .463Reviewer 1Check p-value for erectile dysfunction comparison at 5 years.
“At 5 years, 86/296 (29.1%) CRT and 78/296 (26.4%) SBRT participants reported grade ≥2 CTCAE erectile dysfunction (p=0.46).”
Taken as given: The numbers 86 and 78 are the event counts.; The denominators are 296 for both groups.; The test is a chi-square test for 2x2 table.Method: Used pChi2x2 with cell counts: CRT events=86, CRT non-events=210, SBRT events=78, SBRT non-events=218.How we recomputed it: pChi2x2(86, 210, 78, 218)
- lowinternal contradictionThe abstract reports 5-year biochemical/clinical failure-free rate for CRT as 94.6% and SBRT as 95.8%, but the results section reports 'Five-year biochemical failure event-free rates' with the same numbers. However, the abstract says 'biochemical/clinical failure free-rate' while the results say 'biochemical failure event-free rates'. This is a minor inconsistency in terminology.
Abstract: '5-year biochemical/clinical failure free-rate (95% CI) was CRT: 94.6% (91.9%, 96.4%) vs SBRT: 95.8% (93.3%, 97.4%).' Results: 'Five-year biochemical failure event-free rates (95% CI) were 94.6% (91.9, 96.4) for CRT and 95.8% (93.3, 97.4) for SBRT.'
Abstractreviewer’s wording - lowinternal contradictionThe abstract states '874 patients were randomised from 38 centers (CRT=441, SBRT=433)', but the results section says '424/441 randomized to CRT and 414/433 randomized to SBRT received their allocated treatment; 25 received neither study treatment'. The numbers are consistent.
“874 patients were randomised from 38 centers (CRT=441, SBRT=433)”
AbstractFind in source
Overstated conclusions
4 findings · worst criticalConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Significance claim flips when recomputedRecomputed
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
- SIGNIFICANCE OVERSTATEDINCONSISTENTreported p = .004 · recomputed p = .145Reviewers 1, 2Check p-value for non-inferiority from HR and 90% CI.Reported as statistically significant, but recomputing from the paper’s own numbers gives p ≥ 0.05 — the result may not be significant as claimed.
“SBRT was non-inferior to CRT with an unadjusted HR 0.73 (90%CI 0.48, 1.12), p-value for non-inferiority=0.004”
Taken as given: The HR is 0.73.; The 90% CI is (0.48, 1.12).; The p-value is for a one-sided non-inferiority test, but pCI computes a two-sided p from the CI; the reported p is one-sided, so this check is approximate.Method: Used pCI to derive a two-sided p from the HR and 90% CI, then halved for one-sided comparison.How we recomputed it: pCI(0.73, 0.48, 1.12, 1)
6 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 2The reduction in treatment fractions will alleviate burden on healthcare systems.The claim is based on extrapolation from the number of fractions saved, but the paper does not provide a formal health-economic analysis.Evidence: Discussion mentions 'reduce approximately 72,000 fractions across the UK'.
“Transitioning these patients to a five-fraction regimen could reduce approximately 72,000 fractions across the UK.”
Discussion ¶6Find in source - supportedReviewers 1, 2Five-fraction SBRT is non-inferior to CRT for biochemical/clinical failure.The primary endpoint analysis shows non-inferiority with HR 0.73 and p=0.004, supporting the claim.Evidence: Unadjusted HR 0.73 (90% CI 0.48, 1.12), p=0.004 for non-inferiority.
“Five-fraction SBRT is non-inferior to CRT for biochemical/clinical failure”
ConclusionFind in source - supportedReviewers 1, 2SBRT is an efficacious treatment option for low/intermediate risk localized prostate cancer.The high 5-year biochemical control rates and non-inferiority support efficacy.Evidence: 5-year biochemical failure-free rates of 95.8% for SBRT.
“and is an efficacious treatment option for patients with low/intermediate risk localized prostate cancer as defined in this trial eligibility.”
ConclusionFind in source - supportedReviewers 1, 2SBRT results in higher cumulative incidence of late grade ≥2 GU toxicity compared to CRT.The cumulative incidence of late GU toxicity is significantly higher for SBRT (26.9% vs 18.3%, p<0.001).Evidence: Cumulative incidence of late RTOG grade ≥2 GU toxicity: CRT 18.3% vs SBRT 26.9%, HR 1.59, p<0.001.
“For RTOG GU, incidence of late grade ≥2 events to 5 years was 18.3% (95%CI 14.8, 22.5%) and 26.9% (95%CI 22.8, 31.5%) for CRT and SBRT, respectively (HR 1.59 (95%CI 1.18, 2.12), p<0.001).”
Results ¶5Find in source - supportedReviewer 1There is no significant difference in GI toxicity between SBRT and CRT.The cumulative incidence of GI toxicity is similar between arms (p=0.94).Evidence: Cumulative incidence of late RTOG grade ≥2 GI toxicity: CRT 10.2% vs SBRT 10.7%, HR 1.03, p=0.94.
“For RTOG GI, incidence rates of late grade ≥2 to 5 years were 10.2% (95%CI 7.7, 13.5%) and 10.7% (95%CI 8.1, 14.2%) for CRT and SBRT, respectively (HR 1.03 (95%CI 0.68, 1.56), p=0.94.”
Results ¶6Find in source - supportedReviewer 1The results align with those of the HYPO-RT-PC trial.The paper discusses similarities and differences with HYPO-RT-PC, and the non-inferiority finding is consistent.Evidence: Discussion compares PACE-B results to HYPO-RT-PC.
“Its results align with those of the HYPO-RT-PC phase 3 non-inferiority trial”
Discussion ¶2Find in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is biochemical/clinical failure, which is a surrogate for clinically meaningful outcomes like metastasis or survival. The paper does not provide evidence linking biochemical failure to patient-important outcomes in this context, nor does it demonstrate target engagement for the surrogate at the tested dose.
“The primary endpoint was freedom from biochemical/clinical failure with a critical hazard ratio for non-inferiority of 1.45.”
- INADEQUATEEffect sizeThe reported effect is a small absolute difference of 1.43% in biochemical failure-free rates at 5 years, with confidence intervals crossing zero. No minimal clinically important difference is provided, and the effect is not anchored to clinical meaningfulness.
“The estimated absolute difference in the proportion of participants event free in the SBRT group compared with that in the CRT group at 5 years was: 1.43% (90% CI: -0.60, 2.78).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior studies on hypofractionation and SBRT, and the rationale for testing non-inferiority is clearly linked to the potential benefits of SBRT. The limitations of prior research are implicitly addressed by the trial design, though not explicitly discussed in the introduction.
“Stereotactic body radiotherapy (SBRT) builds on these developments to allow ultra-hypofractionated radiotherapy to be delivered with precision.”
“Stereotactic body radiotherapy (SBRT) builds on these developments to allow ultra-hypofractionated radiotherapy to be delivered with precision.”
Randomization method and unit are clearly described (central computer-generated permuted blocks, stratified). Blinding is not applicable as it is an open-label trial, but this is stated. Power analysis is detailed with sample size calculation. Inclusion/exclusion criteria are pre-specified. Outlier handling is addressed through ITT and per-protocol analyses. Controls are inherent in the comparator arm. Independent replication is not applicable for a single pivotal trial.
“Randomization was performed centrally by the Institute of Cancer Research Clinical Trials and Statistics Unit (ICR-CTSU) using computer generated random permuted blocks (size 4 and 6), stratified by NCCN risk group (low vs intermediate) and randomizing center.”
“Treatment was not masked.”
“A non-inferiority margin of 6% at 5 years (critical hazard ratio (HR) 1.45; selected based on expert clinical opinion), 80% power, 5% one-sided significance and a 10% loss to follow-up allowance gave a sample size of 858 patients.”
“Randomization was performed centrally by the Institute of Cancer Research Clinical Trials and Statistics Unit (ICR-CTSU) using computer generated random permuted blocks (size 4 and 6), stratified by NCCN risk group (low vs intermediate) and randomizing center.”
“Treatment was not masked.”
“A non-inferiority margin of 6% at 5 years (critical hazard ratio (HR) 1.45; selected based on expert clinical opinion), 80% power, 5% one-sided significance and a 10% loss to follow-up allowance gave a sample size of 858 patients.”
Sex is reported (all male), age and health status are reported, and demographics are detailed in Table 1. Sex justification is not applicable as the study is in prostate cancer (male-only). Species/strain and housing conditions are not applicable for a human trial.
“Median age was 69.8 years (IQR 65.4, 74.0), median PSA ng/mL was 8.0 (IQR 5.9, 11.0)”
“Median age was 69.8 years (IQR 65.4, 74.0), median PSA ng/mL was 8.0 (IQR 5.9, 11.0)”
The trial was approved by the London Chelsea Research Ethics Committee with a protocol number, and written informed consent was obtained. Regulatory compliance with Good Clinical Practice is stated.
“PACE is an investigator-initiated trial approved by the London Chelsea Research Ethics Committee (11/LO/1915) in the UK and the relevant institutional review boards in Ireland and Canada.”
“Participants were recruited by their clinical teams and provided written, informed consent before enrolment.”
“The trial was conducted in accordance with the principles of Good Clinical Practice.”
“PACE is an investigator-initiated trial approved by the London Chelsea Research Ethics Committee (11/LO/1915) in the UK and the relevant institutional review boards in Ireland and Canada.”
“Participants were recruited by their clinical teams and provided written, informed consent before enrolment.”
“The trial was conducted in accordance with the principles of Good Clinical Practice.”
The radiotherapy regimens are described in detail (dose, fractions, delivery). Software used for analysis is identified (Stata version 17.0). No antibodies, cell lines, or other biological reagents are used, so those are not applicable.
“36.25Gy in five fractions over 1-2 weeks (daily or alternate days) was delivered to 95% of the planning target volume”
“Analyses are based on a data snapshot taken on 11th September 2023 and were conducted using Stata version 17.0.”
“36.25Gy in five fractions over 1-2 weeks (daily or alternate days) was delivered to 95% of the planning target volume”
“Analyses are based on a data snapshot taken on 11th September 2023 and were conducted using Stata version 17.0.”
All statistical tests are named, assumptions are verified (proportional hazards assessed), exact p-values are reported, effect sizes with confidence intervals are provided, software is identified, and data presentation is appropriate for a clinical trial. Mathematical plausibility is not applicable for large-N continuous outcomes.
“Kaplan Meier methods were used to estimate event rates. Estimates of treatment effect were made using unadjusted and adjusted (NCCN risk group) Cox regression models.”
“SBRT was non-inferior to CRT with an unadjusted HR 0.73 (90%CI 0.48, 1.12), p-value for non-inferiority=0.004”
“The estimated absolute difference in the proportion of participants event free in the SBRT group compared with that in the CRT group at 5 years was: 1.43% (90% CI: -0.60, 2.78).”
“Kaplan Meier methods were used to estimate event rates. Estimates of treatment effect were made using unadjusted and adjusted (NCCN risk group) Cox regression models.”
“SBRT was non-inferior to CRT with an unadjusted HR 0.73 (90%CI 0.48, 1.12), p-value for non-inferiority=0.004”
“Five-year biochemical failure event-free rates (95% CI) were 94.6% (91.9, 96.4) for CRT and 95.8% (93.3, 97.4) for SBRT.”
The paper mentions the protocol is available online and at NEJM.org, but does not provide a clear data availability statement for the trial data. No repository deposit or accession numbers are provided. Code sharing is not applicable as no custom code is mentioned.
“The protocol is available online () and at NEJM.org (https://www.nejm.org/)”
“The protocol is available online () and at NEJM.org (https://www.nejm.org/)”
The trial is registered (NCT01584258). Methods are detailed enough for replication. Limitations are discussed in the Discussion. Conclusions are proportional to the evidence. Funding and COI are disclosed.
“ClinicalTrials.gov registration: NCT01584258”
“The Sponsor (The Royal Marsden NHS Foundation Trust) received funding from Accuray Incorporated for study management, international study coordination and analysis.”
“ClinicalTrials.gov registration: NCT01584258”
“The Sponsor (The Royal Marsden NHS Foundation Trust) received funding from Accuray Incorporated for study management, international study coordination and analysis.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 14 references by DOI: 10 verified — 4 no DOI (shown, not verified).
- NO DOIProstate Cancer - StatisticsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINational Prostate Cancer Audit: State of the Nation ReportNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILong-Term Outcomes of NRG/RTOG 0126, a Randomized Trial of High Dose (79.2Gy) vs. Standard Dose (70.2Gy) Radiation Therapy (RT) for Men with Localized Prostate CancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILong-term results of dose escalation (80 vs 70 Gy) combined with long-term androgen deprivation in high-risk prostate cancers: GETUG-AFU 18 randomized trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly typo, consistency, grammar.
- MINORtypoAbstract, Results“p<0.001) and for gastrointestinal toxicity was CRT: 10.2% (95%CI 7.7, 13.5%) vs. SBRT: 10.7% (95%CI 8.1, 14.2%) (p=0.94).”→ Add a space after 'p' and before '<' for consistency: 'p < 0.001'.Minor formatting inconsistency in p-value reporting.
- MINORconsistencyMethods, Statistical analysis“The apriori defined time point of primary interest was 5 years.”→ Change 'apriori' to 'a priori' (two words).Spelling of 'a priori'.
- MINORgrammarDiscussion, paragraph 2“The PACE-B biochemical failure-free rates of 95% and 96% for CRT and SBRT, respectively, were achieved without ADT and exceeded the expectations of the trial design.”→ Consider rephrasing for clarity: 'The PACE-B biochemical failure-free rates of 95% and 96% for CRT and SBRT, respectively, were achieved without ADT and exceeded the expectations of the trial design.'The sentence is grammatically correct but could be clearer.
- MINORtypoAbstract, Results“p<0.001) and for gastrointestinal toxicity was CRT: 10.2% (95%CI 7.7, 13.5%) vs. SBRT: 10.7% (95%CI 8.1, 14.2%) (p=0.94).”→ Add a space before the parenthesis for consistency: 'p<0.001) and for gastrointestinal toxicity was CRT: 10.2% (95%CI 7.7, 13.5%) vs. SBRT: 10.7% (95%CI 8.1, 14.2%) (p=0.94).'Minor formatting inconsistency.
- MINORconsistencyResults, paragraph 3“p-value for non-inferiority=0.004”→ Use consistent formatting for p-values, e.g., 'p-value for non-inferiority = 0.004'.Spacing around equals sign.
- MINORclarityMethods, Statistical analysis“The apriori defined time point of primary interest was 5 years.”→ Change 'apriori' to 'a priori'.Common typo.
The published work is largely robust, but the non-inferiority p-value inconsistency is a substantive concern that warrants a correction or independent re-analysis. An informed reader should weigh this when interpreting the primary endpoint. The lack of a data availability statement is a reporting gap that could be addressed in a correction or data-sharing addendum.
- 1.CRITICALstatisticsResolve the statistics inconsistency that flips a significance claim: Recomputed 8 tests: 6 consistent, 2 inconsistent (1 change significance at p<.05); 1 recomputed directly from the reported test statistics, 7 via agent-written checks.Demonstrable critical failure — blocks the verdict from passing.
- 2.HIGHstatisticsRe-analyze and correct the non-inferiority p-value for the primary endpoint (HR 0.73, 90% CI 0.48–1.12) in the Results section; the reported p=0.004 is inconsistent with the recomputed p=0.145.The non-inferiority conclusion hinges on this p-value; an incorrect p-value could mislead readers about the primary outcome.
- 3.HIGHstatisticsVerify and correct the p-value for the HR 1.59 (95% CI 1.18–2.12) reported as p<0.001; the recomputed p is 0.0019.Although the difference is small, accurate p-values are essential for transparency and reproducibility.
- 4.HIGHdata codeAdd a specific data availability statement in the Methods or a dedicated section, stating where de-identified patient-level data can be accessed (e.g., via a data access committee or a repository) and under what conditions.The current statement only mentions protocol availability, which is inadequate for a data-driven clinical trial.
- 5.MEDIUMreportingMention adherence to a reporting guideline such as CONSORT in the Methods or a checklist submission statement.Reporting guidelines improve transparency and are expected for randomized trials.
- 6.MEDIUMdata codeConsider depositing the statistical analysis code (e.g., Stata do-files) in a public repository like Zenodo or GitHub with a DOI.Sharing code enhances reproducibility and allows independent verification of the statistical results.
- 7.MEDIUMcopyeditFix the typo 'apriori' to 'a priori' in the Methods, Statistical analysis section.Correct spelling is expected in a published manuscript.
- 8.MEDIUMcopyeditStandardize p-value formatting throughout the manuscript (e.g., add spaces around '<' and '=').Consistent formatting improves readability and professionalism.
- 9.LOWcopyeditClarify the sentence in the Discussion, paragraph 2 about biochemical failure-free rates for clarity.The sentence is grammatically correct but could be clearer to readers.
- 10.LOWreportingIn the Discussion, explicitly address how limitations of prior research (e.g., lack of long-term toxicity data) are mitigated by this trial's design.This would strengthen the scientific premise and address a sub-criterion that was reported but inadequate.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.