Sasanlimab plus BCG in BCG-naive, high-risk non-muscle invasive bladder cancer: the randomized phase 3 CREST trial.
Shore ND, Powles TB, Bedke J, Galsky MD, Palou Redorta J, Ku JH, Kretkowski M, Xylinas E, Alekseev B, Ye D, Guerrero-Ramos F, Briganti A, Kulkarni GS, Brinkmann J, Calella AM, Cesari R, Eccleston A, Michelon E, Vermette J, Wei C, Steinberg GD
- DOI
- 10.1038/s41591-025-03738-z
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/b0a3cd17-5f4e-46d7-a111-94dc0d395c2c is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 2 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is event-free survival (EFS), a composite of recurrence of high-grade disease, progression, persistence of CIS, or death. EFS is a surrogate for long-term clinical outcomes such as overall survival or progression to muscle-invasive disease. The paper does not provide evidence linking EFS to a validated clinical outcome in this setting, nor does it demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure relationship) for sasanlimab. The OS interim analysis showed no difference between arms, and the BICR sensitivity analysis for EFS was not statistically significant.
“The primary endpoint was investigator-assessed event-free survival (EFS) for Arm A versus Arm C; ... The trial met its primary endpoint with a statistically significant and clinically meaningful prolongation of EFS (Arm A versus Arm C); hazard ratio, 0.68…”
- 02Treatment effect not shown to be clinically meaningful
The reported effect is a hazard ratio of 0.68 for EFS, with 36-month EFS rates of 82.1% vs 74.8% (absolute difference 7.3%). While statistically significant, the clinical meaningfulness is not anchored to a minimal clinically important difference or a validated threshold for EFS in NMIBC. The OS interim analysis showed no difference (HR 1.13), and the BICR sensitivity analysis was not significant (HR 0.75, P=0.0517), raising uncertainty about the robustness of the effect.
“The risk of experiencing an EFS event was 32% lower in Arm A versus Arm C (stratified hazard ratio (HR), 0.68 (95% confidence interval (CI): 0.49–0.94); one-sided P = 0.0095) ... The probability of being event free at 36 months was 82.1% for Arm A and 74.8%…”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported phase 3 randomized trial. The paper demonstrates strong methodological rigor in design, statistical analysis, and reporting, with minor gaps in explicit discussion of prior limitations and outlier handling.
Both reviewers classified the study as interventional and agreed on all dimensions. Minor divergences on outlier handling and reporting guideline were resolved by weighing the specific evidence; the paper's overall status remains pass. Statistics verification covered only a subset of tests (5 checked, all consistent); other statistics remain unverified.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks.
- CONSISTENTreported p = .009 · recomputed p = .010Reviewer 1Primary EFS HR for Arm A vs Arm C
“stratified hazard ratio (HR), 0.68 (95% confidence interval (CI): 0.49–0.94); one-sided P = 0.0095”
Taken as given: The HR is 0.68 with 95% CI 0.49-0.94.; The p-value is one-sided, so the two-sided p is halved.; The CI is a 95% confidence interval for the hazard ratio.Method: Recomputed two-sided p from HR and CI using normal approximation, then halved for one-sided.How we recomputed it: pCI(0.68, 0.49, 0.94, 1)/2 - UNCOMPUTABLEreported p = .844 · recomputed p = .157Reviewers 1, 2EFS HR for Arm B vs Arm C
“stratified HR, 1.16; 95% CI: 0.87–1.55; one-sided P = 0.8439”
Taken as given: The HR is 1.16 with 95% CI 0.87-1.55.; The p-value is one-sided, so the two-sided p is halved.; The CI is a 95% confidence interval for the hazard ratio.Method: Recomputed two-sided p from HR and CI using normal approximation, then halved for one-sided.How we recomputed it: pCI(1.16, 0.87, 1.55, 1)/2 - CONSISTENTreported p = .679 · recomputed p = .636Reviewer 2OS HR for Arm A vs Arm C
“stratified HR for Arm A versus Arm C, 1.13 (95% CI: 0.68–1.87); one-sided P = 0.6791”
Taken as given: The HR is 1.13 with 95% CI 0.68-1.87.; The CI is two-sided at 95%.; The p-value is one-sided, so the two-tailed p is doubled.Method: Compute two-tailed p from HR and CI using pCI, then halve for one-sided.How we recomputed it: pCI(1.13, 0.68, 1.87, 1) - CONSISTENTreported p = .604 · recomputed p = .398Reviewers 1, 2OS HR for Arm B vs Arm C
“stratified HR for Arm B versus Arm C, 1.07 (95% CI: 0.64–1.79); one-sided P = 0.6043”
Taken as given: The HR is 1.07 with 95% CI 0.64-1.79.; The p-value is one-sided, so the two-sided p is halved.; The CI is a 95% confidence interval for the hazard ratio.Method: Recomputed two-sided p from HR and CI using normal approximation, then halved for one-sided.How we recomputed it: pCI(1.07, 0.64, 1.79, 1)/2 - UNCOMPUTABLEreported p = .052 · recomputed p = .113Reviewer 2BICR EFS HR for Arm A vs Arm C
“stratified HR = 0.75; 95% CI: 0.52–1.06; one-sided P = 0.0517”
Taken as given: The HR is 0.75 with 95% CI 0.52-1.06.; The CI is two-sided at 95%.; The p-value is one-sided, so the two-tailed p is doubled.Method: Compute two-tailed p from HR and CI using pCI, then halve for one-sided.How we recomputed it: pCI(0.75, 0.52, 1.06, 1)
- lowinternal contradictionThe abstract states 'one-sided P = 0.0095' for the primary endpoint, but the sensitivity analysis by BICR was not statistically significant (P = 0.0517). This is not a contradiction but a potential concern about the robustness of the primary endpoint.
The trial met its primary endpoint with a statistically significant and clinically meaningful prolongation of EFS (Arm A versus Arm C); hazard ratio, 0.68 (95% confidence interval: 0.49–0.94); one-sided P = 0.0095. ... Although the sensitivity analysis of EFS by BICR assessment was not statistically significant at the one-sided 0.025 significance level (stratified HR = 0.75; 95% CI: 0.52–1.06; one-sided P = 0.0517)
Abstractreviewer’s wording - lowinternal contradictionThe abstract states 1,055 patients were randomized, but the safety analysis set includes 1,047 patients who received at least one dose. This is expected due to early dropouts and is not a contradiction.
1,055 patients with BCG-naive high-risk NMIBC ... were randomized ... Of 1,055 randomized patients, 1,047 received at least one dose of trial treatment
Abstractreviewer’s wording
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
7 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Sasanlimab plus BCG-I+M significantly improves event-free survival versus BCG-I+M alone in patients with BCG-naive high-risk NMIBC.The primary endpoint was met with a statistically significant HR of 0.68 (95% CI 0.49-0.94, one-sided P=0.0095), supporting the claim.Evidence: Primary endpoint analysis: HR 0.68, 95% CI 0.49-0.94, one-sided P=0.0095.
“The trial met its primary endpoint with a statistically significant and clinically meaningful prolongation of EFS (Arm A versus Arm C); hazard ratio, 0.68 (95% confidence interval: 0.49–0.94); one-sided P = 0.0095.”
AbstractFind in source - supportedReviewer 1EFS benefit was observed across prespecified subgroups, including CIS and T1.Subgroup analyses show HRs favoring Arm A in CIS and T1 subgroups, though confidence intervals are wide.Evidence: Subgroup analyses: CIS HR 0.53 (95% CI 0.29-0.98); T1 HR 0.63 (95% CI 0.41-0.96).
For patients with CIS at randomization ... the unstratified HR was 0.53 (95% CI: 0.29–0.98) ... In the subgroup of patients with T1 tumor ... unstratified HR was 0.63 (95% CI: 0.41–0.96)
Resultsreviewer’s wording - supportedReviewers 1, 2Sasanlimab is the first anti-PD-1 antibody to show a clinically meaningful prolongation of EFS when combined with BCG-I+M versus SOC in patients with BCG-naive high-risk NMIBC.The claim is based on the trial's positive result and the absence of prior similar approvals, as stated in the paper.Evidence: The trial's positive primary endpoint and the statement that no improvements have been observed in decades.
“To our knowledge, sasanlimab is the first anti-PD-1 antibody to show a clinically meaningful prolongation of EFS when combined with BCG-I+M versus SOC in patients with BCG-naive high-risk NMIBC.”
AbstractFind in source - supportedReviewer 1The safety profile of the combination is consistent with the known profiles of each agent.Safety data show expected TRAEs and irAEs, with no new safety signals.Evidence: Safety analysis: TRAEs and irAEs consistent with known profiles; no treatment-related deaths in Arms A and C.
“The observed safety profile was consistent with the known safety profile for each individual agent.”
ResultsFind in source - supportedReviewers 1, 2Sasanlimab combined with BCG-I did not result in prolongation of EFS versus BCG-I+M.The secondary endpoint for Arm B vs Arm C was not significant, supporting the claim.Evidence: EFS for Arm B vs Arm C: HR 1.16, 95% CI 0.87-1.55, one-sided P=0.8439.
EFS was not significantly different for Arm B versus Arm C (stratified HR, 1.16; 95% CI: 0.87–1.55; one-sided P = 0.8439)
Resultsreviewer’s wording - supportedReviewer 2The EFS benefit was observed across prespecified subgroups, including CIS and T1.The paper reports subgroup analyses showing HRs below 1 for CIS and T1 subgroups, supporting the claim.Evidence: Subgroup analyses: CIS HR 0.53 (95% CI 0.29-0.98), T1 HR 0.63 (95% CI 0.41-0.96).
“EFS benefit for Arm A versus Arm C was observed across prespecified subgroups, including carcinoma in situ (CIS) and T1.”
ResultsFind in source - supportedReviewer 2The safety profile of the combination was consistent with the known profiles.The safety data show expected TRAEs and irAEs, consistent with known profiles of sasanlimab and BCG.Evidence: Safety results: TRAEs, irAEs, and no new safety signals.
“The safety profile of the combination was consistent with the known profiles.”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is event-free survival (EFS), a composite of recurrence of high-grade disease, progression, persistence of CIS, or death. EFS is a surrogate for long-term clinical outcomes such as overall survival or progression to muscle-invasive disease. The paper does not provide evidence linking EFS to a validated clinical outcome in this setting, nor does it demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure relationship) for sasanlimab. The OS interim analysis showed no difference between arms, and the BICR sensitivity analysis for EFS was not statistically significant.
“The primary endpoint was investigator-assessed event-free survival (EFS) for Arm A versus Arm C; ... The trial met its primary endpoint with a statistically significant and clinically meaningful prolongation of EFS (Arm A versus Arm C); hazard ratio, 0.68 (95% confidence interval: 0.49–0.94); one-sided P = 0.0095.”
- INADEQUATEEffect sizeThe reported effect is a hazard ratio of 0.68 for EFS, with 36-month EFS rates of 82.1% vs 74.8% (absolute difference 7.3%). While statistically significant, the clinical meaningfulness is not anchored to a minimal clinically important difference or a validated threshold for EFS in NMIBC. The OS interim analysis showed no difference (HR 1.13), and the BICR sensitivity analysis was not significant (HR 0.75, P=0.0517), raising uncertainty about the robustness of the effect.
“The risk of experiencing an EFS event was 32% lower in Arm A versus Arm C (stratified hazard ratio (HR), 0.68 (95% confidence interval (CI): 0.49–0.94); one-sided P = 0.0095) ... The probability of being event free at 36 months was 82.1% for Arm A and 74.8% for Arm C.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on BCG as standard of care, the unmet need, and the rationale for combining PD-1 inhibition with BCG based on preclinical and clinical evidence. The hypothesis follows logically from the cited evidence. Limitations of prior research are implicitly addressed by the design of a phase 3 trial, though not explicitly discussed.
“Clinical trials have examined the efficacy and safety of PD-(L)1 inhibitors in BCG-naive and post-BCG NMIBC settings”
“Exposure to BCG is associated with increased PD-L1 expression in preclinical models and tumors from patients with high-risk NMIBC”
“Clinical trials have examined the efficacy and safety of PD-(L)1 inhibitors in BCG-naive and post-BCG NMIBC settings”
“Exposure to BCG is associated with increased PD-L1 expression in preclinical models and tumors from patients with high-risk NMIBC”
Randomization method is described (1:1:1) with stratification by CIS and geographic region. The unit of randomization is the patient. Blinding is not applicable as the trial is open-label, but the paper provides a rationale and mitigation via BICR. Power analysis is reported with assumptions and sample size. Inclusion/exclusion criteria are described. Outlier handling is not explicitly addressed, but the analysis population (ITT) and safety set are defined. Controls are inherent in the comparator arm. Independent replication is not applicable for a single pivotal trial.
“Patients were randomized 1:1:1 to receive sasanlimab in combination with BCG induction and maintenance (BCG-I+M; Arm A), sasanlimab with BCG induction only (BCG-I; Arm B) or BCG-I+M (Arm C). Randomization was stratified by the presence of CIS (yes or no) and geographic region”
“Under the assumptions of an HR of 0.69 and a median EFS of 24 months in Arm C, 389 EFS events would be required for each comparison to provide 90% power”
“One limitation of the trial was the open-label design, which was mitigated by the retrospective BICR of tumor biopsy and imaging”
“Patients were randomized 1:1:1 to receive sasanlimab in combination with BCG induction and maintenance (BCG-I+M; Arm A), sasanlimab with BCG induction only (BCG-I; Arm B) or BCG-I+M (Arm C). Randomization was stratified by the presence of CIS (yes or no) and geographic region”
“Under the assumptions of an HR of 0.69 and a median EFS of 24 months in Arm C, 389 EFS events would be required for each comparison to provide 90% power”
“One limitation of the trial was the open-label design, which was mitigated by the retrospective BICR of tumor biopsy and imaging”
Sex is reported (81.8% male). Age is reported (median 67 years). Demographics include race and ethnicity. Disease characteristics such as T stage and CIS presence are reported. Species/strain and housing conditions are not applicable for a human trial.
“The median age was 67 years (range, 31–91); 81.8% of patients were men”
“61.2% of patients were White, 35.5% were Asian and 0.9% were Black or African American”
“54.2% had T1 tumor as the highest grade; and 25.5% had CIS with or without papillary tumors”
“The median age was 67 years (range, 31–91); 81.8% of patients were men”
“61.2% of patients were White, 35.5% were Asian and 0.9% were Black or African American”
The paper states approval by institutional review boards/ethics committees at each site, compliance with the Declaration of Helsinki and GCP, and written informed consent from patients. The protocol was also approved by HGRAC in China. Regulatory compliance is explicitly stated.
“it was approved by the institutional review board or ethics committee at each site”
“Patients provided written informed consent before trial entry.”
“conducted in accordance with the principles of the Declaration of Helsinki, Good Clinical Practice guidelines and applicable regulatory requirements”
Sasanlimab is identified as a humanized anti-PD-1 antibody, and BCG is named. The PD-L1 assay (VENTANA SP263) is specified. Software used for analysis (SAS 9.4) is identified. No cell lines, antibodies, or organisms are used, so those criteria are not applicable.
“Sasanlimab is a humanized, monoclonal antibody specific for human PD-1”
“PD-L1 status was assessed by the VENTANA PD-L1 immunohistochemistry SP263 assay”
“Analyses were performed using SAS version 9.4 software.”
“Subcutaneous sasanlimab (300 mg) was administered in a 2-ml prefilled syringe on day 1 of each 4-week cycle, for up to 25 cycles. Intravesical BCG induction occurred as one dose weekly via instillation for six consecutive weeks”
“PD-L1 status was assessed by the VENTANA PD-L1 immunohistochemistry SP263 assay”
“Analyses were performed using SAS version 9.4 software.”
Statistical tests are named (stratified log-rank test, Cox proportional hazards model, Mantel-Haenszel test). Assumptions are handled by design (stratified Cox model). Exact p-values are reported (e.g., P = 0.0095). Effect sizes with 95% CIs are reported. Software is identified. Data presentation includes Kaplan-Meier curves and forest plots. Mathematical plausibility checks were not performed due to large N and continuous outcomes.
“Comparisons between Arms A and B versus Arm C were conducted using the stratified log-rank test.”
“one-sided P = 0.0095”
“stratified hazard ratio (HR), 0.68 (95% confidence interval (CI): 0.49–0.94)”
“Comparisons between Arms A and B versus Arm C were conducted using the stratified log-rank test.”
“one-sided P = 0.0095”
“stratified hazard ratio (HR), 0.68 (95% confidence interval (CI): 0.49–0.94)”
The data availability statement provides a concrete mechanism for requesting deidentified participant data via a secure portal, with conditions and timeframe. Repository deposit and accession numbers are not applicable for patient-level data. Code sharing is not applicable as no bespoke code is mentioned.
“Upon reasonable request and subject to review, Pfizer will provide the data that support the findings of this trial.”
“The deidentified participant data will be made available to researchers whose proposals meet the research criteria and other conditions and, for which an exception does not apply, via a secure portal.”
“Data may be requested from Pfizer trials 24 months after trial completion. The deidentified participant data will be made available to researchers whose proposals meet the research criteria and other conditions and, for which an exception does not apply, via a secure portal.”
The trial is registered (NCT04165317). Methods are detailed enough for replication. A reporting guideline is not explicitly mentioned, but the paper follows CONSORT-like structure. All pre-specified outcomes are reported, including negative results (Arm B vs C). Limitations are discussed. Conclusions are proportional to evidence. Funding and COI are disclosed.
“ClinicalTrials.gov identifier: NCT04165317”
“One limitation of the trial was the open-label design”
“The CREST trial was funded by Pfizer.”
“ClinicalTrials.gov identifier: NCT04165317”
“One limitation of the trial was the open-label design”
“The CREST trial was funded by Pfizer.”
Registered (2 IDs: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 25 references by DOI: 2 verified — 23 no DOI (shown, not verified).
- NO DOIGlobal cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countriesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDiagnosis and treatment of non-muscle invasive bladder cancer: AUA/SUO Guideline: 2024 AmendmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEuropean Association of Urology Guidelines on non-muscle-invasive bladder cancer (TaT1 and carcinoma in situ)—a summary of the 2024 Guidelines UpdateNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINew perspectives in the medical treatment of non-muscle-invasive bladder cancer: immune checkpoint inhibitors and beyondNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPembrolizumab monotherapy for the treatment of high-risk non-muscle-invasive bladder cancer unresponsive to BCG (KEYNOTE-057): an open-label, single-arm, multicentre, phase 2 studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAbstract 2667: In vitro properties and pre-clinical activity of PF-06801591, a high-affinity engineered anti-human PD-1No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAssessment of subcutaneous vs intravenous administration of anti–PD-1 antibody PF-06801591 in patients with advanced solid tumorsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOI1055P Updated results of subcutaneous (SC) anti-programmed cell death 1 (PD-1) receptor antibody PF-06801591 for locally advanced or metastatic non-small cell lung cancer (NSCLC) or urothelial carcinoma (UC)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEnhanced expression of PD-L1 in non-muscle-invasive bladder cancer after treatment with Bacillus Calmette-GuerinNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPD-L1 expression in high-risk non-muscle-invasive bladder cancer is influenced by intravesical Bacillus Calmette–Guérin (BCG) therapyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBacillus Calmette–Guérin and anti-PD-L1 combination therapy boosts immune response against bladder cancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITumor immunotherapy resistance: revealing the mechanism of PD-1 / PD-L1-mediated tumor immune escapeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInterpreting the significance of changes in health-related quality-of-life scoresNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe interpretation of scores from the EORTC quality of life questionnaire QLQ-C30No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIComparative outcomes of primary versus recurrent high-risk non–muscle-invasive and primary versus secondary muscle-invasive bladder cancer after radical cystectomy: results from a retrospective multicenter studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPopulation-based outcome of muscle-invasive bladder cancer following radical cystectomy: who can benefit from adjuvant chemotherapy?No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFailure to achieve a complete response to induction BCG therapy is associated with increased risk of disease worsening and death in patients with high risk non-muscle invasive bladder cancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExpert consensus document: consensus statement on best practice management regarding the use of intravesical immunotherapy with BCG for bladder cancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntegrating the PD-L1 prognostic biomarker in non-muscle invasive bladder cancer in clinical practice—a comprehensive review on state-of-the-art advances and critical issuesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIManagement of immune-related adverse events in patients treated with immune checkpoint inhibitor therapy: ASCO Guideline UpdateNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHealth-related quality-of-life assessment of patients with solid tumors on immuno-oncology therapiesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOI244MO Primary results from IMscin002: a study to evaluate patient (pt)- and healthcare professional (HCP)-reported preferences for atezolizumab (atezo) subcutaneous (SC) vs intravenous (IV) for the treatment of NSCLCNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDefinitions, end points, and clinical trial designs for bladder cancer: recommendations from the Society for Immunotherapy of Cancer and the International Bladder Cancer GroupNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://clinicaltrials.gov/study/NCT04165317LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.pfizer.com/science/clinical-trials/trial-data-and-resultsLIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORtypoAbstract“Bacillus Calmette–Guérin (BCG) induction and maintenance (I+M) after transurethral resection of bladder tumor is standard of care (SOC) in high-risk non-muscle invasive bladder cancer (NMIBC).”→ Consider adding 'the' before 'standard of care'.Minor grammatical improvement.
- MINORconsistencyResults, Patients“The median age was 67 years (range, 31–91); 81.8% of patients were men; and 61.2% of patients were White, 35.5% were Asian and 0.9% were Black or African American.”→ Ensure consistent use of semicolons and commas in lists.Minor punctuation inconsistency.
- MINORconsistencyAbstract“one-sided P = 0.0095”→ Consider using 'P = 0.0095 (one-sided)' for consistency with other p-value reporting.Minor style inconsistency.
- MINORclarityResults, Safety“Treatment-related adverse events (TRAEs) of any grade occurred in 87.1% of patients in Arm A, 79.0% in Arm B and 70.2% in Arm C”→ Consider adding 'respectively' for clarity.Minor clarity issue.
The published work is robust and well-reported. An informed reader should weigh the open-label design (mitigated by BICR) and the non-significant BICR sensitivity analysis (P=0.0517) when interpreting the primary endpoint. No erratum is warranted based on this audit.
- 1.MEDIUMreportingIn the Introduction, add a sentence explicitly discussing limitations of prior studies (e.g., small sample sizes, lack of randomized data) to strengthen the scientific premise.Both reviewers flagged that limitations of prior research are not explicitly addressed, which is a minor reporting gap.
- 2.MEDIUMreportingIn the Methods, clarify how outliers and missing data (especially for PROs and secondary endpoints) were handled beyond defining ITT and safety populations.Reviewer 1 rated outlier handling as inadequate; explicit description would improve methodological transparency.
- 3.MEDIUMreportingIn the Methods, add a statement about the absence of blinding as a limitation, even though it is discussed in the Discussion.Reviewer 2 suggested this to improve completeness of the Methods section.
- 4.MEDIUMreportingIn the Methods, provide more detail on the randomization implementation (e.g., central randomization system).Reviewer 2 noted that the randomization method is described but implementation details are sparse.
- 5.MEDIUMreportingIn the Results, report the number of patients with PD-L1 unknown status in the baseline table, not just percentages.Reviewer 2 suggested this to improve clarity of the baseline characteristics.
- 6.MEDIUMreportingIn the Discussion, add a statement about the generalizability of the results to other populations (e.g., non-White, female).Reviewer 2 suggested this to address potential limitations in external validity.
- 7.MEDIUMdata codeProvide the full protocol and statistical analysis plan as supplementary material, with a direct link.Reviewer 2 noted that the protocol is mentioned but not linked; making it accessible would enhance transparency.
- 8.LOWcopyeditIn the Abstract, add 'the' before 'standard of care' for grammatical correctness.Copyedit flagged a minor grammatical issue.
- 9.LOWcopyeditIn the Results, ensure consistent use of semicolons and commas in lists (e.g., demographic data).Copyedit flagged a punctuation inconsistency.
- 10.LOWcopyeditIn the Abstract, change 'one-sided P = 0.0095' to 'P = 0.0095 (one-sided)' for consistency with other p-value reporting.Copyedit flagged a style inconsistency.
- 11.LOWcopyeditIn the Results, add 'respectively' after the list of TRAE percentages for clarity.Copyedit flagged a clarity issue.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.