Quemliclustat and chemotherapy with or without zimberelimab in metastatic pancreatic adenocarcinoma: a randomized phase 1 trial.
Wainberg ZA, Manji GA, Bahary N, Ulahannan SV, Pant S, Spigel DR, Uboha NV, Oberstein PE, Saeed A, Beagle B, Kim JY, Wang N, Weeder B, Shitole S, Mrouj K, Scott JR, Ensign LG, DiRenzo DM, Walters MJ, Wu W, Kaplan A, Cho S, Kabbarah O, O'Reilly EM
- DOI
- 10.1038/s41591-026-04283-z
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/205d7963-eb0b-4385-9636-bd5cda6075f4 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingStudy design partially met−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on clinical outcomes (ORR, PFS, OS) from a phase 1b trial, but the trial lacks a concurrent randomized control group. The OS benefit is supported by a post hoc synthetic control arm comparison, which is a surrogate for a randomized controlled trial. Additionally, the biomarker analyses (NR4A expression) are mechanistic surrogates, and the paper does not provide validated evidence linking NR4A expression to clinical outcomes beyond this single study.
“In a phase 1b trial, patients with treatment-naive metastatic pancreatic adenocarcinoma received the CD73 inhibitor quemliclustat plus gemcitabine and nab-paclitaxel with or without the anti-PD1 antibody zimberelimab, showing encouraging clinical response…”
- 02Treatment effect not shown to be clinically meaningful
The reported median OS of 15.7 months in the Quemli100 cohort is compared to historical benchmarks (9.2 and 8.7 months), but the comparison is not from a randomized controlled trial. The effect size is presented as encouraging but lacks a formal statistical anchor to a minimal clinically important difference. The post hoc SCA analysis shows a median OS improvement of 5.9 months, but this is from a synthetic control arm, not a prospective randomized comparison.
“Compared to the SCA, patients treated with the quemliclustat combinations demonstrated an increase in median OS of 5.9 months (hazard ratio = 0.634 (95% CI: 0.471–0.854); P = 0.003)”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This phase 1b trial is methodologically sound for an early-phase oncology study, with clear scientific rationale, detailed reporting of demographics, ethics approvals, and a concrete data availability statement. The main weaknesses are the open-label design, lack of formal power analysis, and minor reporting gaps in statistical assumptions and exact p-values. Overall, the paper is robust and transparent, with only minor issues that do not undermine its conclusions.
Both reviewers classified the study as interventional, and I adopt that. The evaluation covered all eight dimensions; several sub-criteria were marked not applicable (e.g., species/strain, housing, IACUC, controls, independent replication) due to the human clinical trial context. The statistics verification covered only a subset of tests (8 reported with test statistics/CI), so the absence of errors does not confirm overall statistical correctness. The reviewers diverged on study design (warn vs. pass); I weighed the open-label design and lack of formal power analysis as genuine limitations, leading to a warn.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 8 tests: 8 consistent, 0 inconsistent; 6 recomputed directly from the reported test statistics, 2 via agent-written checks.
- CONSISTENTreported p = .024 · recomputed p = .023Recomputed hazard ratio 0.732 (95% CI 0.56–0.96), reported p=0.0238
“hazard ratio = 0.732 (95% CI: 0.56–0.96); P = 0.0238”
Taken as given: 0.56–0.96 is a two-sided 95% confidence interval for the hazard ratio of 0.732, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.0238 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.732, 0.56, 0.96, 1) - CONSISTENTreported p = .024 · recomputed p = .021Recomputed hazard ratio 0.678 (95% CI 0.49–0.95), reported p=0.0239
“hazard ratio = 0.678 (95% CI: 0.49–0.95); P = 0.0239”
Taken as given: 0.49–0.95 is a two-sided 95% confidence interval for the hazard ratio of 0.678, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.0239 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.678, 0.49, 0.95, 1) - CONSISTENTreported p = .003 · recomputed p = .004Recomputed hazard ratio 0.42 (95% CI 0.23–0.76), reported p=0.0034
“hazard ratio = 0.42 (95% CI: 0.23–0.76); P = 0.0034”
Taken as given: 0.23–0.76 is a two-sided 95% confidence interval for the hazard ratio of 0.42, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.0034 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.42, 0.23, 0.76, 1) - CONSISTENTreported p = .015 · recomputed p = .017Recomputed hazard ratio 0.41 (95% CI 0.20–0.86), reported p=0.015
“hazard ratio = 0.41 (95% CI: 0.20–0.86); P = 0.015”
Taken as given: 0.20–0.86 is a two-sided 95% confidence interval for the hazard ratio of 0.41, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.015 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.41, 0.2, 0.86, 1) - CONSISTENTreported p = .073 · recomputed p = .079Recomputed hazard ratio 0.49 (95% CI 0.22–1.08), reported p=0.073
“hazard ratio = 0.49 (95% CI: 0.22–1.08); P = 0.073”
Taken as given: 0.22–1.08 is a two-sided 95% confidence interval for the hazard ratio of 0.49, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.073 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.49, 0.22, 1.08, 1) - CONSISTENTreported p = .004 · recomputed p = .008Recomputed hazard ratio 0.24 (95% CI 0.08–0.67), reported p=0.0035
“hazard ratio = 0.24 (95% CI: 0.08–0.67); P = 0.0035”
Taken as given: 0.08–0.67 is a two-sided 95% confidence interval for the hazard ratio of 0.24, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.0035 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.24, 0.08, 0.67, 1) - CONSISTENTreported p = .003 · recomputed p = .003Reviewer 1Check p-value for OS comparison between Quemli100 and SCA using reported HR and CI.
“Compared to the SCA, patients treated with the quemliclustat combinations demonstrated an increase in median OS of 5.9 months (hazard ratio = 0.634 (95% CI: 0.471–0.854); P = 0.003)”
Taken as given: The hazard ratio is 0.634 and the 95% CI is 0.471-0.854.; The p-value is two-sided and derived from the CI.Method: Recomputed p-value from HR and 95% CI using the pCI function assuming a log-normal distribution.How we recomputed it: pCI(0.634, 0.471, 0.854, 1) - CONSISTENTreported p = .110 · recomputed p = .061Reviewer 1Check p-value for PFS comparison between Quemli100 and SCA using reported HR and CI.
“Median PFS was not significantly different between the two arms (Quemli100, 6.3 months (95% CI: 5.4–7.7); SCA, 5.5 months (95% CI: 4.4–6.6); P = 0.110)”
Taken as given: The HR for PFS is approximately 0.78 (derived from 22% reduction in risk).; The 95% CI is approximately 0.60-1.01.; The p-value is two-sided.Method: Recomputed p-value from HR and CI using pCI function.How we recomputed it: pCI(0.78, 0.6, 1.01, 1)
- lowinternal contradictionIn Table 1, the pooled Q+G/nP+Z column has n=93, but the sum of the two arms (29+61) is 90, and the pooled column includes patients from the non-randomized cohort.
“Pooled Q + G/nP + Z ( n = 93)”
Table 1Find in source - lowinternal contradictionThe abstract states 'randomized phase 1 trial' but the study includes a non-randomized cohort in the expansion phase.
“Quemliclustat and chemotherapy with or without zimberelimab in metastatic pancreatic adenocarcinoma: a randomized phase 1 trial”
AbstractFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
4 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2The NR4A expression signature may represent a reasonable surrogate for assessing adenosine levels in the TME.The paper provides in vitro and in vivo evidence linking NR4A to adenosine, but the surrogate validity is not fully established.Evidence: In vitro experiments show NR4A upregulation by adenosine and downregulation by quemliclustat; in vivo, NR4A expression decreases post-treatment.
“Thus, an NR4A expression signature may represent a reasonable surrogate for assessing adenosine levels in the TME.”
DiscussionFind in source - supportedReviewers 1, 2Quemliclustat combined with G/nP with or without zimberelimab shows encouraging clinical response rates and survival in patients with mPDAC.The reported ORR and OS are consistent with the evidence presented, though the lack of a concurrent control limits definitive conclusions.Evidence: Reported ORR of 38% and 25% in the two arms, and median OS of 19.4 and 14.6 months.
“Clinical response rates and survival outcomes were encouraging.”
AbstractFind in source - supportedReviewers 1, 2High tumor NR4A expression is associated with improved OS in ARC-8 but not in external cohorts.The paper provides HRs and p-values for the ARC-8 BEP and shows no association in PRINCE and MORPHEUS cohorts.Evidence: HR=0.678 (95% CI: 0.49-0.95), P=0.0239 for OS in ARC-8; no significant association in external cohorts.
“The NR4A family expression significantly correlated with survival benefit in the BEP (PFS: hazard ratio = 0.732 (95% CI: 0.56–0.96); P = 0.0238; OS: hazard ratio = 0.678 (95% CI: 0.49–0.95); P = 0.0239) but not in the PRINCE G/nP + nivo or MORPHEUS G/nP clinical cohorts”
ResultsFind in source - supportedReviewers 1, 2Maximal downregulation of NR4A expression after treatment is associated with T cell activation and improved OS.The paper shows significant upregulation of T cell activation signatures in the maximal decrease subgroup and a significant OS benefit (HR=0.24).Evidence: OS HR=0.24 (95% CI: 0.08-0.67), P=0.0035; T cell activation signatures upregulated in maximal decrease subgroup.
“Maximal decrease in NR4A family expression after treatment was associated with a positive trend toward improved PFS (hazard ratio = 0.49 (95% CI: 0.22–1.08); P = 0.073) and was significantly associated with improved OS (hazard ratio = 0.24 (95% CI: 0.08–0.67); P = 0.0035)”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on clinical outcomes (ORR, PFS, OS) from a phase 1b trial, but the trial lacks a concurrent randomized control group. The OS benefit is supported by a post hoc synthetic control arm comparison, which is a surrogate for a randomized controlled trial. Additionally, the biomarker analyses (NR4A expression) are mechanistic surrogates, and the paper does not provide validated evidence linking NR4A expression to clinical outcomes beyond this single study.
“In a phase 1b trial, patients with treatment-naive metastatic pancreatic adenocarcinoma received the CD73 inhibitor quemliclustat plus gemcitabine and nab-paclitaxel with or without the anti-PD1 antibody zimberelimab, showing encouraging clinical response rates and survival in quemliclustat-treated patients.”
- INADEQUATEEffect sizeThe reported median OS of 15.7 months in the Quemli100 cohort is compared to historical benchmarks (9.2 and 8.7 months), but the comparison is not from a randomized controlled trial. The effect size is presented as encouraging but lacks a formal statistical anchor to a minimal clinically important difference. The post hoc SCA analysis shows a median OS improvement of 5.9 months, but this is from a synthetic control arm, not a prospective randomized comparison.
“Compared to the SCA, patients treated with the quemliclustat combinations demonstrated an increase in median OS of 5.9 months (hazard ratio = 0.634 (95% CI: 0.471–0.854); P = 0.003)”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Study-design details incomplete (controls, blinding, power)Assessed
The introduction cites multiple prior studies on CD73, adenosine signaling, and the failure of existing therapies, establishing the scientific foundation. The rationale for targeting CD73 with quemliclustat is clearly articulated, and the study objectives follow logically. Limitations of prior transcriptional signatures are acknowledged, and the paper addresses them by proposing NR4A family expression as a more robust biomarker.
“However, these transcriptional signatures have limited predictive value, in part because they may not adequately reflect the cellular heterogeneity in the TME.”
“In the ARC-8 phase 1b study, we evaluated the safety and tolerability of quemliclustat combined with standard-of-care G/nP with or without the anti-PD-1 antibody zimberelimab, in patients with treatment-naive mPDAC.”
“However, these transcriptional signatures have limited predictive value, in part because they may not adequately reflect the cellular heterogeneity in the TME.”
The dose-escalation phase uses a standard 3+3 design, and the expansion phase includes a randomized 2:1 cohort with permuted block randomization. However, the study is open-label, and no blinding is described. The sample size justification is based on an estimation framework rather than formal power analysis, which is acceptable for a phase 1b trial but limits the ability to detect small effects. Inclusion/exclusion criteria are described, and the safety-evaluable population is defined. Outlier handling is not explicitly addressed, but the analysis population is defined.
“patients were enrolled and randomized 2:1 using the permuted block method”
“The sample size justification was based on an estimation framework, and the study was designed for descriptive statistical analysis rather than formal statistical hypothesis testing involving power and type I error considerations.”
“patients were enrolled and randomized 2:1 using the permuted block method”
“The sample size justification was based on an estimation framework, and the study was designed for descriptive statistical analysis rather than formal statistical hypothesis testing involving power and type I error considerations.”
The paper reports age, sex, race, and ECOG PS for all cohorts in Table 1. Both sexes are enrolled, so a scientific justification for single-sex is not required. Health status is implied by inclusion criteria (ECOG PS 0-1). Species/strain and housing conditions are not applicable as this is a human trial.
“Male | 2 (50) | 4 (67) | 2 (67) | 3 (50) | 1 (33) | 12 (55) | | Female | 2 (50) | 2 (33) | 1 (33) | 3 (50) | 2 (67) | 10 (45)”
The Methods state that the study was conducted in conformance with the Declaration of Helsinki and other guidelines, and that the protocol was approved by the local ethics committee at each site. All patients provided written informed consent. This satisfies the requirements for human research.
“The study protocol was approved by the local ethics committee at each site (Supplementary Table ).”
“All patients provided written informed consent before any study procedures”
“The study was conducted in full conformance with the Declaration of Helsinki, the Council for International Organizations of Medical Sciences International Ethical Guidelines, institutional review board regulations and all other applicable local regulations.”
“The study protocol was approved by the local ethics committee at each site (Supplementary Table ).”
“All patients provided written informed consent before any study procedures”
“The study was conducted in full conformance with the Declaration of Helsinki, the Council for International Organizations of Medical Sciences International Ethical Guidelines, institutional review board regulations and all other applicable local regulations.”
The drugs are named with doses and schedules. Cell lines are identified as purchased from ATCC, and specific reagents (e.g., AMP, EHNA, NECA) are named with suppliers. Software tools are mentioned (e.g., SAS v.9.4) but not all versions are given. Since this is a clinical trial, antibodies and cell line authentication are not applicable.
“Patients received quemliclustat intravenously (25 mg, 50 mg, 75 mg, 100 mg or 125 mg) every 2 weeks, G/nP (gemcitabine 1,000 mg m − 2 and nab-paclitaxel 125 mg m − 2 ) intravenously on days 1, 8 and 15 of a 28-day cycle and zimberelimab 240 mg intravenously every 2 weeks.”
“Cell lines were purchased from the American Type Culture Collection and cultured based on the supplier’s recommendations.”
“The data were extracted and standardized to ADaM datasets in SAS v.9.4.”
“Patients received quemliclustat intravenously (25 mg, 50 mg, 75 mg, 100 mg or 125 mg) every 2 weeks, G/nP (gemcitabine 1,000 mg m − 2 and nab-paclitaxel 125 mg m − 2 ) intravenously on days 1, 8 and 15 of a 28-day cycle and zimberelimab 240 mg intravenously every 2 weeks.”
“Cell lines were purchased from the American Type Culture Collection and cultured based on the supplier’s recommendations.”
“The data were extracted and standardized to ADaM datasets in SAS v.9.4.”
The paper names tests such as log-rank test and Cox proportional hazards model. Effect sizes are reported with 95% CIs. P-values are often reported as exact values (e.g., P = 0.003) but sometimes as thresholds (e.g., P < 0.0001). The statistical software is identified (SAS v.9.4). Data presentation includes Kaplan-Meier curves and forest plots. Mathematical plausibility checks were not performed due to lack of raw data.
“P values were calculated using a log-rank test between groups. HRs and 95% CIs were calculated using Cox proportional hazards model.”
“Median OS was significantly longer in the Quemli100 arm (15.7 months (95% CI: 12.4–20.9)) versus the SCA (9.8 months (95% CI: 7.8–11.4)) ( P = 0.003)”
“P values were calculated using a log-rank test between groups. HRs and 95% CIs were calculated using Cox proportional hazards model.”
“Median OS was 15.7 months (95% CI: 12.4–20.9)”
“The data were extracted and standardized to ADaM datasets in SAS v.9.4.”
The data availability statement describes a managed access process for deidentified participant data through a named platform (trials.arcusbio.com). This is adequate for patient-level data. No raw sequencing data are deposited, but that is not required for this type of trial. Code sharing is not applicable as no bespoke code is described.
“Arcus Biosciences will provide access to individual deidentified participant data and related study documents (protocols, statistical analysis plans and clinical study reports) upon request from qualified researchers and subject to certain criteria, conditions and exceptions. For information on the process or to submit a request, visit https://trials.arcusbio.com/our-transparency-policy”
“Arcus Biosciences will provide access to individual deidentified participant data and related study documents (protocols, statistical analysis plans and clinical study reports) upon request from qualified researchers and subject to certain criteria, conditions and exceptions. For information on the process or to submit a request, visit https://trials.arcusbio.com/our-transparency-policy”
Methods are comprehensive, including dosing, assessments, and statistical analysis. Limitations are explicitly discussed, including the lack of a concurrent control group and the post hoc nature of biomarker analyses. Conclusions are generally proportional, though some statements about NR4A as a surrogate are cautious. Funding and COI are disclosed. Trial registration is provided. No reporting guideline is mentioned, but this is not critical for a phase 1b trial.
“ClinicalTrials.gov identifier: NCT04104672”
“As a phase 1b trial, the ARC-8 study was designed to evaluate the safety and tolerability of quemliclustat combined with G/nP with or without zimberelimab. The findings from ARC-8 may not be generalizable to the broader patient population, despite the sample size being large for an early phase trial.”
“This work was supported by Arcus Biosciences. No grant number is applicable for any funding received.”
“ClinicalTrials.gov identifier: NCT04104672”
“As a phase 1b trial, the ARC-8 study was designed to evaluate the safety and tolerability of quemliclustat combined with G/nP with or without zimberelimab. The findings from ARC-8 may not be generalizable to the broader patient population, despite the sample size being large for an early phase trial.”
“This work was supported by Arcus Biosciences. No grant number is applicable for any funding received.”
Registered (2 IDs: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 65 references by DOI: 62 verified — 3 no DOI (shown, not verified).
- NO DOICommon Terminology Criteria for Adverse Events (CTCAE) version 5.0No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFastQCNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIsurvminer: drawing survival curves using ‘ggplot2ʼNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
3 data/code links checked; 3 live.
- datahttps://trials.arcusbio.com/our-transparency-policyLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- dataGEOLIVEHTTP 200https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE202051Resolves to GEO (data repository).
- codeGitHubLIVEHTTP 200https://github.com/ParkerICI/prince-trial-dataResolves to GitHub (code repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAbstract“Quemliclustat and chemotherapy with or without zimberelimab in metastatic pancreatic adenocarcinoma: a randomized phase 1 trial”→ Consider adding 'a' before 'randomized' for grammatical correctness.Minor grammatical issue.
- MINORconsistencyTable 1“Pooled Q + G/nP + Z ( n = 93)”→ Ensure the pooled column is consistently defined across tables.Pooled column definition is clear but could be repeated in footnotes.
- MINORclarityMethods, Statistical analysis“The sample size justification was based on an estimation framework, and the study was designed for descriptive statistical analysis rather than formal statistical hypothesis testing involving power and type I error considerations.”→ Consider simplifying the sentence for clarity.Long sentence but clear.
- MINORconsistencyTable 1“Pooled Q + G/nP + Z ( n = 93)”→ Ensure the pooled column is consistently labeled across tables.Minor inconsistency in labeling.
- MINORtypoExtended Data Fig. 9“NR4A1 negativ e”→ Change to 'NR4A1 negative'.Typographical error.
The published work is robust and transparent, with no critical integrity concerns. An informed reader should weigh the open-label design and the post hoc nature of biomarker analyses when interpreting efficacy claims. The minor internal contradiction in the abstract (randomized vs. non-randomized cohort) and the lack of a reporting guideline are minor reporting gaps that could warrant a correction or clarification, but they do not undermine the study's validity.
- 1.HIGHreportingIn the Abstract, clarify that the trial is 'phase 1b' and specify that only the dose-expansion phase is randomized, to resolve the internal contradiction with the non-randomized cohort.The abstract states 'randomized phase 1 trial' but the study includes a non-randomized cohort, which is a low-severity internal contradiction that could confuse readers.
- 2.HIGHstatisticsIn the Methods/Statistical analysis, add a statement on verification of statistical assumptions (e.g., proportional hazards for Cox models) and describe how outliers were handled.Both reviewers flagged the lack of explicit assumption verification and outlier handling as reporting gaps that affect reproducibility.
- 3.MEDIUMstatisticsProvide exact p-values for all comparisons currently reported as thresholds (e.g., P < 0.0001) in the text and figures.Threshold-only p-values limit readers' ability to assess the strength of evidence precisely.
- 4.MEDIUMreportingAdd a statement referencing a reporting guideline such as CONSORT in the Methods or a dedicated section.Explicit adherence to a reporting guideline enhances transparency and is a common reviewer request.
- 5.MEDIUMdata codeConsider depositing de-identified biomarker data (e.g., RNA-seq) in a public repository with accession numbers.While the data availability statement is adequate, depositing raw biomarker data would facilitate independent verification and reproducibility.
- 6.MEDIUMotherIn the Methods/Cell culture experiments, add authentication details for cell lines (e.g., STR profiling) and mycoplasma testing results.Reviewer 2 noted cell line authentication and mycoplasma testing are not reported, which is a minor gap for the biomarker experiments.
- 7.LOWcopyeditFix the typo in Extended Data Fig. 9: change 'NR4A1 negativ e' to 'NR4A1 negative'.Typographical error that should be corrected for professionalism.
- 8.LOWcopyeditIn the Abstract, add 'a' before 'randomized' for grammatical correctness.Minor grammatical issue flagged by the copyedit pass.
- 9.LOWcopyeditEnsure the pooled column in Table 1 is consistently labeled and defined across all tables.Minor consistency issue in table labeling that could confuse readers.
- 10.LOWreportingIn the Discussion, explicitly acknowledge the lack of a concurrent control group and discuss the potential for bias in the synthetic control arm comparison.While limitations are discussed, making this explicit would strengthen the transparency of the efficacy comparisons.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.