Quemliclustat and chemotherapy with or without zimberelimab in metastatic pancreatic adenocarcinoma: a randomized phase 1 trial.
Wainberg ZA, Manji GA, Bahary N, Ulahannan SV, Pant S, Spigel DR, Uboha NV, Oberstein PE, Saeed A, Beagle B, Kim JY, Wang N, Weeder B, Shitole S, Mrouj K, Scott JR, Ensign LG, DiRenzo DM, Walters MJ, Wu W, Kaplan A, Cho S, Kabbarah O, O'Reilly EM
- DOI
- 10.1038/s41591-026-04283-z
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/499c3dcf-3be9-4b64-af29-a01af34fc1e8 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ReportingKey resources partially met−0.25★
- ReportingData & code availability partially met−0.25★
- References were not verified against Crossref/OpenAlex.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run on this paper: the pass that reads its reported means did not complete. No reported mean was checked for arithmetic impossibility.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The manuscript is a well-designed and transparently reported phase 1b exploratory trial with a strong scientific rationale, sound ethics documentation, and careful, estimation-framed statistical reporting; the 10 machine-recomputed tests were all consistent. Its main weaknesses are in resource authentication (unidentified mIF/ISH antibodies and probes, no cell-line STR authentication or mycoplasma testing) and data/code availability (no sequencing accession numbers, no shared analysis code), plus minor reporting gaps (no CONSORT reference, disclosed-but-unreported secondary endpoints, and a low-severity percentage rounding inconsistency).
Three independent reviewer runs (same model) were synthesized; reviewers converged on six dimensions and diverged on study design (1 warn vs 2 pass) and data code availability (1 fail vs 2 warn), both resolved to the majority with reasoning preserved. Statistics verification recomputed only 10 of the reported tests (those with a statistic + df or effect + CI); threshold-only, q-value, and resampling/exact p-values were not machine-verified and are not endorsed. Citation check was not run (0 references checked). All 5 data/code links verified live; no retracted or not-found references were flagged.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 10 tests: 10 consistent, 0 inconsistent; 7 recomputed directly from the reported test statistics, 3 via agent-written checks.
- CONSISTENTreported p = .024 · recomputed p = .023Recomputed hazard ratio 0.732 (95% CI 0.56–0.96), reported p=0.0238
“hazard ratio = 0.732 (95% CI: 0.56–0.96); P = 0.0238”
Taken as given: 0.56–0.96 is a two-sided 95% confidence interval for the hazard ratio of 0.732, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.0238 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.732, 0.56, 0.96, 1) - CONSISTENTreported p = .024 · recomputed p = .021Recomputed hazard ratio 0.678 (95% CI 0.49–0.95), reported p=0.0239
“hazard ratio = 0.678 (95% CI: 0.49–0.95); P = 0.0239”
Taken as given: 0.49–0.95 is a two-sided 95% confidence interval for the hazard ratio of 0.678, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.0239 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.678, 0.49, 0.95, 1) - CONSISTENTreported p = .003 · recomputed p = .004Recomputed hazard ratio 0.42 (95% CI 0.23–0.76), reported p=0.0034
“hazard ratio = 0.42 (95% CI: 0.23–0.76); P = 0.0034”
Taken as given: 0.23–0.76 is a two-sided 95% confidence interval for the hazard ratio of 0.42, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.0034 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.42, 0.23, 0.76, 1) - CONSISTENTreported p = .015 · recomputed p = .017Recomputed hazard ratio 0.41 (95% CI 0.20–0.86), reported p=0.015
“hazard ratio = 0.41 (95% CI: 0.20–0.86); P = 0.015”
Taken as given: 0.20–0.86 is a two-sided 95% confidence interval for the hazard ratio of 0.41, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.015 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.41, 0.2, 0.86, 1) - CONSISTENTreported p = .073 · recomputed p = .079Recomputed hazard ratio 0.49 (95% CI 0.22–1.08), reported p=0.073
“hazard ratio = 0.49 (95% CI: 0.22–1.08); P = 0.073”
Taken as given: 0.22–1.08 is a two-sided 95% confidence interval for the hazard ratio of 0.49, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.073 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.49, 0.22, 1.08, 1) - CONSISTENTreported p = .004 · recomputed p = .008Recomputed hazard ratio 0.24 (95% CI 0.08–0.67), reported p=0.0035
“hazard ratio = 0.24 (95% CI: 0.08–0.67); P = 0.0035”
Taken as given: 0.08–0.67 is a two-sided 95% confidence interval for the hazard ratio of 0.24, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.0035 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.24, 0.08, 0.67, 1) - CONSISTENTreported p = .003 · recomputed p = .003Recomputed hazard ratio 0.634 (95% CI 0.471–0.854), reported p=0.003
“hazard ratio = 0.634 (95% CI: 0.471–0.854); P = 0.003”
Taken as given: 0.471–0.854 is a two-sided 95% confidence interval for the hazard ratio of 0.634, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.003 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.634, 0.471, 0.854, 1) - CONSISTENTreported p = .015 · recomputed p = .017Reviewer 1Compute p-value from HR and 95% CI for OS comparison between NR4A high and low groups (HR=0.41, CI=0.20-0.86).
“Similarly, OS was significantly longer for the NR4A high group (hazard ratio = 0.41 (95% CI: 0.20–0.86); P = 0.015)”
Taken as given: The hazard ratio of 0.41 and its 95% confidence interval (0.20, 0.86) are correctly reported from a Cox proportional hazards model.; The CI is two-sided at the 95% level.; The p-value corresponds to the test of the null hypothesis that the true HR equals 1.Method: pCI function from the sandbox, which computes a two-sided p-value from an estimate and its 95% confidence interval on the log scale for a ratio (log=1).How we recomputed it: pCI(0.41, 0.20, 0.86, 1) - CONSISTENTreported p < .003 · recomputed p = .003Reviewer 2SCA comparison: OS HR 0.634 (95% CI 0.471–0.854) for Quemli100 vs synthetic control arm.
“Median OS was significantly longer in the Quemli100 arm (15.7 months (95% CI: 12.4–20.9)) versus the SCA (9.8 months (95% CI: 7.8–11.4)) ( P = 0.003)”
Taken as given: The HR 0.634 and 95% CI 0.471–0.854 are the log-scale effect estimate and confidence interval from the Cox model.; The p-value is two-sided, derived from the Wald test on the log hazard ratio.; The CI is a two-sided 95% confidence interval.Method: Recomputed two-sided p from the reported HR and 95% CI on the log scale using a normal approximation of the Wald test.How we recomputed it: pCI(0.634, 0.471, 0.854, 1) - CONSISTENTreported p = .015 · recomputed p = .017Reviewer 3Recompute p for OS HR of NR4A-high vs NR4A-low (bestcut) from reported HR and 95% CI.
“Similarly, OS was significantly longer for the NR4A high group (hazard ratio = 0.41 (95% CI: 0.20–0.86); P = 0.015)”
Taken as given: the 0.41, 0.20 and 0.86 are the HR and 95% CI bounds of the same OS comparison; the CI is two-sided at 95%; the HR is a ratio, so log=1Method: Two-tailed Wald z-test p derived from the log-HR and its 95% CI.How we recomputed it: pCI(0.41, 0.20, 0.86, 1)
- lowinternal contradictionPercentages in the BEP molecular-subtype breakdown sum to more than 100% (83% classical + 18% basal-like = 101%), indicating a rounding inconsistency in the presented counts.
“We classified patients in the ARC-8 BEP into classical molecular subtype ( n = 66 (83%)) or basal-like molecular subtype ( n = 14 (18%))”
ResultsFind in source
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
12 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewer 2Clinical response rates and survival outcomes in quemliclustat-treated patients were encouraging.The ORR and OS in the Quemli100 cohort are reported (confirmed ORR 29%, median OS 15.7 months), but these are single-arm descriptive results with no concurrent control, and the randomized comparison showed a numerically lower ORR in the triple arm; the claim is cautiously worded ('encouraging').Evidence: Efficacy results for the randomized arms and pooled Quemli100 cohort (Table 3), and the SCA comparison.
“Clinical response rates and survival outcomes were encouraging.”
AbstractFind in source - partialReviewer 2High tumor NR4A expression is associated with improved OS in ARC-8 but not in external cohorts, indicating the signature is predictive of benefit from a quemliclustat-containing regimen rather than prognostic.The OS association in the ARC-8 BEP (HR 0.678, P=0.0239) and the non-association in PRINCE/MORPHEUS are reported, supporting the claim that the signature is not generally prognostic; however, the claim that it is predictive of quemliclustat-specific benefit is an inference from the negative external-cohort results and the small BEP subgroup, not a randomized test of the biomarker.Evidence: BEP survival analyses (HR 0.678 for OS, P=0.0239) and external cohort analyses (not significant).
“These findings suggest that NR4A family expression is predictive of patient benefit from a quemliclustat-containing regimen rather than prognostic of general mPDAC outcomes.”
ResultsFind in source - supportedReviewer 1Quemliclustat combined with gemcitabine/nab-paclitaxel with or without zimberelimab shows encouraging clinical response rates and survival in patients with metastatic pancreatic adenocarcinoma.The paper presents ORR, DCR, PFS, and OS data with confidence intervals, supporting the claim of encouraging clinical activity.Evidence: Table 3 reports ORR of 38% and 25% in the two arms; median OS of 19.4 and 14.6 months in the randomized arms; Kaplan-Meier curves in Figure 1.
In a phase 1b trial, patients with treatment-naive metastatic pancreatic adenocarcinoma received the CD73 inhibitor quemliclustat plus gemcitabine and nab-paclitaxel with or without the anti-PD1 antibody zimberelimab, showing encouraging clinical response rates and survival in quemliclustat-treated patients.
Abstractreviewer’s wording - supportedReviewer 1NR4A family gene expression is upregulated by adenosine in vitro and by chemotherapy in human PDACs.In vitro experiments with AMP and NECA show upregulation of NR4A family members, and snRNA-seq data from chemotherapy-treated patients show increased expression.Evidence: Figure 2a,b show RT-PCR and RNA-seq data; analysis of public snRNA-seq data (Figure 2c) shows upregulation after chemotherapy.
NR4A family expression is upregulated by adenosine in the major cellular components comprising the TME... Upregulation of NR4A family expression after addition of AMP was confirmed... across all cell types.
Resultsreviewer’s wording - supportedReviewer 1High tumor NR4A expression was associated with improved OS in ARC-8 but not in two external cohorts (PRINCE and Morpheus-PDAC).Forest plots show significant association in ARC-8 (HR 0.678, p=0.0239) and non-significant in PRINCE and Morpheus, supporting the claim.Evidence: Figure 2d,e show forest plots; reported HRs and p-values.
The NR4A family expression significantly correlated with survival benefit in the BEP (PFS: hazard ratio = 0.732 (95% CI: 0.56–0.96); P = 0.0238; OS: hazard ratio = 0.678 (95% CI: 0.49–0.95); P = 0.0239) but not in the PRINCE G/nP + nivo or MORPHEUS G/nP clinical cohorts.
Resultsreviewer’s wording - supportedReviewer 1Spatial tissue analyses revealed a scarcity of activated T cells near regions with high NR4A1 expression.ISH and mIF data show fewer IFNγ+ T cells within 50 μm of NR4A1 high regions, and proximity analysis quantifies this difference.Evidence: Extended Data Figures 9 and 10, with quantitative analysis showing significantly fewer IFNγ+ cells near NR4A1 high cells.
There were significantly fewer IFNγ transcripts per cell in NR4A1 high versus NR4A1 low ROIs... Spatial analyses showed that IFNγ+ T cells were distributed significantly further away from cells expressing high levels of NR4A1.
Resultsreviewer’s wording - supportedReviewer 1In paired biopsies, maximal downregulation of NR4A expression was associated with T cell activation and improved OS.Paired biopsy analysis shows that patients with maximal NR4A decrease have upregulation of T cell activation signatures and significantly improved OS (HR 0.24, p=0.0035).Evidence: Figure 3a-d show NR4A downregulation and T cell activation; Figure 3e-f show Kaplan-Meier curves for OS.
Maximal decrease in NR4A family expression after treatment was associated with a positive trend toward improved PFS... and was significantly associated with improved OS (hazard ratio = 0.24 (95% CI: 0.08–0.67); P = 0.0035).
Resultsreviewer’s wording - supportedReviewer 2Quemliclustat combined with G/nP with or without zimberelimab shows an encouraging safety profile consistent with G/nP alone.The safety data across all arms are presented consistently with historical G/nP, and the paper transparently reports TEAEs, grade ≥3 events, SAEs, and treatment-related deaths.Evidence: Safety tables (Table 2) and text reporting TEAE rates, grade ≥3 events, SAEs, and that most events were chemotherapy-related.
“In all treatment arms, the safety profile was consistent with that of G/nP.”
AbstractFind in source - supportedReviewer 2Treatment with a quemliclustat regimen downregulates tumor NR4A expression and increases T cell activation in the TME.Paired biopsy analyses (n=37) show significant NR4A downregulation post-treatment (P=0.0092) and upregulation of T cell activation signatures in the maximal-decrease subgroup, directly supporting the claim.Evidence: Paired pre/post biopsy RNA-seq analyses and T cell activation signature assessments.
“Compared to pretreatment levels, tumor NR4A expression was significantly downregulated posttreatment with quemliclustat ( P = 0.0092)”
ResultsFind in source - supportedReviewer 2The magnitude of NR4A downregulation after treatment is associated with improved overall survival in ARC-8.The maximal-decrease subgroup showed significantly improved OS (HR 0.24, P=0.0035), directly supporting the claim, though the subgroup definition (median split) and small sample (n=37 paired) are acknowledged limitations.Evidence: Kaplan-Meier OS analysis of maximal vs minimal NR4A decrease subgroups.
Maximal decrease in NR4A family expression after treatment was associated with ... significantly associated with improved OS (hazard ratio = 0.24 (95% CI: 0.08–0.67); P = 0.0035)
Resultsreviewer’s wording - supportedReviewer 2Spatial analyses reveal a scarcity of activated (IFNγ+) T cells near NR4A1-high regions, consistent with an immunosuppressed tumor microenvironment.The spatial ISH/mIF analyses quantitatively show fewer IFNγ+ T cells in proximity to NR4A1-high cells across multiple distance metrics, directly supporting the claim.Evidence: Proximity and nearest-neighbor analyses of NR4A1 and IFNγ spatial distribution across 150 ROIs and 41–71 biopsies.
Spatial analyses showed that IFNγ + T cells were distributed significantly further away from cells expressing high levels of NR4A1 ... compared to cells with low NR4A1 expression ( P < 0.0001)
Resultsreviewer’s wording - supportedReviewer 3Quemliclustat combined with G/nP with or without zimberelimab showed encouraging clinical response rates and survival in patients with mPDAC.The reported response rates (confirmed ORR 38% and 25%) and median OS (19.4 and 14.6 months) in the randomized arms support the 'encouraging' characterization, appropriately framed as a phase 1b result.Evidence: Confirmed ORR 38% (21–58) and 25% (15–37); median OS 19.4 and 14.6 months in the randomized arms.
“In the randomized arms, the confirmed objective response rate (ORR) was 38% (95% CI: 21–58) in the Q + G/nP arm and 25% (95% CI: 15–37) in the Q + G/nP + Z arm”
AbstractFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary efficacy claim is based on hard clinical outcomes, specifically overall survival (OS) and progression-free survival (PFS), with a synthetic control arm comparison showing significantly longer OS. OS is a validated hard clinical endpoint, not a surrogate.
“Median OS was significantly longer in the Quemli100 arm (15.7 months (95% CI: 12.4–20.9)) versus the SCA (9.8 months (95% CI: 7.8–11.4)) (P = 0.003)”
- ADEQUATEEffect sizeThe effect size is clinically material: median OS improvement of 5.9 months (HR 0.634, 95% CI 0.471–0.854, P=0.003) versus a propensity-matched synthetic control arm, and also compared favorably to historical benchmarks (NAPOLI 3 median OS 9.2 months; MPACT 8.7 months). This is anchored to meaningful clinical benefit in metastatic pancreatic cancer.
“Compared to the SCA, patients treated with the quemliclustat combinations demonstrated an increase in median OS of 5.9 months (hazard ratio = 0.634 (95% CI: 0.471–0.854); P = 0.003)”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
3 integrity concerns flagged (0 high).
- lowotherThe results text states 116 patients were enrolled in the dose-expansion phase, which is consistent (29+61=90 randomized plus the non-randomized arm), but the randomized-arm Ns (29 and 61) do not sum to the 90 planned randomized patients; the paper does not explicitly clarify the discrepancy between the planned 90 and the actual 90 versus the 116 total expansion enrollment.
116 patients were enrolled ... patients were enrolled and randomized 2:1 to receive ... (Q + G/nP arm; n = 29) or with zimberelimab (Q + G/nP + Z arm; n = 61)
Statistical analysisreviewer’s wording - lowotherThe post hoc SCA OS comparison (HR 0.634, P=0.003) is presented as a headline efficacy result despite being a non-prespecified exploratory comparison against non-concurrent historical controls; the paper does disclose this limitation, so this is a mild interpretive concern rather than a validity threat.
“Median OS was significantly longer in the Quemli100 arm (15.7 months (95% CI: 12.4–20.9)) versus the SCA (9.8 months (95% CI: 7.8–11.4)) ( P = 0.003)”
DiscussionFind in source
Reporting gaps
2 findings · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
- Key resources under-identified (antibodies, cell lines, RRIDs)Assessed
The introduction extensively cites prior research on CD73 and adenosine in cancer, the rationale for targeting CD73 in PDAC, and the limitations of existing adenosine signatures. The hypothesis that CD73 inhibition may benefit patients with mPDAC is logically derived from the cited evidence. The paper acknowledges that previous transcriptional signatures had limited predictive value, and proposes NR4A family expression as a novel surrogate, thereby addressing gaps in prior research.
“However, these transcriptional signatures have limited predictive value, in part because they may not adequately reflect the cellular heterogeneity in the TME.”
“However, these transcriptional signatures have limited predictive value, in part because they may not adequately reflect the cellular heterogeneity in the TME.”
“Quemliclustat is a potent and selective small‑molecule inhibitor of soluble and cell‑bound CD73 that is being studied in multiple tumor types”
“Multiple phase 3 trials of new therapeutic modalities have failed to improve on the standard care for mPDAC”
“these transcriptional signatures have limited predictive value, in part because they may not adequately reflect the cellular heterogeneity in the TME”
Randomization used the permuted block method at the patient level. The open-label design is explicitly stated (phase 1b dose-finding context). The sample-size justification is based on an estimation framework with a stated ORR-difference assumption (90% CI 3–37%), appropriate for a descriptive phase 1b study rather than a confirmatory power calculation. Eligibility criteria are given in text and fully in Supplementary Table. The analysis population (safety-evaluable) is defined. For a human phase 1b RCT, replicate_distinction, bench controls, and independent internal replication are not applicable.
“patients were enrolled and randomized 2:1 using the permuted block method”
“The sample size justification was based on an estimation framework, and the study was designed for descriptive statistical analysis rather than formal statistical hypothesis testing involving power and type I error considerations.”
“patients were enrolled and randomized 2:1 using the permuted block method to receive the quemliclustat RP2D combined with G/nP with or without zimberelimab.”
“The sample size justification was based on an estimation framework, and the study was designed for descriptive statistical analysis rather than formal statistical hypothesis testing involving power and type I error considerations.”
“patients were enrolled and randomized 2:1 using the permuted block method”
“the study was designed for descriptive statistical analysis rather than formal statistical hypothesis testing involving power and type I error considerations”
Table 1 reports age (median, range), sex, race, ECOG PS, liver metastasis, prior therapy, and time since diagnosis per arm. Both sexes are enrolled in all cohorts, so sex_justified is not applicable. Health status is captured via ECOG PS and eligibility criteria. Species/strain and housing criteria are not applicable to a human trial.
“Sex was recorded as a binary variable based on self-reported biological characteristics.”
The Methods state the study was conducted in conformance with the Declaration of Helsinki and approved by the local ethics committee at each site. Written informed consent was obtained from all patients. Regulatory compliance is explicitly stated.
“All patients provided written informed consent before any study procedures”
“The study was conducted in full conformance with the Declaration of Helsinki, the Council for International Organizations of Medical Sciences International Ethical Guidelines, institutional review board regulations and all other applicable local regulations.”
“The study protocol was approved by the local ethics committee at each site”
“All patients provided written informed consent before any study procedures”
“The study was conducted in full conformance with the Declaration of Helsinki, the Council for International Organizations of Medical Sciences International Ethical Guidelines, institutional review board regulations and all other applicable local regulations.”
Investigational products (quemliclustat, zimberelimab, gemcitabine, nab-paclitaxel) are named with doses and regimens, and bench reagents are identified with vendors and catalog numbers (e.g., 'AMP (Thermo Fisher Scientific, J61643.06)', 'EHNA (Sigma-Aldrich, 324630)'). Cell lines are attributed to ATCC but no STR authentication is stated. Mycoplasma testing is not mentioned. The mIF/CD3/CD8/PanCK antibodies are not identified with vendors/clones in the available text. Statistical software (SAS v.9.4) and R packages are identified. Approximately 2 of 5 applicable sub-criteria are adequate, so the dimension warrants a warning.
“Cell lines were purchased from the American Type Culture Collection”
“Data were extracted and standardized to ADaM datasets in SAS v.9.4”
“Cell lines were purchased from the American Type Culture Collection and cultured based on the supplier’s recommendations.”
“Cells were incubated in the presence of AMP (Thermo Fisher Scientific, J61643.06) and EHNA (Sigma-Aldrich, 324630) or NECA (Sigma-Aldrich, E2387)”
“The data were extracted and standardized to ADaM datasets in SAS v.9.4.”
“Human T cells were isolated from healthy donor blood using EasySep Human T Cell Isolation Kits (STEMCELL Technologies, 17592 and 17953)”
“Cell lines were purchased from the American Type Culture Collection and cultured based on the supplier’s recommendations.”
“The data were extracted and standardized to ADaM datasets in SAS v.9.4.”
The paper specifies tests (log-rank, Cox, Wilcoxon, etc.) and provides exact p-values and hazard ratios with 95% CIs. Kaplan-Meier curves and tables show adequate data presentation. Mathematical plausibility checks of reported percentages are consistent. However, no assumption verification for Cox models is provided, and only SAS v.9.4 is mentioned, not other software.
“PFS: hazard ratio = 0.732 (95% CI: 0.56–0.96); P = 0.0238”
“The analysis was not prespecified in the protocol and did not include formal futility boundaries; therefore, it was descriptive in nature and not intended to support definitive conclusions regarding efficacy or futility.”
“The NR4A family expression significantly correlated with survival benefit in the BEP (PFS: hazard ratio = 0.732 (95% CI: 0.56–0.96); P = 0.0238; OS: hazard ratio = 0.678 (95% CI: 0.49–0.95); P = 0.0239)”
“ORR was defined as the percentage of patients with a best overall response of complete or partial response and summarized with two-sided 95% CIs using the Clopper−Pearson method”
The data availability statement names a concrete access route: individual deidentified participant data available on request via a stated process (trials.arcusbio.com transparency policy) with conditions/exceptions, which counts as adequate managed access. However, the trial generated depositable RNA-seq data (transcriptome sequencing on 80 baseline tumors and 37 paired biopsies), yet no sequencing accession numbers (e.g., GEO) are given. Custom biomarker-analysis code is not shared or referenced. repository_deposit is not applicable for the identifiable patient-level data itself.
“Arcus Biosciences will provide access to individual deidentified participant data and related study documents (protocols, statistical analysis plans and clinical study reports) upon request from qualified researchers and subject to certain criteria, conditions and exceptions.”
“Arcus Biosciences will provide access to individual deidentified participant data and related study documents (protocols, statistical analysis plans and clinical study reports) upon request from qualified researchers and subject to certain criteria, conditions and exceptions.”
The study is registered at ClinicalTrials.gov (NCT04104672). Methods detail dosing, schedules, and assessments. Limitations (no concurrent control, phase 1b design) are discussed. Conclusions are appropriately cautious. However, the paper does not mention adherence to a specific reporting guideline like CONSORT, and some secondary endpoints (pharmacokinetics, immunogenicity) are listed but not reported in this manuscript.
“ClinicalTrials.gov identifier: NCT04104672”
“Additional planned secondary endpoints not reported in this paper are plasma concentration and pharmacokinetic parameters for quemliclustat, serum concentration and pharmacokinetic parameters for zimberelimab and number and percentage of patients who develop antidrug antibodies to zimberelimab.”
“There was no concurrent, randomized control group.”
“Additional planned secondary endpoints not reported in this paper are plasma concentration and pharmacokinetic parameters for quemliclustat”
Registered (2 IDs: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
5 data/code links checked; 5 live.
- datahttps://clinicaltrials.gov/study/NCT04104672LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://trials.arcusbio.com/our-transparency-policyLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT06608927LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- dataGEOLIVEHTTP 200https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE202051Resolves to GEO (data repository).
- codeGitHubLIVEHTTP 200https://github.com/ParkerICI/prince-trial-dataResolves to GitHub (code repository).
Copyediting
8 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 8 minor suggestions below.
8 copyedit issues flagged: mostly consistency, clarity, other.
- MINORconsistencyTitle“Quemliclustat”→ Use lowercase 'quemliclustat' for consistency with the rest of the text, or maintain capital if it is a proper noun, but ensure consistency throughout.The title capitalizes 'Quemliclustat' while the abstract and body use lowercase.
- MINORconsistencyTable 1“Quemli 25 mg”→ Consider spelling out 'Quemliclustat' or defining the abbreviation earlier, though it is commonly used in the field.The abbreviation 'Quemli' is used in tables but not defined in the table footnotes.
- MINORclarityTable 1, footnote c“c Stage not specified.”→ Clarify what 'stage not specified' refers to (likely time since initial diagnosis), as the phrasing is ambiguous.Footnote is terse and could be misread.
- MINORconsistencyResults, Biomarker analysis“basal-like molecular subtype ( n = 14 (18%))”→ Verify the percentage: 14 of 80 BEP patients is 17.5%, reported as 18%.Rounding inconsistency of 0.5 percentage point; not a substantive error.
- MINORotherExtended Data Fig. 4 legend“a , RNA-seq and b , spatial analysis from the ARC-8 study.”→ Consider spelling out the full arm labels in the breakdown for accessibility.Minor legibility point.
- MINORconsistencyAuthor contributions (P.E.O.)“project administration, project administration, resources”→ Remove the duplicated 'project administration' entry.Duplicate phrase in author contribution list.
- MINORconsistencyAuthor contributions (B.B.)“project administration, project administration, supervision”→ Remove the duplicated 'project administration' entry.Duplicate phrase in author contribution list.
- MINORconsistencyExtended Data Fig. 3 caption“Schematic in c created in BioRender; Direnzo, D.”→ Use 'DiRenzo, D.' to match the author byline spelling.Author name spelling inconsistency between caption and byline.
This published work is methodologically robust for a phase 1b exploratory trial: design, ethics, and statistical reporting are sound, and the recomputed statistics were consistent. An informed reader should weigh the following: the key-resource identification gaps (mIF/ISH antibody and probe identifiers, cell-line authentication, mycoplasma testing), the absence of deposited sequencing data and analysis code, and the exploratory, non-prespecified nature of the headline SCA survival comparison (HR 0.634, P=0.003) against non-concurrent historical controls — which the paper does disclose but which should temper how strongly the biomarker and synthetic-control efficacy claims are interpreted. None of these rise to a validity threat warranting retraction, but they warrant a correction/companion data deposit by the authors.
- 1.HIGHdata codeDeposit the transcriptomic data (baseline BEP RNA-seq n=80 and 37 paired pre/posttreatment biopsies) in a public repository (e.g., GEO/EGA) and report the accession number(s) in the Data availability section.The paper generated depositable sequencing data but provides no accession numbers, so the data cannot be independently re-analyzed.
- 2.HIGHrigorAdd vendor/catalog/RRID/clone and dilution identifiers for the multiplex immunofluorescence antibodies (CD3, CD8, PanCK, CD4) and the NR4A ISH probe in the spatial biomarker Methods.The mIF and ISH assays are central to the biomarker claims but their antibodies/probes are not identified, preventing replication.
- 3.HIGHrigorAdd cell-line STR authentication and mycoplasma-testing statements (method and result) for PANC-1, MIA PaCa-2, HCT-116, and NCI-H650 in the Cell culture experiments section.ATCC provenance is stated but authentication and mycoplasma status are not, which the rigor standards require for in vitro cell lines.
- 4.HIGHdata codeShare the custom analysis code (NR4A ssGSEA scoring, spatial ISH/mIF quantification, propensity-score matching, synthetic-control analysis) in a versioned public repository with a persistent identifier, referenced in the Data availability or Methods.The post hoc biomarker and SCA analyses rely on custom code that is not available, limiting transparency and reproducibility.
- 5.HIGHreportingReference the CONSORT reporting checklist (or state reporting-summary availability) for the randomized dose-expansion cohort in the Methods.The paper reports no reporting-guideline adherence, a standard expectation for a randomized controlled cohort.
- 6.HIGHreportingState where the disclosed-but-unreported secondary endpoints (PK parameters for quemliclustat/zimberelimab and antidrug antibodies) will be published, or include them in a supplement.Pre-specified secondary endpoints are listed as 'not reported in this paper' without a stated location, leaving the all_outcomes_reported gap open.
- 7.HIGHstatisticsCorrect the molecular-subtype percentage inconsistency in the BEP breakdown (83% classical + 18% basal-like = 101%; 14/80 = 17.5% reported as 18%).The percentages sum to over 100% and one value is mis-rounded, an internal inconsistency that undermines the reported subtype distribution.
- 8.HIGHreportingClarify the dose-expansion enrollment numbers in the Results: 29 + 61 = 90 randomized patients versus the 116 total dose-expansion enrollment, and reconcile with the planned 90.The randomized-arm Ns do not obviously reconcile with the stated total enrollment, and the paper does not explicitly explain the discrepancy.
- 9.MEDIUMstatisticsAdd a statement on verification of the proportional-hazards assumption for the Cox models in the Methods/Statistical analysis section.The Cox models are used for key survival claims but no assumption check is reported, a gap flagged by one reviewer.
- 10.MEDIUMreportingExplicitly label the SCA OS comparison (HR 0.634, P=0.003) in the Results/Discussion as a non-prespecified exploratory analysis against non-concurrent historical controls.Although the limitation is disclosed, the headline framing should more strongly signal that this is exploratory and not a confirmatory efficacy result.
- 11.MEDIUMrigorReport patient weight in the baseline demographics (Table 1) to complete the age/weight/health characterization.Weight is part of the standard baseline biological-variable reporting and is currently absent.
- 12.LOWcopyeditRemove the duplicated 'project administration' entries in the author contributions for P.E.O. and B.B.The duplicate phrase is a copyedit error in the contribution list.
- 13.LOWcopyeditCorrect 'Direnzo, D.' to 'DiRenzo, D.' in the Extended Data Fig. 3 caption to match the author byline spelling.The caption and byline spell the author name differently.
- 14.LOWcopyeditStandardize 'Quemliclustat'/'quemliclustat' capitalization in the title to be consistent with the abstract and body, and define the 'Quemli' abbreviation in the Table 1 footnotes.Inconsistent capitalization of the investigational drug name and an undefined table abbreviation are minor consistency issues.
- 15.LOWcopyeditClarify Table 1 footnote c ('Stage not specified') to state what it refers to (e.g., time since initial diagnosis).The terse footnote is ambiguous and could be misread.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.