Phase 3 Trial of Cabozantinib to Treat Advanced Neuroendocrine Tumors.
Chan JA, Geyer S, Zemla T, Knopp MV, Behr S, Pulsipher S, Ou FS, Dueck AC, Acoba J, Shergill A, Wolin EM, Halfdanarson TR, Konda B, Trikalinos NA, Tawfik B, Raj N, Shaheen S, Vijayvergia N, Dasari A, Strosberg JR, Kohn EC, Kulke MH, O'Reilly EM, Meyerhardt JA
- DOI
- 10.1056/NEJMoa2403991
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/1099ca0c-ae0d-4cc4-89cd-85bb43b320d7 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×4−2★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ReportingData & code availability partially met−0.25★
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is progression-free survival (PFS) by blinded independent central review, which is a surrogate for overall survival. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD data) nor cite validated evidence linking PFS to a meaningful clinical outcome in this setting. Overall survival data are immature and not significantly different.
“The primary endpoint was progression-free survival by blinded independent central review.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a rigorously designed and well-reported phase 3 clinical trial with strong methodology across most dimensions. The main weakness is the lack of a clear data availability statement and repository deposit, which is a common but important reporting gap.
Both reviewers independently scored all eight dimensions and agreed on all statuses, with only minor differences in evidence detail. The study type is interventional, consistent across reviewers. Non-applicable sub-criteria (e.g., animal housing, cell line authentication) were excluded from scoring.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 4 tests: 4 consistent, 0 inconsistent; 2 recomputed directly from the reported test statistics, 2 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Recomputed HR 0.38 (95% CI 0.25–0.59), reported p<0.001
“HR = 0.38, 95% CI: 0.25–0.59; p<0.001”
Taken as given: 0.25–0.59 is a two-sided 95% confidence interval for the HR of 0.38, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p<0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.38, 0.25, 0.59, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Recomputed HR 0.23 (95% CI 0.12–0.42), reported p<0.001
“HR = 0.23, 95% CI: 0.12–0.42; p<0.001”
Taken as given: 0.12–0.42 is a two-sided 95% confidence interval for the HR of 0.23, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p<0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.23, 0.12, 0.42, 1) - CONSISTENTreported p = .050 · recomputed p = .098Reviewer 2P-value for epNET ORR comparison (5% vs 0%)
“Partial responses were observed in 5% of patients in the cabozantinib group (95% CI: 2%−10%) compared with 0% (95% CI: 0%−5%) in the placebo group (p=0.05).”
Taken as given: The number of partial responses in the cabozantinib group is 7 (5% of 134).; The number of partial responses in the placebo group is 0.; The total numbers are 134 and 69 for cabozantinib and placebo, respectively.; The p-value is two-sided.Method: Recomputed two-sided Fisher's exact test p-value from the 2x2 table of response counts.How we recomputed it: pFisher2x2(7, 127, 0, 69, 0) - CONSISTENTreported p = .010 · recomputed p = .008Reviewer 2P-value for pNET ORR comparison (19% vs 0%)
“Partial responses were observed in 19% of patients in the cabozantinib group (95% CI: 10%−30%) compared with 0% (95% CI: 0%−11%) in the placebo group (p=0.01).”
Taken as given: The number of partial responses in the cabozantinib group is 12 (19% of 64).; The number of partial responses in the placebo group is 0.; The total numbers are 64 and 31 for cabozantinib and placebo, respectively.; The p-value is two-sided.Method: Recomputed two-sided Fisher's exact test p-value from the 2x2 table of response counts.How we recomputed it: pFisher2x2(12, 52, 0, 31, 0)
- lowinternal contradictionThe abstract states 'Grade 3 or higher adverse events were noted in 62–65% of patients treated with cabozantinib compared to 23–27% with placebo.' The results section reports 62% and 65% for cabozantinib and 27% and 23% for placebo in the epNET and pNET cohorts, respectively. This is consistent, but the abstract uses ranges that could be misread.
“Grade 3 or higher adverse events were noted in 62–65% of patients treated with cabozantinib compared to 23–27% with placebo.”
AbstractFind in source - lowinternal contradictionIn the epNET cohort, the number of patients with disease progression or death (71+40=111) matches the reported progression events, but the percentages (53% and 58%) do not match the denominators (134 and 69) exactly.
“71 participants (53%) in the cabozantinib group and 40 (58%) in the placebo group had disease progression or had died.”
ResultsFind in source - lowinternal contradictionThe abstract states '203 patients with extra-pancreatic NET and 95 patients with pancreatic NET' but the results section mentions misallocation of patients between cohorts, which could affect the exact numbers.
“The trial enrolled two independent cohorts of patients, including 203 patients with extra-pancreatic NET and 95 patients with pancreatic NET”
AbstractFind in source - lowinternal contradictionThe results state '111 progression events by BICR were reported' for the epNET cohort, but the planned number of events was 164. This is not a contradiction but a note that the trial was stopped early.
“Of the 203 patients in this cohort, 111 progression events by BICR were reported”
ResultsFind in source
Overstated conclusions
1 finding · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
5 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Cabozantinib significantly improves progression-free survival in patients with previously treated, progressive advanced extra-pancreatic or pancreatic NET.The primary endpoint PFS is significantly improved in both cohorts with HRs and p-values.Evidence: Stratified HR 0.38 (95% CI 0.25-0.59, p<0.001) for epNET and HR 0.23 (95% CI 0.12-0.42, p<0.001) for pNET.
“Cabozantinib significantly improves progression-free survival in patients with previously treated, progressive advanced extra-pancreatic or pancreatic NET.”
ConclusionFind in source - supportedReviewers 1, 2Adverse events were consistent with the known safety profile of cabozantinib.The safety data show expected adverse events like hypertension, fatigue, and diarrhea, consistent with prior studies.Evidence: Grade 3 or higher adverse events in 62-65% of cabozantinib-treated patients, with common events listed.
“Adverse events were consistent with the known safety profile of cabozantinib.”
ConclusionFind in source - supportedReviewer 1Cabozantinib is a new treatment option for patients with advanced NET after progression on prior therapy.The trial demonstrates PFS benefit, supporting its use as a treatment option.Evidence: PFS benefit in both cohorts, with manageable safety profile.
“The results support the use of cabozantinib as a new treatment option for patients with advanced extra-pancreatic NET or pancreatic NET whose disease has progressed after or who have experienced intolerance of at least one other line of therapy for their disease, not including somatostatin analogs.”
Discussion ¶2Find in source - supportedReviewer 2Cabozantinib resulted in a higher confirmed response rate than placebo in both cohorts.The response rates are higher with cabozantinib in both cohorts, with p-values of 0.05 and 0.01.Evidence: In epNET, ORR 5% vs 0% (p=0.05). In pNET, ORR 19% vs 0% (p=0.01).
“Cabozantinib resulted in a higher confirmed response rate than placebo.”
ResultsFind in source - supportedReviewer 2No overall survival difference between the treatment groups has been observed to date.The OS data are immature and show no significant difference, as reported.Evidence: In epNET, median OS 21.9 vs 19.7 months (HR 0.86, 95% CI 0.56-1.31). In pNET, median OS 40 vs 31.1 months (HR 0.95, 95% CI 0.45-2.00).
“No overall survival difference between the treatment groups has been observed to date; however, overall survival data were not mature at the time of analyses and may have been impacted by crossover and the high rate of treatment with subsequent anticancer therapies.”
Discussion ¶1Find in source
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary endpoint is progression-free survival (PFS) by blinded independent central review, which is a surrogate for overall survival. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD data) nor cite validated evidence linking PFS to a meaningful clinical outcome in this setting. Overall survival data are immature and not significantly different.
“The primary endpoint was progression-free survival by blinded independent central review.”
- ADEQUATEEffect sizeThe effect sizes are large and clinically meaningful: median PFS improved from 3.9 to 8.4 months in epNET (HR 0.38) and from 4.4 to 13.8 months in pNET (HR 0.23), with highly significant p-values. These are substantial improvements over placebo and are anchored to a clinically relevant outcome (progression-free survival).
“median progression-free survival with cabozantinib was 8.4 months compared to 3.9 months with placebo (stratified hazard ratio [HR], 0.38; 95% confidence interval [CI] 0.25–0.59, p<0.001). In patients with pancreatic NET, median progression-free survival was 13.8 months compared to 4.4 months with placebo (stratified HR, 0.23; 95% CI 0.12–0.42, p<0.001).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior work on NET pathogenesis, antiangiogenic agents, and a phase II trial of cabozantinib, establishing the premise. The rationale linking cabozantinib's mechanism to NET treatment is clear. Limitations of prior research are implicitly addressed by the trial design, though not explicitly discussed.
“Cabozantinib is an oral small-molecule inhibitor of multiple tyrosine kinases including VEGF receptors, MET, AXL, and RET. Clinical activity of cabozantinib was demonstrated in a phase II trial that enrolled patients with advanced epNET or pNET.”
“Based on these results, we conducted a randomized, double-blinded, placebo-controlled study to evaluate the efficacy of cabozantinib in patients with previously treated, progressive epNET or pNET (CABINET; Alliance A021602).”
“Based on these results, we conducted a randomized, double-blinded, placebo-controlled study to evaluate the efficacy of cabozantinib in patients with previously treated, progressive epNET or pNET (CABINET; Alliance A021602).”
Randomization method (2:1 ratio) and stratification factors are described. Blinding is double-blinded with BICR. Power analysis is detailed with target sample sizes and hazard ratios. Inclusion/exclusion criteria are pre-specified. Outlier handling is not explicitly described, but the analysis population (ITT) and missing data approach are defined. Controls are appropriate (placebo). Independent replication is not applicable for a single pivotal trial.
“Patients continued blinded treatment until disease progression, unacceptable toxicity, or withdrawal of consent.”
“Patients continued blinded treatment until disease progression, unacceptable toxicity, or withdrawal of consent.”
Sex is reported for both cohorts. Age and health status (ECOG PS) are reported. Demographics are comprehensive. Species/strain and housing conditions are not applicable for a human trial.
“Female Sex, n (%) | 74 (55) | 31 (45)”
“Age, years, median (range) | 66 (28–86) | 66 (30–82)”
“Female Sex, n (%) | 74 (55) | 31 (45)”
“Age, years, median (range) | 66 (28–86) | 66 (30–82)”
The protocol was approved by the NCI central IRB. Written informed consent was obtained from all participants. The trial was conducted in accordance with the Declaration of Helsinki and ICH-GCP guidelines.
“The protocol was approved by the NCI central institutional review board (CIRB).”
“All participants provided written informed consent before enrollment.”
“The trial was conducted in accordance with the principles of the Declaration of Helsinki and the International Council for Harmonisation Good Clinical Practice guidelines.”
“The protocol was approved by the NCI central institutional review board (CIRB).”
“All participants provided written informed consent before enrollment.”
“The trial was conducted in accordance with the principles of the Declaration of Helsinki and the International Council for Harmonisation Good Clinical Practice guidelines.”
Cabozantinib is named with dose and regimen. The placebo is described. Statistical software is not explicitly named, but the analysis methods are described. No other biological/chemical resources are used.
“Patients received cabozantinib (60 mg) or placebo orally once daily.”
“Patients received cabozantinib (60 mg) or placebo orally once daily.”
“Trial registration number: NCT03375320 (https://clinicaltrials.gov/ct2/show/NCT03375320)”
Tests are named (stratified log-rank, Cox regression, chi-square). Assumptions are handled by design (stratified Cox). Exact p-values are reported (p<0.001). Effect sizes with CIs are reported. Software is not explicitly identified. Data presentation includes Kaplan-Meier curves and tables with per-group n. Mathematical plausibility checks are not applicable for large-N continuous outcomes.
“Hazard ratios (HR) were estimated through stratified Cox regression models and stratified log-rank tests calculated using stratification factors at randomization per the protocol design.”
“stratified HR = 0.38, 95% CI: 0.25–0.59; p<0.001”
“Median progression-free survival was 8.4 months (95% CI: 7.6–12.7 months) with cabozantinib and 3.9 months (95% CI: 3.0–5.7 months) with placebo.”
“Hazard ratios (HR) were estimated through stratified Cox regression models and stratified log-rank tests calculated using stratification factors at randomization per the protocol design.”
“stratified HR = 0.38, 95% CI: 0.25–0.59; p<0.001”
“Median progression-free survival was 8.4 months (95% CI: 7.6–12.7 months) with cabozantinib and 3.9 months (95% CI: 3.0–5.7 months) with placebo.”
The paper mentions that the protocol is available at NEJM.org, but does not provide a clear data availability statement for the clinical data. No repository deposit or accession numbers are provided. Code sharing is not applicable as no bespoke code is described.
“The protocol is available at nejm.org (https://nejm.org)”
Methods are detailed enough for replication. Trial registration number is provided (NCT03375320). Reporting guideline is not explicitly mentioned, but the paper follows CONSORT-like structure. All pre-specified outcomes are reported. Limitations are discussed. Conclusions are proportional. Funding and COI are disclosed.
“Trial registration number: NCT03375320 (https://clinicaltrials.gov/ct2/show/NCT03375320)”
“The early termination of the trial based on interim analysis results could potentially lead to overestimation of treatment effect.”
“Trial registration number: NCT03375320 (https://clinicaltrials.gov/ct2/show/NCT03375320)”
“The early termination of the trial based on interim analysis results could potentially lead to overestimation of treatment effect.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 31 references by DOI: 28 verified — 3 no DOI (shown, not verified).
- NO DOIPhase II trial of cabozantinib in patients with carcinoid and pancreatic neuroendocrine tumors (pNET)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA phase II/III randomized double-blind study of octreotide acetate with axitinib versus octreotide acetate with placebo in patients with advanced G1-G2 NETs of non-pancreatic origin (AXINET trial-GETNE-1107)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIProspective randomized phase II trial of pazopanib versus placebo in patients with progressive carcinoid tumors (CARC)(Alliance A021202)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly typo, clarity, consistency.
- MINORtypoAbstract, Results“62–65%”→ Use consistent formatting for ranges (e.g., 62%-65%)Inconsistent use of en dash and percent sign.
- MINORconsistencyResults, epNET cohort“111 progression events by BICR were reported”→ Clarify whether this is the number of events or the number of patients with events.Potential ambiguity in event counting.
- MINORclarityMethods, Statistical Analysis“Two-sided P values are reported in accordance with Journal policy”→ Specify the journal policy or provide a reference.Could be clearer for readers.
- MINORtypoAbstract, Results“62–65%”→ Use consistent en-dash or hyphen; consider '62%–65%'.Minor formatting inconsistency.
- MINORclarityMethods, Statistical Analysis“Two-sided P values are reported in accordance with Journal policy”→ Consider specifying the journal name for clarity.Minor clarity issue.
The published work is robust and methodologically sound, with no major integrity concerns. An informed reader should weigh the minor reporting gaps (data availability, statistical software identification, reporting guideline) and the early termination of the trial as potential limitations, but these do not undermine the core findings.
- 1.HIGHdata codeAdd a clear data availability statement in the Methods or a dedicated section, specifying how de-identified patient data can be accessed (e.g., via the Alliance or a data-sharing platform like Vivli) and any conditions.The current statement is vague or absent, and a clear data availability statement is expected for clinical trial publications.
- 2.HIGHdata codeDeposit the statistical analysis code (e.g., SAS or R scripts) in a public repository with a DOI to enhance reproducibility.Sharing analysis code improves transparency and allows independent verification of the reported results.
- 3.MEDIUMreportingMention adherence to a reporting guideline such as CONSORT in the Methods or cover letter, and include the CONSORT flow diagram in the supplement.Explicitly stating reporting guideline adherence demonstrates transparency and completeness.
- 4.MEDIUMstatisticsIdentify the statistical software (e.g., SAS version, R version) used for analyses in the Statistical Analysis section.Naming the software version is a standard reporting requirement and aids reproducibility.
- 5.MEDIUMreportingClarify the handling of misallocated patients in the primary analysis (e.g., sensitivity analysis) in the Statistical Analysis section.The paper mentions misallocation of patients between cohorts, and clarifying the analysis approach would improve transparency.
- 6.LOWcopyeditFix the inconsistent formatting of percentage ranges in the Abstract (e.g., change '62–65%' to '62%–65%').Consistent formatting improves readability and professionalism.
- 7.LOWcopyeditClarify whether '111 progression events by BICR' refers to the number of events or the number of patients with events in the Results section.Ambiguity in event counting could confuse readers.
- 8.LOWcopyeditSpecify the journal policy or provide a reference for the statement 'Two-sided P values are reported in accordance with Journal policy' in the Methods.Clarifying the policy reference improves transparency.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.