Adjuvant nivolumab and relatlimab in stage III/IV melanoma: the randomized phase 3 RELATIVITY-098 trial.
Long GV, Garnett-Benson C, Dolfi S, Ascierto PA, Guo J, Tarhini AA, Chandra S, Muñoz-Couselo E, Del Vecchio M, de Melo AC, Callahan M, Gogas H, Dummer R, Schadendorf D, Koelblinger P, Quereux G, Thomas I, Yu JX, Fisher A, Wang B, Djidel P, Chouzy A, Semaan M, Chen B, Cheong AMY, Tawbi HA
- DOI
- 10.1038/s41591-025-04032-8
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/f85abdb9-db23-47af-9fbb-b52e1bc8aeb4 is authoritative.
How this rating was calculated
Started at 5★ — no deductions. Nothing the checks ran surfaced a material problem.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported phase 3 randomized controlled trial. The paper demonstrates strong methodological rigor across all eight dimensions, with clear randomization, blinding, power analysis, and comprehensive reporting of demographics, ethics, resources, statistics, data availability, and transparency. The only minor issues are a typo in the trial name and a few reporting enhancements that could be made.
Both reviewers independently scored all eight dimensions and agreed on every status, indicating high sampling stability. The study type is interventional (phase 3 RCT). Non-applicable sub-criteria (e.g., animal-related, cell line authentication) were excluded from scoring. The statistics verification component recomputed only 3 reported tests consistently; other statistics were not machine-verified and should not be assumed correct.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 3 tests: 3 consistent, 0 inconsistent; 2 recomputed directly from the reported test statistics, 1 via agent-written checks.
- CONSISTENTreported p = .928 · recomputed p = .919Recomputed hazard ratio 1.01 (95% CI 0.83–1.22), reported p=0.928
“hazard ratio = 1.01; 95% confidence interval: 0.83–1.22; P = 0.928”
Taken as given: 0.83–1.22 is a two-sided 95% confidence interval for the hazard ratio of 1.01, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.928 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.01, 0.83, 1.22, 1) - CONSISTENTreported p = .006 · recomputed p = .004Recomputed hazard ratio 0.75 (95% CI 0.62–0.92), reported p=0.006
“hazard ratio = 0.75; 95% confidence interval: 0.62–0.92; P = 0.006”
Taken as given: 0.62–0.92 is a two-sided 95% confidence interval for the hazard ratio of 0.75, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.006 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.75, 0.62, 0.92, 1) - CONSISTENTreported p = .928 · recomputed p = .919Reviewers 1, 2Primary RFS hazard ratio p-value
“hazard ratio for nivolumab plus relatlimab versus nivolumab of 1.01 (95% confidence interval: 0.83–1.22; P = 0.928)”
Taken as given: The hazard ratio is 1.01 and the 95% CI is 0.83-1.22.; The p-value is two-sided from a log-rank test, approximated by the CI-based method.Method: Approximated two-sided p-value from the hazard ratio and 95% CI using the normal approximation.How we recomputed it: pCI(1.01, 0.83, 1.22, 1)
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
4 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2The absence of macroscopic tumor and reduced peripheral LAG-3+ T cells may explain the lack of added benefit.The translational data support a hypothesis, but the paper acknowledges limitations and the need for direct testing.Evidence: Comparative biomarker analyses between RELATIVITY-098 and RELATIVITY-047 show lower LAG-3+ T cells in adjuvant setting.
“The absence of macroscopic tumor and reduced peripheral LAG-3 + T cells may explain the lack of added benefit of nivolumab plus relatlimab over nivolumab in resected versus metastatic melanoma.”
AbstractFind in source - supportedReviewers 1, 2Nivolumab plus relatlimab did not significantly improve RFS compared to nivolumab in resected stage III/IV melanoma.The primary endpoint result directly supports this claim.Evidence: Primary RFS analysis: HR 1.01, 95% CI 0.83-1.22, P=0.928.
“There was no difference in RFS for nivolumab plus relatlimab versus nivolumab (hazard ratio = 1.01; 95% confidence interval: 0.83–1.22; P = 0.928)”
AbstractFind in source - supportedReviewer 1Higher LAG-3 and CD8 expression in baseline tumors enriched for RFS benefit in both arms.The paper presents correlative data supporting this claim.Evidence: Extended Data Fig. 5 shows RFS association with baseline TME biomarkers.
“When correlating tumor expression with RFS, higher LAG-3 and CD8 expression in resected tumors enriched for RFS benefit in both the nivolumab plus relatlimab and nivolumab arms (Extended Data Fig. ).”
ResultsFind in source - supportedReviewer 2The safety profile of nivolumab plus relatlimab in RELATIVITY-098 was similar to that in RELATIVITY-047, with no new safety signals.Safety data are presented and compared to prior trial, supporting the claim.Evidence: Safety summary table and discussion of TRAEs.
“The safety profile of nivolumab plus relatlimab in RELATIVITY-098 was similar to that in RELATIVITY-047, with no new safety signals”
DiscussionFind in source
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites standard-of-care adjuvant therapies and their limitations (e.g., 5-year RFS rates around 50%), and the rationale for the trial is logically derived from the RELATIVITY-047 results showing benefit in advanced melanoma. The paper acknowledges the unmet need for more efficacious adjuvant regimens. Limitations of prior research are implicitly addressed by the trial design, though not explicitly discussed as a separate section.
“However, approximately 50% of patients have disease recurrence, with 5-year RFS rates of 50% for nivolumab , 55% for pembrolizumab and 52% for dabrafenib plus trametinib .”
“Nivolumab plus relatlimab fixed-dose combination (FDC; 480-mg nivolumab plus 160-mg relatlimab intravenously every 4 weeks) was approved for treatment of advanced melanoma based on the phase 2/3 RELATIVITY-047 trial .”
“Given the unmet need for more efficacious adjuvant regimens for completely resected melanoma, the phase 3 RELATIVITY-098 trial compared adjuvant treatment with nivolumab plus relatlimab FDC to nivolumab in patients after complete resection of stage III/IV melanoma.”
Randomization method (permuted blocks within strata) and unit (patient) are stated. Blinding is described (double-blind). Power analysis is provided with sample size, effect size, alpha, and power. Inclusion/exclusion criteria are prespecified. Outlier handling is addressed through the ITT analysis population and safety population definitions. Controls are inherent in the comparator arm. Independent replication is not applicable for a single pivotal trial.
“Randomization was carried out via permutated blocks within each stratum.”
“The sponsor, participants, investigator and site staff were blinded to the study therapy administered.”
“An approximate sample size of 1,050 patients was planned to achieve the required 410 RFS events and show a significant difference in investigator-assessed RFS with a two-sided α of 0.05 using a stratified log-rank test with at least 90% statistical power when the average hazard ratio of nivolumab and relatlimab FDC versus nivolumab was 0.72 and an assumed cure rate of 0.52 in the nivolumab arm.”
“Randomization was carried out via permutated blocks within each stratum. The sponsor, participants, investigator and site staff were blinded to the study therapy administered.”
“An approximate sample size of 1,050 patients was planned to achieve the required 410 RFS events and show a significant difference in investigator-assessed RFS with a two-sided α of 0.05 using a stratified log-rank test with at least 90% statistical power when the average hazard ratio of nivolumab and relatlimab FDC versus nivolumab was 0.72 and an assumed cure rate of 0.52 in the nivolumab arm.”
“Eligible patients were at least 12 years of age with an Eastern Cooperative Oncology Group (ECOG) performance status of 0–1 and stage IIIA (>1-mm tumor in lymph node), stage IIIB/C/D or stage IV (no evidence of disease) melanoma (per the American Joint Committee on Cancer (AJCC) Cancer Staging Manual (eighth edition) (AJCC-8)) completely resected <90 days from randomization.”
Sex is reported for both arms (male/female percentages). Age is reported as median and range. Health status is captured via ECOG performance status and disease stage. Demographics include geographic region. Species/strain and housing conditions are not applicable for a human trial.
“Male | 327 (59.8) | 315 (57.7) | | Female | 220 (40.2) | 231 (42.3)”
“Median age (range) — years | 59 (18‒89) | 59 (19‒92)”
“Median age (range) — years | 59 (18‒89) | 59 (19‒92) | | Sex — no. (%) | | Male | 327 (59.8) | 315 (57.7) | | Female | 220 (40.2) | 231 (42.3)”
“Geographic region — no. (%) | | United States/Canada | 54 (9.9) | 63 (11.5) | | Australia | 67 (12.2) | 57 (10.4) | | Europe | 318 (58.1) | 319 (58.4) | | Latin America | 76 (13.9) | 72 (13.2) | | China | 32 (5.9) | 35 (6.4)”
The methods state that the protocol was reviewed by the institutional review board or independent ethics committee for each trial site, and all patients provided written informed consent. The trial was conducted in accordance with ICH-GCP guidelines. This satisfies the requirements for human research.
“The protocol and amendments for this trial (available in the ) were reviewed by the institutional review board or independent ethics committee for each trial site, and all patients provided written informed consent before enrollment.”
“The trial was conducted in accordance with International Council for Harmonization Good Clinical Practice guidelines”
“The protocol and amendments for this trial (available in the ) were reviewed by the institutional review board or independent ethics committee for each trial site, and all patients provided written informed consent before enrollment.”
“The trial was conducted in accordance with International Council for Harmonization Good Clinical Practice guidelines”
The drugs are named with manufacturer (Bristol Myers Squibb) and dosing regimen. The software used for data collection and analysis is identified (Medidata Classic Rave, SAS, R). Antibodies, cell lines, and mycoplasma testing are not applicable as this is a clinical trial without wet-lab assays.
“Data collection software used was Medidata Classic Rave (version 2025). All clinical analyses were performed using SAS software (SAS Institute) and R version 4.3.1.”
“Data collection software used was Medidata Classic Rave (version 2025). All clinical analyses were performed using SAS software (SAS Institute) and R version 4.3.1.”
The primary analysis uses a two-sided log-rank test and stratified Cox model, with HR and 95% CI reported. The paper reports exact p-values (e.g., P = 0.928). Assumptions are handled by design (stratified log-rank, Cox proportional hazards). Software is identified. Data presentation includes Kaplan-Meier curves and forest plots. Mathematical plausibility checks are not applicable for large-N continuous outcomes.
“RFS was compared using a two-sided log-rank test on all randomized patients; hazard ratio and confidence intervals were estimated using a stratified Cox proportional hazards model and survival curves using Kaplan–Meier methodology.”
“hazard ratio for nivolumab plus relatlimab versus nivolumab of 1.01 (95% confidence interval: 0.83–1.22; P = 0.928)”
“hazard ratio for nivolumab plus relatlimab versus nivolumab of 1.01 (95% confidence interval: 0.83–1.22; P = 0.928)”
“The primary endpoint of RFS was not statistically significantly associated in the intention-to-treat population, with a hazard ratio for nivolumab plus relatlimab versus nivolumab of 1.01 (95% confidence interval: 0.83–1.22; P = 0.928)”
“RFS was compared using a two-sided log-rank test on all randomized patients; hazard ratio and confidence intervals were estimated using a stratified Cox proportional hazards model and survival curves using Kaplan–Meier methodology.”
The data availability statement provides a clear mechanism for requesting data through BMS and Vivli, with conditions. This is adequate for patient-level data. Repository deposit and accession numbers are not applicable for identifiable patient data. Code sharing is not applicable as no bespoke code is mentioned.
The trial is registered (NCT05002569). Methods are detailed enough for replication. Limitations are discussed, including the lack of CODEX analyses in RELATIVITY-098. Conclusions are proportional, noting the lack of benefit and potential reasons. Funding and competing interests are disclosed. Reporting guideline is not explicitly mentioned but the paper follows CONSORT-like flow diagram.
“ClinicalTrials.gov identifier: NCT05002569 (https://clinicaltrials.gov/study/NCT05002569) .”
“Although informative, these analyses have inherent limitations preventing adequate conclusions to be drawn, including that multiplex CODEX analyses were not available for the RELATIVITY-098 study.”
“The RELATIVITY-098 trial was funded by Bristol Myers Squibb, which collected the data and analyzed the data in collaboration with the authors.”
“ClinicalTrials.gov identifier: NCT05002569 (https://clinicaltrials.gov/study/NCT05002569) .”
“Although informative, these analyses have inherent limitations preventing adequate conclusions to be drawn, including that multiplex CODEX analyses were not available for the RELATIVITY-098 study.”
“The RELATIVITY-098 trial was funded by Bristol Myers Squibb, which collected the data and analyzed the data in collaboration with the authors.”
Registered (2 IDs: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 19 references by DOI: 16 verified — 3 no DOI (shown, not verified).
- NO DOIAdjuvant nivolumab versus ipilimumab in resesected stage III/IV melanoma: 7-year results from CheckMate 238No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMerck provides update on phase 3 KeyVibe-010 trial evaluating an investigational coformulation of vibostolimab and pembrolizumab as adjuvant treatment for patients with resected high-risk melanomaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICancer Staging Manual, Eighth ednNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
3 data/code links checked; 3 live.
- datahttps://clinicaltrials.gov/study/NCT05002569LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.bms.com/researchers-and-partners/independent-research/data-sharing-request-process.htmlLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://vivli.org/ourmember/bristol-myers-squibb/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
1 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 1 minor suggestion below.
1 copyedit issue flagged: mostly typo.
- MINORtypoDiscussion, last paragraph“RELATVITY-098”→ RELATIVITY-098Typo in trial name.
The published work is robust and well-reported. An informed reader should weigh the minor reporting gaps (e.g., lack of explicit CONSORT statement, limited detail on randomization implementation) and the typo, but these do not undermine the validity of the findings. No erratum is warranted beyond correcting the typo.
- 1.MEDIUMcopyeditCorrect the typo 'RELATVITY-098' to 'RELATIVITY-098' in the Discussion, last paragraph.A misspelling of the trial name in a published article is a minor but visible error that should be corrected.
- 2.MEDIUMreportingExplicitly state adherence to CONSORT guidelines in the Methods or Reporting Summary.While the paper includes a CONSORT diagram, explicitly naming the guideline strengthens reporting transparency.
- 3.MEDIUMreportingProvide a more detailed description of the randomization implementation (e.g., interactive response technology) in the Methods.Detailing the randomization implementation enhances reproducibility and transparency.
- 4.MEDIUMstatisticsClarify the handling of missing data for the primary endpoint analysis (e.g., censoring rules) in the Statistical analysis section.Explicitly describing censoring and missing data handling helps readers assess the robustness of the primary analysis.
- 5.MEDIUMstatisticsInclude a statement on whether any analyses were adjusted for multiple comparisons in the translational biomarker analyses.The paper notes no adjustment was made; stating this explicitly avoids ambiguity about the exploratory nature of these analyses.
- 6.LOWdata codeProvide the full protocol and statistical analysis plan as supplementary material or a direct link.The protocol is referenced but not directly linked; making it available enhances transparency and reproducibility.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.