Neoadjuvant pembrolizumab, dabrafenib and trametinib in BRAF(V600)-mutant resectable melanoma: the randomized phase 2 NeoTrio trial.
Long GV, Carlino MS, Au-Yeung G, Spillane AJ, Shannon KF, Gyorki DE, Hsiao E, Kapoor R, Thompson JR, Batula I, Howle J, Ch'ng S, Gonzalez M, Saw RPM, Pennington TE, Lo SN, Scolyer RA, Menzies AM
- DOI
- 10.1038/s41591-024-03077-5
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/e4581a52-a41b-4970-a10a-fb0315137aec is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsOverstated claim−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×18−0.25★
- ReportingData & code availability partially met−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on pathological response rate, a surrogate endpoint. Although the paper discusses the correlation between pathological response and survival, it does not provide validated evidence linking pathological response to long-term clinical outcomes in this specific context, nor does it demonstrate target engagement at the tested dose. The paper itself notes that pathological response is a surrogate and that longer follow-up is needed.
“Pathological response is currently the best surrogate marker for survival following neoadjuvant therapy”
- 02Treatment effect not shown to be clinically meaningful
The primary reported effect is the pathological response rate, which is a surrogate. The effect sizes (e.g., 55%, 50%, 80%) are not anchored to a minimal clinically important difference or to a clear clinical benefit. The paper acknowledges that the trial was not powered for comparisons and that longer follow-up is needed, so the clinical meaningfulness of the effect sizes is not established.
“The NeoTrio trial was not powered to make statistical comparisons between the three arms, which limits the interpretation of the findings.”
- 03Printed percentage does not match its own count
36% does not match the reported count 7/20
“sensitivity (36%)”
Efficacy: responseFind in source - 04Printed percentage does not match its own count
92% does not match the reported count 18/20
“specificity (92%)”
Efficacy: responseFind in source - 05Printed percentage does not match its own count
22% does not match the reported count 4/20
“sensitivity (22%)”
Efficacy: responseFind in source - 06Printed percentage does not match its own count
92% does not match the reported count 18/20
“specificity (92%)”
Efficacy: responseFind in source
15 further findings of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper is a well-conducted randomized phase 2 trial with strong scientific premise, rigorous design, and transparent reporting. The main weaknesses are a vague data availability statement, lack of explicit reporting guideline adherence, and a concerning rate of statistical inconsistencies flagged by automated verification.
Both reviewers agreed on study type (interventional) and on all dimension statuses except statistical analysis, where the deterministic verification provided additional evidence that led to a downgrade from pass to warn. The statistics verification covered only a subset of tests (those with test statistics + df or effect estimates + CI); threshold-only p-values and exact tests were not machine-verifiable. The copyedit pass flagged minor internal inconsistencies in Table 3 and a typo in the funding acknowledgement.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks. 18 printed percentages that do not match their own count.
- PERCENT36% does not match the reported count 7/20
“sensitivity (36%)”
Efficacy: responseFind in source - PERCENT92% does not match the reported count 18/20
“specificity (92%)”
Efficacy: responseFind in source - PERCENT22% does not match the reported count 4/20
“sensitivity (22%)”
Efficacy: responseFind in source - PERCENT92% does not match the reported count 18/20
“specificity (92%)”
Efficacy: responseFind in source - PERCENT91% does not match the reported count 18/20
“specificity (91%)”
Efficacy: responseFind in source - PERCENT71% does not match the reported count 14/20
“80% and 71% with concurrent treatment”
Efficacy: recurrence and survivalFind in source - PERCENT89% does not match the reported count 18/20
“89% and 66% with pembrolizumab alone”
Efficacy: recurrence and survivalFind in source - PERCENT66% does not match the reported count 13/20
“89% and 66% with pembrolizumab alone”
Efficacy: recurrence and survivalFind in source - PERCENT84% does not match the reported count 17/20
“84% and 75% with concurrent treatment”
Efficacy: recurrence and survivalFind in source - PERCENT96% does not match the reported count 19/20
“96% and 96% for MPR”
Efficacy: recurrence and survivalFind in source - PERCENT96% does not match the reported count 19/20
“96% and 96% for MPR”
Efficacy: recurrence and survivalFind in source - PERCENT89% does not match the reported count 18/20
“89% and 89%, respectively, for RECIST CR”
Efficacy: recurrence and survivalFind in source - PERCENT89% does not match the reported count 18/20
“89% and 89%, respectively, for RECIST CR”
Efficacy: recurrence and survivalFind in source - PERCENT33% does not match the reported count 7/20
“100% and 33% for EORTC CMR”
Efficacy: recurrence and survivalFind in source - PERCENT89% does not match the reported count 18/20
“89% and 79% for EORTC PMR”
Efficacy: recurrence and survivalFind in source - PERCENT79% does not match the reported count 16/20
“89% and 79% for EORTC PMR”
Efficacy: recurrence and survivalFind in source - PERCENT76% does not match the reported count 15/20
“95% and 76% with pembrolizumab alone”
Efficacy: recurrence and survivalFind in source - PERCENT89% does not match the reported count 18/20
“95% and 89% with sequential treatment”
Efficacy: recurrence and survivalFind in source
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Check p-value for HRQOL GHS change in pembrolizumab arm at week 12
“Global Health Score (GHS) mean change: −17.70 (95% CI, −25.10 to −10.30), P < 0.001”
Taken as given: The CI is a 95% confidence interval.; The sample size for the pembrolizumab arm is 20, so df=19.; The t-statistic is approximated as the mean divided by the standard error, where SE is derived from the CI width.Method: Approximate t-statistic from CI and compute two-tailed p-value using t-distribution.How we recomputed it: pT(-17.70/ ( (25.10-10.30)/ (2*1.96) ), 19)
- lowinternal contradictionIn the text, it says 'including five CMR in the concurrent arm' but Table 3 shows CMR for concurrent arm as 3 (15%). This discrepancy may be due to different denominators or a typo.
“including five CMR in the concurrent arm”
ResultsFind in source - lowinternal contradictionIn Table 3, the PR row for pembrolizumab arm shows 3 (25) but 3/20 is 15%, not 25%. This may be a typographical error.
“PR | 3 (25) | 10 (50) | 7 (35)”
Table 3Find in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions overstated beyond the evidenceAssessed
3 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated).
- overstatedReviewers 1, 2Immunotherapy and targeted therapy should not be combined in the neoadjuvant setting for melanoma.The conclusion is based on a small, non-comparative trial with short follow-up; the authors themselves note the trial was not powered for comparisons and longer follow-up is needed, so the claim is somewhat strong.Evidence: The paper presents higher toxicity and a suggestion of reduced durability of response in targeted therapy arms, but the trial was not designed to compare arms.
“Pending longer follow-up, we suggest that immunotherapy and targeted therapy should not be combined in the neoadjuvant setting for melanoma.”
DiscussionFind in source - supportedReviewers 1, 2Concurrent therapy met the primary endpoint with a pathological response rate of 80%.The reported rate of 16/20 (80%) with 95% CI 60-97% supports the claim that the primary endpoint was met.Evidence: Table 3 shows pathological response rate of 80% (16/20) for concurrent therapy.
“The pathological response rate was 55% (11/20; including six pathological complete responses (pCRs)) with pembrolizumab, 50% (10/20; three pCRs) with sequential therapy and 80% (16/20; ten pCRs) with concurrent therapy, which met the primary outcome in each arm.”
AbstractFind in source - supportedReviewers 1, 2Recurrences after major pathological response were more common in the targeted therapy arms.The paper reports that no patient with MPR in the pembrolizumab arm recurred, while one each in the targeted therapy arms did, supporting the claim.Evidence: Results section states: 'In the pembrolizumab monotherapy arm, no patient with MPR recurred. One patient with MPR recurred in each of the targeted therapy arms.'
In the pembrolizumab monotherapy arm, no patient with MPR recurred. One patient with MPR recurred in each of the targeted therapy arms.
Resultsreviewer’s wording
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on pathological response rate, a surrogate endpoint. Although the paper discusses the correlation between pathological response and survival, it does not provide validated evidence linking pathological response to long-term clinical outcomes in this specific context, nor does it demonstrate target engagement at the tested dose. The paper itself notes that pathological response is a surrogate and that longer follow-up is needed.
“Pathological response is currently the best surrogate marker for survival following neoadjuvant therapy”
- INADEQUATEEffect sizeThe primary reported effect is the pathological response rate, which is a surrogate. The effect sizes (e.g., 55%, 50%, 80%) are not anchored to a minimal clinically important difference or to a clear clinical benefit. The paper acknowledges that the trial was not powered for comparisons and that longer follow-up is needed, so the clinical meaningfulness of the effect sizes is not established.
“The NeoTrio trial was not powered to make statistical comparisons between the three arms, which limits the interpretation of the findings.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites multiple studies (Keynote-022, COMBI-i, IMspire150, SECOMBIT, DREAMseq, SWOG S1801) and discusses the mechanisms by which targeted therapy may enhance immunotherapy. It also acknowledges limitations of prior work, such as the lack of benefit of neoadjuvant targeted therapy alone and the need for better combinations. The hypothesis follows logically from the cited evidence.
“Immune changes early during targeted therapy suggest the mechanisms of each drug class could work synergistically.”
“However, the most effective way to combine these treatments is unknown.”
Randomization method is described (web-based, permuted blocks, stratified by BRAF mutation subtype). Blinding is not applicable as it is an open-label trial, but this is stated. Power analysis is provided for the single-arm design. Inclusion/exclusion criteria are detailed. Outlier handling is not explicitly described, but the analysis population (ITT) is defined. Controls are not applicable as this is a non-comparative trial with three active arms. Independent replication is not applicable for a single trial.
Sex is reported (42% female). Age is reported (median 53). Demographics include ECOG status and BRAF mutation subtype. Species/strain and housing are not applicable for a human trial. Sex justification is not applicable as both sexes are enrolled.
The study states it was performed in accordance with the Declaration of Helsinki and Good Clinical Practice, and that the protocol was approved by the human research ethics committee at each participating institution. Written informed consent was obtained. Regulatory compliance is stated.
“The study was performed in accordance with the Declaration of Helsinki and Good Clinical Practice guidelines. Patients provided written informed consent. The study protocol was approved by the human research ethics committee at each participating institution.”
“The study protocol was approved by the human research ethics committee at each participating institution.”
“Patients provided written informed consent.”
“The study was performed in accordance with the Declaration of Helsinki and Good Clinical Practice guidelines.”
The investigational products (pembrolizumab, dabrafenib, trametinib) are named with manufacturers and dosing regimens. Statistical software (SAS 9.4, R 4.1.3) is identified. Antibodies, cell lines, and mycoplasma testing are not applicable as this is a clinical trial without wet-lab assays.
“Study drugs were supplied by Merck Sharp & Dohme (pembrolizumab) and Novartis (dabrafenib and trametinib).”
“All statistical analyses were performed using SAS (version 9.4) and R (version 4.1.3).”
“Study drugs were supplied by Merck Sharp & Dohme (pembrolizumab) and Novartis (dabrafenib and trametinib).”
“All statistical analyses were performed using SAS (version 9.4) and R (version 4.1.3).”
Statistical tests are named (Clopper-Pearson exact CIs, Kaplan-Meier, mixed linear modeling). Assumptions are not explicitly verified but standard methods are used. Exact p-values are reported for HRQOL (e.g., P < 0.05, P < 0.001). Effect sizes with CIs are reported for response rates and HRQOL changes. Software is identified. Data presentation includes per-group n and CIs. Mathematical plausibility checks were performed and no issues found.
“The primary and secondary response outcomes were summarized using frequency and proportion by arm along with the two-sided 95% Clopper–Pearson exact CIs.”
“Global Health Score (GHS) mean change: −10.83 (95% CI, −19.63 to −2.01), P < 0.05”
“A pathological response was achieved in 55% (11/20; 95% CI, 36–83) of patients in the pembrolizumab arm”
“The primary and secondary response outcomes were summarized using frequency and proportion by arm along with the two-sided 95% Clopper–Pearson exact CIs.”
“Global Health Score (GHS) mean change: −10.83 (95% CI, −19.63 to −2.01), P < 0.05”
The data availability statement says 'De-identified data are available on reasonable request and after signing of a data transfer agreement with Melanoma Institute Australia.' This is a managed access route but lacks specific conditions or a timeframe. No repository deposit or accession numbers are provided, which is acceptable for patient-level data. Code sharing is not applicable as no custom code is mentioned.
“De-identified data are available on reasonable request and after signing of a data transfer agreement with Melanoma Institute Australia.”
“De-identified data are available on reasonable request and after signing of a data transfer agreement with Melanoma Institute Australia.”
The trial is registered (NCT02858921). Methods are detailed enough for replication. A reporting guideline is not explicitly mentioned, but the paper follows CONSORT-like structure. All pre-specified outcomes are reported, including negative results. Limitations are discussed. Conclusions are proportional to the evidence. Funding and COI are disclosed.
“ClinicalTrials.gov registration: NCT02858921”
“The NeoTrio trial was not powered to make statistical comparisons between the three arms, which limits the interpretation of the findings.”
“ClinicalTrials.gov registration: NCT02858921”
“The NeoTrio trial was not powered to make statistical comparisons between the three arms, which limits the interpretation of the findings.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 48 references by DOI: 46 verified — 2 no DOI (shown, not verified).
- NO DOIMelanoma of the SkinNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMeasuring quality of life in patients with melanoma: development of the FACT-melanoma subscaleNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://clinicaltrials.gov/study/NCT02858921LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT02858921LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORconsistencyResults, Efficacy: response“A metabolic response ... was observed in 40% (8/20) of patients treated with pembrolizumab, 50% (10/20) of those treated with sequential therapy and 95% (19/20) of those treated with concurrent therapy”→ Check the percentage for concurrent therapy: 19/20 is 95%, but the text says 95% (19/20) which is correct, but earlier in the same paragraph it says 'including five CMR in the concurrent arm' which may be inconsistent with Table 3 showing 3 CMR in concurrent arm.Potential inconsistency in CMR counts between text and table.
- MINORtypoAbstract“NMHRC”→ Should be NHMRC (National Health and Medical Research Council).Typo in funding acknowledgement.
- MINORclarityResults, Efficacy: response“PR | 3 (25) | 10 (50) | 7 (35)”→ The percentage for pembrolizumab arm PR is 3/20=15%, not 25%. Check if this is a typo.Potential arithmetic error in Table 3.
- MINORconsistencyResults, Efficacy: response“A metabolic response ... was observed in 40% (8/20) of patients treated with pembrolizumab, 50% (10/20) of those treated with sequential therapy and 95% (19/20) of those treated with concurrent therapy”→ Check the percentage for concurrent therapy: 19/20 is 95%, but the text says 95% (19/20) which is correct; however, earlier in the same paragraph it says 'including five CMR in the concurrent arm' but the table shows 3 CMR in concurrent arm. Verify consistency.Potential inconsistency in CMR counts between text and table.
- MINORclarityResults, Efficacy: response“PR | 3 (25) | 10 (50) | 7 (35)”→ The percentage for PR in pembrolizumab arm is 25% but 3/20 is 15%. Verify the denominator.Potential arithmetic inconsistency in Table 3.
The published work is generally robust, but readers should weigh the high rate of statistical inconsistencies flagged by automated verification and the internal inconsistencies in Table 3. An erratum or independent re-analysis may be warranted to confirm the reported response rates and HRQOL findings.
- 1.HIGHstatisticsInvestigate the 18 of 19 recomputable tests that were flagged as inconsistent by the automated verification; re-check the reported test statistics, degrees of freedom, and effect estimates against the raw data, and correct any errors or clarify the discrepancies in a correction or erratum.A high rate of recomputation inconsistencies is a validity threat that could undermine confidence in the reported results.
- 2.HIGHcopyeditCorrect the internal inconsistency in Table 3: the PR row for the pembrolizumab arm shows 3 (25) but 3/20 is 15%, not 25%; verify the denominator and correct the percentage.An arithmetic error in a key efficacy table is a concrete reporting error that readers will notice.
- 3.HIGHcopyeditResolve the discrepancy between the text stating 'including five CMR in the concurrent arm' and Table 3 showing 3 CMR (15%) for the concurrent arm; correct the text or the table to be consistent.Internal contradictions between text and tables undermine the paper's credibility.
- 4.HIGHreportingExplicitly state adherence to CONSORT reporting guidelines in the Methods or a separate Reporting Summary section.The paper follows CONSORT-like structure but does not explicitly cite the guideline, which is a reporting transparency gap.
- 5.HIGHdata codeImprove the data availability statement by specifying a data access committee or platform (e.g., Vivli), the conditions for access, and a timeframe for response.The current statement is vague and lacks a concrete mechanism, which limits reproducibility.
- 6.MEDIUMcopyeditFix the typo 'NMHRC' to 'NHMRC' in the funding acknowledgement.A misspelled funding body name is a minor but easily correctable error.
- 7.MEDIUMstatisticsReport exact p-values for all secondary endpoints, not just HRQOL, where applicable.Threshold-only p-values (e.g., 'P < 0.05') are imprecise and hinder verification.
- 8.MEDIUMstatisticsAdd a statement on how outliers were handled in the statistical analysis, even if none were excluded.Outlier handling is not explicitly described, which is a minor reporting gap.
- 9.MEDIUMstatisticsClarify the assumptions of the statistical models used, particularly for the mixed linear modeling of HRQOL data.Assumption verification is not explicitly reported, which is a minor gap.
- 10.LOWreportingAdd a note on the open-label design rationale in the Methods to strengthen blinding justification.The open-label design is stated but not justified, which is a minor transparency issue.
- 11.LOWdata codeInclude a statement on the absence of custom code or provide code if used for analyses.Clarifying code availability is a nice-to-have for reproducibility.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.