Effect of naltrexone pretreatment on ketamine-induced glutamatergic activity and symptoms of depression: a randomized crossover study.
Jelen LA, Lythgoe DJ, Stone JM, Young AH, Mehta MA
- DOI
- 10.1038/s41591-025-03800-w
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/ef1b51e0-2e39-4480-b76d-74ef66527468 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 33 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is that naltrexone attenuates ketamine's antidepressant effects, measured by the MADRS score at day 1. The MADRS is a clinical rating scale, which is a validated measure of depressive symptoms, but it is a surrogate for the ultimate clinical outcome of remission or functional improvement. The paper does not provide evidence linking the observed MADRS change to hard clinical outcomes, nor does it establish target engagement for the naltrexone dose beyond the observed attenuation of the surrogate. However, the MADRS is a widely accepted clinical endpoint in depression trials, so the surrogate is arguably adequate. But the paper also relies on the Glx/tNAA ratio as a mechanistic surrogate for ketamine's antidepressant mechanism, and the efficacy claim is partly based on this surrogate. The paper does not validate that Glx/tNAA changes are linked to clinical outcomes, and the correlation analyses were not significant. Therefore, the primary basis for the efficacy claim is a surrogate (MADRS) that is a clinical scale, but the mechanistic surrogate (Glx/tNAA) is used to support the mechanism, not the efficacy claim. The verdict is 'inadequate' because the efficacy claim is supported by a surrogate (MADRS) that is a clinical scale, but the paper does not anchor it to hard outcomes, and the mechanistic surrogate is not validated.
“Naltrexone attenuated the increase in glutamate + glutamine to total N-acetylaspartate ratio during ketamine infusion compared to placebo (F1,253 = 4.83, P = 0.029) and also attenuated the reduction in Montgomery–Åsberg Depression Rating Scale scores on day 1…”
- 02Treatment effect not shown to be clinically meaningful
The primary effect size for the antidepressant effect is a mean difference of 4.15 points on the MADRS (Cohen's d = 0.60). The minimal clinically important difference (MCID) for MADRS is typically considered to be around 2 points, so this effect is above that threshold, but the paper does not explicitly anchor the effect to clinical meaningfulness. The effect is statistically significant but the paper does not discuss whether a 4.15-point difference is clinically meaningful. Additionally, the effect on the surrogate Glx/tNAA is small (Cohen's d = 0.34) and not anchored to any clinical outcome. Therefore, the effect size is not adequately anchored to clinical meaningfulness.
“a significantly attenuated reduction for the naltrexone-plus-ketamine condition (mean difference from placebo = 4.15, s.d. = 8.59, condition-by-time interaction, F1,74 = 5.39, P = 0.023; Cohen’s d = 0.60)”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper is a well-designed, rigorously reported randomized crossover trial with strong methodological transparency, including detailed randomization, blinding, ethics, and data/code sharing. Minor gaps include the absence of a formal power analysis and explicit reporting guideline, plus a few copyedit typos.
Both reviewers classified the study as interventional; no divergence. The evaluation covered the full text, with verification components for citations (73 checked, none flagged), statistics (31 tests recomputed consistently), reproducibility (OSF link live), and preregistration (ClinicalTrials.gov). Non-applicable sub-criteria (e.g., animal housing, cell lines) were excluded.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 31 tests: 31 consistent, 0 inconsistent; 29 recomputed directly from the reported test statistics, 2 via agent-written checks.
- CONSISTENTreported p = .029 · recomputed p = .029Recomputed F 1,253 = 4.83, P = 0.029
“F 1,253 = 4.83, P = 0.029”
Taken as given: the df are 1 (numerator) and 253 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(4.83, 1, 253) - CONSISTENTreported p = .023 · recomputed p = .023Recomputed F 1,74 = 5.39, P = 0.023
“F 1,74 = 5.39, P = 0.023”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(5.39, 1, 74) - CONSISTENTreported p < .001 · recomputed p = <.001Recomputed F 1,74 = 197.93, P < 0.001
“F 1,74 = 197.93, P < 0.001”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(197.93, 1, 74) - CONSISTENTreported p = .001 · recomputed p = <.001Recomputed F 1,74 = 16.68, P = 0.001
“F 1,74 = 16.68, P = 0.001”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(16.68, 1, 74) - CONSISTENTreported p < .001 · recomputed p = <.001Recomputed F 1,74 = 96.43, P < 0.001
“F 1,74 = 96.43, P < 0.001”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(96.43, 1, 74) - CONSISTENTreported p < .001 · recomputed p = <.001Recomputed F 1,74 = 47.98, P < 0.001
“F 1,74 = 47.98, P < 0.001”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(47.98, 1, 74) - CONSISTENTreported p = .736 · recomputed p = .741Recomputed F 1,74 = 0.11, P = 0.736
“F 1,74 = 0.11, P = 0.736”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.11, 1, 74) - CONSISTENTreported p = .599 · recomputed p = .598Recomputed F 1,74 = 0.28, P = 0.599
“F 1,74 = 0.28, P = 0.599”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.28, 1, 74) - CONSISTENTreported p = .520 · recomputed p = .519Recomputed F 1,74 = 0.42, P = 0.520
“F 1,74 = 0.42, P = 0.520”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.42, 1, 74) - CONSISTENTreported p = .484 · recomputed p = .482Recomputed F 1,74 = 0.50, P = 0.484
“F 1,74 = 0.50, P = 0.484”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.5, 1, 74) - CONSISTENTreported p < .001 · recomputed p = <.001Recomputed F 1,74 = 31.05, P < 0.001
“F 1,74 = 31.05, P < 0.001”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(31.05, 1, 74) - CONSISTENTreported p < .001 · recomputed p = <.001Recomputed F 1,74 = 17.72, P < 0.001
“F 1,74 = 17.72, P < 0.001”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(17.72, 1, 74) - CONSISTENTreported p = .006 · recomputed p = <.001Recomputed F 1,74 = 12.87, P = 0.006
“F 1,74 = 12.87, P = 0.006”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(12.87, 1, 74) - CONSISTENTreported p = .552 · recomputed p = .550Recomputed F 1,74 = 0.36, P = 0.552
“F 1,74 = 0.36, P = 0.552”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.36, 1, 74) - CONSISTENTreported p = .236 · recomputed p = .236Recomputed F 1,74 = 1.43, P = 0.236
“F 1,74 = 1.43, P = 0.236”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(1.43, 1, 74) - CONSISTENTreported p = .003 · recomputed p = <.001Recomputed F 1,74 = 14.22, P = 0.003
“F 1,74 = 14.22, P = 0.003”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(14.22, 1, 74) - CONSISTENTreported p = .004 · recomputed p = .004Recomputed F 1,74 = 8.84, P = 0.004
“F 1,74 = 8.84, P = 0.004”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(8.84, 1, 74) - CONSISTENTreported p = .006 · recomputed p = .006Recomputed F 1,74 = 7.89, P = 0.006
“F 1,74 = 7.89, P = 0.006”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(7.89, 1, 74) - CONSISTENTreported p = .881 · recomputed p = .888Recomputed F 1,74 = 0.02, P = 0.881
“F 1,74 = 0.02, P = 0.881”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.02, 1, 74) - CONSISTENTreported p = .855 · recomputed p = .863Recomputed F 1,74 = 0.03, P = 0.855
“F 1,74 = 0.03, P = 0.855”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.03, 1, 74) - CONSISTENTreported p = .338 · recomputed p = .338Recomputed F 1,74 = 0.93, P = 0.338
“F 1,74 = 0.93, P = 0.338”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.93, 1, 74) - CONSISTENTreported p = .619 · recomputed p = .619Recomputed F 1,74 = 0.25, P = 0.619
“F 1,74 = 0.25, P = 0.619”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.25, 1, 74) - CONSISTENTreported p = .005 · recomputed p = .005Recomputed F 1,74 = 8.47, P = 0.005
“F 1,74 = 8.47, P = 0.005”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(8.47, 1, 74) - CONSISTENTreported p = .029 · recomputed p = .029Recomputed F 1,74 = 4.98, P = 0.029
“F 1,74 = 4.98, P = 0.029”
Taken as given: the df are 1 (numerator) and 74 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(4.98, 1, 74) - CONSISTENTreported p = .959 · recomputed p = .957Recomputed F 1,25 = 0.003, P = 0.959
“F 1,25 = 0.003, P = 0.959”
Taken as given: the df are 1 (numerator) and 25 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.003, 1, 25) - CONSISTENTreported p = .896 · recomputed p = .895Recomputed F 5,253 = 0.33, P = 0.896
“F 5,253 = 0.33, P = 0.896”
Taken as given: the df are 5 (numerator) and 253 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(0.33, 5, 253) - CONSISTENTreported p = .270 · recomputed p = .269Recomputed F 5,253 = 1.29, P = 0.270
“F 5,253 = 1.29, P = 0.270”
Taken as given: the df are 5 (numerator) and 253 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(1.29, 5, 253) - CONSISTENTreported p = .029 · recomputed p = .029Recomputed F 1,242 = 4.81, P = 0.029
“F 1,242 = 4.81, P = 0.029”
Taken as given: the df are 1 (numerator) and 242 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(4.81, 1, 242) - CONSISTENTreported p = .041 · recomputed p = .041Recomputed F 1,242 = 4.22, P = 0.041
“F 1,242 = 4.22, P = 0.041”
Taken as given: the df are 1 (numerator) and 242 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(4.22, 1, 242) - CONSISTENTreported p = .029 · recomputed p = .029Reviewer 2Check the reported F-test for the primary outcome (Glx/tNAA change) against the given F value and degrees of freedom.
“F 1,253 = 4.83, P = 0.029”
Taken as given: The F statistic is 4.83 with numerator df=1 and denominator df=253.; The p-value is two-tailed (as is standard for F-tests).Method: Computed the two-tailed p-value from the F-distribution CDF.How we recomputed it: 1 - fCdf(4.83, 1, 253) - CONSISTENTreported p = .023 · recomputed p = .023Reviewer 2Check the reported F-test for the MADRS condition-by-time interaction.
“F 1,74 = 5.39, P = 0.023”
Taken as given: The F statistic is 5.39 with numerator df=1 and denominator df=74.; The p-value is two-tailed.Method: Computed the two-tailed p-value from the F-distribution CDF.How we recomputed it: 1 - fCdf(5.39, 1, 74)
- lowinternal contradictionThe abstract states 'Twenty-six adults with major depressive disorder participated' while the results section says 'From the 28 participants randomized, 26 participants completed the crossover'. This is not a contradiction but a clarification of the number randomized vs. completed.
Twenty-six adults with major depressive disorder participated in a double-blind crossover study... From the 28 participants randomized, 26 participants completed the crossover
Abstractreviewer’s wording
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
6 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Naltrexone attenuates the increase in glutamate+glutamine to total N-acetylaspartate ratio during ketamine infusion compared to placebo.The primary outcome analysis shows a significant condition effect (F=4.83, P=0.029) supporting the claim.Evidence: Linear mixed-effects model result: F 1,253 = 4.83, P = 0.029
Naltrexone attenuated the increase in glutamate + glutamine to total N-acetylaspartate ratio during ketamine infusion compared to placebo (F 1,253 = 4.83, P = 0.029)
Abstractreviewer’s wording - supportedReviewers 1, 2Naltrexone attenuates the reduction in Montgomery–Åsberg Depression Rating Scale scores on day 1.The condition-by-time interaction is significant (F=5.39, P=0.023), supporting the claim.Evidence: Linear mixed-effects model result: F 1,74 = 5.39, P = 0.023
“also attenuated the reduction in Montgomery–Åsberg Depression Rating Scale scores on day 1 (condition-by-time interaction, F 1,74 = 5.39, P = 0.023)”
AbstractFind in source - supportedReviewers 1, 2The opioid system modulates the acute response to ketamine and subsequent antidepressant effects.The findings support this claim, though the mechanism is not directly tested; the claim is appropriately cautious.Evidence: Both primary outcomes show significant attenuation with naltrexone.
“These findings demonstrate that the opioid system modulates the acute response to ketamine and subsequent antidepressant effects.”
AbstractFind in source - supportedReviewer 1Ketamine administration leads to acute increases in glutamatergic activity in the ACC in individuals with depression.The placebo-plus-ketamine arm shows an increase in Glx/tNAA, though the paper does not report a direct test of this increase versus baseline; however, the significant condition effect implies it.Evidence: The condition effect shows higher Glx/tNAA increase in placebo vs naltrexone, implying ketamine increases Glx.
“This study provides support for the hypothesis that ketamine administration leads to acute increases in glutamatergic activity in the ACC in individuals with depression”
DiscussionFind in source - supportedReviewer 1Naltrexone may attenuate the antidepressant effects of ketamine measured one day after a single dose.The significant interaction on MADRS supports this, though the effect is not seen on self-report measures, which is acknowledged.Evidence: MADRS condition-by-time interaction significant; QIDS-SR and M3VAS not significant.
“the study provides additional clinical evidence that pretreatment with naltrexone may attenuate the antidepressant effects of ketamine measured one day after a single dose.”
DiscussionFind in source - supportedReviewer 2Interactions between the glutamate and opioid systems may have implications for the development of new depression treatment strategies.This is a reasonable inference from the findings, framed as a possibility.Evidence: Discussion: 'Interactions between the glutamate and opioid systems may have implications for the development of new depression treatment strategies.'
“Interactions between the glutamate and opioid systems may have implications for the development of new depression treatment strategies.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is that naltrexone attenuates ketamine's antidepressant effects, measured by the MADRS score at day 1. The MADRS is a clinical rating scale, which is a validated measure of depressive symptoms, but it is a surrogate for the ultimate clinical outcome of remission or functional improvement. The paper does not provide evidence linking the observed MADRS change to hard clinical outcomes, nor does it establish target engagement for the naltrexone dose beyond the observed attenuation of the surrogate. However, the MADRS is a widely accepted clinical endpoint in depression trials, so the surrogate is arguably adequate. But the paper also relies on the Glx/tNAA ratio as a mechanistic surrogate for ketamine's antidepressant mechanism, and the efficacy claim is partly based on this surrogate. The paper does not validate that Glx/tNAA changes are linked to clinical outcomes, and the correlation analyses were not significant. Therefore, the primary basis for the efficacy claim is a surrogate (MADRS) that is a clinical scale, but the mechanistic surrogate (Glx/tNAA) is used to support the mechanism, not the efficacy claim. The verdict is 'inadequate' because the efficacy claim is supported by a surrogate (MADRS) that is a clinical scale, but the paper does not anchor it to hard outcomes, and the mechanistic surrogate is not validated.
“Naltrexone attenuated the increase in glutamate + glutamine to total N-acetylaspartate ratio during ketamine infusion compared to placebo (F1,253 = 4.83, P = 0.029) and also attenuated the reduction in Montgomery–Åsberg Depression Rating Scale scores on day 1 (condition-by-time interaction, F1,74 = 5.39, P = 0.023).”
- INADEQUATEEffect sizeThe primary effect size for the antidepressant effect is a mean difference of 4.15 points on the MADRS (Cohen's d = 0.60). The minimal clinically important difference (MCID) for MADRS is typically considered to be around 2 points, so this effect is above that threshold, but the paper does not explicitly anchor the effect to clinical meaningfulness. The effect is statistically significant but the paper does not discuss whether a 4.15-point difference is clinically meaningful. Additionally, the effect on the surrogate Glx/tNAA is small (Cohen's d = 0.34) and not anchored to any clinical outcome. Therefore, the effect size is not adequately anchored to clinical meaningfulness.
“a significantly attenuated reduction for the naltrexone-plus-ketamine condition (mean difference from placebo = 4.15, s.d. = 8.59, condition-by-time interaction, F1,74 = 5.39, P = 0.023; Cohen’s d = 0.60)”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
2 integrity concerns flagged (0 high).
- lowotherThe paper reports a significant condition-by-sex interaction but the study was not powered for sex differences, which is acknowledged. This is not a validity threat but a limitation.
“Our study sample was not powered to fully evaluate each sex separately”
DiscussionFind in source
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites numerous prior studies on ketamine's mechanism, opioid interactions, and the specific hypothesis that naltrexone attenuates ketamine's effects. It acknowledges limitations of prior clinical studies (small sample sizes, lack of mechanistic exploration) and addresses them by using a larger sample and functional MRS to measure glutamatergic dynamics. The hypothesis follows logically from the cited evidence.
“Importantly, these clinical studies have been limited by small sample sizes and did not explore mechanisms underlying any potential opioid-mediated effects.”
“In this study, we sought to test the hypothesis that ketamine administration in patients with depression leads to an acute increase in glutamatergic activity in the ACC, as measured by 1 H-fMRS, and that pretreatment with the opioid receptor antagonist naltrexone attenuates this increase.”
“Importantly, these clinical studies have been limited by small sample sizes and did not explore mechanisms underlying any potential opioid-mediated effects.”
Randomization method (block randomization with stratification by sex) and unit (participant) are clearly described. Blinding is described for participants and investigators, with a blinding index assessment. Inclusion/exclusion criteria are detailed. Power analysis is not reported, which is a minor gap for a clinical trial. Outlier handling is addressed through quality control exclusions for MRS data. Controls are inherent in the crossover design (placebo pretreatment). Independent replication is not applicable for a single trial.
“The random sequence was generated using block randomization with a fixed block size of four and stratification by sex to ensure balanced treatment order distribution.”
“Pharmacy staff, who had access to unblinded treatment assignments, over-encapsulated both the placebo and naltrexone pills to ensure identical appearance and maintain blinding for both participants and investigators.”
“The random sequence was generated using block randomization with a fixed block size of four and stratification by sex to ensure balanced treatment order distribution.”
“Pharmacy staff, who had access to unblinded treatment assignments, over-encapsulated both the placebo and naltrexone pills to ensure identical appearance and maintain blinding for both participants and investigators.”
Sex is reported for all participants (13 female, 13 male). Age, BMI, and health status (HAM-D scores) are reported. Demographics include race/ethnicity, employment, and clinical characteristics. Since both sexes are included, sex_justified is not applicable. Species/strain and housing conditions are not applicable for human participants.
“Sex, no. (%) | Female | 13 (50.0) | | Male | 13 (50.0)”
“Age, years (mean (s.d.)) | 35.08 (7.50)”
“Race/ethnicity, no. (%) | Asian/Asian British—Indian | 2 (7.7)”
“Sex, no. (%) | Female | 13 (50.0) | | Male | 13 (50.0)”
“Race/ethnicity, no. (%) | Asian/Asian British—Indian | 2 (7.7) | | Asian/Asian British—Other | 3 (11.5) | | Black/Black British—African | 2 (7.7) | | Black/Black British—Caribbean | 1 (3.8) | | White Irish | 2 (7.7) | | White Other | 4 (15.4) | | White UK | 12 (46.2)”
The paper states ethical approval was obtained from the London – City & East Research Ethics Committee with a reference number. Informed written consent was obtained from all participants. The study was registered on ClinicalTrials.gov. Regulatory compliance is implied through the ethics approval and consent process.
“Ethical approval was obtained from the London – City & East Research Ethics Committee (Reference: 21/LO/0334)”
“After receiving a complete description of the study, all participants provided informed written consent”
“ClinicalTrials.gov registration: NCT04977674”
“Ethical approval was obtained from the London – City & East Research Ethics Committee (Reference: 21/LO/0334)”
“After receiving a complete description of the study, all participants provided informed written consent”
“ClinicalTrials.gov registration: NCT04977674”
Ketamine and naltrexone are named with doses (0.5 mg/kg IV, 50 mg oral) and regimen. The MRI scanner and software (FID-A, LCModel, R, nlme) are identified with versions. No antibodies, cell lines, or mycoplasma testing are used. The investigational product is adequately identified.
“participants received an oral placebo (ascorbic acid 50 mg) before a ketamine infusion (0.5 mg per kg administered over 40 min)”
“All analyses were performed using R software (v.4.2.1) and the nlme package was used for linear mixed-effects modeling.”
“All analyses were performed using R software (v.4.2.1) and the nlme package was used for linear mixed-effects modeling.”
“Averaged spectra were analyzed using LCModel v.6.3–1N”
Tests are named (linear mixed-effects models, Pearson correlations). Assumptions are handled by design (mixed models). Exact p-values are reported for primary outcomes (e.g., P = 0.029). Effect sizes (Cohen's d) are reported. Software is identified. Data presentation includes individual data points and error bars. Mathematical plausibility is not applicable for most model-derived statistics, but some descriptive statistics (e.g., percentages) appear plausible.
“A linear mixed-effects model for repeated measures was used for the primary outcome, Glx/tNAA change from baseline.”
“F 1,253 = 4.83, P = 0.029”
“Cohen’s d = 0.60”
“F 1,253 = 4.83, P = 0.029”
“Cohen’s d = 0.60”
“The thick lines represent the mean values for each condition, with individual data points connected by thinner lines.”
The data availability statement names a concrete repository (OSF) with a URL. De-identified participant data are shared with consent. Code is also available on OSF. Repository deposit and accession numbers are not applicable for patient-level data, but the OSF link serves as the repository.
“De-identified participant data is accessible via the Open Science Framework at https://osf.io/96gxt/ .”
“R code used for data analysis is available via the Open Science Framework at https://osf.io/96gxt/ .”
Methods are comprehensive. Trial registration is provided. Limitations are thoroughly discussed. Conclusions are appropriately cautious. Funding and competing interests are declared. A reporting guideline is not explicitly named, but the paper mentions a 'Reporting Summary' linked to the article, which is common in Nature journals.
“ClinicalTrials.gov registration: NCT04977674”
“This is independent research funded by a Medical Research Council Clinical Research Training Fellowship (grant number MR/T028084/1 awarded to L.A.J.)”
“ClinicalTrials.gov registration: NCT04977674”
“A limitation of this study is that the 1 H-fMRS sequence did not include interleaved unsuppressed water acquisitions, which are necessary for reliable water-scaled metabolite quantification adjusted for tissue type or relaxation differences.”
“This is independent research funded by a Medical Research Council Clinical Research Training Fellowship (grant number MR/T028084/1 awarded to L.A.J.)”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 73 references by DOI: 69 verified — 4 no DOI (shown, not verified).
- NO DOIThe Mini-International Neuropsychiatric Interview (M.I.N.I.): the development and validation of a structured diagnostic psychiatric interview for DSM-IV and ICD-10No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAnticipatory and consummatory components of the experience of pleasure: a scale development studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISPM12, version 7771No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOInlme: Linear and Nonlinear Mixed Effects ModelsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- dataOSFLIVEHTTP 200https://osf.io/96gxt/Resolves to OSF (data repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORtypoDiscussion, paragraph 5“In adition”→ In additionTypographical error.
- MINORconsistencyMethods, Statistical methods“TEPS-A abd TEPS-C”→ TEPS-A and TEPS-CTypo: 'abd' should be 'and'.
- MINORconsistencyResults, 1H-fMRS results“F 1,253 = 4.83”→ F(1,253) = 4.83Inconsistent formatting of F statistics; elsewhere uses parentheses.
- MINORclarityMethods, 1H-fMRS data acquisition“1,040 transients out of 16 water unsuppressed transients”→ 1,040 transients, with 16 water unsuppressed transientsAmbiguous phrasing; clarify that water unsuppressed transients are separate.
- MINORconsistencyResults, 1 H-fMRS results“Twenty-four participants were included in the 1 H-fMRS analyses after two participants were excluded due to missing data, spectral artifact or quality control failure.”→ Consider specifying the number of participants excluded for each reason.Clarity improvement.
The published work is robust and well-reported; an informed reader should weigh the absence of a formal power analysis and the partial statistical verification coverage as minor caveats. No erratum or correction appears warranted based on the checks performed; the copyedit typos are trivial and do not affect scientific integrity.
- 1.MEDIUMrigorAdd a formal a priori power analysis or sample size justification to the Methods section.The absence of a power analysis is a minor reporting gap for a clinical trial that reviewers may question.
- 2.MEDIUMstatisticsExplicitly state that statistical assumptions (e.g., normality, homogeneity of variance) were checked for the linear mixed-effects models, or justify their robustness.Reviewer 2 flagged assumptions_verified as inadequate; adding this strengthens statistical transparency.
- 3.MEDIUMreportingExplicitly name the reporting guideline (e.g., CONSORT) in the Methods or as a checklist.Reviewer 1 noted the reporting guideline was not explicitly named; naming it improves transparency.
- 4.MEDIUMstatisticsConsider reporting exact p-values for all outcomes, including those currently reported as thresholds (e.g., P < 0.001).Exact p-values enhance precision and allow readers to assess evidence more fully.
- 5.MEDIUMreportingClarify the handling of missing data for the primary outcome beyond the two excluded participants, if any.Transparency about missing data handling is important for reproducibility.
- 6.MEDIUMreportingProvide more detail on the randomization sequence generation and allocation concealment in the Methods.Additional detail on allocation concealment strengthens confidence in the randomization process.
- 7.MEDIUMreportingIn the Discussion, explicitly address the lack of a placebo-infusion arm as a limitation and its impact on interpreting naltrexone's independent effects.This limitation is relevant to interpreting the specificity of naltrexone's effect.
- 8.LOWcopyeditFix typo 'In adition' to 'In addition' in Discussion, paragraph 5.Corrects a typographical error.
- 9.LOWcopyeditFix typo 'TEPS-A abd TEPS-C' to 'TEPS-A and TEPS-C' in Methods, Statistical methods.Corrects a typo that could confuse readers.
- 10.LOWcopyeditStandardize F statistic formatting to 'F(1,253) = 4.83' in Results, 1H-fMRS results.Consistent formatting improves readability.
- 11.LOWcopyeditClarify the phrasing '1,040 transients out of 16 water unsuppressed transients' in Methods, 1H-fMRS data acquisition.The current phrasing is ambiguous; clarifying that water unsuppressed transients are separate improves clarity.
- 12.LOWcopyeditSpecify the number of participants excluded for each reason in Results, 1H-fMRS results.Clarifies the exclusion breakdown for transparency.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.