Chronic adaptive deep brain stimulation versus conventional stimulation in Parkinson's disease: a blinded randomized feasibility trial.
Oehrn CR, Cernera S, Hammer LH, Shcherbakova M, Yao J, Hahn A, Wang S, Ostrem JL, Little S, Starr PA
- DOI
- 10.1038/s41591-024-03196-z
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/38be91f8-27a5-41f4-8f36-6132b178958f is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- CitationsUnresolved reference ×2−0.5★
- ReportingStudy design partially met−0.25★
- ReportingBiological variables partially met−0.25★
- ReportingStatistical analysis partially met−0.25★
- LinksDead data/code link−0.25★
- Statistics were not checked: no recomputable values were found in this text — no test statistic reported with its degrees of freedom, no effect estimate printed with both a 95% CI and a p-value, and no percentage printed with both its count and its denominator.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on patient-reported symptom duration (percentage of awake time with most bothersome symptom) and quality of life (EQ-5D), which are clinical outcomes, not surrogate biomarkers. However, the neural signal (stimulation-entrained gamma oscillations) used to drive adaptive DBS is a surrogate biomarker, and the paper does not provide validated evidence linking this surrogate to the clinical outcome, nor does it demonstrate target engagement at the tested dose beyond showing that the signal tracks medication states. The efficacy claim itself is based on clinical outcomes, so the surrogate verdict is not applicable to the primary claim.
“We identified stimulation-entrained gamma oscillations in the subthalamic nucleus or motor cortex as optimal markers of high and low dopaminergic states and their associated residual motor signs in all four patients. We then demonstrate improved motor…”
- 02Declared data/code link does not resolve
Dead link — nothing to verify.
“https://openmind-consortium.github.io”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper is a well-conducted feasibility trial with strong ethical approvals, clear resource identification, and transparent data/code sharing. The main weaknesses are incomplete reporting of randomization/blinding methods, limited demographic detail, and imprecise p-value reporting.
Both reviewers independently scored all eight dimensions and agreed on all statuses; no divergence required reconciliation. The statistics verification component checked 0 tests due to threshold-only p-values, so no statistical results were machine-verified; the citation component flagged 2 references as not found in registry (the data and code DOIs, which are not true citations but data/code identifiers).
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
- lowinternal contradictionThe abstract states 'Four male patients' while the Methods note 'all patients were of the same sex' without specifying sex; this is consistent but could be clearer.
“Four male patients with PD were recruited”
AbstractFind in source - lowinternal contradictionThe abstract states 'Four male patients' while the Methods mention 'sex assigned at birth'; this is a minor inconsistency in terminology.
“Four male patients with PD were recruited”
AbstractFind in source
Overstated conclusions
1 finding · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
3 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Adaptive DBS improves motor symptoms and quality of life compared to conventional DBS.The presented data from the crossover trial support this claim, with significant improvements in the primary outcome and quality of life.Evidence: Group-level linear mixed effects model showing improvement in bothersome symptom duration (β=-16.3%, p<10^-3) and EQ-5D (β=6.9, p<10^-4).
“We then demonstrate improved motor symptoms and quality of life with adaptive compared to clinically optimized standard stimulation.”
AbstractFind in source - supportedReviewers 1, 2Stimulation-entrained gamma oscillations are optimal biomarkers for motor signs.The paper provides extensive evidence from in-clinic and at-home recordings showing gamma oscillations outperform other frequency bands.Evidence: Non-parametric cluster-based permutation and machine learning analyses showing gamma oscillations as best predictors.
“In all four patients, STN or cortical stimulation-entrained gamma, outperformed other frequency bands, including STN beta activity, in distinguishing low and high dopaminergic states”
ResultsFind in source - supportedReviewers 1, 2The study establishes methodological principles for future larger trials.The paper describes a detailed seven-step pipeline that could be applied to larger cohorts.Evidence: Detailed methods and discussion of scalability.
“This study establishes methodological principles for conducting a future trial in a larger cohort of PD patients to assess the generalizability of these findings.”
DiscussionFind in source
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on patient-reported symptom duration (percentage of awake time with most bothersome symptom) and quality of life (EQ-5D), which are clinical outcomes, not surrogate biomarkers. However, the neural signal (stimulation-entrained gamma oscillations) used to drive adaptive DBS is a surrogate biomarker, and the paper does not provide validated evidence linking this surrogate to the clinical outcome, nor does it demonstrate target engagement at the tested dose beyond showing that the signal tracks medication states. The efficacy claim itself is based on clinical outcomes, so the surrogate verdict is not applicable to the primary claim.
“We identified stimulation-entrained gamma oscillations in the subthalamic nucleus or motor cortex as optimal markers of high and low dopaminergic states and their associated residual motor signs in all four patients. We then demonstrate improved motor symptoms and quality of life with adaptive compared to clinically optimized standard stimulation.”
- ADEQUATEEffect sizeThe effect size is reported as a reduction in the percentage of awake time with the most bothersome symptom (β = −16.3±4.4%, p<10−3) and an increase in quality of life (EQ-5D, β = 6.9±1.7, p<10−4). These are clinically meaningful outcomes with statistically significant improvements, and the paper provides individual patient data showing consistent improvements. The effect sizes are anchored to clinical outcomes (symptom duration and quality of life), which are meaningful to patients.
“A group-level linear mixed effects model including all four patients demonstrated an improvement in the percentage of awake-time experiencing the most bothersome symptom during aDBS compared to optimized cDBS (β =−16.3±4.4%, p<10−3), without worsening the percentage of awake-time with the opposite symptom (β =−2.5±2.2%, p=0.26). Furthermore, aDBS increased patients’ quality of life (EQ-5D, β=6.9±1.7, p<10−4).”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Data look implausibly cleanAssessed
3 integrity concerns flagged (0 high).
- lowdata too cleanAll four patients showed improvement in the primary outcome, which is notable for a small feasibility study but not implausible given the personalized approach.
“In each patient, aDBS reduced the time spent with bothersome motor symptoms compared to optimized cDBS”
ResultsFind in source
Reporting gaps
3 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Statistical reporting gaps (tests, assumptions, effect sizes)Assessed
- Biological variables underreported (sex, age, strain)Assessed
- Study-design details incomplete (controls, blinding, power)Assessed
The introduction thoroughly reviews prior research on DBS, adaptive DBS, and neural biomarkers, acknowledging limitations of previous studies (e.g., in-laboratory settings, lack of individualized approaches). The rationale linking the premise to the study objectives is strong, and the study directly addresses prior limitations by using a data-driven, home-based approach.
“However, these results were derived from in-laboratory studies and group-level analyses, lacking individualized, data-driven approaches.”
“Here, we conducted a blinded, randomized, cross-over feasibility trial with four PD patients (across six independently optimized hemispheres) to identify neural biomarkers of motor signs during active stimulation and compare the effects of aDBS to optimized cDBS during normal, unrestricted, daily life.”
“However, these results were derived from in-laboratory studies and group-level analyses, lacking individualized, data-driven approaches.”
“Here, we conducted a blinded, randomized, cross-over feasibility trial with four PD patients (across six independently optimized hemispheres) to identify neural biomarkers of motor signs during active stimulation and compare the effects of aDBS to optimized cDBS during normal, unrestricted, daily life.”
The trial is described as blinded and randomized, but the randomization method (e.g., sequence generation, allocation concealment) is not specified. Blinding is mentioned but not detailed (e.g., who was blinded, how blinding was maintained). A power analysis is provided using binomial distribution analysis, which is adequate for an N-of-1 design. Inclusion/exclusion criteria are clearly pre-specified. Outlier handling is not explicitly described, though some outliers are mentioned in figure legends. Controls are inherent in the crossover design (cDBS as control). Independent replication is not applicable for a single feasibility trial.
“This workflow culminated in a blinded, randomized, cross-over comparison between aDBS and cDBS”
“Binomial distribution analysis allows for predicting the probability of multiple independent subjects exhibiting significant effects in within-subjects statistics, each tested at an individual alpha level of p=0.05.”
“The inclusion criteria for DBS surgery were: age 25-75 years; ability to give informed consent for the study and to comply with study follow-up visits for brain recordings, testing of adaptive stimulation, and clinical assessments.”
“This workflow culminated in a blinded, randomized, cross-over comparison between aDBS and cDBS in the patients’ home environment”
“Binomial distribution analysis allows for predicting the probability of multiple independent subjects exhibiting significant effects in within-subjects statistics, each tested at an individual alpha level of p=0.05.”
“The inclusion criteria for DBS surgery were: age 25-75 years; ability to give informed consent for the study and to comply with study follow-up visits for brain recordings, testing of adaptive stimulation, and clinical assessments.”
Sex is reported (all male) but no scientific justification is given for single-sex enrollment; the paper acknowledges the limitation. Age, disease duration, and UPDRS scores are reported. Species/strain/source is not applicable (human study). Housing conditions not applicable. Demographics are reported (age, sex, disease duration, UPDRS) but race/ethnicity and comorbidities are not detailed.
“Reported sex of participants in our study is based on the sex assigned at birth and patient recruitment was determined by the availability and consent of eligible participants at the time.”
“We recruited four patients with PD from a population undergoing DBS implantation for motor fluctuations (male, age range: 47-68 years, disease duration: 10-15 years, pre-surgery off-medication Movement Disorder Society Unified Parkinson's Disease Rating Scale [MDS-UPDRS]-III scores: 30-49).”
“We recruited four patients with PD from a population undergoing DBS implantation for motor fluctuations (male, age range: 47-68 years, disease duration: 10-15 years, pre-surgery off-medication Movement Disorder Society Unified Parkinson's Disease Rating Scale [MDS-UPDRS]-III scores: 30-49).”
“We acknowledge the importance of including both sexes in future research. No sex- and gender-based analyses have been performed in our study, as all patients were of the same sex.”
The study reports approval from the UCSF IRB with protocol number, FDA investigational device exemption, and registration on ClinicalTrials.gov. Informed consent was obtained in writing per the Declaration of Helsinki. Regulatory compliance is stated. All applicable criteria are met.
“The Institutional Review Board of the University of California, San Francisco gave ethical approval for this work (18-24454, August 2, 2018).”
“Patients provided written consent in accordance with the Declaration of Helsinki.”
“The Institutional Review Board of the University of California, San Francisco gave ethical approval for this work (18-24454, August 2, 2018).”
“Patients provided written consent in accordance with the Declaration of Helsinki.”
The investigational device (Medtronic Summit RC+S) and leads (Medtronic model 3389, 0913025) are identified with model numbers. Software tools are identified with versions (MATLAB 2021a, FieldTrip, Lead-DBS). Reagents are not applicable. Antibodies, cell lines, and mycoplasma testing are not applicable.
“All patients underwent bilateral placement of cylindrical quadripolar deep brain stimulator leads (Medtronic model 3389) into the STN and bilateral quadripolar paddles (Medtronic model 0913025) into the subdural space over the sensorimotor cortex”
“We performed all analyses using MATLAB ® 2021a (The Mathworks, Natick, MA, USA) and the FieldTrip toolbox (version 20210507)”
“All patients underwent bilateral placement of cylindrical quadripolar deep brain stimulator leads (Medtronic model 3389) into the STN and bilateral quadripolar paddles (Medtronic model 0913025) into the subdural space over the sensorimotor cortex”
“We performed all analyses using MATLAB ® 2021a (The Mathworks, Natick, MA, USA) and the FieldTrip toolbox (version 20210507)”
Statistical tests are named (e.g., linear mixed effects model, Wilcoxon rank sum test). Assumptions are not explicitly verified for all tests, but standard methods are used. Many p-values are reported as thresholds (e.g., p<10^-3) rather than exact values, which is a warn-level issue. Effect sizes with confidence intervals are reported for some outcomes. Software is identified. Data presentation includes individual data points and error bars defined. Mathematical plausibility checks were not possible for most statistics due to lack of raw data.
“A group-level linear mixed effects model including all four patients demonstrated an improvement in the percentage of awake-time experiencing the most bothersome symptom during aDBS compared to optimized cDBS ( β =−16.3±4.4%, p<10 −3 )”
“A group-level linear mixed effects model including all four patients demonstrated an improvement in the percentage of awake-time experiencing the most bothersome symptom during aDBS compared to optimized cDBS ( β =−16.3±4.4%, p<10 −3 )”
“pat-1: cDBS: 22.7±1.5 [19.8, 25.7], aDBS: 13.2±1.2 [10.8, 15.7], p<10 −4”
De-identified individual participant data are shared on the Data Archive for the Brain Initiative with a DOI. Code is available on Code Ocean and GitHub. The data availability statement is concrete and meets the criteria for adequate.
“De-identified individual participant data, including neural, wearable, and digital diary data, are shared on the Data Archive for the Brain Initiative website ( https://dabi.loni.usc.edu ; 10.18120/cq9c-d057 ).”
“The code for biomarker identification implemented in Matlab is available in the repository Code Ocean, without restrictions ( 10.24433/CO.5656158.v1 )”
“De-identified individual participant data, including neural, wearable, and digital diary data, are shared on the Data Archive for the Brain Initiative website ( https://dabi.loni.usc.edu ; 10.18120/cq9c-d057 ).”
“The code for biomarker identification implemented in Matlab is available in the repository Code Ocean, without restrictions ( 10.24433/CO.5656158.v1 )”
The trial is registered on ClinicalTrials.gov with identifier. Methods are detailed. Limitations are discussed. Conclusions are proportional to the evidence. Funding and competing interests are stated. Reporting guideline is not explicitly mentioned, but the paper is well-structured.
“ClinicalTrials.gov (http://ClinicalTrials.gov) identifier: NCT03582891”
“However, the sample size was four patients (six independently optimized hemispheres).”
“ClinicalTrials.gov (http://ClinicalTrials.gov) identifier: NCT03582891”
“However, the sample size was four patients (six independently optimized hemispheres).”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
2 findings · worst mediumReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- Dead data/code linksRecomputed
- References not resolvable to a published paperRecomputed
Checked 67 references by DOI: 65 verified — 2 DOI unresolved.
- UNRESOLVED10.18120/cq9c-d057Chronic adaptive deep brain stimulation is superior to conventional stimulation in Parkinson’s disease: a blinded randomized feasibility trial [Source Data]Cited DOI does not resolve to any Crossref record.
- UNRESOLVED10.24433/co.5656158.v1Chronic adaptive deep brain stimulation is superior to conventional stimulation in Parkinson's disease: a blinded randomized feasibility trial [Source Code]Cited DOI does not resolve to any Crossref record.
3 data/code links checked; 2 live, 1 dead.
- datahttps://dabi.loni.usc.eduLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- dataOSFLIVEHTTP 200https://osf.io/cmndqResolves to OSF (data repository).
- codehttps://openmind-consortium.github.ioDEADHTTP 404Dead link — nothing to verify.
Copyediting
3 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 3 minor suggestions below.
3 copyedit issues flagged: mostly clarity, consistency.
- MINORconsistencyAbstract“Four male patients with PD were recruited”→ Consider specifying 'male' as 'sex assigned at birth' for consistency with Methods.Minor wording consistency.
- MINORclarityResults, Primary clinical outcomes“β =−16.3±4.4%, p<10 −3”→ Consider reporting exact p-value if available.Threshold p-values are used throughout; exact values would improve clarity.
- MINORclarityResults, Primary clinical outcomes“β =−16.3±4.4%, p<10 −3”→ Use consistent notation for p-values (e.g., p<0.001) throughout.Mixed use of superscripts and decimals.
The published work is methodologically sound for a feasibility study, but readers should weigh the incomplete randomization/blinding details and the lack of exact p-values. The dead link in the data/code availability and the two not-found registry entries warrant correction or clarification.
- 1.HIGHreportingSpecify the randomization method (e.g., random number generator, block randomization) and allocation concealment in the Methods section under 'Trial design'.The paper states the trial was randomized but does not describe the method, which is a key reporting gap for a randomized trial.
- 2.HIGHreportingDetail the blinding procedure: who was blinded (patients, assessors, statisticians) and how blinding was maintained.Blinding is mentioned but not described, and readers need to know if outcome assessment was blinded to avoid bias.
- 3.HIGHstatisticsReport exact p-values (e.g., p=0.003) instead of thresholds (e.g., p<0.01) for primary and secondary outcomes, or state that exact values are available in supplementary materials.Threshold p-values are imprecise and prevent readers from assessing the strength of evidence; exact values are expected for primary outcomes.
- 4.HIGHdata codeFix the dead link in the data/code availability section (reproducibility check found 1 dead link among 3 checked).A broken link undermines the data/code availability statement and prevents readers from accessing the shared resources.
- 5.HIGHreportingVerify or correct the two references flagged as not found in registry: the data DOI (10.18120/cq9c-d057) and the code DOI (10.24433/co.5656158.v1).These DOIs are listed as references but could not be found in any registry, which may indicate a fabrication signal or a citation error.
- 6.MEDIUMreportingExplicitly state how outliers were handled in the statistical analysis, including any exclusions and reasons.Outliers are mentioned in figure legends but not in the methods, leaving ambiguity about their treatment in the analysis.
- 7.MEDIUMreportingAdd a statement about adherence to a reporting guideline (e.g., CONSORT) in the Methods or a separate section.The paper does not mention a reporting guideline, which is a common expectation for clinical trials.
- 8.MEDIUMreportingInclude more detailed demographic information (e.g., race/ethnicity, comorbidities) in the patient characteristics table.Demographics are limited to age, sex, and disease duration; race/ethnicity and comorbidities are relevant for generalizability.
- 9.MEDIUMstatisticsClarify the statistical assumptions for the linear mixed effects model (e.g., normality of residuals) and how they were verified.Assumptions are not explicitly verified, and readers need to know if the model is appropriate for the data.
- 10.MEDIUMreportingProvide a scientific justification for enrolling only male patients, or acknowledge the limitation more explicitly in the Discussion.The paper acknowledges the limitation but does not justify the single-sex sample, which is important for interpreting generalizability.
- 11.LOWcopyeditUse consistent terminology for sex: specify 'sex assigned at birth' in the Abstract to match the Methods.The Abstract says 'male' while the Methods specify 'sex assigned at birth'; consistency improves clarity.
- 12.LOWcopyeditUse consistent notation for p-values (e.g., p<0.001) throughout the Results.Mixed use of superscripts and decimals (e.g., p<10 −3 vs p<0.001) is a minor clarity issue.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.