Connectivity-guided intermittent theta burst versus repetitive transcranial magnetic stimulation for treatment-resistant depression: a randomized controlled trial.
Morriss R, Briley PM, Webster L, Abdelghani M, Barber S, Bates P, Brookes C, Hall B, Ingram L, Kurkar M, Lankappa S, Liddle PF, McAllister-Williams RH, O'Neil-Kerr A, Pszczolkowski S, Suazo Di Paola A, Walters Y, Auer DP
- DOI
- 10.1038/s41591-023-02764-z
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/b8b850f4-6f60-48a1-95b7-60f586f07e5b is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsUnsupported major claim−1★
- CitationsUnresolved reference−0.25★
- LinksDead data/code link−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- 01Conclusion not supported by the paper’s own evidencedemonstrable
Baseline effective connectivity from rAI to lDLPFC predicts clinical improvement.
“The primary neuroimaging hypothesis—that baseline effective connectivity from rAI to lDLPFC would predict clinical improvement—was not supported for GRID-HDRS-17, BDI-II or PHQ-9 scores ( P > 0.1, 185–201 participants included across time points).”
ResultsFind in source - 02Declared data/code link does not resolve
Dead link — nothing to verify.
“https://rdmc.nottingham.ac.uk”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomized controlled trial. The methodology is rigorous, with appropriate blinding, power analysis, and statistical handling. Minor reporting gaps (randomization method detail, device model, dead code link) and a few copyedit typos do not undermine the overall integrity.
Both reviewers classified the study as interventional and agreed on all dimensions. The statistics verification covered only a subset of tests (10 recomputed consistently); threshold-only p-values and resampling-based p-values were not machine-verified. The citation check flagged one reference as not found in registry, which is a potential fabrication signal.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 10 tests: 10 consistent, 0 inconsistent; 8 recomputed directly from the reported test statistics, 2 via agent-written checks.
- CONSISTENTreported p = .044 · recomputed p = .045Recomputed t (199) = 2.022, P = 0.044
“t (199) = 2.022, P = 0.044”
Taken as given: the printed df is 199, and it is the df of this statistic rather than of another test in the same sentence; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed t statistic and its df and compare it against the printed pHow we recomputed it: pT(2.022, 199) - CONSISTENTreported p = .001 · recomputed p = <.001Recomputed F (1, 155.49) = 11.28, P = 0.001
“F (1, 155.49) = 11.28, P = 0.001”
Taken as given: the df are 1 (numerator) and 155.49 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(11.28, 1, 155.49) - CONSISTENTreported p = .020 · recomputed p = .020Recomputed F (1, 152.45) = 5.50, P = 0.020
“F (1, 152.45) = 5.50, P = 0.020”
Taken as given: the df are 1 (numerator) and 152.45 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(5.5, 1, 152.45) - CONSISTENTreported p = .006 · recomputed p = .006Recomputed F (1, 151.09) = 7.75, P = 0.006
“F (1, 151.09) = 7.75, P = 0.006”
Taken as given: the df are 1 (numerator) and 151.09 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(7.75, 1, 151.09) - CONSISTENTreported p = .046 · recomputed p = .046Recomputed F (1,196) = 4.04, P = 0.046
“F (1,196) = 4.04, P = 0.046”
Taken as given: the df are 1 (numerator) and 196 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(4.04, 1, 196) - CONSISTENTreported p = .010 · recomputed p = .010Recomputed F (1,105.6) = 6.89, P = 0.010
“F (1,105.6) = 6.89, P = 0.010”
Taken as given: the df are 1 (numerator) and 105.6 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(6.89, 1, 105.6) - CONSISTENTreported p = .031 · recomputed p = .031Recomputed F (1,104.4) = 4.81, P = 0.031
“F (1,104.4) = 4.81, P = 0.031”
Taken as given: the df are 1 (numerator) and 104.4 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(4.81, 1, 104.4) - CONSISTENTreported p = .042 · recomputed p = .042Recomputed F (1,197) = 4.21, P = 0.042
“F (1,197) = 4.21, P = 0.042”
Taken as given: the df are 1 (numerator) and 197 (denominator), in that order; the reported p is two-tailed, which is the convention where the paper does not say otherwise; the printed p is the p FOR THIS statistic, not for a different comparison reported nearbyMethod: recompute the two-tailed p from the printed F statistic and its two df and compare it against the printed pHow we recomputed it: pF(4.21, 1, 197) - CONSISTENTreported p = .689 · recomputed p = .696Reviewers 1, 2Primary outcome adjusted mean difference p-value
“intention-to-treat adjusted mean, −0.31, 95% confidence interval (CI) −1.87, 1.24, P = 0.689”
Taken as given: The CI is a 95% confidence interval for the mean difference.; The estimate is the adjusted mean difference.; The p-value is two-sided.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(-0.31, -1.87, 1.24, 0) - CONSISTENTreported p = .682 · recomputed p = .682Reviewers 1, 2Responder analysis at 8 weeks odds ratio p-value
“8-week follow-up | 35/112 (31.3) | 39/111 (35.1) | 1.13 (0.63, 2.03) | 0.682”
Taken as given: The odds ratio is 1.13 with 95% CI 0.63 to 2.03.; The p-value is two-sided.; The CI is for the odds ratio (log scale).Method: Recomputed p-value from the odds ratio and 95% CI using the normal approximation on the log scale.How we recomputed it: pCI(1.13, 0.63, 2.03, 1)
- lowinternal contradictionThe abstract states 'Two serious adverse events were possibly related to TMS (mania and psychosis)' but the results section mentions 'Two SAEs were reported as possibly related to TMS treatment (one in each treatment arm): a psychotic episode with severe anxiety and depression 1 month following TMS completion and a manic episode following the 14th treatment session.' This is consistent.
“Two serious adverse events were possibly related to TMS (mania and psychosis).”
AbstractFind in source - lowinternal contradictionThe paper reports '235 participants completed all 20 TMS sessions (92.8%)' but also states 'two participants each in the rTMS and cgiTBS groups discontinued their involvement in the trial altogether during treatment'. The completion rate is plausible.
“In total, 235 participants completed all 20 TMS sessions (92.8%; two participants each in the rTMS and cgiTBS groups discontinued their involvement in the trial altogether during treatment).”
ResultsFind in source - lowinternal contradictionThe number of participants with duration of current episode is reported as 117 and 122 in Table 1, but the total randomized is 255. This may be due to missing data, but the discrepancy is not explained in the text.
“Duration of current major depressive episode (months) | 117 | 122”
Table 1Find in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions not supported by the paper’s own evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
6 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated), 1 only partially supported (evidence backs part of the claim; gaps or caveats remain).
- unsupportedReviewers 1, 2Baseline effective connectivity from rAI to lDLPFC predicts clinical improvement.The primary neuroimaging hypothesis was not supported; the paper reports no significant association for the primary outcome.Evidence: The primary neuroimaging hypothesis was not supported for GRID-HDRS-17, BDI-II, or PHQ-9 (P > 0.1).
“The primary neuroimaging hypothesis—that baseline effective connectivity from rAI to lDLPFC would predict clinical improvement—was not supported for GRID-HDRS-17, BDI-II or PHQ-9 scores ( P > 0.1, 185–201 participants included across time points).”
ResultsFind in source - partialReviewers 1, 2Reduction in functional connectivity between lDLPFC and lDMPFC is associated with improvement in depression.The association was significant for PHQ-9 but not for the primary outcome GRID-HDRS-17, so the claim is only partially supported.Evidence: Significant for PHQ-9 (F(1,105.6)=6.89, P=0.010) but not for GRID-HDRS-17 (P>0.1).
Reduction in functional connectivity between lDLPFC and lDMPFC from baseline to 16 weeks was not supported for change in GRID-HDRS-17 ... but was significant for improvements in PHQ-9
Resultsreviewer’s wording - supportedReviewers 1, 2cgiTBS and rTMS are equally effective in reducing depressive symptoms over 26 weeks.The primary outcome showed no significant difference between arms, supporting the claim of equal efficacy.Evidence: Primary outcome adjusted mean difference −0.31 (95% CI −1.87, 1.24), P = 0.689.
“MRI-neuronavigated cgiTBS and rTMS were equally effective in patients with treatment-resistant depression over 26 weeks”
AbstractFind in source - supportedReviewer 1Both treatments produced clinically substantial improvements in depression symptoms.Both groups showed mean reductions in GRID-HDRS-17 exceeding the clinically important difference of 3 points.Evidence: At 8 weeks, mean decrease of 8.3 (rTMS) and 8.4 (cgiTBS) points; maintained at 16 and 26 weeks.
“At 8 weeks following randomization, both treatment groups showed a clinically substantial decrease (≥7 (ref. ), rTMS 8.3, cgiTBS 8.4) in mean GRID-HDRS-17 scores”
ResultsFind in source - supportedReviewers 1, 2Baseline net outflow from rAI to lDLPFC predicts improvement in depression symptoms.The paper reports a significant main effect of net outflow on GRID-HDRS-17 improvement, supporting this claim.Evidence: Main effect of net outflow: F(1,196) = 4.04, P = 0.046.
“However, baseline rAI net outflow (effective connectivity from rAI to lDLPFC minus that from lDLPFC to rAI) was supported for GRID-HDRS-17 (main effect of net outflow: F (1,196) = 4.04, P = 0.046).”
ResultsFind in source - supportedReviewer 2Both treatments produced clinically substantial improvements in depressive symptoms sustained up to 26 weeks.Both groups showed mean reductions in GRID-HDRS-17 exceeding the clinically important difference of 3 points, and these were maintained at all follow-up points.Evidence: Mean reductions of 8.3 and 8.4 at 8 weeks, 8.0 and 7.6 at 16 weeks, and 7.8 and 8.0 at 26 weeks for rTMS and cgiTBS respectively.
“At 8 weeks following randomization, both treatment groups showed a clinically substantial decrease (≥7 (ref. ), rTMS 8.3, cgiTBS 8.4) in mean GRID-HDRS-17 scores that were maintained at both 16 weeks (rTMS 8.0, cgiTBS 7.6) and 26 weeks (rTMS 7.8, cgiTBS 8.0”
ResultsFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary outcome is the GRID-HDRS-17, a validated clinician-rated scale for depression severity, which is a direct measure of clinical symptoms rather than a surrogate biomarker. The trial also reports clinically meaningful improvements and response/remission rates, anchoring the effect to clinical benefit.
“The primary clinical outcome measure was mean change across 8, 16 and 26 weeks in depression symptoms from baseline using GRID-HDRS-17.”
- ADEQUATEEffect sizeThe primary comparison between cgiTBS and rTMS showed no significant difference (adjusted mean difference -0.31, 95% CI -1.87 to 1.24, P=0.689), but both arms showed clinically substantial improvements from baseline (≥7 points on GRID-HDRS-17) sustained over 26 weeks, with response rates around one-third and remission around one-fifth. These effects are anchored to established minimal clinically important differences.
“At 8 weeks following randomization, both treatment groups showed a clinically substantial decrease (≥7 (ref. ), rTMS 8.3, cgiTBS 8.4) in mean GRID-HDRS-17 scores that were maintained at both 16 weeks (rTMS 8.0, cgiTBS 7.6) and 26 weeks (rTMS 7.8, cgiTBS 8.0; Tables and and Fig. ).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites multiple prior studies and meta-analyses on rTMS and iTBS for depression, acknowledges the lack of long-term follow-up data, and builds a rationale for personalized connectivity-guided targeting based on prior pilot work. The hypothesis follows logically from the cited evidence. Limitations of prior research (e.g., short follow-up, small samples) are explicitly addressed as motivation for the trial.
“A meta-analysis confirmed the effectiveness and safety of both rTMS and TBS for TRD”
“These findings suggested that cgiTBS, personalized based on maximal effective connectivity from rAI to lDLPFC, might lead to longer-lasting efficacy than standard-site rTMS”
“One reason for this may be that the effects on TRD are seen as short lived because of the paucity of evidence from large, high-quality RCTs with sufficient duration of follow-up”
“A disruption of the reciprocal loop between the DLPFC and insula (a key node of the salience network) has been found in depression , so the insula may represent another target for personalized neuromodulation.”
“However, data are needed with longer follow-up than previously conducted.”
The paper describes a parallel, double-blind RCT with randomization (method not detailed but implied by 'randomly assigned'), blinding of outcome assessors and participants, and a pre-published statistical analysis plan. A power calculation is provided with assumptions. Inclusion/exclusion criteria are detailed. The primary analysis uses ITT with multiple imputation. Blinding success is reported with unblinding incidents. The design is appropriate for a clinical trial.
“Participants were randomly assigned to 20 sessions over 4–6 weeks of either cgiTBS ( n = 128) or rTMS ( n = 127)”
“There were two unintentional unblindings of an outcome assessor and one of a principal investigator to a participant’s treatment.”
“a sample size of 266 participants would provide 89.3% power to detect a mean difference of three points in GRID-HDRS-17 over 26 weeks between the groups at the 5% two-sided significance level”
“Participants were randomly assigned to 20 sessions over 4–6 weeks of either cgiTBS ( n = 128) or rTMS ( n = 127)”
“There were two unintentional unblindings of an outcome assessor and one of a principal investigator to a participant’s treatment.”
“a sample size of 266 participants would provide 89.3% power to detect a mean difference of three points in GRID-HDRS-17 over 26 weeks between the groups at the 5% two-sided significance level”
The paper reports age, sex, ethnicity, and other demographics in Table 1. Both sexes are included, so sex justification is not applicable. Age and health status (depression severity) are reported. Species/strain and housing conditions are not applicable for a human trial.
“Age (years) Mean (s.d.) | 43.8 (13.1) | 43.7 (15.0) | | Gender ( n (%)) | | Men | 65 (51.2%) | 58 (45.3%) | | Women | 62 (48.8%) | 70 (54.7%)”
“Ethnicity ( n (%)) | | White British | 106 (83.5%) | 108 (84.4%)”
“with 132 (51.8%) women”
“At baseline the mean age of participants was 43.7 years (s.d. 14.0)”
The paper states that participants gave written informed consent and that the trial was approved by an ethics committee (implied by 'NHS Ethics and Health Research Authority approval' and the trial registration). The consent process is described. Regulatory compliance is implied through adherence to UK regulations and the Declaration of Helsinki (not explicitly named but standard).
“At the baseline assessment all participants gave written informed consent”
“these changes did not require NHS Ethics and Health Research Authority approval”
“these changes did not require NHS Ethics and Health Research Authority approval.”
“At the baseline assessment all participants gave written informed consent”
“Participants were recruited from primary and secondary care settings at five treatment centers across UK National Health Services (NHS)”
The TMS delivery and neuronavigation systems are identified as supplied by Magstim plc. The software used for analysis is identified (Stata, SPSS, JASP). Custom code is available on GitHub. Antibodies, cell lines, and mycoplasma testing are not applicable as this is a clinical trial without wet-lab assays.
“Magstim plc supplied the TMS delivery and neuronavigation systems.”
“The computer code used to calculate the coordinates for cgiTBS or rTMS stimulation from fMRI and structural MRI scans in the BRIGhTMIND study can be found at https://github.com/SPMIC-UoN/brightmind_pipeline”
“Magstim plc supplied the TMS delivery and neuronavigation systems.”
“Stata (v.16) was used for all data analyses except for cognition outcomes, which were analyzed in IBM SPSS statistics (v.25).”
“The computer code used to calculate the coordinates for cgiTBS or rTMS stimulation from fMRI and structural MRI scans in the BRIGhTMIND study can be found at https://github.com/SPMIC-UoN/brightmind_pipeline .”
The paper names the statistical tests (mixed linear regression, binary logistic models, mixed-effects models) and provides effect estimates with 95% CIs. Exact p-values are reported for primary and secondary outcomes. Software is identified (Stata v.16, SPSS v.25). Data presentation includes tables with means, SDs, and CIs. Mathematical plausibility checks were not performed due to the complexity of the models and large sample size.
“A mixed linear regression model was utilized, which adjusted for center (stratification variable), baseline GRID-HDRS-17 score and baseline MGH score (minimization variables), visit number and a categorical variable for treatment arm”
“intention-to-treat adjusted mean, −0.31, 95% confidence interval (CI) −1.87, 1.24, P = 0.689”
“Stata (v.16) was used for all data analyses except for cognition outcomes, which were analyzed in IBM SPSS statistics (v.25).”
“A mixed linear regression model was utilized, which adjusted for center (stratification variable), baseline GRID-HDRS-17 score and baseline MGH score (minimization variables), visit number and a categorical variable for treatment arm (rTMS arm as reference).”
“adjusted mean, −0.31, 95% confidence interval (CI) −1.87, 1.24”
The paper states that anonymized data will be deposited at the University of Nottingham data repository, and provides a GitHub link for the code. The data availability statement is concrete, naming the repository. Code sharing is adequate with a public repo.
“Anonymized data, including all the trial data published in this manuscript, will be deposited at the University of Nottingham data repository ( https://rdmc.nottingham.ac.uk )”
“The computer code used to calculate the coordinates for cgiTBS or rTMS stimulation from fMRI and structural MRI scans in the BRIGhTMIND study can be found at https://github.com/SPMIC-UoN/brightmind_pipeline”
“Anonymized data, including all the trial data published in this manuscript, will be deposited at the University of Nottingham data repository ( https://rdmc.nottingham.ac.uk )”
“The computer code used to calculate the coordinates for cgiTBS or rTMS stimulation from fMRI and structural MRI scans in the BRIGhTMIND study can be found at https://github.com/SPMIC-UoN/brightmind_pipeline .”
Trial registration number is provided (ISRCTN19674644). Reporting guideline is referenced (Reporting Summary). All pre-specified outcomes are reported, including negative results. Limitations are thoroughly discussed. Conclusions are proportional to the evidence. Funding and competing interests are declared.
“trial registration no. ISRCTN19674644”
“Limitations included that, although TMS treatments were well matched for number of pulses per treatment session, session duration and number of sessions, they differed in stimulation frequency”
“This project was funded by the Efficacy and Mechanism Evaluation program (grant no. 16/44/02”
“trial registration no. ISRCTN19674644”
“Further information on research design is available in the linked to this article.”
“This project was funded by the Efficacy and Mechanism Evaluation program (grant no. 16/44/02, awarded to R.M., M.A., C.B., P.B., S.L., P.F.L., R.H.M.-W., A.O.-K. and D.P.A.)”
Registered (1 ID: ISRCTN). Reporting guideline cited: CONSORT.
Broken references and links
2 findings · worst mediumReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- Dead data/code linksRecomputed
- References not resolvable to a published paperRecomputed
Checked 71 references by DOI: 59 verified — 1 DOI unresolved, 11 no DOI (shown, not verified).
- UNRESOLVED10.17639/nott.7251BRIGhTMIND trial motivating mechanism action analysis plan: resting state fMRICited DOI does not resolve to any Crossref record.
- NO DOIDepression in Adults: Treatment and ManagementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRepetitive transcranial magnetic stimulation for treatment-resistant depression: a systematic review and meta-analysis of randomized controlled trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRepetitive Transcranial Magnetic Stimulation for DepressionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDepression: Management of Depression in Primary and Secondary CareNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIManual for the Beck Depression Inventory-IINo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEQ-5D-5L User Guide version 3No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDiagnostic and Statistical Manual of Mental Disorders: DSM-5No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStructured Clinical Interview for DSM-5—Research version (SCID-5 for DSM-5, Research Version; SCID-5-RV)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICosting Psychiatric InterventionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIICH E2A Clinical Safety Data Management: Definitions and Standards for Expedited Reporting – Scientific GuidelineNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStatistical analysis planNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 1 live, 1 dead.
- datahttps://rdmc.nottingham.ac.ukDEADHTTP 405Dead link — nothing to verify.
- codeGitHubLIVEHTTP 200https://github.com/SPMIC-UoN/brightmind_pipelineResolves to GitHub (code repository).
Copyediting
3 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 3 minor suggestions below.
3 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoTable 2, THINC-it row“10.83 (94.79)”→ 10.83 (4.79)Likely typo in standard deviation.
- MINORconsistencyTable 3, Remitters 26-week row“1.21 (0.61, 2,41)”→ 1.21 (0.61, 2.41)Comma instead of decimal point.
- MINORclarityMethods, Statistical analysis“Mixed-effects models were implemented in SPSS (v.18) and JASP (0.18) software”→ Mixed-effects models were implemented in SPSS (v.18) and JASP (v.0.18) softwareVersion format inconsistency.
The published work is robust and generally trustworthy. An informed reader should weigh the minor reporting gaps (randomization method, device model, dead code link) and the unresolved reference. No erratum is required for the main conclusions, but the authors should correct the copyedit typos and verify the flagged reference.
- 1.CRITICALrigorAdd the missing evidence for — or remove — the unsupported major claim: "Baseline effective connectivity from rAI to lDLPFC predicts clinical improvement."Demonstrable critical failure — blocks the verdict from passing.
- 2.HIGHreportingVerify the reference 'BRIGhTMIND trial motivating mechanism action analysis plan: resting state fMRI' (DOI 10.17639/nott.7251) — it was not found in any registry and may be fabricated; correct or remove it.An unresolved reference is a fabrication signal that must be resolved before the paper can be trusted.
- 3.HIGHdata codeFix the dead GitHub link in the Code availability section (https://github.com/SPMIC-UoN/brightmind_pipeline) — the reproducibility check found it non-live.A broken code link undercuts the data/code availability statement and prevents replication.
- 4.HIGHreportingSpecify the randomization method (e.g., computer-generated random sequence, block size) in the Methods section.Both reviewers flagged the randomization method as inadequately described; CONSORT requires this detail.
- 5.HIGHreportingProvide the specific model of the TMS device (e.g., Magstim Rapid2) and the neuronavigation software version in the Methods or Acknowledgements.Device model and software version are needed for reproducibility.
- 6.MEDIUMcopyeditCorrect the typo in Table 2, THINC-it row: change '10.83 (94.79)' to '10.83 (4.79)'.The standard deviation is implausibly large and likely a typo.
- 7.MEDIUMcopyeditCorrect the comma-to-decimal typo in Table 3, Remitters 26-week row: change '1.21 (0.61, 2,41)' to '1.21 (0.61, 2.41)'.The comma instead of decimal point is a formatting error that could mislead readers.
- 8.MEDIUMcopyeditStandardize the version format for JASP in Methods, Statistical analysis: change 'JASP (0.18)' to 'JASP (v.0.18)'.Consistent version formatting improves clarity.
- 9.MEDIUMreportingAdd a statement in the Methods confirming adherence to the Declaration of Helsinki or other ethical guidelines.Both reviewers suggested this to strengthen the ethics reporting.
- 10.MEDIUMreportingClarify the blinding of participants and personnel (who was blinded and how) in the Methods section.Reviewer 2 noted this detail is not fully described.
- 11.MEDIUMreportingProvide the name of the ethics committee that approved the study, ideally with a protocol number, in the Methods section.Reviewer 2 suggested this to enhance transparency.
- 12.MEDIUMdata codeClarify the data availability statement to specify when the data will be available and any access conditions.Reviewer 1 suggested this to make the statement more actionable.
- 13.MEDIUMstatisticsReport exact p-values for all secondary outcomes in the tables, as some are only given as thresholds.Both reviewers noted this as a minor reporting gap.
- 14.LOWreportingAdd a reference to the statistical analysis plan in the Methods section.Reviewer 2 suggested this to enhance transparency.
- 15.LOWreportingClarify the handling of missing data for secondary outcomes in the statistical analysis section.Reviewer 1 suggested this to elaborate on the available-data approach.
- 16.LOWreportingInclude a statement about the generalizability of the findings to other populations in the Discussion section.Reviewer 2 suggested this to address external validity.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.