Accelerated continuous theta burst stimulation targeting left primary motor cortex for children with autism spectrum disorder: multicentre randomised sham controlled trial.
Tan H, Ren T, Cao A, Fang S, Deng L, Hu B, Wang M, Cheng Y, Zhang X, Li Y, Zhang Y, Zhang L, Chen L, Zhou W, Zhang Q, Li J, Zhou X, Langley C, Luo Q, Zhang J, Sahakian BJ, Yuan TF, Li F
- DOI
- 10.1136/bmj-2025-086295
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/f1f9742c-7da6-4b6f-abb5-cceb63b3e61c is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- CitationsUnresolved reference−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on the Social Responsiveness Scale, second edition (SRS-2), a caregiver-reported questionnaire that measures social communication impairment. This is a surrogate outcome, not a hard clinical endpoint. The paper does not provide evidence of target engagement at the tested dose (e.g., neurophysiological measures such as motor evoked potentials or functional imaging) nor does it cite validated evidence linking changes in SRS-2 scores to long-term clinical outcomes. Although the authors estimate a minimal clinically important difference (MCID) of 5.61 points, this is a post hoc anchor-based estimate from previous trials and not a validated link to clinical meaningfulness.
“The primary outcome was defined as the change in the total score of the Social Responsiveness Scale, second edition (SRS-2, school age version) from T0 to T1 and from T0 to T2.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted multicentre randomised sham-controlled trial with rigorous design, clear reporting, and adequate ethical and data-sharing practices. Minor reporting issues (threshold p-values, a typographical error in Table 3) do not affect the overall integrity of the study.
Both reviewers classified the study as interventional, and this was adopted. The evaluation covered the full text, with verification components checking citations, statistics, reproducibility, and preregistration. Non-applicable sub-criteria (e.g., animal housing, cell lines) were excluded from scoring.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 4 tests: 4 consistent, 0 inconsistent; 4 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Primary outcome at T1: mean difference -6.25, 95% CI -8.69 to -3.81
“mean difference in SRS-2 scores of −6.25 (95% confidence interval (CI) −8.69 to −3.81; Cohen’s d −0.92; P<0.001) at T1”
Taken as given: The CI is a 95% confidence interval for the mean difference.; The estimate is the mean difference (-6.25).; The p-value is two-sided.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(-6.25, -8.69, -3.81, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Primary outcome at T2: mean difference -6.17, 95% CI -8.65 to -3.70
“−6.17 (−8.65 to −3.70; −0.90; P<0.001) at T2”
Taken as given: The CI is a 95% confidence interval for the mean difference.; The estimate is the mean difference (-6.17).; The p-value is two-sided.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(-6.17, -8.65, -3.70, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1CGI-I improvement at T1: 71.7% vs 46.5%
“71.7% of participants in the active group reporting improvement compared with 46.5% in the sham group at T1 (P<0.001)”
Taken as given: The percentages are based on the mITT population of 99 per group.; The counts are approximated as 71 and 46 events out of 99 per group.; The test is a two-sided chi-square test.Method: Recomputed p-value using Pearson's chi-square test on the approximated 2x2 table.How we recomputed it: pChi2x2(71, 28, 46, 53) - CONSISTENTreported p < .008 · recomputed p = .007Reviewer 1CGI-I improvement at T2: 84.8% vs 68.7%
“84.8% versus 68.7% at T2 (P=0.008)”
Taken as given: The percentages are based on the mITT population of 99 per group.; The counts are approximated as 84 and 68 events out of 99 per group.; The test is a two-sided chi-square test.Method: Recomputed p-value using Pearson's chi-square test on the approximated 2x2 table.How we recomputed it: pChi2x2(84, 15, 68, 31)
- lowinternal contradictionThe abstract reports '167 boys and 33 girls' (total 200), but Table 1 reports 87 males in active and 80 in sham, totaling 167 males, consistent. However, the abstract also states '83.5% male' which is consistent with 167/200. No contradiction.
“200 children aged 4-10 years with autism spectrum disorder (167 boys and 33 girls)”
Table 1Find in source - lowinternal contradictionTable 3 reports Tinnitus as '0 (0) | 2 (0)' for active and sham groups, but the sham group percentage should be 2.0% (2/99), not 0%.
“Tinnitus | 0 (0) | 2 (0) | 0.50”
Table 3Find in source - lowinternal contradictionIn Table 3, the sham group's tinnitus count is 2 but the percentage is listed as 0%, which is inconsistent.
“Tinnitus | 0 (0) | 2 (0) | 0.50”
Table 3Find in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1These findings support a-cTBS as a viable and scalable therapeutic option for children with autism spectrum disorder.The efficacy and safety data support viability, but scalability is inferred from the protocol's simplicity and high adherence, not directly measured.Evidence: Discussion on scalability and adherence rates.
“These findings support a-cTBS as a viable and scalable therapeutic option for children with autism spectrum disorder.”
ConclusionFind in source - partialReviewer 2The a-cTBS protocol improved language abilities.Language improvements were observed on the MAIN measure, but not on other language measures (PPVT, CCDI).Evidence: MAIN showed significant improvements (Cohen's d 0.12-0.47; all P<0.02), but PPVT and CCDI showed no significant between-group differences.
“Among the 119 participants (60.1%) with sufficient expressive language ability who completed the MAIN test, those in the a-cTBS group showed greater improvements in narrative production and comprehension (Cohen’s d 0.12-0.47; all P<0.02; ).”
ResultsFind in source - supportedReviewer 1A five day a-cTBS protocol targeting the left primary motor cortex significantly improved social communication in children with autism spectrum disorder.The primary outcome showed a significant improvement in SRS-2 scores with a mean difference of -6.25 (95% CI -8.69 to -3.81, P<0.001) at post-intervention and -6.17 at one month follow-up, supporting the claim.Evidence: Primary outcome results in the abstract and results section.
“A five day a-cTBS protocol targeting the left primary motor cortex significantly improved social communication in children with autism spectrum disorder”
ConclusionFind in source - supportedReviewer 1The protocol showed a favourable safety profile.Adverse events were mostly mild to moderate and resolved spontaneously, with only one moderate event, supporting the claim.Evidence: Safety results in the results section and Table 3.
“showed a favourable safety profile”
ConclusionFind in source - supportedReviewer 1The protocol showed feasibility and efficacy in young children and people with intellectual disability.Subgroup analyses showed consistent treatment effects across age and intellectual disability subgroups, supporting the claim.Evidence: Subgroup analyses in the results section.
“Our protocol showed feasibility and efficacy in young children and people with intellectual disability—populations traditionally underrepresented in neuromodulation studies.”
DiscussionFind in source - supportedReviewer 1The average treatment effects exceeded the estimated MCID, suggesting clinical importance.The mean differences (-6.25 and -6.17) exceed the estimated MCID of 5.61, supporting the claim.Evidence: Primary outcome results and MCID estimation in the methods.
“In this study, the average treatment effects of a-cTBS exceeded the estimated MCID, suggesting its clinical importance.”
DiscussionFind in source - supportedReviewer 2The a-cTBS protocol significantly improved social communication impairment in children with ASD.The primary outcome showed a significant reduction in SRS-2 scores compared to sham, with effect sizes exceeding the estimated MCID.Evidence: Primary outcome results: mean difference -6.25 (95% CI -8.69 to -3.81; P<0.001) at T1 and -6.17 (-8.65 to -3.70; P<0.001) at T2.
“Compared with the sham group, the a-cTBS group showed significantly greater reductions in SRS-2 scores post-intervention (−6.25, 95% confidence interval −8.69 to −3.81; Cohen’s d −0.92; P<0.001) and at one month follow-up (−6.17, −8.65 to −3.70; −0.90; P<0.001).”
AbstractFind in source - supportedReviewer 2The a-cTBS protocol showed a favourable safety profile.Adverse events were mostly mild, with one moderate event, and all resolved spontaneously.Evidence: Safety results: adverse events more frequent in active group (54.5% vs 29.3%), but all mild except one moderate; all resolved.
“All adverse events were mild except for one moderate event; all resolved spontaneously.”
ResultsFind in source - supportedReviewer 2The a-cTBS protocol is feasible and scalable.High adherence (96.5% completed) and no need for neuronavigation support feasibility and scalability.Evidence: Adherence rate of 96.5% and the protocol's simplicity without neuronavigation.
“The protocol achieved a high adherence rate, with 193 out of 200 (96.5%) enrolled participants completing all sessions.”
DiscussionFind in source
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on the Social Responsiveness Scale, second edition (SRS-2), a caregiver-reported questionnaire that measures social communication impairment. This is a surrogate outcome, not a hard clinical endpoint. The paper does not provide evidence of target engagement at the tested dose (e.g., neurophysiological measures such as motor evoked potentials or functional imaging) nor does it cite validated evidence linking changes in SRS-2 scores to long-term clinical outcomes. Although the authors estimate a minimal clinically important difference (MCID) of 5.61 points, this is a post hoc anchor-based estimate from previous trials and not a validated link to clinical meaningfulness.
“The primary outcome was defined as the change in the total score of the Social Responsiveness Scale, second edition (SRS-2, school age version) from T0 to T1 and from T0 to T2.”
- ADEQUATEEffect sizeThe reported treatment effect on the primary outcome (SRS-2) was a mean difference of -6.25 points (95% CI -8.69 to -3.81) at post-intervention and -6.17 points (95% CI -8.65 to -3.70) at one month follow-up, with Cohen's d around -0.90. The authors state that these effects exceed their estimated MCID of 5.61 points, providing an anchor for clinical meaningfulness. The effect size is statistically significant and the magnitude is explicitly compared to a clinically meaningful threshold.
“The average treatment effects were higher than the estimated MCID on SRS-2 (−5.61 points, supplementary methods).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on rTMS in ASD and acknowledges limitations of previous studies, including exclusion of young children and those with intellectual disability. The rationale links these gaps to the study's objectives and hypothesis. The paper also describes how the new protocol addresses prior limitations.
“However, current empirical evidence remains limited and inconclusive.”
“To address these challenges, we developed and evaluated a rTMS protocol adapted for young children with ASD and children with co-occurring intellectual disability.”
“This design delivers a high cumulative dose of stimulation within a condensed time frame.”
“However, current empirical evidence remains limited and inconclusive.”
“To address these challenges, we developed and evaluated a rTMS protocol adapted for young children with ASD and children with co-occurring intellectual disability.”
“The trial included two key autism subpopulations that were frequently excluded from rTMS studies: children with intellectual disability and children of young age.”
Randomization method (block randomization, block length=4) and unit (participant) are stated. Blinding of participants and evaluators is described, with operators unmasked. Power analysis is reported with effect size, alpha, and power. Inclusion/exclusion criteria are pre-specified. Outlier handling is addressed through the modified intention-to-treat population and multiple imputation. Controls (sham) are described. Independent replication is not applicable for a single pivotal trial.
“using a block randomisation sequence (block length=4) generated by an independent coordinator in SPSS version 25.0”
“Participants and evaluators were masked to interventions.”
“Using a two sided α=0.025 to account for two primary outcomes, and assuming 80% power and 10% attrition, the required sample size was 186 participants.”
“using a block randomisation sequence (block length=4) generated by an independent coordinator in SPSS version 25.0”
“Participants and evaluators were masked to interventions.”
“Using a two sided α=0.025 to account for two primary outcomes, and assuming 80% power and 10% attrition, the required sample size was 186 participants.”
Sex is reported (167 boys and 33 girls). Age is reported (mean 6.5±1.6 years). Demographics include sex, age, and intellectual disability status. Species/strain and housing conditions are not applicable for a human trial. Sex justification is not applicable because both sexes are enrolled.
“200 children aged 4-10 years with autism spectrum disorder (167 boys and 33 girls)”
“50.5% of participants had co-occurring intellectual disability (full scale intelligence quotient <70)”
“200 eligible participants (mean age 6.5±1.6 years; 83.5% male)”
“Eligible participants were children with ASD aged 4-10 years with a full scale intelligence quotient of 50 or higher.”
The trial protocol was approved by named ethics committees of all participating sites, with a protocol number for one site. Written informed consent was obtained from legal guardians. Regulatory compliance with the Declaration of Helsinki and Good Clinical Practice is stated.
“The trial protocol was approved by the ethics committees of Xinhua Hospital affiliated to Shanghai Jiao Tong University School of Medicine (XHEC-C-2023-043-4) and each participating site”
“The study procedures and objectives were explained in person to the legal guardians of all participants, who provided written informed consent.”
“The study was conducted in accordance with the Good Clinical Practice and the principles of the Declaration of Helsinki.”
“The trial protocol was approved by the ethics committees of Xinhua Hospital affiliated to Shanghai Jiao Tong University School of Medicine (XHEC-C-2023-043-4) and each participating site”
“The study procedures and objectives were explained in person to the legal guardians of all participants, who provided written informed consent.”
“The study was conducted in accordance with the Good Clinical Practice and the principles of the Declaration of Helsinki.”
The investigational product (M-100 Ultimate, Shenzhen Yingchi Technology Co., China) is identified with manufacturer and model. Statistical software (R version 4.5.2) and SPSS version 25.0 are identified. No antibodies, cell lines, or organisms are used, so those criteria are not applicable.
“All sites used the same model of pulsed magnetic stimulation device (M-100 Ultimate, Shenzhen Yingchi Technology Co., China).”
“All statistical analyses were performed using R version 4.5.2 (R Foundation for Statistical Computing, Vienna, Austria).”
“All sites used the same model of pulsed magnetic stimulation device (M-100 Ultimate, Shenzhen Yingchi Technology Co., China).”
“All statistical analyses were performed using R version 4.5.2 (R Foundation for Statistical Computing, Vienna, Austria).”
Statistical tests are named (linear regression, ordinal logistic regression). Assumptions are handled by design (mixed models, multiple imputation). Exact p-values are reported for primary outcomes. Effect sizes with CIs are reported. Software is identified. Data presentation includes per-group n and CIs. Mathematical plausibility is not applicable for large-N continuous outcomes.
“the group differences in changes were assessed using linear regression”
“−6.25, 95% confidence interval −8.69 to −3.81; Cohen’s d −0.92”
“the group differences in changes were assessed using linear regression”
“mean difference in SRS-2 scores of −6.25 (95% confidence interval (CI) −8.69 to −3.81; Cohen’s d −0.92; P<0.001)”
A data availability statement is present, naming a public repository (OSF) with a URL. Code is stated to be in the supplementary appendix. Repository deposit is applicable and adequate. Accession numbers are not applicable for this type of data. Code sharing is adequate as the code is in the appendix.
“The data underlying the findings in this paper are openly and publicly available and can be found at: https://osf.io/tg9me/overview”
“The code used to analyse the data in the paper can be found in the supplementary appendix.”
“The data underlying the findings in this paper are openly and publicly available and can be found at: https://osf.io/tg9me/overview”
“The code used to analyse the data in the paper can be found in the supplementary appendix.”
The trial is registered (NCT05927792). CONSORT guidelines are followed. All outcomes are reported, including negative results. Limitations are discussed. Conclusions are proportional. Funding and conflicts of interest are disclosed.
“ClinicalTrials.gov NCT05927792”
“This study followed the CONSORT (consolidated standards of reporting trials) guidelines”
“Several limitations warrant consideration when interpreting the results.”
“Trial registration ClinicalTrials.gov NCT05927792”
“This study followed the CONSORT (consolidated standards of reporting trials) guidelines”
“Several limitations warrant consideration when interpreting the results.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 70 references by DOI: 64 verified — 1 DOI unresolved, 5 no DOI (shown, not verified).
- UNRESOLVED10.1016/s2215-0366(24Research prioritiesCited DOI does not resolve to any Crossref record.
- NO DOIDiagnostic and statistical manual of mental disorders: DSM-5No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIManual of the Wechsler Intelligence Scale for Children-RevisedNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWechsler Preschool and Primary Scale of IntelligenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISocial Responsiveness ScaleNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISample Size Calculations in Clinical ResearchNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- dataOSFLIVEHTTP 200https://osf.io/tg9me/overviewResolves to OSF (data repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity.
- MINORconsistencyTable 3, Tinnitus row“0 (0) | 2 (0)”→ Change to '0 (0) | 2 (2.0)' for consistency with other rows.The sham group percentage is missing.
- MINORclarityAbstract, Results“Cohen’s d 0.12-0.47; all P<0.02”→ Consider reporting exact p-values for clarity.P-values are reported as thresholds.
- MINORconsistencyData availability statement“https://osf.io/tg9me/overview (10.1177/1362361320967790)”→ Remove the DOI or clarify its relevance.The DOI appears unrelated to the OSF link.
- MINORconsistencyTable 3“Tinnitus | 0 (0) | 2 (0) | 0.50”→ The sham group percentage for tinnitus should be 2.0% (2/99), not 0%.The percentage appears to be a typographical error.
- MINORclarityAbstract, Results“Secondary outcomes also favoured a-cTBS, with significant improvements observed in language abilities (Cohen’s d 0.12-0.47; all P<0.02; measured by Multilingual Assessment Instrument for Narratives).”→ Clarify that the language improvements were only on the MAIN measure, not all language measures.The abstract could be misinterpreted as all secondary outcomes were significant.
The published work is robust and well-reported. An informed reader should weigh the minor reporting inconsistencies (threshold p-values, Table 3 typo) and the unresolved reference, but these do not undermine the main conclusions. No erratum is urgently required, though correcting the Table 3 percentage and clarifying the abstract would improve clarity.
- 1.HIGHcopyeditCorrect the tinnitus row in Table 3: change the sham group percentage from '0' to '2.0%' (2/99) to match the count.The current entry '2 (0)' is internally inconsistent and could mislead readers about adverse event rates.
- 2.HIGHreportingVerify or correct the reference 'Research priorities' (DOI 10.1016/s2215-0366(24) which could not be found in any registry; if it is a real paper, provide the complete DOI and confirm it exists.An unresolved reference is a potential fabrication signal and must be resolved before the paper is relied upon.
- 3.MEDIUMstatisticsReport exact p-values instead of thresholds (e.g., P<0.001) for primary and secondary outcomes in the abstract and results.Threshold p-values are less informative and were flagged by both reviewers and the copyedit pass as a clarity issue.
- 4.MEDIUMreportingClarify in the abstract that the language improvements were observed on the MAIN measure only, not all language measures.The current wording could be misinterpreted as all secondary outcomes being significant, which overstates the findings.
- 5.MEDIUMdata codeProvide the statistical analysis code in a version-controlled public repository (e.g., GitHub) with a DOI, in addition to the supplementary appendix.This would enhance reproducibility and is a common expectation for data-driven papers.
- 6.MEDIUMreportingAdd a more detailed description of the multiple imputation procedure (variables used, number of imputations) in the main text.Currently only in the supplementary, this detail is important for assessing the robustness of the primary analysis.
- 7.MEDIUMreportingReport the results of the masking integrity assessment in the main text rather than only in the supplementary.Blinding integrity is a key quality indicator for sham-controlled trials and should be transparently reported.
- 8.MEDIUMreportingDiscuss the potential impact of the baseline SRS-2 imbalance on the primary outcome in more detail, even though sensitivity analyses were performed.This addresses a potential confound that readers may question.
- 9.MEDIUMreportingReport the MCID estimation method in the main text, as it is currently only in the supplementary methods.The minimal clinically important difference is central to interpreting the clinical significance of the results.
- 10.LOWdata codeRemove or clarify the DOI (10.1177/1362361320967790) in the data availability statement, as it appears unrelated to the OSF link.The unrelated DOI could confuse readers about the data location.
- 11.LOWreportingAdd a statement about the availability of the study protocol and statistical analysis plan in a public repository.This would further enhance transparency and adherence to open science practices.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.