Accelerated continuous theta burst stimulation targeting left primary motor cortex for children with autism spectrum disorder: multicentre randomised sham controlled trial.
Tan H, Ren T, Cao A, Fang S, Deng L, Hu B, Wang M, Cheng Y, Zhang X, Li Y, Zhang Y, Zhang L, Chen L, Zhou W, Zhang Q, Li J, Zhou X, Langley C, Luo Q, Zhang J, Sahakian BJ, Yuan TF, Li F
- DOI
- 10.1136/bmj-2025-086295
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/90a2e37c-0444-44b5-b039-a5dd65188dee is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsOverstated claim−0.5★
- ReportingStatistical analysis partially met−0.25★
- ReportingData & code availability partially met−0.25★
- CitationsUnresolved reference−0.25★
- 01Conclusion reaches beyond the evidence
This protocol represents a major advancement towards equitable autism care worldwide.
“By addressing key limitations of conventional rTMS, this protocol represents a major advancement towards equitable autism care worldwide.”
ConclusionFind in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed and transparently reported multicentre RCT of a-cTBS for autism. The study shows strong methodological rigor in design, ethics, and reporting, but has two minor weaknesses: imprecise p-value reporting for the primary outcome and analysis code being only in the supplementary appendix.
Both independent reviewer runs (same model, independently sampled) agreed on all dimension statuses and assigned identical overall scores. The verification components checked 8 statistics (all consistent), 70 citations (1 not found in registry), and 1 data link (live). The claim audit flagged one overstated claim. The copyedit pass identified 5 minor issues.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 8 tests: 8 consistent, 0 inconsistent; 8 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary outcome: mean difference in SRS-2 change from baseline to post-intervention (T1): active vs sham
“mean difference in SRS-2 scores of −6.25 (95% confidence interval (CI) −8.69 to −3.81; Cohen’s d −0.92; P<0.001)”
Taken as given: The estimate is the mean difference from a linear regression model.; The 95% CI is based on a normal approximation.; The CI is two-sided and symmetric.Method: pCI function with estimate and CI, assuming normal distribution; two-tailed test.How we recomputed it: pCI(-6.25, -8.69, -3.81, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary outcome: mean difference in SRS-2 change from baseline to one-month follow-up (T2): active vs sham
“−6.17 (−8.65 to −3.70; −0.90; P<0.001)”
Taken as given: The estimate is the mean difference from a linear regression model.; The 95% CI is based on a normal approximation.; The CI is two-sided and symmetric.Method: pCI function with estimate and CI, assuming normal distribution; two-tailed test.How we recomputed it: pCI(-6.17, -8.65, -3.70, 0) - CONSISTENTreported p = .090 · recomputed p = .087Reviewer 2Masking integrity: proportion of caregivers who believed their child received active treatment, active vs sham group
“Masking was effective, with 82.8% (82/99) of caregivers in the active group and 72.7% (72/99) in the sham group believing their child received active treatment (P=0.09).”
Taken as given: 82 and 72 are the counts of caregivers believing active treatment in the active (n=99) and sham (n=99) groups; the non-event cells are derived as 99−82=17 and 99−72=27; the reported P=0.09 corresponds to a two-tailed Pearson chi-square test on the 2×2 table (df=1)Method: Pearson chi-square (two-tailed, df=1) from cell counts; recomputed ≈0.087, consistent with reported 0.09.How we recomputed it: pChi2x2(82,17,72,27) - CONSISTENTreported p = .005 · recomputed p = .005Reviewer 2Adverse event: restlessness frequency, active vs sham group at n=99 each
“| Restlessness | 39 (39.4) | 21 (21.2) | 0.005 |”
Taken as given: 39 and 21 are the event counts in the active (n=99) and sham (n=99) groups (n from Table 3 header); non-event cells derived as 99−39=60 and 99−21=78; the reported P=0.005 is a two-tailed Pearson chi-square (df=1) without continuity correctionMethod: Pearson chi-square (two-tailed, df=1) from cell counts; recomputed ≈0.0054, consistent with 0.005.How we recomputed it: pChi2x2(39,60,21,78) - CONSISTENTreported p = .010 · recomputed p = .014Reviewer 2Adverse event: scalp discomfort frequency, active vs sham group
“| Scalp discomfort | 15 (15.2) | 4 (4.0) | 0.01 |”
Taken as given: 15 and 4 are event counts in the active (n=99) and sham (n=99) groups; non-event cells derived as 99−15=84 and 99−4=95; reported P=0.01 is a two-tailed Fisher's exact test (consistent with the AE table's method)Method: Fisher's exact test, two-tailed, from cell counts; recomputed ≈0.012, rounds to 0.01.How we recomputed it: pFisher2x2(15,84,4,95,0) - CONSISTENTreported p = .680 · recomputed p = .683Reviewer 2Adverse event: dizziness frequency, active vs sham group
“| Dizziness | 4 (4.0) | 2 (2.0) | 0.68 |”
Taken as given: 4 and 2 are event counts in the active (n=99) and sham (n=99) groups; non-event cells derived as 99−4=95 and 99−2=97; reported P=0.68 is a two-tailed Fisher's exact testMethod: Fisher's exact test, two-tailed; recomputed ≈0.68, matches.How we recomputed it: pFisher2x2(4,95,2,97,0) - CONSISTENTreported p > .990 · recomputed p = 1.000Reviewer 2Adverse event: poor sleep frequency, active vs sham group
“| Poor sleep | 1 (1.0) | 2 (2.0) | >0.99 |”
Taken as given: 1 and 2 are event counts in the active (n=99) and sham (n=99) groups; non-event cells derived as 99−1=98 and 99−2=97; reported '>0.99' corresponds to a two-tailed Fisher's exact p≈1.0Method: Fisher's exact test, two-tailed; recomputed ≈1.0, consistent with '>0.99'.How we recomputed it: pFisher2x2(1,98,2,97,0) - CONSISTENTreported p > .990 · recomputed p = 1.000Reviewer 2Adverse event: transient limb spasm frequency, active vs sham group
“| Transient limb spasm | 1 (1.0)* | 0 (0) | >0.99 |”
Taken as given: 1 and 0 are event counts in the active (n=99) and sham (n=99) groups; non-event cells derived as 99−1=98 and 99−0=99; reported '>0.99' corresponds to a two-tailed Fisher's exact p≈1.0Method: Fisher's exact test, two-tailed; recomputed ≈1.0, consistent with '>0.99'.How we recomputed it: pFisher2x2(1,98,0,99,0)
- lowinternal contradictionTable 3 reports 2 tinnitus events in the sham group but a percentage of 0; 2/99 = 2.0%, so '(0)' is a presentation error.
“| Tinnitus | 0 (0) | 2 (0) | 0.50 |”
Table 3Find in source
Overstated conclusions
2 findings · worst mediumConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
11 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated).
- overstatedReviewer 2This protocol represents a major advancement towards equitable autism care worldwide.The claim reaches beyond the evidence: a single 5-day trial with one-month follow-up in a predominantly male sample, with a post hoc MCID, cannot substantiate a global equity claim.Evidence: The trial shows efficacy and feasibility in a mostly male (83.5%) sample with one-month follow-up; limitations acknowledge expectancy bias and lack of longer-term data.
“By addressing key limitations of conventional rTMS, this protocol represents a major advancement towards equitable autism care worldwide.”
ConclusionFind in source - partialReviewer 2a-cTBS is a viable and scalable therapeutic option for children with ASD.Efficacy and feasibility are supported, but 'scalable' is inferred from protocol simplicity and durability beyond one month is untested.Evidence: Primary efficacy plus 96.5% completion of the five-day course and no neuronavigation requirement; follow-up only to one month.
“The five day a-cTBS protocol targeting the left M1 provides a feasible, effective, and scalable therapeutic option for children with ASD, including those with intellectual disability.”
ConclusionFind in source - partialReviewer 2The average treatment effects of a-cTBS exceeded the estimated MCID, suggesting its clinical importance.Effects (−6.25/−6.17) exceed the estimated MCID (5.61), but the MCID is an acknowledged post hoc anchor-based estimate, and the margin is small.Evidence: Effects compared against 'an estimated MCID of 5.61 points' derived post hoc from three prior trials (n=121).
“This analysis yielded an estimated MCID of 5.61 points.”
MethodsFind in source - supportedReviewer 1The five-day a-cTBS protocol significantly improves social communication impairment in children with ASD.The primary outcome shows a statistically significant and clinically meaningful reduction in SRS-2 scores compared to sham, with robust sensitivity analyses.Evidence: Primary outcome results: mean difference −6.25 (95% CI −8.69 to −3.81, P<0.001) at T1 and −6.17 (95% CI −8.65 to −3.70, P<0.001) at T2, both exceeding the estimated MCID of 5.61.
We observed a significant treatment effect of a-cTBS on social communication impairment, with a mean difference in SRS-2 scores of −6.25 (95% confidence interval (CI) −8.69 to −3.81; Cohen’s d −0.92; P<0.001) at T1 and −6.17 (−8.65 to −3.70; −0.90; P<0.001) at T2.
Resultsreviewer’s wording - supportedReviewer 1The protocol improves language abilities in children with ASD.The MAIN test showed significant improvements in narrative production and comprehension, though other language measures (PPVT, CCDI) did not show significant differences.Evidence: Table 2: MAIN production and comprehension scores show significant improvements with Cohen's d 0.12-0.47 and all P<0.02.
Among the 119 participants (60.1%) with sufficient expressive language ability who completed the MAIN test, those in the a-cTBS group showed greater improvements in narrative production and comprehension (Cohen’s d 0.12-0.47; all P<0.02).
Table 2reviewer’s wording - supportedReviewer 1The protocol is feasible and safe in children with ASD, including those with intellectual disability.High adherence (96.5% completed), mild to moderate adverse events, and only one discontinuation related to intervention support feasibility and safety.Evidence: Results: 193/200 completed the intervention; adverse events were mild to moderate and resolved spontaneously.
“All adverse events were mild except for one moderate event; all resolved spontaneously.”
ResultsFind in source - supportedReviewer 1The protocol is effective in young children and those with intellectual disability.Subgroup analyses show consistent treatment effects across age and intellectual disability subgroups, with no significant interaction.Evidence: Subgroup analyses: Interaction tests showed no evidence of effect modification by intellectual disability status or age group at either follow-up visit.
Interaction tests showed no evidence of effect modification by intellectual disability status or age group at either follow-up visit.
Resultsreviewer’s wording - supportedReviewer 2A five-day a-cTBS protocol targeting the left primary motor cortex significantly improved social communication in children with autism spectrum disorder.The primary outcome was significantly improved at both time points with CIs excluding zero, directly supporting the claim.Evidence: Mean difference in SRS-2 of −6.25 (95% CI −8.69 to −3.81; P<0.001) at T1 and −6.17 (−8.65 to −3.70; P<0.001) at T2, in the mITT population.
“the a-cTBS group showed significantly greater reductions in SRS-2 scores post-intervention (−6.25, 95% confidence interval −8.69 to −3.81; Cohen’s d −0.92; P<0.001) and at one month follow-up (−6.17, −8.65 to −3.70; −0.90; P<0.001)”
AbstractFind in source - supportedReviewer 2The protocol proved extremely feasible, with high adherence rates (96% active, 97% sham).The reported 193/200 (96.5%) completion of the full intervention supports the feasibility claim.Evidence: 193 participants (96.5%) completed the full five-day intervention course.
“The protocol proved extremely feasible, with high adherence rates (96% active, 97% sham).”
DiscussionFind in source - supportedReviewer 2Safety profiles were favourable, with only mild to moderate adverse events typical of rTMS.The adverse-event table shows all events mild except one moderate transient limb spasm, all resolving spontaneously.Evidence: Table 3 adverse events; one moderate right upper limb spasm that resolved; all others mild and resolved without intervention.
“All adverse events were mild except for one moderate event; all resolved spontaneously.”
Table 3Find in source - supportedReviewer 2Treatment effects were consistent across subgroups, including young children and those with intellectual disability.Interaction tests showed no effect modification by intellectual disability status or age group, supporting consistency, though interaction power is limited.Evidence: Interaction tests showed no evidence of effect modification by intellectual disability status or age group at either follow-up visit.
“Interaction tests showed no evidence of effect modification by intellectual disability status or age group at either follow-up visit”
ResultsFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary outcome is the SRS-2 total score, a validated caregiver-reported scale for social communication impairment in ASD. The paper further anchors the effect to an estimated minimal clinically important difference, providing a clinical interpretation.
“The primary outcome was defined as the change in the total score of the Social Responsiveness Scale, second edition (SRS-2, school age version) from T0 to T1 and from T0 to T2.”
- ADEQUATEEffect sizeThe treatment effect (mean difference −6.25 and −6.17 on SRS-2) is statistically significant and exceeds the estimated MCID of 5.61, providing an explicit anchor for clinical meaningfulness.
“The average treatment effects were higher than the estimated MCID on SRS-2 (−5.61 points, supplementary methods).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
2 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Statistical reporting gaps (tests, assumptions, effect sizes)Assessed
- Data/code availability incompleteAssessed
The Introduction cites prior rTMS studies in ASD, notes their limitations (exclusion of young children and those with intellectual disability, need for neuronavigation, long treatment duration), and describes how the a-cTBS protocol addresses these gaps. The rationale is logically linked to the objectives.
“Current empirical evidence remains limited and inconclusive.”
“To address these challenges, we developed and evaluated a rTMS protocol adapted for young children with ASD and children with co-occurring intellectual disability.”
“current empirical evidence remains limited and inconclusive”
“Accumulating evidence highlighted the roles of M1 in regulating action understanding, language processing, and social emotional functions—domains frequently affected in autism.”
Randomization was stratified by IQ and site using block randomization (block length 4) generated by an independent coordinator. Participants and evaluators were masked; sham coils were indistinguishable. Power analysis was based on a pilot trial (α=0.025, 80% power, 10% attrition), yielding 186 needed; 200 enrolled. Inclusion and exclusion criteria were pre-specified. The modified intention-to-treat analysis and multiple imputation for missing data are clearly defined.
“Participants and evaluators were masked to interventions.”
“using a block randomisation sequence (block length=4) generated by an independent coordinator in SPSS version 25.0”
“Using a two sided α=0.025 to account for two primary outcomes, and assuming 80% power and 10% attrition, the required sample size was 186 participants.”
“sham coils were used that were visually indistinguishable from the active coils”
Table 1 reports sex, age, IQ, ADOS scores, comorbidities, medications, and site distribution. Sex is reported for both groups. Age is given as mean and SD. Health status is captured through intellectual disability, ADHD, medication use, and behavioural interventions.
“200 eligible participants (mean age 6.5±1.6 years; 83.5% male)”
The trial protocol was approved by the ethics committees of all three participating sites, with a protocol number provided (XHEC-C-2023-043-4). Written informed consent was obtained from legal guardians. The study was conducted in accordance with Good Clinical Practice and the Declaration of Helsinki.
“The trial protocol was approved by the ethics committees of Xinhua Hospital affiliated to Shanghai Jiao Tong University School of Medicine (XHEC-C-2023-043-4) and each participating site in accordance with the Declaration of Helsinki.”
“Written informed consent was obtained from the legal guardians of all participants after a thorough in-person explanation of the study procedures.”
“The study was conducted in accordance with the Good Clinical Practice and the principles of the Declaration of Helsinki.”
“The trial protocol was approved by the ethics committees of Xinhua Hospital affiliated to Shanghai Jiao Tong University School of Medicine (XHEC-C-2023-043-4) and each participating site in accordance with the Declaration of Helsinki.”
“Written informed consent was obtained from the legal guardians of all participants after a thorough in-person explanation of the study procedures.”
“The study was conducted in accordance with the Good Clinical Practice and the principles of the Declaration of Helsinki.”
The transcranial magnetic stimulation device is identified by model and manufacturer (M-100 Ultimate, Shenzhen Yingchi Technology Co., China). Statistical software R version 4.5.2 and SPSS version 25.0 are named. No antibodies, cell lines, or animal reagents are used.
“All sites used the same model of pulsed magnetic stimulation device (M-100 Ultimate, Shenzhen Yingchi Technology Co., China).”
“All statistical analyses were performed using R version 4.5.2 (R Foundation for Statistical Computing, Vienna, Austria).”
“All sites used the same model of pulsed magnetic stimulation device (M-100 Ultimate, Shenzhen Yingchi Technology Co., China).”
“This protocol delivered a total of 1800 pulses over 120 s per session. Stimulation sessions were performed hourly, with 10 sessions per day (18 000 pulses per day) over five consecutive days, for a total of 90 000 pulses.”
“All statistical analyses were performed using R version 4.5.2 (R Foundation for Statistical Computing, Vienna, Austria).”
Most statistical reporting is adequate: tests are named (linear regression, ordinal logistic regression), effect sizes with 95% CIs are provided, software is identified, and data presentation follows clinical trial standards. Assumptions are not explicitly verified but the methods are standard. The primary outcome p-value is reported as a threshold (P<0.001) rather than an exact value, which constitutes imprecise reporting.
“mean difference in SRS-2 scores of −6.25 (95% confidence interval (CI) −8.69 to −3.81; Cohen’s d −0.92; P<0.001)”
“All statistical analyses were performed using R version 4.5.2 (R Foundation for Statistical Computing, Vienna, Austria).”
“the a-cTBS group showed significantly greater reductions in SRS-2 scores post-intervention (−6.25, 95% confidence interval −8.69 to −3.81; Cohen’s d −0.92; P<0.001)”
“| Dizziness | 4 (4.0) | 2 (2.0) | 0.68 |”
“| Tinnitus | 0 (0) | 2 (0) | 0.50 |”
The data availability statement includes a public repository URL (OSF). The data are openly accessible. However, the code is described as being in the supplementary appendix, which is not a version-controlled public repository with a permanent identifier. This meets the criterion for a data availability statement and repository deposit, but code sharing is inadequate.
“The data underlying the findings in this paper are openly and publicly available and can be found at: https://osf.io/tg9me/overview”
“The code used to analyse the data in the paper can be found in the supplementary appendix.”
“The data underlying the findings in this paper are openly and publicly available and can be found at: https://osf.io/tg9me/overview”
“The code used to analyse the data in the paper can be found in the supplementary appendix.”
Methods are complete and replicable. Clinical trial registration (NCT05927792) is provided. CONSORT reporting guideline is referenced. All pre-specified primary and secondary outcomes are reported, including non-significant results. Limitations are thoroughly discussed in the Discussion section. Conclusions are appropriately framed. Funding sources and competing interests are declared.
“ClinicalTrials.gov NCT05927792”
“Several limitations warrant consideration when interpreting the results.”
“was preregistered on ClinicalTrials.gov ( NCT05927792 (https://clinicaltrials.gov/ct2/show/NCT05927792) )”
“This study followed the CONSORT (consolidated standards of reporting trials) guidelines”
“No significant between group differences were observed for Vineland-3, CCDI, or PPVT scores”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 70 references by DOI: 63 verified — 1 DOI unresolved, 5 no DOI (shown, not verified), 1 lookup failed.
- UNRESOLVED10.1016/s2215-0366(24Research prioritiesCited DOI does not resolve to any Crossref record.
- NOT CHECKED10.1002/aur.2954A lack of efficacy of continuous theta burst stimulation over the left dorsolateral prefrontal cortex in autism: A double blind randomized sham-controlled trial[crossref] rate_limited 429 https://api.crossref.org/works/10.1002%2Faur.2954?mailto=editorial%40alpha1science.com: HTTP 429
- NO DOIDiagnostic and statistical manual of mental disorders: DSM-5No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIManual of the Wechsler Intelligence Scale for Children-RevisedNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWechsler Preschool and Primary Scale of IntelligenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISocial Responsiveness ScaleNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISample Size Calculations in Clinical ResearchNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- dataOSFLIVEHTTP 200https://osf.io/tg9me/overviewResolves to OSF (data repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, other, typo.
- MINORconsistencyTable 3“Tinnitus | 0 (0) | 2 (0) | 0.50”→ Change to 'Tinnitus | 0 (0) | 2 (2.0) | 0.50'The count is 2 but the percentage is 0, which is inconsistent.
- MINORotherData Availability Statement“https://osf.io/tg9me/overview (10.1177/1362361320967790)”→ Remove the DOI if it is not associated with the dataset, or clarify its purpose.The DOI appears to be from a different article and may be a formatting error.
- MINORconsistencyTable 3, Tinnitus row“| Tinnitus | 0 (0) | 2 (0) | 0.50 |”→ Change '2 (0)' to '2 (2.0)' since 2/99 = 2.0%.The sham group has 2 tinnitus events but the percentage is printed as 0, which is internally inconsistent.
- MINORotherData Availability Statement / Associated Data“https://osf.io/tg9me/overview (10.1177/1362361320967790)”→ Remove the stray DOI '10.1177/1362361320967790' or replace it with the actual OSF DOI for the dataset.The parenthetical DOI appears to belong to a journal article, not the OSF data repository.
- MINORtypoMethods, Statistical analysis“These modifications were summarised at the end of the statistical analysis plan appendixand did not alter the conclusions”→ Insert a space: 'plan appendix and did not alter'.Missing space in 'appendixand'.
Published work is methodologically robust. An informed reader should weigh the imprecise p-value reporting (threshold only) and the code availability (appendix only); these do not undermine the validity of the findings but warrant a correction for the exact p-values and a repository deposit for the code. The overstated claim about global equity should be tempered in any future correspondence.
- 1.HIGHreportingReport exact p-values (e.g., P=0.0002) for the two primary outcomes instead of the P<0.001 thresholds in the Abstract and Results (Primary efficacy outcome) section.Exact p-values enable precise interpretation and are required for transparency; threshold-only reporting is a common journal critique.
- 2.HIGHdata codeDeposit the analysis code in a version-controlled public repository (e.g., GitHub, Zenodo) with a persistent identifier and update the Data Availability Statement accordingly.Code in a supplementary appendix is not version-controlled and lacks a permanent identifier, which harms reproducibility.
- 3.HIGHreportingCorrect the Table 3 tinnitus cell from '2 (0)' to '2 (2.0)' in the sham group.The percentage is internally inconsistent with the count (2/99 = 2.0%); this is a data presentation error that could mislead readers.
- 4.HIGHreportingRemove the stray DOI '10.1177/1362361320967790' from the Data Availability Statement, or replace it with the actual OSF DOI for the dataset.The DOI appears to belong to a different article and may confuse readers about the correct data repository.
- 5.HIGHotherTemper the claim 'This protocol represents a major advancement towards equitable autism care worldwide' to reflect the single 5-day trial with one-month follow-up in a predominantly male sample.The claim overstates the evidence; a single trial cannot substantiate a global equity claim.
- 6.HIGHotherVerify the reference 'Research priorities' (DOI 10.1016/s2215-0366(24) that was not found in any registry; if it cannot be confirmed, correct or remove it.A reference that cannot be located in Crossref/OpenAlex may be fabricated or contain a typo; it should be verified to maintain integrity.
- 7.MEDIUMcopyeditInsert a space in 'appendixand' in the Methods, Statistical analysis section.The typo 'appendixand' should be 'appendix and' for readability.
- 8.MEDIUMreportingAdd a persistent DOI for the OSF data repository alongside the URL in the Data Availability Statement.A DOI provides a permanent and citable link to the dataset, improving discoverability and reproducibility.
- 9.MEDIUMstatisticsState that assumptions for linear regression (e.g., normality of residuals, homoscedasticity) were checked, or justify why they are considered satisfied, in the Statistical analysis section.Explicit assumption verification strengthens the validity of the analysis and addresses a common reviewer expectation.
- 10.MEDIUMreportingConsider adding a statement about the minimal clinically important difference (MCID) in the Abstract or main results to contextualize the primary outcome effect size.MCID provides clinical context for the observed effect size, aiding interpretation of the SRS-2 change.
- 11.LOWreportingProvide a rationale for the choice of Cohen's d interpretation (e.g., small, medium, large) in the context of ASD research to aid clinical interpretation.Standard effect size benchmarks may not be appropriate for ASD outcomes; a domain-specific rationale would strengthen the interpretation.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.