Tai chi or cognitive behavioural therapy for treating insomnia in middle aged and older adults: randomised non-inferiority trial.
Siu PM, Yu DJ, Yu AP, Recchia F, Li SX, Chan RN, Fong DY, Chan DK, Hui SS, Chung KF, Woo J, Wang C, Irwin MR
- DOI
- 10.1136/bmj-2025-084320
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/61e562dc-9682-4c89-8b96-97306daddef6 is authoritative.
How this rating was calculated
- ReportingStatistical analysis partially met−0.25★
- ReportingData & code availability partially met−0.25★
- No reported statistical tests were found to recompute.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 14 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper is a well-conducted randomised non-inferiority trial with strong reporting transparency, ethical compliance, and a clear scientific premise. The main weaknesses are imprecise p-value reporting in secondary analyses (thresholds instead of exact values) and analysis code shared only in supplementary files rather than a version-controlled repository.
Evaluated using the eight-dimension rigor framework. The key resources dimension was scored not applicable because the study is a behavioural intervention with no biological or chemical resources. The two independent reviewers largely agreed; the divergence on statistical analysis was resolved by weighting the specific evidence of threshold p-values. The statistics verification component found no recomputable tests (coverage limited to tests with full statistics; primary analysis is by estimation). No citation concerns were flagged.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
8 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewers 1, 2The findings support the use of tai chi as an alternative approach for the long term management of chronic insomnia in middle aged and older adults.The month-15 non-inferiority supports a long-term alternative role, but the claim rests on a single-site trial showing inferiority at month 3 and a follow-up partially confounded by continued practice; the conclusion is appropriately hedged but goes somewhat beyond the non-inferiority-at-one-timepoint result.Evidence: Non-inferiority at month 15 only; 31 (36.5%) of 85 tai chi participants continued practice during follow-up.
“This finding supports the use of tai chi as an alternative approach for the long term management of chronic insomnia in middle aged and older adults.”
ConclusionFind in source - supportedReviewer 1Tai chi was inferior to CBT-I at month 3 because the upper confidence limit of the between-group difference exceeded the non-inferiority margin.The reported between-group difference of 4.52 with upper confidence limit 5.81 exceeds the pre-specified 4-point margin, directly supporting inferiority.Evidence: Between-group difference 4.52 (−∞ to 5.81) at month 3, with upper limit 5.81 > margin of 4.
“Tai chi was deemed inferior to CBT-I at month 3 because the upper confidence limit exceeded the non-inferiority margin.”
AbstractFind in source - supportedReviewer 1Tai chi was non-inferior to CBT-I at month 15.The month-15 between-group difference of 0.68 with upper confidence limit 2.00 falls within the 4-point margin, supporting non-inferiority.Evidence: Between-group difference 0.68 (−∞ to 2.00) at month 15, with upper limit 2.00 < margin of 4.
“At this point, tai chi was considered non-inferior to CBT-I because the upper confidence limit fell within the non-inferiority margin.”
AbstractFind in source - supportedReviewer 1Results from the intention-to-treat analysis were consistent with the per protocol findings.The ITT multiple-imputation analysis reproduces the same pattern (inferior at month 3, non-inferior at month 15).Evidence: ITT MI: month-3 difference 3.85 (−∞ to 5.46) inferior; month-15 difference 0.71 (−∞ to 2.28) non-inferior.
“Results of the intention-to-treat analysis were consistent with the per protocol findings.”
ResultsFind in source - supportedReviewer 1No adverse events occurred during the intervention period.The paper explicitly states no adverse events were observed in either group.Evidence: Statement in Results and Conclusion sections.
“No adverse events were observed in either group during the intervention period.”
ResultsFind in source - supportedReviewer 2Tai chi was inferior to CBT-I at month 3 but non-inferior at month 15.The non-inferiority analysis directly supports this claim: at month 3 the upper CI (5.81) exceeded the 4-point margin, and at month 15 the upper CI (2.00) fell within it.Evidence: Primary outcome per-protocol analysis: between-group difference 4.52 (−∞ to 5.81) at month 3 and 0.68 (−∞ to 2.00) at month 15.
Tai chi was deemed inferior to CBT-I at month 3 because the upper confidence limit exceeded the non-inferiority margin. ... At this point, tai chi was considered non-inferior to CBT-I because the upper confidence limit fell within the non-inferiority margin.
Abstractreviewer’s wording - supportedReviewer 2Tai chi and CBT-I had comparable benefits on subjective sleep parameters, quality of life, mental health, and physical activity level.No statistically significant group-by-time interaction effects were observed across the secondary outcomes, supporting comparable benefits.Evidence: Tables 2 and 3 report non-significant group-by-time interaction P values for all secondary outcomes.
No statistically significant group by time interaction effects were observed in all secondary outcomes, indicating that there were no between group differences in the changes in these outcomes over time between the tai chi and CBT-I groups.
Conclusionreviewer’s wording - supportedReviewer 2The observed effect size of the tai chi intervention on ISI was large (Cohen's d 1.50) and comparable to previous studies.Cohen's d of 1.50 for the tai chi group at month 3 is presented in Table 2 and consistent with the reported mean change and dispersion.Evidence: Table 2 reports Cohen's d −1.50 for tai chi ISI at month 3.
“The observed effect size of the tai chi intervention on ISI was large (Cohen’s d 1.50) and comparable to the effect sizes reported in previous studies”
Table 2Find in source
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
2 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Statistical reporting gaps (tests, assumptions, effect sizes)Assessed
- Data/code availability incompleteAssessed
Prior work is cited for both the prevalence/burden of chronic insomnia and the beneficial effects of tai chi, with acknowledged limitations of previous studies (mostly passive-control or exercise-comparator designs). The rationale and non-inferiority hypothesis follow logically from the cited evidence, and the specific gap (lack of direct comparison with the first-line treatment) is identified and addressed by the study design.
“Previous clinical trials and meta-analyses support the beneficial effects of tai chi in middle aged and older adults with insomnia, with improvements that can be sustained for up to 24 months.”
“However, most studies used passive control designs (eg, stretching) or compared it with other exercise modalities (eg, aerobic exercise), while direct comparisons with first line insomnia treatments (eg, CBT-I) in middle aged and older adults with primary chronic insomnia are currently lacking.”
“We hypothesised that tai chi is non-inferior to CBT-I for treating insomnia in middle aged and older adults, and sustaining the sleep improvements 12 months after the intervention.”
“However, most studies used passive control designs (eg, stretching) or compared it with other exercise modalities (eg, aerobic exercise), while direct comparisons with first line insomnia treatments (eg, CBT-I) in middle aged and older adults with primary chronic insomnia are currently lacking.”
“We hypothesised that tai chi is non-inferior to CBT-I for treating insomnia in middle aged and older adults, and sustaining the sleep improvements 12 months after the intervention.”
“Previous clinical trials and meta-analyses support the beneficial effects of tai chi in middle aged and older adults with insomnia, with improvements that can be sustained for up to 24 months.”
Randomisation used a sealed-envelope online generator with block sizes of four to six at the individual-participant level. Assessor blinding is described with an explicit rationale for not blinding participants/instructors. The a priori sample-size calculation specifies 95% power, a 0.025 alpha, a 4-point non-inferiority margin, and a 20% dropout allowance. Inclusion/exclusion criteria are detailed, and the per-protocol/ITT populations plus LOCF and multiple-imputation sensitivity analyses are pre-specified. The bench-science sub-criteria (replicate_distinction, controls, independent_replication) are not applicable to a single pivotal human trial.
“Randomisation was conducted using an online random generator ( https://www.sealedenvelope.com/ ) with block sizes of four to six.”
“Outcome assessors were blinded to the group allocation and participants were instructed not to disclose their group allocation to the outcome assessors during the outcome assessments.”
“The sample size estimation was based on a primary comparison of perceived insomnia severity measured by ISI between the two groups using a 95% power and 0.05/2=0.025 maximum chance of committing false type I errors to account for multiplicity.”
“Randomisation was conducted using an online random generator ( https://www.sealedenvelope.com/ ) with block sizes of four to six.”
“Outcome assessors were blinded to the group allocation and participants were instructed not to disclose their group allocation to the outcome assessors during the outcome assessments.”
“The sample size estimation was based on a primary comparison of perceived insomnia severity measured by ISI between the two groups using a 95% power and 0.05/2=0.025 maximum chance of committing false type I errors to account for multiplicity.”
Table 1 reports sex (77% and 84% female), age (mean 64.83 and 63.76 years), education, marital status, income, and physical/mental comorbidities. Participants are described as 'ethnic Chinese' aged ≥50 with DSM-5 chronic insomnia. Species/strain/housing criteria are not applicable to a human trial.
“Female, n (%) | 77 (77) | 84 (84)”
“Age (years) | 64.83 (6.31) | 63.76 (6.15)”
“participants needed to be middle aged and older adults aged ≥50 years who were ethnic Chinese”
“Female, n (%) | 77 (77) | 84 (84)”
“Age (years) | 64.83 (6.31) | 63.76 (6.15)”
“participants needed to be middle aged and older adults aged ≥50 years who were ethnic Chinese and had a diagnosis of chronic insomnia”
The study involves non-public human subject-level data, so ethics applies. The IRB approval names the institution and gives a protocol number (UW 18-621). Written informed consent is explicitly described. Regulatory compliance with the Declaration of Helsinki is stated by name. All applicable sub-criteria are adequate.
“The study was approved by the University of Hong Kong/Hospital Authority Hong Kong West Institutional Review Board (IRB approval No UW 18-621).”
“Verbal and written information on the study was provided and written informed consent was obtained before baseline outcome assessments.”
“The study was conducted in accordance with the Declaration of Helsinki.”
“The study was approved by the University of Hong Kong/Hospital Authority Hong Kong West institutional review board (IRB approval No UW 18-621).”
“Written informed consent was obtained on a voluntary basis before baseline assessments.”
“The study was conducted in accordance with the Declaration of Helsinki.”
The interventions are mind-body exercise and psychotherapy; no investigational product, antibody, cell line, organism, or bench reagent is involved. The statistical software (SAS) is assessed under the statistical dimension, and no bespoke analysis code repository is pertinent to resources. Per the applicability rules, a behavioural/exercise trial with no drug, biologic, device, or bench reagent scores the dimension not_applicable.
Named tests (GEE, logistic regression) and software (SAS OnDemand) are reported, and the clinical-trial idiom for data presentation (flow chart, per-group n, CIs, Cohen's d) is met. The primary outcome is presented with effect estimates and 95% CIs as is standard for non-inferiority trials. However, many significant group/time effects in Tables 2 and 3 are reported as '<0.001*' rather than exact p-values, which is imprecise reporting and triggers a warn-level rating. No arithmetically implausible summary statistics were found; baseline percentages in Table 1 sum correctly.
“We used generalised estimating equations analyses to examine the treatment effects on the quantitative secondary outcomes, adjusting for baseline values. Pairwise comparisons were performed using linear contrasts.”
“the tai chi group showed a reduction of 6.67 points (95% confidence interval 5.61 to 7.73) in ISI scores”
“All statistical analyses were performed using SAS OnDemand for Academics (SAS Institute).”
“We used generalised estimating equations analyses to examine the treatment effects on the quantitative secondary outcomes, adjusting for baseline values.”
“At month 3, the tai chi group showed a reduction of 6.67 points (95% confidence interval 5.61 to 7.73) in ISI scores, while the CBT-I group had a reduction of 11.19 (10.06 to 12.32), resulting in a between group difference of 4.52 (−∞ to 5.81).”
The data-availability statement is concrete, naming an open figshare DOI (10.6084/m9.figshare.29203967.v3), so data_availability_statement and repository_deposit are adequate. The code is described as available in the supplementary files and from the corresponding author, which falls short of the 'version-controlled public repo with a permanent identifier' standard, so code_sharing is reported_but_inadequate. Accession_numbers is not applicable for a clinical dataset without sequencing. With 2 of 3 applicable criteria adequate, the dimension is rated warn.
“The data underlying the findings in this paper are openly and publicly available and can be found here: https://doi.org/10.6084/m9.figshare.29203967.v3”
“The code used to analyse the data in the paper can be found in the supplementary files.”
“The data underlying the findings in this paper are openly and publicly available and can be found here: https://doi.org/10.6084/m9.figshare.29203967.v3”
“The code used to analyse the data in the paper can be found in the supplementary files.”
The trial is prospectively registered (NCT04384822) and reports adherence to the CONSORT non-inferiority extension. Methods are detailed enough to replicate (session structure, instructor qualifications, scales). Negative/null results are reported (no group-by-time interactions). Conclusions are appropriately hedged—acknowledging inferiority at month 3 while claiming non-inferiority at month 15. Funding sources and a competing-interest statement are provided.
“This randomised, assessor blinded, non-inferiority trial adhered to the Consolidated Standards of Reporting Trials (CONSORT) extension for non-inferiority and equivalence trials.”
“This study has several potential limitations. A large proportion (77.5%, n=155) of the study participants were older adults aged ≥60 years, which might limit the generalisability of our results when considering younger populations.”
“This randomised, assessor blinded, non-inferiority trial adhered to the Consolidated Standards of Reporting Trials (CONSORT) extension for non-inferiority and equivalence trials.”
“This study was supported by General Research Fund of Research Grants Council, Hong Kong University Grants Committee (project No 17112819), and Seed Fund for Basic Research of the University of Hong Kong.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 44 references by DOI: 43 verified — 1 no DOI (shown, not verified).
- NO DOIEffectiveness of exercise, cognitive behavioral therapy, and pharmacotherapy on improving sleep in adults with chronic insomnia: a systematic review and network meta-analysis of randomized controlled trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- dataFigshareLIVEHTTP 202https://doi.org/10.6084/m9.figshare.29203967.v3Resolves to Figshare (data repository).
- datahttps://clinicaltrials.gov/ct2/show/NCT04384822LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
2 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 2 minor suggestions below.
2 copyedit issues flagged: mostly consistency.
- MINORconsistencyThroughout (Introduction, Methods)“middle aged and older adults”→ Use a consistent hyphenation, e.g., 'middle-aged and older adults,' throughout.Hyphenation is inconsistent ('middle aged' vs 'middle-aged').
- MINORconsistencyCorrespondence line vs Data availability statement“pmsiu@hku.hk (correspondence) vs pmsiu@hku.edu.hk (data availability)”→ Use a single consistent email address for the corresponding author in both the correspondence line and the data availability statement.Minor inconsistency, likely the same mailbox under different domains.
The published paper is methodologically robust. An informed reader should note that the exact p-values for significant group/time effects in Tables 2 and 3 are reported only as thresholds, and that the analysis code is available only in supplementary files rather than a version-controlled repository with a permanent identifier. These are transparency gaps that could be addressed with a correction or by depositing the code on a platform like Zenodo, but they do not undermine the core conclusions.
- 1.HIGHstatisticsProvide exact p-values (e.g., P=0.0004) for the significant group/time effects in Tables 2 and 3 instead of the threshold '<0.001', or add a footnote explaining that threshold values are used for very small p-values.Threshold p-values are imprecise reporting and prevent readers from assessing the precise strength of evidence; exact values would improve transparency and enable independent verification.
- 2.HIGHdata codeDeposit the analysis code in a version-controlled public repository (e.g., GitHub, Zenodo) with a permanent identifier (DOI) and update the Data Availability Statement to cite it.Code in supplementary files is not version-controlled and lacks a persistent identifier, reducing reproducibility and long-term accessibility.
- 3.MEDIUMreportingAdd baseline body weight and/or BMI to Table 1 as a continuous variable, not just subgroup categories.Baseline weight/BMI is a standard demographic variable; its absence is a minor gap in the reporting of participant characteristics.
- 4.MEDIUMstatisticsSpecify the version of SAS OnDemand for Academics used in the statistical analysis section.Version identification aids reproducibility in case of software changes.
- 5.MEDIUMreportingClarify the exact denominator for the insomnia remission and treatment response rates at month 3 and month 15 (e.g., completers vs. ITT) to make the reported percentages auditable.The denominator is not explicitly stated, which could confuse readers about the population used for these outcomes.
- 6.MEDIUMreportingIn the limitations, explicitly state that the non-inferiority conclusion at month 15 may be partly driven by continued self-practice of tai chi, and frame this as a potential confounder of the long-term comparison.Acknowledging this confounder strengthens the limitations discussion and helps readers interpret the long-term results.
- 7.LOWcopyeditUse consistent hyphenation for 'middle-aged' throughout the manuscript (e.g., 'middle-aged and older adults' instead of inconsistent 'middle aged').Minor consistency issue that improves readability.
- 8.LOWcopyeditUse a single consistent email address for the corresponding author in both the correspondence line and the data availability statement.Minor inconsistency that could cause confusion; ensure the same email is used throughout.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.