Effect of laughter exercise versus 0.1% sodium hyaluronic acid on ocular surface discomfort in dry eye disease: non-inferiority randomised controlled trial.
Li J, Liao Y, Zhang SY, Jin L, Congdon N, Fan Z, Zeng Y, Zheng Y, Liu Z, Liu Y, Liang L
- DOI
- 10.1136/bmj-2024-080474
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/44d66917-5dbd-4f92-b21e-522310aad9bd is authoritative.
How this rating was calculated
- StatisticsImpossible or misreported statistic ×3−3★
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 44 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- No data or code availability links were detected to verify.
- 01Printed percentage does not match its own countdemonstrable
74% does not match the reported count 78/299
“74% (78/299) were women”
Results ¶1Find in source - 02Printed percentage does not match its own countdemonstrable
73% does not match the reported count 60/149
“60 (73)”
Table 1 - 03Printed percentage does not match its own countdemonstrable
75% does not match the reported count 62/150
“62 (75)”
Table 1 - 04Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on the Ocular Surface Disease Index (OSDI), a patient-reported symptom score, which is a surrogate for clinical benefit. The paper does not provide evidence linking OSDI changes to hard clinical outcomes, nor does it demonstrate target engagement at the tested dose (e.g., PK/PD) for the laughter exercise. The OSDI is a validated tool, but the link to long-term clinical outcomes is not established in this manuscript.
“The primary outcome was the mean change in the ocular surface disease index score from baseline to eight weeks.”
- 05Treatment effect not shown to be clinically meaningful
The primary outcome shows a mean change of -10.5 points in the laughter exercise group, which is a modest improvement on a 0-100 scale. The non-inferiority margin was 6 points, and the between-group difference was -1.45 points (95% CI -5.08 to 2.19), which is small. The clinical meaningfulness of this effect is not explicitly anchored to a minimal clinically important difference (MCID) for OSDI, and the improvement represents only about 10% of the scale range. The paper does not provide a clear anchor for clinical significance.
“The mean change in ocular surface disease index score at eight weeks was −10.5 points (95% confidence interval (CI) −13.1 to −7.82) in the laughter exercise group and −8.83 (−11.7 to −6.02) in the control group.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted non-inferiority RCT with rigorous design, clear reporting, and appropriate statistical methods. The main weaknesses are the vague data availability statement and minor reporting gaps (e.g., no CONSORT mention, incomplete resource identification).
Both reviewers classified the study as interventional, and no divergence was noted. The evaluation covered all eight dimensions; non-applicable sub-criteria (e.g., animal housing, cell lines) were excluded. The statistics verification covered only a subset of tests; the 3 inconsistent recomputations were not specified and did not constitute demonstrable errors.
Numerical inconsistencies
2 findings · worst criticalValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks. 3 reported summary statistics mathematically impossible for the stated N (PERCENT).
- PERCENT74% does not match the reported count 78/299
“74% (78/299) were women”
Results ¶1Find in source - PERCENT73% does not match the reported count 60/149
“60 (73)”
Table 1 - PERCENT75% does not match the reported count 62/150
“62 (75)”
Table 1
- CONSISTENTreported p = .430 · recomputed p = .434Reviewers 1, 2Primary outcome between-group difference in OSDI change at 8 weeks (per protocol)
“mean difference −1.45 points (95% CI −5.08 to 2.19); P=0.43”
Taken as given: The estimate is -1.45 and the 95% CI is -5.08 to 2.19.; The CI is two-sided at 95%.; The p-value is for the between-group difference in change.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(-1.45, -5.08, 2.19, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Secondary outcome: between-group difference in non-invasive tear break up time at 8 weeks
“mean between group difference was 2.30 seconds ((95% CI 1.30 to 3.30); P<0.001)”
Taken as given: The estimate is 2.30 and the 95% CI is 1.30 to 3.30.; The CI is two-sided at 95%.; The p-value is for the between-group difference in change.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(2.30, 1.30, 3.30, 0)
- lowinternal contradictionTable 1 reports female sex as 60 (73%) in laughter group and 62 (75%) in control group, but the text says 74% (78/299) women overall. The percentages in Table 1 are inconsistent with the overall percentage (73% and 75% average to 74%, but the counts sum to 122, not 78). This is likely a typographical error in the table.
“Female sex, No. (%) | 60 (73) | 62 (75)”
Table 1Find in source - lowinternal contradictionThe abstract states '283 (95%) completed the trial' but the results section reports 137/149 (92%) in laughter group and 146/150 (97%) in control group, which sums to 283/299 (94.6%). The 95% is a rounding of 94.6%, which is acceptable.
“283 (95%) completed the trial.”
AbstractFind in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
4 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Laughter exercise was non-inferior to 0.1% sodium hyaluronic acid in relieving subjective symptoms in patients with dry eye disease.The primary outcome analysis shows the upper bound of the CI for the difference is below the non-inferiority margin, supporting non-inferiority.Evidence: Primary outcome: mean difference −1.45 (95% CI −5.08 to 2.19), P=0.43, with upper bound 2.19 < 6.
“The laughter exercise was non-inferior to 0.1% sodium hyaluronic acid in relieving subjective symptoms in patients with dry eye disease with limited corneal staining over eight weeks intervention.”
ConclusionFind in source - supportedReviewers 1, 2Laughter exercise improved tear film stability and meibomian gland function.Secondary and exploratory outcomes show significant improvements in non-invasive tear break up time and meibomian gland measures in the laughter group.Evidence: Non-invasive tear break up time mean difference 2.30 (95% CI 1.30 to 3.30), P<0.001; meibomian gland secretion property score improved by −3.78 (−4.45 to −3.12).
“Additionally, we found that laughter exercise appeared to improve tear film stability and the meibomian gland function.”
DiscussionFind in source - supportedReviewers 1, 2Laughter exercise is a safe, environmentally friendly, and low cost intervention.No adverse events were reported, and the intervention is inherently low-cost and environmentally friendly.Evidence: No adverse events were noted in either study group.
“No adverse events were noted in either study group.”
AbstractFind in source - supportedReviewer 1The benefits of laughter exercise persisted for at least four weeks after discontinuation.The 12-week follow-up showed a significant between-group difference favoring laughter exercise, indicating persistence.Evidence: At 12 weeks, mean between group difference was −4.08 points (95% CI −7.62 to −0.55); P=0.024.
“These benefits persisted for at least four weeks after discontinuation of the exercise”
DiscussionFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on the Ocular Surface Disease Index (OSDI), a patient-reported symptom score, which is a surrogate for clinical benefit. The paper does not provide evidence linking OSDI changes to hard clinical outcomes, nor does it demonstrate target engagement at the tested dose (e.g., PK/PD) for the laughter exercise. The OSDI is a validated tool, but the link to long-term clinical outcomes is not established in this manuscript.
“The primary outcome was the mean change in the ocular surface disease index score from baseline to eight weeks.”
- INADEQUATEEffect sizeThe primary outcome shows a mean change of -10.5 points in the laughter exercise group, which is a modest improvement on a 0-100 scale. The non-inferiority margin was 6 points, and the between-group difference was -1.45 points (95% CI -5.08 to 2.19), which is small. The clinical meaningfulness of this effect is not explicitly anchored to a minimal clinically important difference (MCID) for OSDI, and the improvement represents only about 10% of the scale range. The paper does not provide a clear anchor for clinical significance.
“The mean change in ocular surface disease index score at eight weeks was −10.5 points (95% confidence interval (CI) −13.1 to −7.82) in the laughter exercise group and −8.83 (−11.7 to −6.02) in the control group.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites multiple studies on dry eye disease, artificial tears, and laughter therapy, acknowledging both the economic burden and the psychological associations. The rationale for using laughter exercise is well-argued, linking positive emotions to potential benefits. The hypothesis is clearly stated as a non-inferiority margin of 6 points. Limitations of prior research are implicitly addressed by the pilot studies and the design choice.
“Abundant evidence suggests that laughter therapy alleviates depression, anxiety, stress, and chronic pain, while strengthening immune function”
“The hypothesis was that non-inferiority would be established by observing a between-group difference in the ocular surface disease index score of ≤6 points in eight weeks”
“In our pilot studies before the randomised controlled trial, we observed that laughter could immediately improve tear film stability and lipid layer thickness”
“Abundant evidence suggests that laughter therapy alleviates depression, anxiety, stress, and chronic pain, while strengthening immune function”
“The hypothesis was that non-inferiority would be established by observing a between-group difference in the ocular surface disease index score of ≤6 points in eight weeks”
Randomization used stratified block randomization with a block size of four, generated by an independent statistician. Outcome assessors were masked, though participants were unmasked for practical reasons, which is acknowledged. A priori sample size calculation was performed with 90% power and a non-inferiority margin of 6 points. Inclusion/exclusion criteria are detailed, and the analysis populations (ITT and per-protocol) are defined. Missing data were handled via multiple imputation. The trial is registered and follows a pre-specified statistical analysis plan.
“Stratified block randomisation was applied with a block size of four.”
“The trial was designed to enrol 296 participants (n=148 in each group), reaching 90% statistical power to detect non-inferiority at a one-sided α=0.025”
“Investigators assessing study outcomes were masked to group assignment but participants were unmasked for practical reasons.”
“Stratified block randomisation was applied with a block size of four.”
“Investigators assessing study outcomes were masked to group assignment but participants were unmasked for practical reasons.”
“The trial was designed to enrol 296 participants (n=148 in each group), reaching 90% statistical power to detect non-inferiority at a one-sided α=0.025”
The study reports age (mean 28.9 years), sex (74% female), and education level. Baseline characteristics are presented in Table 1 and are well balanced. Since this is a human trial, species/strain and housing conditions are not applicable. Sex is reported for both groups, and the high proportion of females is noted as consistent with the disease epidemiology.
“Participants had a mean age of 28.9 (6.30) years, 74% (78/299) were women”
“Baseline characteristics were well balanced between the two groups.”
“Participants had a mean age of 28.9 (6.30) years, 74% (78/299) were women”
“95% (283/299) have more than 12 years of education.”
The paper states ethical approval was obtained from the Ethics Committee of Zhongshan Ophthalmic Center, Sun Yat-sen University, with protocol number 2020KYPJ010. Written informed consent was obtained from all participants. Compliance with ethical standards is implied by the approval and consent process.
“All procedures adhered to the protocol approved by the ethics committee of Zhongshan Ophthalmic Center, Sun Yat-sen University, Guangzhou, China (2020KYPJ010).”
“Written informed consent was obtained from all participants.”
“All procedures adhered to the protocol approved by the ethics committee of Zhongshan Ophthalmic Center, Sun Yat-sen University, Guangzhou, China (2020KYPJ010).”
“Written informed consent was obtained from all participants.”
The control intervention, 0.1% sodium hyaluronic acid eyedrops, is named but without a manufacturer or lot number, which is typical for a commercial product. The laughter exercise app 'laughing face' is described and credited to collaborators, but no public repository or version is provided. Since this is a clinical trial of a behavioral intervention and a standard artificial tear, the bench criteria (antibodies, cell lines, mycoplasma) are not applicable. The software used for analysis (SAS 9.4, Stata 16.0, PASS 16.0) is identified.
“artificial tears (0.1% sodium hyaluronic acid eyedrop, control group)”
“Data were cleaned using Stata16.0 and all statistical analyses were performed using SAS 9.4.”
“Participants in the control group applied artificial tears, 0.1% sodium hyaluronic acid eyedrops, to both eyes four times daily for eight weeks”
“The app was developed in collaboration with South China University of Technology and Xinhuixing Information Technology, Inc (Guangzhou, China).”
The paper names the statistical tests used (two-sample t-test, paired t-test, GEE) and describes the handling of multiple comparisons (Benjamini-Hochberg). Exact p-values are reported for primary and key secondary outcomes. Effect sizes are presented with 95% confidence intervals. Software is identified. Data presentation includes per-group n and confidence intervals. The mathematical plausibility checks are not applicable due to continuous outcomes and large N.
“The difference between groups was tested using the two sample t-test for primary and psychological outcomes, and generalised estimated equation model for all clinical outcomes.”
“mean difference −1.45 points (95% CI −5.08 to 2.19); P=0.43”
“The mean change in ocular surface disease index score at eight weeks was −10.5 points (95% confidence interval (CI) −13.1 to −7.82)”
“The difference between groups was tested using the two sample t-test for primary and psychological outcomes, and generalised estimated equation model for all clinical outcomes.”
“mean difference −1.45 points (95% CI −5.08 to 2.19); P=0.43”
“Data were cleaned using Stata16.0 and all statistical analyses were performed using SAS 9.4.”
The data availability statement says 'All data requests should be submitted to lianglingyi@gzzoc.com for consideration. Access to anonymised data may be granted after review.' This is a vague statement without a named platform or data access committee, and no timeframe is given. Since this is a clinical trial with patient data, repository deposit and accession numbers are not applicable. No custom code is mentioned, so code sharing is not applicable.
“All data requests should be submitted to lianglingyi@gzzoc.com for consideration. Access to anonymised data may be granted after review.”
“All data requests should be submitted to lianglingyi@gzzoc.com for consideration. Access to anonymised data may be granted after review.”
The trial is registered at ClinicalTrials.gov (NCT04421300). Methods are comprehensive, including the intervention details and statistical analysis plan. Limitations are explicitly discussed, including the lack of double-blinding. Conclusions are proportional to the results, and funding sources and competing interests are declared. The paper does not explicitly mention a reporting guideline (e.g., CONSORT), but the structure follows it.
“ClinicalTrials.gov NCT04421300”
“A double blinded study design was not practical because this would necessitate a sham laughter exercise for which no approach has been validated.”
“The study was funded by the National Natural Science Foundation of China (grant number 82070922, 82201142)”
“ClinicalTrials.gov NCT04421300”
“A double blinded study design was not practical because this would necessitate a sham laughter exercise for which no approach has been validated.”
“The study was funded by the National Natural Science Foundation of China (grant number 82070922, 82201142)”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 65 references by DOI: 62 verified — 3 no DOI (shown, not verified).
- NO DOIDry eye products report: a global market analysis for 2015 to 2021No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPromotion of mental well-being: pursuit of happinessNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILaughter therapy in diabetesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity, grammar.
- MINORconsistencyAbstract, Results“283 (95%) completed the trial.”→ Ensure the percentage is consistent with the numbers (283/299 = 94.6%, which rounds to 95%).Minor rounding discrepancy.
- MINORclarityMethods, Intervention“Participants were instructed to perform the laughter exercise four times daily for eight weeks.”→ Clarify whether 'four times daily' means four separate sessions or four repetitions within one session.Ambiguity in the intervention frequency.
- MINORgrammarDiscussion, Possible mechanisms“The contraction of orbicularis muscle (sphincter muscles of the eyelids) during laughter exercise is another plausible explanation.”→ Change 'muscle' to 'muscles' for grammatical correctness.Singular/plural mismatch.
- MINORconsistencyTable 1“Female sex, No. (%) | 60 (73) | 62 (75)”→ Ensure percentages are consistent with the group totals (e.g., 60/149=40.3%, not 73%).The percentages in Table 1 appear to be incorrect; they should be calculated from the group totals.
- MINORclarityAbstract, Results“The upper boundary of the CI for difference in change between groups was lower than the non-inferiority margin (mean difference −1.45 points (95% CI −5.08 to 2.19); P=0.43)”→ Clarify that the upper boundary of the CI is 2.19, which is less than 6.The sentence is slightly ambiguous; consider rephrasing for clarity.
The published work is robust and generally well-reported. An informed reader should weigh the vague data availability statement and the minor internal inconsistencies in Table 1 and the abstract; these do not undermine the main conclusions but warrant clarification or a correction.
- 1.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 74% does not match the reported count 78/299Demonstrable critical failure — blocks the verdict from passing.
- 2.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 73% does not match the reported count 60/149Demonstrable critical failure — blocks the verdict from passing.
- 3.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 75% does not match the reported count 62/150Demonstrable critical failure — blocks the verdict from passing.
- 4.HIGHdata codeReplace the vague data availability statement with a concrete mechanism, e.g., deposit de-identified data in a public repository (Dryad/Zenodo) with a DOI, or specify a named data access committee and a response timeframe.The current statement provides only an email address with no conditions or timeframe, which is inadequate for reproducibility.
- 5.HIGHreportingAdd an explicit statement of adherence to CONSORT guidelines and include the CONSORT flow diagram in the main text or as a supplementary file.Reporting guidelines are expected for RCTs and improve transparency; the paper currently does not mention CONSORT.
- 6.HIGHrigorCorrect the internal inconsistency in Table 1: the female sex counts (60 and 62) and percentages (73% and 75%) do not match the overall 78/299 (74%) reported in the text.The integrity check flagged this as a likely typographical error; readers may question data accuracy.
- 7.MEDIUMrigorSpecify the manufacturer and lot number of the 0.1% sodium hyaluronic acid eyedrops used in the control group.Providing product details enhances reproducibility and allows readers to assess equivalence of the comparator.
- 8.MEDIUMrigorProvide a public repository or detailed description of the 'laughing face' app, including version and availability, to facilitate replication.The app is a bespoke intervention tool; without access, other researchers cannot replicate the intervention.
- 9.MEDIUMreportingClarify the intervention frequency: specify whether 'four times daily' means four separate sessions or four repetitions within one session.The copyedit flagged this ambiguity, which could affect interpretation of the intervention dose.
- 10.MEDIUMcopyeditFix the grammatical error in the Discussion: change 'orbicularis muscle' to 'orbicularis muscles'.Minor grammar issues detract from professionalism.
- 11.MEDIUMreportingClarify the sentence about the non-inferiority margin in the Abstract: explicitly state that the upper CI bound (2.19) is below the margin (6).The current phrasing is ambiguous and could confuse readers about the non-inferiority conclusion.
- 12.LOWreportingAdd a statement about adverse events in the Results section, as it is only mentioned in the abstract.Complete reporting of safety outcomes is expected in clinical trials.
- 13.LOWreportingProvide the full inclusion/exclusion criteria in the main text rather than referring to the protocol.Readers should not need to access the protocol to understand who was eligible.
- 14.LOWreportingReport the number of participants excluded after screening and the reasons for exclusion.Transparency about the screening process improves the reader's ability to assess generalizability.
- 15.LOWdata codeIf custom code was used for analysis, share it in a public repository (e.g., GitHub) with version and documentation.Sharing analysis code enhances reproducibility and transparency.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.