Cognitive behavioral therapy skills via a smartphone app for subthreshold depression among adults in the community: the RESiLIENT randomized controlled trial.
Furukawa TA, Tajika A, Toyomoto R, Sakata M, Luo Y, Horikoshi M, Akechi T, Kawakami N, Nakayama T, Kondo N, Fukuma S, Kessler RC, Christensen H, Whitton A, Nahum-Shani I, Lutz W, Cuijpers P, Wason JMS, Noma H
- DOI
- 10.1038/s41591-025-03639-1
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/dc9fb81a-9dad-430c-9f20-fe38153a5203 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingEthical approvals partially met−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is the change in PHQ-9 score, a self-reported symptom severity scale, which is a surrogate for clinical depression. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure) nor cite validated evidence linking PHQ-9 change to hard clinical outcomes such as prevention of major depressive episodes. The claim of efficacy rests on this surrogate without such validation.
“The primary outcome is the change in PHQ-9 score from baseline to week 6.”
- 02Treatment effect not shown to be clinically meaningful
The reported effect sizes (SMDs) range from -0.16 to -0.67, which are small to moderate. The paper does not anchor these to a minimal clinically important difference (MCID) for PHQ-9 or to clinical meaningfulness. The absolute change in PHQ-9 scores (e.g., -1.07 to -1.29 points) is small relative to the baseline mean of 8.07, and no MCID is cited.
“with effect sizes ranging between −0.67 (95% confidence interval: −0.81 to −0.53) and −0.16 (−0.30 to −0.02) for changes in depressive symptom severity from baseline to week 6, as measured with the Patient Health Questionnaire-9 scores.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed and well-reported randomized controlled trial with a strong scientific premise, rigorous statistical methods, and transparent data/code sharing. The main weakness is the absence of an explicit ethics approval statement and insufficient detail on informed consent and randomization method, which are fixable reporting gaps.
Both reviewers agreed on all dimensions; no divergence to reconcile. The study is an interventional RCT; non-applicable sub-criteria (e.g., animal housing, cell line authentication) were excluded. The statistics verification covered only a subset of tests (2 recomputed consistently); the rest are unverified but not flagged as problematic.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Check p-value for BA main effect in Trial 1 (SMD -0.38, 95% CI -0.48 to -0.27)
“The standardized mean difference (SMD) was −0.38 (95% confidence interval (CI): −0.48 to −0.27, P = 5.3 × 10 −13 ) for BA”
Taken as given: The effect size is a standardized mean difference (continuous outcome).; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed two-sided p-value from the reported SMD and 95% CI using the normal approximation.How we recomputed it: pCI(-0.38, -0.48, -0.27, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Check p-value for CR main effect in Trial 1 (SMD -0.27, 95% CI -0.37 to -0.16)
“−0.27 (−0.37 to −0.16, P = 2.9 × 10 −7 ) for CR”
Taken as given: The effect size is a standardized mean difference.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed two-sided p-value from the reported SMD and 95% CI using the normal approximation.How we recomputed it: pCI(-0.27, -0.37, -0.16, 0)
- lowinternal contradictionThe abstract states effect sizes ranging between −0.67 and −0.16, but the text in the Discussion says 'effect sizes ranging between 0.67 and 0.16' (positive values). This is likely a sign error in the Discussion.
“with effect sizes ranging between 0.67 and 0.16 for the primary outcome of depression at week 6.”
Discussion ¶1Find in source - lowinternal contradictionIn Table 5 footnote, the component for BA+BI is incorrectly described as 'ns + ba + at' instead of 'ns + ba + bi'.
“BA + BI of ns + ba + at”
Table 5Find in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
6 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2All included CBT skills and their combinations differentially beat all three control conditions.The primary and sensitivity analyses show all skills superior to delayed treatment, and sensitivity analyses with stricter controls show superiority for most comparisons.Evidence: Table 2 and Extended Data Tables 1-2 show SMDs with CIs below zero for all skills against all controls.
“The study showed that all included CBT skills and their combinations differentially beat all three control conditions of delayed treatment, health information or self-check”
AbstractFind in source - supportedReviewers 1, 2The specific efficacies of individual CBT skills were demonstrated.The factorial design and component analyses provide direct evidence for skill-specific effects.Evidence: Table 2 and Table 5 show significant main effects for each skill.
“our study demonstrated specific efficacies of individual cognitive or behavioral skills directly in a single randomized trial”
DiscussionFind in source - supportedReviewers 1, 2Combining two skills did not double the efficacy due to antagonistic interactions.Interaction terms were significant and effect sizes for combinations were not additive.Evidence: Table 3 shows significant interaction p-values and smaller than additive SMDs.
“Administering two skills did not double the efficacy of providing one skill because the interaction was antagonistic.”
ResultsFind in source - supportedReviewer 1The waiting list control overestimates intervention efficacy.The comparison of control conditions shows delayed treatment yields larger effect sizes than health information or self-check.Evidence: Table 4 shows differences between controls.
“the waiting list overestimates the efficacy of the intervention by SMDs of −0.35 (−0.50 to −0.21) or −0.18 (−0.33 to −0.04)”
DiscussionFind in source - supportedReviewers 1, 2The results contrast with the Dodo bird verdict.Differential efficacies among skills and between one- and two-skill interventions contradict the verdict.Evidence: Tables 2-4 show varying effect sizes across skills.
Our results contrast with the so-called Dodo bird verdict, which claims that all bona fide psychotherapies have comparable efficacies.
Discussionreviewer’s wording - supportedReviewer 2The waiting list control overestimates efficacy and should not be used in confirmatory trials.The study directly compared waiting list to other controls and found it produced larger effect sizes.Evidence: Table 4 shows delayed treatment (waiting list) yields larger SMDs than health information or self-check.
“In comparison with these, the waiting list overestimates the efficacy of the intervention by SMDs of −0.35 (−0.50 to −0.21) or −0.18 (−0.33 to −0.04) respectively, and should not be used in confirmatory trials of interventions”
Discussion ¶7Find in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is the change in PHQ-9 score, a self-reported symptom severity scale, which is a surrogate for clinical depression. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure) nor cite validated evidence linking PHQ-9 change to hard clinical outcomes such as prevention of major depressive episodes. The claim of efficacy rests on this surrogate without such validation.
“The primary outcome is the change in PHQ-9 score from baseline to week 6.”
- INADEQUATEEffect sizeThe reported effect sizes (SMDs) range from -0.16 to -0.67, which are small to moderate. The paper does not anchor these to a minimal clinically important difference (MCID) for PHQ-9 or to clinical meaningfulness. The absolute change in PHQ-9 scores (e.g., -1.07 to -1.29 points) is small relative to the baseline mean of 8.07, and no MCID is cited.
“with effect sizes ranging between −0.67 (95% confidence interval: −0.81 to −0.53) and −0.16 (−0.30 to −0.02) for changes in depressive symptom severity from baseline to week 6, as measured with the Patient Health Questionnaire-9 scores.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Ethics/consent reporting incompleteAssessed
The introduction cites numerous studies on subthreshold depression, CBT, and iCBT, and clearly identifies the gap: unknown contribution of individual CBT skills. The rationale for the master protocol and factorial design is well explained. Limitations of prior research (e.g., small dismantling studies) are acknowledged and addressed by the large sample size and factorial design.
“Subthreshold depression, or mild depression falling short of the diagnostic threshold for major depressive disorder, is prevalent, persistent and disabling.”
“Subthreshold depression, or mild depression falling short of the diagnostic threshold for major depressive disorder, is prevalent, persistent and disabling.”
Randomization method is not explicitly described in the text (likely in protocol), but the trial is described as randomized with equal allocation. Blinding is not applicable for a behavioral intervention where participants know their allocation; the paper does not mention blinding of outcome assessors, but outcomes are self-reported. Power analysis is met (preplanned sample size achieved). Inclusion/exclusion criteria are implied (subthreshold depression, PHQ-9 >4). Outlier handling is addressed via prespecified exclusion of shift workers in trial 4. Controls are well-defined (delayed treatment, health information, self-check). Independent replication is not applicable for a single trial.
“5,364 were judged eligible, provided informed consent and were randomized to one of the 12 intervention or control arms”
“leaving 3,936 as the intention-to-treat cohort for this study, which met the preplanned sample size requirement.”
“Participants working on a three-shift schedule were excluded in the primary and sensitivity analyses for trial 4”
“Participants were randomly allocated in equal proportions to one of these 12 arms.”
“leaving 3,936 as the intention-to-treat cohort for this study, which met the preplanned sample size requirement.”
“Two participants later withdrew their consent, and 1,426 had baseline Patient Health Questionnaire-9 (PHQ-9) scores of ≤4, leaving 3,936 as the intention-to-treat cohort”
Sex is reported (51% male, 49% female, 0.5% other). Age is reported (mean 43.1, SD 10.7). Health status is reported via various clinical scales and comorbidities. Demographics include marital status, education, employment, and deprivation index. Species/strain and housing are not applicable for a human trial.
The paper mentions informed consent was obtained ('provided informed consent') but does not name an ethics committee or provide an approval number. The trial is registered (UMIN000047124), but registration does not substitute for an ethics statement. Regulatory compliance is not explicitly stated. Given the human subjects, this is a reporting gap.
“5,364 were judged eligible, provided informed consent and were randomized”
“5,364 were judged eligible, provided informed consent and were randomized”
“UMIN Clinical Trials Registry UMIN000047124”
The smartphone app is the investigational product; it is described in terms of its CBT skills but not with a manufacturer or version (it is a custom app). Statistical software (R) is identified, and code is shared on GitHub. No antibodies, cell lines, or other wet-lab reagents are used.
“We therefore developed a self-help smartphone CBT app that contained modules for five representative CBT skills”
“R code files used in the statistical analyses are available on GitHub at https://github.com/nomahi/RESiLIENT”
“We therefore developed a self-help smartphone CBT app that contained modules for five representative CBT skills”
The primary analysis uses MMRM, which is named. Assumptions are handled by the model (mixed model). Exact p-values are reported (e.g., P = 5.3 × 10^-13). Effect sizes with 95% CIs are reported throughout. Statistical software (R) is identified. Data presentation includes per-group n and CIs. Mathematical plausibility is not applicable for large-N continuous outcomes.
“The standardized mean difference (SMD) was −0.38 (95% confidence interval (CI): −0.48 to −0.27, P = 5.3 × 10 −13 ) for BA”
“−0.38 (−0.48 to −0.27)”
“in a mixed-model repeated measures analysis (MMRM)”
“P = 5.3 × 10 −13”
“The standardized mean difference (SMD) was −0.38 (95% confidence interval (CI): −0.48 to −0.27”
The data availability statement specifies that deidentified individual participant data will be available on UMIN-ICDR 24 months after publication, with a clear request procedure. This is a managed-access route, which is adequate for patient data. Code is shared on GitHub. Repository deposit and accession numbers are not applicable for patient-level data.
“Deidentified individual participant data and data dictionary will be made available 24 months after publication on UMIN-ICDR, an individual case data repository managed by the Japanese University Hospital Medical Information Network (UMIN) Center”
“R code files used in the statistical analyses are available on GitHub at https://github.com/nomahi/RESiLIENT”
“Deidentified individual participant data and data dictionary will be made available 24 months after publication on UMIN-ICDR”
“R code files used in the statistical analyses are available on GitHub at https://github.com/nomahi/RESiLIENT”
Methods are detailed enough for replication. The trial is registered (UMIN000047124). A CONSORT diagram is included. All prespecified outcomes appear reported. Limitations are discussed. Conclusions are proportional. Funding and competing interests are disclosed.
“(UMIN Clinical Trials Registry UMIN000047124”
“Fig. 1 Consolidated Standards of Reporting Trials (CONSORT) diagram.”
“This study is not without limitations.”
“UMIN Clinical Trials Registry UMIN000047124”
“Fig. 1 Consolidated Standards of Reporting Trials (CONSORT) diagram.”
“This study is not without limitations.”
Registered (1 ID: UMIN-CTR). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 78 references by DOI: 7 verified — 71 no DOI (shown, not verified).
- NO DOIThe prevalence and risk of developing major depression among individuals with subthreshold depression in the general populationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDetailed course of depressive symptoms and risk for developing depression in late adolescents with subthreshold depression: cohort studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDepressive spectrum states in a population-based cohort of 70-year olds followed over 9 yearsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICourse and risk factors of functional impairment in subthreshold depression and anxietyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA comparative study of nonspecific depressive symptoms and minor depression regarding functional impairment and associated characteristics in primary careNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISubsyndromal depression: prevalence, use of health services and quality of life in an Australian populationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDifferential mortality rates in major and subthreshold depression: meta-analysis of studies that measured bothNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEconomic costs of minor depression: a population-based studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe risk of developing major depression among individuals with subthreshold depression: a systematic review and meta-analysis of longitudinal cohort studiesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPsychotherapy for subclinical depression: meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPsychological interventions to prevent the onset of major depression in adults: a systematic review and individual participant data meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPsychological intervention in individuals with subthreshold depression: individual participant data meta-analysis of treatment effects and moderatorsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIManagement of depression in adults: a reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITime for united action on depression: a Lancet–World Psychiatric Association CommissionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffectiveness and acceptability of cognitive behavior therapy delivery formats in adults with depression: a network meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInternet-based cognitive behavioral therapy for depression: a systematic review and individual patient data network meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDismantling, optimising, and personalising internet cognitive behavioural therapy for depression: a systematic review and component network meta-analysis using individual participant dataNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFour 2 × 2 factorial trials of smartphone CBT to reduce subthreshold depression and to prevent new depressive episodes among adults in the community-RESiLIENT trial (Resilience Enhancement with Smartphone in LIving ENvironmenTs): a master protocolNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn overview of platform trials with a checklist for clinical readersNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIControl conditions for randomised trials of behavioural interventions in psychiatry: a decision frameworkNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDifferent control conditions can produce different effect estimates in psychotherapy trials for depressionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIComponents of smartphone cognitive-behavioural therapy for subthreshold depression among 1093 university students: a factorial trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPsychotherapies for depression: a network meta-analysis covering efficacy, acceptability and long-term outcomes of all main treatment typesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPredictors of treatment dropout in self-guided web-based interventions for depression: an ‘individual patient data’ meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAdherence to internet-based and face-to-face cognitive behavioural therapy for depression: a meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWaiting list may be a nocebo condition in psychotherapy trials: a contribution from network meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA meta-analysis of outcome studies comparing bona fide psychotherapies: empirically, ‘all must have prizes’No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIComponent studies of psychological treatments of adult depression: a systematic review and meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDismantling cognitive-behaviour therapy for panic disorder: a systematic review and component network meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIComponents and delivery formats of cognitive behavioral therapy for chronic insomnia in adults: a systematic review and component network meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRoB 2: a revised tool for assessing risk of bias in randomised trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISelf-reported versus clinician-rated symptoms of depression as outcome measures in psychotherapy research on depression: a meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe relations between observer-rating and self-report of depressive symptomatologyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA meta-analysis of antidepressant outcome under ‘blinder’ conditionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISelf-reports vs clinician ratings of efficacies of psychotherapies for depression: a meta-analysis of randomized trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInvestigation of active ingredients within internet-delivered cognitive behavioral therapy for depression: a randomized optimization trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICONSORT 2010 statement: updated guidelines for reporting parallel group randomised trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICONSORT statement for randomized trials of nonpharmacologic treatments: a 2017 update and a CONSORT extension for nonpharmacologic trial abstractsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReporting of factorial randomized trials: extension of the CONSORT 2010 statementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIValidation and utility of a self-report version of PRIME-MD: the PHQ primary care study. Primary Care Evaluation of Mental Disorders. Patient Health QuestionnaireNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAssociations of all-cause mortality with census-based neighbourhood deprivation and population density in Japan: a multilevel survival analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScreening for alcohol problems in primary care: a systematic reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScreening for alcohol abuse using the CAGE questionnaireNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIConstruction of the Big Five Scales of personality trait terms and concurrent validity with NPINo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDevelopment of a short form of the Japanese Big-Five Scale, and a test of its reliability and validityNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIValidation of the Japanese Big Five Scale Short Form in a university student sampleNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRelationship of childhood abuse and household dysfunction to many of the leading causes of death in adults. The Adverse Childhood Experiences (ACE) StudyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIProviding patient progress feedback and clinical support tools to therapists: is the therapeutic process of patients on-track to recovery enhanced in psychosomatic in-patient therapy under the conditions of routine practice?No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDevelopment and validation of the cognitive behavioural therapy skills scale among college studentsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe PHQ-9: validity of a brief depression severity measureNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA brief measure for assessing generalized anxiety disorder: the GAD-7No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIValidation of the Insomnia Severity Index as an outcome measure for insomnia researchNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIProcedural validity of the computerized version of the Composite International Diagnostic Interview (CIDI-Auto) in the anxiety disordersNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOne-year test–retest reliability of a Japanese web-based version of the WHO Composite International Diagnostic Interview (CIDI) for major depression in a working populationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Work and Social Adjustment Scale: a simple measure of impairment in functioningNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe World Health Organization Health and Work Performance Questionnaire (HPQ)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIConstruct validity and test–retest reliability of the World Mental Health Japan version of the World Health Organization Health and Work Performance Questionnaire Short Version: a preliminary studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn ultra-short measure for work engagementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDevelopment and preliminary testing of the new five-level version of EQ-5D (EQ-5D-5L)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDeveloping a Japanese version of the EQ-5D-5L value setNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFurther validation of the Warwick–Edinburgh Mental Well-being Scale (WEMWBS) in the UK veterinary profession: Rasch analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIShort Warwick–Edinburgh Mental Well-being Scale (SWEMWBS): performance in a clinical sample in relation to PHQ-9 and GAD-7No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe client satisfaction questionnaire. Psychometric properties and correlations with service utilization and psychotherapy outcomeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Insomnia Severity Index: psychometric indicators to detect insomnia cases and evaluate treatment responseNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISmartphone cognitive behavioral therapy as an adjunct to pharmacotherapy for refractory depression: randomized controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAccounting for dropout bias using mixed-effects modelsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe effect of correlation structure on treatment contrasts estimated from incomplete clinical trial data with likelihood-based repeated measures compared with last observation carried forward ANOVANo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISmall sample inference for fixed effects from restricted maximum likelihoodNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn improved approximation to the precision of fixed effects from restricted maximum likelihoodNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMethods and Applications of Statistics in Clinical Trials, Volume 1 and Volume 2: Concepts, Principles, Trials, and DesignsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDesign and Analysis of Clinical ExperimentsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://www.umin.ac.jp/icdr/index.htmlLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/nomahi/RESiLIENTResolves to GitHub (code repository).
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORtypoExtended Data Table 8“−02.9 (−0.57 to −0.00)”→ Change to −0.29 (−0.57 to −0.00)Likely typo: missing decimal point.
- MINORconsistencyAbstract“effect sizes ranging between −0.67 ... and −0.16”→ Consider using absolute values for clarity.Negative effect sizes are standard but may be confusing.
- MINORclarityTable 5 footnote“BA + CR of ns + ba + cr , BA + PS of ns + ba + ps , BA + AT of ns + ba + at and BA + BI of ns + ba + at”→ Correct the last term to 'ns + ba + bi'.Typo: 'at' should be 'bi'.
- MINORconsistencyAbstract“effect sizes ranging between −0.67 ... and −0.16”→ Consider clarifying that these are absolute values.Effect sizes are negative, but the range is given as positive numbers.
The published work is robust and generally trustworthy, but readers should weigh the missing ethics approval statement and the minor internal inconsistencies (sign error in Discussion, typo in Table 5 footnote). These do not invalidate the findings but warrant a correction or clarification from the authors.
- 1.HIGHethicsAdd an explicit ethics approval statement in the Methods, naming the IRB/ethics committee and approval number.A human trial without a stated ethics approval is a serious reporting gap that could undermine trust in the study's conduct.
- 2.HIGHreportingCorrect the sign error in the Discussion where effect sizes are reported as positive (0.67 to 0.16) instead of negative (−0.67 to −0.16).Internal contradiction between Abstract and Discussion could confuse readers and raise concerns about data integrity.
- 3.HIGHcopyeditFix the typo in Table 5 footnote: change 'ns + ba + at' to 'ns + ba + bi' for the BA+BI component.Incorrect component description could mislead readers about the trial arms.
- 4.HIGHreportingDescribe the randomization method (e.g., computer-generated random sequence) in the Methods.Randomization method is a key methodological detail that should be explicitly reported for reproducibility.
- 5.HIGHreportingState whether blinding was used or explain why it was not feasible for this app-based trial.Blinding status is important for assessing risk of bias in an interventional study.
- 6.HIGHethicsProvide details on the informed consent process (e.g., written, electronic) in the Methods.Informed consent is a fundamental ethical requirement; the current mention is insufficient.
- 7.HIGHethicsInclude a statement of regulatory compliance (e.g., Declaration of Helsinki) in the Methods.Explicit compliance with ethical standards strengthens the study's credibility.
- 8.MEDIUMreportingSpecify the version and platform of the smartphone app in the Methods to fully identify the investigational product.Full identification of the intervention is needed for replication and interpretation.
- 9.MEDIUMcopyeditFix the typo in Extended Data Table 8: change '−02.9' to '−0.29'.A missing decimal point could be misread as an impossible effect size.
- 10.MEDIUMcopyeditClarify in the Abstract that effect sizes are negative values (e.g., use absolute values or state 'ranging from −0.67 to −0.16').The current phrasing may confuse readers about the direction of effects.
- 11.LOWreportingEnsure all supplementary tables are referenced in the main text.Unreferenced supplementary material can be overlooked by readers.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.