Self-help mobile messaging intervention for depression among older adults in resource-limited settings: a randomized controlled trial.
Scazufca M, Nakamura CA, Seward N, Didone TVN, Moretti FA, Oliveira da Costa M, Queiroz de Souza CH, Macias de Oliveira G, Souza Dos Santos M, Pereira LA, Mendes de Sá Martins M, van de Ven P, Hollingworth W, Peters TJ, Araya R
- DOI
- 10.1038/s41591-024-02864-4
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/35073240-3b78-4104-b9ee-e50f08702500 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsOverstated claim−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 30 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is improvement in depressive symptomatology defined as PHQ-9 score < 10, which is a symptom scale, not a hard clinical outcome. The paper does not provide evidence of target engagement (e.g., PK/PD) or a validated link between PHQ-9 improvement and long-term clinical outcomes such as function or quality of life. The PHQ-9 is a screening tool, and the threshold of 10 is a pragmatic choice, but the paper does not cite validation that this surrogate is a reliable predictor of clinical benefit.
“The primary outcome was improvement from depressive symptomatology (PHQ-9 < 10) at 3 months.”
- 02Treatment effect not shown to be clinically meaningful
The primary effect is an absolute difference of 10.2 percentage points (42.4% vs 32.2%) in the proportion achieving PHQ-9 < 10. The paper acknowledges this is lower than the prespecified target difference of 15 percentage points. The effect is statistically significant but not anchored to a minimal clinically important difference (MCID) for PHQ-9 in this population. The authors themselves describe the effect as 'small and short-term'.
“The estimated intervention effect (10.2 percentage points in absolute terms) was lower than that specified as the original target difference (15 percentage points).”
- 03Conclusion reaches beyond the evidence
The intervention effect is clinically meaningful.
“The estimated intervention effect (10.2 percentage points in absolute terms) was lower than that specified as the original target difference (15 percentage points).”
DiscussionFind in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported pragmatic RCT of a mobile messaging intervention for depression in older adults. The design, ethics, and statistical reporting are strong, with the main weakness being vague data and code availability statements. Minor copyedit issues and a slightly overstated claim about clinical meaningfulness are noted.
Both reviewers classified the study as interventional (RCT), so no divergence. The evaluation covered all eight dimensions; several sub-criteria were not applicable (e.g., animal-related, cell lines). The statistics verification checked only a subset of tests (2 reported with sufficient detail); the rest are unverified. The claim audit flagged one overstated claim.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks.
- CONSISTENTreported p = .019 · recomputed p = .020Reviewers 1, 2Primary outcome adjusted OR p-value from CI
“The adjusted odds ratio (OR) for this improvement, after imputing missing values, was 1.57 (95% CI = 1.07–2.29; P = 0.019).”
Taken as given: The OR is 1.57 with 95% CI 1.07-2.29.; The CI is two-sided at 95%.; The p-value is from a Wald test on the log-odds scale.Method: Recomputed two-sided p-value from the reported OR and 95% CI using the normal approximation for the log-odds ratio.How we recomputed it: pCI(1.57, 1.07, 2.29, 1) - CONSISTENTreported p = .016 · recomputed p = .017Reviewers 1, 2Secondary outcome reduction at 3 months adjusted OR p-value from CI
“Secondary outcome: reduction in depressive symptomatology at 3 months e | 95/257 (37.0) | 73/270 (27.0) | 9.9 (2.0–17.9) | 1.58 (1.08–2.29) | 0.016”
Taken as given: The OR is 1.58 with 95% CI 1.08-2.29.; The CI is two-sided at 95%.; The p-value is from a Wald test on the log-odds scale.Method: Recomputed two-sided p-value from the reported OR and 95% CI using the normal approximation for the log-odds ratio.How we recomputed it: pCI(1.58, 1.08, 2.29, 1)
- lowinternal contradictionThe text says '41 (13.8%) intervention participants were lost to follow-up, 35 (11.5%) of which were controls.' This implies 41+35=76 lost, but the total lost is 603-527=76. Consistent.
“41 (13.8%) intervention participants were lost to follow-up, 35 (11.5%) of which were controls.”
ResultsFind in source - lowinternal contradictionThe abstract states '451 (74.8%) women' while Table 1 shows 225/298 (75.5%) in intervention and 226/305 (74.1%) in control, which sums to 451/603 (74.8%). This is consistent.
“451 (74.8%) women”
Table 1Find in source
Overstated conclusions
4 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
4 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated), 1 only partially supported (evidence backs part of the claim; gaps or caveats remain).
- overstatedReviewers 1, 2The intervention effect is clinically meaningful.The observed absolute difference (10.2 percentage points) was lower than the pre-specified target difference (15 percentage points), and the authors acknowledge this, so claiming clinical meaningfulness is overstated.Evidence: Discussion: 'The estimated intervention effect (10.2 percentage points in absolute terms) was lower than that specified as the original target difference (15 percentage points).'
“The estimated intervention effect (10.2 percentage points in absolute terms) was lower than that specified as the original target difference (15 percentage points).”
DiscussionFind in source - partialReviewers 1, 2The intervention is feasible and acceptable for older adults in low-resource settings.High message opening rates (75.8% opened at least 36 messages) suggest feasibility, but acceptability is not directly measured; the claim is partially supported.Evidence: Exploratory analyses: '226 (75.8%) intervention participants ‘opened’ at least 36 of the 48 messages'.
“Equally important, this study demonstrated the feasibility and acceptability of a simple and affordable intervention that can provide help to a large proportion of older adults with depressive symptoms in Brazil.”
DiscussionFind in source - supportedReviewers 1, 2The Viva Vida intervention improved depressive symptomatology at 3 months compared to control.The primary outcome analysis shows a statistically significant adjusted OR of 1.57 (95% CI 1.07-2.29, P=0.019), supporting the claim.Evidence: Primary outcome result: adjusted OR = 1.57, 95% CI = 1.07–2.29, P = 0.019.
“In the intervention arm, 109 of 257 (42.4%) participants had an improved depressive symptomatology, compared with 87 of 270 (32.2%) participants in the control arm (adjusted odds ratio = 1.57; 95% confidence interval = 1.07–2.29; P = 0.019).”
AbstractFind in source - supportedReviewers 1, 2No severe adverse events related to trial participation were observed.The safety section reports no severe adverse events related to participation, and all events were considered unrelated.Evidence: Safety section: 'No severe adverse events related to trial participation were observed.'
“No severe adverse events related to trial participation were observed.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is improvement in depressive symptomatology defined as PHQ-9 score < 10, which is a symptom scale, not a hard clinical outcome. The paper does not provide evidence of target engagement (e.g., PK/PD) or a validated link between PHQ-9 improvement and long-term clinical outcomes such as function or quality of life. The PHQ-9 is a screening tool, and the threshold of 10 is a pragmatic choice, but the paper does not cite validation that this surrogate is a reliable predictor of clinical benefit.
“The primary outcome was improvement from depressive symptomatology (PHQ-9 < 10) at 3 months.”
- INADEQUATEEffect sizeThe primary effect is an absolute difference of 10.2 percentage points (42.4% vs 32.2%) in the proportion achieving PHQ-9 < 10. The paper acknowledges this is lower than the prespecified target difference of 15 percentage points. The effect is statistically significant but not anchored to a minimal clinically important difference (MCID) for PHQ-9 in this population. The authors themselves describe the effect as 'small and short-term'.
“The estimated intervention effect (10.2 percentage points in absolute terms) was lower than that specified as the original target difference (15 percentage points).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites global and local burden of depression in older adults, limitations of existing task-sharing programs, and the lack of evidence for self-help digital interventions in LMICs. It explicitly states the need for scalable solutions and builds the rationale for the Viva Vida intervention based on prior work (PROACTIVE). Limitations of prior research are acknowledged (e.g., need for health professional support, lack of evidence in older adults in LMICs).
“but there were no local studies that had investigated a clinically meaningful threshold.”
The study is a pragmatic, single-blind, individually randomized controlled trial with 1:1 allocation. Randomization used randomly permuted blocks with random block sizes, stratified by age, sex, and baseline PHQ-9 severity. Allocation was concealed via REDCap. Research assistants were blinded to allocation; participants could not be masked due to intervention nature. Sample size calculation was provided (440-500 for 80-85% power to detect 15 percentage point difference). Inclusion/exclusion criteria were pre-specified. Missing data handling was pre-specified in the SAP, with imputation and complete case analyses. The trial is registered (ReBEC).
“The allocation sequence was generated using randomly permuted blocks with random block sizes of six, eight or ten by research team members not involved in data collection (C.A.N. and T.J.P.).”
“Research assistants involved in recruitment and follow-up data collection were blinded to trial allocation.”
“With an assumed attrition of 25%, 440–500 randomized individuals would yield 80–85% power to detect a 15 percentage point difference in depression improvement (PHQ-9 < 10) rates between the control and intervention arms at 3 months (25% versus 40%) using a two-sided 5% alpha.”
“The allocation sequence was generated using randomly permuted blocks with random block sizes of six, eight or ten by research team members not involved in data collection (C.A.N. and T.J.P.).”
“Research assistants involved in recruitment and follow-up data collection were blinded to trial allocation.”
“With an assumed attrition of 25%, 440–500 randomized individuals would yield 80–85% power to detect a 15 percentage point difference in depression improvement (PHQ-9 < 10) rates between the control and intervention arms at 3 months (25% versus 40%) using a two-sided 5% alpha.”
The paper reports sex (74.8% women), age groups, education, income, hypertension, diabetes, and pharmacological treatment for depression in Table 1. Age and sex are used as stratification variables. Health status is captured via self-reported comorbidities. Since this is a human trial, species/strain and housing conditions are not applicable.
“Female | 225/298 (75.5) | 226/305 (74.1)”
“Hypertension (self-reported) | 207/298 (69.5) | 222/305 (72.8)”
“Female | 225/298 (75.5) | 226/305 (74.1)”
“Hypertension (self-reported) | 207/298 (69.5) | 222/305 (72.8)”
The study was approved by a named ethics committee (Comissão para Análise de Projetos de Pesquisa, ref: 4.097.596) and authorized by the Guarulhos Health Secretary. Informed verbal consent was obtained and audio-recorded. The trial was registered. Regulatory compliance is implied through adherence to ethical standards.
“This study was approved by the ethics committee of the Hospital das Clínicas da Faculdade de Medicina da Universidade de São Paulo (Comissão para Análise de Projetos de Pesquisa, ref: 4.097.596, first approved 10 March 2021)”
“Informed verbal consent was obtained before the screening assessment and when participants were invited to the trial, both conducted by phone. Consent was audio-recorded after authorization from the older adult.”
“This study was approved by the ethics committee of the Hospital das Clínicas da Faculdade de Medicina da Universidade de São Paulo (Comissão para Análise de Projetos de Pesquisa, ref: 4.097.596, first approved 10 March 2021)”
“Informed verbal consent was obtained before the screening assessment and when participants were invited to the trial, both conducted by phone. Consent was audio-recorded after authorization from the older adult.”
The intervention is described in detail (48 audio/visual messages, content, delivery schedule). The control is a single audio message. Software used includes REDCap and Stata v.17. No antibodies, cell lines, or other bench reagents are applicable. The intervention is the key resource and is adequately described.
“A total of 48 audio or visual messages were automatically sent 4 days a week for 6 weeks (one in the morning, one in the afternoon).”
“Statistical tests were two-sided and all analyses were conducted using Stata v.17 (StataCorp LLC).”
“A total of 48 audio or visual messages were automatically sent 4 days a week for 6 weeks (one in the morning, one in the afternoon).”
“all analyses were conducted using Stata v.17 (StataCorp LLC).”
The paper names statistical tests (logistic regression, linear regression, CACE analysis), reports exact p-values (e.g., P = 0.019), and provides effect sizes with 95% CIs. Assumptions are addressed (Box-Tidwell test, residual plots). Software is identified. Data presentation includes per-group n and percentages. Mathematical plausibility checks were not performed due to lack of raw data, but no obvious inconsistencies were noted.
“The adjusted odds ratio (OR) for this improvement, after imputing missing values, was 1.57 (95% CI = 1.07–2.29; P = 0.019).”
“10.2 (2.0–18.4) | 1.57 (1.07–2.29) | 0.019”
“The adjusted odds ratio (OR) for this improvement, after imputing missing values, was 1.57 (95% CI = 1.07–2.29; P = 0.019).”
“The Box–Tidwell test was run after the logistic regression models to test whether the logit transform was a linear function of the predictors for the different models. Normality assumptions for linear regression models were evaluated through residual plots.”
The data availability statement says de-identified individual participant data will be made available 24 months after publication, with proposals directed to the corresponding author. This is a managed-access statement but lacks details on the review process or timeframe. Code availability similarly states code will be made available 24 months after publication, but no repository or identifier is given. No accession numbers are provided.
“De-identified individual participant data and the data dictionary will be made available 24 months after publication. Proposals with specific aims and an analysis plan should be directed to the corresponding author (M.S.).”
“The code for the data analysis will also be made available 24 months after publication.”
“De-identified individual participant data and the data dictionary will be made available 24 months after publication. Proposals with specific aims and an analysis plan should be directed to the corresponding author (M.S.).”
“The code for the data analysis will also be made available 24 months after publication.”
The trial is registered (ReBEC RBR-4c94dtn). A CONSORT diagram is provided. All pre-specified outcomes are reported, including null results. Limitations are discussed in detail. Conclusions are proportional to the evidence. Funding and competing interests are declared.
“Brazilian Registry of Clinical Trials registration: ReBEC ( RBR-4c94dtn (https://ensaiosclinicos.gov.br/rg/RBR-4c94dtn) ).”
“In addition to the issue of the lower magnitude of the observed difference in absolute terms (and some of the values in the relevant CI) compared with the target difference as discussed above, this study has some important limitations.”
“Brazilian Registry of Clinical Trials registration: ReBEC ( RBR-4c94dtn (https://ensaiosclinicos.gov.br/rg/RBR-4c94dtn) ).”
“this study has some important limitations.”
Registration stated in text, but no registry ID was detected. Reporting guideline cited: CONSORT.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 42 references by DOI: 36 verified — 6 no DOI (shown, not verified).
- NO DOIPopulation Ages 65 and Above, Total—Low & Middle Income, High IncomeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICenso 2022: Número de Pessoas com 65 Anos ou Mais de Idade Cresceu 57,4% em 12 AnosNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISelf-Care Interventions for HealthNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPesquisa Nacional de SaúdeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStatistical analysis plan for the PRODIGITAL-D individually randomised controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMultiple Imputation for Nonresponse in SurveysNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly clarity, consistency, typo.
- MINORconsistencyAbstract“n = 298) or a single message ( n = 305)”→ Ensure consistent spacing around 'n' in all instances.Minor formatting inconsistency.
- MINORtypoExtended Data Table 6 footnote“imputed seperately”→ Change to 'separately'.Typographical error.
- MINORclarityDiscussion“Albeit at about half the level that we accounted for in our sample size calculation for the primary outcome at 3 months (about 12.5% overall compared with the 25% allowed for), a further limitation is that there is a risk of bias through attrition.”→ Rephrase for clarity, e.g., 'A further limitation is the risk of bias through attrition, albeit at about half the level accounted for in our sample size calculation (about 12.5% overall compared with the 25% allowed for).'Awkward sentence structure.
- MINORclarityDiscussion, paragraph 8“Albeit at about half the level that we accounted for in our sample size calculation for the primary outcome at 3 months (about 12.5% overall compared with the 25% allowed for), a further limitation is that there is a risk of bias through attrition.”→ Rephrase for clarity: 'A further limitation is the risk of bias through attrition, albeit at about half the level accounted for in our sample size calculation (about 12.5% overall compared with the 25% allowed for).'Awkward sentence structure.
The published work is methodologically robust and generally trustworthy, but readers should weigh the vague data/code availability and the slightly overstated claim of clinical meaningfulness. An erratum or clarification could address the availability statements and temper the claim.
- 1.HIGHdata codeIn the Data availability statement, specify a concrete access mechanism such as a named repository (e.g., Dryad, Zenodo) or a managed-access platform (e.g., Vivli) with conditions and a clear timeline, instead of only directing proposals to the corresponding author.The current statement is vague and does not meet reproducibility standards; a concrete route is needed for readers to access the data.
- 2.HIGHdata codeIn the Code availability statement, provide a public repository (e.g., GitHub, Zenodo) with a DOI or permanent identifier for the analysis code.Stating code will be available '24 months after publication' without a repository is insufficient for reproducibility.
- 3.HIGHrigorIn the Discussion, temper the claim that the intervention effect is clinically meaningful, given the observed absolute difference (10.2 percentage points) was lower than the pre-specified target (15 percentage points).The claim audit flagged this as overstated; the evidence does not fully support clinical meaningfulness as originally claimed.
- 4.MEDIUMcopyeditFix the typo 'seperately' to 'separately' in the Extended Data Table 6 footnote.Typographical errors undermine professionalism and clarity.
- 5.MEDIUMcopyeditRephrase the awkward sentence in the Discussion about attrition risk for clarity, e.g., 'A further limitation is the risk of bias through attrition, albeit at about half the level accounted for in our sample size calculation (about 12.5% overall compared with the 25% allowed for).'The current sentence structure is confusing and could be misinterpreted.
- 6.LOWcopyeditEnsure consistent spacing around 'n' in the Abstract (e.g., 'n = 298' vs 'n = 305').Minor formatting inconsistency that should be corrected for polish.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.