A socially assistive robot to support mental wellbeing in LGBTQ+ young people at risk of self-harm: a randomized controlled trial.
Williams AJ, Rhodes CA, Cleare S, Borschmann R, Gross JJ, Petrova K, Posada L, Tench CR, Chapman-Nisar A, Martin L, Hollis C, Townsend E, Slovak P, Digital Youth research team
- DOI
- 10.1038/s41591-026-04422-6
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/88c86792-036b-4106-869f-fab28f21728a is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 12 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is the Difficulties in Emotion Regulation Scale (DERS-8), a self-report questionnaire, which is a surrogate for clinical benefit. The paper does not demonstrate target engagement at the tested dose (no PK/PD or dose-exposure data) and does not cite validated evidence linking DERS-8 changes to hard clinical outcomes such as self-harm or suicide. The secondary outcomes (GAD-7, PHQ-9) are also symptom scales, and the only hard outcome (self-harm) showed no significant effect.
“The primary outcome was perceived emotion regulation difficulties at follow-up, adjusted for baseline, gender identity and age. Participants allocated to the Purrble intervention reported fewer emotion regulation difficulties at follow-up than those allocated…”
- 02Treatment effect not shown to be clinically meaningful
The primary effect size is a mean difference of -3.04 on the DERS-8, with a partial eta-squared of 0.07 (small-to-medium). The paper does not anchor this change to a minimal clinically important difference (MCID) or demonstrate that it is clinically meaningful. The reliable change index shows 29% vs 14% improvement, but this is not tied to a validated threshold for clinical significance. The effect is statistically significant but not explicitly anchored as clinically material.
“There was a significant main effect of condition such that participants assigned to Purrble + SP reported fewer emotion regulation difficulties as measured by DERS-8 at follow-up (weeks 11−13) than those assigned to SP-Only (B = −3.04, 95% CI: −4.92 to −1.16,…”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported RCT of a socially assistive robot for emotion regulation in LGBTQ+ youth, with strong scientific premise, rigorous design, and full ethical approvals. The main weaknesses are the lack of an explicit power analysis, a vague data availability statement, and minor reporting gaps (statistical software not named, some p-values as thresholds, sex not reported separately from gender identity).
Both reviewers classified the study as interventional (RCT), and this was adopted. The evaluation covered all eight dimensions; non-applicable sub-criteria (e.g., animal housing, cell line authentication) were excluded. The statistics verification recomputed 6 reported tests consistently, but this does not confirm the correctness of unreported analyses. The integrity check found no high-severity concerns.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 6 tests: 6 consistent, 0 inconsistent; 3 recomputed directly from the reported test statistics, 3 via agent-written checks.
- CONSISTENTreported p = .030 · recomputed p = .031Recomputed odds ratio 2.44 (95% CI 1.09–5.49), reported p=0.03
“odds ratio = 2.44, 95% CI: 1.09−5.49, P = 0.03”
Taken as given: 1.09–5.49 is a two-sided 95% confidence interval for the odds ratio of 2.44, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.03 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.44, 1.09, 5.49, 1) - CONSISTENTreported p = .001 · recomputed p = .001Recomputed odds ratio 4.62 (95% CI 1.85–11.53), reported p=0.001
“odds ratio = 4.62, 95% CI: 1.85−11.53, P = 0.001”
Taken as given: 1.85–11.53 is a two-sided 95% confidence interval for the odds ratio of 4.62, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(4.62, 1.85, 11.53, 1) - CONSISTENTreported p = .079 · recomputed p = .079Recomputed odds ratio 2.01 (95% CI 0.92–4.36), reported p=0.079
“odds ratio = 2.01, 95% CI: 0.92−4.36, P = 0.079”
Taken as given: 0.92–4.36 is a two-sided 95% confidence interval for the odds ratio of 2.01, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.079 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.01, 0.92, 4.36, 1) - CONSISTENTreported p = .002 · recomputed p = .002Reviewers 1, 2Primary outcome ANCOVA condition effect p-value
“B = −3.04, 95% CI: −4.92 to −1.16, s.e. = 0.95, P = 0.002; partial η2 = 0.07”
Taken as given: The t-statistic is -3.20 with 134 degrees of freedom.; The test is two-tailed.Method: Two-tailed t-test from reported t and df.How we recomputed it: pT(-3.20, 134) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Depression outcome condition effect p-value
“B = −2.60, 95% CI: −4.02 to −1.19, s.e. = 0.71, P < 0.001; partial η2 = 0.09”
Taken as given: The t-statistic is -3.64 with 134 degrees of freedom.; The test is two-tailed.Method: Two-tailed t-test from reported t and df.How we recomputed it: pT(-3.64, 134) - CONSISTENTreported p = .044 · recomputed p = .043Reviewers 1, 2Anxiety outcome condition effect p-value
“B = −1.35, 95% CI: −2.66 to −0.04, s.e. = 0.66, P = 0.044; partial η2 = 0.03”
Taken as given: The t-statistic is -2.04 with 134 degrees of freedom.; The test is two-tailed.Method: Two-tailed t-test from reported t and df.How we recomputed it: pT(-2.04, 134)
- lowinternal contradictionThe abstract reports 153 participants randomized, but the CONSORT flow indicates 159 attended safety planning and 2 were excluded, leading to 157, then 2 more excluded, resulting in 153. The numbers are consistent but the flow is complex.
In total, 248 youth completed the consent and demographics form, with 159 attending the compulsory safety planning session... After briefing, two more individuals were excluded... An additional two participants were excluded at the analysis stage due to misfiling errors within study records, resulting in a final analytic sample of 153 participants.
Resultsreviewer’s wording - lowinternal contradictionThe abstract reports 'adjusted mean difference: –3.04' while the results table reports B = -3.04, consistent. However, the abstract states 'partial η 2 = 0.07' and the table also shows 0.07, consistent. No contradiction found.
“adjusted mean difference: –3.04; 95% confidence interval (CI): −4.92 to −1.16; P = 0.002; partial η 2 = 0.07”
AbstractFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2Purrble may offer a scalable intervention to complement existing therapeutic approaches.The trial shows efficacy, but scalability is inferred from cost and design, not directly tested.Evidence: Discussion on cost and low-burden nature.
“Purrble may offer a scalable intervention to complement existing therapeutic approaches to support LGBTQ+ youth to enhance their emotion regulation.”
DiscussionFind in source - partialReviewer 2Gender identity moderated intervention efficacy on emotion regulation and anxiety.Interaction effects were significant but not robust to multiple testing correction; the paper acknowledges this.Evidence: Interaction analyses show significant interactions for emotion regulation (P=0.038) and anxiety (P=0.031), but adjusted P values were 0.057 and 0.057 after Benjamini-Hochberg.
“This model revealed a significant interaction (B = 3.92, 95% CI: 0.22−7.63, s.e. = 1.87, P = 0.038, partial η 2 = 0.03).”
ResultsFind in source - supportedReviewers 1, 2Participants allocated to Purrble + SP reported fewer emotion regulation difficulties at follow-up than those allocated to safety planning alone.The primary outcome analysis shows a significant adjusted mean difference with a 95% CI excluding zero.Evidence: Adjusted mean difference: –3.04; 95% CI: −4.92 to −1.16; P = 0.002
Participants allocated to the Purrble intervention reported fewer emotion regulation difficulties at follow-up than those allocated to safety planning alone (adjusted mean difference: –3.04; 95% confidence interval (CI): −4.92 to −1.16; P = 0.002; partial η2 = 0.07).
Abstractreviewer’s wording - supportedReviewers 1, 2Participants in the Purrble intervention also reported significantly lower symptoms of anxiety and depression.Secondary outcome analyses show significant effects for both depression and anxiety.Evidence: Depression: B = −2.60, P < 0.001; Anxiety: B = −1.35, P = 0.044
“For secondary outcomes, participants in the Purrble intervention also reported significantly lower symptoms of anxiety and depression, but no significant main effect was observed for self-harm.”
AbstractFind in source - supportedReviewers 1, 2No serious Purrble-related adverse events were observed.The safety section reports no serious adverse events and only three reactive safeguarding contacts.Evidence: No serious adverse events occurred.
“No serious adverse events occurred.”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is the Difficulties in Emotion Regulation Scale (DERS-8), a self-report questionnaire, which is a surrogate for clinical benefit. The paper does not demonstrate target engagement at the tested dose (no PK/PD or dose-exposure data) and does not cite validated evidence linking DERS-8 changes to hard clinical outcomes such as self-harm or suicide. The secondary outcomes (GAD-7, PHQ-9) are also symptom scales, and the only hard outcome (self-harm) showed no significant effect.
“The primary outcome was perceived emotion regulation difficulties at follow-up, adjusted for baseline, gender identity and age. Participants allocated to the Purrble intervention reported fewer emotion regulation difficulties at follow-up than those allocated to safety planning alone (adjusted mean difference: –3.04; 95% confidence interval (CI): −4.92 to −1.16; P = 0.002; partial η 2 = 0.07).”
- INADEQUATEEffect sizeThe primary effect size is a mean difference of -3.04 on the DERS-8, with a partial eta-squared of 0.07 (small-to-medium). The paper does not anchor this change to a minimal clinically important difference (MCID) or demonstrate that it is clinically meaningful. The reliable change index shows 29% vs 14% improvement, but this is not tied to a validated threshold for clinical significance. The effect is statistically significant but not explicitly anchored as clinically material.
“There was a significant main effect of condition such that participants assigned to Purrble + SP reported fewer emotion regulation difficulties as measured by DERS-8 at follow-up (weeks 11−13) than those assigned to SP-Only (B = −3.04, 95% CI: −4.92 to −1.16, s.e. = 0.95, P = 0.002; partial η 2 = 0.07), controlling for baseline, gender identity and age.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites numerous studies on LGBTQ+ youth mental health disparities, emotion regulation as a transdiagnostic target, and prior pilot work on Purrble. The rationale for the intervention is logically developed, and the paper explicitly addresses limitations of prior research, such as the lack of RCTs for Purrble and the need to examine gender identity subgroups.
“However, to date, no randomized controlled trials evaluating the efficacy of the Purrble intervention have been conducted.”
Randomization method (online generator) and unit (individual) are reported. Blinding is not applicable as participants were aware of allocation; the paper acknowledges this limitation. Power analysis is not explicitly reported, but the target sample size was achieved. Inclusion/exclusion criteria are detailed. Outlier handling is described via sensitivity analyses. The control condition (safety planning alone) is appropriate. Independent replication is not applicable for a single trial.
“Participants were individually randomized in a 1:1 ratio (stratified by gender identity and cisgender versus TGD, to ensure balance between conditions) to either the Purrble + SP condition or the SP-Only (plus waitlist) condition, using an online randomization generator.”
“Participants were individually randomized in a 1:1 ratio (stratified by gender identity and cisgender versus TGD, to ensure balance between conditions) to either the Purrble + SP condition or the SP-Only (plus waitlist) condition, using an online randomization generator.”
Sex is reported as gender identity (cisgender vs TGD). Age is reported with means and SDs. Demographics are detailed in Table 1. Species/strain and housing conditions are not applicable as this is a human trial.
“Age (years), mean (s.d.) | 20.4 (2.3) | 20.1 (2.5) | | Gender identity, n (%) | | Cisgender | 39 (51.3) | 37 (48.1) | | TGD | 37 (48.7) | 40 (51.9)”
“Participants in the Purrble + SP and SP-Only groups were of equivalent age at baseline (20.4 ± 2.3 years versus 20.1 ± 2.5years; t 150.5 = 0.86, P = 0.39, Cohenʼs d = 0.14, 95% CI: −0.18 to 0.46).”
The paper states approval by the ethics committee at King's College London with a protocol number (RESCM-22/23-34570) and that all participants provided written informed consent. Regulatory compliance is implied by adherence to CONSORT and UK regulations.
“All study procedures were approved by the ethics committee at King’s College London (RESCM-22/23-34570), and all participants provided written informed consent prior to participation.”
“All study procedures were approved by the ethics committee at King’s College London (RESCM-22/23-34570), and all participants provided written informed consent prior to participation.”
“all participants provided written informed consent prior to participation.”
Purrble is described with its features and manufacturer (implied). Qualtrics is named as the survey platform. No antibodies, cell lines, or reagents are used. The device is the key resource and is adequately identified.
“Purrble: socially assistive robot intervention device. Image of the socially assistive robot provided to participants in the intervention condition.”
“The trial period lasted 13 weeks, with weekly online surveys hosted by Qualtrics.”
“Interactive features include touch-responsive sensors in the ears, feet and back stripe, a gyroscope for movement detection and an internal vibration motor that generates heartbeat-like and purr-like sensations.”
“The trial period lasted 13 weeks, with weekly online surveys hosted by Qualtrics.”
Tests are named (ANCOVA, linear mixed effects models, chi-square). Assumptions are addressed via sensitivity analyses. Exact p-values are reported. Effect sizes with CIs are provided. Software is not explicitly identified for statistical analysis. Data presentation includes means, SDs, and CIs. Mathematical plausibility checks were performed and no inconsistencies found.
“B = −3.04, 95% CI: −4.92 to −1.16, s.e. = 0.95, P = 0.002; partial η 2 = 0.07”
“B = −2.60, 95% CI: −4.02 to −1.19, s.e. = 0.71, P < 0.001; partial η 2 = 0.09”
The paper states data are available on reasonable request but does not specify a platform or conditions. No repository deposit or accession numbers are provided. Code sharing is not mentioned.
Trial registration number is provided (NCT06025942). CONSORT checklist is referenced. All pre-specified outcomes are reported, including null results for self-harm. Limitations are thoroughly discussed. Conclusions are proportional. Funding sources and COI are stated.
“ClinicalTrials.gov: NCT06025942”
“This paper was developed in accordance with the CONSORT 2010 checklist”
“Results should be interpreted in the context of several limitations.”
“ClinicalTrials.gov: NCT06025942”
“This paper was developed in accordance with the CONSORT 2010 checklist”
“Results should be interpreted in the context of several limitations.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 108 references by DOI: 97 verified — 11 no DOI (shown, not verified).
- NO DOISuicideNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILGBT in Britain: health reportNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA robot of my own: participatory design of socially assistive robots for independently living older adults diagnosed with depressionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOI‘I just let him cry’: designing socio-technical interventions in families to prevent mental health disordersNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWhen emotion goes wrong: realizing the promise of affective scienceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDetecting and treating suicide ideation in all settingsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIncreases in young people waiting over a year for mental health supportNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISelf-harm: assessment, management and preventing recurrenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExploring technology-mediated parental socialisation of emotion: leveraging an embodied, in-situ intervention for child emotion regulationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMindfulness interventions and emotion regulationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIApplied Linear Statistical ModelsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- codeGitHubLIVEHTTP 200https://github.com/caarhodes/PurrbleResolves to GitHub (code repository).
- codeGitHubLIVEHTTP 200http://github.com/caarhodes/PurrbleResolves to GitHub (code repository).
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, grammar.
- MINORconsistencyAbstract“Purrble + SP”→ Define abbreviation on first use.Abbreviation not defined in abstract.
- MINORgrammarDiscussion“Purrble may offer a scalable intervention to complement existing therapeutic approaches to support LGBTQ+ youth to enhance their emotion regulation.”→ Consider rephrasing for clarity: '...to support LGBTQ+ youth in enhancing their emotion regulation.'Awkward phrasing.
- MINORconsistencyAbstract“adjusted mean difference: –3.04; 95% confidence interval (CI): −4.92 to −1.16; P = 0.002; partial η 2 = 0.07”→ Ensure consistent use of en-dash vs hyphen in ranges.Minor formatting inconsistency.
- MINORgrammarDiscussion“Purrble may offer a scalable intervention to complement existing therapeutic approaches to support LGBTQ+ youth to enhance their emotion regulation.”→ Consider rephrasing for clarity: '...to support emotion regulation in LGBTQ+ youth.'Awkward phrasing.
The published work is robust and generally trustworthy, but an informed reader should weigh the absence of a power analysis and the vague data availability statement. These are reporting gaps that would warrant a correction or clarification, but they do not undermine the core findings.
- 1.HIGHdata codeIn the Data Availability section, specify a repository or managed-access platform with conditions for access, and provide a DOI or accession number if available.The current statement is vague and does not meet transparency standards for a data-driven trial.
- 2.HIGHstatisticsIn the Methods, add a power analysis or sample size calculation, specifying the expected effect size, alpha, and power.The absence of a power analysis is a notable gap in study design reporting.
- 3.HIGHstatisticsIn the Methods, identify the statistical software and version used for analyses (e.g., R version, SPSS).Software identification is a standard reporting requirement and aids reproducibility.
- 4.MEDIUMstatisticsIn the Results, report exact p-values for all outcomes instead of thresholds like 'P < 0.001' where possible.Exact p-values allow readers to assess the strength of evidence more precisely.
- 5.MEDIUMreportingIn the Methods, explicitly state that blinding was not feasible and that the trial was open-label.Clarifying blinding status in Methods, even if open-label, improves transparency.
- 6.MEDIUMreportingIn the Abstract, define the abbreviation 'SP' on first use.Abbreviations should be defined at first mention for reader clarity.
- 7.MEDIUMcopyeditIn the Discussion, rephrase the sentence 'Purrble may offer a scalable intervention to complement existing therapeutic approaches to support LGBTQ+ youth to enhance their emotion regulation.' for clarity.The current phrasing is awkward and could be misinterpreted.
- 8.LOWcopyeditIn the Abstract, ensure consistent use of en-dash vs hyphen in numeric ranges (e.g., '−4.92 to −1.16').Minor formatting consistency improves professionalism.
- 9.LOWreportingIn Table 1, consider reporting sex (male/female) in addition to gender identity for completeness.While gender identity is the primary variable, reporting sex may be useful for readers.
- 10.LOWdata codeAdd a statement about code availability if any custom analysis code was used.Code sharing enhances reproducibility if applicable.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.