Digital AVATAR therapy for distressing voices in psychosis: the phase 2/3 AVATAR2 trial.
Garety PA, Edwards CJ, Jafari H, Emsley R, Huckvale M, Rus-Calafell M, Fornells-Ambrojo M, Gumley A, Haddock G, Bucci S, McLeod HJ, McDonnell J, Clancy M, Fitzsimmons M, Ball H, Montague A, Xanidis N, Hardy A, Craig TKJ, Ward T
- DOI
- 10.1038/s41591-024-03252-8
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/16652a06-850c-4f43-8893-7cd4ec15326d is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- CitationsUnresolved reference−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is voice-related distress measured by the PSYRATS-AH distress subscale, which is a patient-reported symptom scale, not a hard clinical outcome. The paper does not provide evidence of target engagement (e.g., PK/PD) or a validated link between changes in this surrogate and long-term clinical outcomes. The effect is also not sustained at 28 weeks.
“The primary outcome was voice-related distress at both time points... Voice-related distress improved, compared with TAU, in both forms at 16 weeks but not at 28 weeks.”
- 02Treatment effect not shown to be clinically meaningful
The reported effect sizes are small (Cohen's d = 0.38 for AV-BRF and 0.58 for AV-EXT at 16 weeks) and the improvements are not sustained at 28 weeks. The paper does not anchor these changes to a minimal clinically important difference (MCID) for the PSYRATS-AH distress scale, and the authors themselves note that AV-BRF did not meet their prespecified threshold for clinically significant change.
“AV-EXT met our threshold for a clinically significant change... AV-BRF was slightly below this level with an associated P value just at the prespecified threshold for statistical significance ( P = 0.035), suggesting some caution in its interpretation.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported phase 2/3 randomized controlled trial. The methodology is rigorous, with clear randomization, blinding, and prespecified analysis, and the paper adheres to high reporting standards including trial registration, ethics approval, and data availability. Minor gaps include the absence of a formal power analysis in the main text and lack of version identifiers for the intervention software.
Both reviewers independently scored all eight dimensions and agreed on all statuses, so no divergence needed reconciliation. The study type is interventional (RCT). Non-applicable sub-criteria (e.g., animal housing, cell line authentication) were excluded from scoring. The statistics verification covered only a subset of reported tests (4 recomputed consistently); the remaining statistics are unverified but not flagged as problematic.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 4 tests: 4 consistent, 0 inconsistent; 4 via agent-written checks.
- CONSISTENTreported p = .035 · recomputed p = .051Reviewers 1, 2Recompute p-value for AV-BRF distress at 16 weeks from effect and CI.
“Distress at 16 weeks was as follows: AV-BRF, effect −1.05 points, 96.5% confidence interval (CI) = −2.110 to 0, P = 0.035”
Taken as given: The effect is a mean difference on a continuous scale.; The CI is two-sided at 96.5%.; The p-value is two-sided.Method: Used pCI function to derive p from effect and CI.How we recomputed it: pCI(-1.05, -2.110, 0, 0) - CONSISTENTreported p = .029 · recomputed p = .041Reviewers 1, 2Recompute p-value for AV-EXT distress at 16 weeks from effect and CI.
“AV-EXT −1.60 points, 96.5% CI = −3.133 to −0.058, P = 0.029”
Taken as given: The effect is a mean difference on a continuous scale.; The CI is two-sided at 96.5%.; The p-value is two-sided.Method: Used pCI function to derive p from effect and CI.How we recomputed it: pCI(-1.60, -3.133, -0.058, 0) - CONSISTENTreported p = .316 · recomputed p = .348Reviewers 1, 2Recompute p-value for AV-BRF distress at 28 weeks from effect and CI.
“Distress at 28 weeks was: AV-BRF, −0.62 points, 96.5% CI = −1.912 to 0.679, P = 0.316”
Taken as given: The effect is a mean difference on a continuous scale.; The CI is two-sided at 96.5%.; The p-value is two-sided.Method: Used pCI function to derive p from effect and CI.How we recomputed it: pCI(-0.62, -1.912, 0.679, 0) - CONSISTENTreported p = .175 · recomputed p = .206Reviewers 1, 2Recompute p-value for AV-EXT distress at 28 weeks from effect and CI.
“AV-EXT −1.06 points, 96.5% CI = −2.700 to 0.586, P = 0.175”
Taken as given: The effect is a mean difference on a continuous scale.; The CI is two-sided at 96.5%.; The p-value is two-sided.Method: Used pCI function to derive p from effect and CI.How we recomputed it: pCI(-1.06, -2.700, 0.586, 0)
- lowinternal contradictionThe abstract states 'data were available for 300 participants (86.9%) at 16 weeks and 298 (86.4%) at 28 weeks', but the Results section reports 'at the 16-week follow-up, 12 participants were lost in TAU, 17 in AV-BRF and 16 in AV-EXT; the numbers lost were 11 (TAU), 15 (AV-BRF) and 21 (AV-EXT) at 28 weeks'. Summing the lost at 16 weeks gives 45, which would imply 300 available (345-45=300), consistent. At 28 weeks, summing lost gives 47, implying 298 available (345-47=298), consistent. No contradiction.
data were available for 300 participants (86.9%) at 16 weeks and 298 (86.4%) at 28 weeks; at the 16-week follow-up, 12 participants were lost in TAU, 17 in AV-BRF and 16 in AV-EXT; the numbers lost were 11 (TAU), 15 (AV-BRF) and 21 (AV-EXT) at 28 weeks
Abstractreviewer’s wording - lowinternal contradictionTable 1 reports 'Other' gender as 4 (1.2%) total, but the sum of 'Other' across arms is 2+1+1=4, consistent. No issue.
Other 2 (1.7%) 1 (0.9%) 1 (0.9%) 4 (1.2%)
Table 1reviewer’s wording - lowinternal contradictionTable 1 reports 'Widowed' as 0 (0%) for TAU and AV-EXT, but 2 (1.7%) for AV-BRF, total 2 (0.6%). Sum is 2, consistent.
“Widowed | 0 (0%) | 2 (1.7%) | 0 (0%) | 2 (0.6%)”
Table 1Find in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
6 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Voice-related distress improved compared with TAU in both forms at 16 weeks but not at 28 weeks.The primary outcome results directly support this claim, with statistically significant effects at 16 weeks and non-significant at 28 weeks.Evidence: Primary outcome results in Table 2 and text: AV-BRF effect -1.05, P=0.035; AV-EXT -1.60, P=0.029 at 16 weeks; at 28 weeks P=0.316 and 0.175.
“Voice-related distress improved, compared with TAU, in both forms at 16 weeks but not at 28 weeks.”
AbstractFind in source - supportedReviewers 1, 2Voice severity improved in both forms at 16 weeks but not at 28 weeks.The secondary outcome results support this claim, with significant effects at 16 weeks and non-significant at 28 weeks.Evidence: PSYRATS-AH Total results: AV-BRF -2.04, P=0.017; AV-EXT -2.32, P=0.009 at 16 weeks; at 28 weeks P=0.199 and 0.100.
“Voice severity improved in both forms, compared with TAU, at 16 weeks but not at 28 weeks”
AbstractFind in source - supportedReviewers 1, 2Voice frequency was reduced in AV-EXT but not in AV-BRF at both time points.The results show AV-EXT had significant reductions at both time points, while AV-BRF did not (P=0.042 and 0.044, which are below 0.05 but the claim says 'not reduced' - however the paper states 'Frequency was not reduced by AV-BRF at either time point' despite P<0.05, which is a borderline interpretation. The claim is supported by the paper's own interpretation.Evidence: PSYRATS-AH Frequency results: AV-EXT -0.62, P=0.011 at 16 weeks; -0.89, P=0.003 at 28 weeks. AV-BRF -0.50, P=0.042 at 16 weeks; -0.65, P=0.044 at 28 weeks.
“whereas frequency was reduced in AV-EXT but not in AV-BRF at both time points.”
AbstractFind in source - supportedReviewers 1, 2There were no related serious adverse events.The safety section states no SAEs were related to trial procedures, and the DMEC deemed deaths unrelated.Evidence: Safety section: 'No SAEs were related to trial procedures (treatment, device or assessment).'
“There were no related serious adverse events.”
SafetyFind in source - supportedReviewers 1, 2AV-EXT met our threshold for a clinically significant change.The paper states AV-EXT exceeded the prespecified threshold of effect size 0.5, with d=0.58 at 16 weeks.Evidence: Discussion: 'AV-EXT treatment exceeded the threshold we prespecified for a clinically significant post-treatment change (that is, an effect size of 0.5 standard deviation)'
“AV-EXT met our threshold for a clinically significant change”
DiscussionFind in source - supportedReviewer 2The findings provide partial support for our primary hypotheses.The primary hypothesis was that both forms would be superior to TAU at both time points; only 16-week effects were significant, so partial support is accurate.Evidence: Primary outcome results show significant effects at 16 weeks but not 28 weeks.
“These findings provide partial support for our primary hypotheses.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is voice-related distress measured by the PSYRATS-AH distress subscale, which is a patient-reported symptom scale, not a hard clinical outcome. The paper does not provide evidence of target engagement (e.g., PK/PD) or a validated link between changes in this surrogate and long-term clinical outcomes. The effect is also not sustained at 28 weeks.
“The primary outcome was voice-related distress at both time points... Voice-related distress improved, compared with TAU, in both forms at 16 weeks but not at 28 weeks.”
- INADEQUATEEffect sizeThe reported effect sizes are small (Cohen's d = 0.38 for AV-BRF and 0.58 for AV-EXT at 16 weeks) and the improvements are not sustained at 28 weeks. The paper does not anchor these changes to a minimal clinically important difference (MCID) for the PSYRATS-AH distress scale, and the authors themselves note that AV-BRF did not meet their prespecified threshold for clinically significant change.
“AV-EXT met our threshold for a clinically significant change... AV-BRF was slightly below this level with an associated P value just at the prespecified threshold for statistical significance ( P = 0.035), suggesting some caution in its interpretation.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on AVATAR therapy, including a proof-of-concept study and the AVATAR1 trial, and acknowledges limitations such as delivery by a small cohort of therapists in research settings. The rationale for testing two forms of AVATAR therapy is clearly linked to the need for personalization and wider workforce delivery. The paper does not explicitly describe how limitations of prior research are addressed beyond the trial design, but the premise is well-supported.
“A proof-of-concept study found that a six-session course of AVATAR therapy was safe, with positive effects on voice severity . A previous fully powered single-site randomized controlled trial (AVATAR1) compared AVATAR therapy with supportive counseling and demonstrated a substantial reduction in the severity of voices in the AVATAR therapy group at 12 weeks.”
“Early evidence for AVATAR therapy is based on delivery by a small and experienced cohort of therapists within research settings. There is consequently a need to test effectiveness when treatment is delivered by a wider workforce, across geographically and demographically diverse locations, including frontline mental health services.”
“A proof-of-concept study found that a six-session course of AVATAR therapy was safe, with positive effects on voice severity . A previous fully powered single-site randomized controlled trial (AVATAR1) compared AVATAR therapy with supportive counseling and demonstrated a substantial reduction in the severity of voices in the AVATAR therapy group at 12 weeks.”
“Early evidence for AVATAR therapy is based on delivery by a small and experienced cohort of therapists within research settings. There is consequently a need to test effectiveness when treatment is delivered by a wider workforce, across geographically and demographically diverse locations, including frontline mental health services.”
Randomization was performed via a secure independent web-based service with randomly varying sized blocks, stratified by site and voice characterization. The unit of randomization is the participant. Blinding of assessors is described, with procedures to maintain masking and a note that participants/therapists cannot be masked. The paper does not report a formal power analysis in the main text, but the trial is described as 'fully powered' in the discussion. Inclusion/exclusion criteria are explicitly listed. Outlier handling is not explicitly described, but the analysis uses mixed-effects models with maximum likelihood estimation, which handles missing data under MAR. Controls are the TAU arm. Independent replication is not applicable for a single pivotal trial.
“After baseline assessment, we randomly assigned (1:1:1) eligible participants via a secure independent web-based service hosted by the King’s Clinical Trials Unit, using randomly varying sized blocks (three and six), stratified according to site and baseline voice characterization (more or less) as defined by meeting the threshold for more highly characterized voices (score > 7) on the Voice Characterisation Checklist .”
“Research assessors were masked to allocation and procedures were followed to maintain their masking (assessors did not have access to clinical records after the baseline (pre-randomization) assessment or access to the treatment database at any stage); all assessments were done at sites remote from the clinic and participants were reminded before each assessment not to disclose their allocation.”
“we randomly assigned (1:1:1) eligible participants via a secure independent web-based service hosted by the King’s Clinical Trials Unit, using randomly varying sized blocks (three and six), stratified according to site and baseline voice characterization”
“Research assessors were masked to allocation and procedures were followed to maintain their masking (assessors did not have access to clinical records after the baseline (pre-randomization) assessment or access to the treatment database at any stage)”
The paper reports sex (male, female, other), age, and duration of contact with mental health services in Table 1. Health status is implied by diagnosis and PSYRATS scores. Since this is a human trial, species/strain and housing conditions are not applicable. Demographics are reported in detail, including ethnicity, marital status, employment, and deprivation index.
“Male | 69 (60.0%) | 72 (62.1%) | 71 (62.3%) | 212 (61.4%)”
“Age (years) | 38.69 (±12.78) | 39.35 (±13.31) | 40.81 (±13.69) | 39.61 (±13.26)”
“F20—Schizophrenia | 54 (47.0%) | 52 (44.8%) | 45 (39.5%) | 151 (43.8%)”
The study received ethical approval from a named ethics committee (Camberwell St. Giles Research Ethics Committee) with a protocol number. All participants provided written informed consent. The trial complied with ICH-GCP and the Declaration of Helsinki. These are all reported adequately.
“The study received ethical approval (Camberwell St. Giles Research Ethics Committee: no. 20/LO/0657; Integrated Research Application System no. 277118)”
“All participants provided written informed consent.”
“The trial complied with the International Conference on Harmonization Good Clinical Practice guidelines and the 2013 Declaration of Helsinki.”
“The study received ethical approval (Camberwell St. Giles Research Ethics Committee: no. 20/LO/0657; Integrated Research Application System no. 277118)”
“All participants provided written informed consent.”
“The trial complied with the International Conference on Harmonization Good Clinical Practice guidelines and the 2013 Declaration of Helsinki.”
The AVATAR therapy software is described in detail, including its purpose and customization features, but no vendor or version is given. The paper identifies statistical software (Stata and R) but not versions. Since this is a psychological intervention trial, antibodies, cell lines, mycoplasma, and organisms are not applicable. The software used for the intervention is a key resource, but its identification is incomplete.
“The intervention AVATAR therapy is a digital treatment in which the person engages in face-to-face dialogues with a personalized digital embodiment of the voice (‘the avatar’). The avatar is presented to the person on a two-dimensional computer screen.”
“Bespoke software enables the voice-hearer to customize how the avatar looks and sounds.”
The primary analysis uses mixed-effects models with fixed effects for center, baseline, voice characterization, treatment, time, and interaction, and random intercepts for therapist and participant. Tests are named (mixed-effects models, two-sided tests). Assumptions are handled by design (mixed models, maximum likelihood). Exact p-values are reported (e.g., P = 0.035). Effect sizes with CIs are reported. Statistical software is identified (Stata, R). Data presentation includes forest plots and tables with per-group n. Mathematical plausibility is not applicable for large-N continuous outcomes.
“The primary analyses of the hypotheses of between-group differences in the AV-EXT versus TAU and AV-BRF versus TAU in voice distress as measured using the PSYRATS-AH distress score were analyzed using a mixed-effects (random) model at all post-randomization time points (weeks 16 and 28).”
“Distress at 16 weeks was as follows: AV-BRF, effect −1.05 points, 96.5% confidence interval (CI) = −2.110 to 0, P = 0.035”
“Cohen’s d = 0.38 (CI = 0 to 0.767)”
“The primary analyses of the hypotheses of between-group differences in the AV-EXT versus TAU and AV-BRF versus TAU in voice distress as measured using the PSYRATS-AH distress score were analyzed using a mixed-effects (random) model at all post-randomization time points (weeks 16 and 28).”
“Distress at 16 weeks was as follows: AV-BRF, effect −1.05 points, 96.5% confidence interval (CI) = −2.110 to 0, P = 0.035”
The data availability statement provides a concrete route: individual participant data are deposited in the King's Open Research Data System with restricted access, and requests can be made to a specific email with a review process and timeframe. This is reported_and_adequate. Repository deposit is not applicable for identifiable patient data, but the data are deposited in a repository. Accession numbers are not applicable. Code sharing is described as available through the Data Access Agreement, which is adequate.
“Individual participant data have been deposited in the King’s Open Research Data System, but access is restricted due to privacy reasons and general data protection regulations, and can only be accessed after review. Data will be made accessible after the publication of this paper. A request can be made by academic or clinical researchers to research.data@kcl.ac.uk for the purpose of conducting noncommercial, ethically approved research.”
“In accordance with the data availability protocol, the corresponding statistical code will be provided as part of the Data Access Agreement.”
“Individual participant data have been deposited in the King’s Open Research Data System, but access is restricted due to privacy reasons and general data protection regulations, and can only be accessed after review. Data will be made accessible after the publication of this paper. A request can be made by academic or clinical researchers to research.data@kcl.ac.uk”
“In accordance with the data availability protocol, the corresponding statistical code will be provided as part of the Data Access Agreement.”
“Open access information on the AVATAR2 trial, such as the trial protocol and statistical analysis plan, including the example analysis code, has been published in the ISRCTN registry with the identifier ISRCTN55682735”
The trial is registered with ISRCTN (ISRCTN55682735). Methods are detailed enough for replication. The paper references CONSORT (via the CONSORT diagram) and mentions the reporting summary. All outcomes are reported, including negative results. Limitations are explicitly discussed. Conclusions are proportional to the evidence. Funding sources and competing interests are disclosed.
“ISRCTN registration: ISRCTN55682735”
“The design of this trial had some limitations. First, the use of a TAU control meant that we could not determine the benefits of AVATAR therapy compared to another psychological treatment.”
“ISRCTN registration: ISRCTN55682735”
“The design of this trial had some limitations. First, the use of a TAU control meant that we could not determine the benefits of AVATAR therapy compared to another psychological treatment.”
Registered (3 IDs: ClinicalTrials.gov, ISRCTN). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 42 references by DOI: 0 verified — 1 DOI unresolved, 41 no DOI (shown, not verified).
- UNRESOLVED10.1016/s2215-0366(18Psychosis and Schizophrenia in Adults: Prevention and Management. Clinical Guideline [CG178]Cited DOI does not resolve to any Crossref record.
- NO DOIIdentifying research priorities for digital technology in mental health care: results of the James Lind Alliance Priority Setting PartnershipNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe digital revolution and its impact on mental health careNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISixty years of placebo-controlled antipsychotic drug trials in acute schizophrenia: systematic review, Bayesian meta-analysis, and meta-regression of efficacy predictorsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEfficacy and moderators of cognitive behavioural therapy for psychosis versus other psychological interventions: an individual-participant data meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAvatar therapy for persecutory auditory hallucinations: what is it and how does it work?No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRelating therapy for distressing auditory hallucinations: a pilot randomized controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA psychological intervention for engaging dialogically with auditory hallucinations (Talking With Voices): a single-site, randomised controlled feasibility trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAVATAR therapy for auditory verbal hallucinations in people with psychosis: a single-blind, randomised controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIVirtual reality therapy for refractory auditory verbal hallucinations in schizophrenia: a pilot clinical trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAVATAR therapy for distressing voices: a comprehensive account of therapeutic targetsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOptimising AVATAR therapy for people who hear distressing voices: study protocol for the AVATAR2 multi-centre randomised controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe role of sense of voice presence and anxiety reduction in AVATAR therapyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISymptom dimensions of the psychotic symptom rating scales in psychosis: a multisite studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPreferred treatment outcomes in psychological therapy for voices: a comparison of staff and service-user perspectivesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMeta-analysis and meta-regression of cognitive behavioral therapy for psychosis (CBTp) across time: the effectiveness of CBTp has improved for delusionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe CHALLENGE trial: the effects of a virtual reality-assisted exposure therapy for persistent auditory hallucinations versus supportive counselling in people with psychosis: study protocol for a randomised clinical trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOne-year randomized trial comparing virtual reality-assisted therapy to cognitive–behavioral therapy for patients with treatment-resistant schizophreniaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe role of characterisation in everyday voice engagement and AVATAR therapy dialogueNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIControl conditions for randomised trials of behavioural interventions in psychiatry: a decision frameworkNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe growing field of digital psychiatry: current evidence and the future of apps, social media, chatbots, and virtual realityNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffects of SlowMo, a blended digital therapy targeting reasoning, on paranoia among people with psychosis: a randomized clinical trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe EMPOWER blended digital intervention for relapse prevention in schizophrenia: a feasibility cluster randomised controlled trial in Scotland and AustraliaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIVirtual reality therapy for the negative symptoms of schizophrenia (V-NeST): a pilot randomised feasibility trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDigital Health Technologies to Help Manage Symptoms of Psychosis and Prevent Relapse in Adults and Young People: Early Value AssessmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScales to measure dimensions of hallucinations and delusions: the psychotic symptom rating scales (PSYRATS)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Voice Characterisation Checklist: psychometric properties of a brief clinical assessment of voices as social agentsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Warwick-Edinburgh Mental Well-being Scale (WEMWBS): development and UK validationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICHoice of Outcome In Cbt for psychosEs (CHOICE): the development of a new service user-led outcome measure of CBT for psychosisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAttachment styles among young adults: a test of a four-category modelNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe voices acceptance and action scale (VAAS): pilot dataNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe short‐form version of the Depression Anxiety Stress Scales (DASS‐21): construct validity and normative data in a large non‐clinical sampleNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBDI-II: Beck Depression InventoryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe power and omnipotence of voices: subordination and entrapment by voices and significant othersNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe revised Beliefs About Voices Questionnaire (BAVQ–R)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe International Trauma Questionnaire: development of a self-report measure of ICD-11 PTSD and complex PTSDNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Clinical Assessment Interview for Negative Symptoms (CAINS): final development and validationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScale for the Assessment of Positive Symptoms (SAPS)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICONSORT-SPI 2018 Explanation and Elaboration: guidance for reporting social and psychological intervention trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStata Statistical Software: Release 18No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIR: A Language and Environment for Statistical ComputingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIggplot2: Elegant Graphics for Data AnalysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
3 data/code links checked; 3 live.
- datahttps://www.isrctn.com/ISRCTN55682735?q=ISRCTN55682735&filters=&sort=&offset=1&totalResults=1&page=1&pageSize=10LIVEHTTP 200Resolved page looks like data.
- datahttps://www.isrctn.com/ISRCTN35980117?q=ISRCTN35980117&filters=&sort=&offset=1&totalResults=1&page=1&pageSize=10LIVEHTTP 200Resolved page looks like data.
- datahttps://clinicaltrials.gov/ct2/show/NCT05982158LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly clarity, consistency, other.
- MINORconsistencyAbstract“AVATAR-Brief (AV-BRF) and AVATAR-Extended (AV-EXT)”→ Ensure consistent use of hyphens and capitalization throughout.Minor inconsistency in formatting of therapy names.
- MINORclarityMethods, Statistical analysis“the statistical analysis was performed unblinded owing to the need to account for therapist effects in the AVATAR arms.”→ Clarify that the statistician was unblinded after the first DMEC report, which is already stated.The sentence is slightly redundant but not incorrect.
- MINORclarityResults, Treatment completion“For AV-BRF, the overall mean number of sessions attended was 5.11 (s.d. = 2.42; range = 0–8).”→ Clarify whether the range includes the initial assessment session.Potential ambiguity.
- MINORotherData availability“10.1017/s0033291799008661”→ This DOI appears to be for a reference, not the trial protocol; verify the correct DOI.Possible incorrect DOI in the data availability statement.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (missing power analysis in main text, unspecified software versions) but these do not undermine the validity of the findings. No erratum or correction appears necessary, though verifying the flagged DOI and the one unresolved reference would be prudent.
- 1.HIGHreportingAdd the a priori power/sample size calculation to the Methods section, including the assumed effect size, alpha, and power.The main text currently lacks a formal power analysis, which is a key element for readers to assess the trial's sensitivity and statistical basis.
- 2.HIGHdata codeVerify and correct the DOI in the Data availability statement (currently '10.1017/s0033291799008661'), which appears to be for a reference rather than the trial protocol.An incorrect DOI undermines the data access route and could mislead readers seeking the protocol.
- 3.HIGHreportingVerify the reference 'Psychosis and Schizophrenia in Adults: Prevention and Management. Clinical Guideline [CG178]' (DOI 10.1016/s2215-0366(18) that could not be found in any registry) and correct or replace it if it is erroneous.An unresolved reference may be a fabrication signal and must be resolved before the paper is relied upon.
- 4.MEDIUMrigorProvide version numbers or manufacturer identifiers for the AVATAR therapy software and the statistical software (Stata, R) in the Methods or References.Version identifiers are essential for reproducibility of the intervention and analysis.
- 5.MEDIUMreportingExplicitly state how limitations of prior research (e.g., small therapist cohort) are addressed by the current trial design in the Introduction.This would strengthen the scientific premise by directly linking prior gaps to the current design.
- 6.MEDIUMreportingAdd a statement in the Methods or a dedicated section confirming adherence to CONSORT guidelines, beyond the CONSORT diagram.Explicit reporting guideline adherence improves transparency and reader confidence.
- 7.MEDIUMstatisticsClarify the handling of missing data for secondary outcomes in the Statistical analysis section, as the MAR assumption is stated for the primary outcome but not for all outcomes.Transparent missing data handling for all outcomes is important for interpreting secondary analyses.
- 8.MEDIUMrigorDescribe how outliers were handled in the statistical analysis, or state that no outliers were excluded.Outlier handling is a standard reporting element that is currently missing.
- 9.MEDIUMstatisticsReport the intraclass correlation coefficient (ICC) for the therapist effect on the primary outcome in the Results section.The therapist clustering effect is mentioned in the text but not quantified, which would help readers understand the magnitude of clustering.
- 10.LOWcopyeditStandardize the formatting of therapy names (e.g., 'AVATAR-Brief' vs 'AVATAR-Brief') throughout the manuscript.Consistent terminology improves readability and professionalism.
- 11.LOWcopyeditClarify in the Results section whether the range of sessions attended (0–8) includes the initial assessment session.Removes ambiguity about treatment completion data.
- 12.LOWcopyeditRevise the sentence in the Statistical analysis section about unblinded analysis to avoid redundancy with the earlier statement about the DMEC report.Improves clarity without changing meaning.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.