Virtual reality-based versus standard cognitive behavioral therapy for paranoia in schizophrenia spectrum disorders: a randomized controlled trial.
Jeppesen UN, Vernal DL, Due AS, Mariegaard LS, Pinkham AE, Austin SF, Vos M, Christensen MJ, Hansen NK, Smith LC, Hjorthøj C, Veling W, Nordentoft M, Glenthøj LB
- DOI
- 10.1038/s41591-025-03880-8
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/4d5faff9-966a-4b40-be1a-fbd8c8428907 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is the Ideas of Persecution subscale of the Green Paranoid Thoughts Scale (GPTS), a self-report questionnaire measuring paranoid ideation, which is a surrogate for clinical outcomes such as functioning or quality of life. The paper does not provide evidence linking changes in this scale to hard clinical outcomes, nor does it demonstrate target engagement at the tested dose beyond the intervention itself.
“The primary outcome was the GPTS subscale Ideas of Persecution (self-report questionnaire, score range 16–80), measured at treatment cessation.”
- 02Treatment effect not shown to be clinically meaningful
The primary between-group effect is not statistically significant (Cohen's d = 0.04, P = 0.77) and is not anchored to a clinically meaningful difference. The within-group reductions are large (d = 0.97 for VR-CBTp, d = 0.75 for CBTp) but are not compared to a minimal clinically important difference, and the paper notes a possible floor effect, suggesting limited clinical meaningfulness.
“There was not a statistically significant between-group difference on the primary outcome at endpoint (effect estimate: 2% in favor of VR-CBTp; 95% confidence interval: −11% to +17%; Cohen’s d = 0.04; P = 0.77, based on exponentiated log-transformed data).”
- 03Other integrity concern
Trial NCT04902066 was first submitted to ClinicalTrials.gov on 2021-04-19, after the registered study start date of 2021-04-09. Retrospective registration means the protocol and outcomes were not on the public record before the study ran, which is what prospective registration exists to establish.
NCT04902066
reviewer’s wording
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomized controlled trial. The study design is rigorous with appropriate randomization, blinding, and power analysis. Reporting is comprehensive with data and code availability, trial registration, and disclosure of funding and conflicts.
Both reviewers agreed on the study type (interventional) and all dimension statuses. The statistics verification covered only a subset of tests (5 recomputed consistently); other tests were not machine-verified. The preregistration check noted that registration was retrospective, which is a minor transparency concern.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 5 tests: 5 consistent, 0 inconsistent; 5 via agent-written checks.
- CONSISTENTreported p = .770 · recomputed p = .777Reviewer 1Primary outcome p-value from exponentiated log-transformed effect estimate and CI
“The exponentiated log-transformed effect estimate showed a 2% lower (that is better) score in the VR-CBTp group (95% CI 11% lower for CBTp to 17% lower for VR-CBTp; Cohen’s d = 0.04; P = 0.77) (Table ).”
Taken as given: The effect estimate is 1.02 (exponentiated ratio).; The 95% CI is 0.89 to 1.17.; The p-value is two-sided from the log-transformed linear regression.Method: Used pCI function with log=1 for ratio, assuming the CI is for the exponentiated estimate.How we recomputed it: pCI(1.02, 0.89, 1.17, 1) - CONSISTENTreported p = .520 · recomputed p = .517Reviewer 1Secondary outcome GPTS Ideas of Persecution at follow-up p-value from adjusted mean difference and CI
“There was no statistically significant between-group difference on the GPTS subscale Ideas of Persecution at follow-up (adjusted mean difference 1.20, 95% CI −2.43 to 4.83; Cohen’s d = 0.08; P = 0.52) (Fig. and Table ).”
Taken as given: The adjusted mean difference is 1.20.; The 95% CI is -2.43 to 4.83.; The p-value is two-sided from the linear regression.Method: Used pCI function with log=0 for difference.How we recomputed it: pCI(1.20, -2.43, 4.83, 0) - CONSISTENTreported p = .013 · recomputed p = .013Reviewer 1Exploratory outcome COGDIS total at treatment cessation p-value from adjusted mean difference and CI
“The VR-CBTp group demonstrated a lower total score in the Cognitive Disturbances Scale (COGDIS) (adjusted mean difference 2.58, 95% CI 0.55–4.62; Cohen’s d = 0.31; P = 0.013).”
Taken as given: The adjusted mean difference is 2.58.; The 95% CI is 0.55 to 4.62.; The p-value is two-sided from the linear regression.Method: Used pCI function with log=0 for difference.How we recomputed it: pCI(2.58, 0.55, 4.62, 0) - CONSISTENTreported p = .009 · recomputed p = .009Reviewer 2Loss to follow-up at treatment cessation p-value
“At treatment cessation, when the primary outcome was measured, 9 participants (7%) in the VR-CBTp group and 23 participants (18%) in the CBTp group were lost to follow-up, a difference that was statistically significant ( P = 0.009)”
Taken as given: The numbers 9 and 23 are the counts lost to follow-up in each group.; The group totals are 126 and 128, respectively.; The test used is a chi-square test for independence.Method: Recomputed p-value using Pearson's chi-square test on the 2x2 table.How we recomputed it: pChi2x2(9, 117, 23, 105) - CONSISTENTreported p = .076 · recomputed p = .076Reviewer 2Loss to follow-up at follow-up p-value
“At follow-up, 21 participants (17%) were lost to follow-up in the VR-CBTp group and 33 (26%) in the CBTp group, which was not statistically significant ( P = 0.076)”
Taken as given: The numbers 21 and 33 are the counts lost to follow-up in each group.; The group totals are 126 and 128, respectively.; The test used is a chi-square test for independence.Method: Recomputed p-value using Pearson's chi-square test on the 2x2 table.How we recomputed it: pChi2x2(21, 105, 33, 95)
- lowinternal contradictionThe abstract reports 'effect estimate: 2% in favor of VR-CBTp; 95% confidence interval: −11% to +17%' while the results section reports 'exponentiated log-transformed effect estimate showed a 2% lower (that is better) score in the VR-CBTp group (95% CI 11% lower for CBTp to 17% lower for VR-CBTp)'. The direction of the CI is described differently but may be consistent.
“effect estimate: 2% in favor of VR-CBTp; 95% confidence interval: −11% to +17%”
AbstractFind in source - lowinternal contradictionThe paper reports that the first assessment at treatment cessation was 29 June 2021 and the final follow-up assessment occurred on 10 August 2024, when we reached a total of 256 participants. However, the total number of participants included in analyses is 254, and the text earlier says the target was 256.
“First assessment at treatment cessation was 29 June 2021 and the final follow-up assessment occurred on 10 August 2024, when we reached a total of 256 participants.”
ResultsFind in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
6 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2VR-CBTp was not superior to CBTp in reducing paranoia.The primary outcome analysis showed no statistically significant difference, and sensitivity analyses were consistent.Evidence: Primary outcome effect estimate 2% (95% CI -11% to +17%, P=0.77); sensitivity analyses and per-protocol analyses showed no significant differences.
“In conclusion, VR-CBTp was not superior to CBTp in reducing schizophrenia-spectrum-disorders-related paranoia.”
AbstractFind in source - supportedReviewer 1No deaths or violent incidents involving law enforcement occurred during the study.The safety monitoring reported no such events.Evidence: Safety section states 'Neither deaths nor violent incidents involving law enforcement were reported during the study'.
“Neither deaths nor violent incidents involving law enforcement were reported during the study, and there were no instances in which participation in our study was linked to a suicide attempt.”
Safety sectionFind in source - supportedReviewers 1, 2VR-CBTp produced a large reduction in paranoia (Cohen's d=0.97) while CBTp yielded a moderate reduction (d=0.75).The within-group effect sizes are reported from post hoc analysis, but the paper notes they are not directly comparable due to different scales.Evidence: Post hoc analysis reported within-group effect sizes.
“VR-CBTp produced a large reduction (Cohen’s d = 0.97; 29.8% reduction), whereas CBTp yielded a moderate reduction ( d = 0.75; 26.3% reduction).”
Post hoc analysesFind in source - supportedReviewer 1The study contrasts with previous trials using wait-list or passive controls.The paper notes that previous trials used wait-list controls, while this trial used an active comparator.Evidence: Abstract states 'in contrast to previous trials using wait-list or passive controls'.
“A randomized controlled trial found no difference in paranoid ideations between virtual reality-based and gold-standard cognitive behavioral therapy for patients with schizophrenia spectrum disorders, in contrast to previous trials using wait-list or passive controls.”
AbstractFind in source - supportedReviewer 2The trial found no difference in paranoid ideations between VR-CBTp and CBTp, in contrast to previous trials using wait-list or passive controls.The primary outcome and sensitivity analyses support this claim.Evidence: Primary outcome non-significant; sensitivity analyses consistent.
“A randomized controlled trial found no difference in paranoid ideations between virtual reality-based and gold-standard cognitive behavioral therapy for patients with schizophrenia spectrum disorders, in contrast to previous trials using wait-list or passive controls.”
AbstractFind in source - supportedReviewer 2The VR-CBTp group had a lower COGDIS total score at treatment cessation (P = 0.013).This exploratory finding was statistically significant in the primary analysis but not in sensitivity analysis.Evidence: Exploratory outcome COGDIS total score.
“The VR-CBTp group demonstrated a lower total score in the Cognitive Disturbances Scale (COGDIS) (adjusted mean difference 2.58, 95% CI 0.55–4.62; Cohen’s d = 0.31; P = 0.013).”
Exploratory outcomesFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is the Ideas of Persecution subscale of the Green Paranoid Thoughts Scale (GPTS), a self-report questionnaire measuring paranoid ideation, which is a surrogate for clinical outcomes such as functioning or quality of life. The paper does not provide evidence linking changes in this scale to hard clinical outcomes, nor does it demonstrate target engagement at the tested dose beyond the intervention itself.
“The primary outcome was the GPTS subscale Ideas of Persecution (self-report questionnaire, score range 16–80), measured at treatment cessation.”
- INADEQUATEEffect sizeThe primary between-group effect is not statistically significant (Cohen's d = 0.04, P = 0.77) and is not anchored to a clinically meaningful difference. The within-group reductions are large (d = 0.97 for VR-CBTp, d = 0.75 for CBTp) but are not compared to a minimal clinically important difference, and the paper notes a possible floor effect, suggesting limited clinical meaningfulness.
“There was not a statistically significant between-group difference on the primary outcome at endpoint (effect estimate: 2% in favor of VR-CBTp; 95% confidence interval: −11% to +17%; Cohen’s d = 0.04; P = 0.77, based on exponentiated log-transformed data).”
Data authenticity concerns
1 finding · worst mediumAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
3 integrity concerns flagged (0 high).
- mediumotherTrial NCT04902066 was first submitted to ClinicalTrials.gov on 2021-04-19, after the registered study start date of 2021-04-09. Retrospective registration means the protocol and outcomes were not on the public record before the study ran, which is what prospective registration exists to establish.
NCT04902066
reviewer’s wording
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on CBTp, VR interventions, and symptom-specific approaches, acknowledging both strengths and limitations. The rationale linking VR's controlled environments to improved exposure therapy is well-argued, and the hypothesis follows logically. Limitations of prior studies (e.g., wait-list controls) are explicitly noted.
“A previous umbrella review of meta-analyses concluded that CBTp for delusions and other psychotic symptoms yields small to medium effect sizes compared to treatment as usual (TAU) at treatment cessation; however, these effects were not maintained after 6–12 months.”
“VR employs computer-generated simulations to immerse users in interactive three-dimensional environments, typically via a headset that tracks movement and dynamically adjusts the scene in real time.”
“A randomized controlled trial found no difference in paranoid ideations between virtual reality-based and gold-standard cognitive behavioral therapy for patients with schizophrenia spectrum disorders, in contrast to previous trials using wait-list or passive controls.”
“A previous umbrella review of meta-analyses concluded that CBTp for delusions and other psychotic symptoms yields small to medium effect sizes compared to treatment as usual (TAU) at treatment cessation; however, these effects were not maintained after 6–12 months.”
“A randomized controlled trial found no difference in paranoid ideations between virtual reality-based and gold-standard cognitive behavioral therapy for patients with schizophrenia spectrum disorders, in contrast to previous trials using wait-list or passive controls.”
The trial is described as assessor-masked and randomized, with a power analysis targeting 256 participants. Inclusion/exclusion criteria are detailed, and the analysis population (ITT) is defined. Blinding is described as assessor-masked, and the randomization method is mentioned (set up by a statistician).
“C.H. set up the randomization program and served as the independent trial statistician, solely responsible for setting up the randomization module and conducting all statistical analyses, ensuring objectivity throughout the study.”
“This assessor-masked, randomized parallel group superiority trial investigated the efficacy of VR-CBTp compared to standard CBTp.”
“However, five withdrew their consent later in the study (VR-CBTp: n = 2, CBTp: n = 3), including two who withdrew late in the study period, preventing us from reaching the target sample size of 256.”
“C.H. set up the randomization program and served as the independent trial statistician, solely responsible for setting up the randomization module and conducting all statistical analyses, ensuring objectivity throughout the study.”
“This assessor-masked, randomized parallel group superiority trial investigated the efficacy of VR-CBTp compared to standard CBTp.”
“However, five withdrew their consent later in the study (VR-CBTp: n = 2, CBTp: n = 3), including two who withdrew late in the study period, preventing us from reaching the target sample size of 256.”
Sex is reported for both groups (57.1% female in VR-CBTp, 57.8% in CBTp). Age is reported as median with IQR. Health status is captured via diagnosis (ICD-10 F20-29) and symptom severity. Demographics include education, occupational status, and substance use. Since both sexes are enrolled, sex_justified is not applicable.
“Female | 72 (57.1%) | 74 (57.8%) | 146 (57.5%)”
“Age, years | 27.7 (23.0–35.0) n = 126 | 26.5 (22.7–31.8) n = 128 | 26.8 (22.8–33.1) n = 254”
“F20 Schizophrenia | 94 (74.6%) | 90 (70.3%) | 184 (72.4%)”
“Female | 72 (57.1%) | 74 (57.8%) | 146 (57.5%)”
“Age, years | 27.7 (23.0–35.0) n = 126 | 26.5 (22.7–31.8) n = 128”
“F20 Schizophrenia | 94 (74.6%) | 90 (70.3%) | 184 (72.4%)”
The paper mentions that the research ethics committee approved the study and that participants completed informed consent. The protocol was published in Trials, and the trial was registered. The ethics statement is adequate.
“Serious adverse events were continuously monitored throughout the trial and reported to the Principal Investigator (PI), the lead therapist, the data monitoring committee and the research ethics committee.”
“A total of 259 participants completed the informed consent process and were enrolled and randomized.”
“Due to the Danish Archives Act (Arkivloven), the Danish Archives Executive Order (Arkivbekendtgørelsen), the General Data Protection Regulation (Databeskyttelsesforordningen) and the Danish Data Protection Act (Databeskyttelsesloven), access is restricted.”
“A total of 259 participants completed the informed consent process and were enrolled and randomized.”
“Serious adverse events were continuously monitored throughout the trial and reported to the Principal Investigator (PI), the lead therapist, the data monitoring committee and the research ethics committee.”
The paper describes the interventions (VR-CBTp and CBTp) in detail, including the VR environment and therapeutic approach. Software tools such as E-Prime and REDCap are mentioned. No antibodies, cell lines, or organisms are used, so those criteria are not applicable.
“Participants were randomized to receive ten sessions of VR-CBTp or CBTp, both on top of treatment as usual.”
“U.N.J. calculated intraclass correlations and conducted IBT calculations on E-prime data.”
“Psychology Software Tools. E-Prime ® Stimulus Presentation Software https://pstnet.com/products/e-prime/ (2025).”
The paper reports effect estimates with 95% CIs and p-values, uses linear regression models with multiple imputations, and describes sensitivity analyses. The primary outcome is reported with an effect estimate and CI, and the analysis is based on ITT. The paper also reports exact p-values for many outcomes.
“Cohen’s d = 0.04; P = 0.77”
“effect estimate: 2% in favor of VR-CBTp; 95% confidence interval: −11% to +17%”
“As residual plots indicated a non-normal distribution on the primary outcome, the GPTS subscale Ideas of Persecution at treatment cessation, we applied log transformation, which improved model fit.”
“The analysis is a linear regression model based on the ITT principle and handled with multiple imputations.”
“effect estimate: 2% in favor of VR-CBTp; 95% confidence interval: −11% to +17%; Cohen’s d = 0.04; P = 0.77”
Data availability statement is concrete: deidentified data available through Danish National Archives with a clear access procedure. Code is available on Codeberg. Clinical trial registration is provided. Repository deposit is not applicable for patient-level data, but the data access route is adequate.
“All deidentified trial data are available through the Danish National Archives (Rigsarkivet), a public data repository, for an unlimited period.”
“The code is available at https://codeberg.org/VIRTU/faceyourfears .”
“ClinicalTrials.gov registration: NCT04902066 (https://clinicaltrials.gov/ct2/show/NCT04902066) .”
The trial is registered (NCT04902066). A reporting summary is mentioned. All outcomes, including negative and null results, are reported. Limitations are discussed. Conclusions are proportional to the evidence. Funding and COI are disclosed.
“ClinicalTrials.gov registration: NCT04902066 (https://clinicaltrials.gov/ct2/show/NCT04902066) .”
“Further information on research design is available in the linked to this article.”
“The study funders were TrygFoundation (ID: 148727) (M.N.), Independent Research Fund Denmark (0134-00066B) (M.N.), Research Fund of the Mental Health Services—Capital Region of Denmark (PhD grant) (L.B.G. and U.N.J.), Research and Fund for Health Research 2019—Capital Region of Denmark (A6622) (L.B.G.), Innovation Fund North Denmark Region (2022-0010) (M.J.C.), Psychiatry Research Fund North Denmark Region (1-45-72-3778-24) (D.L.V. and M.J.C.), The M. L. Jørgensen and Gunnar Hansen Fund (2022-0019) (M.J.C.) and The A. P. Moller Foundation (L-2021-00244) (M.J.C.) and they had no role in the study design, data collection, analysis or interpretation and no role in writing the article.”
“Further information on research design is available in the linked to this article.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 72 references by DOI: 62 verified — 10 no DOI (shown, not verified).
- NO DOIAdvances in the use of virtual reality to treat mental health conditionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAssessing depression in schizophrenia: the Calgary depression scaleNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEQ-5D-5LNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe experience of presence: factor analytic insightsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDigital Technologies for Managing Symptoms of Psychosis and Preventing Relapse: Early Value AssessmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPsychosis and Schizophrenia in AdultsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIManagement of Schizophrenia (SIGN 131)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICognitive Therapy for Delusions, Voices, and ParanoiaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIE-Prime ® Stimulus Presentation SoftwareNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe development of markers for the Big-Five factor structureNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
3 data/code links checked; 3 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT04902066LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.rigsarkivet.dkLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codehttps://codeberg.org/VIRTU/faceyourfearsLIVEHTTP 200Resolved page looks like code.
Copyediting
7 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 7 minor suggestions below.
7 copyedit issues flagged: mostly consistency, clarity, grammar.
- MINORconsistencyTable 1, Big-5 row“3.14 (2.99–3.30%) n = 126”→ Remove the percentage sign from the CI values; it appears to be a formatting artifact.The percentage sign is incorrectly appended to the confidence interval values.
- MINORconsistencyTable 2, CDSS row“0.99 a (0.23) | (0.63–1.57%) a ; 0.97 | 0 b”→ Clarify the footnote markers and ensure the CI is correctly formatted.The footnote markers and CI formatting are confusing.
- MINORgrammarResults, Exploratory outcomes“which improved their model fit’s.”→ Change to 'which improved their model fit'.Apostrophe misuse.
- MINORclarityResults, Patient disposition“First assessment at treatment cessation was 29 June 2021 and the final follow-up assessment occurred on 10 August 2024, when we reached a total of 256 participants.”→ Clarify that the total of 256 refers to the target sample size, not the number of assessments.The sentence is ambiguous.
- MINORconsistencyTable 1 footnote“n ( n %), indicating the frequency and its corresponding percentage of the total; n ( n – n %) indicates the mean value and its 95% CI; n ( n – n ) indicates the median and its 25th and 75th percentiles”→ Clarify the notation for mean and median, as the current phrasing is confusing.The notation for mean and median is ambiguous.
- MINORtypoTable 2, CDSS row“0.99 a (0.23) | (0.63–1.57%) a ; 0.97 | 0 b”→ Check the p-value and CI; the CI appears to be for a log-transformed estimate, but the p-value seems inconsistent.Potential inconsistency in the reported p-value and CI.
- MINORclarityResults, Primary outcomes“The exponentiated log-transformed effect estimate showed a 2% lower (that is better) score in the VR-CBTp group (95% CI 11% lower for CBTp to 17% lower for VR-CBTp; Cohen’s d = 0.04; P = 0.77)”→ Clarify the direction of the effect estimate, as the phrasing is confusing.The direction of the effect is unclear.
The published work is robust and well-reported. An informed reader should weigh the retrospective registration and the minor internal inconsistencies (e.g., sample size discrepancy) as minor concerns, but they do not undermine the overall validity. No erratum is warranted for the core findings.
- 1.HIGHreportingClarify in the Results section (Patient disposition) that the total of 256 refers to the target sample size, not the number of assessments, and reconcile with the 254 participants analyzed.The current sentence is ambiguous and conflicts with the reported analysis sample of 254.
- 2.HIGHreportingClarify the direction of the effect estimate in the Results (Primary outcomes) to match the abstract's phrasing, ensuring the CI direction is unambiguous.The abstract and results describe the CI direction differently, which could confuse readers.
- 3.MEDIUMreportingAdd a note in the Methods or Data Availability that trial registration was retrospective, and explain any protocol changes made after registration.Transparency about retrospective registration is important for readers assessing the pre-specification of outcomes.
- 4.MEDIUMcopyeditFix the formatting artifact in Table 1 (Big-5 row) by removing the stray percentage sign from the CI values.The percentage sign is incorrectly appended and could be misinterpreted.
- 5.MEDIUMcopyeditClarify the footnote markers and CI formatting in Table 2 (CDSS row) to ensure the p-value and CI are correctly presented.The current formatting is confusing and may lead to misreading of the statistics.
- 6.MEDIUMcopyeditCorrect the grammar in Results (Exploratory outcomes): change 'which improved their model fit’s' to 'which improved their model fit'.Apostrophe misuse is a minor but visible error.
- 7.MEDIUMcopyeditClarify the notation in the Table 1 footnote for mean and median values to avoid ambiguity.The current phrasing is confusing and could be misinterpreted.
- 8.MEDIUMreportingSpecify the VR hardware and software versions used in the Methods to enhance reproducibility.The current description lacks specific versions, which limits replication.
- 9.MEDIUMreportingProvide more detail on the randomization sequence generation (e.g., block size, stratification factors) in the Methods.The current description is vague and does not fully describe allocation concealment.
- 10.MEDIUMreportingExplicitly name the statistical software and version used for analyses in the Methods.The software is implied but not explicitly stated, which is a reproducibility gap.
- 11.LOWreportingConsider reporting the number of participants with missing data for each outcome at each time point in the main text.Currently this is only in supplementary tables, which reduces transparency.
- 12.LOWreportingDiscuss the generalizability of findings given the exclusion of non-Danish speakers in the Discussion.This limitation is noted in protocol deviations but not discussed in the main text.
- 13.LOWreportingClarify the definition of 'treatment as usual' and any standardization across sites in the Methods.This is important for understanding the comparator and generalizability.
- 14.LOWreportingAddress the potential impact of the higher dropout rate in the CBTp group on the interpretation of results in the Discussion.Even with ITT analyses, differential dropout can bias results.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.