A socially assistive robot to support mental wellbeing in LGBTQ+ young people at risk of self-harm: a randomized controlled trial.
Williams AJ, Rhodes CA, Cleare S, Borschmann R, Gross JJ, Petrova K, Posada L, Tench CR, Chapman-Nisar A, Martin L, Hollis C, Townsend E, Slovak P, Digital Youth research team
- DOI
- 10.1038/s41591-026-04422-6
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/b93488d1-1751-44e9-80f9-d1115be95d01 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingKey resources not met−0.5★
- ReportingData & code availability not met−0.5★
- ReportingEthical approvals partially met−0.25★
- References were not verified against Crossref/OpenAlex.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on a surrogate endpoint: the Difficulties in Emotion Regulation Scale (DERS-8), a self-report questionnaire measuring perceived emotion regulation difficulties. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose–exposure data) nor cite validated evidence linking changes in DERS-8 to a hard clinical outcome such as self-harm, suicide, or hospitalization. The secondary outcomes (PHQ-9, GAD-7) are also symptom scales, not hard clinical endpoints.
“The primary outcome was perceived emotion regulation difficulties at follow-up, adjusted for baseline, gender identity and age.”
- 02Treatment effect not shown to be clinically meaningful
The primary effect is a mean difference of –3.04 on the DERS-8 (range 8–40). The paper does not anchor this to a minimal clinically important difference (MCID) or a normative reference value. The effect is presented as statistically significant but lacks explicit clinical meaningfulness. The reliable change index (29% vs 14%) is provided but still not anchored to a clinically meaningful threshold.
“adjusted mean difference: –3.04; 95% confidence interval (CI): −4.92 to −1.16; P = 0.002; partial η 2 = 0.07”
- 03Key resources not identified
The investigational product (Purrble) is not fully identified (no manufacturer, source, or catalog number), and statistical software is not identified. This is a key reporting gap.
“weekly online surveys hosted by Qualtrics”
MethodsFind in source - 04Data and code not shared
No data availability statement is provided. No data or code repositories are reported. This is a significant omission for a clinical trial.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper presents a well-conducted RCT with a strong scientific premise, clear study design, and transparent reporting of outcomes. However, it has several reporting gaps: no power analysis, no data availability statement, no statistical software or device source identified, and missing statements on regulatory compliance and conflicts of interest. The statistical recomputations are consistent.
Two independent runs of the same AI model were used; they converged on most dimensions but diverged on ethical approvals (pass vs. warn) and on the checklist for biological variables (sex reporting). The synthesis weighted the stricter evidence where applicable. Verification components covered 10 tests (all consistent), 5 reproducibility links (4 live), and 0 flagged citations.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 10 tests: 10 consistent, 0 inconsistent; 3 recomputed directly from the reported test statistics, 7 via agent-written checks.
- CONSISTENTreported p = .030 · recomputed p = .031Recomputed odds ratio 2.44 (95% CI 1.09–5.49), reported p=0.03
“odds ratio = 2.44, 95% CI: 1.09−5.49, P = 0.03”
Taken as given: 1.09–5.49 is a two-sided 95% confidence interval for the odds ratio of 2.44, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.03 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.44, 1.09, 5.49, 1) - CONSISTENTreported p = .001 · recomputed p = .001Recomputed odds ratio 4.62 (95% CI 1.85–11.53), reported p=0.001
“odds ratio = 4.62, 95% CI: 1.85−11.53, P = 0.001”
Taken as given: 1.85–11.53 is a two-sided 95% confidence interval for the odds ratio of 4.62, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(4.62, 1.85, 11.53, 1) - CONSISTENTreported p = .079 · recomputed p = .079Recomputed odds ratio 2.01 (95% CI 0.92–4.36), reported p=0.079
“odds ratio = 2.01, 95% CI: 0.92−4.36, P = 0.079”
Taken as given: 0.92–4.36 is a two-sided 95% confidence interval for the odds ratio of 2.01, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.079 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.01, 0.92, 4.36, 1) - CONSISTENTreported p = .002 · recomputed p = .002Reviewer 1Primary outcome ANCOVA: condition effect t = -3.20, df = 134, reported P = 0.002
“Condition | −3.04** | −4.92 | −1.16 | 0.95 | −3.20 | 0.002 | 0.07 | 0.01 | 0.17”
Taken as given: The t statistic is 3.20 (absolute value) with 134 residual degrees of freedom.; The p-value is two-tailed.Method: Two-tailed t-test p-value from t=3.20, df=134.How we recomputed it: pT(3.20,134) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Depression ANCOVA: condition effect t = -3.64, df = 134, reported P < 0.001
“Condition | −2.60*** | −4.02 | −1.19 | 0.71 | −3.64 | <0.001 | 0.09 | 0.02 | 0.19”
Taken as given: The t statistic is 3.64 (absolute value) with 134 residual degrees of freedom.; The p-value is two-tailed.Method: Two-tailed t-test p-value from t=3.64, df=134.How we recomputed it: pT(3.64,134) - CONSISTENTreported p = .038 · recomputed p = .038Reviewer 1Emotion regulation interaction ANCOVA: condition x gender identity t = 2.10, df = 134, reported P = 0.038
“Condition × gender identity | 3.92* | 0.22 | 7.63 | 1.87 | 2.10 | 0.038 | 0.03 | 0.00 | 0.11”
Taken as given: The t statistic is 2.10 with 134 residual degrees of freedom.; The p-value is two-tailed.Method: Two-tailed t-test p-value from t=2.10, df=134.How we recomputed it: pT(2.10,134) - CONSISTENTreported p = .750 · recomputed p = .740Reviewer 1Attrition chi-square test: chi-square = 0.11, df = 1, reported P = 0.75
“Binary attrition rates did not differ significantly by condition, χ 2 1 = 0.11, P = 0.75”
Taken as given: The chi-square statistic is 0.11 with 1 degree of freedom.; The test is two-tailed (usual for chi-square).Method: Chi-square p-value from chi2=0.11, df=1.How we recomputed it: pChi2(0.11,1) - CONSISTENTreported p = .030 · recomputed p = .032Reviewer 1Reliable improvement in emotion regulation: Fisher exact test from 2x2 table (Purrble+SP 22/76 vs SP-Only 11/77), reported P = 0.03
“Emotion regulation | Purrble + SP | 76 | ... | 22 (28.9%) | 2.44 [1.09−5.49] | 0.03 | ...”
Taken as given: The 22 and 11 are the numbers of reliable improvement events in Purrble+SP and SP-Only, respectively.; The 76 and 77 are the group totals, so non-events are 54 and 66.; The p-value is two-sided from Fisher's exact test.Method: Two-sided Fisher's exact test from cell counts (22,54,11,66).How we recomputed it: pFisher2x2(22,54,11,66,0) - CONSISTENTreported p = .002 · recomputed p = .002Reviewer 2Verification of primary outcome ANCOVA t-value and p-value for emotion regulation.
“B = −3.04, 95% CI: −4.92 to −1.16, s.e. = 0.95, P = 0.002; t = −3.20, df = 134”
Taken as given: The reported t value is -3.20 with 134 degrees of freedom.; The test is two-tailed.; The p-value is two-sided from the t-distribution.Method: Two-tailed t-test from t-distribution with 134 df.How we recomputed it: 2 * (1 - tCdf(3.20, 134)) - CONSISTENTreported p = .044 · recomputed p = .043Reviewer 2Verification of anxiety ANCOVA t-value and p-value.
“B = −1.35, 95% CI: −2.66 to −0.04, s.e. = 0.66, P = 0.044; t = −2.04, df = 134”
Taken as given: The reported t value is -2.04 with 134 degrees of freedom.; The test is two-tailed.; The p-value is two-sided from the t-distribution.Method: Two-tailed t-test from t-distribution with 134 df.How we recomputed it: 2 * (1 - tCdf(2.04, 134))
- lowinternal contradictionThe abstract states 'adjusted mean difference: –3.04' while the results table reports B = -3.04; these are consistent, but the abstract uses 'mean difference' whereas the model is an ANCOVA coefficient, which is a subtle terminology difference.
“adjusted mean difference: –3.04”
AbstractFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
7 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewer 1Gender identity moderated the effects of Purrble + SP on emotion regulation difficulties and anxiety.The interaction terms were significant in the models (P = 0.038 for emotion regulation, P = 0.031 for anxiety) but none remained significant after Benjamini–Hochberg correction, so the claim is not robust.Evidence: Interaction analyses in Results and Discussion.
“This model revealed a significant interaction (B = 3.92, 95% CI: 0.22−7.63, s.e. = 1.87, P = 0.038, partial η 2 = 0.03).”
ResultsFind in source - partialReviewer 2Taken together, these findings position Purrble as a possible adjunct to usual care or waitlist periods for cisgender sexual minority youth while underscoring the need for targeted co-design and evaluation to meet the distinct needs of TGD youth.The claim is partially supported because the cisgender benefit is from post hoc analyses that were not robust to multiplicity correction, but the pattern is consistent. The paper acknowledges the need for further research.Evidence: Post hoc interaction analyses showed significant condition × gender identity interaction for emotion regulation (P=0.038) and anxiety (P=0.031), but not after Benjamini-Hochberg correction. Simple slopes showed benefit only for cisgender participants.
“Taken together, these findings position Purrble as a possible adjunct to usual care or waitlist periods for cisgender sexual minority youth while underscoring the need for targeted co-design and evaluation to meet the distinct needs of TGD youth.”
Discussion ¶4Find in source - supportedReviewers 1, 2Participants allocated to the Purrble intervention reported fewer emotion regulation difficulties at follow-up than those allocated to safety planning alone.The ANCOVA result (B = −3.04, 95% CI −4.92 to −1.16, P = 0.002) directly supports this claim.Evidence: Primary outcome ANCOVA result in Results section.
There was a significant main effect of condition such that participants assigned to Purrble + SP reported fewer emotion regulation difficulties as measured by DERS-8 at follow-up ... (B = −3.04, 95% CI: −4.92 to −1.16, s.e. = 0.95, P = 0.002; partial η 2 = 0.07)
Resultsreviewer’s wording - supportedReviewer 1Participants in the Purrble intervention also reported significantly lower symptoms of anxiety and depression.The ANCOVA models for depression (B = −2.60, P < 0.001) and anxiety (B = −1.35, P = 0.044) support this claim.Evidence: Secondary outcome ANCOVA results.
There was a significant main effect of condition such that participants assigned to Purrble + SP reported lower depressive symptom severity ... (B = −2.60, 95% CI: −4.02 to −1.19, s.e. = 0.71, P < 0.001; partial η 2 = 0.09)
Resultsreviewer’s wording - supportedReviewer 1No significant main effect was observed for self-harm.The paper explicitly states no significant main effects on self-harm were observed.Evidence: Results, Secondary outcomes
No significant main effects on self-harm were observed as measured by SHQ screening questions.
Resultsreviewer’s wording - supportedReviewers 1, 2Purrble may offer a scalable intervention to complement existing therapeutic approaches.The claim is phrased as a suggestion ('may offer') and is consistent with the positive primary outcome and the discussion of cost and low burden.Evidence: Abstract and Discussion.
“Purrble may offer a scalable intervention to complement existing therapeutic approaches to support LGBTQ+ youth to enhance their emotion regulation.”
AbstractFind in source - supportedReviewer 2For secondary outcomes, participants in the Purrble intervention also reported significantly lower symptoms of anxiety and depression.Depression (PHQ-9) showed a significant effect (P<0.001); anxiety (GAD-7) showed a significant effect (P=0.044). Self-harm was not significant, which is acknowledged.Evidence: Depression: B=−2.60, 95% CI −4.02 to −1.19, P<0.001; Anxiety: B=−1.35, 95% CI −2.66 to −0.04, P=0.044.
“For secondary outcomes, participants in the Purrble intervention also reported significantly lower symptoms of anxiety and depression, but no significant main effect was observed for self-harm.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on a surrogate endpoint: the Difficulties in Emotion Regulation Scale (DERS-8), a self-report questionnaire measuring perceived emotion regulation difficulties. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose–exposure data) nor cite validated evidence linking changes in DERS-8 to a hard clinical outcome such as self-harm, suicide, or hospitalization. The secondary outcomes (PHQ-9, GAD-7) are also symptom scales, not hard clinical endpoints.
“The primary outcome was perceived emotion regulation difficulties at follow-up, adjusted for baseline, gender identity and age.”
- INADEQUATEEffect sizeThe primary effect is a mean difference of –3.04 on the DERS-8 (range 8–40). The paper does not anchor this to a minimal clinically important difference (MCID) or a normative reference value. The effect is presented as statistically significant but lacks explicit clinical meaningfulness. The reliable change index (29% vs 14%) is provided but still not anchored to a clinically meaningful threshold.
“adjusted mean difference: –3.04; 95% confidence interval (CI): −4.92 to −1.16; P = 0.002; partial η 2 = 0.07”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
3 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data and code not sharedAssessed
- Key resources not identifiedAssessed
- Ethics/consent reporting incompleteAssessed
Prior research is extensively cited, including statistics on self-harm in LGBTQ+ youth, reviews of digital interventions, and pilot work on Purrble. The rationale linking emotion regulation to the intervention is clearly articulated, and the paper explicitly addresses a limitation of prior research by examining gender identity moderation. All three sub-criteria are reported adequately.
“However, to date, no randomized controlled trials evaluating the efficacy of the Purrble intervention have been conducted.”
Randomization method (online generator, stratified by gender identity) and unit (individual) are described. Blinding is addressed as an open-label design with rationale (participants aware of allocation). Inclusion/exclusion criteria are pre-specified. Outlier handling (Mahalanobis distance) and missing-data sensitivity analyses are reported. The power analysis is not reported in the text, only that the target minimum sample size was achieved, which is a gap.
“Within each block, participants were individually randomized in a 1:1 ratio (stratified by gender identity and cisgender versus TGD, to ensure balance between conditions) to either the Purrble + SP condition or the SP-Only (plus waitlist) condition, using an online randomization generator.”
“Participants were individually randomized in a 1:1 ratio (stratified by gender identity and cisgender versus TGD, to ensure balance between conditions) to either the Purrble + SP condition or the SP-Only (plus waitlist) condition, using an online randomization generator.”
Age is reported as mean and SD. Gender identity is reported (49.7% cisgender, 50.3% TGD). Demographics include sexual orientation and race/ethnicity. Health status is reflected in baseline DERS-8, PHQ-9, and GAD-7 scores. Since the sample includes both cisgender and TGD participants, a single-sex justification is not required.
The paper states approval by King's College London ethics committee (RESCM-22/23-34570) and that all participants provided written informed consent. However, it does not mention adherence to the Declaration of Helsinki, ICH-GCP, or other specific regulations. The CONSORT checklist is mentioned but not a regulatory compliance statement.
“All study procedures were approved by the ethics committee at King’s College London (RESCM-22/23-34570)”
“all participants provided written informed consent prior to participation.”
“All participants were aged 16 years or older and provided their own informed consent, in line with UK regulations.”
“All study procedures were approved by the ethics committee at King’s College London (RESCM-22/23-34570)”
“all participants provided written informed consent prior to participation.”
Purrble is described as a socially assistive robot but no vendor, model, or source is provided. The paper mentions 'Purrble devices were dispatched' but does not state where they were obtained from. Statistical software used for analyses (e.g., R, SPSS, Stata) is not named or versioned. Qualtrics is mentioned for surveys but without version. Other software (e.g., for randomization) is not identified.
“weekly online surveys hosted by Qualtrics”
Tests (ANCOVA, linear mixed effects models, chi-square) are named. Exact p-values are given (e.g., P = 0.002, P = 0.044). Effect sizes with 95% CIs are reported. Data presentation includes tables with means, SDs, n, and CIs, and a CONSORT diagram. Software is not identified. Mathematical plausibility checks did not reveal demonstrable errors.
“Condition | −3.04** | −4.92 | −1.16 | 0.95 | −3.20 | 0.002 | 0.07 | 0.01 | 0.17”
The paper does not include a data availability statement. There is no mention of deposition of data in a public repository, no accession numbers, and no code sharing. The trial registration number is provided, but that is not a data availability statement.
Trial registration (NCT06025942) is reported. CONSORT 2010 checklist is referenced. All pre-specified outcomes, including non-significant self-harm results, are reported. Limitations are extensively discussed. Conclusions are cautious. Funding is disclosed (MRC grants), but no explicit COI statement is present in the excerpt.
“ClinicalTrials.gov: NCT06025942 (http://clinicaltrials.gov/study/NCT06025942)”
“This paper was developed in accordance with the CONSORT 2010 checklist”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
5 data/code links checked; 4 live.
- datahttp://clinicaltrials.gov/study/NCT06025942LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT06025942LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://bmjopen.bmj.com/content/bmjopen/14/1/e079801.full.pdfUNVERIFIEDHTTP 403Liveness indeterminate — content not checked.
- codeGitHubLIVEHTTP 200https://github.com/caarhodes/PurrbleResolves to GitHub (code repository).
- codeGitHubLIVEHTTP 200http://github.com/caarhodes/PurrbleResolves to GitHub (code repository).
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, clarity, punctuation.
- MINORconsistencyAbstract and Results“adjusted mean difference: –3.04; 95% CI: −4.92 to −1.16; P = 0.002”→ Consider using consistent notation for negative signs (en dash vs hyphen).Minor formatting inconsistency.
- MINORclarityTable 2“Data are presented as mean (s.d.) or frequency (%). n indicates the number of participants.”→ Clarify that the reliable change percentages are based on the same n as the overall condition.Potential ambiguity in table footnote.
- MINORconsistencyReferences, e.g., ref 19“Kazdin, A. E. & Kitt, E. R. Influence of a socially assistive robot on mood, anxiety and arousal in children. Prof. Psychol. Res. Pract. 49 , 48–56 (2018).”→ Remove extra space before the comma: '49, 48–56'.Extra space before comma in journal volume.
- MINORpunctuationResults, Table 3 footnote“The test statistic reported for each coefficient is t with 134 residual degrees of freedom for all models.”→ Consider adding a period after 't' for clarity: 't statistic' or 't-value'.Minor clarity issue.
This paper is published but has reporting gaps that a reader should weigh: no power analysis, no data availability statement, no statistical software or device source identified, and no explicit COI or regulatory compliance statement. The statistical recomputations are consistent (10/10), and the scientific premise is strong. The most consequential gaps are the missing data availability statement and the lack of key resource identification (Purrble source, statistical software). A correction or erratum may be warranted to add these missing details.
- 1.HIGHdata codeAdd a data availability statement to the Methods section, specifying whether and how data can be accessed (e.g., repository, managed access, or contact author).The paper reports no data availability statement, which is a major omission for a clinical trial and undermines reproducibility.
- 2.HIGHreportingIdentify the statistical software (e.g., R version, SPSS version, Stata) used for all analyses in the Methods/Statistical analysis section.Without software identification, readers cannot evaluate the computational environment or reproduce the analyses.
- 3.HIGHotherProvide the manufacturer, model, and source (vendor) of the Purrble device in the Methods/Intervention section.The investigational product is not fully identified, which is a key resource reporting gap.
- 4.HIGHotherAdd a formal a priori power analysis (effect size, alpha, power, target sample size) to the Methods/Statistical analysis section.The paper states the target sample size was achieved but does not report the power calculation, which is essential for evaluating the study's design.
- 5.HIGHotherAdd a Conflicts of Interest statement (e.g., 'The authors declare no competing interests') in a dedicated section.A COI statement is required by most journals and is currently missing.
- 6.HIGHethicsAdd an explicit statement of compliance with the Declaration of Helsinki or ICH-GCP in the Ethical Approval section.The paper only mentions UK regulations for consent; a clear statement of adherence to an international ethical framework is expected.
- 7.HIGHreportingAdd biological sex (male/female) to Table 1 alongside gender identity, or justify why it is not reported.Standard reporting expectations include biological sex; the paper currently reports only gender identity, which is a minor gap.
- 8.HIGHstatisticsReport statistical assumptions testing (normality, homogeneity of variance) for the ANCOVA models, or note that robust methods were used.Assumptions verification is not mentioned, which is a standard reporting element for parametric tests.
- 9.HIGHdata codeConsider adding a code-sharing statement and, if possible, depositing analysis code in a public repository (e.g., GitHub, Zenodo) with a permanent identifier.Sharing code enhances reproducibility; currently, no code availability is mentioned.
- 10.MEDIUMcopyeditStandardize the notation for negative signs (en dash vs hyphen) in the Abstract and Results section.The copyedit report flagged inconsistent use of negative signs, which is a minor formatting issue.
- 11.MEDIUMcopyeditClarify the table footnote in Table 2 to indicate that the reliable change percentages are based on the same n as the overall condition.The copyedit report noted potential ambiguity in the table footnote.
- 12.MEDIUMcopyeditRemove the extra space before the comma in the journal volume number in reference 19 (and check other references for consistency).A formatting inconsistency was noted in the references.
- 13.MEDIUMcopyeditAdd a period after 't' in Table 3 footnote for clarity (e.g., 't statistic' or 't-value').Minor clarity issue in the table footnote.
- 14.LOWreportingConsider reporting the exact sample size calculation from the protocol in the manuscript to strengthen transparency.This is a nice-to-have addition beyond the current reporting.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.