Extended-release ketamine tablets for treatment-resistant depression: a randomized placebo-controlled phase 2 trial.
Glue P, Loo C, Fam J, Lane HY, Young AH, Surman P, BEDROC study investigators
- DOI
- 10.1038/s41591-024-03063-x
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/33f2947f-5ac4-4737-b85a-6f91a4cc1847 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped)−0.25★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 10 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy endpoint is the change in MADRS score, a clinician-rated depression severity scale. While MADRS is a validated clinical scale, it is a surrogate for the clinical outcome of depression remission/response. The paper does not provide evidence linking MADRS changes to hard clinical outcomes, nor does it demonstrate target engagement at the tested dose beyond the clinical response itself. The efficacy claim rests on this surrogate measure without a validated link to long-term functional outcomes or mortality.
“The primary endpoint was least square mean change in MADRS for each active treatment compared with placebo at 13 weeks”
- 02Treatment effect not shown to be clinically meaningful
The primary effect size is a mean difference of -6.1 points on the MADRS scale for the 180 mg dose vs placebo. While the paper claims this exceeds the minimal clinically important difference (MCID), the MCID for MADRS is typically around 2-3 points, so 6.1 is above that. However, the effect is presented as a group mean difference, and the clinical meaningfulness is not robustly anchored to individual-level response or remission rates. The paper does not report the proportion of patients achieving response or remission at the primary endpoint, and the effect size is modest relative to the baseline severity (mean MADRS ~30).
“the least square mean difference of MADRS score for the 180 mg tablet group and placebo was −6.1 (95% confidence interval 1.0 to 11.16, P = 0.019) at 13 weeks”
- 03Printed percentage does not match its own count
11.6% does not match the reported count 26/231
“26 participants (11.6%)”
Safety outcomes, open-label enrichment…Find in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted phase 2 randomized placebo-controlled trial with rigorous design, clear reporting, and appropriate statistical methods. The main weaknesses are a vague data availability statement and the lack of named statistical software, both minor reporting gaps. The copyedit issues are minor typos and consistency problems.
Both reviewers agreed on all dimensions and study type (interventional). The statistics verification component checked only 4 tests (3 consistent, 1 inconsistent) and cannot verify threshold-only p-values or resampling-based tests; the inconsistent test is not specified and does not constitute a demonstrable error. The integrity concern about the wide CI is low severity and not a validity threat.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks. 1 printed percentage that does not match its own count.
- PERCENT11.6% does not match the reported count 26/231
“26 participants (11.6%)”
Safety outcomes, open-label enrichment…Find in source
- CONSISTENTreported p = .019 · recomputed p = .019Reviewer 1Primary outcome p-value for 180 mg vs placebo from ANCOVA
“the least square mean difference of MADRS score for the 180 mg tablet group and placebo was −6.1 (95% confidence interval 1.0 to 11.16, P = 0.019)”
Taken as given: The estimate is -6.1 and the 95% CI is (1.0, 11.16).; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed two-tailed p from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(-6.1, 1.0, 11.16, 0) - CONSISTENTreported p = .046 · recomputed p = .046Reviewers 1, 2P-value for response rate comparison (120 mg vs placebo) using Fisher's exact test
“significant for only the 120 mg dose group for treatment response (48% versus 24.3%, P = 0.046; Extended Data Tables and )”
Taken as given: The 120 mg group has n=31 and 48% response, giving 15 responders and 16 non-responders.; The placebo group has n=37 and 24.3% response, giving 9 responders and 28 non-responders.; The test is Fisher's exact test (two-tailed).Method: Recomputed two-tailed Fisher's exact test from the 2x2 table.How we recomputed it: pFisher2x2(15, 16, 9, 28, 0) - CONSISTENTreported p = .019 · recomputed p = .019Reviewer 2Primary endpoint: 180 mg vs placebo difference in MADRS change
“the least square mean difference of MADRS score for the 180 mg tablet group and placebo was −6.1 (95% confidence interval 1.0 to 11.16, P = 0.019)”
Taken as given: The estimate is -6.1 and the 95% CI is 1.0 to 11.16.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed two-tailed p from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-6.1, 1.0, 11.16, 0)
- lowinternal contradictionThe abstract reports a 95% CI of '1.0 to 11.16' for the primary difference, but the lower bound is positive while the p-value is 0.019, which is consistent. However, the CI is unusually wide and asymmetric around the estimate (-6.1), which may warrant explanation.
“the least square mean difference of MADRS score for the 180 mg tablet group and placebo was −6.1 (95% confidence interval 1.0 to 11.16, P = 0.019)”
AbstractFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1The extended-release oral dosage ketamine formulation may be advantageous compared with intranasal or intravenous dosing, in terms of reduced intensity of dissociation, lower risk of abuse, reduced frequency and intensity of sedative and cardiovascular side effects, and improved convenience for administration in the community.The study shows low dissociation and cardiovascular effects, but the claim about lower abuse risk is speculative and not directly measured.Evidence: Low dissociation (CADSS <1), no blood pressure changes, and home dosing; but no direct abuse liability assessment.
“Use of an extended-release oral dosage ketamine formulation may be advantageous compared with intranasal or intravenous dosing, in terms of reduced intensity of dissociation, lower risk of abuse, reduced frequency and intensity of sedative and cardiovascular side effects, and improved convenience for administration in the community.”
DiscussionFind in source - partialReviewer 2The extended-release formulation overcomes many limitations of IV or intranasal ketamine.The study shows reduced side effects and home dosing, but direct comparison with other routes is not made.Evidence: Discussion notes reduced dissociation and blood pressure changes, but acknowledges no direct comparison.
“Use of an extended-release oral dosage ketamine formulation may be advantageous compared with intranasal or intravenous dosing, in terms of reduced intensity of dissociation, lower risk of abuse, reduced frequency and intensity of sedative and cardiovascular side effects, and improved convenience for administration in the community.”
DiscussionFind in source - supportedReviewer 1R-107 tablets were effective, safe and well tolerated in a patient population with TRD, enriched for initial response to R-107 tablets.The primary endpoint was met with a statistically significant difference for the 180 mg dose, and safety data show minimal side effects.Evidence: Primary outcome: LS mean difference -6.1 (95% CI 1.00 to 11.16, P=0.019); safety outcomes show no blood pressure changes, minimal sedation and dissociation.
“R-107 tablets were effective, safe and well tolerated in a patient population with TRD, enriched for initial response to R-107 tablets.”
AbstractFind in source - supportedReviewer 1The 180 mg dose given twice weekly showed statistically significant and clinically meaningful improvement in depressive symptoms based on MADRS score compared with placebo.The primary analysis shows a statistically significant difference of 6.1 points, which exceeds the MCID threshold mentioned.Evidence: Primary outcome: LS mean difference -6.1 (95% CI 1.00 to 11.16, P=0.019).
“the 180 mg dose given twice weekly showed statistically significant and clinically meaningful improvement in depressive symptoms based on MADRS score compared with placebo, with a group-treatment difference of 6.1.”
DiscussionFind in source - supportedReviewer 1Relapse rates during double-blind treatment showed a dose response from 70.6% for placebo to 42.9% for 180 mg.The relapse rates are reported and show a dose-response trend, though not all pairwise comparisons were statistically significant.Evidence: Relapse rates: placebo 70.6%, 180 mg 42.9% (from abstract).
“Relapse rates during double-blind treatment showed a dose response from 70.6% for placebo to 42.9% for 180 mg.”
AbstractFind in source - supportedReviewer 1Tolerability was excellent, with no changes in blood pressure, minimal reports of sedation and minimal dissociation.Safety data support this claim with mean blood pressure changes near zero and low CADSS scores.Evidence: Mean blood pressure changes -1.2/-0.1 mmHg; mean CADSS scores <1; sedation reported by only 5 participants.
“Tolerability was excellent, with no changes in blood pressure, minimal reports of sedation and minimal dissociation.”
AbstractFind in source - supportedReviewer 2R-107 tablets were effective in reducing depressive symptoms in TRD patients.The primary endpoint was met with a statistically significant difference for the 180 mg dose.Evidence: Primary outcome: difference of -6.1 (95% CI 1.0 to 11.16, P=0.019).
“R-107 tablets were effective, safe and well tolerated in a patient population with TRD, enriched for initial response to R-107 tablets.”
AbstractFind in source - supportedReviewer 2R-107 tablets were safe and well tolerated.Safety data show minimal side effects and no significant changes in blood pressure.Evidence: Safety outcomes: no changes in blood pressure, minimal sedation and dissociation.
“Tolerability was excellent, with no changes in blood pressure, minimal reports of sedation and minimal dissociation.”
AbstractFind in source - supportedReviewer 2The enrichment design reduced study failure rates.The design rationale is supported by cited literature and the study met its primary endpoint.Evidence: Introduction cites failure rates and the study's success.
“Failure rates can be reduced by using an enrichment design, in which nonresponders to acute treatment are excluded, followed by a subsequent relapse-prevention phase in treatment responders”
IntroductionFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy endpoint is the change in MADRS score, a clinician-rated depression severity scale. While MADRS is a validated clinical scale, it is a surrogate for the clinical outcome of depression remission/response. The paper does not provide evidence linking MADRS changes to hard clinical outcomes, nor does it demonstrate target engagement at the tested dose beyond the clinical response itself. The efficacy claim rests on this surrogate measure without a validated link to long-term functional outcomes or mortality.
“The primary endpoint was least square mean change in MADRS for each active treatment compared with placebo at 13 weeks”
- INADEQUATEEffect sizeThe primary effect size is a mean difference of -6.1 points on the MADRS scale for the 180 mg dose vs placebo. While the paper claims this exceeds the minimal clinically important difference (MCID), the MCID for MADRS is typically around 2-3 points, so 6.1 is above that. However, the effect is presented as a group mean difference, and the clinical meaningfulness is not robustly anchored to individual-level response or remission rates. The paper does not report the proportion of patients achieving response or remission at the primary endpoint, and the effect size is modest relative to the baseline severity (mean MADRS ~30).
“the least square mean difference of MADRS score for the 180 mg tablet group and placebo was −6.1 (95% confidence interval 1.0 to 11.16, P = 0.019) at 13 weeks”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior research on ketamine's antidepressant effects, oral dosing, and the prodrug hypothesis, and explains the rationale for an extended-release formulation. It also addresses limitations of prior work, such as high failure rates in acute trials and the need for enrichment designs. The hypothesis follows logically from the cited evidence.
“We chose this design owing to observations that acute antidepressant clinical trials in non-TRD depression have high failure rates (inability to separate clinical response between active and placebo arms), as high as 50% (refs. , ).”
“We hypothesized that an extended-release tablet formulation of ketamine could be an effective and well-tolerated treatment option for patients with TRD.”
“We chose this design owing to observations that acute antidepressant clinical trials in non-TRD depression have high failure rates (inability to separate clinical response between active and placebo arms), as high as 50% (refs. , ).”
“We hypothesized that an extended-release tablet formulation of ketamine could be an effective and well-tolerated treatment option for patients with TRD.”
Randomization was by an automated integrated web response system, and blinding was described for patients and all personnel. A power analysis was provided with effect size, alpha, and power. Inclusion/exclusion criteria were pre-specified. The enrichment design and double-blind phase are well described. Outlier handling is addressed via the pre-specified analysis population and imputation method.
“Randomization was by an automated integrated web response system.”
“All patients, and all people involved in the conduct of the clinical trial were blinded to treatment allocation.”
“The sample size calculation was based on the superiority of R-107 to placebo by a magnitude of six MADRS units, using an s.d. of change in MADRS of 7.5 units, a two-sided type 1 error of 0.05 and a power of 80%.”
“Randomization was by an automated integrated web response system.”
“All patients, and all people involved in the conduct of the clinical trial were blinded to treatment allocation.”
“The sample size calculation was based on the superiority of R-107 to placebo by a magnitude of six MADRS units, using an s.d. of change in MADRS of 7.5 units, a two-sided type 1 error of 0.05 and a power of 80%.”
The paper reports sex (M/F) and age for each treatment group in Table 1. Health status is implied by inclusion/exclusion criteria and baseline MADRS scores. Demographics are reported in Table 1. Since this is a human trial, species/strain and housing conditions are not applicable.
“Sex (M/F) | 22/15 | 18/16 | 18/16 | 13/18 | 21/11”
“Age, years | Mean (s.d.) | 43.7 (15.43) | 44.6 (12.89) | 42.5 (15.80) | 47.2 (13.80) | 46.8 (11.90)”
“Sex (M/F) | 22/15 | 18/16 | 18/16 | 13/18 | 21/11”
“We screened adult self-reported male and female patients (18–80 years)”
“Age, years | Mean (s.d.) | 43.7 (15.43) | 44.6 (12.89) | 42.5 (15.80) | 47.2 (13.80) | 46.8 (11.90)”
The paper states that the trial was conducted in accordance with the Declaration of Helsinki and Good Clinical Practice, and that the protocol and consent forms were approved by local or national ethics committees. Written informed consent was obtained from patients. The registration number is provided.
“The protocol, consent forms and associated documents were approved by local or national ethics committees.”
“Patients who provided written informed consent were eligible to enter screening.”
“The trial was conducted in accordance with the ethical principles stated in the Declaration of Helsinki and Good Clinical Practice quality standards”
“The protocol, consent forms and associated documents were approved by local or national ethics committees.”
“Patients who provided written informed consent were eligible to enter screening.”
“The trial was conducted in accordance with the ethical principles stated in the Declaration of Helsinki and Good Clinical Practice quality standards”
The investigational product is named as R-107 extended-release ketamine tablets, with doses and regimen described. The manufacturer (Douglas Pharmaceuticals) is mentioned in the acknowledgements. Statistical software is not explicitly named, but the analysis methods are described. Since this is a drug trial, bench resources are not applicable.
“receive double-blind R-107 doses of 30, 60, 120 or 180 mg, or placebo, twice weekly for a further 12 weeks.”
“The study was sponsored by Douglas Pharmaceuticals.”
“extended-release ketamine tablets (R-107)”
“The study was sponsored by Douglas Pharmaceuticals.”
The primary analysis used ANCOVA with baseline MADRS as covariate, and the paper reports exact p-values and 95% CIs for the primary and secondary outcomes. The sample size calculation is described. The paper reports effect sizes with confidence intervals. Statistical software is not explicitly named, but the methods are standard. Data presentation includes tables and Kaplan-Meier curves.
“This was evaluated with analysis of covariance, with dose as a factor and baseline MADRS as a covariate.”
“the least square mean difference of MADRS score for the 180 mg tablet group and placebo was −6.1 (95% confidence interval 1.0 to 11.16, P = 0.019)”
“This was evaluated with analysis of covariance, with dose as a factor and baseline MADRS as a covariate.”
“the least square mean difference of MADRS score for the 180 mg tablet group and placebo was −6.1 (95% confidence interval 1.0 to 11.16, P = 0.019)”
The data availability statement says deidentified individual participant data will be made available 24 months after publication, but only directs proposals to a contact person without specifying a platform or review process. This is a vague statement, not a concrete managed-access route. No code sharing is mentioned.
“Deidentified individual participant data and the data dictionary will be made available 24 months after publication. Proposals with specific aims and an analysis plan should be directed to P.S.”
“Deidentified individual participant data and the data dictionary will be made available 24 months after publication. Proposals with specific aims and an analysis plan should be directed to P.S.”
The trial is registered (ACTRN12618001042235). The paper states it is reported in accordance with CONSORT 2010. All outcomes are reported, including negative results. Limitations are discussed in the Discussion. Conclusions are proportional to the evidence. Funding and competing interests are disclosed.
“ClinicalTrials.gov registration: ACTRN12618001042235”
“and is reported in accordance with the CONSORT 2010 statement.”
“There are several important limitations to the trial.”
“ClinicalTrials.gov registration: ACTRN12618001042235”
“and is reported in accordance with the CONSORT 2010 statement.”
“There are several important limitations to the trial.”
Registered (1 ID: ANZCTR). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 28 references by DOI: 23 verified — 5 no DOI (shown, not verified).
- NO DOI( R , S )-Ketamine metabolites ( R , S )-norketamine and ( 2S , 6S )-hydroxynorketamine increase the mammalian target of rapamycin functionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExtended release pharmaceutical formulation and methods of treatmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExpectation, the placebo effect and the response to treatmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe clinical global impressions scale: applying a research tool in clinical practiceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIColumbia-Suicide Severity Rating Scale (C-SSRS)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link found; not probed for liveness in this run.
- datahttps://www.anzctr.org.au/Trial/Registration/TrialReview.aspx?id=375359&isReview=trueUNVERIFIEDHTTP 403Liveness indeterminate — content not checked.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAbstract“pronged absorption phase”→ prolonged absorption phaseTypo in the introduction.
- MINORconsistencyResults, Primary outcome“least square mean difference of MADRS score for the 180 mg tablet group and placebo was −6.1 (95% confidence interval 1.0 to 11.16, P = 0.019)”→ Ensure the CI is reported as (1.00 to 11.16) for consistency with other CIs.CI reported as 1.0 to 11.16 in abstract but 1.00 to 11.16 in results.
- MINORclarityResults, Secondary efficacy outcomes“A total of 132 participants (57.1%) of the 231 enrolled in the enrichment phase achieved remission with as MADRS total score ≤10 at day 8.”→ Remove 'with as' and rephrase to 'achieved remission, defined as MADRS total score ≤10 at day 8.'Grammatical error.
- MINORtypoIntroduction, paragraph 2“Due to its pronged absorption phase”→ Due to its prolonged absorption phaseTypo: 'pronged' should be 'prolonged'.
- MINORconsistencyAbstract“least square mean change”→ least squares mean changeInconsistent use of 'least square' vs 'least squares'.
- MINORclarityResults, Primary outcome“The 120 mg and 180 mg dose groups had lower mean reductions (<10 points) compared with lower-dose groups.”→ The 120 mg and 180 mg dose groups had smaller mean reductions (<10 points) compared with lower-dose groups.Clarify 'lower mean reductions' to avoid ambiguity.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (vague data access, unnamed software) and the single inconsistent recomputed statistic as caveats, but none warrant an erratum or independent re-analysis. The copyedit issues are trivial and do not affect scientific integrity.
- 1.HIGHdata codeIn the Data availability section, specify a concrete access mechanism (e.g., a data access committee, a repository like Vivli or YODA, or a named platform) and the conditions for access, to move from 'available on request' to a fully adequate statement.The current statement is vague and does not provide a clear route for researchers to obtain the data, which is a common reviewer concern.
- 2.HIGHstatisticsIn the Methods, identify the statistical software used (e.g., SAS version, R version) to satisfy the software_identified criterion.Naming the software improves reproducibility and is a standard expectation for clinical trial reports.
- 3.MEDIUMdata codeIn the Data availability statement, clarify whether the data dictionary will be deposited in a repository or only provided on request, and provide a timeline for when the data will be accessible.Clarifying the data dictionary availability and timeline strengthens the data sharing commitment.
- 4.MEDIUMreportingIn the Methods, explicitly state that no statistical assumptions were violated or describe how violations were handled, to strengthen the assumptions_verified criterion.Explicitly addressing assumptions adds rigor to the statistical analysis reporting.
- 5.MEDIUMreportingIn the Discussion, explicitly address the potential for unblinding due to the known psychoactive effects of ketamine, even though blinding was described.Acknowledging this risk enhances transparency about a potential source of bias.
- 6.MEDIUMreportingIn the Discussion, consider adding a sentence acknowledging the lack of independent replication and the need for confirmatory trials.Explicitly noting the need for replication is important for a phase 2 trial.
- 7.MEDIUMreportingIn the Methods, consider adding a statement about the availability of the statistical analysis plan or protocol in a public repository to enhance transparency.Public availability of the SAP/protocol supports reproducibility and transparency.
- 8.LOWcopyeditFix the typo 'pronged absorption phase' to 'prolonged absorption phase' in the Abstract and Introduction.Correcting typos improves professionalism and readability.
- 9.LOWcopyeditFix the grammatical error in Results, Secondary efficacy outcomes: 'achieved remission with as MADRS total score ≤10' should be 'achieved remission, defined as MADRS total score ≤10'.Correcting grammar improves clarity.
- 10.LOWcopyeditStandardize the CI reporting: use '1.00 to 11.16' consistently in the Abstract and Results.Consistent formatting avoids confusion.
- 11.LOWcopyeditStandardize 'least square mean' to 'least squares mean' throughout the manuscript.Consistent terminology is expected in scientific writing.
- 12.LOWcopyeditClarify 'lower mean reductions' to 'smaller mean reductions' in Results, Primary outcome.Avoids ambiguity about the direction of the effect.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.