Full-spectrum extract from Cannabis sativa DKJ127 for chronic low back pain: a phase 3 randomized placebo-controlled trial.
Karst M, Meissner W, Sator S, Keßler J, Schoder V, Häuser W
- DOI
- 10.1038/s41591-025-03977-0
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/891c58e3-bb00-4521-ad64-b8ca43a03760 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×2−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 36 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on patient-reported pain intensity (NRS) and neuropathic pain symptoms (NPSI), which are subjective symptom scales, not hard clinical outcomes. Although these are standard endpoints in pain trials, they are surrogate measures for the ultimate clinical benefit of improved function and quality of life. The paper does not provide evidence linking changes in NRS or NPSI to validated long-term clinical outcomes, nor does it demonstrate target engagement at the tested dose beyond the observed symptomatic effects.
“The primary endpoint of phase A was a change in mean numeric rating scale (NRS) pain intensity, with a change in total neuropathic pain symptom inventory (NPSI) score as a key secondary endpoint”
- 02Treatment effect not shown to be clinically meaningful
The primary effect size is a mean difference of -0.6 NRS points (95% CI -0.9 to -0.3) compared to placebo. While statistically significant, this difference is below the commonly accepted minimal clinically important difference (MCID) of 1.0 point for chronic pain. The paper does not anchor this effect to a clinically meaningful threshold, and although responder analyses (≥30% reduction) are provided, the primary endpoint's effect size is small.
“mean difference (MD) versus placebo = −0.6, 95% confidence interval (CI) = −0.9 to −0.3; P < 0.001”
- 03Printed percentage does not match its own count
32.2% is unattainable for n=390 (nearest: 32.1, 32.3%)
“≥50% pain responder at week 15 (%) | 32.2”
Table 2Find in source - 04Printed percentage does not match its own count
23.4% is unattainable for n=425 (nearest: 23.3, 23.5%)
“PGIC—participants with improvement of symptoms at visit A6 (%) | 23.4”
Table 2Find in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported phase 3 randomized controlled trial of a cannabis extract for chronic low back pain. The methods are rigorous, with clear randomization, blinding, sample size justification, and comprehensive reporting of demographics, ethics, and data availability. Minor reporting inconsistencies and copyedit issues exist but do not undermine the overall integrity.
Both reviewers classified the study as interventional and agreed on all dimensions. The statistics verification component recomputed only a subset of tests (5 total, 3 consistent, 2 inconsistent with no decision errors); this does not constitute a full validation of all statistical results. The citation check found no retracted or non-existent references. The integrity check flagged two low-severity internal contradictions (enrollment vs. analysis population, and a p-value discrepancy) that are addressed in the action items.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks. 2 printed percentages that do not match their own count.
- PERCENT32.2% is unattainable for n=390 (nearest: 32.1, 32.3%)
“≥50% pain responder at week 15 (%) | 32.2”
Table 2Find in source - PERCENT23.4% is unattainable for n=425 (nearest: 23.3, 23.5%)
“PGIC—participants with improvement of symptoms at visit A6 (%) | 23.4”
Table 2Find in source
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Primary endpoint MD and CI
“VER-01 demonstrated a greater pain reduction compared to placebo with a mean difference (MD) of −0.6 (95% confidence interval (CI) = −0.9 to −0.3; P < 0.001).”
Taken as given: The MD is the difference in mean change from baseline between groups.; The CI is a 95% confidence interval for the MD.; The p-value is two-sided.Method: Recomputed p-value from the reported MD and 95% CI using the normal approximation.How we recomputed it: pCI(-0.6, -0.9, -0.3, 0) - CONSISTENTreported p < .017 · recomputed p = .016Reviewers 1, 2Key secondary endpoint MD and CI
“The mean NPSI total score decreased by −14.4 (s.e. = 3.3) points from baseline in the VER-01 arm compared to −7.2 (s.e. = 2.8) in the placebo arm, with an MD of −7.3 (95% CI = −13.2 to −1.3; P = 0.017).”
Taken as given: The MD is the difference in mean change from baseline between groups.; The CI is a 95% confidence interval for the MD.; The p-value is two-sided.Method: Recomputed p-value from the reported MD and 95% CI using the normal approximation.How we recomputed it: pCI(-7.3, -13.2, -1.3, 0) - CONSISTENTreported p > .288 · recomputed p = .287Reviewers 1, 2Phase D HR and CI
“Time to treatment failure did not differ significantly between VER-01 and placebo in phase D (hazard ratio = 0.75, 95% CI = 0.44–1.27; P = 0.288)”
Taken as given: The HR is the hazard ratio for VER-01 vs placebo.; The CI is a 95% confidence interval for the HR.; The p-value is two-sided.Method: Recomputed p-value from the reported HR and 95% CI using the normal approximation on the log scale.How we recomputed it: pCI(0.75, 0.44, 1.27, 1)
- lowinternal contradictionThe abstract states 820 participants were enrolled, but the efficacy analysis included 815 participants. This is likely due to 5 participants not receiving any dose, but the discrepancy is not explicitly explained in the text.
It enrolled 820 adults with CLBP (VER-01, n = 394; placebo, n = 426) ... A total of 815 participants were included in the efficacy analysis of phase A
Abstractreviewer’s wording - lowinternal contradictionThe text reports a p-value of 0.019 for the ≥30% RMDQ responder, while Table 2 reports P < 0.001 for the same endpoint.
“Post hoc analysis revealed that 51.7% of participants in the VER-01 arm achieved an improvement of the RMDQ score by at least 30% from baseline compared to 42.2% in the placebo arm ( P = 0.019).”
Table 2Find in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2VER-01 is safe and well-tolerated.While the paper reports no serious safety signals, the higher rate of AEs and discontinuations due to AEs in the VER-01 arm warrants caution.Evidence: TEAEs 83.3% vs 67.3%; discontinuation due to AEs 17.3% vs 3.5%.
“VER-01 was well-tolerated, with no signs of dependence or withdrawal.”
AbstractFind in source - supportedReviewers 1, 2VER-01 reduces chronic low back pain compared to placebo.The primary endpoint was met with a statistically significant difference in NRS pain reduction.Evidence: Primary endpoint: MD = -0.6, 95% CI -0.9 to -0.3, P < 0.001.
“The study met its primary endpoint in phase A, with a mean pain reduction of −1.9 NRS points in the VER-01 group (mean difference (MD) versus placebo = −0.6, 95% confidence interval (CI) = −0.9 to −0.3; P < 0.001).”
AbstractFind in source - supportedReviewers 1, 2VER-01 improves physical function and sleep quality.Secondary endpoints showed significant improvements in sleep quality and physical function.Evidence: Sleep quality MD = -0.7 (95% CI -1.0 to -0.3; P < 0.001); RMDQ MD = -1.1 (95% CI -1.8 to -1.1; P < 0.001).
“A phase 3 trial found that VER-01, a full-spectrum cannabis extract, reduces chronic low back pain, improves physical function and sleep and shows no signs of dependence or withdrawal.”
AbstractFind in source - supportedReviewers 1, 2VER-01 shows no signs of dependence or withdrawal.The paper reports no AEs indicative of abuse, dependence, or withdrawal, and no withdrawal symptoms on the CWS.Evidence: No AEs indicative of drug abuse, dependence or withdrawal were reported; no withdrawal symptoms as measured with the CWS.
“No AEs indicative of drug abuse, dependence or withdrawal, as classified under the respective standardized MedDRA queries, were reported (Supplementary Tables and ).”
ResultsFind in source - supportedReviewers 1, 2VER-01 shows potential as a new, safe and effective treatment for CLBP.The evidence supports efficacy and safety, though long-term comparative data are lacking.Evidence: Primary and secondary endpoints met; safety profile acceptable.
“VER-01 shows potential as a new, safe and effective treatment for CLBP.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on patient-reported pain intensity (NRS) and neuropathic pain symptoms (NPSI), which are subjective symptom scales, not hard clinical outcomes. Although these are standard endpoints in pain trials, they are surrogate measures for the ultimate clinical benefit of improved function and quality of life. The paper does not provide evidence linking changes in NRS or NPSI to validated long-term clinical outcomes, nor does it demonstrate target engagement at the tested dose beyond the observed symptomatic effects.
“The primary endpoint of phase A was a change in mean numeric rating scale (NRS) pain intensity, with a change in total neuropathic pain symptom inventory (NPSI) score as a key secondary endpoint”
- INADEQUATEEffect sizeThe primary effect size is a mean difference of -0.6 NRS points (95% CI -0.9 to -0.3) compared to placebo. While statistically significant, this difference is below the commonly accepted minimal clinically important difference (MCID) of 1.0 point for chronic pain. The paper does not anchor this effect to a clinically meaningful threshold, and although responder analyses (≥30% reduction) are provided, the primary endpoint's effect size is small.
“mean difference (MD) versus placebo = −0.6, 95% confidence interval (CI) = −0.9 to −0.3; P < 0.001”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior work on the prevalence and burden of CLBP, the risks of NSAIDs and opioids, and the limitations of existing cannabis-based medicine studies (small samples, short durations, inconsistent dosing). It explicitly states that the study addresses a critical gap by providing a large-scale, placebo-controlled phase 3 trial with a chemically well-defined extract. The rationale linking the premise to the study objectives is logical and well-supported.
“By providing a large-scale, placebo-controlled phase 3 trial of adequate duration using a chemically well-defined, full-spectrum cannabis extract in CLBP, this study addresses a critical gap in the clinical research of cannabis-based pharmacotherapy in chronic pain.”
“The limitations of existing treatments and the stagnation in the development of new analgesics have fueled growing public and scientific interest in the use of cannabis-based medicines for the management of chronic pain.”
“By providing a large-scale, placebo-controlled phase 3 trial of adequate duration using a chemically well-defined, full-spectrum cannabis extract in CLBP, this study addresses a critical gap in the clinical research of cannabis-based pharmacotherapy in chronic pain.”
Randomization method (computer-generated, block size 4, stratified by neuropathic pain component) and unit (participant) are reported. Blinding of participants, investigators, and site personnel is described. A priori power analyses are provided for the primary and key secondary endpoints. Inclusion/exclusion criteria are detailed, and the analysis population (full analysis set) is defined. Outlier handling is addressed through imputation strategies for intercurrent events. Controls (placebo) are appropriate. Independent replication is not applicable for a single pivotal trial.
“The sample size calculation was based on an assumed treatment difference of 0.6 NRS points, an s.d. of 2.5 points, a two-sided significance level of 5% and 90% statistical power.”
“This study enrolled individuals who were at least 18 years of age and diagnosed with CLBP (low back pain for at least 3 months), with or without a neuropathic pain component.”
“The sample size calculation was based on an assumed treatment difference of 0.6 NRS points, an s.d. of 2.5 points, a two-sided significance level of 5% and 90% statistical power.”
Sex is reported for both groups (57.4% female in VER-01, 55.8% in placebo). Age, BMI, and health status (comorbidities like hypertension, diabetes, obesity) are reported. Demographics include race. Species/strain and housing conditions are not applicable for a human trial.
“The most frequent concurrent diseases were hypertension (35.3%) and obesity (32.0%).”
The methods state that the trial was approved by ethics committees in each country (with a supplementary table listing them) and that written informed consent was obtained from all participants. It also states compliance with the Declaration of Helsinki and ICH-GCP guidelines. This meets the criteria for adequate reporting.
“The trial was approved by ethics committees in each country (Supplementary Table ), and written informed consent was obtained from all participants.”
“The trial was conducted at 66 outpatient sites and university-based hospitals in Germany and Austria in accordance with the principles of the Declaration of Helsinki, the Good Clinical Practice guidelines of the International Council for Harmonization and applicable regulatory requirements.”
The investigational product is described in detail (composition, THC content, manufacturing). The placebo is described. Statistical software (SAS 9.4) is identified. No antibodies, cell lines, or mycoplasma testing are applicable. The product is adequately identified for a clinical trial.
“Each dose unit (119 µl) of the finished product VER-01 contains 50 µl of the full-spectrum extract, delivering 2.5 mg THC, 0.1 mg cannabigerol and 0.02 mg cannabidiol, with sesame oil as excipient.”
“All statistical analyses were conducted by an independent clinical research organization with SAS software, version 9.4 (SAS Institute).”
“Each dose unit (119 µl) of the finished product VER-01 contains 50 µl of the full-spectrum extract, delivering 2.5 mg THC, 0.1 mg cannabigerol and 0.02 mg cannabidiol, with sesame oil as excipient.”
Statistical tests are named (ANCOVA, chi-squared, t-test, Wilcoxon, Cox proportional hazards). Assumptions are handled through pre-specified models and imputation strategies. Exact p-values are reported for primary and secondary endpoints. Effect sizes with confidence intervals are provided. Software (SAS 9.4) is identified. Data presentation includes per-group n, means, SDs, and CIs. Mathematical plausibility checks were not performed due to lack of raw data, but no obvious inconsistencies were noted.
“The primary endpoint of phase A was tested by an analysis of covariance model, with treatment as the main effect and with baseline characteristics (presence of a neuropathic pain component; NRS morning pain intensity, age, sex and country) as covariates.”
“VER-01 demonstrated a greater pain reduction compared to placebo with a mean difference (MD) of −0.6 (95% confidence interval (CI) = −0.9 to −0.3; P < 0.001).”
“VER-01 demonstrated a greater pain reduction compared to placebo with a mean difference (MD) of −0.6 (95% confidence interval (CI) = −0.9 to −0.3; P < 0.001).”
“P value for the two-sided chi-squared test testing the null hypothesis that responder status and treatment group are independent.”
The data availability statement describes a managed access process: requests via a website, review of proposal, and data sharing through a secure platform within 3 months. This is adequate for patient-level data. Code availability mentions the software used (ClinCase, SAS) but does not provide a public repository; however, for a clinical trial, custom code is not typically shared, and the statement is adequate.
“Access to anonymized individual data and blank case report forms that underlie the results reported in this article can be requested by qualified researchers for academic purposes. Vertanical provides access within 3 months following review and approval of a research proposal, statistical analysis plan and execution of a data access agreement.”
“Analyses were performed with SAS software, version 9.4 (SAS Institute), in keeping with the statistical analysis plan.”
“Access to anonymized individual data and blank case report forms that underlie the results reported in this article can be requested by qualified researchers for academic purposes. Vertanical provides access within 3 months following review and approval of a research proposal, statistical analysis plan and execution of a data access agreement.”
“Analyses were performed with SAS software, version 9.4 (SAS Institute), in keeping with the statistical analysis plan.”
The trial is registered (NCT04940741, EudraCT). Methods are detailed enough for replication. A CONSORT flow diagram is provided. All pre-specified outcomes are reported, including negative results (phase D). Limitations are explicitly discussed. Conclusions are proportional to the evidence. Funding and competing interests are disclosed.
“ClinicalTrials.gov registration: NCT04940741 (https://clinicaltrials.gov/study/NCT04940741)”
“ClinicalTrials.gov registration: NCT04940741 (https://clinicaltrials.gov/study/NCT04940741)”
Registered (2 IDs: ClinicalTrials.gov, EudraCT). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 53 references by DOI: 46 verified — 7 no DOI (shown, not verified).
- NO DOIWHO guideline for non-surgical management of chronic primary low back pain in adults in primary and community care settingsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBotanical drug development: guidance for industryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGuideline on the clinical development of medicinal products intended for the treatment of painNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMulticentre, randomized, open-label study to prove an additional benefit of the full-spectrum cannabis extract ver-01 over opioids in the treatment of patients with chronic non-specific low back painNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDraft guidance for industry on analgesic indications: developing drug and biological products; availabilityNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe MOS 36-item short-form health survey (SF-36). I. Conceptual framework and item selectionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHow to Score Version 2 of the SF-36® Health SurveyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://vertanical.com/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, grammar.
- MINORconsistencyAbstract“Pain further decreased to −2.9 NRS points in phase B, with effects sustained through phase C.”→ Consider specifying the time point for the −2.9 reduction (e.g., 'at the end of phase B') for clarity.The abstract mentions phase B and C but does not specify the exact time point for the −2.9 reduction.
- MINORgrammarDiscussion“Additionally, participants in the VER-01 arm required substantially lower rescue medication use less rescue medication.”→ Revise to 'Additionally, participants in the VER-01 arm required substantially less rescue medication.'Redundant phrasing.
- MINORconsistencyResults, Efficacy“The difference between VER-01 and placebo was significantly in favor of VER-01 in every single study week of the 12-week treatment phase”→ Consider rephrasing to 'The difference between VER-01 and placebo was significantly in favor of VER-01 at every study week of the 12-week treatment phase' for clarity.Awkward phrasing.
- MINORconsistencyResults, Efficacy“Post hoc analysis revealed that 51.7% of participants in the VER-01 arm achieved an improvement of the RMDQ score by at least 30% from baseline compared to 42.2% in the placebo arm ( P = 0.019).”→ Ensure the p-value is consistent with Table 2, which reports P < 0.001 for the ≥30% RMDQ responder.Potential inconsistency between text and table.
- MINORgrammarDiscussion, paragraph 1“Additionally, participants in the VER-01 arm required substantially lower rescue medication use less rescue medication.”→ Revise to: 'Additionally, participants in the VER-01 arm required substantially less rescue medication.'Redundant phrasing.
- MINORconsistencyResults, Efficacy“The difference between VER-01 and placebo was significantly in favor of VER-01 in every single study week”→ Consider rephrasing to 'was significantly in favor of VER-01 at every study week' for clarity.Awkward phrasing.
The published work is robust and generally well-reported, but an informed reader should weigh the minor internal inconsistencies (enrollment count, p-value discrepancy) and the lack of public code sharing. These do not invalidate the conclusions but warrant clarification or correction via an erratum or author note.
- 1.HIGHreportingReconcile the discrepancy between the abstract's 820 enrolled participants and the 815 in the efficacy analysis by explicitly stating that 5 participants were excluded (e.g., never received a dose) in the Results section.The internal contradiction could confuse readers and reviewers about the analysis population.
- 2.HIGHstatisticsCorrect the p-value for the ≥30% RMDQ responder in the Results text (P = 0.019) to match Table 2 (P < 0.001), or clarify if these are different analyses.Inconsistent p-values for the same endpoint undermine statistical credibility.
- 3.MEDIUMdata codeDeposit the statistical analysis code (e.g., SAS macros) in a public repository (e.g., Zenodo, GitHub) with a DOI, and link it in the Code availability section.Sharing code enhances reproducibility, which is currently limited to naming the software.
- 4.MEDIUMreportingAdd an explicit statement that the trial followed CONSORT guidelines and provide the completed CONSORT checklist as supplementary material.Explicit adherence to reporting guidelines strengthens transparency.
- 5.MEDIUMdata codeClarify the data access process by specifying the exact criteria for 'qualified researchers' and the timeline for response to requests.Reduces ambiguity for potential data requesters.
- 6.MEDIUMreportingPublish the full study protocol as a supplementary file or provide a persistent link to it.Full protocol access improves transparency and reproducibility.
- 7.MEDIUMreportingAcknowledge the lack of a formal blinding assessment as a limitation and suggest future studies include a blinding questionnaire.Addresses a known limitation that could affect interpretation of results.
- 8.LOWcopyeditFix the redundant phrasing in the Discussion: 'required substantially lower rescue medication use less rescue medication' → 'required substantially less rescue medication'.Grammar error detracts from professionalism.
- 9.LOWcopyeditRephrase 'in every single study week' to 'at every study week' in the Results section for clarity.Awkward phrasing reduces readability.
- 10.LOWcopyeditSpecify the time point for the −2.9 NRS reduction in the Abstract (e.g., 'at the end of phase B').Clarifies the temporal context of the reported effect.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.