BCMA-directed mRNA CAR T cell therapy for myasthenia gravis: a randomized, double-blind, placebo-controlled phase 2b trial.
Vu T, Durmus H, Rivner M, Shroff S, Ragole T, Myers B, Pasnoor M, Small G, Karam C, Vullaganti M, Peltier A, Sahagian G, Feinberg MH, Slanksy A, Barnett-Tapia C, Siddiqi Z, Gwathmey K, Badruddoja MA, Kamboh H, Ruggerie RN, Fedak RR, Stewart CA, Kurtoglu M, Kalayoglu M, Singer M, Jewell CM, Miljkovic MD, Dimachkie M, Mozaffar T, Howard JF Jr, MG-001 Study Team
- DOI
- 10.1038/s41591-025-04171-y
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/4bee52dc-1762-4dc6-a4f1-629dd6ada414 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is the MGC score, a composite of patient-reported and provider-assessed items, which is a clinical scale, not a surrogate biomarker. However, the efficacy claim also relies on secondary endpoints like MG-ADL and QMG, which are also clinical scales. The paper does not present a surrogate biomarker as the primary basis for efficacy; instead, it uses validated clinical outcome measures. Therefore, the surrogate criterion is not applicable in the traditional sense, but since the primary endpoint is a clinical scale, it is considered adequate.
“The primary endpoint was a ≥5-point improvement in the MG Composite (MGC) score at month 3.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomized phase 2b trial of an mRNA CAR T cell therapy for myasthenia gravis. The design, ethics, statistical reporting, and data availability are all strong, with no major rigor gaps identified. Minor copyedit inconsistencies (rounding, allocation description) are the only issues, and they do not undermine the scientific integrity.
Both reviewers independently scored all dimensions as pass with high confidence, and their evidence was consistent. The study is interventional (randomized controlled trial). Non-applicable criteria (e.g., animal housing, cell line authentication) were excluded. The statistics verification covered only 4 tests with sufficient detail; other reported statistics were not machine-verified and should not be assumed correct.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 4 tests: 4 consistent, 0 inconsistent; 4 via agent-written checks.
- CONSISTENTreported p = .047 · recomputed p = .047Reviewer 1Primary endpoint: proportion of MGC responders at month 3 (Descartes-08 vs placebo) using two-sided two-independent-sample proportion test (Wald chi-squared).
“At month 3, the proportion of patients achieving a ≥5-point improvement in the MGC score was significantly higher for those treated with Descartes-08 compared to placebo in the overall population (66.7% ( n = 10/15) versus 27.3% ( n = 3/11), P = 0.0472)”
Taken as given: The numbers 10/15 and 3/11 are the responder counts and group totals for Descartes-08 and placebo, respectively.; The test is a two-sided Pearson chi-squared test on the 2x2 table (10,5,3,8).Method: Pearson chi-squared test on 2x2 table (10,5,3,8) using pChi2x2.How we recomputed it: pChi2x2(10,5,3,8) - CONSISTENTreported p = .041 · recomputed p = .051Reviewer 1Secondary endpoint: MG-ADL change at month 3 in AChR-positive subgroup (Descartes-08 vs placebo).
“There was a significant and clinically meaningful reduction in the mean (s.d.) MG-ADL score at month 3 for the Descartes-08 group compared to the placebo group (−3.4 (2.84) versus 0.6 (2.93), P = 0.0409)”
Taken as given: The t-statistic is approximately 2.1 (derived from the difference in means and SDs, but not directly reported).; Degrees of freedom = 17 (n1+n2-2 = 11+8-2 = 17).; The test is a two-sample t-test (two-sided).Method: Two-sample t-test approximation using t-statistic 2.1 and df=17.How we recomputed it: pT(2.1,17) - CONSISTENTreported p = .047 · recomputed p = .047Reviewer 2Primary endpoint: proportion of MGC responders at month 3 (Descartes-08 vs placebo) in overall population.
“At month 3, the proportion of patients achieving a ≥5-point improvement in the MGC score was significantly higher for those treated with Descartes-08 compared to placebo in the overall population (66.7% ( n = 10/15) versus 27.3% ( n = 3/11), P = 0.0472)”
Taken as given: The numbers 10/15 and 3/11 are the responder counts and group totals for Descartes-08 and placebo, respectively.; The test used is a two-sided Pearson chi-squared test (or equivalent) on the 2x2 table.; The p-value is two-tailed.Method: Recomputed using Pearson's chi-squared test on the 2x2 table (10,5,3,8).How we recomputed it: pChi2x2(10,5,3,8) - CONSISTENTreported p = .041 · recomputed p = .051Reviewer 2Secondary endpoint: MG-ADL change at month 3 in AChR-positive population.
“There was a significant and clinically meaningful reduction in the mean (s.d.) MG-ADL score at month 3 for the Descartes-08 group compared to the placebo group (−3.4 (2.84) versus 0.6 (2.93), P = 0.0409)”
Taken as given: The means and SDs are for the AChR-positive subgroup.; The test is a two-sample t-test (Welch's or Student's) comparing the two means.; The degrees of freedom are approximated as 17 (n1+n2-2 = 11+8-2 = 17).; The p-value is two-tailed.Method: Recomputed using a two-sample t-test with the given means, SDs, and sample sizes (11 and 8), assuming equal variances (df=17).How we recomputed it: pT(2.1,17)
- lowinternal contradictionThe abstract reports infusion-related reactions in 80.0% (n=16/20) for Descartes-08 and 56.3% (n=9/16) for placebo, but the safety population is 20 and 16, respectively, which is consistent. However, the efficacy population is 15 and 11, and the abstract does not clarify which population is used for safety.
“Infusion-related reactions were the most common adverse events reported (Descartes-08, 80.0% ( n = 16/20); placebo, 56.3% ( n = 9/16)).”
AbstractFind in source - lowinternal contradictionThe abstract reports 33.0% and 55.60% for MSE achievement, while the results section reports 33.3% (4/12) and 55.6% (5/9). These are minor rounding inconsistencies.
“33.0% of patients achieved minimum symptom expression (MSE) (MG-ADL score ≤1) by month 6, which was sustained through month 12. Among biologic-naive patients, 55.60% achieved MSE by month 6”
AbstractFind in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2A single course of six once-weekly infusions of Descartes-08 resulted in sustained clinically meaningful responses among patients with gMG.The claim is supported by the primary endpoint and follow-up data, but the small sample size and lack of formal statistical testing for secondary endpoints and long-term outcomes warrant caution.Evidence: Primary endpoint met; sustained improvements in MGC, MG-ADL, QMG through month 12; MSE achieved by a third of patients.
“In summary, a single course of six once-weekly infusions of Descartes-08 was well tolerated and resulted in sustained clinically meaningful responses among patients with gMG.”
DiscussionFind in source - partialReviewer 2Descartes-08 results in sustained clinically meaningful responses through month 12.The claim is supported by descriptive data showing maintained improvements, but the lack of formal statistical testing at later time points and the small sample size limit the strength of the evidence.Evidence: Mean changes from baseline in MGC, MG-ADL, QMG at month 4 and month 12 are reported; 83.0% of patients achieving a sustained response at month 12.
“with 83.0% of patients achieving a sustained and clinically meaningful response at month 12.”
AbstractFind in source - supportedReviewers 1, 2Descartes-08 significantly improves MGC score at month 3 compared to placebo.The primary endpoint was met with a statistically significant difference (P=0.0472) and the effect size with CI is provided.Evidence: Primary outcome: 66.7% vs 27.3% responders, P=0.0472, difference in proportions 0.39 (95% CI 0.01-0.77).
“At month 3, the proportion of patients achieving a ≥5-point improvement in the MGC score was significantly higher for those treated with Descartes-08 compared to placebo in the overall population (66.7% ( n = 10/15) versus 27.3% ( n = 3/11), P = 0.0472)”
AbstractFind in source - supportedReviewers 1, 2Descartes-08 is generally safe and well tolerated.Safety data show similar AE rates between groups, mostly mild/moderate, with no grade 4/5 events.Evidence: AE rates: 85% vs 81.3% any AE; no grade 4/5 events; most AEs mild/moderate.
“Adverse event (AE) rates were similar between groups, with 17 of 20 (85.0%) participants in the Descartes-08 group and 13 of 16 (81.3%) participants in the placebo group reporting at least one AE during the study”
ResultsFind in source - supportedReviewers 1, 2Descartes-08 may provide a valuable new treatment opportunity in gMG.The claim is appropriately cautious, acknowledging the small sample size and need for phase 3 confirmation.Evidence: Discussion states the results indicate a valuable new treatment opportunity may be achievable.
“The results of this trial indicate that a valuable new treatment opportunity in gMG—a brief course of treatment leading to at least a year-long benefit—may be achievable.”
DiscussionFind in source
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary endpoint is the MGC score, a composite of patient-reported and provider-assessed items, which is a clinical scale, not a surrogate biomarker. However, the efficacy claim also relies on secondary endpoints like MG-ADL and QMG, which are also clinical scales. The paper does not present a surrogate biomarker as the primary basis for efficacy; instead, it uses validated clinical outcome measures. Therefore, the surrogate criterion is not applicable in the traditional sense, but since the primary endpoint is a clinical scale, it is considered adequate.
“The primary endpoint was a ≥5-point improvement in the MG Composite (MGC) score at month 3.”
- ADEQUATEEffect sizeThe primary effect size is the proportion of MGC responders (≥5-point improvement) at month 3: 66.7% in the Descartes-08 group vs 27.3% in placebo (P=0.0472). The MGC scale has a validated minimal clinically important difference of 3 points, and the threshold of 5 points is higher, indicating a clinically meaningful improvement. Additionally, secondary outcomes show mean changes in MG-ADL and QMG that exceed their respective MCIDs (2 and 3 points). The effect is statistically supported and anchored to clinical meaningfulness.
“At month 3, the proportion of patients achieving a ≥5-point improvement in the MGC score was significantly higher for those treated with Descartes-08 compared to placebo in the overall population (66.7% (n=10/15) versus 27.3% (n=3/11), P=0.0472)”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on MG pathophysiology, BCMA as a therapeutic target, and prior clinical data from phase 1b/2a. It acknowledges limitations of existing treatments and prior CAR T approaches. The rationale linking BCMA-expressing plasma cells to MG and the mechanism of Descartes-08 is well articulated.
“In an open-label, multicenter, phase 1b/2a trial of Descartes-08 in gMG (MG-001 parts 1 and 2), patients who received six once-weekly doses without preconditioning chemotherapy in an outpatient setting experienced a robust improvement in gMG symptom severity”
“The secretion of AChR autoantibodies is primarily driven by pathogenic B cell maturation antigen (BCMA)-expressing plasmablasts and long-lived plasma cells”
“Conventional treatments for gMG include chronic broad immunosuppression with corticosteroids and nonsteroidal immunosuppressive therapy; however, these are often insufficient for complete symptom control and can result in considerable toxicity due to off-target effects”
“nonintegrating BCMA-directed CAR T cell therapies may circumvent this toxicity due to the lack of requirement for chemotherapy.”
“many patients continue to experience incomplete disease control.”
Randomization used a computer-generated permuted-block scheme without stratification, with concealed block sizes. Blinding was maintained for participants, investigators, and outcome assessors. A power analysis was performed (80% power to detect 47% difference). Inclusion/exclusion criteria are detailed. The primary analysis population (mITT) and missing data handling are pre-specified. The trial is registered (NCT04146051).
“Randomization was performed centrally using a computer-generated permuted-block scheme without stratification.”
“investigators, participants, outcome assessors and all other study personnel remained blinded to the treatment assignment throughout the trial.”
“We estimated that 15 participants per treatment group would provide 80% power to detect a difference of 47% in responders between Descartes-08 and placebo.”
“Randomization was performed centrally using a computer-generated permuted-block scheme without stratification.”
“investigators, participants, outcome assessors and all other study personnel remained blinded to the treatment assignment throughout the trial.”
“We estimated that 15 participants per treatment group would provide 80% power to detect a difference of 47% in responders between Descartes-08 and placebo.”
Table 1 provides detailed baseline demographics including sex, age, weight, ethnicity, MGFA class, antibody status, and disease duration. Both sexes are enrolled, so sex justification is not applicable. Age, weight, and health status are reported. Species/strain and housing are not applicable for a human trial.
“Sex, n (%) | Female | 10 (66.7) | 6 (54.5) | 16 (61.5)”
“Mean age, years (s.d.) | 56.7 (16.39) | 59.0 (13.96) | 57.7 (15.16)”
“Ethnicity, n (%) | White, non-Hispanic | 13 (86.7) | 11 (100.0) | 24 (92.3)”
“Sex, n (%) | Female | 10 (66.7) | 6 (54.5) | 16 (61.5)”
“Mean age, years (s.d.) | 56.7 (16.39) | 59.0 (13.96) | 57.7 (15.16)”
“Ethnicity, n (%) | White, non-Hispanic | 13 (86.7) | 11 (100.0) | 24 (92.3)”
The trial was approved by named IRBs and ethics committees (WCG IRB, University of Toronto, University of Alberta, Istanbul University). Written informed consent was obtained from all participants. Compliance with Declaration of Helsinki and ICH-GCP is stated.
“The trial was approved by regulatory authorities in each country (USA, Canada and Türkiye), the central and local institutional review boards in the USA (Western Institutional Review Board-Copernicus Group, WCG IRB, Puyallup, WA), and the ethics committee at each site in other countries”
“All participants provided written informed consent before any study-related activities.”
“The MG-001 part 3 trial was performed in accordance with the principles of the Declaration of Helsinki and the International Council for Harmonization E6 guidelines for Good Clinical Practice.”
“All participants provided written informed consent before any study-related activities.”
“The MG-001 part 3 trial was performed in accordance with the principles of the Declaration of Helsinki and the International Council for Harmonization E6 guidelines for Good Clinical Practice.”
Descartes-08 is described as an autologous BCMA-directed mRNA CAR T cell therapy, with dose specified (52.5 × 10^6 viable CAR+ cells per kg ± 45%). The placebo is described as matched in appearance. Statistical software (SAS, Mathematica) is identified. No antibodies, cell lines, or other bench reagents are used, so those criteria are not applicable.
“The dose of Descartes-08 was 52.5 × 10 6 viable CAR + cells per kg ± 45% per infusion”
“The placebo was matched to Descartes-08 in appearance and supplied in identical containers.”
“Statistical analyses were conducted using SAS (version 9.2 or higher) and Mathematica (version 11.0 or higher), where applicable.”
“Descartes-08, an autologous BCMA-directed mRNA chimeric antigen receptor T cell therapy”
“The dose of Descartes-08 was 52.5 × 10 6 viable CAR + cells per kg ± 45% per infusion”
“Statistical analyses were conducted using SAS (version 9.2 or higher) and Mathematica (version 11.0 or higher)”
The primary analysis used a two-independent-sample proportion test with Wald chi-squared test, and exact p-values are reported (e.g., P = 0.0472). Effect sizes with 95% CIs are provided (difference in proportions 0.39, 95% CI 0.01–0.77). Missing data imputation is described. Statistical software is identified. Data presentation includes per-group n and dispersion measures.
“66.7% versus 27.3%, P = 0.0472”
“Difference in proportions (95% CI) | 0.39 (0.01–0.77)”
“Difference in proportions (95% CI) | 0.39 (0.01–0.77)”
The data availability statement specifies that anonymized trial-level data will be provided upon request to qualified researchers after review of a proposal and execution of a data sharing agreement, with a response timeframe (30 business days) and contact email. This is a concrete managed-access route, adequate for patient-level data. No code is shared, but no bespoke code is mentioned.
“Access to anonymized trial-level data (analysis datasets) and/or the study protocol will be provided upon request to qualified researchers conducting independent, rigorous research, after the review and approval of a research proposal and statistical analysis plan, as well as the execution of a data sharing agreement.”
“Data requests can be submitted at any time and will be responded to within 30 business days of submission.”
“Access to anonymized trial-level data (analysis datasets) and/or the study protocol will be provided upon request to qualified researchers conducting independent, rigorous research, after the review and approval of a research proposal and statistical analysis plan, as well as the execution of a data sharing agreement. Data requests can be submitted at any time and will be responded to within 30 business days of submission.”
“ClinicalTrials.gov identifier: NCT04146051”
Methods are detailed enough for replication. The trial is registered (NCT04146051). A reporting summary is mentioned. All outcomes are reported, including secondary and post hoc analyses. Limitations are extensively discussed. Conclusions are appropriately cautious, acknowledging the small sample size and exploratory nature. Funding and competing interests are disclosed.
“ClinicalTrials.gov identifier: NCT04146051”
“There are several important limitations to note. First, although randomized and placebo-controlled, this was a phase 2 study with a limited sample size and modest power (80%) to detect differences in the primary outcome”
“This work was supported by Cartesian Therapeutics, which provided the study product and was involved in the design, execution and analysis of the study.”
“ClinicalTrials.gov identifier: NCT04146051”
“There are several important limitations to note. First, although randomized and placebo-controlled, this was a phase 2 study with a limited sample size and modest power (80%) to detect differences in the primary outcome”
“This work was supported by Cartesian Therapeutics, which provided the study product and was involved in the design, execution and analysis of the study.”
Registered (3 IDs: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 67 references by DOI: 67 verified.
Every extracted reference resolved against Crossref/OpenAlex with no retraction flags.
3 data/code links checked; 3 live.
- datahttps://clinicaltrials.gov/study/NCT04146051LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT07089121?term=NCT07089121LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT06799247?term=NCT06799247LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
9 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 9 minor suggestions below.
9 copyedit issues flagged: mostly consistency, clarity, other.
- MINORconsistencyAbstract“33.0% of patients achieved minimum symptom expression (MSE) (MG-ADL score ≤1) by month 6”→ In the Results section, the percentage is reported as 33.3% (4/12). Ensure consistency between abstract and results.Abstract states 33.0% while results state 33.3%.
- MINORconsistencyAbstract“55.60% achieved MSE by month 6”→ In the Results section, the percentage is reported as 55.6% (5/9). Ensure consistency.Abstract uses 55.60% while results use 55.6%.
- MINORconsistencyResults, Participant characteristics“randomized 1:1 to receive six once-weekly intravenous infusions of Descartes-08 or placebo”→ The actual allocation was 15:11 (Descartes-08:placebo) in the efficacy population, not 1:1. Clarify that randomization was 1:1 but actual allocation was unequal due to dropouts.The text says 1:1 but the numbers are 15 vs 11.
- MINORconsistencyResults, Safety“17 of 20 (85.0%) participants in the Descartes-08 group and 13 of 16 (81.3%) participants in the placebo group reporting at least one AE”→ Check that the percentages are correct: 17/20 = 85%, 13/16 = 81.25% (rounds to 81.3%). This is fine.No issue.
- MINORconsistencyAbstract“33.0% of patients achieved minimum symptom expression (MSE) (MG-ADL score ≤1) by month 6”→ Ensure consistency with the Results section where 33.3% is reported.Abstract states 33.0% while Results state 33.3%.
- MINORconsistencyAbstract“55.60% achieved MSE by month 6”→ Use consistent decimal formatting (55.6% vs 55.60%).Inconsistent decimal places.
- MINORconsistencyResults, Participant characteristics“randomized 1:1”→ Clarify that the actual allocation was not 1:1 due to the mITT population.The randomization was 1:1, but the efficacy population was not balanced (15 vs 11).
- MINORclarityDiscussion“a short course of mRNA CAR T cell therapy to be disease-modifying”→ Revise to 'a short course of mRNA CAR T cell therapy to be disease-modifying' (missing verb).Grammatical issue.
- MINORotherData availability“trials@cartesiantx.com”→ Ensure this email is correct and active.Potential typo in email address.
The published work is robust and well-reported; an informed reader should weigh the small sample size and modest power as the main limitations, which the authors appropriately acknowledge. The minor copyedit inconsistencies (e.g., 33.0% vs 33.3% in abstract) are trivial and do not warrant a correction, but the authors may consider a clarification of the 1:1 randomization vs actual allocation for completeness.
- 1.MEDIUMcopyeditIn the Abstract, change '33.0%' to '33.3%' and '55.60%' to '55.6%' to match the Results section.The abstract and results report slightly different percentages for the same endpoints, which could confuse readers.
- 2.MEDIUMcopyeditIn the Results, Participant characteristics, clarify that randomization was 1:1 but the efficacy population was 15:11 due to dropouts, to avoid implying the actual allocation was unequal.The text says 'randomized 1:1' but the numbers are 15 vs 11, which is a potential inconsistency.
- 3.MEDIUMreportingIn the Discussion, fix the grammatical error: 'a short course of mRNA CAR T cell therapy to be disease-modifying' should be revised to include a verb (e.g., 'has the potential to be disease-modifying').The sentence is missing a verb, which affects clarity.
- 4.LOWdata codeIn the Data availability section, verify that the email address trials@cartesiantx.com is correct and active.A typo in the contact email could prevent data access requests from being received.
- 5.LOWreportingConsider adding a CONSORT flow diagram with explicit numbers for each stage (screened, enrolled, randomized, treated, analyzed) to enhance transparency.A flow diagram would clarify participant disposition and the difference between safety and efficacy populations.
- 6.LOWstatisticsConsider reporting exact p-values for secondary endpoints (e.g., MG-ADL change at month 3) with more decimal places for transparency.Exact p-values aid interpretation and reproducibility.
- 7.LOWreportingConsider adding a sensitivity analysis for the primary endpoint using different imputation methods.Sensitivity analyses would strengthen confidence in the primary result given the small sample size.
- 8.LOWreportingConsider discussing the potential for unblinding due to the characteristic adverse event profile (fevers, chills) and its impact on the results.This is a relevant limitation that could affect the interpretation of the primary outcome.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.