Measurable Residual Disease-Guided Therapy for Chronic Lymphocytic Leukemia.
Munir T, Girvan S, Cairns DA, Bloor A, Allsup D, Varghese AM, Gohil S, Paneesha S, Pettitt A, Eyre T, Fox CP, Forconi F, Kennedy B, Balotis C, Pemberton N, Sheehy O, Gribben J, Elmusharaf N, Gatto S, Preston G, Schuh A, Walewska R, Duley L, Webster N, Dalal S, Rawstron A, Howard D, Hockaday A, Jackson S, Greatorex N, Bell S, Stones D, Brown JM, Patten PEM, Hillmen P, UK CLL Trials Group
- DOI
- 10.1056/NEJMoa2504341
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/f505727f-e41f-44bd-b1d5-f41acf0b2df1 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ReportingData & code availability partially met−0.25★
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint for the ibrutinib-venetoclax vs ibrutinib comparison is undetectable measurable residual disease (uMRD) in bone marrow, a surrogate biomarker. The paper does not provide evidence of target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking uMRD to clinical outcomes in this context. Although the trial also reports progression-free and overall survival as secondary endpoints, the primary efficacy claim for this comparison is based on uMRD.
“The primary endpoint for ibrutinib-venetoclax vs ibrutinib was uMRD (sensitivity 10-4) in BM at two years after initiation of therapy”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted phase III randomized trial with rigorous design, clear reporting of demographics, ethics, and statistical methods. The main weaknesses are the absence of a data availability statement, lack of explicit statistical software identification, and minor copyedit issues including an internal inconsistency in death counts and impossible percentages in supplementary tables.
Both reviewers agreed on all dimensions and study type (interventional). The statistics verification recomputed 5 tests consistently, but coverage is limited to tests with test statistics/df or effect estimates with CIs; threshold-only p-values and exact p-values were not machine-verified. The integrity check flagged an internal contradiction in death counts and impossible percentages in supplementary tables, which are not reflected in the dimension scores but are noted as action items.
Numerical inconsistencies
2 findings · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
- Mathematically impossible statisticAssessed
Recomputed 5 tests: 5 consistent, 0 inconsistent; 5 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary endpoint: uMRD in BM within 2 years, ibrutinib-venetoclax vs FCR (172/260 vs 127/263).
“172/260 ibrutinib-venetoclax (66.2%), 127/263 FCR (48.3%) (p<0.001)”
Taken as given: The numbers 172 and 260 are the event count and group total for ibrutinib-venetoclax.; The numbers 127 and 263 are the event count and group total for FCR.; The test is a two-sided chi-square test on the 2x2 table.Method: Pearson chi-square test on the 2x2 table (172, 88, 127, 136).How we recomputed it: pChi2x2(172, 88, 127, 136) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary endpoint: uMRD in BM within 2 years, ibrutinib-venetoclax vs ibrutinib (172/260 vs 0/263).
“172/260 ibrutinib-venetoclax (66.2%), 0/263 ibrutinib (0%) (p<0.001)”
Taken as given: The numbers 172 and 260 are the event count and group total for ibrutinib-venetoclax.; The numbers 0 and 263 are the event count and group total for ibrutinib.; The test is a two-sided chi-square test on the 2x2 table.Method: Pearson chi-square test on the 2x2 table (172, 88, 0, 263).How we recomputed it: pChi2x2(172, 88, 0, 263) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Powered secondary endpoint: PFS HR for ibrutinib-venetoclax vs FCR (0.13, 95% CI 0.08-0.21).
“Hazard ratio (HR) for progression-free survival was: 0.13 (95% CI, 0.08 to 0.21, p<0.001) for ibrutinib-venetoclax vs FCR”
Taken as given: The HR is a ratio (log scale).; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Compute p-value from the HR and 95% CI using the normal approximation on the log scale.How we recomputed it: pCI(0.13, 0.08, 0.21, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Powered secondary endpoint: PFS HR for ibrutinib-venetoclax vs ibrutinib (0.29, 95% CI 0.17-0.49).
“HR for progression-free survival (ibrutinib-venetoclax vs. ibrutinib) was 0.29 (95%CI, 0.17 to 0.49; P<0.001)”
Taken as given: The HR is a ratio (log scale).; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Compute p-value from the HR and 95% CI using the normal approximation on the log scale.How we recomputed it: pCI(0.29, 0.17, 0.49, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2HR for OS ibrutinib-venetoclax vs FCR
“HR for overall survival was: 0.26 (95% CI, 0.13 to 0.50) for ibrutinib-venetoclax vs FCR”
Taken as given: The HR is 0.26 with 95% CI 0.13 to 0.50.; The CI is two-sided at 95%.; The p-value is two-sided.Method: Recomputed p from HR and CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.26, 0.13, 0.50, 1)
- lowinternal contradictionThe abstract reports 76 deaths (39 FCR, 26 ibrutinib, 11 ibrutinib-venetoclax), but the Results section reports 11 (4.2%) ibrutinib-venetoclax, 26 (9.9%) ibrutinib, and 39 (14.8%) FCR deaths, which sums to 76. However, the Safety section states 'Eleven deaths were seen in participants treated with ibrutinib-venetoclax, 25 with ibrutinib and 37 with FCR', which sums to 73 and conflicts with the earlier counts.
Eleven deaths were seen in participants treated with ibrutinib-venetoclax, 25 with ibrutinib and 37 with FCR
Resultsreviewer’s wording - lowimpossible statisticIn Supplementary Table S20, the percentage for 'Other' in the 4-5 years column is 115.2%, which exceeds 100% and is impossible for a proportion.
53 (115.2%)
Table S20reviewer’s wording
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Conclusions only partially backed by the presented evidenceAssessed
8 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1MRD-guided ibrutinib-venetoclax is more durable than fixed-duration therapy.The paper compares to historical fixed-duration trials, but this is a cross-trial comparison with caveats.Evidence: Discussion states MRD responses are more durable than reported in fixed-duration ibrutinib-venetoclax, but no direct comparison is made.
In FLAIR, MRD responses are more durable than reported in fixed-duration ibrutinib-venetoclax, possibly due to the personalized approach and longer exposure to ibrutinib-venetoclax.
Discussion ¶1reviewer’s wording - partialReviewer 2MRD-guided ibrutinib-venetoclax is more effective than continuous BTK inhibitors.The claim is supported by cross-trial comparisons, but the paper acknowledges the need for prospective head-to-head comparison.Evidence: The discussion compares results to other trials but states 'Prospective comparison of the ibrutinib-venetoclax time-limited treatment duration vs continuous BTK inhibitors is needed to definitively address the relative activity.'
With the caveats of cross-trial comparisons, the results from MRD-guided ibrutinib-venetoclax suggests improved survival outcomes over continuous BTK inhibitors.
Discussion ¶3reviewer’s wording - supportedReviewers 1, 2Ibrutinib-venetoclax leads to higher rates of uMRD compared to ibrutinib alone or FCR.The primary endpoint data directly support this claim.Evidence: 172/260 (66.2%) achieved uMRD in BM vs 0/263 (0%) and 127/263 (48.3%) with p<0.001.
Within two-years, 299 participants achieved uMRD in BM: 127/263 FCR (48.3%), 0/263 ibrutinib (0%) and 172/260 ibrutinib-venetoclax (66.2%) (p<0.001).
Abstractreviewer’s wording - supportedReviewer 1Ibrutinib-venetoclax improves progression-free survival compared to ibrutinib alone or FCR.The HRs and 5-year PFS estimates support this claim.Evidence: HR 0.13 (95% CI 0.08-0.21) vs FCR and 0.29 (95% CI 0.17-0.49) vs ibrutinib; 5-year PFS 93.9% vs 79.0% and 58.1%.
5-year progression-free survival was 58.1% with FCR, 79.0% with ibrutinib and 93.9% with ibrutinib-venetoclax.
Abstractreviewer’s wording - supportedReviewer 1Ibrutinib-venetoclax improves overall survival compared to ibrutinib alone or FCR.The HRs for OS support this claim, though the comparison vs ibrutinib has a wide CI.Evidence: HR 0.26 (95% CI 0.13-0.50) vs FCR and 0.41 (95% CI 0.20-0.83) vs ibrutinib; 5-year OS 95.9% vs 90.5% and 86.5%.
HR for overall survival was: 0.26 (95% CI, 0.13 to 0.50) for ibrutinib-venetoclax vs FCR; 0.41 (95% CI, 0.20 to 0.83) for ibrutinib-venetoclax vs ibrutinib
Abstractreviewer’s wording - supportedReviewer 1The unmutated IGHV subgroup benefits most from ibrutinib-venetoclax.Subgroup analyses show larger HRs in unmutated IGHV for PFS and OS.Evidence: HR for PFS 0.20 (95% CI 0.08-0.48) vs ibrutinib and 0.07 (95% CI 0.03-0.15) vs FCR in unmutated IGHV.
In participants with unmutated IGHV, progression-free survival was longer in patients treated with ibrutinib-venetoclax than ibrutinib alone (HR for PFS, 0.20; 95%CI, 0.08 to 0.48; Fig. 2B) or FCR (HR, 0.07; 95%CI, 0.03 to 0.15; Fig. 2B).
Resultsreviewer’s wording - supportedReviewer 2Ibrutinib-venetoclax improves progression-free survival compared to ibrutinib and FCR.The HRs and 5-year PFS estimates support this claim.Evidence: HR 0.13 (95% CI 0.08-0.21) vs FCR and 0.29 (95% CI 0.17-0.49) vs ibrutinib, both p<0.001.
Hazard ratio (HR) for progression-free survival was: 0.13 (95% CI, 0.08 to 0.21, p<0.001) for ibrutinib-venetoclax vs FCR; 0.29 (95% CI, 0.17 to 0.49, p<0.001) for ibrutinib-venetoclax vs ibrutinib
Abstractreviewer’s wording - supportedReviewer 2Ibrutinib-venetoclax improves overall survival compared to FCR.The HR for OS vs FCR is 0.26 (95% CI 0.13-0.50), supporting the claim.Evidence: HR 0.26 (95% CI 0.13-0.50) for ibrutinib-venetoclax vs FCR.
HR for overall survival was: 0.26 (95% CI, 0.13 to 0.50) for ibrutinib-venetoclax vs FCR
Abstractreviewer’s wording
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary endpoint for the ibrutinib-venetoclax vs ibrutinib comparison is undetectable measurable residual disease (uMRD) in bone marrow, a surrogate biomarker. The paper does not provide evidence of target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking uMRD to clinical outcomes in this context. Although the trial also reports progression-free and overall survival as secondary endpoints, the primary efficacy claim for this comparison is based on uMRD.
“The primary endpoint for ibrutinib-venetoclax vs ibrutinib was uMRD (sensitivity 10-4) in BM at two years after initiation of therapy”
- ADEQUATEEffect sizeThe effect sizes for the key clinical outcomes are large and clinically meaningful: 5-year PFS 93.9% vs 79.0% vs 58.1% for ibrutinib-venetoclax, ibrutinib, and FCR, respectively; HR for PFS 0.13 (95% CI 0.08-0.21) for ibrutinib-venetoclax vs FCR and 0.29 (95% CI 0.17-0.49) vs ibrutinib. Overall survival also improved with HR 0.26 (95% CI 0.13-0.50) vs FCR. These are anchored to survival outcomes and are statistically supported.
“5-year progression-free survival was 58.1% with FCR, 79.0% with ibrutinib and 93.9% with ibrutinib-venetoclax. Hazard ratio (HR) for progression-free survival was: 0.13 (95% CI, 0.08 to 0.21, p<0.001) for ibrutinib-venetoclax vs FCR; 0.29 (95% CI, 0.17 to 0.49, p<0.001) for ibrutinib-venetoclax vs ibrutinib”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The paper cites prior studies (GLOW, CAPTIVATE, CLARITY) and explains the biological rationale for combining ibrutinib and venetoclax. It acknowledges limitations of prior work, such as fixed-duration therapy and the need for individualized approaches. The premise is well-supported by the cited evidence.
Randomization method and unit are clearly described (computer-generated minimization, 1:1:1). Blinding is not applicable as it is an open-label trial, which is stated. Power analysis is described for the primary endpoint (O'Brien-Fleming alpha-spending) and the powered secondary endpoint (90 events). Inclusion/exclusion criteria are summarized and detailed in the appendix. Outlier handling is addressed through the pre-specified analysis population (ITT) and safety population. Controls are the comparator arms (ibrutinib and FCR). Independent replication is not applicable for a single pivotal trial.
The paper reports sex (male/female), age (median and IQR), ethnicity, and other baseline characteristics in Table 1. Since both sexes are enrolled, sex_justified is not applicable. Age, weight, and health status are reported (age, performance status, creatinine clearance). Demographics are adequately reported.
“Gender Male 187 (71.1%) 186 (70.7%) 186 (71.5%) 559 (71.1%) Female 76 (28.9%) 77 (29.3%) 74 (28.5%) 227 (28.9%)”
“Ethnicity White 240 (91.3%) 241 (91.6%) 235 (90.4%) 716 (91.1%)”
“Gender Male 187 (71.1%) 186 (70.7%) 186 (71.5%) 559 (71.1%) Female 76 (28.9%) 77 (29.3%) 74 (28.5%) 227 (28.9%)”
“Ethnicity White 240 (91.3%) 241 (91.6%) 235 (90.4%) 716 (91.1%)”
The paper states that a national ethics committee and institutional review boards approved the protocol, and participants provided written informed consent. Compliance with the Declaration of Helsinki is stated. This meets the criteria for human research.
“Participants provided written informed consent.”
“The trial was performed in accordance with the Declaration of Helsinki.”
“Participants provided written informed consent.”
“The trial was performed in accordance with the Declaration of Helsinki.”
The trial uses ibrutinib and venetoclax as investigational products; they are named with manufacturers (Johnson and Johnson, AbbVie) and dosing regimens. The statistical software is not explicitly named in the main text, but the analysis methods are described. Since this is a drug trial, the bench criteria (antibodies, cell lines, mycoplasma, organisms) are not applicable. Reagents are scored against the investigational products, which are adequately identified.
“Ibrutinib monotherapy was administered orally 420 mg/day for six years.”
Tests are named (binary logistic penalized regression, Cox proportional hazards, Kaplan-Meier). Assumptions are handled by design (Cox model, penalized regression). Exact p-values are reported for primary and powered secondary endpoints (e.g., p<0.001). Effect sizes are reported with 95% CIs. Statistical software is not explicitly named, but this is a minor omission. Data presentation includes Kaplan-Meier curves and forest plots, which is appropriate for a clinical trial. Mathematical plausibility is not applicable for large-N continuous outcomes.
No data availability statement is present in the paper. The protocol is mentioned as available at NEJM.org, but this does not constitute a data availability statement for the trial data. No repository deposits or accession numbers are provided. Code sharing is not applicable as no bespoke code is described.
Trial registration numbers are provided. Methods are detailed enough for replication. Limitations are discussed (e.g., open-label design, cross-trial comparisons, need for longer follow-up). Conclusions are proportional to the evidence. Funding sources and conflicts of interest are disclosed. Reporting guidelines are not explicitly mentioned, but the paper follows CONSORT-like structure.
“Primary financial support was from Cancer Research UK (C18027/A15790).”
“Primary financial support was from Cancer Research UK (C18027/A15790).”
Registered (2 IDs: ISRCTN). No reporting guideline cited.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 20 references by DOI: 18 verified — 2 no DOI (shown, not verified).
- NO DOIU.S. Population Data 1969 - 2022 with Other SoftwareNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA multiple testing procedure for clinical trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
13 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 13 minor suggestions below.
13 copyedit issues flagged: mostly typo, consistency.
- MINORtypoAbstract, Results“1 27 /263 FCR ( 48 . 3 %)”→ Remove extra spaces: '127/263 FCR (48.3%)'Spacing errors in numbers.
- MINORtypoAbstract, Results“0 /263 ibrutinib (0%)”→ Remove extra spaces: '0/263 ibrutinib (0%)'Spacing error.
- MINORtypoAbstract, Results“1 72 /260 ibrutinib - venetoclax ( 66.2 %)”→ Remove extra spaces: '172/260 ibrutinib-venetoclax (66.2%)'Spacing error.
- MINORtypoAbstract, Results“0. 64 (95% CI, 0. 39 to 1.05 )”→ Remove extra spaces: '0.64 (95% CI, 0.39 to 1.05)'Spacing error.
- MINORconsistencyResults, Safety“Eleven deaths were seen in participants treated with ibrutinib-venetoclax, 25 with ibrutinib and 37 with FCR”→ Ensure consistency with earlier death counts (11, 26, 39) and Table S23.Potential inconsistency in death counts.
- MINORtypoSupplementary Table S20“53 (115.2%)”→ Check percentage calculation; 115.2% is impossible.Percentage exceeds 100%.
- MINORtypoSupplementary Table S21“29 (120.8%)”→ Check percentage calculation; 120.8% is impossible.Percentage exceeds 100%.
- MINORconsistencyResults, Safety“742 of 756 (98.1%) in the safety population reported at least one adverse event.”→ Verify safety population total (756) vs. sum of group sizes (239+260+257=756).Safety population total is consistent.
- MINORtypoReferences“0.1056/nejm200012283432602.”→ Remove stray DOI fragment.Stray text at end of references.
- MINORtypoAbstract, Results“1 27 /263 FCR ( 48 . 3 %)”→ 127/263 FCR (48.3%)Spacing and formatting inconsistency.
- MINORconsistencyResults, Efficacy“0/263 ibrutinib (0%)”→ 0/263 ibrutinib (0%)Consistent formatting of fractions.
- MINORtypoResults, Safety“70 / 25 7 [ 27.2 %]”→ 70/257 [27.2%]Spacing issue.
- MINORconsistencyTable 2“FCR (n=239) I (n=260) I+V (n=257)”→ Ensure consistent use of n vs N.Minor inconsistency in capitalization.
The published work is methodologically robust and the main conclusions are supported by the reported analyses. An informed reader should weigh the missing data availability statement and the minor internal inconsistencies (death counts, impossible percentages) as reporting gaps that warrant correction or clarification, but they do not undermine the primary efficacy findings.
- 1.HIGHrigorReconcile the death counts: the Safety section states 'Eleven deaths were seen in participants treated with ibrutinib-venetoclax, 25 with ibrutinib and 37 with FCR' (sum 73), conflicting with the abstract and Results which report 11, 26, and 39 (sum 76). Correct the inconsistent numbers in the Safety section or elsewhere.An internal contradiction in a key safety outcome is a validity threat that could mislead readers and warrants a correction.
- 2.HIGHrigorCorrect the impossible percentages in Supplementary Tables S20 and S21 (e.g., 115.2% and 120.8%) which exceed 100% for a proportion.Impossible statistics indicate a calculation or reporting error that undermines the credibility of the supplementary data.
- 3.HIGHdata codeAdd a data availability statement to the manuscript (e.g., in Methods or a dedicated section) specifying how de-identified patient data can be accessed (e.g., via a data access committee or repository like YODA/Vivli) and the conditions for access.The absence of a data availability statement is a reporting gap for a clinical trial and is required by most journals and funders.
- 4.HIGHstatisticsExplicitly name the statistical software used (e.g., SAS, R) with version numbers in the Statistical Analysis section.Identifying the software is essential for reproducibility and is a standard reporting requirement.
- 5.MEDIUMreportingMention adherence to a reporting guideline such as CONSORT in the Methods or a separate section, and provide the checklist as supplementary material.Explicitly referencing CONSORT improves transparency and is expected for randomized trials.
- 6.MEDIUMcopyeditFix the spacing errors in the Abstract Results (e.g., '1 27 /263 FCR ( 48 . 3 %)' should be '127/263 FCR (48.3%)', '1 72 /260 ibrutinib - venetoclax ( 66.2 %)' should be '172/260 ibrutinib-venetoclax (66.2%)', '0. 64 (95% CI, 0. 39 to 1.05 )' should be '0.64 (95% CI, 0.39 to 1.05)').Spacing errors in key results detract from professionalism and readability.
- 7.MEDIUMcopyeditRemove the stray DOI fragment at the end of the References ('0.1056/nejm200012283432602.').Stray text in references is a formatting error that should be corrected.
- 8.MEDIUMcopyeditEnsure consistent formatting of fractions and percentages throughout the manuscript (e.g., '70 / 25 7 [ 27.2 %]' should be '70/257 [27.2%]').Consistent formatting improves clarity and avoids confusion.
- 9.LOWcopyeditStandardize the use of 'n' vs 'N' in Table 2 (e.g., 'FCR (n=239) I (n=260) I+V (n=257)').Minor capitalization inconsistency is a polish issue.
- 10.LOWdata codeConsider depositing the statistical analysis code (if any) in a public repository with a DOI to enhance reproducibility.Sharing analysis code would strengthen reproducibility, though it is not mandatory.
- 11.LOWreportingClarify the availability of the trial protocol and statistical analysis plan (e.g., at NEJM.org) in a dedicated statement.Providing a clear link to the protocol and SAP improves transparency.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.