Neoadjuvant nivolumab with or without relatlimab in resectable non-small-cell lung cancer: a randomized phase 2 trial.
Schuler M, Cuppens K, Plönes T, Wiesweg M, Du Pont B, Hegedus B, Köster J, Mairinger F, Darwiche K, Paschen A, Maes B, Vanbockrijck M, Lähnemann D, Zhao F, Hautzel H, Theegarten D, Hartemink K, Reis H, Baas P, Schramm A, Aigner C
- DOI
- 10.1038/s41591-024-02965-0
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/7150f091-4ade-4080-8cf4-4cb9575e4d7b is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×2−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on pathological response (major pathological response, MPR) and radiographic response rates, which are surrogate endpoints. The study does not demonstrate target engagement at the tested dose (no PK/PD data) and does not cite validated evidence linking MPR to long-term clinical outcomes in this setting. The paper itself acknowledges limitations in assessing clinical efficacy.
“The rates of major pathological responses (MPR, ≤10% viable tumor cells) were 27% and 30% ... With a median duration of follow-up of 12 months, rates of DFS and overall survival (OS) at 12 months were 89% and 93% with nivolumab monotherapy, and 93% and 100%…”
- 02Treatment effect not shown to be clinically meaningful
The reported effect sizes (e.g., MPR rates of 27% and 30%) are presented without anchoring to a minimal clinically important difference or to the normal/reference value. The study is not powered for efficacy and the authors state it was not designed for formal statistical comparison. The clinical meaningfulness of these response rates is not established.
“The study was not designed for formal statistical comparison of both treatment arms. ... the moderate sample size and study design preclude formal assessment of clinical efficacy.”
- 03Printed percentage does not match its own count
30% does not match the reported count 9/29
“and 30%”
Results - 04Printed percentage does not match its own count
89% does not match the reported count 27/30
“rates of DFS and overall survival (OS) at 12 months were 89% and 93% with nivolumab monotherapy”
ResultsFind in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported phase 2 randomized trial with strong scientific premise, clear design, and comprehensive reporting of ethics, resources, data, and code. Minor reporting gaps include lack of explicit reporting guideline, incomplete power analysis, and some ambiguous phrasing in the abstract and results.
Both reviewers classified the study as interventional, and I adopt that classification. The evaluation covered the full text, including methods, results, and supplementary materials. Non-applicable criteria (e.g., animal housing, cell line authentication) were excluded. The statistics verification component checked only 2 tests, both inconsistent, but these are not listed in the inconsistent tests array, so they are not treated as demonstrable errors.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
2 printed percentages that do not match their own count.
- PERCENT30% does not match the reported count 9/29
“and 30%”
Results - PERCENT89% does not match the reported count 27/30
“rates of DFS and overall survival (OS) at 12 months were 89% and 93% with nivolumab monotherapy”
ResultsFind in source
- lowinternal contradictionThe abstract reports major pathological response rates of 27% and 30% for arms A and B, while the results section states 'The rates of major pathological responses (MPR, ≤10% viable tumor cells) were 27% and 30%' - consistent. However, the abstract also reports 'objective radiographic responses were achieved in 27% and 10% (nivolumab) and in 30% and 27% (nivolumab and relatlimab) of patients, respectively.' This phrasing is confusing: it lists four numbers for two arms, likely meaning 27% and 10% for arm A and 30% and 27% for arm B, but the order is ambiguous.
“Major pathological (≤10% viable tumor cells) and objective radiographic responses were achieved in 27% and 10% (nivolumab) and in 30% and 27% (nivolumab and relatlimab) of patients, respectively.”
AbstractFind in source - lowinternal contradictionThe abstract states 'In 100% (nivolumab) and 90% (nivolumab and relatlimab) of patients, tumors and lymph nodes were pathologically completely resected.' The results section reports 'Complete surgical resection (R0) was achieved in 57 patients (95%)'. The abstract refers to pathological complete resection of tumors and lymph nodes, which may differ from R0 resection. This is not necessarily contradictory but could be clarified.
“In 100% (nivolumab) and 90% (nivolumab and relatlimab) of patients, tumors and lymph nodes were pathologically completely resected.”
AbstractFind in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
6 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2The study establishes the feasibility and safety of dual targeting of PD-1 and LAG-3 before lung cancer surgery.The primary endpoint (surgery within 43 days) was met by all patients, and safety data show manageable toxicity, supporting feasibility and safety.Evidence: All 60 patients proceeded to surgery within the protocol-defined time frame; grade ≥3 treatment-emergent AEs in 10% and 13%.
“This study establishes the feasibility and safety of dual targeting of PD-1 and LAG-3 before lung cancer surgery.”
AbstractFind in source - supportedReviewers 1, 2Curative resection was achieved in 95% of patients.The R0 resection rate of 95% is directly reported.Evidence: Complete surgical resection (R0) was achieved in 57 patients (95%).
Complete surgical resection (R0) was achieved in 57 patients (95%).
Resultsreviewer’s wording - supportedReviewers 1, 2Major pathological responses were achieved in 27% (nivolumab) and 30% (nivolumab+relatlimab).These rates are reported in the results section.Evidence: The rates of major pathological responses (MPR, ≤10% viable tumor cells) were 27% and 30%.
The rates of major pathological responses (MPR, ≤10% viable tumor cells) were 27% and 30%.
Resultsreviewer’s wording - supportedReviewers 1, 2Disease-free survival and overall survival rates at 12 months were 89% and 93% (nivolumab), and 93% and 100% (nivolumab+relatlimab).These survival rates are reported in the results section.Evidence: With a median duration of follow-up of 12 months, rates of DFS and OS at 12 months were 89% and 93% with nivolumab monotherapy, and 93% and 100% with nivolumab plus relatlimab.
With a median duration of follow-up of 12 months, rates of DFS and overall survival (OS) at 12 months were 89% and 93% with nivolumab monotherapy, and 93% and 100% with nivolumab plus relatlimab.
Resultsreviewer’s wording - supportedReviewers 1, 2Both treatments were safe with grade ≥3 treatment-emergent adverse events reported in 10% and 13% of patients per study arm.Safety data are reported in Table 2 and text.Evidence: Grade ≥3 treatment-emergent AEs were 10% (arm A) and 13% (arm B).
“Both treatments were safe with grade ≥3 treatment-emergent adverse events reported in 10% and 13% of patients per study arm.”
AbstractFind in source - supportedReviewers 1, 2Exploratory analyses provided insights into biological processes triggered by preoperative immunotherapy.The paper presents exploratory immune phenotyping, gene expression, and genomic analyses that provide such insights.Evidence: Immune cell phenotyping, gene expression profiling, and whole-exome sequencing analyses are reported.
“Exploratory analyses provided insights into biological processes triggered by preoperative immunotherapy.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on pathological response (major pathological response, MPR) and radiographic response rates, which are surrogate endpoints. The study does not demonstrate target engagement at the tested dose (no PK/PD data) and does not cite validated evidence linking MPR to long-term clinical outcomes in this setting. The paper itself acknowledges limitations in assessing clinical efficacy.
“The rates of major pathological responses (MPR, ≤10% viable tumor cells) were 27% and 30% ... With a median duration of follow-up of 12 months, rates of DFS and overall survival (OS) at 12 months were 89% and 93% with nivolumab monotherapy, and 93% and 100% with nivolumab plus relatlimab.”
- INADEQUATEEffect sizeThe reported effect sizes (e.g., MPR rates of 27% and 30%) are presented without anchoring to a minimal clinically important difference or to the normal/reference value. The study is not powered for efficacy and the authors state it was not designed for formal statistical comparison. The clinical meaningfulness of these response rates is not established.
“The study was not designed for formal statistical comparison of both treatment arms. ... the moderate sample size and study design preclude formal assessment of clinical efficacy.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction extensively cites prior work on neoadjuvant ICI, including SWOG S1801, NEOSTAR, and CheckMate 816, and discusses the rationale for dual checkpoint blockade. It acknowledges limitations of prior studies, such as the toxicity of chemoimmunotherapy and the need for better combinations. The hypothesis logically follows from the cited evidence.
“Although this approach resulted in impressive histopathological response rates and improved event-free survival, combined chemoimmunotherapy may obscure the contribution of the ICI component at the single patient level.”
“Based on their distinct and potentially synergistic mode of action, combined targeting of the immune checkpoints LAG-3 and PD-1 is a rational choice to overcome immune resistance in NSCLC.”
“Although this approach resulted in impressive histopathological response rates and improved event-free survival, combined chemoimmunotherapy may obscure the contribution of the ICI component at the single patient level.”
Randomization was performed via an interactive web response system (1:1) without stratification or blinding, which is appropriate for an open-label feasibility trial. The primary endpoint (surgery within 43 days) is clearly defined, and a sample size rationale is provided (up to 30 evaluable patients per arm). Inclusion/exclusion criteria are detailed. Blinding is not applicable as it is an open-label design, and the paper states this. Power analysis is not a formal statistical power calculation but a feasibility-based sample size, which is acceptable for a phase 2 trial.
“Patients were randomly assigned (1:1) via an interactive web response system provided by Alcedis GmbH”
“Based on published results of a study with preoperative nivolumab each study arm included up to 30 evaluable patients with the expectation that at least 26 of 30 patients treated in each study arm will undergo curatively intended surgery within 6 weeks of initiation of study treatment.”
“Patients were randomly assigned (1:1) via an interactive web response system provided by Alcedis GmbH ( https://www.alcedis.de/en ); there was no stratification or blinding.”
“Based on published results of a study with preoperative nivolumab each study arm included up to 30 evaluable patients with the expectation that at least 26 of 30 patients treated in each study arm will undergo curatively intended surgery within 6 weeks of initiation of study treatment.”
Table 1 reports sex (female/male counts), age (median and range), ECOG performance status, histology, clinical stage, PD-L1 status, and smoking status. Both sexes are enrolled, so sex_justified is not applicable. Age and health status (ECOG PS) are reported. Demographics are adequate for a clinical trial.
“n (female, male) | 30 (15, 15) | 30 (13, 17) | | Age in years, median (range) | 64 (43–77) | 67 (43–81)”
“ECOG PS (0, 1) | 28, 2 | 28, 2”
“n (female, male) | 30 (15, 15) | 30 (13, 17) | | Age in years, median (range) | 64 (43–77) | 67 (43–81)”
“Histology | | Adenocarcinoma | 13 | 15 | | Squamous cell carcinoma | 10 | 9 | | Adenosquamous carcinoma | 2 | 2 | | Other | 5 | 4”
The paper names the Ethics Committee of the Medical Faculty of the University Duisburg-Essen (approval 19-8828-AF) and other ethics committees for each site, with protocol numbers. It states that all patients provided written informed consent. Regulatory compliance with the Declaration of Helsinki and ICH-GCP is explicitly mentioned.
“the Ethics Committee of the Medical Faculty of the University Duisburg-Essen, Essen, Germany, granted primary approval on 10 September 2019 (19-8828-AF).”
“All patients provided written informed consent before enrollment.”
“The study was conducted according to the principles of the Declaration of Helsinki and the International Conference on Harmonization Good Clinical Practice guidelines.”
“The Ethics Committee of the Medical Faculty of the University Duisburg-Essen, Essen, Germany, granted primary approval on 10 September 2019 (19-8828-AF).”
“All patients provided written informed consent before enrollment.”
“The study was conducted according to the principles of the Declaration of Helsinki and the International Conference on Harmonization Good Clinical Practice guidelines.”
The drugs are named with manufacturer (Bristol Myers Squibb) and dosing (240 mg and 80 mg). Software tools for analysis are identified with versions (e.g., EdgeR v.3.40.0, Snakemake workflows) and code repositories are provided. Antibodies and reagents used in exploratory assays are identified with catalog numbers (e.g., PD-L1 clone 22C3, DAKO/Agilent M3653).
“two doses of nivolumab (240 mg every 14 days per intravenous infusion, arm A) or nivolumab and relatlimab (240 and 80 mg, respectively, every 14 days per intravenous infusion, arm B)”
“Further data analysis was performed using our open-source Snakemake workflow dna-seq-varlociraptor (v.3.24, https://github.com/snakemake-workflows/dna-seq-varlociraptor )”
“PD-L1 expression by tumor cells was assessed locally using the primary antibody clone 22C3 (DAKO/Agilent M3653)”
“This manuscript reports results from arms A and B of the study, which treated patients with two doses of nivolumab (240 mg every 14 days per intravenous infusion, arm A) or nivolumab and relatlimab (240 and 80 mg, respectively, every 14 days per intravenous infusion, arm B).”
“Further data analysis was performed using our open-source Snakemake workflow dna-seq-varlociraptor (v.3.24, https://github.com/snakemake-workflows/dna-seq-varlociraptor )”
The paper names tests (e.g., Wilcoxon matched pairs signed-rank test, log-rank test, EdgeR quasi-likelihood F-test). Exact p-values are reported for exploratory analyses (e.g., P = 0.04, P = 0.068). Effect sizes are presented as rates and survival percentages with confidence intervals implied in Kaplan-Meier curves. Statistical software is identified (EdgeR, etc.). Data presentation includes individual patient dots in figures and per-group n. Mathematical plausibility is not applicable for large-N continuous outcomes.
“Wilcoxon matched pairs signed-rank test was applied for statistical comparison.”
“Comparable effects were observed in responders treated with nivolumab monotherapy ( n = 13, P = 0.04) and nivolumab plus relatlimab ( n = 13, P = 0.068)”
“differential expression analysis was performed using the quasi-likelihood F -test approach of EdgeR (two-sided, v.3.40.0)”
“Wilcoxon matched pairs signed-rank test was applied for statistical comparison.”
“Comparable effects were observed in responders treated with nivolumab monotherapy ( n = 13, P = 0.04) and nivolumab plus relatlimab ( n = 13, P = 0.068)”
The data availability statement describes a managed-access process via the sponsor's Data Access Committee with a response timeframe. De-identified raw sequencing data are deposited in EGA with accession EGAS00001007753. Code is available in Zenodo with DOIs. For patient-level clinical data, managed access is appropriate and adequately described.
“Requests should be submitted to the Office of Data Governance of the study sponsor, University Hospital Essen ( https://www.uk-essen.de/ ), which also serves as Data Access Committee (DAC). Responses can be expected within 4 weeks.”
“De-identified raw data from gene expression profiling and whole-exome sequencing have been deposited in the European Genome-Phenome Archive (EGA) with accession number EGAS00001007753”
“The Snakemake workflows for whole-exome sequencing analysis and NanoString nCounter gene expression analysis can be found at https://zenodo.org/records/10838511”
“De-identified raw data from gene expression profiling and whole-exome sequencing have been deposited in the European Genome-Phenome Archive (EGA) with accession number EGAS00001007753 (https://ega-archive.org/studies/EGAS00001007753) . Requests should be submitted to the Office of Data Governance of the study sponsor, University Hospital Essen ( https://www.uk-essen.de/ ), which also serves as Data Access Committee (DAC). Responses can be expected within 4 weeks.”
“The Snakemake workflows for whole-exome sequencing analysis and NanoString nCounter gene expression analysis can be found at https://zenodo.org/records/10838511 (ref. ) and https://zenodo.org/doi/10.5281/zenodo.10838907 (ref. ).”
The trial is registered with ClinicalTrials.gov identifier NCT04205552. Methods are comprehensive. Limitations are explicitly discussed, including sample size, lack of formal statistical comparison, and potential bias from PD-L1 imbalance. Conclusions are appropriately cautious. Funding sources and competing interests are disclosed.
“ClinicalTrials.gov Indentifier: NCT04205552”
“First, the moderate sample size and study design preclude formal assessment of clinical efficacy, and appreciation of an additional contribution of relatlimab to pathological and radiographic response rates and survival endpoints.”
“M.S., K.C., B.H., J.K., F.M., A.P., B.M., H.R., P.B., A.S. and C.A. received institutional funding, paid to the University Hospital Essen, from Bristol Myers Squibb to support the study NEOpredict-Lung.”
“ClinicalTrials.gov Indentifier: NCT04205552 (https://clinicaltrials.gov/ct2/show/NCT04205552)”
“First, the moderate sample size and study design preclude formal assessment of clinical efficacy, and appreciation of an additional contribution of relatlimab to pathological and radiographic response rates and survival endpoints.”
“M.S., K.C., B.H., J.K., F.M., A.P., B.M., H.R., P.B., A.S. and C.A. received institutional funding, paid to the University Hospital Essen, from Bristol Myers Squibb to support the study NEOpredict-Lung.”
Registered (2 IDs: ClinicalTrials.gov, EudraCT). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 52 references by DOI: 50 verified — 2 no DOI (shown, not verified).
- NO DOIALINA: efficacy and safety of adjuvant alectinib versus chemotherapy in patients with early-stage ALK+ non-small cell lung cancer (NSCLC)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINeoadjuvant nivolumab (N) + ipilimumab (I) vs chemotherapy (C) in the phase III CheckMate 816 trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
6 data/code links checked; 6 live.
- dataEGALIVEHTTP 200https://ega-archive.org/studies/EGAS00001007753Resolves to EGA (data repository).
- datahttps://clinicaltrials.gov/ct2/show/NCT04205552LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.uk-essen.de/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- dataZenodoLIVEHTTP 200https://zenodo.org/records/10838511Resolves to Zenodo (data repository).
- dataZenodoLIVEHTTP 200https://zenodo.org/doi/10.5281/zenodo.10838907Resolves to Zenodo (data repository).
- codeGitHubLIVEHTTP 200https://github.com/snakemake-workflows/dna-seq-varlociraptorResolves to GitHub (code repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORtypoAbstract“Indentifier”→ IdentifierTypo in 'Identifier'.
- MINORconsistencyResults, Secondary outcomes“pathological responses (≤50% viable tumor cells) were observed in 60% and 72% of resected tumors and lymph nodes, respectively.”→ Clarify whether 'respectively' refers to arms A and B.Ambiguous reference for 'respectively'.
- MINORclarityMethods, Statistical analyses“All secondary parameters were evaluated in an explorative or descriptive manner, providing means, medians, ranges, standard deviations and/or confidence intervals.”→ Specify which parameters used which summary statistics.Vague description of statistical summaries.
- MINORconsistencyResults, Secondary outcomes“rates of DFS and overall survival (OS) at 12 months were 89% and 93% with nivolumab monotherapy, and 93% and 100% with nivolumab plus relatlimab”→ Clarify that 89% and 93% refer to DFS and OS respectively, and similarly for the combination arm.The sentence could be misread as DFS=89% and OS=93% for arm A, and DFS=93% and OS=100% for arm B, which is correct but could be clearer.
- MINORclarityMethods, Statistical analyses“All secondary parameters were evaluated in an explorative or descriptive manner, providing means, medians, ranges, standard deviations and/or confidence intervals.”→ Specify which parameters were evaluated with which statistics.Vague description of statistical methods.
The published work is robust and well-reported, with minor reporting gaps that do not undermine its conclusions. An informed reader should weigh the lack of a formal power calculation and the absence of an explicit reporting guideline, but these are not validity threats. The copyedit issues are minor and do not warrant a correction.
- 1.HIGHreportingAdd an explicit statement of adherence to a reporting guideline (e.g., CONSORT) in the Methods or Reporting Summary, and include a completed checklist as supplementary material.The paper does not explicitly mention a reporting guideline, which is a standard expectation for clinical trials and would strengthen transparency.
- 2.HIGHstatisticsAdd a formal power analysis or sample size justification with effect size, alpha, and power in the Statistical analyses section.The current sample size rationale is feasibility-based; a formal power calculation would clarify the study's ability to detect clinically meaningful differences.
- 3.MEDIUMstatisticsProvide exact p-values for all statistical comparisons, especially in figures, rather than only thresholds like 'P < 0.05'.Exact p-values allow readers to assess the strength of evidence more precisely.
- 4.MEDIUMstatisticsDescribe how outliers were handled in the statistical analysis, or state that no outliers were excluded.The paper does not explicitly describe outlier handling, which is a common reviewer request for transparency.
- 5.MEDIUMstatisticsInclude a statement on verification of statistical assumptions (e.g., normality, proportional hazards) for the tests used.Verification of assumptions is important for the validity of the statistical tests employed.
- 6.MEDIUMreportingClarify the blinding status: since the trial is open-label, state explicitly that blinding was not feasible and why.Explicitly stating the lack of blinding and its rationale addresses a potential reviewer concern.
- 7.MEDIUMdata codeIn the Data Availability section, specify the conditions for data access more concretely (e.g., proposal requirements, approval criteria).More concrete access conditions would improve the reproducibility of the data sharing process.
- 8.MEDIUMreportingAdd a section on patient and public involvement if applicable, or state that it was not done.Patient and public involvement is increasingly expected in clinical research reporting.
- 9.MEDIUMreportingConsider reporting the study according to the CONSORT extension for pilot/feasibility trials.This extension is specifically designed for feasibility studies and would improve reporting completeness.
- 10.MEDIUMreportingIn the Discussion, explicitly address the lack of formal statistical comparison between arms as a limitation and suggest future confirmatory trials.This would strengthen the interpretation of the results and set expectations for future work.
- 11.LOWcopyeditFix the typo 'Indentifier' to 'Identifier' in the Abstract.Correcting the typo improves professionalism and readability.
- 12.LOWcopyeditClarify the ambiguous 'respectively' in the Results section regarding pathological responses (60% and 72%) to specify which arm each percentage refers to.Ambiguous phrasing can lead to misinterpretation of the results.
- 13.LOWcopyeditClarify the sentence about DFS and OS rates to explicitly state that 89% and 93% refer to DFS and OS respectively for arm A, and similarly for arm B.This prevents potential misreading of the survival outcomes.
- 14.LOWcopyeditSpecify which secondary parameters were evaluated with which summary statistics (means, medians, ranges, SD, CI) in the Statistical analyses section.The current description is vague and could be more precise.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.