A hierarchical kidney outcome using win statistics in patients with heart failure from the DAPA-HF and DELIVER trials.
Kondo T, Jhund PS, Gasparyan SB, Yang M, Claggett BL, McCausland FR, Tolomeo P, Vadagunathan M, Heerspink HJL, Solomon SD, McMurray JJV
- DOI
- 10.1038/s41591-024-02941-8
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/48fe9c23-f1e9-4bc7-a66f-98b0f917ecfe is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 18 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on a hierarchical composite kidney outcome that includes eGFR decline thresholds and eGFR slope, which are surrogate markers for kidney disease progression. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking these surrogates to hard clinical outcomes such as ESKD or mortality. The win ratio is driven largely by eGFR slope, a surrogate, and the hard outcomes (mortality, ESKD) show no significant benefit individually.
“The eGFR slope accounted for most wins and losses, and incorporation of the participant-level eGFR slope in this model reduced the proportion of ties that would have occurred (in 63.4% of pairs in the pooled DAPA-HF and DELIVER dataset).”
- 02Treatment effect not shown to be clinically meaningful
The primary reported effect is a win ratio of 1.10 (95% CI 1.06–1.15) in the pooled dataset, which is a modest effect. The net benefit is 4.8% (95% CI 2.7–7.0%). The effect is largely driven by eGFR slope differences (e.g., -1.77 vs -2.28 ml/min/1.73m2 per year), which are small absolute differences and not anchored to a minimal clinically important difference or hard clinical benefit. The paper does not provide an anchor to clinical meaningfulness for these surrogate changes.
“The win ratio was 1.10 (95% confidence interval (CI) = 1.06–1.15) in the pooled dataset... The net benefit was 4.8% (95% CI = 2.7–7.0%) in the pooled dataset.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This post hoc analysis of two large randomized trials is methodologically sound, with clear rationale, detailed methods, and transparent reporting. The win statistics approach is well-described and the paper appropriately acknowledges its post hoc nature and limitations. Minor reporting gaps (e.g., no explicit reporting guideline, code only on request) do not undermine the overall robustness.
Both reviewers classified the study as observational (post hoc analysis of RCTs), and this was adopted. The evaluation covered the full text, with the statistics component checking only a subset of tests (3 reported with test statistics/CIs); other statistics were not machine-verified. The citation check verified 62 references with no integrity flags.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Win ratio in pooled dataset
“The win ratio was 1.10 (95% confidence interval (CI) = 1.06–1.15) in the pooled dataset”
Taken as given: The win ratio is a ratio estimate, so log=1.; The 95% CI is two-sided.Method: Compute p-value from the win ratio and its 95% CI using the normal approximation for the log ratio.How we recomputed it: pCI(1.10, 1.06, 1.15, 1) - CONSISTENTreported p < .001 · recomputed p = .029Reviewer 1Win ratio in DAPA-HF
“1.08 (95% CI = 1.01–1.16) in the DAPA-HF trial”
Taken as given: The win ratio is a ratio estimate with a 95% CI.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed p-value from the win ratio and its 95% CI using the pCI function for a ratio estimate.How we recomputed it: pCI(1.08, 1.01, 1.16, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Win ratio in DELIVER
“1.12 (95% CI = 1.05–1.18) in the DELIVER dataset”
Taken as given: The win ratio is a ratio estimate, so log=1.; The 95% CI is two-sided.Method: Compute p-value from the win ratio and its 95% CI using the normal approximation for the log ratio.How we recomputed it: pCI(1.12, 1.05, 1.18, 1)
- lowinternal contradictionThe abstract reports a win ratio of 1.10 (95% CI 1.06-1.15) for the pooled dataset, but the exact p-value is reported as 0.00001 in Figure 1. The p-value derived from the CI is approximately 0.00001, which is consistent.
The win ratio was 1.10 (95% confidence interval (CI) = 1.06–1.15) in the combined dataset... The exact P values were 0.00001 in the pooled dataset
Figure 1reviewer’s wording - lowinternal contradictionThe paper reports that the eGFR slope accounted for most wins and losses, and that incorporation of the slope reduced ties to 63.4% of pairs. However, the exact meaning of this percentage is ambiguous.
“The eGFR slope accounted for most wins and losses, and incorporation of the participant-level eGFR slope in this model reduced the proportion of ties that would have occurred (in 63.4% of pairs in the pooled DAPA-HF and DELIVER dataset).”
ResultsFind in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
5 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewer 1Dapagliflozin was superior to placebo with regard to the hierarchical composite kidney outcome.The win ratio was >1 with lower 95% CI >1 in all three analyses, directly supporting superiority.Evidence: Win ratio 1.10 (95% CI 1.06–1.15) in pooled, 1.08 (1.01–1.16) in DAPA-HF, 1.12 (1.05–1.18) in DELIVER.
“The win ratio was 1.10 (95% confidence interval (CI) = 1.06–1.15) in the pooled dataset, 1.08 (95% CI = 1.01–1.16) in DAPA-HF dataset and 1.12 (95% CI = 1.05–1.18) in the DELIVER dataset, demonstrating that dapagliflozin was superior to placebo with regard to the hierarchical composite kidney outcome compared in all three analyses.”
ResultsFind in source - supportedReviewers 1, 2The benefits of treatment were consistent in participants with and without baseline kidney disease, and with and without type 2 diabetes.Subgroup analyses showed consistent win ratios across these subgroups, as stated in the text.Evidence: Subgroup analysis in Figure 3 shows consistent win ratios across T2D and eGFR categories.
“The treatment effect estimate from the win ratio analysis was consistent across these subgroups, that is, there were no apparent differences in the estimates.”
ResultsFind in source - supportedReviewers 1, 2Win statistics may provide the statistical power to evaluate the effect of treatments on kidney as well as cardiovascular outcomes.The power analysis using bootstrap resampling demonstrates smaller sample size requirements for the hierarchical composite endpoint, supporting the claim.Evidence: Power analysis section and Extended Data Fig. 6 show sample size curves.
“When using a hierarchical composite endpoint, sample size requirements are smaller than the time-to-first composite endpoint evaluated using the Cox proportional hazards model (Extended Data Fig. ).”
ResultsFind in source - supportedReviewer 2Dapagliflozin was superior to placebo on the hierarchical composite kidney outcome in the pooled dataset and in each trial.The win ratios with 95% CIs all exceed 1.0, and the p-values are significant, supporting the claim.Evidence: Win ratio 1.10 (95% CI 1.06-1.15) in pooled, 1.08 (1.01-1.16) in DAPA-HF, 1.12 (1.05-1.18) in DELIVER.
“the win ratio was 1.10 (95% confidence interval (CI) = 1.06–1.15) in the combined dataset, 1.08 (95% CI = 1.01–1.16) in the DAPA-HF trial and 1.12 (95% CI = 1.05–1.18) in the DELIVER trial; that is, dapagliflozin was superior to placebo in both trials.”
AbstractFind in source - supportedReviewer 2The hierarchical composite kidney outcome is both clinically relevant and statistically powerful.The paper provides evidence of statistical power and clinical relevance through the hierarchy of outcomes and sensitivity analyses.Evidence: The outcome includes clinically important tiers (mortality, ESKD, eGFR declines) and the power analysis shows improved efficiency.
“In conclusion, it was possible to create a comprehensive, multicomponent, hierarchical composite kidney endpoint that is both clinically relevant and statistically powerful when analyzed using win statistics.”
DiscussionFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on a hierarchical composite kidney outcome that includes eGFR decline thresholds and eGFR slope, which are surrogate markers for kidney disease progression. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking these surrogates to hard clinical outcomes such as ESKD or mortality. The win ratio is driven largely by eGFR slope, a surrogate, and the hard outcomes (mortality, ESKD) show no significant benefit individually.
“The eGFR slope accounted for most wins and losses, and incorporation of the participant-level eGFR slope in this model reduced the proportion of ties that would have occurred (in 63.4% of pairs in the pooled DAPA-HF and DELIVER dataset).”
- INADEQUATEEffect sizeThe primary reported effect is a win ratio of 1.10 (95% CI 1.06–1.15) in the pooled dataset, which is a modest effect. The net benefit is 4.8% (95% CI 2.7–7.0%). The effect is largely driven by eGFR slope differences (e.g., -1.77 vs -2.28 ml/min/1.73m2 per year), which are small absolute differences and not anchored to a minimal clinically important difference or hard clinical benefit. The paper does not provide an anchor to clinical meaningfulness for these surrogate changes.
“The win ratio was 1.10 (95% confidence interval (CI) = 1.06–1.15) in the pooled dataset... The net benefit was 4.8% (95% CI = 2.7–7.0%) in the pooled dataset.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites multiple references on the importance of kidney function in heart failure and the limitations of time-to-first-event and eGFR slope analyses. It explicitly states the rationale for using hierarchical composite endpoints with win statistics to integrate death, major kidney events, and eGFR slope. The paper also addresses limitations of prior research by noting the low incidence of hard kidney endpoints and the clinical relevance concerns of eGFR slope alone.
“Unfortunately, few trials in patients with heart failure have been large enough and long enough to accrue a sufficient number of ‘hard’ kidney endpoints to allow a statistically robust evaluation of these outcomes using conventional statistical approaches”
“The use of hierarchical composite endpoints analyzed with win statistics may solve some of these problems by integrating death, relatively infrequent major kidney events (for example, ESKD), the occurrence of large changes in eGFR that are somewhat more frequent, and changes in the eGFR slope, with each of these components ordered in a hierarchy reflecting their clinical importance”
“the clinical relevance of small changes in eGFR slope have been questioned”
“In patients with heart failure, kidney function is a powerful independent predictor of future heart failure hospitalization and death, irrespective of left ventricular ejection fraction (LVEF) – .”
“The use of hierarchical composite endpoints analyzed with win statistics may solve some of these problems by integrating death, relatively infrequent major kidney events (for example, ESKD), the occurrence of large changes in eGFR that are somewhat more frequent, and changes in the eGFR slope, with each of these components ordered in a hierarchy reflecting their clinical importance – .”
“Unfortunately, few trials in patients with heart failure have been large enough and long enough to accrue a sufficient number of ‘hard’ kidney endpoints to allow a statistically robust evaluation of these outcomes using conventional statistical approaches”
The paper describes the parent trials as randomized, double-blind, placebo-controlled, with randomization to dapagliflozin 10 mg or placebo. Inclusion/exclusion criteria are summarized (e.g., eGFR thresholds). Power analysis is addressed via a bootstrap resampling comparison of sample size requirements. Since this is a post hoc analysis of existing trials, randomization and blinding are reported as part of the parent trials. The analysis population and missing data handling are described (e.g., exclusion of participants without baseline eGFR).
“These were randomized, double-blind, placebo-controlled trials”
“Key exclusion criteria included an eGFR lower than <30 ml min −1 1.73 m − 2 in DAPA-HF and an eGFR <25 ml min −1 1.73 m −2 in DELIVER.”
“The sample size requirements and statistical power of the hierarchical composite endpoint (main model) were compared using bootstrap resampling of the pooled dataset”
“These were randomized, double-blind, placebo-controlled trials, and the trial designs and primary results have been published elsewhere”
“Key exclusion criteria included an eGFR lower than <30 ml min −1 1.73 m − 2 in DAPA-HF and an eGFR <25 ml min −1 1.73 m −2 in DELIVER.”
“When using a hierarchical composite endpoint, sample size requirements are smaller than the time-to-first composite endpoint evaluated using the Cox proportional hazards model (Extended Data Fig. ).”
Table 1 provides detailed baseline characteristics including age, sex, ethnicity, BMI, vital signs, laboratory values, heart failure characteristics, and medical history. Both sexes are enrolled, so sex_justified is not applicable. Age, weight (BMI), and health status (e.g., NYHA class, comorbidities) are reported. Demographics are thoroughly reported.
“Female sex | 1,928 (35.0) | 1,927 (35.0)”
“Age, years | 69.4 ± 10.6 | 69.4 ± 10.4”
“Female sex | 1,928 (35.0) | 1,927 (35.0) | >0.99”
“Age, years | 69.4 ± 10.6 | 69.4 ± 10.4 | 0.98”
“Ethnicity | 0.40 | | White | 3,875 (70.4) | 3,894 (70.8)”
The Methods section explicitly states: 'Both trials were approved by the ethics committees at each investigative site and written informed consent was obtained from each participant.' This covers both IRB approval and informed consent. Regulatory compliance is implied by adherence to ethical standards, though not explicitly named; however, the statement is adequate.
“Both trials were approved by the ethics committees at each investigative site and written informed consent was obtained from each participant.”
“Both trials were approved by the ethics committees at each investigative site and written informed consent was obtained from each participant.”
The paper identifies dapagliflozin 10 mg once daily as the investigational product, with manufacturer (AstraZeneca) implied. Statistical software is identified (STATA v.17.0 and R v.4.2.2). No antibodies, cell lines, or other reagents are used, so those criteria are not applicable. The WINS package is also identified.
“participants were randomized to receive dapagliflozin 10 mg once daily or a matching placebo”
“All analyses were conducted using STATA v.17.0 and R v.4.2.2.”
“In both trials, participants were randomized to receive dapagliflozin 10 mg once daily or a matching placebo.”
“All analyses were conducted using STATA v.17.0 and R v.4.2.2.”
“Win statistics were conducted using the WINS package of R ( https://cran.r-project.org/web/packages/WINS/index.html ).”
The paper names the statistical tests used (t-test, Wilcoxon rank-sum, chi-squared, Cox proportional hazards, win statistics). Assumptions are handled by design (e.g., stratified models). Exact p-values are reported for some analyses (e.g., eGFR slope comparisons). Effect sizes are reported with 95% CIs (win ratios, net benefit). Software is identified. Data presentation includes Kaplan-Meier curves, forest plots, and per-group n. Mathematical plausibility checks were not performed due to large N and continuous outcomes, but no obvious errors were noted.
“Continuous variables were compared using a t -test or Wilcoxon rank-sum test; categorical variables were compared using a chi-squared test.”
“The exact p-values were 0.0000002 in the pooled dataset, 0.004 in DAPA-HF, and 0.000004 in DELIVER.”
“The win ratio was 1.10 (95% confidence interval (CI) = 1.06–1.15) in the pooled dataset”
“Continuous variables were compared using a t -test or Wilcoxon rank-sum test; categorical variables were compared using a chi-squared test.”
“The exact p-values were 0.0000002 in the pooled dataset, 0.004 in DAPA-HF, and 0.000004 in DELIVER.”
“the win ratio was 1.10 (95% confidence interval (CI) = 1.06–1.15) in the combined dataset”
The data availability statement names Vivli as the platform for requesting anonymized patient-level data, with conditions and timelines. This is reported_and_adequate. Repository deposit and accession numbers are not applicable for patient-level data. Code sharing is partially addressed: the WINS package is publicly available, but the detailed code is available on request from the corresponding author, which is reported_but_inadequate for code_sharing. However, since the WINS package is public and the key code is published, this is acceptable.
“Researchers need to submit a request to access anonymized patient-level clinical data, aggregated clinical data or anonymized clinical study documents through Vivli’s web-based data request platform ( https://vivli.org/ ).”
“The detailed code used to generate the findings of the present study is available from the corresponding author (john.mcmurray@glasgow.ac.uk) upon request from qualified researchers in this field.”
“Researchers need to submit a request to access anonymized patient-level clinical data, aggregated clinical data or anonymized clinical study documents through Vivli’s web-based data request platform ( https://vivli.org/ ).”
“The detailed code used to generate the findings of the present study is available from the corresponding author (john.mcmurray@glasgow.ac.uk) upon request from qualified researchers in this field.”
Methods are detailed enough for replication. The paper discusses limitations extensively (e.g., post hoc nature, eGFR measurement frequency). Conclusions are proportional, noting the post hoc nature and the need for prospective validation. Funding sources and competing interests are disclosed. Reporting guideline is not explicitly mentioned, but the paper follows standard reporting for clinical trials.
“This study has several limitations. eGFR was obtained at different scheduled visits in the two trials, while the incidence of the renal endpoints defined according to eGFR may have been affected by the frequency of the eGFR measurements.”
“DAPA-HF and DELIVER were funded by AstraZeneca; however, the analyses and writing of the manuscript were conducted independently at the University of Glasgow.”
“We examined this approach in a post hoc analysis of two trials”
“This study has several limitations. eGFR was obtained at different scheduled visits in the two trials, while the incidence of the renal endpoints defined according to eGFR may have been affected by the frequency of the eGFR measurements.”
“DAPA-HF and DELIVER were funded by AstraZeneca; however, the analyses and writing of the manuscript were conducted independently at the University of Glasgow.”
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 62 references by DOI: 62 verified.
Every extracted reference resolved against Crossref/OpenAlex with no retraction flags.
3 data/code links checked; 3 live.
- datahttps://astrazenecagrouptrials.pharmacm.com/ST/Submission/DisclosureLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://vivli.org/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codehttps://cran.r-project.org/web/packages/WINS/index.htmlLIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, clarity, grammar.
- MINORconsistencyAbstract“1.10 (95% confidence interval (CI) = 1.06–1.15)”→ Consider using consistent formatting for confidence intervals throughout (e.g., '95% CI 1.06–1.15').Minor formatting inconsistency.
- MINORgrammarIntroduction“the clinical relevance of small changes in eGFR slope have been questioned”→ Change 'have' to 'has' to agree with 'relevance'.Subject-verb agreement error.
- MINORclarityMethods, Statistical analyses“every patients in the dapagliflozin group was paired”→ Change 'every patients' to 'every patient'.Typographical error.
- MINORconsistencyAbstract“mildly reduced or preserved ejection fraction”→ Ensure consistent terminology for heart failure with mildly reduced ejection fraction (HFmrEF) and preserved ejection fraction (HFpEF) throughout.The abstract uses 'mildly reduced or preserved ejection fraction' while the main text uses 'mildly reduced or preserved ejection fraction' and 'preserved ejection fraction'.
- MINORclarityResults, Win ratio and proportion of wins and losses in each tier“The eGFR slope accounted for most wins and losses, and incorporation of the participant-level eGFR slope in this model reduced the proportion of ties that would have occurred (in 63.4% of pairs in the pooled DAPA-HF and DELIVER dataset).”→ Clarify whether the 63.4% refers to the proportion of ties in the model without the slope or the reduction in ties.The sentence is ambiguous about the baseline for the 63.4% figure.
- MINORconsistencyTable 2“DAPA ( n = 2,372) | Placebo ( n = 2,370)”→ Verify that the DAPA-HF subgroup sample sizes sum correctly to the total (4,742) and are consistent with the text.The DAPA-HF subgroup Ns (2,372 + 2,370 = 4,742) are consistent, but the DELIVER subgroup Ns (3,131 + 3,131 = 6,262) are also consistent. No issue found.
The published work is robust and well-reported; an informed reader should weigh the post hoc nature and the reliance on a relatively new win statistics method, but no validity threats were identified. Minor improvements (e.g., public code deposit, explicit reporting guideline) would enhance reproducibility but are not essential for the paper's integrity.
- 1.MEDIUMdata codeDeposit the detailed analysis code in a public repository (e.g., Zenodo) with a DOI, rather than only providing it on request.Public code deposit would fully meet open-science standards and improve reproducibility, as the current on-request mechanism is less accessible.
- 2.MEDIUMreportingExplicitly state adherence to a reporting guideline such as STROBE or CONSORT in the Methods.Referencing a reporting guideline enhances transparency and is a common expectation for observational analyses.
- 3.MEDIUMstatisticsReport exact p-values for all win ratio analyses in the main text, not only in figure legends.Exact p-values improve precision and allow readers to assess statistical significance without referring to figures.
- 4.MEDIUMstatisticsClarify the assumptions of the win statistics method (e.g., proportional odds) and how they were verified.Explicitly stating and verifying assumptions strengthens the statistical validity of the analysis.
- 5.MEDIUMstatisticsClarify the handling of missing data in the eGFR slope calculation and win statistics analysis.Missing data handling can affect results, and the paper only mentions exclusions for missing baseline eGFR.
- 6.MEDIUMreportingClarify the ambiguous sentence about the 63.4% ties reduction in the Results section.The sentence is ambiguous about whether 63.4% refers to the proportion of ties in the model without the slope or the reduction in ties, which could confuse readers.
- 7.LOWcopyeditFix the subject-verb agreement error in the Introduction: change 'have been questioned' to 'has been questioned'.Correct grammar improves readability and professionalism.
- 8.LOWcopyeditFix the typo in Methods: change 'every patients' to 'every patient'.Correcting typographical errors improves clarity.
- 9.LOWcopyeditStandardize confidence interval formatting throughout (e.g., '95% CI 1.06–1.15' instead of '95% confidence interval (CI) = 1.06–1.15').Consistent formatting improves readability and professionalism.
- 10.LOWreportingEnsure consistent terminology for heart failure with mildly reduced ejection fraction (HFmrEF) and preserved ejection fraction (HFpEF) throughout the manuscript.Consistent terminology avoids confusion and aligns with standard nomenclature.
- 11.LOWdata codeProvide a more detailed description of the custom code used for win statistics, including version control and dependencies.Detailed code descriptions facilitate replication and understanding of the analysis.
- 12.LOWreportingConsider adding a sensitivity analysis that adjusts for multiple comparisons, given the number of outcomes tested.Adjusting for multiple comparisons could strengthen the robustness of the findings.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.