Combined endurance and resistance exercise training in heart failure with preserved ejection fraction: a randomized controlled trial.
Edelmann F, Wachter R, Duvinage A, Mueller S, Fegers-Wustrow I, Schwarz S, Christle JW, Pieske-Kraigher E, Seyfarth M, Knapp M, Dörr M, Nolte K, Düngen HD, Herrmann-Lingen C, Esefeld K, Hagendorff A, Haykowsky MJ, Hasenfuss G, Holzendorf V, Prettin C, Mende M, Pieske B, Halle M
- DOI
- 10.1038/s41591-024-03342-7
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/ba0f7033-6604-457f-8c1e-1872f1e9ef3c is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is a modified Packer score that includes hard outcomes (mortality, hospitalizations) but also surrogate components (peak VO2, E/e', NYHA class, GSA). The efficacy claim is primarily based on improvements in peak VO2 and NYHA class, which are surrogate or intermediate outcomes. The paper does not provide evidence linking these surrogates to hard clinical outcomes in this context, and target engagement at the tested dose is not established.
“clinically relevant differences favoring the ET group as compared to the UC group were observed for the following secondary endpoints: changes in peak VO2 (mean difference, 1.3 ml kg−1 min−1 (95% confidence interval (CI): 0.4–2.1)) and NYHA class (odds ratio…”
- 02Treatment effect not shown to be clinically meaningful
The primary endpoint was not met (P=0.17). The secondary improvements in peak VO2 (1.3 ml/kg/min) and NYHA class are presented as clinically meaningful, but the paper does not anchor these to an established minimal clinically important difference or demonstrate that they translate to improved hard outcomes. The effect size is small relative to the disease burden and the primary endpoint was negative.
“the modified Packer score showed an improvement in 33 ET patients (20.5%) and in 13 UC patients (8.1%) and showed a worsening in 69 ET patients (42.9%) and in 71 UC patients (44.1%) (Kendall’s tau-b = −0.073, P = 0.17).”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported multicenter RCT of combined exercise training in HFpEF. The main methodological strength is the rigorous design (randomization, sample size calculation, ITT analysis, CONSORT adherence, trial registration). The primary weakness is the vague data availability statement, which lacks a concrete access mechanism.
Both reviewers classified the study as interventional (RCT) and agreed on all dimension statuses; no divergence to reconcile. Non-applicable criteria (e.g., animal housing, cell line authentication) were excluded. The statistics verification recomputed only 5 of the reported tests; the remaining analyses were not machine-verified and should not be assumed correct.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 5 tests: 5 consistent, 0 inconsistent; 5 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1P-value for the clinical score OR (2.43, 95% CI 1.55-3.81) is reported as <0.001. The p-value can be approximated from the CI.
“OR (95% CI) for better category in exercise training | 2.43 (1.55–3.81) | <0.001”
Taken as given: The OR is from an ordinal regression model.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Approximated two-sided p-value from the OR and 95% CI using the normal approximation for the log-odds ratio.How we recomputed it: pCI(2.43, 1.55, 3.81, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1P-value for peak VO2 OR (1.09, 95% CI 1.03-1.14) is reported as <0.001. The p-value can be approximated from the CI.
“OR (95% CI) for better category in exercise training | 1.09 (1.03–1.14) | <0.001”
Taken as given: The OR is from an ordinal regression model.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Approximated two-sided p-value from the OR and 95% CI using the normal approximation for the log-odds ratio.How we recomputed it: pCI(1.09, 1.03, 1.14, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1P-value for NYHA class OR (5.89, 95% CI 3.08-11.25) is reported as <0.001. The p-value can be approximated from the CI.
“OR (95% CI) for better category in exercise training | 5.89 (3.08–11.25) | <0.001”
Taken as given: The OR is from an ordinal regression model.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Approximated two-sided p-value from the OR and 95% CI using the normal approximation for the log-odds ratio.How we recomputed it: pCI(5.89, 3.08, 11.25, 1) - CONSISTENTreported p = .940 · recomputed p = .938Reviewer 1P-value for GSA OR (1.02, 95% CI 0.62-1.69) is reported as 0.94. The p-value can be approximated from the CI.
“OR (95% CI) for better category in exercise training | 1.02 (0.62–1.69) | 0.94”
Taken as given: The OR is from an ordinal regression model.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Approximated two-sided p-value from the OR and 95% CI using the normal approximation for the log-odds ratio.How we recomputed it: pCI(1.02, 0.62, 1.69, 1) - CONSISTENTreported p = .390 · recomputed p = .392Reviewer 1P-value for septal E/e' OR (1.20, 95% CI 0.79-1.82) is reported as 0.39. The p-value can be approximated from the CI.
“OR (95% CI) for better category in exercise training | 1.20 (0.79–1.82) | 0.39”
Taken as given: The OR is from an ordinal regression model.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Approximated two-sided p-value from the OR and 95% CI using the normal approximation for the log-odds ratio.How we recomputed it: pCI(1.20, 0.79, 1.82, 1)
- lowinternal contradictionThe abstract reports 192 females (59.6%) and 130 males (40.4%), but Table 1 shows 100 females (62.1%) in ET and 92 (57.1%) in UC, totaling 192 females and 130 males, which is consistent. No contradiction.
“192 females (59.6%) and 130 males (40.4%)”
Table 1Find in source - lowinternal contradictionThe number of patients with available data for secondary endpoints is not fully reported in the main text, but this is a reporting gap, not a validity threat.
The number of patients with available data for the secondary endpoints is shown in Fig. or Supplementary Table.
Resultsreviewer’s wording
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
9 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewer 1The primary endpoint (modified Packer score) was not significantly different between ET and UC.The primary analysis showed no significant difference (Kendall's tau-b = -0.073, P=0.17), which directly supports the claim.Evidence: Kendall's tau-b = -0.073, P = 0.17
“the modified Packer score showed an improvement in 33 ET patients (20.5%) and in 13 UC patients (8.1%) and showed a worsening in 69 ET patients (42.9%) and in 71 UC patients (44.1%) (Kendall’s tau-b = −0.073, P = 0.17).”
AbstractFind in source - supportedReviewer 1ET improved peak VO2 compared to UC at 12 months.The mean difference of 1.3 ml/kg/min with 95% CI 0.4-2.1 (excluding zero) supports a significant improvement.Evidence: Mean difference 1.3 ml/kg/min (95% CI 0.4-2.1)
changes in peak VO2 (mean difference, 1.3 ml kg−1 min−1 (95% confidence interval (CI): 0.4–2.1))
Abstractreviewer’s wording - supportedReviewer 1ET improved NYHA class compared to UC at 12 months.The OR of 7.77 with 95% CI 3.73-16.21 strongly supports a significant improvement.Evidence: OR = 7.77 (95% CI: 3.73-16.21)
“NYHA class (odds ratio = 7.77 (95% CI: 3.73–16.21))”
AbstractFind in source - supportedReviewers 1, 2No significant differences were observed for other secondary endpoints (E/e', GSA, time to cardiovascular hospitalization, all-cause mortality).The paper reports non-significant results for these endpoints, consistent with the claim.Evidence: Reported non-significant p-values and CIs for E/e', GSA, and time to events.
“No significant between-group differences were observed for other secondary endpoints, including change in E/e′, change in GSA, time to cardiovascular hospitalization or all-cause mortality.”
AbstractFind in source - supportedReviewer 1The win ratio analysis favored ET over UC.The post hoc win ratio was 1.32 (95% CI 1.01-1.73, P=0.04), which supports a significant benefit, though it is post hoc.Evidence: Win ratio 1.32 (95% CI 1.01-1.73, P=0.04)
“Post hoc evaluation of the components of the primary endpoint using a more widely applied hierarchical win ratio revealed significantly more wins with ET than with UC (win ratio: 1.32 (95% CI: 1.01–1.73), P = 0.04; Extended Data Fig. ).”
ResultsFind in source - supportedReviewer 1Adherence to ET was associated with better outcomes.The paper reports a significant association between adherence and the modified Packer score (P=0.002) and change in peak VO2 (P<0.001), supporting the claim, though the authors note it is observational.Evidence: Test for trend in proportions P=0.002; ANOVA P<0.001
Adherence to ET was significantly associated with the modified Packer score (test for trend in proportions: P = 0.002; Extended Data Fig. ) and the change in peak VO2 ( P < 0.001; Extended Data Table ).
Resultsreviewer’s wording - supportedReviewer 21 year of combined endurance and resistance ET did not result in a significantly better modified Packer score compared to UC.The primary analysis (Kendall's tau-b = -0.073, P = 0.17) directly supports this claim.Evidence: Primary endpoint analysis: Kendall's tau-b = -0.073, P = 0.17.
“Although the primary endpoint was not met”
AbstractFind in source - supportedReviewer 2ET resulted in improvements in important clinical parameters, such as peak VO2 and NYHA class, as compared to UC.The paper reports statistically significant improvements in peak VO2 (mean difference 1.3 ml/kg/min, 95% CI 0.4-2.1) and NYHA class (OR 7.77, 95% CI 3.73-16.21) at 12 months, supporting this claim.Evidence: Secondary endpoint analyses: peak VO2 mean difference 1.3 (95% CI 0.4-2.1); NYHA class OR 7.77 (95% CI 3.73-16.21).
clinically relevant differences favoring the ET group as compared to the UC group were observed for the following secondary endpoints: changes in peak VO2 (mean difference, 1.3 ml kg−1 min−1 (95% confidence interval (CI): 0.4–2.1)) and NYHA class (odds ratio = 7.77 (95% CI: 3.73–16.21)).
Abstractreviewer’s wording - supportedReviewer 2The modified Packer score showed an improvement in 20.5% of ET patients and in 8.1% of UC patients.These percentages are directly reported in the results section and are consistent with the patient counts (33/161 and 13/161).Evidence: Results section: 'the modified Packer score showed improvement in 20.5% of patients in the ET arm and in 8.1% of patients in the UC arm'.
“the modified Packer score showed improvement in 20.5% of patients in the ET arm and in 8.1% of patients in the UC arm”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is a modified Packer score that includes hard outcomes (mortality, hospitalizations) but also surrogate components (peak VO2, E/e', NYHA class, GSA). The efficacy claim is primarily based on improvements in peak VO2 and NYHA class, which are surrogate or intermediate outcomes. The paper does not provide evidence linking these surrogates to hard clinical outcomes in this context, and target engagement at the tested dose is not established.
“clinically relevant differences favoring the ET group as compared to the UC group were observed for the following secondary endpoints: changes in peak VO2 (mean difference, 1.3 ml kg−1 min−1 (95% confidence interval (CI): 0.4–2.1)) and NYHA class (odds ratio = 7.77 (95% CI: 3.73–16.21)).”
- INADEQUATEEffect sizeThe primary endpoint was not met (P=0.17). The secondary improvements in peak VO2 (1.3 ml/kg/min) and NYHA class are presented as clinically meaningful, but the paper does not anchor these to an established minimal clinically important difference or demonstrate that they translate to improved hard outcomes. The effect size is small relative to the disease burden and the primary endpoint was negative.
“the modified Packer score showed an improvement in 33 ET patients (20.5%) and in 13 UC patients (8.1%) and showed a worsening in 69 ET patients (42.9%) and in 71 UC patients (44.1%) (Kendall’s tau-b = −0.073, P = 0.17).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites multiple prior studies on exercise training in HFpEF, noting both strengths (e.g., improvements in peak VO2) and limitations (e.g., short-term studies, lack of long-term data on combined training). The rationale linking the premise to the study's objectives is clearly stated, and the hypothesis follows from the cited evidence. Limitations of prior research, such as the short duration of most studies and the lack of a composite endpoint, are addressed by the design of the current trial.
“However, most studies have been short term (≤6 months) or limited to endurance training”
“Therefore, Packer proposed combining symptom severity outcomes and clinical events in a hierarchical, three-level ordinal score to evaluate the efficacy of HF therapies”
Randomization was performed via a web-based system using Pocock's minimization algorithm, stratified by center, NYHA class, and peak VO2. Blinding was applied to core laboratory assessments (echocardiography and CPET) and the endpoint committee, though participants and investigators were not blinded (inherent to exercise trials). A sample size calculation was provided with assumptions and power. Inclusion/exclusion criteria were described, and the primary analysis was intention-to-treat. Missing data were handled via multiple imputation as a sensitivity analysis.
“only data analyzed in the blinded echocardiography core laboratory were used for statistical analyses”
“Assuming an improved modified Packer score in 20% of ET patients and in 5% of UC patients and a worsened modified Packer score in 15% of ET patients and in 25% of UC patients, a power of 98% and a significance level of α = 0.05 results in 160 patients per treatment group.”
“Assuming an improved modified Packer score in 20% of ET patients and in 5% of UC patients and a worsened modified Packer score in 15% of ET patients and in 25% of UC patients, a power of 98% and a significance level of α = 0.05 results in 160 patients per treatment group.”
“a full description of inclusion and exclusion criteria is shown in Supplementary Table”
Table 1 provides a comprehensive baseline table with sex (female/male), age (mean and s.d.), BMI, NYHA class, smoking status, comorbidities (hypertension, atrial fibrillation, etc.), and medications. The study enrolled 60% female and 40% male participants, so a single-sex justification is not needed. Age and health status are reported. Species/strain and housing conditions are not applicable for a human clinical trial.
“Female | 100 (62.1) | 92 (57.1) | | Male | 61 (37.9) | 69 (42.9) | | Age, years; mean (s.d.) | 69.1 (7.4) | 70.1 (7.1)”
“322 patients (mean age, 70 years; 60% female, 86% New York Heart Association (NYHA) class II, 99% Caucasian)”
“Age, years; mean (s.d.) | 69.1 (7.4) | 70.1 (7.1)”
“322 patients (mean age, 70 years; 60% female, 86% New York Heart Association (NYHA) class II, 99% Caucasian)”
The methods state: 'The study was performed in accordance with the Declaration of Helsinki and was approved by the ethics committee of the University Medical Center Göttingen and the local ethics committees at all participating sites. All participants provided written informed consent.' This satisfies both IRB approval and informed consent requirements.
“The study was performed in accordance with the Declaration of Helsinki and was approved by the ethics committee of the University Medical Center Göttingen and the local ethics committees at all participating sites. All participants provided written informed consent.”
“The study was performed in accordance with the Declaration of Helsinki and was approved by the ethics committee of the University Medical Center Göttingen and the local ethics committees at all participating sites.”
“All participants provided written informed consent.”
The exercise intervention is described in detail: endurance training on a bicycle ergometer with progressive intensity and duration, and resistance training on weight machines with specified sets, repetitions, and intensity. The software used for statistical analysis is identified (SPSS version 24, R versions 3 and 4). Antibodies, cell lines, mycoplasma testing, and organisms are not applicable for this human clinical trial without wet-lab components.
“Data were extracted from the database and prepared for analysis with SPSS software version 24 and later (IBM). Descriptive statistics and the primary analysis were performed with SPSS. Linear mixed models and ordinal regression analyses were performed with R (versions 3 and 4)”
The primary analysis uses Kendall's tau-b, a non-parametric test. Secondary analyses use ordinal regression, linear mixed models, and Cox regression. Effect sizes with 95% CIs are reported for key outcomes (e.g., mean difference in peak VO2, OR for NYHA class). Exact p-values are provided for most analyses (e.g., P = 0.17 for primary, P < 0.001 for clinical score). Software is identified. Data presentation includes CONSORT diagram, Kaplan-Meier curves, and forest plots. The paper acknowledges that secondary endpoints were not adjusted for multiple comparisons and should be interpreted as exploratory. Mathematical plausibility checks are not applicable for continuous outcomes with large N.
“The main analysis of the primary endpoint was performed using the test of Kendall’s tau”
“Kendall’s tau-b = −0.073, P = 0.17”
The data availability statement says that individual participant data cannot be shared for legal reasons, but aggregated data may be shared upon reasonable request after consultation with data protection officers. This is a managed-access statement, but it does not specify a platform or a clear timeframe for the initial response (only 'within 4 weeks' after the request is made). The statement is therefore considered reported_but_inadequate because it lacks a concrete mechanism. No code was used for data collection, so code_sharing is not applicable.
“No computer code was used to collect the data in this study.”
“No computer code was used to collect the data in this study.”
The paper states it follows CONSORT guidelines. The trial is registered with ISRCTN (number provided). Funding sources and grant numbers are listed, and a detailed conflicts of interest statement is provided. Limitations are discussed in the Discussion section, including the asymmetric definition of the modified Packer score, lack of differentiation between endurance and resistance training effects, low adherence, and the inability to blind participants. Conclusions are proportional: the primary endpoint was not met, but secondary improvements are noted with appropriate caveats. Methods are detailed, including the exercise prescription, endpoint definitions, and statistical analysis plan.
“ISRCTN registration: ISRCTN86879094 (https://www.isrctn.com/ISRCTN86879094)”
“In this paper, we followed the CONSORT guidelines for reporting randomized controlled trials.”
“Our study has several limitations, including the asymmetric definition of the modified Packer score.”
“ISRCTN registration: ISRCTN86879094 (https://www.isrctn.com/ISRCTN86879094)”
“In this paper, we followed the CONSORT guidelines for reporting randomized controlled trials.”
“Our study has several limitations, including the asymmetric definition of the modified Packer score.”
Registered (1 ID: ISRCTN). Reporting guideline cited: CONSORT.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 51 references by DOI: 46 verified — 5 no DOI (shown, not verified).
- NO DOISequential treatment assignment with balancing for prognostic factors in the controlled clinical trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAnswering the call for a standard reliability measure for coding dataNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIContent Analysis: An Introduction to Its MethodologyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOImice: Multivariate Imputation by Chained Equations in RNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMultiple imputation after 18+ yearsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
3 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 3 minor suggestions below.
3 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAbstract“V ˙ O 2”→ Ensure the math notation for VO2 is rendered consistently as 'VO2' or 'V̇O2'.The LaTeX rendering may appear as 'V ˙ O 2' in some formats.
- MINORconsistencyTable 2, footnote“Exact P values for all-cause death, clinical score, peak VO2 and NYHA class were 1.00, 0.0001, 0.0003 and 0.0000001, respectively.”→ Consider using scientific notation for very small p-values (e.g., 1 × 10^-7) for readability.The p-value 0.0000001 is correct but could be formatted more clearly.
- MINORclarityResults, Primary endpoint“Regarding the components of the primary endpoint, no significant difference was observed between groups for number of all-cause deaths (P > 0.99) or hospitalizations (potentially) related to exercise or HF (P = 0.54), whereas significant group differences favoring ET were observed for the clinical score (odds ratio (OR) = 2.43 (95% confidence interval (CI): 1.55–3.81))”→ The phrase 'hospitalizations (potentially) related to exercise or HF' could be simplified to 'hospitalizations potentially related to exercise or HF' for clarity.Minor stylistic issue.
The published work is methodologically robust and transparent, with only a minor data-sharing transparency gap. An informed reader should weigh the vague data availability statement and the fact that only a subset of statistical tests were independently verified; neither issue undermines the core findings, but a correction or clarification of the data access route would strengthen the record.
- 1.HIGHdata codeIn the Data availability section, replace the vague 'upon reasonable request' statement with a concrete managed-access route, such as a named data-sharing platform (e.g., Vivli, YODA) or a data access committee with defined conditions and a clear timeline for the initial response.The current statement lacks a concrete mechanism, which is a transparency gap that reviewers and readers may flag.
- 2.HIGHdata codeConsider depositing aggregated, de-identified data in a public repository (e.g., Zenodo, Figshare) with a DOI, if legally and ethically permissible, and reference it in the Data availability section.A public repository deposit would satisfy the repository_deposit criterion and enhance reproducibility.
- 3.MEDIUMreportingIn the Abstract, add the exact p-value for the primary analysis of the modified Packer score (currently only 'P = 0.17' appears in the Results section).Reporting the exact p-value in the abstract improves transparency and consistency with the results section.
- 4.MEDIUMstatisticsIn the Methods, Statistical analysis section, clarify how missing data were handled for the primary endpoint (the multiple imputation analysis is described for secondary endpoints but not explicitly for the primary).Clarifying the missing-data handling for the primary endpoint strengthens the statistical reporting.
- 5.MEDIUMreportingIn the Methods, Study design section, provide a more detailed description of the blinding of outcome assessors (e.g., echocardiography core lab, endpoint committee), even though participants cannot be blinded.Detailed blinding information is a CONSORT requirement and improves methodological transparency.
- 6.MEDIUMreportingIn the Results, Secondary endpoints section, report the number of patients with available data for each secondary endpoint at each time point in the main text or a main table, rather than only in a supplementary figure.Reporting the number of patients with available data is essential for interpreting secondary analyses.
- 7.LOWcopyeditIn the Abstract, ensure the math notation for VO2 is rendered consistently as 'VO2' or 'V̇O2' (the LaTeX rendering may appear as 'V ˙ O 2').Consistent notation improves readability and avoids confusion.
- 8.LOWcopyeditIn Table 2 footnote, consider using scientific notation for very small p-values (e.g., 1 × 10^-7) for readability.Scientific notation improves clarity for extremely small p-values.
- 9.LOWcopyeditIn Results, Primary endpoint, simplify the phrase 'hospitalizations (potentially) related to exercise or HF' to 'hospitalizations potentially related to exercise or HF' for clarity.Minor stylistic improvement for readability.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.