Supervised, structured and individualized exercise in metastatic breast cancer: a randomized controlled trial.
Hiensch AE, Depenbusch J, Schmidt ME, Monninkhof EM, Pelaez M, Clauss D, Gunasekara N, Zimmer P, Belloso J, Trevaskis M, Rundqvist H, Wiskemann J, Müller J, Sweegers MG, Fremd C, Altena R, Gorecki M, Bijlsma R, van Leeuwen-Snoeks L, Ten Bokkel Huinink D, Sonke G, Lahuerta A, Mann GB, Francis PA, Richardson G, Malter W, van der Wall E, Aaronson NK, Senkus E, Urruticoechea A, Zopf EM, Bloch W, Stuiver MM, Wengstrom Y, Steindorf K, May AM
- DOI
- 10.1038/s41591-024-03143-y
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/19c21b3e-b394-4c03-9111-dbfc29e123da is authoritative.
How this rating was calculated
- StatisticsImpossible or misreported statistic ×2−2★
- CitationsUnresolved reference ×4−1★
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×3−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 4 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Printed percentage does not match its own countdemonstrable
42% does not match the reported count 43/179
“43 (42.0)”
Table 1 - 02Printed percentage does not match its own countdemonstrable
24.6% does not match the reported count 46/179
“46 (24.6)”
Table 1 - 03Efficacy rests on an unvalidated surrogate endpoint
The primary outcomes are patient-reported measures of physical fatigue (EORTC QLQ-FA12) and health-related quality of life (EORTC QLQ-C30 summary score). These are subjective, patient-reported outcomes that serve as surrogates for clinical benefit. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking these specific surrogate measures to hard clinical outcomes such as survival or disease progression. Although MIDs are mentioned for some secondary outcomes, they are not available for the primary outcomes, and the link to clinical meaningfulness is not established.
“Exercise resulted in significant positive effects on both primary outcomes. Physical fatigue was significantly lower (−5.3 (95% confidence interval (CI), −10.0 to −0.6), Bonferroni–Holm-adjusted P = 0.027; Cohen's effect size, 0.22) and HRQOL significantly…”
- 04Treatment effect not shown to be clinically meaningful
The reported effect sizes are small (Cohen's d = 0.22 for physical fatigue and 0.33 for HRQOL). The between-group differences are 5.3 points on a 0-100 scale for fatigue and 4.8 points for HRQOL. The paper does not provide a minimal clinically important difference (MID) for the primary outcomes, and the effect sizes are below the commonly accepted threshold for a small effect (0.2-0.5). The clinical meaningfulness is not explicitly anchored, as the MIDs are only available for secondary outcomes, not the primary ones.
“Physical fatigue was significantly lower (−5.3 (95% confidence interval (CI), −10.0 to −0.6), Bonferroni–Holm-adjusted P = 0.027; Cohen's effect size, 0.22) and HRQOL significantly higher (4.8 (95% CI, 2.2–7.4), Bonferroni–Holm-adjusted P = 0.0003; effect…”
- 05Printed percentage does not match its own count
64% does not match the reported count 114/179
“114 (64.0)”
Table 1 - 06Printed percentage does not match its own count
58.4% does not match the reported count 104/179
“104 (58.4)”
Table 1
1 further finding of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomized controlled trial. The methodology is rigorous, with clear randomization, power analysis, and appropriate statistical methods. Minor reporting issues (a potential percentage typo in Table 1 and a few copyedit items) do not undermine the overall integrity.
Both reviewers classified the study as interventional, which is adopted. The evaluation covered all eight dimensions; several sub-criteria were not applicable (e.g., animal housing, cell line authentication) due to the human trial nature. The statistics verification recomputed only a subset of tests; the remaining were not machine-verifiable but no errors were found.
Numerical inconsistencies
3 findings · worst criticalValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks. 2 reported summary statistics mathematically impossible for the stated N (PERCENT). 3 printed percentages that do not match their own count.
- PERCENT42% does not match the reported count 43/179
“43 (42.0)”
Table 1 - PERCENT24.6% does not match the reported count 46/179
“46 (24.6)”
Table 1 - PERCENT64% does not match the reported count 114/179
“114 (64.0)”
Table 1 - PERCENT58.4% does not match the reported count 104/179
“104 (58.4)”
Table 1 - PERCENT8.9% does not match the reported count 7/80
“fatigue (8.9%)”
Safety sectionFind in source
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Check p-value for HRQOL primary outcome at 6 months from reported BGD and CI.
“The exercise group reported significantly better HRQOL than the control group at 6 months (between-group difference (BGD), 4.8 (95% CI, 2.2–7.4); Bonferroni–Holm-adjusted P = 0.0003; effect size (ES), 0.33).”
Taken as given: The BGD is the estimate and the CI is a 95% confidence interval.; The CI is two-sided.; The p-value is two-tailed.Method: Recomputed two-tailed p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(4.8, 2.2, 7.4, 0) - CONSISTENTreported p = .027 · recomputed p = .027Reviewers 1, 2Check p-value for physical fatigue primary outcome at 6 months from reported BGD and CI.
“At 6 months, the exercise group also reported significantly lower physical fatigue levels compared to the control group (BGD, −5.3 (95% CI, −10.0 to −0.6); Bonferroni–Holm-adjusted P = 0.027; ES, 0.22)”
Taken as given: The BGD is the estimate and the CI is a 95% confidence interval.; The CI is two-sided.; The p-value is two-tailed.Method: Recomputed two-tailed p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-5.3, -10.0, -0.6, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Check p-value for HRQOL summary score at 6 months from reported BGD and CI.
“between-group difference (BGD), 4.8 (95% CI, 2.2–7.4); Bonferroni–Holm-adjusted P = 0.0003”
Taken as given: The BGD is the estimated difference in means.; The CI is a 95% confidence interval.; The p-value is two-sided.; The Bonferroni-Holm adjustment is applied to the raw p-value.Method: Compute p-value from estimate and 95% CI using normal approximation, then apply Bonferroni-Holm correction for two primary outcomes.How we recomputed it: pCI(4.8, 2.2, 7.4, 0)
- lowinternal contradictionIn Table 1, the percentage for 'Higher education' in the control group is listed as 42.0%, which seems inconsistent with the count of 43 out of 179 (which would be 24.0%).
“Higher education | 37 (20.8) | 43 (42.0)”
Table 1Find in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
6 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Exercise resulted in significant positive effects on both primary outcomes (physical fatigue and HRQOL) at 6 months.The reported between-group differences and adjusted p-values support this claim.Evidence: BGD for HRQOL: 4.8 (95% CI 2.2-7.4), adjusted P=0.0003; BGD for physical fatigue: -5.3 (95% CI -10.0 to -0.6), adjusted P=0.027.
“Exercise resulted in significant positive effects on both primary outcomes.”
AbstractFind in source - supportedReviewers 1, 2Supervised exercise should be recommended as part of supportive care for patients with MBC.The positive effects on primary and secondary outcomes, along with safety data, support this recommendation.Evidence: Significant improvements in HRQOL, fatigue, physical functioning, pain, and dyspnea; only two SAEs unrelated to bone metastases.
“These results demonstrate that supervised exercise has positive effects on physical fatigue and HRQOL in patients with MBC and should be recommended as part of supportive care.”
AbstractFind in source - supportedReviewer 1The exercise program was safe and well tolerated.Only two SAEs occurred, both unrelated to bone metastases, and high attendance/compliance rates support tolerability.Evidence: Two SAEs (fractures) not related to bone metastases; attendance rate 77%.
“In total, two exercise-related serious adverse events (SAEs) were reported: a wrist fracture and a sacral stress fracture. These were not related to bone metastases.”
ResultsFind in source - supportedReviewer 1The effects on fatigue and social functioning were clinically relevant as they exceeded MIDs.The paper states that BGDs for fatigue, social functioning, and role functioning were larger than published MIDs.Evidence: BGD for fatigue at 6 months: -8.0 (95% CI -12.7 to -3.4), MID=8; social functioning BGD: 5.5 (95% CI 0.2-10.8), MID=7; role functioning BGD: 7.6 (95% CI 2.1-13.1), MID=4.
“The BGDs observed in our study for fatigue, social functioning and role functioning were larger than the published MIDs.”
Discussion ¶5Find in source - supportedReviewer 2The exercise program was safe and well-tolerated in patients with MBC, including those with stable bone metastases.The low number of SAEs and high attendance/compliance rates support this claim.Evidence: Two SAEs (fractures) not related to bone metastases; attendance rate 77%; compliance rates 59-100%.
“In the current study, we found statistically significant effects of exercise on several HRQOL outcomes, including fatigue.”
Discussion ¶5Find in source - supportedReviewer 2The observed effects are clinically relevant, as several BGDs exceeded published minimally important differences (MIDs).The paper explicitly states that BGDs for fatigue, social functioning, and role functioning were larger than the MIDs.Evidence: BGDs for fatigue, social functioning, and role functioning exceeded MIDs.
“The BGDs observed in our study for fatigue, social functioning and role functioning were larger than the published MIDs.”
Discussion ¶6Find in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcomes are patient-reported measures of physical fatigue (EORTC QLQ-FA12) and health-related quality of life (EORTC QLQ-C30 summary score). These are subjective, patient-reported outcomes that serve as surrogates for clinical benefit. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking these specific surrogate measures to hard clinical outcomes such as survival or disease progression. Although MIDs are mentioned for some secondary outcomes, they are not available for the primary outcomes, and the link to clinical meaningfulness is not established.
“Exercise resulted in significant positive effects on both primary outcomes. Physical fatigue was significantly lower (−5.3 (95% confidence interval (CI), −10.0 to −0.6), Bonferroni–Holm-adjusted P = 0.027; Cohen's effect size, 0.22) and HRQOL significantly higher (4.8 (95% CI, 2.2–7.4), Bonferroni–Holm-adjusted P = 0.0003; effect size, 0.33) in the exercise group than in the control group at 6 months.”
- INADEQUATEEffect sizeThe reported effect sizes are small (Cohen's d = 0.22 for physical fatigue and 0.33 for HRQOL). The between-group differences are 5.3 points on a 0-100 scale for fatigue and 4.8 points for HRQOL. The paper does not provide a minimal clinically important difference (MID) for the primary outcomes, and the effect sizes are below the commonly accepted threshold for a small effect (0.2-0.5). The clinical meaningfulness is not explicitly anchored, as the MIDs are only available for secondary outcomes, not the primary ones.
“Physical fatigue was significantly lower (−5.3 (95% confidence interval (CI), −10.0 to −0.6), Bonferroni–Holm-adjusted P = 0.027; Cohen's effect size, 0.22) and HRQOL significantly higher (4.8 (95% CI, 2.2–7.4), Bonferroni–Holm-adjusted P = 0.0003; effect size, 0.33) in the exercise group than in the control group at 6 months.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on exercise in cancer survivors, acknowledges the scarcity of evidence in metastatic breast cancer, and justifies the need for a high-quality trial. The rationale links the premise to the study objectives, and limitations of prior research are addressed by designing an adequately powered trial.
“Hence, a high-quality and adequately powered study is needed to assess the effects of exercise in patients with MBC.”
“However, the guidelines acknowledge that, in the context of MBC, the evidence of the effects of exercise is scarce, and no recommendations can be provided.”
“Hence, a high-quality and adequately powered study is needed to assess the effects of exercise in patients with MBC.”
“These were primarily feasibility studies and were not powered to detect statistically significant differences in fatigue or HRQOL between groups.”
Randomization was centralized with a blocked computer-generated sequence, stratified by center and therapy line. Blinding was not possible due to the nature of the intervention, but this is acknowledged. A power analysis was performed with a target sample size of 350. Inclusion/exclusion criteria were pre-specified, and the analysis population (ITT) and missing data handling (mixed models, multiple imputation) are defined.
“Owing to the nature of the intervention, participants, local clinicians and study nurses, and investigators were not blinded to group assignment after randomization.”
“Randomization was performed centrally using a blocked computer-generated sequence and was stratified by study center and therapy line (first-line or second-line vs. third-line treatment or a later line of treatment).”
“Owing to the nature of the intervention, participants, local clinicians and study nurses, and investigators were not blinded to group assignment after randomization.”
“With n = 139 patients per group ( n = 278 in total), a mean standardized ES of at least 0.35 could be detected with a power of at least 78% or 82% at a nominal two-sided significance level of 2.5% for each outcome separately”
Sex is reported (99.4% female), age is reported (mean 55.4 years), and extensive demographics are provided in Table 1. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing are not applicable for a human trial.
“The mean age of the participants was 55.4 years (s.d., 11.1), and the majority were female (99.4%)”
“The mean age of the participants was 55.4 years (s.d., 11.1), and the majority were female (99.4%)”
The study was approved by the institutional review board of the University Medical Center Utrecht and local boards. Written informed consent was obtained from all patients. Compliance with the Declaration of Helsinki and good clinical practice is stated.
“The study was approved by the institutional review board of the University Medical Center Utrecht, the Netherlands (19-524/M), and by the local ethical review boards of all participating institutions.”
“All patients provided written informed consent before enrollment.”
“The study was conducted in accordance with standards of good clinical practice and the Declaration of Helsinki.”
“The study was approved by the institutional review board of the University Medical Center Utrecht, the Netherlands (19-524/M), and by the local ethical review boards of all participating institutions.”
“All patients provided written informed consent before enrollment.”
“The study was conducted in accordance with standards of good clinical practice and the Declaration of Helsinki.”
The exercise program is described in detail, including components, intensity, and supervision. The activity tracker (Fitbit Inspire HR) and exercise app are named. Statistical software (R v4.2.2) is identified. No antibodies, cell lines, or organisms are used, so those criteria are not applicable.
“The multimodal exercise program consisted of resistance, aerobic and balance exercises (Extended Data Table ).”
“All statistical analyses were performed using R v4.2.2.”
“The multimodal exercise program consisted of resistance, aerobic and balance exercises”
“All statistical analyses were performed using R v4.2.2.”
The primary analysis uses mixed models for repeated measures, adjusted for baseline and stratification factors. Effect sizes and 95% CIs are reported for all outcomes. P-values are reported for primary outcomes with Bonferroni-Holm adjustment. Software is identified. Data presentation includes per-group n and confidence intervals. Mathematical plausibility checks were not applicable due to continuous outcomes and large N.
“For the primary outcomes, linear mixed-effects models were used to assess exercise effects on physical fatigue and HRQOL separately while taking the hierarchical structure of the data into account.”
“The exercise group reported significantly better HRQOL than the control group at 6 months (between-group difference (BGD), 4.8 (95% CI, 2.2–7.4); Bonferroni–Holm-adjusted P = 0.0003; effect size (ES), 0.33).”
“Modeling assumptions were examined and met.”
“For the primary outcomes, linear mixed-effects models were used to assess exercise effects on physical fatigue and HRQOL separately”
“between-group difference (BGD), 4.8 (95% CI, 2.2–7.4); Bonferroni–Holm-adjusted P = 0.0003; effect size (ES), 0.33”
“All statistical analyses were performed using R v4.2.2.”
The data availability statement provides a concrete route: pseudonymized data will be made available through the Digital Research Environment after proposal approval, with a data access agreement and processing within 6 weeks. The statistical code is available on GitHub. Repository deposit and accession numbers are not applicable for patient-level data.
“Pseudonymized data (including data dictionaries) will be made available through the Digital Research Environment, which is a trusted digital research environment that can be accessed at https://mydre.org . This will be carried out after the review and approval of a methodologically sound proposal by the General Assembly of PREFERABLE, with a signed data access agreement, which is in line with Ethics Committee requirements (The Ethics Committee of University Medical Center Utrecht, The Netherlands). Requests will be processed within 6 weeks.”
“The statistical coding is available on GitHub at https://github.com/AnoukHiensch/PREFERABLE-EFFECT.git .”
“Pseudonymized data (including data dictionaries) will be made available through the Digital Research Environment, which is a trusted digital research environment that can be accessed at https://mydre.org .”
“The statistical coding is available on GitHub at https://github.com/AnoukHiensch/PREFERABLE-EFFECT.git .”
The trial is registered (NCT04120298). A reporting summary is mentioned. All outcomes are reported, including non-significant ones. Limitations are discussed in detail. Conclusions are proportional to the evidence. Funding and COI are stated.
“ClinicalTrials.gov Identifier: NCT04120298 (https://clinicaltrials.gov/ct2/show/NCT04120298) .”
“We were also unable to blind participants to their respective study arm. This may have motivated patients in the control arm to voluntarily increase their physical activity levels, especially as all patients had received general advice on physical activity and a fitness tracker.”
“This study received funding from the European Union’s Horizon 2020 research and innovation program (no. 825677; main applicant, A.M.M.) and the National Health and Medical Research Council of Australia (2018/GNT1170698; main applicant, E.M.Z.).”
“ClinicalTrials.gov Identifier: NCT04120298”
“We were also unable to blind participants to their respective study arm.”
“This study received funding from the European Union’s Horizon 2020 research and innovation program (no. 825677; main applicant, A.M.M.) and the National Health and Medical Research Council of Australia (2018/GNT1170698; main applicant, E.M.Z.).”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 48 references by DOI: 44 verified — 4 DOI unresolved.
- UNRESOLVED10.1056/nejmoa2202802Pembrolizumab plus chemotherapy in advanced triple-negative breast cancerCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1080/09593985.2017.1422160Physical activity and advanced cancer: the views of chartered physiotherapists in IrelandCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1007/s11845-017-1666-3Physical activity and advanced cancer: the views of oncology and palliative care physicians in IrelandCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1007/s11764-015-0433-6Validation of the Godin-Shephard leisure-time physical activity questionnaire classification coding system using accelerometer assessment among breast cancer survivorsCited DOI does not resolve to any Crossref record.
3 data/code links checked; 3 live.
- datahttps://mydre.orgLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT04120298LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/AnoukHiensch/PREFERABLE-EFFECT.gitResolves to GitHub (code repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly typo, consistency, grammar.
- MINORtypoAbstract“Bonferroni–Holm-adjusted P = 0.027; Cohen's effect size, 0.22”→ Consider using 'Cohen's d' for consistency.Minor terminology consistency.
- MINORconsistencyTable 1“Higher education | 37 (20.8) | 43 (42.0)”→ Check if the percentage for control group should be 24.0% instead of 42.0%.Potential typo in percentage.
- MINORgrammarMethods, Statistical analysis“Modeling assumptions were examined and met.”→ Consider specifying which assumptions were examined.Could be more specific.
- MINORtypoAuthor list“Hiensch Anouk E Depenbusch Johanna Schmidt Martina E Monninkhof Evelyn M Pelaez Mireia Clauss Dorothea Gunasekara Nadira Zimmer Philipp Belloso Jon Trevaskis Mark Rundqvist Helene Wiskemann Joachim Müller Jana Sweegers Maike G Fremd Carlo Altena Renske Gorecki Maciej Bijlsma Rhodé van Leeuwen-Snoeks Lobke ten Bokkel Huinink Daan Sonke Gabe Lahuerta Ainhara Mann G Bruce Francis Prudence A Richardson Gary Malter Wolfram van der Wall Elsken Aaronson Neil K Senkus Elzbieta Urruticoechea Ander Zopf Eva M Bloch Wilhelm Stuiver Martijn M Wengstrom Yvonne Steindorf Karen May Anne M”→ Format author names with proper punctuation and affiliations.Author list appears as a continuous string without separators.
- MINORconsistencyTable 1“Higher education | 37 (20.8) | 43 (42.0)”→ Check the percentage for control group higher education; 43/179 = 24.0%, not 42.0%.Potential typo in percentage.
The published work is robust and well-reported. An informed reader should weigh the minor Table 1 percentage inconsistency and the four references not found in registries, which may warrant verification or correction. No substantive methodological flaws were identified.
- 1.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 42% does not match the reported count 43/179Demonstrable critical failure — blocks the verdict from passing.
- 2.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 24.6% does not match the reported count 46/179Demonstrable critical failure — blocks the verdict from passing.
- 3.HIGHreportingVerify and correct the percentage for 'Higher education' in the control group in Table 1 (currently 42.0%, should be 24.0% based on 43/179).The percentage is internally inconsistent with the reported count, which could mislead readers and undermine data credibility.
- 4.HIGHreportingVerify the four references not found in registries (Pembrolizumab plus chemotherapy..., Physical activity and advanced cancer: the views of chartered physiotherapists..., Physical activity and advanced cancer: the views of oncology and palliative care physicians..., Validation of the Godin-Shephard...) and correct or replace them if they are erroneous or fabricated.References that cannot be located in any registry may be fabricated or contain errors, which is a serious integrity concern.
- 5.MEDIUMcopyeditFormat the author list with proper punctuation and affiliations instead of a continuous string.The current formatting is unprofessional and may cause indexing or citation errors.
- 6.MEDIUMreportingSpecify which modeling assumptions were examined in the statistical analysis section (e.g., normality, homoscedasticity).The statement 'Modeling assumptions were examined and met' is vague; specifying them enhances transparency and reproducibility.
- 7.MEDIUMreportingAdd a statement in the main text (or supplement) that the study followed CONSORT reporting guidelines.Explicitly stating adherence to reporting guidelines improves transparency and reader confidence.
- 8.MEDIUMreportingReport exact p-values for secondary outcomes (or at least provide them in a supplement) to facilitate meta-analyses.Providing exact p-values, even if not adjusted, allows readers to interpret the strength of evidence and aids future meta-analyses.
- 9.MEDIUMreportingProvide a more detailed description of the exercise app and its development to enhance replicability.The exercise app is a key component of the intervention; more detail would allow others to replicate the intervention accurately.
- 10.MEDIUMreportingProvide a more detailed breakdown of reasons for dropout at each time point.Detailed dropout reasons improve transparency and help readers assess potential attrition bias.
- 11.LOWreportingInclude the full statistical analysis plan as a supplementary file.A full SAP enhances transparency and allows readers to verify that the analysis was pre-specified.
- 12.LOWreportingReport the number of participants who declined to participate and their reasons to address potential selection bias.This information helps readers assess the generalizability of the findings.
- 13.LOWcopyeditUse 'Cohen's d' consistently instead of 'Cohen's effect size' in the abstract.Consistent terminology improves clarity and professionalism.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.