Effectiveness of shared decision making strategies for stroke prevention among patients with atrial fibrillation: cluster randomized controlled trial.
Ozanne EM, Barnes GD, Brito JP, Cameron KA, Cavanaugh KL, Greene T, Jackson EA, Montori VM, Steinberg BA, Witt DM, Noseworthy P, Passman RS, Kansal P, Crossley G, Roden DM, Christensen JT, Ariotti A, Jones AE, Bardsley T, Wu C, Fagerlin A, STEP-UP Writing Group
- DOI
- 10.1136/bmj-2024-079976
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/26faa9b0-6b01-477c-b4b0-656ac65de8aa is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcomes are surrogate measures of shared decision making (OPTION12 score, knowledge, decisional conflict) rather than hard clinical outcomes. The paper does not demonstrate target engagement at the tested dose (not applicable for a behavioral intervention) nor cite validated evidence linking these surrogates to clinical outcomes such as stroke prevention or patient-important outcomes. The claim of effectiveness rests on these surrogates without establishing a validated link to clinical benefit.
“Primary outcome measures were quality of shared decision making measured by OPTION12, knowledge of atrial fibrillation and its management, and decisional conflict.”
- 02Treatment effect not shown to be clinically meaningful
The reported effect sizes are small relative to the scale ranges and lack anchoring to clinically meaningful thresholds. For example, the OPTION12 score improved by 12.1 points on a 0-100 scale, knowledge odds ratio 1.68, and decisional conflict reduced by 6.3 points on a 0-100 scale. The paper does not provide a minimal clinically important difference or other anchor to establish that these changes are clinically material. The discussion mentions a threshold for decisional conflict (<25) but does not use it to interpret the effect size.
“Compared with usual care, the combined use of both the patient decision aid and the encounter decision aid improved the quality of shared decision making (adjusted mean difference 12.1 (95% confidence interval (CI) 8.0 to 16.2; P<0.001), improved patients’…”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted cluster randomized trial with clear randomization, sample size justification, and transparent reporting. The main weaknesses are a vague data availability statement, lack of statistical software identification, and minor copyedit issues. Overall, the paper is methodologically sound and the findings are credible.
Both reviewers classified the study as interventional, which is appropriate for a cluster RCT. The key resources dimension was rated not applicable by one reviewer and pass by the other; I adopted 'pass' because the decision aids are investigational products. The statistics verification checked 6 tests, all consistent, but coverage is limited to tests with test statistics or effect estimates with CIs; other p-values were not machine-verified.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 6 tests: 6 consistent, 0 inconsistent; 1 recomputed directly from the reported test statistics, 5 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Recomputed odds ratio 1.68 (95% CI 1.35–2.09), reported p<0.001
“odds ratio 1.68 (95% CI 1.35 to 2.09; P<0.001”
Taken as given: 1.35–2.09 is a two-sided 95% confidence interval for the odds ratio of 1.68, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p<0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.68, 1.35, 2.09, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary comparison for OPTION12: estimated difference 12.1, 95% CI 8.0 to 16.2, p<0.001
“adjusted mean difference 12.1 (95% confidence interval (CI) 8.0 to 16.2; P<0.001)”
Taken as given: The estimate is a mean difference.; The CI is two-sided at 95%.Method: Recomputed p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(12.1, 8.0, 16.2, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary comparison for decisional conflict: estimated difference -6.3, 95% CI -9.6 to -3.1, p<0.001
“adjusted mean difference −6.3 (95% CI −9.6 to −3.1; P<0.001)”
Taken as given: The estimate is a mean difference.; The CI is two-sided at 95%.Method: Recomputed p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-6.3, -9.6, -3.1, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Primary comparison for OPTION12: EDA and PDA vs control
“estimated difference 12.1, 95% confidence interval (CI) 8.0 to 16.2; P<0.001”
Taken as given: The estimate is a mean difference with a 95% CI.; The CI is two-sided.Method: Recomputed p-value from the estimate and 95% CI using normal approximation.How we recomputed it: pCI(12.1, 8.0, 16.2, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Primary comparison for knowledge: EDA and PDA vs control
“adjusted odds ratio comparing the proportion of correct responses on the knowledge score 1.68, 95% CI 1.35 to 2.09; P<0.001”
Taken as given: The estimate is an odds ratio with a 95% CI.; The CI is two-sided.Method: Recomputed p-value from the log odds ratio and its 95% CI using normal approximation.How we recomputed it: pCI(1.68, 1.35, 2.09, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Primary comparison for decisional conflict: EDA and PDA vs control
“estimated difference −6.3, 95% CI −9.6 to −3.1, P<0.001”
Taken as given: The estimate is a mean difference with a 95% CI.; The CI is two-sided.Method: Recomputed p-value from the estimate and 95% CI using normal approximation.How we recomputed it: pCI(-6.3, -9.6, -3.1, 0)
- lowinternal contradictionThe abstract states '1117 participants across six sites were included in the analysis' but the results section reports 1214 enrolled and 97 excluded, which sums to 1117, consistent. No contradiction found.
“1117 participants across six sites were included in the analysis.”
AbstractFind in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
11 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewer 1Combined use of both decision aids improves quality of shared decision making, knowledge, and reduces decisional conflict compared to usual care.The primary analysis shows statistically significant improvements for all three outcomes.Evidence: Adjusted mean difference 12.1 (95% CI 8.0 to 16.2; P<0.001) for OPTION12; odds ratio 1.68 (95% CI 1.35 to 2.09; P<0.001) for knowledge; adjusted mean difference −6.3 (95% CI −9.6 to −3.1; P<0.001) for decisional conflict.
“Compared with usual care, the combined use of both the patient decision aid and the encounter decision aid improved the quality of shared decision making (adjusted mean difference 12.1 (95% confidence interval (CI) 8.0 to 16.2; P<0.001), improved patients’ knowledge (odds ratio 1.68 (95% CI 1.35 to 2.09; P<0.001), and reduced patients’ decisional conflict (adjusted mean difference −6.3 (95% CI −9.6 to −3.1; P<0.001).”
AbstractFind in source - supportedReviewer 1Encounter decision aid alone improves all three outcomes compared to usual care.Secondary analysis shows significant improvements for all three outcomes.Evidence: OPTION12 estimated difference 12.9 (8.6 to 17.1; P<0.001); knowledge OR 1.41 (1.11 to 1.79; P=0.003); decisional conflict −5.8 (−9.3 to −2.4; P<0.001).
EDA v control 12.9 (8.6 to 17.1) ... P<0.001 ... 1.41 (1.11 to 1.79) ... 0.003 ... −5.8 (−9.3 to −2.4) ... <0.001
Table 3reviewer’s wording - supportedReviewer 1Patient decision aid alone improves shared decision making and knowledge but not decisional conflict.Secondary analysis shows significant improvements for OPTION12 and knowledge, but not for decisional conflict.Evidence: OPTION12 estimated difference 3.8 (1.1 to 6.4; P<0.001); knowledge OR 1.68 (1.24 to 2.28; P<0.001); decisional conflict −2.6 (−6.8 to 1.6; P=0.11).
PDA v control 3.8 (1.1 to 6.4) ... <0.001 ... 1.68 (1.24 to 2.28) ... <0.001 ... −2.6 (−6.8 to 1.6) ... 0.11
Table 3reviewer’s wording - supportedReviewer 1No important differences were observed in treatment choices or satisfaction.Secondary outcomes show no significant differences across groups.Evidence: Table 5 shows no significant differences in agreement, clinician recommendation, or satisfaction.
“No statistically significant differences when comparing clinician-patient agreement on treatment choice by group”
Table 5Find in source - supportedReviewer 1The study is the largest randomized controlled trial of decision aids and the first to directly compare complementary patient and encounter decision aids.The paper claims this and no contradicting evidence is presented.Evidence: Statement in Discussion.
“To our knowledge, this is the largest randomized controlled trial of decision aids and the first to directly compare the use of a complementary patient decision aid and encounter decision aid.”
DiscussionFind in source - supportedReviewer 2The combined use of both the patient decision aid and the encounter decision aid improved the quality of shared decision making.The primary comparison shows a statistically significant improvement in OPTION12 score.Evidence: Adjusted mean difference 12.1 (95% CI 8.0 to 16.2; P<0.001) for EDA and PDA vs control.
“the combined use of both the patient decision aid and the encounter decision aid improved the quality of shared decision making (adjusted mean difference 12.1 (95% confidence interval (CI) 8.0 to 16.2; P<0.001)”
AbstractFind in source - supportedReviewer 2The combined use of both decision aids improved patients' knowledge.The primary comparison shows a statistically significant improvement in knowledge.Evidence: Odds ratio 1.68 (95% CI 1.35 to 2.09; P<0.001) for EDA and PDA vs control.
“improved patients’ knowledge (odds ratio 1.68 (95% CI 1.35 to 2.09; P<0.001)”
AbstractFind in source - supportedReviewer 2The combined use of both decision aids reduced patients' decisional conflict.The primary comparison shows a statistically significant reduction in decisional conflict.Evidence: Adjusted mean difference −6.3 (95% CI −9.6 to −3.1; P<0.001) for EDA and PDA vs control.
“reduced patients’ decisional conflict (adjusted mean difference −6.3 (95% CI −9.6 to −3.1; P<0.001)”
AbstractFind in source - supportedReviewer 2The encounter decision aid alone versus usual care showed statistically significant improvements for all three outcomes.Secondary comparisons show significant improvements for EDA vs control on all three outcomes.Evidence: Table 3: EDA vs control OPTION12 diff 12.9 (8.6 to 17.1) P<0.001; knowledge OR 1.41 (1.11 to 1.79) P=0.003; decisional conflict diff −5.8 (−9.3 to −2.4) P<0.001.
“Statistically significant improvements were also observed with the encounter decision aid alone versus usual care for all three outcomes”
AbstractFind in source - supportedReviewer 2The patient decision aid alone versus usual care showed statistically significant improvements for quality of shared decision making and knowledge.Secondary comparisons show significant improvements for PDA vs control on OPTION12 and knowledge, but not decisional conflict.Evidence: Table 3: PDA vs control OPTION12 diff 3.8 (1.1 to 6.4) P<0.001; knowledge OR 1.68 (1.24 to 2.28) P<0.001; decisional conflict diff −2.6 (−6.8 to 1.6) P=0.11.
“with the patient decision aid alone versus usual care for quality of shared decision making and knowledge”
AbstractFind in source - supportedReviewer 2No important differences were observed in treatment choices for stroke prevention or in participants' satisfaction.Secondary outcomes show no significant differences in treatment choice or satisfaction.Evidence: Table 5 shows no significant differences in agreement, clinician recommendation, or satisfaction.
“No important differences were observed in treatment choices for stroke prevention or in participants’ satisfaction.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcomes are surrogate measures of shared decision making (OPTION12 score, knowledge, decisional conflict) rather than hard clinical outcomes. The paper does not demonstrate target engagement at the tested dose (not applicable for a behavioral intervention) nor cite validated evidence linking these surrogates to clinical outcomes such as stroke prevention or patient-important outcomes. The claim of effectiveness rests on these surrogates without establishing a validated link to clinical benefit.
“Primary outcome measures were quality of shared decision making measured by OPTION12, knowledge of atrial fibrillation and its management, and decisional conflict.”
- INADEQUATEEffect sizeThe reported effect sizes are small relative to the scale ranges and lack anchoring to clinically meaningful thresholds. For example, the OPTION12 score improved by 12.1 points on a 0-100 scale, knowledge odds ratio 1.68, and decisional conflict reduced by 6.3 points on a 0-100 scale. The paper does not provide a minimal clinically important difference or other anchor to establish that these changes are clinically material. The discussion mentions a threshold for decisional conflict (<25) but does not use it to interpret the effect size.
“Compared with usual care, the combined use of both the patient decision aid and the encounter decision aid improved the quality of shared decision making (adjusted mean difference 12.1 (95% confidence interval (CI) 8.0 to 16.2; P<0.001), improved patients’ knowledge (odds ratio 1.68 (95% CI 1.35 to 2.09; P<0.001), and reduced patients’ decisional conflict (adjusted mean difference −6.3 (95% CI −9.6 to −3.1; P<0.001).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior research on decision aids and identifies a gap in comparative effectiveness. The rationale linking the premise to the study objectives is clear. Limitations of prior research are implicitly addressed by the study design, though not explicitly detailed.
“Evidence has shown the effectiveness of a patient decision aid or encounter decision aid in improving shared decision making outcomes in a clinical setting ; however, no reliable estimate exists of the comparative effectiveness of these different types of decision aids in supporting shared decision making in practice.”
“The aim of this study was to evaluate the effectiveness of a patient decision aid and an encounter decision aid in promoting high quality shared decision making for stroke prevention in the care of patients with non-valvular atrial fibrillation at risk of stroke.”
“Evidence has shown the effectiveness of a patient decision aid or encounter decision aid in improving shared decision making outcomes in a clinical setting ; however, no reliable estimate exists of the comparative effectiveness of these different types of decision aids in supporting shared decision making in practice.”
“The aim of this study was to evaluate the effectiveness of a patient decision aid and an encounter decision aid in promoting high quality shared decision making for stroke prevention in the care of patients with non-valvular atrial fibrillation at risk of stroke.”
Randomization method and unit are clearly described (patients and clinicians randomized, stratified by site and other factors). Blinding is described as not possible for participants, but outcome assessors were not blinded (noted as a limitation). Power analysis is provided with assumptions. Inclusion/exclusion criteria are detailed. Outlier handling is addressed through imputation and intention-to-treat analysis. Controls are the usual care group. Independent replication is not applicable for a single trial.
“Patients were randomly assigned with equal allocation to either the use of the patient decision aid or usual care. Clinicians were randomly assigned with equal allocation to either use the encounter decision aid with all study participants or a usual care arm in which they did not use the encounter decision aid for any study participants.”
“Participants could not be blinded to allocation owing to the nature of the intervention.”
“Under these assumptions, the sample size of 1200 would provide 80% power with a two sided α of 0.05 to detect mean differences, expressed as fractions of one standard deviation, of 0.40, 0.33, and 0.33 for the OPTION12, knowledge, and decisional conflict scores, respectively”
“A study coordinator allocated all participants by using the REDCap randomization function.”
“Participants could not be blinded to allocation owing to the nature of the intervention.”
“Under these assumptions, the sample size of 1200 would provide 80% power with a two sided α of 0.05 to detect mean differences, expressed as fractions of one standard deviation, of 0.40, 0.33, and 0.33 for the OPTION12, knowledge, and decisional conflict scores, respectively”
Sex, age, race/ethnicity, education, and health insurance are reported in Table 1. Age and sex are reported for the cohort. Since both sexes are enrolled, sex justification is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“The mean age of the patients was 69 (standard deviation 9) years”
“The mean age of the patients was 69 (standard deviation 9) years”
The study has ethical approval from a named IRB with a protocol number, and informed consent was obtained from all participants. Regulatory compliance is implied through adherence to ethical standards, though not explicitly named.
“The study has ethical approval from the University of Utah Institutional Review Board (IRB_00124127).”
“All participants provided informed consent before taking part in the study.”
“The study has ethical approval from the University of Utah Institutional Review Board (IRB_00124127).”
“All participants provided informed consent before taking part in the study.”
The decision aids are described in detail, though not with a manufacturer/catalog number as they are custom tools. Statistical software is not explicitly named, but the analysis methods are described. No wet-lab reagents are used.
“The patient decision aid was designed as an interactive, non-linear online tool”
“The patient decision aid includes an explanation of atrial fibrillation and how it affects a patient’s life and features the CHA 2 DS 2 -VASc and HAS-BLED calculators”
Statistical tests are named (linear mixed effects models, generalized linear mixed effects models). Assumptions are handled through model design. Exact p-values are reported for primary outcomes. Effect sizes with confidence intervals are provided. Software is not explicitly identified. Data presentation includes means, SDs, and CIs. Mathematical plausibility checks were not performed due to lack of raw data.
“We did the co-primary statistical analyses for the OPTION12 scale and Decisional Conflict Scale by using separate analyses of linear mixed effects models”
“We did the co-primary statistical analyses for the OPTION12 scale and Decisional Conflict Scale by using separate analyses of linear mixed effects models”
The data availability statement says de-identified data will be posted to ClinicalTrials.gov, but no timeline or accession number is provided. Additional data may be shared on reasonable request, but no mechanism is specified. No code is shared.
“De-identified data will be posted to ClinicalTrials.gov according to our sponsors’ requirements and timeline. Additional data may be shared on reasonable request.”
“De-identified data will be posted to ClinicalTrials.gov according to our sponsors’ requirements and timeline. Additional data may be shared on reasonable request.”
Trial registration is provided. CONSORT diagram is included. All outcomes are reported, including non-significant ones. Limitations are discussed. Conclusions are proportional to evidence. Funding and COI statements are provided.
“Trial registration ClinicalTrials.gov NCT04357288”
“Fig 2 CONSORT (Consolidated Standards of Reporting Trials) diagram.”
“Observers assessing the OPTION12 outcome measure on video recorded encounters cannot be blinded to allocation; their unblinded assessments could be biased in favor of decision aids”
“Trial registration ClinicalTrials.gov NCT04357288”
“Fig 2 CONSORT (Consolidated Standards of Reporting Trials) diagram.”
“Observers assessing the OPTION12 outcome measure on video recorded encounters cannot be blinded to allocation; their unblinded assessments could be biased in favor of decision aids”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 36 references by DOI: 31 verified — 5 no DOI (shown, not verified).
- NO DOIAtrial fibrillation and stroke: epidemiologyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINewly Detected Atrial Fibrillation: AAFP Updates Guideline on Pharmacologic ManagementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDecision aids for people facing health treatment or screening decisionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILinear Mixed Models for Longitudinal DataNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDigital Readiness GapsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT04357288LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORtypoAbstract“Ozanne Elissa M associate professor”→ Format author names consistently, e.g., 'Elissa M Ozanne, associate professor'Author list formatting is inconsistent.
- MINORconsistencyMethods, Randomization and blinding“Participants could not be blinded to allocation owing to the nature of the intervention.”→ Clarify whether outcome assessors were blinded.Blinding of assessors is not explicitly stated.
- MINORclarityResults, Table 1“Positive for limited reading ability 31/291 (11)”→ Define 'limited reading ability' in the table footnote.Term not defined.
- MINORconsistencyAbstract“P<0.001”→ Use consistent formatting for p-values (e.g., P<0.001 vs p<0.001).Inconsistent capitalization of P in p-values.
- MINORclarityMethods, Sample size“The minimum detectable effect sizes were 0.41, 0.34, and 0.34, respectively, for encounter decision aid versus usual care and 0.29, 0.29, and 0.29, respectively, for patient decision aid versus usual care, using the α levels described below.”→ Clarify which outcome each effect size corresponds to.Ambiguous which effect size is for which outcome.
- MINORgrammarIntroduction“as many as 50% of at risk patients with atrial fibrillation who are given a prescription for an oral anticoagulants do not start therapy”→ Change 'an oral anticoagulants' to 'an oral anticoagulant'.Subject-verb agreement error.
The published work is robust and credible, with no major integrity concerns. An informed reader should weigh the minor reporting gaps (vague data availability, missing software identification) and the unblinded outcome assessors as limitations, but these do not undermine the main conclusions. No erratum is warranted based on the checks performed.
- 1.HIGHdata codeIn the Data availability statement, specify the exact timeline for posting de-identified data to ClinicalTrials.gov and the process for requesting additional data (e.g., contact email or data access committee).The current statement is vague and does not meet common reproducibility standards.
- 2.HIGHstatisticsIn the Methods, Statistical analyses section, name the statistical software and version used (e.g., R 4.2, SAS 9.4).Software identification is a standard reporting requirement and was flagged by both reviewers.
- 3.HIGHdata codeShare analysis code in a public repository (e.g., GitHub) and link it in the Data availability statement.Code sharing enhances reproducibility and was recommended by both reviewers.
- 4.MEDIUMreportingIn the Methods, Randomization and blinding section, explicitly state whether outcome assessors were blinded (or not) and any measures taken to mitigate bias.The copyedit pass and reviewers noted that assessor blinding is not clearly described.
- 5.MEDIUMreportingIn the Introduction, add a sentence discussing limitations of prior research that this study addresses.Both reviewers flagged that limitations of prior work are not explicitly discussed.
- 6.MEDIUMcopyeditIn the Abstract, standardize p-value formatting (e.g., use 'P<0.001' consistently) and fix the author name formatting (e.g., 'Elissa M Ozanne, associate professor').Minor copyedit issues affect professionalism and consistency.
- 7.MEDIUMcopyeditIn Table 1, define 'limited reading ability' in a footnote.The term is not defined and could be ambiguous to readers.
- 8.MEDIUMcopyeditIn the Methods, Sample size section, clarify which outcome each minimum detectable effect size corresponds to.The current text is ambiguous about which effect size is for which outcome.
- 9.LOWcopyeditIn the Introduction, fix the grammar error: change 'an oral anticoagulants' to 'an oral anticoagulant'.Subject-verb agreement error.
- 10.LOWethicsIn the Ethics statements, add an explicit statement of adherence to the Declaration of Helsinki or equivalent ethical framework.One reviewer noted regulatory compliance is only implied; explicit statement strengthens the ethics section.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.