Anticoagulation to prevent ischemic stroke and neurocognitive impairment in atrial fibrillation: the BRAIN-AF randomized clinical trial.
Rivard L, Khairy P, Talajic M, Tardif JC, Healey JS, Black SE, Andrade JG, Field TS, Nault I, Bherer L, Massoud F, Nattel S, Lanthier S, Racine N, Roux JF, Greiss I, Macle L, Guerra PG, Tadros R, Mayrand H, Gosselin G, Conen D, Bocti C, Chayer C, Deschaintre Y, Sandhu RK, Manlucu J, Khaykin Y, Verma A, Mondésert B, Dyrda K, Cadrin-Tourigny J, Thibault B, Raymond-Paquin A, Aguilar M, Brouillette J, Roussin A, Robillard A, Tremblay-Gravel M, David LP, Cossette M, Parkash R, Guertin MC, Roy D, BRAIN-AF investigators
- DOI
- 10.1038/s41591-025-04101-y
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/fae7681b-034c-46e7-aa8b-bc259d78934f is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ReportingData & code availability partially met−0.25★
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The BRAIN-AF trial is a rigorously designed and transparently reported randomized controlled trial: randomization, blinding, power analysis, ethics approval, demographic reporting, and statistical reporting all meet high standards, and the controlled-access data statement is appropriate for identifiable patient-level data. The primary outcomes are reported with named tests, exact p-values, and effect sizes with CIs, and the statistics verification recomputed all checkable values consistently. The main weaknesses are minor: analysis code is not shared, the reporting checklist is not explicitly named, and two small internal inconsistencies (the primary-analysis denominator count and the stroke event count) are left unexplained.
Both reviewers sampled the same model independently and converged on 7 of 8 dimensions (all pass); they diverged only on data code availability, where I weighed toward warn because the trial's custom SAS analysis code is an applicable, unreported artifact. Statistics coverage is partial (only 3 tests with test-statistic+df or effect+CI were recomputed; threshold-only and resampling/exact p-values were not machine-verified), and the citation check found no retracted or unlocatable references.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks.
- CONSISTENTreported p = .460 · recomputed p = .443Reviewer 1Verify the reported p-value for the primary efficacy outcome (HR=1.10, 95% CI 0.86-1.40) using the confidence interval method.
“hazard ratio of 1.10 (95% confidence interval (CI) 0.86–1.40; P = 0.46 by generalized log-rank test)”
Taken as given: The 95% CI is two-sided at the 0.05 level.; The HR is a log-hazard ratio, so log=1 is appropriate.; The CI is not adjusted for multiple comparisons.Method: pCI is used to compute the two-tailed p-value from the point estimate and 95% confidence interval for a log-scale ratio.How we recomputed it: pCI(1.10, 0.86, 1.40, 1) - CONSISTENTreported p = .270 · recomputed p = .286Reviewers 1, 2Verify the reported p-value for major bleeding (HR=0.41, 95% CI 0.08-2.12) using the confidence interval method.
“Major bleeding—no. (%) 2 (0.3) 5 (0.8) 0.41 (0.08, 2.12) 0.27”
Taken as given: The 95% CI is two-sided at the 0.05 level.; The HR is a log-hazard ratio, so log=1 is appropriate.; The CI is not adjusted for multiple comparisons.Method: pCI is used to compute the two-tailed p-value from the point estimate and 95% confidence interval for a log-scale ratio.How we recomputed it: pCI(0.41, 0.08, 2.12, 1) - CONSISTENTreported p = .460 · recomputed p = .443Reviewer 2Primary composite outcome: hazard ratio 1.10 (95% CI 0.86–1.40), P = 0.46
“yielding a hazard ratio of 1.10 (95% confidence interval (CI) 0.86–1.40; P = 0.46 by generalized log-rank test”
Taken as given: the HR 1.10 and 95% CI 0.86–1.40 are the two-sided estimates for the primary composite outcome; the HR is approximately log-normally distributed; the reported p (0.46) was computed from the interval-censored proportional hazards model, so the CI-derived two-tailed p is an approximationMethod: Two-tailed p derived from the hazard ratio and its 95% confidence interval on the log scale (pCI with log=1).How we recomputed it: pCI(1.10, 0.86, 1.40, 1)
- lowinternal contradictionThe text reports 13 strokes among the 256 primary-outcome events, whereas Table 2 lists 9 + 7 = 16 stroke events; the discrepancy is partly explained by the table footnote that some participants experienced more than one component, but the first-event classification (13) versus all-events count (16) is not reconciled in the text.
stroke for 13 (5.1%) and TIA for 9 (3.5%) ... | Stroke | 9 (1.5) | 7 (1.1) |
Table 2reviewer’s wording - lowinternal contradictionThe primary-efficacy analysis denominators (610 and 623) sum to 1,233, two fewer than the 1,235 participants stated as randomized, with no stated reason for the two exclusions in the provided text.
“the primary outcome occurred in 130 of 610 participants (21.3%), compared with 126 of 623 participants (20.2%)”
ResultsFind in source
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewers 1, 2The results provide the strongest evidence to date in support of current guidelines that do not recommend anticoagulation in patients without conventional risk factors for stroke.The trial's null result supports the guideline position in this low-risk population, but 'strongest evidence to date' is a comparative claim about the broader literature that the paper itself does not substantiate.Evidence: Null primary outcome and low stroke/TIA rate (<0.7% per year).
“This low rate, combined with the lack of efficacy of rivaroxaban in improving cognitive outcomes, provides the strongest evidence to date in support of current guidelines that do not recommend anticoagulation in patients without conventional risk factors for stroke.”
DiscussionFind in source - supportedReviewer 1No benefit of anticoagulation was observed across sensitivity analyses.The primary result (HR 1.10, 95% CI 0.86-1.40, p=0.46) shows no statistically significant benefit, and the sensitivity analyses (censored and on-treatment) are consistent, with overlapping confidence intervals and similar p-values.Evidence: Primary efficacy outcome: HR 1.10 (0.86-1.40), p=0.46; Censored analysis: HR not explicitly given but p=0.44; On-treatment analysis: p=0.41.
“No benefit of anticoagulation was observed across sensitivity analyses (censored and on-treatment analyses) or in any predefined subgroup.”
Figure 2Find in source - supportedReviewer 1The trial was terminated early due to futility.The conditional power analysis showing a 1.2% probability of achieving statistical significance under the original alternative hypothesis, and the DSMB recommendation based on the pre-specified futility boundary, support this claim.Evidence: Conditional power analysis: 1.2% probability of a statistically significant treatment effect if continued.
“The conditional power analysis based on the alternative hypothesis underlying the sample size calculation indicated a 1.2% probability of achieving a statistically significant treatment effect had the trial been continued to its planned completion.”
ResultsFind in source - supportedReviewer 1Despite the high incidence of cognitive decline, the BRAIN-AF trial was stopped early for futility.The paper reports a high incidence of cognitive decline (20.7% primary outcome, 6.4% annual rate in placebo) and the early stopping for futility, which is consistent with the results.Evidence: Primary outcome occurred in 20.7% of participants; annual rate 6.4% in placebo; trial stopped for futility after interim analysis.
“Despite the high incidence of cognitive decline observed among patients with AF and low stroke risk, the BRAIN-AF trial, which tested a low dose of rivaroxaban to prevent stroke, transient ischemic attack and cognitive decline in patients with prior AF, was stopped early due to futility.”
DiscussionFind in source - supportedReviewer 1The lack of rivaroxaban's superiority was not attributable to an insufficient event rate.The observed event rate in the placebo group (6.4% per year) aligned with the assumptions used in the sample size calculation, so the null result is not due to a lower-than-expected event rate.Evidence: Annualized rate in placebo: 6.4%; trial design assumed 6% per year. The primary outcome occurred in 20.7% of participants.
“In the BRAIN-AF trial, the primary outcome occurred in more than 20% of participants, with an annual incidence of 6.4% in the placebo group, aligning with the trial’s sample size and power calculations. The lack of rivaroxaban’s superiority was, therefore, not attributable to an insufficient event rate.”
Discussion ¶3Find in source - supportedReviewer 2Treatment with rivaroxaban showed no benefit in reducing cognitive decline, stroke or transient ischemic attack compared to placebo.The presented primary result (HR 1.10, 95% CI 0.86–1.40, P=0.46) and consistent sensitivity and subgroup analyses directly support the null conclusion.Evidence: HR 1.10 (0.86–1.40), P=0.46; consistent censored (P=0.44) and on-treatment (P=0.41) analyses; no subgroup interaction significant.
“treatment with the anticoagulant rivaroxaban showed no benefit in reducing cognitive decline, stroke or transient ischemic attack when compared to placebo”
AbstractFind in source - supportedReviewer 2The trial was stopped early due to futility with a low conditional probability of success had it continued.The DSMB recommendation, the 256 of 410 events at termination, and the stated 1.2% conditional power directly back the futility claim.Evidence: Conditional power analysis indicated a 1.2% probability of achieving a statistically significant treatment effect if continued.
“Conditional power analysis indicated a 1.2% probability of achieving a statistically significant treatment effect if the trial had been continued to its planned total of 410 events.”
AbstractFind in source - supportedReviewer 2The lack of rivaroxaban superiority was not attributable to an insufficient event rate.The observed placebo event rate (6.4% per year) aligned with the sample-size assumption (6% per year), supporting that the null result was not due to too few events.Evidence: Primary outcome annual rate 6.4% in placebo vs the assumed 6%; 256 events occurred by termination.
The lack of rivaroxaban's superiority was, therefore, not attributable to an insufficient event rate.
Discussionreviewer’s wording - supportedReviewer 2BRAIN-AF is the largest randomized controlled trial designed to determine whether oral anticoagulation could prevent cognitive decline in patients with AF.This is a framing claim about trial scope; the paper presents a 1,235-participant randomized trial designed for this purpose, consistent with the statement.Evidence: 1,235 participants randomized across 53 sites; composite cognitive endpoint.
“BRAIN-AF is the largest randomized controlled trial designed to determine whether oral anticoagulation could prevent cognitive decline in patients with AF.”
DiscussionFind in source
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites multiple observational studies and reviews that establish the link between AF and cognitive decline, and the potential benefit of anticoagulation. It acknowledges the limitations of prior work (observational, not randomized) and explicitly states that the trial aims to address this gap. The hypothesis follows logically from the cited evidence.
“Observational studies suggest that anticoagulation may reduce the risk of cognitive decline in patients with AF and elevated thromboembolic risk, implicating subclinical cerebral emboli as a potential mechanistic link.”
“Observational studies and real-world healthcare data further suggest that anticoagulation may reduce cognitive decline and dementia risk in AF patients at high stroke risk”
“We, therefore, conducted the BRAIN-AF trial to determine whether a low dose of rivaroxaban (15 mg once daily) reduces a composite endpoint of cognitive decline, stroke or transient ischemic attack (TIA) compared to placebo in participants with AF who have no established indication for thromboprophylaxis”
Randomization was performed via an interactive web response system using permuted blocks, with double-blinding (patients and investigators). An a priori sample size calculation was provided (1,424 participants, 85% power, alpha=0.05). Detailed inclusion and exclusion criteria are listed. Pre-specified outlier handling (ITT, on-treatment, censored analyses) is described. Independent replication is not applicable for a single pivotal trial.
“Based on these assumptions, a sample size of 1,424 patients was required to detect a relative risk reduction of 30% with a power of 85% and a two-sided significance level of 0.05.”
“Efficacy outcomes were analyzed based on the original randomization groups for all participants, including those who discontinued the trial early or switched to open-label anticoagulant therapy after developing a new clinical indication for oral anticoagulation post-randomization.”
“randomized in a 1:1 ratio to receive either rivaroxaban 15 mg once daily orally or a matching placebo through an interactive web response system (IWRS) using a permuted block randomization method”
“a sample size of 1,424 patients was required to detect a relative risk reduction of 30% with a power of 85% and a two-sided significance level of 0.05”
“Patients and investigators remained blinded to treatment assignments until data analysis”
Sex is reported (25.6% female), and both sexes are included, so sex justification is not applicable. Age, weight (BMI), and health status (e.g., CHA2DS2-VA score, LVEF, creatinine clearance) are reported in Table 1. Demographics include age, sex, race, education, depression, alcohol use, etc. The species/strain and housing criteria are not applicable for a human study.
“Female sex—no. (%) | 316 (25.6) | 163 (26.7) | 153 (24.5)”
“Race or ethnic group—no. (%) a | | Caucasian | 1,181 (95.6)”
The study was approved by the institutional review board of the Montreal Heart Institute (MP-33-2014-1559) and each participating center. All patients provided written informed consent. Compliance with the Declaration of Helsinki is stated.
“The study was authorized by Health Canada and the institutional review boards of the Montreal Heart Institute (MP-33-2014-1559) and each participating center”
“All patients provided written, informed consent before any trial activities.”
“conducted in accordance with the principles of the Declaration of Helsinki.”
“the study was authorized by Health Canada and the institutional review boards of the Montreal Heart Institute (MP-33-2014-1559) and each participating center and was conducted in accordance with the principles of the Declaration of Helsinki”
“All patients provided written informed consent to participate.”
Rivaroxaban 15 mg daily is the investigational product, with in-kind support from Bayer Inc. Placebo is matching but not further described, which is acceptable. Software (SAS version 9.4) is identified. Antibodies, cell lines, organisms, and mycoplasma testing are not applicable.
“rivaroxaban 15 mg once daily orally”
“All statistical analyses were conducted using SAS software version 9.4.”
“rivaroxaban 15 mg once daily orally or a matching placebo”
“All statistical analyses were conducted using SAS software version 9.4.”
The primary analysis uses a generalized log-rank test and proportional hazards model for interval-censored data, both clearly named. Exact p-values (e.g., P=0.46) and 95% CIs are reported. SAS 9.4 is identified. Data presentation includes Kaplan-Meier curves, forest plots, and per-group event counts. Assumptions are not explicitly tested but handled by standard methods. Mathematical plausibility checks revealed no inconsistencies; percentages sum to ~100% with rounding noted.
“All statistical analyses were conducted using SAS software version 9.4.”
“yielding a hazard ratio of 1.10 (95% confidence interval (CI) 0.86–1.40; P = 0.46 by generalized log-rank test”
The data availability statement describes a concrete route for controlled access to deidentified individual-level data, which is adequate for patient-level data. However, no code repository is provided for the statistical analysis code, which limits reproducibility. Repository deposit and accession numbers are not applicable for the patient-level data.
“Qualified academic investigators may request controlled on-site access to analyze these data after completion of all prespecified primary and secondary analyses and beginning 12 months after publication, for a period of 36 months.”
Methods are thorough and replicable. The trial is registered at ClinicalTrials.gov (NCT02387229). A reporting summary is linked. All pre-specified outcomes are reported, including negative and null results. Limitations are discussed extensively. Conclusions are proportional to the evidence. Funding sources and competing interests are fully disclosed.
“ClinicalTrials.gov registration: NCT02387229 (https://clinicaltrials.gov/study/NCT02387229)”
“This trial was funded by the Canadian Institutes of Health Research (CIHR, grant no. PJT-175207, L.R.), Canadian Stroke Prevention Network (CSPIN, L.R.), Montreal Heart Institute Foundation (L.R.) and BAYER Inc”
“ClinicalTrials.gov registration: NCT02387229”
“Further information on research design is available in the linked to this article.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 50 references by DOI: 49 verified — 1 no DOI (shown, not verified).
- NO DOIBeck Depression Inventory: ManualNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://clinicaltrials.gov/study/NCT02387229LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT02387229LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
3 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 3 minor suggestions below.
3 copyedit issues flagged: mostly clarity, consistency.
- MINORclarityTable 1“Body mass index—m kg − 2”→ Use standard units 'kg/m²' for BMI.Likely a conversion artifact; the value column is correct but the unit string reads as 'm kg − 2'.
- MINORclarityTable 1“Left atrial volume—ml m − 2”→ Use standard units 'ml/m²' for the indexed left atrial volume.Same unit-formatting artifact as BMI row.
- MINORconsistencyResults, Efficacy outcomes“the primary outcome occurred in 130 of 610 participants (21.3%), compared with 126 of 623 participants (20.2%)”→ Explain why the primary-analysis denominators (610 and 623) total 1,233 rather than the 1,235 randomized (e.g., two participants without post-baseline data).Appears in abstract-derived text and Results; not contradicted elsewhere but left unexplained.
This is a robust published trial: an informed reader can rely on its core rigor dimensions, and no evidence of fabrication, retracted citations, or invalid statistics was found. The two low-severity internal inconsistencies (primary-analysis denominators summing to 1,233 vs. 1,235 randomized, and text reporting 13 strokes vs. Table 2 listing 16 stroke events) warrant a correction/clarification from the authors and should be weighed when interpreting the primary analysis, as should the absence of shared analysis code.
- 1.HIGHrigorClarify in the Results/figure-1 flow diagram why the primary-efficacy denominators (610 and 623, summing to 1,233) differ from the 1,235 randomized participants, e.g., state that two participants were excluded for lack of post-baseline cognitive assessment.The unexplained two-participant discrepancy was independently flagged by a reviewer, the copyedit pass, and the integrity check, so a reader cannot reconcile the reported analysis population without a stated reason.
- 2.HIGHrigorReconcile the stroke event count discrepancy between the text (13 strokes among the 256 primary-outcome events) and Table 2 (9 + 7 = 16 stroke events), clarifying the first-event classification versus the all-events count.The integrity check flagged this internal contradiction; although partly explained by the table footnote that some participants had more than one component, the text does not reconcile the 13 vs. 16 figures.
- 3.MEDIUMreportingName the CONSORT 2010 reporting checklist explicitly (with completion status) in the Reporting summary section instead of only referencing the generic 'Reporting summary'.Reviewer 2 flagged that the reporting guideline is referenced only generically, and naming the checklist is the expected standard for a randomized trial.
- 4.MEDIUMdata codeDeposit the custom SAS 9.4 analysis code in a public repository (e.g., GitHub) or describe it in the Methods to complement the controlled-access data statement.The trial's SAS analyses are an applicable artifact whose absence limits reproducibility; this was the sole basis for the warn on data_code_availability.
- 5.MEDIUMcopyeditFix the BMI unit string in Table 1 from 'm kg − 2' to 'kg/m²'.The copyedit pass flagged this as a likely conversion artifact that renders the unit unreadable.
- 6.MEDIUMcopyeditFix the left atrial volume unit string in Table 1 from 'ml m − 2' to 'ml/m²'.The copyedit pass flagged the same unit-formatting artifact as the BMI row.
- 7.LOWdata codeDeposit non-identifiable aggregate data (baseline characteristics, event rates) in a public repository such as Dryad or Zenodo to complement the controlled-access statement.Reviewer 1 suggested this as a low-cost way to broaden access beyond the on-site SERIANT mechanism for readers who only need aggregate results.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.