Consumer wearable devices for evaluation of heart rate control using digoxin versus beta-blockers: the RATE-AF randomized trial.
Gill SK, Barsky A, Guan X, Bunting KV, Karwath A, Tica O, Stanbury M, Haynes S, Folarin A, Dobson R, Kurps J, Asselbergs FW, Grobbee DE, Camm AJ, Eijkemans MJC, Gkoutos GV, Kotecha D, BigData@Heart Consortium, cardAIc group, RATE-AF trial team
- DOI
- 10.1038/s41591-024-03094-4
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/8279ed23-1ebc-4bfd-9a0c-752809c2f28f is authoritative.
How this rating was calculated
- StatisticsImpossible or misreported statistic ×2−2★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped)−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 22 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Printed percentage does not match its own countdemonstrable
59.3% does not match the reported count 16/28
“16 patients (59.3%) in the digoxin group”
DiscussionFind in source - 02Printed percentage does not match its own countdemonstrable
54.2% does not match the reported count 13/25
“and 13 (54.2%) in the beta-blocker group”
DiscussionFind in source - 03Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is that digoxin and beta-blockers have equivalent effects on heart rate in atrial fibrillation, based on wearable heart rate measurements. Heart rate is a surrogate biomarker for clinical outcomes in AF. The paper does not demonstrate target engagement at the tested dose (e.g., no PK/PD or dose-exposure relationship linking the measured heart rate to drug effect) nor does it cite validated evidence linking heart rate control to hard clinical outcomes in this context. The claim of equivalence is based solely on the surrogate heart rate measurement.
“The results of this study indicate that digoxin and beta-blockers have equivalent effects on heart rate in atrial fibrillation at rest and on exertion”
- 04Treatment effect not shown to be clinically meaningful
The primary reported effect is a null result (no significant difference in heart rate between digoxin and beta-blockers), with a regression coefficient of 1.22 bpm (95% CI -2.82 to 5.27; P=0.55). The paper interprets this as equivalence, but no minimal clinically important difference (MCID) for heart rate in AF is provided, and the confidence interval is wide, spanning a range that could include clinically meaningful differences. The effect size is not anchored to any clinical or biological meaningfulness.
“The unadjusted regression coefficient for digoxin versus beta-blockers was 1.22 (95% CI −2.82 to 5.27; P = 0.55)”
- 05Printed percentage does not match its own count
90.4% is unattainable for n=53 (nearest: 88.7, 90.6%)
“90.4% of participants using this every day for the last 7 days before the interim review”
ResultsFind in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper is a well-conducted randomized controlled trial substudy with clear scientific premise, rigorous design, and transparent reporting. Minor reporting gaps exist in data availability specifics and statistical verification, but overall the methodology is sound.
Both reviewers classified the study as interventional, and this was adopted. The evaluation covered all eight dimensions; several sub-criteria were not applicable (e.g., animal-related items, cell line authentication). The statistics verification component checked only a subset of tests, and the inconsistent findings were not detailed, so they were not treated as demonstrable errors.
Numerical inconsistencies
2 findings · worst criticalValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Printed percentage does not match its own countRecomputed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks. 2 reported summary statistics mathematically impossible for the stated N (PERCENT). 1 printed percentage that does not match its own count.
- PERCENT90.4% is unattainable for n=53 (nearest: 88.7, 90.6%)
“90.4% of participants using this every day for the last 7 days before the interim review”
ResultsFind in source - PERCENT59.3% does not match the reported count 16/28
“16 patients (59.3%) in the digoxin group”
DiscussionFind in source - PERCENT54.2% does not match the reported count 13/25
“and 13 (54.2%) in the beta-blocker group”
DiscussionFind in source
- CONSISTENTreported p = .550 · recomputed p = .554Reviewers 1, 2P-value for the unadjusted regression coefficient for digoxin vs beta-blockers
“The unadjusted regression coefficient for digoxin versus beta-blockers was 1.22 (95% CI −2.82 to 5.27; P = 0.55)”
Taken as given: The estimate is a mean difference (not a ratio).; The 95% CI is two-sided.; The p-value is two-tailed.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(1.22, -2.82, 5.27, 0) - CONSISTENTreported p = .750 · recomputed p = .753Reviewers 1, 2P-value for the adjusted regression coefficient for digoxin vs beta-blockers
“adjusted 0.66 (95% CI −3.45 to 4.77; P = 0.75)”
Taken as given: The estimate is a mean difference (not a ratio).; The 95% CI is two-sided.; The p-value is two-tailed.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(0.66, -3.45, 4.77, 0)
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
4 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Digoxin and beta-blockers have equivalent effects on heart rate in atrial fibrillation at rest and on exertion.The claim is supported by the primary analysis showing no significant difference in heart rate between groups.Evidence: Regression coefficient 1.22 (95% CI −2.82 to 5.27; P = 0.55) and adjusted 0.66 (95% CI −3.45 to 4.77; P = 0.75).
“The results of this study indicate that digoxin and beta-blockers have equivalent effects on heart rate in atrial fibrillation at rest and on exertion”
AbstractFind in source - supportedReviewers 1, 2Wearable device data could predict NYHA functional class similarly to standard clinical measures.The claim is supported by the F1 score comparison showing no significant difference between the wearable CNN and conventional model.Evidence: F1 score 0.56 (95% CI 0.41 to 0.70) versus 0.55 (95% CI 0.41 to 0.68); P = 0.88.
“wearable device data could predict New York Heart Association functional class 5 months after baseline assessment similarly to standard clinical measures”
AbstractFind in source - supportedReviewers 1, 2The wearable neural network was self-training and designed to address missing data.The methods describe a self-supervised CNN that uses missing data as a third channel, supporting the claim.Evidence: Description of the self-supervising CNN and missing data channel in Methods.
“This missing data was neither dropped nor imputed, but used as a third time-series channel alongside heart rate and step count.”
MethodsFind in source - supportedReviewer 2The study provides robust information on the value and limitations of consumer wearables in clinical research.The claim is supported by the study's design and results, including high data volume and discussion of limitations.Evidence: The study collected over 140 million data points and discusses limitations such as missing data and limited diversity.
“Embedded in a randomized trial, the study provides robust information on the value, but also the limitations of using these devices.”
Discussion ¶1Find in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is that digoxin and beta-blockers have equivalent effects on heart rate in atrial fibrillation, based on wearable heart rate measurements. Heart rate is a surrogate biomarker for clinical outcomes in AF. The paper does not demonstrate target engagement at the tested dose (e.g., no PK/PD or dose-exposure relationship linking the measured heart rate to drug effect) nor does it cite validated evidence linking heart rate control to hard clinical outcomes in this context. The claim of equivalence is based solely on the surrogate heart rate measurement.
“The results of this study indicate that digoxin and beta-blockers have equivalent effects on heart rate in atrial fibrillation at rest and on exertion”
- INADEQUATEEffect sizeThe primary reported effect is a null result (no significant difference in heart rate between digoxin and beta-blockers), with a regression coefficient of 1.22 bpm (95% CI -2.82 to 5.27; P=0.55). The paper interprets this as equivalence, but no minimal clinically important difference (MCID) for heart rate in AF is provided, and the confidence interval is wide, spanning a range that could include clinically meaningful differences. The effect size is not anchored to any clinical or biological meaningfulness.
“The unadjusted regression coefficient for digoxin versus beta-blockers was 1.22 (95% CI −2.82 to 5.27; P = 0.55)”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on wearable technology, digoxin's efficacy, and the need for longer-term assessment. The hypothesis follows logically from the cited evidence. Limitations of prior research (e.g., acute studies only) are acknowledged and addressed by the longer-term design.
“digoxin has typically been considered a poorly effective drug for controlling heart rate, particularly on exertion, although this is based on acute studies only”
“we hypothesized that a wrist-worn wearable could: (1) address whether digoxin is inferior to beta-blockers for longer-term heart rate control in patients with AF at rest and on exertion; (2) adjust for differences in individual physical activity; and (3) explore whether wearable sensor data are comparable with conventional measurements for the prediction of clinical progress”
“digoxin has typically been considered a poorly effective drug for controlling heart rate, particularly on exertion, although this is based on acute studies only”
“we hypothesized that a wrist-worn wearable could: (1) address whether digoxin is inferior to beta-blockers for longer-term heart rate control in patients with AF at rest and on exertion; (2) adjust for differences in individual physical activity; and (3) explore whether wearable sensor data are comparable with conventional measurements for the prediction of clinical progress”
“Wearables offer an opportunity to assess each patient in their own environment, with longer-term evaluation better reflecting the extended time taken to achieve therapeutic benefit from digoxin”
Randomization used a computer-generated minimization algorithm, and the trial was open-label with blinded endpoint assessment. A sample size calculation was performed post hoc but justified the target enrollment. Inclusion/exclusion criteria were defined, and the analysis was intention-to-treat. Controls are inherent in the randomized comparison. Independent replication is not applicable for a single trial.
“Randomization was completed using a computer-generated minimization algorithm to ensure treatment arms were balanced for gender and AF symptoms”
“prospective, randomized, open-label, blinded end-point trial”
“a sample size of 40 participants would provide 90% power over 20 weeks to detect a 1/3 s.d. difference in heart rate (2 bpm) between digoxin and beta-blockers”
“Randomization was completed using a computer-generated minimization algorithm to ensure treatment arms were balanced for gender and AF symptoms”
“The RATE-AF trial was a prospective, randomized, open-label, blinded end-point trial”
“a sample size of 40 participants would provide 90% power over 20 weeks to detect a 1/3 s.d. difference in heart rate (2 bpm) between digoxin and beta-blockers”
Sex is reported for all participants, and both sexes are included, so sex justification is not applicable. Age, weight (not explicitly but BMI is mentioned), and health status are reported. Demographics include age, sex, and ethnicity (in discussion). Species/strain and housing are not applicable for a human study.
“mean age at randomization of 75.6 years (s.d. 8.4; range 61 to 90 years) and 40% women”
“the most common comorbidities being hypertension (74%) and heart failure (45%)”
“mean age at randomization of 75.6 years (s.d. 8.4; range 61 to 90 years) and 40% women”
“with the most common comorbidities being hypertension (74%) and heart failure (45%)”
The study reports approval from a named ethics committee (East Midlands–Derby Research Ethics Committee) with protocol number, and informed consent was obtained. Regulatory compliance is implied through Health Research Authority and MHRA approvals.
“Ethical approval was obtained from the East Midlands–Derby Research Ethics Committee (16/EM/0178)”
“were asked to sign an optional form to indicate informed consent”
“Ethical approval was obtained from the East Midlands–Derby Research Ethics Committee (16/EM/0178)”
“were asked to sign an optional form to indicate informed consent”
The drugs are named with dosing regimens. The wearable device (Fitbit Charge 2) and smartphone (Samsung A6) are identified. Software platforms (RADAR-base, REDCap, Stata, Python, TensorFlow) are named with versions where applicable. No antibodies, cell lines, or mycoplasma testing are applicable.
“Each participant was randomized to either digoxin 62.5–250 µg or bisoprolol 1.25–10 mg once daily”
“Each participant was randomized to either digoxin 62.5–250 µg or bisoprolol 1.25–10 mg once daily”
“wrist-worn Fitbit Charge 2 wearable device”
“Statistical analyses were performed using Stata v.17 (StataCorp LP)”
Tests are named (t-test, Kruskal-Wallis, Spearman, generalized linear models). Assumptions are handled by design (nonparametric tests where needed). Exact p-values are reported. Effect sizes with 95% CIs are provided. Software is identified. Data presentation includes individual trajectories and per-group n. Mathematical plausibility checks were not possible for all values but no inconsistencies were found.
“The unadjusted regression coefficient for digoxin versus beta-blockers was 1.22 (95% CI −2.82 to 5.27; P = 0.55)”
The data availability statement provides a clear access route with conditions and timeframes. Code is available on GitHub. Repository deposit is not applicable for patient-level data due to privacy, but summary data are available on request.
“Summary anonymized wearable sensor data are available for noncommercial purposes on request to the corresponding author (d.kotecha@bham.ac.uk; 60 days response time for decisions).”
“The machine learning framework for embedding multichannel time-series data is made freely available at: https://github.com/gkoutos-group/wearable_data_embedding”
“Summary anonymized wearable sensor data are available for noncommercial purposes on request to the corresponding author (d.kotecha@bham.ac.uk; 60 days response time for decisions).”
“The machine learning framework for embedding multichannel time-series data is made freely available at: https://github.com/gkoutos-group/wearable_data_embedding”
The trial is registered (NCT02391337). The MI-CLAIM checklist is referenced. All outcomes are reported, including null results. Limitations are discussed. Conclusions are proportional. Funding and COI are declared.
“ClinicalTrials.gov identifier: NCT02391337”
“The study is reported according to the Minimum Information about Clinical Artificial Intelligence Modeling (MI-CLAIM) checklist”
“ClinicalTrials.gov identifier: NCT02391337”
“The study is reported according to the Minimum Information about Clinical Artificial Intelligence Modeling (MI-CLAIM) checklist”
“The wearables were implemented post randomization and hence there is a risk of residual confounding”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 33 references by DOI: 33 verified.
Every extracted reference resolved against Crossref/OpenAlex with no retraction flags.
5 data/code links checked; 5 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT02391337LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.hra.nhs.uk/planning-and-improving-research/application-summaries/research-summaries/rate-af/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://fundingawards.nihr.ac.uk/award/CDF-2015-08-074LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.clinicaltrialsregister.eu/ctr-search/trial/2015-005043-13/resultsLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/gkoutos-group/wearable_data_embeddingResolves to GitHub (code repository).
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORtypoTable 1, Age row“77.2. (8.3)”→ Remove extra period: '77.2 (8.3)'Extra period after the number.
- MINORconsistencyAbstract vs. Results“n = 143,379,796”→ Ensure consistent use of commas in large numbers throughout.Large numbers are formatted with commas in some places and without in others.
- MINORconsistencyAbstract“n = 143,379,796”→ Ensure consistent use of commas in large numbers throughout.Large numbers are formatted with commas in some places but not others.
- MINORclarityMethods, Wearables substudy“Following work led by the PPI team that showed SF-36 to be as suboptimal measure of assessment”→ Change 'as suboptimal' to 'a suboptimal'Grammatical error.
The published work is robust and generally well-reported. An informed reader should weigh the minor statistical inconsistencies flagged by the verification component and the lack of a specified open-access repository timeline, but these do not undermine the core findings. No erratum is warranted based on the available evidence.
- 1.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 59.3% does not match the reported count 16/28Demonstrable critical failure — blocks the verdict from passing.
- 2.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 54.2% does not match the reported count 13/25Demonstrable critical failure — blocks the verdict from passing.
- 3.HIGHstatisticsIn the Statistical analysis section, clarify the handling of missing data and the proportion of missingness, and specify how the neural network incorporated missing data as a channel.Reviewers noted that missing data handling is mentioned but not detailed, which is important for reproducibility and interpretation.
- 4.HIGHdata codeIn the Data availability section, specify the intended open-access repository and expected deposit date for the summary anonymized data.Reviewer 2 flagged repository_deposit as inadequate; providing a concrete plan enhances transparency and reproducibility.
- 5.MEDIUMreportingIn the Results section, report the exact p-values for subgroup analyses (e.g., P = 0.48, 0.47, 0.97) in the main text rather than only in the supplementary.Exact p-values improve transparency and allow readers to assess the strength of subgroup findings.
- 6.MEDIUMreportingIn the Methods section, add a CONSORT-style flow diagram for the substudy to improve transparency.A flow diagram clarifies participant flow and attrition, which is a standard expectation for trial reports.
- 7.MEDIUMreportingIn the Methods section, provide the full statistical analysis plan as a supplementary file to enhance reproducibility.A detailed analysis plan allows readers to verify that the analyses were pre-specified and reduces the risk of selective reporting.
- 8.MEDIUMotherIn the Methods section, report the exact version numbers for Python, scikit-learn, and TensorFlow used in the analysis.Exact software versions are essential for reproducibility of the machine learning analyses.
- 9.MEDIUMstatisticsIn the Statistical analysis section, add a statement on the handling of outliers in the statistical analysis.Reviewer 2 noted that outlier handling is not explicitly described for the statistical analyses, which is important for robustness.
- 10.MEDIUMreportingIn the Discussion, discuss the potential impact of the open-label design on the results and how blinding of endpoint assessment mitigated bias.Addressing the open-label design strengthens the interpretation of the findings.
- 11.MEDIUMreportingIn the Abstract, consider adding a statement on the generalizability of the findings given the limited ethnic diversity.Highlighting the limitation in the abstract helps readers interpret the applicability of the results.
- 12.LOWcopyeditIn Table 1, Age row, remove the extra period: change '77.2. (8.3)' to '77.2 (8.3)'.Typographical error that should be corrected for professionalism.
- 13.LOWcopyeditThroughout the manuscript, ensure consistent use of commas in large numbers (e.g., '143,379,796' vs '143379796').Consistency in number formatting improves readability and avoids confusion.
- 14.LOWcopyeditIn Methods, Wearables substudy, change 'as suboptimal' to 'a suboptimal'.Grammatical error that should be fixed.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.