Artificial intelligence for individualized treatment of persistent atrial fibrillation: a randomized controlled trial.
Deisenhofer I, Albenque JP, Busch S, Gitenay E, Mountantonakis SE, Roux A, Horvilleur J, Bakouboula B, Oza S, Abbey S, Theodore G, Lepillier A, Guyomar Y, Bessiere F, Jan Smit J, Mohr Durdez T, Milpied P, Appetiti A, Guerrero D, De Potter T, De Chillou C, Goldbarg S, Verma A, Hummel JD, TAILORED-AF Investigators
- DOI
- 10.1038/s41591-025-03517-w
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/97bbb254-0875-4047-80ee-9c27a0651a29 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×6−3★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ReportingData & code availability partially met−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy endpoint is freedom from documented AF, which is a clinical outcome, but the trial's main claim of superiority is based on this endpoint. However, the secondary endpoint of freedom from any atrial arrhythmia did not show a significant difference after one procedure, and the primary endpoint does not capture all arrhythmias. The surrogate issue arises because the primary endpoint is a clinical outcome, but the trial also relies on surrogate markers like AF termination and electrogram dispersion as mechanistic proxies. The primary endpoint is a hard clinical outcome, so the surrogate verdict is not applicable. However, the efficacy claim is based on a clinical endpoint, so the surrogate is adequate.
“The primary efficacy endpoint was freedom from documented AF with or without antiarrhythmic drugs at 12 months after a single ablation procedure.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted, rigorously reported randomized controlled trial with strong methodology, clear ethical approvals, and transparent reporting. The main weakness is the vague data availability statement and lack of code sharing, which limits reproducibility.
Both reviewers independently scored all eight dimensions and agreed on every status; no divergence to reconcile. The statistics verification recomputed 12 tests (all consistent) but covers only a subset of reported analyses; the remaining statistics are unverified. The citation check found no retracted or unresolved references.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 11 tests: 11 consistent, 0 inconsistent; 11 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Primary endpoint HR from mITT population
“hazard ratio (HR), 0.3; 95% CI, 0.21–0.57; log-rank P < 0.0001”
Taken as given: The HR is for the tailored vs anatomical arm.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed p-value from HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.3, 0.21, 0.57, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary endpoint HR in PP population
“All patients in the PP population (HR, 0.29; 95% CI, 0.16–0.51).”
Taken as given: The HR is for the tailored vs anatomical arm.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed p-value from HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.29, 0.16, 0.51, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary endpoint HR in subgroup with AF duration >=6 months (PP)
“Patients with persistent AF with a duration of AF greater or equal to 6 months in the PP population (HR, 0.18; 95% CI, 0.08–0.40).”
Taken as given: The HR is for the tailored vs anatomical arm.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed p-value from HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.18, 0.08, 0.40, 1) - CONSISTENTreported p > .070 · recomputed p = .212Reviewer 1Secondary endpoint: freedom from any atrial arrhythmia after one or two procedures (mITT)
“76% versus 71%; HR, 0.77; 95% CI, 0.51–1.16; log-rank P = 0.07”
Taken as given: The HR is for the tailored vs anatomical arm.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed p-value from HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.77, 0.51, 1.16, 1) - CONSISTENTreported p > .160 · recomputed p = .711Reviewer 1Secondary endpoint: freedom from any atrial arrhythmia after one procedure (mITT)
“60% versus 60%; HR, 0.94; 95% CI, 0.68–1.31; log-rank P = 0.16”
Taken as given: The HR is for the tailored vs anatomical arm.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed p-value from HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.94, 0.68, 1.31, 1) - CONSISTENTreported p > .160 · recomputed p = .273Reviewer 1Secondary endpoint: freedom from any atrial arrhythmia after one procedure (PP)
“Outcome after a single procedure for all patients in the PP population (HR, 0.81; 95% CI, 0.56–1.19).”
Taken as given: The HR is for the tailored vs anatomical arm.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed p-value from HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.81, 0.56, 1.19, 1) - CONSISTENTreported p > .050 · recomputed p = .055Reviewer 1Secondary endpoint: freedom from any atrial arrhythmia after one or two procedures (PP)
“Outcome after one or two procedures for all patients in the PP population (HR, 0.62; 95% CI, 0.38–1.01).”
Taken as given: The HR is for the tailored vs anatomical arm.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed p-value from HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.62, 0.38, 1.01, 1) - CONSISTENTreported p < .050 · recomputed p = .038Reviewer 1Secondary endpoint: freedom from any atrial arrhythmia after one procedure in >=6 months subgroup (PP)
“Outcome after a single procedure for patients with a duration of AF greater or equal to 6 months in the PP population (HR, 0.60; 95% CI, 0.37–0.97).”
Taken as given: The HR is for the tailored vs anatomical arm.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed p-value from HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.60, 0.37, 0.97, 1) - CONSISTENTreported p < .050 · recomputed p = .020Reviewer 1Secondary endpoint: freedom from any atrial arrhythmia after one or two procedures in >=6 months subgroup (PP)
“Outcome after one or two procedures for patients with a duration of AF greater or equal to 6 months in the PP population (HR, 0.50; 95% CI, 0.28–0.90).”
Taken as given: The HR is for the tailored vs anatomical arm.; The CI is a 95% confidence interval.; The p-value is two-sided.Method: Recomputed p-value from HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.50, 0.28, 0.90, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Acute AF termination by ablation (Fisher's exact test)
“Acute AF termination by ablation | 122/186 (66%) | 26/169 (15%) | <2.2 × 10 –16”
Taken as given: The table is 2x2 with counts: tailored events=122, tailored non-events=64 (186-122), anatomical events=26, anatomical non-events=143 (169-26).; The p-value is two-sided.Method: Recomputed two-sided Fisher's exact test from the 2x2 table.How we recomputed it: pFisher2x2(122, 64, 26, 143, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Acute sinus rhythm conversion by ablation (Fisher's exact test)
“Acute sinus rhythm conversion by ablation | 100/187 (53%) | 23/172 (13%) | 3.0 × 10 –16”
Taken as given: The table is 2x2 with counts: tailored events=100, tailored non-events=87 (187-100), anatomical events=23, anatomical non-events=149 (172-23).; The p-value is two-sided.Method: Recomputed two-sided Fisher's exact test from the 2x2 table.How we recomputed it: pFisher2x2(100, 87, 23, 149, 0) - UNCOMPUTABLEreported p = .070 · recomputed p = .212Reviewer 2Secondary endpoint HR and CI for freedom from any atrial arrhythmia after one or two procedures in mITT
“76% versus 71%; HR, 0.77; 95% CI, 0.51–1.16; log-rank P = 0.07”
Taken as given: The HR is 0.77 with 95% CI 0.51-1.16.; The CI is two-sided at 95%.; The p-value is two-sided from a log-rank test.Method: Recomputed p-value from HR and CI using pCI function.How we recomputed it: pCI(0.77, 0.51, 1.16, 1)
- lowinternal contradictionThe results section says 'A total of 370 patients underwent catheter ablation' but the flow diagram says 'Of the 374 patients randomized to either treatment, 188 were assigned to a tailored cardiac ablation procedure and 186 were assigned to a standard-of-care PVI procedure. Before the ablation procedure, one patient in the tailored arm and three patients in the anatomical arm withdrew their consent.' This implies 374-4=370 underwent ablation, which is consistent.
A total of 370 patients underwent catheter ablation ... Of the 374 patients randomized to either treatment, 188 were assigned to a tailored cardiac ablation procedure and 186 were assigned to a standard-of-care PVI procedure. Before the ablation procedure, one patient in the tailored arm and three patients in the anatomical arm withdrew their consent.
Figure 1reviewer’s wording - lowinternal contradictionThe primary endpoint analysis in the mITT population is stated as n=357, but the sum of the two arms (180+177) is 357, which is consistent. However, the text says 'Of the 357 patients in the mITT population, 43 and 45 were excluded because of important protocol deviations in the tailored and anatomical arms, respectively.' This implies 180+177=357, but 43+45=88 excluded, leaving 269 in PP, which matches. No contradiction.
“The primary analysis in the mITT population consisted of 180 patients and 177 patients in the tailored and anatomical arms, respectively. Of the 357 patients in the mITT population, 43 and 45 were excluded because of important protocol deviations in the tailored and anatomical arms, respectively. A total of 137 and 132 patients were included in the PP population in the tailored and anatomical arms, respectively.”
Figure 1Find in source - lowinternal contradictionThe results section says 'In the tailored arm, AF terminated by ablation directly to sinus rhythm in 27 patients (22%) and to regular atrial tachycardias (ATs) in 95 patients (78%).' The sum 27+95=122, which matches the number of patients with AF termination (122). However, the percentages 22% and 78% sum to 100%, but 27/122=22.1% and 95/122=77.9%, consistent.
“In the tailored arm, AF terminated by ablation directly to sinus rhythm in 27 patients (22%) and to regular atrial tachycardias (ATs) in 95 patients (78%).”
ResultsFind in source - lowinternal contradictionThe safety endpoint table (Table 3) lists 'Subjects with minor procedure-related complications | 15 d | 6 d | 0.07' and footnote d explains that some patients had multiple events, so the sum of components (9+4+3+1+1=18) exceeds the total (15). This is explained.
Subjects with minor procedure-related complications | 15 d | 6 d | 0.07 ... d Two patients had vascular access complications and fluid overload events; one patient who had a mild pericardial effusion subsequently had a fluid overload event; accordingly, the individual components add to more than the total number of patients with any event.
Table 3reviewer’s wording - lowinternal contradictionThe results section says 'By contrast, in the anatomical arm, 29 of the 38 repeat procedures (76%) were performed for AF.' This is consistent with 38 repeat procedures in anatomical arm.
“By contrast, in the anatomical arm, 29 of the 38 repeat procedures (76%) were performed for AF.”
ResultsFind in source - lowinternal contradictionThe primary endpoint HR is reported as 0.3 in the text and 0.34 in Figure 2a. This may be due to rounding or different populations (mITT vs all who underwent ablation).
hazard ratio (HR), 0.3; 95% CI, 0.21–0.57; log-rank P < 0.0001 ... a , All patients in the mITT population (HR, 0.34; 95% CI, 0.21–0.57).
Figure 2reviewer’s wording
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1The tailored procedure results in a higher rate of freedom from any atrial arrhythmia after a mean of 1.26 ablation procedures.The secondary endpoint of freedom from any atrial arrhythmia after one or two procedures was not significantly different in the mITT population (P=0.07), but was significant in the PP population. The claim is partially supported.Evidence: Secondary endpoint: 76% vs 71% (mITT, P=0.07); PP population significant.
“Patients in the tailored arm also experienced a higher rate of freedom from any arrhythmia after a mean of 1.26 ablation procedures (47 out of 180 patients required a second procedure).”
DiscussionFind in source - partialReviewer 2The AI algorithm provides objective, reproducible, and reliable identification of ablation target areas.The algorithm's performance is validated on test datasets, but the claim of reproducibility across centers is inferred from the trial results, not directly measured.Evidence: Algorithm performance on testing datasets (AUC 0.94) and multi-annotation agreement.
“On the single-annotated electrogram testing dataset, the algorithm achieved high performance, achieving a receiver operating characteristic (ROC) area under the curve (AUC) score of 0.94”
ResultsFind in source - supportedReviewer 1AI-guided ablation of spatio-temporal dispersion areas in addition to PVI is superior to PVI alone in eliminating AF at 1-year follow-up in patients with persistent and long-standing persistent AF.The primary endpoint result (88% vs 70%, HR 0.3, p<0.0001) directly supports this claim.Evidence: Primary efficacy endpoint: 88% vs 70%, HR 0.3, 95% CI 0.21-0.57, log-rank P < 0.0001.
“These results show that AI-guided ablation of spatio-temporal dispersion areas in addition to PVI is superior to PVI alone in eliminating AF at 1-year follow-up in patients with persistent and long-standing persistent AF.”
AbstractFind in source - supportedReviewer 1The AI algorithm achieved high performance in detecting spatio-temporal dispersion.The algorithm's AUC of 0.94 and the monotonic increase in probability with expert agreement support this claim.Evidence: Algorithm performance: AUC 0.94; probabilities 0.02 to 0.83 with increasing expert agreement.
“the algorithm achieved high performance, achieving a receiver operating characteristic (ROC) area under the curve (AUC) score of 0.94”
ResultsFind in source - supportedReviewer 1The tailored procedure is particularly beneficial in patients with prolonged persistent AF (≥6 months).Prespecified subgroup analyses show larger differences in this subgroup, supporting the claim.Evidence: Subgroup analysis: 23% difference in mITT, up to 30% in PP for primary endpoint; secondary endpoints also favor tailored arm.
“In the prespecified subgroup of patients with an AF duration of ≥6 months in the mITT population ( n = 196), the difference in treatment success was even more substantial, with a 23% difference between the two arms”
ResultsFind in source - supportedReviewers 1, 2The tailored procedure is safe, with no significant difference in the composite safety endpoint.The safety endpoint did not differ between arms, and the paper reports no significant differences in adverse events.Evidence: Safety endpoint: 8/187 (4%) vs 5/183 (3%), P > 0.05.
“Death, cerebrovascular events, or major treatment-related serious adverse events (composite safety endpoint) occurred in 8/187 (4%) patients who underwent a tailored cardiac ablation procedure and 5/183 (3%) patients who underwent PVI only ( P > 0.05 for all comparisons; Table ).”
ResultsFind in source - supportedReviewer 1The AI-driven software provided objective, reproducible, and reliable identification of ablation target areas across all 26 centers and 51 operators.The paper provides evidence of algorithm performance and consistent use across centers, though direct reproducibility data are not shown.Evidence: Algorithm performance and multi-center use.
“This led to an objective, reproducible, and reliable identification of ablation target areas for individual patients across all 26 centers and 51 operators.”
DiscussionFind in source - supportedReviewer 2AI-guided ablation of spatio-temporal dispersion areas in addition to PVI is superior to PVI alone in eliminating AF at 1-year follow-up.The primary endpoint was met with a significant difference (88% vs 70%, HR 0.3, P<0.0001), supporting the claim.Evidence: Primary efficacy endpoint results in mITT population.
“One year post-procedure, the trial met its primary efficacy endpoint, which was achieved in 88% of patients in the tailored arm compared with 70% of patients in the anatomical arm (log-rank P < 0.0001 for superiority).”
AbstractFind in source - supportedReviewer 2The tailored procedure is particularly beneficial in patients with longer AF duration (≥6 months).Prespecified subgroup analysis showed a larger difference (23% difference in mITT, up to 30% in PP), supporting the claim.Evidence: Subgroup analysis of patients with AF duration ≥6 months.
In the prespecified subgroup of patients with an AF duration of ≥6 months in the mITT population ( n = 196), the difference in treatment success was even more substantial, with a 23% difference between the two arms.
Resultsreviewer’s wording
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary efficacy endpoint is freedom from documented AF, which is a clinical outcome, but the trial's main claim of superiority is based on this endpoint. However, the secondary endpoint of freedom from any atrial arrhythmia did not show a significant difference after one procedure, and the primary endpoint does not capture all arrhythmias. The surrogate issue arises because the primary endpoint is a clinical outcome, but the trial also relies on surrogate markers like AF termination and electrogram dispersion as mechanistic proxies. The primary endpoint is a hard clinical outcome, so the surrogate verdict is not applicable. However, the efficacy claim is based on a clinical endpoint, so the surrogate is adequate.
“The primary efficacy endpoint was freedom from documented AF with or without antiarrhythmic drugs at 12 months after a single ablation procedure.”
- ADEQUATEEffect sizeThe primary endpoint showed a significant difference: 88% vs 70% freedom from AF, with HR 0.3 (95% CI 0.21-0.57), log-rank P<0.0001. This is a clinically meaningful difference of 18 percentage points, and the effect size is anchored to a hard clinical outcome. The effect is statistically supported and clinically material.
“The estimated probability of treatment success in the modified-intention-to-treat (mITT) population (n = 357) was 88% for the tailored procedure and 70% for the anatomical procedure (hazard ratio (HR), 0.3; 95% CI, 0.21–0.57; log-rank P < 0.0001).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior studies on PVI and 'PVI plus' strategies, acknowledges their limitations, and explains the need for objective, reproducible electrogram detection. The rationale linking AI-based detection of spatio-temporal dispersion to improved outcomes is clearly stated. Limitations of prior research are addressed by the AI approach.
“A key challenge in incorporating electrogram-based ablation strategies into larger trials has been the need to ensure objectivity, consistency, and most importantly, reproducibility in the detection and adjudication of intracardiac electrograms across different physicians and centers.”
“The purpose of the TAILORED-AF randomized controlled trial was to evaluate whether a tailored cardiac-ablation procedure targeting AI-detected areas harboring spatio-temporal dispersion, in addition to PVI, is more effective than an anatomical PVI-only procedure in patients with persistent and long-standing persistent AF.”
“Previous studies investigating ‘PVI plus’ ablation methods, such as PVI plus standardized anatomically defined atrial areas (for instance, the posterior wall of the left atrium) or PVI plus targeting magnetic resonance imaging (MRI)-detected areas of atrial fibrosis, have failed to demonstrate a clear advantage over a more comprehensive ablation strategy”
“A key challenge in incorporating electrogram-based ablation strategies into larger trials has been the need to ensure objectivity, consistency, and most importantly, reproducibility in the detection and adjudication of intracardiac electrograms across different physicians and centers.”
“The purpose of the TAILORED-AF randomized controlled trial was to evaluate whether a tailored cardiac-ablation procedure targeting AI-detected areas harboring spatio-temporal dispersion, in addition to PVI, is more effective than an anatomical PVI-only procedure in patients with persistent and long-standing persistent AF.”
Randomization used a random permuted block method with stratification by AF type and site. Blinding of patients and endpoint assessors is described. A priori sample size calculation with power and alpha is provided. Inclusion/exclusion criteria are detailed, and the mITT and PP populations are defined. Outlier handling is addressed through pre-specified analysis populations and missing data handling.
“Randomization was performed using a random permuted block method through the electronic case report form, with block sizes randomly chosen as two, four, and six. Randomization was stratified according to AF type and site.”
“double-blind (study subject and endpoint assessor)”
“a total of 292 participants were needed for the study to have a power of 80% at a one-sided alpha level of 0.025.”
“Randomization was performed using a random permuted block method through the electronic case report form, with block sizes randomly chosen as two, four, and six. Randomization was stratified according to AF type and site.”
“double-blind (study subject and endpoint assessor)”
“a total of 292 participants were needed for the study to have a power of 80% at a one-sided alpha level of 0.025. Assuming a dropout rate of 22% (no index ablation performed or loss to follow-up), 374 participants were required.”
The baseline table reports age, sex, and a comprehensive list of comorbidities and cardiac parameters. Both sexes are enrolled, so sex justification is not applicable. Demographics are adequately reported.
“Sex, female | 77 (21%) | 42 (23%) | 35 (19%)”
“Age (years) | 65.7 ± 8.5 | 66.4 ± 8.5 | 64.9 ± 8.5”
“Hypertension b | 227 (61%) | 118 (63%) | 109 (60%)”
The paper lists specific IRBs that approved the trial, including WIRB and others. Informed consent is mentioned. Compliance with the Declaration of Helsinki is stated. The trial is registered (NCT04702451).
“was approved by the institutional or ethics review board at each center: Western IRB (WIRB), Ascension St. Vincent IRB, Rhode Island Hospital IRB (United States), CPP SUD-EST IV (France), Technische Universität München (TUM) Ethikkommission, Landesärztekammer Baden-Württemberg Ethikkommission (Germany), OLV Ziekenhuis vzw Ethisch Comite (Belgium), and Brabant Medical Ethics Committee (Netherlands).”
“After providing written informed consent, adults with symptomatic persistent or long-standing persistent AF”
“It was conducted in accordance with principles of the Declaration of Helsinki.”
“was approved by the institutional or ethics review board at each center: Western IRB (WIRB), Ascension St. Vincent IRB, Rhode Island Hospital IRB (United States), CPP SUD-EST IV (France), Technische Universität München (TUM) Ethikkommission, Landesärztekammer Baden-Württemberg Ethikkommission (Germany), OLV Ziekenhuis vzw Ethisch Comite (Belgium), and Brabant Medical Ethics Committee (Netherlands).”
“After providing written informed consent, adults with symptomatic persistent or long-standing persistent AF”
“It was conducted in accordance with principles of the Declaration of Helsinki.”
The investigational product (Volta AF-Xplorer) is named with manufacturer and description. Ablation catheters and mapping systems are identified (e.g., Carto3, EnSite, Rhythmia). Software tools are named with versions (SAS 9.4, R 4.3.2). Antibodies, cell lines, and mycoplasma testing are not applicable as this is a clinical trial without wet-lab assays.
“The Volta AF-Xplorer (previously VX1, Volta Medical) software is an embedded, expertise-based AI tool”
“using three commercially available navigation systems (Carto3, Biosense Webster; EnSite Precision or EnSite X, Abbott Medical; and Rhythmia HDx, Boston Scientific; Extended Data Table )”
“Data were analyzed with SAS Software, version 9.4 (SAS Institute), and R Statistical Software, version 4.3.2 (R Core Team).”
“The Volta AF-Xplorer (previously VX1, Volta Medical) software is an embedded, expertise-based AI tool”
“using three commercially available navigation systems (Carto3, Biosense Webster; EnSite Precision or EnSite X, Abbott Medical; and Rhythmia HDx, Boston Scientific; Extended Data Table )”
“Data were analyzed with SAS Software, version 9.4 (SAS Institute), and R Statistical Software, version 4.3.2 (R Core Team).”
The primary analysis uses Cox proportional hazards and log-rank tests, with HRs and 95% CIs reported. Exact p-values are given (e.g., P < 0.0001). Software is identified. Data presentation includes Kaplan-Meier curves and per-group n. Mathematical plausibility checks were not applicable due to large N and continuous outcomes.
“Primary and secondary endpoints were analyzed using a Cox proportional-hazards model, with treatment group and a dichotomous effect for AF duration (<6 months or ≥6 months) as independent terms”
“log-rank P < 0.0001”
“hazard ratio (HR), 0.3; 95% CI, 0.21–0.57”
“A two-sided log-rank test at a significance alpha level of 0.05 was used, and Kaplan–Meier curves were generated.”
“hazard ratio (HR), 0.3; 95% CI, 0.21–0.57; log-rank P < 0.0001”
“Data were analyzed with SAS Software, version 9.4 (SAS Institute), and R Statistical Software, version 4.3.2 (R Core Team).”
The data availability statement says 'All supporting data are available in the article and the Supplementary Information' and that source data will not be shared due to patient privacy, which is a legitimate reason but lacks a concrete access mechanism. No repository deposit or accession numbers are provided. Code is not shared; the AI code is proprietary. For a clinical trial, managed access is acceptable, but the statement is vague.
“All supporting data are available in the article and the Supplementary Information. Source data will not be shared owing to patient privacy obligations applicable to the sponsor under the European Union General Data Protection Regulation (GDPR)”
“No custom code was used for analysis of the numerical data reported in this manuscript. The AI system used is a deterministic (locked) medical device commercially available in the European Union and the United States. The AI-based code for EGM adjudication that forms the core of the medical device is intellectual property of Volta Medical and is not publicly available.”
“All supporting data are available in the article and the Supplementary Information. Source data will not be shared owing to patient privacy obligations applicable to the sponsor under the European Union General Data Protection Regulation (GDPR) and in particular owing to the obligations of privacy to which the sponsor committed itself in the informed consent form signed by the patients.”
“No custom code was used for analysis of the numerical data reported in this manuscript. The AI system used is a deterministic (locked) medical device commercially available in the European Union and the United States. The AI-based code for EGM adjudication that forms the core of the medical device is intellectual property of Volta Medical and is not publicly available.”
The trial is registered (NCT04702451). Methods are comprehensive. Limitations are explicitly discussed. Conclusions are proportional to the evidence. Funding and competing interests are disclosed. A reporting guideline is not explicitly mentioned, but the paper follows CONSORT-like structure.
“ClinicalTrials.gov identifier: NCT04702451”
“Our trial has several limitations: (1) the primary endpoint focuses only on freedom from AF, and does not account for freedom from any atrial arrhythmia after one procedure.”
“ClinicalTrials.gov identifier: NCT04702451”
“Our trial has several limitations: (1) the primary endpoint focuses only on freedom from AF, and does not account for freedom from any atrial arrhythmia after one procedure.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 34 references by DOI: 34 verified.
Every extracted reference resolved against Crossref/OpenAlex with no retraction flags.
2 data/code links checked; 2 live.
- datahttps://clinicaltrials.gov/study/NCT04702451LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT04702451LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
18 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 18 minor suggestions below.
18 copyedit issues flagged: mostly consistency, clarity.
- MINORconsistencyAbstract“n = 187, 23% women) or to a conventional PVI-only procedure (anatomical arm, n = 183, 19% women)”→ Ensure consistent use of 'women' vs 'female' throughout.Table 1 uses 'Sex, female' while abstract uses 'women'.
- MINORclarityResults, Primary efficacy endpoint“The estimated probability of treatment success in the modified-intention-to-treat (mITT) population ( n = 357) was 88% for the tailored procedure and 70% for the anatomical procedure (hazard ratio (HR), 0.3; 95% CI, 0.21–0.57; log-rank P < 0.0001)”→ Clarify that HR is for the tailored vs anatomical arm.HR direction is not explicitly stated.
- MINORconsistencyResults, Primary efficacy endpoint“In the anatomical arm, the primary endpoint was achieved in 124 patients and was not achieved in 53; 6 had no follow-up data beyond the blanking period.”→ Check that 124+53+6 = 183, which is the anatomical arm size.124+53+6=183, consistent.
- MINORconsistencyResults, Primary efficacy endpoint“In the tailored arm, the primary endpoint was achieved in 158 and was not achieved in 22; 7 had no follow-up data beyond the blanking period.”→ Check that 158+22+7 = 187, which is the tailored arm size.158+22+7=187, consistent.
- MINORconsistencyResults, Primary efficacy endpoint“The estimated probability of treatment success in the modified-intention-to-treat (mITT) population ( n = 357) was 88% for the tailored procedure and 70% for the anatomical procedure”→ Verify that 88% and 70% correspond to 158/180 and 124/177.158/180=87.8%, 124/177=70.1%, consistent.
- MINORconsistencyResults, Primary efficacy endpoint“The difference between the two arms was confirmed in the PP population ( n = 269; Fig. ) and in the 306 patients (32 versus 19 in the tailored and anatomical arms, respectively; P = 0.07) that were not taking antiarrhythmic drugs at 12 months ( n = 306; 86% versus 68%; HR, 0.35; 95% CI, 0.21–0.59; log-rank P < 0.001).”→ Clarify the numbers 32 and 19: are these the number of events?The sentence is confusing; 32 and 19 likely refer to recurrences.
- MINORconsistencyResults, Secondary efficacy endpoints“In the mITT population, more patients in the tailored arm achieved freedom from any atrial arrhythmia at 12 months after one or two ablation procedures (76% versus 71%; HR, 0.77; 95% CI, 0.51–1.16; log-rank P = 0.07)”→ Check that 76% and 71% are consistent with the numbers of patients.No exact counts given, but plausible.
- MINORconsistencyResults, Index procedure characteristics“The periprocedural AF termination rate was significantly higher in the tailored arm than in the anatomical arm (66% versus 15%; P < 0.001).”→ Verify that 66% and 15% correspond to 122/186 and 26/169.122/186=65.6%, 26/169=15.4%, consistent.
- MINORconsistencyResults, Index procedure characteristics“Acute AF termination by ablation | 122/186 (66%) | 26/169 (15%)”→ Check that denominators are correct.Denominators differ from arm sizes; likely due to missing data.
- MINORconsistencyResults, Safety endpoint“Death, cerebrovascular events, or major treatment-related serious adverse events (composite safety endpoint) occurred in 8/187 (4%) patients who underwent a tailored cardiac ablation procedure and 5/183 (3%) patients who underwent PVI only ( P > 0.05 for all comparisons; Table ).”→ Verify that 8/187=4.3% and 5/183=2.7%.Consistent.
- MINORconsistencyTable 3“Subjects with major procedure-related complications | 5 | 5 c | >0.99”→ Check footnote c explains why components sum to more than total.Footnote c explains overlap.
- MINORconsistencyTable 3“Subjects with minor procedure-related complications | 15 d | 6 d | 0.07”→ Check footnote d explains overlap.Footnote d explains overlap.
- MINORconsistencyResults, Repeat procedures“During the study period, a total of 88 patients (50/180 versus 38/177 in the tailored and anatomical arms, respectively; P = 0.18) underwent a repeat catheter ablation procedure.”→ Verify that 50+38=88.50+38=88, consistent.
- MINORconsistencyResults, Repeat procedures“The number of repeat procedures considered in the multiple-procedures endpoint was not significantly different between the arms (47/180 versus 34/177; P = 0.13).”→ Verify that 47+34=81, not 88.47+34=81, but 88 patients had repeat procedures; the difference is due to exclusions.
- MINORconsistencyResults, Repeat procedures“In the anatomical arm, repeat procedures included a roof line and mitral line ( n = 13; 36%), posterior wall isolation ( n = 9; 25%), posterior wall isolation and mitral line ( n = 3; 8%), roof line ( n = 4; 11%), or mitral line ( n = 1; 3%), averaging 1.6 anatomical lines per patient. No information was provided for 2 patients, and the remaining 6 patients (17%) just had PV re-isolation and/or other ATs mapped and ablated.”→ Check that the percentages sum to 100%.36+25+8+11+3+17=100, consistent.
- MINORconsistencyAbstract“log-rank P < 0.0001 for superiority”→ Consider reporting exact p-value if available, e.g., P < 0.0001 is acceptable but exact value may be preferred.P-value reported as threshold, which is common in medical literature.
- MINORclarityResults, Primary efficacy endpoint“The estimated probability of treatment success in the modified-intention-to-treat (mITT) population ( n = 357) was 88% for the tailored procedure and 70% for the anatomical procedure (hazard ratio (HR), 0.3; 95% CI, 0.21–0.57; log-rank P < 0.0001)”→ Clarify that HR is for the tailored vs anatomical arm (i.e., HR < 1 favors tailored).HR direction is implied but not explicitly stated.
- MINORconsistencyResults, Primary efficacy endpoint“In the anatomical arm, the primary endpoint was achieved in 124 patients and was not achieved in 53; 6 had no follow-up data beyond the blanking period.”→ Check that 124+53+6 = 183, which matches the anatomical arm N.Numbers sum correctly.
The published work is methodologically robust and the primary findings are well supported. An informed reader should weigh the limited data/code availability and the proprietary AI algorithm when assessing reproducibility; these do not undermine the core conclusions but warrant caution in independent verification.
- 1.HIGHdata codeIn the Data availability section, provide a concrete managed-access mechanism (e.g., a data access committee or platform like Vivli) with conditions and timeframe for requesting de-identified source data.The current statement is a blanket refusal that does not meet transparency expectations for a clinical trial and limits reproducibility.
- 2.HIGHdata codeIn the Code availability section, clarify that no custom analysis code was used and consider sharing the statistical analysis scripts (SAS/R) in a public repository to enhance reproducibility.Even though the AI code is proprietary, the statistical analysis code can be shared without compromising intellectual property.
- 3.MEDIUMreportingIn the Methods or a Reporting Summary, explicitly state adherence to the CONSORT reporting guideline and provide a completed CONSORT checklist.Explicit guideline adherence improves transparency and is expected for randomized trials.
- 4.MEDIUMreportingIn the Results, clarify the direction of the primary hazard ratio (e.g., 'HR < 1 favors the tailored arm') to avoid ambiguity.The HR direction is implied but not explicitly stated, which could confuse readers.
- 5.MEDIUMreportingIn the Results, clarify the numbers '32 versus 19' in the antiarrhythmic drug subgroup analysis to indicate they are recurrence events.The sentence is confusing and could be misinterpreted.
- 6.MEDIUMreportingIn the Results, reconcile the primary HR reported as 0.3 in the text and 0.34 in Figure 2a, explaining any difference in population or rounding.An apparent discrepancy between text and figure may raise concerns about consistency.
- 7.MEDIUMreportingIn the Abstract and Table 1, standardize the terminology for sex (use 'female' consistently instead of mixing 'women' and 'female').Consistent terminology improves clarity and professionalism.
- 8.MEDIUMreportingIn the Results, report exact p-values for secondary endpoints that are currently given only as thresholds (e.g., P < 0.001).Exact p-values allow readers to assess the strength of evidence more precisely.
- 9.LOWreportingIn the Discussion, add a statement on whether adverse events were adjudicated by an independent committee.Independent adjudication of safety endpoints strengthens the credibility of safety findings.
- 10.LOWreportingIn the Discussion, address the generalizability of the findings given the exclusion of patients with long-standing persistent AF in the US.Acknowledging this limitation helps readers interpret the applicability of the results.
- 11.LOWotherIn the Data availability section, specify exactly which data are available in the article and supplementary, and what would be shared upon request.Reducing ambiguity about data availability helps readers know what they can access.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.