Vernakalant versus procainamide for rapid cardioversion of patients with acute atrial fibrillation (RAFF4): randomised clinical trial.
Stiell IG, Taljaard M, Eagles D, Yadav K, Vadeboncoeur A, Hohl CM, Archambault PM, Birnie D, Brown E, Campbell SG, Chen Y, Clement CM, Cournoyer A, de Wit K, Emond M, Macle L, McRae AD, Mercier E, Morris J, Mohamad G, Nemnom MJ, Nicholls SG, Pare D, Parkash R, Sivilotti M, Thavorn K, Perry JJ
- DOI
- 10.1136/bmj-2025-085632
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/4907544c-8f8c-4b90-8b12-259bb03d07cf is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 26 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted randomized open-label trial with rigorous design, clear reporting, and appropriate statistical methods. The paper demonstrates strong adherence to reporting guidelines and transparency, with only minor copyedit issues and a low-severity internal inconsistency in table labeling.
Both reviewers independently scored all dimensions as pass with high confidence, and no divergence was found. The study type is interventional (randomized controlled trial). Non-applicable sub-criteria (e.g., animal housing, cell line authentication) were excluded from scoring. The statistics verification covered only a subset of reported tests (those with test statistics + df or effect estimates + CI); other statistics remain unverified.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 11 tests: 11 consistent, 0 inconsistent; 6 recomputed directly from the reported test statistics, 5 via agent-written checks.
- CONSISTENTreported p = .006 · recomputed p = .005Recomputed adjusted odds ratio 1.87 (95% CI 1.2–2.9), reported p=0.006
“adjusted odds ratio 1.87, 95% confidence interval 1.2 to 2.9, P=0.006”
Taken as given: 1.2–2.9 is a two-sided 95% confidence interval for the adjusted odds ratio of 1.87, not a range, an IQR, or a different interval level; the adjusted odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.006 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.87, 1.2, 2.9, 1) - CONSISTENTreported p = .033 · recomputed p = .037Recomputed odds ratio 0.62 (95% CI 0.39–0.96), reported p=0.033
“odds ratio 0.62, 95% confidence interval 0.39 to 0.96, P=0.033”
Taken as given: 0.39–0.96 is a two-sided 95% confidence interval for the odds ratio of 0.62, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.033 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.62, 0.39, 0.96, 1) - CONSISTENTreported p = .001 · recomputed p = <.001Recomputed adjusted odds ratio 3.1 (95% CI 1.7–5.5), reported p=0.001
“adjusted odds ratio 3.1, 95% confidence interval 1.7 to 5.5, P=0.001”
Taken as given: 1.7–5.5 is a two-sided 95% confidence interval for the adjusted odds ratio of 3.1, not a range, an IQR, or a different interval level; the adjusted odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(3.1, 1.7, 5.5, 1) - CONSISTENTreported p = .006 · recomputed p = .006Recomputed adjusted odds ratio 1.87 (95% CI 1.20–2.91), reported p=0.006
“adjusted odds ratio 1.87, 95% confidence interval 1.20 to 2.91, P=0.006”
Taken as given: 1.20–2.91 is a two-sided 95% confidence interval for the adjusted odds ratio of 1.87, not a range, an IQR, or a different interval level; the adjusted odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.006 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.87, 1.2, 2.91, 1) - CONSISTENTreported p = .009 · recomputed p = .006Recomputed adjusted odds ratio 1.81 (95% CI 1.2–2.8), reported p=0.009
“adjusted odds ratio 1.81, 95% confidence interval 1.2 to 2.8, P=0.009”
Taken as given: 1.2–2.8 is a two-sided 95% confidence interval for the adjusted odds ratio of 1.81, not a range, an IQR, or a different interval level; the adjusted odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.009 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.81, 1.2, 2.8, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Recomputed adjusted odds ratio 3.07 (95% CI 1.71–5.50), reported p<0.001
“adjusted odds ratio 3.07, 95% confidence interval 1.71 to 5.50, P<0.001”
Taken as given: 1.71–5.50 is a two-sided 95% confidence interval for the adjusted odds ratio of 3.07, not a range, an IQR, or a different interval level; the adjusted odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p<0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(3.07, 1.71, 5.5, 1) - CONSISTENTreported p = .006 · recomputed p = .006Reviewers 1, 2Primary outcome odds ratio p-value
“adjusted odds ratio 1.87, 95% confidence interval 1.2 to 2.9, P=0.006”
Taken as given: The odds ratio is 1.87 with 95% CI 1.20 to 2.91.; The CI is two-sided at 95%.; The p-value is two-sided.Method: Recomputed two-sided p-value from the odds ratio and its 95% confidence interval using the normal approximation.How we recomputed it: pCI(1.87, 1.20, 2.91, 1) - CONSISTENTreported p = .005 · recomputed p = .004Reviewer 1Primary outcome absolute difference p-value
“adjusted absolute difference 15.0%, 95% confidence interval 4.6% to 25.0%, P=0.005”
Taken as given: The absolute difference is 15.0% with 95% CI 4.6% to 25.0%.; The CI is two-sided at 95%.; The p-value is two-sided.Method: Recomputed two-sided p-value from the absolute difference and its 95% confidence interval using the normal approximation.How we recomputed it: pCI(15.0, 4.6, 25.0, 0) - CONSISTENTreported p = .001 · recomputed p = <.001Reviewer 1Time to conversion mean difference p-value
“mean difference −22.9, 95% confidence interval −29.9 to −16.0, P<0.001”
Taken as given: The mean difference is -22.9 with 95% CI -29.9 to -16.0.; The CI is two-sided at 95%.; The p-value is two-sided.Method: Recomputed two-sided p-value from the mean difference and its 95% confidence interval using the normal approximation.How we recomputed it: pCI(-22.9, -29.9, -16.0, 0) - CONSISTENTreported p = .033 · recomputed p = .037Reviewers 1, 2Electrical cardioversion odds ratio p-value
“odds ratio 0.62, 95% confidence interval 0.39 to 0.96, P=0.033”
Taken as given: The odds ratio is 0.62 with 95% CI 0.39 to 0.96.; The CI is two-sided at 95%.; The p-value is two-sided.Method: Recomputed two-sided p-value from the odds ratio and its 95% confidence interval using the normal approximation.How we recomputed it: pCI(0.62, 0.39, 0.96, 1) - CONSISTENTreported p = .110 · recomputed p = .082Reviewer 2Secondary outcome: adverse event during or after infusion.
“Adverse event during or after infusion | 43 (25.0) | 31 (17.4) | 0.11”
Taken as given: The numbers 31 and 43 are the event counts in the vernakalant and procainamide groups, respectively.; The group totals are 178 and 172, respectively.; The p-value is from a chi-square test without continuity correction.Method: Pearson chi-square test on the 2x2 table (31, 147, 43, 129).How we recomputed it: pChi2x2(31, 147, 43, 129)
- lowinternal contradictionThe per-protocol analysis table (Table 3) reports n=342 for the conversion time analysis, but the per-protocol population is 342 (168+174). However, the ITT table (Table 2) also reports n=342 for the same row, which is inconsistent with the ITT population of 350.
“Conversion time (min) among randomised patients (n=342; start of infusion to conversion or censoring), restricted mean survival time (standard error)”
Table 2Find in source
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
5 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Vernakalant is superior to procainamide for conversion to sinus rhythm within 30 minutes.The primary outcome shows a statistically significant difference with a 15% absolute improvement, supported by the adjusted analysis.Evidence: Primary outcome: 62.4% vs 48.3%, adjusted absolute difference 15.0%, 95% CI 4.6% to 25.0%, P=0.005.
“For the primary outcome of conversion success, vernakalant was more effective (62.4% v 48.3%; adjusted absolute difference 15.0%, 95% confidence interval 4.6% to 25.0%, P=0.005; adjusted odds ratio 1.87, 95% confidence interval 1.2 to 2.9, P=0.006).”
AbstractFind in source - supportedReviewers 1, 2Vernakalant leads to faster time to conversion.The time to conversion was significantly shorter in the vernakalant group, supported by the mean difference and hazard ratio.Evidence: Time to conversion: 21.8 vs 44.7 minutes, mean difference -22.9, 95% CI -29.9 to -16.0, P<0.001.
“With vernakalant, time to conversion was faster (21.8 v 44.7 minutes; mean difference −22.9, 95% confidence interval −29.9 to −16.0, P<0.001)”
AbstractFind in source - supportedReviewers 1, 2Fewer patients in the vernakalant group required electrical cardioversion.The difference in electrical cardioversion rates was statistically significant.Evidence: Attempted electrical cardioversion: 33.7% vs 44.2%, odds ratio 0.62, 95% CI 0.39 to 0.96, P=0.033.
“fewer patients underwent attempted electrical cardioversion (33.7% v 44.2%; odds ratio 0.62, 95% confidence interval 0.39 to 0.96, P=0.033).”
AbstractFind in source - supportedReviewers 1, 2Vernakalant is particularly effective in patients younger than 70 years.The subgroup analysis shows a significant interaction, supporting the claim.Evidence: Subgroup analysis: patients <70 years, 73.3% vs 47.2%, adjusted odds ratio 3.1, 95% CI 1.7 to 5.5, P=0.001, interaction P=0.005.
“Subgroup analysis strongly favoured vernakalant for conversion in patients younger than 70 years (73.3% v 47.2%; adjusted odds ratio 3.1, 95% confidence interval 1.7 to 5.5, P=0.001, interaction P=0.005).”
AbstractFind in source - supportedReviewers 1, 2Vernakalant is safe and well tolerated.Adverse events were similar between groups and generally mild, supporting the safety claim.Evidence: Adverse events during or after infusion: 17.4% vs 25.0%, P=0.11; no significant differences in serious events.
“Adverse events during or after the drug infusion were similar in both groups (17.4% v 25.0%; absolute difference −7.6%, 95% confidence interval −16.1% to 0.95%, P=0.11; ).”
ResultsFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary outcome is conversion to sinus rhythm, a clinical outcome directly reflecting restoration of normal heart rhythm, which is a hard clinical endpoint in the context of acute atrial fibrillation. The trial also reports secondary outcomes such as time to conversion, need for electrical cardioversion, and discharge home, all clinically meaningful. No surrogate biomarker is used as the basis for the efficacy claim.
“The primary outcome was conversion to and maintenance of sinus rhythm for at least 30 minutes at any time after randomisation until 30 minutes after completion of the drug infusion.”
- ADEQUATEEffect sizeThe primary effect size is an absolute difference of 15.0% in conversion rate (62.4% vs 48.3%), which was predefined as the minimum clinically important difference. The effect is statistically significant and anchored to a clinically meaningful threshold. Additionally, time to conversion was faster by 22.9 minutes, and fewer patients required electrical cardioversion, supporting clinical meaningfulness.
“Based on a poll of investigators, an absolute difference of 15% was considered the minimum clinically important difference.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior trials (RAFF2) and observational studies, acknowledges the limitations of existing treatments, and clearly states the hypothesis and objective. The rationale links the need for a head-to-head comparison to the lack of direct evidence.
“In the recent RAFF2 trial, we showed good outcomes for acute atrial fibrillation managed with rhythm control in the emergency department.”
“We hypothesise that vernakalant has the advantages of a higher conversion rate, more rapid administration and onset, and fewer adverse events, and could therefore give clinicians an important alternative for pharmacological cardioversion of acute atrial fibrillation.”
“We are aware of no previous studies directly comparing intravenous procainamide with intravenous vernakalant.”
“In the recent RAFF2 trial, we showed good outcomes for acute atrial fibrillation managed with rhythm control in the emergency department.”
“We hypothesise that vernakalant has the advantages of a higher conversion rate, more rapid administration and onset, and fewer adverse events, and could therefore give clinicians an important alternative for pharmacological cardioversion of acute atrial fibrillation.”
“We are aware of no previous studies directly comparing intravenous procainamide with intravenous vernakalant.”
Randomization was computer-generated, stratified, and concealed. The open-label design is justified. Sample size calculation is provided. Inclusion/exclusion criteria are detailed. The analysis population (ITT and per-protocol) is defined, addressing outlier handling.
“The 1:1 allocation sequence was computer generated by a statistician at the Methods Center of the Ottawa Hospital Research Institute.”
“We chose an open label approach because both drugs are available for routine use in Canada and other countries, and we wanted to mirror actual clinical use.”
“A total of 340 patients (170 per group) achieves 80% power to detect an absolute difference in conversion rates of 15%.”
“The 1:1 allocation sequence was computer generated by a statistician at the Methods Center of the Ottawa Hospital Research Institute.”
“We chose an open label approach because both drugs are available for routine use in Canada and other countries, and we wanted to mirror actual clinical use.”
“A total of 340 patients (170 per group) achieves 80% power to detect an absolute difference in conversion rates of 15%.”
Age, sex, and health status (comorbidities) are reported in Table 1. Both sexes are enrolled, so sex justification is not applicable. Species/strain and housing are not applicable.
“Male | 114 (66.3) | 108 (60.7)”
“Age (years), mean (SD) | 62.4 (15.2) | 63.5 (15.0)”
The paper states approval from local research ethics boards and Clinical Trials Ontario. Informed consent is described (verbal or written). Regulatory compliance is implied through adherence to guidelines.
“The study had approval from the local research ethics boards and Clinical Trials Ontario.”
“Eligible patients were approached to provide verbal or written informed consent according to local research ethics board requirements.”
“The study had approval from the local research ethics boards and Clinical Trials Ontario.”
“Eligible patients were approached to provide verbal or written informed consent according to local research ethics board requirements.”
Both vernakalant and procainamide are named with doses and infusion protocols. No other biological or chemical resources are used. Software used for analysis is identified (SAS 9.4).
“Patients randomised to the vernakalant group received an initial infusion of 3 mg/kg over 10 minutes by a preprogrammed intravenous pump.”
“We conducted all analyses using Statistical Analysis Software (SAS) version 9.4.”
“Patients randomised to the vernakalant group received an initial infusion of 3 mg/kg over 10 minutes by a preprogrammed intravenous pump.”
“We conducted all analyses using Statistical Analysis Software (SAS) version 9.4.”
Tests are named (logistic regression, linear regression, Cox proportional hazards). Exact p-values are reported. Effect sizes with 95% CIs are provided. Software is identified. Data presentation includes per-group n and appropriate measures. Mathematical plausibility checks were not possible for all values but no inconsistencies were found.
“The primary analysis used multiple logistic regression analysis controlling for the stratification variables”
“P=0.005; adjusted odds ratio 1.87, 95% confidence interval 1.2 to 2.9, P=0.006”
“adjusted absolute difference 15.0%, 95% confidence interval 4.6% to 25.0%”
“The primary analysis used multiple logistic regression analysis controlling for the stratification variables”
“adjusted absolute difference 15.0%, 95% confidence interval 4.6% to 25.0%, P=0.005”
The data availability statement names a public repository (FRDR) with a DOI. The protocol is publicly available. No custom code is mentioned, so code sharing is not applicable.
“The authors will make the data available in a publicly accessible repository, the Canadian Federated Research Data Repository (FRDR: https://www.frdr-dfdr.ca/repo/dataset/9547ab29-882b-4c2a-a0e1-68ebb596d9f4 ).”
“The authors will make the data available in a publicly accessible repository, the Canadian Federated Research Data Repository (FRDR: https://www.frdr-dfdr.ca/repo/dataset/9547ab29-882b-4c2a-a0e1-68ebb596d9f4 ).”
The trial is registered (NCT04485195). CONSORT compliance is stated. All outcomes are reported. Limitations are discussed. Conclusions are proportional. Funding and COI are declared.
“Trial registration ClinicalTrials.gov NCT04485195”
“This manuscript is compliant with the CONSORT (consolidated standards of reporting trials) statement for reporting randomised trials.”
“Many patients refused to participate, mostly because they had a strong preference for procainamide or electrical cardioversion.”
“Trial registration ClinicalTrials.gov NCT04485195”
“This manuscript is compliant with the CONSORT (consolidated standards of reporting trials) statement for reporting randomised trials.”
“Funding: Peer reviewed grants received from the Canadian Institutes of Health Research and the Accelerating Clinical Trials (ACT) Consortium (Canada).”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 44 references by DOI: 38 verified — 6 no DOI (shown, not verified).
- NO DOIEmergency department management and 1-year outcomes of patients with atrial flutterNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICanadian Cardiovascular Society atrial fibrillation guidelines 2010: management of recent-onset atrial fibrillation and flutter in the emergency departmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOutcomes for emergency department patients with recent-onset atrial fibrillation and flutter treated in Canadian hospitalsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEmergency department services in Ontario 1993-2000No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEfficacy of agents for pharmacologic conversion of atrial fibrillation and subsequent maintenance of sinus rhythm: a meta-analysis of clinical trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRapid cardioversion of recent-onset atrial fibrillation in the emergency department with vernakalant: insights from the Multinational Spectrum RegistryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://www.frdr-dfdr.ca/repo/dataset/9547ab29-882b-4c2a-a0e1-68ebb596d9f4LIVEHTTP 200Resolved page looks like data.
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORconsistencyAbstract, Results“adjusted odds ratio 1.87, 95% confidence interval 1.2 to 2.9”→ Use consistent decimal places: 1.20 to 2.91The abstract uses one decimal place while the results section uses two.
- MINORclarityMethods, Data analysis“The primary analysis used multiple logistic regression analysis controlling for the stratification variables (age, first or repeat episode as fixed effects, and site as a random effect) and the following prespecified prognostic variables: sex, time from onset, and history of heart failure, all specified as fixed effects.”→ Clarify that age and first/repeat episode are also fixed effects, as written it is ambiguous.The sentence structure could be clearer.
- MINORtypoAuthor affiliations“Département de médecine familiale et de médicine d’urgence”→ Change 'médicine' to 'médecine'.Typographical error in French text.
- MINORconsistencyTable 2 and Table 3“Conversion time (min) among randomised patients (n=342; start of infusion to conversion or censoring), restricted mean survival time (standard error)”→ Ensure the n=342 is consistent with the per-protocol population (342) and clarify that this row is from the per-protocol analysis.The n=342 appears in both ITT and per-protocol tables, but the ITT population is 350. This may be a labeling inconsistency.
- MINORclarityData availability statement“Requests for secondary use of the data should be addressed to the corresponding author.”→ Clarify the process for data access requests, including any review criteria.The statement is adequate but could be more specific.
The published work is robust and well-reported. An informed reader should weigh the minor internal inconsistency in table labeling (n=342 in ITT table) and the open-label design as potential limitations, but these do not undermine the main conclusions. No erratum is warranted for the identified issues, though the authors may consider clarifying the table labeling.
- 1.HIGHreportingClarify the n=342 labeling in Table 2 (ITT analysis) to indicate it refers to the per-protocol population, or correct the population count to be consistent with the ITT population of 350.The internal inconsistency between the ITT population (350) and the n=342 reported in the ITT table could confuse readers and may warrant a correction.
- 2.MEDIUMcopyeditStandardize decimal places for the adjusted odds ratio in the Abstract (1.87, 95% CI 1.2 to 2.9) to match the Results section (e.g., 1.20 to 2.91).Inconsistent decimal precision between abstract and results is a minor copyedit issue that could be corrected for consistency.
- 3.MEDIUMcopyeditClarify the sentence in Methods, Data analysis, to explicitly state that age and first/repeat episode are fixed effects, as the current wording is ambiguous.Ambiguity in the description of the statistical model could lead to misinterpretation of the analysis.
- 4.MEDIUMcopyeditFix the typo in the French author affiliation: change 'médicine' to 'médecine'.Correcting the typographical error improves professionalism and accuracy.
- 5.MEDIUMdata codeAdd a statement about the availability of statistical analysis code, even if not applicable, to enhance reproducibility.Providing code or explicitly stating its unavailability helps readers assess reproducibility.
- 6.MEDIUMreportingReport the number of patients screened and excluded at each stage in the CONSORT flow diagram more explicitly in the text.A detailed CONSORT flow diagram is a key transparency element for randomized trials.
- 7.MEDIUMstatisticsReport exact p-values for all secondary outcomes in the tables, as some are only reported as thresholds.Exact p-values allow readers to assess the strength of evidence for secondary outcomes.
- 8.MEDIUMstatisticsAdd a sensitivity analysis for the primary outcome using multiple imputation for missing data, if any.Sensitivity analyses strengthen the robustness of the primary finding.
- 9.LOWreportingClarify the process for data access requests in the Data Availability Statement, including any review criteria.A more detailed data access procedure improves transparency and usability.
- 10.LOWreportingAdd a statement about the generalizability of the findings to other healthcare settings, given the open-label design.Discussing generalizability helps readers interpret the applicability of the results.
- 11.LOWstatisticsReport the intraclass correlation coefficient for the site random effect to assess clustering.Reporting the ICC provides insight into the degree of clustering by site.
- 12.LOWreportingAdd a pre-specified analysis plan for the cost-effectiveness analysis, which is mentioned but not detailed.Detailing the cost-effectiveness analysis plan improves transparency and completeness.
- 13.LOWreportingReport the number of patients who withdrew consent or were lost to follow-up in more detail.Detailed attrition information is important for assessing potential bias.
- 14.LOWstatisticsAdd a statement about the handling of missing data for the 30-day follow-up outcomes.Clarifying missing data handling for follow-up outcomes is important for interpretation.
- 15.LOWreportingReport the exact p-values for the subgroup analyses in the text, as some are only reported in the table.Exact p-values for subgroup analyses facilitate interpretation of subgroup effects.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.