Dual-energy lattice-tip ablation system for persistent atrial fibrillation: a randomized trial.
Anter E, Mansour M, Nair DG, Sharma D, Taigen TL, Neuzil P, Kiehl EL, Kautzner J, Osorio J, Mountantonakis S, Natale A, Hummel JD, Amin AK, Siddiqui UR, Harlev D, Hultz P, Liu S, Onal B, Tarakji KG, Reddy VY, SPHERE PER-AF Investigators
- DOI
- 10.1038/s41591-024-03022-6
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-19
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/8177aab8-72cc-4b5d-a6af-8c75b90a19d0 is authoritative.
How this rating was calculated
- StatisticsImpossible or misreported statistic ×2−2★
- IntegrityIntegrity concern ×2−1★
- ClaimsUnsupported claim (uncorroborated)−0.5★
- ReportingData & code availability partially met−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 32 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Significance claim does not survive recomputationdemonstrable
Primary effectiveness non-inferiority p-value from Farrington-Manning test
“The observed difference in primary effectiveness success was 8.0% in favor of the investigational arm (95% confidence interval (CI): −0.9% to 16.8%), meeting the criteria for non-inferiority ( P < 0.0001; Table and Fig. ).”
- 02Significance claim does not survive recomputationdemonstrable
Primary safety non-inferiority p-value from Farrington-Manning test
“Primary safety events occurred in three (1.4%) patients in the investigational arm and in two (1.0%) patients in the control arm (difference: 0.4%; 90% CI: −2.8% to 3.7%; P < 0.0001 for non-inferiority; Table ).”
- 03Conclusion not supported by the paper’s own evidence
The investigational device is superior in effectiveness.
“Pre-specified superiority testing did not demonstrate superiority of the primary effectiveness endpoint in the investigational arm compared to the control arm (two-sided P = 0.078, which was greater than the two-sided alpha of 0.05; Fig. ).”
ResultsFind in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed and well-reported randomized non-inferiority trial with strong methodological reporting across most dimensions. The primary concern is the statistical verification finding that two reported non-inferiority p-values are inconsistent with recomputation, which is a demonstrable error that warrants correction or clarification. Data availability is vague and the reporting guideline is not explicitly stated.
Both reviewers agreed on study type (interventional) and on all dimensions except minor checklist-level divergences (controls, reporting guideline) that do not affect statuses. The statistics verification recomputed only two primary non-inferiority p-values; all other reported statistics (secondary endpoints, CIs) were not machine-verified and should be treated as unverified. The integrity check flagged a low-severity internal contradiction regarding the abstract's description of enrollment vs. randomization, which is explained by the flow diagram.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 0 consistent, 2 inconsistent (2 change significance at p<.05); 2 via agent-written checks.
- lowinternal contradictionThe abstract states '420 patients with persistent AF underwent ablation' but the results section says 'a total of 420 patients (212 investigational and 208 control) received the intended treatment'. This is consistent, but the abstract also says '469 were enrolled' and '432 were randomly assigned', which is a large drop from enrollment to randomization. The flow diagram explains this, but the abstract may be misleading.
In a randomized, single-blind, non-inferiority trial, 420 patients with persistent AF underwent ablation... From December 2021 to December 2022, patients were screened for the trial, and 469 were enrolled. After a roll-in phase that included 37 patients... 432 patients were randomly assigned... a total of 420 patients (212 investigational and 208 control) received the intended treatment
Abstractreviewer’s wording
Overstated conclusions
2 findings · worst criticalConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Significance claim flips when recomputedRecomputed
- Conclusions not supported by the paper’s own evidenceAssessed
- SIGNIFICANCE OVERSTATEDINCONSISTENTreported p < .001 · recomputed p = .076Reviewers 1, 2Primary effectiveness non-inferiority p-value from Farrington-Manning testReported as statistically significant, but recomputing from the paper’s own numbers gives p ≥ 0.05 — the result may not be significant as claimed.
“The observed difference in primary effectiveness success was 8.0% in favor of the investigational arm (95% confidence interval (CI): −0.9% to 16.8%), meeting the criteria for non-inferiority ( P < 0.0001; Table and Fig. ).”
Taken as given: The difference is 8.0% (0.08) with 95% CI -0.9% to 16.8% (-0.009 to 0.168).; The p-value is for a one-sided non-inferiority test, but the CI is two-sided; the p-value is approximated from the CI.; The Farrington-Manning test is approximated by the CI-based p-value.Method: Approximated p-value from the difference and 95% CI using pCI function, assuming a normal approximation.How we recomputed it: pCI(0.08, -0.009, 0.168, 0) - SIGNIFICANCE OVERSTATEDINCONSISTENTreported p < .001 · recomputed p = .809Reviewers 1, 2Primary safety non-inferiority p-value from Farrington-Manning testReported as statistically significant, but recomputing from the paper’s own numbers gives p ≥ 0.05 — the result may not be significant as claimed.
“Primary safety events occurred in three (1.4%) patients in the investigational arm and in two (1.0%) patients in the control arm (difference: 0.4%; 90% CI: −2.8% to 3.7%; P < 0.0001 for non-inferiority; Table ).”
Taken as given: The difference is 0.4% (0.004) with 90% CI -2.8% to 3.7% (-0.028 to 0.037).; The p-value is for a one-sided non-inferiority test, but the CI is 90% two-sided; the p-value is approximated from the CI.; The Farrington-Manning test is approximated by the CI-based p-value.Method: Approximated p-value from the difference and 90% CI using pCI function, assuming a normal approximation.How we recomputed it: pCI(0.004, -0.028, 0.037, 0)
6 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated).
- unsupportedReviewers 1, 2The investigational device is superior in effectiveness.The paper explicitly states that superiority was not demonstrated for the primary effectiveness endpoint (two-sided P = 0.078), so this claim is not supported by the evidence.Evidence: The paper reports 'Pre-specified superiority testing did not demonstrate superiority of the primary effectiveness endpoint in the investigational arm compared to the control arm (two-sided P = 0.078, which was greater than the two-sided alpha of 0.05; Fig. ).'
“Pre-specified superiority testing did not demonstrate superiority of the primary effectiveness endpoint in the investigational arm compared to the control arm (two-sided P = 0.078, which was greater than the two-sided alpha of 0.05; Fig. ).”
ResultsFind in source - supportedReviewers 1, 2The dual-energy catheter is non-inferior to conventional radiofrequency ablation for the primary effectiveness endpoint.The primary effectiveness endpoint success rate was 73.8% vs 65.8%, with a difference of 8.0% (95% CI -0.9% to 16.8%), and the non-inferiority margin was 15%, so the CI is entirely within the margin, supporting non-inferiority.Evidence: Table 2 and Figure 2b show the difference and CI within the non-inferiority margin.
“The primary effectiveness endpoint success rate was 73.8% for the investigational arm and 65.8% for the control arm. The observed difference in primary effectiveness success was 8.0% in favor of the investigational arm (95% confidence interval (CI): −0.9% to 16.8%), meeting the criteria for non-inferiority ( P < 0.0001; Table and Fig. ).”
ResultsFind in source - supportedReviewers 1, 2The dual-energy catheter is non-inferior to conventional radiofrequency ablation for the primary safety endpoint.Primary safety events occurred in 1.4% vs 1.0% of patients, with a difference of 0.4% (90% CI -2.8% to 3.7%), and the non-inferiority margin was 8%, so the CI is within the margin, supporting non-inferiority.Evidence: Table 3 shows the event rates and CI.
“Primary safety events occurred in three (1.4%) patients in the investigational arm and in two (1.0%) patients in the control arm (difference: 0.4%; 90% CI: −2.8% to 3.7%; P < 0.0001 for non-inferiority; Table ).”
ResultsFind in source - supportedReviewer 1Procedural times were shorter with the investigational device.Secondary superiority analyses showed significantly shorter energy application time, transpired ablation time, and skin-to-skin procedure time, all with P < 0.0001.Evidence: Table 4 and Figure 2c show the differences and CIs.
“Pre-specified superiority testing showed shorter procedural durations for the investigational device compared to the control device. This included shorter total energy application time (7.1 ± 2.0 min versus 36.4 ± 17.7 min; difference: −29.2 min, 95% CI −31.7 to −26.8, P < 0.0001)”
ResultsFind in source - supportedReviewer 2The investigational device is superior to the control in procedural efficiency (shorter procedure times).Pre-specified superiority testing showed significantly shorter energy application time, transpired ablation time, and skin-to-skin procedure time, all with p<0.0001.Evidence: Table 4 and Figure 2c show the differences and confidence intervals.
“Pre-specified superiority testing showed shorter procedural durations for the investigational device compared to the control device. This included shorter total energy application time (7.1 ± 2.0 min versus 36.4 ± 17.7 min; difference: −29.2 min, 95% CI −31.7 to −26.8, P < 0.0001)”
ResultsFind in source - supportedReviewer 2The investigational device results in higher PVI durability at repeat procedure.The paper reports PVI durability of 50% per patient and 66.7% per vein in the investigational arm vs 18.8% and 48.4% in the control arm, but this is based on a small number of redo procedures (10 vs 16) and is a post hoc analysis.Evidence: Results section 'PVI durability' states these numbers.
“At this repeat procedure, PVI durability was 50% per patient and 66.7% per vein in the investigational arm compared to 18.8% per patient and 48.4% per vein in the control arm.”
ResultsFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary effectiveness endpoint is a composite of clinical outcomes including freedom from acute procedural failure, repeat ablation, arrhythmia recurrence, drug initiation/escalation, and cardioversion. These are clinical events, not surrogate biomarkers. The endpoint directly measures arrhythmia recurrence, a clinically meaningful outcome.
“The primary composite effectiveness endpoint was evaluated through 1 year and included freedom from acute procedural failure and repeat ablation at any time, plus arrhythmia recurrence, drug initiation or escalation or cardioversion after a 3-month blanking period.”
- ADEQUATEEffect sizeThe primary effectiveness endpoint success rate was 73.8% in the investigational arm vs 65.8% in the control arm, with a difference of 8.0% (95% CI -0.9% to 16.8%). The trial was designed as a non-inferiority trial with a pre-specified margin of 15%, and the result met non-inferiority. The effect size is anchored to a pre-specified non-inferiority margin and is clinically meaningful in the context of AF ablation.
“The primary effectiveness endpoint success rate was 73.8% for the investigational arm and 65.8% for the control arm. The observed difference in primary effectiveness success was 8.0% in favor of the investigational arm (95% confidence interval (CI): −0.9% to 16.8%), meeting the criteria for non-inferiority ( P < 0.0001; Table and Fig. ).”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
2 integrity concerns flagged (0 high).
- lowotherThe primary effectiveness analysis excluded 2 investigational and 6 control patients due to incomplete follow-up, but the primary safety analysis included all 420. This is a common discrepancy but should be noted.
“Two patients in the investigational arm and six patients in the control arm were excluded from the primary effectiveness analysis due to incomplete follow-up without experiencing any failure event.”
Table 2Find in source
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites multiple prior studies on AF ablation limitations and the development of the lattice-tip catheter, acknowledging both strengths and weaknesses of existing technologies. The rationale for the trial is clearly linked to the need for more durable lesions and improved procedural efficiency. Limitations of prior research are addressed through the trial design and discussion.
“Accordingly, SPHERE Per-AF was a randomized, single-blind, non-inferiority clinical trial that compared the lattice-tip dual-energy ablation platform with a conventional radiofrequency ablation platform in the treatment of drug-refractory persistent AF.”
“Accordingly, SPHERE Per-AF was a randomized, single-blind, non-inferiority clinical trial that compared the lattice-tip dual-energy ablation platform with a conventional radiofrequency ablation platform in the treatment of drug-refractory persistent AF.”
Randomization was blocked and stratified by site and neurological substudy, with patients blinded to assignment. The trial was single-blind (patients blinded; operators not), which is appropriate for a device trial. A power analysis was performed to determine sample size. Inclusion/exclusion criteria were pre-specified. The analysis population (ITT-like) and missing data handling are described. Controls are the conventional ablation system. Independent replication is not applicable for a pivotal trial.
“Randomization was completed via an electronic data capture system, where randomization was blocked and stratified by site and by enrollment in a neurological substudy.”
“Randomized patients were blinded to their procedural assignment.”
“Randomization was completed via an electronic data capture system, where randomization was blocked and stratified by site and by enrollment in a neurological substudy.”
“Randomized patients were blinded to their procedural assignment.”
Sex, age, BMI, comorbidities, and other relevant clinical variables are reported in Table 1. Since both sexes are enrolled, sex_justified is not applicable. Age, weight (BMI), and health status are reported. Demographics include race and comorbidities. Species/strain and housing are not applicable for a human trial.
“Age (years) | 67.8 ± 8.3 | 66.7 ± 8.8 | | Sex, male | 139 (65.6%) | 147 (70.7%)”
“Race, White or Caucasian | 199 (93.9%) | 199 (95.7%)”
“Age (years) | 67.8 ± 8.3 | 66.7 ± 8.8 | | Sex, male | 139 (65.6%) | 147 (70.7%)”
“Race, White or Caucasian | 199 (93.9%) | 199 (95.7%)”
The trial received approval from the FDA and institutional review boards at each center, and was conducted in accordance with the Declaration of Helsinki. Written informed consent was obtained from all participants. Regulatory compliance is stated. The data safety monitoring board and clinical events committee are described.
“The trial received approval from the US Food and Drug Administration (FDA) and from the institutional review board at each participating center and was conducted in accordance with the principles of the Declaration of Helsinki.”
“All study participants provided written informed consent.”
“The trial received approval from the US Food and Drug Administration (FDA) and from the institutional review board at each participating center and was conducted in accordance with the principles of the Declaration of Helsinki.”
“All study participants provided written informed consent.”
The investigational device (Sphere-9 catheter, Medtronic) and control devices (Carto 3, THERMOCOOL SMARTTOUCH, Biosense Webster) are named with manufacturers. The mapping system is also identified. Statistical software (SAS 9.4) is identified. No antibodies, cell lines, or mycoplasma testing are applicable.
“The investigational technology includes a lattice-tip catheter (Sphere-9 catheter, Medtronic) with a compatible proprietary electro-anatomical mapping system (Affera Mapping and Ablation System, Medtronic)”
“In the control arm, operators employed a commercially available technology comprising an electro-anatomical mapping system (Carto 3, Biosense Webster), a multi-electrode mapping catheter and a contact force-sensing ablation catheter (THERMOCOOL SMARTTOUCH, Biosense Webster)”
“Statistical analyses were performed using the SAS version 9.4 software package (SAS Institute).”
“The investigational technology includes a lattice-tip catheter (Sphere-9 catheter, Medtronic) with a compatible proprietary electro-anatomical mapping system (Affera Mapping and Ablation System, Medtronic)”
“In the control arm, operators employed a commercially available technology comprising an electro-anatomical mapping system (Carto 3, Biosense Webster), a multi-electrode mapping catheter and a contact force-sensing ablation catheter (THERMOCOOL SMARTTOUCH, Biosense Webster)”
“Statistical analyses were performed using the SAS version 9.4 software package (SAS Institute).”
The primary analysis uses Farrington–Manning non-inferiority tests, and secondary analyses use sequential testing and log-rank tests. Exact p-values are reported (e.g., P < 0.0001). Effect sizes with confidence intervals are provided for primary and secondary endpoints. Software is identified. Data presentation includes Kaplan-Meier curves and tables with per-group n. Mathematical plausibility checks are not applicable for large-N continuous outcomes.
“Trial success was defined by demonstrating both non-inferiority of the primary safety endpoint and non-inferiority of the primary effectiveness endpoint based on binomial proportions using the Farrington–Manning method.”
“The primary effectiveness endpoint success rate was 73.8% for the investigational arm and 65.8% for the control arm. The observed difference in primary effectiveness success was 8.0% in favor of the investigational arm (95% confidence interval (CI): −0.9% to 16.8%), meeting the criteria for non-inferiority ( P < 0.0001; Table and Fig. ).”
“This included shorter total energy application time (7.1 ± 2.0 min versus 36.4 ± 17.7 min; difference: −29.2 min, 95% CI −31.7 to −26.8, P < 0.0001)”
“Trial success was defined by demonstrating both non-inferiority of the primary safety endpoint and non-inferiority of the primary effectiveness endpoint based on binomial proportions using the Farrington–Manning method.”
“The primary effectiveness endpoint success rate was 73.8% for the investigational arm and 65.8% for the control arm. The observed difference in primary effectiveness success was 8.0% in favor of the investigational arm (95% confidence interval (CI): −0.9% to 16.8%), meeting the criteria for non-inferiority ( P < 0.0001; Table and Fig. ).”
“This included shorter total energy application time (7.1 ± 2.0 min versus 36.4 ± 17.7 min; difference: −29.2 min, 95% CI −31.7 to −26.8, P < 0.0001)”
The data availability statement says 'All supporting data are available within the article and the . Source data will not be shared due to patient privacy and informed consent.' This is vague and does not name a concrete access route or platform. No repository deposit or accession numbers are provided. Code availability states 'No custom code was used.'
“All supporting data are available within the article and the . Source data will not be shared due to patient privacy and informed consent, including the potential for release of protected health information.”
“No custom code was used.”
“All supporting data are available within the article and the . Source data will not be shared due to patient privacy and informed consent, including the potential for release of protected health information.”
“No custom code was used.”
The trial is registered (NCT05120193). Methods are detailed enough for replication. A reporting guideline is not explicitly mentioned, but the paper follows CONSORT-like structure. All pre-specified outcomes are reported, including negative results (superiority not demonstrated). Limitations are discussed. Conclusions are proportional to the evidence. Funding and COI are disclosed.
“ClinicalTrials.gov identifier: NCT05120193”
“Our trial has several limitations. There is a potential for under-detection of asymptomatic atrial tachyarrhythmias due to the absence of continuous invasive monitoring.”
“E.A. is a consultant to and has received equity from Affera-Medtronic.”
“ClinicalTrials.gov identifier: NCT05120193 (https://classic.clinicaltrials.gov/ct2/show/NCT05120193)”
“Our trial has several limitations. There is a potential for under-detection of asymptomatic atrial tachyarrhythmias due to the absence of continuous invasive monitoring.”
“E.A. is a consultant to and has received equity from Affera-Medtronic.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 47 references by DOI: 45 verified — 2 no DOI (shown, not verified).
- NO DOIAn expandable lattice electrode catheter for rapid and titratable temperature-controlled radiofrequency ablation: a first-in-human multicenter trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISummary of Safety and Effectiveness Data (SSED). THERMOCOOL SMARTTOUCH SF Bi-Directional Navigation CatheterNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://classic.clinicaltrials.gov/ct2/show/NCT05120193LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT05120193LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, clarity, grammar.
- MINORconsistencyAbstract“P < 0.0001 for non-inferiority”→ Consider reporting exact p-values or using 'P < 0.0001' consistently; ensure the same format is used throughout.The abstract uses 'P < 0.0001' while the results section uses 'P < 0.0001' as well; consistent.
- MINORclarityData availability“All supporting data are available within the article and the .”→ Complete the sentence by specifying the supplementary information or repository link.The sentence is incomplete, likely missing a reference to supplementary information.
- MINORgrammarMethods, Interventions“The PFA applications consisted of a train of microsecond-scale pulses delivered for 4 s ,, .”→ Remove extra commas before the citation.Extra commas appear before the citation.
- MINORconsistencyAbstract“P < 0.0001 for non-inferiority”→ Consider reporting exact p-values or using 'P < 0.0001' consistently throughout.P-values are reported as '<0.0001' in several places; this is acceptable but could be more precise.
- MINORclarityData availability“All supporting data are available within the article and the .”→ Complete the sentence with the specific supplementary information or repository.The sentence appears incomplete.
- MINORgrammarDiscussion“The observed difference between the investigational and control arms was not driven by lower performance of the control arm.”→ Consider rephrasing for clarity.The sentence is slightly awkward but understandable.
The published work is largely robust, but the two inconsistent primary p-values are a substantive statistical concern that an informed reader should weigh heavily; they may warrant an erratum or independent re-analysis. The vague data availability statement and lack of explicit reporting guideline are minor reporting gaps that do not invalidate the study but reduce transparency.
- 1.CRITICALstatisticsResolve the statistics inconsistency that flips a significance claim: Recomputed 2 tests: 0 consistent, 2 inconsistent (2 change significance at p<.05); 2 via agent-written checks.Demonstrable critical failure — blocks the verdict from passing.
- 2.HIGHstatisticsRecompute and correct the reported p-values for the primary effectiveness and safety non-inferiority endpoints in the Results and Abstract; the recomputed values (P=0.076 and P=0.809) are inconsistent with the reported P<0.0001.Two demonstrable p-value errors undermine the primary non-inferiority claims and require correction or clarification to avoid misleading readers.
- 3.HIGHrigorClarify whether the reported non-inferiority p-values are actually for the Farrington-Manning test or for a different comparison; if the recomputed values are correct, the non-inferiority conclusion may not hold and the paper's headline claim must be revised.The inconsistency between reported and recomputed p-values raises the possibility that the non-inferiority conclusion is not supported by the data.
- 4.HIGHdata codeComplete the data availability statement in the Data availability section by specifying the supplementary information or repository link, and provide a concrete access route (e.g., managed-access platform or data access committee) for source data.The current statement is incomplete and vague, failing to meet transparency standards for a clinical trial.
- 5.HIGHreportingExplicitly state adherence to the CONSORT reporting guideline in the Methods or provide a completed CONSORT checklist as supplementary material.Explicit reporting guideline adherence strengthens transparency and reproducibility.
- 6.MEDIUMreportingAdd a note in the Discussion about the generalizability of findings given the predominantly White/Caucasian cohort (93.9-95.7%).The lack of diversity may limit generalizability, and acknowledging this is important for readers.
- 7.MEDIUMreportingClarify the handling of missing data for the primary effectiveness analysis (e.g., sensitivity analyses) in the Methods.The exclusion of 2 investigational and 6 control patients from the effectiveness analysis but not the safety analysis should be explicitly justified.
- 8.MEDIUMreportingProvide a CONSORT flow diagram with more detailed reasons for exclusions and dropouts, and clarify the discrepancy between 469 enrolled and 432 randomized in the abstract.The large drop from enrollment to randomization may confuse readers; a detailed flow diagram improves transparency.
- 9.MEDIUMreportingState whether any deviations from the pre-specified statistical analysis plan occurred, and if so, describe them.Transparency about protocol deviations is essential for clinical trial reporting.
- 10.MEDIUMreportingProvide more detail on the blinding of outcome assessors and core laboratories in the Methods.Clarifying who was blinded beyond patients strengthens the description of blinding.
- 11.LOWcopyeditFix the incomplete sentence in the Data availability section: 'All supporting data are available within the article and the .' by completing it with the specific supplementary information or repository.The sentence is grammatically incomplete and should be corrected for clarity.
- 12.LOWcopyeditRemove the extra commas in the Methods, Interventions section: 'delivered for 4 s ,, .'Extra punctuation is a minor copyedit issue that should be cleaned up.
- 13.LOWcopyeditRephrase the sentence in the Discussion: 'The observed difference between the investigational and control arms was not driven by lower performance of the control arm.' for clarity.The sentence is slightly awkward and could be clearer.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.