Stepwise dual antiplatelet therapy de-escalation in patients after drug coated balloon angioplasty (REC-CAGEFREE II): multicentre, randomised, open label, assessor blind, non-inferiority trial.
Gao C, Zhu B, Ouyang F, Wen S, Xu Y, Jia W, Yang P, He Y, Zhong Y, Zhou Y, Guo Z, Shen G, Ma L, Xu L, Xue Y, Hu T, Wang Q, Liu Y, Zhang R, Liu J, Jiang Z, Xia J, Garg S, van Geuns RJ, Capodanno D, Onuma Y, Wang D, Serruys P, Tao L, REC-CAGEFREE II Investigators
- DOI
- 10.1136/bmj-2024-082945
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/024adbf9-39f1-4519-b75f-b32ee3a9c91f is authoritative.
How this rating was calculated
- StatisticsImpossible or misreported statistic−1★
- IntegrityIntegrity concern ×2−1★
- ReportingData & code availability partially met−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- 01Significance claim does not survive recomputationdemonstrable
Primary endpoint non-inferiority p-value from reported difference and CI
“The 0.36% difference in the cumulative event rate and the upper boundary of the one sided 95% CI 2.47% met the prespecified criteria of 3.2% for non-inferiority (P non-inferiority =0.013, and ).”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed and transparently reported randomised non-inferiority trial with strong methodology, clear ethical approvals, and comprehensive reporting. The main weaknesses are a vague data-sharing statement and a potential inconsistency in the primary non-inferiority p-value that warrants verification.
Both reviewers agreed on study type (interventional) and on all dimensions except statistical analysis, where the deterministic recomputation found an inconsistency that overrides the reviewers' pass. The statistics component only checked a subset of reported tests; other statistics remain unverified.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 2 consistent, 1 inconsistent (1 change significance at p<.05); 3 via agent-written checks.
- CONSISTENTreported p = .008 · recomputed p = .008Reviewers 1, 2BARC type 3 or 5 bleeding p-value from reported difference and CI
“BARC type 3 or 5 bleeding occurred in four versus 16 participants (0.4% v 1.6%, difference −1.19% (95% CI −2.07% to −0.31%), P=0.008)”
Taken as given: The difference is -1.19 percentage points.; The 95% CI is two-sided.; The p-value is two-sided.Method: Recomputed p-value from the reported difference and 95% CI using the pCI function.How we recomputed it: pCI(-1.19, -2.07, -0.31, 0) - CONSISTENTreported p = .004 · recomputed p = .004Reviewer 2Win ratio p-value
“win ratio 1.43 (95% CI 1.12 to 1.83), P=0.004”
Taken as given: The win ratio is 1.43.; The 95% CI is two-sided and on the log scale.; The p-value is two-sided.Method: Using pCI with log=1 for a ratio to derive a two-sided p.How we recomputed it: pCI(1.43, 1.12, 1.83, 1)
- lowinternal contradictionThe abstract reports 74.9% men, while Table 1 shows 74.6% in the stepwise group and 75.3% in the standard group; the overall percentage is consistent with the weighted average, so this is not a contradiction.
“74.9% were men”
Table 1Find in source - lowinternal contradictionThe abstract reports 20.6% at high bleeding risk, while Table 1 shows 21.0% and 20.2% in the two groups; the overall percentage is plausible but not exactly matching.
20.6% were at high bleeding risk (Abstract); 197/936 (21.0) and 190/939 (20.2) (Table 1)
Table 1reviewer’s wording
Overstated conclusions
2 findings · worst criticalConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Significance claim flips when recomputedRecomputed
- Conclusions only partially backed by the presented evidenceAssessed
- SIGNIFICANCE OVERSTATEDINCONSISTENTreported p = .013 · recomputed p = .738Reviewers 1, 2Primary endpoint non-inferiority p-value from reported difference and CIReported as statistically significant, but recomputing from the paper’s own numbers gives p ≥ 0.05 — the result may not be significant as claimed.
“The 0.36% difference in the cumulative event rate and the upper boundary of the one sided 95% CI 2.47% met the prespecified criteria of 3.2% for non-inferiority (P non-inferiority =0.013, and ).”
Taken as given: The difference is 0.36 percentage points.; The 95% CI is two-sided and symmetric.; The p-value is for a one-sided non-inferiority test.Method: Recomputed p-value from the reported difference and 95% CI using the pCI function, assuming a normal approximation.How we recomputed it: pCI(0.36, -1.75, 2.47, 0)
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 2The results are generalisable to the broader DCB-treated population.The study population is limited to Chinese patients and those with paclitaxel-coated balloons, and the authors acknowledge this limitation.Evidence: The study was conducted only in China, and only paclitaxel-coated balloons were used.
“Finally, this study was only conducted in China with an East Asian population and therefore extrapolating these results to other ethnic groups warrants further investigation.”
LimitationsFind in source - supportedReviewers 1, 2Stepwise DAPT de-escalation is non-inferior to standard 12-month DAPT for net adverse clinical events.The primary endpoint result meets the pre-specified non-inferiority margin, and the sensitivity analyses support this conclusion.Evidence: Primary endpoint occurred in 87 (8.9%) vs 84 (8.6%), difference 0.36%, upper 95% CI 2.47%, P non-inferiority =0.013.
“At 12 months, the primary endpoint occurred in 87 (8.9%) participants in the stepwise de-escalation group and 84 (8.6%) in the standard group (difference 0.36%; upper boundary of the one sided 95% CI 2.47%; P non-inferiority =0.013).”
AbstractFind in source - supportedReviewers 1, 2Stepwise DAPT de-escalation reduces BARC type 3 or 5 bleeding compared with standard DAPT.The difference is statistically significant and the CI excludes zero.Evidence: BARC type 3 or 5 bleeding occurred in 4 vs 16 participants (0.4% v 1.6%, difference −1.19% (95% CI −2.07% to −0.31%), P=0.008).
“BARC type 3 or 5 bleeding occurred in four versus 16 participants (0.4% v 1.6%, difference −1.19% (95% CI −2.07% to −0.31%), P=0.008)”
AbstractFind in source - supportedReviewers 1, 2Stepwise DAPT de-escalation is associated with more wins in the hierarchical composite endpoint.The win ratio analysis shows a statistically significant benefit, though the clinical significance is modest.Evidence: Win ratio 1.43 (95% CI 1.12 to 1.83), P=0.004.
“win ratio 1.43 (95% CI 1.12 to 1.83), P=0.004”
AbstractFind in source - supportedReviewer 1The results support the use of stepwise DAPT de-escalation as a viable option for DCB-treated ACS patients.The conclusion is consistent with the non-inferiority finding and the secondary endpoint results, though the open-label design and single-country population are limitations.Evidence: Non-inferiority met; secondary endpoints show reduced bleeding without significant increase in ischaemic events.
“Among participants with acute coronary syndrome who could be treated by drug coated balloons exclusively, a stepwise DAPT de-escalation was non-inferior to 12 month DAPT for net adverse clinical events.”
ConclusionFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary endpoint is a composite of hard clinical outcomes (all cause death, stroke, myocardial infarction, revascularisation, and BARC type 3 or 5 bleeding), not a surrogate biomarker. The trial is a non-inferiority study comparing two antiplatelet strategies, and the primary endpoint directly measures clinical events.
“The primary endpoint was net adverse clinical events (all cause death, stroke, myocardial infarction, revascularisation, and Bleeding Academic Research Consortium (BARC) type 3 or 5 bleeding) at 12 months in the intention-to-treat population.”
- ADEQUATEEffect sizeThe primary endpoint shows non-inferiority with a difference of 0.36% and an upper 95% CI of 2.47%, below the prespecified margin of 3.2%. The effect is statistically supported and anchored to a clinically meaningful non-inferiority margin. Secondary endpoints show significant reductions in bleeding events, with effect sizes that are clinically relevant.
“At 12 months, the primary endpoint occurred in 87 (8.9%) participants in the stepwise de-escalation group and 84 (8.6%) in the standard group (difference 0.36%; upper boundary of the one sided 95% CI 2.47%; P non-inferiority =0.013).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The paper cites prior trials (BASKET-SMALL 2, DEBUT, AGENT IDE) and real-world registries to establish the premise that DCB-treated patients may need less intense antiplatelet therapy. It explicitly states that randomised data on optimal DAPT for DCB are lacking, providing a logical rationale for the study. Limitations of prior work are indirectly addressed by noting the absence of dedicated trials, though not deeply critiqued.
“Among the randomised studies investigating DCBs, such as the BASKET-SMALL 2 trial for de novo small-vessel disease, the DEBUT trial for patients with high bleeding risk, and the AGENT IDE trial for in-stent restenosis, nearly half of the participants had acute coronary syndrome.”
“randomised data investigating the optimal DAPT regimen for the patients receiving DCB is lacking.”
“Patients who receive exclusive treatment with DCBs may have the theoretical advantage of adopting a low intensity antiplatelet regimen because of the absence of a metallic scaffold and polymer inside the coronary artery, as well as the shorter local retention of the anti-proliferative drug.”
“Among the randomised studies investigating DCBs, such as the BASKET-SMALL 2 trial for de novo small-vessel disease, the DEBUT trial for patients with high bleeding risk, and the AGENT IDE trial for in-stent restenosis, nearly half of the participants had acute coronary syndrome.”
“However, despite extensive research on the optimal antiplatelet strategy for patients with acute coronary syndrome treated with drug eluting stents, randomised data investigating the optimal DAPT regimen for the patients receiving DCB is lacking.”
“We aimed to evaluate a stepwise DAPT de-escalation strategy compared with standard 12 months DAPT with respect to clinical outcomes, including both ischaemic and bleeding events.”
Randomisation used a web-based centralised system with computer-generated dynamic permuted blocks, stratified by site and lesion type. Blinding is described: patients and investigators were not masked, but the clinical event committee and statisticians were masked. A priori power analysis is reported with assumptions and non-inferiority margin. Inclusion/exclusion criteria are pre-specified and listed in the appendix. Outlier handling is addressed through ITT and per-protocol analyses, and missing data are handled by censoring at last contact. Controls are inherent in the comparator arm (standard DAPT). Independent replication is not applicable for a single pivotal trial.
“Randomisation sequences were computer generated with the dynamic permuted block method, with block sizes of two or four, and stratified by site and the type of lesion being treated (de novo or in-stent restenosis).”
“members of the independent clinical event committee who adjudicated the endpoints and statisticians who developed the statistical programmes were masked to treatment allocation.”
“Considering an anticipated 5% patient attrition rate, 1908 patients were required for the study to have 80% power to show non-inferiority with a 5% one sided type I error rate.”
“Randomisation sequences were computer generated with the dynamic permuted block method, with block sizes of two or four, and stratified by site and the type of lesion being treated (de novo or in-stent restenosis).”
“Patients and the investigators were not masked to treatment allocation; however, members of the independent clinical event committee who adjudicated the endpoints and statisticians who developed the statistical programmes were masked to treatment allocation.”
“Considering an anticipated 5% patient attrition rate, 1908 patients were required for the study to have 80% power to show non-inferiority with a 5% one sided type I error rate.”
The paper reports sex (74.9% men), age (mean 59.2 years), and extensive demographics and comorbidities in Table 1. Age and health status are reported. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“Age, years; mean (SD) | 59.4 (10.7) | 59.0 (11.0) | | Sex: | | Female | 248/975 (25.4) | 240/973 (24.7) | | Male | 727/975 (74.6) | 733/973 (75.3)”
“Overall, the mean age of patients was 59.2 years; 74.9% of the patients were men, 30.5% had diabetes, 8.8% had history of a stroke, 13.2% had history of a myocardial infarction, 32.2% had history of a percutaneous coronary intervention, and 20.6% were defined as at high bleeding risk”
“Overall, the mean age of patients was 59.2 years; 74.9% of the patients were men”
“Overall, the mean age of patients was 59.2 years; 74.9% of the patients were men, 30.5% had diabetes, 8.8% had history of a stroke, 13.2% had history of a myocardial infarction, 32.2% had history of a percutaneous coronary intervention, and 20.6% were defined as at high bleeding risk”
“Age, years; mean (SD) | 59.4 (10.7) | 59.0 (11.0)”
The protocol was approved by the ethics committee of Xijing Hospital (ID: KY20212080-F-1) and responsible ethics committees in all participating centres. Written informed consent was obtained from all patients. The trial was conducted in accordance with the Declaration of Helsinki and Good Clinical Practice guidelines.
“the protocol was approved by the ethics committee of Xijing Hospital (ID: KY20212080-F-1) and responsible ethics committees in all participating centres.”
“Written informed consent was obtained from all patients.”
“The trial was conducted in accordance with the Declaration of Helsinki and Good Clinical Practice guidelines”
“The trial was conducted in accordance with the Declaration of Helsinki and Good Clinical Practice guidelines, and the protocol was approved by the ethics committee of Xijing Hospital (ID: KY20212080-F-1) and responsible ethics committees in all participating centres.”
“Written informed consent was obtained from all patients.”
“Clinicaltrials.gov NCT04971356 (https://clinicaltrials.gov/ct2/show/NCT04971356)”
The trial uses paclitaxel-coated balloons (manufacturer Yinyi Biotech) and antiplatelet drugs (aspirin, ticagrelor, clopidogrel) with doses specified. The DCB brands are summarised in table S2. Statistical software (R version 4.2.1) is identified. Bench criteria (antibodies, cell lines, mycoplasma) are not applicable. Reagents are adequately identified for a clinical trial.
“For maintenance, aspirin was prescribed at 100 mg daily and ticagrelor was prescribed at 90 mg twice daily.”
“The analysis was done using R statistical software version 4.2.1 (R Project for statistical computing).”
“brands and features of DCBs used are summarised in table S2”
“Yinyi Biotech manufactures paclitaxel coated balloons”
“For maintenance, aspirin was prescribed at 100 mg daily and ticagrelor was prescribed at 90 mg twice daily.”
“The analysis was done using R statistical software version 4.2.1 (R Project for statistical computing).”
The paper names statistical tests (Kaplan-Meier, Greenwood's method, approximate z test, win ratio) and reports exact p-values and 95% CIs. Assumptions are handled by design (covariate-adjusted analysis, competing risk sensitivity). Software is identified. Data presentation includes Kaplan-Meier curves and per-group n. Mathematical plausibility checks: primary endpoint counts (87/975=8.9%, 84/973=8.6%) are consistent; BARC 3/5 bleeding (4 vs 16) yields difference -1.19% (95% CI -2.07 to -0.31) which is plausible. No demonstrable errors.
“The cumulative event rate was estimated at 360 days by the Kaplan-Meier method, with the standard error of difference calculated using Greenwood's method and P value calculated using an approximate z test.”
“BARC type 3 or 5 bleeding | 4 (0.4) | 16 (1.6) | −1.19 (−2.07 to −0.31) | 0.008”
“The analysis was done using R statistical software version 4.2.1 (R Project for statistical computing).”
“The cumulative event rate was estimated at 360 days by the Kaplan-Meier method, with the standard error of difference calculated using Greenwood's method and P value calculated using an approximate z test.”
“P non-inferiority =0.013”
“BARC type 3 or 5 bleeding occurred in four versus 16 participants (0.4% v 1.6%, difference −1.19% (95% CI −2.07% to −0.31%), P=0.008)”
The data availability statement says patient-level data 'will not be made publicly available but will be available for data sharing on request for collaboration on specific projects.' This is reported_but_inadequate because it lacks a named platform, conditions, or timeframe. Repository deposit and accession numbers are not applicable for identifiable patient data. Code sharing is not applicable as no bespoke code is mentioned.
“Patient level data collected for this study will not be made publicly available but will be available for data sharing on request for collaboration on specific projects.”
“Patient level data collected for this study will not be made publicly available but will be available for data sharing on request for collaboration on specific projects.”
The trial is registered at ClinicalTrials.gov (NCT04971356). Methods are detailed enough for replication. All pre-specified outcomes are reported, including negative results. Limitations are explicitly discussed. Conclusions are proportional to the evidence. Funding sources and competing interests are declared.
“Clinicaltrials.gov NCT04971356 (https://clinicaltrials.gov/ct2/show/NCT04971356)”
“This study has several limitations. Firstly, the sample size calculation for non-inferiority was based on a one sided α of 5%; nevertheless, the sensitivity analysis using a one sided α of 2.5% still showed non-inferiority.”
“The study received unrestricted grant support from Yinyi Biotech (Dalian, China).”
“Clinicaltrials.gov NCT04971356 (https://clinicaltrials.gov/ct2/show/NCT04971356)”
“This study has several limitations. Firstly, the sample size calculation for non-inferiority was based on a one sided α of 5%; nevertheless, the sensitivity analysis using a one sided α of 2.5% still showed non-inferiority.”
“The study received unrestricted grant support from Yinyi Biotech (Dalian, China).”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 44 references by DOI: 44 verified.
Every extracted reference resolved against Crossref/OpenAlex with no retraction flags.
1 data/code link checked; 1 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT04971356LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly clarity, consistency, typo.
- MINORconsistencyAbstract, Results“P non-inferiority =0.013”→ Use consistent formatting for p-values, e.g., 'P=0.013'.Minor formatting inconsistency.
- MINORclarityMethods, Statistical analysis“The treatment difference was defined as the stepwise DAPT de-escalation group minus standard DAPT group.”→ Clarify the direction of the difference to avoid ambiguity.Could be clearer.
- MINORclarityTable 2, footnote“The listed percentages were estimated with the use of the Kaplan-Meier method, so values may not be calculated mathematically.”→ Clarify that percentages are Kaplan-Meier estimates and not simple proportions.This is a helpful note but could be more explicit.
- MINORtypoDiscussion, Comparison with other studies“the mean device length was 33 mm”→ Ensure consistency with earlier 'Total DCB length' reporting.Potential inconsistency in terminology.
The published work is generally robust, but the primary non-inferiority p-value inconsistency is a validity concern that an informed reader should weigh; an erratum or independent re-analysis may be warranted. The vague data-sharing statement is a transparency gap.
- 1.CRITICALstatisticsResolve the statistics inconsistency that flips a significance claim: Recomputed 3 tests: 2 consistent, 1 inconsistent (1 change significance at p<.05); 3 via agent-written checks.Demonstrable critical failure — blocks the verdict from passing.
- 2.HIGHstatisticsRe-verify the primary non-inferiority p-value (reported P=0.013) against the reported difference and CI; if the recomputed value (P=0.738) is correct, issue a correction or re-analysis.The deterministic recomputation found a large discrepancy that could affect the trial's primary conclusion.
- 3.HIGHdata codeEnhance the data availability statement to specify a concrete access mechanism (e.g., data-access committee or platform like Vivli) and include conditions and a timeframe for requests.The current 'on request' statement is vague and does not meet transparency standards for reproducibility.
- 4.HIGHreportingExplicitly state adherence to CONSORT in the methods or provide a completed CONSORT checklist.The trial is a randomised controlled trial and explicit CONSORT adherence would strengthen reporting transparency.
- 5.MEDIUMreportingClarify the race/ethnicity breakdown of participants in the baseline table, as it is collected but not reported.Enhances demographic transparency and generalisability assessment.
- 6.MEDIUMdata codeAdd a statement about the availability of the statistical analysis code, even if not publicly shared.Improves reproducibility and transparency of the analysis.
- 7.MEDIUMstatisticsReport exact p-values for all secondary endpoints in the main text, as some are only given as '<0.001'.Provides more precise information for readers and facilitates verification.
- 8.MEDIUMreportingClarify the handling of missing data for covariates in the adjusted analysis.The current description is incomplete and could affect reproducibility.
- 9.MEDIUMcopyeditStandardise p-value formatting in the abstract (e.g., 'P=0.013' instead of 'P non-inferiority =0.013').Minor formatting inconsistency that could be cleaned up.
- 10.MEDIUMcopyeditClarify the direction of the treatment difference in the statistical analysis section.The current wording is ambiguous and could confuse readers.
- 11.LOWcopyeditClarify the Table 2 footnote that percentages are Kaplan-Meier estimates and not simple proportions.The note is helpful but could be more explicit to avoid misinterpretation.
- 12.LOWcopyeditEnsure consistent terminology for device length (e.g., 'mean device length' vs 'Total DCB length').Potential inconsistency in terminology could confuse readers.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.