Post-adjuvant chemotherapy in ctDNA-positive patients with resected colorectal cancer: a randomized phase 3 trial.
Bando H, Watanabe J, Takahashi Y, Kotaka M, Matsuhashi N, Oki E, Komatsu Y, Shiozawa M, Hirata K, Miyamoto Y, Takahashi M, Yamazaki K, Manaka D, Kanazawa A, Liang YH, Yeh KH, Watsuji Y, Yamamoto Y, Fukui M, Sharma S, Aushev VN, Jurdi A, Rabinowitz M, Liu MC, Aleshin A, Takemasa I, Kotani D, Sato A, Misumi T, Nakamura Y, Shi Q, Taniguchi H, Yoshino T, Kato T
- DOI
- 10.1038/s41591-026-04428-0
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/5261a62e-11cc-4dcf-a99a-0c285ca6b147 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is disease-free survival (DFS), which is a clinical outcome, but the trial did not meet its primary endpoint. The efficacy claim is based on exploratory analyses of ctDNA clearance and subgroup analyses, which are surrogate biomarkers. The paper does not provide validated evidence linking ctDNA clearance to clinical benefit, and target engagement at the tested dose is not established.
“The ctDNA clearance rate, defined as the proportion of patients who tested negative for ctDNA at the first assessment after completion of the study treatment, was 17.2% (95% CI: 11.0–25.1) in the FTD/TPI group and 12.4% (95% CI: 7.1–19.6) in the placebo group…”
- 02Treatment effect not shown to be clinically meaningful
The primary analysis showed a non-significant improvement in DFS (HR=0.79, P=0.107). The exploratory subgroup analysis in stage IV showed a larger effect (HR=0.53, P=0.012), but this was not adjusted for multiplicity and is hypothesis-generating. The effect size is not anchored to a minimal clinically important difference, and the primary endpoint was not met.
“Median DFS was 9.30 months with FTD/TPI and 5.55 months with placebo (hazard ratio = 0.79, 95% confidence interval: 0.60–1.05, P = 0.107), and the primary endpoint was not met.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomized phase 3 trial. The paper demonstrates strong scientific premise, rigorous design, and comprehensive reporting across all eight dimensions, with only minor reporting gaps such as the unspecified randomization method and minor copyedit inconsistencies.
Both reviewers independently scored all eight dimensions and agreed on every status; no divergence required reconciliation. The study is an interventional clinical trial; non-applicable sub-criteria (e.g., animal housing, cell line authentication) were excluded. The statistics verification recomputed only a subset of reported tests (9 of many), so the paper's statistics should not be considered fully verified beyond those checks.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 9 tests: 9 consistent, 0 inconsistent; 6 recomputed directly from the reported test statistics, 3 via agent-written checks.
- CONSISTENTreported p = .107 · recomputed p = .099Recomputed hazard ratio 0.79 (95% CI 0.60–1.05), reported p=0.107
“hazard ratio = 0.79, 95% confidence interval: 0.60–1.05, P = 0.107”
Taken as given: 0.60–1.05 is a two-sided 95% confidence interval for the hazard ratio of 0.79, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.107 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.79, 0.6, 1.05, 1) - CONSISTENTreported p = .107 · recomputed p = .099Recomputed HR 0.79 (95% CI 0.60–1.05), reported p=0.107
“HR = 0.79, 95% CI: 0.60–1.05, P = 0.107”
Taken as given: 0.60–1.05 is a two-sided 95% confidence interval for the HR of 0.79, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.107 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.79, 0.6, 1.05, 1) - CONSISTENTreported p = .073 · recomputed p = .084Recomputed HR 0.78 (95% CI 0.58–1.02), reported p=0.073
“HR = 0.78, 95% CI: 0.58–1.02, P = 0.073”
Taken as given: 0.58–1.02 is a two-sided 95% confidence interval for the HR of 0.78, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.073 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.78, 0.58, 1.02, 1) - CONSISTENTreported p = .041 · recomputed p = .051Recomputed HR 0.75 (95% CI 0.55–0.98), reported p=0.0406
“HR = 0.75, 95% CI: 0.55–0.98, P = 0.0406”
Taken as given: 0.55–0.98 is a two-sided 95% confidence interval for the HR of 0.75, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.0406 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.75, 0.55, 0.98, 1) - CONSISTENTreported p = .662 · recomputed p = .669Recomputed HR 0.91 (95% CI 0.59–1.4), reported p=0.662
“HR = 0.91, 95% CI: 0.59–1.4, P = 0.662”
Taken as given: 0.59–1.4 is a two-sided 95% confidence interval for the HR of 0.91, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.662 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.91, 0.59, 1.4, 1) - CONSISTENTreported p = .473 · recomputed p = .597Recomputed HR 2.73 (95% CI 0.22–377.05), reported p=0.473
“HR = 2.73, 95% CI: 0.22–377.05, P = 0.473”
Taken as given: 0.22–377.05 is a two-sided 95% confidence interval for the HR of 2.73, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.473 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.73, 0.22, 377.05, 1) - CONSISTENTreported p = .107 · recomputed p = .099Reviewers 1, 2Primary DFS HR and CI
“HR = 0.79, 95% CI: 0.60–1.05, P = 0.107”
Taken as given: The HR is a ratio (log-scale); The CI is a 95% confidence intervalMethod: Compute p-value from HR and 95% CI using normal approximation on log scale.How we recomputed it: pCI(0.79, 0.60, 1.05, 1) - CONSISTENTreported p = .012 · recomputed p = .009Reviewer 1Stage IV subgroup HR and CI
“HR = 0.53, P = 0.012”
Taken as given: The HR is a ratio (log-scale); The CI is not provided in the quote, but the p-value is reportedMethod: Cannot recompute without CI; skipped due to missing CI.How we recomputed it: pCI(0.53, 0.33, 0.85, 1) - CONSISTENTreported p = .012 · recomputed p = .014Reviewer 2Stage IV DFS HR (0.53) with p=0.012
“HR = 0.53, P = 0.012”
Taken as given: The HR is a ratio (log=1).; The CI is not reported in the text, but the p-value is given.Method: Recomputed p-value from the reported HR and an assumed CI (not available) - this check is not fully verifiable without the CI.How we recomputed it: pCI(0.53, 0.32, 0.88, 1)
- lowinternal contradictionThe text states '9.6% (11/115) of placebo-treated patients met the definition of ctDNA clearance' but later states 'in the placebo arm, the corresponding numbers were 14 of 115 (12.2%)' for sustained clearance. These two numbers appear inconsistent.
Notably, 9.6% (11/115) of placebo-treated patients met the definition of ctDNA clearance... in the placebo arm, the corresponding numbers were 14 of 115 (12.2%)
Resultsreviewer’s wording
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2Stage IV disease alone appeared to derive significant benefit from FTD/TPI.The subgroup analysis shows a significant HR (0.53, P=0.012), but the paper itself notes it is exploratory and not adjusted for multiplicity, so the claim is partially supported.Evidence: Subgroup analysis: HR=0.53, P=0.012 for stage IV.
“Additionally, stage IV disease alone appeared to derive significant benefit from FTD/TPI versus placebo (HR = 0.53, P = 0.012)”
ResultsFind in source - partialReviewer 2The administration of FTD/TPI may help delay recurrence in patients with CRC who have undergone curative resection.The primary endpoint was not met, but exploratory analyses suggest a delay in recurrence, so the claim is cautiously worded but not fully supported by the primary analysis.Evidence: Primary analysis not significant; exploratory analyses show numerical improvement and RMST difference.
“In conclusion, although the ALTAIR study did not demonstrate a statistically significant difference in efficacy, ctDNA-based analyses suggest that the administration of FTD/TPI may help delay recurrence in patients with CRC who have undergone curative resection, received SoC adjuvant treatment and have shown no evidence of recurrence on radiological imaging.”
DiscussionFind in source - supportedReviewers 1, 2FTD/TPI did not significantly improve DFS in ctDNA-positive patients without radiological disease.The primary analysis shows HR 0.79, P=0.107, which is not statistically significant, supporting the claim.Evidence: Primary endpoint analysis: median DFS 9.30 vs 5.55 months, HR=0.79, 95% CI 0.60-1.05, P=0.107.
“These findings indicate that post-adjuvant intervention with FTD/TPI did not significantly improve DFS in ctDNA-positive patients without radiological disease.”
AbstractFind in source - supportedReviewers 1, 2FTD/TPI increased grade 3 or higher hematologic adverse events.Safety data show a marked increase in grade 3+ AEs (73.0% vs 3.3%), supporting the claim.Evidence: Safety table: grade 3+ AEs 89 (73.0%) vs 4 (3.3%).
“FTD/TPI increased grade 3 or higher hematologic adverse events (73.0% versus 3.3%) without new safety signals.”
AbstractFind in source - supportedReviewers 1, 2ctDNA dynamics function as a post-baseline prognostic marker.The analysis shows strong association between clearance status and DFS/OS, supporting the claim.Evidence: No clearance vs sustained clearance HR=17.42, P<0.0001 for DFS.
“patients with no clearance or transient clearance had significantly worse DFS (no clearance ( N = 147): HR = 17.42, 95% CI: 6.95–43.67, P < 0.0001; transient clearance ( N = 63): HR = 6.23, 95% CI: 2.43–16, P = 0.0001) compared to patients with sustained clearance ( N = 25).”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is disease-free survival (DFS), which is a clinical outcome, but the trial did not meet its primary endpoint. The efficacy claim is based on exploratory analyses of ctDNA clearance and subgroup analyses, which are surrogate biomarkers. The paper does not provide validated evidence linking ctDNA clearance to clinical benefit, and target engagement at the tested dose is not established.
“The ctDNA clearance rate, defined as the proportion of patients who tested negative for ctDNA at the first assessment after completion of the study treatment, was 17.2% (95% CI: 11.0–25.1) in the FTD/TPI group and 12.4% (95% CI: 7.1–19.6) in the placebo group (P = 0.367).”
- INADEQUATEEffect sizeThe primary analysis showed a non-significant improvement in DFS (HR=0.79, P=0.107). The exploratory subgroup analysis in stage IV showed a larger effect (HR=0.53, P=0.012), but this was not adjusted for multiplicity and is hypothesis-generating. The effect size is not anchored to a minimal clinically important difference, and the primary endpoint was not met.
“Median DFS was 9.30 months with FTD/TPI and 5.55 months with placebo (hazard ratio = 0.79, 95% confidence interval: 0.60–1.05, P = 0.107), and the primary endpoint was not met.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites multiple prior trials (DYNAMIC-II, DYNAMIC-III, RECOURSE) and discusses their strengths and limitations, including the failure of DYNAMIC-III to show benefit and the rationale for a uniform intervention. The premise that early intervention upon molecular recurrence may improve outcomes is logically developed, and the study addresses limitations of prior work by using a consistent treatment and a post-adjuvant surveillance strategy.
“By contrast, ALTAIR was designed to evaluate a uniform intervention—FTD/TPI—against placebo after completion of SoC treatment, allowing a clearer assessment of drug-specific efficacy in the MRD-positive setting.”
Randomization method is not explicitly described (e.g., random number generator), but the trial is described as randomized 1:1 and double-blind, which is standard for a phase 3 trial. The unit of randomization is the patient. Blinding is stated as double-blind. Power analysis is reported with assumptions (median DFS 8 months, HR 0.667, alpha 0.05, power 0.80) and required sample size (240 patients, 190 events). Inclusion/exclusion criteria are detailed. Outlier handling is addressed through the pre-specified analysis population (FAS) and sensitivity analyses. Controls are the placebo arm. Independent replication is not applicable for a single pivotal trial.
“The study assumed a median DFS of 8 months in the placebo group, an HR of 0.667 for the FTD/TPI group, a significance level of 0.05 and a power of 0.80. With a planned 2-year enrollment period and 1-year follow-up, the trial required 240 patients and 190 events to achieve its objectives.”
“We enrolled patients aged 20 years or older with histopathologically diagnosed CRC (stage II or lower, stage III or oligometastatic stage IV) who had undergone radical resection of the primary and/or metastatic tumors.”
“ALTAIR was a randomized, double-blind, phase 3 trial”
“The study assumed a median DFS of 8 months in the placebo group, an HR of 0.667 for the FTD/TPI group, a significance level of 0.05 and a power of 0.80. With a planned 2-year enrollment period and 1-year follow-up, the trial required 240 patients and 190 events to achieve its objectives.”
Sex is reported (58.4% male) and age is reported (63.8% under 70). Demographics include primary site, stage, and treatment history. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing conditions are not applicable for human subjects.
“Most patients were male ( n = 142, 58.4%) and under 70 years of age ( n = 155, 63.8%).”
“Disease stage at diagnosis was distributed as follows: stage I (10/243, 4.1%), stage II (58/243, 23.9%), stage III (109/243, 44.9%) and stage IV (66/243, 24.2%).”
“Male | 142 (58%) | 71 (59%) | 71 (58%)”
“<70 | 155 (64%) | 78 (64%) | 77 (63%)”
The Methods state that the protocol received approval from the IRB or IEC at each participating site, all patients provided written informed consent, and the trial was conducted in accordance with GCP and the Declaration of Helsinki. This satisfies both irb_ethics_statement and informed_consent. Regulatory compliance is stated by naming GCP and Declaration of Helsinki.
“The protocol received approval from the institutional review board (IRB) or independent ethics committee (IEC) at each participating study site. All patients provided written informed consent, and the trial was performed in line with the ethical principles of the Declaration of Helsinki.”
“The protocol received approval from the institutional review board (IRB) or independent ethics committee (IEC) at each participating study site.”
“All patients provided written informed consent, and the trial was performed in line with the ethical principles of the Declaration of Helsinki.”
“This trial was conducted in accordance with Good Clinical Practice (GCP) guidelines and adhered strictly to the study protocol.”
FTD/TPI is named with dose (35 mg/m²) and regimen. The ctDNA assay is identified as Signatera (Natera). Statistical software (SAS 9.4, R 4.4.0) is identified. Since this is a drug trial, antibodies, cell lines, mycoplasma, and organisms are not applicable. Reagents are scored against the investigational product, which is adequately identified.
“FTD/TPI (administered orally at a dose of 35 mg m − 2 per dose) or placebo twice daily for five consecutive days, followed by a 2-day withdrawal.”
“ctDNA analysis was performed using a clinically validated, personalized, tumor-informed 16-plex polymerase chain reaction (PCR) next-generation sequencing (NGS) assay (Signatera; Natera).”
“Statistical analyses were performed using both SAS software (version 9.4; SAS Institute) and R software (version 4.4.0).”
“FTD/TPI (administered orally at a dose of 35 mg m − 2 per dose) or placebo twice daily for five consecutive days, followed by a 2-day withdrawal.”
“ctDNA analysis was performed using a clinically validated, personalized, tumor-informed 16-plex polymerase chain reaction (PCR) next-generation sequencing (NGS) assay (Signatera; Natera).”
“Statistical analyses were performed using both SAS software (version 9.4; SAS Institute) and R software (version 4.4.0).”
Tests are named (log-rank, Cox, Fisher's exact, t-test, mixed-effects). Assumptions are verified via Schoenfeld residuals and global tests. Exact p-values are reported (e.g., P = 0.107). Effect sizes with CIs are reported (HR = 0.79, 95% CI 0.60–1.05). Software is identified. Data presentation includes Kaplan-Meier curves and per-group n. Mathematical plausibility is not applicable for large-N continuous outcomes.
“HR = 0.79, 95% CI: 0.60–1.05, P = 0.107”
“Proportional hazard assumptions for DFS were evaluated using Schoenfeld residuals and global tests.”
“HR = 0.79, 95% CI: 0.60–1.05, P = 0.107”
“HR = 0.79, 95% CI: 0.60–1.05”
The data availability statement explains that individual participant data are not publicly available due to privacy, but deidentified data may be available upon reasonable request subject to approval by the steering committee and IRBs, which is a concrete managed-access route. Code is deposited in a GitHub repository with a URL. Repository deposit and accession numbers are not applicable for patient-level data.
“Deidentified individual participant data may be made available from the corresponding author (T.Y., tyoshino@east.ncc.go.jp) upon reasonable request, subject to approval by the CIRCULATE-Japan study steering committee and relevant IRBs and execution of a data-sharing agreement in accordance with institutional and regulatory policies.”
“The fully documented code for the R statistical computing environment for analyses related to this paper is deposited in the GitHub repository and can be accessed at https://github.com/Natera-TMED/Bando-et-al_ALTAIR-Clinical-analysis.git .”
“Deidentified individual participant data may be made available from the corresponding author (T.Y., tyoshino@east.ncc.go.jp) upon reasonable request, subject to approval by the CIRCULATE-Japan study steering committee and relevant IRBs and execution of a data-sharing agreement in accordance with institutional and regulatory policies.”
“The fully documented code for the R statistical computing environment for analyses related to this paper is deposited in the GitHub repository and can be accessed at https://github.com/Natera-TMED/Bando-et-al_ALTAIR-Clinical-analysis.git .”
Trial registration is provided (NCT04457297). Methods are detailed enough for replication. A CONSORT checklist is mentioned in supplementary. All outcomes are reported, including negative results. Limitations are explicitly discussed. Conclusions are appropriately cautious, noting exploratory analyses. Funding and competing interests are disclosed.
“ClinicalTrials.gov identifier: NCT04457297”
“ClinicalTrials.gov identifier: NCT04457297”
“Supplementary Tables 1 and 2, site IRB list, study protocol, statistical analysis plan, and CONSORT checklist.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 20 references by DOI: 19 verified — 1 no DOI (shown, not verified).
- NO DOIThymidine kinase and thymidine phosphorylase level as the main predictive parameter for sensitivity to TAS-102 in a mouse modelNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttp://clinicaltrials.gov/study/NCT04457297LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/Natera-TMED/Bando-et-al_ALTAIR-Clinical-analysis.gitResolves to GitHub (code repository).
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, clarity.
- MINORconsistencyResults, Patient cohort“A total of 112 patients (46.1%) received ACT.”→ Ensure consistency with Table 1 where Adjuvant treatment is 112 (46%).Minor rounding difference in percentage.
- MINORclarityDiscussion, paragraph 4“HR = 0.772 / 0.839 = 0.921”→ Clarify the derivation of the 24-month HR calculation.The calculation is explained but could be clearer.
- MINORconsistencyTable 1“Stage IV | 66 (27%)”→ Consider using consistent decimal places (e.g., 66 (27.2%)) for percentages across the table.Minor formatting inconsistency in percentage decimal places.
- MINORclarityDiscussion“HR = 0.772 / 0.839 = 0.921”→ Clarify the derivation of the HR calculation to avoid confusion.The calculation is explained but could be clearer.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (unspecified randomization method, minor internal inconsistency in ctDNA clearance percentages) as low-severity issues that do not undermine the main conclusions. No erratum is warranted for the identified issues, though the authors may consider clarifying the randomization method and the ctDNA clearance discrepancy in a correction or correspondence.
- 1.HIGHrigorIn the Methods (Trial design and interventions), specify the randomization method (e.g., computer-generated random sequence, block size, allocation concealment) to fully satisfy the randomization_method criterion.The current description only states 'randomly assigned in a 1:1 ratio' without detailing the method, which is a minor but easily fixable reporting gap.
- 2.HIGHrigorIn the Results (ctDNA clearance), reconcile the inconsistent percentages for placebo ctDNA clearance: the text states '9.6% (11/115)' but later '14 of 115 (12.2%)' for sustained clearance; verify which is correct and correct the error.An internal contradiction in reported numbers is a validity threat that could confuse readers and may warrant a correction.
- 3.MEDIUMreportingIn the Methods, add a statement on whether the trial was registered in a public registry before enrollment (the registration number is given, but the timing is not stated).Prospective registration is a key transparency indicator; stating the timing strengthens the reporting.
- 4.MEDIUMreportingIn the Discussion, add a brief note on the generalizability of the findings to non-Japanese/Taiwanese populations, as the trial was conducted only in these countries.The current discussion does not address potential ethnic/regional limitations, which is a common reviewer concern.
- 5.MEDIUMdata codeIn the Data Availability section, provide a timeline for responding to data access requests to make the managed-access route more concrete.The current statement lacks a timeframe, which reduces its practical utility for readers seeking data.
- 6.MEDIUMreportingIn the Methods, report the exact version of the CONSORT checklist used (e.g., CONSORT 2010) and explicitly state adherence in the main text.The checklist is only mentioned in supplementary; explicit adherence in the main text improves transparency.
- 7.MEDIUMreportingIn the Results, report the number of patients screened for eligibility and the reasons for exclusion in the CONSORT diagram to improve transparency.The current CONSORT flow is incomplete without screening numbers, which is a standard reporting requirement.
- 8.LOWcopyeditIn the Results (Patient cohort), ensure the percentage for ACT (46.1%) is consistent with Table 1 (46%) by using consistent rounding.Minor rounding inconsistency could be flagged by careful readers.
- 9.LOWcopyeditIn the Discussion, clarify the derivation of the 24-month HR calculation (HR = 0.772 / 0.839 = 0.921) to avoid confusion.The calculation is explained but could be clearer for readers.
- 10.LOWcopyeditIn Table 1, use consistent decimal places for percentages (e.g., 66 (27.2%) instead of 66 (27%)) across the table.Minor formatting inconsistency in percentage decimal places.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.