Fecal microbiota transplantation plus pembrolizumab and axitinib in metastatic renal cell carcinoma: the randomized phase 2 TACITO trial.
Porcari S, Ciccarese C, Heidrich V, Rondinella D, Quaranta G, Severino A, Arduini D, Buti S, Fornarini G, Primi F, Stumbo L, Giannarelli D, Giudice GC, Damassi A, Giron Berríos JR, Punčochář M, Barbazuk TB, Piccinno G, Pinto F, Armanini F, Asnicar F, Schinzari G, Derosa L, Kroemer G, Sanguinetti M, Masucci L, Gasbarrini A, Tortora G, Cammarota G, Zitvogel L, Segata N, Iacovelli R, Ianiro G
- DOI
- 10.1038/s41591-025-04189-2
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/06e380b0-dde0-4601-9654-c8afb812ab60 is authoritative.
How this rating was calculated
- StatisticsStatistic did not reproduce ×3−1.5★
- StatisticsPrinted percentage does not match its own count (capped)−0.25★
- ReportingBiological variables partially met−0.25★
- ReportingKey resources partially met−0.25★
- References were not verified against Crossref/OpenAlex.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run on this paper: the pass that reads its reported means did not complete. No reported mean was checked for arithmetic impossibility.
- 01Significance claim does not survive recomputation
Recomputed hazard ratio 0.36 (95% CI 0.13–0.99), reported p=0.167
“hazard ratio = 0.36, 95% CI: 0.13–0.99, P = 0.167”
- 02Significance claim does not survive recomputation
Recomputed hazard ratio 0.36 (95% CI 0.13–0.99), reported p=0.146
“hazard ratio = 0.36 (95% CI: 0.13–0.99), P = 0.146”
- 03Significance claim does not survive recomputation
Recomputed HR 0.36 (95% CI 0.13–0.99), reported p=0.146
“HR 0.36, 95% CI 0.13-0.99, p = 0.146”
- 04Printed percentage does not match its own count
67% does not match the reported count 33/50
“33 of 50 patients (67%)”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed randomized double-blind placebo-controlled phase 2a trial with strong reporting, registration, and data availability. Its main weaknesses are incomplete identification of investigational products, missing race/ethnicity and weight/performance-status demographics, and a demonstrable inconsistency in reported p-values for a key hazard ratio.
Two independent reviewer runs were synthesized; they diverged on biological variables (pass vs warn) which we resolved to warn, and both missed the statistical inconsistency found by the verification component, which we incorporated. The statistics component covered only 13 of reported tests; everything not recomputed remains unverified.
Numerical inconsistencies
1 finding · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
Recomputed 12 tests: 9 consistent, 3 inconsistent (3 change significance at p<.05); 8 recomputed directly from the reported test statistics, 4 via agent-written checks. 1 printed percentage that does not match its own count.
- PERCENT67% does not match the reported count 33/50
“33 of 50 patients (67%)”
- CONSISTENTreported p = .048 · recomputed p = .049Recomputed hazard ratio 0.48 (95% CI 0.23–0.99), reported p=0.048
“hazard ratio = 0.48, 95% CI: 0.23–0.99, P = 0.048”
Taken as given: 0.23–0.99 is a two-sided 95% confidence interval for the hazard ratio of 0.48, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.048 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.48, 0.23, 0.99, 1) - CONSISTENTreported p = .180 · recomputed p = .248Recomputed hazard ratio 0.66 (95% CI 0.33–1.35), reported p=0.18
“hazard ratio = 0.66, 95% CI: 0.33–1.35, P = 0.18”
Taken as given: 0.33–1.35 is a two-sided 95% confidence interval for the hazard ratio of 0.66, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.18 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.66, 0.33, 1.35, 1) - CONSISTENTreported p = .146 · recomputed p = .158Recomputed hazard ratio 0.51 (95% CI 0.20–1.30), reported p=0.146
“hazard ratio = 0.51, 95% CI: 0.20–1.30, P = 0.146”
Taken as given: 0.20–1.30 is a two-sided 95% confidence interval for the hazard ratio of 0.51, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.146 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.51, 0.2, 1.3, 1) - CONSISTENTreported p = .180 · recomputed p = .248Recomputed HR 0.66 (95% CI 0.33–1.35), reported p=0.18
“HR: 0.66, 95% CI: 0.33-1.35, p = 0.18”
Taken as given: 0.33–1.35 is a two-sided 95% confidence interval for the HR of 0.66, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.18 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.66, 0.33, 1.35, 1) - CONSISTENTreported p = .146 · recomputed p = .158Recomputed HR 0.51 (95% CI 0.20–1.30), reported p=0.146
“HR 0.51, 95% CI: 0.20-1.30, p = 0.146”
Taken as given: 0.20–1.30 is a two-sided 95% confidence interval for the HR of 0.51, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.146 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.51, 0.2, 1.3, 1) - CONSISTENTreported p = .113 · recomputed p = .075Reviewer 1Primary endpoint: 12-month PFS comparison (d-FMT vs p-FMT) using Fisher's exact test.
“d-FMT: 16/23 patients, 70%; p-FMT: 9/22 patients, 41%; P = 0.053”
Taken as given: The numbers 16 and 9 are the progression-free counts in each arm.; The denominators 23 and 22 are the total patients per arm.; The test is two-sided Fisher's exact test.Method: Two-sided Fisher's exact test computed from the 2x2 contingency table (16,7,9,13).How we recomputed it: pFisher2x2(16,7,9,13,0) - CONSISTENTreported p = .035 · recomputed p = .035Reviewer 1Median PFS comparison using log-rank test (reported as Breslow test).
“hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035”
Taken as given: The test statistic is chi-square with 1 degree of freedom from a log-rank test.; The reported p-value is from the Breslow test (which is a generalized Wilcoxon test).Method: The p-value is taken as reported; we cannot recompute without raw data. The expression is approximate based on the chi-square distribution.How we recomputed it: pChi2(4.44,1) - CONSISTENTreported p = .053 · recomputed p = .053Reviewer 2Primary endpoint comparison: 12-month PFS proportion (70% vs 41%)
“70% versus 41% for d-FMT versus p-FMT, respectively, P = 0.053”
Taken as given: The numbers 16 and 7 are the progression-free and progressed counts in the d-FMT arm (FAS); The numbers 9 and 13 are the progression-free and progressed counts in the p-FMT arm (FAS); The chi-square test is two-sidedMethod: Pearson chi-square test for 2x2 contingency tableHow we recomputed it: pChi2x2(16,7,9,13) - CONSISTENTreported p = .035 · recomputed p = .027Reviewer 2Median PFS hazard ratio, 90% CI
“hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035”
Taken as given: The HR is from a Cox proportional hazards model; The 90% CI is used (the paper states 90% CI for HR); The p-value is two-sided from the Cox modelMethod: P-value from estimate and confidence interval on log scaleHow we recomputed it: pCI(0.5, 0.27, 0.92, 1)
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Significance claim flips when recomputedRecomputed
- Conclusions only partially backed by the presented evidenceAssessed
- MORE SIGNIFICANT ON RECHECKINCONSISTENTreported p = .167 · recomputed p = .049Recomputed hazard ratio 0.36 (95% CI 0.13–0.99), reported p=0.167Recomputing from the paper’s own numbers lands below p = 0.05 — more significant than the printed value. Usually benign (the reported figure is conservative), but the two don’t match.
“hazard ratio = 0.36, 95% CI: 0.13–0.99, P = 0.167”
Taken as given: 0.13–0.99 is a two-sided 95% confidence interval for the hazard ratio of 0.36, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.167 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.36, 0.13, 0.99, 1) - MORE SIGNIFICANT ON RECHECKINCONSISTENTreported p = .146 · recomputed p = .049Recomputed hazard ratio 0.36 (95% CI 0.13–0.99), reported p=0.146Recomputing from the paper’s own numbers lands below p = 0.05 — more significant than the printed value. Usually benign (the reported figure is conservative), but the two don’t match.
“hazard ratio = 0.36 (95% CI: 0.13–0.99), P = 0.146”
Taken as given: 0.13–0.99 is a two-sided 95% confidence interval for the hazard ratio of 0.36, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.146 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.36, 0.13, 0.99, 1) - MORE SIGNIFICANT ON RECHECKINCONSISTENTreported p = .146 · recomputed p = .049Recomputed HR 0.36 (95% CI 0.13–0.99), reported p=0.146Recomputing from the paper’s own numbers lands below p = 0.05 — more significant than the printed value. Usually benign (the reported figure is conservative), but the two don’t match.
“HR 0.36, 95% CI 0.13-0.99, p = 0.146”
Taken as given: 0.13–0.99 is a two-sided 95% confidence interval for the HR of 0.36, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.146 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.36, 0.13, 0.99, 1)
10 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewer 1Acquisition of specific donor strains (e.g., Blautia wexlerae) is associated with 12-month PFS.The association is significant but based on small numbers and multiple testing; the paper appropriately cautions about exploratory nature.Evidence: Results, Donor microbiome engraftment effects on clinical outcomes.
The acquisition of the Blautia wexlerae (SGB4837) strain from the donor at week 1 was positively associated with 12-month PFS (percentage of recipients acquiring strain with versus without PFS > 12 months: 50% versus 0%, P = 0.047).
Figure 4Dreviewer’s wording - partialReviewer 2Acquisition or loss of specific strains, but not total engraftment, was associated with the primary endpointSome specific strain associations were significant (e.g., B. wexlerae acquisition, A. massiliensis acquisition), but total engraftment (DoSER) was not associated. The associations are exploratory with multiple testing.Evidence: B. wexlerae acquisition: 50% vs 0%, P=0.047; A. massiliensis acquisition: 0% vs 57%, P=0.006; DoSER not associated (P=0.15, 0.53).
The acquisition of the Blautia wexlerae (SGB4837) strain from the donor at week 1 was positively associated with 12-month PFS (percentage of recipients acquiring strain with versus without PFS > 12 months: 50% versus 0%, P = 0.047)... on the other hand, acquisition of the donor strain of a yet-to-be-described species (SGB14845)... was inversely associated with 12-month PFS (7% versus 71%, P = 0.006)
Resultsreviewer’s wording - supportedReviewer 1Donor FMT significantly improved median PFS compared to placebo FMT in patients with mRCC.The paper reports a significant difference in median PFS (24.0 vs 9.0 months, HR=0.50, P=0.035) in the FAS population.Evidence: Figure 2a and related text in Results.
“Median PFS was significantly improved in the d-FMT arm (24.0 months, 95% CI: 8.0–40.0 months) compared with the p-FMT arm (9.0 months, 95% CI: 2.2–15.2 months) (hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035).”
Figure 2AFind in source - supportedReviewer 1The primary endpoint (12-month PFS) was not met but showed a trend favoring d-FMT.The paper states the primary endpoint was not met (70% vs 41%, P=0.053), which is correctly interpreted.Evidence: Results section: 'the proportion of patients without progression or death 12 months after randomization was higher in the d-FMT arm than in the p-FMT arm (d-FMT: 16/23 patients, 70%; p-FMT: 9/22 patients, 41%; P = 0.053).'
“the proportion of patients without progression or death 12 months after randomization was higher in the d-FMT arm than in the p-FMT arm (d-FMT: 16/23 patients, 70%; p-FMT: 9/22 patients, 41%; P = 0.053).”
ResultsFind in source - supportedReviewer 1Donor FMT is safe with no FMT-related serious adverse events.Safety data show no FMT-related SAEs, only one grade 3 TRAE in the placebo arm.Evidence: Safety assessment section.
“No deaths related to experimental treatments were reported. No transmission of any infectious agent after d-FMT was observed.”
ResultsFind in source - supportedReviewer 1Donor FMT leads to microbiome changes, including increased α-diversity and donor strain engraftment.The paper demonstrates significant increases in α-diversity and β-diversity, and higher DoSER in the d-FMT arm.Evidence: Results, Assessment of microbiome changes and Quantification of donor microbiome engraftment.
we observed an increase in Shannon α-diversity versus baseline at week 1 (P = 0.05)... The DoSER in patients receiving d-FMT was consistently higher than in the p-FMT arm throughout the whole study period (P < 0.001 across all timepoints).
Figure 3reviewer’s wording - supportedReviewer 2Donor FMT was superior to placebo in significantly improving the median PFSThe paper reports a median PFS of 24.0 vs 9.0 months, HR=0.50, P=0.035, which supports this claim.Evidence: HR=0.50, 90% CI 0.27-0.92, P=0.035
Median PFS was significantly improved in the d-FMT arm (24.0 months, 95% confidence interval (CI): 8.0–40.0 months) compared with the p-FMT arm (9.0 months, 95% CI: 2.2–15.2 months) (hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035)
Resultsreviewer’s wording - supportedReviewer 2The primary endpoint (12-month PFS) was not metThe paper states P=0.053, which is above the prespecified significance level (likely 0.05).Evidence: 70% vs 41%, P=0.053
“the proportion of patients without progression or death 12 months after randomization was higher in the d-FMT arm than in the p-FMT arm (d-FMT: 16/23 patients, 70%; p-FMT: 9/22 patients, 41%; P = 0.053)”
ResultsFind in source - supportedReviewer 2Microbiome analysis confirmed donor strain engraftment and increased α-diversityThe paper shows significantly higher α-diversity and DoSER in d-FMT arm, supporting the claim.Evidence: Increased Shannon α-diversity at multiple timepoints, DoSER consistently higher in d-FMT arm (P<0.001).
we observed an increase in Shannon α-diversity versus baseline at week 1 (P = 0.05), week 4 (P < 0.001), week 12 (P = 0.02) and week 24 (P = 0.048) follow-ups... The DoSER in patients receiving d-FMT was consistently higher than in the p-FMT arm throughout the whole study period (P < 0.001 across all timepoints)
Resultsreviewer’s wording - supportedReviewer 2Our findings support the safety and potential efficacy of selected donor FMT to enhance ICI-based treatment in mRCCThe paper reports safety data (no FMT-related SAEs) and efficacy signals (improved median PFS, ORR trends), supporting the claim.Evidence: No FMT-related SAEs; median PFS improved; ORR 52% vs 32%.
“Our findings support the safety and potential efficacy of selected donor FMT to enhance ICI-based treatment in mRCC, which deserves further investigations.”
Discussion ¶1Find in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointPrimary endpoint is 12-month progression-free survival (PFS), a hard clinical outcome based on tumor progression or death.
“The primary endpoint was the rate of patients free from disease progression at 12 months after randomization (12-month progression-free survival (PFS)).”
- ADEQUATEEffect sizeMedian PFS improved from 9.0 to 24.0 months (HR=0.50, p=0.035), a clinically meaningful improvement.
“Median PFS was significantly improved in the d-FMT arm (24.0 months, 95% CI: 8.0–40.0 months) compared with the p-FMT arm (9.0 months, 95% CI: 2.2–15.2 months) (hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
2 findings · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Biological variables underreported (sex, age, strain)Assessed
- Key resources under-identified (antibodies, cell lines, RRIDs)Assessed
The introduction cites key studies on ICI efficacy in RCC, the association of gut microbiome with ICI response, and prior FMT trials. It explicitly states that no RCT has evaluated FMT in mRCC, providing a rationale for the study. The limitations of prior work (e.g., lack of RCTs, focus on melanoma) are acknowledged.
“However, thus far, no randomized controlled trials (RCTs) have demonstrated the efficacy of FMT in mRCC.”
“Our phase 2a placebo-controlled RCT aims to evaluate whether FMT from patients with mRCC with complete response to ICIs was effective in improving response to combined first-line therapy with pembrolizumab and axitinib in patients with mRCC.”
“However, thus far, no randomized controlled trials (RCTs) have demonstrated the efficacy of FMT in mRCC.”
Randomization was performed using an online random number generator with permuted blocks of size 4, allocation concealed. Blinding was maintained for patients, study staff, and outcome assessors. A power analysis was reported with assumptions. Inclusion/exclusion criteria were detailed, and analysis populations (ITT, FAS, per-protocol, safety) were pre-defined. The primary endpoint was 12-month PFS. The study is registered at ClinicalTrials.gov.
“A total of 50 patients is required to enter this two-treatment, parallel-design study.”
“A total of 50 patients is required to enter this two-treatment, parallel-design study. The probability is 80% that the study detects a treatment difference at a one-sided 5.0% significance level, if the true hazard ratio is 0.436.”
Table 1 reports sex and median age with range. Health status is indicated by IMDC risk class, but weight or performance status (e.g., Karnofsky score) is not tabulated. No race/ethnicity data are reported. For a human study, these are important biological variables.
“Most participants were male (73%), and the median age at the time of treatment initiation was 62 years (range, 41–79 years).”
“the median age at the time of treatment initiation was 62 years (range, 41–79 years)”
The study was approved by the institutional review board/ethics committee (ID 2664). All patients gave written informed consent. The study was conducted in accordance with the Declaration of Helsinki and ICH-GCP guidelines. ClinicalTrials.gov registration is provided.
“All enrolled patients gave their written informed consent to participate in the study.”
“The study was approved by the institutional review board (IRB)/local ethics committee (ID: 2664)”
“All enrolled patients gave their written informed consent to participate in the study.”
“The study was conducted in accordance with the Declaration of Helsinki and International Conference on the Harmonization of Good Clinical Practice guidelines”
The drugs pembrolizumab and axitinib are named without manufacturer or lot numbers; the FMT is described but donor identification is limited. Software tools are well-identified with versions. No custom code was shared publicly. According to the scoring rule, the trial's investigational product is a scored resource, and the absence of manufacturer details makes it reported_but_inadequate.
“The combination of pembrolizumab (a programmed cell death protein 1 (PD-1) inhibitor monoclonal antibody) with axitinib (a vascular endothelial growth factor receptor (VEGFR) tyrosine kinase inhibitor)”
“we enrolled patients with metastatic, histologically confirmed RCC who were eligible to receive pembrolizumab and axitinib as first-line therapy.”
“IBM-SPSS version 28.0 statistical software, GraphPad Prism version 10, R version 4.4.2 (‘survival’ and ‘survminer’ packages) and Python version 3.10.12 (‘scikit-bio’ and ‘scipy’ packages) were used for the analyses.”
The paper reports exact p-values (e.g., P=0.053, P=0.035) and hazard ratios with 95% CIs. Kaplan-Meier curves and box plots are shown with error bars defined. Software is identified. The Cox proportional hazards assumption was checked. No arithmetic inconsistencies were detected in the checked statistics.
“hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035”
“IBM-SPSS version 28.0 statistical software, GraphPad Prism version 10, R version 4.4.2”
“70% versus 41% for d-FMT versus p-FMT, respectively, P = 0.053”
“P < 0.001 across all timepoints starting from week 1 follow-up”
“hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035”
The shotgun metagenomic data are available at ENA under PRJEB94043. A data availability statement describes access to patient-level data upon request. Although no custom code repository is provided, the majority of applicable criteria are adequate.
“PRJEB94043 (http://www.ebi.ac.uk/ena/data/view/PRJEB94043)”
Methods are comprehensive. The trial is registered (NCT04758507) and a CONSORT checklist is provided. All pre-specified primary and secondary outcomes are reported, including negative results. Limitations are discussed (sample size, single donor). Conclusions are proportional to the evidence. Funding and competing interests are disclosed.
“ClinicalTrials.gov identifier: NCT04758507”
“This study has some limitations. The sample size was relatively small, with only 45 patients evaluated for the primary endpoint.”
“ClinicalTrials.gov identifier: NCT04758507”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
4 data/code links checked; 4 live.
- dataENALIVEHTTP 200http://www.ebi.ac.uk/ena/data/view/PRJEB94043Resolves to ENA (data repository).
- datahttps://clinicaltrials.gov/study/NCT04758507?term=ianiro&;rank=2#study-planLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.sealedenvelope.comLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/SegataLab/preprocessingResolves to GitHub (code repository).
Copyediting
3 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 3 minor suggestions below.
3 copyedit issues flagged: mostly consistency, other, clarity.
- MINORconsistencyAbstract“P = 0.053”→ Consider reporting exact p-value as P=0.053 (consistent with other sections).This is fine as is.
- MINORotherMethods: Statistical analysis“IBM-SPSS version 28.0 statistical software, GraphPad Prism version 10, R version 4.4.2”→ Consider adding version numbers for Python packages used.Python version is given elsewhere.
- MINORclarityMethods, Study design and approvals“The study was conducted in accordance with the Declaration of Helsinki and International Conference on the Harmonization of Good Clinical Practice guidelines as well as in compliance with local and institutional regulations.”→ Consider adding a comma after 'guidelines' for clarity.Readability improvement.
The published work is methodologically robust overall, but the inconsistent p-values for the HR 0.36 (reported 0.167/0.146 vs recomputed 0.0485) constitute a decision error that warrants an erratum or independent re-analysis; readers should weigh this against the otherwise strong design and reporting. Minor gaps in reagent identification and demographics are also worth noting.
- 1.HIGHstatisticsIn the Results, re-verify the reported p-values for the hazard ratio 0.36 (95% CI 0.13–0.99): reported p=0.167 and p=0.146 differ from the recomputed p=0.0485, which changes the significance decision.This is a demonstrable statistical error that affects the study's conclusions and warrants a correction or erratum.
- 2.HIGHstatisticsReplace threshold-only p-values such as 'P < 0.001' with exact values where available in the Results.Reviewer 2 noted imprecise reporting; exact values support reproducibility and reader verification.
- 3.HIGHotherAdd manufacturer, dose, and regimen details for pembrolizumab and axitinib in the Methods section.Both reviewers flagged incomplete identification of investigational products, which hampers reproducibility.
- 4.HIGHotherProvide details of FMT capsule preparation including excipients and storage conditions in the Methods.Improves reproducibility of the intervention and addresses Reviewer 1's concern.
- 5.HIGHreportingReport race/ethnicity and weight or performance status (e.g., Karnofsky) in Table 1.Reviewer 2 identified these as missing biological variables; adds completeness for readers.
- 6.MEDIUMdata codeDeposit custom analysis code (e.g., strain engraftment pipeline) in a public repository with a persistent identifier.Both reviewers noted no code repository; increases reproducibility.
- 7.MEDIUMreportingAdd a statement on clinical dataset availability under a data use agreement in the Data Availability section.Reviewer 1 noted that patient-level data access policy could be more explicit.
- 8.MEDIUMcopyeditAdd version numbers for the Python packages used in statistical analysis in the Methods.Copyedit flagged missing package versions, improving reproducibility.
- 9.LOWcopyeditAdd a comma after 'guidelines' in the Methods sentence about compliance with the Declaration of Helsinki and ICH-GCP.Minor readability improvement flagged by copyedit.
- 10.LOWcopyeditConsider harmonizing the reporting of P=0.053 in the Abstract with other sections.Copyedit noted consistency; already acceptable but could be standardized.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.