Fecal microbiota transplantation plus pembrolizumab and axitinib in metastatic renal cell carcinoma: the randomized phase 2 TACITO trial.
Porcari S, Ciccarese C, Heidrich V, Rondinella D, Quaranta G, Severino A, Arduini D, Buti S, Fornarini G, Primi F, Stumbo L, Giannarelli D, Giudice GC, Damassi A, Giron Berríos JR, Punčochář M, Barbazuk TB, Piccinno G, Pinto F, Armanini F, Asnicar F, Schinzari G, Derosa L, Kroemer G, Sanguinetti M, Masucci L, Gasbarrini A, Tortora G, Cammarota G, Zitvogel L, Segata N, Iacovelli R, Ianiro G
- DOI
- 10.1038/s41591-025-04189-2
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/e5d8b01c-ff9c-45f9-8077-b011271f560c is authoritative.
How this rating was calculated
- StatisticsImpossible or misreported statistic ×3−3★
- StatisticsStatistic did not reproduce ×3−1.5★
- IntegrityIntegrity concern ×2−1★
- StatisticsPrinted percentage does not match its own count (capped) ×6−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- 01Printed percentage does not match its own countdemonstrable
67% does not match the reported count 33/50
“33 of 50 patients (67%)”
- 02Printed percentage does not match its own countdemonstrable
48% does not match the reported count 10/25
“p-FMT 48%”
ITT population - 03Printed percentage does not match its own countdemonstrable
67% does not match the reported count 33/50
“33 of 50 patients (67%)”
ITT populationFind in source - 04Significance claim does not survive recomputation
Recomputed hazard ratio 0.36 (95% CI 0.13–0.99), reported p=0.167
“hazard ratio = 0.36, 95% CI: 0.13–0.99, P = 0.167”
- 05Significance claim does not survive recomputation
Recomputed HR 0.36 (95% CI 0.13–0.99), reported p=0.146
“HR 0.36, 95% CI 0.13-0.99, p = 0.146”
- 06Significance claim does not survive recomputation
Recomputed hazard ratio 0.36 (95% CI 0.13–0.99), reported p=0.146
“hazard ratio = 0.36 (95% CI: 0.13–0.99), P = 0.146”
6 further findings of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper reports a well-designed phase 2a RCT with strong methodological reporting across most dimensions. The primary concern is the statistical verification finding 12 of 19 recomputed tests inconsistent, including 3 decision errors, which warrants a correction or independent re-analysis. Copyedit issues are minor.
Both reviewers agreed on study type (interventional). The evaluation is based on the full text. The statistics verification covered only 19 tests with sufficient information for recomputation; threshold-only and resampling-based p-values were not checked. No cell-line or animal-subject issues were applicable.
Numerical inconsistencies
3 findings · worst criticalValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 10 tests: 7 consistent, 3 inconsistent (3 change significance at p<.05); 8 recomputed directly from the reported test statistics, 2 via agent-written checks. 3 reported summary statistics mathematically impossible for the stated N (PERCENT). 6 printed percentages that do not match their own count.
- PERCENT67% does not match the reported count 33/50
“33 of 50 patients (67%)”
- PERCENT48% does not match the reported count 10/25
“p-FMT 48%”
ITT population - PERCENT28% does not match the reported count 7/22
“28% (7/22 patients)”
ORR ITTFind in source - PERCENT67% does not match the reported count 33/50
“33 of 50 patients (67%)”
ITT populationFind in source - PERCENT44% does not match the reported count 10/22
“p-FMT: 44%”
Safety assessment - PERCENT20% does not match the reported count 5/23
“d-FMT: 20%”
Safety assessment - PERCENT12% does not match the reported count 3/23
“d-FMT: 12%”
Safety assessment - PERCENT12% does not match the reported count 3/23
“d-FMT: 12%”
Safety assessment - PERCENT8% does not match the reported count 2/22
“p-FMT: 8%”
Safety assessment
- CONSISTENTreported p = .048 · recomputed p = .049Recomputed hazard ratio 0.48 (95% CI 0.23–0.99), reported p=0.048
“hazard ratio = 0.48, 95% CI: 0.23–0.99, P = 0.048”
Taken as given: 0.23–0.99 is a two-sided 95% confidence interval for the hazard ratio of 0.48, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.048 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.48, 0.23, 0.99, 1) - CONSISTENTreported p = .180 · recomputed p = .248Recomputed hazard ratio 0.66 (95% CI 0.33–1.35), reported p=0.18
“hazard ratio = 0.66, 95% CI: 0.33–1.35, P = 0.18”
Taken as given: 0.33–1.35 is a two-sided 95% confidence interval for the hazard ratio of 0.66, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.18 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.66, 0.33, 1.35, 1) - CONSISTENTreported p = .180 · recomputed p = .248Recomputed HR 0.66 (95% CI 0.33–1.35), reported p=0.18
“HR: 0.66, 95% CI: 0.33-1.35, p = 0.18”
Taken as given: 0.33–1.35 is a two-sided 95% confidence interval for the HR of 0.66, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.18 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.66, 0.33, 1.35, 1) - CONSISTENTreported p = .146 · recomputed p = .158Recomputed HR 0.51 (95% CI 0.20–1.30), reported p=0.146
“HR 0.51, 95% CI: 0.20-1.30, p = 0.146”
Taken as given: 0.20–1.30 is a two-sided 95% confidence interval for the HR of 0.51, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.146 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.51, 0.2, 1.3, 1) - CONSISTENTreported p = .146 · recomputed p = .158Recomputed hazard ratio 0.51 (95% CI 0.20–1.30), reported p=0.146
“hazard ratio = 0.51, 95% CI: 0.20–1.30, P = 0.146”
Taken as given: 0.20–1.30 is a two-sided 95% confidence interval for the hazard ratio of 0.51, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.146 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.51, 0.2, 1.3, 1) - CONSISTENTreported p = .053 · recomputed p = .075Reviewers 1, 2Primary endpoint comparison (12-month PFS) using Fisher's exact test on reported counts.
“d-FMT: 16/23 patients, 70%; p-FMT: 9/22 patients, 41%; P = 0.053”
Taken as given: The counts are 16 events in d-FMT (n=23) and 9 events in p-FMT (n=22).; The test is two-sided Fisher's exact test.; The table is 2x2 with events and non-events.Method: Two-sided Fisher's exact test on the 2x2 table (16,7,9,13).How we recomputed it: pFisher2x2(16, 7, 9, 13, 0) - CONSISTENTreported p = .035 · recomputed p = .027Reviewers 1, 2Secondary endpoint median PFS comparison using log-rank test (approximated from HR and CI).
“hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035”
Taken as given: The hazard ratio is 0.50 with a 90% confidence interval (0.27, 0.92).; The p-value is derived from the log-rank test, approximated from the CI.; The CI is two-sided at 90%.Method: Approximated two-sided p-value from the hazard ratio and its 90% CI using the normal approximation.How we recomputed it: pCI(0.50, 0.27, 0.92, 1)
- lowinternal contradictionThe safety analysis population is described as 49 patients, but the adverse event rates are reported with denominators 25 and 24, which sum to 49. This is consistent, but the text says '49 patients (all randomized patients except one patient who withdrew his consent)' while the denominators suggest 25 and 24, which is fine.
“Overall, 49 patients (all randomized patients except one patient who withdrew his consent) were included in the safety analysis”
ResultsFind in source - lowinternal contradictionThe abstract reports ORR as 52% for d-FMT and 32% for placebo, but the ITT analysis reports ORR as 52% (13/25) for d-FMT and 28% (7/22) for p-FMT. The discrepancy may be due to different analysis populations (FAS vs ITT).
The ORR was 52% of patients in the d-FMT arm and 32% of patients receiving placebo. ... The ORR in the ITT population was 52% (13/25 patients, 95% CI: 0.33–0.70) in the d-FMT arm and 28% (7/22 patients, 95% CI: 0.14–0.55) in the p-FMT arm.
Abstractreviewer’s wording
Overstated conclusions
1 finding · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Significance claim flips when recomputedRecomputed
- MORE SIGNIFICANT ON RECHECKINCONSISTENTreported p = .167 · recomputed p = .049Recomputed hazard ratio 0.36 (95% CI 0.13–0.99), reported p=0.167Recomputing from the paper’s own numbers lands below p = 0.05 — more significant than the printed value. Usually benign (the reported figure is conservative), but the two don’t match.
“hazard ratio = 0.36, 95% CI: 0.13–0.99, P = 0.167”
Taken as given: 0.13–0.99 is a two-sided 95% confidence interval for the hazard ratio of 0.36, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.167 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.36, 0.13, 0.99, 1) - MORE SIGNIFICANT ON RECHECKINCONSISTENTreported p = .146 · recomputed p = .049Recomputed HR 0.36 (95% CI 0.13–0.99), reported p=0.146Recomputing from the paper’s own numbers lands below p = 0.05 — more significant than the printed value. Usually benign (the reported figure is conservative), but the two don’t match.
“HR 0.36, 95% CI 0.13-0.99, p = 0.146”
Taken as given: 0.13–0.99 is a two-sided 95% confidence interval for the HR of 0.36, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.146 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.36, 0.13, 0.99, 1) - MORE SIGNIFICANT ON RECHECKINCONSISTENTreported p = .146 · recomputed p = .049Recomputed hazard ratio 0.36 (95% CI 0.13–0.99), reported p=0.146Recomputing from the paper’s own numbers lands below p = 0.05 — more significant than the printed value. Usually benign (the reported figure is conservative), but the two don’t match.
“hazard ratio = 0.36 (95% CI: 0.13–0.99), P = 0.146”
Taken as given: 0.13–0.99 is a two-sided 95% confidence interval for the hazard ratio of 0.36, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.146 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.36, 0.13, 0.99, 1)
9 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Donor FMT significantly improved median PFS compared to placebo.The reported median PFS difference (24.0 vs 9.0 months) with HR 0.50 and P=0.035 supports this claim.Evidence: Median PFS was significantly improved in the d-FMT arm (24.0 months, 95% CI: 8.0–40.0 months) compared with the p-FMT arm (9.0 months, 95% CI: 2.2–15.2 months) (hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035).
“Median PFS was significantly improved in the d-FMT arm (24.0 months, 95% CI: 8.0–40.0 months) compared with the p-FMT arm (9.0 months, 95% CI: 2.2–15.2 months) (hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035).”
ResultsFind in source - supportedReviewer 1The primary endpoint of 12-month PFS was not met.The reported P-value of 0.053 is above the conventional 0.05 threshold, supporting the claim that the primary endpoint was not met.Evidence: The proportion of patients without progression or death 12 months after randomization was higher in the d-FMT arm than in the p-FMT arm (d-FMT: 16/23 patients, 70%; p-FMT: 9/22 patients, 41%; P = 0.053).
“The proportion of patients without progression or death 12 months after randomization was higher in the d-FMT arm than in the p-FMT arm (d-FMT: 16/23 patients, 70%; p-FMT: 9/22 patients, 41%; P = 0.053).”
ResultsFind in source - supportedReviewer 1Donor FMT increased alpha-diversity and beta-diversity compared to placebo.The reported increases in Shannon diversity and Bray-Curtis dissimilarity with significant p-values support this claim.Evidence: After treatments, we observed an increase in Shannon α-diversity versus baseline at week 1 (P = 0.05), week 4 (P < 0.001), week 12 (P = 0.02) and week 24 (P = 0.048) follow-ups in the d-FMT group. Also, significantly higher Bray-Curtis dissimilarity in d-FMT versus p-FMT at week 4 (P = 0.015).
“After treatments, we observed an increase in Shannon α-diversity versus baseline at week 1 ( P = 0.05), week 4 ( P < 0.001), week 12 ( P = 0.02) and week 24 ( P = 0.048) follow-ups (Fig. ) and significantly increased species richness versus baseline at week 4 follow-up ( P = 0.001; Extended Data Fig. ) in the d-FMT group.”
ResultsFind in source - supportedReviewer 1Donor strain engraftment was higher in the d-FMT arm.The DoSER was consistently higher in the d-FMT arm with P < 0.001 across timepoints, supporting this claim.Evidence: The DoSER in patients receiving d-FMT was consistently higher than in the p-FMT arm throughout the whole study period ( P < 0.001 across all timepoints starting from week 1 follow-up).
The DoSER in patients receiving d-FMT was consistently higher than in the p-FMT arm throughout the whole study period ( P < 0.001 across all timepoints starting from week 1 follow-up).
Resultsreviewer’s wording - supportedReviewers 1, 2Acquisition or loss of specific strains, but not total engraftment, was associated with the primary endpoint.The paper reports associations for specific strains (e.g., B. wexlerae, A. massiliensis) and no association for DoSER, supporting this claim.Evidence: The acquisition of the Blautia wexlerae (SGB4837) strain from the donor at week 1 was positively associated with 12-month PFS (50% versus 0%, P = 0.047). On the other hand, acquisition of the donor strain of a yet-to-be-described species (SGB14845) of the family Oscillospiraceae (7% versus 71%, P = 0.006) was inversely associated. Also, DoSER was not associated with PFS (P = 0.15 and P = 0.53).
“The acquisition of the Blautia wexlerae (SGB4837) strain from the donor at week 1 was positively associated with 12-month PFS (percentage of recipients acquiring strain with versus without PFS > 12 months: 50% versus 0%, P = 0.047; Fig. ).”
ResultsFind in source - supportedReviewers 1, 2Donor FMT is safe with no FMT-related SAEs.The safety data reported no FMT-related serious adverse events, supporting this claim.Evidence: No deaths related to experimental treatments were reported. No transmission of any infectious agent after d-FMT was observed.
“No deaths related to experimental treatments were reported. No transmission of any infectious agent after d-FMT was observed.”
ResultsFind in source - supportedReviewer 2Donor FMT did not significantly improve the primary endpoint of 12-month PFS.The claim is supported by the reported P=0.053, which is above the conventional threshold.Evidence: The proportion of patients without progression or death 12 months after randomization was higher in the d-FMT arm than in the p-FMT arm (d-FMT: 16/23 patients, 70%; p-FMT: 9/22 patients, 41%; P = 0.053).
“The proportion of patients without progression or death 12 months after randomization was higher in the d-FMT arm than in the p-FMT arm (d-FMT: 16/23 patients, 70%; p-FMT: 9/22 patients, 41%; P = 0.053).”
ResultsFind in source - supportedReviewer 2Donor FMT increased α-diversity and β-diversity compared with placebo.The claim is supported by reported increases in Shannon diversity and Bray-Curtis dissimilarity in the d-FMT arm.Evidence: After treatments, we observed an increase in Shannon α-diversity versus baseline at week 1 (P = 0.05), week 4 (P < 0.001), week 12 (P = 0.02) and week 24 (P = 0.048) follow-ups ... and significantly higher Bray–Curtis dissimilarity between posttreatment and baseline microbiomes in d-FMT versus p-FMT patients at the week 4 follow-up (P = 0.015).
After treatments, we observed an increase in Shannon α-diversity versus baseline at week 1 (P = 0.05), week 4 (P < 0.001), week 12 (P = 0.02) and week 24 (P = 0.048) follow-ups
Resultsreviewer’s wording - supportedReviewer 2Donor strain engraftment was higher in the d-FMT arm than in the p-FMT arm.The claim is supported by the DoSER measurements showing consistently higher engraftment in d-FMT.Evidence: The DoSER in patients receiving d-FMT was consistently higher than in the p-FMT arm throughout the whole study period (P < 0.001 across all timepoints starting from week 1 follow-up).
The DoSER in patients receiving d-FMT was consistently higher than in the p-FMT arm throughout the whole study period (P < 0.001 across all timepoints starting from week 1 follow-up).
Resultsreviewer’s wording
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary endpoint is 12-month progression-free survival (PFS), a hard clinical outcome. Although the primary endpoint was not met (P=0.053), the secondary endpoint of median PFS was significantly improved (HR=0.50, P=0.035). PFS is a validated clinical endpoint in oncology, and the trial also reports overall survival and objective response rate. The efficacy claim is based on these clinical outcomes, not solely on a surrogate biomarker.
“The primary endpoint was the rate of patients free from disease progression at 12 months after randomization (12-month progression-free survival (PFS)).”
- ADEQUATEEffect sizeThe effect size for median PFS is substantial: 24.0 months in the d-FMT arm versus 9.0 months in the placebo arm, with a hazard ratio of 0.50 (P=0.035). This represents a clinically meaningful improvement in a hard endpoint, and the magnitude is anchored to clinical benefit. The primary endpoint difference (70% vs 41%) was borderline significant, but the secondary endpoint provides strong evidence of efficacy.
“Median PFS was significantly improved in the d-FMT arm (24.0 months, 95% CI: 8.0–40.0 months) compared with the p-FMT arm (9.0 months, 95% CI: 2.2–15.2 months) (hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior evidence on microbiome-ICI interactions, including antibiotic use and microbial taxa associations, and notes the lack of RCTs in mRCC. The rationale for using FMT from complete responders is logically connected to the study objectives. Limitations of prior work are implicitly addressed by the randomized design and focus on mRCC.
“Our phase 2a placebo-controlled RCT aims to evaluate whether FMT from patients with mRCC with complete response to ICIs was effective in improving response to combined first-line therapy with pembrolizumab and axitinib in patients with mRCC.”
“However, thus far, no randomized controlled trials (RCTs) have demonstrated the efficacy of FMT in mRCC.”
“Our phase 2a placebo-controlled RCT aims to evaluate whether FMT from patients with mRCC with complete response to ICIs was effective in improving response to combined first-line therapy with pembrolizumab and axitinib in patients with mRCC.”
“However, thus far, no randomized controlled trials (RCTs) have demonstrated the efficacy of FMT in mRCC.”
Randomization used random permuted blocks with a block size of four, and the sequence was hidden. Blinding was double-blind with placebo capsules identical in appearance and blinded outcome assessors. Sample size calculation was based on a superiority hypothesis with 80% power. Inclusion/exclusion criteria were pre-specified, and analysis populations (ITT, FAS, per-protocol) were defined. Outlier handling is addressed through predefined withdrawal criteria and per-protocol analysis.
“An online random number generator software ( https://www.sealedenvelope.com ) was used to provide random permuted blocks with a block size of four and an equal allocation ratio; the sequence was hidden until the interventions were assigned.”
“Placebo capsules were made of cellulose and were identical in appearance to the FMT capsules, to ensure the blinding of patients and study staff.”
“A total of 50 patients is required to enter this two-treatment, parallel-design study. The probability is 80% that the study detects a treatment difference at a one-sided 5.0% significance level, if the true hazard ratio is 0.436.”
“An online random number generator software ( https://www.sealedenvelope.com ) was used to provide random permuted blocks with a block size of four and an equal allocation ratio; the sequence was hidden until the interventions were assigned.”
“To mask treatments to recipients, both infusate bottles and syringes were covered with dark-colored paper before the infusion, and the patients were sedated.”
“A total of 50 patients is required to enter this two-treatment, parallel-design study. The probability is 80% that the study detects a treatment difference at a one-sided 5.0% significance level, if the true hazard ratio is 0.436.”
Sex is reported (73% male), age is reported (median 62 years), and demographics include IMDC risk groups and histology. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“Most participants were male (73%), and the median age at the time of treatment initiation was 62 years (range, 41–79 years).”
“According to the International Metastatic RCC Database Consortium (IMDC) risk model, 69% of participants had intermediate-prognosis and poor-prognosis disease.”
“Most participants were male (73%), and the median age at the time of treatment initiation was 62 years (range, 41–79 years).”
“Clear cell | 21 (91%) | 19 (86%)”
The study was approved by the institutional review board (IRB)/local ethics committee with ID 2664, and all patients gave written informed consent. Compliance with the Declaration of Helsinki and ICH-GCP is stated. This is a human trial, so iacuc_statement is not applicable.
“The study was approved by the institutional review board (IRB)/local ethics committee (ID: 2664)”
“All enrolled patients gave their written informed consent to participate in the study.”
“The study was conducted in accordance with the Declaration of Helsinki and International Conference on the Harmonization of Good Clinical Practice guidelines”
“The study was approved by the institutional review board (IRB)/local ethics committee (ID: 2664)”
“All enrolled patients gave their written informed consent to participate in the study.”
“The study was conducted in accordance with the Declaration of Helsinki and International Conference on the Harmonization of Good Clinical Practice guidelines”
The FMT product is described in detail including donor screening, preparation, and administration. Key reagents for DNA extraction and sequencing are identified with vendors. Software tools are named with versions. Antibodies, cell lines, and mycoplasma testing are not applicable as this is a clinical trial without wet-lab assays.
“pembrolizumab + axitinib”
“IBM-SPSS version 28.0 statistical software, GraphPad Prism version 10, R version 4.4.2 (‘survival’ and ‘survminer’ packages) and Python version 3.10.12 (‘scikit-bio’ and ‘scipy’ packages) were used for the analyses.”
“Patients received a styrofoam box, filled with dry ice and 10 capsule containers. Each container included 15 capsules.”
“DNA extraction consisted of sample homogenization followed by DNA isolation with the PowerSoil Pro DNA Isolation Kit (Qiagen).”
“MetaPhlAn 4 (version 4.1)”
Tests are named (Breslow test, Cox proportional hazard, chi-square, Mann-Whitney, Wilcoxon signed-rank, Fisher's exact, PERMANOVA). Assumptions for Cox are checked. Exact p-values are reported (e.g., P = 0.035). Effect sizes with CIs are provided (HR with 95% CI). Software is identified. Data presentation includes Kaplan-Meier curves and per-group n. Mathematical plausibility is not applicable for large-N continuous outcomes.
“Survival curves were estimated with the Kaplan–Meier method and compared with the Breslow test.”
“hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035”
“Median PFS was significantly improved in the d-FMT arm (24.0 months, 95% CI: 8.0–40.0 months) compared with the p-FMT arm (9.0 months, 95% CI: 2.2–15.2 months)”
“Survival curves were estimated with the Kaplan–Meier method and compared with the Breslow test.”
“hazard ratio = 0.50, 90% CI: 0.27–0.92, P = 0.035”
“Median PFS was significantly improved in the d-FMT arm (24.0 months, 95% CI: 8.0–40.0 months) compared with the p-FMT arm (9.0 months, 95% CI: 2.2–15.2 months)”
The metagenomic data are deposited in the ENA with accession PRJEB94043. Patient-level data are available under a data use agreement for IRB-approved research, with a stated response timeframe. Code availability is not explicitly stated, but the preprocessing pipeline is available on GitHub.
“The shotgun metagenomic data generated in this study are available at the European Nucleotide Archive under accession number PRJEB94043”
“The minimum dataset, without individual patient data, used for the primary, secondary and post hoc analyses, may be shared under a data use agreement for IRB-approved research.”
“available at https://github.com/SegataLab/preprocessing”
“The shotgun metagenomic data generated in this study are available at the European Nucleotide Archive under accession number PRJEB94043”
“The minimum dataset, without individual patient data, used for the primary, secondary and post hoc analyses, may be shared under a data use agreement for IRB-approved research.”
“available at https://github.com/SegataLab/preprocessing”
The trial is registered at ClinicalTrials.gov (NCT04758507). A CONSORT checklist is provided. All outcomes are reported, including negative results. Limitations are explicitly discussed. Conclusions are proportional, noting the primary endpoint was not met. Funding and competing interests are disclosed.
“ClinicalTrials.gov identifier: NCT04758507”
“This study was conducted by following CONSORT guidelines , and a CONSORT checklist is provided in Supplementary Table .”
“This study has some limitations. The sample size was relatively small, with only 45 patients evaluated for the primary endpoint.”
“prospectively registered at ClinicalTrials.gov (registration identifier: NCT04758507”
“This study was conducted by following CONSORT guidelines , and a CONSORT checklist is provided in Supplementary Table .”
“This study has some limitations. The sample size was relatively small, with only 45 patients evaluated for the primary endpoint.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 65 references by DOI: 1 verified — 64 no DOI (shown, not verified).
- NO DOIA gut microbial signature for combination immune checkpoint blockade across cancer typesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAntibiotics as deep modulators of gut microbiota: between good and evilNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA systematic review and meta-analysis evaluating the impact of antibiotic use on the clinical outcomes of cancer patients treated with immune checkpoint inhibitorsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINegative association of antibiotics on clinical activity of immune checkpoint inhibitors in patients with advanced renal cell and non-small-cell lung cancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAnticancer immunotherapy by CTLA-4 blockade relies on the gut microbiotaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFecal microbiota transplant promotes response in immunotherapy-refractory melanoma patientsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFecal microbiota transplant overcomes resistance to anti-PD-1 therapy in melanoma patientsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFecal microbiota transplantation improves anti-PD-1 inhibitor efficacy in unresectable or metastatic solid cancers refractory to anti-PD-1 inhibitorNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFecal microbiota transplantation plus anti-PD-1 immunotherapy in advanced melanoma: a phase I trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntegrating taxonomic, functional, and strain-level profiling of diverse microbial communities with bioBakery 3No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExtending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAkkermansia beyond muciniphila—emergence of new species Akkermansia massiliensis sp. nov.No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPangenomic analysis identifies correlations between Akkermansia species and subspecies and human health outcomesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRandomised clinical trial: faecal microbiota transplantation by colonoscopy vs. vancomycin for the treatment of recurrent Clostridium difficile infectionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRandomised clinical trial: faecal microbiota transplantation by colonoscopy plus vancomycin for the treatment of severe refractory Clostridium difficile infection—single versus multiple infusionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFaecal microbiota transplantation for the treatment of diarrhoea induced by tyrosine-kinase inhibitors in patients with metastatic renal cell carcinomaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIncidence of bloodstream infections, length of hospital stay, and survival in patients with recurrent Clostridioides difficile infection treated with fecal microbiota transplantation or antibiotics: a prospective cohort studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFecal microbiota transplantation for recurrent Clostridioides difficile infection in patients with concurrent ulcerative colitisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA necessary discussion after transmission of multidrug-resistant organisms through faecal microbiota transplantationsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISystematic review: the global incidence of faecal microbiota transplantation-related adverse events from 2000 to 2020No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGut-microbiome signatures predicting response to neoadjuvant chemoradiotherapy in locally advanced rectal cancer: a systematic reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOral administration of Blautia wexlerae ameliorates obesity and type 2 diabetes via metabolic remodeling of the gut microbiotaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntestinal Akkermansia muciniphila predicts clinical response to PD-1 blockade in patients with advanced non-small-cell lung cancerNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICustom scoring based on ecological topology of gut microbiota associated with cancer immunotherapy outcomeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICross-cohort gut microbiome associations with immune checkpoint inhibitor response in advanced melanomaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISurrogate markers of intestinal dysfunction associated with survival in advanced cancersNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICorrection: Experimental evaluation of ecological principles to understand and modulate the outcome of bacterial strain competition in gut microbiomesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPathogenic Escherichia coliNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGenomic features and prevalence of Ruminococcus species in humans are associated with age, lifestyle, and diseaseNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIVariability of strain engraftment and predictability of microbiome composition after fecal microbiota transplantation across different diseasesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDonor screening for fecal microbiota transplantation with a direct stool testing-based strategy: a prospective cohort studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe use of faecal microbiota transplantation (FMT) in Europe: a Europe-wide surveyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICabozantinib and nivolumab with or without live bacterial supplementation in metastatic renal cell carcinoma: a randomized phase 1 trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINivolumab plus ipilimumab with or without live bacterial supplementation in metastatic renal cell carcinoma: a randomized phase 1 trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFirst-in-class Microbial Ecosystem Therapeutic 4 (MET4) in combination with immune checkpoint inhibitors in patients with advanced solid tumors (MET4-IO trial)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICONSORT 2025 statement: updated guideline for reporting randomized trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAssessing engraftment following fecal microbiota transplantNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINew response evaluation criteria in solid tumours: revised RECIST guideline (version 1.1)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInternational consensus conference on stool banking for faecal microbiota transplantation in clinical practiceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReorganisation of faecal microbiota transplant services during the COVID-19 pandemicNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEfficacy of different faecal microbiota transplantation protocols for Clostridium difficile infection: a systematic review and meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILaboratory handling practice for faecal microbiota transplantationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffect of oral capsule vs colonoscopy-delivered fecal microbiota transplantation on recurrent Clostridium difficile infection: a randomized clinical trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFecal microbiota transplantation capsules with targeted colonic versus gastric delivery in recurrent Clostridium difficile infection: a comparative cohort analysis of high and low doseNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe person-to-person transmission landscape of the gut and oral microbiomesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGlobal genetic diversity of human gut microbiome species is related to geographic location and host healthNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIdentification and assembly of genomes and genetic elements in complex metagenomic samples without using reference genomesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe dynamics of the human infant gut microbiome in development and in progression toward type 1 diabetesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDynamics and stabilization of the human gut microbiome during the first year of lifeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICharacterization of the gut microbial community of obese patients following a weight-loss intervention using whole metagenome shotgun sequencingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIVariation in microbiome LPS immunogenicity contributes to autoimmunity in humansNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntegrated multi-omics of the human gut microbiome in a case study of familial type 1 diabetesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMaturation of the infant microbiome community structure and function across multiple body sites and in relation to mode of deliveryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStudying vertical microbiome transmission from mothers to infants by strain-level metagenomic profilingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA novel Ruminococcus gnavus clade enriched in inflammatory bowel disease patientsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISubspecies in the global human gut microbiomeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDynamics of metatranscription in the inflammatory bowel disease gut microbiomeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStability of the human faecal microbiome in a cohort of adult menNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMother-to-infant microbial transmission from different body sites shapes the developing infant gut microbiomeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStrain-level analysis of mother-to-child bacterial transmission during the first few months of lifeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBirth mode is associated with earliest strain-conferred gut microbiome functions and immunostimulatory potentialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMulti-omics of the gut microbial ecosystem in inflammatory bowel diseasesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStunted microbiota and opportunistic pathogen colonization in caesarean-section birthNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA new method for non-parametric multivariate analysis of varianceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
4 data/code links checked; 4 live.
- dataENALIVEHTTP 200http://www.ebi.ac.uk/ena/data/view/PRJEB94043Resolves to ENA (data repository).
- datahttp://clinicaltrials.gov/study/NCT04758507LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/study/NCT04758507LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/SegataLab/preprocessingResolves to GitHub (code repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORconsistencyAbstract“P = 0.053”→ Consider reporting the exact p-value as 0.053 or as 'P = 0.053' consistently.The p-value is reported as 0.053 in the abstract and results; ensure consistency.
- MINORtypoMethods, Analysis sets“consent withrawal”→ Change to 'consent withdrawal'.Typographical error.
- MINORconsistencyResults, Safety assessment“28% (7/25) of the patients in the d-FMT group and in 16% (4/24) of the patients in the p-FMT group”→ Ensure the denominators (25 and 24) are consistent with the safety population described elsewhere.The safety population is stated as 49 patients, but the denominators here sum to 49; verify consistency.
- MINORconsistencyAbstract“P = 0.053”→ Consider reporting the p-value as 0.053 or 0.05 to two decimal places for consistency with other p-values.Minor inconsistency in decimal places.
- MINORclarityResults, Microbiome changes“we observed a significant increase in both α-diversity and β-diversity in the d-FMT arm compared with the p-FMT arm after treatments (secondary outcomes of our study).”→ Clarify that these are secondary outcomes and specify the timepoints.Could be clearer.
The published work is methodologically strong in design and reporting, but the statistical inconsistencies (12/19 tests inconsistent, 3 decision errors) are a substantive concern. An informed reader should weigh these errors heavily; a correction or independent re-analysis of the affected endpoints is warranted.
- 1.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 67% does not match the reported count 33/50Demonstrable critical failure — blocks the verdict from passing.
- 2.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 48% does not match the reported count 10/25Demonstrable critical failure — blocks the verdict from passing.
- 3.HIGHstatisticsRe-analyze and correct the three tests where the reported p-value crosses the significance threshold upon recomputation (e.g., reported p=0.167 recomputed as p=0.049).These decision errors directly affect the interpretation of the corresponding endpoints and may change the conclusions.
- 4.HIGHstatisticsProvide a full corrected statistical analysis for all 19 tests flagged as inconsistent, and explain the source of the discrepancies (e.g., rounding, software version, or analytical error).12 of 19 recomputed tests were inconsistent, undermining confidence in the reported numerical results.
- 5.HIGHreportingIn the Data Availability section, clarify the exact terms of the data use agreement and the process for requesting the minimum dataset.The current description is vague; reviewers may require more detail to assess reproducibility.
- 6.HIGHreportingIn the Methods, specify the version of the CONSORT checklist used (e.g., CONSORT 2010) and ensure the checklist is provided as a supplementary file.The paper references a CONSORT checklist but does not state the version, which is a minor reporting gap.
- 7.HIGHreportingIn the Discussion, explicitly acknowledge the post hoc nature of the subgroup analysis and the risk of type I error due to multiple testing.Post hoc analyses without correction for multiplicity can be misleading; transparency is essential.
- 8.MEDIUMreportingIn the Abstract, report the exact p-value for the primary endpoint (P=0.053) rather than a threshold, to avoid ambiguity.The borderline p-value is more informative when reported exactly.
- 9.MEDIUMreportingIn the Methods, describe the blinding of the statisticians who performed the analyses.Blinding of analysts is a best practice to reduce bias; its absence is a minor reporting gap.
- 10.MEDIUMreportingIn the Results, report the number of patients who discontinued treatment due to adverse events in each arm.This enhances transparency about tolerability and is standard for RCTs.
- 11.MEDIUMreportingIn the Discussion, address the potential impact of the single-donor FMT design on generalizability and suggest future multi-donor strategies.Single-donor FMT limits generalizability; acknowledging this strengthens the discussion.
- 12.LOWcopyeditFix the typo 'consent withrawal' to 'consent withdrawal' in Methods, Analysis sets.Typographical error detracts from professionalism.
- 13.LOWcopyeditClarify the timepoints for the reported microbiome diversity changes in Results, Microbiome changes.The current phrasing is vague; specifying timepoints improves clarity.
- 14.LOWreportingConsider providing a data dictionary for the microbiome data deposited in ENA.A data dictionary facilitates reuse and reproducibility.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.