Genomically matched therapy in advanced solid tumors: the randomized phase 2 ROME trial.
Marchetti P, Curigliano G, Biffoni M, Lonardi S, Scagnoli S, Fornaro L, Guarneri V, De Giorgi U, Ascierto PA, Blandino G, D'Amati G, Aglietta M, Cremolini C, Conte P, Crimini E, Ceracchi M, Pisegna S, Verkhovskaia S, Bordonaro R, Bracarda S, Butturini G, Del Mastro L, DeCensi A, Fabbri A, Fenocchio E, Gori S, Metro G, Pessino A, Pozzessere D, Puglisi F, Tamberi S, Zambelli A, Marino D, Capoluongo E, Cappuzzo F, Cerbelli B, Giannini G, Malapelle U, Mazzuca F, Nuti M, Pruneri G, Simmaco M, Strigari L, Tonini G, Martini N, Botticelli A, ROME trial investigators consortia
- DOI
- 10.1038/s41591-025-03918-x
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/2ce5cbb4-0f0e-4a44-8673-e49d4adf6142 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsOverstated claim−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- LinksDead data/code link−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is ORR, a surrogate for clinical benefit. The paper does not demonstrate target engagement at the tested doses (no PK/PD data) and does not cite validated evidence linking ORR to improved survival or quality of life in this context. The claim of improved outcomes rests on this surrogate without establishing a validated surrogate-to-clinical-outcome link.
“Overall response rate (ORR) was the primary endpoint”
- 02Treatment effect not shown to be clinically meaningful
The primary effect is a modest absolute improvement in ORR of 7.5 percentage points (17.5% vs 10%) and a median PFS improvement of 0.7 months (3.5 vs 2.8 months). These are small fractions of the reference values and are not anchored to a minimal clinically important difference or demonstrated biological meaningfulness. The paper itself describes the ORR improvement as 'modest in magnitude'.
“The improvement in ORR, the primary endpoint of the study, observed in the TT arm compared to the SoC arm within the ITT population (17.5% versus 10%), is clinically relevant and methodologically robust, although modest in magnitude.”
- 03Conclusion reaches beyond the evidence
The ROME trial provides the strongest evidence supporting MTB implementation in clinical practice.
“The ROME trial provides, to our knowledge, the strongest evidence supporting MTB implementation in clinical practice, thanks to its agnostic and randomized design.”
DiscussionFind in source - 04Declared data/code link does not resolve
Dead link — nothing to verify.
“https://www.esmo.org/guidelines/esmo-scale-for-clinical-actionability-of-molecular-targets-escat”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The ROME trial is a well-designed, randomized phase 2 study with clear scientific premise, rigorous methodology, and transparent reporting. Minor issues include inconsistent crossover rates and a typo, but these do not undermine the core findings.
Both reviewers classified the study as interventional, and no divergence was noted. The evaluation covered the full text, with N/A for non-applicable items (e.g., animal housing, cell lines). The statistics verification covered only a subset of tests (3 recomputed, all consistent); other statistics remain unverified.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 1 recomputed directly from the reported test statistics, 2 via agent-written checks.
- CONSISTENTreported p = .532 · recomputed p = .515Recomputed HR 0.92 (95% CI 0.72–1.19), reported p=0.5322
“HR = 0.92, 95% CI: 0.72–1.19, P = 0.5322”
Taken as given: 0.72–1.19 is a two-sided 95% confidence interval for the HR of 0.92, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.5322 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.92, 0.72, 1.19, 1) - CONSISTENTreported p = .029 · recomputed p = .029Reviewers 1, 2Primary ORR comparison using chi-square test
“The P value from the chi-square test ( P = 0.0294; Extended Data Table )”
Taken as given: The ORR counts are 35/200 in TT and 20/200 in SoC.; The test is a two-sided Pearson chi-square test on the 2x2 table.Method: Recomputed two-sided Pearson chi-square p-value from cell counts (35,165,20,180).How we recomputed it: pChi2x2(35,165,20,180) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2PFS hazard ratio p-value from CI
“with a hazard ratio (HR) of 0.66 (95% CI: 0.53–0.82, P = 0.0002; Fig. )”
Taken as given: The HR is 0.66 with 95% CI 0.53-0.82.; The p-value is two-sided from the log-rank test, approximated from the CI.Method: Approximated two-sided p-value from HR and 95% CI using normal approximation on log scale.How we recomputed it: pCI(0.66,0.53,0.82,1)
- lowinternal contradictionCrossover rate is reported inconsistently: 52% in abstract, 59% in results, 58.8% in discussion.
with a 52% crossover rate. ... A high crossover rate (59%) from SoC to TT ... The substantial crossover rate (58.8%) from the SoC arm to the TT arm
Discussionreviewer’s wording - lowinternal contradictionFigure 2a legend reports P=0.002 for PFS, but the text reports P=0.0002 for the same comparison.
Fig. 2 Secondary endpoints in the ITT population. a , PFS ( P = 0.002). ... PFS was significantly improved in the TT arm ... with a hazard ratio (HR) of 0.66 (95% CI: 0.53–0.82, P = 0.0002
Figure 2Areviewer’s wording
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions overstated beyond the evidenceAssessed
6 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated).
- overstatedReviewers 1, 2The ROME trial provides the strongest evidence supporting MTB implementation in clinical practice.While the trial is randomized and positive, claiming 'strongest evidence' is a strong assertion not directly supported by comparative analysis of all MTB studies.Evidence: Discussion statement
“The ROME trial provides, to our knowledge, the strongest evidence supporting MTB implementation in clinical practice, thanks to its agnostic and randomized design.”
DiscussionFind in source - supportedReviewers 1, 2TT achieved a significantly higher ORR compared to SoC.The primary endpoint ORR was significantly higher in TT (17.5% vs 10%, P=0.0294), supported by the reported chi-square test.Evidence: ORR results with p-value
“TT achieved a significantly higher ORR (17.5% versus 10%; P = 0.0294)”
AbstractFind in source - supportedReviewers 1, 2TT improved median PFS compared to SoC.PFS was significantly longer in TT (3.5 vs 2.8 months, HR=0.66, P=0.0002), supported by Kaplan-Meier analysis.Evidence: PFS results with HR and p-value
“improved median PFS (3.5 months versus 2.8 months; hazard ratio = 0.66 (0.53–0.82), P = 0.0002)”
AbstractFind in source - supportedReviewers 1, 2Median OS was similar between arms.OS was not significantly different (HR=0.92, P=0.5322), consistent with the claim.Evidence: OS results with HR and p-value
“Median OS was similar, with a 52% crossover rate.”
AbstractFind in source - supportedReviewers 1, 2TT showed superior 12-month PFS rates.12-month PFS was 22.0% vs 8.3%, supporting the claim.Evidence: 12-month PFS rates
“TT also showed superior 12-month PFS rates (22.0% versus 8.3%).”
AbstractFind in source - supportedReviewers 1, 2Grade 3/4 adverse events were similar between arms.Grade 3/4 AEs were 40% vs 52%, supporting the claim of similarity.Evidence: Safety results
“Grade 3/4 adverse events were also similar (40% TT versus 52% SoC).”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is ORR, a surrogate for clinical benefit. The paper does not demonstrate target engagement at the tested doses (no PK/PD data) and does not cite validated evidence linking ORR to improved survival or quality of life in this context. The claim of improved outcomes rests on this surrogate without establishing a validated surrogate-to-clinical-outcome link.
“Overall response rate (ORR) was the primary endpoint”
- INADEQUATEEffect sizeThe primary effect is a modest absolute improvement in ORR of 7.5 percentage points (17.5% vs 10%) and a median PFS improvement of 0.7 months (3.5 vs 2.8 months). These are small fractions of the reference values and are not anchored to a minimal clinically important difference or demonstrated biological meaningfulness. The paper itself describes the ORR improvement as 'modest in magnitude'.
“The improvement in ORR, the primary endpoint of the study, observed in the TT arm compared to the SoC arm within the ITT population (17.5% versus 10%), is clinically relevant and methodologically robust, although modest in magnitude.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites multiple prior studies (SHIVA, MOSCATO-01, TAPUR, NCI-MATCH, SAFIR02-BREAST) and frameworks (ESCAT, ETAC-S), acknowledging inconsistent outcomes and gaps. The rationale for the trial—evaluating MTB-guided TT versus SoC in a randomized setting—follows logically from the cited evidence. Limitations of prior research (e.g., lack of randomized evidence, variability in actionability) are explicitly addressed.
“The nationwide ROME trial was designed to provide comprehensive genomic profiling and TT for patients with solid tumors across multiple medical oncology units in Italy. This trial aims to evaluate the outcomes of a TT established after discussions in the MTB compared to the standard of care (SoC).”
“However, definitive evidence demonstrating the superiority of these approaches over standard therapies is lacking, with inconsistent outcomes reported across randomized agnostic precision oncology clinical trials.”
“However, definitive evidence demonstrating the superiority of these approaches over standard therapies is lacking, with inconsistent outcomes reported across randomized agnostic precision oncology clinical trials.”
“The nationwide ROME trial was designed to provide comprehensive genomic profiling and TT for patients with solid tumors across multiple medical oncology units in Italy.”
Randomization method is described (SAS PROC PLAN, independent statistician, 1:1 ratio). Power analysis is detailed with assumptions (20% vs 5% ORR, α=0.10, β=0.20, 86 per stratum, 344 total, 384 with dropout, 400 final). Inclusion/exclusion criteria are extensive and prespecified. Outlier handling is addressed via ITT analysis and predefined SAP. Controls are inherent in the SoC arm. Independent replication is not applicable for a single pivotal trial. Blinding is not reported, but the open-label design is stated; however, no explicit rationale for open-label is given, which is a minor gap.
“The randomization list was generated with a dedicated SAS program using the PROC PLAN procedure by an independent statistician.”
“To detect a 15% difference between the two arms, assuming an α of 0.10 and a β of 0.20 (80% power) and employing a one-sided chi-square test, a total of 86 patients (43 in each arm of the four cohorts) were required.”
“The ROME trial was a multicenter, randomized, open-label phase 2 study comparing tailored treatment (TT) to standard of care (SoC)”
“The randomization list was generated with a dedicated SAS program using the PROC PLAN procedure by an independent statistician.”
“To detect a 15% difference between the two arms, assuming an α of 0.10 and a β of 0.20 (80% power) and employing a one-sided chi-square test, a total of 86 patients (43 in each arm of the four cohorts) were required.”
“Eligible patients had failed at least one, but no more than two, prior lines of systemic therapy.”
Sex is reported for the ITT population (48% male, 52% female). Age (median 61, range 22-85) and ECOG PS are reported. Demographics include ethnicity breakdown. Species/strain/housing are not applicable for a human trial. Sex justification is not applicable as both sexes are enrolled.
“Median (range) | 61 (22–85) | 60 (34–84) | 62 (22–85)”
“Caucasian | 392 (98.1) | 196 (98.0) | 196 (98.0)”
“Male | 192 (48.0) | 100 (50.0) | 92 (46.0)”
“Median (range) | 61 (22–85) | 60 (34–84) | 62 (22–85)”
“Caucasian | 392 (98.1) | 196 (98.0) | 196 (98.0)”
The study was approved by the institutional ethics committee of the coordinating center (Sapienza no. rif. C.E. 5575) and by each participating center. AIFA authorized the trial. All patients signed informed consent. Adherence to the Declaration of Helsinki is stated.
“The study was approved by the institutional ethics committee of the coordinating center (Sapienza no. rif. C.E. 5575; February 2020) and by the ethics committee of each participating center.”
“All patients signed the specifically conceived informed consent form (ICF).”
“The trial adhered to the principles of the Declaration of Helsinki regarding research involving human subjects.”
“The study was approved by the institutional ethics committee of the coordinating center (Sapienza no. rif. C.E. 5575; February 2020) and by the ethics committee of each participating center.”
“All patients signed the specifically conceived informed consent form (ICF).”
“The trial adhered to the principles of the Declaration of Helsinki regarding research involving human subjects.”
The trial uses multiple targeted therapies and immunotherapies; the drugs are named with manufacturers (e.g., Roche, Novartis, Pfizer). NGS tests (FoundationOne CDx and FoundationOne Liquid CDx) are identified with provider. Statistical software (SAS v9.4, R v4.3.3) is identified. Antibodies, cell lines, and mycoplasma testing are not applicable.
“Molecular profiling was performed using FoundationOne CDX and FoundationOne Liquid CDX tests on tissue and liquid biopsies.”
“All analyses were conducted using SAS software (v.9.4).”
“Erlotinib, pertuzumab, vemurafenib, trastuzumab emtansine, alectinib, vismodegib, cobimetinib, atezolizumab, trastuzumab, ipatasertib (GDC-0068), entrectinib and pralsetinib were provided by Roche;”
“Erlotinib, pertuzumab, vemurafenib, trastuzumab emtansine, alectinib, vismodegib, cobimetinib, atezolizumab, trastuzumab, ipatasertib (GDC-0068), entrectinib and pralsetinib were provided by Roche; everolimus, lapatinib and alpelisib were provided by Novartis;”
“Molecular profiling was performed using FoundationOne CDX and FoundationOne Liquid CDX tests on tissue and liquid biopsies.”
“All analyses were conducted using SAS software (v.9.4).”
Tests are named (CMH, chi-square, log-rank, Breslow-Day). Exact p-values are given for primary and key secondary endpoints (e.g., P=0.0294, P=0.0002). Effect sizes with 95% CIs are reported (HR=0.66, 95% CI 0.53-0.82). Software identified (SAS v9.4, R v4.3.3). Data presentation includes Kaplan-Meier curves and forest plots. Mathematical plausibility checks: ORR percentages are consistent with counts (e.g., 17.5% of 200 = 35 patients, matching 6 CR + 29 PR). However, some p-values in figures are reported as thresholds (e.g., P<0.0001), which is acceptable for very small values.
“The P value from the chi-square test ( P = 0.0294; Extended Data Table ) is nearly identical to the stratified CMH result”
“with a hazard ratio (HR) of 0.66 (95% CI: 0.53–0.82, P = 0.0002; Fig. )”
“All analyses were conducted using SAS software (v.9.4).”
“The primary endpoint analysis conducted using the stratified analysis by Cochran–Mantel–Haenszel (CMH) test demonstrated a significant difference in ORR between groups ( P = 0.0285; Extended Data Table ).”
“PFS was significantly improved in the TT arm (median 3.5 months, 95% CI: 3.0–4.8) compared to the SoC arm (2.8 months, 95% CI: 2.5–3.2), with a hazard ratio (HR) of 0.66 (95% CI: 0.53–0.82, P = 0.0002; Fig. ).”
“Kaplan–Meier survival analysis, accompanied by the log-rank test, was employed to compare OS and PFS across arms and subgroups.”
The data availability statement describes a managed-access process: requests reviewed by steering committee, data access agreement, secure platform, within 4-8 weeks, considered within 12 months. This is reported_and_adequate. Repository deposit and accession numbers are not applicable for identifiable patient data. Code sharing is not applicable as no custom code is mentioned (only standard software).
“Individual deidentified participant data generated during the current study are available upon reasonable request from academic or qualified clinical researchers affiliated with recognized institutions, strictly for the purpose of conducting non-commercial, ethically approvable research aligned with the original scope of the trial.”
“Data will be shared via a secure data-sharing platform within 4–8 weeks of approval, contingent upon data volume and complexity. Data requests will be considered within 12 months of manuscript publication.”
“Individual deidentified participant data generated during the current study are available upon reasonable request from academic or qualified clinical researchers affiliated with recognized institutions, strictly for the purpose of conducting non-commercial, ethically approvable research aligned with the original scope of the trial.”
“The trial registration, study protocol and methodological details are publicly accessible through ClinicalTrials.gov (accession identifier NCT04591431”
Trial registration is provided (NCT04591431, EudraCT 2018-002190-21). CONSORT checklist is mentioned in supplementary. All outcomes are reported, including negative OS result. Limitations are discussed (crossover, subgroup sizes, TMB variability). Conclusions are proportional, acknowledging exploratory nature of subgroups. Funding sources and COI are detailed.
“The trial is registered on ClinicalTrials.gov with identifier NCT04591431”
“Supplementary Table 1: Type of previous treatments received by patients in the ITT population; Supplementary Table 2: Crossover rate; Supplementary Table 3: Reasons for patients not undergoing crossover (SoC arm); Supplementary Table 4: Extension of tumor primary site listed as ‘Other’ in Table 1; Protocol V4; Protocol Appendix 2; and CONSORT checklist.”
“The substantial crossover rate (58.8%) from the SoC arm to the TT arm after disease progression, although ethically appropriate, likely contributed to this lack of OS benefit.”
“The trial is registered on ClinicalTrials.gov with identifier NCT04591431”
“Supplementary Table 1: Type of previous treatments received by patients in the ITT population; Supplementary Table 2: Crossover rate; Supplementary Table 3: Reasons for patients not undergoing crossover (SoC arm); Supplementary Table 4: Extension of tumor primary site listed as ‘Other’ in Table 1; Protocol V4; Protocol Appendix 2; and CONSORT checklist.”
“The substantial crossover rate (58.8%) from the SoC arm to the TT arm after disease progression, although ethically appropriate, likely contributed to this lack of OS benefit.”
Registered (3 IDs: ClinicalTrials.gov, EudraCT). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst mediumReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- Dead data/code linksRecomputed
Checked 39 references by DOI: 39 verified.
Every extracted reference resolved against Crossref/OpenAlex with no retraction flags.
7 data/code links checked; 6 live, 1 dead.
- datahttps://clinicaltrials.gov/study/NCT04591431?cond=rome%20trial&rank=1LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT04591431LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.clinicaltrialsregister.eu/ctr-search/search?query=2018-002190-21LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.ncbi.nlm.nih.gov/clinvar/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.oncokb.org/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://cancer.sanger.ac.uk/cosmicLIVEHTTP 200Resolved page looks like data.
- datahttps://www.esmo.org/guidelines/esmo-scale-for-clinical-actionability-of-molecular-targets-escatDEADHTTP 404Dead link — nothing to verify.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, typo.
- MINORconsistencyAbstract“crossover rate (52%)”→ Ensure consistency with the 58.8% crossover rate reported in the Discussion.Abstract states 52% crossover, while Discussion states 58.8%.
- MINORtypoCompeting interests“Bristol Myers Squibb snd Eli Lilly”→ Change 'snd' to 'and'.Typographical error.
- MINORconsistencyResults, Efficacy and survival outcomes“A high crossover rate (59%) from SoC to TT”→ Reconcile with the 58.8% mentioned in Discussion.Crossover rate reported as 59% here and 58.8% in Discussion.
- MINORconsistencyAbstract“crossover rate (52%)”→ Ensure consistency with the 58.8% reported in the Discussion.The abstract states 52% crossover, while the Discussion states 58.8%.
- MINORtypoCompeting interests“snd”→ Change to 'and'.Typo in 'snd'.
- MINORconsistencyResults, Efficacy and survival outcomes“P = 0.002”→ Verify if this is the correct p-value for PFS; the text reports HR=0.66 with P=0.0002.Figure 2a legend states P=0.002, but the text reports P=0.0002 for the same comparison.
The published work is robust and well-reported. An informed reader should weigh the minor internal inconsistencies (crossover rate, p-value discrepancy) and the overstated claim of 'strongest evidence' when interpreting the results. No erratum is strictly required, but correcting these inconsistencies would improve clarity.
- 1.HIGHreportingReconcile the crossover rate across the abstract (52%), results (59%), and discussion (58.8%) to a single consistent value.Inconsistent reporting of a key trial metric undermines reader trust and could be flagged as an internal contradiction.
- 2.HIGHreportingReconcile the PFS p-value discrepancy: text reports P=0.0002 while Figure 2a legend reports P=0.002 for the same comparison.A mismatch between text and figure for a primary endpoint p-value is a factual inconsistency that should be corrected.
- 3.HIGHreportingTemper the claim 'The ROME trial provides the strongest evidence supporting MTB implementation in clinical practice' to reflect that it is one of several randomized trials, not necessarily the strongest.The claim is overstated given the lack of a comparative analysis across all MTB studies; over-claiming is a common reviewer objection.
- 4.MEDIUMcopyeditFix the typo 'snd' to 'and' in the Competing interests section.Typographical errors in a published manuscript are unprofessional and easily corrected.
- 5.MEDIUMreportingAdd a rationale for the open-label design in the Methods (e.g., impracticality of blinding due to different treatment modalities).Providing a rationale for the lack of blinding strengthens the study design reporting.
- 6.MEDIUMstatisticsClarify how the assumptions of the CMH test and log-rank test were verified (e.g., proportional hazards) in the Statistical analysis section.Explicitly addressing assumption verification enhances the statistical rigor and reproducibility.
- 7.MEDIUMreportingProvide exact p-values in figures (e.g., Fig. 2) instead of thresholds like P<0.0001 where feasible.Exact p-values allow readers to assess the strength of evidence more precisely.
- 8.MEDIUMdata codeSpecify the name of the secure data-sharing platform in the Data availability statement.Naming the platform makes the data access route more concrete and actionable.
- 9.MEDIUMreportingConsider making the full statistical analysis plan (SAP) publicly available rather than only upon reasonable request.Public SAP increases transparency and reduces concerns about post hoc analyses.
- 10.MEDIUMreportingDiscuss the potential implications of the higher proportion of women in the TT arm on the results.Addressing baseline imbalances, even if not statistically significant, preempts reader concerns about confounding.
- 11.MEDIUMreportingClarify how the protocol amendment in 2024 (adding drugs) affected the predefined statistical analysis plan.Transparency about protocol changes is essential for interpreting the results as prespecified.
- 12.MEDIUMreportingReport the number of patients who received each specific targeted therapy in the TT arm.Granular treatment data improves interpretability and reproducibility.
- 13.MEDIUMstatisticsPerform a sensitivity analysis adjusting for crossover to assess the robustness of the OS result.Given the high crossover rate, a sensitivity analysis would strengthen the OS conclusion.
- 14.LOWreportingProvide exact p-values for all secondary endpoints in the text, not just thresholds like '<0.0001'.Exact p-values for secondary endpoints improve transparency and allow readers to assess evidence strength.
- 15.LOWreportingConsider moving the CONSORT flow diagram into the main text for easier access.A CONSORT diagram in the main text improves readability and adherence to reporting guidelines.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.