Efficacy and safety of intravenous induction and subcutaneous maintenance therapy with guselkumab for patients with Crohn's disease (GALAXI-2 and GALAXI-3): 48-week results from two phase 3, randomised, placebo and active comparator-controlled, double-blind, triple-dummy trials.
Panaccione R, Feagan BG, Afzali A, Rubin DT, Reinisch W, Panés J, Danese S, Hisamatsu T, Terry NA, Salese L, Van Rampelbergh R, Sahoo A, Vetter ML, Yee J, Han C, Frustaci ME, Wan KYY, Yang Z, Johanns J, Andrews JM, D'Haens GR, Sands BE, GALAXI 2 & 3 Study Group
- DOI
- 10.1016/S0140-6736(25)00681-6
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/00ca4ce0-b143-4515-b460-e7d9bd48c471 is authoritative.
How this rating was calculated
- CitationsUnresolved reference ×3−0.75★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on composite endpoints that include clinical response/remission (CDAI) and endoscopic response/remission (SES-CD). Endoscopic response is a surrogate marker for long-term disease control. Although the paper cites STRIDE-II consensus linking endoscopic healing to clinical outcomes, it does not provide validated evidence directly linking the specific surrogate (SES-CD endoscopic response) to hard clinical outcomes in this context. Target engagement at the tested dose is not explicitly demonstrated in this paper (PK/PD data are mentioned but not linked to efficacy).
“The coprimary endpoints assessing the long-term efficacy of guselkumab compared with placebo were (1) clinical response at week 12 and clinical remission at week 48 and (2) clinical response at week 12 and endoscopic response at week 48.”
- 02Treatment effect not shown to be clinically meaningful
The primary reported effects are absolute differences in composite endpoint rates compared with placebo (e.g., 43% for clinical response/remission, 33% for endoscopic response). These are statistically significant but not anchored to a minimal clinically important difference or to a clear biological/clinical meaningfulness threshold. The paper does not provide an explicit anchor for what constitutes a clinically meaningful effect size for these composite endpoints.
“adjusted treatment difference 43% [95% CI 32–54] in the guselkumab 200 mg group and 38% [27–49] in the guselkumab 100 mg group; p<0·0001”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a rigorously designed and transparently reported pair of phase 3 randomized controlled trials. The methods are comprehensive, ethics and data availability are well documented, and the statistical reporting is appropriate. Minor reporting gaps (explicit CONSORT statement, copyedit typos) and three references not found in registries are the only concerns.
Both reviewers agreed on study type (interventional) and on all dimension statuses. The only divergence was on the reporting guideline sub-criterion (reported but inadequate vs not reported), which does not affect the overall pass. The statistics verification covered only 2 tests; many p-values are threshold-only and not machine-verifiable. The citation check flagged 3 references not found in registries, which are integrity concerns but do not affect dimension statuses.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Recompute p-value for clinical response at week 12 and clinical remission at week 48 in GALAXI-2 guselkumab 200 mg vs placebo.
“80 (55%) of 146 participants in the guselkumab 200 mg group ... and nine (12%) of 76 in the placebo group (adjusted treatment difference 43% [95% CI 32–54] ... p<0·0001)”
Taken as given: The numbers 80 and 66 are the event and non-event counts in the guselkumab 200 mg group (146 total).; The numbers 9 and 67 are the event and non-event counts in the placebo group (76 total).; The p-value is from a chi-squared test without continuity correction.Method: Pearson's chi-squared test on the 2x2 table (80,66,9,67).How we recomputed it: pChi2x2(80, 66, 9, 67) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Recompute p-value for clinical response at week 12 and clinical remission at week 48 in GALAXI-3 guselkumab 200 mg vs placebo.
“72 (48%) of 150 participants in the guselkumab 200 mg group ... and nine (13%) of 72 in the placebo group (35% [24–46] ... p<0·0001)”
Taken as given: The numbers 72 and 78 are the event and non-event counts in the guselkumab 200 mg group (150 total).; The numbers 9 and 63 are the event and non-event counts in the placebo group (72 total).; The p-value is from a chi-squared test without continuity correction.Method: Pearson's chi-squared test on the 2x2 table (72,78,9,63).How we recomputed it: pChi2x2(72, 78, 9, 63)
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
4 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Both guselkumab regimens were superior to placebo for the coprimary endpoints.The presented results show statistically significant differences with p<0.0001 and confidence intervals excluding zero.Evidence: Results section reports adjusted treatment differences and p-values for both coprimary endpoints in both trials.
Both guselkumab regimens were superior to placebo for clinical response at week 12 and clinical remission at week 48 in GALAXI-2 (adjusted treatment difference 43% [95% CI 32–54] ... p<0·0001)
Results ¶4reviewer’s wording - supportedReviewer 1Guselkumab was superior to ustekinumab for endoscopic outcomes at week 48.The paper reports prespecified, multiplicity-controlled analyses showing superiority for endoscopic endpoints.Evidence: Discussion states 'Both guselkumab dose regimens were also superior to ustekinumab at week 48 in prespecified, multiplicity-controlled, treat-through analyses of the pooled GALAXI-2 and GALAXI-3 dataset for all endoscopic endpoints'.
“Both guselkumab dose regimens were also superior to ustekinumab at week 48 in prespecified, multiplicity-controlled, treat-through analyses of the pooled GALAXI-2 and GALAXI-3 dataset for all endoscopic endpoints”
Discussion ¶1Find in source - supportedReviewers 1, 2Guselkumab was well tolerated with a safety profile consistent with its approved indications.Safety data show no deaths and similar adverse event rates across groups, supporting the claim.Evidence: Results report serious adverse events and no deaths; Discussion states safety profile consistent with approved uses.
“No deaths were reported.”
ResultsFind in source - supportedReviewer 2Guselkumab showed superiority to ustekinumab at week 48 for endoscopic outcomes.The paper reports prespecified pooled analyses showing superiority for endoscopic endpoints, with multiplicity control.Evidence: Discussion states 'both guselkumab dose regimens were also superior to ustekinumab at week 48 in prespecified, multiplicity-controlled, treat-through analyses of the pooled GALAXI-2 and GALAXI-3 dataset for all endoscopic endpoints'.
“Both guselkumab dose regimens were also superior to ustekinumab at week 48 in prespecified, multiplicity-controlled, treat-through analyses of the pooled GALAXI-2 and GALAXI-3 dataset for all endoscopic endpoints”
Discussion ¶1Find in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on composite endpoints that include clinical response/remission (CDAI) and endoscopic response/remission (SES-CD). Endoscopic response is a surrogate marker for long-term disease control. Although the paper cites STRIDE-II consensus linking endoscopic healing to clinical outcomes, it does not provide validated evidence directly linking the specific surrogate (SES-CD endoscopic response) to hard clinical outcomes in this context. Target engagement at the tested dose is not explicitly demonstrated in this paper (PK/PD data are mentioned but not linked to efficacy).
“The coprimary endpoints assessing the long-term efficacy of guselkumab compared with placebo were (1) clinical response at week 12 and clinical remission at week 48 and (2) clinical response at week 12 and endoscopic response at week 48.”
- INADEQUATEEffect sizeThe primary reported effects are absolute differences in composite endpoint rates compared with placebo (e.g., 43% for clinical response/remission, 33% for endoscopic response). These are statistically significant but not anchored to a minimal clinically important difference or to a clear biological/clinical meaningfulness threshold. The paper does not provide an explicit anchor for what constitutes a clinically meaningful effect size for these composite endpoints.
“adjusted treatment difference 43% [95% CI 32–54] in the guselkumab 200 mg group and 38% [27–49] in the guselkumab 100 mg group; p<0·0001”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction and 'Research in context' section cite prior trials and identify gaps (e.g., lack of head-to-head comparisons, randomized withdrawal designs). The rationale for guselkumab and the study objectives follow logically. Limitations of prior research are explicitly discussed, such as the difficulty applying randomized withdrawal results to practice.
“To date, most trials assessing the efficacy of maintenance therapy with biological agents in patients with Crohn’s disease did not include statistically robust, head-to-head comparisons with established therapies.”
“New treatment options that confer increased efficacy with an acceptable safety profile are needed.”
“The use of a treat-through study design, in which clinical outcomes at a specific timepoint were not a prerequisite for maintenance therapy, more closely mimicked clinical practice than previous studies.”
“To date, most trials assessing the efficacy of maintenance therapy with biological agents in patients with Crohn’s disease did not include statistically robust, head-to-head comparisons with established therapies.”
“New treatment options that confer increased efficacy with an acceptable safety profile are needed.”
“The use of a treat-through study design, in which clinical outcomes at a specific timepoint were not a prerequisite for maintenance therapy, more closely mimicked clinical practice than previous studies.”
Randomization used a centralized computer-generated schedule with permuted blocks and stratification variables. Blinding was comprehensive (participants, investigators, site personnel, funder, and central readers). A priori power analysis is provided. Inclusion/exclusion criteria are detailed. The analysis population (primary analysis population) and missing data handling are pre-specified. Outlier handling is addressed through the pre-specified analysis population and missing data imputation rules.
“A centralised computer-generated schedule prepared before the start of the study under the supervision of the sponsor was used to randomly assign participants to study treatment.”
“Participants, investigators, site personnel, and the sponsor were masked to study treatment until all participants had either completed the follow-up visit at week 48 or terminated study participation before week 48.”
“A centralised computer-generated schedule prepared before the start of the study under the supervision of the sponsor was used to randomly assign participants to study treatment.”
“Participants, investigators, site personnel, and the sponsor were masked to study treatment until all participants had either completed the follow-up visit at week 48 or terminated study participation before week 48.”
Sex, age, race, ethnicity, and disease characteristics (CDAI, SES-CD, CRP, etc.) are reported in Table 2. Both sexes are enrolled, so sex justification is not applicable. Age and health status are reported. Demographics are comprehensive.
“Male 41 (54%) 87 (60%) 69 (48%) 84 (59%) 47 (65%) 91 (61%) 85 (59%) 84 (57%)”
“Race Asian 17 (22%) 28 (19%) 34 (24%) 32 (22%) 18 (25%) 28 (19%) 38 (27%) 22 (15%)”
“Sex Male 41 (54%) 87 (60%) 69 (48%) 84 (59%) 47 (65%) 91 (61%) 85 (59%) 84 (57%)”
“Race Asian 17 (22%) 28 (19%) 34 (24%) 32 (22%) 18 (25%) 28 (19%) 38 (27%) 22 (15%)”
The protocol was approved by an institutional review board or ethics committee at each site, and the study was conducted in compliance with the Declaration of Helsinki and Good Clinical Practice. Written informed consent was obtained from all participants. Regulatory compliance is explicitly stated.
“The study protocol was approved by an institutional review board or ethics committee at each study site and both studies were conducted in compliance with the Declaration of Helsinki, Good Clinical Practice guidelines, and applicable local regulations.”
“Before any study procedure, participants provided written informed consent.”
“The study protocol was approved by an institutional review board or ethics committee at each study site”
“Before any study procedure, participants provided written informed consent.”
“both studies were conducted in compliance with the Declaration of Helsinki, Good Clinical Practice guidelines, and applicable local regulations.”
The drugs are named with manufacturer (Johnson & Johnson) and dosing regimens are fully described. No bench reagents or cell lines are used, so those criteria are not applicable. Statistical software (SAS) is identified.
“200 mg intravenous guselkumab at weeks 0, 4, and 8, then 200 mg subcutaneous guselkumab every 4 weeks from week 12 to week 44 (guselkumab 200 mg group)”
“All statistical analyses were performed with SAS Studio (version 3.8 on SAS 9.4 M6).”
“200 mg intravenous guselkumab at weeks 0, 4, and 8, then 200 mg subcutaneous guselkumab every 4 weeks from week 12 to week 44”
“All statistical analyses were performed with SAS Studio (version 3.8 on SAS 9.4 M6).”
Tests are named (e.g., Chi-squared test, Mantel-Haenszel common risk difference). Assumptions are handled by design (pre-specified analysis model). Exact p-values are reported (e.g., p<0·0001). Effect sizes with 95% CIs are reported. Software is identified. Data presentation includes per-group n and CIs. Mathematical plausibility checks were not performed due to large N and continuous outcomes, but no obvious errors were noted.
“Adjusted treatment differences and associated 95% CIs and p values were based on the common risk difference by use of Mantel–Haenszel stratum weights and the Sato variance estimator”
“adjusted treatment difference 43% [95% CI 32–54]”
“Adjusted treatment differences and associated 95% CIs and p values were based on the common risk difference by use of Mantel–Haenszel stratum weights and the Sato variance estimator”
“adjusted treatment difference 43% [95% CI 32–54]”
The data sharing policy is described, and requests can be submitted through the Yale Open Data Access Project, which is a managed-access platform. This is adequate for patient-level data. No raw data repository deposit or accession numbers are applicable due to privacy. No custom code is mentioned, so code sharing is not applicable.
The trials are registered (NCT03466411). Methods are detailed. Limitations are discussed. Conclusions are proportional to the evidence. Funding and conflicts of interest are disclosed. A reporting guideline is not explicitly mentioned, but the paper follows CONSORT-like reporting with a trial profile.
“The GALAXI-2 and GALAXI-3 trials are registered with ClinicalTrials.gov (NCT03466411).”
“As with all studies, the GALAXI studies had limitations. Adults with moderately to severely active Crohn’s disease were eligible, so the results might not be generalisable to children, adolescents, or patients with mild forms of the disease.”
“Funding Johnson & Johnson.”
“The GALAXI-2 and GALAXI-3 trials are registered with ClinicalTrials.gov (NCT03466411).”
“As with all studies, the GALAXI studies had limitations.”
“Funding Johnson & Johnson.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 24 references by DOI: 18 verified — 3 DOI unresolved, 3 no DOI (shown, not verified).
- UNRESOLVED10.2147/itt.s247129All are equal, some are more equal: targeting IL 12 and 23 in IBD–a clinical perspectiveCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1016/s2468-1253(23)00323-5Efficacy and safety of 48 weeks of guselkumab for patients with Crohn’s disease: maintenance results from the phase 2, randomised, double-blind GALAXI-1 trialCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1016/s0140-6736(24)01534-4IL-23 inhibition for chronic inflammatory diseaseCited DOI does not resolve to any Crossref record.
- NO DOITREMFYA (guselkumab) prescribing informationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntroductory guide: MedDRA version 26.0No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntroductory guide for standardised MedDRA queries (SMQs) version 26.0No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
2 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 2 minor suggestions below.
2 copyedit issues flagged: mostly typo, consistency.
- MINORtypoAbstract, Methods“treat- through trials”→ treat-through trialsHyphenation inconsistency.
- MINORconsistencyResults, paragraph 1“1048 participants were enrolled, randomly assigned, treated, and followed up until week 48”→ 1048 participants were enrolled and randomly assigned, treated, and followed up until week 48Possible redundancy.
The published work is robust and well-reported; an informed reader should weigh the three references not found in registries (potential fabrication) and the minor reporting gaps (no explicit CONSORT statement, copyedit typos). No erratum is warranted for the main results, but the reference issues should be investigated.
- 1.HIGHotherVerify or correct the reference 'All are equal, some are more equal: targeting IL 12 and 23 in IBD–a clinical perspective' (DOI 10.2147/itt.s247129) which was not found in any registry.A reference that cannot be located in Crossref/OpenAlex may be fabricated or have an incorrect DOI, which is an integrity concern.
- 2.HIGHotherVerify or correct the reference 'Efficacy and safety of 48 weeks of guselkumab for patients with Crohn’s disease: maintenance results from the phase 2, randomised, double-blind GALAXI-1 trial' (DOI 10.1016/s2468-1253(23)00323-5) which was not found in any registry.A reference that cannot be located in Crossref/OpenAlex may be fabricated or have an incorrect DOI, which is an integrity concern.
- 3.HIGHotherVerify or correct the reference 'IL-23 inhibition for chronic inflammatory disease' (DOI 10.1016/s0140-6736(24)01534-4) which was not found in any registry.A reference that cannot be located in Crossref/OpenAlex may be fabricated or have an incorrect DOI, which is an integrity concern.
- 4.MEDIUMreportingAdd an explicit statement in the Methods (or a checklist) that the trials follow the CONSORT reporting guideline.Both reviewers noted the absence of an explicit reporting guideline statement; adding it would fully satisfy reporting transparency.
- 5.MEDIUMcopyeditFix the hyphenation inconsistency in the Abstract, Methods: change 'treat- through trials' to 'treat-through trials'.Consistent terminology improves readability and professionalism.
- 6.MEDIUMcopyeditRevise the sentence in Results, paragraph 1 to remove redundancy: '1048 participants were enrolled and randomly assigned, treated, and followed up until week 48'.The original phrasing is redundant and can be streamlined for clarity.
- 7.LOWdata codeAdd a brief statement in the Data sharing section clarifying that statistical analysis code is not publicly available (if that is the case).Clarifying code availability, even if not applicable, preempts reader queries and enhances reproducibility transparency.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.