Benralizumab versus placebo for hypereosinophilic syndrome: a randomized, placebo-controlled phase 3 trial.
Ogbogu PU, Roufosse F, Akuthota P, Kuna P, Groh M, Reiter A, Yokota A, Siddiqui SH, Mutsaers PGNJ, Li B, Khoury P, Bahadori LM, Bednarczyk A, Bouma G, Brooks LG, Ferreira J, Grindebacke H, Ho CN, Jain P, Palmer RL, Jison ML, Klion AD, NATRON study group
- DOI
- 10.1038/s41591-026-04315-8
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/53ad7c8c-0543-4158-a417-3ccca3f97977 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- StatisticsStatistic did not reproduce−0.5★
- ClaimsOverstated claim−0.5★
- ReportingKey resources partially met−0.25★
- ReportingStatistical analysis partially met−0.25★
- LinksDead data/code link−0.25★
- References were found (29) but none could be checked — every lookup failed or lacked a DOI.
- 01Printed percentage does not match its own count
92.5% does not match the reported count 62/66
“92.5% (62/66)”
92.5% (62/66) of placebo-treated patien…Find in source - 02Conclusion reaches beyond the evidence
The safety of benralizumab was consistent with its known profile, including in adolescents.
“Safety and tolerability results were consistent with the known safety profile for benralizumab, including in adolescents.”
Discussion ¶2Find in source - 03Declared data/code link does not resolve
Dead link — nothing to verify.
“https://vivli.org/members/enquiries-about-studies-not-listed-on-the-vivli-platform/”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a methodologically strong, well-designed, prospectively registered phase 3 RCT with thorough ethics reporting, transparent data-availability routes, and a headline result (HR 0.35, p=0.0024) that survives the checks performed. Its weaknesses are reporting-quality issues rather than validity threats: missing SAS version, unverified model assumptions, two minor internal inconsistencies (placebo completion percentage and Table 1 race denominators), a slightly overbroad adolescent safety claim, and several truncated cross-references.
Two independent evaluator runs of the same model scored all eight dimensions; the only pass/warn split (statistical analysis) was resolved via checklist-level scoring and corroborating verification components. Citation audit: 29 references checked, 0 retracted, 0 not-found. Statistics audit: 12 tests recomputed, 11 consistent, 1 minor inconsistency (completion percentage), 0 decision errors; tests without full statistics (threshold-only p-values, resampling/exact p-values) were not machine-verified. Reproducibility audit: 6/7 links live, 1 dead. Claim audit: 1 overstated claim (adolescent safety). Copyedit: 11 minor issues.
Numerical inconsistencies
2 findings · worst highValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 11 tests: 11 consistent, 0 inconsistent; 2 recomputed directly from the reported test statistics, 9 via agent-written checks. 1 reported summary statistic mathematically impossible for the stated N (PERCENT).
- PERCENT92.5% does not match the reported count 62/66
“92.5% (62/66)”
92.5% (62/66) of placebo-treated patien…Find in source
- CONSISTENTreported p = .002 · recomputed p = .002Recomputed hazard ratio 0.35 (95% CI 0.18–0.69), reported p=0.0024
“hazard ratio 0.35, 95% CI 0.18 to 0.69, P = 0.0024”
Taken as given: 0.18–0.69 is a two-sided 95% confidence interval for the hazard ratio of 0.35, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.0024 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.35, 0.18, 0.69, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Recomputed RR 0.34 (95% CI 0.18–0.63), reported p=0.0008
“RR 0.34, 95% CI 0.18 to 0.63, P = 0.0008”
Taken as given: 0.18–0.63 is a two-sided 95% confidence interval for the RR of 0.34, not a range, an IQR, or a different interval level; the RR is a RATIO measure, so the interval is symmetric on the log scale; p=0.0008 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.34, 0.18, 0.63, 1) - CONSISTENTreported p = .003 · recomputed p = .004Reviewer 1Key secondary: proportion of patients with flare or withdrawal (OR 0.31, 95% CI 0.14-0.69, P=0.0033).
“OR 0.31, 95% CI 0.14 to 0.69, P = 0.0033”
Taken as given: The CI is two-sided at 95%; The OR is on the log scaleMethod: pCI function from sandbox, log=1How we recomputed it: pCI(0.31, 0.14, 0.69, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Key secondary: time to first hematologic relapse (HR 0.08, 95% CI 0.03-0.20, P<0.0001).
“HR 0.08, 95% CI 0.03 to 0.20, P < 0.0001”
Taken as given: The CI is two-sided at 95%; The HR is on the log scaleMethod: pCI function from sandbox, log=1; p < 0.0001 confirmedHow we recomputed it: pCI(0.08, 0.03, 0.20, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Additional secondary: proportion of patients with AEC <500 cells/µL (OR 87.9, 95% CI 26.1-296.0, P<0.0001).
“OR 87.9, 95% CI 26.1 to 296.0, P < 0.0001”
Taken as given: The 2x2 table is (61,6) for benralizumab and (8,58) for placebo; Fisher's exact test is appropriateMethod: Fisher's exact test (two-tailed) via pFisher2x2How we recomputed it: pFisher2x2(61, 6, 8, 58, 0) - CONSISTENTreported p = .003 · recomputed p = .004Reviewer 2Key secondary: flare-or-withdraw OR 0.31 (95% CI 0.14–0.69), reported P = 0.0033
“22.4% (15/67) versus 45.5% (30/66), respectively (odds ratio (OR) 0.31, 95% CI 0.14 to 0.69, P = 0.0033)”
Taken as given: 95% CI is two-sided; CI bounds 0.14 and 0.69 are on the OR (ratio) scale; p derived from a Wald test on log OR using the CI width; p is two-tailedMethod: Two-tailed normal p from log OR; SE = (ln(0.69) − ln(0.14)) / 3.92.How we recomputed it: pCI(0.31, 0.14, 0.69, 1) - CONSISTENTreported p = .002 · recomputed p = .002Reviewer 2Key secondary: PROMIS Fatigue LS mean difference −4.72 (95% CI −7.64 to −1.80), reported P = 0.0017
“LS means difference –4.72 (–7.64 to –1.80) | 0.0017”
Taken as given: CI bounds −7.64 and −1.80 are the 95% CI of the LS mean difference; p derived from a Wald test on the difference using the CI width; p is two-tailedMethod: Two-tailed normal p from the difference; SE = ((−1.80) − (−7.64)) / 3.92 = 1.49.How we recomputed it: pCI(-4.72, -7.64, -1.80, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Additional secondary: AEC <500 OR 87.87 (95% CI 26.09–295.97), reported P < 0.0001
“91.0% (61/67) compared to 12.1% (8/66) in the placebo group (OR 87.87, 95% CI 26.09 to 295.97, nominal P < 0.0001)”
Taken as given: 95% CI is two-sided; CI bounds 26.09 and 295.97 are on the OR (ratio) scale; p derived from a Wald test on log OR using the CI width; reported p is a floor ('<0.0001'); recomputed p expected to be below 0.0001Method: Two-tailed normal p from log OR; SE = (ln(295.97) − ln(26.09)) / 3.92.How we recomputed it: pCI(87.87, 26.09, 295.97, 1) - CONSISTENTreported p = .003 · recomputed p = .006Reviewer 2Key secondary: flare-or-withdraw 15/67 vs 30/66, reported P = 0.0033 (CMH per Table 2)
“Proportion of patients experiencing a flare or withdrawing by week 24, n (%) b | 15 (22.4%) | 30 (45.5%) | OR 0.31 (0.14 to 0.69) | 0.0033”
Taken as given: Events: 15 in benralizumab (n=67), 30 in placebo (n=66); non-events 52 and 36; Unstratified Fisher exact test is an approximation to the region-stratified CMH test actually used; Two-tailed testMethod: Two-tailed Fisher's exact test on the 2×2 cell counts.How we recomputed it: pFisher2x2(15, 52, 30, 36, 0) - CONSISTENTreported p = .005 · recomputed p = .007Reviewer 2Additional secondary: OCS increase 17/67 vs 32/66, reported P = 0.005
“Proportion of patients requiring an increase in systemic corticosteroids during the double-blind period, n (%) b,f | 17 (25.4%) | 32 (48.5%) | OR 0.35 (0.16 to 0.73) | 0.005”
Taken as given: 17/67 benralizumab and 32/66 placebo patients required an OCS increase (non-events 50 and 34); Unstratified Fisher exact test approximates the CMH test; Two-tailed testMethod: Two-tailed Fisher's exact test on the 2×2 cell counts.How we recomputed it: pFisher2x2(17, 50, 32, 34, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Additional secondary: SF-36v2 PCS LS mean difference 4.54 (95% CI 1.9–7.1), reported P = 0.0008
“Change from baseline in SF-36v2 PCS at week 24, LS mean (95% CI) g | 6.7 (4.9 to 8.5) | 2.2 (0.3 to 4.0) | LS means difference 4.54 (1.9 to 7.1) | 0.0008”
Taken as given: CI bounds 1.9 and 7.1 are the 95% CI of the LS mean difference; p derived from a Wald test on the difference using the CI width; Two-tailed testMethod: Two-tailed normal p from the difference; SE = (7.1 − 1.9) / 3.92.How we recomputed it: pCI(4.54, 1.9, 7.1, 0)
- lowinternal contradictionRace percentages in Table 1 are based on a subset of patients (58 out of 67 for benralizumab, 57 out of 66 for placebo), but the table header reports n=67 and n=66, and the missing data are not indicated. This is a minor reporting inconsistency.
White: 42 (72.4%) [benralizumab, n=67]
Table 1reviewer’s wording - lowinternal contradictionThe placebo completion percentage 92.5% does not match its own numerator/denominator: 62/66 = 93.94% ≈ 93.9%. One of the two numbers is a typo, but the discrepancy is small and does not affect any result or conclusion.
“In total, 97% (65/67) of benralizumab-treated patients and 92.5% (62/66) of placebo-treated patients completed the 24-week double-blind period”
ResultsFind in source
Overstated conclusions
2 findings · worst mediumConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
11 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated).
- overstatedReviewer 2The safety of benralizumab was consistent with its known profile, including in adolescents.The adult/overall 24-week safety data support the general statement, but 'including in adolescents' rests on only 4 enrolled adolescents (3 benralizumab) — too thin a base for a specific adolescent safety claim.Evidence: AE rates 64.2% vs 66.7%, SAEs 7.5% vs 7.6%, one death (sepsis, investigator-assessed as unrelated); adolescent outcomes are not separately reported.
“Safety and tolerability results were consistent with the known safety profile for benralizumab, including in adolescents.”
Discussion ¶2Find in source - partialReviewer 1Benralizumab is effective across HES subtypes and organ systems.The paper shows consistent treatment effect across subgroups in Supplementary Figure 3, but the sample size is small for some subgroups (e.g., lymphocytic HES), limiting conclusions.Evidence: Results, primary endpoint, Supplementary Figure 3
Across predefined subgroups, the treatment effect of benralizumab versus placebo was consistent and favored benralizumab (Supplementary Fig. 3).
Results ¶2reviewer’s wording - supportedReviewers 1, 2Benralizumab significantly reduces the risk of first HES flare compared to placebo.The primary endpoint analysis provides strong evidence (HR 0.35, 95% CI 0.18-0.69, P=0.0024).Evidence: Results, primary endpoint, Table 2, Figure 2
“Benralizumab significantly reduced the risk of first flare versus placebo (hazard ratio 0.35, 95% CI 0.18 to 0.69, P = 0.0024).”
AbstractFind in source - supportedReviewer 1Benralizumab reduces the annualized rate of HES flares.The annualized flare rate analysis shows a significant 66% reduction (RR 0.34, 95% CI 0.18-0.63, P=0.0008).Evidence: Results, key secondary endpoints, Table 2, Figure 3b
Significantly fewer HES flares occurred in those treated with benralizumab versus placebo (0.41 versus 1.23 flares per year, respectively), with a 66% reduction in the annualized rate of flares (RR 0.34, 95% CI 0.18 to 0.63, P = 0.0008).
Results ¶3reviewer’s wording - supportedReviewer 1Benralizumab delays time to first hematologic relapse.The analysis shows a 92% reduction in risk (HR 0.08, 95% CI 0.03-0.20, P<0.0001), with clear separation from week 4.Evidence: Results, key secondary endpoints, Table 2, Figure 3c
There was a significant delay in the time to first hematologic relapse: benralizumab patients were 92% less likely to relapse at any given point... (HR 0.08, 95% CI 0.03 to 0.20, P < 0.0001).
Results ¶4reviewer’s wording - supportedReviewer 1Benralizumab improves fatigue as measured by PROMIS Fatigue.The LS mean difference at week 24 is -4.72 (95% CI -7.64 to -1.80, P=0.0017), which is statistically significant and clinically meaningful.Evidence: Results, key secondary endpoints, Table 2, Figure 3d
“Fatigue was significantly improved in the benralizumab group versus the placebo group. The least squares (LS) mean difference in Patient-Reported Outcomes Measurement Information System (PROMIS) Fatigue scores between groups was –4.72 (95% CI –7.64 to –1.80, P = 0.0017) at week 24.”
Results ¶5Find in source - supportedReviewer 1Benralizumab has a safety profile consistent with its known profile.AE rates are similar between groups (64.2% vs 66.7%), and no new safety signals are identified. The one death is not attributed to treatment.Evidence: Results, safety section, Table 3
“Safety was consistent with the known safety profile of benralizumab.”
Results ¶14Find in source - supportedReviewer 2Benralizumab reduced the annualized rate of flares versus placebo.The negative binomial model yields RR 0.34 with 95% CI 0.18–0.63, P = 0.0008, and model-estimated rates (0.41 vs 1.23 per year) are reported; the claim is directly backed.Evidence: Annualized flare rate: RR 0.34, 95% CI 0.18 to 0.63, P = 0.0008; 0.41 vs 1.23 flares per year.
“with a 66% reduction in the annualized rate of flares (RR 0.34, 95% CI 0.18 to 0.63, P = 0.0008)”
ResultsFind in source - supportedReviewer 2Benralizumab significantly improved fatigue.The PROMIS Fatigue difference is a pre-specified key secondary endpoint that passed the hierarchical gate; LS mean difference −4.72 (95% CI −7.64 to −1.80, P = 0.0017) supports the claim.Evidence: PROMIS Fatigue LS mean difference −4.72 (95% CI −7.64 to −1.80, P = 0.0017) at week 24.
“Fatigue was significantly improved in the benralizumab group versus the placebo group.”
ResultsFind in source - supportedReviewer 2Benralizumab demonstrated efficacy and safety in the treatment of HES.The claim is supported for the studied 24-week double-blind period: primary and all key secondary endpoints met the hierarchical testing strategy, and safety was descriptively comparable to placebo; the authors appropriately defer long-term outcomes to the OLE.Evidence: Primary and key secondary endpoints all statistically significant; AE/SAE rates similar between groups.
“Benralizumab demonstrated efficacy and safety in the treatment of HES, with evidence of clinical benefit apparent at the earliest visits in the trial”
DiscussionFind in source - supportedReviewer 2These results may expand the potential therapeutic options for patients with this disease.A modest, well-calibrated inference from the positive primary and secondary endpoint results; it is framed as a possibility, not an overclaim.Evidence: Positive primary and key secondary endpoints in a placebo-controlled phase 3 trial.
“These results may therefore expand the potential therapeutic options for patients with this disease.”
DiscussionFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- N/ASurrogate endpointPrimary endpoint is time to first HES flare, defined as clinical manifestations or lab abnormalities leading to treatment escalation or hospitalization. This is a clinical composite outcome, not a surrogate.
“primary endpoint was time to first HES flare. A flare was defined as HES clinical manifestation or laboratory abnormality resulting in an increase of OCS ≥10 mg day−1 prednisone equivalent for ≥2 days, or an increase or addition of a new cytotoxic and/or immunosuppressive therapy, or hospitalization.”
- ADEQUATEEffect sizeThe primary endpoint shows a 65% reduction in risk of first HES flare (HR 0.35, 95% CI 0.18-0.69, P=0.0024). Absolute event rates: 19.4% vs 42.4%. This is a clinically meaningful reduction.
“Benralizumab significantly reduced the risk of first flare versus placebo (hazard ratio 0.35, 95% CI 0.18 to 0.69, P = 0.0024).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
2 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Statistical reporting gaps (tests, assumptions, effect sizes)Assessed
- Key resources under-identified (antibodies, cell lines, RRIDs)Assessed
The introduction discusses the limitations of current treatments for HES and the potential of benralizumab based on its mechanism and prior phase 2 data. It cites key studies on mepolizumab and benralizumab, and clearly links the premise to the study objectives.
Randomization is central and computer-generated with permuted block design, stratified by region and flare status. Double-blind design is described with explicit unblinding exceptions. A power analysis is provided. Inclusion and exclusion criteria are detailed. The primary analysis uses the full analysis set (ITT).
“All patients were centrally assigned to a randomized study treatment using Interactive Web Response Systems (IWRS)/Interactive Voice Response Systems (IVRS).”
“The randomization sequence was computer-generated centrally using a permuted block design with a fixed block size of 4 and stratified by geographic region (North America, Europe, Asia and rest of the world) and HES flare status at screening.”
“It was estimated that approximately 38 first HES flare events during the double-blind period were required to detect a statistically significant difference between treatment groups at the two-sided 5% significance level with approximately 80% power if the true treatment effect is an HR of 0.389”
Table 1 and the Results report sex (61.7% female), age (median 51.0 years, range 14–87, including 4 adolescents), race, region, HES subtype, organ involvement, and background therapy. Health status is characterized in depth via AEC, prior flares, OCS dose, and organ systems. Weight is not reported, which is not material for this patient-level trial; both sexes enrolled, so sex_justified is not applicable.
“Overall, 62% (82/133) of patients were female, 38% (51/133) were male and the median (range) age was 51.0 (14.0–87.0) years.”
“White | 42 (72.4%) | 45 (78.9%) | 87 (75.7%)”
“The median (range) time since diagnosis was 1.9 (0.1–32.3) years, and patients had experienced a median (range) of 2 (0–12) HES flares in the 12 months before enrollment.”
The paper states that the protocol was reviewed and approved by Independent Ethics Committees/Institutional Review Boards (listed in supplementary). All patients provided written informed consent. Compliance with Declaration of Helsinki and ICH-GCP is stated.
“All patients provided written informed consent.”
“The protocol, protocol amendments and any other relevant documents were reviewed and approved by the Independent Ethics Committees/Institutional Review Boards listed in the .”
“All patients provided written informed consent.”
“The trial was conducted in accordance with the ethical principles of the Declaration of Helsinki and is consistent with the International Council for Harmonisation Good Clinical Practice guidelines”
Benralizumab is identified as a 30 mg dose via accessorized prefilled syringe, with manufacturer (AstraZeneca) and regimen. This is adequate. However, the statistical software (SAS System) is mentioned without a version number, which is reported but inadequate.
“All data analyses were performed with SAS System (SAS Institute Inc.) software.”
“benralizumab 30 mg (accessorized prefilled syringe) subcutaneously every 4 weeks or a matching placebo”
“All data analyses were performed with SAS System (SAS Institute Inc.) software.”
“PK was assessed through benralizumab serum concentrations and immunogenicity assessed through ADA assays and neutralizing antibody testing, as previously described”
All statistical tests are named (log-rank, Cox, logistic, negative binomial, MMRM). Exact p-values and CIs are reported. However, assumptions such as proportional hazards for Cox are not discussed. The SAS version is not given. The race percentages in Table 1 are based on a subset of patients, but the denominator is not indicated, which is a minor plausibility issue.
“Time-to-event endpoints were analyzed using a stratified log-rank test, adjusted for region. HRs and 95% CIs were estimated using a Cox proportional hazards model with treatment group and region as covariates.”
“Time to first HES flare, n events (%) | 13 (19.4%) | 28 (42.4%) | HR 0.35 (0.18 to 0.69) a | 0.0024”
“Time-to-event endpoints were analyzed using a stratified log-rank test, adjusted for region. HRs and 95% CIs were estimated using a Cox proportional hazards model with treatment group and region as covariates.”
“In total, 97% (65/67) of benralizumab-treated patients and 92.5% (62/66) of placebo-treated patients completed the 24-week double-blind period”
The data availability statement provides two concrete routes: through AstraZeneca's data sharing policy (with URL) and through Vivli (with URL). This meets the standard for managed access. No sequence data warrants accession numbers, and no custom code was written, so those are not applicable.
“Data underlying the findings described in this Article can be requested in accordance with AstraZeneca’s data sharing policy available via AstraZeneca at https://astrazenecagrouptrials.pharmacm.com/ST/Submission/Disclosure . Data for studies directly listed on Vivli are available via Vivli at https://www.vivli.org .”
“Data for studies directly listed on Vivli are available via Vivli at https://www.vivli.org”
“The full protocol is publicly available at https://www.astrazenecaclinicaltrials.com/study/D3254C00001/”
All applicable sub-criteria are adequately addressed: methods are comprehensive, trial registration is provided, a reporting summary is linked, all outcomes (including non-significant) are reported, limitations are discussed, conclusions are proportional, and funding and conflicts of interest are disclosed.
“ClinicalTrials.gov identifier: NCT04191304”
“The study is sponsored and funded by AstraZeneca (Södertälje, Sweden).”
“This trial is registered with ClinicalTrials.gov ( NCT04191304 (https://clinicaltrials.gov/ct2/show/NCT04191304) )”
“First, the sample size was small owing to the rarity of the disease.”
“The study is sponsored and funded by AstraZeneca (Södertälje, Sweden).”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
1 finding · worst mediumReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- Dead data/code linksRecomputed
Checked 29 references by DOI: 0 verified — 29 no DOI (shown, not verified).
- NO DOIProposed refined diagnostic criteria and classification of eosinophil disorders and related syndromesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIApproach to the patient with suspected hypereosinophilic syndromeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEpidemiology, clinical picture and long-term outcomes of FIP1L1-PDGFRA -positive myeloid neoplasm with eosinophilia: data from 151 patientsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBiologics and hypereosinophilic syndromes: knowledge gaps and controversiesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHypereosinophilic syndrome in Europe: retrospective study of treatment patterns, clinical manifestations, and healthcare resource utilizationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMepolizumab therapy improves the most bothersome symptoms in patients with hypereosinophilic syndromeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMepolizumab incompletely suppresses clinical flares in a pilot study of episodic angioedema with eosinophiliaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAnti-IL-5 (mepolizumab) therapy induces bone marrow eosinophil maturational arrest and decreases eosinophil progenitors in the bronchial mucosa of atopic asthmaticsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEfficacy and safety of mepolizumab in hypereosinophilic syndrome: a phase III, randomized, placebo-controlled trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITreatment of patients with the hypereosinophilic syndrome with mepolizumabNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn international, retrospective study of off-label biologic use in the treatment of hypereosinophilic syndromesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMEDI-563, a humanized anti-IL-5 receptor alpha mAb with enhanced antibody-dependent cell-mediated cytotoxicity functionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffects of benralizumab on airway eosinophils in asthmatic patients with sputum eosinophiliaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBenralizumab for PDGFRA-negative hypereosinophilic syndromeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEosinophil depletion with benralizumab for eosinophilic esophagitisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBenralizumab depletes IL-5Rα-bearing cells in skin lesions of patients with atopic dermatitisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBenralizumab for eosinophilic gastritis: a single-site, randomised, double-blind, placebo-controlled, phase 2 trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEfficacy and safety of benralizumab for patients with severe asthma uncontrolled with high-dosage inhaled corticosteroids and long-acting β(2)-agonists (SIROCCO): a randomised, multicentre, placebo-controlled phase 3 trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBenralizumab, an anti-interleukin-5 receptor α monoclonal antibody, as add-on treatment for patients with severe, uncontrolled, eosinophilic asthma (CALIMA): a randomised, double-blind, placebo-controlled phase 3 trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOral glucocorticoid-sparing effect of benralizumab in severe asthmaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOral corticosteroid elimination via a personalised reduction algorithm in adults with severe, eosinophilic asthma treated with benralizumab (PONENTE): a multicentre, open-label, single-arm studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBenralizumab versus mepolizumab for eosinophilic granulomatosis with polyangiitisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILong-term efficacy and safety of benralizumab treatment for PDGFRA-negative hypereosinophilic syndromeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISymptom assessment in hypereosinophilic syndrome: toward development of a patient-reported outcomes toolNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBurden of hypereosinophilic syndromes in the United States: patients’ perspectiveNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISingle-center off-label benralizumab use for refractory hypereosinophilic syndrome demonstrates satisfactory safety and efficacyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHypereosinophilia and hypereosinophilic syndromes: first findings from a nationwide multicenter cohortNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITwo-year efficacy and safety of anti-interleukin-5/receptor therapy for eosinophilic granulomatosis with polyangiitisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISelection of a ligand-binding neutralizing antibody assay for benralizumab: comparison with an antibody-dependent cell-mediated cytotoxicity (ADCC) cell-based assayNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
7 data/code links checked; 6 live, 1 dead.
- datahttps://astrazenecagrouptrials.pharmacm.com/ST/Submission/DisclosureLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.vivli.orgLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://vivli.org/members/enquiries-about-studies-not-listed-on-the-vivli-platform/DEADHTTP 404Dead link — nothing to verify.
- datahttps://vivli.org/ourmember/astrazeneca/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT04191304LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttp://clinicaltrials.gov/ct2/show/NCT04191304LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.astrazenecaclinicaltrials.com/study/D3254C00001/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
11 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 11 minor suggestions below.
11 copyedit issues flagged: mostly clarity, typo, consistency.
- MINORtypoData Availability Statement“is also avialable via Vivli”→ is also available via VivliTypo in 'avialable'.
- MINORtypoAuthor list“Universtié de Versailles Saint-Quentin-en-Yvelines”→ Université de Versailles Saint-Quentin-en-YvelinesTypo in 'Universtié'.
- MINORconsistencyTable 1, Race“Race percentages are based on a subset of patients without explicit denominator”→ Add footnote indicating the number of patients with race data (e.g., 'n=58 for benralizumab group')The percentages for race do not sum to 100% of the group N, implying missing data; the denominator should be stated.
- MINORtypoAffiliation 5 (author list)“Universtié de Versailles-Saint-Quentin-en-Yvelines”→ Université de Versailles-Saint-Quentin-en-YvelinesMisspelling of 'Université'.
- MINORtypoData availability“The AstraZeneca Vivli member page outlining further details is also avialable via Vivli”→ is also available via VivliMisspelling of 'available'.
- MINORconsistencyResults, Patient disposition“92.5% (62/66) of placebo-treated patients completed the 24-week double-blind period”→ 93.9% (62/66) — 62/66 = 93.94%, which rounds to 93.9%The printed percentage is inconsistent with its own numerator/denominator.
- MINORclarityResults, Safety“please refer to the for further details.”→ please refer to the Supplementary Information for further details.Truncated cross-reference: the target of 'the' is missing.
- MINORclarityMethods, Statistical analysis“Please refer to the for further details on the sensitivity analysis.”→ Please refer to the Supplementary Methods for further details on the sensitivity analysis.Truncated cross-reference.
- MINORclarityMethods, Outcomes“please refer to the for further details.”→ please refer to the Supplementary Methods for further details.Truncated cross-reference.
- MINORclarityReporting summary“Further information on research design is available in the linked to this article.”→ Further information on research design is available in the linked Reporting Summary to this article.Missing noun after 'the linked'.
- MINORclarityMethods, Study design“The end-of-study definition is described in the .”→ The end-of-study definition is described in the Supplementary Methods.Truncated cross-reference.
The published findings are robust overall: the design, registration, and transparency are strong, and the primary result is supported by the recomputed statistics (11/12 consistent). An informed reader should weigh the minor internal inconsistencies (92.5% vs 93.9% completion; Table 1 race denominators), the unverified model assumptions and missing SAS version, one dead external link, and the overstated adolescent safety claim; none of these currently undermines the primary conclusion, but a corrigendum correcting the percentage and repairing the truncated cross-references would be appropriate.
- 1.HIGHstatisticsCorrect the placebo completion percentage in Results/Patient disposition: 62/66 = 93.9%, not 92.5%, and reconcile the text with the Figure 1 flow counts.A verifiable internal inconsistency in the Results is the single inconsistency found in the statistics audit and will be caught by an informed reader.
- 2.HIGHrigorTemper the Discussion claim that benralizumab safety was consistent with its known profile 'including in adolescents': only 4 adolescents (3 benralizumab) were enrolled, so either report their safety outcomes explicitly or qualify the claim.The claim audit rated this statement as overstated because an n of 4 cannot support an adolescent-specific safety generalization.
- 3.HIGHcopyeditRepair the truncated cross-references ('please refer to the for further details', 'described in the .', 'linked to this article') in Methods (Outcomes, Statistical analysis, Study design), Results (Safety), and the Reporting summary by naming the target (Supplementary Methods/Information, Reporting Summary).Dead cross-references make methods, ethics listings, and reporting checklists unauditable; they were flagged by both the copyedit pass and Reviewer 2.
- 4.HIGHdata codeIdentify and fix the single dead external link among the 7 checked (protocol/registration/data pages) and re-verify the remaining links resolve.A broken link undercuts the data-availability and protocol-access statements that reviewers will follow.
- 5.MEDIUMstatisticsAdd an explicit statement in Methods/Statistical analysis on checking model assumptions (e.g., proportional hazards for Cox models, residual checks for MMRM), or cite where these were assessed and satisfied.Reviewer 1 flagged assumptions as unverified; Reviewer 2 accepted them without evidence, so making the assessment explicit removes ambiguity.
- 6.MEDIUMstatisticsAdd the SAS version (e.g., 'SAS 9.4') to the 'All data analyses were performed with SAS System (SAS Institute Inc.) software' sentence in Methods/Statistical analysis.Software without a version lowers both the key-resources and statistical-analysis ratings and impedes reproducibility.
- 7.MEDIUMreportingSpecify the PK, ADA, and neutralizing-antibody assay platforms, vendors, and validation/performance data in Methods/Outcomes instead of 'as previously described'.Immunogenicity and PK methods need to be self-contained for the trial to be auditable.
- 8.MEDIUMreportingAdd a Table 1 footnote stating the denominators for race percentages (e.g., n=58 for benralizumab, n=57 for placebo) and noting missing race data.Percentages that do not sum to the group N are an internal-consistency issue flagged by both reviewers and the integrity check.
- 9.MEDIUMethicsCite the IEC/IRB names and approval identifiers in the main text (or explicitly reference supplementary listing entries) and repair the truncated 'listed in the .' sentence in Methods/Study design.Ethics approval must be auditable without relying on a broken cross-reference to the supplement.
- 10.MEDIUMreportingName the reporting guideline (e.g., CONSORT 2010) with item-level compliance in addition to the linked Nature Portfolio reporting summary.Explicitly naming the guideline makes the reporting-checklist mechanism clear to readers and reviewers.
- 11.MEDIUMdata codeAdd a dataset identifier (e.g., a Vivli study-page DOI) to the data-availability statement and note that no custom code was used (SAS analyses), so code sharing is not applicable.A specific identifier strengthens the managed-access route and preempts ambiguity about code availability.
- 12.LOWcopyeditFix the typos 'avialable' → 'available' in the Data Availability statement and 'Universtié' → 'Université' in the author affiliations.These are minor copyedit issues flagged by the copyedit pass.
- 13.LOWreportingReport weight or BMI as a baseline characteristic in Table 1, or note its omission in the Limitations.Weight was the only biological-variable gap noted by Reviewer 1 and would complete the biological-variables reporting.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.