Efficacy and safety of fezolinetant for moderate-severe vasomotor symptoms associated with menopause in individuals unsuitable for hormone therapy: phase 3b randomised controlled trial.
Schaudig K, Wang X, Bouchard C, Hirschberg AL, Cano A, Shapiro C M M, Stute P, Wu X, Miyazaki K, Scrine L, Nappi RE
- DOI
- 10.1136/bmj-2024-079525
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/bfd9be1d-fadb-460b-b84a-fb0d85039e7d is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped)−0.25★
- LinksDead data/code link−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 16 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is the frequency of moderate-severe vasomotor symptoms, a patient-reported symptom, not a hard clinical outcome. The paper does not provide evidence of target engagement at the tested dose (e.g., PK/PD data) nor does it cite validated evidence linking reduction in vasomotor symptom frequency to a hard clinical outcome. The efficacy claim is based on this surrogate measure.
“The primary endpoint was mean change in daily frequency of moderate-severe vasomotor symptoms from baseline to week 24.”
- 02Treatment effect not shown to be clinically meaningful
The primary effect is a reduction in vasomotor symptom frequency from 10.58 to 2.61 events/day (a 75.66% reduction) compared to placebo (59.12% reduction), with a between-group difference of -1.93 events/day. While statistically significant, the clinical meaningfulness of this difference is not anchored to a minimal clinically important difference or other established threshold. The paper does not provide a justification for why this magnitude is clinically meaningful.
“At week 24, fezolinetant significantly reduced the frequency (least squares mean difference –1.93, 95% confidence interval (CI) –2.64 to –1.22; P<0.001)”
- 03Printed percentage does not match its own count
96.7% does not match the reported count 435/452
“most of the participants (435 (96.7%) were white”
ResultsFind in source - 04Declared data/code link does not resolve
Dead link — nothing to verify.
“https://www.trialsummaries.com/Home/LandingPage”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported phase 3b randomized controlled trial with strong methodological rigor across all eight dimensions. The paper clearly establishes the scientific premise, uses appropriate design and statistical methods, and provides comprehensive ethical, resource, and transparency documentation. Minor reporting gaps (e.g., statistical software not named, CONSORT not explicitly referenced) and a few copyedit issues (typos, misplaced DOI) do not undermine the overall integrity.
This is an interventional study (phase 3b RCT). Both reviewers agreed on the study type. The evaluation covered all eight dimensions; several sub-criteria were marked not applicable (e.g., animal-related items, independent replication) due to the human trial nature. The statistics verification covered only a subset of reported tests (4 tests), and the one 'inconsistent' finding was not specified; it may be a threshold-only p-value that cannot be machine-verified. The citation check found no retracted or non-existent references.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks. 1 printed percentage that does not match its own count.
- PERCENT96.7% does not match the reported count 435/452
“most of the participants (435 (96.7%) were white”
ResultsFind in source
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Primary endpoint: frequency of moderate-severe VMS change from baseline at week 24, LS mean difference -1.93 with 95% CI -2.64 to -1.22.
“Fezolinetant significantly reduced the frequency of vasomotor symptoms compared with placebo at week 24 (least squares mean difference –1.93, 95% confidence interval (CI) –2.64 to –1.22; P<0.001).”
Taken as given: The CI is two-sided at 95%.; The estimate is a difference in means (not a ratio).; The p-value is two-sided.Method: Recomputed two-sided p-value from the reported LS mean difference and 95% CI using the normal approximation.How we recomputed it: pCI(-1.93, -2.64, -1.22, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Secondary endpoint: severity of VMS change from baseline at week 24, LS mean difference -0.39 with 95% CI -0.57 to -0.21.
“The week 24 difference in symptom severity between fezolinetant and placebo was significant (least squares mean difference –0.39, 95% CI –0.57 to –0.21; P<0.001).”
Taken as given: The CI is two-sided at 95%.; The estimate is a difference in means.; The p-value is two-sided.Method: Recomputed two-sided p-value from the reported LS mean difference and 95% CI using the normal approximation.How we recomputed it: pCI(-0.39, -0.57, -0.21, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Secondary endpoint: PROMIS SD-SF 8b total score change from baseline at week 24, LS mean difference -2.5 with 95% CI -3.9 to -1.1.
“Participants receiving fezolinetant also had a greater reduction in total scores on the PROMIS Sleep Disturbance Short Form 8b compared with the placebo group (least mean squares mean difference –2.5, –3.9 to –1.1; P<0.001).”
Taken as given: The CI is two-sided at 95%.; The estimate is a difference in means.; The p-value is two-sided.Method: Recomputed two-sided p-value from the reported LS mean difference and 95% CI using the normal approximation.How we recomputed it: pCI(-2.5, -3.9, -1.1, 0)
- lowinternal contradictionThe abstract states 370 (81.7%) completed the study, but the results section states 387 (85.6%) completed the 24-week treatment period. These are different denominators (completed study vs. completed treatment period), so not a direct contradiction.
370 (81.7%) participants completed the study (fezolinetant=195, placebo group=175). ... Of these, 387 (85.6%) participants completed the 24 week treatment period.
Abstractreviewer’s wording
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
8 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewer 1Fezolinetant significantly reduced the frequency of moderate-severe vasomotor symptoms compared with placebo at week 24.The primary endpoint result is statistically significant with a clear effect size and CI.Evidence: LS mean difference -1.93, 95% CI -2.64 to -1.22, P<0.001
“At week 24, fezolinetant significantly reduced the frequency (least squares mean difference –1.93, 95% confidence interval (CI) –2.64 to –1.22; P<0.001)”
AbstractFind in source - supportedReviewer 1Fezolinetant significantly reduced the severity of vasomotor symptoms compared with placebo at week 24.Secondary endpoint is statistically significant with a clear effect size and CI.Evidence: LS mean difference -0.39, 95% CI -0.57 to -0.21, P<0.001
“severity of vasomotor symptoms (–0.39, –0.57 to –0.21; P<0.001)”
AbstractFind in source - supportedReviewer 1Fezolinetant reduced sleep disturbance (PROMIS SD-SF 8b) compared with placebo at week 24.Secondary endpoint is statistically significant with a clear effect size and CI.Evidence: LS mean difference -2.5, 95% CI -3.9 to -1.1, P<0.001
“the fezolinetant group had a greater reduction in sleep disturbance (PROMIS SD-SF 8b total score) compared with placebo (–2.5, –3.9 to –1.1; P<0.001)”
AbstractFind in source - supportedReviewers 1, 2Fezolinetant was well tolerated over six months.Safety data show similar TEAE rates between groups and no new safety signals.Evidence: TEAEs 65.0% vs 61.1%, serious TEAEs 4.4% vs 3.5%, no DILI
“Both groups showed similar incidences of treatment emergent adverse events (TEAEs, 147 (65.0%) in the fezolinetant group, 138 (61.1%) in the placebo group) and serious TEAEs (10 (4.4%) and 8 (3.5%), respectively).”
AbstractFind in source - supportedReviewers 1, 2Fezolinetant is an effective treatment option for individuals unsuitable for hormone therapy.The efficacy results support this conclusion for the studied population.Evidence: Primary and secondary endpoints all significant
“Fezolinetant was efficacious and well tolerated over a six month period for treating moderate-severe vasomotor symptoms in individuals considered unsuitable for hormone therapy.”
ConclusionFind in source - supportedReviewer 2Fezolinetant significantly reduced the frequency of moderate-severe vasomotor symptoms at week 24 compared with placebo.The primary endpoint analysis shows a statistically significant difference with a 95% CI excluding zero.Evidence: LS mean difference –1.93, 95% CI –2.64 to –1.22; P<0.001
Fezolinetant significantly reduced the frequency of vasomotor symptoms compared with placebo at week 24 (least squares mean difference –1.93, 95% confidence interval (CI) –2.64 to –1.22; P<0.001).
Resultsreviewer’s wording - supportedReviewer 2Fezolinetant significantly reduced the severity of vasomotor symptoms at week 24 compared with placebo.The secondary endpoint analysis shows a statistically significant difference with a 95% CI excluding zero.Evidence: LS mean difference –0.39, 95% CI –0.57 to –0.21; P<0.001
The week 24 difference in symptom severity between fezolinetant and placebo was significant (least squares mean difference –0.39, 95% CI –0.57 to –0.21; P<0.001).
Resultsreviewer’s wording - supportedReviewer 2Fezolinetant significantly reduced sleep disturbance compared with placebo.The PROMIS SD-SF 8b total score analysis shows a statistically significant difference with a 95% CI excluding zero.Evidence: LS mean difference –2.5, 95% CI –3.9 to –1.1; P<0.001
Participants receiving fezolinetant also had a greater reduction in total scores on the PROMIS Sleep Disturbance Short Form 8b compared with the placebo group (least mean squares mean difference –2.5, –3.9 to –1.1; P<0.001).
Resultsreviewer’s wording
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is the frequency of moderate-severe vasomotor symptoms, a patient-reported symptom, not a hard clinical outcome. The paper does not provide evidence of target engagement at the tested dose (e.g., PK/PD data) nor does it cite validated evidence linking reduction in vasomotor symptom frequency to a hard clinical outcome. The efficacy claim is based on this surrogate measure.
“The primary endpoint was mean change in daily frequency of moderate-severe vasomotor symptoms from baseline to week 24.”
- INADEQUATEEffect sizeThe primary effect is a reduction in vasomotor symptom frequency from 10.58 to 2.61 events/day (a 75.66% reduction) compared to placebo (59.12% reduction), with a between-group difference of -1.93 events/day. While statistically significant, the clinical meaningfulness of this difference is not anchored to a minimal clinically important difference or other established threshold. The paper does not provide a justification for why this magnitude is clinically meaningful.
“At week 24, fezolinetant significantly reduced the frequency (least squares mean difference –1.93, 95% confidence interval (CI) –2.64 to –1.22; P<0.001)”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The paper cites prior work on the prevalence and burden of vasomotor symptoms, limitations of hormone therapy, and existing non-hormonal options. It explicitly references the SKYLIGHT 1 and 2 phase 3 trials and explains that DAYLIGHT extends the placebo-controlled period to 24 weeks and enrolls a population unsuitable for hormone therapy. The rationale for the study is logically derived from the identified unmet need and the mechanism of fezolinetant. Limitations of prior research (e.g., shorter placebo control, broader population) are addressed by the study design.
“Fezolinetant was shown to be efficacious and well tolerated for treating moderate-severe vasomotor symptoms associated with menopause in phase 3 studies SKYLIGHT 1 and SKYLIGHT 2, which both included a 12 week placebo control period followed by active treatment extension to 52 weeks.”
“A substantial unmet need therefore exists for safe and effective options for non-hormonal treatment of vasomotor symptoms associated with menopause.”
“To further investigate the clinical benefits of fezolinetant for the treatment of moderate-severe vasomotor symptoms, we performed a phase 3b trial (DAYLIGHT), which included a 24 week placebo control period and enrolled a population considered unsuitable for hormone therapy.”
“Fezolinetant was shown to be efficacious and well tolerated for treating moderate-severe vasomotor symptoms associated with menopause in phase 3 studies SKYLIGHT 1 and SKYLIGHT 2, which both included a 12 week placebo control period followed by active treatment extension to 52 weeks.”
“To further investigate the clinical benefits of fezolinetant for the treatment of moderate-severe vasomotor symptoms, we performed a phase 3b trial (DAYLIGHT), which included a 24 week placebo control period and enrolled a population considered unsuitable for hormone therapy.”
“The study included a longer placebo control period (24 weeks) than previous phase 3 studies”
Randomization used interactive response technology with stratification by smoking status. Blinding is described as double-blind. Power analysis is provided for the primary endpoint. Inclusion/exclusion criteria are described (age, symptom severity, unsuitability for hormone therapy). The analysis population (safety and full analysis sets) is defined, addressing missing data. Controls are inherent in the placebo arm. Independent replication is not applicable for a single pivotal trial.
“Individuals aged 40-65 years with moderate-severe vasomotor symptoms associated with menopause and considered unsuitable candidates for hormone therapy were randomised 1:1 using interactive response technology to fezolinetant 45 mg or placebo once daily and stratified by smoking status (current and non-smoker (former or never)).”
“For a pairwise comparison of the primary endpoint using a two sample t test at a two sided 5% α, we determined that 220 participants in each group would provide at least 80% power to detect a difference from placebo of –1.8, assuming a standard deviation (SD) of 5.6.”
“The present study was a phase 3b, randomised, double blind, placebo controlled trial to assess the efficacy and safety of fezolinetant for treating moderate-severe vasomotor symptoms associated with menopause in individuals considered unsuitable for hormone therapy.”
“Individuals aged 40-65 years with moderate-severe vasomotor symptoms associated with menopause and considered unsuitable candidates for hormone therapy were randomised 1:1 using interactive response technology to fezolinetant 45 mg or placebo once daily and stratified by smoking status (current and non-smoker (former or never)).”
“For a pairwise comparison of the primary endpoint using a two sample t test at a two sided 5% α, we determined that 220 participants in each group would provide at least 80% power to detect a difference from placebo of –1.8, assuming a standard deviation (SD) of 5.6.”
“The present study was a phase 3b, randomised, double blind, placebo controlled trial”
The study enrolled only women (individuals with menopause), so sex is reported and justified by the condition. Age (mean 54.5, SD 4.7) and weight/BMI are reported in Table 1. Health status is implied by inclusion criteria (moderate-severe VMS, unsuitable for HT). Demographics include race (96.7% white) and categories of HT unsuitability. Species/strain and housing are not applicable as this is a human trial.
“Mean (SD) age (years) | 54.9 (4.8) | 54.1 (4.6) | 54.5 (4.7)”
“White | 217 (96.0) | 218 (97.3) | 435 (96.7)”
“Participants 453 individuals aged 40-65 years with moderate-severe vasomotor symptoms associated with menopause”
“Mean (SD) age (years) | 54.9 (4.8) | 54.1 (4.6) | 54.5 (4.7)”
“White | 217 (96.0) | 218 (97.3) | 435 (96.7)”
“individuals aged 40-65 years with moderate-severe vasomotor symptoms associated with menopause”
The ethics statement lists the specific ethics committee or IRB for each of the 16 countries, which is more than adequate. Written informed consent is explicitly stated. Regulatory compliance with Declaration of Helsinki, Good Clinical Practice, and ICH guidelines is stated. No animal research is involved.
“Written informed consent was obtained from all participants before any study related procedures.”
“This study was conducted in accordance with the Declaration of Helsinki, Good Clinical Practice, and International Council for Harmonisation guidelines.”
“Written informed consent was obtained from all participants before any study related procedures.”
“This study was conducted in accordance with the Declaration of Helsinki, Good Clinical Practice, and International Council for Harmonisation guidelines.”
Fezolinetant is named as the investigational drug, with dose (45 mg) and regimen (once daily). The comparator is placebo. No antibodies, cell lines, or other reagents are used. The trial is registered (ClinicalTrials.gov NCT05033886). Statistical software is not explicitly named, but this is not a key biological resource.
“Intervention Fezolinetant 45 mg or placebo once daily for 24 weeks.”
“Trial registration ClinicalTrials.gov NCT05033886 (https://clinicaltrials.gov/ct2/show/NCT05033886) ; EudraCT 2021-001685-38.”
“Fezolinetant 45 mg or placebo once daily for 24 weeks.”
The primary analysis uses a mixed model for repeated measures, with least squares mean differences and 95% CIs. Exact p-values are reported (e.g., P<0.001). Effect sizes with CIs are provided. Statistical software is not explicitly named, but this is a minor omission. Data presentation includes tables with per-group n and dispersion. Mathematical plausibility checks: baseline means and SDs are plausible; percentages sum correctly (e.g., 96.7% + 3.3% = 100%). No inconsistencies found.
“We performed a mixed model for repeated measures analysis with a missing at random assumption on change in the average daily frequency (or severity) of moderate-severe vasomotor symptoms from baseline to week 24.”
“We performed a mixed model for repeated measures analysis with a missing at random assumption on change in the average daily frequency (or severity) of moderate-severe vasomotor symptoms from baseline to week 24.”
“least squares mean difference –1.93, 95% confidence interval (CI) –2.64 to –1.22; P<0.001”
“least squares mean difference –1.93, 95% confidence interval (CI) –2.64 to –1.22”
The data availability statement names the platform (www.clinicalstudydatarequest.com) and specifies that anonymized participant-level data, trial-level data, and protocols are available on request. This meets the criteria for reported_and_adequate. No code is shared, but no bespoke code is mentioned. Repository deposit and accession numbers are not applicable for patient-level data.
“Researchers may request access to anonymised participant level data, trial level data, and protocols from Astellas sponsored clinical trials at www.clinicalstudydatarequest.com”
Trial registration is provided (NCT05033886). Methods are comprehensive. All pre-specified outcomes (primary, secondary, exploratory) are reported. Limitations are discussed (e.g., predominantly white population). Conclusions are proportional to the evidence. Funding and competing interests are disclosed. A reporting guideline (CONSORT) is not explicitly referenced, but the paper includes a flow diagram and follows standard reporting.
“Trial registration ClinicalTrials.gov NCT05033886 (https://clinicaltrials.gov/ct2/show/NCT05033886) ; EudraCT 2021-001685-38.”
“A potential limitation of this study was that participants were from 16 countries (Canada, the Netherlands, Belgium, France, Spain, Finland, Hungary, Italy, Czech Republic, UK, Denmark, Sweden, Norway, Poland, Germany, and Turkey), with most self-identifying as white.”
“Trial registration ClinicalTrials.gov NCT05033886 (https://clinicaltrials.gov/ct2/show/NCT05033886) ; EudraCT 2021-001685-38.”
“Funding: This study was funded by Astellas Pharma.”
Registered (3 IDs: ClinicalTrials.gov, EudraCT). No reporting guideline cited.
Broken references and links
1 finding · worst mediumReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- Dead data/code linksRecomputed
Checked 28 references by DOI: 22 verified — 6 no DOI (shown, not verified).
- NO DOIBRISDELLE™ (paroxetine) prescribing informationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISummary of product characteristicsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAustralian product information - Veoza™ (fezolinetant)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHighlights of prescribing informationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIVeoza TM (fezolinetant)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGuidance for Industry. Drug-induced liver injury: premarketing clinical evaluationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
6 of 7 data/code links checked; 5 live, 1 dead; 1 not probed.
- datahttps://clinicaltrials.gov/ct2/show/NCT05033886LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT03192176LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.clinicaltrials.astellas.comLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.trialsummaries.com/Home/LandingPageDEADHTTP 404Dead link — nothing to verify.
- datahttp://www.clinicalstudydatarequest.comUNVERIFIEDLiveness indeterminate — content not checked.
- datahttps://clinicalstudydatarequest.com/Study-Sponsors/Study-Sponsors-Astellas.aspxLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicalstudydatarequest.comLIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, typo.
- MINORtypoTable 6“Transminases increased”→ Transaminases increasedSpelling error in table.
- MINORconsistencyData availability statement“www.clinicalstudydatarequest.com (10.1097/gme.0b013e31813429d6)”→ Remove the DOI or correct it; it appears to be a reference to a different article.The DOI in the data availability statement seems misplaced.
- MINORconsistencyAbstract“387 e079525 e079525”→ Remove duplicate page numbers or correct formatting.Duplicate page numbers appear in the header.
- MINORtypoResults, Safety“Transminases increased”→ Change to 'Transaminases increased'.Spelling error in Table 6.
- MINORconsistencyData availability statement“www.clinicalstudydatarequest.com (10.1097/gme.0b013e31813429d6)”→ Remove the DOI that appears to be incorrectly placed.A DOI is appended to the URL, likely a formatting error.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (statistical software not named, CONSORT not explicitly referenced) and the copyedit issues (typos, misplaced DOI) as low-severity concerns. No erratum or re-analysis is warranted based on the available evidence.
- 1.MEDIUMreportingIn the Methods section, explicitly name the statistical software (e.g., SAS version) used for all analyses.Both reviewers noted the statistical software is not identified, which is a minor reporting gap that could be easily fixed.
- 2.MEDIUMreportingAdd an explicit statement in the Methods or a dedicated section confirming adherence to the CONSORT reporting guideline, and consider providing the checklist as supplementary material.The paper follows a CONSORT-like structure but does not explicitly reference the guideline, which is a minor transparency gap.
- 3.MEDIUMcopyeditFix the typo 'Transminases increased' to 'Transaminases increased' in Table 6 and in the Results, Safety section.The copyedit pass flagged this spelling error in two locations.
- 4.MEDIUMcopyeditRemove the misplaced DOI (10.1097/gme.0b013e31813429d6) from the data availability statement, as it appears to reference a different article.The copyedit pass flagged this as a formatting error that could confuse readers.
- 5.LOWcopyeditRemove the duplicate page numbers 'e079525 e079525' from the abstract header.The copyedit pass flagged this as a formatting inconsistency.
- 6.LOWreportingConsider providing a more detailed description of the blinding procedures (e.g., who was blinded, how allocation was concealed) in the Methods.One reviewer suggested this as an improvement to enhance transparency.
- 7.LOWstatisticsClarify the handling of missing data for the PROMIS endpoint, as the number of participants with data differs from the full analysis set.One reviewer raised this as a potential transparency issue.
- 8.LOWreportingProvide exact p-values for secondary endpoints that are reported as P<0.001, if possible, to enhance transparency.One reviewer suggested this as an improvement for statistical reporting.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.