A Randomized Trial of Shunting for Idiopathic Normal-Pressure Hydrocephalus
Luciano MG, Williams MA, Hamilton MG, Katzen HL, Dasher NA, Moghekar A, Hua J, Malm J, Eklund A, Alpert Abel N, Raslan AM, Elder BD, Savage JJ, Barrow DL, Shahlaie K, Jensen H, Zwimpfer TJ, Wollett J, Hanley DF, Holubkov R, PENS Trial Investigators and the Adult Hydrocephalus Clinical Research Network.
- DOI
- 10.1056/nejmoa2503109
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/1e07899e-17b1-4181-8da7-adbad27155a6 is authoritative.
How this rating was calculated
- ReportingData & code availability not met−0.5★
- ReportingEthical approvals partially met−0.25★
- ReportingKey resources partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run on this paper: the pass that reads its reported means did not complete. No reported mean was checked for arithmetic impossibility.
- 01Data and code not shared
No data availability statement is present in the paper. Although the trial is registered, the paper does not state where the data can be accessed, which is a significant transparency gap.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a methodologically strong, well-powered, double-blind placebo-controlled RCT of shunt surgery for iNPH, with a clear scientific rationale, robust design, adequate biological-variable reporting, consistent recomputation of the 7 machine-checkable statistics, and transparent reporting of outcomes, limitations, funding, and COI. The main weaknesses are the complete absence of a data availability statement (a fail) and under-specification of the randomization method, analysis software, and an explicit regulatory-compliance framework (warn-level gaps).
Both reviewers were synthesized per dimension; they agreed on 7 of 8 statuses and diverged only on ethical approvals (warn vs pass), which I resolved to warn while preserving the pass-dissenter's evidence. Statistics verification covered only the 7 tests reported with test statistic + df or effect + CI — everything else (including 'NS' p-values) is unverified, not confirmed. No retracted or unresolvable references (44 checked); both reproducibility links are live; the trial is registered (ClinicalTrials.gov NCT05081128).
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 7 tests: 7 consistent, 0 inconsistent; 7 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary gait velocity treatment difference: p from reported estimate and 95% CI
“treatment difference=0.21 m/s (95% confidence interval 0.12 to 0.31; P<0.001)”
Taken as given: The CI is a two-sided 95% confidence interval for the treatment difference.; The estimate and CI are on the linear (additive) scale.; The p-value is derived from a normal approximation using SE = (high-low)/(2*1.96).Method: Two-sided p computed from the reported estimate and 95% CI under a normal approximation.How we recomputed it: pCI(0.21, 0.12, 0.31, 0) - CONSISTENTreported p = .003 · recomputed p = .003Reviewer 1Tinetti secondary outcome: p from reported treatment difference and 95% CI
“Tinetti | 18.8 ± 6.1 | 20.3 ± 5.1 | 0.5 ± 5.3 | 2.9 ± 3.8 | 3.1 (1.0, 5.1) | 0.003”
Taken as given: The CI is a two-sided 95% confidence interval for the treatment difference.; The effect is on the additive scale.; The p-value is from the linear regression model, approximated as normal.Method: Two-sided p computed from the reported estimate and 95% CI under a normal approximation.How we recomputed it: pCI(3.1, 1.0, 5.1, 0) - CONSISTENTreported p = .030 · recomputed p = .028Reviewers 1, 2Falls comparison: Fisher's exact test with mid-p correction
“More participants reported falls for Placebo (23/50; 46%) than Open Shunt (12/49; 24.5%) (P=0.03).”
Taken as given: The 23 and 50 are fall events and total participants in the Placebo arm.; The 12 and 49 are fall events and total participants in the Open Shunt arm.; The analysis used Fisher's exact test with mid-p correction, as stated in the Table 4 footnote.Method: Two-sided Fisher's exact test with mid-p correction, computed from the event counts and group totals.How we recomputed it: pFisher2x2(23, 50-23, 12, 49-12, 1) - CONSISTENTreported p = .002 · recomputed p = .002Reviewer 1Positional headaches comparison: Fisher's exact test with mid-p correction
“More participants reported positional headaches suggesting low CSF pressure for Open Shunt (29/49; 59.2%) than Placebo (14/50; 28.0%) (P=0.002).”
Taken as given: The 14 and 50 are headache events and total participants in the Placebo arm.; The 29 and 49 are headache events and total participants in the Open Shunt arm.; The analysis used Fisher's exact test with mid-p correction.Method: Two-sided Fisher's exact test with mid-p correction, computed from the event counts and group totals.How we recomputed it: pFisher2x2(14, 50-14, 29, 49-29, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Primary outcome: gait velocity treatment difference from linear regression
“treatment difference=0.21 m/s (95% confidence interval 0.12 to 0.31; P<0.001)”
Taken as given: The 95% CI is two-sided and approximately normal; The estimate is the mean differenceMethod: p from 95% CI using normal approximationHow we recomputed it: pCI(0.21, 0.12, 0.31, 0) - CONSISTENTreported p = .003 · recomputed p = .003Reviewer 2Secondary outcome: Tinetti treatment difference from linear regression
“Tinetti (2.9 vs 0.5; P=0.003)”
Taken as given: The 95% CI is two-sided and approximately normal; The estimate is the mean differenceMethod: p from 95% CI using normal approximationHow we recomputed it: pCI(3.1, 1.0, 5.1, 0) - CONSISTENTreported p = .040 · recomputed p = .036Reviewer 2Subdural hematoma/hemorrhage comparison: Placebo vs Open Shunt
“There were more subdural hematomas (SDH) for Open Shunt (n=6; 12%) than Placebo (n=1; 2.0%) (P=0.04)”
Taken as given: The table is 2x2: Placebo events=1, non-events=49; Open events=6, non-events=43; Mid-p correction is usedMethod: Fisher's exact test with mid-p correctionHow we recomputed it: pFisher2x2(1, 49, 6, 43, 1)
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
8 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewer 1The combined effect of improved gait velocity and a lower rate of falls in the Open Shunt group suggests that shunt surgery may have a broad beneficial clinical impact on the health of elderly patients with iNPH.The paper shows improved gait velocity and fewer falls, but 'broad beneficial clinical impact' is an extrapolation beyond the 3-month outcomes measured, despite the cautious 'suggests' language.Evidence: Gait velocity improvement and falls 24.5% vs 46.0% in Open Shunt versus Placebo; however, Open Shunt had more subdural hematomas and positional headaches.
“The combined effect of improved gait velocity and a lower rate of falls in the Open Shunt group suggests that shunt surgery may have a broad beneficial clinical impact on the health of elderly patients with iNPH.”
DiscussionFind in source - supportedReviewer 1Shunting for iNPH resulted in significant improvements in gait velocity and a measure of gait and balance, but not measures of cognition or incontinence within 3 months.The primary gait velocity endpoint and the Tinetti secondary endpoint are significant, while MoCA and OABQsf are non-significant, exactly as stated.Evidence: Primary: treatment difference 0.21 m/s (95% CI 0.12 to 0.31; P<0.001). Tinetti: difference 3.1 (1.0 to 5.1; P=0.003). MoCA and OABQsf were not significant.
“For patients responsive to temporary CSF drainage, shunting for iNPH resulted in significant improvements in gait velocity and a measure of gait and balance, but not measures of cognition or incontinence within 3 months.”
ConclusionFind in source - supportedReviewer 1The PENS Trial provides evidence that shunt surgery is effective for improving gait velocity at 3 months in patients with iNPH selected for shunt surgery in accordance with the International iNPH Guidelines.The primary endpoint result directly supports this claim for the selected patient population.Evidence: Primary outcome gait velocity change was 0.23 vs 0.03 m/s with treatment difference 0.21 m/s (P<0.001).
“The PENS Trial provides evidence that shunt surgery is effective for improving gait velocity at 3 months in patients with idiopathic normal pressure hydrocephalus selected for shunt surgery in accordance with the International iNPH Guidelines.”
ConclusionFind in source - supportedReviewers 1, 2Open shunting appeared to reduce lateral ventricular volume, consistent with a functioning shunt.The imaging analysis shows a significant decrease in lateral ventricular volume for Open Shunt versus Placebo, consistent with the claim.Evidence: Lateral ventricle volume difference -21.0 mL (95% CI -33.9 to -8.1).
“Open shunting appeared to reduce lateral ventricular volume, consistent with a functioning shunt.”
DiscussionFind in source - supportedReviewer 2Shunt surgery is effective for improving gait velocity at 3 months in patients with iNPH.The primary outcome shows a statistically and clinically significant improvement in gait velocity for the Open Shunt group compared to Placebo, with a treatment difference of 0.21 m/s (95% CI 0.12 to 0.31, P<0.001).Evidence: Table 3, primary outcome: gait velocity change 0.23 vs 0.03 m/s, difference 0.21 m/s (95% CI 0.12 to 0.31, P<0.001).
“Gait velocity increased for Open Shunt (0.23 ± 0.23 m/s; n=49) and was unchanged for Placebo (0.03 ± 0.23 m/s; n=49); treatment difference=0.21 m/s (95% confidence interval 0.12 to 0.31; P<0.001).”
AbstractFind in source - supportedReviewer 2Shunting for iNPH resulted in significant improvements in gait and balance, but not cognition or incontinence within 3 months.The secondary outcome Tinetti (gait and balance) showed a significant improvement (P=0.003), while MoCA (cognition) and OABQsf (incontinence) did not reach significance after multiplicity adjustment.Evidence: Table 3: Tinetti treatment difference 3.1 (95% CI 1.0 to 5.1, P=0.003); MoCA 1.2 (95% CI 0.1 to 2.2, NS); OABQsf -1.9 (95% CI -4.0 to 0.1, NS).
A significant treatment difference favoring Open Shunt vs Placebo was seen for the Tinetti (2.9 vs 0.5, P = 0.003), but not the MoCA (1.3 vs 0.3) or OABQsf (−3.3 vs −1.5).
Resultsreviewer’s wording - supportedReviewer 2More participants reported falls for Placebo than Open Shunt.The safety data show a statistically significant lower rate of falls in the Open Shunt group (24.5%) compared to Placebo (46.0%, P=0.03).Evidence: Table 4: Participants reporting falls: Placebo 23/50 (46.0%), Open Shunt 12/49 (24.5%), P=0.03.
“More participants reported falls for Placebo (23/50; 46%) than Open Shunt (12/49; 24.5%) (P=0.03).”
ResultsFind in source - supportedReviewer 2There were more subdural hematomas for Open Shunt than Placebo.The safety data show a higher rate of subdural hematoma/hemorrhage in the Open Shunt group (12.2%) compared to Placebo (2.0%, P=0.04).Evidence: Table 4: Participants with subdural hematoma/hemorrhage: Open Shunt 6/49 (12.2%), Placebo 1/50 (2.0%), P=0.04.
“There were more subdural hematomas (SDH) for Open Shunt (n=6; 12%) than Placebo (n=1; 2.0%) (P=0.04).”
ResultsFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointPrimary endpoint is gait velocity, a validated functional clinical measure (hard clinical outcome), not a surrogate biomarker.
“The primary outcome was gait velocity change 3 months after surgery.”
- ADEQUATEEffect sizeEffect size (0.23 m/s) exceeds established substantial meaningful change (0.10 m/s) and large MCID (0.22 m/s), and is anchored to clinical meaningfulness.
“The mean gait velocity change for Open Shunt (0.23 m/s) is over twice the substantial meaningful change in the elderly and exceeds the “large” minimum clinically important difference of 0.22 m/s in Parkinsonism.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
3 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data and code not sharedAssessed
- Ethics/consent reporting incompleteAssessed
- Key resources under-identified (antibodies, cell lines, RRIDs)Assessed
The introduction cites the 2005 International iNPH Guidelines, the Japanese Guidelines, previous small trials, and a 2024 Cochrane review noting the need for larger studies. It acknowledges skepticism about shunt effectiveness and a call for a moratorium. The study objectives directly follow from these gaps.
“a 2024 Cochrane review noted, “there is a need for similar studies to increase the certainty of the findings presented””
“The effectiveness of shunting is questioned due to the absence of adequately powered, randomized trials.”
“This is an international, multi-center, prospective, double-blind, randomized, placebo-controlled trial.”
“the clinical effectiveness of shunt surgery for iNPH has been subject to skepticism to the point of a call for a moratorium on shunt surgery because of variability of study results, questionable durability of benefit, the risks of surgery, and the potential for a strong placebo effect.”
“The effectiveness of shunting is questioned due to the absence of adequately powered, randomized trials.”
The trial is described as double-blind, randomized, placebo-controlled. Randomization occurred in a 1:1 ratio but the specific method (e.g., computer-generated) is not stated. Blinding of participants and personnel except shunt adjusters is described. A priori power analysis (>90% power for 0.2 m/s difference) is provided. Inclusion/exclusion criteria are detailed. Missing data are handled via multiple imputation, and per-protocol and as-treated analyses are performed. Control is via placebo setting. Independent replication is not applicable for a single pivotal trial.
“Randomization occurred immediately before surgery in a 1:1 ratio.”
“Treatment assignments were concealed from study participants and personnel, except for investigators responsible for shunt adjustments, and were disclosed only for patient safety reasons.”
“A sample size of 100 participants was specified, providing >90% power to detect a between group gait velocity difference of 0.2 m/s (SD=0.29 m/s), accounting for interim analyses and attrition.”
“Randomization occurred immediately before surgery in a 1:1 ratio.”
“Treatment assignments were concealed from study participants and personnel, except for investigators responsible for shunt adjustments, and were disclosed only for patient safety reasons.”
“A sample size of 100 participants was specified, providing >90% power to detect a between group gait velocity difference of 0.2 m/s (SD=0.29 m/s), accounting for interim analyses and attrition.”
Table 1 presents sex (48.5% female, 51.5% male), age (mean 75.0 ± 5.7 years), race, ethnicity, education, and comorbidities (Table 2). The study includes both sexes, so a single-sex justification is not needed. Age and health status are adequately reported. Demographics are comprehensive.
“Participants were 51.5% male and 48.5% female, with mean age 75.0 ± 5.7 years, and 63% with a Bachelor’s degree or higher.”
“Any comorbidities | 46 (92.0%) | 33 (67.3%) | 79 (79.8%)”
“Participants were 51.5% male and 48.5% female”
“mean age 75.0 ± 5.7 years”
IRB approval is adequately reported with a protocol identifier. Consent is reported as obtained. However, the paper does not state adherence to the Declaration of Helsinki, ICH-GCP, or the Common Rule, so the regulatory_compliance sub-criterion is not met.
“For U.S. sites, the central Institutional Review Board (IRB) was at Johns Hopkins Medicine (IRB00305245), with local IRBs at the Canadian and Swedish sites.”
“125 (53%) provided consent and 99 were randomized”
“For U.S. sites, the central Institutional Review Board (IRB) was at Johns Hopkins Medicine (IRB00305245), with local IRBs at the Canadian and Swedish sites.”
“Of 237 approached for the study, 125 (53%) provided consent”
The shunt valve is identified by name, manufacturer, and location: 'Codman Certas Plus with SiphonGuard®, Integra LifeSciences, Princeton, NJ'. However, the statistical software (e.g., SAS, R) and the MRI analysis software are not named. No other biological resources (antibodies, cell lines, etc.) are used, so those are not applicable.
“The study intervention was the initial setting of a commercially available noninvasively adjustable shunt valve (Codman Certas Plus with SiphonGuard ® , Integra LifeSciences, Princeton, NJ)”
“MRI evaluation of lateral ventricular volume and the Evans ratio was performed on de-identified T1 images using segmentation from spatially localized atlas network tiles and a deep-learning “brain extraction tool”.”
The primary analysis uses linear regression with adjustment for covariates. Tests are named (linear regression, Fisher's exact test). Exact p-values are given for the primary outcome (P<0.001) and secondary outcomes (P=0.003). Effect sizes with 95% CIs are reported for all outcomes. Data presentation includes tables with means, SDs, and per-group n, and a figure showing individual data points. Assumptions are not explicitly tested, but the use of robust methods is acceptable for a clinical trial. Mathematical plausibility checks confirmed no obvious errors.
“Three-month outcomes were analyzed using linear regression with treatment arm as the predictor, adjusting for baseline values and prespecified covariates”
“Gait Velocity (m/s) | 0.82 ± 0.27 | 0.87 ± 0.25 | 0.03 ± 0.23 | 0.23 ± 0.23 | 0.21 (0.12, 0.31) | <0.001”
“MoCA | 20.8 ± 4.1 | 21.4 ± 4.0 | 0.3 ± 2.9 | 1.3 ± 2.1 | 1.2 (0.1, 2.2) | NS”
The paper does not include a data availability statement. There is no mention of a repository deposit, accession numbers, or a mechanism for requesting data. Code sharing is not applicable as no custom code is described. Given that this is a clinical trial, the absence of a data availability statement is a notable omission.
“Trial registration number: PENS ClinicalTrials.gov (http://ClinicalTrials.gov) number, NCT05081128”
Methods are described in sufficient detail. The trial is registered (NCT05081128). All pre-specified outcomes are reported, including negative results (MoCA, OABQsf). Limitations are discussed in a dedicated section. Conclusions are proportional to the evidence. Funding sources and conflicts of interest are disclosed. No reporting checklist is referenced, but the paper follows CONSORT-like structure.
“Trial registration number: PENS ClinicalTrials.gov (http://ClinicalTrials.gov) number, NCT05081128”
“The major strengths of the study include its blinded design, statistical power, rigorous training of clinical assessors, and quality control procedures.”
“A potential limitation is the 3-month primary timepoint, which is standard for iNPH research. This study does not yet address the long-term outcomes.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 44 references by DOI: 39 verified — 5 no DOI (shown, not verified).
- NO DOIBDI-II, Beck Depression Inventory: Manual.No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAssessment of older people: self-maintaining and instrumental activities of daily living.No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA multiple testing procedure for clinical trials.No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA Simple Sequentially Rejective Multiple Test Procedure.No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe combination of probabilities arising from data in discrete distributions.No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT05081128LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttp://ClinicalTrials.govLIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, other, grammar.
- MINORotherFigure 1 legend“Click or tap here to enter text.”→ Remove the placeholder text left by the template.Template artifact in figure legend.
- MINORgrammarMethods, Statistical Analysis“A post-hoc logistic regression analysis of dichotomized 3-month mRS (values grouped as 0-2 and 3-6) was also performed controlling for treatment arm as the predictors, adjusting for baseline mRS.”→ Change 'as the predictors' to 'as the predictor'.Singular/plural disagreement.
- MINORconsistencyTable 3“NS”→ Report exact P values for MoCA and OABQsf, or state that they are reported by CI without hypothesis testing.Imprecise P-value reporting for secondary outcomes.
- MINORconsistencyTable 1, Header“Assigned treatment group”→ Consider capitalizing consistently as 'Assigned Treatment Group' for uniformity with other headers.Minor capitalization inconsistency.
In this post-publication audit, the published trial is robust in design, power, and transparency, and all 7 recomputable statistics are consistent; an informed reader should weigh the absence of a data availability statement and the incomplete identification of randomization method and software as the most substantive gaps. The missing explicit Helsinki/GCP compliance statement is a minor reporting gap that would warrant a small correction only if the authors wish to state the governing framework. No retracted citations, dead links, or statistical inconsistencies were found, so none of these gaps undermines the core conclusions.
- 1.HIGHdata codeAdd a data availability statement specifying a concrete managed-access route (e.g., de-identified individual participant data available on request to the corresponding author or via a data-access committee, with conditions and timeframe) and note it in a published correction.The complete absence of a data availability statement is the single fail dimension and a significant transparency gap for a clinical trial.
- 2.HIGHrigorDescribe the randomization method in the Randomization/Blinding subsection (e.g., central web-based system, block sizes, stratification variables) — currently only '1:1 ratio' is stated.Both reviewers flagged the missing randomization mechanism as a reporting gap that an informed reader cannot verify allocation concealment integrity without.
- 3.HIGHrigorIdentify the statistical software (name and version, e.g., SAS or R) used for all analyses in the Statistical Analysis section.Both reviewers noted the analysis software is unnamed, which hampers reproducibility of the computations.
- 4.MEDIUMrigorIdentify the deep-learning 'brain extraction tool' used for MRI segmentation by name and version (or RRID) in the Outcome Measures section.The MRI analysis pipeline is under-specified, which limits reproducibility of the imaging-derived outcomes (ventricular volume, Evans ratio).
- 5.MEDIUMethicsAdd an explicit regulatory-compliance statement naming the governing framework (Declaration of Helsinki and/or ICH-GCP) in the Ethics/Approvals section, or confirm in a correction that the listed IRB approvals constitute the compliance statement.An explicit framework statement is currently absent; Reviewer 1 scored this as a warn-level gap while Reviewer 2 considered IRB approval sufficient, so a one-line clarification resolves the ambiguity.
- 6.MEDIUMstatisticsReport exact P values for MoCA and OABQsf in Table 3, or explicitly state that these outcomes are reported by effect estimate and CI without hypothesis testing.Two secondary outcomes are shown only as 'NS', which is imprecise reporting the copyedit pass also flagged.
- 7.MEDIUMreportingReference the CONSORT reporting checklist (in Methods or supplementary material) to document reporting completeness.Both reviewers noted no reporting guideline is cited, which is expected for a published RCT.
- 8.MEDIUMcopyeditRemove the template placeholder text 'Click or tap here to enter text.' from the Figure 1 legend.This is a visible template artifact in the published figure legend that should be corrected.
- 9.LOWcopyeditCorrect the grammatical error in Methods, Statistical Analysis: 'controlling for treatment arm as the predictors' → 'as the predictor'.Singular/plural disagreement flagged by the copyedit pass.
- 10.LOWcopyeditCapitalize the Table 1 header 'Assigned treatment group' consistently (e.g., 'Assigned Treatment Group').Minor capitalization inconsistency flagged by the copyedit pass.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.