Trial of N-Acetyl-l-Leucine in Niemann-Pick Disease Type C.
Bremova-Ertl T, Ramaswami U, Brands M, Foltan T, Gautschi M, Gissen P, Gowing F, Hahn A, Jones S, Kay R, Kolnikova M, Arash-Kaps L, Marquardt T, Mengel E, Park JH, Reichmannová S, Schneider SA, Sivananthan S, Walterfang M, Wibawa P, Strupp M, Martakis K
- DOI
- 10.1056/NEJMoa2310151
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/5c510431-4470-4884-aa96-f23b0a9082e2 is authoritative.
How this rating was calculated
- CitationsUnresolved reference ×4−1★
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 16 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is the SARA total score, a clinical rating scale for ataxia severity. Although it is a functional scale, it has not been validated in Niemann-Pick disease type C, and the paper explicitly states there is no validated biomarker or surrogate endpoint for the disease. The efficacy claim rests on a change in this scale without established clinical meaningfulness for this condition.
“The SARA is a validated clinical scale that measures the severity of neurologic signs and symptoms with internal consistency in patients with spinocerebellar ataxias but has not been validated in patients with Niemann–Pick disease type C.”
- 02Treatment effect not shown to be clinically meaningful
The primary effect is a mean difference of -1.28 points on the SARA total score (range 0-40). The minimal clinically important difference for SARA in Niemann-Pick disease type C is not established, and the paper does not anchor this change to clinical meaningfulness. The effect size is small relative to the scale range.
“least-squares mean difference, −1.28 points; 95% confidence interval, −1.91 to −0.65; P<0.001”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported phase 3 randomized controlled trial. The main methodological strengths include rigorous randomization, blinding, pre-specified analysis, and comprehensive reporting of demographics and ethics. The primary weakness is the vague data sharing statement, which lacks details on how to access the data.
Both reviewers classified the study as interventional, which is adopted. The evaluation covers all eight dimensions; several sub-criteria were marked not applicable due to the human clinical trial context (e.g., no animal housing, no cell lines). The statistics verification checked only 1 test and found it consistent; this does not confirm overall statistical correctness.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Primary endpoint p-value from reported least-squares mean difference and 95% CI
“least-squares mean difference, −1.28 points; 95% confidence interval, −1.91 to −0.65; P<0.001”
Taken as given: The CI is a 95% confidence interval for the difference in means.; The estimate is the least-squares mean difference.; The p-value is two-sided.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(-1.28, -1.91, -0.65, 0)
- lowinternal contradictionThe abstract states 60 patients were enrolled, but the results section mentions 59 patients completed the trial, which is explained by one withdrawal.
A total of 60 patients 5 to 67 years of age were enrolled. ... A total of 59 patients (98%) completed the trial.
Abstractreviewer’s wording
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
3 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Treatment with NALL for 12 weeks led to better neurologic status than placebo.The primary endpoint showed a statistically significant improvement in SARA score with NALL compared to placebo.Evidence: Primary endpoint: least-squares mean difference −1.28 points (95% CI −1.91 to −0.65; P<0.001).
“Among patients with Niemann–Pick disease type C, treatment with NALL for 12 weeks led to better neurologic status than placebo.”
ConclusionFind in source - supportedReviewers 1, 2The results for secondary end points were generally supportive of the primary analysis.Secondary endpoints showed improvements in the same direction, though not all were statistically significant.Evidence: Secondary endpoints: CGI-I, mDRS, SCAFI showed differences favoring NALL, with CIs mostly excluding zero.
“The results for the secondary end points were generally supportive of the findings in the primary analysis, but these were not adjusted for multiple comparisons.”
AbstractFind in source - supportedReviewers 1, 2The incidence of adverse events was similar with NALL and placebo.Adverse event rates were comparable between groups, with no treatment-related serious adverse events.Evidence: Safety results: 79 adverse events in 36 patients on NALL vs 75 in 30 on placebo; no treatment-related SAEs.
The incidence of adverse events was similar with NALL and placebo, and no treatment-related serious adverse events occurred.
Resultsreviewer’s wording
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is the SARA total score, a clinical rating scale for ataxia severity. Although it is a functional scale, it has not been validated in Niemann-Pick disease type C, and the paper explicitly states there is no validated biomarker or surrogate endpoint for the disease. The efficacy claim rests on a change in this scale without established clinical meaningfulness for this condition.
“The SARA is a validated clinical scale that measures the severity of neurologic signs and symptoms with internal consistency in patients with spinocerebellar ataxias but has not been validated in patients with Niemann–Pick disease type C.”
- INADEQUATEEffect sizeThe primary effect is a mean difference of -1.28 points on the SARA total score (range 0-40). The minimal clinically important difference for SARA in Niemann-Pick disease type C is not established, and the paper does not anchor this change to clinical meaningfulness. The effect size is small relative to the scale range.
“least-squares mean difference, −1.28 points; 95% confidence interval, −1.91 to −0.65; P<0.001”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior research on NALL's mechanism and animal models, and a phase 2b trial, providing a strong foundation. The rationale linking NALL's effects to the study objectives is clear. Limitations of prior work are implicitly addressed by the phase 3 design, though not explicitly discussed.
Randomization was performed by a CRO using computerized interactive response technology, with 1:1 allocation. Blinding was double-blind with patients, families, investigators, and sponsor representatives unaware of assignments. A sample size calculation was provided. Inclusion/exclusion criteria were pre-specified. Missing data handling was described. The trial is a crossover design with appropriate analysis.
“Randomization was performed by Medpace, a clinical research organization, with the use of computerized interactive response technology.”
“Randomization was performed by Medpace, a clinical research organization, with the use of computerized interactive response technology.”
The paper reports age group, sex, race, age at diagnosis, disease duration, and miglustat use. Both sexes are included, so sex justification is not applicable. Age and health status are reported. Species/strain and housing are not applicable for a human trial.
“Sex — no. (%) Female 27 (45) Male 33 (55)”
“Age group — no. (%) Pediatric, <18 yr 23 (38) Adult, ≥18 yr 37 (62)”
“Sex — no. (%) Female 27 (45) Male 33 (55)”
“Age group — no. (%) Pediatric, <18 yr 23 (38) Adult, ≥18 yr 37 (62)”
“Race or ethnic group — no. (%)† American Indian or Alaska native 0 Asian 0 Black or African American 2 (3) Native Hawaiian or other Pacific Islander 0 White 54 (90) Other 4 (7)”
The trial was approved by a central research ethics committee or institutional review board at each center. Written informed consent was obtained. The trial was registered (ClinicalTrials.gov and EudraCT). Regulatory compliance is implied by the approval and registration.
“The trial was approved by a central research ethics committee or an institutional review board at each center.”
“The trial was approved by a central research ethics committee or an institutional review board at each center.”
The investigational product NALL is named with dose and regimen. The placebo is described as matching. Statistical software (SAS 9.4) is identified. No antibodies, cell lines, or other biological resources are used, so those criteria are not applicable.
The primary analysis used a mixed-effects model with ANCOVA, and the paper reports exact p-values (P<0.001) and 95% CIs. Secondary endpoints are reported with CIs without p-values, which is appropriate. Software is identified. Data presentation includes individual patient data in figures and per-group n. Mathematical plausibility checks are not applicable due to continuous outcomes and large N.
The paper states that a data sharing statement is available with the full text, but the specific mechanism is not described in the provided text. No repository or accession numbers are given. Code sharing is not applicable.
“A data sharing statement provided by the authors is available with the full text of this article at NEJM.org.”
“A data sharing statement provided by the authors is available with the full text of this article at NEJM.org.”
The trial is registered (NCT05163288). Methods are detailed. Limitations are explicitly discussed. Conclusions are appropriately cautious. Funding and COI are disclosed. Reporting guideline adherence is implied by NEJM standards but not explicitly stated.
“ClinicalTrials.gov number, NCT05163288”
“Supported by IntraBio.”
“ClinicalTrials.gov number, NCT05163288”
“Supported by IntraBio.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 19 references by DOI: 14 verified — 4 DOI unresolved, 1 no DOI (shown, not verified).
- UNRESOLVED10.1046/j.0953-816x.2000.01435.xIn vitro effects of acetyl- dl -leucine (tanganil) on central vestibular neurons and vestibulo-ocular networks of the guinea-pigCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1038/s41598-021-88693-4N -acetyl- l -leucine improves functional recovery and attenuates cortical cell death and neuroinflammation after traumatic brain injury in miceCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1186/s13063-023-07377-0N-acetyl-L-leucine for Niemann–Pick type C: a multinational double-blind randomized placebo-controlled crossover studyCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1212/01.wnl.0000324860.76232.6aSCA Functional Index: a useful compound performance measure for spinocerebellar ataxiaCited DOI does not resolve to any Crossref record.
- NO DOIEQ-5D instruments: about EQ-5D-5LNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
2 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 2 minor suggestions below.
2 copyedit issues flagged: mostly typo, consistency.
- MINORtypoAbstract, Results“The mean (±SD) change from baseline in the SARA total score was −1.97±2.43 points after 12 weeks of receiving NALL and −0.60±2.39 points after 12 weeks of receiving placebo”→ Ensure consistent use of spaces around ± symbols.Minor formatting inconsistency.
- MINORconsistencyTable 1 footnote“The mean baseline scores are shown for 60 patients before first dose of NALL and for 59 patients before first dose of placebo.”→ Clarify why one patient is missing for placebo baseline.The footnote explains the missing patient, but it could be clearer.
The published paper is robust and well-reported. An informed reader should note the vague data sharing statement and the lack of explicit reporting guideline adherence as minor transparency gaps, but these do not undermine the core findings. No erratum or correction appears warranted based on this audit.
- 1.HIGHdata codeProvide a detailed data sharing statement in the manuscript, specifying the mechanism for requesting de-identified patient data (e.g., via a data access committee or a platform like Vivli) and the conditions for access.The current statement is vague and does not meet the standard for data availability in clinical trials.
- 2.HIGHreportingVerify the four references not found in the registry (DOIs: 10.1046/j.0953-816x.2000.01435.x, 10.1038/s41598-021-88693-4, 10.1186/s13063-023-07377-0, 10.1212/01.wnl.0000324860.76232.6a) and correct or replace them if they are incorrect or fabricated.References that cannot be located in any registry are a potential fabrication signal and must be resolved.
- 3.MEDIUMreportingExplicitly state adherence to a reporting guideline such as CONSORT in the methods or acknowledgments.Enhances transparency and allows readers to verify that reporting standards were followed.
- 4.MEDIUMreportingInclude the protocol and statistical analysis plan as supplementary materials or provide a link to them.Allows readers to verify that the analysis was pre-specified and reduces concerns about selective reporting.
- 5.MEDIUMreportingClarify the handling of missing data for secondary endpoints, as the paper notes no imputation was performed for these.Provides a complete picture of how missing data were addressed across all analyses.
- 6.MEDIUMreportingConsider reporting the results of the mSARA analysis with a p-value or confidence interval, as it was a primary endpoint in the US.Ensures consistency in reporting of primary endpoints across regulatory regions.
- 7.MEDIUMreportingAdd a note on the generalizability of the findings given the exclusion of patients younger than 4 years and those with advanced disease.Helps readers understand the population to which the results apply.
- 8.LOWcopyeditEnsure consistent use of spaces around ± symbols in the abstract (e.g., '−1.97±2.43' should be '−1.97 ± 2.43').Minor formatting inconsistency that should be corrected for professional presentation.
- 9.LOWcopyeditClarify the Table 1 footnote explaining why one patient is missing for placebo baseline scores.The current footnote could be clearer to avoid confusion.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.