Intrathecal onasemnogene abeparvovec in treatment-naive patients with spinal muscular atrophy: a phase 3, randomized controlled trial.
Proud CM, Vũ DC, Wilmshurst JM, Sanmaneechai O, Gulati S, Xiong H, Moreno HC, Tay SKH, Thong MK, Born AP, Banzzatto Ortega A, Jong YJ, Al-Muhaizea MA, Lee AW, Visootsak J, Tauscher-Wisniewski S, Alecu I, Parlikar R, Finkel RS, STEER Study Group
- DOI
- 10.1038/s41591-025-04103-w
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/e32281f8-46c9-4a37-8c4c-ffb114eafb82 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 4 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy endpoint is the change from baseline in Hammersmith Functional Motor Scale-Expanded (HFMSE) score, which is a functional motor scale, not a hard clinical outcome. The paper does not provide evidence of target engagement at the tested dose (e.g., PK/PD data) nor does it cite validated evidence linking HFMSE changes to long-term clinical outcomes such as survival or need for ventilation. The claim of clinical benefit is based on a surrogate motor function scale.
“Primary efficacy endpoint was change from baseline in Hammersmith Functional Motor Scale-Expanded (HFMSE) score.”
- 02Treatment effect not shown to be clinically meaningful
The primary effect size is a least squares mean difference of 1.88 points on the HFMSE scale (0-66 points), which is a small fraction of the scale range. The paper mentions a 3-point threshold for responder analysis and a 1.5-point minimal clinically important difference, but the observed difference of 1.88 is below the 3-point threshold and only slightly above the 1.5-point MCID. The clinical meaningfulness is not robustly established, especially given the small magnitude relative to the scale.
“least squares mean difference, 1.88 (95% confidence interval: 0.51−3.25); P = 0.0074”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted Phase 3 randomized, sham-controlled, double-blind trial with rigorous design, thorough ethical documentation, and transparent reporting. Minor reporting gaps include lack of explicit statistical software identification and explicit reporting guideline statement, but these do not undermine the overall robustness.
Both reviewers independently scored all dimensions as pass with high confidence, and their evidence was consistent. The statistics verification checked 5 tests, all consistent; citation check found no retracted or missing references; integrity check found only low-severity internal consistency observations that were resolved as expected. The paper is a clinical trial, so many sub-criteria (e.g., animal housing, cell line authentication) are not applicable.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 5 tests: 5 consistent, 0 inconsistent; 2 recomputed directly from the reported test statistics, 3 via agent-written checks.
- CONSISTENTreported p = .088 · recomputed p = .088Recomputed odds ratio 2.03 (95% CI 0.90–4.57), reported p=0.0879
“odds ratio: 2.03 (95% confidence interval: 0.90−4.57); P = 0.0879”
Taken as given: 0.90–4.57 is a two-sided 95% confidence interval for the odds ratio of 2.03, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.0879 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.03, 0.9, 4.57, 1) - CONSISTENTreported p = .645 · recomputed p = .647Recomputed odds ratio 1.27 (95% CI 0.46–3.56), reported p=0.6448
“odds ratio: 1.27 (95% confidence interval: 0.46−3.56), P = 0.6448”
Taken as given: 0.46–3.56 is a two-sided 95% confidence interval for the odds ratio of 1.27, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.6448 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.27, 0.46, 3.56, 1) - CONSISTENTreported p = .007 · recomputed p = .007Reviewers 1, 2Primary endpoint LSM difference p-value from CI
“least squares mean difference, 1.88 (95% confidence interval: 0.51−3.25); P = 0.0074”
Taken as given: The CI is a 95% two-sided confidence interval for the LSM difference.; The estimate is the LSM difference (1.88).; The p-value is two-sided.Method: Recomputed two-sided p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(1.88, 0.51, 3.25, 0) - CONSISTENTreported p = .012 · recomputed p = .012Reviewers 1, 2Secondary endpoint RULM LSM difference p-value from CI
“LSM difference, 1.52 (95% confidence interval: 0.34−2.71); P = 0.0122”
Taken as given: The CI is a 95% two-sided confidence interval for the LSM difference.; The estimate is the LSM difference (1.52).; The p-value is two-sided.Method: Recomputed two-sided p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(1.52, 0.34, 2.71, 0) - CONSISTENTreported p = .088 · recomputed p = .088Reviewers 1, 2Secondary endpoint HFMSE responder odds ratio p-value from CI
“odds ratio: 2.03 (95% confidence interval: 0.90−4.57); P = 0.0879”
Taken as given: The CI is a 95% two-sided confidence interval for the odds ratio.; The estimate is the odds ratio (2.03).; The p-value is two-sided.Method: Recomputed two-sided p-value from the log odds ratio and its 95% CI using the normal approximation.How we recomputed it: pCI(2.03, 0.90, 4.57, 1)
- lowinternal contradictionThe text states 'One participant in each group experienced an AE that led to discontinuation from the study.' but Table 2 shows 'Any AE leading to study discontinuation' as 1 (1.3) for OAV101 IT and 1 (2.0) for sham, which is consistent. However, the text later says 'One participant from the sham group discontinued due to an SAE of pneumonia aspiration, and one participant in the OAV101 IT group completed Period 1 but did not meet eligibility criteria for Period 2 and was, therefore, discontinued from the study.' This is consistent.
“One participant in each group experienced an AE that led to discontinuation from the study.”
ResultsFind in source - lowinternal contradictionThe abstract states 126 patients received OAV101 IT (n=75) or sham (n=51), but the results section says 122 completed Period 1. This is expected due to discontinuations.
In total, 126 patients received OAV101 IT ( n = 75) or a sham procedure ( n = 51), and 122 participants (OAV101 IT, n = 72; sham, n = 50) completed the 52-week Period 1.
Abstractreviewer’s wording
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
6 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1OAV101 IT provides benefit across a broad age range.The primary endpoint was met in the overall population, but subgroup analyses were exploratory and not powered.Evidence: Subgroup analyses show numerical improvements but not all statistically significant.
“In both the 2 to <5 years and 5 to <18 years age subgroups, the percentage of participants who achieved a ≥3-point increase in HFMSE score was greater in the OAV101 IT group compared with the sham group.”
DiscussionFind in source - partialReviewer 2OAV101 IT is a one-time treatment option for a broad range of SMA patients.The study shows benefit in the 2-<18 age group, but long-term durability and broader generalizability are not fully established.Evidence: Primary endpoint met in the overall population; long-term follow-up is ongoing.
“OAV101 IT offers a fixed dose that reduces systemic viral vector exposure, making it a potential one-time treatment for a broad range of patients with SMA regardless of age and weight.”
DiscussionFind in source - supportedReviewers 1, 2OAV101 IT significantly improves motor function compared with sham.The primary endpoint was met with a statistically significant LSM difference.Evidence: Primary endpoint: LSM difference 1.88 (95% CI 0.51-3.25), P=0.0074.
“The primary endpoint was met: patients treated with OAV101 IT demonstrated a significant increase in HFMSE score compared with sham (least squares mean difference, 1.88 (95% confidence interval: 0.51−3.25); P = 0.0074).”
AbstractFind in source - supportedReviewer 1Safety profile is acceptable with similar AE rates.AE, SAE, and AESI incidences were similar between groups.Evidence: Table 2 shows similar AE rates; no deaths.
“Overall incidence of adverse events (AEs), serious adverse events (SAEs) and adverse events of special interest (AESI) was similar between groups.”
AbstractFind in source - supportedReviewers 1, 2Secondary endpoints favored OAV101 IT but did not reach statistical significance.The paper reports nominal p-values and states significance was not met per the prespecified strategy.Evidence: Secondary endpoints reported with nominal p-values; e.g., RULM P=0.0122.
The secondary endpoints did not achieve statistical significance according to the prespecified multiple testing strategy.
Resultsreviewer’s wording - supportedReviewer 2The overall safety findings were acceptable.Similar incidences of AEs, SAEs, and AESI between groups support this claim.Evidence: Table 2 shows similar AE/SAE/AESI rates.
“The overall safety findings were acceptable, with similar incidences of AEs, SAEs and AESI in the OAV101 IT and sham groups.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy endpoint is the change from baseline in Hammersmith Functional Motor Scale-Expanded (HFMSE) score, which is a functional motor scale, not a hard clinical outcome. The paper does not provide evidence of target engagement at the tested dose (e.g., PK/PD data) nor does it cite validated evidence linking HFMSE changes to long-term clinical outcomes such as survival or need for ventilation. The claim of clinical benefit is based on a surrogate motor function scale.
“Primary efficacy endpoint was change from baseline in Hammersmith Functional Motor Scale-Expanded (HFMSE) score.”
- INADEQUATEEffect sizeThe primary effect size is a least squares mean difference of 1.88 points on the HFMSE scale (0-66 points), which is a small fraction of the scale range. The paper mentions a 3-point threshold for responder analysis and a 1.5-point minimal clinically important difference, but the observed difference of 1.88 is below the 3-point threshold and only slightly above the 1.5-point MCID. The clinical meaningfulness is not robustly established, especially given the small magnitude relative to the scale.
“least squares mean difference, 1.88 (95% confidence interval: 0.51−3.25); P = 0.0074”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on SMA treatments and the STRONG trial, acknowledges the limitations of existing therapies, and justifies the need for a one-time gene therapy. The rationale links the unmet need to the study objectives and hypothesis.
“As such, there remains a need for a safe and effective, one-time gene therapy option that provides sustained expression of SMN protein.”
“Challenges associated with chronic administration include the burden of repeated administration and potential non-compliance.”
“As such, there remains a need for a safe and effective, one-time gene therapy option that provides sustained expression of SMN protein.”
“Challenges associated with chronic administration include the burden of repeated administration and potential non-compliance.”
Randomization was stratified by age and baseline HFMSE score, with a 3:2 allocation. Blinding of participants, caregivers, evaluators, and investigators is described. The sample size was pre-specified (125 participants). Inclusion/exclusion criteria are detailed. The analysis population and missing-data approach are defined.
“Randomization was stratified by age and the highest pretreatment HFMSE score at screening as follows to ensure that both OAV101 IT and sham groups were balanced in terms of baseline characteristics of age and motor function: age 2 to <5 years, HFMSE score ≤15 or >15; age 5 to <13 years, HFMSE score ≤10 or >10; age 13 to <18 years, no stratification based on HFMSE score.”
“Participants (and their caregivers), clinical evaluators who performed the HFMSE and the RULM assessments and study investigators were blinded.”
“The STEER study aimed to enroll 125 participants to receive OAV101 IT ( n ≈ 75) or sham procedure ( n ≈ 50).”
“Randomization was stratified by age and the highest pretreatment HFMSE score at screening as follows to ensure that both OAV101 IT and sham groups were balanced in terms of baseline characteristics of age and motor function: age 2 to <5 years, HFMSE score ≤15 or >15; age 5 to <13 years, HFMSE score ≤10 or >10; age 13 to <18 years, no stratification based on HFMSE score.”
“Participants (and their caregivers), clinical evaluators who performed the HFMSE and the RULM assessments and study investigators were blinded.”
“The STEER study aimed to enroll 125 participants to receive OAV101 IT ( n ≈ 75) or sham procedure ( n ≈ 50).”
The paper reports age, sex, SMN2 copy number, and baseline motor function scores. Since both sexes are enrolled, sex justification is not applicable. Age and health status are reported.
“Male | 34 (45.3) | 28 (54.9) | 62 (49.2)”
“Mean (s.d.) age at dosing was 5.89 (3.58; range, 2.1–16.6) years for the OAV101 IT group and 5.87 (3.05; range, 2.4–14.2) years for the sham group.”
“Male | 34 (45.3) | 28 (54.9) | 62 (49.2)”
“Mean (s.d.) age at dosing was 5.89 (3.58; range, 2.1–16.6) years for the OAV101 IT group and 5.87 (3.05; range, 2.4–14.2) years for the sham group.”
The methods list numerous named IRBs that approved the study, and written informed consent from parents/guardians is stated. Compliance with ICH-GCP and Declaration of Helsinki is declared.
“Written informed consent was obtained from parents or legal guardians (and assent as appropriate) before participation.”
“STEER was undertaken in accordance with the International Council for Harmonisation E6 Guidelines for Good Clinical Practice with the ethical principles in accordance with the Declaration of Helsinki.”
“Written informed consent was obtained from parents or legal guardians (and assent as appropriate) before participation.”
“STEER was undertaken in accordance with the International Council for Harmonisation E6 Guidelines for Good Clinical Practice with the ethical principles in accordance with the Declaration of Helsinki.”
The drug is named, with dose (1.2 × 10^14 vector genomes) and route (intrathecal). The sham procedure is described. No other key biological/chemical resources are used, so other sub-criteria are not applicable.
“OAV101 IT was delivered as a single IT injection under sedation/anesthesia.”
“evaluating the efficacy and safety of OAV101 IT (1.2 × 10 14 vector genomes)”
“OAV101 IT delivers the SMN transgene via an AAV9 vector and is identical to that used in the STRONG study and the current intravenous formulation.”
“OAV101 IT (1.2 × 10 14 vector genomes)”
The primary analysis uses a mixed model with repeated measurements, and the paper reports LSM differences with 95% CIs and p-values. Secondary endpoints are reported with nominal p-values. The paper reports by estimation with CIs, so exact p-values are not required for all endpoints. Software is not explicitly named, but the analysis is standard.
“The primary efficacy endpoint was analyzed using a mixed model with repeated measurements, with the observed change from baseline in HFMSE score at all post-baseline visits (through the end of Follow-Up Period 1) as the dependent variable.”
“least squares mean difference, 1.88 (95% confidence interval: 0.51−3.25); P = 0.0074”
“The primary efficacy endpoint was analyzed using a mixed model with repeated measurements, with the observed change from baseline in HFMSE score at all post-baseline visits (through the end of Follow-Up Period 1) as the dependent variable.”
“least squares mean difference, 1.88 (95% confidence interval: 0.51−3.25)”
The data availability statement names a managed-access platform (ClinicalStudyDataRequest) with conditions for access. Since patient-level data are identifiable, repository deposit and accession numbers are not applicable. No bespoke code is mentioned.
“Novartis is committed to sharing with qualified external researchers access to patient-level data and supporting clinical documents from eligible studies. These requests are reviewed and approved by an independent review panel on the basis of scientific merit.”
“This trial data availability is according to the criteria and process described at https://www.clinicalstudydatarequest.com/”
The trial is registered (NCT05089656). Limitations are discussed. Conclusions are proportional. Funding and competing interests are disclosed. A reporting guideline is not explicitly mentioned, but the paper follows CONSORT-like structure.
“Trial registration: ClinicalTrials.gov identifier: NCT05089656”
“Possible limitations of this study include broad inclusion of all patients regardless of baseline HFMSE and wide age range in the eligibility criteria.”
“Novartis Pharma AG sponsored this clinical trial.”
“Trial registration: ClinicalTrials.gov identifier: NCT05089656”
“Possible limitations of this study include broad inclusion of all patients regardless of baseline HFMSE and wide age range in the eligibility criteria.”
“This study was supported by Novartis Pharma AG, which was involved in the study design, data collection, data analysis, data interpretation and writing of all related reports and publications.”
Registered (3 IDs: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 27 references by DOI: 22 verified — 5 no DOI (shown, not verified).
- NO DOISpinal muscular atrophyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISpinraza prescribing informationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEvrysdi prescribing informationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIZolgensma prescribing informationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIZolgensma summary of product characteristicsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link found; not probed for liveness in this run.
- datahttps://www.clinicalstudydatarequest.com/UNVERIFIEDLiveness indeterminate — content not checked.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAuthor list“Vũ Dũng Chí Wilmshurst Jo M.”→ Ensure author names are correctly formatted and separated.Possible missing comma or line break.
- MINORconsistencyAbstract“least squares mean difference, 1.88 (95% confidence interval: 0.51−3.25); P = 0.0074”→ Use consistent formatting for confidence intervals (e.g., en dash).Minor formatting inconsistency.
- MINORclarityResults, Safety“The most frequent SAEs were pneumonia (12.0% versus 13.7%) and vomiting (4.0% versus 0%) in the OAV101 IT group and pneumonia (12.0% versus 13.7%) and lower respiratory tract infection (2.7% versus 7.8%) in the sham group”→ Clarify the sentence structure to avoid ambiguity.The sentence lists SAEs for both groups in a confusing manner.
- MINORtypoAuthor list“Vũ Dũng Chí Wilmshurst Jo M.”→ Separate author names with commas for clarity.Author names appear concatenated.
- MINORconsistencyAbstract“least squares mean difference, 1.88 (95% confidence interval: 0.51−3.25); P = 0.0074”→ Ensure consistent use of 'least squares mean difference' vs 'LSM difference'.Minor inconsistency in terminology.
- MINORclarityResults, Safety“The most frequent SAEs were pneumonia (12.0% versus 13.7%) and vomiting (4.0% versus 0%) in the OAV101 IT group and pneumonia (12.0% versus 13.7%) and lower respiratory tract infection (2.7% versus 7.8%) in the sham group”→ Clarify which percentages correspond to which group.The sentence is confusingly structured.
The published work is robust and well-reported. An informed reader should weigh the minor reporting gaps (statistical software not named, reporting guideline not explicitly stated) as non-substantive. No erratum or re-analysis is warranted based on the evidence reviewed.
- 1.MEDIUMstatisticsIn the Methods, Data analysis section, explicitly name the statistical software (e.g., SAS version) used for all analyses.Naming the software enhances reproducibility and is a standard expectation for clinical trial reporting.
- 2.MEDIUMreportingIn the Methods or Reporting Summary, state explicit adherence to CONSORT guidelines and provide the checklist.Explicitly citing the reporting guideline improves transparency and aligns with journal requirements.
- 3.MEDIUMreportingIn the Methods, Data analysis section, specify the software used for the mixed model and logistic regression analyses.Clarifying the software for specific analyses aids reproducibility.
- 4.MEDIUMdata codeIn the Data availability section, clarify the timeframe for data access requests.Providing a timeline helps readers understand when data will be available.
- 5.MEDIUMreportingIn the Discussion, add a sentence on the generalizability of the results to broader SMA populations.Addressing generalizability strengthens the interpretation of the findings.
- 6.LOWcopyeditIn the Author list, separate author names with commas and ensure correct formatting (e.g., 'Vũ Dũng Chí Wilmshurst Jo M.').Correct author formatting is essential for indexing and professional presentation.
- 7.LOWcopyeditIn the Abstract, standardize the formatting of confidence intervals (e.g., use en dash consistently) and terminology ('least squares mean difference' vs 'LSM difference').Consistent formatting improves readability and professionalism.
- 8.LOWcopyeditIn the Results, Safety section, clarify the sentence listing SAEs to avoid ambiguity about which percentages correspond to which group.Clearer sentence structure prevents misinterpretation of safety data.
- 9.LOWreportingIn the Methods, specify the version of MedDRA used for coding AEs (already mentioned as 27.1, but could be more explicit).Explicitly stating the MedDRA version enhances methodological transparency.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.