p75 neurotrophin receptor modulation in mild to moderate Alzheimer disease: a randomized, placebo-controlled phase 2a trial.
Shanks HRC, Chen K, Reiman EM, Blennow K, Cummings JL, Massa SM, Longo FM, Börjesson-Hanson A, Windisch M, Schmitz TW
- DOI
- 10.1038/s41591-024-02977-w
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/e3b5deee-2551-4be8-bcea-99860f754a0a is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The efficacy claim is based on surrogate biomarkers (CSF Aβ42, Aβ40, SNAP25, NG, YKL40, sMRI gray matter volume, FDG PET glucose metabolism) without demonstrating target engagement at the tested dose or citing validated evidence linking these surrogates to clinical outcomes. The paper acknowledges lack of direct p75NTR engagement markers and no significant cognitive effects.
“Currently, in vivo markers of direct p75 NTR engagement in humans are lacking. However, the CSF, sMRI and PET biomarkers included in this trial were prespecified on the basis of preclinical research examining pathways and mechanisms regulated by p75 NTR”
- 02Treatment effect not shown to be clinically meaningful
The reported effects are small percentage changes in biomarkers (e.g., Aβ42 -6.98%, SNAP25 -19.20%, NG -9.17%) with no anchor to clinical meaningfulness. The paper states no significant cognitive effects and acknowledges the study was not powered for cognitive outcomes.
“Although the effect was not statistically significant, 26-week treatment with LM11A-31 produced up to 50% slowing of cognitive decline relative to that observed in the placebo group”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported phase 2a randomized controlled trial of LM11A-31 in Alzheimer disease. The paper demonstrates strong scientific premise, rigorous design, and transparent reporting, with only minor gaps in data/code availability and some statistical reporting imprecision.
Both reviewers classified the study as interventional, and this was adopted. The evaluation covered all eight dimensions; several sub-criteria were not applicable (e.g., animal housing, cell line authentication) due to the human clinical trial nature. The reviewers diverged slightly on exact p-value reporting and software identification, but these were minor and did not affect the overall status.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks.
- CONSISTENTreported p < .050 · recomputed p = .033Reviewers 1, 2Check p-value for nasopharyngitis odds ratio
“nasopharyngitis, 5.41 (1.15 to 25.52); diarrhea, 12.22 (1.54 to 97.00); P < 0.05 for each”
Taken as given: The odds ratio is 5.41 with 95% CI 1.15 to 25.52.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed p-value from the odds ratio and its 95% CI using the pCI function.How we recomputed it: pCI(5.41, 1.15, 25.52, 1) - CONSISTENTreported p < .050 · recomputed p = .018Reviewers 1, 2Check p-value for diarrhea odds ratio
“nasopharyngitis, 5.41 (1.15 to 25.52); diarrhea, 12.22 (1.54 to 97.00); P < 0.05 for each”
Taken as given: The odds ratio is 12.22 with 95% CI 1.54 to 97.00.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed p-value from the odds ratio and its 95% CI using the pCI function.How we recomputed it: pCI(12.22, 1.54, 97.00, 1)
- lowinternal contradictionThe abstract states 242 participants, while the results state 241 were randomized. This is likely due to the safety population vs ITT population distinction, but it could be confusing.
242 participants with mild to moderate AD ... 241 were successfully randomized and accounted for in the intention-to-treat (ITT) population.
Abstractreviewer’s wording
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
4 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2LM11A-31 slows progression of pathophysiological features of AD.Significant drug-placebo differences were found on some biomarkers, but not on all, and the effects are exploratory.Evidence: Significant differences in CSF Aβ42, Aβ40, SNAP25, NG, YKL40, and imaging measures.
“significant drug–placebo differences were found, consistent with the hypothesis that LM11A-31 slows progression of pathophysiological features of AD”
AbstractFind in source - supportedReviewers 1, 2LM11A-31 is safe and tolerable in patients with mild to moderate AD.The primary endpoint of safety was met, with detailed AE reporting and DSMB conclusion.Evidence: The trial met its primary endpoint of safety and tolerability; DSMB concluded no overall safety concerns.
“This trial met its primary endpoint of safety and tolerability.”
AbstractFind in source - supportedReviewers 1, 2No significant effect of active treatment was observed on cognitive tests.The paper reports no significant differences on cognitive tests, consistent with the claim.Evidence: No significant differences on NTB, ADAS-Cog-13, MMSE, or Amunet.
“no significant effect of active treatment was observed on cognitive tests”
AbstractFind in source - supportedReviewers 1, 2Targeting p75NTR with LM11A-31 warrants further investigation in larger-scale clinical trials.The exploratory findings provide a rationale for further trials, and the claim is appropriately cautious.Evidence: The safety profile and exploratory biomarker findings support further investigation.
“these results suggest that targeting p75 NTR with LM11A-31 warrants further investigation in larger-scale clinical trials of longer duration”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe efficacy claim is based on surrogate biomarkers (CSF Aβ42, Aβ40, SNAP25, NG, YKL40, sMRI gray matter volume, FDG PET glucose metabolism) without demonstrating target engagement at the tested dose or citing validated evidence linking these surrogates to clinical outcomes. The paper acknowledges lack of direct p75NTR engagement markers and no significant cognitive effects.
“Currently, in vivo markers of direct p75 NTR engagement in humans are lacking. However, the CSF, sMRI and PET biomarkers included in this trial were prespecified on the basis of preclinical research examining pathways and mechanisms regulated by p75 NTR”
- INADEQUATEEffect sizeThe reported effects are small percentage changes in biomarkers (e.g., Aβ42 -6.98%, SNAP25 -19.20%, NG -9.17%) with no anchor to clinical meaningfulness. The paper states no significant cognitive effects and acknowledges the study was not powered for cognitive outcomes.
“Although the effect was not statistically significant, 26-week treatment with LM11A-31 produced up to 50% slowing of cognitive decline relative to that observed in the placebo group”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
2 integrity concerns flagged (0 high).
- lowotherThe paper reports a significant effect on diastolic blood pressure (P=0.036) but dismisses it as not clinically significant. This is a potential multiple comparisons issue, but the paper acknowledges the exploratory nature.
Longitudinal changes in diastolic blood pressure differed significantly among the three groups (P = 0.036).
Resultsreviewer’s wording
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction thoroughly reviews prior research on p75NTR as a therapeutic target, including preclinical studies of LM11A-31, and acknowledges limitations of existing AD therapies. The rationale for targeting p75NTR is logically developed, and the study hypothesis follows directly from the cited evidence. The paper also addresses limitations of prior research by noting the lack of human testing and the need for a phase 2a trial.
“Over the past two decades, multiple lines of evidence have converged on the p75 neurotrophin receptor (p75 NTR ) as a promising deep biology target for modifying neuronal dysfunction and degeneration in AD.”
“we hypothesized that modulation of p75 NTR using LM11A-31 in persons with AD would be well tolerated and would slow AD progression, as measured by biomarkers of synaptic function, degeneration and glial activation”
“Despite its fundamental functional role in neural and developmental cell biology, the therapeutic potential for targeted modulation of p75 NTR in humans has not been tested.”
“A limitation of these strategies is that they each target a narrow set of AD-related pathophysiological processes.”
The paper describes a randomized, double-blind, placebo-controlled, parallel-group design with 1:1:1 allocation. Randomization was performed using a list developed by Data Magik, with treatment center as a stratification variable. Blinding was maintained for sponsor, site personnel, participants, and caregivers. A priori power calculations were performed, and the sample size was adjusted based on a blinded review. Inclusion/exclusion criteria are detailed, and the analysis population (ITT) is defined. Outlier handling is described for biomarker analyses, and the trial includes appropriate controls (placebo). Independent replication is not applicable for a single pivotal trial.
“Participants were randomized 1:1:1 into placebo, 200 mg LM11A-31 or 400 mg LM11A-31. The randomization list was developed by Data Magik and was structured to allow for a total of at least 240 participants (80 per group), with treatment center as the only stratification variable.”
“The sponsor’s personnel, study sites’ personnel, participants and caregivers were blinded to the assigned treatment.”
“These analyses determined that 51 participants per group were required to demonstrate an effect size of 0.56 between either dose of LM11A-31 and placebo with 80% power and a type 1 error rate of 0.05 (two-tailed)”
“Participants were randomized 1:1:1 into placebo, 200 mg LM11A-31 or 400 mg LM11A-31. The randomization list was developed by Data Magik and was structured to allow for a total of at least 240 participants (80 per group), with treatment center as the only stratification variable.”
“The sponsor’s personnel, study sites’ personnel, participants and caregivers were blinded to the assigned treatment.”
“These analyses determined that 51 participants per group were required to demonstrate an effect size of 0.56 between either dose of LM11A-31 and placebo with 80% power and a type 1 error rate of 0.05 (two-tailed), resulting in an initial target of 60 participants per arm for a total of 180 participants.”
The paper reports age, sex, race, MMSE score, CSF Aβ42, and APOE4 status in Table 1. Health status is addressed through inclusion criteria and baseline characteristics. Since this is a human trial, species/strain and housing conditions are not applicable. Sex is reported for both sexes, so a justification for single-sex is not applicable.
“Age | 72 (8.00) | 72 (7.75) | 72 (8.00) | H = 0.81 | 0.67 |”
“Males, n (%) | 35 (43.2) | 38 (48.7) | 40 (48.2) | χ 2 = 0.60 | 0.74 |”
“General health status acceptable for participation in a 26-week clinical trial”
“Age | 72 (8.00) | 72 (7.75) | 72 (8.00) | H = 0.81 | 0.67 |”
“Screening MMSE | 22 (4.00) | 22 (5.00) | 23 (4.00) | H = 3.19 | 0.20 |”
The paper explicitly names the IRBs that approved the study for each country (e.g., IRB00002556 for Austria) and states that the trial was conducted in accordance with the Declaration of Helsinki and ICH-GCP. Informed consent is described as obtained before the screening visit. Regulatory compliance is stated by naming the Declaration of Helsinki and ICH-GCP.
“The IRBs approving the trial were IRB00002556 (Austria), IRB00002091 (Czech Republic), IRB00007525 (Germany), IRB00004959 (Sweden) and IRB00002590 (Spain).”
“Informed consent was obtained before the screening visit.”
“The trial was conducted in accordance with the Declaration of Helsinki and Good Clinical Practice of the International Council for Harmonization of Technical Requirements for Pharmaceuticals for Human Use (ICH-GCP).”
“The IRBs approving the trial were IRB00002556 (Austria), IRB00002091 (Czech Republic), IRB00007525 (Germany), IRB00004959 (Sweden) and IRB00002590 (Spain).”
“Informed consent was obtained before the screening visit.”
“The trial was conducted in accordance with the Declaration of Helsinki and Good Clinical Practice of the International Council for Harmonization of Technical Requirements for Pharmaceuticals for Human Use (ICH-GCP).”
The investigational product LM11A-31 is described with its formulation (oral capsules) and dosing regimen (200 mg or 400 mg twice daily). The placebo is also described. Software tools are identified, including SPM12 and CAT12, with a link to preprocessing scripts. Antibodies, cell lines, and mycoplasma testing are not applicable as this is a clinical trial without wet-lab assays.
“A single administration of medication consisted of the following: two capsules of 200 mg placebo (placebo group), one capsule of 200 mg LM11A-31 and one capsule of 200 mg placebo (200 mg LM11A-31 group) or two capsules of 200 mg LM11A-31 (400 mg LM11A-31 group).”
“Wrapper scripts to call SPM12 and CAT12 functions for MRI and PET preprocessing are available at https://github.com/hayleyshanks/Longitudinal-MRI-PET-preproc”
“A single administration of medication consisted of the following: two capsules of 200 mg placebo (placebo group), one capsule of 200 mg LM11A-31 and one capsule of 200 mg placebo (200 mg LM11A-31 group) or two capsules of 200 mg LM11A-31 (400 mg LM11A-31 group).”
“Wrapper scripts to call SPM12 and CAT12 functions for MRI and PET preprocessing are available at https://github.com/hayleyshanks/Longitudinal-MRI-PET-preproc”
The paper names statistical tests (e.g., Wilcoxon rank sum, Kruskal-Wallis, Fisher's exact, voxel-wise ANOVA) and reports effect sizes with 95% CIs for key outcomes. P-values are reported as exact values for many comparisons, though some are reported as thresholds (e.g., P < 0.05). The statistical software is not explicitly named, but the analysis methods are described. Data presentation includes box plots with per-group n. Mathematical plausibility is not applicable for large-N continuous outcomes.
“Significant differences in median change between placebo and LM11A-31 groups were investigated using Wilcoxon rank sum tests with 95% bootstrap CIs from 5,000 bootstrap iterations.”
“The difference in median annual percent change of Aβ42 in the LM11A-31 group relative to the placebo group was −6.98% (95% CI, −14.22% to −1.45%).”
“P < 0.05 for each”
“Significant differences in median change between placebo and LM11A-31 groups were investigated using Wilcoxon rank sum tests with 95% bootstrap CIs from 5,000 bootstrap iterations.”
“LM11A-31 significantly slowed longitudinal increases in Aβ42 compared to placebo (Fig. ; P rank sum = 0.037).”
“The difference in median annual percent change of Aβ42 in the LM11A-31 group relative to the placebo group was −6.98% (95% CI, −14.22% to −1.45%).”
The data availability statement says data can be shared in compliance with regulations but does not specify a concrete mechanism or platform, only directing requests to corresponding authors. This is reported_but_inadequate. The custom code for majority count statistics is not publicly available, but wrapper scripts for MRI/PET preprocessing are on GitHub. Repository deposit and accession numbers are not applicable for patient-level data.
“Data files containing pseudonymized participant data (baseline characteristics, raw data used to conduct primary and exploratory endpoint analyses reported in this article) can be shared in compliance with current data protection regulations by the EU. All requests for data access should be directed to the corresponding authors.”
“The custom code to perform the majority count statistics and Monte Carlo simulations (Fig. ) developed by K.C. is not publicly available but may be made available to qualified researchers on reasonable request to the corresponding authors.”
“The custom code to perform the majority count statistics and Monte Carlo simulations (Fig. ) developed by K.C. is not publicly available but may be made available to qualified researchers on reasonable request to the corresponding authors.”
Methods are detailed enough for replication. The trial is registered with EU and ClinicalTrials.gov numbers. A reporting summary is mentioned. All outcomes, including non-significant ones, are reported. Limitations are discussed. Conclusions are proportional to the evidence. Funding and competing interests are disclosed.
“EU Clinical Trials registration: 2015-005263-16 (https://www.clinicaltrialsregister.eu/ctr-search/search?query=eudract_number:2015-005263-16) ; ClinicalTrials.gov registration: NCT03069014 (https://clinicaltrials.gov/ct2/show/NCT03069014) .”
“By design, this phase 2a safety trial had several limitations for detecting cognitive effects, including a small number of participants and a relatively short 26-week study duration .”
“This trial was sponsored and funded by PharmatrophiX (Menlo Park, California) and the National Institute on Aging (NIA AD pilot trial 1R01AG051596).”
“This trial was sponsored and funded by PharmatrophiX (Menlo Park, California) and the National Institute on Aging (NIA AD pilot trial 1R01AG051596).”
Registered (2 IDs: ClinicalTrials.gov, EudraCT). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 102 references by DOI: 101 verified — 1 no DOI (shown, not verified).
- NO DOIMATLAB code to perform longitudinal structural MRI and PET analysesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
3 data/code links checked; 3 live.
- datahttps://adni.loni.usc.edu/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttp://adni.loni.usc.edu/wp-content/uploads/how_to_apply/ADNI_Acknowledgement_List.pdfLIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/hayleyshanks/Longitudinal-MRI-PET-preprocResolves to GitHub (code repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity, other.
- MINORconsistencyAbstract“242 participants”→ Ensure consistency with the number of participants randomized (241) and safety population (242) throughout the text.The abstract states 242 participants, but the results state 241 were randomized. This is a minor inconsistency.
- MINORclarityResults, Primary outcome“P < 0.05 for each”→ Consider reporting exact p-values for the odds ratios to improve precision.Reporting exact p-values would be more informative.
- MINORconsistencyAbstract“242 participants”→ Ensure consistency with the number of randomized participants (241) mentioned in the results.The abstract states 242 participants, but the results state 241 were randomized. Clarify the safety vs ITT population.
- MINORclarityResults, Participant disposition“Of these individuals, 221 completed the study as outlined in the protocol and 211 completed the study at the 26-week visit”→ Clarify the difference between 'completed the study as outlined' and 'completed the 26-week visit'.The distinction is not immediately clear; consider defining these terms.
- MINORotherTable 2“freq”→ Define 'freq' in the table footnote.The footnote defines 'freq' as total number of events, but it could be clearer.
The published work is robust and generally well-reported, but readers should weigh the vague data availability statement and the lack of public custom code as limitations for reproducibility. The minor inconsistency in participant numbers (242 vs 241) and threshold p-values are minor reporting issues that could warrant a correction or clarification.
- 1.HIGHdata codeIn the Data availability section, specify a concrete data access mechanism (e.g., a data access committee or a platform like Vivli) and the conditions for access, rather than just directing requests to corresponding authors.The current statement is vague and does not meet reproducibility standards for a data-driven clinical trial.
- 2.HIGHdata codeIn the Code availability section, make the custom code for majority count statistics and Monte Carlo simulations publicly available in a repository with a DOI, or provide a detailed algorithm description.The custom code is essential for reproducing the primary analysis and is currently only available on request.
- 3.MEDIUMstatisticsIn the Statistical analysis section, explicitly name the statistical software used (e.g., R version, SPSS) and provide version numbers for SPM12 and CAT12.Software identification is a standard reproducibility requirement and is currently missing.
- 4.MEDIUMstatisticsIn the Results section, report exact p-values for all comparisons instead of thresholds like 'P < 0.05'.Exact p-values improve precision and allow readers to assess the strength of evidence.
- 5.MEDIUMreportingIn the Abstract and Results, clarify the discrepancy between 242 participants (abstract) and 241 randomized (results), specifying the safety vs ITT populations.The inconsistency could confuse readers and should be resolved for clarity.
- 6.MEDIUMreportingIn the Results, clarify the difference between 'completed the study as outlined' and 'completed the 26-week visit'.The distinction is not immediately clear and should be defined for transparency.
- 7.MEDIUMreportingIn the Methods or supplementary, reference the CONSORT checklist and provide a CONSORT flow diagram with exact numbers for each stage, including screen failures.Explicitly following CONSORT improves reporting completeness and is expected for RCTs.
- 8.MEDIUMreportingIn the Methods, provide more detail on the randomization sequence generation (e.g., random number generator used) and clarify outlier handling for the primary analysis.These details are important for reproducibility and are currently only partially described.
- 9.LOWreportingIn the Data availability statement, clarify whether data can be shared with researchers outside the EU and under what conditions.The current statement is ambiguous about international data sharing.
- 10.LOWreportingIn Table 2, define 'freq' more clearly in the footnote.The current definition could be clearer to avoid misinterpretation.
- 11.LOWreportingIn the Discussion, consider discussing the potential impact of the lack of diversity in the study population on generalizability.Addressing diversity would strengthen the discussion of limitations.
- 12.LOWreportingIn the Data availability statement, consider providing a data dictionary for the variables.A data dictionary would facilitate data reuse and reproducibility.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.