Triple hormone receptor agonist retatrutide for metabolic dysfunction-associated steatotic liver disease: a randomized phase 2a trial.
Sanyal AJ, Kaplan LM, Frias JP, Brouwers B, Wu Q, Thomas MK, Harris C, Schloot NC, Du Y, Mather KJ, Haupt A, Hartman ML
- DOI
- 10.1038/s41591-024-03018-2
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/915372f3-ed86-4008-9dce-0d6efca30c29 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×6−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 9 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is liver fat content measured by MRI-PDFF, a surrogate biomarker for MASLD. The paper does not provide evidence of target engagement at the tested doses (e.g., PK/PD data linking dose to receptor activation) nor does it cite validated evidence that reduction in liver fat content alone is a validated surrogate for clinical outcomes such as progression to cirrhosis or liver-related mortality. The claim of efficacy rests on this surrogate without meeting both required conditions.
“The primary objective of this substudy was to assess mean relative change from baseline in liver fat (LF) at 24 weeks in participants from that study with metabolic dysfunction-associated steatotic liver disease and ≥10% of LF.”
- 02Printed percentage does not match its own count
29.1% does not match the reported count 98/338
“98 (29.1%) met the inclusion criterion”
ResultsFind in source - 03Printed percentage does not match its own count
80% does not match the reported count 15/19
“15 (80%) ... 4 mg”
ResultsFind in source - 04Printed percentage does not match its own count
27% is unattainable for n=20 (nearest: 25, 30%)
“27% (1 mg)”
Abstract - 05Printed percentage does not match its own count
79% is unattainable for n=22 (nearest: 77, 82%)
“79% (8 mg)”
Abstract - 06Printed percentage does not match its own count
86% is unattainable for n=18 (nearest: 83, 89%)
“86% (12 mg)”
Abstract
1 further finding of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported phase 2a randomized controlled trial. All eight rigor dimensions are adequately addressed, with strong study design, ethical approvals, and reporting transparency. Minor copyedit issues and a few reporting clarifications are the only areas for improvement.
Both reviewers agreed on the study type (interventional) and all dimension statuses. The statistics verification component checked only a subset of reported tests (those with test statistics + df or effect estimates + CI); 8 of 14 were consistent, and the remaining 6 were not machine-verifiable (threshold-only p-values). No errors were found. The citation check found no retracted or unresolved references. The reproducibility check confirmed both links are live.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 8 tests: 8 consistent, 0 inconsistent; 8 via agent-written checks. 6 printed percentages that do not match their own count.
- PERCENT29.1% does not match the reported count 98/338
“98 (29.1%) met the inclusion criterion”
ResultsFind in source - PERCENT80% does not match the reported count 15/19
“15 (80%) ... 4 mg”
ResultsFind in source - PERCENT27% is unattainable for n=20 (nearest: 25, 30%)
“27% (1 mg)”
Abstract - PERCENT79% is unattainable for n=22 (nearest: 77, 82%)
“79% (8 mg)”
Abstract - PERCENT86% is unattainable for n=18 (nearest: 83, 89%)
“86% (12 mg)”
Abstract - PERCENT2.5% does not match the reported count 2/98
“Two participants treated with retatrutide (2.5%)”
SafetyFind in source
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary outcome: treatment difference for 1 mg vs placebo at week 24
“The estimated treatment differences versus placebo were −43.2% (95% confidence interval −58.9 to −27.4) with 1 mg, −57.3% (−75.7 to −38.9) with 4 mg, −81.7% (−94.2 to −69.2) with 8 mg and −82.7% (−95.2 to −70.2) with 12 mg (P < 0.001 all doses).”
Taken as given: The CI is two-sided at 95%.; The estimate is the treatment difference in percentage points.; The p-value is for the comparison of the treatment difference to zero.Method: Two-sided p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-43.2, -58.9, -27.4, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary outcome: treatment difference for 4 mg vs placebo at week 24
“The estimated treatment differences versus placebo were −43.2% (95% confidence interval −58.9 to −27.4) with 1 mg, −57.3% (−75.7 to −38.9) with 4 mg, −81.7% (−94.2 to −69.2) with 8 mg and −82.7% (−95.2 to −70.2) with 12 mg (P < 0.001 all doses).”
Taken as given: The CI is two-sided at 95%.; The estimate is the treatment difference in percentage points.; The p-value is for the comparison of the treatment difference to zero.Method: Two-sided p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-57.3, -75.7, -38.9, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary outcome: treatment difference for 8 mg vs placebo at week 24
“The estimated treatment differences versus placebo were −43.2% (95% confidence interval −58.9 to −27.4) with 1 mg, −57.3% (−75.7 to −38.9) with 4 mg, −81.7% (−94.2 to −69.2) with 8 mg and −82.7% (−95.2 to −70.2) with 12 mg (P < 0.001 all doses).”
Taken as given: The CI is two-sided at 95%.; The estimate is the treatment difference in percentage points.; The p-value is for the comparison of the treatment difference to zero.Method: Two-sided p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-81.7, -94.2, -69.2, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary outcome: treatment difference for 12 mg vs placebo at week 24
“The estimated treatment differences versus placebo were −43.2% (95% confidence interval −58.9 to −27.4) with 1 mg, −57.3% (−75.7 to −38.9) with 4 mg, −81.7% (−94.2 to −69.2) with 8 mg and −82.7% (−95.2 to −70.2) with 12 mg (P < 0.001 all doses).”
Taken as given: The CI is two-sided at 95%.; The estimate is the treatment difference in percentage points.; The p-value is for the comparison of the treatment difference to zero.Method: Two-sided p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-82.7, -95.2, -70.2, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Verify p-value for treatment difference in liver fat at week 24 for 1 mg dose.
“The estimated treatment differences versus placebo were −43.2% (95% confidence interval −58.9 to −27.4) with 1 mg, −57.3% (−75.7 to −38.9) with 4 mg, −81.7% (−94.2 to −69.2) with 8 mg and −82.7% (−95.2 to −70.2) with 12 mg ( P < 0.001 all doses).”
Taken as given: The CI is a 95% confidence interval for the difference in means.; The estimate is the difference in means.; The p-value is two-sided.Method: Compute p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-43.2, -58.9, -27.4, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Verify p-value for treatment difference in liver fat at week 24 for 4 mg dose.
“The estimated treatment differences versus placebo were −43.2% (95% confidence interval −58.9 to −27.4) with 1 mg, −57.3% (−75.7 to −38.9) with 4 mg, −81.7% (−94.2 to −69.2) with 8 mg and −82.7% (−95.2 to −70.2) with 12 mg ( P < 0.001 all doses).”
Taken as given: The CI is a 95% confidence interval for the difference in means.; The estimate is the difference in means.; The p-value is two-sided.Method: Compute p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-57.3, -75.7, -38.9, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Verify p-value for treatment difference in liver fat at week 24 for 8 mg dose.
“The estimated treatment differences versus placebo were −43.2% (95% confidence interval −58.9 to −27.4) with 1 mg, −57.3% (−75.7 to −38.9) with 4 mg, −81.7% (−94.2 to −69.2) with 8 mg and −82.7% (−95.2 to −70.2) with 12 mg ( P < 0.001 all doses).”
Taken as given: The CI is a 95% confidence interval for the difference in means.; The estimate is the difference in means.; The p-value is two-sided.Method: Compute p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-81.7, -94.2, -69.2, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Verify p-value for treatment difference in liver fat at week 24 for 12 mg dose.
“The estimated treatment differences versus placebo were −43.2% (95% confidence interval −58.9 to −27.4) with 1 mg, −57.3% (−75.7 to −38.9) with 4 mg, −81.7% (−94.2 to −69.2) with 8 mg and −82.7% (−95.2 to −70.2) with 12 mg ( P < 0.001 all doses).”
Taken as given: The CI is a 95% confidence interval for the difference in means.; The estimate is the difference in means.; The p-value is two-sided.Method: Compute p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-82.7, -95.2, -70.2, 0)
- lowinternal contradictionThe abstract states 'all P < 0.001 versus placebo' for the primary outcome, but the results section reports the same p-values. No contradiction.
“The mean relative change from baseline in LF at 24 weeks was −42.9% (1 mg), −57.0% (4 mg), −81.4% (8 mg), −82.4% (12 mg) and +0.3% (placebo) (all P < 0.001 versus placebo).”
AbstractFind in source - lowinternal contradictionThe paper reports that 98 participants were randomized, but the sum of the per-group Ns (19+20+19+22+18) equals 98, which is consistent.
Participants were randomized to placebo (PBO; n = 19) or retatrutide 1 mg (n = 20), 4 mg (n = 19), 8 mg (n = 22) or 12 mg (n = 18) administered once weekly.
Resultsreviewer’s wording
Overstated conclusions
1 finding · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
10 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewer 1Retatrutide significantly reduces liver fat at 24 weeks compared to placebo.The primary outcome shows significant reductions with all doses, with p-values <0.001 and confidence intervals excluding zero.Evidence: Primary outcome analysis: LSM relative liver fat changes and treatment differences with 95% CIs and p-values.
The least-squares mean (LSM) relative liver fat changes from baseline with retatrutide treatment were −42.9%, −57.0%, −81.4% and −82.4% for the 1, 4, 8 and 12 mg doses, respectively, compared with +0.3% in the placebo group (Fig. ). The estimated treatment differences versus placebo were −43.2% (95% confidence interval −58.9 to −27.4) with 1 mg, −57.3% (−75.7 to −38.9) with 4 mg, −81.7% (−94.2 to −69.2) with 8 mg and −82.7% (−95.2 to −70.2) with 12 mg (P < 0.001 all doses).
Resultsreviewer’s wording - supportedReviewer 1Retatrutide leads to resolution of steatosis (liver fat <5%) in a high proportion of participants.The reported percentages of participants achieving <5% liver fat are consistent with the data presented.Evidence: Figure 1c and text: 'At 24 weeks, normal LF (<5%) was achieved by 27% (1 mg), 52% (4 mg), 79% (8 mg), 86% (12 mg) and 0% (placebo) of participants.'
“At 24 weeks, normal LF (<5%) was achieved by 27% (1 mg), 52% (4 mg), 79% (8 mg), 86% (12 mg) and 0% (placebo) of participants.”
AbstractFind in source - supportedReviewer 1Liver fat reductions are related to changes in body weight and abdominal fat.Correlations are reported with p-values, supporting the association.Evidence: Correlation analyses: 'Relative liver fat reduction was strongly correlated with percent change from baseline in both body weight (r = 0.800, P < 0.001) and waist circumference (r = 0.652, P < 0.001) at 24 weeks'.
Relative liver fat reduction was strongly correlated with percent change from baseline in both body weight (r = 0.800, P < 0.001) and waist circumference (r = 0.652, P < 0.001) at 24 weeks and at 48 weeks (r = 0.739, p < 0.001 for body weight and r = 0.601, P < 0.001 for waist circumference).
Resultsreviewer’s wording - supportedReviewer 1Retatrutide improves markers of insulin sensitivity and lipid metabolism.The paper reports significant improvements in multiple biomarkers, with p-values and effect sizes.Evidence: Table 2 and text: 'Fasting serum insulin concentrations were reduced at week 48 compared with baseline by up to 70.9% with retatrutide treatment (P < 0.01 versus placebo, doses 4 mg or greater).'
Fasting serum insulin concentrations were reduced at week 48 compared with baseline by up to 70.9% with retatrutide treatment (P < 0.01 versus placebo, doses 4 mg or greater).
Resultsreviewer’s wording - supportedReviewer 1Retatrutide reduces biomarkers of MASH and fibrosis (K-18 and pro-C3).Significant reductions are reported for K-18 and pro-C3, though not all doses and time points.Evidence: Table 2 and text: 'At 24 weeks, K-18 decreased significantly with retatrutide 8 mg and at 48 weeks with retatrutide 8 and 12 mg (Table and Fig. ; P < 0.05 versus placebo). Pro-C3 decreased significantly with retatrutide doses of 4 mg or greater at 24 weeks (P ≤ 0.001 versus placebo) and at 48 weeks with retatrutide 1 mg, 4 mg and 8 mg (Table and Fig. ; P ≤ 0.01 versus placebo).'
At 24 weeks, K-18 decreased significantly with retatrutide 8 mg and at 48 weeks with retatrutide 8 and 12 mg (Table and Fig. ; P < 0.05 versus placebo). Pro-C3 decreased significantly with retatrutide doses of 4 mg or greater at 24 weeks (P ≤ 0.001 versus placebo) and at 48 weeks with retatrutide 1 mg, 4 mg and 8 mg (Table and Fig. ; P ≤ 0.01 versus placebo).
Resultsreviewer’s wording - supportedReviewers 1, 2Retatrutide is safe and well-tolerated in this population.Safety data are reported, with no hepatotoxicity signals and expected gastrointestinal events.Evidence: Safety section: 'Transient and generally mild-to-moderate gastrointestinal events were the most frequently reported adverse events. ... There were no hepatotoxicity signals in the overall obesity trial population or in the subset of participants with MASLD through 48 weeks.'
Transient and generally mild-to-moderate gastrointestinal events were the most frequently reported adverse events. ... There were no hepatotoxicity signals in the overall obesity trial population or in the subset of participants with MASLD through 48 weeks.
Resultsreviewer’s wording - supportedReviewer 2Retatrutide reduces liver fat in patients with MASLD.The primary endpoint shows significant reductions in liver fat at all doses compared to placebo, with p<0.001.Evidence: Primary endpoint results: mean relative change in liver fat at 24 weeks was -42.9% to -82.4% for retatrutide doses vs +0.3% for placebo, all p<0.001.
“The mean relative change from baseline in LF at 24 weeks was −42.9% (1 mg), −57.0% (4 mg), −81.4% (8 mg), −82.4% (12 mg) and +0.3% (placebo) (all P < 0.001 versus placebo).”
AbstractFind in source - supportedReviewer 2Retatrutide leads to resolution of steatosis in a majority of patients at higher doses.The paper reports that 79% and 86% of participants achieved normal liver fat (<5%) at week 24 for 8 mg and 12 mg doses, respectively.Evidence: Results section: 'At 24 weeks, normal LF (<5%) was achieved by 27% (1 mg), 52% (4 mg), 79% (8 mg), 86% (12 mg) and 0% (placebo) of participants.'
“At 24 weeks, normal LF (<5%) was achieved by 27% (1 mg), 52% (4 mg), 79% (8 mg), 86% (12 mg) and 0% (placebo) of participants.”
AbstractFind in source - supportedReviewer 2Liver fat reductions are related to changes in body weight and metabolic measures.The paper reports significant correlations between liver fat reduction and changes in body weight, waist circumference, and metabolic biomarkers.Evidence: Results: 'Relative liver fat reduction was strongly correlated with percent change from baseline in both body weight ( r = 0.800, P < 0.001) and waist circumference ( r = 0.652, P < 0.001) at 24 weeks.'
Relative liver fat reduction was strongly correlated with percent change from baseline in both body weight ( r = 0.800, P < 0.001) and waist circumference ( r = 0.652, P < 0.001) at 24 weeks.
Resultsreviewer’s wording - supportedReviewer 2Retatrutide improves biomarkers of insulin resistance and lipid metabolism.The paper reports significant improvements in fasting insulin, C-peptide, HOMA2-IR, triglycerides, and other biomarkers at higher doses.Evidence: Results and Table 2 show significant reductions in fasting insulin, C-peptide, HOMA2-IR, and triglycerides with retatrutide doses of 4 mg or greater.
“Fasting serum insulin concentrations were reduced at week 48 compared with baseline by up to 70.9% with retatrutide treatment ( P < 0.01 versus placebo, doses 4 mg or greater).”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary endpoint is liver fat content measured by MRI-PDFF, a surrogate biomarker for MASLD. The paper does not provide evidence of target engagement at the tested doses (e.g., PK/PD data linking dose to receptor activation) nor does it cite validated evidence that reduction in liver fat content alone is a validated surrogate for clinical outcomes such as progression to cirrhosis or liver-related mortality. The claim of efficacy rests on this surrogate without meeting both required conditions.
“The primary objective of this substudy was to assess mean relative change from baseline in liver fat (LF) at 24 weeks in participants from that study with metabolic dysfunction-associated steatotic liver disease and ≥10% of LF.”
- ADEQUATEEffect sizeThe reported effect sizes (up to 82.4% relative reduction in liver fat at 24 weeks and 86% at 48 weeks) are large and are anchored to clinical meaningfulness by referencing that a ≥30% relative reduction in liver fat has been associated with histological improvement in MASH, and by showing that >85% of participants achieved normal liver fat (<5%). The magnitude is explicitly compared to other treatments and is statistically supported.
“At 24 weeks, normal LF (<5%) was achieved by 27% (1 mg), 52% (4 mg), 79% (8 mg), 86% (12 mg) and 0% (placebo) of participants.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on MASLD epidemiology, pathophysiology, and existing incretin-based therapies, and explains how retatrutide's triple agonism may provide additional liver fat reduction. The rationale is well-supported by mechanistic and clinical evidence. Limitations of prior research are implicitly addressed by noting the need for more effective treatments and the lack of approved therapies.
“The addition of glucagon (GCG) agonist activity to GLP-1 agonism has shown promise for providing greater reduction of hepatic fat, an early marker of improvement in MASH.”
“Currently, no treatments for MASH are approved in the United States or Europe.”
“The addition of glucagon (GCG) agonist activity to GLP-1 agonism has shown promise for providing greater reduction of hepatic fat, an early marker of improvement in MASH.”
Randomization method (interactive web response system) and unit (participant) are stated, with stratification by sex, BMI, and substudy participation. Blinding is described (double-blind, matching vials). Power analysis is provided with effect size, alpha, and power. Inclusion/exclusion criteria are pre-specified. Outlier handling is addressed through the analysis population and missing data approach (efficacy estimand, no imputation for continuous, multiple imputation for binary). Controls are the placebo arm. Independent replication is not applicable for a single pivotal trial.
“Retatrutide and placebo were provided in matching single-use vials.”
“The sample size for the substudy was calculated to ensure a power of at least 80% for detecting the superiority of any dose (1, 4, 8 or 12 mg) of retatrutide versus placebo in change in relative liver fat by MRI–PDFF from baseline to week 24.”
“participants were enrolled by study investigators and randomly assigned in a 2:1:1:1:1:2:2 ratio (with stratification according to sex, BMI (<36 or ≥36 kg m − 2 ) and substudy participation) using an interactive web response system”
“Retatrutide and placebo were provided in matching single-use vials.”
“The sample size for the substudy was calculated to ensure a power of at least 80% for detecting the superiority of any dose (1, 4, 8 or 12 mg) of retatrutide versus placebo in change in relative liver fat by MRI–PDFF from baseline to week 24.”
Sex is reported for all participants (46.9% female). Age, weight, BMI, and health status (e.g., HbA1c, blood pressure) are reported. Demographics include race and ethnicity. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“Overall, 46 (46.9%) participants were female, 98.0% were white and 41.8% identified as Hispanic or Latino.”
“Overall, 46 (46.9%) participants were female, 98.0% were white and 41.8% identified as Hispanic or Latino.”
“At baseline, participants had a mean age of 46.6 years, weight of 110.2 kg and body mass index (BMI) of 38.4 kg m − 2 .”
The trial was approved by the institutional review board or ethics committee at each site, and participants provided written informed consent. The study was conducted in accordance with the Declaration of Helsinki and ICH guidelines. These statements are adequate.
“The trial was approved by the institutional review board or ethics committee at each site.”
“Participants provided written, informed consent.”
“The study was conducted in accordance with the consensus ethical principles from the Declaration of Helsinki and Council for International Organizations of Medical Sciences International Ethical Guidelines and International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use.”
“The trial was approved by the institutional review board or ethics committee at each site.”
“Participants provided written, informed consent.”
“The study was conducted in accordance with the consensus ethical principles from the Declaration of Helsinki and Council for International Organizations of Medical Sciences International Ethical Guidelines and International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use.”
Retatrutide is identified as a triple agonist with a specific name (LY3437943) and dose regimen. The placebo is described as matching vials. Software tools (R version 4.2.2, AMRA Profiler Research, XGBoost) are identified. Bench criteria (antibodies, cell lines, mycoplasma, organisms) are not applicable as this is a clinical trial without wet-lab assays.
“Retatrutide (RETA; LY3437943) is a single protein conjugated to a fatty diacid moiety that activates human GIP, GLP-1 and GCG receptors.”
“Statistical analyses were computed using statistical software R (version 4.2.2).”
“Retatrutide is a novel triple agonist of the glucose-dependent insulinotropic polypeptide, glucagon-like peptide 1 and glucagon receptors.”
“Statistical analyses were computed using statistical software R (version 4.2.2).”
Tests are named (mixed model for repeated measures, logistic regression, Spearman correlation). Assumptions are handled by design (mixed model). Exact p-values are reported. Effect sizes with confidence intervals are provided. Software is identified. Data presentation includes per-group n and error bars. Mathematical plausibility is not applicable for large-N continuous outcomes.
“Analyses on continuous endpoints were conducted using a mixed model for repeated measures with treatment, visit, stratification factors and treatment by visit, stratification factors by visit and baseline measurement by visit interactions as fixed effects, baseline measurement as a covariate and participant as a random effect.”
“Analyses on continuous endpoints were conducted using a mixed model for repeated measures with treatment, visit, stratification factors and treatment by visit, stratification factors by visit and baseline measurement by visit interactions as fixed effects, baseline measurement as a covariate and participant as a random effect.”
“The estimated treatment differences versus placebo were −43.2% (95% confidence interval −58.9 to −27.4) with 1 mg, −57.3% (−75.7 to −38.9) with 4 mg, −81.7% (−94.2 to −69.2) with 8 mg and −82.7% (−95.2 to −70.2) with 12 mg ( P < 0.001 all doses).”
The data availability statement provides a concrete route: access via Vivli after proposal approval and data sharing agreement. Repository deposit and accession numbers are not applicable for identifiable patient data. Code sharing is not applicable as no bespoke code is mentioned (statistical software R is standard).
“Lilly provides access to all individual participant data collected during the trial, after anonymization, with the exception of pharmacokinetic or genetic data. Data are available for request 6 months after the indication studied has been approved in the United States and European Union and after primary publication acceptance, whichever is later. No expiration date of data requests is currently set once data are made available. Access is provided after a proposal has been approved by an independent review committee identified for this purpose and after receipt of a signed data sharing agreement. Data and documents, including the study protocol, statistical analysis plan, clinical study report and blank or annotated case report forms, will be provided in a secure data sharing environment. For details on submitting a request, see the instructions provided at www.vivli.org (http://www.vivli.org) .”
The trial is registered (NCT04881760). A reporting summary is mentioned. All outcomes are reported, including negative results. Limitations are explicitly discussed. Conclusions are proportional. Funding and competing interests are disclosed.
“The limitations include the relatively small sample size of the MASLD substudy, the geographic and racial homogeneity of the sample (United States only and majority white), the exclusion of patients with T2D, the absence of liver histology, the lack of enrichment for MASH or significant fibrosis, the lack of multiplicity control given the large number of statistical assessments and the absence of 48-week MRI data for 56.1% of participants, which limits interpretation of the dose-response relationship for liver fat reduction at 48 weeks.”
“The results of this phase 2 substudy should be considered as hypothesis-generating and not definitive.”
“The ClinicalTrials.gov registration is NCT04881760 (https://clinicaltrials.gov/ct2/show/NCT04881760) .”
“The limitations include the relatively small sample size of the MASLD substudy, the geographic and racial homogeneity of the sample (United States only and majority white), the exclusion of patients with T2D, the absence of liver histology, the lack of enrichment for MASH or significant fibrosis, the lack of multiplicity control given the large number of statistical assessments and the absence of 48-week MRI data for 56.1% of participants, which limits interpretation of the dose-response relationship for liver fat reduction at 48 weeks.”
“M.L.H., B.B., C.H., K.J.M., A.H., Q.W., Y.D. and N.C.S. are employees and shareholders of Eli Lilly and Company.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 48 references by DOI: 46 verified — 2 no DOI (shown, not verified).
- NO DOIThe neo-epitope specific PRO-C3 ELISA measures true formation of type III collagen associated with liver and muscle parametersNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIXGBoost: a scalable tree boosting systemNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT04881760LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttp://www.vivli.orgLIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
8 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 8 minor suggestions below.
8 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORtypoTable 2, Fasting insulin, CFB week 24, RETA 1 mg“18.9 (2. 5)”→ 18.9 (2.5)Extra space in the standard deviation.
- MINORtypoTable 2, Fasting insulin, CFB week 48, RETA 12 mg“−70. 9 (5.2)”→ −70.9 (5.2)Extra space in the value.
- MINORtypoTable 2, HOMA2-IR (C-peptide), Baseline, RETA 1 mg“1. 9 (0.2)”→ 1.9 (0.2)Extra space in the value.
- MINORconsistencyExtended Data Fig. 3 caption“Results shown as LS means ± SE.”→ Results shown as LSM ± s.e.m. for consistency with main text.Inconsistent abbreviation for least squares mean.
- MINORconsistencyExtended Data Fig. 4 caption“Results shown graphically as LS means ± 1.96 × SE.”→ Results shown as LSM ± 95% CI for consistency.Inconsistent presentation of error bars.
- MINORconsistencyAbstract“The mean relative change from baseline in LF at 24 weeks was −42.9% (1 mg), −57.0% (4 mg), −81.4% (8 mg), −82.4% (12 mg) and +0.3% (placebo) (all P < 0.001 versus placebo).”→ Ensure consistent use of 'LF' abbreviation; define it at first use.LF is defined in the abstract but not in the main text.
- MINORconsistencyTable 2“Data are LSMs (s.e.m.) or geometric mean (s.e.m.) from analysis of variance.”→ Clarify which rows are geometric means vs LSMs.The table mixes LSM and geometric mean without clear indication.
- MINORclarityMethods, Statistical analysis“No imputation was considered for missing data.”→ Clarify that this applies to continuous endpoints, as imputation is described for binary endpoints.Potential confusion with later imputation description.
The published work is robust and well-reported. An informed reader should weigh the minor reporting clarifications (e.g., missing data handling for continuous endpoints, software version details) and the copyedit issues, but none of these undermine the core findings. No erratum or re-analysis is warranted based on the current evidence.
- 1.MEDIUMreportingIn the Methods, Statistical analysis section, clarify that 'No imputation was considered for missing data' applies to continuous endpoints, and specify the imputation method for binary endpoints.The current wording is ambiguous and could confuse readers about the missing data handling for the primary continuous outcome.
- 2.MEDIUMreportingIn the Methods, specify the exact version and source of the AMRA Profiler Research software and the XGBoost algorithm.Providing version numbers improves resource identification and reproducibility.
- 3.MEDIUMreportingIn the Results, report the number of participants with liver fat data at each time point in the main text, not just in figure legends.This improves clarity and transparency regarding missing data at each visit.
- 4.MEDIUMreportingIn the Abstract, define the abbreviation 'LF' at first use and ensure consistent use throughout the main text.The abbreviation is used in the abstract but not defined in the main text, which could confuse readers.
- 5.MEDIUMreportingIn Table 2, clarify which rows are geometric means versus LSMs, as the table mixes both without clear indication.This will prevent misinterpretation of the reported values.
- 6.LOWcopyeditFix the extra spaces in Table 2 values: '18.9 (2. 5)' → '18.9 (2.5)', '−70. 9 (5.2)' → '−70.9 (5.2)', and '1. 9 (0.2)' → '1.9 (0.2)'.These are minor typos that should be corrected for professional presentation.
- 7.LOWcopyeditStandardize the abbreviation for least squares mean in Extended Data Fig. 3 and Fig. 4 captions to 'LSM ± s.e.m.' for consistency with the main text.Inconsistent abbreviations across figures can confuse readers.
- 8.LOWreportingIn the Introduction, consider adding a brief subsection explicitly discussing limitations of prior research and how this study addresses them.This would strengthen the scientific premise, as noted by one reviewer.
- 9.LOWreportingIn the Data Availability statement, consider adding a note about the availability of the statistical analysis code, even if not publicly shared.This would enhance transparency, as suggested by one reviewer.
- 10.LOWreportingIn the Discussion, consider adding a paragraph on the generalizability of the findings given the predominantly white, US-based sample and the exclusion of patients with T2D.This would address a potential limitation more explicitly.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.