Orforglipron for maintenance of body weight reduction: the double-blind, randomized phase 3b ATTAIN-MAINTAIN trial.
Aronne LJ, Horn DB, le Roux CW, Chao AM, Ho W, Halpern B, Griffin R, Xie C, Valderas EG, Lee CJ, Ribeiro A, Hyman DM, Glass L, Xavier N
- DOI
- 10.1038/s41591-026-04386-7
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/9ee8793f-7aa6-4fc1-afd8-bff4a6061976 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsOverstated claim−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is percent maintenance of body weight reduction, which is a surrogate for long-term health outcomes. The paper does not provide evidence linking this surrogate to hard clinical outcomes, nor does it demonstrate target engagement at the tested dose beyond stating that orforglipron is a GLP-1 receptor agonist.
“the primary endpoint was the percent maintenance of body weight reduction in SURMOUNT-5 for participants who achieved a body weight plateau.”
- 02Treatment effect not shown to be clinically meaningful
The effect sizes are reported as percentages of weight maintenance (e.g., 74.7% vs 49.2% in cohort 1), but there is no anchor to a minimal clinically important difference or to long-term health benefits. The absolute weight changes are small (approximately 5 kg in cohort 1 and 1 kg in cohort 2) and may not be clinically meaningful.
“Participants treated with orforglipron demonstrated average reductions in body weight from week 0 to 52 in ATTAIN-MAINTAIN of approximately 5 kg (5%) in cohort 1 and 1 kg (1%) in cohort 2.”
- 03Conclusion reaches beyond the evidence
Switching to orforglipron results in both cohorts ending at the same body weight of 95.9 kg, suggesting a potential biological control of obesity.
“An additional finding is that, regardless of the initial intervention, switching to orforglipron resulted in both cohorts ending at the same body weight of 95.9 kg. This finding with orforglipron could suggest that there may be a biological component that…”
Discussion ¶4Find in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported phase 3b randomized controlled trial. The manuscript demonstrates strong methodological rigor in design, statistical analysis, and transparency, with minor reporting gaps such as lack of explicit reporting guideline adherence and minor copyedit issues.
Both reviewers classified the study as interventional, and this was adopted. The evaluation covered all eight dimensions; several sub-criteria were marked not applicable due to the human clinical trial context (e.g., species/strain, housing, IACUC, cell line authentication). The statistics verification component recomputed only 3 tests (all consistent), so the broader statistical analysis remains unverified. The citation check found no retracted or unresolved references.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary endpoint treatment difference in cohort 1 (orforglipron vs placebo) p-value from CI.
“resulting in an estimated treatment difference of MBE 25.5% (95% confidence interval 14.5 to 36.5); P < 0.001; treatment-regimen estimand) at week 52.”
Taken as given: The estimate is 25.5 and the 95% CI is 14.5 to 36.5.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed two-tailed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(25.5, 14.5, 36.5, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Primary endpoint treatment difference in cohort 2 (orforglipron vs placebo) p-value from CI.
“resulting in an estimated treatment difference of MBE 41.7 (95% confidence interval 24.4 to 59.0); P < 0.001; treatment-regimen estimand) at week 52.”
Taken as given: The estimate is 41.7 and the 95% CI is 24.4 to 59.0.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed two-tailed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(41.7, 24.4, 59.0, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 2Primary endpoint for cohort 1: estimated treatment difference of 25.5% with 95% CI 14.5 to 36.5, P < 0.001. The p-value is derived from an ANCOVA model; the CI does not include zero, consistent with P < 0.001.
“resulting in an estimated treatment difference of MBE 25.5% (95% confidence interval 14.5 to 36.5); P < 0.001”
Taken as given: The estimate is 25.5, the lower bound is 14.5, the upper bound is 36.5.; The CI is two-sided at 95%.; The p-value is from the same ANCOVA model that produced the CI.Method: pCI function: two-tailed p from estimate and 95% CI assuming normal approximation.How we recomputed it: pCI(25.5, 14.5, 36.5, 0)
- lowinternal contradictionIn the abstract, cohort 2 orforglipron group N is 105, but in the safety table (Table 2) the orforglipron N is 105, consistent. However, the total N in Table 2 for cohort 2 is 171, which matches the randomized total.
Cohort 2: N = 171 (N = 105 orforglipron 36 mg or MTD, N = 66 placebo) ... Table 2: Orforglipron N = 105
Table 2reviewer’s wording - lowinternal contradictionIn the abstract, cohort 1 orforglipron group N is 125, but in the safety table (Table 2) the orforglipron N is 124. This may be due to a participant excluded from the safety analysis set.
Cohort 1: N = 205 (N = 125 orforglipron 36 mg or maximum tolerated dose (MTD), N = 80 placebo) ... Table 2: Orforglipron N = 124
Table 2reviewer’s wording
Overstated conclusions
4 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated), 2 only partially supported (evidence backs part of the claim; gaps or caveats remain).
- overstatedReviewer 2Switching to orforglipron results in both cohorts ending at the same body weight of 95.9 kg, suggesting a potential biological control of obesity.The observation that both cohorts reached the same mean weight is interesting but speculative; the paper itself acknowledges this is a hypothesis requiring further study, and the claim is presented as a suggestion rather than a definitive conclusion.Evidence: The text reports that both cohorts had a mean body weight of 95.9 kg at week 52 with orforglipron, but this is a single observation from a non-randomized comparison across cohorts.
“An additional finding is that, regardless of the initial intervention, switching to orforglipron resulted in both cohorts ending at the same body weight of 95.9 kg. This finding with orforglipron could suggest that there may be a biological component that controls the disease of obesity at this body weight. Further studies are needed to explore this potential biological control of obesity.”
Discussion ¶4Find in source - partialReviewers 1, 2Orforglipron is a globally scalable option for minimizing weight changes after injectable therapy.The efficacy and safety data support the potential, but the claim of global scalability is an inference based on the oral formulation and not directly tested in this trial.Evidence: The paper discusses the oral nonpeptide formulation and its potential to overcome barriers to injectable therapy, but no direct evidence of global scalability is presented.
“These data demonstrate orforglipron’s potential as a globally scalable option for minimizing weight changes after injectable therapy.”
AbstractFind in source - partialReviewer 2Orforglipron preserves cardiometabolic benefits (e.g., waist circumference, blood pressure, lipids) achieved with prior injectable therapy.The paper presents descriptive data on cardiometabolic parameters (e.g., HbA1c, waist circumference) showing preservation, but these are exploratory endpoints without formal statistical testing for all parameters.Evidence: Extended Data Figures 3 and 4 show trends for lipids, blood pressure, and glycemic parameters. The text states 'similar preservation of reductions' but does not provide p-values for these comparisons.
“In addition, other cardiometabolic risk factors demonstrated similar preservation of reductions at the end of ATTAIN-MAINTAIN. As an example, in cohort 1, participants subsequently randomized to orforglipron had a mean baseline HbA1c of 5.6% at the beginning of SURMOUNT-5. At the beginning of ATTAIN-MAINTAIN, after weight reduction with injectables, the HbA1c improved to a mean of 5.2%. At 52 weeks, the mean HbA1c remained at 5.2%”
Figure 3Find in source - supportedReviewer 1Orforglipron maintains a significantly higher percentage of body weight reduction compared with placebo in both cohorts.The primary endpoint results show a statistically significant and clinically meaningful difference favoring orforglipron in both cohorts.Evidence: Primary endpoint: cohort 1 MBE 74.7% vs 49.2%, ETD 25.5% (95% CI 14.5-36.5), P<0.001; cohort 2 MBE 79.3% vs 37.6%, ETD 41.7 (95% CI 24.4-59.0), P<0.001.
“Cohort 1 participants who achieved body weight plateau maintained a model-based estimate (MBE) of 74.7% (s.e.m. 4.05) of body weight reduction with orforglipron compared with an MBE of 49.2% (s.e.m. 3.92) with placebo, resulting in an estimated treatment difference of MBE 25.5% (95% confidence interval 14.5 to 36.5); P < 0.001; treatment-regimen estimand) at week 52.”
AbstractFind in source - supportedReviewer 1All key secondary endpoints were met.The paper reports that key secondary endpoints, including maintenance of ≥80% weight reduction, were met with statistical significance.Evidence: Key secondary endpoints: cohort 1 43.7% vs 16.4% maintained ≥80% weight reduction, risk difference 27.3 (95% CI 14.1-40.6), P<0.001; cohort 2 55.0% vs 6.9%, risk difference 48.1 (95% CI 35.6-60.5), P<0.001.
“All key secondary endpoints were met.”
AbstractFind in source - supportedReviewer 1Switching to orforglipron preserves cardiometabolic benefits achieved with injectable therapy.The paper reports sustained improvements in waist circumference, blood pressure, lipids, and glycemic parameters, supporting the claim.Evidence: Results show preservation of reductions in waist circumference, HbA1c, lipids, and blood pressure at week 52.
“In addition, other cardiometabolic risk factors demonstrated similar preservation of reductions at the end of ATTAIN-MAINTAIN.”
ResultsFind in source - supportedReviewer 1The trial's limitations include the absence of a comparator arm involving continued use of injectable obesity-management medications and the trial's 1-year duration.The limitations are explicitly acknowledged in the discussion.Evidence: The discussion states these limitations.
“Trial limitations include the absence of a comparator arm involving continued use of injectable obesity-management medications and the trial’s 1-year duration.”
DiscussionFind in source - supportedReviewer 2Orforglipron maintains a significantly greater proportion of body weight reduction compared with placebo in participants previously treated with tirzepatide or semaglutide.The primary endpoint results for both cohorts show statistically significant and clinically meaningful differences favoring orforglipron over placebo.Evidence: Primary endpoint: cohort 1 ETD 25.5% (95% CI 14.5 to 36.5; P < 0.001); cohort 2 ETD 41.7% (95% CI 24.4 to 59.0; P < 0.001).
“Cohort 1 participants who achieved body weight plateau maintained a model-based estimate (MBE) of 74.7% (s.e.m. 4.05) of body weight reduction with orforglipron compared with an MBE of 49.2% (s.e.m. 3.92) with placebo, resulting in an estimated treatment difference of MBE 25.5% (95% confidence interval 14.5 to 36.5); P < 0.001”
AbstractFind in source - supportedReviewer 2Orforglipron is generally well-tolerated with a safety profile similar to injectable GLP-1 receptor agonists.Safety data show that adverse events were mostly mild to moderate gastrointestinal effects, with low discontinuation rates, consistent with the known profile of GLP-1 RAs.Evidence: Table 2: AEs leading to discontinuation were 7.3% (cohort 1) and 4.8% (cohort 2) with orforglipron. Most GI AEs were mild to moderate.
“The most frequently reported AEs with orforglipron were gastrointestinal disorders of nausea, constipation, vomiting or diarrhea. The overall incidence of GI AEs including nausea, vomiting or diarrhea in the first 4 weeks of the trial for both cohorts was 10.5% and 9.5% in the tirzepatide and semaglutide cohorts, respectively. Most GI AEs were mild to moderate in severity”
Table 2Find in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is percent maintenance of body weight reduction, which is a surrogate for long-term health outcomes. The paper does not provide evidence linking this surrogate to hard clinical outcomes, nor does it demonstrate target engagement at the tested dose beyond stating that orforglipron is a GLP-1 receptor agonist.
“the primary endpoint was the percent maintenance of body weight reduction in SURMOUNT-5 for participants who achieved a body weight plateau.”
- INADEQUATEEffect sizeThe effect sizes are reported as percentages of weight maintenance (e.g., 74.7% vs 49.2% in cohort 1), but there is no anchor to a minimal clinically important difference or to long-term health benefits. The absolute weight changes are small (approximately 5 kg in cohort 1 and 1 kg in cohort 2) and may not be clinically meaningful.
“Participants treated with orforglipron demonstrated average reductions in body weight from week 0 to 52 in ATTAIN-MAINTAIN of approximately 5 kg (5%) in cohort 1 and 1 kg (1%) in cohort 2.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on the chronic nature of obesity, the need for sustained therapy, and the efficacy of orforglipron in the ATTAIN program. It acknowledges limitations of prior work, such as the lack of studies on switching from injectable to oral therapy. The rationale logically links these gaps to the study objectives.
“Incretins have improved the management of obesity and its related complications, but maintaining these health benefits requires ongoing administration, which can be challenging.”
“This is the initial clinical trial investigating the switch from injectable incretin-based therapy to an oral OMM for the maintenance of body weight reduction, and as a result, the study evaluated various endpoints, acknowledging that the clinical relevance of these endpoints may not be appreciated until trial completion.”
Randomization method is described (computer-generated random sequence via interactive web response system), with stratification factors. Blinding is clearly stated for investigators, site staff, clinical monitors, and participants. A priori power analysis is provided with effect size, alpha, and power. Inclusion/exclusion criteria are pre-specified, and the analysis population (modified ITT) is defined. Outlier handling is addressed through the use of estimands and imputation for missing data. Controls are appropriate (placebo comparator). Independent replication is not applicable for a single pivotal trial.
“Participants were randomized in a 3:2 ratio to receive once-daily orforglipron (36 mg or MTD of 24 or 36 mg), or placebo, using a computer-generated random sequence from a Lilly interactive web response system.”
“Study investigators, site staff, clinical monitors and participants were blinded to the study intervention until study completion.”
“We estimated that a sample size of 150 participants was required for each cohort to ensure that 118 participants reached a plateau in body weight. This sample size provides approximately 90% power to detect a 10% treatment difference for the primary endpoint.”
“Participants were randomized in a 3:2 ratio to receive once-daily orforglipron (36 mg or MTD of 24 or 36 mg), or placebo, using a computer-generated random sequence from a Lilly interactive web response system.”
“Study investigators, site staff, clinical monitors and participants were blinded to the study intervention until study completion.”
“We estimated that a sample size of 150 participants was required for each cohort to ensure that 118 participants reached a plateau in body weight. This sample size provides approximately 90% power to detect a 10% treatment difference for the primary endpoint.”
Table 1 provides detailed demographics: age, sex, race, body weight, BMI, waist circumference, blood pressure, lipid parameters, and comorbidities. Both sexes are enrolled (62.9% female in cohort 1, 68.4% in cohort 2), so no single-sex justification is needed. Age and health status (e.g., prediabetes, HbA1c) are reported. Species/strain and housing conditions are not applicable for a human trial.
“Female sex, n (%) | 79 (63.2) | 50 (62.5) | 129 (62.9) | 73 (69.5) | 44 (66.7) | 117 (68.4)”
“Age, years | 49.0 (12.6) | 47.7 (11.8) | 48.5 (12.2) | 48.6 (11.8) | 48.7 (13.5) | 48.6 (12.4)”
“Female sex, n (%) | 79 (63.2) | 50 (62.5) | 129 (62.9) | 73 (69.5) | 44 (66.7) | 117 (68.4)”
“Age, years | 49.0 (12.6) | 47.7 (11.8) | 48.5 (12.2) | 48.6 (11.8) | 48.7 (13.5) | 48.6 (12.4)”
“White | 94 (76.4) | 55 (69.6) | 149 (73.8) | 74 (71.2) | 55 (83.3) | 129 (75.9)”
The trial protocol was approved by the Advarra central institutional review board, and participants provided written informed consent. Compliance with the Declaration of Helsinki and ICH-GCP is stated. The trial is registered on ClinicalTrials.gov.
“The trial protocol was approved by the Advarra central institutional review board.”
“All participants provided signed informed consent before trial participation.”
“The trial was conducted in compliance with the Declaration of Helsinki and the Good Clinical Practice guidelines of the International Council for Harmonization.”
“The trial protocol was approved by the Advarra central institutional review board.”
“All participants provided signed informed consent before trial participation.”
“The trial was conducted in compliance with the Declaration of Helsinki and the Good Clinical Practice guidelines of the International Council for Harmonization.”
Orforglipron is described as a once-daily oral nonpeptide GLP-1 receptor agonist, with doses specified (1, 3, 6, 12, 24, 36 mg capsules, with tablet equivalents). The placebo is matching. Statistical software is identified as R version 4.4.2. No antibodies, cell lines, or organisms are used, so those criteria are not applicable.
“Statistical analyses were computed using R version 4.4.2.”
“Statistical analyses were computed using R version 4.4.2.”
The primary analysis uses ANCOVA with model-based estimates, and p-values are reported as exact values (e.g., P < 0.001). Effect sizes are reported with 95% confidence intervals. The software (R 4.4.2) is identified. Data presentation includes per-group n and s.e.m. in figures. Assumptions verification is not explicitly discussed, but the use of MMRM and ANCOVA with stratification factors is standard for this type of trial. Mathematical plausibility is not applicable as the primary endpoint is a model-based estimate, not raw integer data.
“P < 0.001; treatment-regimen estimand) at week 52.”
“estimated treatment difference of MBE 25.5% (95% confidence interval 14.5 to 36.5)”
“resulting in an estimated treatment difference of MBE 25.5% (95% confidence interval 14.5 to 36.5); P < 0.001”
“resulting in an estimated treatment difference of MBE 25.5% (95% confidence interval 14.5 to 36.5)”
“Statistical analyses were computed using R version 4.4.2.”
The data availability statement provides a concrete mechanism for accessing individual participant data through Vivli, with conditions and timeframe. Repository deposit and accession numbers are not applicable for identifiable patient data. Code sharing is not applicable as no bespoke code is mentioned.
“Lilly provides access to all individual participant data collected during the trial, after anonymization, with the exception of pharmacokinetic or genetic data. Data are available to request 6 months after the indication studied has been approved in the USA and European Union and after primary publication acceptance, whichever is later. No expiration date of data requests is currently set once data are made available. Access is provided after a proposal has been approved by an independent review committee identified for this purpose and after receipt of a signed data sharing agreement. Data and documents, including the study protocol, statistical analysis plan, clinical study report, blank or annotated case report forms, will be provided in a secure data sharing environment. For details on submitting a request, see the instructions provided at www.vivli.org (http://www.vivli.org/) .”
Methods are detailed enough for replication, including dosing, randomization, and statistical analysis. The trial is registered (NCT06584916). Limitations are discussed (e.g., US-only sites, predominantly white population, lack of injectable comparator, 1-year duration). Conclusions are proportional to the evidence. Funding (Eli Lilly) and competing interests are disclosed. A CONSORT diagram is provided, but the paper does not explicitly state adherence to a reporting guideline.
“ClinicalTrials.gov registration: NCT06584916 (https://clinicaltrials.gov/study/NCT06584916) .”
“Trial limitations include study sites located only in the USA and a predominantly white study population, although the study includes a more representative US population with a higher Black or African American and Hispanic or Latino population. Additional limitations were the lack of a comparator arm that included continuing injectable OMM and a trial duration of 1 year.”
“ClinicalTrials.gov registration: NCT06584916 (https://clinicaltrials.gov/study/NCT06584916)”
“Trial limitations include study sites located only in the USA and a predominantly white study population, although the study includes a more representative US population with a higher Black or African American and Hispanic or Latino population. Additional limitations were the lack of a comparator arm that included continuing injectable OMM and a trial duration of 1 year.”
“This study was funded by Eli Lilly and Company.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 16 references by DOI: 16 verified.
Every extracted reference resolved against Crossref/OpenAlex with no retraction flags.
2 data/code links checked; 2 live.
- datahttps://clinicaltrials.gov/study/NCT06584916LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttp://www.vivli.org/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
8 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 8 minor suggestions below.
8 copyedit issues flagged: mostly consistency, clarity.
- MINORconsistencyAbstract“resulting in an estimated treatment difference of MBE 25.5% (95% confidence interval 14.5 to 36.5); P < 0.001; treatment-regimen estimand) at week 52.”→ Remove the extra closing parenthesis after '36.5'.Extra parenthesis in the abstract.
- MINORconsistencyResults, Cohort 2“resulting in an estimated treatment difference of MBE 41.7 (95% confidence interval 24.4 to 59.0); P < 0.001; treatment-regimen estimand) at week 52.”→ Add '%' after 41.7 for consistency with cohort 1.Missing percentage sign.
- MINORclarityDiscussion, paragraph 2“This preserved the previously achieved weight reduction with an average difference of approximately 3 kg.”→ Clarify what 'this' refers to.Ambiguous pronoun.
- MINORconsistencyAbstract, Results“resulting in an estimated treatment difference of MBE 25.5% (95% confidence interval 14.5 to 36.5); P < 0.001; treatment-regimen estimand) at week 52.”→ Remove the extra closing parenthesis after '36.5' or ensure consistent punctuation.The parenthesis structure is slightly off; likely a typo.
- MINORconsistencyResults, Cohort 1“resulting in an estimated treatment difference of MBE 25.5% (95% confidence interval 14.5 to 36.5); P < 0.001; treatment-regimen estimand) at week 52.”→ Consider rewriting as '...resulting in an estimated treatment difference of 25.5% (95% CI 14.5 to 36.5; P < 0.001; treatment-regimen estimand) at week 52.'The semicolon before 'treatment-regimen estimand' is unusual; a comma or restructuring would improve clarity.
- MINORclarityResults, Cohort 1“All participants in cohort 1 had a mean MBE percent change from SURMOUNT-5 baseline body weight of –16.5% (s.e.m. 0.90) with orforglipron and 12.6% (s.e.m. 0.91) with placebo with an ETD relative to placebo of –3.9 percentage points (95% CI –6.1 to –1.8; P < 0.001).”→ The placebo value of 12.6% appears to be a positive number (weight gain) but is not clearly labeled as such; consider adding 'weight gain' or clarifying the direction.The context suggests placebo participants regained weight, but the sign is ambiguous without careful reading.
- MINORconsistencyResults, Cohort 2“All participants in cohort 2 had a mean MBE percent change from SURMOUNT-5 baseline body weight of –14.9% (s.e.m. 0.94) with orforglipron and –7.9% (SE 0.81) with placebo”→ Use consistent abbreviation for standard error (s.e.m. vs SE) throughout the manuscript.In the same sentence, 's.e.m.' is used for orforglipron and 'SE' for placebo; should be consistent.
- MINORclarityDiscussion, paragraph 2“Mean absolute change in weight may be an important complementary endpoint to facilitate a clearer and fuller understanding of maintenance for the clinical provider.”→ Consider rephrasing for clarity: 'Mean absolute change in weight may serve as an important complementary endpoint to provide clinicians with a clearer understanding of weight maintenance.'The sentence is grammatically correct but slightly awkward.
The published work is robust and well-reported. An informed reader should weigh the minor reporting gaps (e.g., explicit CONSORT adherence, statistical assumption verification) and the copyedit issues, but none threaten the validity of the conclusions. No erratum is warranted based on the current evidence.
- 1.HIGHreportingAdd an explicit statement of adherence to CONSORT reporting guidelines in the Methods section, and consider submitting the CONSORT checklist as supplementary material.Both reviewers noted that the reporting guideline is not explicitly named, which is a transparency gap for a randomized trial.
- 2.HIGHstatisticsAdd a brief discussion of statistical assumption verification (e.g., normality, homogeneity of variance) for the ANCOVA/MMRM models, or note that the methods are robust to violations.Reviewer 2 flagged that assumptions verification is not explicitly discussed, which is a common reviewer concern.
- 3.MEDIUMcopyeditFix the extra closing parenthesis in the abstract and Results section: '...MBE 25.5% (95% confidence interval 14.5 to 36.5); P < 0.001; treatment-regimen estimand) at week 52.' should be '...MBE 25.5% (95% CI 14.5 to 36.5; P < 0.001; treatment-regimen estimand) at week 52.'The copyedit pass flagged this as a punctuation error that could confuse readers.
- 4.MEDIUMcopyeditAdd '%' after 41.7 in the Cohort 2 result: '...MBE 41.7 (95% confidence interval 24.4 to 59.0)' should be '...MBE 41.7% (95% CI 24.4 to 59.0)'.The copyedit pass noted the missing percentage sign, which is inconsistent with Cohort 1 reporting.
- 5.MEDIUMcopyeditClarify the ambiguous pronoun in Discussion paragraph 2: 'This preserved the previously achieved weight reduction...' — specify what 'this' refers to.The copyedit pass flagged this as a clarity issue that could hinder understanding.
- 6.MEDIUMcopyeditClarify the direction of weight change for the placebo group in Cohort 1: '...and 12.6% (s.e.m. 0.91) with placebo' — add 'weight gain' or clarify the sign.The copyedit pass noted that the positive value is ambiguous without context.
- 7.MEDIUMcopyeditStandardize the abbreviation for standard error: use either 's.e.m.' or 'SE' consistently throughout the manuscript, especially in Results Cohort 2.The copyedit pass flagged inconsistent use of 's.e.m.' and 'SE' in the same sentence.
- 8.MEDIUMcopyeditRephrase the sentence in Discussion paragraph 2: 'Mean absolute change in weight may be an important complementary endpoint to facilitate a clearer and fuller understanding of maintenance for the clinical provider.' to improve clarity.The copyedit pass noted the sentence is grammatically correct but awkward.
- 9.LOWreportingConsider adding a more detailed discussion of how limitations of prior research were addressed in the introduction.Reviewer 1 suggested this to strengthen the scientific premise, though Reviewer 2 found it adequate.
- 10.LOWreportingClarify the definition of 'plateau' and its implications for the primary endpoint in the Methods.Reviewer 1 suggested this to improve methodological clarity.
- 11.LOWreportingConsider reporting the number of participants who discontinued due to adverse events in the abstract for completeness.Reviewer 1 suggested this to enhance the abstract's completeness.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.