Effects of semaglutide with and without concomitant SGLT2 inhibitor use in participants with type 2 diabetes and chronic kidney disease in the FLOW trial.
Mann JFE, Rossing P, Bakris G, Belmar N, Bosch-Traberg H, Busch R, Charytan DM, Hadjadj S, Gillard P, Górriz JL, Idorn T, Ji L, Mahaffey KW, Perkovic V, Rasmussen S, Schmieder RE, Pratley RE, Tuttle KR
- DOI
- 10.1038/s41591-024-03133-0
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/5a7a2773-b57e-449c-83fd-5c6db16a4e6b is authoritative.
How this rating was calculated
Started at 5★ — no deductions. Nothing the checks ran surfaced a material problem.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 8 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a rigorously reported prespecified subgroup analysis of a large, randomized, double-blind, placebo-controlled trial (FLOW). All eight rigor dimensions are adequately addressed, with clear reporting of design, ethics, statistical methods, data availability, and transparency. No copyedit issues or integrity concerns were identified.
Both reviewers independently scored all eight dimensions and agreed on every status; no divergence was present. The study type is interventional (randomized controlled trial subgroup analysis). Non-applicable sub-criteria (e.g., animal housing, cell line authentication) were excluded from scoring. The statistics verification recomputed only a subset of reported tests (6 tests, all consistent); other statistics were not machine-verified and should not be interpreted as confirmed.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 5 tests: 5 consistent, 0 inconsistent; 4 recomputed directly from the reported test statistics, 1 via agent-written checks.
- CONSISTENTreported p = .755 · recomputed p = .764Recomputed hazard ratio 1.07 (95% CI 0.69–1.67), reported p=0.755
“hazard ratio 1.07; 95% confidence interval: 0.69, 1.67; P = 0.755”
Taken as given: 0.69–1.67 is a two-sided 95% confidence interval for the hazard ratio of 1.07, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.755 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.07, 0.69, 1.67, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Recomputed HR 0.73 (95% CI 0.63–0.85), reported p<0.001
“HR 0.73; 95% CI: 0.63, 0.85; P < 0.001”
Taken as given: 0.63–0.85 is a two-sided 95% confidence interval for the HR of 0.73, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p<0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.73, 0.63, 0.85, 1) - CONSISTENTreported p = .532 · recomputed p = .527Recomputed HR 1.18 (95% CI 0.71–1.98), reported p=0.532
“HR 1.18; 95% CI: 0.71, 1.98; P = 0.532”
Taken as given: 0.71–1.98 is a two-sided 95% confidence interval for the HR of 1.18, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.532 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.18, 0.71, 1.98, 1) - CONSISTENTreported p = .003 · recomputed p = .004Recomputed HR 0.75 (95% CI 0.61–0.90), reported p=0.003
“HR 0.75; 95% CI: 0.61, 0.90; P = 0.003”
Taken as given: 0.61–0.90 is a two-sided 95% confidence interval for the HR of 0.75, not a range, an IQR, or a different interval level; the HR is a RATIO measure, so the interval is symmetric on the log scale; p=0.003 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.75, 0.61, 0.9, 1) - UNCOMPUTABLEreported p = .237 · recomputed p = .052Reviewer 1eGFR slope difference in SGLT2i subgroup: 0.75, 95% CI -0.01 to 1.50, p interaction 0.237
“Treatment differences favoring semaglutide for total estimated glomerular filtration rate slope (ml min −1 /1.73 m 2 /year) were 0.75 (−0.01, 1.5) in the SGLT2i subgroup”
Taken as given: The difference is on the linear scale.; The CI is two-sided at 95%.Method: Recomputed p from difference and 95% CI using normal approximation.How we recomputed it: pCI(0.75, -0.01, 1.50, 0) - CONSISTENTreported p = .003 · recomputed p = .004Reviewer 2Kidney-specific outcome HR in non-SGLT2i subgroup: p-value from HR and 95% CI
“HR 0.75; 95% CI: 0.61, 0.90; P = 0.003”
Taken as given: The HR is a ratio (log scale).; The CI is a 95% confidence interval.Method: Two-tailed p-value derived from the estimate and confidence interval assuming a normal distribution for the log HR.How we recomputed it: pCI(0.75, 0.61, 0.90, 1)
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
4 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2The benefits of semaglutide in reducing kidney outcomes were consistent in participants with/without baseline SGLT2i use.The primary outcome HRs and eGFR slope differences show no significant interaction, supporting consistency.Evidence: Primary outcome HR 1.07 (0.69-1.67) in SGLT2i subgroup vs 0.73 (0.63-0.85) in non-SGLT2i subgroup, P interaction 0.109; eGFR slope differences 0.75 and 1.25, P interaction 0.237.
“The benefits of semaglutide in reducing kidney outcomes were consistent in participants with/without baseline SGLT2i use”
AbstractFind in source - supportedReviewers 1, 2Semaglutide benefits on major CV events and all-cause death were similar regardless of SGLT2i use.Interaction p-values for MACE and all-cause death are high (0.741 and 0.901), indicating no significant heterogeneity.Evidence: P interaction 0.741 for MACE and 0.901 for all-cause death.
“Semaglutide benefits on major CV events and all-cause death were similar regardless of SGLT2i use ( P interaction 0.741 and 0.901, respectively).”
AbstractFind in source - supportedReviewers 1, 2The use of an SGLT2 inhibitor did not impact the overall benefits of semaglutide on kidney and cardiovascular outcomes.The overall trial showed benefit, and subgroup analyses showed no significant interaction, supporting the claim.Evidence: Overall primary outcome HR 0.76 (0.66-0.88) and subgroup interaction p-values all non-significant.
“In a prespecified analysis of the FLOW trial, the use of an SGLT2 inhibitor did not impact the overall benefits of semaglutide on kidney and cardiovascular outcomes”
AbstractFind in source - supportedReviewers 1, 2Power was limited to detect smaller but clinically relevant effects.The paper explicitly acknowledges limited power in the SGLT2i subgroup due to small sample size.Evidence: Discussion states 'power was limited due to the low use of SGLT2i at trial entry'.
“power was limited due to the low use of SGLT2i at trial entry”
DiscussionFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary efficacy claim is based on a hard clinical composite outcome (kidney failure, ≥50% eGFR reduction, kidney death, or CV death) and other clinical outcomes (MACE, all-cause death). Although eGFR slope and UACR are surrogate measures, they are secondary/supportive and the main claim rests on the hard composite. Target engagement is supported by the randomized controlled design and dose, and the surrogate (eGFR slope) is a validated marker of kidney disease progression.
“The primary outcome was a composite of kidney failure, ≥50% estimated glomerular filtration rate reduction, kidney death or CV death.”
- ADEQUATEEffect sizeThe primary effect is a 24% relative risk reduction in the composite outcome (HR 0.76, 95% CI 0.66-0.88), which is statistically significant and clinically meaningful. The eGFR slope difference of 1.16 ml/min/1.73m2/year is also presented as clinically relevant. The effect is anchored to hard clinical outcomes and the discussion explicitly states 'clinically meaningful benefit'.
“The risk of the primary outcome was 24% lower in all participants treated with semaglutide versus placebo (95% confidence interval: 34%, 12%).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites multiple trials and meta-analyses establishing the independent benefits of both drug classes and notes the lack of direct trials on their combination. The rationale for the prespecified subgroup analysis is clearly linked to the premise. Limitations of prior work (e.g., rare combined use in prior trials) are acknowledged.
“There are no clinical trials that have directly examined the combination of GLP-1 RAs plus SGLT2i on major kidney and CV outcomes in participants with T2D and CKD.”
“There are no clinical trials that have directly examined the combination of GLP-1 RAs plus SGLT2i on major kidney and CV outcomes in participants with T2D and CKD.”
Randomization method (central interactive web response system) and unit (participant) are reported. Blinding is described as double-blind. Power analysis is reported for the primary outcome (90% power to detect 20% RRR). Inclusion/exclusion criteria are detailed. Outlier handling is addressed via multiple imputation for missing data. Controls are inherent in the placebo comparator. Independent replication is not applicable for a single pivotal trial.
“participants with T2D and CKD were randomly assigned double-blind in a 1:1 ratio to receive semaglutide 1 mg per week subcutaneously or matching placebo using a central interactive web response system”
“Inclusion criteria: Informed consent obtained before any trial-related activities.”
“participants with T2D and CKD were randomly assigned double-blind in a 1:1 ratio to receive semaglutide 1 mg per week subcutaneously or matching placebo using a central interactive web response system”
“Inclusion criteria: Informed consent obtained before any trial-related activities.”
Sex is reported (percentage female). Age is reported (mean eGFR, median UACR, mean HbA1c). Demographics are described in terms of age, sex, and clinical parameters. Species/strain and housing are not applicable for a human trial.
“Those reporting SGLT2i use tended to be younger, less frequently female, with higher eGFR and lower systolic blood pressure.”
“randomized 3,533 participants (mean estimated glomerular filtration rate (eGFR) of 47.0 ml min −1 /1.73 m 2 , median urine albumin‐to-creatinine ratio (UACR) 568 mg g −1 , mean glycated hemoglobin (HbA 1c ) 7.8%)”
“Those reporting SGLT2i use tended to be younger, less frequently female, with higher eGFR and lower systolic blood pressure.”
“randomized 3,533 participants (mean estimated glomerular filtration rate (eGFR) of 47.0 ml min −1 /1.73 m 2 , median urine albumin‐to-creatinine ratio (UACR) 568 mg g −1 , mean glycated hemoglobin (HbA 1c ) 7.8%)”
The paper states that all participants provided written informed consent and the protocol was approved by national and institutional ethical and regulatory authorities. This satisfies both IRB approval and informed consent requirements. Regulatory compliance is implied by adherence to ethical standards.
“the protocol was approved by both national and institutional ethical and regulatory authorities.”
“All participants provided written informed consent”
“All participants provided written informed consent, and the protocol was approved by both national and institutional ethical and regulatory authorities.”
Semaglutide is named with dose (1.0 mg once weekly) and route (subcutaneous). The placebo is described as matching. Statistical software (SAS 9.4) is identified. No other biological/chemical resources are used.
“receive semaglutide 1 mg per week subcutaneously or matching placebo”
“All statistical analyses were performed with SAS software, version 9.4 (SAS Institute).”
“receive semaglutide 1 mg per week subcutaneously or matching placebo”
“All statistical analyses were performed with SAS software, version 9.4 (SAS Institute).”
Tests are named (Cox proportional hazards, ANCOVA, etc.). Assumptions are handled by design (e.g., stratified Cox). Exact p-values are reported for primary and secondary outcomes. Effect sizes with CIs are reported. Software is identified. Data presentation includes per-group n and CIs. Mathematical plausibility checks were not possible for all numbers, but no inconsistencies were found.
“Time-to-event endpoints were analyzed using a stratified Cox proportional hazards model with randomized treatment group (semaglutide or placebo) as a fixed factor.”
“HR 0.73; 95% CI: 0.63, 0.85; P < 0.001”
“Time-to-event endpoints were analyzed using a stratified Cox proportional hazards model with randomized treatment group (semaglutide or placebo) as a fixed factor.”
“HR 0.73; 95% CI: 0.63, 0.85; P < 0.001”
The data availability statement specifies that data will be shared with bona fide researchers who submit a research proposal approved by an independent review board, with de-identified data, and provides a URL for access requests. This meets the standard for clinical trial data.
“Data will be shared with bona fide researchers who submit a research proposal approved by the independent review board. Individual participant data will be shared in datasets in a de-identified and anonymized format. Data will be made available after research completion and approval of the product and product use in the European Union and the United States. Information about data access request proposals can be found at https://www.novonordisk-trials.com/ .”
“Data will be shared with bona fide researchers who submit a research proposal approved by the independent review board. Individual participant data will be shared in datasets in a de-identified and anonymized format.”
Trial registration number is provided (NCT03819153). Funding and COI are disclosed. Limitations are discussed, including limited power in the SGLT2i subgroup. Conclusions are appropriately cautious, noting the inability to detect small interactions. Reporting guideline adherence is implied by the journal's reporting summary.
“ClinicalTrials.gov identifier: NCT03819153”
“Our analysis has limitations, mainly the limited power in the groups with baseline SGLT2i use.”
“This study was funded by Novo Nordisk A/S”
“ClinicalTrials.gov identifier: NCT03819153”
“Our analysis has limitations, mainly the limited power in the groups with baseline SGLT2i use.”
“This study was funded by Novo Nordisk A/S”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 23 references by DOI: 22 verified — 1 no DOI (shown, not verified).
- NO DOIKDIGO 2012 clinical practice guideline for the evaluation and management of chronic kidney diseaseNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://www.novonordisk-trials.com/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/study/NCT03819153LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
None foundWording, consistency and formatting errors that need correcting before submission.
Checked — nothing surfaced.
No copyedit issues found.
The published work is robust and well-reported; an informed reader should weigh the limited power in the SGLT2i subgroup (acknowledged by the authors) and the fact that only a subset of statistics was independently recomputed. No erratum or correction appears warranted based on this audit.
- 1.MEDIUMreportingIn the Statistical analysis section, add a sentence explicitly stating that subgroup analyses were not adjusted for multiplicity and that results should be interpreted as exploratory.Reviewer 2 noted that the handling of multiplicity in subgroup analyses is not described; adding this clarifies the inferential status of the subgroup findings.
- 2.MEDIUMreportingIn the Methods or Acknowledgements, add an explicit statement of adherence to the CONSORT reporting guideline for this subgroup analysis.Reviewer 2 suggested this to strengthen reporting transparency, even though the journal's reporting summary is referenced.
- 3.MEDIUMdata codeIn the Data availability section, add a statement about the availability of the statistical analysis code, if any, or explicitly state that code is not shared.Reviewer 1 suggested this to enhance reproducibility; even if code is not shared, stating so is transparent.
- 4.LOWreportingIn the Results section, report the number of participants with missing data for each outcome and the imputation details more explicitly in the main text.Reviewer 2 suggested this to improve transparency about missing data handling, which is currently only briefly mentioned.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.