Effect of dapagliflozin on metabolic dysfunction-associated steatohepatitis: multicentre, double blind, randomised, placebo controlled trial.
Lin J, Huang Y, Xu B, Gu X, Huang J, Sun J, Jia L, He J, Huang C, Wei X, Chen J, Chen X, Zhou J, Wu L, Zhang P, Zhu Y, Xia H, Wen G, Liu Y, Liu S, Zeng Y, Zhou L, Jia H, He H, Xue Y, Wu F, Zhang H
- DOI
- 10.1136/bmj-2024-083735
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/046a4c9c-c453-454b-ba79-22b146332527 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 49 reported means were read, and their group size is not stated where the values are printed. These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is a histological surrogate (NAS-based MASH improvement) rather than a hard clinical outcome. Although the paper cites FDA guidance that such endpoints are 'reasonably likely to predict long term clinical benefit', it does not provide validated evidence linking the specific surrogate (NAS change) to clinical outcomes such as cirrhosis or mortality. Target engagement at the tested dose is not demonstrated (no PK/PD data).
“The primary endpoint was MASH improvement (defined as a decrease of at least 2 points in non-alcoholic fatty liver disease activity score (NAS) or a NAS of ≤3 points) without worsening of liver fibrosis (defined as without increase of fibrosis stage) at 48…”
- 02Treatment effect not shown to be clinically meaningful
The primary effect is reported as a risk ratio of 1.73 (53% vs 30%) for MASH improvement, but the absolute difference is 23%. The clinical meaningfulness is not anchored to a minimal clinically important difference or a hard outcome. The effect on NAS is a mean difference of -1.39, but the clinical significance of this change is not established.
“MASH improvement without worsening of fibrosis was reported in 53% (41/78) of participants in the dapagliflozin group and 30% (23/76) in the placebo group (risk ratio 1.73 (95% confidence interval (CI) 1.16 to 2.58); P=0.006).”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported multicentre RCT of dapagliflozin in MASH. The trial design is rigorous, statistical methods are appropriate, and reporting is transparent with only minor gaps (no explicit CONSORT mention, a possible group-label swap in one sentence, and a minor typo).
Both reviewers classified the study as interventional, which is adopted. The evaluation covers all eight dimensions; several sub-criteria were marked not applicable (e.g., cell line authentication, housing conditions) as this is a human clinical trial. The reviewers agreed on all dimensions, so no divergence to resolve.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 7 tests: 7 consistent, 0 inconsistent; 3 recomputed directly from the reported test statistics, 4 via agent-written checks.
- CONSISTENTreported p = .006 · recomputed p = .007Recomputed risk ratio 1.73 (95% CI 1.16–2.58), reported p=0.006
“risk ratio 1.73 (95% confidence interval (CI) 1.16 to 2.58); P=0.006”
Taken as given: 1.16–2.58 is a two-sided 95% confidence interval for the risk ratio of 1.73, not a range, an IQR, or a different interval level; the risk ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.006 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.73, 1.16, 2.58, 1) - CONSISTENTreported p = .010 · recomputed p = .016Recomputed risk ratio 2.91 (95% CI 1.22–6.97), reported p=0.01
“risk ratio 2.91 (95% CI 1.22 to 6.97); P=0.01”
Taken as given: 1.22–6.97 is a two-sided 95% confidence interval for the risk ratio of 2.91, not a range, an IQR, or a different interval level; the risk ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.01 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.91, 1.22, 6.97, 1) - CONSISTENTreported p = .001 · recomputed p = .002Recomputed risk ratio 2.25 (95% CI 1.35–3.75), reported p=0.001
“risk ratio 2.25 (95% CI 1.35 to 3.75); P=0.001”
Taken as given: 1.35–3.75 is a two-sided 95% confidence interval for the risk ratio of 2.25, not a range, an IQR, or a different interval level; the risk ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.25, 1.35, 3.75, 1) - CONSISTENTreported p = .006 · recomputed p = .005Reviewer 1Primary endpoint risk ratio p-value
“MASH improvement without worsening of fibrosis was reported in 53% (41/78) of participants in the dapagliflozin group and 30% (23/76) in the placebo group (risk ratio 1.73 (95% confidence interval (CI) 1.16 to 2.58); P=0.006).”
Taken as given: The 41 and 23 are the event counts in the dapagliflozin and placebo groups, respectively.; The 78 and 76 are the group totals.; The p-value is from a chi-square test (or CMH) on the 2x2 table.; The test is two-sided.Method: Pearson chi-square test on the 2x2 table (41,37,23,53) using pChi2x2.How we recomputed it: pChi2x2(41, 37, 23, 53) - CONSISTENTreported p = .010 · recomputed p = .009Reviewers 1, 2MASH resolution secondary endpoint p-value
“MASH resolution without worsening of fibrosis occurred in 23% (18/78) of participants in the dapagliflozin group and 8% (6/76) in the placebo group (risk ratio 2.91 (95% CI 1.22 to 6.97); P=0.01).”
Taken as given: The 18 and 6 are the event counts.; The 78 and 76 are the group totals.; The p-value is from a chi-square test (or CMH).; The test is two-sided.Method: Pearson chi-square test on the 2x2 table (18,60,6,70) using pChi2x2.How we recomputed it: pChi2x2(18, 60, 6, 70) - CONSISTENTreported p = .001 · recomputed p = <.001Reviewers 1, 2Fibrosis improvement secondary endpoint p-value
“Fibrosis improvement without worsening of MASH was reported in 45% (35/78) of participants in the dapagliflozin group, as compared with 20% (15/76) in the placebo group (risk ratio 2.25 (95% CI 1.35 to 3.75); P=0.001).”
Taken as given: The 35 and 15 are the event counts.; The 78 and 76 are the group totals.; The p-value is from a chi-square test (or CMH).; The test is two-sided.Method: Pearson chi-square test on the 2x2 table (35,43,15,61) using pChi2x2.How we recomputed it: pChi2x2(35, 43, 15, 61) - CONSISTENTreported p = .006 · recomputed p = .005Reviewer 2Primary endpoint: MASH improvement without worsening of fibrosis (dapagliflozin vs placebo) using Cochran-Mantel-Haenszel test stratified by diabetes status.
“MASH improvement without worsening of fibrosis was reported in 53% (41/78) of participants in the dapagliflozin group and 30% (23/76) in the placebo group (risk ratio 1.73 (95% confidence interval (CI) 1.16 to 2.58); P=0.006).”
Taken as given: The 41 and 23 are the event counts in the dapagliflozin and placebo groups, respectively.; The 78 and 76 are the total numbers in each group.; The test is two-sided and uses a chi-square approximation to the CMH test.; The stratification by diabetes status is ignored in this approximation.Method: Pearson chi-square test on the 2x2 table (41,37,23,53) to approximate the CMH p-value.How we recomputed it: pChi2x2(41, 37, 23, 53)
- lowinternal contradictionIn the Results section, the sentence 'worsening of fibrosis occurred in 5% (4/76) in the dapagliflozin group, and 22% (17/78) in the placebo group' appears to have the group labels swapped, as the dapagliflozin group has 78 participants and placebo has 76.
“Among all trial participants, worsening of fibrosis occurred in 5% (4/76) in the dapagliflozin group, and 22% (17/78) in the placebo group at week 48”
ResultsFind in source - lowinternal contradictionThe percentages for fibrosis stages in the Results text (33%, 45%, 19%) sum to 97%, not 100%, and do not include F0 or F4, which are reported in Table 1.
“A total of 33% participants (51/154) had stage F1 fibrosis, 45% (70/154) had stage F2, and 19% (29/154) had stage F3.”
ResultsFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2The effect of dapagliflozin on MASH improvement is largely mediated by weight loss.The mediation analysis is exploratory and not fully detailed; the claim is presented as suggestive, not definitive.Evidence: Mediation analysis mentioned in Results and Discussion.
“In addition, weight loss was found to largely mediate the total effect of dapagliflozin on the higher proportion of MASH improvement and MASH resolution, but not fibrosis improvement in the mediation analysis (supplementary table 6).”
ResultsFind in source - supportedReviewers 1, 2Dapagliflozin improves MASH without worsening of fibrosis compared with placebo.The primary endpoint result (53% vs 30%, RR 1.73, P=0.006) directly supports this claim.Evidence: Primary endpoint result in Results section.
“MASH improvement without worsening of fibrosis was reported in 53% (41/78) of participants in the dapagliflozin group and 30% (23/76) in the placebo group (risk ratio 1.73 (95% confidence interval (CI) 1.16 to 2.58); P=0.006).”
AbstractFind in source - supportedReviewers 1, 2Dapagliflozin leads to MASH resolution without worsening of fibrosis.The secondary endpoint result (23% vs 8%, RR 2.91, P=0.01) supports this claim.Evidence: Secondary endpoint result in Results section.
“MASH resolution without worsening of fibrosis occurred in 23% (18/78) of participants in the dapagliflozin group and 8% (6/76) in the placebo group (risk ratio 2.91 (95% CI 1.22 to 6.97); P=0.01).”
AbstractFind in source - supportedReviewers 1, 2Dapagliflozin improves fibrosis without worsening of MASH.The secondary endpoint result (45% vs 20%, RR 2.25, P=0.001) supports this claim.Evidence: Secondary endpoint result in Results section.
“Fibrosis improvement without worsening of MASH was reported in 45% (35/78) of participants in the dapagliflozin group, as compared with 20% (15/76) in the placebo group (risk ratio 2.25 (95% CI 1.35 to 3.75); P=0.001).”
AbstractFind in source - supportedReviewers 1, 2Dapagliflozin is safe and well tolerated in this population.Safety data show similar adverse event rates and no serious events in the dapagliflozin group.Evidence: Safety section in Results.
“Adverse events were reported in 56% (44/78) of the participants in the dapagliflozin group, compared with 64% (49/76) of those in the placebo group (supplementary table 7).”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is a histological surrogate (NAS-based MASH improvement) rather than a hard clinical outcome. Although the paper cites FDA guidance that such endpoints are 'reasonably likely to predict long term clinical benefit', it does not provide validated evidence linking the specific surrogate (NAS change) to clinical outcomes such as cirrhosis or mortality. Target engagement at the tested dose is not demonstrated (no PK/PD data).
“The primary endpoint was MASH improvement (defined as a decrease of at least 2 points in non-alcoholic fatty liver disease activity score (NAS) or a NAS of ≤3 points) without worsening of liver fibrosis (defined as without increase of fibrosis stage) at 48 weeks.”
- INADEQUATEEffect sizeThe primary effect is reported as a risk ratio of 1.73 (53% vs 30%) for MASH improvement, but the absolute difference is 23%. The clinical meaningfulness is not anchored to a minimal clinically important difference or a hard outcome. The effect on NAS is a mean difference of -1.39, but the clinical significance of this change is not established.
“MASH improvement without worsening of fibrosis was reported in 53% (41/78) of participants in the dapagliflozin group and 30% (23/76) in the placebo group (risk ratio 1.73 (95% confidence interval (CI) 1.16 to 2.58); P=0.006).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The paper cites prior research on SGLT2 inhibitors and MASH, acknowledges inconsistent findings from two small trials, and states that no trial has been conducted in biopsy-diagnosed MASH. The rationale linking SGLT2 inhibition to MASH pathophysiology is logical and the hypothesis follows directly. Limitations of prior studies (small sample sizes, possible absence of key histological features) are explicitly noted.
“Two small clinical trials have reported inconsistent findings of SGLT2 inhibitors (ipragliflozin and tofogliflozin) on liver histological features among people with diabetes and MASLD, but no trial has been conducted among participants with biopsy diagnosed MASH.”
“Therefore, high quality evidence is required to develop clinical guidelines for SGLT2 inhibitors in MASH treatment.”
“These studies did not provide adequate information needed to formulate evidence based clinical guidance for MASH management because of the limited sample sizes and the possible absence of ballooning, lobular inflammation or fibrosis in the MASLD population.”
“Two small clinical trials have reported inconsistent findings of SGLT2 inhibitors (ipragliflozin and tofogliflozin) on liver histological features among people with diabetes and MASLD, but no trial has been conducted among participants with biopsy diagnosed MASH.”
“Therefore, high quality evidence is required to develop clinical guidelines for SGLT2 inhibitors in MASH treatment.”
Randomisation was computer-generated and stratified by diabetes status; allocation concealment via coded containers; participants, investigators, site personnel, and pathologists were masked. A sample size calculation with assumptions is provided. Inclusion/exclusion criteria are described, and the analysis population (ITT) and missing-data handling (non-responder imputation) are defined. The trial is a single pivotal trial, so independent replication is not applicable.
“The randomisation schedules were generated using a computer program centrally and stratified by the presence of type 2 diabetes.”
“Participants, investigators, site personnel, and pathologists were masked to treatment assignments.”
“We estimated that a sample size of 148 participants (74 per group) would provide the trial with greater than 90% power to detect a difference of 27% between dapagliflozin and placebo for the primary endpoint at a two sided significance level of 0.05”
“The randomisation schedules were generated using a computer program centrally and stratified by the presence of type 2 diabetes.”
“Participants, investigators, site personnel, and pathologists were masked to treatment assignments.”
“We estimated that a sample size of 148 participants (74 per group) would provide the trial with greater than 90% power to detect a difference of 27% between dapagliflozin and placebo for the primary endpoint at a two sided significance level of 0.05, assuming that 20% of the participants in the placebo group would reach the primary endpoint and 47% of the participants in the dapagliflozin group would meet the primary endpoint, as well as an anticipated dropout rate of 20%.”
Sex is reported for each group (64/76 male in placebo, 67/78 in dapagliflozin). Age, BMI, and health status (diabetes, dyslipidaemia, hypertension) are reported in Table 1. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing are not applicable for a human trial.
“The mean age of the participants was 35.1 (standard deviation (SD) 10.2) years, mean body mass index was 29.2 (SD 4.3)”
“Of them, 85% (131/154) were male, 85% (131/154) had dyslipidaemia, and 45% (69/154) had type 2 diabetes.”
“The mean age of the participants was 35.1 (standard deviation (SD) 10.2) years, mean body mass index was 29.2 (SD 4.3), and mean NAS was 6.0 (SD 1.1).”
The paper states that the protocol was approved by the review board at each participating centre, and participants provided written informed consent. It also states compliance with the Declaration of Helsinki and ICH-GCP. This satisfies the requirements for human research.
“The DEAN trial protocol was approved by the review board at each participating centre.”
“Participants provided written informed consent.”
“The trial was conducted in accordance with the principles of the Declaration of Helsinki, the International Council for Harmonization, Good Clinical Practice guidelines, and all relevant regulations.”
“The DEAN trial protocol was approved by the review board at each participating centre.”
“Participants provided written informed consent.”
“The trial was conducted in accordance with the principles of the Declaration of Helsinki, the International Council for Harmonization, Good Clinical Practice guidelines, and all relevant regulations.”
Dapagliflozin is identified as AstraZeneca product, 10 mg, with dose and regimen. Placebo is described as matching. Statistical software (SAS 9.4) is identified. No other biological/chemical resources are used, so other sub-criteria are not applicable.
“Participants were randomly assigned to receive 10 mg dapagliflozin (AstraZeneca; IN, USA) or matching placebo once daily in a 1:1 ratio.”
“Statistical analyses were conducted using SAS software, version 9.4 (SAS Institute).”
“Participants were randomly assigned to receive 10 mg dapagliflozin (AstraZeneca; IN, USA) or matching placebo once daily in a 1:1 ratio.”
“Statistical analyses were conducted using SAS software, version 9.4 (SAS Institute).”
The primary analysis uses Cochran-Mantel-Haenszel with stratification, and continuous endpoints use mixed models. Exact p-values are reported (e.g., P=0.006). Effect sizes with 95% CIs are provided throughout. Software is identified. Data presentation includes per-group n and CIs. Mathematical plausibility checks: the reported percentages and counts are consistent (e.g., 53% of 78 = 41.34, reported 41/78; 30% of 76 = 22.8, reported 23/76, which is plausible rounding). No arithmetic errors detected.
“We used the Cochran-Mantel-Haenszel method, controlling for the randomisation stratification factor (baseline diabetes status), for analysis of the primary and secondary endpoints.”
“risk ratio 1.73 (95% confidence interval (CI) 1.16 to 2.58); P=0.006”
“Mean difference of NAS was −1.39 (95% CI −1.99 to −0.79); P<0.001”
“We used the Cochran-Mantel-Haenszel method, controlling for the randomisation stratification factor (baseline diabetes status), for analysis of the primary and secondary endpoints.”
“MASH improvement without worsening of fibrosis was reported in 53% (41/78) of participants in the dapagliflozin group and 30% (23/76) in the placebo group (risk ratio 1.73 (95% confidence interval (CI) 1.16 to 2.58); P=0.006).”
The data availability statement provides a concrete route: data are available in a GitHub repository (https://github.com/Yurence/DEAN_trial.git). The code used for analysis is also in supplemental files. Since the data are not individual patient-level but likely aggregate, repository deposit is adequate. No accession numbers are applicable.
“The data underlying the findings of this paper are available in: https://github.com/Yurence/DEAN_trial.git .”
“The codes used to analyse the data can be found in the supplemental files.”
“The data underlying the findings of this paper are available in: https://github.com/Yurence/DEAN_trial.git .”
“The codes used to analyse the data can be found in the supplemental files.”
Trial registration number is provided (NCT03723252). Methods are comprehensive. Limitations are explicitly discussed (Chinese population, male predominance, primary endpoint choice). Conclusions are proportional to the results. Funding and competing interests are disclosed. Reporting guideline (CONSORT) is not explicitly mentioned but the paper follows it; this is a minor gap.
“Trial registration ClinicalTrials.gov NCT03723252 .”
“Secondly, the current trial was conducted in a Chinese population, which limits the broader generalisability of the conclusion.”
“Trial registration ClinicalTrials.gov NCT03723252 .”
“Secondly, the current trial was conducted in a Chinese population, which limits the broader generalisability of the conclusion.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 36 references by DOI: 35 verified — 1 no DOI (shown, not verified).
- NO DOINoncirrhotic nonalcoholic steatohepatitis with liver fibrosis: developing drugs for treatment: guidance for industryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- codeGitHubLIVEHTTP 200https://github.com/Yurence/DEAN_trial.gitResolves to GitHub (code repository).
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAbstract, Results“Mean difference of NAS was −1.39 (95% CI −1.99 to −0.79); P<0.001).”→ Remove the extra closing parenthesis: 'Mean difference of NAS was −1.39 (95% CI −1.99 to −0.79); P<0.001.'Extra parenthesis after the CI.
- MINORconsistencyResults, Participants“A total of 33% participants (51/154) had stage F1 fibrosis, 45% (70/154) had stage F2, and 19% (29/154) had stage F3.”→ Ensure the percentages sum to 100% (33+45+19=97) and clarify if F0 and F4 are included.Percentages do not sum to 100; likely due to rounding or missing F0/F4.
- MINORclarityMethods, Statistical analysis“We used the Cochran-Mantel-Haenszel method, controlling for the randomisation stratification factor (baseline diabetes status), for analysis of the primary and secondary endpoints.”→ Clarify that the method was used for binary endpoints, and specify the test for continuous endpoints.The sentence is clear but could be more explicit.
- MINORtypoAbstract, Results“Mean difference of NAS was −1.39 (95% CI −1.99 to −0.79); P<0.001).”→ Remove the extra closing parenthesis: 'Mean difference of NAS was −1.39 (95% CI −1.99 to −0.79); P<0.001).'Extra parenthesis after the CI.
- MINORconsistencyResults, Participants“A total of 33% participants (51/154) had stage F1 fibrosis, 45% (70/154) had stage F2, and 19% (29/154) had stage F3.”→ The percentages sum to 97%, not 100%. Consider adding a note about missing or other stages.Percentages do not sum to 100% (33+45+19=97).
- MINORclarityMethods, Statistical analysis“We used the Cochran-Mantel-Haenszel method, controlling for the randomisation stratification factor (baseline diabetes status), for analysis of the primary and secondary endpoints.”→ Clarify that the CMH test was used for binary endpoints, and specify the exact test for continuous endpoints.The sentence is clear but could be more precise.
The published paper is robust and methodologically sound. An informed reader should weigh the minor reporting issues (missing CONSORT reference, possible group-label swap in one sentence) as low-severity; none warrant an erratum or independent re-analysis, but the authors should consider issuing a correction for the group-label swap if confirmed.
- 1.HIGHreportingCorrect the possible group-label swap in the Results sentence: 'worsening of fibrosis occurred in 5% (4/76) in the dapagliflozin group, and 22% (17/78) in the placebo group' — the dapagliflozin group has 78 participants and placebo has 76, so the labels may be reversed.This internal contradiction could mislead readers about the direction of the effect and should be verified and corrected.
- 2.HIGHreportingExplicitly state adherence to CONSORT reporting guidelines in the Methods or a dedicated section.Both reviewers noted this omission; referencing CONSORT strengthens transparency and is expected for a clinical trial.
- 3.MEDIUMcopyeditFix the extra closing parenthesis in the Abstract Results: 'Mean difference of NAS was −1.39 (95% CI −1.99 to −0.79); P<0.001).' should read 'Mean difference of NAS was −1.39 (95% CI −1.99 to −0.79); P<0.001.'Typographical error that could cause confusion.
- 4.MEDIUMreportingClarify the fibrosis stage percentages in the Results text: 33% (F1) + 45% (F2) + 19% (F3) = 97%, not 100%. Add a note about F0 and F4 stages or explain rounding.The copyedit and integrity checks flagged this inconsistency; readers may question the completeness of the data.
- 5.MEDIUMreportingAdd a statement about the availability of the full protocol and statistical analysis plan (SAP) in a public repository (e.g., ClinicalTrials.gov) to enhance transparency.Both reviewers suggested this; it improves reproducibility and trust.
- 6.MEDIUMreportingIn the Data Availability Statement, clarify the license and any access conditions for the GitHub repository.Ensures that readers know the terms under which data and code can be reused.
- 7.MEDIUMreportingAdd a note in the Methods about how missing data for secondary endpoints were handled beyond the primary non-responder imputation.Both reviewers noted this gap; it improves methodological transparency.
- 8.MEDIUMreportingIn the Discussion, explicitly acknowledge that the primary endpoint was changed from the originally registered endpoint, and discuss the potential impact on interpretation.Both reviewers flagged this as a transparency issue; readers should be aware of any changes to the primary outcome.
- 9.LOWreportingConsider reporting the number of participants who were screened and excluded at each stage in the CONSORT flow diagram to improve transparency.Enhances the completeness of the trial reporting.
- 10.LOWreportingIn the safety section, provide more detailed tabulation of adverse events by severity and relatedness to treatment.Improves the completeness of safety reporting.
- 11.LOWstatisticsConsider adding a sensitivity analysis using a different imputation method (e.g., multiple imputation with chained equations) to test robustness.Strengthens the robustness of the primary analysis.
- 12.LOWreportingAdd a statement about the use of a data monitoring committee and any interim analyses, if applicable.Provides additional context about trial oversight.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.