A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trial
Xu Z, Ren F, Wang P, Cao J, Tan C, Ma D, Zhao L, Dai J, Ding Y, Fang H, Li H, Liu H, Luo F, Meng Y, Pan P, Xiang P, Xiao Z, Rao S, Satler C, Liu S, Lv Y, Zhao H, Chen S, Cui H, Korzinkin M, Gennert D, Zhavoronkov A.
- DOI
- 10.1038/s41591-025-03743-2
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/dc17aebb-3ae8-4e41-8ad2-ec092c41534f is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×6−3★
- ClaimsOverstated claim ×3−1.5★
- ReportingStatistical analysis not met−0.5★
- ReportingKey resources partially met−0.25★
- References were not verified against Crossref/OpenAlex.
- No reported statistical tests were found to recompute.
- 01Statistical reporting inadequate
Statistical tests are named and effect estimates are reported with 95% CIs, but the paper contains demonstrable arithmetic/percentage errors in the TEAE reporting, and several p-values are reported only as thresholds.
“3 in 60 mg QD (20.4%)”
ResultsFind in source - 02Internal contradictions in the reported numbers
ALT-increase counts in the AE narrative contradict Table 2: the text lists ALT increase as '6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%)' (and lists 60 mg QD twice), while Table 2 shows ALT increase of 0, 0, 1, 0 across the four arms.
“alanine aminotransferase (ALT) increase (1 in placebo (5.9%), 1 in 30 mg QD (5.6%), 1 in 30 mg BID (5.6%), 6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%))”
Table 2Find in source - 03Internal contradictions in the reported numbers
The text reports two different numbers for ALT increase in the 60 mg QD group (6 and 3), while Table 2 shows only 1 case. This is a clear internal inconsistency that needs correction.
“alanine aminotransferase (ALT) increase (1 in placebo (5.9%), 1 in 30 mg QD (5.6%), 1 in 30 mg BID (5.6%), 6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%)).”
ResultsFind in source - 04Internal contradictions in the reported numbers
The ALT increase sentence lists two separate counts for the same 60 mg QD arm (6 events and 3 events), which is internally inconsistent as written.
ALT increase (1 in placebo (5.9%), 1 in 30 mg QD (5.6%), 1 in 30 mg BID (5.6%), 6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%))
Resultsreviewer’s wording - 05Mathematically impossible statistic
The reported percentage 20.4% for 3 of 18 patients in the 60 mg QD group is arithmetically impossible (3/18 = 16.7%).
“3 in 60 mg QD (20.4%)”
ResultsFind in source - 06Conclusion reaches beyond the evidence
The AI-discovered agent is 'safe and effective' for IPF.
“Preliminary results from a phase 2a trial involving 71 patients suggest that a new agent, discovered and designed with artificial intelligence assistance, is safe and effective for the treatment of idiopathic pulmonary fibrosis.”
AbstractFind in source
3 further findings of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
Automated narrative synthesis was unavailable for this run; this verdict is the majority vote of the independent reviewers, with their own suggestions listed as action items.
Fallback synthesis: the narrative synthesizer did not return a usable result, so the report is assembled directly from the reviewers.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
- Mathematically impossible statisticAssessed
- mediuminternal contradictionALT-increase counts in the AE narrative contradict Table 2: the text lists ALT increase as '6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%)' (and lists 60 mg QD twice), while Table 2 shows ALT increase of 0, 0, 1, 0 across the four arms.
“alanine aminotransferase (ALT) increase (1 in placebo (5.9%), 1 in 30 mg QD (5.6%), 1 in 30 mg BID (5.6%), 6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%))”
Table 2Find in source - mediuminternal contradictionThe text reports two different numbers for ALT increase in the 60 mg QD group (6 and 3), while Table 2 shows only 1 case. This is a clear internal inconsistency that needs correction.
“alanine aminotransferase (ALT) increase (1 in placebo (5.9%), 1 in 30 mg QD (5.6%), 1 in 30 mg BID (5.6%), 6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%)).”
ResultsFind in source - mediuminternal contradictionThe ALT increase sentence lists two separate counts for the same 60 mg QD arm (6 events and 3 events), which is internally inconsistent as written.
ALT increase (1 in placebo (5.9%), 1 in 30 mg QD (5.6%), 1 in 30 mg BID (5.6%), 6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%))
Resultsreviewer’s wording - mediumimpossible statisticThe reported percentage 20.4% for 3 of 18 patients in the 60 mg QD group is arithmetically impossible (3/18 = 16.7%).
“3 in 60 mg QD (20.4%)”
ResultsFind in source - lowimpossible statisticThe hypokalemia count/percentage pair '3 in 60 mg QD (20.4%)' is arithmetically inconsistent: 3/18 = 16.7%, so 20.4% cannot be derived from the stated count and group size.
“hypokalemia (2 in placebo (11.8%), 3 in 30 mg QD (16.7%), 5 in 30 mg BID (27.8%) and 3 in 60 mg QD (20.4%))”
ResultsFind in source
Overstated conclusions
2 findings · worst mediumConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
12 major claims checked against the paper's own evidence: 3 not fully backed by the presented evidence (unsupported or overstated).
- overstatedReviewer 1The AI-discovered agent is 'safe and effective' for IPF.Safety is supported, but 'effective' overreaches: the primary endpoint was safety, and the efficacy signal was exploratory/secondary with no confirmatory between-group p-value.Evidence: No primary efficacy endpoint tested with a p-value; FVC improvement reported only as a per-group estimate with overlapping CIs.
“Preliminary results from a phase 2a trial involving 71 patients suggest that a new agent, discovered and designed with artificial intelligence assistance, is safe and effective for the treatment of idiopathic pulmonary fibrosis.”
AbstractFind in source - overstatedReviewer 2Rentosertib is 'safe and effective for the treatment of idiopathic pulmonary fibrosis'.Safety is supported, but 'effective' overreaches: the FVC improvement at 60 mg QD is a secondary endpoint showing a trend (+98.4 ml, CI 10.9-185.9 vs placebo -20.3 ml) with no reported between-group significance, in a small phase 2a cohort.Evidence: Safety data (TEAE rates similar across arms) plus a secondary FVC trend at the highest dose; no primary-efficacy endpoint or significance test for FVC is reported.
“is safe and effective for the treatment of idiopathic pulmonary fibrosis”
AbstractFind in source - overstatedReviewer 3The study represents a revolutionary shift in the streamlining of drug discovery.The paper presents a single example of AI-discovered drug reaching phase 2a, which is promising but not yet a 'revolutionary shift' as it has not completed phase 3 or been approved.Evidence: The claim is made in the introduction but not supported by the phase 2a results alone.
Our generative AI-powered approach streamlined preclinical candidate nomination to a mere 18 months and completion of phase 0/1 clinical testing to under 30 months from the initiation of target discovery, representing a revolutionary shift in the streamlining of drug discovery.
Introductionreviewer’s wording - partialReviewer 160 mg rentosertib QD increased FVC compared with placebo.The 60 mg arm showed a within-group mean FVC increase of +98.4 ml with 95% CI excluding zero, but the between-group difference vs placebo was not statistically tested and the CIs overlap.Evidence: Mean change +98.4 ml (95% CI 10.9 to 185.9) for 60 mg QD vs −20.3 ml (95% CI −116.1 to 75.6) for placebo; no p-value for the comparison.
“We observed increased forced vital capacity at the highest dosage with a mean change of +98.4 ml (95% confidence interval 10.9 to 185.9) for patients in the 60 mg rentosertib QD group, compared with −20.3 ml (95% confidence interval −116.1 to 75.6) for the placebo group.”
AbstractFind in source - partialReviewer 160 mg QD without SOC antifibrotic therapy significantly improved FVC.The subgroup mean change has a CI excluding zero, but the subgroup analysis is post-hoc, small, and lacks multiplicity control, so 'significant' is not fully established.Evidence: Subgroup result: +187.8 ml, 95% CI 68.6 to 306.9 ml.
“Patients receiving 60 mg rentosertib QD not concurrently taking SOC antifibrotic therapy exhibited significant improvement in FVC (+187.8 ml, 95% CI 68.6 to 306.9 ml)”
ResultsFind in source - partialReviewer 2Downregulation of the IPF-associated serum protein profile supports the hypothesis that inhibition of TNIK modulates IPF pathophysiological pathways.The proteomic data show dose- and time-dependent downregulation of fibrosis-associated proteins, but the inference that this is specifically due to TNIK inhibition is indirect (no direct TNIK-modulation readout in patients).Evidence: Differential protein abundance analyses, ECM-organization pathway enrichment, and inverse correlations with FVC change (Fig. 4, Extended Data Figs. 6-8).
“Downregulation of this IPF-associated protein profile supports the hypothesis that inhibition of TNIK modulates IPF pathophysiological pathways”
Discussion ¶5Find in source - partialReviewer 3Downregulation of fibrosis-associated proteins supports that TNIK inhibition modulates IPF pathways.The proteomic data show correlation but not causation; the mechanistic link is inferred from pathway enrichment and prior knowledge.Evidence: Results: 'Pathway enrichment analysis identified extracellular matrix organization as the most downregulated Reactome pathway gene set' and correlation with FVC change.
“Downregulated proteins associated with 30 mg BID and 60 mg QD treatment include known fibrosis-associated proteins such as MMP10, PTPRZ1, COL1A1, FAP, FN1, ROBO2, ASPN and LTBP2”
ResultsFind in source - supportedReviewer 1Rentosertib is safe and well tolerated in patients with IPF.The primary safety endpoint (TEAE rates) was similar across arms, and treatment-related SAEs were low, directly supporting the claim.Evidence: Primary endpoint TEAE rates 72.2%, 83.3%, 83.3%, and 70.6% across arms; low treatment-related SAE rates.
“These results suggest that targeting TNIK with rentosertib is safe and well tolerated and warrants further investigation in larger-scale clinical trials of longer duration.”
AbstractFind in source - supportedReviewer 1Downregulation of IPF-associated proteins supports TNIK pathway modulation.The proteomic data directly show dose- and time-dependent downregulation of fibrosis-associated proteins and ECM pathway enrichment, supporting the mechanistic claim.Evidence: Olink proteomics: 22 high-confidence differentially abundant proteins at 60 mg QD; COL1A1, FAP, FN1, MMP10 decrease with dose/time; Reactome ECM organization enrichment.
“Downregulation of this IPF-associated protein profile supports the hypothesis that inhibition of TNIK modulates IPF pathophysiological pathways and points to a potential serum protein signature as a biomarker for response to rentosertib treatment.”
DiscussionFind in source - supportedReviewer 2Targeting TNIK with rentosertib is safe and well tolerated and warrants further investigation.The safety/tolerability data (TEAE rates and SAE rates comparable across arms, detailed AE tables) adequately back this claim, which is also appropriately hedged.Evidence: Primary safety endpoint (TEAE percentages across arms) and treatment-related SAE rates reported in Results and Table 2.
“These results suggest that targeting TNIK with rentosertib is safe and well tolerated and warrants further investigation in larger-scale clinical trials of longer duration.”
AbstractFind in source - supportedReviewer 2Treatment with 60 mg rentosertib QD over 12 weeks was associated with a trend toward an increase in FVC.The reported +98.4 ml change (95% CI 10.9 to 185.9) at 60 mg QD versus -20.3 ml in placebo supports a trend claim, and the paper labels it a trend rather than proof of efficacy.Evidence: FVC change with 95% CI for the 60 mg QD and placebo groups in Results and Fig. 2.
“Treatment with 60 mg rentosertib QD over 12 weeks was associated with a trend toward an increase in FVC in patients with IPF.”
Discussion ¶2Find in source - supportedReviewer 2This is the first phase 2a trial of an AI-discovered/AI-designed drug, and the first report of AI-enabled discovery of both a target and a compound.The paper consistently frames rentosertib as first-in-class AI-generated and cites its own prior work; the novelty claim is presented as the study's premise and is internally consistent (though not independently verifiable here).Evidence: Background statements and citation of prior phase 0/1 results (NCT05154240, CTR20221542).
“This was also, importantly, the first reported instance of AI platform-enabled discovery of both a disease-associated target and a compound for that target.”
Main, paragraph 2Find in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointPrimary efficacy endpoint is FVC, a validated clinical outcome in IPF, not a surrogate biomarker.
“FVC is the gold-standard metric for assessing the lung function of patients with IPF and response to therapeutic intervention”
- ADEQUATEEffect sizeReported mean FVC increase of +98.4 ml (95% CI 10.9 to 185.9) in 60 mg QD group, percentage change of 2.82% meeting MCID of 2–6%.
“Patients receiving 60 mg rentosertib QD experienced a mean increase in FVC percentage change of 2.82% by our modeling, meeting the minimal clinically important difference (MCID) reported for FVC in IPF of 2–6%”
Data authenticity concerns
1 finding · worst mediumAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
6 integrity concerns flagged (0 high).
- mediumotherTrial NCT05938920 was first submitted to ClinicalTrials.gov on 2023-06-28, after the registered study start date of 2023-06-19. Retrospective registration means the protocol and outcomes were not on the public record before the study ran, which is what prospective registration exists to establish.
NCT05938920
reviewer’s wording
Reporting gaps
2 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Statistical reporting inadequateAssessed
- Key resources under-identified (antibodies, cell lines, RRIDs)Assessed
The introduction cites prior studies on TNIK as a target for IPF, AI-driven drug discovery, and limitations of current therapies. It acknowledges that no AI-discovered drug has progressed through phase 3, providing a clear rationale for the trial. The hypothesis that TNIK inhibition may treat IPF follows from the cited evidence.
“Despite these advancements, few AI-discovered or AI-designed drugs have reached clinical trials.”
“Clinical trials investigating these two drugs in patients with IPF have demonstrated only a slowing of disease progression”
“Clinical trials investigating these two drugs in patients with IPF have demonstrated only a slowing of disease progression”
“AI-discovered drugs have experienced similar levels of phase 2 trial failure as non-AI-discovered drugs”
“We reported the development and positive phase 0 and 1 trial results of rentosertib”
“Despite these advancements, few AI-discovered or AI-designed drugs have reached clinical trials. AI-discovered drugs have experienced similar levels of phase 2 trial failure as non-AI-discovered drugs”
Randomization was performed via interactive response technology in a 1:1:1:1 ratio. The study is double-blind, with blinding of subjects, investigators, and staff. Inclusion and exclusion criteria are detailed. Missing data were handled by multiple imputation. A sample size calculation was provided for safety (90% probability to detect an AE with 15% rate) but not for efficacy endpoints.
“Adults with IPF were randomly assigned in a 1:1:1:1 ratio via interactive response technology to receive oral rentosertib”
“Subjects, investigators, site study staff, reviewers and everyone involved in study conduct or analysis were blinded with regard to the randomized treatment assignments until after data freeze.”
“although a sample size calculation based on statistical power considerations was not performed”
“Adults with IPF were randomly assigned in a 1:1:1:1 ratio via interactive response technology”
“Subjects, investigators, site study staff, reviewers and everyone involved in study conduct or analysis were blinded with regard to the randomized treatment assignments until after data freeze.”
“although a sample size calculation based on statistical power considerations was not performed”
“Adults with IPF were randomly assigned in a 1:1:1:1 ratio via interactive response technology”
“Subjects, investigators, site study staff, reviewers and everyone involved in study conduct or analysis were blinded with regard to the randomized treatment assignments until after data freeze”
“Given an approximate sample size of 15 subjects per treatment arm, there exists a 90% probability of observing at least one AE if the true population rate is approximately 15%, which was sufficient to assess the feasibility of safety parameters, although a sample size calculation based on statistical power considerations was not performed.”
Table 1 reports sex, age, weight, BMI, smoking history, and baseline lung function (FVC, DLCO, LCQ) for all treatment groups. All participants are male (90.1%) and Asian, but this is a single-country study and does not require justification for a single-sex study as it is not by design.
“Male sex – no. (%) | 18 (100.0) | 15 (83.3) | 16 (88.9) | 15 (88.2) | 64 (90.1)”
“Age – years, mean ± s.d. | 65.8 ± 7.0 | 67.2 ± 7.8 | 65.7 ± 6.8 | 68.3 ± 5.3 | 66.7 ± 6.7”
“Male sex – no. (%) | 18 (100.0) | 15 (83.3) | 16 (88.9) | 15 (88.2) | 64 (90.1)”
“Male sex – no. (%) | 18 (100.0) | 15 (83.3) | 16 (88.9) | 15 (88.2) | 64 (90.1)”
“Age – years, mean ± s.d. | 65.8 ± 7.0 | 67.2 ± 7.8 | 65.7 ± 6.8 | 68.3 ± 5.3 | 66.7 ± 6.7”
The paper states that the institutional review board or ethics committee at participating centers approved the protocols, and all patients provided written informed consent. The trial was conducted in accordance with the Declaration of Helsinki and ICH-GCP guidelines.
“The institutional review board or ethics committee at participating centers approved protocols and adhered to local laws before initiation of the clinical trial.”
“All patients in this study provided written informed consent.”
“The trial was conducted following the principles outlined in the Declaration of Helsinki and the International Council for Harmonization guidelines for Good Clinical Practice.”
“The institutional review board or ethics committee at participating centers approved protocols and adhered to local laws before initiation of the clinical trial.”
“All patients in this study provided written informed consent.”
“The trial was conducted following the principles outlined in the Declaration of Helsinki and the International Council for Harmonization guidelines for Good Clinical Practice.”
“The institutional review board or ethics committee at participating centers approved protocols and adhered to local laws before initiation of the clinical trial.”
“All patients in this study provided written informed consent.”
“The trial was conducted following the principles outlined in the Declaration of Helsinki and the International Council for Harmonization guidelines for Good Clinical Practice.”
Rentosertib is named with dose and regimen, and provided by Insilico Medicine. However, statistical software such as lme4 and clusterProfiler are mentioned without version numbers. SpiroSphere devices are identified but no version. Other resources like Olink panel are identified but not with catalog numbers. The drug is adequately identified, but software identification is incomplete.
“The study treatment rentosertib was provided by Insilico Medicine or a designated contract research organization, packaged and labeled in accordance with the principles of Good Manufacturing Practice.”
“The lme4 package was used for the generalized linear mixed model analysis.”
“using the Olink Explore 3072 panel”
“The study treatment rentosertib was provided by Insilico Medicine or a designated contract research organization”
“Custom code used to analyze Olink proteomics data is available via GitHub at https://github.com/HUICUI1992/Code-for-NM”
“oral rentosertib (at a dose of 30 mg (QD), 30 mg (BID) or 60 mg (QD))”
“The lme4 package was used for the generalized linear mixed model analysis.”
ANCOVA, MMRM, Spearman, Pearson, paired t-test, and generalized linear mixed models are named; effect sizes are reported with CIs. However, two TEAE percentage errors are present: 3/18 is reported as 20.4% (actual 16.7%), and the ALT increase listing includes two separate entries for the same 60 mg QD arm (6 events at 33.3% and 3 events at 16.7%). These are demonstrable numeric errors, warranting a fail despite otherwise adequate reporting. Software versions are absent, and many p-values are given as thresholds ('P < 0.05').
“3 in 60 mg QD (20.4%)”
“two-sided P = 0.0495”
“patients receiving 60 mg rentosertib QD showed improved FVC, with and +98.4 ml (95% CI 10.9 to 185.9)”
“1, 8 and 22 high-confidence ( P adj < 0.05) differentially abundant proteins”
“hypokalemia (2 in placebo (11.8%), 3 in 30 mg QD (16.7%), 5 in 30 mg BID (27.8%) and 3 in 60 mg QD (20.4%))”
“our modeling indicated a significant increase of least-squares mean change of the LCQ scores of patients receiving 60 mg rentosertib compared with those receiving placebo (two-sided P = 0.0495).”
“patients receiving 60 mg rentosertib QD showed improved FVC, with and +98.4 ml (95% CI 10.9 to 185.9)”
The data availability statement names a repository and accession, describes a managed-access route for the Olink data, and states that the protocol and SAP are available on request. The custom code used for Olink proteomics analysis is available via a public GitHub repository. All applicable sub-criteria are adequate.
“Complete deidentified Olink proteomics data, and FVC data, have been deposited at the OMIX database under accession codes at OMIX008341 ( https://ngdc.cncb.ac.cn/omix/release/OMIX008341 ).”
“Custom code used to analyze Olink proteomics data is available via GitHub at https://github.com/HUICUI1992/Code-for-NM .”
“Study protocol and statistical analysis plan will be provided in a secure data sharing environment upon academic or research request.”
“Complete deidentified Olink proteomics data, and FVC data, have been deposited at the OMIX database under accession codes at OMIX008341 ( https://ngdc.cncb.ac.cn/omix/release/OMIX008341 )”
“Custom code used to analyze Olink proteomics data is available via GitHub at https://github.com/HUICUI1992/Code-for-NM”
“Custom code used to analyze Olink proteomics data is available via GitHub at https://github.com/HUICUI1992/Code-for-NM”
Methods are detailed enough for replication. The trial is registered at ClinicalTrials.gov (NCT05938920). All primary and secondary outcomes are reported. Limitations are discussed (small cohort, homogeneity, short follow-up). Conclusions are proportional. Funding and competing interests are disclosed. However, no explicit statement of adherence to CONSORT or other reporting guidelines is provided, only a vague reference to a 'Reporting summary'.
“ClinicalTrials.gov registration number: NCT05938920”
“Further information on research design is available in the linked to this article.”
“Preliminary results from a phase 2a trial involving 71 patients suggest that a new agent, discovered and designed with artificial intelligence assistance, is safe and effective for the treatment of idiopathic pulmonary fibrosis.”
“ClinicalTrials.gov registration number: NCT05938920”
“The limitations of this study include the small cohort size of each arm, the geographical and demographic homogeneity of the participants (all were residents of China of similar race) and a short period of follow-up”
“is safe and effective for the treatment of idiopathic pulmonary fibrosis”
“ClinicalTrials.gov registration number: NCT05938920 (https://clinicaltrials.gov/study/NCT05938920)”
“The limitations of this study include the small cohort size of each arm, the geographical and demographic homogeneity of the participants (all were residents of China of similar race) and a short period of follow-up”
“Reporting summary Further information on research design is available in the linked to this article.”
Registered (2 IDs: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
6 data/code links checked; 6 live.
- datahttps://clinicaltrials.gov/study/NCT05938920LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://ngdc.cncb.ac.cn/omix/release/OMIX008341LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://ngdc.cncb.ac.cn/omix/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT05154240LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttp://www.chinadrugtrials.org.cn/LIVEHTTP 202Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/HUICUI1992/Code-for-NMResolves to GitHub (code repository).
Copyediting
1 finding · worst lowWording, consistency and formatting errors that need correcting before submission.
- Wording or formatting errors that need correctingAssessed
13 copyedit issues flagged (3 major): mostly consistency, grammar, clarity.
- MAJORconsistencyResults, Primary safety endpoint“3 in 60 mg QD (20.4%)”→ Change to '3 in 60 mg QD (16.7%)'3/18 = 16.7%, not 20.4%.
- MAJORconsistencyResults, Primary safety endpoint“ALT increase (1 in placebo (5.9%), 1 in 30 mg QD (5.6%), 1 in 30 mg BID (5.6%), 6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%))”→ Reconcile the duplicate '60 mg QD' entries; one count likely belongs to a different group or subcategory.The same arm appears twice with two different event counts.
- MAJORconsistencyResults, Primary safety endpoint vs Table 2“alanine aminotransferase (ALT) increase (1 in placebo (5.9%), 1 in 30 mg QD (5.6%), 1 in 30 mg BID (5.6%), 6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%))”→ Reconcile narrative counts with Table 2 ALT-increase row (0, 0, 1, 0) and remove the duplicated 60 mg QD entryText and Table 2 disagree; '60 mg QD' appears twice for ALT increase.
- MINORgrammarResults, Secondary lung function endpoints“with and +98.4 ml”→ Delete the stray 'and'.Typographical error.
- MINORgrammarResults, Pharmacokinetics“indicates an increased exposure to rentosertib with increases time on treatment”→ Change to 'with increased time on treatment'.Grammar error.
- MINORclarityResults, Secondary lung function endpoints“with and +98.4 ml (95% CI 10.9 to 185.9)”→ Remove 'and' -> 'with +98.4 ml (95% CI 10.9 to 185.9)'Typo: 'with and'.
- MINORgrammarResults, Pharmacokinetics“indicates an increased exposure to rentosertib with increases time on treatment”→ Change to 'with increased time on treatment'Ungrammatical phrase.
- MINORconsistencyResults, Primary safety endpoint“3 in 60 mg QD (20.4%)”→ Change percentage to 16.7% (3/18)3/18 = 16.7%, not 20.4%.
- MINORconsistencyExtended Data Fig. 5“n = 10 at week 10 in 60 QD”→ Change to 'n = 10 at week 12 in 60 QD'Timepoint inconsistency appearing in both panels a and b.
- MINORconsistencyResults, secondary lung function endpoints“forced expiry in 1 s”→ Change to 'forced expiratory volume in 1 second' (FEV1) for consistency with standard terminology.The term 'forced expiry' is used once; elsewhere 'FEV1' is used correctly.
- MINORclarityResults, secondary lung function endpoints“patients receiving 60 mg rentosertib QD not concurrently taking SOC antifibrotic therapy exhibited significant improvement”→ Add 'who were' before 'not concurrently' for clarity: 'patients receiving 60 mg rentosertib QD who were not concurrently taking SOC antifibrotic therapy'.Minor grammatical improvement.
- MINORotherTable 2 and Results text“ALT increase: 6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%)”→ Verify and correct the numbers for ALT increase in the 60 mg QD group; the text gives two different counts (6 and 3) for the same group, while the table shows 1. This is a data inconsistency that needs resolution.See integrity_concerns for more details.
- MINORpunctuationResults, secondary lung function endpoints“with and +98.4 ml”→ Remove 'and': 'with +98.4 ml' or 'with an increase of +98.4 ml'.Typo: 'with and' should be 'with'.
- 1.MEDIUMrigorCorrect the TEAE percentage error in Results: change '3 in 60 mg QD (20.4%)' to '3 in 60 mg QD (16.7%)' (3/18 = 16.7%).
- 2.MEDIUMrigorReconcile the ALT increase listing in Results: '6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%)' duplicates the same arm; determine which count belongs to another group or clarify the subcategories.
- 3.MEDIUMrigorName the central/lead IRB and each participating center's ethics committee, and include protocol approval numbers, in Methods so the ethics approval is auditable.
- 4.MEDIUMrigorState the randomization sequence generation method (e.g., computer-generated random numbers, block/stratification) in Methods, in addition to 'interactive response technology'.
- 5.MEDIUMrigorReplace the safety-feasibility sample size note with a formal power/sample-size statement for the primary safety endpoint, including alpha and power, or clearly label it as a feasibility calculation in the protocol.
- 6.MEDIUMrigorReport exact p-values or effect estimates with CIs for all inferential tests, including the Spearman/Pearson correlations in figure legends, instead of 'P < 0.05' thresholds.
- 7.MEDIUMrigorAdd versions for R, lme4, clusterProfiler, and the Olink panel in Methods, and provide a software environment file (renv/conda) in the GitHub repository.
- 8.MEDIUMrigorExplicitly cite the CONSORT extension for pilot/phase 2 trials in the reporting summary, rather than a generic 'Reporting Summary' link.
- 9.MEDIUMrigorAdd a Funding statement with grant numbers and sponsor support in the Acknowledgements/Funding section; competing interests are already declared.
- 10.MEDIUMrigorTone down the abstract's 'safe and effective' to 'safe and well tolerated, with a preliminary efficacy signal', since the primary endpoint was safety and the FVC result was exploratory.
- 11.MEDIUMrigorAdd a formal a priori power/sample-size calculation (with effect size, alpha, power) for the primary safety endpoint, or state explicitly that the trial is feasibility-only; edit the Methods 'Statistical analysis' and 'Study design' sections.
- 12.MEDIUMrigorReport exact p-values (2-3 significant figures) instead of thresholds such as 'P < 0.05' and 'P adj < 0.05' throughout the Results (proteomics and correlation analyses).
- 13.MEDIUMrigorResolve the contradiction in ALT-increase counts between the Results narrative ('6 in 60 mg QD (33.3%) and 3 in 60 mg QD (16.7%)') and Table 2 (ALT increase row: 0, 0, 1, 0); correct the duplicated 60 mg QD entry.
- 14.MEDIUMrigorCorrect the hypokalemia percentage in the Results AE narrative: '3 in 60 mg QD (20.4%)' should be 16.7% (3/18).
- 15.MEDIUMrigorSpecify versions of the statistical software/packages used (R, lme4, clusterProfiler, lm) in the Methods 'Statistical analysis' section.
- 16.MEDIUMrigorSoftening the abstract's 'safe and effective' to 'safe and potentially effective' to match the phase 2a secondary-endpoint trend evidence.
- 17.MEDIUMrigorFix the typo 'with and +98.4 ml' in the Results 'Secondary lung function endpoints' section.
- 18.MEDIUMrigorCorrect 'n = 10 at week 10 in 60 QD' to 'week 12' in Extended Data Fig. 5 (appears twice).
- 19.MEDIUMrigorAdd version numbers for all software (R, lme4, clusterProfiler, etc.) in the Methods section to improve key_resources reporting.
- 20.MEDIUMrigorExplicitly state adherence to the CONSORT 2010 statement for randomized trials and provide a completed checklist as supplementary material to enhance reporting_transparency.
- 21.MEDIUMrigorProvide a formal sample size calculation for efficacy endpoints (e.g., FVC) based on expected effect size and power, not just for safety, to strengthen study_design.
- 22.MEDIUMrigorClarify the potential impact of partial unblinding by bioanalytics staff who could identify placebo samples; discuss how this was managed to avoid bias.
- 23.MEDIUMrigorDeposit the study protocol and statistical analysis plan in a public repository (e.g., ClinicalTrials.gov or a data repository) rather than only on request, to improve data availability.
- 24.MEDIUMrigorIn the limitations, explicitly discuss the generalizability of findings given the single-country, all-Asian, predominantly male cohort, and the potential impact on external validity.
- 25.MEDIUMrigorFor the claim of 'revolutionary shift' in drug discovery, consider toning down to avoid overstatement, as a single successful phase 2a trial does not yet constitute a paradigm shift.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.