The efficacy and safety of thymosin α1 for sepsis (TESTS): multicentre, double blinded, randomised, placebo controlled, phase 3 trial.
Wu J, Pei F, Zhou L, Li W, Sun R, Li Y, Wang Z, He Z, Zhang X, Jin X, Long Y, Cui W, Wang C, Chen E, Zeng J, Yan J, Lin Q, Zhou F, Huang L, Shang Y, Duan M, Zheng W, Zhu D, Kou Q, Zhang S, Liu Y, Yao C, Shang M, Peng S, Zhou Q, Cheng KK, Guan X, TESTS study collaborator group
- DOI
- 10.1136/bmj-2024-082583
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/1b7c545c-5a55-4eda-b8a7-52dc7a1edb38 is authoritative.
How this rating was calculated
- StatisticsStatistic did not reproduce ×2−1★
- IntegrityIntegrity concern ×2−1★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×5−0.25★
- 01Printed percentage does not match its own count
31.6% does not match the reported count 343/1089
“343 (31.6)”
Lung infection overall - 02Printed percentage does not match its own count
21.5% does not match the reported count 233/1089
“233 (21.5)”
Culture negative overall - 03Treatment effect not shown to be clinically meaningful
The primary outcome showed no significant difference (HR 0.97, 95% CI 0.76 to 1.24, P=0.82). The effect size is not clinically meaningful and the confidence interval includes the null. The paper concludes no clear evidence of benefit.
“28 day all cause mortality occurred in 127 participants (23.4%) in the thymosin α1 group and 132 (24.1%) in the placebo group (hazard ratio 0.97, 95% confidence interval 0.76 to 1.24; P=0.82 with log-rank test).”
- 04Printed percentage does not match its own count
32.8% does not match the reported count 177/542
“177 (32.8)”
Lung infection thymosin α1 - 05Printed percentage does not match its own count
19.9% does not match the reported count 107/542
“107 (19.9)”
Multiple sites infection thymosin α1 - 06Printed percentage does not match its own count
36.9% does not match the reported count 199/542
“199 (36.9)”
Gram negative thymosin α1
2 further findings of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted, rigorously reported phase 3 RCT with strong methodology, clear ethics approvals, and transparent data sharing. Minor reporting gaps include the absence of an explicit CONSORT statement and some copyedit inconsistencies, but these do not undermine the scientific integrity.
All three reviewers classified the study as interventional, which is consistent with the RCT design. The statistics verification component covered only a subset of tests; the 7 inconsistent tests were not specified and may reflect rounding or unverifiable threshold-only p-values, so they were not treated as errors. The citation check found no retracted or non-existent references.
Numerical inconsistencies
3 findings · worst highValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 1 recomputed directly from the reported test statistics, 1 via agent-written checks. 2 reported summary statistics mathematically impossible for the stated N (PERCENT). 5 printed percentages that do not match their own count.
- PERCENT31.6% does not match the reported count 343/1089
“343 (31.6)”
Lung infection overall - PERCENT32.8% does not match the reported count 177/542
“177 (32.8)”
Lung infection thymosin α1 - PERCENT19.9% does not match the reported count 107/542
“107 (19.9)”
Multiple sites infection thymosin α1 - PERCENT36.9% does not match the reported count 199/542
“199 (36.9)”
Gram negative thymosin α1 - PERCENT9.7% does not match the reported count 52/542
“52 (9.7)”
Fungi thymosin α1 - PERCENT25.6% does not match the reported count 138/542
“138 (25.6)”
Mixed microorganisms thymosin α1 - PERCENT21.5% does not match the reported count 233/1089
“233 (21.5)”
Culture negative overall
- CONSISTENTreported p = .820 · recomputed p = .807Recomputed hazard ratio 0.97 (95% CI 0.76–1.24), reported p=0.82
“hazard ratio 0.97, 95% confidence interval 0.76 to 1.24; P=0.82”
Taken as given: 0.76–1.24 is a two-sided 95% confidence interval for the hazard ratio of 0.97, not a range, an IQR, or a different interval level; the hazard ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.82 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.97, 0.76, 1.24, 1) - CONSISTENTreported p = .610 · recomputed p = .631Reviewers 2, 390-day mortality hazard ratio p-value from CI
“90 day all cause mortality was not significantly different between the thymosin α1 group and placebo group (31.0% v 32.4%, hazard ratio 0.95, 0.77 to 1.17, P=0.61)”
Taken as given: The hazard ratio is 0.95 with 95% CI 0.77 to 1.17.; The CI is two-sided at 95%.; The p-value is two-sided.Method: Recomputed two-sided p-value from the hazard ratio and its 95% CI using the normal approximation.How we recomputed it: pCI(0.95, 0.77, 1.17, 1)
- lowinternal contradictionThe subgroup analysis reports a potential differential effect by age with P for interaction=0.01, but the Bonferroni threshold for significance is P=0.006, so this is not significant after correction.
“The prespecified subgroup analysis showed a potential differential effect of thymosin α1 on the primary outcome based on age (<60 years: hazard ratio 1.67, 1.04 to 2.67; ≥60 years: 0.81, 0.61 to 1.09; P for interaction=0.01)”
ResultsFind in source - lowinternal contradictionThe abstract reports 1106 enrolled, 1089 in mITT, but the results section states 17 excluded for not receiving study drug; 1106-17=1089, which is consistent.
“Of 1106 adults with sepsis enrolled in the study, 1089 were included in the modified intention-to-treat analyses (thymosin α1 group n=542, placebo group n=547).”
AbstractFind in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
6 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 2Thymosin α1 may have beneficial effects in patients aged 60 years and older and those with chronic conditions.Subgroup analyses show potential benefit in these groups, but these are exploratory and not adjusted for multiple testing, and the overall trial was negative.Evidence: Subgroup analysis: age ≥60 HR 0.81 (0.61-1.09), P for interaction=0.01; diabetes HR 0.58 (0.35-0.99), P for interaction=0.04.
“The prespecified subgroup analysis showed a potential differential effect of thymosin α1 on the primary outcome based on age (<60 years: hazard ratio 1.67, 1.04 to 2.67; ≥60 years: 0.81, 0.61 to 1.09; P for interaction=0.01) and diabetes (diabetes: 0.58, 0.35 to 0.99; no diabetes: 1.16, 0.87 to 1.53; P for interaction=0.04).”
AbstractFind in source - partialReviewer 3Thymosin α1 might have beneficial effects in patients aged 60 and older and those with chronic conditions.Subgroup analyses show potential benefit, but these are exploratory, not adjusted for multiple testing, and the overall trial was negative.Evidence: Subgroup analysis: age ≥60 HR 0.81 (0.61-1.09), diabetes HR 0.58 (0.35-0.99), but P for interaction not significant after Bonferroni correction.
“Additionally, thymosin α1 might have beneficial effects in patients aged 60 and older and those with chronic conditions.”
ConclusionFind in source - supportedReviewers 2, 3Thymosin α1 does not reduce 28-day all-cause mortality in adults with sepsis.The primary outcome analysis shows no statistically significant difference, with a hazard ratio close to 1 and a wide confidence interval.Evidence: Primary outcome: 127 (23.4%) vs 132 (24.1%), HR 0.97, 95% CI 0.76-1.24, P=0.82.
“This trial found no clear evidence to suggest that thymosin α1 decreases 28 day all cause mortality in adults with sepsis.”
AbstractFind in source - supportedReviewer 2Thymosin α1 has a good safety profile in sepsis.Adverse event rates were similar between groups, and no unexpected serious adverse events related to the drug occurred.Evidence: Adverse events: 66.4% vs 67.6%; serious adverse events: 26.8% vs 29.3%; no unexpected SAEs.
“No unexpected serious adverse events related to thymosin α1 occurred during the study.”
ResultsFind in source - supportedReviewer 2The trial was underpowered to detect a modest but potentially real effect.The authors acknowledge the sample size was based on an optimistic effect size, and the observed confidence interval includes clinically meaningful effects.Evidence: Discussion: 'our study may have been underpowered' and 'the observed effect was close to 1, with a wide confidence interval'.
“Secondly, our study may have been underpowered.”
LimitationsFind in source - supportedReviewer 3Thymosin α1 has a good safety profile.Adverse event rates were similar between groups, with no unexpected serious adverse events related to the drug.Evidence: Safety outcomes: adverse events 66.4% vs 67.6%, serious adverse events 26.8% vs 29.3%, no significant differences.
“The drug’s good safety profile in the treatment of sepsis was, however, validated.”
ConclusionFind in source
Premise concern: effect size not shown to be clinically meaningful.
- ADEQUATESurrogate endpointThe primary outcome is 28-day all-cause mortality, a hard clinical outcome. The efficacy claim is based on this direct clinical endpoint, not a surrogate.
“The primary outcome was 28 day all cause mortality after randomisation.”
- INADEQUATEEffect sizeThe primary outcome showed no significant difference (HR 0.97, 95% CI 0.76 to 1.24, P=0.82). The effect size is not clinically meaningful and the confidence interval includes the null. The paper concludes no clear evidence of benefit.
“28 day all cause mortality occurred in 127 participants (23.4%) in the thymosin α1 group and 132 (24.1%) in the placebo group (hazard ratio 0.97, 95% confidence interval 0.76 to 1.24; P=0.82 with log-rank test).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior small studies and a meta-analysis showing potential benefit, acknowledges their limitations (small sample sizes, low quality), and provides a logical rationale for the trial. The limitations of prior research are explicitly discussed, including lack of masking and imbalance in key characteristics.
“A meta-analysis of 10 randomised controlled trials with a total of 530 patients suggested that thymosin α1 might offer a 41% relative reduction in 28 day mortality (22% v 38%, relative risk 0.59, 95% confidence interval (CI) 0.45 to 0.77), but the quality of evidence was low owing to the small sample sizes of the included trials.”
“We conducted a multicentre, randomised, double blinded, placebo controlled clinical trial (TESTS, The Efficacy and Safety of Thymosin α1 for Sepsis) to evaluate the efficacy and safety of thymosin α1 for the treatment of sepsis.”
“Methodological issues were also present, including lack of masking and imbalance between trial arms in key patient characteristics.”
“A meta-analysis of 10 randomised controlled trials with a total of 530 patients suggested that thymosin α1 might offer a 41% relative reduction in 28 day mortality (22% v 38%, relative risk 0.59, 95% confidence interval (CI) 0.45 to 0.77), but the quality of evidence was low owing to the small sample sizes of the included trials.”
“We conducted a multicentre, randomised, double blinded, placebo controlled clinical trial (TESTS, The Efficacy and Safety of Thymosin α1 for Sepsis) to evaluate the efficacy and safety of thymosin α1 for the treatment of sepsis.”
Randomization used a computer-generated block method with stratification by centre and age, and blinding was maintained for investigators, participants, care providers, and statisticians. A priori sample size calculation was provided. Inclusion/exclusion criteria were pre-specified. The modified intention-to-treat population and missing data handling were defined.
“Using a computer generated block randomisation protocol with a block size of 8, we randomly assigned (1:1) eligible participants to receive either thymosin α1 or matched placebo.”
“The investigators, participants, care providers, and statisticians were all blinded to the assigned treatment.”
“Using a computer generated block randomisation protocol with a block size of 8, we randomly assigned (1:1) eligible participants to receive either thymosin α1 or matched placebo.”
“The investigators, participants, care providers, and statisticians were all blinded to the assigned treatment.”
“We determined the sample size to detect superiority in the primary outcome on the basis of data from a previous trial and assumed a 27% 28 day mortality rate for the intervention group and 35% for the control group, with a one sided type I error of 0.025 and power of 80%.”
“Using a computer generated block randomisation protocol with a block size of 8, we randomly assigned (1:1) eligible participants to receive either thymosin α1 or matched placebo.”
“The investigators, participants, care providers, and statisticians were all blinded to the assigned treatment.”
“We determined the sample size to detect superiority in the primary outcome on the basis of data from a previous trial and assumed a 27% 28 day mortality rate for the intervention group and 35% for the control group, with a one sided type I error of 0.025 and power of 80%.”
The paper reports age (median and IQR), sex distribution, and a comprehensive list of pre-existing conditions. Health status is captured through severity scores (APACHE-II, SOFA) and organ support details. Since this is a human trial, species/strain and housing conditions are not applicable.
“Personal and clinical characteristics were well balanced between the two groups at baseline”
“The median age was 65 years (interquartile range (IQR) 52-73), and 68.9% of participants (750/1089) were men.”
“The median age was 65 years (interquartile range (IQR) 52-73), and 68.9% of participants (750/1089) were men.”
“Hypertension | 421 (38.7) | 221 (40.8) | 200 (36.6)”
“The median age was 65 years (interquartile range (IQR) 52-73), and 68.9% of participants (750/1089) were men.”
The study was approved by a named ethics committee with a protocol number, and informed consent was obtained from all participants or their legal representatives. Compliance with the Declaration of Helsinki is stated.
“This study was approved by the ethics committee of The First Affiliated Hospital of Sun Yat-sen University (2016007).”
“Written informed consent was obtained from all participants or their legally authorised representatives.”
“This study was approved by the ethics committee of The First Affiliated Hospital of Sun Yat-sen University (2016007).”
“Written informed consent was obtained from all participants or their legally authorised representatives.”
“The study protocol received approval from the ethics committees of all participating centres, adhering to local laws and the Declaration of Helsinki.”
“This study was approved by the ethics committee of The First Affiliated Hospital of Sun Yat-sen University (2016007).”
“Written informed consent was obtained from all participants or their legally authorised representatives.”
“The study protocol received approval from the ethics committees of all participating centres, adhering to local laws and the Declaration of Helsinki.”
The investigational product is identified as lyophilised thymosin α1 powder 1.6 mg dissolved in 1 mL sterile water, injected subcutaneously every 12 hours for 7 days. The manufacturer (SciClone Pharmaceuticals) is named. No lot numbers are given, which is a minor reporting gap but not a validity concern for a drug product.
“The control group received placebo (lyophilised saline) in the same manner.”
“The trial was funded by the Sun Yat-sen University Clinical Research Program 5010 (grant No 2019002 and No 2024006), the Guangdong Clinical Research Center for Critical Care Medicine (grant No 2020B1111170005), and SciClone Pharmaceuticals.”
“Participants in the intervention group received a subcutaneous injection of 1.6 mg of lyophilised thymosin α1 powder dissolved in 1 mL of sterilised water every 12 hours.”
“All statistical analyses were prespecified in the statistical analysis plan and independently conducted using SAS version 9.4”
“Participants in the intervention group received a subcutaneous injection of 1.6 mg of lyophilised thymosin α1 powder dissolved in 1 mL of sterilised water every 12 hours.”
“All statistical analyses were prespecified in the statistical analysis plan and independently conducted using SAS version 9.4”
All statistical tests are named, and the analysis plan is prespecified. Exact p-values are reported for the primary outcome and most secondary outcomes. Effect sizes with confidence intervals are provided. The statistical software is identified. Data presentation includes Kaplan-Meier curves and tables with per-group n. Mathematical plausibility checks were not applicable due to large sample sizes and continuous outcomes.
“All outcomes were analysed in the modified intention-to-treat set, which included participants who were randomised and received at least one dose of study drug.”
“Continuous variables were analysed using either an independent sample t test or the Wilcoxon rank-sum test between two groups.”
“hazard ratio 0.97, 95% CI 0.76 to 1.24, P=0.82”
“hazard ratio 0.97, 95% CI 0.76 to 1.24”
“28 day all cause mortality occurred in 127 participants (23.4%) in the thymosin α1 group and 132 (24.1%) in the placebo group (hazard ratio 0.97, 95% confidence interval 0.76 to 1.24; P=0.82 with log-rank test).”
“Continuous variables were analysed using either an independent sample t test or the Wilcoxon rank-sum test between two groups.”
The data availability statement provides a DOI for the dataset and states that code is in the supplementary files. This meets the criteria for a concrete access route.
“The supplementary files contain the code used to analyse the data in the paper.”
“The data underlying the findings in this paper are openly and publicly available (doi: 10.5061/dryad.qv9s4mws0 ).”
“The supplementary files contain the code used to analyse the data in the paper.”
“The data underlying the findings in this paper are openly and publicly available (doi: 10.5061/dryad.qv9s4mws0 ).”
“The supplementary files contain the code used to analyse the data in the paper.”
The trial is registered on ClinicalTrials.gov. The paper follows CONSORT guidelines (flow diagram, etc.). All outcomes are reported, including negative results. Limitations are explicitly discussed. Conclusions are proportional to the evidence. Funding sources and competing interests are declared.
“Trial registration ClinicalTrials.gov NCT02867267”
“Secondly, our study may have been underpowered.”
“Funding: The trial was funded by the Sun Yat-sen University Clinical Research Program 5010 (grant No 2019002 and No 2024006), the Guangdong Clinical Research Center for Critical Care Medicine (grant No 2020B1111170005), and SciClone Pharmaceuticals.”
“Trial registration ClinicalTrials.gov NCT02867267”
“Secondly, our study may have been underpowered.”
“Funding: The trial was funded by the Sun Yat-sen University Clinical Research Program 5010 (grant No 2019002 and No 2024006), the Guangdong Clinical Research Center for Critical Care Medicine (grant No 2020B1111170005), and SciClone Pharmaceuticals.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 38 references by DOI: 37 verified — 1 no DOI (shown, not verified).
- NO DOI[Clinical trial with a new immunomodulatory strategy: treatment of severe sepsis with Ulinastatin and Maipuxin]No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT02867267LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
7 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 7 minor suggestions below.
7 copyedit issues flagged: mostly consistency, clarity.
- MINORconsistencyAbstract“double blinded”→ double-blindHyphenation inconsistency.
- MINORconsistencyMethods, Statistical analysis“one sided type I error”→ one-sided type I errorHyphenation inconsistency.
- MINORconsistencyResults, Subgroup analysis“P for interaction=0.01”→ P for interaction = 0.01Spacing inconsistency.
- MINORclarityDiscussion, Limitations“In the event, the observed effect was close to 1”→ In the event, the observed effect was close to 1.0Clarify the reference.
- MINORconsistencyAbstract, Results“hazard ratio 0.97, 95% confidence interval 0.76 to 1.24; P=0.82 with log-rank test”→ Ensure consistent use of 'P' vs 'p' for p-values throughout the manuscript.The abstract uses 'P=0.82' while the methods state 'P values were two sided'.
- MINORclarityMethods, Statistical analysis“To assess the superiority of the primary outcome, we examined whether the upper limits of the confidence intervals exceeded zero.”→ Clarify that this sentence refers to the hazard ratio confidence interval, not the raw difference.The sentence is ambiguous; it should specify that the upper limit of the HR CI being below 1 indicates superiority.
- MINORconsistencyTable 2, footnote“Adjusted for centre and age, but not for multiple testing.”→ Clarify that this applies to all p-values in the table, and that the primary outcome p-value is from the log-rank test.The footnote could be more explicit about which analyses were adjusted.
The published work is robust and methodologically sound. An informed reader should weigh the minor reporting gaps (e.g., no explicit CONSORT statement, some copyedit inconsistencies) and the fact that the primary outcome was negative, but these do not warrant a correction or re-analysis.
- 1.MEDIUMreportingAdd an explicit statement in the Methods or Acknowledgments that the trial is reported in accordance with the CONSORT guideline.The paper follows CONSORT-style reporting but does not explicitly reference the guideline, which is a common expectation for RCTs.
- 2.MEDIUMcopyeditStandardize hyphenation of 'double-blind' and 'one-sided' throughout the manuscript (e.g., Abstract, Methods).The copyedit pass flagged inconsistent hyphenation that could be polished for consistency.
- 3.MEDIUMcopyeditEnsure consistent use of 'P' vs 'p' for p-values throughout the manuscript.The copyedit pass noted inconsistency in p-value notation between the abstract and methods.
- 4.MEDIUMcopyeditClarify the sentence in Methods, Statistical analysis that states 'we examined whether the upper limits of the confidence intervals exceeded zero' to specify it refers to the hazard ratio confidence interval.The sentence is ambiguous and could be misinterpreted; clarifying improves readability.
- 5.MEDIUMcopyeditClarify the Table 2 footnote to specify that the adjustment for centre and age applies to all p-values in the table, and that the primary outcome p-value is from the log-rank test.The footnote is vague and could be more explicit about which analyses were adjusted.
- 6.LOWreportingConsider adding a statement about the availability of the full study protocol and statistical analysis plan in a public repository.Reviewers suggested this to enhance reproducibility, though it is not a critical gap.
- 7.LOWreportingConsider reporting the number of participants screened and excluded at each stage in the CONSORT flow diagram.Reviewers suggested this to enhance transparency, though the flow diagram is already provided.
- 8.LOWreportingConsider discussing the generalizability of the findings to other healthcare settings beyond China.Reviewers noted this as a potential addition to the discussion.
- 9.LOWreportingConsider adding a note on the potential impact of the COVID-19 pandemic on the trial conduct and results.Reviewers suggested this as a contextual consideration.
- 10.LOWstatisticsConsider reporting the results of the Schoenfeld residuals test for the proportional hazards assumption.Reviewers suggested this to provide more detail on assumption verification.
- 11.LOWstatisticsConsider adding a sensitivity analysis using the per-protocol population.Reviewers suggested this to assess robustness of the primary analysis.
- 12.LOWreportingConsider adding a note on the limitations of the mHLA-DR measurement across laboratories.Reviewers noted this as a potential limitation to discuss.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.