Effects of a personalized nutrition program on cardiometabolic health: a randomized controlled trial.
Bermingham KM, Linenberg I, Polidori L, Asnicar F, Arrè A, Wolf J, Badri F, Bernard H, Capdevila J, Bulsiewicz WJ, Gardner CD, Ordovas JM, Davies R, Hadjigeorgiou G, Hall WL, Delahanty LM, Valdes AM, Segata N, Spector TD, Berry SE
- DOI
- 10.1038/s41591-024-02951-6
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/20901ffa-61e0-456d-aac8-8a8fd8b0bb78 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on changes in serum triglycerides (TG) and LDL cholesterol, which are surrogate biomarkers for cardiovascular outcomes. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking these surrogates to hard clinical outcomes. The effect on LDL-C was not significant, and the TG reduction, while statistically significant, is a biomarker change presented as proof of clinical benefit.
“Primary outcomes were serum low-density lipoprotein cholesterol and TG concentrations at baseline and at 18 weeks.”
- 02Treatment effect not shown to be clinically meaningful
The primary reported effect is a small reduction in triglycerides (mean difference = -0.13 mmol/L) and no significant change in LDL-C. The TG reduction is a small fraction of the baseline value (1.35 mmol/L) and is not anchored to a minimal clinically important difference or hard clinical outcome. The weight loss (2.46 kg) is below the 5% threshold considered clinically meaningful, as acknowledged in the discussion.
“The mean difference in changes between the groups was −0.13 mmol l−1 (log-transformed, 95% CI = −0.07 to −0.01, P = 0.016)”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported RCT with strong methodological rigor across most dimensions. The main weakness is the lack of a public code repository and incomplete data accession details, which limits full reproducibility.
Both reviewers independently scored all eight dimensions and agreed on all statuses; no divergence was present. The study is a human interventional trial, so animal-related and cell-line criteria were marked not applicable. Statistical verification was limited to one recomputed test; other statistics were not machine-verified.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks.
- CONSISTENTreported p = .521 · recomputed p = .514Reviewers 1, 2Primary outcome LDL-C between-group difference p-value from CI
“Differences in LDL-C concentrations between groups were not significant: −0.04 mmol l −1 (95% CI = −0.16 to 0.08, P = 0.521”
Taken as given: The estimate is the mean difference in changes (-0.04).; The 95% CI is on the original scale.; The p-value is two-sided.Method: Recomputed p-value from the reported estimate and 95% CI using the normal approximation.How we recomputed it: pCI(-0.04, -0.16, 0.08, 0)
- lowinternal contradictionThe abstract states 'n = 347' randomized, but the CONSORT diagram and results mention 347 randomized; however, the per-protocol analysis includes 225 participants, and the microbiome analysis includes 118 and 112 for control and PDP, respectively. These numbers are consistent with expected dropouts and missing data, so not a concern.
Participants ( n = 347), aged 41–70 years ... were randomized to the PDP ( n = 177) or control ( n = 170).
Abstractreviewer’s wording
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
7 major claims checked against the paper's own evidence: 3 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1The PDP improved gut microbiome composition (beta-diversity) compared to control.Beta-diversity differed between groups at week 18 (KSp=0.04), and favorable species increased in PDP, but the clinical significance is unclear and the analysis was exploratory.Evidence: Bray-Curtis dissimilarity KSp=0.04; favorable species summed abundance change 0.48 vs -0.73 (MWWp=0.015).
“Comparing beta-diversity dissimilarities across the control and PDP groups at week 18 showed a statistically significant difference (Kolmogorov–Smirnov stochasticity parameter, KSp = 0.04).”
ResultsFind in source - partialReviewers 1, 2The PDP reduced LDL-C in highly adherent participants.Subgroup analysis showed significant LDL-C reduction in highly adherent PDP vs low adherent, but the primary analysis showed no overall effect, and the subgroup was not pre-specified.Evidence: LDL-C change -0.20 vs 0.07 mmol/L (P=0.019) in high vs low adherence.
“greater reductions in LDL-C (−0.20 ± 0.48 versus 0.07 ± 0.56 mmol l −1 , P = 0.019)”
ResultsFind in source - partialReviewer 2The PDP improved gut microbiome composition, specifically increasing favorable species.The paper shows a significant difference in beta-diversity and an increase in some favorable species, but the clinical relevance is unclear and the effect on unfavorable species was not significant.Evidence: Beta-diversity KSp=0.04; 8 of 15 favorable species increased in PDP vs none in control; summed abundance change 0.48 vs -0.73 (MWWp=0.015).
“Notably, among the 15 favorable species, we found eight species in the PDP group showing an increase in terms of relative abundance at the endpoint”
ResultsFind in source - supportedReviewers 1, 2The personalized dietary program led to significant improvements in cardiometabolic health compared to standard dietary advice.The primary outcome (TG) showed a significant reduction, and several secondary outcomes (weight, waist circumference, HbA1c, diet quality) improved significantly.Evidence: Primary outcome TG mean difference -0.13 mmol/L (P=0.016); secondary outcomes weight -2.46 kg, waist -2.35 cm, HbA1c -0.05%, HEI +7.08.
“Following a personalized diet led to some improvements in cardiometabolic health compared to standard dietary advice.”
AbstractFind in source - supportedReviewers 1, 2The PDP led to greater reductions in body weight and waist circumference than control.Both weight and waist circumference showed significant between-group differences with CIs excluding zero.Evidence: Weight difference -2.46 kg (95% CI -3.67 to -1.25); waist -2.35 cm (95% CI -4.07 to -0.63).
“Reductions in body weight, waist circumference and glycated hemoglobin (HbA1c), and increases in diet quality (HEI score), were significantly greater after the PDP than with the control diet”
ResultsFind in source - supportedReviewer 1The PDP improved diet quality (HEI score) compared to control.HEI score increased significantly more in PDP group.Evidence: HEI difference 7.08 (95% CI 5.02 to 9.15).
“diet quality (HEI score): 7.08 (95% CI = 5.02 to 9.15)”
ResultsFind in source - supportedReviewer 2The PDP improved subjective feelings of energy, sleep, mood, and hunger.Self-reported improvements were significantly higher in the PDP group.Evidence: 43% vs 11% energy, 35% vs 9% sleep, 33% vs 15% mood, 22% vs 14% hunger (P<0.01).
“a greater proportion of PDP participants reported improvements in energy level (43% versus 11%), sleep quality (35% versus 9%), general mood (33% versus 15%) and reduced hunger levels (22% versus 14%) compared with controls ( P < 0.01 for all)”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on changes in serum triglycerides (TG) and LDL cholesterol, which are surrogate biomarkers for cardiovascular outcomes. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking these surrogates to hard clinical outcomes. The effect on LDL-C was not significant, and the TG reduction, while statistically significant, is a biomarker change presented as proof of clinical benefit.
“Primary outcomes were serum low-density lipoprotein cholesterol and TG concentrations at baseline and at 18 weeks.”
- INADEQUATEEffect sizeThe primary reported effect is a small reduction in triglycerides (mean difference = -0.13 mmol/L) and no significant change in LDL-C. The TG reduction is a small fraction of the baseline value (1.35 mmol/L) and is not anchored to a minimal clinically important difference or hard clinical outcome. The weight loss (2.46 kg) is below the 5% threshold considered clinically meaningful, as acknowledged in the discussion.
“The mean difference in changes between the groups was −0.13 mmol l−1 (log-transformed, 95% CI = −0.07 to −0.01, P = 0.016)”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
3 integrity concerns flagged (0 high).
- lowotherThe paper reports a significant improvement in subjective energy, sleep, mood, and hunger, but these are self-reported and not pre-specified primary outcomes; the magnitude of differences (e.g., 43% vs 11%) seems large for a dietary intervention, but not impossible.
“a greater proportion of PDP participants reported improvements in energy level (43% versus 11%), sleep quality (35% versus 9%), general mood (33% versus 15%) and reduced hunger levels (22% versus 14%) compared with controls ( P < 0.01 for all)”
ResultsFind in source - lowotherThe paper reports a significant difference in beta-diversity between groups using a Kolmogorov–Smirnov stochasticity parameter (KSp = 0.04), but the clinical relevance of this microbiome change is not established.
“Comparing beta-diversity dissimilarities across the control and PDP groups at week 18 showed a statistically significant difference (Kolmogorov–Smirnov stochasticity parameter, KSp = 0.04).”
ResultsFind in source
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites observational research and prior RCTs on personalized nutrition, acknowledges limitations of single-axis personalization, and hypothesizes that a multilevel approach will improve efficacy. The rationale is well-linked to the study objectives. Limitations of prior work are addressed by the multilevel design.
“Observational research supports the application of personalized nutrition , but there are few randomized controlled trials designed to test the efficacy of personalized nutrition programs compared to standard dietary advice on health outcomes.”
“Therefore, we hypothesized that a multilevel approach to personalization encompassing multiple factors contributing to intraindividual and interindividual variability in nutritional responses to diet will improve the efficacy of advice to elicit a meaningful impact on health outcomes.”
“Personalized nutrition approaches and corresponding studies typically use a single axis of personalization but reported low correlations between biomarkers, for example, triglycerides (TGs) and glucose, suggesting that a prediction algorithm using a multilevel approach to personalization may yield superior results.”
“Observational research supports the application of personalized nutrition , but there are few randomized controlled trials designed to test the efficacy of personalized nutrition programs compared to standard dietary advice on health outcomes.”
“Therefore, we hypothesized that a multilevel approach to personalization encompassing multiple factors contributing to intraindividual and interindividual variability in nutritional responses to diet will improve the efficacy of advice to elicit a meaningful impact on health outcomes.”
“Personalized nutrition approaches and corresponding studies typically use a single axis of personalization but reported low correlations between biomarkers, for example, triglycerides (TGs) and glucose, suggesting that a prediction algorithm using a multilevel approach to personalization may yield superior results.”
Randomization was performed using a minimization program (MinimPy) with stratification factors. The unit of randomization is the individual participant. Blinding is partially addressed: the between-group analysis was performed by a blinded researcher, and group allocation was concealed. A power analysis is reported (n=150 per group, 90% power, P<0.05). Inclusion/exclusion criteria are detailed. Outlier handling is addressed through the ITT and per-protocol analyses and missing-data approaches. Controls are appropriate (USDA dietary advice). Independent replication is not applicable for a single pivotal trial.
“a minimization-randomization program (MinimPy v.0.3, Python Package Index; pypi.org/project/MinimPy/ (https://pypi.org/project/MinimPy/) ) was used for treatment allocation.”
“The between-group analysis was performed by a blinded researcher. Group allocation was concealed by labeling the groups with nonidentifying terms.”
“The study was powered on a sample size of 150 participants per group ( n = 300) at 90% power and P < 0.05, to detect a 0.21 mmol l −1 between-group difference in TG (endpoint change from baseline).”
“a minimization-randomization program (MinimPy v.0.3, Python Package Index; pypi.org/project/MinimPy/ (https://pypi.org/project/MinimPy/) ) was used for treatment allocation.”
“The study was powered on a sample size of 150 participants per group ( n = 300) at 90% power and P < 0.05, to detect a 0.21 mmol l −1 between-group difference in TG (endpoint change from baseline).”
“The between-group analysis was performed by a blinded researcher. Group allocation was concealed by labeling the groups with nonidentifying terms.”
Sex is reported (86% female). Age, BMI, and health status are reported. Demographics include ethnicity, education, and menopausal status. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing are not applicable for a human trial.
“In total, 86% of participants were female”
“mean ± s.d. age of 52 ± 7.5 years, body mass index (BMI) of 34 ± 5.8 kg m − 2”
“In total, 86% of participants were female”
“mean ± s.d. age of 52 ± 7.5 years, body mass index (BMI) of 34 ± 5.8 kg m − 2”
The paper states ethical approval was obtained through Advarra IRB with protocol number, and all participants provided written informed consent. Regulatory compliance with good clinical practice and the Declaration of Helsinki is stated.
“Ethical approval for the trial was obtained through the Advarra IRB (IRB no. 00000971; protocol no. 00044316).”
“All participants provided written informed consent”
“the study was carried out in accordance with good clinical practice and the Declaration of Helsinki (2013).”
“Ethical approval for the trial was obtained through the Advarra IRB (IRB no. 00000971; protocol no. 00044316).”
“All participants provided written informed consent”
“the study was carried out in accordance with good clinical practice and the Declaration of Helsinki (2013).”
The investigational product (PDP) is described in detail, including the ZOE 2022 algorithm and app. Software tools are identified with versions (R, Python, MinimPy, MetaPhlAn). Reagents for stool collection and DNA extraction are identified with catalog numbers. Antibodies, cell lines, and mycoplasma testing are not applicable.
“DNA/RNA SheildTM Fecal Collection Tube (Zymo Research) containing buffer (catalog no. R1101, Zymo Research).”
“Analyses were carried out using v.4.0.2 of R and Python v.3.9.7. Pandas v.1.1.3, NumPy v.1.23.5 and SciPy v.1.11.1 were used”
“Species-level profiling of the 815 samples was performed with both MetaPhlAn 3.0 (ref. ) and MetaPhlAn 4.0 (ref. ).”
“A personalized ZOE food quality score was computed using the ZOE 2022 algorithm for each food item consumed by the PDP participants.”
“Analyses were carried out using v.4.0.2 of R and Python v.3.9.7. Pandas v.1.1.3, NumPy v.1.23.5 and SciPy v.1.11.1 were used to manage and preprocess data.”
“The DNA was first isolated using the ZymoBIOMICS 96 MagBead DNA Kit (Zymo Research).”
The primary analysis used repeated measures models with interaction terms, and p-values are reported exactly (e.g., P = 0.016). Effect sizes are reported with 95% CIs. Software is identified. Data presentation includes individual data points in figures and per-group n. Assumptions are handled by log-transformation and normality testing. Mathematical plausibility checks were not possible for all values, but no obvious errors were found.
“mean difference in changes between the groups was −0.13 mmol l −1 (log-transformed, 95% CI = −0.07 to −0.01, P = 0.016”
“body weight: −2.46 kg (95% CI = −3.67 to −1.25)”
“Analyses were carried out using v.4.0.2 of R and Python v.3.9.7.”
“P = 0.016 for the interaction between diet group, time-adjusted for age and sex”
“mean difference in changes between the groups was −0.13 mmol l −1 (log-transformed, 95% CI = −0.07 to −0.01, P = 0.016”
“Reductions in body weight, waist circumference and glycated hemoglobin (HbA1c), and increases in diet quality (HEI score), were significantly greater after the PDP than with the control diet; differences between treatments were as follows: body weight: −2.46 kg (95% CI = −3.67 to −1.25); waist circumference: −2.35 cm (95% CI = −4.07 to −0.63); HbA1c: −0.05% (95% CI = −0.01 to −0.001); and diet quality (HEI score): 7.08 (95% CI = 5.02 to 9.15).”
The data availability statement provides a concrete route for data access (proposal to scientific advisory board, contact email). Microbiome data will be uploaded to EBI. However, code availability is only 'available upon request' with no public repository or permanent identifier, which is inadequate for code sharing.
“The study data can be released to bona fide researchers submitting a research proposal approved by a subpanel of our scientific advisory board.”
“The microbiome data will be uploaded onto the EBI website ( www.ebi.ac.uk/ (https://www.ebi.ac.uk/) ).”
“The scripts for the statistical analysis are freely available upon request to ZOE Ltd.”
“The study data can be released to bona fide researchers submitting a research proposal approved by a subpanel of our scientific advisory board.”
“The microbiome data will be uploaded onto the EBI website ( www.ebi.ac.uk/ (https://www.ebi.ac.uk/) ).”
“The scripts for the statistical analysis are freely available upon request to ZOE Ltd. Application is via data.papers@joinzoe.com. Code will be made available within 2 months of the request.”
The trial is registered on ClinicalTrials.gov (NCT05273268). Methods are comprehensive. Limitations are discussed, including lack of matching for contact/intensity and inability to capture physical activity. Conclusions are generally proportional, though some claims about microbiome benefits are somewhat strong. Funding and competing interests are disclosed.
“ClinicalTrials.gov registration: NCT05273268 (https://clinicaltrials.gov/ct2/show/NCT05273268)”
“Limitations include that we could not accurately capture changes in physical activity status.”
“ClinicalTrials.gov registration: NCT05273268 (https://clinicaltrials.gov/ct2/show/NCT05273268) .”
“Limitations include that we could not accurately capture changes in physical activity status. Furthermore, although reflective of how the advice is delivered in real life, the USDA recommended diet was delivered via leaflet and video, and was intentionally not matched for contact or intensity with the PDP group.”
“This research was funded by ZOE Ltd; the study funder contributed, as part of the scientific advisory board, to study design, data collection and analysis, and the writing of the manuscript.”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 47 references by DOI: 46 verified — 1 no DOI (shown, not verified).
- NO DOIBasal Metabolic Rate: Review and Prediction, Together with an Annotated Bibliography of Source MaterialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
4 data/code links checked; 4 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT05273268LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://www.ebi.ac.uk/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/SegataLab/preprocessingResolves to GitHub (code repository).
- codehttps://pypi.org/project/MinimPy/LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORtypoAbstract“The results demonstrates”→ The results demonstrateSubject-verb agreement error.
- MINORconsistencyResults, Primary outcomes“log-transformed 95% confidence interval = −0.07 to −0.01”→ log-transformed 95% CI = −0.07 to −0.01Inconsistent abbreviation of confidence interval.
- MINORclarityMethods, Statistical analysis“The model evaluates the interaction between time (within-subject factor) and diet treatment (between-subject factor) with diet treatment, time, age and sex included as fixed effects along with a random effect for participants.”→ The model evaluates the interaction between time (within-subject factor) and diet treatment (between-subject factor), with diet treatment, time, age, and sex as fixed effects and a random effect for participants.Run-on sentence; consider splitting for clarity.
- MINORconsistencyResults, Participant characteristics“In total, 86% of participants were female”→ Consider reporting the exact number of female participants for clarity.Percentage is fine, but exact count is available in Table 1.
- MINORclarityMethods, Statistical analysis“The ITT cohort was restricted to 118 and 112 individuals for the control and PDP groups, respectively.”→ Clarify that this restriction is for the microbiome analysis only.This could be misinterpreted as the ITT cohort for all analyses.
The published work is robust and generally trustworthy, but readers should weigh the limited code availability and missing microbiome accession numbers as minor reproducibility concerns. No erratum is warranted based on the current evidence, though providing the code and accession numbers would strengthen the record.
- 1.HIGHdata codeDeposit the statistical analysis code in a public repository (e.g., Zenodo or GitHub) with a DOI, and update the Code availability statement to include the repository link and permanent identifier.The current 'available upon request' is vague and does not meet reproducibility standards.
- 2.HIGHdata codeProvide the EBI accession number for the microbiome data once deposited, and include it in the Data availability statement.Without an accession number, the data deposit is not verifiable.
- 3.MEDIUMstatisticsReport exact p-values for all secondary outcomes instead of thresholds (e.g., P < 0.05) in the Results section.Exact p-values improve precision and transparency.
- 4.MEDIUMreportingClarify the blinding status of participants and outcome assessors in the Methods; currently only the between-group analysis is stated as blinded.Full blinding details are expected for a rigorous RCT.
- 5.MEDIUMreportingAdd an explicit statement that the CONSORT checklist is followed, and consider providing it as a supplementary file.Explicit reference to the reporting guideline enhances transparency.
- 6.MEDIUMreportingTemper the claim about microbiome benefits in the Discussion, as the clinical relevance of the observed changes is not established.Overstating exploratory findings can mislead readers.
- 7.MEDIUMstatisticsAdd a note in the Methods about how missing data were handled in the repeated measures model (e.g., mixed-effects model handles missing data).Clarifying missing data handling is important for reproducibility.
- 8.MEDIUMdata codeSpecify the exact version of the ZOE algorithm used and whether it is publicly available for replication.The algorithm is a key resource; version and availability are needed for reproducibility.
- 9.LOWcopyeditFix the subject-verb agreement error in the Abstract: 'The results demonstrates' should be 'The results demonstrate'.Correct grammar is expected in a published manuscript.
- 10.LOWcopyeditStandardize the abbreviation of confidence interval in the Results section (e.g., use '95% CI' consistently).Consistent terminology improves readability.
- 11.LOWcopyeditClarify in the Methods that the ITT cohort restriction to 118 and 112 individuals applies only to the microbiome analysis.Prevents misinterpretation of the ITT cohort for all analyses.
- 12.LOWreportingConsider reporting the exact number of female participants in the Results for clarity, in addition to the percentage.Exact counts are more informative than percentages alone.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.