Androgen receptor pathway inhibitors and taxanes in metastatic prostate cancer: an outcome-adaptive randomized platform trial.
De Laere B, Crippa A, Discacciati A, Larsson B, Persson M, Johansson S, D'hondt S, Bergström R, Chellappa V, Mayrhofer M, Banijamali M, Kotsalaynen A, Schelstraete C, Vanwelkenhuyzen JP, Hjälm-Eriksson M, Pettersson L, Ullén A, Lumen N, Enblad G, Thellenberg Karlsson C, Jänes E, Sandzén J, Schatteman P, Nyre Vigmostad M, Olsson M, Ghysel C, Sautois B, De Roock W, Van Bruwaene S, Anden M, Verbiene I, De Maeseneer D, Everaert E, Darras J, Aksnessether BY, Luyten D, Strijbos M, Mortezavi A, Oldenburg J, Ost P, Eklund M, Grönberg H, Lindberg J
- DOI
- 10.1038/s41591-024-03204-2
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/4d627ac3-2234-4f07-8f7e-5ca9b6148f7f is authoritative.
How this rating was calculated
- CitationsCitations & links (capped) ×8−1★
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
Citations & links are capped at −1★ combined, however many are flagged.
- No reported statistical tests were found to recompute.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is time to no longer clinically benefitting (NLCB), a composite of PSA, radiologic, and clinical progression, which is a surrogate for overall survival. The paper does not provide evidence of target engagement at the tested doses (e.g., PK/PD) nor does it cite validated evidence linking NLCB to overall survival in this setting. Although overall survival is reported as a secondary endpoint, the primary efficacy claim is based on NLCB.
“The primary endpoint was the time to no longer clinically benefitting (NLCB).”
- 02Other integrity concern
Trial NCT03903835 was first submitted to ClinicalTrials.gov on 2019-03-29, after the registered study start date of 2019-02-01. Retrospective registration means the protocol and outcomes were not on the public record before the study ran, which is what prospective registration exists to establish.
NCT03903835
reviewer’s wording
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper is a well-conducted, biomarker-driven adaptive platform trial with strong methodological rigor, clear reporting of ethics, data/code availability, and statistical methods. Minor gaps include lack of explicit reporting guideline, missing race/ethnicity data, and unverified model assumptions, but these do not undermine the overall validity.
Both reviewers classified the study as interventional; no divergence. The evaluation covered the full text, with N/A for animal-related and bench-resource criteria. The statistics verification component found no recomputable tests, so statistical correctness is not claimed beyond what was reported. The integrity check flagged retrospective registration (submitted after study start), which is a transparency concern but not a validity threat.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Conclusions only partially backed by the presented evidenceAssessed
8 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1Biomarker signatures identify subgroups with differential treatment effects.The paper reports differential STRs across biomarker subgroups, but some analyses have wide credible intervals and are exploratory, so the claim is partially supported.Evidence: STR ratios for AR (SNV/GSR)-negative/TP53 wild-type and TMPRSS2-ERG positive show larger effects, but TP53-altered shows no difference.
“Comparing the STRs, the effect of ARPIs versus taxanes was 44% (STR ratio 1.44, 90% CrI 1.05, 1.95) higher in AR (SNV/GSR)-negative and TP53 wild-type patients compared to patients with alterations in these genes.”
ResultsFind in source - supportedReviewer 1ARPIs demonstrate longer time to NLCB compared to taxanes and physician's choice in the biomarker-unselected population.The primary endpoint analysis shows STRs of 1.50 and 1.60 with 90% CrIs excluding 1, supporting the claim.Evidence: STR for NLCB: ARPIs vs physician's choice 1.50 (1.20-1.86); ARPIs vs taxanes 1.60 (1.28-2.01).
“The STR for the time to NLCB for ARPIs was 1.50 (90% credible intervals (CrI) 1.20, 1.86) compared to the physician’s choice (median 11.1 versus 7.4 months) and 1.60 (90% CrI 1.28, 2.01) compared to taxanes (median 11.1 versus 6.9 months; Fig. , Table and Supplementary Fig. ).”
ResultsFind in source - supportedReviewers 1, 2ARPIs demonstrate longer overall survival compared to taxanes and physician's choice.Overall survival STRs are 1.77 and 1.78 with 90% CrIs excluding 1, supporting the claim.Evidence: STR for OS: ARPIs vs physician's choice 1.77 (1.29-2.51); ARPIs vs taxanes 1.78 (1.28-2.61).
“The STR for overall survival in the biomarker-unselected ‘all’ patients group was 1.77 (90% CrI 1.29, 2.51) for ARPIs compared to physician’s choice (median 38.7 versus 21.8 months) and 1.78 (90% CrI 1.28, 2.61) compared to the taxane arm (median 38.7 versus median 21.7 months; Fig. and Table ).”
ResultsFind in source - supportedReviewer 1The trial is the first outcome-adaptive platform trial in prostate cancer.The paper states this as a novel design, and the description supports it.Evidence: Abstract states 'ProBio is the first outcome-adaptive platform trial in prostate cancer'.
“ProBio is the first outcome-adaptive platform trial in prostate cancer utilizing a Bayesian framework to evaluate efficacy within predefined biomarker signatures across systemic treatments.”
AbstractFind in source - supportedReviewer 2ARPIs demonstrate ~50% longer time to NLCB compared to taxanes in the biomarker-unselected population.The claim is supported by the reported STR of 1.60 (90% CrI 1.28, 2.01) and median times of 11.1 vs 6.9 months.Evidence: Table 2 and Figure 2a show STR 1.60 (90% CrI 1.28, 2.01) for ARPIs vs taxanes in all patients.
“ARPIs demonstrated ~50% longer time to NLCB compared to taxanes (median, 11.1 versus 6.9 months)”
AbstractFind in source - supportedReviewer 2The largest increase in time to NLCB was observed in AR-negative and TP53 wild-type patients and TMPRSS2-ERG fusion-positive patients.The claim is supported by STRs of 1.76 and 1.80 for these subgroups, with 90% CrI excluding 1.0.Evidence: Table 2 shows STR 1.76 (90% CrI 1.26, 2.51) for AR-negative/TP53 wild-type and 1.80 (90% CrI 1.21, 2.64) for TMPRSS2-ERG fusion-positive patients.
“Biomarker signature findings suggest that the largest increase in time to NLCB was observed in AR (single-nucleotide variant/genomic structural rearrangement)-negative and TP53 wild-type patients and TMPRSS2–ERG fusion-positive patients”
AbstractFind in source - supportedReviewer 2No difference between ARPIs and taxanes was observed in TP53-altered patients.The claim is supported by STR of 1.05 (90% CrI 0.81, 1.38) for ARPIs vs taxanes in TP53-altered patients, with CrI crossing 1.0.Evidence: Table 2 shows STR 1.05 (90% CrI 0.81, 1.38) for ARPIs vs taxanes in TP53-altered patients.
“whereas no difference between ARPIs and taxanes was observed in TP53 -altered patients.”
AbstractFind in source - supportedReviewer 2ProBio is the first outcome-adaptive platform trial in prostate cancer utilizing a Bayesian framework to evaluate efficacy within predefined biomarker signatures.The claim is supported by the description of the trial design and the Bayesian framework, and the paper does not contradict this novelty claim.Evidence: The paper describes the adaptive randomization and Bayesian Weibull model in the Methods.
“ProBio is the first outcome-adaptive platform trial in prostate cancer utilizing a Bayesian framework to evaluate efficacy within predefined biomarker signatures across systemic treatments.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary endpoint is time to no longer clinically benefitting (NLCB), a composite of PSA, radiologic, and clinical progression, which is a surrogate for overall survival. The paper does not provide evidence of target engagement at the tested doses (e.g., PK/PD) nor does it cite validated evidence linking NLCB to overall survival in this setting. Although overall survival is reported as a secondary endpoint, the primary efficacy claim is based on NLCB.
“The primary endpoint was the time to no longer clinically benefitting (NLCB).”
- ADEQUATEEffect sizeThe effect sizes are reported as survival time ratios (STR) with credible intervals, e.g., STR 1.50 (90% CrI 1.20, 1.86) for ARPIs vs physician's choice, and median survival times (11.1 vs 7.4 months). These are clinically meaningful improvements and are statistically supported.
“The STR for the time to NLCB for ARPIs was 1.50 (90% credible intervals (CrI) 1.20, 1.86) compared to the physician’s choice (median 11.1 versus 7.4 months)”
Data authenticity concerns
1 finding · worst mediumAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
1 integrity concern flagged (0 high).
- mediumotherTrial NCT03903835 was first submitted to ClinicalTrials.gov on 2019-03-29, after the registered study start date of 2019-02-01. Retrospective registration means the protocol and outcomes were not on the public record before the study ran, which is what prospective registration exists to establish.
NCT03903835
reviewer’s wording
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior retrospective analyses suggesting varying therapeutic benefits by tumor genotype, acknowledges the small number of predictive biomarkers, and describes the limitations of traditional clinical trials in evaluating biomarker-treatment combinations. The rationale for the ProBio platform trial is logically linked to these gaps, and the paper explains how the adaptive design and prospective ctDNA profiling address prior limitations.
“ProBio uses a number of design elements to maximize the information gained from each patient entering the trial in order to rapidly sieve through multiple hypotheses.”
Randomization is stratified on 16 biomarker subgroup combinations with adaptive probabilities. Physicians are blinded to biomarker results. Inclusion/exclusion criteria are described, including a protocol amendment limiting enrollment to first-line mCRPC. A power analysis was performed via simulation studies, with type I error <10% and power 65-83%. The primary endpoint (time to NLCB) and secondary endpoints are defined. Outlier handling is not explicitly discussed but the analysis uses a Bayesian model with vague priors, which is robust. Controls are the physician's choice arm. Replicate distinction and independent replication are not applicable for this human RCT.
“Randomization was stratified on these subgroups, leading to varying randomization probabilities across different subgroup levels.”
“Participating physicians, both in the control and investigational treatment arms, are blinded to the biomarker results.”
“Randomization was stratified on these subgroups, leading to varying randomization probabilities across different subgroup levels.”
“Participating physicians, both in the control and investigational treatment arms, are blinded to the biomarker results.”
The study includes only male patients (sex is reported). Age is reported as median and IQR in Table 1. Health status is captured by ECOG performance status. Demographics include age, ECOG, PSA levels, and prior treatments. Race/ethnicity is not reported, which is noted as a limitation in the Discussion. Species/strain and housing conditions are not applicable for a human trial.
“Age at study entry (median, range), years | 71 (68, 76) | 71 (65, 75) | 69 (64, 73) | 71 (68, 75) | 69 (66, 73) | 68 (66, 72)”
“the absence of data on race and ethnicity may impact the generalizability of our findings”
The paper states approval by ethics boards in Sweden, Belgium, Norway, and Switzerland with specific IDs. Written informed consent is mentioned. Regulatory compliance is implied through adherence to local guidelines and the Declaration of Helsinki (not explicitly named but implied).
“approved by ethics boards in Sweden (ID: 2018/2206–32; 22 October 2018), Belgium (ID: BC-06057; 20 March 2020), Norway (ID: Søknadsnummer 81005/58005; 24 June 2020) and Switzerland (ID: BASEC 2021–02495; 01 March 2022)”
“At inclusion and upon receiving the patient’s written informed consent, the patient is enrolled in the electronic case report form system”
“approved by ethics boards in Sweden (ID: 2018/2206–32; 22 October 2018), Belgium (ID: BC-06057; 20 March 2020), Norway (ID: Søknadsnummer 81005/58005; 24 June 2020) and Switzerland (ID: BASEC 2021–02495; 01 March 2022).”
“At inclusion and upon receiving the patient’s written informed consent, the patient is enrolled in the electronic case report form system”
The investigational products are identified: abiraterone acetate (1000 mg daily), enzalutamide (160 mg daily), docetaxel (75 mg/m2), cabazitaxel (20-25 mg/m2). The software used for analysis is R version 4.2.2, and the electronic case report form system is SMART-TRIAL. Antibodies, cell lines, mycoplasma testing, and organisms are not applicable for this human drug trial.
“ARPIs (either 1,000 mg of abiraterone acetate or 160 mg of enzalutamide daily), taxanes (either docetaxel at a dosage of 75 mg or cabazitaxel at a dosage of 20–25 mg per square meter of body-surface area intravenously every 3 weeks)”
“All analyses were performed in R (version 4.2.2)”
“ARPIs (either 1,000 mg of abiraterone acetate or 160 mg of enzalutamide daily), taxanes (either docetaxel at a dosage of 75 mg or cabazitaxel at a dosage of 20–25 mg per square meter of body-surface area intravenously every 3 weeks)”
“All analyses were performed in R (version 4.2.2)”
The primary analysis uses a Bayesian Weibull accelerated failure time model, with STRs and 90% CrI reported. Effect sizes are reported with credible intervals. Software (R 4.2.2) is identified. Data presentation includes Kaplan-Meier curves, posterior survival curves, and per-group Ns in tables. Exact p-values are not applicable as the analysis is Bayesian (posterior probabilities of superiority are reported). Assumptions of the Weibull model are not explicitly verified, but this is standard for such models. Mathematical plausibility checks are not applicable due to the Bayesian framework and continuous outcomes.
“We utilized a Bayesian Weibull accelerated failure time model estimated on all data (that is, first and second randomization) for the time to NLCB and on first randomization data for overall survival.”
“The STR for the time to NLCB for ARPIs was 1.50 (90% credible intervals (CrI) 1.20, 1.86)”
“We utilized a Bayesian Weibull accelerated failure time model estimated on all data (that is, first and second randomization) for the time to NLCB”
“The STR for the time to NLCB for ARPIs was 1.50 (90% credible intervals (CrI) 1.20, 1.86)”
“All analyses were performed in R (version 4.2.2)”
The data availability statement explains that individual patient data cannot be deposited publicly due to Swedish law, but provides a concrete access route via the corresponding author with conditions and a timeframe. Code for statistical analysis is available on GitHub.
“Requests will be processed within 1–2 months upon submission of a complete and compliant auxiliary research-specific study protocol.”
“Code for the statistical analysis is available via GitHub at https://github.com/alecri/arpi_all/”
“Code for the statistical analysis is available via GitHub at https://github.com/alecri/arpi_all/”
The trial is registered (NCT03903835). The primary and secondary endpoints are reported, and additional planned endpoints (QoL, cost-effectiveness) are noted as not reported here. Limitations are extensively discussed. Conclusions are proportional to the evidence. Funding sources and competing interests are declared.
“ClinicalTrials.gov registration: NCT03903835”
“Some limitations need to be considered when interpreting our results.”
“The ProBio consortium thanks the following funding organizations for the following awarded research grants: ALF Medicine (to H.G., 20190087), Swedish Cancer Society (to H.G., 211610P), Swedish Research Council (to H.G., 2021-00331), Krebsliga beider Basel (to A.M., KLbB-5580-02-2022) and Kom Op Tegen Kanker/Stand up to Cancer – Flemish Cancer Society (to P.O., STI.VLK.2020.0006.01; to B.D.L., STI.VLK.2022.0005.01).”
“ClinicalTrials.gov registration: NCT03903835 (https://clinicaltrials.gov/study/NCT03903835)”
“Some limitations need to be considered when interpreting our results. Firstly, the chosen endpoint of NLCB as the termination point for treatment may lead to bias in open-label trials.”
Registered (1 ID: ClinicalTrials.gov). Reporting guidelines cited: CONSORT, SPIRIT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 55 references by DOI: 46 verified — 8 DOI unresolved, 1 no DOI (shown, not verified).
- UNRESOLVED10.1007/s11523-020-00720-wReal-world outcomes in first-line treatment of metastatic castration-resistant prostate cancer: the prostate cancer registryCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1158/1078-0432.ccr-18-2026TP53 outperforms other androgen receptor biomarkers to predict abiraterone or enzalutamide outcome in metastatic castration-resistant prostate cancerCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1158/1078-0432.ccr-22-2138SPOP mutations as a predictive biomarker for androgen receptor axis-targeted therapy in de novo metastatic castration-sensitive prostate cancerCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1038/s41571-023-00822-2Progression-free survival, disease-free survival and other composite end points in oncology: improved reporting is neededCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1186/s13063-020-04510-5The ProBio trial: molecular biomarkers for advancing personalized treatment decision in patients with metastatic castration-resistant prostate cancerCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1158/1078-0432.ccr-20-4091Evolution of castration-resistant prostate cancer in ctDNA during sequential androgen receptor pathway inhibitionCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.1158/1078-0432.ccr-22-1726Olaparib efficacy in patients with metastatic castration-resistant prostate cancer and BRCA1, BRCA2, or ATM alterations identified by testing circulating tumor DNACited DOI does not resolve to any Crossref record.
- UNRESOLVED10.6004/jnccn.2022.0063NCCN Guidelines Insights: Prostate Cancer, version 1.2023Cited DOI does not resolve to any Crossref record.
- NO DOIEAU GuidelinesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
4 data/code links checked; 4 live.
- datahttps://clinicaltrials.gov/study/NCT03903835LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT03903835LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- codeGitHubLIVEHTTP 200https://github.com/alecri/arpi_all/Resolves to GitHub (code repository).
- codeGitHubLIVEHTTP 200https://github.com/ClinSeq/jumbleResolves to GitHub (code repository).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoTable 3“permament”→ permanentTypo in table header.
- MINORconsistencyAbstract“~50% longer time to NLCB”→ Consider using exact STR values for consistency with results.Approximate percentage may be less precise than the reported STR.
- MINORclarityDiscussion“the NLCB endpoint might have been more impacted by, for example, PSA-only progressive disease in ARPI-treated patients.”→ Clarify the direction of potential bias.Sentence could be clearer.
- MINORtypoTable 3, header“permament discontinuation”→ permanent discontinuationTypo in table header.
- MINORconsistencyTable 2, footnote“Posterior probability of superiority (PPS) calculated as STR exceeding one”→ Consider clarifying that PPS is the probability that STR > 1.Minor clarity issue.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (no explicit CONSORT statement, missing race/ethnicity data, unverified model assumptions) and the retrospective registration as transparency issues, but none warrant an erratum or independent re-analysis. The copyedit issues are trivial typos and clarity improvements.
- 1.HIGHreportingAdd an explicit statement referencing the CONSORT reporting guideline in the Methods or a reporting summary section.The paper does not explicitly reference a reporting guideline, which is a standard expectation for clinical trials and improves transparency.
- 2.HIGHethicsAdd a statement confirming compliance with the Declaration of Helsinki in the Methods/ethics section.Regulatory compliance is implied but not explicitly stated; making it explicit strengthens the ethics documentation.
- 3.HIGHstatisticsExplicitly verify and report assumptions of the Weibull accelerated failure time model (e.g., proportional hazards) or note that they were checked.Reviewer 2 flagged that model assumptions are not explicitly verified; documenting this adds rigor to the statistical analysis.
- 4.MEDIUMreportingClarify in the Discussion that additional planned secondary endpoints (QoL, cost-effectiveness) will be reported in future publications.Reviewer 1 noted ambiguity about unreported endpoints; a clear statement avoids the impression of selective reporting.
- 5.MEDIUMreportingInclude race/ethnicity data in Table 1 or discuss the limitation more prominently in the Discussion.Missing race/ethnicity data affects generalizability; both reviewers flagged this as a limitation.
- 6.MEDIUMstatisticsProvide a more detailed description of outlier handling or note that the Bayesian model is robust to outliers.Reviewer 2 rated outlier handling as inadequate; clarifying this addresses a methodological concern.
- 7.MEDIUMreportingClarify the direction of potential bias in the Discussion sentence about NLCB endpoint and PSA-only progressive disease.The copyedit flagged this sentence as unclear; clarifying improves reader understanding.
- 8.MEDIUMreportingClarify in Table 2 footnote that PPS is the probability that STR > 1.The copyedit noted this clarity issue; a precise definition aids interpretation.
- 9.LOWcopyeditFix the typo 'permament' to 'permanent' in Table 3 header.Copyedit flagged this typo; correcting it improves professionalism.
- 10.LOWcopyeditConsider using exact STR values instead of approximate percentage in the Abstract for consistency with results.Copyedit noted the approximate percentage may be less precise; using exact values improves consistency.
- 11.LOWdata codeConsider depositing de-identified aggregate data in a public repository if permitted by ethical approvals.Reviewer 2 suggested this to enhance data availability; while current managed access is adequate, public deposit would increase transparency.
- 12.LOWotherAdd a note on the statistical software version for the Bayesian analysis (e.g., R package used).Reviewer 1 suggested this to improve reproducibility; specifying the package version aids replication.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.