Patient, family caregiver, and economic outcomes of an integrated screening and novel stepped collaborative care intervention in the oncology setting in the USA (CARES): a randomised, parallel, phase 3 trial.
Steel JL, George CJ, Terhorst L, Yabes JG, Reyes V, Zandberg DP, Nilsen M, Kiefer G, Johnson J, Marsh C, Bierenbaum J, Tageja N, Krauze M, VanderWeele R, Goel G, Ramineni G, Antoni M, Vodovotz Y, Walker J, Tohme S, Billiar T, Geller DA
- DOI
- 10.1016/S0140-6736(24)00015-1
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/0d2e799e-958a-4c20-aa0e-d1822f5c0cf2 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsOverstated claim−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is health-related quality of life (HRQOL) measured by the FACT-General, which is a patient-reported outcome. While HRQOL is a clinical outcome, it is a surrogate for more definitive outcomes like survival or disease progression. The paper does not provide evidence linking improvements in HRQOL to hard clinical outcomes, nor does it demonstrate target engagement at the tested dose (e.g., PK/PD) for the intervention. The efficacy claim is based on a surrogate measure without a validated link to clinical benefit.
“The primary outcome was health-related quality of life in patients at 6 months.”
- 02Treatment effect not shown to be clinically meaningful
The reported effect size for the primary outcome (HRQOL) is small (Cohen's d = 0.09 for between-group difference at 6 months). The paper claims this is clinically meaningful, but the effect size is below the commonly accepted threshold for a minimal clinically important difference (MCID) for HRQOL measures. No anchor to clinical meaningfulness is provided beyond the authors' assertion.
“Patients in the stepped collaborative care group had a greater 0–6-month improvement in health-related quality of life than patients in the standard-of-care group (p=0·013, effect size 0·09).”
- 03Conclusion reaches beyond the evidence
The intervention is cost-saving and could shift practice.
“The findings of this study will advance the implementation of guideline concordant care (screening and treatment) and has the potential to shift the practice of screening and treatment paradigm nationwide, improving outcomes for patients diagnosed with cancer.”
InterpretationFind in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported randomized phase 3 trial with rigorous design, clear ethics approval, and appropriate statistical methods. The main weakness is the vague data sharing statement, which lacks a concrete access mechanism or repository deposit. Minor copyedit issues and an overstated claim about shifting practice nationwide are also noted.
Both reviewers independently scored all eight dimensions and agreed on all statuses, so no divergence needed reconciliation. The study is interventional (randomized controlled trial). Non-applicable criteria (e.g., animal housing, cell line authentication) were excluded. The statistics verification covered only 2 tests; the rest were not machine-verifiable, so statistical correctness is not fully confirmed.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks.
- CONSISTENTreported p = .440 · recomputed p = .787Reviewers 1, 2Check p-value for 30-day re-admission comparison (13% vs 14%) using Fisher's exact test.
“30-day re-admission rates did not differ significantly between stepped collaborative care and standard of care (≥1 re-admission in 31 [13%] of 237 patients vs 31 [14%] of 222 patients; p=0·44).”
Taken as given: The numbers 31 and 237 are the event count and total for the stepped collaborative care group.; The numbers 31 and 222 are the event count and total for the standard of care group.; The test used is Fisher's exact test (two-tailed).Method: Fisher's exact test on 2x2 table (31, 206, 31, 191).How we recomputed it: pFisher2x2(31, 206, 31, 191) - CONSISTENTreported p = .050 · recomputed p = .051Reviewers 1, 2Check p-value for 90-day re-admission comparison (5% vs 10%) using Fisher's exact test.
“fewer 90-day re-admissions (5% vs 10%; p=0·050)”
Taken as given: The 5% and 10% are the re-admission rates for stepped collaborative care and standard of care, respectively.; The denominators are the group totals (237 and 222).; The event counts are approximated as 12 and 22 (5% of 237 ≈ 11.85, 10% of 222 ≈ 22.2).; The test used is Fisher's exact test (two-tailed).Method: Fisher's exact test on 2x2 table (12, 225, 22, 200).How we recomputed it: pFisher2x2(12, 225, 22, 200)
- lowinternal contradictionThe number of family caregivers enrolled is reported as 190 in the abstract, but the results section states 318 consented and 190 were assigned (91+99). The abstract says '190 family caregivers were enrolled' which matches the assigned number, but the results say '318 family caregivers consented' which is a different number. This is not a contradiction but a difference between consented and randomized.
Abstract: '190 family caregivers were enrolled.' Results: '318 family caregivers consented to participate in the study; 91 were assigned to standard of care and 99 to stepped collaborative care.'
Abstractreviewer’s wording - lowinternal contradictionThe number of patients enrolled is reported as 459 in the abstract, but the results state 735 were enrolled and 459 were randomly assigned. This is a difference between enrolled and randomized, which is expected.
Abstract: '459 patients and 190 family caregivers were enrolled.' Results: '735 were enrolled. 459 patients were randomly assigned.'
Abstractreviewer’s wording
Overstated conclusions
4 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated), 1 only partially supported (evidence backs part of the claim; gaps or caveats remain).
- overstatedReviewer 1The intervention is cost-saving and could shift practice.While cost savings are reported, the claim that it 'has the potential to shift the practice of screening and treatment paradigm nationwide' is an extrapolation beyond the single-trial evidence.Evidence: Cost savings of $4 million estimated; but this is a single trial and generalizability is not proven.
“The findings of this study will advance the implementation of guideline concordant care (screening and treatment) and has the potential to shift the practice of screening and treatment paradigm nationwide, improving outcomes for patients diagnosed with cancer.”
InterpretationFind in source - partialReviewers 1, 2The intervention reduces health-care use and costs.The paper reports lower activity-based costs and fewer re-admissions, but these are tertiary outcomes and not all comparisons were significant (e.g., 30-day re-admissions).Evidence: Activity-based costs lower by $17,085 per patient per year; fewer 90-day re-admissions (5% vs 10%, p=0.050); 30-day re-admissions not significantly different.
“Patients in the stepped collaborative care group had US$17 085·04/patient per year lower activity-based costs than those in standard-of-care group”
Results ¶8Find in source - supportedReviewers 1, 2The integrated screening and stepped collaborative care intervention improves health-related quality of life in cancer patients compared to standard of care.The primary outcome showed a statistically significant improvement in HRQOL at 6 months (p=0.013) with a small effect size, supporting the claim.Evidence: Primary outcome analysis: p=0.013, effect size 0.09.
“Patients in the stepped collaborative care group had a greater 0–6-month improvement in health-related quality of life than patients in the standard-of-care group (p=0·013, effect size 0·09).”
AbstractFind in source - supportedReviewers 1, 2The intervention improves family caregiver outcomes.Multivariate analysis showed significant differences in caregiver outcomes (p=0.027), primarily driven by lifetime CVD risk, supporting the claim.Evidence: Multivariate analysis of caregiver outcomes: χ-bar-squared 9.00, p=0.027, effect size 0.41.
Multivariate analyses using general linear mixed effects of caregiver outcomes ... showed that groups differed significantly (χ-bar-squared 9·00, p=0·027; effect size 0·41).
Results ¶5reviewer’s wording - supportedReviewer 2The intervention is cost-saving.The paper reports total savings of $4.0 million, but this is based on activity-based costing and may not generalize; however, the claim is supported by the presented data.Evidence: Total savings estimated at $4,049,145 million if payers cover treatment cost.
“The total savings for the stepped collaborative care group was $4 049 145 million if the payers and patients covered the treatment cost”
Results ¶8Find in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is health-related quality of life (HRQOL) measured by the FACT-General, which is a patient-reported outcome. While HRQOL is a clinical outcome, it is a surrogate for more definitive outcomes like survival or disease progression. The paper does not provide evidence linking improvements in HRQOL to hard clinical outcomes, nor does it demonstrate target engagement at the tested dose (e.g., PK/PD) for the intervention. The efficacy claim is based on a surrogate measure without a validated link to clinical benefit.
“The primary outcome was health-related quality of life in patients at 6 months.”
- INADEQUATEEffect sizeThe reported effect size for the primary outcome (HRQOL) is small (Cohen's d = 0.09 for between-group difference at 6 months). The paper claims this is clinically meaningful, but the effect size is below the commonly accepted threshold for a minimal clinically important difference (MCID) for HRQOL measures. No anchor to clinical meaningfulness is provided beyond the authors' assertion.
“Patients in the stepped collaborative care group had a greater 0–6-month improvement in health-related quality of life than patients in the standard-of-care group (p=0·013, effect size 0·09).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior research on cancer-related symptoms, the ineffectiveness of standard screening and referral, and the evidence base for collaborative care. It acknowledges limitations of prior work (e.g., only one UK study showing cost-effectiveness, guidelines not integrating screening and treatment) and explains how the current study addresses these gaps. The hypothesis follows logically from the cited evidence.
“Decades of research have shown the effectiveness and cost-effectiveness of the collaborative care approach, which is widely implemented across primary care practices in the USA.”
“This study aimed to test the efficacy of an integrated screening and novel stepped collaborative care intervention versus standard of care (ie, screening and treatment referral) for patients with cancer and comorbid symptoms, such as depression, pain, and fatigue.”
“Decades of research have shown the effectiveness and cost-effectiveness of the collaborative care approach, which is widely implemented across primary care practices in the USA.”
“This study aimed to test the efficacy of an integrated screening and novel stepped collaborative care intervention versus standard of care (ie, screening and treatment referral) for patients with cancer and comorbid symptoms, such as depression, pain, and fatigue.”
Randomization used a central permuted block design stratified by sex and prognostic status. Masking of biostatisticians, oncologists, and outcome assessors was described. A priori power analysis was provided with effect sizes and sample size calculations. Inclusion/exclusion criteria were pre-specified. Outlier handling is addressed through intention-to-treat analysis and sensitivity analyses for missing data. Controls are inherent in the standard-of-care comparator. Independent replication is not applicable for a single pivotal trial.
“The biostatisticians, oncologists, and outcome assessors were masked to patient allocation until 12-month outcome data were collected.”
“a sample size of 364 (182 participants per group) would have 80% power to detect a between-group effect (Cohen’s d ) as small as 0·15 and 95% power to detect a within-group effect (Cohen’s d ) as small as 0·10.”
“The biostatisticians, oncologists, and outcome assessors were masked to patient allocation until 12-month outcome data were collected.”
“a sample size of 364 (182 participants per group) would have 80% power to detect a between-group effect (Cohen’s d ) as small as 0·15 and 95% power to detect a within-group effect (Cohen’s d ) as small as 0·10.”
Sex is reported for patients and caregivers. Age and health status (cancer type, stage, comorbidities) are reported. Demographics include race, education, income, and marital status. Since both sexes are enrolled, sex justification is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“Patients’ mean age was 65·70 years (SD 11·42), and of 459 patients, 201 (44%) were male and 258 (56%) were female.”
“White | 218 (92%) | 207 (93%) | 95 (96%) | 85 (93%)”
“Stage IV | 105 (44%) | 86 (39%)”
“Patients’ mean age was 65·70 years (SD 11·42), and of 459 patients, 201 (44%) were male and 258 (56%) were female.”
The study received approval from the University of Pittsburgh Institutional Review Board with a protocol number (PRO15030290). Written informed consent was obtained from patients. Regulatory compliance is implied through IRB approval and adherence to ethical standards, though not explicitly naming a framework like the Declaration of Helsinki.
“Ethics approval was received from the University of Pittsburgh Institutional Review Board (PRO15030290).”
“Patients provided written informed consent.”
“Ethics approval was received from the University of Pittsburgh Institutional Review Board (PRO15030290).”
“Patients provided written informed consent.”
The intervention is described with details on delivery, frequency, and content. Pharmacotherapy is mentioned but not specified as a named drug. Statistical software (SAS version 9.4) is identified. No antibodies, cell lines, or organisms are used, so those criteria are not applicable. The intervention itself is the key resource and is adequately described.
“The stepped collaborative care intervention was once weekly CBT for approximately 50–60 min from a care coordinator based on the patients’ preference for treatment (eg, psychotherapy, pharmacotherapy) and delivery (eg, telephone or videoconferencing).”
“Linear mixed effect models were performed using SAS (PROC MIXED; version 9.4)”
“The stepped collaborative care intervention was once weekly CBT for approximately 50–60 min from a care coordinator based on the patients’ preference for treatment (eg, psychotherapy, pharmacotherapy) and delivery (eg, telephone or videoconferencing).”
“Linear mixed effect models were performed using SAS (PROC MIXED; version 9.4)”
The primary analysis used linear mixed effects models with details on covariance structure and degrees of freedom. Exact p-values are reported (e.g., p=0·013). Effect sizes with confidence intervals are provided. Software is identified. Data presentation includes per-group n and error bars. Mathematical plausibility checks were not possible for all results due to complex models, but no obvious errors were found.
“Linear mixed effect models were performed using SAS (PROC MIXED; version 9.4)”
“p=0·013, effect size 0·09”
“Error bars are 95% CI.”
“Linear mixed effect models were performed using SAS (PROC MIXED; version 9.4) to incorporate all available data and test hypotheses”
“p=0·013, effect size 0·09”
“Error bars are 95% CI.”
The data sharing statement says de-identified data will be available on request from the principal investigator, but does not specify a platform, conditions, or timeframe. This is reported_but_inadequate. No code sharing is mentioned, and no repository deposit or accession numbers are provided.
“De-identified data will be available on request from the principal investigator ( steejl@upmc.edu ) to those with appropriate qualifications, approved aims and hypotheses, and a signed data access agreement.”
“De-identified data will be available on request from the principal investigator ( steejl@upmc.edu ) to those with appropriate qualifications, approved aims and hypotheses, and a signed data access agreement.”
The trial is registered with ClinicalTrials.gov (NCT02939755). Methods are detailed enough for replication. Limitations are discussed, including COVID-19 pandemic, missing data, and generalizability. Conclusions are proportional to the evidence. Funding and conflicts of interest are declared. No reporting guideline checklist is mentioned, but this is not mandatory for a trial report.
“This trial was registered with ClinicalTrials.gov (https://clinicaltrials.gov/) ( NCT02939755 (https://clinicaltrials.gov/ct2/show/NCT02939755) ).”
“Although this study has many strengths, there were also limitations. The trial was partly conducted during the COVID-19 pandemic.”
“JLS receives royalties from Springer for the books Living Donor Advocacy and Psychological Aspects of Cancer .”
“This trial was registered with ClinicalTrials.gov (https://clinicaltrials.gov/) ( NCT02939755 (https://clinicaltrials.gov/ct2/show/NCT02939755) ).”
“The trial was partly conducted during the COVID-19 pandemic. Sensitivity analyses were performed to test outcomes for patients who were treated before versus during the COVID-19 pandemic (March 2020); we did not observe any differences in outcomes.”
“JLS receives royalties from Springer for the books Living Donor Advocacy and Psychological Aspects of Cancer .”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 30 references by DOI: 23 verified — 7 no DOI (shown, not verified).
- NO DOINIH State-of-the-Science Statement on symptom management in cancer: pain, depression, and fatigueNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHospital Readmissions Reduction Program (HRRP)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe CES-D scale: a self-report depression scale for research in the general populationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPain assessment: global use of the Brief Pain InventoryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIModeling the mean: analyzing response profilesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOn some characteristics of Gaussian covariance functionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn improved approximation to the precision of fixed effects from restricted maximum likelihoodNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly clarity, grammar, consistency.
- MINORconsistencyAbstract, Findings“p=0·013, effect size 0·09”→ Ensure effect size notation is consistent (e.g., Cohen's d) throughout.Effect size is reported without specifying the type in the abstract.
- MINORclarityResults, paragraph 3“p difference =0·013, effect size 0·09”→ Clarify what 'p difference' refers to (e.g., p-value for between-group difference).The term 'p difference' is not standard; consider rephrasing.
- MINORgrammarDiscussion, paragraph 2“Although all caregiver outcomes differed significantly between the stepped collaborative care and standard-of-care groups with a small-to-moderate effect size, reduced lifetime risk of CVD seemed to drive the differences between groups.”→ Consider splitting this long sentence for clarity.Long sentence with multiple clauses.
- MINORclarityResults, paragraph 3“The distribution of symptoms by treatment group is shown in the (p 5).”→ Complete the sentence with the figure/table reference.Incomplete sentence referencing supplementary material.
- MINORgrammarDiscussion, paragraph 2“Although all caregiver outcomes differed significantly between the stepped collaborative care and standard-of-care groups with a small-to-moderate effect size, reduced lifetime risk of CVD seemed to drive the differences between groups.”→ Consider rephrasing for clarity.Long sentence with multiple clauses.
The published work is robust and well-reported, with only minor reporting gaps. An informed reader should weigh the vague data sharing statement and the slightly overstated claim about shifting practice nationwide. No erratum is warranted, but the authors could improve transparency by depositing data in a repository and tempering the claim.
- 1.HIGHdata codeIn the Data sharing section, specify a concrete data access mechanism such as depositing de-identified data in a repository (e.g., Vivli, Dryad, Zenodo) with conditions and a timeframe, rather than only 'on request'.The current vague 'on request' statement is inadequate for reproducibility and is the only major rigor gap.
- 2.HIGHdata codeAdd a statement about analysis code availability in the Data sharing section, either providing a link to a repository or stating that no custom code was used.Code sharing is not mentioned, which limits reproducibility.
- 3.HIGHrigorIn the Discussion, temper the claim that the intervention 'has the potential to shift the practice of screening and treatment paradigm nationwide' to reflect that this is a single-trial finding requiring replication.The claim audit flagged this as overstated; over-claiming is a common reviewer objection.
- 4.MEDIUMreportingAdd an explicit statement of compliance with the Declaration of Helsinki or other regulatory framework in the Methods/ethics section.Reviewer 1 noted that regulatory compliance is implied but not explicitly named; adding it strengthens the ethics reporting.
- 5.MEDIUMreportingMention adherence to the CONSORT reporting guideline in the Methods or acknowledgments.The paper does not reference a reporting guideline, which is a minor transparency gap for a trial report.
- 6.MEDIUMreportingClarify the specific pharmacotherapy agents used in the intervention, if any, in the Methods/Procedures section.The intervention description mentions pharmacotherapy but does not name specific drugs, which limits replicability.
- 7.MEDIUMcopyeditIn the Abstract and Results, specify the effect size type (e.g., Cohen's d) when reporting 'effect size 0·09'.The copyedit pass flagged inconsistent effect size notation; specifying the type improves clarity.
- 8.MEDIUMcopyeditIn Results, paragraph 3, clarify the term 'p difference' to 'p-value for between-group difference'.The copyedit pass noted that 'p difference' is non-standard and could confuse readers.
- 9.MEDIUMcopyeditIn Results, paragraph 3, complete the sentence 'The distribution of symptoms by treatment group is shown in the (p 5).' with the correct figure/table reference.The copyedit pass flagged an incomplete sentence referencing supplementary material.
- 10.LOWcopyeditIn Discussion, paragraph 2, split the long sentence about caregiver outcomes and CVD risk into shorter sentences for clarity.The copyedit pass flagged the sentence as overly long and complex.
- 11.LOWreportingClarify the role of the funder in the Role of the funding source section.Reviewer 2 suggested clarifying funder role to ensure transparency.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.