Symptom monitoring with electronic patient-reported outcomes during cancer treatment: final results of the PRO-TECT cluster-randomized trial.
Basch E, Schrag D, Jansen J, Henson S, Ginos B, Stover AM, Carr P, Spears PA, Jonsson M, Deal AM, Bennett AV, Thanarajasingam G, Rogak L, Reeve BB, Snyder C, Bruner D, Cella D, Kottschade LA, Perlmutter J, Geoghegan C, Given B, Mazza GL, Miller R, Strasser JF, Zylla DM, Weiss A, Blinder VS, Wolf AP, Dueck AC
- DOI
- 10.1038/s41591-025-03507-y
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/982f5e00-abc1-4544-90ff-1d13d4c6d492 is authoritative.
How this rating was calculated
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on secondary outcomes including time to deterioration of physical function, symptom control, and HRQL, which are patient-reported outcome measures (surrogates for clinical benefit). The paper does not provide evidence of target engagement at the tested dose (since it's a behavioral intervention, not a drug) nor does it cite validated evidence linking these PRO measures to hard clinical outcomes. The primary outcome of overall survival was not significant, so the efficacy claim rests on these surrogate endpoints.
“Benefits also significantly favored PRO for delayed deterioration of physical function (median 12.6 versus 8.5 months, HR 0.73; P = 0.002), symptoms (12.7 versus 9.9, HR 0.69; P < 0.001) and HRQL (15.6 versus 12.2, HR 0.72; P = 0.001)”
- 02Treatment effect not shown to be clinically meaningful
The reported effect sizes for the surrogate outcomes (e.g., HR 0.73 for physical function deterioration) are presented as statistically significant but are not anchored to a minimal clinically important difference or to a clear biological/clinical meaningfulness. The paper does not provide a threshold for clinical significance for these HRs, and the absolute differences in median times (e.g., 4.1 months for physical function) are not explicitly interpreted as clinically meaningful. The primary outcome (overall survival) showed no effect, so the meaningfulness of the secondary effects is not established.
“Time to deterioration in physical function was statistically significantly longer in the PRO group compared to the control group, with a median time of 12.6 months versus 8.5 months (HR 0.73; P = 0.002).”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported multicenter cluster-randomized trial. The methods are rigorous, with clear randomization, power analysis, and prespecified criteria, and the reporting is thorough with trial registration, data availability, and limitations discussed. Minor reporting gaps include lack of explicit blinding rationale, outlier handling, and a Declaration of Helsinki statement, plus a few copyedit issues.
Both reviewers independently scored all eight dimensions as pass, with high agreement. The study is an interventional cluster-randomized trial; no divergence in study type. Non-applicable criteria (e.g., animal housing, cell line authentication) were excluded. The statistics verification covered only 2 tests due to limited machine-verifiable statistics; the rest were not independently recomputed.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks.
- CONSISTENTreported p = .860 · recomputed p = .909Reviewers 1, 2Overall survival HR p-value from CI
“There was no statistically significant difference between study groups in this outcome at 2 years with an HR (HR) of 0.99 (95% confidence interval (CI) 0.83–1.17, P = 0.86).”
Taken as given: The HR is 0.99 and the 95% CI is 0.83 to 1.17.; The CI is two-sided at 95%.; The p-value is two-sided.Method: Recomputed p-value from the reported HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.99, 0.83, 1.17, 1) - CONSISTENTreported p = .030 · recomputed p = .034Reviewers 1, 2Emergency department visit HR p-value from CI
“The time to first emergency department visit was statistically significantly prolonged (improved) in the PRO intervention group compared to control, with an HR of 0.84 (95% CI, 0.71–0.98, two-sided P = 0.03).”
Taken as given: The HR is 0.84 and the 95% CI is 0.71 to 0.98.; The CI is two-sided at 95%.; The p-value is two-sided.Method: Recomputed p-value from the reported HR and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.84, 0.71, 0.98, 1)
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
6 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Symptom monitoring with electronic patient-reported outcomes improves clinical outcomes, the patient experience and utilization of services.The claim is supported by the trial's secondary outcomes showing significant benefits in emergency visits, physical function, symptom control, and HRQL, as well as high patient satisfaction.Evidence: Secondary outcomes: time to first emergency visit HR 0.84 (95% CI 0.71-0.98, P=0.03); time to deterioration in physical function HR 0.73 (P=0.002); symptom control HR 0.69 (P<0.001); HRQL HR 0.72 (P=0.001); patient satisfaction percentages.
“These findings demonstrate that symptom monitoring with PRO meaningfully improves clinical outcomes, the patient experience and utilization of services and should be included as a standard part of quality cancer clinical care.”
AbstractFind in source - supportedReviewers 1, 2There was no difference in overall survival between PRO and usual care.The primary outcome analysis showed no statistically significant difference, with HR 0.99 (95% CI 0.83-1.17, P=0.86).Evidence: Primary outcome: HR 0.99 (95% CI 0.83-1.17, P=0.86).
“Among 1,191 enrolled patients, there was no difference in survival (hazard ratio (HR) 0.99 (95% confidence interval (CI), 0.83–1.17); P = 0.86).”
AbstractFind in source - supportedReviewers 1, 2Time to first emergency visit was significantly prolonged with PRO compared to usual care.The secondary outcome analysis showed a statistically significant HR of 0.84 (95% CI 0.71-0.98, P=0.03).Evidence: Secondary outcome: HR 0.84 (95% CI 0.71-0.98, P=0.03).
“Time to first emergency visit was significantly prolonged with PRO compared to usual care (HR 0.84 ((95% CI, 0.71–0.98); P = 0.03), with a 6.1% reduction in the cumulative incidence of emergency visits and fewer mean visits at 12 months with PRO (1.02 versus 1.30; P < 0.001).”
AbstractFind in source - supportedReviewers 1, 2Benefits significantly favored PRO for delayed deterioration of physical function, symptoms, and HRQL.The time-to-deterioration analyses showed statistically significant improvements in all three domains, with HRs ranging from 0.69 to 0.73.Evidence: Time to deterioration: physical function HR 0.73 (P=0.002), symptom control HR 0.69 (P<0.001), HRQL HR 0.72 (P=0.001).
“Benefits also significantly favored PRO for delayed deterioration of physical function (median 12.6 versus 8.5 months, HR 0.73; P = 0.002), symptoms (12.7 versus 9.9, HR 0.69; P < 0.001) and HRQL (15.6 versus 12.2, HR 0.72; P = 0.001), which remained significant when considering deaths in analyses.”
AbstractFind in source - supportedReviewers 1, 2Most patients felt that PRO improved discussions with the care team, made them feel more in control, and would recommend it.The satisfaction survey results show high percentages supporting these claims.Evidence: Satisfaction: 77.0% (188/244) improved discussions, 84.0% (205/244) felt more in control, 91.4% (223/244) would recommend.
“Most patients felt that PRO improved discussions with the care team (77.0% (188/244)), made them feel more in control of their care (84.0% (205/244)) and would recommend it to other patients (91.4% (223/244)).”
AbstractFind in source - supportedReviewer 1Future studies of PRO in clinical care should focus on these outcomes rather than mortality as primary endpoints.The paper's findings of no survival benefit but significant quality-of-life benefits support this recommendation.Evidence: The trial showed no survival difference but significant benefits in secondary outcomes.
“Future studies of PRO in clinical care should focus on these outcomes rather than mortality as primary endpoints.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on secondary outcomes including time to deterioration of physical function, symptom control, and HRQL, which are patient-reported outcome measures (surrogates for clinical benefit). The paper does not provide evidence of target engagement at the tested dose (since it's a behavioral intervention, not a drug) nor does it cite validated evidence linking these PRO measures to hard clinical outcomes. The primary outcome of overall survival was not significant, so the efficacy claim rests on these surrogate endpoints.
“Benefits also significantly favored PRO for delayed deterioration of physical function (median 12.6 versus 8.5 months, HR 0.73; P = 0.002), symptoms (12.7 versus 9.9, HR 0.69; P < 0.001) and HRQL (15.6 versus 12.2, HR 0.72; P = 0.001)”
- INADEQUATEEffect sizeThe reported effect sizes for the surrogate outcomes (e.g., HR 0.73 for physical function deterioration) are presented as statistically significant but are not anchored to a minimal clinically important difference or to a clear biological/clinical meaningfulness. The paper does not provide a threshold for clinical significance for these HRs, and the absolute differences in median times (e.g., 4.1 months for physical function) are not explicitly interpreted as clinically meaningful. The primary outcome (overall survival) showed no effect, so the meaningfulness of the secondary effects is not established.
“Time to deterioration in physical function was statistically significantly longer in the PRO group compared to the control group, with a median time of 12.6 months versus 8.5 months (HR 0.73; P = 0.002).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites multiple prior trials and observational studies (STAR, CAPRI, Canadian population-based study) and discusses their strengths and limitations. The rationale for the trial is clearly linked to the objective of evaluating PRO symptom monitoring in a national US community oncology setting. The paper also addresses limitations of prior research, such as the single-center nature of STAR and the need for a multicenter trial.
“This trial has several limitations. Benefits may not have been experienced by all patients, and future research could aim to identify if certain subpopulations may particularly benefit.”
Randomization method (permuted blocks, stratified by rural/urban) and unit (oncology practices) are clearly described. Blinding is not explicitly described, but the cluster-randomized design makes blinding of practices infeasible; the paper does not state whether outcome assessors were blinded, which is a minor gap. A power analysis is provided with effect size, alpha, and power. Inclusion/exclusion criteria are prespecified. Outlier handling is not explicitly discussed, but the analysis population is defined (all randomized patients). Controls are the usual care group. Independent replication is not applicable for a single pivotal trial.
“Participating practices were randomly assigned 1:1 to electronic PRO symptom monitoring (intervention) or usual care (control) using permuted blocks with block sizes of 2 or 4 and stratified by rural/urban based on US Census Bureau criteria.”
“Each practice could enroll up to 50 consecutively approached adults (aged 21 years or older) with any type of metastatic cancer receiving outpatient systemic antineoplastic treatment (including immuno-therapy, targeted oral therapy or chemotherapy), if they understood English, Spanish or Mandarin.”
“Participating practices were randomly assigned 1:1 to electronic PRO symptom monitoring (intervention) or usual care (control) using permuted blocks with block sizes of 2 or 4 and stratified by rural/urban based on US Census Bureau criteria.”
Sex, age, race, ethnicity, education, employment, rural location, marital status, technology use, and cancer type are reported in Table 1. Health status is implied by cancer type and line of therapy. Species/strain and housing conditions are not applicable for a human trial. Demographics are comprehensive.
“Age, median (range), years | 64 (29–89) | 62 (28–93) | | Sex | | Female | 359/593 (60.5%) | 335/597 (56.1%)”
The paper states that all patients signed written informed consent and that the protocol was approved by the University of North Carolina IRB (IRB Number 17–1864), Quorum central IRB (IRB number 32498), and Advarra central IRB (IRB number Pro00043507). Regulatory compliance is implied by adherence to IRB approvals, though no explicit statement of compliance with the Declaration of Helsinki is present, but this is not required for adequacy.
“The protocol and consent were approved by the Institutional Review Board (IRB) of the University of North Carolina (IRB Number 17–1864), as well as the Quorum central IRB (IRB number 32498) and the Advarra central IRB (IRB number Pro00043507).”
“All patient participants signed written informed consent.”
“The protocol and consent were approved by the Institutional Review Board (IRB) of the University of North Carolina (IRB Number 17–1864), as well as the Quorum central IRB (IRB number 32498) and the Advarra central IRB (IRB number Pro00043507).”
“All patient participants signed written informed consent.”
The PRO monitoring system is described in detail, including the PRO-CTCAE item library and other instruments. The EORTC QLQ-C30 is identified. Statistical software (SAS v9.4) is named. No antibodies, cell lines, or mycoplasma testing are applicable. The PRO system is not a drug/device but a software intervention; it is adequately described.
“Statistical testing was two sided, with P values < 0.05 considered statistically significant, and carried out in SAS v9.4 (SAS Institute).”
“Statistical testing was two sided, with P values < 0.05 considered statistically significant, and carried out in SAS v9.4 (SAS Institute).”
All statistical tests are named (Cox regression, Fine-Gray competing risk, mixed models). Assumptions are handled by design (e.g., competing risk for death). Exact p-values are reported for primary and secondary outcomes. Effect sizes with confidence intervals are reported. Software is identified. Data presentation includes Kaplan-Meier curves and per-group n. Mathematical plausibility checks were not possible for most outcomes due to continuous data and model-based estimates, but no obvious errors were found.
“The outcome of overall survival was analyzed via Cox regression using prespecified covariates of months since initial diagnosis to development of metastases, months since initial diagnosis to first systemic cancer treatment, line of sustemic cancer treatment, months since developing metastases to date of trial enrollment and a random effect for site clustering.”
“There was no statistically significant difference between study groups in this outcome at 2 years with an HR (HR) of 0.99 (95% confidence interval (CI) 0.83–1.17, P = 0.86).”
“there was no difference in survival (hazard ratio (HR) 0.99 (95% confidence interval (CI), 0.83–1.17); P = 0.86).”
The data availability statement provides a specific mechanism for requesting deidentified participant data via the Alliance's data sharing process, including a URL and contact email. This is a managed-access route, which is appropriate for patient-level data. No code was generated for analysis, so code sharing is not applicable.
“Individual deidentified participant data for this trial, including data dictionaries and all variables from analyses in this publication, are available through the Alliance for Clinical Trials in Oncology. Data Sharing Requests may be submitted at: https://www.allianceforclinicaltrialsinoncology.org/main/public/standard.xhtml?path=%2FPublic%2FDatasharing”
“Individual deidentified participant data for this trial, including data dictionaries and all variables from analyses in this publication, are available through the Alliance for Clinical Trials in Oncology. Data Sharing Requests may be submitted at: https://www.allianceforclinicaltrialsinoncology.org/main/public/standard.xhtml?path=%2FPublic%2FDatasharing”
The trial is registered at ClinicalTrials.gov (NCT03249090). Methods are detailed enough for replication. A reporting guideline is referenced (Nature Portfolio Reporting Summary). All prespecified outcomes are reported, including negative results (overall survival). Limitations are thoroughly discussed. Conclusions are proportional to the evidence. Funding and competing interests are disclosed.
“ClinicalTrials.gov (http://ClinicalTrials.gov) registration: NCT03249090 (https://clinicaltrials.gov/ct2/show/NCT03249090)”
“This trial has several limitations. Benefits may not have been experienced by all patients, and future research could aim to identify if certain subpopulations may particularly benefit.”
“ClinicalTrials.gov (http://ClinicalTrials.gov) registration: NCT03249090 (https://clinicaltrials.gov/ct2/show/NCT03249090)”
“This trial has several limitations. Benefits may not have been experienced by all patients, and future research could aim to identify if certain subpopulations may particularly benefit.”
“This trial was funded by the Patient-Centered Outcomes Research Institute (PCORI) IHS-1511-33392, with research support from this grant provided to all authors listed on this manuscript except V.B. and A.W.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 44 references by DOI: 39 verified — 5 no DOI (shown, not verified).
- NO DOIThe PROTEUS Guide to Implementing Patient-Reported Outcomes in Clinical Practice covers design, implementation, and management of PRO systems in clinical careNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEstablishing Effective Patient Navigation Programs in Oncology: Proceedings of a WorkshopNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPatient-Reported Outcomes Core of the University of North CarolinaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPatient-Reported Outcomes version of the Common Terminology Criteria for Adverse EventsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe EORTC QLQ-C30 Scoring Manual (3rd Edition)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://www.allianceforclinicaltrialsinoncology.org/main/public/standard.xhtml?path=%2FPublic%2FDatasharingLIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORtypoMethods, Statistical analysis“line of sustemic cancer treatment”→ line of systemic cancer treatmentTypo: 'sustemic' should be 'systemic'.
- MINORconsistencyResults, Emergency department visits“The proportion of patients with zero, one, two, three and four or more emergency department visits was 315/593 (53.1%), 136/593 (22.9%), 74/593 (12.5%), 30/593 (5.1%) and 38/593 (6.4%), respectively, in the PRO group”→ Check that the percentages sum to 100% (53.1+22.9+12.5+5.1+6.4 = 100.0).Percentages sum correctly, but the counts sum to 593, which is correct.
- MINORconsistencyAbstract vs. Results“77.0% (188/244)”→ Ensure consistency in satisfaction percentages between abstract and results.The abstract reports 77.0% (188/244) for improved discussions, while the results section reports 72.5% (359/495) at 3 months and 77.0% (188/244) at off-study; the abstract uses the off-study value without specifying the time point.
- MINORclarityMethods, Statistical analysis“The study database was frozen on 4 October 2022.”→ Consider adding a sentence explaining the database freeze in relation to the analysis.The sentence is clear but could be expanded for context.
The published work is robust and well-reported; an informed reader should weigh the minor reporting gaps (blinding rationale, outlier handling, explicit ethics framework statement) as areas for potential clarification but they do not undermine the study's conclusions. No erratum appears warranted based on the checks performed; the copyedit issues are minor and do not affect scientific integrity.
- 1.HIGHreportingAdd an explicit statement in the Methods (Randomization or Limitations) explaining that blinding of practices and patients was not feasible due to the cluster-randomized design, and acknowledge this as a potential source of bias in the Discussion.Both reviewers flagged the lack of explicit blinding rationale as a minor gap; adding it improves transparency and addresses a common reviewer concern.
- 2.HIGHreportingAdd a sentence in the Methods (Statistical analysis) describing how outliers and missing data were handled, or state that no outliers were excluded and missing data were handled as per the analysis population.Both reviewers noted outlier handling was not explicitly discussed; clarifying this strengthens the statistical reporting.
- 3.MEDIUMethicsAdd a statement in the Ethics section confirming compliance with the Declaration of Helsinki or other relevant ethical guidelines.Reviewer 2 flagged regulatory_compliance as reported_but_inadequate due to lack of explicit ethical framework statement; adding it resolves the minor gap.
- 4.MEDIUMcopyeditFix the typo 'sustemic' to 'systemic' in the Methods, Statistical analysis section.Copyedit pass flagged this typo; correcting it improves professionalism.
- 5.MEDIUMcopyeditClarify in the Abstract that the 77.0% (188/244) satisfaction figure refers to the off-study assessment, not the 3-month assessment, to avoid ambiguity.Copyedit pass noted inconsistency between abstract and results regarding the time point for satisfaction percentages.
- 6.LOWdata codeConsider providing the statistical analysis code (e.g., SAS scripts) in a public repository to enhance reproducibility.Reviewer 2 suggested sharing code; while not required, it would strengthen reproducibility.
- 7.LOWreportingAdd a note in the Data Availability section about the expected timeframe for responding to data requests.Reviewer 2 suggested this to set expectations for data requesters.
- 8.LOWreportingConsider including a CONSORT flow diagram in the main text to enhance reporting transparency.Reviewer 1 suggested this; while not essential, it improves adherence to reporting guidelines.
- 9.LOWreportingProvide the statistical analysis plan as a supplementary file to enhance transparency.Reviewer 1 suggested this; it would allow readers to verify prespecified analyses.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.