Symptom Screening Linked to Care Pathways for Pediatric Patients With Cancer: A Randomized Clinical Trial.
Dupuis LL, Vettese E, Grimes AC, Beauchemin MP, Klesges LM, Baggott C, Demedis J, Aftandilian C, Freyer DR, Crellin-Parsons N, Orgel E, Dickens D, Kelly KM, Kyono W, Walsh A, Sherani F, Cannone D, Orsey AD, King AA, Yu L, Woods-Swafford W, Bradfield SM, Roth ME, Esbenshade AJ, Caywood EH, Agarwal V, Nagasubramanian R, Tomlinson GA, Sung L
- DOI
- 10.1001/jama.2024.19585
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/c5d2eefc-fa14-4ba3-9130-894ff78cad93 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingEthical approvals partially met−0.25★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 4 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is the total SSPedi score, a patient-reported symptom burden measure, which is a surrogate for clinical benefit. The paper does not demonstrate target engagement at the tested dose (since it's a behavioral intervention, not a drug) and does not cite validated evidence linking SSPedi score changes to hard clinical outcomes. The claim of improved symptom scores is presented as evidence of efficacy, but the surrogate is not validated as a surrogate for long-term clinical outcomes.
“The primary outcome was self-reported total SSPedi score at week 8 (range, 0-60; higher scores indicate more bothersome).”
- 02Treatment effect not shown to be clinically meaningful
The adjusted mean difference in SSPedi score is -3.8 points on a 0-60 scale, which is a small fraction of the scale range (about 6.3%). The paper does not anchor this difference to a minimal clinically important difference or other clinical meaningfulness, and the difference is small relative to the scale.
“adjusted mean difference, −3.8 [95% CI, −6.4 to −1.2]”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted cluster randomized trial with a clear premise, appropriate design, and transparent reporting of funding and registration. The main weaknesses are missing ethics approval/consent statements, a vague data sharing statement, and incomplete reporting of statistical details (exact p-values, software).
Both reviewers agreed on all dimensions; no divergence to reconcile. The statistics verification covered only 4 tests with test statistics/CIs; other p-values were not machine-verified. The citation check found no retracted or non-existent references.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 4 tests: 4 consistent, 0 inconsistent; 4 via agent-written checks.
- CONSISTENTreported p < .004 · recomputed p = .004Reviewer 1Check adjusted mean difference p-value from CI
“adjusted mean difference, −3.8 [95% CI, −6.4 to −1.2]”
Taken as given: The CI is a 95% confidence interval for the adjusted mean difference.; The estimate is -3.8 and the CI is symmetric on the linear scale.Method: Two-sided p-value derived from the confidence interval using normal approximation.How we recomputed it: pCI(-3.8, -6.4, -1.2, 0) - CONSISTENTreported p < .040 · recomputed p = .038Reviewer 1Check rate ratio p-value for ED visits
“rate ratio, 1.72 [95% CI, 1.03-2.87]”
Taken as given: The CI is a 95% confidence interval for the rate ratio.; The estimate is 1.72 and the CI is on the log scale.Method: Two-sided p-value derived from the confidence interval using log-normal approximation.How we recomputed it: pCI(1.72, 1.03, 2.87, 1) - CONSISTENTreported p < .050 · recomputed p = .004Reviewer 2Adjusted mean difference for primary outcome
“adjusted mean difference, −3.8 [95% CI, −6.4 to −1.2]”
Taken as given: The CI is a 95% confidence interval.; The estimate is a mean difference (not a ratio).Method: Two-sided p-value derived from the confidence interval using normal approximation.How we recomputed it: pCI(-3.8, -6.4, -1.2, 0) - CONSISTENTreported p < .050 · recomputed p = .038Reviewer 2Rate ratio for emergency department visits
“rate ratio, 1.72 [95% CI, 1.03-2.87]”
Taken as given: The CI is a 95% confidence interval.; The estimate is a rate ratio (log scale).Method: Two-sided p-value derived from the confidence interval using normal approximation on the log scale.How we recomputed it: pCI(1.72, 1.03, 2.87, 1)
- lowinternal contradictionThe abstract reports 221 participants at intervention sites and 224 at control sites, totaling 445, which matches the total. However, the results section states 445 participants were enrolled, which is consistent.
“221 participants were enrolled at intervention sites and 224 participants at control sites.”
Abstract
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
6 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1Symptom screening improves symptom-specific interventions.The abstract mentions increased symptom-specific interventions, but no specific data are provided in the abstract; the full text may contain details.Evidence: Not detailed in abstract; likely in results.
“Symptom screening with symptom feedback and symptom management care pathways was associated with improved symptom scores and increased symptom-specific interventions.”
Conclusion - supportedReviewers 1, 2Symptom screening with symptom feedback and care pathways improves total SSPedi scores compared with usual care.The primary outcome shows a statistically significant adjusted mean difference of -3.8 with 95% CI excluding zero.Evidence: Adjusted mean difference -3.8 (95% CI -6.4 to -1.2) in total SSPedi score.
“The total 8-week SSPedi score (range, 0-60) was significantly better with symptom screening compared with usual care (7.9 vs 11.4, respectively; adjusted mean difference, −3.8).”
Abstract - supportedReviewer 1Symptom screening reduces bothersome individual symptoms.The paper reports that 12 of 15 symptoms were statistically significantly reduced, supporting the claim.Evidence: 12 of 15 symptoms statistically significantly reduced.
“Symptom screening was associated with significantly better 8-week total SSPedi scores (adjusted mean difference, −3.8 [95% CI, −6.4 to −1.2]) and less bothersome individual symptoms, with 12 of 15 symptoms being statistically significantly reduced.”
Abstract - supportedReviewers 1, 2Symptom screening increases emergency department visits.The rate ratio of 1.72 with 95% CI 1.03-2.87 indicates a statistically significant increase, supporting the claim.Evidence: Rate ratio 1.72 (95% CI 1.03-2.87) for ED visits.
“There were significantly more emergency department visits in the symptom screening group (rate ratio, 1.72 [95% CI, 1.03-2.87]).”
Abstract - supportedReviewer 2Symptom screening is associated with less bothersome individual symptoms, with 12 of 15 symptoms being statistically significantly reduced.The paper reports that 12 of 15 symptoms were significantly reduced, which supports the claim.Evidence: Reported in Results: '12 of 15 symptoms being statistically significantly reduced.'
“Symptom screening was associated with significantly better 8-week total SSPedi scores (adjusted mean difference, −3.8 [95% CI, −6.4 to −1.2]) and less bothersome individual symptoms, with 12 of 15 symptoms being statistically significantly reduced.”
Results - supportedReviewer 2There is no difference in fatigue or quality of life.The paper states there was no difference in fatigue or quality of life, which is a negative result reported.Evidence: Reported in Results: 'There was no difference in fatigue or quality of life.'
“There was no difference in fatigue or quality of life.”
Results
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is the total SSPedi score, a patient-reported symptom burden measure, which is a surrogate for clinical benefit. The paper does not demonstrate target engagement at the tested dose (since it's a behavioral intervention, not a drug) and does not cite validated evidence linking SSPedi score changes to hard clinical outcomes. The claim of improved symptom scores is presented as evidence of efficacy, but the surrogate is not validated as a surrogate for long-term clinical outcomes.
“The primary outcome was self-reported total SSPedi score at week 8 (range, 0-60; higher scores indicate more bothersome).”
- INADEQUATEEffect sizeThe adjusted mean difference in SSPedi score is -3.8 points on a 0-60 scale, which is a small fraction of the scale range (about 6.3%). The paper does not anchor this difference to a minimal clinically important difference or other clinical meaningfulness, and the difference is small relative to the scale.
“adjusted mean difference, −3.8 [95% CI, −6.4 to −1.2]”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
2 findings · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
- Ethics/consent reporting incompleteAssessed
The introduction cites prior research on symptom burden in pediatric cancer and the lack of evidence for routine symptom screening. The rationale linking symptom screening to improved outcomes is logical and the hypothesis follows from the cited evidence. Limitations of prior research are implicitly addressed by the study design, though not explicitly detailed.
“Pediatric patients with cancer commonly experience severely bothersome symptoms. The effectiveness of routine symptom screening with symptom feedback and symptom management care pathways is unknown.”
“To determine whether thrice-weekly symptom screening with symptom feedback and management care pathways, compared with usual care, improves overall self-reported symptom scores measured by the Symptom Screening in Pediatrics Tool (SSPedi) in pediatric patients with cancer.”
“Pediatric patients with cancer commonly experience severely bothersome symptoms.”
“The effectiveness of routine symptom screening with symptom feedback and symptom management care pathways is unknown.”
The trial is a cluster randomized trial with sites randomized to intervention or control. Randomization method is not explicitly described but is implied. Blinding is not mentioned, but for a behavioral intervention, blinding of participants is infeasible; however, outcome assessors could have been blinded. Power analysis is not reported. Inclusion/exclusion criteria are stated. Outlier handling is not explicitly addressed. Controls are the usual care group. Independent replication is not applicable for a single trial.
“Twenty sites were randomized to provide symptom screening (n = 10) vs usual care (n = 10)”
“Patients newly diagnosed with cancer aged 8 to 18 years receiving any cancer treatment were included.”
“This cluster randomized trial enrolled participants between July 2021 and August 2023 from 20 pediatric cancer centers in the US.”
“Patients newly diagnosed with cancer aged 8 to 18 years receiving any cancer treatment were included.”
The study reports age and sex of participants. Since both sexes are enrolled, sex_justified is not applicable. Age and health status are reported. Demographics include age and sex but not race/ethnicity or comorbidities. Species/strain and housing conditions are not applicable for a human trial.
“A total of 445 participants (median [range] age, 14.8 [8.1-18.9] years; 58.9% males) were enrolled.”
“median [range] age, 14.8 [8.1-18.9] years”
The paper does not mention an IRB approval or ethics committee statement. It also does not describe informed consent. Since this is a human interventional trial, these are required. The absence of these statements is a significant reporting gap, though it does not necessarily indicate misconduct.
The intervention is described as symptom screening with care pathways, but no specific product or manufacturer is applicable. The SSPedi tool is named. Statistical software is not explicitly identified. No antibodies, cell lines, or organisms are used. The trial uses a behavioral intervention, so most bench criteria are not applicable. The intervention itself is adequately described.
“Symptom screening included providing thrice-weekly symptom screening prompts to participants, email alerts to the health care team, and locally adapted symptom management care pathway implementation.”
“Symptom screening included providing thrice-weekly symptom screening prompts to participants, email alerts to the health care team, and locally adapted symptom management care pathway implementation.”
The primary analysis uses adjusted mean difference with 95% CI, which is a complete reporting style. Tests are not explicitly named but are implied. Assumptions are not explicitly verified. Exact p-values are not reported for all outcomes; some are reported as thresholds. Effect sizes with CIs are reported. Statistical software is not identified. Data presentation includes means and SDs, but individual data points are not shown. Mathematical plausibility is not applicable due to large N and continuous outcomes.
“adjusted mean difference, −3.8 [95% CI, −6.4 to −1.2]”
“with 12 of 15 symptoms being statistically significantly reduced”
“adjusted mean difference, −3.8 [95% CI, −6.4 to −1.2]”
The paper mentions a Data Sharing Statement but does not provide details in the text. It likely refers to a supplementary document. No repository deposit or accession numbers are provided. No code sharing is mentioned. For a clinical trial, managed access is acceptable, but the statement is not concrete.
“Data Sharing Statement: See .”
“Data Sharing Statement: See .”
The trial is registered with ClinicalTrials.gov. Methods are described in sufficient detail. Reporting guidelines are not mentioned. All outcomes appear to be reported. Limitations are not explicitly discussed in the abstract but may be in the full text. Conclusions are proportional to the evidence. Funding and COI are disclosed.
“ClinicalTrials.gov Identifier: NCT04614662”
“The funding for this study was provided by a project grant from the Canadian Institutes of Health Research (PJT-169165) and the National Institutes of Health (R01CA251112).”
“Trial Registration ClinicalTrials.gov Identifier: NCT04614662”
“Funding/Support: The funding for this study was provided by a project grant from the Canadian Institutes of Health Research (PJT-169165) and the National Institutes of Health (R01CA251112).”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 1 reference by DOI: 1 verified.
Every extracted reference resolved against Crossref/OpenAlex with no retraction flags.
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAuthor list“Dupuis L. Lee RPh PhD”→ Correct the name to 'Lee Dupuis, RPh, PhD'Name appears reversed.
- MINORconsistencyKey Points“Findings The total 8-week SSPedi score (range, 0-60) was significantly better with symptom screening compared with usual care (7.9 vs 11.4, respectively; adjusted mean difference, −3.8).”→ Add 95% CI for consistency with abstract.CI omitted in Key Points.
- MINORclarityData Sharing Statement“Data Sharing Statement: See .”→ Provide the actual data sharing statement or a link.Incomplete sentence.
- MINORtypoAuthor list“Chlidren's Oncology Group”→ Children's Oncology GroupTypo in conflict of interest disclosure.
- MINORconsistencyAbstract“The mean (SD) number of emergency department visits was 0.77 (1.12) in the symptom screening group and 0.45 (0.81) in the usual care group.”→ Ensure consistency in reporting of emergency department visits across abstract and results.The abstract reports ED visits, but the results section may have more detail.
The published work is robust in design and analysis, but readers should weigh the missing ethics approval/consent statements and the vague data sharing statement. These are reporting gaps that warrant a correction or clarification from the authors, but they do not undermine the core findings.
- 1.HIGHethicsAdd an explicit ethics approval statement in the Methods, naming the IRB and protocol number.A clinical trial without a stated ethics approval is a serious reporting gap that readers and journals will question.
- 2.HIGHethicsDescribe the informed consent process in the Methods, including whether written consent was obtained or waived.Informed consent is a fundamental ethical requirement for human research; its absence is a major omission.
- 3.HIGHdata codeReplace the incomplete 'Data Sharing Statement: See .' with a concrete statement specifying how data can be accessed (e.g., repository link or contact for data access committee).The current statement is incomplete and does not meet transparency standards for a data-driven clinical trial.
- 4.HIGHstatisticsProvide exact p-values for all primary and secondary outcomes in the Results, instead of only thresholds.Exact p-values allow readers to assess the strength of evidence and are expected in clinical trial reporting.
- 5.HIGHstatisticsIdentify the statistical software and version used for analysis in the Methods.Software identification is part of reproducible research and is commonly required by journals.
- 6.HIGHotherReport the power analysis or sample size calculation in the Methods, including assumed effect size and power.A power analysis is essential for interpreting the study's ability to detect effects and is a standard reporting requirement.
- 7.HIGHotherSpecify the randomization method (e.g., computer-generated random sequence) and whether allocation was concealed.Details on randomization and allocation concealment are critical for assessing risk of bias.
- 8.HIGHotherState whether outcome assessors were blinded to group allocation, or provide a rationale for lack of blinding.Blinding of outcome assessors reduces bias; its absence should be disclosed and justified.
- 9.MEDIUMreportingMention adherence to a reporting guideline such as CONSORT in the Methods or Acknowledgments.Reporting guidelines improve completeness and are often required by journals.
- 10.MEDIUMreportingAdd a limitations section discussing potential biases, such as lack of blinding and cluster-level randomization.A thorough limitations discussion helps readers interpret the findings appropriately.
- 11.MEDIUMotherReport race/ethnicity and comorbidities in the baseline characteristics table.These demographic variables are important for generalizability and are often expected in clinical trials.
- 12.MEDIUMstatisticsClarify the handling of missing data and outliers in the statistical analysis section.Transparency about missing data and outlier handling is essential for reproducibility.
- 13.LOWcopyeditCorrect the author name 'Dupuis L. Lee RPh PhD' to 'Lee Dupuis, RPh, PhD' in the author list.The name appears reversed, which is a typographical error that should be fixed.
- 14.LOWcopyeditCorrect the typo 'Chlidren's Oncology Group' to 'Children's Oncology Group' in the conflict of interest disclosure.A simple typo that should be corrected for professionalism.
- 15.LOWcopyeditAdd the 95% CI to the Key Points finding for consistency with the abstract.The Key Points omit the CI that is reported in the abstract, which is a consistency issue.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.