Personalized Patient Data and Behavioral Nudges to Improve Adherence to Chronic Cardiovascular Medications: A Randomized Pragmatic Trial.
Ho PM, Glorioso TJ, Allen LA, Blankenhorn R, Glasgow RE, Grunwald GK, Khanna A, Magid DJ, Marrs JC, Novins-Montague S, Orlando S, Peterson P, Plomondon ME, Sandy LM, Saseen JJ, Trinkley KE, Vaughn S, Waughtal J, Bull S
- DOI
- 10.1001/jama.2024.21739
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/d8d81623-0a0c-4ff8-9f55-791c845f5889 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- StatisticsPrinted percentage does not match its own count (capped)−0.25★
- ReportingEthical approvals partially met−0.25★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 4 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- No data or code availability links were detected to verify.
- 01Printed percentage does not match its own count
47% does not match the reported count 4351/9501
“47% female [n = 4351]”
Results
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted pragmatic randomized trial with a clear premise, sound design, and transparent reporting of results, funding, and registration. The main weaknesses are missing ethics approval and consent statements, a vague data sharing statement, and incomplete reporting of randomization, blinding, power analysis, and statistical software.
Both reviewers classified the study as interventional and agreed on all dimension statuses; no divergence required resolution. The statistics verification checked only a subset of reported tests (4 tests, 3 consistent, 1 unspecified inconsistency), so the statistical analysis is not fully verified. The copyedit pass flagged minor typographical and formatting issues.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 3 via agent-written checks. 1 printed percentage that does not match its own count.
- PERCENT47% does not match the reported count 4351/9501
“47% female [n = 4351]”
Results
- CONSISTENTreported p = .020 · recomputed p = .027Reviewers 1, 2Check p-value for generic reminder vs usual care adjusted difference
“mean proportion of days covered was 2.2 percentage points (95% CI, 0.3-4.2; P = .02) higher for generic reminder”
Taken as given: The estimate is 2.2 percentage points.; The 95% CI is 0.3 to 4.2.; The p-value is two-sided.Method: Recomputed p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(2.2, 0.3, 4.2, 0) - CONSISTENTreported p = .040 · recomputed p = .039Reviewers 1, 2Check p-value for behavioral nudge vs usual care adjusted difference
“2.0 percentage points (95% CI, 0.1-3.9; P = .04) higher for behavioral nudge”
Taken as given: The estimate is 2.0 percentage points.; The 95% CI is 0.1 to 3.9.; The p-value is two-sided.Method: Recomputed p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(2.0, 0.1, 3.9, 0) - CONSISTENTreported p = .020 · recomputed p = .018Reviewers 1, 2Check p-value for behavioral nudge + chatbot vs usual care adjusted difference
“2.3 percentage points (95%, 0.4-4.2; P = .02) higher for behavioral nudge + chatbot”
Taken as given: The estimate is 2.3 percentage points.; The 95% CI is 0.4 to 4.2.; The p-value is two-sided.Method: Recomputed p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(2.3, 0.4, 4.2, 0)
- lowinternal contradictionThe abstract reports a p-value of .06 for the unadjusted comparison of mean PDC across groups, but the adjusted analyses show p-values of .02, .04, and .02. This is not contradictory but may confuse readers.
“At 12 months, the mean proportion of days covered was 62.0% for generic reminder, 62.3% for behavioral nudge, 63.0% for behavioral nudge + chatbot, and 60.6% for usual care ( P = .06).”
Abstract
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
3 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Text message reminders did not improve medication adherence or reduce clinical events at 12 months.The primary outcome showed no statistically significant differences after multiple comparisons correction, and no differences in clinical events.Evidence: Adjusted mean differences were small and not significant after correction; no differences in clinical events.
“Text message reminders targeting patients who delay refilling their cardiovascular medications did not improve medication adherence based on pharmacy refill data or reduce clinical events at 12 months.”
Conclusion - supportedReviewers 1, 2The three text messaging strategies tested did not increase refill adherence at 12 months.The adjusted differences were small and not significant after multiple comparisons correction.Evidence: Adjusted mean differences ranged from 2.0 to 2.3 percentage points with p-values .02-.04, but not significant after correction.
“the 3 text messaging medication refill reminder strategies tested (generic reminders, behavioral nudge reminders, and behavioral nudge reminders plus a fixed-message chatbot) did not increase refill adherence at 12 months or reduce clinical events.”
Key Points Finding - supportedReviewers 1, 2Additional interventions need to be rigorously tested to improve adherence.This is a reasonable conclusion given the null results.Evidence: The trial showed no benefit, suggesting need for further research.
“Additional interventions need to be rigorously tested to try to improve adherence to chronic cardiovascular medications given the growing incidence of cardiovascular conditions.”
Key Points Meaning
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
2 findings · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
- Ethics/consent reporting incompleteAssessed
The introduction cites prior work on medication adherence and text messaging, acknowledges that text messaging is often not rigorously tested, and provides a logical rationale for the trial. Limitations of prior research are implicitly addressed by the rigorous randomized design.
“Text messaging is increasingly used to change patient behavior but often not rigorously tested.”
“To compare different types of text messaging strategies with usual care to improve medication refill adherence among patients nonadherent to cardiovascular medications.”
“Text messaging is increasingly used to change patient behavior but often not rigorously tested.”
“To compare different types of text messaging strategies with usual care to improve medication refill adherence among patients nonadherent to cardiovascular medications.”
The trial is described as a patient-level randomized pragmatic trial with 4 groups. Randomization method is not explicitly detailed but is implied by the pragmatic design. Blinding is not mentioned, but in pragmatic trials blinding is often not feasible; however, the paper does not state this. Power analysis is not reported, but the large sample size (9501) suggests adequate power. Inclusion/exclusion criteria are described. Outlier handling is not explicitly discussed, but the analysis uses standard methods. Controls are the usual care group. Independent replication is not applicable for a single trial.
“Patient-level randomized pragmatic trial between October 2019 to April 2022 at 3 US health care systems”
“Adult (18 to <90 years) patients were eligible based on diagnosis of 1 or more cardiovascular condition(s) and prescribed medication to treat the condition.”
“Patients who did not opt out and had a 7-day refill gap were randomized to 1 of 4 study groups.”
“Adult (18 to <90 years) patients were eligible based on diagnosis of 1 or more cardiovascular condition(s) and prescribed medication to treat the condition.”
The paper reports mean age, sex distribution, and race/ethnicity percentages. Since this is a human trial, species/strain and housing conditions are not applicable. Age and health status are implied by the inclusion criteria (cardiovascular conditions).
“baseline characteristics across the 4 groups were comparable (mean age, 60 years; 47% female [n = 4351]; 16% Black [n = 1517]; 49% Hispanic [n = 4564])”
“baseline characteristics across the 4 groups were comparable (mean age, 60 years; 47% female [n = 4351]; 16% Black [n = 1517]; 49% Hispanic [n = 4564]).”
The paper does not mention an IRB approval or informed consent. Although it is a pragmatic trial using electronic health records and pharmacy data, it still involves human subjects and requires ethics oversight. The absence of an ethics statement is a concern.
The trial uses text messaging as the intervention, which is not a biological or chemical resource. There are no antibodies, cell lines, organisms, or reagents. The software used for randomization or analysis is not described, but it is not a key biological resource.
“Generic text message refill reminders (generic reminder); behavioral nudge text refill reminders (behavioral nudge); behavioral nudge text refill reminders plus a fixed-message chatbot (behavioral nudge + chatbot); usual care.”
The paper reports adjusted mean differences with 95% CIs and p-values. The primary analysis uses a linear model for proportion of days covered. The p-values are reported as exact values (e.g., P = .02). The paper does not explicitly name the statistical software, but it is implied. Data presentation includes means and CIs. Mathematical plausibility checks are not applicable due to large sample size and continuous outcomes.
“mean proportion of days covered was 2.2 percentage points (95% CI, 0.3-4.2; P = .02) higher for generic reminder”
“In adjusted analysis, when compared with usual care, mean proportion of days covered was 2.2 percentage points (95% CI, 0.3-4.2; P = .02) higher for generic reminder, 2.0 percentage points (95% CI, 0.1-3.9; P = .04) higher for behavioral nudge, and 2.3 percentage points (95%, 0.4-4.2; P = .02) higher for behavioral nudge + chatbot, none of which were statistically significant after multiple comparisons correction.”
The paper mentions a 'Data Sharing Statement: See .' but the actual statement is not included in the text provided. This is vague and does not specify a concrete access route. No repository or code is mentioned.
“Data Sharing Statement: See .”
“Data Sharing Statement: See .”
The paper includes a trial registration number, funding sources, and conflict of interest disclosures. It discusses limitations implicitly through the conclusion. The conclusions are proportional to the findings. However, it does not explicitly reference a reporting guideline like CONSORT.
“Trial Registration ClinicalTrials.gov Identifier: NCT03973931”
“Funding/Support: This work was supported within the National Institutes of Health (NIH) Pragmatic Trials Collaboratory by cooperative agreement UG3HL144163 from the NHLBI.”
“Trial Registration ClinicalTrials.gov Identifier: NCT03973931”
“Funding/Support: This work was supported within the National Institutes of Health (NIH) Pragmatic Trials Collaboratory by cooperative agreement UG3HL144163 from the NHLBI.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 1 reference by DOI: 1 verified.
Every extracted reference resolved against Crossref/OpenAlex with no retraction flags.
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly typo, consistency.
- MINORtypoKey Points section“test message reminders”→ text message remindersTypo in the Key Points question.
- MINORconsistencyResults section“95%, 0.4-4.2”→ 95% CI, 0.4-4.2Missing 'CI' in the confidence interval notation.
- MINORtypoKey Points“test message reminders”→ text message remindersTypo in the Key Points question.
- MINORconsistencyResults“95%, 0.4-4.2”→ 95% CI, 0.4-4.2Inconsistent formatting of confidence interval.
The published work is generally robust, but an informed reader should weigh the missing ethics approval/consent statement and the vague data sharing statement as reporting gaps that could warrant a correction or clarification. The minor copyedit issues (typos, CI formatting) are trivial and do not affect the scientific validity.
- 1.HIGHethicsAdd an explicit ethics approval statement in the Methods section, naming the IRB and protocol number, and describe the informed consent or waiver process.The paper currently lacks any ethics approval or consent statement, which is a critical reporting gap for a human trial.
- 2.HIGHdata codeProvide a detailed data availability statement in the manuscript, specifying where and how data can be accessed (e.g., a repository or data access committee).The current statement is vague ('See .') and does not provide a concrete access route, which is a reporting gap for a data-driven paper.
- 3.HIGHreportingDescribe the randomization method (e.g., computer-generated random sequence) and any blinding or lack thereof with rationale in the Methods section.The randomization method and blinding are not explicitly described, which is a major reviewer comment for a randomized trial.
- 4.HIGHstatisticsInclude a power analysis or sample size justification in the Methods to demonstrate the trial was adequately powered.The paper does not report a power analysis, which is a major reviewer comment for a clinical trial.
- 5.HIGHstatisticsMention the statistical software used for analyses (e.g., SAS, R) in the Methods.The statistical software is not identified, which is a reporting gap that affects reproducibility.
- 6.MEDIUMreportingReference a reporting guideline (e.g., CONSORT) in the manuscript to enhance transparency.The paper does not explicitly reference a reporting guideline, which is a minor transparency gap.
- 7.MEDIUMstatisticsClarify the handling of missing data and outliers in the statistical analysis section.The paper does not explicitly discuss outlier handling or missing data, which is a minor reporting gap.
- 8.MEDIUMreportingDiscuss limitations more explicitly, including potential biases and generalizability.Limitations are discussed but could be more explicit, which is a minor reviewer comment.
- 9.LOWcopyeditFix the typo 'test message reminders' to 'text message reminders' in the Key Points section.This is a minor typo that should be corrected for professionalism.
- 10.LOWcopyeditFix the confidence interval formatting '95%, 0.4-4.2' to '95% CI, 0.4-4.2' in the Results section.This is a minor formatting inconsistency that should be corrected for clarity.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.