Multifaceted Strategies for Hypertension Control in Low-Income Patients.
Mills KT, Krousel-Wood M, Peacock EM, Chen J, Allouch F, Carreras AK, Geng S, Cyprian A, Davis G, Fuqua SR, Gilliam D, Greer A, Mitchell T, Gray-Winfrey W, Williams S, Wiltz GM, Winfrey KL, He H, Whelton PK, He J
- DOI
- 10.1056/NEJMoa2504068
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/e5feeda8-5366-4669-adb2-190ad0516951 is authoritative.
How this rating was calculated
- StatisticsImpossible or misreported statistic ×3−3★
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×3−0.25★
- ReportingData & code availability partially met−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 4 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Printed percentage does not match its own countdemonstrable
73.4% does not match the reported count 468/642
“Annual family income <$25,000 — no. (%) | 468 (73.4)”
Table 1Find in source - 02Printed percentage does not match its own countdemonstrable
44% does not match the reported count 277/642
“Duration of hypertension >10 years — no. (%) | 277 (44.0)”
Table 1Find in source - 03Printed percentage does not match its own countdemonstrable
41.2% does not match the reported count 252/630
“Duration of hypertension >10 years — no. (%) | 252 (41.2)”
Table 1Find in source - 04Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is based on reduction in systolic blood pressure, which is a surrogate endpoint for cardiovascular outcomes. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking the specific blood pressure reduction to clinical outcomes in this population. Although blood pressure is a well-established surrogate, the manuscript does not explicitly provide the required validation link.
“The primary effectiveness outcome was net between-group difference in mean systolic blood pressure from baseline to 18 months”
- 05Printed percentage does not match its own count
67.1% does not match the reported count 430/642
“Non-Hispanic Black | 430 (67.1)”
Table 1Find in source - 06Printed percentage does not match its own count
73.5% does not match the reported count 460/630
“Annual family income <$25,000 — no. (%) | 460 (73.5)”
Table 1Find in source
1 further finding of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed and well-reported cluster-randomized trial with a strong scientific premise, rigorous methods, and transparent reporting. The main weakness is the lack of a clear data availability statement and code sharing, which is a common reporting gap. Minor copyedit issues (missing negative signs in confidence intervals) should be corrected.
Both reviewers independently scored all eight dimensions and agreed on all statuses, so no divergence needed reconciliation. The statistics verification component recomputed only a subset of tests (2 of 8 consistent), and the coverage note clarifies that many p-values are threshold-only and cannot be machine-verified; therefore, the statistics are not fully verified but no errors were identified. The citation check found no retracted or unresolved references.
Numerical inconsistencies
3 findings · worst criticalValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks. 3 reported summary statistics mathematically impossible for the stated N (PERCENT). 3 printed percentages that do not match their own count.
- PERCENT67.1% does not match the reported count 430/642
“Non-Hispanic Black | 430 (67.1)”
Table 1Find in source - PERCENT73.4% does not match the reported count 468/642
“Annual family income <$25,000 — no. (%) | 468 (73.4)”
Table 1Find in source - PERCENT73.5% does not match the reported count 460/630
“Annual family income <$25,000 — no. (%) | 460 (73.5)”
Table 1Find in source - PERCENT83.5% does not match the reported count 535/642
“No private health insurance — no. (%) | 535 (83.5)”
Table 1Find in source - PERCENT44% does not match the reported count 277/642
“Duration of hypertension >10 years — no. (%) | 277 (44.0)”
Table 1Find in source - PERCENT41.2% does not match the reported count 252/630
“Duration of hypertension >10 years — no. (%) | 252 (41.2)”
Table 1Find in source
- CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Primary outcome net difference p-value from CI
“with a net difference of −6.4 mmHg (95% CI, −9.0 to −3.8; P<0.001)”
Taken as given: The CI is a 95% confidence interval for the mean difference.; The estimate is the mean difference (-6.4).; The CI is symmetric on the linear scale.Method: Two-sided p-value derived from the estimate and 95% CI using normal approximation.How we recomputed it: pCI(-6.4, -9.0, -3.8, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Fidelity score net difference p-value from CI
“with a between-group difference of 0.7 (95% CI, 0.6 to 0.8; P<0.001)”
Taken as given: The CI is a 95% confidence interval for the mean difference.; The estimate is the mean difference (0.7).; The CI is symmetric on the linear scale.Method: Two-sided p-value derived from the estimate and 95% CI using normal approximation.How we recomputed it: pCI(0.7, 0.6, 0.8, 0)
- lowinternal contradictionIn the Results section, the 95% CI for the primary net difference is printed as '9.0 to 3.8' and for diastolic as '5.1 to 1.8', missing the minus signs, which contradicts the Abstract and Table 2.
“net difference of −6.4 mmHg (95% CI, 9.0 to 3.8; P< 0.001)”
ResultsFind in source - lowinternal contradictionIn the Results section, the confidence intervals for the net differences in systolic and diastolic blood pressure are missing negative signs, which contradicts the values in Table 2.
“net difference of −6.4 mmHg (95% CI, 9.0 to 3.8; P< 0.001)”
ResultsFind in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Conclusions only partially backed by the presented evidenceAssessed
4 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1The intervention is scalable to other primary care settings to improve hypertension control in underserved populations.The trial demonstrates effectiveness in FQHCs, but scalability is inferred from the design and partnership, not directly tested; the claim is reasonable but extends beyond the direct evidence.Evidence: Discussion states the strategy is scalable, but no direct scalability data are presented.
“This proven, multifaceted strategy is scalable to other primary care settings to improve hypertension control in underserved populations.”
Discussion ¶2Find in source - partialReviewer 2The trial is the first to show that a multifaceted, team-based strategy effectively lowers blood pressure among low-income patients receiving care at FQHCs.The claim of being 'first' is not directly evidenced by the paper's data; it is an assertion based on the authors' knowledge of prior literature. The effectiveness is supported, but the novelty claim is not verifiable from the presented evidence.Evidence: The paper cites prior studies but does not provide a systematic search to prove it is the first.
“Our trial is the first to show that a multifaceted, team-based strategy effectively lowers blood pressure among low-income patients receiving care at FQHCs.”
Discussion ¶3Find in source - supportedReviewers 1, 2The multifaceted, team-based implementation strategy significantly lowered systolic blood pressure among low-income patients with hypertension compared to enhanced usual care.The primary outcome shows a statistically significant net reduction of -6.4 mmHg (95% CI -9.0 to -3.8, P<0.001), directly supporting the claim.Evidence: Primary outcome: net difference in systolic BP change -6.4 mmHg (95% CI -9.0 to -3.8, P<0.001).
“Compared to enhanced usual care, the multifaceted, team-based implementation strategy significantly lowered systolic blood pressure among low-income patients with hypertension.”
ConclusionFind in source - supportedReviewers 1, 2The intervention improved fidelity to hypertension treatment compared to enhanced usual care.The primary implementation outcome (fidelity score) showed a significant between-group difference of 0.7 (95% CI 0.6 to 0.8, P<0.001), supporting the claim.Evidence: Primary implementation outcome: fidelity score difference 0.7 (95% CI 0.6 to 0.8, P<0.001).
“The mean fidelity score over the 18-month follow-up was 2.8 (95% CI, 2.7 to 2.9) in the intervention group and 2.1 (95% CI, 2.0 to 2.2) in the control group, with a between-group difference of 0.7 (95% CI, 0.6 to 0.8; P<0.001).”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary efficacy claim is based on reduction in systolic blood pressure, which is a surrogate endpoint for cardiovascular outcomes. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD) nor cite validated evidence linking the specific blood pressure reduction to clinical outcomes in this population. Although blood pressure is a well-established surrogate, the manuscript does not explicitly provide the required validation link.
“The primary effectiveness outcome was net between-group difference in mean systolic blood pressure from baseline to 18 months”
- ADEQUATEEffect sizeThe effect size is a net reduction of 6.4 mmHg in systolic blood pressure, which is statistically significant and clinically meaningful. The manuscript anchors this to prior trials and meta-analyses showing similar reductions are associated with reduced cardiovascular events, and the effect is consistent across subgroups.
“net difference of −6.4 mmHg (95% CI, −9.0 to −3.8; P<0.001)”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior research on hypertension prevalence, disparities, and the effectiveness of multicomponent strategies, including SPRINT. It identifies a gap in data on implementing such strategies in resource-constrained FQHC settings. The rationale linking the premise to the study objectives is explicit, and the study directly addresses the identified gap.
“The Systolic Blood Pressure Intervention Trial (SPRINT) demonstrated that intensive antihypertensive treatment targeting a systolic blood pressure <120 mm Hg significantly reduced cardiovascular disease events and all-cause mortality, compared to a target systolic blood pressure of <140 mm Hg.”
“However, data on implementing multifaceted blood pressure interventions in resource-constrained primary care settings, such as Federally Qualified Health Centers (FQHCs), which primarily serve low-income populations, are limited.”
“We tested the effectiveness and implementation of a multifaceted, team-based strategy to deliver an intensive blood pressure control protocol adapted from SPRINT among low-income patients receiving care at FQHC clinics.”
“The Systolic Blood Pressure Intervention Trial (SPRINT) demonstrated that intensive antihypertensive treatment targeting a systolic blood pressure <120 mm Hg significantly reduced cardiovascular disease events and all-cause mortality, compared to a target systolic blood pressure of <140 mm Hg.”
“There is a critical need to develop and test effective, scalable implementation strategies to translate evidence-based interventions into real-world primary care, particularly for underserved populations.”
“However, data on implementing multifaceted blood pressure interventions in resource-constrained primary care settings, such as Federally Qualified Health Centers (FQHCs), which primarily serve low-income populations, are limited.”
Randomization method (SAS-generated sequence, stratified by FQHC organization) and unit (clinic) are reported. Blinding is described as unblinded with a rationale. Power analysis is detailed with effect size, alpha, power, and assumptions. Inclusion/exclusion criteria are pre-specified. Outlier handling is addressed via multiple imputation and complete-case analysis. Controls (enhanced usual care) are appropriate. Independent replication is not applicable for a single pivotal trial.
“FQHC clinics were randomized in a 1:1 ratio to either a multifaceted team-based implementation strategy or enhanced usual care based on a random allocation sequence generated using SAS software, with stratification by FQHC organization.”
“This sample size provided 80% statistical power to detect a 5.0 mmHg difference in mean systolic blood pressure change over 18 months at a two-sided significance level of 0.05, with a follow-up rate of 85%.”
“Due to the nature of the cluster design and intervention program, study participants, healthcare providers, health coaches, and research staff who collected study data were unblinded.”
“FQHC clinics were randomized in a 1:1 ratio to either a multifaceted team-based implementation strategy or enhanced usual care based on a random allocation sequence generated using SAS software, with stratification by FQHC organization.”
“Due to the nature of the cluster design and intervention program, study participants, healthcare providers, health coaches, and research staff who collected study data were unblinded.”
“This sample size provided 80% statistical power to detect a 5.0 mmHg difference in mean systolic blood pressure change over 18 months at a two-sided significance level of 0.05, with a follow-up rate of 85%.”
Sex, age, race/ethnicity, income, education, employment, and health insurance are reported in Table 1. Health status (comorbidities, blood pressure) is also reported. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“Women — no. (%) | 377 (58.7) | 344 (54.6)”
“History of diabetes — no. (%) | 255 (39.7) | 248 (39.4)”
“Women — no. (%) | 377 (58.7) | 344 (54.6)”
“History of diabetes — no. (%) | 255 (39.7) | 248 (39.4)”
The study protocol was approved by the Tulane University Institutional Review Board. Informed consent was obtained from all participants. Regulatory compliance is implied through IRB approval and trial registration, though not explicitly named as a framework.
“Informed consent was obtained from all participants during screening visits.”
“Informed consent was obtained from all participants during screening visits.”
The intervention is described in detail (stepped-care protocol, health coaching, home BP monitoring). The blood pressure device is identified (HEM-907 XL, Omron Healthcare). Software used for analysis is identified (SAS V9.4, R V4.4.2). No antibodies, cell lines, or organisms are used.
“Three blood pressure measurements were taken at each visit using a standardized protocol with an automated device (HEM-907 XL, Omron Healthcare) and appropriately sized cuffs.”
“Statistical analyses were performed using SAS V9.4 (SAS Institute, Cary, NC) and R V4.4.2 (R Foundation) for calculating intraclass correlation coefficients.”
“Three blood pressure measurements were taken at each visit using a standardized protocol with an automated device (HEM-907 XL, Omron Healthcare) and appropriately sized cuffs.”
“Statistical analyses were performed using SAS V9.4 (SAS Institute, Cary, NC) and R V4.4.2 (R Foundation) for calculating intraclass correlation coefficients.”
Tests are named (linear mixed-effects, generalized linear mixed-effects). Assumptions are handled via model design (random effects, covariance structures). Exact p-values are reported for primary outcomes. Effect sizes with 95% CIs are reported throughout. Software is identified. Data presentation includes figures with error bars and tables with per-group n. Mathematical plausibility is not applicable for large-N continuous outcomes.
“We tested the difference in mean systolic blood pressure from baseline to 18 months between the intervention and control groups using measurements taken at 0, 6, 12, and 18 months in a linear mixed-effects regression analysis.”
“with a net difference of −6.4 mmHg (95% CI, −9.0 to −3.8; P<0.001)”
“Change in systolic BP from baseline to 18-month visit, mm Hg | −15.5 (−17.4, −13.6) | −9.1 (−11.0, −7.2) | −6.4 (−9.0, −3.8)”
“We tested the difference in mean systolic blood pressure from baseline to 18 months between the intervention and control groups using measurements taken at 0, 6, 12, and 18 months in a linear mixed-effects regression analysis.”
“with a net difference of −6.4 mmHg (95% CI, −9.0 to −3.8; P<0.001)”
The paper mentions supplementary material and a protocol available online, but does not provide a clear data availability statement with a repository or access mechanism. No code is shared. For a clinical trial, managed access is acceptable, but the statement is not explicit.
“Supplementary Material supplement”
Trial registration number is provided (NCT03483662). Methods are comprehensive. Limitations are discussed, including lack of blinding and extended enrollment. Conclusions are proportional to the evidence. Funding and COI are stated.
“This study has several limitations related to its cluster design. The extended enrollment period prevented recruitment of all participants before randomization.”
“Research reported in this publication was supported by the National Heart, Lung, and Blood Institute of the National Institutes of Health under Award Number R01HL133790.”
“This study has several limitations related to its cluster design.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 29 references by DOI: 26 verified — 3 no DOI (shown, not verified).
- NO DOINew medication adherence scale versus pharmacy fill rates in seniors with hypertensionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIApplied Longitudinal AnalysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Surgeon General’s Call to Action to Control HypertensionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT03483662LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
3 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 3 minor suggestions below.
3 copyedit issues flagged: mostly consistency, typo.
- MINORtypoAuthor affiliations“Cliton, LA”→ Clinton, LALikely misspelling of city name.
- MINORconsistencyResults, Effectiveness Outcomes“net difference of −6.4 mmHg (95% CI, 9.0 to 3.8; P< 0.001)”→ net difference of −6.4 mmHg (95% CI, −9.0 to −3.8; P<0.001)CI bounds missing negative signs; inconsistent with Table 2.
- MINORconsistencyResults, Effectiveness Outcomes“net difference of −3.4 mmHg (95% CI, 5.1 to 1.8)”→ net difference of −3.4 mmHg (95% CI, −5.1 to −1.8)CI bounds missing negative signs; inconsistent with Table 2.
The published work is robust and well-reported, with only minor reporting gaps (data availability statement, code sharing) and copyedit issues (missing negative signs in CIs). An informed reader should weigh the lack of a data availability statement and the unverified subset of statistics, but these do not undermine the main conclusions. A correction for the CI sign errors is warranted.
- 1.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 73.4% does not match the reported count 468/642Demonstrable critical failure — blocks the verdict from passing.
- 2.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 44% does not match the reported count 277/642Demonstrable critical failure — blocks the verdict from passing.
- 3.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 41.2% does not match the reported count 252/630Demonstrable critical failure — blocks the verdict from passing.
- 4.HIGHcopyeditIn the Results section, correct the 95% CI for the primary net difference to read '−9.0 to −3.8' (currently '9.0 to 3.8') to match the Abstract and Table 2.The missing negative signs are an internal contradiction that could mislead readers and undermine the reported effect direction.
- 5.HIGHcopyeditIn the Results section, correct the 95% CI for the diastolic net difference to read '−5.1 to −1.8' (currently '5.1 to 1.8') to match Table 2.The missing negative signs are an internal contradiction that could mislead readers and undermine the reported effect direction.
- 6.HIGHdata codeAdd an explicit data availability statement in the Data Availability section, specifying how de-identified data can be accessed (e.g., via a repository or managed access process).The current paper lacks a clear data availability statement, which is a reporting gap that reviewers and readers expect for a clinical trial.
- 7.HIGHdata codeProvide a statement on whether analysis code is available and where, or explicitly state that it is not available.Code sharing enhances reproducibility, and its absence is a common reviewer concern.
- 8.MEDIUMreportingMention adherence to the CONSORT extension for cluster randomized trials in the Methods or supplement.Explicitly referencing the reporting guideline would strengthen the transparency of the trial reporting.
- 9.MEDIUMdata codeConsider depositing the statistical analysis code (e.g., SAS macros) in a public repository to enhance reproducibility.Depositing code would address the code-sharing gap and improve the paper's reproducibility.
- 10.LOWcopyeditCorrect the typo in the author affiliations: change 'Cliton, LA' to 'Clinton, LA'.A misspelled city name is a minor but visible error that should be fixed.
- 11.LOWdata codeClarify the data availability for the supplementary materials, including a persistent identifier if deposited.Providing a persistent identifier for supplementary materials would improve accessibility and transparency.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.