Heterogeneous effects of Medicaid coverage on cardiovascular risk factors: secondary analysis of randomized controlled trial.
Inoue K, Athey S, Baicker K, Tsugawa Y
- DOI
- 10.1136/bmj-2024-079377
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/c822e534-dc11-4902-ac91-e0ac6d7626a5 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- LinksDead data/code link−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 33 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy claim is that Medicaid coverage lowers systolic blood pressure in a subgroup. Blood pressure is a surrogate biomarker for cardiovascular outcomes, not a hard clinical outcome. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure) nor cite validated evidence linking the observed blood pressure reduction to a clinical outcome. The effect is presented as a biomarker change without establishing a validated surrogate-to-clinical-outcome link.
“Medicaid coverage significantly lowered systolic blood pressure (−2.93 mmHg (95% confidence interval −5.82 to −0.32)) for people predicted to benefit highly.”
- 02Treatment effect not shown to be clinically meaningful
The primary reported effect is a reduction of 2.93 mmHg in systolic blood pressure among a subgroup. This is a small fraction of the normal/reference value and is below the established minimal clinically important difference for blood pressure. The authors themselves acknowledge the effect size may be of limited clinical significance for any individual, and they do not anchor the effect to a clinically meaningful threshold.
“Although the effect size may be of limited clinical significance for any individual, at a broad population level that includes individuals who are both hypertensive and normotensive, the findings may be of public health importance for policy interventions.”
- 03Declared data/code link does not resolve
Dead link — nothing to verify.
“https://www.nber.org/research/data/oregon-health-insurance-experiment”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper is a methodologically sound secondary analysis of a well-known randomized trial, using causal forests to estimate heterogeneous treatment effects of Medicaid on cardiovascular risk factors. It reports clear methods, adequate biological variables, ethical approvals, and transparent reporting. The main weaknesses are minor reporting gaps (no power analysis, no explicit outlier handling, no code sharing) and a small internal inconsistency in the reported average treatment effect.
Both reviewers classified the study as observational (secondary analysis of an RCT), and this was adopted. The evaluation covered all eight dimensions; key resources was not applicable. The statistics verification covered only 2 tests (1 consistent, 0 inconsistent) due to limited reporting of test statistics; most results remain unverified. The reproducibility check found the data link dead, which is a concern for data availability.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks.
- CONSISTENTreported p = .006 · recomputed p = .011Reviewers 1, 2Check p-value for difference in systolic blood pressure between high benefit group and overall population.
“adjusted difference −2.30 mmHg (95% CI −4.22 to −0.66), P=0.006”
Taken as given: The reported difference is the estimate.; The 95% CI is two-sided.; The p-value is two-tailed.Method: Recomputed p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-2.30, -4.22, -0.66, 0) - UNCOMPUTABLEreported p = .008 · recomputed p = .019Reviewers 1, 2Check p-value for difference in diastolic blood pressure between high benefit group and overall population.
“−1.46 mmHg (−2.75 to −0.31), P=0.008”
Taken as given: The reported difference is the estimate.; The 95% CI is two-sided.; The p-value is two-tailed.Method: Recomputed p-value from estimate and 95% CI using normal approximation.How we recomputed it: pCI(-1.46, -2.75, -0.31, 0)
- lowinternal contradictionThe average treatment effect for systolic blood pressure is reported as −0.62 mmHg in the Results and Table 3, but the Discussion states −0.52 mmHg.
“which was six times larger than the average treatment effect (−0.52 mmHg) observed in the original Oregon health insurance experiment.”
DiscussionFind in source
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
4 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Medicaid coverage significantly lowered systolic blood pressure for people predicted to benefit highly.The paper reports a statistically significant reduction in systolic blood pressure in the high benefit group with a 95% CI excluding zero.Evidence: Table 3: Local average treatment effect for high benefit group: −2.93 mmHg (95% CI −5.82 to −0.32).
“Medicaid coverage significantly lowered systolic blood pressure (−2.93 mmHg (95% confidence interval −5.82 to −0.32)) for people predicted to benefit highly.”
Results ¶3Find in source - supportedReviewers 1, 2No evidence showed that Medicaid coverage lowered HbA1c for people with high predicted benefits.The paper reports a null effect with a CI including zero.Evidence: Table 3: Local average treatment effect for HbA1c in high benefit group: 0.01% (95% CI −0.09 to 0.13).
We found no evidence that the Medicaid coverage lowered HbA 1c among individuals with high predicted benefit for HbA 1c compared with the overall population (0.01% v 0.00%, adjusted difference 0.01% (95% CI −0.08% to 0.10%), P=0.83).
Results ¶4reviewer’s wording - supportedReviewers 1, 2Individuals with high predicted benefits were more likely to have no or low prior healthcare charges.The paper shows lower baseline charges in the high benefit group compared to the overall population.Evidence: Table 2: Sum of total charges in high benefit group for SBP: $488 vs overall $2114.
Individuals predicted to benefit highly (ie, conditional local average treatment effect (<0)) for systolic blood pressure (n=8593) or HbA 1c (n=6937) were less likely to have a history of hypertension diagnosis and had lower total and emergency department charges at baseline than those with lower predicted benefit.
Results ¶2reviewer’s wording - supportedReviewers 1, 2The findings suggest that Medicaid coverage leads to improved blood pressure for some people, but those benefits may be diluted by individuals who did not benefit.The paper's results support this interpretation, as the average effect is null but a subgroup shows benefit.Evidence: Overall average effect is −0.62 (95% CI −3.16 to 1.73), while high benefit group shows −2.93 (95% CI −5.82 to −0.32).
“Our findings suggest that Medicaid coverage leads to improved blood pressure for some people, but those benefits may be diluted by individuals who did not experience benefits as well as by the inclusion of a population of individuals who were hypertensive and normotensive.”
ConclusionFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy claim is that Medicaid coverage lowers systolic blood pressure in a subgroup. Blood pressure is a surrogate biomarker for cardiovascular outcomes, not a hard clinical outcome. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure) nor cite validated evidence linking the observed blood pressure reduction to a clinical outcome. The effect is presented as a biomarker change without establishing a validated surrogate-to-clinical-outcome link.
“Medicaid coverage significantly lowered systolic blood pressure (−2.93 mmHg (95% confidence interval −5.82 to −0.32)) for people predicted to benefit highly.”
- INADEQUATEEffect sizeThe primary reported effect is a reduction of 2.93 mmHg in systolic blood pressure among a subgroup. This is a small fraction of the normal/reference value and is below the established minimal clinically important difference for blood pressure. The authors themselves acknowledge the effect size may be of limited clinical significance for any individual, and they do not anchor the effect to a clinically meaningful threshold.
“Although the effect size may be of limited clinical significance for any individual, at a broad population level that includes individuals who are both hypertensive and normotensive, the findings may be of public health importance for policy interventions.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior RCTs (RAND, IRS outreach, Oregon) and notes their limitations, including lack of individual-level health outcomes and average effect estimates that may obscure subgroup benefits. The rationale for using machine learning to detect heterogeneity is well articulated, linking the premise to the study objectives. Limitations of prior work are addressed by applying novel methods to a randomized design.
“Several randomized controlled trials have investigated insurance coverage in the United States of America. The RAND health insurance experiment was conducted in the 1970s-80s with the primary aim of studying the price elasticity of demand for healthcare services and implications for health outcomes, but the effect of having health insurance itself was not studied.”
“Recent rapid advancements in machine learning techniques have enabled nuanced estimation of how treatment effects vary based on individuals’ observable characteristics, so-called heterogeneous treatment effects.”
“The randomized controlled trial design used in the Oregon health insurance experiment eliminated such biases. However, some subgroups in the Oregon health insurance experiment might have had an improvement in cardiovascular risk factors, while the average treatment effect was diluted by other subgroups who did not benefit from Medicaid coverage.”
“The results showed improvements in access to care and outcomes, including depression, but showed, on average, no evidence that Medicaid coverage improved physical health, including cardiovascular risk factors such as blood pressure and hemoglobin A 1c (HbA 1c ) concentrations.”
“Recent rapid advancements in machine learning techniques have enabled nuanced estimation of how treatment effects vary based on individuals’ observable characteristics, so-called heterogeneous treatment effects.”
“Some studies using observational or quasi-experimental designs have found that Medicaid coverage is associated with an improved health status, including lower risk of mortality, but such studies are subject to confounding factors and omitted variable bias.”
The study uses data from a randomized controlled trial with a lottery-based assignment, which serves as the randomization method. The unit of randomization is the individual. Blinding is not applicable as this is a secondary analysis of an open-label trial. Power analysis is not reported, but the study is a secondary analysis of a large trial. Inclusion/exclusion criteria are described (low-income, uninsured adults). Outlier handling is not explicitly discussed, but missing data imputation is mentioned. Controls are inherent in the randomized design. Independent replication is not applicable as this is a secondary analysis.
“This study leveraged the random assignment of access to Medicaid insurance coverage in 2008 for low income adults (defined as less than the federal poverty line) in Oregon who were uninsured.”
“Across a total of 12 229 participants who responded to the survey (effective response rate, 73%), this study included 12 134 individuals with whose outcome data were available.”
“Missing data for these covariates at baseline were imputed using a random forest approach.”
“This study leveraged the random assignment of access to Medicaid insurance coverage in 2008 for low income adults (defined as less than the federal poverty line) in Oregon who were uninsured.”
“Across a total of 12 229 participants who responded to the survey (effective response rate, 73%), this study included 12 134 individuals with whose outcome data were available.”
“Missing data for these covariates at baseline were imputed using a random forest approach.”
Sex is reported in Table 1 (female/male percentages). Age is reported as mean (SD). Health status is captured through baseline diagnoses (e.g., hypertension, diabetes). Demographics include race/ethnicity and education. Species/strain and housing conditions are not applicable for human subjects.
“Female | 3296 (56.9) | 3564 (56.2)”
“Age, mean (SD), years | 40.56 (11.68) | 40.96 (11.71)”
“Hypertension | 1054 (18.2) | 1146 (18.1)”
“Female | 3296 (56.9) | 3564 (56.2)”
“Age, mean (SD), years | 40.56 (11.68) | 40.96 (11.71)”
“Hypertension | 1054 (18.2) | 1146 (18.1)”
The study reports IRB approval from UCLA (protocol number 24-000623) and notes that the Oregon Health Insurance Experiment received approvals from several IRBs and all participants provided written consent. Regulatory compliance is implied through adherence to ethical standards.
“The protocol for this study was approved by the institutional review board at University of California, Los Angeles, USA (institutional review board number 24-000623).”
“The Oregon health insurance experiment has received approvals from several institutional review boards, and all participants provided written consent during the in-person survey.”
“The protocol for this study was approved by the institutional review board at University of California, Los Angeles, USA (institutional review board number 24-000623).”
“The Oregon health insurance experiment has received approvals from several institutional review boards, and all participants provided written consent during the in-person survey.”
The study does not use antibodies, cell lines, organisms, or reagents. The only software used is R, which is identified. Since no investigational product or bench reagents are involved, the dimension is not applicable.
“All statistical analyses were conducted using R, version 4.1.1 (R Project for Statistical Computing).”
“All statistical analyses were conducted using R, version 4.1.1 (R Project for Statistical Computing).”
The paper names the causal forest algorithm and instrumental variable regression. Effect sizes are reported with 95% CIs. P-values are reported for some comparisons. Software (R version 4.1.1) is identified. Data presentation includes tables with per-group n and means/SDs. Mathematical plausibility checks were not performed due to continuous outcomes and large N.
“We built the causal forest algorithm with an instrumental variable regression (ie, instrumental variable forests; instrumental_forest function in grf package in R)”
“Local average treatment effect (95% CI) | −0.62 (−3.16 to 1.73)”
“We built the causal forest algorithm with an instrumental variable regression (ie, instrumental variable forests; instrumental_forest function in grf package in R)”
“Medicaid coverage significantly lowered systolic blood pressure (−2.93 mmHg (95% confidence interval −5.82 to −0.32)) for people predicted to benefit highly.”
“adjusted difference −2.30 mmHg (95% CI −4.22 to −0.66), P=0.006”
The data availability statement provides a concrete route: all data are available online from the National Bureau of Economic Research's Public Use Data Archive with a URL. No code is shared, but the analysis is based on standard R packages; code sharing is not explicitly required for this type of study.
“All data used in this study are available online from the National Bureau of Economic Research’s Public Use Data Archive and can be accessed at https://www.nber.org/research/data/oregon-health-insurance-experiment”
“All data used in this study are available online from the National Bureau of Economic Research’s Public Use Data Archive and can be accessed at https://www.nber.org/research/data/oregon-health-insurance-experiment”
Methods are detailed enough for replication. The trial is registered (AEARCTR-0000028). No specific reporting guideline is mentioned, but the paper follows standard reporting. All outcomes are reported, including null results. Limitations are extensively discussed. Conclusions are proportional to evidence. Funding and COI are disclosed.
“The Oregon health insurance experiment was registered at the American Economic Association’s registry for randomized controlled trials (registration number AEARCTR-0000028).”
“Our study has limitations. Firstly, the causal forest model evaluated heterogeneity based on measured covariates, and other unmeasured characteristics may also be important.”
“This study was supported by the Japan Society for the Promotion of Science (22K17392 and 23KK0240; PI, Inoue), the Japan Science and Technology Agency (JST, JPMJPR23R2; PI, Inoue), National Institutes of Health (NIH) (P01AG005842 and R01AG034151; PI, Baicker), and Gregory Annenberg Weingarten, GRoW @ Annenberg (PI, Tsugawa).”
“The Oregon health insurance experiment was registered at the American Economic Association’s registry for randomized controlled trials (registration number AEARCTR-0000028).”
“Our study has limitations. Firstly, the causal forest model evaluated heterogeneity based on measured covariates, and other unmeasured characteristics may also be important.”
“This study was supported by the Japan Society for the Promotion of Science (22K17392 and 23KK0240; PI, Inoue), the Japan Science and Technology Agency (JST, JPMJPR23R2; PI, Inoue), National Institutes of Health (NIH) (P01AG005842 and R01AG034151; PI, Baicker), and Gregory Annenberg Weingarten, GRoW @ Annenberg (PI, Tsugawa).”
Registration stated in text, but no registry ID was detected. No reporting guideline cited.
Broken references and links
1 finding · worst mediumReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- Dead data/code linksRecomputed
Checked 25 references by DOI: 23 verified — 2 no DOI (shown, not verified).
- NO DOIFree for All?: Lessons from the Rand Health Insurance ExperimentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIGeneric machine learning inference on heterogenous treatment effects in randomized experimentsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 0 live, 1 dead.
- datahttps://www.nber.org/research/data/oregon-health-insurance-experimentDEADHTTP 404Dead link — nothing to verify.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly grammar, consistency, typo.
- MINORtypoAbstract, Results“hemoglobin A 1c (HbA 1c ) concentrations”→ Remove extra space before closing parenthesis: 'hemoglobin A 1c (HbA 1c) concentrations'Inconsistent spacing around parentheses.
- MINORconsistencyDiscussion, Policy implications“−0.52 mmHg”→ Verify the average treatment effect value; elsewhere it is reported as −0.62 mmHg.Potential inconsistency in reported average treatment effect.
- MINORgrammarMethods, Statistical analyses“we further divided each subsample into two parts.”→ Consider rephrasing for clarity: 'we further divided each subsample into two parts' is acceptable.Minor style issue.
- MINORgrammarAbstract, Results“In the in-person interview survey, mean systolic blood pressure was 119 (standard deviation 17) mmHg and mean HbA 1c concentrations was 5.3% (standard deviation 0.6%).”→ Change 'concentrations was' to 'concentration was'.Subject-verb agreement error.
- MINORgrammarMethods, Study sample“this study included 12 134 individuals with whose outcome data were available.”→ Change 'with whose' to 'whose'.Unnecessary preposition.
- MINORconsistencyDiscussion, Policy implications“which was six times larger than the average treatment effect (−0.52 mmHg) observed in the original Oregon health insurance experiment.”→ Verify the average treatment effect value; the abstract reports −0.62 mmHg.Potential inconsistency in reported average effect.
The published work is generally robust, but an informed reader should weigh the minor reporting gaps and the internal inconsistency in the reported average treatment effect. The dead data link is a concrete issue that warrants correction or clarification. No major validity threats were identified, but the unverified statistics and lack of code sharing limit full reproducibility.
- 1.HIGHreportingReconcile the reported average treatment effect for systolic blood pressure: the Results/Table 3 report −0.62 mmHg while the Discussion reports −0.52 mmHg; correct the inconsistent value.An internal contradiction in a headline effect size undermines reader trust and could warrant an erratum.
- 2.HIGHdata codeFix or replace the dead data availability link (https://www.nber.org/research/data/oregon-health-insurance-experiment) to ensure the data are actually accessible.A data availability statement undercut by a broken link is a reporting gap that reviewers and readers will catch.
- 3.MEDIUMreportingAdd a power analysis or sample size justification in the Methods section, even if it is a secondary analysis.The absence of any power analysis leaves the reader unable to assess the study's ability to detect heterogeneity.
- 4.MEDIUMreportingExplicitly state how outliers were handled in the statistical analysis, or note that no outliers were excluded.Outlier handling is a standard reporting element that is currently missing.
- 5.MEDIUMdata codeShare the analysis code in a public repository (e.g., GitHub) to enhance reproducibility.Code sharing is not reported, and providing it would allow independent verification of the causal forest analysis.
- 6.MEDIUMreportingMention adherence to a reporting guideline (e.g., STROBE) in the Methods or as a checklist.Reporting guidelines improve transparency and are expected for observational studies.
- 7.MEDIUMstatisticsProvide more detail on the verification of model assumptions (e.g., instrumental variable validity, causal forest assumptions) in the statistical analysis section.Assumptions are currently rated as 'reported_but_inadequate'; explicit verification would strengthen the analysis.
- 8.LOWcopyeditFix the subject-verb agreement error in the Abstract: change 'concentrations was' to 'concentration was'.Grammar error in a key summary section.
- 9.LOWcopyeditFix the unnecessary preposition in Methods, Study sample: change 'with whose' to 'whose'.Grammar error that affects readability.
- 10.LOWcopyeditRemove extra space before closing parenthesis in 'HbA 1c )' throughout the paper.Inconsistent spacing around parentheses is a minor typographical issue.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.