Pragmatic Trial of Hospitalization Rate in Chronic Kidney Disease.
Vazquez MA, Oliver G, Amarasingham R, Sundaram V, Chan K, Ahn C, Zhang S, Bickel P, Parikh SM, Wells B, Miller RT, Hedayati S, Hastings J, Jaiyeola A, Nguyen TM, Moran B, Santini N, Barker B, Velasco F, Myers L, Meehan TP, Fox C, Toto RD, ICD-Pieces Study Group
- DOI
- 10.1056/NEJMoa2311708
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/7a2ce24f-34eb-4be3-b351-d1a52aa6f219 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 20 reported means were read, and their group size is not stated where the values are printed. These checks need the count the mean was averaged over, so none was performed.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted pragmatic cluster-randomized trial with rigorous design, clear reporting of demographics, ethical approvals, and statistical methods. The main weaknesses are the vague data availability statement, lack of software version identification, and minor copyedit issues such as a missing numeric difference in the primary outcome text.
Both reviewers classified the study as interventional, and this was adopted. The evaluation covered all eight dimensions; several sub-criteria were marked not applicable (e.g., animal-related items, blinding in an open-label trial). The statistics verification checked only 2 tests, so the overall statistical soundness is not fully confirmed beyond those checks.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks.
- CONSISTENTreported p = .580 · recomputed p = .568Reviewer 1Primary outcome hospitalization rate comparison using reported counts and totals.
“Hospitalization for any cause was reported in 1139 of 5508 patients (20.7%; 95% confidence interval [CI], 19.7 to 21.8) in the intervention group and in 1160 of 5492 patients (21.1%; 95% CI, 20.1 to 22.2) in the usual-care group, for a difference of percentage points (95% CI, −2.0 to 1.1; P = 0.58)”
Taken as given: The counts 1139 and 1160 are the number of hospitalizations in each group.; The totals 5508 and 5492 are the group sizes.; The p-value is from a two-sided test comparing two proportions.Method: Pearson chi-square test for 2x2 table using cell counts.How we recomputed it: pChi2x2(1139, 5508-1139, 1160, 5492-1160) - CONSISTENTreported p = .580 · recomputed p = .568Reviewer 2Primary outcome: hospitalization at 1 year, intervention vs usual care, using a generalized linear mixed model. The reported p-value is 0.58. We can approximate using a chi-square test on the 2x2 table of events/total.
“Hospitalization for any cause was reported in 1139 of 5508 patients (20.7%; 95% confidence interval [CI], 19.7 to 21.8) in the intervention group and in 1160 of 5492 patients (21.1%; 95% CI, 20.1 to 22.2) in the usual-care group, for a difference of percentage points (95% CI, −2.0 to 1.1; P = 0.58).”
Taken as given: The 1139 and 1160 are the event counts in the intervention and usual-care groups respectively.; The 5508 and 5492 are the total patients in each group.; A chi-square test approximates the mixed model result (the paper used a mixed model, but the p-value is similar).Method: Pearson chi-square test on 2x2 table of events vs non-events.How we recomputed it: pChi2x2(1139, 5508-1139, 1160, 5492-1160)
- lowinternal contradictionThe text in the Results section for the primary outcome omits the numeric difference (0.4 percentage points) that appears in the abstract, which could be a typographical omission.
“for a difference of percentage points (95% CI, −2.0 to 1.1; P = 0.58)”
ResultsFind in source - lowinternal contradictionThe baseline table shows body-mass index for usual care as '33±7.4' without a decimal, while all other values have one decimal, suggesting a formatting inconsistency.
“Body-mass index | 33.4±7.6 | 33±7.4”
Table 1Find in source
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
9 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewer 1The intervention did not reduce hospitalization at 1 year.The primary outcome analysis shows no significant difference between groups, supporting the claim.Evidence: Primary outcome: 20.7% vs 21.1%, difference 0.4 percentage points, P=0.58.
“the use of an EHR-based algorithm and practice facilitators embedded in primary care clinics did not translate into reduced hospitalization at 1 year.”
ConclusionFind in source - supportedReviewer 1Secondary outcomes were similar between groups.All secondary outcomes show overlapping confidence intervals and no significant differences, supporting the claim.Evidence: Secondary outcomes table shows similar rates with CIs overlapping zero.
Secondary outcomes were similar in the two groups.
Resultsreviewer’s wording - supportedReviewer 1Acute kidney injury was more common in the intervention group.The adverse events table shows a higher rate in the intervention group (12.7% vs 11.3%), supporting the claim.Evidence: Table 4: Acute kidney injury 701 (12.7) vs 619 (11.3).
“Acute kidney injury was the most common adverse event and was reported in 701 of 5508 patients (12.7%) in the intervention group and in 619 of 5492 patients (11.3%) in the usual-care group.”
ResultsFind in source - supportedReviewer 1The trial enrolled a representative population with high inclusion of underrepresented groups.The baseline characteristics show a diverse population with about 20% Black and 20% Hispanic patients, supporting the claim.Evidence: Table 1 shows race and ethnicity distributions.
“Our trial design enabled the participation of a high fraction of underrepresented groups, thus making our findings generalizable.”
Discussion ¶3Find in source - supportedReviewer 1The intervention improved process-of-care measures.The implementation table shows higher rates of updated problem lists, patient education, and new ACE inhibitor/ARB prescriptions in the intervention group, supporting the claim.Evidence: Table 3 shows higher percentages in intervention for several process measures.
“More patients in the intervention group had an updated problem list, had adopted targets for blood pressure and glycated hemoglobin, and had documentation of the receipt of education regarding the kidney-dysfunction triad.”
ResultsFind in source - supportedReviewer 2The intervention did not reduce hospitalization at 1 year compared to usual care.The primary outcome analysis directly supports this claim with a non-significant p-value and overlapping confidence intervals.Evidence: Primary outcome: 20.7% vs 21.1%, difference -0.4 percentage points (95% CI -2.0 to 1.1, P=0.58).
“The hospitalization rate at 1 year was 20.7% (95% confidence interval [CI], 19.7 to 21.8) in the intervention group and 21.1% (95% CI, 20.1 to 22.2) in the usual-care group (between-group difference, 0.4 percentage points; P = 0.58).”
AbstractFind in source - supportedReviewer 2Secondary outcomes (ED visits, readmissions, cardiovascular events, dialysis, death) were similar between groups.All secondary outcomes show overlapping confidence intervals and small absolute differences, supporting similarity.Evidence: Table 2 and Results section report point estimates and 95% CIs for each secondary outcome, all with overlapping intervals.
Secondary outcomes were similar in the two groups.
Resultsreviewer’s wording - supportedReviewer 2The intervention improved process-of-care measures (e.g., problem list updates, goal setting, patient education).Table 3 shows higher percentages in the intervention group for several process measures, based on a 10% random sample chart review.Evidence: Table 3: e.g., 'Met all criteria: problem list updated and patient education received' 64% vs 39%.
“More patients in the intervention group had an updated problem list, had adopted targets for blood pressure and glycated hemoglobin, and had documentation of the receipt of education regarding the kidney-dysfunction triad.”
ResultsFind in source - supportedReviewer 2Adverse events were similar between groups except for acute kidney injury, which was higher in the intervention group.Table 4 shows similar rates for most adverse events, with acute kidney injury at 12.7% vs 11.3%.Evidence: Table 4: Acute kidney injury 701/5508 (12.7%) vs 619/5492 (11.3%).
“Acute kidney injury was the most common adverse event and was reported in 701 of 5508 patients (12.7%) in the intervention group and in 619 of 5492 patients (11.3%) in the usual-care group.”
ResultsFind in source
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior studies on the burden of the triad and the potential of practice facilitators, acknowledges that large-scale implementation trials are lacking, and states the hypothesis. Limitations of prior research are implicitly addressed by the trial's pragmatic design, though not explicitly enumerated.
“We conducted the Improving Chronic Disease Management with Pieces (ICD-Pieces) trial to test the hypothesis that an intervention combining information technology with practice facilitators could reduce the hospitalization rate among patients with the kidney-dysfunction triad.”
“Despite a growing array of guideline-directed therapies for patients with the kidney-dysfunction triad, the results of large trials to examine the implementation of guideline-directed therapy to reduce morbidity and mortality in this population are lacking.”
“Chronic kidney disease, type 2 diabetes, and hypertension (the kidney-dysfunction triad) represent chronic conditions that often coexist and place patients at high risk for major cardiovascular events and kidney failure.”
“We conducted the Improving Chronic Disease Management with Pieces (ICD-Pieces) trial to test the hypothesis that an intervention combining information technology with practice facilitators could reduce the hospitalization rate among patients with the kidney-dysfunction triad.”
“Despite a growing array of guideline-directed therapies for patients with the kidney-dysfunction triad, the results of large trials to examine the implementation of guideline-directed therapy to reduce morbidity and mortality in this population are lacking.”
Randomization method (permuted-block with variable block sizes) and unit (practice) are reported. A power analysis with assumptions is provided. Inclusion/exclusion criteria are clearly defined. Outlier handling is not explicitly discussed but the primary analysis uses a mixed model that accounts for clustering. Controls are the usual-care group. Independent replication is not applicable for a single pragmatic trial.
“Practices were stratified according to health system and were randomly assigned either to the trial intervention or to usual care. Randomization was conducted with the use of a permuted-block approach with variable block sizes.”
“To detect this difference in hospitalization, we determined that the enrollment of 10,991 patients would be needed to provide the trial with 80% power at a two-sided significance level of 0.05.”
“Practices were stratified according to health system and were randomly assigned either to the trial intervention or to usual care. Randomization was conducted with the use of a permuted-block approach with variable block sizes.”
“Exclusion criteria were pregnancy, incarceration, acute kidney injury, rapidly progressive glomerulonephritis, chronic kidney disease stage 5, end-stage kidney disease, end-stage organ damage, or life expectancy of less than 2 years.”
Table 1 provides age, sex, race, ethnicity, blood pressure, glycated hemoglobin, eGFR, proteinuria, BMI, weight, cholesterol, and comorbidities. Sex is reported as male sex percentage. Since both sexes are enrolled, no single-sex justification is needed. Species/strain and housing conditions are not applicable for a human trial.
“Male sex — no. (%) | 2958 (53.7) | 2951 (53.7)”
“Age — yr | 68.1±10.4 | 68.9±10.3”
“Male sex — no. (%) | 2958 (53.7) | 2951 (53.7)”
“Age — yr | 68.1±10.4 | 68.9±10.3”
The paper states that the institutional review board at each health system approved the trial, including a waiver of informed consent. This satisfies irb_ethics_statement and informed_consent (waiver justified). Regulatory compliance is not explicitly named but the trial is NIH-funded and follows standard pragmatic trial regulations; however, a specific framework (e.g., Declaration of Helsinki) is not cited, making this reported_but_inadequate.
“The institutional review board at each health system approved the conduct of the trial, which included a waiver of informed consent.”
“The institutional review board at each health system approved the conduct of the trial, which included a waiver of informed consent.”
“The institutional review board at each health system approved the conduct of the trial, which included a waiver of informed consent.”
“The institutional review board at each health system approved the conduct of the trial, which included a waiver of informed consent.”
The intervention is described as an EHR-based algorithm and practice facilitators. The EHR systems are named (Parkland Health, Texas Health Resources, VANTX, ProHealth Physicians). No specific software versions or RRIDs are provided. Reagents, antibodies, cell lines, and organisms are not applicable. The trial uses no traditional biological/chemical resources, but the intervention itself is a scored resource; it is adequately described.
“The intervention consisted of two components: first, an algorithm identified patients in the electronic health record in real time, and then practice facilitators assisted primary care providers in implementing evidence-based interventions.”
“The intervention consisted of two components: first, an algorithm identified patients in the electronic health record in real time, and then practice facilitators assisted primary care providers in implementing evidence-based interventions.”
The primary test is named (generalized linear mixed model). Exact p-value is reported for the primary outcome (P=0.58). Effect sizes with 95% CIs are reported for all outcomes. Software is not identified. Data presentation includes per-group Ns and 95% CIs. Mathematical plausibility checks: the primary outcome counts (1139/5508 = 20.68%, reported as 20.7%; 1160/5492 = 21.12%, reported as 21.1%) are consistent. The total enrolled (5690 + 5492 = 11182) matches the stated total. The opt-out (182) reduces intervention to 5508, which is consistent.
“Hospitalization for any cause was reported in 1139 of 5508 patients (20.7%; 95% confidence interval [CI], 19.7 to 21.8) in the intervention group and in 1160 of 5492 patients (21.1%; 95% CI, 20.1 to 22.2) in the usual-care group, for a difference of percentage points (95% CI, −2.0 to 1.1; P = 0.58)”
“The primary outcome of hospitalization at 1 year was compared between the two groups by means of a generalized linear mixed-model approach”
“The primary outcome of hospitalization at 1 year was compared between the two groups by means of a generalized linear mixed-model approach, which includes random effects to account for within-practice correlation.”
“Emergency department visits occurred in 24.3% (95% CI, 23.2 to 25.4) of the patients in the intervention group and in 22.6% (95% CI, 21.5 to 23.7) of those in the usual-care group, for a difference of 1.7% percentage points (95% CI, −0.2 to 3.3).”
The paper states 'A data sharing statement provided by the authors is available with the full text of this article at NEJM.org.' This is vague and does not specify a repository or managed-access platform. No code sharing is mentioned. For a clinical trial, managed access is acceptable but the statement here is insufficiently concrete.
“A data sharing statement provided by the authors is available with the full text of this article at NEJM.org”
Trial registration number is provided. Methods are detailed enough for replication. Limitations are discussed in the Discussion. Conclusions are proportional to the null result. Funding sources and conflict of interest disclosures are provided. No reporting guideline (e.g., CONSORT) is referenced.
“Our trial also has some limitations. The relatively high rate of uptake of process-of-care interventions at baseline in the usual-care group may have limited the effect of the intervention on the primary outcome.”
“Our trial also has some limitations. The relatively high rate of uptake of process-of-care interventions at baseline in the usual-care group may have limited the effect of the intervention on the primary outcome.”
“Supported by the NIH Pragmatic Trials Collaboratory by cooperative agreement (UH3DK104655) with the National Institute of Diabetes and Digestive and Kidney Diseases.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 40 references by DOI: 36 verified — 4 no DOI (shown, not verified).
- NO DOI2022 USRDS annual data report: epidemiology of kidney disease in the United StatesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFequently asked questions — 2 midnight inpatient admission guidance and patient status reviews for admissions on or after October 1, 2013No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISample size determination for clustered outcomesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDigital health’s role in managing chronic diseaseNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 of 2 data/code links checked; 1 live; 1 not probed.
- datahttps://NEJM.orgUNVERIFIEDHTTP 403Liveness indeterminate — content not checked.
- datahttps://clinicaltrials.gov/ct2/show/NCT02587936LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, typo.
- MINORtypoMethods, Trial Procedures“nonsteriodal antiinflammatory drugs”→ nonsteroidal anti-inflammatory drugsMisspelling of 'nonsteroidal'.
- MINORconsistencyResults, Primary Outcome“for a difference of percentage points”→ for a difference of 0.4 percentage pointsThe difference value is missing in the text.
- MINORconsistencyTable 1“Body-mass index | 33.4±7.6 | 33±7.4”→ 33.0±7.4Inconsistent decimal formatting for the usual-care group.
- MINORconsistencyResults, Primary Outcome“for a difference of percentage points (95% CI, −2.0 to 1.1; P = 0.58)”→ for a difference of -0.4 percentage points (95% CI, −2.0 to 1.1; P = 0.58)The difference value (-0.4) is missing from the sentence.
The published work is robust and generally well-reported, but an informed reader should weigh the vague data availability statement and the minor copyedit inconsistencies. No erratum is warranted for the core findings, but the authors should consider issuing a correction for the missing numeric difference in the primary outcome text and the BMI formatting inconsistency.
- 1.HIGHcopyeditIn the Results, Primary Outcome section, insert the missing numeric difference '0.4 percentage points' into the sentence 'for a difference of percentage points (95% CI, −2.0 to 1.1; P = 0.58)'.The omission of the effect size is a factual inconsistency that could confuse readers and should be corrected.
- 2.HIGHcopyeditIn Table 1, standardize the body-mass index for the usual-care group from '33±7.4' to '33.0±7.4' to match the decimal formatting of other values.Inconsistent decimal formatting is a minor but visible reporting error that should be fixed for clarity.
- 3.HIGHdata codeReplace the vague data sharing statement with a concrete access route, such as a managed-access repository (e.g., Vivli, YODA) or a DOI, in the Data Availability section.The current statement referencing NEJM.org is insufficiently specific for readers to access the data, undermining reproducibility.
- 4.MEDIUMstatisticsIdentify the statistical software and version used for analysis (e.g., SAS 9.4, R 4.0) in the Statistical Analysis section.Providing software version details enhances transparency and reproducibility of the statistical methods.
- 5.MEDIUMreportingReference the CONSORT extension for cluster-randomized trials in the Methods section and indicate that the checklist was completed.Explicitly citing the reporting guideline demonstrates adherence to best practices and improves completeness.
- 6.MEDIUMethicsAdd an explicit statement of compliance with a regulatory framework such as the Declaration of Helsinki or ICH-GCP in the Ethics section.While IRB approval is stated, naming the framework clarifies the ethical standards followed.
- 7.MEDIUMstatisticsAdd a brief note on how outliers or missing data were handled in the primary analysis, even if the mixed model is robust to missingness.Clarifying missing data and outlier handling strengthens the statistical reporting.
- 8.MEDIUMdata codeInclude a statement on code sharing, even if no custom code was used, to clarify that analysis was performed with standard statistical software.A clear code-sharing statement prevents ambiguity about the availability of analysis code.
- 9.LOWcopyeditCorrect the typo 'nonsteriodal antiinflammatory drugs' to 'nonsteroidal anti-inflammatory drugs' in the Methods, Trial Procedures section.Fixing the misspelling improves professionalism and readability.
- 10.LOWreportingIn the Discussion, explicitly address how the study addresses limitations of prior research, such as the lack of large implementation trials.Explicitly framing the study's contribution relative to prior gaps strengthens the scientific premise.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.