Artificial intelligence guided screening for cardiomyopathies in an obstetric population: a pragmatic randomized clinical trial.
Adedinsewo DA, Morales-Lara AC, Afolabi BB, Kushimo OA, Mbakwem AC, Ibiyemi KF, Ogunmodede JA, Raji HO, Ringim SH, Habib AA, Hamza SM, Ogah OS, Obajimi G, Saanu OO, Jagun OE, Inofomoh FO, Adeolu T, Karaye KM, Gaya SA, Alfa I, Yohanna C, Venkatachalam KL, Dugan J, Yao X, Sledge HJ, Johnson PW, Wieczorek MA, Attia ZI, Phillips SD, Yamani MH, Tobah YB, Rose CH, Sharpe EE, Lopez-Jimenez F, Friedman PA, Noseworthy PA, Carter RE, SPEC-AI Nigeria Investigators
- DOI
- 10.1038/s41591-024-03243-9
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/5b722775-1751-437e-b1cc-c9e210475ddb is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is detection of left ventricular systolic dysfunction (LVSD) defined as LVEF < 50% on echocardiography, which is a surrogate for clinical outcomes. The paper does not provide evidence linking LVSD detection to improved clinical outcomes, and target engagement at the tested dose is not applicable since it's a screening intervention. The claim of improved diagnosis is based on a surrogate biomarker without demonstrating that this leads to improved patient outcomes.
“The primary end point was identification of LVSD during the study period.”
- 02Treatment effect not shown to be clinically meaningful
The primary effect is an increase in detection of LVSD from 2.0% to 4.1% (absolute difference 2.1 percentage points). While statistically significant, the clinical meaningfulness is not anchored to any established minimal clinically important difference or patient-centered outcome. The number needed to screen is 47, but the benefit of detecting these cases is not shown to translate into improved outcomes.
“Using the AI-enabled digital stethoscope, the primary study end point was met with detection of 24 out of 587 (4.1%) versus 12 out of 608 (2.0%) patients with LVSD (intervention versus control odds ratio 2.12, 95% CI 1.05–4.27; P = 0.032).”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported pragmatic randomized trial with a strong scientific premise, rigorous design, and appropriate statistical analysis. The main weakness is the vague data availability statement and lack of code sharing, which limits reproducibility.
Both reviewers independently scored all eight dimensions and agreed on every status; no divergence to reconcile. The statistics verification recomputed 6 tests (all consistent) but does not cover all reported statistics; the citation check found no retracted or unresolved references; the reproducibility check found both links live.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 6 tests: 6 consistent, 0 inconsistent; 5 recomputed directly from the reported test statistics, 1 via agent-written checks.
- CONSISTENTreported p = .032 · recomputed p = .036Recomputed odds ratio 2.12 (95% CI 1.05–4.27), reported p=0.032
“odds ratio 2.12, 95% CI 1.05–4.27; P = 0.032”
Taken as given: 1.05–4.27 is a two-sided 95% confidence interval for the odds ratio of 2.12, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.032 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.12, 1.05, 4.27, 1) - CONSISTENTreported p = .125 · recomputed p = .130Recomputed odds ratio 1.75 (95% CI 0.85–3.62), reported p=0.125
“odds ratio 1.75, 95% CI 0.85–3.62; P = 0.125”
Taken as given: 0.85–3.62 is a two-sided 95% confidence interval for the odds ratio of 1.75, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.125 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.75, 0.85, 3.62, 1) - CONSISTENTreported p = .227 · recomputed p = .232Recomputed odds ratio 1.57 (95% CI 0.75–3.29), reported p=0.227
“odds ratio 1.57, 95% CI 0.75–3.29; P = 0.227”
Taken as given: 0.75–3.29 is a two-sided 95% confidence interval for the odds ratio of 1.57, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.227 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.57, 0.75, 3.29, 1) - CONSISTENTreported p = .976 · recomputed p = 1.000Recomputed OR 1.00 (95% CI 0.74–1.35), reported p=0.976
“OR 1.00, 95% CI 0.74–1.35 P = 0.976”
Taken as given: 0.74–1.35 is a two-sided 95% confidence interval for the OR of 1.00, not a range, an IQR, or a different interval level; the OR is a RATIO measure, so the interval is symmetric on the log scale; p=0.976 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1, 0.74, 1.35, 1) - CONSISTENTreported p = .621 · recomputed p = .639Recomputed OR 1.10 (95% CI 0.74–1.64), reported p=0.621
“OR 1.10, 95% CI 0.74–1.64, P = 0.621”
Taken as given: 0.74–1.64 is a two-sided 95% confidence interval for the OR of 1.10, not a range, an IQR, or a different interval level; the OR is a RATIO measure, so the interval is symmetric on the log scale; p=0.621 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.1, 0.74, 1.64, 1) - CONSISTENTreported p = .026 · recomputed p = .026Reviewers 1, 2All-cause mortality: HR 4.20, 95% CI 1.18-14.87, P=0.026
“All-cause mortality d | 12/587 | 3/608 | 4.20 (1.18, 14.87) | 0.026”
Taken as given: The hazard ratio is 4.20 with 95% CI 1.18-14.87.; The CI is a 95% confidence interval for the hazard ratio.; The p-value is two-sided.Method: Recomputed p-value from the reported hazard ratio and 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(4.20, 1.18, 14.87, 1)
- lowinternal contradictionThe number of LVSD cases at study end for the digital stethoscope is 24, but the text says 'three additional LVSD cases were identified with two in the intervention arm and one in the control arm' after 22 at entry, which sums to 24, but the control arm total is 12, which is 11+1, consistent.
“At the study end, three additional LVSD cases were identified with two in the intervention arm and one in the control arm, resulting in a total of 24 cases (4.1%) of LVSD (LVEF < 50%) identified in the intervention arm compared to 12 (2.0%) in the control arm”
ResultsFind in source - lowinternal contradictionThe abstract reports 1,232 randomized, but the main text says 1,232 randomized and 1,195 completed baseline. The number 1,196 appears only in the summary paragraph.
“A total of 1,232 (616 in each arm) participants were randomized and 1,195 participants (587 intervention arm and 608 control arm) completed the baseline visit”
AbstractFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
7 major claims checked against the paper's own evidence: 2 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2AI-based electrocardiogram screening proved accurate in detecting cardiomyopathies.The 12-lead AI-ECG did not reach statistical significance for the primary outcome (OR 1.75, P=0.125), but the diagnostic performance (AUC) was high.Evidence: Primary outcome for 12-lead ECG: OR 1.75, 95% CI 0.85-3.62, P=0.125; AUC for LVEF<50% was 0.928.
“AI-based electrocardiogram screening proved accurate in detecting cardiomyopathies and suggests that it could improve detection of these conditions.”
AbstractFind in source - partialReviewer 1The higher all-cause mortality in the intervention arm is likely due to differential ascertainment.The paper offers plausible explanations but does not provide direct evidence to confirm differential ascertainment.Evidence: Discussion of potential explanations: increased contact with healthcare services, mortality ascertainment prone to error.
Potential explanations or hypothesis for the observed higher mortality in the intervention arm include (1) a differential observation of mortality in the intervention arm due to increased contact with healthcare services...
Discussion ¶7reviewer’s wording - supportedReviewers 1, 2AI-guided screening using a digital stethoscope improved the diagnosis of pregnancy-related cardiomyopathy.The primary outcome was met with a statistically significant odds ratio of 2.12 (95% CI 1.05-4.27, P=0.032) for the digital stethoscope.Evidence: Primary outcome analysis: 24/587 vs 12/608, OR 2.12, 95% CI 1.05-4.27, P=0.032.
“In pregnant and postpartum women, AI-guided screening using a digital stethoscope improved the diagnosis of pregnancy-related cardiomyopathy.”
AbstractFind in source - supportedReviewer 1The study demonstrates a high prevalence of LVSD in an obstetric population in Nigeria.The prevalence of LVSD in the intervention arm was 4.1% (24/587), which is higher than previously reported estimates.Evidence: Primary outcome: 24/587 (4.1%) in intervention arm.
“We demonstrate a high prevalence of LVSD, in an obstetric population in Nigeria that supports the need for screening.”
Discussion ¶1Find in source - supportedReviewers 1, 2The AI-enabled digital stethoscope detected cardiomyopathy with high sensitivity, specificity and negative predictive value.The diagnostic performance metrics reported in Table 3 show high sensitivity (95.7%) and NPV (99.8%) for LVEF<50% using max prediction.Evidence: Table 3: Max prediction for LVEF<50%: sensitivity 95.7%, specificity 82.0%, NPV 99.8%.
“We found that an AI-enabled digital stethoscope that analyzes single-lead ECG and phonocardiogram recordings, as well as an AI-enabled 12-lead ECG detected the presence of cardiomyopathy with high sensitivity, specificity and negative predictive value.”
Discussion ¶1Find in source - supportedReviewer 2The NNS to detect one additional case of LVSD was 47.The NNS is directly reported in the results.Evidence: Results: 'The estimated number needed to screen (NNS) to detect one additional case of LVSD was 47.'
“The estimated number needed to screen (NNS) to detect one additional case of LVSD was 47.”
ResultsFind in source - supportedReviewer 2AI-guided screening doubled the diagnosis of pregnancy-related cardiomyopathy compared to usual care.The odds ratio of 2.12 indicates a doubling of detection, and the result was statistically significant.Evidence: Primary outcome OR 2.12 (95% CI 1.05-4.27).
“AI-guided screening with a digital stethoscope doubled the diagnosis of pregnancy-related cardiomyopathy when compared to usual obstetric care”
DiscussionFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is detection of left ventricular systolic dysfunction (LVSD) defined as LVEF < 50% on echocardiography, which is a surrogate for clinical outcomes. The paper does not provide evidence linking LVSD detection to improved clinical outcomes, and target engagement at the tested dose is not applicable since it's a screening intervention. The claim of improved diagnosis is based on a surrogate biomarker without demonstrating that this leads to improved patient outcomes.
“The primary end point was identification of LVSD during the study period.”
- INADEQUATEEffect sizeThe primary effect is an increase in detection of LVSD from 2.0% to 4.1% (absolute difference 2.1 percentage points). While statistically significant, the clinical meaningfulness is not anchored to any established minimal clinically important difference or patient-centered outcome. The number needed to screen is 47, but the benefit of detecting these cases is not shown to translate into improved outcomes.
“Using the AI-enabled digital stethoscope, the primary study end point was met with detection of 24 out of 587 (4.1%) versus 12 out of 608 (2.0%) patients with LVSD (intervention versus control odds ratio 2.12, 95% CI 1.05–4.27; P = 0.032).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites multiple prior studies (retrospective and pilot prospective) showing AI-ECG effectiveness, and notes the high incidence of peripartum cardiomyopathy in Nigeria. The rationale for the trial is clearly stated: to determine whether AI-guided screening improves detection beyond standard care. Limitations of prior work (e.g., small pilot samples, retrospective designs) are implicitly addressed by conducting a large prospective randomized trial.
“A retrospective study (area under the curve (AUC) = 0.89) and a pilot prospective study among pregnant and postpartum women in the United States showed AI-based screening to be effective (AUC = 1.00 using a 12-lead ECG and 0.98 using a digital stethoscope) in identifying pregnancy-related LVSD with LVEF < 45%.”
“To address this question, we conducted an open-label, randomized, pragmatic clinical trial among pregnant and postpartum women to evaluate whether AI-guided screening (using a digital stethoscope and 12-lead ECG) improves the diagnosis of pregnancy-related LVSD in an obstetric population in Nigeria compared to usual care.”
“A retrospective study (area under the curve (AUC) = 0.89) and a pilot prospective study among pregnant and postpartum women in the United States showed AI-based screening to be effective (AUC = 1.00 using a 12-lead ECG and 0.98 using a digital stethoscope) in identifying pregnancy-related LVSD with LVEF < 45%.”
“it remains unknown whether AI-guided screening improves cardiomyopathy detection in obstetric patients beyond the current standard of care.”
“we conducted an open-label, randomized, pragmatic clinical trial among pregnant and postpartum women to evaluate whether AI-guided screening (using a digital stethoscope and 12-lead ECG) improves the diagnosis of pregnancy-related LVSD in an obstetric population in Nigeria compared to usual care.”
Randomization used dynamic minimization with site stratification via a web-based application. The unit of randomization is the individual participant. A power analysis is reported (848 women for 80% power, increased to 1,200). Inclusion/exclusion criteria are clearly listed. The mITT analysis set is defined, and an ITT sensitivity analysis is performed. Blinding is not applicable as the trial is open-label; this is stated in the title and methods. Outlier handling is addressed through the mITT definition and conservative assumptions for poor-quality AI predictions.
“Randomization was performed in real time using dynamic minimization with the study site as a stratification factor, through a web-based application (iMedidata).”
“Randomization was performed in real time using dynamic minimization with the study site as a stratification factor, through a web-based application (iMedidata).”
The study reports age, race (all Black), ethnicity, weight, height, blood pressure, heart rate, hemoglobin, and comorbidities such as hypertensive disorders. Sex is inherently female due to the obstetric population, and the study is single-sex by design, which is justified by the condition being pregnancy-related. Demographics are comprehensive.
“Age, years | 1,195 | 31 (27–35) | 31 (26–35)”
“Inclusion criteria were female, aged 18–49 years, pregnant or within 12 months postpartum”
“The median age was 31 years and all women identified as Black. Fifty-five percent were of Yoruba ethnicity, 28% Hausa, 11% Igbo and 6% were from other ethnic groups.”
“At baseline, 150 (12.5%) women had been diagnosed with a hypertensive disorder of pregnancy (which includes chronic hypertension, gestational hypertension, pre-eclampsia and eclampsia) during the index pregnancy.”
The methods state that the study was approved by the Mayo Clinic institutional review board and local ethics research committees at all participating sites. Informed consent is described as written or oral, in accordance with local approvals. Regulatory compliance is implied through adherence to local ethics committee approvals and the Declaration of Helsinki is not explicitly named, but the approval statements are sufficient.
“The study was approved by the Mayo Clinic institutional review board as well as local ethics research committees at all participating sites in Nigeria.”
“All study participants provided written or oral informed consent in accordance with local ethics research committee approvals.”
“The study was approved by the Mayo Clinic institutional review board as well as local ethics research committees at all participating sites in Nigeria.”
“All study participants provided written or oral informed consent in accordance with local ethics research committee approvals.”
The digital stethoscope (Eko DUO) and 12-lead ECG machine (GE Marquette 2000) are identified. The AI algorithms are named (US FDA-cleared 12-lead AI-ECG, Mayo Clinic model, digital stethoscope model) with version numbers (e.g., lvef_v2.2.0, ELEFT 7.2.0). Statistical software R v.4.1.2 is identified. Since this is a device trial, antibodies, cell lines, mycoplasma, and organisms are not applicable.
“Statistical analyses were performed using R v.4.1.2.”
“The US FDA-cleared model had an in-built ECG data quality check. As such ECGs deemed to be of insufficient quality did not have AI predictions generated (algorithm version lvef_v2.2.0).”
“The 12-lead ECGs were acquired in a standard fashion as with routine clinical care, in a supine or semi-recumbent position for 10 s on standard ECG paper, at a sampling rate of 500 Hz using a GE Marquette 2000 ECG machine (GE Healthcare).”
The primary analysis uses logistic regression with Pearson chi-squared test, and p-values are reported exactly (e.g., P = 0.032). Effect sizes are reported as odds ratios with 95% CIs. Assumptions are handled by design (logistic regression, Cox for mortality). Data presentation includes CONSORT diagram, forest plots, and tables with per-group n. Mathematical plausibility checks: the reported percentages and counts appear consistent (e.g., 24/587 = 4.1%, 12/608 = 2.0%).
“The odds ratio and 95% large sample CI was estimated using a logistic regression model and statistical significance was assessed with a Pearson chi-squared test at the α = 0.05 level of significance (two-sided).”
“intervention versus control odds ratio 2.12, 95% CI 1.05–4.27; P = 0.032”
“The odds ratio and 95% large sample CI was estimated using a logistic regression model and statistical significance was assessed with a Pearson chi-squared test.”
“odds ratio 2.12, 95% CI 1.05–4.27; P = 0.032”
“Statistical analyses were performed using R v.4.1.2.”
The data availability statement says data can be made available upon request with an analysis plan, but it does not specify a repository or a managed-access platform. It mentions a Material Transfer Agreement and a 1-month response timeframe, which is somewhat concrete but still relies on contacting the corresponding author. Code is not shared because it is proprietary, which is understandable but limits reproducibility.
“The underlying data supporting the findings of this study can be made available to clinical investigators and researchers upon request. Written requests for data sharing including an analysis plan will be required before approval.”
“The code itself cannot be shared because it is proprietary intellectual property that has been licensed to Anumana and Eko Health.”
“The underlying data supporting the findings of this study can be made available to clinical investigators and researchers upon request. Written requests for data sharing including an analysis plan will be required before approval.”
“The code itself cannot be shared because it is proprietary intellectual property that has been licensed to Anumana and Eko Health.”
The trial is registered (NCT05438576). The CONSORT-AI extension is mentioned. All prespecified outcomes are reported, including exploratory ones. Limitations are thoroughly discussed, including referral bias, attrition, and the definition of LVSD. Conclusions are proportional, acknowledging the higher mortality in the intervention arm and the need for further study. Funding sources and competing interests are disclosed.
“ClinicalTrials.gov registration: NCT05438576”
“The CONSORT-AI Extension guideline was utilized in reporting study design and results.”
“The key limitations of this study were introduced by the pragmatic clinical trial design and enrolling study participants at teaching hospitals with a licensed cardiologist and echocardiography capabilities.”
“ClinicalTrials.gov registration: NCT05438576”
“The CONSORT-AI Extension guideline was utilized in reporting study design and results.”
“The key limitations of this study were introduced by the pragmatic clinical trial design and enrolling study participants at teaching hospitals with a licensed cardiologist and echocardiography capabilities.”
Registered (1 ID: ClinicalTrials.gov). Reporting guidelines cited: CONSORT, STARD.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 38 references by DOI: 3 verified — 35 no DOI (shown, not verified).
- NO DOIEpidemiology of peripartum cardiomyopathy: incidence, predictors, and outcomesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIClinical features and outcomes of peripartum cardiomyopathy in nigeriaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIncidence, clinical characteristics, and risk factors of peripartum cardiomyopathy in Nigeria: results from the PEACE RegistryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWorldwide incidence of peripartum cardiomyopathy and overall maternal mortalityNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPregnancy-related cardiovascular deaths in California: beyond peripartum cardiomyopathyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPeripartum cardiomyopathy: from genetics to managementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOITeam-based care of women with cardiovascular disease from pre-conception through pregnancy and postpartum: JACC Focus Seminar 1/5No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn artificial intelligence-enabled ECG algorithm for the identification of patients with atrial fibrillation during sinus rhythm: a retrospective analysis of outcome predictionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDetection of hypertrophic cardiomyopathy using a convolutional neural network-enabled electrocardiogramNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIElectrocardiogram screening for aortic valve stenosis using artificial intelligenceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDeep learning electrocardiographic analysis for detection of left-sided valvular heart diseaseNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScreening for cardiac contractile dysfunction using an artificial intelligence-enabled electrocardiogramNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence-enabled electrocardiograms for identification of patients with low ejection fraction: a pragmatic, randomized clinical trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDetecting cardiomyopathies in pregnancy and the postpartum period with an electrocardiogram-based deep learning modelNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArtificial intelligence based screening for cardiomyopathy in an obstetric population: a pilot studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDevelopment and validation of an electrocardiographic artificial intelligence model for detection of peripartum cardiomyopathyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn artificial intelligence electrocardiogram analysis for detecting cardiomyopathy in the peripartum periodNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPoint-of-care screening for heart failure with reduced ejection fraction using artificial intelligence during ECG-enabled stethoscope examination in London, UK: a prospective, observational, multicentre studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAutomated detection of low ejection fraction from a one-lead electrocardiogram: application of an AI algorithm to an electrocardiogram-enabled digital stethoscopeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICardiovascular diseases in nigeria: current status, threats, and opportunitiesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFirst trimester preeclampsia screening and predictionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScreening for depression among the general adult population and in women during pregnancy or the first-year postpartum: two systematic reviews to inform a guideline of the canadian task force on preventive health careNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe mathematical limitations of fetal echocardiography as a screening tool in the setting of a normal second-trimester ultrasoundNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffect of second-trimester sonographic cervical length on the risk of spontaneous preterm delivery in different risk groups: a prospective observational multicenter studyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Mobile Economy: Sub Saharan Africa 2022No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPeripartum cardiomyopathy: JACC state-of-the-art reviewNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIVital signs: pregnancy-related deaths, United States, 2011–2015, and strategies for prevention, 13 states, 2013–2017No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRacial and ethnic disparities in maternal mortality in the united states using enhanced vital records, 2016‒2017No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILong-term outcomes of women with peripartum cardiomyopathy having subsequent pregnanciesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIConsequences of maternal mortality on infant and child survival: a 25-year longitudinal analysis in Butajira Ethiopia (1987–2011)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIClinical trials overview: from explanatory to pragmatic clinical trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEko Low Ejection Fraction Tool (ELEFT): Reduced Ejection Fraction Machine Learning-based Notification SoftwareNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIScreening for peripartum cardiomyopathies using artificial intelligence in Nigeria (SPEC-AI Nigeria): clinical trial rationale and designNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extensionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAge and sex estimation using artificial intelligence from standard 12-lead ECGsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://clinicaltrials.gov/study/NCT05438576LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/ct2/show/NCT05438576LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORconsistencyAbstract“1,232 (616 in each arm) participants were randomized and 1,195 participants (587 intervention arm and 608 control arm) completed the baseline visit”→ Ensure the numbers are consistent throughout the paper.The abstract states 1,232 randomized, but the main text says 1,232 randomized and 1,195 completed baseline. This is consistent, but the abstract also mentions 1,196 in the summary paragraph, which is inconsistent.
- MINORconsistencyAbstract, summary paragraph“In this pragmatic, randomized clinical trial involving 1,196 pregnant and postpartum women”→ Change to 1,195 to match the mITT analysis set.The number 1,196 appears only in the summary paragraph and conflicts with the 1,195 reported elsewhere.
- MINORclarityMethods, Statistical analysis“The estimates were rounded up to 500 per group to account for uncertainties in the calculations.”→ Clarify that the sample size was increased to 500 per group, and later to 1,200 total.The sentence is slightly confusing because it says 'rounded up to 500 per group' but the total was later increased to 1,200.
- MINORconsistencyAbstract“1,232 (616 in each arm) participants were randomized and 1,195 participants (587 intervention arm and 608 control arm) completed the baseline visit”→ Ensure consistent use of 'intervention arm' vs 'intervention group' throughout.Minor inconsistency in terminology.
- MINORtypoResults, Sample characteristics“1,195 (587 in the intervention arm and 608 in the control arm) completed baseline assessments”→ Check for duplicate '1,195' in text.Potential duplication.
- MINORclarityMethods, Statistical analysis“The estimates were rounded up to 500 per group to account for uncertainties in the calculations.”→ Clarify that the sample size was increased to 500 per group, then later to 600 per group.Could be clearer.
The published work is robust and well-reported; an informed reader should weigh the limited data/code availability and the minor internal inconsistency in the participant count (1,196 vs 1,195) as the main caveats. No erratum is warranted for the core findings, but the authors should consider clarifying the data access mechanism and correcting the count discrepancy.
- 1.HIGHdata codeIn the Data Availability section, specify a concrete managed-access platform (e.g., Vivli or YODA) or a named data access committee with clear conditions and a defined timeline for response.The current statement relies on contacting the corresponding author and lacks a transparent, independent review process, which limits reproducibility.
- 2.HIGHdata codeConsider depositing de-identified aggregate data or summary statistics in a public repository (e.g., Figshare or Zenodo) to enhance transparency while protecting patient privacy.Even if raw data cannot be shared, aggregate data would improve the reproducibility of the reported results.
- 3.HIGHreportingCorrect the participant count inconsistency in the Abstract summary paragraph: change '1,196' to '1,195' to match the mITT analysis set reported elsewhere.The number 1,196 appears only in the summary paragraph and conflicts with the 1,195 reported in the abstract and main text, which could confuse readers.
- 4.MEDIUMreportingClarify the sample size calculation in Methods, Statistical analysis: state that the estimate was rounded up to 500 per group and later increased to 600 per group (1,200 total) to account for uncertainties.The current sentence 'rounded up to 500 per group' is confusing given the final sample size of 1,200.
- 5.MEDIUMreportingEnsure consistent terminology throughout the manuscript: use either 'intervention arm' or 'intervention group' consistently.Minor inconsistency in terminology could distract readers and reduce clarity.
- 6.MEDIUMreportingIn the CONSORT flow, report the number of participants screened but not enrolled, and the number who withdrew consent or were excluded for each reason.Improves transparency and completeness of the participant flow, as recommended by CONSORT.
- 7.MEDIUMreportingClarify the role of the data safety monitoring board (or lack thereof) and whether any interim analyses were performed.Readers need to know how safety was monitored in this open-label trial.
- 8.MEDIUMreportingProvide more detail on the handling of missing data for secondary outcomes and any sensitivity analyses performed.Missing data handling is a common reviewer concern and affects the robustness of secondary findings.
- 9.MEDIUMreportingClarify the definition of 'clinical recognition' in the control arm and how it was ascertained to reduce potential bias.The primary outcome depends on this definition, and ambiguity could affect interpretation.
- 10.MEDIUMreportingReport the intraclass correlation coefficient for echocardiogram readings between local and central review.Provides evidence of the reliability of the outcome measurement.
- 11.LOWreportingSpecify the exact version of the CONSORT-AI checklist used and where it is available.Enhances transparency and allows readers to verify adherence to the reporting guideline.
- 12.LOWreportingInclude a statement on whether any adverse events were related to the AI screening itself.Addresses a potential safety concern specific to the intervention.
- 13.LOWdata codeConsider sharing the statistical analysis code (even if the AI algorithms are proprietary) to improve reproducibility.The statistical code is not proprietary and sharing it would allow independent verification of the analyses.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.