Infant mental health services for birth and foster families of maltreated pre-school children in foster care (BeST(?)): a cluster-randomized phase 3 clinical effectiveness trial.
Crawford K, Young R, Wilson P, Deidda M, Forde M, Millar S, McConnachie A, Boyd K, McIntosh E, Ougrin D, Henderson M, Gillberg C, Kainth G, Turner F, Sonuga-Barke EJS, Fitzpatrick B, Minnis H
- DOI
- 10.1038/s41591-025-03534-9
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/6bf7cdc9-9f7c-4259-8779-c88049e6ed41 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- StatisticsStatistic did not reproduce−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 12 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Printed percentage does not match its own count
72.8% does not match the reported count 156/216
“156 (72.8%)”
Results
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported cluster-randomized trial of a complex intervention (NIM) versus services as usual for maltreated pre-school children in foster care. The design, statistical methods, and reporting are rigorous, with only minor reporting gaps (explicit regulatory compliance statement, code sharing) and a few copyedit inconsistencies.
Both reviewers classified the study as interventional (cluster-randomized trial); no divergence. The evaluation covered the full text, including methods, results, tables, and supplementary materials. Non-applicable sub-criteria (e.g., animal housing, cell lines) were excluded. The statistics verification covered only a subset of reported tests (5 tests; 4 consistent, 1 inconsistent but unspecified); the absence of a GRIM finding is not a check that passed. Citation verification found no retracted or unresolved references.
Numerical inconsistencies
2 findings · worst highValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 4 tests: 4 consistent, 0 inconsistent; 4 via agent-written checks. 1 reported summary statistic mathematically impossible for the stated N (PERCENT).
- PERCENT72.8% does not match the reported count 156/216
“156 (72.8%)”
Results
- CONSISTENTreported p = .170 · recomputed p = .175Reviewers 1, 2Primary outcome SDQ-TD mean difference p-value
“SDQ-TD score a | 11.5 (7.6) | 11.1 (7.2) | 1.44 | (−0.63, 3.53) | 0.17”
Taken as given: The estimate is a mean difference (not a ratio).; The 95% CI is two-sided.; The p-value is two-sided.Method: Recomputed p-value from the reported mean difference and 95% CI using the normal approximation.How we recomputed it: pCI(1.44, -0.63, 3.53, 0) - CONSISTENTreported p = .330 · recomputed p = .331Reviewers 1, 2Secondary outcome PIRGAS p-value
“PIRGAS score a | 83.66 (12.73) | 84.83 (9.85) | −2.82 | (−8.53, 2.84) | 0.33”
Taken as given: The estimate is a mean difference.; The 95% CI is two-sided.; The p-value is two-sided.Method: Recomputed p-value from the reported mean difference and 95% CI using the normal approximation.How we recomputed it: pCI(-2.82, -8.53, 2.84, 0) - CONSISTENTreported p = .630 · recomputed p = .649Reviewers 1, 2Secondary outcome PedsQL p-value
“PedsQL score a | 87.2 (12.3) | 87.3 (12.4) | −0.05 | (−0.27, 0.16) | 0.63”
Taken as given: The estimate is a mean difference.; The 95% CI is two-sided.; The p-value is two-sided.Method: Recomputed p-value from the reported mean difference and 95% CI using the normal approximation.How we recomputed it: pCI(-0.05, -0.27, 0.16, 0) - CONSISTENTreported p = .930 · recomputed p = .915Reviewers 1, 2Secondary outcome TTPLS HR p-value
“Time (years) to permanent legal status (HR) b | 1.38 (0.67) | 1.32 (0.73) | 0.98 | (0.68, 1.43) | 0.93”
Taken as given: The estimate is a hazard ratio (log scale).; The 95% CI is two-sided.; The p-value is two-sided.Method: Recomputed p-value from the reported hazard ratio and 95% CI using the normal approximation on the log scale.How we recomputed it: pCI(0.98, 0.68, 1.43, 1)
- lowinternal contradictionThe abstract reports 286 families (79.4%) followed up at 2.5 years, but the results section says 285 families (79.2%) with 374 children were followed up at 15 months and 286 families (79.4%) with 367 children at 2.5 years. The numbers are consistent but the abstract omits the 15-month follow-up.
In total, 286 families (149 NIM and 137 SAU, 367 children) were followed-up (79.4%). ... 285 families (79.2%) with 374 children were followed-up 15 months after entry to the study, and 286 families (79.4%) with 367 children were followed-up 2.5 years after entry to the study.
Abstractreviewer’s wording - lowinternal contradictionThe abstract states 382 families with 488 children were randomized, but the results state 360 families (464 children) were randomized after excluding 22 families (24 children). The abstract may be referring to the total consented, not randomized.
“382 families with 488 0–5-year-old children, entering foster care, were randomized to the New Orleans Intervention Model (NIM) or social work services as usual (SAU).”
AbstractFind in source
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
6 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2No intervention effect of NIM on child mental health as measured by SDQ-TD at 2.5 years.The primary outcome analysis shows no statistically significant difference, with a p-value of 0.17 and a confidence interval crossing zero.Evidence: Table 2: SDQ-TD mean difference 1.44 (95% CI -0.63 to 3.53), P=0.17.
“Intention-to-treat analysis found no intervention effect of NIM: mean (s.d.) SDQ-TD NIM, 11.5 (7.6); SAU, 11.1 (7.2); adjusted mean difference (NIM − SAU), 1.4; 95% confidence interval (−0.63, 3.53); P = 0.17.”
AbstractFind in source - supportedReviewers 1, 2No within-trial effects for primary or secondary outcomes were observed.All secondary outcomes (PedsQL, PIRGAS, TTPLS) show non-significant p-values and confidence intervals crossing zero.Evidence: Table 2: PedsQL p=0.63, PIRGAS p=0.33, TTPLS HR 0.98 (95% CI 0.68-1.43), p=0.93.
“No within-trial effects for primary or secondary outcomes were observed.”
AbstractFind in source - supportedReviewers 1, 2NIM could be delivered to only 66.4% of eligible families.The compliance rate is reported in the results section.Evidence: Results: 'Of the 223 children from 180 families randomized to NIM, 148 (66.4%) were from families deemed to be compliant with the intervention.'
Of the 223 children from 180 families randomized to NIM, 148 (66.4%) were from families deemed to be compliant with the intervention.
Results ¶2reviewer’s wording - supportedReviewers 1, 2Girls randomized to NIM had significantly poorer SDQ scores compared to SAU.The subgroup analysis shows a significant interaction with a p-value of 0.0217 and a confidence interval not crossing zero.Evidence: Results: 'there was a significant interaction (P = 0.0217) between sex and arm of trial such that girls randomized to NIM had significantly poorer SDQ scores compared to SAU (2.96; 95% CI (0.12, 5.82))'
there was a significant interaction ( P = 0.0217) between sex and arm of trial such that girls randomized to NIM had significantly poorer SDQ scores compared to SAU (2.96; 95% confidence interval (CI) (0.12, 5.82))
Resultsreviewer’s wording - supportedReviewers 1, 2The rate of permanent placement was more than four times higher in London than in Glasgow.The hazard ratio of 4.36 with a highly significant p-value supports this claim.Evidence: Results: 'HR = 4.36; 95% CI (2.96, 6.40); P < 0.0001'
“A large and statistically significant difference was observed in the rate at which children were placed permanently with a legal order across 2.5 years in England versus Scotland: HR = 4.36; 95% CI (2.96, 6.40); P < 0.0001.”
ResultsFind in source - supportedReviewer 2The UK legal context prevented delivery of NIM to all eligible families.The paper provides qualitative process evaluation and compliance data supporting this claim.Evidence: Discussion and process evaluation findings
“Despite this, NIM could be delivered to only 66.4% of eligible families.”
Discussion ¶2Find in source
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior research on early intervention, foster care mental health, and the limited success of previous trials. It identifies a gap: only one multimodal intervention (NIM) targeting both birth and foster families was found. The rationale for the trial is clearly linked to this gap, and the paper acknowledges limitations of prior work (e.g., small effect sizes, focus solely on foster families).
“We found only one intervention aiming to improve the mental health of infants and young children placed in foster care through such a multimodal approach.”
“There is no previous randomized controlled trial (RCT) evidence for NIM.”
“This has led us to ask the following research question: Is NIM effective in improving the mental health, relationship quality and placement stability of infants and pre-school children placed in foster care?”
Randomization used a mixed minimization/randomization method with stratification and blocks of 10, and the unit of randomization is the family. Blinding is described: researchers, data collectors, and statisticians were blinded, while participants were not (single-blind). A power calculation was performed (target 462 families for 90% power). Inclusion/exclusion criteria are pre-specified. Outlier handling is addressed through the pre-specified analysis population (ITT) and missing data approaches. Controls are inherent in the SAU comparator arm. Independent replication is not applicable for a single pivotal trial.
“Random allocation was performed, through a web portal requiring a login and password with access rights allocated to relevant Clinical Trials Unit staff, using a mixed minimization/randomization method, stratified within the study site using an a priori schedule for each site in blocks of 10.”
“Masking was assured for two secondary outcome measures, TTPLS and PIRGAS, because data were collected (TTPLS) and rated (PIRGAS) by individuals with no contact with participants or other trial procedures.”
“Random allocation was performed, through a web portal requiring a login and password with access rights allocated to relevant Clinical Trials Unit staff, using a mixed minimization/randomization method, stratified within the study site using an a priori schedule for each site in blocks of 10.”
“Masking was assured for two secondary outcome measures, TTPLS and PIRGAS, because data were collected (TTPLS) and rated (PIRGAS) by individuals with no contact with participants or other trial procedures. Participants were aware of the intervention to which they had been allocated; data collectors, researchers and statisticians were not.”
The study reports sex (48% female overall), age (mean and range), and demographics (ethnicity, deprivation index) in Table 1. Health status is indirectly captured through reasons for coming into care and baseline SDQ scores. Species/strain and housing conditions are not applicable for a human trial.
“In total, 360 families (464 children; 48% female) were randomized: 180 to NIM and 180 to SAU.”
“Mean age (s.d.; range) at baseline (years) | 2.00 (1.64; 0.01–5.38) | 1.98 (1.65; 0.01–5.40)”
“Sex (% female) | 96 (43.4%) | 108 (50.7)”
“Mean age (s.d.; range) at baseline (years) | 2.00 (1.64; 0.01–5.38) | 1.98 (1.65; 0.01–5.40)”
The paper states approval by 'West of Scotland Research Ethics Service, Committee 3 (15/WS/0280)'. Informed consent is described: initial consent and later deferred consent procedures are detailed. Regulatory compliance is not explicitly named (e.g., Declaration of Helsinki), but the ethics approval and consent process imply compliance; however, the absence of an explicit statement is minor.
“The trial was approved by West of Scotland Research Ethics Service, Committee 3 (15/WS/0280).”
“From January 2012 to June 2015, participants were asked for consent, assessed at baseline and then randomized.”
“The trial was approved by West of Scotland Research Ethics Service, Committee 3 (15/WS/0280).”
NIM is described as a complex intervention with components and actors. SAU is described as three models of social work services. Statistical software (R version 4.1.1) and packages (lme4, survival) are identified. No antibodies, cell lines, or organisms are used, so those are not applicable.
“NIM is delivered by a multidisciplinary infant mental health team comprising psychologists, a psychiatrist, therapists and social workers who assess the mental health and relationship quality of children younger than 5 years of age upon reception into foster care.”
“The primary analysis, conducted in R version 4.1.1 (2021-08-10) for Windows (with ‘lme4’ and ‘survival’ for TTPLS)”
“NIM is delivered by a multidisciplinary infant mental health team comprising psychologists, a psychiatrist, therapists and social workers who assess the mental health and relationship quality of children younger than 5 years of age upon reception into foster care.”
“The primary analysis, conducted in R version 4.1.1 (2021-08-10) for Windows (with ‘lme4’ and ‘survival’ for TTPLS)”
Tests are named (generalized linear mixed effects regression, Cox proportional hazard). Assumptions are handled by design (mixed models, Cox). Exact p-values are reported (e.g., P = 0.17). Effect sizes with 95% CIs are reported for all outcomes. Software is identified. Data presentation includes per-group n and dispersion. Mathematical plausibility checks: the primary outcome means and SDs are plausible; subgroup counts in Table 1 sum correctly (e.g., ethnicity percentages sum to 100% within rounding).
“SDQ-TD score a | 11.5 (7.6) | 11.1 (7.2) | 1.44 | (−0.63, 3.53) | 0.17”
“Mean difference (or HR) | 95% CI | P value”
“The primary analysis, conducted in R version 4.1.1 (2021-08-10) for Windows (with ‘lme4’ and ‘survival’ for TTPLS) , , used a generalized linear mixed effects regression model for the primary outcome measure (SDQ-TD), with random effects terms for both family and child”
“SDQ-TD score a | 11.5 (7.6) | 11.1 (7.2) | 1.44 | (−0.63, 3.53) | 0.17”
“PIRGAS score a | 83.66 (12.73) | 84.83 (9.85) | −2.82 | (−8.53, 2.84) | 0.33”
The data availability statement is concrete: it names the TMG, the secure platform (Robertson Centre for Biostatistics), and the process for requests. It states de-identified data will be available within 6 months. Repository deposit and accession numbers are not applicable for identifiable patient data. Code sharing is not reported, but the statistical methods are described in detail.
“Data requests will be considered by the Trial Management Group (TMG), which includes representatives of the sponsor, the University of Glasgow, senior investigators independent of the research team and the chief investigator.”
“Data requests will be considered by the Trial Management Group (TMG), which includes representatives of the sponsor, the University of Glasgow, senior investigators independent of the research team and the chief investigator.”
“Data access will be provided through the secure analytical platform of the Robertson Centre for Biostatistics.”
Trial registration numbers are provided (NCT02653716 and NCT01485510). CONSORT framework is used. All pre-specified outcomes are reported, including null results. Limitations are thoroughly discussed. Conclusions are proportional to the evidence. Funding sources and COI are stated.
“ClinicalTrials.gov registration: NCT02653716 (https://clinicaltrials.gov/study/NCT02653716)”
“We used the CONSORT framework to report the findings (see the checklist in ).”
“Smaller numbers of families recruited in London limited our ability to examine subgroup differences by site, and, because of heterogeneity across study sites, future trials will be required to determine whether NIM might be effective in certain contexts.”
“ClinicalTrials.gov registration: NCT02653716 (https://clinicaltrials.gov/study/NCT02653716)”
“We used the CONSORT framework to report the findings (see the checklist in ).”
“Smaller numbers of families recruited in London limited our ability to examine subgroup differences by site, and, because of heterogeneity across study sites, future trials will be required to determine whether NIM might be effective in certain contexts.”
Registered (2 IDs: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 33 references by DOI: 27 verified — 6 no DOI (shown, not verified).
- NO DOIPunishing the poor? Child welfare and protection under neoliberalismNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPrioritising early childhood to promote the nation’s health, wellbeing and prosperityNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWhere are the children?: addiction workers’ knowledge of clients’ offspring and related risksNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIChild maltreatmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIModeling Survival Data: Extending the Cox ModelNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIApplied Econometrics with RNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://clinicaltrials.gov/study/NCT02653716LIVEHTTP 200Resolves, but the content could not be matched to the paper.
- datahttps://clinicaltrials.gov/study/NCT01485510?cond=Evaluation%20of%20the%20new%20orleans&rank=1LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORtypoAbstract“BeST ?”→ BeST?The question mark is likely a placeholder for a registered trademark symbol.
- MINORconsistencyResults, paragraph 1“285 families (79.2%) with 374 children were followed-up 15 months after entry to the study, and 286 families (79.4%) with 367 children were followed-up 2.5 years after entry to the study.”→ Ensure consistency in the number of families followed up at each time point.The numbers differ slightly; verify the correct counts.
- MINORclarityMethods, Statistical analyses“The primary analysis, conducted in R version 4.1.1 (2021-08-10) for Windows (with ‘lme4’ and ‘survival’ for TTPLS)”→ Clarify that 'survival' is an R package.The sentence is slightly ambiguous.
- MINORconsistencyResults, paragraph 1“382 families with 488 children consented”→ 382 families with 488 children consentedLater text says 360 families (464 children) were randomized; the discrepancy is explained by exclusions, but the flow could be clearer.
- MINORclarityMethods, Statistical analyses“The primary analysis, conducted in R version 4.1.1 (2021-08-10) for Windows (with ‘lme4’ and ‘survival’ for TTPLS) , , used”→ Remove extra commas and ensure proper citation formatting.There are stray commas and citation placeholders.
The published work is robust and generally trustworthy, with minor reporting gaps that do not undermine the main conclusions. An informed reader should weigh the missing explicit regulatory compliance statement and the lack of public code sharing as minor transparency limitations; the one inconsistent statistic (unspecified) warrants a check by the authors to rule out a typo, but no decision errors were found.
- 1.HIGHethicsAdd an explicit statement of regulatory compliance (e.g., 'The trial was conducted in accordance with the Declaration of Helsinki') in the Methods section.Both reviewers flagged the absence of an explicit regulatory compliance statement, which is a standard ethical reporting requirement.
- 2.HIGHdata codeShare the analysis code in a public repository (e.g., GitHub) with a DOI, or at least provide the statistical analysis plan as a supplementary file with version control.Code sharing enhances reproducibility and is a common reviewer expectation for clinical trials.
- 3.HIGHstatisticsInvestigate the one inconsistent statistic flagged by the verification component (not specified in the report) and correct any typo or clarify the discrepancy.An unexplained inconsistency in a reported test statistic could undermine reader confidence, even if it is likely a typographical error.
- 4.MEDIUMreportingClarify the data availability statement by specifying the expected timeline for data access decisions and any criteria for approval.Reviewer 2 noted the statement lacks a timeline and approval criteria, which would improve transparency for potential data requesters.
- 5.MEDIUMreportingIn the Discussion, explicitly acknowledge the lack of a pre-specified plan for handling the sex interaction finding to avoid potential overinterpretation.Reviewer 2 suggested this to ensure the sex-related finding is not overinterpreted without a pre-specified analysis plan.
- 6.MEDIUMreportingProvide more detail on the blinding of outcome assessors for the primary outcome (SDQ) to ensure complete transparency.Reviewer 2 noted that blinding details for the primary outcome are less explicit than for secondary outcomes.
- 7.MEDIUMreportingConsider reporting the intracluster correlation coefficient for secondary outcomes as well.Reviewer 2 suggested this to provide a more complete picture of clustering effects.
- 8.MEDIUMreportingIn the CONSORT flow diagram, ensure all numbers reconcile (e.g., 382 families randomized vs. 360 after exclusions) with clear annotations.The copyedit and integrity checks noted a potential discrepancy between consented and randomized numbers that could confuse readers.
- 9.MEDIUMstatisticsAdd a statement on whether any analyses were adjusted for multiple comparisons, as the paper notes no adjustments were made.Reviewer 2 flagged this as a transparency issue for the interpretation of secondary outcomes.
- 10.MEDIUMreportingSpecify the version of the statistical analysis plan and where it can be accessed.Reviewer 2 suggested this to improve reproducibility and version control.
- 11.LOWcopyeditFix the typo in the Abstract: 'BeST ?' should be 'BeST?' (likely a placeholder for a registered trademark symbol).Copyedit flagged this as a minor typo that should be corrected for professionalism.
- 12.LOWcopyeditEnsure consistency in the number of families followed up at each time point (285 vs 286 families) in the Results section.Copyedit flagged a minor inconsistency that could confuse readers.
- 13.LOWcopyeditClarify that 'survival' is an R package in the Methods section.Copyedit noted the sentence is slightly ambiguous.
- 14.LOWcopyeditRemove stray commas and citation placeholders in the Methods section (e.g., 'with ‘lme4’ and ‘survival’ for TTPLS) , , used').Copyedit flagged formatting issues that should be cleaned up.
- 15.LOWdata codeConsider including a data-sharing agreement template or reference to facilitate data access requests.Reviewer 2 suggested this to streamline the data access process.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.