Coaching inexperienced clinicians before a high stakes medical procedure: randomized clinical trial.
Flynn SG, Park RS, Jena AB, Staffa SJ, Kim SY, Clarke JD, Pham IV, Lukovits KE, Huang SX, Sideridis GD, Bernier RS, Fiadjoe JE, Weinstock PH, Peyton JM, Stein ML, Kovatsis PG
- DOI
- 10.1136/bmj-2024-080924
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/d877cb41-3c6c-426f-aa82-42273165819b is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- ClaimsOverstated claim−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ReportingData & code availability partially met−0.25★
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is first attempt success rate of infant intubation, which is a procedural success metric, not a hard clinical outcome. The paper does not provide evidence linking first attempt success to long-term clinical outcomes, nor does it demonstrate target engagement for the intervention (coaching) beyond the immediate procedural success. The claim of improved patient safety is based on this surrogate.
“Primary outcome was the first attempt success rate of intraoperative infant intubation.”
- 02Conclusion reaches beyond the evidence
Just-in-time training could improve high stakes procedures more broadly.
“Our study’s findings also raise whether just-in-time training might be useful in procedures other than infant intubations, such as central lines and chest tubes.”
DiscussionFind in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed and clearly reported randomized clinical trial evaluating just-in-time coaching for infant intubation. The main methodological strengths are rigorous randomization, appropriate statistical analysis, and transparent reporting. The primary weakness is the vague data availability statement, which lacks a concrete access mechanism or timeframe.
This synthesis is based on two independent reviewer runs that agreed on all dimensions, plus a copyedit pass and verification components. The statistics verification recomputed only a subset of reported tests (6), all consistent; the remainder are unverified. The citation check found no retracted or unresolved references. The claim audit flagged one overstatement in the discussion.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 6 tests: 6 consistent, 0 inconsistent; 5 recomputed directly from the reported test statistics, 1 via agent-written checks.
- CONSISTENTreported p = .001 · recomputed p = <.001Recomputed odds ratio 2.42 (95% CI 1.45–4.04), reported p=0.001
“odds ratio 2.42 (95% confidence interval 1.45 to 4.04), P=0.001”
Taken as given: 1.45–4.04 is a two-sided 95% confidence interval for the odds ratio of 2.42, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.42, 1.45, 4.04, 1) - CONSISTENTreported p = .001 · recomputed p = <.001Recomputed odds ratio 3.18 (95% CI 1.62–6.24), reported p=0.001
“odds ratio 3.18 (95% CI 1.62 to 6.24), P=0.001”
Taken as given: 1.62–6.24 is a two-sided 95% confidence interval for the odds ratio of 3.18, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(3.18, 1.62, 6.24, 1) - CONSISTENTreported p = .001 · recomputed p = <.001Recomputed odds ratio 2.58 (95% CI 1.48–4.5), reported p=0.001
“odds ratio 2.58 (95% CI 1.48 to 4.5), P=0.001”
Taken as given: 1.48–4.5 is a two-sided 95% confidence interval for the odds ratio of 2.58, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(2.58, 1.48, 4.5, 1) - CONSISTENTreported p = .220 · recomputed p = .224Recomputed odds ratio 0.57 (95% CI 0.23–1.41), reported p=0.22
“odds ratio 0.57 (95% CI 0.23 to 1.41), P=0.22”
Taken as given: 0.23–1.41 is a two-sided 95% confidence interval for the odds ratio of 0.57, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.22 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.57, 0.23, 1.41, 1) - CONSISTENTreported p = .020 · recomputed p = .020Recomputed odds ratio 3.53 (95% CI 1.22–10.2), reported p=0.02
“odds ratio 3.53 (95% CI 1.22 to 10.2), P=0.02”
Taken as given: 1.22–10.2 is a two-sided 95% confidence interval for the odds ratio of 3.53, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.02 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(3.53, 1.22, 10.2, 1) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Primary outcome risk ratio p-value from reported RR and 95% CI
“1.12 (1.05 to 1.19), P<0.001”
Taken as given: The RR is 1.12 with 95% CI 1.05 to 1.19.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed two-tailed p from the reported RR and 95% CI using the normal approximation for the log risk ratio.How we recomputed it: pCI(1.12, 1.05, 1.19, 1)
- lowinternal contradictionThe abstract states 153 trainees were analyzed, but the results section says 172 were randomized and 5 withdrawn, leaving 167, then 14 excluded, leaving 153. This is consistent, but the abstract omits the intermediate steps.
“172 were randomized, and 153 were subsequently analyzed.”
AbstractFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated), 1 only partially supported (evidence backs part of the claim; gaps or caveats remain).
- overstatedReviewers 1, 2Just-in-time training could improve high stakes procedures more broadly.The study only tested infant intubation; generalizing to other procedures is speculative.Evidence: Discussion section speculates on broader applicability.
“Our study’s findings also raise whether just-in-time training might be useful in procedures other than infant intubations, such as central lines and chest tubes.”
DiscussionFind in source - partialReviewers 1, 2Complications were lower in the treatment group.Complication rates were lower but the difference was not statistically significant.Evidence: Secondary outcome: complication rate 2.75% vs 4.71%, OR 0.57 (0.23-1.41), P=0.22.
“The overall complication rate was 2.75% (7/255) in the treatment group and 4.71% (16/340) in the control group (odds ratio 0.57 (95% CI 0.23 to 1.41), P=0.22”
ResultsFind in source - supportedReviewers 1, 2Just-in-time training increased first attempt success of infant intubation.The primary outcome shows a significant improvement with OR 2.42 (95% CI 1.45-4.04), P=0.001.Evidence: Primary outcome result: 91.4% vs 81.6%, OR 2.42 (1.45-4.04), P=0.001.
“In modified intention-to-treat analysis, first attempt success was 91.4% (212/232) in the trainee treatment group and 81.6% (231/283) in the control group (odds ratio 2.42 (95% confidence interval 1.45 to 4.04), P=0.001).”
AbstractFind in source - supportedReviewers 1, 2Just-in-time training decreased cognitive load.NASA-TLX scores were significantly lower for mental demand, temporal demand, effort, and frustration.Evidence: Secondary outcomes: significant differences in NASA-TLX domains.
“Mental workload scores, measured by the NASA cognitive task load index, were significantly lower for mental demand (coefficient −9.5 (95% CI −16 to −3), P=0.004), temporal demand (−9.1 (−16.1 to −2.1), P=0.01), effort (−10.1 (−16 to −4.4), P=0.001), and frustration (−7.1 (−12.6 to −1.7), P=0.01) in the treatment group.”
ResultsFind in source - supportedReviewers 1, 2Just-in-time training improved competency metrics.The treatment group had better airway views, fewer advancement maneuvers, fewer technical difficulties, and faster intubation times.Evidence: Secondary outcomes: significant differences in technical skill metrics.
“The treatment group had more modified Cormack-Lehane grade 1 views (the best possible airway view) for video laryngoscopy than the control group, half the number of endotracheal tube advancement maneuvers, fewer technical difficulties during laryngoscopy, and faster intubation times”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit.
- INADEQUATESurrogate endpointThe primary outcome is first attempt success rate of infant intubation, which is a procedural success metric, not a hard clinical outcome. The paper does not provide evidence linking first attempt success to long-term clinical outcomes, nor does it demonstrate target engagement for the intervention (coaching) beyond the immediate procedural success. The claim of improved patient safety is based on this surrogate.
“Primary outcome was the first attempt success rate of intraoperative infant intubation.”
- ADEQUATEEffect sizeThe effect size is a 10 percentage point absolute improvement in first attempt success (91.4% vs 81.6%), with an odds ratio of 2.42 (95% CI 1.45 to 4.04). The paper explicitly states this is clinically meaningful, citing the harms associated with multiple intubation attempts. The effect is statistically supported and anchored to clinical meaningfulness.
“The improvement in first attempt success by 10 percentage points is clinically meaningful, considering the harms associated with multiple tracheal intubation attempts and the many trainee intubations performed yearly.”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
2 integrity concerns flagged (0 high).
- lowotherThe complication rates are reported as 2.75% (7/255) and 4.71% (16/340), but the denominators (255 and 340) are larger than the number of intubations (232 and 283), suggesting they are based on attempts, not intubations. This is explained in Table 3 but could be confusing.
“The overall complication rate was 2.75% (7/255) in the treatment group and 4.71% (16/340) in the control group”
ResultsFind in source
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites examples from sports and music, and discusses prior studies on just-in-time training for intubation, noting their non-randomized designs and timing differences. The rationale linking the premise to the study objectives is logical and the hypothesis follows from the cited evidence. Limitations of prior research are addressed by the randomized design and the timing of coaching within one hour.
“Previously, just-in-time simulation for tracheal intubation in the pediatric intensive care unit was compared to historical controls, and no difference in first attempt success rate was found. However, that study was not randomized, and in it, training could occur up to 24 hours before the clinical encounter, compared with training that occurred within an hour of intubation in our study.”
“Therefore, we conducted a randomized clinical trial to assess whether coaching inexperienced clinicians just before a procedure could improve the quality of procedural care.”
“Previously, just-in-time simulation for tracheal intubation in the pediatric intensive care unit was compared to historical controls, and no difference in first attempt success rate was found. However, that study was not randomized, and in it, training could occur up to 24 hours before the clinical encounter”
“Therefore, we conducted a randomized clinical trial to assess whether coaching inexperienced clinicians just before a procedure could improve the quality of procedural care.”
Randomization was stratified by trainee role and block randomized with a block size of four, using SAS PROC PLAN. The unit of randomization is the trainee. Blinding was not possible, but the paper explicitly states this and provides a rationale. A power analysis was conducted for the primary outcome. Inclusion/exclusion criteria are clearly defined. The analysis population (modified intention-to-treat) is described, and protocol deviations are reported. The trial is registered and has IRB approval.
“We stratified trainees by role (fellow, resident, student registered nurse anesthetist (SRNA)) and prospectively block randomized to the treatment or control group for tracheal intubation of children aged ≤12 months.”
“Using a χ 2 test with a 5% two sided α, a sample size of 200 intubations per group (400 total) provided 80% power to detect a difference of 80% versus 90% (10% absolute difference) in first attempt success rates.”
“Finally, masking of participants was impossible, given the study’s nature.”
“We stratified trainees by role (fellow, resident, student registered nurse anesthetist (SRNA)) and prospectively block randomized to the treatment or control group”
“Using a χ 2 test with a 5% two sided α, a sample size of 200 intubations per group (400 total) provided 80% power to detect a difference of 80% versus 90% (10% absolute difference) in first attempt success rates.”
“Finally, masking of participants was impossible, given the study’s nature.”
The study reports patient age, weight, ASA classification, and other intubation-level data. Trainee demographics are also reported. Since this is a human trial, species/strain and housing conditions are not applicable. Sex of patients is not explicitly reported, but this is not a critical omission for this type of study.
“Patient age (months) | 6 (3-9) | 6 (3-9) | 0.02 | | Patient weight (kg) | 7.2 (5.4-8.7) | 7.3 (5.5-8.8) | <0.01”
“Trainee type | | Resident | 45 (64.3) | 53 (63.9) | 0.08”
The paper states that the institutional review board of Boston Children's Hospital (P00034169) approved the study. Informed consent was obtained from trainees. The trial was conducted according to the Declaration of Helsinki. These meet the criteria for adequate reporting.
“The institutional review board of Boston Children’s Hospital (P00034169) approved the study.”
“Informed consent was obtained before trainee participation.”
“The institutional review board of Boston Children’s Hospital (P00034169) approved the study.”
“Informed consent was obtained before trainee participation.”
The intervention is a coaching session, which is described in detail. The manikin used is not specified by brand/model, but this is not a key biological/chemical resource. Statistical software (SAS, Stata) and REDCap are identified. No antibodies, cell lines, or reagents are used.
“We did statistical analyses using Stata (version 16 0.0, StataCorp, College Station, TX).”
“using the PROC PLAN procedure in SAS (version 9.4, SAS Institute, Cary, NC).”
“The treatment group received a standardized coaching session (supplementary figure S1) on an infant manikin within one hour of patient intubation”
“We did statistical analyses using Stata (version 16 0.0, StataCorp, College Station, TX).”
The paper names the statistical tests used (GEE, median regression, Fisher's exact test, mixed effects ordinal logistic regression). Exact p-values are reported (e.g., P=0.001). Effect sizes with 95% CIs are provided. Statistical software is identified. Data presentation includes per-group n and appropriate measures. Mathematical plausibility checks were not possible for all values, but no obvious errors were found.
“Multivariable generalized estimating equations modeling was implemented with a logit link (to estimate odds ratios) or log link (to estimate risk ratios) and binomial family to account for multiple intubations per participant”
“odds ratio 2.42 (95% CI 1.45 to 4.04), P=0.001”
“Multivariable generalized estimating equations modeling was implemented with a logit link (to estimate odds ratios) or log link (to estimate risk ratios) and binomial family”
“odds ratio 2.42 (95% CI 1.45 to 4.04), P=0.001”
The data availability statement says data will be made available on reasonable request after review by SGF and other study team members, with a data use agreement required. This is vague and does not specify a platform, conditions, or timeframe. No code is shared. Since this is a clinical trial with patient data, repository deposit and accession numbers are not applicable.
“Research data will be made available after publication on reasonable request after review by SGF and other study team members. A data use agreement will be required before the release of data and the institutional review board’s approval as appropriate.”
“Research data will be made available after publication on reasonable request after review by SGF and other study team members. A data use agreement will be required before the release of data and the institutional review board’s approval as appropriate.”
The trial is registered with ClinicalTrials.gov (NCT04472195). Methods are detailed enough for replication. Limitations are explicitly discussed. Conclusions are proportional to the evidence. Funding and competing interests are disclosed. No reporting guideline is explicitly mentioned, but the paper follows CONSORT-like structure.
“Trial registration ClinicalTrials.gov NCT04472195”
“Our study had several limitations. First, just-in-time training could slow workflow.”
“Trial registration ClinicalTrials.gov NCT04472195”
“Our study had several limitations. First, just-in-time training could slow workflow.”
“Funding: Funded by the Anesthesia Research Distinguished Trailblazer Award, Department of Anesthesiology, Critical Care and Pain Medicine, Boston Children’s Hospital (297600-000-847).”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 47 references by DOI: 46 verified — 1 no DOI (shown, not verified).
- NO DOIPeak: Secrets From the New Science of ExpertiseNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORtypoMethods, Statistical analysis“Stata (version 16 0.0, StataCorp, College Station, TX)”→ Stata (version 16.0, StataCorp, College Station, TX)Missing decimal point in version number.
- MINORconsistencyResults, first paragraph“250 trainees were assessed for eligibility (), of whom 172 trainees were randomized”→ Add reference to CONSORT diagram figure.Missing figure reference.
- MINORclarityMethods, Statistical analysis“We did a number-needed-to-treat analysis at the number of intubations level.”→ Clarify that NNT is calculated per intubation, not per trainee.Could be ambiguous.
- MINORconsistencyResults, Secondary outcomes“The overall complication rate was 2.75% (7/255) in the treatment group and 4.71% (16/340) in the control group”→ Ensure denominators are consistent with the number of attempts (255 vs 340) and clarify if these are per-attempt rates.Denominators differ from the intubation counts (232 and 283) but are explained as attempts.
The published work is robust and well-reported, with only minor reporting gaps. An informed reader should weigh the vague data availability statement and the lack of a CONSORT checklist as minor limitations, but neither undermines the core findings. No erratum or re-analysis appears warranted based on the checks performed.
- 1.HIGHdata codeReplace the vague data availability statement with a concrete plan, e.g., deposit de-identified data in a repository like Dryad or Vivli, or specify a managed-access platform with conditions and a timeframe.The current statement lacks a concrete access mechanism or timeframe, which limits reproducibility and is a common reviewer concern.
- 2.HIGHdata codeShare analysis code in a public repository (e.g., GitHub) with a DOI to enhance reproducibility.No code is currently shared, which limits the ability of others to reproduce the analysis.
- 3.HIGHreportingAdd an explicit statement about adherence to CONSORT reporting guidelines, and provide the CONSORT checklist as supplementary material.The paper follows a CONSORT-like structure but does not explicitly mention the guideline, which is expected for a randomized trial.
- 4.MEDIUMreportingReport patient sex in Table 1 to fully characterize the study population.Sex is a basic demographic variable that is currently missing, and its absence is a minor reporting gap.
- 5.MEDIUMreportingClarify the manikin model and manufacturer used for coaching sessions to improve replicability.The manikin is not specified by brand/model, which could hinder replication of the intervention.
- 6.MEDIUMcopyeditFix the typo in the Stata version number in Methods, Statistical analysis: change '16 0.0' to '16.0'.The missing decimal point is a minor typo that could cause confusion.
- 7.MEDIUMcopyeditAdd a reference to the CONSORT flow diagram in the Results, first paragraph where it says '250 trainees were assessed for eligibility'.The missing figure reference makes it harder for readers to locate the flow diagram.
- 8.MEDIUMcopyeditClarify in Methods, Statistical analysis that the number-needed-to-treat analysis is calculated at the intubation level, not the trainee level.The current wording is ambiguous and could be misinterpreted.
- 9.MEDIUMcopyeditClarify in Results, Secondary outcomes that the complication rates (2.75% and 4.71%) are per-attempt rates, not per-intubation rates, and ensure denominators are consistent.The denominators (255 and 340) differ from the intubation counts (232 and 283), which could confuse readers.
- 10.MEDIUMreportingTemper the claim in the Discussion that just-in-time training could improve high-stakes procedures more broadly, as the study only tested infant intubation.The claim is overstated; generalizing to other procedures is speculative and not directly supported by the evidence.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.