Time restricted eating and exercise training before and during pregnancy for people with increased risk of gestational diabetes: single centre randomised controlled trial (BEFORE THE BEGINNING)
Sujan MJ, Skarstad HM, Rosvold G, Fougner SL, Follestad T, Salvesen KÅ, Moholdt T.
- DOI
- 10.1136/bmj-2024-083398
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/ec6a1561-4246-433e-a8b7-c439f6120c69 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×5−2.5★
- StatisticsPrinted percentage does not match its own count (capped)−0.25★
- ReportingEthical approvals partially met−0.25★
- ReportingStatistical analysis partially met−0.25★
- References were not verified against Crossref/OpenAlex.
- 01Printed percentage does not match its own count
86% does not match the reported count 146/166
“146/166 participants (86%)”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper presents a well-designed RCT with adequate data sharing and reporting of methods, but contains several numerical inconsistencies in the Discussion (incorrect percentages, mismatched counts) and missing ethics details (informed consent, regulatory compliance). These issues are minor individually but collectively warrant a correction or erratum.
Two independent runs of the same model were synthesised; they agreed on most dimensions, diverging mainly on the statistical analysis dimension (one saw a pass, the other a warn due to a numerical error). The copyedit and integrity components confirmed the numerical inconsistencies. The statistics component checked only a subset of reported tests; no claim of overall correctness is made.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 5 tests: 5 consistent, 0 inconsistent; 5 via agent-written checks. 1 printed percentage that does not match its own count.
- PERCENT86% does not match the reported count 146/166
“146/166 participants (86%)”
- CONSISTENTreported p = .080 · recomputed p = .076Reviewers 1, 2Primary outcome: two-hour plasma glucose difference at week 28
“mean difference 0.48 mmol/L, 95% confidence interval −0.05 to 1.01, P=0.08”
Taken as given: The 95% confidence interval is two-sided.; The estimate is a mean difference on an additive scale (not log).; The p-value is derived from a normal approximation of the estimate and CI.Method: Two-tailed p-value computed from estimate and 95% CI using normal approximation.How we recomputed it: pCI(0.48, -0.05, 1.01, 0) - CONSISTENTreported p = .002 · recomputed p = .002Reviewer 1Weight gain difference at week 28
“estimated mean weight gain in the intervention group at gestational week 28 was 2.0 kg lower (95% confidence interval −3.3 to −0.8, P=0.002)”
Taken as given: The 95% confidence interval is two-sided.; The estimate is a mean difference on an additive scale.; The p-value is from a normal approximation.Method: Two-tailed p-value from estimate and 95% CI.How we recomputed it: pCI(-2.0, -3.3, -0.8, 0) - CONSISTENTreported p = .008 · recomputed p = .005Reviewer 1Fat mass gain difference at week 28
“fat mass gain was 1.5 kg lower (−2.5 to −0.4, P=0.008)”
Taken as given: The 95% confidence interval is two-sided.; The estimate is a mean difference on an additive scale.; The p-value is from a normal approximation.Method: Two-tailed p-value from estimate and 95% CI.How we recomputed it: pCI(-1.5, -2.5, -0.4, 0) - CONSISTENTreported p = 1.000 · recomputed p = 1.000Reviewers 1, 2GDM prevalence at week 12 (Fisher's exact test)
“At gestational week 12, 3/51 participants (5.9%) in each group fulfilled the criteria for GDM diagnosis (P=1.00)”
Taken as given: The 3 and 51 are the event count and total in each group.; The test is Fisher's exact test (two-sided).; Non-events are 51-3=48 in each group.Method: Two-sided Fisher's exact test on 2x2 table.How we recomputed it: pFisher2x2(3,48,3,48,0) - CONSISTENTreported p = .570 · recomputed p = .486Reviewers 1, 2GDM prevalence at week 28 (chi-square test)
“The corresponding numbers at gestational week 28 were 8/49 participants (16.3%) in the intervention group and 6/52 participants (11.5%) in the control group (P=0.57; supplementary table 1)”
Taken as given: The 8 and 49 are the event count and total in the intervention group.; The 6 and 52 are the event count and total in the control group.; The test is Pearson chi-square (2x2).; Non-events are 49-8=41 and 52-6=46.Method: Pearson chi-square test on 2x2 table.How we recomputed it: pChi2x2(8,41,6,46)
- lowinternal contradictionDiscussion states '146/166 participants (86%) having a body mass index greater than 25', but Table 1 reason-for-inclusion BMI≥25 sums to 142 (74 control + 68 intervention) and 146/166 = 88%, not 86%.
“with 146/166 participants (86%) having a body mass index greater than 25”
DiscussionFind in source - lowinternal contradictionHOMA2-B control baseline mean/SD differs between Table 1 (168.2 (59.2)) and Table 2 (167.8 (58.9)); likely due to differing n (83 vs 166) but not explained.
“HOMA2-B | 168.2 (59.2) | 164.9 (69.2)”
Table 1Find in source - lowinternal contradictionDiscussion states '148/166 (89%) being white', but Table 1 Europe ethnic-origin counts sum to 146 (74 control + 72 intervention) = 88%.
“148/166 (89%) being white”
DiscussionFind in source - lowinternal contradictionThe fraction 146/166 is reported as 86%, but 146/166 = 87.9%, which would round to 88%.
“146/166 participants (86%) having a body mass index greater than 25”
DiscussionFind in source
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
8 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewer 2A prepregnancy intervention must induce substantial weight loss to impact glucose tolerance in pregnancy.This is an inference drawn from comparing Price et al. (9.2 kg weight loss, improved glucose) with the present trial (no weight loss, no effect); it is appropriately hedged ('it seems likely') but goes beyond the data presented, and the comparison rests on cross-study differences.Evidence: Discussion contrast with Price and colleagues' 12-week very low energy diet trial.
“Based on the findings from Price and colleagues and our study, it seems likely that a prepregnancy intervention must induce substantial weight loss to impact glucose tolerance in pregnancy.”
DiscussionFind in source - supportedReviewers 1, 2The intervention had no significant effect on two-hour plasma glucose level in an oral glucose tolerance test at gestational week 28.The primary outcome result (mean difference 0.48, 95% CI -0.05 to 1.01, P=0.08) directly supports this claim.Evidence: Primary outcome result in Results section.
“The intervention had no significant effect on two hour plasma glucose level in an oral glucose tolerance test at gestational week 28 (mean difference 0.48 mmol/L, 95% confidence interval −0.05 to 1.01, P=0.08).”
AbstractFind in source - supportedReviewer 1The intervention reduced weight and fat mass gain at gestational week 28.Secondary outcome analyses show significant reductions in weight gain (-2.0 kg, P=0.002) and fat mass gain (-1.5 kg, P=0.008).Evidence: Secondary outcomes in Results.
The estimated mean weight gain in the intervention group at gestational week 28 was 2.0 kg lower (95% confidence interval −3.3 to −0.8, P=0.002), and fat mass gain was 1.5 kg lower (−2.5 to −0.4, P=0.008) than the control group.
Resultsreviewer’s wording - supportedReviewer 1Participants adhered to the time-restricted eating intervention before pregnancy.Adherence data show that 41/83 (49%) adhered to a ≤10-hour eating window in prepregnancy, supporting the claim of good adherence.Evidence: Adherence section in Results.
“In the prepregnancy period, 41/83 participants (49%) in the intervention group adhered to a ≤10 hour eating window.”
ResultsFind in source - supportedReviewer 1The intervention had no significant effect on glycaemic control in late pregnancy.No significant between-group differences were found in any glycaemic outcome, consistent with this conclusion.Evidence: Primary and secondary glycaemic outcomes in Results.
“A combination of time restricted eating and exercise training started before and continued throughout pregnancy had no significant effect on glycaemic control in late pregnancy.”
ConclusionFind in source - supportedReviewer 2A combination of time-restricted eating and exercise started before and continued throughout pregnancy had no significant effect on glycaemic control in late pregnancy.The null primary result and the absence of significant improvement in secondary glycaemic outcomes (fasting glucose, insulin, HbA1c, HOMA indices) support this conclusion.Evidence: Primary outcome result and secondary glycaemic outcomes (Table 2).
“A combination of time restricted eating and exercise training started before and continued throughout pregnancy had no significant effect on glycaemic control in late pregnancy.”
ConclusionFind in source - supportedReviewer 2The intervention reduced body weight and fat mass gain at gestational week 28.The stated weight gain difference of 2.0 kg (P=0.002) and fat mass difference of 1.5 kg (P=0.008) directly support this claim.Evidence: Results: 'The estimated mean weight gain in the intervention group at gestational week 28 was 2.0 kg lower (95% confidence interval −3.3 to −0.8, P=0.002), and fat mass gain was 1.5 kg lower (−2.5 to −0.4, P=0.008)'.
“Although the intervention reduced body weight and fat mass gain at gestational week 28, there was no significant effect on GDM incidence.”
ResultsFind in source - supportedReviewer 2Participants were able to adhere to the ≤10 hour time-restricted eating intervention before pregnancy.Reported prepregnancy adherence (49% with ≤10 hour window; average window 9.9 h/day) supports feasibility of the dietary component before pregnancy.Evidence: Adherence results: 'In the prepregnancy period, the intervention group had an average eating window of 9.9 hours (standard deviation 1.2)' and '41/83 participants (49%) in the intervention group adhered to a ≤10 hour eating window'.
“The participants were able to adhere to the ≤10 hour time restricted eating intervention before pregnancy, with a slight increase in the time window of energy intake during pregnancy.”
ResultsFind in source
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Methods and results do not matchAssessed
5 integrity concerns flagged (0 high).
- lowmethod result mismatchThe week-28 GDM comparison (8/49 vs 6/52) is reported as a chi-square test with P=0.57, but the stated counts yield a chi-square p of approximately 0.49; the reported p may reflect a different test N or a correction.
“The corresponding numbers at gestational week 28 were 8/49 participants (16.3%) in the intervention group and 6/52 participants (11.5%) in the control group (P=0.57; supplementary table 1)”
ResultsFind in source
Reporting gaps
2 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Statistical reporting gaps (tests, assumptions, effect sizes)Assessed
- Ethics/consent reporting incompleteAssessed
The introduction cites conflicting prior evidence and highlights the prepregnancy window as an opportunity, directly motivating the trial. The hypothesis is explicitly stated. Limitations of prior work (late intervention start, scarce prepregnancy data) are addressed.
“Most clinical trials of lifestyle interventions in pregnancy have started the intervention around 16-20 weeks of gestation, leaving a missed window of opportunity to implement lifestyle changes and improve glycaemic control.”
“we hypothesised that combined time restricted eating and exercise training started before pregnancy and continued throughout pregnancy would improve maternal glucose tolerance by gestational week 28 in people at increased risk of GDM.”
“evidence on the specific components of prepregnancy interventions and their effectiveness remains scarce.”
“The Finnish Gestational Diabetes Prevention Study (RADIEL) showed that a diet and exercise intervention started before 20 weeks of gestation could reduce the incidence of GDM by 39% in those at high risk. In contrast, some large lifestyle intervention trials, such as the UPBEAT study and the LIMIT trial, did not show the effectiveness”
“However, evidence on the specific components of prepregnancy interventions and their effectiveness remains scarce.”
“we hypothesised that combined time restricted eating and exercise training started before pregnancy and continued throughout pregnancy would improve maternal glucose tolerance by gestational week 28 in people at increased risk of GDM.”
Randomisation used a computer random number generator (WebCRF3) with variable block sizes and stratification by prior GDM. The open-label design is explicitly stated. A priori sample size calculation is provided with effect size, SD, power, and alpha. Pre-specified inclusion/exclusion criteria and analysis population (ITT and per-protocol) are described.
“The study personnel used WebCRF3, a computer random number generator developed and administered at the Clinical Research Unit (Klinforsk, NTNU/St Olav’s Hospital, Trondheim, Norway) to randomly allocate participants using various block sizes.”
“Neither participants nor study personnel were masked.”
“Sample size calculation for a two sided t test to detect a 1.0 mmol/L difference between the groups, using a standard deviation of 1.3, a power of 0.90, and a significance level of 0.05, yields 37 participants in each group at gestational week 28.”
“Neither participants nor study personnel were masked.”
“Sample size calculation for a two sided t test to detect a 1.0 mmol/L difference between the groups, using a standard deviation of 1.3, a power of 0.90, and a significance level of 0.05, yields 37 participants in each group at gestational week 28.”
All participants are female (pregnancy trial), so sex is effectively reported. Age, weight, BMI, and other health measures are in Table 1. Demographics (ethnicity, education) are also reported. Single-sex justification is inherent to the pregnancy context, so sex_justified is not applicable.
“Age, years | 30.3 (3.2) | 30.2 (3.1) | | Weight, kg | 81.5 (13.2) | 81.1 (16)”
“Ethnic origin, n (%) | | Europe | 74 (89) | 72 (87)”
“We also sent electronic invitations to participate in the trial to all women aged 20-35 years in Trondheim and the surrounding area”
The ethics approval statement names the Regional Committees for Medical and Health Research Ethics in Norway and gives protocol number REK 143756. However, no informed consent is described, and no explicit regulatory compliance statement (e.g., Declaration of Helsinki) is present. These gaps are fixable reporting omissions.
“The Regional Committees for Medical and Health Research Ethics in Norway approved the study (REK 143756).”
“The Regional Committees for Medical and Health Research Ethics in Norway approved the study (REK 143756).”
The intervention is time-restricted eating and exercise, not a drug or device. The glucose drink and ELISA kits are measurement tools, not investigational products. Per the applicability guide, such trials are not_applicable for this dimension.
“The intervention consisted of exercise training and time restricted eating, started before pregnancy and continued throughout pregnancy.”
“The intervention consisted of time restricted eating and exercise training, and spanned from inclusion before pregnancy and throughout pregnancy.”
“Statistical analyses were performed using IBM SPSS Statistics 29.0 and STATA MP version 18.”
Tests are named (linear mixed models, Fisher's exact, chi-square, t-test), assumptions checked via QQ plots and bootstrap for non-normal residuals, software identified, and effect sizes with 95% CIs given throughout. The primary outcome p-value is consistent with its CI. The only issue is a percentage of 86% for 146/166, which should be 88%, a minor rounding discrepancy.
“mean difference 0.48 mmol/L, 95% confidence interval −0.05 to 1.01, P=0.08”
“146/166 participants (86%) having a body mass index greater than 25”
“We used linear mixed models to estimate differences in primary and secondary continuous outcomes between groups, with time and the interaction between time and group as fixed effects variables”
“We checked the normality of residuals by visually inspecting QQ plots. For variables that were not normally distributed, bias corrected and accelerated bootstrap confidence intervals based on 3000 bootstrap samples were calculated.”
“mean difference 0.48 mmol/L, 95% confidence interval −0.05 to 1.01, P=0.08”
The data availability statement specifies a public repository (Zenodo) with a DOI, and states that individual deidentified participant data and statistical codes are available. This satisfies all applicable criteria.
“All individual deidentified participant data and statistical codes are available on Zenodo data repository. The code used to analyse the data in the paper can be found in the supplementary files.”
“https://doi.org/10.5281/zenodo.15675472”
Trial registration is provided (NCT04585581). Methods are detailed. Limitations are discussed in a dedicated section. Funding and COI statements are present. However, no reporting guideline (e.g., CONSORT) is referenced, and the paper states that some secondary outcomes will be reported in separate publications.
“Trial registration ClinicalTrials.gov NCT04585581”
“Here we report the main secondary maternal cardiometabolic outcomes and will report the remaining secondary outcomes in separate publications.”
“Here we report the main secondary maternal cardiometabolic outcomes and will report the remaining secondary outcomes in separate publications.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
2 data/code links checked; 2 live.
- dataZenodoLIVEHTTP 200https://doi.org/10.5281/zenodo.15675472Resolves to Zenodo (data repository).
- datahttps://clinicaltrials.gov/ct2/show/NCT04585581LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, other.
- MINORconsistencyDiscussion, Comparisons with other studies“146/166 participants (86%)”→ 146/166 participants (88%)146/166 = 87.9%, which rounds to 88%, not 86%.
- MINORotherHeader/abstract“e083398 e083398”→ Remove the duplicated page/article identifier string.BMJ template artifact repeated at the top of the text.
- MINORconsistencyData Availability Statement“(10.3390/ijerph18094582)”→ Remove the unrelated DOI appended after the Zenodo DOI.An unrelated journal DOI is appended to the data availability statement; likely a copy-paste artifact.
- MINORconsistencyDiscussion“146/166 participants (86%) having a body mass index greater than 25”→ Correct to match Table 1 (142/166) or restate the percentage (88%).Numerator does not match the Table 1 sum and the stated percentage.
- MINORconsistencyDiscussion“148/166 (89%) being white”→ Correct to match Table 1 Europe sum (146/166 = 88%).Numerator does not match the Table 1 sum.
- MINORconsistencyTable 1 vs Table 2“HOMA2-B | 168.2 (59.2) | 164.9 (69.2)”→ Reconcile control baseline mean/SD with Table 2 (167.8 (58.9)) or note the differing n.Minor numeric difference between the two tables for the control group baseline.
In this post-publication audit, the paper is robust in design and data sharing, but the numerical inconsistencies in the Discussion (incorrect percentages, mismatched counts) and missing informed consent statement are concrete issues that an informed reader should weigh. The authors should consider issuing a correction or erratum for the numerical errors and adding the missing ethics details to the published version if possible.
- 1.HIGHstatisticsCorrect the percentage in Discussion: '146/166 participants (86%)' to 88% (146/166 = 87.9%).This is a demonstrable numerical error that could mislead readers.
- 2.HIGHstatisticsCorrect the statement '148/166 (89%) being white' to match Table 1 Europe sum (146/166 = 88%).Inconsistent counts between text and table undermine data integrity.
- 3.HIGHethicsAdd an informed consent statement to the Ethics statements section, detailing how consent was obtained (written/oral) or why it was waived.Missing informed consent is a reporting gap that could raise ethical concerns.
- 4.HIGHethicsAdd a regulatory compliance statement (e.g., 'conducted in accordance with the Declaration of Helsinki') to the Ethics statements section.The absence of a compliance framework is a standard reporting omission.
- 5.HIGHstatisticsReconcile the HOMA2-B control baseline mean/SD between Table 1 (168.2 (59.2)) and Table 2 (167.8 (58.9)) and explain the discrepancy.Minor numeric differences between tables may indicate data handling issues.
- 6.HIGHstatisticsVerify the chi-square test for the week-28 GDM comparison (8/49 vs 6/52, reported P=0.57) and correct if the p-value is inaccurate (recomputed ~0.49).An inconsistent p-value could affect the interpretation of a secondary outcome.
- 7.MEDIUMreportingReference the CONSORT reporting checklist in the Methods or a dedicated section.Adherence to reporting guidelines is expected for RCTs and improves transparency.
- 8.MEDIUMreportingProvide a clear statement of which secondary outcomes are deferred and where they will be reported, with reference to the registered protocol.Selective outcome reporting is a potential bias; transparency is needed.
- 9.MEDIUMotherAdd a brief rationale for the open-label design (e.g., infeasibility of masking a lifestyle intervention) to the Methods.Justifying the lack of blinding helps readers assess potential bias.
- 10.LOWcopyeditRemove the duplicated page/article identifier string from the header/abstract (BMJ template artifact).Minor formatting artifact that does not affect content but should be cleaned.
- 11.LOWcopyeditRemove the unrelated DOI (10.3390/ijerph18094582) appended after the Zenodo DOI in the Data Availability Statement.Copy-paste artifact that does not belong in the data availability statement.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.