Time restricted eating and exercise training before and during pregnancy for people with increased risk of gestational diabetes: single centre randomised controlled trial (BEFORE THE BEGINNING).
Sujan MJ, Skarstad HM, Rosvold G, Fougner SL, Follestad T, Salvesen KÅ, Moholdt T
- DOI
- 10.1136/bmj-2024-083398
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/77d225f2-682b-4d73-a501-d6b566cf8f91 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- StatisticsPrinted percentage does not match its own count (capped)−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 12 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Printed percentage does not match its own count
86% does not match the reported count 146/166
“146/166 participants (86%)”
DiscussionFind in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomised controlled trial of a prepregnancy lifestyle intervention. The main methodological strengths are the rigorous randomisation, pre-specified sample size, detailed statistical methods, and open data/code availability. The primary weakness is the lack of an explicit informed consent statement, and there are minor reporting and copyedit issues.
Both reviewers classified the study as interventional (RCT), so no divergence on study type. The statistics verification covered only 5 tests (those with test statistics/df or effect estimates with CIs); the remaining analyses (e.g., bootstrap, exact tests) were not machine-verified and should not be assumed correct. The citation check found no retracted or unresolved references. The reproducibility check confirmed the Zenodo link is live.
Numerical inconsistencies
2 findings · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 4 tests: 4 consistent, 0 inconsistent; 4 via agent-written checks. 1 printed percentage that does not match its own count.
- PERCENT86% does not match the reported count 146/166
“146/166 participants (86%)”
DiscussionFind in source
- CONSISTENTreported p = .080 · recomputed p = .076Reviewer 1Primary outcome p-value from linear mixed model (two-hour glucose at week 28).
“mean difference 0.48 mmol/L, 95% confidence interval −0.05 to 1.01, P=0.08”
Taken as given: The estimate is the mean difference (0.48) and the CI is 95% two-sided.; The CI is symmetric on the linear scale (log=0).Method: Recomputed p-value from the reported estimate and 95% CI using the normal approximation.How we recomputed it: pCI(0.48, -0.05, 1.01, 0) - CONSISTENTreported p = 1.000 · recomputed p = 1.000Reviewers 1, 2GDM prevalence at week 12 (Fisher's exact test).
“At gestational week 12, 3/51 participants (5.9%) in each group fulfilled the criteria for GDM diagnosis (P=1.00).”
Taken as given: The numbers 3 and 51 are the event count and group total for the intervention group.; The numbers 3 and 51 are the event count and group total for the control group.; The test is two-sided Fisher's exact test.Method: Recomputed Fisher's exact test from the 2x2 table (3,48,3,48).How we recomputed it: pFisher2x2(3, 48, 3, 48, 0) - CONSISTENTreported p = .570 · recomputed p = .486Reviewers 1, 2GDM prevalence at week 28 (chi-square test).
“The corresponding numbers at gestational week 28 were 8/49 participants (16.3%) in the intervention group and 6/52 participants (11.5%) in the control group (P=0.57; supplementary table 1).”
Taken as given: The numbers 8 and 49 are the event count and group total for the intervention group.; The numbers 6 and 52 are the event count and group total for the control group.; The test is Pearson's chi-square test.Method: Recomputed chi-square test from the 2x2 table (8,41,6,46).How we recomputed it: pChi2x2(8, 41, 6, 46) - CONSISTENTreported p = .080 · recomputed p = .076Reviewer 2Primary outcome: two-hour plasma glucose difference at week 28
“mean difference 0.48 mmol/L, 95% confidence interval −0.05 to 1.01, P=0.08”
Taken as given: The estimate is a mean difference from a linear mixed model.; The 95% CI is two-sided.; The p-value is for the test of the difference being zero.Method: Recomputed p-value from the reported estimate and 95% CI using the normal approximation.How we recomputed it: pCI(0.48, -0.05, 1.01, 0)
- lowinternal contradictionThe abstract states '31/83 participants (37%) in the intervention group adhered to prespecified criteria' but the per-protocol analysis section states 'Thirty one of 83 participants (37%) satisfied both criteria'. This is consistent, but the abstract also says '24/55 participants (44%) in the intervention group who became pregnant fulfilled these criteria' which matches the per-protocol analysis. No contradiction.
“31/83 participants (37%) in the intervention group adhered to prespecified criteria, whereas 24/55 participants (44%) in the intervention group who became pregnant fulfilled these criteria.”
AbstractFind in source - lowinternal contradictionThe paper reports '167 participants were enrolled' and '111 became pregnant', but the flowchart and text indicate that 166 were included in ITT analyses. The exclusion of one participant with prepregnancy diabetes explains the difference. This is correctly reported.
“We excluded data from one participant in the intervention group because of prepregnancy diabetes, so that data from 166 participants were included in the intention-to-treat analyses”
ResultsFind in source - lowinternal contradictionThe percentage 146/166 is stated as 86%, but 146/166 = 88%. This is a minor arithmetic inconsistency.
“with 146/166 participants (86%) having a body mass index greater than 25”
DiscussionFind in source
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
8 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2The intervention had no significant effect on two-hour plasma glucose at gestational week 28.The primary outcome analysis shows a mean difference of 0.48 mmol/L with 95% CI crossing zero and P=0.08, supporting the claim of no significant effect.Evidence: Primary outcome result: mean difference 0.48 mmol/L, 95% CI −0.05 to 1.01, P=0.08.
“The intervention had no significant effect on two hour plasma glucose level in an oral glucose tolerance test at gestational week 28 (mean difference 0.48 mmol/L, 95% confidence interval −0.05 to 1.01, P=0.08).”
AbstractFind in source - supportedReviewers 1, 2The intervention reduced weight and fat mass gain at gestational week 28.The estimated mean weight gain was 2.0 kg lower and fat mass gain 1.5 kg lower in the intervention group, with CIs excluding zero and P<0.01, supporting the claim.Evidence: Secondary outcomes: weight gain difference −2.0 kg (95% CI −3.3 to −0.8, P=0.002); fat mass gain −1.5 kg (95% CI −2.5 to −0.4, P=0.008).
The estimated mean weight gain in the intervention group at gestational week 28 was 2.0 kg lower (95% confidence interval −3.3 to −0.8, P=0.002), and fat mass gain was 1.5 kg lower (−2.5 to −0.4, P=0.008) than the control group.
Resultsreviewer’s wording - supportedReviewer 1The intervention did not significantly affect GDM incidence.GDM rates at week 28 were 16.3% vs 11.5% with P=0.57, indicating no significant difference.Evidence: GDM at week 28: 8/49 (16.3%) vs 6/52 (11.5%), P=0.57.
“The corresponding numbers at gestational week 28 were 8/49 participants (16.3%) in the intervention group and 6/52 participants (11.5%) in the control group (P=0.57; supplementary table 1).”
ResultsFind in source - supportedReviewer 1Participants adhered to time-restricted eating before pregnancy.The average eating window was 9.9 hours/day and 49% adhered to ≤10 hours, supporting the claim of good adherence.Evidence: Adherence data: average eating window 9.9 hours (SD 1.2); 41/83 (49%) adhered to ≤10 hours.
In the prepregnancy period, the intervention group had an average eating window of 9.9 hours (standard deviation 1.2) ... In the prepregnancy period, 41/83 participants (49%) in the intervention group adhered to a ≤10 hour eating window.
Resultsreviewer’s wording - supportedReviewer 1The intervention did not significantly affect time to pregnancy.Time to pregnancy was 112 vs 84 days with P=0.10, indicating no significant difference in the ITT analysis.Evidence: Time to pregnancy: mean 112 days (SD 105) vs 84 days (69), P=0.10.
“Time to pregnancy did not differ significantly between groups, with a mean of 112 days (standard deviation 105) in the intervention group and 84 days (69) in the control group (P=0.10).”
ResultsFind in source - supportedReviewer 2The intervention did not significantly improve secondary glycaemic outcomes.All secondary glycaemic outcomes (fasting glucose, insulin, HbA1c, HOMA2-IR, HOMA2-B) showed no statistically significant between-group differences.Evidence: Table 2 shows non-significant p-values for all glycaemic outcomes at all time points.
“The intervention did not significantly improve secondary glycaemic outcomes (fasting glucose, fasting insulin, HbA1c, HOMA2-IR, and HOMA2-B) compared with the control group.”
ResultsFind in source - supportedReviewer 2The intervention was feasible and adherence was acceptable before pregnancy but declined during pregnancy.Adherence data show good prepregnancy adherence (e.g., 49% adhered to eating window) but declining during pregnancy (e.g., PAI points decreased).Evidence: Adherence section: prepregnancy eating window 9.9 hours, 41/83 (49%) adhered; PAI points decreased during pregnancy.
“In the prepregnancy period, 41/83 participants (49%) in the intervention group adhered to a ≤10 hour eating window. The proportion of pregnant participants who adhered to a ≤10 hour eating window was 23/55 (42%) in the first trimester, 17/55 (31%) in the second trimester, and 21/55 (38%) in the third trimester of pregnancy.”
ResultsFind in source - supportedReviewer 2The per-protocol analysis showed longer time to pregnancy in the intervention group.The per-protocol analysis reported a significantly longer time to pregnancy (48 days, 95% CI 15 to 81, P=0.005), but this was not seen in the intention-to-treat analysis.Evidence: Per protocol analyses: 'time to pregnancy which was significantly longer in the intervention group (48 days, 95% confidence interval 15 to 81, P=0.005)'.
“The results from the per protocol analyses were similar to those from the intention-to-treat analyses (supplementary tables 6-9), except for time to pregnancy which was significantly longer in the intervention group (48 days, 95% confidence interval 15 to 81, P=0.005; supplementary table 7).”
ResultsFind in source
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites multiple relevant studies (RADIEL, UPBEAT, LIMIT, Prepare) and systematic reviews, and it explicitly notes the lack of evidence on specific prepregnancy intervention components. The hypothesis follows logically from the cited evidence that prepregnancy lifestyle changes may improve glycaemic control. The paper also addresses limitations of prior research by highlighting the missed window of opportunity in pregnancy and the need for prepregnancy interventions.
“The Finnish Gestational Diabetes Prevention Study (RADIEL) showed that a diet and exercise intervention started before 20 weeks of gestation could reduce the incidence of GDM by 39% in those at high risk. In contrast, some large lifestyle intervention trials, such as the UPBEAT study and the LIMIT trial, did not show the effectiveness of a diet and exercise intervention in pregnancy for GDM prevention.”
“However, evidence on the specific components of prepregnancy interventions and their effectiveness remains scarce.”
“In the BEFORE THE BEGINNING trial, we hypothesised that combined time restricted eating and exercise training started before pregnancy and continued throughout pregnancy would improve maternal glucose tolerance by gestational week 28 in people at increased risk of GDM.”
“The Finnish Gestational Diabetes Prevention Study (RADIEL) showed that a diet and exercise intervention started before 20 weeks of gestation could reduce the incidence of GDM by 39% in those at high risk. In contrast, some large lifestyle intervention trials, such as the UPBEAT study and the LIMIT trial, did not show the effectiveness of a diet and exercise intervention in pregnancy for GDM prevention.”
“In the BEFORE THE BEGINNING trial, we hypothesised that combined time restricted eating and exercise training started before pregnancy and continued throughout pregnancy would improve maternal glucose tolerance by gestational week 28 in people at increased risk of GDM.”
“Most clinical trials of lifestyle interventions in pregnancy have started the intervention around 16-20 weeks of gestation, leaving a missed window of opportunity to implement lifestyle changes and improve glycaemic control.”
Randomisation was computer-generated with concealed allocation and stratification by previous GDM. Blinding was not performed, but the paper explicitly states 'Neither participants nor study personnel were masked' and provides a rationale for an open-label design. A sample size calculation was performed and revised during the trial. Inclusion/exclusion criteria were pre-specified and modifications were documented. The analysis population (ITT) and missing-data handling are described. Controls (standard care) are appropriate. Independent replication is not applicable for a single pivotal trial.
“The study personnel used WebCRF3, a computer random number generator developed and administered at the Clinical Research Unit (Klinforsk, NTNU/St Olav’s Hospital, Trondheim, Norway) to randomly allocate participants using various block sizes.”
“Neither participants nor study personnel were masked.”
“Sample size calculation for a two sided t test to detect a 1.0 mmol/L difference between the groups, using a standard deviation of 1.3, a power of 0.90, and a significance level of 0.05, yields 37 participants in each group at gestational week 28.”
“The study personnel used WebCRF3, a computer random number generator developed and administered at the Clinical Research Unit (Klinforsk, NTNU/St Olav’s Hospital, Trondheim, Norway) to randomly allocate participants using various block sizes.”
“Neither participants nor study personnel were masked.”
“Sample size calculation for a two sided t test to detect a 1.0 mmol/L difference between the groups, using a standard deviation of 1.3, a power of 0.90, and a significance level of 0.05, yields 37 participants in each group at gestational week 28.”
The study enrolled only women (sex is implied by the population and confirmed by the baseline table). Age, weight, BMI, and health status are reported in Table 1. Demographics including education, ethnicity, and parity are reported. Since the study is human, species/strain and housing conditions are not applicable. Sex justification is not applicable because the study is in pregnant people (both sexes not applicable).
“Age, years | 30.3 (3.2) | 30.2 (3.1) | | Weight, kg | 81.5 (13.2) | 81.1 (16) | | Body mass index | 29.1 (4.5) | 29.2 (4.9)”
“Education level, n (%) | | Compulsory schooling | 0 (0) | 1 (1) | | Completed upper secondary school | 8 (10) | 10 (12) | | Completed university education, <4 years | 28 (34) | 23 (28) | | Completed university education, ≥4 years | 46 (56) | 47 (58)”
“Age, years | 30.3 (3.2) | 30.2 (3.1) | | Weight, kg | 81.5 (13.2) | 81.1 (16) | | Body mass index | 29.1 (4.5) | 29.2 (4.9)”
“Education level, n (%) | | Compulsory schooling | 0 (0) | 1 (1) | | Completed upper secondary school | 8 (10) | 10 (12) | | Completed university education, <4 years | 28 (34) | 23 (28) | | Completed university education, ≥4 years | 46 (56) | 47 (58)”
The paper states 'The Regional Committees for Medical and Health Research Ethics in Norway approved the study (REK 143756)', which is a named ethics body with a protocol number. Informed consent is not explicitly described in the text, but it is standard for such trials and the paper mentions 'informed consent' in the context of the trial registration? Actually, the paper does not explicitly state that informed consent was obtained. However, the ethics approval implies consent was obtained. The paper also states compliance with the Declaration of Helsinki? It does not explicitly mention it, but the ethics approval is sufficient. Regulatory compliance is implied by the ethics approval.
“The Regional Committees for Medical and Health Research Ethics in Norway approved the study (REK 143756).”
“The Regional Committees for Medical and Health Research Ethics in Norway approved the study (REK 143756).”
The intervention components (time-restricted eating and exercise) are described in detail, including the PAI metric and smartwatches used. The glucose drink (Glucosepro) is identified with manufacturer. Statistical software (IBM SPSS Statistics 29.0 and STATA MP version 18) is identified. Antibodies, cell lines, and mycoplasma testing are not applicable as this is a clinical trial without wet-lab assays.
“participants consumed a premade drink of 75 g glucose diluted in 250 mL water (Glucosepro, Finnamedical, Finland)”
“Statistical analyses were performed using IBM SPSS Statistics 29.0 and STATA MP version 18.”
“We counselled the participants in the intervention group to restrict their daily time window of energy intake to ≤10 hours, ending no later than 7 pm, for a minimum of five days per week throughout the study period.”
“Statistical analyses were performed using IBM SPSS Statistics 29.0 and STATA MP version 18.”
The paper names the statistical tests (linear mixed models, Fisher's exact, chi-square, t-test) and describes how assumptions were checked (QQ plots, bootstrap for non-normal). Exact p-values are reported (e.g., P=0.08). Effect sizes with 95% CIs are reported throughout. Software is identified. Data presentation includes per-group n and error bars defined. Mathematical plausibility checks were not performed due to continuous data and large N.
“We used linear mixed models to estimate differences in primary and secondary continuous outcomes between groups, with time and the interaction between time and group as fixed effects variables, and subject (participant ID) as random effect.”
“mean difference 0.48 mmol/L, 95% confidence interval −0.05 to 1.01, P=0.08”
“We checked the normality of residuals by visually inspecting QQ plots. For variables that were not normally distributed, bias corrected and accelerated bootstrap confidence intervals based on 3000 bootstrap samples were calculated.”
“We used linear mixed models to estimate differences in primary and secondary continuous outcomes between groups, with time and the interaction between time and group as fixed effects variables, and subject (participant ID) as random effect.”
“mean difference 0.48 mmol/L, 95% confidence interval −0.05 to 1.01, P=0.08”
“0.48 (−0.05 to 1.01) | 0.08”
The data availability statement names a specific repository (Zenodo) with a DOI, and states that individual deidentified participant data and statistical codes are available. This is reported_and_adequate. Repository deposit and accession numbers are covered by the Zenodo DOI. Code sharing is also covered by the statement that statistical codes are available on Zenodo.
“All individual deidentified participant data and statistical codes are available on Zenodo data repository. The code used to analyse the data in the paper can be found in the supplementary files. The data underlying the findings in this paper are openly and publicly available and can be found here: https://doi.org/10.5281/zenodo.15675472”
The trial is registered at ClinicalTrials.gov (NCT04585581). Methods are detailed enough for replication. A reporting guideline (CONSORT) is not explicitly mentioned, but the paper follows a structured format typical of RCT reports. All pre-specified outcomes are not fully reported (some secondary outcomes deferred to separate publications), but the main ones are. Limitations are thoroughly discussed. Conclusions are proportional to the findings. Funding and COI are disclosed.
“Trial registration ClinicalTrials.gov NCT04585581”
“Most participants were well educated, potentially resulting in healthy volunteer bias as they were likely knowledgeable about their health and wellbeing.”
“The trial was funded by the Novo Nordisk Foundation (NNF19SA058975), the Liaison Committee for education, research, and innovation in Central Norway, and the Joint Research Committee between St Olav’s Hospital and the Faculty of Medicine and Health Sciences, NTNU (FFU).”
“Trial registration ClinicalTrials.gov NCT04585581 (https://clinicaltrials.gov/ct2/show/NCT04585581)”
“Most participants were well educated, potentially resulting in healthy volunteer bias as they were likely knowledgeable about their health and wellbeing.”
“The trial was funded by the Novo Nordisk Foundation (NNF19SA058975), the Liaison Committee for education, research, and innovation in Central Norway, and the Joint Research Committee between St Olav’s Hospital and the Faculty of Medicine and Health Sciences, NTNU (FFU).”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 65 references by DOI: 62 verified — 3 no DOI (shown, not verified).
- NO DOIIDF Diabetes Atlas 2021 10th editionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIManagement of Diabetes in PregnancyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDiagnostic criteria and classification of hyperglycaemia first detected in pregnancyNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- dataZenodoLIVEHTTP 200https://doi.org/10.5281/zenodo.15675472Resolves to Zenodo (data repository).
Copyediting
7 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 7 minor suggestions below.
7 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORconsistencyAbstract, Results“31/83 participants (37%) in the intervention group adhered to prespecified criteria, whereas 24/55 participants (44%) in the intervention group who became pregnant fulfilled these criteria.”→ Clarify that the first number refers to all intervention participants and the second to those who became pregnant.The sentence is slightly confusing but not incorrect.
- MINORtypoMethods, Statistical analysis“bias corrected and accelerated confidence intervals”→ Use 'bias-corrected and accelerated' with hyphens.Hyphenation is missing.
- MINORclarityResults, Participants“We ended inclusion of new participants when 47 participants in each group reached gestational week 12, according to our sample size calculations.”→ Consider rephrasing for clarity: 'We stopped recruiting when 47 participants in each group had reached gestational week 12.'The sentence is understandable but could be clearer.
- MINORconsistencyAbstract, Results“31/83 participants (37%) in the intervention group adhered to prespecified criteria, whereas 24/55 participants (44%) in the intervention group who became pregnant fulfilled these criteria.”→ Clarify that the first number refers to all intervention participants, and the second to those who became pregnant.The sentence is slightly confusing but not incorrect.
- MINORclarityMethods, Statistical analysis“We checked the normality of residuals by visually inspecting QQ plots.”→ Specify which residuals (e.g., from the linear mixed models) were checked.Minor clarity improvement.
- MINORtypoTable 2, HbA1c row“32.0 (3.19)”→ Use consistent decimal places (e.g., 3.2) for SDs.Inconsistent decimal places in table.
- MINORconsistencyDiscussion, Comparisons with other studies“146/166 participants (86%) having a body mass index greater than 25”→ Verify the denominator: 146/166 is 88%, not 86%.Potential arithmetic inconsistency in percentage.
The published work is robust and generally trustworthy, but an informed reader should weigh the minor reporting gaps: the absence of an explicit informed consent statement, the lack of an explicit CONSORT adherence statement, and a minor arithmetic inconsistency (146/166 reported as 86% instead of 88%). These do not undermine the main conclusions but would warrant a correction or clarification from the authors.
- 1.HIGHethicsAdd an explicit statement in the Ethics statements section describing how informed consent was obtained (e.g., written informed consent from all participants) or provide a waiver justification.Informed consent is a required element for human research; its absence is a reporting gap that both reviewers flagged and that an informed reader would expect to see.
- 2.HIGHreportingCorrect the arithmetic inconsistency in the Discussion: 146/166 is 88%, not 86%.An incorrect percentage is a factual error that could mislead readers and warrants a correction.
- 3.MEDIUMreportingAdd an explicit statement of adherence to the CONSORT reporting guideline in the Methods or a dedicated section, and indicate where the checklist is available.Explicitly naming the reporting guideline improves transparency and is expected for RCT reports.
- 4.MEDIUMreportingClarify in the Results or Discussion that some pre-specified secondary outcomes are deferred to separate publications, and consider providing a summary of all outcomes in the supplementary material.Readers should know which outcomes are reported here and where to find the rest, to avoid selective reporting concerns.
- 5.MEDIUMstatisticsIn the Discussion, explicitly address the potential for type II error due to the reduced sample size (167 vs. planned 260) and its impact on the null result.The trial was underpowered relative to the original plan; acknowledging this helps readers interpret the null findings appropriately.
- 6.MEDIUMreportingAdd a note on the generalizability of the findings given the predominantly white, well-educated sample.The sample is not representative of all at-risk populations; this limitation should be explicitly stated.
- 7.MEDIUMdata codeSpecify the license and any access conditions for the Zenodo data and code repository.Clear licensing terms are important for reuse and reproducibility.
- 8.LOWcopyeditClarify the abstract sentence about adherence: '31/83 participants (37%) in the intervention group adhered to prespecified criteria, whereas 24/55 participants (44%) in the intervention group who became pregnant fulfilled these criteria.'The sentence is confusing because it is unclear that the first number refers to all intervention participants and the second to those who became pregnant.
- 9.LOWcopyeditUse consistent hyphenation in 'bias-corrected and accelerated' in the Methods, Statistical analysis section.Minor typographical consistency improves readability.
- 10.LOWcopyeditRephrase the sentence in Results, Participants: 'We ended inclusion of new participants when 47 participants in each group reached gestational week 12' to 'We stopped recruiting when 47 participants in each group had reached gestational week 12.'The original phrasing is understandable but could be clearer.
- 11.LOWcopyeditSpecify which residuals (e.g., from the linear mixed models) were checked for normality in the Methods, Statistical analysis section.Clarifies the statistical methods for readers.
- 12.LOWcopyeditUse consistent decimal places for standard deviations in Table 2 (e.g., change '32.0 (3.19)' to '32.0 (3.2)').Consistent formatting in tables improves clarity.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.