Long acting progestogens versus combined oral contraceptive pill for preventing recurrence of endometriosis related pain: the PRE-EMPT pragmatic, parallel group, open label, randomised controlled trial.
Cooper KG, Bhattacharya S, Daniels JP, Horne AW, Clark TJ, Saridogan E, Cheed V, Pirie D, Melyda M, Monahan M, Roberts TE, Cox E, Stubbs C, Middleton LJ, PRE-EMPT Collaborative Group
- DOI
- 10.1136/bmj-2023-079006
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/034e0611-f28d-48c4-9729-eb0ce32f1f71 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 64 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is the pain domain of the EHP-30 questionnaire, a patient-reported outcome measure, which is a surrogate for clinical benefit. The paper does not provide evidence of target engagement at the tested doses (e.g., PK/PD data) nor does it cite validated evidence linking changes in EHP-30 pain scores to hard clinical outcomes such as reduced need for surgery or improved long-term function. Although the trial also reports a secondary outcome of treatment failure (further surgery or second-line treatment), the primary efficacy claim is based on the EHP-30 pain score, which is a surrogate.
“The primary outcome was pain measured three years after randomisation using the pain domain of the Endometriosis Health Profile 30 (EHP-30) questionnaire.”
- 02Treatment effect not shown to be clinically meaningful
The primary outcome shows a 40% improvement in pain scores from baseline in both groups, but the between-group difference is not statistically significant (adjusted mean difference −0.8, 95% CI −5.7 to 4.2, P=0.76). The paper does not anchor this effect size to a minimal clinically important difference (MCID) for the EHP-30 pain domain, nor does it provide evidence that a 40% improvement is clinically meaningful. The effect size is presented as a positive result without explicit clinical meaningfulness.
“At three years, there was no difference in pain scores between the groups (adjusted mean difference −0.8, 95% confidence interval −5.7 to 4.2, P=0.76), which had improved by around 40% in both groups compared with preoperative values (an average of 24 and 23…”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomised controlled trial with rigorous design, clear ethical approvals, and appropriate statistical methods. The main weaknesses are the vague data availability statement, lack of code sharing, and minor reporting inconsistencies such as the unstated statistical software and a date discrepancy.
Both reviewers independently scored all eight dimensions and agreed on every status, so no divergence needed reconciliation. The study type is interventional (RCT). Non-applicable sub-criteria (e.g., animal housing, cell lines) were excluded from scoring.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 2 tests: 2 consistent, 0 inconsistent; 2 via agent-written checks.
- CONSISTENTreported p = .760 · recomputed p = .751Reviewers 1, 2Primary outcome adjusted mean difference p-value
“adjusted mean difference −0.8, 95% CI −5.7 to 4.2; P=0.76”
Taken as given: The estimate is -0.8 and the 95% CI is -5.7 to 4.2.; The CI is two-sided at 95%.; The estimate is a mean difference (not a ratio).Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(-0.8, -5.7, 4.2, 0) - CONSISTENTreported p = .050 · recomputed p = .056Reviewer 2Hazard ratio for treatment failure p-value
“hazard ratio 0.67, 95% confidence interval 0.44 to 1.00”
Taken as given: The CI is a 95% confidence interval for the hazard ratio.; The p-value is two-sided.; The hazard ratio is on a log scale.Method: Recomputed p-value from the hazard ratio and 95% CI using the log-normal approximation.How we recomputed it: pCI(0.67, 0.44, 1.00, 1)
- lowinternal contradictionThe abstract states '73 v 97' for surgical procedures or second line treatments, while the results section states '73 v 97 events, occurring in 50 v 61 women'. The numbers are consistent but the phrasing could be clarified.
“Women randomised to a long acting progestogen underwent fewer surgical procedures or second line treatments compared with those randomised to the combined oral contraceptive pill group (73 v 97; hazard ratio 0.67, 95% confidence interval 0.44 to 1.00).”
AbstractFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
6 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2The reduced risk of repeat surgery might make LAPs preferable for some.The reduced risk is supported by the hazard ratio, but the CI includes 1.0, so the claim is somewhat cautious.Evidence: Hazard ratio 0.67, 95% CI 0.44 to 1.00.
“the reduced risk of repeat surgery for endometriosis and hysterectomy might make long acting reversible progestogens preferable for some”
ConclusionFind in source - supportedReviewers 1, 2There is no difference in pain scores between LAP and COCP at three years.The primary outcome analysis shows no statistically significant difference, with a small adjusted mean difference and wide CI including zero.Evidence: Adjusted mean difference −0.8, 95% CI −5.7 to 4.2, P=0.76.
“At three years, there was no difference in pain scores between the groups (adjusted mean difference −0.8, 95% confidence interval −5.7 to 4.2, P=0.76)”
AbstractFind in source - supportedReviewers 1, 2Both groups showed around 40% improvement in pain compared with preoperative levels.The paper reports average improvements of 24 and 23 points, which are approximately 40% of baseline scores.Evidence: Both groups showing a similar reduction of around 40% (on average, 24 points for LAP group and 23 points for COCP group) compared with preoperative values.
“with both groups showing a similar reduction of around 40% (on average, 24 points for LAP group and 23 points for COCP group) compared with preoperative values”
ResultsFind in source - supportedReviewer 1Women randomised to LAP underwent fewer surgical procedures or second line treatments compared with COCP.The hazard ratio of 0.67 with 95% CI 0.44 to 1.00 indicates a reduction, though the CI just touches 1.0, making it borderline.Evidence: 73 v 97; hazard ratio 0.67, 95% confidence interval 0.44 to 1.00.
“Women randomised to a long acting progestogen underwent fewer surgical procedures or second line treatments compared with those randomised to the combined oral contraceptive pill group (73 v 97; hazard ratio 0.67, 95% confidence interval 0.44 to 1.00)”
AbstractFind in source - supportedReviewer 1Postoperative prescription of LAP or COCP results in similar levels of improvement in endometriosis related pain at three years.The primary outcome shows no difference, and secondary outcomes are largely consistent with this.Evidence: Primary outcome result and secondary outcome analyses.
“Postoperative prescription of a long acting progestogen or the combined oral contraceptive pill results in similar levels of improvement in endometriosis related pain at three years”
ConclusionFind in source - supportedReviewer 2Women in the LAP group underwent fewer surgical procedures or second line treatments.The hazard ratio of 0.67 (95% CI 0.44 to 1.00) indicates a 33% reduction in time to treatment failure, supporting the claim, though the CI includes 1.00.Evidence: Results: 'Fewer women required additional treatment in the LAP group compared with the COCP group (73 v 97 events... hazard ratio 0.67, 95% CI 0.44 to 1.00)'
“Fewer women required additional treatment in the LAP group compared with the COCP group (73 v 97 events, occurring in 50 v 61 women because of several repeat interventions in some participants; supplementary table 14), translating to a 33% reduction in time to treatment failure (; hazard ratio 0.67, 95% CI 0.44 to 1.00).”
ResultsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is the pain domain of the EHP-30 questionnaire, a patient-reported outcome measure, which is a surrogate for clinical benefit. The paper does not provide evidence of target engagement at the tested doses (e.g., PK/PD data) nor does it cite validated evidence linking changes in EHP-30 pain scores to hard clinical outcomes such as reduced need for surgery or improved long-term function. Although the trial also reports a secondary outcome of treatment failure (further surgery or second-line treatment), the primary efficacy claim is based on the EHP-30 pain score, which is a surrogate.
“The primary outcome was pain measured three years after randomisation using the pain domain of the Endometriosis Health Profile 30 (EHP-30) questionnaire.”
- INADEQUATEEffect sizeThe primary outcome shows a 40% improvement in pain scores from baseline in both groups, but the between-group difference is not statistically significant (adjusted mean difference −0.8, 95% CI −5.7 to 4.2, P=0.76). The paper does not anchor this effect size to a minimal clinically important difference (MCID) for the EHP-30 pain domain, nor does it provide evidence that a 40% improvement is clinically meaningful. The effect size is presented as a positive result without explicit clinical meaningfulness.
“At three years, there was no difference in pain scores between the groups (adjusted mean difference −0.8, 95% confidence interval −5.7 to 4.2, P=0.76), which had improved by around 40% in both groups compared with preoperative values (an average of 24 and 23 points for long acting progestogen and combined oral contraceptive pill groups, respectively).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites population-based data on recurrence rates and guidelines recommending hormonal treatments, and clearly states the uncertainty about which regimen is better. The rationale for comparing LAPs and COCP is well articulated, and the trial aims to address this gap. Limitations of prior research are implicitly addressed by the trial's design, though not explicitly discussed in the introduction.
“It is unclear as to which of these two treatment regimens is better at preventing the recurrence of endometriosis related pain after surgical treatment.”
“Population based data from Scotland shows that, after initial surgery for endometriosis, 62% of treated women have at least one repeat operation, 45% have two or more, and nearly 25% need surgical removal of their ovaries, often combined with a hysterectomy.”
“It is unclear as to which of these two treatment regimens is better at preventing the recurrence of endometriosis related pain after surgical treatment.”
Randomization was performed using a central internet service with minimization variables, and the unit of randomization is the participant. Blinding was not feasible due to the nature of interventions, and this is explicitly stated. A priori sample size calculation is provided with effect size, alpha, and power. Inclusion/exclusion criteria are pre-specified. Outlier handling is addressed through the analysis population and missing data methods. Controls are inherent in the comparator arm. Independent replication is not applicable for a single pivotal trial.
“Randomisation occurred intraoperatively or immediately postoperatively using a central internet randomisation service provided by the Birmingham Clinical Trials Unit.”
“Participants and investigators were not blinded to treatment allocation owing to the substantial differences in formulations and their routes of delivery.”
“To detect an eight point difference on the EHP-30 pain domain with 90% power (P=0.05) and assuming a standard deviation of 22 points required 160 participants per group, 320 in total.”
“Randomisation occurred intraoperatively or immediately postoperatively using a central internet randomisation service provided by the Birmingham Clinical Trials Unit.”
“To detect an eight point difference on the EHP-30 pain domain with 90% power (P=0.05) and assuming a standard deviation of 22 points required 160 participants per group, 320 in total.”
“Participants and investigators were not blinded to treatment allocation owing to the substantial differences in formulations and their routes of delivery.”
Sex is reported (all female), age is reported with mean and SD, and demographics including ethnicity are reported. Since the study includes both sexes (all female), sex_justified is not applicable. Species/strain and housing conditions are not applicable for a human trial.
The trial received clinical trial authorisation from the MHRA and a favourable ethical opinion from the East of Scotland Ethics Committee, with protocol numbers. Written informed consent was obtained from all participants. Regulatory compliance is stated through the trial authorisation and ethical approval.
“The protocol (supplementary material 1) received clinical trial authorisation (EudraCT 2013-001984-21) from the Medicines and Healthcare products Regulatory Authority and a favourable ethical opinion from the East of Scotland Ethics Committee (14/ES1004).”
“all participants provided written informed consent.”
“The protocol (supplementary material 1) received clinical trial authorisation (EudraCT 2013-001984-21) from the Medicines and Healthcare products Regulatory Authority and a favourable ethical opinion from the East of Scotland Ethics Committee (14/ES1004).”
“all participants provided written informed consent.”
The investigational products (DMPA, LNG-IUS, COCP) are named with doses and regimens. No antibodies, cell lines, or mycoplasma testing are applicable. Software tools are not specifically identified, but the statistical analysis is described; however, the specific software is not named, which is a minor gap.
“In the LAP group, the options were DMPA, administered at a dose of 150 mg in an aqueous suspension by intramuscular injection every three months, or LNG-IUS that delivers a daily dose of 20 µg of levonorgestrel for five years.”
The primary analysis uses a mixed effects linear regression model, and secondary outcomes use appropriate regression models. Exact p-values are reported for the primary outcome (P=0.76). Effect sizes are reported with 95% CIs. Statistical software is not identified. Data presentation includes means, SDs, and CIs, and per-group n are stated. Mathematical plausibility checks were not performed due to continuous data and large N.
“For the primary outcome (EHP-30 pain scores at three years), a mixed effects linear regression model for repeated measures calculated the adjusted difference between group means, along with 95% confidence intervals (CIs).”
“adjusted mean difference −0.8, 95% CI −5.7 to 4.2; P=0.76”
“For the primary outcome (EHP-30 pain scores at three years), a mixed effects linear regression model for repeated measures calculated the adjusted difference between group means, along with 95% confidence intervals (CIs).”
“adjusted mean difference −0.8, 95% CI −5.7 to 4.2; P=0.76”
The data availability statement provides an email address for data requests, which is a concrete route but lacks details on conditions and timeframe. No repository deposit or accession numbers are provided, and no code sharing is mentioned. For a clinical trial, managed access is acceptable, but the statement is somewhat vague.
“All data requests should be submitted to bctudatashare@contacts.bham.ac.uk for consideration. Access to anonymised data might be granted after review.”
“All data requests should be submitted to bctudatashare@contacts.bham.ac.uk for consideration. Access to anonymised data might be granted after review.”
The trial is registered with ISRCTN, and the CONSORT checklist is referenced. All outcomes are reported, including negative results. Limitations are discussed in detail. Conclusions are proportional to the evidence. Funding and competing interests are declared.
“Trial registration ISRCTN registry ISRCTN97865475.”
“We used the CONSORT (Consolidated Standards of Reporting Trials) checklist when writing this report.”
“The predominance of white women in the recruited sample limits our confidence about extrapolating the results to women from other ethnic groups.”
“Trial registration ISRCTN registry ISRCTN97865475.”
“We used the CONSORT (Consolidated Standards of Reporting Trials) checklist when writing this report.”
“The three year follow-up period and the pragmatic design meant that relatively few women continued on their initially allocated drug, changing or stopping their treatments depending on their circumstances, including changes in reproductive plans.”
Registered (2 IDs: ISRCTN, EudraCT). Reporting guideline cited: CONSORT.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 25 references by DOI: 23 verified — 2 no DOI (shown, not verified).
- NO DOIApplied mixed models in medicineNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILaparoscopic surgery for endometriosisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
4 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 4 minor suggestions below.
4 copyedit issues flagged: mostly consistency, clarity.
- MINORconsistencyAbstract, Results“73 v 97”→ Use 'vs' instead of 'v' for consistency.Minor style inconsistency.
- MINORclarityMethods, Statistical analysis“All participants recruited from 23 October 2015 were included in the final analysis population, along with 92 from the internal pilot phase who were randomised to combinations of treatments that only included LAPs and COCP.”→ Clarify the date discrepancy: recruitment start date is stated as 23 November 2015 elsewhere.Potential inconsistency in dates.
- MINORconsistencyAbstract, Results“73 v 97; hazard ratio 0.67, 95% confidence interval 0.44 to 1.00”→ Ensure consistent use of 'v' vs 'vs' throughout.Minor stylistic inconsistency.
- MINORclarityMethods, Statistical analysis“All participants recruited from 23 October 2015 were included in the final analysis population, along with 92 from the internal pilot phase who were randomised to combinations of treatments that only included LAPs and COCP.”→ Clarify the exact date and inclusion criteria for pilot phase participants.Potential ambiguity in the date and inclusion criteria.
The published work is robust and generally trustworthy, but an informed reader should weigh the vague data availability statement and the lack of code sharing as minor reproducibility limitations. The copyedit-flagged date discrepancy (23 October vs 23 November 2015) and the 'v' vs 'vs' inconsistency are minor and do not undermine the conclusions, but could warrant a correction if confirmed.
- 1.HIGHdata codeExpand the Data Availability Statement to specify the review process, conditions, and timeframe for data access, and consider depositing anonymised data in a recognised repository (e.g., Zenodo, Dryad).The current statement is vague and lacks details, which limits reproducibility and transparency.
- 2.HIGHdata codeAdd a statement about code sharing or provide analysis code in a public repository.No code is shared, which hampers independent verification of the analyses.
- 3.HIGHreportingIdentify the statistical software and version used for analyses in the Methods section.Both reviewers noted that the software is not identified, which is a reproducibility gap.
- 4.MEDIUMcopyeditClarify the recruitment start date discrepancy: the Methods state '23 October 2015' while elsewhere it is '23 November 2015'.The copyedit pass flagged this potential inconsistency, which could confuse readers.
- 5.MEDIUMcopyeditStandardise the use of 'v' vs 'vs' in the Abstract and Results sections.Minor stylistic inconsistency that should be corrected for professional presentation.
- 6.MEDIUMreportingExplicitly discuss limitations of prior research in the introduction to strengthen the scientific premise.Both reviewers noted this gap, and addressing it would improve the framing of the trial's rationale.
- 7.MEDIUMstatisticsProvide more detail on how missing data were handled for secondary outcomes, beyond the sensitivity analysis for the primary outcome.Reviewer 2 suggested this to improve transparency of the statistical methods.
- 8.MEDIUMstatisticsClarify whether any adjustment for multiple testing was made for secondary outcomes.Reviewer 2 raised this as a potential concern for interpretation of secondary analyses.
- 9.LOWreportingInclude a statement on whether the data monitoring committee charter or statistical analysis plan is publicly available.Reviewer 2 suggested this to enhance transparency of trial governance.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.