Impact of population based breast density notification: multisite parallel arm randomised controlled trial in BreastScreen.
Nickel B, Ormiston-Smith N, Cvejic E, Isautier J, Hammerton L, Baker K, Legerton P, Vardon P, McInally Z, Robertson S, McCaffery K, Houssami N
- DOI
- 10.1136/bmj-2024-083649
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/7743b7bf-e1b5-44c2-9b46-e1d9963aaa32 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ReportingData & code availability partially met−0.25★
- No data or code availability links were detected to verify.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomized controlled trial with rigorous design, clear ethical approvals, and sound statistical analysis. The main weakness is the vague data availability statement and lack of a persistent repository for code/data, along with minor reporting gaps (statistical software not named, CONSORT not explicitly referenced).
Both reviewers classified the study as interventional (RCT) and agreed on all dimension statuses. The statistics verification recomputed 11 tests, all consistent; however, this covers only a subset of reported analyses, so the paper's statistics should not be considered fully verified. The citation check found no retracted or non-existent references.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 11 tests: 11 consistent, 0 inconsistent; 1 recomputed directly from the reported test statistics, 10 via agent-written checks.
- CONSISTENTreported p < .001 · recomputed p = <.001Recomputed odds ratio 1.36 (95% CI 1.14–1.63), reported p<0.001
“odds ratio 1.36, 95% CI 1.14 to 1.63; P<0.001”
Taken as given: 1.14–1.63 is a two-sided 95% confidence interval for the odds ratio of 1.36, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p<0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.36, 1.14, 1.63, 1) - CONSISTENTreported p = .005 · recomputed p = .006Reviewers 1, 2Recompute p-value for intervention 1 vs control on feeling anxious from reported OR and 95% CI.
“intervention 1: odds ratio 1.30, 95% confidence interval (CI) 1.08 to 1.57 (P=0.005)”
Taken as given: The odds ratio is 1.30 with 95% CI 1.08 to 1.57.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Compute p-value from the reported odds ratio and its 95% confidence interval using the normal approximation for the log odds ratio.How we recomputed it: pCI(1.30, 1.08, 1.57, 1) - CONSISTENTreported p = .007 · recomputed p = .008Reviewers 1, 2Recompute p-value for intervention 2 vs control on feeling anxious from reported OR and 95% CI.
“intervention 2: odds ratio 1.28, 1.07 to 1.54 (P=0.007)”
Taken as given: The odds ratio is 1.28 with 95% CI 1.07 to 1.54.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Compute p-value from the reported odds ratio and its 95% confidence interval using the normal approximation for the log odds ratio.How we recomputed it: pCI(1.28, 1.07, 1.54, 1) - CONSISTENTreported p = .001 · recomputed p = <.001Reviewers 1, 2Recompute p-value for intervention 1 vs control on feeling confused from reported OR and 95% CI.
“intervention 1: 11.5%, odds ratio 1.92, 1.58 to 2.33 (P<0.001)”
Taken as given: The odds ratio is 1.92 with 95% CI 1.58 to 2.33.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Compute p-value from the reported odds ratio and its 95% confidence interval using the normal approximation for the log odds ratio.How we recomputed it: pCI(1.92, 1.58, 2.33, 1) - CONSISTENTreported p = .001 · recomputed p = <.001Reviewers 1, 2Recompute p-value for intervention 2 vs control on feeling confused from reported OR and 95% CI.
“intervention 2: 9.0%, odds ratio 1.76, 1.46 to 2.13 (P<0.001)”
Taken as given: The odds ratio is 1.76 with 95% CI 1.46 to 2.13.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Compute p-value from the reported odds ratio and its 95% confidence interval using the normal approximation for the log odds ratio.How we recomputed it: pCI(1.76, 1.46, 2.13, 1) - CONSISTENTreported p = .059 · recomputed p = .065Reviewers 1, 2Recompute p-value for intervention 1 vs control on feeling informed from reported OR and 95% CI.
“intervention 1 (odds ratio 0.83, 0.68 to 1.01; P=0.059)”
Taken as given: The odds ratio is 0.83 with 95% CI 0.68 to 1.01.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Compute p-value from the reported odds ratio and its 95% confidence interval using the normal approximation for the log odds ratio.How we recomputed it: pCI(0.83, 0.68, 1.01, 1) - CONSISTENTreported p = .022 · recomputed p = .023Reviewers 1, 2Recompute p-value for intervention 2 vs control on feeling informed from reported OR and 95% CI.
“intervention 2 (odds ratio 0.80, 0.66 to 0.97; P=0.022)”
Taken as given: The odds ratio is 0.80 with 95% CI 0.66 to 0.97.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Compute p-value from the reported odds ratio and its 95% confidence interval using the normal approximation for the log odds ratio.How we recomputed it: pCI(0.80, 0.66, 0.97, 1) - CONSISTENTreported p = .001 · recomputed p = <.001Reviewers 1, 2Recompute p-value for intervention 1 vs control on planning to talk to GP (Yes vs No) from reported RRR and 95% CI.
“intervention 1: 22.8%, relative risk ratio 2.08, 95% CI 1.59 to 2.73 (P<0.001)”
Taken as given: The relative risk ratio is 2.08 with 95% CI 1.59 to 2.73.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Compute p-value from the reported relative risk ratio and its 95% confidence interval using the normal approximation for the log relative risk.How we recomputed it: pCI(2.08, 1.59, 2.73, 1) - CONSISTENTreported p = .001 · recomputed p = <.001Reviewers 1, 2Recompute p-value for intervention 2 vs control on planning to talk to GP (Yes vs No) from reported RRR and 95% CI.
“intervention 2: 19.4%, relative risk ratio 1.71, 1.31 to 2.25 (P<0.001)”
Taken as given: The relative risk ratio is 1.71 with 95% CI 1.31 to 2.25.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Compute p-value from the reported relative risk ratio and its 95% confidence interval using the normal approximation for the log relative risk.How we recomputed it: pCI(1.71, 1.31, 2.25, 1) - CONSISTENTreported p = .040 · recomputed p = .044Reviewers 1, 2Recompute p-value for intervention 1 vs control on planning to go for extra tests (Yes vs No) from reported RRR and 95% CI.
“intervention 1: 2.7%, relative risk ratio 2.09, 1.02 to 4.29 (P=0.04)”
Taken as given: The relative risk ratio is 2.09 with 95% CI 1.02 to 4.29.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Compute p-value from the reported relative risk ratio and its 95% confidence interval using the normal approximation for the log relative risk.How we recomputed it: pCI(2.09, 1.02, 4.29, 1) - CONSISTENTreported p = .010 · recomputed p = .014Reviewers 1, 2Recompute p-value for intervention 2 vs control on planning to go for extra tests (Yes vs No) from reported RRR and 95% CI.
“intervention 2: 3.2%, relative risk ratio 2.38, 1.19 to 4.75 (P=0.01)”
Taken as given: The relative risk ratio is 2.38 with 95% CI 1.19 to 4.75.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Compute p-value from the reported relative risk ratio and its 95% confidence interval using the normal approximation for the log relative risk.How we recomputed it: pCI(2.38, 1.19, 4.75, 1)
- lowinternal contradictionThe abstract reports 3107 women randomised (1030 control, 1003 intervention 1, 1074 intervention 2), but the sum is 3107 (1030+1003+1074=3107). However, the numbers in the results section for the analysis population are 802, 776, 823, which sum to 2401, consistent. No contradiction found.
“3107 women (1030 control, 1003 intervention 1, and 1074 intervention 2) were randomised, and 2401 women (802 control, 776 intervention 1, and 823 intervention 2)”
AbstractFind in source - lowinternal contradictionIn Table 1, the percentages for 'Highest level of education achieved' do not sum to 100% for each group (e.g., control: 20.1+30.6+49.3=100.0, but intervention 1: 23.2+23.5+53.4=100.1, intervention 2: 22.3+27.7+50.1=100.1). This is likely due to rounding and is not a significant concern.
High school or below 153 (20.1) | Diploma or certificate 233 (30.6) | Bachelor’s degree or above 375 (49.3)
Table 1reviewer’s wording
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
5 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Women notified of their dense breasts felt significantly more anxious and confused compared to control.The claim is directly supported by the reported odds ratios and confidence intervals for both intervention groups.Evidence: Odds ratios for anxiety: intervention 1 OR 1.30 (95% CI 1.08-1.57), intervention 2 OR 1.28 (1.07-1.54); for confusion: intervention 1 OR 1.92 (1.58-2.33), intervention 2 OR 1.76 (1.46-2.13).
“Compared with the control group, women who were notified of their dense breasts reported feeling significantly more anxious (intervention 1: odds ratio 1.30, 95% confidence interval (CI) 1.08 to 1.57; intervention 2: odds ratio 1.28, 1.07 to 1.54) and confused (intervention 1: odds ratio 1.92, 1.58 to 2.33; intervention 2: odds ratio 1.76, 1.46 to 2.13)”
AbstractFind in source - supportedReviewers 1, 2Notified women did not feel more informed to make decisions about their breast health.The claim is supported by the odds ratios for feeling informed, which are below 1 and not significant for intervention 1, and significant but below 1 for intervention 2.Evidence: Odds ratios for feeling informed: intervention 1 OR 0.83 (0.68-1.01), P=0.059; intervention 2 OR 0.80 (0.66-0.97), P=0.022.
“Notified women did not feel more informed (intervention 1: odds ratio 0.83, 0.68 to 1.01; intervention 2: odds ratio 0.80, 0.66 to 0.97).”
AbstractFind in source - supportedReviewers 1, 2Women notified of breast density had significantly higher intentions to talk to their general practitioner about their screening results.The claim is supported by the relative risk ratios for GP consultation intention, which are significantly elevated in both intervention groups.Evidence: Relative risk ratios for GP consultation: intervention 1 RRR 2.08 (1.59-2.73), intervention 2 RRR 1.71 (1.31-2.25).
“had significantly higher intentions to talk to their general practitioner about their screening results (intervention1: relative risk ratio 2.08, 95% CI 1.59 to 2.73; intervention 2: relative risk ratio 1.71, 1.31 to 2.25)”
AbstractFind in source - supportedReviewers 1, 2Most women did not intend to have supplemental screening.The claim is supported by the high percentages of women reporting 'No' to supplemental screening intention in all groups.Evidence: Percentages not intending supplemental screening: control 91.3%, intervention 1 78.9%, intervention 2 81.4%.
“However, most women did not intend to have supplemental screening (control: 91.3%; intervention 1: 78.9%; intervention 2: 81.4%).”
AbstractFind in source - supportedReviewers 1, 2Notification of breast density as part of population based breast screening may have adverse outcomes including additional consultation burden on general practitioners.The claim is supported by the increased intentions to consult GPs and the discussion of potential burden, though it is framed as 'may' and is a reasonable interpretation.Evidence: Increased GP consultation intentions and reliance on GPs for supplemental screening advice.
“Notification of breast density as part of population based breast screening may have adverse outcomes including additional consultation burden on general practitioners to advise women.”
DiscussionFind in source
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Other integrity concernAssessed
3 integrity concerns flagged (0 high).
- lowotherThe paper reports that 52 (6.3%) women out of 823 in intervention 2 accessed and watched the video in full, but it does not report how many accessed the video partially or clicked the link. This is a reporting gap but not a validity threat.
“Only 52 (6.3%) women out of 823 in intervention 2 accessed and watched the video in full.”
Results ¶6Find in source
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior research on breast density notification, including post-legislation cross-sectional studies in the US, and acknowledges the lack of robust evidence on population-level benefits versus harms. The rationale for the trial is clearly stated: to investigate the effect of notifying women of dense breasts on psychosocial outcomes and health service use intentions, and to compare written versus video information modes. Limitations of prior research are addressed by noting the lack of randomized controlled trial evidence and the need for better clarity and balance in educational materials.
“Most of the evidence on the outcomes of density notification comes from post-legislation cross sectional studies in the US.”
“The aim of this trial was to investigate the effect of notifying women participating in population based breast cancer screening that they have dense breasts on their psychosocial outcomes and intention for health services use and to determine whether using different modes of information provision (written versus video) alters these effects.”
“despite the lack of robust evidence on whether the benefits of notifying women at the population level outweigh the potential harms.”
“Most of the evidence on the outcomes of density notification comes from post-legislation cross sectional studies in the US.”
“The aim of this trial was to investigate the effect of notifying women participating in population based breast cancer screening that they have dense breasts on their psychosocial outcomes and intention for health services use”
Randomization method is described (simple randomization using Oracle DBMS_RANDOM package) with equal allocation. Blinding of participants is described, though assessor blinding is not explicitly mentioned (but outcomes are self-reported, so assessor blinding is less critical). A priori sample size calculation is provided with effect size, alpha, and power. Inclusion/exclusion criteria are pre-specified. Outlier handling is addressed through complete case analysis and sensitivity analysis with multiple imputation. Controls are appropriate (standard care). Independent replication is not applicable for a single pivotal trial.
“we managed a simple randomisation with equal (1:1:1) allocation by using the Oracle DBMS_RANDOM package following screening results.”
“To achieve 80% power to detect differences in proportions as small as 10% (at a Bonferroni adjusted α level of 0.017 to permit all three possible pairwise comparisons) in the primary outcomes between the control arm and each of the intervention arms, we needed a sample of 373 women randomised to each arm (total n=1119).”
“Women were excluded if they did not have dense breasts (BI-RADS density A (almost entirely fatty) and B (scattered areas of fibroglandular density)), were recalled with screen detected abnormalities, reported a breast symptom or a clinical concern was noted by the radiographer at their screening episode, had a personal history of breast cancer, needed an interpreter (self-report on screening booking information or identified by the screening service), were unable to consent to breast screening, or did not have an active mobile phone number or email address.”
“we managed a simple randomisation with equal (1:1:1) allocation by using the Oracle DBMS_RANDOM package following screening results.”
“To achieve 80% power to detect differences in proportions as small as 10% (at a Bonferroni adjusted α level of 0.017 to permit all three possible pairwise comparisons) in the primary outcomes between the control arm and each of the intervention arms, we needed a sample of 373 women randomised to each arm (total n=1119).”
“Women were excluded if they did not have dense breasts (BI-RADS density A (almost entirely fatty) and B (scattered areas of fibroglandular density)), were recalled with screen detected abnormalities, reported a breast symptom or a clinical concern was noted by the radiographer at their screening episode, had a personal history of breast cancer, needed an interpreter (self-report on screening booking information or identified by the screening service), were unable to consent to breast screening, or did not have an active mobile phone number or email address.”
Sex is reported (all female participants). Age is reported with mean and standard deviation. Demographics including education, income, language, Aboriginal and Torres Strait Islander status, and family history are reported in Table 1. Since the study includes both sexes (all female), sex_justified is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“The mean age at baseline in the total study population was 57.4 (standard deviation 9.9).”
“Women aged 40 years and older who had booked mammographic screening at one of the 13 participating BreastScreen Queensland (BSQ) sites”
“The mean age at baseline in the total study population was 57.4 (standard deviation 9.9).”
The paper names the ethics committees (Gold Coast HHS Ethics Committee, Sunshine Coast, and Metro North) with protocol numbers. Informed consent is described (all participants provided informed consent). Regulatory compliance is implied through adherence to ethical standards, though not explicitly named; however, the ethics approval statement is sufficient.
“Gold Coast Hospital and Health Service (HHS) Ethics Committee (HREC/2023/QGC/8977), Sunshine Coast (SSA/2023/QSC/89770), and Metro North (HREC/2023/QGC/8977) Hospital and Health Service.”
“All participants provided informed consent to participate in the trial.”
“Gold Coast Hospital and Health Service (HHS) Ethics Committee (HREC/2023/QGC/8977), Sunshine Coast (SSA/2023/QSC/89770), and Metro North (HREC/2023/QGC/8977) Hospital and Health Service.”
“All participants provided informed consent to participate in the trial.”
The investigational product (breast density notification) is described in detail, including its content and development. The software used for density measurement (Volpara Health software) and for data collection (Qualtrics) are identified. No antibodies, cell lines, or organisms are used, so those criteria are not applicable.
“The breast density notification (identical in both intervention 1 and intervention 2) advised women in writing in their screening results letter that no sign of breast cancer was found on their recent mammogram screen, but their screen showed that their breasts were dense (supplementary materials).”
“automated volumetric density measures using Volpara Health software”
“collected via the online questionnaire (supplementary materrials) using Qualtrics”
“automated volumetric density measures using Volpara Health software”
“collected via the online questionnaire (supplementary materrials) using Qualtrics”
Statistical tests are named (ordinal and multinomial logistic regression). Assumptions are verified (Brant-Wald test for proportional odds). Exact p-values are reported. Effect sizes with confidence intervals are reported. Statistical software is not explicitly identified, but the analysis methods are described. Data presentation includes per-group n and percentages. Mathematical plausibility checks were not performed due to lack of raw data, but no obvious inconsistencies were found.
“We used generalised linear models (ordinal or multinomial logistic regression, as appropriate, presented as odds ratios or relative risk ratios, respectively) to analyse outcome data across study arms as complete case analyses.”
“We assessed the proportional odds assumption for ordinal logistic regression by using a Brant-Wald test.”
“We used generalised linear models (ordinal or multinomial logistic regression, as appropriate, presented as odds ratios or relative risk ratios, respectively) to analyse outcome data across study arms as complete case analyses.”
The data availability statement says 'The code used to analyse the data in the paper can be found in the supplementary materials.' This is vague and does not provide a clear route for data access. No repository deposit or accession numbers are provided. Code is shared in supplementary materials, which is not a version-controlled public repository.
“The code used to analyse the data in the paper can be found in the supplementary materials.”
“The code used to analyse the data in the paper can be found in the supplementary materials.”
The trial is registered (ACTRN12623000001695). Funding sources and competing interests are declared. Limitations are discussed in detail. Conclusions are proportional to the evidence. Methods are described in sufficient detail for replication. Reporting guideline (CONSORT) is implied through the flow diagram, though not explicitly referenced.
“Trial registration Australian New Zealand Clinical Trials Registry (ACTRN12623000001695).”
“Funding: This work was supported by a National Health and Medical Research Council (NHMRC) Emerging Leader Research Fellowship (1194108) to BN and a National Breast Cancer Foundation chair in breast cancer prevention grant (EC-21-001) to NH.”
“The new findings from this trial should be interpreted in the context of its limitations.”
“Trial registration Australian New Zealand Clinical Trials Registry (ACTRN12623000001695).”
“Although all women received the same density notification integrated in the mammogram results letter, a limitation is the low proportion of women in the video intervention arm who viewed the additional information.”
Registered (1 ID: ANZCTR). Reporting guideline cited: CONSORT.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 48 references by DOI: 42 verified — 6 no DOI (shown, not verified).
- NO DOIACR BI-RADS® Atlas, Breast Imaging Reporting and Data SystemNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIBreastScreen Australia Position Statement on Mammographic (Breast) Density and ScreeningNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImportant Information: Final Rule to Amend the Mammography Quality Standards Act (MQSA)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRelationships among breast cancer perceived absolute risk, comparative risk, and worriesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWellbeing measures in primary health care/the DEPCARE project: report on a WHO meetingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWomen’s responses to information on mammographic breast densityNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORtypoMethods, Outcomes“supplementary materrials”→ supplementary materialsTypographical error.
- MINORconsistencyResults, paragraph 2“P<.001”→ P<0.001Inconsistent formatting of p-values; elsewhere uses P<0.001.
- MINORconsistencyDiscussion, paragraph 1“supplementary ultrasonography or MRI”→ supplemental ultrasonography or MRIInconsistent use of 'supplementary' vs 'supplemental'.
- MINORclarityDiscussion, Strengths and limitations“Although only a third of women invited to participate in the trial consented, participation rates in cancer screening trials vary widely, with opt-in consent leading to lower participation rates.”→ Consider rephrasing for clarity: 'Although only a third of invited women consented, participation rates in cancer screening trials vary widely, and opt-in consent typically leads to lower participation.'Slightly awkward phrasing.
- MINORconsistencyResults, paragraph 3“P<.001”→ P<0.001Inconsistent formatting of p-values.
- MINORclarityDiscussion, Policy and practice implications“screenng”→ screeningTypographical error.
The published paper is methodologically robust and its conclusions are supported by the reported analyses. An informed reader should weigh the minor reporting gaps (data availability, statistical software, CONSORT reference) and the fact that only a subset of statistics were independently recomputed; none of these undermine the core findings, but a correction or clarification of the data availability statement would improve reproducibility.
- 1.HIGHdata codeIn the Data availability statement, specify a concrete access route for the de-identified dataset (e.g., a data access committee with contact details and conditions) and provide a persistent identifier (DOI) for the analysis code in a public repository such as Zenodo or GitHub.The current statement only mentions code in supplementary materials, which is inadequate for reproducibility and does not meet common journal requirements for clinical trials.
- 2.HIGHreportingIn the Methods or a dedicated section, explicitly state adherence to the CONSORT reporting guideline and provide a completed CONSORT checklist as a supplementary file.The flow diagram implies CONSORT but the guideline is not referenced, which is a reporting transparency gap that reviewers may flag.
- 3.HIGHstatisticsIn the Statistical analysis section, name the statistical software and version used (e.g., R version 4.3.1, Stata 18) for all analyses.Both reviewers noted that statistical software is not identified, which hampers reproducibility.
- 4.MEDIUMethicsIn the Ethics statements, add an explicit statement of compliance with the Declaration of Helsinki or equivalent regulatory standards.Reviewer 2 flagged regulatory compliance as not explicitly stated; adding this clarifies ethical oversight.
- 5.MEDIUMrigorIn the Methods, clarify whether outcome assessors were blinded to group allocation, even though outcomes are self-reported.Blinding of assessors is not explicitly described; clarifying this addresses a potential source of bias.
- 6.MEDIUMrigorIn the Results or Discussion, report the number of women in intervention 1 who accessed the written information, and in intervention 2 who partially viewed the video, to aid interpretation of engagement.The paper reports only full video views; partial access data would help interpret the intervention's reach.
- 7.MEDIUMreportingIn the Methods, provide more detail on the development and validation of the study-specific questionnaire measures used as outcomes.Reviewer 2 suggested this; it would strengthen the validity of the outcome measures.
- 8.LOWcopyeditFix the typo 'supplementary materrials' to 'supplementary materials' in Methods, Outcomes.Typographical error that should be corrected.
- 9.LOWcopyeditStandardize p-value formatting to 'P<0.001' throughout the Results (currently 'P<.001' appears in paragraphs 2 and 3).Inconsistent formatting of p-values is a minor copyedit issue.
- 10.LOWcopyeditIn the Discussion, replace 'supplementary ultrasonography or MRI' with 'supplemental ultrasonography or MRI' for consistency.Inconsistent use of 'supplementary' vs 'supplemental'.
- 11.LOWcopyeditIn the Discussion, Strengths and limitations, rephrase the sentence about participation rates for clarity: 'Although only a third of invited women consented, participation rates in cancer screening trials vary widely, and opt-in consent typically leads to lower participation.'The original phrasing is slightly awkward and could be clearer.
- 12.LOWcopyeditIn the Discussion, Policy and practice implications, fix the typo 'screenng' to 'screening'.Typographical error.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.