Effectiveness of screening and ultra-brief intervention for hazardous drinking in primary care: pragmatic cluster randomised controlled trial.
So R, Kariyama K, Oyamada S, Matsushita S, Nishimura H, Tezuka Y, Sunami T, Furukawa TA, Sahker E, Kawaguchi M, Kobashi H, Nishina S, Otsuka Y, Kanda H, Tsujimoto Y, Horie Y, Yoshiji H, Yuzuriha T, Nouso K
- DOI
- 10.1136/bmj-2024-083985
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/c5ded045-df5b-4ec8-acb9-bd87f5a4ecde is authoritative.
How this rating was calculated
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingStatistical analysis partially met−0.25★
- CitationsUnresolved reference−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 9 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is total alcohol consumption (g/4wk), a self-reported behavioral measure, which is a surrogate for clinical outcomes such as mortality, liver disease, or quality of life. The paper does not demonstrate target engagement at the tested dose (e.g., no PK/PD or dose-exposure data) and does not cite validated evidence linking this specific measure to hard clinical outcomes. Although the paper mentions WHO drinking risk levels as meaningful surrogate markers, it does not provide validation linking the primary outcome to clinical benefit.
“Total alcohol consumption and drinking risk level, which can be calculated from total alcohol consumption, are meaningful surrogate markers that reflect patients’ health and quality of life.”
- 02Treatment effect not shown to be clinically meaningful
The primary outcome shows a null result with a difference of 27.8 g/4wk (95% CI −149.7 to 205.4, P=0.75) and Hedges' g of 0.02, which is not statistically significant and far below any clinically meaningful threshold. The paper does not anchor this effect size to a minimal clinically important difference or demonstrate biological/clinical meaningfulness. The secondary outcome of readiness to change shows a small effect (Hedges' g 0.21 at 12 weeks) but is not linked to clinical benefit.
“The difference between groups was 27.8 g/4wk (95% CI −149.7 to 205.4), with a Hedges’ g of 0.02 (95% CI −0.10 to 0.14).”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported pragmatic cluster randomised controlled trial. The design is rigorous, ethical approvals are clear, and data/code are openly available. Minor reporting gaps (e.g., missing table references, unit inconsistencies) and a single unresolved reference do not undermine the core findings.
Both reviewers classified the study as interventional, which is adopted. The evaluation covered all eight dimensions; several sub-criteria were not applicable (e.g., animal-related, cell line authentication). The reviewers diverged on statistical analysis (warn vs. pass), but the combined evidence supports 'pass'.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 8 tests: 8 consistent, 0 inconsistent; 8 via agent-written checks.
- CONSISTENTreported p = .750 · recomputed p = .759Reviewer 1Primary outcome difference at 24 weeks (ITT population)
“At 24 weeks, the difference in total alcohol consumption between the ultra-brief intervention group (1046.9 g/4 weeks (g/4wk), 95% confidence interval (CI) 918.3 to 1175.4) and control group (1019.0 g/4wk, 893.5 to 1144.6) was 27.8 g/4wk (−149.7 to 205.4, P=0.75), with a Hedges’ g of 0.02 (95% CI −0.10 to 0.14).”
Taken as given: The reported difference of 27.8 g/4wk is the point estimate.; The 95% CI for the difference is -149.7 to 205.4.; The p-value is derived from a z-test or similar large-sample test where the standard error can be estimated from the CI width.; The test is two-sided.Method: Two-sided z-test from point estimate and 95% CIHow we recomputed it: 2*(1-normalCdf(Math.abs(27.8 / ((205.4 - (-149.7)) / (2 * 1.96)))) ) - CONSISTENTreported p = .490 · recomputed p = .499Reviewer 1Primary outcome difference at 12 weeks (ITT population)
“At 12 weeks, the difference in total alcohol consumption between the intervention group (1034.1 g/4wk, 919.6 to 1148.7) and control group (979.3 g/4wk, 866.1 to 1092.4) was 54.9 g/4wk (−104.1 to 213.9, P=0.49), with a Hedges’ g of 0.04 (−0.08 to 0.16).”
Taken as given: The reported difference of 54.9 g/4wk is the point estimate.; The 95% CI for the difference is -104.1 to 213.9.; The p-value is derived from a z-test or similar large-sample test where the standard error can be estimated from the CI width.; The test is two-sided.Method: Two-sided z-test from point estimate and 95% CIHow we recomputed it: 2*(1-normalCdf(Math.abs(54.9 / ((213.9 - (-104.1)) / (2 * 1.96)))) ) - CONSISTENTreported p < .010 · recomputed p = <.001Reviewer 1Readiness to change drinking behaviour difference at 12 weeks (ITT population)
“Readiness to change drinking behaviour, converted into numerical scores (higher scores indicating greater readiness to change), was higher in the ultra-brief intervention group compared with control group at both 12 weeks (difference 0.25 (95% CI 0.12 to 0.39); Hedges’ g 0.21 (95% CI 0.10 to 0.33)) and 24 weeks (difference 0.19 (0.05 to 0.32); Hedges’ g 0.16 (0.05 to 0.28)) (, ).”
Taken as given: The reported difference of 0.25 is the point estimate.; The 95% CI for the difference is 0.12 to 0.39.; The p-value is derived from a z-test or similar large-sample test where the standard error can be estimated from the CI width.; The test is two-sided.Method: Two-sided z-test from point estimate and 95% CIHow we recomputed it: 2*(1-normalCdf(Math.abs(0.25 / ((0.39 - 0.12) / (2 * 1.96)))) ) - CONSISTENTreported p < .010 · recomputed p = .006Reviewer 1Readiness to change drinking behaviour difference at 24 weeks (ITT population)
“Readiness to change drinking behaviour, converted into numerical scores (higher scores indicating greater readiness to change), was higher in the ultra-brief intervention group compared with control group at both 12 weeks (difference 0.25 (95% CI 0.12 to 0.39); Hedges’ g 0.21 (95% CI 0.10 to 0.33)) and 24 weeks (difference 0.19 (0.05 to 0.32); Hedges’ g 0.16 (0.05 to 0.28)) (, ).”
Taken as given: The reported difference of 0.19 is the point estimate.; The 95% CI for the difference is 0.05 to 0.32.; The p-value is derived from a z-test or similar large-sample test where the standard error can be estimated from the CI width.; The test is two-sided.Method: Two-sided z-test from point estimate and 95% CIHow we recomputed it: 2*(1-normalCdf(Math.abs(0.19 / ((0.32 - 0.05) / (2 * 1.96)))) ) - CONSISTENTreported p = .750 · recomputed p = .759Reviewer 2Primary outcome difference at 24 weeks (ITT): p-value from CI
“The difference between groups was 27.8 g/4wk (95% CI −149.7 to 205.4), with a Hedges’ g of 0.02 (95% CI −0.10 to 0.14).”
Taken as given: The CI is a 95% confidence interval for the difference.; The difference is normally distributed (or the CI is derived from a t-distribution).Method: Two-tailed p-value derived from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(27.8, -149.7, 205.4, 0) - CONSISTENTreported p = .490 · recomputed p = .499Reviewer 2Secondary outcome difference at 12 weeks (ITT): p-value from CI
“At 12 weeks, the difference in total alcohol consumption between the intervention group (1034.1 g/4wk, 919.6 to 1148.7) and control group (979.3 g/4wk, 866.1 to 1092.4) was 54.9 g/4wk (−104.1 to 213.9, P=0.49)”
Taken as given: The CI is a 95% confidence interval for the difference.; The difference is normally distributed.Method: Two-tailed p-value derived from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(54.9, -104.1, 213.9, 0) - CONSISTENTreported p < .010 · recomputed p = <.001Reviewer 2Readiness to change drinking behaviour at 12 weeks: p-value from CI
“Readiness to change drinking behaviour, converted into numerical scores (higher scores indicating greater readiness to change), was higher in the ultra-brief intervention group compared with control group at both 12 weeks (difference 0.25 (95% CI 0.12 to 0.39); Hedges’ g 0.21 (95% CI 0.10 to 0.33))”
Taken as given: The CI is a 95% confidence interval for the difference.; The difference is normally distributed.Method: Two-tailed p-value derived from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(0.25, 0.12, 0.39, 0) - CONSISTENTreported p < .010 · recomputed p = .006Reviewer 2Readiness to change drinking behaviour at 24 weeks: p-value from CI
“and 24 weeks (difference 0.19 (0.05 to 0.32); Hedges’ g 0.16 (0.05 to 0.28))”
Taken as given: The CI is a 95% confidence interval for the difference.; The difference is normally distributed.Method: Two-tailed p-value derived from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(0.19, 0.05, 0.32, 0)
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
9 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewer 1This trial found no evidence to support the effectiveness of a doctor delivered ultra-brief intervention for hazardous drinking compared with simplified assessment only in primary care in Japan.The primary outcome analysis for total alcohol consumption at 24 weeks showed a non-significant difference between groups (P=0.75), directly supporting this conclusion.Evidence: Results section, primary outcome analysis, Table 2
“This trial found no evidence to support the effectiveness of a doctor delivered ultra-brief intervention for hazardous drinking compared with simplified assessment only in primary care in Japan.”
ConclusionFind in source - supportedReviewer 1The findings of this trial did not support the effectiveness of ultra-brief intervention on total alcohol consumption at 12 and 24 weeks compared with simplified assessment only among patients with hazardous drinking in Japanese primary care settings.The results for total alcohol consumption at both 12 and 24 weeks showed non-significant differences between the intervention and control groups, confirming the lack of effectiveness.Evidence: Results section, Table 2, showing P=0.49 at 12 weeks and P=0.75 at 24 weeks for total alcohol consumption.
“The findings of this trial did not support the effectiveness of ultra-brief intervention on total alcohol consumption at 12 and 24 weeks compared with simplified assessment only among patients with hazardous drinking in Japanese primary care settings.”
Discussion ¶1Find in source - supportedReviewer 1Readiness to change drinking behaviour was more favourable in the ultra-brief intervention group at both 12 weeks and 24 weeks.The statistical analysis showed significant differences in readiness to change drinking behaviour in favor of the intervention group at both 12 weeks (P<0.01) and 24 weeks (P<0.01), with positive Hedges' g values.Evidence: Results section, Table 4, showing P<0.01 for readiness to change drinking behaviour at both time points.
“However, readiness to change drinking behaviour was more favourable in the ultra-brief intervention group at both 12 weeks and 24 weeks.”
Discussion ¶1Find in source - supportedReviewer 1The absence of similar improvements in readiness to change diet or smoking behaviour suggests the ultra-brief intervention specifically affected readiness to change drinking behaviour.The results for readiness to change diet and smoking behavior showed no significant differences between groups, supporting the specificity of the intervention's effect on drinking behavior readiness.Evidence: Results section, paragraph 5, stating mean scores for readiness to change diet and smoking behavior were not significantly different between groups.
“The absence of similar improvements in readiness to change diet or smoking behaviour suggests the ultra-brief intervention specifically affected readiness to change drinking behaviour.”
Discussion ¶1Find in source - supportedReviewer 1From a public health perspective, our results suggest that widely implementing ultra-brief intervention in primary care may not lead to clinically relevant benefits to reduce alcohol consumption.Given the null findings for the primary outcome of total alcohol consumption, this public health implication is directly supported by the study's results.Evidence: Results section, primary outcome analysis, Table 2, showing no significant reduction in alcohol consumption.
“From a public health perspective, our results suggest that widely implementing ultra-brief intervention in primary care may not lead to clinically relevant benefits to reduce alcohol consumption.”
Implications, paragraph 1Find in source - supportedReviewer 2The ultra-brief intervention was not effective in reducing alcohol consumption compared with simplified assessment only.The primary outcome analysis shows a non-significant difference with a small effect size, supporting the null finding.Evidence: Primary outcome at 24 weeks: difference 27.8 g/4wk (95% CI −149.7 to 205.4, P=0.75), Hedges' g 0.02.
“At 24 weeks, the difference in total alcohol consumption between the ultra-brief intervention group (1046.9 g/4 weeks (g/4wk), 95% confidence interval (CI) 918.3 to 1175.4) and control group (1019.0 g/4wk, 893.5 to 1144.6) was 27.8 g/4wk (−149.7 to 205.4, P=0.75), with a Hedges’ g of 0.02 (95% CI −0.10 to 0.14).”
AbstractFind in source - supportedReviewer 2The ultra-brief intervention improved readiness to change drinking behaviour.Secondary outcome analyses show statistically significant improvements in readiness to change at both time points.Evidence: Readiness to change drinking behaviour was higher in the intervention group at 12 weeks (difference 0.25, 95% CI 0.12 to 0.39) and 24 weeks (difference 0.19, 95% CI 0.05 to 0.32).
Readiness to change drinking behaviour, converted into numerical scores (higher scores indicating greater readiness to change), was higher in the ultra-brief intervention group compared with control group at both 12 weeks (difference 0.25 (95% CI 0.12 to 0.39); Hedges’ g 0.21 (95% CI 0.10 to 0.33)) and 24 weeks (difference 0.19 (0.05 to 0.32); Hedges’ g 0.16 (0.05 to 0.28)).
Resultsreviewer’s wording - supportedReviewer 2The findings challenge the syllogistic reasoning that ultra-brief intervention would be effective.The null result on the primary outcome directly contradicts the assumption that ultra-brief intervention is effective, supporting this claim.Evidence: The primary outcome was null, and the discussion explicitly states this challenges the syllogistic reasoning.
“These findings challenge the syllogistic reasoning that ultra-brief intervention would be effective based on the reported effectiveness of standard brief intervention and the assumed equivalence between the two approaches.”
Discussion ¶1Find in source - supportedReviewer 2The ultra-brief intervention specifically affected readiness to change drinking behaviour, not diet or smoking.The results show significant effects on drinking readiness but not on diet or smoking readiness, supporting specificity.Evidence: Readiness to change diet and smoking showed non-significant differences, while drinking readiness showed significant differences.
“The absence of similar improvements in readiness to change diet or smoking behaviour suggests the ultra-brief intervention specifically affected readiness to change drinking behaviour.”
Discussion ¶1Find in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is total alcohol consumption (g/4wk), a self-reported behavioral measure, which is a surrogate for clinical outcomes such as mortality, liver disease, or quality of life. The paper does not demonstrate target engagement at the tested dose (e.g., no PK/PD or dose-exposure data) and does not cite validated evidence linking this specific measure to hard clinical outcomes. Although the paper mentions WHO drinking risk levels as meaningful surrogate markers, it does not provide validation linking the primary outcome to clinical benefit.
“Total alcohol consumption and drinking risk level, which can be calculated from total alcohol consumption, are meaningful surrogate markers that reflect patients’ health and quality of life.”
- INADEQUATEEffect sizeThe primary outcome shows a null result with a difference of 27.8 g/4wk (95% CI −149.7 to 205.4, P=0.75) and Hedges' g of 0.02, which is not statistically significant and far below any clinically meaningful threshold. The paper does not anchor this effect size to a minimal clinically important difference or demonstrate biological/clinical meaningfulness. The secondary outcome of readiness to change shows a small effect (Hedges' g 0.21 at 12 weeks) but is not linked to clinical benefit.
“The difference between groups was 27.8 g/4wk (95% CI −149.7 to 205.4), with a Hedges’ g of 0.02 (95% CI −0.10 to 0.14).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Statistical reporting gaps (tests, assumptions, effect sizes)Assessed
The paper discusses the global public health concern of harmful alcohol use and the recommendation of brief interventions. It highlights mixed results for ultra-brief interventions in different settings and explicitly states the lack of randomized controlled trials directly investigating ultra-brief intervention over assessment-only control in primary care settings.
“Available evidence, however, remains inconclusive as to whether the ultra-brief intervention and more time intensive brief intervention are equally effective or equally ineffective compared with assessment only control.”
“As no randomised controlled trial has directly investigated the effectiveness of ultra-brief intervention over assessment only control in primary care settings, we designed and conducted a large scale pragmatic cluster randomised controlled trial in primary care settings in Japan.”
“As no randomised controlled trial has directly investigated the effectiveness of ultra-brief intervention over assessment only control in primary care settings, we designed and conducted a large scale pragmatic cluster randomised controlled trial in primary care settings in Japan.”
“The Screening and Intervention Programme for Sensible drinking (SIPS) study, a large scale cluster randomised controlled trial, found that the group assigned to an ultra-brief intervention, comprising a leaflet with feedback on screening results, showed comparable reductions in hazardous drinking to the groups assigned to more time intensive brief interventions.”
“As no randomised controlled trial has directly investigated the effectiveness of ultra-brief intervention over assessment only control in primary care settings, we designed and conducted a large scale pragmatic cluster randomised controlled trial in primary care settings in Japan.”
“Available evidence, however, remains inconclusive as to whether the ultra-brief intervention and more time intensive brief intervention are equally effective or equally ineffective compared with assessment only control.”
Cluster randomization used a computer-generated random sequence with block randomization. Blinding of participants and outcome assessors was maintained. Sample size calculation was pre-specified with assumptions. Inclusion/exclusion criteria were defined. Outlier handling is addressed through ITT and per-protocol analyses. Controls are appropriate (simplified assessment only). Independent replication is not applicable for a single pivotal trial.
“We allocated clinic clusters to the two arms of the study using a block randomisation method without any stratification or matching.”
“We allocated clinic clusters to the two arms of the study using a block randomisation method without any stratification or matching.”
“This was determined based on an assumed effect size of 0.25 standardised mean difference for the ultra-brief intervention, with a two sided α level of 0.05, a power of 0.80, 40 clusters, and an intracluster correlation coefficient between 0.03 and 0.04.”
“We analysed observed cases only, without imputation for missing data, under the assumption that missing data occurred at random.”
“We allocated clinic clusters to the two arms of the study using a block randomisation method without any stratification or matching. The statistician (SO), who was not involved in the recruitment of clusters, generated a random sequence on a computer.”
“Participants and staff who collected participant reported outcomes remained blinded to assignment.”
“We set the sample size for the cluster randomised controlled trial at 1125 participants. This was determined based on an assumed effect size of 0.25 standardised mean difference for the ultra-brief intervention, with a two sided α level of 0.05, a power of 0.80, 40 clusters, and an intracluster correlation coefficient between 0.03 and 0.04.”
The paper provides detailed baseline characteristics of participants, including age group, mean age with standard deviation, percentage of men, and medical history (hypertension, hyperuricaemia, diabetes, dyslipidaemia, liver diseases, digestive diseases).
“% men | 342 (64) | 417 (69) | 759 (67)”
“Mean (SD) age* (years) | 57.3 (12.6) | 58.0 (12.7) | 57.6 (12.6)”
“Medical history: | | Hypertension | 287 (54) | 347 (58) | 634 (56)”
“% men | 342 (64) | 417 (69) | 759 (67)”
“Mean (SD) age* (years) | 57.3 (12.6) | 58.0 (12.7) | 57.6 (12.6)”
“Hypertension | 287 (54) | 347 (58) | 634 (56)”
The study protocol was approved by a named IRB (Kurihama Medical and Addiction Centre) with a registration number. Written informed consent was obtained from all participants. Regulatory compliance with the Declaration of Helsinki is stated.
“The study protocol was approved by the institutional review board of the Kurihama Medical and Addiction Centre on 22 May 2023 (registration No 423).”
“The entire study adhered to the Declaration of Helsinki, and we obtained written informed consent from all participants.”
“The entire study adhered to the Declaration of Helsinki, and we obtained written informed consent from all participants.”
“The study protocol was approved by the institutional review board of the Kurihama Medical and Addiction Centre on 22 May 2023 (registration No 423).”
“The entire study adhered to the Declaration of Helsinki, and we obtained written informed consent from all participants.”
The investigational product (alcohol information leaflet) is described in detail, including its content and development. Statistical software (SAS 9.4 and R 4.2.3) is identified. No antibodies, cell lines, or organisms are used, so those criteria are not applicable.
“We used SAS software version 9.4 (SAS Institute, Cary, NC) and R version 4.2.3.”
“The double sided leaflet is in a simple and compact format, using easy-to-understand illustrations and colours (see supplementary files 1-4).”
“The template of the oral message was: “Mr./Ms. [Name], you might drink too much.” (Feedback); “I recommend you calculate your alcohol consumption on your own.” (Advice); “The information here (in the leaflet) will be beneficial for you.” (chance of effectiveness); “You can easily apply these recommendations (in the leaflet) starting today.” (Assurance of feasibility); and “I look forward to hearing your thoughts when we meet next time” (Motivation through commitment).”
“The double sided leaflet is in a simple and compact format, using easy-to-understand illustrations and colours (see supplementary files 1-4).”
“We used SAS software version 9.4 (SAS Institute, Cary, NC) and R version 4.2.3.”
The paper names general linear mixed effects models and specifies SAS and R versions. Effect sizes (Hedges' g) with 95% CIs are reported. However, many p-values are given as 'P=0.75' or 'P<0.01' rather than exact values with 2-3 significant figures. While the model structure is described, explicit verification of assumptions (e.g., normality of residuals, appropriateness of unstructured variance-covariance matrix) is not detailed.
“A general linear mixed effects model was used to estimate the least square means of total alcohol consumption in the past four weeks for each group at each time point, along with the point estimates and 95% confidence intervals (CIs) for the differences between groups.”
“At 24 weeks, the difference in total alcohol consumption between the ultra-brief intervention group (1046.9 g/4 weeks (g/4wk), 95% confidence interval (CI) 918.3 to 1175.4) and control group (1019.0 g/4wk, 893.5 to 1144.6) was 27.8 g/4wk (−149.7 to 205.4, P=0.75)”
“with a Hedges’ g of 0.02 (95% CI −0.10 to 0.14).”
“An unstructured variance-covariance matrix was assumed for the correlation between time points for the dependent variable, and the Kenward-Roger method was used to calculate the degrees of freedom.”
“A general linear mixed effects model was used to estimate the least square means of total alcohol consumption in the past four weeks for each group at each time point, along with the point estimates and 95% confidence intervals (CIs) for the differences between groups.”
“with a Hedges’ g of 0.02 (95% CI −0.10 to 0.14).”
The data availability statement provides a concrete route: a Dryad DOI. This satisfies the data_availability_statement and repository_deposit criteria. Accession numbers are not applicable for this type of data. Code sharing is covered by the same repository.
“The datasets and statistical codes can be found at https://doi.org/10.5061/dryad.866t1g22m .”
“The datasets and statistical codes can be found at https://doi.org/10.5061/dryad.866t1g22m .”
“The datasets and statistical codes can be found at https://doi.org/10.5061/dryad.866t1g22m .”
“The datasets and statistical codes can be found at https://doi.org/10.5061/dryad.866t1g22m .”
The paper is registered with UMIN Clinical Trials Registry, states adherence to CONSORT guidelines, and explicitly discusses limitations including potential biases and the nature of the control condition. Funding sources and conflicts of interest are clearly declared. Importantly, it transparently reports changes from the original statistical analysis plan and identifies post-hoc analyses.
“Trial registration UMIN Clinical Trials Registry UMIN000051388.”
“In this paper we report the results in accordance with the CONSORT (Consolidated Standards of Reporting Trials) 2010 extension for cluster randomised controlled trials.”
“However, some limitations should be acknowledged. Baseline data on total alcohol consumption and readiness to change drinking behaviour were not collected. Although this was to avoid the potential effect of intensive screening, it became difficult to balance between group differences in these outcomes at baseline due to random error.”
“Funding: The EASY (Education on Alcohol after Screening to Yield moderated drinking) study was conducted with a grant from the Japan Agency for Medical Research and Development (AMED) (23he0122018j0003), awarded to the consortium comprising CureApp, the Kurihama Medical and Addiction Centre, and Okayama City Hospital.”
“Of note, among the outcomes analysed, only total alcohol consumption was prespecified in the protocol, statistical analysis plan version 1.0, and trial registry. Analyses of readiness to change drinking behaviour, as well as diet and smoking behaviour, were inadvertently omitted from both the protocol and the trial registration and are therefore considered post hoc analyses.”
“Trial registration UMIN Clinical Trials Registry UMIN000051388.”
“In this paper we report the results in accordance with the CONSORT (Consolidated Standards of Reporting Trials) 2010 extension for cluster randomised controlled trials.”
“However, some limitations should be acknowledged. Baseline data on total alcohol consumption and readiness to change drinking behaviour were not collected.”
Registered (1 ID: UMIN-CTR). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 54 references by DOI: 46 verified — 1 DOI unresolved, 7 no DOI (shown, not verified).
- UNRESOLVED10.13140/rg.2.2.29007.48808International guide for monitoring alcohol consumption and related harmCited DOI does not resolve to any Crossref record.
- NO DOIGlobal Status Report on Alcohol and Health 2018No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Alcohol Use Disorders Identification Test: Guidelines for use in primary careNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMotivating young adults for treatment and lifestyle changeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIR: A Language and Environment for Statistical ComputingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRevised Cochrane risk of bias tool for randomized trials (RoB 2): additional considerations for cluster-randomized trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIASK (Nonprofit Organization): A Website for Disseminating Information on Preventing and Supporting Recovery from Alcohol, Drug, and Other Dependency IssuesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAlcohol consumption in Japan: different culture, different rulesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
2 data/code links checked; 2 live.
- datahttps://www.umin.ac.jp/ctr/LIVEHTTP 200Resolved page looks like data.
- dataDryadLIVEHTTP 200https://doi.org/10.5061/dryad.866t1g22mResolves to Dryad (data repository).
Copyediting
15 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 15 minor suggestions below.
15 copyedit issues flagged: mostly grammar, consistency, clarity.
- MINORconsistencyAbstract, Results“g/4 weeks (g/4wk)”→ g/4wkThe unit 'g/4 weeks' is sometimes written out and sometimes abbreviated as 'g/4wk'; consistent abbreviation would improve readability.
- MINORclarityAbstract, Results“At 24 weeks, the difference in total alcohol consumption between the ultra-brief intervention group (1046.9 g/4 weeks (g/4wk), 95% confidence interval (CI) 918.3 to 1175.4) and control group (1019.0 g/4wk, 893.5 to 1144.6) was 27.8 g/4wk (−149.7 to 205.4, P=0.75), with a Hedges’ g of 0.02 (95% CI −0.10 to 0.14).”→ Consider rephrasing to clarify that the difference of 27.8 g/4wk is the point estimate, and the CI is for this difference, not for the individual group means.The phrasing could be slightly clearer about which CI belongs to which value.
- MINORgrammarMethods, paragraph 1“Settings and procedures presents a summary of the trial procedures...”→ Settings and procedures present a summary of the trial procedures...Subject-verb agreement: 'Settings and procedures' is plural.
- MINORconsistencyMethods, paragraph 3“The statistician (SO), who was not involved in the recruitment of clusters, generated a random sequence on a computer. Using this random sequence, another researcher (SM), also not involved in the recruitment of clusters, allocated the clusters to the study groups before each site initiated the screening and intervention.”→ Ensure consistent use of full names or initials for authors throughout the methods section for clarity.Authors are referred to by initials, but it's not immediately clear who SO and SM are without checking the author list.
- MINORclarityMeasures, Screening survey, paragraph 4“We developed these question and response options based on the Japanese National Health and Nutrition Survey, although their psychometric reliability and validity have not yet been confirmed.”→ Clarify if 'these question and response options' refers to the readiness to change drinking behavior questions specifically, or all questions developed based on the survey.Ambiguity in the scope of the unconfirmed psychometric properties.
- MINORconsistencyOutcomes, paragraph 1“total alcohol consumption is equivalent to the average alcohol consumption specified as the primary outcome in the trial registration”→ Ensure consistent terminology for the primary outcome throughout the paper and registration.Slight variation in wording for the primary outcome between the paper and registration.
- MINORgrammarResults, paragraph 1“presents the baseline characteristics of the participants.”→ Table 1 presents the baseline characteristics of the participants.Missing reference to Table 1.
- MINORgrammarResults, paragraph 2“and present the total alcohol consumption at 12 and 24 weeks...”→ Table 2 presents the total alcohol consumption at 12 and 24 weeks...Missing reference to Table 2.
- MINORgrammarResults, paragraph 3“also includes results from the sensitivity analysis...”→ Table 2 also includes results from the sensitivity analysis...Missing reference to Table 2.
- MINORgrammarResults, paragraph 4“When categorising total alcohol consumption by WHO drinking risk levels in the ITT population at 24 weeks, 263 participants (58%) in the ultra-brief intervention group were classified as abstinent or low risk and 108 (24%) as medium risk ().”→ Add a table reference here, e.g., 'Table 3'.Missing table reference.
- MINORgrammarResults, paragraph 5“Additionally, at 12 weeks 192 (46%) participants in the ultra-brief intervention group reported “Intending to improve” compared with 174 (37%) in the control group, and 68 (16%) versus 47 (10%), respectively, reported “Already working on improvement.” At 24 weeks, 197 (44%) participants in the ultra-brief intervention group and 206 (41%) in the control group reported “Intending to improve,” while 78 (17%) in the ultra-brief intervention group and 53 (10%) in the control group reported “Already working on improvement” ().”→ Add a table reference here, e.g., 'Table 5'.Missing table reference.
- MINORgrammarResults, paragraph 6“Subgroup analysis and present the results of the subgroup analysis...”→ Figures 5 and 6 present the results of the subgroup analysis...Missing reference to Figures 5 and 6.
- MINORclarityDiscussion, Possible reasons for null findings, paragraph 5“The low intensity of the training for doctors, which involved only watching a video, might have reduced the effect of the ultra-brief intervention owing to insufficient confidence or motivation of the doctors in dealing with alcohol related issues.”→ Consider if 'owing to' could be replaced with 'due to' for slightly more direct phrasing.Minor stylistic suggestion.
- MINORconsistencyAbstract, Results“1046.9 g/4 weeks (g/4wk)”→ Use consistent unit notation: 'g/4wk' throughout.Minor inconsistency in unit abbreviation.
- MINORclarityMethods, Statistical analysis“We analysed observed cases only, without imputation for missing data, under the assumption that missing data occurred at random.”→ Clarify that this is a complete-case analysis and discuss potential bias if missing not at random.Could be clearer about the missing data assumption.
The published work is robust and well-reported. An informed reader should weigh the minor reporting gaps (e.g., missing table references, unit inconsistencies) and the single unresolved reference, but these do not affect the validity of the conclusions. No erratum is warranted for the core findings.
- 1.HIGHreportingVerify and correct the reference 'International guide for monitoring alcohol consumption and related harm' (DOI 10.13140/rg.2.2.29007.48808) which was not found in any registry; if it cannot be verified, remove or replace it.An unresolved reference may indicate a fabrication signal and must be addressed.
- 2.MEDIUMcopyeditAdd missing table/figure references in the Results section (e.g., Table 3, Table 5, Figures 5 and 6) where data are presented.Missing references reduce clarity and hinder the reader's ability to locate the data.
- 3.MEDIUMcopyeditStandardize the unit notation for alcohol consumption to 'g/4wk' throughout the manuscript, including the abstract.Inconsistent unit abbreviations (g/4 weeks vs g/4wk) are a minor but avoidable distraction.
- 4.MEDIUMreportingAdd a CONSORT flow diagram as a figure to enhance transparency of participant flow.A flow diagram is a standard element of CONSORT reporting and improves clarity.
- 5.MEDIUMreportingProvide the statistical analysis plan as a supplementary file for full transparency.Sharing the SAP allows readers to verify that analyses were pre-specified.
- 6.MEDIUMreportingReport the number of participants with missing data for each outcome in the tables.Transparent missing data reporting helps readers assess potential bias.
- 7.MEDIUMreportingReport the intracluster correlation coefficient for the primary outcome in the abstract or main text.The ICC is a key design parameter for cluster trials and its reporting is informative.
- 8.MEDIUMreportingReport the number of clinics that declined to participate and reasons for non-participation.This information is relevant for assessing generalizability and selection bias.
- 9.MEDIUMreportingReport baseline characteristics of the clusters (e.g., size, location) to assess comparability.Cluster-level comparability is important in cluster randomized trials.
- 10.MEDIUMreportingReport the results of the per-protocol analysis for the secondary outcomes.Per-protocol results complement the ITT analysis and provide insight into treatment effects.
- 11.LOWcopyeditClarify the scope of the psychometric properties statement in the Measures section (whether it applies to all questions or just readiness to change).Ambiguity in this statement could confuse readers about the validity of the measures.
- 12.LOWcopyeditEnsure consistent use of author initials (e.g., SO, SM) throughout the methods section, or spell out full names at first mention.Consistent naming improves readability.
- 13.LOWreportingDiscuss the potential impact of the inadvertent omission of a response option in the follow-up survey on the readiness to change outcome.This limitation could affect the interpretation of that outcome.
- 14.LOWreportingReport the exact p-values for the subgroup analyses in the forest plots.Exact p-values allow readers to assess the strength of subgroup findings.
- 15.LOWreportingAdd a statement about the availability of the study protocol.Providing access to the full protocol enhances transparency.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.