Personal protective effect of wearing surgical face masks in public spaces on self-reported respiratory symptoms in adults: pragmatic randomised superiority trial.
Solberg RB, Fretheim A, Elgersma IH, Fagernes M, Iversen BG, Hemkens LG, Rose CJ, Elstrøm P
- DOI
- 10.1136/bmj-2023-078918
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/81c00a06-e1f5-49cf-a499-c32e1abe67e9 is authoritative.
How this rating was calculated
- StatisticsImpossible or misreported statistic ×2−2★
- IntegrityIntegrity concern ×2−1★
- CitationsUnresolved reference ×3−0.75★
- StatisticsStatistic did not reproduce−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×6−0.25★
A demonstrable critical failure caps the rating at the minimum, regardless of the deductions above.
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 2 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Printed percentage does not match its own countdemonstrable
28.4% does not match the reported count 655/2262
“655 (28.4)”
Table 2 - 02Printed percentage does not match its own countdemonstrable
5.8% does not match the reported count 114/2262
“114 (5.8)”
Table 2 - 03Reported statistic does not recompute
Recomputed odds ratio 0.71 (95% CI 0.57–0.87), reported p<0.001
“odds ratio 0.71, 95% CI 0.57 to 0.87; P<0.001”
- 04Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is self-reported respiratory symptoms consistent with a respiratory infection, which is a surrogate for actual infection. The paper does not provide evidence that this surrogate is validated to predict hard clinical outcomes, nor does it demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure).
“The primary outcome was self-reported respiratory symptoms consistent with a respiratory infection.”
- 05Treatment effect not shown to be clinically meaningful
The absolute risk reduction is 3.2% (from 12.2% to 8.9%), which is a small absolute effect. The paper does not anchor this to a minimal clinically important difference or demonstrate that this magnitude is clinically meaningful.
“The absolute risk difference was −3.2% (95% CI −5.2% to −1.3%; P<0.001).”
- 06Printed percentage does not match its own count
50.3% does not match the reported count 1133/2262
“1133 (50.3)”
Table 2
5 further findings of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This pragmatic randomised trial is methodologically robust, with clear randomisation, blinding of researchers, a pre-specified power analysis, and transparent reporting including trial registration and data/code sharing. Minor copyedit issues (typos, truncated percentages, missing citation) and three references not found in registries are the main concerns for a post-publication audit.
Both reviewers classified the study as interventional, which is adopted. The evaluation covers all eight rigor dimensions; no dimensions were excluded as not applicable. The statistics verification component checked 14 tests, finding minor rounding discrepancies but no decision errors.
Numerical inconsistencies
4 findings · worst criticalValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Reported statistics do not recomputeRecomputed
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 6 tests: 5 consistent, 1 inconsistent; 5 recomputed directly from the reported test statistics, 1 via agent-written checks. 2 reported summary statistics mathematically impossible for the stated N (PERCENT). 6 printed percentages that do not match their own count.
- PERCENT50.3% does not match the reported count 1133/2262
“1133 (50.3)”
Table 2 - PERCENT13.2% does not match the reported count 295/2262
“295 (13.2)”
Table 2 - PERCENT26.2% does not match the reported count 598/2313
“598 (26.2)”
Table 2 - PERCENT42.7% does not match the reported count 994/2313
“994 (42.7)”
Table 2 - PERCENT28.4% does not match the reported count 655/2262
“655 (28.4)”
Table 2 - PERCENT5.8% does not match the reported count 114/2262
“114 (5.8)”
Table 2 - PERCENT63.7% does not match the reported count 1447/2262
“1447 (63.7)”
Table 2 - PERCENT34.8% does not match the reported count 780/2262
“780 (34.8)”
Table 2
- CONSISTENTreported p = .820 · recomputed p = .829Recomputed odds ratio 1.07 (95% CI 0.58–1.98), reported p=0.82
“odds ratio 1.07, 95% CI 0.58 to 1.98; P=0.82”
Taken as given: 0.58–1.98 is a two-sided 95% confidence interval for the odds ratio of 1.07, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.82 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.07, 0.58, 1.98, 1) - CONSISTENTreported p = .001 · recomputed p = <.001Recomputed odds ratio 0.71 (95% CI 0.58–0.87), reported p=0.001
“odds ratio 0.71, 95% CI 0.58 to 0.87; P=0.001”
Taken as given: 0.58–0.87 is a two-sided 95% confidence interval for the odds ratio of 0.71, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.71, 0.58, 0.87, 1) - INCONSISTENTreported p < .001 · recomputed p = .001Recomputed odds ratio 0.71 (95% CI 0.57–0.87), reported p<0.001
“odds ratio 0.71, 95% CI 0.57 to 0.87; P<0.001”
Taken as given: 0.57–0.87 is a two-sided 95% confidence interval for the odds ratio of 0.71, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p<0.001 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.71, 0.57, 0.87, 1) - CONSISTENTreported p = .006 · recomputed p = .006Recomputed odds ratio 0.76 (95% CI 0.62–0.92), reported p=0.006
“odds ratio 0.76, 95% CI 0.62 to 0.92; P=0.006”
Taken as given: 0.62–0.92 is a two-sided 95% confidence interval for the odds ratio of 0.76, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.006 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(0.76, 0.62, 0.92, 1) - CONSISTENTreported p = .970 · recomputed p = 1.000Recomputed odds ratio 1.00 (95% CI 0.81–1.22), reported p=0.97
“odds ratio 1.00, 95% CI 0.81 to 1.22; P=0.97”
Taken as given: 0.81–1.22 is a two-sided 95% confidence interval for the odds ratio of 1.00, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.97 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1, 0.81, 1.22, 1) - CONSISTENTreported p = .970 · recomputed p = 1.000Reviewers 1, 2Secondary outcome sick leave odds ratio p-value from CI
“Self-reported sick leave (complete case analysis) | 194/1630 (11.9) | 208/1740 (11.9) | 1.00 (0.81 to 1.22) | 0.97”
Taken as given: The odds ratio is 1.00 with 95% CI 0.81 to 1.22.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed two-tailed p-value from the reported odds ratio and 95% CI using the normal approximation for the log odds ratio.How we recomputed it: pCI(1.00, 0.81, 1.22, 1)
- lowinternal contradictionThe text states 'No statistically significant effect was found on self-reported covid-19' with OR 1.07, but the absolute risk difference is reported as 0.1% with CI -6.0 to 8.0, which is consistent.
“No statistically significant effect was found on self- reported (marginal odds ratio 1.07, 95% CI 0.58 to 1.98; P=0.82)”
AbstractFind in source - lowinternal contradictionThe abstract reports 4647 randomised, but the results section reports 5086 read consent and 4647 consented/randomised. This is consistent, but the flow diagram is not shown in text.
4647 adults aged ≥18 years: 2371 were assigned to the intervention arm and 2276 to the control arm. ... 5086 individuals read the consent form. Of these, 4647 (91%) provided consent, completed the baseline form, and were randomised
Abstractreviewer’s wording
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
6 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Wearing a surgical face mask in public spaces over 14 days reduces the risk of self-reported respiratory symptoms.The primary outcome analysis shows a statistically significant reduction with OR 0.71 (95% CI 0.58-0.87), supported by sensitivity analyses.Evidence: Primary outcome: OR 0.71, 95% CI 0.58-0.87, P=0.001; absolute risk difference -3.2%.
“Wearing a surgical face mask in public spaces over 14 days reduces the risk of self-reported symptoms consistent with a respiratory infection, compared with not wearing a surgical face mask.”
ConclusionFind in source - supportedReviewers 1, 2The effect size was moderate.The absolute risk reduction of 3.2% is moderate, as stated.Evidence: Absolute risk difference -3.2% (95% CI -5.2% to -1.3%).
“the effect size was moderate.”
Discussion ¶1Find in source - supportedReviewers 1, 2Wearing face masks in public spaces was safe and generally well tolerated.Adverse effects were reported by 3.4% of participants, mostly minor, supporting the claim.Evidence: 155 participants (3.4%) reported adverse effects, with most being unpleasant comments.
“Wearing face masks in public spaces was safe and generally well tolerated.”
Discussion ¶1Find in source - supportedReviewer 1Our findings provide a more precise estimate of effect compared with earlier face mask trials.The confidence interval is narrower than the Danish trial's, supporting the claim.Evidence: OR 0.71 (95% CI 0.58-0.87) vs Danish trial OR 0.82 (95% CI 0.54-1.23).
“Compared with the earlier face mask trials, our findings provide a more precise estimate of effect.”
DiscussionFind in source - supportedReviewer 1The results support the claim that face masks may be an effective measure to reduce the incidence of self-reported respiratory symptoms.The primary outcome and sensitivity analyses support this claim, though the outcome is self-reported.Evidence: Primary outcome OR 0.71 (95% CI 0.58-0.87).
“The results support the claim that face masks may be an effective measure to reduce the incidence of self-reported respiratory symptoms consistent with respiratory tract infections”
Discussion ¶1Find in source - supportedReviewer 2The trial was sufficiently powered.The trial enrolled 4647 participants, exceeding the required 2692, and the primary outcome was statistically significant.Evidence: Power calculation required 2692; actual enrolment was 4647.
“Unlike most earlier trials of face mask, our study was sufficiently powered”
What this study addsFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is self-reported respiratory symptoms consistent with a respiratory infection, which is a surrogate for actual infection. The paper does not provide evidence that this surrogate is validated to predict hard clinical outcomes, nor does it demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure).
“The primary outcome was self-reported respiratory symptoms consistent with a respiratory infection.”
- INADEQUATEEffect sizeThe absolute risk reduction is 3.2% (from 12.2% to 8.9%), which is a small absolute effect. The paper does not anchor this to a minimal clinically important difference or demonstrate that this magnitude is clinically meaningful.
“The absolute risk difference was −3.2% (95% CI −5.2% to −1.3%; P<0.001).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites systematic reviews of observational studies and a Cochrane review of randomised trials, acknowledging the discrepant findings and methodological limitations. The rationale for the trial follows logically from the identified evidence gaps, and the study design addresses prior limitations such as insufficient power and low adherence.
“Systematic reviews of observational studies have reported an association between wearing face masks and lower risk of respiratory infections. On the basis of findings from 10 randomised trials, however, the authors of a recent Cochrane review concluded that use of a face mask in the community had little or no effect on risk of developing a respiratory viral infection.”
“Several factors could explain the seemingly discrepant findings from observational studies and randomised trials, including the higher risk of bias inherent to observational studies, insufficient power of the randomised controlled trials, or low adherence to the intervention.”
“Systematic reviews of observational studies have reported an association between wearing face masks and lower risk of respiratory infections. On the basis of findings from 10 randomised trials, however, the authors of a recent Cochrane review concluded that use of a face mask in the community had little or no effect on risk of developing a respiratory viral infection.”
“Several factors could explain the seemingly discrepant findings from observational studies and randomised trials, including the higher risk of bias inherent to observational studies, insufficient power of the randomised controlled trials, or low adherence to the intervention.”
Randomisation used a computer-generated pseudorandom sequence via an independent web tool, with 1:1 allocation. Blinding of participants was not possible (open-label), but researchers and statistician were blinded. A priori power calculation was provided. Inclusion criteria were defined, and no exclusion criteria were applied. Missing data handling was pre-specified with multiple imputation and sensitivity analyses. The trial is a single pivotal trial, so independent replication is not applicable.
“We used Nettskjema, an independent web based survey tool, to randomise participants using a computer generated pseudorandom sequence over which we had no influence.”
“The researchers and study statistician were blinded to intervention allocation throughout the trial, and all main analyses were performed blinded.”
“We calculated that a minimum of 2692 participants (1346 in each arm) would be required to detect a risk reduction of 30% from an assumed 10% risk of infection in the control arm to a 7% risk in the intervention arm, with a two sided α of 0.05 (significance criterion) and 80% power.”
“We used Nettskjema, an independent web based survey tool, to randomise participants using a computer generated pseudorandom sequence over which we had no influence.”
“The researchers and study statistician were blinded to intervention allocation throughout the trial, and all main analyses were performed blinded.”
“We calculated that a minimum of 2692 participants (1346 in each arm) would be required to detect a risk reduction of 30% from an assumed 10% risk of infection in the control arm to a 7% risk in the intervention arm, with a two sided α of 0.05 (significance criterion) and 80% power.”
Sex, age, and various health-related characteristics (vaccination status, household composition, etc.) are reported in baseline tables. Since both sexes were enrolled, sex justification is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“Female sex | 1423 (61.5) | 1365 (60.3) | | Mean (SD) age (years) | 51 (15.2) | 51 (16.3)”
“Female sex | 1423 (61.5) | 1365 (60.3) | | Mean (SD) age (years) | 51 (15.2) | 51 (16.3)”
The study was approved by a named ethics committee (Regional Ethics Committee South East Norway, reference 36544). Informed consent was obtained via an online consent form. Compliance with the Declaration of Helsinki is stated.
“This study was approved by the Regional Ethics Committee South East Norway (reference 36544).”
“provide written informed consent (online consent form)”
“The trial was performed according to a published protocol, with exceptions (see Protocol Amendments section), and the principles outlined in the Declaration of Helsinki.”
“This study was approved by the Regional Ethics Committee South East Norway (reference 36544).”
“provide written informed consent (online consent form)”
The surgical face mask is described as 'three ply, disposable, surgical face masks (type II/IIR, compliant with the EN 14683 standard)' and provided by pharmacies. Statistical software R version 4.2.2 is identified. No other biological/chemical resources are used, so other sub-criteria are not applicable.
“These participants collected a pack of 50 three ply, disposable, surgical face masks (type II/IIR, compliant with the EN 14683 standard) from their nearest pharmacy”
“We conducted all analyses using R version 4.2.2.”
“three ply, disposable, surgical face masks (type II/IIR, compliant with the EN 14683 standard)”
“We conducted all analyses using R version 4.2.2.”
The primary analysis used unadjusted logistic regression to estimate marginal odds ratios, with multiple imputation for missing data. Tests are named, and effect sizes with 95% CIs are reported. Exact p-values are given for primary and secondary outcomes. Software is identified. Data presentation includes per-group n and percentages. Mathematical plausibility checks were not applicable due to large N and continuous outcomes.
“We estimated marginal odds ratios using unadjusted logistic regression for all outcomes following the intention-to-treat principle.”
“odds ratio 0.71, 95% CI 0.58 to 0.87; P=0.001”
“absolute risk difference −3.2%, 95% CI −5.2% to −1.3%; P<0.001”
“We estimated marginal odds ratios using unadjusted logistic regression for all outcomes following the intention-to-treat principle.”
“odds ratio 0.71, 95% CI 0.58 to 0.87; P=0.001”
The data availability statement provides a specific GitHub repository for the anonymised dataset and statistical codes. This is a concrete access route, satisfying the criterion. Repository deposit and accession numbers are not applicable for patient-level data, but the GitHub link serves as a repository.
“The final anonymised trial dataset and statistical codes will be freely available to the public through GitHub ( https://github.com/folkehelseinstituttet/2024-facemask-trial-bmj ).”
“The final anonymised trial dataset and statistical codes will be freely available to the public through GitHub ( https://github.com/folkehelseinstituttet/2024-facemask-trial-bmj ).”
The trial is registered (ClinicalTrials.gov NCT05690516). CONSORT guidelines are followed. All pre-specified outcomes are reported, including non-significant ones. Limitations are thoroughly discussed. Conclusions are proportional to the evidence. Funding and competing interests are declared.
“Trial registration ClinicalTrials.gov NCT05690516 (https://clinicaltrials.gov/ct2/show/NCT05690516)”
“We followed the Consolidated Standards of Reporting Trials (CONSORT) guidelines (see supplementary material, table 1).”
“Our trial has several limitations. Firstly, outcome data were missing for 13.7% and 20.7% of the participants in the control arm and intervention arm, respectively.”
“Trial registration ClinicalTrials.gov NCT05690516”
“We followed the Consolidated Standards of Reporting Trials (CONSORT) guidelines”
Registered (1 ID: ClinicalTrials.gov). Reporting guideline cited: CONSORT.
Broken references and links
1 finding · worst lowReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
- References not resolvable to a published paperRecomputed
Checked 24 references by DOI: 16 verified — 3 DOI unresolved, 5 no DOI (shown, not verified).
- UNRESOLVED10.5281/zenodo.7648390Study protocol: The protective effect of face mask wearing against respiratory tract infections: a pragmatic randomized trialCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.5281/zenodo.7984532Statistical analysis plan: The effect of face mask wearing against respiratory tract infections - a pragmatic randomised trialCited DOI does not resolve to any Crossref record.
- UNRESOLVED10.5281/zenodo.8089320Blinded assessment of face-masks study resultsCited DOI does not resolve to any Crossref record.
- NO DOICOVID-19 DASHBOARD 2023No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICoronavirus disease (COVID-19): MasksNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINettskjemaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMultiple imputation for nonresponse in surveysNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPartial identification of probability distributionsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- codeGitHubLIVEHTTP 200https://github.com/folkehelseinstituttet/2024-facemask-trial-bmjResolves to GitHub (code repository).
Copyediting
7 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 7 minor suggestions below.
7 copyedit issues flagged: mostly consistency, typo, clarity.
- MINORtypoAbstract, Results“self- reported (marginal odds ratio 1.07, 95% CI 0.58 to 1.98; P=0.82)”→ Remove space: 'self-reported'Minor spacing issue.
- MINORconsistencyTable 2, 'No of close daily contacts at work'“1-4 | 293 (12.6) | 306 (13.)”→ Change '13.' to '13.5' or '13.5%' for consistency.Percentage appears truncated.
- MINORclarityMethods, Statistical analysis“In another non-prespecified analysis, we also considered three less extreme scenarios using a method similar to the mean score method suggested by White et al ().”→ Add the citation for White et al.Citation placeholder is empty.
- MINORconsistencyTable 3, footnote“Scenario 2§§ j”→ Remove stray 'j'.Typographical error.
- MINORtypoAbstract, Results“self- reported”→ self-reportedMissing hyphen.
- MINORconsistencyTable 2“1-4 | 293 (12.6) | 306 (13.)”→ 13.5Incomplete percentage.
- MINORclarityMethods, Statistical analysis“We also performed complete case analyses, including participants with complete data at baseline and follow-up.”→ We also performed complete case analyses, including participants with complete data at baseline and follow-up.Redundant phrasing.
The published paper is robust overall. An informed reader should note the three references not found in registries (potential fabrication signals) and the minor copyedit issues, but these do not undermine the core findings. No erratum or re-analysis is warranted based on this audit.
- 1.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 28.4% does not match the reported count 655/2262Demonstrable critical failure — blocks the verdict from passing.
- 2.CRITICALstatisticsCorrect or explain the statistically impossible value: PERCENT: 5.8% does not match the reported count 114/2262Demonstrable critical failure — blocks the verdict from passing.
- 3.HIGHreportingVerify the three references not found in any registry: (1) 'Study protocol: The protective effect of face mask wearing against respiratory tract infections: a pragmatic randomized trial' (DOI 10.5281/zenodo.7648390), (2) 'Statistical analysis plan: The effect of face mask wearing against respiratory tract infections - a pragmatic randomised trial' (DOI 10.5281/zenodo.7984532), and (3) 'Blinded assessment of face-masks study results' (DOI 10.5281/zenodo.8089320). If they are legitimate, ensure they are correctly indexed; if not, correct or remove them.References that cannot be located in any registry are a fabrication signal and must be resolved for integrity.
- 4.MEDIUMcopyeditFix the truncated percentage in Table 2: change '306 (13.)' to '306 (13.5)' or the correct value.Incomplete data in a table undermines trust in reporting accuracy.
- 5.MEDIUMcopyeditAdd the missing citation for 'White et al.' in the Methods, Statistical analysis section where the mean score method is described.An empty citation placeholder is a reporting gap that should be filled.
- 6.MEDIUMcopyeditRemove the stray 'j' in Table 3 footnote: 'Scenario 2§§ j' should be 'Scenario 2§§'.Typographical errors reduce professionalism and clarity.
- 7.MEDIUMcopyeditFix the spacing in 'self- reported' in the Abstract Results to 'self-reported'.Minor typo that should be corrected for consistency.
- 8.LOWreportingConsider adding a statement on whether the trial was conducted in accordance with ICH-GCP or other regulatory standards, beyond the Declaration of Helsinki.Strengthens regulatory compliance reporting for a clinical trial.
- 9.LOWdata codeConsider specifying the license under which the dataset and code are shared on GitHub, and provide a DOI for the repository.A persistent identifier and license improve long-term accessibility and reuse.
- 10.LOWreportingReport the exact number of participants screened but not randomised (e.g., those who read consent but did not consent) in the CONSORT flow diagram.Full transparency in participant flow is a CONSORT recommendation.
- 11.LOWreportingProvide a brief justification for the lack of a data monitoring committee, as this is mentioned but not explained.Clarifies the trial's oversight structure.
- 12.LOWreportingAdd a note on whether any adverse events were serious or required medical attention, as only general adverse effects are described.Provides a complete safety profile.
- 13.LOWreportingConsider reporting the number of participants who used public transport or attended events as a potential confounder in the analysis.Addresses a potential source of confounding in a pragmatic trial.
- 14.LOWreportingProvide a more detailed description of the multiple imputation model, including the variables included and the number of imputations.Enhances reproducibility of the primary analysis.
- 15.LOWreportingConsider reporting the results of the adjusted analysis in the main text, as it is only mentioned in the supplementary material.Key sensitivity analyses are more visible in the main text.
- 16.LOWdata codeAdd a statement about the availability of the statistical analysis plan or protocol in the data availability statement.Provides a complete record of pre-specified analyses.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.