Clinical effectiveness of an online supervised group physical and mental health rehabilitation programme for adults with post-covid-19 condition (REGAIN study): multicentre randomised controlled trial.
McGregor G, Sandhu H, Bruce J, Sheehan B, McWilliams D, Yeung J, Jones C, Lara B, Alleyne S, Smith J, Lall R, Ji C, Ratna M, Ennis S, Heine P, Patel S, Abraham C, Mason J, Nwankwo H, Nichols V, Seers K, Underwood M
- DOI
- 10.1136/bmj-2023-076506
- Record issued
- 2026-08-16
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/01835a6b-820a-4aca-ab63-a6c1233fa8ed is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×2−1★
- StatisticsStatistic did not reproduce−0.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- StatisticsPrinted percentage does not match its own count (capped) ×12−0.25★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 32 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- No data or code availability links were detected to verify.
- 01Printed percentage does not match its own count
44% does not match the reported count 108/220
“108 (44)”
Table 3 - 02Efficacy rests on an unvalidated surrogate endpoint
The primary endpoint is the PROMIS-PROPr score, a patient-reported outcome measure (a surrogate for health-related quality of life). The paper does not demonstrate target engagement (PK/PD or dose-exposure) at the tested dose, nor does it cite validated evidence linking the PROMIS-PROPr score to a hard clinical outcome. The claim of clinical effectiveness rests on this surrogate without meeting both required conditions.
“The primary outcome was health related quality of life using the patient reported outcomes measurement information system (PROMIS) preference (PROPr) score at three months.”
- 03Treatment effect not shown to be clinically meaningful
The primary effect size is an adjusted mean difference of 0.03 (95% CI 0.01 to 0.05) in PROPr score, which the authors themselves note is smaller than the suggested minimally important difference of 0.04. The effect is not anchored to a clinically meaningful threshold; the authors rely on post-hoc analyses (CACE and NNT) to argue for clinical importance, but the primary result is below the established MCID.
“Our observed differences of 0.03 (95% confidence interval 0.01 to 0.05) at three months and 0.03 (0.01 to 0.06) at 12 months are smaller than this suggestion.”
- 04Printed percentage does not match its own count
88% does not match the reported count 508/585
“508/585 (88%)”
- 05Printed percentage does not match its own count
88% does not match the reported count 508/585
“508/585; 88%”
ResultsFind in source - 06Printed percentage does not match its own count
34% does not match the reported count 77/212
“77 (34)”
Table 3
9 further findings of this severity or below — every one is in the sections below, filed under its error type.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported pragmatic RCT of an online rehabilitation programme for post-covid-19 condition. The main methodological strength is the rigorous design with adequate randomization, blinding of outcome assessment, and pre-specified analyses. The primary weakness is the vague data availability statement and lack of code sharing, which limits reproducibility.
Both reviewers independently scored all eight dimensions and agreed on every status; no divergence required reconciliation. The statistics verification component covered only a subset of reported tests (those with test statistics/df or effect estimates with CIs); 13 of 17 tests were not machine-verifiable, so the absence of detected errors does not confirm overall statistical correctness. The citation check found no retracted or unresolved references.
Numerical inconsistencies
3 findings · worst highValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Summary statistic impossible for the stated N (GRIM/GRIMMER)Recomputed
- Printed percentage does not match its own countRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 3 tests: 3 consistent, 0 inconsistent; 1 recomputed directly from the reported test statistics, 2 via agent-written checks. 1 reported summary statistic mathematically impossible for the stated N (PERCENT). 12 printed percentages that do not match their own count.
- PERCENT88% does not match the reported count 508/585
“508/585 (88%)”
- PERCENT88% does not match the reported count 508/585
“508/585; 88%”
ResultsFind in source - PERCENT34% does not match the reported count 77/212
“77 (34)”
Table 3 - PERCENT38% does not match the reported count 85/214
“85 (38)”
Table 3 - PERCENT29% does not match the reported count 65/206
“65 (29)”
Table 3 - PERCENT35% does not match the reported count 78/216
“78 (35)”
Table 3 - PERCENT17% does not match the reported count 39/216
“39 (17)”
Table 3 - PERCENT8% does not match the reported count 20/220
“20 (8)”
Table 3 - PERCENT34% does not match the reported count 81/216
“81 (34)”
Table 3 - PERCENT24% does not match the reported count 60/220
“60 (24)”
Table 3 - PERCENT30% does not match the reported count 72/216
“72 (30)”
Table 3 - PERCENT44% does not match the reported count 108/220
“108 (44)”
Table 3 - PERCENT11% does not match the reported count 27/220
“27 (11)”
Table 3
- CONSISTENTreported p = .010 · recomputed p = .008Recomputed odds ratio 1.66 (95% CI 1.14–2.41), reported p=0.01
“odds ratio 1.66, 95% confidence interval 1.14 to 2.41; P=0.01”
Taken as given: 1.14–2.41 is a two-sided 95% confidence interval for the odds ratio of 1.66, not a range, an IQR, or a different interval level; the odds ratio is a RATIO measure, so the interval is symmetric on the log scale; p=0.01 is the p for THIS estimate, not for another comparison in the same sentenceMethod: back the two-tailed p out of the log-scale CI width and compare it against the printed pHow we recomputed it: pCI(1.66, 1.14, 2.41, 1) - UNCOMPUTABLEreported p = .020 · recomputed p = .003Reviewers 1, 2Primary outcome adjusted mean difference at 3 months
“adjusted mean difference in PROPr score 0.03 (95% confidence interval 0.01 to 0.05), P=0.02”
Taken as given: The estimate is 0.03 and the 95% CI is 0.01 to 0.05.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(0.03, 0.01, 0.05, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Secondary outcome: fatigue subscore at 3 months
“fatigue (2.50 (1.19 to 3.81), P<0.001)”
Taken as given: The estimate is 2.50 and the 95% CI is 1.19 to 3.81.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(2.50, 1.19, 3.81, 0) - CONSISTENTreported p = .010 · recomputed p = .007Reviewers 1, 2Secondary outcome: pain interference subscore at 3 months
“pain interference (1.80 (0.50 to 3.11), P=0.01)”
Taken as given: The estimate is 1.80 and the 95% CI is 0.50 to 3.11.; The CI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed p-value from the estimate and 95% CI using the normal approximation.How we recomputed it: pCI(1.80, 0.50, 3.11, 0)
- lowinternal contradictionThe number of participants with primary outcome data at 3 months is 237+248=485, but the total randomised is 585. The difference (100) is not explained in the text, but likely due to loss to follow-up. This is not a contradiction.
“Primary outcome data were collected from 237/298 (80%) in the REGAIN intervention group and 248/287 (86%) participants in the usual care group”
Results ¶2Find in source - lowinternal contradictionThe abstract states '585 adults (26-86 years)' but the Methods state 'Participants were adults (26-86 years)' - consistent. No contradiction found.
“585 adults (26-86 years)”
AbstractFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewers 1, 2The intervention is clinically effective.The effect size (0.03) is below the suggested minimal important difference of 0.04, but the CACE analysis shows a larger effect. The claim of clinical effectiveness is partially supported.Evidence: Adjusted mean difference 0.03 (95% CI 0.01 to 0.05) at 3 months; CACE analysis 0.05 (0.01 to 0.09).
“Our observed differences of 0.03 (95% confidence interval 0.01 to 0.05) at three months and 0.03 (0.01 to 0.06) at 12 months are smaller than this suggestion.”
Discussion ¶2Find in source - supportedReviewers 1, 2The REGAIN intervention improves health-related quality of life at 3 months compared with usual care.The primary outcome analysis shows a statistically significant adjusted mean difference of 0.03 (95% CI 0.01 to 0.05, P=0.02), supporting the claim.Evidence: Adjusted mean difference in PROPr score 0.03 (95% CI 0.01 to 0.05), P=0.02 at 3 months.
“Compared with usual care, the REGAIN intervention led to improvements in health related quality of life (adjusted mean difference in PROPr score 0.03 (95% confidence interval 0.01 to 0.05), P=0.02) at three months”
AbstractFind in source - supportedReviewers 1, 2The effect is sustained at 12 months.The 12-month analysis shows a similar adjusted mean difference of 0.03 (95% CI 0.01 to 0.06, P=0.02), supporting the claim.Evidence: Adjusted mean difference in PROPr score 0.03 (95% CI 0.01 to 0.06), P=0.02 at 12 months.
“Effects were sustained at 12 months (0.03 (0.01 to 0.06), P=0.02).”
AbstractFind in source - supportedReviewer 1The intervention is safe, with only one serious adverse event possibly related.The paper reports 21 serious adverse events, with only one possibly related to the intervention, supporting the claim.Evidence: Of 21 serious adverse events, only one was possibly related to the REGAIN intervention.
“Of 21 serious adverse events, only one was possibly related to the REGAIN intervention.”
AbstractFind in source - supportedReviewer 2The intervention was safe.Only one serious adverse event was possibly related to the intervention, and no post-exertional symptom exacerbation was observed, supporting the safety claim.Evidence: Adverse events data: 21 serious adverse events, only one possibly related; no post-exertional symptom exacerbation.
“Of the 21 serious adverse events, 19 concerned admission to hospital or prolongation of admission, and two involved persistent or major disability or incapacity. Only one serious adverse event (syncope with vomiting 24 hours after a live exercise session) was possibly related to the REGAIN intervention.”
Results ¶7Find in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary endpoint is the PROMIS-PROPr score, a patient-reported outcome measure (a surrogate for health-related quality of life). The paper does not demonstrate target engagement (PK/PD or dose-exposure) at the tested dose, nor does it cite validated evidence linking the PROMIS-PROPr score to a hard clinical outcome. The claim of clinical effectiveness rests on this surrogate without meeting both required conditions.
“The primary outcome was health related quality of life using the patient reported outcomes measurement information system (PROMIS) preference (PROPr) score at three months.”
- INADEQUATEEffect sizeThe primary effect size is an adjusted mean difference of 0.03 (95% CI 0.01 to 0.05) in PROPr score, which the authors themselves note is smaller than the suggested minimally important difference of 0.04. The effect is not anchored to a clinically meaningful threshold; the authors rely on post-hoc analyses (CACE and NNT) to argue for clinical importance, but the primary result is below the established MCID.
“Our observed differences of 0.03 (95% confidence interval 0.01 to 0.05) at three months and 0.03 (0.01 to 0.06) at 12 months are smaller than this suggestion.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites epidemiological data on long covid prevalence and symptoms, and notes that only small quasi-experimental studies have investigated exercise-based rehabilitation, with no high-quality definitive evidence. The rationale for a multicomponent physical and mental health rehabilitation programme is logically linked to the biopsychosocial model and evidence from other long-term conditions. The paper does not explicitly discuss limitations of prior research in detail, but the gap in evidence is clearly stated.
“Across the World Health Organization European Region during the first two years of the covid-19 pandemic, more than 17 million people may have experienced covid-19 symptoms lasting more than four weeks.”
“The biopsychosocial model of care may contribute to improved outcomes for people with post-covid-19 condition. Multicomponent physical and mental health rehabilitation can improve breathlessness, fatigue, and quality of life in other long term conditions.”
“To date, only small quasi-experimental studies have investigated exercise based rehabilitation interventions for people with post-covid-19 condition, and no high quality definitive evidence exists as to the potential benefits or harms of physical and mental health rehabilitation interventions.”
“Across the World Health Organization European Region during the first two years of the covid-19 pandemic, more than 17 million people may have experienced covid-19 symptoms lasting more than four weeks.”
“To date, only small quasi-experimental studies have investigated exercise based rehabilitation interventions for people with post-covid-19 condition, and no high quality definitive evidence exists as to the potential benefits or harms of physical and mental health rehabilitation interventions.”
“The biopsychosocial model of care may contribute to improved outcomes for people with post-covid-19 condition. Multicomponent physical and mental health rehabilitation can improve breathlessness, fatigue, and quality of life in other long term conditions.”
Randomization used a centralised computer-generated sequence with minimisation, stratified by age, level of hospital care, and mental health symptomatology. Participants and practitioners were not masked, but outcome assessments were completed online or by staff blind to allocation. A sample size calculation was performed with 90% power and 5% type I error, accounting for clustering. Inclusion/exclusion criteria were pre-specified. Outlier handling is addressed through the pre-specified analysis population and multiple imputation for missing data. Controls are appropriate (usual care group). Independent replication is not applicable for a single pivotal trial.
“participants were randomly allocated (1:1.03) to the REGAIN intervention or to usual care by a centralised computer generated randomisation sequence using a bespoke web based system, administered independently by Warwick Clinical Trials Unit.”
“Follow-up outcome assessments were completed by participants online, or, in a small number of cases, over the telephone by a member of the trial team, blind to treatment allocation.”
“Assuming an intracluster coefficient of 0.01, 90% power, and type I error rate of 5%, with a 10% loss to follow-up, we determined that 535 participants would be required.”
“participants were randomly allocated (1:1.03) to the REGAIN intervention or to usual care by a centralised computer generated randomisation sequence using a bespoke web based system, administered independently by Warwick Clinical Trials Unit.”
“Follow-up outcome assessments were completed by participants online, or, in a small number of cases, over the telephone by a member of the trial team, blind to treatment allocation.”
“Assuming an intracluster coefficient of 0.01, 90% power, and type I error rate of 5%, with a 10% loss to follow-up, we determined that 535 participants would be required.”
The paper reports sex (52% female), age (mean 56 years), ethnicity (88% white), and various comorbidities. Age and sex are reported for the study population. Since both sexes are enrolled, sex_justified is not applicable. Demographics are adequately reported. Species/strain and housing conditions are not applicable for a human trial.
“Mean age of the study sample was 56 (SD 12) years, more than half were female participants (305/585; 52%)”
“most were of white ethnicity (517/585, 88%)”
“The most common pre-existing medical conditions related to chest or breathing (444/585; 76%) and musculoskeletal conditions (275/585; 47%)”
“Mean age of the study sample was 56 (SD 12) years, more than half were female participants (305/585; 52%)”
“most were of white ethnicity (517/585, 88%)”
“The most common pre-existing medical conditions related to chest or breathing (444/585; 76%) and musculoskeletal conditions (275/585; 47%)”
The study was approved by a named research ethics committee (East of England, Cambridge South Research Ethics Committee, reference 20/EE/0235) and the Health Research Authority. Informed consent is implied through the description of online consent procedures. Regulatory compliance is stated through adherence to the COPI regulations and good clinical practice.
“This study was approved by the East of England, Cambridge South Research Ethics Committee (reference 20/EE/0235) and the Health Research Authority/Health and Care Research Wales on 6 November 2020.”
“Participants then completed an online consent form and baseline outcomes questionnaire before randomisation.”
“The NHS Digital “Digi-Trials” service was approved to identify and invite potential participants in accordance with Regulation 3(4) of the Health Service (Control of Patient Information, COPI) Regulations 2002”
“This study was approved by the East of England, Cambridge South Research Ethics Committee (reference 20/EE/0235) and the Health Research Authority/Health and Care Research Wales on 6 November 2020.”
“Participants then completed an online consent form and baseline outcomes questionnaire before randomisation.”
“The NHS Digital “Digi-Trials” service was approved to identify and invite potential participants in accordance with Regulation 3(4) of the Health Service (Control of Patient Information, COPI) Regulations 2002”
The REGAIN intervention is described in detail, including the delivery platform (Zoom/Beam) and the workbook URL. Statistical software (Stata version 17 and R version 4) is identified. No antibodies, cell lines, or mycoplasma testing are applicable. The intervention is a complex behavioural programme, not a drug or device, but it is still scored as the investigational product.
“The REGAIN intervention comprised an eight week, online, home based, supervised, group rehabilitation programme (see supplementary figure S1), supported by a workbook for participants ( https://wrap.warwick.ac.uk ).”
“All analyses were conducted using Stata version 17 and R version 4.”
“delivered through Zoom using the Beam platform ( https://www.beamfeelgood.com )”
“supported by a workbook for participants ( https://wrap.warwick.ac.uk )”
“Participants in the usual care group received best practice usual care, consisting of a 30 minute, online, one-to-one consultation with a trained practitioner.”
The primary analysis used a partially nested heteroscedastic model, with adjustments for baseline and stratification variables. Tests are named (e.g., Mann-Whitney for normality violations). Exact p-values are reported for primary and secondary outcomes. Effect sizes with 95% confidence intervals are reported throughout. Statistical software is identified. Data presentation includes per-group n and confidence intervals. Mathematical plausibility checks were not possible for all values, but no obvious errors were found.
“For the primary outcome (PROPr score) we performed a partially nested heteroscedastic model to compare health related quality of life at three months between the REGAIN intervention group and usual care group, producing unadjusted and adjusted estimates.”
“adjusted mean difference in PROPr score 0.03 (95% confidence interval 0.01 to 0.05), P=0.02”
“depression (1.39 (0.06 to 2.71), P=0.04), fatigue (2.50 (1.19 to 3.81), P<0.001), and pain interference (1.80 (0.50 to 3.11), P=0.01)”
“For the primary outcome (PROPr score) we performed a partially nested heteroscedastic model to compare health related quality of life at three months between the REGAIN intervention group and usual care group, producing unadjusted and adjusted estimates.”
“adjusted mean difference in PROPr score 0.03 (95% confidence interval 0.01 to 0.05), P=0.02”
“All analyses were conducted using Stata version 17 and R version 4.”
The data availability statement says 'Data are available on reasonable request from wctudataaccess@warwick.ac.uk' but does not specify conditions or timeframe, which is inadequate. No repository deposit or accession numbers are provided, and no code sharing is mentioned. For a clinical trial, managed access is acceptable, but the statement lacks detail.
“Data are available on reasonable request from wctudataaccess@warwick.ac.uk .”
“Data are available on reasonable request from wctudataaccess@warwick.ac.uk .”
The trial is registered (ISRCTN11466448). Methods are detailed and reproducible. The paper follows CONSORT guidelines (implied by the flow diagram). All pre-specified outcomes are reported, including null results. Limitations are discussed in the discussion section. Conclusions are proportional to the evidence. Funding and competing interests are declared.
“Trial registration ISRCTN registry ISRCTN11466448.”
“Limitations include the inability of trial participants or practitioners delivering the intervention to be masked to treatment allocation.”
“This trial was funded by the UK National Institute for Health and Care Research Health Technology Assessment Programme.”
“Trial registration ISRCTN registry ISRCTN11466448.”
“Limitations include the inability of trial participants or practitioners delivering the intervention to be masked to treatment allocation.”
“This trial was funded by the UK National Institute for Health and Care Research Health Technology Assessment Programme.”
Registered (1 ID: ISRCTN). No reporting guideline cited.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 29 references by DOI: 24 verified — 5 no DOI (shown, not verified).
- NO DOIAt least 17 million people in the WHO European Region experienced long COVID in the first two years of the pandemic; millions may have to live with it for years to comeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPrevalence of ongoing symptoms following coronavirus (COVID-19) infection in the UK: 30 March 2023No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIYour COVID RecoveryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAssessing psychological trauma and PTSD - The Impact of Event Scale-RevisedNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntegrated addendum to ICH E6 (R1): guideline for good clinical practiceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoAbstract, Results“P<0.001”→ P<0.001Missing space after 'P' in some instances.
- MINORconsistencyTable 2, footnote“ITT=intention to treat; CACE=complier average causal effect”→ ITT=intention-to-treat; CACE=complier average causal effectHyphenation inconsistency.
- MINORclarityDiscussion, paragraph 3“The PROPr score for health related quality of life is calculated from seven PROMIS subscores, and in addition we measured four separate PROMIS subscales.”→ The PROPr score is calculated from seven PROMIS subscores; we also measured four separate PROMIS subscales.Slightly awkward phrasing.
- MINORtypoAbstract, Results“P<0.001”→ Ensure consistent spacing in p-values (e.g., P<0.001 vs P < 0.001).Minor formatting inconsistency.
- MINORconsistencyTable 2 footnote“ITT=intention to treat; CACE=complier average causal effect”→ Define all abbreviations consistently in each table footnote.Abbreviations are defined but could be standardized.
- MINORclarityMethods, Statistical analysis“We checked normality assumptions and used the Mann-Whitney test to test the treatment effect (unadjusted).”→ Clarify which outcome this applies to, as the primary analysis uses a mixed model.Potential ambiguity.
The published work is robust and well-reported; an informed reader should weigh the vague data availability statement and lack of code sharing as the main reproducibility limitations. No erratum or correction appears warranted based on the checks performed, though the authors could strengthen the data access statement and share analysis code to improve transparency.
- 1.HIGHdata codeIn the Data availability statement, specify the conditions and timeframe for data access, e.g., 'Data are available on reasonable request from wctudataaccess@warwick.ac.uk after approval of a proposal and with a signed data access agreement.'The current statement is vague and does not meet common reproducibility standards for clinical trials.
- 2.HIGHdata codeDeposit de-identified aggregate data or statistical analysis code in a public repository (e.g., Zenodo) with a DOI.Sharing data and code enhances reproducibility and is increasingly expected for funded trials.
- 3.MEDIUMreportingExplicitly state in the Methods that the study is reported following CONSORT guidelines and provide the checklist as supplementary material.Explicit reporting guideline adherence improves transparency and completeness.
- 4.MEDIUMreportingIn the Discussion, expand on the limitations of the PROPr score's minimal important difference, as the observed effect is below the suggested threshold.This contextualizes the clinical significance of the primary outcome.
- 5.MEDIUMstatisticsClarify in the Methods, Statistical analysis, which outcome the Mann-Whitney test applies to, as the primary analysis uses a mixed model.The current wording is ambiguous and could confuse readers about the analysis plan.
- 6.LOWcopyeditEnsure consistent spacing in p-values (e.g., P<0.001 vs P < 0.001) throughout the Abstract and Results.Minor formatting consistency improves professional presentation.
- 7.LOWcopyeditStandardize hyphenation in Table 2 footnote: change 'ITT=intention to treat' to 'ITT=intention-to-treat'.Consistency in abbreviations avoids confusion.
- 8.LOWcopyeditRephrase the Discussion sentence about PROPr score for clarity: 'The PROPr score is calculated from seven PROMIS subscores; we also measured four separate PROMIS subscales.'Improves readability and precision.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.