Provision of knee bracing for knee osteoarthritis (PROP OA): multicentre, parallel group, superiority, statistician blinded, randomised controlled trial.
Holden MA, Nicholls E, Abdali Z, Birrell F, Borrelli B, Callaghan M, Dziedzic K, Felson D, Foster NE, Halliday N, Ingram C, Jinks C, Jowett S, Peat G, PROP OA trial team
- DOI
- 10.1136/bmj-2025-086005
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/d461c528-c31d-4ce5-be28-cf29e3293309 is authoritative.
How this rating was calculated
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- Statistics were not checked: no recomputable values were found in this text — no test statistic reported with its degrees of freedom, no effect estimate printed with both a 95% CI and a p-value, and no percentage printed with both its count and its denominator.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is KOOS-5, a patient-reported composite score of pain, symptoms, activities of daily living, sport/recreation, and quality of life. This is a surrogate for clinical benefit. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure) and does not cite validated evidence linking KOOS-5 changes to hard clinical outcomes. The minimal clinically important difference (MCID) of 8 points is mentioned, but the observed effect (3.39) is below this threshold, and the paper does not provide a validated link between KOOS-5 and long-term clinical outcomes.
“The primary outcome was a composite patient reported Knee Osteoarthritis Outcomes Score (KOOS)-5 (0-100) at six months after randomisation.”
- 02Treatment effect not shown to be clinically meaningful
The primary effect size is an adjusted mean difference of 3.39 points on KOOS-5 (0-100 scale), with an effect size of 0.24. This is below the predefined minimal clinically important difference of 8 points, and the paper acknowledges the effect is 'small' and 'very small' at 12 months. The effect is not anchored to a clinically meaningful threshold; instead, the paper argues for clinical importance based on secondary outcomes and responder rates, but the primary outcome does not meet the MCID.
“The treatment effect for the primary outcome did not reach the predefined minimum clinically important difference of eight points on KOOS-5, used to inform our sample size calculation.”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported randomised controlled trial. The design is rigorous, with appropriate randomisation, blinding of the statistician, sample size calculation, and comprehensive reporting of demographics, ethics, resources, statistics, and data/code availability. Minor reporting gaps include lack of explicit CONSORT adherence and a few copyedit issues.
This is a post-publication audit of a published RCT. Both reviewers independently scored all eight dimensions and agreed on all statuses. The statistics verification component checked 0 tests (none recomputable), so no statistical errors were found, but this does not confirm correctness of unreported tests. The citation check found no retracted or unresolved references. The reproducibility check confirmed the data link is live.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
5 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Adding compartment specific knee bracing and an adherence intervention to advice, written information, and exercise instruction resulted in small improvements in patient reported outcomes among individuals with knee osteoarthritis.The primary outcome showed a statistically significant adjusted mean difference of 3.39 (95% CI 0.96 to 5.82) at six months, supporting the claim of small improvements.Evidence: Primary outcome result: adjusted mean difference 3.39, 95% CI 0.96 to 5.82; effect size 0.24.
“Adding compartment specific knee bracing and an adherence intervention to advice, written information, and exercise instruction resulted in small improvements in patient reported outcomes among individuals with knee osteoarthritis.”
ConclusionFind in source - supportedReviewers 1, 2This safe intervention offers a potential treatment option for this common condition.Adverse events were minor and expected, and the intervention was acceptable, supporting the safety and potential as a treatment option.Evidence: Adverse events were minor and expected; acceptability ratings were higher in the AIE+B group.
“This safe intervention offers a potential treatment option for this common condition.”
ConclusionFind in source - supportedReviewer 1The trial is the largest and provides more certainty that adding bracing leads to small additional benefits.The paper compares with previous trials and notes its larger sample size and longer follow-up, supporting the claim of added certainty.Evidence: Discussion states PROP OA is the largest trial with longer follow-up and good follow-up rates.
“Compared with previous randomised controlled trials, PROP OA is the largest, included a broad suite of outcomes, and followed participants over the longer term with good follow-up rates.”
DiscussionFind in source - supportedReviewer 2The largest effects observed were for pain reduction (KOOS pain adjusted mean difference at six months 6.13, 95% CI 3.36 to 8.91; effect size 0.39).The reported effect size and CI support the claim that pain reduction was the largest effect among secondary outcomes.Evidence: Table 4 shows KOOS pain adjusted mean difference 6.13 (95% CI 3.36 to 8.91) at six months, which is the largest among KOOS subscales.
“The largest effects observed were for pain reduction (KOOS pain (0-100) adjusted mean difference at six months 6.13, 95% CI 3.36 to 8.91; effect size 0.39).”
AbstractFind in source - supportedReviewer 2Secondary outcomes showed the benefits of AIE+B over AIE that diminished over time.The treatment effect at 12 months was smaller and no longer significant for the primary outcome, and several secondary outcomes showed non-significant differences at 12 months.Evidence: Primary outcome at 12 months: adjusted mean difference 2.67 (95% CI −0.24 to 5.57), not significant.
“Secondary outcomes showed the benefits of AIE+B over AIE that diminished over time.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is KOOS-5, a patient-reported composite score of pain, symptoms, activities of daily living, sport/recreation, and quality of life. This is a surrogate for clinical benefit. The paper does not demonstrate target engagement at the tested dose (e.g., PK/PD or dose-exposure) and does not cite validated evidence linking KOOS-5 changes to hard clinical outcomes. The minimal clinically important difference (MCID) of 8 points is mentioned, but the observed effect (3.39) is below this threshold, and the paper does not provide a validated link between KOOS-5 and long-term clinical outcomes.
“The primary outcome was a composite patient reported Knee Osteoarthritis Outcomes Score (KOOS)-5 (0-100) at six months after randomisation.”
- INADEQUATEEffect sizeThe primary effect size is an adjusted mean difference of 3.39 points on KOOS-5 (0-100 scale), with an effect size of 0.24. This is below the predefined minimal clinically important difference of 8 points, and the paper acknowledges the effect is 'small' and 'very small' at 12 months. The effect is not anchored to a clinically meaningful threshold; instead, the paper argues for clinical importance based on secondary outcomes and responder rates, but the primary outcome does not meet the MCID.
“The treatment effect for the primary outcome did not reach the predefined minimum clinically important difference of eight points on KOOS-5, used to inform our sample size calculation.”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites prior work and identifies gaps: 'Internationally, clinical guidelines offer conflicting recommendations on the use of knee bracing for knee osteoarthritis, and evidence from high quality randomised controlled trials is sparse.' The rationale links this to the trial's objective. Limitations of prior research are addressed in the discussion, noting previous trials were 'hampered by targeting only one knee compartment for all participants, small sample sizes, risk of bias, lack of follow-up beyond three months, and heterogeneity.'
“Internationally, clinical guidelines offer conflicting recommendations on the use of knee bracing for knee osteoarthritis, and evidence from high quality randomised controlled trials is sparse.”
“The Provision of Braces for Patients with Knee Osteoarthritis (PROP OA) trial was designed to answer the need for a large, independent, high quality randomised controlled trial with outcomes for >6 months.”
“Before PROP OA, randomised controlled trials of knee bracing for knee osteoarthritis were hampered by targeting only one knee compartment for all participants, small sample sizes, risk of bias, lack of follow-up beyond three months, and heterogeneity.”
“Internationally, clinical guidelines offer conflicting recommendations on the use of knee bracing for knee osteoarthritis, and evidence from high quality randomised controlled trials is sparse.”
“Before PROP OA, randomised controlled trials of knee bracing for knee osteoarthritis were hampered by targeting only one knee compartment for all participants, small sample sizes, risk of bias, lack of follow-up beyond three months, and heterogeneity.”
Randomisation used a 'computerised web based randomisation service and random number generator, stratified by clinic site, predominant compartmental distribution of knee osteoarthritis... and the presence of instability (buckling), with a 1:1 allocation with random permuted blocks of sizes 2, 4, and 6.' The unit is the individual participant. Blinding: 'the trial statistician was masked to treatment allocation' and the impossibility of masking participants/physiotherapists is stated. Power analysis: 'powered to detect an effect size between groups of 0.35... with two sided 5% significance and 90% power.' Inclusion/exclusion criteria are detailed. Outlier handling is addressed via the prespecified analysis population and multiple imputation for missing data. Controls: the AIE group serves as the comparator. Independent replication is not applicable for a single pivotal trial.
“We used a computerised web based randomisation service and random number generator, stratified by clinic site, predominant compartmental distribution of knee osteoarthritis (based on a combination of clinical assessment and radiographic presentation), and the presence of instability (buckling), with a 1:1 allocation with random permuted blocks of sizes 2, 4, and 6.”
“Although masking participants or physiotherapists to treatment allocation was not possible, the trial statistician was masked to treatment allocation.”
“The trial was powered to detect an effect size between groups of 0.35 (small-to-medium effect) in KOOS-5 at six months with two sided 5% significance and 90% power.”
“We used a computerised web based randomisation service and random number generator, stratified by clinic site, predominant compartmental distribution of knee osteoarthritis (based on a combination of clinical assessment and radiographic presentation), and the presence of instability (buckling), with a 1:1 allocation with random permuted blocks of sizes 2, 4, and 6.”
“Although masking participants or physiotherapists to treatment allocation was not possible, the trial statistician was masked to treatment allocation.”
“The trial was powered to detect an effect size between groups of 0.35 (small-to-medium effect) in KOOS-5 at six months with two sided 5% significance and 90% power.”
Sex is reported (46% female) and both sexes are enrolled, so sex_justified is not applicable. Age (mean 64, SD 9) and health status (comorbidities, BMI) are reported. Demographics include ethnicity, deprivation, and education. Species/strain and housing conditions are not applicable for a human trial.
“466 participants (mean age 64 (standard deviation 9) years; 46% female participants) were randomised”
“466 participants (mean age 64 (standard deviation 9) years; 46% female participants) were randomised”
The paper states: 'The study was approved by North West Preston Research Ethics Committee, the Health Research Authority, and Health and Care Research in Wales (research ethics committee reference 19/NW/0183; Integrated Research Application System reference 247370).' Informed consent is described: 'The physiotherapist obtained written informed consent from participants before collection of self-reported baseline data.' Regulatory compliance is covered by the HRA approval and adherence to ethical principles.
“The study was approved by North West Preston Research Ethics Committee, the Health Research Authority, and Health and Care Research in Wales (research ethics committee reference 19/NW/0183; Integrated Research Application System reference 247370).”
“The physiotherapist obtained written informed consent from participants before collection of self-reported baseline data”
“The study was approved by North West Preston Research Ethics Committee, the Health Research Authority, and Health and Care Research in Wales (research ethics committee reference 19/NW/0183; Integrated Research Application System reference 247370).”
“The physiotherapist obtained written informed consent from participants before collection of self-reported baseline data”
The braces are named: 'patellofemoral (Bioskin Q Brace), tibiofemoral unloading (first choice brace was Össur Unloader One and second choice brace was Donjoy Nano), or a neutral stabilising knee brace (Össur Formfit Knee Hinged).' Statistical software is identified: 'Data were analysed with Stata version 18.0.' Antibodies, cell lines, mycoplasma, and organisms are not applicable. Reagents are not applicable beyond the braces.
“Participants were then given a patellofemoral (Bioskin Q Brace), tibiofemoral unloading (first choice brace was Össur Unloader One and second choice brace was Donjoy Nano), or a neutral stabilising knee brace (Össur Formfit Knee Hinged).”
“Data were analysed with Stata version 18.0.”
“Participants were then given a patellofemoral (Bioskin Q Brace), tibiofemoral unloading (first choice brace was Össur Unloader One and second choice brace was Donjoy Nano), or a neutral stabilising knee brace (Össur Formfit Knee Hinged).”
“Data were analysed with Stata version 18.0.”
Tests are named: 'Longitudinal mixed models were used to estimate treatment effects... with results presented as adjusted mean differences and effect sizes, or adjusted odds ratios... with 95% confidence intervals.' Assumptions are addressed: 'All model assumptions were largely satisfied in the data, despite the raw scores for some outcomes not following a normal distribution.' Exact p-values are not reported; instead, effect estimates with 95% CIs are given, which is acceptable for estimation-based reporting. Effect sizes and CIs are reported. Software is identified. Data presentation includes per-group n and means with SDs. Mathematical plausibility is not applicable for large-N continuous outcomes.
“Longitudinal mixed models were used to estimate treatment effects for primary and secondary outcomes at the three, six, and, 12 month follow-up periods, with results presented as adjusted mean differences and effect sizes, or adjusted odds ratios (AIE+B v AIE) for continuous and categorical outcomes, respectively, with 95% confidence intervals (CIs).”
“All model assumptions were largely satisfied in the data, despite the raw scores for some outcomes not following a normal distribution.”
“Longitudinal mixed models were used to estimate treatment effects for primary and secondary outcomes”
“adjusted mean difference 3.39, 95% confidence interval (CI) 0.96 to 5.82; effect size 0.24”
“All model assumptions were largely satisfied in the data, despite the raw scores for some outcomes not following a normal distribution.”
The data availability statement provides a direct link: 'The data underlying the findings in this paper are openly and publicly available and can be found here: https://doi.org/10.21252/9tpn-8970.' Code is also shared: 'The code used to analyse the data in the paper can be found in the supplementary files.' Repository deposit is satisfied by the DOI. Accession numbers are not applicable for this type of data. Code sharing is adequate.
“The data underlying the findings in this paper are openly and publicly available and can be found here: https://doi.org/10.21252/9tpn-8970 .”
“The code used to analyse the data in the paper can be found in the supplementary files.”
“The data underlying the findings in this paper are openly and publicly available and can be found here: https://doi.org/10.21252/9tpn-8970 .”
“The code used to analyse the data in the paper can be found in the supplementary files.”
Trial registration is provided (ISRCTN28555470). Methods are comprehensive. Reporting guideline adherence is implied by the structured abstract and CONSORT-like flow diagram, though not explicitly named. All outcomes are reported, including negative results. Limitations are thoroughly discussed. Conclusions are proportional. Funding and competing interests are disclosed.
“Trial registration ISRCTN28555470.”
“The PROP OA trial had some limitations. Although our analyses were undertaken masked to treatment allocation, in response to the covid-19 pandemic, outcome data were collected for some participants over the telephone by an unblinded trial manager.”
“This study was funded by the National Institute of Health and Care Research (NIHR) Health Technology Assessment programme (16/160/03) and Keele University.”
“Trial registration ISRCTN28555470.”
“The PROP OA trial had some limitations. Although our analyses were undertaken masked to treatment allocation, in response to the covid-19 pandemic, outcome data were collected for some participants over the telephone by an unblinded trial manager.”
“This study was funded by the National Institute of Health and Care Research (NIHR) Health Technology Assessment programme (16/160/03) and Keele University.”
Registered (1 ID: ISRCTN). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 34 references by DOI: 30 verified — 4 no DOI (shown, not verified).
- NO DOIOsteoarthritis in over-16s: diagnosis and managementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIValidation study of WOMAC: a health status instrument for measuring clinically important patient relevant outcomes to antirheumatic drug therapy in patients with osteoarthritis of the hip or kneeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIOutcome variables for osteoarthritis clinical trials: The OMERACT-OARSI set of responder criteriaNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIStata Statistical Software: Release 18No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://doi.org/10.21252/9tpn-8970LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
3 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 3 minor suggestions below.
3 copyedit issues flagged: mostly typo, consistency, clarity.
- MINORtypoTable 5 footnote“past six months””→ past six monthsStray quotation mark.
- MINORconsistencyData availability statement“https://doi.org/10.21252/9tpn-8970 (10.1093/rheumatology/kew201)”→ Remove extraneous DOI or clarify its relevance.An extra DOI appears in the data availability statement that may confuse readers.
- MINORclarityAbstract, Results“401 (86%), 394 (85%), and 370 (79%) participants followed up with analysable data at three, six, and 12 months, respectively.”→ Consider rephrasing for clarity: 'follow-up data were available for 401 (86%), 394 (85%), and 370 (79%) participants at three, six, and 12 months, respectively.'Minor grammatical improvement.
The published work is robust and well-reported. An informed reader should weigh the minor reporting gaps (lack of explicit CONSORT adherence, extraneous DOI in data availability statement, and minor copyedit issues) as low-impact. No erratum or re-analysis is warranted based on the available evidence.
- 1.MEDIUMreportingAdd an explicit statement of adherence to CONSORT guidelines in the Methods or a dedicated reporting checklist section.Both reviewers noted the absence of an explicit CONSORT statement, which is a standard expectation for RCT reporting and would strengthen transparency.
- 2.MEDIUMdata codeRemove the extraneous DOI (10.1093/rheumatology/kew201) from the Data availability statement, or clarify its relevance.The copyedit pass flagged this as confusing; a clean data availability statement is important for reader trust.
- 3.MEDIUMcopyeditFix the stray quotation mark in the Table 5 footnote: change 'past six months”' to 'past six months'.Minor typo that should be corrected for professional presentation.
- 4.MEDIUMcopyeditRephrase the Abstract Results sentence for clarity: 'follow-up data were available for 401 (86%), 394 (85%), and 370 (79%) participants at three, six, and 12 months, respectively.'The copyedit pass suggested this grammatical improvement for readability.
- 5.LOWstatisticsConsider reporting exact p-values alongside confidence intervals for key outcomes to facilitate meta-analyses.Both reviewers suggested this as an optional enhancement; estimation-based reporting is acceptable but p-values aid meta-analysis.
- 6.LOWdata codeAdd a more detailed description of the statistical code repository (e.g., version control, license) in the Data availability statement.Reviewer 2 suggested this to improve code sharing transparency.
- 7.LOWreportingClarify the handling of the unblinded trial manager's role in outcome collection in the Limitations section to address potential bias.Both reviewers noted this limitation; a more explicit discussion of its potential impact would strengthen the paper.
- 8.LOWdata codeConsider adding a statement about the availability of the statistical analysis plan in a public repository beyond the ISRCTN registry.Reviewer 2 suggested this to enhance transparency of the analysis plan.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.