A Randomized Trial of Physical Therapy for Meniscal Tear and Knee Pain.
Katz JN, Collins JE, Bisson L, Jones MH, Irrgang JJ, Selzer F, Safran-Norton CE, Spindler KP, Yang HY, Shrestha S, Bennell KL, Sullivan JK, Kluczynski MA, Arant K, Opare-Addo M, Huizinga JL, Zimmerman Z, Sople D, Tonsoline P, Kale M, Wind WM Jr, Chen AF, Freitas M, Lesniak B, Jordan K, Matzkin EG, Dawson C, Farrow L, Musahl V, Leddy JJ, Martin SD, Losina E
- DOI
- 10.1056/NEJMoa2503385
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/4c88bf4d-28fb-492b-be60-0e3e9df0544e is authoritative.
How this rating was calculated
- ReportingData & code availability not met−0.5★
- ReportingEthical approvals partially met−0.25★
- References were found (35) but none could be checked — every lookup failed or lacked a DOI.
- 01Data and code not shared
No data availability statement is provided, which is required for clinical trials.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The TeMPO four-arm RCT is methodologically rigorous: sound randomization and blinding, a priori power analysis, detailed intervention descriptions (including a sham PT arm), trial registration, exact p-values, and effect sizes with confidence intervals. The principal reporting gaps are the complete absence of a data availability statement, an unnamed statistical software package, and inexplicit ethics-approval/regulatory-compliance language; copyedit issues are minor and chiefly consistency-related.
Three independent reviewer runs were synthesized; they agreed on premise, design, biological variables, statistics, and transparency, and diverged only on ethical approvals (warn vs pass), key resources (not applicable vs pass), and data code availability (fail vs warn) — resolved by weighing the underlying evidence and applying the stated scoring rules. Statistical verification covered only 1 machine-checkable test (consistent); threshold-only and resampling-based p-values could not be verified, so the paper's statistics are not certified as correct beyond what was checked. Key resources was scored not applicable for this behavioral/exercise trial with no biological/chemical resource, investigational drug/device, or bespoke software.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks.
- CONSISTENTreported p = .970 · recomputed p = .958Reviewer 1Recompute p-value for the primary comparison Home Exercise vs Home Exercise + Text Messages from the reported difference and 98.3% CI.
“The difference in three-month change between Home Exercise versus Home Exercise + text messages was −0.1 points (98.3% CI −3.8, 3.7)”
Taken as given: The CI is a two-sided 98.3% confidence interval; The estimate is a linear regression coefficient; The CI is based on a normal approximationMethod: Two-tailed p-value derived from the confidence interval width using the normal distribution (pCI function).How we recomputed it: pCI(-0.1, -3.8, 3.7, 0)
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewers 1, 2The addition of in clinic PT (standard or sham) appeared to be associated with slightly greater pain improvement at 6 months compared to home exercises with no in-clinic PT.The 6-month difference is reported as 4.1 points (95% CI 0.7, 7.6) in a secondary analysis without multiplicity adjustment, and the paper itself cautions against definitive interpretation.Evidence: Secondary analysis: 'At 6 months, the difference in KOOS Pain from baseline between the Standard PT + Home Exercise + Text Messages and Home Exercise arms was 4.1 (95% CI 0.7, 7.6).'
“The addition of in clinic PT (standard or sham) appeared to be associated with slightly greater pain improvement at 6 months compared to home exercises with no in-clinic PT.”
DiscussionFind in source - partialReviewer 3The addition of in-clinic PT (standard or sham) appeared to be associated with slightly greater pain improvement at 6 months compared to home exercises with no in-clinic PT.This finding comes from a secondary analysis without adjustment for multiplicity, and the authors caution against definitive interpretation.Evidence: At 6 months, the difference in KOOS Pain between Standard PT + Home Exercise + Text Messages and Home Exercise was 4.1 (95% CI 0.7, 7.6).
“At 6 months, the difference in KOOS Pain from baseline between the Standard PT + Home Exercise + Text Messages and Home Exercise arms was 4.1 (95% CI 0.7, 7.6).”
ResultsFind in source - supportedReviewers 1, 2For patients with degenerative meniscal tear and knee pain, the addition of physical therapy or text messages to encourage adherence to home exercises was not superior in reducing pain to a home exercise program alone.The primary outcome shows no significant difference between arms, with small effect sizes and confidence intervals crossing zero, directly supporting this claim.Evidence: Table 2 and primary analysis results: differences in KOOS Pain change of 2.5 points (98.3% CI -1.3, 6.2) and -0.1 points (98.3% CI -3.8, 3.7).
“For patients with degenerative meniscal tear and knee pain, the addition of physical therapy or text messages to encourage adherence to home exercises was not superior in reducing pain to a home exercise program alone.”
ConclusionFind in source - supportedReviewer 1KOOS Pain scores in the Standard PT + Home Exercise + Text Messages and Sham PT + Home Exercise + Text Messages arms were virtually identical across all time points.The data show nearly identical values at all time points, with the difference at 3 months being 0.7 points (95% CI -3.7, 2.3) and overlapping at 6 and 12 months.Evidence: Table 2 and Figure 1: differences between sham and standard PT are small and non-significant across time points.
“KOOS Pain scores in the Standard PT + Home Exercise + Text Messages and Sham PT + Home Exercise + Text Messages arms were virtually identical across all time points.”
DiscussionFind in source - supportedReviewers 1, 3Motivational text messages were not associated with differences in adherence to home exercises nor in pain outcomes.The comparison between Home Exercise and Home Exercise + Text Messages shows no difference in pain improvement (-0.1 points, 98.3% CI -3.8, 3.7) and adherence rates were similar (77% vs 80%).Evidence: Primary analysis and adherence results: 'The difference between Home Exercises and Home Exercises + Text Messages was −0.1 (98.3% CI −3.8, 3.7)' and 'The mean proportion of weeks in which participants exercised at least three times was 77% for Home Exercise, 80% for Home Exercise + Text Messages'.
“Motivational text messages were not associated with differences in adherence to home exercises nor in pain outcomes.”
AbstractFind in source - supportedReviewer 2Adverse events were rare, generally minor, and evenly distributed overall across arms.Table 3 shows similar rates of adverse events across arms, with most events being minor.Evidence: Table 3: Any adverse event rates: 21.1%, 21.6%, 21.4%, 16.0% across arms.
“Adverse events were rare, generally minor, and evenly distributed overall across arms.”
AbstractFind in source - supportedReviewer 2Contextual effects are likely to explain the small apparent differences in pain between standard PT with home exercises versus home exercises alone over 12 months.The similar outcomes between standard PT and sham PT support this interpretation.Evidence: Results: 'KOOS Pain scores in the Standard PT + Home Exercise + Text Messages and Sham PT + Home Exercise + Text Messages were nearly identical at all timepoints.'
“These findings suggest that contextual effects are likely to explain the small apparent differences in pain between standard PT with home exercises versus home exercises alone over 12 months.”
DiscussionFind in source - supportedReviewers 2, 3The addition of physical therapy or text messages to encourage adherence to home exercises was not superior in reducing pain to a home exercise program alone.The primary outcome results show no statistically significant or clinically important differences between arms, supporting this claim.Evidence: Primary analysis: differences in KOOS Pain change at 3 months between arms are small and not significant (e.g., Standard PT vs Home Exercise: 2.5 points, 98.3% CI -1.3 to 6.2, p=0.11).
We did not observe meaningful differences in the three primary contrasts.
Resultsreviewer’s wording - supportedReviewer 3Sham PT and Standard PT produced virtually identical KOOS Pain scores.The paper reports that KOOS Pain scores in the two in-clinic PT arms were nearly identical at all timepoints.Evidence: KOOS Pain scores in the Standard PT + Home Exercise + Text Messages and Sham PT + Home Exercise + Text Messages were nearly identical at all timepoints.
“KOOS Pain scores in the Standard PT + Home Exercise + Text Messages and Sham PT + Home Exercise + Text Messages were nearly identical at all timepoints.”
DiscussionFind in source
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
2 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data and code not sharedAssessed
- Ethics/consent reporting incompleteAssessed
Prior work is cited extensively, including the prevalence of meniscal tears, prior RCTs comparing surgery to PT, and treatment guidelines. The paper notes that it is unclear whether improvements from PT are due to physiological effects or therapist interaction. The TeMPO trial is designed to test whether adding text reminders or in-clinic PT provides additional benefit over home exercise alone, and whether standard PT is more effective than sham, directly addressing the limitations of prior research.
“Several randomized controlled trials (RCTs) reported that participants randomized to arthroscopic partial meniscectomy reported similar pain and function after one year compared to those randomized to in-clinic physical therapy (PT), home exercises, or both.”
“It is unclear whether improvements following PT in these trials arose from physiological effects of exercises and/or interaction with physical therapists.”
“TeMPO (Treatment of Meniscal Problems in Osteoarthritis) was a RCT designed to address whether adding text reminders to exercise or adding in-clinic PT result in greater pain relief than home exercises alone.”
“Several randomized controlled trials (RCTs) reported that participants randomized to arthroscopic partial meniscectomy reported similar pain and function after one year compared to those randomized to in-clinic physical therapy (PT), home exercises, or both.”
“It is unclear whether improvements following PT in these trials arose from physiological effects of exercises and/or interaction with physical therapists.”
“Sham PT included elements not known to have physiologic benefit including 1) assessment of knee symptoms (5 minutes); 2) ultrasound of knee region with intensity set to 0 (12 minutes); 3) inert lotion applied gently along mid-thigh and distal tibia (5 minutes); and 3) sham manual therapy, consisting of minimal force to non-articular areas of the knee, without joint mobilization (8 minutes).”
“Several randomized controlled trials (RCTs) reported that participants randomized to arthroscopic partial meniscectomy reported similar pain and function after one year compared to those randomized to in-clinic physical therapy (PT), home exercises, or both.”
“It is unclear whether improvements following PT in these trials arose from physiological effects of exercises and/or interaction with physical therapists.”
“TeMPO (Treatment of Meniscal Problems in Osteoarthritis) was a RCT designed to address whether adding text reminders to exercise or adding in-clinic PT result in greater pain relief than home exercises alone.”
Randomization used varying block sizes stratified by site and KL grade. Assessors were blinded. A sample size calculation was performed assuming 80% power, alpha 0.0167, to detect a 5.3-point difference. The primary analysis used multiple imputation and complete case analysis. Eligibility criteria are clearly specified.
“randomized 1:1:1:1 to four arms in varying blocks of 4 and 8, stratified by site and KL grade (0-2 vs. 3).”
“Personnel who assessed participants were blinded to treatment assignment.”
“We powered TeMPO to detect an effect of 0.33 SD, which equates to 5.3 points on the KOOS Pain scale, given baseline SD of 16.”
“randomized 1:1:1:1 to four arms in varying blocks of 4 and 8, stratified by site and KL grade (0-2 vs. 3).”
“Personnel who assessed participants were blinded to treatment assignment.”
“We powered TeMPO to detect an effect of 0.33 SD, which equates to 5.3 points on the KOOS Pain scale, given baseline SD of 16. . Assuming 80% power and Type I error of 0.0167, each arm required 194 subjects.”
“randomized 1:1:1:1 to four arms in varying blocks of 4 and 8, stratified by site and KL grade (0-2 vs. 3).”
“Personnel who assessed participants were blinded to treatment assignment.”
Sex is reported for each arm in Table 1. Age and BMI are reported as means with SDs. Health status is captured through KL grade and baseline KOOS scores. Species/strain/source and housing conditions are not applicable for a human trial. Demographics (race, ethnicity, education) are reported in Table 1. Sex justification is not applicable as both sexes are enrolled.
“Age (mean, SD) | 58.8 (8.1) | 58.9 (7.5) | 59.5 (7.5) | 59.4 (8.1) |”
“Female | 114 (52%) | 131 (59%) | 132 (60%) | 128 (58%)”
“Age (mean, SD) | 58.8 (8.1) | 58.9 (7.5) | 59.5 (7.5) | 59.4 (8.1)”
“White | 180 (88%) | 201 (92%) | 192 (88%) | 195 (90%)”
“Table 1: Baseline features of TeMPO study participants according to randomization arm”
The paper states that 'Sites ceded oversight to the Mass General Brigham IRB,' indicating approval. However, there is no mention of obtaining informed consent from participants or adherence to the Declaration of Helsinki or other regulations. For a clinical trial, both are required.
“Sites ceded oversight to the Mass General Brigham IRB.”
“Sites ceded oversight to the Mass General Brigham IRB.”
“Clinical Trials.gov (http://ClinicalTrials.gov) NCT03059004”
“Sites ceded oversight to the Mass General Brigham IRB.”
The interventions are home exercise, text messages, and physical therapy (standard and sham). No investigational drug, biologic, or device is involved. The only materials mentioned are ankle weights and instructional pamphlets, which are not considered key biological or chemical resources. Therefore, this dimension is not applicable.
“Standard PT , each session followed an unsupervised warm-up on an exercise bicycle and included: 1) manual therapy -- soft tissue and joint mobilization and stretching of tissues around the knee (5 minutes); and 2) therapist-directed strengthening and functional exercises, targeting the gluteus maximus and medius, hamstrings, and quadriceps muscles (25 minutes).”
The primary analysis uses linear regression with Bonferroni correction, and tests are named. Assumptions are addressed through the use of robust methods (multiple imputation, mixed models). Exact p-values are reported for primary comparisons (e.g., p=0.97, 0.11, 0.12). Effect sizes with confidence intervals are the primary reporting method. Software is not explicitly identified. Data presentation includes tables with per-group n, means, SDs, and CIs. Mathematical plausibility checks are not applicable as the data are continuous and N is large.
“Difference in ΔKOOS Pain (98.3% CI) | P-value | | A | B | | Home Exercise | Home Exercise + Text Messages | −17.1 | −17.0 | −0.1 (−3.8, 3.7) | | 0.97 |”
“We used a Bonferroni-corrected p-value of 0.0167 for these 3 contrasts.”
“The difference in three-month change in KOOS Pain between the Standard PT + Home Exercise + Text Messages and the Home Exercise arms was 2.5 points (98.3% confidence interval (CI) −1.3, 6.2)”
The paper does not contain any statement about data sharing or availability. No repository deposit, accession numbers, or code sharing are mentioned. This is a significant omission for a clinical trial.
Methods are detailed enough for replication. The trial is registered (NCT03059004). No reporting guideline is explicitly mentioned (e.g., CONSORT), but the paper follows standard clinical trial reporting. All pre-specified outcomes appear to be reported. Limitations are discussed (generalizability, COVID impact). Conclusions are proportional, noting that the findings should not be interpreted as definitive. Funding sources and COI are disclosed.
“Clinical Trials.gov (http://ClinicalTrials.gov) NCT03059004”
“Funding: Supported by NIH/NIAMS U01AR071658; R21AR076156; P30AR072577; K01AR075879 (Dr. Collins).”
“Clinical Trials.gov (http://ClinicalTrials.gov) NCT03059004”
“We note several limitations. Generalizability is limited by the small number of Black, Asian, and Hispanic participants”
“Funding: Supported by NIH/NIAMS U01AR071658; R21AR076156; P30AR072577; K01AR075879 (Dr. Collins).”
“Clinical Trials.gov (http://ClinicalTrials.gov) NCT03059004”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 35 references by DOI: 0 verified — 35 no DOI (shown, not verified).
- NO DOIIncidental Meniscal Findings on Knee MRI in Middle Aged and Elderly PersonsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe clinical importance of meniscal tears demonstrated by magnetic resonance imaging in osteoarthritis of the kneeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIncrease in outpatient knee arthroscopy in the United States: a comparison of National Surveys of Ambulatory Surgery, 1996 and 2006No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArthroscopic partial meniscectomy for meniscal tears of the knee: a systematic review and meta-analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArthroscopic or conservative treatment of degenerative medial meniscal tears: a prospective randomised trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIs arthroscopic surgery beneficial in treating non-traumatic, degenerative medial meniscal tears? A five year follow-upNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISurgery versus physical therapy for a meniscal tear and osteoarthritisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA comparative study of meniscectomy and nonoperative treatment for degenerative horizontal tears of the medial meniscusNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExercise therapy versus arthroscopic partial meniscectomy for degenerative meniscal tear in middle aged patients: randomised controlled trial with two year follow-upNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEffect of Early Surgery vs Physical Therapy on Knee Function Among Patients With Nonobstructive Meniscal Tears: The ESCAPE Randomized Clinical TrialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIFive-year outcome of operative and nonoperative management of meniscal tear in persons older than forty-five yearsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISurgical Management of Degenerative Meniscus Lesions: The 2016 ESSKA Meniscus ConsensusNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPosition Statement From the Australian Knee Society on Arthroscopic Surgery of the Knee, Including Reference to the Presence of Osteoarthritis or Degenerative Joint Disease: Updated October 2016No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPosition Statement of the Arthroscopy Association of Canada (AAC) Concerning Arthroscopy of the Knee Joint-September 2017No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIArthroscopic meniscal surgery: a national society treatment guideline and consensus statementNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDutch Guideline on Knee Arthroscopy Part 1, the meniscus: a multidisciplinary review by the Dutch Orthopaedic AssociationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAAOS clinical practice guideline summary: management of osteoarthritis of the knee (nonarthroplasty)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDegenerative meniscus tears-assimilation of evidence and consensus statements across three continents: state of the artNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe TeMPO trial (treatment of meniscal tears in osteoarthritis): rationale and design features for a four arm randomized controlled clinical trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA consensus-based process identifying physical therapy and exercise treatments for patients with degenerative meniscal tears and knee OA: the TeMPO physical therapy interventions and home exercise programNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIThe Theory of planned behaviorNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA comparison of the theory of planned behavior and the theory of reasoned actionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHow individuals, environments, and health behaviors interactNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIKnee Injury and Osteoarthritis Outcome Score (KOOS)—Development of a Self-Administered Outcome MeasureNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReliability and validity of the EuroQol in patients with osteoarthritis of the kneeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntrarater Reliability and Agreement of Recommended Performance-Based Tests and Common Muscle Function Tests in Knee OsteoarthritisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIReliability and measurement error of the Osteoarthritis Research Society International (OARSI) recommended performance-based tests of physical function in people with hip and knee osteoarthritisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMeaningful thresholds for patient-reported outcomes following interventions for anterior cruciate ligament tear or traumatic meniscus injury: a systematic review for the OPTIKNEE consensusNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMultiple imputation using chained equations: Issues and guidance for practiceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInference and missing dataNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIMultiple imputation for nonresponse in surveysNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILongitudinal data analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExercise for osteoarthritis of the kneeNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPain in clinical trials for knee osteoarthritis: estimation of regression to the meanNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIExamination of overall treatment effect and the proportion attributable to contextual effect in osteoarthritis: meta-analysis of randomised controlled trialsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- datahttps://clinicaltrials.gov/ct2/show/NCT03059004LIVEHTTP 200Resolves, but the content could not be matched to the paper.
Copyediting
6 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 6 minor suggestions below.
6 copyedit issues flagged: mostly consistency, clarity, grammar.
- MINORconsistencyTable 2 header and text“Home Exercises”→ Home ExerciseIn Table 2, the arm label appears as 'Home Exercises' in the row header but as 'Home Exercise' in the text and other tables. Standardize to 'Home Exercise'.
- MINORconsistencyTable 2 header“ΔKOOS Pain for A | ΔKOOS Pain for B”→ Use consistent notation: 'ΔKOOS Pain (A)' and 'ΔKOOS Pain (B)' for clarity.Minor formatting issue.
- MINORclarityMethods, Study interventions“Sham PT included elements not known to have physiologic benefit including 1) assessment of knee symptoms (5 minutes); 2) ultrasound of knee region with intensity set to 0 (12 minutes); 3) inert lotion applied gently along mid-thigh and distal tibia (5 minutes); and 3) sham manual therapy”→ Change 'and 3)' to 'and 4)' for correct numbering.Numbering error: two items are labeled '3)'.
- MINORgrammarResults, Primary outcome“The difference in three-month change in KOOS Pain between the Standard PT + Home Exercise + Text Messages and the Home Exercise arms was 2.5 points (98.3% confidence interval (CI) −1.3, 6.2), as was the difference between Standard PT + Home Exercise + Text Messages and Home Exercises + Text Messages (2.5 points, 98.3% CI −1.4, 6.5).”→ Consider rephrasing for clarity: 'The difference in three-month change in KOOS Pain was 2.5 points (98.3% CI −1.3, 6.2) between the Standard PT + Home Exercise + Text Messages and Home Exercise arms, and 2.5 points (98.3% CI −1.4, 6.5) between Standard PT + Home Exercise + Text Messages and Home Exercise + Text Messages.'The sentence is grammatically correct but could be clearer.
- MINORconsistencyThroughout“Use of 'subjects' and 'participants' interchangeably”→ Standardize to 'participants' throughout.The paper uses both terms; 'participants' is preferred for human research.
- MINORclarityTable 1“Hispanic or Latino row: 'No' and 'Yes' labels”→ Clarify that percentages are within each arm.Table 1 is clear but could be slightly more explicit about the denominator.
The published trial is robust and well reported on the core rigor dimensions; an informed reader should weigh the missing data-availability statement, the unspecified analysis software, and the inexplicit IRB-approval/regulatory language — none of which undermine the reported findings but all of which warrant a correction/clarification or caution. The one machine-checkable statistic recomputed consistently, and no retracted or unverifiable references were found, so nothing here suggests the conclusions are invalid.
- 1.HIGHdata codeAdd a data availability statement (via erratum/corrigendum) describing how de-identified individual participant data can be accessed (e.g., controlled-access repository, or reasonable request to the corresponding author with a data-sharing agreement and timeframe).A data-driven clinical trial with no data-sharing statement has the most consequential reporting gap; readers cannot verify or reuse the data.
- 2.HIGHstatisticsIdentify the statistical software and version used for all analyses (e.g., SAS 9.4, R 4.x, Stata) in the Statistical Analysis section.Unnamed analysis software hinders reproducibility of the reported primary and secondary analyses.
- 3.HIGHethicsAdd an explicit regulatory-compliance statement (e.g., 'conducted in accordance with the Declaration of Helsinki') and state the IRB approval explicitly with a protocol number.The current 'Sites ceded oversight to the Mass General Brigham IRB' language is not an explicit approval statement and lacks a protocol identifier, and no compliance statement is present.
- 4.MEDIUMdata codeDeposit any custom analysis code in a public repository (e.g., Zenodo/GitHub) with a DOI and cite it in the paper.Code sharing would make the reported analyses fully reproducible and complement the data availability statement.
- 5.MEDIUMreportingReference the CONSORT 2010 checklist and include the completed checklist as supplementary material.CONSORT is the standard reporting guideline for RCTs and is currently not referenced, a minor transparency gap.
- 6.MEDIUMcopyeditFix the numbering error in the sham PT description by changing the second 'and 3)' to 'and 4)' (Methods, Study interventions).Two items are currently both labeled '3)', a concrete error a reader will notice.
- 7.MEDIUMcopyeditStandardize the arm label to 'Home Exercise' throughout, including the Table 2 row header that currently reads 'Home Exercises'.Inconsistent arm naming across tables and text is a consistency defect.
- 8.LOWstatisticsAdd a brief statement on how outliers were handled (even if none were excluded) in the Statistical Analysis section.Reviewer 2 flagged that outlier handling is not explicitly reported; multiple imputation addresses missing data but not outliers.
- 9.LOWreportingClarify the unit of randomization explicitly ('Participants were randomized individually') in the Recruitment and randomization section.Though implied by the design, an explicit statement removes any ambiguity about the randomization unit.
- 10.LOWcopyeditStandardize terminology to 'participants' throughout (currently 'subjects' and 'participants' are used interchangeably).Consistent terminology is preferred for human research reporting.
- 11.LOWcopyeditClarify in Table 1 that percentages in the Hispanic/Latino row are within each arm (state the denominator).Minor clarity improvement about the percentage denominator for that row.
- 12.LOWcopyeditRephrase the long primary-outcome sentence in the Results for clarity (split into two sentences with the CI attached to each comparison).The sentence is grammatically correct but difficult to parse as written.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.