Reliability of urological telesurgery compared with local surgery: multicentre randomised controlled trial.
Wang Y, Xia D, Xu W, Rexiati M, Zhao W, Huang Q, Shi T, Wang B, Wang S, Tai S, Qiao B, Zhang Y, Ye S, Zhang X, Mao J, Zhu Y, Wang H, Ma S, Yang C, Fu W, Song T, Ai Q, Song Y, Xu L, Yang G, Gao Y, Niu S, Guo J, Liu G, Xiang X, Liang C, Ma X, Li H, Zhang X, TeleS Research Group
- DOI
- 10.1136/bmj-2024-083588
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/e711a267-3acb-4e9c-83b1-fee3a80e8be8 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×6−3★
- ClaimsOverstated claim−0.5★
- ReportingEthical approvals partially met−0.25★
- ReportingData & code availability partially met−0.25★
- 01Internal contradictions in the reported numbers
The NASA-TLX surgeon workload adjusted mean difference is −6.26 with 95% CI (−31.52 to −6.41); the point estimate is 0.15 units below the upper CI bound and 25 units above the lower bound, which is not credible for a standard 95% CI and suggests an error in either the estimate or the interval.
“The task load scores of surgeons in the telesurgery and local surgery groups were 29.0 and 48.0, respectively (adjusted mean difference −6.26, 95% CI −31.52 to −6.41; P=0.004; adjusted Cohen’s d=−0.93).”
Table 2Find in source - 02Internal contradictions in the reported numbers
The QoR-15 6-week adjusted mean difference has 95% CI (0.94 to 31.82), which excludes zero, yet the reported P is 0.06; a 95% CI excluding zero implies a two-sided p<0.05.
“6 weeks | 147.0 (140.0-150.0) | 149.0 (144.0-150.0) | 8.18 (0.94 to 31.82) | 0.53 | 0.06”
Table 2Find in source - 03Conclusion reaches beyond the evidence
Secondary outcomes, including operative basic data, complications, early recovery, oncological outcome, and medical team workload, did not differ substantially between the two groups.
“Secondary outcomes, including operative basic data, complications, early recovery, oncological outcome, and medical team workload, did not differ substantially between the two groups.”
AbstractFind in source - 04Methods and results do not match
The complications row lists 'odds ratio 0.031 (−0.03 to 0.09)', but an odds ratio cannot have a negative lower bound; the interval is that of a risk difference, and the footnote attributes the calculation to Fisher's exact test, which would yield p≈1.0 for 1/31 vs 0/31, not the reported P=0.61.
“No (%) complications and adverse events | 1/31 (3) | 0 | 0.031 (−0.03 to 0.09) | 0.03 | 0.61”
Table 2Find in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The trial is well designed and generally transparently reported: a registered, powered, single-blind multicentre RCT with pre-specified Bayesian analysis, ITT with multiple imputation, named ethics approvals, and open data with a DOI. Its main weaknesses are the multiple demonstrable internal errors in the Table 2 secondary-statistics rows, together with two reporting gaps (missing regulatory-compliance and CONSORT statements) and analysis code confined to supplementary materials.
Two independent reviewer runs (same model) agreed on all eight dimensions; the synthesis downgraded statistical analysis from pass to fail because the copyedit and integrity-verification components documented concrete internal contradictions that satisfy the 'one or more demonstrable errors' fail condition. Machine-statistics verification covered only one recomputable test (consistent); Bayesian posterior probabilities and Fisher's exact p-values in Table 2 are not machine-verifiable, so most reported statistics remain unverified. The citation check found none of 34 references retracted or missing from registries.
Numerical inconsistencies
1 finding · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks.
- CONSISTENTreported p = .610 · recomputed p = 1.000Reviewer 2Recompute Fisher's exact test p-value for the complication comparison (telesurgery 1/31 vs local 0/31) to verify the reported P=0.61.
“No (%) complications and adverse events | 1/31 (3) | 0 | 0.031 (−0.03 to 0.09) | 0.03 | 0.61 |”
Taken as given: The telesurgery group had 1 complication out of 31 (a=1, b=30).; The local surgery group had 0 complications out of 31 (c=0, d=31).; The reported p-value is two-tailed from Fisher's exact test.Method: Two-tailed Fisher's exact test from the 2x2 cell counts using pFisher2x2.How we recomputed it: pFisher2x2(1,30,0,31)
- mediuminternal contradictionThe NASA-TLX surgeon workload adjusted mean difference is −6.26 with 95% CI (−31.52 to −6.41); the point estimate is 0.15 units below the upper CI bound and 25 units above the lower bound, which is not credible for a standard 95% CI and suggests an error in either the estimate or the interval.
“The task load scores of surgeons in the telesurgery and local surgery groups were 29.0 and 48.0, respectively (adjusted mean difference −6.26, 95% CI −31.52 to −6.41; P=0.004; adjusted Cohen’s d=−0.93).”
Table 2Find in source - mediuminternal contradictionThe QoR-15 6-week adjusted mean difference has 95% CI (0.94 to 31.82), which excludes zero, yet the reported P is 0.06; a 95% CI excluding zero implies a two-sided p<0.05.
“6 weeks | 147.0 (140.0-150.0) | 149.0 (144.0-150.0) | 8.18 (0.94 to 31.82) | 0.53 | 0.06”
Table 2Find in source - lowinternal contradictionThe text states 'Only one failure was observed in the local surgery group', but the ITT table reports local success as 34/36 (94.44%), i.e., two failures in that group; one may have been imputed, but this is not explained.
“Only one failure was observed in the local surgery group owing to a surgical robotic malfunction.”
ResultsFind in source - lowinternal contradictionTable 2 reports the local positive surgical margin as 5/32 (16%), but the per-protocol local group has n=31 and the Discussion cites 16.1% (which equals 5/31), indicating a denominator inconsistency.
“No (%) positive surgical margin | 1/32 (3) | 5/32 (16)”
Table 2Find in source
Overstated conclusions
2 findings · worst mediumConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions overstated beyond the evidenceAssessed
- Conclusions only partially backed by the presented evidenceAssessed
6 major claims checked against the paper's own evidence: 1 not fully backed by the presented evidence (unsupported or overstated).
- overstatedReviewers 1, 2Secondary outcomes, including operative basic data, complications, early recovery, oncological outcome, and medical team workload, did not differ substantially between the two groups.The abstract's 'did not differ substantially' is contradicted by the paper's own statistically significant surgeon workload difference, so the claim is overbroad.Evidence: Table 2 shows surgeon NASA-TLX workload with adjusted mean difference −6.26 (95% CI −31.52 to −6.41) and P=0.004, a statistically significant difference.
“Secondary outcomes, including operative basic data, complications, early recovery, oncological outcome, and medical team workload, did not differ substantially between the two groups.”
AbstractFind in source - partialReviewer 2The decrease in the workload of surgeons in the telesurgery group may be due to the inability to be masked.The data show a significant difference, but the causal explanation (masking) is speculative and not tested.Evidence: The NASA-TLX surgeon scores were 29.0 vs 48.0 (adjusted mean difference -6.26, 95% CI -31.52 to -6.41, P=0.004).
“The decrease in the workload of surgeons in the telesurgery group may be due to the inability to be masked.”
DiscussionFind in source - supportedReviewers 1, 2Telesurgery was not inferior to local surgery in terms of the probability of surgical success in the intention-to-treat population.The presented Bayesian analysis directly supports non-inferiority within the pre-specified margin.Evidence: Primary outcome in Table 2 and Results: adjusted difference 0.02 (95% CrI −0.03 to 0.15), posterior probability 0.99; lower CrI above the pre-specified −0.1 margin.
“Telesurgery was not inferior to local surgery in terms of the probability of surgical success in the intention-to-treat population, accounting for clustering by surgeon (success probability difference 0.02 (95% credible interval −0.03 to 0.15) with bayesian posterior probability of 0.99 for non-inferiority).”
AbstractFind in source - supportedReviewer 1The reliability of telesurgery was non-inferior to that of local robotic surgery according to the non-inferiority margin of a 0.1 reduction in success probability.The conclusion matches the reported primary-outcome analysis.Evidence: Same primary outcome results in ITT and per-protocol analyses, with posterior probabilities >0.98.
“The reliability of telesurgery was non-inferior to that of local robotic surgery according to the non-inferiority margin of a 0.1 reduction in success probability.”
ConclusionFind in source - supportedReviewers 1, 2The telesurgery system was stable with a distance from 1000 km to 2800 km.The monitoring data directly support the stability claim.Evidence: Telesurgery monitoring data report stable latency (20.1-47.5 ms) and very low frame loss (0-1.5) across the four pathways.
“The telesurgery system was stable with a distance from 1000 km to 2800 km, a mean round trip network latency of 20.1-47.5 ms, and frame loss of 0-1.5 per telesurgery.”
AbstractFind in source - supportedReviewer 1This trial provides important evidence and reference for future larger cohort studies to explore the comprehensive benefits of telesurgery.This is a measured, forward-looking conclusion consistent with the reported results and limitations.Evidence: The non-inferiority result and the acknowledged small sample size and need for larger cohorts.
“This trial provides important evidence and reference for future larger cohort studies to explore the comprehensive benefits of telesurgery in clinical application.”
ConclusionFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary endpoint is a hard clinical outcome: the probability of surgical success, defined by pre-specified criteria including absence of conversion, major injury, or system malfunction. This is not a surrogate biomarker or mechanistic proxy.
“The primary outcome was the probability of success of surgery, which the research team defined on the basis of the characteristics of telesurgery. The success was confirmed according to the following determination points: the surgical process was carried out according to the planned steps; no obvious injury to large blood vessels or adjacent organs occurred during the surgery; no conversion of the surgical method occurred...”
- ADEQUATEEffect sizeThe effect size is the difference in success probability (0.02; 95% CrI −0.03 to 0.15) with a pre-specified non-inferiority margin of 0.1, which the authors explicitly state was chosen as clinically relevant. The posterior probability of non-inferiority was 0.99, and the effect is anchored to this clinical threshold.
“The pre-specified non-inferiority margin was an absolute reduction in probability of 0.1... The lower boundaries of the 95% credible intervals were all above the pre-specified non-inferiority margin of −0.1, and posterior probabilities for non-inferiority were both higher than 0.98, indicating that telesurgery was not inferior to standard local robotic surgery with high probability.”
Data authenticity concerns
2 findings · worst mediumAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Methods and results do not matchAssessed
- Implausibly large reported effectAssessed
6 integrity concerns flagged (0 high).
- mediummethod result mismatchThe complications row lists 'odds ratio 0.031 (−0.03 to 0.09)', but an odds ratio cannot have a negative lower bound; the interval is that of a risk difference, and the footnote attributes the calculation to Fisher's exact test, which would yield p≈1.0 for 1/31 vs 0/31, not the reported P=0.61.
“No (%) complications and adverse events | 1/31 (3) | 0 | 0.031 (−0.03 to 0.09) | 0.03 | 0.61”
Table 2Find in source - lowimplausible effectReported standard deviations for round-trip network latency are implausibly precise (e.g., 47.5 (SD 0.08) ms, 22.8 (0.05) ms, 20.1 (0.06) ms), suggesting either extremely stable measurements or a possible unit/typographic error.
“The mean round trip network latencies of Beijing-Urumqi (2800 km), Beijing-Hangzhou (1300 km), Beijing-Harbin (1250 km), and Beijing-Hefei (1000 km) telesurgery pathways were 47.5 (SD 0.08) ms, 30.6 (5.7) ms, 22.8 (0.05) ms, and 20.1 (0.06) ms, respectively.”
ResultsFind in source
Reporting gaps
2 findings · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
- Ethics/consent reporting incompleteAssessed
Prior work on telesurgery is cited (e.g., 2001 trans-Atlantic cholecystectomy, urological exploratory trials), and its limitations are acknowledged ('focused on feasibility in narrow indications, and no clinical evidence has been established'). The rationale logically links the premise to the non-inferiority hypothesis. The study is positioned as the first RCT in the field.
“Despite these prospects, the clinical validity of telesurgery remains unproved.”
“Early milestone, such as the 2001 trans-Atlantic cholecystectomy, prioritised technical novelty over clinical rigour.”
“However, previous studies, including our exploratory trial, focused on feasibility in narrow indications, and no clinical evidence has been established to support further research or the wider application of telesurgery.”
“As this study is the first randomised controlled trial of telesurgery, no established reports were available for reference for determining the non-inferiority margin.”
All applicable design elements for a human RCT are reported adequately: randomization method and unit, blinding (participants, follow-up specialists, statisticians; medical team unblinded with justification), power analysis (α=0.025, power=0.8, n=34/group), inclusion/exclusion criteria, and missing-data handling via multiple imputation. The comparator arm is the local-surgery group, and independent replication is not applicable for a single pivotal trial.
“The participants were randomly assigned in a 1:1 ratio to undergo either telesurgery or standard local robotic surgery, using stratified randomisation with random block sizes of four.”
“Assuming a 10% attrition rate, we determined that a sample size of 34 participants per group was needed.”
“Participants, follow-up specialists (independent research nurses or doctors), and independent statisticians were masked to the form of surgery performed (telesurgery or local surgery).”
“The participants were randomly assigned in a 1:1 ratio to undergo either telesurgery or standard local robotic surgery, using stratified randomisation with random block sizes of four.”
“As masking of the medical team was not feasible, the determination of surgical success was conducted by the surgical team according to pre-specified criteria to minimise subjectivity.”
Table 1 reports sex (e.g., male sex 28/36 vs 26/36), age (median and IQR), height, weight, and BMI, plus tumour characteristics and patient location. Both sexes were enrolled, so a single-sex justification is not applicable. Health status is captured via inclusion criteria and tumour characteristics; comorbidities are addressed in exclusion criteria. Race/ethnicity is omitted, a minor inadequacy for the demographics criterion in the Chinese context.
“Male sex, total | 28 (78) | 26 (72)”
“Median (IQR) age, years | 61.0 (57.5-68.0) | 65.0 (56.5-70.0)”
“Male sex, total | 28 (78) | 26 (72) | 24 (75) | 22 (71) |”
“Median (IQR) age, years | 61.0 (57.5-68.0) | 65.0 (56.5-70.0) | 61.0 (57.5-68.0) | 62.0 (53.0-70.5) |”
The paper states approval by the ethical review committee of Chinese PLA General Hospital (ID:2023-021) and the other four sites, and that participants signed both surgical and research consent. However, there is no sentence naming a regulatory framework such as the Declaration of Helsinki or ICH-GCP. The absence of this sub-criterion triggers a warning per the scoring rule, although the core ethics reporting is adequate.
“This study was approved by the ethical review committee of Chinese PLA General Hospital (ID:2023-021) and the other four clinical trial sites.”
“Participants signed a surgical informed consent and a separate research consent before enrolment.”
“This study was approved by the ethical review committee of Chinese PLA General Hospital (ID:2023-021) and the other four clinical trial sites.”
“Participants signed a surgical informed consent and a separate research consent before enrolment.”
The telesurgery system is identified by model and manufacturer: 'four arm, multi-port Surgical Robotic System (MP1000, Edge Medical Co, Shenzhen, China)'. Statistical software is named (SAS 9.4, RStudio, GraphPad Prism 8.0). Antibodies, cell lines, mycoplasma testing, and organisms are not applicable to this device-trial design. The local-surgery comparator device is not explicitly named, but the back-up console described implies the same system used locally; this is a minor gap but does not warrant a downgrade.
“We used a four arm, multi-port Surgical Robotic System (MP1000, Edge Medical Co, Shenzhen, China) as the robotic subsystem.”
“We used SAS version 9.4 or RStudio for analyses.”
“We used a four arm, multi-port Surgical Robotic System (MP1000, Edge Medical Co, Shenzhen, China) as the robotic subsystem.”
“We used SAS version 9.4 or RStudio for analyses.”
All applicable sub-criteria are adequate: tests are named (Bayesian mixed-effects logistic regression, mixed-effects linear regression, Fisher's exact test); assumptions are handled by design (Bayesian and mixed models); exact p-values are reported; effect sizes with 95% CIs are given; software and versions are identified; data presentation includes per-group n, CONSORT flow, and appropriate measures of dispersion. Mathematical plausibility checks (e.g., percentages in Table 1) are consistent.
“For the primary outcome, we derived the differences in the probability of success between the groups along with their 95% credible interval from the posterior distributions of the bayesian mixed effect logistic regression with penalised priors accounting for clustering by surgeon, taking the intercept of surgeon as a random effect.”
“All tests were two tailed with significance level α=0.05.”
“We used SAS version 9.4 or RStudio for analyses.”
“We used Fisher’s exact test to analysis the risk difference of the Clavin-Dindo complications without stratification.”
The data availability statement is explicit and names a public repository with a DOI (Dryad: https://doi.org/10.5061/dryad.t4b8gtjg8). The code is stated to be in the supplementary materials, which lacks version control and a persistent identifier, so code_sharing is reported_but_inadequate. Sequencing accession numbers are not applicable; managed data access is not needed because the data are open.
“The data underlying the findings in this paper are openly and publicly available and can be found at Dryad: https://doi.org/10.5061/dryad.t4b8gtjg8 .”
“The code used to analyse the data in the paper can be found in the supplementary materials.”
“The data underlying the findings in this paper are openly and publicly available and can be found at Dryad: https://doi.org/10.5061/dryad.t4b8gtjg8 .”
“The code used to analyse the data in the paper can be found in the supplementary materials.”
Trial registration is given (ChiCTR2300077721). Methods are sufficiently detailed for replication. All secondary outcomes are reported, including null results. Limitations are discussed in a dedicated section. Conclusions are proportional. Funding and competing interests are declared. The only missing element is an explicit reporting-guideline statement; however, the paper includes a CONSORT-style flow diagram (Fig 3), so practical adherence is present. Since 6 of 7 applicable criteria are adequate, the dimension passes.
“Trial registration ChiCTR.org ChiCTR2300077721.”
“Trial registration ChiCTR.org ChiCTR2300077721.”
“Firstly, clinical adoption of telesurgery remains limited, and a notable proportion of trial participants were recruited from non-local regions.”
Registered (1 ID: Chinese Clinical Trial Registry). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 34 references by DOI: 33 verified — 1 no DOI (shown, not verified).
- NO DOICancer statistics, 2024No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- dataDryadLIVEHTTP 200https://doi.org/10.5061/dryad.t4b8gtjg8Resolves to Dryad (data repository).
Copyediting
1 finding · worst lowWording, consistency and formatting errors that need correcting before submission.
- Wording or formatting errors that need correctingAssessed
13 copyedit issues flagged (4 major): mostly consistency, typo, punctuation.
- MAJORconsistencyTable 2, QoR-15 6 weeks“8.18 (0.94 to 31.82) | 0.53 | 0.06”→ Recheck the p-value and CI; a 95% CI excluding zero cannot accompany P=0.06.Internal inconsistency.
- MAJORconsistencyTable 2, NASA-TLX surgeon“−6.26 (−31.52 to −6.41) | −0.93 | 0.004”→ Verify the adjusted difference and CI; the CI is highly asymmetric and not centered on the estimate.CI/estimate mismatch.
- MAJORconsistencyTable 2, positive surgical margin“5/32 (16)”→ Confirm the denominator; per-protocol local n=31 and Discussion cites 16.1% (5/31).Denominator inconsistency.
- MAJORconsistencyTable 2, complications“odds ratio 0.031 (−0.03 to 0.09)”→ If this is a risk difference, label it as such; odds ratios cannot have negative bounds.Mislabeled effect measure.
- MINORtypoStatistical analysis“We used Fisher’s exact test to analysis the risk difference of the Clavin-Dindo complications”→ We used Fisher’s exact test to analyse the risk difference of the Clavien-Dindo complicationsVerb form and misspelling of Clavien.
- MINORtypoDiscussion“early recovery, complications, and ontological outcomes showed no statistically significant differences”→ early recovery, complications, and oncological outcomes showed no statistically significant differencesoncological misspelled as ontological.
- MINORpunctuationOutcomes“or WeChat social media..”→ or WeChat social media.Double period.
- MINORgrammarResults, Primary outcome“The probability of surgical success in the telesurgery and local surgery groups in the intention-to-treat population were 100% and 94.44%, respectively.”→ The probability ... was 100% and 94.44%, respectively.Subject-verb agreement.
- MINORtypoDiscussion, Comparison with other studies“0.1 reduction in the probability success of”→ 0.1 reduction in the probability of successMissing 'of'.
- MINORtypoMethods, Statistical analysis“We used Fisher’s exact test to analysis the risk difference of the Clavin-Dindo complications without stratification.”→ Change 'analysis' to 'analyse' and 'Clavin-Dindo' to 'Clavien-Dindo'.Both a typo and a spelling error for the classification system.
- MINORgrammarResults, Primary outcome“The probability of surgical success in the telesurgery and local surgery groups in the intention-to-treat population were 100% and 94.44%, respectively.”→ Change 'were' to 'was' to agree with the singular subject 'probability'.Subject-verb agreement.
- MINORpunctuationMethods, Outcomes“A blinded independent research nurse or doctor at each trial site ascertained postoperative secondary outcomes through medical records, in-person patient interviews, or WeChat social media..”→ Remove the extra period at the end of the sentence.Double period.
- MINORclarityTable 2 footnote“Posterior probability of bayesian mixed effects logistic regression was for non-inferiority.”→ Revise to complete the sentence, e.g., 'Posterior probability of non-inferiority from the Bayesian mixed-effects logistic regression was calculated.'Incomplete sentence.
In this post-publication audit, the trial is methodologically robust at the design level, but an informed reader should weigh the Table 2 secondary statistics very cautiously: several rows contain demonstrable internal errors (CI/P contradiction, CI not containing the estimate, a mislabelled effect measure, a denominator mismatch) that warrant a published correction and independent re-analysis before those results are relied upon. The primary non-inferiority finding is not implicated by the flagged rows; the unflagged secondary rows should still be treated as unverified until a corrected table is issued.
- 1.HIGHstatisticsRecompute and correct the QoR-15 6-week row in Table 2 (and the corresponding Results text): the reported 95% CI (0.94–31.82) excludes zero, which cannot coexist with P=0.06; publish a correction with a mutually consistent CI and P.A 95% CI excluding zero implies two-sided P<0.05, so the reported pair is mathematically impossible and misleads readers about a secondary endpoint.
- 2.HIGHstatisticsCorrect the NASA-TLX surgeon workload row in Table 2: the point estimate (−6.26) lies outside its own 95% CI (−31.52 to −6.41); re-verify the estimate and interval and correct both in a correction.A CI that does not contain its point estimate is impossible for a standard confidence interval and signals an output or transcription error.
- 3.HIGHstatisticsRe-label and re-compute the complications effect measure in Table 2: an 'odds ratio' with a negative lower bound (−0.03 to 0.09) is impossible; if it is a risk difference, label it as such and recompute the P value, since a Fisher's exact test on the stated counts would not yield P=0.61.The negative lower bound and the method/result mismatch make the reported effect measure invalid as printed.
- 4.HIGHstatisticsReconcile the positive surgical margin denominator in Table 2 (5/32) with the per-protocol local group size (n=31) and the Discussion's 16.1% (5/31), and state the correct fraction in a correction.An inconsistent denominator in a reported event count undermines confidence in the oncological/complication outcomes.
- 5.HIGHstatisticsClarify the local-group failure count: the text says only one failure was observed in the local surgery group, but the ITT table reports 34/36 success (two failures); state explicitly whether one failure was imputed for a missing outcome and align the wording.The text and table disagree about a primary-outcome event count, which is essential to interpreting the non-inferiority result.
- 6.HIGHreportingTemper the Abstract's claim that secondary outcomes 'did not differ substantially': surgeon workload differed significantly (NASA-TLX P=0.004), so revise the sentence to state which outcomes differed and which did not.The current wording overstates null findings and is contradicted by the paper's own statistically significant result.
- 7.HIGHethicsAdd an explicit statement of adherence to a recognised regulatory framework (Declaration of Helsinki and/or ICH-GCP) via an erratum or the repository study record; the named ethics approvals and consent are present, but the compliance statement is missing.The ethics dimension is currently incomplete because the regulatory-compliance sub-criterion is not reported, a fixable reporting gap.
- 8.HIGHdata codeDeposit the analysis code in a version-controlled public repository (e.g., GitHub with a Zenodo DOI) and update the Data availability statement to cite it; supplementary-material code lacks version control and a persistent identifier.The data are openly available with a DOI, but the code that produced the results is not persistently citable, limiting reproducibility.
- 9.HIGHstatisticsVerify every remaining Table 2 row for internal consistency (CI contains the estimate; CI excluding zero implies P<0.05; correct effect-measure labels) and publish one consolidated correction for all affected rows.Multiple independent indicators of error in a single table raise the possibility of additional transcription or output errors that a systematic re-check would catch.
- 10.MEDIUMreportingAdd an explicit CONSORT statement and a completed CONSORT checklist (e.g., in supplementary or repository records); a flow diagram alone is not a checklist.The reporting-guideline sub-criterion is not met, though the trial flow diagram suggests practical adherence to CONSORT.
- 11.MEDIUMreportingReport race/ethnicity and a full comorbidity profile in the baseline table or supplementary material, or explicitly justify their omission.The demographics sub-criterion is currently only partially reported, and completeness would strengthen the baseline comparability assessment.
- 12.MEDIUMcopyeditFix the copyediting errors in a correction: 'analyse' not 'analysis', 'Clavien-Dindo' not 'Clavin-Dindo', 'oncological' not 'ontological', remove the double period in the Outcomes section, change 'were' to 'was' for 'the probability of surgical success', restore 'of' in 'probability of success', and complete the Table 2 footnote sentence.These errors, while minor, reduce the paper's professional quality and one misnames a standard surgical complication classification system.
- 13.LOWotherVerify and clarify the network-latency standard deviations (e.g., 47.5 (SD 0.08) ms, 22.8 (0.05) ms): SDs of ~0.05–0.08 ms around means of 20–50 ms are implausibly precise; check units and correct the reported dispersion.Implausibly precise dispersion values may indicate a unit or typographical error and invite reader distrust of the technical measurements.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.