Reliability of urological telesurgery compared with local surgery: multicentre randomised controlled trial.
Wang Y, Xia D, Xu W, Rexiati M, Zhao W, Huang Q, Shi T, Wang B, Wang S, Tai S, Qiao B, Zhang Y, Ye S, Zhang X, Mao J, Zhu Y, Wang H, Ma S, Yang C, Fu W, Song T, Ai Q, Song Y, Xu L, Yang G, Gao Y, Niu S, Guo J, Liu G, Xiang X, Liang C, Ma X, Li H, Zhang X, TeleS Research Group
- DOI
- 10.1136/bmj-2024-083588
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-21
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/58edc796-b8d8-401b-b0b0-d0dbb693ccaa is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- StatisticsEffect estimate outside its own confidence interval ×2−1★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 2 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- 01Effect estimate falls outside its own confidence interval
mean difference 1.60 lies outside its own 95% CI [-5.87, 0.53] — an interval cannot exclude the estimate it is the interval for, so at least one of these printed numbers is wrong
“mean difference 1.60, 95% CI −5.87 to 0.53”
- 02Effect estimate falls outside its own confidence interval
mean difference -6.26 lies outside its own 95% CI [-31.52, -6.41] — an interval cannot exclude the estimate it is the interval for, so at least one of these printed numbers is wrong
“mean difference −6.26, 95% CI −31.52 to −6.41”
- 03Efficacy rests on an unvalidated surrogate endpoint
The primary outcome is the probability of surgical success, which is a composite of procedural criteria (planned steps, no major injury, no conversion, no postponement due to malfunction). This is a surrogate for clinical benefit, not a hard clinical outcome. The paper does not provide evidence linking this composite to patient-important outcomes, nor does it demonstrate target engagement at the tested dose (e.g., PK/PD) for the intervention. The surrogate is a technical/process measure, and its clinical meaningfulness is not validated.
“The primary outcome was the probability of success of surgery, which the research team defined on the basis of the characteristics of telesurgery. The success was confirmed according to the following determination points: the surgical process was carried out…”
- 04Treatment effect not shown to be clinically meaningful
The primary effect size is a difference in success probability of 0.02 (95% CrI -0.03 to 0.15) in the ITT analysis, with a non-inferiority margin of 0.1. The observed difference is small and the confidence interval includes the possibility of a reduction in success probability up to 0.15, which exceeds the margin. The effect is not anchored to a minimal clinically important difference; the margin was chosen by expert discussion without established reference. The result is presented as non-inferior, but the magnitude of the effect is not shown to be clinically meaningful.
“The estimated difference in the probability of success in the intention-to-treat and per protocol populations was 0.02 (95% credible interval (CrI) −0.03 to 0.15; bayesian posterior probability 0.99 for non-inferiority) and 0.003 (−0.001 to 0.03; bayesian…”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and transparently reported non-inferiority RCT comparing telesurgery with local robotic surgery. The study design is rigorous, with clear randomisation, blinding, sample size calculation, and pre-specified criteria. The paper is strong on ethics, data/code availability, and reporting transparency, with only minor copyedit issues and a few suggested improvements.
Both reviewers classified the study as interventional, and this was adopted. The evaluation covered all eight dimensions; several sub-criteria were marked not applicable (e.g., species/strain, housing, cell lines) due to the human trial nature. The statistics verification component checked only a subset of tests; the two inconsistent recomputations were not detailed and do not constitute demonstrable errors in the paper's own reporting.
Numerical inconsistencies
2 findings · worst highValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Effect estimate outside its own confidence intervalRecomputed
- Internal contradictions in the reported numbersAssessed
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks. 2 effect estimates printed outside their own confidence interval.
- CI_COHERENCEmean difference 1.60 lies outside its own 95% CI [-5.87, 0.53] — an interval cannot exclude the estimate it is the interval for, so at least one of these printed numbers is wrong
“mean difference 1.60, 95% CI −5.87 to 0.53”
- CI_COHERENCEmean difference -6.26 lies outside its own 95% CI [-31.52, -6.41] — an interval cannot exclude the estimate it is the interval for, so at least one of these printed numbers is wrong
“mean difference −6.26, 95% CI −31.52 to −6.41”
- UNCOMPUTABLEreported p = .060 · recomputed p = .299Reviewer 1Check the p-value for the QoR-15 score at 6 weeks.
“The QoR-15 scale score at baseline, four weeks, and six weeks for the telesurgery and local surgery groups were 147.5 and 148.0 (adjusted mean difference 1.60, 95% CI −5.87 to 0.53; P=0.10; adjusted Cohen’s d=−0.44), 146.0 and 145.0 (adjusted mean difference 2.07, −5.08 to 3.22; P=0.65; adjusted Cohen’s d=−0.11), and 147.00 and 149.00 (adjusted mean difference 8.18, 0.94 to 31.82; P=0.06; adjusted Cohen’s d=0.53) , respectively.”
Taken as given: The reported adjusted mean difference is 8.18 and the 95% CI is 0.94 to 31.82.; The p-value is two-tailed.Method: Recomputed the two-tailed p-value from the CI difference and effect size using the pCI function.How we recomputed it: pCI(8.18, 0.94, 31.82, 0) - CONSISTENTreported p = .070 · recomputed p = .122Reviewer 2Check p-value for positive surgical margin odds ratio using the reported OR and 95% CrI.
“The positive margin rates of telesurgery and local surgery were 3% and 16%, respectively (odds ratio 13.41, 95% CrI 0.80 to 575.37; two tailed weighted posterior probability 0.07) ().”
Taken as given: The odds ratio is 13.41.; The 95% CrI is (0.80, 575.37).; The CrI is two-sided at 95%.; The p-value is two-tailed.Method: Recomputed two-tailed p-value from the log odds ratio and its 95% CI using the normal approximation (pCI function with log=1).How we recomputed it: pCI(13.41, 0.80, 575.37, 1)
- lowinternal contradictionThe abstract states 'Telesurgery was not inferior to local surgery in terms of the probability of surgical success in the intention-to-treat population, accounting for clustering by surgeon (success probability difference 0.02 (95% credible interval −0.03 to 0.15) with bayesian posterior probability of 0.99 for non-inferiority).' The results section states 'The probability of surgical success in the telesurgery and local surgery groups in the intention-to-treat population were 100% and 94.44%, respectively.' The difference between 100% and 94.44% is 5.56%, which is not 0.02. This is a potential inconsistency, but the 0.02 is the adjusted difference from the Bayesian model, not the raw difference.
Telesurgery was not inferior to local surgery in terms of the probability of surgical success in the intention-to-treat population, accounting for clustering by surgeon (success probability difference 0.02 (95% credible interval −0.03 to 0.15) with bayesian posterior probability of 0.99 for non-inferiority). The probability of surgical success in the telesurgery and local surgery groups in the intention-to-treat population were 100% and 94.44%, respectively.
Abstractreviewer’s wording - lowinternal contradictionThe abstract states 'A total of 72 participants were enrolled' and 'nine (12.5%) withdrew', but the per-protocol population has 32 and 31 participants, summing to 63, which is 72 minus 9. This is consistent, but the text says 'nine (12.5%) withdrew' while 9/72 = 12.5% exactly, which is fine.
“Among the 381 patients screened, 309 were excluded and 72 participants were randomised (; supplementary figures A and B), of whom nine (12.5%) withdrew (four for telesurgery and five for local surgery). Finally, 32 participants (17 prostatectomies and 15 partial nephrectomies) underwent telesurgery, and 31 participants (16 prostatectomies and 15 partial nephrectomies) underwent local surgery.”
ResultsFind in source
Overstated conclusions
3 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
- Conclusions only partially backed by the presented evidenceAssessed
5 major claims checked against the paper's own evidence: 1 only partially supported (evidence backs part of the claim; gaps or caveats remain); the rest adequately supported.
- partialReviewer 1Secondary outcomes, including operative basic data, complications, early recovery, oncological outcome, and medical team workload, did not differ substantially between the two groups.The paper reports no statistically significant differences for most secondary outcomes, but the surgeon's workload (NASA-TLX) was significantly lower in the telesurgery group (P=0.004).Evidence: The secondary outcomes in Table 2.
“Secondary outcomes, including operative basic data, complications, early recovery, oncological outcome, and medical team workload, did not differ substantially between the two groups.”
AbstractFind in source - supportedReviewers 1, 2Telesurgery is non-inferior to local surgery in terms of the probability of surgical success.The primary outcome analysis shows a success probability difference of 0.02 (95% CrI -0.03 to 0.15) with a posterior probability of 0.99 for non-inferiority, which is above the pre-specified margin of -0.1.Evidence: The primary outcome analysis in the Results section and Table 2.
“Telesurgery was not inferior to local surgery in terms of the probability of surgical success in the intention-to-treat population, accounting for clustering by surgeon (success probability difference 0.02 (95% credible interval −0.03 to 0.15) with bayesian posterior probability of 0.99 for non-inferiority).”
AbstractFind in source - supportedReviewers 1, 2The telesurgery system was stable with low latency and frame loss.The paper reports mean round trip network latencies of 20.1-47.5 ms and frame loss of 0-1.5 per telesurgery, which are presented as stable.Evidence: The telesurgery monitoring data in the Results section.
“The telesurgery system was stable with a distance from 1000 km to 2800 km, a mean round trip network latency of 20.1-47.5 ms, and frame loss of 0-1.5 per telesurgery.”
ResultsFind in source - supportedReviewers 1, 2The reliability of telesurgery was non-inferior to that of local robotic surgery according to the non-inferiority margin of a 0.1 reduction in success probability.The primary outcome analysis supports this claim, with the lower bound of the credible interval above the non-inferiority margin.Evidence: The primary outcome analysis in the Results section.
“The reliability of telesurgery was non-inferior to that of local robotic surgery according to the non-inferiority margin of a 0.1 reduction in success probability.”
ConclusionFind in source - supportedReviewer 2Secondary outcomes did not differ substantially between the two groups.Most secondary outcomes show no statistically significant differences, except surgeon workload (P=0.004) and positive margin rate (OR 13.41, but with wide CI). The claim is supported for most outcomes.Evidence: Table 2 and text.
“Secondary outcomes, including operative basic data, complications, early recovery, oncological outcome, and medical team workload, did not differ substantially between the two groups.”
AbstractFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary outcome is the probability of surgical success, which is a composite of procedural criteria (planned steps, no major injury, no conversion, no postponement due to malfunction). This is a surrogate for clinical benefit, not a hard clinical outcome. The paper does not provide evidence linking this composite to patient-important outcomes, nor does it demonstrate target engagement at the tested dose (e.g., PK/PD) for the intervention. The surrogate is a technical/process measure, and its clinical meaningfulness is not validated.
“The primary outcome was the probability of success of surgery, which the research team defined on the basis of the characteristics of telesurgery. The success was confirmed according to the following determination points: the surgical process was carried out according to the planned steps; no obvious injury to large blood vessels or adjacent organs occurred during the surgery; no conversion of the surgical method occurred... no postponement due to surgical system malfunction occurred.”
- INADEQUATEEffect sizeThe primary effect size is a difference in success probability of 0.02 (95% CrI -0.03 to 0.15) in the ITT analysis, with a non-inferiority margin of 0.1. The observed difference is small and the confidence interval includes the possibility of a reduction in success probability up to 0.15, which exceeds the margin. The effect is not anchored to a minimal clinically important difference; the margin was chosen by expert discussion without established reference. The result is presented as non-inferior, but the magnitude of the effect is not shown to be clinically meaningful.
“The estimated difference in the probability of success in the intention-to-treat and per protocol populations was 0.02 (95% credible interval (CrI) −0.03 to 0.15; bayesian posterior probability 0.99 for non-inferiority) and 0.003 (−0.001 to 0.03; bayesian posterior probability >0.99 for non-inferiority), respectively.”
Data authenticity concerns
1 finding · worst lowAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
- Data look implausibly cleanAssessed
3 integrity concerns flagged (0 high).
- lowdata too cleanThe primary outcome success probability in the telesurgery group is 100% (36/36) in the ITT population, which is unusually perfect but plausible given the small sample and the definition of success.
“No (%) in intention-to-treat population | 36/36 (100) | 34/36 (94.44)”
Table 2Find in source
Reporting gaps
None foundRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
Checked — nothing surfaced.
The introduction cites the historical development of telesurgery, including the 2001 trans-Atlantic cholecystectomy, and acknowledges that previous studies, including the authors' own exploratory trial, focused on feasibility in narrow indications. The rationale for the non-inferiority design is clearly linked to the lack of robust clinical evidence. The limitations of prior research are explicitly addressed by designing a randomised controlled trial.
“Early milestone, such as the 2001 trans-Atlantic cholecystectomy, prioritised technical novelty over clinical rigour.”
“However, previous studies, including our exploratory trial, focused on feasibility in narrow indications, and no clinical evidence has been established to support further research or the wider application of telesurgery.”
“Therefore, we designed this randomised controlled trial in patients undergoing robotic urological procedures (radical prostatectomy or partial nephrectomy) to determine whether telesurgery is non-inferior to standard local surgery in the probability of achieving surgical success”
“Early milestone, such as the 2001 trans-Atlantic cholecystectomy, prioritised technical novelty over clinical rigour.”
“However, previous studies, including our exploratory trial, focused on feasibility in narrow indications, and no clinical evidence has been established to support further research or the wider application of telesurgery.”
“This study was the first randomised controlled trial in the field of telesurgery to analyse the difference in reliability between· telesurgery and local surgery.”
The trial is described as a multicentre, single-blind, non-inferiority RCT. Randomisation was central, stratified by surgery type, with random block sizes of four. Blinding of participants, follow-up specialists, and statisticians is described, with a rationale for not masking the surgical team. A sample size calculation was performed using a simulation method, specifying alpha, power, and non-inferiority margin. Inclusion and exclusion criteria are clearly pre-specified. The analysis population (ITT and per-protocol) and missing data handling (multiple imputation) are defined.
“The participants were randomly assigned in a 1:1 ratio to undergo either telesurgery or standard local robotic surgery, using stratified randomisation with random block sizes of four.”
“Participants, follow-up specialists (independent research nurses or doctors), and independent statisticians were masked to the form of surgery performed (telesurgery or local surgery). As masking of the medical team was not feasible, the determination of surgical success was conducted by the surgical team according to pre-specified criteria to minimise subjectivity.”
“The participants were randomly assigned in a 1:1 ratio to undergo either telesurgery or standard local robotic surgery, using stratified randomisation with random block sizes of four.”
“Participants, follow-up specialists (independent research nurses or doctors), and independent statisticians were masked to the form of surgery performed (telesurgery or local surgery).”
The study reports sex (male), age, height, weight, and body mass index in the baseline characteristics table. Health status is addressed through inclusion/exclusion criteria and disease characteristics (prostate cancer, renal tumour). Demographics are reported in Table 1. Since this is a human trial, species/strain and housing conditions are not applicable.
“Male sex, total | 28 (78) | 26 (72) | 24 (75) | 22 (71)”
“Median (IQR) age, years | 61.0 (57.5-68.0) | 65.0 (56.5-70.0) | 61.0 (57.5-68.0) | 62.0 (53.0-70.5)”
“The inclusion criteria were age 18-80 years; body mass index 18-30; diagnosis of renal tumour or prostate cancer and fit to undergo urological laparoscopic surgery”
“Male sex, total | 28 (78) | 26 (72) | 24 (75) | 22 (71)”
“Median (IQR) age, years | 61.0 (57.5-68.0) | 65.0 (56.5-70.0) | 61.0 (57.5-68.0) | 62.0 (53.0-70.5)”
“Prostate cancer | 20 (56) | 20 (56) | 17 (53) | 16 (52)”
The paper states that the study was approved by the ethical review committee of Chinese PLA General Hospital (ID:2023-021) and the other four trial sites. It also states that participants signed a surgical informed consent and a separate research consent. The trial is registered (ChiCTR2300077721).
“This study was approved by the ethical review committee of Chinese PLA General Hospital (ID:2023-021) and the other four clinical trial sites.”
“Participants signed a surgical informed consent and a separate research consent before enrolment.”
“Trial registration ChiCTR.org ChiCTR2300077721.”
“This study was approved by the ethical review committee of Chinese PLA General Hospital (ID:2023-021) and the other four clinical trial sites.”
“Participants signed a surgical informed consent and a separate research consent before enrolment.”
The surgical robotic system (MP1000, Edge Medical Co, Shenzhen, China) is identified. The telecommunication providers are named. Statistical software (SAS 9.4, RStudio, GraphPad Prism 8.0) is identified. Since this is a device trial, bench criteria like antibodies and cell lines are not applicable.
“We used a four arm, multi-port Surgical Robotic System (MP1000, Edge Medical Co, Shenzhen, China) as the robotic subsystem.”
“We used SAS version 9.4 or RStudio for analyses. We visualised telesurgery monitoring data by using RStudio or GraphPad Prism 8.0.”
“We used a four arm, multi-port Surgical Robotic System (MP1000, Edge Medical Co, Shenzhen, China) as the robotic subsystem.”
“We used SAS version 9.4 or RStudio for analyses.”
The primary analysis uses Bayesian mixed-effects logistic regression, and the paper reports the difference in success probability with a 95% credible interval and posterior probability. Secondary outcomes are analysed with mixed-effects linear regression and Bayesian logistic regression. Fisher's exact test is used for complications. Effect sizes (adjusted Cohen's d) and confidence intervals are reported. The paper reports exact p-values for secondary outcomes. The statistical software is identified. The paper reports the primary outcome by estimation (credible interval) and posterior probability, which is a complete way to report for a Bayesian analysis.
“For the primary outcome, we derived the differences in the probability of success between the groups along with their 95% credible interval from the posterior distributions of the bayesian mixed effect logistic regression with penalised priors accounting for clustering by surgeon, taking the intercept of surgeon as a random effect.”
“We used Fisher’s exact test to analysis the risk difference of the Clavin-Dindo complications without stratification.”
The data availability statement names a concrete repository (Dryad) with a DOI. The code is stated to be in the supplementary materials. This is a clear and concrete access route.
“The data underlying the findings in this paper are openly and publicly available and can be found at Dryad: https://doi.org/10.5061/dryad.t4b8gtjg8 .”
“The code used to analyse the data in the paper can be found in the supplementary materials.”
“The data underlying the findings in this paper are openly and publicly available and can be found at Dryad: https://doi.org/10.5061/dryad.t4b8gtjg8”
“The code used to analyse the data in the paper can be found in the supplementary materials.”
The trial is registered (ChiCTR2300077721). The methods are detailed. The discussion includes a 'Strengths and limitations of study' section. The conclusions are proportional to the evidence. Funding sources and competing interests are declared.
“Despite the ideal performance of telesurgery in this trial, several concerns and limitations remain.”
“Funding: This study was supported by grants from National Natural Science Foundation of China (M-0735), Noncommunicable Chronic Diseases-National Science and Technology Major Project (NO.2024ZD0536000), and Beijing Natural Science Foundation (NO.7254459).”
“Trial registration ChiCTR.org ChiCTR2300077721.”
“Thirdly, although the sample size was estimated on the basis of the predefined non-inferiority margin, we acknowledge that the relatively small cohort in this trial may have limited the ability to detect statistically significant differences in certain outcomes.”
“Funding: This study was supported by grants from National Natural Science Foundation of China (M-0735), Noncommunicable Chronic Diseases-National Science and Technology Major Project (NO.2024ZD0536000), and Beijing Natural Science Foundation (NO.7254459).”
Registered (1 ID: Chinese Clinical Trial Registry). No reporting guideline cited.
Broken references and links
None foundReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Checked — nothing surfaced.
Checked 34 references by DOI: 33 verified — 1 no DOI (shown, not verified).
- NO DOICancer statistics, 2024No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
1 data/code link checked; 1 live.
- dataDryadLIVEHTTP 200https://doi.org/10.5061/dryad.t4b8gtjg8Resolves to Dryad (data repository).
Copyediting
10 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 10 minor suggestions below.
10 copyedit issues flagged: mostly typo, consistency, grammar.
- MINORtypoAbstract“bayesian posterior probability of 0.99 for non-inferiority”→ Bayesian posterior probability of 0.99 for non-inferiorityCapitalize 'Bayesian'.
- MINORgrammarMethods, Statistical analysis“We used Fisher’s exact test to analysis the risk difference”→ We used Fisher’s exact test to analyze the risk differenceIncorrect verb form.
- MINORtypoDiscussion“ontological outcomes”→ oncological outcomesTypo.
- MINORconsistencyTable 2“P (difference in probability of success >−0.1)”→ P (difference in probability of success > -0.1)Inconsistent spacing around the minus sign.
- MINORclarityMethods, Statistical analysis“We did both intention-to-treat analysis and per protocol analysis.”→ We performed both intention-to-treat and per-protocol analyses.Clarity and consistency.
- MINORtypoAbstract, Results“bayesian posterior probability of 0.99 for non-inferiority”→ Capitalize 'Bayesian' for consistency.Inconsistent capitalization of 'Bayesian' throughout the manuscript.
- MINORconsistencyTable 2, footnote“P (difference in probability of success >−0.1)”→ Clarify that this is a posterior probability, not a frequentist p-value.The column header uses 'P' but the values are Bayesian posterior probabilities.
- MINORgrammarMethods, Statistical analysis“We used Fisher’s exact test to analysis the risk difference”→ Change 'to analysis' to 'to analyze'.Grammatical error.
- MINORconsistencyDiscussion, Comparison with other studies“reliability between· telesurgery and local surgery”→ Remove the stray middle dot.Typographical artifact.
- MINORclarityTable 2, footnote“P value was “two tailed weighted posterior probability””→ Clarify the meaning of 'two tailed weighted posterior probability'.Unclear terminology.
The published work is robust and well-reported. An informed reader should weigh the small sample size and the fact that the primary outcome was 100% in the telesurgery group, which may limit generalisability. Minor copyedit issues and the lack of an explicit CONSORT statement are not substantive concerns.
- 1.MEDIUMreportingAdd an explicit statement of adherence to the CONSORT reporting guideline in the Methods or a dedicated section.Both reviewers noted that the reporting guideline is not explicitly mentioned, which is a minor transparency gap.
- 2.MEDIUMethicsAdd a statement on regulatory compliance (e.g., Declaration of Helsinki) in the Ethics section.Reviewer 2 noted that regulatory compliance is implied but not explicitly stated.
- 3.MEDIUMdata codeConsider providing the statistical code in a public repository (e.g., GitHub) with a DOI, in addition to supplementary materials.Reviewer 2 suggested that code sharing could be enhanced beyond supplementary materials.
- 4.MEDIUMstatisticsClarify the handling of missing data for secondary outcomes in the statistical analysis section.Reviewer 2 noted that some secondary outcomes have missing data (e.g., QoR-15 at 6 weeks) and the handling is not fully described.
- 5.MEDIUMreportingDiscuss the potential impact of the high withdrawal rate (12.5%) on the generalisability of the findings in more detail.Reviewer 2 suggested that the withdrawal rate could be discussed more thoroughly.
- 6.MEDIUMreportingProvide a more detailed description of the blinding of outcome assessors for secondary outcomes.Reviewer 2 noted that the paper only mentions masking of follow-up specialists, not outcome assessors for secondary outcomes.
- 7.MEDIUMreportingConsider reporting the results of the per-protocol analysis for secondary outcomes.Reviewer 2 noted that only the primary outcome is reported for both ITT and per-protocol populations.
- 8.MEDIUMotherAdd a note on the validation of the surgical robotic system (MP1000) used in the trial.Reviewer 2 suggested that the domestically produced system's validation could be clarified.
- 9.MEDIUMreportingInclude a statement on patient and public involvement in the trial design.Reviewer 2 noted that the paper notes it was not routine practice, but a statement would be transparent.
- 10.MEDIUMreportingClarify the definition of 'surgical success' in the abstract to ensure consistency with the Methods.Reviewer 2 suggested that the abstract definition could be clearer.
- 11.LOWcopyeditCapitalize 'Bayesian' consistently throughout the manuscript (e.g., in the Abstract and Methods).Copyedit flagged inconsistent capitalization of 'Bayesian'.
- 12.LOWcopyeditFix the grammatical error 'to analysis' to 'to analyze' in the Methods, Statistical analysis section.Copyedit flagged the incorrect verb form.
- 13.LOWcopyeditCorrect the typo 'ontological outcomes' to 'oncological outcomes' in the Discussion.Copyedit flagged the typo.
- 14.LOWcopyeditRemove the stray middle dot in 'reliability between· telesurgery and local surgery' in the Discussion.Copyedit flagged the typographical artifact.
- 15.LOWcopyeditClarify the meaning of 'two tailed weighted posterior probability' in Table 2 footnote.Copyedit flagged the unclear terminology.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.