Evaluate teaching quality of physical education using a hybrid multi-criteria decision-making framework
Chen Z, Luo S.
- DOI
- 10.1371/journal.pone.0280845
- Record issued
- 2026-08-05
- Engine
- 7.15.0
- Exported
- 2026-09-19
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/a632bdfc-71b4-4117-b76f-e45ea9c04d7d is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ReportingStudy design partially met−0.25★
- Statistics were not checked: no recomputable values were found in this text — no test statistic reported with its degrees of freedom, no effect estimate printed with both a 95% CI and a p-value, and no percentage printed with both its count and its denominator.
- No data or code availability links were detected to verify.
- 01Internal contradictions in the reported numbers
At ζ=0.6 in Table 9, the reported RI values (A1≈0.260, A2≈0.257, A3≈-0.327, A4≈-0.843) imply ranking A1≻A2≻A3≻A4 under the paper's rule that larger RI means better, yet the table lists A3 as the best college with ranking A3≻A1≻A2≻A4.
“ζ = 0.6 | (0.216,0.273, 0.165,0.158, 0.188) | RI ( A 1 ) ≈ 0.260, RI ( A 2 ) ≈ 0.257, RI ( A 3 ) ≈ -0.327, RI ( A 4 ) ≈ -0.843. | A 3 ≻ A 1 ≻ A 2 ≻ A 4 | A 3”
Table 9
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
The paper proposes a novel hybrid MCDM framework with a clear scientific premise and adequate reporting transparency, but it is undermined by unresolved internal contradictions in key tables (Table 9 and Table 10) and a lack of code sharing, which limit reproducibility and trust in the illustrative results.
The three independent reviewers largely agreed on the strong dimensions (scientific premise, reporting transparency) and N/A designations, but diverged on study design (N/A vs. warn) and key resources/data code availability (N/A vs. fail/warn). The synthesized judgment follows the scoring rules strictly, resolving the N/A cases, and the warnings reflect the most impactful gaps. The statistics component verified no inferential tests. The citation check found no retracted or missing references.
Numerical inconsistencies
1 finding · worst mediumValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
- mediuminternal contradictionAt ζ=0.6 in Table 9, the reported RI values (A1≈0.260, A2≈0.257, A3≈-0.327, A4≈-0.843) imply ranking A1≻A2≻A3≻A4 under the paper's rule that larger RI means better, yet the table lists A3 as the best college with ranking A3≻A1≻A2≻A4.
“ζ = 0.6 | (0.216,0.273, 0.165,0.158, 0.188) | RI ( A 1 ) ≈ 0.260, RI ( A 2 ) ≈ 0.257, RI ( A 3 ) ≈ -0.327, RI ( A 4 ) ≈ -0.843. | A 3 ≻ A 1 ≻ A 2 ≻ A 4 | A 3”
Table 9 - lowinternal contradictionThe Total gap (TG) values in Table 10 do not consistently follow the stated formula TG = ∑|h_i - H_i|/H_i. For example, for Approach 1 (A4≻A1≻A2≻A3) versus the optimal order A1≻A2≻A4≻A3, the formula yields approximately 2.17, not 4.00; several other rows also appear inconsistent.
T G = ∑ i = 1 m | h i − H i H i | ... Approach 1 ... 4.00
Table 10reviewer’s wording - lowinternal contradictionCriterion C4 is named 'Teaching effect' in Table 1 but is referred to as 'teaching link' in the case study text, which is inconsistent.
“they recognize the most important criterion C 2 and the least important criterion C 4”
Case study, Illustration, paragraph 2
Overstated conclusions
1 finding · worst lowConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Conclusions only partially backed by the presented evidenceAssessed
9 major claims checked against the paper's own evidence: all adequately supported.
- partialReviewers 1, 3The proposed MAIRCA-ELECTRE approach is feasible and the most appropriate among these approaches.The comparison supports feasibility and superiority on this single illustrative dataset, but 'most appropriate' is a generalization from one example, and the Table 10 Total gap values are not consistently reproduced from Eq. (17).Evidence: Table 10 compares seven approaches; the proposed method's ranking matches the aggregate optimal ranking and has the smallest reported Total gap (0.00).
“Overall, results show that the proposed MAIRCA-ELECTRE approach is feasible and the most appropriate among these approaches.”
Discussions - partialReviewers 1, 3The main advantage of the approach is that it can properly deal with non-compensatory criteria.The paper argues that ELECTRE handles non-compensatory criteria, but it does not provide a direct test or demonstration of this property beyond the design rationale.Evidence: The discussion states that ELECTRE can effectively handle non-compensatory criteria, but no dedicated experiment isolates this behavior.
“The main advantage of such approach is that it can properly deal with non-compensatory criteria during the teaching quality evaluation process.”
Discussions, paragraph 1 - supportedReviewers 1, 3The proposed approach is practicable and can provide instructions for teaching-quality evaluation of physical education.The illustrative case and sensitivity/comparison analyses provide adequate evidence that the algorithm is operable and yields a coherent ranking.Evidence: A worked case study (Section 4) with four alternatives and five criteria, plus sensitivity and comparison analyses (Section 5), demonstrates the algorithm's applicability.
“Results show that our approach is practicable and can provide instructions for the teaching quality evaluation of physical education.”
Abstract - supportedReviewer 1MAIRCA has neither been extended in a picture fuzzy environment nor utilized for teaching-quality evaluation.The literature review provides reasonable support for the claimed gap based on the cited works.Evidence: The introduction lists numerous MAIRCA extensions (rough, hesitant-fuzzy, fuzzy, intuitionistic-fuzzy) and none in a picture-fuzzy or teaching-quality context.
“However, the MAIRCA has neither been extended in a picture fuzzy environment, nor been utilized to deal with teaching quality evaluation issues.”
Introduction - supportedReviewer 2The proposed hybrid MAIRCA-ELECTRE method is feasible for evaluating teaching quality of physical education.The paper presents a detailed case study with four alternatives and five criteria, demonstrating the step-by-step application of the method and obtaining a ranking order, which supports feasibility.Evidence: Section 4, Case study, Tables 2-8 and ranking results.
“In this section, a case of teaching quality evaluation of physical education is investigated.”
Section 4, Case study, paragraph 1 - supportedReviewer 2The method can handle non-compensatory criteria during the evaluation process.The ELECTRE method, which is designed for non-compensatory criteria, is used in Phase IV. The paper explicitly states this advantage and the comparison analyses show that other methods (which assume compensation) yield different results.Evidence: Section 3, Phase IV: 'The ELECTRE III method is adopted for obtaining the alternatives’ ranks.'; Section 5.2: 'Approach 1, 2, 3, 4 and 6 cannot deal with the situation where evaluation criteria are not compensatory.'
“In this phase, the ELECTRE III method is adopted for obtaining the alternatives’ ranks.”
Section 3, Phase IV - supportedReviewer 2The proposed MAIRCA-ELECTRE approach outperforms other existing methods in terms of total gap from the optimal ranking.The comparison analysis shows that the proposed method has the smallest total gap (TG=0.00) compared to six other approaches, as shown in Table 10.Evidence: Table 10, last column: Total gap TG values.
“Our approach (MAIRCA-ELECTRE) achieves the smallest gap, followed by Approach 6.”
Table 10 - supportedReviewer 2Picture fuzzy numbers can accommodate diverse viewpoints and describe decision makers’ ambiguity, uncertainty, and inconsistency sufficiently.The paper uses PFNs to represent evaluation data (Table 2) and explains that PFNs have four membership functions to capture different attitudes. The case study illustrates this by showing how decisions are aggregated.Evidence: Section 2, Definition 1 and Section 4, Step 1: 'PFNs are finally used to describe the evaluation results.'
“The most obvious feature of PFNs is it that it contains four different membership functions, which can accommodate diverse viewpoints and describe decision makers’ ambiguity, uncertainty and inconsistency sufficiently.”
Section 2, Definition 1 - supportedReviewer 3The proposed MAIRCA-ELECTRE approach is more pertinent than six other approaches in handling teaching quality evaluation issues.The comparison analysis shows that the proposed method's ranking matches the optimal ranking order derived from a scoring mechanism, and it has the smallest total gap.Evidence: Table 10 lists ranking orders of six approaches and the proposed method; the scoring mechanism yields the optimal ranking A1≻A2≻A4≻A3, which matches the proposed method.
“It is clear that the ranking order with MAIRCA-ELECTRE approach is the same with the optimal ranking order.”
Comparison analyses, paragraph 2
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Study-design details incomplete (controls, blinding, power)Assessed
The introduction cites multiple relevant studies on physical education teaching quality evaluation and picture fuzzy decision-making. It acknowledges limitations such as qualitative methods or information loss, and explains how the hybrid MAIRCA-ELECTRE method with PFNs addresses these gaps. The rationale logically links the premise to the study objectives.
“However, these methods are qualitative or applied only to certain decision-making environments.”
“However, they may lead to information loss or distortion during the aggregation process.”
“Nevertheless, the classical MAIRCA approach with exact numbers cannot capture decision makers’ ambiguity and vagueness.”
“However, these methods are qualitative or applied only to certain decision-making environments.”
“The major innovation and contribution are: First, PFNs are utilized to depict complex and fuzzy evaluation information.”
“However, the MAIRCA has neither been extended in a picture fuzzy environment, nor been utilized to deal with teaching quality evaluation issues.”
The four phases of the proposed framework are fully described. Sensitivity analysis and comparison with six alternative approaches serve as an adequate baseline/ablation analogue. However, the case study is introduced as a 'Suppose' scenario with no rationale for choosing four colleges, five criteria, and ten decision makers, and no data-curation or inclusion/exclusion criteria are reported. Randomization and blinding are not applicable to a computational illustration.
“Suppose there are four physical education colleges { A 1 , A 2 , A 3 , A 4 }, the proposed methodology in Section 4 is adopted to evaluate their teaching quality.”
“ten decision makers (including four government representatives related to physical education, four professionals and scholars in the field of physical education, and two sports practitioners) are organized to make assessments”
“To testify the sensitivity of the proposed method, the variations of ranking results are analyzed by changing the criteria weights”
“Several different decision-making approaches are applied to deal with the teaching quality evaluation of physical education.”
The paper is a computational MCDM framework applied to a hypothetical teaching quality evaluation problem. It does not involve any animals, human subjects, or biological materials. Therefore, all sub-criteria are not applicable.
The study does not involve any human participants, animals, or non-public subject-level data. It is a purely computational methodology paper with a synthetic illustrative example. Therefore, all ethical approval sub-criteria are not applicable.
The paper proposes a mathematical framework and does not report any antibodies, cell lines, organisms, reagents, or custom software. The methods are described algorithmically, but no specific software tool is identified. Since no key biological or chemical resources are used, the dimension is not_applicable.
There are no hypothesis tests, p-values, confidence intervals, or effect-size estimates. Spearman's rho is computed only as a descriptive similarity measure (λ and ℏ) without a test, so this dimension is not applicable.
“the Spearman’s rank-correlation test [] is adopted with the following two parameters: ƛ = 1 − 6 ⋅ ∑ i = 1 m ( d i s ( A i ) ) 2 m ( m 2 − 1 )”
The paper includes a data availability statement stating that all relevant data are within the manuscript. The data used in the case study are fully presented in tables. No additional repository deposit or code sharing is applicable since the data are fully included and no custom code was developed (the method is described mathematically).
“Data Availability All relevant data are within the manuscript.”
“Data Availability All relevant data are within the manuscript.”
“Data Availability All relevant data are within the manuscript.”
The methods are described comprehensively in Sections 3 and 4. The case study is presented with all steps and tables. Limitations are discussed in the conclusion (e.g., equal expert weights, complete rationality). The conclusions are proportional to the evidence. Funding sources and a competing interests statement are provided. Trial registration is not applicable.
“experts are assumed to have equal weights, which may be not appropriate in some cases.”
“The authors have declared that no competing interests exist.”
“One limitation of the model may be that all decision makers are assumed to have complete rationality and equal weights, which should be overcome by introducing other theories in the future.”
“However, there are still some limitations of the proposed methodology.”
“The authors have declared that no competing interests exist.”
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 41 references by DOI: 3 verified — 38 no DOI (shown, not verified).
- NO DOIAssessment for learning in physical education: The what, why and howNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntegrating assessment into physical education teachingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEvaluation of physical education teaching quality in colleges based on the hybrid technology of data mining and hidden Markov modelNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIResearch on teaching quality evaluation model of physical education based on simulated annealing algorithmNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEvaluation of physical education teaching based on web embedded system and virtual realityNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOICombining psychology, a Game Sense Approach and the Aboriginal game Buroinjin to teach quality physical educationNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEvaluation of teaching quality of public physical education in colleges based on the fuzzy evaluation theoryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIIntuitionistic fuzzy similarity and information measures with physical education teaching quality assessmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIResearch on the teaching quality evaluation of physical education with intuitionistic fuzzy TOPSIS methodNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPicture fuzzy setsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOILikelihood-based hybrid ORESTE method for evaluating the thermal comfort in underground minesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPicture fuzzy aggregation operators and their application to multiple attribute decision makingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPicture fuzzy Dombi aggregation operators: application to MADM processNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISome picture fuzzy Bonferroni mean operators with their application to multicriteria decision makingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPicture fuzzy tensor and its application in multi-attribute decision makingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIInterval-valued picture fuzzy Maclaurin symmetric mean operator with application in multiple attribute decision-makingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn integrated EDAS-ELECTRE method with picture fuzzy information for cleaner production evaluation in gold minesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA multi-criteria decision-making framework for risk ranking of energy performance contracting project under picture fuzzy environmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIDifferent approaches to multi-criteria group decision making problems for picture fuzzy environmentNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn extended bidirectional projection method for picture fuzzy MAGDM and its application to safety assessment of construction projectNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIHybrid PSO-WDBA method for the site selection of tailings pondNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEvaluating public transport service quality using picture fuzzy analytic hierarchy process and linear assignment modelNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIERP selection using picture fuzzy CODAS methodNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA fuzzy-multi attribute decision making approach for efficient service selection in cloud environmentsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEvaluating the satisfaction level of citizens in municipality services by using picture fuzzy VIKOR method: 2014–2019 period analysisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISelection of railway level crossings for investing in security equipment using hybrid DEMATEL-MARICA modelNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOINew hybrid multi-criteria decision-making DEMATELMAIRCA model: sustainable selection of a location for the development of multimodal logistics centreNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEvaluating the performance of suppliers based on using the R’AMATEL-MAIRCA method for green supply chain implementation in electronics industryNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn integrated MC-HFLTS&MAIRCA method and application in cargo distribution companiesNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn integrated approach for fuzzy failure modes and effects analysis using fuzzy AHP and fuzzy MAIRCANo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImplementation of MCDM-based integrated approach to identifying the uncertainty factors on the constructional projectNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA fuzzy AHP-MAIRCA Model for overtourism assessment: The Case of Malaga ProvinceNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIEvaluating biological inspiration for biologically inspired design: An integrated DEMATEL‐MAIRCA based on fuzzy rough numbersNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIAn aggregation technique for optimal decision-making in materials selectionNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA neutrosophic normal cloud and its application in decision-makingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOISustainable supplier selection in healthcare industries using a new MCDM method: Measurement of alternatives and ranking according to COmpromise solution (MARCOS)No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIA novel IMF SWARA-FDWGA-PESTEL analysis for assessment of healthcare systemNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIImpact of the number of vehicles on traffic safety: multiphase modelingNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
1 finding · worst lowWording, consistency and formatting errors that need correcting before submission.
- Wording or formatting errors that need correctingAssessed
12 copyedit issues flagged (2 major): mostly typo, other, consistency.
- MAJORotherReferences section, trailing line“t of alternatives and ranking according to COmpromise solution (MARCOS) . Computers & Industrial Engineering , 140 , 106231 .”→ Restore the full reference title; likely 'Technique of alternatives and ranking according to COmpromise solution (MARCOS)'.The reference appears truncated/garbled, making it unverifiable.
- MAJORconsistencyCase study, Illustration, paragraph 2“they recognize the most important criterion C 2 and the least important criterion C 4”→ they recognize the most important criterion C 2 and the least important criterion C 5C4 is 'Teaching effect', not 'Teaching link' (C5). The naming is inconsistent with Table 1.
- MINORtypoIntroduction, Williams et al. sentence“discussed the teach quality of physical education”→ change 'teach quality' to 'teaching quality'Simple typo in a key term.
- MINORtypoIntroduction, Gireesha et al. sentence“to assess could service”→ change 'could service' to 'cloud service'Context indicates 'cloud service'.
- MINORgrammarDiscussions“can effectively handle evaluation criteria with non-compensate.”→ change 'with non-compensate' to 'with non-compensatory criteria'Incomplete grammatical phrase.
- MINORtypoSection 5, Sensitivity analyses“Toss explore the relationship between ranking orders under different ζ valuesss”→ change 'Toss' to 'To' and 'valuesss' to 'values'Typo in section heading/introductory sentence.
- MINORtypoSection 5, Approach 6 description“the better the physical education college A iss”→ change 'A iss' to 'A i'Corrupted subscript in text.
- MINORtypoSection 5.1, Sensitivity analyses, paragraph beginning 'Toss explore'“Toss explore the relationship between ranking orders...”→ Change to 'To explore the relationship between ranking orders...'Typo: 'Toss' should be 'To'.
- MINORotherSection 5.1, Sensitivity analyses, last line of text before references“sss”→ Remove extraneous 'sss'Appears to be a stray character.
- MINORconsistencySection 2, Preliminaries, Definition 6“d ( α 1 , α 2 ) = ( | a 1 − a 2 | 2 + | b 1 − b 2 | 2 + | c 1 − c 2 | 2 + | e 1 − e 2 | 2 ) .”→ Add square root symbol: d ( α 1 , α 2 ) = sqrt( | a 1 − a 2 |^2 + | b 1 − b 2 |^2 + | c 1 − c 2 |^2 + | e 1 − e 2 |^2 ).The formula is missing the square root; can be clarified.
- MINORclaritySection 5.1, Sensitivity analyses, Table 9 footnote“ƛ and ℏ values are computed...”→ Define ƛ and ℏ in the text before the table, or in the legend.The symbols are defined in equations (15)-(16) but not in the table caption.
- MINORtypoSensitivity analyses, paragraph 3“Toss explore the relationship”→ To explore the relationshipTypo: 'Toss' should be 'To'.
This published paper has significant internal contradictions in its main results tables that warrant a correction or erratum. An informed reader should withhold trust in the illustrative rankings until the authors explain or correct the inconsistencies in Table 9 (ζ=0.6 row) and Table 10 (Total gap values). The absence of shared code also hinders independent verification. The paper's methodological framework remains of interest, but the numerical demonstration is unreliable in its current form.
- 1.HIGHrigorCorrect the internal contradiction in Table 9 at ζ=0.6: the reported RI values (A1≈0.260, A2≈0.257, A3≈-0.327, A4≈-0.843) imply ranking A1≻A2≻A3≻A4, but the table lists A3 as the best. Recompute the RI values or re-rank and update the sensitivity discussion.This inconsistency undermines the credibility of the sensitivity analysis and the main illustrative results.
- 2.HIGHrigorRecompute the Total gap (TG) values in Table 10 using Eq. (17) and correct entries that are inconsistent with the formula (e.g., Approach 1 should yield ~2.17, not 4.00).The TG values are central to the comparison analysis; incorrect values misrepresent the relative performance of competing methods.
- 3.HIGHreportingIssue a correction or erratum to address the two internal contradictions above (Table 9 and Table 10) and clarify which values are correct.Published results with clear numerical errors require formal correction to maintain scientific integrity.
- 4.HIGHdata codeDeposit the code or a detailed spreadsheet implementing the MAIRCA-ELECTRE calculations in a public repository (e.g., Zenodo, GitHub) and update the data availability statement to include the repository link.Reproducibility of the numerical results is impossible without the computational steps; code sharing is expected for a computational method paper.
- 5.HIGHreportingAdd a justification for the choice of 4 alternatives, 5 criteria, and 10 decision makers in the case study, or explicitly state that the example is purely illustrative and the numbers are arbitrary.The lack of rationale for the dataset dimensions leaves the reader wondering whether the results are sensitive to these choices.
- 6.HIGHreportingCorrect the garbled MARCOS reference in the References section (missing title) and ensure all references are complete and verifiable.An unverifiable reference is a citation integrity concern and may indicate sloppy proofreading.
- 7.MEDIUMcopyeditFix the typo 'teach quality' to 'teaching quality' in the Introduction.Standard copyedit issue; the key term is misspelled.
- 8.MEDIUMcopyeditFix 'could service' to 'cloud service' in the Introduction (Gireesha et al. sentence).Contextually clear typo that should be corrected for clarity.
- 9.MEDIUMcopyeditFix 'non-compensate' to 'non-compensatory criteria' in the Discussions section.Incomplete grammatical phrase hinders readability.
- 10.MEDIUMcopyeditFix 'Toss explore' to 'To explore' and 'valuesss' to 'values' in Section 5.1 (Sensitivity analyses).Multiple typos in the same sentence suggest a need for thorough proofreading.
- 11.MEDIUMcopyeditRemove the stray 'sss' at the end of the Sensitivity analyses section before the references.Likely a copy-paste artifact that should be deleted.
- 12.MEDIUMotherConfirm criterion C4 name: Table 1 lists 'Teaching effect' but the case study text refers to it as 'teaching link'; use consistent terminology throughout.Inconsistent naming confuses the reader and may indicate a deeper error in the criteria set.
- 13.LOWreportingAdd a statement clarifying whether the decision makers are real or hypothetical; if hypothetical, remove the implication of empirical recruitment that would require ethics approval.Avoids any ambiguity about the nature of the case study.
- 14.LOWreportingConsider adding a reporting guideline checklist (e.g., a general transparency checklist) for computational methods, or state that none is applicable.Enhances completeness, though not strictly required for a methodological paper.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.