Second-Line Antiretroviral Therapy for Children Living with HIV in Africa.
Musiime V, Bwakura-Dangarembizi M, Szubert AJ, Mumbiro V, Mujuru HA, Kityo CM, Lugemwa A, Doerholt K, Chabala C, Makumbi S, Mulenga V, McIlleron H, Burger D, Natukunda E, Shakeshaft C, Linda KJ, Nathoo K, Monkiewicz L, Yawe I, Kapasa M, Nyathi M, Lungu J, Nduna B, Ndebele W, South A, Mwamabazi M, Musoro G, Griffiths A, Zyambo K, Nazzinda R, Zimba K, Zhang Y, Walker S, Turkova A, Walker AS, Bamford A, Gibb DM, CHAPAS-4 Trial Team
- DOI
- 10.1056/NEJMoa2404597
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/3af7ead5-d707-4892-8145-c21a19bbd57e is authoritative.
How this rating was calculated
- IntegrityIntegrity concern ×3−1.5★
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run: 8 reported means were read, and their group size is not stated where the values are printed (this source has no machine-readable table structure). These checks need the count the mean was averaged over, so none was performed.
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy endpoint is virological suppression (viral load <400 copies/mL) at week 96, which is a surrogate marker for clinical outcomes such as disease progression and mortality. The manuscript does not provide evidence linking this surrogate to hard clinical outcomes in this pediatric population, nor does it demonstrate target engagement (e.g., PK/PD) at the tested doses. The claim of 'effective' is based on this surrogate without a validated surrogate-to-clinical-outcome link.
“Primary endpoint was week-96 viral load (VL)<400copies/mL”
- 02Treatment effect not shown to be clinically meaningful
The reported effect sizes are small absolute differences in virological suppression (e.g., +6.3% for TAF vs SOC, +9.7% for DTG vs LPV/r+ATV/r) and are not anchored to a minimal clinically important difference or to clinical meaningfulness. The manuscript does not state what magnitude of difference would be clinically meaningful, and the differences are presented as statistically significant without explicit clinical interpretation.
“TAF/FTC was superior to SOC (adjusted difference [95% CI] VL<400copies/mL +6.3%[2.0%,10.6%],p=0.004)”
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-designed and rigorously reported randomized clinical trial (CHAPAS-4) with strong methodology across most dimensions: clear scientific premise, robust study design, adequate reporting of biological variables, ethics, key resources, statistics, and transparency. The main weakness is the lack of a formal data availability statement and repository deposit, which is a common but important reporting gap for a data-driven clinical trial.
This is a post-publication audit of a published randomized controlled trial. Both reviewers independently scored all eight dimensions and agreed on all statuses; no divergence required reconciliation. The statistics verification covered only a subset of reported tests (7 of many), and the citation check found no retracted or unresolved references. The integrity check flagged minor internal inconsistencies (e.g., SAE percentage 3% vs 3.2%) that are not validity threats.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 7 tests: 7 consistent, 0 inconsistent; 7 via agent-written checks.
- CONSISTENTreported p = .004 · recomputed p = .004Reviewer 1Primary backbone comparison: TAF vs SOC at week-96 VL<400 copies/mL
“At week-96, 406/454(89.4%) TAF/FTC vs. 378/454(83.3%) SOC had VL <400copies/mL (adjusted difference +6.3% [95% confidence interval (CI) +2.0%,+10.6%]; p=0.004)”
Taken as given: The adjusted difference is 6.3% (0.063) with 95% CI 2.0% to 10.6%.; The p-value is two-sided from a logistic regression model.; The CI is a 95% confidence interval.Method: Recomputed p-value from the reported estimate and 95% CI using the normal approximation (pCI function).How we recomputed it: pCI(0.063, 0.020, 0.106, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewer 1Primary anchor comparison: DTG vs LPV/r+ATV/r combined at week-96 VL<400 copies/mL
“DTG was superior to LPV/r and ATV/r arms combined (adjusted difference +9.7% [95%CI +4.8%,+14.5%]; p<0.001).”
Taken as given: The adjusted difference is 9.7% (0.097) with 95% CI 4.8% to 14.5%.; The p-value is two-sided from a logistic regression model.; The CI is a 95% confidence interval.Method: Recomputed p-value from the reported estimate and 95% CI using the normal approximation (pCI function).How we recomputed it: pCI(0.097, 0.048, 0.145, 0) - CONSISTENTreported p = .040 · recomputed p = .040Reviewer 1Secondary anchor comparison: DRV/r vs LPV/r+ATV/r combined at week-96 VL<400 copies/mL
“DRV/r was not superior to LPV/r and ATV/r combined as the comparison did not meet pre-specified significance (adjusted difference +5.6% [+0.3%,+11.0%]; p=0.04 vs. threshold p=0.03 from multiple comparisons).”
Taken as given: The adjusted difference is 5.6% (0.056) with 95% CI 0.3% to 11.0%.; The p-value is two-sided from a logistic regression model.; The CI is a 95% confidence interval.Method: Recomputed p-value from the reported estimate and 95% CI using the normal approximation (pCI function).How we recomputed it: pCI(0.056, 0.003, 0.110, 0) - CONSISTENTreported p = .004 · recomputed p = .004Reviewer 2Backbone primary endpoint: TAF/FTC vs SOC difference in VL<400 at week 96
“At week-96, 406/454(89.4%) TAF/FTC vs. 378/454(83.3%) SOC had VL <400copies/mL (adjusted difference +6.3% [95% confidence interval (CI) +2.0%,+10.6%]; p=0.004)”
Taken as given: The adjusted difference is 0.063 (6.3%); The 95% CI is (0.020, 0.106); The p-value is two-sided from a normal approximation using the CIMethod: Recomputed p from estimate and CI using normal approximation (pCI).How we recomputed it: pCI(0.063, 0.020, 0.106, 0) - CONSISTENTreported p = .001 · recomputed p = <.001Reviewer 2Anchor primary endpoint: DTG vs LPV/r+ATV/r combined difference in VL<400 at week 96
“DTG was superior to LPV/r and ATV/r arms combined (adjusted difference +9.7% [95%CI +4.8%,+14.5%]; p<0.001)”
Taken as given: The adjusted difference is 0.097 (9.7%); The 95% CI is (0.048, 0.145); The p-value is two-sided from a normal approximation using the CIMethod: Recomputed p from estimate and CI using normal approximation (pCI).How we recomputed it: pCI(0.097, 0.048, 0.145, 0) - CONSISTENTreported p = .040 · recomputed p = .040Reviewer 2Anchor primary endpoint: DRV/r vs LPV/r+ATV/r combined difference in VL<400 at week 96
“DRV/r was not superior to LPV/r and ATV/r combined as the comparison did not meet pre-specified significance (adjusted difference +5.6% [+0.3%,+11.0%]; p=0.04 vs. threshold p=0.03 from multiple comparisons).”
Taken as given: The adjusted difference is 0.056 (5.6%); The 95% CI is (0.003, 0.110); The p-value is two-sided from a normal approximation using the CIMethod: Recomputed p from estimate and CI using normal approximation (pCI).How we recomputed it: pCI(0.056, 0.003, 0.110, 0) - CONSISTENTreported p = .330 · recomputed p = .327Reviewer 2Anchor non-inferiority: ATV/r vs LPV/r difference in VL<400 at week 96
“ATV/r was non-inferior to LPV/r (adjusted difference +3.4% [-3.4%,+10.2%]; p=0.33).”
Taken as given: The adjusted difference is 0.034 (3.4%); The 95% CI is (-0.034, 0.102); The p-value is two-sided from a normal approximation using the CIMethod: Recomputed p from estimate and CI using normal approximation (pCI).How we recomputed it: pCI(0.034, -0.034, 0.102, 0)
- lowinternal contradictionThe abstract states 'DRV/r was not superior (+5.6%[+0.3%,+11.0%],p=0.04 vs. multiple-comparison adjusted threshold p=0.03)' but the results section says 'p=0.04 vs. threshold p=0.03 from multiple comparisons'. This is consistent, but the abstract's phrasing 'not superior' might be misinterpreted as a negative result when it is actually a borderline non-significant result.
“DRV/r was not superior (+5.6%[+0.3%,+11.0%],p=0.04 vs. multiple-comparison adjusted threshold p=0.03)”
AbstractFind in source - lowinternal contradictionThe abstract states '29(3%) had serious adverse events' while Table 2 shows 29 (3.2%) with 31 events. The percentage differs slightly (3% vs 3.2%).
Abstract: '29(3%) had serious adverse events'; Table 2: 'Serious adverse event 14 (3.0%) 14 | 15 (3.3%) 17 | ... | 29 (3.2%) 31'
Table 2reviewer’s wording - lowinternal contradictionThe abstract reports '497(54.1%) male' while Table 1 shows '497 (54.1%)' for total male, consistent. No issue.
Abstract: '497(54.1%) male'; Table 1: 'Male ... 497 (54.1%)'
Table 1reviewer’s wording
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
5 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2TAF/FTC is superior to SOC for virological suppression at week-96.The primary endpoint analysis shows a statistically significant adjusted difference of +6.3% (95% CI 2.0-10.6%, p=0.004), supporting superiority.Evidence: Adjusted difference +6.3% [95% CI +2.0%,+10.6%]; p=0.004
“At week-96, 406/454(89.4%) TAF/FTC vs. 378/454(83.3%) SOC had VL <400copies/mL (adjusted difference +6.3% [95% confidence interval (CI) +2.0%,+10.6%]; p=0.004)”
ResultsFind in source - supportedReviewers 1, 2DTG is superior to LPV/r and ATV/r combined for virological suppression.The pre-specified comparison shows a statistically significant adjusted difference of +9.7% (95% CI 4.8-14.5%, p<0.001), supporting superiority.Evidence: Adjusted difference +9.7% [95%CI +4.8%,+14.5%]; p<0.001
“DTG was superior to LPV/r and ATV/r arms combined (adjusted difference +9.7% [95%CI +4.8%,+14.5%]; p<0.001).”
ResultsFind in source - supportedReviewers 1, 2DRV/r is not superior to LPV/r and ATV/r combined.The adjusted difference of +5.6% (95% CI 0.3-11.0%) has p=0.04, which does not meet the pre-specified multiple-comparison threshold of p=0.03, so the claim of non-superiority is supported.Evidence: Adjusted difference +5.6% [+0.3%,+11.0%]; p=0.04 vs. threshold p=0.03
“DRV/r was not superior to LPV/r and ATV/r combined as the comparison did not meet pre-specified significance (adjusted difference +5.6% [+0.3%,+11.0%]; p=0.04 vs. threshold p=0.03 from multiple comparisons).”
ResultsFind in source - supportedReviewers 1, 2ATV/r is non-inferior to LPV/r.The adjusted difference of +3.4% (95% CI -3.4% to +10.2%) is within the pre-specified 12% non-inferiority margin, supporting non-inferiority.Evidence: Adjusted difference +3.4% [-3.4%,+10.2%]; p=0.33
“ATV/r was non-inferior to LPV/r (adjusted difference +3.4% [-3.4%,+10.2%]; p=0.33).”
ResultsFind in source - supportedReviewers 1, 2Second-line ART including TAF/FTC and DTG are effective for children without evident safety concerns.The trial demonstrates superior virological suppression with TAF/FTC and DTG, with no significant safety differences across arms, supporting the claim.Evidence: Primary and secondary outcomes show efficacy and safety; only one death and no between-arm differences in serious AEs.
“Second-line ART including TAF/FTC and DTG are effective for children without evident safety concerns.”
ConclusionFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy endpoint is virological suppression (viral load <400 copies/mL) at week 96, which is a surrogate marker for clinical outcomes such as disease progression and mortality. The manuscript does not provide evidence linking this surrogate to hard clinical outcomes in this pediatric population, nor does it demonstrate target engagement (e.g., PK/PD) at the tested doses. The claim of 'effective' is based on this surrogate without a validated surrogate-to-clinical-outcome link.
“Primary endpoint was week-96 viral load (VL)<400copies/mL”
- INADEQUATEEffect sizeThe reported effect sizes are small absolute differences in virological suppression (e.g., +6.3% for TAF vs SOC, +9.7% for DTG vs LPV/r+ATV/r) and are not anchored to a minimal clinically important difference or to clinical meaningfulness. The manuscript does not state what magnitude of difference would be clinically meaningful, and the differences are presented as statistically significant without explicit clinical interpretation.
“TAF/FTC was superior to SOC (adjusted difference [95% CI] VL<400copies/mL +6.3%[2.0%,10.6%],p=0.004)”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites prior research on pediatric ART, TAF, DTG, and boosted PIs, acknowledging gaps in pediatric data. The rationale for comparing TAF vs SOC and different anchor drugs is clearly linked to the study objectives. Limitations of prior research (e.g., small single-arm trials, adult data) are addressed by the trial's design and discussion.
“There are minimal data on TAF in African children; the first paediatric pharmacokinetic data showed tenofovir concentrations equivalent to those safe and effective in adults.”
“CHAPAS-4 compared efficacy, safety and tolerability of different second-line anchor drugs combined with TAF-based or SOC backbone in African children aged 3-15 years.”
“concerns about bone and renal toxicity and lack of paediatric formulations limit paediatric TDF use.”
“A tenofovir disoproxil fumarate (TDF)-based backbone is recommended for first and second-line ART for adolescents >30kg; INSTI-based regimens including tenofovir demonstrate robust efficacy when compared to ritonavir-boosted PI-based regimens including zidovudine in adult second-line trials.”
“There are minimal data on TAF in African children; the first paediatric pharmacokinetic data showed tenofovir concentrations equivalent to those safe and effective in adults.”
“The superior virological suppression of 89.4% at 96 weeks observed with TAF/FTC is comparable to the 93-100% reported in four small single-arm paediatric trials of TAF.”
Randomization method (computer-generated permuted blocks) and unit (individual child) are clearly described. The open-label design is stated with a rationale (objective primary endpoint). Power analysis is provided for both backbone and anchor comparisons. Inclusion/exclusion criteria are pre-specified. Outlier handling is addressed via ITT and per-protocol analyses. Controls are inherent in the factorial design (comparator arms). Independent replication is not applicable for a single pivotal trial.
“A computer-generated sequential randomisation list with variably sized permuted blocks was prepared by the trial statistician and incorporated securely into an online database.”
“For the backbone randomisation, assuming 80.0%-87.5% SOC achieved VL<400 copies/ml at week-96, 920 children provided ≥95% power to demonstrate TAF was non-inferior (10% margin) (two-sided alpha=0.05), assuming 2.5% loss-to-follow-up”
“The open-label design of the trial could have potentially introduced bias; however the primary endpoint (VL) was objective.”
“A computer-generated sequential randomisation list with variably sized permuted blocks was prepared by the trial statistician and incorporated securely into an online database.”
“The open-label design of the trial could have potentially introduced bias; however the primary endpoint (VL) was objective.”
Sex is reported (54.1% male). Age, weight, height, BMI, CD4 count, and WHO stage are reported in baseline table. Demographics include age, sex, and clinical characteristics. Species/strain and housing are not applicable for a human trial. Sex justification is not applicable as both sexes are enrolled.
“497(54.1%) children were male; median age 10 years (IQR 8,13)”
“497(54.1%) children were male; median age 10 years (IQR 8,13)”
The trial was approved by named ethics committees in Uganda, Zambia, Zimbabwe, and UK. Informed consent and assent are described. Regulatory compliance is implied by adherence to national guidelines and the Declaration of Helsinki (not explicitly named but implied).
“Guardians provided written informed consent, with additional assent from older children, according to national guidelines.”
“Guardians provided written informed consent, with additional assent from older children, according to national guidelines.”
The investigational drugs (TAF/FTC, ABC/3TC, ZDV/3TC, DTG, DRV/r, ATV/r, LPV/r) are named with manufacturers (e.g., Gilead, ViiV, Janssen). Doses and regimens are implied. Statistical software (Stata version 17.0) is identified. Antibodies, cell lines, and mycoplasma testing are not applicable as this is a clinical trial without wet-lab assays.
“Participants were randomised to one of two backbones (TAF/FTC or standard-of-care (SOC) (abacavir (ABC)/3TC or ZDV/3TC, whichever not used first-line)) and simultaneously to one of four anchor drugs (DTG, DRV/r, ATV/r, LPV/r).”
“Analyses were intention-to-treat using Stata (version 17.0).”
“European Developing Country Clinical Trial Partnership (funder), and pharmaceutical companies donating additional funding (Gilead Sciences, Johnson and Johnson) and drugs (ViiV Healthcare, Gilead Sciences, Johnson and Johnson, CIPLA)”
“Analyses were intention-to-treat using Stata (version 17.0).”
Statistical tests are named (logistic regression, Cox regression, GEE). Assumptions are handled by design (ITT, per-protocol). Exact p-values are reported (e.g., p=0.004). Effect sizes with 95% CIs are reported. Software is identified. Data presentation includes per-group n and CIs. Mathematical plausibility is not applicable for large-N continuous outcomes.
“Primary endpoint analyses used logistic regression (adjusting for stratification factors), then marginal estimation of risk differences.”
“At week-96, 406/454(89.4%) TAF/FTC vs. 378/454(83.3%) SOC had VL <400copies/mL (adjusted difference +6.3% [95% confidence interval (CI) +2.0%,+10.6%]; p=0.004)”
“Primary endpoint analyses used logistic regression (adjusting for stratification factors), then marginal estimation of risk differences.”
“adjusted difference +6.3% [95% confidence interval (CI) +2.0%,+10.6%]; p=0.004”
“DTG was superior to LPV/r and ATV/r arms combined (adjusted difference +9.7% [95%CI +4.8%,+14.5%]; p<0.001)”
No data availability statement is present. The protocol is available at a URL, but raw data are not deposited in a public repository. For a clinical trial, managed access is acceptable, but no mechanism is described. Code sharing is not applicable as no custom code is mentioned.
“protocol: www.mrcctu.ucl.ac.uk/studies/all-studies/c/chapas-4”
“Full study details can be found in the protocol at nejm.org.”
Trial is registered (ISRCTN22964075). Reporting guideline (CONSORT) is implied by the CONSORT flow diagram. All pre-specified outcomes are reported. Limitations are discussed (open-label, generalisability). Conclusions are proportional to evidence. Funding and COI are disclosed.
“(ISRCTN22964075)”
“The open-label design of the trial could have potentially introduced bias; however the primary endpoint (VL) was objective.”
“The main funding for this study is provided by the European and Developing Countries Clinical Trials Partnership.”
“(ISRCTN22964075)”
“A limitation is that CHAPAS-4 does not provide direct evidence to inform anchor/backbone choice in this situation; however, safety and efficacy could be inferred (given lack of evidence of interaction) and they will undoubtedly remain important future options.”
“The CHAPAS-4 trial was sponsored by University College London (UCL), with central management by the Medical Research Council (MRC) Clinical Trials Unit at UCL supported by MRC core funding (MC_UU_00004/03).”
Registered (1 ID: ISRCTN). Reporting guideline cited: CONSORT.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 22 references by DOI: 19 verified — 3 no DOI (shown, not verified).
- NO DOIDolutegravir with recycled NRTIs is noninferior to PI-based ART: VISEND trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIConsolidated guidelines on HIV prevention, testing, treatment, service delivery and monitoring: recommendations for a public health approachNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPriorities for antiretroviral drug optimization in adults and children: report of a CADO, PADO and HIVResNet joint meeting, 27 September–15 October 2021No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
3 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 3 minor suggestions below.
3 copyedit issues flagged: mostly typo, consistency, grammar.
- MINORtypoAbstract, Results“17,573copies/ml[5,549-55,700]”→ Add space: '17,573 copies/ml [5,549-55,700]'Missing space before bracket.
- MINORconsistencyMethods, paragraph 2“protocol: www.mrcctu.ucl.ac.uk/studies/all-studies/c/chapas-4 (http://www.mrcctu.ucl.ac.uk/studies/all-studies/c/chapas-4)”→ Use a single URL format.Duplicate URL with and without http.
- MINORgrammarDiscussion, paragraph 4“These findings, alongside the additional benefits of smaller pill size, once-daily administration, lower cost and lower risk of hypersensitivity, make TAF a valuable second-line option.”→ Consider rephrasing for clarity: 'These findings, along with the additional benefits of smaller pill size, once-daily administration, lower cost, and lower risk of hypersensitivity, make TAF a valuable second-line option.'Minor punctuation.
The published work is methodologically robust and the findings are well-supported. An informed reader should weigh the absence of a data availability statement and the minor internal inconsistencies (e.g., SAE percentage) as reporting gaps, but these do not undermine the main conclusions. No erratum is warranted for the core results, though the authors should consider issuing a data availability statement and clarifying the SAE percentage discrepancy.
- 1.HIGHdata codeAdd a data availability statement to the manuscript (e.g., in the Methods or a dedicated section) specifying how de-identified individual participant data can be accessed, such as via a managed access committee or a repository like ClinicalStudyDataRequest.The paper currently lacks a data availability statement, which is a key reporting requirement for clinical trials and a common reason for reader concern about reproducibility.
- 2.HIGHdata codeDeposit the statistical analysis code (e.g., Stata do-files) in a public repository such as Zenodo with a DOI, and reference it in the manuscript.Sharing analysis code enhances reproducibility and is a low-cost way to strengthen the paper's transparency.
- 3.MEDIUMreportingClarify the dosing regimens for each investigational drug in the main text or supplement, as the current text implies but does not fully detail them.Complete methods are essential for replication and for readers to assess the clinical applicability of the results.
- 4.MEDIUMreportingExplicitly state adherence to CONSORT reporting guidelines in the Methods or acknowledgments, and provide a CONSORT checklist as supplementary material.Explicitly confirming CONSORT adherence and providing the checklist helps readers verify reporting completeness.
- 5.MEDIUMreportingClarify the regulatory compliance framework (e.g., Declaration of Helsinki) in the ethics section.Explicitly naming the ethical framework strengthens the ethics reporting.
- 6.MEDIUMreportingReconcile the discrepancy in the serious adverse event percentage between the abstract (3%) and Table 2 (3.2%) and ensure consistent reporting.Internal inconsistencies, even minor, can undermine reader trust in the accuracy of the reported data.
- 7.LOWcopyeditFix the missing space in the abstract: change '17,573copies/ml[5,549-55,700]' to '17,573 copies/ml [5,549-55,700]'.Minor typographical errors detract from the paper's professionalism.
- 8.LOWcopyeditStandardize the protocol URL format in Methods, paragraph 2, to use a single consistent format (e.g., with or without 'http://').Duplicate URL formats are a minor consistency issue.
- 9.LOWcopyeditRephrase the sentence in Discussion, paragraph 4, for clarity: 'These findings, along with the additional benefits of smaller pill size, once-daily administration, lower cost, and lower risk of hypersensitivity, make TAF a valuable second-line option.'Minor punctuation and phrasing improvements enhance readability.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.