Second-Line Antiretroviral Therapy for Children Living with HIV in Africa.
Musiime V, Bwakura-Dangarembizi M, Szubert AJ, Mumbiro V, Mujuru HA, Kityo CM, Lugemwa A, Doerholt K, Chabala C, Makumbi S, Mulenga V, McIlleron H, Burger D, Natukunda E, Shakeshaft C, Linda KJ, Nathoo K, Monkiewicz L, Yawe I, Kapasa M, Nyathi M, Lungu J, Nduna B, Ndebele W, South A, Mwamabazi M, Musoro G, Griffiths A, Zyambo K, Nazzinda R, Zimba K, Zhang Y, Walker S, Turkova A, Walker AS, Bamford A, Gibb DM, CHAPAS-4 Trial Team
- DOI
- 10.1056/NEJMoa2404597
- Record issued
- 2026-08-10
- Engine
- 7.29.0
- Exported
- 2026-09-22
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/2ccea516-3f15-456c-ae4d-f3732ff0778a is authoritative.
How this rating was calculated
- ClaimsEfficacy rests on an unvalidated surrogate endpoint−0.5★
- ClaimsTreatment effect not shown to be clinically meaningful−0.5★
- ReportingData & code availability not met−0.5★
- ReportingEthical approvals partially met−0.25★
- ReportingKey resources partially met−0.25★
- The numeric-impossibility checks (GRIM/GRIMMER/DEBIT/SPRITE) did not run on this paper: the pass that reads its reported means did not complete. No reported mean was checked for arithmetic impossibility.
- No data or code availability links were detected to verify.
- 01Efficacy rests on an unvalidated surrogate endpoint
The primary efficacy endpoint is virological suppression (viral load <400 copies/mL) at week 96, which is a surrogate marker for clinical outcomes such as disease progression and mortality. The manuscript does not provide evidence linking this surrogate to hard clinical outcomes in this pediatric population, nor does it demonstrate target engagement (e.g., PK/PD) at the tested doses beyond stating that TAF concentrations were equivalent to those in adults. Thus, the efficacy claim rests on a surrogate without a validated surrogate-to-clinical-outcome link.
“Primary endpoint was week-96 viral load (VL)<400copies/mL”
- 02Treatment effect not shown to be clinically meaningful
The primary effect sizes are reported as absolute differences in the proportion of children achieving VL<400 copies/mL: TAF/FTC vs SOC +6.3% (95% CI 2.0-10.6%) and DTG vs LPV/r+ATV/r +9.7% (95% CI 4.8-14.5%). These differences, while statistically significant, are not anchored to a minimal clinically important difference or to clinical meaningfulness. The manuscript does not specify what magnitude of difference in virological suppression would be considered clinically material, and the non-inferiority margins (10% and 12%) are not explicitly justified as clinically meaningful thresholds. Therefore, the effect sizes lack an explicit anchor to clinical/biological meaningfulness.
“TAF/FTC was superior to SOC (adjusted difference [95% CI] VL<400copies/mL +6.3%[2.0%,10.6%],p=0.004) ... DTG was superior (+9.7%[+4.8%,+14.5%],p<0.001) to LPV/r and ATV/r arms combined”
- 03Data and code not shared
No data availability statement is present in the provided text; the protocol URL is given but no route to individual patient data or analytical outputs is stated.
“The trial was approved by ethics committees in Uganda, Zambia, Zimbabwe, and UK (protocol: www.mrcctu.ucl.ac.uk/studies/all-studies/c/chapas-4 (http://www.mrcctu.ucl.ac.uk/studies/all-studies/c/chapas-4) ).”
Methods ¶1Find in source
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This paper reports a well-conducted factorial RCT with strong design, statistical analysis, and reporting transparency. However, two notable gaps exist: no data availability statement and no explicit regulatory compliance framework. The key resources dimension is borderline due to deferred dosing details.
Both independent reviewer runs converged on most dimensions; the only divergence was on key resources, resolved to warn because anchor-drug dosing is deferred to the protocol. The statistics verification component checked 7 tests and found all consistent; the citation component found no retracted or unresolved references. The copyedit pass identified minor typographical and formatting issues.
Numerical inconsistencies
None foundValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
Checked — nothing surfaced.
Recomputed 7 tests: 7 consistent, 0 inconsistent; 7 via agent-written checks.
- CONSISTENTreported p = .004 · recomputed p = .004Reviewer 1Backbone primary comparison: TAF/FTC vs SOC risk difference for VL<400 at week 96
“adjusted difference [95% CI] VL<400copies/mL +6.3%[2.0%,10.6%],p=0.004”
Taken as given: The 95% CI is two-sided and the p-value corresponds to the same estimate and CI; The estimate is on the risk-difference (additive) scale; The normal approximation used to derive the p from the CI matches the reported analysisMethod: Two-tailed normal approximation p computed from estimate and 95% CI via pCI.How we recomputed it: pCI(0.063, 0.020, 0.106, 0) - CONSISTENTreported p < .001 · recomputed p = <.001Reviewers 1, 2Anchor comparison: DTG vs LPV/r+ATV/r combined risk difference
“DTG was superior to LPV/r and ATV/r arms combined (adjusted difference +9.7% [95%CI +4.8%,+14.5%]; p<0.001)”
Taken as given: The 95% CI is two-sided and the p-value corresponds to the same estimate and CI; The estimate is on the risk-difference (additive) scale; The normal approximation used to derive the p from the CI matches the reported analysisMethod: Two-tailed normal approximation p computed from estimate and 95% CI via pCI; reported as p<0.001, so the check confirms the boundary.How we recomputed it: pCI(0.097, 0.048, 0.145, 0) - CONSISTENTreported p = .040 · recomputed p = .040Reviewers 1, 2Anchor comparison: DRV/r vs LPV/r+ATV/r combined risk difference
“adjusted difference +5.6% [+0.3%,+11.0%]; p=0.04 vs. threshold p=0.03 from multiple comparisons”
Taken as given: The 95% CI is two-sided and the p-value corresponds to the same estimate and CI; The estimate is on the risk-difference (additive) scale; The normal approximation used to derive the p from the CI matches the reported analysisMethod: Two-tailed normal approximation p computed from estimate and 95% CI via pCI.How we recomputed it: pCI(0.056, 0.003, 0.110, 0) - CONSISTENTreported p = .330 · recomputed p = .327Reviewers 1, 2Anchor comparison: ATV/r vs LPV/r non-inferiority risk difference
“ATV/r was non-inferior to LPV/r (adjusted difference +3.4% [-3.4%,+10.2%]; p=0.33)”
Taken as given: The 95% CI is two-sided and the p-value corresponds to the same estimate and CI; The estimate is on the risk-difference (additive) scale; The normal approximation used to derive the p from the CI matches the reported analysisMethod: Two-tailed normal approximation p computed from estimate and 95% CI via pCI.How we recomputed it: pCI(0.034, -0.034, 0.102, 0) - CONSISTENTreported p = .002 · recomputed p = .002Reviewer 1Backbone per-protocol comparison: TAF/FTC vs SOC risk difference
“403/449(89.8%) TAF/FTC vs. 370/445(83.1%) SOC had VL <400copies/mL (adjusted difference +6.8%[+2.4%,+11.1%]; p=0.002)”
Taken as given: The 95% CI is two-sided and the p-value corresponds to the same estimate and CI; The estimate is on the risk-difference (additive) scale; The normal approximation used to derive the p from the CI matches the reported analysisMethod: Two-tailed normal approximation p computed from estimate and 95% CI via pCI.How we recomputed it: pCI(0.068, 0.024, 0.111, 0) - CONSISTENTreported p < .004 · recomputed p = .004Reviewer 2Backbone ITT: TAF/FTC vs SOC risk difference at week 96
“adjusted difference +6.3% [95% confidence interval (CI) +2.0%,+10.6%]; p=0.004”
Taken as given: The 95% CI is two-sided; The estimate is a risk difference (additive scale), so log=0; The CI is symmetric on the linear probability scale as produced by marginal estimationMethod: Two-tailed p from estimate and 95% CI on the additive scale.How we recomputed it: pCI(0.063, 0.020, 0.106, 0) - CONSISTENTreported p < .002 · recomputed p = .002Reviewer 2Backbone per-protocol: TAF/FTC vs SOC risk difference
“adjusted difference +6.8%[+2.4%,+11.1%]; p=0.002”
Taken as given: The 95% CI is two-sided; The estimate is a risk difference (additive scale)Method: Two-tailed p from estimate and 95% CI on the additive scale.How we recomputed it: pCI(0.068, 0.024, 0.111, 0)
Overstated conclusions
2 findings · worst highConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
- Efficacy rests on an unvalidated surrogate endpointAssessed
- Treatment effect not shown to be clinically meaningfulAssessed
9 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2TAF/FTC was non-inferior and superior to standard-of-care backbone for viral suppression at week 96The presented primary and per-protocol analyses directly support superiority with the effect estimate, CI and p-value.Evidence: 406/454 (89.4%) vs 378/454 (83.3%), adjusted difference +6.3% [2.0%,10.6%], p=0.004; per-protocol +6.8% [2.4%,11.1%], p=0.002.
“At week-96, 406/454(89.4%) TAF/FTC vs. 378/454(83.3%) SOC had VL <400copies/mL (adjusted difference +6.3% [95% confidence interval (CI) +2.0%,+10.6%]; p=0.004)”
ResultsFind in source - supportedReviewers 1, 2DTG was superior to LPV/r and ATV/r arms combinedThe headline anchor comparison is directly supported by the reported estimate, CI and p-value.Evidence: Adjusted difference +9.7% [+4.8%,+14.5%]; p<0.001.
“DTG was superior to LPV/r and ATV/r arms combined (adjusted difference +9.7% [95%CI +4.8%,+14.5%]; p<0.001)”
AbstractFind in source - supportedReviewer 1DRV/r achieved higher virological suppression but could not be declared superiorThe claim accurately reflects the result, which narrowly missed the prespecified multiple-comparison threshold.Evidence: Adjusted difference +5.6% [+0.3%,+11.0%]; p=0.04 vs threshold p=0.03.
“DRV/r was not superior to LPV/r and ATV/r combined as the comparison did not meet pre-specified significance (adjusted difference +5.6% [+0.3%,+11.0%]; p=0.04 vs. threshold p=0.03 from multiple comparisons).”
AbstractFind in source - supportedReviewer 1ATV/r was non-inferior to LPV/rThe estimate and CI lie within the prespecified 12% non-inferiority margin.Evidence: Adjusted difference +3.4% [-3.4%,+10.2%]; p=0.33.
“ATV/r was non-inferior to LPV/r (adjusted difference +3.4% [-3.4%,+10.2%]; p=0.33).”
AbstractFind in source - supportedReviewers 1, 2TAF/FTC and DTG are effective for children without evident safety concernsEfficacy is supported by the primary comparisons, and the safety data (one death, 3.2% serious AEs, no between-arm differences) support the absence of evident safety concerns.Evidence: Superior suppression for TAF/FTC and DTG; 29 (3.2%) serious AEs with p>0.1 across arms; bone health similar.
“Second-line ART including TAF/FTC and DTG are effective for children without evident safety concerns.”
ConclusionFind in source - supportedReviewer 1There was no evidence of bone toxicity with TAF and no excess weight gain with any backbone/anchor combinationThe bone and weight outcomes are reported as secondary analyses supporting the absence of these harms.Evidence: Bone health similar between backbone arms; no excess weight gain with any combination including DTG+TAF/FTC.
“Bone health was similar between backbone arms, irrespective of anchor drug.”
AbstractFind in source - supportedReviewer 2DRV/r could not be declared superior to LPV/r and ATV/r combined.The paper reports the comparison did not reach the pre-specified multiple-comparison threshold, consistent with the claim.Evidence: Adjusted difference +5.6% [0.3%,11.0%], p=0.04 vs. threshold p=0.03.
“DRV/r was not superior to LPV/r and ATV/r combined as the comparison did not meet pre-specified significance (adjusted difference +5.6% [+0.3%,+11.0%]; p=0.04 vs. threshold p=0.03 from multiple comparisons)”
ResultsFind in source - supportedReviewer 2ATV/r was non-inferior to LPV/r.The CI lies within the pre-specified 12% margin, supporting non-inferiority.Evidence: Adjusted difference +3.4% [-3.4%,10.2%], p=0.33.
“ATV/r was non-inferior to LPV/r (adjusted difference +3.4% [-3.4%,+10.2%]; p=0.33)”
ResultsFind in source - supportedReviewer 2No excess weight gain was observed with any backbone/anchor combination, including DTG+TAF/FTC.The paper reports no evidence of excess weight gain and contrasts this with adult data; the claim is consistent with the presented results.Evidence: Weight/BMI-for-age Z-score analyses and explicit statement of no interaction by backbone.
“we observed no excessive weight-gain with any anchor/backbone combination, including DTG+TAF/FTC”
DiscussionFind in source
Premise concern: surrogate not validated for clinical benefit; effect size not shown to be clinically meaningful.
- INADEQUATESurrogate endpointThe primary efficacy endpoint is virological suppression (viral load <400 copies/mL) at week 96, which is a surrogate marker for clinical outcomes such as disease progression and mortality. The manuscript does not provide evidence linking this surrogate to hard clinical outcomes in this pediatric population, nor does it demonstrate target engagement (e.g., PK/PD) at the tested doses beyond stating that TAF concentrations were equivalent to those in adults. Thus, the efficacy claim rests on a surrogate without a validated surrogate-to-clinical-outcome link.
“Primary endpoint was week-96 viral load (VL)<400copies/mL”
- INADEQUATEEffect sizeThe primary effect sizes are reported as absolute differences in the proportion of children achieving VL<400 copies/mL: TAF/FTC vs SOC +6.3% (95% CI 2.0-10.6%) and DTG vs LPV/r+ATV/r +9.7% (95% CI 4.8-14.5%). These differences, while statistically significant, are not anchored to a minimal clinically important difference or to clinical meaningfulness. The manuscript does not specify what magnitude of difference in virological suppression would be considered clinically material, and the non-inferiority margins (10% and 12%) are not explicitly justified as clinically meaningful thresholds. Therefore, the effect sizes lack an explicit anchor to clinical/biological meaningfulness.
“TAF/FTC was superior to SOC (adjusted difference [95% CI] VL<400copies/mL +6.3%[2.0%,10.6%],p=0.004) ... DTG was superior (+9.7%[+4.8%,+14.5%],p<0.001) to LPV/r and ATV/r arms combined”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
3 findings · worst highRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data and code not sharedAssessed
- Ethics/consent reporting incompleteAssessed
- Key resources under-identified (antibodies, cell lines, RRIDs)Assessed
Prior work is cited extensively (ODYSSEY, NADIA, DAWNING, VISEND, adult TAF trials) with strengths and weaknesses acknowledged (e.g., LPV/r unpalatability, DRV/r cost, TAF bone/renal safety rationale). The premise that backbone and anchor drug choice is unresolved is explicitly stated and links directly to the trial objectives. Limitations of the evidence base and of the trial itself are discussed in the Discussion.
“There are minimal data on TAF in African children; the first paediatric pharmacokinetic data showed tenofovir concentrations equivalent to those safe and effective in adults.”
“INSTI-based regimens including tenofovir demonstrate robust efficacy when compared to ritonavir-boosted PI-based regimens including zidovudine in adult second-line trials.”
“Lopinavir (LPV) is the only paediatric ritonavir co-formulated boosted PI but requires twice-daily dosing and is unpalatable; ritonavir-boosted darunavir (DRV/r) and atazanavir (ATV/r) are dosed once-daily but paediatric FDCs are unavailable and DRV/r is relatively costly.”
“There are minimal data on TAF in African children; the first paediatric pharmacokinetic data showed tenofovir concentrations equivalent to those safe and effective in adults.”
“hypothesising TAF would be non-inferior to SOC (10% margin), DTG and DRV/r superior to LPV/r and ATV/r combined, and ATV/r non-inferior to LPV/r (12% margin)”
Randomization used a computer-generated sequence with variably sized permuted blocks, stratified by centre and first-line NRTI, with allocation concealed until eligibility confirmed. Open-label design is stated and justified (objective primary endpoint). Power analysis is fully described (≥95% power backbone, 88%/89% anchor, with margins and alpha). Inclusion/exclusion criteria are pre-specified. Analysis population (ITT, with per-protocol for non-inferiority) addresses missing-data/outlier handling. Replicate/controls/replication are n/a for a single pivotal human trial.
“A computer-generated sequential randomisation list with variably sized permuted blocks was prepared by the trial statistician and incorporated securely into an online database. Allocation was concealed until eligibility was confirmed by local centre staff, who then randomised.”
“assuming 80.0%-87.5% SOC achieved VL<400 copies/ml at week-96, 920 children provided ≥95% power to demonstrate TAF was non-inferior (10% margin) (two-sided alpha=0.05)”
“The open-label design of the trial could have potentially introduced bias; however the primary endpoint (VL) was objective.”
“A computer-generated sequential randomisation list with variably sized permuted blocks was prepared by the trial statistician and incorporated securely into an online database. Allocation was concealed until eligibility was confirmed by local centre staff, who then randomised.”
“The open-label design of the trial could have potentially introduced bias; however the primary endpoint (VL) was objective.”
“assuming 80.0%-87.5% SOC achieved VL<400 copies/ml at week-96, 920 children provided ≥95% power to demonstrate TAF was non-inferior (10% margin) (two-sided alpha=0.05)”
Sex (497 male, 54.1%), age (median 10, IQR 8-13), weight, height, CD4, VL, WHO stage and time on first-line ART are all reported in Table 1. Both sexes enrolled, so sex_justified is n/a. Demographics include age and sex; race/ethnicity is implicit (three African countries) and comorbidity data are limited but WHO staging is given.
“497(54.1%) children were male; median age 10 years (IQR 8,13)”
“497(54.1%) children were male; median age 10 years (IQR 8,13); 777(84.5%) were WHO stage 1/2.”
“Values are n (%) or median (IQR)”
IRB approval is reported with named committees (JREC Uganda, UNZABREC Zambia, Zimbabwe committee, and UK). Informed consent and assent are described. However, regulatory_compliance is only 'according to national guidelines' — no named recognized framework such as the Declaration of Helsinki is cited, making this criterion reported_but_inadequate and triggering a warn overall.
“Guardians provided written informed consent, with additional assent from older children, according to national guidelines.”
“according to national guidelines”
“The trial was approved by ethics committees in Uganda (Joint Research Ethics Committee (JREC)), Zambia (University of Zambia Biomedical Research Ethics Committee (UNZABREC))”
“Guardians provided written informed consent, with additional assent from older children, according to national guidelines.”
For a drug trial, bench criteria (antibodies, cell lines, mycoplasma, organisms) are n/a. reagents_identified (scored against investigational products): the drugs are named and donors listed (ViiV, Gilead, J&J, CIPLA) and the TAF/FTC FDC is given as 15mg/120mg, but anchor-drug doses/regimens are not specified in text ('Full study details can be found in the protocol'), so it is reported_but_inadequate. software_tools_identified is adequate (Stata version 17.0). 1 of 2 applicable adequate = 50% → warn.
“a new paediatric TAF/emtricitabine(FTC) fixed-dose-combination (FDC) (15mg/120mg)”
“pharmaceutical companies donating additional funding (Gilead Sciences, Johnson and Johnson) and drugs (ViiV Healthcare, Gilead Sciences, Johnson and Johnson, CIPLA)”
“Analyses were intention-to-treat using Stata (version 17.0).”
“Analyses were intention-to-treat using Stata (version 17.0).”
“a new paediatric TAF/emtricitabine(FTC) fixed-dose-combination (FDC) (15mg/120mg) has been developed”
Tests are named (logistic regression with marginal risk differences, Cox regression, GEE for continuous outcomes). Assumptions are handled by design for a large trial. Exact p-values are given (e.g., p=0.004, p<0.001, p=0.04, p=0.33). Effect sizes reported with 95% CIs throughout. Software (Stata 17.0) identified. Data presentation follows the clinical idiom (per-group n, CIs, CONSORT flow). mathematical_plausibility is n/a for large-N continuous outcomes; the arithmetic I checked (baseline counts and percentages in Tables 1-2) is internally consistent.
“adjusted difference [95% CI] VL<400copies/mL +6.3%[2.0%,10.6%],p=0.004”
“89% power to detect 10% higher suppression in each of DTG and DRV/r than LPV/r and ATV/r combined (two-sided alpha=0.03; as multiple comparisons)”
“Primary endpoint analyses used logistic regression (adjusting for stratification factors), then marginal estimation of risk differences.”
“DTG was superior to LPV/r and ATV/r arms combined (adjusted difference +9.7% [95%CI +4.8%,+14.5%]; p<0.001).”
The paper references the protocol at www.mrcctu.ucl.ac.uk/studies/all-studies/c/chapas-4 and 'the protocol at nejm.org', but there is no explicit data availability statement describing where or how the trial data can be accessed. Repository deposit and accession numbers are not applicable to identifiable paediatric patient data, and no bespoke code repository is mentioned. Because the one applicable criterion (data availability statement) is not reported, the dimension fails. Confidence is moderate because the provided text may be truncated.
“The trial was approved by ethics committees in Uganda, Zambia, Zimbabwe, and UK (protocol: www.mrcctu.ucl.ac.uk/studies/all-studies/c/chapas-4 (http://www.mrcctu.ucl.ac.uk/studies/all-studies/c/chapas-4) ).”
Trial registration number is given. Methods are detailed enough to replicate, with full details deferred to the protocol. Limitations are thoroughly discussed (open-label, generalisability, high CD4 at enrolment, pill burden). Conclusions are appropriately hedged and disease-relevant. Funding and conflict-of-interest disclosures are present. A CONSORT flow diagram is labelled, but no explicit CONSORT checklist submission statement is made, so reporting_guideline is reported_but_inadequate. 6 of 7 applicable criteria adequate = 86% → pass.
“CHAPAS-4(ISRCTN22964075)”
“DRV/r was not superior to LPV/r and ATV/r combined as the comparison did not meet pre-specified significance (adjusted difference +5.6% [+0.3%,+11.0%]; p=0.04 vs. threshold p=0.03 from multiple comparisons).”
“A limitation is that CHAPAS-4 does not provide direct evidence to inform anchor/backbone choice in this situation; however, safety and efficacy could be inferred (given lack of evidence of interaction)”
“CHAPAS-4(ISRCTN22964075) was a randomised, open-label trial with a 2x4 factorial design.”
“A limitation is that CHAPAS-4 does not provide direct evidence to inform anchor/backbone choice in this situation; however, safety and efficacy could be inferred”
“The main funding for this study is provided by the European and Developing Countries Clinical Trials Partnership.”
Registered (1 ID: ISRCTN). Reporting guideline cited: CONSORT.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 22 references by DOI: 19 verified — 3 no DOI (shown, not verified).
- NO DOIDolutegravir with recycled NRTIs is noninferior to PI-based ART: VISEND trialNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIConsolidated guidelines on HIV prevention, testing, treatment, service delivery and monitoring: recommendations for a public health approachNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIPriorities for antiretroviral drug optimization in adults and children: report of a CADO, PADO and HIVResNet joint meeting, 27 September–15 October 2021No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
7 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 7 minor suggestions below.
7 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORconsistencyHealth economic analysis, Results“saving $190.77 compared to ATZ/r”→ Change 'ATZ/r' to 'ATV/r' (atazanavir/ritonavir) for consistency throughout the paper.Drug abbreviation typo; 'ATZ/r' is used nowhere else.
- MINORconsistencyDiscussion, paragraph 7“superiority of DTG vs. LVP/r in the DAWNING trial”→ Change 'LVP/r' to 'LPV/r' (lopinavir/ritonavir).Drug abbreviation typo; 'LPV/r' is used elsewhere.
- MINORclarityAbstract, Results“baseline VL 17,573copies/ml[5,549-55,700]”→ Add spaces: '17,573 copies/ml [5,549-55,700]'.Missing spaces around the bracket and unit.
- MINORclarityResults, paragraph 1“CD4 count 669cells/mm 3 [413-971]”→ Add spaces and superscript: 'CD4 count 669 cells/mm3 [413-971]'.Formatting of unit and bracket spacing.
- MINORconsistencyResults, health economic analysis“saving $190.77 compared to ATZ/r”→ Change 'ATZ/r' to 'ATV/r' for consistency with the rest of the manuscript.Non-standard abbreviation used only once.
- MINORtypoDiscussion, DAWNING reference“non-inferiority of DTG vs. LVP/r in the NADIA trial”→ Change 'LVP/r' to 'LPV/r'.Transposed letters in lopinavir abbreviation.
- MINORclarityEthical approval section (end of text)“Zimbabwe (Joint Research Ethics Committee University of”→ Complete the truncated sentence; the ethics approval text appears cut off.Manuscript/OCR truncation; should be completed with the full committee name and approval reference.
For post-publication audit, the paper is methodologically sound. The most consequential gaps are the missing data availability statement (a common but important omission for a data-driven trial) and the lack of a named regulatory compliance framework. An informed reader should weigh these as reporting completeness issues, not validity threats. A correction or erratum could address the minor typos.
- 1.HIGHdata codeAdd a data availability statement describing the access route for individual participant data (e.g., managed access via the MRC CTU data-sharing committee, with conditions and timeframe).The paper currently lacks any data availability statement, which is a standard requirement for a data-driven trial and is flagged as a fail in the audit.
- 2.HIGHreportingComplete the truncated ethics committee name in the Ethical approval section: 'Zimbabwe (Joint Research Ethics Committee University of ...)' – add the full committee name and approval reference.The sentence is cut off, leaving the ethics approval information incomplete and potentially confusing.
- 3.HIGHethicsAdd an explicit statement of compliance with a recognized regulatory framework (e.g., 'conducted in accordance with the Declaration of Helsinki and ICH-GCP') in the Ethical approval section.The current text only references 'national guidelines'; naming a recognized framework is standard for clinical trial reporting and is flagged as inadequate.
- 4.HIGHrigorSpecify the dose/regimen and formulation for each anchor drug (DTG, DRV/r, ATV/r, LPV/r) and the SOC backbone in the Methods or a supplementary table, rather than deferring entirely to the protocol.The current description is incomplete, making the key resources dimension borderline. Providing the dosing details supports replication and clarity.
- 5.HIGHreportingAdd a sentence stating that the manuscript adheres to the CONSORT 2010 statement and note that the checklist is available as supplementary material.A CONSORT flow diagram is present, but the paper does not explicitly reference the checklist, which is a reporting guideline requirement.
- 6.MEDIUMcopyeditCorrect the drug abbreviation 'ATZ/r' to 'ATV/r' in the health economic analysis results.The non-standard abbreviation 'ATZ/r' appears only once; it should be consistent with the rest of the paper ('ATV/r').
- 7.MEDIUMcopyeditCorrect the drug abbreviation 'LVP/r' to 'LPV/r' in the Discussion paragraph referencing the DAWNING trial.The typo 'LVP/r' should be 'LPV/r' (lopinavir/ritonavir) for consistency.
- 8.LOWcopyeditAdd spaces around units and brackets in the Abstract and Results (e.g., '17,573 copies/ml [5,549-55,700]' and '669 cells/mm3 [413-971]').Minor formatting issues improve readability and consistency.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.