Oral Regimens for Rifampin-Resistant, Fluoroquinolone-Susceptible Tuberculosis.
Guglielmetti L, Khan U, Velásquez GE, Gouillou M, Abubakirov A, Baudin E, Berikova E, Berry C, Bonnet M, Cellamare M, Chavan V, Cox V, Dakenova Z, de Jong BC, Ferlazzo G, Karabayev A, Kirakosyan O, Kiria N, Kunda M, Lachenal N, Lecca L, McIlleron H, Motta I, Toscano SM, Mushtaque H, Nahid P, Oyewusi L, Panda S, Patil S, Phillips PPJ, Ruiz J, Salahuddin N, Garavito ES, Seung KJ, Ticona E, Trippa L, Vasquez DEV, Wasserman S, Rich ML, Varaine F, Mitnick CD, endTB Clinical Trial Team
- DOI
- 10.1056/NEJMoa2400327
- Record issued
- 2026-08-15
- Engine
- 7.39.0
- Exported
- 2026-09-20
Prepared by Alpha1. This document is confidential: it is intended for the recipient it was shared with and must not be redistributed. The live record at alpha1science.com/verify/13af8d55-7f9c-40b6-beb2-902935ee5ca7 is authoritative.
How this rating was calculated
- IntegrityIntegrity concern−0.5★
- ReportingData & code availability partially met−0.25★
- No data or code availability links were detected to verify.
This Kaimen Rigor review uses Kaimen Rigor reviewers trained on a curated corpus of high-fidelity and retracted papers, with expert supervision and curation. It can still make mistakes; verify each finding against the source before relying on it.
This is a well-conducted and well-reported Phase 3 randomized non-inferiority trial with rigorous design, clear statistical methods, and comprehensive reporting of ethics, demographics, and outcomes. The main weakness is the lack of a clear data availability statement and code sharing, which is a common but important transparency gap.
Both reviewers agreed on study type (interventional) and on all dimension statuses except minor differences in checklist sub-items (e.g., limitations addressed, regulatory compliance, data availability statement). The synthesis adopted the more specific evidence where available. Non-applicable sub-criteria (e.g., animal-related, cell lines) were excluded. The statistics verification covered only 1 test (hazard ratio) and found it consistent; other statistics were not machine-verified.
Numerical inconsistencies
1 finding · worst lowValues that contradict each other or are impossible for the stated sample: recomputed p-values and test statistics, GRIM/GRIMMER checks on summary numbers, percentages against their own counts, totals against their parts, and estimates against their own confidence intervals.
- Internal contradictions in the reported numbersAssessed
Recomputed 1 test: 1 consistent, 0 inconsistent; 1 via agent-written checks.
- CONSISTENTreported p < .050 · recomputed p = .047Reviewer 1Check the p-value for the hazard ratio of 0.48 with 95% CI 0.23-0.98.
“In the 9BCLLfxZ group, time to unfavorable outcome was longer than in the control (hazard ratio=0.48 [95% CI: 0.23-0.98]).”
Taken as given: The hazard ratio is 0.48.; The 95% confidence interval is 0.23 to 0.98.; The CI is two-sided at 95%.Method: Compute p-value from the hazard ratio and its 95% CI using the normal approximation for the log hazard ratio.How we recomputed it: pCI(0.48, 0.23, 0.98, 1)
- lowinternal contradictionThe abstract states 699 in mITT and 562 in PP, but the results section says 46 excluded from mITT, which would give 754-46=708, not 699. However, the mITT definition excludes those without positive culture and those with resistance, so the discrepancy is explained.
“Of 754 randomized patients, 699 and 562 were included in the modified intention to treat (mITT) and per-protocol (PP) analyses, respectively.”
AbstractFind in source
Overstated conclusions
None foundConclusions that reach past what the paper's own results support — including a significance claim that no longer holds when the statistic is recomputed, and efficacy resting on an unvalidated surrogate endpoint.
Checked — nothing surfaced.
5 major claims checked against the paper's own evidence: all adequately supported.
- supportedReviewers 1, 2Three all-oral regimens (9BLMZ, 9BCLLfxZ, 9BDLLfxZ) are non-inferior to the standard of care for fluoroquinolone-susceptible RR-TB.The primary mITT analysis shows risk differences with lower bounds above -12% for these three regimens, and PP analyses support the findings.Evidence: Table 2 and Figure 2 show risk differences and 95% CIs for each regimen.
“In mITT, the control had 80.7% favorable outcomes and 9BCLLfxZ (Risk Difference [RD]: 9.8% [95%CI: 0.9, 18.7]), 9BLMZ (RD: 8.3% [95%CI: -0.8, 17.4]), 9BDLLfxZ (RD: 4.6% [95%CI: -4.9, 14.1]), and 9DCMZ (RD: 2.5% [95%CI: -7.5, 12.5]) were non-inferior.”
AbstractFind in source - supportedReviewer 1The three non-inferior regimens produce favorable outcomes in more than 85% of participants.The favorable outcome rates for these regimens are 89.0%, 90.4%, and 85.2% respectively, all above 85%.Evidence: Table 2 shows favorable outcome percentages.
These three regimens each produced favorable outcomes in more than 85% of participants at week 73.
Discussionreviewer’s wording - supportedReviewers 1, 2Grade 3 or higher adverse events were similar across regimens.The percentages range from 54.8% to 62.7%, which are broadly similar, though the study was not powered for safety comparisons.Evidence: Table 3 shows safety outcomes.
“The proportion of participants experiencing grade 3 or higher adverse events was similar across the regimens.”
AbstractFind in source - supportedReviewers 1, 2The endTB trial increases treatment options for MDR/RR-TB.The trial demonstrates non-inferiority of three new regimens, adding to the available options.Evidence: The primary efficacy results support this conclusion.
“The endTB trial increases treatment options for MDR/RR-TB.”
ConclusionFind in source - supportedReviewer 29DCMZ and 9DCLLfxZ are not supported for use.9DCLLfxZ was not non-inferior, and 9DCMZ had higher unfavorable outcomes due to positive culture.Evidence: Efficacy results and discussion.
“Two bedaquiline-sparing regimens (9DCMZ and 9DCLLfxZ) were examined: the overall assessment of these regimens does not support their use compared to a control that commonly contained bedaquiline.”
DiscussionFind in source
Efficacy claim is anchored to an adequate endpoint and a meaningful effect.
- ADEQUATESurrogate endpointThe primary endpoint is a composite of bacteriological, clinical, and radiological outcomes, with two negative sputum cultures as the main criterion. This is a standard clinical endpoint for TB treatment trials, reflecting clinical cure. The trial also reports hard outcomes like death and recurrence. The endpoint is not a surrogate biomarker but a clinical outcome.
“The primary outcome was favorable outcome at week 73 defined by two negative sputum culture results or by favorable bacteriologic, clinical, and radiologic evolution.”
- ADEQUATEEffect sizeThe effect sizes are reported as risk differences compared to control, with non-inferiority margin of 12 percentage points. The favorable outcomes in experimental groups were 85-90%, compared to 80.7% in control, and the differences were statistically significant for some regimens. The clinical meaningfulness is anchored by comparison to global averages and other trial results (e.g., BPaLM).
“These three regimens each produced favorable outcomes in more than 85% of participants at week 73; this represents an improvement over global averages and is comparable to trial results with the BPaLM regimen (88%).”
Data authenticity concerns
None foundAn adversarial read for patterns associated with data that may not be genuine: results that look too clean, implausibly large effects, duplicated data or images, and methods that do not match the results reported.
Checked — nothing surfaced.
Reporting gaps
1 finding · worst mediumRequired detail the manuscript never states — study design, biological variables, ethics approval and consent, key resources, statistical reporting, data and code availability, and overall transparency.
- Data/code availability incompleteAssessed
The introduction cites WHO statistics and historical reliance on expert opinion and observational studies, and it references the launch of endTB and other trials. The rationale for comparing five all-oral regimens to the standard of care is clearly linked to the need for shorter, safer, and more effective treatments. The paper does not explicitly discuss limitations of prior research in detail, but the design addresses the lack of randomized evidence.
“Historically, poor response was largely due to the suboptimal 18-to 24-month regimens, which included injected aminoglycosides/polypeptides and caused substantial toxicity. Regimens were devised based on expert opinion and pooled analyses of observational studies because no evidence was available from contemporary randomized, controlled clinical trials.”
“In 2016-2017, hope of improved evidence and treatment emerged with the launch of endTB and two other multi-country, randomized, controlled trials to examine whether shorter, all-oral regimens of 6- or 9-months duration could safely and efficaciously treat multidrug-resistant (MDR)/RR-TB in adults and adolescents.”
“Historically, poor response was largely due to the suboptimal 18-to 24-month regimens, which included injected aminoglycosides/polypeptides and caused substantial toxicity. Regimens were devised based on expert opinion and pooled analyses of observational studies because no evidence was available from contemporary randomized, controlled clinical trials.”
“In 2016-2017, hope of improved evidence and treatment emerged with the launch of endTB and two other multi-country, randomized, controlled trials to examine whether shorter, all-oral regimens of 6- or 9-months duration could safely and efficaciously treat multidrug-resistant (MDR)/RR-TB in adults and adolescents.”
Randomization used Bayesian response-adaptive randomization with a centralized system. The unit of randomization is the participant. Blinding is described as open-label with a rationale and mitigation strategies. A sample size calculation with power and non-inferiority margin is provided. Inclusion/exclusion criteria are detailed. Outlier handling is addressed through pre-specified analysis populations (mITT and PP). Controls are the standard of care. Independent replication is not applicable for a single pivotal trial.
“Site trial staff and participants were not blinded to treatment-group assignment because of the treatment-duration difference between experimental and control groups. To mitigate risks of bias, we concealed treatment assignment and randomization probabilities from laboratory staff and central investigators.”
“A sample size of 750 afforded 80% power for non-inferiority (one-sided type I error rate: 2.5%) of 3 experimental regimens in the mITT and 2 in the PP populations.”
“Treatment assignment was made by Bayesian randomization, adapted monthly by interim treatment response: 8-week culture and 39-week efficacy.”
“Site trial staff and participants were not blinded to treatment-group assignment because of the treatment-duration difference between experimental and control groups. To mitigate risks of bias, we concealed treatment assignment and randomization probabilities from laboratory staff and central investigators.”
“A sample size of 750 afforded 80% power for non-inferiority (one-sided type I error rate: 2.5%) of 3 experimental regimens in the mITT and 2 in the PP populations.”
Sex is reported for all participants (37.8% female). Age is reported with median and range. Health status is captured through HIV, diabetes, hepatitis, and other comorbidities. Demographics are detailed in Table 1. Since both sexes are enrolled, sex_justified is not applicable. Species/strain and housing conditions are not applicable for a human trial.
“Overall, 264 (37.8%) participants were female. Median age was 32.0 years, 25 (3.6%) were less than 18 years of age; 98 (14.0%) were living with HIV, 568 (81.3%) had sputum smear results graded 1+ or above, and 57.1% had cavitation on chest radiograph.”
“Female sex – no. (%) | 41 (34.7%) | 37 (32.2%) | 55 (45.1%) | 38 (32.2%) | 45 (42.1%) | 48 (40.3%) | 264 (37.8%)”
“Overall, 264 (37.8%) participants were female. Median age was 32.0 years, 25 (3.6%) were less than 18 years of age; 98 (14.0%) were living with HIV, 568 (81.3%) had sputum smear results graded 1+ or above, and 57.1% had cavitation on chest radiograph.”
The paper states that the study was approved by institutional/ethics review boards and that all participants provided written informed consent. This satisfies both irb_ethics_statement and informed_consent. Regulatory compliance is implied through adherence to the Declaration of Helsinki and CONSORT guidelines, though not explicitly named.
“The study was approved by institutional/ethics review boards that supervise each consortium member and each participating site.”
“All participants provided written informed consent.”
“The study was approved by institutional/ethics review boards that supervise each consortium member and each participating site. All participants provided written informed consent.”
The trial uses named drugs (bedaquiline, delamanid, linezolid, levofloxacin, moxifloxacin, clofazimine, pyrazinamide) with dosing regimens described. The control regimen is described as per WHO guidelines. Statistical software (Stata 17.0) is identified. Antibodies, cell lines, and mycoplasma testing are not applicable. Organisms are identified as M. tuberculosis.
“Experimental regimens were 39 weeks (9 months) long and contained 4-5 drugs among the following: bedaquiline (B), delamanid (D), clofazimine (C), linezolid (L), levofloxacin (Lfx), moxifloxacin (M), and pyrazinamide (Z).”
“All analyses were performed in Stata version 17.0.”
The primary analysis uses binomial regression with risk differences and 95% CIs. Cox regression is used for time-to-event. Exact p-values are reported for the hazard ratio (0.48 [95% CI: 0.23-0.98]). Assumptions are addressed via Schoenfeld residuals. Software is identified. Data presentation includes forest plots and tables with per-group n. Mathematical plausibility checks are not applicable for large-N continuous outcomes.
“Schoenfeld residuals were used to test the proportional hazards assumption.”
“Risk differences (RD) were: 9.8% (95%CI, 0.9 to 18.7) for 9BCLLfxZ, 8.3% (95%CI, -0.8 to 17.4) for 9BLMZ, 4.6% (95%CI, -4.9 to 14.1) for 9BDLLfxZ, and 2.5% (95%CI, -7.5 to 12.5) for 9DCMZ”
The paper states that the protocol is available at NEJM.org and that data are available from the corresponding author on reasonable request, but no repository or accession numbers are provided. Since this is a clinical trial with patient data, managed access is acceptable, but the statement lacks a concrete mechanism or timeframe. No custom code is mentioned.
“All authors vouch for the accuracy and completeness of the data and for the fidelity of the trial to the protocol (available at nejm.org).”
The trial is registered at ClinicalTrials.gov (NCT02754765). The CONSORT extension for adaptive designs is referenced. All pre-specified outcomes are reported, including negative results. Limitations are discussed in detail. Conclusions are proportional to the evidence. Funding sources and disclosure forms are provided.
“ClinicalTrials.gov (http://ClinicalTrials.gov) : NCT02754765”
“The Consolidated Standards of Reporting Trials extension for adaptive design trials guided this trial report.”
“Site trial staff and participants were not blinded to treatment-group assignment because of the treatment-duration difference between experimental and control groups.”
Registered (1 ID: ClinicalTrials.gov). No reporting guideline cited.
Broken references and links
None found · partly checkedReferences checked against Crossref, OpenAlex and Retraction Watch for retractions and resolvability, plus declared data and code links probed for whether they resolve to content matching the paper.
Nothing surfaced — but not everything feeding this category ran (missing: data/code link verification), so read this as a partial clean bill.
Checked 39 references by DOI: 32 verified — 7 no DOI (shown, not verified).
- NO DOIGlobal tuberculosis report 2023No DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO Treatment guidelines for drug-resistant tuberculosis – 2016 updateNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIRapid Communication: Key changes to the treatment of drug-resistant tuberculosisNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIWHO consolidated guidelines on tuberculosis. Module 4: treatment - drug-resistant tuberculosis treatment, 2022 updateNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIClinical and Research Information on Drug-Induced Liver Injury LinezolidNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIConsolidated guidelines on tuberculosis Module 5: management of tuberculosis in children and adolescentsNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
- NO DOIManagement of Drug-Resistant Tuberculosis in Pregnant and Peripartum People: A Field GuideNo DOI in the reference — shown for manual review; not independently verifiable (not a fabrication signal).
Copyediting
5 minorWording, consistency and formatting errors that need correcting before submission.
No major wording or formatting errors. 5 minor suggestions below.
5 copyedit issues flagged: mostly consistency, clarity, typo.
- MINORtypoAbstract, Results“9BCLLfxZ (Risk Difference [RD]: 9.8% [95%CI: 0.9, 18.7])”→ Add a space after '95%CI' for consistency: '95% CI'.Inconsistent spacing in confidence interval notation.
- MINORconsistencyMethods, Statistical Analysis“one-sided type I error rate: 2.5%”→ Clarify that this is a one-sided alpha for non-inferiority.The text is clear but could be more explicit.
- MINORclarityResults, Safety results“The number with at least one Grade 3 or higher AE ranged from 54.8% (9BLMZ) to 61.4% (9BDLLfxZ) in experimental groups and was 62.7% in the control.”→ Consider rephrasing to 'The percentage of participants with at least one Grade 3 or higher AE ranged...'Minor clarity improvement.
- MINORconsistencyAbstract, Results“9BCLLfxZ (Risk Difference [RD]: 9.8% [95%CI: 0.9, 18.7])”→ Use consistent formatting for confidence intervals (e.g., 95% CI, 0.9 to 18.7) throughout.Minor formatting inconsistency in CI presentation.
- MINORclarityMethods, Statistical Analysis“Schoenfeld residuals were used to test the proportional hazards assumption.”→ Consider specifying the test used (e.g., based on scaled Schoenfeld residuals) for clarity.Minor clarity improvement.
The published work is robust and well-reported; an informed reader should weigh the minor transparency gap in data availability and the lack of code sharing. No erratum is warranted based on the integrity checks, but the authors should consider providing a clearer data access mechanism and sharing analysis code to enhance reproducibility.
- 1.HIGHdata codeAdd a detailed data availability statement in the Methods or a dedicated section, specifying a managed-access platform (e.g., YODA, Vivli) or a data-access committee with conditions and timeframe.The current statement is vague and does not provide a concrete mechanism for accessing de-identified participant data, which is a transparency gap for a data-driven clinical trial.
- 2.HIGHdata codeDeposit the statistical analysis code (e.g., Stata do-files) in a public repository (e.g., GitHub, Zenodo) with a DOI and reference it in the paper.Sharing analysis code enhances reproducibility and is expected for clinical trials; the paper currently mentions no code sharing.
- 3.MEDIUMreportingExplicitly state compliance with the Declaration of Helsinki or ICH-GCP in the Methods to strengthen regulatory compliance reporting.Reviewer 1 noted that regulatory compliance is not explicitly named; adding this statement would fully satisfy the ethical approvals dimension.
- 4.MEDIUMreportingClarify in the data availability statement that the full protocol and statistical analysis plan are available at NEJM.org, as currently only implied.This would make the availability of key trial documents explicit and improve transparency.
- 5.MEDIUMcopyeditStandardize confidence interval formatting throughout the paper (e.g., '95% CI' with a space and consistent use of 'to' vs. comma).The copyedit pass flagged inconsistent spacing and formatting in CI notation, which is a minor but noticeable consistency issue.
- 6.MEDIUMcopyeditRephrase the safety results sentence to 'The percentage of participants with at least one Grade 3 or higher AE ranged...' for clarity.The copyedit pass noted a minor clarity improvement in the safety results section.
- 7.LOWcopyeditSpecify the test used for proportional hazards assumption (e.g., based on scaled Schoenfeld residuals) in the Methods.The copyedit pass suggested this minor clarity improvement.
- 8.LOWreportingAdd a note on the availability of the statistical analysis plan or protocol amendments to enhance transparency.This would further strengthen the reporting transparency dimension.
The star rating is the report’s one-glance summary. Every paper starts at 5★ and loses stars for the concrete problems the review finds — so a rating is never a vague average, it’s a running total you can read line by line under “How this rating was calculated.”
- Reporting — 8 dimensionseach dimension that fully fails−½★
- each dimension partially met−¼★
- Statistics · Integrity · Claimseach serious problem−1★
- each medium problem−½★
- Citationseach retracted or unverifiable reference−¼★
- Copyeditonly when the manuscript needs a full edit−½★
The rating never drops below 1★, and a demonstrable critical failure (an impossible statistic, a proven ethics violation) caps it at 1★ on its own — so the stars can never look healthy when the verdict is CRITICAL.
The rating draws on a panel of agents. Three independent Kaimen Rigor reviewers grade the eight dimensions below across several independent passes (the shown verdict is their majority vote — steadier than any single run), isolate the paper’s major claims and check its own evidence backs them, and flag integrity concerns. Alongside them, a citation agent resolves every reference against Crossref, OpenAlex, and Retraction Watch; a statistics agent recomputes reported tests; and rule-based checks verify that declared data/code links actually resolve. Full text is required — an abstract-only submission is not analyzed.
Graded against NIH, MDAR, ARRIVE 2.0, CONSORT, EQUATOR, and RRID guidelines. A dimension that doesn’t apply to the study type is skipped, never penalized.
This Kaimen Rigor review is model-assisted and is not a substitute for formal expert review. It complements human evaluation by surfacing potential methodological concerns — verify each finding against the source.