Point-of-care early infant HIV diagnosis at birth in a pragmatic cluster-randomized trial in Mozambique and Tanzania: A comparative cost and cost-effectiveness studyElsbernd, Kira;Sabi, Issa;Jani, Ilesh V.;Mudenyanga, Chishamiso;Boniface, Siriel;Mahumane, Arlete;Lequechane, Joaquim;Chale, Falume;Meggi, Bindiya;Pereira, Kassia;Edom, Raphael;Lwilla, Anange F.;Buck, W. Chris;Ntinyinya, Nyanda Elias;Hoelscher, Michael;Baernighausen, Till;Kroidl, Arne;Kohler, Stefan;Consortium, the LIFE Study
doi: 10.1371/journal.pmed.1005069pmid: 42090430
Background Timely access to early infant diagnosis (EID) is crucial for newborns with HIV, as late diagnosis can delay lifesaving antiretroviral treatment (ART). We assessed the comparative cost and cost-effectiveness of integrating point-of-care EID at birth into routine care in primary healthcare settings. Methods and findings This pre-specified secondary analysis was nested in the cluster-randomized LIFE study conducted at 28 primary healthcare facilities in Mozambique and Tanzania from October 2019 to September 2021. We estimated the health system cost of point-of-care birth plus 4–8-week HIV testing (very early infant diagnosis; VEID) compared to standard-of-care (SoC) testing at 4–8 weeks only, both with immediate ART initiation. We assessed the cost-effectiveness of VEID relative to SoC with respect to ART initiation within one week of life using Bayesian hierarchical models. As this is an intermediate outcome, incremental cost-effectiveness ratios (ICERs) cannot be directly compared to available life-year-based cost-effectiveness thresholds. To contextualize results, we derived the minimum life-years gained per early ART initiation required for VEID to meet standard thresholds in a break-even analysis. VEID was associated with a higher cost and resulted in earlier ART initiation than SoC in both countries. In Mozambique, VEID increased the proportion of infants initiating ART within one week of life by 90.0 (95% CrI [67.5, 98.5]) percentage points at an incremental cost of $2,632 (95% CrI [$2,249, $3,062]) per infant with HIV. In Tanzania, VEID increased early ART initiation by 59.9 (95% CrI [20.9, 89.5]) percentage points at an incremental cost of $6,263 (95% CrI [$5,394, $7,243]) per infant with HIV. The ICER was $2,924 and $10,458 in Mozambique and Tanzania, respectively and was sensitive to intrauterine transmission rate. These findings were limited by the lack of long-term health outcome data and reliance on an intermediate outcome. Based on the break-even analysis, we estimated that VEID would need to yield 6–32 life-years gained per additional early ART initiation to meet standard thresholds. Conclusions Adding birth testing improved early ART initiation but was unlikely to be cost-effective relative to standard thresholds given current prices, vertical transmission rates, and knowledge of long-term health benefits. Cost-effectiveness could be achieved at current costs if early ART translates to substantial long-term health benefits or if targeted to infants at high risk of vertical transmission. Why was this study done? Newborns who acquire HIV before birth are at high risk of illness and death unless they start treatment quickly. Testing for HIV at birth with rapid, same-day point-of-care (PoC) tests could identify these newborns sooner than the usual test at 4–8 weeks, allowing earlier treatment. Adding PoC testing at birth requires extra resources, and it is unclear whether the health gains justify the higher costs in low-resource countries. What did the researchers do and find? We evaluated PoC birth testing plus the routine 4–8 weeks test (called VEID) against the standard approach of testing only at 4–8 weeks (called SoC) within a pragmatic cluster-randomized trial at 28 primary healthcare facilities in Mozambique and Tanzania between 2019 and 2021. Infants born in VEID sites more often received PoC testing, more often received treatment, and started treatment earlier than in SoC sites. To be cost-effective, VEID would need to produce large long-term health gains. What do these findings mean? Adding birth PoC testing can speed up lifesaving treatment for infants with HIV, but universal roll-out in settings with low HIV transmission is unlikely to be cost-effective without targeted use, lower test prices, sharing of testing resources across programs, or large long-term benefits. A limitation of our study is that infants were followed only for the first few months of life, meaning we could not directly measure the long-term health benefits of starting treatment earlier. Instead, we estimated the minimum survival benefit required (6–32 additional years of life) for VEID to be considered cost-effective at current prices. Introduction Timely access to HIV early infant diagnosis (EID) could improve the health outcomes of the 1.3 million children born to mothers living with HIV each year globally [1]. In 2023, approximately 120,000 of these children acquired HIV through gestation, birth, or breastfeeding. Early diagnosis is required for early antiretroviral treatment (ART) initiation, which is especially critical for neonates acquiring HIV in-utero, half of whom die before two years of age without treatment [2]. The World Health Organization (WHO) currently recommends that all infants born to mothers living with HIV receive an EID test by two months of life and all infants diagnosed with HIV immediately initiate ART [3]. Late diagnostic testing and frequent loss to retention after birth cause delays in access to ART [4], often past a peak in HIV-related mortality reported at 2–3 months of age [5]. Same-day point-of-care (PoC) EID has improved access and decreased time to treatment initiation by streamlining EID and linkage to care processes [6–9] and is cost-effective compared to laboratory-based testing [10]. Yet in 2023, only 67% of infants exposed to HIV were tested in the first two months of life and only 57% of children 0–14 years living with HIV were on ART [1]. Testing at birth, in addition to the standard 4–8 weeks of age, offers the possibility to identify infants acquiring HIV in-utero earlier. Immediate ART in the first weeks of life has the potential to reduce early mortality and morbidity, prevent or lessen the development of long-lasting viral reservoirs, and improve viral control [11–15]. EID programs will likely need to consider same-day test-and-treat in the first week of life to reduce persistently high early HIV-related mortality among infants. PoC EID is now widely available, but few sub-Saharan African countries have implemented birth testing into routine practice, partially due to uncertainty around costs and cost-effectiveness. A modelling study for South Africa suggests that adding birth testing to standard EID schedules is cost-effective [16]. Cost-effectiveness analyses informed by pragmatic trials of PoC EID at birth could provide additional support for program-level implementation, especially from resource-poor, rural, and peri-urban settings where effective EID programs are most needed. This study was conducted within a pragmatic trial of PoC EID at birth in public primary healthcare facilities in Mozambique and Tanzania, where high HIV prevalence and persistent vertical transmission [1] suggest that adding birth test-and-treat could have a meaningful impact on HIV-related early infant mortality and morbidity in these settings. The objective of the study was to estimate the cost and cost-effectiveness of offering PoC EID and immediate ART to infants exposed to HIV under routine conditions at birth plus 4–8 weeks of age (very early infant diagnosis; VEID) compared with the standard-of-care (SoC) at 4–8 weeks of age only. Cost-effectiveness was evaluated with respect to early ART initiation and EID uptake, which are intermediate outcomes on the causal pathway to reduced early infant mortality [14]. Methods Study setting This trial-based cost and cost-effectiveness analysis was nested in the LIFE study (NCT04032522), a pragmatic cluster-randomized trial which took place at 28 primary healthcare facilities in the Sofala and Manica provinces of Mozambique and the Mbeya and Songwe regions of Tanzania (7 sites per country per arm). The LIFE study enrolled 6,602 infants born to women living with HIV from October 2019 to September 2021: 3,294 in VEID sites and 3,308 in SoC sites. Among 125 infants diagnosed with HIV until 12 weeks of age, the study demonstrated a clinically relevant but not significant reduction in mortality up to 6 months of age with VEID. A total of 65 infants (52.0%) were diagnosed with HIV at birth. Vertical transmission was 1.89% (95% confidence interval [CI] [1.58, 2.25]) overall with 86.4% of infants (108) from Mozambique and 13.6% (17) from Tanzania [15]. Ethical considerations Ethical approvals for the LIFE study were obtained from the Comité Institucional de Bioética para a Saúde of the Instituto Nacional de Saúde and Comité Nacional de Bioética em Saúde (Ref No. 509/CNBS/20) in Mozambique, the Mbeya Medical Research and Ethics Committee (Ref No. SZEC-2439/R.E/V. 1/90) and the Medical Research Coordinating Committee of the National Institute for Medical Research (Ref No. NIMR/HQ/R.8a/Vol. IX/3071) in Tanzania, and the Ethics Committee of the Ludwig Maximilians University Hospital (Ref No. 19–441) in Germany. All study participants provided written informed consent for themselves and their infants. Testing and treatment procedures Half of the sites implemented PoC EID within 72 hours of birth and at 4–8 weeks of age (VEID) and the other half at 4–8 weeks of age only (SoC) (Fig 1). Infants with negative or unknown HIV status at birth started post-natal prophylaxis and received PoC EID at follow-up visits at 4–8 and 12 weeks of age (plus 4-week window period). HIV–positive results were confirmed by a second PoC EID test or, if a valid result could not be obtained on site, HIV-DNA performed at a central laboratory from dried blood spots. Infants diagnosed with HIV were immediately started on ART. Dried blood spots were collected for all infants at SoC sites at birth for retrospective analysis of HIV status at birth and in the case of death or loss to follow-up with unknown HIV status. Download: PNG larger image TIFF original image Fig 1. Testing and treatment algorithms. HIV-exposed infants received their first HIV test in the maternity ward (VEID) or the pediatric clinic (SoC), where follow-up visits were also performed. (A) In Mozambique, each clinic had one mPIMA analyzer. (B) In Tanzania, each health facility had one Xpert analyzer. Infants with HIV–positive test results were immediately initiated on ART. Infants with HIV–negative test results were initiated on post-natal prophylaxis according to country guidelines: AZT + NVP for all infants in Mozambique and high-risk infants in Tanzania and NVP only for low-risk infants in Tanzania.VEID, very early infant (HIV) diagnosis; SoC = standard of care; PoC EID, point-of-care early infant (HIV) diagnosis; AZT, zidovudine; 3TC, lamivudine; NVP, nevirapine; ABC, abacavir. Created in BioRender. Hoelscher, M. (2025) https://BioRender.com/t4btr5b. https://doi.org/10.1371/journal.pmed.1005069.g001 Following routine guidelines, all infants exposed to HIV in Mozambique were given enhanced post-natal prophylaxis with zidovudine (AZT) syrup for 6 weeks plus nevirapine (NVP) syrup for 12 weeks [17]. In Tanzania, only high-risk infants (i.e., mother newly diagnosed with HIV, not on ART or on ART for less than four weeks, or with high viral load in the last four weeks), were given enhanced post-natal prophylaxis. Standard post-natal prophylaxis with NVP syrup for 6 weeks was given to low-risk infants [18]. ART for infants with HIV aged 0–4 weeks and at least 2 kg consisted of AZT, lamivudine (3TC), and NVP syrups. At 4 weeks of age and at least 3 kg, infants were given abacavir (ABC)/3TC dispersible tablets and lopinavir/ritonavir (LPV/r) granules. Nurses were trained in sample collection and use of the PoC analyzers, performed all PoC testing and pre- and post-test counseling, and initiated ART with physician support. Testing platforms The Abbott mPIMA HIV-1/2 Detect and Cepheid Xpert HIV-1 Qual were used for PoC EID testing in Mozambique and Tanzania, respectively, reflecting preexisting local HIV program preferences. The mPIMA analyzer is also approved for HIV viral load and can be procured with an external battery to bridge power outages. The Xpert analyzer can run HIV viral load, tuberculosis diagnosis, and other assays (though equipment sharing across programs is not common practice in the study setting). In Tanzania, two variations of the Xpert analyzer were available: Xpert II with the capacity to run two and Xpert IV four assays concurrently. In general, Xpert IV analyzers were placed at higher volume sites. In Tanzania, PoC analyzers were located in the pediatric outpatient clinic or health facility laboratory. In Mozambique, PoC analyzers were placed at the pediatric outpatient clinic, and an additional analyzer was placed at the maternity ward in VEID sites. Both analyzers have comparable diagnostic performance and run-time for EID [19,20]. Therefore, we expect that the choice of platform did not affect linkage to ART or other health outcomes. Costing approach We conducted a micro-costing study from the health system perspective. We estimated implementation and operations costs as they occurred in the LIFE study, including fixed costs (equipment purchase, installation, and maintenance; utilities; communication; training; and facility upgrades) and variable costs (consumables, labor, and ART) for VEID and SoC approaches. Direct client and societal costs were excluded from this analysis, as the study was designed to focus on direct budgetary and resource allocation impact on the healthcare provider. The time horizon for individual-level cost and effectiveness outcomes corresponded to the 12-week follow-up period, during which clinical pathways and resource use differed between arms. Beyond this period, costs and clinical management were assumed to be equivalent in both arms. Infants lost to follow-up after birth in the SoC arm incurred no intervention-related resource use and were therefore treated as structural zeros with respect to costs. We first identified resources used for VEID from study documents, local HIV testing and care guidelines, and communication with implementing partners. To measure quantities of resources consumed, we used data from the LIFE study and usage data from the PoC analyzers. Consumables, labor, and ART were collected using a bottom-up approach. Time required for nurses to conduct PoC EID tests and perform HIV-related counselling was collected as anonymous self-reported average active work time once they had logged sufficient experience with PoC testing (>50 tests/individual). Fixed costs were allocated to EID testing according to observed testing volume; consumables, labor, and ART costs were allocated directly per test or per infant. Facility upgrades included the purchase of cabinets to store reagents and the installation of air conditioning units in some sites. Fixed costs were shared with routine services, including HIV viral load monitoring. Shared resources (e.g., analyzers, utilities, and communication) were allocated proportionally based on observed PoC EID testing volume relative to total analyzer use. Prices for resources were collected from Ministry of Health budgets and expenditure reports, purchase tenders, manufacturer contracts, and government salary scales. Test cartridges were procured at a flat rate inclusive of shipping, customs clearance, distribution, and administration costs. Capital costs were discounted at 3% per year and amortized using equivalent annual cost—PoC platform costs over a 5-year life span according to manufacturer specifications and other capital costs over 10 years. Cost data collected for 2020 and 2021 were converted to nominal 2020 United States dollars (US$) using World Bank Global Economic Monitor annual mean exchange rates [21]. At the time of the study, these were 66.8 Mozambique Metical and 2304.4 Tanzanian Shillings per 1 US$ for 2020 and 64.4 Mozambique Metical per 1 US$ for 2021. No costs paid in Tanzanian Shillings were collected for 2021. Further costing details are provided in S1 Text. Outcomes We assessed the proportion of infants with HIV initiating ART within 1 week of life, defined as trial-documented initiation of ART on or before day seven after birth, as the primary effectiveness outcome. We also assessed the proportion of infants exposed to HIV receiving an EID test within eight weeks of life as a secondary effectiveness outcome. EID testing was documented in study visit records and confirmed with PoC analyzer logs. Incremental costs, incremental effectiveness, and incremental cost-effectiveness ratios (ICERs) per additional infant initiating ART within 1 week and per additional infant receiving a PoC EID test within 8 weeks were estimated relative to the SoC. In the absence of robust data quantifying the incremental survival benefit associated with starting ART in the first week relative to starting at 4–8 weeks of life, we conducted a break-even analysis to derive the minimum life-years that would need to be gained per additional infant initiating ART within one week of life for our intervention to be considered cost-effective given standard cost-effectiveness thresholds. Statistical analysis Frequentist methods were used for descriptive summaries and unadjusted comparisons between study arms. Descriptive cost per test and per infant were reported as testing volume-weighted mean and bootstrapped 95% CI per country and study arm to reflect typical costs across sites. To show the influence of testing volume on cost per test, we additionally presented unweighted medians by low-, medium-, and high-volume sites. Summary measures for effectiveness outcomes were reported as proportions for the probability of EID testing and ART initiation and median and range for age at first EID test and ART initiation. Incremental costs and incremental effectiveness were estimated using multivariate Bayesian hierarchical models to account for skewed distribution of cost data and clustering at the health facility level, with results reported as posterior mean and 95% credible interval (CrI). Costs were modeled on the natural scale using a hurdle Gamma specification to accommodate observed zeros arising from infants who did not receive EID testing due to loss to follow-up, while effectiveness outcomes were modeled as Bernoulli processes. Models were fit separately for 1-week ART and 8-week PoC EID effectiveness outcomes with country- and site-level random effects to capture intra-cluster correlation and differences in costs, PoC testing platforms, prophylaxis protocols, and transmission rates. For the 1-week ART model, site-level random effects were not included due to sparse outcome variation within clusters, which precluded reliable estimation of additional variance components. The country was included as a fixed effect to account for the systematic differences in context. Informative priors were specified for study arm effects to reflect expected differences between VEID and SoC arms. Prior sensitivity analyses using weakly informative priors were conducted to assess the robustness of model estimates to prior specification. Model convergence was confirmed (all R� ≈ 1), and posterior predictive checks indicated good agreement between observed and replicated data. Full model specification, priors, and diagnostics are provided in S1 Text. ICER point estimates were calculated as the ratio of mean incremental cost to mean incremental effect from the posterior distributions. To reflect uncertainty around these estimates, we used posterior draws of incremental cost and incremental effect jointly sampled from the Bayesian hierarchical models in probabilistic sensitivity analyses. Uncertainty was visualized in cost-effectiveness planes showing additional costs versus additional benefit for VEID over SoC. The probability of each PoC EID approach being cost-effective at a range of WTP thresholds is shown in cost-effectiveness acceptability curves. To contextualize results, we referenced empirically derived country-level cost-effectiveness thresholds from Pichon-Riviere and colleagues [22], which are based on health expenditure per capita and life expectancy growth and expressed as cost per life-year gained. These were $189 per life-year in Mozambique and $316 per life-year in Tanzania (2020 US$). For comparability with previous studies, we also considered GDP per capita as a willingness-to-pay (WTP) threshold ($462 in Mozambique and $1,117 in Tanzania per life-year gained) [21]. As ICERs in this analysis are expressed in cost per additional infant initiating ART within 1 week of life rather than per life-year gained, they cannot be directly compared to these thresholds. Rather than model long-term health outcomes directly, we conducted a break-even analysis to derive the minimum life-years that would need to be gained per additional infant initiating ART within one week of life for VEID to meet these thresholds, calculated as the posterior draw of incremental cost divided by the WTP threshold. Results are reported as median and 95% CrI. Cost data was compiled in Microsoft Excel (Microsoft Corp.). Statistical analysis was performed using R version 4.2.3 (R Foundation for Statistical Computing, Vienna, Austria). Analyses followed a pre-specified health economic analysis plan. This study is reported as per the Consolidated Health Economic Evaluation Reporting Standards 2022 (CHEERS 2022) Statement [23] (S1 Checklist). Figures were created with R and BioRender.com. Sensitivity analyses Deterministic sensitivity analyses were utilized to explore the impact of selected drivers of cost per test not explicitly parameterized in the Bayesian hierarchical models. We chose ranges reflective of plausible uncertainty and policy-relevant scenarios. Consumable prices were varied by ±50% to account for uncertainty in procurement and distribution costs, testing volume from −50% to 200% of observed levels to reflect variability across sites and potential scale-up, analyzer life span from 4 to 15 years for a conservative lower bound and assumptions that PoC platforms can last longer with maintenance, and discount rate from 0% to 6% in line with standard guidance [24]. This analysis was designed to decompose cost-input uncertainty and identify key drivers of unit costs across policy-relevant scenarios, and is distinct from the probabilistic sensitivity analysis, which addresses decision uncertainty. We additionally varied the intrauterine transmission rate across a wide range of values within a probabilistic sensitivity analysis framework to assess how cost-effectiveness would vary across epidemiological contexts and to identify threshold levels of transmission at which VEID would meet common cost-effectiveness benchmarks. Additional probabilistic uncertainty around ICERs was captured in the main analysis and visualized in cost-effectiveness planes and acceptability curves. Here, we focused on scenario-driven variation. Study setting This trial-based cost and cost-effectiveness analysis was nested in the LIFE study (NCT04032522), a pragmatic cluster-randomized trial which took place at 28 primary healthcare facilities in the Sofala and Manica provinces of Mozambique and the Mbeya and Songwe regions of Tanzania (7 sites per country per arm). The LIFE study enrolled 6,602 infants born to women living with HIV from October 2019 to September 2021: 3,294 in VEID sites and 3,308 in SoC sites. Among 125 infants diagnosed with HIV until 12 weeks of age, the study demonstrated a clinically relevant but not significant reduction in mortality up to 6 months of age with VEID. A total of 65 infants (52.0%) were diagnosed with HIV at birth. Vertical transmission was 1.89% (95% confidence interval [CI] [1.58, 2.25]) overall with 86.4% of infants (108) from Mozambique and 13.6% (17) from Tanzania [15]. Ethical considerations Ethical approvals for the LIFE study were obtained from the Comité Institucional de Bioética para a Saúde of the Instituto Nacional de Saúde and Comité Nacional de Bioética em Saúde (Ref No. 509/CNBS/20) in Mozambique, the Mbeya Medical Research and Ethics Committee (Ref No. SZEC-2439/R.E/V. 1/90) and the Medical Research Coordinating Committee of the National Institute for Medical Research (Ref No. NIMR/HQ/R.8a/Vol. IX/3071) in Tanzania, and the Ethics Committee of the Ludwig Maximilians University Hospital (Ref No. 19–441) in Germany. All study participants provided written informed consent for themselves and their infants. Testing and treatment procedures Half of the sites implemented PoC EID within 72 hours of birth and at 4–8 weeks of age (VEID) and the other half at 4–8 weeks of age only (SoC) (Fig 1). Infants with negative or unknown HIV status at birth started post-natal prophylaxis and received PoC EID at follow-up visits at 4–8 and 12 weeks of age (plus 4-week window period). HIV–positive results were confirmed by a second PoC EID test or, if a valid result could not be obtained on site, HIV-DNA performed at a central laboratory from dried blood spots. Infants diagnosed with HIV were immediately started on ART. Dried blood spots were collected for all infants at SoC sites at birth for retrospective analysis of HIV status at birth and in the case of death or loss to follow-up with unknown HIV status. Download: PNG larger image TIFF original image Fig 1. Testing and treatment algorithms. HIV-exposed infants received their first HIV test in the maternity ward (VEID) or the pediatric clinic (SoC), where follow-up visits were also performed. (A) In Mozambique, each clinic had one mPIMA analyzer. (B) In Tanzania, each health facility had one Xpert analyzer. Infants with HIV–positive test results were immediately initiated on ART. Infants with HIV–negative test results were initiated on post-natal prophylaxis according to country guidelines: AZT + NVP for all infants in Mozambique and high-risk infants in Tanzania and NVP only for low-risk infants in Tanzania.VEID, very early infant (HIV) diagnosis; SoC = standard of care; PoC EID, point-of-care early infant (HIV) diagnosis; AZT, zidovudine; 3TC, lamivudine; NVP, nevirapine; ABC, abacavir. Created in BioRender. Hoelscher, M. (2025) https://BioRender.com/t4btr5b. https://doi.org/10.1371/journal.pmed.1005069.g001 Following routine guidelines, all infants exposed to HIV in Mozambique were given enhanced post-natal prophylaxis with zidovudine (AZT) syrup for 6 weeks plus nevirapine (NVP) syrup for 12 weeks [17]. In Tanzania, only high-risk infants (i.e., mother newly diagnosed with HIV, not on ART or on ART for less than four weeks, or with high viral load in the last four weeks), were given enhanced post-natal prophylaxis. Standard post-natal prophylaxis with NVP syrup for 6 weeks was given to low-risk infants [18]. ART for infants with HIV aged 0–4 weeks and at least 2 kg consisted of AZT, lamivudine (3TC), and NVP syrups. At 4 weeks of age and at least 3 kg, infants were given abacavir (ABC)/3TC dispersible tablets and lopinavir/ritonavir (LPV/r) granules. Nurses were trained in sample collection and use of the PoC analyzers, performed all PoC testing and pre- and post-test counseling, and initiated ART with physician support. Testing platforms The Abbott mPIMA HIV-1/2 Detect and Cepheid Xpert HIV-1 Qual were used for PoC EID testing in Mozambique and Tanzania, respectively, reflecting preexisting local HIV program preferences. The mPIMA analyzer is also approved for HIV viral load and can be procured with an external battery to bridge power outages. The Xpert analyzer can run HIV viral load, tuberculosis diagnosis, and other assays (though equipment sharing across programs is not common practice in the study setting). In Tanzania, two variations of the Xpert analyzer were available: Xpert II with the capacity to run two and Xpert IV four assays concurrently. In general, Xpert IV analyzers were placed at higher volume sites. In Tanzania, PoC analyzers were located in the pediatric outpatient clinic or health facility laboratory. In Mozambique, PoC analyzers were placed at the pediatric outpatient clinic, and an additional analyzer was placed at the maternity ward in VEID sites. Both analyzers have comparable diagnostic performance and run-time for EID [19,20]. Therefore, we expect that the choice of platform did not affect linkage to ART or other health outcomes. Costing approach We conducted a micro-costing study from the health system perspective. We estimated implementation and operations costs as they occurred in the LIFE study, including fixed costs (equipment purchase, installation, and maintenance; utilities; communication; training; and facility upgrades) and variable costs (consumables, labor, and ART) for VEID and SoC approaches. Direct client and societal costs were excluded from this analysis, as the study was designed to focus on direct budgetary and resource allocation impact on the healthcare provider. The time horizon for individual-level cost and effectiveness outcomes corresponded to the 12-week follow-up period, during which clinical pathways and resource use differed between arms. Beyond this period, costs and clinical management were assumed to be equivalent in both arms. Infants lost to follow-up after birth in the SoC arm incurred no intervention-related resource use and were therefore treated as structural zeros with respect to costs. We first identified resources used for VEID from study documents, local HIV testing and care guidelines, and communication with implementing partners. To measure quantities of resources consumed, we used data from the LIFE study and usage data from the PoC analyzers. Consumables, labor, and ART were collected using a bottom-up approach. Time required for nurses to conduct PoC EID tests and perform HIV-related counselling was collected as anonymous self-reported average active work time once they had logged sufficient experience with PoC testing (>50 tests/individual). Fixed costs were allocated to EID testing according to observed testing volume; consumables, labor, and ART costs were allocated directly per test or per infant. Facility upgrades included the purchase of cabinets to store reagents and the installation of air conditioning units in some sites. Fixed costs were shared with routine services, including HIV viral load monitoring. Shared resources (e.g., analyzers, utilities, and communication) were allocated proportionally based on observed PoC EID testing volume relative to total analyzer use. Prices for resources were collected from Ministry of Health budgets and expenditure reports, purchase tenders, manufacturer contracts, and government salary scales. Test cartridges were procured at a flat rate inclusive of shipping, customs clearance, distribution, and administration costs. Capital costs were discounted at 3% per year and amortized using equivalent annual cost—PoC platform costs over a 5-year life span according to manufacturer specifications and other capital costs over 10 years. Cost data collected for 2020 and 2021 were converted to nominal 2020 United States dollars (US$) using World Bank Global Economic Monitor annual mean exchange rates [21]. At the time of the study, these were 66.8 Mozambique Metical and 2304.4 Tanzanian Shillings per 1 US$ for 2020 and 64.4 Mozambique Metical per 1 US$ for 2021. No costs paid in Tanzanian Shillings were collected for 2021. Further costing details are provided in S1 Text. Outcomes We assessed the proportion of infants with HIV initiating ART within 1 week of life, defined as trial-documented initiation of ART on or before day seven after birth, as the primary effectiveness outcome. We also assessed the proportion of infants exposed to HIV receiving an EID test within eight weeks of life as a secondary effectiveness outcome. EID testing was documented in study visit records and confirmed with PoC analyzer logs. Incremental costs, incremental effectiveness, and incremental cost-effectiveness ratios (ICERs) per additional infant initiating ART within 1 week and per additional infant receiving a PoC EID test within 8 weeks were estimated relative to the SoC. In the absence of robust data quantifying the incremental survival benefit associated with starting ART in the first week relative to starting at 4–8 weeks of life, we conducted a break-even analysis to derive the minimum life-years that would need to be gained per additional infant initiating ART within one week of life for our intervention to be considered cost-effective given standard cost-effectiveness thresholds. Statistical analysis Frequentist methods were used for descriptive summaries and unadjusted comparisons between study arms. Descriptive cost per test and per infant were reported as testing volume-weighted mean and bootstrapped 95% CI per country and study arm to reflect typical costs across sites. To show the influence of testing volume on cost per test, we additionally presented unweighted medians by low-, medium-, and high-volume sites. Summary measures for effectiveness outcomes were reported as proportions for the probability of EID testing and ART initiation and median and range for age at first EID test and ART initiation. Incremental costs and incremental effectiveness were estimated using multivariate Bayesian hierarchical models to account for skewed distribution of cost data and clustering at the health facility level, with results reported as posterior mean and 95% credible interval (CrI). Costs were modeled on the natural scale using a hurdle Gamma specification to accommodate observed zeros arising from infants who did not receive EID testing due to loss to follow-up, while effectiveness outcomes were modeled as Bernoulli processes. Models were fit separately for 1-week ART and 8-week PoC EID effectiveness outcomes with country- and site-level random effects to capture intra-cluster correlation and differences in costs, PoC testing platforms, prophylaxis protocols, and transmission rates. For the 1-week ART model, site-level random effects were not included due to sparse outcome variation within clusters, which precluded reliable estimation of additional variance components. The country was included as a fixed effect to account for the systematic differences in context. Informative priors were specified for study arm effects to reflect expected differences between VEID and SoC arms. Prior sensitivity analyses using weakly informative priors were conducted to assess the robustness of model estimates to prior specification. Model convergence was confirmed (all R� ≈ 1), and posterior predictive checks indicated good agreement between observed and replicated data. Full model specification, priors, and diagnostics are provided in S1 Text. ICER point estimates were calculated as the ratio of mean incremental cost to mean incremental effect from the posterior distributions. To reflect uncertainty around these estimates, we used posterior draws of incremental cost and incremental effect jointly sampled from the Bayesian hierarchical models in probabilistic sensitivity analyses. Uncertainty was visualized in cost-effectiveness planes showing additional costs versus additional benefit for VEID over SoC. The probability of each PoC EID approach being cost-effective at a range of WTP thresholds is shown in cost-effectiveness acceptability curves. To contextualize results, we referenced empirically derived country-level cost-effectiveness thresholds from Pichon-Riviere and colleagues [22], which are based on health expenditure per capita and life expectancy growth and expressed as cost per life-year gained. These were $189 per life-year in Mozambique and $316 per life-year in Tanzania (2020 US$). For comparability with previous studies, we also considered GDP per capita as a willingness-to-pay (WTP) threshold ($462 in Mozambique and $1,117 in Tanzania per life-year gained) [21]. As ICERs in this analysis are expressed in cost per additional infant initiating ART within 1 week of life rather than per life-year gained, they cannot be directly compared to these thresholds. Rather than model long-term health outcomes directly, we conducted a break-even analysis to derive the minimum life-years that would need to be gained per additional infant initiating ART within one week of life for VEID to meet these thresholds, calculated as the posterior draw of incremental cost divided by the WTP threshold. Results are reported as median and 95% CrI. Cost data was compiled in Microsoft Excel (Microsoft Corp.). Statistical analysis was performed using R version 4.2.3 (R Foundation for Statistical Computing, Vienna, Austria). Analyses followed a pre-specified health economic analysis plan. This study is reported as per the Consolidated Health Economic Evaluation Reporting Standards 2022 (CHEERS 2022) Statement [23] (S1 Checklist). Figures were created with R and BioRender.com. Sensitivity analyses Deterministic sensitivity analyses were utilized to explore the impact of selected drivers of cost per test not explicitly parameterized in the Bayesian hierarchical models. We chose ranges reflective of plausible uncertainty and policy-relevant scenarios. Consumable prices were varied by ±50% to account for uncertainty in procurement and distribution costs, testing volume from −50% to 200% of observed levels to reflect variability across sites and potential scale-up, analyzer life span from 4 to 15 years for a conservative lower bound and assumptions that PoC platforms can last longer with maintenance, and discount rate from 0% to 6% in line with standard guidance [24]. This analysis was designed to decompose cost-input uncertainty and identify key drivers of unit costs across policy-relevant scenarios, and is distinct from the probabilistic sensitivity analysis, which addresses decision uncertainty. We additionally varied the intrauterine transmission rate across a wide range of values within a probabilistic sensitivity analysis framework to assess how cost-effectiveness would vary across epidemiological contexts and to identify threshold levels of transmission at which VEID would meet common cost-effectiveness benchmarks. Additional probabilistic uncertainty around ICERs was captured in the main analysis and visualized in cost-effectiveness planes and acceptability curves. Here, we focused on scenario-driven variation. Results Age of early infant HIV diagnostic testing and ART initiation In VEID sites, 100% of infants exposed to HIV had a PoC EID test by the recommended 8 weeks of age (median 1 day, interquartile range [IQR]: 0–1 days), compared with 84.4% in SoC sites (median 4.7 weeks, IQR: 4.4–6.4) (Fig 2A and Fig AA in S1 Text). Proportions of infants receiving PoC EID testing by eight weeks in SoC sites differed substantially by country with 91% in Mozambique and 75% in Tanzania. Repeat testing due to errors or confirmatory testing was performed for 9.5% versus 3.7% in Mozambique and 9.9% versus 8.7% in Tanzania of infants for VEID and SoC approaches, respectively. Download: PNG larger image TIFF original image Fig 2. Time from birth to PoC EID testing and treatment initiation. (A) First EID test among all HIV-exposed infants; proportion receiving an EID test indicated in vertical axis labels. (B) ART initiation among infants diagnosed with HIV up to 16 weeks; proportion initiating ART at all indicated in vertical axis labels. Box plots show median (center line), 25% and 75% (box limits), and 5% and 95% (whiskers); violin plots show data distribution. VEID, very early infant (HIV) diagnosis; SoC = standard of care; PoC EID, point-of-care early infant (HIV) diagnosis; ART, antiretroviral treatment. Created in BioRender. Hoelscher, M. (2025) https://BioRender.com/u260ruo. https://doi.org/10.1371/journal.pmed.1005069.g002 In VEID sites, 98.6% of infants diagnosed with HIV started ART compared to 89.3% in SoC sites, and among infants with an HIV–positive result available at birth (VEID sites), 89.5% started ART within the first week of life. The median age at ART initiation was 0.9 (range: 0–14) weeks in VEID sites and 4.7 (range: 4–16) weeks in SoC sites (Fig 2B and Fig AB in S1 Text). Reasons for not starting ART were death or loss to follow-up. Reasons for delayed treatment initiation at birth were not meeting the minimum weight requirement for neonatal antiretroviral dosing and a delay in confirmatory HIV testing. Costs of very early infant diagnosis versus standard-of-care Cost per test was $39.81 (95% CI [$37.44, $43.70]) for the VEID approach and $41.97 (95% CI [$39.04, $46.58]) for the SoC approach using mPIMA in Mozambique. Using Xpert in Tanzania, these costs were $33.91 (95% CI [$30.85, $40.92]) for VEID and $39.90 (95% CI [$34.30, $49.75]) for SoC. Lower cost per test for the VEID approach reflects higher testing volumes, since infants were tested twice by design and fixed costs were allocated across more tests. Cost per infant exposed to HIV, which incorporates PoC EID testing and, if appropriate, earlier ART initiation was $90.13 (95% CI [$88.76, $91.47]) for VEID versus $39.70 (95% CI [$39.00, $40.42]) for SoC in Mozambique and $67.81 (95% CI [$66.26, $69.43]) for VEID versus $35.92 (95% CI [$34.67, $37.20]) for SoC in Tanzania. Cost inputs used to derive unit costs per test and per infant are listed in Table 1. Download: PNG larger image TIFF original image Table 1. Costs of inputs. Costs expressed in 2020 US$ by category for each country/PoC testing platform. https://doi.org/10.1371/journal.pmed.1005069.t001 Consumables, primarily the test cartridge, accounted for 63% in Mozambique and 58% in Tanzania of the cost per test (Fig 3, Fig B and Table A in S1 Text). Equipment and overhead costs, by contrast, varied substantially by testing volume. We categorized health facilities into high (>12 tests per week), medium (5–12 tests per week) and low (<5 tests per week) volume for ease of comparison (S1 Text). On average in Mozambique, four (29%) sites ran over 15 tests weekly and two (14%) ran less than five tests weekly. In Tanzania, only one (7%) site ran over 15 tests weekly and six (43%) ran less than five tests weekly. Cost per test varied by 33% in Mozambique and 63% in Tanzania between high and low volume sites, with the greater variability in Tanzania driven by a larger number of low volume sites compared to Mozambique. Labor and overhead costs were small in comparison to consumable and equipment costs in both countries. Download: PNG larger image TIFF original image Fig 3. PoC EID test cost components. (A) Cost estimates expressed in 2020 US$. (B) Proportion of total cost per test. Sites are categorized as high (>12 tests per week), medium (5–12 tests per week), and low (<5 tests per week) testing volume. Equipment includes amortized costs of initial purchase, installation, and yearly maintenance of PoC analyzers. Overhead includes apportioned costs of electricity, communications, and facility upgrades. Overhead and labor costs were < $3 per test (<6% of cost per test) across all sites. https://doi.org/10.1371/journal.pmed.1005069.g003 Cost-effectiveness of very early infant HIV diagnosis VEID was associated with higher costs and improved effectiveness outcomes in both countries (Fig 4, Tables B-C in S1 Text). In Mozambique, VEID increased the proportion of infants initiating ART within one week of life by 90.0 (95% CrI [67.5, 98.5]) percentage points at an incremental cost of $2,632 (95% CrI [$2,249, $3,062]) per infant with HIV. In Tanzania, the incremental cost was higher at $6,263 (95% CrI [$5,394, $7,243]), with a 59.9 (95% CrI [20.9, 89.5]) percentage point increase in the proportion initiating ART within one week. The corresponding ICER with respect to additional infants initiating ART within one week of life was $2,924 in Mozambique and $10,458 in Tanzania. Download: PNG larger image TIFF original image Fig 4. Cost-effectiveness of VEID with respect to PoC EID testing and early ART. Cost-effectiveness planes presenting the incremental costs and (A and B) early ART benefit (% of eligible infants with HIV initiated on ART within 1 week of life; eligible infants include those meeting clinical criteria for treatment, e.g., weight) or 8-week EID benefit (%) for VEID (E and F). Points represent estimates of costs and health benefit for VEID relative to SoC at the origin. A random sample of 500 points is plotted. All estimates indicate that VEID is more expensive and results in increased health benefit (upper right quadrant). (C and D; G and H) Cost-effectiveness acceptability curves for each outcome in A-B and E-F, respectively, showing cumulative probabilities of each testing strategy being cost-effective at a particular willingness to pay value. Costs are presented as 2020 US$. PoC EID, point-of-care early infant (HIV) diagnosis; ART, antiretroviral treatment; SoC = standard of care; VEID, very early infant (HIV) diagnosis. https://doi.org/10.1371/journal.pmed.1005069.g004 Based on the break-even analysis, VEID would need to achieve 15.24 (95% CrI [12.63, 21.37]) life-years gained per additional infant initiating ART within 1 week of life in Mozambique and 32.17 (95% CrI [21.15, 94.84]) in Tanzania to meet empirically derived cost-effectiveness thresholds [22] (Fig 5). At WTP thresholds based on 1x GDP per capita, the corresponding required life-years gained were 6.24 (95% CrI [5.17, 8.74]) in Mozambique and 9.10 (95% CrI [5.98, 26.83]) in Tanzania. Download: PNG larger image TIFF original image Fig 5. Required life-years per early ART initiation for cost-effectiveness. Shaded ribbons for Mozambique and Tanzania indicate 95% credible interval (CrI) across posterior draws from the Bayesian hierarchical model for 1-week ART initiation. Empirically derived cost-effectiveness thresholds, based on health expenditure and life expectancy growth [21] and expressed in 2020 US$, were $189 in Mozambique and $316 in Tanzania. 2020 gross domestic product per capita-based thresholds were $462 in Mozambique and $1,117 in Tanzania. Higher required life-years reflect scenarios in which greater health gains per infant would be needed to approach the indicated thresholds. https://doi.org/10.1371/journal.pmed.1005069.g005 Implementation of VEID also increased uptake of EID. In Mozambique, VEID increased the proportion of infants receiving an EID test within eight weeks of life by 8.0 (95% CrI [5.4, 11.2]) percentage points at an incremental cost of $50.72 (95% CrI [$43.72, $58.65]) per infant exposed to HIV. In Tanzania, the corresponding increase was 20.9 (95% CrI [14.9, 28.0]) percentage points in the proportion of infants receiving an EID test within eight weeks of life at an incremental cost of $28.52 (95% CrI [$24.39, $33.19]) per infant exposed to HIV. The ICER with respect to additional infants receiving an EID test within eight weeks of life was $635.47 and $136.19 in Mozambique and Tanzania, respectively. Sensitivity analysis Cost per test was most sensitive to relative changes in the price of consumables, followed by testing volume and PoC analyzer life span (Fig 6 and Fig C in S1 Text). We observed moderate variability in baseline cost across health facilities resulting from different testing volumes and 5% probability of zero cost across all sites (i.e., infants not receiving testing mostly in SoC sites due to loss to follow-up after birth). Site-level variation was modest with an intra-cluster correlation of 0.033 (95% CrI [0.018, 0.059]) for cost per infant and 0.140 (95% CrI [0.066, 0.258]) for PoC EID within eight weeks. Sites with higher baseline costs tended to experience smaller increases in cost per additional test as testing volumes increased (posterior correlation ρ = −0.25), consistent with economies of scale. Due to low number of infants diagnosed with HIV in Tanzania, uncertainty in effectiveness outcomes was high; combined with uncertainty in costs, this resulted in substantial uncertainty in the ICER with respect to ART initiation in the first week of life in Tanzania. Download: PNG larger image TIFF original image Fig 6. Deterministic one-way sensitivity analysis showing influence of key assumptions on cost per PoC EID test. Costs expressed in 2020 US$. Per test cost difference is shown on the x-axis and parameters with ranges are shown on the y-axis. High and low colors indicate the direction in which the parameter varies. This analysis addresses cost-input uncertainty rather than decision uncertainty; see Figure 4 for probabilistic sensitivity analysis of incremental cost-effectiveness. https://doi.org/10.1371/journal.pmed.1005069.g006 Intrauterine transmission rates observed in the LIFE study were 1.01% (95% CrI [0.65, 1.51]) in SoC sites and 1.48% (95% CrI [1.01, 2.06]) in VEID sites in Mozambique and 0.43% (95% CrI [0.16, 0.85]) in SoC sites and 0.50% (95% CrI [0.19, 0.94]) in VEID sites in Tanzania. The ICER with respect to additional infants initiating ART within one week of life was sensitive to intrauterine transmission rate, ranging from $2,394 to $4,009 in Mozambique and from $5,993 to $27,424 in Tanzania between the 5th and 95th percentiles of estimated intrauterine transmission. In Mozambique, intrauterine transmission would need to exceed 34% to meet the cost-effectiveness threshold from [22] or 14% to meet the GDP-based threshold (Fig 7). In Tanzania, the corresponding intrauterine transmission rates were 20% and 5.8%, respectively, indicating that VEID becomes cost-effective only above specific thresholds of early vertical transmission risk. These thresholds are based on cost per additional infant initiating ART within 1 week of life and do not directly incorporate downstream life-years gained, which are explored separately above. Download: PNG larger image TIFF original image Fig 7. Sensitivity analysis of the incremental cost effectiveness ratio (ICER) per infant with HIV on ART within 1 week across a range of intrauterine transmission rates, based on posterior estimates from the Bayesian hierarchical model. Lines show the posterior mean ICER, with shaded areas representing the 95% credible intervals. Dashed horizontal lines indicate reference willingness to pay (WTP) thresholds: empirically derived cost-effectiveness thresholds based on health expenditure and life expectancy growth from [21], expressed in 2020 US$ ($189 for Mozambique and $316 for Tanzania) and gross domestic product per capita in 2020 ($462 for Mozambique and $1,117 for Tanzania). https://doi.org/10.1371/journal.pmed.1005069.g007 Age of early infant HIV diagnostic testing and ART initiation In VEID sites, 100% of infants exposed to HIV had a PoC EID test by the recommended 8 weeks of age (median 1 day, interquartile range [IQR]: 0–1 days), compared with 84.4% in SoC sites (median 4.7 weeks, IQR: 4.4–6.4) (Fig 2A and Fig AA in S1 Text). Proportions of infants receiving PoC EID testing by eight weeks in SoC sites differed substantially by country with 91% in Mozambique and 75% in Tanzania. Repeat testing due to errors or confirmatory testing was performed for 9.5% versus 3.7% in Mozambique and 9.9% versus 8.7% in Tanzania of infants for VEID and SoC approaches, respectively. Download: PNG larger image TIFF original image Fig 2. Time from birth to PoC EID testing and treatment initiation. (A) First EID test among all HIV-exposed infants; proportion receiving an EID test indicated in vertical axis labels. (B) ART initiation among infants diagnosed with HIV up to 16 weeks; proportion initiating ART at all indicated in vertical axis labels. Box plots show median (center line), 25% and 75% (box limits), and 5% and 95% (whiskers); violin plots show data distribution. VEID, very early infant (HIV) diagnosis; SoC = standard of care; PoC EID, point-of-care early infant (HIV) diagnosis; ART, antiretroviral treatment. Created in BioRender. Hoelscher, M. (2025) https://BioRender.com/u260ruo. https://doi.org/10.1371/journal.pmed.1005069.g002 In VEID sites, 98.6% of infants diagnosed with HIV started ART compared to 89.3% in SoC sites, and among infants with an HIV–positive result available at birth (VEID sites), 89.5% started ART within the first week of life. The median age at ART initiation was 0.9 (range: 0–14) weeks in VEID sites and 4.7 (range: 4–16) weeks in SoC sites (Fig 2B and Fig AB in S1 Text). Reasons for not starting ART were death or loss to follow-up. Reasons for delayed treatment initiation at birth were not meeting the minimum weight requirement for neonatal antiretroviral dosing and a delay in confirmatory HIV testing. Costs of very early infant diagnosis versus standard-of-care Cost per test was $39.81 (95% CI [$37.44, $43.70]) for the VEID approach and $41.97 (95% CI [$39.04, $46.58]) for the SoC approach using mPIMA in Mozambique. Using Xpert in Tanzania, these costs were $33.91 (95% CI [$30.85, $40.92]) for VEID and $39.90 (95% CI [$34.30, $49.75]) for SoC. Lower cost per test for the VEID approach reflects higher testing volumes, since infants were tested twice by design and fixed costs were allocated across more tests. Cost per infant exposed to HIV, which incorporates PoC EID testing and, if appropriate, earlier ART initiation was $90.13 (95% CI [$88.76, $91.47]) for VEID versus $39.70 (95% CI [$39.00, $40.42]) for SoC in Mozambique and $67.81 (95% CI [$66.26, $69.43]) for VEID versus $35.92 (95% CI [$34.67, $37.20]) for SoC in Tanzania. Cost inputs used to derive unit costs per test and per infant are listed in Table 1. Download: PNG larger image TIFF original image Table 1. Costs of inputs. Costs expressed in 2020 US$ by category for each country/PoC testing platform. https://doi.org/10.1371/journal.pmed.1005069.t001 Consumables, primarily the test cartridge, accounted for 63% in Mozambique and 58% in Tanzania of the cost per test (Fig 3, Fig B and Table A in S1 Text). Equipment and overhead costs, by contrast, varied substantially by testing volume. We categorized health facilities into high (>12 tests per week), medium (5–12 tests per week) and low (<5 tests per week) volume for ease of comparison (S1 Text). On average in Mozambique, four (29%) sites ran over 15 tests weekly and two (14%) ran less than five tests weekly. In Tanzania, only one (7%) site ran over 15 tests weekly and six (43%) ran less than five tests weekly. Cost per test varied by 33% in Mozambique and 63% in Tanzania between high and low volume sites, with the greater variability in Tanzania driven by a larger number of low volume sites compared to Mozambique. Labor and overhead costs were small in comparison to consumable and equipment costs in both countries. Download: PNG larger image TIFF original image Fig 3. PoC EID test cost components. (A) Cost estimates expressed in 2020 US$. (B) Proportion of total cost per test. Sites are categorized as high (>12 tests per week), medium (5–12 tests per week), and low (<5 tests per week) testing volume. Equipment includes amortized costs of initial purchase, installation, and yearly maintenance of PoC analyzers. Overhead includes apportioned costs of electricity, communications, and facility upgrades. Overhead and labor costs were < $3 per test (<6% of cost per test) across all sites. https://doi.org/10.1371/journal.pmed.1005069.g003 Cost-effectiveness of very early infant HIV diagnosis VEID was associated with higher costs and improved effectiveness outcomes in both countries (Fig 4, Tables B-C in S1 Text). In Mozambique, VEID increased the proportion of infants initiating ART within one week of life by 90.0 (95% CrI [67.5, 98.5]) percentage points at an incremental cost of $2,632 (95% CrI [$2,249, $3,062]) per infant with HIV. In Tanzania, the incremental cost was higher at $6,263 (95% CrI [$5,394, $7,243]), with a 59.9 (95% CrI [20.9, 89.5]) percentage point increase in the proportion initiating ART within one week. The corresponding ICER with respect to additional infants initiating ART within one week of life was $2,924 in Mozambique and $10,458 in Tanzania. Download: PNG larger image TIFF original image Fig 4. Cost-effectiveness of VEID with respect to PoC EID testing and early ART. Cost-effectiveness planes presenting the incremental costs and (A and B) early ART benefit (% of eligible infants with HIV initiated on ART within 1 week of life; eligible infants include those meeting clinical criteria for treatment, e.g., weight) or 8-week EID benefit (%) for VEID (E and F). Points represent estimates of costs and health benefit for VEID relative to SoC at the origin. A random sample of 500 points is plotted. All estimates indicate that VEID is more expensive and results in increased health benefit (upper right quadrant). (C and D; G and H) Cost-effectiveness acceptability curves for each outcome in A-B and E-F, respectively, showing cumulative probabilities of each testing strategy being cost-effective at a particular willingness to pay value. Costs are presented as 2020 US$. PoC EID, point-of-care early infant (HIV) diagnosis; ART, antiretroviral treatment; SoC = standard of care; VEID, very early infant (HIV) diagnosis. https://doi.org/10.1371/journal.pmed.1005069.g004 Based on the break-even analysis, VEID would need to achieve 15.24 (95% CrI [12.63, 21.37]) life-years gained per additional infant initiating ART within 1 week of life in Mozambique and 32.17 (95% CrI [21.15, 94.84]) in Tanzania to meet empirically derived cost-effectiveness thresholds [22] (Fig 5). At WTP thresholds based on 1x GDP per capita, the corresponding required life-years gained were 6.24 (95% CrI [5.17, 8.74]) in Mozambique and 9.10 (95% CrI [5.98, 26.83]) in Tanzania. Download: PNG larger image TIFF original image Fig 5. Required life-years per early ART initiation for cost-effectiveness. Shaded ribbons for Mozambique and Tanzania indicate 95% credible interval (CrI) across posterior draws from the Bayesian hierarchical model for 1-week ART initiation. Empirically derived cost-effectiveness thresholds, based on health expenditure and life expectancy growth [21] and expressed in 2020 US$, were $189 in Mozambique and $316 in Tanzania. 2020 gross domestic product per capita-based thresholds were $462 in Mozambique and $1,117 in Tanzania. Higher required life-years reflect scenarios in which greater health gains per infant would be needed to approach the indicated thresholds. https://doi.org/10.1371/journal.pmed.1005069.g005 Implementation of VEID also increased uptake of EID. In Mozambique, VEID increased the proportion of infants receiving an EID test within eight weeks of life by 8.0 (95% CrI [5.4, 11.2]) percentage points at an incremental cost of $50.72 (95% CrI [$43.72, $58.65]) per infant exposed to HIV. In Tanzania, the corresponding increase was 20.9 (95% CrI [14.9, 28.0]) percentage points in the proportion of infants receiving an EID test within eight weeks of life at an incremental cost of $28.52 (95% CrI [$24.39, $33.19]) per infant exposed to HIV. The ICER with respect to additional infants receiving an EID test within eight weeks of life was $635.47 and $136.19 in Mozambique and Tanzania, respectively. Sensitivity analysis Cost per test was most sensitive to relative changes in the price of consumables, followed by testing volume and PoC analyzer life span (Fig 6 and Fig C in S1 Text). We observed moderate variability in baseline cost across health facilities resulting from different testing volumes and 5% probability of zero cost across all sites (i.e., infants not receiving testing mostly in SoC sites due to loss to follow-up after birth). Site-level variation was modest with an intra-cluster correlation of 0.033 (95% CrI [0.018, 0.059]) for cost per infant and 0.140 (95% CrI [0.066, 0.258]) for PoC EID within eight weeks. Sites with higher baseline costs tended to experience smaller increases in cost per additional test as testing volumes increased (posterior correlation ρ = −0.25), consistent with economies of scale. Due to low number of infants diagnosed with HIV in Tanzania, uncertainty in effectiveness outcomes was high; combined with uncertainty in costs, this resulted in substantial uncertainty in the ICER with respect to ART initiation in the first week of life in Tanzania. Download: PNG larger image TIFF original image Fig 6. Deterministic one-way sensitivity analysis showing influence of key assumptions on cost per PoC EID test. Costs expressed in 2020 US$. Per test cost difference is shown on the x-axis and parameters with ranges are shown on the y-axis. High and low colors indicate the direction in which the parameter varies. This analysis addresses cost-input uncertainty rather than decision uncertainty; see Figure 4 for probabilistic sensitivity analysis of incremental cost-effectiveness. https://doi.org/10.1371/journal.pmed.1005069.g006 Intrauterine transmission rates observed in the LIFE study were 1.01% (95% CrI [0.65, 1.51]) in SoC sites and 1.48% (95% CrI [1.01, 2.06]) in VEID sites in Mozambique and 0.43% (95% CrI [0.16, 0.85]) in SoC sites and 0.50% (95% CrI [0.19, 0.94]) in VEID sites in Tanzania. The ICER with respect to additional infants initiating ART within one week of life was sensitive to intrauterine transmission rate, ranging from $2,394 to $4,009 in Mozambique and from $5,993 to $27,424 in Tanzania between the 5th and 95th percentiles of estimated intrauterine transmission. In Mozambique, intrauterine transmission would need to exceed 34% to meet the cost-effectiveness threshold from [22] or 14% to meet the GDP-based threshold (Fig 7). In Tanzania, the corresponding intrauterine transmission rates were 20% and 5.8%, respectively, indicating that VEID becomes cost-effective only above specific thresholds of early vertical transmission risk. These thresholds are based on cost per additional infant initiating ART within 1 week of life and do not directly incorporate downstream life-years gained, which are explored separately above. Download: PNG larger image TIFF original image Fig 7. Sensitivity analysis of the incremental cost effectiveness ratio (ICER) per infant with HIV on ART within 1 week across a range of intrauterine transmission rates, based on posterior estimates from the Bayesian hierarchical model. Lines show the posterior mean ICER, with shaded areas representing the 95% credible intervals. Dashed horizontal lines indicate reference willingness to pay (WTP) thresholds: empirically derived cost-effectiveness thresholds based on health expenditure and life expectancy growth from [21], expressed in 2020 US$ ($189 for Mozambique and $316 for Tanzania) and gross domestic product per capita in 2020 ($462 for Mozambique and $1,117 for Tanzania). https://doi.org/10.1371/journal.pmed.1005069.g007 Discussion Focusing on primary healthcare settings in Mozambique and Tanzania, this trial-based economic evaluation estimated the cost and cost-effectiveness of offering PoC HIV diagnostic services at birth and 4–8 weeks of age compared to the SoC at 4–8 weeks only. We evaluated intermediate outcomes, ART initiation within 1 week of life and EID uptake within 8 weeks of life, which are important for reducing early mortality and morbidity. In the context of uncertain global funding for HIV programs, particularly recent funding freezes and proposed cuts to the President’s Emergency Plan for AIDS Relief (PEPFAR), the landscape of HIV epidemiology and vertical transmission may change substantially [25,26]. Amid this uncertainty, generating evidence on the cost and cost-effectiveness of VEID is crucial for guiding investment and design decisions in EID programs, especially as transmission rates and resource availability evolve. VEID at birth increased the proportions of infants tested and initiated on ART and reduced the age at ART initiation. Previous studies also show PoC testing of newborns in similar settings results in earlier ART initiation [27–29], including a multi-country observational study which reported that 92.3% of infants with HIV diagnosed at the PoC initiated ART within 60 days of sample collection [9]. In our study, the majority of infants delivered at VEID sites received a PoC HIV test in the first 24 hours of life and initiated ART within the first week of life. Initiating ART shortly after birth may suppress viral replication and inhibit the establishment of long-lasting viral reservoirs, which could slow disease progression and reduce mortality and morbidity [12,13], as also observed in the LIFE study [15]. From a cost perspective, our findings were consistent with other PoC EID studies using comparable testing platforms and including equipment purchase [30]. Cost per test was driven by reagent costs, apart from at sites with very low testing volumes. These results echo advocacy initiatives, such as the Time for $5 campaign, which emphasize the importance of lowering reagent prices for sustainable access in low- and middle-income countries [31]. We included equipment, overhead, labor, and consumable costs in our calculations, which explains higher cost per test compared to studies omitting equipment procurement and set-up [32,33]. Without initial equipment investment (i.e., in scenarios where existing PoC testing infrastructure allows for repurposing of analyzers for PoC EID or integration with other programs), cost per test can be reduced by up to 43% for mPIMA in Mozambique and 47% for Xpert in Tanzania. However, as manufacturers specify relatively short analyzer lifespans, estimating costs with initial equipment investment for EID remains important for comprehensive budget planning. Further, cross-utilization of PoC infrastructure across programs (e.g., HIV viral load monitoring, tuberculosis diagnosis) may be a viable approach to increase testing volume and decrease costs, especially at low-volume sites. At 70% utilization, cost per test was reduced by up to 4% for mPIMA in Mozambique and 27% for Xpert in Tanzania, highlighting an advantage of the Xpert platform’s broad multiplex diagnostic capability. In both countries, leveraging existing PoC infrastructure and increasing PoC analyzer utilization can roughly halve costs and yield substantial efficiency gains. Previous modelling studies have demonstrated that PoC compared to laboratory-based EID is cost-effective, increases the proportion of infants rapidly initiated on ART, and improves life expectancy across a range of programmatic and economic conditions [32,34]. While we did not assess cost-effectiveness of PoC versus laboratory-based EID, our estimated costs of PoC EID were comparable to costs of laboratory-based EID tests reported in other studies [30]. Instead, we compared PoC birth plus 4–8 week testing with PoC 4–8 week testing alone, consistent with WHO guidelines [3]. This likely underestimated the added impact of VEID, as SoC also benefited from shorter PoC turnaround times compared with central laboratory-based testing in our study. A model-based analysis of birth plus 6-week testing in South Africa improved survival and was deemed cost-effective as long as uptake was high [16]. Our study adds real-world experience related to test frequency and timing from two countries with different intrauterine transmission rates using different PoC testing platforms. Although early ART initiation is an intermediate outcome, it is an important precursor to survival benefit [14,35]. The survival benefits required for VEID to be considered cost-effective based on the break-even analysis, 15–32 life-years gained per additional infant initiating ART within 1 week using cost-effectiveness thresholds derived from Pichon Riviere and colleagues [22] or 6–9 life-years gained using GDP-based thresholds, are substantially higher than those demonstrated to date. In the CHER trial [14], which remains the primary benchmark, early versus deferred ART was associated with roughly 0.8 life-years gained over 5 years, based on reported deaths and follow-up time per arm. However, CHER compared 6-week to median 7-month ART initiation, while our study compared birth to 4–8-week ART initiation. Importantly, HIV-related infant mortality peaks around 2–3 months [5], suggesting that ART at birth could be underestimated. While early mortality reductions in the LIFE study were not sustained through 18 months, most infants received LPV/r and viral suppression was low and associated with poor clinical outcomes overall [15]. Differences in treatment regimens, adherence patterns, background mortality, and health system context across settings and time introduce structural uncertainty that may influence the magnitude of survival benefits achievable with ART initiation in the first week of life. With dolutegravir-containing ART now available for infants and expected to address common barriers to adherence due to once daily dosing and improved palatability, greater longer-term health benefits may be achievable than observed in earlier cohorts. Differences in comparator, follow-up, and treatment context caution against directly extrapolating CHER outcomes to our setting and highlight that long-term benefits of VEID are still uncertain. The life-years estimates above are a derived benchmark intended to contextualize the magnitude of survival benefit required for cost-effectiveness. The cost-effectiveness thresholds used are intended to approximate health system opportunity costs and should be considered alongside local budget constraints, feasibility, and equity priorities when informing policy decisions. Beyond survival, transmission dynamics were central to the cost-effectiveness of VEID. Mozambique and Tanzania observed average intrauterine transmission rates of 1.3% and 0.5%, respectively, and the incremental cost per additional infant initiating ART within one week of life was sensitive to this. Our analysis indicates that intrauterine transmission would need to exceed 30% in Mozambique and 20% in Tanzania using an empirically derived cost-effectiveness threshold, or 14% and 5.8% using a GDP-based threshold, for universal VEID to be cost-effective. Settings with low intrauterine transmission rates may alternatively consider targeted VEID based on maternal risk factors (e.g., lack of ART, high viral load) [36], combined with efforts to maintain PoC analyzer utilization, as a more feasible path to affordability in low- and middle-income countries. In settings with higher intrauterine transmission or where survival benefits of VEID exceed current evidence, universal VEID may be more justifiable. Where and to whom VEID is offered may also have important equity implications. The benefits of VEID are likely concentrated among infants at higher risk of vertical transmission and in health facilities with sufficient capacity to reliably deliver timely testing and treatment. In settings with lower testing volumes or weaker follow-up systems, costs and benefits may differ, underscoring the importance of pairing VEID implementation with health system strengthening so that it reduces disparities in infant HIV testing and linkage to care. Retention in care, adherence to treatment, and ensuring follow-up HIV testing for infants while they remain at risk for acquiring HIV, which depend on strong health systems and support services, are other important considerations when evaluating the cost-effectiveness of PoC EID approaches. Very early initiation of ART can only be translated into longer-term health benefits if infants continue receiving effective treatment. Thus, supportive interventions addressing other challenges such as inadequate family support, lack of disclosure, and postpartum maternal depression are also needed [37]. Additionally, all infants should receive follow-up testing to identify those with undetectable virus at birth or acquiring HIV during the breastfeeding period [3,16], especially considering potential disruptions to maternal ART access due to funding cuts. We assumed that infants lost to follow-up did not incur further program-related costs. If unobserved services were received elsewhere, this assumption could bias costs downward. However, loss to follow-up was low (<5%) in our study. While full population coverage and universal follow-up testing would slightly increase program costs, our estimates reflect what a well-run program could realistically achieve. A strength of this study is that it is primarily informed by a trial combining PoC and birth EID designed to mirror routine care and align with country-specific HIV program priorities. We provided contextualized information about resources required to scale-up EID programs, about which little was known [16]. The limitations include that uncertainty estimates for many cost inputs were lacking. LIFE study recruitment spanned a time period when prices and health-seeking behavior may have been influenced by the COVID-19 pandemic. We mitigated price fluctuations by relying mostly on pre-pandemic costs and did not observe a notable decline in 4–8-week visit attendance during lockdown periods. Costs related to confirmatory central laboratory testing and broader shared infrastructure were not included, which may underestimate total programmatic costs, especially at low-volume sites. Our probabilistic sensitivity analysis propagates uncertainty arising from observed variation in costs and effectiveness outcomes captured by the Bayesian hierarchical models. We did not additionally model uncertainty in individual cost inputs, as the analysis was based on observed resource use and expenditures. In addition, scarce long-term health data for infants with HIV—particularly survival times or quality-of-life measures—make assessing the efficiency of interventions challenging. Given the absence of a significant difference in longer-term health outcomes in the LIFE study [15], we restricted our analysis to a 12-week time horizon and assessed cost-effectiveness using intermediate outcomes relevant to EID programs as key indicators of health benefit. We referenced empirically derived cost-effectiveness thresholds as well as GDP per capita [21,22,38], however, these thresholds, expressed in terms of life-years gained rather than additional early ART initiations, should be used to contextualize results and not to directly interpret ICERs. Our break-even analysis provided an informative benchmark of the survival benefits required for cost-effectiveness, but these estimates were derived from posterior draws and depend on assumptions about WTP. This approach excludes explicit examination of cost-utility in terms of quality- or disability-adjusted life years. Dynamic, longer-term modeling frameworks could more explicitly account for longitudinal clinical transitions and age-dependent mortality, though such approaches would require substantial extrapolation beyond the observed data. Finally, direct client and broader societal costs (e.g., schooling gains) were beyond the scope of this work. From a healthcare system perspective, universally offering PoC EID at birth to infants exposed to HIV was more expensive and resulted in more frequent and earlier ART initiation. Early ART initiation could help reduce persistently high HIV-related mortality and morbidity among infants. Our analysis suggests that VEID would need to generate substantial long-term health benefits to meet standard cost-effectiveness thresholds at current costs. The limited available evidence linking early ART initiation to longer-term health benefits suggests that these benefits may be smaller than required for VEID to be cost-effective when offered universally. The cost-effectiveness of VEID increases if ART initiation at birth translates to long-term survival benefits or if VEID can be effectively targeted to infants at high risk of HIV acquisition. In settings with high vertical transmission rates or where weak HIV programs lead to significant gaps in infant retention in care, EID programs could consider adding PoC birth testing with careful consideration of PoC infrastructure utilization to improve affordability. In settings with low vertical transmission rates, the additional benefit of universal birth testing is limited and targeted testing of high-risk infants at birth may be a more practical approach. Supporting information S1 Text. Supplementary material. Fig A: Violin plots of per site time from birth to PoC HIV EID testing and treatment initiation. (A) First EID test among all HIV-exposed infants; (B) ART initiation among infants diagnosed with HIV up to 16 weeks of age. Only sites with infants with HIV are shown. Black triangle markers represent individual data. VEID, very early infant (HIV) diagnosis; SoC = standard of care; EID, early infant (HIV) diagnosis; ART, antiretroviral treatment. Created in BioRender. Hoelscher, M. (2025) https://BioRender.com/u260ruo. Fig B: Per site point-of-care early infant (HIV) diagnosis test cost components. (A) Cost estimates expressed in 2020 US$. (B) Proportion of total cost per test. Fig C: Per site deterministic one-way sensitivity analysis showing influence of key assumptions on cost per test. Per test cost difference is shown on the x-axis and parameters with ranges evaluated are shown on the y-axis. High and low colors indicate the direction in which the parameter varies. Table A: Point-of-care early infant (HIV) diagnosis test cost components. Cost estimates are expressed in 2020 US$. Sites are categorized as high (>12 tests per week), medium (5–12 tests per week), and low (<5 tests per week) testing volume. Equipment includes amortized costs of initial purchase, installation, and yearly maintenance of PoC analyzers. Overhead includes apportioned costs of electricity, communications, and facility upgrades. Table B: Modeled incremental cost and effect estimates for 1-week ART uptake per infant diagnosed with HIV. Estimated means and 95% CrI shown. SoC = standard of care; VEID, very early infant (HIV) diagnosis; CrI = credible interval. Table C: Modeled incremental cost and effect estimates for 8-week PoC EID uptake per infant exposed to HIV. Estimated means and 95% CrI shown. SoC = standard of care; VEID, very early infant (HIV) diagnosis; CrI = credible interval. Fig D: Histogram of median PoC tests per ISO week with divisions for low-, medium-, and high-volume sites shown by dashed lines. PoC = point-of-care; ISO, International Organization for Standardization. Fig E: Model parameter traces for A) 1-week ART uptake model and B) 8-week PoC EID uptake model. Burn-in phase not included. Parameters defined in 1. ART, antiretroviral treatment; PoC EID, point-of-care early infant (HIV) diagnosis. Fig F: Posterior predictive plots for A) Cost, B) 1-week ART uptake and C) 8-week PoC EID uptake. y refers to the observed data, yrep to the simulated data from the posterior predictions. ART, antiretroviral treatment; PoC EID, point-of-care early infant (HIV) diagnosis. Fig G: Posterior distributions and scatter plots for 1-week ART uptake model. ART, antiretroviral treatment. Fig H Posterior distributions and scatter plots for 8-week PoC EID uptake model. PoC EID, point-of-care early infant (HIV) diagnosis. Fig I: Distributions of intrauterine transmission probabilities (defined as probability of positive HIV test result at birth) generated from 10,000 Bayesian bootstrap samples per country and study arm. SoC = standard of care; VEID, very early infant (HIV) diagnosis. Table D: Summary of model parameters and posterior estimates. Table E: Posterior estimates for primary informative directional and sensitivity weak neutral prior specifications. Parameter definitions can be found in Table D. https://doi.org/10.1371/journal.pmed.1005069.s001 (PDF) S1 Checklist. CHEERS 2022 Checklist. Husereau D, Drummond M, Augustovski F, de Bekker-Grob E, Briggs AH, Carswell C, Caulley L, Chaiyakunapruk N, Greenberg D, Loder E, Mauskopf J, Mullins CD, Petrou S, Pwu RF, Staniszewska S; CHEERS 2022 ISPOR Good Research Practices Task Force. Consolidated Health Economic Evaluation Reporting Standards 2022 (CHEERS 2022) statement: updated reporting guidance for health economic evaluations. BMJ. 2022 Jan 11;376:e067975. doi: https://doi.org/10.1136/bmj-2021-067975. PMID: 3550177145; PMCID: PMC8749494. This checklist is licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0). https://doi.org/10.1371/journal.pmed.1005069.s002 (DOCX) Acknowledgments The authors gratefully acknowledge the infants and families who participated in the study and the dedicated healthcare personnel at the study sites. They also thank Martina Penazzato and Lara Vojnov from the World Health Organization, Landon Myer and Lynn Horn from the University of Cape Town, and Karim Manji from the Muhimbili University of Health and Allied Sciences in Dar es Salaam for external expert advice. The study was conducted at health facilities within the Mozambican and Tanzanian National HIV programs. The Clinton Health Access Initiative (CHAI) provided infrastructural support at the study sites and contingency supplies for pediatric antiretroviral drugs. We thank Lise Ellyin and Helder Mendes from CHAI Mozambique, and Patricia Mbago and Esther Mtumbuka from CHAI Tanzania for their support. The authors also acknowledge the LIFE Study Consortium: Araújo Patricio, Dadirai Mutsaka, Lara Samuel, Sergey Bocharnikov, Timothy Bollinger, Wilson Simbine, Abhishek Bakuli, Cornelia Lueer, Elmar Saathoff, Fidelina Zekoll, Mariana Mueller, Friedrich Rieß, Otto Geisenberger, Nuno Taveira, Rute Marcelino, Absalao Zumba, Daniel Machavae, Adolfo Vubil, Ana Duajá, Jacinto Adolfo Ndarissone, Joao Manuel, Maria Maviga, Nalia Ismael, Jorge Morais, Nedio Mabunda, Adolfo Vubil, Fatima Mecupa, Amina de Sousa, Abisai Kisinda, Chacha Mangu, Doreen Pamba, Festina Paschal, Hellen Mahiga, Janeth Stephen, Lilian Njovu, Margareth Haule, Oliver Lyoba, Theodora Mbunda, Nhamo Chiwerengo, and Willyhelmina Olomi.
Availability, appeal, and addictiveness by design: Tobacco and nicotine industry deliberate targeting of youthMaddox, Raglan;Freeman, Becky;Pisinger, Charlotta;Banks, Emily
doi: 10.1371/journal.pmed.1005133pmid: 42213812
Introduction The 2026 World No Tobacco Day theme is “unmasking the appeal”, raising awareness of tobacco and nicotine product marketing and design, including features that increasingly appeal to adolescents. This is especially the case for e-cigarettes/vapes; vaping prevalence at ages 13–15 is on average nine times that of adults (Fig 1), and nicotine dependence among young people is increasing [1,2]. Download: PNG larger image TIFF original image Fig 1. Average global prevalence of current e-cigarette use among children aged 13–15 years and adults, from countries with available data. Data adapted from WHO global estimates [1]. https://doi.org/10.1371/journal.pmed.1005133.g001 Tackling appeal is therefore imperative to reversing this trend. This means appreciating that appeal goes beyond marketing, and is deliberately built into product supply and design, market environments, and regulatory conditions that serve the structural production and sustainability aims of the tobacco and nicotine industry. Tobacco and nicotine products intentionally attract attention, reduce aversion, and sustain use, particularly among young people who represent the next generation of consumers as smoking rates decline [3]. The commercial tobacco and nicotine industry generates profit through the creation and maintenance of addiction, and lifetime nicotine addiction largely occurs when use starts while the brain is developing, in childhood and adolescence. Youth uptake must therefore be understood not as an unintended outcome, but as a fundamental industry aim, necessary for industry profits and survival, and a predictable outcome of systems that maximize availability, appeal, and addictiveness. Understanding these key drivers is essential for effective regulations, policies, and programs to reduce tobacco use and nicotine addiction and protect the health of young people. This is especially important given the growing evidence regarding the adverse health impacts of e-cigarettes, particularly for youth, including addiction, toxicity through inhalation, poisoning, burns and injuries, lung injury, increased smoking uptake, and concerning findings relating to cardiovascular, respiratory, and carcinogenic effects [2,4]. Furthermore, the long-term impacts of e-cigarettes on many important clinical conditions remain uncertain [2]. The evolution of the industry narrative Industry messaging about the purported benefits of vaping and other nicotine products has shifted from individual smoking cessation-oriented claims toward broader and often unsubstantiated narratives of population benefit that position youth uptake as incidental, unavoidable, or even protective against future smoking [5,6]. This framing relocates responsibility from industry practices onto individual behavior while preserving market legitimacy and obscuring the structural and commercial drivers of nicotine dependence [5,6]. Consistent with longstanding industry strategies to manufacture doubt and delay regulation, contemporary “harm reduction” narratives emphasize uncertainty, personal choice, and the inevitability of nicotine use, while minimizing the role of product design, marketing, and commercial systems in producing addiction [5,6]. Importantly, this framing also reproduces colonial patterns of reasoning, recasting structurally produced harms as matters of individual responsibility, which legitimizes continued expansion into younger populations [5,6]. Industry and investor communications continue to frame younger generations as critical to the long-term sustainability of nicotine markets, reinforcing the commercial importance of ongoing youth uptake and dependence. Why youth uptake is predictable E-cigarette appeal arises from interacting features across product design, market environments, and regulatory conditions. These features operate synergistically, creating conditions in which initiation, dependence, and sustained use are likely [6]. At the product level, flavors [7], sensory appeal [8], device aesthetics, nicotine salts, digital integration, and misleading labeling all increase product palatability and attractiveness among young people. For example, nicotine salts allow smoother inhalation and higher nicotine delivery, accelerating dependence, particularly among people who have not previously used nicotine [8]. Devices increasingly incorporate social and digital features, embedding nicotine use within everyday routines and online environments. These features also increase visibility, peer diffusion, and social normalization among young people. Higher nicotine concentration product use has increased substantially over time: from 2017 to 2022, the share of products containing ≥5% nicotine strength increased by 1,486% [9]. Products are marketed as long-lasting and continuously available for use, and use patterns often include ubiquitous “grazing”: repeatedly inhaling small series of puffs throughout the day. Rapid product innovation and designs intended to facilitate frequent and sustained nicotine exposure reinforce these patterns [8,9]. Products are also designed to be easy to use and integrated throughout normal daily routines and activities [8,9]. Product designs, affordability, widespread availability, and equivalence framings (e.g., puff-to-cigarette comparisons) normalize experimentation, uptake, consumption, and sustained nicotine use among young people. This pattern of continuous product change reflects a deliberate strategy to sustain use and dependence [6,9]. Misleading descriptors such as “tobacco-free” further obscure risk. Promotion occurs within highly permissive digital and retail environments. Social media marketing, influencer partnerships, and youth-oriented branding increase exposure and normalize use at scale [10], often beyond the reach of existing regulation. These dynamics are enabled by regulatory gaps. Products frequently enter markets without robust pre-market evidence, standards remain inconsistent, and enforcement is often weak. These conditions also contribute to unregulated and illicit markets, which are a predictable consequence of oversupply, weak regulation, and profit-driven distribution [11]. Taken together, these conditions make youth uptake of products an entirely foreseeable and system-driven outcome. Importantly, exposure to appealing product environments is not evenly distributed. Youth and communities experiencing structural disadvantage are more heavily targeted and less protected by regulation. This includes Indigenous peoples, racialized populations, LGBTQA+ communities, and those facing socioeconomic disadvantage, reflecting patterns shaped by colonization, racism, and commercial targeting [12]. Describing these groups as “vulnerable” risks obscuring the structural drivers of harm [13]. The issue is not inherent susceptibility, but a failure of systems to provide adequate protection, allowing predatory industries to exploit these gaps [12]. Policy implications: Regulating design and supply, not just marketing Product design and other factors contributing to appeal should be recognized as key determinants of health, not as neutral features. If appeal is engineered, policy must address the systems that produce and sustain it. This requires adopting comprehensive tobacco advertising, promotion, and sponsorship bans, as well as regulating product design, supply, and the broader commercial environment [12,14]. Priorities include: pre-market product standards, where the tobacco industry is legally required to prove product safety, as in Norway; restrictions on flavors and design features (including bans); stronger control of retail and digital environments (e.g., strict tobacco retailer licensing systems, limiting retailer density, strong enforcement, and substantial penalties for violations); and full implementation of WHO Framework Convention on Tobacco Control, including high tobacco taxes, comprehensive advertising and sponsorship bans, 100% smoke-free public spaces, large graphic health warnings, restrictions on youth access, and measures to combat illicit trade [14]. This also includes Article 5.3 protections against industry interference, which recognize the fundamental conflict between public health and tobacco industry interests and require governments to protect health policy from the commercial and other vested interests of the tobacco industry through transparency, limiting interactions to those strictly necessary for regulation, and rejecting partnerships or voluntary agreements with tobacco and nicotine companies [14]. A precautionary approach is needed, placing the burden of proof of product safety on industry rather than on populations exposed to harms. Conclusions Unmasking appeal requires recognizing that attractiveness is not incidental but a deliberate strategy within commercial and regulatory systems that enable this harm. Youth uptake is not a failure of individual choice or a lack of awareness of harms, but a predictable outcome of an industry structured and incentivised to generate and sustain addiction. Effective protection depends on governing commercial practices, and more fundamentally, on addressing the systems that allow harmful products to be created, promoted, and made widely available [6,11,12,14]. This shifts policy from managing risk to addressing its source: the commercial tobacco and nicotine industries.
Novel symptoms associated with eclampsia could improve detection and save livesBeardmore-Gray, Alice;Shennan, Andrew
doi: 10.1371/journal.pmed.1005123pmid: 42166530
A woman dies every 2 min during pregnancy or childbirth [1]. Hypertensive disorders of pregnancy, including pre-eclampsia, are the leading direct cause after post-partum hemorrhage. Of these deaths, 95% occur in low- and middle-income countries (LMICs), and the regions of sub-Saharan Africa and South Asia account for 87% of all global maternal deaths [1]. This stark inequality is projected to deepen over the next decade, as funding scarcity, conflict, climate change, and a shifting geopolitical landscape threaten hard-won gains in women’s health [2]. Pre-eclampsia, a multi-system endothelial disorder caused by placental dysfunction, is progressive, and the onset of severe features is difficult to predict [3]. This is especially true in low-resource contexts where limited access to laboratory testing and obstetric ultrasound is further compounded by health workforce shortages, resulting in a lack of adequate surveillance for women at high-risk of complications and poor quality care [4]. Eclampsia is a manifestation of severe pre-eclampsia with an associated mortality rate of up to 7% [5]. It is frequently associated with additional complications including intracerebral hemorrhage, pulmonary oedema, placental abruption, and stillbirth. Current management includes administration of magnesium sulfate, which more than halves the risk of eclampsia and is indicated for both primary prevention and recurrence. It should be given to all women who have severe pre-eclampsia or eclampsia, whether antenatal or postnatal, and is recognized by the World Health Organization and United Nations to be a priority drug. While its exact mechanism of action is unknown, it is thought to prevent and control seizures by stabilizing neuronal networks and reducing cerebral vasospasm and ischemia. Understanding which women with pre-eclampsia are most at risk of progressing to eclampsia and other severe complications is critical to informing women’s and clinicians’ decision-making and guiding clinical management. While stabilizing blood pressure is important, delivery is currently the only intervention shown to reduce the risk of adverse maternal outcomes [6] and fetal death [7], and should be recommended immediately if severe features, including eclampsia, develop [8]. Accurately triaging women to identify those most at risk, in order to optimize timing of birth, with consideration of antenatal corticosteroids if delivery is planned prior to 34 weeks’ gestation, remains a critical challenge. Although there have been advances in the prediction and diagnosis of pre-eclampsia, there are still no prognostic tools that reliably predict the onset of complications once pre-eclampsia has been diagnosed. Moreover, clinical prediction tools that rely on laboratory tests, sonographer expertise, and biomarkers may be challenging and costly to implement in already fragile healthcare systems within LMICs. Clinical decision-making around ongoing management is therefore often based on the presence or absence of poorly defined “severe maternal symptoms” [8,9], which have limited prediction. This is confounded by clinical signs also being unreliable: A 2018 study demonstrated that high blood pressure was not significantly associated with eclampsia in an LMIC country (South Africa) [10], yet it is often used to define the severity of pre-eclampsia, particularly in low-resource settings. A woman with severe pre-eclampsia may need to be transferred from remote and rural settings in LMICs without diagnostic tools to a higher-level care facility in order to access life-saving interventions [11]. Thus, better prognostic tools are needed. The process of history taking has been described as “the most powerful and sensitive and most versatile instrument available to the physician” and has been reported to provide 60%–80% of the information that is relevant for a diagnosis [12]. Midwives and obstetricians have traditionally been taught to ask women with pre-eclampsia if they have a headache, visual changes, and/or epigastric pain as part of their clinical history and routine symptoms enquiry. However, these previously identified prodromal symptoms, based on a small number of methodologically limited studies, have shown only modest associations with eclampsia and do not accurately predict its onset, nor reliably rule it out if absent [13]. In a recent PLOS Medicine study [14], Hastie and colleagues have identified 10 novel prodromal symptoms exhibiting far stronger associations with eclampsia than the traditionally associated symptoms of headache, visual changes, and epigastric pain (Fig 1). This case-control study prospectively enrolled women over a 5-year period, with eclampsia, pre-eclampsia or normotensive pregnancies in South Africa and Pakistan and asked whether they experienced 20 neurological symptoms. For those women who experienced eclampsia (n = 341), this was within 7 days of the seizure. They compared the likelihood of symptoms occurring before eclampsia, compared to being present with pre-eclampsia. Ten symptoms had odds ratios (ORs) over 10 for eclampsia (whereas ORs for existing symptoms are all lower than 10: headache (OR 2.26), visual changes (OR 5.73), and epigastric pain (OR 2.25)). The newly identified prodromal symptoms are summarized in Fig 1. Download: PNG larger image TIFF original image Fig 1. Ten novel prodromal symptoms of eclampsia. Novel symptoms (as identified in [14]) and their odds ratios (ORs) include: twitching/jerking limbs (OR 42.03), affected hearing (OR 33.12), altered mind state (OR 33.60), impaired speech (OR 33.12), feelings of doom (OR 23.71), severe vertigo (OR 26.59), confusion (20.52), jitters (OR 18.16), difficulty concentrating (OR 15.81), and weakness/paralysis (OR 10.49). https://doi.org/10.1371/journal.pmed.1005123.g001 Notably, only 2.4% of women who experienced eclampsia did not report any of the full list of 20 prodromal symptoms screened for in this study, highlighting the importance of continually re-assessing women’s symptoms. Furthermore, those who had eclampsia were younger, had a lower Body Mass Index, were more likely to be nulliparous, book later for antenatal care, or have not received any antenatal care. This represents a high-risk group of women who require heightened surveillance and pro-active inclusion within healthcare systems. This study is strengthened by its sites in two countries which represent regions of the world with a high burden of pre-eclampsia and maternal death, thus increasing the generalizability of their findings across different populations and healthcare systems. Eclampsia, while serious, is a relatively rare complication, and therefore to have evaluated a cohort of 341 women who experienced eclampsia is impressive and increases the robustness of their findings. Expert neurology input into this study also highlights the importance of integrating obstetric and medical teams as part of providing holistic care and developing a broad understanding of this multi-system disorder, which also increases future risk of cardiovascular and cerebrovascular disease. While it is possible that women affected by eclampsia may have recalled their symptoms differently, the study investigators were able to mitigate the risk of recall bias by ensuring that the entire symptoms list was screened for within all groups, including women with pre-eclampsia and normotensive pregnancies. Moreover, given the rare occurrence of eclampsia, a prospective screening study would likely not have been feasible. The proportion of women who did not receive antenatal care differed between the two study sites (6.6% in South Africa versus 43.4% in Pakistan). Future work should focus on evaluating these symptoms in an even larger cohort of women across multiple different care settings, to establish whether screening for these prodromal symptoms is similar in all settings. Correlation of symptom severity with blood pressure and other markers of pre-eclampsia severity (biochemistry, ultrasound findings, placental growth factor levels) could also add further information to a future predictive model. While acknowledging that this was not intended to be a diagnostic prediction study, and that further prospective validation of these symptoms’ predictive ability is warranted, maternity care providers should implement screening for these 10 symptoms identified as being strongly associated with eclampsia into clinical protocols. It is important that educators and policymakers take these findings into account when updating training curriculums and guidelines, and that further research evaluates effective dissemination and implementation of such materials. Women’s health remains neglected, under-funded, and under-prioritized, and this contributes to poor reproductive health outcomes for women. Women spend nearly 25% more of their lives in poor health than men, yet less than 5% of global health research funding is directed toward women-specific conditions beyond oncology. This study therefore offers a timely and relevant contribution to the existing evidence, questioning previous practice and expanding our understanding of a condition which remains a leading cause of maternal death globally. The identification of these 10 prodromal symptoms, which may predict the onset of eclampsia, offers real-world clinical utility and has the potential to offer a simple, low-cost screening tool that could prompt earlier intervention and thus reduce maternal and perinatal mortality. In an era increasingly dominated by high-tech solutions—artificial intelligence, digital health, and precision medicine—this study reminds us that the most transformative interventions are often the simplest. A better history, grounded in robust evidence, costs nothing and can be implemented anywhere. We look forward to seeing how these findings will be taken forwards and ultimately incorporated into routine clinical practice.
Adherence to voluntary UK sugar, salt, and calorie reduction targets in the highest-grossing restaurant chains: A cross-sectional studyO’Hagan, Alice;Pechey, Rachel;Forde, Hannah;Bandy, Lauren
doi: 10.1371/journal.pmed.1004681pmid: 42085343
Background To address high rates of diet-related disease, the UK Government has a series of voluntary targets for retailers, manufacturers, and the out-of-home sector (e.g., restaurants), to reduce the sugar, salt, and calorie content of food products. The sugar targets were intended to be met in 2020, the salt targets in 2024, and the calorie targets in 2025 (extended from 2024 due to Covid-19). There is limited evidence for how the out-of-home sector is performing against these targets, and individual company responses have not been evaluated. This study aimed to assess adherence to UK Government’s sugar, salt, and calorie reduction targets for menu items offered by the 21 highest-grossing restaurant chains in 2024. Methods and findings Nutritional information was collected from restaurants’ online menus. Mean/median sugar, salt, and calorie content, per 100 g and per serving, was calculated for each restaurant and food subcategory. Sugar, salt, and calorie content for each menu item was compared against the UK Government’s targets, and the proportion of menu items meeting (i) each and (ii) every applicable target, was calculated for each restaurant and food subcategory. Three thousand ninety-nine menu items were included. Across all restaurants, 61% of menu items met their calorie targets, 58% met their salt targets, 36% met their sugar targets, and 43% met all of their applicable targets. Six of the 12 food subcategories, and nine of the 21 restaurants, had over 50% of menu items meeting all of their applicable targets. Menu items from Papa John’s were the lowest adhering for the calorie (35%) and salt (8%) targets, and menu items from Burger King, KFC, Nando’s, and Vintage Inns were the lowest adhering for the sugar targets (0%). Menu items from pizza restaurants had the lowest adherence to all applicable targets (32% overall) out of all the restaurant types, but items offered by restaurants with similar menu foci were also found to vary in their adherence. We were unable to account for heterogeneity in item-level sales due to the lack of accessible sales data from the out-of-home food sector, and therefore we could only assess performance against the targets for available items as opposed to purchased items. Conclusions Our findings suggest that while menu items from certain restaurant types appear to perform worse than others against the sugar, salt, and calorie targets, items from restaurants with similar menu portfolios also vary in their adherence, highlighting the potential for restaurants to improve the nutritional quality of their products without changing their menu focus. Our study demonstrates that there is low adherence to voluntary schemes across the out-of-home sector, and therefore mandatory regulations may be a more effective approach to improving the nutritional quality of out-of-home food. Introduction The purchasing and consumption of foods high in energy, saturated fat, free sugars, and salt, is associated with an increased risk of obesity and diet-related non-communicable diseases (NCDs) [1]. Globally, diet-related NCD prevalence is high, with approximately 40% of cardiovascular disease mortality between 1990 and 2019 attributable to dietary risk factors [2]. The same is true for the UK specifically, with poor diet being a leading cause of death and ill health [3]. Food reformulation initiatives are a common approach to NCD prevention across the globe; 68 countries have a salt reduction policy in place [4], with countries such as the UK, Brazil, New Zealand, and USA, having further initiatives for sugar reduction [5], with the majority of such schemes being voluntary [4,5]. Modelling studies highlight the potential for such policies to significantly reduce the incidence of diet-related diseases, such as obesity and cardiovascular disease [6,7]. The UK Government first introduced their salt reduction programme in 2004, following advice from the Scientific Advisory Committee on Nutrition to reduce the population average salt intake by approximately 10%, to 6 g a day. The first reduction and reformulation targets for industry were published in 2006 and have since been revised on several occasions, with the most recent revision in 2020 (for targets to be met by 2024) covering 84 category-specific salt targets for grocery foods (e.g., Ready Meals), and 24 for the out-of-home (OOH) sector (e.g., Seasoned Fries) [8]. Evidence suggests that the salt reduction programme has been successful, with average urinary sodium reducing by approximately 2% each year from 2004 to 2011, and significant reductions being observed in the salt content of food products sold by UK retailers [9]. Complementary to the salt reduction programme, the UK Government published targets for sugar and calories in 2017 and 2020, respectively [10,11]. The sugar reduction programme aimed to achieve a 20% reduction in the sugar content of 14 food categories that contribute most to sugar intake in children, across all sectors of the food industry, by 2020 [10]. The calorie reduction programme aimed to achieve a 10% reduction by retailers and manufacturers across nine food categories, and a 20% reduction by the OOH sector across seven food categories (with four common categories), in the calorie content of products sold, by 2025 (extended from 2024 due to the Covid-19 pandemic) [11]. Governmental progress reports indicate that within the retailer and manufacturer sector, there has been a 3.5% reduction in the sugar content of products sold between 2015 and 2020 [12], and category-level decreases of up to 2.4% in the calorie content of products sold between 2017 and 2021 [13]. Their findings suggest comparatively less progress has been made in the OOH sector, with only a 0.2% reduction in the sugar content of products sold between 2017 and 2020, and category-level increases of up to 2.3% in the calorie content of products sold between 2017 and 2021 [12,13]. However, since the most recent data included in the sugar, salt, and calorie reduction progress reports was from 2020, 2018, and 2021 respectively, these conclusions may have since changed. A considerable proportion of UK individuals’ weekly food intake comes from takeaway or restaurant meals [14]. The UK Diet and Nutrition Survey finds 72% of respondents reported purchasing food or drink from an OOH establishment in the past week, with the largest proportion of these consumers being children aged 11–18 years old [15]. Eating OOH is becoming more popular, with a 159% increase in per-person expenditure on eating out from 2021 to 2022 in the UK [16], and an additional 3.5% increase to 2023 [17], potentially due to its perceived lower cost and higher convenience than preparing food at home [18]. The OOH sector is dominated by a small number of multinational companies who operate chained food outlets (hereafter ‘restaurants’; see Table 1). In 2023, sales from chained food service outlets reached £35.1 billion in the UK, increasing by nearly £8 billion over the previous two years [19], showcasing the significant influence these companies have over diet and health. Download: PNG larger image TIFF original image Table 1. A summary of characteristics for each restaurant, including their affiliated company, type, 2022 sales value, number of menu items included, and number of subcategories included. Restaurants presented in descending order by 2022 Sales Value. https://doi.org/10.1371/journal.pmed.1004681.t001 The UK Government’s sugar, salt, and calorie reduction targets are voluntary, leaving the responsibility for meeting the targets to individual companies, yet there is limited evidence for how food companies have responded to the reduction programmes. One study looking at company-level adherence to the sugar reduction targets among UK manufacturers found that of the top 10 companies across 5 target food categories, just under half met the interim 5% sugar reduction targets for 2018 [20]. Similar work conducted in the OOH sector found that only four out of 48 brands reduced the sugar content of their desserts by at least 20% from 2018 to 2020, and only half of which significantly reduced their calorie content as well [21]. There is a need for a more comprehensive assessment of company-level adherence to the targets. Furthermore, fewer studies in the broader literature have assessed the nutritional quality of foods in the OOH sector compared to the retailer and manufacturer sector, due to OOH nutrition data being less easily accessible. Evaluating industry’s response to the reduction programmes will improve transparency around companies’ commitment to population health, help with enforcement of comparable schemes in the future, and underpin other approaches to incentivising healthier food provision (e.g., investment decision-making based on food healthiness). The introduction of mandatory reporting of healthy sales from large companies, as pledged in the NHS 10 Year Plan [22], will contribute to improving transparency and potentially incite greater commitment from companies to meet nutritional guidelines. This study aimed to assess the nutritional content and adherence to the UK Government’s sugar, salt, and calorie reduction targets, for food items offered by the highest-grossing restaurant chains in the UK, in 2024. Methods We assessed the nutritional content of food items offered in 2024 by the 21 highest-grossing restaurant chains in the UK. We collected nutrition information for restaurants’ menu items directly from their websites. Each menu item was categorised by the sugar, salt and calorie reduction target groups and their nutrient values compared to the relevant targets. The average kcal, salt, sugar per 100 g and per serving, was also calculated for each restaurant and food subcategory. The study protocol was published on Open Science Framework prior to data analysis [23] (S1 File). This study is reported as per the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guideline (S1 Table—STROBE checklist). Ethical approval was not required for this study as it does not involve human participants. Identifying restaurants This study focuses on the highest-grossing restaurant chains in the UK, as these restaurants have the largest impact on diets within the OOH food sector. We identified the 21 highest-grossing chained restaurants in the UK using the latest (2022) sales data from Euromonitor International, a privately-owned market research company. Euromonitor’s data portal Passport GMID was accessed via the Bodleian Library, University of Oxford [24]. We selected restaurants that Euromonitor defined as “Consumer Foodservice”, including “cafés/bars, full-service restaurants, limited-service restaurants, self-service cafeterias and street stalls/kiosks”. “Full-service” refers to sit-down restaurants with table service, while “limited-service” refers to fast food or take-away outlets. We excluded restaurants that had separate menus for each individual location to avoid introducing considerable heterogeneity to nutritional information at the restaurant level. Restaurants that only provided nutrition information per ingredient (e.g., bun, patty, cheese) rather than per item (e.g., complete burger) were also excluded, as it was unclear what a full menu item would consist of. We initially aimed for a sample of 20 restaurants. Subway was excluded at the start of the data collection phase as it provided nutrition information for individual ingredients only. However, new menu-item nutrition information was subsequently published after our data collection had started, and therefore Subway was reinstated and a sample of 21 restaurants’ data was collected. Data collection Menu data was collected directly from restaurants’ UK websites in February and March of 2024 (May 2024 for Subway). It was collected primarily through PDF menu documents or menu webpages. Where a restaurant provided multiple PDFs, data was only collected from those labelled as containing ‘core’ (or equivalent) menu items. Limited-time offer menu items (e.g., seasonal) were only included in data collection if they were present in the ‘core’ menus. For pizza restaurants that offered multiple crust options, only nutrition information from the default option (e.g., classic crust) was collected, but all size variations (e.g., small, medium, large) were included. Meals that combined several food items that could be purchased separately, were also included as their own menu item. Data in PDFs were extracted either by copying and pasting the data into Excel or by using an online data extraction software (Smallpdf [25]). Where data had to be extracted from menu webpages, we used the web scraping tool Octoparse (v8 desktop [26]) to extract the menu data into an Excel sheet. The extracted data was crosschecked against the source material by AOH. The data collection approach for each restaurant is detailed in S2 Table. The data extracted included restaurant name, product name, nutritional information, and serving size, where available. Nutritional information was collected per 100 g and per serving wherever given, and included kcal, kJ, fat, saturated fat, carbohydrates, sugar, protein, fibre, salt, and sodium. Categorising menu items All menu items were categorised twice. First, author AOH assigned all menu items to one of the following 12 common subcategories: Pizzas, Burgers, Chicken, Other Mains, Children’s Meals, Salads, Sandwiches, Potato Sides, Other Sides, Breakfast Items, Desserts, and Sauces. This provided a consistent framework in which we could compare the nutritional content of similar menu items across restaurants. Hereafter, this categorisation will be referred to as an item’s ‘subcategory’. Categorisation criteria for each subcategory is detailed in S3 Table. Second, menu items were also matched to the OOH food categories as defined by the technical guidance for the sugar, salt, and calorie reduction targets [8,10,11]. The target-specific categories are not inclusive of all menu items and the guidance states that there should be no overlap between the calorie and sugar targets. Therefore, a menu item could be (i) ineligible for all three target types, (ii) eligible for only one of the three target types, (iii) eligible for a salt and calorie target but not a sugar target, or (iv) eligible for a salt and sugar target but not a calorie target. To align with the guidance that an item should not have both a calorie and a sugar target, we first categorised items by the sugar targets, and then categorised the remaining items by the calorie targets, as the sugar reduction target programme preceded the calorie target programme. The salt reduction programme provided separate targets for the OOH sector and the retailer and manufacturer sector. In line with the technical guidance for the salt targets, menu items were first categorised against the OOH targets, and any remaining uncategorised items were categorised against the retailer and manufacturer targets. A random 10% of menu items’ categorisations were checked by two co-authors (LB and HF) who had not been involved in the original categorisation process. A small subset of menu items was re-categorised and changes were applied to the complete dataset where relevant, followed by a final round of checking by author AOH. We divided our 21 restaurants into the following five ‘restaurant types’, based on the predominant main meal subcategory that the restaurant offered: Burger restaurants, Chicken restaurants, Pizza restaurants, Sandwich restaurants, and Other Mains restaurants. The 12 subcategories were also divided into ‘Mains’ and ‘Sides/Extras’, with ‘Desserts’ in its own category. Identifying restaurants This study focuses on the highest-grossing restaurant chains in the UK, as these restaurants have the largest impact on diets within the OOH food sector. We identified the 21 highest-grossing chained restaurants in the UK using the latest (2022) sales data from Euromonitor International, a privately-owned market research company. Euromonitor’s data portal Passport GMID was accessed via the Bodleian Library, University of Oxford [24]. We selected restaurants that Euromonitor defined as “Consumer Foodservice”, including “cafés/bars, full-service restaurants, limited-service restaurants, self-service cafeterias and street stalls/kiosks”. “Full-service” refers to sit-down restaurants with table service, while “limited-service” refers to fast food or take-away outlets. We excluded restaurants that had separate menus for each individual location to avoid introducing considerable heterogeneity to nutritional information at the restaurant level. Restaurants that only provided nutrition information per ingredient (e.g., bun, patty, cheese) rather than per item (e.g., complete burger) were also excluded, as it was unclear what a full menu item would consist of. We initially aimed for a sample of 20 restaurants. Subway was excluded at the start of the data collection phase as it provided nutrition information for individual ingredients only. However, new menu-item nutrition information was subsequently published after our data collection had started, and therefore Subway was reinstated and a sample of 21 restaurants’ data was collected. Data collection Menu data was collected directly from restaurants’ UK websites in February and March of 2024 (May 2024 for Subway). It was collected primarily through PDF menu documents or menu webpages. Where a restaurant provided multiple PDFs, data was only collected from those labelled as containing ‘core’ (or equivalent) menu items. Limited-time offer menu items (e.g., seasonal) were only included in data collection if they were present in the ‘core’ menus. For pizza restaurants that offered multiple crust options, only nutrition information from the default option (e.g., classic crust) was collected, but all size variations (e.g., small, medium, large) were included. Meals that combined several food items that could be purchased separately, were also included as their own menu item. Data in PDFs were extracted either by copying and pasting the data into Excel or by using an online data extraction software (Smallpdf [25]). Where data had to be extracted from menu webpages, we used the web scraping tool Octoparse (v8 desktop [26]) to extract the menu data into an Excel sheet. The extracted data was crosschecked against the source material by AOH. The data collection approach for each restaurant is detailed in S2 Table. The data extracted included restaurant name, product name, nutritional information, and serving size, where available. Nutritional information was collected per 100 g and per serving wherever given, and included kcal, kJ, fat, saturated fat, carbohydrates, sugar, protein, fibre, salt, and sodium. Categorising menu items All menu items were categorised twice. First, author AOH assigned all menu items to one of the following 12 common subcategories: Pizzas, Burgers, Chicken, Other Mains, Children’s Meals, Salads, Sandwiches, Potato Sides, Other Sides, Breakfast Items, Desserts, and Sauces. This provided a consistent framework in which we could compare the nutritional content of similar menu items across restaurants. Hereafter, this categorisation will be referred to as an item’s ‘subcategory’. Categorisation criteria for each subcategory is detailed in S3 Table. Second, menu items were also matched to the OOH food categories as defined by the technical guidance for the sugar, salt, and calorie reduction targets [8,10,11]. The target-specific categories are not inclusive of all menu items and the guidance states that there should be no overlap between the calorie and sugar targets. Therefore, a menu item could be (i) ineligible for all three target types, (ii) eligible for only one of the three target types, (iii) eligible for a salt and calorie target but not a sugar target, or (iv) eligible for a salt and sugar target but not a calorie target. To align with the guidance that an item should not have both a calorie and a sugar target, we first categorised items by the sugar targets, and then categorised the remaining items by the calorie targets, as the sugar reduction target programme preceded the calorie target programme. The salt reduction programme provided separate targets for the OOH sector and the retailer and manufacturer sector. In line with the technical guidance for the salt targets, menu items were first categorised against the OOH targets, and any remaining uncategorised items were categorised against the retailer and manufacturer targets. A random 10% of menu items’ categorisations were checked by two co-authors (LB and HF) who had not been involved in the original categorisation process. A small subset of menu items was re-categorised and changes were applied to the complete dataset where relevant, followed by a final round of checking by author AOH. We divided our 21 restaurants into the following five ‘restaurant types’, based on the predominant main meal subcategory that the restaurant offered: Burger restaurants, Chicken restaurants, Pizza restaurants, Sandwich restaurants, and Other Mains restaurants. The 12 subcategories were also divided into ‘Mains’ and ‘Sides/Extras’, with ‘Desserts’ in its own category. Analysis Data analyses were conducted in R (version 4.4.2) [27], using the following packages: readxl [28], dplyr [29], tidyr [30], stringr [31], tidyverse [32], zoo [33], writexl [34], reshape [35], ggplot2 [36]. Missing data We excluded menu items where all three of the nutrients of interest (kcal, salt, and sugar), were missing. For menu items that were missing two or fewer of the nutrients of interest, where possible, we calculated the missing values using the following formulas: In order to assess an item’s adherence to the Government’s reduction targets, nutrition information was needed per 100 g for the sugar targets, per serving for the calorie targets, and in both formats for the salt targets. If per 100 g, per serving, or serving size information was missing, the missing value was calculated from the provided information using the following formulas: Here, ‘Serving Size’ refers to the weight of the menu item in grams, and ‘Per 100 g’ and ‘Per Serving’ refer to nutrient content (e.g., salt content in 100 g or in a single serving of a menu item). Where it was not possible to calculate serving size (only per 100 g or only per serving nutrient information was provided), the mean serving size for the menu item’s subcategory was used instead. For example, if the serving size for a burger was not provided, the mean serving size of menu items within the ‘Burger’ subcategory was used instead. Average nutrient content The mean and median kcal, sugar, salt, and fat content was calculated for each restaurant and subcategory. Averages were calculated using nutrient information provided per 100 g, per reported serving size, and per subcategory average serving size. ‘Reported serving size’ refers to the serving size of a menu item as reported by the restaurant, whereas ‘subcategory average serving size’ refers to the average serving size of menu items within a subcategory (e.g., the average serving size for menu items within the ‘Burger’ subcategory). For Pizzas, average nutrient content per serving was calculated using the serving sizes provided by individual restaurants. Papa John’s provided per serving information per slice, Pizza Hut and Domino’s provided per serving information per person sharing (e.g., medium pizza is shared between two people, so one pizza would count as two servings), and the remaining restaurants provided per serving information per whole pizza. Adherence to targets The sugar, salt, and calorie content of each product was compared to their matched category’s target value. S4 Table shows the range of target values that menu items had to meet, within each subcategory. Where a menu item’s sugar, salt, or calorie content was equal to or lower than the target value, they were deemed to have met the target. The proportion of each restaurant and subcategory’s menu items that met the targets was expressed as a percentage of their total number of menu items eligible for the respective target. Sensitivity analyses We repeated the primary analyses for mean nutrient content and target adherence, but with limited-time offer menu items that featured on the main menu, excluded. We repeated the primary analyses, but for items where the subcategory average serving size was used to impute nutritional content (as the restaurant did not provide it), we instead used the lower quartile and upper quartile subcategory serving sizes to impute nutritional content, to check the robustness of our findings to altering our imputation strategy. Deviations from protocol We had planned to use ANOVAs to determine whether the average nutrient content differed significantly across restaurants and subcategories, and logistic regression models to determine whether a menu item’s restaurant or subcategory was a predictor of their adherence to the sugar, salt, and calorie reduction targets. We have opted instead for a descriptive presentation of the results, as the dataset is not a random sample from a wider population of menu items. This means that any differences in nutrient content or target adherence that we observed between restaurants or subcategories, are in fact real, and not attributable to sampling variation. As a result, inferential statistical testing is not warranted in this case. As a sensitivity analysis, we planned to repeat the primary analysis for target adherence but using ‘as sold’ nutrition information instead of ‘per serving’ information. For example, if a restaurant reports that a menu item contains two servings, our primary analysis would have assessed target adherence based on a single serving (being ‘per serving’), while our sensitivity analysis would assess it as the whole menu item (‘as sold’). The aim of this analysis was to remove the influence of individual restaurants’ reporting of serving size, to allow for a consistent assessment of nutritional quality across restaurants. On examination of the data, we found for Pizzas in particular, the suggested serving sizes and reporting of per serving nutrition information lacked consistency across restaurants. For example, Papa John’s provided their nutrition information per slice without explicitly stating how many slices equate to a single serving, while Pizza Express provided information per whole pizza with no suggestions for how many servings it contained. The technical guidance for the salt and calorie targets attempts to account for this, with the salt targets being applied per slice or per whole pizza depending on the style of pizza (takeaway or Italian-style), and the calorie targets being applied per serving for ‘sharing’ pizzas (defined as large and above, or 11.5″ and above). Our analysis followed this guidance, with assumptions for serving size being made to apply the calorie targets where information was not provided. For example, Papa John’s large and extra-large pizzas were assumed to contain three and four servings respectively, as they did not provide their own suggestions, and Italian-style pizzas were assumed to be a single serving. For analyses with the outcome of average nutrient content per serving, serving sizes for Pizzas were as provided by the restaurants. We planned to conduct a sensitivity analysis in which we would repeat all primary analyses, but for menu items that did not have a serving size reported, we would use an applicable serving size from the Food Standard’s Agency ‘Food Portion Sizes’ handbook [37]. Upon further investigation, we concluded that the suggested serving sizes from this handbook would not be applicable to OOH menu items, as they were often provided per meal component rather than complete meal (e.g., suggested serving size provided for a burger patty and a bun, rather than a whole burger), and therefore this analysis was not conducted. The serving size imputation sensitivity analysis was not pre-registered, but was included to test the robustness of our approach to dealing with missing data. The planned exploratory analysis comparing the same menu items reported in 2022 with 2024 was not conducted due to the small sample size of products present on menus in both years. As the targets are food only, and the same top-selling brands of drinks appear on the majority of restaurant’s menus, drinks were excluded. The planned analysis using the UK Ofcom/FSA Nutrient Profile Model as an additional assessment of nutritional quality, will be presented in a separate paper. Missing data We excluded menu items where all three of the nutrients of interest (kcal, salt, and sugar), were missing. For menu items that were missing two or fewer of the nutrients of interest, where possible, we calculated the missing values using the following formulas: In order to assess an item’s adherence to the Government’s reduction targets, nutrition information was needed per 100 g for the sugar targets, per serving for the calorie targets, and in both formats for the salt targets. If per 100 g, per serving, or serving size information was missing, the missing value was calculated from the provided information using the following formulas: Here, ‘Serving Size’ refers to the weight of the menu item in grams, and ‘Per 100 g’ and ‘Per Serving’ refer to nutrient content (e.g., salt content in 100 g or in a single serving of a menu item). Where it was not possible to calculate serving size (only per 100 g or only per serving nutrient information was provided), the mean serving size for the menu item’s subcategory was used instead. For example, if the serving size for a burger was not provided, the mean serving size of menu items within the ‘Burger’ subcategory was used instead. Average nutrient content The mean and median kcal, sugar, salt, and fat content was calculated for each restaurant and subcategory. Averages were calculated using nutrient information provided per 100 g, per reported serving size, and per subcategory average serving size. ‘Reported serving size’ refers to the serving size of a menu item as reported by the restaurant, whereas ‘subcategory average serving size’ refers to the average serving size of menu items within a subcategory (e.g., the average serving size for menu items within the ‘Burger’ subcategory). For Pizzas, average nutrient content per serving was calculated using the serving sizes provided by individual restaurants. Papa John’s provided per serving information per slice, Pizza Hut and Domino’s provided per serving information per person sharing (e.g., medium pizza is shared between two people, so one pizza would count as two servings), and the remaining restaurants provided per serving information per whole pizza. Adherence to targets The sugar, salt, and calorie content of each product was compared to their matched category’s target value. S4 Table shows the range of target values that menu items had to meet, within each subcategory. Where a menu item’s sugar, salt, or calorie content was equal to or lower than the target value, they were deemed to have met the target. The proportion of each restaurant and subcategory’s menu items that met the targets was expressed as a percentage of their total number of menu items eligible for the respective target. Sensitivity analyses We repeated the primary analyses for mean nutrient content and target adherence, but with limited-time offer menu items that featured on the main menu, excluded. We repeated the primary analyses, but for items where the subcategory average serving size was used to impute nutritional content (as the restaurant did not provide it), we instead used the lower quartile and upper quartile subcategory serving sizes to impute nutritional content, to check the robustness of our findings to altering our imputation strategy. Deviations from protocol We had planned to use ANOVAs to determine whether the average nutrient content differed significantly across restaurants and subcategories, and logistic regression models to determine whether a menu item’s restaurant or subcategory was a predictor of their adherence to the sugar, salt, and calorie reduction targets. We have opted instead for a descriptive presentation of the results, as the dataset is not a random sample from a wider population of menu items. This means that any differences in nutrient content or target adherence that we observed between restaurants or subcategories, are in fact real, and not attributable to sampling variation. As a result, inferential statistical testing is not warranted in this case. As a sensitivity analysis, we planned to repeat the primary analysis for target adherence but using ‘as sold’ nutrition information instead of ‘per serving’ information. For example, if a restaurant reports that a menu item contains two servings, our primary analysis would have assessed target adherence based on a single serving (being ‘per serving’), while our sensitivity analysis would assess it as the whole menu item (‘as sold’). The aim of this analysis was to remove the influence of individual restaurants’ reporting of serving size, to allow for a consistent assessment of nutritional quality across restaurants. On examination of the data, we found for Pizzas in particular, the suggested serving sizes and reporting of per serving nutrition information lacked consistency across restaurants. For example, Papa John’s provided their nutrition information per slice without explicitly stating how many slices equate to a single serving, while Pizza Express provided information per whole pizza with no suggestions for how many servings it contained. The technical guidance for the salt and calorie targets attempts to account for this, with the salt targets being applied per slice or per whole pizza depending on the style of pizza (takeaway or Italian-style), and the calorie targets being applied per serving for ‘sharing’ pizzas (defined as large and above, or 11.5″ and above). Our analysis followed this guidance, with assumptions for serving size being made to apply the calorie targets where information was not provided. For example, Papa John’s large and extra-large pizzas were assumed to contain three and four servings respectively, as they did not provide their own suggestions, and Italian-style pizzas were assumed to be a single serving. For analyses with the outcome of average nutrient content per serving, serving sizes for Pizzas were as provided by the restaurants. We planned to conduct a sensitivity analysis in which we would repeat all primary analyses, but for menu items that did not have a serving size reported, we would use an applicable serving size from the Food Standard’s Agency ‘Food Portion Sizes’ handbook [37]. Upon further investigation, we concluded that the suggested serving sizes from this handbook would not be applicable to OOH menu items, as they were often provided per meal component rather than complete meal (e.g., suggested serving size provided for a burger patty and a bun, rather than a whole burger), and therefore this analysis was not conducted. The serving size imputation sensitivity analysis was not pre-registered, but was included to test the robustness of our approach to dealing with missing data. The planned exploratory analysis comparing the same menu items reported in 2022 with 2024 was not conducted due to the small sample size of products present on menus in both years. As the targets are food only, and the same top-selling brands of drinks appear on the majority of restaurant’s menus, drinks were excluded. The planned analysis using the UK Ofcom/FSA Nutrient Profile Model as an additional assessment of nutritional quality, will be presented in a separate paper. Results A total of 3,099 menu items across 21 restaurants were included in this study. Originally, 5,435 menu items were collected. Of the 2,336 items excluded, 453 were duplicates, 103 did not have the required nutritional information, 825 were pizzas with non-default crust options, and 955 were drinks. One menu item was missing a sugar value, so this item was excluded from analyses where sugar content is the outcome. For 1,630 out of 3,099 menu items, no serving size was provided and only one format of nutrient content was provided (either per 100 g or per serving), and therefore the subcategory average serving size was used in calculating the missing nutrient content. S5 and S6 Tables provide the number of affected products, and the mean, median, and lower and upper quartiles for serving size (g) (from the 1,469 where it was provided/calculated), for each restaurant and subcategory. The mean number of menu items per restaurant was 148, although this varied widely from 40 items for Burger King to 330 items for Pizza Hut (Table 1). The number of subcategories offered by each restaurant ranged from 5 for Starbucks and Burger King, to 11 for Leon. Some restaurants had large menus that focussed on a smaller number of subcategories, such as Pizza Hut and Domino’s, while others had smaller menus with more diverse types of items, such as Leon (Fig 1). Download: PNG larger image TIFF original image Fig 1. The number and proportion of menu items by restaurant and subcategory. The values in boxes show the number of menu items belonging to that restaurant-subcategory. Main meals are coloured in blue, sides/extras are coloured in red, and desserts are coloured in purple. The darkness of each colour indicates the proportion of menu items from each restaurant that belong to that subcategory (calculated separately for mains, sides/extras, and desserts, with the ‘total’ rows calculated across mains, sides/extras, and desserts), where a darker colour indicates a higher proportion. KFC and Nando’s were characterised as ‘Chicken’ restaurant type despite majority of main meal menu items belonging to ‘Sandwich’ or ‘Burger’ subcategories, as these menu items were majority chicken burgers and chicken wraps. The same principle was applied to restaurants with the majority of their main menu items being ‘Breakfast Items’. https://doi.org/10.1371/journal.pmed.1004681.g001 Average nutrient content By subcategory. Per 100 g, mean nutrient content across all menu items was 277 kcal, 1.1 g salt, and 9.5 g sugar. Per serving, mean nutrient content across all menu items was 450 kcal, 2.0 g salt, and 10.9 g sugar. Desserts had the highest mean calorie (409 kcal) and sugar (34.2 g) content per 100 g across all subcategories, while Sauces had the highest salt content per 100 g (2.2 g) (Fig 2). Per serving, Other Mains had the highest calorie content (756 kcal) and joint highest salt content with Pizzas (3.4 g), and Desserts had the highest sugar content (26.2 g) (Fig 3). Means, Medians, and Standard Deviations for nutrient content per 100 g, per reported serving, and per subcategory average serving, for each subcategory, are found in S7–S10 Tables. Download: PNG larger image TIFF original image Fig 2. The distribution of kcal, salt, and sugar content per 100 g for menu items by subcategory. Subcategories are ordered ascendingly by median kcal per 100 g. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Main meal subcategories are coloured in blue, Side categories are coloured in red, and Desserts are coloured in purple. https://doi.org/10.1371/journal.pmed.1004681.g002 Download: PNG larger image TIFF original image Fig 3. The distribution of kcal, salt, and sugar content per serving for menu items by subcategory. Subcategories are ordered ascendingly by median kcal per serving. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Main meal subcategories are coloured in blue, Side categories are coloured in red, and Desserts are coloured in purple. https://doi.org/10.1371/journal.pmed.1004681.g003 By restaurant. Menu items from Caffé Nero had the highest mean calorie content per 100 g (368 kcal) across all restaurants, while menu items from Prezzo had the highest salt content (1.8 g), and menu items from Costa had the highest sugar content (22.0 g) (Fig 4). Per serving, menu items from Prezzo had the highest calorie (644 kcal) and salt content (3.7 g) and menu items from Harvester had the highest sugar content (17.2 g) (Fig 5). Means, Medians, and Standard Deviations for nutrient content per 100 g, per reported serving, and per subcategory average serving, for each restaurant, are found in S11–S14 Tables. Download: PNG larger image TIFF original image Fig 4. The distribution of kcal, salt, and sugar content per 100 g for menu items by restaurant. Restaurants are ordered ascendingly by median kcal per 100 g. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Burger restaurants are coloured in purple, Chicken restaurants in yellow, Pizza restaurants in Blue, Other Main restaurants in green, and Sandwich restaurants in pink. https://doi.org/10.1371/journal.pmed.1004681.g004 Download: PNG larger image TIFF original image Fig 5. The distribution of kcal, salt, and sugar content per serving for menu items by restaurant. Restaurants are ordered ascendingly by median kcal per serving. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Burger restaurants are coloured in purple, Chicken restaurants in yellow, Pizza restaurants in Blue, Other Main restaurants in green, and Sandwich restaurants in pink. https://doi.org/10.1371/journal.pmed.1004681.g005 Averaging across menu items within restaurant groups, menu items from the Sandwich group had the highest mean calorie (293 kcal) and sugar content per 100 g (13.4 g), and menu items from the Pizza group had the highest mean salt content (1.4 g). Menu items from the Pizza group had the highest mean calorie (515 kcal) and salt content per serving (2.6 g), and menu items from the Other Mains group had the highest mean sugar content (12.6 g). Means, Medians, and Standard Deviations for nutrient content per 100 g, per reported serving, and per subcategory average serving, for each restaurant group, are found in S15–S18 Tables. Target adherence Across all restaurants and subcategories, 61% of menu items met their calorie targets (n = 1300/2148), 58% met their salt targets (n = 1348/2344), 36% met their sugar targets (n = 207/578), and 43% met all of their applicable targets (n = 1271/2951). By subcategory. Six out of the 12 subcategories had over 50% of their menu items meeting all applicable targets. Salads had the highest adherence to all applicable targets at 96% (n = 44/46) but were only eligible for the calorie targets (therefore having 96% adherence for calorie targets as well). Breakfast Items had the second highest adherence to all applicable targets at 66% (n = 125/190), while being eligible for all three target types. Excluding Other Sides and Children’s Meals (which had 100% adherence to the sugar targets but each had only one eligible item), the highest sugar target adherence was seen for Breakfast Items at 74% (n = 56/76), which also had the highest salt target adherence at 82% (n = 126/154) (Fig 6). Download: PNG larger image TIFF original image Fig 6. The proportion of menu items that met sugar, salt, calorie, and all applicable targets, for each subcategory. Values in brackets show the total number of menu items that were eligible for the given target, in that subcategory. Subcategories are ordered descending by the proportion of menu items meeting the respective target, with ‘All Subcategories’ as the top bar. The dark orange bars indicate the proportion of menu items meeting the target, and the light orange bars indicate the proportion of menu items not meeting the target. https://doi.org/10.1371/journal.pmed.1004681.g006 By restaurant. Nine of the 21 restaurants had over 50% of their menu items meeting all applicable targets (Fig 7). Download: PNG larger image TIFF original image Fig 7. The proportion of menu items that met sugar, salt, calorie, and all applicable targets, for each restaurant. Values in brackets show the total number of menu items that were eligible for the given target, in that restaurant. Restaurants are ordered descending by the proportion of menu items meeting the respective target, with ‘All Restaurants’ as the top bar. The dark blue bars indicate the proportion of menu items meeting the target, and the light blue bars indicate the proportion of menu items not meeting the target. https://doi.org/10.1371/journal.pmed.1004681.g007 Menu items from the Pizza restaurant group had the lowest combined adherence to all applicable targets at 32% (n = 310/961), to the salt targets at 49% (n = 409/832), and to the calorie targets at 53% (n = 449/846). Adherence to sugar targets was based on lower product numbers, with menu items from the Chicken restaurant group having the lowest adherence at 0%, but with only 23 eligible items. Menu items from the Pizza restaurant group were the second lowest adhering to the sugar targets at 23%, with 65 eligible items. Menu items from the Burger restaurant group had the highest combined adherence to all applicable targets at 59% (n = 88/149), and to the salt targets at 80% (n = 92/115), while menu items from the Chicken group had the highest combined adherence to the calorie targets at 78% (n = 97/124). Menu items from the Burger restaurant group had the highest adherence to the sugar targets at 53%, but with only 36 eligible items, of which 32 were from McDonald’s. Full results for adherence to each target type by restaurant group can be found in S19 Table. Sensitivity analysis excluding limited-time offer menu items We excluded 37 limited-time offer menu items, spanning four restaurants (McDonald’s, Burger King, Pret, and KFC) and eight unique subcategories, for this analysis. The average nutrient content across all menu items did not differ from the primary analysis, except for calorie content per serving which was 1 kcal higher in the sensitivity analysis (451 kcal from 450 kcal). Overall adherence to sugar, salt, and all applicable targets did not differ from the primary analysis, but the proportion of menu items meeting calorie targets dropped by 1% (from 61% to 60%). S20 and S21 Tables provide the average nutrient content and target adherence for the four restaurants with limited-time menu items (McDonald’s, Burger King, Pret, and KFC), including and excluding limited-time offer items. Sensitivity analysis using lower and upper quartiles for serving size Table 2 provides the average nutrient content across all menu items when the subcategory average, lower quartile, and upper quartile serving sizes were used to replace missing serving sizes. S22–S25 Tables provide the equivalent values for each subcategory and restaurant. Download: PNG larger image TIFF original image Table 2. Average nutrient content across all menu items when the subcategory average, lower quartile, and upper quartile serving sizes were used to replace missing serving sizes. https://doi.org/10.1371/journal.pmed.1004681.t002 Table 3 provides the proportion of menu items meeting sugar, salt, calorie, and all applicable targets across all menu items, when the subcategory average, lower quartile, and upper quartile serving sizes were used to replace missing serving size. S26 and S27 Tables provide the equivalent values for each subcategory and restaurant. Download: PNG larger image TIFF original image Table 3. Target adherence across all menu items when the subcategory average, lower quartile, and upper quartile serving sizes were used to replace missing serving sizes. https://doi.org/10.1371/journal.pmed.1004681.t003 Average nutrient content By subcategory. Per 100 g, mean nutrient content across all menu items was 277 kcal, 1.1 g salt, and 9.5 g sugar. Per serving, mean nutrient content across all menu items was 450 kcal, 2.0 g salt, and 10.9 g sugar. Desserts had the highest mean calorie (409 kcal) and sugar (34.2 g) content per 100 g across all subcategories, while Sauces had the highest salt content per 100 g (2.2 g) (Fig 2). Per serving, Other Mains had the highest calorie content (756 kcal) and joint highest salt content with Pizzas (3.4 g), and Desserts had the highest sugar content (26.2 g) (Fig 3). Means, Medians, and Standard Deviations for nutrient content per 100 g, per reported serving, and per subcategory average serving, for each subcategory, are found in S7–S10 Tables. Download: PNG larger image TIFF original image Fig 2. The distribution of kcal, salt, and sugar content per 100 g for menu items by subcategory. Subcategories are ordered ascendingly by median kcal per 100 g. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Main meal subcategories are coloured in blue, Side categories are coloured in red, and Desserts are coloured in purple. https://doi.org/10.1371/journal.pmed.1004681.g002 Download: PNG larger image TIFF original image Fig 3. The distribution of kcal, salt, and sugar content per serving for menu items by subcategory. Subcategories are ordered ascendingly by median kcal per serving. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Main meal subcategories are coloured in blue, Side categories are coloured in red, and Desserts are coloured in purple. https://doi.org/10.1371/journal.pmed.1004681.g003 By restaurant. Menu items from Caffé Nero had the highest mean calorie content per 100 g (368 kcal) across all restaurants, while menu items from Prezzo had the highest salt content (1.8 g), and menu items from Costa had the highest sugar content (22.0 g) (Fig 4). Per serving, menu items from Prezzo had the highest calorie (644 kcal) and salt content (3.7 g) and menu items from Harvester had the highest sugar content (17.2 g) (Fig 5). Means, Medians, and Standard Deviations for nutrient content per 100 g, per reported serving, and per subcategory average serving, for each restaurant, are found in S11–S14 Tables. Download: PNG larger image TIFF original image Fig 4. The distribution of kcal, salt, and sugar content per 100 g for menu items by restaurant. Restaurants are ordered ascendingly by median kcal per 100 g. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Burger restaurants are coloured in purple, Chicken restaurants in yellow, Pizza restaurants in Blue, Other Main restaurants in green, and Sandwich restaurants in pink. https://doi.org/10.1371/journal.pmed.1004681.g004 Download: PNG larger image TIFF original image Fig 5. The distribution of kcal, salt, and sugar content per serving for menu items by restaurant. Restaurants are ordered ascendingly by median kcal per serving. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Burger restaurants are coloured in purple, Chicken restaurants in yellow, Pizza restaurants in Blue, Other Main restaurants in green, and Sandwich restaurants in pink. https://doi.org/10.1371/journal.pmed.1004681.g005 Averaging across menu items within restaurant groups, menu items from the Sandwich group had the highest mean calorie (293 kcal) and sugar content per 100 g (13.4 g), and menu items from the Pizza group had the highest mean salt content (1.4 g). Menu items from the Pizza group had the highest mean calorie (515 kcal) and salt content per serving (2.6 g), and menu items from the Other Mains group had the highest mean sugar content (12.6 g). Means, Medians, and Standard Deviations for nutrient content per 100 g, per reported serving, and per subcategory average serving, for each restaurant group, are found in S15–S18 Tables. By subcategory. Per 100 g, mean nutrient content across all menu items was 277 kcal, 1.1 g salt, and 9.5 g sugar. Per serving, mean nutrient content across all menu items was 450 kcal, 2.0 g salt, and 10.9 g sugar. Desserts had the highest mean calorie (409 kcal) and sugar (34.2 g) content per 100 g across all subcategories, while Sauces had the highest salt content per 100 g (2.2 g) (Fig 2). Per serving, Other Mains had the highest calorie content (756 kcal) and joint highest salt content with Pizzas (3.4 g), and Desserts had the highest sugar content (26.2 g) (Fig 3). Means, Medians, and Standard Deviations for nutrient content per 100 g, per reported serving, and per subcategory average serving, for each subcategory, are found in S7–S10 Tables. Download: PNG larger image TIFF original image Fig 2. The distribution of kcal, salt, and sugar content per 100 g for menu items by subcategory. Subcategories are ordered ascendingly by median kcal per 100 g. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Main meal subcategories are coloured in blue, Side categories are coloured in red, and Desserts are coloured in purple. https://doi.org/10.1371/journal.pmed.1004681.g002 Download: PNG larger image TIFF original image Fig 3. The distribution of kcal, salt, and sugar content per serving for menu items by subcategory. Subcategories are ordered ascendingly by median kcal per serving. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Main meal subcategories are coloured in blue, Side categories are coloured in red, and Desserts are coloured in purple. https://doi.org/10.1371/journal.pmed.1004681.g003 By restaurant. Menu items from Caffé Nero had the highest mean calorie content per 100 g (368 kcal) across all restaurants, while menu items from Prezzo had the highest salt content (1.8 g), and menu items from Costa had the highest sugar content (22.0 g) (Fig 4). Per serving, menu items from Prezzo had the highest calorie (644 kcal) and salt content (3.7 g) and menu items from Harvester had the highest sugar content (17.2 g) (Fig 5). Means, Medians, and Standard Deviations for nutrient content per 100 g, per reported serving, and per subcategory average serving, for each restaurant, are found in S11–S14 Tables. Download: PNG larger image TIFF original image Fig 4. The distribution of kcal, salt, and sugar content per 100 g for menu items by restaurant. Restaurants are ordered ascendingly by median kcal per 100 g. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Burger restaurants are coloured in purple, Chicken restaurants in yellow, Pizza restaurants in Blue, Other Main restaurants in green, and Sandwich restaurants in pink. https://doi.org/10.1371/journal.pmed.1004681.g004 Download: PNG larger image TIFF original image Fig 5. The distribution of kcal, salt, and sugar content per serving for menu items by restaurant. Restaurants are ordered ascendingly by median kcal per serving. Mid-lines represent the median, the box represents the inter-quartile range, and whiskers represent the range. Burger restaurants are coloured in purple, Chicken restaurants in yellow, Pizza restaurants in Blue, Other Main restaurants in green, and Sandwich restaurants in pink. https://doi.org/10.1371/journal.pmed.1004681.g005 Averaging across menu items within restaurant groups, menu items from the Sandwich group had the highest mean calorie (293 kcal) and sugar content per 100 g (13.4 g), and menu items from the Pizza group had the highest mean salt content (1.4 g). Menu items from the Pizza group had the highest mean calorie (515 kcal) and salt content per serving (2.6 g), and menu items from the Other Mains group had the highest mean sugar content (12.6 g). Means, Medians, and Standard Deviations for nutrient content per 100 g, per reported serving, and per subcategory average serving, for each restaurant group, are found in S15–S18 Tables. Target adherence Across all restaurants and subcategories, 61% of menu items met their calorie targets (n = 1300/2148), 58% met their salt targets (n = 1348/2344), 36% met their sugar targets (n = 207/578), and 43% met all of their applicable targets (n = 1271/2951). By subcategory. Six out of the 12 subcategories had over 50% of their menu items meeting all applicable targets. Salads had the highest adherence to all applicable targets at 96% (n = 44/46) but were only eligible for the calorie targets (therefore having 96% adherence for calorie targets as well). Breakfast Items had the second highest adherence to all applicable targets at 66% (n = 125/190), while being eligible for all three target types. Excluding Other Sides and Children’s Meals (which had 100% adherence to the sugar targets but each had only one eligible item), the highest sugar target adherence was seen for Breakfast Items at 74% (n = 56/76), which also had the highest salt target adherence at 82% (n = 126/154) (Fig 6). Download: PNG larger image TIFF original image Fig 6. The proportion of menu items that met sugar, salt, calorie, and all applicable targets, for each subcategory. Values in brackets show the total number of menu items that were eligible for the given target, in that subcategory. Subcategories are ordered descending by the proportion of menu items meeting the respective target, with ‘All Subcategories’ as the top bar. The dark orange bars indicate the proportion of menu items meeting the target, and the light orange bars indicate the proportion of menu items not meeting the target. https://doi.org/10.1371/journal.pmed.1004681.g006 By restaurant. Nine of the 21 restaurants had over 50% of their menu items meeting all applicable targets (Fig 7). Download: PNG larger image TIFF original image Fig 7. The proportion of menu items that met sugar, salt, calorie, and all applicable targets, for each restaurant. Values in brackets show the total number of menu items that were eligible for the given target, in that restaurant. Restaurants are ordered descending by the proportion of menu items meeting the respective target, with ‘All Restaurants’ as the top bar. The dark blue bars indicate the proportion of menu items meeting the target, and the light blue bars indicate the proportion of menu items not meeting the target. https://doi.org/10.1371/journal.pmed.1004681.g007 Menu items from the Pizza restaurant group had the lowest combined adherence to all applicable targets at 32% (n = 310/961), to the salt targets at 49% (n = 409/832), and to the calorie targets at 53% (n = 449/846). Adherence to sugar targets was based on lower product numbers, with menu items from the Chicken restaurant group having the lowest adherence at 0%, but with only 23 eligible items. Menu items from the Pizza restaurant group were the second lowest adhering to the sugar targets at 23%, with 65 eligible items. Menu items from the Burger restaurant group had the highest combined adherence to all applicable targets at 59% (n = 88/149), and to the salt targets at 80% (n = 92/115), while menu items from the Chicken group had the highest combined adherence to the calorie targets at 78% (n = 97/124). Menu items from the Burger restaurant group had the highest adherence to the sugar targets at 53%, but with only 36 eligible items, of which 32 were from McDonald’s. Full results for adherence to each target type by restaurant group can be found in S19 Table. By subcategory. Six out of the 12 subcategories had over 50% of their menu items meeting all applicable targets. Salads had the highest adherence to all applicable targets at 96% (n = 44/46) but were only eligible for the calorie targets (therefore having 96% adherence for calorie targets as well). Breakfast Items had the second highest adherence to all applicable targets at 66% (n = 125/190), while being eligible for all three target types. Excluding Other Sides and Children’s Meals (which had 100% adherence to the sugar targets but each had only one eligible item), the highest sugar target adherence was seen for Breakfast Items at 74% (n = 56/76), which also had the highest salt target adherence at 82% (n = 126/154) (Fig 6). Download: PNG larger image TIFF original image Fig 6. The proportion of menu items that met sugar, salt, calorie, and all applicable targets, for each subcategory. Values in brackets show the total number of menu items that were eligible for the given target, in that subcategory. Subcategories are ordered descending by the proportion of menu items meeting the respective target, with ‘All Subcategories’ as the top bar. The dark orange bars indicate the proportion of menu items meeting the target, and the light orange bars indicate the proportion of menu items not meeting the target. https://doi.org/10.1371/journal.pmed.1004681.g006 By restaurant. Nine of the 21 restaurants had over 50% of their menu items meeting all applicable targets (Fig 7). Download: PNG larger image TIFF original image Fig 7. The proportion of menu items that met sugar, salt, calorie, and all applicable targets, for each restaurant. Values in brackets show the total number of menu items that were eligible for the given target, in that restaurant. Restaurants are ordered descending by the proportion of menu items meeting the respective target, with ‘All Restaurants’ as the top bar. The dark blue bars indicate the proportion of menu items meeting the target, and the light blue bars indicate the proportion of menu items not meeting the target. https://doi.org/10.1371/journal.pmed.1004681.g007 Menu items from the Pizza restaurant group had the lowest combined adherence to all applicable targets at 32% (n = 310/961), to the salt targets at 49% (n = 409/832), and to the calorie targets at 53% (n = 449/846). Adherence to sugar targets was based on lower product numbers, with menu items from the Chicken restaurant group having the lowest adherence at 0%, but with only 23 eligible items. Menu items from the Pizza restaurant group were the second lowest adhering to the sugar targets at 23%, with 65 eligible items. Menu items from the Burger restaurant group had the highest combined adherence to all applicable targets at 59% (n = 88/149), and to the salt targets at 80% (n = 92/115), while menu items from the Chicken group had the highest combined adherence to the calorie targets at 78% (n = 97/124). Menu items from the Burger restaurant group had the highest adherence to the sugar targets at 53%, but with only 36 eligible items, of which 32 were from McDonald’s. Full results for adherence to each target type by restaurant group can be found in S19 Table. Sensitivity analysis excluding limited-time offer menu items We excluded 37 limited-time offer menu items, spanning four restaurants (McDonald’s, Burger King, Pret, and KFC) and eight unique subcategories, for this analysis. The average nutrient content across all menu items did not differ from the primary analysis, except for calorie content per serving which was 1 kcal higher in the sensitivity analysis (451 kcal from 450 kcal). Overall adherence to sugar, salt, and all applicable targets did not differ from the primary analysis, but the proportion of menu items meeting calorie targets dropped by 1% (from 61% to 60%). S20 and S21 Tables provide the average nutrient content and target adherence for the four restaurants with limited-time menu items (McDonald’s, Burger King, Pret, and KFC), including and excluding limited-time offer items. Sensitivity analysis using lower and upper quartiles for serving size Table 2 provides the average nutrient content across all menu items when the subcategory average, lower quartile, and upper quartile serving sizes were used to replace missing serving sizes. S22–S25 Tables provide the equivalent values for each subcategory and restaurant. Download: PNG larger image TIFF original image Table 2. Average nutrient content across all menu items when the subcategory average, lower quartile, and upper quartile serving sizes were used to replace missing serving sizes. https://doi.org/10.1371/journal.pmed.1004681.t002 Table 3 provides the proportion of menu items meeting sugar, salt, calorie, and all applicable targets across all menu items, when the subcategory average, lower quartile, and upper quartile serving sizes were used to replace missing serving size. S26 and S27 Tables provide the equivalent values for each subcategory and restaurant. Download: PNG larger image TIFF original image Table 3. Target adherence across all menu items when the subcategory average, lower quartile, and upper quartile serving sizes were used to replace missing serving sizes. https://doi.org/10.1371/journal.pmed.1004681.t003 Discussion This study shows that only 43% of menu items from the highest-grossing UK restaurant chains had met all of their reduction targets at the start of 2024, indicating low adherence with the reduction programmes from the OOH sector. The majority of menu items met the calorie (61%) and salt (58%) reduction targets, however, only 36% of menu items met sugar targets. Heterogeneity in adherence was observed across food categories, with Desserts having the lowest proportion of menu items meeting all applicable targets at 22%, and Salads the highest at 96%. Heterogeneity in target adherence was also observed across restaurants, with Papa John’s having the lowest proportion of menu items that met all their applicable targets at 8%, and Subway the highest at 76%. A similar number of restaurants had over 50% of their menu items meeting the calorie and salt targets (17/21 and 18/21, respectively), but salt target adherence varied much more widely from 8%–89%, compared to 35%–89% for calories. Subcategories were not always consistent in their performance across the target types, for example, Children’s Meals had 95% of menu items meeting calorie targets, but 62% meeting salt targets. To our knowledge, only two studies have looked at company-level performance against the reduction targets, both focussing on the sugar reduction programme only, with one looking at manufacturers of grocery foods and the other at OOH. Bandy and colleagues [20] found that in 2018, just under half (24/50) of the best-selling manufacturers across five food categories (biscuits and cereal bars, breakfast cereals, chocolate confectionery, sugar confectionery, and yoghurts) met the intermediary 2018 sugar targets, and four companies had already met the 2020 targets. Our study found no companies had 100% of menu items meeting their sugar targets, potentially indicating greater target adherence from the retailer grocery food sector compared to OOH. However, our study included a wider range of food categories, used nutritional data from 2024 (six years later), and Bandy and colleagues used sales-weighted averages for sugar content (by brand sales volume), which we did not have the relevant data to replicate. Pepper and colleagues [21] found that 4/48 OOH companies met the 20% reduction target for Desserts in 2020, which given that Bandy and colleagues found four manufacturers had already met the 20% reduction target by 2018, this could provide another indication that the grocery sector is more engaged with the programme than OOH. Pepper and colleagues observed large variation in sugar and calorie content between companies with different styles of food, but also between similar chains, which corroborates with findings from our study across all three target nutrients (sugar, salt, and calories) and beyond just Desserts. Together, these findings suggest that while adherence to the targets may be low overall, there are examples of OOH companies that are performing well, and performance is not constrained by the type of cuisine being offered. The most recent sugar reduction progress report found no OOH companies met the 20% reduction target by 2020 [12], and while this is consistent with our findings, it was based on data from only five companies, and four of which were sandwich/café restaurants (Costa, Pret, Greggs, and Starbucks). The most recent salt reduction progress report had insufficient data to present company-level performance for the OOH sector, although it reported that overall, 74% of OOH products met their salt target [38]. Our study found only 58% of menu items met their salt target, potentially due to the progress report using the ‘maximum’ target values while we used the ‘average’ target values, which were less lenient (e.g., Dips had a ‘maximum’ target of 0.9g per 100 g but an ‘average’ target of 0.75 g). There is no relevant comparison between our findings and those in the calorie reduction progress report, as this report did not include any company-level analyses [13]. Reduction target programmes in other countries have been found to be effective in reducing the sugar, salt, calorie, and fat content of food products. One systematic review found that of 26 studies evaluating government-set reduction targets across 15 different countries, 22 found improvements in nutritional quality [39]. However, the vast majority of these studies only focussed on salt, and reported the percentage change in salt content over time rather than adherence to the targets specifically, so our findings may not be directly comparable. A study assessing OOH foods in the US found significantly fewer fast-food meals met the American Heart Association’s calorie guidelines in 2015, 2016, and 2017 compared to 2008, with no changes observed for saturated fat or sodium [40], suggesting that our findings are indicative of a wider global trend for poor nutritional quality in the OOH food sector. The restaurants included in our study are owned by multinational companies operating on a global scale, therefore our findings can provide insight into the nutritional quality of OOH foods beyond a UK context. We evaluated all three reduction programmes within a single study, allowing us to compare adherence across the targets, both by individual company and overall. By providing an overview of restaurants’ whole menus, we were able to demonstrate that menu items offered by restaurants with similar menu portfolios displayed variable adherence to the reduction targets. Therefore, companies should not have to change the types of foods they offer in order to improve the nutritional quality of their menus, making the shift towards a healthier OOH sector a more achievable goal for industry. We were not able to account for heterogeneity in item-level sales due to the lack of accessible sales data from the OOH sector. It is possible that healthier menu items (items that did meet their sugar, salt, and calorie targets) are responsible for a smaller proportion of sales, thus making little difference to diet-related health outcomes. The technical guidance for the reduction targets highlights that applying the targets based on sales-weighted averages is the gold standard approach, however, this is not always possible, particularly for the OOH sector (noted in the Government’s salt reduction progress report [38]). Greater transparency from companies in regards to the proportion of their sales that come from healthy and less healthy foods would permit a more holistic analysis, and could be encouraged by governments mandating the reporting of this information. The time and resource constraints of manually collecting data from individual websites limited our sample size of menu items, and meant that the completeness and accuracy of the data was largely dependent on restaurants’ own reporting. For example, we found instances of Toby Carvery underreporting kcal (detailed further in S2 Table), which we left as reported as we did not have the capacity to laboratory test all menu items to verify the information. Only five restaurants reported serving size for all of their menu items, therefore we often had to use subcategory average serving sizes to calculate per serving or per 100 g information where missing, which may have differed to the true value given the variation we observed in nutritional content within subcategories. Not all menu items would have been included for each company, for example, we evaluated Burger King’s online menu rather than their more extensive ‘Nutrition Explorer’, due to the less complete nutritional information in the latter, and Subway now (as of 2025) have ‘Cookies and Sweet Treats’ on their online menu which was not published during our data collection. These limitations reflect the lack of standardisation in reporting nutritional information across the OOH sector, highlighting another gap that government regulation could address. We only assessed adherence at one time point, therefore we cannot determine which restaurants or subcategories have shown the most or least improvement in nutritional quality over time, which could provide further insight into which areas of the sector need more stringent monitoring. Our data collection took place predominantly in February/March 2024, so it is possible that we would have seen better adherence to the salt targets if we had collected data later in 2024 (as the targets had to be met in 2024, with no specification of which day or month), and better adherence to the calorie targets if we had collected data later in 2025 (as the targets had to be met in 2025). However, with the lack of more recent governmental progress reports to monitor target adherence, our study provides a useful benchmark for how restaurants were performing against the targets in early 2024, which future monitoring of the targets could compare against. The NHS 10 Year Health Plan for England [22] outlines plans for the introduction of mandatory reporting of healthy sales from large companies, with a further proposal to use this reporting to inform mandatory targets for healthy sales. Our study highlights that there is low adherence in the OOH sector with current voluntary regulation, and that monitoring of target adherence is largely limited by the lack of available sales data, which the two policies proposed in the NHS 10 Year Plan would directly address. This study highlights the importance of mandatory reporting and targets, demonstrating overall low adherence to the targets in these voluntary programmes, with research from other countries also evidencing the increased effectiveness of mandatory (e.g., maximum limits or required declaration of salt content) versus voluntary (e.g., suggested limit on salt content) nutrition policies in inciting reformulation [39]. With no regular publication of progress reports to monitor companies’ adherence to the targets, there are minimal incentives for the companies to work towards them. More regular and granular reporting of adherence to the targets by individual restaurants in the OOH sector, alongside the introduction of mandatory reporting for healthy sales for example, might lead to greater public scrutiny and thus greater adherence with the targets. While the voluntary nature of the targets may be a contributor to the low target adherence, our study was able to demonstrate that restaurants with varying cuisines are able to meet the targets, highlighting their attainability across the OOH sector. The overall lower adherence to the sugar targets compared to the calorie and salt targets could be a result of the deadline for the sugar targets being in 2020, versus 2024/5 for the calorie and salt targets. It is possible that adherence to the sugar targets has dropped since the deadline was reached, as there has been no subsequent revisions to the sugar targets to encourage food companies to maintain their adherence to the programme. Alternatively, the sugar reduction targets were set per 100 g, while the calorie and salt targets were primarily per serving (all calorie targets per serving, OOH-specific salt targets per serving), meaning progress towards the calorie and salt targets could be made through reducing menu items’ serving size, while the sugar targets could only be met through reformulation, which is potentially more time and resource intensive. Unlike the retail sector, there is no mandatory reporting standard for nutritional information per 100 g or serving size for the OOH sector, resulting in inconsistent and limited data being used to monitor adherence to the targets. Mandating standardised nutrition reporting would improve transparency in the OOH sector and make tracking compliance to the targets easier and more accurate, potentially inciting greater compliance from companies. Modelling work by Shangguan and colleagues [7] found the drop from 100% to 50% industry compliance with US National Salt and Sugar Reduction targets approximately halved the averted cardiovascular disease events, QALYs gained, and net savings for healthcare observed over a lifetime, highlighting the importance of meeting the targets in full in order to bring about substantial benefits to public health. Changes could be made to the targets themselves to encourage greater adherence. All menu items had to be manually categorised into the relevant target categories to compare their nutrient content to the target value, which was time and resource-intensive, and open to subjectivity and error. Using comparable categories across targets, or setting targets based on a single holistic measure of nutritional quality rather than individual nutrients (e.g., using the UK Nutrient Profiling Model), could remove some of the burden of self-monitoring placed onto companies, permitting greater progress. The new policy proposals in the NHS 10 Year Plan suggest there is an appetite from the UK Government to impose stricter regulation on the food sector, highlighting the importance of research to inform careful and realistic target setting, particularly within under-regulated areas of the food sector such as OOH. In conclusion, our findings suggest there has been low adherence against the UK Government’s reduction targets from the OOH sector, which could suggest that mandatory regulations may be a more effective approach to improving the nutritional quality of OOH food. While we found menu items from certain restaurant types to perform worse against the targets than others, menu items from restaurants with similar portfolios were also found to vary in target adherence, suggesting that companies should not have to change the focus of their menus in order to meet the targets, making them more attainable. This study highlights the need for standardised reporting of nutritional and serving size information from the OOH sector, alongside accessible sales data, to aid monitoring companies’ performance against the targets, and in turn incite greater adherence from industry with the reduction programmes. Supporting information S1 File. Study protocol. https://doi.org/10.1371/journal.pmed.1004681.s001 (PDF) S1 Table. STROBE checklist. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) Statement: Guidelines for Reporting Observational Studies von Elm E, Altman DG, Egger M, Pocock SJ, Gøtzsche PC, et al. (2007) The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) Statement: Guidelines for Reporting Observational Studies. PLOS Medicine 4(10): e296. https://doi.org/10.1371/journal.pmed.0040296. https://doi.org/10.1371/journal.pmed.1004681.s002 (PDF) S2 Table. Overview of data collection approach and completeness of collected data, for each restaurant. Restaurants in descending order by number of menu items. https://doi.org/10.1371/journal.pmed.1004681.s003 (PDF) S3 Table. Categorisation criteria for the 12 broad food subcategories. https://doi.org/10.1371/journal.pmed.1004681.s004 (PDF) S4 Table. The range of calorie, salt (per 100 g and per serving), and sugar target values set for menu items within each subcategory. https://doi.org/10.1371/journal.pmed.1004681.s005 (PDF) S5 Table. The number of products per subcategory where the subcategory mean serving size had to be used to calculate either per 100 g or per serving nutrient information. Subcategories are listed in descending order by the proportion of products belonging to that category where the subcategory mean serving size had to be used. https://doi.org/10.1371/journal.pmed.1004681.s006 (PDF) S6 Table. The number of products per restaurant where the subcategory mean serving size had to be used to calculate either per 100 g or per serving nutrient information. Restaurants are listed in descending order by the proportion of products belonging to that category where the subcategory mean serving size had to be used. https://doi.org/10.1371/journal.pmed.1004681.s007 (PDF) S7 Table. The mean, median, and standard deviation, for kcal per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each subcategory. In descending order by Mean kcal per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s008 (PDF) S8 Table. The mean, median, and standard deviation, for Salt per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each subcategory. In descending order by Mean Salt per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s009 (PDF) S9 Table. The mean, median, and standard deviation, for Sugar per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each subcategory. In descending order by Mean Sugar per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s010 (PDF) S10 Table. The mean, median, and standard deviation, for Fat per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each subcategory. In descending order by Mean Fat per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s011 (PDF) S11 Table. The mean, median, and standard deviation, for kcal per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each restaurant. In descending order by Mean kcal per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s012 (PDF) S12 Table. The mean, median, and standard deviation, for Salt per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each restaurant. In descending order by Mean Salt per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s013 (PDF) S13 Table. The mean, median, and standard deviation, for Sugar per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each restaurant. In descending order by Mean Sugar per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s014 (PDF) S14 Table. The mean, median, and standard deviation, for Fat per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each restaurant. In descending order by Mean Fat per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s015 (PDF) S15 Table. The mean, median, and standard deviation, for kcal per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each restaurant group. In descending order by Mean kcal per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s016 (PDF) S16 Table. The mean, median, and standard deviation, for Salt per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each restaurant group. In descending order by Mean Salt per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s017 (PDF) S17 Table. The mean, median, and standard deviation, for Sugar per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each restaurant group. In descending order by Mean Sugar per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s018 (PDF) S18 Table. The mean, median, and standard deviation, for Fat per 100 g, per recommended serving, and per subcategory average serving, across all menu items in each restaurant group. In descending order by Mean Fat per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s019 (PDF) S19 Table. The proportion of menu items meeting sugar, salt, calorie, and all applicable targets, for each restaurant group. In descending order by proportion of menu items meeting all applicable targets. https://doi.org/10.1371/journal.pmed.1004681.s020 (PDF) S20 Table. Mean nutrient content per 100 g and per serving for restaurants with limited time menu items, including (as per primary analysis) and excluding the limited time offer items. https://doi.org/10.1371/journal.pmed.1004681.s021 (PDF) S21 Table. The proportion of menu items meeting sugar, salt, calorie, and all applicable targets, for restaurants with limited time menu items, including and excluding limited time offer items. https://doi.org/10.1371/journal.pmed.1004681.s022 (PDF) S22 Table. Mean nutrient content per 100 g for each subcategory when the subcategory average (as per the primary analysis), lower quartile, and upper quartile, were used to replace missing serving size. Subcategories are listed in descending order by mean kcal per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s023 (PDF) S23 Table. Mean nutrient content per serving for each subcategory when the subcategory average (as per the primary analysis), lower quartile, and upper quartile, were used to replace missing serving size. Subcategories are listed in descending order by mean kcal per serving. https://doi.org/10.1371/journal.pmed.1004681.s024 (PDF) S24 Table. Mean nutrient content per 100 g for each restaurant when the subcategory average (as per the primary analysis), lower quartile, and upper quartile, were used to replace missing serving size. Restaurants are listed in descending order by mean kcal per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s025 (PDF) S25 Table. Mean nutrient content per serving for each restaurant when the subcategory average (as per the primary analysis), lower quartile, and upper quartile, were used to replace missing serving size. Restaurants are listed in descending order by mean kcal per 100 g. https://doi.org/10.1371/journal.pmed.1004681.s026 (PDF) S26 Table. The proportion of menu items meeting sugar, salt, calorie, and all applicable targets for each subcategory, when the subcategory average (as per the primary analysis), lower quartile, and upper quartile, were used to replace missing serving size. Subcategories are listed in descending order by proportion of menu items meeting all applicable targets when the average serving size was used to replace missing values. https://doi.org/10.1371/journal.pmed.1004681.s027 (PDF) S27 Table. The proportion of menu items meeting sugar, salt, calorie, and all applicable targets for each restaurant, when the subcategory average (as per the primary analysis), lower quartile, and upper quartile, were used to replace missing serving size. Restaurants are listed in descending order by proportion of menu items meeting all applicable targets when the average serving size was used to replace missing values. https://doi.org/10.1371/journal.pmed.1004681.s028 (PDF) Acknowledgments The views expressed are those of the author(s) and not necessarily those of the NIHR or the Department of Health and Social Care.
Conditional cash transfer and mortality among interpersonal violence victims: A cohort studyBonfim, Camila;Alves, Flávia;Barreto, Maurício L.;Patel, Vikram;Machado, Daiane Borges
doi: 10.1371/journal.pmed.1004673pmid: 42090627
Background Interpersonal violence is a significant public health issue, increasing mortality risks for those affected. While Cash Transfer Programs offer health benefits, their role in addressing the needs of interpersonal violence victims remains unclear. This study aims to examine the association between Brazil’s Bolsa Família Program (BFP) participation and reduced mortality rates among interpersonal violence victims. Methods and findings This cohort study was conducted using data from 100 Million Brazilian Cohort, which were linked with interpersonal violence registries (2011−2015). All individuals with a record of interpersonal violence following their registration in Brazil’s primary social assistance system during the study period were included. The primary outcome was overall mortality, while secondary outcomes comprised deaths due to natural and unnatural causes, as recorded in the Mortality Information System (SIM) and classified according to the International Classification of Diseases, 10th Revision (ICD-10). We used Cox proportional hazards models with propensity score-based method to analyze overall mortality and competing risk models to assess specific causes of death, estimating the association between BFP participation and mortality rates. A total of 29,075 individuals who were victims of interpersonal violence were followed throughout the five-year period. A total of 990 individuals died from overall causes. BFP participation was associated with an 18% reduction in overall mortality rate (hazard ratio, HR 0.82, 95% CI [0.70,0.95]; p = 0.011) and a 66% reduction in mortality rate from natural causes (HR 0.34 [95% CI 0.28, 0.41]; p < 0.001). This sample includes only individuals who seek healthcare services, which may overrepresent more severe cases of interpersonal violence. Conclusions The association between BFP participation and lower mortality rates, especially from natural causes, among interpersonal violence victims highlights that such programs may be associated with reductions in poverty, improvements in health outcomes, and increased survival in vulnerable populations. Why was this study done? Interpersonal violence is a major public health problem associated with increased risk of morbidity and mortality. Although cash transfer programs are known to improve health outcomes and reduce social inequalities, their role in addressing the specific needs of victims of interpersonal violence remains unclear. What did the researchers do and find? This cohort study followed nearly 30,000 victims of interpersonal violence from the 100 Million Brazilian Cohort over a 5-year period. There was an association between the Bolsa Família Program (BFP) and a reduction in all-cause mortality (18%) This was primarily driven by a significant decrease in deaths from natural causes (66%). What do these findings mean? Our findings underscore the potential of cash transfers in reducing mortality rate among people victimized by interpersonal violence even when these programs were not specifically designed for them. Cash transfer programs could be used as a tool for prevention, reducing the mortality burden among victims of interpersonal violence. Hence, it is imperative for health administrators to advocate for the allocation of funds to support victims of interpersonal violence while concurrently addressing their health needs. The main limitation of this study is that it was restricted to individuals who sought healthcare services, which likely represent the most severe cases of interpersonal violence and may have led to misclassification bias. Introduction Interpersonal violence is a multifactorial public health problem encompassing a wide range of relational contexts and manifestations. It may involve family members, intimate partners, friends, acquaintances, or strangers, and includes various forms such as child maltreatment, youth violence, violence against women, and elder abuse [1]. Its consequences extend across multiple domains, including physical, psychological, financial, sexual, neglect-related, and other dimensions [2,3]. It affects millions of people worldwide, ranking as the third leading cause of disability-adjusted life years and the second leading cause of years of life lost due to premature mortality in 2019 [4]. In addition to experiencing violence, these individuals also face reduced life expectancy [4]. However, what is associated with lower mortality among individuals exposed to violence is still not well-established [5]. This is a global issue, with a higher impact in countries such as Brazil, where the mortality rate due to violence is among the highest worldwide, reaching 22.4 per 100,000 inhabitants [5]. Although previous studies have estimated the global prevalence of interpersonal violence [4,6,7], assessing its health burden remains challenging due to its complexity and frequent underreporting, as many victims do not seek healthcare after violent episodes [2]. Interpersonal violence has been associated with social inequalities related to poverty, such as limited employment opportunities, restricted access to education, and gender, racial, and income disparities [2,5], highlighting heterogeneity according to the vulnerability of the at-risk population. These inequalities further increase vulnerability to the negative health outcomes observed in victims of interpersonal violence [8]. They are at higher risk of non-communicable diseases such as cardiovascular disease, cancer, respiratory problems, and diabetes as well as psychiatric disorders [2,9]. Exposure to violence can also lead to damage to the nervous, endocrine, and immune systems in addition to genetic changes associated with the environment [9]. Furthermore, the victims commonly have worse health habits such as alcohol abuse, tobacco use, and physical inactivity [2]. These factors contribute to increased mortality risk [4,6,10]. Considering interpersonal violence an important risk factor for poor health across the life course, preventing this public health problem has been a goal of various initiatives [2,6]. Programs that reduce social inequalities such as Cash Transfer Programs (CTPs) have had an association with decreased violence risk [2,5]. CTPs are social protection programs designed to reduce poverty and vulnerability through cash-based benefits [11]. They can be unconditional or conditional, the latter can include attendance at health appointments, access to social services, school attendance when the family have children or adolescents, prenatal appointments for pregnant and mandatory vaccinations for children [11]. CTPs such as Bolsa Família Program (BFP) have been associated with improved health outcomes among beneficiaries [12]. This is one of the largest conditional cash transfer initiatives globally [12]. Since its implementation in 2004, it has played a significant role in reducing poverty and extreme poverty among Brazilian families [13]. Although originally designed as a social equity policy, evidence indicates that the BFP is also associated with improvements in health outcomes and lower mortality [12]. These associations may reflect not only the income transfer itself but also from the program’s conditionalities [12]. By providing financial support to families, the program contributes to improved living conditions and fosters engagement with healthcare services and school attendance [12]. It is already known that these programs, especially conditional CTPs, may be associated with lower mortality among people hospitalized with psychiatric disorders [14], homicide [15], and suicide [16] rates in the general population. However, whether conditional CTPs would reduce mortality rates among people victimized by interpersonal violence, remains unanswered. The role of BFP in health and mortality has been studied, and its main mechanisms have been explored [12,17,18]. Regarding specific populations, such as individuals exposed to interpersonal violence, we believe that the role of BFP operates through secondary prevention, by facilitating referral to health services, as well as through the well-documented improvement of living conditions reported in other studies [12,19]. We hypothesized that the reduction in mortality may occur by improving access to health services and mitigating physiological stress responses to violence as well as facilitating connection to other social protection programs, promoting social inclusion, better living conditions, and improved physical and mental health [2,5]. These individuals face dual vulnerability, as they live in poverty [20] and have been victims of violence [4,5] both strong risk factors for mortality. The literature has significant limitations, as none of the available studies have evaluated competing risks of death in this specific population. Assessing this is crucial, as mortality risks vary depending on the context [21]. Therefore, this study aimed to test the association between receiving benefits from the Brazilian conditional CTP, the Bolsa Familia Program (BFP), and decreased mortality rates among victims of interpersonal violence. We hypothesized that BFP would be associated with a reduction in mortality rates, controlling for competing death risks, among these individuals. Method Study design, setting, and data source This study followed a prospective protocol, with detailed descriptions of the study design and procedures published elsewhere [22]. This study used a cohort design from the 100 million Brazilian Cohort baseline, a dynamic cohort developed by the Center for Data Integration and Knowledge in Health (CIDACS) [23,24] based on a non-deterministic linkage between administrative health and social assistance databases [25]. The 100 million Brazilian Cohort baseline was developed to investigate social determinants and the impact of social programs and policies on various health contexts in Brazil [24]. This dataset is based on information from over 131 million individuals who were registered between 2001 and 2,018 in CadÚnico, the primary system for applying for social benefits in Brazil which include the BFP [23]. It includes socio-economic and sociodemographic information from poorer Brazilian individuals and their families who apply for social programs such as BFP [24]. To qualify and register with CadÚnico, families must have a per capita income of up to half a minimum wage or a total family income of up to three minimum wages [24]. We selected a subset of just over 23 million for this study based on the period during which all data on interpersonal violence and BFP were available (January, 1, 2011, to December, 31, 2015). The methods and analyses were detailed in the research protocol published previously [22]. This study is reported as the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guideline (S1 STROBE Checklist). Interpersonal violence records were extracted from two administrative databases, the National Disease Notification System (SINAN) (Method A in S1 Appendix) and the Hospitalization Information System (SIH) (Method A in S1 Appendix). SINAN is the system for mandatory reporting of interpersonal violence since 2011 [26,27]. This system includes physical, sexual, psychological, psychological abuses, neglect and other relational contexts perpetrated by family members, intimate partners, friends, or strangers [26]. The data is registered by professionals working in health facilities, and it has improved its reporting over the years. Individuals hospitalized due to interpersonal violence were also identified through the SIH. SIH is the system responsible for recording most of Brazilian hospital admissions [28,29], including those due to violent causes, according to the International Classification of Diseases, 10th Revision (ICD-10) [30]. In this study, we included hospital admissions for aggression (codes X85-Y09) which are recorded as secondary diagnoses, as well as other external causes. However, to capture all possible hospitalizations related to this cause within the system, we also include records where it appears as the primary diagnosis. We also used Mortality Information System (SIM), which comprises all deaths in the country and uses a mandatory certificate internationally recognized as a high-quality system [31] (Method A in S1 Appendix). SIM registers the cause of death according to the ICD-10 [32]. All the health systems use standardized forms completed by health professionals [22]. Finally, this study extracted information from the BFP database, the main poverty and extreme poverty alleviation program implemented by the Brazilian government in 2004 [24]. All participants benefiting from the BFP are registered in the CadÚnico. These databases were linked using a tool developed by CIDACS, which uses five identifiers (name, gender, year of birth, name of the mother and municipality of residency) [22]. Reliability analyses showed very high sensitivity and specificity of the linkage [25] (SMethod B in S1 Appendix). Additional information on data governance and the linkage process has been published elsewhere [25]. This study was approved by the ethics committees of the Federal University of Bahia (UFBA – registration number: 1023107) and Gonçalo Muniz Institute at the Oswaldo Cruz Foundation (registration number: 1.612.302). Participants Individuals from the 100 million Brazilian cohort who were reported in SINAN (N = 67,487) or SIH (N = 11,286) as having experienced interpersonal violence (N = 78,038) from January 1, 2011, to December 31, 2015. The SINAN system collects all violence-related information reported in the country. To ensure an unbiased sample, we focused on those who became victims of violence without prior exposure to program intervention. Therefore, we included only individuals who were registered in CadÚnico after a violent event (N = 30,323). Individuals who were already receiving the benefit before the violent event may have a lower chance of hospitalization or notification of violence, considering the association between receiving the benefit and improved health [14–18,32,33]. As individuals with multiple records of interpersonal violence may have a higher risk of mortality [2], we removed duplicate entries (1,174 cases, 4%) to mitigate biases in the association measurement. Finally, we removed individuals who had inconsistent data on dates, such as death before CadÚnico registration or the same start and end date for receiving BFP (n = 74, < 1% of the total sample). Therefore, the study included 29,075 participants, of whom 14,856 (51%) were BFP recipients (Fig 1). Download: PNG larger image TIFF original image Fig 1. Flowchart of the study population. Abbreviations: SIH, Hospitalization Information System (acronymous in Portuguese); SINAN, National Disease Notification System (acronymous in Portuguese); BFP, Bolsa Familia Program. https://doi.org/10.1371/journal.pmed.1004673.g001 Follow-up For the subset of BFP beneficiaries, (A) individuals were followed from the moment they received the benefit after notification/hospitalization, starting on January 1, 2011. Follow-up ended at the earliest occurrence of either: (B) the individual’s death from any cause, or (C) December 31, 2015. For the subset of non-beneficiaries, (A) follow-up began from the moment they were registered in CadÚnico following notification/hospitalization, starting on January 1, 2011. Follow-up ended at the earliest occurrence of either: (B) the individual’s death from any cause, or (C) December 31, 2015. Variables Exposure. The Brazilian cash transfer Program BFP aims to lift people out of extreme poverty or poverty [34]. In addition to providing income to families in poverty, the BFP aims to integrate public policies, improve access to basic rights such as health and education [34]. The benefits range from BRL 41.00 (USD 10.00) to BRL 300.00 (USD 75.00) per individual, based on 2015 values adjusted for inflation [22]. Eligibility depends on a monthly household income of less than BRL 70.00 (USD 17.00), or BRL 140.00 (USD 34.00) for households with a child, teenager, or pregnant woman [23]. Recipients of the BFP must meet specific conditionalities to maintain their benefits, including minimum school attendance, keeping up with vaccinations, and monitoring the growth of young children [23]. Additionally, pregnant or breastfeeding women are required to follow a specified health and nutrition protocol [23]. The conditionalities are grounded in the notion that tying benefits to constructive behaviors can further improve families’ chances of breaking free from poverty through educational attainment or health improvements [23]. We considered the beneficiary group (exposed) individuals who receive the benefit after the interpersonal violence event over the period of the study. The non-beneficiary group (unexposed) was composed of individuals who were also victims of interpersonal violence and did not receive the benefit during the same period. We highlighted that both groups were identified through the CadÚnico. This system includes individuals who meet the criteria for poverty or extreme poverty and who may apply for various social benefits such as social benefits for people with disabilities, housing for low-income families, BFP, among others [19]. BFP has more stringent eligibility standards and focuses on families that constitute a defined subset within CadÚnico [23]. Therefore, not all individuals meet the eligibility criteria for participation in the BFP. Outcomes. Our primary outcome was a record of overall mortality in the SIM. Secondary outcomes included natural and unnatural causes of death. Natural causes of death were defined as all causes of mortality, except for suicide and external causes (X60-Y09), according to ICD-10 [30]. Unnatural causes included external causes such as accidents, homicide, suicide, and other external causes of death. Statistical analysis Following our study protocol [22] and other studies using the 100 million Brazilian cohort [14–18,32,33], we used a propensity score (PS)-based method [35,36] to identify the association between BFP and reduction of mortality rates. By applying stabilized IPTW, we aimed to balance differences between groups [35]. Although this cohort contains complete data, beneficiaries and non-beneficiaries may still differ because eligibility for the BFP depends on income and additional socioeconomic criteria that may introduce imbalance [23]. The use of PS estimation allowed us to model the probability of receiving the BFP based on observable baseline sociodemographic characteristics, thereby reducing potential selection bias, especially given that eligibility could not be fully determined using income cutoff alone. This methodological approach is consistent with prior studies [14–18,32,33]. First, we ran a logistic regression to estimate the PS using baseline covariates, which are associated with BFP according to the literature review [14–18,32,33] (Table A in S1 Appendix). The following covariates were considered to estimate the PS: sex, age, education level, race, location of residence (rural × urban), living conditions (which included water supply, waste, sanitation, and construction materials), isolation (single person in the household), Brazilian regions of residence, and year of CadÚnico registration. We also evaluated the common support graph and compared the range of PS between the BFP and non-BFP groups (Fig A in S1 Appendix). Second, the stabilized Inverse Probability of Treatment Weighting (IPTW) approach was applied [35]. We applied the stabilized IPTW weights for non-beneficiaries using the formula (1 − Pt)/(1 − Psmul) and for beneficiaries using the formula Pt/Psmul, where “Pt” is the marginal probability of treatment in the population and “Psmul” is the PS obtained from the multivariable logistic regression adjusted for covariates [35]. To improve the accuracy and the precision of the estimation, in addition to stabilized IPTW, weight truncation was performed based on distribution for the 99th percentile [14,32,35]. Third, standardized differences among covariates of beneficiaries and non-beneficiaries before and after applying stabilized IPTW weighting were calculated to assess covariate balance, and changes in absolute values greater than 10% were considered acceptable [36]. Finally, we used stabilized IPTW-weighted Cox proportional hazards regression analysis, adjusted for the year of notification or hospitalization due to interpersonal violence, to examine the association between BFP and the overall mortality rate in the final model [35,36]. To address the potential bias of overestimating the risk of the event of interest due to competing risks in traditional Cox regression, we employed competing risk models using the Fine-Gray subdistribution hazard model [37]. For this analysis, each specific outcome was considered a failure, other causes were treated as competing risks, and individuals who remained alive were censored. A stratified analysis was also performed by sex, considering that gender differences [2,5] may influence mortality risk (TableB in S1 Appendix). Sensitivity analyses We conducted the following analysis to assess the robustness of the results. First, to test whether our results are not an artifact created with the stabilized IPTW weights, we used another PS-based method, the Kernel Matching approach [35] (Table 3). Second, we repeated the analysis using Poisson models with Incidence Rate Ratios estimation and 95% CI (Table C in S1 Appendix). Third, to address potential selection bias, we conducted the analysis with the entire population victimized by interpersonal violence during the study period, including those registered in CadÚnico before the violent event (Table D in S1 Appendix). The main analyses were additionally replicated by treating missing data as a separate category, with the aim of assessing potential biases arising from information loss (Table E in S1 Appendix). Furthermore, we ran crude Cox model without propensity adjustment to evaluate if the results are an artifact created with the IPTW weights (Table F in S1 Appendix). Fourth, we ran Cox proportional hazards models with time-varying covariates [38] considering BFP as a time-varying covariate to account for possible differences in the duration of benefit receipt and immortal time bias [38] (Table G in S1 Appendix). Fifth, to assess whether age modified the association between exposure and mortality [39], we compared models with and without an age–exposure interaction term using the Akaike (AIC) and Bayesian (BIC) Information Criteria. According to established guidelines, an AIC difference of less than two indicates no substantial improvement in model fit [40] (Table H in S1 Appendix). Finally, to check the broad age range within the 25–59-year category of age, we conducted an additional sensitivity analysis introducing a more granular age stratification (Table I in S1 Appendix). Stata version 15.0 was used for the data analysis. Study design, setting, and data source This study followed a prospective protocol, with detailed descriptions of the study design and procedures published elsewhere [22]. This study used a cohort design from the 100 million Brazilian Cohort baseline, a dynamic cohort developed by the Center for Data Integration and Knowledge in Health (CIDACS) [23,24] based on a non-deterministic linkage between administrative health and social assistance databases [25]. The 100 million Brazilian Cohort baseline was developed to investigate social determinants and the impact of social programs and policies on various health contexts in Brazil [24]. This dataset is based on information from over 131 million individuals who were registered between 2001 and 2,018 in CadÚnico, the primary system for applying for social benefits in Brazil which include the BFP [23]. It includes socio-economic and sociodemographic information from poorer Brazilian individuals and their families who apply for social programs such as BFP [24]. To qualify and register with CadÚnico, families must have a per capita income of up to half a minimum wage or a total family income of up to three minimum wages [24]. We selected a subset of just over 23 million for this study based on the period during which all data on interpersonal violence and BFP were available (January, 1, 2011, to December, 31, 2015). The methods and analyses were detailed in the research protocol published previously [22]. This study is reported as the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guideline (S1 STROBE Checklist). Interpersonal violence records were extracted from two administrative databases, the National Disease Notification System (SINAN) (Method A in S1 Appendix) and the Hospitalization Information System (SIH) (Method A in S1 Appendix). SINAN is the system for mandatory reporting of interpersonal violence since 2011 [26,27]. This system includes physical, sexual, psychological, psychological abuses, neglect and other relational contexts perpetrated by family members, intimate partners, friends, or strangers [26]. The data is registered by professionals working in health facilities, and it has improved its reporting over the years. Individuals hospitalized due to interpersonal violence were also identified through the SIH. SIH is the system responsible for recording most of Brazilian hospital admissions [28,29], including those due to violent causes, according to the International Classification of Diseases, 10th Revision (ICD-10) [30]. In this study, we included hospital admissions for aggression (codes X85-Y09) which are recorded as secondary diagnoses, as well as other external causes. However, to capture all possible hospitalizations related to this cause within the system, we also include records where it appears as the primary diagnosis. We also used Mortality Information System (SIM), which comprises all deaths in the country and uses a mandatory certificate internationally recognized as a high-quality system [31] (Method A in S1 Appendix). SIM registers the cause of death according to the ICD-10 [32]. All the health systems use standardized forms completed by health professionals [22]. Finally, this study extracted information from the BFP database, the main poverty and extreme poverty alleviation program implemented by the Brazilian government in 2004 [24]. All participants benefiting from the BFP are registered in the CadÚnico. These databases were linked using a tool developed by CIDACS, which uses five identifiers (name, gender, year of birth, name of the mother and municipality of residency) [22]. Reliability analyses showed very high sensitivity and specificity of the linkage [25] (SMethod B in S1 Appendix). Additional information on data governance and the linkage process has been published elsewhere [25]. This study was approved by the ethics committees of the Federal University of Bahia (UFBA – registration number: 1023107) and Gonçalo Muniz Institute at the Oswaldo Cruz Foundation (registration number: 1.612.302). Participants Individuals from the 100 million Brazilian cohort who were reported in SINAN (N = 67,487) or SIH (N = 11,286) as having experienced interpersonal violence (N = 78,038) from January 1, 2011, to December 31, 2015. The SINAN system collects all violence-related information reported in the country. To ensure an unbiased sample, we focused on those who became victims of violence without prior exposure to program intervention. Therefore, we included only individuals who were registered in CadÚnico after a violent event (N = 30,323). Individuals who were already receiving the benefit before the violent event may have a lower chance of hospitalization or notification of violence, considering the association between receiving the benefit and improved health [14–18,32,33]. As individuals with multiple records of interpersonal violence may have a higher risk of mortality [2], we removed duplicate entries (1,174 cases, 4%) to mitigate biases in the association measurement. Finally, we removed individuals who had inconsistent data on dates, such as death before CadÚnico registration or the same start and end date for receiving BFP (n = 74, < 1% of the total sample). Therefore, the study included 29,075 participants, of whom 14,856 (51%) were BFP recipients (Fig 1). Download: PNG larger image TIFF original image Fig 1. Flowchart of the study population. Abbreviations: SIH, Hospitalization Information System (acronymous in Portuguese); SINAN, National Disease Notification System (acronymous in Portuguese); BFP, Bolsa Familia Program. https://doi.org/10.1371/journal.pmed.1004673.g001 Follow-up For the subset of BFP beneficiaries, (A) individuals were followed from the moment they received the benefit after notification/hospitalization, starting on January 1, 2011. Follow-up ended at the earliest occurrence of either: (B) the individual’s death from any cause, or (C) December 31, 2015. For the subset of non-beneficiaries, (A) follow-up began from the moment they were registered in CadÚnico following notification/hospitalization, starting on January 1, 2011. Follow-up ended at the earliest occurrence of either: (B) the individual’s death from any cause, or (C) December 31, 2015. Variables Exposure. The Brazilian cash transfer Program BFP aims to lift people out of extreme poverty or poverty [34]. In addition to providing income to families in poverty, the BFP aims to integrate public policies, improve access to basic rights such as health and education [34]. The benefits range from BRL 41.00 (USD 10.00) to BRL 300.00 (USD 75.00) per individual, based on 2015 values adjusted for inflation [22]. Eligibility depends on a monthly household income of less than BRL 70.00 (USD 17.00), or BRL 140.00 (USD 34.00) for households with a child, teenager, or pregnant woman [23]. Recipients of the BFP must meet specific conditionalities to maintain their benefits, including minimum school attendance, keeping up with vaccinations, and monitoring the growth of young children [23]. Additionally, pregnant or breastfeeding women are required to follow a specified health and nutrition protocol [23]. The conditionalities are grounded in the notion that tying benefits to constructive behaviors can further improve families’ chances of breaking free from poverty through educational attainment or health improvements [23]. We considered the beneficiary group (exposed) individuals who receive the benefit after the interpersonal violence event over the period of the study. The non-beneficiary group (unexposed) was composed of individuals who were also victims of interpersonal violence and did not receive the benefit during the same period. We highlighted that both groups were identified through the CadÚnico. This system includes individuals who meet the criteria for poverty or extreme poverty and who may apply for various social benefits such as social benefits for people with disabilities, housing for low-income families, BFP, among others [19]. BFP has more stringent eligibility standards and focuses on families that constitute a defined subset within CadÚnico [23]. Therefore, not all individuals meet the eligibility criteria for participation in the BFP. Outcomes. Our primary outcome was a record of overall mortality in the SIM. Secondary outcomes included natural and unnatural causes of death. Natural causes of death were defined as all causes of mortality, except for suicide and external causes (X60-Y09), according to ICD-10 [30]. Unnatural causes included external causes such as accidents, homicide, suicide, and other external causes of death. Exposure. The Brazilian cash transfer Program BFP aims to lift people out of extreme poverty or poverty [34]. In addition to providing income to families in poverty, the BFP aims to integrate public policies, improve access to basic rights such as health and education [34]. The benefits range from BRL 41.00 (USD 10.00) to BRL 300.00 (USD 75.00) per individual, based on 2015 values adjusted for inflation [22]. Eligibility depends on a monthly household income of less than BRL 70.00 (USD 17.00), or BRL 140.00 (USD 34.00) for households with a child, teenager, or pregnant woman [23]. Recipients of the BFP must meet specific conditionalities to maintain their benefits, including minimum school attendance, keeping up with vaccinations, and monitoring the growth of young children [23]. Additionally, pregnant or breastfeeding women are required to follow a specified health and nutrition protocol [23]. The conditionalities are grounded in the notion that tying benefits to constructive behaviors can further improve families’ chances of breaking free from poverty through educational attainment or health improvements [23]. We considered the beneficiary group (exposed) individuals who receive the benefit after the interpersonal violence event over the period of the study. The non-beneficiary group (unexposed) was composed of individuals who were also victims of interpersonal violence and did not receive the benefit during the same period. We highlighted that both groups were identified through the CadÚnico. This system includes individuals who meet the criteria for poverty or extreme poverty and who may apply for various social benefits such as social benefits for people with disabilities, housing for low-income families, BFP, among others [19]. BFP has more stringent eligibility standards and focuses on families that constitute a defined subset within CadÚnico [23]. Therefore, not all individuals meet the eligibility criteria for participation in the BFP. Outcomes. Our primary outcome was a record of overall mortality in the SIM. Secondary outcomes included natural and unnatural causes of death. Natural causes of death were defined as all causes of mortality, except for suicide and external causes (X60-Y09), according to ICD-10 [30]. Unnatural causes included external causes such as accidents, homicide, suicide, and other external causes of death. Statistical analysis Following our study protocol [22] and other studies using the 100 million Brazilian cohort [14–18,32,33], we used a propensity score (PS)-based method [35,36] to identify the association between BFP and reduction of mortality rates. By applying stabilized IPTW, we aimed to balance differences between groups [35]. Although this cohort contains complete data, beneficiaries and non-beneficiaries may still differ because eligibility for the BFP depends on income and additional socioeconomic criteria that may introduce imbalance [23]. The use of PS estimation allowed us to model the probability of receiving the BFP based on observable baseline sociodemographic characteristics, thereby reducing potential selection bias, especially given that eligibility could not be fully determined using income cutoff alone. This methodological approach is consistent with prior studies [14–18,32,33]. First, we ran a logistic regression to estimate the PS using baseline covariates, which are associated with BFP according to the literature review [14–18,32,33] (Table A in S1 Appendix). The following covariates were considered to estimate the PS: sex, age, education level, race, location of residence (rural × urban), living conditions (which included water supply, waste, sanitation, and construction materials), isolation (single person in the household), Brazilian regions of residence, and year of CadÚnico registration. We also evaluated the common support graph and compared the range of PS between the BFP and non-BFP groups (Fig A in S1 Appendix). Second, the stabilized Inverse Probability of Treatment Weighting (IPTW) approach was applied [35]. We applied the stabilized IPTW weights for non-beneficiaries using the formula (1 − Pt)/(1 − Psmul) and for beneficiaries using the formula Pt/Psmul, where “Pt” is the marginal probability of treatment in the population and “Psmul” is the PS obtained from the multivariable logistic regression adjusted for covariates [35]. To improve the accuracy and the precision of the estimation, in addition to stabilized IPTW, weight truncation was performed based on distribution for the 99th percentile [14,32,35]. Third, standardized differences among covariates of beneficiaries and non-beneficiaries before and after applying stabilized IPTW weighting were calculated to assess covariate balance, and changes in absolute values greater than 10% were considered acceptable [36]. Finally, we used stabilized IPTW-weighted Cox proportional hazards regression analysis, adjusted for the year of notification or hospitalization due to interpersonal violence, to examine the association between BFP and the overall mortality rate in the final model [35,36]. To address the potential bias of overestimating the risk of the event of interest due to competing risks in traditional Cox regression, we employed competing risk models using the Fine-Gray subdistribution hazard model [37]. For this analysis, each specific outcome was considered a failure, other causes were treated as competing risks, and individuals who remained alive were censored. A stratified analysis was also performed by sex, considering that gender differences [2,5] may influence mortality risk (TableB in S1 Appendix). Sensitivity analyses We conducted the following analysis to assess the robustness of the results. First, to test whether our results are not an artifact created with the stabilized IPTW weights, we used another PS-based method, the Kernel Matching approach [35] (Table 3). Second, we repeated the analysis using Poisson models with Incidence Rate Ratios estimation and 95% CI (Table C in S1 Appendix). Third, to address potential selection bias, we conducted the analysis with the entire population victimized by interpersonal violence during the study period, including those registered in CadÚnico before the violent event (Table D in S1 Appendix). The main analyses were additionally replicated by treating missing data as a separate category, with the aim of assessing potential biases arising from information loss (Table E in S1 Appendix). Furthermore, we ran crude Cox model without propensity adjustment to evaluate if the results are an artifact created with the IPTW weights (Table F in S1 Appendix). Fourth, we ran Cox proportional hazards models with time-varying covariates [38] considering BFP as a time-varying covariate to account for possible differences in the duration of benefit receipt and immortal time bias [38] (Table G in S1 Appendix). Fifth, to assess whether age modified the association between exposure and mortality [39], we compared models with and without an age–exposure interaction term using the Akaike (AIC) and Bayesian (BIC) Information Criteria. According to established guidelines, an AIC difference of less than two indicates no substantial improvement in model fit [40] (Table H in S1 Appendix). Finally, to check the broad age range within the 25–59-year category of age, we conducted an additional sensitivity analysis introducing a more granular age stratification (Table I in S1 Appendix). Stata version 15.0 was used for the data analysis. Results We identified 78,038 records of interpersonal violence in either the SIH or SINAN databases between 2011 and 2015 in the baseline of the study. After applying the exclusion and inclusion criteria, the study sample comprised 29,075 individuals who were registered at CadÚnico following a single hospitalization or notification of interpersonal violence (Fig 1). Most of these records were in SINAN (78.75%). Of these, 14,856 (51.11%) were BFP beneficiaries. The average time of receiving BFP after the violent event was 1.71 years (SD = 1.10). Before weighting using stabilized IPTW, beneficiaries, compared to non-beneficiaries, were more likely to be female, aged 25–59 years old, had a high school level of education, were brown/ mixed race, lived in urban areas and in Southeast Brazil, had good household conditions, lived with someone else and were registered at CadÚnico in 2014. After using the stabilized IPTW, the groups became more balanced in most of the covariates (Table 1). Download: PNG larger image TIFF original image Table 1. Study population characteristics overall and by Bolsa Família Program participation before and after applying stabilized IPTW, 2011–2015. https://doi.org/10.1371/journal.pmed.1004673.t001 During the follow-up, 990 deaths were identified, the majority of which were due to natural causes in both BFP groups (BFP: 33.71%; Non-BFP: 66.29%; p < 0.001). BFP beneficiaries had lower mortality rates for: overall mortality (1349.04 [95% CI 1223.55, 1487.39]) and natural causes (796.70 [95% CI 701.64, 904.63]), both estimated for 100,000 inhabitants. Females and younger beneficiaries showed lower mortality rates except for unnatural causes of death (Table 2). Download: PNG larger image TIFF original image Table 2. Mortality rates through receipt of the Bolsa Familia Program, 2011–2015. https://doi.org/10.1371/journal.pmed.1004673.t002 Download: PNG larger image TIFF original image Table 3. Association of Bolsa Família Program participation with mortalities rates, 2011–2015. https://doi.org/10.1371/journal.pmed.1004673.t003 The receipt of BFP benefits was associated with a lower overall mortality (HR, hazard ratio 0.82 [95% CI 0.70,0.95]; p = 0.011). Using the Fine-Gray subdistribution hazard model, the subdistribution hazard ratio for natural causes was 0.34 ([95% CI 0.28,0.41]; p < 0.001), while that for unnatural causes of death was 0.90 ([95% CI 0.67,1.19]; p = 0.461) among the beneficiary group. Therefore, the BFP was associated with a 66% decrease in the natural causes of deaths and a 10% decrease for unnatural causes (non-significant) among individuals who either experienced a competing risk or were still alive. Robustness checks using time-varying and stabilized IPTW Cox models confirmed these findings, suggesting a lower risk of mortality among BFP beneficiaries (Table D in S1 Appendix). The stratified analysis by sex showed that the BFP reduced mortality mainly among women, including both all-cause (HR 0.42 [95% CI 0.34,0.52]; p ≤ 0.001) and natural causes (HR 0.61 [95% CI 0.46,0.80]; p = 0.001), whereas among men, the reduction was observed only for natural causes (HR 0.69 [95% CI 0.52,0.91]; p = 0.009) (Table B in S1 Appendix). In general, the sensitivity analyses confirmed the findings (Tables 3 and A–I in S1 Appendix). Discussion This is the first study to estimate the role of a conditional CTP on mortality among individuals with a documented history of interpersonal violence. Using longitudinal data from over 29,000 low-income individuals followed for up to five years, we observed an 18% reduction in overall mortality and a 66% reduction in mortality from natural causes as well as 10% from unnatural causes (non-significant) among program participants. These findings support the hypothesis that CTPs, primarily designed to alleviate poverty, can also improve the living conditions of individuals who have experienced violence, thereby preventing health deterioration and reducing mortality. While there is some divergence in the literature regarding the role of conditional CTPs on violence [5], the role of such programs on mortality in the general population is well-established [5,41,42]. A multinational study that used data from 7 million people living in 37 Latin American countries showed that CTPs reduced overall mortality [43]. While preventing violence remains a global challenge, identifying measures that mitigate adverse outcomes among victims is crucial. To our knowledge, no previous studies have specifically evaluated the contribution of CTPs on mortality in populations who have been victims of interpersonal violence. Our study addresses a critical evidence gap by providing the first empirical support that participation in the BFP is associated with a substantial reduction in mortality among individuals exposed to interpersonal violence populations that not only endure the immediate harm of violence but also face long-term risks to their life expectancy. These findings highlight the potential of social protection policies to mitigate the lethal consequences of violence. These findings are particularly striking, given that individuals affected by violence exhibit significantly higher mortality rates compared to the general population, primarily due to violent causes [44]. The BFP might affect mortality in people who have been victims of interpersonal violence through different processes. First, through the conditionalities, these programs are associated with improved access to health services, which in turn may be linked to fewer deleterious changes in the nervous, endocrine, and immune systems related to the stress of violence2, thereby reducing the risk of mortality. Additionality due to the conditionalities, these individuals may have greater access to other social Brazilian programs which promote inclusion and social protection for people in vulnerability and social risk [45]. Second, the role of the Family Health Strategy in monitoring families who receive benefits through home visits and follow-ups, can be highly relevant in reporting violence, facilitating subsequent treatment, and monitoring other health conditions [46]. Third, the benefit can also promote a greater sense of social inclusion and citizenship by providing access to social, educational, and health services, reducing the vulnerability associated with poverty and violence [5,47]. Fourth, financial support can improve living conditions and nutrition, increase resources, and reduce stress, which can contribute to improving physical and mental health [5,41]. Surprisingly, the BFP was not associated with non-natural causes of death. This finding might have happened because victims of interpersonal violence have an increased risk of dying from violent causes, given that recipients often reside in areas characterized by high rates of urban violence [6], and that a set of systematic public actions, together with the cash transfer, may be required for this already vulnerable population [2]. The findings further indicate that beneficiaries may experience improvements in overall health through the conditionalities of the BFP; however, the program may be insufficient to prevent deaths from violent causes [2]. BFP was also associated with reduced mortality especially among females, possibly by mitigating household-level financial stress [15]. Furthermore, previous research has highlighted the beneficial contribution of CTPs on women’s health, particularly in programs such as the BFP, which target pregnant and breastfeeding women [24,48]. Evidence from Low- and Middle-Income Countries (LMICs) indicates that cash transfers improve women’s health and increase their use of health services, particularly for maternal and childcare [48]. Another factor that may influence women’s health outcomes is the continued presence of the perpetrator within the household. Evidence from previous studies indicates that CTPs can enhance women’s autonomy and bargaining power [41], primarily because, in general, women are prioritized in the receipt of the social benefit [23]. Income transfer may alter household power or control dynamics, leading to separation processes when women are exposed to domestic violence [49]. However, it is also possible that perpetrators may appropriate the financial resources provided, thereby sustaining the cycle of domestic violence [41]. This dynamic is reflected in evidence indicating that CTPs have not led to a reduction in femicide rates [49], underscoring the need for complementary public policies specifically aimed at addressing perpetrator behavior and preventing severe forms of violence. This study has strengths and limitations. To our knowledge, it is the first to examine the association between BFP and overall mortality in victims of interpersonal violence, using nationwide linked datasets. Additionally, since the 100 Million Brazilian Cohort includes data from Brazil’s poorest population, this study uniquely highlights the potential contribution of a nationwide social program like BFP in reducing mortality rates among those in socioeconomic vulnerability. In addition, we used a robust analytical approach with a PS-based method, as well as various methodological strategies to check the robustness of our data. Moreover, this study took advantage of estimating survival using a competing risks model. Classical survival analysis models tend to overestimate survival probabilities and underestimate the risks of death, as the presence of competing risks is not considered in the analyses. This highlights the relevance of the competing risks model used in this study. Finally, using administrative data linked with a robust and accurate method reduced recall bias commonly associated with primary data collection. With regards to limitations, the datasets used in this study were composed of administrative data. Although these data are widely used for many studies [14–18,32,33], they were not designed for research purposes, and some missing data was observed in the covariates. However, the main data used in this study related to outcomes and exposure were fully complete. Second, considering registration on interpersonal violence from SINAN was not mandatory before 2011, we could only include individuals victimized from 2011 onward, which only enabled us to investigate the short-term contributions of BFP. A longer follow-up may be necessary to fully assess the impact of receiving BFP on mortality long-term, especially among individuals with repeated exposure to violent events. Furthermore, our study only included individuals that seek healthcare services, which likely represent the most severe types of interpersonal violence, potentially resulting in misclassification bias. Additionally, less visible forms of violence, such as psychological violence, may be underreported due to stigma and the difficulty healthcare professionals face in identifying them [50]. Third, we were unable to isolate the contribution of other smaller interventions that might also target low-income families registered at CadÚnico. Similarly, it is difficult to disentangle the role of receiving the benefit itself from the conditionalities attached to it. Therefore, it remains unclear whether the observed associations are attributable to the financial component per se or to behavioral changes required to maintain eligibility for the program. Fourth, although we controlled for sociodemographic covariates, we must consider the influence of unmeasured confounders, especially those related to behavioral factors, such as alcohol abuse. In addition, the covariates were available only at the baseline of the study, which may limit their contribution to the outcome, especially age [39]. Finally, this study may have biases due to the computational complexity involved in the linkage process and the absence of a unique number that identifies individuals across the health and social systems. However, the data used in this study presented high sensitivity and specificity in validation processes which allowed us to find that these errors are probably nondifferential [25]. Considering that we used data from the 100 million Brazilian cohort, which includes individuals with socioeconomic difficulties who applied for social benefits, our findings cannot be generalized for all Brazilian people. The cohort has a specific profile being predominantly composed by women and younger individuals [23]. It is also noteworthy that the profile of the population in this study differed from other studies that used SINAN [51], which is related to the focus of this study on individuals registered in CadÚnico after the violent event. Our findings revealed the potential contribution that a conditional CTP may have as a public policy for preventing negative health consequences related to interpersonal violence, in a specific population with greater vulnerability due to living in poverty and being victimized by violence. Our study showed that a considerable number of mortalities could be avoided by participation in the BFP. Therefore, governments should integrate cash transfers into efforts to prevent deaths among those vulnerabilized by interpersonal violence, especially conditional programs that improve access to health and social services. Furthermore, based on these findings, it is argued that governments should implement conditional CTPs worldwide, given the significant contribution interpersonal violence has on public health. Conditional CTPs should be advisable for governments to implement than non-conditional programs, as they have the potential to constructive behaviors, mitigating adverse health outcomes associated with interpersonal violence. Supporting information S1 Appendix. Method A – Description of the main datasets used in the study. Method B – Description of the linkage. Table A – Logistic regression to estimate propensity scores for receiving Bolsa Familia Program according to covariates. Table B – Association of Bolsa Família Program participation with overall mortality by sex (2011–2015). Table C – Association of Bolsa Família Program participation with the outcomes using Poisson model. Table D – Association of Bolsa Família Program participation with the outcomes considering overall population who were victim of interpersonal violence (2011–2015). Table E – Association of Bolsa Família Program participation with overall mortality considering missing as a category (2011–2015). Table F – Crude association of Bolsa Família Program participation with overall mortality (2011–2015). Table G – Association of Bolsa Família Program participation with overall mortality using time-varying BFP status (2011–2015). Table H – Information Criteria–Based Model Comparison Evaluating the Inclusion of an Interaction Term. Table I – Association of Bolsa Família Program participation with overall mortality considering other age categorization in the propensity score estimation (2011–2015). Fig A – Distribution of the propensity score in the sample, 2011–2015. https://doi.org/10.1371/journal.pmed.1004673.s001 (DOCX) S1 STROBE Checklist. The STROBE checklist is best used in conjunction with this article (freely available on the Web sites of PLoS Medicine at http://www.plosmedicine.org/, Annals of Internal Medicine at http://www.annals.org/, and Epidemiology at http://www.epidem.com/). Information on the STROBE Initiative is available at www.strobe-statement.org. https://doi.org/10.1371/journal.pmed.1004673.s002 (DOC) Acknowledgments We extend our gratitude to the data production team and all collaborators at CIDACS/FIOCRUZ for their efforts in developing the 100 million Brazilian Cohort.
U = U for all: Advancing equity in HIV preventionTorres, Thiago S.;Luz, Paula M.
doi: 10.1371/journal.pmed.1005090pmid: 42166695
Advances in antiretroviral therapy (ART) against HIV have contributed to global rates of HIV suppression rising from 40% in 2015 to 73% in 2024 [1,2]. Adherence to ART leads to HIV viral suppression to undetectable levels, with studies demonstrating that sexual transmission risk is zero when viral suppression is sustained [3], producing both individual health benefits and important population-level prevention effects. These effects are reflected in the “Undetectable = Untransmittable” (U = U) message, first coined in 2016, as well as the World Health Organization’s “Zero Risk” policy brief (www.iasociety.org/zero-risk-transmitting-hiv). Awareness, understanding, and acceptance of U = U is associated with health outcomes, including stigma reduction, enhanced sexual intimacy, the alleviation of physiological distress (such as anxiety, shame, and guilt), and increased HIV service engagement [4]. In this sense, U = U links viral suppression to dignity, empowerment, and quality of life. In this Perspective, we examine the persistent barriers to U = U awareness and acceptance across diverse populations and interventions to promote equitable implementation and broader public health impact. Despite the strength of the evidence, U = U literacy remains highly unequal across regions and population groups. A meta-analysis involving ~227,000 participants indicated that while U = U awareness is high among people living with HIV (PLHIV), it remains moderate among gay and other men who have sex with men (MSM) and low in the general population [5]. A study from Brazil, for instance, showed that 79% of PLHIV perceived U = U as completely accurate (undetectable viral load means no HIV transmission; 4-points Likert scale), contrasted with only 44% of MSM not living with HIV and 17% of the general population [6]. Geographical disparities have also been shown with studies conducted in high-income settings reporting greater U = U awareness and understanding than those conducted in low-income settings [5]. In sub-Saharan Africa, which accounts for more than 60% of the PLHIV globally and where women and girls accounted for ~63% of new HIV diagnoses in 2024 [1], literature suggests limited diffusion of U = U [4]. In a study from Uganda, for instance, men perceived that it was very unlikely that a couple could have different HIV statuses, even if the partner living with HIV was on ART [4]. In a qualitative study among PLHIV from Rwanda, most participants expressed skepticism or reluctance to accept U = U [7]. At the individual level, U = U understanding might also be influenced by social determinants of health, such as income and education. A study from Brazil that included 23,981 sexual and gender diverse participants (21% PLHIV, 72% HIV–negative, and 7% HIV unknown) showed that higher income and education was significantly associated with higher HIV knowledge, and higher HIV knowledge was associated with a higher odds of perceiving U = U as completely accurate [8]. Meanwhile, in a Canadian study, PLHIV reporting lower education and unemployment were less likely to report awareness, acceptance, and positive impacts of U = U in their lives [9]. These results suggest that U = U literacy is not distributed evenly, such that different regions and populations do not benefit from the social and preventive meaning of U = U and remain vulnerable to outdated fears about HIV transmission. Similarly, although awareness and acceptability of U = U among healthcare providers have increased over time, there is ample evidence to suggest that healthcare providers do not fully understand U = U or remain reluctant to communicate it [4]. Many healthcare providers remain hesitant to share the U = U message, sometimes preferring softer or ambiguous language such as virtually impossible and negligible [10]. A qualitative study in Malawi highlighted that stakeholders expressed uncertainty of how to communicate the message without causing harm, meaning that while they acknowledged the scientific reality of U = U, they feared that patients would misinterpret viral suppression as being cured, leading them to relax, stop taking medication, or start sleeping around and potentially re-infect themselves [11]. Likewise, a qualitative study conducted in the United States and Australia showed that providers engaged with HIV treatment and prevention used ambiguous or inaccurate messaging regarding U = U [12]. These findings reflect persistent paternalism, fear of patient misunderstanding, and misplaced concerns about patient behavior such as engagement in condomless sex. However, when providers withhold or dilute the message, they inadvertently preserve stigma and limit patient autonomy. Provider bias therefore becomes more than a communication problem; it becomes a barrier to empowerment. If U = U is to function as both a clinical principle and a stigma-reduction strategy, providers need structured support to routinely deliver clear non‑stigmatizing U = U messages that are tailored to patients’ knowledge levels and supplemented with visual aids [12]. Provider education is especially important because the credibility of U = U often depends on whether it is introduced as a routine part of HIV care or framed as an exceptional or controversial message. In practice, provider hesitation can delay patient understanding, weaken trust, and reduce the likelihood that U = U will be shared beyond the clinic. More studies are needed to inform how to best communicate the U = U message and the actual impact of disseminating U = U messaging on behaviors and clinical outcomes [4]. Scaling U = U equity requires interventions tailored for different populations and levels of access. For example, sub-Saharan Africa bears a higher HIV burden across multiple demographics, whereas in the Americas and Western Europe, the epidemic disproportionately affects key populations such as MSM, and education must be tailored accordingly. In South Africa, where 74% of 7.8 million PLHIV have viral suppression [1], a study described a structured co‑creation process of a U = U messaging tool that involved PLHIV, lay counselors, primary‑care clinicians, and representatives from local HIV advocacy organizations [13]. The web‑ and mobile‑based tool called “Undetectable and You” integrated first‑person video testimonials, narrative framing, and simplified explanations of viral suppression grounded in locally validated idioms and explicitly targeted culturally embedded beliefs, HIV‑related stigma, and trust dynamics within the South African primary‑care settings. Though the impact of this campaign is still to be shown, the effort exemplifies how multiple actors need to be actively engaged and trained to compose accurate, culturally relevant information to help disseminate the U = U message within their own networks and strengthen trust in the intervention. Beyond the content, campaign dissemination strategies must target population and providers level gaps: targeted social media ads on platforms for youth, radio spots for low-literacy/rural groups, partnerships with influencers and advocacy networks, posters with QR code linking to educational videos in healthcare facilities and free online U = U training modules for providers. Peer-led education is particularly important for reaching populations that may distrust formal health institutions or have limited contact with specialized HIV services. U = U should also be integrated into LGBTQIAPN-focused and transgender-affirming health services, where culturally competent, nonjudgmental care can improve trust, normalize viral suppression as part of routine care. For broader reach, U = U messaging should be embedded in primary care, sexual health services and in school-based education. Ultimately, achieving U = U equity will require more than biomedical evidence alone. Although the scientific basis is strong, the reach and impact of the message still depend on social trust, accessible communication, and supportive legal and policy environments. Persistent stigma, unequal access to information, punitive HIV-related laws, and broader violations of LGBTQIAPN+ and women’s rights continue to limit who can fully benefit from U = U [14]. These challenges are compounded by political instability, rising conservatism, and shrinking HIV funding, which weaken prevention systems and narrow the space for rights-based health communication. Closing these gaps will require coordinated action across health systems, community organizations, and policymakers to ensure that U = U is communicated clearly, understood broadly, and supported by laws and services that align with current science.
Differences in tuberculosis prevalence by sex in low- and middle-income countries over 1993–2025: A systematic review and meta-analysisSwartwood, Nicole A.;Singh, Nanki;Mortazavi, Seyed Alireza;Can, Melike Hazal;Cui, Hening;Ryuk, Do Kyung;MacPherson, Peter;Horton, Katherine C.;Menzies, Nicolas A.
doi: 10.1371/journal.pmed.1005114pmid: 42172294
Background Global and national initiatives to combat tuberculosis (TB) have expanded over recent years. Despite this, the TB burden remains high in some population groups, with men recognized as having elevated TB risks. Summary measures of sex differences in TB prevalence were last estimated in 2016. Since then, many additional prevalence surveys have been conducted, including in the highest TB burden countries. We conducted a systematic review of sex-stratified TB prevalence survey data published over 1993–2025, to provide updated estimates of male-to-female (M:F) TB prevalence ratios and determine whether sex-related disparities in TB burden have closed over time. Methods and findings We identified surveys reporting community-representative, sex-stratified estimates of pulmonary TB prevalence in low- and middle-income countries (LMICs), including surveys from an earlier review (covering January 1993–March 2016) and a new systematic review (covering 1st December 2015–13th October 2025). This review was prospectively registered with PROSPERO (CRD42024503853) and included searches of PubMed, Embase, Global Health, the Cochrane Library, Africa Index Medicus, LILACS, and SciELO. We extracted data on bacteriologically confirmed and smear-positive TB prevalence among adults (aged ≥ 15 years), stratified by sex. Risk of bias was evaluated using eight criteria specific to prevalence surveys. We fit multi-level Bayesian regression models with study- and country-level random effects to estimate the M:F ratio of TB prevalence (male prevalence divided by female prevalence), overall and for key subgroups. In meta-regression analyses, we estimated how prevalence ratios varied over time and according to known TB risk factors and TB case definitions. We identified 10,124 publications and extracted data from 100 eligible studies representing 102 unique prevalence surveys and 4,658,310 participants (45.6% male) in 33 LMICs. TB prevalence was higher in men than women in 90/102 of the included surveys, with a pooled M:F prevalence ratio of 2.02 (95% credible interval (CrI): 1.71, 2.34) for bacteriologically confirmed TB and 2.38 (95% CrI: 1.91, 2.90) for smear-positive TB. Time trend analyses showed a 2.0% (95% CrI: −0.2, 4.5%) average annual change in the M:F ratio of bacteriologically confirmed TB over the study period. The M:F prevalence ratio was estimated to be higher for countries with greater excess HIV prevalence among men, and countries with greater gender equity (as measured by the United Nation’s Gender Development Index). The estimated M:F prevalence ratio was also higher for surveys that did not restrict testing to individuals reporting TB symptoms. Study limitations include heterogeneity in survey methods and definitions, as well as limited data from the Americas, Eastern Mediterranean, and Europe WHO world regions and post-COVID-19 period. Conclusions Men in LMICs consistently experience TB at a higher prevalence than women. Time trend estimates are uncertain, but consistent with widening sex differences in TB prevalence over the last three decades, despite efforts to address the risk factors underlying this excess TB burden. Why was the study done? Previous studies, including a 2016 systematic review and meta-analysis, have identified substantial sex differences in tuberculosis (TB) burden, with higher TB infection, prevalence, and mortality consistently observed among men in low- and middle-income countries. The change in these sex differences over time has not been previously estimated. Recent improvements in TB case detection efforts, diagnostics, and treatment approaches may differentially affect TB epidemiology among men and women. Additionally, newly conducted TB prevalence surveys have increased the duration and diversity of evidence to inform trends in TB burden. What did the researchers do and find? We systematically identified 102 national and sub-national TB prevalence surveys conducted in low- and middle-income countries and used Bayesian meta-regression models to estimate male-to-female TB prevalence risk ratios overall and among key demographic and epidemiological subgroups. We also estimated how prevalence risk ratios varied according to known TB risk factors, TB case definitions, and over time (1994–2024). Adult men have over twice the prevalence of pulmonary TB as compared to women in low- and middle-income countries (LMICs). Increased prevalence was associated with greater excess HIV prevalence among men, and countries with greater gender equity (as measured by the United Nation’s Gender Development Index). The estimated M:F prevalence ratio was higher among surveys that did not restrict testing to individuals reporting TB symptoms. Temporal analysis suggested that male-to-female inequalities in TB prevalence might be growing, particularly in the World Health Organization (WHO) Africa region. What do these findings mean? Despite global commitments to gender equity in health, men in LMICs continue to bear a disproportionate burden of TB. If sex-related inequalities in TB burden are growing, developing effective strategies to reduce men’s risk of TB and to engage men in TB prevention and care will be essential to end TB. Limitations include limited prevalence survey data in the Americas, Eastern Mediterranean, and Europe WHO world regions, methodological heterogeneity across surveys, especially among subnational surveys, and few prevalence surveys from the post-COVID-19 period. Introduction Despite sustained global efforts to reduce tuberculosis (TB), it remains the leading cause of death from a single infectious agent globally [1]. In 2024, men were estimated to account for approximately 60% of TB incidence among adults and TB deaths among HIV–negative adults worldwide [1]. The mechanisms underpinning these differences, and the extent to which they vary across time and settings, are not well understood. Gender—defined as the set of socially constructed roles, behaviors, and norms related to perceived sex—is a key determinant of health and shapes TB susceptibility, exposure risks, engagement with prevention and care, and outcomes [2–5]. Gendered health behaviors and systemic healthcare access barriers can lead to delayed diagnosis and worse care engagement among men, contributing to ongoing transmission and under-detection. In addition, sex-assortative mixing among men—particularly in workplaces, social settings, and high-risk congretate settings such as prisons or homeless shelters—can amplify transmission within male networks [2,3]. Men also have a higher prevalence of risk factors that increase both susceptibility to infection (e.g., tobacco use, occupational, and ambient air pollution) and progression from infection to TB disease (e.g., untreated HIV, diabetes, malnutrition, and smoking). Together, sex and gender interact with structural determinants, such as occupational exposures, unequal health system access, and prevailing societal norms, to produce a persistent excess burden of TB among men [4,5]. Conventionally, measures of disease burden are reported by biological sex, not gender. However, given the considerable overlap of sex and gender identities within populations, sex-stratified outcomes will reflect the joint impact of these two related, but distinct, factors. While routinely reported, TB case notifications have marked limitations as measures of TB burden, as they can be affected by incomplete case detection and reporting. National TB prevalence surveys are designed to generate population-representative estimates of TB burden unaffected by these potential biases and have been conducted in a number of high-burden countries with the support of the World Health Organization (WHO). Historically, these surveys have reported higher TB prevalence in men than in women in low- and middle-income countries (LMICs) [6]. A 2016 systematic review and meta-analysis by Horton and colleagues estimated that men in LMICs had more than twice the prevalence of TB compared with women [7]. TB epidemiology is not static and will change under the influence of multiple factors. During the COVID-19 pandemic, many TB resources were reallocated to pandemic response—an effort that led to the pronounced decreases in TB case detection and increases in TB deaths [8,9]. Most recently, global and national commitments to TB elimination have been renewed through accelerated case detection and introduction of new diagnostics and treatment approaches [10]. There has also been a scale-up of interventions to address key risk factors for TB, including increased treatment of HIV [11]. Alongside these changes, a substantial number of new prevalence surveys have been conducted since 2016, including in some of the world’s highest TB burden countries (e.g., India, South Africa) [1]. Sex differences in TB prevalence might have changed in response to these factors, or as a result of the increasing recognition of sex as a TB risk factor [1,12], most notably the addition of men to WHO’s list of populations vulnerable to TB [13]. However, empirical analyses have not described whether sex differences in TB burden have changed over time. In this study, we undertook a systematic review and meta-analysis of sex-disaggregated TB prevalence surveys conducted in LMICs, published between 1st January 1993 and 13th October 2025. Based on this review, we investigated whether sex differences in TB prevalence (operationalized by the male-to-female ratio of TB prevalence) have changed over time. We also explored how this ratio is correlated with key socio-demographic determinants and approaches used to determine TB prevalence. Download: PNG larger image TIFF original image Table 1. Multiplicative effect of standardized3 covariates on male-to-female ratio of bacteriologically-confirmed TB prevalence estimated from nationally representative surveys (N = 38)4. https://doi.org/10.1371/journal.pmed.1005114.t001 Methods We conducted a systematic review and meta-analysis of sex-stratified TB prevalence surveys conducted among people living in LMICs (as defined by the World Bank) [14], We registered the review and meta-analysis with the International Prospective Register of Systematic Reviews (PROSPERO) protocol number CRD42024503853 [15]. We followed the Meta-analysis of Observational Studies in Epidemiology (MOOSE) guidelines for the conduct of the review (S1 Checklist) and the Preferred Reporting Items for Systematic Reviews and Meta-analyses (PRISMA) guidelines for preparing the manuscript (S2 Checklist) [16,17]. We used the Covidence online systematic review tool to manage the review [18]. Search strategy We designed the search strategy as an update of a previous systematic review and meta-analysis by Horton and colleagues [7]. Our search included a combination of MeSH terms and keywords specifically related to “tuberculosis,” “prevalence,” and “low- and middle-income countries (LMICs).” We searched four electronic databases—Global Health, the Cochrane Database of Systematic Reviews, PubMed, and Embase—between 1st December 2015 and 13th October 2025 for English language publications. The full search strategy is reported in Table A in S1 Appendix. We additionally evaluated all studies included in the previous analysis for inclusion [7] and compared our included studies against the national TB prevalence surveys listed in WHO Global Tuberculosis Report 2025, adding any survey reports that were not captured by our search [1]. We also examined the citations of all identified publications for additional relevant material. Finally, we screened published abstracts from the 2023 Union World Conference on Lung Health for recent prevalence surveys [19]. Inclusion and exclusion criteria We reviewed all articles and included TB prevalence surveys that reported sex-stratified estimates of adult (aged ≥ 15 years) pulmonary TB prevalence from a population-based sample, conducted in a LMIC (based on 2024 World Bank classification). Extra-pulmonary TB was excluded due to the diagnostic complexity of this condition, which often necessitates invasive diagnostic procedures not typically available in community survey settings. We also excluded studies that only evaluated care-seeking persons, surveys solely reliant on self-reported data without laboratory confirmation, surveys conducted following a community-based active case finding intervention in the same setting (as the intervention may have altered local TB prevalence), and surveys of children aged 14 years or younger, as determining TB status can be challenging in this group. Detailed inclusion and exclusion criteria can be found in Tables B and C in S1 Appendix. Our review team consisted of nine people (NS, NAS, AM, MHC, HC, DKR, PM, KCH, and NAM). Two reviewers (of NS, NAS, AM, MHC, HC, and DKR, allocated at random) independently assessed the titles and abstracts of all studies to identify those requiring full text review. Discrepancies were resolved through consensus discussions, mediated by a third reviewer (NAS, PM, KCH, or NAM). Two reviewers (of NS, NAS, AM, MHC, HC, DKR, PM, and KCH) independently evaluated full texts for study inclusion, with discrepancies being resolved through discussion as outlined above. Data extraction Two reviewers independently extracted data on study methodology, risk of bias, and TB prevalence using a pretested extraction form (S1 Form). Discrepancies between reviewers were resolved by a third reviewer. Where multiple studies reported on the same TB prevalence survey, we merged the studies and extracted all available data; where conflicting data were reported, we extracted data from the most recent publication. Where studies reported multiple prevalence surveys, data were extracted separately for each survey. Therefore, the unit of analysis in our meta-analysis was the prevalence survey as opposed to the publication. Risk of bias We assessed the risk of bias for each included study using eight criteria adapted from the Hoy and colleagues tool for the appraisal of prevalence surveys [20]. This tool examines study population selection, nonresponse bias, data collection methodology, and case definition parameters. The assessment criteria are listed in the extraction form (S1 Form, Section 8). Reviewers were also asked to provide an overall assessment of risk of bias (“low”, “moderate”, or “high”) based on the eight criteria. To evaluate the robustness of our primary English-language search strategy, we also searched three databases—Africa Index Medicus, LILACS, and SciELO—for French, Portuguese, and Spanish publications over the full study period (1st January 2009–13th October 2025). The corresponding search strings are in Table D in S1 Appendix. Non-English titles and abstracts were translated to English using DeepL, a neural machine translation (NMT) tool; studies identified for full-text review were similarly translated. Definitions To be eligible for quantitative analysis, a survey was required to report a measure of bacteriologically confirmed TB. Bacteriologically confirmed TB was defined as TB confirmed through bacteriological methods, requiring at least one positive result of smear microscopy, culture, or a molecular diagnostic (e.g., Xpert MTB/RIF). Where reported, we adopted the study definition of bacteriologically confirmed TB. In the absence of a study definition, we used estimates of culture-positive TB prevalence or smear-positive TB prevalence, prioritized in that order. Smear-positive TB was defined as TB confirmed through sputum smear microscopy, regardless of other diagnostic results. Where studies reported more than one prevalence measure, adjusted TB prevalence estimates with 95% confidence intervals were the preferred measure used for synthesis, followed by crude prevalence estimates with 95% confidence intervals, then counts of individuals with TB and total number tested, or summary prevalence estimates without a confidence interval (Fig A in S1 Appendix). Survey year was set to the end year of data collection; where the end year was missing, we set the survey year to the publication year less the mean observed publication delay in our data, which was 4 years. Participants were classified as male or female by the underlying prevalence review, the basis for which was not commonly reported. As such, we adopted these labels as reported for our analysis. Data synthesis and statistical analysis Reported data from all surveys were converted to prevalence per 100,000 persons. We took one of two approaches to calculate the standard error of each survey-reported prevalence estimate. If confidence intervals were reported, we used these to back-calculate the standard error; for estimates without confidence intervals, we applied the normal approximation method of the binomial standard error. Standard errors were estimated on the log scale. All crude prevalence estimates (i.e., those not already adjusted for features of the survey design) had a variance correction factor applied in order to approximate the additional sampling variance due to complex survey designs. Our primary outcome was the male-to-female (M:F) ratio of bacteriologically confirmed TB prevalence, calculated as TB prevalence among males divided by TB prevalence among females. We used Bayesian multilevel meta-regression models to synthesize evidence from the pooled survey data, with models constructed for the log of the M:F prevalence ratio. This log transformation was adopted to improve the symmetry and homoscedasticity of residuals for the fitted model. This approach also produced a multiplicative relationship between predictor variables and the M:F prevalence ratio, allowing results to be interpreted as risk ratios. We adopted an identity link function and specified a normal likelihood function for the survey data. The standard deviation of individual observations (the log of study-specific M:F prevalence ratios) was calculated from study-reported standard errors. We chose our priors to be weakly informative, based on published guidance [21], and tested the impact of alternative prior specifications on the study results. To estimate overall and study-specific M:F prevalence ratios, we fit models with study- and country-level random effects. To estimate M:F prevalence ratios for each WHO region, we fit models with study- and region-level random effects. For each model, random effects and associated variance parameters were assigned priors and estimated simultaneously, such that uncertainty in these values was fully propagated into posterior inference. We examined these results to report the relative magnitude of random effect variance at the country (or region) and study level. All model equations and priors are reported in Table E in S1 Appendix. To explore how the M:F prevalence ratio changed over time, we revised the regression models to include a term for study year and added random effects for the interaction of this time trend with study country or WHO region, respectively. From the results of these models, we calculated the estimated annual percentage change (EAPC) in the M:F ratio of bacteriologically confirmed TB prevalence, overall and by WHO region. We also calculated the probability that these EAPC estimates were positive (P(EAPC > 0)), as an indicator of whether the magnitude of the M:F prevalence ratio was increasing over time. We extended these regression models to investigate univariable and multivariable country-level exploratory associations between several TB determinants and the M:F ratio of bacteriologically confirmed TB. We included covariates for survey year, United Nations’ Gender Development Index (GDI) [22]. and absolute difference in country-specific male and female levels of prevalence of alcohol use disorder, type II diabetes mellitus, HIV/AIDS, underweight as measured through low body mass index (BMI), and smoking, obtained from the Global Burden of Diseases Project (GBD) [23,24]. Covariates were matched to each survey based on survey country and end year, and standardized globally to a zero mean and unit standard deviation. As such, coefficient estimates reflect the change in the log M:F prevalence ratio associated with a 1 standard deviation change in the predictor. GBD covariates were missing for one survey (2.5%) and GDI was missing for five surveys (12.5%); where available (one of GBD and three of GDI missing surveys), these surveys were assigned the value corresponding to the nearest available year. We list the affected surveys in Table F in S1 Appendix. We fitted these models using data from 38 of 40 nationally-representative surveys (GDI was not available for the Democratic People’s Republic of Korea or Eritrea, and the corresponding surveys were omitted from univariable and multivariable analyses). We applied the main effect model to subgroups of the overall survey data, such as surveys stratified by survey representativeness (national or subnational), risk of bias (low or not low), use of symptom screening (yes or no), and whether bacteriologically confirmed TB required a positive Xpert MTB/RIF or culture, to explore potential differences in the M:F ratio of bacteriologically confirmed TB prevalence. For countries represented by five or more surveys, we estimated country-specific M:F prevalence ratios. We also examined how our results compared with the Horton and colleagues 2016 review by replicating our main effect model for those only those studies included in the previous study. Finally, as an alternative examination of the change in M:F ratios of bacteriologically confirmed TB prevalence over the study period, we fit the main effect model to extracted data grouped into five-year bands (e.g., 2000–2004, 2005–2009, etc.) by survey end year. We tested several alternative model specifications. We revised the main effects and time trend models to include binary covariates representing variation in the use of more sensitive TB diagnostics (i.e., Xpert MTB/RIF or culture) in each survey, and whether symptoms were required for sputum collection. We also explored alternative hierarchical structures, including nested country-random effects within WHO region-random effects. All analyses were conducted in R (version 4.4.2) using the `brms` and `tidybayes` packages [25–27]. The model ran for 10,000 iterations (2,500 burn-in) on 4 chains. Convergence was assessed through inspection of each model’s effective sample size (ESS), potential scale reduction factor (R), and posterior predictive checks. For all outcomes, we report the posterior mean and equal-tailed 95% credible interval (representing an interval with 95% probability of containing the true value) obtained from summarizing 10,000 posterior draws per parameter. Search strategy We designed the search strategy as an update of a previous systematic review and meta-analysis by Horton and colleagues [7]. Our search included a combination of MeSH terms and keywords specifically related to “tuberculosis,” “prevalence,” and “low- and middle-income countries (LMICs).” We searched four electronic databases—Global Health, the Cochrane Database of Systematic Reviews, PubMed, and Embase—between 1st December 2015 and 13th October 2025 for English language publications. The full search strategy is reported in Table A in S1 Appendix. We additionally evaluated all studies included in the previous analysis for inclusion [7] and compared our included studies against the national TB prevalence surveys listed in WHO Global Tuberculosis Report 2025, adding any survey reports that were not captured by our search [1]. We also examined the citations of all identified publications for additional relevant material. Finally, we screened published abstracts from the 2023 Union World Conference on Lung Health for recent prevalence surveys [19]. Inclusion and exclusion criteria We reviewed all articles and included TB prevalence surveys that reported sex-stratified estimates of adult (aged ≥ 15 years) pulmonary TB prevalence from a population-based sample, conducted in a LMIC (based on 2024 World Bank classification). Extra-pulmonary TB was excluded due to the diagnostic complexity of this condition, which often necessitates invasive diagnostic procedures not typically available in community survey settings. We also excluded studies that only evaluated care-seeking persons, surveys solely reliant on self-reported data without laboratory confirmation, surveys conducted following a community-based active case finding intervention in the same setting (as the intervention may have altered local TB prevalence), and surveys of children aged 14 years or younger, as determining TB status can be challenging in this group. Detailed inclusion and exclusion criteria can be found in Tables B and C in S1 Appendix. Our review team consisted of nine people (NS, NAS, AM, MHC, HC, DKR, PM, KCH, and NAM). Two reviewers (of NS, NAS, AM, MHC, HC, and DKR, allocated at random) independently assessed the titles and abstracts of all studies to identify those requiring full text review. Discrepancies were resolved through consensus discussions, mediated by a third reviewer (NAS, PM, KCH, or NAM). Two reviewers (of NS, NAS, AM, MHC, HC, DKR, PM, and KCH) independently evaluated full texts for study inclusion, with discrepancies being resolved through discussion as outlined above. Data extraction Two reviewers independently extracted data on study methodology, risk of bias, and TB prevalence using a pretested extraction form (S1 Form). Discrepancies between reviewers were resolved by a third reviewer. Where multiple studies reported on the same TB prevalence survey, we merged the studies and extracted all available data; where conflicting data were reported, we extracted data from the most recent publication. Where studies reported multiple prevalence surveys, data were extracted separately for each survey. Therefore, the unit of analysis in our meta-analysis was the prevalence survey as opposed to the publication. Risk of bias We assessed the risk of bias for each included study using eight criteria adapted from the Hoy and colleagues tool for the appraisal of prevalence surveys [20]. This tool examines study population selection, nonresponse bias, data collection methodology, and case definition parameters. The assessment criteria are listed in the extraction form (S1 Form, Section 8). Reviewers were also asked to provide an overall assessment of risk of bias (“low”, “moderate”, or “high”) based on the eight criteria. To evaluate the robustness of our primary English-language search strategy, we also searched three databases—Africa Index Medicus, LILACS, and SciELO—for French, Portuguese, and Spanish publications over the full study period (1st January 2009–13th October 2025). The corresponding search strings are in Table D in S1 Appendix. Non-English titles and abstracts were translated to English using DeepL, a neural machine translation (NMT) tool; studies identified for full-text review were similarly translated. Definitions To be eligible for quantitative analysis, a survey was required to report a measure of bacteriologically confirmed TB. Bacteriologically confirmed TB was defined as TB confirmed through bacteriological methods, requiring at least one positive result of smear microscopy, culture, or a molecular diagnostic (e.g., Xpert MTB/RIF). Where reported, we adopted the study definition of bacteriologically confirmed TB. In the absence of a study definition, we used estimates of culture-positive TB prevalence or smear-positive TB prevalence, prioritized in that order. Smear-positive TB was defined as TB confirmed through sputum smear microscopy, regardless of other diagnostic results. Where studies reported more than one prevalence measure, adjusted TB prevalence estimates with 95% confidence intervals were the preferred measure used for synthesis, followed by crude prevalence estimates with 95% confidence intervals, then counts of individuals with TB and total number tested, or summary prevalence estimates without a confidence interval (Fig A in S1 Appendix). Survey year was set to the end year of data collection; where the end year was missing, we set the survey year to the publication year less the mean observed publication delay in our data, which was 4 years. Participants were classified as male or female by the underlying prevalence review, the basis for which was not commonly reported. As such, we adopted these labels as reported for our analysis. Data synthesis and statistical analysis Reported data from all surveys were converted to prevalence per 100,000 persons. We took one of two approaches to calculate the standard error of each survey-reported prevalence estimate. If confidence intervals were reported, we used these to back-calculate the standard error; for estimates without confidence intervals, we applied the normal approximation method of the binomial standard error. Standard errors were estimated on the log scale. All crude prevalence estimates (i.e., those not already adjusted for features of the survey design) had a variance correction factor applied in order to approximate the additional sampling variance due to complex survey designs. Our primary outcome was the male-to-female (M:F) ratio of bacteriologically confirmed TB prevalence, calculated as TB prevalence among males divided by TB prevalence among females. We used Bayesian multilevel meta-regression models to synthesize evidence from the pooled survey data, with models constructed for the log of the M:F prevalence ratio. This log transformation was adopted to improve the symmetry and homoscedasticity of residuals for the fitted model. This approach also produced a multiplicative relationship between predictor variables and the M:F prevalence ratio, allowing results to be interpreted as risk ratios. We adopted an identity link function and specified a normal likelihood function for the survey data. The standard deviation of individual observations (the log of study-specific M:F prevalence ratios) was calculated from study-reported standard errors. We chose our priors to be weakly informative, based on published guidance [21], and tested the impact of alternative prior specifications on the study results. To estimate overall and study-specific M:F prevalence ratios, we fit models with study- and country-level random effects. To estimate M:F prevalence ratios for each WHO region, we fit models with study- and region-level random effects. For each model, random effects and associated variance parameters were assigned priors and estimated simultaneously, such that uncertainty in these values was fully propagated into posterior inference. We examined these results to report the relative magnitude of random effect variance at the country (or region) and study level. All model equations and priors are reported in Table E in S1 Appendix. To explore how the M:F prevalence ratio changed over time, we revised the regression models to include a term for study year and added random effects for the interaction of this time trend with study country or WHO region, respectively. From the results of these models, we calculated the estimated annual percentage change (EAPC) in the M:F ratio of bacteriologically confirmed TB prevalence, overall and by WHO region. We also calculated the probability that these EAPC estimates were positive (P(EAPC > 0)), as an indicator of whether the magnitude of the M:F prevalence ratio was increasing over time. We extended these regression models to investigate univariable and multivariable country-level exploratory associations between several TB determinants and the M:F ratio of bacteriologically confirmed TB. We included covariates for survey year, United Nations’ Gender Development Index (GDI) [22]. and absolute difference in country-specific male and female levels of prevalence of alcohol use disorder, type II diabetes mellitus, HIV/AIDS, underweight as measured through low body mass index (BMI), and smoking, obtained from the Global Burden of Diseases Project (GBD) [23,24]. Covariates were matched to each survey based on survey country and end year, and standardized globally to a zero mean and unit standard deviation. As such, coefficient estimates reflect the change in the log M:F prevalence ratio associated with a 1 standard deviation change in the predictor. GBD covariates were missing for one survey (2.5%) and GDI was missing for five surveys (12.5%); where available (one of GBD and three of GDI missing surveys), these surveys were assigned the value corresponding to the nearest available year. We list the affected surveys in Table F in S1 Appendix. We fitted these models using data from 38 of 40 nationally-representative surveys (GDI was not available for the Democratic People’s Republic of Korea or Eritrea, and the corresponding surveys were omitted from univariable and multivariable analyses). We applied the main effect model to subgroups of the overall survey data, such as surveys stratified by survey representativeness (national or subnational), risk of bias (low or not low), use of symptom screening (yes or no), and whether bacteriologically confirmed TB required a positive Xpert MTB/RIF or culture, to explore potential differences in the M:F ratio of bacteriologically confirmed TB prevalence. For countries represented by five or more surveys, we estimated country-specific M:F prevalence ratios. We also examined how our results compared with the Horton and colleagues 2016 review by replicating our main effect model for those only those studies included in the previous study. Finally, as an alternative examination of the change in M:F ratios of bacteriologically confirmed TB prevalence over the study period, we fit the main effect model to extracted data grouped into five-year bands (e.g., 2000–2004, 2005–2009, etc.) by survey end year. We tested several alternative model specifications. We revised the main effects and time trend models to include binary covariates representing variation in the use of more sensitive TB diagnostics (i.e., Xpert MTB/RIF or culture) in each survey, and whether symptoms were required for sputum collection. We also explored alternative hierarchical structures, including nested country-random effects within WHO region-random effects. All analyses were conducted in R (version 4.4.2) using the `brms` and `tidybayes` packages [25–27]. The model ran for 10,000 iterations (2,500 burn-in) on 4 chains. Convergence was assessed through inspection of each model’s effective sample size (ESS), potential scale reduction factor (R), and posterior predictive checks. For all outcomes, we report the posterior mean and equal-tailed 95% credible interval (representing an interval with 95% probability of containing the true value) obtained from summarizing 10,000 posterior draws per parameter. Results Of the 10,124 publications screened by title and abstract, 216 had their full text reviewed (Fig 1). Table G in S1 Appendix lists the publications excluded at the full-text review stage with their corresponding exclusion reasons. The most common exclusion reasons were wrong outcome (39%) and wrong study population (34%). We identified 100 English-language publications that described 102 unique surveys reporting sex-stratified bacteriologically confirmed TB prevalence estimates [28–127]. We did not find any relevant surveys in the examined French, Portuguese, or Spanish studies. The characteristics of included surveys and participants are reported in Tables H and I in S1 Appendix. These surveys were conducted in 33 countries across five WHO world regions: 16 in Africa region, seven in the South-East Asia region, six in the Western Pacific region, two in the Eastern Mediterranean region, and two in the region of the Americas (Fig 2). All included countries were classified as WHO high TB incidence countries. Overall, 100/102 surveys reported total participant numbers, giving 4,658,310 participants; 96/102 surveys reported sex-stratified participant numbers, with 45.6% male participants. Download: PNG larger image TIFF original image Fig 1. PRISMA diagram of study selection. https://doi.org/10.1371/journal.pmed.1005114.g001 Download: PNG larger image TIFF original image Fig 2. Distribution and counts of included prevalence surveys. Panel A: Geographic distribution of surveys, including the total number of surveys in each country. Panel B: Number of subnational and national TB prevalence surveys per country. Note: Map was generated using the maps R package [156]. https://doi.org/10.1371/journal.pmed.1005114.g002 Among the identified surveys, 90 (88%) reported greater bacteriologically confirmed TB prevalence among men than women; all of the 40 nationally representative surveys reported greater male bacteriologically confirmed TB prevalence. The overall pooled M:F ratio of bacteriologically confirmed TB was 2.02 (95% credible interval: 1.71, 2.34; Fig 3). Posterior predictive checks did not reveal any major systematic discrepancies between model estimates and the underlying data. Details on model performance can be found in Table J and Fig B in S1 Appendix. Table K in S1 Appendix contains estimated M:F ratios from an alternative model hierarchical structure. We also estimated the impact of requiring symptoms for sputum collection and differential diagnostic algorithms on the male-to-female ratio of bacteriologically confirmed TB prevalence; these estimates are reported in Table L in S1 Appendix. Surveys that required symptoms for sputum collection had lower M:F ratios compared to surveys that did not (multiplicative effect: 0.64; 95% CrI: 0.51, 0.80). Surveys that used Xpert MTB/RIF or sputum culture as a component of the diagnostic cascade had higher M:F ratios compared to surveys that only used sputum smear microscopy (multiplicative effect: 1.48; 95% CrI: 1.12, 1.94). Download: PNG larger image TIFF original image Fig 3. Male-to-female ratios of bacteriologically confirmed (n = 102) TB prevalence by world region, as compared with ratios calculated from survey reported data. https://doi.org/10.1371/journal.pmed.1005114.g003 In each WHO region, bacteriologically confirmed TB prevalence was higher among males compared with females, ranging from 1.76 (95% CrI: 1.17, 2.49) in the Eastern Mediterranean to 2.91 (95% CrI: 2.51, 3.33) in South-East Asia. Posterior summaries indicated that the between-country and between-region variance exceeded within-county or within-region (i.e., survey-level) variance, respectively (Fig C and Table M in S1 Appendix). This reallocation of variance is consistent with country and region random effects capturing genuine structural differences in TB epidemiology and health system context, rather than representing arbitrary aggregation. Of the identified surveys, 64 also reported sex-stratified smear-positive TB prevalence. All these studies reported the number of participants, for a total of 2,825,955 (45.6% male). The estimated M:F ratio of smear-positive TB was higher than the bacteriologically confirmed TB estimate for the pooled survey data (2.38; 95% CrI: 1.91, 2.90) and for each world region, when estimated using only those surveys that reported both bacteriologically confirmed and smear-positive TB prevalence (Table N in S1 Appendix). The M:F ratios of smear-positive TB prevalence varied among WHO regions—Americas: 1.26 (95% CrI: 0.51, 2.42), Africa: 1.82 (95% CrI: 1.42, 2.26), Eastern Mediterranean: 2.09 (95% CrI: 1.20, 3.31), Western Pacific: 2.58 (95% CrI: 1.95, 3.33), and South-East Asia: 3.65 (95% CrI: 2.97, 4.38). Survey year, defined as the last year of data collection, ranged from 1994 to 2024; five surveys did not report the survey end year (two in Africa, two in South-East Asia, and one in Western Pacific regions). We estimated a likely increasing, but uncertain, time trend in the M:F ratio of bacteriologically confirmed TB (Fig 4). On average, the M:F prevalence ratio increased by 2.0% (95% CrI: −0.2, 4.5%) annually, with a probability of positive direction (PPD) of 0.96. The Africa region had the greatest annual rate of increase, with an average annual change of 2.9% (95% CrI: 0.2, 6.0%) and a PPD of 0.98, while trends in other regions were more uncertain (Table O in S1 Appendix). In exploratory subgroup analyses of 5-year bands, the M:F ratio ranged from 1.68 (95% CrI: 1.01, 2.58) during 2005–2010 to 2.84 (95% CrI: 1.50, 4.64) during 2020–2024; however, these years were among those with the fewest number of prevalence surveys (2005–2010: 18 total, 6 nationally-representative; 2020–2024: 6 total; 2 nationally-representative). We found no evidence that the use of Xpert MTB/RIF or culture or requiring symptoms for sputum collection modified the estimated annual percentage change (Table P and Fig D in S1 Appendix). Download: PNG larger image TIFF original image Fig 4. Estimated annual male-to-female ratios of bacteriologically confirmed TB between 1994 and 2020. Panel A: All included surveys. Panel B and C: World regions. Notes: Shaded area represents 95% credible interval. Points on the plot show the empirical male-to-female prevalence ratio for each study, with size corresponding to 1/the standard error of the log odds ratio of male-to-female prevalence, which is used to weight the relative influence of each data point on the regression. Due to the small number of surveys and subsequent uncertainty, we do not present temporal analysis for the Americas (N = 2) or Eastern Mediterranean (N = 4) regions. https://doi.org/10.1371/journal.pmed.1005114.g004 In univariable exploratory analyses using data from 38 of 40 nationally-representative surveys, we estimated a positive association between the M:F ratio and GDI (multiplicative effect: 1.15; 95% CrI: 1.03, 1.28; Table 1). In multivariable analysis, the uncertainty in this estimate widened (1.16; 95% CrI: 1.00, 1.34). We also estimated a positive association with excess HIV/AIDS prevalence in men (1.18; 95% CrI: 1.03, 1.35). Together, the included covariates accounted for 60% (95% CrI: 38%, 78%) of the variation observed, as measured by Bayes R2. These estimates were robust to prior choice (Table Q in S1 Appendix). We assessed the correlation between covariates and found minimal evidence of strong correlation. The strongest correlation was observed between alcohol use disorder and GDI (0.36; 95% CrI: 0.10, 0.58); Fig E in S1 Appendix presents the estimated covariate correlations. We also estimated M:F ratios for each of the countries included in the multivariable analysis (Table R in S1 Appendix); these ranged from 1.80 (95% CrI: 1.15, 2.69) in Eswatini to 3.79 (95% CrI: 2.48, 5.57) in Viet Nam. In subgroup analysis, surveys that were nationally representative (N = 40), with low risk of bias (N = 46), or that did not require symptom screening for sputum collection (N = 66) all had M:F ratios of bacteriologically confirmed TB prevalence greater than 2.0 (Table 2). All of the five countries that had five or more surveys—India (N = 27), Ethiopia (N = 11), China (N = 6), Bangladesh (N = 5), Viet Nam (N = 5)—had mean M:F ratios greater than 1.0, but only India, Bangladesh, and China had M:F ratios for which the 95% credible interval excluded the null (1.0), with 3.28 (95% CrI: 2.68, 3.91), 3.25 (95% CrI: 1.54, 5.57), and 2.65 (95% CrI: 1.77, 3.57) times higher prevalence among males compared to females, respectively. We also found that surveys published after the Horton and colleagues (2016) review had a higher pooled M:F ratio of 2.34 (95% CrI: 1.95, 2.79) compared with those included in the previous review, which had a pooled M:F ratio of 1.79 (95% CrI: 1.39, 2.25). Download: PNG larger image TIFF original image Table 2. Estimated male-to-female ratios of bacteriologically-confirmed TB among key survey subgroups. https://doi.org/10.1371/journal.pmed.1005114.t002 Among the 102 surveys included in this analysis, 46, 41, and 14 were assessed to have low, moderate, and high risk of bias, respectively; one survey did not report sufficient information to assess the risk of bias. Of the assessed criteria, surveys were evaluated to have minimal nonresponse bias (51/102 surveys) or presented sufficient information to evaluate study population representativeness (49/102 surveys) less frequently than the other criteria. Fig F in S1 Appendix shows the distribution of studies across bias assessment criteria across the overall bias assessment levels. We found no evidence of publication bias as evaluated via the doi plot and an estimated LFK index of 0.08 (value < |1| indicative of no asymmetry; Fig G in S1 Appendix) [128–130]. Discussion This systematic review and meta-analysis synthesized data on more than 4 million participants in 102 community-representative prevalence surveys from five global regions. We found that bacteriologically confirmed TB prevalence was over twice as high among men compared with women in LMICs. Strikingly, 90 of 102 surveys had an estimated TB prevalence among men that exceeded that among women, with a pooled M:F prevalence ratio of 2.02 (95% CrI: 1.71, 2.34). These findings are consistent with other studies, including the previous systematic review we updated in this study [7]. As the identified surveys collected data across a 30-year (1994–2024) period, we were able to analyze the change in M:F prevalence ratios over time. These estimates suggest that male-to-female inequalities in TB prevalence appear to be growing (albeit with substantial uncertainty), even with increased acknowledgement of men’s disproportionate TB burden by WHO and other public health organizations [12,131]. This rising trend was clearer in WHO Africa region, with an estimated 3% annual increase in the M:F prevalence ratio. These trends reflect the persistent (and potentially growing) impact of factors driving sex differences in TB prevalence. Despite strong evidence pointing to excess TB burden among men, and increasing recognition of this excess burden [1,12,13], efforts to address this difference have been insufficient to produce statistically discernable reductions in the M:F ratio of bacteriologically confirmed TB prevalence. To further investigate the change in the M:F ratio over time, we also analyzed surveys in 5-year subgroups. Of note, while the estimated difference in TB prevalence was among surveys with data collection coinciding with the COVID-19 pandemic (2020–2024), this estimate pooled the results of only six (two nationally-representative) surveys. Therefore, the estimated M:F ratio for this subgroup should be interpreted with caution. The differential impact of the pandemic on TB burden by sex is varied and likely context-specific. A study of TB notifications during the COVID-19 pandemic found evidence supporting differences in the association between missed or delayed diagnoses due to COVID-19, with 22.5% of the 40 examined countries showing a greater impact among women and a similar percentage showing greater impact among men [132]. In further subgroup analyses and alternative model specification, we estimated the sex differences among surveys that required symptoms for sputum collection and those that did not. We found evidence that requiring symptoms was associated with a substantial reduction in the estimated M:F ratio of bacteriologically confirmed TB. This finding suggests that symptom reporting is a sex-differential measurement mechanism. Consequently, approaches to TB detection that require symptom reporting may lead to further under-detection of TB among men. This finding is consistent with previous literature reporting differential symptom-reporting behaviors by sex [133,134]. Improving the uptake of TB prevention and care among men is essential to ending the TB epidemic and ensuring that the global TB response is person-centered and gender-responsive. In exploratory multivariable analyses, we found a significant association between higher male HIV prevalence (relative to females) and higher M:F TB prevalence ratios. This finding suggests that in countries with a higher excess HIV prevalence among men, the M:F TB prevalence ratio might be elevated. On an individual level, HIV increases the risk of progression from Mycobacterium tuberculosis infection to TB disease [135–137], the severity of TB disease [138,139], and TB-associated mortality [140,141]. While the majority of people living with HIV now receive anti-retroviral therapy (ART), UNAIDS estimates suggest that men are less likely to know their HIV status, be enrolled in ART or achieve viral suppression [11]. ART is a proven tool in the reduction of TB incidence among people living with HIV [142,143]; therefore, increasing the uptake and continuation among men could reduce TB prevalence in this population as well. We also estimated a positive association between the M:F prevalence ratio and the GDI. The GDI is calculated as the ratio of a country’s human development index (HDI) among women compared with the HDI among men. A higher value of the GDI indicates lower inequities in health, education, and command of financial resources faced by women. The positive relationship estimated in our analyses suggests greater M:F ratios of bacteriologically confirmed TB prevalence in countries with more equitable gender development. The mechanisms generating this relationship are unclear, and it could be induced by multiple other factors correlated with both GDI distribution and TB risk factors. For this reason, any causal interpretation of these associations is speculative. We leave these questions for future research to further elucidate the drivers of sex differences in TB. Structural and cultural factors may also contribute to men’s higher TB prevalence beyond the risk factor differences we were able to include in our analysis. Previous research has highlighted how gender norms and expectations influence healthcare access and utilization among men [144]. Men have lower rates of healthcare attendance compared to women, which may be due to perceived stigma, weakness, or lack of control [145] and competing priorities, particularly in settings where men are expected to fulfill provider roles [133]. Adapting the healthcare experience to be gender-responsive and male-friendly may increase TB diagnoses and successful treatment among men [146]. Qualitative studies of male TB survivors, stakeholders, and healthcare workers have identified preferences for modified TB screening interventions that incorporate outreach to male-dominated workplaces and sociocultural settings, as well as the creation of male-only spaces and communication materials [147,148]. These priorities align with WHO’s identification of “men in settings where healthcare access is not tailored to their needs” as a key TB vulnerable population in 2025 [13]. If sex differences in TB prevalence are growing, the interventions aimed at reducing excess TB burden among men have increasing urgency and impact. Community-based interventions have the potential to substantially improve men’s engagement with the TB care cascade. Previous research has investigated the feasibility of screening at occupation centers or key transportation hubs [149–151]. Additional interventions have further attempted to “meet men where they are” through outreach to male-dominated socio-cultural centers, such as bars, betting halls, or churches [152]. Implementing a combination of health facility- and community-based interventions may increase men’s access to diagnosis and treatment, thereby reducing TB burden among men. TB intervention development can be further informed through previous research done to increase men’s engagement in other disease care cascades, such as HIV testing and treatment and diabetes management [153–155]. Evaluation of the implementation of these community-based strategies and developing guidance for male-focused TB screening could be incorporated into strategies to reduce TB incidence in LMICs. This study has several limitations. The geographic distribution of included surveys was uneven, with over-representation of large, high TB burden countries. While this may bias global generalizability, such concentration arguably reflects an appropriate focus of prevalence surveys on settings with the greatest burden and need. Heterogeneity in survey methods—such as differences in sampling strategies, screening algorithms, and diagnostic tools—introduces additional nonsampling variability that may affect cross-survey comparability. Our definition of bacteriologically confirmed TB was dependent on the methods employed in each survey and could not be fully standardized. To mitigate this, we prioritized test results from more sensitive diagnostics (e.g., Xpert MTB/RIF and culture) when available, although some variation remains. The use of symptom-based screening could also introduce bias in sex-stratified analyses, as the prevalence of TB-related symptoms and their reporting rates may differ by sex, potentially affecting the male-to-female prevalence ratio. We addressed this by attempting to quantify this difference in the included surveys. Furthermore, the estimated M:F ratios of TB prevalence may reflect social and behavioral factors, as discussed above, or an imbalance of demographic or health risk factors, such as age or HIV status, either in the underlying population or survey sample. While we did detect a relationship between excess HIV and TB prevalence among men, our covariate analyses did not account for measurement error in country-level TB determinants, which may have affected coefficient estimates. As with any systematic review, our estimates may reflect publication and language biases. Our data were restricted to prevalence surveys with publicly available data, either through formal published reports or manuscripts indexed in the electronic databases searched. We did not find evidence of effect size bias in our dataset, suggesting the impact of any publication bias is small. Our data were further limited to only English language studies; to our knowledge, only one prevalence survey was excluded due to this reason. A robustness check found no additional relevant surveys in Spanish, French, or Portuguese. Finally, despite the large overall sample size, uncertainty remains substantial for many of the relationships examined. The 95% credible intervals for several estimates were inclusive of the null relationship or overlapping in subgroup analyses; these findings are directionally consistent with existing evidence but should be interpreted cautiously due to the uncertainty of these estimates. In summary, we found strong evidence that adult men have over twice the prevalence of pulmonary TB as compared to women in LMICs. These estimated differences in male-to-female bacteriologically confirmed TB prevalence may have increased over time, in particular in WHO Africa region. If sex-related inequalities in TB burden are growing, developing effective strategies to reduce men’s risk of TB and to engage men in TB prevention and care will be essential to end TB. Supporting information S1 Appendix. Supplementray Tables A–R and Figures A–G. Table A: Full search strings for each database in the systematic review. Table B: Inclusion and exclusion criteria for systematic review study selection. Table C: Full text review exclusion reasons. Table D: Search strings for publications in French, Spanish, and Portuguese as a robustness check of the search strategy. Fig A: Decision tree for reported estimates used in meta-analysis of sex-stratified prevalence bacteriological positive TB estimates. Table E: Model equations and corresponding priors. Table F: Surveys with missing covariate data and the corresponding management approach. Table G: Studies excluded at full-text review with their corresponding exclusion reasons. Table H: Selected characteristics and reported data of 102 TB prevalence surveys included in quantitative analysis. Table I: Cumulative characteristics of study participants in the included prevalence surveys (N=102). Note: not all surveys included all reported characteristics. Bacteriologically-confirmed TB count is the sum of reported bacteriologically-confirmed case counts and case counts estimated from reported prevalence risks and survey participants. Of 102 surveys, 82 reported case counts, 16 were estimated, and 4 were missing the information to estimate case counts. Similarly, among 64 surveys reporting smear-positive TB prevalence, 54 reported case counts and 10 surveys had case counts estimated. Culture-positive TB case counts were reported by 27 surveys. Table J: Model performance statistics. Fig B: Density and trace plots for selected model fits. Model 1: i. Main effect. Model 2: ii. Main effect across world regions. Model 3: iii. Univariable analysis of the study end year on the main effect (fit to all surveys). Model 4: iv. Univariable analysis of study end year on main effect across world regions (fit to all surveys). Model 35: v. Multi-variable analysis of covariates on main effect. Model 36: vi. Main effect fit to smear-positive TB prevalence ratios. Model 37: vii. Main effect across world regions fit to smear-positive TB prevalence ratios. Table K: Male-to-female ratio of bacteriologically-confirmed TB prevalence as estimated by alternative model hierarchical structures. Model descriptions: 1. Main model specification with country- and survey-level random effects. 2. Alternative model with country-level random effects nested in WHO region random-effects and survey-level random effects. Table L: Estimated impact of screening and diagnostic algorithms on the male-to-female ratio of bacteriologically-confirmed TB prevalence. Note: coefficients are reported on the natural log-scale. Model descriptions: 1. Main effects model with country- and survey-level random effects. 2. Main model with an additional binary variable describing whether a survey required symptoms for sputum collection. 3. Main model with an additional binary variable describing whether Xpert or culture was required for a bacteriologically confirmed diagnosis covariate. That is, an individual must receive an Xpert or culture test and at least one of these must be positive for a bacteriologically-confirmed TB diagnosis. 4. Main model with an additional binary variable describing whether Xpert and/or culture testing were conducted. That is, the survey used Xpert and/or culture (in addition to or instead of sputum smear microscopy), and a positive result on any of these tests was sufficient for bacteriologically-confirmed TB diagnosis. Fig C: Posterior heterogeneity across main effect models with country or WHO region random effects. Table M: Posterior summary of variance in main effect models with country or WHO region random effects. Note: World Health Organization (WHO). Table N: Male-to-female prevalence ratio estimates for bacteriologically-confirmed TB and smear-positive TB by world region. Notes: Bacteriologically-confirmed (BC); smear-positive (SP). Estimates represent the posterior mean (posterior median; 95% credible interval). Table O: Estimated annual percentage change in male-to-female ratios of bacteriologically-confirmed TB prevalence by world region. Table P: Estimated effect of screening and diagnostic algorithms on the temporal model of male-to-female ratio of bacteriologically confirmed TB prevalence. Note: coefficients are reported on the natural log-scale. Model descriptions: 1. Main temporal model specification with survey- and country-level random effects. 2. Main temporal model with an additional binary variable describing whether a survey required symptoms for sputum collection. 3. Main temporal model with an additional binary variable describing whether a survey required symptoms for sputum collection and accounts for an interaction between this variable and the study end year. 4. Main temporal model with an additional binary variable describing whether Xpert or culture was required for a bacteriologically confirmed diagnosis. That is, an individual must receive an Xpert or culture test and at least one of these be positive for a bacteriologically-confirmed TB diagnosis. 5. Main temporal model with an additional binary variable describing whether Xpert or culture was required for a bacteriologically confirmed diagnosis and accounts for an interaction between this variable and study end year. 6. Main temporal model with an additional binary variable describing whether Xpert and/or culture testing were available. That is, the survey used Xpert and/or culture (in addition to or instead of sputum smear microscopy), and a positive result on any of these tests was sufficient for bacteriologically-confirmed TB diagnosis. 7. Main temporal model with an additional binary variable describing whether Xpert and/or culture testing were conducted and accounts for an interaction between this variable and study end year. Fig D: Trends in the male-to-female ratio of bacteriologically-confirmed TB prevalence as estimated by alternative model specifications. i. Symptoms required versus symptoms not required. ii. Xpert or culture is required versus not required for bacteriologically-confirmed TB diagnosis. iii. Xpert or culture used versus not used for bacteriologically-confirmed TB diagnostic pathway. Model descriptions: i. Main model with a binary variable for whether symptoms were required for sputum collection. Accounts for an interaction with survey end year. ii. Main model with a binary variable for whether Xpert or culture is required for bacteriologically confirmed diagnosis covariate. That is, an individual must receive an Xpert or culture test and at least one of these is required for a bacteriologically-confirmed TB diagnosis. Accounts for an interaction with the survey end year. iii. Main model with a binary variable for Xpert or culture used for diagnosis covariate. That is, an individual must receive an Xpert or culture test, but a smear-positive test was sufficient for a bacteriologically-confirmed TB diagnosis. Accounts for an interaction with survey end year. Fig E: Posterior distribution of the correlation of model covariates. Note: the dot represents the mean of the posterior distribution, and the black line represents the 95% credible interval. Table Q: Estimated impact of alternative priors on the covariate model of the male-to-female ratio of bacteriologically-confirmed TB prevalence. Note: Coefficients are reported on the natural log-scale. Model descriptions: All models have the same model specification, which includes study- and survey-level random effects. These models are all fit to nationally-representative survey data. Models differ only in their prior on the covariate regression coefficients: 1. Normal distribution with mean = 0 and standard deviation = 1. 2. Normal distribution with mean = 0 and standard deviation = 10. 3. Cauchy distribution with location = 0 and scale = 1.25. Table R: Country-level estimates of male-to-female ratios of bacteriologically-confirmed TB prevalence as calculated from the multivariate regression model. Fig F: Distribution of overall risk of bias by each assessed criterion. Fig G: Doi plot to evaluate potential publication bias among included prevalence surveys. Note: The figure above is used to examine symmetry across effect sizes. This symmetry is quantified through the LFK index, where | LFK value < 1 indicates no evidence of asymmetry. https://doi.org/10.1371/journal.pmed.1005114.s001 (DOCX) S1 Form. Extraction form for systematic review of tuberculosis prevalence in low- and middle-income countries. https://doi.org/10.1371/journal.pmed.1005114.s002 (DOCX) S1 Checklist. MOOSE checklist. https://doi.org/10.1371/journal.pmed.1005114.s003 (DOCX) S2 Checklist. PRISMA 2020 checklist (with PRISMA 2020 Abstract checklist). From: Page MJ, McKenzie JE, Bossuyt PM, Boutron I, Hoffmann TC, Mulrow CD, and colleagues. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. MetaArXiv. 2020, September 14. https://doi.org/10.31222/osf.io/v7gm2. For more information, visit: www.prisma-statement.org. https://doi.org/10.1371/journal.pmed.1005114.s004 (DOCX) Acknowledgments We thank Stephanie Su for assistance in study coordination.
Climate change and non-communicable diseases: An invisible syndemicParameswaran, Gokul;Al-Kindi, Sadeer;Rajagopalan, Sanjay
doi: 10.1371/journal.pmed.1005082pmid: 42101971
Planetary disruption and disease: An integrated framework Climate change generates multiscale disruptions across different planetary systems—the atmosphere, hydrosphere, pedosphere, cryosphere, biosphere, geosphere, and anthroposphere—that create interacting, cascading, and compounding exposures (Fig 1). The effects of disruptions on NCDs can be both direct and indirect and may operate across varying time scales. The most established direct pathway operates through atmospheric warming. Heat, an independent cardiometabolic stressor, has driven a 63% rise in heat-related mortality since the 1990s, exceeding 546,000 deaths annually [2]. Warming simultaneously degrades air quality by accelerating pollutant formation, hindering pollutant dispersal, and creating drier conditions that intensify wildfires and dust storms [5]. Together, these exposures interact to produce multiplicative increases in mortality, driving multi-organ NCD burden across cardiovascular, respiratory, and neurological systems [5]. In one of the first studies to causally link climate change with wildfire smoke and premature mortality, warming climate conditions are projected to cause ~71,000 excess annual deaths and nearly 1.9 million cumulative excess deaths between 2026–2055 [6]. Beyond physical health, increasing frequency of extreme weather events worsens mental health by increasing mood disorders and trauma-related illness [7]. Indirect pathways, meanwhile, operate primarily through disruption of the systemic infrastructure protecting health. Hydrosphere and pedosphere disruptions characterized by altered precipitation, drought, and soil degradation erode freshwater availability and nutritional quality. The resulting climate-related shifts in agricultural output risk increasing reliance on calorie-dense ultra-processed foods, accelerating metabolic disease, and potentially driving over 500,000 deaths annually by 2050, with many of these attributable to CVD and cancer [8,9]. Cryosphere changes, including glacial melt and sea level rises, may result in population displacement, severing communities from stable food and healthcare access. These pressures converge on the anthroposphere with urban systems, healthcare infrastructure, and supply chains buckling under compounding climate stressors. Behavioral adaptations compound these risks, as hostile temperatures may discourage physical activity and increase harmful indoor exposures, such as fungal spores from mould or volatile organic compounds from paints [10]. Collectively, these interacting disruptions transform climate change from a singular threat into a syndemic driver, overwhelming public health models focused on isolated risk factors than integrated planetary-scale exposures. Download: PNG larger image TIFF original image Fig 1. The climate-NCD syndemic iceberg: Hidden dimensions of climate-driven disease burden. Climate change generates a constellation of environmental stressors—including extreme heat, wildfire smoke, water insecurity, food system disruption, urban heat islands, and healthcare disruption—that drive the visible burden of infectious disease outbreaks. Beneath the surface lies a far larger, largely invisible syndemic of climate-driven non-communicable diseases, encompassing cardiovascular disease, cancer, diabetes, chronic respiratory disease, dementia, and mental illness. Figure created using Nano Banana AI. https://doi.org/10.1371/journal.pmed.1005082.g001 Climate vulnerability: Disproportionate burdens and cumulative disadvantage Climate change disproportionately burdens socioeconomically disadvantaged populations, constituting a fundamental equity crisis. Most Global South countries sit at the geographic fault lines of climate disaster, where rising sea levels, shifting weather patterns, and extreme events cascade across ecosystems to expose systemic vulnerabilities [1]. Reliance on climate-sensitive sectors, fragile infrastructure, and chronically underfunded health systems limits capacity to respond to disasters or adapt to shifting planetary systems [1]. These asymmetries translate climate shocks into compounding NCD burdens across already vulnerable regions of the planet, despite historically minimal emissions. In the U.S., formerly redlined neighborhoods—systematically denied investment based on race or ethnicity—often feature dense impervious surfaces, limited tree canopy, aging housing, and proximity to major roadways, amplifying pollution exposures and fostering obesogenic environments [11]. Occupational exposures further compound risk, as lower-income communities predominantly work outdoor or manual labor jobs with high heat exposure [5]. National analyses have shown higher climate vulnerability measures, integrating cumulative environmental stressors, pollution, and adaptive capacity, are associated with greater risk of cardiometabolic disease, independent of other baseline social and infrastructural features [11]. Why we fail to act: Barriers to prevention Despite mounting evidence linking climate change to NCD mortality, there is a large gap in implementing adaptation and mitigation strategies. NCDs operate through diffuse and delayed pathways obscuring causation and enabling political gridlock. Partisan polarization has led to trust deficits fueling climate denialism and obstructing interventions. Recent U.S. decisions exemplify this: withdrawal from the Paris accord, rollback of environmental regulations, reduced authority of the Environmental Protection Agency, and halting renewable investments have reversed American decarbonization efforts [12]. Policy resistance arises from misalignment between short-term political horizons and long latency of environmental exposures. Individuals often underestimate environmental risks given the protracted development of NCDs and near invisibility of these exposures. Unlike pandemic preparedness, which commands emergency appropriations, climate-NCD prevention receives no comparable funding, despite vastly exceeding ID mortality. Economic barriers further impede action, as adaptation requires substantial upfront investment, including retrofitting power/transportation networks, building climate-resilient infrastructure, or providing community education, which only gradually accrue benefits. This mismatch discourages long-term political commitment, especially when climate adaptation competes with more immediate priorities. Towards a multi-level adaptation strategy Current and projected climate impacts demand urgent adaptation strategies within communities, environmental infrastructure, and healthcare systems [3]. Community interventions must include targeted education for vulnerable populations, including teaching practical avoidance strategies at peak exposures. In addition, subsidizing adaptive technologies like air conditioning or air purifiers for high-risk households or communities may prove cost-effective. Infrastructure interventions will require interdisciplinary collaboration, such as healthcare professionals engaging with urban planners in designing climate-resilient environments through increased vegetation, shaded pathways, and reflective roofing. Early-warning systems integrated with public cooling/clean-air centers would also provide critical refuge during extreme events. In healthcare settings, environmental risk assessments should be incorporated into patient counseling, especially for vulnerable populations on drugs affecting thermoregulation [5]. Healthcare systems must also strengthen disaster preparedness through climate-resilient infrastructure, ensuring supply chain continuity and expanding decentralized infrastructure via mobile clinics to maintain access during emergencies [5]. Reframing climate change as a health imperative makes visible its direct role in shaping the global burden of NCDs, and anchors climate action within clinical, public health, and policy decision-making. Because climate harms are unequally distributed, only an equity-centered approach that meaningfully engages the Global South and structurally disadvantaged populations may generate durable solutions [1]. Conclusions The climate–NCD syndemic is likely the defining public health issue of the 21st century. Meeting this challenge will require not only innovation, but ethical leadership and political courage to confront entrenched systems that sustain delay. Anchoring climate action in the universal imperative of health may be the clearest pathway to bridging ideological divisions and catalyzing sustained collective action.