Abstract
Estimating causal effects in the presence of spillover among individuals within a social network poses challenges due to missing information. Spillover effects refer to the impact of an intervention on individuals not directly exposed themselves but connected to intervention recipients within the network. In network-based studies, outcomes may be missing due to study termination or participant dropout, termed censoring. We introduce an inverse probability censoring weighted estimator which extends the inverse probability weighted estimator for network-based observational studies to handle possible outcome censoring. We prove the consistency and asymptotic normality of the proposed estimator and derive a closed-form estimator for its asymptotic variance. Applying the inverse probability censoring weighted estimator, we assess spillover effects in a network-based study of a nonrandomized intervention with outcome censoring. A simulation study evaluates the finite-sample performance of the inverse probability censoring weighted estimator, demonstrating its effectiveness with sufficiently large sample sizes and number of connected subnetworks. We then employ the method to assess spillover effects of community alerts on self-reported human immunodeficiency virus risk behavior among people who inject drugs and their contacts in the Transmission Reduction Intervention Project (TRIP), from 2013 to 2015, Athens, Greece. Results suggest that community alerts may help reduce human immunodeficiency virus risk behavior for both the individuals who receive them and others in their network, possibly through shared information. In this study, we found that the risk of human immunodeficiency virus behavior was reduced by increasing the proportion of a participant’s immediate contacts exposed to community alerts.
Keywords
Introduction
Estimating causal effects in the presence of spillover (dissemination) among individuals in a network is often complicated by missing information. We need to consider missing outcome data when evaluating spillover within networks due to the dependencies between individuals. Ignoring missing data could not only introduce selection bias but also distort the network structure, resulting in possible unmeasured confounding or mismeasured exposures and covariates. There are various sources of missingness in the data, such as missing outcomes, baseline covariates, exposure status, and network connections. However, we focus on missing outcomes due to potential censoring, for example, the differential loss to follow-up due to study dropout. This censoring can occur when participants who drop out differ from those who do not with respect to their exposure or outcome (or both).
One common approach for estimating causal effects is to exclude all observations with missing outcomes, known as complete case analysis. The complete case estimator is consistent if the observations are missing completely at random (MCAR), indicating that missingness is unrelated to any measured and unmeasured variables. When the MCAR assumption does not hold, and the missing at random (MAR) assumption is appropriate for the observed data, several approaches such as maximum likelihood, multiple imputation, fully Bayesian, and inverse probability weighting1,2 can be applied in this case.
Social networks can be represented as graphs illustrating connections among individuals, such as those delineated by engagement in human immunodeficiency virus (HIV) risk behaviors (e.g. sexual or injection behaviors) among people who inject drugs (PWID). The exposure of one individual could affect the outcome of another, a phenomenon known as interference. 3 In this context, a key interest lies in the spillover (or indirect) effect, often conceptualized as a contrast in average potential outcomes if an individual is unexposed, comparing different exposure vectors for their spillover set (e.g. neighbors). In network-based studies, the spillover (or interference) set can be defined as an individual’s neighbors, namely those who share a connection (edge or link) with that individual. Therefore, the assumption of partial interference, which posits that spillover is possible within a set or group of individuals but not between groups,4,5 may not hold. That is, while no interference occurs between individuals in different groups, interference is possible among individuals within the same group defined by the interference set. Additionally, the collection of interference sets forms a partition of the study population.
Inverse probability weighted (IPW) estimators for causal effects in observational studies in the presence of interference are often formulated under the assumption of partial interference. 3 Alternatively, some approaches define interference based on spatial proximity or network ties,6,7 allowing for possibly overlapping interference sets. Liu et al. 7 proposed an IPW estimator for a generalized interference set that permits overlaps between interference sets; however, the asymptotic variance was estimated under the assumption of partial interference defined by larger groupings or clusters. Forastiere et al. 6 quantified effects using a subclassification estimator with a generalized propensity score, employing a bootstrapping procedure with resampling either at the individual-level or the cluster-level to quantify variance. Nevertheless, these methods either hinge on partial interference defined by larger clusters or resort to bootstrapping for variance estimation. In practice, overlooking overlapping interference sets during variance estimation can yield inaccurate inference, and resampling approaches may be computationally intensive.
To address these challenges, the study by Lee et al. 8 utilized the IPW estimators proposed by Forastiere et al. 6 and Liu et al., 7 and introduced closed-form variance estimators capable of incorporating overlapping interference sets. Through statistical simulations, their proposed variance estimators demonstrated greater efficiency in network-based studies by leveraging additional information on connections between individuals. McNealis et al. 9 proposed doubly robust estimators that include models for both treatment and outcome, ensuring consistency and asymptotic normality as long as at least one of the models is correctly specified. Additionally, alongside the IPW approach, van der Laan 10 proposed a targeted maximum likelihood estimator for estimating causal effects in time invariant and longitudinal network studies. Their focus was on specific types of causal quantities, particularly the counterfactual mean under stochastic intervention on the unit-specific treatment nodes. When the outcome of interest is subject to possible right censoring (i.e. events occurring after a certain period of follow-up are missing), Chakladar et al. 11 and Loh et al. 12 extended IPW estimators under partial interference in observational studies to accommodate censoring using inverse probability of censoring weighted (IPCW). In these methods, survival models, such as proportional hazards frailty models and accelerated failure time models, were employed to estimate censoring weights. Moreover, a grouping of observations was utilized to define the interference set (e.g. study clusters and geographic location), allowing for spillover within but not between clusters. However, these approaches did not account for spillover between individuals with connections (i.e., edges) within a cluster or component, such as those observed connections with HIV risk transmission risk in network-based studies. 13 Liu et al. 7 proposed generalized weighted type estimators with neighbor-level exposure weights, relaxing the partial interference assumption. However, this method did not incorporate censoring weights, and the variance was estimated assuming partial interference.
In this work, we developed methods to estimate causal effects in network-based observational studies in the presence of missing binary or continuous outcomes due to possible censoring. In this study, we employed a IPW estimator by Lee et al., 8 which allows for spillover to each participant from their immediate connections in the network (i.e., neighbors) to quantify causal effects in network-based observational studies with a nonrandomized intervention. More precisely, each participant has a unique interference set determined by the observed network and interference sets can now overlap and no longer partition the network. We extended their IPW estimator to a setting with missing outcomes due to censoring using IPCW, where we consider two different censoring mechanism: (1) censoring indicators are independent across participants conditional on baseline covariates and exposures or (2) censoring indicators are correlated between participants within a connected subnetwork or component in the network. Components are defined as subsets of connected participants, but not connected with participants in other components. We used the network structures including network connections between individuals and the independence of components to calculate a novel closed-form variance estimator by applying M-estimation. We evaluated finite-sample performance and the impact of the dependency assumption between censoring indicators of the proposed estimators in a simulation study.
We employed the proposed IPCW estimator to evaluate the direct and spillover effects of community alerts on self-reported HIV risk behavior at 6 months among PWID and their contacts in the Transmission Reduction Intervention Project (TRIP) from 2013 to 2015 in Athens, Greece.14–16 Previous work found evidence of a possibly meaningful spillover effect of community alerts on HIV risk behavior in TRIP; however, the prior study used a complete case analysis and did not consider the impact of missing outcomes on the validity of the analysis. 8 PWID often participate in behaviors with HIV transmission risk and are part of sexual, social, and drug use networks. In these networks, nodes represent individuals and edges between nodes represent the relationship between two people, such as sexual, social or drug use contacts, referred to as HIV risk networks. These individuals may or may not engage in behavior that increases HIV transmission risk in their partnership, including sharing equipment for injection drug use or condomless sex. In networks of PWIDs, educational and treatment-based interventions often have spillover causal effects, and intervention effects frequently depend on the network structure.17,18
The proposed IPCW estimator is developed in Section 2. In Section 3, the results of a simulation study are reported that evaluates the performance of IPCW in finite-sample settings. In Section 4, we report the results that used the IPCW estimator to analyze the spillover effects of community alerts on self-reported HIV risk behavior among PWID in TRIP. We then discuss the performance of the IPCW estimators in the simulation study, limitation of the method, and future directions in Section 5.
Methods
In this section, we propose an IPCW estimator for network-based studies with binary or continuous outcomes subject to possible censoring due to differential loss to follow-up due to study drop out. If participants who drop out of the study differ from those who did not with respect to their outcome or exposure (or both), then this may result in selection bias, even under the null. 19 There are several types of missing mechanisms including MCAR, MAR, and missing not at random (MNAR). In the case of MCAR, the missing data mechanism is unrelated to any measured or unmeasured study variable. That is, the participants with completely observed data are assumed to be a random sample of all the participants assigned a particular intervention, which is often an unrealistic assumption in practice. MAR occurs when the missingness can be accounted for by fully observed study variables. 20 In other words, censoring indicators are independent conditional on baseline variables (e.g. covariates and exposures). MNAR assumes the missingness depends on underlying variables that we do not observe, such as participants with more frequent injection drug use with others are more likely to drop out of the study. In this article, we focus on the MAR assumption, but the censoring models could be extended to some cases of MNAR.21,22
The rest of this section is organized as follows: We first state the notation in Section 2.1 and the assumptions in Section 2.2. We then define the potential outcome framework and the population average causal effects of interest in Section 2.3. In Section 2.4, we define the inverse probability censoring weight estimator with censoring weights and neighbor-level exposure weights. The large sample variance estimators are derived using M-estimation theory 23 in Section 2.5.
Notation
A network is a structure consisting of a set of nodes,

Left panel is the Athens people who inject drugs (PWID) network. Right panel is enlarged one of the network components.
The degree of unit
The proposed approach requires identification assumptions in this observational network design.
The neighbor interference assumption
6
assumes that the outcome of participant
Let
We assume that the intervention assignment mechanism does not affect the outcome beyond the assigned interventions themselves.
6
That is, if multiple versions of the intervention exist, they are considered irrelevant to the causal contrasts of interest. We assume a single version of the intervention exposure and a single version of no exposure. For each
There are no multiple versions of this intervention or exposure, and if there were multiple versions of exposure, those are assumed to be irrelevant for the causal effect of interest. 24
The conditional exchangeability assumption,
8
Assumption 3 states that exposure is independent of potential outcomes, conditional on the pre-exposure covariates of participant
Positivity assumption for the intervention exposure,
7
Positivity assumes that each participant has a non-zero probability of being assigned every possible intervention exposure status given every combination for the levels of the observed covariates.
Conditional independent censoring
Assumption 5 states that the censoring of participant
(Identification) Under Assumptions 1–3, we have
The proof of this proposition is in Supplemental Appendix A.
We define individual average potential outcomes under counterfactual intervention assignment strategies specific to the individual and their neighbors, rather than network-wide estimands. We take the expectation with respect to the population of individuals from which our network components were sampled. In particular, we consider strategies that assign an individual
Recall that
We also consider counterfactual treatment strategies that assign the intervention to both individual and their contacts using a Bernoulli assignment strategy with probability
Note that this definition is similar to that by Hudgens and Halloran,
4
but here the Bernoulli assignment is applied to the neighbors of unit
Different contrasts of these population average causal effects are often of interest in public health research. In particular, spillover effects can assess the public health impact of an intervention in the network among individuals who were not exposed to intervention but who are connected to recipients. An example of this type of intervention would be one delivered through contact tracing, where an individual and their contacts possibly received the intervention. We may be interested in the population-level impact of this intervention delivered via contact tracing. We defined these effects under allocation strategies
According to Assumption 1, spillover is assumed to be possible to a participant from their neighbors in the network (i.e. a participant’s immediate or first-degree contacts). The spillover set or interference set of participant
In a network-based study, data are ascertained from a network with
For the causal effects of interest, we define two estimators. In this article, our estimands are not defined at the component level, but instead average individual-specific estimands, so we average over all individuals in the study (or take the expectation with respect to the population of individuals from which our network components was sampled). We are scientifically interested in the participant-average estimand, so we define our estimator similarly for this parameter. Under Assumptions 1–5, an extension of the nearest neighbor IPW estimator
8
of the study population average outcome for exposure
Under allocation strategies
If the exposure propensity scores and censoring weights are known, then the IPCW estimator is unbiased:
The proof of this proposition is in Supplemental Appendix A. We can easily get unbiased estimators for the causal effects because the causal effects are different contrasts of the average potential outcomes.
Following the approach of Lee et al.,
8
we used a mixed-effects model to estimate the exposure propensity score for each participant and their interference set
Different censoring mechanisms require distinct model specifications to appropriately account for underlying dependencies and variability. When censoring indicators exhibit multiple sources of correlation among participants such as shared characteristics within a network, spatial clustering, or unobserved heterogeneity at different levels, models incorporating multiple random effects may be necessary to effectively capture these dependencies.
In this article, we consider a censoring mechanism in which the censoring indicators of participants within a component may be correlated, allowing us to demonstrate the estimators and estimate the model parameters. That is, if
This mechanism states the censoring of participant
To account for potential correlations in the censoring mechanisms among participants within the same component, the conditional probability of censoring can be modeled using a mixed-effects model with a logit link function.
The vectors
The large sample variance estimators can be derived using M-estimation theory.
23
We assume that an observed network can be expressed as the union of connected subnetworks, referred to as components. Given a network with
We denote
Let
We use
Given the set of observed variables
Let
Under suitable regularity conditions that as
Under suitable regularity conditions and due to the unbiased estimating equations,
More details of Proposition 3 can be found in Supplemental Appendix A. A consistent estimator of the variance for the IPCW estimator is given in Supplemental Appendix A. This variance estimator can be used to construct Wald-type CIs for the estimated direct, spillover, total, and overall effects.
A simulation study was conducted to evaluate the finite-sample performance of the IPCW estimators and their corresponding closed-form variance estimators. We focused on assessing finite-sample bias and the coverage of the associated 95% Wald-type CIs. The observed network was generated as follows: we created
Given a fixed, randomly generated network, we simulated a total of 1000 datasets following the steps described below. This design reflects the applied setting in which researchers typically work with a single observed network. Our aim was to evaluate the performance of the proposed estimators conditional on an observed network structure, isolating variability due to the exposure, outcome, and censoring processes. The network characteristics (number of components and number of nodes in each component), parameters in the outcome model, and censoring models were informed by estimates in the TRIP data. We assumed that the censoring indicators were conditionally independent given their own and their neighbors’ covariates and exposures. We also did not find significant associations between the censoring indicators and neighbors’ covariates, their own and their neighbors’ exposures in TRIP analysis. Therefore, we did not include them in our simulation study. The data was generated in the following steps: Generate a baseline covariate Simulate the random effects for the censoring model, The censoring indicators are generated as
We then generate the potential outcomes
The observed exposures are generated as
Based on the observed exposures generated in Step 5, the observed outcome for each individual
For each simulated data set, the
We compared the performance of the estimators for networks with 10, 50, 100, and 200 components. Figure 2 illustrated the absolute value of average bias and ECPs of We considered a scenario with a 50 component network, taking into account component sizes of 20 and 30, which are larger than those in the main scenario, to evaluate the effect on performance when varying the component sizes. The results indicate that networks with larger components exhibited smaller bias, which may be attributed to the larger number of participants in the entire study. We evaluated the impact on the performance of the estimators when there was a higher level of dependency between censoring indicators for the mixed-effects model approach. We used the same simulated regular network with degree 4 and 100 components to compare the performance of the average outcome estimators. The censoring indicators were generated by
We considered the network structure from our motivating study TRIP. The TRIP network consisted of 10 components with 277 nodes and 542 edges. Based on the main simulation scenarios results, a small number of components may result in poor finite-sample performance of variance estimators. In addition, if there is substantial variation in the observed component sizes, the average component size used to derive the estimator of the variance may not be appropriate unless the average potential outcomes in the component are independent of component size.
7
To increase the number of components based on network structure and generate components that were more comparable in size for estimation of the variance of the estimated causal effects, we employed an efficient modularity-based, fast greedy, approach to detect communities to further divide some large connected components of the TRIP network into a total of 24 smaller and denser components (Tables 3 and 4). By ignoring sparser edges between components, we treated the obtained communities as independent units to possibly improve the estimation of the variance with more components similar in size. Furthermore, community detection reduces the magnitude of the degree distribution resulting in estimators with lower variance. Importantly, we still defined the interference sets using the neighbors for point estimation of the causal effects. The results illustrate that the variance estimators were conservative on original TRIP network (10 components) and closer to or below the nominal level for most scenarios with the 24 components network under both censoring mechanisms (Tables 3 and 4). In network data sets with few components or large variability in the size of components, sensitivity analyses with different community detection methods in the network may be appropriate to assess the impact on statistical inference (Puleo et al., 2023).

The average absolute value of bias (left) and empirical coverage probability (ECP) (right) on networks with 10, 50, 100, and 200 components using logistic regression censoring model (top) and mixed-effects censoring model (bottom).
Simulation results of IPCW estimators using the logistic regression censoring model (left) and the mixed-effects censoring model (right) under allocation strategies of 25%, 50%, and 75% for a network with 50 components, each with a component size of 20.
IPCW: inverse probability censoring weighted; ESE: empirical standard error; ASE: average estimated standard error; ECP: empirical coverage probability.
Simulation results of IPCW estimators using the logistic regression censoring model (left) and the mixed-effects censoring model (right) under allocation strategies of 25%, 50%, and 75% for a network with 50 components, each with a component size of 30.
IPCW: inverse probability censoring weighted; ESE: empirical standard error; ASE: average estimated standard error; ECP: empirical coverage probability.
Simulation results from 1000 simulated datasets of IPCW estimators using logistic regression censoring model on original TRIP network (10 components) (left) and the network divided by community detection into 24 components (right).
IPCW: inverse probability censoring weighted; TRIP: Transmission Reduction Intervention Project; ESE: empirical standard error; ASE: asymptotic standard error; ECP: empirical coverage probability.
Simulation results from 1000 simulated datasets of IPCW estimators using mixed effects censoring model on original TRIP network (10 components) (left) and the network divided by community detection into 24 components (right).
IPCW: inverse probability censoring weighted; TRIP: Transmission Reduction Intervention Project; ESE: empirical standard error; ASE: asymptotic standard error; ECP: empirical coverage probability.
Participants in the TRIP study were PWID and their contacts who lived in Athens, Greece between 2013 and 2015. PWID who participated in the ARISTOTLE project at HIV testing centers in Athens were initially recruited into the TRIP study if they were found to be recently infected with HIV. ARISTOTLE was a community-based program aiming to mitigate HIV transmission among PWID by implementing a care system that involved reaching out to high-risk PWID, engaging them in HIV testing, and initiating HIV care, opioid use disorder treatment, and antiretroviral therapy.26,27 In TRIP, each newly diagnosed individual was asked to identify their recent sexual and drug use partners in the last 6 months. These partners were then recruited and asked to identify their sexual and drug use partners, who were also recruited and contacts to other individuals already recruited in the study were also ascertained. If any of these partners were determined to be recently infected with HIV, then their contacts and the contacts of their contacts (i.e. two waves of contact tracing) were recruited as well, including their possible connections to the other participants in the study. Complete details about the TRIP study design can be found in previously published papers.15,28
In addition to HIV testing, the study provided access to treatment as prevention (TasP), referrals for medical care, and distributed community alerts to inform community members about temporary increases in the risk for HIV acquisition. For example, a community alert would be distributed to those in close proximity in the observed network to a recently infected participant. These alerts included paper flyers given to participants and posted in a location frequented by members of the local PWID community. We considered those who received the alerts from the study staff or flyers to be exposed to the community alert, while the remaining participants were not exposed but could have possibly received this information from their exposed neighbors. All participants completed computer-assisted interviews and also had their HIV status ascertained. They provided demographic information, answered questions about engagement in risk behaviors, HIV status, substance use, access to care, HIV knowledge, stigma, injection norms, and their opinions on the project. Follow-up interviews were conducted with participants about six months after they completed their baseline interview. 15
We applied the IPCW estimators to evaluate the causal effects of community alerts on HIV risk behavior ascertained by the report of risk behavior at the 6-month follow-up visit. The exposure community alerts was defined as receiving a study flyer, either from study staff or viewing one posted in the community. We determined the outcome of HIV risk behavior ascertained at the 6-month visit as a self-report of shared drug equipment (e.g. needles and syringes) in the last 6 months. We consider the report of any injection HIV risk behavior at a 6-month visit as a binary outcome. The following pre-exposure baseline covariates are included in the adjusted models for both the exposure and censoring based on expert knowledge and were known or suspected risk factors for the outcome: HIV status, shared drug equipment (i.e. self-report of sharing or being shared drug equipment (e.g. needles and syringes)) in the last 6 months prior to baseline, the calendar date of the first interview (binary: before or after ARISTOLE program ended), education (primary school, high school, and post-high school), and employment status (employed, unemployed/looking for a job, cannot work because of health reason, etc.) (Table 5). The causal effects were estimated separately using the logistic regression censoring model under the assumption that censoring mechanisms are independent across participants and the mixed-effects censoring model under the assumption that there is a correlation between censoring of participants within a component. The results showed that the participant’s, their neighbors’ exposure to community alerts, and neighbors’ baseline covariates were not statistically significant associated with censoring indicators.
The network structure in TRIP included 356 participants and 542 shared connections. One of the participant was recruited twice as a network member of a recent seed and as a network member of a control seeds with long-term HIV infection. In our analysis, we only used the information for this participant corresponding to their records as a network member of a recent seed. A total of 79 participants were isolates (i.e. not sharing connection with other network members) and removed for our analysis as spillover is not possible for isolates. If these individuals were included, the assumption is made that having no contacts is the same as having contacts with no exposure. In TRIP, these are two different scenarios with different HIV transmission risks, so we exclude isolates from the analysis. In addition, two participants were removed due to missing values on HIV risk behavior in the past 6 months reported at baseline. The final TRIP network had 10 unique components (component sizes were 2, 2, 2, 2, 2, 4, 5, 7, 10, and 239) with 275 participants and 540 shared connections among those individuals after excluding isolates. There were 56 participants (21%) who were lost to follow-up by 6 months. Among the 275 participants in TRIP, 29 participants (11%) received a community alert about an increased risk for HIV infection. The point estimates and corresponding 95% Wald-type CIs use both censoring models under allocation strategies 25%, 50%, and 75%, representing low, moderate and high coverage strategies. The mixed-effects censoring model did not detect random effects in each component on the original TRIP network. Therefore, the results using either censoring model are identical (Table 6).
Descriptive statistics of network characteristics and baseline variables for study participants after excluding isolates in the Transmission Reduction Intervention Project, Athens, Greece, from 2013 to 2015.
Descriptive statistics of network characteristics and baseline variables for study participants after excluding isolates in the Transmission Reduction Intervention Project, Athens, Greece, from 2013 to 2015.
The estimated RD and 95% Wald-type CIs of the effects of community alerts at baseline on HIV risk behavior at 6 months on the original TRIP network (10 components) (left) and estimated using completed cases (
RD: risk difference; CIs: confidence intervals; HIV: human immunodeficiency virus; TRIP: Transmission Reduction Intervention Project.
These results indicate that the risk of HIV behavior was reduced by increasing the proportion of a participant’s neighbors exposed to community alerts, in addition to a participant’s exposure. The estimated direct effect under allocation strategy 75% was
In this article, we extended the neighbor inverse probability weighted estimator 8 to allow for possible censoring of outcomes in network-based studies. The two IPCW estimators were obtained by including a censoring weight derived from two different censoring models, namely logistic regression and mixed-effects model, and an inverse probability weighted estimator that assumes the interference set comprises the first-degree connections for each participants. The additional simulation scenario of varying variances of random effects in the censoring weighted model suggested that the performance in terms of empirical coverage probabilities of both censoring models, with and without correlations of censoring indicators, was comparable. The main difference between the logistic model with fixed effects only and logistic mixed model is the model-based standard error but not the point estimation, while IPCW point estimators and variance estimators only use the information from the point estimation of fixed and random effects of the censoring models. In terms of estimating the average outcomes and their corresponding closed-form variances, both censoring models provided similar results for the simulation scenario described in Section 3. When one anticipates correlation of the censoring indicators within a network component either due to prior knowledge or estimated in the study data, the mixed effects censoring model can provide the information on the correlations between censoring indicators, a measure of the association between the missingness of participant outcomes. Therefore, the selection of the censoring model is based on what information is needed and if studying the association of the censoring mechanism between participants is of substantive interest.
We demonstrated both IPCW estimators to be consistent and asymptotically normal. We also derived a consistent estimator of the asymptotic variance. The simulation study showed that the IPCW estimator performed well in finite samples given a large number (
Using two different assumptions about the correlation of censoring mechanisms, we applied the IPCW estimators to assess the causal effects of community alerts on HIV risk behavior at a 6-month follow-up in TRIP. The TRIP network contained 79 isolates, which could be relevant for both direct and overall effects. However, since we are focusing on evaluating spillover effects, which are not possible for isolates, we excluded these 79 isolates from our analysis. Further, based on the results from additional simulation scenario 3, we observed that ECPs were over-estimated for 10 components of the TRIP network and slightly under-estimated for 24 components. As a result, we focused on the original TRIP 10-component network for our analysis. The mixed effects censoring model did not detect any random effects for each component, resulting in identical outcomes when using both the logistic censoring model and the mixed effects model. This outcome may be due to discordant random effects within the connected components of the network, which likely canceled each other out. The estimated CIs were quite wide when considering the original TRIP network, possibly due to the variability in the sizes of the components. In this study, community alerts were found to be protective, with potential spillover effects from neighbors to unexposed participants. The spillover effects ranged between
There are several interesting directions for future research. One prominent area involves the choice and correct specification of the censoring model. Although the additional simulation scenario varying the variance of random effects in the censoring weight model produced comparable results in our study setting, results may differ in other contexts, for example, in studies involving participants grouped in families, where dependency in censoring indicators among closely connected individuals may need to be modeled to improve variance estimation. 31 Another important direction is to explore alternative censoring mechanisms, such as MNAR. Graphical models known as ”missingness graphs”—causal directed acyclic graphs offer a useful framework for representing MNAR mechanisms and may support the development of adjusted estimators for causal parameters using partially observed data.21,22,32 Additionally, extending the simulation framework to incorporate varying network structures across replicates would allow evaluation of the estimators under structural variation, further aligning with real-world sampling scenarios. Finally, since the consistency of IPCW estimators depends on the correct specification of both the censoring and exposure weight models, assessing model fit and conducting sensitivity analyses is crucial to ensuring the validity of the proposed method. 33 It is worth noting that, due to the large sample properties, we require the network has fairly large number of components. The other direction is related to evaluating the accuracy of the variance estimator when the sample size or number of components is small. 34 Based on the simulation study, the CI coverage levels can be below the nominal level in a network with a small number of network components or above the nominal level when the variability of the size of components is high. Future research could include developing a methodology that has reasonable finite-sample performance when the network has a small number of components or small number of participants. Additionally, the random effect in the same component was assumed homogeneous in this study. However, this may not hold when some of components in the network have large size. Future work could involve a revised M-estimation procedure for the variance to use individually weighted estimators that are consistent when there are varying component sizes, extending results from a two-stage randomized trial to this network setting. 25
Supplemental Material
sj-pdf-1-smm-10.1177_09622802251382586 - Supplemental material for Assessing spillover effects: Handling missing outcomes in network-based studies
Supplemental material, sj-pdf-1-smm-10.1177_09622802251382586 for Assessing spillover effects: Handling missing outcomes in network-based studies by TingFang Lee, Ashley L Buchanan, Natallia Katenka, Laura Forastiere, M Elizabeth Halloran and Georgios Nikolopoulos in Statistical Methods in Medical Research
Footnotes
Acknowledgements
These findings are presented on behalf of the Transmission Reduction Intervention Project (TRIP). We would like to thank all of the TRIP investigators, data management teams, and participants who contributed to this project. The project described was supported by award number DP2DA046856 by the Avenir Award Program for Research on Substance Abuse and HIV/AIDS (DP2) and award number R01DA058994 from the National Institute on Drug Abuse of the National Institutes of Health (NIH), the NIH National Institute of Mental Health award number R01MH134715, the NIH National Institute on Drug Abuse award number DP1DA034989, which funded Preventing HIV Transmission by Recently-Infected Drug Users, the NIH National Institute on Drug Abuse award number P30DA011041, which supported the Center for Drug Use and HIV Research, and the NIH National Institute of Allergy and Infectious Diseases award number R01AI085073, which supported the project Causal Inference in Infectious Disease Prevention Studies. Special thanks to Dr Samuel R Friedman (Department of Population Health, New York University Grossman School of Medicine) for his generous support and comments. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the National Institutes of Health (NIH) through the following awards: the Avenir Award Program for Research on Substance Abuse and HIV/AIDS, National Institute on Drug Abuse [DP2DA046856]; the National Institute on Drug Abuse [R01DA058994, DP1DA034989, P30DA011041]; the National Institute of Mental Health [R01MH134715]; and the National Institute of Allergy and Infectious Diseases [R01AI085073].
Declaration of conflicting interest
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
