Abstract
A randomized trial comparing survival in hemodialysis and peritoneal dialysis remains a utopian aspiration. Dialysis is still relatively rare on a population basis, and a natural tension exists between desirability and feasibility in terms of quality of evidence. In practice, it is very difficult to perform prospective comparisons with large groups of contemporary representative subjects, and much of the literature comes from retrospective national registries. This article considers several questions to address when trying to compare the outcomes of peritoneal dialysis and hemodialysis.
Prognostic similarity at baseline is a fundamental issue. Traditionally, adjustment for known prognostic factors has been used in an attempt to minimize the bias caused by nonrandom treatment assignment. Propensity scores have been suggested to be superior, and matched-case analysis may also be a useful method for comparison. Other questions include, when, in relation to starting dialysis, to start the observation clock; the definition and handling of switches of dialysis therapy; and the decision to censor at transplantation.
Finally, comparisons are complicated by hazards ratios that vary over time, and time-segmented analysis is obligatory. Many types of analytical approaches are needed to begin to appreciate outcome disparities between dialysis therapies.
The decision to choose one form of dialysis should weigh cultural and lifestyle factors. Very few studies have compared the impact of dialysis therapies on these outcomes, which are often described as “soft” to distinguish them from “hard” outcomes, such as death. Quality-of-life studies, which are also frequently dismissed as nonscience or soft-science, consistently show that living with end-stage renal disease is truly “hard.” Relatively few rigorous quality-of-life comparisons have been performed. A Medline survey performed on 11 July 2003, limited to humans, produced 683 citations with the search terms “comparison,” “hemodialysis,” “peritoneal dialysis,” “(survival or mortality).” Substituting “quality of life” for “(survival or mortality)” yielded 27 citations. At face value, it seems that medical researchers consider quantity to be 683/27 or 25.3 times more important that quality of life. This review is guilty of examining quantity of life in isolation.
A randomized trial comparing survival with both modes of therapy remains a utopian aspiration. The logistic and ethical barriers remain formidable. Using the search terms “randomized,” “comparison,” “hemodialysis,” “peritoneal dialysis,” “(survival or mortality)” on Medline yields a single study, which compares hemofiltration and PD in infection-related acute renal failure (2). Observational studies — always a poor substitute for trials — remain the most feasible tool for comparison. As with all studies, data acquisition that is planned to occur at regular intervals is desirable. It is important to characterize the patients comprehensively, especially initially. A large sample size is useful because associations can be generalized back to a greater segment of the target population, and because numerical estimates of association can be made more precise. Usually these ideals can only be met with a prospective inception-cohort design and large numbers of centers. Prospective comparisons of HD and PD are relatively rare in practice. They usually include small numbers of centers and several years of follow-up to generate sufficient statistical power for meaningful comparison. Hemodialysis and PD change palpably from year to year, which means that generalizability of single-center studies of long duration must always be questioned. Retrospective mining of large registries, typically at the national level, can help address the issue of generalizability to the present, but the quality of initial subject assessment is always open to question. Thus, even without mentioning the many difficulties unique to HD/PD comparisons, it becomes quickly apparent that ideal studies may be impossible, “good” studies may not be so good, and “bad” studies may not be so bad. In practice, “might” often defeats “right,” and large-scale retrospective studies remain the most practical approach to comparing outcomes.
Quite apart from the underlying study design, observational comparisons of PD and HD involve several challenges. For example, HD is much more likely to be used as initial therapy when the onset of end-stage renal disease is rapid. Peritoneal dialysis patients often use HD as a temporary rescue treatment and are much more likely to switch permanently to HD than vice versa. Several of these factors tend to bias toward the null when initial treatment modes are compared using a philosophy of intention-to-treat. In addition, non-constant mortality PD-to-HD hazards ratios have also been seen in most studies, usually favoring PD initially, and HD later on (3,4). The analytical problems divide into two broad categories: the initial similarity of the groups compared and the analytical approach.
Prognostic Similarity at Dialysis Inception
Prognostic comparability at the beginning of dialysis therapy is a fundamental issue. Random assignment to HD or PD is the only valid method that naturally tends to generate groups of patients that are similar with respect to both known and unknown factors. Observational comparisons can never address the issue of unmeasured prognostic factors. At the very least, they must attempt to account for differences in known, or suspected, prognostic factors.
Usually, suspected prognostic factors are incorporated along with mode of dialysis therapy in multivariate regression models that are applied to the combined population of PD and HD patients. The proportional hazards model in which survival time is the outcome studied is the most frequent statistical technique, followed by Poisson regression, where death rates per unit of time are the dependent variable. Traditionally, each adjustment variable is used in the model as a single entity. A difference in therapy outcome that persists after adjusting for ever more prognostic factors is more convincing evidence implicating the therapy itself. Accommodating more and more adjustment variables is necessarily a drain on statistical power.
Even with large data sets, PD and HD may be very different in terms of actual covariate patterns, and traditional adjustment approaches may not eliminate bias. Other approaches have been used to lower the likelihood of nonrandom treatment assignment, including propensity scores (5). The propensity score for a given patient is defined as the probability of receiving one of the two treatments based on that patient's baseline characteristics. A typical procedure would use multivariate logistic regression, with “assignment to PD: Yes (1) or No (0)?” as the dependent variable. Regression coefficients from this model are then used to calculate the probability of using PD for each patient, based on patient-level covariate data. At their most basic, propensity scores are an economical way to condense several variables into one variable without unduly sacrificing discriminatory power. Typically, the propensity score is then used along with the actual treatment assigned in analyses that use regression adjustment, stratification, or matching.
All the multivariate techniques discussed so far tacitly assume that covariate distributions are similar in PD and HD. It is likely that there are patients whose comorbidity profiles essentially rule out using one form of therapy. Matched-pair analysis is another potentially useful method to compare outcomes when extreme distribution dissimilarity is thought to be a problem. In the United States, for example, HD is used over 10 times more frequently than PD. It is possible to match a given PD patient with a similar HD patient, followed by outcome comparisons. The sample size and computing requirements needed for this approach are highly dependent on the number and type of matching variables chosen. Comparing outcomes before and after matching is a logical approach to estimate the impact of nonoverlapping comorbidity profiles.
When does the Observation Clock Start Ticking?
Hemodialysis is virtually always used as first therapy when renal failure develops acutely. Acutely declining renal function often develops in the face of other life-threatening organ dysfunction. Retrospective registry studies, which are often based on administrative claims data, are unlikely to be able to completely quantify these effects. Studies that begin follow-up at first dialysis may not be a valid comparison. Many studies use 3 months as a starting point to get around this problem. This may seem reasonable, but is arbitrary. Clearly, there is no biological reason to believe measurable therapy-related adverse outcomes could not happen before 3 months. In practice, the safest approach may be to test the sensitivity of outcomes to variable study start times.
Most analyses comparing dialysis therapies are highly artificial in the sense of assuming that choice exists for patients whose first presentation could be as late as the first dialysis session. There is a vast literature suggesting that patients who present late with end-stage renal disease have poor outcomes (6). If a randomized trial was being envisaged, it is likely that the only settings where a true choice could be made would be in the setting of known, slowly progressive kidney disease managed in a nephrology clinic. Observational studies of patients in multidisciplinary chronic kidney disease clinics, followed from two time points, could provide very useful reallife comparisons: after information on dialysis type has been assimilated and a preference indicated, and at the time of first dialysis.
Dealing with Switches of Dialysis Therapy
Intention-to-treat is a safe approach to comparing therapies. It espouses the philosophy that outcomes are compared according to initially assigned therapy, irrespective of time spent on that treatment. Applied strictly to HD and PD mortality, only the date of death (or final follow-up) is needed in outcome comparison. In as-treated analysis, outcomes are assigned to the actual treatment in use when the event takes place. A number of variants of as-treated analysis have been used. One counts follow-up and events in each discrete period spent on that treatment. For example, a patient starting with PD who switches annually until death at the end of the fourth year contributes 2 years of follow-up and no events to the PD tally, and 2 years of follow-up and one event to the HD tally. Another common variant stops follow-up at the first treatment switch. In our hypothetical patient, 1 year of PD with no events is added to the PD experience and nothing is added to the HD experience. These forms of as-treated analysis are clearly very different. The first involves much more computation and cannot be readily analyzed with standard survival methods, relying instead on Poisson regression techniques applied to rates, or time-dependent Cox regression techniques applied to survival times. “As sequentially treated” and “as first treated” might be more accurate descriptions of the two approaches. Hybrids exist that are intermediate between intention-to-treat and as-treated. One approach, which has been termed “integrated care,” compares outcomes in the four groups formed by dichotomizing the parameters “initial mode of therapy” and “switch of therapy before a specified time” (7).
It remains a fact that patients initially assigned to PD are much more likely to switch to HD than vice versa. In Canada for example, technique failure rates of 154.0 per 1000 patient-years have been reported, suggesting that approximately half of all PD patients will switch to HD over their expected survival time (8). Assuming that switching is a neutral event, restricting oneself to the question, Which initial therapy shows better survival rates? limits the possibility of finding out which is actually a better therapy. Playing out simple scenarios can be instructive in this regard. Two very different therapies can be chosen, A and B. Most A patients switch to B, whereas B never switches to A. Risk estimates are ratios between B and A, with A as reference category:
scenario I: staying on A causes higher mortality than staying on B;
scenario II: staying on A causes lower mortality than staying on B.
With scenario I, intention-to-treat mortality risk ratios must be less than as-first-treated mortality risk ratios, and the disparity must widen as the switch rate increases. The exact opposite must apply with scenario II. When comparing PD and HD, intention-to-treat mortality comparisons alone are unrealistic. To get an accurate picture, switch rates should be known and both intention-to-treat and as-treated analyses presented. When switch rates are high, mortality differences in an intention-to-treat analysis are especially noteworthy.
Defining a Switch of Dialysis Therapy
Most of the previous section deals with the pluses and minuses of different strategies for dealing with switches of dialysis therapy. The definition of a switch can lead to several analytical challenges. Several key questions come to mind: Do I define a switch as occurring instantaneously, or do I wait for a specified period of time? Is the reason for the switch a prognostic indicator too? If an event happens soon after a switch, is this tallied against the old or the new therapy? One approach has been to define switches of therapy as occurring when a predefined interval of time has elapsed, with events attributed to the old therapy until that time. This approach is somewhat arbitrary. One obvious problem is that rapidly lethal events could as easily cause a switch of therapy as vice versa. There is no simple solution to this dilemma. One practical approach is to test the sensitivity of the mortality associations to different time intervals for switch definition, both before and afterward.
It is useful to check the prognostic associations of switching. There are many ways to do this. One method compares the outcomes of the four groups defined by combining mode of dialysis and occurrence of a switch (9). Another useful approach is to consider a switch as a conditional time-dependent covariate in an analysis, with baseline mode of dialysis as a fixed covariate.
Censoring at Transplantation
A pure intention-to-treat analysis ignores transplantation; transplantation is associated with much higher survival rates than dialysis therapy, and imbalanced transplantation rates could have profound effects on dialysis mortality comparisons (10). A recent analysis from the United States showed higher transplantation rates in PD patients, slightly offset by higher rates of early graft dysfunction (11). In contrast, a single-center European study showed a lower incidence and the severity of delayed recovery of renal function after renal transplantation with PD (12). Because of its impact on survival, transplantation should be considered when comparing dialysis therapies. This is typically addressed with analyses that either ignore transplantation or use it as a censoring event. The disparity between the two types of estimate allows an indirect estimate of transplantation inequities on outcome. Unfortunately, advocating both methods adds another multiplier to the number of analyses needed to make a “big picture” comparison of HD and PD.
Non-Proportional Mortality Hazards
Relatively speaking, mortality hazards ratios of PD versus HD do not remain constant over time. Typically, an initial tendency for mortality comparisons to be more favorable toward PD is reversed after approximately 2 years (4,9). The underlying basis for these trends is unknown. Less severe comorbidity at baseline and slower loss of residual renal function have been postulated as underlying mechanisms. Most studies suggest that temporal non-proportionality reaches clinically important levels, which leads to several analytical challenges. Mortality comparisons based on prevalent samples become difficult to interpret. Simple adjustment for vintage in prevalent outcome comparisons may be somewhat simplistic, as this fails to account for how timing of end-stage renal disease, transplantation, death, and switches of dialysis therapy led to the prevalent mix of PD and HD.
Non-proportional hazards lead to other implications. Outcomes should be compared in discrete intervals of time, which places demands on statistical power. It also adds yet another multiplier to the number of analyses that are needed to get a reasonable picture of the relative mortality of PD and HD. Nonproportional hazards also have clinical implications, suggesting that there may be an ideal sequence of dialysis therapies that optimizes outcome.
Analysis in discrete time intervals can lead to pitfalls. For example, imagine 10 discrete time intervals are used and two therapies compared. If the hazards ratio for B relative to A is 0.80 for the first interval, followed by 1.20 in each of the subsequent 9 intervals, one might be tempted to conclude A is better than B. If, for arguments purposes, the vast majority of patients are censored in the first interval, a given patient at time zero might have a better overall outcome with B. Also, absolute mortality risks may change with time: in dialysis populations, mortality rates are much higher in the first 6 months of treatment than vice versa. Thus, both unsegmented analysis and analyses in discrete time frames should be presented to get an overall picture of the comparative merits of the therapies.
Subgroups and Interactions
Choice of subgroups and interactions is a difficult area. With multiple levels of multiple variables, there is obvious potential for both indecision and abuse. In general, the choices should be based on some combination of biological rationale and historical precedence in the literature, which is akin to prespecification of a limited number of subgroup comparisons in randomized trials. Most recent registry studies have tended to define subgroups on the basis of age, gender, or diabetic status, and, more recently, the presence or absence of coronary artery disease (13). Even when very large sample sizes are available, multiple subgroup analyses can create difficulties, and even misconceptions. For example, four subgroups are created when two binary variables are combined (as was done for the variable diabetes and coronary disease in the aforementioned study). If for example, the findings suggest excess mortality for one form of dialysis in three of the four subgroups, this might be interpreted as even stronger evidence than finding a mortality difference in one or two subgroups. Similarly, the presence of one very large risk difference in one subgroup analysis can be misinterpreted. Two questions need to be addressed: What are the relative sizes of the subgroups? and What is the correlation between the factors defining the subgroups? Thus in the recent paper by Ganesh and colleagues (13), three of the four subgroups defined by the variables diabetes and coronary artery disease were associated with higher mortality rates in PD patients. This “3 of 4” result might overshadow the fact that no difference was found in the largest cell, non-diabetes and non-coronary artery disease, which accounted for almost half of all patients. A skeptic would be quick to point out that diabetes and coronary artery disease are strongly correlated in dialysis populations. When similar analyses were performed in two subgroups, those with and those without coronary artery disease (with adjustment for diabetes mellitus), use of PD patients among the 25% of patients with coronary artery disease was clearly associated with worse outcomes; statistically marginal, and clinically very small differences were found in the 75% of patients without coronary artery disease. An analysis including all patients, adjusting for diabetes and coronary disease, favored PD in the first 6 months and HD subsequently; no corresponding unsegmented (with respect to time) analysis was presented, making it difficult to determine whether early benefit was outweighed by later risk.
Analytical edifices built on subgroups have shaky foundations; the most credible findings come from analyses that include all subjects, with adjustment for the factors later used to define subgroups. This type of analysis should be seen as a prerequisite before subgroup analyses can be entertained.
Conclusion
Reasonable studies should include the following: the best possible attempt to minimize nonrandom treatment assignment at baseline, intention-to-treat and as-treated analysis, analyses with and without censoring at transplant, sensitivity analysis with respect to timing of study start, sensitivity analysis with respect to the time intervals used to define a switch of dialysis therapy, and analysis in discrete intervals of time. Comparing outcomes of peritoneal dialysis and hemodialysis is challenging.
