Abstract
Given globalization trends in the conduct of clinical trials, the external validity of trial results across geographic regions is questioned. The objective of this study was to examine the efficacy of treatment in acute mania in bipolar disorder across regions and to explain potential differences by differences in patient characteristics. We performed a meta-analysis of individual patient data from 12 registration studies for the indication acute manic episode of bipolar disorder. Patients (n = 3207) were classified into one of three geographic regions: Europe (n = 981), USA (n = 1270), and other regions (n = 956). Primary outcome measures were mean symptom change score on the Young Mania Rating Scale (YMRS) from baseline to endpoint and responder status (50% improvement form baseline). Effect sizes were significantly smaller in the USA (g = 0.203, 95% confidence interval (CI) 0.062–0.344; odds ratio (OR) 1.406, 95% CI 0.998–1.980) than in Europe (g = 0.476, 95% CI 0.200–0.672; OR 2.380, 95% CI 1.682–3.368) or other regions (g = 0.533, 95% CI 0.399–0.667; OR 2.300, 95% CI 1.800–2.941). Regional differences in age, gender, initial severity, body mass index, placebo response, discontinuation rate, and type of compound could not explain the geographic differences in effect. Less severe symptoms at baseline in the US patients did explain some of the difference in responder status between patients in Europe and the USA. These findings suggest that the results of studies involving patients with acute mania cannot be extrapolated across geographic regions. Similar findings have been identified in schizophrenia, contraceptive, and in cardiovascular trials. Therefore, this finding may indicate a more general problem regarding the generalizability of pharmacological trials over geographic regions.
Introduction
The function of regulatory authorities is to protect and promote public health by deciding on the market authorization of pharmaceutical products (Gispen-de Wied and Leufkens, 2013). To this end, they assess, among other things, the quality of clinical trial data and the relevance of the results for their domestic market (FDA, 2007a; Rehnquist, 2003). Until a few years ago, clinical trials were mainly carried out in North America, Western Europe, South Africa, and Australia (Thiers and Sinskey, 2008), but nowadays clinical trials, like the economy, are global.
As discussed by Thiers and Sinskey, the globalization of clinical trials has both advantages and disadvantages. Potential benefits include dissemination of medical knowledge and effective medical practice to other parts of the world and greater access to high-quality medical care for patients worldwide. An important disadvantage is the difficulty of drawing valid scientific conclusions based on pooled data from ethnically and culturally diverse populations (Thiers and Sinskey, 2008).
The globalization of clinical trials means that regulatory bodies, such as the European Medicines Agency (EMA) and the US Food and Drug Administration (FDA), now have to evaluate the relevance of studies from different regions and involving populations that may differ from those in their own region (CHMP, 2008; FDA, 2012). The question arises whether the results of studies from one geographic region can be extrapolated to another region, bearing in mind there might be differences in intrinsic factors (patient related) and extrinsic factors (environment and culture related) (CHMP, 2008; ICH, 1998). In a recent study, Mattila et al. (2014) found differences in the efficacy of atypical antipsychotics for the treatment of acute psychotic episodes in patients with schizophrenia from North America, Europe, and the rest of the world. The question arises whether similar differences exist in the efficacy of treatment for other major psychiatric disorders, e.g. bipolar disorder.
According to the DSM-5, bipolar disorder is characterized by manic, depressive, and mixed episodes, and its diagnosis is based on the occurrence of at least one acute hypomanic episode (American Psychiatric Association, 2014). Bipolar disorder has a lifetime prevalence of approximately 1% (Merikangas et al., 2007) and is accompanied by a high level of social dysfunction, comorbidity, and suicide (American Psychiatric Association, 2014; Rosa et al., 2009; Sanchez-Moreno et al., 2009; Sierra et al., 2005). Drugs used to treat the acute manic episode in bipolar disorder are commonly grouped into (atypical) antipsychotics and (anticonvulsant) mood stabilizers.
Very little is known about differences in the efficacy and effect size of these drugs in patients from different geographic regions. In a recent meta-analysis, Vieta et al. (2011) found significant differences in the baseline characteristics and in the mean change score from baseline to follow-up between patients from the USA, India, and Russia. These findings suggest that there are differences in efficacy between countries. However, regulatory bodies such as the EMA are not country oriented but region oriented and organized, and it is thus essential to obtain evidence about inter-regional variations in effect.
The aim of this study was twofold: to investigate whether there are differences across geographic regions (USA, Europe, other regions) in the efficacy and effect size of medications for the treatment of the acute manic episode of bipolar disorder, and to investigate whether possible regional differences in effect can be explained by regional differences in baseline characteristics, placebo response or compound distributed.
Experimental procedures
Selection of studies
We included all studies (n = 12) submitted to the Dutch Medicines Evaluation Board during an 11 year period as part of market authorization application for the indication acute manic episode of bipolar disorder. All studies were double-blind randomized, placebo-controlled trials involving patients diagnosed with DSM-IV bipolar disorder. Pharmaceutical companies provided raw data in order to enable an individual patient data meta-analysis.
The drugs investigated were antipsychotics and (anticonvulsant) mood stabilizers. Active comparators were included and analyzed as treatment. For reasons of confidentiality, medications are referred to as compounds A to G. We restricted the analyses to treatment groups that were given a proven effective dose of the medication, as indicated in the Summary of Product Characteristics (SmPC) if the drug was registered for an acute manic episode. If the drug was not registered for this indication, expert consensus agreed on what would constitute an effective dose taking into account the doses mentioned in SmPCs for related disorders.
Instruments
The severity of the acute manic episode of bipolar disorder at baseline and at study endpoint was assessed with two interview-based questionnaires. The Young Mania Rating Scale (YMRS) comprises 11 items: seven items are scored on a 0–4 scale and four are scored on a 0–8 scale. Total scores range from 0 (no symptoms) to 60 (severe symptoms) (Young et al., 1978). The Mania Rating Scale from the Schedule for Affective Disorders and Schizophrenia – Change Version (MRS from SADS-C) comprises 11 items: one item is scored on a 0–2 scale and ten items are scored on a 0–5 scale (higher score indicates higher severity). Total scores range from 0 (no symptoms) to 52 (severe symptoms) (Spitzer and Endicott, 1987).
Outcome measures
We used two efficacy outcomes: the standardized difference in mean change score on the YMRS or the MRS from baseline to follow-up and the difference in percentage responders. A patient was considered a responder if his/her score on the YMRS or MRS decreased by 50% or more from baseline to follow-up.
The endpoint was defined as the three-week post-baseline assessment, since this is the time point recommended in the EMA Committee for Proprietary Medicinal Products (CPMP) guideline on the clinical investigation of medicinal products for the treatment and prevention of bipolar disorder (CHMP, 2001). As studies involved patients from more than one region, we categorized the patients (who were from 33 countries) into three “regions:” Europe (n = 17), USA (n = 1), and other regions (n = 15). This classification was based mainly on the country grouping used by the World Health Organization (WHO, 2013). We classified studies that included patients from more than one region into separate sub-studies per region. These sub-studies are referred to as USA, European, and other studies.
Statistical analysis
To answer the first research question, we used a two-step, random effects individual patient data meta-analysis. We used random effect rather than a fixed effect meta-analysis because the included studies had been performed by independently operating companies who examined different medications, tested at different times and in different populations.
In the first step, we calculated the total scores on the respective questionnaires at baseline and week three. If outcome data at week three were missing, we used data for week four and if data for week four was missing, last observation carried forward analysis was used to impute the missing outcome data for week three. In the second step, we performed a meta-analysis on the outcomes of step one. We used Hedges’ g as the effect size for the continuous outcome (difference in mean total score from baseline to follow-up) and the odds ratio (OR) as the effect measure for the dichotomous outcome (responder). The interpretation of Hedges’ g is as follows: 0.20–0.30 a “small effect,” around 0.50 a “moderate effect,2 and >0.80 a “large effect” (Cohen, 1988). In addition to the effect size and 95% confidence intervals (CIs), we calculated the 95% prediction interval (PI) for Hedges’ g and the OR (Borenstein et al., 2009). In a random effects meta-analysis, the 95% PI indicates the upper and lower bound of the effect that may be expected when a new study is performed.
To answer the second research question, we used individual patient data. To assess the effect of potential explanatory variables such as age, gender, body mass index (BMI), ethnicity, initial severity, study year, discontinuation rate, and placebo response on outcome, we used a mixed effects linear regression analysis with a random intercept for study and mixed effect logistic regression analysis with a random intercept for study.
The analyses of step one in the two-step meta-analysis were performed with SPSS version 20 (SPSS20) and those of step two were performed with Comprehensive Meta-Analysis, version two (CMA2). All mixed effects regression analyses for the second research question were performed with the xtmixed and xtmemixed programs of STATA 12.
Results
Of the 12 studies, four were performed in only one region (three in the USA and one in the other regions), and of the seven medications, one was investigated in only one region. Subdividing the 12 studies per medication and region resulted in 36 sub-studies, of which 13 were performed in Europe (Belgium, Bulgaria, Croatia, Czech Republic, Estonia, France, Germany, Greece, Hungary, Latvia, Lithuania, Poland, Romania, Russia, Turkey, Ukraine, United Kingdom), 9 in the USA, and 15 in the other regions (Argentina, Australia, Chile, China, Hong Kong, India, Indonesia, Korea, Malaysia, New Zealand, Philippines, Singapore, South Africa, Taiwan, Tunisia). The total number of patients in the included studies was 3207: 981 in Europe, 1270 in the USA, and 956 in the other regions. Table 1 presents the patient characteristics per sub-study (stratified by study, medication type, and region). Table 2 presents the overall patient characteristics stratified by region. Results for the two-stage meta-analysis are presented in the Forest plots of Figures 1 and 2.
Patient characteristics per sub-study.
= Compound; 2 = MRS questionnaire (rest = YMRS questionnaire).
Patient characteristics per region.
= Compound.

Flow chart of mean change score (g) by region: meta-analysis.

Flow chart of responder status (OR) by region: meta-analysis.
Effect in terms of mean change scores
The combined Hedges’ g (95% CI) for all studies was 0.396 (0.306–0.486, tau: 0.234), resulting in a 95% PI of −0.089–0.881. Hedges’ g (95% CI) differed across the regions: Europe 0.476 (0.253–0.699), USA 0.203 (0.033–0.373), and other regions 0.533 (0.385–0.681). These differences were statistically significant (p < 0.01). This can be interpreted as a small effect in the USA and a moderate effect in Europe and other regions (Figure 1).
Effect in terms of responder status
The combined OR (95% CI) for all studies was 2.048 (1.721–2.432, tau: 0.133), resulting in a 95% PI of 0.954–4.390. The ORs (95% CI) differed across the regions: Europe 2.380 (1.682–3.368), USA 1.406 (0.998–1.980), and other regions 2.300 (1.800–2.941). These differences were statistically significant (p < 0.05) (Figure 2).
Differences in baseline characteristics across regions
There were significant differences at baseline in age (p < 0.001), BMI (p < 0.01), initial severity (p < 0.01), gender (p < 0.001), and ethnicity (p < 0.001) between the US and European studies. Only severity measured with the MRS was not significantly different (n = 433, p = 0.673). All variables were significantly different in the US and other region studies (p < 0.05).
Patients in the USA were younger (39.68) than patients in Europe (42.12) and older than the patients in the other regions (35.59). They had a higher BMI (29.22) than patients in Europe (25.90) and the other regions (22.86) and were less severely ill at baseline according to the YMRS and MRS scores (28.24, 25.35) than patients in Europe (29.54, 27.90) and the other regions (33.62, 30.27). The US studies included more men (0.56) than the European studies (0.46) and fewer than the other region studies (0.60). Ethnicity was strongly related to region: the proportion of Caucasians, African Americans, and Asians was 99.2%, 0.3%, and 0.4%, respectively, in Europe; 66.8%, 28.2%, and 0.6%, respectively, in the USA; and 8.2%, 3.1%, 43.2%, respectively, in the other regions (Table 3).
Differences in baseline characteristics and placebo response across geographic regions (p-value).
N = 2764; 2 N = 443.
Differences in placebo response across regions
Patients in the other regions had a higher placebo response than patients in the USA in terms of both mean change score (−10.132 vs. −7.561; p = 0.017) and responder status (0.423 vs. 1.962; p = 0.021). There was no significant difference in the placebo response between patients in Europe and the USA (mean change score −7.161 vs. −7.561, respectively; p = 0.702) or in responder status (0.392 vs. 0.423, respectively; p = 0.862).
Differences in study year across regions
All studies were performed between 1996 and 2007. Studies including US patients were performed in the period 1996–2007 and studies including patients from Europe and other regions were performed in the period 1998–2007.
Differences in discontinuation rates across regions
There was a significant difference (p < 0.001) in discontinuation rates between patients in the US and patients in Europe and patients in the other regions, showing a higher discontinuation rate in patients in the USA (40%) compared to patients in Europe (21%) or the other regions (21%).
Potential explanatory variables
Mean change score
After adjusting for the potential explanatory variables age, gender, initial severity, BMI, placebo response, study year, and discontinuation rate separately, we still found significant differences across regions (p < 0.001). After simultaneous adjustment for the combination of the variables age, gender, severity, placebo response, study year, and discontinuation rate, the significant differences across geographic regions remained (p < 0.001). Since ethnicity was strongly confounded with region, we were only able to assess the potential confounding effect of African Americans on YMRS data for 1097 US patients in a post-hoc analysis (Table 4). Adjustment for this potential confounder resulted in an unadjusted effect estimate −2.531 (SE: 0.597, p < 0.001) versus an unadjusted effect estimate of −2.521 (SE: 0.597, p < 0.001), showing that the proportion African Americans did not influence the mean change score in patients in the USA. This indicated that differences in ethnic composition between the regions could not explain the differences in mean change score between the regions.
IPD analysis effects of potential confounders on mean change score (g) by region in YMRS questionnaire (SE).
Analyses only for studies using YMRS questionnaire. N = 2764; 2BMI could only be analyzed in N = 2186; 3Combined adjustment for all variables except for BMI.
Mixed model linear regression analyses with random intercept for study and heterogenetic level one variance. Dependent variable = mean change score. Independent variables are Region and Treatment (control group is reference category). Region is a categorical variable (USA = 0, Europe = 1, Other = 2) represented in the model by two dummy variables, Region(1) contrasting Europe vs. USA and Region(2) contrasting Other vs. USA. In each column the results are presented for an analysis with the variable mentioned in the column header added as covariate.
Likelihood ratio test model with vs. model without treatment by region interaction χ2(2) = 6.39, p-value = 0.041.
LRT = likelihood ratio test comparing model with and without the treatment by region interaction terms.
Responders
Age, gender, BMI, placebo response, and discontinuation rate did not explain the observed geographic differences in effect size. However, after adjustment for initial severity the difference between responders from Europe and the USA was no longer significant (p = 0.082), but the difference between responders from the other regions and the USA remained significant (p = 0.037). Nevertheless, the same tendency for a higher proportion of responders in European studies compared with US studies remained. After simultaneous adjustment for age, gender, severity, placebo response, study year, and discontiunation rate, the significant differences across regions remained (p < 0.01) (Table 5). Since ethnicity was strongly confounded with region (see above), we assessed the potential confounding effect of African Americans in the US studies (n = 1270 patients) in a post-hoc analysis. This resulted in an adjusted OR of 1.474 (95% CI: 1.159–1.875) versus an unadjusted OR of 1.460 (95% CI: 1.149–1.857), indicating that differences in ethnic composition between the regions could not explain the differences in the proportion of responders between the regions.
IPD analysis effects of potential confounders on effect modification response (OR) by region in YMRS questionnaire (p-value).
Analyses based on all mania studies n=3207; 2analyses only based on the YMRS questionnaire studies N = 2764; 3BMI could only be analyzed in N = 2186; 4Combined adjustment for all variables except for BMI.
Mixed model logistic regression analyses with random intercept for study. Dependent variable = Responder. Independent variables are Region and Treatment (control group is reference category). Region is a categorical variable (USA = 0, Europe = 1, Other = 2) represented in the model by two dummy variables, Region(1) contrasting Europe vs. USA and Region(2) contrasting Other vs. USA. In each column the results are presented for an analysis with the variable mentioned in the column header added as covariate.
Likelihood ratio test model with vs. model without treatment by region interaction χ2(2) = 6.39, p-value =0.041
LRT = likelihood ratio test comparing model with and without the treatment by region interaction terms; bthese analyses only used studies that used the YMRS subjects (we needed an initial severity score).
Subgroup antipsychotics
To analyze whether geographic differences could be caused by a different distribution of compound type to the three regions, we analyzed the main group of compounds, the antipsychotics. This led to an overall combined Hedges’ g (SE) of 0.482 (0.047). Hedges’ g differed significantly (p < 0.01) across the regions: Europe 0.543 (0.101), USA 0.248 (0.081), and other regions 0.627 (0.070). However, in this subgroup analysis, the difference in responder status (OR, 95% CI) across regions was no longer significant (p = 0.134), although the size and direction of the differences was very similar to that found in the main analyses: USA 1.564 (0.915–2.678), Europe 2.465 (1.518–4.002), and other regions 2.562 (1.830–3.588).
Discussion
We investigated whether results from efficacy studies of drugs for the treatment of acute mania can be extrapolated across geographic regions, and whether potential differences in efficacy can be ascribed to regional differences in baseline characteristics, placebo response, discontinuation rate, and compound distributed. We found significant differences in drug effect size across regions, with European studies (g = 0.476, OR = 2.380) and other region studies (g = 0.533, OR = 2.300) consistently reporting larger effects than US studies (g = 0.203, OR = 1.406). Although we found significant regional differences in baseline characteristics (age, gender, ethnicity, BMI, and initial severity), and significant differences in placebo response and discontinuation rates between patients in the USA and other regions, these differences did not explain the regional differences in mean change score from baseline to follow-up. However, differences in responder status in the US and European studies could be partly explained by the lower initial severity of manic episodes in the US patients compared with that in the European patients, but this was not the case for differences in responder status between the US and Other region studies.
Our results are consistent with the findings of Mattila et al. (2014), who showed a tendency to regional differences in the efficacy of atypical antipsychotics in the acute treatment of patients with schizophrenia in North America versus Europe and the rest of the world. They found a smaller effect size in US studies (Mattila et al., 2014). Furthermore, our findings are consistent with those of Vieta et al. (2011), who also found that patients treated for acute mania in the USA were younger, had a higher BMI, and had less severe symptoms than patients in Russia and India, and that, as in our study, none of the baseline characteristics was a significant predictor of change in MRS score. Unlike Vieta et al. (2011) but in line with other reports, we did not find a higher placebo response in the US studies (Watsky et al., 2009).
Geographic differences in drug efficacy have been reported for other health issues/conditions, including cardiovascular disease (Blair et al., 2008; Mentz et al., 2012; Stough et al., 2007) and contraception in contraceptive trials (Grubb et al., 2008; Lete et al., 2014). Despite the efforts of Blair et al. (2008) in the EVEREST study to select a fairly homogenous population of patients with heart failure, important differences in etiology, severity, management, and outcomes were found (Blair et al., 2008). A review by Mentz et al. (2012) also found differences in outcome and baseline characteristics of patients with different types of cardiovascular disorders (heart failure, acute coronary syndromes, hypertension, and atrial fibrillation) across geographic regions (Stough et al., 2007). Similarly, efficacy of contraceptives frequently showed lower efficacy in clinical trials conducted in US patients compared to trials conducted in Europe (Grubb et al., 2008; Lete et al., 2014), an issue that was addressed by the FDA (FDA, 2007b).
The current study had two limitations. First, the p-value of responder status of the US studies was only marginally significant (p = 0.051), unlike that of the European and other region studies (p < 0.001). In order to determine whether this was caused by a difference in drug distributed, we separately analyzed the main group of compounds, the antipsychotics. This led to an even larger difference in the magnitude of the mean change score between the US, European, and other region studies. However, the difference in effect size of responder status across regions was no longer significant, although the differences were similar. The second limitation is the limited ability to explain regional differences, because of the limited availability of information about baseline variables in our database. Possible potential explanatory factors not investigated in the current study include genetic differences (intrinsic factor) and differences in available treatment/medical services, differences in the patient recruitment, benefits of participation, differences in interpretation of disease-specific symptoms, duration and onset of illness (Swann et al., 2001; Vieta et al., 2011; Welge et al., 2004; Zarate et al., 1998), and history of prior (pharmacological) treatments/medication and hospitalization (Swann et al., 2001; Vieta et al., 2011; Welge et al., 2004). Vieta et al. (2011) did not find the last four variables to be significantly associated with treatment outcome.
To our knowledge, this is the first study to evaluate geographic differences in the efficacy of drugs for the treatment of acute mania in patients with bipolar disorder, based on individual patient data meta-analysis and meta-regression-analysis. The effects of treatment were significantly smaller in the USA (small effect) than in Europe and other regions (moderate effect), and these differences could not be explained by differences in baseline characteristics, study year, differences in placebo response or compound distributed across geographic regions. These findings may affect the requirements of regulatory authorities with regard to regional extrapolation of study results. The development of new medicines, and not only for psychiatric patients, has become a global activity. This involves many different continents, ethnic populations, or health systems. The interpretation of study results by regulators, payers and clinical practice, therefore implies complex questions regarding the external validity of the findings and the consistency of study results across subsets. These questions will remain on the table, because research, including our study, pertinently reveals that geography matters when it comes to the interpretation of clinical data from studies in different psychiatric settings. With regard to geographic differences in efficacy for drug treatment of patients with acute mania, we found that geographic differences could only partly be explained. But we also saw that albeit differences, overall data showed efficacy of drug treatment in acute manic patients both in US and European/other patients. This would therefore probably result in a positive benefit/risk balance in patients of all three regions. However, in drugs for a less severe indication, for chronic use, or with smaller effect sizes and more severe adverse events, this judgment could of course change. Therefore, whether the FDA can and will accept European data or whether European regulators will do vice versa, remains a question of judgment and regulatory policies and future research could help in this judgment. In regulatory terms, we can only aim for a prescription drug label that is rooted in the best science available, which is not the same as the best thinkable science. Our study shows that there is scientific room for extrapolation, not always, and sometimes with obvious restrictions.
Footnotes
Acknowledgements
We thank Dr Armand Voorschuur at Nefarma, the association for innovative medicines in The Netherlands, for his facilitation of the current collaboration. We also thank the pharmaceutical companies AstraZeneca, Eli Lilly, Glaxo Wellcome Inc., and Johnson & Johnson for offering their raw individual patient data for study.
Declaration of Conflicting Interests
The authors declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: W van den Brink has received honoraria from Lundbeck, Merck Serono, Schering-Plough, Reckitt Benckiser, Pfizer, and Eli Lilly; speaker’s fees from Lundbeck; investigator-initiated industry grants from Alkermes, Neurotech, and Eli Lilly; is a consultant to Lundbeck, Merck Serono, Schering-Plough, and Teva; and has performed paid expert testimony for Schering-Plough.
HGM Leufkens has received unrestricted research funding from the Netherlands Organization for Health Research and Development (ZonMw), the private-public funded Top Institute Pharma (
, includes co-funding from universities, government and industry) and Innovative Medicines Initiative, the EU 7th Framework Program (FP7), the Dutch Medicines Evaluation Board, the Dutch National Health Care Institute, and the Dutch Ministry of Health.
D Denys is a member of the advisory board of Lundbeck. He receives occasional consultant fees from Medtronic for educational purposes.
CCM Welten, MWJ Koeter, T Wohlfarth, JG Storosum, and CC Gispen-de Wied declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
