Abstract
Using data scraping techniques to gather data from a variety of previously disjointed sources—some proprietary and some publicly available—this research applies the analytical techniques of data visualization and machine learning to (1) gain exploratory insights into the drivers of prescription drug list prices and (2) test how well these variables impact prices directly and interact to predict pricing. Specifically, this inductive analysis considers characteristics related to the brand (i.e., manufacturer, brand/generic classification), product attributes (i.e., dosing levels, amount of active ingredient), the condition for which the drug is recommended (i.e., therapeutic class, subclass, and pricing tier), and market factors (i.e., number of drugs in class and approval year). Through these analytic analyses, the authors seek to cut through some of the opacity of pharmaceutical drug list prices to consider the drivers of drug prices, evaluate how these insights might drive marketplace and policy solutions, and spark future research inquiries in this area.
Keywords
Prescription drug prices in the United States continue to spur conversation, with consumers and policy makers voicing concerns about prices paid, as well as the steep price increases observed over time. Growth in drug prices has outpaced the cost of living in the United States: in the first six months of 2019, the average price for 3,400 pharmaceutical drugs increased 10.5%—five times the rate of inflation (Picchi 2019). With approximately 60% of Americans on at least one prescription medication, the price of prescription drugs is of utmost importance for the physical and financial well-being of consumers (Kaiser Family Foundation 2019). Understanding the drivers of pharmaceutical prices is critical to informing discussions related to ethical concerns surrounding escalating prices and crafting marketplace and policy solutions that address current price levels and growth over time.
Though the U.S. government does not regulate pharmaceutical pricing, it has responded to public concerns regarding escalating prices by enacting initiatives aimed at investigating prices and proposing solutions aimed at curbing growth. In 2018, the Department of Health and Human Services released the American Patients First Blueprint, which identifies challenges in the American drug market. In addition, both the U.S. House of Representatives and the U.S. Senate have drafted bills aimed to address issues related to pharmaceutical prices (Department of Health and Human Services 2018a; Lovelace 2019). One key proposal emerging from these policy conversations centers on promoting price transparency for pharmaceuticals. On the one hand, advocates for greater transparency argue that it provides payers (e.g., hospitals or insurance companies) with a stronger basis for negotiating price and will hold manufacturers accountable for prices charged—potentially resulting in lower prices. On the other hand, some experts argue that price transparency may yield higher prices, due perhaps in part to the additional administrative costs associated with data collection, reporting, and communicating this information (Coukell and Shih 2016). Though the effect of increased transparency on price is unclear, calls for increased disclosure (e.g., in direct-to-consumer television ads and in state filings specifying price and price increases; Department of Health and Human Services 2018a) spotlight the common belief that pharmaceutical pricing practices are hidden beneath a shroud of secrecy known only to the manufacturer.
Drug list prices, which form the basis for prices paid throughout the pharmaceutical supply chain, are set by the manufacturer. Though the tenets of pricing theory would hold that these prices cover costs (e.g., research and development, production) and also account for manufacturer profit, market factors are also at play, and a clear explanation of the elements that contribute to price is neither provided by manufacturers nor required by policy (American Medical Association 2019). Most Americans support the publication of manufacturers’ list prices, thought there is little research that explores the impact of drug transparency on payers in the pharmaceutical supply chain, including intermediaries and patients (Kirzinger et al. 2019). Moreover, given the disjointed nature of data on pharmaceutical drugs and drug pricing, there are few inquiries in the literature that explore the manner in which multiple variables directly and interactively impact pharmaceutical prices.
In this inquiry, we use an analytics-based approach to explore the key inputs to pharmaceutical pricing and aim to answer the following guiding research question: Can analytical techniques uncover deeper insights into the drivers of drug prices? Using data scraping techniques to gather a rich data set from a variety of previously disjointed sources—some proprietary and some publicly available—our work applies data visualization and machine learning techniques to (1) gain exploratory insights into the drivers of prescription drug list prices and (2) test the extent to which these variables impact prices directly and interact to predict drug prices. Specifically, we explore the data to glean insights as to how drug prices vary according to characteristics related to the brand (i.e., brand name, brand/generic classification), product attributes (i.e., dosing levels, amount of active ingredient), the condition for which the drug is recommended (i.e., therapeutic class, subclass, and pricing tier), and market factors (i.e., number of drugs in class and approval year).
Our inductive, diagnostic approach is appropriate for this inquiry, as much of the research that explores the factors underlying pharmaceutical list prices reveals findings that are inconsistent at face level (e.g., the entry of generic drugs can result in branded drugs increasing or decreasing in price), and most existing papers test the role of one key variable as a driver of price. While some work has considered a broader set of factors (e.g., Iacocca, Sawhill, and Zhao 2013), we know of no inquiry that uses analytical methods to explore these relationships. The use of analytical techniques allows us to adopt an inductive approach to explore how the aforementioned factors influence price, which we see as a key contribution of this work. Rather than testing formal hypotheses, we follow the tenets of analytics by allowing insights to emerge from patterns observed in the data (via visualization) and then testing the relationships between variables using predictive techniques (via machine learning).
Our inquiry seeks to cut through some of the opacity in drug pricing by considering which relationships hold true across the data and which do not. In presenting our findings, we seek to spark conversations on pricing and price transparency. Although disclosures may increase knowledge regarding pharmaceutical prices, the more precise drivers of price may not be revealed, and an understanding of whether prices are in line with the public’s—or, for that matter, academic researchers’—understanding of the marketing-based drivers of price is an important starting point for crafting solutions aimed at setting and communicating prices for patients and payers. In addition to impacting policy, our insights can provide drug manufacturers and other players in the pharmaceutical supply chain with knowledge that will enable them to better align incentives, which can benefit society through improved price-setting standards.
Prescription Drug Prices in the Pharmaceutical Supply Chain
P rices Paid Throughout the Pharmaceutical Supply Chain
Very few pharmaceutical drugs are distributed directly to the consumer, with most goods flowing through a supply chain (see Figure 1). Drug manufacturers sell drugs to wholesale distributors, who then supply product to retailers (e.g., pharmacies, hospitals, other care facilities such as clinics or physician offices) from which consumers obtain the product recommended by their provider. Pharmacy benefit managers (PBMs) and health plans also provide critical services in this supply chain (The Commonwealth Fund 2019). Although the physical flow of goods is relatively straightforward, Figure 1 shows that the financial transactions are more complicated, with the prices different entities pay varying widely throughout the supply chain.

The pharmaceutical supply chain (from The Commonwealth Fund 2019).
Although the way in which individuals pay for prescriptions is heterogeneous (e.g., insurance copayment, Medicare Part D, out-of-pocket), the prices—as well as markups throughout the supply chain—are based on the manufacturer’s list price (also known as wholesale acquisition cost [WAC]; Department of Health and Human Services 2019). WAC is used as a benchmark for prices in the pharmaceutical market and is what news outlets publish when reporting drug prices (Gencarelli 2005). It is the price published in direct-to-consumer advertising for drugs that have a greater than $35 per month cost to the consumer under a proposed rule set by the Centers for Medicare and Medicaid Services, and it is the amount paid by a patient for a drug not listed on their insurance formulary or for patients who have not yet met an insurance plan deductible (Department of Health and Human Services 2018b).
WAC is set by the manufacturer and does not include any type of discount. Discounts are often negotiated by parties in the distribution channel (e.g., PBMs, see Figure 1) and may further impact the cost for patients (U.S. Code § 1395w–3a). The result of pricing negotiations within the pharmaceutical supply chain is that most providers (e.g., pharmacies) do not pay the full list price to the wholesaler (Iacocca and Mahar 2019). Though information on discounts and rebates between the manufacturer and other payers in the system is proprietary, it is estimated that they amount to approximately 28% of the list price (Department of Health and Human Services 2016).
The price that consumers pay is a function of the retailer, but it is also driven by insurance coverage (for those who have it) and is impacted by negotiations with PBMs, who work closely with insurers and public health programs (e.g., Medicare and Medicaid) to specify which drugs are covered by prescription plans and what the consumer will pay (and the retailer will receive) at point of purchase (Kaiser Family Foundation 2005; The Commonwealth Fund 2019; Curtiss, Lettrick, and Fairman 2010). Despite this variability in the price different entities pay, the importance of drug list prices (i.e., WAC) is undisputed, as these prices form the basis of all supply chain negotiations and thus impact the final cost borne by the consumer (Entis 2019). Moreover, insured consumers often assume some cost either as a direct result of purchase or through insurance premiums that rise with prescription costs (Rivers, Nall, and Frimpong 2006). High prices can be particularly significant for the 47% of Americans with high deductible plans, under which they pay drug list prices until their deductible is met, and can be financially devastating to the approximately 12.1% of the population that is uninsured (Berchick, Hood, and Barnett 2018; Department of Health and Human Services 2019).
Drivers of Price Throughout the Pharmaceutical Supply Chain
Our analysis looks at the relationship between WAC and multiple variables related to (1) the brand, including brand patent protection and the manufacturer; (2) the product, in terms of the amount of the active ingredient and number of dosing levels on the market; (3) the condition the drug aims to treat, with respect to the drug classification and product tier; and (4) factors that define the market in which the drug competes, including the number of competitors and the approval date for the drug, which reflects its age in the marketplace. Further discussion on the data is provided in the following section. Next, we review extant research that considers these and related factors. As revealed when reviewing the literature, the relationships between these variables and price are sometimes at odds with what we would predict using economic theories of pricing, and some relationships reveal inconsistent results across published studies.
Brand elements: patent protection and manufacturer
A critical factor in the pricing of a pharmaceutical drug is whether the formulation is patent protected. A standard patent protects a brand for 20 years, although companies can file for extensions for the time it takes to apply for FDA approval (Kesselheim, Avorn, and Sarpatwari 2016) or secondary patents when a drug is repurposed (e.g., when a company develops a new form of administration or determines a new use for an existing medication via research and testing; The Economist 2019). This creates conditions of market exclusivity and generally higher prices due to a lack of competition. Though it seems logical to posit that brand prices fall when generic substitutes enter the market, research in this area shows otherwise. While the entry of generics provides lower cost alternatives for consumers in the form of generic drugs, the downward pressure on prices is typically felt by other generic (vs. brand name) competitors (Berndt et al. 2007; Frank and Salkever 1997). In fact, the prices of brand name drugs typically increase when generic substitutes enter the market as brands focus on selling to a smaller segment of brand-loyal consumers (Iacocca, Sawhill, and Zhao 2015). This finding conflicts with the tenet that enhanced competition should exert downward pressure on prices, and so exploring the relationship between generic and branded products and price in the market is critical.
In addition to the brand name that identifies a drug, the manufacturer brand name (e.g., Bristol-Myers Squibb, Pfizer) may also be of consequence in pricing. Some experts argue for the importance of manufacturer-level branding that can be used to link the multiple pharmaceutical brands sold by one firm (e.g., Pfizer branding its Gardasil and Januvia products with the Pfizer brand name; Schuiling and Moss 2004; Moss and Schuiling 2004). This strategy of “umbrella branding” is frequently used—particularly by global firms—in consumer goods industries (e.g., food, cars) and allows new and existing products to capitalize on the brand awareness, image, and trust of a large and well-known brand (Schuiling and Moss 2004); however, it should be noted that there are also risks with this strategy, as negative publicity at the corporate level can similarly impact the individual brands. Herein, we seek to explore the effect of the manufacturer’s brand name on drug prices to see if brands impact pricing. Perhaps well-known manufacturers confer equity to their specific drug brands, driving higher prices. However, some work suggests that the development of new products under new brand names (which would relate to price via newer treatments in emerging therapeutic classes; see the subsequent discussion) in the pharmaceutical industry is fairly widely dispersed and has continued to grow over time (DiMasi 2000). This suggests the importance of exploring the role of the manufacturer as a driver of price.
Product attributes: amount of active ingredient and dosing level
Individual product attributes related to active ingredient levels and dosing level may vary by particular items in the product line. The amount of active ingredient—typically measured in milligrams (mg)—of the product refers to the different amounts of active ingredient in one tablet or capsule, whereas the dosing level refers to the number of different dosing strength options; for instance, Coumadin—a blood thinner produced by Bristol-Myers Squibb—is available in at least eight different dosage levels, ranging from 1 to 10 mg. When pricing different formulations, manufacturers tend to use either a flat pricing strategy, whereby they price all dosing strengths the same, or a strategy in which they charge a price proportionate to the amount of active ingredient contained in the product (Jönnson 2001). In our analysis, we consider both milligrams and the number of dosing levels as potential drivers of price. In addition, from a policy perspective, these two factors matter as consumers (and physicians) become increasingly concerned about the price of pharmaceutical drugs. On the one hand, flat pricing is in line with the idea that physicians may start out at the lowest dose needed, ramping consumers up in milligrams as required to treat their condition. On the other hand, some argue that such pricing may lead to behaviors like splitting higher dosages (e.g., breaking a pill in half) to design a more cost-effective means of treatment (Jönnson 2001).
Condition factors: drug classification and product tier
Drug classification systems work to categorize pharmaceuticals into a hierarchy that reflects the manner in which a drug works, the physiological effects of the drug, and the therapeutic indication, which describes the primary condition or pathology the drug is used to treat (Mahoney and Evans 2008). In our inquiry, we focus on therapeutic class (e.g., heart and circulatory, central nervous system) and subclass (e.g., beta blockers, depression), given the documented role of these indications in pharmaceutical pricing and their correspondence to the underlying condition the drug treats. We also consider the formulary tier in which a medication is classified. Although these vary by health plan and provider, formulary tiers categorize drugs on the basis of factors like cost and therapeutic class. Tier 1 drugs are usually generic, have lower copays, and are automatically approved for coverage. Tier 2 typically consists of brand name drugs with higher costs, and Tier 3 is limited to drugs that treat specialized conditions, as well as those for which there are no lower-cost substitute products (Torrey 2019).
Extant research demonstrates that the key determinant in the pricing of new pharmaceutical drugs is the therapeutic value of the drug (Lu and Comanor 1998). A 2009 report on escalating drug prices produced by the U.S. Government Accountability Office found that for the majority of pharmaceutical drugs that experience extraordinary price growth (between 100–499%), price increases can be attributed to a lack of therapeutically equivalent drugs and limited competition (Government Accountability Office 2009). Indeed, there are certain classes of drugs that by nature of the severity of the symptoms and/or conditions they treat may carry higher price tags (Danzon and Taylor 2009; e.g., cancer drugs). In line with the economic argument to price drugs on the basis of the marginal value provided to the patient using the drug (see Pauley 2017), the concept of value-based pricing for pharmaceutical drugs—although ethically questionable—is one that has been discussed among policy makers and consumers (Feldman 2019). In exploring the role of drug class and tier in driving WAC, we seek to uncover if there is any relationship between condition and pharmaceutical pricing.
Market factors: competitive landscape and approval date
Another issue impacting pharmaceutical drug pricing is the number of drugs that compete in a particular category, or therapeutic class. As previously discussed, patents provide protection for formulations. Although patents cover the molecular entities that constitute pharmaceutical drugs, “me too” drugs that offer a similar therapeutic advantage—though with a different composition—work to create competition for prescription drugs even when those drugs are patent protected (Kessler et al. 1994). Sometimes these options offer benefits via enhanced efficacy or patient tolerability, and a pricing analysis reveals that these options, though not first to market, are often priced at a higher level than pioneering drugs in the category (Jena et al. 2009). As one example, Gleevec was introduced by Novartis in 2001 as a revolutionary treatment for leukemia with a list price of $26,000 per year; now, several drugs compete in the same category with prices of about $150,000 annually—far exceeding the cost of the pioneer (Rosenthal 2018). This runs contrary to the general belief that enhanced competition should drive prices down; indeed, many suggestions for drawing down current prescription prices set forth in the 2018 American Patients First Blueprint are built on increasing competition (Department of Health and Human Services 2018a). We explore the number of drugs in a therapeutic class to explore whether and when increased competition results in lower or higher drug prices.
We also consider the FDA approval date in our analyses, as data demonstrate interesting trends related to a drug’s age and its market price. Here, pricing theory might suggest that older, established entrants would be able to maintain a higher price. That said, physicians or patients may infer that newer entrants are of superior quality, which suggests that prices of existing drugs will decrease with new market entrants. Trend data demonstrate that of the 49 top-selling brand name drugs in the U.S. market, 48 saw regular annual or even biannual price increases from 2012 through 2017 (Wineinger, Zhang, and Topal 2019). Examples of older drugs undergoing dramatic price increases receive immediate attention—such as the now infamous cases of Daraprim, a 62-year-old off-patent drug used to treat parasitic infections that increased in price from $13.50 to $750 per tablet after being acquired by a pharmaceutical start-up, or standard Epi-Pens used to treat severe and life-threatening allergic reactions rising in cost from $100 to $600 in 2009 (Long 2016; Pollack 2015). Overall, the public has witnessed dramatic increases for older, off-patent drugs (Pollack 2015). Therefore, we explore the relationship between current price and number of years on the market.
Methodological Approach and Data
We used an analytics-based approach to examine the following factors as potential drivers of pharmaceutical drug prices: brand/generic identification and the role of the manufacturer (brand elements), milligrams of active ingredient per dose and the number of dosing levels (product attributes), drug classification and product tier (condition factors), and the number of drugs in a therapeutic class in addition to years on the market (market factors).
Analytics-Based Approach
Using analytics was appropriate in this context for two key reasons. First, our data set was heterogeneous, having been collected via data scraping from multiple publicly available and proprietary sources. In addition, rather than testing formal hypotheses using a standard deductive method, we adopted an inductive approach aimed at providing insights based on more complex relationships in the data; in other words, we used a bottom-up approach common in the analytics ecosystem, which allowed relevant insights to emerge from patterns observed in the data. Our methods allowed us to explore not only direct relationships (e.g., the impact of brand/generic classification on WAC) but also multilevel interactions between the variables. It is the multilayered understanding of ways in which our variables interact in driving price that we believe will yield the richest insights into understanding the drivers of pharmaceutical prices.
The analytical techniques used in this inquiry fall under the categories of visualization and machine learning. Visualization is a go-to first step in analytics, often referred to as the descriptive phase (what happened?) of analytics. It is also often used in the diagnostic phase to explain why relationships are observed. Although our inquiry into drug pricing is exploratory, guiding questions from the literature include the following: Do average drug prices decrease as the number of available drugs in a therapeutic class increase? Are generic drugs always priced lower than branded ones? Do average drug prices decrease as more doses of a drug become available? Is pricing a function of the amount of active ingredient present in the formulation? Are older drugs less expensive than newer drugs? How do branding elements and the condition a drug aims to treat impact these relationships?
It is important to note that the exploratory phase of analytics may reveal insights beyond those provided by the answers to these questions—this is one of the key benefits of such analysis. After we gained insights through the exploratory process of visualization, we then used machine learning in the predictive phase of analytics. Machine learning differs from traditional predictive models (e.g., regression) in the interpretation. Traditional models use a sample to predict the behavior of a larger population. Machine learning models use a sample to predict behavior during a specific instance. Thus, machine learning models use data to predict what will happen for a specific event and further drive conversation (or policy) on how to influence the prediction. We focus on the accuracy of our results through the following question: Do drug prices follow enough logic that machine learning can accurately predict prices using the drug’s characteristics?
If we find a model that predicts drug prices accurately, policy makers can identify characteristics that are associated with lower drug prices. That said, we may find no such model for drug price prediction. This would suggest that factors that are generally accepted as lowering drug prices do not, in fact, play a role. Through this process of elimination, policy makers can remove some of the opacity surrounding drug prices—a critical first step in wrangling them.
Data Description
Our data set contains 2,088 of the most frequently prescribed drugs. We collected data from both publicly available and proprietary sources. The price (WAC 1 ), dosing levels, and milligrams of active ingredient (denoted by mg) for each drug and manufacturer were provided by a major retail pharmacy at a single point in time at the conclusion of 2015. We used multiple public sources for the remaining information (see Table 1). Where permitted, we scraped data using R programming language, and the remainder of the data were collected manually. In total, we utilized seven disjointed data sets that were cleaned and merged using Tableau software, an international leading visualization tool that uses best practices for analytic analysis.
Data Source Information.
Note here that all drugs were assigned a single therapeutic class on the basis of the primary condition that they treat. We collected data on the primary therapeutic class and subclass (a more specific condition than the therapeutic class). Our data set includes 14 therapeutic classes and 85 subtherapeutic classes. We also gathered the tier that each drug belongs to as well as the number of competitors from each therapeutic class. In addition, the patent filing and approval dates for branded drugs, as well as the number of drugs in a class and brand/generic classification, is captured in the data set. A full description of the data and an example of the variables appears in Table 1.
Technique 1: Visualization
Despite its perceived simplicity, visualization remains one of the most valuable tools in data analytics. As data sets become larger, visualization becomes even more valuable because it easily gives the analyst descriptive characteristics of the data. Furthermore, visualization provides output that is easy for readers to understand; in fact, organizations identify the ability to visualize and effectively communicate data as one of the most valuable analytical techniques used to understand complex problems (LaValle et. al 2011). Advances in computing power and visualization software enable researchers to gain more insight on data sets through more robust visualizations. We created visualizations and dashboards using Tableau software. Dashboards allow us to combine multiple visualizations into one frame so they can be viewed side by side, allowing complex interrelationships between the data to be understood. Data visualization is critical to our inductive approach to considering drug prices, as it allows us to explore both direct and interactive effects of variables rather than test a priori hypotheses.
The purpose of visualization in this inquiry is to uncover patterns in pricing when looking across different variables. Although much of our data is publicly available, it is fragmented and therefore difficult to access and understand in its entirety. This is where visualization shows its strength—allowing data sets to be joined, filtered, and cleaned. Many machine learning techniques lead to high predictive power (as discussed in the following section), but interpretation and understanding are key strengths of visualization techniques. Here, we rely on visualization to glean insights into the relationships between the variables identified in our data set.
Results and Discussion
The top panel of the dashboard presented in Figure 2 shows the average WAC for branded and generic drugs by therapeutic class (e.g., cancer drugs, blood modifying drugs). The lower panel shows the number of branded and generic drugs in each therapeutic class. Through this and Figure 3, which shows the total drugs in each class, one can quickly see that there is an overall inverse relationship between the number of drugs in the therapeutic class and brand drug prices. With some exceptions, more drugs being in a therapeutic class generally leads to lower prices. Four therapeutic classes price their branded drugs far higher than the other classes: cancer drugs, blood modifying drugs, central nervous system, and anti-infective agents. This is perhaps not surprising, as these categories are associated with life-sustaining benefits (although this observation is not meant to reduce the value of prescription drugs in other categories). However, the most surprising observation is in the genitourinary therapeutic class, in which the average price of generic drugs exceeds the average price of branded drugs.

Brand and generic price and quantity comparison (N = 1,682).

Total drugs in therapeutic class, sorted by price (descending) (N = 1,682).
The dashboard presented in Figure 4 drills deeper into the genitourinary therapeutic class by separating the price by tier and comparing prices for individual branded and generic drugs. Surprisingly, many generic drugs are in Tier 1 and are more expensive than the branded drugs in Tiers 2 and 3. Drugs in Tier 1 are considered preferred drugs and are more likely to be covered by insurance plans. An advantage of our method is that it allowed us to dive deeper into this perhaps unexpected relationship. When adding the subtherapeutic class variable, we discovered that most of these higher-priced generics were in the urinary tract spasms class.

Instances in which generic prices are higher than brand prices (N = 102).
Upon further investigation, we found additional instances in which generic prices were more expensive than branded drug prices when holding their tier status constant or, in other words, when Tier 1 drugs were notably higher priced than Tier 2 and Tier 3 drugs. Figure 4 shows three additional instances in which there are many generic drugs that are priced higher than branded drugs within a subtherapeutic class, including depression, heart rhythm, and fluid retention drugs. Interestingly, these are all therapeutic classes that require ongoing treatment for the condition.
Figure 5 shows an interesting relationship between factors that affect brand and generic prices by highlighting the range of prices at different dosing levels. It appears that the number of dosing levels plays a more significant role in the prices of branded drugs than generic drugs. For branded (vs. generic) drugs, list prices are reduced and stabilize to a greater extent when more than one dosing level is introduced. Our data also suggest that the relationship between price and dosing level may not be linear—this can be seen when we look beyond one dosing level for branded drugs. The WAC prices for drugs with lower dosing levels are similar and much lower than for drugs with multiple dosing levels. In general, we would expect additional analysis to reveal that the number of dosing levels are inversely related to prices.

Dosing level versus price for brand and generic drugs (N = 1,982).
Visualization of the relationship between age and price tells a somewhat surprising story—that the age of the drug does not influence the price (see Figure 6). It should be noted that there is a spike in prices for branded and generic drugs approved after 1990. Indeed, in 1992, the FDA instituted a fee structure for pharmaceutical manufacturers that fueled the drug approval process. This allowed more drugs to enter the market as potential grew (Frank 2018). If a price increase was due to a blockbuster drug, we would not see the price jump in generics at the same time. Visualization did not reveal insights related to the manufacturer or milligrams of active ingredient.

Price of brand and generics by age (N = 1,760).
Our visualization looked at variables in isolation as well as together to reveal insights into the relationship between various factors and drug prices. In summary, our visualization analysis uncovered some key insights. With respect to competition, across therapeutic classes, our analysis shows that more competition (as measured by the number of brands in a therapeutic class) leads to lower pricing. Moreover, support for the commonly held perception that branded drugs have substantially higher prices than generic counterparts is seen in four therapeutic classes: cancer drugs, blood modifying drugs, central nervous system, and anti-infective agents. Considering the conditions they treat, this pricing strategy may not be surprising, as innovation may be particularly relevant in these categories given that they treat specific and life-threatening conditions. Interestingly though, prices for generic drugs were higher than for branded drugs in the genitourinary, depression, psychotic bipolar disorders, heart rhythm, and fluid retention therapeutic classes.
When exploring the relationship of branded versus generic drug prices by tier, we find that generic drugs are more expensive in Tier 1 than in Tiers 2 or 3. This is important, as Tier 1 drugs are those typically covered by insurance plans. That said, generic drugs are perhaps more likely to be found in this category. Lastly, our analysis related to dosing levels reveals that more dosing levels lowers the price of branded drugs, whereas more dosing levels does not appear to impact the pricing of generic drugs. A summary of these findings can be found in Table 5.
Technique 2: Unsupervised Machine Learning
Machine learning is an application of analytics that utilizes patterns in the data to “learn” about relationships in a data set. Rather than relying on a specific set of predetermined rules or specifications, machine learning techniques build models on the basis of a sample from a data set (i.e., training data) and use what is learned from patterns in the data to make predictions. The main difference between machine learning and statistics is the focus of analysis: statistics generates inferences about a population on the basis of a sample, whereas machine learning predicts individual records and uses these predictions to identify patterns (Shmueli, Bruce, and Patel 2016).
The goal of unsupervised machine learning is to learn the natural structure of the data without explicitly provided labels (Soni 2018). Unsupervised machine learning is common for exploratory analysis. By contrast, supervised machine learning uses data to learn the route from the inputs to the output. Machine learning intentionally employs several techniques, and we first report the results of a cluster analysis. The goal of clustering in our analysis is to group drugs on the basis of similar characteristics. Because clustering is an unbiased technique, it allows us to see if drugs that are clustered together align in terms of a specific variable; for example, clustering would allow us to see whether a majority of drugs in the high-price cluster are in the cancer therapeutic class, or whether high-priced drugs are more likely to be branded (vs. generic) offerings.
The clustering technique applied in this analysis uses the k-means algorithm, where k is the number of clusters. The algorithm is as follows: Randomly select k centers within the data. Assign each data point to the cluster that minimizes the distance to the center. Recompute the center of each cluster by taking the mean of all the vectors in the cluster. Return to Step 2.
The algorithm continues until the values for the center of each cluster stabilize (i.e., when there is very little change in Step 4). We followed industry practice and tried values between 2 and 8 for k. The final model included the value of k that balanced the between-cluster distance and within-cluster distance. We hoped for clearly defined clusters; thus, we desired a large between-cluster distance and a small within-cluster distance. Consistent with best practices, it should be noted that our cluster analysis did not consider categorical variables or variables with missing values.
Results and Discussion
The ideal number of clusters was found to be three when comparing between-cluster and within-cluster distances. Table 2 summarizes the results of the analysis and shows the average value for each cluster. The shading shown in the table is a result of conditional formatting, highlighting trends in the data. Here, lighter shading represents values below average, and darker shading represents values above average. This helps us characterize the clusters as follows:
Cluster 1 – High-priced drugs: newest, highest milligrams of active ingredient, fewest dosing levels Cluster 2 – Low-priced drugs: oldest drugs, many competitors Cluster 3 – Moderate-priced drugs: many dosing levels
Results of Cluster Analysis for Branded and Generic Drugs.
Notes: Lighter shading represents values below average, and darker shading represents values above average.
Not surprisingly, most (57.6%) of the generic drugs are in Cluster 2, and most of the brand drugs (71.6%) are in Cluster 1. Consistent with marketing literature, we observe that increased competition and age are clearly associated with lower drug prices.
For comparison purposes, we performed a second cluster analysis, this time exclusively of branded drugs. Again, the ideal number of clusters was found to be three. Table 3 shows the results, and the clusters can be defined as follows:
Cluster 1 – Low-priced branded drugs: oldest, lowest milligrams of active ingredient, many dosing levels Cluster 2 – Moderate-priced branded drugs: average pricing and characteristics Cluster 3 – High-priced branded drugs: newest, highest milligrams of active ingredient, few dosing levels
Results of Cluster Analysis for Branded Drugs.
Notes: Lighter shading represents values below average, and darker shading represents values above average.
When looking exclusively at branded drugs, it is interesting to note that the number of drugs in the therapeutic class does not drive prices. For example, Cluster 1 is the least expensive group, yet it has the least number of competitors of any cluster. By contrast, the number of dosing level options does appear to inversely relate to prices, as drugs with more options tend to be clustered in the lower-priced groups. Finally, when looking at the therapeutic classes of the drugs in each cluster, we find that only half of the therapeutic classes are represented in Cluster 1: anti-infective agents, blood modifying drugs, central nervous system, hormone and diabetes, neuromuscular drugs, and pain relief. This indicates that there are characteristics within those therapeutic classes that are similar. Despite Figure 2 showing that blood-modifying, central nervous system, and anti-infective therapeutic classes have some of the highest WAC prices on average, cluster analysis reveals that there is at least one drug in each of these classes that has a lower price—so much so that it is clustered into the low price group (Cluster 1). Building on these insights, we now move on to supervised machine learning aimed at predicting drug prices.
Technique 3: Supervised Machine Learning
Using an analytics-based approach, different machine learning techniques are used and then compared to determine which algorithm performs best for a specific data set. Here, we conducted and compared three techniques: k-nearest neighbors, bootstrap forest, and artificial neural networks. To perform all supervised machine learning techniques, we first had to randomly partition our data set so that 25% of the original data set became a validation data set. The remaining 75% of the data became the training data set.
k-Nearest Neighbors
The goal of k-nearest neighbors (k-NN) is to classify (predict) categorical (quantitative) variables using an algorithm that relies on finding “similar” records in training data (Shmueli, Bruce, and Patel 2016). k-NN is a supervised, nonparametric method in which k-records are identified as similar to a new record that we wish to classify/predict. The training data refer to the portion of the data that is removed from the data set and used to build the model. By partitioning the data, some of the data can be used to build the model while separate data can be used to test/validate the model. The “neighbors” for each record are determined on the basis of their Euclidean distance between each record in the validation set and every record in the training set.
The value of k is chosen on the basis of model performance after trying different values of k. The final value of k balances overfitting and underfitting the data. As k increases, the risk of overfitting due to noise in the training data decreases; however, as k decreases, we increase the risk of missing useful information in the predictors and approach a method similar to that of the naïve rule (i.e., k = 1). The ideal value of k balances these risks and is selected on the basis of the lowest misclassification rate. Typically, k is between 1 and 20.
Extant studies demonstrate the potential use of k-NN in the health care industry (e.g., breast cancer diagnoses, medical records; Sarkar and Leong 2000; Shee, Cheruiyot, and Kimani 2014). Building on this, we first used k-NN to classify drugs as generic or branded. The logic here was to evaluate if prices and other factors can accurately predict brand/generic classification. We next used k-NN to predict the drug prices. If drugs were priced according to characteristics in our data set, k-NN would have a high degree of predicting the price of similar drugs. This would be valuable for policy makers when determining if prices are in line with expectations (i.e., if they follow the model). Practically, using a technique like k-NN may alert drug manufacturers—as well as policy makers, industry trade associations, and watchdog groups—when a price of a drug is out of an existing range. On the one hand, this information might be used as a means of self-auditing for manufacturers and could offer a gauge for when to address pricing issues. On the other hand, k-NN could also flag high-priced drugs, in turn prompting further investigations by the FDA, states’ attorneys general, or other investigatory bodies.
We selected the predictor variables on the basis of the results from the visualizations, and we included the manufacturer, price, brand/generic classification, subtherapeutic class, tier, number of drugs in the class, dosing levels, milligrams of active ingredient, and approval year. Figure 7 shows that the lowest misclassification rate is when k = 8.

Misclassification rate for different levels of k in k-NN analysis (N = 2,088).
Using k = 8, Figure 8 shows the confusion matrix for the validation data (N = 522), detailing the misclassification rate. Interestingly, the model misclassified branded drugs as generics far more often than misclassifying generic drugs as branded (75 vs. 39 instances). That is, 32.61% of the branded drugs in the validation set had characteristics more like the generic drugs. Our model was much more successful at classifying generic drugs according to the variables included, with a misclassification rate of only 13.59%.

Confusion matrix (N = 522).
Using k-NN for predicting drug prices yielded an R-square value of .34. Unlike regression, interpreting effect size for each predictor did not provide relevant insights about the relationship between variables. Our k-NN analysis revealed that (1) some branded drugs (approximately 33%) have characteristics that are more similar to generic drugs and vice versa (i.e., approximately 14% of generics are similar to branded drugs) and (2) the predictability of k-NN for drug prices seems relatively low, indicating that prices are not derived on the basis of similarities to other drugs in terms of criteria including manufacturer, therapeutic class, dosing level, milligrams of active ingredient, or approval year. While k-NN did not explicitly point toward clear indicators that drive drug prices, policy makers can understand from this analysis that the analyzed characteristics do not tell the whole story. For example, it would be erroneous to point to a single manufacturer for having higher drug prices, or to define a certain therapeutic class as being lower in price. A key question that does arise from these findings relates to why some generic drugs have characteristics (including prices) that are more like branded drugs.
Like other machine learning models, k-NN is known for its high accuracy, but it lacks in interpretability of the model. Alone, we are not sure if this model is the best option for predicting drug prices. We proceed with two additional supervised machine learning techniques so that the three models can be compared. Next, we consider the method of bootstrap forest.
Bootstrap Forest
Bootstrap forest is a popular ensemble method, meaning that it combines multiple machine learning techniques to produce more robust models. In the case of bootstrap forest, multiple decision trees are combined to produce a stronger tree than any of the individual trees. The purpose of bootstrap forest in our analysis is to predict drug prices. The predicted price for an individual drug is the average of that drug’s predicted value across multiple decision tree models.
Bootstrapping means that samples are drawn with replacement. By training each tree on a different sample, the combination of all trees reduces the error that may be present in single trees. A strength of this method is the ability to limit overfitting. Whereas a single decision tree needs pruning to avoid overfitting, bootstrap forest has enough trees to aggregate the results and leave all trees unpruned. In addition, bootstrap forests partition the data and train using different samples of the data (see Hastie, Tibshirani, and Friedman 2009 for more information on methodological detail). Trees are fitted as follows: For each tree, select a random sample of n observations; Select a random set of predictors for each split; Continue splitting until a specified stopping rule is achieved; Repeat for multiple trees.
We used the same predictors that were used in k-NN so that the methods are comparable: the manufacturer, brand/generic classification, subtherapeutic class, tier, number of drugs in the class, dosing levels, milligrams of active ingredient, and approval year. We grew 100 trees using the algorithm described previously and then averaged for a final forest. We used the default settings for the minimum (five) and maximum (2,000) splits per tree, and the minimum number of observations per split was kept at the default of five. Early stopping was permitted if continued tree growth did not improve the R-square value of the model.
Each tree is grown through a process of splitting the variables to reduce impurity within the branch. That is, a variable is selected to split if it yields the greatest benefit toward the objective of purity, whereby each branch predicts prices with certainty. A “split” defines a rule, and the collective set of rules on a path (from beginning to end) gives the prediction. For example, a single tree may produce a rule that states “when it is a branded drug, in the subtherapeutic class of thyroid disorders, produced by AstraZeneca, and approved in 1998, the price is predicted to be $48.”
Figure 9 shows the cumulative validation report, in which the validation statistic shows the R-square value as more trees are added to the model. The final model shows that 18 trees were grown to yield an R-square value of .54. The number of splits shown in Figure 9 is a count of how many times that variable was selected to split at a branch within the forest. The column contribution report (see Figure 9) suggests that the therapeutic class is the strongest predictor of drug prices, followed by manufacturer, age, and the amount of active ingredient (in mg). That is, the primary purpose of the drug explains the variability in the cost of drugs more so than any other variable in the data set. Interestingly, brand/generic classification was not a top driver of price, and we acknowledge here that there may be a relationship between branded versus generic status and manufacturer, as some manufacturers may primarily (or exclusively) produce branded or generic drugs. These findings are significant and build on our prior analyses to consider the relative contribution of our set of variables to pricing. Next, we perform the final supervised machine learning technique, artificial neural networks, and compare our methods.

Results of bootstrap forest analysis (N = 2,088).
Artificial Neural Networks
Artificial neural networks (ANN) is a data-driven machine learning method known for its high predictive power. ANN can capture complex relationships between predictors. Unlike regression, ANN does not require the user to define the relationship between predictors because ANN learns the relationship from the data. ANN has had successful applications in the financial and engineering industries (Kaastra and Boyd 1996), and in the pharmaceutical sphere, players and policy makers might use ANN for detecting prices that are out of line with predictions.
ANN consists of hidden nodes that are nonlinear functions of the input variables. Although the hidden nodes can model complex relationships, they do not yield easily interpretable results because there is no direct path between the inputs and prediction. Like many machine learning techniques, the model may be highly useful for predicting a drug’s price, but insight into the contributing factors may not be clear. However, one advantage of ANN is understanding the degree of interaction between predictor variables. Indeed, ANN can detect complex linear and nonlinear relationships between the variables to give the researcher greater understanding of the data.
As in the other supervised machine learning techniques, we used eight variables to predict drug prices (manufacturer, brand/generic classification, subtherapeutic class, tier, number of drugs in the class, dosing levels, milligrams of active ingredient, and approval year). We selected the option to transform the covariates to mitigate the effect of outliers and heavily skewed data. The training model used least absolute deviations on the continuous variables to mitigate the effect of outliers. Four nodes were specified for the hidden layer (see Figure 10).

Results of artificial neural network.
As seen in Figure 10, a four-node hidden layer yielded an R-square value of .36. That is, ANN can predict (and therefore validate) drug prices by explaining 35% of the variation in the observed prices. However, ANN does not effectively explain the rationale of pricing.
Comparing Supervised Machine Learning Models
Our next step was to compare the results to determine the techniques that were most valuable in predicting drug prices. Table 4 compares the R-square values. In summary, the k-NN and ANN techniques did not have that highest predictive power (as determined by the R-square value) and yielded little insight because of their “black box” characteristics. Bootstrap forest resulted in the highest predictive power among the machine learning techniques, with an R-square value of .54. Because of its relative accuracy in predicting prices, we rely on the insights from bootstrap forest moving forward when synthesizing our findings.
Model Comparison.
Discussion: Pharmaceutical Pricing Insights and Implications
Synthesizing Insights from the Analytical Models
By synthesizing our results from visualization, clustering, and bootstrap forest, we are able to present a variety of inductive insights related to the pricing of pharmaceuticals. Table 5 summarizes the variables that each technique found to contribute to drug prices.
Summary of Relevant Variables in Analysis and Key Insights for Future Research.
Notes: Not all techniques examined all variables. Blank areas in the table indicate that the variable was not examined using the technique.
When considering factors that contribute to drug prices, the variables that were studied here are key pricing facets, according to the literature. As a result of our analytics-based techniques, we found that 53% of the variability of drug prices can be attributed to these factors. We expand on the findings from our analysis and propose directions for deeper inquiry based on our findings.
Insights related to brand elements
With respect to patent protection, our clustering and visualization analyses found that branded drug prices are usually higher than generic drug prices, as one would expect. However, the genitourinary, depression, psychotic bipolar disorders, heart rhythm, and fluid retention therapeutic classes had instances in which generic prices were higher than branded prices, suggesting that general beliefs about the brand–price relationship may not hold across all therapeutic classes. This knowledge gives policy makers a clear focal point for further examination, as understanding the pricing of drugs on the basis of patent status would require consideration of why this pattern does not hold true across all drug types, and whether there are specific factors related to these therapeutic categories or conditions that drive these effects. Future research might consider whether the observed differences in price across class reflects a meaningful distinction, perhaps related to research and development expenditures or differential value in treatment, which would support a value-based model of pricing (Feldman 2019; Pauley 2017). Although our methods—by definition—do not provide insight as to why differences occur, exploring outlier pricing patterns, such as the one observed here, would be a good starting point for policy-based research inquiries to delve deeper into the specific mechanisms underlying this effect.
Via supervised machine learning—and more specifically, our bootstrap forest analysis—our findings demonstrate differences in prices between manufacturers. Clarifying the role of manufacturer in price is a critical direction for future research. Taking a branding perspective, these price differences may be due to stronger brand names serving as a promise of quality and safety, thus affecting consumers’ willingness to pay. From a more operational standpoint, these price differences could be due to the types of drugs a manufacturer produces. For example, a manufacturer may concentrate on developing drugs in a category in which prices are high due to perceptions of value (e.g., cancer drugs), or perhaps they develop branded versus generic drugs. In addition, the price differences could be due to other expenditures, such as research and development or marketing and advertising, which vary considerably between manufacturers (Capella et al. 2009; Kanski 2019). We believe this to be a focal point for research aimed at providing insight for policy makers by clarifying the drivers of price in relation to branding elements. Because the relationship between these variables and price does not always align with economic theory—particularly our findings related to patent protection—it is important to understand these relationships so that they might drive policy, including consumer-facing policy efforts aimed at promoting transparency.
Insights related to product attributes
Prices in this complex industry do follow some economic pricing principles. For instance, the amount of a drug’s active ingredient was shown to be a significant factor in drug prices in both our clustering analysis and our results using the bootstrap forest technique, supporting the notion that active ingredients themselves contribute to drug prices. If true, this would rule out manufacturers using a flat pricing strategy across the drugs explored here (Jönnson 2001). Though this might not be perceived as delivering high value to patients (i.e., some patients get more of the drug for the same price), it is perhaps a worthwhile strategy from a consumer welfare standpoint, particularly for drugs that have nefarious side effects or addictive properties. In such cases, this variable pricing strategy would remove incentives to provide larger dose scripts to patients as a means of lowering costs.
Our data also reveal that the number of dosing levels plays a role in drug prices, especially for branded drugs. Interestingly, both our cluster analysis and visualization show that higher-priced drugs have fewer dosing levels. Further analysis could explore the relationship between prices and additional dosing levels to determine if there is a causal relationship. Given our findings related to the amount of active ingredient, offering intermediate levels may be one way to reduce prices for patients who can tolerate lower doses of a particular drug.
Insights related to condition factors
Consistent with prior work, our analysis demonstrates that therapeutic class is one of the largest drivers of price for branded drugs, accounting for 33.5% of the variance in our bootstrap forest model. To put this in context, our visualization shows that cancer drugs, blood modifying drugs, central nervous system, and anti-infective drugs have substantially higher brand prices than other therapeutic classes. Indeed, there are some therapeutic classes that tend to see high prices and growth due at least in part to the nature of innovation in that product class (Government Accountability Office 2009). Future research could investigate specific characteristics of these therapeutic classes to understand the factors that drive such variability among categories. Research and development expenditures may play a role, as the pursuit of treatments and cures for diseases with a strong negative impact on quality of life are important not only from a consumer welfare perspective but also from a profitability standpoint. Certainly, extraordinary prices for life-saving treatments are also of concern to the American public, particularly for those who lack comprehensive prescription insurance coverage.
Insights related to market factors
With respect to market factors, we found the age of the drug—or years since patent approval—to be a significant factor in drug prices in both the bootstrap forest and clustering techniques. Older drugs, and especially those approved after 1990, are less expensive than more recently approved drugs. This may be related to the pace of innovation and also the number of lower-price competitors, which generally increases as drug patents expire and generic competitors enter the marketplace. Supporting this, our findings also demonstrate an inverse relationship between number of drugs in therapeutic class and price, which is in line with enhanced competition leading to lower prices. This supports the logic underlying current policy initiatives aimed at increasing competition in the pharmaceutical market as a means of curbing price growth (e.g., American Patients First Blueprint; Department of Health and Human Services 2018a), and it suggests that policy efforts aimed at fostering competition might lower prices.
Key Takeaways for Consumers, Manufacturers, and Policy Stakeholders
Interesting observations emerge from this inquiry that may provide valuable insights to players across the pharmaceutical supply chain, as well as policy makers. In our discussion, we consider how our findings may be of use to policy makers in the pharmaceutical industry and providers and payers within the pharmaceutical system (see Figure 1), as well as state and federal regulators and industry watchdog groups.
Implications for the pharmaceutical industry
In light of calls for increased transparency related to pricing, we suggest that manufacturers continue a dialogue with policy makers and patients to communicate the manner in which pricing is set and to explain the mechanisms underlying pricing differentials. Although publishing the price of a drug (as is suggested by current policy proposals; Kirzinger et al. 2019) would increase transparency, it would not yield further insight into how prices are set. In fact, it might lead to confusion for patients, as this research points to departures from tenets of pricing theory that may be well understood by consumers and other intermediaries in the pharmaceutical supply chain (e.g., generic drugs are not universally priced lower than branded counterparts).
Moreover, there are myriad complexities of price setting in this industry. The prices paid by payers in the pharmaceutical supply chain (e.g., pharmacies, other dispensaries) vary widely and are also different from the price paid by the end user. In reality, the list price of a drug (i.e., WAC) represents the starting point for a series of complex negotiations throughout the supply chain, all of which impact the prices that consumers ultimately pay. Perhaps some consumers pay little attention to the prices of drugs, particularly when some or all of that cost is borne by insurance; that said, uninsured consumers and the growing number of patients on high-deductible plans may be likely to consider prices more closely. Because of this, additional programs are likely needed to further educate the public on the factors that drive the prices in the upstream supply chain. As such, knowledge about pricing and transparency surrounding list price may result in informed consumers that can put pressure on manufacturers to keep prices in check—perhaps by exploring lower-cost comparable options. In addition, consumers are not the only paying party in the system, so to the extent that transparency throughout the supply chain enhances the negotiating power of payers, it may also result in lower prices to parties farther downstream (e.g., patients). Such free market mechanisms may put pressure on parties in the channel of distribution, resulting in lower prices throughout the system.
It is important to note that pharmaceutical companies are entering the conversation about increasing drug prices and taking action. For instance, Merck has committed to publishing “useful” information about pricing and payer discounts, as well as working to ensure that their average prices do not increase at a rate greater than annual inflation (Frazier 2020). Opportunities for enhanced communication, like editorials or white papers, might allow industry players to participate in conversations with members of the pharmaceutical supply chain and to educate the public on the drivers of price.
Implications for the providers and payers
Providers and payers (e.g., PBMs, pharmacies) play a key role in the pharmaceutical supply chain. Arguments related to transparency often center on providing supply chain participants with a stronger basis for negotiating price, and ample research across decision making contexts demonstrates that buyers are able to negotiate better prices when more and better pricing information is available (e.g., Grennan 2013). Here, our results provide some specific guidance related to this negotiation. First, we suggest that insurance companies might more carefully investigate blanket rules related to drug coverage in their policies. As our results demonstrate, generic drugs may not always be less expensive than their branded counterparts, which suggests that policies covering generics, for example, may contribute to additional healthcare costs. That said, collaboration with pharmaceutical companies and others in the pharmaceutical supply chain may help reduce systems-level costs. At least one major pharmaceutical manufacturer, for instance, does not provide coupons for its products when generic competitors enter the market (Frazier 2020). Collaboration with members of the supply chain in this way may help realign incentives, but it is important to note that the complexities of pricing in this industry argue against most “one-size fits all” strategies.
Again, payers in the pharmaceutical supply chain are not likely to pay WAC for a drug but rather a price in which rebates or other discounts that vary by transaction are factored in. This underscores the point that the provision of list prices does not tell the entire story of pharmaceutical pricing and that to truly achieve the negotiating power associated with transparency, information related to rebates should also be made clear to members of the system. Addressing these concerns, and more specifically providing communication summarizing which price benefits are passed on to the consumer and how (e.g., via lower prices paid at the pharmacy or lower insurance premiums), might help foster a dialogue about mutually beneficial pricing in the system.
Taking this into the examination room or pharmacy, where patients enter the picture, providers should consider the costs associated with a treatment to determine if there is a similar course of action available at a lower price. If the goal is to reduce costs in the health care system, perhaps there needs to be an incentive system that supports the choice of lower-priced treatments that are equivalent in terms of efficacy. That said, it is not clear that physicians are prepared to have these conversations, and indeed even arriving at the price paid by patients might be a complicated task for care providers.
Implications for regulators and industry watchdog groups
In the United States, prices are not controlled by the government, yet there is a continual call by policy makers for efforts aimed at lowering the price of prescription drugs. The findings of our inquiry provide insight for some specific, evidence-based policy recommendations, as well as broader considerations for policy makers. The task of public policy is to create a mutually beneficial set of incentives whereby prices are lowered for payers and/or patients without hampering competitive markets or the quest for innovation. This is largely discussed as a bipartisan goal, but the solution is complex due to the complicated nature of the industry.
Perhaps one of the most meaningful insights for policy makers is the lack of clear factors that contribute to drug pricing. That said, our findings do provide insight regarding the drivers of price. As noted in our results, as drugs age, they are associated with lower prices; however, there are additional incentives that could be put in place to lower prices, such as enforcing patents. Because patents allow for extension, companies are sometimes able to delay the entry of generic counterparts. Policy makers might evaluate these situations, again recognizing the need not only to reward innovation (e.g., repurposing an existing drug for an alternative treatment) but also to promote the entry of lower-priced competitors. Along these lines, an additional viable route in the patent arena would be to work for more rapid approvals of generics, biosimilars, and me-too drugs so that these drugs enter the market and create competition, driving down prices (Eichler et al. 2016). We recommend that policy makers consider ways to support this suggestion, of course recognizing the key goal of consumer safety.
If we are to assume that branded drug prices (instead of generic) are a primary concern, we would look to the cluster analysis and see that the lowest-priced branded drugs have the highest dosing options. Perhaps there could be incentives for companies to explore the medical necessity of additional dosing options, such as the FDA expediting the approval of additional dosing options of branded drugs or lowering other barriers to entry. Indeed, the speed of approval must be balanced against the need to proceed with high standards for approval that work to ensure the safety and efficacy of approved drugs and dosages.
Collaboration between policy makers and players in the pharmaceutical supply chain must continue to facilitate an understanding of prices. Research suggests that most of the earnings for branded drugs are retained within the supply chain (Vandervelde and Brownlee 2020). Along these lines, two key questions facing regulators and industry watchdog groups relate to understanding the following: What price is too high, and what amount of profit is fair? These are particularly relevant questions at the time of this writing, when many industry players are rushing to bring COVID-19 vaccines and treatments to market, and consumers and policy makers are anxiously awaiting the pricing of these pharmaceuticals, which are critical to the health and well-being of people and the global economy. Although the pricing of consumer products might be based on factors including the costs associated with production, profit margin, consumer value, and the price of substitutes, the societal good associated with drugs makes an assessment of fairness complicated (and hotly debated; Moon et al. 2020).
At present, 79% of Americans believe that the prices of pharmaceutical drugs are unreasonable and that regulatory solutions should be enacted to draw prices down (Kaiser Family Foundation 2019). As COVID-19 vaccines roll out at the time of this writing, it will be interesting to observe whether this dependence on pharmaceutical and biotech companies alters attitudes toward these companies and/or pharmaceutical price perceptions and price sensitivity. On the one hand, it is possible that current perceptions related to and driven by perceived high prices will grow as consumers recognize their reliance on pharmaceutical products and see vaccination programs as a way for companies to derive profit from a global pandemic that has negatively impacted so many aspects of life. On the other hand, perhaps individual reliance on companies to provide a solution that allows for a return to prepandemic life will be seen as an effort to promote common good, in turn enhancing corporate image and perhaps altering beliefs—at least in part—that drugs are universally priced too high.
We believe that our results provide guidance for how analytics-based analyses can trigger investigations into pricing for cases in which prices appear to exceed norms. For example, our k-NN analysis provides for an understanding of when and which drug prices do not follow patterns of pricing in the industry or within a particular drug class, as dictated by a predetermined set of variables (such as the ones used in this analysis). Bootstrap forest allows us to use characteristics of a drug to determine the expected price. Such tools might allow policy makers and watchdog groups to monitor and determine when prices fall outside of pricing norms and develop a gauge of what constitutes a fair profit in conjunction with consideration of costs. Perhaps prices that fall outside of parameters could trigger investigations. Similarly, such methods could be used for industry trend observation and perhaps even the monitoring of competitors’ prices throughout the supply chain to serve as a form of industry self-regulation.
Key Takeaways of an Analytics-Based Approach to Pricing
This inquiry into the drivers of pharmaceutical pricing uses of tools from analytics to explore which factors both independently and interactively influence price, revealing interesting relationships that may drive future research and provide for a baseline understanding for those involved in pricing and paying for pharmaceuticals throughout the supply chain. While extant literature has considered the manner in which individual factors influence price, inquiries into the set of drivers that explain pricing variability are limited. We see analytics-based methodologies as critical in exploring these variables, in part to generate future research and to drive policy efforts aimed at addressing the rising cost of pharmaceuticals.
Moreover, we acknowledge that, due to our approach, we are limited in our understanding by the variables in our data set. Exploring WAC as a gauge for prescription drug prices indeed ignores some of the intricacies of pricing throughout the pharmaceutical supply chain, which is defined by closed-door payer negotiations that include rebates and coupons for payers from PBMs to patients. Further analysis that considers differential pricing mechanisms would be beneficial, though the data might perhaps be even more difficult to come by given the secrecy inherent in these negotiations.
Furthermore, although our pricing data come from Blue Cross Blue Shield—one of largest insurance companies in the United States—other insurance companies may have varying tier structures for a number of drugs. Future studies could compare tier status across insurance companies and examine if our results hold. Moreover, our pricing data represent one point in time, which allows us to pinpoint price drivers but not to observe trends over time. Future research with time series data would allow these factors to be examined in relation to annual drug price increases and would offer the potential for event analysis to uncover disruptions to pricing in the industry.
While wide-ranging and obtained from a variety of sources, there certainly may be other variables that act as drivers of pricing. One critical factor to consider moving forward would be the role of research and development, as extant work shows that drug prices are associated with increased corporate R&D spending (Giaccotto, Santerre, and Vernon 2005). That said, R&D expenses are also well-guarded in the industry, with some claiming that studies examining R&D expenditures consider only those that are particularly high (Feldman 2019). Future work might consider these expenditures in an analysis aimed at specifying their role in pricing.
Another potential driver of costs is expenditures in marketing and advertising, with research demonstrating escalating expenditures in in this area, including direct-to-consumer advertising (e.g., television advertisements) and marketing efforts aimed at physicians and health care professionals (e.g., advertisements in medical journals; Schwartz and Woloshin 2019). Although existing studies do not support the contention that marketing activities raise the price of pharmaceutical products (Capella et al. 2009), future work might include this variable to explore interactions between marketing expenditures and other product classification variables.
All that said, we believe that there are advantages to using analytics-based models in pricing exploration, as well as in other inquiries in the policy sphere. Rising drug prices are widely acknowledged as an issue in the United States, but there is much finger-pointing as to who is responsible for both high prices and increases over time. An inductive approach allows one to remain agnostic about the drivers of price and therefore to explore which factors and interactions are, in fact, contributing to pricing. Similar inquiries might start to uncover key insights in other arenas in which pricing is unclear, such as medical and hospital services, food deserts, and lending fees, to name only a few. Regardless of context, the field of analytics provides innovative techniques for exploring important issues for consumers and society—and to move us from data to decisions in the policy sphere. We believe the techniques outlined here can generate conversation, yield insights, and help in setting research agendas for key policy issues.
Footnotes
Special Issue Guest Coeditors
Brennan Davis, Dhruv Grewal, and Steve Hamilton
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
