Abstract

The Marketing Science Institute, representing a broad array of businesses, has declared “Capturing Information to Fuel Growth” as one of their five major priorities for academic research. As Wedel and Kannan (2016, p. 97) state, “Although big data’s potential may have been over-hyped initially…it is becoming clear that the availability of big data is spawning data-driven decision cultures…providing [companies] with competitive advantages, and having a significant impact on their financial performance.” This commentary recognizes the huge potential of big data and its analytical partner, artificial intelligence (AI), and provides some thoughts about its use in the context of marketing and public policy with emphases on its current role in health care, marketing tactics, the privacy issues it raises, and its corresponding limitations, as well as some suggestions for ways to increase its value in research and practice.
A Historical Perspective on Big Data
“Big data” is a loosely used term that is often operationally defined as having numerous columns and rows in databases, usually at the level of the unit used in the analysis, typically an individual or company, augmented with data on what else is going on at the time (i.e., the context; Bradlow et al. 2017). To a certain extent, big data has been analyzed for decades. For example, often due to limited computer capacity, researchers combined (aggregated) individual data points on variables of interest into averages at different levels of predictors (independent variables), essentially capturing variation/heterogeneity of response based on homophily in terms of the levels of the predictors. Alternatively, researchers simply sampled from big data (sometimes repeatedly to test robustness) and analyzed small data sets to draw inferences about big data.
Researchers have always used different approaches to analyzing big data. One trend that has emerged in this regard has been the use of meta-analyses. By combining studies, each of which employed multiple data points, one combines different estimates of relations across different samples which, in aggregate, represent at least pretty big data. For example, Keller and Lehmann (2008) used a meta-analysis across studies to develop a website at the Centers for Disease Control to predict the effectiveness of health communications. Similarly, Keller, Lehmann, and Milligan (2009) explored the effectiveness of corporate wellness programs via meta-analysis. The main difference in method between this and big data as it is now employed is that big data offers more opportunities to uncover higher-order interactions, which may capture important relations or simply noise (i.e., be driven by a few data points). One cost is that these newer methods and their outputs are often more difficult to describe (and are purely predictive) and thus harder to sell to managers and public policy makers. Of course, this may be a function of age, such that subsequent generations will be more receptive to relying on predictions with less explanation or causal logic.
Bigger Data and Better Analytics
Having large data sets and mechanical means to analyze them is not new. Rather, data have gradually become available at a more micro level about more things. Similarly, analysis has become simpler; rather than programming in Fortran, researchers now use packages or program in R (many of which, ironically, have Fortran as their basis). Nielsen TV data and scanner panel data were considered big at one time, and software packages such as the UCLA BioMedical programs provided statistical analysis capabilities in the 1960s. The current excitement about big data is driven by the ability to get more micro- (individual-) level data more frequently and connect these to other sources of data. Similarly, the term “analytics” is used to describe the ability to handle more detailed data faster and uncover complex relations (basically higher-order interactions). In other words, it is more of a change in degree than in kind (for example, adding more rows and columns to an existing database as opposed to a completely new data source). Nonetheless, it does present new opportunities to learn about social problems and potential cures for them as well as how marketers can contribute to public policy decision making and implementations.
The growth of analytics and big data in marketing is intimately tied to firms’ access to individual-level transaction or personal data and the computer power to analyze it. Big data enables firms to make their offerings to customers more effective. That is one reason many firms offer loyalty programs, which not only reward “loyal” behavior but also serve as a vehicle for collecting detailed individual-level data. As firms use loyalty programs, they generate more data that make these programs more effective, which encourages expanding them, which generates even more data, and so on. For example, Kopalle et al. (2012) used data from program members in the hotel industry to establish program reward requirements that are projected to increase revenues by 9%.
Big Data, Analytics, and Health Care
The most obvious use for big data is for the diagnosis and treatment of diseases. However, this runs up against several issues including privacy, proprietary interests, and incompatibility of health care record-keeping systems (even within hospitals). Researchers are building AI algorithms that enable autonomous robots to perform routine tasks such as stitching wounds, thus freeing up surgeons for less routine procedures and improving patient outcomes. A major limitation of big data and analytics is that it requires a vast amount of data to “learn” in a training module and “test” in out-of-sample data. When things change rapidly, as they have in the COVID-19 pandemic, limited data are available for AI algorithms to learn from. In such situations, the solution lies in the cooperation among multiple disciplines (e.g., business, computer science, medicine, engineering) whose data from multiple sources can be combined and analyzed for potential patterns and solutions. It is particularly important for health care companies, researchers, and government agencies such as the Centers for Disease Control and Prevention to be aware when the data on which analytics and AI algorithms are trained no longer represent the world we live in. How much and how to use data from the “old” world is a key question. The pandemic and the corresponding business challenges firms face in applying AI provides an important case study about collecting and using (or at least weighting most heavily) the more recent and relevant data on testing, tracking, hygiene, and social distancing and combining them with health records, identifying those at high risk and using life cycle, lifestyle, and demographic information as well as AI to predict and prevent future outbreaks, and develop vaccines and treatments. In this regard, smartphones and wearables, which store a lot of information about their users (e.g., location history, contacts) may be trained using data and analytics—via apps that can sense various things such as voice and breath—to detect whether a person is infected or especially prone to be infected with the virus. Recent advances in the development of new medical products show the promise of cancer screening using AI to analyze tens of millions of sites in the human genome sequence and detect cancer well before symptoms show up. A by-product of such an application would hopefully be the reduction of false positives and false negatives and an increase in the accuracy of cancer screening results in terms of metrics such as recall ([True Positives]/[True Positives + False Negatives]) and precision ([True Positives]/[True Positives + False Positives]).
Google has now partnered with a few large hospital systems and health care providers and achieved the ability to view or analyze tens of millions of patient health records in approximately three-quarters of U.S. states (Copeland, Mattioli, and Evans 2020). In some instances, companies give Google access to personally identifiable health information without the knowledge of patients or doctors. The goal behind this is for Google to develop a searchable tool that can be used by doctors and nurses to provide better health care delivery for their patients. Such data may be used to train algorithms to detect cancer and other diseases and internal injuries. From a public policy standpoint, the key goal is the overall societal good of making people healthier and not firm profitability. Nonetheless, this has set off privacy concerns, with U.S. senators, governmental organizations, and privacy watchdogs questioning Google’s expansion into the electronic medical records arena, which constitutes perhaps the most closely guarded personal information. While patients commonly believe the Health Insurance Portability and Accountability Act (HIPAA) prevents doctors from sharing their data, the rules are written broadly enough for health care providers to share identifiable data with many business associates for quality assurance purposes. If medical providers can learn from data, so can machines. As long as the learning is used to pursue science and develop cures for people with complex illnesses, then society is overall better off with the sharing of such data. From a public policy perspective, it will be a win-win situation. Nonetheless, it is important that public policy makers ensure that there are regulations in place so that the patient-level electronic medical records data are not misused. There is also the important marketing problem of getting patients and the public to adhere to the findings (as demonstrated by, e.g., the reluctance to wear masks in response to COVID-19). In other words, gaining acceptance of the results of big data analysis is a major task.
According to the Centers for Disease Control and Prevention, there are 2,445 diseases in the world, with 638 symptoms, 510 signs a doctor may elicit, and 2,671 laboratory and diagnostic reports (https://www.cdc.gov/diseasesconditions/az/z.html). Going forward, big data and machine learning analytics can develop AI models to help physicians and patients in terms of linking symptoms with diseases, thus enhancing health care delivery efficiency and reducing costs. Biometric measurement tools with internet-connected devices such as the Apple Watch can detect signs of a person’s impending cardiac arrest early enough so that they can reach a hospital before total heart failure. However, such tools may lull individuals into a false sense of security if they rely completely on a device to signal problems or suggest positive actions. An important current application of data from devices is their use for contact tracing after potential exposure to COVID-19. This involves the interesting ethical issue of whether sharing this information should be mandatory or consent-based. It is important for public policy makers to more fully understand the social implications of such a transition toward digital ecosystems and how to balance human and artificial importance.
Another application in health care of analytics is the use of analytics in drug discovery, which involves using quantum computing to run simulations of protein interactions that are too complex for current computers to work through. Quantum computing is expected to solve standard encryption problems in hours that would take typical computers years to decode. Specifically, quantum computers are expected to perform mathematical operations in a fraction of time that would have taken a supercomputer take years to complete.
Marketing Tactics (Programs) and Big Data
A less prominent, but important, aspect is the use of marketing to steer customers/patients toward adopting and continuing/adhering to sound health practices and treatment regimens (i.e., personalized persuasion instead of just personalized medicine). For most commercial products, the risks of a nonoptimal purchase are mainly financial and inconvenience. While this applies to many medical conditions (e.g., minor acne, mild allergies), for others the risks are much greater. Individual data allow for more tailored messaging and programs to encourage healthy behaviors. Many of these pertain to exercise, sleep, and eating behaviors, which can be monitored unobtrusively. Here, encouraging/increasing perceived and self-efficacy seems relatively uncontroversial to the extent it makes individuals aware of problems they may have and procedures they can use to address them. What is less clear is who gets to decide what behaviors are good and bad and the extent to which “tricks” (nudges) should be used to change behavior. Put differently, knowing more about what can be done to change behavior does not resolve issues of whether we should focus on changing specific behaviors or how to do so. Relatedly, we need to know how changing one behavior (e.g., eating sensibly and regularly) affects others (e.g., engaging in risky physical activities to compensate for the loss of excitement and stimulation obtained from unhealthy eating).
An important marketing application that leverages big data and analytics and has public policy implications involves using AI-based predictive models to determine which individuals to provide, for example, loans (e.g., housing, automobile, home equity) and the terms (e.g., interest rate, down payment, time period) of those loans, rental housing, or even visas to enter a particular country. These predictive models are trained using past data. If there were discrimination in the data on the part of human beings, the same discrimination may be exhibited by the machine-learning-based AI models when they are used to make predictions. Although price discrimination may be generally legal in its implementation, public policy makers should pay close attention to determine whether machine-based discrimination, especially when there is little or no human supervision, is both legal and ethical. We believe that there should be human supervision prior to the implementation of predictive analytics-based decisions recommended by AI.
Customer Data and Privacy
Because rewards are calculated on the basis of transaction data, customers are motivated to self-identify when they make a large variety of purchases. The implications of the resulting big data and analytics on data privacy are apparent and palpable. Firms can collect mounds of useful customer data; for example, a retailer collects data on its customers online, offline, and via its mobile app. The completeness of customer-level data—both transactional and personal—intensifies the danger of infringing on customer privacy. The omnipresence of loyalty programs gathering data from members plays squarely to these concerns. It is important that firms use the data carefully so that consumers do not feel that their privacy is being violated. Even for customers who voluntarily provide their data, there is the potential for (and manifestation of) increased concerns about misuse of the information and loss of control over how it is collected and disseminated. The European Union is becoming more stringent in terms of how firms can collect and use data on their customers. It is likely that restrictions will be placed on facial recognition tools and disclosure requirements developed regarding which data may be used for AI development.
Recently, homes have been transformed into multipurpose spaces—for work, school, leisure, exercise, cooking, and sleeping—as more people sheltered, studied, and worked from home during the COVID-19 crisis. People in general have been sharing a lot of personal things (e.g., photos, birthdays), which puts individuals at risk partly because they may be clues to their passwords, because passwords are often based on hobbies and names of loved ones or pets. In general, consumers need to be more aware of how exposed their data are online and take privacy more seriously (or at least give it up more knowledgably).
Consumers have also become more comfortable with robot-enabled home technologies and virtual assistants. This will, of course, tend to reduce their privacy concerns as consumers become more focused on other things in their lives. Given this relaxation of privacy concerns, one big need for policy makers is to learn how to use big data and analytics to catch online predators by training analytics-based AI systems to learn how predators operate online (e.g., the types of words and patterns they use to lure innocent victims, their modus operandi for gaining the trust of their victims), recognizing that the methods that predators use evolve over time.
Another important application stems from the fact that cars collect information through sensors in brake pedals, seat belts, tire pressure, windshield wipers, wheel position, odometer, ignition, GPS, and so on. The navigation system in a car records its location every few minutes even when the system is not in use. On Google, it is easy to track the locations a person has visited. Much of this information may be used to help improve the car’s performance, refine features, identify quality problems, create more personalized offerings to the driver, and more. Again, it is important that consumers be given the opportunity to give explicit consent for their technology to collect and share such data.
Limitations of Big Data and Analytics
Big data and complex analyses are useful for many purposes related to social welfare. In the broad area of health, they can help with diagnoses, the identification of (more) individualized treatments, and the creation of programs to modify behavior. In terms of marketing, they can identify more effective ways to impact individual behavior. Where they fall short, however, is dealing with multiple objectives. For example, people are concerned with both longevity and quality of life. Quality of life, in turn, has multiple components, including physical strength and stamina, mental acuity, stress, and happiness, all of which are related to each other. How one best measures the softer elements is unclear. Even more unclear is how one weights/combines them with other concerns, such as economic welfare. In essence, there are an infinite number of combinations of outcomes that produce a Pareto optimal result. While big data/analytics can optimize a given objective function, they are not designed to choose the “correct” objective function. Of course, one can use a method such as conjoint analysis to assess trade-offs, but quantifying attributes such as “happiness” and “feeling energetic” is problematic. This same issue applies to the different aspects of climate and the environment as well: for example, how should we trade off carbon emissions, the health of various species, and food production to feed undernourished individuals? Establishing their relative weights and the functional form of their relations is even more difficult because, among other reasons, they change over time and depend on their current and past levels as well as how those levels compare with the levels of others (both in general and those with whom one is in close contact). Furthermore, data at the individual level are often limited, and thus researchers must use data on others to reach stable results, which makes the individual-level analysis less individual.
At the societal level, the COVID-19 crisis has made it salient that there is a trade-off between economic and physical health; related concerns such as mental health exist, as do issues of equality and access. While analytics and AI can crunch data, they do not collect/create it. The closest procedures we have for doing this involve analyzing unstructured (first text, now visual) data. It would be interesting to explore whether analyzing big data on individuals both helps identify health issues and provides a way to “nudge” people toward corrective actions. For example, the CORD-19 open research data set (https://www.semanticscholar.org/cord19) put together by the Semantic Scholar team at the Allen Institute for AI is a free resource of scholarly articles about the novel coronavirus for use by the global research community.
Another question is how much input to allow humans (in particular experts or those charged with making a decision) to have in the process. AI has an adoption problem. Put simply, people tend not to want to allow robots and algorithms to make their decisions. Part of this is self-interest related to job security while another part relates to a sense of loss of “agency”/autonomy. Although humans are more willing to give up control over objective tasks than subjective ones, they show a noticeable reluctance to totally relinquish control. In essence, these findings support conjectures made in 1990s that the routine, repetitive marketing decisions (e.g., promotions) would be automated by 2020. How to incorporate prior opinions/theory is also an interesting area to explore. While researchers have explored how to ex post combine model prescriptions and managerial insight, an interesting and seemingly superior approach involves incorporating human intuition/theory directly in the modeling process. This makes analysis essentially Bayesian rather than purely data/case based.
Footnotes
Special Issue Guest Coeditors
Brennan Davis, Dhruv Grewal, and Steve Hamilton
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
