Abstract
The production of large, shareable datasets is increasingly prioritized for a wide range of research purposes. In biomedicine, especially in the United States, calls to enhance representation of historically underrepresented populations in databases that integrate genomic, health history, demographic and lifestyle data have also increased in order to support the goals of precision medicine. Understanding the assumptions and values that shape the design of such datasets and the practices through which they are constructed are a pressing area of social inquiry. We examine how diversity is conceptualized in U.S. precision medicine research initiatives, specifically attending to how measures of diversity, including race, ethnicity, and medically underserved status, are constructed and harmonized to build commensurate datasets. In three case studies, we show how symbolic embrace of both diversity and harmonization efforts can compromise the utility of diversity data. Although big data and diverse population representation are heralded as the keys to unlocking the promises of precision medicine research, these cases reveal core tensions between what kinds of data are seen as central to ‘the science’ and which are marginalized.
The past two decades in biomedicine have been marked by the search for more individualized, precise therapeutic interventions. Since the mapping of the human genome at the end of the 20th century, the promise of leveraging genetic information for biomedical advances has captivated research institutions, industry, and researchers (Green et al., 2011, but see Evans et al., 2011). ‘Precision medicine’ has been heralded as the next generation of biomedicine that will produce targeted diagnostics and therapies for individuals and population groups. To achieve precision medicine goals, national and international health agencies, industry, and researchers increasingly prioritize the production of large, shareable datasets and biobanks (National Institutes of Health, 2015). However, having large datasets alone is insufficient to achieve the goals of precision medicine: Researchers contend that sufficient population representation is also requisite (Bentley et al., 2017; Hindorff et al., 2018). It is at this particular intersection, where the goals of ever-larger and more diverse datasets meet, that issues of defining, standardizing, harmonizing, or otherwise operationalizing ‘diversity’ are made visible and their downstream consequences become especially salient.
Biomedicine’s shortfalls in terms of broad representation are both persistent and well known. Medical research in the United States has long relied on data skewed toward white, middle-aged men, building normative standards based on the exclusion of women and racial and ethnic minorities (Epstein, 2007). This over-representation continues, decades after the National Institutes of Health (NIH) Revitalization Act of 1993, which institutionalized the inclusion of women and racial minorities in federally funded research (Hwang & Brawley, 2022). Recent estimates show a stark lack of diversity in the genomics context: as of 2016, 81% of genomic study participants were of ‘European ancestry’ (Popejoy & Fullerton, 2016). In response to criticisms, recent efforts focus on assembling datasets that reflect the diversity of the U.S. (All of Us Research Program, 2019; S. S. J. Lee et al., 2019).
Constructing large-scale datasets also requires harmonization and standardization processes to ensure that data is commensurate (Espeland & Stevens, 1998; Timmermans & Epstein, 2010). Either study measures must be standardized at the outset of data collection by appealing to pre-existing protocols and coordinating measurement across study sites, or existing datasets must be harmonized and repackaged into similar units. Harmonization—the process of making heterogeneous units compatible— is necessary in the absence of standardization; in science, harmonization enables data that was measured differentially and/or drawn from different sources to be combined for analysis. Harmonization thus becomes an important site for social inquiry because it involves decisions about making initially dissimilar entities comparable, exposing social assumptions and values in the process (Espeland & Stevens, 1998). Once harmonized, these assumptions and decisions are rendered invisible, with downstream consequences for the quality and utility of data produced (Desrosières, 2000).
In this article, we trace how genomic datasets are constructed in the context of twin mandates to produce harmonized databases and advance diversity in precision medicine research. We follow how ‘diversity’ measures—such as racial or ethnic identity and medically underserved status—are harmonized. Our cases are embedded in U.S.-based consortia that have been federally funded, in part, to increase the diversity of genomic databases.
Yet, even in these contexts, attention to diversity-related issues remain unprioritized, as both structural and individual constraints make diversity work marginal (see Berrey, 2015; Bowman Williams & Cox, 2022). We contend that harmonization was undervalued in part because investigators saw it as competing with scientific freedom and also, for diversity measures in particular, because researchers only partially attended to what they term ‘social’ factors (Ackerman et al., 2016). As we will show, harmonizing diversity measures was not seen as central to the science in these settings, or the study site specific research questions; rather its treatment was handled, as one researcher put it ‘on the side of everything else’, in terms of both time and effort and a more distant relationship to the research aims. This relegation has broader consequences for both our understanding of the persistent inequalities in health and biomedicine’s ability to intervene on them, and points to the need for greater accountability to promote serious attention to these issues.
Setting the stakes: Reckoning with standards, race, and big data
In this article, we bring together three literatures: standardization and harmonization, race in genomics, and the emergence of big data. Although there has been some recent research on the role of race in big data, especially around algorithmic fairness and annotation (e.g., Hutchinson et al., 2021), and some research on race and standards, relating to the history of the census and from whence the Office of Management and Budget (OMB) race categories emerged (e.g., Mora, 2014; Snipp, 2003), research to date has paid limited attention to either harmonization as a general process, or the harmonization of sociodemographic data more specifically. Given recent investment in ever-larger biomedical science consortia and data aggregation—and the ensuing harmonization that must occur for this data to be useful (e.g., Harris et al., 2012)—our understanding of harmonization processes and how they shape both the data and downstream knowledge produced is increasingly critical.
Standardization and harmonization
In its ideal form, standardization is a ‘process of constructing uniformities across time and space, through the generation of agreed-upon rules’ that seek to ‘render the world equivalent across cultures, time, and geography’ (Timmermans & Epstein, 2010, p. 71). Harmonization work is often done in concert with standards and potentially creates new ones: While developing harmonized measures, investigators routinely look to external standards and practices to inform their work (e.g., Firnkorn et al., 2015). However, harmonization itself has not been seen as a particular area of expertise; harmonization work in data science fields has been both highly technical and specific to the field in which it emerges, leading to a lack of general guidance around harmonization processes (Cheng et al., 2024). Harmonization also has been notably absent from discussions in the growing field of the ‘sociology of quantification’ (see Berman & Hirschman, 2018; Mennicken & Espeland, 2019).
Standardization is taken to be a bedrock value in science because it ensures that data collection is systematic and facilitates comparison across studies. However, in research settings, existing standards can be experienced as a constraint on autonomy and actively or passively resisted (C. Lee & Skrentny, 2010; Rosemann & Chaisinthop, 2016). Researchers may trust standards, but only to a point (Timmermans, 2015); in practice, competing standards can co-exist with researchers switching between them when it suits their purposes (Rosemann, 2014). This echoes research on categorical ambiguity, which highlights instances when a lack of clear definition and inconsistent terminology can be both organizationally and individually useful (Mora, 2014; Panofsky & Bliss, 2017). Similarly, sociological work on ‘decoupling’ suggests that organizations may adopt prevailing norms and standards publicly, to boost legitimacy, but may be unwilling or unable to comply in practice (Meyer & Rowan, 1977; for a recent review, see Brunsson et al., 2012).
Overall, this literature suggests that researchers draw on existing standards when required but are likely to passively resist or only superficially comply when they are perceived as an unnecessarily constraining. In our empirical setting, however, one might expect to see harmonization of diversity measures handled differently, both because of longstanding debates about the use of race in genomics and because of the perceived value and utility of big data in precision medicine research.
Race and genomics
Biomedical research in general, and genomics in particular, has drawn considerable scrutiny regarding its use and misuse of the concept of ‘race’. Historical presumptions of presumed inferiority and innate biological propensities, well known cases of unethical medical experimentation, along with past eugenic policies of sterilization and elimination of people deemed ‘unfit,’ provide grounds for concern whenever ideas about race and medicine meet (S. S. J. Lee et al., 2019; Roberts, 2011). Calls for greater clarity about the use of racial and ethnic categories in genomic studies have only increased in recent years (Mauro et al., 2022) and inspired the recent convening of an expert committee for the National Academies of Sciences (NASEM, 2023). This context of heightened scrutiny has prompted widespread concern amongst researchers about how to get race ‘right’.
Yet the answer to whether and how race should be accounted for in biomedical research remains highly debated. The debate is often presented as one between ‘essentialists’, who perceive that important biological or genetic variation is associated with the major ‘races’ as they have been understood historically, and ‘constructivists’, who see race primarily as a form of social and political status based more in persistent beliefs about biological difference than in the reality of human diversity (Morning, 2011). However, the debate can also be cast as between researchers who point to observed differences in responses to treatment and rates of disease and attribute them to race, and those who see the same variation and point instead to racism (Yudell et al., 2016, 2020). Although few scholars stake out extreme positions at either pole, the tension between them remains: Should the field be trying to ‘move beyond’ race as a population descriptor, as the National Human Genome Research Institute recently put it in one of ten ‘bold predictions’ for human genomics (Green et al., 2020)? Or should the field be trying to take race—and racism—more seriously?
A major barrier to being able to simply ‘move beyond’ race is how race-based practices are embedded throughout the research life course from participant recruitment to the translation of results to clinical practice (Bentz et al., 2024). Since the 1993 Revitalization Act, researchers who receive U.S. federal funding have had to provide demographic breakdowns of the participants they enroll to demonstrate inclusion of women and racial minorities. The policy did not stipulate how research analyses should be carried out, and federal guidelines for racial data collection have always been framed as a minimum standard (meaning that researchers could always collect more, or more detailed information, as long as they could also report data in terms of the official scheme). However, in the years since, genomics researchers have become habituated to stratifying their samples and analyzing their data using U.S. Census-style racial categories (Shim et al., 2014). When samples are sent to labs to be genotyped, it is common to label them in racial or ethnic terms, which are later translated into the racialized categories used in databases (Popejoy et al., 2020). It is also common for researchers to run their assays with race-specific parameters, to train their models on the racialized populations represented in reference databases such as HapMap, and to build population stratification based on race and ethnicity into their analyses (Fujimura & Rajagopalan, 2011; Rajagopalan & Fujimura, 2018). Recent efforts in genomics have questioned the utility of some of these practices (e.g., Lewis et al., 2022), and made attempts to de-racialize population labeling (Wojcik et al., 2019), but the overall taken-for-grantedness of using racial categories in genomics research remains.
After the wave of racial justice protests in the summer 2020, the perceived pressure to account for race and racism in research only increased. Corporations and universities raced to publicly condemn police brutality and anti-Black racism. A new wave of academic journal guidelines began to appear, with specific stipulations for how authors should discuss the concept of race in their manuscripts (Brothers et al., 2021; Flanagin et al., 2021; Nature Human Behavior, 2022). Dozens of think pieces were published (e.g., Chokshi et al., 2022; Tsai, 2021), along with retrospective analyses highlighting continued shortcomings (Byeon et al., 2021; Martinez et al., 2023). Of course, these are only the most recent articulations of similar concerns raised by scholars for decades, from the lack of attention to how race is being measured and why it is being collected (Kaplan & Bennett, 2003), to the uncritical use of race as a ‘control’ in statistical analyses (LaVeist, 1994), and the continued inequitable representation in genomic databases (Sirugo et al., 2019). The harmonization work we observed took place against the backdrop of these long-standing critiques, debates, and newly heightened pressures.
Big data in biomedical research
The introduction of big data in biomedical research promises to make medicine more precise, based on individual health needs rather than generalizations made about groups or populations. For example, efforts to mine electronic health records and integrate disparate record systems enable an ever-expanding amount of data on patients to be captured (Cruz, 2022). The ability to aggregate, link, analyze, visualize, and make available these larger datasets is key to precision efforts in public health and medicine (Kenney & Mamo, 2020).
Scholars have underscored diversity and equity concerns that arise with big data initiatives—specifically around what data are collected, on whom, who has access, and who benefits and profits from their amalgamation. Of note for this article, health-related data are constituted into clinical algorithms and decision-making tools that mirror and even amplify the biases of the data going into them (Eubanks, 2018; Obermeyer et al., 2019; Vyas et al., 2020). Calls for more attention to transparency in data creation and algorithmic bias has raised the stakes on the importance of how measures of diversity are managed methodologically (Hutchinson et al., 2021). We add to these conversations empirically by examining how diversity measures are harmonized in precision medicine research.
Data and methods
Our data come from a multi-sited qualitative project that investigates how commitments to diversity and inclusion shape research practices in genomic medicine, including how working definitions of diversity are interpreted and operationalized. We followed five studies across three NIH-funded consortia. We selected consortia at different points in the research process to investigate how diversity and inclusion were operationalized across the research life course. This allowed us to observe how decisions about diversity measures, for instance, were discussed prior to data collection, during data collection, and during analysis. Individual study sites were selected to represent geographic diversity within the U.S. and because they made commitments to involve diverse participant populations. In this sense, they are potential best-case-scenarios for taking ‘diversity work’ seriously.
These consortia and the studies that comprised them were focused on collecting and/or analyzing genomic, clinical, behavioral, and other health-related data among study participants, especially populations underrepresented in biomedical research. Generally speaking, the studies were interested in assessing the value added of genomics integration in biomedical and health-related studies. Study teams were comprised of interdisciplinary researchers, including geneticists, clinicians, data scientists, and ethical, legal, and social implications (ELSI) researchers. Some teams also hired consultants who specialized in research with underrepresented groups; such consultants were typically involved in the creation of recruitment and community engagement strategies (see also Epstein, 2008).
In this article, we focus on data from two of the three consortia. One consortium, which we call Existing Cohort Consortium (ECC) was an initiative to leverage longitudinal epidemiological datasets from established studies. ECC studies were selected to augment their data with genomic sequencing, with the explicit goal of increasing diversity of sequencing databases. The second consortium, which we call New Cohort Consortium (NCC), was focused on demonstrating the clinical utility of genomics and aimed for ‘enhanced’ diversity in enrollment, defined in the request for applications (RFA) as a particular threshold of participants qualifying as one of the following: racial or ethnic minority populations, underserved populations, or populations that experience poor medical outcomes.
In federally funded biomedical research, many initiatives are organized as research consortia, in which studies are carried out at multiple research centers and scientific teams are increasingly large. In some cases, the same study protocol is carried out at all research sites; in others there is considerable flexibility across sites. ECC and NCC are examples of the latter, where the nature of the research is somewhat similar but each study site involves distinct populations, research protocols, and study designs. Within the consortium framework, sites share data through centralized data coordination hubs. ECC and NCC differ, however, in terms of the timing of their harmonization efforts; all harmonization of ECC data that we observed was post hoc while NCC offered the potential to standardize practices at the outset rather than having to harmonize after data collection had already occurred.
Both consortia used a working group model, in which interested researchers organize around a specific topic area and meet virtually at a given frequency. These included working groups for genomics, ethical, legal, and social implications (ELSI), specific diseases (e.g., diabetes), harmonization of measures, social determinants of health, and clinical utility. Groups usually had a chair or co-chairs who led meetings, interfaced with NIH program officers regarding the group’s progress, and handled the administration of the group for a set period (e.g., one year). Both consortia also held annual or semi-annual meetings, in which the full consortium assembled for multiple days.
Data and analysis
Data, consisting of observations, interviews, and documents, were collected over a three-year period, between January 2019 and December 2021. Institutional review board approval was obtained from University of California, San Francisco, and Columbia University. For each site, we conducted observations during multiple time points over the period of data collection. Ethnographic observation (totaling over 450 hours) included in-person and virtual observations of: working group meetings, individual study team meetings, annual consortium-wide scientific gatherings, and professional society meetings. Events were attended by one or more members of the research team. Detailed fieldnotes were taken during observations and, where possible, consortia recordings were obtained and transcribed.
In-depth interviews (n = 125) were conducted with 102 purposively recruited study investigators, research staff, program officers, and study participants involved in our study sites. A subset of interviewees (n = 23) were interviewed twice. Interviews followed a semi-structured, open-ended format and were conducted via Zoom or in person. Interview topics included consortium and study site activities, team organization and governance, community and stakeholder engagement, diversity, recruitment, enrollment, harmonization, and broader reflections. Interviewees provided written consent and were offered a $50 gift card for each completed interview. All interviews were recorded and professionally transcribed verbatim.
Documents (n = 76) included RFAs, grant proposal excerpts, study instruments (e.g., demographic questionnaires, consent forms, recruitment materials), and published materials from sites and consortium. Publicly available documents were accessed by members of the research team. Site-specific instruments were requested from study leadership at sites or the consortium level.
Interview transcripts, observation fieldnotes, and documents were uploaded to Dedoose, a qualitative analysis software, for coding. We developed a codebook using a modified constructivist grounded theory approach (Charmaz, 2014). First, the research team generated a priori codes based on study aims, the existing literature, and issues that had arisen repeatedly throughout data collection. Then, team members coded a sample of documents, interview transcripts, and fieldnotes to test and adapt the initial codes and to inductively generate additional codes. To maximize intra- and inter-coder reliability, our team periodically jointly coded the same data and discrepancies were discussed and reconciled. This article draws on data coded for mentions of diversity measures, standardization, and harmonization. Analytic memos were written throughout data collection and analysis. In-process analysis informed the design of interview guides for subsequent interviews to follow up on specific aspects of data harmonization discussed in earlier interviews and fieldwork.
Interviewees often did not recall the nuances of specific measures and harmonization processes. In these instances, as well as where interviewees offered conflicting accounts of activities, we triangulated with our observation records as well as requested study materials from consortium and specific sites.
In what follows, we offer three cases from our empirical data that illustrate the efforts investigators undertook to harmonize diversity measures like race, ethnicity, and medically underserved status. We show that because these efforts were lower priority, constrained for time, and under-resourced, the resulting harmonized measures provided only surface-level fidelity to the social differences they sought to represent, and fell short of the analytic heft needed to be fully integrated into analyses. We conclude by discussing what these cases—potentially best-case scenarios because of their commitment to recruiting and enrolling diverse study participants—show us about the consequences of how contemporary precision medicine initiatives are harmonizing diversity data. Given the timeline of our own study in which data collection ended before sites fully completed their analyses or published results, the full extent of the consequences of harmonization in these consortia remain to be seen.
Case 1: Kicking the can down the line
For NCC, ‘enhanced diversity’ in participant enrollment was quickly narrowed to race and ethnicity. Although the RFA laid out multiple criteria by which sites could meet their diversity goals, investigators told us that they designed their studies with racial and ethnic diversity specifically in mind. One investigator explained: [Our] team was definitely under the impression that … funding was going to be driven in large part by diversity with respect to underrepresented minority groups from a race and ethnicity perspective. … Representation of rural southerners was something that we as a team saw as a strength of the proposal. But on top of all of that was, ‘Yes, we do think this is really scientifically interesting and valuable, and we also think it fits within the strict definition from the RFA, but we better have strong representation of underrepresented groups from a race and ethnicity perspective.’
This investigator and their team perceived race and ethnicity would be the most important definition of diversity for funders, even when other demographic characteristics such as socioeconomic status, ability, sex, and geographic diversity, were relevant to study goals. This demonstrates both the power and the limits of stressing diversity in the RFA (S. S. J. Lee et al., 2022), and is emblematic of sociological research that has shown how diversity tends to be interpreted narrowly as referring to racial and ethnic minorities (Berrey, 2015). Investigators took the requirement seriously but saw the task as more about getting ‘diverse’ people in the proverbial door than about how to record data and conduct analysis in a way consciously aimed to leverage that diversity to make scientific advances.
The consortia’s harmonized variable was often pointed to in working group meetings as a settled, stable object. However, when we asked researchers about this process, we were struck by their accounts: Many could not recall how the process had gone, or what had been decided. Even those people described as intimately involved either could not or would not discuss with us how the harmonization had been carried out. Many directed us to ask others or stated with considerable hesitation that they ‘thought they’d settled on [official Census-style categories],’ but that they would need to check. Several investigators expressed frustration, including one who lamented: ‘[E]very survey asks race, isn’t there some kind of standard question that we can just use? No, there wasn’t.’ Overall, interviewees seemed to remember it as an arduous process, saying things like: ‘[I]t was really a year of discussion.’
Maintaining site-level flexibility
Harmonization efforts began in the first year of NCC’s funding period. A ‘Measures Working Group’ (MWG) was tasked with developing a consortium-wide set of survey measures that all sites would use in their data collection. Harmonized measures were designated as ‘required’, ‘recommended’, or ‘optional’ for sites to collect, and many were intended to complement additional measures that individual sites were collecting. After designating race and ethnicity as required, MWG passed the task of harmonizing it to another group, one they deemed to have the relevant expertise: the Ethics and Diversity Working Group (EDWG), which was comprised of ELSI researchers. When we asked investigators why EDWG was assigned this task, one put it simply: ‘Honestly, I don’t know. It just came from on high that this was something we [EDWG] had to figure out.’
Within EDWG, months of debate ensued. Several sites wanted more granular measures and others felt they could not hold off on beginning data collection while harmonization at the consortium level was sorted out. One investigator described the challenge: Our studies just had very different goals, and [one] site in particular wanted extremely detailed information about race and ethnicity. For our site, we just are high level five categories [referring to OMB 1997 categories], that’s good enough. I think just trying to meet everyone’s needs with a single question wouldn’t really have worked. [It] was one of those cases where [that site] collects extremely detailed data, that they then collapsed into a small number of categories for the harmonized measure.
Harmonizing in this way seemed to satisfy many sites because of their heterogenous goals with respect to race and ethnicity, specifically. Another investigator explained, Sites could choose at which level they wanted to ask the question. So there may be the possibility across sites to look at them more granularly for those who asked the more granular [responses]. But then the benefit of having a higher-level category is, first of all, simplification. If you’re going to do any sort of analysis, it’s hard to look at 16 variations of Asian. Your sample size just gets way too small. And so being able to roll up can provide some sort of, at least crude way to look at differences across groups.
The decision to allow for flexibility across studies within the consortium in terms of the number of response options and their level of detail addressed one set of concerns. Other issues, such as from whom such data should be collected and which categories or measures (e.g., skin color, known ancestry) would be of most value given study aims were raised in group discussions but left largely unresolved. Some site investigators took issue with the federal standard for racial categories, opting to collect their data through open-ended self-identification and thought this would be a good approach for the consortium. When reflecting on the harmonization process, one investigator said: We did bring them to the consortium. We said, ‘Hey, here’s how we did it and here’s how we try to allow for self-identity … Not just give them some categories.’ But there wasn’t a huge amount of traction. … I felt a little bit like we had to ram through our [race and ethnicity] measures [in the consortium] faster than I felt comfortable with.
Along with others, this investigator was disconcerted with the consortium-wide harmonized measure, but ultimately recognized that they needed to move on for sake of time. At this point in the study life course, the full range of what sites decided to do was deemed acceptable: Site choices were constrained only insofar as they had to figure out how to ‘roll up’ the data they collected into the minimum categories designated for eventual consortium-wide analyses. The result was the appearance of agreement about racial and ethnic categorization at the consortium level that papered over wide-ranging site-specific practices.
Harmonizing across study stages
At the same time that NCC was harmonizing a race measure that could balance the aims of individual study sites with those of the broader consortium, sites needed to adhere to enrollment reporting requirements for the NIH. To ensure compliance with the 1993 Revitalization Act, quarterly reports are submitted using NIH-provided templates modeled after federal standards for racial and ethnic data collection put forth by the OMB. The required report response categories are shown in the first row of Table 1.
Comparison of ‘race/ethnicity’ measures used for enrollment and analysis.
Funders viewed quarterly reports as serving a regulatory purpose, rather than a scientific one. As one explained, The harmonized measures was the data needed to be in one place so that people could analyze them together. And until they were in that one place, NIH didn’t really have a way of keeping track [of study enrollments]. So we did have, sort of, two parallel ways of tracking diversity, as you mentioned, race/ethnicity, specifically. … We tried to make it clear to the investigators that scientifically they could basically, not do whatever they wanted, but they had much more flexibility than what we wanted to see for the NIH, kind of OMB reporting requirements. I think it did confuse people. The fact that we kept insisting on these QR [quarterly report] templates, and then did the harmonized measures a different way. But I think the harmonized measures were really intended to serve the science and the research questions.
This distinction between ‘the science’ and administrative requirements like the quarterly report was frequently underscored in our interviews, though what exactly was needed for ‘the science’ was never specified. Funders explained their broad RFA language was purposeful; they deferred to investigators to define what was needed for their scientific aims (S. S. J. Lee et al., 2022). Yet the unintended consequence of this vagueness was that it perpetuated ambiguity around expectations for not only operationalizing diversity but also bringing it to bear in analyses.
Although the NIH-provided categories are intended for the purposes of tracking study enrollment and are not required for use in data analysis, previous research has shown the unintended consequence of mandated enrollment reporting: that investigators come to treat OMB race categories as a default (Epstein, 2007; Shim et al., 2014). The operationalization of race and ethnicity in this consortium further supports these findings. The harmonized measure that the consortium ultimately adopted is substantially similar to the OMB standards and quarterly report template, despite investigators’ concerns about existing measures. The primary difference is the consortium measure involved incorporating the category ‘Hispanic/Latino(a)’ into a combined list of racial and ethnic categories rather than treating Hispanic origin as a separate question (see comparison in Table 1).
The similarity between the two measurement approaches is striking. On one hand, multiple investigators expressed dissatisfaction with standard measures, the consortium had the opportunity to depart from existing standards, and funders believed they had encouraged researchers to do so. On the other hand, investigators were more interested in maintaining site flexibility so they could fulfill their scientific aims related to integrating genomics into clinical care, which were seen as separate from funder mandates for diverse participant enrollment. Moreover, as genomics and ethics researchers, they did not have the expertise to forge a different path. Cross-consortium analyses were not an immediate priority and, as such, this solution was considered good enough at the time.
Reckoning with race and racism
The treatment of the harmonized race variable as an effectively settled object changed after the summer of 2020, amidst the perception of renewed pressure to ‘get race right.’ Following public protest about police violence and the public murder of Black Americans, a provocative article called out biomedical researchers for their failure to ‘interrogate racism as a critical driver of racial health inequities,’ and for operationalizing race in ways that reproduced racism (Boyd et al., 2020). Many biomedical journals responded with additional critical perspectives as well as guidelines for the reporting and use of race in research (e.g., Adkins-Jackson et al., 2022; Brothers et al., 2021; Cell Editorial Team, 2020; Khazanchi et al., 2020).
The call for new publication standards reverberated across the health sciences, and its impact was felt in each of the consortia we observed. In the NCC, researchers revisited how they were reporting race in consortium-wide studies and how individual sites were handling race in their analyses. In an EDWG meeting, attendees deliberated whether the research they were doing was implicated in reproducing racism. In meetings, they spoke candidly, recognizing that they not only uncritically treated ‘white’ as the reference category in past analyses but also failed to recognize ‘that their findings were racialized’. Others noted how sometimes they used softer terms like ‘underrepresented’, and did not explicitly acknowledge racism in their publications. One investigator posed the question to the group ‘How do we ensure our manuscripts do not increase harm?’ Yet, given where they were in their projects—projects that had not been designed to examine racism—the researchers were unsure how to best respond.
At a consortium-wide meeting, the EDWG and MWG groups held a joint event to discuss forthcoming guidelines for a leading journal in the field, and their implications for planned analyses. The discussion took place during a breakout session, scheduled opposite another session involving the working group on genomic sequencing, meaning that it was not possible to attend both sessions. This scheduling decision was emblematic of the obvious but generally unspoken split between ‘the science’ and the ‘diversity work’ and how it was taken up by different people, with different levels of power, in the consortium (see also, Ahmed, 2012). 1
During the joint meeting, one of the co-chairs displayed a slide with the following points to direct discussion about NCC’s response:
• The categories are not the problem.
• Trust is not the problem.
• Racism is the problem.
• How categories are used and the trustworthiness of the system are the problem.
• What can we do [in NCC]?
Then, an editor of a prominent journal in the field presented forthcoming best practices. They included, among others, that (1) race should only be used only as a sociopolitical category; (2) genetic ancestry should not be conflated with ‘sociopolitical race’, (3) variables related to race, as with all scientific variables, should be described in detail; (4) given underrepresentation of ‘Black, Latinx, Asian, and other non-white populations’ in genomics research, editors and reviewers should ‘prioritize manuscripts with strong representation of these groups, even when findings replicate earlier findings in white populations;’ and (5) avoid structuring figures, tables, and findings in a way that makes white populations the ‘normal’ population against which other populations are compared.
This presentation sparked a lively conversation about whose responsibility—funders? investigators? journals?—it is to ensure that best practices are being followed, as well as whether and how the consortium was following these guidelines. And while the co-chair had explicitly noted that ‘the categories are not the problem,’ the ensuing conversation quickly turned to just that: whether the consortium’s use of OMB categories had been the right choice:
This exchange highlights four points of tension: (1) the desire to revisit previous decisions given the sociopolitical environment they were now in and its ramifications for publishing; (2) confusion remained about how they decided to harmonize the race/ethnicity variable and why; (3) the development of their harmonized measure did not resolve ongoing disagreement about the best way to measure this variable among investigators; and (4) a lack of clarity on the harmonized measure and the quarterly report measure and their differences.
Following this exchange, another investigator jumped in to ask the journal editor who had just presented on best practices, ‘What should we do?’ The editor responded half in jest, saying, ‘I just wrote to you in chat asking not to ask that question. Just kidding. You measured it the way you did, and ought to report it that way.’ A senior investigator put in plainly toward the end of the session: ‘I think this reflects [how] early on we struggled with this, we came up with something, and also kicked the can down the line. And now we are really dealing with this and have to grapple with it.’
In the moment of racial reckoning, it was clear that the consortium’s own practices—dealing with race and ethnicity measures ‘on the side of everything else’, as opposed to more careful consideration up front about how such measures might be deployed in service of consortium-wide science—were part of the problem. As researchers became concerned about publishing their findings in this climate, they realized the consequences of their earlier actions. The consortium had settled on using OMB categories and a ‘roll up’ approach to reporting race and ethnicity data that were collected in varying ways across study sites. This allowed them to appeal to external standards on a subject they felt ill-equipped to address and maintained flexibility across the divergent practices of individual study sites. It also provided something that would look, at least superficially, like a harmonized variable. In doing so, as the senior investigator noted, they were really just ‘kicking the can down the line’, so that they could move their projects forward. The decision was deemed good enough at the time but came under renewed scrutiny amidst a changing sociopolitical environment. As the group tried to get race ‘right’ in this new moment, they were constrained by their earlier decisions, ones made in an environment in which they had been underprioritized.
Case 2: Reframing shortcomings as strengths
Our second case follows efforts to harmonize another diversity measure, an indicator of medically underserved status (MUS), which unfolded over several months in later years of our observation period. As in the first case, here we observed the creation of a harmonized measure that does not impinge on individual study autonomy or priorities. Also similar were disagreements about existing standard measures, tensions around who, if anyone—in the consortium—has requisite expertise, and an overall sense that these discussions are largely disconnected from the primary genomic work of the consortium.
About a year into consortium work, an investigator in the EDWG asked how sites measured MUS given that it appeared in the RFA definition of diversity but had not been explicitly addressed in harmonization efforts up to that point. The group, in consultation with NIH program officers, decided to add it to their agenda and asked sites to provide details about their MUS-related measures, which prompted concerns among investigators about being asked to collect more data. A program officer tried to assuage these fears in a working group meeting: Just to underscore, we’re not interested in making people start over in the way they are counting underserved for what the grants stated, not trying to put new burdens on the sites, not trying to put any site at a disadvantage. Instead, we’re trying to create a whole greater than each part, to enhance the value of the data.
The program officer attempted to assure others that reporting requirements for participant enrollment would not change, instead hoping a harmonized approach would provide added value for the consortium.
Debating standard measures
After canvassing the study sites, EDWG discovered there were no consistent measurement criteria for MUS. Sites had collected as many as six different measures from individual insurance status to the socioeconomic profile of one’s zipcode, but few of them were collected across all sites. Some study sites had looked to existing standards for guidance but found such measures wanting. An investigator explained, We went around and around and around and there were, I think the NIH has certain categories that they suggest such as the HRSA [Health Resources and Services Administration] measures. There are neighborhoods that they measure how far you are from a primary care provider, but [another investigator and I] mapped our own where we live and we found that we’re both in underserved areas. He was mapping it [from his office] in a hospital, and then I live in [major city], so that didn’t make sense to us, but we did end up using some of those measures because we just didn’t know what to do. We were at a loss.
While investigators did not feel that they had the expertise to decide how to measure MUS on their own, they also had qualms about whether existing standard measures—which tend to focus on geography and insurance status—were a valid substitute, given the seeming nonsensicality of being inside a major medical center that was in an ‘underserved’ area. Moreover, as the following exchange at a working group meeting shows, the various measures that sites were using could capture some aspect of the lack of access to healthcare resources, but was not specific nor clear enough to say how this lack of access or what kinds of resources specifically were most significant. 3
In this exchange, the researchers navigate a series of tensions: (1) recognition of the difficulty of pivoting after all the study sites had been recruiting participants for some time and measuring MUS differently, (2) concern that local differences might lead to over- or under-estimating the outcome of interest, (3) questioning whether those present in the group have the expertise to capture MUS in a value-added way, and (4) realizing that different measures might be necessary for different research questions. At this point, the conversation turned toward what was manageable within the consortium work:
As conversations about these measures continued, there seemed to be a desire to rein in the scope, ensuring that any steps forward would be something that ‘as a consortium we are able to do’. Implicit here is that there is not the time nor resources to do what might be best practice, but rather that they needed to move forward with the kinds of data they were collecting, and to agree on something they deemed manageable given the work they had been funded to do. But the deeper concern here, regarding ‘reinventing sociology’, highlighted a consistent tension in the group: that those being asked to take up this work did not have the requisite expertise to do so. One of the investigators explained how he felt that it was important work, and was glad he had been involved in it, but: I actually said this numerous times while we were working on this, I’m not an expert on this. We have anthropologists, sociologists, social scientists of various types, and they can theoretically walk circles around me in terms of understanding what’s going on here, how to describe it and what the theoretical basis is.
Although this investigator highlighted his lack of expertise on this topic, he was frequently referred to by others as the resident expert, perhaps primarily because he was leading the effort.
Constructing a conceptual compromise
Over the following several months, as the EDWG discussed how to harmonize MUS, efforts pivoted toward developing a conceptual model. Group members debated how geographic and race variables should be combined and what other dimensions contributed to being medically underserved. Ultimately, language barriers, income, insurance status, place of residence, and race and ethnicity were chosen as component measures. However, since not all sites had measured every component, the framework allowed for heterogenous definitions to be captured within one measure, masking that sites had taken different approaches and may have had missing data on one or more dimensions altogether. Further, because not every site had collected data on every component measure, their presence or absence could not be summed into an ordinal scale.
This led to a key compromise: If any of the barriers that contributed to the risk of being underserved were present for a research participant, that individual would be coded as MUS; that is, to achieve harmonization across sites, a concept that the group agreed was multifactorial and non-dichotomous was ultimately rendered in the database as a binary variable. One of the investigators most closely involved explained the final product this way: [The framework] is not designed to say there is one kind of underserved and this is what it is and everybody’s using that [definition]. It’s really meant to show underserved as kind of a collection of closely related ideas, that would include experiencing barriers to access, or maybe even at risk for experiencing barriers to access.
The MUS measure was ultimately operationalized as a simple threshold: It boiled down to whether any of the focal factors were deemed present at all for a given participant. In this way, it performed the harmonization work necessary to bring the disparate measures of underserved status together. But in so doing it lost its specificity, flattened local difference, and conflated and obscured the mechanisms through which individuals and groups become medically underserved.
Despite these limitations, the framework was operationalized into a measure to be included in the centralized data hub for the consortium. Moreover, because its development was also documented in a manuscript for publication, other investigators pointed to the MUS framework as a critical demonstration and part of the consortium’s attention to diversity. In this sense, the MUS case departs from NCC’s harmonized race measure. In both cases, we observed how researchers without expertise on these topics become the ones leading harmonization efforts. In both cases there also was recognition that the work being done was partial, a proxy at best, but that they needed to make decisions and move on with the work of the projects given time and funding constraints. However, in the MUS case, harmonization work was treated as something important and successful—a centralized measure used across the consortium as well as a scientific product. Thus, while the limitations of the measure constrained its substantive contribution to ‘the science’, the underserved framework became a demonstration of the consortium’s concern for healthcare access—a dimension of diversity and focus for the consortium from the get-go. The harmonization work invested in the measure served important performative functions for the consortium and funders.
Case 3: Inconsistency between and within diversity objects
Our final case considers ECC harmonization activities, in which preexisting studies were brought together to perform various ‘-omics’ analyses. This consortium’s structure contrasts with our previous two cases; here, all harmonization was post-hoc, as the studies were legacy epidemiological cohorts. Its data center was tasked with unifying measures that had been collected at different times using widely varying methods. Staff focused on harmonizing genetic and biomedical data across the consortium. We were told that the only ‘social variable’ harmonized was race, again highlighting that this aspect of participant diversity was the one that mattered most. However, constructing its harmonized variable was not the only time ECC was forced to grapple with race and ethnicity.
We compare ECC’s approach to harmonizing race and ethnicity across three diversity objects: a public-facing pie chart, the harmonized variable, and a set of research guidelines for how to treat the concepts of race and ancestry. Each object offers inconsistent information about what race, ethnicity, and ancestry are in the ECC as well as how they should be used. They underscore how substantial collective effort can be put into harmonization-related tasks without meaningfully operationalizing race, ethnicity or ancestry for research purposes.
Diversity on display
The ‘pie chart’ had become a public symbol of this consortium before we began our observations. It had a prominent place on ECC’s website and was used often in presentations that featured the consortium’s data, intended as an accessible visual summary of the diversity of ECC data. One investigator explained, Our NIH staff always wants to see an updated version of it. … For our audience of funders … I think they really like that. When you look at the pie chart you can say, ‘Oh, [ECC] is only 40% European ancestry. Meaning, look at all these other groups we have so well-represented. I mean, if that pie chart was like the pie charts people show about how bad genetics research is for not being diverse, if it was like one of those 98% European pie charts, I’m sure they wouldn’t want to show it. But I think that it’s probably the best single artifact of diversity in [ECC].
Over time, as the consortium expanded, the wedges of the pie changed in both size and label. The categories featured at the end of our observation period are shown in the first row of Table 2; one researcher explained their evolution:
Two ways of harmonizing race in the same consortium.
Early on when we had just a small number of studies … we were able to be more specific. Right now, [several studies are] in the ‘Other’ wedge, which is also not great, to other them in that way … there’s a tension. You don’t want to group them inappropriately. You don’t want them to be not captured either or labeled in this nebulous other wedge. It’s hard to try and honor all those different pieces.
Incorporating ‘all those different pieces’ without creating individual wedges for each is the essence of harmonization work, which proved challenging for the data center for two key reasons: Not all studies had obtained self-identification data on race, ethnicity or ancestry from their participants and those that had did not collect the data using the same categories or methods.
Given this built-in flexibility, researchers we spoke with expressed discomfort with categorizing all study populations into the pie chart’s ‘ancestry’ and ‘ethnicity’ categories, particularly for studies carried out in non-US settings where understandings of race or ethnicity differed. One researcher explained their unease: When we first started getting studies from Brazil, I felt very silly asking them, ‘Can you please give us a breakdown? Here are the categories we’ve been using’ and they’re the U.S. Census categories. And they’re like, ‘Well, these don’t really make sense for us, but I guess you can put us under Hispanic/Latino. We don’t really like that, but whatever.’ It was interesting to try and fit people into those, I guess, wedges in the pie, when it didn’t feel ideal. But I guess … having some information is better than none. And so, we recognized it wasn’t perfect.
Here, the researcher felt uncomfortable using categories that did not align with investigators’ understanding of their study populations, but describes the tradeoff as ‘having some information is better than none’. Another compromise was for the ECC to categorize participants of studies joining the ECC who did not fit the chosen classification scheme to be included in the ‘Other/Unknown’ wedge.
On a closer look, the pie chart is also emblematic of slippages between race and genetic ancestry, a slippage that researchers themselves acknowledged (a point to which we return below). Under the pie chart, a note reads: ‘ancestry’ does not come from inferring genetic ancestry, but by using a ‘combination of self-identified or ascriptive race/ethnicity categories, study inclusion criteria, or other demographic information provided by study investigators’.
Thus, as we saw in our previous two cases, deference to the preferences and practices of individual studies was given primacy over creating a truly harmonized image of the racial, ethnic and/or ancestral diversity of study participants. The pie chart also was used externally to signal the value of this consortium in fulfilling its mission to diversify genomics datasets. Its flexible and evolving categorization scheme, regarded as inaccurate by some and insufficient by others, is tolerated in the service of this larger goal.
Distilling diversity for analysis
The same challenges that the data center encountered in maintaining the pie chart had to be confronted again when the time came to create a harmonized variable to be included in the consortium-wide database. As before, the work was primarily taken on by data center staff, who periodically consulted with principal investigators from the individual studies. In other words, although the harmonized data was designed to be used for consortium-wide analyses, the most senior investigators who would be leading those analyses contributed little to this aspect of the decision-making. Also, as before, the process was continually modified as new studies were added. One analyst explained why they felt this made harmonizing race uniquely challenging, It’s very different because race is such a societal or cultural concept or construct. You can’t really combine, it’s really hard to have a race variable that is universal and spans all the studies and topics. So we discovered this as more and more studies came in and decided to modify our harmonization. So we [data center staff] changed to having a race U.S. variable where we thought very specifically about which studies are appropriate to include in it and which studies should be excluded.
Continued unease about lumping U.S. and non-U.S. studies together using the same ethnic/ancestry categories drove the staff away from simply re-creating the pie chart classification scheme for the harmonized variable. It did not, however, lead them to abandon the use of OMB racial categories for the purposes of harmonization (see Table 2). Data center staff decided that only US-based studies reporting race, or recruiting based on race, would be included in the harmonized measure. This meant that this measure would be missing from all non-U.S. studies.
Even for the U.S.-based data that were harmonized, inconsistencies persisted. As consortium documentation explains, in their description of the harmonized measure, the values ‘are largely concordant’ with OMB categories, but ‘ultimately rely on if and how each study asked subjects about their race and ethnicity’. That is, data that went into the harmonized variable could be explicit self-identification or implicit recruitment criteria or a blanket classification made by the principal investigator of a given study. For example, if a study was designed to represent a cohort of Black Americans, that racial label was applied to all participants whether they were ever asked to self-identify or not. Similarly, participant data were recoded to match current U.S. Census-style categories even if they were not originally collected that way. When actual data collection and categories used conflicted with the harmonized scheme, the consortium document explained how this was handled: Some studies included the option ‘Asian or Pacific Islander’ (in contrast to separate options of ‘Native Hawaiian/Pacific Islander’ and ‘Asian’). In these cases, the value assigned for this variable was ‘Asian’. If a component study variable included the option of Hispanic/Latino, ‘Other’ was assigned for this variable unless there was additional information to allow for a specific race value to be assigned. In addition to ‘Other’ being assigned in this way, some studies had questionnaires where ‘Other’ was an option to select for race. In these cases, the value ‘Other’ was used for these subjects in this harmonized variable.
These decisions result in participants being represented differently between the pie chart and ECC’s harmonized variable. For example, although ‘Hispanic/Latino’ gets its own wedge in the chart, it does not appear in the harmonized race variable. Similarly, ‘Asian’ has its own wedge in the pie chart but, as the database documentation acknowledges, the harmonized Asian category may include studies that used a combined category of ‘Asian or Pacific Islander’. That is, even within the same consortium, different classification schemes are adopted for different diversity-related aims. The primary purpose of the pie chart was to show database diversity beyond just ‘European ancestry’ but when it came to creating a variable for use in analysis, concerns about some aspects of data comparability (but not others) led many of the cases to be classified in the database simply as ‘missing’.
Attempts to correct the use of race and ancestry
Dissatisfaction with such experiences, combined with growing national discussions about structural racism, led a group of ECC junior investigators and research staff to propose a set of guidelines on the use and reporting of race, ethnicity, and ancestry. Notably, these guidelines do not conform to ECC’s actual practices, nor were they adopted as an official consortium protocol.
4
Instead, they reflect the difficulty DC staff had while working with race and ethnicity data, and tensions they were noticing. An investigator explained: [There is a] general tendency for genetic researchers to conflate genetic ancestry with race and ethnicity … we do that in the way that we draw plots and the way that we talk about things or you read a paper and they say we studied European ancestry individuals and there was no discussion of does that mean people who have self-identified as white or does that mean you ran a principal component analysis. … We’re layering on all these different studies that have different study specific concerns about how their groups might be described or labeled. … When you get [so many] studies together, first of all, how is everyone going to know about all of those, especially when people may be analyzing data and aren’t working very closely with the study investigators from those different studies…also when you have people who have conflicting preferences.
The guidelines put forward a number of recommendations, including to: (1) explicitly distinguish between variables that derive from non-genetic, self-reported information versus genetically inferred information, (2) avoid terminology that is linked to hierarchical racial ordering (e.g., Caucasian), (3) avoid using race or ethnicity as a proxy for genetic ancestry, and (4) articulate the motivation for the use of race, ethnicity, or ancestry in a given study. These guidelines have since been pointed to as ‘best practices’ for the handling of race and ethnicity data in the broader genomics community and in other consortia in our study. 5 However, the suggestions themselves break little new ground, as they echo the recommendations of Boyd et al. (2020) and have been articulated in countless previous perspectives pieces over the past several decades (Mauro et al., 2022). But they also have had little obvious impact on research practices within ECC. For example, the pie chart—which seemingly violates recommendations 1 and 3—remains prominently displayed, though with a disclaimer that reads: ‘Please note that while groupings may correlate to some extent with genetic ancestry, [ECC] recommends distinguishing between genetically and non-genetically inferred descriptions in analyses and publications.’ Thus, appearance of diversity through continued display of the pie chart was prioritized over revisiting the underlying data harmonization process that created it.
Discussion and implications
Harmonization is a prerequisite to bringing together existing datasets and has become a key practice in precision medicine research. Some have criticized this escalating data aggregation in biomedical research, arguing that true precision will be found in studies with more local specificity (e.g., Ossorio, 2022). Our aim here is not to debate the relative advantages or disadvantages of aggregated versus more local datasets, but rather to underscore the need for more intentional harmonization efforts when data aggregation occurs. Epstein (2008) argued for greater attention to ‘recruitmentology’ as an emerging ‘ancillary science’ that followed inclusion mandates in biomedical research, and we make a similar case for harmonization in the wake of renewed diversity efforts in precision medicine and biomedical research more broadly. The stakes for when and how harmonization is done will only increase, as interest in dataset aggregation grows and with precision initiatives expanding to include global research consortia (e.g., Harris et al., 2012; Xuan et al., 2020). The harmonization of diversity measures and their implications for understanding human difference and health inequities will only become more consequential, as well (see Aspinall, 2007; Mulinari & Bredström, 2024).
But despite its increasing importance in biomedical research, we found harmonization work in general was materially undervalued, accomplished as ‘service work’, and often left up to individual initiatives. Funders emphasized the importance of harmonization, viewing it as central to the future utility of databases, and felt that it had been explicitly funded as part of the project work. Many investigators disagreed, lamenting the amount of work harmonization entailed and seeing it as distracting from individual study goals. Participants spoke about the challenges of harmonizing across studies that were ‘not even looking at the same populations or the same topics’. One investigator explained the tradeoffs: Being able to pull data and do analyses that have greater numbers, and sample size, and potential statistical power, I think is very important and a huge bonus. … [NCC] was set up as individual sites doing their own studies, asking their own research questions, with this added layer of data sharing and collaboration. That’s a little bit harder because it’s in the middle. You don’t have complete control over how you design your study and what you focus your attention on. But there’s also not complete harmonization in the data that you’re collecting. You could think of it as a happy medium, or you could think of it as no man’s land.
This tension is significant: Without ‘complete harmonization,’ the value of the data was compromised. The challenge of collaborating may have contributed to harmonization being seen as secondary work (many claimed it was not a line item in their budgets, for instance), with investigators ultimately adopting a ‘good enough’ approach. Ultimately, as previous research on standardization would have predicted, investigators treated harmonization efforts as an unnecessary constraint on their scientific aims and seemed to see harmonization of more ‘social’ data such as race and ethnicity, specifically, as diverting from rather than furthering their core efforts (see Ackerman et al., 2016).
In the cases we present here, harmonization of diversity measures was considered secondary to ‘the science’ despite funder mandates prioritizing population diversity in precision medicine research databases. The consortia we observed often embraced diversity but paid less attention to inequality, mirroring phenomena that scholars have observed in education, workplaces, and everyday life (Bell & Hartmann, 2007; Berrey, 2015; Bowman Williams & Cox, 2022). As they crafted grant proposals, researchers frequently read into what they thought funders would ‘count’ in terms of diversity efforts and pointed to their budgets to justify what work was and was not being supported during their study periods. As we have described elsewhere, enhancing diversity is accomplished through funding projects at already well-resourced institutions with ‘baked in’ diversity in patient populations (Shim et al., 2022) and through instrumental approaches to diversity, in which teams hire lower-ranking study staff who ‘look like’ the target populations in attempts to boost study enrollment (Jeske et al., 2022). Collectively, these findings indicate diversity measure harmonization efforts tend to be superficial and concerned primarily with appearances and surface-level compliance with funder mandates. This raises critical questions about downstream consequences of biomedical research that utilizes the data generated not only though consortia like these, in which attention was ostensibly paid to social measures, but also through other research efforts that are not mandated or incentivized to attend to such measures at all.
Building accountability
The decoupling of what consortia participants said about diversity and what they did was enabled by a relative lack of accountability. Study sites were held to recruitment benchmarks related to diversity—i.e., to representation by U.S. Census-style racial categories— but not to attending to measurement issues throughout data collection and analysis. A growing number of explicit journal guidelines related to the use of race in research (as discussed in Case 1) promise some accountability for transparency and conceptual rigor in the publication process. But, as our data showed, by then, it was often too late to reverse the relative disinvestment in collecting, managing, and optimizing ‘social’ data. Indeed, similar guidelines around the role of race in research have been promoted previously to little effect (Martinez et al., 2023; Mauro et al., 2022), suggesting that intervening only at the publishing stage is unlikely to produce lasting results.
Our research suggests an important role for funders in not only articulating diversity recruitment mandates but requiring clear conceptual frameworks at the grant application stage. Making these research design decisions subject to review would provide a stronger incentive for the various activities that fall under ‘diversity work’ to be incorporated more centrally into study designs. Precision medicine researchers could also benefit from consulting social science research that conceptualizes race as multidimensional (e.g., Roth, 2016), demonstrates different methodological approaches and their consequences for research conclusions (e.g., Guluma & Saperstein, 2022; Saperstein et al., 2016), and highlights the role of structural racism in health (e.g., Adkins-Jackson et al., 2022). If funding were contingent on review of measurement and harmonization plans, with the work of harmonization included in grant funding, our observations suggest it would be less likely to be done by people without relevant expertise and ‘on the side of everything else’.
These recommendations align with the recent U.S. National Academies of Sciences, Engineering, and Medicine report on the ‘Use of race and ancestry as population descriptors in genomics research’ (NASEM, 2023). The report urges researchers to justify the selection of population descriptors as it relates to the goals of specific types of studies and stop using race as a proxy for human genetic variation. It also calls on genomics research organizations to offer more appropriate cross-disciplinary training and create mechanisms for accountability, including at both the funding and publication stages. Our own call for deeper integration of social science perspectives into the design, conduct, and analysis of studies in precision medicine research aligns with this NASEM recommendation (see Reardon et al., 2023).
As researchers rely on ostensibly harmonized ‘big data’ generated by consortia like the ones we studied, the consequences of sidelining both diversity and harmonization will continue to reverberate. Continued reflection on the specific practices through which concepts like ‘race’ and ‘diversity’ are deployed is critical, particularly in fields like genomics that have been used to support and promote racist beliefs and eugenic policies (Comfort, 2018; Duster, 1990/2003). Without stronger systems of accountability shaping decisions earlier in the research lifecourse, for studies across disciplines, the failure to treat diversity work as integral to ‘the science’—even while diversity is promoted as a crucial scientific benefit—will likely persist.
Footnotes
Acknowledgements
We thank our study participants for sharing their time, perspectives, and experiences with us during observations and interviews. We also appreciate the input and collaboration of members, including Stefanie M. Fullerton, Emily Vasquez, Nicole Foti, and Michael Bentz for their data collection and analysis contributions.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Research reported in this publication was supported by the National Human Genome Research Institute of the National Institutes of Health under award number R01HG010330. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
