Abstract
This paper offers an examination of morphosyntactic factors that are generally understood to measure grammatical integration—and therefore used to help determine the status of other-language-origin nouns as borrowings or code-switches—through the lens of discourse, semantics, and lexical patterns. A total of 820 lone English-origin nouns surrounded by otherwise Spanish discourse are compared to Spanish and English nouns from the recorded speech of the same bilingual speakers in New Mexico. The semantic domains most open to English-origin nouns include both those traditionally expected, such as technology, and those generally thought to be unborrowable, such as kinship terms. In the case of determiner patterning, lone English-origin nouns’ propensity to occur with indefinite articles or as bare is linked to use in a nonreferential predicating function. Regarding gender, the preference for masculine assignment for lone English-origin nouns is tied to both nonreferentiality and the general patterns found in Spanish. The impact is felt here not from English, but from the conventions of the local community. Among their many functions, these nouns are best suited in this community for naming kin, classifying individuals as belonging to a certain occupation, and creating verbal compounds. It is argued that the morphosyntactic patterns found reflect the community norms, in which English-origin nouns tend to perform certain discourse functions. Systematic quantitative analysis thus reveals the powerful role of discourse referentiality of nominal forms, in tandem with local practices.
Introduction
The literature on contact, including that focused on Spanish, is replete with commentary on the use of lexical items from one language in the discourse of another, as in (1), in which English-origin cards is surrounded by Spanish-language discourse.
(1)
Francisco (TAP) (H) mucha gente juega muchos
no?‘(TAP) (H) lots of people play a lot of
right?’
The use of single other-language-origin content words is the most frequent code-mixing phenomenon in contact situations (Jake, Myers-Scotton, & Gross, 2002, p. 72; Poplack & Meechan, 1998, p. 127). Scholars have generally focused on two main questions: a) why speakers use certain lexical items in certain contexts; and b) how and to what extent speakers incorporate these items into the grammar of the language surrounding them. Regarding the motivations behind such usage, it is often suggested that some new items enter a language to refer to new cultural concepts, constituting what Smead (2000) terms “unique loans” (p. 292), as in (2), in which quarter refers to a culture-specific item.
(2)
Anita .. papá le pagaba a mijo un
‘.. dad used to pay my son a
Clearly, however, not all such items represent a lexical gap. In (3), for example, Francisco considers both running and correr as options in talking about winning a basketball game, and he uses them both within two contiguous Intonation Units (represented on separate lines in this and following examples).
(3)
Francisco .. (H) … en todo el año de=l–
.. State Champion[ship],
Gabriel [aquí] [2en Taos2]?
Francisco [2I mean en2] --
(H) porque=,
.. era puro
gana asi[na a puro
Gabriel [oh yeah].
[2yeah2].
Francisco [2(H)2] puro ‘.. (H) … all year during the,
.. State Champion[ship],
[here] [2in Taos2]?
[2I mean in2] --
(H) because,
.. it was just
you win [like that just
[oh yeah].
[2yeah2].
[2(H)2] just
Regardless of the general speaker motivations behind code-mixing, the grammatical status of single items is superficially ambiguous. Do they represent one-word code-switches, in which two grammars are juxtaposed, or are these items grammatically incorporated into the language that surrounds them, and thus best understood as borrowings (Poplack & Meechan, 1998)? In an earlier corpus of New Mexican and Southern Colorado Spanish (NMCOSS) (Bills & Vigil, 2008), Torres Cacoullos and Aaron (2003, p. 292) found that when incorporating English-origin nouns into their Spanish-language discourse, the speakers drew on Spanish, not English, grammar. This was evidenced in the patterns of bare (i.e. determiner-less) English-origin nouns in Spanish discourse, which followed those of bare Spanish-origin nouns in Spanish discourse and were unlike those of bare English-origin nouns in the English discourse of the same speakers. With this, they conclude that, on the whole, these items were treated as instances of borrowing, not code-switching, the latter of which is understood as “the juxtaposition of sentences or sentence fragments, each of which is internally consistent with the morphological and syntactic (and, optionally phonological) rules of its lexifier language” (Poplack, 1993, p. 255).
A reliable view of how English and Spanish are combined in discourse can only be obtained through systematic quantitative examination of actual usage in a particular bilingual community. This paper investigates the status of single English-origin nouns in bilingual discourse in New Mexico, as nouns were not only the most numerous single-word other-language items, but are also an open word class that is particularly susceptible to borrowing. However, unlike the earlier NMCOSS corpus, the data examined here also include copious unambiguous multi-word code-switching. Through quantitative exploration of morphosyntactic, semantic, and discourse factors, I address the question of whether these nouns are best understood as grammatically integrated or non-integrated into the recipient language (that is, as borrowings or single-word code-switches). At the same time, these factors reveal community-specific practices that intersect with—and illuminate—the issue of morphosyntactic integration. The results will illustrate the power local practices have in defining linguistic norms.
Data and methods
Participants and corpus
The data for this study are drawn from the New Mexico Spanish-English Bilingual corpus (NMSEB) (Torres Cacoullos & Travis, in preparation), constituted by approximately 340,000 words and 29 hours of spontaneous speech. The participants are speakers of a Spanish that has been spoken in New Mexico for centuries, and not descendants of recent immigrants. The Northern New Mexican communities represented in this corpus have been in contact with English for 150 years, and all of the speakers acquired both Spanish and English during childhood (Bills & Vigil, 2008). This diachronic depth affords us a wealth of information regarding the long-term effects of code-mixing within a community. NMSEB offers a particularly unique perspective on US Spanish phenomena, as it comprises spontaneous use of English and Spanish by a community of bilingual speakers who regularly use both languages, including during the same conversation (for more details on the corpus, see Torres Cacoullos & Travis, 2015; Travis & Torres Cacoullos, 2013). For this study, a sub-corpus of 18 interviews was examined, or approximately 180,000 words encompassing 17.5 hours of recorded speech from 20 speakers.
Single English-origin nouns occurred in 15 of these 18 interviews, and in the speech of 16 of the 20 speakers. (Interviews 4, 10 and 12 contained none, though these interviews did contain multi-word code-switches to English; see Wilson & Dumont, 2015. Interviewer data were not included.) 2 Among the speakers who did use single English-origin items, rates of use were highly variable, with normalized frequencies ranging from 6 to 215 tokens per 10,000 words (see Appendix).
Data: Lone English-origin nouns
The data include all English-origin nouns and compound nouns surrounded by Spanish discourse by the same speaker (N=1,112 nouns, from 2,449 English-origin words, including adjectives, verbs, and discourse markers), irrespective of Intonation Unit breaks. From this initial extraction, proper nouns, including names and places (N=629), as in (4), and nouns that acted as a title, and which were thereby connected to a proper noun (N=32), as in (5), were excluded from the study because proper nouns may not be subject to the same processes of integration as common nouns (Poplack, Sankoff, & Miller, 1988, p. 99, n. 8).
(4)
Anita … no tenemos que tenerle miedo a
‘… we don’t have to be afraid of
(5)
Sandra y mi
‘and my
Also excluded were examples in which the noun was used in a metalinguistic context in which the noun itself was the topic of conversation (N=18), as in (6).
(6)
Manuel … allá les dicen
‘… there they call them
Numbers, when treated as nominals as in (7), were included in this study; however, adjectival uses, as in (8), were excluded. Finally, items in which the noun of interest was truncated (N=12) were excluded.
(7) (8)
Monica Jared paga four hundred.
‘... Jared pays four hundred.’
Monica me retiré cuando tenía
‘I retired when I was
Beyond these exclusions, the English-origin lone noun dataset was further refined to ensure, first, that the noun in question would not generally be considered part of a monolingual Spanish speaker’s lexicon, and second, that the compounds were conventionalized compounds among monolingual American English speakers (and therefore could be treated as single words). To satisfy the first criterion, the English-origin nouns that were listed in the most recent edition of the Diccionario de la Real Academia (DRAE, http://www.rae.es/rae.html) were excluded from these data (99 lexical types, N=102; see Appendix). Retained were cases where the definition provided in the DRAE was entirely different from the semantics found in the data (N=29). These uses were not judged to be equivalent to the DRAE listings, and included nouns such as brecas ‘brakes,’ carton (as a container), chanza ‘chance,’ craques ‘cracks,’ cricket (the insect), home (rest home), light (as in rays of light), machine, post, robe, semi (truck), trust, and yarda ‘yard’ (as in a piece of land next to a house).
To satisfy the second criterion, that the compounds were conventionalized compounds among monolingual American English speakers (and therefore could be treated as single items), I only included compounds found in Merriam-Webster’s English Dictionary (http://www.merriam-webster.com/), like alarm clock, chain saw, day care, duct tape, ghost town, mental illness, and monkey bars (see Appendix for excluded compounds). This left 820 lone English-origin nouns.
Though adjectives in general were excluded from this study, a handful (N=11) of lone adjectives were found with articles, and therefore functioning as nouns, as in (9). 3 This structure is generally acquired between the second and third year among monolingual Spanish speakers, which may be taken to indicate a certain level of syntactic complexity (Snyder, Senghas, & Inman, 2001, p. 162). Its presence in these data—both with Spanish-origin and English-origin nouns—speaks to the extent to which English-origin words have been incorporated into everyday language.
(9)
Ivette como que les ponía más atención a los
‘It’s like she paid more attention to the
((graders)) than the
Baseline comparisons
There are several measures to determine whether an other-language-origin item is a borrowing or a code-switch, the former of which involves recourse to only one grammar, and the latter of which involves the juxtaposition of two grammars. Items that are frequent and that occur among many speakers (i.e. are diffuse) are expected to function as borrowings, but the primary criterion by which borrowing status is determined is on the basis of morphosyntactic integration: an established borrowing like troca ‘truck’ should pattern no differently from a Spanish-origin noun like camioneta ‘truck.’ Indeed, this morphosyntactic criterion applies regardless of whether the item is established; that is, any item that shows clear signs of morphosyntactic integration can be seen as a borrowing.
However, with languages as typologically similar as English and Spanish, morphosyntactic integration is often not easy to identify. This is particularly true because in many instances Spanish and English grammars show variability, such that any contrast is probabilistic, not absolute. To determine the grammatical status of lone other-language-origin nouns in code-mixed discourse on a larger scale, it is necessary to identify a conflict site, i.e. “a form or class of forms which differs functionally, structurally, and/or quantitatively across comparison varieties” (Poplack & Meechan, 1998, p. 132; Poplack & Tagliamonte, 2001, p. 101). In other words, the distributional patterns of contentious items (i.e. those that are neither frequent nor diffuse) must be compared with those of other items whose grammatical status is clear. Nouns in monolingual stretches of speech (e.g. Spanish-origin nouns surrounded by Spanish) are unambiguous. Moreover, established borrowings, as determined by diffusion within the speech community, also demonstrate morphosyntactic integration.
To identify conflict sites and examine distributional patterns, comparison datasets must be prepared. In this case, the NMSEB offers us the unique opportunity to explore the patterns found within the two languages of the speech community itself, instead of resorting to an idealized monolingual norm or some other regional variety. The use of Spanish and English from the same speakers allows us to avoid the “comparative fallacy” (Bley-Vroman, 1983), since we are assuming that the most appropriate measure of language norms, and therefore of integration or lack thereof, is within the bilingual community itself (see Poplack & Meechan, 1998; Pires & Rothman, 2009; Torres Cacoullos & Aaron, 2003).
To create comparison datasets for each of the languages in contact, a roughly equivalent number of Spanish-origin nouns surrounded by Spanish discourse (N=856) and up to the same number of English-origin nouns surrounded by English discourse (N=608) was extracted from the speakers, taking into account the number of occurrences provided by each in the lone English-origin dataset. 4 This ensured that the monolingual data were comparable with the English-origin data in terms of any individual biases or tendencies. For these datasets, extraction began at the 15-minute mark for each interview.
Results
Frequency distribution of lone English-origin nouns
The 820 lone English-origin nouns represent a total of 405 types. Of the 820 occurrences, 30% (N=244) appeared only once in these data (these will be referred to as “Singleton” lone items), and another 18% (N=148, with 74 types) appeared only twice. Another 23% (N=186, with 52 types) appeared three to five times, and 11% (N=92, 13 types) appeared six to nine times. The remaining 20% of the data (N=150) includes six types, which occurred between 12 and 51 times. These six types were: uranium (N=12), troca/truck (N=15), grandpa (N=19), daddy (N=25), grandma (N=28), and dad (N=51). Figure 1 shows this distribution.

Breakdown of lone English-origin nouns according to number of occurrences (N=820). For example, 30% of all tokens are of types that occurred only once.
However, given that 70 of the repeated items, such as uranium (N=12), were simply repeated multiple times by the same speaker, a better gauge for how widespread the English-origin words were is how many speakers used them. There were 314 types (N=433) used by only one speaker; of these, 244 occurred only once. There were 52 types (N=154) produced by two speakers. Finally, 12 types (N=51) were produced by three speakers, five types (N=53) by four speakers, and six types (N=129; grandpa, mom, weekend, troca ‘truck’, grandma, and dad) by six or more speakers. Those produced by three or more speakers will be considered “Diffuse” in this paper. Figure 2 shows this distribution both in terms of number of occurrences and number of types.

Breakdown of lone English-origin nouns according to number of speakers (for “Occurrences,” N=820; for “Types,” N=389). For example, 129 of all 820 occurrences (i.e. 16% of the data) were produced by six or more speakers; however, these 129 occurrences are made up of only six types (i.e. distinct lexemes), making up under 2% of the 433 lexical items.
To recap, I will call those produced only once “Singletons,” and those produced by three or more speakers “Diffuse.” This allows us to clearly distinguish between those items that are commonly used within the community and are therefore likely loanwords (and may be transmitted by other speakers), from those whose loanword status is not established (and which therefore may entail active access to English). In the case of Diffuse items, we would expect these to pattern like Spanish-origin nouns (which, for the most part, they do). In contrast, the Singleton items may pattern like English nouns, revealing them to be, in the aggregate, single-word code-switches. Alternatively, they may pattern like the Diffuse English-origin nouns and like the Spanish nouns, revealing their aggregate status as nonce loans that are instantly integrated morphosyntactically (Poplack, 2012; Poplack & Meechan, 1998).
Semantic domains
It has often been argued that lexical borrowings tend to be drawn from certain semantic domains, particularly those that have to do with new cultural items that are associated with the contact culture. Smead (2000) identifies two types of loanwords: “unique” loanwords, defined earlier, and “synonymic” loanwords, i.e. “those that compete for the same semantic space with a native language term” (p. 292). Teschner’s (1974) annotated bibliography reveals several semantic domains that appear more likely to have English-origin words, including education, food, sports, and technology. It has often been assumed that highly frequent or core vocabulary items (Swadesh, 1971) are not likely to be borrowed (e.g. Smead, 2000, p. 282). However, Myers-Scotton and Okeju (1973) argue that (based on evidence from the East African language Ateso) “borrowings within the core vocabulary itself are also very common, given sufficiently extensive contact with another culture” (pp. 872–873).
The preliminary categories coded here were gleaned from the previous literature. These were then supplemented as the analysis was guided by the data. The 17 semantic categories coded are: kinship terms (e.g. dad); everyday items (e.g. bag); events and places (e.g. birthday); person (not kin) (e.g. firefighter); year or number; work- or money-related (e.g. spending money); related to the land or earth (e.g. mountain); technology (e.g. treadmill); academia (or school) (e.g. homework); food, drink, and smoke (e.g. popcorn); vehicle or transport-related (e.g. troca); institution (e.g. social services); animal (e.g. bird); abstract time/space concept (e.g. weekend); health and body (e.g. eyedrops); domestic life (e.g. sewing); and linguistics or language (e.g. Spanish). Nominalized adjectives and other items that did not fall into one of these domains were coded separately.
The tendency for both culture-specific and more universal (or “core”) terms to be borrowed in this community (cf. Torres Cacoullos & Aaron, 2003) is seen in Table 1. The last column shows the proportion of each domain that is produced as lone English items (based on a comparison with the Spanish and English samples described above).
Semantic domains per lone English-origin nouns vs. Spanish and English nouns in corresponding samples.
This analysis shows that English-origin nouns do indeed occur relatively more often in particular semantic domains. These domains, however, are not all on the list of usual suspects, which, as mentioned, include technology, food items, and academia. The most favorable contexts for English-origin lone nouns in these data are kinship terms (which were produced as lone English-origin nouns 57% of the time), years and numbers (56%), technology (40%), and vehicles (52%). The items referring to events, technology, or vehicles, as in (10), (11), and (12), may indeed represent cultural novelties (though in the case of vehicles, 79% (15/19) of the Diffuse items are troca). However, the category of kinship belongs to the core vocabulary (Smead, 2000, p. 282) and cannot in this case be said to be culturally specific or represent lexical gaps.
(10) (11) (12)
Anita <VOX .. anda llévanos pa’l
‘<VOX .. come on take us to the
Francisco y tienen un
‘and they have a big
Manuel .. (H) .. y esto no lo puedo cargar en el
‘.. (H) .. and I can’t carry this onto the
The English-origin kinship terms found in these data exhibit certain lexical patterns that provide evidence that the conventionalization of English-origin nouns occurs at both the lexical and semantic levels (cf. Poplack et al., 1988), since we find both dad and grandma (but not other kinship terms) among the most frequent. For example, dad and daddy combined occurred 76 times as lone nouns, compared with four occurrences of papá among the Spanish nouns. In contrast, mom/momma occurred only eight times in the lone nouns, and mamá 15 times in the Spanish nouns. Grandma had 30 occurrences; abuela only one. Grandpa occurred 20 times, abuelo three. There were no occurrences of brother or sister, while mentions of hermano, hermana, and hermanos were made 11 times. Son and daughter did not appear (though son-in-law occurred once); hija and hijo occurred 18 times combined.
These patterns are mirrored in the corpus as a whole: English terms are used over 1.5 times as often as Spanish terms to refer to dads (305 vs. 182 tokens), while for moms, the breakdown is reversed (178 tokens of English terms vs. 219 in Spanish). 5 Given the overwhelming preponderance of kinship terms, they will be excluded from the analyses of morphosyntactic and discourse patterns in the Diffuse dataset, as they tend to display peculiar tendencies that would skew the results (for example, they show a greater tendency than other nouns to be referential, to occur in subject position, and with possessive marking).
In the datasets in this study, while the core items of kinship show evidence of high rates of English-origin nouns, the core items in the domain of abstract time or space, such as day, are much less likely to appear in English within Spanish discourse (2%, N=15), when compared with these types of items in Spanish (12%, N=101). Over half of these English-origin items are weekend(s) (N=8); items like day and year did not appear among the lone nouns, though there were 14 and 23 occurrences of their Spanish counterparts, día and año, in the Spanish, respectively. Nonetheless, the relatively high rate of specific years and numbers in the English-origin data shows that use of English in this related context is a local community norm. Indeed, the 36 occurrences of years (some as Singletons, some repeated across speakers, some Diffuse), as shown in (13), (14), and (15), are best understood as instances of a construction in which it is established practice to use English in this community.
(13)
Susan … nació en
‘he was born in eighteen sixty-nine,’
(14)
Francisco (TSK) (H) en el
‘(TSK) (H) in
(15)
Ivette … (H) se me hace que en ..
‘…(H) I think in ..
A third area that at first blush seems to be disproportionately favorable to lone nouns is that of vehicles. However, 15 of the 23 occurrences are represented by troca, which indicates that this is a lexical effect, rather than a semantic class effect. Thus, while some use of English-origin nouns are domain-specific (e.g. numbers, kinship), there are also lexical effects.
Gender
Gender has been a primary target for the examination of the status of lone items because it can be understood as one indication of morphosyntactic integration: if the gender patterns are similar to those of Spanish-origin nouns, then it can be argued that they are morphosyntactically integrated, and thus not code-switches (e.g. Poplack, Pousada, & Sankoff, 1982; Zamora, 1975; Zamora Munnt & Btjar, 1987). A preference for masculine gender assignment has been suggested in several previous studies, among L2 learners (Martinez-Gibson, 2011), bilinguals (Chaston, 1996; Clegg, 2010: 16; Garcia, 1998; Montes-Alcalá & Lapidus Shin, 2011; Sánchez, 1995: 134-137), and monolinguals (Banfield, 1994; Pérez-Pereira, 1991; Smith, Nix, Davey, López Ornat, & Messer, 2003).
To determine gender for both Spanish-origin and English-origin nouns, overt cues beyond the noun itself were used. These cues included the gender of determiners and adjectival modifiers, as in (16), in which machine was coded as feminine due to the feminine marking on otra. If no overt cues were present, the noun was coded as unmarked for gender. In cases in which the semantics of the noun dictated one or the other gender, such as girls, a separate code was used, to ensure that all nouns coded as having been assigned a gender were, at least in theory, sites at which either gender could have been assigned. 6 Overt cue-based gender coding was necessary across all datasets to ensure comparability of the coding.
(16)
Ivette tenías que irte pa’ otra
‘you had to go to another
Table 2 and Figure 3 show the results regarding gender assignment for these data. Note that if we compare only masculine and feminine, we get a rate of 14% (16/117) feminine for Singletons. This is nearly identical to the results reported by others for other bilingual populations (Mota, cited in Otheguy & Lapidus, 2003, pp. 214–215; Otheguy & Lapidus, 2003; Poplack, 1982). Also notable here is the high rate of unmarked nouns, at 51% for Singletons. A superficial analysis of Singletons alone might lead us to take this as evidence of their lack of grammatical integration. However, a comparison with the rate of unmarked Spanish nouns, also high at 39%, discourages such an interpretation. This slightly elevated rate of the absence of gender marking can be attributed not to a lack of grammatical integration, but rather to the greater proportion of the Singletons used in nonreferential contexts, as will be discussed below.
Distribution of gender marking on nouns in Singleton, Diffuse, and Spanish datasets.

Distribution of gender marking on nouns in Singleton, Diffuse, and Spanish datasets (percentages).
The preference for masculine gender reported in other studies of English-origin nouns in Spanish is found among these speakers as well, for both Singleton and Diffuse lone items, at 42% and 44%, respectively, compared with 30% for the Spanish nouns dataset. In contrast, only around 7% of the Singleton data was assigned feminine gender, compared with 21% for the Diffuse lone items and 32% for the Spanish nouns data. Figure 3 illustrates the contrast between Spanish and Singletons; Diffuse lone English-origin items fall in the middle.
The preference for masculine, however, likely has nothing to do with code-mixing tendencies per se, but may rather simply follow from patterns and preferences that are internal to Spanish. In a connectionist analysis of a longitudinal database of parental production, Smith et al. (2003) noted that “while regular feminine nouns were slightly more frequent than regular masculine nouns, irregular masculine nouns outnumbered irregular feminine nouns by roughly 2 to 1” (p. 306). Based on this input, which preserved type and token frequencies, a computer-generated model produced a similar bias toward masculine gender assignment to novel words, suggesting that the frequency distribution has a direct role in gender assignment. 7 When applied to lone English-origin items, such experiential patterns would likewise lead to a preference for masculine gender assignment.
Discourse function
Lacking from most previous studies of English-origin nouns in Spanish is consideration of patterns of referentiality (Hopper & Thompson, 1984; Thompson, 1997). This turns out to be a particularly profitable context where we may be able to measure the extent to which lone English-origin nouns are grammatically integrated, and can therefore be considered to be functioning as Spanish nouns and not as code-switches to English.
The data were coded for referentiality, which, in Hopper and Thompson’s (1984, p. 711) terms, refers to the “manipulability” of the referent; that is, “a noun phrase is referential when it is used to speak about an object as an object, with continuous identity over time” (Du Bois, 1980, p. 208). Thompson and Hopper (2001) identified three contexts in which nouns perform nonreferential functions: to form intransitive predicates with semantically weak verbs, to classify, and to orient. These contexts are illustrated in (17), (18), and (19), respectively.
(17)
Ivette .. él me daba
‘.. he would give me a
(18)
Susan …(1.1) era también
‘…(1.1) he was a
(19)
Manuel …(0.8) tenemos un toothpick,
…(1.0) es u=n,
.. credit card,
ése,
casi lo pudiera acarrear en la
.. pero no te dejan.‘…(0.8) we have a toothpick,
…(1.0) it’s a,
.. credit card,
that one,
you could almost carry it in your
.. but they don’t let you.’
Referentiality is key to the study of single other-language-origin nouns, as Torres Cacoullos and Aaron (2003) found that nonce-loan nouns in New Mexican Spanish often occurred in predicating functions, and were disproportionately frequent in classifying contexts, where they designated occupation or social status as predicate nominals. This led them to conclude that “It is not lack of grammatical integration, but these nonreferential uses in recipient-language predicates, that is manifested in bare nonce-loan nouns” (p. 293).
If we examine Table 3 and Figure 4, we see that both lone nouns datasets (Singleton and Diffuse) showed higher rates than both English nouns and Spanish nouns in the predicating role. This means that the former were more likely to appear as nonreferential objects of semantically weak verbs, thereby forming verbal compounds (Thompson & Hopper, 2001), as in (17). However, given that both Diffuse and Singleton lone nouns demonstrate this tendency, and this is not a conflict site between Spanish and English because nouns from both languages readily occur in this context, this cannot be understood as a measurement of grammatical integration. Instead, we may posit that this reflects the community practice of creating novel compounds, perhaps at times to express culturally specific concepts, as in (20). This could potentially be linked to the similar community practice of creating verbal compounds with hacer ‘do’ + V, which is a novel bilingual device (Wilson, 2013; Wilson & Dumont, 2015).
(20)
Mónica .. están agarrando
‘.. they are getting
Distribution of nonreferential functions of nouns in Singleton, Diffuse, Spanish, and English datasets.

Distribution of nonreferential functions of nouns in Singleton, Diffuse, Spanish, and English datasets (percentages).
In fact, in terms of discourse referentiality, the only notable difference between English nouns and Spanish nouns is found in the classifying role, as in (18). Here, we find that English nouns and Singletons are more likely to fulfill this function. However, 22% (9/41) of the classifying Singletons refer to occupation, compared with 9% (8/85) of the English nouns, 9% (6/69) of the Spanish nouns, and none of the Diffuse nouns. The nonreferential use of Singletons in occupation-related classifying function was also found in Torres Cacoullos and Aaron (2003), and may explain in part some of the gender-marking prevarication noted by Otheguy and Lapidus (2003) in this context. For example, in (21), the lone item does not actually refer to the person herself, but only a category to which she belongs. Since the category itself—social worker—is unmarked for gender (i.e. both women and men can fulfill this role), its gender assignment would not necessarily be linked to that of the woman, who represents only one example of somebody in that category.
(21)
¿Y tu mamá? Ella es un social worker, una trabajadora social…
‘And your mother? She is
The explanation for the elevated occurrence of English nouns in classifying role, however, is not the same as for Singletons. Instead, it comes from variations of the expression that’s the thing, which accounted for another 9% (8/85) of the English nouns data; in contrast, cosa ‘thing’ occurs only once in classifying position in the Spanish nouns dataset. Finally, overall, Singletons are less likely to perform orienting functions (24%) than nouns in all other datasets; examples include numbers and words such as north, winter, month, and weekend.
A final site in which the ramifications of the discourse functions of English-origin nouns in New Mexico can be found is in the determiner patterns. Determiners have been cited as an ideal area in which transfer might be identified (e.g. Montrul & Ionin, 2010). However, the scant evidence that exists on the question of transfer involving determiners among Spanish-English bilinguals offers mixed results. One difficulty lies in the relative scarcity of obvious conflict sites in English and Spanish.
Two clear loci of inter-lingual difference do exist, however. First, in contexts of inalienable possession, as in (23), Spanish generally has a definite article, while English would use a possessive pronoun. Second, in generic contexts, as in (24), Spanish generally has a plural definite article, while English nouns in these contexts tend to be bare. Unfortunately, these two contexts were rare in the Singleton data (N=1 for inalienable possessions; N=20 for generics), and thus could not be analyzed here. In order to gain a fuller understanding of conflict sites between typologically similar languages, then, we must move beyond these clearer cases. Corpus studies are particularly well suited for more fine-grained analysis.
(23)
Mónica le duele el
‘
(24)
Rafael … y qué tan ra- --
rápido estabas yendo?‘… and how qui- -- quickly were you going?
.. qué tan veloz?
Ivette .. pues,
no se me hace que venía muy recio porque me pude ir pa’ atrás pa’ adentro.
… y luego,
después me acuerdo que,
… % .. tenía tanto miedo a las .. how fast?
.. well,
I didn’t want to go that quickly because I could go back inside.
… and then,
afterwards I remember that,
… % .. I was so afraid of ((the))
Determiner types considered here included bare (25), definite (26), indefinite (27), possessive (28), demonstrative (29), and quantifier (30).
(25)
Fabiola qué
‘what
(26)
Monica .. está yendo
‘he’s going to
(27)
Sandra … (1.4) y entró con
‘...(1.4) and he came in with
(28)
Monica .. Todavía está en
‘.. She is still in
(29)
Bartolomé … (H) y nosotros agarramos parte de
‘… (H) and we would grab part of
(30)
Francisco (TAP) (H) mucha gente juega
no?‘(TAP) (H) a lot of people play
right?’
Table 4 and Figure 5 show the distribution of determiner types for all four datasets. Regarding definiteness, it has been proposed that definite articles are more frequent, and indefinite articles less frequent, in Spanish than in English (see, e.g. Vargas-Barón, 1952, p. 410). This characterization is borne out here: the Spanish data show 43% definite articles, compared with 30% in English; conversely, Spanish shows 7% indefinite articles, compared with 16% in English. The Diffuse items line up with Spanish in terms of definite articles, at 43%, and are closer to Spanish than to English in indefinite article rates. Singletons, however, as shown clearly in Figure 5, are in line with English for both, at 29% definite articles and 17% for indefinite articles. This suggests that, at least sometimes, speakers may treat these items like English in terms of (Spanish) article usage. However, if possessives and definite articles are combined into one category, the English looks much like the Spanish, suggesting that the “conflict” may not be as great as it appears at first glance.
Distribution of determiner types of nouns in Singleton, Diffuse, Spanish, and English datasets.

Distribution of determiner types of nouns in Singleton, Diffuse, Spanish, and English datasets (percentages).
One context in which Singletons do indeed seem to align with English is in the predicating context. Only 1% (15/159) of the predicating nouns in Spanish occurred with indefinite articles, as in (31); much more common was zero marking, as in (32), produced by the same speaker only moments earlier. This compares with 19% (12/64) indefinite articles among predicating Singletons and 24% (23/95) in English, as shown in Figure 6. Notably, the indefinite article with Singletons in this context, as in (33), tends to occur with culture-specific nouns, such as guzzler, recreational vehicle, and B-47. Thus, the predicating noun construction may be a context in which we may see some one-word code-switches or transfer in this particular construction, pending investigation with more robust numbers.
(31)
Fabiola tiene una
‘does he have
(32)
Fabiola …(1.2) y él no tiene
‘…(1.2) and he doesn’t have ((a))
(33)
Pedro .. antes le ponían un
‘.. before they would put a

Percentage of indefinite articles among predicating nouns in Singletons, English nouns, and Spanish nouns (N=64, 95, and 159, respectively).
Nevertheless, Singletons do not generally align with English when it comes to bare marking. Singletons clearly stand out; at 47% bare, these nouns are marked with zero about 16 percentage points more than nouns in the other datasets, all of which hover near 30%. It is worth noting that 13% (15/114) of the bare singletons are numbers or years, compared with 7% (13/176) of the English nouns. So, while the prevalence of numbers in the Singletons may play some part in the elevated rate of bare nouns, the overall trend goes beyond lexical tendencies: it is linked to nonreferentiality. Nonreferentiality is common among bare nouns, both in Singletons, at 75% (N=86/114), and in Spanish nouns, at 83% (N=184/253). The higher proportion of nonreferential contexts with Singletons, then, would explain the higher rate of bare nouns in this dataset (see also Torres Cacoullos and Aaron 2003).
In sum, an examination of morphosyntactic factors that are generally understood to measure grammatical integration—and are therefore used to help determine the status of other-language-origin nouns as borrowings or code-switches—through the lens of discourse, semantics, and lexical patterns, has revealed the fingerprint of this community’s code-mixing norms. The morphosyntactic patterns found are in fact reflections not of a lack of grammatical integration, but rather of the use of English-origin words to perform certain discourse functions. These norms are established locally; there is nothing about English nouns in particular that makes them more or less suitable in these contexts. As we saw in section 3.2 above, the power of the community norm in linguistic usage is found at the semantic and lexical levels as well, with years and certain kinship terms generally expressed as English-origin nouns.
Discussion
While some of the evidence here suggests that Singletons pattern like English and Diffuse nouns pattern like Spanish, a greater part of the evidence extends beyond the question of grammatical integration: the distributional patterns show that where English-origin nouns are different from Spanish it is not because they function like English nouns in English discourse, but rather because they serve specific, locally determined discourse functions that have been conventionalized within this community. For instance, the semantic domains found to be most open to English-origin nouns included both those traditionally expected, such as technology, and those generally thought to be unborrowable (Smead, 2000, p. 282), such as kinship terms. In fact, the nouns in the latter group were the most prevalent overall. In terms of discourse function, both Singletons and Diffuse English-origin nouns were found to be used more often to create verbal compounds, a trend not shared in either English or Spanish. While Singletons showed similar rates to English for classifying roles, diverging from Spanish and Diffuse nouns, this effect was due to different underlying causes: Singletons are often used to classify individuals according to their occupation, while in English expressions like that is the thing are common.
Utilizing the comparative method (Poplack & Meechan, 1998) to examine patterns of English-origin nouns occurring in a stretch of Spanish discourse and both their Spanish and English-origin counterparts in stretches of monolingual discourse in a corpus of New Mexican speech, this study has examined various semantic and discourse factors relevant to the occurrence of English-origin nouns in Spanish discourse. Community norms were revealed through certain contexts in which English-origin items distinguish themselves distributionally from both English and Spanish, illustrating the tendency to use certain types of English-origin nouns in certain discourse-functional contexts in this community.
In the case of determiner patterning, it was found that the English-origin nouns (occurring in a stretch of Spanish discourse) were more likely than both Spanish and English nouns (occurring in monolingual discourse) to be bare. These findings are similar to those found in an earlier, less-bilingual corpus (with less English and less multi-word switching) from the same community, in which Singleton English-origin nouns patterned like Spanish but were more frequently used nonreferentially (Torres Cacoullos & Aaron, 2003). However, here it was also found that Singletons occurred with indefinite articles in the predicating function, at a rate that aligns with English rather than Spanish.
Spontaneously used other-language-origin nouns generally behave like their recipient-language counterparts though in particular constructions they may constitute one-word code-switches (cf. Poplack & Meechan, 1995). Nonetheless, the impact is felt here not from English, but from the conventions of the local community. Among their many functions, these nouns are best suited in this community for naming kin, classifying individuals as belonging to a certain occupation, and creating verbal compounds. These results remind us that language is a local phenomenon, molded by speakers in everyday life as they build a grammar that works in their lives.
Footnotes
Appendix
Table A shows the production rate of each speaker in each interview, including both the raw number of tokens of lone English-origin nouns and a standardized number, which shows how many such tokens each speaker produced per 10,000 words.
The excluded noun types found in DRAE are: alumni, area, babysitter, base, baseball, basket, basketball, bingo, boom, bus, carbon, carro, chef, chequeadita, claustrophobia, closet, estaca ‘stake’, football ( listed as fútbol ), honor, kinder, kindergarten, lonche (lunch), magazine, material, miss, mister, mobile (móvil), nylon (nailon), panels, panties, party, pajamas, pizza, plataforme (plataforma), rails, record(s), renta ‘rent’, rifle, satin, shorts (short), supervisor, test(s), tour, tractor, and whiskey.
The excluded English compound types (N=102) are: a lot later, antique mall, antique stores, atomic bomb test, baby grade, baseball team, big star, birthday party, bus depot, business manager, camper trailer, card game, check stubs, child support, egg farm, electric heaters, first/second/third/fourth/fifth/sixth/eighth/ninth grade(s)/graders, first holy communion, fishing reports, fried egg, fun pocket, game wardens, grocery cart, hard metal, heavy equipment, ice fishing, last one, last period, lunch bag, major news, mechanic school, mid high, nylon dresses, power saw, rain gear, rock wall, roe sacks, second time, security guard, senior citizen center, senior year, sesame truck, sewing kit, shower stall, state championship, state fair, third one, third prize, tire tube, trophy fee, twenty minutes, VIN number, and Vista volunteers.
Acknowledgements
I would like to thank Meagan Day for her able assistance with some of the coding for this project. I would also like to thank Rena Torres Cacoullos and Catherine Travis for their generous input, and Naomi Shin for invaluable consultation regarding gender and code-mixing.
Funding
This work was partially supported by National Science Foundation (NSF) grant #1019112/1019122 awarded to Rena Torres Cacoullos and Catherine E. Travis to support the development of the NMSEB corpus.
