Abstract
Pitch syntax is an important part of musical syntax. It is a complex hierarchical system that involves generative production and perception based on pitch. Because hierarchical systems are also present in language grammar, the processing of a pitch hierarchy is predominantly explained by the activity of cognitive mechanisms that are not solely specific to music. However, in contrast to the processing of language grammar, which is mainly cognitive in nature, the processing of pitch syntax includes subtle emotional sensations that are often described in terms of tension and resolution or instability and stability. This difference suggests that the very nature of pitch syntax may be evolutionarily older than grammar in language, and has served another adaptive function. The aim of this paper is to indicate that the recognition of pitch structure may be a separate ability, rather than merely being part of general syntactic processing. It is also proposed that pitch syntax has evolved as a specific tool for social bonding in which subtle emotions of tension and resolution are indications of mutual trust. From this perspective, it is considered that musical pitch started to act as a medium of communication by the means of spectral synchronization between the brains of hominins. Pitch syntax facilitated spectral synchronization between performers of a well-established, enduring, communal ritual and in this way increased social cohesion. This process led to the evolution of new cortico-subcortical pathways that enabled the implicit learning of pitch hierarchy and the intuitive use of pitch structure in music before language, as we know it now, began.
Introduction
Vocalization is an important way for all primates to communicate (Fedurek & Slocombe, 2011), especially Homo sapiens, as we use our voice to cry, to laugh, to sing, and to speak. Of all the different forms of human vocalization, speech seems to be the most complex and have the most capacity to convey information. These characteristics are possible because speech represents an example of a Humboldt system, one that is based on a restricted set of discrete units organized by combinatorial rules (Chomsky, 1965). Such a system enables information to be coded as chains of units. However, speech is not the only example of a natural recursive system used by human beings. Another example of a Humboldt system is music (Merker, 2002). In music, the discrete units on which the system is based are rhythmic measures and pitch classes (Merker, 2002, 2003) or pitch intervals (Jacoby et al., 2019). However, discrete pitches can be used independently of rhythmic units measured by reference to musical pulse, as in non-metrical music or music in “free rhythm,” as played and sung in many non-Western cultures (Clayton, 1996). This suggests that musical pitch is probably processed by a cognitive mechanism separate from the one that processes rhythm syntax. Moreover, pitch classes can be organized in a syntactically complex way (Lerdahl & Jackendoff, 1983) similar to phonemes and words in speech (Rizzi, 2009). Nonetheless, while grammar, in the case of speech, is more or less connected to referential meaning (Dor, 2000; but see Fitch, 2017), there is no such connection in music (Lerdahl, 2013; but see Koelsch, 2011b). For example, we can infer a new meaning from concatenated language units, as is the case when the meaning of a word is changed by adding a prefix. Such semantic concatenation is generally lacking in music. Instead, in the case of pitch syntax, pitches arranged in different pitch contexts can elicit different kinds of subtle emotions often called tonal “qualia” (Huron, 2006, pp. 144–147; Margulis, 2014, p. 30). These subtle emotional states seem to be caused by different levels of uncertainty and surprise, which represent the prospective and retrospective states of pitch expectations respectively (cf. Cheung et al., 2019).
As it may be difficult to determine a direct relationship between pitch syntax and the proposed adaptive values of music, pitch syntax has often been described as a byproduct of language cognition (Lerdahl & Jackendoff, 1983) or of some other more general cognitive mechanism (Huron, 2006; Krumhansl, 1990; Krumhansl & Cuddy, 2010). Fitch coined the term “dendrophilia” to describe such a multi-domain ability to interpret stimuli in terms of tree-like structures (Fitch, 2014, 2017). According to Fitch, dendrophilia is not restricted only to language and music, but can also be found in the visual arts, and characterizes human thought more generally (Fitch, 2017). In fact, there are studies showing that the same brain structures are activated during the processing of both language grammar and music syntax (Fedorenko et al., 2009; Kunert et al., 2015; Maess et al., 2001; but see Rogalsky et al., 2011). Moreover, there is also evidence of interaction between the processing of language and music syntax (Fedorenko et al., 2009; Koelsch et al., 2005; Slevc et al., 2009; Steinbeis & Koelsch, 2008). However, this interaction may be due to the use of a similar amount of attentional resources in both tasks, rather than the activity of a shared syntactic processor (Perruchet & Poulin-Charronnat, 2013). Also, the fact that these same brain structures are active during the processing of language and music syntax suggests that they may be parts of different circuits that perform different functions. The question is, however, whether the syntactic function related to pitch has been naturally selected, or whether it is just a cultural invention, the result of using an existing cognitive mechanism to carry out a new, socially learned task (cf. Patel, 2008). This view would support the former explanation, which is consistent with the idea of evolution as “tinkering” (Jacob, 1977, p. 1164), as it is easier to adapt an existing structure to serve a new function, where possible, than to build it from scratch. From this perspective, the processing of pitch syntax can be viewed as a mechanism that evolved separately from language grammar, in response to different selective pressures, even though the same brain areas are involved in both tasks.
Pitch as a Feature of Human Vocalization
The term “pitch” refers to a human auditory sensation that is the result of the interpretation of a periodic sound wave by the nervous system. A periodic sound wave consists of various acoustic parameters, of which the main one influencing our sensation of pitch is fundamental frequency (F0). Although pitch as a psychoacoustic percept is described as “that attribute of auditory sensation in terms of which sounds may be ordered on a scale extending from low to high” 1 (Stainsby & Cross, 2012, p. 47), the interpretation of sound frequency by the nervous system is much more complex than this unidimensional sensation (Rakowski, 1999, 2009), depending on both the external acoustic context and internal cognitive factors. As we perceive the sounds of human vocal expression, other than detecting simply whether a particular sound is lower or higher than another, we attribute additional qualities to our sensation of pitch depending on what kind of vocalization we are hearing, for example crying, speech, or singing. In the case of hearing a crying baby or speech prosody, changes in sound frequency are interpreted as forms of emotional expression (Reybrouck & Podlipniak, 2019) and thus emotional intensity (arousal) and valence (Bachorowski, 1999; Jaquet et al., 2012; Zeskind, 2013; Zimmermann et al., 2013). This kind of vocal communication, in which parameters such as volume, pitch, and tempo change continuously, makes use of “expressive dynamics” (Merker, 2003, p. 405) or is called “affective prosody” (Zimmermann et al., 2013, p. 117) and is not exclusive to Homo sapiens (Merker, 2003).
Pitch can be experienced as a continuous sensation, as described above, but it can also be perceived as a series of discrete pitches. In tonal languages, for example, different changes of F0 within a vowel (or between vowels) are important cues for lexical recognition, creating the so-called “lexically contrastive tones” (Gandour, 1978, p. 64) that are phonologically as important as phonemic contrasts. Similar cues can be observed in certain whistle languages (Meyer, 2008) in which contrasts between tones are imitated by means of whistling. However, even during the perception of lexical tones in speech, representations of pitch are coarse in contrast to the fine representations of pitch classes in music (Zatorre & Baum, 2012). In fact, in the Western tradition, professional singing requires a high level of pitch accuracy that is related to the Western musical system in which every pitch class is defined by precise tuning: for example, A4 = 440 Hz. Apart from this, there are people who speak with normal prosody but lose the ability to sing even familiar melodies (Peretz, 2006).
Some scientists have therefore suggested that the processing of fine-grained pitch in music is a domain-specific ability (Peretz et al., 2002; Peretz & Coltheart, 2003; Peretz & Zatorre, 2005). Yet there is more to the difference between the processing of pitch in speech and in music than “coarse” and “fine” pitch representations, respectively. On the one hand, there is no well-established standard of “precise” intonation in many examples of ethnic music (Ambrazevicius, 2006, 2014). In addition, although variations of F0 in speech are more continuous than they are in singing, even sounds with varying F0 can be recognized by the human listener as discrete pitches (Deutsch et al., 2011). This ability must be cognitive, as sounds consisting of varying frequencies are often interpreted in this way. For example, when we listen to the song of a starling we hear discrete pitches (Shannon, 2016) that are in fact not a part of the structure of the song as it is recognized by starlings themselves (Bregman et al., 2016). On the other hand, Patel (2008) describes an example of the use of musical pitch intervals as surrogates for discrete-level tones in the Benčnon tone language. It would therefore seem that fine-grained pitch intervals can be used not only as part of a musical code (Rakowski, 2009) but also as part of a lexical code. The two codes function in quite different ways, however.
The Uniqueness of Pitch Syntax
Broadly speaking, syntax is the system whereby signals are formed and mapped onto meaningful representations in a process involving the combination of a finite number of units into an infinite variety of communicative signals (Hilliard & White, 2009). As a musical pitch space (such as an octave, in Western classical music) consists of a limited set of pitches, perceived as discrete from which an infinite number of melodies can be composed, the organization of musical pitch can be described as syntactical. Pitch, or tonal, hierarchy, the basis for the syntactical ordering of pitches in tonal music 2 (Krumhansl, 1990), makes the use of pitch in music unique. Although the musical scales used in different cultures vary in terms of their intervals, even in cultures that use scales in which pitches are equidistant from each other, individual pitches are experienced in relation to one (Ross & Knight, 2019) or more pitch centers (cf. Ambrazevičius & Wiśniewska, 2009) that underpin the organization of every pitch hierarchy. In fact, there are certain similarities between tonal hierarchy and grammatical hierarchy in language (Lerdahl & Jackendoff, 1983; Sloboda, 1986). However, these two systems have different functions. The main difference is that pitch hierarchy is unconnected to conceptual meaning (Asano & Boeckx, 2015; Lerdahl, 2013), unlike grammar hierarchy and semantics in language (Bickerton, 2009). Rather, pitch hierarchy is based on the subtle emotional sensations that accompany the perception of particular pitch classes and which are elicited according to those that precede them in a particular sequence of pitch classes (Huron, 2006). In other words, units of speech are mapped onto conceptual representations (Hilliard & White, 2009), whereas units of music are mapped onto subtle, preconceptual emotional states. There is nothing resembling an emotion-based hierarchy such as this in any other form of human vocal expression (Jackendoff & Lerdahl, 2006). Another unique characteristic of pitch syntax is the fact that it seems to be untranslatable into other modalities, as is possible with language grammar or rhythm syntax in music. While both grammatical and rhythmic relations can be expressed by the means of movement (gestures), as in sign language and dance, respectively, tonal relations can only be felt while listening to or performing music. Although the syntactic relations in rhythm syntax are also felt, there is a difference between pitch and rhythm tree-like hierarchies in terms of their dependence on time.
While the hierarchy of musical meter is strictly dependent on the timing of beats (London, 2012), the hierarchy of pitches seems to be restricted solely by the span of working memory (Woolhouse et al., 2016), which is most probably determined by the number of pitches rather than time (Ding et al., 2018). As a result, pitch hierarchies can be more elaborate than rhythm hierarchies, since they are based on the spectral processing of sound that can be, to some extent, independent of the passing of time. In Western tonal music, the independence of pitch hierarchy from time can lead to redefinition, as in the case of modulation, which usually produces an emotional effect. The potentially greater complexity of a pitch hierarchy in comparison with a time-dependent rhythm hierarchy suggests that pitch syntax is evolutionarily younger than rhythm syntax. However, the unimodality of pitch syntax combined with its lack of connection with any propositional meaning suggests that pitch syntax is probably more conservative and evolutionarily ancient than grammar in language, whose hierarchy is based on morphology. Nonetheless, while grammar has been treated as an essential component of language (Chomsky, 1965; Pinker, 2009), musical syntax is usually described as the result of the specific use of cognitive mechanisms that evolved because of adaptive functions unrelated to music (Huron, 2006; Lerdahl & Jackendoff, 1983).
For instance, according to Patel’s “shared syntactic integration hypothesis” (Patel, 2003, p. 678), the processing of pitch relations in music, and grammar in language, involves domain-specific representations (pitches in the case of music; words in the case of language), which are the subjects of similar cognitive operations that share neural cortical resources. As mentioned before, neuroimaging studies reveal shared neural activity during the processing of syntax in music and speech (Brown et al., 2006; Fedorenko et al., 2009; Maess et al., 2001; Patel et al., 1998). In addition, research shows that the processing of pitch syntax violations interferes with the processing of syntactic violations in language (Slevc et al., 2009), which also supports Patel’s view (2003, 2013).
However, the processing of syntactic violations in language does not interfere with the processing of timbral violations in music (Slevc et al., 2009). This poses the question as to why we draw on the same neural resources when we process grammar in language and pitch in music but not when we process timbre in music. Patel has argued that, if musical pitch abilities represent an adaptation, then they should develop much earlier than has been observed. In his opinion, children’s sensitivity to tonality is slower to develop than their syntactic language abilities (Patel, 2008), based on evidence that 5-year-olds are better at detecting out-of-key than within-key changes of melody (Trainor & Trehub, 1994). However, more recent research has shown that the rate and sequence of the development of the syntactic abilities that are necessary to recognize music and language are in fact comparable (Brandt et al., 2012; Corrigall & Trainor, 2009). For instance, a tacit knowledge of pitch syntax has been observed in 30-month-old children (Jentschke et al., 2014). Human sensitivity to pitch syntax seems to begin immediately after birth (Perani et al., 2010), or possibly even earlier: Lecanuet et al. (2000) report that the fetus in the womb can be observed to detect changes of musical pitch during the third trimester of development. Woodward (2019) argues that the higher-order processing of sound probably begins at the same time, during the third trimester. Prenatal exposure to melodies during this period results in the creation of neural representations that last at least four months after birth (Partanen et al., 2013). Perani et al. (2010) presented simple tonal melodies and syntactically disturbed versions of the same melodies to 1–3-day-old neonates, who exhibited predominantly right-hemispheric activity while processing the simple melodies and reduced activation in the right hemisphere while processing the syntactically disturbed versions. It has also been observed that newborn babies process pitch intervals in a similar way to adults (Stefanics et al., 2009). All these results suggest that certain developmental predispositions are related to the processing of pitch syntax. The fact that the same cortical structures are involved in the processing of both language and pitch syntaxes does not mean that the two kinds of processing ability cannot be separate adaptive mechanisms (Peretz, 2011).
Another explanation of the recognition of pitch hierarchy focuses on a general mechanism of prediction (Huron, 2006). According to Huron, disparate emotional reactions to diverse pitch classes are the result of different probabilities of occurrence that are calculated by the human nervous system based on previous experience. In other words, by using statistical learning, our brains can create expectations concerning every subsequent pitch class in a musical sequence, and the position of a particular pitch class in the pitch hierarchy depends on how likely it is to occur. Although Western pitch hierarchy fits neatly into the probabilities of pitch occurrence calculated from the database of the Western repertoire, this theory does not explain its uniqueness. After all, a similar mechanism of prediction based on statistical learning is involved in the processing of speech phonotactics (Adriaans & Kager, 2010). Phonemes in speech, like pitch classes in music, are distributed with different statistical frequency. If so, something similar to pitch hierarchy should exist in the domain of phonemes. But, so far as it is known, there is no such “phoneme hierarchy” based on emotional qualia in speech cognition. Pitch syntax and phonotactics would seem to be fundamentally different in this respect, although more research is needed to explain why and how. This does not mean that tacit statistical knowledge and predictive strategies do not play an important role in both. Statistical learning is related to both the activity of cortical areas (including the medial temporal lobe and the dorsolateral prefrontal, cingulate, and sensory-motor cortices) and the striatum (Frost et al., 2015; Turk-Browne et al., 2009; Wang et al., 2017), suggesting cortico-subcortical processing. Similar patterns of activity have been observed during both the processing of speech (Chan et al., 2013; Friederici & Kotz, 2003) and music (Geiser et al., 2012; Salimpoor et al., 2011; Salimpoor et al., 2013). According to some scientists (e.g., Gorzelańczyk, 2011; Koziol & Budding, 2009; Strick et al., 1995), subcortical structures such as the putamen, caudate, and nucleus accumbens are parts of functionally different cortico-subcortical loops. This view is consistent with the claim that the striatum plays a significant role in regulating the salience of subsequent musical events (e.g., pitches), which resolves uncertainty (Cheung et al., 2019). Therefore, the processing of music and language syntaxes should be understood as the result of a complex interdependence between cortical and subcortical brain areas (Lieberman, 2002; Ullman, 2006) rather than being the result of cortical activity alone (Patel, 2003, 2013). Nevertheless, because of observable differences between behaviors representing the understanding of pitch syntax and rhythm syntax in music and phonotactics and grammar in speech, it is reasonable to assume that these two kinds of syntax depend on functionally different loops that have evolved as the result of different selective pressures.
Many characteristics of pitch syntax suggest that it is an adaptation. First it is ubiquitous. Although the universality of chroma octave equivalence has been questioned (Erickson, 1986; Jacoby et al., 2019; Voisin, 1994; Will, 1997), all cultures have music consisting of pitch intervals (Harwood, 1976; Brown & Jordania, 2011; Nettl, 2000) that are organized according to culture-specific rules. Second, the most obvious expression of syntactically organized pitches is in singing (cf. Bannan, 2019), which is a natural form of activity among tribal societies (Morley, 2013). Third, although not all people sing (or play music) in contemporary Western societies, we are all equally good at discriminating pitch structure violations (Peretz et al., 2003) and at detecting phonological and grammatical errors (Batterink & Neville, 2013; Kröger et al., 2016). This efficiency is a result of the implicit learning of tonal rules (Ettlinger et al., 2011; Tillmann, 2005; Tillmann et al., 2000), which resembles the learning of grammatical rules in language. This means that even “mere exposure” (Zajonc, 1968, p. 2) to tonal music is enough to learn its syntactic rules (Tillmann, 2005; Tillmann et al., 2000). In other words, human beings do not need any formal teaching to recognize hierarchical relationships between pitches and thus so-called out-of-key notes. Given that they can learn such a complex and specific system of communication suggests that human beings have a natural proclivity for learning languages described by Darwin as an “instinctive tendency to speak” (1871, p. 55).
Fourth, another indication that the recognition of pitch syntax can be driven by certain genetic factors is the fact that some people with congenital amusia do not recognize out-of-key notes (Ayotte et al., 2002). Congenital amusia is most probably hereditary (Peretz, 2008; Peretz et al., 2007), although there is debate as to which particular genes are related to musical abilities (Tan et al., 2014). Fifth, the specificity of pitch syntax processing is also observed in the domain of memory. It seems as if human beings are born with a kind of implicit mnemotechnic that facilitates the memorizing of pitch sequences. Sequences of pitch classes ordered according to tonal syntactic rules are remembered better than sequences of words, digits, and nonrepresentational figures (Steinke et al., 1997). Sixth, another characteristic of pitch syntax is recursion (Koelsch, 2011a; Rohrmeier, 2011; Woolhouse et al., 2016). On the one hand, recursion as a feature of language has been indicated as the only human-specific part of the faculty of language (Hauser et al., 2002). On the other hand, recursion is observed in different human cognitive systems (Pinker & Jackendoff, 2005). Nonetheless, the recursion which characterizes pitch syntax seems to be unique because it is based on a preconceptual tension–relaxation framework (Jackendoff & Lerdahl, 2006). The same framework seems to enable the experience of “long-distance dependencies” between pitches (key relationships) in music (Koelsch et al., 2013; Woolhouse et al., 2016). Although there are similarities between language and pitch syntax in this respect, long-distance dependencies in language (Bickerton, 2009) are connected with conceptual grammatical and semantic categories, whereas the recognition of pitch dependencies is related to subtle preconceptual sensations of tension and relaxation.
The Adaptive Value of Pitch Syntax
If pitch syntax is a unique feature of a species-specific communicative system, then it should play an important adaptive role in human evolution. Unfortunately, both music and speech syntaxes are more complex than any rules of sound organization observed so far among songbirds and other animals (Fitch & Zuberbühler, 2013), which makes it impossible to assume that the adaptive function of pitch syntax is analogous to the organization of sound in the animal kingdom. Nonetheless, it is often argued that music possesses an adaptive value (Darwin, 1871; Hagen & Bryant, 2003; Miller, 2000; Mithen, 2009; Mithen, 2006; Morley, 2013; Peretz, 2006; Roederer, 1984; Wang, 2015). While the influence of sexual selection on the evolution of music remains debatable (Mosing et al., 2015, Ravignani, 2018), an increasing number of studies indicate that the origins of music may be related to the social character of our species (Koelsch, 2013; Pearce et al., 2015; Pearce et al., 2017; Reddish et al., 2013; Tarr et al., 2014; Weinstein et al., 2016). However, music is an amalgam of many different features that influence various cognitive mechanisms (Zimmermann et al., 2013), such as simple reflexes in response to sudden changes in the volume of sound or the advanced processing of multifactorial structural relations, both temporal, as in the case of rhythm, and spectral or spectro-temporal, as far as pitch and timbre are concerned. These cognitive mechanisms evolved at different times in response to different selective pressures (Merker, 2003; Mithen, 2006). Therefore, the “social bonding” effects (Tarr et al., 2014; Weinstein et al., 2016) of music that have been observed must be the result of features specific to music, since other forms of human vocal expression are less efficient at enhancing social bonds. Unfortunately, there is no agreement as to the particular features that constitute music (Brown et al., 1999). One example of a very effective tool for social bonding is rhythm that enables movement to be synchronized to a beat (Dunbar et al., 2012; Fitch, 2013; Kirschner & Tomasello, 2010; McNeill, 1995; Reddish et al., 2013; Tarr et al., 2014). However, although rhythm can play an important role in certain forms of social bonding, it does not explain the widespread use of pitch in bonding rituals. The fact that the evolution of the vocal control of pitch (Bannan, 2012; Morley, 2013) coincided with the increasing size of hominin groups (Dunbar, 2014) suggests that pitch had an important role in hominin vocalization, which could have been related to the need to resolve social tensions.
If vocal “grooming” lay behind the origins of speech and music (Dunbar, 1996), then ritual song could have evolved as a tool for social bonding that complemented speech as the size of hominin groups increased. Conversations usually involve just a few individuals (Dunbar, 2014), whereas singing, especially in tribal societies, is a group activity (Blacking, 1973) in which all members of a tribe can participate. Therefore, while speech seems to be the tool that is most effective for sustaining close interpersonal relationships (Dunbar, 2014), and to establish social hierarchy, as the higher someone is in the social hierarchy the more people want to talk with them, singing is more egalitarian because all who sing feel that they are equally important members of the group (Pearce et al., 2017). Even the collective recitation of prayer often leads to the impression of melody similar to the result of the speech-to-song illusion (Deutsch et al., 2011), which could have been the source of many religious oral traditions such as Gregorian or Vedic chanting. Weinstein et al. (2016) have shown that choral singing facilitates social bonding, the more so in large choirs than small choirs. In a study by Stewart and Lonsdale (2016), players of team sports and choral singers reported higher levels of psychological wellbeing than solo singers. The choral singers regarded their choirs to be more “coherent” than did the sports players their teams. The authors therefore attribute the experience of wellbeing associated with choral singing to group membership rather than the activity itself. This may indicate that group singing is more effective than team sports for social bonding. Choral singing has also been shown to create social bonds faster than other group activities such as creative writing and crafting (Pearce et al., 2015), suggesting that participation in collective singing is a good way for coalitions to be formed quickly. However, when someone joins a choir for the first time, they need to be able to learn unfamiliar sequences of pitches. To do so it is essential that they possess appropriate pitch expectations in the form of a tonal hierarchy based on emotional qualia (Huron, 2006; Margulis, 2014). These qualia, linked to particular degrees of the scale, facilitate the prediction of pitches, which in turn accelerates the learning of pitch sequences. It has been suggested that social cohesion can be increased by aligning brain states (Bharucha et al., 2011). In the domain of pitch, such alignments are achieved by a kind of spectral synchronization whereby performers and listeners (and of course choral singers are both, simultaneously) experience similar neural representations of the same sound spectra. If so, the alignment of the brain states of choral singers, which involves not only the processing of spectral but also of emotional and even motor information, could be the main mechanism responsible for the observed socializing effect of collective singing, via the elicitation of similar motivations.
The Evolution of Pitch Syntax
Although pitch syntax appears to lack an obvious precursor within the primate lineage, it could have had its roots in the vocalizations of the last common ancestor of both chimpanzees and hominins. In fact, changes of F0 are an important part of a primate’s affective displays, which indicates that pitch was most probably an important part of our ancestor’s vocalizations. However, in the majority of nonhuman primate calls, pitch is used as part of an affective vocalization (Ackermann et al., 2014; Briefer, 2012; Bryant, 2013; Hauser, 1993; Snowdon & Teie, 2013) representing the expressive dynamics described above (Merker, 2003). Therefore, nonhuman primate calls, comprising continuous or “analogue” sounds, are the behavioral counterparts of human laughing, crying, and moaning, rather than speech and music, which are digital and combinatorial (Fitch & Zuberbühler, 2013; Merker, 2002). Based on the supposition that analogue and digital forms of vocal communication are processed in different brain circuits, dual-pathway model of human acoustic communication has been proposed (Ackermann et al., 2014). In this model, human nonverbal affective vocalizations, including speech prosody, belong to a separate domain, evolutionarily older than that of syntactically organized articulate speech. In a similar vein, the “quartet theory of human emotions” has been proposed (Koelsch et al., 2015, p. 19), in which evolutionarily older pre-verbal feelings are transformed into evolutionarily younger symbolic codes such as language. While these models explain the complexity of language, highlighting the different evolutionary ages of its different components, they do not characterize the whole spectrum of human vocal communication. As has been emphasized, pitch syntax, like language syntax, is digital (in the sense that the user codes discrete units of information), recursive, volitional, and implicit. This being the case, there must either be a speech-specific circuit that has been recruited to operate on pitch, or there are two functionally different circuits operating on phonemes (words, sentences) and pitches (phrases, melodies), which have evolved as a result of different selective pressures. The first explanation leads us to ask why the language-specific pathway is used to operate on phonemes but not on pitch intervals and dynamics. Although pitch contours (or levels) can be used in speech as syntactically valid units, for example as lexical and grammatical tones in tonal languages, they never create self-reliant systems (even in tonal languages, pitch cues are always combined with other spectro-temporal characteristics; that is, lexical and/or grammatical tones combined with consonants and vowels (Best, 2019)), as in the case of the pitch hierarchy in music. Another unanswered question is why a pitch hierarchy has never been mapped on to referential meaning. In other words, why is there no language that combines pitch hierarchy with semantics, even though pitch can be used as a distinctive feature of the discrete components of articulate speech, as in a lexical and grammatical tone in tonal languages? So far, the explanation that there is a speech-specific circuit recruited to operate on pitch offers no convincing answers to these questions; the second explanation is more convincing.
The most obvious behavioral difference between Homo sapiens and other primates in relation to vocalization is vocal learning (i.e., our ability to produce sounds that we hear with our own voice) (Fitch & Jarvis, 2013; Janik & Slater, 1997). It is necessary to speak and sing (two human-specific forms of vocalization, according to Merker, 2012), which are partly learned and partly volitional. The appearance of vocal learning in the human lineage was made possible by changes in hominins’ vocal tracts and vocal control attributable to the increased innervation of the larynx (Fitch, 2000; Simonyan, 2014), and cerebral innovations related, at least in part, to the development of new cortico-subcortical connections (Fitch & Jarvis, 2013; Jürgens, 2002, 2009) that supports the view, referred to above, that interactions between cortical and subcortical regions are crucial for the processing of pitch syntax. According to Merker (2009) the evolution of vocal learning is strictly related to the ritual culture that enabled the evolution of language and music. Another possibility is that vocal control gave an advantage to infants who competed for care in cooperative breeding (Zuberbühler, 2012), which does not exclude the possibility that the communication between caregivers and infants took a ritual form (Bannan & Woodward, 2009; Dissanayake, 2008; Trevarthen, 2008). Whatever the reason for the evolution of vocal control among our ancestors, the learning of both speech and singing (i.e., vocal expressions of language and music) is based on the imitation of sounds in which selected spectral and spectro-temporal cues play an important role.
Likewise, tonal and harmonic relationships are imitated in proto-lingual mother-infant interactions (Van Puyvelde et al., 2010; Van Puyvelde et al., 2015), and similarities have been observed between the speech intonation of adults and musical intervals (Bowling et al., 2012; Ross et al., 2007). Yet an “interval-like” intonation in speech is not the main carrier of referential meaning, which suggests that the tonal and harmonic features of speech may be the remnants of a pre-verbal protolanguage (Brown, 2000, 2017). By contrast, chimpanzees have been observed to use their lips so as to imitate certain sound parameters for the purpose of communicating referential meaning (Kalan et al., 2015; Watson et al., 2015). Therefore, while complex, humanlike laryngeal control is not needed to imitate certain sounds to use them as symbols of referential meaning, singing is impossible without vocal learning. This learning takes place socially, in what might be described as ritual circumstances, and involves developing the ability to control and match other individuals’ use of not only pitch and rhythm, but also dynamic intention (appropriate loudness) and timbre. Also, the fact that certain aspects of meaning attributed to consonants in human speech seem to be independent of cultural influence (Fort et al., 2015) suggests that the evolution of referential communication by means of manipulating the spectral characteristics of sounds had begun before the evolution of vocal learning. It is worth mentioning that human beings are not equally good at imitating all kinds of sounds (Jackendoff & Lerdahl, 2006). Unlike grey parrots, for example, which are equally good at learning to produce species- and hetero-specific sounds (Pepperberg, 2012), people are better at learning to produce species-specific vocalizations. They are also quite good at imitating the calls of animals, particularly imitative hunting calls (e.g., Nazina, 2009), but learning to do so is relatively hard, in comparison to the way in which speech and singing skills appear to be acquired spontaneously. This means that there are some constraints on human vocal learning that enable people to be very good at learning to produce consonants and vowels but not the timbres of musical instruments (or even—for most of us—the timbres of other people’s voices) or environmental sounds. Human beings are also very good at imitating F0 accurately with their voices, and their ability to sustain the vocalization of a particular sound frequency is limited only by their need to breathe.
Such elaborate control of pitch, in terms of both the frequency of the sound and its duration, is not necessary for speech (Bannan, 2012), even in the case of tonal language. Purves (2017) speculates that our sense of tonality evolved as a tool for recognizing conspecific vocalizations, but pitch control requires such costly use of energy that the likelihood of such an elaborate learning system evolving solely to recognize conspecifics is minimal; after all, this can be accomplished by simpler means. The explanation that vocal control of pitch evolved in hominins because of the adaptive value of speech also seems unconvincing.
Human vocal control, which is extraordinary among primates, suggests that pitch must have become an important part of hominin rituals, even as it is today. The essence of ritual is that its form is independent of its function (Merker, 2009). Such independence explains the lack of any connection between the frequencies of sounds, which can be observed and measured, that are interpreted as discrete pitches, and the function of pitch syntax. For example, the identity of a national anthem depends on the sequence of intervals between pitches that, if memorized by members of a particular nation, facilitates national identity. But the melodies of different national anthems (or any other songs reflecting national identity) may be composed of different intervals according to the rules of culture-specific pitch syntaxes. All these anthems, however, fulfil the same – or at least broadly similar – function. Another crucial characteristic of ritual is the importance of being able to recognize the ritual from its form (Merker, 2005, 2009). In this respect the imitation of pitch sequences is even more restrictive than the imitation of word sequences. While replacing one word in a sentence by another does not hamper communication so long as the meaning of the sentence is preserved, a melody in which one note is replaced by another is usually perceived as a different pattern of notes altogether. Thus melodies are processed and stored in memory as sequences of items in which every item must be present and in its correct place; as Tillman and Dowling (2007), and Dowling and Tillmann (2014) found, listeners’ memory for melodies, like poetry but not prose, improves over time.
Hence, a discrete category of pitch—in addition to other categories including rhythm, rhyme, and assonance—seems to be essential to human ritual culture. However, the efficiency of pitch processing as a tool for engaging in ritual depends on the organization of pitches within a culture-specific pitch syntax. After all, not all pitch sequences are remembered equally well (Vuvan et al., 2014). This means that the evolutionary advantage of individuals who participated in rituals producing social bonds must have depended on their privileged recognition of, and better memory for, relationships between pitches rather than isolated categories of pitch. Although it is known that rhesus monkeys (Maccaca mulata) can recognize that two simple tonal melodies are the same provided one of them is a transposition of the other by either one or two octaves, but not if it is transposed by half an octave or one-and-a-half octaves (Wright et al., 2000), studies conducted so far have failed to show that any nonhuman primates are able to recognize pitch syntax, despite the fact that they can use simple syntactic rules restricted to neighboring sounds or affixations to them (Fitch & Hauser, 2004; Ouattara et al., 2009). Nonhuman species of primates are also able to discriminate between changes of pitch direction (Brosch et al., 2004), and to categorize sounds (Wright et al., 1990). These observations allow us to hypothesize that after the split with the chimpanzee lineage, our ancestors must have been able to discriminate pitch even categorically before they started to use pitch as a way of organizing sequences syntactically. If pitch syntax enables easier spectral synchronization during communal singing, and communal singing fosters social bonds, then those hominins who were born with the predisposition to use pitch syntax gained an advantage over those who were less gifted. However, according to classical Darwinism, the appearance of a new trait of an organism is the result of accidental mutation or recombination that has occurred in one individual.
How, therefore, can we explain the adaptive value of the ability to organize pitches syntactically, if no-one else in the community has acquired the ability to use pitch sequences?
A possible explanation is the Baldwinian model of evolution (Podlipniak, 2017) in which a new trait appears in the beginning as a phenotypic adaptation that then comes under genetic control – canalization (Baldwin, 1896a, 1896b; Dor & Jablonka, 2010). In the case of a behavioral trait, a new behavior is elicited by the process of learning (Simpson, 1953). This means that a new behavior must be invented and then shared throughout the whole population. If the learned behavior is adaptive, lasts many generations, and the process of learning is time-consuming, sooner or later an accidental genetic change will appear that enables an individual to learn faster, more easily, and more efficiently in terms of energy consumption (Godfrey-Smith, 2003). This kind of evolution has been proposed as the mechanism that led to the appearance of language grammar (Dor & Jablonka, 2000) and pitch centricity (Podlipniak, 2016). Changizi (2011) has proposed a scenario for the appearance of music that is similar, to some extent, by suggesting that the cultural evolution of music consisted of “harnessing” the mechanisms of our auditory system, the function of which was to identify the actions of other people, albeit with no canalization phase. The Baldwinian scenario proposed here starts with a similar cultural harnessing whereby pitch sequences had to be invented initially by hominins and used, most probably, as a part of social rituals. This invention was possible because of the sensitivity of the hominin auditory system to the harmonic series, which allowed our ancestors to recognize and sense pitch. Such sensitivity is observed among many species of mammals, including primates (Bendor & Wang, 2005; Oxenham, 2018) and is important for the processing of nonverbal affective vocalizations. The volitional control of these vocalizations allowed hominins to sustain pitch levels, leading to the appearance of monotony (cf. Bannan, 2012; 2019). At this stage, however, it would have been hard to vocalize on a monotone; if hominins were able to vocalize within a relatively stable compass of pitches and were also sensitive to the harmonic series, they could have invented simple pitch sequences. Once one hominin had invented such a sequence to be used in rituals, the other members of the tribe would have to have had accepted it, and to imitate it so as to learn it, rather as they learned how to make tools via group imitation (Shipton, 2010).
Learning would have been difficult and time-consuming for individuals with limited vocal control and lacking the ability of contemporary human beings to memorize and recall tonal melodies. Yet they would have had to acquire group-specific rituals, as individuals moved between groups as a result of exogamy, and in response to intergroup competition. Even today, people show that they identify with a particular group by collectively singing the songs that are popular with that group. As reported by Weinstein et al. (2016), such singing can increase individuals’ sense of closeness to other singers even though they are unfamiliar or unknown to them. Pre-lingual hominins probably used rituals based on pitch syntax and would recognize each other as members of their group by their ability to take part in the ritual. If the ritual included singing, then participants who were quicker to learn the ritual pitch sequence would have had an advantage over those who were slower to learn it. Pitch sequences would have differed from one group of hominins to another, according to how many times each pitch occurred in the sequence. This would have imposed specific demands on the members of the group, that is, to recognize and to learn pitch regularities as quickly as possible. This hypothetical scenario meets the conditions of the Baldwinian model of evolution in so far as it took many generations to learn how to identify and memorize pitch sequences. It would have been time-consuming and difficult work, but it fulfilled an important biological function (Baldwin, 1896a, 1896b; Godfrey-Smith, 2003). Sooner or later, however, a genetic predisposition would have appeared, making the learning of pitch sequences instinctive. As those who possessed this novel predisposition would have had advantages over who did not possess it, so a new, instinctive, behavioral trait would have spread throughout the whole population.
Conclusion
In this article I have put forward the argument that pitch syntax is essential to vocal communication although, unlike phonotactics and grammar, it did not evolve because of the function of speech.
Rather, it appeared in the Homo sapiens lineage once hominins had come to possess vocal control of sound frequency. Rituals based on pitch syntax were a tool for enhancing social acceptance and communicated membership of a particular group by providing preconceptual emotional cues. These cues were more subtle than those evoking the emotions experienced during rhythmic entrainment (Becker, 2004), which suggests that the function of pitch syntax may be more elaborate, perhaps encouraging fellow singers to trust each other. Nevertheless, these preconceptual emotional cues probably became new elements of the hominin consciousness. Derived from responses to tonal qualia, awareness of pitch syntax represents one of the final stages in the transition between preconceptual, emotional consciousness and the conceptually complex consciousness that Ginsburg and Jablonka (2019, p. 19) have referred to as the Aristotelian “sensitive soul” and the “rational soul” respectively. Of course, this proposed evolutionary scenario is speculative, and more research is needed to support the notion that pitch syntax had a specific function. First, the socializing effects of tonal and non-tonal (pitch-free) or atonal music could be compared, as the former uses the kind of pitch syntax that is recognized implicitly, and the latter do not. Comparison could be conducted first of all between the socializing effect of tonal music (possessing implicitly recognizable pitch syntax) and music without pitch or atonal music (without implicitly recognizable pitch syntax) could be conducted. If pitch syntax evolved as an elaborate tool for groups of human beings to sing simple tonal melodies together, then such melodies should be more efficient for social bonding than other forms of synchronized sounds. It is worth noting that Western music, based on harmony, is an exception rather the rule (Jackendoff & Lerdahl, 2006); unlike melody, harmony is not universal. It would therefore seem better to explore the evolutionary origins of music for social bonding using simple tunes than complex harmonic music. Conversely, however, it would be useful to investigate the possible role of harmonic factors in the mapping of speech units on to conceptual representations, and musical units onto subtle preconceptual emotional states. Given that sensitivity to the harmonic series is crucial for the recognition of musical pitch and vowels in speech, this could be achieved by comparing the processing of melodic, harmonic and speech sequences. It would also be helpful in determining the evolutionary stage during which the divergence between speech and song took place.
Another field of research that is very promising for identifying the specific function of pitch syntax is genetics. A better understanding of the role of particular genes in the development of cortico-subcortical loops may shed light on the evolutionary origin of human vocal behavior and the possible differences between the processing of pitch in speech and in singing. A better understanding, also, of the role in the processing of pitch syntax and speech of limbic (e.g., Gebauer et al., 2012) and dorsolateral-prefrontal loops (e.g., Schaal et al., 2017) and the auditory cortex (e.g., Kell et al., 2018; Norman-Haignere et al., 2015) could explain to what extent the two kinds of processing are independent. Cross-cultural studies designed to explicate the role of environmental influences in the development of brain specialization related to pitch processing would be most informative in this respect. So far as the prelingual character of pitch syntax is concerned, studies focusing on speech and music deficits in dementia and other conditions are useful. Although some research shows that the loss of speech in dementia can occur without loss of response to pitch and music (Baird & Samson, 2015; Cuddy & Duffin, 2005; Golden et al., 2017), suggesting that the latter may be an older system, further systematic investigation is necessary to explain how and why grammar and pitch are processed in different ways.
If accepted, the argument that pitch syntax is a prelingual evolutionary innovation could change the popular understanding of human musicality. Although music may be viewed from the perspective of Western aesthetic philosophy as a very broad phenomenon that may be composed of any kind of sounds, from a worldwide perspective, music is usually recognized intuitively as something different both from speech and from environmental sounds. Even so, music for untuned percussion or clapping, for example, is recognized as such. All the same, pitch syntax is a universal trait of human vocal communication. Perhaps musical rhythm, pitch syntax, and speech are separate forms of human behavior that are sometimes used together, as in the case of songs, and sometimes separately, as in the case of music for untuned percussion, music in free rhythm, and speech.
Footnotes
Acknowledgements
I should like to thank the reviewers for their helpful comments and suggestions. I should particularly like to thank Jane Ginsborg and Nicholas Bannan, whose advice was invaluable. I would also like to thank Peter Kośmider-Jones for his language consultation.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
