Abstract
Voices are present in most communications. Yet, the literature on voice persuasion is astonishingly limited and fragmented, focusing on certain voice characteristics (e.g. pitch), contexts, and providing mixed results. This research attempts to integrate the various constructs and mechanisms involved in voice persuasion as a result of the cross-fertilization of the disciplines having studied voice (psychoacoustics, cognitive psychology, anthropology, psycho-sociology, marketing, and politics). Study 1 manipulates via acoustic software the key voice characteristics (i.e. pitch, roughness, and brightness) and gender of a speaker heard in a radio advertisement for a neutral, non-gendered product category. Study 2 explores a potential boundary condition of the effects of voice, the presence of context-specific expectations toward the speaker (i.e. gender and competence level), by manipulating the voice of a political candidate. The effects of the voice characteristics are consistent in both contexts: speakers with low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices are the most effective. Speakers with high-pitched, dull, and smooth voices are perceived as the most competent. Finally, speaker gender plays a secondary persuasive role; listener gender only plays a role in the absence of context-specific expectations toward the speaker. Implications for voice and speaker persuasion as well as for voice casting and coaching are discussed.
Introduction
Voices can be heard in most communications. Visible or invisible (i.e. voice-over), belonging to celebrities (e.g. endorsers, political candidates, athletes, or artists) or to common people (e.g. advertising speakers, representatives, coworkers, volunteers, or interviewers), voices are pervasive and multifold. In fact, evidence from psychology suggests that voice is the first thing humans recognize and respond to, even in utero (Kisilevsky et al., 2003). A speaker’s voice can be a powerful persuasion tool that conveys rich emotions and imagery (Mehrabian and Wiener, 1967), influencing listeners’ behavior in many contexts (e.g. marketing: Wiener and Chartrand, 2014; politics: Klofstad, 2016). Therefore, practitioners need to understand how to select the right voice or how to change an existing voice with a voice coach or a sound engineer. This is not an easy task, considering the multidimensional nature of the vocal sound, including the three fundamental notions of pitch, timbre brightness, and timbre roughness (Belin et al., 2011).
Unfortunately, the current state of research offers little guidance. Prior research having actually investigated the persuasive outcomes of a speaker’s voice is astonishingly limited (e.g. approximately five studies in marketing) and highly fragmented. Such studies have focused on different voice characteristics or variables (i.e. pitch: Zoghaib, 2017; Chattonadhyay et al., 2003; Klofstad, 2016; Oksenberg et al., 1986; timbre: Wiener and Chartrand, 2014; speaker gender: Whipple and McManamon, 2002), on different mediators and mechanisms (perceived voice masculinity, Puts et al., 2011; perceived voice arousal, Laukka et al., 2005; perceived speaker competence and warmth, Zuckerman, Hodgins, and Miyake), on different outcomes (i.e. speaker perceptions, attitude, or intentions), and on different contexts (e.g. advertising, political campaign, phone survey, and non-contextualized interpersonal discussion), providing mixed results.
This research aims to suggest an integrative approach to voice persuasion based on the cross-fertilization between the previous disciplines as well as on research on non-vocal speaker persuasion (e.g. endorsement: Albert et al., 2017; Knoll and Matthes, 2017; political communication: Capelli et al., 2012). This approach has enabled the identification of three key voice characteristics (voice pitch, brightness, and roughness), two potential moderating variables (speaker and listener gender), four mediating variables (perceived voice masculinity, perceived voice arousal, perceived speaker competence, and perceived speaker warmth), and two outcome variables (attitude toward the speaker and behavioral intentions).
The first part of the research presents the cross-disciplinary literature review leading to the formulation of the hypotheses. The second part presents the overall methodology, which entails three experiments manipulating the voice of a radio advertisement via acoustic software in three different contexts. The methodology thereby addresses two potential boundary conditions identified in the literature (i.e. absence vs presence of specific speaker expectations; peripheral vs central processing). More precisely, the third part presents the first experiment (Study 1), which explores a context with no specific speaker expectations (i.e. an advertisement for a non-gendered, neutral product). The fourth part presents a second experiment (Study 2), whose context generates some speaker expectations (i.e. an advertisement for a political candidate). Before presenting Study 2, a preliminary study explores a context generating some speaker expectations and central processing (i.e. an advertisement for real political candidates). Implications for research on voice and speaker persuasion as well as for practitioners are discussed in the final part. In particular, recommendations are formulated for voice casting, coaching, and morphing relative to voice gender parity and to the effectiveness of each voice characteristic.
Theoretical background and hypotheses
Prior research on voice is fragmented per discipline, research field, and context. This cross-disciplinary literature review presents the insights and gaps of prior research by following the different parts of the voice persuasion process (Figure 1), leading to the formulation of hypotheses.

The persuasive effects of a speaker’s voice characteristics.
Speaker voice: The importance of pitch, brightness, and roughness
Prior research in psychoacoustics and cognitive psychology agrees upon the fact that the two most perceptually salient elements of voice are pitch and timbre (Baumann and Belin, 2010).
Voice pitch
In acoustic terms, pitch is the most audible and stable frequency (fundamental frequency, F0, in Hertz, Hz) emitted by the vocal folds while vibrating (Titze, 1994). Pitch can vary from 40 Hz for very low-pitched voices to 600 Hz for very high-pitched voices, with an average pitch band of 100–220 Hz (Huang et al., 2001). For instance, James Earl Jones’ (Darth Vador) low-pitched voice is around 85 Hz (Mayew et al., 2013), Sean Connery’s medium-pitched voice is around 150 Hz, and Barbara Streisand’s high-pitched voice is around 220 Hz (Wake Forest University Baptist Medical Center, 2002).
Voice timbre
Timbre represents the way the vocal tract filters or enhances the frequencies emitted by the vocal folds (Titze, 1994). According to the well-documented source-filter model of speech production, the source of the vocal sound (i.e. the vocal folds) and the filter changing this sound (i.e. the vocal tract) are independent (Fant, 1981). Research in music neuropsychology seems to corroborate this model as pitch and timbre are treated as two different pieces of information by different brain areas (Krumhansl and Iverson, 1992).
While the notion of pitch is easy to define in a perceptual way (i.e. perceived highness), there is no satisfying perceptual definition of timbre. The American National Standards Institute (ANSI, 1973) actually defines timbre by the negative as an “auditory sensation in terms of which a listener can judge that two sounds similarly presented and having the same loudness and pitch are dissimilar” (p. 56). This “ill-defined wastebasket category” (Bregman, 1990: 92) leads to a plethora of timbre dimensions in the psychoacoustic literature. Nevertheless, the review of prior studies on timbre reveals two main dimensions, presented in the following points.
Voice brightness
Brightness (equally referred as sharpness, Von Bismarck, 1974) is considered as the most prominent dimension of timbre for many authors (Knoeferle et al., 2015; Von Bismarck, 1974). In acoustic terms, brightness is the centroid (fc, in Hz) of the frequencies emitted by the voice. Bright (vs dull) voices have a higher concentration of frequencies above (below) 1500 Hz (Castellengo et al., 1996).
In perceptual terms, brightness is the sense of acuity or sharpness of a sound, often described as dull, soft, dark, and resonant versus bright, sharp, and not resonant (Bänziger et al., 2014; Caclin et al., 2005; Von Bismarck, 1974). A review of brightness descriptors indicates that, most often, two descriptors are mixed to describe or measure the perceived brightness, namely dull and soft versus bright and sharp.
Voice roughness
The second most salient and most studied dimension of timbre (Knoeferle et al., 2015; Von Bismarck, 1974) is roughness. From an acoustic point of view, roughness results from irregular vocal fold vibrations (i.e. modulation, also referred as vibrato or tremolo), which are characterized by the rate of the modulation (i.e. modulation frequency, fam, in Hz) and the extent of the modulation (i.e. modulation depth, m, ranging from 0% to 100%; Fastl and Zwicker, 2007).
From a perceptual point of view, roughness is the perceived harshness, abruptness, and irregularity of a voice, often described as smooth or regular versus rough or irregular (Rabinov et al., 1995; Von Bismarck, 1974). Voice roughness is generally considered as a key element of voice quality, along with, among others, voice creakiness (i.e. vocal fry), hoarseness, and jitter (Kempster et al., 2009).
Other less studied dimensions of timbre can be found in the literature (fullness, richness, warmth, color, density; Moravec and Štepánek, 2003; Von Bismarck, 1974), but many authors have shown that they can be aggregated with the previous dimensions of timbre (i.e. brightness and roughness) after having conducted reduction analyses (Von Bismarck, 1974). Notably, some studies aggregate timbre dimensions even further to form supra categories such as voice quality. This is not the objective pursued by this research.
To summarize, pitch, brightness, and roughness constitute the key characteristics of human voice (Baumann and Belin, 2010; Belin et al., 2011; McAdams et al., 1995). Surprisingly, to our knowledge, no research has considered the effects of these voice characteristics simultaneously, apart from prior research on voice sensory perception. Conversely, prior voice persuasion research has mostly focused on the voice’s pitch or gender.
Voice gender
A speaker’s voice characteristics contribute to the identification of a speaker and, in particular, his or her gender (Pernet and Belin, 2012; Puts et al., 2012). Specifically, numerous studies in cognitive psychology indicate that listeners can correctly identify a speaker’s gender solely based on his or her voice characteristics (Pernet and Belin, 2012). Therefore, the manipulation of the voice characteristics can induce the perception of a specific gender; this manipulation is, for that matter, at the heart of most treatments and voice coaching exercises for transgender patients (Pernet and Belin, 2012).
More specifically, prior research in anthropology and cognitive psychology indicates that there is a biological dimorphism of the male and female vocal apparatus (Belin et al., 2011; Puts et al., 2012). Males tend to have bigger vocal folds and a longer vocal tract (Pernet and Belin, 2012), which results in lower pitch and lower brightness.
The link between roughness and gender is less obvious. There is no biological reason linking roughness and gender but some sociocultural factors (Eddins and Shrivastav, 2013). Timbre dimensions such as roughness or creakiness would be “social markers of maleness” (Klatt and Klatt, 1990: 821) employed to sound more socially desirable by both male and female subjects (Eddins and Shrivastav, 2013). For instance, several studies have found that creaky voices were more employed by male subjects (Podesva, 2013). However, there is a structural trend in young female adults to use a creaky, rough voice, initiated two decades ago and then perpetuated by numerous female celebrities such as Kim Kardashian, Katy Perry, Britney Spears, or Zooey Deschanel (Fessenden, 2011; Mendoza-Denton, 2007; Yuasa, 2010).
To conclude, the manipulation of a speaker’s gender and voice characteristics seems somewhat redundant, especially if pitch and brightness are manipulated. Therefore, we expect that the effects of speaker gender will play a redundant role with those of the voice characteristics. The following hypotheses thus focus on a speaker’s voice characteristics.
The persuasive role of voice masculinity
A first body of research suggests that voice persuasion can be explained by the perceived masculinity of the voice, which would affect perceptions as well as attitude and intentions.
Antecedents of perceived voice masculinity
According to prior research in anthropology, people share a common, implicit knowledge of the aforementioned physiological and sociocultural factors differentiating male and female voices (Pernet and Belin, 2012; Puts et al., 2012). Because of this knowledge, people implicitly associate low- (vs high-) pitched and dull (vs bright) voices with masculinity (Feinberg et al., 2008; Puts et al., 2012). We expect the same association phenomenon for roughness and anticipate that smooth (vs rough) voices will be perceived as more masculine based on recent research in anthropology and linguistics (Fessenden, 2011; Mendoza-Denton, 2007; Yuasa, 2010).
Outcomes of perceived voice masculinity: The role of gender stereotypes
Prior research in psycho-sociology indicates that perceived voice masculinity has a positive effect on perceived competence and attitude toward the speaker (Ko et al., 2009; Zuckerman et al., 1993) because of gender stereotypes (Stern, 1988).
In addition, prior research on speaker persuasion indicates that attitude toward the speaker and speaker competence also directly affect behavioral intentions (Knoll and Matthes, 2017). Thus, perceived voice masculinity could trigger positive hierarchical, cognitive, affective, and conative effects. These positive effects would occur even if the voices perceived as masculine (vs feminine) are also perceived as less warm (Berry, 1992; Zuckerman et al., 1993).
One study has found evidence of such hierarchical effects for pitch. Political candidates with low- (vs high-) pitched voices are perceived as more competent, preferred, and receive more votes (Klofstad, 2016).
To our knowledge, the effects of brightness and roughness on speaker persuasion have not been studied. As we expected dull (vs bright) and smooth (vs rough) voices to be perceived as more masculine, such voices should induce more favorable responses, except for perceived warmth.
Outcomes of perceived voice masculinity: The role of sexual preferences
In addition, the implicit knowledge of the physiological differences between male and female voices could not only trigger hierarchical persuasive effects but also have direct effects on attitude and intentions. In particular, research in anthropology on sexual selection theory indicates that female listeners would prefer speakers with masculine voice characteristics, such as low-pitched and dull voices, whereas male listeners would prefer speakers with feminine voice characteristics, such as high-pitched and bright voices (Pernet and Belin, 2012; Puts et al., 2011; Puts et al., 2012; Re et al., 2012). Prior studies suggest that this sexual selection process might also apply to voice roughness. For instance, creaky male voices would induce the most favorable responses from female listeners (Wiener and Chartrand, 2014).
However, little research has actually managed to show that male listeners prefer feminine voice characteristics. For instance, voice creakiness has no effect on male listeners (Wiener and Chartrand, 2014). The review of prior research on voice persuasion having studied the effects of listener gender seems to indicate that listener gender moderates the intensity of the persuasive outcomes but not their direction (i.e. positive or negative).
Implications relative to the role of voice masculinity
To conclude, the previous body of research considers that perceived voice masculinity plays an important role in voice persuasion. This approach is further referred to as the voice masculinity hypothesis.
More precisely, previous studies on gender stereotypes suggest that low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices have a positive effect on perceived competence, a negative effect on perceived warmth, and positive effect on attitude toward the speaker and behavioral intentions. According to prior research on sexual selection, the two latest outcomes would be stronger (weaker) for female (male) listeners.
However, several studies contradict this voice masculinity hypothesis. First, telemarketers with high- (vs low-) pitched voices have positive effects on perceived competence, attitude, and intentions (for both male and female speakers: Chebat et al., 2007; for female speakers: Oksenberg et al., 1986). Moreover, Anolli and Ciceri (2002) have found that, during a rendezvous, the male partners who started talking with a higher pitch were preferred and had greater mating chances with female partners.
The previous body of research does not allow us to formulate hypotheses relative to the effects of a speaker’s voice characteristics on speaker perception and persuasion. However, the following hypothesis can be formulated relative to the effects of voice on perceived voice masculinity:
H1. Low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices are perceived as more masculine.
To conclude, the previous mixed results suggest that the persuasive effects of a speaker’s voice characteristics might not – or not entirely – be explained by the notion of voice masculinity. According to Berlyne’s (1974) theory of aesthetic preference, the attitude toward a sensory cue largely depends on the arousal (i.e. a subjective sense of energy, Mehrabian and Russell, 1974) it induces. Thus, the arousal induced by voices could explain some of the previous results, as exposed in the following point.
The persuasive role of voice arousal
Antecedents of voice arousal
There is a consensus among research in psychoacoustics and cognitive psychology indicating that high (vs low), bright (vs dull), and rough (vs smooth) voices are more arousing (Laukka et al., 2005; Scherer, 1986), that is to say, inducing a subjective sense of energy (Mehrabian and Russell, 1974). In this research, the subjective sense of energy induced by a speaker’s voice characteristics is further referred to as “voice arousal.”
Outcomes of voice arousal: The role of acoustic inferences
In addition, arousing voices such as high (vs low), bright (vs dull), and rough (vs smooth) voices would induce stronger perceptions of excitement and joy and weaker perceptions of calm, sadness, and warmth (Banse and Scherer, 1996; Bänziger et al., 2014; Breitenstein, Van Lancker and Daum, 2001; Laukka et al., 2005). Thus, listeners would infer meanings based on the acoustic properties of voices. Such inferences can also be found for environmental sounds. For instance, people make inferences relative to an object’s size based on the sound it makes (Lowe and Haws, 2017).
Being able to induce excitement could be valuable for a speaker in certain contexts. For instance, in the context of a phone survey, interviewers with high-pitched voices induce more favorable perceptions of enthusiasm (Oksenberg et al., 1986) and competence as well as lower refusal rates (Chebat et al., 2007; Oksenberg et al., 1986). Thus, in a context of advertising clutter, listeners might not expect a specific competence level in a speaker, except for delivering information in an efficient way.
Outcomes of voice arousal: The role of aesthetic preference
According to Berlyne’s (1974) theory of aesthetic preference, people would prefer moderately arousing sounds. Thus, voice arousal would have an inverted U-shaped effect on attitude toward the voice and, potentially, on attitude toward the speaker as well as on behavioral intentions (respectively, based on the Affect Transfer Theory, MacKenzie et al., 1986, and the Theory of Reasoned Action, Fishbein and Ajzen, 1975).
In addition, prior research has shown that high- (vs low-) pitched, bright (vs dull), and rough (vs smooth) voices have a positive effect on perceived voice arousal. Thus, a voice with heterogeneous levels of pitch, brightness, or roughness (e.g. with low pitch, high brightness, and high roughness, that is, a low, bright, and rough voice) would induce a moderate voice arousal, and, potentially, have the most favorable effects on attitude toward the speaker and on behavioral intentions. This effect would, however, only occur if there is a significant interaction effect between the voice characteristics. Potential interaction effects will be tested in the results section of each experiment.
Implications relative to the role of voice arousal
To conclude, the previous body of research considers that voice arousal plays an important, positive role in voice persuasion. This approach is further referred to as the voice arousal hypothesis.
More precisely, previous studies on acoustic inferences suggest that high (vs low), bright (vs dull), and rough (vs smooth) voices have a positive effect on perceived speaker competence, a negative effect on perceived speaker warmth, and positive effects on attitude toward the speaker as well as on behavioral intentions. According to prior research on aesthetic preference, a voice characteristic could moderate the effects of another characteristic on the two latest outcomes.
The voice arousal hypothesis (high, bright, and rough voices are more effective) contradicts the voice masculinity hypothesis (low, dull, and smooth voices are more effective). Again, the previous literature review is inconclusive relative to the effects of a speaker’s voice characteristics on speaker perceptions and speaker persuasion. However, there is a consensus on the effects of the voice characteristics on voice arousal:
H2. High- (vs low-) pitched, bright (vs dull), and rough (vs smooth) voices have a positive effect on voice arousal.
To conclude, prior research on voice persuasion as a whole has provided mixed results. However, most of the studies presented in the review have explored different contexts. Considering the role of context in voice persuasion might reconcile the voice masculinity and voice arousal hypotheses.
The role of context in voice persuasion
Contextual expectations toward the speaker
Several studies have shown that the context in which a speaker is heard could play an important role in the persuasion process. For instance, voice persuasion would only occur when a specific gender or competence level is expected in a speaker (Whipple and McManamon, 2002). More precisely, a female (vs male) voice is more effective in an advertisement for a female perfume in a self-gift scenario but a male (vs female) voice is more effective for a female perfume in a gift-buying scenario; voice gender would have no effect for a non-gendered perfume.
Thus, the expectations toward a speaker’s gender and competence level could change the direction of the effects of a speaker’s voice. In addition, the presence (vs absence) of speaker expectations could affect the significance of a speaker’s voice. Therefore, the presence (vs absence) of speaker expectations could constitute a boundary condition of voice persuasion.
Informational context of the advertisement
In addition, prior research on persuasion (Petty et al., 1983) has shown that peripheral persuasion (e.g. of music and voices) only occurs when the central information is not processed. Chattonadhyay et al. (2003) have found evidence of such a phenomenon in voice persuasion. Voice persuasion was significant when the central message was not processed (because of a fast speech rate) and non-significant when the central message was processed (because of a normal and moderate speech rate).
Moreover, the persuasion of a source (e.g. an endorser) is more effective when the source is unknown (Cacioppo et al., 1992; Knoll and Matthes, 2017) because it is difficult to change prior associations. To summarize, the peripheral persuasion of a source (e.g. a speaker’s voice) would be more effective when undisturbed by central processing, especially the processing of prior associations with the source.
Previous studies indicate that the type of processing (peripheral vs central) could affect the significance of voice persuasion. Therefore, peripheral (vs central) processing would constitute a second boundary condition of voice persuasion. Voice persuasion and peripheral processing would occur with unknown, fast-speaking voices.
Implications for the voice persuasion process
Previous studies on speaker persuasion suggest that the presence of speaker expectations (presence vs absence) and the type of processing (peripheral vs central) represent two boundary conditions of voice persuasion. However, a large number of studies in psychology indicate that a speaker’s voice has a great influence on our emotions and perceptions, which can be stronger than a speaker’s actual words (Mehrabian and Wiener, 1967). Thus, we anticipate that voice persuasion will have no boundaries. This assumption will be tested as a result of the methodology presented in the following point.
Nevertheless, the previous studies suggest that speaker expectations can change the effects of a speaker’s voice. Speaker expectations might explain the mixed results of prior research on voice persuasion and reconcile the voice masculinity and the voice arousal hypotheses. Specifically, in the absence (presence) of specific expectations toward a speaker’s gender or competence level, voice arousal (perceived voice masculinity) would play a greater persuasive role.
Following this line of reasoning, some hypotheses can be formulated concerning the direct effects of a speaker’s voice characteristics.
In a context where no specific gender or competence level is expected in a speaker,
H3.1. High- (vs low-) pitched, bright (vs dull), and rough (vs smooth) voices have a positive effect on perceived speaker competence.
H4.1. High- (vs low-) pitched, bright (vs dull), and rough (vs smooth) voices have a negative effect on perceived speaker warmth.
H5.1. High- (vs low-) pitched, bright (vs dull), and rough (vs smooth) voices have a positive effect on attitude toward the speaker.
We anticipate that a context where no specific speaker gender or competence level is expected gives the opportunity for the implicit sexual selection process to occur. As we have considered that high- (vs low-) pitched, bright (vs dull), and rough (vs smooth) voices were perceived as more feminine (H1), we anticipate the following moderating effect:
H6.1. High-pitched, bright, and rough voices have a more positive effect on the attitude toward the speaker for male than for female listeners (a); conversely, low-pitched, dull, and smooth voices have a more positive effect on attitude toward the speaker for female than for male listeners (b).
H7.1. High- (vs low-) pitched, bright (vs dull), and rough (vs smooth) voices have a positive effect on behavioral intentions.
In a context where a male gender or a high competence level is expected in a speaker,
H3.2. Low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices have a positive effect on perceived speaker competence.
H4.2. Low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices have a negative effect on perceived speaker warmth.
H5.2. Low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices have a positive effect on attitude toward the speaker.
We anticipate that a context where a specific speaker gender is expected no longer gives the opportunity for the sexual selection process to occur. Thus,
H6.2. Listener gender does not moderate the effects of a speaker’s voice characteristics on attitude toward the speaker.
H7.2. Low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices have a positive effect on behavioral intentions.
All the indirect effects mentioned in the literature and depicted in Figure 1 will be tested and reported in the results section of each experiment in order to explain the results. However, we will not formulate hypotheses on these links.
In addition, the objective of Figure 1 is to display the relationships and hypotheses tested in this research. However, considering the exploratory nature of this research, a global model has not been tested. Future research could test the model suggested in this research in different contexts and eventually add context-specific mediators and moderators.
Overall methodology
Overview
Study 1 explores the effects of a speaker’s voice in a radio advertisement for a male and female fashion brand. According to a pretest, this context does not generate any expectations toward a speaker’s gender or competence level.
Study 2 then explores a first boundary condition (i.e. the presence of expectations toward a speaker’s gender or competence level) by replicating Study 1 in a context with specific expectations toward a speaker’s gender or competence level: a radio advertisement for a political candidate. In this context, listeners tend to expect a male, competent political persona according to a pretest and to prior research in politics (Klofstad, 2016).
A preliminary study is conducted before Study 2 in order to explore another potential boundary condition identified in the literature (i.e. the type of processing), operationalized with the constructs of speaker familiarity and speech rate. While Study 1 and Study 2 resort to anonymous voices talking at a moderate rate, this preliminary experiment studies the voices of well-known presidential candidates speaking at a slow rate. This context favors central processing, especially of prior associations with the speaker.
If the presence of specific speaker expectations truly is a boundary condition, then only Study 2 and the preliminary study should induce significant results (speaker expectations are absent in Study 1).
If peripheral processing truly is a boundary condition, then none of the studies should have significant results. Indeed, all the voices employed in the experiments talk at a moderate or slow rate, which facilitates central processing.
Design and participants
The experiments all manipulate the pitch, brightness, and roughness of voices (2 × 2 × 2) heard in a radio advertisement. Study 1 also manipulates speaker gender (2 × 2 × 2 × 2). However, the results of Study 1 indicate that the interaction between a speaker’s gender and voice characteristics is not significant. In addition, the results indicate that speaker gender plays a redundant and secondary role in comparison with the voice characteristics. Thus, Study 2 and the preliminary study measure the perceived gender but do not manipulate it.
In the three experiments, the respondents are exposed to all versions of a radio advertisement. Thus, the experiments all possess a within-subjects design, for multiple reasons. First, it was important to have the same design between experiments to be able to compare results. In addition, the natural exposure to a presidential candidate’s speech is in batch. All candidates have their share of voice. Moreover, audio perception can differ a lot between participants, based on numerous demographic factors (gender, age; Kellaris and Altsech, 1992) as well as on individual preferences (e.g. need for stimulation, Daucé and Rieunier, 2002). A within-subjects design limits these potential influences. Consequently, numerous studies involving listening tasks resort to this design (e.g. Knoeferle et al., 2015). However, because of the within-subjects design, it was important to suppress order effects by automatically randomizing the order of presentation of the advertisements in all experiments.
Manipulations
In the three experiments, respondents are exposed to different radio advertisements, which are strictly similar except for the voice. In Study 1 and Study 2, the different voice versions are obtained via acoustic software following the same procedure. In Study 2’s preliminary study (with well-known voices), the voice versions are obtained by measuring the voice characteristics with objective measures and by coding them according to two levels of pitch, brightness, and roughness.
In Study 1 and Study 2, the different versions of the voice are obtained in the following manner. First, the voice of a voice actor is recorded in a soundproof room using a professional microphone and solid-state recorder. Second, two musicologists verify aurally and visually via acoustic software (Audacity and Melodyne) that the message clarity is high (clear enunciation, pronunciation, and stresses; absence of non-speech noise). Third, the voice is then averaged via acoustic software, which means that the voice characteristics are modified in order to obtain average levels (i.e. average pitch, brightness, and roughness, among other characteristics) before re-synthetizing the desired voice characteristics. Fourth, to obtain two levels of pitch, the frequencies of the original voice (F0 = 150 Hz) are leveled up (+70 Hz; F0 = 220 Hz) or down (−60 Hz; F0 = 90 Hz), which remains in line with the human average pitch band (Huang et al., 2001). To obtain two levels of brightness, the two previous stimuli are further modified by boosting the intensity of the frequencies below (above) the cut-off frequency of 1500 Hz by 12 dB. To obtain two levels of roughness, a vibrato is applied to the four previous stimuli (fam = 70 Hz, m = 100%).
Study 1 also manipulates speaker gender by recording two voices (a male and a female voice) at the beginning of the procedure. Study 2 only uses one voice. Therefore, a total of 16 and 8 voice versions are obtained in Study 1 and 2, respectively.
Procedure and measures
The same procedure is used for all experiments. First, listening instructions are provided. Studies including listening tasks are faced with the problem of varying listening levels between respondents. To tackle with this issue, the following measures were taken: (1) sound source settings – the online audio players included in the questionnaires of this research were set at a loud but comfortable level of 80 dB, above the universally acceptable range of sound level (Sato et al., 2007); (2) listening conditions – to control the listening conditions as much as possible, the participants were instructed in the invitation email that they had to complete the study on their computer without using headphones; (3) healthy sample – the respondents did not report known auditory problems. Those who did were excluded from the sample; (4) listening level adjustment – prior research indicates that respondents’ equipment as well as the nature of instructional set influence listener performance (Hochberg, 1975). It is higher when listeners are asked to adjust their speakers for comfortable speech intelligibility (vs speech loudness). After listening to a first sound (a voice pronouncing a word), respondents were asked to adjust their speaker to a comfortable speech intelligibility; (5) listening task – the respondents completed a second sound test (they had to recognize the sound of a ringtone). Those who failed the test were excluded from the sample; (6) replay option – during the questionnaire, the respondents could play the audio extract as many times as needed by clicking on the play button (most listeners only played it once with an average of 1.13 times).
This procedure still contains certain biases but controls the major issues associated with listening tasks based on psychoacoustic research. Most prior studies involving listening tasks have focused on the first two measures.
After the listening instructions, a specific scenario is described for each experiment. After having listened to a first advertisement, the variables are measured using well-established scales (Supplemental Appendix 1) successfully pre-tested on 10 respondents. This process is repeated for all the advertisements, which are automatically randomized. It is not possible for the respondents to go back to previous pages of the questionnaire, and respondents received a financial incentive based on the duration of the questionnaire (on average: 10 minutes) and on the average incentive given by international panels for their panelists (approx. 0.1€ per minute, source: Research Now).
Analyses
The three experiments resort to repeated-measures univariate analyses of variance (ANOVAs for repeated measures in SPSS v. 19.0.0) to assess the extent to which a speaker’s voice pitch, brightness, and roughness have an influence on each measured variable. This analysis is employed because the dependent variables are measured on multiple occasions (Tabachnick and Fidell, 2007). The post hoc comparisons of the repeated-measures ANOVAs are corrected with the Bonferroni procedure. This procedure adjusts the p values to account for multiple testing.
In addition, a mixed ANOVA (within-subjects factors: voice characteristics; between-subjects factor: listener gender) was conducted to assess whether listener gender moderates the effects of a speaker’s voice characteristics on attitude toward the speaker.
Finally, the effects between the intermediate variables were tested with multiple regressions. Potential non-linear effects were assessed with a curve estimation procedure (SPSS v. 19.0.0) including linear, logarithmic, quadratic, and cubic regression models.
Study 1: Voice persuasion without specific speaker expectations
Design and participants
The objective of this first experiment is to test the effects of a speaker’s pitch, brightness, roughness, and gender (2 × 2 × 2 × 2) in a first context, an advertisement for a male and female brand of clothes.
A sample of respondents (N = 597; 53% female) was exposed via an online questionnaire to the 16 experimental conditions (within-subjects design). The respondents were recruited by an international Internet market research company (i.e. Research Now, awarded the international quality leader in digital data collection by the 2016 Annual Survey of Market Research Professionals; source: marketresearchcareers.com) via a quota sampling method based on the socio-demographic and geographic distribution of the French population in 2012 (source: official national statistics issued by the National Institute of Statistics and Economic Studies, INSEE, 2012) with a confidence interval of 95%.
Stimuli
Manipulations
The respondents were exposed to different radio advertisements that were strictly similar except for the voice. To create the 16 voice versions, a male and a female voice actor were recruited. They recorded the same sentence with the same tone and at the same moderate speech rate (150 syllables per minute (SPM); total duration: 9 seconds): “Jarro. Clothes, accessories, beauty, and more for men and women. Discover our new collection on www.jarro.com and in our stores in Paris.” These voice actors were siblings and had perceptually similar voices, with average characteristics (medium-pitched, with an average brightness and roughness). Their voice characteristics were, however, averaged via acoustic software. The 16 versions were obtained as described in the overall methodology.
Pretests
The neutrality of each component of the advertisement (product category, brand name, message, and the two original voices) was pretested on several groups of respondents.
The neutrality (in terms of expectations - related to the gender or the competence level of the speaker) was measured via an open-ended item (“Describe in a few words what comes to your mind about …” Clothes/The brand Jarro/This message/This voice) and two closed items (“To what extent do you associate [Clothes/The brand Jarro/This message/This voice] with a certain gender?,” “To what extent do you associate [Clothes/The brand Jarro/This message/This voice] with a certain competence level?,” 7-point Likert-type scales).
The free associations were coded by two independent researchers as non-neutral (vs neutral) whenever they contained words from the same semantic field as gender, sex, or competence. The percentage of neutral associations formed a neutrality score. The mean of the two items was calculated to form a second neutrality score. Each advertising cue had satisfying scores of neutrality (product category: N = 28, score = 92%, M = 6.6; brand name: N = 28, score = 91%, M = 6.7; voices: N = 10, score = 96%, M = 5.8; and message: N = 42, score = 95%, M = 6.7).
In addition, the realism of the stimuli was pretested. First, we verified that the voices were perceived as belonging to different persons (“These voices belong to different persons,” 7-point Likert-type scale, N = 17, M = 6.8). In addition, the advertisements were perceived as realistic (“This advertisement is credible,” 7-point Likert-type scale, N = 52, M = 6.3; Table 1).
Study 1 – description of the vocal stimuli.
Hz: Hertz (frequency measure); <CF/>CF: frequencies below/above the cut-off frequency of 1500 Hz; dB: decibel (loudness measure).
Procedure and measures
After the general and listening instructions (cf. Overall methodology), participants were told that a brand required their help in order to select its coming advertisement. This type of study is not unusual for such respondents (i.e. panelists), who are frequently recruited for completing advertising pre- and post-tests. After each advertisement, perceived pitch, perceived brightness, perceived roughness, perceived speaker gender, perceived voice masculinity, voice arousal, perceived speaker competence, perceived speaker warmth, attitude toward the speaker, and behavioral intentions were measured (Supplemental Appendix 1).
Results
Manipulation checks
The multiple comparisons of four-way (pitch × brightness × roughness × speaker gender) repeated-measures ANOVA confirm the efficacy of the manipulation (Table 2): low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices are perceived as significantly less high, bright, and rough, respectively (pitch: MLow pitch = 2.48 vs MHigh pitch = 3.73, F(1, 596) = 1182.273, p < 0.001; brightness: MDull = 4.54 vs MBright = 5.35, F(1, 596) = 395.727; p < 0.001; roughness: MSmooth = 2.78 vs MRough = 3.48, F(1, 596) = 338.157, p < 0.001).
Study 1 – persuasive effects of a speaker’s voice characteristics in an advertisement for a non-gendered product.
ANOVA: analysis of variance.
n = 597; the mean values in boldface are higher than others within the same condition.
p < 0.05; **p < 0.01; ***p < 0.001.
Male (vs female) voices are significantly more attributed to a male gender (p < 0.001). However, the results of a follow-up multiple regression analysis including pitch, brightness, roughness, and speaker gender into the regression model indicate that speaker gender is the worse predictor of voice gender attribution. Its effect actually becomes non-significant in comparison with the other factors (β = −0.01, t(596) = −0.61, p = 0.545).
In addition, the results indicate that voice brightness moderates the effect of voice pitch on perceived voice pitch (F(1, 596) = 210.030, p < 0.001; Figure 2). When a speaker’s voice is high-pitched, it is perceived as higher pitched with a bright timbre (MHigh pitch × bright = 4.41) than with a dull one (MHigh pitch × dull = 3.31). In the following analyses, we tried to observe whether this interaction had any impact on voice persuasion.

Study 1 – interactive effects involved in voice persuasion.
Effects of speaker voice on voice perceptions
The results of three-way repeated-measures ANOVAs (Table 2) indicate that low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices are perceived as more masculine (pitch: MLow pitch = 4.32 vs MHigh pitch = 3.37, F(1, 596) = 595.034, p < 0.001; brightness: MDull = 3.88 vs MBright = 3.82, F(1, 596) = 4.098; p = 0.04; roughness: MSmooth = 4.26 vs MRough = 3.46, F(1, 596) = 303.785, p < 0.001), as expected (H1). There is no interaction effect of the voice characteristics on perceived voice masculinity.
Increasing pitch, brightness, and roughness levels have positive effects on perceived voice arousal (pitch: MLow pitch = 2.80 vs MHigh pitch = 3.56, F(1, 596) = 364.247, p < 0.001; brightness: MDull = 3.14 vs MBright = 3.22, F(1, 596) = 12.509; p < 0.001; roughness: MSmooth = 2.95 vs MRough = 3.39, F(1, 596) = 119.815, p < 0.001), as expected (H2). There is no interaction effect of the voice characteristics on perceived voice arousal.
Direct effects of speaker voice on speaker perception and persuasion
In a context deprived of specific speaker expectations, we have expected arousing voices (i.e. high, bright, or rough) to be more effective than less arousing ones (i.e. low, dull, or smooth), although they weaken perceived speaker warmth.
Contrary to our expectations (H3.1), there is a significant interaction effect between voice pitch and roughness on perceived speaker competence (F(1, 596) = 29.530, p < 0.001). Follow-up simple effect tests indicate that, when a speaker has a smooth voice, a high pitch leads to greater perceived speaker competence than a low pitch (MSmooth × High pitch = 3.64 vs MSmooth × Low pitch = 3.16, p < 0.001). The other simple effects are not significant. Moreover, dull (vs bright) voices are perceived as slightly more competent (MDull = 3.29 vs MBright = 3.13, F(1, 596) = 8.656; p = 0.01). The predicting contribution of brightness is, however, the lowest of the voice characteristics (brightness: β = −0.04 vs 0.10 for pitch and −0.09 for roughness).
As expected (H4.1), high- (vs low-) pitched, bright (vs dull), and rough (vs smooth) voices have negative effects on perceived speaker warmth (pitch: MLow pitch = 2.82 vs MHigh pitch = 2.52, F(1, 596) = 65.123, p < 0.001; brightness: MDull = 2.87 vs MBright = 2.34, F(1, 596) = 173.890; p < 0.001; roughness: MSmooth = 2.80 vs MRough = 2.55, F(1, 596) = 35.486, p < 0.001).
In addition, high- (vs low-) pitched, bright (vs dull), and rough (vs smooth) voices have negative effects on attitude toward the speaker (pitch: MLow pitch = 4.07 vs MHigh pitch = 3.77, F(1, 596) = 51.574, p < 0.001; brightness: MDull = 4.16 vs MBright = 3.51, F(1, 596) = 222.540; p < 0.001; roughness: MSmooth = 4.26 vs MRough = 3.58, F(1, 596) = 189.420, p < 0.001) as well as on behavioral intentions (pitch: MLow pitch = 3.92 vs MHigh pitch = 3.56, F(1, 596) = 85.422, p < 0.001; brightness: MDull = 4.00 vs MBright = 3.31, F(1, 596) = 264.350; p < 0.001; roughness: MSmooth = 4.10 vs MRough = 3.39, F(1, 596) = 247.328, p = 0.02), which contradicts H5.1 and H7.1. The interactive effects of the voice characteristics on attitude and intentions are non-significant.
To summarize, high, bright, and rough voices have negative effects but for one (i.e. high-pitched, smooth voices induce the highest mean for perceived speaker competence). In addition, high, bright, and rough voices are also the most arousing and perceived as the least masculine. The analyses of the indirect effects might provide some explanations: Are high, bright, and rough voices less effective because they are perceived as less masculine, too arousing, or both? Is it because of another reason? Why is the effect on perceived speaker competence different from the others (curvilinear)?
Indirect persuasive effects: The role of perceived voice masculinity
Prior literature has suggested a positive effect of perceived voice masculinity on perceived speaker competence, a negative effect on perceived speaker warmth, and a positive effect on attitude toward the speaker.
Contrary to expectations, a curve estimation procedure reveals an inverted U-relationship between perceived voice masculinity and perceived speaker competence (R2 = 0.02 vs R2 = 0.01 for the linear regression model, F(2, 595) = 72.989, p < 0.001) and between voice masculinity and perceived speaker warmth (R2 = 0.07 vs R2 = 0.05 for the linear regression model, F(2, 595) = 241.868, p < 0.001). A multiple regression analysis indicates that listener gender moderates the effect of perceived voice masculinity on attitude toward the speaker (β = 0.14, t(596) = 2.78, p = 0.005).
To summarize, speakers with a masculine voice are preferred, and those with a moderately masculine voice are perceived as the most competent and warm. The positive effect of voice masculinity on attitude toward the speaker might explain the positive effects of low, dull, and smooth voices (they are more masculine). In addition, the effect of voice masculinity on perceived competence (curvilinear) might explain why high and smooth voices were perceived as the most competent (they are moderately masculine). However, the effect of voice masculinity on speaker warmth (curvilinear) does not explain the direct effects of the voice characteristics on warmth (linear). Yet, warmth is a crucial element of social judgment (Fiske et al., 2007). Therefore, the previous results suggest that perceived voice masculinity does not fully explain voice persuasion. Another piece of the puzzle could be voice arousal, as we had expected in a context with no specific speaker expectations.
Indirect persuasive effects: The role of perceived voice arousal
Prior research has suggested a positive effect of voice arousal on perceived speaker competence, a negative effect on perceived speaker warmth, and a positive effect (or, possibly, a curvilinear effect, theory of aesthetic preference) on attitude toward the speaker.
Contrary to expectations, a curve estimation procedure testing the effect of perceived voice arousal on perceived speaker competence shows that this relationship is curvilinear (inverted U-shaped: R2 = 0.02 vs R2 = 0.00 for the linear regression model, F(2, 595) = 70.473, p < 0.001).
As expected, multiple regressions indicate a significant negative effect of perceived voice arousal on perceived speaker warmth (β = −0.06, t(596) = −4.57, p < 0.001). The curve estimation procedure shows that the linear model has a more superior contribution than the quadratic model (respectively, R2 = 0.01 vs R2 = 0.00).
In addition, contrary to expectations, the results indicate a significant negative effect of perceived voice arousal on attitude toward the speaker (β = −0.22, t(596) = −17.99, p < 0.001). The curve estimation procedure shows that the quadratic model has a weaker contribution than the linear one.
The previous results indicate that voice arousal mostly has negative effects but for one: voice arousal has a curvilinear effect on perceived speaker competence. Thus, voice arousal seems to explain all the direct effects of speaker voice (contrary to voice masculinity): high, bright, and rough voices are the most arousing, and the least effective, except for their effect on perceived speaker competence (high-pitched and smooth voices are moderately arousing and induce the highest mean of perceived speaker competence).
To conclude, it seems that high, bright, or rough (vs low, dull, or smooth) voices are less effective because they are too arousing, not because they are less masculine.
Indirect persuasive effects: The role of speaker perceptions
The results relative to the speaker perception and persuasion are fully consistent with prior research on speaker persuasion. More precisely, there are positive effects of perceived speaker competence and warmth on attitude toward the speaker (competence: β = 0.31, t(596) = 25.86, p < 0.001; warmth: β = 0.35, t(596) = 29.65, p < 0.001) as well as on behavioral intentions (competence: β = 0.25, t(596) = 20.59, p < 0.001; warmth: β = 0.42, t(596) = 36.05, p < 0.001). In addition, there is a significant effect of attitude toward the speaker on behavioral intentions (β = 0.72, t(596) = 82.53, p < 0.001).
The previous results show that all the direct and indirect effects of a speaker’s voice characteristics are significant, which indicates the presence of multiple mediations in the sense of Baron and Kenny (1986; a, b, and c paths are significant for each mediation). These mediations and, more generally, the relationships forming a potential voice persuasion model were not further investigated given the exploratory nature of this research.
Moderating effects of speaker gender
A four-way (pitch × brightness × roughness × speaker gender) repeated-measures ANOVA (Table 2) indicates that speaker gender does not moderate the effects of the voice characteristics on attitude toward the speaker. However, speaker gender has a main effect on attitude toward the speaker (MMale speaker = 3.91 vs MFemale speaker = 3.68, F(1, 596) = 14.468; p < 0.001).
A follow-up multiple regression analysis including pitch, brightness, roughness, and speaker gender into the regression model indicates that speaker gender is the worse predictor of attitude toward the speaker. Its effect actually becomes non-significant in comparison with the other factors (β = 0.10, t(596) = 1.123, p = 0.262).
Interestingly, the same effects are observed for the perceived voice masculinity measure: there is no interaction effect between speaker gender and voice characteristics but a significant main effect of speaker gender (MMale speaker = 3.80 vs MFemale speaker = 3.55, F(1, 596) = 20.889; p < 0.001). However, when speaker gender is integrated into a regression model along with the voice characteristics, this effect becomes non-significant (β = −0.07, t(596) = −0.80, p = 0.419). These results seem rather counterintuitive at first glance but they are perceptible in everyday life: some female voices may sound masculine and male ones feminine, depending on their acoustic characteristics, not on their gender.
To summarize, in this context, speaker gender seems to play a redundant, secondary role in comparison with that of the voice characteristics.
Moderating effects of listener gender
A mixed ANOVA (within-subjects: pitch, brightness, and roughness; between-subjects: listener gender) indicates significant interaction effects between listener gender and all voice characteristics on attitude toward the speaker (Figure 2). Specifically, male listeners have a more favorable attitude than female listeners toward speakers with high-pitched voices (MMale listener = 3.79 vs MFemale listener = 3.59, F(1, 596) = 5.772; p = 0.02), bright voices (MMale listener = 3.66 vs MFemale listener = 3.39, F(1, 596) = 15.973; p < 0.001), and rough voices (MMale listener = 3.60 vs MFemale listener = 3.39, F(1, 596) = 3.877; p = 0.049). Follow-up tests indicate that the other simple effects are not significant. These results partially confirm H6.1 because the effects of low, dull, and smooth voices are not more favorable for female than for male listeners.
We also completed a mixed ANOVA (within-subjects: speaker gender; between-subjects: listener gender) to test the interaction effect between listener gender and speaker gender on attitude toward the speaker. The results indicate that the interaction is significant (F(1, 596) = 7.576; p = 0.01): male (female) listeners have a more favorable attitude toward female (male) voices than female (male) listeners. However, the results of multiple regression analyses show that this interaction (listener gender × speaker gender) becomes non-significant when the regression model also contains listener gender × pitch, listener gender × brightness, and listener gender × roughness interaction factors (p > 0.40). Thus, the interaction effects between listener gender and the voice characteristics explain the effects on attitude toward the speaker to a larger extent than the interaction between listener gender and speaker gender. To conclude, in a context with no specific speaker expectations, listener gender plays a role in voice persuasion. However, the sexual selection theory is not fully corroborated.
Brief discussion
Main results
In the context of an advertisement for a non-gendered product, this study shows that speakers with low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices have positive effects on perceived speaker warmth, attitude toward the speaker, and behavioral intentions. However, speakers with high and smooth voices are perceived as the most competent. Speaker gender plays a minor, redundant role, whereas listener gender moderates the effects of voice on attitude toward the speaker (feminine voice characteristics are more favorable for male than for female listeners). The effects of voice arousal seem to explain all the results. However, contrary to our hypotheses and the voice arousal hypothesis, the effects of voice arousal are, mostly, negative.
Contributions
These results fill several gaps. To our knowledge, brightness, roughness, and voice arousal were never studied in voice persuasion research. Pitch, brightness, and roughness equally contribute to the voice persuasion process. In particular, brightness contributes to gender identification and perceived voice masculinity to the same extent as pitch. Rough (vs smooth) voices are perceived as more feminine, which contradicts several studies (Podesva, 2013) and confirm others (Fessenden, 2011; Mendoza-Denton, 2007; Yuasa, 2010). In addition, roughness explains most of the effects of voice on competence. Finally, voice arousal plays a central persuasive role.
Conversely, the results show that speaker gender plays a minor, secondary role, although it is the most studied voice variable (Whipple and McManamon, 2002; Wiener and Chartrand, 2014). In fact, the perceptions of a voice’s gender and masculinity are primarily induced by a voice’s pitch and brightness (not its gender), which confirms prior research on voice gender categorization. For this reason and others exposed in the following section, the second experiment will manipulate the voice characteristics and not the voice gender (measured a posteriori).
Contrary to speaker gender, listener gender does play a moderating role in voice persuasion, at least in the absence of specific speaker expectations. Male listeners have a greater tolerance than female listeners for speakers with high-pitched, bright, and rough voices, which is consistent with the sexual selection theory (Puts et al., 2012). These results, however, do not completely confirm this theory (i.e. masculine voices are not more effective for female than for male listeners).
In addition, Study 1 simultaneously integrates the various variables identified in the multidisciplinary literature as key variables of the voice persuasion process. This approach has two implications. First, the cross-fertilization between various disciplines enables the understanding of the underlying mechanisms of voice persuasion. For instance, speakers with the least arousing voices are preferred and speakers with moderately arousing or moderately masculine voices are perceived as the most competent. Second, this approach explains some mixed results from prior research. For instance, prior research has found opposite effects of pitch on perceived speaker competence. The interaction between pitch and roughness might explain these results (these studies have not considered roughness).
Finally, Study 1 shows that voice persuasion is significant even in a context, which generates no specific speaker expectations (contrary to Whipple and McManamon, 2002) and where speakers talk at a moderate speech rate (contrary to Chattonadhyay et al., 2003).
Implications for the present research
The previous results show that a speaker’s voice characteristics and listener gender play an important role in persuasion. The results would mostly be explained by voice arousal, not by voice masculinity. However, our hypotheses suggest that context might change the direction of the effects, the roles of voice arousal and masculinity, and the significance of the moderation of listener gender. The results of Study 1 were observed in a particular context, deprived of specific expectations toward a speaker’s gender or competence level. The following study explores a context, which typically generates such expectations in listeners (Klofstad, 2016): a political campaign.
Study 2: Voice persuasion with specific speaker expectations
Before presenting the main study, a preliminary study explores a boundary condition of voice persuasion, that of the type of processing. Prior research has found that speaker familiarity and speech rate could affect the type of information processing. For instance, voice persuasion would not occur with familiar speakers or with slow speech rate, because the type of information processing would then be central. This could be an issue, especially in the context of a presidential campaign, where speakers tend to be seen, well-known, liked or disliked, and to talk with a slow speech rate.
Preliminary study
Design and participants
A sample of 83 respondents (Mage = 39 years, SD = 13.86; 62% female; 87% college-educated) were exposed to the voices of the 10 candidates for the 2012 French presidential election (N Arthaud, F Bayrou, J Cheminade, N Dupont-Aignan, F Hollande, E Joly, M Le Pen, JL Mélenchon, P Poutou, and N Sarkozy).
Stimuli
The same and unique sentence was extracted from the speech where candidates announced their candidacy (“My dear compatriots, I announce my candidacy for the presidential election”). The messages’ durations were similar (from 9.7 to 10.2 seconds). Candidates spoke at the same, slow speed rate (the number of SPM ranged from 150 to 170; for instance, Steeve Jobs’ and Al Gore’s speech rate in TED Talks ranged from 200 to 230, Dlugan, 2012). Loudness was kept constant by normalizing the intensity of the recordings in Audacity.
Objective measures of voice characteristics
Candidates’ voice pitch, brightness, and roughness were coded according to objective measures (F0, fc, fam, and m) as well as visual and aural estimations made by two musicologists using the spectrogram of each voice while pronouncing the vowel [a] from “compatriots” (acoustic software: Audacity and Melodyne). As an illustration, J Cheminade was the lowest voice of all, E Joly the highest, F Hollande the brightest, P Poutou the dullest, F Bayrou the smoothest, and JL Mélenchon the roughest.
Procedure and perceptual measures
In line with real listening conditions where the candidates are generally introduced, the listeners were told that these voices belonged to candidates for the 2012 French presidential election. The objective was to facilitate candidate recognition and central processing. In addition, the data were collected just before the 2012 elections and participants were not aware that the purpose of the study was solely related to voice perception. This context placed the listeners in a mind-set where their attention was more focused on the central message and on the candidate persona than on his or her voice characteristics. In order to verify that listeners processed the central information, a memorization test was successfully conducted: all respondents could correctly recall the message and recall the voices they had heard (prompted recognition: 98% on average). Spontaneous recognition varied but mostly exceeded 40%.
After the same instructions as in the previous studies, participants listened to a first audio extract, which they could play as many times as needed. Prior attitude toward the candidate, candidate recognition, perceived pitch, brightness, and roughness, and vote intentions were then measured (Supplemental Appendix 1).
Results
The results of three-way (voice pitch × brightness × roughness) repeated-measures analyses of covariance (ANCOVA) with prior attitude toward the candidate and candidate recognition as covariates indicate that low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices are perceived as significantly less high, bright, and rough, respectively, (pitch: MLow pitch = 2.10 vs MHigh pitch = 3.17, F(1, 82) = 14.915, p < 0.001; brightness: MDull = 4.05 vs MBright = 5.62, F(1, 82) = 56.419; p < 0.001; roughness: MSmooth = 3.22 vs MRough = 4.04, F(1, 82) = 33.519, p < 0.001).
These results indicate that, even with prior attitude toward the speaker and candidate recognition as covariates, the voice characteristics are still perceived with the same precision as in Study 1.
Interestingly, our results also show that a candidate’s voice pitch, brightness, and roughness significantly influence vote intentions (pitch: MLow pitch = 2.41 vs MHigh pitch = 2.12, F(1, 82) = 8.264, p = 0.004; brightness: MDull = 2.40 vs MBright = 2.14, F(1, 82) = 7.182, p = 0.008; roughness: MSmooth = 2.39 vs MRough = 2.14, F(1, 82) = 7.611, p = 0.006), even if prior attitude toward the candidate represents the most significant factor (multiple regression model including all the previous factors, p < 0.001). These results complete a study conducted on the voice pitch of real, American presidential candidates (Klofstad, 2016). However, the results of this study are to be considered with caution. This preliminary study only aims at exploring the capacity of voices to influence perception when central information is processed, especially prior associations with the speaker. Future research could replicate this design using more vocal stimuli and by coding other voice variables such as intonation.
To conclude, voices can still be persuasive when central information is processed, even those of famous political persona. Political leaders are, in fact, increasingly convinced of the power of voice, as witnessed by the current trend to resort to voice coaches among presidential candidates (e.g. JP Laffont for E Macron and M Beacco for F Hollande).
Main study
Design and participants
The main study manipulates the pitch, brightness, and roughness (2 × 2 ×2) of a fictitious candidate using the same stimuli creation process, study procedure, and measures (objective and perceptual) as Study 1. The respondents are exposed to the eight experimental conditions (within-subjects design), which actually corresponds to the natural exposure to candidates’ speeches during a political campaign (citizens are exposed to the speech of all candidates, who all legally have their share of voice). In addition, the data were collected just before the 2012 French presidential election, which enhanced the realism of the within-subjects’ exposure.
The sample of respondents (N = 505; 53% female) was recruited by the same international Internet market research company (Research Now) as Study 1 via a quota sampling method based on the socio-demographic and geographic distribution of the French population in 2012 (INSEE, 2012) with a confidence interval of 95%.
Stimuli
Study 1 resorted to a male voice and a female voice as a basis for the other vocal stimuli. Even if the male and the female voices were similar (medium-pitched with a moderate brightness and roughness), some differences still remained (other unstudied dimensions of timbre, slight pronunciation differences, etc.). To avoid this bias, the following experiment only uses one androgynous voice as a basis for the manipulations. An androgynous voice can be obtained via acoustic gender-morphing software (Pernet and Belin, 2012). To obtain such an androgynous voice, any voice can be recorded (male or female) and then modified via acoustic software.
In this experiment, the voice of a female voice actor recorded the following message: “my dear compatriots, I urge you to vote for me this Sunday.” This message is politically relevant (French citizens vote on Sundays) yet partisan-neutral, as recommended by Klofstad (2016).
Then, the software automatically modified a set of voice characteristics (including pitch, brightness, and roughness) to obtain an averaged, androgynous voice throughout the recording. A manipulation check verified that respondents (N = 24; 55% female) did not associate this voice with a particular gender (“this voice belongs to a man” = 27%, “this voice belongs to a woman” = 25%, “I don’t know if this voice belongs to a man or a woman” = 48%).
Based on this androgynous morphed voice, different combinations of pitch, brightness, and roughness (2 × 2 × 2) were then re-synthesized (cf. Overall methodology – Manipulations) to obtain eight versions of the message (Table 3).
Study 2 – description of the vocal stimuli.
Hz: Hertz (frequency measure); <CF/>CF: frequencies below/above the cut-off frequency of 1500 Hz; dB: decibel (loudness measure).
Respondents being exposed to all the vocal stimuli, we verified that they perceived them as different. Another sample of respondents (N = 25; 58% female) perceived the eight voices as belonging to different persons (“These voices belong to different persons,” 7-point Likert-type scale, M = 6.80) and as realistic (“This voice is credible for a political candidate,” 7-point Likert-type scale, M > 6.07).
Finally, one objective of Study 2 is to explore a speaker’s voice persuasiveness when a specific gender or competence level is expected. Prior research indicates that people clearly expect candidates to be male and competent (Klofstad, 2016). Nevertheless, we controlled on the previous sample of respondents that they typically expected male, competent candidates (“Imagine a candidate for an election. This candidate is …” male/female, Mmale = 70%; “This candidate has a high level of competence in the political field,” 7-point Likert-type scale, M = 6.1).
Procedure and measures
After the sound tests, the participants were told that they would listen to speeches of political candidates, without mentioning which type of political candidates. After each extract, the same variables as Study 1 were measured (Supplemental Appendix 1).
Manipulation checks
The multiple comparisons of three-way (pitch × brightness × roughness) repeated-measures ANOVA confirm the efficacy of the manipulation (Table 4): low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices are perceived as significantly less high, bright, and rough, respectively, (pitch: MLow pitch = 2.69 vs MHigh pitch = 3.57, F(1, 504) = 386.666, p < 0.001; brightness: MDull = 4.22 vs MBright = 5.07, F(1, 504) = 294.734; p < 0.001; roughness: MSmooth = 2.74 vs MRough = 3.29, F(1, 504) = 145.628, p < 0.001).
Study 2 – persuasive effects of a speaker’s voice characteristics in a political communication.
ANOVA: analysis of variance.
n = 505; the mean values in boldface are higher than others within the same condition.
p < 0.05; **p < 0.01; ***p < 0.001.
In addition, the results indicate that voice brightness moderates the effect of voice pitch on perceived voice pitch (F(1, 504) = 73.047, p < 0.001), as in Study 1. When a speaker’s voice is high-pitched, it is perceived as higher pitched with a bright timbre (MHigh pitch × bright = 4.26) than with a dull one (MHigh pitch × dull = 2.82). However, the results reported below indicate that this perceptual moderating effect has no implication whatsoever.
Finally, there is a significant interaction effect of the voice characteristics on perceived speaker gender (p < 0.001), male gender being primarily associated with low, dull, and smooth voices. This result confirms prior research indicating that a speaker’s voice characteristics serve as cue for voice gender categorization (Pernet and Belin, 2012).
Effects of speaker voice on voice perceptions
The results of the three-way repeated-measures ANOVAs (Table 4) indicate that low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices are perceived as more masculine (pitch: MLow pitch = 4.19 vs MHigh pitch = 3.25, F(1, 504) = 365.465, p < 0.001; brightness: MDull = 4.07 vs MBright = 3.37, F(1, 504) = 192.654; p = 0.04; roughness: MSmooth = 4.01 vs MRough = 3.43, F(1, 504) = 125.396, p < 0.001). There is no interaction effect of the voice characteristics on perceived voice masculinity. These results are consistent with H1 and Study 1.
Moreover, high- (vs low-) pitched, bright (vs dull), and rough (vs smooth) voices have positive effects on perceived voice arousal (pitch: MLow pitch = 3.18 vs MHigh pitch = 3.47, F(1, 504) = 39.419, p < 0.001; brightness: MDull = 3.13 vs MBright = 3.51, F(1, 504) = 51.498; p < 0.001; roughness: MSmooth = 3.06 vs MRough = 3.58, F(1, 504) = 118.291, p < 0.001). There is no interaction effect of the voice characteristics on perceived voice arousal. These results are consistent with H2 and Study 1.
Direct effects of speaker voice on speaker perception and persuasion
In a context generating specific speaker expectations, we have hypothesized that masculine voice characteristics (i.e. low, dull, or smooth) would be more effective than feminine ones, except for the effects on warmth.
Contrary to expectations (H3.2), there is a significant interaction effect between voice pitch and roughness on perceived speaker competence (F(1, 504) = 52.811, p < 0.001). Follow-up simple effect tests indicate that, when a speaker has a smooth voice, a high pitch leads to greater perceived speaker competence than a low pitch (MSmooth × High pitch = 4.00 vs MSmooth × Low pitch = 3.35, p < 0.001). The effect of voice brightness on perceived speaker competence is not significant.
Moreover, low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices have positive effects on perceived speaker warmth (pitch: MLow pitch = 2.90 vs MHigh pitch = 2.69, F(1, 504) = 18.082, p < 0.001; brightness: MDull = 3.14 vs MBright = 2.46, F(1, 504) = 232.069; p < 0.001; roughness: MSmooth = 3.02 vs MRough = 2.57, F(1, 504) = 76.456, p < 0.001), which contradicts H4.2.
As expected, low- (vs high-) pitched, dull (vs bright), and smooth (vs. rough) voices have positive effects on attitude toward the speaker (pitch: MLow pitch = 4.34 vs MHigh pitch = 3.83, F(1, 504) = 88.828, p < 0.001; brightness: MDull = 4.43 vs MBright = 3.74, F(1, 504) = 162,755; p < 0.001; roughness: MSmooth = 4.18 vs MRough = 3.99, F(1, 504) = 11,701, p = 0.001; H5.2) as well as on behavioral intentions (pitch: MLow pitch = 4.18 vs MHigh pitch = 3.64, F(1, 504) = 103.738, p < 0.001; brightness: MDull = 4.34 vs MBright = 3.48, F(1, 504) = 282.519; p < 0.001; roughness: MSmooth = 3.97 vs MRough = 3.84, F(1, 504) = 5,548, p = 0.02; H7.2). The interactive effects of the voice characteristics on attitude and intentions are not significant.
To summarize, low, dull, and smooth voices have positive effects but for one: high-pitched, smooth voices induce the highest mean for perceived speaker competence. In addition, these results are fully consistent with Study 1, despite the change in context. However, the underlying mechanisms behind the results of Study 2 are different, as presented below.
Indirect persuasive effects: The role of perceived voice masculinity
Prior research has suggested a positive effect of perceived voice masculinity on perceived speaker competence, a negative one on perceived speaker warmth, and a positive effect on attitude toward the speaker.
The results of the multiple regressions (Table 4) indicate positive effects of perceived voice masculinity on perceived speaker competence (β = 0.44, t(504) = 30.35, p < 0.001), perceived speaker warmth (β = 0.19, t(504) = 12.02, p < 0.001), and attitude toward the speaker (β = 0.42, t(504) = 29.07, p < 0.001).
These results show that perceived voice masculinity mainly has positive effects in a context where a male gender or a high level of competence is expected. They seem to explain the positive effects of low, dull, and smooth voices, which are perceived as masculine.
In addition, these results differ from those of Study 1, where the effects of voice masculinity were not positive but either curvilinear or moderated. Therefore, the results of Study 1 and Study 2 suggest that perceived voice masculinity plays a greater role when a male speaker gender or a high level of competence is expected.
Indirect persuasive effects: The role of perceived voice arousal
Prior research has suggested that voice arousal would have a positive effect on perceived speaker competence, a negative effect on perceived speaker warmth, and a positive (or curvilinear) effect on attitude toward the speaker.
Contrary to expectations, a curve estimation procedure indicates that voice arousal has an inverted U-shaped effect on perceived speaker competence (R2 = 0.01 vs 0.00 for the linear model, F(2, 503) = 23.075, p < 0.001), speaker warmth (R2 = 0.013 vs 0.012 for the linear model, F(2, 503) = 305.561, p < 0.001), and on attitude toward the speaker (R2 = 0.01 vs 0.00 for the linear model, F(2, 503) = 10.505, p < 0.001).
The results show that voice arousal has inverted-U shaped effects. These results do not explain the positive results of low, dull, and smooth voices, which induce a low voice arousal. However, the curvilinear effect of voice arousal on perceived speaker competence might explain the interactive effect of pitch and brightness on perceived speaker competence.
In addition, these results differ from Study 1. In the previous study, voice arousal mostly had negative effects, which explained the positive effects of low, dull, and smooth voices (they are less arousing).
To conclude, the results of Study 1 and Study 2 suggest that voice arousal explains voice persuasion in the absence of specific speaker expectations whereas perceived voice masculinity explains voice persuasion in the presence of specific speaker expectations.
Indirect persuasive effects: The role of speaker perceptions
There are positive effects of perceived speaker competence and warmth on attitude toward the speaker (competence: β = 0.27, t(504) = 17.47, p < 0.001; warmth: β = 0.34, t(504) = 23.01, p < 0.001) as well as on behavioral intentions (competence: β = 0.20, t(504) = 12.84, p < 0.001; warmth: β = 0.43, t(504) = 30.02, p < 0.001). In addition, there is a significant effect of attitude toward the speaker on behavioral intentions (β = 0.73, t(504) = 66.67, p < 0.001).
The previous results indicate that all the direct and indirect effects of a speaker’s voice characteristics are significant (except the link brightness – competence), which indicates the presence of multiple mediations in the sense of Baron and Kenny (1986; a, b, and c paths are significant for each mediation).
These results are consistent with prior literature and with Study 1.
Moderating effects
Contrary to study 1, the results of a mixed ANOVA show that there are no interactive effects of listener gender and a speaker’s voice characteristics on attitude toward the speaker. Male and female listeners have the same attitude toward a speaker’s voice characteristics (which is more positive for low, dull, and smooth voices). These results confirm H6.2.
Brief discussion
These results are consistent with those of Study 1 despite the change in context. Whatever the context, speakers with low, dull, and smooth voices are more persuasive. In addition, speakers with high and smooth voices are perceived as more competent.
However, some of the results of Study 1 and Study 2 differ. Voice persuasion is explained by different mechanisms in the absence versus presence of specific speaker expectations, by voice arousal and voice masculinity, respectively. The effects of voice arousal and voice masculinity are different from what prior research has suggested, based on the context. In the absence of specific speaker expectations, voice arousal mostly has negative effects and voice masculinity curvilinear effects. These results contradict the voice arousal hypothesis, which suggests positive effects of voice arousal based on acoustic inferences. The results also contradict the voice masculinity hypothesis, which suggests a positive effect of voice masculinity based on gender stereotypes. Conversely, in the presence of specific speaker expectations, voice arousal mostly has curvilinear effects and voice masculinity positive ones. These results contradict the voice arousal hypothesis and partially contradict the voice masculinity hypothesis, which suggests a negative effect of voice masculinity on perceived speaker warmth. This result could be explained by the fact that masculine voices are perceived as more benevolent (Brown et al., 1973), which is a valuable quality especially in the political field.
In addition, listener gender moderates the effects of speaker voice in Study 1 but not in Study 2. As we had expected, the presence of expectations toward a speaker’s gender and competence level undermined the influence of sexual preferences.
Finally, the preliminary study and Study 2 show that voice persuasion still occurs with central processing.
General discussion
Conceptual contributions
The results of the three experiments show that speakers with low- (vs high-) pitched, dull (vs bright), and smooth (vs rough) voices have positive effects on perceived speaker warmth, attitude toward the speaker, and behavioral intentions. Speakers with high and smooth voices are perceived as the most competent. In addition, listener gender moderates the effect of speaker voice on attitude toward the speaker.
This research fills several gaps of prior research as a result of the cross-fertilization of the various disciplines having studied voice perception, voice persuasion, and speaker persuasion. Prior research mostly considered one voice characteristic and/or speaker gender and never considered voice arousal. Our results indicate that pitch, brightness, and roughness equally contribute to voice persuasion, that voice arousal and voice masculinity both play a central persuasive role, whereas speaker gender plays a redundant role with a speaker’s voice characteristics. In particular, pitch and brightness are the two first cues enabling to categorize voice gender, confirming prior research in cognitive psychology. Roughness also contributes to voice gender categorization, rough voices being perceived as more feminine in today’s sociocultural context.
In addition, the consideration of multiple mediating variables and contexts explain the mechanism of voice persuasion and some of the mixed results from prior research. In both contexts, moderately arousing voices (e.g. high and smooth) lead to a stronger perception of competence. In the absence of specific speaker expectations, the persuasiveness of a speaker’s voice is explained by low voice arousal and by listeners’ sexual preferences (feminine voice characteristics are more favorable for male than for female listeners). Conversely, in the presence of speaker expectations, voice persuasion is explained by perceived voice masculinity which, notably, has a positive influence on perceived speaker warmth. These mechanisms explain why prior research has found different persuasive effects for pitch (Chebat et al., 2007; Klofstad, 2016; Oksenberg et al., 1986), creakiness (Wiener and Chartrand, 2014; Yuasa, 2010), voice masculinity (Zuckerman, Hodgins, and Miyake), listener gender, and speaker gender (Whipple and McManamon, 2002; Wiener and Chartrand, 2014).
Moreover, Study 1 and Study 2 show that speaker expectations and central processing do not constitute boundary conditions for voice persuasion (contrary to Chattonadhyay et al., 2003 and Whipple and McManamon, 2002), which still occurs with, mostly, the same effects.
Finally, the present research shows that voice’s persuasion power is persistent (e.g. across context and gender – listener’s and speaker’s) and multifold (e.g. cognitive, affective, and conative influence). Therefore, the effects of voice should be studied or at least controlled in research related to source persuasion (spokespeople, endorsers, candidates, etc.) or interpersonal persuasion (between coworkers, interviewers/interviewees, etc.).
Managerial implications
This research has several implications for the selection of voices in a communication context. Low, dull, and smooth voices constitute a safe choice in various contexts. In particular, speakers with such voices induce a low arousal, which would be appreciated in the context of advertising clutter. They are perceived as the warmest, which can be reassuring in a difficult political context or in an advertisement for a product associated with certain risks. They also induce the most favorable attitude and intentions.
However, high-pitched and smooth voices should be favored in order to demonstrate the competence of a speaker or of the entity represented by the speaker. In addition, rough voices are perceived as more feminine, in line with the current trend in young females to use a creaky voice (Fessenden, 2011; Mendoza-Denton, 2007; Yuasa, 2010). Such voices can be selected to represent a female target.
The previous voice characteristics can be used in communications via multiple ways. First, a specific voice actor can be selected via agencies’ voice catalogs. Pitch and brightness can be evaluated aurally or visually via free, easy to use acoustic software (e.g. Audacity). In addition, the voice of a current speaker can be modified via coaching, voice morphing, or acoustic effects.
Finally, the results demonstrate that speaker gender plays a minor persuasive role in comparison with the voice characteristics. Therefore, it is important to evaluate a speaker’s voice characteristics, and there is no particular reason to hire a male rather than a female voice actor. The previous recommendations go against current casting practices for two reasons. The same voice-overs can be heard in multiple advertisements for totally different brands (the French voice actors dubbing Bruce Willis and Julia Roberts for instance), which demonstrates that the actual voice characteristics of the voice actor are not considered. In addition, female voice actors are underrepresented in advertising (Whipple and McManamon, 2002) or cinema, which continuously resorts to the same male voices for action movie trailers (such as Don LaFontaine, the movie trailer voice-over of Star Wars, Terminator, or Independence Day, nicknamed “the voice of god” and “Thunder throat”).
Limitations and future research
There are several limitations in this investigation of a speaker’s voice. Given the exploratory nature of the research, a complete model of voice persuasion was not tested. In addition, the within-subjects design of the studies induces several biases. For instance, the perception of respondents was relative, not absolute. Even if contamination and listening biases have been controlled and results have been corrected for repeated measures (Bonferroni procedure), this type of exposure might have affected perception. Moreover, the scales used to measure the mediating variables contained one item. For instance, arousal could have been measured using the Pleasure Arousal Dominance scale of Mehrabian and Russell (1974). Furthermore, as in most prior studies, we did not control for the sexual orientation of listeners. Finally, we only considered two contexts (i.e. no expectations vs expectations: male gender and high competence level). Another context could be investigated, where a female gender is expected or where warmth is more valued than competence. In addition, future research could manipulate context as a moderating variable.
Areas for future research are broad. Prior research has shown that the verbal communication style could affect source persuasiveness (Capelli et al., 2012). Future research could explore the interaction between vocal and verbal cues. Furthermore, the interplay between the voice and the advertising music might be of interest, previous research on advertising music having also emphasized the role played by arousal (Galan, 2009). For instance, according to Coutinho and Dibben (2013), the emotional expressiveness of music and voice are close. This remains to be studied in a marketing context. Finally, previous research has shown that voice perception (Sauter et al., 2010) and source persuasion (Winterich et al., 2018) varied across cultures. Future research could compare the persuasive effects of voices across cultures to examine the relevance of local versus global speakers (for instance, for global brands or international institutions).
Supplemental Material
Supplemental_Material – Supplemental material for Persuasion of voices: The effects of a speaker’s voice characteristics and gender on consumers’ responses
Supplemental material, Supplemental_Material for Persuasion of voices: The effects of a speaker’s voice characteristics and gender on consumers’ responses by Alice Zoghaib in Recherche et Applications en Marketing (English Edition)
Footnotes
Appendix
Study 1 and Study 2 measures.
| Measures | Item | Modalities | Source |
|---|---|---|---|
| Pitch | “This voice is…” | 1 = low-pitched to 7 = high-pitched | 1 |
| Brightness | “This voice is…” | 1 = dull, soft to 7 = bright, sharp | |
| Roughness | “This voice is…” | 1 = smooth, regular to 7 = rough, irregular | |
| Speaker gender | “This speaker is a…” | 1 = man, 2 = woman | |
| Voice arousal | “This voice is …” | 1 = calm to 7 = excited | 2 |
| Voice masculinity | “This voice is masculine” | 1 = strongly disagree to 7 = strongly agree | 3 |
| Speaker warmth | “This speaker is warm” | ||
| Speaker competence | “This speaker is competent” | ||
| Attitude toward the speaker | “I like this speaker” | 4 | |
| Attitude toward the candidate | “I like this candidate” | ||
| Purchase intentions | “I want to buy this brand” | ||
| Vote intentions | “I want to vote for this candidate” | ||
| Candidate recognition | “Which presidential candidate is it?”, coded as accurate or inaccurate recollection. | ||
1 = Huang et al., 2001, Von Bismarck, 1974; 2 = Bänziger, Patel and Scherer, 2014; Laukka, Juslin and Bresin, 2005; 3 = Zuckerman, Hodgins and Miyake, 1993; and 4 = MacKenzie, Lutz and Belch (1986).
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
