Abstract
Having access to people’s narratives and their meaning cannot solely be detected through the verbal content of discourse. The short transcript that is analyzed highlights that important emotional dimensions can be captured when considering voice characteristics (pitch, loudness, speech rate, and resonance) together with nonverbal vocalizations (e.g. laughs) in addition to the more classical approach centered only on a verbal approach. The text analyzed is related to a woman who discovered later in her life that her father was involved in collaboration with the Nazis during WWII. The results show that most of the woman’s verbal language is descriptive. This contrasts with the vocal analysis suggesting emotional experiences evidenced by variations in pitch, loudness, and speech rate. In some instances, two vocal indicators of emotion were present simultaneously. They always involved speech rate, together with loudness, resonance, and pitch. Moments of dissociation between the verbal and the vocal dimensions of emotions are presented, suggesting that the woman still experiences strong emotions related to her past without being aware of them or without being willing to disclose them. The discussion highlights the relevance of analyzing jointly the verbal content together with the prosody and nonverbal vocalizations when examining autobiographical narratives.
- So he might potentially have been siding with the German army.
- Yes.
- But you don’t know if he left for the Eastern Front or if he stayed somewhere else?
- And so there would have been a moment where he was in the combats and received some shrapnel. 1
- Yes, exactly.
This is one example of an interview a researcher from my lab had with a Belgian woman, whose father was an active member of the collaboration with Germany during WWII.
However, the woman never got formal confirmation for that. This excerpt highlights that she only has indirect evidence of her father’s involvement in the collaboration. The following excerpt describes a discussion the woman had with her brother, who was likely to have more information. Because the woman does not have contact anymore with her brother, she cannot get additional information to support her assumptions. The original language of the transcript was French. It was translated into English for the special issue this article is part of.
- But do you know this story about dad’s pension from Germany? (. . .)
- Yes, because dad asked me to correct a letter that he wrote in German so that his pension could continue and so I helped him.
- But did you ask him any questions?
- No, I didn’t ask, I translated.
- But you, do you know something?
- No. I don’t know anything.
A very usual way to analyze an interview like this one would consist of a verbal content analysis with a narrative highlighting the strong assumptions about the woman’s father being actively involved in collaboration during WWII. The excerpts are too short, however, to provide full access to the emotional responses experienced by the woman in such a situation. The overall context of the interview, related to the difficult past of the father, is expected to trigger emotional responses. In addition, the team of researchers received a brief contextual background before reading the interview. It mentioned that the woman discovered more than 40 years ago that her father lost his civil rights but that she is still uncertain about the exact nature of his behaviors that can account for the court order he received. The discovery of a socially unacceptable past, together with uncertainty and ambiguity of the situation, are central dimensions that predict the experience of strong negative emotions as we will explain later. However, our central claim is that using only the verbal material described above could not properly account for the woman’s emotional feelings states.
Before going further in the description of my proposal to consider non-verbal dimensions of emotions such as the voice, I would like to disclose different personal statements that will help the reader understand my expertise and the way I can approach the interview of the woman. I am a Belgian psychology researcher who examines the role of emotional responses in the formation of individual and collective memories. My work includes studies conducted in naturalistic environments in which I analyze the factors that explain the formation of flashbulb memories. These are detailed memories of the reception context where one learns about a surprising and consequential event (Luminet and Curci, 2018). The goal of those studies is to understand when and how emotional factors affect personal and collective memories. I am also involved for more than 10 years in research examining the transmission of historical events across generations, in which, however, the role of emotions in the potential facilitation process of memory transmission was never considered. This background information indicates that I am familiar with this type of memory account from the remote past. But importantly, before the workshop in Trento during which I was first exposed to the interview, I did not know the story of the woman. It involves that my level of familiarity with the story is similar to that of one of the other contributors to this special issue.
The goal of this article is to emphasize the role of non-verbal emotional factors in shaping the content of personal memories. Hopefully, this contribution will broaden and improve the type of information the reading of this short excerpt could bring regarding the emotions experienced by the narrator. Coming back to the examples provided above, despite strong assumptions that the woman would experience intense negative emotions due to the context of likely reprehensible acts committed by her father, together with her uncertainty about its content, the transcript reveals a very descriptive language the woman uses to narrate her assumptions about her father’s collaborative acts. She does not use a single emotional word and she doesn’t make comments about some major family events that usually trigger important emotional reactions such as the death of her father or the break-up with her brother. Based on those verbal dimensions, one could conclude that the woman’s discourse is factual and does not contain emotional features. It is therefore remarkable that she can remember so well details from her past, knowing that the emotional nature of events is a central dimension to predicting better memory consolidation and retrieval (Emotion-Enhanced Memory effect or EEM; Ack Baraly et al., 2017). The presence in the woman’s narrative of very detailed memories of events that happened a long time ago would therefore represent a strong assumption that she experienced emotions. However, locating them will require using other sources than verbal language. The empirical part of the article will show that, in addition to a verbal analysis, it is crucial to consider non-verbal dimensions of emotional experiences, such as voice indicators that can represent a complementary way to grasp the emotional content of a narrative.
The next section will provide definitions of emotion and the facilitating effect of emotions on memory. Then, we highlight the importance of considering emotional responses through non-verbal indicators. The article then focuses on one aspect of those non-verbal dimensions, the vocal features, which are available together with the written transcript. The next section argues that a combined analysis of vocal and non-vocal indicators of emotion is crucial to detect potential dissociations between these two levels. Finding dissociations (e.g. factual discourse together with vocal emotional responses) is one fruitful approach to reveal dysfunctional ways of processing emotion that can impact long-term well-being (Constantinou et al., 2014; Peasley-Miklus et al., 2016). The discussion will argue that a complete analysis of emotions involved in a narrative should contain both vocal and non-vocal dimensions to fully understand the emotional and mnemonic impact of the narrative.
Defining emotion
Emotions are episodic, biologically based patterns of perception, experience, physiology, action, and communication that occur in response to specific physical and social challenges and opportunities (Keltner and Gross, 1999; Luminet and Cordonnier, 2024). An emotional episode is accompanied by the subjective experience of a mental state that is very distinct from our usual mental states. Certain symptoms inform us about this singular state, such as the awareness of bodily changes (sweating, palpitations, respiratory oppression, etc.), together with behavioral and expressive manifestations (face tensing up, hands shaking, etc.) (Frijda, 1986). The full emotional response includes three interconnected components: physiological (e.g. heart rate variability, skin conductance), behavior-expressive (e.g. facial responses, vocal expressions, postures, and gestures), and cognitive-experiential (e.g. appraisals, action readiness, feeling states).
Emotions are central to human beings’ lives as they fulfill various functions such as directing, facilitating, and maintaining people’s attention (Luminet and Cordonnier, 2024). Another essential function fulfilled by emotions is to facilitate communication. The emotional information conveyed by the face, the sound of the voice, or the body posture are fundamental for social interactions. While facial expressions are the main sources of communication of emotions, people can also express them through other channels such as posture, which refers to the position of their body (Luminet and Cordonnier, 2024).
One important additional function of emotion is to fix memories in people’s minds. Having memories of intense emotional events that happened in the past helps get better prepared when similar situations will occur in the future. If someone cannot keep memory traces of emotional episodes, they will be deprived of efficient ways to act and react (Luminet, 2022).
The intensity of the emotions felt when an event occurs and the ability of the person to retrieve it later are strongly associated. More specifically, the EEM effect shows that encoding, consolidation, and retrieval are enhanced when the information at encoding is emotional as opposed to neutral (Ack Baraly et al., 2017). The magnitude of EEM is affected by the valence people attribute to the event (from very pleasant to very unpleasant). Unpleasant events lead to more consistent memories and include more details and sensory dimensions than positive memories (Holland and Kensinger, 2010; Kensinger and Ford, 2020). Experiencing negative emotions also favors more specific memories than events experienced as pleasant. The memories held by the woman being mainly negative would even increase the assumption that she experienced some intense emotional episodes.
Non-verbal dimensions of emotion
While two out of the three components of emotional responses are nonverbal (physiological and behavioral-expressive), the majority of studies only consider the verbal one (cognitive-experiential). Within the behavioral-expressive component, most of the research is focused on the analysis of facial expressions, although there is a significant growth in methods for analyzing the emotional dimensions of the voice. Other channels are posture and gestures (e.g. Bänziger et al., 2014, for a psychological perspective; Ruusuvuori, 2012, for a conversation analysis perspective; see also Ochs, 1979, on the importance of non-verbal dimensions to fully understand people’s language).
The focus of this article is restricted to the vocal aspect of emotion. One approach is to ask judges to make subjective assessments of vocal excerpts. In addition to observational reports by judges, there is also a huge development of techniques that can measure acoustic parameters than can be measured through sophisticated software (e.g. Bänziger et al., 2014; Juslin and Laukka, 2003; Kamiloğlu et al., 2020; Scherer et al., 2015). Using those measures is beyond the goals of this article and is therefore not considered further. However, it is important to note that both perceived vocal features and acoustic parameters provide a high level of differentiation across emotionally sampled voices, highlighting that both approaches are valid (Bänziger et al., 2014).
Importantly, emotions in the voice can be expressed and measured in several ways, including verbal and non-verbal approaches (e.g. Bänziger et al., 2014; Ruusuvuori, 2012). Semantic information refers to the verbal content of speech, such as the meaning of sentences like ’I am sad’ or ’I am excited.’ This aspect will be later detailed in the language analysis section. Regarding the non-verbal dimension of the voice, two aspects are usually considered: speech prosody and vocalizations. Dimensions of prosody will be considered below. Vocalizations include, for instance, laughs and cheers usually related to positive emotions, and screams that are mostly linked with negative emotions. The majority of previous studies in psychology only focused on speech prosody and thus neglected vocalizations, which constitute, however, an important non-linguistic way of expressing emotions that will be considered in this article.
Recent studies established a list of prosody features that can be reliably rated by non-expert (i.e. naïve) raters 2 and that are involved in emotional expression and communication in everyday speech (e.g. Bänziger et al., 2014; Kamiloğlu et al., 2020). There are three dimensions for which there is a high agreement. It includes pitch (low vs high), loudness (weak vs strong), and speech rate (slow vs fast). We provide details about the content of these three dimensions here below. We also present a fourth one (vocal quality and resonance) that, although less often considered, was relevant for some moments of the woman’s interview.
The first dimension, pitch, refers to the relative highness or lowness of a tone as perceived by the ear. It varies as a function of the valence of the emotion experienced. For example, the voice of someone happy sounds more upbeat and has a higher pitch. Conversely, when someone is sad or upset, their voice may sound lower in pitch and have a more subdued tone. The second dimension is related to the loudness of the voice, sometimes also referred to as intensity or volume. When someone is excited or angry, their voice gets louder and more forceful. In contrast, when someone is feeling down or tired, their voice will become softer and less energetic. The third dimension indicates that changes in speech rate and rhythm can also reflect the presence of experienced emotions. For instance, when someone is anxious, they usually speak faster and have irregular speech patterns. In contrast, when someone is relaxed, their speech is slower and more steady. Finally, the fourth dimension, vocal quality and resonance, illustrates that when someone is feeling confident or enthusiastic, their voice sounds clear, vibrant, and resonant. Conversely, when someone is experiencing fear or sadness, their voice sounds strained, weak, or shaky.
When considering assessment by judges, emotional dimensions within the voice are limited to the dimensions of valence (e.g. degree of (un)pleasantness) and arousal (e.g. level of intensity). Asking to evaluate the presence of specific emotions would be unreliable (Banse and Scherer, 1996; Juslin & Laukka, 2003; Kamiloğlu et al., 2020). Our analysis will therefore be limited to these two dimensions. To access emotion-specific responses (e.g. sadness, anger, . . .), specific material is required to analyze the acoustic features. This goes beyond the scope of this article.
A combined analysis of verbal and non-verbal indicators of emotion
This section is divided into three different parts. In the first one, we examine the verbal characteristics of the discourse, highlighting the dominance of a descriptive discourse. In the next one, we show that examining the four categories of voice characteristics presented above provides a valuable complementary broader perspective on the emotional tone of the transcript. In the last part, the vocal analysis will be compared with the verbal one. This combined analysis will allow the detection of potential dissociations between the verbal and the vocal levels of analysis.
Verbal analysis
A method often considered in psychology is using software that analyzes the emotional tone of a text, such as LIWC (Tausczik and Pennebaker, 2010). The present text, however, is too short to provide a reliable evaluation of the proportion of words related to emotional states using such software. Moreover, even with a longer text, the goal is to compare different texts and their degree of emotionality, while in the present case, the goal is descriptive. Therefore, I use a manual analysis of the presence of emotional words in the interview. This lexical analysis highlights the mainly descriptive nature of the language used by the woman in different episodes that were, however, expected to have an emotional nature due to the context of her history. I will illustrate this with three examples.
The first one is when the woman related her father explaining the presence of remains of shell waste (“shrapnels”) in the skin below his eyes to her children (lines 18–40). She seems quite distant when disclosing that episode. Although she emphasizes that her children were laughing regarding the difficulty of pronouncing the word “shrapnels,” she does not express any feeling (being positive or negative) about that description.
Another episode is the description of the conflict with her brother that is narrated in a purely descriptive way, despite a context involving no communication and no more relations from anyone else in the family with her brother after the death of their father (lines 41–58). One striking moment is when she narrates the moment she asked her brother about him being informed that their father received pension benefits from Germany (see the full excerpt in the introduction paragraph of the article). The narration is purely descriptive, even though the existence of pension benefits sent by Germany was a clear indication of the father’s collaboration activities. In addition, she concludes the narrative with the descriptive sentence “So that’s it” (line 58).
A third noticeable episode of using descriptive language when narrating an emotional topic is the reference to the potential existence of her half-sister (lines 65–66), which is simply mentioned without any further elaboration. Revealing her existence has two important emotional implications. The first one is that their father was unfaithful to her mother. The second is the existence of a secret regarding the existence of their half-sister.
In contrast with these three examples, there are two moments when an emotional language was present. In the first one, the woman mentioned that her brother might be acting “in a kind of provocation“ (line 69), suggesting that he could feel angry about the situation. Another moment is when the woman repeats the same word several times (“I I I I I send a letter saying . . .,” line 48). Such type of word repetition suggests that she was not at ease or embarrassed with the situation of her father needing a letter to acknowledge his right to receive a German pension.
This section highlights that a purely verbal analysis of the woman’s discourse suggests a dominant, neutral attitude toward her father’s collaboration activities during WWII that is reflected in the descriptive language she uses.
Vocal analysis
I would now like to show that the analysis of the voice as a central non-verbal indicator offers a complementary perspective on the presence of emotional responses in the interview. The transcription, based on conversational analysis, gave important information, including length of pauses, changes in intonation, or overlapping speech. It thus offered a unique opportunity to include an analysis of nonverbal dimensions of the discourse. Another strength of the method is that as readers we were provided with a minimal amount of context, which allowed us to pay more attention to nonverbal aspects. Thus, the experimental device created favorable conditions to consider the voice characteristics.
This part considers two aspects. The first one illustrates some moments during which one single emotional characteristic of the voice can be detected. The second one presents sequences during which two emotional characteristics of the voice are present almost simultaneously. To fully comprehend these vocal characteristics, the reader is advised to hear the audio recording. 3 Although most readers would not understand its content in French, they can probably notice the changes in the voice in the specific moments described below.
Regarding the pitch dimension, at the start of the sequence, the interviewer adopts a supportive/empathetic tone. But some of the sentences of the woman are pronounced defensively, in particular lines 9 and 10, with the voice going up (rising tone). This gives the impression that she is pushed to explain obvious things to the interviewer and that she is annoyed because the interviewer doesn’t know them. Another moment during which the tone of the voice varies very rapidly is found in lines 21–22 when the woman describes the unusual sound made by the “shrapnels” under the skin of her father. Variations in loudness can also be detected. After lines 9–10 mentioned above, the woman calms down for a bit, but suddenly she uses a very assertive voice when she explains how well her children remember the story of the “shrapnels” (line 18). Another noticeable change in loudness happens in the other direction with a sudden decrease in her voice intensity when she makes references to her brother whom she does not see anymore (line 41). Finally, the dimension of speech rate can be observed just after the episode described above (lines 9–10), when the woman calms down very rapidly, which translates into a much slower speech rate (line 11).
There were also moments in the interview during which two voice indicators changed simultaneously. Interestingly, it was always a combination including speech rate. The combination of changes in speech rate and loudness is particularly noticeable from line 24 until line 36 during which the woman speaks both very fast and in an intense way about the “shrapnels” episode. It also occurs in line 71, when the woman suddenly slows down and says at a very low volume, almost whispering, “I don’t know.” Lines 42–58 illustrate changes in speech rate together with vocal quality/resonance. This is an episode during which the woman relates that she has no more contact with her brother. In this part, her sentences are disclosed quietly, with a relatively low speed and including a warm resonance of her voice. On lines 68–70, there is a combination of pitch together with speech rate. In that part, her voice is getting at the same time higher and the speed is faster.
One goal of the analysis was also to examine potential moments of dissociation between the content of the discourse and nonverbal vocalizations, which is another important indicator in addition to the four prosody indices (see above). In lines 18–19, it is noticeable that while the woman talks about the serious topic of her father being a collaborator, she laughs twice when disclosing this information. This suggests a moment when she is not at ease with the situation.
Discussion
Throughout different examples of verbal and non-verbal aspects of the narrative, the results showed situations in which a lack of correspondence was observed between emotional versus non-emotional characteristics. This lack of correspondence is often referred to as dissociation in psychology. Dissociation can take different forms. Specifically, there can be simultaneous activation of one dimension, while there is deactivation of another, or deactivation or activation of one component, with no activation of other components. Regarding the present situation, the dissociation occurs between a verbal discourse that is factual and mostly deprived of emotional content together with variations in different dimensions of the voice, with the indication that these markers were activated on several occasions.
Examining dissociations between emotional responses is one way to catch potential dysfunctions in people’s processing of emotion (e.g. Luminet et al., 2021). Most studies examine the dissociation between physiological and cognitive-experiential components of emotions (e.g. Peasley-Miklus et al., 2016). The present observation suggests extending the study of dissociation across emotion components. Examining dissociations between emotional responses can reveal some important outcomes. For instance, chronic dissociation between emotional components can result in the long-term in negative physical and mental health outcomes (e.g. Schäflein et al., 2018). Psychological scholars in affective science argue that a full understanding of the functionality of emotion responses must include a dynamic analysis of the convergence versus divergence between indicators of the emotion components (e.g. Mauss et al., 2005). The different examples of dissociations I presented strongly suggest that the woman is still experiencing strong emotions related to her past but that she is not necessarily aware of those emotions or that she does not want to disclose them.
The analyses presented in this article highlight that considering together verbal and nonverbal aspects of a narrative brings important additional information than an analysis that focuses on only one level. An unidimensional perspective, as is often the case in memory studies, could not allow us to distinguish whether the woman was unconcerned and uninvolved by the actions of her father or if the descriptive nature of her language reflected active emotion regulation mechanisms such as avoidance to keep the emotional impact of the family past aside.
Integrating a multi-modal approach to emotion—in this case considering the link between the cognitive-experiential (language) and the behavioral-expressive (vocal) components of emotion—provides a new angle of analysis in which we can suggest that the woman is feeling strong emotional reactions in relation with her family past. Of course, the relatively crude analysis of voice indicators that we provided based simply on one judgment, although from a specialist of emotions, cannot allow us to distinguish between the specific emotional feelings the woman experienced while talking. However, it helped solve the dilemma that can emerge when only a textual analysis of someone’s narrative is considered. As noted above, the interpretation of vocal markers is both subjective and context-dependent. It means that it cannot be the sole source of information but it is a rich complementary source that can sometimes lead to different conclusions than a single component analysis.
The present article suggests different avenues for the future. The first one involves a content analysis of language that relies on valid tools, complemented by an analysis of prosody indicators and nonverbal vocalizations. The second one requires a fine-grained analysis of nonverbal behaviors such as the voice, by combining a computer-based analysis of acoustic parameters, together with a judgment of voice characteristics, ideally by at least two native speaker judges. The third one involves asking a judge who does not understand the language of the narrative to interpret it only from a non-verbal perspective. Such a triangulation of information can allow for a full description of the emotional nature of narratives and assess their impact on individual and collective memories.
Finally, emotional markers of the voice can vary across cultures, but also across individuals. The interpretation of these markers is subjective and context-dependent. Despite the improvements in the field of voice analysis and emotion recognition, accurately identifying specific emotions solely based on vocal cues remains a challenging task (e.g. Kamiloğlu et al., 2020). Despite these current limitations, we believe that when considering different indices of emotion in voices, there is sufficient evidence to suggest that several parts of the excerpt are revealing the presence of emotion in the woman’s discourse.
Footnotes
This article is part of the special issue “Micro-Memories,” edited by Thomas Van de Putte and William Hirst. Readers are advised to read the introduction to the special issue first to understand the exercise individual contributors were asked to conduct.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Fund for Scientific Research (FRS-FNRS), Belgium.
