Abstract
Ample scientific research has confirmed significant linguistic differences between truthful and deceptive discourse in both laboratory and field experiments. That literature is reviewed, followed by presentation of an experiment that tested the effects of veracity on a wide array of linguistic indicators and tested which effects were moderated by motivation and modality. A 2 (veracity: truthful/deceptive) × 2 (incentives: high/low) × 3 (modality: FtF/audio/text) factorial experiment revealed that linguistic indicators of quantity, immediacy, vividness/dominance, specificity, complexity, diversity, and hedging/uncertainty were all affected by veracity, and veracity interacted with motivation in the latter four cases. Only personalism and affect failed to differ between truth and deception. Modality also affected language use but did not interact with veracity. Four linguistic indicators together successfully classified 76% of text-based deception and 76% to 78% of truthful responses from text, audio, and face-to-face interaction. The importance of context in predicting linguistic patterns is emphasized.
Keywords
A well-established finding in the literature on deception detection is that human judges are only slightly better than chance accuracy when telling lies from truths (Aamodt & Custer, 2006; C. F. Bond & DePaulo, 2008a; Hartwig & Bond, 2011). Yet when making judgments along a rating scale, humans reliably differentiate truthful from deceptive messages (C. F. Bond & DePaulo, 2008b; Burgoon, Buller, Afifi, White, & Hamel, 2008; Burgoon, Buller, Ebesu, & Rockwell, 1994). One reason people may be sensitive to shades of difference in message veracity, despite their reluctance to label the utterances as lies and truths, is that they detect differences in the language used in such utterances. And, just as people may be less able to state the rules of grammar than to recognize an ungrammatical utterance, they may rely on such linguistic differences in formulating their judgments, despite being less able to explicitly articulate the indicators contributing to their judgments than their implicit knowledge warrants (Hartwig & Bond, 2011).
In the practitioner worlds of forensics and law enforcement, several tools such as the Scientific Content Analysis technique (Driscoll, 1994) and Statement Validity Analysis (Undeutsch, 1989) have been developed for ferreting out truthful from deceptive statements. Among the linguistic features they utilize are pronoun usage, verb tense, amount of detail, and verbatim recollections. Scientific evidence from both laboratory and field experiments has confirmed that truthful and deceptive discourse differ linguistically (e.g., Burgoon & Qin, 2006; Fuller, Biros, Burgoon, & Nunamaker, 2013; Hancock, Curry, Goorha, & Woodworth, 2007; Newman, Pennebaker, Berry, & Richards, 2003; Porter & Yuille, 1996; Sporer, 1997; Zhou, Burgoon, Nunamaker, & Twitchell, 2004). Yet language features have not been prominent in meta-analyses and summaries of deceptive communication (DePaulo et al., 2003; Hartwig & Bond, 2011; Vrij, 2000, 2008), partly because some summaries predated the accelerating empirical attention to linguistic features, partly because the number of studies has been insufficient to warrant their inclusion, and partly because moderators have produced such variability in findings that their utility as reliable indicators has been questioned. For example, DePaulo et al. (2003) included only nine linguistic indicators in their meta-analysis and two—response length and self-references—were highly variable.
Notwithstanding, the most current and complete meta-analysis (Hauch, Blandón-Gitlin, Masip, & Sporer, 2015), which included 44 independent data sets, identified many relevant linguistic features. Reviewed below, some findings were consistent with predictions, but others showed conflicts even among measures of the same construct. For example, deceivers used fewer words but longer sentences (both quantity measures); fewer exclusion words but not fewer causation terms (both complexity measures); and more future tense verbs but not fewer past tense or present tense verbs (measures of nonimmediacy). Other recent experiments following essentially the same experimental protocol (Braun & Van Swol, 2016; Braun, Van Swol, & Vang, 2015; Van Swol & Braun, 2014; Van Swol, Braun, & Malhotra, 2012) and not included in the meta-analysis likewise produced mixed results in that differences in pronoun, specificity and affect usage appeared in one investigation but not the others, and several other linguistic indicators failed to yield significant results.
One approach to reconciling these disparate findings is to group linguistic indicators into theoretically meaningful classes to see if more clarity emerges. That is the approach taken here.
The second approach taken here was to consider whether clear patterns emerge if certain moderators are taken into account. Clearly, the task of predicting whether an utterance or text segment is deceptive would be much simpler if indicators remained stable across contextual features such as topic genres, speaker motivation, and transmission modality. In the case of nonverbal behavior, there has been a longstanding, if misplaced, belief that many indicators of deceit are universal, stable, and hence context independent (see, e.g., Frank & Ekman, 1997; but for a counterpoint, see Burgoon & Buller, 2008). In the case of linguistic indicators, research has shown they differ depending on such factors as the type of deception under examination (e.g., fabrication, concealment, or equivocation; Buller, Burgoon, Buslig, & Roiger, 1994, 1996), event type, intensity of involvement, and emotional valence, among others. Zhou, Burgoon, et al. (2004) and Burgoon and Qin (2006) also showed that linguistic indicators are highly dynamic across the course of an interview. The inconsistent results across indicators, tests, and time point to the likelihood that linguistic patterns are highly context dependent and that moderator variables account for the contradictory findings.
Under examination here were two moderators that are centrally important in criminal justice, forensic interviewing, and deception detection contexts: the degree to which a sender is motivated to put forward a favorable impression, and the modality in which communication takes place. These two moderators are also essential for inclusion in any research design testing for motivation impairment effects (MIEs), discussed below, which was a major impetus for this investigation. Before taking up these factors, a brief description of linguistic indicators and their associated classes is presented.
Linguistic Categories and Indicators
By linguistics are meant lexical and syntactic features of language that are independent of content. Although the actual semantic content of a segment of text or discourse is excluded, metacontent features such as sensory details can be counted using dictionaries. It is possible, for example, to measure how many terms refer to sight, sound, touch, and so forth without knowing the specific sights, sounds, and touches being described. Linguistic features can be extracted from written text or from transcribed oral discourse. Although much of the impetus for analyzing such features automatically has been their application to text-based communication such as e-mails, they are also applicable to spoken language.
The features examined here were ones that have been hypothesized as relevant to deception and impression management. They have been tested under varying paradigms and in multiple laboratory and field contexts. The current classes of linguistic features and their associated indicators, shown in Table 1, are derived from the Zhou, Burgoon, et al. (2004), Zhou, Twitchell, Qin, Burgoon, and Nunamaker (2004), Burgoon and Qin (2006), Fuller et al. (2013), and Burgoon, Mayew, et al. (2015) experiments, with some modifications and regroupings due to multicollinearity among features. Some, but not all, were included in the recent meta-analysis by Hauch et al. (2015) on linguistic features that could be measured automatically by computer tools.
Linguistic Classes, Definitions and Indicators.
Quantity
One the most studied linguistic features is the quantity or length of an utterance, measured in such forms as number of syllables, words, verbs, content words, sentences, or talk time. Longer utterances not only convey more information but also signal how forthcoming a speaker is. The quantity or completeness of an utterance is also one of four conversational maxims of cooperative discourse (Grice, 1989): Interlocutors are expected to provide as much information as the conversational circumstance dictates. The early received wisdom was that truthful communicators would be forthcoming and deceivers would instead opt for reticence so as to abbreviate the conversation, limit the amount of incriminating information they disclosed and reduce the risk of their prevarications being detected (Buller & Burgoon, 1994; ten Brinke & Porter, 2012). Meta-analyses addressed this general pattern under the umbrella term of reticence, although the findings were mixed depending on what indicator was measured (DePaulo et al., 2003; Hartwig & Bond, 2011).
One potential reason for mixed findings is that there are many occasions when deceivers may intentionally opt for longer messages. For example, interactive experiments have found that when deceivers are attempting to be persuasive or are afforded the opportunity to plan, rehearse, or edit their utterances, those utterances are actually longer than those of truth tellers, whether measured as words, verbs, sentences, or talk time (Burgoon, Mayew, et al., 2015; Burgoon, Wilson, Hass, & Schuetzler, 2015; Dunbar et al., 2014; Van Swol, Braun, et al., 2012; Zhou, Burgoon et al., 2004). Greater loquacity may be part of a strategy to be more persuasive and dominant. Thus, quantity of verbiage can be deployed in contradictory ways—by saying very little or saying a lot. Consequently, the circumstances under which deception occurs must be taken into account before predicting whether deception will be associated with shorter or longer messages.
Specificity
Closely related to quantity is specificity, which refers to the amount of detail present in an utterance and thus is a concrete operationalization of the quantity of information present. Systems for coding legal testimony such as Statement Validity Analysis (Undeutsch, 1989), Criteria-Based Content Analysis (Steller & Köhnken, 1989; Vrij, 2005, 2008), and Reality Monitoring (Johnson & Raye, 1981) incorporate sensory details, spatiotemporal references, and contextual embedding as important elements for determining the veracity of depositions and courtroom testimony. Truth is expected to be populated with more specific details on the assumption that people retelling real events will have mental access to far more facets of context than those relating imagined events. Terms related to place and time, to the five senses, and to numbers, for example, add concrete details to a narrative. From a cooperative discourse perspective (Grice, 1989), reticence would also be associated with a paucity of specific details. Both information manipulation theory (McCornack, 1992) and information management theory (Burgoon, Buller, Guerrero, Afifi, & Feldman, 1996) posit that one way to accomplish deception is to manipulate the quantity of details.
Complexity
Linguistic complexity can take one of two forms: lexical complexity or syntactic complexity (Burgoon, Buller, et al., 1996). The former concerns the use of simple or difficult words, as measured by number of letters or syllables. The latter concerns the use of simple sentences or sentences with multiple phrases and clauses, as measured by punctuation, number of conjunctions, or syntactic parsing of sentence structure. Together, the use of more difficult, polysyllabic vocabulary and more complicated syntax affects how easy or difficult an utterance is to comprehend. To the extent that deception requires more cognitive effort, it may deplete cognitive resources, resulting in a speaker or writer devolving to less complex and more understandable language (Qin, Burgoon, & Nunamaker, 2004). Hauch et al. (2015) did not find use of simpler lexicon nor more exclusion words by deceivers; they did not measure syntactic complexity.
Diversity
Diversity refers to how unique or redundant each lexical item is. It is often measured as the type-token ratio, which is the percentage of distinct words divided by total number of words, or its converse, redundancy, which measures how many terms or phrases are used more than once. Repeated words and phrases produce nondiverse language. As with complexity, repetitive language may be indicative of deception diminishing the mental resources needed to access more advanced vocabulary and to construct more complex sentence structure, leading a communicator to default to overlearned and frequently used language choices. Repetition also gives the appearance of fulfilling the conversational requirements to give complete answers without actually doing so. Hauch et al. (2015) reported fewer diverse content words and a somewhat lower type-token ratio for deceivers than truth tellers.
Hedging/Uncertainty
Hedging and uncertainty refer to words, phrases, and sentence constructions that introduce vagueness, evasiveness, or ambiguity in meaning. Deceivers may use such language to evade detection by signaling indecision or lack of commitment to what they are saying and by failing to give definitive, verifiable responses (Bachenko, Fitzpatrick, & Schonwetter, 2008; Duran, Hall, McCarthy, & McNamara, 2010). Conversely, certainty language conveys confidence, something that should typically be more true of truth tellers than deceivers (Fuller, Biros, & Wilson, 2009), unless deceivers are adopting a persuasive and assertive stance.
Hedging and uncertainty can be instantiated in many ways. Modal verbs (auxiliary verbs) can either convey certainty or uncertainty. Examples of modals that convey uncertainty are “might” and “could.” Indefinite pronouns refer to other unspecified pronouns (such as “it”). Demonstratives present a vague referent for the subject of a sentence (such as “this,” “that one,” and “these”). Conjunctions may add vagueness by creating compound constructions. Future tense verbs are also less definitive than past or present tense verbs, which are open to verification. In addition to parsers that capture these various parts of speech, Loughran and McDonald (2011) created a dictionary of uncertainty words (such as “approximate,” “depend,” “fluctuate,” and “possibly”), weak modal verbs, and adverbs that create tentativeness. All these terms were collected under the umbrella of hedging/uncertainty in the current investigation.
Of course, the opposite of uncertainty is certainty, and other measures are available to capture this polar opposite of the construct. Strong modals (e.g., “must”), terms that demonstrate confidence in a statement, and intense adverbs like “definitely” all convey certainty and would counteract the use of hedging or uncertainty terms. Although truth tellers should be more confident in their answers than deceivers, during persuasive deception, speakers might purposely pepper their responses with more certainty language to bolster their credibility.
When uncertainty and certainty are measured using dictionaries, they must be measured separately. In the Hauch et al. (2015) meta-analysis, deceivers used slightly fewer tentative words but not more modal verbs or certainty terms—again, mixed findings that might be accounted for by key contextual factors.
Immediacy and Personalism
Immediacy concerns the degree to which language creates a psychological sense of closeness or distance. Wiener and Mehrabian (1968) proposed several linguistic features that make a statement more verbally nonimmediate: passive voice, past or future tense, modifiers, and pronouns that separate speaker from listener or remove the speaker from an utterance so that it is “depersonalized.” Constructions such as “you and I” are more distant than “we,” for example. And use of third-person pronouns removes the speaker entirely from a narrative. Nonimmediate and impersonal pronouns are thought to be a way for deceivers to disassociate themselves from an account or to blur personal accountability (Newman et al., 2003). Hauch et al. (2015) reported in their meta-analysis more frequent use of first-person singular pronouns by truth tellers and more frequent use of second and third-person pronouns by deceivers, as predicted. However, the effect sizes were quite small and heterogeneity was quite high, indicating mixed findings across the 22 estimates they examined.
Affect
Most perspectives on deception feature affect as a major etiology of deception displays. Deceivers are thought to experience emotions of guilt, anxiety, and fear of detection when lying that prompt uncontrollable and uncontrolled nonverbal displays of the same (Buller & Burgoon, 1996; Ekman, 2009; Hancock et al., 2007). Whether these emotional states also generate negatively valenced language use is less clear, given that language choice is volitional and can be managed strategically. On the one hand, guilt is suspected to be reflected in the amount of negative emotion words (e.g., sad, bad, worthless; G. D. Bond & Lee, 2005; Vrij, Edward, Roberts, & Bull, 2000). On the other hand, when attempting to be persuasive, deceivers may also use more positive and hyperbolic language to win over the receiver. For example, in their analysis of fraudulent and nonfraudulent statements during a quarterly earnings report, Burgoon, Mayew, et al. (2015) found more positive affect and less negative affect language in the fraudulent statements, and affect was a significant predictor in classifying the deceptive and nondeceptive statements. The Hauch et al. (2015) meta-analysis found greater use of negative emotion words and a near-significant trend toward more positive emotion words by liars, leading to the tentative conclusion that compared with truth tellers, liars use more emotion words in general. However, the mixed results suggest that the degree and valence of affect-laden language may be moderated by several factors including the ones examined in this investigation.
Vividness/Dominance
A final category of linguistic features that may be diagnostic concerns the extent to which language conveys vividness, intensity, and power or the lack of these qualities that are often associated with dominance. Several nonverbal indicators of dominance have been investigated but verbal ones have seldom been examined in the context of deception and are not present in the Hauch et al. (2015) meta-analysis. Two prospective measures that convey this quality are expressivity (also referred to as emotiveness) and activation. Deceivers attempting to “fly under the radar” might be expected to use nondominant language, whereas those attempting to be forceful and persuasive might instead adopt more vivid and dominant language.
In summary, the linguistic results across wide-ranging investigations present a mixed picture that implies a lack of a straightforward truth-deception effect unmodified by contextual factors. Thus, an overarching research question was as follows:
Moderators of Linguistic Indicators of Deception
Motivation
In the deception literature, degree of sender motivation has long been considered a major influence on both sender displays and receiver detection accuracy (C. F. Bond & DePaulo, 2008b; Burgoon, 2005; Burgoon & Floyd, 2000; Hauch et al., 2015; Zuckerman & Driver, 1985). A significant dispute among scholars has been whether the majority of extant literature, which is based on mundane, low-stakes lies, is generalizable to circumstances with significant jeopardy or adverse consequences if caught (see Bachenko et al., 2008; Buller et al., 1994; Burgoon & Buller, 2008; Burgoon, Mayew, et al., 2015; DePaulo et al., 2003; Frank & Ekman, 1997; Hartwig & Bond, 2014; Vrij, 2008; Vrij & Mann, 2001). Certainly motivation should exert more influence on performance under high- than low-stakes circumstances, yet even low-stakes situations have shown motivation effects on deception displays and their detection.
Also in dispute is whether motivation facilitates or impairs performance. DePaulo and Kirkendol (1989; DePaulo, Kirkendol, Tang, & O’Brien, 1988) advanced the MIE hypothesis, which argues that motivated lies will be more transparent when judges have access only to nonverbal information but less transparent when they have access to verbal information. The implication is that motivation impairs nonverbal performance but improves verbal performance. The DePaulo and Kirkendol (1989) research did not examine behavior itself, but Burgoon and Floyd (2000) investigated several operationalizations of motivation for their effects on performance and concluded that motivation had positive effects on many aspects of both nonverbal and verbal behavior. (Verbal performance was measured with information management composites of veridicality/personalism and completeness/directness/clarity rather than with linguistic features.) Recent research by Hancock, Woodworth, and Goorha (2010) found that highly motivated liars interacting in an instant-messaging medium (text only) were the most successful in deceiving their partners, but the authors did not examine performance itself. A next logical step is thus to consider whether motivation moderates the use of linguistic indicators.
We contend, first, because higher motivation by definition represents a state of heightened readiness to act, it should be manifested through indicators of elevated arousal, activation, and expressiveness. Regardless of veracity, motivated communicators may become more voluble (measured as quantity), may use more language scaled as high in imagery and activation, and may exhibit more verbal expressivity. Following principles of social facilitation, greater motivation should stimulate communicators to perform well those overlearned aspects of language that people typically “run off” without much forethought, including utterances that are relatively more complex, diverse, and specific. That is, their utterances should run toward bigger words rather than simple vocabulary, complex and compound sentences rather than simple constructions, varied rather than redundant word choice, and more rather than sparser sensory details. Motivated communicators also are likely to be more conscious of their self-presentation and to attempt to put their best face forward by appearing pleasant and friendly. We hypothesize that these predictions regarding motivation should pertain to both truthful and deceptive responding:
Where motivation may interact with veracity to exert an ordinal impact on language is on those aspects of message production that are subject to impairment by excessive cognitive or emotional taxation. Comparatively, motivation may have a felicitous effect on those linguistic elements that communicators are best able and most likely to control and are, therefore, least easily degraded by cognitive or emotional overload. Speculatively, motivated deceivers should be able to produce lengthy messages that are complex, diverse, vivid, and personalized and yet may intentionally include more hedging language. Motivated deceivers may be less likely to monitor or control language that reveals negative affect or uncertainty. Thus, we predict that
Modality
It is well-known that people speak differently than they write, and synchronous communication such as chat differs from asynchronous communication such as e-mails, the former being far more informal, cryptic, and studded with abbreviations than the latter. Thus, modality in itself is bound to produce differences in the language that is used. Less obvious is whether modality should interact with sender veracity to produce differential patterns of features. To the extent that deceivers are strategic in their communication and take advantage of the potentials of a medium to present themselves in the most credible light, it may be that several features of deceptive language vary as a function of modality. For example, deceivers in face-to-face (FtF) interaction must create messages on the fly and may focus more on their nonverbal presentation but also utilize receiver feedback to guide their message production, taking note when receivers seem to be accepting what they are saying and elaborating only when it appears that the receiver is not convinced by what they are saying. By contrast, deceivers in written mode have more opportunity to edit their messages before transmitting and to focus exclusively on what they are saying. Audio communication spares deceivers concerns with their visual self-presentation but they must also produce messages in a timely fashion without benefit of nonverbal feedback from the receiver. Prior research (e.g., Van Swol, Braun, & Kolb, 2015) has shown modality influences what kinds of arguments deceivers make, depending on whether communication is FtF or computer-mediated communication. Thus, we speculate that modality may interact with veracity to affect linguistic features:
Method
Sample
Participants (N = 173) were recruited from a multisectioned, introductory-level communication class for a study investigating “the ability of people to deceive others and escape detection.” They received extra credit for participation and the chance to earn a monetary bonus if successful at convincing an interviewer of their innocence and credibility. Participants who confessed and those who failed to complete all measures were not included in the final sample. Among completions, 64% were female and 36 % were male; 77% were Caucasian, 13% were African American, and 10% were Hispanic/Latino, Pacific Islander, or another ethnicity; mean age was 19.6 years (range = 18-31).
Experimental Design and Procedures
The experimental design was a 2 (veracity: truthful/deceptive) × 2 (incentives: high/low) × 3 (modality: FtF/audio/text) factorial, with participants assigned randomly to conditions. Those assigned to the deception condition were instructed to commit a mock theft of a wallet from their classroom and, when interviewed about the theft immediately after class, to persuade an interviewer that they were not guilty. Thieves had the wallet on their person during the interview. Those assigned to the truthful condition were instructed to be in class on certain days, alerted that a theft might take place, and instructed to be completely truthful during the interview immediately after class.
One of the problems surrounding testing of motivation is that often it has only been inferred rather than manipulated or measured directly, producing empirical findings fraught with inconsistencies and contradictions (Burgoon, 2005; Burgoon & Floyd, 2000). Given that we define motivation as an internal drive an individual possesses, we chose to operationalize it through individual self-report. To instigate differential degrees of motivation, we followed methods common in past research by offering extrinsic rewards in the form of money, coupled with ego-based appeals meant to activate intrinsic motivation (DePaulo, Kashy, Kirkendol, Wyer, & Epstein, 1996; DePaulo & Kirkendol, 1989; DePaulo, Lanier, & Davis, 1983; Van Swol, Malhotra, & Braun, 2012; Zuckerman & Driver, 1985). Those in the high-incentive condition were told that, in addition to course credit, they could earn $10 if the interviewer evaluated them as truthful and $50 if they were the most successful at being evaluated as innocent and credible. Those in the low-incentive condition were told merely to convince the interviewer that they did not take the wallet and were reminded that regardless of how successful they were, they would receive course credit. After the interview, they became eligible for the same monetary rewards as those in the high-motivation condition.
Interviews took place under one of three modalities. In the text condition, interviewer and interviewee were placed in separate rooms equipped with wireless notebook computers and conducted the interview using Microsoft NetMeeting’s chat facility. In the audio condition, interviewer and interviewee were placed in the same separate rooms but communicated via handsets. In the FtF condition, interviewer and interviewee were seated in the same room and the interview was video recorded for later transcription.
On reporting to the research site, interviewees completed a social skills pretest and a written statement about what happened in class. They were then interviewed by one of three trained interviewers who followed a Behavioral Analysis Interview format that is taught to criminal investigators. After background questions related to high school and recent employment, the interview moved to the theft. Interviewees were asked if they committed the theft, to describe their activities on the day in question, to speculate about who might be responsible for the theft, what should be done to punish the perpetrator, and so forth. Two interviewers were trained by the third, who is a certified trainer in interviewing and interrogation. A pilot test provided extensive practice and feedback. All interviewers followed the same script and question sequence. Interviews averaged 10 minutes and were followed by a postinterview questionnaire and debriefing.
Dependent Measures
Video and audio recordings were first transcribed; text chats were retained in verbatim form, with interviewer turns removed. All transcripts were then segmented by turns at talk so that each question could be analyzed separately.
The features examined are ones that can be identified automatically from text. Most of the analysis utilized Agent99 Analyzer (Cao, Crews, Lin, Burgoon, & Nunamaker, 2003), a software tool built on the open-source General Architecture for Text Engineering (GATE; Cunningham, 2002; Cunningham et al., 2005). GATE is a Java-based, object-oriented architecture and development environment for analyzing, processing, or generating natural language. GATE comes with a shallow parser that decomposes text into sentences and words, a part-of-speech tagger, and capacity to compute a variety of counts and calculations. Agent99 Analyzer incorporated several additional linguistic-based algorithms, to which we added features found in the Grammatik tool in WordPerfect and the Whissell (1989) dictionary, which includes over 7,000 words with preassigned, empirically derived values along continua of pleasantness, imagery, and activation. To adjust for differences in utterance length, count variables were standardized by dividing raw counts by the total number of words in an utterance. Additionally, variables such as second-person pronouns that occurred rarely were dropped from analysis.
Results
The research question and all three hypotheses were tested in 2 (veracity: truth/deception) × 2 (motivation: high/low) × 3 (modality: FtF/audio/text) multivariate analyses of variance conducted on each linguistic category, with follow-up analyses of the specific features comprising each category. To improve the efficiency of the tests, where interaction terms had F values less than 1, the terms were removed from the analysis and reduced models are reported.
Research Question: Main Effects of Veracity
Veracity yielded main effects on (a) quantity (univariate only), (b) specificity, (c) hedging/uncertainty, (d) immediacy, and (e) vividness/dominance. Quantity produced a significant univariate effect on word quantity but with a small effect size, F(1, 166) = 4.07, p = .04, partial η2 = .024. Deceivers used fewer words than truth tellers. The specificity multivariate main effect, F(3, 163) = 3.39, p = .019, partial η2 = .059, produced no significant univariates and was overridden by an interaction with motivation (see below). The hedging/uncertainty main effect, F(2, 164) = 3.44, p = .05, partial η2 = .040, produced a univariate effect on modal verbs, but was qualified by an ordinal veracity by motivation interaction (see below). Deceivers used more modal verbs. The immediacy multivariate main effect, F(5, 161) = 2.52, p < .03, partial η2 = .073, was accompanied by a significant univariate on temporal immediacy, with deceivers being more (not less) immediate. Speculatively, this may be a means of avoiding being highly concrete by talking in present tense. The vividness/dominance multivariate main effect, F(3, 163) = 2.92, p = .04, partial η2 = .051, was accompanied by a significant univariate on extreme activation (1 standard deviation or greater). Deceivers used more extreme activation language than truth tellers. The remaining indicators did not produce main effects for veracity. Thus, several linguistic indicators were associated with the veracity of the interviewee’s statements. Deceivers’ utterances were briefer, more uncertain, more temporally immediate (present tense), and “active.”
A supplementary discriminant analysis was conducted to identify which combination of linguistic features best predicted the veracity of responses. Using prior probabilities calculated from actual group size and stepwise method for entering predictors in the model, a four-variable model emerged, F(4, 163) = 7.32, p < .0001, λ − 1 = .152, with these as the predictors: temporal immediacy (immediacy), average word length (complexity), modal verb ratio (hedging/uncertainty), and extreme activation (dominance). Cross-validated results produced 62% accuracy detecting deception and 64% accuracy detecting truth. When the analysis was conducted within modality, cross-validated accuracies improved. As shown in Table 2, accuracy within the text modality was 76% for deception and 78% for truth. Within the audio modality, accuracy for deception detection was worse at 58% but detection accuracy for truth was 75%. Within the FtF modality, deception detection accuracy was 67% and truth detection accuracy was 78%.
Classification Results, Cross-Validated.
Note. For text, 79% of the original cases and 77% of cross-validated cases were classified correctly. For audio, 68% of original cases and 66% of cross-validated cases were classified correctly. For face-to-face, 74% of original cases and 73% of cross-validated cases were classified correctly.
Hypotheses 1 and 2: Motivation Main Effects and Veracity by Motivation Interactions
Hypothesis 1 predicted that motivation exerts a main effect on linguistic features regardless of veracity, and Hypothesis 2 predicted an additional ordinal interaction with veracity on the linguistic features. Although many of the effects fell short of traditional levels of significance, a number of linguistic indicators surfaced as influenced by motivation. Consistent with Hypothesis 1, a multivariate main effect for motivation emerged on (a) quantity, F(2, 165) = 4.76, p = .01, partial η2 = .045 (with significant univariate effects on words, verbs, and sentences). Near-significant multivariate effects emerged on several other measures and merit mention because they were associated with significant univariate effects. This was the case for (b) specificity, F(3, 163) = 2.40, p = .070, partial η2 = .042 (with a univariate effect on modifiers); (c) diversity, F(3, 159) = 2.27, p = .08, partial η2 = .041 (with univariate effects on both lexical diversity and content word diversity); (d) personalism, F(4, 162) = 2.22, p = .07, partial η2 = .052 (with a significant univariate effect on second-person pronouns); and (e) affect, F(3, 164) = 2.60, p = .054, partial η2 = .055 (with a univariate effect on extreme positive pleasantness). Regardless of their truthfulness, highly motivated senders used more words, verbs, and sentences. They also tended to use more modifiers, “you” pronouns, and positive pleasantness, but less diverse language. Motivation thus altered language use in and of itself, although the effects were weak.
Hypothesis 2 predicted that motivation and deception would combine to produce ordinal interactions. Modest motivation by veracity interactions emerged on (a) complexity, F(1, 167) = 3.88, p = .05, partial η2 = .023 (with a univariate interaction on average sentence length); (b) hedging/uncertainty, F(2, 164) = 2.52, p = .03, partial η2 = .041 (with a univariate interaction on impersonal pronouns). Additional near-significant multivariate effects emerged on (c) specificity, F(2, 164) = 2.62, p = .076, partial η2 = .031 (with univariate interactions on specificity and sensory terms); and (d) diversity, F(3, 159) = 2.19, p = .09, partial η2 = .024 (with a univariate interaction on lexical diversity). The results for specificity, sensory terms, lexical diversity, and impersonal pronouns, which are displayed in Figures 1, 2, 4, and 5, were similar in showing that truth tellers did not change very much due to level of motivation, whereas motivation level did affect deceivers. High-motivation deceivers were the least specific, least detailed, least diverse, and most impersonal, whereas low-motivation deceivers were the most so. Only on average sentence length, shown in Figure 3, were high-motivation deceivers more complex than low-motivation ones, but the interaction may have been due more to differences among truth tellers than deceivers. Looked at differently within the levels of motivation, the results add another insight in that in most cases, high-motivation deceivers were similar to high-motivation truth tellers (a convergence predicted by interpersonal deception theory). The five illustrated interactions, although supporting the hypothesized interactions between motivation and deception, do not conform to the predicted ordinal interaction of Hypothesis 2. They also reveal that deceivers’ verbal communication may suffer in the form of a paucity of details, simpler language, and more indefinite or impersonal pronouns only when highly motivated, inasmuch as low-motivation deceivers included the most details, the most complex language, and the fewest impersonal pronouns.

Effects of veracity by motivation interaction on specificity.

Effects of veracity by motivation interaction on sensory terms.

Effects of veracity by motivation interaction on average sentence length.

Effects of veracity by motivation interaction on lexical diversity.

Effects of veracity by motivation interaction on impersonal (indefinite) pronouns.
Modality Effects
Hypothesis 3 proposed that modality would moderate veracity effects. Main effects for modality emerged on (a) quantity, F(6, 318) = 4.48, p < .001, partial η2 = .078 (univariate effects on words, verbs, and sentences); (b) specificity, F(6, 326) = 4.93, p < .001, partial η2 = .083 (univariate effects on modifiers); (c) complexity, F(6, 294) = 2.97, p = .008, partial η2 = .057; (d) diversity, F(6, 318) = 5.12, p < .001, partial η2 = .088 (lexical diversity, redundancy, content word diversity); (e) immediacy, F(10, 322) = 2.09, p = .02, partial η2 = .061 (spatial far ratio); (f) personalism, F(8, 324) = 3.98, p < .001, partial η2 = .089 (first-person plural pronouns, total pronouns); and (g) affect, F(6, 328) = 2.28, p = .036, partial η2 = .040 (no significant univariates). Only hedging/uncertainty and vividness/dominance were not affected by modality. No interactions with veracity emerged. Hypothesis 3 was not supported.
Discussion
A central question when applying linguistic analysis to assess a person’s veracity is whether indicators of truth or deception are context independent or are moderated by factors such as motivation and modality. In answer to the research question of whether deception exerts a main effect on language use, the current investigation discovered that seven classes of features were affected by veracity—quantity, specificity, complexity, diversity, hedging/uncertainty, immediacy, and vividness/dominance. Only personalism and affect were not influenced by veracity. Four classes of indicators (specificity, complexity, diversity, and hedging/uncertainty) were qualified by an interaction with motivation. Thus, linguistic features that could be most diagnostic of deception were often moderated by how motivated the interviewee was.
Supporting Hypothesis 1, five linguistic features produced main effects for motivation. Quantity, specificity, diversity, personalism, and affect all differed according to whether interviewees were given high-motivation incentives. Those who were in the high-motivation condition used more words, sentences, modifiers, second-person pronouns, and positive pleasantness but less diverse language than those in the low-motivation condition. However, the only interactions between motivation and veracity were not ordinal, as predicted by Hypothesis 2, but instead disordinal ones that were modest in magnitude. For indicators related to specificity, complexity, diversity, and uncertainty, motivation exerted more influence on deceivers than truth tellers. High-motivation deceivers were less specific on sensory details and other concrete details, less diverse in their lexicon, and more impersonal in their pronoun use than low-motivation deceivers. The exception was sentence complexity, where high-motivation deceivers were more complex than low-motivation ones. Comparatively, truth tellers differed little as a function of motivation. Looked at within the motivation conditions, results showed that highly motivated deceivers and truth tellers used similar language, a finding consistent with interpersonal deception theory that deceivers attempt to appear normal (i.e., like truth tellers) and converge on the communication patterns of truth tellers. Finally, results for Hypothesis 3 did not support a moderating effect of modality. Although virtually all the linguistic classes were affected by modality, not a single veracity by modality interaction emerged.
Implications for Specific Linguistic Classes
If the nine linguistic classes are considered one by one, further insight can be gleaned into the value of each of these classes of indicators in studying human communication.
Quantity was measured by typical features associated with message length, in this case, words, verbs, and sentences. Received wisdom had been that deceivers would be more reticent than truth tellers so as to limit the amount of discrediting information they might disclose and/or because greater cognitive taxation would limit their ability to give lengthy accounts. Although a meta-analysis by DePaulo et al. (2003) had found that deceivers had less talk time than truth tellers, they had failed to find a significant effect size for quantity when measured as response length. A later meta-analysis by Hartwig and Bond (2011) found that only perceptions of deception and not actual deception were associated with shorter responses, and the linguistic analysis of automatically extracted features by Hauch et al. (2015) found that truth tellers used significantly more words but deceivers used more sentences. However, their moderator analyses also revealed that amount of words had a much stronger effect when examined as a within-subjects variable. In other words, speakers measured under both truth and deception showed a significant change in their usage across the two conditions. They also found that motivation and modality influenced word frequency.
What does the current analysis add to this mixed bag of prior results? Words, verbs, and sentences all differed by level of motivation and modality (with more produced under high motivation and FtF communication). Truth tellers also used more words, verbs, and sentences than deceivers but the difference was only statistically significant for words. Thus, how much people say can be informative about a person’s truthfulness and level of motivation but the magnitude of effect may be weak. Coupled with the Hauch et al. (2015) findings of word count being moderated by motivation level and modality of communication, the effects of these moderators on response length under truth and deception require more unpacking in future research. A last noteworthy observation, and one relevant to most of the linguistic classes, is that form of measurement can determine whether a relationship is uncovered. Just as talk time but not response length was found to be a reliable indicator in the DePaulo et al. (2003) meta-analysis, words were the only significant indicator of veracity in this experiment. It seems advisable in the future to consider collecting multiple measures so that the full extent of the relationship between linguistic quantity and veracity can be assessed.
Inasmuch as quantity indicators were affected by modality, it is worth reminding researchers to adjust linguistic findings according to message length so that results are not a spurious consequence of communicators’ text messages being shorter or FtF messages being longer, for example, or other factors that influence message length.
Specificity was measured by adjectives and adverbs that are modifiers, sensory terms, such as are part of the Criteria-Based Content Analysis and Reality Monitoring coding protocols, and other linguistic forms such as determiners, conjunctions, and numbers that add concrete details to an utterance. In the DePaulo et al. (2003) meta-analysis, deceivers were found to produce fewer details than truth tellers. However, the Hartwig and Bond (2011) meta-analysis found that sensory terms and total details were perceived as related to deception but did not have strong correlations with actual deception. The Hauch et al. (2015) meta-analysis, which measured each type of detail separately, found significant results only for hearing-related terms and quantifiers. Here again, prior findings were inconsistent and not particularly strong when considering the whole complexion of possible features that could produce specificity and contextual embedding.
One explanation for past inconclusive findings surfaced in the present experiment. Many types of utterances do not require giving a detailed narrative and many past experiments entailed either brief utterances or a different genre of speech such as opinions or autobiographical information that would typically be devoid of contextual details. Under those conditions, specificity would not be diagnostic of veracity. By contrast, the current experiment required participants to give a detailed narrative about the conditions surrounding the theft, thus making the amount of specific details a salient aspect of the response.
A second explanation for past mixed findings may reside in the motivation level of the speaker. Here, veracity interacted with motivation on sensory terms and other specificity indicators. Whereas truth tellers’ responses were unaffected by motivation, deceivers’ responses were. Oddly enough, low-motivation deceivers gave the most details and high-motivation deceivers were similar to truth tellers in giving leaner accounts. Although an obvious explanation for this result has not presented itself, it may be connected to high-motivation deceivers giving more restrained responses. It is also worth noting that the similarity between highly motivated deceivers and highly motivated truth tellers implies that specificity may not be a reliable discriminator between truth and deception under circumstances where respondents are presumed to be highly motivated. Specificity was also affected by modality but results differed by indicator.
Complexity, composed of both lexical and syntactic indicators, was expected to differ by veracity. Deceivers were predicted to use more simplistic language and sentence structure than truth tellers. That prediction did not materialize. However, as hypothesized, motivation did interact with deception, on average sentence length. Low-motivation truthful utterances were longer than low-motivation deceptive ones. High-motivation deceivers were quite similar to high-motivation truth tellers, a finding supportive of the proposition that motivation encourages deceivers to mimic truth tellers but one also presaging the unreliability of complexity as a diagnostic discriminator between truth and deception if speakers are motivated. That said, complexity was one of the four linguistic indicators to emerge as significant in the discriminant analysis, an indication that it might play a useful diagnostic role when combined with other indicators. Finally, complexity varied by modality of communication but modality did not interact with veracity. Sentences were the most complex in FtF communication, whereas word choice was most complex in text.
Diversity, which consisted of the ratio of unique words or content words to total words and (the opposite of) redundancy, depended on the interaction of motivation and veracity. Motivation only altered linguistic diversity for deceivers, not truth tellers, and curiously, low-motivation deceivers rather than high-motivation deceivers used a more diverse vocabulary. Also, high-motivation deceivers approximated the linguistic diversity of truth tellers, which would make using diversity as a discriminator quite difficult under high-motivation conditions. All three diversity indicators (lexical diversity, redundancy, and content word diversity) were influenced by a modality main effect. The modality effects should alert researchers that the means of communication can exert significant influence over the variety or redundancy of language that will be seen and could conceivably set floor or ceiling effects, especially with shorter utterances.
Hedging/Uncertainty was measured by dictionaries of words conveying indecision, ambiguity, vagueness, evasiveness, and the like. A main effect for veracity revealed that deceivers used more modal verbs, which are often associated with uncertainty. An interaction between veracity and motivation revealed that motivated deceivers were more impersonal than low-motivation deceivers but also rather similar to high-motivation truth tellers, again reducing the discrimination power of such language under high-motivation conditions. The fact that such language is associated with deception warrants further unpacking the circumstances that might yield noticeable differences between truthful and deceptive responses along the uncertainty–certainty continuum. Given that Burgoon, Mayew et al. (2015) found that certainty and uncertainty were both useful constructs and not just opposite ends of a single dimension bolsters the value of further exploration of this linguistic class. Modality also did not affect this linguistic class, suggesting it may be more robust than others for use across forms of communication.
Immediacy, which in the verbal realm pertains to language that creates psychological closeness or distance, was instantiated through verb tense, voice (passive or active), spatially close or far references and current or not current temporality. Deceivers were expected to use more passive voice and spatially distant language, which convey nonimmediacy. It was thought that deceivers might also use more past or future tense verbs as a distancing mechanism, but there was the possibility of deceivers using present tense as a means of reducing temporal specificity.
Results revealed that veracity affected only one immediacy indicator: temporal immediacy. Contrary to predictions, deceivers were more immediate than truth tellers. Veracity did not interact with motivation. However, temporal immediacy was the first of the four significant predictors in the discriminant analysis, so its utility as a deception indicator warrants further study. Modality of communication also exerted a main effect on language immediacy, with FtF communication producing less temporal immediacy than the two computer-mediated communication modalities.
Personalism, composed of first-, second-, and third-person pronouns, was unaffected by veracity but was affected by motivation and modality main effects. Motivated speakers used more “you” references. Modality also affected personalism on first-person plural pronouns and total pronouns. Fewer pronouns were used in text messages. The failure of any of the pronoun indicators to differ by veracity or its interaction with motivation, coupled with the very weak showing for pronouns in the Hauch et al. (2015) meta-analysis, contributes more evidence of the questionable utility of pronouns as reliable indicators of deception. It may be that the polysemous nature of pronouns adds to the heterogeneity of their patterns of use.
Affect was intended to capture both negatively and positively valenced language related to emotions, mood states, and pleasantness. In light of how prominently affect has been featured in theories and models of deception, and how many validated measures of linguistic affect are available, affect was also expected to be revealed through language. Yet none of the affect indicators were influenced by veracity. Only motivation and modality produced significant relationships. High-motivated speakers used more extreme positive pleasantness than low-motivated speakers. The significant modality multivariate effect was not accompanied by significant univariates, and inspection of the means revealed inconsistent patterns across indicators. Thus, these results do not illuminate the role of affective language in identifying deception other than to add to the still small corpus of mixed findings. It may be that type of discourse investigated here did not lend itself to use of affect-laden language. If so, more theorizing regarding what contexts are likely to yield expressions of affect may need to precede investigations of affect in deceptive discourse.
Vividness/Dominance
In the relational communication literature, dominance is a fundamental dimension along which humans exchange messages and organize societies. Although numerous nonverbal dominance strategies have been identified, far fewer scholarly forays have ventured into linguistic expressions of dominance (an exception being the body of work on powerless language). However, linguistic analogues exist for many nonverbal indicators. For example, being assertive and active rather than passive has as its analogue activation language. Being nonverbally animated and impression leaving has as its analogue expressive and imagistic language. Thus, we expected that motivated deceivers wishing to be persuasive might use vivid and dominant language. Results showed that deceivers, regardless of motivation, used more extreme activation language. Vividness/dominance indicators were unaffected by motivation and modality. Consequently, this linguistic class has potential to span, unmoderated, some key contextual factors. Moreover, its emergence as one of the four significant predictors of veracity in the discriminant analysis bolsters its candidacy in the corpus of valid discriminators of truth and deception. Deeper conceptualization of what constitutes dominant or submissive language and development of appropriate measurement is warranted. At the same time, the circumstances under which deceivers are likely to opt for a submissive strategy versus a dominant one should be explicated if vividness/dominance truly is to be a useful discriminator.
Implications for Linguistic Predictors of Veracity
That language is complex and ofttimes can be used in a contradictory manner should come as no surprise. Take present tense verbs and first-person plural pronouns. On the one hand, expressions like “We always spend quality time together” might convey not only that “we” have a close relational bond (as compared with the pronouns “you and I”) but that the present tense verb “spend” (as compared with “spent” or “will spend”) makes the activity temporally immediate. On the other hand, a murder suspect’s answer of “We go to dinner early on Wednesday evenings” to the question of “Where were you last Wednesday night?” might qualify as nonimmediate and nonspecific by using “we” instead of “I” and “go” instead of “went.”
The high sensitivity of linguistic indicators represents a double-edged sword. With the correct specification of conditions, they could lead to precise assessment of veracity. But they also point to the necessity of knowing the context before making accurate predictions. The sparse knowledge of contextual moderators points to the need for researchers to spell out the conditions under which different linguistic indicators are likely to emerge.
The current investigation also reinforces the importance of testing linguistic features under varying conditions so that greater precision can be attained in predictive models of linguistic usage. Because humans are less able to manage their linguistic selections on the fly, such indicators have great promise for deception detection, but they must be subjected to far more rigorous testing before firm conclusions can be drawn.
Additionally, the constituents of each class merit further conceptualization, given that indicators within each class did not yield parallel results. Some indicators may be expendable or substitutable for one another. Where indicators within a class produced inconsistent results, the class itself may require revision, the end goal being determination of the most parsimonious yet comprehensive meaningful linguistic classes.
Implications for Motivation Impairment
One impetus for the current investigation was to test the dual MIE hypothesis that motivation impairs nonverbal performance but facilitates verbal performance. If the MIE hypothesis is correct, we should expect that highly motivated deceivers utter more credible responses than those with low motivation (which would then make their deception less detectable). In one respect, all the instances in which high-motivation deceivers approximated the language of truth tellers would be suggestive of the MIE being correct, but that would only be a valid conclusion if high-motivation deceivers behaved differently than low-motivation ones.
If longer messages are indicative of more credible messages, then motivated deceivers did not enjoy a facilitated performance. Regardless of motivation, deceivers’ messages were shorter. However, motivation itself elicited more words, verbs, and sentences. Other indicators that responded to high motivation, regardless of veracity, were modifiers, “you” pronouns, positive pleasantness, and diversity but not all in the direction of better performance. High-motivated senders used more modifiers and positive pleasantness, both of which would count as better performance, but also more “you” pronouns and less diverse language, both indicative of less credible performance. Moreover, in those cases where motivation interacted with veracity, high-motivation deception was more impaired than low-motivation deception. High-motivation deceivers were less specific on sensory details and other concrete details, less diverse in their lexicon, and more impersonal in their pronoun use than low-motivation deceivers. The exception was sentence complexity, where motivation had more effect on truth tellers than deceivers and high-motivation deceivers were similar to high-motivation truth tellers. Simply put, these combined results present a much more complicated picture than the prediction of the MIE. Motivation did not uniformly facilitate verbal performance.
Future Directions
Future research should also validate the alignment of specific indicators with their purported classes. If, for example, only one personalism feature is influenced by a given moderator, it raises the question of whether all pronoun forms serve the same function. The pronoun “I” could signal that a speaker is taking responsibility for a statement, or could signal nonimmediacy, as in “you and I” rather than “we.” Moving to n-gram analyses and collocations that consider collections of words in recurring expressions would advance understanding of what functions that utterances perform.
A further fruitful direction for linguistic analysis in deception would be to routinely conduct classification analyses to ascertain which combinations of linguistic indicators together achieve strong discrimination. Here four linguistic indicators successfully classified more deceptive utterances than human judges typically achieve (64% compared with the 47% estimate from meta-analyses). When analyzed by modality, deception detection accuracy actually improved to 76% within text and truth detection accuracy was 76% to 78% within the text, audio, and FtF modalities. Combining linguistic indicators with nonverbal and psychophysiological ones will surely also improve predictive ability.
Conclusion
Nine classes of linguistic indicators were examined for their ability to predict veracity. Linguistic features clearly can signal the veracity of a speaker but those signals must be placed within, and adjusted according to, the context in which they occur. The degree of motivation that the speaker is likely to have, and the modality through which communication is taking place, are just two of the contextual features that must be factored into the interpretation of any linguistic signals. Discovery of reliable predictors of truth and deception will prosper to the extent that language is part of the equation.
Footnotes
Acknowledgements
The author thanks Lauren Hamel, Pete Blair, and Tiantian Qin for their contributions to the conduct of the experiment and preliminary analysis of some of the linguistic data.
Author’s Note
An earlier version of this article was presented as the James J. Bradac Memorial Lecture, University of California Santa Barbara, September 2016, and portions were presented to the European Intelligence and Security Informatics Conference Workshop on Innovation in Border Control Conference, Odense, Denmark (August, 2012).
Declaration of Conflicting Interests
The author declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was partially supported by funding from the U.S. Air Force Office of Scientific Research (Grant #F49620-01-1-0394) and the Center for Identification Technology Research (University of Arizona Site), a National Science Foundation Industry/University Cooperative Research Center (Award IIP-1068026).
