Abstract
Pragmatic inferences require listeners to use alternatives to arrive at the speaker’s intended meaning. Previous research has shown that intonation interacts with alternatives but not how it does so. We present two mouse tracking experiments that test how pitch accents affect the processing of ad hoc scalar implicatures in English. The first shows that L+H* accents facilitate implicatures relative to H* accents. The second replicates this finding and demonstrates that the facilitation is caused by early derivation of the implicature in the L+H* condition. We attribute the effect to a link between L+H* and pragmatic considerations, such as speaker knowledge effects, or the saliency of alternatives relevant to the computation of implicatures. More generally our findings illustrate how intonation interacts at a cognitive level with pragmatic inference.
1 Introduction
The relationship between intonation and meaning is complex. To better understand the systematic contribution of intonation, a central aim in linguistics has been to distinguish the properly linguistic and non-linguistic functions of intonation (Gussenhoven, 2004; Ladd, 1996; Lehiste, 1970). This approach has led to considerable advancements in our understanding of how intonation contributes structurally to referential meaning and discourse comprehension. Our investigation, and that of this special issue, examines cases in which intonational meaning goes beyond reference and also blurs the distinction between linguistic and non-linguistic functions of intonation. That is, we examine how intonation works at the level of inferential meaning. We start by discussing several reasons for shifting away from referential meaning towards pragmatic inferences. We then integrate this into our broader goal, and that of experimental pragmatics in general, which is to develop both linguistic and psychological mechanisms that underlie pragmatic inferences.
Intonation can affect referential meaning and information flow in discourse in very predictable ways. For example, pitch accenting, the placement of pitch contours onto stressed syllables, correlates strongly with availability of referents in common ground, that is, new versus given information (Baumann & Grice, 2006; Clark & Haviland, 1977; Pierrehumbert & Hirschberg, 1990; Prince, 1980). However, intonation clearly affects other aspects of pragmatic communication and these have received much less attention. Consider the following example:
(1a) A: “I saw you stumble out of that new hip bar around the corner. The one with the over priced cocktails and 30 or so craft beers on tap. Looked like you had a little of everything, right?”
(1b) B: “I drank some of the BEERs”
(1c) B: “I drank SOME of the beers”
Here, Speaker A is teasing Speaker B about how drunk she got at a local bar. While the suggestion that she drank a bit of everything is meant to be sarcastic, B understands A’s comment as a request to list what and how much she drank. The example shows two replies that differ in their prosody. In (1b), B uses a pitch accent to increase the prominence of beers. The communicated meaning is that she only drank beers, and did not have any cocktails. In (1c), however, the accent is placed on the quantifier some. In this case the communicated meaning is different. B now implies something about the quantity of available beers that she drank, that is, not all of the 30 beers, but it is harder in this example to derive that she did not drink any cocktails. In other words, different implications arise depending on the intonation.
The inferences in these examples arise because the listener reasons about alternatives to what the speaker said, that is, material that would have been relevant, but that the speaker did not say (Grice, 1975). For example, in (1b), B’s emphasis on beers means that the listener generates alternatives to beer, such as cocktails, and in (1c), the listener generates alternatives to some, such as all. Since the speaker did not say the alternatives, the listener potentially derives an enriched meaning of the utterance, concluding that the alternatives are false. 1 In our study we investigate how intonation alters this process. We present two experiments that test the role of intonation in deriving a specific type of pragmatic inference, namely ad hoc scalar implicatures, such as the inference that the speaker did not drink cocktails, arising from Example (1b). We start out by presenting an overview on the role of intonation in establishing reference and evoking alternatives. Then, we discuss work on the derivation of scalar implicatures and how intonation affects this process.
1.1 Evoking alternatives through intonation
Intonation can perform a variety of linguistic functions across different contexts. We restrict our focus here to how intonation helps listeners evoke linguistic alternatives. The derivation of alternatives is essential for implicatures, in that the computation of implicatures requires listeners to reason about things that the speaker could have said, but did not. We review several pioneering and recent psycholinguistic studies on how different pitch accents can both evoke and integrate alternatives in identifying referents during comprehension. First, we note that one challenging aspect in every empirical investigation of intonational meanings is defining the exact nature of the intonational contour in question; namely these minimal pairs can be hard to uniquely identify in the continuous acoustic stream filled with both other meaningful prosodic parameters as well as noise. Like most psycholinguists, we are ultimately interested in meaning differences introduced by (perceived) categorically different intonational phenomena. Because of this, we limit our investigation to two pitch contours, namely the H* (L-L%) and the L+H* (L-L%) patterns, 2 that have not only received much attention in the experimental phonetic literature, but also whose functions have been investigated under the rigorous scrutiny of psycholinguistic research using various online measures such as eye-tracking (for a review, see Wagner & Watson, 2010; Watson, Gunlogson, & Tanenhaus, 2006).
The so-called H* accent is generally characterized as an increase in the relative prosodic prominence of a stressed syllable (Pierrehumbert, 1980). Dahan, Tanenhaus, and Chambers (2002) have shown that referents marked with such accents facilitate the processing of new discourse referents (as originally predicted by Pierrehumbert & Hirschberg, 1990). A second complex pitch accent, L+H*, has received much attention in experimental phonetics as well as in psycholinguistics. 3 While the H* accent is characterized by a simple high target on the stressed syllable, the L+H* accent starts with a low target followed by a steep rise in pitch contour (late peak). Phonetic studies have shown that the L+H* accent has a higher pitch excursion, longer duration and greater intensity compared to the H* accent (e.g., Bartels & Kingston, 1994; Krahmer & Swerts, 2001). An example of the two accent types is shown in Figure 2, depicting the average fundamental frequency (F0) values for our experimental items. Moreover, psycholinguistic work has shown that listeners distinguish these two accents during online processing. Yet the meanings of these accents seem to overlap in certain cases. For example, Watson, Tanenhaus, and Gunlogson (2008) investigated whether H* and L+H* accents help listeners predict referents among a set of alternatives in a visual world paradigm. In their experiment, listeners heard discourses such as “Click on the camel and the dog. Move the dog to the right of the square. Now, move the camel/candle (L+H* vs. H*) below the triangle.” In this example, the first sentence establishes an alternative set (camel and dog), whereas the last sentence tests whether the pitch accent pushes listeners to contrastive referents (look back towards the camel) or new referents (the candle). This is because camel was previously mentioned and established in a contrastive set, whereas the candle was not. Critically, listeners’ eye-movements anticipated contrastive targets and cohorts upon hearing L+H* on the first (overlapping) syllable. However, they were significantly less likely to look at unmentioned (new) targets and cohorts, such as the candle. When hearing the H* accent on the other hand, listeners were equally likely to look at both new and contrastive targets. This finding supports the idea that H* might be a more general prosodic marker for increased salience of the referent in common ground (a replication of Dahan et al.,’s 2002 findings), whereas the L+H* seems to create a bias towards an alternative referent that has been previously mentioned in the discourse. In addition, an experiment by Ito and Speer (2008) showed that the L+H* led listeners on a garden path such that they anticipated an alternative referent even when the display did not present an item contrasting in the relevant dimension. For example, when the noun phrase blue angel was followed by GREEN ball, listeners erroneously fixated on angels for 200 milliseconds (ms) and only afterwards fixations turned to the target.
As we have seen, the H* and L+H* accents help listeners establish reference, albeit in different ways, by allowing listeners to rapidly integrate alternatives in common ground during online processing. However, how do these accents evoke alternatives when they are not explicitly provided in the context? Braun and Tagliapietra (2010) found that the L+H* primed alternatives to the accented item when no set of elements was contextually-introduced (e.g., the word slipper activated the word flip flop). The H* accent, on the contrary, did not lead to priming effects. These results suggest that the L+H* accent activates alternatives. In addition, Husband and Ferreira (2016) found that L+H* can help restrict the set of alternatives. They showed that while listeners initially activate a broad set of elements, L+H* makes relevant alternatives salient after a delay. Thus, when alternatives are not already provided in the context, listeners need more time to evoke alternatives and converge on the relevant ones. One way to reconcile these findings with those presented above is to assume that the activation of appropriate alternatives underlies two mechanisms: one mechanism that activates a cohort of alternatives, and a second mechanism that selects those that are contextually-appropriate (see especially Gotzner, Wartenburger & Spalek, 2016; Husband & Ferreira, 2016). Intonation seems to play a crucial role in both. If, however, alternatives are already active in the context, the L+H* accent makes these contextual alternatives more salient, as found in a probe recognition experiment by Gotzner, Wartenburger, and Spalek (2013) (see also Fraundorf et al., 2010 demonstrating such effects in long-term memory). To summarize, previous studies indicate that the L+H* can introduce additional alternatives as well as make specific alternatives more salient, depending on whether alternatives are contextually available. We now address how intonation affects how listeners integrate alternatives when deriving pragmatic inferences.
1.2 Scalar implicatures and the role of intonation
The paradigmatic case of how listeners use alternatives to derive pragmatic inferences are scalar implicatures (Gazdar, 1979; Grice, 1969; Horn, 1972). Scalar implicatures arise when a speaker uses a weak expression when a stronger expression would have been relevant. There are a variety of different approaches to how scalar implicatures are derived but most researchers agree on the following: (i) the listener computes the basic meaning of the phrase; (ii) recognizes that an alternative phrase could have been used, but that it was not; and (iii) negates the alternative and combines it with the basic meaning (modern and developed theories can be found in Chierchia, Fox & Spector, 2012; Sauerland, 2004; van Rooij & Schulz, 2004; amongst others). In (1c), for example, the speaker said some, and the listener recognizes that they could have used the stronger expression, all. The listener then negates the alternative (not all) and combines it with the basic meaning of the sentence to form some but not all.
Scalar implicatures are optional components of meaning. The listener is not obliged to derive an implicature in the same way that they are obliged to derive the basic meaning of a sentence. Thus, the listener must make a decision about whether the speaker intends an implicature. This decision depends on a range of pragmatic and semantic factors. From a pragmatic perspective, the goals and knowledge of the speaker are relevant (Grice, 1975). For example, if the speaker is not knowledgeable enough to know whether the stronger expression is true, the listener should not draw conclusions on the basis that they did not use the stronger expression. Structural factors also seem to play a role in that certain linguistic environments block scalar implicatures (see e.g., Chierchia, Fox, & Spector, 2008).
An important question in scalar implicature research has been whether scalar implicatures are derived by default (Levinson, 2000) or whether they are contextually modulated. To this end, most psycholinguistic studies have used expressions in which the relevant alternatives could be stored in the lexicon, such as quantifiers (some) or disjunction (or) (e.g., Bott & Noveck, 2004; Bott, Bailey, & Grodner, 2012; Chevallier et al., 2008; Degen & Tanenhaus, 2014; Huang & Snedecker, 2009; Tomlinson, Bailey, & Bott, 2013; among many others). The consensus is that scalar implicatures are not derived by default. However, there is still much disagreement about the exact mechanism and constraints that underlie their derivation. One open question is the role of the intonation in this process.
Intuitively, intonation and scalar implicatures interact. For example, consider (1c), in which stress on some makes it clearer that the implicature is relevant. Two previous studies have investigated how intonation affects the derivation of scalar implicatures. Chevallier et al. (2008) investigated the effect of intonation on implicatures associated with or (e.g., “You can have the fish or the meat course” implies not both courses). They presented participants with letter strings, such as, “TABLE,” for very brief display times and asked questions about the letters in the word. For the crucial questions, the answer was “false” if the participant derived an implicature, but “true” if they did not. For example, for “TABLE,” the question was, “Is there a T or a B?” The participant answered “false” if they understood or to mean T or B but not both (the implicature; since both T and B were present), or “true” if they understood or to mean T or B and possibly both (the semantic meaning). They found that when or was stressed there were more false responses (implicature responses) than when there was no stress on or. As such, their findings show that intonation increased the rate of implicatures (see Schwarz, Clifton, & Frazier, 2008 for a similar finding). One drawback of this study is that no phonetic or phonological characterization of the stimuli was provided by the authors, hence it is difficult to see which intonational aspects were driving the effect and how this work relates to previous studies concerning the role of pitch accents in alternative activation (e.g., Braun & Tagliapietra, 2010).
A recent study by Gotzner and Spalek (2014), which does provide phonetic and phonological details, tested ad hoc scalar implicatures (Hirschberg, 1991; and see Bott & Chemla, 2016; Katsos & Bishop, 2011, for psycholinguistic examples). For these, the scale is constructed according to the context (on an ad hoc basis, similar to Example 1b). Gotzner and Spalek presented short discourses to participants in which the critical sentence contained an L+H* accent or an H* accent on a noun in subject position (e.g., The JUDGE followed the argument). They asked true/false interpretation questions that were false if the participant derived the implicature and found that the rate of implicature was higher in the L+H* condition than the H* condition (just as in Chevallier et al., 2008).
While both of these studies demonstrate a link between scalar implicatures and pitch accenting, they suffer two limitations. First, the effects of intonation on the derivation of pragmatic inferences might be conflated with lexical retrieval. For example, Chevallier et al. (2008) only investigated lexical scalar implicatures triggered by or. Since or and its alternative and are conceivably linked in the lexicon (e.g., Horn, 1972; Levinson, 2000), their results might be restricted to lexical scalar implicatures. Second, the methodology limits the range of conclusions that can be made about the underlying processing mechanisms. For example, both studies focused on final interpretation judgments and do not report online processing measures, such as response times, eye fixations or mouse paths. Although these are likely to be related, a higher rate of implicature does not necessarily equate to faster processing (see e.g., Bott et al., 2012).
Processing data can shed light on the exact mechanisms involved in the inferential process. In the case of intonation, there are at least three ways pitch accents could affect these mechanisms. First, intonation may only affect how likely a listener is to derive the implicature. For example, L+H* could be an additional signal to the listener that they derive the implicature, in which case L+H* would lead to a greater rate of implicature (as in the results of Chevallier et al., 2008; Gotzner & Spalek, 2014) but not necessarily a processing advantage. Second, L+H* might reduce the number or complexity of the mechanisms used to arrive at the decision. For example, L+H* intonation could speed up how quickly alternatives are negated. Finally, L+H* could direct the listener to derive the implicature at a different point during sentence processing, but otherwise leave the mechanisms involved in derivation unchanged. For example, L+H* might license the listener to derive implicatures early in sentence processing rather than waiting until the speaker has finished the utterance.
In our study, we tested between these possibilities. We compared interpretations of sentences with L+H* against those with H*, and varied whether the sentence gave rise to a scalar implicature. We used ad hoc implicatures so as to avoid lexically generated alternatives. Interpretation judgments were collected using a mouse, and the mouse trajectories provided the processing measure. Experiment 1 tested whether L+H* affects the rate of implicature or the rate and processing of implicatures, and Experiment 2 tested whether L+H* licenses an early implicature.
2 Experiment 1
Participants heard a sentence about a boy called Mark and some objects in his possession. They had to say which of two pictures best matched the meaning of the utterance. In inferential conditions, both pictures were consistent with the basic meaning of the sentence, and participants had to make an inference to form an unambiguous interpretation. One picture was the weak interpretation and one was the strong interpretation. The weak picture included the reference object and an object that was not mentioned in the sentence. For example, the weak picture for “Mark has a candle,” was a candle and a camel. The interpretation was weak because if participants did not enrich “Mark has a candle,” then it was consistent with Mark having a candle and a camel, and indeed any other object. The strong interpretation included only the reference object. For example, the strong picture of “Mark has a candle,” was, simply, a candle (and nothing else). This interpretation required the participant to reason along the lines of, “since the speaker did not say that Mark had two objects, and they are being truthful and informative, they must believe that Mark has only one object,” that is, the interpretation required a scalar implicature. For comparison, we added control conditions, in which one picture was consistent with the sentence and one was not, and so the response was unambiguous. For example, “Mark has a candle,” was accompanied by a picture of a candle on the left and a camel on the right, and so the correct interpretation (the target) was the candle picture.
We also manipulated the pitch accent on the noun (e.g., “candle”), which received either an H* or L+H* accent. Thus, if the L+H* accent interacts with the implicature, there should be a greater effect of the L+H* in the experimental trials compared to the control trials.
Participants made their responses by clicking on the appropriate picture with the mouse. The cursor appeared at the bottom-center of the screen at the beginning of each trial and the participant moved the mouse from the start to the response option (top left or top right corner of the screen—see Figure 1). Ease of processing was measured by the directness of the mouse path from initial position to response (see Spivey, 2007). Mouse-trajectories that are indirect, such as involving a wide arc from start to finish, indicate extensive processing, while those that are direct, for example, a straight line from start to finish, indicate minimal processing (Freeman & Ambady, 2010). We used mouse trajectories as a measure of processing rather than button-press reaction time for the following reasons. First, mouse dynamics provide information about cognitive processes intervening between stimulus and completed action (e.g., Tomlinson et al., 2013), whereas button-presses only provide information about the completed action. Second, and relatedly, mouse trajectories provide a measure of incremental sentence interpretation. Participants change their mouse paths as the interpretation of the sentence changes. Finally, mouse tracking allows multiple response options whereas button-press paradigms become difficult for the participant with more than two.

Trial set up for Experiment 1.
To summarize, there were two interpretation conditions (control and inferential) and two pitch accents (L+H* and H*) crossed in a 2 × 2 design. If processing of scalar implicatures is facilitated by the L +H* accent, we should observe more direct mouse paths in the inferential condition than the control condition than for implicatures with the H* accent.
2.1 Method
2.1.1 Participants
Twenty-eight Cardiff University students participated for course credit or 3 pounds Sterling. The experiment took roughly 15 minutes to complete.
2.1.2 Sentences and pictures
Inferential trials and control trials used the same sentence frame, “Mark has a [A]” where A was an object (see Table 1). The response options differed however. For inferential conditions (see Figure 1), the strong response option was a picture of one object, A, and the weak response option was a picture of two objects, A and B, where A was an object named in the sentence and B was not. The strong response option was consistent with the strong meaning of the sentence, A and nothing else, and the weak response option was consistent with the weak meaning, A regardless of whether something else was present. While both meanings were logically consistent with the sentence we anticipated that participants would generally choose the strong meaning, and so we designated the strong response option as the target. For control conditions (see Figure 1), both response options were pictures of only one object. However, only one response option was a picture of A, the object named in the sentence, and so this was designated the target.
Experiment 1 conditions.
Note: * all items had falling (L–L% boundary tones).
There were 42 experimental items. Roughly half were adapted from Dahan et al. (2002) and the other half (20 items) were created to increase the number of items. Of these, half of the sentence and picture combinations were phonological competitors, for example, candle versus camel and the other half were semantic competitors, for example, pencil versus eraser. This was done to help disguise the purpose of the experiment. Similarly, 41 filler items were created and varied among several dimensions to prevent listeners developing strategies during the experiment.
The pictures consisted of black and white clip art pictures found from an internet search engine. The objects were the same size in one object pictures as in two object pictures. This was done to control the salience of one-object and two-object pictures. Response options were placed in the top left and the top right corners of the screen.
2.1.3 Prosody
A male speaker of British English with no noticeable regional variety recorded the sentences. Sentences were recorded in a sound attenuated booth using a uni-directional microphone and digitized with a Universal Serial Bus sound capture device. All utterances were first recorded in carrier phrase (“Mark has a candle”) in both H* and L+H* forms. A trained phonetician inspected these recordings and made sure that utterances had either H*L-L% patterns or L+H*L-L% patterns. These patterns were further tested by an analysis measuring the stylized F0 difference between the patterns across all items at 5 normalized time points across the stressed syllable as used in Watson et al. (2008) and Spalek, Gotzner, and Wartenburger, (2014) (see Figure 2). The F0 values were measured at 5 equal intervals in the syllable receiving the pitch accent. Figure 2 shows that the H* tone of the L+H* pitch accent is realized as a plateau between the 3rd and the 4th F0 interval. The mean F0 was not significantly different between H* and L+H* items at the first and last intervals, 111.5 versus 118.9 Hz, t = 0.43, p = 0.66; 108.8 vs 122.1 Hz, t = 0.78, p = 0.44, but it was across the 2nd, 3rd, and 4th intervals, t’s > 2.21 and p’s < 0.03. L+H* and H* items differed also in where in the syllable the F0 peak was reached: F0 maxima were found on average 62 ms later in the syllable for L+H* items than H* items, t = 2.71, p < 0.01. The accented syllables for L+H* items had slightly longer durations than accented syllables in H* items (332 ms vs. 319 ms), however this difference was not significant, t = 0.64, p = 0.54. Further acoustic measurements are summarized in Table 2.
Fundamental frequency (F0) and duration parameters for experimental items.

Mean fundamental frequency (F0) values for stressed syllables.
Next, a single sentence frame (“Mark has a”) was spliced into each item, so that each item had the same sentence frame across conditions. 4 This was done to ensure that the acoustic stimuli only differed in pitch accent and not in acoustic information preceding the referent. All items were scaled for intensity at 68 dBs using Praat (Boersema & Weenick, 2015).
2.1.4 Counter-balancing
The assignment of item to condition was counter-balanced so that: (i) each item occurred equally often in each condition across participants; and (ii) participants never saw the same item twice. A Latin square design was used so that participants only saw one version of an item across the 4 cells (H* control, L+H* control, H* inferential, L+H* inferential). This meant there were four counterbalancing lists. Each list had 42 experimental items (10 or 11 from each cell) as well as 41 filler items. The position of the target response was also counterbalanced, so that it appeared equally at the left top and right top positions throughout the experiment.
2.1.5 Filler trials
In addition to experimental trials there were filler trials. Filler trials made up roughly 60% of all trials. All filler sentences had two referents, for example, “Mark has a [A] and a [B],” where A and B were objects. They were included to prevent participants from adopting strategies in which they only considered one of the objects in the sentence. There were three types of filler trials. One type involved a one object versus two object picture display, as in the inferential conditions, but unlike the inferential conditions, participants heard utterances with two referents. This type of filler made it equally possible that a participant would hear a one referent utterance or a two referent utterance when seeing a picture display that had one versus two object response options. Participants also saw picture combinations in which both responses were two object pictures. Half of these trials had cases in which the response options shared the first initial object, and therefore participants had to use the second referent to distinguish between the two responses. The different pitch accents (L+H* and H*) were intermixed amongst them, so that listeners would be equally likely to hear both pitch accents across all picture combinations. 5
2.1.6. Procedure
The experiment was conducted using the Runner program available in the Mousetracker suite (Freeman & Ambady, 2010). The instructions told participants that they were overhearing someone describe which objects a fictitious child (Mark) had on his desk. To start each trial, participants had to click on a START button at the bottom of the screen. Two pictures subsequently appeared at the upper corners and 2 seconds later the audio stimulus was presented. The participants could start moving their mouse at the onset of the audio stimulus and their movements were recorded at the onset of the word “has.”
2.2 Results
Participants’ responses were analyzed for both accuracy and the directness of the mouse-path towards the target response. Accuracy rates were at ceiling (over 97%) for both inference and control conditions (see Table 3). The data files were pre-processed using the Analyzer program, which normalized participants’ raw mouse trajectories into 101 time steps in order to compare the geometrical spatial attraction towards the targets across responses with different response times. Responses with reaction times 3 standard deviations outside of the grand mean of all responses were excluded. This amounted to roughly 1.3% of the entire data.
Accuracy and mean (M) area under the curve (AUC) values for Experiment 1.
The main dependent measure was area under the curve (AUC). This measure calculates the total geometrical area between a straight line from the starting point of a response (the START button) and the correct response and the participant’s actual response. Participants’ mouse trajectories (average y- over x-coordinates as well as the raw trajectories) are shown in Figures 3 and 4, for the control and the inferential conditions respectively.

Raw and average mouse trajectories for the control conditions in Experiment 1. The x-coordinates were transformed by multiplying by a factor of -1 for trials with reversed potions of the target and competitor pictures.

Raw and average mouse trajectories for the inferential conditions in Experiment 1. The x-coordinates were transformed by multiplying by a factor of -1 for trials with reversed potions of the target and competitor pictures.
Two linear mixed-effect regression models were fitted to the data; the first model made use of the 2 × 2 design, and used the two experimental factors as predictors, pitch accent (H* vs. L+H*) and sentence–picture combination (inferential vs. control). Similarly, a simple effects model, in which the four levels were collapsed into one fixed factor (condition), was also constructed to examine pair-wise comparisons and for consistency across the two experiments. 6 All of the models included random slopes per item and per participant (as suggested by Barr, Levy, Scheepers, & Tily, 2013). The formula for the simple effects model as well as the coefficients, t-values, and p-values for the AUC values are summarized in Table 4.
Fixed effect and variance estimates for area under the curve (AUC) values for simple effects models in Experiment 1.
Note: * not all of the models converged using condition in the random slope term. Because of this pitch accent was used this was the comparison of interest across the experiments.
Experiment 2 conditions.
Note: target and competitor images were randomly rotated for screen positions. Only target images and competitor 1 images remained next to each other, although these positions were randomized (e.g., either below/above or left of right).
First, the omnibus mixed-model used two fixed factors (pitch accent and picture condition) to predict the AUC values shown in Table 4. In this model, the factors were sum coded. Critically, there was an interaction of pitch accent and sentence–picture combination, β = 0.15, t = 2.05, p < 0.05. In the inferential condition, listeners’ mouse paths were significantly more direct when hearing L+H* accents than when hearing H* accents. In the control condition, however, mouse paths towards the target image did not differ significantly.
This interpretation was further supported by two simple effects models, which provided pair-wise comparisons in each level of the picture combination factor. In the first model (Model 1 in Table 4), the H* Control condition was used as the reference variable (intercept) in the model to test the effect of pitch accent in the control condition (treatment coding). There was no significant difference for AUC values between H* and L+H* accents in the control condition, β = -0.06, t = 0.68, p = 0.49, suggesting that participants’ mouse paths towards the target image did not differ in the control conditions. In the second model (Model 2 in Table 4), the reference variable (intercept) was changed to the H* inferential condition to test the effect of pitch accent on AUC values in the inferential picture condition. In the inferential conditions, there was a difference between pitch- accents, β = -0.28, t = 3.15, p < 0.002, showing that listeners had more direct mouse paths towards the target picture when hearing the L+H* than the H* accent.
2.3 Discussion
When participants heard sentences where a strong interpretation was relevant they overwhelmingly selected the strong interpretation as the most plausible meaning (97%). Participants thus saw nothing unusual about the task nor had difficulties understanding the intended meaning. More importantly, mouse paths were more direct when referents had L+H* pitch accents than when they had H* pitch- accents. When the strong interpretation was not relevant there was no advantage for the L+H* accent. Our results therefore show that L+H* facilitates the derivation of scalar implicatures.
While our findings are consistent with those of Chevallier et al. (2008), Schwarz et al. (2008), and Gotzner et al. (2013), they also add to them. Previous literature demonstrated that the implicatures were more likely under L+H* but our findings demonstrate that there is also a processing advantage. Furthermore, we demonstrate that facilitation is not restricted to lexical implicatures such as or or some. Instead, L+H* must affect mechanisms that are not related only to lexical retrieval of alternatives. In Experiment 2 we investigate the effects of L+H* on processing in more detail.
3 Experiment 2
There are two explanations for the findings of Experiment 1. The first is that L+H* increased the speed at which the implicature mechanisms were executed. For example, if the process that integrated the negation of the alternative with the basic sentence meaning was faster under L+H*, mouse paths would be more direct to the implicature interpretation under L+H*. The second is that under L+H*, the implicature was derived earlier than under H*. For example, under L+H* the implicature might have been derived when the referent was understood (e. g, “candle,” in “Mark has a candle”) whereas under H*, the implicature may have been delayed in the anticipation of more information as the L+H* might have reinforced nuclear stress and communicated finality. More direct mouse paths would occur under L+H* than H* because information about where to direct the mouse would arise earlier than under H*.
Experiment 2 was designed to test between these accounts. We used a similar task to Experiment 1 but made two changes. First, we added a prepositional phrase onto the sentences from Experiment 1. Objects were either on a table or a shelf, such as, “Mark has candle on the table,” or “Mark has candle on the shelf.” This meant that the reference object and accompanying prosody were no longer sentence final. Furthermore, since the location of the objects varied from trial to trial, participants needed to listen to the complete sentence before they were able to select the correct interpretation (note that increasing the variability in the sentence also meant that we had to increase the number of response options from two to four). Second, in addition to the inferential and control conditions from Experiment 1, we included non-inferential conditions. These were trials in which deriving the implicature mid-sentence would “garden-path” participants to the incorrect sentence interpretation. Sentences were the same as in the other conditions but the pictures indicated either a strong interpretation at an incorrect location, or a weak interpretation at a correct location. For example, “Mark has a candle on the table,” would be accompanied by pictures of a candle on the shelf, and a candle and a camel on the table (and two irrelevant response options).
Deriving the implicature on “candle” mid-sentence would direct participants to the picture of only a candle (on the shelf), whereas if participants waited until the end of the sentence, they would go directly to the correct interpretation, that is, the picture of a candle and a camel (on the table). L+H* or H* accents were placed on the reference object, just as in the other conditions. Thus, if L+H* makes the implicature more likely to be derived mid-sentence, mouse trajectories should first be directed at the strong response option, but subsequently double back towards the weak (and correct) interpretation. Of course, ad hoc implicatures might be derived incrementally even with an unmarked accent (e.g., Breheny, Ferguson, & Katsos, 2013), in which case garden-path effects will be observed for both accents. However, there should be a greater garden-path effect with L+H* accent than the H* accent.
3.1 Method
3.1.1 Participants
Sixty Cardiff University students participated for course credit or 3 pounds Sterling reimbursement. The experiment took roughly 25 minutes to complete.
3.1.2 Design
The pitch factor involved two levels, L+H* and H*, as in Experiment 1. However, there were now three levels of the sentence interpretation: inferential, control, and non-inferential. Thus, it was a 2 × 3 factorial design; Table 3 illustrates the design in more detail.
3.1.3 Sentences and pictures
The sentences for the inferential, control, and non-inferential conditions all followed the same form, “Mark has a [A] on the table,” where A was an object from the list used in Experiment 1. As in Experiment 1, the pictures in the response options distinguished between the conditions, however instead of two pictures at the top corners of the screen, there were four response options in total, for example, on each corner of the screen. There were also filler trials so that participants had to pay attention to all of the objects in the sentence and the prepositional phrase, as we describe below.
The addition of the non-inferential condition and the prepositional phrase meant that we needed to increase the response options from two to four. Thus, there were pictures in all four corners of the screen. In the inferential conditions (see Figure 5), the options were: (i) A and B on the shelf; (ii) A on the table; (iii) C on the shelf; and (iv) C and D on the shelf, where A was the object mentioned in the sentence and B, C and D were objects not mentioned in the sentence. In the non-inferential condition, the options were: (i) A on the shelf; (ii) A and B on the table; and with (iii) and (iv) as above. In each case option (ii) was the target. Finally, in the control conditions (see Figure 5), the options were: (i) A on the shelf; (ii) A on the table; and with (iii) and (iv) as above. Table 3 shows examples. The same 30 filler items from Experiment 1 served as the base for the fillers in this experiment and prepositional phrases were also spliced onto these utterances. Filler trials consisted of different types of picture combinations: roughly one-fourth using the same combination as experimental trials (i) A and B on the shelf, (ii) A on the table, (iii) C on the shelf (iv), C and D on the shelf, and the other three-fourths consisting of different patterns of picture arrangements.

Trial set up for Experiment 2.
3.1.4 Counterbalancing
The assignment of items to conditions was counterbalanced in a similar way to Experiment 1. This meant that there were six counterbalancing lists and each list showed 42 experimental items (7 in each condition) along with 41 filler items. The position of the correct response (top-left, top-right, bottom-right, and bottom-left) was rotated randomly across lists so that participants were equally likely to see the correct response at each of the four corners across the experiment.
3.1.5 Prosody
The same experimental items and fillers were used from Experiment 1. They were modified to fit the four-picture response paradigm by splicing a prepositional phrase indicating the object location (“on the shelf” and “on the table”) to the carrier phrase. The prepositional phrases were recorded by the same speaker and with the carrier phrase. All items were again scaled for intensity using Praat (Boersema & Weenick, 2015) after the prepositional phrases were spliced onto the items and fillers.
3.1.6 Procedure
The experimental procedure was the same as in Experiment 1. However, the number of counterbalanced lists was increased to six. Participants were presented with all four picture targets 2000 ms before the onset of the auditory stimulus. The only difference was that the START button was located in the middle of screen as opposed to the bottom of the screen as in Experiment 1.
3.2 Results
The accuracy rates for the inferential and control conditions used in Experiment 1 were at ceiling (over 97%) but the non-inferential condition had slightly lower accuracy rates, around 90% (see Table 6). The dependent measure and data preprocessing for the mouse-tracking data were the same as in Experiment 1.
Accuracy and mean (M) area under the curve (AUC) values for Experiment 2.
The plots of average response trajectories over x- and y- coordinates for the experimental conditions are shown in Figures 6, 7, and 8. An omnibus mixed-effects model with two factors (pitch- accent, 2 levels, × picture condition, 3 levels), revealed a significant interaction between pitch-accent and picture, β = -0.35, t = 3.62, p < 0.001, as would be predicted if L+H* facilitated processing of the inference sentences and impaired processing of the non-inference sentences. However, a variety of other patterns could also explain the effect. We therefore conducted three, one factor simple effects models to identify our effects in greater detail. The statistics for these models are shown in Table 7. Raw and average mouse trajectories for control conditions in Experiment 2. The x- and y-coordinates were transformed by multiplying by a factor of -1 for trials with reversed potions of the target and competitor pictures. Raw and average mouse trajectories for inferential conditions in Experiment 2. The x- and y-coordinates were transformed by multiplying by a factor of -1 for trials with reversed potions of the target and competitor pictures. Raw and average mouse trajectories for non-inferential conditions in Experiment 2. The x- and y-coordinates were transformed by multiplying by a factor of -1 for trials with reversed potions of the target and competitor pictures.


Fixed effect and variance estimates for area under the curve (AUC) values for simple effects models in Experiment 2.
The simple effects models replicated the findings from Experiment 1 for inferential and control conditions: listeners had more direct mouse trajectories with L+H* patterns than H* patterns in the inferential condition, whereas there was no difference between these two accents in the control conditions (Figures 6 and 7). As in Experiment 1, the L+H* accent in the control condition did not differ from the H* accent in the control condition (see Model 1 in Table 7), β = 0.0003, t = 0.014, p = 0.98. To test whether the L+H* pattern in the inferential condition differed significantly from the H* pattern in the inferential condition, the model was rerun (see Model 2 in Table 7) by setting the H* accent in the inferential condition as the reference variable (intercept). This was indeed the case and replicated the findings from Experiment 1: the AUC values for the L+H* in the inferential condition were significantly smaller than those for the H* condition in inferential condition, β = 0.062, t = 2.72, p < 0.01.
In the non-inferential condition, both pitch accents showed a garden-path effect as can be observed in Figure 8. To test, whether the L+H* accent induced a stronger garden path effect than the H*, a third simple effects model (see Model 3 in Table 7) was constructed by setting the H* accent in the non-inferential condition as the reference variable (intercept). The L+H* accent had significantly larger AUC values than the H* accent in the non-inferential condition, β = 0.075, t = 3.13, p < 0.002. This suggests that the L+H* accent serves as a stronger trigger to derive the implicature than the H* accent.
3.3 Discussion
Experiment 2 showed similar results to Experiment 1. In the inference condition, mouse paths to the strong interpretation were more direct with L+H* than with H* but in the control condition they were not. Experiment 2 therefore provides further evidence that the L+H* accent facilitates the derivation of scalar implicatures rather than merely making them more likely.
In Experiment 2 we added non-inferential conditions. If implicatures were derived incrementally, mouse trajectories should have initially been directed at the strong interpretation, but subsequently redirected towards the weak interpretation. In other words, they should have shown a garden path effect. This was indeed the pattern of results. More importantly, the garden path effect in the L+H* condition was greater than in the H* condition. Consequently, we argue that the L+H* pitch accent facilitates implicature processing by facilitating an earlier derivation of the implicature. We provide more detailed explanations of this effect in the General Discussion section.
Experiment 2 also allows us to exclude an alternative explanation for the facilitative effects of L+H*. In Experiment 1, the steeper fall in F0 under L+H* could have served as a signal to the listener that the speaker was finishing her turn (De Ruiter, Mitter, & Enfeld. 2006). This meant that the more direct mouse paths for L+H* accents could have been due to turn-taking cues instead of mechanisms involved with pragmatic inference. Embedding the referential expression in a prepositional phrase, as we did in Experiment 2, meant that L+H* was no longer at the end of the sentence. Thus, even when participants could no longer use L+H* to predict the end of a speaker’s turn, we still observed more direct mouse paths to the strong response option.
4 General discussion
Our goal was to investigate how pitch accents affect the processing of ad hoc scalar implicatures. In Experiment 1, we found that L+H* facilitates the processing of ad hoc scalar implicatures. In Experiment 2, we found that this effect arises because the implicature is processed earlier under L+H* than under H*. Below we explain how this might work and the implications of our findings.
L+H* encourages early implicature processing
In Experiment 2 we argued the implicature was derived earlier in sentence processing under L+H* than under H*. How might L+H* cause this to happen? We can think of a number of possibilities. The first is that L+H* causes the processor to bypass the normal, pragmatic mechanisms associated with implicatures, such as establishing whether the speaker has sufficient knowledge to intend the implicature (e.g., Grice, 1969; Sauerland, 2004). L+H* might act as a strong signal that the speaker intends the implicature interpretation, almost like an explicit focus operator, such as only, as in “Mark has only a candle.” Support for the similarity between L+H* and explicit only can be found in the offline data by Gotzner and Spalek (2014), who show that implicature rates are equally high with the L+H* and explicit only, and by the observation that it is very difficult to cancel an implicature with an L+H* accent relative to an unmarked accent. For example, (4) is difficult to understand, and far worse than (5).
(4) I drank SOME of the beers. In fact, I drank them all. (!)
(5) I drank some of the beers. In fact, I drank them all.
Crucially, if L+H* bypasses many of the normal pragmatic mechanisms, then the implicature should arise immediately after the pitch accent is recognized, just as we observed, because there would be no need to wait for further evidence of the speaker’s knowledge or other pragmatic information. Further support for this explanation comes from Filik, Patterson, and Liversedge (2009) and Kim, Gunlogson, Tanenhaus, and Runner (2015), who demonstrate that (explicit) only is processed incrementally during sentence comprehension in that listeners generate and disregard alternatives prior to having processed the entire sentence.
While this account of L+H* is consistent with the intuitive difficulty of defeasing L+H* and the offline results of Gotzner and Spalek (2014), there are some difficulties with it. There are no other studies linking L+H* with speaker knowledge effects in implicatures or similar, high level pragmatic factors and, moreover, we have no direct evidence that knowledge effects play a role in our study. Furthermore, L+H* does not behave exactly like only, in that understanding cancelled only statements (e.g., “I drank only some of the beers. In fact, I drank them all”) is nearly impossible, whereas participants in our study did just that in the non-inference conditions. 7
Another possible account of how L+H* causes an early implicature is that L+H* may elevate the salience of the relevant alternatives compared to H*. More salient alternatives might make the implicature more likely overall and difficult to cancel. This account would be consistent with the conclusions of Fraundorf et al. (2010), Gotzner et al. (2013), and Husband and Ferreira (2016), all of whom claim that L+H* makes relevant alternatives salient. If L+H* were directly linked to a mechanism that retrieved alternatives then this account too would predict that the implicature would occur earlier in the sentence than under H*. While this account is parsimonious, in that it draws together previous findings with ours, there are a number of factors that remain to be explained, just as with the knowledge account above. First, Gotzner and Spalek (2014) found very similar inference rates with only as with L+H*. Even though L+H* could raise the saliency of the alternatives by any degree, the level of saliency might still differ for L+H* and only (as was the case in Gotzner et al., 2013). Another is that the alternatives in our task were presumably very salient to participants, even without L+H*. The alternatives were easily derived from the response options and we included filler trials that explicitly introduced the alternatives. Furthermore, rates of implicature were extremely high throughout the experiment, which would not have been the case if the alternatives were difficult to retrieve. Under these circumstances it is difficult to see how L+H* could make the alternatives any more salient than they already were. Finally, there is no experimental evidence that salient alternatives make the implicature difficult to defease. It is an interesting possibility, but one that so far remains to be tested.
We hope that future experiments can clarify how L+H* causes the early derivation of the implicature. Interesting possibilities include testing whether speaker knowledge manipulations affect L+H* marked implicatures differently to H* implicatures, as they do with only and unmodified implicatures (Bergen & Grodner, 2012), and collecting offline interpretation judgments about the defeasibility of L+H* marked implicatures. However, our data do not allow us to categorically distinguish between a knowledge-based account and an alternative-based account of L+H* effects.
4.1 Gradient versus categorical effects of the pitch accents L+H* and H*
While the design of Experiment 2 rules out the possibility that listeners interpret L+H* accents to signal turn finality, it could be the case that the phonetic differences between L+H* and H* (higher overall F0 and larger pitch range for L+H*) could be driving our effects. Because our control conditions rule out a general low-level effect for salience at the level of word recognition, this effect must be more than lexical and referential, hence supporting our general claim that pitch-accents have unique functions during inferential processing. Nonetheless, it is still possible that these effects could have been driven by phonetic and not phonological differences. This criticism, however, would be valid for not only our study, but also for other phonetic and psycholinguistic studies researching contrast effects for the L+H* (Braun & Tagliapietra, 2010; Ladd & Morton, 1997; Watson, Tanenhaus, & Gunlogson, 2008;). Regardless of whether this effect is more phonetic or phonological in nature, both ideas are consistent with our account.
While we do not take a stand on the phonological status of these two accents, we see reason to believe that differences in meaning between L+H* and H* are gradient and not categorical. The presence of the garden path effects in Experiment 2 support Watson et al.’s (2008) conclusion that the meaning differences between L+H* and H* can overlap. We build on their study in that our findings show that these gradient differences in meaning are also found in more complex meanings such as the derivation of pragmatic inferences, not just in establishing reference.
4.2 Intonation and theories of scalar implicatures
Our studies show that pitch accents are integrated incrementally into the derivation of ad hoc scalar implicatures. We now explain how our findings could be accounted for in different theories of implicature. Although these accounts are not processing theories per se, how listeners integrate intonation in pragmatic inferences speaks to both the availability of linguistic alternatives and, especially in our studies, the mechanisms that operate over these alternatives. In other words, we see processing data as acting as constraints on these theories as opposed to explicit tests of the core assumptions of these theories.
In grammatical accounts (Chierchia, 2004, 2008, 2013; Fox, 2006) implicatures arise by insertion of a grammatical operator, a silent counterpart of only. Further, it is assumed that focus activates alternatives and feeds into the implicature generation mechanism (see especially Chierchia, 2013). At first glance, our findings are quite consistent with this account: L+H* could encourage the insertion of the grammatical operator and thereby facilitate the derivation of the implicature. But this would be predicted with any pitch accent that can lead to focus marking. In our experiments, referents were in focus position and received an H* accent, which should be sufficient for focus marking. While the garden path effects from Experiment 2 show that with both accents participants derive the implicature to a high degree, it does not clearly follow from grammatical accounts why different pitch accents should have differing gradient effects on the insertion of covert operators. One way to reconcile this is to assume that the L+H* makes the alternatives more salient, as we describe above, and therefore a covert operator can be inserted earlier. For example, in the version of the grammatical account by Chierchia (2013), it is assumed that once alternatives are active, they have to be consumed/taken up by an overt or covert operator (e.g., only or its silent counterpart). We find it plausible that the L+H* facilitates implicatures because it has increased the salience of one alternative over another. However, it is largely an open question about the nature of the information that makes one alternative more salient than another in the first place.
Neo-Gricean accounts (e.g., Geurts, 2010; Sauerland, 2004) hold that listeners have to complete the so-called epistemic step to arrive at the strengthened interpretation (derive primary and secondary implicatures). The “epistemic step” account holds that listeners consider alternatives that are relevant to a speaker’s conversational goals. Whether listeners choose to derive an inference will depend on listeners recognizing these goals and whether a speaker has the necessary knowledge to intend one alternative over another. This account is usually portrayed as a set of additional considerations above and beyond structural aspects of an utterance such as linguistic scales and information structure (see Breheny et al., 2013). As we discussed previously, the L+H* accent might remove the necessity to derive primary implicatures first or reduce the need to consider contextual factors relevant for deriving the implicature (Breheny et al., 2013; Sauerland, 2004). Both the facilitation by the L+H* accents in Experiment 1 and the garden path effects in Experiment 2 could equally be explained by such an account. Breheny et al. (2013) show that the epistemic step can be done incrementally: listeners do not need to wait until the end of an utterance to derive pragmatic inferences. The Neo-Gricean account seems to differ from grammatical accounts in that grammatical accounts assume a role of speaker knowledge (and relevance) in the derivation of alternatives. However, the mechanisms by which alternatives are disregarded should be impervious to non-linguistic factors. Therefore, future research should examine the role of speaker knowledge in both the derivation of alternatives with L+H* as well as the process of negating alternatives.
A final consideration is whether pitch accents affect different implicatures in the same way. This question relies heavily on whether the same processing mechanisms are used for different types of implicatures (Bott & Chemla, 2016). If intonation affects the processing of different types on implicatures in similar ways, this would suggest a shared mechanism needed to integrate information from pitch accents and boundary tones. As mentioned, intonation has been shown to increase the likelihood of standard cases of scalar implicatures, for example, with the quantifier some (Schwarz, Clifton, & Frazier, 2008) or the connective or (Chevallier et al., 2008) as well as ad hoc scalar implicatures (Gotzner & Spalek, 2014). Unlike our studies, these studies do not seek to test the exact processing mechanisms, which are ultimately responsible for increase in implicature rates. Future research should also examine how pitch accents work across different types of implicatures to better understand how intonation factors into the inferential procedures, which ultimately influence whether listeners opt for one alternative over another.
Footnotes
Acknowledgements
We are grateful for the help of Elena Chepucova for the data collection and stimuli construction. This paper was based on data published in the 35th Annual Proceedings of the Cognitive Science Society (Tomlinson & Bott, 2013).
Funding
John Michal Tomlinson, Jr. and Lewis Bott were supported by Economic and Social Research Council Grant RES-062-23-2410. The first author was also supported by the Alexander von Humboldt Foundation during the preparation of this manuscript.
