Abstract
A dialog consisting of an utterance by one speaker and another speaker’s correction of its content seems intuitively to be made more acceptable when the new information is pitch accented or otherwise focused, and when the utterance and correction have the same syntactic form. Three acceptability judgment studies, one written and two auditory, investigated the interaction of focus (manipulated by sentence position and, in Experiments 2 and 3, pitch accent) and syntactic parallelism. Experiment 1 indicated that syntactic parallelism interacted with position of the new (contrastive) term: nonparallel forms were relatively acceptable when the new term appeared in object position, a position that commonly contains new information (a ‘default focus’ position). Experiments 2 and 3 indicated that presence of a pitch accent and placement in a default focus position had additive effects on acceptability. Surprisingly, spoken dialogs in which the new term appeared in object position were acceptable even when given information carried the most prominent pitch accent. The present studies, and earlier work, suggest that corrected information can be focused either by prosody or position even in spoken English–a language often thought to express focus through pitch accent, not syntactic position.
1 Introduction
Questions and answers are fundamental to language. The semantics of a question, understood as the set of its possible answers (Hamblin, 1973), may underlie the structuring of discourse as in the Question under Discussion approach (Clifton & Frazier, 2012; Ginzberg, 2012; Roberts, 1996/2012, among many others). Further, discourse coherence appears to be enhanced when a question and its answer are similar in form. For instance, in studies of question–answer pairs involving fragment answers to questions, it is clear that answers typically have a parallel syntax to the question. In the classic study by Levelt and Kelter (1982), Dutch shopkeepers were called and asked either the Dutch equivalent of At what time do you close? or What time do you close? The shopkeepers supplied a prepositional answer At 6 o’clock more often if they had been asked a prepositional question. It is important to note, however, that the alternative answer (At what time do you close? 6 o’clock.) is not ungrammatical, merely less common.
Explicit question–answer pairs are not unique in favoring commonality of form (‘parallelism’). Many studies have shown that the comprehension of phrases and sentences is speeded when they are parallel in syntax to a preceding phrase or sentence (Branigan, Pickering, Liversedge, Stewart, & Urbach, 1995; Branigan, Pickering, & Stewart, 1994; Boudewyn, Zirnstein, Swaab, & Traxler, 2013; Frazier et al., 1984; Frazier, Munn, & Clifton, 2000; Sturt, Keller, & Dubey, 2010; Tooley & Bock, 2014), a phenomenon that may motivate speakers to ‘align’ syntactically with their interlocutors (Pickering & Garrod, 2004; but cf. Healey, Purver, & Howes, 2014, for evidence that in natural dialog, factors other than syntactic persistence dominate the choice of form). In the present research, we investigate how the choice of syntactic and prosodic structure affects the naturalness of discourses in which these generalizations are violated. We examine discourses consisting of an utterance and a correction of the utterance, where the corrective portion of the second utterance must receive focus.
There is a deep similarity between the semantics of questions and the semantics of focus. We view focus as a semantic concept. A focused constituent introduces a set of alternative propositions, derived by replacing the focused constituent with a variable whose interpretation is derived by substituting the value of contextually available alternatives. This set contains the contextually supplied alternatives as well as the ordinary semantic value of the focused constituent (roughly, its truth conditions) (Rooth 1985, 1992). Similarly, the meaning of a question is the set of its possible answers, and the answer is a selection from this set. A correction can be seen as supplying both a question and its answer simultaneously. We may think of the dialogue John left. No, Bill left. as essentially having the structure A: John left. B: X left, but was it John? No, BILL left. Semantically, the corrected constituent is focused, introducing a set of alternatives as well as its ordinary semantic value. In a spoken utterance, a focused constituent obligatorily receives a pitch accent (possibly a contrastive pitch accent; Chafe, 1976; Katz & Selkirk, 2011). Some languages, such as Hungarian (E. Kiss, 1995), provide a special syntactic position for a focused constituent, and even in English, which depends largely on prosody to convey focus (Vallduvi, 1993), focused constituents by default occur at the end of a clause (Birner & Ward, 1998; Cinque, 1993; Selkirk, 1984). There is some evidence that, compared to sentences that provide new information, sentences that present a correction incur additional processing difficulty (Benatar & Clifton, 2014). Here we examine how focus, conveyed by pitch accents and by syntactic position, affects the naturalness of corrective sentences, especially when such utterances are not syntactically parallel to a preceding sentence.
Throughout this paper, by ‘correction’ what we mean is that one speaker’s utterance is ‘corrected’ by supplying different information, that is, making a different assertion. We do not mean ‘correction’ in the sense of making a grammatical repair to a sentence containing a grammatical error (although of course the relation between the mechanisms used in the two cases is an interesting topic in its own right). Intuitively, it is natural for the correction to repeat the form of the original utterance, replacing the constituent containing incorrect information, as in (1a). However, as an anonymous reviewer pointed out, Speaker B in (1) has a wide variety of options for making the correction (e.g., No she didn’t; No, Susie did; No, It was Susie; No, I thought Susie did, etc.). One possibility is the nonparallel correction in (1b), which in certain circumstances, can be quite natural. We hypothesize that a nonparallel correction can be felicitous if its focus structure is appropriate to the focus requirements of the correction.
(1) Speaker A: Mary brought the pie Speaker B: a) No, Susie brought the pie. b) No, the pie was brought by Susie.
We suggest that, because of their very clear focus requirements, correctives provide a good vehicle for studying the role of focus in dialogs and how focus can be expressed in both written and spoken language. To our knowledge, the psycholinguistic literature has not contained much discussion of the role of focus in comprehending corrections. One exception is Bornkessel and Schlesewsky (2006), who investigated the processing of focus of various types in German. In German, there is typically an acceptability penalty for scrambled sentences (in which a noun phrase other than the subject appears in clause-initial position). When the sentence appears in an appropriate discourse context (in particular, one that makes the scrambled term given in the sense used by Halliday, 1967), the acceptability penalty may be attenuated or eliminated (Meng, Bader, & Bayer, 1999, as cited in Bornkessel & Schlesewsky, 2006). However, despite the effect of context on judgments of sentence acceptability, Event Related Potentials (ERPs) indicate the continued presence of a scrambling penalty. A frontal-central negativity is observed when the scrambled word is read, despite the word being given in context (Bornkessel, Schlesewsky, & Friederici, 2003). Further, focusing the scrambled term (using a contextual question) did not eliminate this negativity (Bornkessel et al., 2003; Bornkessel & Schlesewsky, 2006), but resulted in a parietal positivity. In contrast, when the scrambled term was a correction of a prior assertion, the frontal-central negativity in the ERP data was eliminated, indicating the disappearance of the local cost of scrambling (Bornkessel & Schlesewsky, 2006). This suggests that a predictable relation exists between a sentence and its correction and that the focusing requirements of this relation overcome the local processing cost of scrambled sentences.
The following experiments examine the role of focus in determining the acceptability of both written and spoken corrective sentences. They concentrate in particular on how an appropriate focus structure (or information structure; Vallduvi & Engdahl, 1996) affects the acceptability of corrections that are nonparallel to the sentence they correct. Intuition, as well as experimental evidence (Bock & Mazzella, 1983, inter alia), makes it clear that in spoken language, processing of even parallel sentences (corrections or otherwise) is facilitated by having a pitch accent on the focused term (or on an argument of the term; Birch & Clifton, 1995); sentence (1a) sounds good with pitch accent on Susie, bad on pie. The auditory experiments reported below (Experiments 2 and 3) extend this observation to nonparallel corrective sentences.
The experiments also investigate whether pitch accents are the only way to convey focus in English. English is commonly thought of as being a language where (in speech) focus is primarily signaled by pitch accent.
1
An alternative is that English, like many other languages, contains syntactic positions that are relatively appropriate for focused material. In English, default focus falls on the final constituent of a sentence, or (equivalently in the materials we will be studying) in the predicate (Birner & Ward, 1998; Cinque, 1993; Selkirk, 1984; cf. the discussion of how given information tends to appear early in a sentence and new information late; Haviland & Clark, 1974). Some evidence exists that both readers and listeners seem to treat material that appears in object position as being focused. For instance, Carlson, Dickey, Frazier, and Clifton (2009) (see also Frazier & Clifton, 1998) examined ambiguous sluicing sentences like (2) in an auditory questionnaire:
(2) The captain talked with the co-pilot, but we couldn’t find out who else. a. We couldn’t find out who else talked with the co-pilot. (subject antecedent) b. We couldn’t find out who else the captain talked with. (object antecedent)
Participants were presented with auditory versions of (2) with the matrix subject or the object (or both or neither) accented, and were asked to choose between two answers, one of which (2a) indicated that the subject was the presumably focused antecedent of the sluice (who else…), and the other (2b) indicated that it was the object. Both accent and syntactic position affected answer choice: there was a bias toward choosing the object as the antecedent, which was counteracted by an equally strong bias to choose the accented term as antecedent. Under the assumption that the preferred antecedent of the sluice is the focused constituent in the earlier part of the sentence (see discussion in Carlson et al., 2009), this indicates that both pitch accent and syntactic position play a role in assignment of focus.
Setting clefts aside, some evidence does exist, however, suggesting that only pitch accent, and not sentence position, conveys focus in spoken English. Bock and Mazzella (1983) investigated exchanges including sentences like those in (3), examining the role of parallelism and pitch accent on comprehension time. Either the agent or the patient of the second sentence was accented.
(3) a. Evelyn kissed Jeremy. Rhonda kissed Jeremy too. b. Evelyn kissed Jeremy. Jeremy was kissed by Rhonda too.
Comprehension time was faster when the second sentence was active (and parallel to the first) than when it was passive, and faster when the pitch accent fell on the new term (Rhonda in the example) than on the given term. There was a marginal interaction between these factors, with the effect of accent placement being larger for passives than for actives, an effect that is difficult to interpret given the generally longer times for passives than actives. Of particular relevance to the question of how focus is conveyed, there was no significant interaction between the appropriateness of accent placement (new versus given) and the position of the new item (subject versus by-object). Bock and Mazzella do not deny that in written English new information generally appears late in a sentence (Halliday, 1970; Haviland & Clark, 1974). However, they cautiously suggest that in spoken English, ‘intonation is the primary indicator of information structure’ (p. 72). In fact, their data provide no evidence that sentence position of the new term plays a role in facilitating comprehension of the sentences in the exchanges they studied.
The studies reported below explore two alternative hypotheses about the appropriate expression of focus in dialogs that clearly require appropriate focus marking: (1) in spoken English, focus is expressed primarily or solely by pitch accent; (2) in both written and spoken English, sentence position as well as pitch accent can play a role in the expression of focus. In addition, the studies allow us to observe the interplay of pitch accent and position: Do they interact or are their effects additive? Experiment 1 is a written study of dialogs consisting of an assertion and a correction, designed to see if syntactic position of the corrected term affects acceptability judgments in discourses where the corrective sentence is either active or passive, and thus parallel or not parallel to the initial sentence. Experiments 2 and 3 are auditory studies designed to see if the presumed effect of pitch accent on the given versus the contrastive term in nonparallel responses is modulated by the syntactic position of the contrastive term in spoken English.
2 Experiment 1
In the first experiment, readers judged the naturalness of short discourses like those in Table 1. Each dialog appeared in four versions. In the syntactically parallel versions, both the initial sentence and its correction were in the active voice; in the nonparallel versions, the correction was in the passive voice. 2 This parallelness manipulation was crossed with a position manipulation. In English, the default focus falls on the final constituent of a sentence, where new information typically appears, the grammatical object or the agentive by-phrase in the current cases (Birner & Ward, 1998; Carlson et al., 2009; Haviland & Clark, 1974; Selkirk, 1984). The corrective phrase appears in sentence final position, where it would be by default interpreted as focused, or in sentence initial subject position, where given or topical material typically appears, generally without focus (except in the case of ‘contrastive topics’). We predicted that syntactically parallel corrective sentences would be preferred over syntactically nonparallel ones, but further, that this preference would be modulated by position/default focus. When the corrective term appears in a default focus position, a reader will implicitly interpret it as focused. Since the corrective term, as the answer to an implicit question, needs to be focused to be felicitous, placing it as the final phrase of a sentence provides some justification for using a passive correction to an active sentence, potentially ameliorating the penalty for having the correction be nonparallel to the initial sentence.
Example of dialogs, Experiment 1 (A = Speaker A; B = Speaker B).
2.1 Method
2.1.1 Materials
Sixteen two-sentence discourses were constructed, with four versions of each, as illustrated in Table 1. The four versions were defined by the factorial combination of Parallel (active correction) versus NonParallel (passive correction) and Focus Position (corrective term in sentence-final position) versus NonFocus Position (corrective term in subject position). All experimental items appear in the Appendix. These experimental items were combined with four practice items and 48 sentences or two-sentence discourses from other unrelated experiments (involving the naturalness of non-corrective continuations of a discourse and the plausibility of single sentences).
2.1.2 Participants and procedures
Forty-eight University of Massachusetts undergraduates were tested in individual half-hour sessions, receiving course extra credit for their participation. They were instructed to rate the second sentence of a two-person dialog on a five-point naturalness scale. They were told that if the sentence felt natural, and was easy to understand, they should give it a high naturalness rating, but if a sentence was even a little bit difficult or if there was anything odd about it, they should give it a lower rating. Following four practice items, they received the 64 items (including the 16 experimental sentences) in an individually randomized order. The four versions of each item were distributed across four lists using a Latin square design, so that each participant saw four items in each of the four versions, and each item was shown to 12 participants in each version. Items were presented, and responses recorded, using the program Linger (Rohde, 2003). The first sentence was identified as coming from SPEAKER A, and the second from SPEAKER B. Subjects were told to read each sentence or dialog for understanding, and when they had understood it, to press the space bar on a computer keyboard. They then saw a five-point rating scale that asked them to judge how natural SPEAKER B’s response was, with 1 = ‘very unnatural’ and 5 = ‘very natural.’ Ratings were recorded by the computer.
2.2 Results
The mean ratings appear in Table 2. They were analyzed with a linear mixed model analysis with random intercepts and interacting slopes for participants and items, using the statistical programming language R, version 2.15 (R Development Core Team, 2012) and the lme4 package, version 0.999999. Sum coding was used, so that main effects can be interpreted as they are in familiar analyses of variance (ANOVAs). The results are straightforward: each main effect and the interaction was fully significant (since df in this analysis is large but uncertain, so that the t distribution approximates the z distribution, a t > 2.0 will be taken to be a significant effect; Baayen, Davidson, & Bates, 2008): parallelness (b = −0.58, SE = 0.07, t = −7.77), position (b = 0.27, SE = 0.04, t = 6.07), and the interaction (b = 0.40, SE = 0.08, t = 4.87). Corrections that were syntactically parallel to the initial sentence were preferred, as were corrections where the correction term appeared in the default focus position, but the interaction indicated that the latter effect was limited to the nonparallel corrections. 3
Mean naturalness ratings (5 = very natural) and SEs, Experiment 1.
2.3 Discussion
The results of the experiment suggest that parallelism is not actually the only factor in determining a well-formed correction. What matters is information structure: the corrected information should be in the same position as the material being corrected or the corrected information should be in default focus position. What is not acceptable is if the corrected information is not in its original position within the syntax and not moved to the position where focus normally resides (in the predicate, not the subject, Carlson et al., 2009, a.o.).
In examples like those tested, correcting a constituent in object position by placing the constituent in subject position creates a drop in acceptability/naturalness because the speaker has done an odd thing: placing the corrected information neither in its original position nor in the position of (informational) focus, as in the NonParallel, NonFocus version in Table 1. On the other hand, when the incorrect constituent was in subject position originally, placing the correction in object position does not seem to incur much of a penalty, as in the NonParallel, Focus position example in Table 1. Apparently, it is acceptable to focus the correction by placing it in the default focus position. Indeed, that may be the role of scrambling in the German study discussed in the introduction (Bornkessel & Schlesewsky, 2006).
In a written experiment, pitch accents are not explicitly provided, although the reader presumably supplies them perhaps in an implicit prosodic representation (Bader, 1998; Fodor, 1998). Because informational focus, that is, focus associated with discourse new material, is typically in the predicate (Carlson et al., 2009, and references therein), readers may expect a pitch accent and thus place an implicit nuclear accent on the predicate, resulting in the relatively high acceptability of the NonParallel Focus condition. Experiment 2 asks whether a similar phenomenon occurs in a spoken dialog: will the expectation that focus appears in predicate position result in acceptability of an utterance where the corrected term does appear in the predicate, even though the subject and not the predicate is accented?
3 Experiment 2
Experiment 2 investigated discourses like those in Table 3. In all versions, the first sentence was in the active voice and the second sentence, the correction, was in the passive voice. What varied is whether the new (correction) term appeared in object (of the passive by-phrase) or in subject position, and whether it was accented or not.
Example of dialogs, Experiment 2 (A = Speaker A; B = Speaker B).
Note: UPPERCASE indicates a pitch accent. The experimental conditions were the same as the Experiment 1 nonparallel conditions with the addition of the accent-placement factor.
If in Experiment 1 the lowered naturalness judgments of the nonparallel example with the correction in subject position were due to the lack of an implicit pitch accent/focus in subject position, the unnaturalness should disappear if an accent is explicitly provided. In general, information that is new, including information that corrects a previous assertion, must be focused (Schwarszchild, 1999; Selkirk, 1984) and, thus, in spoken language, it should be accented (perhaps receiving a ‘contrastive focus’ accent; Chafe, 1976; see Katz & Selkirk, 2011, for additional discussion). If the results are governed by this principle and this principle alone, then the correction sentences in the Accented column of Table 3 should be rated as highly acceptable and those in the NonAccented column should not be.
However, if our account of Experiment 1 is correct, one response that was rated unacceptable in Experiment 1 (the passive response with a correction in subject position) will be rated acceptable when it appears with a pitch accent on the subject, and the response that was relatively acceptable in Experiment 1 (the passive response with a correction in object position) will remain acceptable even when the object is accented. The remaining nonparallel responses will be rated unacceptable, because the correction term is not presented as being focused.
3.1 Method
3.1.1 Materials
The nonparallel versions of the 16 sentence used in Experiment 1 were recorded in a sound-deadened chamber. A male speaker spoke the first sentence, and a female speaker (trained in phonological analysis) replied. She placed a prominent pitch accent on the ‘focused’ term indicated in Table 3. The accent was generally an H* accent (Pierrehumbert, 1981; Pierrehumbert and Hirschberg, 1990), the accent appropriate for new information, but sometimes had more pitch movement than is typical of an H* accent, and could be considered to be an L+H* accent, sometimes referred to as contrastive accent (Katz & Selkirk, 2011). When the corrected (new) but unaccented term was in object position, it consistently received no H* accent. When it was in subject position, the new term was either unaccented (but not reduced) or (presumably because of phonological factors; Ladd, 1996; Shattuck-Hufnagel, Ostendorf, & Ross, 1994; Selkirk, 2000; Terken & Hirschberg, 1994) received a small H* accent. Typical waveforms appear in Figure 1, where the CAPITALIZED word received a pitch accent. These dialogs were combined with 113 other sentences and two-sentence discourses, fillers and, discourses from other unrelated experiments, some of which required naturalness ratings similar to those of the experimental sentences, and some of which required comprehension questions to be answered.

Illustrative waveforms and pitch tracks, Experiments 2 and 3: (a) by-object accented; (b) subject accented.
3.1.2 Participants and procedures
Forty-eight University of Massachusetts undergraduates were tested, as in Experiment 1. The procedures used were the same as in Experiment 1, except that the participants, seated in a sound-deadened chamber, heard the dialogs played over speakers by a computer, rather than reading them. Once they heard and understood the dialog, they pressed the space bar on a computer keyboard. They then saw a seven-point naturalness judgment question similar to the five-point scale used in Experiment 1, and reported their judgment by pressing a number key on the computer keyboard.
3.2 Results
The mean naturalness ratings appear in Table 4. The data were analyzed as in Experiment 1, using sum (ANOVA-style) contrasts for the two fixed effects of position of the corrected term (object of by-phrase versus subject) and accent on the corrected term (present or absent). Interacting random slopes by subjects and items were included in the model. Ratings were higher when the corrected term carried a pitch accent than when it did not (b = 0.41, SE = 0.08, t = 4.92) and when it appeared in object position than in subject position (b = 0.34, SE = 0.08, t = −4.14), while the interaction between the two factors was nonsignificant (b = 0.05, SE = 0.07, t = 0.71). These results confirm the predicted importance of an accent, as well as the importance of subject versus object position that was observed in Experiment 1. Overall, placing a pitch accent on the corrected term increased naturalness ratings, and ratings were higher when the corrected term appeared as the final object of the by-phrase than when it appeared in subject position. The two effects did not interact.
Mean naturalness ratings (7 = very natural) and SEs, Experiment 2.
However, the results did not fully support our expectation that the presence of a focusing pitch accent on the corrected term would make it completely natural. A contrast of the two conditions where the corrected term was Accented (conducted using a linear mixed model like the one described above but with treatment coding in which Accented was the baseline for the Accented/Not Accented factor) indicated that judgments were higher when the corrected term appeared in object position than when it appeared in subject position (5.97 versus 5.20; b = −0.77, SE = 0.15, t = −5.08). Further, a similar difference appeared when the corrected term was not Accented (5.06 versus 4.48; b = −0.58, SE = 0.24, t = −2.43). One could simply note that these two significant effects reflected the fact that the main effect of correction position did not interact with the presence of pitch accent, indicating independent contributions of syntactic position and pitch accent. However, this conclusion suggests that the effect of syntactic position cannot be reduced to position’s role in guiding implicit prosody (with predicate position encouraging an implicit pitch accent), as suggested in our discussion of Experiment 1. Implicit prosody, and thus default focus position, should have played no role when an explicit pitch accent is present, but it apparently did. We will present other possibilities in the Discussion, and go on to test one of them in Experiment 3.
3.3 Discussion
One possible reason for the persisting influence of syntactic position of the contrasting term in Experiment 2, despite the presence of an explicit pitch accent elsewhere, is that assignment of focus is not exhaustively determined (in English) by the location of a pitch accent. It may be that identification of likely newness imparts focus, even to a non-accented term, and since new information tends to come late in sentences (Haviland & Clark, 1974), focus is assigned to a non-accented final term. This imputation of focus could account for the relative acceptability of the discourses with a non-accented contrastive term in object position, but it clashes with the widespread observation that new or contrastive terms must have focus (Schwarszchild, 1999, Selkirk, 1984, among others). An alternative possibility is that, in the process of arriving at an acceptability rating, the experimental subjects might have implicitly repeated the sentence to themselves, and judged the result of this repetition. Given the observed tendency for speakers to place nuclear accent on the final phrase of a sentence, these implicit repetitions might have tended to have an apparent pitch accent on the final phrase, the agent, increasing the apparent acceptability of sentences in which the new, contrastive term was in final position. Experiment 3 was intended to minimize this possibility, by requiring speeded acceptability judgments, which presumably would not allow time for this implicit recitation. 4
4 Experiment 3
4.1 Method
4.1.1 Materials
The 16 nonparallel correction sentences used in Experiment 2 were used in Experiment 3 (see Table 3 for an example). They were combined with a total of 36 other sentences from unrelated experiments, 16 of which were possibly confusing sentences with reduced relative clauses and 20 of which were syntactically simple but (in half the instances) contained an unusual prosody. These filler sentences were all included in the filler sentences used in Experiment 2. Eight practice discourses were presented before the experimental items began.
4.1.2 Participants and procedures
Forty-eight University of Massachusetts undergraduates were tested in individual 15-minute sessions, receiving course credit. The procedures were generally the same as in Experiment 2, except that participants made speeded acceptability judgments. They were instructed to judge each sentence (or for the experimental items, the second sentence of the two-person dialog) for whether or not it sounded ‘normal and natural.’ They were told to respond with as little delay as possible. Every trial began with the participant seeing the following on a computer terminal: ‘Press any key to hear next item, and then press up-arrow if it sounds OK and down-arrow if it doesn’t.’ When the participant pressed a key on the computer keyboard, a sentence was presented over speakers. When the sentence ended, the message ‘OK up-arrow, Not OK down-arrow’ appeared on the computer screen, and the subject’s response and reaction time was recorded. The program PsychoPy (Pierce, 2007) was used to present materials, and four different counterbalancing conditions were used so that each item was tested equally often in each of its four versions, and each subject saw four items in each version.
4.2 Results
The mean proportion of ‘acceptable’ responses, and the mean reaction time (for all responses) appear in Table 5 (together with standard errors).
Top panel: mean proportion ‘Acceptable;’. Bottom panel: RT (se of ‘Acceptable’ responses, s, Experiment 3. Standard errors in parentheses.
The data were analyzed as in Experiment 2, except that the logits of the proportion ‘acceptable’ responses were analyzed (Jaeger, 2008). The proportion of acceptable results was very similar to Experiment 2. A logistic mixed effects analysis (with random interacting slopes and intercepts, and sum coding of the fixed effects) indicated significantly more acceptances when the corrected term carried a pitch accent than when it did not (b = 0.87, SE = 0.21, z = 4.08, p < .001) and when the corrected term was in object position (by-agent) than when it was in subject position (b = 0.99, SE = 0.13, z = 7.43, p < .001). The interaction was thoroughly nonsignificant (b = 0.01, SE = 0.13, z = 0.11).
The ‘acceptable’ reaction times are of interest primarily because they indicated that subjects generally responded in less time than it would have taken to rehearse the sentence (a grand mean of 1.3 s, timed from the end of the sentence). Numerically, they seem to pattern very much like the proportion of ‘acceptable’ data. A linear mixed model analysis (sum contrasts, random slopes, and intercepts) indicated that ‘acceptable’ responses were made more quickly when the object/agent was new than when the subject was new (b = −0.15, SE = 0.06, t = −2.47), but the effect of which term was accented was not significant (b = 0.10, SE = 0.08, t = −1.38) nor was the interaction (b = 0.05, SE = 0.07, t = 0.70). Overall, ‘acceptable’ responses were faster than ‘unacceptable’ responses, 1.26 versus 1.43 s.
4.3 Discussion
Experiment 3 reinforces the suggestion made in discussing the results of Experiment 2: focus in spoken English is not determined entirely by pitch accent, even implicit pitch accent. Position in a sentence is associated with newness, hence focus, and thus presenting an unaccented new (contrastive) term in final predicate position is not as unacceptable as one might have expected, given its lack of accent. As in Experiment 2, presence of a pitch accent on a new term and appearance of the new term in predicate position had independent, additive effects on the acceptability of spoken sentences. It may well be correct that new and contrastive terms must have focus (Schwarszchild, 1999; Selkirk, 1984), but a pitch accent may not be the only way to convey focus in English.
5 General discussion
The account we have provided makes a variety of predictions concerning how syntactic structure, pitch accent, default focus position, and discourse structuring interact. Turning first to syntactic parallelism between antecedent sentence and correction, the results of the experiments suggest that parallelism in the structure of the antecedent and correction is not required. The speaker may deviate from the structure of the antecedent, but not randomly. Deviations must be motivated. In the present examples, syntactic deviations from the structure of the antecedent are relatively acceptable providing that they serve to place the corrected constituent in a position where focus is expected, for example, in the predicate, especially as the final constituent. In other words, information structure determines whether a nonparallel syntax is acceptable for a correction.
Accent alone does not govern the felicity of a discourse. Although it is easy to imagine a theory in which the syntax of a correction matters only when the prosody does not clearly mark the corrected constituent, this is not what was observed in Experiments 2 and 3. Both the prosody and the syntax matter. Moreover, they seem to function independently rather than, for example, the prosody mattering most when the syntax does not mark the focus of the sentence. The fact that both prosody and syntactic position matter with respect to the processing of focus is not a new observation. In the study of sluicing (Carlson et al., 2009; Frazier & Clifton, 1998), elided constituents prefer focused antecedents and both the presence of a pitch accent and the syntactic position of a constituent clearly matter. Surprisingly, in the present data, as in the auditory sluicing data, the syntactic position has an effect even when there is a very prominent clearly marked pitch accent that might be taken to indicate a salient focus.
What is not clear at present is whether listeners treat syntactic position on a par with semantic focusing, for example, in Cutler and Fodor’s (1979) classic study where semantic focus, provided by a preceding question, behaved on a par with a pitch accent. In this case, various types of information would feed a listener’s assumptions about whether the speaker focused a particular constituent. (See Du Bois (2003) for discussion of cross-language corpus counts indicating that by far the most common clause type is one with only a single new/lexical argument, and data indicating that there is generally a particular position where that new argument is introduced.)
As an alternative to treating syntactic position on a par with semantic focusing, it is possible that the expectation that a focused constituent will appear in a particular position in a sentence might lead to perception of the expected prominence. On a par with other illusions, for example, the experience of hearing a period of silence between sentences when no period of acoustic silence is present, listeners might perceive the input as being somewhere between the objective input and their expectations. In short, they may hear a constituent as being more prominent than it is under circumstances where prominence is expected (cf. Cole, Mo, & Hasegawa-Johnson, 2010). In the present study and the auditory study in Frazier and Clifton (1998), listeners responded as if a constituent late in the clause was focused despite the presence of a very prominent pitch accent elsewhere in the sentence. This makes us lean toward an account based on focus perception rather than on pitch accent or prominence perception.
Why did Bock and Mazzella (1983) not observe an effect of position as well as of pitch accent in the listening studies described in the introduction? In principle, it might be the nature of their materials where information in a second sentence was additive, rather than corrective (recall that Bornkessel & Schlesewsky, 2006, discussed in the Introduction, showed that corrective focus had different effects from question answering focus). Alternatively, the fact that ‘repeated names’ appeared in their materials, even in subject position where pronouns are highly preferred, might have led listeners to accommodate unusual discourse circumstances to countenance the unusual information structure. Although the goal should certainly be to construct a general theory of focus and focus processing, not separate theories for corrections, for sluicing, for additive sentences, etc., corrections such as those investigated in the present experiments and ellipsis structures, such as sluicing, in principle might have different conditions. Further work would be needed to insure that a single account is possible at least for all structures within a given language.
Generally, it is assumed that English is a language in which focus is expressed by pitch accents, not by syntactic position (e.g., Vallduvi, 1993, Vallduvi and Engdahl, 1996). The present results raise the possibility that English is more like languages generally believed to use both syntactic position and pitch accent to mark focus, such as Spanish and Italian (Face & D’Imperio, 2005). We do not suggest that there is no difference between English and these other languages with respect to the expression of focus, but simply note that showing positional effects of focus in English, in particular for constituents in final position, makes English look less purely like a ‘pitch accent’ language.
In summary, it seems that both prosody and syntax independently influence the acceptability of a correction. One obvious question is whether prosody and syntax contribute to processing other structures in a similar manner or whether there is something special about corrections. As already mentioned, in studies of sluicing (Carlson et al., 2009; Frazier & Clifton, 1998) both prosody and syntactic position contribute to determining focus. What these studies suggest is that in any kind of what we might think of as an ‘anaphoric’ structure, be they ellipsis structures, (presumably) question–answer pairs, or corrections, listeners and readers are expecting structural deviation from the first sentence to occur primarily when there is some reason for it. Placing a focused constituent into a position where focused constituents usually occur counts as a good reason.
Footnotes
Appendix
Materials used in Experiment 1 are presented. Experiments 2 and 3 used the nonparallel (a) and (c) versions of these items, and manipulated whether the subject (theme) or by-object (agent) received an H* pitch accent.
Acknowledgements
We would like to thank Amanda Rysling, Kiah Atkinson, Christopher Cuna, Jennifer Dimian, Adina Gallili, Brittany Stepton, and Rose Underhill for assistance in conducting the reported experiments.
Funding
This research was supported in part by Grant Number HD18708 from NICHD to the University of Massachusetts. The contents of this paper are solely the responsibility of the authors and do not necessarily represent the official views of NICHD or NIH.
