Abstract
Prior research has shown that task instructions influence the locations and durations of eye fixations during scene viewing. These task-related changes in gaze patterns are likely to be associated with a top-down influence of attention. Presently we applied a saccadic-inhibition manipulation in order to detect another expected manifestation of top-down attention: perceptual enhancement. Participants viewed eight-item arrays containing photographs from two categories of scenes. Four of the photos depicted natural landscapes (“nature”) and the other four depicted urban environments (“buildings”). Participants were instructed to memorize scenes from one of the categories in preparation for a later recognition memory test. During eye fixations the border around the fixated scene flickered briefly from black to white with a random interval between flickers ranging from 400 to 600 ms. We computed the likelihood of a saccade being initiated in the period following the flicker. Consistent with prior research, we found a strong saccadic inhibition effect with maximum saccadic inhibition occurring approximately 97 ms following the flicker. Importantly, the saccadic inhibition effect was stronger and longer lasting when the subject’s eyes were fixated on a relevant scene compared to an irrelevant scene. These findings are consistent with perceptual enhancement as a result of top-down attention.
The relation between visual attention and eye-movement control constitutes an important focus for research on the viewing of natural scenes. High-velocity eye movements (referred to as saccades) serve to align the fovea (the high-acuity area of the retina) with areas of interest in a scene in order to facilitate the acquisition of detailed visual information during the periods between saccades in which the eyes are relatively still (referred to as fixations). While certain basic information about scenes (i.e., “gist”) is processed extremely rapidly prior to the initiation of saccadic eye movements (Biederman, 1981; Biederman, Mezzanotte, & Rabinowitz, 1982; Castelhano & Henderson, 2007, 2008; Greene & Oliva, 2009; Rousselet, Joubert, & Fabre-Thorpe, 2005), subsequent scene processing is typically accompanied by eye fixations to various areas in the scene (Henderson, 2003; Henderson & Hollingworth, 1999).
A variety of bottom-up and top-down attentional mechanisms have been shown to influence eye-movement control in scene viewing (for reviews see Baluch & Itti, 2011; Tatler, Hayhoe, Land, & Ballard, 2011). Specifically, bottom-up factors are typically assumed to produce rapid and involuntary shifting of attention and higher likelihood of fixations being directed towards salient visual features in the scene. In contrast, top-down factors are thought to produce volitional shifts of attention towards scene locations and features that are relevant to behavioural goals as determined by task demands and observers’ knowledge. Demonstrating the importance of bottom-up factors, models that compute saliency based on low-level stimulus characteristics such as contrast, luminance, and edge density have been shown to fairly accurately predict variations in the spatial distribution of eye movements during scene viewing (Itti & Koch, 2001; Itti, Koch, & Niebur, 1998; Parkhurst, Law, & Niebur, 2002). However, since the classic studies by Buswell (1935) and Yarbus (1967), there have been many demonstrations that top-down factors can have strong influences on eye movements during scene viewing. For example, studies investigating top-down attentional influences have manipulated factors such as scene context or task instructions (e.g., Ballard & Hayhoe, 2009; Castelhano, Mack, & Henderson, 2009; DeAngelus & Pelz, 2009; Henderson, Brockmole, Castelhano, & Mack, 2007; Henderson, Malcolm, & Schandl, 2009; Land & Hayhoe, 2001; Rao, Zelinsky, Hayhoe, & Ballard, 2002; Torralba, Oliva, Castelhano, & Henderson, 2006) and demonstrated effects on the location and duration of fixations during scene viewing. Consequently, top-down attentional mechanisms were incorporated into several recent frameworks for modelling eye-movement control in scene viewing including the contextual guidance model (Torralba et al., 2006), the controlled random-walk with inhibition for saccade planning model (CRISP; Nuthmann & Henderson, 2012; Nuthmann, Smith, Engbert, & Henderson, 2010) and the target acquisition model (TAM; Zelinsky, 2008). Furthermore, saliency models that predict the spatial distribution of fixations within scenes have been expanded to incorporate higher order stimulus features, such as objects, faces, and so on (Ehinger, Hidalgo-Sotelo, Torralba, & Oliva, 2009; Navalpakkam & Itti, 2005; Peters & Itti, 2007; for a review see Baluch & Itti, 2011).
In the present investigation, we further explored the impact of top-down attentional influences in the context of a multi-scene viewing task. Specifically, in the present design subjects viewed arrays of eight scenes drawn from two scene categories and were instructed to memorize scenes from one of the two categories in preparation for a later memory test. This instructional manipulation defined half of the scenes in the array as relevant while the remaining scenes were irrelevant to the task. In prior studies using a similar paradigm we demonstrated that viewing times and fixation locations were strongly influenced by the task relevance of scene stimuli, with relevant scenes being fixated more often and for longer duration than irrelevant scenes (Glaholt & Reingold, 2009a, 2009b, 2011, 2012). Building upon these findings, the present study was aimed at investigating the extent to which top-down attentional influences produce perceptual enhancement during fixations on relevant scenes as compared to irrelevant scenes. This prediction is motivated by a growing number of behavioural and neurophysiological studies demonstrating that the allocation of visual attention results in perceptual enhancement of stimuli within the attended region (e.g., Bashinski & Bacharach, 1980; Briand & Klein, 1987; Downing, 1988; Eriksen & Yeh, 1985; Jonides, 1980; for reviews see Carrasco, 2011, 2014; Treue, 2004). In particular, for simple visual discrimination tasks the allocation of visual attention has been shown to improve contrast sensitivity and spatial resolution at the attended region, by way of increased signal strength and noise reduction (Carrasco, 2011, 2014). Moreover, attention has been shown to affect the appearance of the attended stimulus material (Carrasco, 2014; Treue, 2004), such as increasing perceived brightness (Liu, Abrams, & Carrasco, 2009) and contrast (Stormer, McDonald, & Hillyard, 2009). These findings have led to the hypothesis that an increase in the strength of a percept due to attentional enhancement might be indistinguishable from that resulting from increased signal strength (Carrasco, 2014; Reynolds & Chelazzi, 2004; Treue, 2004).
Therefore, in the context of these findings, the goal of the present study was to look for evidence of perceptual enhancement resulting from top-down attentional modulation during scene viewing. To do this we combined the multi-scene viewing task used in our previous studies with the saccadic inhibition paradigm (Reingold & Stampe, 1999, 2000, 2002, 2003, 2004; see also Bompas & Sumner, 2011; Glaholt, Rayner, & Reingold, 2012, 2014; Henderson & Pierce, 2008; Henderson & Smith, 2009; Luke, Nuthmann, & Henderson, 2013; Nuthmann & Henderson, 2012; Slattery, Angele, & Rayner, 2011; Yang, 2009).The saccadic inhibition paradigm was designed to explore the effects of task-irrelevant, sudden-onset visual events on saccades produced during reading, visual search, scene processing, and other basic oculomotor tasks. For example, in a study by Reingold and Stampe (2004), while participants were reading text for comprehension, the text screen was replaced for 33 ms by a black screen at intervals that varied randomly between 300 and 400 ms, resulting in the subjective experience of a flicker. The authors documented a decrease in saccadic frequency as early as 60–70 ms following the onset of the flicker, and this oculomotor effect was referred to as saccadic inhibition. This effect was hypothesized to be a low-level reflex-like oculomotor phenomenon mediated by the superior colliculus (Reingold & Stampe, 2002). Saccadic inhibition has also been investigated in the context of the orienting response (Graupner, Velichkovsky, Pannasch, & Marx, 2007; Pannasch, Dornhoefer, Unema, & Velichkovsky, 2001) and has been linked to a phenomenon known as the remote distractor effect (McSorley, McCloy, & Lyne, 2012; Pannasch & Velichkovsky, 2009; Walker & Benson, 2013, 2014; Walker, Deubel, Schneider, & Findlay, 1997; Walker, Kentridge, & Findlay, 1995). More specifically, the remote distractor effect is observed in a paradigm where subjects are required to make saccades to targets in the visual field. The appearance of an irrelevant distractor shortly after the onset of the saccadic target produces an increase in the latency of the saccade to the target. The relationship between saccadic inhibition and the remote distractor effect is the subject of controversy and is beyond the scope of the present paper, but briefly some studies have suggested that the remote distractor effect is caused by saccadic inhibition (Buonocore & McIntosh, 2008; McIntosh & Buonocore, 2014), while others have argued that they constitute dissociable behavioural phenomena (Walker & Benson, 2013, 2014).
However, most importantly for the present study, saccadic inhibition can provide an indirect measure of perceptual enhancement resulting from the deployment of attention. The prior study of the saccadic inhibition effect by Reingold and Stampe (2004) demonstrated stronger and longer lasting inhibition when the flicker occurred in an attended location than in an unattended location. Such attentional modulation of the strength and duration of saccadic inhibition was attributed to an increase in the saliency of the flicker due to perceptual enhancement within the attended region (for another demonstration of attentional modulation of the strength of saccadic inhibition see Buonocore & McIntosh, 2013). The influence of attention has also been implicated in a scene-viewing distractor paradigm (Pannasch, Schulz, & Velichkovsky, 2011; see also Graupner et al., 2007) where participants studied scenes in preparation for a memory test. Irrelevant distractors appeared during a subset of fixations and resulted in a lengthening of the duration of those fixations. Pannasch et al. (2011) found this effect to be greater for “focal” fixations that were preceded by short saccades than for “ambient” fixations preceded by long saccades. This difference in the distractor effect was attributed to differences in “focal” and “ambient” modes of visual attention during scene viewing.
The goal of the present study was to use saccadic inhibition to provide evidence of perceptual enhancement as a result of top-down attention during scene viewing. Accordingly, subjects viewed multi-scene arrays composed of task-relevant and task-irrelevant scenes, and the border around the fixated scene flickered at random intervals. We computed several measures of saccadic inhibition as a function of the task-relevance of the fixated scene. Given the sensitivity of saccadic inhibition to attentional modulation, we hypothesized that saccadic inhibition would be stronger when it occurred during fixations on relevant than on irrelevant scenes within the multi-scene array, and consequently the strength of the observed saccadic inhibition would constitute an indirect measure of perceptual enhancement due to the influence of top-down attention.
Experimental study
Method
Subjects
The subjects were 12 undergraduate students at the University of Toronto at Mississauga, and each received $10 for their participation. All subjects had normal or corrected-to-normal vision and were naive with respect to the purpose of the experiment.
Apparatus
Eye movements were measured with an SR Research EyeLink 1000 system with a sampling rate of 1000 Hz. Viewing was binocular, but only the right eye was monitored. Following calibration, gaze-position error was less than 0.5°. The stimuli were presented on a 21″ ViewSonic monitor with a refresh rate of 85 Hz and a screen resolution of 1600 × 1200 pixels (38° × 28.5° of visual angle). Subjects were seated 60 cm from the display, and a chinrest with a head support was used to minimize head movement.
Materials and design
Stimuli were colour images drawn from the Corel Image Database and were selected such that they fell into one of two scene categories. The “nature” category was composed of natural outdoor scenes with no man-made objects, and the “buildings” category was composed of scenes that contained one or more buildings. Each scene measured 520 × 350 pixels (12.35° × 8.31°), and in each trial the subjects were shown a composite stimulus array containing four scenes from the nature category and four scenes from the buildings category that were randomly assigned to occupy one of eight stimulus locations (see Figure 1). To create the arrays, scenes were overlaid on a black background, and individual scenes were not repeated across trials. Subjects were required to commit the scenes from one category (i.e., nature or buildings) to memory in preparation for a later memory test. The assignment of the relevant category (nature or buildings) was fixed for each subject and was counterbalanced across subjects.

A sample stimulus for a memory-encoding trial. Participants were instructed to memorize either the “Nature” or the “Building” images from each array in preparation for a later memory test. The border around the fixated image flickered according to a randomly determined interval.
Saccadic inhibition was induced during memory encoding trials via a random-interval flicker. Each flicker event was scheduled to follow the prior flicker event (or else the beginning of the trial) by a randomly determined interval that was between 400 and 600 ms. The flicker was local to the stimulus location being fixated by the subject: The border of the fixated stimulus location (6 pixels in thickness) turned from black to white for two screen refreshes and then back to black again. If the subject was not looking at one of the eight stimulus locations when the timer for the upcoming flicker event expired, that flicker event was skipped, and the subsequent flicker event was scheduled. Accordingly, flicker events occurred during fixations on scenes from both stimulus categories (relevant and irrelevant).
Procedure
A 9-point calibration procedure was performed at the beginning of the experiment, followed by a 9-point calibration accuracy test. Calibration was repeated if any point was in error by more than 1° or if the average error for all points was greater than 0.5°. Subjects were told that the images within each display were either natural scenes (with no buildings) or scenes containing buildings. They were instructed to memorize the scenes from one category in preparation for an upcoming memory test, and to ignore the scenes from the other category because their memory for those images would never be tested (the relevant category was nature for half the subjects and buildings for the other half). Subjects were also told that the border around the images would flicker occasionally, but that they should ignore this and focus on the memory task. Trials were self-paced, and subjects terminated the trial by fixating the centre of the screen and pressing a button on a response pad. Subjects completed 18 blocks of 12 memory-encoding trials, during which the saccadic inhibition manipulation took place (see Materials and Design). After each block of 12 encoding trials, subjects made six 2-alternative forced-choice recognition decisions. On each of these recognition trials, a pair of scenes from the relevant category was shown. One of these scenes was presented during the previous block of trials (i.e., old), and the second scene was not used in any of the trials in the experiment (i.e., new). On half of the recognition trials, the old scene was presented on the left side of screen, and on the other half of trials it appeared on the right side of the screen. Subjects indicated which item was old by pressing a button on the response pad.
Results
Of central interest was the effect of stimulus relevance on saccadic inhibition. In particular, we hypothesized that the task-relevance of fixated material would influence the magnitude of the saccadic inhibition effect in response to the flicker. In order to test this hypothesis we measured the saccadic inhibition effect using a methodology similar to that of Reingold and Stampe (2004). In particular, we analysed the probability of saccades being initiated in the period following the flicker during memory encoding trials. We then conducted a quantitative analysis of the magnitude and duration of the inhibition effect as a function of the task-relevance of fixated material.
Saccadic frequency plots
The times of display changes corresponding to the flicker and the onset time of all saccades were extracted from the eye tracker data files. Saccadic frequency plots were created separately depending on the category of image that the subject fixated at the time when the flicker occurred. In particular, for each subject one of the stimulus categories (nature or buildings) was relevant in the context of the memory task, and the other category was irrelevant. Our analysis focused on the 300 ms following flicker onset, and accordingly we computed the frequency of saccades occurring in of each 60 five-millisecond time bins following the flicker. For each subject, saccade likelihood was averaged across flickers within a trial and across trials. For the average subject there were 5023 saccades following flickers while fixating relevant scenes, and 698 saccades following flickers while fixating irrelevant scenes. Smoothing with a five-bin moving average was applied to the resulting saccadic frequency time course functions.
Baseline
For each saccade frequency plot we computed the average saccade frequency over the five time bins prior to the flicker onset and the five time bins following the flicker event (50 ms in total). This was used as an estimate of the baseline or expected saccadic frequency. Each of the 60 time bins for the 300 ms post-flicker interval was then divided by this baseline saccadic frequency, thereby converting the saccadic frequency plot units into the ratio of measured to expected saccadic frequency. The strength of saccadic inhibition, which is the proportion of expected saccades eliminated by inhibition, can be computed for any bin after the flicker onset by subtracting the normalized saccadic frequency from 1.0. The normalized saccadic frequency plot for each condition, averaged across subjects, is presented in Figure 2.

Plot of saccadic frequency following the flicker onset for fixations on relevant and irrelevant scenes, averaged across subjects.
Measures of inhibition
On inspection of Figure 2, a saccadic inhibition pattern is apparent: Saccadic frequency drops sharply with a maximum inhibition occurring around 100 ms following the onset of the flicker. Importantly, this reduction in saccadic frequency was larger for fixations on relevant scenes than for those on irrelevant scenes. The first step in analysing the saccadic frequency plot for each subject was to locate the minimum saccadic frequency (the “dip”). This was achieved by finding the lowest three-bin average in the interval of 50–175 ms following the onset of the flicker. The latency to maximum saccadic inhibition (LMAX) was then computed as the time interval from the onset of display change to the minimum saccadic frequency. Because the histogram had been normalized, the magnitude of inhibition (henceforth, magnitude) was simply computed as 1.0 minus the average normalized saccadic frequency at LMAX. A magnitude of 1.0 entails that no saccades were observed in the bin corresponding to the centre of the bottom of the dip. As a measure of the temporal onset of the inhibition effect, we computed the latency to 50% of maximum inhibition (L50%). This was defined as the latency from the display change at which inhibition first reached 50% of its maximum strength. To minimize the effects of noise, we actually computed this measure by averaging the latency of all bins prior to the centre of the dip in which inhibition was between 33% and 67% of its maximum strength. Similarly, the time at which inhibition returned to 50% of its maximum strength following the centre of the dip was computed, and the difference between L50% and this value was defined as the duration of inhibition (henceforth, duration). That is, duration corresponds to the period in which inhibition remained above 50% of its maximum strength. The values for these measures, averaged across subjects, are presented in Table 1.
Quantitative measures of saccadic inhibition for fixations on relevant and irrelevant scenes, averaged across subjects.
Note: Standard errors are in parentheses.
p < .05. **p < .01.
As can be seen in the table, the saccadic inhibition effect for relevant and irrelevant images differed according to these quantitative measures. In particular, for relevant images compared to irrelevant images, the saccadic inhibition effect had a later maximum, t(11) = 2.25, SE = 2.22, p <.05, a greater magnitude, t(11) = 3.15, SE = 0.032, p<.01, and a longer duration, t(11) = 2.92, SE =1.92, p<.05. L50%, however, did not differ between conditions, t(11) = 1.29, SE = 1.45, ns; there was no evidence for a difference in the onset of the inhibition. Therefore, consistent with our hypothesis, the saccadic inhibition effect in response to the flicker was stronger during fixations on relevant scenes than during those on irrelevant scenes.
General discussion
In the present study we employed a combination of a multi-scene viewing task (Glaholt & Reingold, 2012) and the saccadic inhibition paradigm (Reingold & Stampe, 1999, 2000, 2002, 2003, 2004) in order to obtain evidence for perceptual enhancement resulting from top-down influences of visual attention during scene viewing. In prior studies using multi-scene arrays (Glaholt & Reingold, 2009a, 2009b, 2011, 2012) we demonstrated that viewers fixate task-relevant scenes more frequently, and for longer duration, than task-irrelevant scenes. Because the relevance of scenes in this paradigm is solely defined by task instructions, it follows that the observed differences in the location and durations of fixations constitute top-down, rather than bottom-up, attentional influences. Similarly, the present findings of stronger and longer lasting saccadic inhibition for fixations on relevant scenes than for those on irrelevant scenes provide additional support for the impact of top-down attentional mechanisms in scene viewing tasks. Taken together, the results from studies using the multi-scene viewing task as well as other demonstrations that task relevance influences the location of fixations during scene viewing (e.g., Castelhano et al., 2009; Henderson, 2003; Henderson & Ferreira, 2004; Henderson et al., 2007, 2009; Navalpakkam & Itti, 2005; Torralba et al., 2006) strongly support the importance of incorporating top-down attentional mechanisms into models of scene viewing (e.g., Ehinger et al., 2009; Nuthmann & Henderson, 2012; Nuthmann et al., 2010; Torralba et al., 2006; Tsotsos, 2011; Zelinksy, 2008).
The finding of stronger and longer lasting inhibition for relevant than for irrelevant scenes is also consistent with prior demonstrations of attentional modulation of the magnitude of the saccadic inhibition effect (Buonocore & McIntosh, 2013; Reingold & Stampe, 2003, 2004). Furthermore, similar to prior demonstrations, the attentional modulation effect on the magnitude of saccadic inhibition was significant at approximately 100 ms post-flicker, indicating an extremely rapid top-down influence. Thus, it appears that the allocation of visual attention intensifies the saliency of the flicker, resulting in stronger and longer lasting saccadic inhibition. More generally, this perceptual enhancement effect provides further evidence for a top-down selection mechanism that operates early in visual processing. The existence of such a mechanism is supported by behavioural studies demonstrating enhanced perceptual discrimination due to the allocation of visual attention (e.g., Bashinski & Bacharach, 1980; Briand & Klein, 1987; Carrasco, 2011, 2014; Downing, 1988; Eriksen & Yeh, 1985; Jonides, 1980; Treue, 2004), and also by demonstrations that the retinotopic representations of attended regions exhibit increased functional magnetic resonance imaging (fMRI) activation in humans and increased neural firing in primate single-unit recordings (for reviews see Carrasco, 2011, 2014; Treue, 2004). Importantly, the present investigation extends this pattern of findings by demonstrating perceptual enhancement due to the allocation of visual attention in the context of scene viewing and by further illustrating the usefulness of the saccadic inhibition phenomenon as an indirect measure of the extent to which visual attention is deployed.
Footnotes
Disclosure statement
No potential conflict of interest was reported by the authors.
Funding
This research was supported by the Natural Sciences and Engineering Research Council of Canada (NSERC).
