Abstract
Vigilance is the ability to maintain attention on a task over time without becoming distracted. Many high-stakes fields require employees to regularly engage in vigilance tasks (e.g., TSA baggage screening). Understanding strategies to improve accuracy on such tasks can be critical in ensuring efficiency and safety. In the current work, trial-by-trial feedback and extrinsic motivators were tested as potential aids to improve accuracy on vigilance tasks. Measures of state boredom were also collected. Results provide insight that trial-by-trial feedback may be an effective tool to increase accuracy on basic vigilance tasks. Further research is needed to understand the impact of trial-by-trial feedback and extrinsic motivators on state boredom and their impact on applied tasks.
Introduction
Trial-by-Trial Feedback and Sustained Attention
Trial-by-trial feedback can increase learning in some instances, such as with basic visual perceptual tasks, but the relationship depends on the type of feedback (Herzog & Fahle, 1997; Liu et al., 2014). Reliable feedback can improve retention on perceptual learning tasks, but faulty feedback can eliminate evidence of learning (Herzog & Fahle, 1997). The positive impact of trial-by-trial feedback has been primarily studied within this realm of visual perceptual learning, which is learning that occurs with practice of the task (Liu et al., 2014). However, limited research has examined other effects of trial-by-trial feedback on various tasks, such as sustained attention tasks (Robison et al., 2021).
Sustained attention is “a process that enables the maintenance of response persistence and continuous effort over extended periods” (Ko et al., 2017, p. 388). Conversely, vigilance is the ability to sustain attention on a task without becoming distracted by other external factors or influences (Shaw et al., 2010). Vigilance, or performance on a sustained attention task, tends to decline over time (Warm et al., 2008).
Two primary theories explain this phenomenon: Mindlessness Theory (Robertson et al., 1997; Manly et al., 1999) and Resource Theory (Davies & Parasuraman, 1982). Mindlessness Theory explains that the vigilance decrement is due to the monotonous nature of sustained attention tasks. Resource Theory counters this idea, claiming that sustained attention tasks are resource-demanding and stressful (Hancock & Warm, 1989).
Including feedback throughout a sustained attention task has been found to motivate focused participation, potentially by reducing mind-wandering (Robison et al., 2021). Further, no relationship was found between boredom and sustained attention (using a trial-by-trial feedback paradigm) in a sample of undergraduate students (McGough & Mayhorn, 2022). Thus, either boredom is unrelated to sustained attention (unlikely due to the extensive literature providing evidence of this relationship) or including trial-by-trial feedback on a sustained attention paradigm re-attenuates the person to the task at hand, resulting in less boredom from the monotonous nature of the task. This explanation would support Mindlessness Theory and may suggest that including trial-by-trial feedback in sustained attention tasks could reduce the vigilance decrement and improve performance accuracy.
Trial-by-Trial Feedback and Extrinsic Motivation
Combining a distinct goal with feedback has been found to improve performance on sustained attention tasks (Robison et al., 2021). The desire to achieve a goal could be viewed as an intrinsic or extrinsic motivator. On the one hand, an individual may strive to achieve a goal for personal fulfillment. An alternative explanation for working towards a goal could be the desire for recognition. Self-Determination Theory describes these two primary forms of motivation in detail (Deci & Ryan, 1985; Ryan & Deci, 2000). According to the theory, intrinsic motivation is self-driven, while extrinsic motivation involves working towards something due to an external factor (Deci & Ryan, 1985; Ryan & Deci, 2000). Examples of extrinsic motivators include recognition, praise, and monetary compensation. Extrinsic motivators can also be presented as deterrents, such as a warning to avoid a negative outcome or punishment. Adding an extrinsic motivator to an otherwise undesirable or unfulfilling task may facilitate improved performance in various ways (e.g., accuracy, retention, engagement). For example, one study found that a warning prompt (informing the participant that they will be required to restart the task if they score with less than 95% accuracy) resulted in significantly higher accuracy on a sustained attention task (McGough & Mayhorn, 2022). Whether this finding generalizes to tasks outside the laboratory is an important empirical question.
In Application
Airport security is one applied domain where understanding the effects of vigilance is particularly relevant. Since 9/11, the Department of Homeland Security (DHS) has invested billions of dollars in the Transportation Security Agency (TSA)’s ability to ensure the utmost safety of passengers and crew (Transportation Security Agency, n.d.). One direct result of these efforts is the incorporation of X-ray technology to screen passenger baggage for threats. Vigilance is critical in identifying potentially ambiguous threats using this technology (Warm et al., 2008). Previous literature has found performance declines in long baggage screening shift work, and declines were seen as early as 10 minutes after the start of the baggage screening (Meuter et al., 2015). As such, it is essential to understand the effects of vigilance on this sustained attention-based task and identify techniques that can be used to mitigate errors.
Another consideration is how trial-by-trial feedback and extrinsic motivators influence in-the-moment experiences of boredom. Mindlessness Theory argues that monotony and boredom are responsible for performance accuracy declines on sustained attention tasks (Robertson et al., 1997; Manly et al., 1999). Therefore, finding strategies to mitigate boredom, like trial-by-trial feedback and extrinsic caution signals, may help improve accuracy on vigilance-based tasks.
Research Questions and Hypotheses
Based on the literature regarding the effects of trial-by-trial feedback, the following research hypotheses were formulated:
1.) Does trial-by-trial feedback improve accuracy on a basic and applied sustained attention task with and without an indicated performance threshold being required to pass (extrinsic motivator)?
2.) Does trial-by-trial feedback and/or extrinsic motivators influence state boredom?
For hypothesis 1, we expected that the inclusion of trial-by-trial feedback would most improve overall sustained attention performance accuracy on a basic and applied vigilance task when an indicated performance threshold was depicted as an extrinsic motivator. We anticipated that having either condition (trial-by-trial feedback or extrinsic motivation) would also result in improved performance accuracy compared to the control condition. For hypothesis 2, it was expected that including trial-by-trial feedback and an extrinsic motivator as a caution prompt would result in the lowest state boredom levels. Additionally, it was expected that having either condition on its own (trial-by-trial feedback or extrinsic motivation) would also result in lower state boredom levels compared to the control condition.
Method
Participants
An Institutional Review Board approved the current study. Participants were undergraduate students at a large university in the southern United States. All participants were recruited online to fulfill a research requirement in an introductory psychology course. One hundred fifty total participants were recruited for the study. G-Power 3 Software (Erdfelder et al., 1986) was used to conduct a post hoc power analysis to determine the appropriate sample size for analyses. Using an effect size of .50, an alpha of .05, Power set to .80, and group number set to 4, it was determined that at least 48 participants needed to be recruited for adequate power. Thus, our sample size was sufficiently powered for our analyses. For this project, data from 103 participants were assessed.
Tasks
Sustained Attention to Response Task (SART; Robertson et al., 1997)
This online go/no-go task measured sustained attention performance accuracy. Participants viewed a series of number stimuli (0-9). For each number that appeared except for 3, they were instructed to press the SPACE bar on their keyboard. They were asked to withhold their response when they saw the number 3 on their screen. The number appeared on the screen for 250 milliseconds. After 250 milliseconds, the number was replaced by a white cross symbol, a mask covering the number. Participants then had an additional 900 milliseconds to select their response. If they did not select an answer within the time frame, the trial was considered incorrect and automatically continued to the subsequent trial. The task consisted of 18 practice trials and 225 actual trials, which took approximately five minutes to complete. Incorrect responses were coded as a 0, and correct responses were coded as a 1. Overall vigilance performance accuracy was measured through the participant’s total score. A higher score on the SART was indicative of higher accuracy. The SART was administered remotely using PsyToolkit software (Stoet 2010, 2017).
Sustained Attention to Response Task 2 (SART2; Robertson et al., 1997; Stoet 2010, 2017)
In this modified version of the SART, the participant completed the same task as described above. However, this version included visual feedback after each trial. When the participant correctly gave a response or correctly withheld a response, a green circle with a checkmark appeared briefly on the screen. A red X appeared when the participant failed to respond or withhold a response.
Airport Scanner (Kedlin Company)
A baggage screening task was developed using the same instructions and materials initially developed by Pearson (2019). The task was adapted from the iOS game Airport Scanner, available through the Apple App Store and Google Play. During the task, participants viewed sixty images of baggage X-rays on Qualtrics software (Qualtrics, Provo, UT), all of which were images from the Airport Scanner game. Based on their condition, participants received a combination of either the extrinsic motivator caution prompt at the start of the task and/or feedback after each trial (Correct/Incorrect). Participants in each condition viewed the same 60 images. 30% of the images contained contraband that the participant was instructed to identify (Merritt, 2011; Pearson et al., 2019). Participants had three seconds to respond to each trial. If they did not select an answer within three seconds, the trial was considered incorrect and automatically continued to the subsequent trial. Incorrect responses were coded as a 0, and correct responses were coded as a 1. Overall vigilance performance accuracy was measured through the participant’s total score. A higher score on the baggage screening task was indicative of higher accuracy. The baggage screening task was administered remotely using Qualtrics software (Qualtrics, Provo, UT).
Measures
Caution Prompt (McGough & Mayhorn, 2022)
A caution phrase panel was adapted from previous research for the current study. The caution panel was included as the extrinsic motivator in the appropriate conditions, intended to motivate the participant on the task to avoid punishment. The panel appeared on the screen before the participant completed the task, and it read, “
Multidimensional State Boredom Scale (MSBS; Fahlman et. al., 2013)
The MSBS is a self-report measure to assess “in the moment” or state boredom. The measure consists of 29 questions and is broken down into five subscales (Disengagement, High Arousal, Inattention, Low Arousal, and Time Perception). All responses are scored using a 7-point Likert scale ranging from Strongly Disagree (1) to Strongly Agree (7). Participants completed the MSBS three times for the current study. They completed it once at the start of the study. For this paper, only the latter two measurements were examined, taken once after a basic vigilance task (SART) and again after an applied vigilance task (baggage screening task). In response to the basic task, Cronbach’s Alpha value for this questionnaire was .97, indicating excellent internal consistency. For the applied task, Cronbach’s Alpha value was also .97, again indicating excellent internal consistency.
Design
The current study used a 2 x 2 x 2 mixed factorial design. Exposure/non-exposure to the extrinsic motivator caution prompt and presence/lack of trial-by-trial feedback were between-subjects independent variables. Task type (basic or applied) was a within-subjects independent variable. Dependent variables included total score accuracy and state boredom. Demographic information and information about the subjective workload for each task were also collected but were not included in the current analyses due to page limits.
Procedure
Participants for this online study were recruited using the university’s online participant management system (SONA). Each participant completed one of four versions of the study (group 1 = yes trial-by-trial feedback/yes caution, group 2 = yes trial-by-trial feedback/no caution, group 3 = no trial-by-trial feedback/yes caution, group 4 = no trial-by-trial feedback/no caution). When the participants enrolled in the study, they were directed to the university’s Qualtrics server (Qualtrics, Provo, UT). They were assigned a random participant ID number. At this point, they also read and provided informed consent. If the participant was still interested in participating after reading the informed consent, they were redirected to a new link (psytoolkit.org) to begin the study. On PsyToolkit (Stoet 2010, 2017), which is an online psychological experiment repository, participants completed the first survey (MSBS) and the appropriate version of the Sustained Attention to Response Task (basic task) (based on which version they were in). The participants in the caution prompt versions of the study saw a screen that showed the visual caution prompt before beginning the computer task (see Measures). After this task, participants were provided a link to return to Qualtrics to complete the remainder of the study. Within Qualtrics, each participant completed another version of the MSBS that assessed participants’ state boredom experience regarding the previous task. Then, participants completed the baggage screening task (applied task). Trial-by-trial feedback and the extrinsic motivator caution signal were presented for participants in the appropriate versions of the study. After the baggage screening task, participants completed the MSBS once again. At the end of the study, participants were presented with a debriefing statement. In its entirety, the study took approximately 45-60 minutes to complete.
Results
Participants who did not complete the basic and/or applied task were excluded from analyses. Since the basic and applied tasks were timed, individual trials with missing data were coded as incorrect responses. Additionally, any participant with scores three standard deviations above or below the mean for the tasks relevant to that given analysis was excluded from that analysis. Distributions were assessed for normality by visually analyzing each dependent variable’s histograms and the skewness and kurtosis values. All the variables appeared to have normal distributions except for the total score accuracy on the SART, which was negatively skewed.
For each task type, we conducted several one-way ANOVAs. We ran one-way ANOVAs for each outcome variable on the applied task and for the state boredom outcome variable on the basic task. Since the basic task (SART) accuracy variable was negatively skewed, we opted to run a nonparametric Kruskal-Wallis test, given that nonparametric statistics do not assume that the data maintains a normal distribution.
For hypothesis 1, we were interested in understanding the effects of trial-by-trial feedback and caution prompts on performance accuracy. For the basic task, A Kruskal-Wallis test was performed on the accuracy scores of the four groups. Significant differences were found between groups in their accuracy on the task, χ 2 (3, N=101) = 21.52, p<.001). For the ranked dependent variable (SART accuracy), the proportion of variability accounted for by trial-by-trial feedback and/or the caution prompt was .22 (η2 calculated using ranked data).
A pairwise post hoc Dunn’s test with Bonferroni corrections showed that the accuracy scores between groups 1 and 2 significantly differed from groups 3 and 4 (see Table 1, see Figure 1). However, groups 1 and 2 themselves were not significantly different from one another, nor were groups 3 and 4 significantly different from one another. Groups 1 and 2 received trial-by-trial feedback, while groups 3 and 4 received no feedback. Therefore, the groups that received the feedback performed significantly better than those that did not. However, the caution prompt did not significantly affect the accuracy between the groups. We ran a one-way ANOVA for the applied task since the distribution appeared normal. Results showed no significant differences between the groups regarding performance accuracy, F(3,98) = 1.11, p = .35, η2=.03.
Dunn's Multiple Comparison Test - SART Performance Accuracy.
Bonferroni corrections, * p<.05, ** p<.01.

SART Accuracy Between Groups.
For hypothesis 2, we sought to understand the effects of trial-by-trial feedback and caution prompts on state boredom. For the basic task, no significant state boredom differences were reported between the groups, which was contrary to our predictions, F(3,99) = .91, p = .44, η2=.03. Similarly, for the applied task, we did not find any significant differences between the groups in their state boredom reports, F(3, 99) = .96, p = .41, η2=.03.
Discussion
For our first hypothesis, our results provide evidence that the participants who received trial-by-trial feedback performed more accurately on the basic task than the participants who received no feedback. This may provide evidence that trial-by-trial feedback reduces errors on basic vigilance tasks. It also may suggest that trial-by-trial feedback helps keep a participant engaged with a task, resulting in improved accuracy (Robison et al., 2021). However, the caution prompt did not affect accuracy, which was inconsistent with our hypothesis and the findings of our previous study on this topic (McGough & Mayhorn, 2022). These non-significant findings likely indicate that our caution prompt was not an effective motivator, as previous literature provides evidence of external influences' effectiveness as motivators (Deci & Ryan, 1985). Some potential explanations for this discrepancy could be the difference in language used since we referred to the prompt as a “caution” in the current study but as a “warning” in the previous study. If participants felt that “warning” has a more substantial meaning, it may have motivated accurate performance more than when we used the word “caution.” An additional explanation could be the time the basic task was completed in the study. For the current study, the SART was one of the first tasks the participant completed, as opposed to the initial study, where the SART was the last activity they completed. Therefore, it is possible that participants did not consider the caution as heavily since there was less risk of needing to start the task over if they were not yet experiencing fatigue from the study.
When looking at performance accuracy on the applied task, the groups had no significant accuracy differences. This may be because of the applied task itself. The baggage task may have picked up on user errors that were more so due to confusion or genuine errors on the task rather than declines in vigilance. In that case, the caution prompt nor the feedback would have influenced the participants’ accuracy.
For our second hypothesis, we expected the results to provide evidence that trial-by-trial feedback and extrinsic motivators reduce state boredom levels on basic and applied sustained attention tasks. Mindlessness Theory explains that monotony and boredom are potential causes of performance accuracy decrements on sustained attention tasks (Robertson et al., 1997; Manly et al., 1999). Therefore, effectively reducing the state boredom level could improve accuracy over time. However, our findings did not provide evidence that trial-by-trial feedback and extrinsic motivators were worthwhile mitigation strategies to reduce in-the-moment experiences of boredom on our tasks. It may be that the tasks were not long enough to cause the participants to report experiences of boredom, or perhaps trial-by-trial feedback and caution prompts are ineffective strategies to mitigate boredom on the vigilance tasks we used for this study.
The current study has several limitations. With the current design, there was no way to determine when the vigilance decrement occurred. Additionally, participants may have experienced fatigue due to the length of the study. This fatigue may have impacted performance on the tasks. This is particularly relevant since we kept the order of the vigilance tasks the same. All participants completed the basic task first and the applied task second. Future research could partition the sustained attention tasks into several time blocks to assess performance accuracy over time instead of solely looking at total performance accuracy. Additionally, the order of the tasks should be varied to reduce the potential for order effects. Future studies should also utilize longer vigilance tasks to see if longer tasks elicit stronger experiences of boredom, which could influence engagement and accuracy on prolonged tasks. Finally, additional research is needed to explore the perceptual effects of using the word “caution” instead of “warning” and its effect on accuracy and state boredom.
Overall, it is well established that performance accuracy declines over time on vigilance tasks (Davies & Parasuraman, 1982; Robertson et al., 1997). This phenomenon poses significant safety hazards in real-world applications. This study aimed to establish strategies to mitigate various issues on vigilance tasks. Our findings provide evidence for using trial-by-trial feedback to improve accuracy on basic vigilance tasks. This finding may inform employers of one potential strategy that can be utilized to increase performance on certain vigilance tasks in the workplace.
