Abstract
Observers can voluntarily avoid reversals of an ambiguous, reversible figure, extending the duration of an intended percept. This is usually attributed to high-level, top-down attentional processes. However, voluntary control is limited. Reversals occur despite attempts to avoid them. In two experiments, observers demonstrated significant, but limited, voluntary control over Necker cube perception. Cube size and cube completeness, variables associated with stimulus-driven processes involving neural adaptation, influenced the frequency of reversals regardless of observers’ intentions. Results are consistent with the hybrid hypothesis that both top-down and bottom-up processes contribute to Necker-cube perception and support the hypothesis that the contribution of bottom-up processes is responsible for the limitation on voluntary control.
Reversible figures are ambiguous visual patterns that support at least two different, mutually exclusive, perceptual organizations or interpretations. Their defining characteristic is that, during a period of continuous viewing, observers’ perception alternates between possible interpretations of the figure. When one percept is dominant, other percepts are suppressed or inhibited. Reversible figures have a rich history in human experience and scientific pursuit (Wade, 2017). Perhaps the most famous, most studied example is the Necker cube (Figure 1A), which was introduced in 1832 (see Boring, 1942). Interest in reversible figures has been maintained by the belief that their multistable nature provides a useful lens through which the operation of fundamental perceptual mechanisms can be observed (Long & Toppino, 2004).

Complete Cube (A) and Incomplete Cube (B), both with a central fixation cross.
Historically, explanations of figure reversals have emphasized one of two competing views. (See Long & Toppino, 2004, for a review.) In one, it is hypothesized that an ambiguous stimulus can be represented by two or more reciprocally inhibitive neural structures. Reversals are attributed to lower-level, stimulus-driven, or bottom-up processes such as adaptation and recovery of these neural structures (e.g., Blake et al., 2003; Toppino & Long, 2015; Wallis & Ringelhan, 2013). With continuous viewing, the active or dominant neural structures adapt at a rate that is related to the magnitude of their activation (Klink et al., 2008a). Observers experience fluctuating perceptions as the competing underlying neural structures alternately adapt and recover (e.g., Chholak et al., 2020; Dornic, 1967; Lankheet, 2006). 1
Evidence consistent with the neural-adaptation hypothesis includes the fact that the reversal rate (frequency of reversals) is modulated by stimulus factors such as the figure’s size (e.g., Toppino & Long, 1987) and completeness (e.g., Babich & Standing, 1981). When observers simultaneously view multiple copies of a reversible figure (e.g., multiple Necker cubes), the different stimuli can, and often do, reverse independently (e.g., Long & Toppino, 1981). Furthermore, the reversal rate for each of two simultaneously presented copies of a reversible figure does not differ from the reversal rate for a figure presented alone to the same retinal location (Long et al., 1983). In both the single- and multiple-figure conditions, the rate of reversal shows a negatively accelerated increase during an extended period of continuous viewing (Long et al., 1983), although explaining the increase requires the additional assumption that recovery is incomplete on each successive reversal cycle (e.g., Long & Toppino, 2004). Other evidence indicates that this process is localized such that the increased reversal rate resulting from extended exposure is reduced or eliminated if the figure is moved to a retinal location that maps onto a different cortical area (e.g., Blake et al., 2003; Spitz & Lipman, 1962) or if the size of the ambiguous stimulus is changed (Long & Moran, 2007; Toppino & Long, 1987). Also, when observers adapt one perceptual interpretation of a reversible figure by viewing an unambiguous version of the stimulus for an extended period, they perceive the opposite interpretation when they subsequently view the ambiguous version of the stimulus (e.g., Toppino & Long, 2015).
The second theoretical perspective emphasizes high-level, central processes, such as attention, that exert a top-down influence on perception. For example, reversals may be attributed to shifts in attention that activate or enhance one high-level representation and/or inhibit or depress activation of the competing, high-level alternative (e.g., Leopold & Logothetis, 1999; Meng & Tong, 2004). Consistent with this view is the finding that observers often do not perceive figure reversals unless they are aware of their potential occurrence (e.g., Rock et. al., 1994), although contradictory evidence has been reported more recently (e.g., Mitroff et al., 2006). Observers also are more likely to perceive the perceptual alternative that prior experience has set or primed them to expect (e.g., Leeper, 1935; Long et al., 1992). Perhaps the strongest evidence for the contribution of top-down processes, however, is that observers can exert voluntary control over the perception of reversible figures (e.g., Sato et al., 2020; Seth & Reddy, 1979; Toppino, 2003), shortening or lengthening the period during which one or both percepts are dominant.
Many researchers have proposed a hybrid account of figure reversal with a role for both stimulus-driven, bottom-up processes and higher-level, top-down processes (e.g., Hochberg, 1968; Klink et al., 2008a; Kornmeier & Bach, 2012; Long & Toppino, 2004; Peterson & Hochberg, 1983; Toppino & Long, 2005). From this perspective, the primary goal is to understand how lower-order and higher-order processes are coordinated in determining the perception of reversible figures. The present paper takes this approach to studying the limits of voluntary control.
Observers viewing an ambiguous, reversible figure can both induce and suppress reversals voluntarily (e.g., Mathes et al., 2006; Meng, & Tong, 2004; Seth & Reddy, 1979). The primary concern in the present paper is observers’ ability to suppress reversals and the limits on this ability. Both are on display when observers are instructed to hold a particular interpretation of the reversible figure for as much of a viewing period as possible. Voluntary control is demonstrated to the extent that observers in a Hold condition are able to increase the time during a viewing session that a designated interpretation is perceptually dominant relative to participants in a No-Hold control condition in which no attempt is made to influence perception (e.g., Liebert & Burk, 1985; Meng & Tong, 2004; Peterson & Hochberg, 1983; Toppino, 2003; van Ee et al., 2005). However, voluntary control is clearly limited in that observers cannot suppress reversals indefinitely and inevitably perceive the unintended interpretation for a portion of the viewing period. Such involuntary intrusions of the undesired percept occur frequently with ambiguous figures (e.g., Allen et al., 2016).
Both voluntary control and its limits were demonstrated in a series of experiments by Hochberg and Peterson (1987) that were designed to study piecemeal perception. Observers viewed a rotating Necker cube in which apparent reversals of depth are perceived as changes in the direction of rotation. Support for piecemeal perception was obtained when observers’ organization of the cube was shown to be biased by whether they fixated on a subarea of the cube that had been disambiguated by the addition of depth cues or on an unbiased subarea. Subsequent research by Peterson and Gibson (1991) indicated that the critical factor was the focus of spatial attention rather than of fixation per se. Observers also were able to intentionally hold or maintain one perceptual organization of the cube, indicating that the representation of the cube was “malleable” in the authors’ terminology, although observers’ success depended upon whether they fixated the biased or unbiased part of the cube. Importantly, however, the limited nature of voluntary control was apparent in the finding that observers attempting to hold a particular perceptual organization still perceived the cube in the unintended organization at least some of the time, demonstrating what Peterson and Hochberg termed “non-elective instability.”
Although researchers have been aware of the limitation on voluntary control since at least the early work of Breese (1899), they have emphasized the degree to which observers can exert intentional control over the perception of reversible figures and have made little or no attempt to empirically determine the processes responsible for the limitation on observers’ control (e.g., Hochberg & Peterson, 1987; Liebert & Burk, 1985; Meng & Tong, 2004; Phillipson & Harris, 1984; Pitts et al., 2008; Struber & Stadler, 1999; van Ee et al., 2005). However, suggestive evidence about the cause of the limitation of voluntary control was obtained incidentally by Toppino (2003, Experiment 2).
Toppino’s (2003) second experiment was designed to test the focal-feature hypothesis of voluntary control which seems to follow from Hochberg and Peterson’s (1987; Peterson & Hochberg, 1983) work on piecemeal perception. 2 According to the hypothesis, voluntary control is exerted by selectively processing subareas of a figure that are biased toward one or the other possible perceptual alternative (see also García-Pérez, 1989, 1992). Toppino’s research capitalized on previous observations that there are two areas within the Necker cube that each bias a different one of its alternative interpretations (e.g., Kawabata, 1986). Focusing attention on the area near the upper right interior vertex or the lower left interior vertex (see Figure 1A) respectively bias observers to perceive the cube with its front face oriented down to the left or up to the right. Observers were instructed to hold one perceptual alternative (Hold condition) or to view the cube passively with no attempt to control their perception (No-Hold condition). Fixation location also was varied such that observers were instructed to fixate one or the other biased location in the cube. The perceptual alternative favored by the fixation location could be either compatible or incompatible with the alternative observers tried to maintain in the Hold condition. Toppino also varied cube size (Large vs. Small), with the small cube being so tiny that both biased areas were likely to fall within the focus of spatial attention.
Although the results demonstrated that observers could exert voluntary control over their perceptions of the Necker cube, the findings did not support the focal-feature explanation of this phenomenon. Viewing the small cube eliminated the effect of fixation which was apparent with the larger cube, presumably because the focal features were so close together in the small cube that they could not be processed selectively. However, the effect of intentionality was undiminished. This finding, among others, was inconsistent with the focal-feature hypothesis, suggesting that intentionality is mediated by other, most likely central, processes.
Consistent with previous studies (e.g., Hochberg & Peterson, 1987; Liebert & Burk, 1985; van Ee et al., 2005), Toppino (2003, Experiment 2) also clearly demonstrated the limited nature of voluntary control because observers in the Hold condition experienced a substantial number of reversals despite their efforts to prevent reversals and maintain the perception of a single cube orientation. A potential source of the limitation was suggested by the effect of cube size on reversal rates. When observers engaged in passive viewing in the No-Hold condition, small cubes reversed more often than large cubes, replicating a phenomenon that has long been known (e.g., Borsellino et al., 1982; Dugger & Courson, 1968; Washburn et al., 1931). Interestingly, cube size had the same effect in the Hold condition. That is, unintended reversals were experienced more frequently with a small than a large cube. This finding suggests that the intrusion of unintended reversals may be driven by bottom-up processes related to the stimulus variable of cube size.
From a hybrid theoretical perspective (Long & Toppino, 2004; Toppino, 2003), voluntarily maintaining one perception of an ambiguous figure is likely mediated by some sort of top-down attentional process such as self-priming the intended perceptual organization or, perhaps equivalently, trying to fit the intended representation to the ambiguous stimulus (e.g., Hochberg & Peterson, 1987). However, the findings that the rate of reversal is influenced by a stimulus variable such as cube size and that the variable exerted its effect even when observers were instructed to inhibit reversals may reflect the operation of relatively automatic, stimulus-driven processes.
Adapting an account of stimulus-variable effects proposed by Babich and Standing (1981), Toppino (2003) hypothesized that the effect of cube size on reversal rate is mediated by different rates of adaptation and recovery such that the neural structures representing small cubes adapt and recover more quickly than the structures underlying the perception of large cubes. Babich and Standing also had hypothesized that a greater degree of stimulation causes faster adaptation rates. Although research has not worked out the complexities of the underlying neurophysiology, a high-level overview of a possible account of the cube-size effect can be elucidated. That is, compared to a large cube, a small cube is more concentrated in the central area of the visual field where acuity is greater, leading to a more complete, detailed pattern of retinal stimulation which may contribute to greater activation of related cortical structures and, thus, to more rapid adaptation and a faster reversal rate. This hypothesis conforms with findings indicating that centrally presented Necker cubes reverse more rapidly than more peripherally presented cubes of the same size (Babich & Standing, 1981; Toppino, 2003, Experiment 1). The same mechanism also predicts that complete cubes, which are composed of continuous lines, will reverse faster than cubes of the same size that have been rendered incomplete by removing segments of their contours. The latter prediction, which was confirmed by Babich and Standing, is directly relevant to Experiment 2 of the present paper.
Although the hybrid hypothesis is interesting, plausible, and consistent with Toppino’s (2003) findings, it is important to note that he did not set out to investigate the limitation on voluntary control. He varied cube size in Experiment 2 as a way of manipulating spatial attention. The effect of cube size on reversals and its persistence in the face of hold instructions were incidental outcomes, not predicted from the hypotheses being assessed. As a result, it is important to determine that the obtained effects are replicable and not attributable to some extraneous aspect of the experiment. Experiment 1 was designed to achieve this goal. Experiment 2 expanded the research to a different stimulus variable (cube completeness) which was expected to affect reversal rate in the same manner as cube size.
Experiment 1
Toppino (2003) varied cube size and intentionality (hold instructions) in an experiment that was atypical in certain respects and made relatively complex demands on the observer. Unlike most reversible-figure studies, there was no central fixation point. Fixation was focused on non-central locations within the cube that biased the perception of one or the other cube orientation, and this bias was either compatible or incompatible with the bias induced by the hold instructions. Experiment 1 of the present paper replicated the essential features of Toppino’s methodology with a more straightforward set of conditions: Only cube size and hold instructions were varied, and the fixation point was always centrally located.
The dependent variables were the total or cumulative time in a viewing period during which observers experienced a particular orientation of the cube and the number of reversals (i.e., reversal cycles) they experienced during the viewing period. Observers were expected to maintain an interpretation of the cube longer in the Hold condition than in the No-Hold condition, replicating previous research (e.g., Hochberg & Peterson, 1987; Liebert & Burk, 1985; Toppino, 2003). This finding was necessary to establish voluntary control but, otherwise, the duration data were not a central focus of this research. The critical findings involved the reversal data. Observers who viewed the cubes passively in the No-Hold condition were expected to perceive more reversals with a small than a large cube, reflecting a more rapid cycle of neural adaptation and recovery with the smaller cubes. Observers in the Hold condition were expected to experience unintended reversals. And, like their counterparts in the No-Hold condition, the frequency of reversals was expected to be a function of cube size with unintended reversals occurring more often with small cubes. If cube size were to affect reversals comparably in both intentionality conditions, the results would imply that reversals in both intentionality conditions were caused, at least in part, by the same underlying mechanism.
Method
Observers and Design
Forty-eight undergraduate students served as observers, with 24 assigned randomly to each of two Hold-Instruction groups (Hold vs. No Hold) in a 2 × 2 mixed factorial design involving Cube Size (Large vs. Small) as the within-participants variable. According to an a priori power analysis, the sample size was sufficient to detect the cube-size effect (η2p = .267) reported by Toppino (2003, Experiment 2) with an alpha of .05 and a power of .95 in either experimental group.
Materials and Procedure
Observers were seated with heads restrained by a chinrest while viewing a computer monitor from a distance of 50.8 cm. Stimuli were black Necker cubes on a white background with a central fixation cross.
Necker cubes and their possible perceived orientations were explained, using a demonstration cube subtending a visual angle of 8.33°. Then, using the same cube and following the procedures of Toppino (2003) as closely as possible, observers were given four 60-s practice trials, which were separated from one another by 60-s rest intervals. In the first two trials, they familiarized themselves with the reversals of the Necker cube while practicing maintaining fixation and following the response requirement to press and hold a response key when a designated orientation of the cube was perceived and to release it when the alternate orientation was perceived. In the next two practice trials, the requirements of the Hold conditions were added so that observers in the Hold condition practiced maintaining a designated orientation of the cube during the 60-s viewing period. Half of the observers in each Hold-Instruction condition reported when they perceived the down-to-the-left (DL) orientation, whereas the remaining observers reported the up-to-the right (UR) orientation. They were instructed to maintain the orientation about which they were reporting as long as possible and to get it back as quickly as possible whenever the percept flipped to the other orientation. Observers in the No-Hold condition were to view the cube passively, without trying to influence it. The importance of maintaining fixation was emphasized throughout both practice and experimental trials.
Two-min after the last practice trials, observers received final instructions including instructions about the nature of the cubes they would see during the experimental trials. Then they began a series of 12 60-s experimental trials, separated by 120-s rest periods to allow recovery from neural adaptation that may have occurred during the preceding trial. Six trials involved a large (12°) cube and six involved a small (4°) cube. Cubes of each size appeared in a mixed order that was determined randomly for each observer. The sizes of the large and small cubes fell within the range used in previous investigations of the effect of cube size (e.g., Borsellino et al., 1982; Dugger & Courson, 1968; Toppino, 2003; Washburn et al., 1931) and were selected to ensure an effective manipulation of the variable.
Results
Data were analyzed separately for the Time during which observers reported perceiving the designated orientation of the Necker cube and for the Number of Reversal Cycles. The time data comprised each observer’s mean total (i.e., cumulative) time within each 60-s trial during which they depressed the response key, indicating perception of the cube orientation they were reporting (DL or UR). The data for the number of reversal cycles was each observer’s mean number of key presses per 60-s trial. Each key press represents a reversal cycle in which the observers’ percept flipped (key released) and flipped back again (key pressed anew). This measure approximates one half of the reversals observers reported experiencing by pressing and releasing the response key.
Both dependent variables were analyzed by a 2 × 2 (Hold Instructions X Cube Size) mixed ANOVA with repeated measures on the second factor.
Time
The time results are depicted in the top panel of Figure 2. The ANOVA revealed a strong main effect of Hold Instructions, F(1, 46) = 12.469, MSE = 177.076, p = .001, η2p = .213, indicating that observers in the Hold condition experienced the reported orientation longer than observers in the No-Hold condition. Cube size, however, did not generally affect the time the reported orientation was perceived, F(1, 46) < 1.00, MSE = 5.867. η2p <.001, although there was a reliable Hold-Instructions × Cube-Size interaction, F(1, 46) = 6.564, MSE = 5.867, p = .014, η2p = .125. Probing the interaction revealed that the mean total time the designated orientation of the cube was perceived was much larger for the Hold than for the No-Hold condition for both the large cube (MHold = 43.292 vs. MNo Hold = 32.433, t(46) = 3.912, p < .001, d = 1.129) and small cube (MHold = 42.017 vs. MNo Hold = 33.692, t(46) = 3.031, p = .002, d = .875). The interaction can be attributed to the large cube being perceived for slightly more time than the small cube in the Hold conditions, t(23) = 2.095, p = .047, d = .428, whereas there was no reliable difference as a function of cube size in the No-Hold conditions, t(23) = −1.614, p = .120, d = .329.

Total mean dominance time (Top) and mean number of reversal cycles (Bottom) in Experiment 1. Error bars represent standard error of the mean.
Reversal Cycles
The ANOVA indicated a strong main effect of cube size, F(1, 46) = 9.917, MSE = 1.422, p = .003, η2p = .177, with the small cube reversing more often than the large cube. (See the bottom panel of Figure 2.) There were numerically fewer reversals for the Hold than the No-Hold condition, but the effect fell short of significance, F(1, 46) = 3.839, MSE = 21.786, p = .056, η2p = .077. There was no significant interaction, F < 1.000.
Discussion
The present results confirm the incidental findings of Toppino (2003). Observers demonstrated voluntary control in that they reported perceiving the designated cube orientation longer in the Hold than in the No-Hold condition. That voluntary control is limited was demonstrated by the finding that observers in the hold condition could not suppress reversals entirely and perceived the unintended orientation of the cube for an appreciable portion of the viewing period. Most importantly, observers reported small cubes to reverse more often than large cubes, and this effect occurred, regardless of the hold instructions. The results support the hypothesis that a stimulus variable (in this case, cube size) affects the frequency of reversals through relatively automatic processes such as neural adaptation and recovery and that these processes exert a similar effect regardless of whether or not observers are instructed to prevent reversals.
The only difference between the present results and the results of Toppino (2003) was the cube-size × hold-instructions interaction in the time data which was obtained exclusively in the present experiment. However, the interaction does not seem to alter any of the central conclusions from this experiment. Beyond that, it may be advisable to view the interaction with some caution. It has no readily apparent explanation. There was no hint of the interaction in Toppino’s (2003) results. And, there was no hint of a comparable interaction in the results of the present Experiment 2, which paralleled the results of Experiment 1 in other ways.
Experiment 2
The results of Experiment 1 are consistent with the hybrid hypothesis that, although voluntary control of the perception of a reversible figure is mediated by top-down processes, it is limited by the bottom-up processes of neural adaptation and recovery which operate relatively automatically to produce unintended perceptual reversals. In Experiment 1, bottom-up processes were manipulated by means of the stimulus variable of cube size. Small cubes produce a faster perceptual fluctuation rate than large cubes, because the former are thought to lead to a more rapid cycle of adaptation and recovery.
The hybrid hypothesis was further tested in Experiment 2. Hold instructions (Hold vs. No-Hold) were again varied to manipulate intentionality. However, a different stimulus variable, cube completeness, was used to manipulate reversal rate and the underlying adaptation and recovery processes. The incomplete version of the stimulus was created by removing the middle portion of the lines representing the edges of the cube. (See Figure 1.) Previous research has shown that, with continuous viewing, complete cubes fluctuate faster than incomplete cubes (Babich & Standing, 1981; Cornwell, 1976). Babich and Standing (1981) hypothesized that a complete cube reverses more rapidly than an incomplete cube of the same size because the latter “involves a less intense stimulus” (p. 205), leading to a slower adaptation rate. As noted earlier, to the extent that more or less stimulation corresponds to more or less cortical activation, cube completeness plausibly may contribute to differences in the rate of neural adaptation.
According to the hybrid hypothesis, the results should be similar to those obtained in Experiment 1. Observers in the Hold condition should exhibit voluntary control by intentionally increasing the time during which they experience the designated perceptual organization, but a limitation on voluntary control should be evident from the fact that they cannot completely suppress the unintended percept. Critically, complete cubes were expected to reverse more rapidly than incomplete cubes regardless of observers’ intent.
Method
Observers were another sample of 48 undergraduate students. The design was again a 2 × 2 mixed factorial. Hold Instructions (Hold vs. No Hold) was the between-participants variable, but this time the within-participants variable was Cube Completeness (Complete vs. Incomplete). Materials and procedures were the same as in Experiment 1 with the sole exception of the nature of cubes presented during the 12 experimental trials. The cubes on all experimental trials subtended a visual angle of 6.4°. A random half of the trials involved a complete cube, while the other half involved an incomplete cube.
Results and Discussion
The dependent variables (Time and Reversal Cycles) were defined as in Experiment 1, and both were analyzed by a 2 × 2 (Hold Instructions × Cube Completeness) mixed ANOVA with repeated measures on the second factor.
Time
The time results appear in the top panel of Figure 3. The ANOVA revealed a strong main effect of Hold Instructions, F(1, 46) = 18.276, MSE = 196.813, p < .001, η2p = .284. As expected, observers in the Hold condition voluntarily extended the time they perceived the reported orientation compared to observers in the No-Hold condition, but they did not come close to completely eliminating the unintended orientation, highlighting the limited nature of their voluntary control. Neither Cube Completeness nor the Hold X Completeness interaction exerted a significant effect on the cumulative duration of percepts, F’s < 1.000.

Total mean dominance time (Top) and mean number of reversal cycles (Bottom) in Experiment 2. Error bars represent standard error of the mean.
Reversal Cycles
The reversal-cycle results appear in the lower panel of Figure 3. The ANOVA revealed a strong effect of Cube Completeness, F(1, 46) = 17.952, MSE = 2.233, p < .001, η2p = .281. The complete cube reversed more often than the incomplete cube. Although Hold Instructions led to numerically fewer reversals than No-Hold Instructions, the effect was small and non-significant as was the hold-instructions X cube-completeness interaction, both F’s < 1.000. Thus, changes in the number of reversals were caused by the stimulus variable, cube completeness, regardless of whether observers were or were not instructed to prevent perceptual fluctuations.
The present results were very similar to those obtained in the previous experiment with one possible exception. Hold instructions appeared to have had a larger effect on the number of reversal cycles in the first experiment involving cube size than in the present experiment involving cube completeness. There is no obvious explanation for this apparent discrepancy, but the difference may be less than it appears at first glance. The data in both experiments showed a higher reversal rate in the No-Hold than in the Hold condition. Although the effect of hold instructions was noticeably larger and approached significance in Experiment 1, the effect was not significant in either Experiment. To provide a clearer indication of whether the effect of Hold Instructions differed between experiments, the Experiment X Hold-Instructions interaction was assessed in an analysis of the reversal-cycle data that included Experiment (1 vs. 2) as a between-participants factor. The critical result was that the interaction did not approach significance, F (1, 92) = 1.572, MSE = 10.017, p = .213, η2p = .017. Thus, there currently is no clear evidence that the apparent difference in the effect of Hold Instructions in Experiments 1 and 2 is greater than would be expected by chance.
General Discussion
The results of the two experiments are strikingly similar, and in both cases conform to the predictions derived from a hybrid model of reversible-figure perception (e.g., Long and Toppino, 2004). Consistent with the contribution of top-down processes, observers exhibited robust voluntary control over figure reversals by extending the period of time that they perceived the reported orientation of the Necker cube relative to the time the same orientation was perceived by observers in the No Hold condition. However, a severe limit on voluntary control was manifest in that observers in the Hold condition continued to report many reversals in spite of being instructed to avoid them. Consistent with the contribution of bottom-up processes, the stimulus variables, Cube Size and Cube Completeness, strongly affected the number of reversals and, crucially, did so regardless of observers’ intentionality condition. This suggests that the processes activated by the stimulus variables are critical in producing reversals and that the same processes are involved in producing spontaneous reversals that occur with passive viewing in the No-Hold condition and unintended reversals that occur in the Hold condition.
A closer look at the dynamics of perceptual reversals in the present experiments is informative. Due to the occurrence of reversals, observers reported an oscillating sequence of percepts. Periods in which one orientation of the cube dominated alternated with periods in which the opposite orientation dominated. The duration of these individual alternating periods of dominance (called dominance durations or dwell times) was not recorded in the present experiments. However, as will be explained below, a clear picture of how dominance durations varied on the average can by gleaned by considering the effect of the stimulus variables (cube size and cube completeness) on the two dependent variables that actually were measured (total or cumulative time perceiving a particular orientation of the cube and the number of reversal cycles).
The cumulative time that observers perceived the designated orientation of the cube was virtually identical for small and large cubes in Experiment 1 and for complete and incomplete cubes in Experiment 2. This also was true of the time observers perceived the alternate, non-designated orientation of the cube. Simultaneously, however, rapidly reversing small cubes and complete cubes caused observers to experience more reversal cycles during the viewing period compared to slowly reversing large cubes and incomplete cubes. Together, these facts necessarily imply that the stimulus variables (cube size and cube completeness) had a similar effect on the dominance durations of both the designated orientation and the alternate orientation of the cube. That is, the dominance durations for both percepts had to be shorter on the average for small and complete cubes than for large and incomplete cubes. So, faster and slower reversal rates were associated, respectively, with shorter and longer dominance durations. However, when dominance durations were accumulated over the entire viewing period, this effect was obscured. That is, the sum of a larger number of shorter periods of perceptual dominance for the small and complete cubes was approximately equal to the sum of a smaller number of longer periods of dominance for the large and incomplete cubes.
Kornmeier et al. (2009) obtained compatible results using a discontinuous presentation procedure involving brief stimulus presentations separated by short interstimulus intervals (ISIs). With this procedure, stimulus variables that affect reversal rate include the presentation time of successive occurrences of the ambiguous stimulus and the length of the ISI. When these stimulus variables were crossed with intentionality instructions, Kornmeier et al. found that the stimulus variables and intentionality both significantly affected reversal rate, but intentionality did not alter the effect of the stimulus variables. Together, the present results and those of Kornmeier et al. may provide converging evidence for the conclusion that relatively automatic bottom-up processes account for the unintended reversals experienced by observers. However, it is difficult to interpret Kornmeier et al.’s findings in relation to the present results, despite their overall similarity. Unlike reversals obtained with continuous viewing, reversals obtained with discontinuous viewing occur in response to stimulus onsets and may reflect a fundamentally different process than that underlying reversals in continuous viewing (e.g., Klink et al., 2008a).
Voluntary control is widely believed to reflect a top-down, attention-like process by which the perceptual alternative that is dominant is selected and maintained. No peripheral factors such as eye movements and shifts of spatial attention have been found that might mediate the intentionality effects (e.g., Toppino, 2003; van Dam & van Ee, 2005), suggesting that they are attributable to central mechanisms. For example, intentional control may reflect top-down activation of one or the other of two high-level perceptual representations, tipping the balance of activation toward one representation or the other (e.g., Toppino, 2003). Whatever the exact mechanism, observers can both intentionally delay the occurrence of reversals and voluntarily induce reversals before they would ordinarily occur (e.g., Mathes et al., 2006; Meng & Tong, 2004; Seth & Reddy, 1979). This is evident in the present experiments from the fact that observers reported experiencing the non-designated orientation of the cube for less of the viewing period in the Hold conditions than in the No-Hold conditions. See Figures 2 and 3. 3
Bottom-up processes of neural adaptation and recovery are also believed to affect figure reversals. The neural structures underlying the different interpretations of a reversible figure are hypothesized to reciprocally inhibit one another (e.g., Blake et al., 2003; Toppino & Long, 2015). During the time that one interpretation of a reversible figure is perceived, the neural structures underlying that percept are thought to adapt until they are weakened sufficiently that a reversal occurs in which the competing percept assumes dominance. The newly dominant representation begins to adapt while the previously dominant representation recovers until another reversal takes place, and the cyclic process continues (e.g., Dornic, 1967). Modern adaptation theories also have other characteristics that do not bear directly on the present results. For example, to capture the variability that characterizes bistable perception, models often supplement the deterministic effects of adaptation with a stochastic contribution from neural noise (e.g., Chholak et al., 2020; Kang & Blake, 2010). Moreover, the system is thought to have a “cross-inhibitory” architecture in which each perceptual alternative suppresses its rival interpretation via inhibitory pathways that are at least partly independent of one another (e.g., Klink et al., 2008b; Toppino & Long, 2015).
As adaptation was applied in the present experiments, smaller and complete Necker cubes were assumed to result in more rapid adaptation than larger and incomplete cubes, presumably because the former cubes stimulate greater cortical activation. Consequently, smaller and complete cubes were predicted to yield faster reversal rates. In addition, stimulus-driven adaptation processes were hypothesized to be relatively automatic. Consequently, reversals were predicted to occur in the Hold conditions even though observers intended to prevent them, and the stimulus variables of cube size and completeness were predicted to affect reversal rate in the same manner regardless of observers’ intentions. These hypotheses received support from the fact that the results of Experiments 1 and 2 both confirmed the predictions. However, the mechanisms underlying the effects of stimulus size and completeness on reversible-figure perception have received surprisingly little attention from researchers, especially in recent years. There is a clear need for additional research focusing on the processes by which these variables exert their effects on the perception of Necker cubes and other reversible figures as well.
To summarize in closing, the results of the present experiments indicate that top-down and bottom-up processes exert separate effects on the perception of reversible figures. Observers can intentionally choose one percept or the other, but while they are maintaining that percept, they cannot prevent ongoing bottom-up processes. Thus, although they may be able to delay a reversal intentionally, they cannot do so indefinitely, because of the inexorable effect of bottom-up processes. It appears that bottom-up processes may be able to force a reversal on their own, thus limiting observers’ ability to voluntarily maintain a particular percept.
Acknowledgements
The author thanks Gerald Long and Pamela Blewitt for their helpful discussions related to this paper.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
