Abstract
In its current, popular manifestation, Virtual Reality (VR) represents the culmination of more than two centuries of screen practice aimed at creating greater immersion. VR’s optical illusions produce an expanded multisensory immersive experience that enhances the viewer’s interior position within new space. This article questions where embodiment and disembodiment lie in VR’s multisensory optical illusion and whether there is a difference produced by the digital environment versus the photographic, live-action environment? It takes into account our present moment in the history of VR during which the fantasy of total bodily engagement and transference into the “machine” has not yet occurred. In doing so, this article considers the way VR uses synesthetic modes rather than direct sensory stimuli to engage more of the senses.
Virtual Reality’s (VR) current, popular manifestation as a three-dimensional and 360-degree representational system uses head-mounted-displays (HMDs) to shutter its users from the outside world. It brings together more than two centuries of screen practice aimed at creating greater immersion. Building on the capacity of panoramas, dioramas, widescreen cinema, IMAX, and 3D cinema to fill peripheral vision (Acland 1998; Ross 2015; Belton 1992; Griffiths 2008; Huhtamo 2013; MacGowan 1957; Miller 1996) as well as the theme park attraction, Sensorama and 4D cinema’s ability to add additional sensory cues (Darley 2000; Gunning 2006; Ndalianis 2004), immersion is not situated in this practice as the loss of one’s self in story (although that might also occur; Green and Brock 2000). Rather it is “the sensation of entering a space that immediately identifies itself as somehow separate from the world and that eschews conventional modes of spectatorship in favor of a more bodily participation in the experience” (Griffiths 2008, 2). VR’s facility for erasing the frame to enhance the user’s interior position within new space has only previously been hinted at in 3D cinema and 4D theme park attractions. In the latter two, the comingling of the virtual diegesis among the user’s every day, lived world means both worlds remain visible. Within VR representational systems (as opposed to Augmented Reality and Mixed Reality; Billinghurst 2017), that dual visibility is lost when the user dons the HMD and is completely engulfed by visual and audio fields that appear to surround them.
This shift into a distinct visual space has led to utopian predictions that VR will foster new social, cultural, and educational user engagements including, for example, heightened empathy (Milk 2015), the ability to travel to otherwise unreachable destinations (Cheong 1995), and “a renewed respect for innate, intuitive, real-world intelligence over acquired, abstract, symbolic intelligence” (Krueger 1991, 205). There have equally been pessimistic considerations of this new representational space including critiques of VR’s supposed empathetic leanings (Bollmer 2017; Mitchell 2017), concerns with the loss of the human self (Batchen 1998), a fear of new panopticon regimes of surveillance (Hillis 1999), and the negative impact of “the serious contradiction between corporeal reality and artificial image illusion” (Grau 2003, 203). Rather than attempting to reconcile the complex ways in which these different possibilities play out, this article considers a particular angle that has only received limited attention: against the impossibility of transcending their corporeal body, how is the VR user provided with multisensory engagement derived primarily from the HMD’s optical illusion? This focus on the optical illusion by no means negates the importance of audio and haptic input provided by current VR systems but rather considers the way visual data drives synesthetic encounters that transfer feelings across the senses. For example, when hand controllers provide haptic feedback, simple vibrations can suggest the touch of a hard, three-dimensional object but only because of visual stimuli provided by the HMD. Even when “expanded,” VR works provide extra sensory input such as actual sand-covered floors and dusty backpacks in Iñárritu’s Carne y Arena: Virtually present, Physically invisible (2017), the content delivered by the HMD determines how the user feels present in the virtual world. Within this context, the optical illusion plays a significant role in creating an enhanced multisensory experience that builds upon the synesthetic qualities evident in much screen media (Marks 2002; Sobchack 2004) but with new implications for how we might “feel” our way through diegetic space now that distance from the screen is seemingly dissolved. At this moment, when VR is only just beginning to reach mainstream audiences, it remains to be seen whether VR’s corporeality will mean the medium is dismissed as a lesser mode of cultural engagement in the same way that spectatorship of “body genres” (horror, melodrama, and pornography) and 3D cinema have been (Sandifer 2011; L. Williams 1991) or whether VR will represent a paradigm shift for the way we engage with media content.
For any paradigm shift to take place, significant technological work needs to occur. During the first major VR wave at the end of the late twentieth century, Brooks (1999, 16) noted that the challenge, particularly in terms of computer graphic rendering of the virtual world, was to make it “look real, sound real, move and respond to interaction in real time, and even feel real.” Putting aside ontological issues of how we qualify a “realistic experience” and the problematic assumption of universally accepted “realness,” Brooks’s comments concur with much popular writing on the end goal for VR. The attraction of VR, its novel hook, is the ability for its optical illusion to convince us that we have entered an entirely new embodied world. Yet for all that Bazin’s (2009) dream of a Total Cinema is anticipated, we are still, at a technological level, at the very least, far from the state envisioned by popular media and far from leaving behind the knowledge of the illusion (Gunning 1995a) and/or doubled placement in which our body feels located in two places at once (Barker 2009). Instead, even though experiments with VR have been taking place for a number of decades, we appear to be at a nascent moment in the technology’s development akin to the “cinema of attractions” era that characterized early cinema (Gunning 2006). This article takes into account our present moment in the history of VR during which the fantasy of total bodily engagement and transference into the “machine” has not yet occurred, but in a manner similar to that of early cinema, there has been a significant shift in media production and content delivery with implications for how synesthetic transference occurs. By considering our current VR era in this way, this article is able to update the previous literature on VR that was published mainly before the release of publicly affordable HMDs in 2016.
To that end, six VR works are used as case studies. They provide a sample small enough for in-depth analysis while allowing a picture to emerge of the various factors that create synesthetic modes across variant VR works. In particular, the case studies are grouped together in sets that contain comparable thematic elements so that the similar and contrasting ways in which they use these elements can be unpacked: TheBlue: Encounter (2015) and Always Sunny in Philadelphia: Project Badass VR (2017); Richie’s Plank (2016) and Sightline: The Chair (2014); and Everest VR (2016) and The North Face: Nepal (2015). They were chosen for their commercial availability with all accessible to everyday consumers via the gaming platform Steam, the HMDs’ content platforms, and/or the streaming site Jaunt VR. Each case study operates as an entry-level product that does not need prior gaming skills and/or significant financial investment beyond that of the VR viewing technology. In many ways, they position a universal user with no reference to assumed gender, ethnic, or other identity signifiers. However, the need for certain types of mobility in some of them (TheBlue, Richie’s Plank, and Everest VR) implicitly expects an able-bodied user, while extensive use of English dialogue in others (Always Sunny in Philadelphia and The North Face) demands familiarity with this language. In this context, the article generalizes the user experience to develop arguments about the way our current VR operates, but it should be implicit that not every user will engage with these VR works in the same way. A final parameter is that all of the case studies rely on audiovisual data rather than haptic or other sensory feedback. (Although both Richie’s Plank and Everest VR require the use of hand controllers to engage with their content, haptic feedback is only minimally provided by these controllers.) The reason for limiting the case studies in this way is to provide fuller discussion of the way synesthetic transference from audio-visual data to embodied sensation can occur. As haptic feedback begins to be used in an increasingly sophisticated way, particularly as a larger number of HMDs with compatible hand controllers come to market, future analysis of VR works would do well to draw from the scholarship on hand controllers in games consoles (e.g., Crogan 2010; Kirkpatrick 2009).
For the time being, this article draws on its case studies to consider how VR works use synesthetic modes rather than direct sensory stimuli to engage more of the senses. It takes into account some of the specific factors that affect synesthetic transference in the new era of VR—tactile sensorium, doubled embodiment, and live-action versus computer-generated imagery (CGI). Thus, it is able to understand how the optical illusion configures expanded multisensory experiences in a way that sets VR apart from other screen media while demonstrating VR’s reliance on existent audiovisual traditions.
Considering the Embodiment of VR
The desire to produce advanced multisensory audiovisual experiences has a lengthy history, and often the vision for a corporeal VR outstrips technological possibilities. From Stanley G. Weinbaum’s 1935 novella Pygmalion’s Spectacles (2015)—which describes a pair of goggles that allows one to step inside the movies—through to visions of VR in Hollywood blockbusters—such as Total Recall (1990) and Star Trek episodes featuring the holodeck—popular culture has imagined a touchable sensorium that fully envelops the user/participant’s body so that “we are pushed through the medium to sensations that approach direct experience” (Biocca 2002, 102). Contrasting these fantasies, VR’s work with the limitations of contemporary technology has thus far meant a focus on stimulating or simulating separate sensory interactions that may or may not join up to provide a wider embodied experience. For example, the 1962 VR precursor Sensorama Simulator offered a visual approximation of a motorbike and car ride that was supplemented with sound, wind, smell, and a vibrating seat, yet the various devices imparting sensory stimulation were not of a level to convince the rider that they were actually on the vehicle (Guttentag 2010; Rheingold 1991). More recent developments include the VR Sense machine that promises to stimulate multiple senses via fragrance, touch, wind, thermal cooling, and mist (Humphries 2017); experiments at the National University of Singapore to add wind and temperature variations (Revell 2017); the Vaqso clip-on device that can be added to commercial HMDs to provide a variety of smells (Feltham 2017); and the Teslasuit that covers most of the body to simulate touch or pressure (Javelosa 2016). In each case, different approaches stimulate different senses, but there is not the possibility of working with the entire somatic schema to approximate the sensory input we experience outside of VR. Indeed, contemporary systems continue to rely on the synesthetic potential of optical illusions to create corporeal engagement, and the majority of commercially available HMDs such as Oculus, HTC Vive, Playstation, and Samsung Gear do so via visual, audio, and limited haptic data.
Even with these limitations, VR’s corporeal possibilities have long interested scholars, particularly those who engaged with emerging conceptualizations of cyberspace and what it might mean for postmodern subjectivities (e.g., Batchen 1998; Bukatman 1993; Grau 2003; Heim 1998; Hillis 1999). Often writing with regard to exploratory VR systems, they frequently looked forward to what VR might entail when it becomes more ubiquitous, suggesting it might generate problems of perception (Grau 2003, 203); Alternate World Syndrome, a sickness of kinesthetic disconnect (Heim 1998, 52); and the fragmentation of identity (Hillis 1999, 164). Later accounts moved beyond a postmodernism framework, but their publication prior to the release of publicly available HMDs meant they were equally concerned with exploring what VR might offer based on limited access to examples of VR works (Bolter and Grusin 1999; Elsaesser 2014). They acknowledged that the VR they described did not have the technical capacity to fulfil the fantasy of an alternate world that meets the demands of advanced telepresence: “the extent to which one feels present in the mediated environment, rather than in the immediate physical environment” (Steuer 1992, 76). Rather, the VR systems they witnessed contained “many ruptures: slow frame rates, jagged graphics, bright colours, bland lighting and system crashes” (Bolter and Grusin 1999, 22).
In this context, articles by Nash (2018) and Popat (2016) are useful as both are able to refer to more recent, and more technically sophisticated, VR works. In the former’s analysis of 360-degree documentaries and the latter’s analysis of a flying hot air balloon ride, they take forward the textual analysis begun by Grau (2003), Rheingold (1991), and Heim (1998) to analyze VR works that provide users with a greater sense of immersion via more photorealist visual fields. Equally significant, Nash and Popat provide extensive and detailed first-hand accounts of these VR works to engage with the specific ways in which users are transported into visual worlds yet retain an interplay based on their preexisting corporeal subjectivity. This article carries on their work with deeper analysis of the way commercially available VR provides synesthetic modes. Furthermore, it expands analysis of individual VR works in the current era beyond consideration of the 360-degrees documentaries that have thus far garnered most scholarly attention (Bollmer 2017; Herson 2016; Kool 2016; Kubota 2016; Mitchell 2017; Nash 2018).
Synesthetic Encounters
The synesthetic potential of screen media, particularly as it has been addressed in film phenomenology, provides the ground for understanding a number of rich spectatorial interactions that cannot be explained within a narrative framework (Marks 2000). This phenomenological approach has already been applied to other screen media, but there is yet opportunity to think about how it operates in the 360-degree viewing space provided by the new generation of VR works. As noted by Sobchack (2004, 60), “even at the movies our vision and hearing are informed and given meaning by our other modes of sensory access to the world: our capacity not only to see and to hear but also to touch, to smell, to taste, and always to proprioceptively feel our weight, dimension, gravity and movement in the world.” While clinical definitions of synesthesia often point toward extreme permutations in which a person may perceive a sound as a color or a word might have a particular taste, Sobchack takes into account the way cinema’s visual and audio data can draw on our sensory knowledge to stimulate touch, smell, and taste in ways that may be regarded as metaphorical but present themselves to the viewer more directly. In VR, similar processes to cinema take place so that, for example, a snowy environment can give the user a sudden chill, a flower situated in a meadow may bring forth the smell of a summer day, while a rotting food source can result in an acrid taste on the tongue. Although it is dangerous to over-simplify the differences between VR and cinema considering both work with a range of aesthetic, narrative, and technological tendencies, VR’s 360-degree visual field has a distinct ability to shut off the outside world, which, in turn, establishes the potential for synesthetic interactions to take place with greater proximity and in a more enveloping manner. Similarly, the ability to use head-tracked spatial audio whereby surround sound is calibrated in relation to where the user is situated in the virtual world means that the user is continuously situated at the center of the sensorial experience. This spectatorial position—markedly interior to the audio-visual environment—collapses the distance often found in commercial, mainstream media and described by Marks (2000) as optical visuality. Whereas the distance found in optical visuality encourages a type of mastery with the potential to retreat to a safe space provided by that separation, the inner, among-the-visual-field, of VR provides no such escape. In this context, it is not surprising that many recent VR works have drawn upon the horror genre to exploit the potential thrills to be gained when the user is aware they cannot escape (Daniel 2017; Staubli 2017; V. Williams 2018). At the same time, the HMD’s manner of shutting off the outside world also provides the potential for a number of, less frightening, multisensory experiences that occur with limited distraction.
One of the synesthetic transferences that has received significant attention during cinema’s history is the ability for viewers to feel as if they might touch objects placed in front of them, most notably articulated in discussions of 3D cinema (Ross 2013; Johnston 2012; Paul 2004). While the haptic feedback stimulated via VR hand controllers can produce physical sensations of touch that go far beyond the “chimera” of 3D cinema (Paul 2004, 230), even those VR works that do not implement this function are able to work in tactile ways. Two particular examples, TheBlue: Encounter (2015) and Always Sunny in Philadelphia: Project Badass VR (2017) use underwater scenes that point to the possibilities of producing a tactile sensorium that plays on our proprioceptive positioning while working solely with audiovisual data. TheBlue presents a digitally created environment that allows users six-degrees-of-freedom movement whereby they can both look around in all directions and physically move around the objects in its sunken ship environment—an aspect that many consider necessary for “true” VR (Loomis 2016). Always Sunny in Philadelphia, on the other hand, offers a live-action-filmed fixed viewpoint in which users may look around 360-degree space, but they cannot navigate it. Nonetheless, both use their optical illusion to position the user as if they were among the depths of the ocean. TheBlue does so from the outset, allowing the user significant time to take in shoals of fish that dart toward them and, in turn, encourages a dramatic “looming response” (Bottomore 1999) in which users physically recoil from the marine creatures that seem to be coming toward them. Always Sunny in Philadelphia, on the other hand, submerges the user toward the end of its narrative after central character Mac has purposefully driven off the end of a pier on his motorbike. In both cases, the underwater scenes emphasize the materiality of water “in which normative and familiar rules of gravity and bodily movement do not apply,” where “our sensory relationship with the world changes dramatically, including our sense of vision, hearing, touch, smell, and taste, but also our kinesthetic and muscular senses as well as our proprioceptive sense” (Lindner 2012, n.p.). In TheBlue and Always Sunny in Philadelphia, this manifests in the unusual but not entirely unfamiliar sensation of being suspended in water, particularly because the small air bubbles floating through the water give the impression of being in a thick, aqueous space. The way in which the “thick” space surrounds us in a manner not possible in other media affects our sense of proprioceptive engagement: the type of movement that we seem to be able to make through this space as well as how we might feel the materiality of the water. When a large blue whale swims toward the deck and floats close to the user in TheBlue, three-dimensional contouring of its ridged surface combine with the thick space that seems to exist between user and whale to intensify tactile materiality even though waving a hand through the whale will make the illusion apparent.
Although the notable difference between these two VR works is the greater navigability afforded by TheBlue, another significant variance is the way in which different narrative cues are used to situate the user’s presence underwater from a point of view (POV) position. The immovable POV position that the user holds in Always Sunny in Philadelphia is narratively justified when Mac’s friends, Dee and Dennis, discuss the user as a character: a friend of Mac who is strapped to the motorbike and forced to participate in his stunt. During this discussion, Dee and Dennis talk to the camera, providing two of the major forms of direct address that Brown (2012) sees operating in cinema: intimacy and instantiation, the present-ness and immediacy of direct address, which often derive from a present tense-ness. Working in similar modes to the VR documentaries that Nash (2018) discusses, intimacy and instantiation combine to help position an embodied suspension of disbelief so that, rather than operating as the omnipresent viewer common to cinema who can witness all action but from a distance, the VR user is positioned as a character within the diegesis: allowed to feel events as they occur. In this way, narrative helps assert the process of immersion in the environment so that synesthetic interaction is less likely to be diminished by the user questioning the rationale for their interior, immovable placement. While TheBlue does not have the same narrative framework for positioning the user—and the POV position is not fixed in the same way—there is a moment when the whale’s large blinking eye gives the user the briefest sense of direct address before it swims off toward the surface. In this way, there is acknowledgement of the user within the whale’s material space so that their embodiment among the sensual underwater field is relational to their inclusion among diegetic elements. Surpassing the limited use of POV in cinema and television, POV is included as a dominant spectatorial position in these works and in much VR (Hillis 1999). It brings VR more in line with first-person gaming modes and their corresponding corporeal intensity (Grodal 2003) than the incorporeal modes that were established and supported by the camera obscura before going on to dominate monoscopic photographic media (Crary 1992).
Doubled Embodiment
Yet for all that this embodiment—particularly its present tense-ness—is significant in opening up the possibility of deep synesthetic interaction, there is still the doubled placement of being both in the virtual world as well as physically present in the external world. In both TheBlue and Always Sunny in Philadelphia, there is the possibility of momentary absorption—a forgetting of the world external to the VR space—but there is always an imbalance between the tactile environment created in the optical illusion and the physical surfaces exterior to it: “our bodies are both present and absent, experiencing agency and aspects of sensation even though there is no direct contact between flesh and world” (Popat 2016, 359). In the underwater scenes, there is a lack of correlation between the thick liquid materiality circulating below the user and the hard surface of the floor on which their feet actually rest. This is particularly true in TheBlue when the user can, like a ghost, walk through the wooden barriers at the edge of the deck to stand suspended in open water. There is thus a tension between the user investing fully in the thickness of the virtual space by moving limbs and muscles through that space in tandem with how the virtual world suggests they should move (at reduced velocity, imagining the pressure of the water on their skin), and the ability for their body to disregard the synesthetic qualities of the environment and move through it at their own pace and without limitation. Yet it is not as if the body has to operate within one or the other of these states, but rather, there is the opportunity to move along a scale from full absorption in the virtual world to reticent interest in what is occurring within it. More often than not, the liminal sensation of being in between these physical states is a common place to find oneself.
A VR work that has become popular for highlighting this liminal state is Richie’s Plank (2016). When entering the VR world, users are placed at street level in a dense urban environment with numerous skyscrapers to either side including one which has an elevator the user can walk into. Like TheBlue, Richie’s Plank has six-degrees-of-freedom navigability, and once the user has physically stepped inside the elevator, they can use a hand controller to press a button to ascend to the top floor. Spatial audio helps shift the user’s focus from the exterior urban environment (with car, bird, and other city noises) to the interior of the elevator (with “elevator” music). The combination of the two sound scapes during the elevator’s lengthy elevation—along with the slight glimpse of the urban environment through a crack in the elevator doors—provides a strong sensation of upward motion even though the “real world” floor beneath the user’s feet does not move. On the top floor is Richie’s Plank’s key attraction: a wooden plank that stretches out from the elevator with a 160m drop beneath it. While the perceived drop is only an optical illusion, and while a step to the side of the plank would allow the user to feel the hard floor of the space in which they have donned their HMD, Richie’s Plank’s popularity lies in the convincing way in which it instils a sense of height and vertigo in its users. Various online videos (e.g., Rick and morty co-creator terrified in VR 2016) document the range of responses users have to the plank such as those who are unable to conceive of leaving the “safety” of the elevator space, those who use their hands to help feel their way along the floor as they edge toward the end of the plank, and those who confidently walk to the edge of the plank and step off. For those that do step off the edge of the plank, the visual field shifts upward and provides the user with the sensation of moving rapidly downward until they “hit” the ground and the visual field cuts to white. In this instance, many users demonstrate a feeling of visceral sensation akin to that experienced when undergoing a drop on a roller coaster ride. Because those who are able to step off the plank are the most likely to be engaged in perceiving the virtuality of the experience (maintaining the embodied knowledge that they are stepping onto another part of the floor rather than on to open space), the fact that they directly feel the physical sensation of the drop demonstrates the complex, multifaceted way in which doubled embodiment can take place. Although they are fully aware that they are participating in a virtual experience, the visual dominance experienced in VR has the ability to produce involuntary sensory responses that can resonate throughout the body: in the gut, rapid heartbeat drops and acceleration, shortening of breath, and/or muscular tensing.
It is in this context that those who cannot step out from the elevator (and in some extreme circumstances immediately take the HMD off because the illusion affects too deeply their fear of heights) should not be understood as some kind of dupe, akin to the visualization of the rube in early cinema: the “simpleton spectator” who did not understand that cinema was only a representation (Elsaesser 2006). These users know that they can take the HMD off at any moment to break the illusion so any continued investment in the experience is akin to the film spectatorship that Allen (1995, 4) describes: while we know that what we are seeing is only a film, we nevertheless experience that film as a fully realized world. I call this form “projective illusion.” The experience of projective illusion is not one that is imposed upon a passive spectator but an experience into which an active spectator voluntarily enters.
One of the ways that some users proactively negotiate the projective/optical illusion in Richie’s Plank is to request that someone holds their hand as they walk along the plank and, particularly, when they want to step off the end. Recognizing that the floor beneath their feet is somewhat equivalent to the virtual plank on which they are walking, the touch of another person’s hand—no matter that it cannot be seen—offers a greater connection to the physical world outside of the HMD. The reversibility of touching that unseen hand at the same time as being touched by it provides a type of anchor that is normally unnecessary in framed screen media where a glance down, or to the side, shows one’s body outside of diegetic space. This type of embodied negotiation, thus, augments the processes Barker notes in cinema spectatorship whereby “we extend our bodies to the film, and it extends its body to us simultaneously, and in doing so, we agree on certain terms. We commit ourselves to the film’s world without ever abandoning our own world, for the limits of our bodies are never forgotten or confused in the handshake” (Barker 2009, 94). The commitment in VR is often more fully embodied than that which can occur in framed media, and while the limits of the body are equally not forgotten, the limits of what they can undergo (vertigo, for example) are more forcefully tested.
Significant to this commitment is the extent to which we feel our embodiment and hear our embodiment (the sound of our breath, gasps, shrieks, and/or laughs) yet we are not able to see it. When a VR work such as Richie’s Plank is played in current commercial systems, the HMD sufficiently blocks out other visual fields so that no part of the user’s body can be seen. In the diegetic world, the user’s body is also absent, increasing the previously described ghost-like sensation of moving through space without seeing the body (or apparatus) that directs this movement. In Richie’s Plank, this simply means that looking down at the plank, the space above it seems empty. In a VR work such as Always Sunny in Philadelphia, on the other hand, there is an obvious gap on the motorbike seat behind Mac where his friend/the user should be. Some VR works have attempted to overcome this by providing a type of visual body that can substitute for the absent user’s body. While the aim of much VR industry research is to eventually provide a digital avatar whose movements will correspond to those of the user (Bukatman 1993; Hillis 1999; Walser 1991), technological limitations mean that this is still an experimental concept, and in most cases, floating hands are the only indicator of a body that corresponds with the user’s movement. Aiming to provide an intermediary solution, Sightline: The Chair (2014) positions a (headless) relatively immobile, human body under the eyeline of the user at around neck height. Sightline uses a type of head tracking technology to change aspects of the visual field from abstract objects to forest and then cityscapes while the user is looking in the other direction. To simplify this process, the user is instructed to sit down during its run time. This means that the height of the headless body in the diegetic space roughly corresponds to that of the user and the two face the same direction. However, the immobility of the virtual body is at odds with the mobility of the user’s body. In this way, it draws attention to the lack of correlation between user and avatar at the same time as asking the user to imaginatively project into that virtual body. Whereas a missing body can allow users to engage a certain type of suspension of disbelief in regard to the visual positioning of their body in space, the visualized virtual body somewhat disrupts the optical illusion. It draws attention to the contingency of the illusion. Yet rather than being exemplary of a primitive—not yet Total Cinema, state—the virtual body in Sightline interacts with a long history of dealing with the “trick” in audiovisual media in which there are pleasures in viewing the operation of the illusion (Gunning 1995b).
CGI Versus Live-Action Visual Fields
Although the immobility of the virtual body in Sightline is probably the most significant aspect disrupting the optical illusion, the digitally created (as opposed to live-action) visual field also plays a role. While somewhat photorealist, the body in Sightline is nonetheless clearly a digital creation with relatively sharp edges, basic contouring, and limited textural surfaces. This is not uncommon in interactive 360-degree environments where the processing power needed to visualize three-dimensional digital objects means that it is a challenge to present visual fields of the same complexity and detail that viewers are accustomed to in Hollywood CGI productions. Attempts to overcome this can generally be situated in two directions: increasing the technological capacity to create CGI environments that are substantially photorealist, and using photographic technology such as light field capture systems that can provide photographic images of a three-dimensional, 360-degree, navigable environment (Matney 2017). However, the teleological aims of these goals can be at the expense of drawing on productive understandings of the relationship between photography, digital environments, and cinema that have been developed over a number of decades. As Allen (1995, 89) notes when discussing Bazin’s preference for photography as a more realistic medium, “since realism is a matter of artistic convention, the claim that photography is more realistic by virtue of its basis in mechanical reproduction cannot be defended.” In this context, digital environments that are not highly photorealist have captured significant viewer investment in their ontological “realness,” as a space for embodied engagement, across a range of media, including early VR manifestations. Writing in 1999 about his opportunity to participate in flying a Boeing 747 simulation—a digital environment that would now be considered primitive in comparison to recent VR works—Brooks (1999, 20) noted that “so compelling was the illusion that the breaks in presence came as visceral, not intellectual, shocks.” Conversely, although interactivity and navigability in VR are seen as desirable ends, the festival and critical success of a number of photography-based, fixed-viewpoint, 360-degree VR films (Collisions [2016], Clouds over Sidra [2015], The Protectors: Walk in the Rangers’ Shoes [2017]) suggests that this particular form provides spectatorial pleasures that do not necessarily need further technological augmentation. Thinking through this point in terms of the future of VR, Moody (2017, 53) suggests that “the more restrictive pleasures of the 360° film, which are experienced much more passively than many other forms of immersive storytelling, are not barriers to immersion, and that in fact, many of its restrictions are likely to become conventions for VR experiences in general in the future.”
Two VR works, both set in the mountains of Nepal—Everest VR (2016) and The North Face: Nepal (2015)—demonstrate the way in which different uses of digitally created (navigable) and photographically captured (highly realist) environments produce distinct spectatorial and synesthetic pleasures. Combining three hundred thousand high-resolution images of the mountain range with digitally generated 3D mesh and textures, Everest VR is described by its production company Sólfar (2015, n.p.) as “the definitive experience of what it feels like to summit Mount Everest.” It provides spectacular vistas of the mountains at the same time as it produces photorealistic snow and rock beneath the user as they navigate the different scenes that lead them from base camp to summit. Beyond navigation of the mountainous environment, interactivity is available in specific moments when, for example, the user is given the opportunity to climb Lhotse Face. In this instance, the hand controllers simulate ice-axes and, as the user swings them into the virtual ice face in front of them, they gain leverage to ascend the mountain. Although there is no haptic feedback from the controllers, the visual impact of the ice-axe slamming into the snow-covered ice—often dislodging pieces of packed snow—produces the sensation that the user has a physical impact on the environment.
This is distinct from The North Face, a 360-degree film, which, similar to Always Sunny in Philadelphia, offers the user a fixed viewpoint as they are taken through a number of live-action scenes in the Nepali mountains before arriving at the Lobuche peak. What both achieve in their spectacular vistas is what Darley (2000, 17) describes as the aim of much computer image research “the proximate or accurate image: the ‘realisticness’ or resemblance of an image to the phenomenal everyday world that we perceive and experience (partially) through sight.” This is particularly achievable in the landscape vistas as the distance between user and far off mountains means that other sensory input is unnecessary to simulate its resemblance to the phenomenological object with which we believe we are familiar. Yet, when Everest VR allows users the opportunity to “place” their hands on the rungs of a ladder or swing an ice-axe into the snow, the visual feedback provided by the controllers is only a rough approximation rather than a strong resemblance of the phenomenal reality we expect from the cold steel rungs of a ladder high in the mountains or ice crushing beneath an axe. The disconnect between the embodiment of the user’s hand on the controllers compared to the visual scene complicates the potential for presence within the virtual environment. At the same time, it does not mean that no presence is provided, but—similar to viewing the virtual body in Sightline—it is a contingent presence, supported by the other sensory input from visual and sonic fields. We should not assume, as Sólfar (2015, n.p.) does, that “you will leave the experience feeling like you were there,” for there is no singular, universal, user who will undergo Everest VR in the same way. However, the success of Everest VR (the most downloaded VR work when HTC Vive began its subscription service) suggests that there are substantial pleasures in feeling one’s way around an environment that is almost, but not quite, an approximation of some of our world’s most spectacular landscapes.
What then of the 360-degree film where these pleasurable tactile sensations are suggested to a greater degree by higher resolution photographic imagery but the user has less control to navigate and interact with the environment? The North Face relies on a much less physically active user, and it is various cinematic techniques (camera motion/editing) that move the user from place to place rather than six-degrees-of-freedom navigability. Nonetheless, the impact of these movements should not be underestimated. First, the visceral sensations that take place as the fixed viewpoint is moved through the street of a Nepali village recall the intensity of the POV shot moving through space in cinema but with the additional visual plenitude provided by an expanded visual field. Second, there is a perceptual richness to live-action high resolution photographic objects in motion toward the user when, for example, birds flock toward the camera and create dynamic converging motion vectors. They provide a point of contact between user and diegesis that is more difficult to achieve with the lower fidelity objects provided in most interactive, digitally created, VR environments. As Gunning (2010, 261) notes with regard to cinema, “we do not just see motion and we are not simply affected emotionally by its role within a plot; we feel it in our guts or throughout our bodies.” In the 360-degree film, this aspect is equally in play and intensified by the way that high-fidelity, photographically captured objects in motion are relational to a wider field of view in which the user’s own ability to look around can add further vectors and trajectories.
Another aspect aiding the sensory connection the user has to live-action 360-degree environments is the intimate connections that can occur when put into close proximity with other subjects. As noted previously, the direct address mode fosters this intimacy, but scholars have also noted the intimacy produced by the tactile appeal of the close-up in cinema (Balász 1999; Epstein 1977). Although the lack of frame in VR does not allow the exact same technique, positioning the camera near to subjects allows for similar engagement with their face, “producing an intense phenomenological experience of presence, and yet, simultaneously, that deeply experienced entity becomes a sign, a text, a surface that demands to be read” (Doane 2003, 94). In The North Face, this aspect is provided by climber Renan Ozturk who introduces us to the region, often sitting close to the camera so we can see the way he contemplates his surroundings. The perceptual intricacies of the small gestures on his face combine with the textures of his beard, hair, and mountaineering clothes in relation to the elaborate detail provided in the rest of the photographically captured 360-degree visual field. We later move to a scene on the mountain peak where Ozturk and another climber make their way to the summit. Although we no longer have the same proximity to their bodies, the previous intimacy remains and fosters a connection so that, in their arduous ascent, there is the capacity to understand our potential, if not realized, motion through this space with them.
This is quite different from Everest VR where climbers also exist in the visual field but their digital constitution means they have not been rendered with the same degree of photorealism. Their outlines are less intricate, and, significantly, their textural surfaces are more basic. The lack of solidity to the basic textural surfaces on the climbers in Everest VR is further emphasized if the user walks toward and then through the climber. The frame composition of the climber’s surface can be seen, but the inside appears as a hollow structure. Furthermore, their motion in space lacks the complexity of live-action human subjects, bringing them close to the trough of the uncanny valley whereby subjects have left behind the abstraction of animation or other clearly nonhuman forms and are closer to undead subjects such as zombies (Mori 1970). With all that said, however, there are similar pleasures to those provided by the other interactive elements in Everest VR. The ability to move close to, and then around these digital subjects in 360-degree space offers a type of embodied engagement with the space and its inhabitants that is not found in other media. While the bodies in Everest VR may seem perceptually “fake” compared to the photographically shot bodies in The North Face, different, yet no less significant, embodied engagement and sensory connections play out in relation to their different visual fields.
Conclusion
While the case studies I have drawn on were chosen for the ways in which they foreground the issues of synesthesia and embodiment in twenty-first century VR, they are by no means unique. Underwater scenes are common in VR and follow a lineage in 3D cinema where aqueous environments are often introduced that add relatively little to narrative development but produce tactile, seemingly enveloping, visual fields (Ross 2015). It is also common for other VR works to experiment with and push the limits of how they can visually, sonically, and haptically play with our (dis) embodied presence within the virtual world. The various ways in which the case studies produce their visual worlds is typical: from entirely computer-generated data in TheBlue, Richie’s Plank, and Sightline to the use of live-action footage in Always Sunny in Philadelphia and The North Face, as well as the mix of photographed images and CGI elements in Everest VR. In each case, they represent a historical moment during which form and content are malleable to the needs of individual VR works, yet ongoing technological limitations mean the optical illusion is paramount in the creation of synesthetic interactions.
In the popular press, and via accounts originating from promoters of VR systems, there is the sense that VR is on the cusp of entering a consolidation stage during which multisensory realism will be increased via technological updates. There is the strong expectation that current idiosyncratic approaches to content production will streamline into recognizable narrative forms. This forward-looking approach is nothing new. As Bates noted in 1992 “while VR appears to be a first-person, “realistic,” form, we believe that a language of presentation will develop over time” (p. 138). The ongoing expectation that we will soon surpass VR’s current state has led various critics (King 2000; McMahan 2006; Staubli 2017) to elaborate on the parallel between VR and the early film’s “cinema of attractions” phase (Gunning 2006). Whether VR will develop a formal system akin to the Classical Hollywood Cinema that supplanted early cinema, or a more multifarious combination of media styles and innovation similar to the modes employed in commercial video games, has yet to play out. Nonetheless, the comparison with the “cinema of attractions” is useful as this framework rescues early cinema from descriptions as a primitive audiovisual form and instead celebrates the way the attraction posits interactive spectatorial relationships that provide pleasurable engagement. In this context, the diverse ways in which the case studies discussed in this article play with synesthetic encounters to embody and disembody users with differing levels of presence becomes part of the attraction of current VR rather than a shortfall. The continuation of the “cinema of attractions” modes in later media (Strauven 2006)—particularly the acknowledgement of the user and attempts to place them within rather than at a distance from the diegesis—speaks to their enduring quality and provides a platform for VR to make use of these modes. Thus, aspects of the “cinema of attractions” that contribute to the synesthetic pleasures described in this article such as direct address and playful engagement with (dis)embodiment have the capacity to remain key components of VR as it develops further.
Footnotes
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by Victoria University of Wellington funding.
Author Biography
.
