Abstract
Picturebooks featuring same-sex parents, although growing in number, remain underexplored. In this article, the authors look at the covers of four such picturebooks, in particular at the representation of the co-parents and the multimodal workings of image and text. They ask: ‘How can the multimodal relationship between image and written text (the title) on the covers of picturebooks featuring same-sex parents best be described and explained?’ This study is timely in that the image–text relationship is a contested one. Drawing on the notions of modal affordance and epistemological commitment and the Hallidayan functional grammar category of enhancement, the authors use Theo van Leeuwen’s (2008, 1996) Social Actors frameworks, in particular the Visual Representation frameworks, to show that image and text (the title) are not commensurate in the meanings they communicate. Further, rather than one mode being merely supportive of the other, image and text, here, are ‘mutually enhancing’ (see Unsworth and Cléirigh’s contribution to The Routledge Handbook of Multimodal Analysis, edited by Carey Jewitt, 2011). In these picturebooks, gay identities and practices can be – and indeed need to be read through an appreciation of this mutual enhancement, rather than through image or text (title) alone or in parallel. The authors propose that mutual enhancement may be characteristic of a sometime transgressive genre such as picturebooks featuring gay parents.
Keywords
Introduction
Picturebooks are prototypically multimodal. Hassett and Curwood (2009: 272) argue convincingly that each element of a picturebook ‘is a mode of sorts’: further, even those picturebooks with no words may be multimodally associated with parent–child reading practices involving action, for example gesture and touch (the ‘haptic mode’), such as the child or ‘reader-aloud’ pointing at the text. Picturebooks however usually comprise at least the two semiotic modes of image and written language. These have different potentials for meaning-making, the creative synergies and tensions between which have been widely noted in the literature on picturebooks (e.g. Nikolajeva and Scott, 2000; Nodelman, 1988).
The titles of certain monographs on picturebooks reveal the ways that theorists attempt to conceptualise the relationship between written and visual modes: the subtitle of Lewis’s (2001) Reading Contemporary Picturebooks is ‘Picturing Text’, Nodelman’s (1988) book is called Words about Pictures, and Arizpe and Styles’s (2002) book title is Children Reading Pictures: Interpreting Visual Texts. Lewis characterises the relationship as a process in which the written story becomes illustrated; Nodelman implies the words are there to describe the pictures. Arizpe and Styles, however, hint at a need for recursive reading strategies in the case of multimodal texts such as picturebooks. Studies of multimodality in picturebooks accordingly extend beyond the text to the pedagogic value of multimodal texts, and teaching children to read multimodally (Agosto, 1999; Buckland and Croker, 2005/2006; Duncum, 2004; Hassett and Curwood, 2009; Sipe, 2008; see also Unsworth and Cléirigh, 2011, on pedagogically helpful design of labelled diagrams in schoolbooks).
Modes, which are almost certainly multiple for any meaning-making text or event, can be seen both as being and comprising resources for meaning. Modes include action (e.g. gesture, posture), voice (e.g. talk, singing, whispering), other forms of sound (e.g. music, noise from machines), gaze and facial expression (and other non-verbal communication) and texture, as well as image (e.g. colour, shape) and writing (see also Kress, 2010).
These examples remind us that each mode has an associated rich set of ‘modal resources’. In regard to picturebooks, modal resources of writing include ‘syntactic, grammatical and lexical resources, graphic resources such as font type, size, and resources for “framing”, such as punctuation’ (Bezemer and Kress, 2008: 171). On picturebook covers, the data for this study, writing (in the book titles) includes the modal resources of colour, case, graphic variation (e.g. of font and font size) and shape. The sequencing of these resources may create a linear and directional ‘reading path’ in that, in English, competent readers read from left to right, top to bottom (Bezemer and Kress, 2008; Kress, 2011).
The modal resources of image include ‘the position of elements in a framed space, size, color, shape, icons of various kinds – lines, circles – as well as resources such as spatial relation’ (Bezemer and Kress, 2008: 171), and, for picturebook covers, we can add focus (sharp or blurred), background detail, appearance (face, hair, clothing, accessories, posture), action of characters, and indeed representational style: line drawings, photographs, and ‘stylised’ as well as ‘realistic’ images. 1 The interest in images thus lies in who/what is included (and accordingly who not), and how, but also where (foreground/middle-ground/background, and centre/periphery). Such spatial relationships extend to who/what is isolated, and who/what connected, for example through touch, proximity or gaze: who is looking at whom? ‘Vectors’ with semiotic potential can then be identified, i.e. imaginary lines created by direction of gaze or other spatial arrangements, for example an outstretched arm, or a road running through the image.
Multimodality refers to how simultaneously-present different modes and their modal resources combine to communicate a message – it is more than multiple modality. Multimodality in the picturebook – a material, cultural artefact – including the cover, however, appears largely neglected in the academic literature (but see Walsh, 2005, and, recently, Painter et al., 2012; see also the 2012 AERA symposium on ‘Beyond Words: Action and Animation in Young Children’s Reading, Writing, and Playing’), 2 less so in the practitioner literature (see e.g. Hassnett and Curwood, 2009). The data for this study are taken from a small but growing genre of children’s picturebooks: those featuring same-sex parents. There are, we estimate, currently around 70 of these books in existence: most are in English, and most published in America. The genre functions, we suggest, not only to provide reading material for children with same-sex parents, but also to raise awareness of such families, with a view to promoting social understanding or even celebration. In contrast to the range of work on gender representation in children’s books in general (e.g. Davies, 1989; Gooden and Gooden, 2001; Knowles and Malmkjaer, 1996; Sunderland, 2011; Zipes, 1986), there has been little academic exploration of children’s books about families with same-sex parents to date. Two exceptions are Patrick Finnessy’s (2002) defence of these books, and Virginia Wolf’s (1989) essay lamenting the paucity, inaccessibility and lack of artistic merit in those that existed at the time of writing (see also McGlashan and Sunderland, 2011; Sunderland and McGlashan, 2012). We are thus addressing two gaps in the study of children’s fiction: multimodality in picturebook covers, and the ‘same-sex family’ picturebook sub-genre.
Multimodality, we argue, is of particular relevance to picturebooks featuring gay parents. The subject matter is both sensitive and controversial – always potentially subject to negative responses from pressure groups, politicians and even the law. 3 Explicit ideational representation of gay identity may be counterproductive, and so what cannot be specifically written may be suggested visually; what cannot be depicted may be just hinted at in words. A challenge for the writers and illustrators of these books is to represent the parents as partners, not just friends or co-carers. One way to do this is to index physical contact – visually or textually. But not only is physical contact between gay people in public a marked social practice to the extent of often being seen as transgressive and deviant, it is additionally marked representationally in that heterosexual parents are rarely shown in physical contact in the wider picturebook genre (Women on Words and Images, 1975). So do these picturebook covers index physical contact, in which mode(s), and how?
Earlier Work
In our own work (McGlashan and Sunderland, 2011) on a ‘corpus’ of 29 children’s stories involving same-sex parents, we identified three different strategies authors use to represent parents’ gayness in a positive way: ‘different’, ‘backgrounding’ and what we now call ‘upfront’. ‘Upfront’, involving explicit explanation of gayness by a parent to a child, including the word gay, was least frequent; ‘backgrounding’, i.e. showing the family doing things that families (stereo)typically do, with no reference to gayness, the most frequent. In a follow-up study (Sunderland and McGlashan, 2012), we looked linguistically at the book titles and the naming of the parents. Titles of books about two-Mum households tended to be more explicit about the nature of the parents’ social roles and relationship than those titles featuring two-Dad households (e.g. Heather Has Two Mommies cf. Daddy’s Roommate). As regards naming, the most frequent pattern for Mums was nomination + categorisation (e.g. Mom Sara), whereas for Dads it was nomination (e.g. Frank) alone (see Van Leeuwen, 1996, 2008). In some contrast, the linguistic and visual representation of physical contact between the parents was more likely to suggest a physical (i.e. gay) relationship between the Dads than the Mums, and in this sense Mums were thus primarily represented as parents, Dads as partners. In a specific examination of the linguistic text and images in two comparable double-page-spread internal images: Mummy, Mama and Me, and Daddy, Papa and Me, the Dads were again represented multimodally more strongly as in a sexual relationship, something that may have been unintentional on the part of the authors and illustrators.
The notion of modal affordances (Bezemer and Kress, 2008), key to this study, refers to what a mode can (easily) express and represent, with reference to its potentials and constraints. Most broadly: text tells, image depicts. If a character in a children’s book is a cook, the writer may tell us various details about her background, which would be hard for the illustrator; the illustrator might depict her in detail, whereas the writer may not wish or feel able to use the number of words necessary to do this. The illustrator’s decisions about selection extend beyond what/who to include, and how, to spatial arrangements. Illustrators thus often spatially arrange and foreground/background in some detail, whereas writing, as Bezemer and Kress note, ‘affords’ greater generality than does image (see also Martinec and Salway, 2005: 350). This does not mean that an image must be more specific than any associated written text, but that it has the potential to be so because of its different ‘functional specializations’. Accordingly, in a multimodal text, meaning may be distributed differently across modes, and, because of its associated affordance, each mode in a ‘multimodal ensemble’ can (and may) realise different communicative functions (see Jewitt, 2011a, 2011b). Genre is also important here. As Van Leeuwen (2008: 136) observes: In many contexts of communication, the division of labor between word and image is more or less [that] words provide the facts, the explanations, the ‘things that need to be said in so many words’; images provide interpretations … [but] in many domains of science and technology, visualizations are seen as the most complete and explicit way of explaining things, and words become supplements, comments, footnotes, labels.
So what is the likely modal ‘division of labour’ for picturebooks, and picturebooks featuring same-sex parents specifically, bearing in mind publishers’ need for caution here?
In addition to spatial arrangements, an illustrator must also consider specificity of appearance: what a pictured character looks like (hair, clothes, adornment). Kress (2003: 3) refers here to the epistemological commitment of modes, i.e. an ‘unavoidable affordance’, exemplifying this with the cell/nucleus relationship in the context of a science lesson: in writing there is ‘commitment to the naming of a relation … “the cell owns a nucleus”’; with image there is ‘commitment to a location in space … “this is where it goes”’ (see also Kress, 2010). Bezemer and Kress (2008: 176) more usefully (for us) exemplify epistemological commitment in relation to the production of a particular image: The author might have given a written description like, ‘Sitting in the late autumn sunshine, Sam and Bill share a bench in the park’. An illustrator or designer might have been asked to ‘draw across’, to transduct, the written description into the mode of image. Now the illustrator has to ask ‘How close to each other were they sitting?’; Was Bill to the left or right of Sam?’ The translator/transductor has to be precise, whether she or he wishes to do so or not.
The author could have indicated Sam and Bill’s proximity on the bench but their illustrator must show this precisely. Given the possibility of physical contact as a way to index gay identity, the notion of ‘visualised proximity’ is very relevant to any visual representation of gay parents.
Taking as given the importance of modal resources, modal affordances and epistemological commitment, as well as that of modal interaction, the question of the nature of the interaction between text and image, i.e. how different simultaneously-present modes combine to communicate a message, remains. This area of study is young and proposed analytical systems for text–image relationships are contested (see e.g. Martinec and Salway, 2005; Royce, 2007; Unsworth and Cléirigh, 2011). Although we would expect a measure of coherence, including specific cohesive ties between some words of the text and some elements of the image, there still exists a range of actual possibilities (i.e. different text–image relationships) and different analytical possibilities. Royce (2007) for example identifies cohesive ties between image and text through relations not only of synonymy, but also meronymy (the relation between part and whole). Martinec and Salway (2005) consider the relative status of text and image, the higher status mode being the weightier for meaning. One mode, either text or image, might then simply ‘support’ the other, while being ‘anchored’ by it (Barthes, 1977), or text–image relations may be equal (‘relay’). Unworth and Cléirigh (2011: 153) however challenge the whole value of looking at relative status of modes, arguing for ‘the reciprocity of the different affordances of image and text’, i.e. that the different modes are mutually informing. Kress and Van Leeuwen (2006: 177) adopt a similar stance: ‘the parts should be looked upon as interacting with and affecting one another’ (emphasis added). While this stance has analytic potential, it may be that actual text–image relations also vary with genre – in this case, that of a sometime transgressive picturebook subgenre.
Like others exploring multimodality, Unworth and Cléirigh (2011) draw on Hallidayan systemic functional grammar (Halliday, 2002, 2004), interpreting SFG as a semiotic rather than linguistic theory (Pagani, 2010). While not all SFG terminology is relevant to non-linguistic modes, much is. In particular, Halliday (2004: 377) utilises and adapts the term Expansion, relevant to clauses and clause complexes, which ‘relates phenomena as being of the same order of experience’, i.e. the phenomena are not spoken or thought about (which would constitute a higher order of experience). The type of Expansion in question is what Halliday (2004) calls enhancement, i.e. one clause ‘enhances the meaning of another by qualifying it in one of a number of possible ways: by reference to time, place, manner, cause or condition’ (p. 410). Halliday exemplifies this with cause in the paratactic ‘John was scared, so he ran away’ and the hypotactic ‘John ran away, because he was scared’ (p. 380). Interestingly, he also uses the notion of ‘x (‘is multiplied by’)’ (p. 377) for enhancement, suggesting ‘more than the sum of its parts’, which we suggest is particularly relevant to text–image relationships. Halliday’s list of the qualifications of enhancement remains open-ended, and below we propose possible additions of event and co-participants.
Accepting that enhancement can apply semiotically to image as well as written text, we can then look at multimodal texts and consider the possibility of what Unsworth and Cléirigh (2011) call mutual enhancement of text and image. In looking at image–text interaction on picturebook covers, we are thus contributing to the debate on how the semantics of image and language ‘co-articulate’ (Unsworth and Cléirigh, 2011; see also Page, 2010).
This Study
In addition to our use of the concepts of modal affordance, epistemological commitment and (mutual) enhancement, this study also draws on a particular approach to multimodality, and combines this with discourse analysis. Jewitt (2011b) identifies three approaches to multimodality: social semiotic, multimodal discourse analysis and multimodal interaction analysis (see also Bezemer and Jewitt, 2010), although these are not clear-cut, and, indeed, still in the making. The first two have both been shaped by social semiotic theories of communication, including Hallidayan SFG, which allows images as well as linguistic units to have ideational, interpersonal and textual metafunctions.
Social semiotics draws on the traditional semiotic Sausseaurean notion of the sign, constituted by the signifier (e.g. a word, gesture, red traffic light) and the signified (what that word, gesture or red traffic light means) (e.g. Peirce, 1931). A signifier can be symbolic (e.g. a red rose to mean love); iconic, based on resemblance (e.g. a smiley face to mean happiness); or indexical, where there is a cause–effect relationship between signifier and signified (e.g. a footprint). Social semiotics usually entails a refusal to privilege language (here, the picturebook cover titles) over image (see Kress and Van Leeuwen, 2006). We follow Kress (2010) in seeing social semiotics as additionally concerned with how signs are made – here, through authorial and illustrator choices – rather than how they are used, and that … signs are motivated, not arbitrary relations of meaning and form; the motivated relation of a form and a meaning is based on and arises out of the interest of sign-makers; the forms/signifiers which are used in the making of signs are made in social interaction and become part of the semiotic resources of a culture. (Kress, 2010: 54, original emphases)
Here, it is the illustrators and writers, the sign-makers, who have ‘interests’ (see also Kress, 1993) and their selected signifiers are hence potentially highly motivated – presumably reflecting these authors’ and illustrators’ desire to represent the co-parents in these picturebooks in a positive, progressive light. However, as Jewitt (2011b: 31) points out, we also need to ‘understand the social beyond its articulation through the individual sign-maker’ for this study; this means taking on board the diversity of social attitudes to gay and lesbian practice, identity and desire, which create one aspect of the ‘social interaction’ which shapes the making of these signs. In this respect, while the cover is a key way publishers flag up what the book is about (and an important ‘unit of analysis’ on the social basis that covers are designed as ‘marketing units’, for example on websites and in catalogues), it is also a marketing tightrope as, from a commercial perspective, it needs to have ‘shelf appeal’, and not damage sales or attract opprobrium. Accordingly, it is almost invariably the publisher rather than the writer or illustrator who has the final say (which sometimes impacts on the illustrations within the book). 4
By discourse analysis – introduced into our study given the written titles on these picturebook covers – we mean socially situated but close analysis of language, as well as the provisional identification of relevant discourses (see e.g. Sunderland, 2004), of which these titles and their constituent elements can be seen as linguistic ‘traces’. While we are concerned with the titles’ potential ideational and interpersonal meanings, as these interact with the images, in accordance with social semiotics we are not privileging the titles in terms of meaning potential, and indeed devote more space to the rich visuals than to the language of these ‘tiny texts’ (see Sunderland, 2012).
Our broad research question, then, is ‘How can the multimodal relationship between image and written text (the title) on the covers of picturebooks featuring gay parents best be described and explained?’ In our analysis of selected covers, we describe and interpret what is ‘going on’ multimodally, using Van Leeuwen’s (2008) representational frameworks that conceptualise and analyse discourse (including image) as both a form and recontextualization of social practice. For the images, we make substantial use of the Visual Representation of Social Actors frameworks. For the titles, we draw on selected categories of the Representation of Social Actors/Social Action frameworks (Van Leeuwen, 1995, 1996, 2008) to explore how the ‘participants [here, gay parents] of social practices can be represented in English discourse’ (Van Leeuwen, 2008: 23, original emphasis). Interpretation using these frameworks needs to be fine-grained and sensitive, and any meanings that image, written text and their interaction multimodally suggest to us remain open to different readings.
Our data come from a wider dataset of 25 picturebooks (not collections) featuring same-sex parents, which were all the books of this small genre in our possession at the time of writing, selected from all those picturebooks featuring gay parents in existence we were aware of (around 70) by the very practical criteria of cost and availability. The dataset picturebooks all (a) depict adults in gay/lesbian/homosexual relationship, and (b) show the gay couple caring for a child or children. Those that we have not yet collected, we believe, are in these ways no different from the more available ones. Uncle Bobby’s Wedding is included in the 25 because of Uncle Bobby’s quasi-fatherly relationship with his niece (whose biological father is ‘Radically excluded’ (Van Leeuwen, 1996, 2008) from the text). All the books are written for young children (although target age is never specified), and feature the co-parents and a child or children for whom they more-or-less permanently care.
The focus of the study is four ‘telling’ covers: Uncle Bobby’s Wedding (author/illustrator Sarah Brannen), Mom and Mum are getting Married! (author Ken Setterington, illustrator Alice Priestley), and the pair Mummy, Mama and Me and Daddy, Papa and Me (author Lesléa Newman, illustrator Carol Thompson). Our selection criteria was that both co-parents should be represented, but that the selected books should vary in their ‘explicitness’ of the image and title as regards the subject matter. In the event, on each cover the co-parents are in equal focus, i.e. are both ‘salient’, and all four covers depict a smiling ‘Actor’ (Van Leeuwen, 1996, 2008) who is looking directly ‘out of the picture’ at the viewer. All four also feature a visual boundary between image and title.
We now introduce the Visual Social Actor Networks (to which we suggest a development), and selected categories of the Social Actors network, then move on to our multimodal analysis of the book covers themselves, using Unsworth and Cléirigh’s (2011) notion of multimodal mutual enhancement and Kress’s (2003) notion of epistemological commitment.
Visual Analysis
Van Leeuwen’s Visual Representation of Social Actors framework (2008: 147) is twofold: the ideational ‘Visual Social Actor Network’ (‘How are people depicted?’) and the more interpersonal ‘Representation and Viewer Network’ (‘How are the depicted people related to the viewer?’). The Visual Social Actor network is shown in Figure 1:

Visual Social Actor Network (Van Leeuwen, 2008).
Exclusion and Inclusion refer to who is included/represented and who not (but could logically have been) – in this case, for the wider dataset, whether either or both parents appear on the book covers, and with whom. The how of Inclusion has three dimensions. If Actors are involved in action of some sort, is this as ‘Agent’, who ‘acts on’, or ‘Patient’, the ‘acted upon’? For example, someone represented pushing a child in a swing is an Agent, while the child is likely to be a Patient (although there are degrees here – the child may also be contributing to the motion). Generic/Specific and Individual/Group dimensions of Inclusion are related. ‘Generic’ means stereotypical representation of a social group member: this stereotyping may be ‘Cultural’ (e.g. same-sex parents being fastidious about their dress) or ‘Biological’ (e.g. black people being shown with frizzy hair or very white teeth). ‘Group’ means similar/different visual representation of members of particular groups (i.e. Homogenization/Differentiation, respectively; Van Leeuwen’s example is photographs of Allied soldiers from the first Gulf War being portrayed as individuals, Iraqi soldiers as groups). The visual representation of gay parents could indeed be (culturally) stereotypical, and there is always the potential for gay co-parents to be depicted similarly (stereotypically or otherwise).
The ‘Representation and Viewer Network’ (Van Leeuwen, 2008), which considers represented ‘relationships’ between people in an image and the viewer (interpersonal meaning) through their imaginary contact (see Kress and Van Leeuwen, 2006), is shown in Figure 2:

Representation and Viewer Network (Van Leeuwen, 2008).
The three Representation categories of Distance, Relation and Interaction each entail imaginary ‘vectors’ (or their absence) between those represented and viewers. ‘Distance’ refers to whether individuals are shown using a long shot or close-up (or are in the ‘middle distance’). Van Leeuwen (2008: 138) claims that ‘People shown in a “long shot”, from far away, are shown as if they are strangers; people shown in a “close-up” are shown as if they are “one of us”.’
‘Relation’ has two dimensions: ‘Involvement’ and ‘Power’, both functions of the camera angle. ‘Involvement’ is suggested by the ‘horizontal angle’: ‘whether we see a person frontally or from the side, or perhaps somewhere in between’ (p. 139), so that frontally suggests involvement, obliquely, detachment. Van Leeuwen also proposes that when the viewer looks down on the Actors, he or she ‘exert[s] imaginary symbolic power over that person’, occupying ‘the kind of “high” position which, in real life, would be created by stages, pulpits, balconies, and other devices for literally elevating people in order to show their social elevation’ (p. 139).
In contrast: ‘To look up at someone signifies that the someone has symbolic power over the viewer, whether as authority, role model, or something else. To look at someone from eye-level signals equality’ (p. 139).
‘Interaction’ entails either ‘Direct address’ or ‘Indirect address’, referring to whether the represented Actor is looking at the viewer. Van Leeuwen claims, strongly, that: If [the depicted people] do not look at us, they are, as it were, offered to our gaze as a spectacle for dispassionate scrutiny. The picture makes us look at them as we would look at people who are not aware we are looking at them, as ‘voyeurs’, rather than interactants. If they do look at us, if they do address us with their look, the picture articulates a kind of visual ‘you’, a symbolic demand. The people in the picture want something from us … (pp. 140–141)
He concedes, though (referring to a case of ‘Relation’), that ‘Just what this means precisely will, of course, be coloured by the specific context’ (p. 139).
Bringing Distance/Relation/ Interaction together, Van Leeuwen then identifies three possible ‘strategies’ for ‘visually representing people as “others”, as “not like us”: distanciation, disempowerment and objectivation’ (p. 141). If illustrators of picturebooks featuring gay parents wish (as we assume they do) to present these parents as, basically, ‘like us’, they might then use strategies of, say, proximisation, empowerment and interaction (our terms). While it is unlikely that illustrators would intentionally deploy distanciation, disempowerment or objectivation, analytical vigilance is apt, given the discourses of heteronormativity (and sometimes homophobia) to which we are all subject (see e.g. Talbot, 2010) – as well as the marketing ‘tightrope’ noted above.
We see Van Leeuwen’s interpretations as possibilities rather than definitive. Space must always be created for wider, contextually-informed understandings and alternative readings: people shown in close-up, for example, may be experienced as menacing rather than ‘one of us’; 5 people who are not looking at us (Indirect address) could be seen not only as objects of our voyeurism but also, say, regal, aloof and/or indifferent.
Multimodal Analysis of the Four Book Covers
Below, then, we use the two Visual Representation of Social Actors frameworks in conjunction with categories from Van Leeuwen’s (1995, 1996, 2008) linguistic Representation of Social Actors/Action frameworks to address ‘the intersemiotic semantic relationships between images and language to show how the visual and verbal modes interact to construct the integrated meanings of multi-modal texts’ (Unsworth and Cléirigh, 2011: 151, emphases added). In doing so, we seek evidence for workings of mutual enhancement of text and image.
To look first, briefly and linguistically (only), at the two-Mum and two-Dad book titles: in Sunderland and McGlashan (2012), we proposed six ‘levels of explicitness’, including ‘Implicit’ (as we read this), in terms of linguistic representation of the sexual identity of the same-sex parents. ‘High’ explicitness, we argued, was achieved not only topically or lexically but through particular linguistic (syntactic) patterning: use of a declarative main clause (vis à vis phrase or subordinate clause) in combination with the social actor category of categorization (vis à vis ‘nomination’; Van Leeuwen, 1996, 2008). We identified as the three ‘most explicit’ titles: Heather has Two Mommies Josh and Jaz have Three Mums Mom and Mum are getting Married!
As declarative main clauses, and unmodalised (hence able to function as propositions), these titles carry a certain epistemic authoritativeness about what is. In addition, the Mums of Heather, Josh and Jaz, and those in the third title, are categorized, i.e. given a ‘label’ associated with ‘identities and functions [social actors] share with others’ (Van Leeuwen, 2008: 40): Mommies, Mums, Mom, Mum. We are thus told about their social role as parents, which would not be the case if they were simply nominated (e.g. Sara). Even these three titles are not however completely explicit: Heather’s Mommies could refer to a biological Mum and a step-Mum, as could Josh and Jaz’s Mums (plus, say, a father’s ex-partner), and Mom and Mum do not have to be marrying each other. (All three are of course two-Mum books, and there are in fact salient differences in gender representation in this dataset in the naming of the parents (see above, also Sunderland and McGlashan, 2012).
Exploring the explicitness of a book cover however also requires consideration of the visual as well as the text (title)–visual multimodal interaction. This we do below, extending the linguistic notion of enhancement from clauses and clause complexes (in Halliday’s linguistic SFG) to nominal (non-clausal) constructions such as noun phrases, characteristic of book titles.
The cover of our first book does not index gayness at all in the (noun phrase) title Uncle Bobby’s Wedding (recall that above we reported Uncle Bobby’s quasi-fatherly relationship with his niece); we also chose this book cover (see Figure 3) because of the title’s low ‘level of explicitness’. Visually, however, and using the Visual Social Actor Network, Included are three Actors: two guinea pigs (one larger and placed to the centre-right) and a (much smaller) squirrel. The choice of animal characters allows the publishers to sidestep the arguably more socially threatening image of (and story about) two male humans marrying. 6 The guinea pigs are depicted as smiling, in black morning suits and bow-ties; the squirrel in a pink dress is holding a bouquet. The guinea pigs are standing very close together. Recall that the ‘epistemological commitment’ of the image mode means that, in this case, the illustrator has to show the Actors’ proximity and what they are wearing. The ‘Action’ consists of the squirrel looking up at the guinea pigs, the larger of whom has his left arm/paw on her left shoulder in an affectionate way: in this sense he is ‘Agent’ and she ‘Patient’. The guinea pigs’ portrayal, not least because of the anthropomorphism characteristic of children’s books about animals, is as Specific as it can be, and they are also depicted as Individuals, being of slightly different size and colouration.

Front cover of Uncle Bobby’s Wedding (Brannen, 2008).
Given the guinea pigs’ masculine formal attire, the squirrel’s pink dress, and her bouquet (a symbolic signifier of a wedding), the reader is likely to experience some form of cohesion between the title and the image, i.e. that the ‘wedding’ in question is that of the two individuals pictured in morning suits, even though both appear to be male; and that the larger one to the centre-right is likely to be Uncle Bobby. In terms of ‘Information value’, Kress and Van Leeuwen (2006: 181) argue that ‘the elements placed on the left are presented as Given, the elements placed on the right as New’, mirroring eye-movements when reading texts in English. The relevance of this claim to this image can however only be seen in relation to the title.
Further visual ‘wedding’ clues are that the image is placed within a literal ‘frame’ with a bow at the bottom (potentially indexing a wedding photo), and the title itself, below the image in large red letters but outside this frame, which is in an ornate font characteristic of a wedding invitation. As the visual elements allow the reading of both gayness and a wedding, the title confirms, indeed, enhances, the meaning of the image as a wedding (as opposed to a formal occasion), i.e. with ‘what’ is going on, the event, and ‘who’, i.e. the co-participants. In turn, the image enhances the meaning of the title as a gay wedding (Uncle Bobby is marrying another male); a further enhancement of ‘what’. As long as a ‘gay wedding’ is part of the reader’s schemata, the text/image interaction achieves explicitness in ways that neither text nor image alone could have done; here, as Nikolajeva and Scott (2000: 226) put it, ‘the visual and the verbal aspects are both essential for full communication.’ (If it were not, the reader might see the three pictured Actors as Uncle Bobby, his best man and a bridesmaid, perhaps finding it slightly odd that Uncle Bobby’s guinea pig bride is apparently not shown.)
The Visual Social Network however as it stands does not include the dimension of ‘Spatial arrangement’ (and hence proximity) of the Included Actors, to which the illustrator is epistemologically committed. The two guinea pigs in this cover image are standing so close together that there is no visible space between them; indeed, their arms/paws could be round each other. The spatial affordance of image appears important to a reading of this book cover, and, we propose, could be fruitfully added to the dimensions of Inclusion in the Visual Social Actor Network.
As regards the interpersonal question: ‘How are the depicted people related to the viewer?’ (Representation and Viewer Network), in terms of Distance, the ‘Uncle Bobby’ actors pretty much fill the cover, and hence are ‘Close’ to us. In terms of ‘Relation’, the angle of view, we are looking directly at all three Actors, at a frontal angle, and neither up nor down at them (although in terms of the previous Network, the squirrel is looking respectfully upwards). In terms of ‘Interaction’, the larger guinea pig is looking directly at the viewer, making this Actor potentially the agent and subject (as well as object) of gaze, another reason for the reader to see the larger guinea pig as the protagonist in the title, i.e. Uncle Bobby.
Drawing on Van Leeuwen’s (2008: 138) interpretation of these parameters, then, as Uncle Bobby is ‘Close’ to the viewer, he is represented as ‘one of us’, and in a way which suggests involvement rather than power. The represented gaze can be interpreted in different ways, but in terms of what the smiling Uncle Bobby ‘wants from us’, one reading is that the viewer is being confronted unaggressively but unashamedly with this gay wedding (with all its heternormative ‘trappings’ of tradition and respectability), and is being asked to ‘deal with it’.
The other three covers, we argue, unlike Uncle Bobby’s Wedding, do signal ‘gay parents’ through their titles, if not unequivocally. As shown above, we included the (clausal) Mom and Mum are getting Married! in the ‘Most explicit’ group as regards sexual identity, the clause being unmodalized and declarative (Sunderland and McGlashan, 2012). In the Social Action network (Van Leeuwen, 1995, 2008), marry represents Material Action. In traditional grammar terms it can be described as an ‘ambitransitive’ verb, i.e. one able to take an object (e.g. She married him), or not (e.g. When she was 25 she married). The verb phrase get married functions similarly, but here it is used ‘non-transactively’ (Van Leeuwen, 2008: 60): linguistically, there is no ‘Goal’ (object). Decontextualised, the title alone does not explicitly show Mom or Mum as ‘hav[ing] an effect on others, or the world’, in part as it does not indicate denotatively who either Mom or Mum are marrying, and indeed heteronormative thinking would suggest ‘a man’ (hence not each other). It does, however, suggest immediacy and excitement through the exclamation mark
The cover of Mom and Mum are getting Married! (see Figure 4) visually depicts a celebration of a happy event, and – because of the petals being scattered which extend across the cover – a wedding.

Front cover of Mom and Mum are getting Married! (Setterington, 2004).
Petals are, we suggest, symbolic signifiers of a wedding in many ‘Western’ countries; further, these petals are all pink, another symbolic signifier, this time of ‘things feminine’ (Koller, 2008). Included are two youngish women, and a boy and a girl, the boy in a suit and tie, the girl in a dress. All are active, apparently moving forward. The children are holding baskets and energetically scattering the petals. The women’s clasped hands, which notably are in the extreme centre of the image, are also carefully framed by the handle of one of the baskets – the most obvious within-image boundary in the four covers. Their smiles suggest agency – that they are acting of their own volition. The women are represented Individually, with different clothes and hairstyles (the representation of which the illustrator was epistemologically committed to, once it was decided that the women would be on the cover). The appearance of the woman on the left is however relatively ‘feminine’ (blonde wavy hair, lilac top, pendant necklace) in traditional terms, whereas the woman on the right has very short, dark hair and is wearing a shirt, dark red waistcoat and chunky necklace. A left-to-right reading would mean that the blonde woman is ‘Given’ information, the dark-haired woman ‘New’ (and indeed we could argue that her depiction is rather ‘socially familiar’ as a stereotypical butch lesbian). While she may also be being positively distanced from traditional (hegemonic) representations of femininity, the representation of the couple together can also then be seen as drawing on butch/femme lesbian stereotypes, i.e. a suggestion of Generic representation.
This depiction (the celebratory nature of the image, the petals, the clasped hands, the rather differently-pictured women) then suggests that, contrary to heteronormative thinking, these women are getting married to each other (not to two men, as heteronormative discourses might lead us to expect). The non-transactivity of getting married in the title is thus addressed, and any ambiguity resolved. The lack of a ‘Goal’ for the Material verb marry is supplied visually.
As with the cover image of Uncle Bobby’s Wedding, spatial affordance is important. In terms of spatial relations, the two women are walking together, leaning in towards each other, and, as indicated, openly holding hands. Given that illustrators are epistemologically committed to showing proximity and possible/actual contact between any two depicted Actors, this again points to ‘Spatial arrangement’ as a fruitful additional parameter of Inclusion in the Visual Social Actor Network.
As regards their depiction in relation to the viewer (Representation and Viewer Network), Mom and Mum occupy the middle ground against a background of trees, and are proximally only slightly more Distant than the children in the foreground. We are looking at them horizontally, again suggesting Involvement, and at eye-level, suggesting Equality. While the woman on the left is gazing diagonally down towards the girl, signalling Indirect address, the other (like Uncle Bobby) is looking directly at the viewer, making her, like him, the subject as well as object of gaze. Despite her smile (again like Uncle Bobby), one reading is that she too is expecting us to ‘deal with’ the situation: note that both of these books focus not just on same-sex relationships, but the institutional, public and positive recognition of these. It may not be a coincidence that, of the two women in the Mom and Mum are getting Married! image, the one whose gaze meets ours is the one more ‘Generically’ (stereotypically) depicted as gay.
The title of the book, above the image and outside the image ‘frame’, occupies relatively little space compared to Uncle Bobby’s Wedding. Nevertheless, Mom, Mum and Married! are at least twice the size of the other words, and are in yellow (lined with the same pink as the petals); the three smaller words are white and unlined. The visual features make clear that Mom and Mum of the title are marrying each other, not men, depicting something like: ‘Yes, they really are – and to each other!’ Again we can see the title and visual as mutually enhancing. The visual enhances the title with ‘to whom’ (i.e. tells us who the co-participants really are), but we can also add Halliday’s ‘where’ (outdoors) and even ‘when’ (judging from the clothes, a warm season). The title in its turn makes it crystal clear that the celebration evident in the visual is a wedding (entailing a socially-recognized marriage) – an enhancement of ‘what’ (again, event).
The last two book covers are Mommy, Mama, and Me, and Daddy, Papa, and Me (see Figures 5 and 6). We see these two titles, both noun phrases, as more explicit in terms of gay sexual identity than Uncle Bobby’s Wedding but less explicit than Mom and Mum are getting Married!, in part because their phrasal nature allows a greater range of possible readings, many of which do not entail the Actors in question being a couple. All the Actor referents in the titles are capitalized, and in red, against a contrasting, paler background. Both titles curve gently above the images. Unlike the other two titles, they share the same frame as the corresponding image, although they are still separated by a boundary created by the names of the author and illustrator.

Front cover of Mommy, Mama, and Me (Newman, 2009a).

Front cover of Daddy, Papa, and Me (Newman, 2009b).
Both covers visually index close physical contact. On the former, two youngish women are shown in profile, looking down at a very young smiling child whom they are holding between them. On the latter, one youngish man is shown in profile, one in half-profile; both are looking up, the former holding a very young smiling child up in the air. Available to the reader is then a potential cohesion between the adults Included in the visuals and those referred to in the titles, i.e. that the adults are Mommy and Mama, and Daddy and Papa. The Actors are also ‘Involved in action’ – as Agents in relation to the children, although one man is particularly active, and the women are more involved in closeness. This echoes the gender representation of parents in children’s fiction more broadly (see e.g. Stephens, 1992; Sunderland, 2011) but indirectly both further index all these adults as parents, albeit same-sex ones. The Actors are also represented Specifically, or at least potentially so: one woman’s ear-piercing, and one man’s somewhat shaven head, which may function as stereotypical ‘gay’ representations in some contexts, are also now familiar more broadly in many ‘Western’ urban contexts. Like Mum and Mom, the women are depicted differently, i.e. Individually, as are the men. Unlike Mom and Mum are getting Married!, no action is lexically or grammatically represented in these titles.
Again, the spatial arrangements and proximity in the images are important. Although all the adults are framed only above the shoulders/neck, projecting downwards, the viewer can ‘see’ that the two women could be hugging each other, and that both arms of one of the men may be round the other (or, at least, that their proximity is such that there is no space between them). Projected, we can thus see both cover images as exemplifying the (gay) ‘family embrace’ also identified in several visuals inside other books (Sunderland and McGlashan, 2012). This is subtly enhanced by the curve of the titles which ‘overarch’ these depicted families.
We can again convincingly propose mutual enhancement of image and text. Once the likelihood of equivalence between those Actors linguistically referenced in the titles and those pictured is established, the titles indicate that those pictured are more than just professional carers (providing an enhancement of ‘who’, i.e. the nature of these co-participants); in turn, the titles, which do not indicate the relationship of Mommy and Mama, or of Daddy and Papa, are enhanced by the visuals, which do suggest the nature of the relationship (an enhancement of ‘how’ these adult co-participants are parents). In a multimodal sense the visuals ‘complete’ the phrasal titles by suggesting main clauses such as Mommy, Mama and Me … are a Happy Family.
Looked at through the lens of the Representation and Viewer Network, being in the foreground of the images, both families are proximally ‘Close’ to the viewer. Their orientation (Relation) is on the horizontal plane, which again suggests viewer involvement. In regard to Interaction, these parents’ gaze is directed at their children, and they are thus addressing the viewer only Indirectly. In contrast (and to the previous two books), both children (wearing tops of an identical and rather striking shade of red) are looking very directly, and rather intently, at the viewer, which suggests not only that each is the Me of the title (to which the left–right Given/New distinction may also be applicable), but that they too are addressing us directly (again, in a subtle, happy, but arguably ‘Deal with it’ sort of way). Being at the top of the image, the child in Daddy … is looking down, but is s/he also looking down at us, and, if so, might we interpret this, provisionally, as ‘Representation has power over viewer’ (Van Leeuwen, 2008: 141)? The cover of Daddy, Papa, and Me would then in this sense have more such ‘power’ than that of Mommy, Mama, and Me, echoing the more general, more explicit ‘dads-as-partners’ representation mentioned earlier.
While neither of these two titles nor the images are explicit in terms of gay identity, when read in conjunction, the mutual enhancement of title and image again construct a much more robust potential reading in terms of co-participants, i.e. that these Actors are indeed co-parents: two Mums and two Dads. This is particularly so when these books are seen as a pair, as they then begin to represent a community of gay co-parents.
For all four book covers, the ‘information value’ of the images, i.e. the placement of elements (see Kress and Van Leeuwen, 2006: 177), what is centred/peripheral, left/right and top/bottom, is of interest given that young children’s book covers are epistemologically committed to both a title and (almost always) an image, and that one is usually on top, one below (as opposed to left/right or centre/periphery). Whereas the Uncle Bobby’s Wedding title is below the image, the other three titles are above. Kress and Van Leeuwen (2006) propose that whatever is on top plays the ‘lead role’, which would then further weaken the already ‘inexplicit’ Uncle Bobby’s Wedding. Because of text–image mutual enhancement, however, the cover as a whole is still relatively explicit in terms of sexual identity.
Conclusion
We hope to have shown that these four examples of text–image interaction are not just cases of ‘the words expand[ing] the picture’ or the ‘pictures amplify[ing] more fully the meaning of the words’ (Nikolajeva and Scott, 2000: 225) but of multimodal mutual enhancement of the small title-texts and associated images, in particular, enhancement of event (what) and with whom (co-participants). While none of the four titles or images alone offers an unequivocal reading to this effect, this enhancement enables us to read these picturebook covers as featuring same-sex parents.
Unsworth and Cléirigh (2011: 154) argue that although text and image ‘interact synergistically’, it is ‘their essential incommensurability that enables new meanings to be made from combinations of modalities in texts’; see also Lemke, 2000). The modal affordances of image are not the same as those of written text. We have shown that while an image cannot state that something is the case, it can depict it through relevant detail, including culturally-understood symbolic signifiers such as petals and bouquets, and, importantly, marked close proximity between social actors, degree of proximity of depicted Actors being something to which the image mode is epistemologically committed. And while a children’s book title cannot conventionally be lengthy, or state prosaically that something is the case, through carefully chosen lexis and syntax, it can make the likelihood of a particular reading of a cover more or less available.
Van Leeuwen’s (2008) Visual Representation of Social Actors frameworks constitutes a valuable way of mapping out which Actors are represented, as what, what they are doing, and relationships between Actors and viewers. However, we hope to have shown that ‘meaning’ here is always contingent, co-constructed, interpretive, and cannot simply be read off from the covers via the frameworks. We also propose the category of ‘Spatial relation’ (with a focus on proximity) as a useful addition to the Visual Social Actor network as it is particularly productive for the co-construction of meaning around (socially marginalized) portrayed relationships. ‘Spatial relation’ includes touch but may even extend to gaze (imagined proximity?) between represented Actors (commensurate with exploration of gaze between Actor and viewer).
While mutual enhancement may be particularly characteristic of the multimodality of this sometime transgressive and hence socially (and educationally) challenging sub-genre, it may be less usual in multimodal texts in other genres. A next step may then be to look at other cases of potentially transgressive text–image combinations – perhaps pedagogic texts on sex-education in contexts where such explicit education is widely deemed inappropriate. A second may be to look at other cases of image accompanied by small amounts of written text; in the field of children’s literature, such a sub-genre might be very early school readers. Both would be of interest for educationalists and multimodal discourse analysts alike.
Footnotes
Acknowledgements
For permission to reproduce the covers of Mommy, Mama and Me and Daddy, Papa and Me, we thank Lesléa Newman, Carol Thompson and Ilse Craane from Tricycle Press. For permission to reproduce the cover of Uncle Bobby’s Wedding we thank Sarah Brannen and Lynn Bennett, and Paula Sadler and Mary Sullivan at Putnam. We would like to acknowledge the illustrator Alice Preistley for permission to reproduce the cover of Mom and Mum Are Getting Married! by Ken Settering and published by Second Story Press, Toronto, Canada (see
, accessed 13 December 2011). Lastly, thanks go to Veronika Koller and Karin Tusting for reading and making insightful insights on an earlier version of this paper.
Notes
Biographical notes
JANE SUNDERLAND is a Senior Lecturer in the Department of Linguistics and English Language, Lancaster University, where she teaches ‘Gender and Language’ on different postgraduate programmes. She is also a past President of the International Gender and Language Association (IGALA). Her interests include gender and language in relation to children’s fiction, Africa, and boys’ literacies – with a particular focus on Harry Potter. Two recent publications are: Language, Gender and Children’s Fiction (Continuum, 2011) and ‘Brown sugar’: The textual construction of femininity in two “tiny texts”’ in Gender and Language, 2012, 6(1): 105–130).
Address: Lancaster University, Linguistics and English Language, Lancaster LA 14YL. [email:
MARK MCGLASHAN is an ESRC-funded doctoral student and tutor at Lancaster University’s Department of Linguistics and English Language. His main research interests include children’s literature (specifically, picturebooks), (critical) discourse analysis, discursive psychology and identity. Recent publications include: ‘The linguistic, visual and multimodal representation of two-Mum and two-Dad families in children’s picturebooks’ with Jane Sunderland, in Language and Literature, 2012, 21(2): 189–210; and ‘The branding of European nationalism: Perpetuation and novelty in racist symbolism’, in Ruth Wodak and John Richardson (eds) Analysing Fascist Discourse: European Fascism in Talk and Text (Routledge, 2013).
Address: as Jane Sunderland. [email:
