Abstract
AI-generated images now function as routine multimodal texts, yet we know little about how they construct meaning beyond minimal prompts. Using an SFL-informed multimodal framework, this study analyses intersemiotic expansion, projection, and affect in a purpose-built corpus of 120 text-to-image outputs produced with a strictly controlled two-line prompt. Four subject types differing in animacy and conventional agency were crossed with six minimally specified predicates spanning material, mental, verbal, and projective processes (five independent generations each). Coding shows that divergence from prompts follows stable semiotic patterns rather than random variation. Action-oriented predicates strongly license specification and addition, whereas predicates foregrounding internal states rely more heavily on projection and selective specification. Subject type modulates these strategies: schema-rich human subjects constrain variation, while biologically misaligned or non-agentive subjects trigger compensatory environmental reconfiguration, explicit projection, and heightened affective loading. Projection peaks where subject plausibility is lowest, and affect intensifies for prompts involving desire or internal states. These findings suggest that AI-generated images operate as multimodal meaning-making artefacts that actively reconstrue ideational, interpersonal, and textual meanings, rendering culturally sedimented visual schemas analytically visible.
Keywords
Introduction
In recent years, AI-generated images have become increasingly visible within everyday multimodal communication across a range of creative, professional, and online social contexts. Advances in text-to-image systems and vision-language models have made it possible to produce visually coherent images from minimal verbal input, often drawing on established artistic and photographic conventions (Zhou et al., 2026). These systems are typically discussed in terms of technical performance or creative potential, yet less attention has been paid to how they construct meaning through the relation between prompt and image. As such images proliferate, they increasingly participate in shaping visual common sense, normalising culturally shared ways of visualising action, relation, and interiority.
At the same time, generated images often diverge in systematic ways from their prompts. Prior research has documented how output varies depending on model architecture and prompt formulation (Hu et al., 2023) as well as persistent limitations in semantic grounding, including instances commonly described as visual hallucination (Chen et al., 2026; Niu et al., 2025). From a social-semiotic perspective, however, such divergences are often not simply failures of alignment. Rather, they can be understood as instances of semiotic expansion, in which ideational, interpersonal, and affective meanings are added, displaced, or reconfigured as language is transformed into image (Cope and Kalantzis, 2024; Halliday and Matthiessen, 2014; Kress and Van Leeuwen, 2021). Analysing these expansions requires attention to how different kinds of depicted entities afford or constrain particular meaning-making strategies. As Kress (2003: 50) notes, “awareness of the affordances of modes and the facilities of media provides competence,” an observation that takes on renewed relevance in the context of contemporary text-to-image generation.
Existing research on AI-generated artefacts has largely relied on large datasets and aggregated evaluation protocols (Otani et al., 2023). While linguistic frameworks such as Systemic Functional Linguistics (SFL) have been productively applied to the analysis of AI-generated written texts (e.g., Mateen, 2025), SFL-informed analyses of AI-generated images remain comparatively rare, particularly with respect to how minimal prompts are expanded into ideational, interpersonal, and affective meanings across different depicted entities. By bringing systemic functional multimodal theory into dialogue with contemporary generative image systems, this study positions AI-generated images as sites where culturally sedimented semiotic resources become empirically observable. Guided by the following research questions, it contributes to ongoing debates in multimodality research concerning how emerging media reshape the distribution of ideational, interpersonal, and textual meanings. 1. How do AI-generated images expand and reconfigure meaning beyond minimal verbal prompts through intersemiotic relations and affect? 2. How do different subject types (person, fish, robot, spoon) condition the ways in which AI-generated images expand and reconfigure meaning beyond minimal verbal prompts?
Literature review
Research on AI-generated artefacts
Much existing research on AI has focused on written text in pedagogical and instructional contexts (Alaqlobi et al., 2024). By contrast, fewer studies address how AI-generated artefacts construct meaning beyond task performance or learning outcomes, particularly where meaning emerges through figurative, affective, or intersemiotic processes. One early study, for instance, investigated metaphor quality in rule-based AI-generated writing (Littlemore et al., 2018), showing that increased novelty does not necessarily correspond to higher perceived quality. Related work on AI-generated images suggests that current systems struggle to produce abstract or metaphorical visual representations without explicit verbal specification, often requiring “a detailed explanation of the implicit meaning” (Chakrabarty et al., 2023:7378). Taken together, these findings point to the importance of interpretability and structured meaning, rather than surface novelty alone.
Arikan and Aram (2025) examined AI creativity in visual art by evaluating generated images using expert ratings of attributes such as novelty, ambiguity, and emotion. They argue that more advanced computational memory mechanisms can foster creative output, “paralleling human artistic processes that rely heavily on accumulated knowledge and experiences” (Arikan and Aram, 2025: 11). While their analysis remains primarily evaluative rather than semiotic, their findings nevertheless indicate that AI-generated images often introduce ambiguity, affect, and narrative structure beyond what is explicitly specified in prompts.
At the same time, approaches that assess the quality of AI-generated images continue to document persistent failures in compositional and relational accuracy (Hu et al., 2023). Some divergences between prompt and output clearly reflect technical inconsistencies or limitations in semantic grounding. However, evaluation-oriented approaches remain largely silent on how other forms of divergence may also function as meaningful semiotic transformations, in which ideational, interpersonal, or affective meanings are expanded beyond what is explicitly specified in the prompt. A survey of 37 studies on automatic and human evaluations of AI-generated artefacts, for example, shows that many human-based assessments in text-to-image research lack reproducibility (Otani et al., 2023). Extending this critique beyond the visual domain, Rohrmeier (2022) argues that creativity constitutes an AI-complete problem, requiring more than surface-level pattern replication. This perspective supports analytical approaches that treat AI outputs as structured meaning-making artefacts rather than as isolated technical products. Accordingly, Tian et al. (2025: 44) propose that evaluating AI-generated images should involve multiple interacting dimensions, including text-to-image alignment, human-perceived aesthetics, and potential biases inherited from training data. Building on this multidimensional view, further studies show that AI-generated images reproduce culturally anchored visual discourses, with the type of depicted subject playing a central role in shaping which meanings are foregrounded or suppressed (e.g., Putland et al., 2025).
Collectively, this body of research suggests that while AI-generated images are frequently evaluated in terms of creativity, quality, or alignment, far less attention has been paid to how they systematically expand and reconfigure meaning. One notable exception is Ghazvineh (2024), who conducted an intersemiotic analysis of ideational meaning in text-prompted images using Kress and Van Leeuwen’s Grammar of Visual Design (Kress and Van Leeuwen, 2021). Ghazvineh demonstrates how visual representations systematically realise processes, participants, and circumstances specified in textual prompts. The present study builds on this work by extending intersemiotic analysis to meaning expansion beyond minimal prompts, the realisation of projection and affect, and a comparative analysis of subject types within an SFL-informed multimodal framework.
SFL for intersemiotic analysis
SFL conceptualizes language, image, and other semiotic modes as socially shaped resources for meaning-making, each with distinct capacities for organising experience, enacting social relations, and structuring discourse (O’Halloran et al., 2019). Within this semiotic theory, meaning is not transmitted intact across modes but is reconstrued through the affordances and constraints of different representational systems (a process Kress (2003) describes as transduction). In the context of text-to-image generation, this implies that visual outputs should not be treated as direct translations of verbal prompts, but rather as outcomes of intersemiotic processes in which meanings are selectively realised, expanded, or transformed. Prompts shape both semantic interpretation and visual outcomes in text-to-image systems, positioning the prompt as an interface where human intentions, linguistic choices, and model capacities intersect (Valdez et al., 2024).
Logico-semantic relations between verbal and visual modes in text-to-image generation (after Bateman, 2014).
Closely related to logico-semantic relations is the realisation of affect, understood as the semiotic construal of emotion, attitude, and intensity (Martin and White, 2005). In SFL-informed multimodal analysis, affect forms part of the interpersonal metafunction and can be expressed through lexical choices, prosody, gesture, colour, salience, composition, and other non-verbal resources (Kress and Van Leeuwen, 2021). Such affective meanings may also be amplified through forms of graduation, including sharpened focus, intensification, repetition, or isolation (see Martin and White, 2005: 137–141), which in visual texts can be realised through compositional prominence, colour saturation, or reiterated semiotic elements. Affective attributes become particularly salient in text-to-image generation, where they contribute to mood, attitude, or emotional orientation even though minimal prompts typically specify little or no evaluative content (Ghazvineh, 2024).
Previous studies demonstrate that prompts function as mediating artefacts that shape both semantic interpretation and aesthetic outcomes in text-to-image systems (e.g., Liu and Chilton, 2022; Putland et al., 2025). Adopting logico-semantic relations in the present study enables a systematic account of how AI-generated images go beyond prompt realisation to perform intersemiotic meaning expansion. This framework also supports comparison across different subject types, making it possible to examine how the nature of the depicted entity conditions the kinds of expansions that are visually instantiated.
Methods
Corpus design and subject selection
Subject types by animacy and conventional agency.
Person represents a schema-licensed human agent for a wide range of activities. Fish was chosen as an animate entity whose biological affordances are misaligned with many human activities, allowing the influence of environmental schemas (e.g., underwater settings) to be examined. Robot represents an inanimate but culturally agentive entity with partial access to human action schemas. Finally, spoon was chosen as an inanimate, non-agentive object serving as a boundary case. The progression of subjects allows comparison across decreasing degrees of semantic plausibility and conventional agency while keeping prompt structure constant.
Six verbal predicates were selected to span material, mental, verbal, and projective processes: ‘exercising’, ‘reading a book’, ‘giving a present to someone’, ‘desiring something’, ‘talking with someone’, and ‘having a dream’. Each predicate was combined with each subject, yielding 24 unique subject-predicate prompts. All images were generated using OpenAI’s text-to-image system accessed via ChatGPT (OpenAI, 2026). 1 To minimize variation introduced by prompt engineering and to foreground model-internal meaning construction, a strictly controlled two-line prompt format was used for all generations:
‘Generate a drawing (1024 × 1024, 1:1 square ratio) of:
A [subject] [predicate]’
The subjects and predicates were not intended as a representative sample but as analytically contrasting cases across animacy and process type, refined through iterative pilot testing and theoretical comparison. The modality cue ‘drawing’ and a fixed square resolution were used to reduce stylistic variability associated with photorealism and aspect ratio. For each subject-predicate combination, five independent generations were produced, resulting in a total corpus of 120 images (6 predicates × 4 subjects × 5 repetitions). Each generation was conducted in a new, independent session, with no carry-over context and no additional meta-instructions. The order of prompts was randomised across sessions. All images were generated within a contiguous 2-day collection window, and date and time were logged for each image. The corpus is therefore treated as a snapshot of model behaviour rather than a longitudinal sample. The complete corpus of 120 generated images is provided as supplemental material.
Coding
Generated images were analysed using an SFL-informed coding scheme targeting logico-semantic expansion, projection, and affect. Coding was conducted with reference to a detailed manual (see supplemental material), and all decisions were recorded in a structured spreadsheet.
Three expansion relations were coded at the level of visually identifiable items. 2 Modifications of the subject itself that narrow or subtype the prompted participant beyond schematic defaults (e.g., wearable accessories, altered material properties, exaggerated form) were coded as specification. Actor-scoped resources that are functionally coupled to the subject’s activity and staged as part of an individualised action setup were coded as addition. Circumstantial elements that situate the activity in terms of place, environment, or atmosphere and are spatially backgrounded or scene-integrated were coded as enhancement. In adapting logico-semantic relations for multimodal analysis, the present study operationalises specification, addition, and enhancement as analytically distinct expansion strategies, without reproducing the full hierarchical organisation of elaboration and extension proposed for linguistic clause complexes in Halliday and Matthiessen (2014: 613–614).
Each identifiable item within an image was counted individually; however, multiple instances of the same or closely related items were grouped and coded as a single object (for a discussion on identifying visual verbal messages, see Royce, 2002). For example, a generated image that produced many trees in the background was coded as one instance of ‘foliage’ under enhancement. Items constitutive of the depicted activity (once visually realised) were excluded from all counts.
Projection was coded dichotomously (No = 0; Yes = 1), indicating whether representational layering was absent or present. Projection was coded as present where layered meaning was suggested through symbolic overlays, image-within-image structures, or overt markers such as thought or speech bubbles (e.g., Unsworth, 2007). When projection was present, only items belonging to the primary scene layer were included in specification, addition, and enhancement counts. Accordingly, items depicted within thought bubbles, for example in prompts such as ‘having a dream’, were not coded for the purposes of this study.
Affect was graded on an ordinal three-level scale capturing the overall degree of interpersonal loading. While holistic, judgements were anchored in recurring semiotic cues (e.g., symbolic overlays, colour saturation, compositional salience, bodily orientation), which were discussed and exemplified during coder training. Affect was coded independently of expansion and projection. The decision to code only affect from the appraisal system in SFL reflects its role as a basic resource for encoding feelings across settings (cf. Oteiza, 2017: 460).
A second coder with training in multimodal discourse analysis was introduced via a brief calibration phase (the first five images were discussed to align use of the coding manual), after which she provided independent ratings for a randomly selected subset of 40 images from the corpus. Interrater reliability for count-based expansion variables (specification, addition, and enhancement) was assessed using intraclass correlation coefficients (ICC, absolute agreement), yielding ICC = 0.918 for specification, 0.873 for addition, and 0.916 for enhancement. Agreement on the presence of projection (coded dichotomously) was assessed using Cohen’s kappa (κ = 1.000), while affect, coded on an ordinal three-level scale (low, medium, high), was assessed using quadratically weighted Cohen’s kappa (κ = 0.635, n = 40). Overall, these results indicate high agreement between raters, and remaining discrepancies were resolved through discussion and consensus.
All images in the corpus are synthetically generated. AI was used exclusively for image generation; all analytical decisions, coding, and interpretation were carried out by the author and a trained co-rater. For each image, prompt text, interface, timestamp, and session independence were logged to support transparency and reproducibility.
Results
This section presents the results of coding logico-semantic relations and affect across the image corpus. General trends across prompt predicates are examined first, followed by subject-specific patterns. Selected images illustrate recurring patterns rather than exceptional cases.
Overall patterns of logico-semantic relations and affect between prompts
Mean logico-semantic expansion counts by prompt.
Images generated in response to the ‘exercising’ prompt consistently depict an athletic subject engaged in a workout scenario. This typically includes gym-related attire and equipment, prompting the generation of items such as sweatbands, smartwatches, water bottles, or personal music devices (a full list of coded items is provided as supplemental material). Such elements appear either on the subject’s body (specification), in close proximity as part of the activity setup (addition), or spatially backgrounded within the environment (enhancement).
Image 1 illustrates the 20 independent generations for all four subject types under the ‘exercising’ prompt. Across all instances, the human action schema gains primacy, with recurrent gym interiors featuring elements such as dumbbell racks, shelving units, exercise machines, and yoga mats. This high degree of visual elaboration draws on culturally entrenched action schemas that operate across both language and image, supporting the plausibility of the depicted activity (Hart and Marmol Queralto, 2021). Subjects that lack conventional animacy or agency are systematically anthropomorphised in order to conform to the prompted activity, a pattern that recurs throughout the corpus. As we will see in the next section, subject type strongly conditions whether elements are realised through specification or introduced as enhancement. Images generated for the ‘exercising’ prompt.
By contrast, in prompts such as ‘giving a present’, the depicted event corresponds to a material process with a clearly foregrounded Goal (the present itself; cf. Halliday and Matthiessen, 2014). Because the process is semantically complete once the transfer of the object is realised, it corresponds to the absence of addition relations for this prompt. At the same time, gift-giving is culturally associated with festive occasions such as Christmas, which typically take place in winter. As a result, some images (discussed further in Image 6) include specified seasonal clothing such as scarves or wool hats (specification), as well as backgrounded decorative elements such as Christmas ornaments (enhancement). Notably, the model shows a strong preference for coffee as a default additional element in scenes involving talking or reading, independent of subject type. Coffee was coded 34 times across these two activity categories.
With regard to projection, prompt type exerts a strong influence on the appearance of explicit projection in the form of speech or thought bubbles for certain subjects. Image 2 presents spoon as the subject for ‘talking with someone’ (left) and fish as the subject for ‘desiring something’ (right). In the talking prompt, projection is realised both as locution (speech bubbles) and as projection of meaning (thought bubbles), consistent with Halliday and Matthiessen’s (2014: 509) account of projection. Examples of projection realised through speech and thought bubbles.
As for affect, its rated degree varies systematically by prompt predicate, reflecting differences in the degree to which interpersonal meaning is foregrounded in the visual realisations. Prompts associated with routine or backgrounded activities tend to cluster at the lower end of the affect scale, whereas prompts invoking desire, interpersonal exchange, or subjective experience more frequently elicit medium to high affective ratings. An example of the latter is shown in Image 3, which illustrates one generation of the prompt ‘a robot desiring something’. One generation for the prompt ‘a robot desiring something’ to illustrate affect.
In Image 3, affect is realised through a combination of symbolic and compositional resources rather than through embodied facial cues. Most prominently, affect is externalised via symbolic elements, including heart-shaped icons integrated into the scene. In the case of the robot subject, some of these resources also fall under specification, such as when the eyes are rendered as heart shapes, simultaneously elaborating the subject and intensifying interpersonal orientation. Additional affective loading is achieved through colour and lighting, with warm, saturated hues dominating the scene and increasing visual salience. Finally, affect is supported through bodily configuration: the subject is positioned in close proximity to the desired object (see intensity in Martin and White, 2005: 140–141), with a posture that draws on culturally licensed human expressions of wanting, thereby reinforcing the prompted mental process.
By contrast, the prompt ‘reading a book’ is consistently associated with low affect scores across all subject types. Images generated for this predicate typically foreground environmental detail and circumstantial context rather than interpersonal engagement, resulting in subdued affective loading. The ‘exercising’ prompt predominantly elicits medium affect, with occasional instances of high affect. Across subjects, exercising is visually construed as an energetic and goal-directed activity (see Image 1), producing moderate interpersonal salience without strongly intensifying evaluative stance.
Taken together, the results show that prompt predicates systematically condition both the type and degree of logico-semantic expansion and affective loading in AI-generated images. Action-oriented prompts such as ‘exercising’ strongly license specification and addition, producing visually dense realisations grounded in culturally entrenched activity schemas. Projection is tightly coupled to prompts involving subjective experience or mental processes, most notably ‘having a dream’. Affect co-varies with these patterns, remaining backgrounded for routinised activities while increasing for prompts that foreground desire, interpersonal exchange, or internal states. Overall, prompt semantics guide not only what kinds of meanings are expanded, but also how strongly interpersonal meaning is foregrounded.
Overall patterns of logico-semantic relations and affect between subjects
Mean logico-semantic expansion counts by subject type.
Clear contrasts also emerge with respect to projection. The person subject exhibits the lowest mean projection score (0.17), despite being the most semantically plausible participant across all prompts. This suggests that highly schema-licensed human subjects rarely require representational layering to sustain interpretive coherence. At the opposite end of the spectrum, spoon shows the highest degree of projection (0.63), compensating for its lack of animacy and agency by externalising mental and interpersonal relations through layered representations.
The person subject represents the most default and schema-rich category, functioning as a baseline for the activities described in the prompts. It is also the subject type that affords comparatively little variation across repeated image generations. As illustrated in Image 4, the prompt ‘talking with someone’ produces multiple settings and configurational variants for the robot subject (e.g. indoor vs outdoor settings, seated vs standing postures), whereas the corresponding person images largely retain a single, prototypical setting (a café), with only minimal perspectival variation. Despite these differences in setting and configuration, the narrative processes depicted across all images remain stable. This stability is achieved through the use of visual vectors (e.g., hand gestures and eye contact), understood as “the lines created by the visual forms in an image” (Bateman, 2014: 59). Variance within the same prompt (‘talking with someone’) between subject types.
This pattern is also reflected in the relatively low variance observed in counts of specification, addition, and enhancement relations for person subjects across repeated generations. By contrast, prompts situated in more variable environments, such as talking with someone in urban outdoor settings, allow for additional specifying elements, including accessories such as backpacks, which were visible in multiple images.
Image 4 also illustrates the compositional principle of Given and New. In many of the generated images, the prompted subject is positioned on the left side of the image, functioning as Given—that is, as information treated as familiar or taken for granted as the point of departure for the message. By contrast, the right side is reserved for New information, most notably the unspecified participant introduced by the prompt ‘talking with someone’ (Kress and Van Leeuwen, 2021: 186–187). A similar left-right distribution is observed for the prompt ‘giving a present to someone’.
Interestingly, this compositional pattern does not consistently hold for the prompt ‘desiring something’ (see Image 3). In these images, the object of desire frequently appears on the left side of the image, occupying the Given position, while the desiring subject is positioned to the right. From a compositional perspective, placing the desired object in the Given position foregrounds it as the initial semiotic departure point of the image, while the desiring subject is introduced as New, visually construing a mental orientation toward an already salient element. This inversion aligns with Kress and Van Leeuwen’s (2021: 187) observation that Given-New structures are not fixed mappings of grammatical roles, but are sensitive to what an image treats as its primary point of semiotic departure.
While sharing animacy with the person subject, the fish subject presents a different case in that it is associated with a distinct embodied schema, namely that of an underwater environment. This schema frequently clashes with prompts designed around human activities. As shown in Image 5, for the prompts ‘having a dream’ (left) and ‘desiring something’ (right), the background consistently shifts to an underwater setting, resulting in the inclusion of corals and other sea creatures, which were coded as enhancement. In addition, the object of desire is reconfigured to align with the fish schema: instead of sweets displayed in a shop window, the desired item appears as a donut attached to a fishing line, resembling bait. Notably, the bait remains semantically aligned with the prompted activity through its association with practices of attraction and pursuit specified by the prompt ‘desiring something’. Schema-driven reconfiguration for fish subjects under ‘having a dream’ and ‘desiring something’.
Another subject-induced difference concerns the use of projection. As discussed earlier, subject plausibility strongly influences the likelihood of explicit projection, with spoon exhibiting the highest mean projection count. This pattern is illustrated in Image 6, which compares person, robot, and spoon as subjects for the prompt ‘giving a present to someone’. The spoon images rely heavily on projection, which is typically realised through thought or speech bubbles, to establish a relationship between the giver and receiver. In contrast, the person and robot conditions predominantly use non-verbal communicative cues, such as body orientation, gaze, or spatial proximity, to construe the interpersonal relationship. Comparison of ‘giving a present to someone’ across person, robot, and spoon subjects.
In several of the person and robot subjects in Image 6, these relations are additionally structured through transactional vectors that organise the act of giving around a visually foregrounded Goal (Kress and Van Leeuwen, 2021: 59–60). The giver frequently functions as Actor through bodily orientation and gesture vectors directed toward the present, while the recipient’s gaze is often oriented toward the transferred object rather than reciprocally toward the giver. This creates an asymmetrical interactional configuration in which the present itself becomes the central interpersonal focus of the scene. By contrast, the spoon subject relies less on embodied vector structures and instead externalises interpersonal relations through projection resources such as speech or thought bubbles.
Subject plausibility also shapes the degree of creative abstraction within the images. For the spoon subject in Image 6, the recipient of the present is frequently transformed into another object (e.g. a fork, a coffee cup, or a small figurine), whereas the recipient remains a human-like agent in the person and robot conditions. A closer examination further shows that gift-giving scenes differ systematically in how social relations are visually configured across subject types. In the person condition, the depicted interaction is frequently realised through configurations associated with romantic exchange, whereas the robot and spoon conditions more often rely on additional contextual cues, such as Christmas or other festive settings, to visually scaffold the act of giving presents. These cues, in turn, correspond to differences in enhancement (e.g., decorations, lights) and specification (e.g., winter clothing consistent with a Christmas schema).
Another contrast concerns how activity plausibility is stabilised across subject types. Image 7 presents a comparison of the prompt ‘reading a book’ across the four subject types examined in this study. For the person subject, the activity is primarily realised through the book itself, but is further stabilised through environmental and functional resources, including additional books placed nearby and the specification of accessories such as glasses. These elements function as schematic cues linking the depicted action to entrenched cultural expectations of reading. The recurrent appearance of such items constitutes a common form of expansion relation for this prompt. Comparison of ‘reading a book’ across subjects.
By contrast, the robot subject is situated within an environment densely populated with reading-related resources, including books, a desk lamp, and other study-associated artefacts. Here, plausibility is supported less through subject-internal specification than through extensive environmental elaboration, indicating a shift toward enhancement-driven stabilisation of the activity. Interestingly, the robot images generated for this predicate also tend to realise slightly more detailed eye configurations than in several of the other prompt categories, allowing gaze direction toward the book to become more visually explicit. This suggests that the visual construal of reading places greater pressure on the explicit realisation of attentional orientation, even for subjects with otherwise minimalist facial features.
The fish subject illustrates a different configuration: while the subject undergoes specification through accessories such as glasses, it remains embedded within its schema-bound underwater environment. This combination indicates an attempt to reconcile the prompted activity with animacy while maintaining environmental coherence. In this case, specification operates alongside enhancement rather than fully overriding the subject’s default schema.
Finally, the spoon subject, which lacks both animacy and agency, undergoes pronounced anthropomorphism. Despite the spoon’s conventional schema as a kitchen utensil, the human activity schema associated with reading dominates, resulting in a visual configuration that closely resembles the person-based realisation.
Taken together, this comparison shows that specification can function independently of anthropomorphism as a means of reinforcing action plausibility. Accessories such as glasses do not primarily individualise the subject, but instead anchor the activity within a familiar cultural schema. At the same time, the relative balance between specification, enhancement, and anthropomorphism shifts as subject plausibility decreases, revealing a progression in which activity schemas increasingly dominate when subject-based grounding becomes tenuous.
Distribution of affect counts by subject type.
The person subject tends to cluster around low (19) to medium affect (7) across most prompt predicates. This reflects the role of the human subject as a highly schema-licensed baseline, where interpersonal meaning can be construed through familiar, routinised configurations without requiring strong affective intensification. Even for prompts involving desire or interpersonal exchange, affect for the person subject rarely reaches the highest level.
The fish subject exhibits a mixed affective profile. For prompts that are weakly interpersonal, such as ‘reading a book’, affect remains consistently low. However, for prompts involving internal states or projected motivation, such as ‘desiring something’, affect more frequently rises to medium or high, indicating that affective intensification supports the stabilisation of mental or interpersonal processes when schematic alignment is strained.
The robot subject shows a broader distribution, with a noticeable shift toward medium and high affect for prompts involving desire or gift-giving. As an inanimate but culturally agentive entity, the robot frequently relies on intensified interpersonal loading to support the construal of motivation, evaluation, or relational orientation. However, the robot’s minimalist eye design constrains more exaggerated facial expression and therefore shifts affective realisation toward symbolic overlays and specification (see Image 3) rather than highly articulated facial cues. Affect functions as a supplementary resource, supporting the interpretation of interpersonal meaning.
Finally, the spoon subject displays the highest relative incidence of high affect (13) across several prompt categories. As a non-animate, non-agentive object, the spoon often requires strong interpersonal loading to support activities involving intention, desire, or social relations. Affect here operates alongside projection as a key semiotic strategy, enabling the construal of interpersonal meaning that cannot be grounded in embodied action or conventional agency.
In summary, subject type emerges as a key conditioning factor in how logico-semantic relations and affect are mobilised. Schema-rich and biologically plausible subjects constrain expansion and rely less on projection or affective intensification, producing relatively stable and recurrent realisations within the generated corpus. As subject plausibility and conventional agency decrease, expansion strategies shift: animate but schematically misaligned subjects preferentially reorganise the environment through enhancement, while non-agentive subjects increasingly rely on projection and heightened affect to support the construal of interpersonal and mental meaning. These patterns show how different forms of expansion and affect function as compensatory semiotic resources, enabling AI-generated images to maintain interpretive coherence under varying degrees of schematic constraint.
Discussion
This study set out to examine how AI-generated images expand and reconfigure meaning beyond minimal verbal prompts through intersemiotic logico-semantic relations and affect, and how these processes are conditioned by subject type. The findings demonstrate that such expansions are neither random nor merely the result of technical misalignment, but instead reflect systematic semiotic strategies through which visual meaning is constructed from underspecified linguistic input.
Across prompt predicates, the distribution of logico-semantic relations shows that different activity types invite distinct modes of visual elaboration. Action-oriented prompts such as ‘exercising’ strongly license specification and addition, resulting in visually dense scenes populated by culturally entrenched activity cues. By contrast, prompts such as ‘reading a book’ favour enhancement through circumstantial detail, stabilising relatively static processes via environment rather than interpersonal engagement. These findings extend prior intersemiotic work showing that visual representations do not merely translate linguistic meaning, but selectively reorganise it in line with the affordances of the visual mode (cf. Ghazvineh, 2024). At the same time, they support broader accounts of language-image relations that emphasise the role of shared action schemas and embodied knowledge in cross-modal meaning-making (Hart and Marmol Queralto, 2021). The variability associated with prompt structure observed here further supports Zhou et al.’s (2026) characterisation of prompting as a contextual rather than directive mechanism, consistent with SFL accounts of meaning as emerging from stratified contextual choices rather than deterministic instruction (Halliday and Matthiessen, 2014).
Crucially, the analysis shows that subject type plays a central role in shaping how these expansions are realised. Schema-rich and biologically plausible subjects, most notably person, constrain variation and require relatively little compensatory elaboration. By contrast, subjects with reduced or ambiguous agency trigger systematic shifts in expansion strategy. For fish, schematic misalignment is primarily resolved through enhancement, with the environment being reconfigured to maintain biological plausibility. For spoon, which lacks both animacy and agency, expansion increasingly relies on specification and projection to externalise mental and interpersonal relations that cannot be grounded in embodied action. These patterns resonate with recent discussions of AI-generated imagery as reproducing culturally sedimented visual discourses rather than generating unconstrained novelty (Putland et al., 2025; Rohrmeier, 2022; Tao et al., 2024).
The analysis of affect further refines this picture by demonstrating that interpersonal meaning is not evenly distributed across prompts or subjects. Rather, affective loading functions as a complementary semiotic resource that co-varies with logico-semantic expansion (cf. Martin and White, 2005: 35). Importantly, affect is frequently realised through symbolic and compositional resources rather than through embodied action alone, particularly for non-human subjects. This supports recent observations that AI-generated images routinely introduce evaluative and emotional orientation beyond what is specified in the prompt, even in the absence of explicit affective cues (Chakrabarty et al., 2023: 7374).
Taken together, these patterns complicate dominant technical accounts of hallucination in text-to-image generation, which tend to frame divergence from prompt content as error or failure (cf. Chen et al., 2026). From a semiotic perspective, many such divergences can instead be understood as motivated expansions that support plausibility, coherence, and interpretability under conditions of underspecification (Cope and Kalantzis, 2024). In this sense, specification, enhancement, projection, and affect operate as structured meaning-making strategies rather than as symptoms of misalignment. This interpretation aligns with broader critiques of purely metric-based evaluation frameworks in AI research, which often overlook how meaning is redistributed across modes.
The recurrent patterns observed across prompts and subjects also foreground the role of cultural bias in AI-generated imagery. Preferences for specific settings, objects, and interpersonal configurations, such as café scenes, coffee as a default beverage, or particular representations of desire, reflect latent cultural assumptions embedded in training data, echoing recent findings on cultural alignment and bias in large language models (Tao et al., 2024). Gender representation also emerged as a recurring pattern across several prompt categories. In the ‘giving a present to someone’ condition, for example, recipients in the person condition were consistently realised as young female figures, often positioned through asymmetrical interpersonal configurations in which visual attention was directed toward the gift itself. Similar tendencies were also visible in prompts involving desire, where interpersonal orientation and desired objects were frequently realised through gendered visual conventions associated with romance, sweetness, or gendered affective imagery. Such patterns suggest that generative systems may reproduce culturally sedimented assumptions concerning gender, affect, and interpersonal exchange, even when prompts remain minimally specified. Viewed through an SFL-informed lens, these biases are not incidental, but become visible through the systematic deployment of logico-semantic, compositional, and affective resources (see Kress, 2003).
From a broader semiotic perspective, the patterns identified in this study support the view that AI-generated images operate as fully multimodal texts rather than as secondary realisations of linguistic input (cf. Zhou et al., 2026). In systemic functional terms, the visual mode actively reinstantiates ideational, interpersonal, and textual meanings using resources that are not isomorphic with language, but functionally equivalent (Halliday and Matthiessen, 2014; O’Halloran et al., 2019). The systematic deployment of expansion, projection, and affect indicates that image generation involves metafunctional choices analogous to those described for language (Kress and Van Leeuwen, 2021). In this respect, the present findings can be mapped onto earlier work on intersemiotic relations (e.g., Lindenberg, 2025; Martinec and Salway, 2005), showing that generative systems do not merely translate meanings across modes, but actively reconstrue them in ways that reflect culturally stabilised semiotic conventions.
Importantly, the observed regularities across prompts and subjects suggest that AI-generated images make visible otherwise tacit principles of meaning-making. As Halliday and Matthiessen (2014) note, meaning is organised through patterned choices within semiotic systems; the consistency with which particular expansion strategies recur here indicates that such choices are not arbitrary, but draw on entrenched semiotic resources. In this sense, AI-generated images can be understood as revealing how contemporary visual culture construes action, interaction, and mental processes, rather than as idiosyncratic artefacts of model behaviour. This aligns with Ghazvineh’s (2024) argument that AI-mediated intersemiosis offers a productive site for examining how meanings are redistributed across modes under conditions of constrained input. Interpreted in this way, the systematic expansion patterns observed in the corpus point to more than model-internal strategies for resolving underspecification. They foreground how AI-generated imagery participates in the circulation and reinforcement of culturally dominant ways of visualising agency, desire, interaction, and plausibility.
Conclusion
This study has examined how AI-generated images expand and reconfigure meaning beyond minimal verbal prompts through intersemiotic logico-semantic relations and affect, and how these processes are conditioned by subject type. Drawing on an SFL-informed multimodal framework, the analysis showed that text-to-image generation systematically mobilises specification, addition, enhancement, and projection to construct visually coherent representations from highly underspecified input. These expansions are not uniform or arbitrary but vary predictably across both prompt predicates and depicted entities. Moreover, the findings suggest that multimodal analyses of AI-generated imagery can contribute to ongoing discussions of cultural and gender bias by revealing how interpersonal meanings and social assumptions become visually stabilised through recurrent semiotic patterns.
By comparing subjects with differing degrees of animacy and agency, the study demonstrated that subject plausibility plays a key role in shaping the form and distribution of visual expansion. Schema-rich human subjects tend to constrain variation, while less plausible or non-agentive subjects prompt compensatory strategies such as environmental reconfiguration or explicit projection. These findings challenge deficit-based accounts of AI image generation that frame divergence from prompts as error, instead highlighting expansion as a form of structured semiotic problem-solving. The framework may also prove useful in educational contexts by supporting critical multimodal literacy approaches to AI-generated imagery and prompting practices.
More broadly, the findings suggest that close analysis of AI-generated images can offer insight into how semantic structures and mental processes are construed visually. Rather than merely reflecting surface prompt content, generated images systematically reorganise ideational, interpersonal, and affective meanings in ways that foreground culturally stabilised schemas, experiential salience, and perceptual anchoring. From this perspective, AI-generated imagery functions as a site where latent assumptions about action, perception, desire, and interaction become visible through compositional and representational choices (for example, by visually marking Given and New relations). This study showed that generative systems not only exploit such affordances, but in doing so make explicit the semiotic principles through which meaning is routinely construed across modes.
Several limitations of the present study should be acknowledged. First, while the distinction between addition and enhancement was operationalised through explicit coding criteria, borderline cases and identification issues inevitably arise, reflecting the gradient nature of circumstantial and actor-scoped meaning in multimodal texts (see Bateman, 2014: 197). Second, the coding of affect relied on a three-level ordinal scale and holistic judgement, which, although theoretically motivated, necessarily involves a degree of interpretive subjectivity. Third, the analysis was based on a relatively small, purpose-built corpus coded by the author and a single trained co-rater, limiting the statistical generalisability of the findings. Finally, the study examined outputs from a single text-to-image model, and the observed patterns may therefore reflect model-specific training data rather than universal properties of text-to-image generation.
Notwithstanding these qualifications, this work contributes a methodological and theoretical framework for analysing AI-generated images as multimodal meaning-making artefacts. By extending logico-semantic relations beyond language to intersemiotic contexts, the study offers tools for understanding how generative systems negotiate plausibility, bias, and creativity. Future research can build on this approach by incorporating more fine-grained analyses of affect, expanding the range of subject types and activities, and systematically comparing expansion patterns across different generative models and prompting conditions.
Supplemental material
Supplemental material - How a text-to-image model adds meaning beyond minimal prompts: An SFL-informed intersemiotic analysis of visual-verbal relations
Supplemental material for How a text-to-image model adds meaning beyond minimal prompts: An SFL-informed intersemiotic analysis of visual-verbal relations by Dennis Lindenberg in Multimodality & Society
Supplemental material
Supplemental material - How a text-to-image model adds meaning beyond minimal prompts: An SFL-informed intersemiotic analysis of visual-verbal relations
Supplemental material for How a text-to-image model adds meaning beyond minimal prompts: An SFL-informed intersemiotic analysis of visual-verbal relations by Dennis Lindenberg in Multimodality & Society
Supplemental material
Supplemental material - How a text-to-image model adds meaning beyond minimal prompts: An SFL-informed intersemiotic analysis of visual-verbal relations
Supplemental material for How a text-to-image model adds meaning beyond minimal prompts: An SFL-informed intersemiotic analysis of visual-verbal relations by Dennis Lindenberg in Multimodality & Society
Footnotes
Acknowledgments
The author would like to thank the anonymous co-rater for assistance with the multimodal coding and reliability assessment.
Ethical considerations
Ethical approval and informed consent were not required for this study, as it is based exclusively on the analysis of synthetically generated images and does not involve human participants or personal data.
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The author declares no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
The data supporting the findings of this study are provided as supplementary materials. Supplementary File A contains the full coding manual used for the intersemiotic analysis, and Supplementary File B contains the spreadsheet documenting all coded items for each image across categories and subject types. Supplementary File C contains the full corpus of generated images, organised by prompt category.
Supplemental material
Supplemental material for this article is available online.
Notes
Author biography
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
