Abstract
Visual working memory (VWM) operates within structured environments, yet it remains debated whether representations are guided primarily by the global scene configuration (gist) or the intrinsic features of objects. This question extends to cognitive ageing, where declines in VWM co-occur with a relative preservation of global information over fine-grained visual details. To disentangle these influences, younger and older adults detected changes to an object’s identity, location, or both, while its semantic consistency with the scene was manipulated. Unlike our previous work, where location changes disrupted the spatial layout, here we preserved the layout by swapping the critical object, thereby isolating memory for object-location binding from sensitivity to global disruptions. Across age groups, conjunctive changes were detected more accurately than single-feature changes. Crucially, detecting location changes was significantly more difficult when the layout was preserved (swaps) than when it was disrupted (displacements). This demonstrates that while layout disruptions provide salient cues, recalling an object’s location from a stable configuration requires retrieving its intrinsic features. This is further supported by a detection advantage for semantically inconsistent objects (e.g., a torch in a bathroom), which was observed specifically in younger adults, suggesting an age-related decline in the strategic use of contextual violations. Eye-tracking during retrieval revealed longer fixations for inconsistent objects, indicating increased effort in integrating them. In contrast, consistent objects were fixated faster, a novel retrieval-phase finding likely driven by memory from their initial encoding as inconsistent. While core attentional mechanisms were preserved with age, our results argue that when global structure remains unaltered, VWM is guided by hierarchical representations of object features, in which semantic meaning plays a central, organising role.
Keywords
Introduction
Encoding, retaining, and binding visual information about the identity (“what”) and location (“where”) of objects is crucial for recognition and recall in most daily tasks. However, how visual objects are integrated into scene representations and retrieved from memory remains debated (Bays et al., 2024; Brady et al., 2011; Ngiam, 2024, for reviews). Discussions about the format of object representation in visual working memory (VWM) primarily contrast two views: one where objects are assumed to be stored in memory as integrated single units (Green & Quilty-Dunn, 2021; Kahneman et al., 1992; Luck & Vogel, 1997; Vogel et al., 2001), and the other postulating a hierarchical representation of independent bundles of features (Brady et al., 2013; Fougnie & Alvarez, 2011; Fougnie et al., 2010; Wang et al., 2017), linked through focused attention (A. Treisman, 1998; A. M. Treisman & Gelade, 1980; Wheeler & Treisman, 2002), where mnemonic precision is dynamically distributed across them. Beyond its intrinsic features, an object’s spatial position is a key component of its memory representation (Golomb et al., 2014). Memory performance typically declines when an object’s location changes between encoding and retrieval, even when location information is task-irrelevant (A. Treisman & Zhang, 2006; Hollingworth, 2007; Olson & Marshuetz, 2005). Likewise, changes to an object’s featural properties can disrupt memory for its original position (Toh et al., 2020), suggesting an inherent mechanism for object-location binding (Postma et al., 2008). This extends to real-world scenes where shifts in object placement attract attention and are rapidly detected (Ryan & Cohen, 2004; Võ et al., 2010), suggesting that the visual system continuously monitors the spatial coherence between incoming input and its short-term memory representation.
In complex naturalistic environments, contextual regularities are also fundamental for forming, accessing, and predicting object representations from VWM (Bar, 2004; Bays et al., 2024; Castelhano & Williams, 2021; Hu & Jacobs, 2021; Kaiser et al., 2015, 2019; Võ, 2021). Objects violating semantic expectations attract early attention (e.g., a torch vs. toothpaste in a bathroom; Allegretti et al., 2025a; Borges et al., 2020; Castelhano & Heaven, 2011; Coco et al., 2020; Gordon, 2004; Loftus & Mackworth, 1978; Malcolm & Henderson, 2010; Spotorno et al., 2014; Underwood & Foulsham, 2006; Underwood et al., 2007, 2008; Võ & Henderson, 2009) and can enhance memory recall (Allegretti et al., 2025b; Biederman et al., 1982; Evans & Wolfe, 2022; Friedman, 1979; Hollingworth & Henderson, 2000; Pezdek et al., 1989). Contextual expectations also constrain where objects are likely to appear (Castelhano & Krzyś, 2020), reflecting spatial predictions that can operate independently of semantic content (Castelhano & Heaven, 2011; Castelhano & Henderson, 2007). Prior knowledge about object-location facilitates scene segmentation (surface guidance framework; Pereira & Castelhano, 2019) and speeds the detection of changes in contextually relevant regions (Rensink et al., 1997).
Spatial information in VWM is not only limited to the position of individual objects but also encompasses the spatial arrangement of multiple items—how objects are positioned and aligned relative to one another (Brady & Alvarez, 2011). This “structural gist” (Vidal et al., 2005) guides access to individual object features (Bar & Ullman, 1996; Blalock & Clegg, 2010; Boduroglu & Shah, 2009; Mou et al., 2008; Williams et al., 2013; but see Woodman et al., 2012 for counterevidence), and its disruption impairs change detection for object colour (Jiang et al., 2000) or location (Mou et al., 2008). In real-world scenes, this structure conveys global layout properties, such as openness, depth, and expansion (Oliva & Torralba, 2001, 2006), which enable VWM to maintain a coherent spatial ensemble independently of detailed object information. Conceptually, this view aligns with visual index theories (e.g., Kahneman et al., 1992; Pylyshyn, 1989), which propose that memory representations are structured along location-based pointers rather than detailed representations of individual objects.
Thus, VWM representations can be understood through two complementary perspectives. On one hand, memory is organised around high-level features of individual objects (identity, position, and semantic consistency) and their contextual co-occurrences (e.g., a fork, knife, and plate on a kitchen table; Biederman, 1981; Fei-Fei et al., 2007; Friedman, 1979; Võ et al., 2019). On the other hand, it relies on global scene layout—the broad spatial structure and low-level geometric properties (Greene & Oliva, 2010; Oliva & Schyns, 2000) that enable rapid gist extraction (Greene & Oliva, 2009; Oliva & Torralba, 2001). This global information drives expectations (e.g., categorisation) even before individual objects are recognised. While object-based and gist-based information likely operate in parallel (Joubert et al., 2007; Kim & Biederman, 2011), their relative weight in VWM representations remains unclear; a primary gap the current study aims to address.
Regardless of how object or scene information shapes memory representations, their source is attentional (Bahle et al., 2018; Chen, 2012; Chun & Turk-Browne, 2007; Cowan et al., 2024; Evans et al., 2011; Kravitz & Behrmann, 2011). Rensink’s (2000) coherence theory proposes that visual information requires attention for memory persistence, avoiding “inattentional amnesia” (Wolfe, 1999). Findings from the flicker paradigm support this: detecting changes in object colour, location, or presence coincides with attention (Rensink et al., 1997). Overt attention is necessary to feed and retain visual memory representations with rich and detailed information about the identity and location of individual objects over multiple fixations (Hollingworth, 2006), even when the objects are no longer in view (Hollingworth & Henderson, 2002). In practice, eye movements can reflect memory-guided behaviour (Ryan et al., 2007; see Hannula et al., 2010; Ryan & Shen, 2020 for reviews), such as indexing the implicit detection of changes even when participants cannot explicitly report them (Hayhoe et al., 1998; Henderson & Hollingworth, 2003; Ryan & Cohen, 2004; Ryan et al., 2000). Such implicit knowledge is also evidenced by global differences in viewing patterns between changed and unchanged scenes (Ramey et al., 2022). Despite broad agreement on the role of attention in memory, the relationship between its explicit deployment and object-based versus gist-based guidance remains largely unresolved.
Understanding this balance is particularly relevant in cognitive ageing, where VWM declines are well documented (e.g., Brockmole & Logie, 2013), but the specific components of object representations most affected remain unclear. Older adults often show greater difficulty than younger adults with maintaining the spatial location of objects compared to their identities (e.g., Chalfonte & Johnson, 1996; Tran et al., 2021) or the binding of these features (e.g., Kessels et al., 2007; Mitchell et al., 2000; Muffato et al., 2019). However, these impairments are often linked to general age-related reductions in VWM capacity (Olson et al., 2004; Pertzov et al., 2015; Read et al., 2016; Thomas et al., 2012), raising the question of whether ageing truly impairs object-based representations or simply reflects broader memory capacity constraints. At the same time, older adults often perform markedly better when information is embedded in meaningful contexts, where long-term semantic knowledge and real-world statistical regularities can scaffold memory performance (Allegretti et al., 2025b; Coco et al., 2022; D’Innocenzo et al., 2022; Mäntylä & Bäckman, 1992; Prull, 2015; Ramzaoui et al., 2021, 2022; Umanath & Marsh, 2014). Converging evidence suggests that, in later adulthood, global or gist-based information is retained more effectively than fine-grained details (Greene & Naveh-Benjamin, 2023; Grilli & Sheldon, 2022), indicating a possible shift towards greater reliance on gist-based guidance. Yet, this line of research has focused almost exclusively on conceptual aspects of episodic information, leaving open the question of whether preserved gist-based representations might also extend to the spatial organisation of visual information.
The present study addresses whether successful access to object information in VWM depends primarily on global spatial structure (gist-based) or on high-level object properties (object-based). We specifically focus on structural gist, defined as the spatial layout and geometric organisation of objects, which is distinct from conceptual gist (Oliva, 2005; Oliva & Torralba, 2006). We build on our previous eye-tracking study (D’Innocenzo et al., 2022), where identity and/or location changes were identified within naturalistic scenes. Unlike our previous work, where location changes disrupted the spatial structure by leaving an empty space, the current study swapped the position of the critical object with another object. This preserved the overall integrity of the scene’s spatial layout.
This manipulation enabled us to test whether an observed advantage for location changes unequivocally reflects memory for object-specific information rather than mere sensitivity to layout disruptions. We tested two competing predictions. Firstly, if short-term access is guided mainly by the spatial layout, then identity and location changes should be detected at similar levels when the global configuration (i.e., “gist representation”) remains unchanged (i.e., through object swaps). This is because the location-based indices remain identical before and after the change. Conversely, if access depends on specific features and their hierarchical organisation, then performance should vary according to the feature involved, even when the global layout is preserved. In this case, semantically inconsistent objects should significantly enhance detection (the “inconsistency advantage”), as evidenced by attentional strategies such as shorter latencies and longer gaze durations (Henderson & Castelhano, 2005; Johansson & Johansson, 2014). Furthermore, we predicted that identity changes would elicit longer gaze durations than location changes, reflecting the cognitive effort required to integrate a new object and its features into an existing VWM representation.
We tested these predictions in younger and older adults to determine which mechanisms remain resilient to ageing. If older adults rely more on global spatial cues, they should be disproportionately impaired at detecting location changes when the global configuration is preserved, relative to conditions in which it is disrupted (as in D’Innocenzo et al., 2022).
Methods
Participants
Twenty-seven younger adults (21 female; age = 24.6 ± 3.1; education = 16.89 ± 2.24 years) and 24 healthy older adults (14 female; age = 69.7 ± 5.88; education = 13.88 ± 3.77 years) participated in the study. Data from one additional younger and six older participants were collected but excluded from the analyses due to machine error (one younger, four older), performance at chance level under a binomial test (one older), or withdrawal during the experiment (one older). An a priori power analysis (G*Power 3.1, Faul et al., 2007; ANOVA: within-between interaction) for our 2 (Group) × 3 (Change Type) × 2 (Consistency) design indicated a required sample of 38 participants to reach a medium effect size (f = 0.25; α = .05; power = 0.99; repeated-measures correlation = .50; ε = 1). Post-hoc power analyses using the simulation method by Kumle et al. (2018) showed that our sample size (N = 51) provided sufficient power to detect the effects of our variables of interest. The full output of the simulations, along with a detailed explanation of this analysis, is reported in Supplemental Material S1. The younger individuals were Italian students from Sapienza University of Rome and participated as volunteers. Older individuals were recruited from a senior recreation centre in Rome, through flyers and word of mouth, and compensated with a €10 book voucher. Eligibility criteria were an age of 18 to 35 (younger group) or over 60 (older group), normal or corrected-to-normal vision, no history of eye surgery or visual impairments (e.g., maculopathy, cataracts, glaucoma, etc.), and no history of neurological or psychiatric disorders. Older participants were assessed on their general cognitive abilities using the Montreal Cognitive Assessment (MoCA; mean score = 25.28 ± 2.04, range = 23–30), a brief standardised neuropsychological screening test for cognitive impairment (Nasreddine et al., 2005). A cut-off score of ⩽17 was used for exclusion, as indicated by Bosco et al. (2017), who reported lower normative MoCA scores in the Italian population. The study was approved by the Ethics Committee for Transdisciplinary Research of Sapienza, University of Rome (CERT_4_186B65F32D9) prior to commencing data collection, and all participants provided written informed consent before participating in the experimental session.
Design
The change detection task design was borrowed from previous studies (Coco et al., 2020, 2022; D’Innocenzo et al., 2022) with participants asked to detect whether an object, either consistent or inconsistent with its embedding scene, changed in its identity (became another object, e.g., a book in a bedroom became a bunch of bananas), location (moved to another position, e.g., the book shifted from the left to the right side), or concurrently both these features (see Figure 1). Changes in object identity always involved a reversal of the object’s semantic consistency between encoding and recall; that is, objects that were consistent at encoding became inconsistent at recall, and vice versa. Unlike our previous studies, changes in object location were implemented by swapping the position of the critical object with another object (the swap object), which was always semantically consistent with the scene context, rather than simply displacing it to a new location. This manipulation prevented empty spaces at the original object location between the encoding and recall phases, thereby forcing participants to rely on object-based memory representations rather than identifying disruptions in the spatial layout of the scene to detect changes.

Illustration of the types of changes implemented across the encoding (top) and recall (bottom) phases. Three types of changes were introduced: identity (the critical object was replaced by another object, either consistent or inconsistent with the scene context; e.g., a book replaced by a bunch of bananas), location (the critical object swapped its position with another object; e.g., a book with a pair of socks), and both (the critical object changed identity and swapped location with another object). The critical and swap objects are highlighted with red and yellow boxes, respectively, for illustration purposes only; these were not visible to participants during the task. The red box also defines the area of interest used to compute the eye-movement measures reported in this study. The image resolution of this example corresponds to that of the stimuli as presented during the experiment.
Stimuli
The stimulus set consisted of 1,089 photographic images of indoor scenes (e.g., bathrooms, bedrooms, and kitchens). These included: (a) 120 experimental (change) items sourced from the VISIONS database (Allegretti et al., 2025a), each used in 8 versions (N = 120 × 8 = 960) where a critical object could be either semantically consistent or inconsistent with the scene, positioned on the left or right side and swapped with another object; (b) 120 filler (no-change) items, balanced for object spatial position and semantic consistency; and (c) 6 practice items (3 change in 2 versions, and 3 no-change, N = 9). Comparability in the low-level visual properties of target and swap objects was established using the Adaptive Whitening Saliency maps provided in the VISIONS database (https://osf.io/pbf4r/). For each of the 960 object pairs, maximum, mean, and summed pixelwise salience values were extracted, together with bounding-box pixel area as an estimate of object size. Paired-samples t-tests indicated no significant differences between the two object types on any measure (all p > .05), confirming that they were well matched.
Apparatus
Eye movements were recorded binocularly using a Gazepoint GP3 HD eye-tracker (Gazepoint Research Inc., Vancouver, BC, Canada) at a sampling rate of 150 Hz. Participants were seated approximately 60 cm from the screen with their heads stabilised on a chinrest. For the younger group, scenes were displayed at 1,920 × 1,080 pixels on a Dell 24-inch touch monitor with a 60-Hz refresh rate (53 cm wide × 29 cm high). For the older group, scenes were presented at a resolution of 1,280 × 1,024 pixels on a 19-inch LG LCD monitor with a refresh rate of 60 Hz (37.8 cm width × 30.3 cm height 1 ). The experiment was implemented using OpenSesame 3.3.14 (Mathôt et al., 2012) and the PyGaze plug-in (Dalmaijer et al., 2014), which enabled the acquisition of eye-tracking data. Raw gaze data were processed into fixations and saccades using a two-means clustering algorithm, which is well-suited for low-frequency eye-tracking data (Hessels et al., 2017). This algorithm was run in MATLAB R2024a(The MathWorks, Inc., Natick, MA, USA).
Experimental Procedure
At the beginning of each experimental session, the eye-tracker was calibrated on nine points, reporting a final visual angle deviation error for the average of both eyes (mean ± standard deviation) of 0.39° ± 0.10° on the x-axis and 0.55° ± 0.18° on the y-axis for the younger group, while for the older group it was 0.53° ± 0.37° on the x-axis and 0.63° ± 0.30° on the y-axis. In each experimental trial, participants viewed a full-screen scene and were instructed to study it (i.e., encoding phase). As in D’Innocenzo et al. (2022), the presentation time of each scene during the encoding phase was controlled by a gaze-contingent mechanism, ensuring that participants fixated on the critical object in most trials. If participants’ gaze entered a predefined area of interest around the critical object, the image remained on screen for an additional 2 s before a 900-ms retention interval with a black central fixation cross was displayed. This 2-s delay prevented participants from potentially associating their last fixation with the changed object and additionally ensured a constant interval between that fixation and the detection phase, minimising variability in the amount of visual information acquired about the critical object before the retention interval. If participants did not fixate on the critical object during the encoding phase within 10 s, the retention interval was triggered nevertheless. 2 Then, the encoding scene appeared again with (or without) one of the three changes described above (refer to the section ‘Design’), and participants were given up to 10 s to indicate whether a change had occurred in the scene by pressing either the s key for change or the n key for no change on the keyboard. The next trial began immediately after a response was recorded. If participants failed to respond within the time limit, a null response was logged, and the next trial began. Each participant completed 6 practice trials, followed by 120 experimental trials and 120 filler trials, all presented in random order. A Latin Square Rotation was used to counterbalance and distribute the experimental conditions across 24 randomisation lists. The task was explained using written instructions and lasted approximately 45 min. After completing the main task, participants performed a brief computerised Corsi block-tapping test approximately 3 min; Fischer, 2001) using PsyToolkit (Stoet, 2017) to confirm that their visuo-spatial working memory was intact and thus unlikely to have influenced their performance in the change detection task. In this task, nine squares appeared on the screen and lit up in a specific sequence. After hearing the word “go,” participants reproduced the sequence by clicking the squares in the same order (i.e., forward span). The sequence length increased progressively, with up to two attempts per level. The test ended when participants failed to reproduce two consecutive sequences correctly (mean younger = 5.56 ± 1.19; mean older = 4.87 ± 1.3, t[49] = −1.957, p = .06). Eye movements were not recorded during this task. Finally, older participants also completed the MoCA test at the start of the experimental session, which lasted approximately 1 hr.
Data Analyses
Data Preprocessing
Analyses focused exclusively on change trials 3 and included manual responses and eye movements collected during the recall phase. Out of a total of 6,120 change trials (i.e., 51 participants × 120 scenes), 48 trials were excluded due to timeouts (younger = 5, older = 43; 0.78%), and 178 trials as they had a response time faster than 1% and slower than 99% of all trials (younger = 101, older = 77; 2.91%), as computed independently for each participant. For the analysis of manual responses, this resulted in 3,134 trials for the younger group (by-participant average of 116 ± 1.59) and 2,760 for the older group (by-participant average of 115 ± 2.23; 4,346 correct trials in total). For the eye-movement analysis, out of the 5,894 remaining trials, we further excluded 201 trials due to machine error (younger = 68, older = 133; 3.41%), 46 trials due to poor validity in the eye-tracking record (younger = 31, older = 15; 0.78%), 43 trials with an excessive number of fixations during the encoding phase (e.g., failures of the parsing algorithm based on the upper quantile of the overall distribution of fixation count, younger = 18, older = 25; 0.73%), 1,684 trials which had no fixation on the critical object during the recall phase (younger = 832, older = 852; 28.57%), and 953 trials in which participants failed to detect the change (younger = 441; older = 512; 16.17%). After exclusions, a total of 1,744 change trials contributed to the eye-movement analyses for younger adults, with a by-participant average of 64.6 ± 11.9, and 1,223 trials for older adults, with a by-participant average of 51 ± 15.17 trials. When grouping the trials contributing to the analysis of the eye-movement responses by the type of change, we have a by-participant average of 24.3 ± 5.13 Identity trials, 19.4 ± 3.85 Location trials, and 20.9 ± 4.58 Both trials for the younger group, and 19.1 ± 5.93 Identity trials, 14.8 ± 5.35 Location, and 17.8 ± 4.78 Both trials for the older adults.
Strategy of Analysis and Measures
Our primary goal was to investigate how the scene’s spatial layout influences change detection performance. To this end, we compared the detection accuracy (a binary variable: 1 = correct, 0 = incorrect) of the present study with that of its sibling dataset (D’Innocenzo et al., 2022), where location changes always implied an empty spot in the original object position, rather than a swap with another object present in the scene. The predictors were Group (Younger, Older; with Older as the reference level), Type of Change (Identity, Location, and Both; with Identity as the reference level), and Dataset (NoSwap = previous study, Swap = current study; with NoSwap as the reference level). In this analysis, Identity was chosen as the reference level for Type of Change, as the only condition implemented identically across both datasets, and thus appropriate for assessing differences in location changes.
We then focused exclusively on the current dataset to determine whether detection performance was further modulated by the Semantic Consistency of the critical object (Consistent, Inconsistent; with Consistent as the reference level). Consistency was determined based on the median ratings of object consistency across the scenes used in the study, as normed in the VISIONS database (Allegretti et al., 2025a). For this analysis, we examined both detection accuracy and, on correct trials only, reaction times (RTs), defined as the interval between the onset of the scene during the recall phase and the participant’s keyboard response.
Finally, on successful trials, we also examined the attentional strategies supporting accurate detection by looking at eye-movement measures capturing two different stages of visual processing of the critical object: (a) the latency to the first fixation, defined as the time occurring between the onset of the scene and the first fixation made on the critical object, which indicates how quickly the object is identified from the periphery of the visual field, and (b) the first-pass gaze duration, which is the sum of the duration of all fixations made during the first inspection of the critical object and points to the effort to integrate it with the overall meaning of the scene. Eye-movement measures were analysed using the same predictors (i.e., Group, Type of Change, and Semantic Consistency). For all analyses conducted on our dataset, Both was set as the reference level for Type of Change to directly assess the contribution of single-features (Identity and Location) against conjunctive-feature changes in allocating attention during successful memory retrieval. To account for general age-related slowing (Faust et al., 1999), RTs and both eye-movement measures were z-scored within each age group before statistical analysis. To complete our analysis, we examined viewing behaviours by analysing the total number of fixations and re-entries into the critical area to assess potential effects beyond first-pass processing. Because these measures yielded complementary results that did not alter our main conclusions, full details are presented in Supplemental Material S4.
Statistical Modelling
We employed linear and generalised linear mixed-effects models (G/LMER), implemented in the lme4 package in R (Bates et al., 2015). The fixed effects were described above, while the random effects included participants (N = 101 for the comparative analysis; N = 51 for the current dataset) and scenes (N = 153 for the comparative analysis; N = 120 for the current dataset). Models were first built with a full fixed and random effect structure (including all main effects and interactions, with random variables as both intercepts and slopes; Barr et al., 2013) and then reduced backwards using the step() function from the lmerTest package (Kuznetsova et al., 2017) to retain the most parsimonious model adequately capturing the data structure (Matuschek et al., 2017). The final model results report fixed-effect estimates, confidence intervals (proxies for effect size; Luke, 2017), t-values (or z-values for binomial outcomes), and p-values calculated using the Satterthwaite (1946) approximation. Pairwise comparisons were computed using the emmeans R package (Lenth et al., 2020), with Tukey’s correction for multiple comparisons applied. Data and R scripts are available on the Open Science Framework: https://osf.io/sx7gw/.
Results
The Role of the Global Scene Configuration on Detection Accuracy
Detection accuracy was significantly higher in younger than in older participants for location changes and for changes in both identity and location (see Figure 2 and Table 1). However, detection accuracy did not differ significantly between the current study (with the critical object swapped) and the previous study (with the critical object displaced). Ageing did not interact with any other predictor, suggesting that the capacity to bind object identity to spatial location remains substantially preserved in older adults, even when the overall scene configuration remains intact (i.e., the swap dataset). Changes in the global scene configuration, however, did impact location-only changes. These were detected significantly more accurately in the no-swap dataset (where layout was disrupted) than in the swap dataset (pairwise comparison: z-ratio = 4.53, p = .001). This confirms that when an object swap preserves the spatial layout, participants must rely on a detailed representation of object-to-object relationships; in contrast, the “empty space” left by a displacement in the no-swap condition provides a strong bottom-up perceptual cue. As expected, no effects of global layout (i.e., no differences across datasets) were observed in the identity-only condition (z-ratio = 0.27, p = 1.00), as this manipulation was identical across both studies. Likewise, no differences were found when both features changed (z-ratio = 2.5, p = .12). This suggests that object identity contributes additively to detection, independently of the scene layout.

Percentage detection accuracy (y-axis) as a function of group (x-axis) across the two datasets (noswap = green circles, swap = warm beige triangles). The three types of change are compared within the panel. The hinges of the boxplots represent the 25th and 75th percentiles of the measure (lower and upper quartiles), while the horizontal line represents the median of the distribution. Each dot indicates the by-participant average for that factor.
Generalised Linear Mixed-Effects Model Output for Detection Accuracy (a Binomial, 0–1, Incorrect vs. Correct Responses) as Predicted by Group (Younger vs. Older, Reference = Older), Type of Change (Identity, Location, Both, Reference = Identity), and the Dataset Compared (Swap, NoSwap, Reference = NoSwap). Participants (101) and the Unique Identifier of the Scene Item (153) were the Random Effects Introduced as Intercept and Slope.
Note. The final model formula in Wilkson notation, resulting from stepwise backwards selection, is: detection accuracy ~ type of change + group + set + type of change:set + (0 + type of change | sceneID) + (0 + type of change | subjectID). SE: standard error; CI: confidence interval. Bold values indicate statistically significant effects (p < .05).
The Influence of Object Semantic Consistency and Spatial Location on Speed and Accuracy of the Change Detection
Focusing on the current (swap) dataset, accuracy was higher for younger adults and, across groups, when both features changed compared to when only one feature changed. We observed a significant interaction between Type of Change and Group: while both groups benefited from conjunctive changes, older adults showed a disproportionately larger benefit when both features changed compared to location-only changes (Older: z = 5.82, p < .001; Younger: z = 5.46, p < .001), indicating a strong reliance on dual cues. Crucially, a three-way interaction revealed that younger participants detected changes involving semantically inconsistent objects more accurately than those involving consistent objects, but only when the location changed (z-ratio = −4.96, p < .001). No such difference was observed when both identity and location were changed (z-ratio = 2.1, p = .62; see Figure 3A and Table 2). This suggests that inconsistent objects, which stand out from the scene context, enhance younger adults’ memory for specific spatial positions. Consistent objects, on the other hand, are harder to spatially track because they blend into the overall gist of the scene. Interestingly, this “inconsistency advantage” was absent in older adults (Location: z-ratio = −1.49, p = .94; Both: z-ratio = −1.00, p = .1). This implies that while violations of contextual expectations persist in the working memory of younger adults across the retention interval, older adults appear to lose this specific object-level information more rapidly.

Percentage detection accuracy (A) and z-scored reaction times (B) as a function of the type of change (x-axis) across the two levels of critical object consistency (consistent = grey circles, inconsistent = red triangles). Older and younger adults are compared within (A). For all boxplots, hinges represent the 25th and 75th percentiles of the measure (lower and upper quartiles), and the horizontal line represents the median of the distribution. Each dot indicates the by-participant average for that factor.
Generalised Linear Mixed-Effects Model Output for Detection Accuracy (a Binomial, 0–1, Incorrect vs. Correct Responses) as Predicted by Group (Younger vs. Older, Reference = Older), Type of Change (Identity, Location, Both, Reference = Both), and the Critical Object’s Consistency (Consistent vs. Inconsistent, Reference = Consistent). Participants (51) and the Unique Identifier of the Scene Item (120) were the Random Effects Introduced as Intercept and Slope.
Note. The final model formula in Wilkson notation, resulting from stepwise backward selection, is: detection accuracy ~ type of change + consistency + group + type of change: consistency + type of change:group + consistency:group + type of change:consistency:group + (0 + type of change | sceneID) + (1 | subjectID). SE: standard error; CI: confidence interval. Bold values indicate statistically significant effects (p < .05).
RTs on correct trials largely mirrored the accuracy pattern: single-feature changes were detected significantly slower than conjunctive changes. We also observed a main effect of Consistency, with slower responses to inconsistent objects. Changes involving both features were detected faster than location-only changes, but only when the object was consistent (i.e., two-way interaction, z-ratio = −3.92, p = .001). This pattern likely reflects the encoding history of the object: a consistent object that undergoes an identity change (in the “Both” condition) was encoded as inconsistent, making the change highly salient compared to a simple location shift of a consistent object (Figure 3B and Table 3).
Linear Mixed-Effects Model Output for Reaction Times (Continuous, z-Scored) as Predicted by Group (Younger vs. Older, Reference = Older), Type of Change (Identity, Location, Both, Reference = Both), and the Critical Object’s Consistency (Consistent vs. Inconsistent, Reference = Consistent). Participants (51) and the Unique Identifier of the Scene Item (120) were the Random Effects Introduced as Intercept and Slope.
Note. The final model formula in Wilkson notation, resulting from stepwise backward selection, is: detection accuracy ~ type of change + consistency + group + type of change:consistency + consistency:group + (1 | sceneID) + (0 + type of change | subjectID). SE: standard error; CI: confidence interval. Bold values indicate statistically significant effects (p < .05).
Attentional Strategies Underlying the Successful Detection of Changes
Analysis of latency to first fixation revealed that objects consistent at retrieval were fixated significantly faster than those that were inconsistent (Figure 4A). This prioritisation was driven by their history: these objects were semantically inconsistent at encoding, and this initial mismatch guided memory-driven attention to their location more rapidly during retrieval. No significant age effects were observed on this measure. Analysis of first-pass gaze duration showed that objects changing only in identity were fixated for longer than those changing in both identity and location (Figure 4B). This suggests that the additional spatial cue in the “Both” condition facilitates faster processing, whereas identity-only changes take longer to resolve the mismatch. Furthermore, inconsistent objects were fixated for longer than consistent ones, reflecting the effort required to integrate objects that violate contextual expectations. This effect was significant for identity changes (z-ratio = 3.26, p = .02) and conjunctive changes (z-ratio = 3.54, p = .006), but not for location-only changes (z-ratio = 0.19, p = 1). This indicates that when an object’s identity is transformed, the VWM representation must be actively updated with novel semantic information, demanding greater attentional resources (Table 4).

Eye movement measures on the critical object. (A) Latency of the first fixation (y-axis) as a function of the two levels of critical object Consistency (x-axis). (B) First-pass gaze duration on the critical object plotted on the y-axis as a function of the Type of Change (x-axis) across the two levels of critical object Consistency (consistent = grey circles, inconsistent = red triangles). The hinges of the boxplots represent the 25th and 75th percentiles of the measure (lower and upper quartiles). The horizontal line represents the median of the distribution instead. Each dot indicates the by-participant average for that factor.
Linear Mixed-Effects Model Output for the Eye-Tracking Measures of Latency to the First Fixation and First-Pass Gaze Duration on the Critical Object. Predictors Entered in the LMER were: Group (Younger vs. Older, Reference = Older), Type of Change (Identity, Location, Both, Reference = Both), and the Critical Object’s Consistency (Consistent vs. Inconsistent, Reference = Consistent). Participants (51) and the Unique Identifier of Scene Item (120) were the Random Effects Introduced as Intercept and Slope.
Note. The final model formulas in Wilkson notation, resulting from stepwise backward selection, are:
(a) Latency to first fixation ~ consistency + (0 + consistency | subjectID) + (1 | sceneID).
(b) First-pass gaze duration ~ type of change + consistency + type of change:consistency + (0 + consistency | subjectID) + (0 + consistency | sceneID).
SE: standard error; CI: confidence interval.
Discussion
The interplay between global scene properties (gist-based accounts; e.g., Greene & Oliva, 2009; Oliva & Torralba, 2006) and individual object features (object-based accounts; e.g., Biederman, 1981; Võ, 2021) in forming and accessing visual short-term memory remains a debated topic (Intraub, 2012; Kravitz & Behrmann, 2011; Markov et al., 2019; Tatler & Land, 2011). Ageing provides a useful opportunity to examine the relative contribution of these sources, as declines in VWM are typically accompanied by a greater loss of fine-grained details relative to more preserved global representations, suggesting an age-related shift in their contribution towards greater reliance on gist-based information (Greene & Naveh-Benjamin, 2023; Grilli & Sheldon, 2022). The present study investigated the impact of global spatial layout on the detection of changes to an object’s identity, location, or both, while also varying its semantic consistency. By analysing change detection accuracy alongside eye-movement patterns during retrieval in younger and healthy older adults, we assessed the extent to which VWM representations depend on global spatial relationships versus independent object features, and how these representational mechanisms are affected by ageing.
Not surprisingly, older adults exhibited a general decline in detection accuracy. This is consistent with age-related reductions in VWM capacity (Brockmole & Logie, 2013; Pertzov et al., 2015; Read et al., 2016; Thomas et al., 2012) and prior research using naturalistic scenes (Costello et al., 2010; Rizzo et al., 2009). Nonetheless, comparing the present results with our earlier study (D’Innocenzo et al., 2022) revealed that the spatial configuration of the scene modulated the detection of location changes to a similar extent across age groups. In that earlier work, location changes created an empty space, disrupting the scene’s layout, and were therefore easier to detect than in the current object-swap design, where the global configuration was preserved.
Relocating an object to a previously unoccupied position likely generates salient perceptual cues by producing discontinuities both at the object (disappearance) and scene level (previously occluded regions become visible). As a result, change detection may depend on identifying global scene differences, without necessarily relying on detailed memory for the object itself, confirming the importance of spatial coherence in organising VWM contents (Hollingworth, 2006, 2007; Hollingworth & Rasmussen, 2010; Jiang et al., 2000; Olson & Marshuetz, 2005; Rensink, 2000, 2002; Simons, 1996) and the evidence that salient perceptual changes (e.g., appearance or disappearance of objects) improve their detection (Cole & Liversedge, 2006; Cole et al., 2004). Conversely, when the spatial configuration remains intact (i.e., via object swapping), successful detection requires greater effort to associate the object’s identity with its location (Irwin & Zelinsky, 2002). This interpretation aligns with ERP data, which show that object displacements elicited early neural responses, while object swaps were associated with later components, reflecting increased demands in retrieving the relational structure of objects within a scene (van Hoogmoed et al., 2012). The lack of age differences in this spatial-layout effect suggests that the short-term maintenance of global spatial structure is preserved in ageing, an ability that, to our knowledge, had not been directly tested before.
Beyond the role of spatial coherence, our study also demonstrates that object-based information is critical for evaluating VWM representations and actively supports successful change detection. Changes involving both object identity and location (i.e., feature conjunction) were detected better and faster than changes to a single feature, independent of ageing, which replicates our previous findings of preserved mnemonic strategies in naturalistic contexts by healthy older adults (Allegretti et al., 2025b; Coco et al., 2022; D’Innocenzo et al., 2022), while arguing against evidence of age-related binding deficits in VWM (Chalfonte & Johnson, 1996; Kessels et al., 2007; Mitchell et al., 2000; Muffato et al., 2019). Critically, this binding advantage persisted even when the spatial layout of the scene was preserved, indicating that the benefit for conjunctive changes does not arise from disruptions to the global scene structure but instead reflects the additive contribution of object identity and location information within VWM. Moreover, the detection performance of objects that changed in identity was not significantly modulated by disruptions (or not) in the spatial layout of the scene, and it benefited from the violation of contextual expectations (i.e., “inconsistency advantage,” Allegretti et al., 2025b; Hollingworth & Henderson, 2000; LaPointe & Milliken, 2016; LaPointe et al., 2013).
Interestingly, however, the effect of object semantics was observed only when the object changed in location and was restricted to younger adults. Semantically inconsistent objects mismatch with the overall scene meaning (Allegretti et al., 2025a, 2025b; Borges et al., 2020; Castelhano & Heaven, 2011; Coco et al., 2020; Davenport & Potter, 2004; Gordon, 2004; Malcolm & Henderson, 2010; Mudrik et al., 2010; Spotorno et al., 2014; Underwood & Foulsham, 2006; Underwood et al., 2007, 2008; Võ & Henderson, 2009), and thus are more easily noticed when moved within the scene, suggesting that spatial memory is not simply tied to the location of an object but rather to the semantic features bundling its identity (Toh et al., 2020), which points to object-based representations of scenes (Biederman, 1981; Henderson, 1992; Thorpe et al., 1996; Võ, 2021; Võ et al., 2019). In contrast, consistent objects blend into the scene context as they provide semantically homogenous information with respect to it (i.e., “less informative,” Loftus & Mackworth, 1978). This proposition is supported by recent fMRI data, which show that searching for consistent objects requires greater activation of the dorsal attention network. This reflects the increased top-down effort to discriminate them from other semantically similar objects within the scene context (Salsano et al., 2025). However, older adults did not benefit from semantic information for any type of change, suggesting age-related declines in maintaining faithful object representations over the retention interval (Brockmole & Logie, 2013; Costello et al., 2010; Rizzo et al., 2009). Thus, the distinctive memory trace that helps younger adults detect changes to inconsistent objects appears to deteriorate more quickly with age, eliminating the typical inconsistency advantage. This pattern also aligns with evidence that older adults may rely more heavily on consolidated expectations to support their attentional and memory processes (e.g., Ramzaoui et al., 2022; Wynn et al., 2020), and so are less sensitive to context violations or less likely to capitalise on the semantic salience that inconsistent objects normally provide.
Taken together, our results on the detection performance in naturalistic scenes lend theoretical support to models of VWM postulating that objects are represented as hierarchically organised bundles of features (Brady et al., 2013; Fougnie & Alvarez, 2011) rather than as integrated single units (Hardman & Cowan, 2015; Luck & Vogel, 1997; Vogel et al., 2001). In fact, if the capacity of recognising changes to objects were independent of their actual features, we should not have observed any significant difference across the conditions of identity or location, and especially, on the semantic fit of the object within the scene. Under such a framework, a change is detected by recalling the object as an indivisible unit; thus, the features that characterise its representation in VWM have an equal weight, a prediction clearly contradicted by our findings. Indeed, the robust binding advantage we observed indicates that each feature contributes independently; thus, the greater the mismatch between the incoming input and its hierarchical representation in VWM, the more likely it is that object changes will be identified. Future studies could examine whether the binding advantage stems from feature integration independently of task demands (i.e., detecting any change in the scene) by manipulating instructions that prioritise either object identity or location, and so assess if there is a mnemonic advantage even when one of the two features is task-irrelevant.
Beyond the manual detection performance, our study examined the role of overt attention in supporting successful memory retrieval by analysing eye movements during the recall phase. Departing from previous literature reporting attentional prioritisation (i.e., earlier latencies for the first fixation) of semantically inconsistent objects relative to consistent ones (see Wu et al., 2014, for a review), we observed the exact opposite here, paralleling our finding of shorter RTs for consistent than inconsistent objects. The explanation of this apparently contradictory finding, however, is straightforward. The prioritisation effect in previous literature was found contingent on the first presentation of the scene during free-viewing (Allegretti et al., 2025a; de Graef et al., 1992; Pedziwiatr et al., 2022), visual search (Borges et al., 2020; Underwood & Foulsham, 2006), or the study phase of memory tasks (Allegretti et al., 2025b; Coco et al., 2020; Loftus & Mackworth, 1978; Spotorno & Tatler, 2017), but in this study, we focused on the recall phase, that is, the second presentation of the scene, which underlies different attentional mechanisms from those operating at encoding (e.g., Damiano & Walther, 2019). Therefore, it is plausible that participants were faster to fixate on consistent objects because they had been encoded as inconsistent during the initial study phase; this early semantic mismatch likely strengthened their memory trace, rendering the objects more salient and drawing attention more rapidly when the scene reappeared at retrieval. To validate this interpretation, we examined eye movements during the encoding phase of successful detections and replicated classic findings of shorter latencies (and longer gaze durations) for inconsistent compared to consistent objects (refer to Supplemental Material S5 for this analysis).
With respect to the type of change, we did not find any difference in the latency of the first fixation related to either identity or location, in contrast to our previous work (Coco et al., 2022; D’Innocenzo et al., 2022, and see Supplemental Material S6 for the analysis comparing eye movement metrics across the two studies). This discrepancy, however, highlights once again the role played by spatial layout, especially in relation to the effects of location. In our earlier no-swap designs, the disruption of global scene structure created a spatial discontinuity that reliably captured attention, drawing gaze towards the now-vacant region and delaying the attentional orienting towards the displaced object at its new position. In the present study, the spatial configuration remained intact due to object swapping, eliminating this bottom-up cue and, consequently, the layout-driven delay previously observed.
Once fixated, the critical object was viewed for longer when only its identity changed than when both its identity and location changed, and this was independent of its consistency with the scene. This finding suggests that when multiple features can be accessed to retrieve an object from memory, attentional demands are reduced, consistent with the notion of independent feature storage in memory (Brady & Alvarez, 2015). Inconsistent objects were fixated for longer than consistent ones during first inspection, replicating prior findings that semantic mismatches demand greater attentional effort (Allegretti et al., 2025a, 2025b; Borges et al., 2020; Coco et al., 2020, 2022; Cornelissen & Võ, 2017; Loftus & Mackworth, 1978; Võ & Henderson, 2009), and this effect was specific to identity changes, where new semantic content had to be integrated, but absent for location changes, likely because in such case the contextual mismatch was already solved at encoding. No significant age-related differences were found in eye-movement patterns during successful detections, suggesting that both younger and healthy older adults effectively use contextual information from naturalistic scenes to guide attention and support memory performance (Allegretti et al., 2025b; Coco et al., 2022; D’Innocenzo et al., 2022; Mäntylä & Bäckman, 1992).
Our findings demonstrate that VWM representations of scenes depend on a dynamic interplay between their spatial layout and the individual features of objects therein, a dynamic that is largely preserved in healthy ageing. On the one hand, when the scene layout is disrupted (as in our 2022 study), location changes are detected more accurately, and during successful detections, first fixations were more rapidly detected towards critical objects that changed only in their identity, effects that do not emerge when the layout is preserved, as in the present study. These findings confirm that spatial coherence contributes to accessing information about individual objects stored in VWM and suggest that structural alterations, that is, the “empty space,” capture bottom-up attention, momentarily delaying orienting. On the other hand, we demonstrated a robust “binding advantage,” where concurrent changes to an object’s identity and location were more reliably detected than single-feature changes, an effect that persisted even when the global scene layout was maintained via object swapping. This indicates that while layout disruptions provide salient perceptual cues for change detection, successful retrieval demands accessing information that is primarily bound to objects. This proposition is further corroborated by the critical role of object semantics: its violation of contextual expectations (i.e., semantic inconsistency) creates a strong mnemonic signal that boosts detection, particularly for location changes. Eye-tracking corroborated these behavioural results: during successful retrieval, inconsistent objects elicited longer gaze durations, reflecting the cognitive effort required to integrate them into the scene representation. Although older adults performed less accurately overall, the underlying pattern of attentional and mnemonic effects closely mirrored that of younger adults, indicating that object-based representational mechanisms remain largely intact in healthy ageing and, for the first time, offering evidence that memory for a scene’s structural gist is preserved across the lifespan.
A primary limitation of this study is the inherent overlap between manipulating object identity and semantic consistency, a constraint imposed by the stimulus database used. Because any change in object identity necessarily alters its semantic fit with the scene, our design cannot fully disentangle the unique contributions of these two factors. To isolate their respective influences, future research should create stimuli where identity can change while semantic consistency is held constant (e.g., swapping one type of fruit for another in a kitchen). Nevertheless, the effects of semantic consistency were also observed in the location condition, where we found that changes to inconsistent objects were better detected than those to consistent ones, even though the objects’ identities remained unchanged between encoding and recognition. This suggests that semantic information contributes to the object’s internal representation independently of changes in identity, influencing how it is accessed during retrieval. We also note that location change manipulation involved swapping the spatial placement of two objects, which can be interpreted as two identity changes occurring in different positions, thereby introducing a potential asymmetry relative to the identity-change condition. Because participants were not asked to specify in their response which object had changed, we cannot determine whether detection was driven by one or both swapped objects. However, if the presence of two changed objects had facilitated the detection, accuracy should have been higher in the location-change condition than in the identity-change condition. This was not observed, suggesting that performance depended mainly on memory for object-location bindings rather than on the number of changed objects.
This work also opens several avenues for future investigation. First, the precise temporal dynamics of how spatial and semantic information are integrated during memory retrieval remain unclear. Employing methods with high temporal resolution, such as co-registered eye-tracking and EEG, could reveal whether the scene’s spatial layout provides an initial scaffold for memory access or if spatial and semantic cues are processed in parallel. If retrieval follows a similar temporal hierarchy to scene perception, in which global scene structure is accessed rapidly and subsequently guides the recognition of local object details, then disrupting the layout should produce early neural responses (~100 ms) time-locked to scene onset, reflecting reinstatement of the scene’s overall structure. Semantic modulation (e.g., N400) tied to specific objects would instead only emerge during fixations on those objects. Second, our paradigm focused on change detection without requiring an explicit identification of what had actually changed. Future studies could incorporate identification tasks to determine whether detection is driven by a general mismatch signal or by conscious access to the specific features that have changed, further clarifying the nature of the underlying VWM representations.
Conclusion
In conclusion, our findings bridge the gap between gist-based and object-based theories by demonstrating that VWM for objects embedded in naturalistic scenes flexibly utilises both the global spatial scaffold and the specific features of individual objects. While layout disruptions provide potent cues, object features have a primary role in guiding attention and memory when the scene structure is stable. The preservation of this dynamic across the adult lifespan underscores its fundamental importance in healthy cognitive ageing. Ultimately, this work refines our understanding of VWM, advancing an integrated model where coherent representations of the world are built from the hierarchical and interactive contributions of both the scene and, especially, objects as their binding units.
Supplemental Material
sj-pdf-1-qjp-10.1177_17470218261445069 – Supplemental material for Object Features Hierarchically Guide Eye-Movement and Short-Term Memory for Naturalistic Scenes with Preserved Spatial Layout Across the Lifespan
Supplemental material, sj-pdf-1-qjp-10.1177_17470218261445069 for Object Features Hierarchically Guide Eye-Movement and Short-Term Memory for Naturalistic Scenes with Preserved Spatial Layout Across the Lifespan by Elena Allegretti and Moreno I. Coco in Quarterly Journal of Experimental Psychology
Footnotes
Acknowledgements
The authors would like to thank Bernard Fan for his assistance in the initial stages of task implementation and Sergio Della Sala for his insightful comments and feedback on the theoretical framework of this work. The authors are also grateful to the senior participants from the Villa Gordiani recreational centre in Rome for generously contributing their time to this research.
Ethical Considerations
The study was performed in accordance with the ethical standards laid down in the 1964 Declaration of Helsinki. Ethical approval was obtained from the Ethics Committee for Transdisciplinary Research of Sapienza, University of Rome (CERT_4_186B65F32D9).
Consent to Participate
All participants provided written informed consent before starting the experimental session, and their responses were collected anonymously to protect their privacy.
Consent for Publication
The authors affirm that all participants provided informed consent for publication of their data.
Author Contributions
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study was supported by the research grant “Ateneo Grande” (RG123188B3F9299C) and “Ateneo Medio” (RM122181673FAF68), awarded to MIC and funded by Sapienza University of Rome.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
Supplemental material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
