Abstract
This commentary builds on Karadöller et al.’s endeavor to include gesture studies on the one hand and sign language studies on the other to highlight the crucial part played by the visual-gestural modality in child language development. While we acknowledge the invaluable contribution of their paper to research on multimodal language development, we question the authors’ phrasing of the relationship between gesture and language, as well as the selection of reviewed studies, arguing that it might narrow perspectives on methodologies, contexts, age groups, and cultural influences. We advocate for a deeper recognition of the parallels within visual-gestural modality across both sign and spoken (multimodal) languages.
We apprehend language use as a shared human craft, rooted in sensorimotor and social connection. This multimodal social skill is shaped and transmitted to children by the expert languagers who surround them (Morgenstern, 2023). It is, therefore, essential for us to expand Cienki’s (2015) insights on multimodal language as well as Ochs’s (2012) framework on how children experience language to trace children’s developmental pathways. From this perspective, Karadöller et al. (2025) offer a particularly insightful exploration of language acquisition – first in spoken and then in sign languages. Their choice to examine visual-gestural expressions in the two types of languages is especially valuable. It enables a comparative analysis that highlights both shared and distinct mechanisms underlying language development across different modalities. The paper provides a thorough overview of a selected choice of focus topics (iconicity, indexicality, simultaneity) while engaging with foundational theoretical constructs that have shaped research on multimodal communication. Since the seminal work of Bates in the 1970s, numerous studies have investigated the development of the visual-gestural modality in hearing children, both as an area of interest in its own right and in relation to vocal development. Simultaneously, research on sign language acquisition in deaf children has advanced significantly. As a result, Karadöller et al.’s proposal is highly topical and challenging at the theoretical and methodological levels. However, we suggest that (1) their positioning on the relationship between gesture and language could be clarified, and (2) the literature review could have been expanded to include topics that further enhance the integrative perspective on the visual-gestural modality.
The relationship between gesture and language
Despite their global stance on including gesture in the analysis of language, the authors’ phrasing of the relationship between gesture and language occasionally either contradicts a multimodal perspective on the nature of language (‘gestures may even be a precursor to language’) or is ambiguous (‘gestures are not entirely independent of language’), and it often reflects consideration of gesture as secondary to speech (‘the development of speech and accompanying gestures’). Some gesture specialists indeed emphasize the imagistic power of gesture as a partner to language, while others focus on the praxic origins of gesture and consider them as part of language (Morgenstern & Goldin-Meadow, 2022).
We lack the terminology to discuss language and its development independently of modality, adopt an a-modal perspective, and give equal consideration to both gestural and vocal expressions. Only recently have researchers proposed terms such as ‘multimodal languages’ (Müller, 2018) or ‘languaging’ (Linell, 2009) to describe language activity regardless of modality. Similarly, the authors’ formulation—linking ‘words’ and ‘sentences’ to speech when they phrase their goal as reviewing the roles of gesture ‘in the context of first word acquisition, sentence production, vocabulary development’ – suggests that a number of specialists do not fully recognize ‘gestures’ as linguistic units capable of conveying complex meanings, morphological features, or syntactic functions, and oppose them to ‘signs’ in sign languages.
We thus need to clarify what we mean by ‘gesture’ as not all authors associate the term with the same concept. Different positions may emerge depending on the approach one adopts and the role given to interaction as a driving force for child language development.
Enriching the scope of the literature overview
Although the final section offers excellent recommendations for future directions, the paper’s overall scope somewhat restricts deeper insights into potential gaps in existing research. It could thus reinforce methodological and theoretical biases that are already present in the literature. We identify five key areas whose inclusion could have further strengthened the paper.
(1) The literature taken into account in the paper seems to have set aside several functional studies – beginning with Bates et al. (1975) on proto-declarative versus proto-imperative pointing and subsequent research. This suggests the selection of a perspective on language acquisition that leans toward a formal approach. Giving more consideration to functional-interactional research could lead to a more comprehensive integration of gesture into language studies, particularly in comparing spoken (or rather, multimodal) languages with sign languages. The authors’ portrayal of pointing, – ranging from a speech precursor to a facilitator and a compensatory mechanism for missing speech –suggests that some of its linguistic and interactional functions may have been overlooked (Morgenstern, 2021). These include anaphora in narratives, turn-taking, and referencing prior speech. Such functions become more evident in spontaneous, multimodal, and multi-party interactions captured in a variety of cultures, where language competence is seen as inherently social (de León & Garcia-Sánchez, 2021).
(2) Since the 1980s, some studies have examined both hearing and deaf signing children (Caselli, 1983; Volterra, 1981). Studying spoken and sign languages together is crucial for understanding the dynamic interplay of action, audio-vocal, and visual-gestural modalities in meaning-making, as separate analyses risk overlooking cross-modal interactions. An integrative approach reveals patterns missed in isolated studies (Boutet et al., 2021; Capirci et al., 2022), tracing the conventionalization of gestures from idiosyncratic use to grammaticalized forms along the pathway outlined by Wilcox (2004).
(3) Meaning emerges dynamically through engagement with the environment, social interactions, shared discourses, and common ground. Capturing this process may require to include research conducted over time and across diverse contexts and cultures (Haviland, 1998), including varied environments, activities, and multi-party interactions (Morgenstern et al., 2021).
(4) While deictic and iconic gestures are the most frequently studied in language acquisition, as highlighted in the paper, broader studies of gesture types have given us valuable insights into multimodal development. Pragmatic or recurrent gestures (Ladewig, 2014), shaped by culture and embodied experiences, are crucial to capture the full range of language practices. Emerging from multimodal interactions over time, these gestures can become conventionalized. Analyzing recurrent gestures (e.g., shrugs and headshakes) in longitudinal data through a multimodal lens reveals meaning-making across body segments beyond just hands or upper limbs and sheds light on the evolution of form-meaning pairings (Andrén, 2014; Beaupoil-Hourdel & Morgenstern, 2021).
(5) Although still limited in number, we suggest also drawing inspiration from studies on older children and on the place and role of gesture beyond the lexical level, as highlighted by Colletta (2021). This would encourage the integration of gestures in syntax and the multilayered dimensions of language throughout the child’s development (Beaupoil-Hourdel, 2021).
Failing to consider these factors and incorporate the relevant literature (which we cannot list exhaustively here) risks reinforcing biases rather than challenging them. As Karadöller et al. (2025) highlight, a systematic study of the place and role of the visual-gestural modality in sign and spoken (multimodal) language development together, over a longer time span and in a variety of contexts, activities, and cultures would open up an innovative research avenue. Doing so would illuminate how visual-gestural linguistic expressions in both types of languages exist on a continuum, with closely interconnected semantic, pragmatic, and formal features, and how interacting human bodies, permeated with affect in a variety of situations and environments, play a significant role in shaping meaning in children’s language development.
Footnotes
Author contributions
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The authors received financial support from the Agence Nationale de la Recherche for their research on Multimodal interaction ANR project-20-CE27-0010-01.
