Abstract
Countertenors occupy an ambivalent position in the cultural politics of voice—audible yet unintelligible, visible yet misrecognized. This article examines how male singers in this vocal register negotiate platformed audibility across algorithmically governed spaces such as TikTok and Instagram. Drawing on qualitative interviews and reflexive thematic analysis, the study foregrounds how voice operates not simply as a medium of expression, but as a contested cultural form shaped through regimes of gendered listening and platform logics of recognition. Participants describe being misheard, disbelieved, or compelled to justify their vocal legitimacy—what this study conceptualizes as auditory disciplining and sonic pre-legitimation. Yet amid these structural pressures, countertenors also craft spaces of affective resonance and situated vocal agency, often outside institutional gatekeeping. By theorizing voice as relational, negotiated, and infrastructurally mediated, the article contributes to broader debates on digital subjectivity, cultural legibility, and the politics of mediated embodiment.
Keywords
Introduction
Voice is not merely sound—it is a socially encoded, affectively charged, and culturally mediated form of expression (Edgar, 2019). Before it is evaluated on technical grounds, the voice is filtered through cultural assumptions about gender, authenticity, and bodily identity (Strand, 1999). These assumptions reflect hegemonic logics that link pitch, timbre, and resonance to binary gender norms (Lavan et al., 2019). Listening, therefore, is not a neutral act but a culturally situated one (Neuenswander et al., 2024), shaped by expectations of which voices “belong” to which bodies (Ardener, 2006).
In digital environments, such assumptions are often intensified. Platform infrastructures—including interface design, feedback systems, and user- generated commentary—not only shape how voices circulate (Nass et al., 1997), but also how they are categorized (Schumacher, 2022), interpreted (Simon, 2019), and rendered meaningful (Horowit-Hendler and Hendler, 2020). Performers are rarely judged by voice alone; instead, sonic, visual, and textual cues are expected to align within intelligible identity categories. When vocal performance disrupts these alignments—particularly by decoupling pitch from normative gender presentation—it frequently provokes confusion, correction, or disbelief (Mahmood and Huang, 2024).
The countertenor voice exemplifies such dissonance. As a male voice in the contralto or mezzo-soprano range, typically achieved through falsetto (Cruz et al., 2024), it challenges dominant assumptions about voice– gender congruence. Although historically rooted in classical music (Giles, 1994), countertenors remain unfamiliar to many digital audiences. On social media, their voices often elicit skeptical reactions—such as “Is this a woman?” or “Is this fake?”—reflecting discomfort with auditory ambiguity (Fugate, 2016).
These tensions become especially pronounced on platformed stages (DeVito et al., 2017), where self-presentation is inherently multimodal. Voice, image, caption, and gesture converge in the production of identity (Dritsas et al., 2025; Kersten and Lotze, 2020). For non-normative vocal performers, this convergence entails additional labor: anticipating misrecognition, explaining vocal range, and managing the affective consequences of being misheard (Scolere et al., 2018). Such labor is shaped by platform-specific dynamics of visibility and audibility (DeVito et al., 2017).
Despite growing research on digital self-presentation, voice remains undertheorized within media and cultural studies. Existing scholarship has largely privileged visuality (Tiainen, 2013), textual discourse (Jin, 2024), or performative aesthetics (Pang, 2022), while often treating sound as secondary. Vocal pedagogy likewise centers on technique (Cobb, 2022) or physiology (Gill and Herbst, 2016), with limited attention to how voices are socially received and interpreted.
This study addresses this gap by examining how countertenors experience, negotiate, and respond to the reception of their voices in platformed contexts. Drawing on in-depth interviews with six singers active on social media, it asks:
How are countertenor voices interpreted, misheard, or corrected in relation to dominant gendered expectations? How do countertenors construct and manage their vocal identities across sonic, visual, and textual modalities? How do these experiences shape their sense of artistic legitimacy and affective belonging?
By foregrounding the cultural labor of “sounding different” (Edgar, 2019; Schäfers, 2017), this study contributes to broader conversations on voice, recognition, and platformed subjectivity by conceptualizing listening as a socially mediated process that governs who is heard and under what conditions voices are deemed legitimate.
Literature review and theoretical framework
Vocality as gendered cultural practice
Vocal expression is widely understood as a culturally embedded practice shaped by social norms and interpretive frameworks (Tiainen, 2013). Feminist media studies and critical voice theory have shown that vocal production and reception are embedded in regimes of identity, power, and representation (Ashby, 2011; Eidsheim, 2009, 2019). Voices do not merely reflect gender; they are heard through cultural logics that associate pitch, timbre, and resonance with gendered meanings (Eidsheim, 2015).
Listeners routinely draw on these logics when making judgments about what they hear (Schumacher, 2022). High-pitched voices are conventionally associated with femininity, while low-pitched voices are linked to masculinity and authority (O’Connor and Barclay, 2017). These associations are culturally produced rather than biologically fixed (Chen et al., 2021). Vocal performances that transgress these norms—such as male countertenors—often provoke discomfort, misrecognition, or dismissal (Tiainen, 2013). Listening thus operates as an active interpretive process that enforces gendered legibility through sound (Eidsheim, 2015). These insights resonate with Butler’s (1990) theory of gender performativity, which conceptualizes gender as the repeated enactment of social norms. Vocal expression functions as one such enactment, actively participating in the constitution of gendered identity (Schlichter, 2011; Jacobs, 2017). Extending this framework, Eidsheim (2019) emphasizes that vocal perception itself is culturally conditioned, shaped by intersecting structures of race, gender, class, and cultural capital. What is perceived as “natural” or “authentic” is therefore inseparable from learned modes of listening (Heidemann, 2016; Donison, 2022).
These interpretive dynamics are further intensified in digital environments, where voices circulate through technologically mediated spaces and are evaluated through metrics, algorithms, and commentary (Bhandari and Bimo, 2022; Baumann et al., 2025). Platforms such as TikTok and Instagram do more than host performances; they shape reception by organizing visibility and engagement. For countertenors, whose vocal range challenges dominant gendered expectations, these contexts often generate confusion, skepticism, or ridicule (Fugate, 2016).
In this sense, the countertenor voice is not simply a physiological anomaly or historical curiosity, but a socially situated vocal identity. It becomes a site where gender, authenticity, and embodiment are negotiated through sound, positioning vocality as a form of identity performance, affective labor, and cultural meaning.
Platformed visibility and self-presentation
If cultural norms shape how voices are interpreted (Zhang and Pell, 2022), digital platforms structure the conditions under which voices become visible, legible, and circulated (Van Dijck and Poell, 2013; Papa and Loannidis, 2023). Platform affordances—such as algorithmic curation, interface design, and engagement metrics—introduce new forms of mediation for vocal performers (Richou, 2025). For countertenors, whose vocality already unsettles normative expectations, these infrastructures heighten both visibility and vulnerability.
On platforms like TikTok and Instagram, voice is embedded within multimodal assemblages in which thumbnails, captions, hashtags, facial expressions, and pinned comments shape reception (Simungala et al., 2024). Visibility thus becomes a strategic negotiation across sonic, visual, and textual registers (Henriksen, 2020), requiring performers to manage coherence across modalities (Marwick and Boyd, 2011).
This negotiation is particularly fraught for countertenors. While vocal distinctiveness may attract algorithmic attention, automated tagging and audience assumptions often lead to misclassification or reductive labeling, reflecting biases embedded in platform architectures (Gu et al., 2023; Gutierrez, 2021). At the same time, platform-specific audio tools and processing standards shape normative ideals of vocal professionalism (Dubiel et al., 2024; McDermott, 2012; Provenzano, 2019; Scerbo et al., 2024). Optimized for algorithmic amplification, these norms may smooth over vocal idiosyncrasies and complicate the legibility of countertenor voices (Du et al., 2021; García-Benito, 2025).
These sonic constraints intersect with the visual dominance of platform interfaces (Yang et al., 2025). Facial expressivity, thumbnail spectacle, and on-screen charisma frequently take precedence over nuanced auditory experience, aligning with broader concerns about audiovisual hierarchy and sensory recognition online (Chen et al., 2024). At a structural level, such dynamics reflect the logic of platform governance, in which expression is classified, monetized, and rendered extractable (Couldry and Mejias, 2019; Kaiser, 2025; Srnicek, 2017). Vocal performance thus becomes data—ranked and circulated according to algorithmic imperatives—rendering countertenor voices vulnerable not because of technical deficiency, but because they resist dominant modes of platform legibility (Stjernquist and Carling, 2024).
Vocal legitimacy and affective belonging
For countertenors, platformed self-presentation entails ongoing negotiation of vocal legitimacy. Legitimacy refers not only to technical competence, but to whether a voice is culturally intelligible and socially acceptable (Airoldi, 2024; Publius, 2015; Schmutz, 2009). A high-pitched male voice often disrupts normative assumptions about masculinity, provoking reactions such as misgendering, curiosity, or dismissal.
These responses function as forms of cultural boundary enforcement (Orbe, 1998), rooted in long-standing associations between voice, gender, and authenticity (Fugate, 2016). Such encounters are affectively embodied, shaping singers’ emotional relationships with their voices in public and mediated settings (Mühlhoff, 2015). Belonging in this context is affective rather than institutional, emerging when one's vocal self-image is recognized and affirmed by others (Pratt, 2023).
Sustaining such belonging requires labor. Peer networks, niche publics, and attuned listeners form affective infrastructures that support expressive agency and resilience (Wright-Mair, 2020). Vocal legitimacy thus remains relational and emergent—not conferred by institutions, but produced through iterative cycles of listening, recognition, and response (Ytre-Arne and Das, 2021). Taken together, these processes frame countertenor vocality through three interwoven dynamics: auditory disciplining, platformed identity negotiation, and the ongoing construction of legitimacy and belonging (Lawy, 2017).
Methodology
Research design and data collection
This study adopts a qualitative research design informed by interpretivist and constructivist epistemologies (Tanlaka and Aryal, 2025; Pitard, 2017), seeking to understand how countertenor singers make sense of their vocal identity and platformed experiences. The research privileges participants’ narratives while attending to the broader structural dynamics that shape their expressive agency.
Empirical data derive from semi-structured interviews with six professional or pre-professional countertenors active on digital platforms (e.g., TikTok, Instagram, Bilibili). Participants were selected through purposive and snowball sampling (Palinkas et al., 2015; Naderifar et al., 2017; Stratton, 2024) from a broader pool identified via platform observation and network referrals. The final sample reflects variation across region, training, platform preference, and audience scale, following a theoretical sampling logic aimed at capturing diverse experiences (Given, 2008; Nyimbili and Nyimbili, 2024).
Interviews were conducted between May and July 2025, in person or via secure video platforms. Sessions lasted 60–90 min and were transcribed verbatim with informed consent. Anonymity was ensured, and ethical approval was obtained from the relevant institutional board. Table 1 summarizes key participant attributes.
Overview of interview participants.
Note. Column headings:
The interview design aimed to elicit in-depth reflections on participants’ vocal training, gendered listening experiences, platform strategies, and affective responses to recognition or misrecognition. A conversational format promoted both thematic consistency and narrative openness. The protocol was structured around four thematic domains: (1) voice formation and identity, (2) gendered listening and misrecognition, (3) platformed self-presentation, and (4) community, belonging, and vocal politics. Rather than following a fixed script, the interviews used open-ended prompts to encourage participants to articulate how they made sense of their voices across personal, social, and digital contexts.
Analytical strategy and theme construction
Data were analyzed using reflexive thematic analysis, following Braun and Clarke’s (2006, 2021) six-phase model. This approach emphasizes the researcher's active role in interpretation, prioritizing contextual meaning over measurement. It is particularly suited to studies exploring identity, affect, and negotiation, where individual experiences resist reduction to fixed categories.
The analysis was abductive and iterative—moving between close reading of transcripts, open inductive coding, and theoretically informed interpretation (Braun and Clarke, 2019). Sensitizing concepts such as gendered voice perception, platformed visibility, and legitimacy served as guides rather than fixed templates. Line-by-line coding was conducted using MAXQDA 2020, and analytic memos were maintained throughout to document emergent insights and reflexive positionality. Over 100 initial codes were generated and axially clustered to identify not only what participants experienced, but how they made sense of those experiences.
Three interlinked thematic clusters were developed to capture the layered dynamics of platformed vocal life:
Table 2 provides an overview of the identified thematic clusters, highlighting key subthemes and illustrative codes that guided the interpretive process. Rather than treating these as discrete domains, the analysis understands them as operating in recursive feedback loops. A voice misheard on the basis of gendered expectations may prompt platformed strategies of self-clarification, which in turn affect feelings of legitimacy or marginality.
Thematic coding overview.
Findings
Auditory disciplining: Performing under the gaze of normative listening
Vocality is a socially structured event shaped by cultural expectations about what voices should sound like, and who is entitled to produce them (Azul and Hancock, 2020). For countertenors, whose vocal range disrupts normative associations between pitch and gender, the act of being heard is frequently complicated by the persistent risk of being misheard. This section examines auditory disciplining as the initial and recurrent moment within a broader feedback loop of platformed audibility: the point at which vocal difference is first registered, classified, and rendered problematic through dominant listening frameworks. Drawing on participants’ accounts, the findings show how misrecognition operates not as an episodic misunderstanding but as a structuring condition that shapes subsequent strategies of self-presentation, explanation, and affective labor.
Mishearing and gender attribution
Participants often described how their singing voice was perceived as female—particularly in audio-only contexts—leading to confusion or disbelief when audiences later encountered their visual appearance. While not always malicious, these mishearings revealed deeply internalized associations between high pitch and femininity, and were often experienced as stubbornly repetitive rather than easily corrected. Many listeners assume I’m a woman when they hear my singing—especially when it's audio-only. When they see I’m a man, they’re surprised, sometimes even disappointed. (Interviewee B)
Importantly, such reactions did not dissipate once visual information was introduced. Instead, the encounter between voice and body often intensified the sense of dissonance, producing surprise, disappointment, or suspicion rather than clarification. In this sense, mishearing functioned less as a perceptual error than as an interpretive judgment grounded in binary gender expectations. As Interviewee D explained: I don’t feel heard—I feel misheard. My voice is often treated as a joke, or something abnormal. (Interviewee D)
Here, being misheard signals not only confusion about gender attribution but a broader denial of vocal seriousness. Listening operates as a classificatory mechanism that sorts voices into intelligible or unintelligible categories, with countertenor vocality frequently falling into the latter. Misrecognition thus becomes a mode of social marginalization, positioning the voice as anomalous and undermining its legitimacy before any aesthetic evaluation can occur.
Voice–body dissonance and the labor of explanation
A recurring theme across interviews was the tension between participants’ vocal output and their perceived gender presentation. This perceived dissonance frequently invited intrusive questions or prompted efforts to justify vocal identity, particularly in non-specialist or casual contexts where countertenor singing was unfamiliar. When I tell people I sing as a countertenor, they often ask, ‘So are you really a man or a fake one?’ Sometimes it's a joke, sometimes it's genuine doubt. (Interviewee E)
Such encounters illustrate how vocal difference is not simply heard but interrogated. Rather than being received as an artistic practice, the voice becomes an object of scrutiny that demands explanation. Participants described having to translate or defend their vocality in real time, often anticipating skepticism even before it was explicitly voiced.
Crucially, this labor of explanation did not always resolve misrecognition. Several participants noted that repeated clarification could itself reinforce a sense of abnormality, marking their voice as something that required justification in the first place. In this way, explanatory labor became folded into the disciplining process: an effort to secure intelligibility that simultaneously confirmed the voice's perceived deviation from normative expectations.
Strategic framing and discursive preemption
In response to recurring misrecognition, several participants described preemptively clarifying their vocal identity through video captions, hashtags, or verbal cues. These strategies were not merely informative but affectively protective, intended to shape the audience's frame of expectation before listening began. I usually label my videos with ‘countertenor’ or ‘male alto’ so that audiences are mentally prepared. Otherwise, they feel deceived. (Interviewee A)
Such framing practices represent a form of discursive preemption: an attempt to stabilize interpretation in advance. While participants reported that these strategies could reduce overt confusion, they also acknowledged their limits. Labeling did not eliminate mishearing so much as manage its anticipated fallout, shifting the burden of intelligibility onto the performer.
This dynamic underscores a structural asymmetry in regimes of listening. Normative voices are presumed coherent and require no advance explanation, whereas countertenor voices must be framed, justified, and contextualized to avoid being dismissed or misinterpreted. Discursive preemption thus operates ambivalently—as both a tactic of survival and a reminder of unequal conditions of audibility.
Taken together, these accounts reveal auditory disciplining as more than a momentary perceptual process. It is a recurring and formative condition that shapes how countertenors anticipate being heard, how they prepare for misrecognition, and how subsequent strategies of platformed self- presentation emerge. In this sense, auditory disciplining functions as the triggering node of a broader feedback loop: an initial misalignment that sets in motion cycles of explanation, adaptation, and affective negotiation that rarely bring resolution, but instead reconfigure the terms under which vocal difference is encountered.
Platformed identity negotiation: Vocal presence in algorithmic spaces
While vocal performance in physical settings is shaped by embodied training and immediate audience response, digital platforms introduce a distinct set of mediating forces. Here, visibility is algorithmically curated, content is persistently archived, and recognition is filtered through both machine vision and user expectation (Magalhães and Yu, 2017; Metzler and Garcia, 2024). For countertenors, navigating these conditions involves not only artistic expression but sustained efforts to manage how vocal difference is anticipated, framed, and evaluated. This section examines platformed identity negotiation as a reparative moment within the broader feedback loop of audibility: an attempt to stabilize interpretation following misrecognition, one that remains provisional and frequently incomplete.
Visual signaling and the politics of appearance
Participants consistently emphasized that visual presentation plays a central role in shaping how their voices are interpreted. Several described adjusting their appearance—clothing, posture, framing, and lighting—to align with or subtly offset gendered assumptions triggered by their falsetto register. I avoid presenting too femininely, because with this voice, it's easier to be heard as a woman. There needs to be some balance. (Interviewee C)
Such practices reflect a careful calibration between voice and appearance, aimed at stabilizing the interpretive frame in which vocality is received. Rather than expressing personal style freely, participants described visual self-presentation as anticipatory work: a way of reducing the likelihood of misgendering before listening occurs. At the same time, these strategies exposed a persistent asymmetry. Visual coherence did not guarantee intelligibility; it merely reduced the intensity of interpretive disruption. Self-styling thus functioned less as self-expression than as a form of pre-emptive defense, shaped by the expectation of being misheard.
Captioning, labeling, and narrative framing
To further manage reception, participants commonly employed textual strategies such as hashtags, pinned comments, and disclaimers. These devices were used to clarify vocal identity, reframe performance, and preempt anticipated audience confusion. I always clarify that it's a male voice or falsetto. Otherwise, many viewers will comment, ‘Oh, I thought it was a woman.’ (Interviewee A)
In this context, textual framing became a form of discursive labor oriented toward repair rather than expression. Participants emphasized that captions did not eliminate misrecognition so much as contain its effects, redirecting confusion into more predictable interpretive channels. Explanation thus operated ambivalently: while it reduced repetitive questioning, it also marked the voice as something that required justification. The act of labeling reinforced the sense that countertenor vocality was not self-evident, but conditionally intelligible within platformed publics.
Algorithmic affordances and performance strategy
Beyond human interpretation, participants highlighted the role of platform architectures in shaping visibility. Algorithmic recommendation systems, thumbnail selection, and genre tagging were widely described as powerful yet opaque forces that structured whether a performance would be seen, heard, or ignored. There are so many uncontrollable factors in video streaming, you never know what will go viral. Sometimes a serious performance gets ignored, but a behind-the-scenes clip blows up. So I started editing more actively—using brighter thumbnails, adding subtitles, even splitting one aria into three shorter parts. (Interviewee F)
In response, countertenors adopted a range of tactical practices: selecting repertoire aligned with viral trends, optimizing thumbnails for recognizability, fragmenting performances for algorithmic circulation, and experimenting with multiple accounts. These adjustments extended vocal labor into the technical and affective domains, demanding ongoing calibration to platform metrics rather than musical criteria alone. Yet participants also noted that algorithmic visibility remained unpredictable. Strategic adaptation did not secure recognition; it merely increased exposure, often intensifying scrutiny or misinterpretation.
Across these domains, platformed self-presentation emerged not as a stable solution but as a continual process of negotiation under conditions of uncertainty. Visual signaling, textual framing, and algorithmic adaptation functioned as provisional repairs rather than definitive resolutions. Together, they illustrate how attempts to manage audibility on digital platforms rarely dissolve misrecognition, instead folding it into new cycles of anticipation, adjustment, and emotional labor.
Claiming vocal legitimacy: Affective belonging beyond sonic norms
Digital platforms may offer visibility, but they do not guarantee legitimacy. For countertenors, whose vocal expression challenges dominant sonic norms, recognition is rarely granted automatically and must instead be actively negotiated—through social validation, selective institutional endorsement, and affective connection. This section examines how participants narrate their pursuit of artistic legitimacy and emotional belonging, not as stable achievements, but as fragile and uneven outcomes that emerge within, rather than outside of, platformed regimes of audibility.
The fragility of recognition
Participants frequently reflected on the uncertainty surrounding their vocal legitimacy. Despite technical proficiency, their voices were sometimes dismissed as unnatural, derivative, or gimmick-driven, particularly in contexts where countertenor singing was unfamiliar. One singer noted: Sometimes I feel the audience is waiting for me to fail. They don’t really see me as a legitimate classical singer. (Interviewee D)
Such accounts underscore that legitimacy extends beyond skill or training. It depends on whether a voice is received as belonging within an established aesthetic and cultural frame. Recognition, in this sense, is affectively charged and perpetually provisional—felt, doubted, and renegotiated across performances rather than secured once and for all. Even moments of apparent acceptance were described as reversible, contingent on audience expectations and interpretive context.
Community as counterpublic
In response to repeated misrecognition, many participants described seeking out affirming spaces—fan communities, small online audiences, and networks of fellow falsettists — where their voices could be received without immediate suspicion. These environments offered support that was both emotional and symbolic. As one performer shared: I can find some people who appreciate my voice, sometimes online fans. That really gives me encouragement. (Interviewee F)
Participants framed these spaces as affective counterpublics: contexts in which vocal difference did not require constant justification. However, such recognition was often localized and uneven. While these communities provided respite from dominant listening regimes, they did not necessarily translate into broader legitimacy across platformed publics. Belonging thus emerged as situational rather than generalizable, intensely meaningful yet bounded in scope.
Reframing difference as agency
Several interviewees described a gradual shift from defensiveness toward a more affirmative relationship with their vocal difference. Rather than treating their voice as something that required continuous explanation, they began to articulate it as a medium of expression on their own terms. One singer put it succinctly: I no longer try to prove my legitimacy. I treat my voice as a mode of expression. (Interviewee C)
This reframing did not eliminate misrecognition, but altered how participants related to it. Agency here was less a matter of overcoming dominant norms than of recalibrating self-understanding in their presence. By investing their vocality with personal and artistic meaning, participants carved out a limited form of autonomy—one that coexisted with, rather than replaced, structural constraints on audibility.
The emotional labor of being heard
The negotiation of legitimacy remained emotionally demanding. Participants spoke candidly about shame, doubt, and exhaustion, alongside moments of pride, joy, and embodied presence. One singer captured this ambivalence poignantly: When I sing that note, I feel like I am there—fully present, whole. (Interviewee B)
Such moments highlight vocal performance as a practice of emotional survival as much as artistic labor. Being heard, in this context, meant being momentarily recognized—not only acoustically, but socially and existentially. Yet participants emphasized that these moments were often fleeting, requiring continual effort to sustain, in environments that routinely reassert normative expectations.
Taken together, these narratives suggest that legitimacy and belonging are not endpoints within the platformed circulation of countertenor voices. They function instead as provisional stabilizations within a recursive feedback loop: temporary sites of affirmation that buffer against misrecognition while remaining vulnerable to its return. In this sense, affective belonging does not resolve the politics of audibility but renders it livable, allowing countertenors to persist vocally within systems that continue to test the limits of their intelligibility.
Discussion
The cultural politics of audibility: Misrecognition, legibility, and platformed listening
Audibility does not precede interpretation; it is produced through culturally conditioned acts of listening that sort, classify, and regulate what can be heard as meaningful. Rather than encountering sound as neutral, listeners interpret voices through what Eidsheim (2015) terms “cultural ear training”—a habituated perceptual regime shaped by gender norms, sonic expectations, and embodied assumptions. From this perspective, audibility is not simply an acoustic condition but a cultural achievement, unevenly distributed across different kinds of voices.
For countertenors, whose vocal registers diverge from dominant mappings of masculinity and pitch, misrecognition is not an occasional anomaly but a structural condition of audibility itself. What participants described as “not being heard” did not merely signal mistaken attribution; it pointed to a deeper process of sonic sorting in which their voices were rendered unintelligible, suspicious, or inauthentic from the outset. Misrecognition, in this sense, operates less as perceptual error than as a mechanism through which listening practices enforce normative boundaries.
This form of enforcement is not reducible to individual misunderstanding. As participants recounted being questioned, corrected, or met with skepticism, their experiences revealed how dominant listening frameworks actively compress vocal difference into socially acceptable categories (Butler, 2011). We conceptualize this process as auditory disciplining: a mode of interpretive regulation through which vocal legitimacy is conditionally granted or withheld (Bourdieu, 2018). Auditory disciplining does not merely foreclose alternative ways of hearing; it generates ongoing demands for explanation, framing, and self-justification, transforming audibility into a site of continuous labor.
Digital platforms intensify these dynamics by embedding voices within algorithmic architectures that privilege legibility, coherence, and predictability (Banet-Weiser, 2012). In platformed environments such as TikTok and Instagram, the pressure to “explain the voice” often emerges before any sound is encountered — through pinned captions (“male alto”), hashtags (“#countertenor”), or thumbnail styling designed to preempt confusion. This anticipatory labor exemplifies what we term sonic pre- legitimation: the effort to secure intelligibility in advance under platform logics that conflate recognition with alignment to normative expectations.
Algorithmic mediation compounds, rather than resolves, this tension. Voices are filtered not only through sociocultural assumptions but also through computational systems optimized for engagement through familiarity and interpretive ease. As one participant described,
A video went viral overnight, but the comments were flooded with doubt—‘This has to be fake,’ ‘What am I listening to?’ I felt seen and erased at the same time. (Interviewee D)
Visibility, in this sense, does not guarantee recognition. Instead, algorithmic amplification can intensify scrutiny, exposing non-normative voices to renewed cycles of doubt and correction.
Taken together, these processes point to a broader cultural politics of audibility in which sounding different becomes a structural liability rather than a neutral variation. Countertenor voices, despite technical proficiency, are rendered suspect because they disrupt dominant logics of vocal–body congruence. Voicing thus becomes inseparable from its conditions of reception: to sound is to anticipate being misheard; to perform is to engage in pre-emptive explanation. Audibility emerges here not as a stable capacity but as a conditional status—one that must be repeatedly negotiated within overlapping regimes of listening and platform governance (Puar, 2018).
From this perspective, voice should not be understood as a fixed attribute or isolated expressive act. It is formed relationally, through recursive interactions with listeners, platforms, and the cultural logics that govern recognition. The countertenor case illuminates how audibility itself operates as a regulatory mechanism: granting conditional access to legibility while continually testing the limits of who, and what kinds of voices, can be heard on their own terms. What is ultimately at stake, then, is not simply whether one is heard, but how audibility is produced, policed, and unevenly distributed in platformed cultural life.
Voicing belonging: Affective labor, counterpublics, and situated vocal agency
Belonging, for the countertenor, is rarely immediate or unconditional. It is not enough to sing well; one must also be heard in ways that affirm rather than correct, that recognize rather than categorize. Across interviews, participants described how their voices frequently provoked confusion or scrutiny before admiration—eliciting questions, disclaimers, or interpretive glances that demanded clarification. They were not simply listened to, but evaluated through expectations that left little space for the kinds of vocality they offered (Butler, 2011). Belonging thus emerged not as a baseline condition, but as something unevenly and provisionally achieved.
Within this context, voicing becomes a form of labor—not only to produce sound, but to actively shape the conditions under which that sound might be received. Participants spoke of carefully selecting repertoire that appeared “serious enough” to justify their register, moderating their expressive choices, or repeatedly managing audience reactions through comment deletion, reply restriction, and explanation. This labor was not merely technical or strategic, but affective in nature (Berlant, 2020): a sustained effort to remain vocally present while absorbing the emotional costs of misrecognition, doubt, and exposure. Affective labor here functioned as a mechanism of endurance rather than resolution.
Importantly, this labor did not always unfold in isolation. Participants often located recognition in spaces beyond formal institutions, including peer networks, niche online audiences, and communities of fellow countertenors encountered through short-form platforms. These spaces operated as affective counterpublics—contexts in which vocal difference did not immediately trigger suspicion or demand justification. As participants described, such environments allowed the voice to arrive without preface, offering moments of being “heard without having to prove anything” (Interviewee A). Yet this form of belonging was often spatially and socially bounded, sustained within sympathetic publics rather than extending across dominant listening regimes.
Over time, these experiences enabled forms of agency that were neither oppositional nor fully emancipatory. Several participants described a shift from apologizing for their voices toward embracing them as expressive signatures, or from attempting to “pass” vocally toward exploring the aesthetic possibilities of their range. This reframing did not dissolve normative expectations, but altered how performers oriented themselves toward them. What emerged was a situated vocal agency: an ability to articulate difference within constraint, forged through ongoing negotiation rather than achieved through transcendence.
Still, the affective terrain of platformed voicing remained uneven. Visibility generated connection but also heightened exposure. Belonging, when it occurred, was meaningful yet fragile, always vulnerable to renewed misrecognition and interpretive correction (Berlant, 2020). The effort to remain vocally and emotionally present could be sustaining, as well as exhausting. In this sense, voicing difference did not guarantee recognition. Instead, it enacted a provisional claim to presence—a declaration, however precarious, that this voice matters and persists, even within environments that continue to test the limits of its audibility.
Conclusion
This article has examined how countertenors navigate the intersections of voice, gender, and platformed visibility, not in order to catalogue another instance of vocal marginalization, but to rethink how audibility itself is produced, regulated, and unevenly distributed. Drawing on interview material and thematic analysis, the study has shown that voice cannot be understood as a neutral medium of expression. Rather, it emerges relationally—through listening practices, interpretive expectations, and platform infrastructures that condition what kinds of voices can be heard as legitimate.
One of the central arguments developed here is that digital platforms do not simply amplify voices; they actively participate in governing audibility. Algorithmic systems privilege legibility, coherence, and recognizability, transforming being heard into a conditional achievement rather than a baseline right. For countertenors, whose vocality unsettles dominant mappings of pitch and gender, this means that audibility is continually tested. Practices of visual framing, textual explanation, and anticipatory clarification function less as paths to recognition than as ongoing efforts to manage misrecognition within platformed regimes of listening (Couldry and Mejias, 2019; Eidsheim, 2015).
At the same time, the article has demonstrated that audibility is not governed solely from above. Affective relations—formed through niche publics, peer networks, and moments of resonant listening—can temporarily stabilize vocal legitimacy. Yet these forms of belonging remain partial and fragile. Rather than resolving the politics of voice, they operate as buffers that make persistent misrecognition livable. Vocal agency, in this sense, is neither a story of emancipation nor of simple resistance, but of situated endurance within unequal structures of recognition (Butler, 2011; Lawy, 2017).
Taken together, these findings suggest a shift in how voice should be approached in cultural analysis. Instead of treating voice primarily as expression or identity, this study proposes understanding audibility as a recursive process of governance—one that operates through both cultural norms and algorithmic mediation. The countertenor emerges here not as a niche or exceptional case, but as an analytic lens that renders audible the conditions under which voices become intelligible, suspect, or dismissible in platformed cultural life.
Future research might extend this perspective by examining how other forms of non-normative vocality—across gender, ability, language, or accent—are similarly shaped by regimes of listening and mediation. More broadly, this study argues that to ask who gets to speak is insufficient without also asking how, under what conditions, and at what cost voices are allowed to be heard. Audibility, as this article has shown, is never simply given. It is negotiated, policed, and unevenly sustained—and it is within these negotiations that the cultural politics of voice unfold.
Footnotes
Acknowledgements
The author would like to thank all interview participants for their valuable contributions to this study.
Ethical statements/consent statements
All participants provided informed consent prior to the interviews. The study was conducted in accordance with institutional ethical guidelines.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
