Abstract
This paper examines the impact of Whisper, an open-source AI automated speech recognition (ASR) software, on captioning work at Emory Libraries and the staff who conduct the work. Captions are essential for searchability, discoverability, and user accessibility; however, providing captions has typically been challenging due to the resources required. Whisper shows potential for enabling proactive captioning of digitized content but also raises questions about the impact of ASR-generated captions on the staff engaged in this and similar work. With data collected from a grant-funded captioning project and an ongoing oral history program, this paper examines this impact and observes that while Whisper reduces the most labor-intensive phase of captioning, this does not necessarily entail a reduction in staff time, due to changing demand and the ongoing need for human review in providing accurate, high-quality captions and transcripts. Whisper's overall impact may be on transforming human labor rather than replacing it.
Keywords
Introduction
The remarkable increase in access to large language models (LLMs)/artificial intelligence (AI) models in recent years has prompted individuals, institutions, and organizations to reassess how they design workflows, establish goals, and assess the impact of these technologies on staff experiences. New and innovative programs are seemingly capable of performing laborious tasks in considerably shorter amounts of time than humans require. The boom in machine learning and programs specializing in specific tasks has prompted some to quickly adopt the new technology, while others may be more hesitant. For this case study, we focus on one such tool, Whisper, an open-source automated speech recognition (ASR) system. We explore the impact of integrating Whisper's machine-generated captions on a captioning and transcription program, especially focusing on staff time and satisfaction with the work. What impact does Whisper have on the output capacity of a captioning program and on staff satisfaction with the work? How much human work is needed to generate quality captions? And, does it change how staff do captioning work? Using the Emory University Libraries’ Oral History Program's transcription system as a model and comparison point, this paper examines the impact of machine-generated captions on the feasibility of proactive captioning within the library and on the staff who may conduct the work. Our study posits that, while Whisper outputs reduce the amount of time required to generate a quality caption and transcript, integrating these outputs does not necessarily entail a reduction in staff time allocated to this work, as the demand for captions and transcripts also changes with the increased efficiency of generating them. Furthermore, human labor is still essential, albeit with a shift from text-creation to text-correction, to providing accurate, high-quality captions and transcripts for increased accessibility and discoverability.

Collection criteria matrix.
Background
Emory University is a private research university located in Atlanta, Georgia, with a student population of 15,889 (Emory University, n.d.). There are 15 libraries across campus, most of which contain audiovisual (AV) material. To meet the research and pedagogical needs of students, Emory Libraries has a robust audiovisual digitization program for special collections materials, which facilitates patron requests for specific items, and supports library preservation priorities, exhibitions, and faculty instruction. The program serves preservation and access functions, preserving content on fragile, at-risk, and obsolete carriers and providing access to this content. For the past decade, Emory Libraries has digitized hundreds of content hours annually at a stable rate. The result is that in 2024, our AV repository contained over 5600 video resources and over 15,000 audio resources, totaling several thousands of hours of content. While digitization of audiovisual material for preservation and basic access has been stable, the lack of captions for most of this material is a growing problem. Although any existing captions are preserved in the digitization process, most audiovisual material held in Emory special collections remains uncaptioned, and so the percentage of material without captions has increased year after year.
Lack of captions is a problem for several reasons, most prominently for accessibility. Captions and transcriptions allow people who are hearing-impaired or who may not be as familiar with the language to understand and make use of the material in question. Providing captions also improves resource discovery and access in catalogs and finding aids, and improves the clarity of the resource. Emory Libraries has AV content that has poor audio quality, background noise, and/or crosstalk. These qualities can make understanding the material difficult. Having captions alleviates some of these problems and allows for a greater use and understanding of these difficult materials. Finally, there is growing expectation among users that AV has captions or transcriptions. Evidence indicates that younger generations prefer to view material with subtitles (Ballard, 2023; Mykhalevych, 2024; StageTEXT, 2023). Potential reasons for this shift include younger generations growing up in a digital environment where captions on devices are more common, as well as multitasking between different devices being normal behavior. In addition to the digital environment, younger generations are also living in a more diverse environment with the globalization of popular media. Students are exposed to entertainment and culture from a variety of countries and languages at an early age and are often used to using subtitles to understand the content they consume.
While it is agreed that having captions is something of significant importance, providing them has typically been challenging due to the resources required. At the main Emory Library, the audiovisual digitization program consists of 1.5 FTE (full-time equivalents) staff, whose responsibilities also encompass conservation work; repair, maintenance, and upgrades of equipment, hardware, and software; and facilitation of streaming video content for academic instruction. Given these responsibilities, the volume of material digitized each year, and the amount of time required to produce accurate and usable captions, it has been time-prohibitive to produce captions by digitization staff.
An alternative to in-house caption creation is outsourcing. There are several outsourcing companies that offer caption services, including Rev,1 3Play Media,2 and Scribie.3 Most of these companies offer both human and AI transcription services with a wide range of pricing based on scale, level of human interaction, and language. Emory has a vendor contract for captioning and transcription; however, captioning the AV content that the library digitizes would cost at minimum tens of thousands of dollars annually, with this cost expected to increase as the rate of digitized AV content increases. Solutions employing ASR offer lower rates but introduce data privacy and content accuracy concerns. At organizations such as Emory, where some materials must remain with the organization, the use of such services may not be feasible. A new approach was needed to address the growing caption problem, one that was not extremely taxing in terms of staff time (in-house production) or funding (outsourcing). Whisper, a free and open-source ASR tool developed by OpenAI, presented a potential opportunity for such new approach.

Whisper total processing times.
Literature review
The increased integration of AI in academic libraries over the past few years has prompted the publication of case studies and articles exploring its impact on the library workforce and labor. Among many changes, the evolution of roles and a demand for new skill sets stand out significantly, as AI systems have been adopted to perform roles that were traditionally done by humans, including cataloging (Deng, 2023) and reference. The introduction and growth of AI chatbots in academic libraries has modified librarian reference responsibilities (Lee, 2024; Sanji et al., 2022) whereby chatbots can provide general information to users before shifting the query to library staff if needed, theoretically freeing up time for library staff. This overall shift of integrating AI into library workflows has pushed library staff to learn about AI topics in order to facilitate new directions in academic librarianship as well as learn new skills (Andersdotter, 2023; Lo, 2024). Relating to this is the nature of AI to excel in routine tasks, allowing human labor to focus on more complex work, work that can be accomplished because it builds upon the output of AI tools and systems (Halaychik, 2024). This is a topic that warrants further discussion and investigation. One of the potential challenges relating to AI and labor within academic libraries is that of job displacement. Many aspects of library work are described as having the potential to be accomplished by AI (Frey and Osborne, 2017) or radically changed (Pence, 2022). Additionally, a major issue with AI usage in libraries is the misinformation that it often provides. Software such as ChatGPT can generate inaccuracies and fake journal studies (Marcus, 2022). Overall, the labor-related impact of AI in academic libraries has been one of opportunities and challenges. It has allowed some work to be automated and accomplished more efficiently, freeing up library staff time for other tasks. However, concerns remain regarding the erroneous output of AI software and the impact of AI implementation on human employment.
Regarding the Whisper software specifically, there has been some research on Whisper within libraries: general comparisons of Whisper to other methods of speech recognition (Rodriguez and Brown, 2023); pilot programs using Whisper for specific content, such as educational videos (Rao, 2023) and emotional support conversations (Qu et al., 2024); as well as issues using Whisper, such as continued “hallucinations” (Koenecke et al., 2024) and underperformance of the program for deaf and hard-of-hearing audiences (Zhao et al., 2024). However, further research and investigation is needed to measure the impact of Whisper on academic libraries and their user communities.

Whisper average processing times.
Catalyst captioning and transcription project
In early 2023, we applied for and were awarded a Lyrasis Catalyst Fund grant to explore Whisper as a potential solution for captioning and transcription. Whisper performs multilingual speech recognition and translation while producing time-synchronized caption files and plain-text transcripts. All computing is done locally; no content or information is uploaded to the cloud or sent to a vendor, avoiding the introduction of data privacy concerns associated with many ASR solutions. Our preliminary testing prior to the grant found that the accuracy and efficiency of Whisper's output, combined with its low cost to operate and its ability to accommodate data privacy needs, suggested it could be a feasible solution to caption digitized content on a large scale within our digitization workflow. Our initial test was conducted on a small set of video files from a collection in which all videos featured a single speaker, with good audio quality, very little background noise, and no music or non-speech sound content. We chose this collection as a baseline for testing in the anticipation that these audio characteristics would be relatively less challenging for Whisper to interpret, and we used word error rate (WER) as our metric for accuracy. WER is the rate of erroneous word substitutions, insertions, and deletions in comparison to a corrected caption file. It is a widely used metric in assessing ASR software, although it does have some weaknesses, in that it does not assess the degree of error within a word or the degree to which the error impacts semantic meaning. In this initial test, Whisper processed content at a rate of 25–30 min per content hour, and the results achieved an average word error rate (WER) of 1.06%. This is better than the WER achieved by other available ASR solutions and is comparable with the accuracy rates guaranteed by human transcription services.
Building on these initial findings, our grant project tested Whisper across a larger and more varied sample of AV content to further assess its accuracy and determine the feasibility of deploying it within our AV digitization workflow. We ran 250 h of audio and video recordings through Whisper. The content was chosen from five collections at Emory Libraries, selected to provide a representative sample of collections with AV material in order to assess Whisper's performance with a variety of sound characteristics that might be encountered in the collections. The material selected included oral history interviews with students at Emory; medical lectures from Emory's School of Medicine; a public-access television show discussing politics and current events; a politician's press conferences; a radio program featuring recorded civil rights speeches and live discussion; and a public-access television show featuring musical performances, interviews, and sketch comedy. These are all primarily English-language materials, although we are interested in Whisper's performance with additional languages as a future avenue of research. The selected materials varied in several ways, including number of speakers; subject matter and use of subject-specific vocabularies; formality of sentence structure; use of regional dialects, slang, and vernaculars; presence or absence of background noise, music, and non-speech sound; and the quality of audio recording. This allowed us to assess Whisper's overall performance, its equity of performance across content with differing characteristics, and the impact of the stricter formatting requirements for captions, as compared to transcripts, on Whisper's output and the editing process (Figure 1).
We ran this material through Whisper on a PC and on a Mac through Whisper.cpp, a C++ version of Whisper which allows the software to make use of the Mac's M chip for greater processing speed. The selection of hardware and software was based on findings in our initial testing, which demonstrated that robust graphics processing unit (GPU) capabilities are essential to efficient processing of the files. For each media file, Whisper generated a time-stamped caption file and an untimed plain-text transcript file. On both the PC running Whisper and the Mac running Whisper.cpp, the files were generated at a rate of approximately 16–18 min per hour of media content (Figures 2 and 3).
We initially hired seven students in October 2023 to review and edit the files generated by Whisper and produce a corrected version. The students were divided into two job classifications: project assistants, who conducted a first review focusing on content accuracy; and project specialists, who conducted a second review focusing on formatting and consistency. The students came mostly from a humanities background, with some in the social sciences and medicine. A few students had had prior experience with transcription as part of their academic work, but the majority were new to transcribing. Most of the undergraduate students were classified as project assistants unless they had had previous experience with transcription; those with previous experience, along with the graduate students, were classified as project specialists. A further five students were hired in 2024, such that over the course of the grant, a total of 12 students (nine undergraduates and three graduate students) worked on the project. Three student editors had graduated from Emory, and two students were unable to continue beyond May 2023 due to internships. As of November 2024, we had seven active student workers (Figure 4).
We provided the students with an orientation to our caption editing software, a style guide addressing the purpose of captioning and transcription work and commonly encountered editing situations, and a guide to each collection in the project, which provided context, topic-specific resources such as controlled vocabularies and name authority resources, and information on any known tendencies of Whisper's output from that collection. After completing orientation, students worked primarily independently, with project management and communication through Airtable and Microsoft Teams.4 All editing and training resources were posted on a Teams channel dedicated to the project. Students scheduled their shifts through a spreadsheet accessed on this channel, posted editing questions through a “Questions Log” document in the same space, and communicated with each other and us through Teams chat for more immediate questions. The project's progress was managed in Airtable, with students selecting media items to work with, recording the total editing time for that file, and documenting any observations about Whisper's performance for that item. Although the majority of the project management and communication were online, all editing work was completed in person in the Robert W. Woodruff Library building at Emory University (Figure 5).
Our workflow included two rounds of editing: the project assistants conducted a first-round review and editing of the Whisper output, focusing on accuracy and clarity, and then the project specialists conducted a second-round review of the file, focusing on quality control (catching any inaccuracies that were missed in the first round), conformance to broadcast formatting standards specified by Federal Communications Commission regulation CEA-608 for closed captioning of television, and overall consistency of editing decisions from file to file. For video files, a caption file was edited and formatted to match broadcast requirements for appearing over video on a screen, with accompanying timestamps to synchronize the appearance of the text on screen with the dialogue as it is spoken. A plain-text transcript without timestamps, intended to be viewed side-by-side with the video content, was then generated from the edited caption file. For audio files, only the plain-text transcript was created, edited, and formatted in paragraphs and without timestamps. The editing was accomplished using CADET, a free caption and description editing software created by the National Center for Accessible Media at GBH, a public broadcasting station (GBH, n.d.). Our project's approach to transcribing speech and non-speech sound followed the guidelines of the Described and Captioned Media Program (n.d.).
Deliverables for the project included hardware and software configuration recommendations; a standardized workflow to generate, process, edit, and deliver captions and transcripts; a style guide with subject-specific guidelines to standardize editing and formatting, incorporating consultation with accessibility advocates and best practices for captioning and transcription; and the establishment of benchmarks for assessing processing time, cost, accuracy rates, and editing time. We sought to understand how Whisper handles audio and video files with varying sound and content characteristics, and also to understand the human labor needed to turn Whisper's output into captions and transcripts that are useful and accurate for end-users.
In the first 14 months of the project, the student captioning team completed first-round edits for 340 media files, totaling 189.5 h of content. This work was accomplished over 809 h, as self-reported by the students, suggesting an average of 4.27 h to edit one hour of media content across all collections. We compared the edited files with Whisper's output using a Python script and calculated that Whisper generated captions and transcripts with an average WER of 5.24 to 24.01%, depending on the collection. These averages are different from the results in our initial test; we drew our initial set of test files from the Surgical Grand Rounds collection, which has remained the collection with which Whisper is most accurate. Of the 73 Surgical Grand Rounds files processed, Whisper returned two outlier files, one captioned entirely in Welsh and the other consisting of just numbers and no words. If these two results were not included in the WER calculation, the average WER for this collection would be 2.57%. This number is more in line with the results from the initial test, although somewhat decreased with the larger number of content hours in the main data set. Other collections have included more varied audio characteristics and have produced correspondingly less accurate results than with Surgical Ground Rounds (Figure 6).
While it is difficult to fully measure the impact of accuracy or caption-formatting requirements on editing time, our student team reported that correcting Whisper's output for caption formatting took up a greater portion of their editing time than correcting for content inaccuracies, resulting in much faster progress editing audio files compared to video files. We have further observed a substantial difference in editing time between collections for which Whisper struggled with formatting and sentence structure, such as oral history interviews, and those for which it produced well-formatted and structured output, such as medical lectures.
Regarding the human impact of this work, students generally reported satisfaction in engaging with the media content and the work. Students tended to work in sessions of two hours at a time, to reduce physical discomfort associated with the time spent at the workstations; although in summer sessions, students chose to work in longer shifts. Students typically completed work on a video recording in a few sessions, and could complete work on an audio recording in a single session; this rate of completion contributed to the students’ sense of accomplishment and satisfaction with the work, in addition to facilitating their engagement with a greater variety of stories and content. Students’ weekly hours varied somewhat over the course of the academic year, with less work scheduled during midterms and finals, and more work tending to occur early in the semester or after completing finals. Student burnout has not been reported. Most students continued on from the previous academic year; they have only left the project due to graduation or internships.
Whisper's impact through November 2024 has been the creation of quality caption files that would not have been created otherwise, as well as an increase in student employment opportunities. Prior to this project, captioning was facilitated on an as-requested basis, with no staff or student time regularly allocated. With the demonstrated output from the project, we expect to continue captioning on an ongoing basis, resulting in the creation of several ongoing student jobs within the Preservation Department. While our collection of data on the qualitative and quantitative impact to human labor continues, our observations during this first year suggest that because Whisper was completing the most labor-intensive portion of the workflow in generating the initial text from scratch, the human workers engaged in captioning were able to focus on the more creative aspects of the job, increasing their satisfaction with the work.

Student workers.

Airtable in use.
Emory Oral History Program transcription
The workflow described above mostly followed the overall model of the Emory Oral History Program (EOHP). The EOHP's transcription system provides a comparison point and a perspective on the impact of machine-generated transcription on previously established transcription workflows within the library and on the staff who conduct the work. The EOHP was established in 2018 in the Emory Libraries to create, curate, and make oral history collections accessible. Following the best practices in the oral history profession, the EOHP established a workflow to process oral history interviews and provide user access as stipulated by interview participants. Within the processing workflow, the program creates interview transcripts that are intended to be used as secondary document sources alongside the original audio or audiovisual object. In this case, the transcripts are not intended to represent original source documents; the spoken word does not transfer seamlessly into written form nor are unspoken forms of communication easily translatable into text. While acknowledging these translation limitations, the EHOP holds that providing an accurate transcript to be used in concert with the oral history interview increases overall accessibility by providing searchable text and clarifying regional dialects, keyword lists, and the spelling of names and places. The EOHP produces transcripts that are different than captions, which are designed to be overlaid on top of visual content and require formatting in small segments of time-stamped text. EOHP transcribers follow a style guide designed by the program to ensure consistency among the many subjective choices in turning spoken word into written text. Emphasis is placed on respecting the words and vernacular of the narrators, including colloquial language. Transcribers are also responsible for entering timestamps when speakers change or when the narrative subject transitions, as well as formatting to apply light grammatical structure (commas, periods, and paragraph divisions).
The EOHP established two different workflow systems to measure the time, labor, and human experience of creating transcripts of oral history interviews. The first model was the EOHP's standard procedure to create transcripts “in-house” relying on student and staff labor through three distinct steps: an initial first pass to create a full written representation of the oral interview; a second pass in which a different transcriber verifies content and reviews stylistic choices; and a third and final pass to address outstanding questions and confirm metadata. For this paper, the workflow described above will be referred to as the “EOHP Human System.” The EOHP only slightly modified its transcription workflow to integrate Whisper-generated scripts, referred to as the “EOHP Whisper System.” In this workflow, staff members were responsible for processing 20 oral history interviews through Whisper's ASR system. The generated scripts were then issued to student transcribers, who inserted the text into an EOHP transcription template that contained only a standard descriptive cover page. The student was also provided an audio or audiovisual copy of the interview and tasked to create a “first pass” transcription by directly editing Whisper script rather than typing text on a blank page. Beyond this phase, the program maintained the same workflow as the EOHP Human System described above, including a second pass and a final verification. These two workflows provide a worthwhile comparative to assess the time and labor required to create an EOHP transcript, and to reflect on the human experience in working with Whisper-generated script compared to transcribing without an AI-generated script.
From spring 2023 until spring 2024, the EOHP employed 14 students and staff in the creation of oral history transcripts. They completed 41 transcripts from interviews that averaged 50 min in length per interview. EOHP interviews vary considerably in content, vernacular, vocabulary, and rate of speech. Accordingly, the transcription time for each interview varied as well, even for those of comparative recorded run time. Of the 41 transcripts, 30 were created using the EOHP Human System, and 11 were created through the EOHP Whisper System.
On average for the EOHP Human System, the first pass required 6 h of labor per hour of content, the second pass required 2.5 h per content hour, and the final verification took .75 h per content hour, totaling 9.25 h of labor to transcribe one content hour. In the EOHP Whisper System, on average, the first pass required 4 h of labor per hour of content, the second pass required 1.5 h per content hour, and the final verification took .75 h per content hour, totaling 6.25 h of labor to transcribe one content hour. This is roughly a 35% decrease in time taken over the fully human derived workflow.
The most impactful takeaway from the perspective of the EOHP is that integrating Whisper scripts significantly reduced the labor hours required to generate oral history transcripts in alignment with the program's guidelines. There is potential for further labor reduction by modifying the workflow. Student transcribers working on a Whisper output to create the “first pass” noted that their main tasks focused on formatting and editing, with scant direct text creation. Students responsible for the “second pass” noted how the text required far fewer corrections or edits, and their efforts instead focused on verification duties that were typically the responsibilities of the final verification pass. This feedback indicates that the workflow can be optimized by folding the second pass and verification pass stages into one instance, eliminating a step in the EOHP's transcription workflow and further decreasing the time required.
Ten of the EOHP's student transcribers reflected on their experiences creating transcripts using the EOHP Human System and the EOHP Whisper System. Reflecting on the EOHP Human System, three students identified additional benefits from transcribing to a blank page, citing improved listening skills and interview strategies, and insights on narrative analysis. Two students articulated their appreciation for the close listening required in creating the original transcribed text, and that the process offered a quiet space to abscond from a busy day. Providing critical feedback, most students highlighted their frustration at the time required and pace of creation when transcribing to a blank page. They noted that working intermittently for multiple shifts on the same interview felt tiresome. Slowing the recording play speed to maintain their typing pace reinforced a sense of slow progress. Further, frequently rewinding a section of an interview to ensure accuracy instilled a sense of repetition and stagnation. These observations led one student to describe the EOHP Human System's first pass as “grunt work,” compared to the EOHP Whisper System.
EOHP's student transcribers working with Whisper outputs to generate transcripts provided mostly positive feedback. All students noted that having interview text as a starting point increased their sense of productivity even if the output required substantial editing and formatting. They described the EOHP Whisper System as more intellectually engaging because the bulk of the work required editorial choices. Half the respondents shared a sense of contentment that they mostly avoided the need to modify the play speed of the interview and were also able to reduce pauses and rewinds. The EOHP student transcribers identified changes in their work tasks when working with Whisper-generated outputs. One area of change included basic formatting tasks, including font size and style, line breaks, sentence structure and grammar, and paragraph breaks. They also noted the common repetition of phrases or questions that needed to be removed. Another area of common tasks focused on the content itself and assessing if Whisper correctly identified words. A few anticipated issues emerged, including frequent incorrect identification of acronyms and the spelling of names and places. Furthermore, student transcribers also identified occasional hallucinations of words and phrases that were not in the original recording. Overall, each student transcriber emphasized that these editorial responsibilities composed the bulk of their tasks, and that human verification is still essential to generate an accurate written representation of the oral history interviews. “It's not great, but it's not bad” captured the assessment of one EOHP student transcriber.
The comparison between the EOHP Human and EOHP Whisper systems provides a useful model to assess worker preference and transcription metrics. The statistics suggest that integrating Whisper outcomes reduced the number of human labor hours required to generate an EOHP quality transcript. Moreover, student transcribers collectively preferred to perform the workflow of the EOHP Whisper System, especially omitting the responsibility of creating a first pass version of a transcript typed on a blank page. What remains uncertain is whether student workers would report these positive sentiments of using Whisper outputs if there were no comparative model. Speculatively, beyond assessing how student workers enjoy the tasks, questions about changes in skill building and learning between the two models are yet to be explored. Beyond the economics of labor costs, there may be a decline in educational or experiential benefits by relying on Whisper outputs and centering editing and formatting as the primary worker responsibilities.

Whisper accuracy rates.
Conclusions
Broadly, integrating Whisper to generate captions and transcripts presents nuanced outcomes regarding human labor in the library. The EOHP does not have a robust backlog of recordings that require transcription. Thus, from this program's perspective, integrating Whisper would reduce the demand for labor to create transcripts over time, and reduce the required number of student employees working on transcription. However, the EOHP is an outlier compared to the audiovisual captioning and transcription needs of Emory Libraries overall, the collections of which contain thousands of hours of digitized audiovisual content requiring captions and transcription, with hundreds of hours more ingested each year. Two simultaneous processes call attention: first, integrating Whisper clearly reduces human labor requirements to generate transcripts and captions; second, human labor is still required to improve the quality of Whisper outputs to reach an acceptable level for the majority of use cases. If the standard quality of Whisper outputs improves further, one can anticipate that the required human labor for a single transcript will decline in response. However, at the same time, as the feasibility of providing captions and transcripts increases, and thus their availability, the demand for captions and transcripts and the expectation that they are available will also increase. As a result of our Catalyst project work, our team has been approached about the possibility of implementing Whisper as part of other AV-related service points within the library, and regarding our digitization and preservation workflow, we anticipate that it will be necessary to continue or increase the level of work in this area to meet changing norms and expectations. Access to tools like Whisper may have further useful applications within library workflows, such as in support of research data management or faculty-led research and digital humanities projects; thus, the demands for human captioning and transcription labor may shift within different areas of the library, decreasing within some workflows but increasing within others. The overall impact of Whisper within the EOHP and Catalyst projects suggests that the paradigm for considering the impact of AI technologies, at least for ASR solutions, may not be one of replacing human work but of transforming that work.
Footnotes
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by a 2023 Lyrasis Catalyst Fund award.
