Abstract
Introduction
Podcasts are a popular medical education tool, especially for the review of journal articles, but production requires significant resources. The objectives of this study were to determine the feasibility of a trained artificial intelligence (AI) model generating well-received educational podcasts based on journal articles and to assess listener comprehension with knowledge-based questions.
Methods
Google Gemini 2.0 Flash AI was trained on 12 Wilderness and Environmental Medicine journal articles to generate 3 distinct educational podcasts (4 articles per podcast). These podcasts featured 2 AI-generated host voices in a journal club format. Participants from the Wilderness Medical Society Student/Resident Committee, the Virginia Tech Carilion Wilderness Medicine Fellowship Program, and a convenience sample of students and residents were randomly assigned to listen to 1 podcast, blinded to its AI origin. They then completed a 5-point Likert perception survey and a knowledge assessment with continuing medical education-style multiple-choice questions based on the discussed articles.
Results
Thirty-one participants completed the study. Participant perception was highly positive: 87.5% agreed or strongly agreed that the content was accurate and relevant (mean Likert score 4.37), and 81.25% agreed or strongly agreed that the podcasts aided journal article review and learning (mean Likert score 4.14). The mean knowledge assessment score was 88.7% correct (SD 5.0%).
Conclusions
Despite limitations, including a small sample size and lack of a control group, AI-generated podcasts were positively received and effectively conveyed educational content, as evidenced by high knowledge assessment scores.
Introduction
Audio lectures and podcasts have become an integral and increasingly popular component of modern medical education. Their asynchronous, on-demand nature allows learners to engage with educational content during commutes, exercise, or other activities, integrating seamlessly into the demanding schedules of medical students, residents, and practicing physicians.1–3 A scoping review by Kelly et al found that learners value podcasts for their portability, efficiency, and ability to combine education with entertainment. 1 Similarly, a systematic review by Caldwell et al confirmed the escalating rate of publications on this topic, particularly since the COVID-19 pandemic, with emergency medicine being a leading specialty in podcast adoption. 2 These reviews indicate that podcasts are frequently used to review journal articles and stay current with evidence-based practices.
Despite the popularity and educational utility of high-quality medical podcasts, their production is a resource-intensive endeavor. It requires a significant investment of time for content research and scripting, technical expertise in recording and editing, and often financial resources for equipment and hosting. 4 This high barrier to entry can limit the creation and dissemination of niche-specific content, such as that for wilderness medicine.
Generative artificial intelligence (AI) presents a potential solution to these production challenges. Recent advancements in large language models and text-to-speech synthesis have enabled the automated creation of sophisticated, human-like content. Studies have begun to explore the use of generative AI in creating educational materials, from personalized text-to-podcast conversions to engaging cinematic clinical narratives.5,6 These tools have the potential to democratize content creation, allowing for the rapid, low-cost production of educational podcasts tailored to specific learning objectives.
However, the feasibility and effectiveness of using AI to generate educational podcasts for a specialized field such as wilderness medicine remain unexplored. This pilot study aimed to address this gap by determining whether a trained AI model could generate a well-received educational podcast in a journal club format based on articles from Wilderness and Environmental Medicine (WEM). We also sought to assess listener comprehension of the AI-generated podcasts through knowledge-based questions.
Methods
Study Design
This was a pilot study using a cross-sectional survey and knowledge assessment to evaluate the feasibility of AI-generated educational podcasts based on wilderness medicine peer-reviewed journal articles.
Setting and Participants
Participants were recruited from 3 distinct groups with an interest in wilderness medicine: the Wilderness Medical Society (WMS) Student/Resident Committee, physicians and advanced practice providers involved with the Virginia Tech Carilion Wilderness Medicine Fellowship Program, and a convenience sample of medical students and residents known to the authors. Participants represented multiple levels of training and medical specialties. Participants from multiple institutions were invited. They were not informed that the study subject was specifically AI podcasts.
Intervention
Google's Gemini AI Studio (Google LLC, Mountainview, CA) was trained using 12 full-text journal articles published in WEM (Table 1). This was the free version of Gemini, only requiring a Google account. Training in this case refers to the use of full-text PDF articles as the foundational material to inform the AI's responses to a given prompt. A prompt is the written instructions or input provided to the AI. The AI then was prompted to create 3 distinct educational podcasts, each ∼20 min in length. Each podcast was designed to review the key findings of 4 of the selected articles. The AI was instructed to adopt a journal club format, featuring a conversational dialogue between 2 distinct AI-generated host voices.
Wilderness & Environmental Medicine Articles Used to Train Each Podcast Episode.
AI Prompt
Once trained on the WEM articles, the AI was instructed with the prompt: “Perform a thorough journal club-style review of these 4 wilderness medicine-related publications, aimed at a physician level of education. Make it a podcast with 2 hosts, some entertaining and witty banter back and forth to keep it interesting, but still professional. Make it around 20 min in length.” All 3 podcasts were output as .wav files and were between 18 and 22 min long.
Data Collection and Main Outcomes Measured
A total of 31 participants were randomly assigned to listen to 1 of the 3 AI-generated podcasts. To prevent bias, participants were blinded to the AI origin of the podcasts. After listening, participants filled out an online questionnaire. The primary outcomes were listener perception and knowledge comprehension.
Listener Perception
A perception survey was administered that included demographic questions and frequency of listening to podcasts to learn about medical topics. The key survey questions were developed by modeling them after previously validated instruments for assessing educational podcasts, specifically the Questionnaire for Assessing Educational Podcasts and the Student Satisfaction with Educational Podcasts Questionnaire.8,9 These frameworks assess dimensions such as content adequacy, relevance, ease of use, and value as an aid to learning. Based on these established domains, our survey used 10 key statements assessed on a 5-point Likert scale (1=strongly disagree; 5=strongly agree):
“The podcast introduced me to new topics.” “The podcast reinforced information I have already learned.” “The content of the podcast is accurate and relevant to wilderness medicine.” “The podcast addresses common challenges encountered in wilderness medicine.” “The presentation format of the podcast is good (the speakers’ engagement, audio quality, clarity, etc).” “The length of the podcast is appropriate for understanding the content.” “The podcast made reviewing the journal articles more enjoyable.” “The podcast was a good aid to review journal articles and learn new information.” “The podcast is something I would recommend to others to learn topics in wilderness medicine.” “The podcast is something I would like to continue listening to.”
Knowledge Comprehension
Twelve questions for each podcast (4 from each article) were developed by the study authors directly from the original source article. Questions that were not covered in the podcast were then eliminated. This was determined by the authors listening to the full podcast and reviewing the questions. A knowledge-assessment test was administered after respondents listened to the podcast, consisting of 6 to 10 continuing medical education-style multiple-choice questions. Questions were like WEM questions and did not assess multilevel comprehension. Descriptive statistics, including means and standard deviations, were used to analyze the quantitative data from the surveys and knowledge assessments.
Results
A total of 31 participants listened to 1 of the 3 AI-generated podcasts and completed the subsequent survey and knowledge assessment. The demographic breakdown of the participants is detailed in Table 2. The cohort was composed primarily of attending physicians (n=13, 41.9%) and advanced practice providers (n=7, 22.6%), with representation from medical students, residents, and paramedics. The average Likert rating for “I frequently use audio sources or podcasts to learn about medical topics” was 3.6, with 48% of participants responding “strongly agree.”
Participant Demographics (N=31).
Listener Perception
Participant perception of the AI-generated podcasts was overwhelmingly positive, as measured by a 10-item survey using a 5-point Likert scale. The mean Likert score, percentage responding “agree” or “strongly agree” for each podcast, and the total are listed in Table 3. The mean scores for each survey question are detailed in Figure 1.

Mean participant perception scores for artificial intelligence-generated podcasts.
Listener Perceptions of Each Podcast and Aggregate.
a Analysis of variance.
Overall, the responses indicated a high level of satisfaction and perceived utility. Six of the 10 questions received a mean score of 4.0 or higher, signifying that participants, on average, agreed or strongly agreed with the statements about the podcasts. Notably, participants strongly endorsed the educational value of the content, agreeing that it was “accurate and relevant” (mean score 4.37) and that it “aided in my review and learning of the journal articles” (mean score 4.14). Even the questions with the lowest mean scores still indicated a neutral to positive perception, with no average score falling into the disagree range.
To assess whether perceptions were consistent across the 3 different podcast versions, an analysis of variance was conducted for each survey item, as seen in Table 3. For half the questions (5 of 10), there were no statistically significant differences in ratings among the 3 groups (P>0.05). This suggests that the positive reception was largely uniform regardless of the specific set of articles presented. For the other 5 questions, a statistically significant variation was found between the groups (P<0.05), suggesting that perceptions of certain aspects, such as specific design elements or engagement, may vary depending on the content. Despite this variability, the overall trend of positive perception was maintained across all podcasts.
Knowledge Assessment
Participants demonstrated strong comprehension of the material presented in the podcasts. The average score on the knowledge-assessment test was 88.7%, with an SD of 5.0% (Table 4).
Postpodcast Comprehension Questions.
Conclusion
This study demonstrates the feasibility of using generative AI to create high-quality, effective educational podcasts for a specialized field such as wilderness medicine. Our findings indicate that AI-generated podcasts were well received by a diverse group of learners and served as an effective tool for knowledge transfer, as evidenced by high scores on the comprehension assessment. The positive perception of the content's accuracy and educational utility, despite its synthetic origin, suggests that AI can successfully replicate the engaging and informative journal club format that is popular in medical education. This may add to already available wilderness medicine education, such as the Wilderness Medicine Podcast.
Educators should be careful to thoroughly review all AI-generated content for accuracy prior to using a learning tool. For these 3 podcasts, the authors reviewed the content and determined it to be accurate. No additional podcasts or revisions were needed. However, this study is not sufficient to make any inferences about the accuracy of AI output that lacks human review.
The results align with existing literature on the benefits of podcasting in medical education while offering a novel solution to the primary barrier of resource-intensive production.1–4 By automating the process of synthesizing and presenting complex information from journal articles, AI offers a scalable and cost-effective method for creating current, accessible educational content. The high knowledge-assessment scores (88.7%) support findings from other studies, such as that of Wolpaw et al, which demonstrated that podcast learning can be equivalent or superior to traditional textbook reading for knowledge acquisition and retention. 7 This may not apply to all listeners based on generational differences.
Limitations
This study has several limitations. The small sample size (N=31) and the use of a convenience sample may limit the generalizability of our findings. Respondents’ age and sex were not collected and could influence results. Furthermore, the study lacked a control group, such as participants who read the articles directly or listened to a human-produced podcast on the same topic, such as the Wilderness Medicine Podcast. Because this was a pilot study addressing the feasibility of an AI podcast, no control was used. This makes it difficult to draw definitive conclusions about the relative efficacy of the AI-generated format. The assessment also only measured immediate knowledge recall, not long-term retention.
Only a small number of questions were generated for each podcast. The questions were generated based on article content, not podcast content. Questions with content not covered in the podcast were not included in the final knowledge-assessment tests, leading to a small and nonuniform number of questions.
Future Directions
Despite these limitations, this study provides a strong foundation for future research. Larger randomized, controlled trials are needed to compare AI-generated podcasts against traditional learning modalities and human-produced podcasts. Comparison with wilderness medicine-specific podcasts, such as the Wilderness Medicine Podcast, will be the next step. Future studies also should explore the potential for personalization, as investigated by Do et al, 5 by tailoring content to a learner's specific knowledge gaps or interests. Investigating long-term knowledge retention and conducting a cost-benefit analysis of AI vs human production would further elucidate the value proposition of this technology.
AI-generated podcasts offer a promising and powerful approach to wilderness medicine education, making it easier, more engaging, and more efficient for learners at all levels to stay current with the latest literature. AI audio formats should be considered as a potentially useful tool for use in wilderness medical education with appropriate oversight and review.
Previous Presentation
Presented at the Wilderness Medical Society Summer Conference Virtual, Three-Minute Thesis Presentations, July 21, 2025.
