Abstract
Citizens’ Assemblies and other deliberative mini-publics (DMPs) are institutions designed to foster democratic deliberation. AI systems, especially large language models (LLMs), have been suggested as tools to scale up the breadth and depth of DMPs. We distinguish six core functions within democratic deliberation – representation, perspective-giving and taking, inquiry, co-creation, integration, and decision-making – and point out some potential benefits of AI in each of these functions. Our case study presents the results of the Citizens’ Assembly on Energy, organised in Finland in 2025, in which an LLM was used in inquiry, co-creation, and integration. We found that while LLMs have some shortcomings in terms of inquiry, especially in creating questions for experts, they can ease the workload of human facilitators and deliberators in co-creation and integration. Our findings contribute to the emerging debate on the optimal division of labour in deliberative processes between human deliberators, facilitators, and generative AI.
Introduction
In recent years, democracy scholars have debated on the challenges and the promise of artificial intelligence in democratic decision-making (Goñi, 2025; Innerarity, 2024; Landemore, 2022; Wihbey, 2024). Commonly stated risks include, for example, the detrimental impact of AI systems on public knowledge and opinion formation. While many authors emphasise such negative impacts of generative AI on democracy (e.g. Coeckelbergh, 2023), AI can arguably also open new possibilities to nurture deliberation and respectful discussion among citizens (Summerfield et al., 2024).
Citizens’ Assemblies (CAs) and other forms of deliberative mini-publics (DMPs) are institutions that aim to increase lay citizens’ opportunities for participation, deliberation, and influence in decision-making (Elstub and Escobar, 2019). They have spread across established democracies during the last two decades, and they have been applied in many domains, from environmental politics to moral questions pertaining to euthanasia and abortion (OECD, 2020). DMPs consist of randomly selected citizens who learn and deliberate on policy topics, usually producing recommendations as policy advice (Setälä and Smith, 2018). Compared to other democratic innovations, the strength of CAs and other DMP formats is particularly in their ability to enhance collective opinion formation through processes of inquiry, exchange of viewpoints, co-creation, and integration (Jäske and Setälä, 2020).
There are already various proposals of using AI technologies to facilitate deliberative processes in DMPs and help scale up their impact (e.g. McKinney, 2024). The purpose of this study is to identify the tasks in which AI technologies can facilitate inclusive processes of citizen deliberation that are consequential in the wider democratic system. We are aware of the criticism of integrating AI technologies with DMPs as ‘techno-solutionism’ that can be associated with simplistic and depoliticised views of democracy (Oleart and Palomo, 2025). Moreover, critics argue that DMPs themselves are ‘shortcuts’ that are harmful of democracy, especially if they replace deliberation and participation in the broader public sphere (Hammond, 2021; Lafont, 2015). It is also necessary to take into account ethical issues and anti-democratic features related to AI technologies themselves. AI companies and data centres have faced criticism for the inequality of their organisations, poor working conditions, and excessive use of energy (Muldoon and Grant, 2024). Furthermore, application of general-purpose LLMs in any democratic process requires serious consideration of the built-in biases of LLM training data and algorithms that can perpetuate existing inequalities in society (Ranjan et al., 2024) and their opaqueness that makes it impossible to scrutinise their analyses. However, the point of departure in this article is that AI technologies can also be used to support human reasoning and communication, not just within DMPs, but also to help DMPs to have a role in ‘deliberation-making’ (Curato and Böker, 2016), in other words, to enhance deliberation in the wider democratic system.
Organising informed, inclusive, and respectful deliberation in DMPs requires effort and resources (Curato et al., 2017). At the same time, DMPs often lack legitimacy and impact because policymakers and large segments of the citizenry remain excluded from and even uninformed of them (Lafont, 2019). There are various proposals for enhancing ‘deliberative impacts’ of DMPs (Dryzek and Goodin, 2006). These include increasing the number of participants (Fishkin et al., 2025), enhancing the visibility of DMPs among the wider public, careful coupling of DMPs with policy processes, and developing the capacity of DMPs to co-produce policy recommendations with justifications (Setälä, 2017). There is empirical evidence that the contents of recommendations and justifications are crucial for the legitimacy of DMPs among the wider public (Goovaerts et al., 2025; Himmelroos et al., 2026). However, the requirement to co-produce written factually ground and well-justified statements makes the task of organising DMPs more challenging, especially when large number of deliberators are involved.
This study addresses these theoretical and practical challenges regarding the use of AI in DMPs in two different ways. First, it introduces a theoretical framework which distinguishes six core functions in DMPs and analyses the potential of AI in each of these functions. Second, our study also presents a unique case of a CA, in which a general-purpose LLM was used to assist information gathering and co-writing in the deliberative process. Our empirical research questions are: Can LLMs assist the CA in building its competence on the subject matter? And can they support the CA in its collaborative task, that is, writing the common statement? We start with a theoretical discussion on the potential of AI technologies in different functions of democratic deliberation. Thereafter, we introduce our case, the Finnish CA on Energy, and explain our data and methods for analysis. Our analysis is based on surveys to participants and facilitators, material produced by DMP participants, as well as discussion transcripts. Our empirical findings are presented in two parts, focusing first on the inquiry function of the CA and then on the co-creation and integration functions of the CA. To conclude, we discuss our findings and their implications for research on democratic innovations and AI technologies developed for the purposes of public deliberation.
Theory: LLMs and six functions of democratic deliberation
Artificial intelligence is a broad, multifaceted concept that may refer to technological constructs, particular methodologies, or a field of study (König et al., 2022). Democratic processes can be designed so that participants interact with various AI systems, such as ChatGPT, that are digital artefacts built on one or several AI technologies, capable of learning and processing input data into outputs. Natural language processing (NLP) technologies such as BERT and GPT models were developed for understanding and classifying language and textual input, making them particularly promising from the perspective of analysing the contents of deliberation and deliberative outputs (Gardazi et al., 2025; Gelauff et al., 2024). More recently, several large language models (LLMs) in GPT, Claude, Gemini, and Llama families have also become capable of generating new high-quality text based on learning the patterns of their input data. This has opened possibilities to use LLMs in deliberative processes (Poole-Dayan et al., 2025; Tessler et al., 2024). Studies have found that LLMs can also increase the quality of online discussions (Heide-Jørgensen et al., 2025) and detect authoritarian discourse (Mochtak, 2025) and thus alleviate the erosion of public sphere.
From a technological perspective, democratic deliberation and LLMs are quite natural companions because both process large amounts of linguistic information (Summerfield et al., 2024). Furthermore, organised deliberation often entails ranking preferences – a task that is quickly performed by machine learning technologies. However, understanding what roles LLMs and other AI tools should play in DMPs requires reviewing them in relation to the normative values of democratic deliberation. These values comprise inclusion, mutual respect, communication free of coercion, mutual justification, and weighing of arguments by their merits, all exercised in the pursuit of collectively formed judgements on matters of public policy (e.g. Bächtiger et al., 2018).
In the theoretical scaffolding that follows, we distinguish six different functions of DMPs that are expected to realise these core values of democratic deliberation and identify potential roles for AI to support these functions. Because the six key functions are anchored in the normative theory of deliberative democracy, they enable us to examine and evaluate how different ways of using AI tools can benefit, but also potentially undermine, the core values of democratic deliberation in DMPs. The framework thus enables a critical analysis of real-life applications.
The conceptual contribution of our framework is that these key functions offer a way to understand the normative implications of the tasks and practices that humans and AI tools actually perform in the process of DMPs (c.f. Warren, 2017). Previous theoretical accounts either explicitly or implicitly focus on temporal phases of DMPs or core design features that are explained in a chronological order from the organiser or commissioner perspective (Curato et al., 2021a; McKinney, 2024). Functions, however, are not necessarily temporally sequential but may be realised throughout the process in varying order and in several phases.
First, DMPs are expected to represent different kinds of citizens. The function of representation is based on the normative principle of inclusion, which emphasises that all citizens bound by or affected by the decision should be able to participate or be represented in the deliberation leading to those decisions (Curato et al., 2017; Young, 2000). Representativeness in DMPs is often pursued in descriptive terms through sortition and stratification so that the participants demographically mirror the wider population. However, it sometimes refers to representation of relevant societal discourses, or even more broadly as representation of affected interests (Dryzek and Niemeyer, 2008). The development of LLM-based interpretation could help the recruitment and representation of affected groups in deliberative processes despite language barriers. Moreover, the use of LLMs in facilitation could help include larger numbers of deliberators, thereby enhancing representativeness (McKinney and Chwalisz, 2025), which confirms interdependencies between functions. Others have explored the use of LLMs in predicting various groups’ preferences (Zambrano et al., 2026) or viewpoints that would otherwise be excluded in deliberative processes (e.g. Fulay et al., 2025). This is particularly important for affected groups that are unable to participate in deliberations (e.g. future generations). However, while LLMs can be used to express and concretise otherwise missing perspectives, their use for representation gives rise to questions regarding the authenticity of representation, agency, and autonomy of those represented.
Second, DMPs have an inquiry function, meaning that normative arguments are weighed in light of empirical knowledge and evidence (c.f. Fischer, 2009). Provision of evidence and critically scrutinising it ensures that the outcomes of DMPs can be expected to be epistemically sound (Curato et al., 2021b), compared to other types of participatory processes collecting more spontaneous views. Inquiry in DMPs is ensured by allowing participants to access trustworthy and balanced evidence beforehand, often my means of information packages. LLMs could potentially help in simplifying and summarising relevant information (McKinney, 2024) and support different learning styles. LLMs can, for example, be used to visualise summarised materials and to re-phrase scientific reports into plain and comprehensible language.
Another way to ensure inquiry are Q&A sessions where participants can hear testimonies and present questions to experts and advocates. LLMs could empower participants and bridge the gap between experts and lay citizens by helping participants formulate questions that, for example, address information gaps discussed in small groups, or critically challenge expert views. It has also been suggested that experts could be replaced by LLMs serving as ‘on tap’ information sources that analyse predetermined background material during the deliberations (McKinney, 2024). Moreover, information-sharing and critical scrutiny of arguments provided by fellow deliberators are crucial aspects of deliberative inquiry, and there is some evidence that LLM applications can enhance critical thinking (Li et al., 2025).
Third, DMPs also foster mutual exchange of arguments among participants who hold different viewpoints and interests. The third function can thereby be labelled as ‘perspective-giving’ and ‘perspective-taking’ (Bruneau and Saxe, 2012). This entails expressing one’s own views, listening to the arguments of others, and seriously considering and weighing them (Niemeyer et al., 2024). From the normative perspective of deliberative democracy, participants’ viewpoints should be heard and considered in equitable terms (cf. Young, 2000), which increases participants’ understanding of others’ views. Some authors have suggested that AI could support perspective-giving and perspective-taking by, for example, moderating citizen deliberations online (Fishkin et al., 2025), measuring the deliberative quality of discussions, or helping participants develop arguments that express their views (McKinney, 2024). LLMs could also encourage reflection among participants or even act as ‘Devil’s advocates’ that challenge participants’ views.
Fourth, DMPs often produce policy recommendations and justifications. While deliberative polls focus on post-deliberation preferences, Citizens’ Juries and Assemblies produce written statements including policy recommendations and justifications. Writing collective statements requires the co-creation policy proposals and factual and normative justifications for proposals (Giraudet et al., 2022). These proposals and arguments are based on citizens’ own perspectives, developed and refined in the course deliberation. Co-creation in DMPs entails individual and collective ideation that is directly consequential in terms of the final output. It must be noticed that while other phases of DMPs may involve similar activities, for example, developing questions for experts as part of inquiry in our case study, these do not directly contribute to the collective output.
LLMs can, for example, assist participants in DMPs in tracking the evolution of their ideas and in capturing viable-but-overlooked original ideas that might be lost in the process of deliberation (Poole-Dayan et al., 2025). An underlying design question that affects the potential of AI in co-creation is the method of note-taking during deliberation. Traditionally, human facilitators have written minutes on participants’ ideas and arguments, which have then been synthetised with participants’ human intelligence. However, the rapid development of domain-specific LLM systems that convert participants’ speech to text in real time, such as Dembrane and DeliberAIde (UNDP, 2025), enables idea extraction from deliberations without the human note-taker.
Fifth, for statement-writing DMPs, another important function is integration, which entails accommodating diverging policy preferences, as well as relevant facts and values, into mutually acceptable and shared views (Mansbridge et al., 2010). LLMs can enhance integration by categorising preliminary outputs and identifying both commonly shared and conflictual viewpoints and arguments. Empirical studies have found that domain-specific LLMs can produce trustworthy summaries and consensus statements that reflect the views of marginal and minority groups (Tessler et al., 2024). Our empirical study presents in a hybrid process where co-creation and integration processes are based on the use of both human and artificial intelligence.
Sixth, many DMPs are ‘decision-oriented’ (Chambers and Warren, 2025), meaning that they aim to produce coherent and concise recommendations for the policy-making process. An important function of DMPs is thus to make collective decisions on common recommendations and justifications that best reflect the reasoning of the group and its understanding of the common good. In practice, therefore, DMPs apply voting mechanisms that provide closure to deliberation. The relationship of decision-making tasks with democratic deliberation is interdependent. While working towards a decision can help cultivate group atmosphere and the quality of deliberation, pressures to make decisions may introduce undesirable group dynamics to the process (Arnesen et al., 2025; Setälä et al., 2010). Because the aggregation of policy preferences as well as the selection or prioritisation of proposals and arguments are based on algorithms (such as voting rules), various AI tools could be helpful in ranking recommendations (e.g. Tessler et al., 2024). In addition to expediting computational processes, AI tools could be used to make aggregation more transparent for participants, for example, through visualisations.
Figure 1 summarises these potential roles for AI in supporting various deliberative functions in DMPs.

Potential benefits of AI tools in different DMP functions.
Case and data
The Finnish CA on Energy took place in January and February 2025. The CA was organised by the authors of this article, who work in the Research Group for Innovating Democracy at the University of Turku. The Assembly had a mandate from the Ministry of the Environment and the Ministry of Economic Affairs and Employment. The aim of the process was to produce a joint statement containing key facts about the Finnish energy system, as well as recommendations concerning various policy measures that affect energy use in households. The statement was expected to feed into multiple policy processes, including the National Climate and Energy Strategy and the Medium-Term Climate Change Policy Plan, which were being prepared in the ministries involved in the process. The statement of the Assembly was received by the ministries in a public event after the CA in March 2025 (the full statement can be found in Supplemental Appendix C). A follow-up event focusing on the impacts of the CA was organised in November 2025. In addition to ministries, these events involved number of stakeholders, facilitating public debate on the role of citizens and consumers in energy transitions.
Finnish policy-making has been characterised as a system of ‘routine corporatism’ (Vesa et al., 2018), where interest groups such as trade unions routinely are heard in policy preparation. In this context, DMPs are rarely used. Moreover, as Ruostetsaari (2017) has pointed out, citizens’ involvement is more limited in energy policy-making than in most other policy sectors in Finland. The need for bringing in citizens’ perspectives on energy issues, however, was particularly timely in Finland in 2025. During the past years, the Finnish energy system has experienced a rapid transition to renewable energy sources, especially wind power (Ministry of Economic Affairs and Employment of Finland, 2026, 178). For citizens, these changes in energy supply have meant significant fluctuations of energy prices, very high electricity prices especially in cold winter days, and a growing need for energy demand side management. The CA was expected to help articulate citizens’ perspectives on energy policies in this new situation.
The participants of the CA were recruited by mailing an invitation to 8000 randomly selected adult citizens residing in Finland. Of all the invitees, 186 volunteered, and approximately 70 people were selected to the Assembly, ensuring representativeness by age, gender, educational background, housing type, and region of residence. Ultimately, 54 people participated in the CA from start to finish. The composition of the CA reflected the wider population quite well, although people living in detached houses were overrepresented, and those living in apartment buildings were underrepresented.
The CA convened six times in total: two evenings online and four days in person over two weekends. The CA’s task consisted of, first, collecting and evaluating key facts about the Finnish energy system and, second, formulating recommendations on acceptable ways to steer energy use in households. Recommendations concerned four issue areas selected in collaboration with the ministries involved: energy demand side management, household emission-reduction policies, energy poverty alleviation, and energy counselling. Both the CA’s workflow and the final statement structure followed this thematic divide. The first three meetings consisted of expert presentations and Q&A’s, followed by drafting an initial set of facts and recommendations. The fourth meeting focused on finalising the key facts, and the last weekend was spent editing recommendations. Deliberations were guided by trained facilitators in alternating small groups and in the plenary. A detailed description of the CA’s recruitment and deliberation processes is presented in the Supplemental Appendix A.
In addition to fulfilling its main task, the CA was designed to test the use of LLMs to support the functions of DMPs. Because we did not want LLM use to compromise the quality or the legitimacy of the CA and its outcomes, only minimal adaptations to the deliberative procedures were made in incorporating LLMs. At the same time, we wanted to test how easily-accessible LLMs could help facilitate the deliberative process and the co-production of a written statement. Therefore, we limited the experimentation to three functions: inquiry, co-creation, and integration. Out of different LLM applications, we opted for OpenAI’s GPT-4o model used with ChatGPT Team subscription, as it was deemed sufficiently developed with an easy-to-use interface, and it enabled retaining ownership and control over the materials processed with it. At the beginning of the Assembly, participants were informed that generative AI would be used to assist them in their work, but that they would have the final decision-making power over the statement’s content.
To support the inquiry function, ChatGPT was utilised to summarise participants’ questions to experts in some of the expert Q&As. During the first and third meetings, small groups could formulate as many questions as they wanted, and ChatGPT was then prompted by the organisers or moderators to formulate a pre-defined number of questions based on the small groups’ input. In the second meeting, small groups themselves decided which two questions they wished to ask. For co-creation and integration, ChatGPT was used between the third and fourth meetings by the organisers in two ways: First, it was prompted to formulate 10 key facts based on the initial set of 29 facts produced by participants during the first meeting (Table A2 in Supplemental Appendix B). Second, it was prompted to thematically cluster draft recommendations (35–37 per theme, 142 in total) written in the second and third meetings (Table A3 in Supplemental Appendix B). The clustering resulted in four to eight categories per theme, and these were further divided into nine subsets by the organisers to correspond with the number of the Assembly’s small groups. All prompts can be found in Supplemental Appendix B. An outline of the CA process is presented in Table 1.
Process design of the Citizens’ Assembly on Energy.
Data
Survey data used in the analysis were collected from the participants of the CA throughout the process. More comprehensive surveys were fielded at the recruitment stage and shortly before the Assembly (T0). These surveys mapped participants’ views and knowledge on energy policies, energy consumption, and experiences with AI, as well as more general background information, prior to the deliberation. Surveys conducted during deliberations (T1–T5) focused on participants’ views on the quality of the deliberation and the use of AI in each meeting of the CA. Questions on participants’ views and knowledge on the energy system were repeated in the post-Assembly survey (T6), conducted after the last meeting, allowing a comparison of views before and after deliberation. In total, 54 participants completed the CA process, but the number of respondents varied from meeting to meeting due to, for example, the absence caused by an illness. In addition to the participant surveys, survey data was collected from the small group facilitators after each meeting (F1–F6). These surveys focused on the facilitators’ experiences with the use of AI and the quality of deliberation in small group discussions. In total, nine facilitators completed these surveys. An overview of the survey data can be found in Table 1.
Analysis
The following analysis section is divided into three parts: first we will inspect the inquiry and co-creation and integration functions in their dedicated segments, after which we present the participants’ and facilitators’ evaluations concerning the impacts LLM use had on deliberation. Our analysis is based on the survey answers and open-ended feedback given by the participants and the facilitators. It should be noted that evaluations given on different days are not entirely comparable, as each day of the Assembly varied in several factors, such as format (online or offline), topic, and aim, in addition to the use of AI.
Inquiry
Regarding the inquiry function, LLM was used to assist participants in crafting questions for the experts. After each day of deliberations, both the participants and the facilitators were asked to evaluate the questions presented to the experts on that day according to multiple criteria. Mean scores and standard deviations regarding the answers are presented in Table 2.
Participants’ and facilitators’ evaluations of questions presented to experts.
Data are presented as mean scores (SD).
During the CA, the participants crafted a total of 254 questions in small groups (see Supplemental Appendix B Table A1). Of these, a total of 57 were posed to the experts during the process. In the parts of the process where the LLM finished the questions for the participants, they were encouraged to come up with basically an unlimited number of questions. Because of this, small groups produced almost twice as many questions in these stages compared to Q&As where LLM was not used. For example, according to the transcripts one participant noted that: ‘we can just lightly throw in some questions, and the AI will filter out the bad ones’.
After the first meeting (T1), the participants’ evaluations were mostly positive; based on survey data, the questions generated by AI were considered to be relevant, useful, and understandable. After the second day (T2), participants also perceived the questions that they had generated themselves as diverse and fairly general, but they were not seen as fully understandable or helpful. After the third day (T3), the evaluations turned more critical. Participants perceived the questions generated by AI to be more one-sided, irrelevant, unhelpful, and confusing than before.
The open-ended feedback from the participants mirrors these evaluations. After the first meeting, the feedback did not focus on the use of AI at all. This is presumably because online deliberation itself was a new experience for most of the participants, even without the use of AI. T3 produced more critical feedback concerning the questions proposed by AI. Several participants criticised AI’s formulation of the questions as ‘too general’, ‘similar’, and ‘vague’. One participant noted that there was a widespread perception in their small group that AI was not helpful in formulating questions and that the session itself did not meet the participants’ information needs. During the small group discussions, participants similarly critiqued LLMs for ‘deleting’ or for ‘rounding off the edges’ of their questions. This is supported by the observations of organisers who found that AI-generated questions were very descriptive and did not challenge the experts in any way. For example, AI changed all ‘why’ questions into descriptive ‘what’ questions.
Based on the transcripts, participants’ reflections on LLM during the small group discussions were mixed. While some participants were satisfied with the outcomes and felt that majority of their small group’s questions were covered, others claimed that LLM had excluded some of their questions. Some even started to plan ways to optimise the probability of their questions being selected or included in the final set, for example, by repeating the same point several times, even though they were not certain about how the LLMs ‘aggregated’ the questions. Participants also reflected that it is preferable to formulate, prioritise, and ask the questions by themselves rather than rely on LLM, even if this results in ‘dumb questions’, since these can be more enjoyable for experts to answer than the ‘smart questions’ generated by LLMs.
The facilitators also considered the LLM-assisted inquiry during the first day (F1) to be successful, giving the questions high evaluations throughout. In the open-ended feedback, one facilitator summarised: ‘The questions generated by AI were quite good at taking into account all the topics that emerged in the discussion, although they remained at a slightly more general level than the original questions formulated in the group’. However, after the third day (F3), the facilitators’ assessments of the quality of the questions plummeted, even more so than those of the participants. For example, the facilitators evaluated the questions as being very one-sided. According to the open-ended feedback, facilitators evaluated the use of LLM in the small groups rather negatively. The facilitators also expressed that, even though the use of LLM made their own work easier by reducing writing-related tasks, the participants showed dissatisfaction and considered the questions they had initially formulated to be more useful. According to the facilitators, the participants felt that the very detailed questions they had produced themselves were not voiced. They also wished that they could have further re-formulated the questions generated by LLMs.
There are likely numerous reasons for the more critical evaluations towards the end of the inquiry phase. First, the topics varied each day. On the first day, the purpose was to learn more about the Finnish energy system in general, and therefore, the questions the participants had formulated were quite descriptive and broad. At this stage, the use of LLMs suited the purpose. On the third day, the topics energy poverty and counselling were more defined and specific. Here, the somewhat abstract and descriptive formulations by LLMs no longer satisfied the participants’ needs and expectations. Second, the rather positive assessments on the first day could have been due to the novelty of the whole deliberative process and the participants’ overall satisfaction. After the second day, when the participants had been able to formulate questions on specific topics themselves in small groups, the questions summarised with the LLM on the third day may have seemed disappointing. The facilitators’ harsh critique of the helpfulness of AI after the third day can be due to their dissatisfaction with the questions formulated and summarised by AI, but also due to the sceptical remarks made by the participants in small groups, which were then transmitted to facilitators’ evaluations
Co-creation and integration
During the co-writing of the citizens’ statement, the participants were asked to evaluate the quality of the facts or recommendations at the end of each day. The criteria for evaluation were the same as those in the previous surveys, with two additional criteria specific to this phase. It should be noted that, because the procedure, the tasks of the participants, and the uses of LLMs were different each day, comparisons between the daily scores should again be made with caution.
At the end of the fourth day, the participants (T4) and the facilitators (F4) were specifically asked to evaluate the preliminary facts summarised by LLMs that they had received at the beginning of the day (see Table 3). Even though participants considered these facts to be unfinished and not very diverse, the truthfulness and helpfulness of the output were rated somewhat high. Participants also accepted some of the LLM-integrated key facts as they were, often expressing surprise: ‘These were surprisingly good!’; ‘Was this the first time I would accept a sentence generated by AI?’ Evaluations given by the facilitators did not differ greatly from those of the participants. The facilitators saw that the preliminary facts formulated by LLMs were quite truthful, diverse, and relevant. They evaluated them as helpful, even if they were perceived to be far from being finished.
Participants’ and facilitators’ evaluations of preliminary facts and recommendations.
Data are presented as mean scores (SD).
Open-ended feedback from the participants on the use of LLMs for formulating facts was scarce. When co-writing the facts, one participant felt that LLM ‘mixed things up’. One participant stated that the use of LLM in the process did not feel ‘smooth’. This evaluation could have been due to the fact that the small group facilitators had to take care of multiple things at once; they had to keep a record of all the facts in a Google sheets document, periodically share the screen with participants to approve edits, and ensure the quality of the discussion in an online environment. During this phase of the work, it became very clear that editing the summaries created by LLMs required a lot of time. In their open-ended feedback, many of the facilitators reported these types of challenges, noting that some participants had concerns about the process and edits made by other groups.
At the end of the fifth day (T5), the participants and the facilitators were asked to evaluate recommendations regarding the four themes discussed. It is worth noting that, at this stage, these evaluations concerned recommendations that were only classified, not modified, by AI, and further reformulations were done by the participants themselves in small groups. The participants particularly praised the diversity of the recommendations. The evaluations of the final day (T6) focused on the final recommendations that ended up in the statement, which received more positive evaluations overall than the unfinished recommendations.
The final versions were seen as understandable, truthful, and more finalised than the preliminary ones. However, it is interesting that the facilitators also viewed the recommendations listed in the statement as slightly more one-sided and general and less relevant than the ones evaluated the day before. This is likely explained in part by the fact that there were fewer recommendations included in the final statement, which could have impacted perceptions regarding diversity and generality. On the other hand, this evaluation can reflect deliberative convergence, a somewhat expected result of integration of different views.
Based on the open-ended feedback, the facilitators were not suggesting that the participants performed worse than the LLM. While the facilitators thought that using LLMs to categorise the recommendations by the topic was successful and made them easier to summarise and possibly combine, as well as to eliminate, overlaps, they also highlighted the importance of reviewing LLM’s work and emphasised that the participants should have the final say. One facilitator expressed that the use of LLM could ‘lose interesting details and perspectives that emerged in discussions’. On the other hand, another facilitator raised an important point about the impact of LLM assistance by wondering whether it decreased the motivation of some participants to carefully engage in deliberation and prepare the facts and recommendations thoroughly when they were going to be revised or reorganised by LLMs in any case.
Evaluations of the impact of the use of LLMs in deliberations
Next, we inspect the participants’ and facilitators’ evaluations of the use of LLM in the deliberations in general. These evaluations were carried out at the end of first (T1), third (T3), and final day (T6) of the Assembly. The respondents were asked to evaluate on a scale of 0–10, whether the use of AI decreased (0) or improved (10) the quality of deliberations; whether it was harmful (0) or helpful (10) for the deliberation; and whether AI should be used less (0) or more (10) in deliberations. The distributions of evaluations are presented in Figure 2 below, using violin charts and boxplots.

Participants’ evaluations (n = 53) of the overall use of AI in the process.
The survey responses show that after the first day (blue colour), the participants’ assessments were on the positive side, with a noticeable number of respondents choosing values in the middle of the scale. LLMs were seen as quite helpful and as improving the quality of the deliberations to some extent. However, after the third day (red colour), the evaluations were more mixed, evident in the notable shift towards lower scores. At the same time, there is a noticeable decrease in the mean scores between T1 and T3 evaluations regarding helpfulness (6.82 > 5.85) and quality (6.45 > 5.74), which show a shift towards more sceptical outlook. While the mean scores between T1 and T3 for the third item, whether deliberation should be used less or more, are exactly the same (5.36), there is considerably more deviation in evaluations after the third day.
Even though the evaluations concerning the use of LLM at T3 were somewhat critical – a finding that mirrors the criticism aimed at the use of LLMs in the inquiry task – the participants’ evaluations were a bit more positive after the process ended. In terms of mean scores, T1 and T6 evaluations regarding helpfulness (6.82 > 6.70) and quality (6.45 > 6.11) are quite close to each other. Still, evaluations were noticeably more mixed in the end of deliberations than at the start.
It is notable that, while the evaluations were mixed and somewhat critical, participants suggested several times during the small group discussions that LLMs could be used to ease the task: ‘How would you phrase that briefly? AI would probably be sharper than me at it’; ‘Put those two sentences into ChatGPT’; ‘You have got the AI there’, ‘Yeah, can you ask it?’ Some even mentioned they ‘harassed’ LLMs with their personal devices to request better terms for the discussed concepts, although these were quite quickly forgotten. LLMs seem to be treated as a convenient shortcut even when its outputs are not perceived as particularly satisfactory.
The facilitators’ evaluations (n = 9) follow a similar trend as those of the participants, but with starker contrasts. After the fairly positive early evaluations, the facilitators’ mean scores decreased significantly between F1 and F3, especially regarding the harmfulness or helpfulness of LLM (8.00 > 3.56). After F3, there were also noticeable decreases in evaluations regarding LLM’s impact on the quality of deliberation (6.56 > 4.78) and whether it should be used more often (4.78 > 2.78). The fact that facilitators’ views became extremely negative after the third day may be due to the concrete experience of being obliged to use the LLM chat tools themselves. This may have concretized the way LLMs reduce the richness and diversity of human deliberation in the task of formulating questions to experts. In addition, they may have felt a desire to protect human interaction in deliberation. However, the facilitators’ evaluation bounced back, and the differences in mean scores between F1 and F6 were considerably smaller (6.56 > 6.00; 8.00 > 6.11; and 4.78 < 4.89, respectively).
As was stated earlier, topic complexity, meeting format, and other different aspects are confounded with the LLM intervention, which itself was carried out a bit differently in different phases of the Assembly. Therefore, interpretations should be carried out with caution. However, our observations from using multiple different sets of data points to a general finding are as follows: the general-purpose LLM GPT-4o struggled to accomplish its task acceptably in the inquiry function, when it had to craft questions to experts regarding specific, rather narrow topics, but performed better when the topic of inquiry was more general, and when it was tasked to aid in co-writing and integration.
Discussion and conclusions
Some deliberative democracy scholars and democracy developers have high hopes for the capability of AI to improve and scale democratic deliberation. We argue that the evaluation of the strengths and weaknesses of AI technologies should be grounded in key normative values of deliberative democracy. We distinguish six different deliberative functions in DMPs that should be performed by humans or AI tools in order to advance the key values and norms of democratic deliberation, namely representation, inquiry, perspective-giving and perspective-taking, co-creation, integration, and decision-making. These six key functions of democratic deliberation are interconnected. Most notably, perspective-giving and perspective-taking, co-creation, and integration require inclusive representation of citizenry (or affected interests). The design features and the practical implementation of DMPs affect the ways in which these key functions play out in the deliberative process. Nevertheless, distinguishing these functions is important for identifying the benefits and drawbacks of AI interventions in deliberative processes.
Our explorative study investigated to what extent AI technologies, and LLMs in particular, can assist in inquiry, co-creation, and integration in a real-world DMP. The study provides a first-of-its-kind process design for AI-assisted DMPs where AI is used during the process ‘on-the-go’. Based on survey data, facilitator feedback, and organiser observations, we found that a commonly used general-purpose LLM, GPT-4o, was considered useful in summarising preliminary facts and categorising preliminary recommendations for deliberations among small groups. Therefore, the first main takeaway from our study is that LLMs can facilitate deliberative functions of co-creation and integration and therefore help DMPs co-produce written statements among a large group of deliberators. This can improve the quality of DMP outputs but also the quality of deliberation by allowing more time for weighing arguments and developing justifications, not just copyediting and clustering them. Consequently, the use of LLMs could help scale up the impacts of DMPs in the broader democratic system by making the viewpoints and justifications accessible for the wider public.
However, LLMs should be used cautiously to integrate preliminary recommendations, and they should not have a final say in writing common statements as they can distort or misunderstand information. In other words, the use of LLMs does not eliminate the need for human work on the contents of the statement. On the other hand, improved assessments of AI at the end of the CA may reflect not just increased satisfaction with AI tools but, more generally, positive sentiments triggered by a successful conclusion of the deliberative process (cf. Morrell et al., 2022).
The second main implication of our findings is that LLMs have some major shortcomings in terms of the inquiry function of DMPs when it comes to forming questions for experts. Our results indicate that LLMs have also limited capability to assist in summarising and selecting questions to be posed to experts. With the prompts we used in this CA, ChatGPT omitted many important details from the participants’ original questions and ignored the most critical ones among the questions created for expert panels. This could have severe consequences for not only the knowledge base participants build from expert hearings but also on the critical potential of DMPs more generally by steering the discussion towards less controversial topics that do not challenge existing authorities. This would further aggravate the concerns that mini-publics have lost the critical edge that deliberative theory started with, being unable to propose radical or novel policies (Frick et al., 2026). In the rapid development of AI-assisted deliberation and LLMs, we must ensure that LLM tools are designed to support critical thinking and help citizens to better articulate their own recommendations in DMPs.
Our results are by no means exhaustive, and much remains to be done to further develop the capacity of AI tools in the different functions of deliberation. A further, more practical takeaway from our study is that prompts for LLMs in DMPs should be designed with the aim to maximise the critical and reflective qualities of the outputs. We found that with a rather simple prompt, in which the basic features of DMPs were described to the LLM, much of the nuance and critical edge was lost. Predicting the impact that different prompts will have on the output is difficult, however. What is going on ‘under the hood’ of LLM interfaces is not visible; therefore, an approach of trial and error may be the only way to develop prompts with a more critical edge and deliberative focus. Future research could therefore systematically focus on deliberative prompt engineering while also enabling it by maintaining transparency regarding the prompts used in experimental or real-life DMPs.
We have theorised possible roles for AI technologies in relation to six deliberative functions in DMPs. While we now have some evidence on three of these functions, many other potential uses require further investigation. For example, could LLMs enable the organisation of multi-lingual assemblies by providing live translation and transcribing. Or could LLMs represent otherwise excluded groups in deliberative processes, and how could we solve the issues of authenticity and autonomy in this case? Furthermore, will AI facilitators be able to understand hidden meanings and non-verbal communication that is crucial for human interaction and, in doing so, enhance perspective-taking among larger groups of deliberating citizens? Finally, how can various AI tools be used in voting in ways that help support the key functions of deliberation, such as perspective-giving and perspective-taking, and avoid undesirable group pressures?
Future empirical studies are needed to gauge the performance of AI in various deliberative functions in DMPs. One limitation of our study is that we have investigated how a commercial, general-purpose LLM functions in a DMP. Our results should not therefore be generalised to cases where custom deliberative LLM tools are used. However, we see that local communities and public authorities are often obliged to use general-purpose models as there are no resources allocated for democracy technologies, which is why it is important to understand the consequences of these models. Another limitation of our case study is that empirically, we rely on subjective assessments of those involved in the DMP, namely citizens, moderators, and organisers. Especially participants’ critical assessments of the use of AI cannot be regarded as an indication that it was harmful by the normative standards of deliberative democracy, but perhaps rather as a positive sign that participants retained their capacity of critical thinking. Further research on similar DMPs should investigate the impact of AI on deliberative quality by analysing transcribed speech acts with discourse quality index (DQI) or similar methods. Moreover, randomised, controlled experiments are needed to explore how different uses of AI affect DMPs’ deliberative processes and their outputs, as well as their impacts and perceptions among the wider public. Critical studies are also needed to identify how AI shapes the motivation and commitment of citizen deliberators and to what extent the introduction of new technologies creates meta-deliberation on the method itself.
Regardless of which function AI is designed to contribute to or perform, it is important that the final version of DMPs’ output remains in the hands of participants. Our results indicate that classifications and summaries that LLMs construct from preliminary questions, facts, or recommendations put forth by participants are by no means exhaustive or trustworthy. Therefore, an iterative process in which human participants check all the outputs of AI is crucial. As an example of ‘techno-solutionism’, governments and public administrators may see AI as a possibility to organise more cost-effective and shorter DMPs. However, because the value of citizen deliberation is in the communication among participants, AI tools should be used to free up time for substantial and consequential deliberation, rather than merely cutting its costs.
Supplemental Material
sj-docx-1-pol-10.1177_02633957261466242 – Supplemental material for Can AI support inquiry, co-creation, and integration functions in Citizens’ Assemblies?
Supplemental material, sj-docx-1-pol-10.1177_02633957261466242 for Can AI support inquiry, co-creation, and integration functions in Citizens’ Assemblies? by Maija Jäske, Katariina Kulha, Mikko Leino, Maija Setälä and Toni Wessman in Politics
Footnotes
Acknowledgements
We want to thank Oona Ylikoski for organisation and communication support in the Citizens’ Assembly process.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The study received funding from the SRC project ‘Fair, flexible and socially resilient energy systems’ (FLAIRE) (decision number 358428). This work was also supported by the European Union (ERC, ADDI, 101166894). Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Council Executive Agency. Neither the European Union nor the granting authority can be held responsible for them.
Statements related to ethics and integrity policies
Data availability statement
Supplemental material
Supplemental material for this article is available online.
Author biographies
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
