Abstract
This study aimed to determine the role of artificial intelligence (AI) techniques in accessing manuscripts by critically evaluating the accuracy and completeness of the information provided by ChatGPT about Islamic manuscripts with the aim of identifying areas of strength and areas where improvement is needed in its performance as an information extraction tool. The manuscripts included in the study (23 Arabic manuscripts) were known and were identified by referring to a number of academicians specializing in Islamic manuscripts. The study used the ChatGPT-3.5 version, and one of its most important results is that the number of manuscripts that ChatGPT provided information about was 14 out of 23 (60.9%), while it did not provide information about nine manuscripts (39.1%). At the level of the descriptive or identifying data of the manuscript, ChatGPT provided four descriptive elements: the title of all 14 manuscripts (100%), the subject of 10 manuscripts (71.4%), the formal features of eight manuscripts (57.1%) and the author of six manuscripts (42.8%).
Introduction
A manuscript is the original copy that the authors wrote in their own handwriting, or allowed to be written, or approved, or what was copied by copyists later in other copies transferred from the original. Manuscripts constitute an important heritage created by the Arab and Islamic civilization in various fields of human knowledge. They are works into which scholars put their ideas, experiences and creativity. Therefore, many cultural and educational institutions in the world have sought to compete in searching for what remains of this heritage with the aim of collecting, preserving and then making it available to researchers, owing to its great importance in scientific research. Manuscripts attract the attention of many Arab and foreign scholars and researchers alike owing to their scientific and artistic value. Islamic manuscripts provide unique visions in various fields such as science, art, religion and philosophy; therefore, many categories of researchers and specialists are interested in them, which enhances efforts to preserve and explore them for future generations. Identifying data related to manuscripts, including detailed information about their content, historical context and subject matter, is essential and facilitates access to them, which enhances scientific research and encourages multidisciplinary studies based on Islamic manuscripts. Arab and Islamic culture has attached importance to books; according to Gacek (2001), from the early first century AH (mid-seventh century AD) to the thirteenth century AH (nineteenth century AD), the Islamic civilization produced tens of thousands of works in Arabic on religious and scientific topics, resulting in millions of manuscript copies that provide insights into various aspects of the Arab and Islamic civilization, including science, art, literature and language. Some of these manuscripts are distinguished by their unique artistic and aesthetic qualities, displaying intricate decorations and drawings that reflect the artistic skills of the scribes (Tawhara, 2022). Libraries around the world contain a treasure of Islamic manuscripts written in Arabic and other languages such as Persian, Ottoman Turkish and Urdu (Khasawneh, 2020).
Islamic manuscripts are of great historical and cultural importance, as they contain the knowledge and teachings of Islam and the history of the Islamic nation. Islamic manuscripts also reflect the civilizational and scientific development of Islamic societies throughout the ages. By preserving manuscripts, this valuable heritage is protected and documented for future generations. The history of preserving Islamic manuscripts dates to the early Islamic era, when manuscripts were kept in mosques, schools and libraries. The importance of preserving manuscripts increased with the spread and expansion of Islam, and this process has faced many challenges throughout the ages, such as the loss of manuscripts during wars and invasions, the deterioration of paper, and the lack of sufficient funding to preserve them well.
With the rapid development in the world nowadays in modern technologies, methods of storage and digital retrieval have been used to collect, preserve and make available manuscript heritage, which has led most Arab countries to carry out digitization projects for their manuscript holdings (Ghazal, 2012). Information specialists consider digitization a modern means to facilitate the preservation, organization and wide availability of texts. The most prominent techniques used in preserving Islamic manuscripts are digital imaging and infrared imaging, chemical analysis techniques for paper and ink, and paper strengthening and restoration techniques. Electronic storage techniques for manuscripts and the creation of their own databases are also used. These techniques are constantly evolving to achieve the best results in preserving Islamic manuscripts.
Recently, the role of AI-based systems in processing and providing access to published information sources or manuscripts and their metadata has gradually emerged. Among these systems, ChatGPT stands out as an advanced language model developed by OpenAI, designed to participate in natural language conversations and generate coherent and contextually relevant text. The ChatGPT model is characterized by its ability to handle complex language tasks such as conversations (Wu et al., 2023). This model has revolutionized the approach of AI in interacting with humans, as it acts as a chatbot trained to generate understandable text from prompts. ChatGPT can produce new content based on its training data (pre-training) and is designed to process natural language, analyze the meanings of sentences (natural language understanding) and generate new sentences based on inputs (natural language generation), resulting in more contextual and flexible responses compared with rule-based natural language processing models. Recently, ChatGPT has gained attention for its ability to provide high-quality responses to human queries in multiple domains (Zhong et al., 2023). The accuracy of ChatGPT in extracting and presenting information from diverse sources is of paramount importance, especially in research environments where data reliability is critical. Understanding ChatGPT's capabilities is essential for researchers, scholars, and information specialists who rely on automated tools to extract and analyze data. As a result, this study focused on evaluating the effectiveness of ChatGPT as a tool for accessing Islamic manuscript data, which is of great importance to those working in fields related to Arab–Islamic heritage.
Study objectives
Some recent studies indicate the ability of AI applications, such as ChatGPT, to provide detailed and accurate answers and use them in research-related tasks such as article writing, idea generation, and information summarization (Kocoń et al., 2023). In addition to its ability to convert manuscripts into readable text automatically, with the growing use of AI applications in scientific research, they can be used in research and access to Islamic manuscripts. By searching online and in databases, researchers in this study concluded that there are few studies that have addressed the use of AI applications in preservation, digitization, and accessibility of Islamic manuscripts. The problem of the study lies in identifying the effectiveness of using ChatGPT as a means of searching for Islamic manuscripts and evaluating the accuracy and completeness of the information provided by the ChatGPT about Islamic manuscripts, which is the main objective of this study. To achieve this goal, the study aimed to find answers to the following questions:
Does ChatGPT provide accurate and complete metadata about Islamic manuscripts? What level of summary does ChatGPT provide about Islamic manuscripts (general summary of the manuscript, summary of chapters, or summary of each chapter)? Does ChatGPT provide access to digital copies of manuscripts? Does ChatGPT provide any other additional data relevant to the search for Islamic manuscripts?
This study supports future developments in the field of natural language processing and information retrieval using AI, by identifying areas of strength and areas for improvement in ChatGPT's performance as an application for AI in extracting information for researchers and those interested in Islamic manuscript heritage preservation and accessibility.
Methodology
The study used a qualitative approach through the content analysis method as a suitable research method for extracting and summarizing data and revealing its strengths and weaknesses of the data (Bowen, 2009) provided by ChatGPT when used to access Islamic manuscripts with the aim of identifying areas for improvement in its performance. Given the researchers’ experience in dealing with Arabic manuscripts and their importance as a source of Islamic heritage, they have focused on them owing to their value in providing a deeper understanding of Islamic culture and history. To assess the accuracy and completeness of the information provided by ChatGPT about the manuscript and its author, the study relied on determining the number of relevant manuscripts retrieved, the number of correct answers, the accuracy of the manuscript's metadata, and the summary provided by ChatGPT about the manuscript, in addition to analyzing errors in the information.
Regarding the questions asked to ChatGPT, the study used a checklist that included four questions related to four aspects: the manuscript's descriptive data (such as title, author, year of writing, scribe, year of copying, place of copying, library where the manuscript is kept, subject, formal features), a summary (keywords, manuscript summary, chapter summaries, content evaluation), the availability of a digital copy (text/image) and any other details provided by the ChatGPT program about the manuscript. The content was analyzed according to these four aspects. ChatGPT was asked about the manuscript using its title, and if the response was not accurate, the manuscript was asked about using its title and the author's name (Figure 1). The researchers depended on the thematic analysis method used in qualitative research, developed by Creswell and Plano (2017), to analyze the content of ChatGPT responses: collecting data in paragraphs or phrases, analyzing the paragraphs and phrases and identifying main ideas and subtopics, summarizing and presenting the results briefly in frequency tables, then analyzing and interpreting the results, and making recommendations based on them.

ChatGPT's response about the metadata for the manuscript of Al-Tabari's interpretation “Jame’ Al-Bayan fi Ta’weel Ai Al-Qur’an”.
The study used GPT-3.5 because it is faster than GPT-4 and more flexible than GPT Base (Ipsen, 2023), and it is also good for most tasks, whether chat or general. The data were analyzed in Tables 3–7 for each of the study questions. In selecting the manuscripts included in the study (Table 1) and (Table 2), it was considered that they are well known, which means that ChatGPT is supposed to be able to provide information about them, through the following steps:
Considering the researchers’ experience and by referring to the databases of a number of indexes of Arab and foreign libraries and digital repositories that include collections of Islamic manuscripts, the study identified a preliminary list of titles of a number of famous Islamic manuscripts for which ChatGPT would be asked to provide identification and descriptive data. The researchers prepared a guiding list that included 30 manuscripts (Table 1). Among the digital repositories of manuscripts reviewed in the preparation of the preliminary list are the State Library of Berlin (2023) and the Leipzig University Library (2023), known as the Albertina Library, which is the central library of the University of Leipzig in Germany, and is famous for its special collections of Oriental manuscripts. The holdings of both libraries can be searched for Oriental manuscripts through the Central Database of Manuscripts in German Libraries (Qalamos, n.d.). Moreover, the researchers reviewed the electronic catalogue of documents and manuscripts held by French libraries, universities and research institutions (Calames, n.d.), the Bodleian Library at Oxford University in the UK (SOLO, n.d.) and Al-Alokah Library (Al-Alokah, n.d.). The proposed list was sent to a number of professors specialized in and concerned with Islamic manuscripts (five specialists) with specializations in manuscripts, literature and language, history, history of science and Islamic sciences from several universities (Cairo University, Al-Azhar University, Beni Suef University, Sultan Qaboos University and Freie Universität Berlin). Those specialists were asked to review the list and choose 25 manuscripts well known among researchers and those interested in Islamic manuscripts that cover various fields such as jurisprudence, hadith, interpretation, doctrine, poetry, medicine, geography and others, and told that they could be guided by the list or add any other manuscripts. They were also asked to consider several factors in their selection of manuscripts, including historical importance, manuscripts written by well-known scholars or important figures in Islamic history, and manuscripts that are famous among researchers and scholars. The lists of manuscripts received from specialists were sorted based on the title and author of each manuscript, and it was found that they agreed in the selection of 23 manuscripts included in this study (Table 1).
Manuscripts included in the study.
The importance of the manuscripts included in the study.
Information provided by ChatGPT about the manuscripts included in the study considering Table 3, among the descriptive data related to the manuscripts, ChatGPT provided only four descriptive elements out of ten, at a rate of 40%. These elements are title, author, topic, and physical features. They are arranged in descending order as follows: title (14 instances) with a percentage of 100%, subject (10 instances) with a percentage of 71.4%, physical features (eight instances) with a percentage of 57.1%, and author (six instances) with a percentage of 42.8%.
Literature review
Searching online and in databases turned up some studies on indexing, preserving, making available and investigating Islamic manuscripts, including studies related to automated indexing of Islamic manuscripts, their digitization, applying optical character recognition (OCR) technology to manuscripts, and using AI technology in preserving and accessing digital manuscripts. However, the search did not find studies on evaluating the accuracy and completeness of the information provided by ChatGPT about Islamic manuscripts, which is the main objective of this study.
Among the studies that focused on automated indexing and databases of Islamic manuscripts is that of Mahmoud et al. (2015), where they presented the design of an electronic program for indexing Arabic manuscripts according to MARC 21 standard and recommended the necessity of using a computerized training method to teach courses on indexing Arab manuscripts, given its superiority over the traditional method, which helps in diversifying training methods. Another study on the reality of the readiness of indexes provided by Arab heritage preservation institutions presented the most important international standards for indexing Arabic manuscripts and index services available on the Internet (Ahmed, 2013). Moreover, a study on the challenges facing the digitization and availability of Omani manuscripts recommended the need to qualify a specialized team, unify indexes, facilitate administrative and legal procedures, and create a website through which a single search strategy can be implemented, and all manuscript collections can be accessed (Al-Hajji, 2016).
Likewise, among the studies related to the digitization of Arab manuscripts is the study of Ahmed (2013), which dealt with the digitization of manuscripts in Algerian universities and recommended the establishment of laboratories that are concerned with scientific research in the field of manuscripts. Also, the study of Al-Hajji (2016) indicated that one of the most prominent challenges facing the digitization of manuscripts is the weakness of the indexes, the need to verify the accuracy of the author's name and to ensure that the book is attributed to him (Al-Hajji, 2016). The UNESCO article on the basic principles of digitization indicates that in order to enhance the dissemination and availability of manuscript heritage and maximize its value for research and study, appropriate descriptors or metadata should be identified for retrieval in preparation for making them available electronically and determining the appropriate file format for display as an image, PDF file or other appropriate format, as well as making the digitized collections available to beneficiaries (IFLA, 2014).
At the level of applying OCR technology in manuscripts, the study of Schoen and Saretto (2022) was able to apply it to manuscripts dating back to the Middle Ages using the Kraken OCR engine, which was trained on copying English manuscripts in the early fifteenth century, achieving an accuracy rate of 97%. Similarly, the study by Memon et al. (2020) indicated that the techniques used in OCR in manuscripts make retrieving the required information easier, and that they mainly depend on extracting features and classifying these features based on patterns, which allows the extraction of textual content for further processing and analysis. HTR technology, “handwritten text recognition”, is used in copying and digitizing handwritten manuscripts, where handwritten manuscripts can be efficiently converted into searchable and editable digital formats, allowing easy access to content and facilitating further analysis and research. Based on the importance of digitizing manuscripts, Hassen and Khemakhem (2023) recommended the creation of a platform for digitizing Islamic manuscripts that facilitates storage, retrieval and processing, as well as rapid access and unlimited storage capacity. Chanda et al. (2018) targeted Arabic handwritten materials on the web as ancient and modern information materials, and reported that they require descriptions in web search tools that are compatible with the nature of handwriting, content and text-based indexing methods to deal with thematic analysis of these materials.
It's worth noting that among the distinguished experiences in digitizing and indexing manuscripts, is that of BULAC (Bibliothèque universitaire des langues et civilisations), a library whose holdings cover the Balkans, Central and Eastern Europe, the Maghreb, the Near East, the Middle East, Central Asia, Africa and Asia. It contains a multi-font index, allowing searches in the alphabet chosen by the researcher and works written in languages other than Latin can be searched in the original font to facilitate searches for users. The library's primary mission is to create rich collections of published and manuscript information sources around the world in many different languages. The idea of digitizing the manuscript collections of the National Library of France (Bibliothèque Nationale de France) dates to 1997, when the virtual library known as Gallica became available to users, providing free and open access to several million restored manuscripts and thousands of images, old records and digital documents from all periods and on all media (Gallica, n.d.).
Among the studies that addressed the effectiveness of ChatGPT in natural language processing and focused on the Arabic language is the study by Khondaker et al. (2023). They conducted a large-scale automated and human evaluation of ChatGPT, covering 44 different language comprehension and generation tasks on over 60 different datasets. The study compared ChatGPT and GPT-4 in Modern Standard Arabic and Arabic Dialects. They found relative shortcomings in both models when handling Arabic dialects, but there was a positive correlation between the human evaluation and the GPT-4 evaluation. The study also found that ChatGPT, despite its large size, struggles with multiple Arabic tasks, but its models can be improved. Likewise, the study by Anwar and Ahyarudin (2023) aimed to develop an AI-based learning system that educational institutions could adopt, enabling them to provide more personalized and effective Arabic language instruction based on teaching methods and individual student responses. This provides an in-depth understanding of how AI can adapt to each student's learning needs in the context of Arabic language education. In addition, Aljanabi's study (2024) provided a comparative evaluation of two AI tools, ChatGPT and Claude, in terms of their ability to accurately analyze Arabic sentences. Five sentences embodying diverse linguistic features were selected, and three Arabic language experts evaluated the analysis outputs. The results revealed a clear performance gap, with Claude outperforming ChatGPT in overall accuracy (72.9% vs. 33.3%). Claude also excelled at handling morphological complexities and basic grammatical relationships, but struggled with idiomatic expressions and ambiguous constructions. In contrast, ChatGPT struggled with complex morphology. The study concluded that it is important to develop specialized Arabic language processing tools, recognizing the potential of general language models and the need for further fine-tuning.
At the level of using AI technology in the preservation and access of digital manuscripts, Kocoń et al. (2023) emphasize the ability of ChatGPT to provide detailed and accurate answers in various fields. Academicians use ChatGPT for research-related tasks such as drafting articles, generating ideas, and summarizing information (Rahman et al., 2023). For medievalists studying late medieval Burgundian manuscripts, ChatGPT provides relatively accurate translations of Old and Middle French source text into modern English (or another common language) (Caers, 2024). Furthermore, Rockenberger (2023) conducted two experiments on automated text recognition in historical documents using ChatGPT 4 and found errors in recognizing lowercase letters, but was able to transcribe the text.
Study results
Manuscripts included in the study and information provided by ChatGPT
Based on Table 1, the fields covered by the manuscripts included in the study varied: interpretation of the Holy Qur’an Tafseer, Prophet Mohamed's Speech “Hadith Shareef”, Islamic history, language and literature, etc. Out of the 23 manuscripts, ChatGPT provided information about only 14, which represents 60.9%. Among the manuscripts about which ChatGPT did not provide information are: Ṣowar al-Kawākeb by Abd al-Rahman al-Sufi, al-Shāfiyah fī al-Taṣrīf by Ibn al-Hājib, Ethaf Al-Akhsa Befadha’el Al-Masjed Al-Aqsa by Ibrahim bin Muhammad al-Asyūti, Moltaqā al-Abḥor by Ibrahim bin Muhammad al-Ḥalabī and Dhahiri Fatwas by Dhahir al-Din al-Marghinānī. Also excluded are Ershad Al-Aql Al-Saleem Ela Mazaya Al-Ketab Al-Kareem by Abū al-Su‘ūd Effendi, Murshid al-Anām ilā Dār al-Salām by Qūrd Effendi Muhammad ibn ‘Umar, Sharḥ Mirāḥ al-Arwāḥ by Ahmed Danqūz and Risālat Hirmis al-Muthalath bi-l-Ḥikmah (author undetermined).
The importance of the manuscripts included in the study
Descriptive information provided by ChatGPT about the manuscripts included in the study
In some cases, ChatGPT did not provide an answer to the question, and when the question was rephrased, it gave an inaccurate response. When asked about the manuscript al-Kāfiyah fī al-Naḥw, it did not provide an answer. On the second attempt, it provided information about another grammar manuscript called al-Kāfiyah by al-Jawharī (Figure 2). However, when asked directly about al-Kāfiyah ibn al-Hājib, an accurate response was obtained. This discrepancy can be attributed to the fact that there may be multiple books with the same title in Arabic pre-modern literature (Figure 3). This suggests that it is preferable to search for a manuscript on ChatGPT by its title and author to ensure the accuracy of the results. In the case of the manuscript Al-Mukhtār min al-Nawādir wa al-Akhbār whose author is unknown, ChatGPT did not give the author's name because unknown authorship is normal.

The second attempt to ask ChatGPT about the manuscript of Al-Kafiya in grammar, but it provided information about another grammar manuscript.

ChatGPT's accuracy improves when asked about a manuscript by title and author name.
Sometimes ChatGPT gave inaccurate answers, as shown in Table 4. This is consistent with Cheng et al.'s study (2023) which indicates that ChatGPT cannot be fully relied upon because it sometimes “hallucinates” and makes errors in reasoning, and GPT-4 suffers from a limited ability to separate facts from statistically supported incorrect data, and that it shows various biases in its well-known outputs. However, OpenAI, the company that developed it, has initiatives to reduce its biases in order to provide guarantees and reasonable default behaviors that reflect the shared values of the community.
Examples of errors in ChatGPT when asked about the metadata of manuscripts.
Providing a summary of the manuscript by ChatGPT
Table 5 shows a classification of the levels of summary provided by ChatGPT about manuscripts (general summary of the manuscript, summary of the chapters of the manuscript, summary of each chapter, manuscript evaluation). Based on Table 5, with regard to the level of summary provided by ChatGPT (general summary of the manuscript, summary of the chapters of the manuscript, summary of each chapter, manuscript evaluation), ChatGPT is limited to providing a general idea or summary of the manuscript topic in 85.7% of cases, and a brief evaluation of the manuscript in 71.4% of the manuscripts included in the study (14 manuscripts). In the case of the manuscript Mukhtaṣar al-Futūḥāt al-Makiyyah, ChatGPT incorrectly guessed the subject of the book and relied on this incorrect guess for the title and summary of the content (Figure 4). As an example of a brief evaluation of a manuscript, ChatGPT mentions the manuscript al-Mukhtār min al-Nawādir wa al-Akhbār as an important historical work dating back to the Middle Ages in the Islamic world. It is distinguished by its combination of historical events, news and popular stories from the Islamic world at that time. The manuscript contains a variety of stories, historical tales, myths and cultural and religious news. It is noted for its engaging style and detailed presentation of events and figures, making it a valuable source for studying history and culture in the Islamic Middle Ages.

ChatGPT relies on incorrectly guessing the title of the manuscript Mukhtaṣar Al-Futūḥat al-Makkiyyah by Sulayman Al-Aqḥiṣārī.
Levels of summarization provided by ChatGPT about manuscripts.
Providing access to manuscripts
Based on Table 6, ChatGPT did not provide any links through which a digital copy of the manuscript can be viewed or downloaded. In only six out of 14 manuscripts included in the study, representing 42.8%, ChatGPT suggested the names of libraries that possess digital copies of the manuscript in question. For example, when asked about the manuscript Anwār al-Tanzīl wa Asrār al-Ta’wīl, ChatGPT apologized for not providing digital copies or direct links to download books or digital copies. Instead, it indicated the possibility of reviewing digital platforms such as electronic libraries, bookstores or using popular search engines to find sites that offer this book in electronic format or allow for the download of the manuscript in PDF format. ChatGPT also reminded users of the importance of respecting copyrights and not downloading books illegally or distributing them without permission from the publisher.
Providing links or locations of manuscripts.
Additional information provided by ChatGPT about Islamic manuscripts
Based on Table 7, the additional information and advice provided by ChatGPT to those searching for Islamic manuscripts can be divided into three categories:
Additional metadata that researchers found to be mostly very general. In the case of the manuscript Nuzhat al-Anām fī Maḥāsin al-Shām, ChatGPT suggested communicating with researchers and experts in the field of Arabic studies and history, which is general and intuitive advice, and it also suggested some questions that can be used to obtain more accurate results. In the case of Dīwān Ibn al-Wardī, ChatGPT advises contacting “academic libraries, research centers, and historical institutions” to obtain a copy of the manuscript or access the printed text. This is a general and unhelpful advice. In the case of the manuscript Tārīkh Ḥukamā’ al-Islām (History of the sages of Islam), ChatGPT offers a very general advice: “If you are looking for sources on Islamic sages and their history, there are a few options you can explore: historical books that deal with Islamic civilization and history, and include a section on Islamic sages. Encyclopedias and academic sources in Islamic studies and history to find articles and sources that talk about Islamic sages.” Sometimes ChatGPT is limited to advising to search on the Internet, and sometimes it suggests the names of specific sites such as the manuscript of al-Tabari's interpretation of the Qur’ān, entitled: Jāmi‘ al-Bayān ‘an Ta’wīl Āyy al-Qur’ān. (“Copies of Jāmi‘ al-Bayan can be found online from several sites and platforms that provide digital copies of this interpretive work. Some of these sites offer printed copies of the book, while others provide digital copies that can be searched and browsed. You can use search engines to find these sites, and then search for Jāmi‘ al-Bayān to find copies available online. Some popular sites that provide a wide library of Islamic books include the ‘Maktabat al-Islam’ site, the ‘Waqfiya’ site, and the ‘Al-Maktaba al-Shamilah’ site …”). The advice provided by ChatGPT regarding summary information is sometimes not useful because it is axiomatic and general, and the readers know it already. For example, in the cases of Tafsīr al-Ṭabarī, Saḥīḥ al-Bukhārī, Tafsīr al-Bayḍāwī, and Dalā’il al-Khayrāt, ChatGPT gave a very general advice such as: “It is recommended to consult reliable sources and recognized references to obtain a comprehensive and reliable summary.” In the case of the Nahj al-Balāgh manuscript, ChatGPT said: “We note that summarizing the Nahj al-Balāgh manuscript will not be a substitute for reading the full text and benefiting from its details and context. It is recommended to consult reliable sources and recognized references to obtain a comprehensive and accurate understanding of the Nahj al-Balāgh manuscript.” ChatGPT also provided similar suggestions in the cases of Mukhtaṣar al-Kāfiyah by Ibn al-Hājib, al-Mukhtār min al-Nawādir wa al-Akhbār, Nuzhat al-Anām fi Maḥāsin al-Shām, Dīwān Ibn al-Wardī, and Tārīkh Ḥukamā’ al-Islam (Figure 5). For tips on obtaining copies of the manuscript, ChatGPT typically does not provide information about the location or availability of the manuscript. However, in the cases of Tafsīr al-Ṭabarī

Advice of ChatGPT to help deepen research on the manuscript topic.
Additional information provided by ChatGPT about manuscripts for Researchers Seeking Manuscripts.
Discussion of study results
In this study, the accuracy and completeness of the information provided by ChatGPT as an AI technology about Islamic manuscripts was analyzed from four aspects: the identification and descriptive data of the manuscript; its summary; access to a digital copy of it; and any other details provided by ChatGPT about the manuscript. The results of the study showed that ChatGPT provided information about 14 out of 23 manuscripts included in the study, at a rate of 60.9%, while it did not provide information about nine manuscripts, at a rate of 36.1%. The researchers in this study consider this percentage to be weak and not commensurate with the importance of Islamic manuscripts, especially since the manuscripts included in this study are well-known, and they are not manuscripts that can be described as unknown to most researchers. At the level of descriptive or identifying data related to the manuscript, ChatGPT provided only four descriptive elements for the manuscripts it provided information about (14 manuscripts, 100%): the title for all manuscripts it provided information about (14 manuscripts, 100%), the subject (10 manuscripts, 71.4%), the formal features (eight manuscripts, 57.1%), and the author (six manuscripts, 42.8%) (see Figure 4). This result is inconsistent with the findings of Johnson et al. (2023), which indicated that ChatGPT provided answers with high degrees of accuracy and completeness when addressing questions of various difficulty levels in the medical field. Their study suggested that integrating language models such as ChatGPT into medical practice yields promising results at an early stage.
In addition, the researchers noticed that ChatGPT did not provide important descriptive data for them as manuscripts but rather treated them as books or general texts; it did not provide important descriptive data related to the indexing of manuscripts, such as the manuscript's copyist, the year of copying, the place of copying, the classification code, and the library that holds the manuscript. Also, ChatGPT didn’t provide any data regarding these descriptive elements either during the search or in response to questions.
Regarding the level of summary provided by ChatGPT about the manuscript (general summary of the manuscript, summary of its chapters, summary of each chapter, evaluation of the manuscript), ChatGPT is limited to providing a general idea or summary of the topic of the manuscript in 85.7% of cases, and a brief evaluation of the manuscript in 71.4% of the manuscripts included in the study (14 manuscripts) (Figure 4). This means that the summaries provided by ChatGPT may not be sufficient for those seeking a comprehensive understanding of the manuscript through its summary. These results are compatible with the study conducted by Lund and Wang (2023) that indicates that while ChatGPT can summarize texts and support researchers by providing answers to specific questions in their field of study, the limited depth of these summaries may not meet the expectations of researchers who need detailed evaluations. Also, ChatGPT did not provide information about the year in which the manuscript was written but rather limited itself to indicating the century in which the author lived. Regarding access to the manuscript, ChatGPT did not provide any links through which a digital copy of the manuscript could be viewed or downloaded. In only six of the 14 manuscripts for which ChatGPT provided information (42.8%), ChatGPT suggested the names of libraries that hold digital copies of the manuscript in question (Figure 6).

Indicators of the accuracy of information provided by ChatGPT about Islamic manuscripts included in this study.
In terms of advice or additional information provided by ChatGPT about the manuscript, it was mostly very general, such as communicating with researchers and experts in the field of Arab studies and history, which is general and clear advice, but sometimes it suggests some questions that can be used to obtain more accurate results. Some of the suggestions provided by ChatGPT may be useful to the average reader and not the researcher specializing in manuscript studies, such as advising users to review historical books on Islamic civilization and history, encyclopedia and academic sources in Islamic studies to find articles and sources related to the subject of the manuscript. As for the advice provided by ChatGPT to obtain copies of the manuscript, it usually does not provide information about the location of the manuscript or the availability of a digital copy of it. However, in some cases, ChatGPT suggests the names of libraries and sites to search for manuscripts. This happened when searching for manuscripts such as Tafsir al-Tabari, Sahih al-Bukhari, Dalā’il al-Khayrat, Nahj al-Balagha, Diwan of Ibn al-Wardi, and Nahj al-Burdah. Researchers believe that this is a positive aspect of the search result using ChatGPT that can be improved by increasing the training models provided to generative AI applications such as ChatGPT, which requires supporting efforts related to making Islamic manuscript heritage available through a thoughtful digitization process that relies on presenting the manuscript in the form of an image first, which highlights the aesthetic values present in Islamic manuscripts (Omzdi, 2024) and benefiting from OCR techniques to increase the retrieval capabilities of manuscripts and maximize the benefit from their content, and benefiting from AI applications in converting manuscript images into texts, which allows access to manuscripts on a wide scale. This trend will increase the capabilities of AI applications to provide more accurate information about Islamic manuscripts. Some recent studies indicate an improvement in the results of applying OCR technology to manuscripts dating back to the Middle Ages, where it achieved an accuracy rate of 97% (Schoen and Saretto, 2022).
Memon et al.’s (2020) study also confirms that the techniques used in OCR in manuscripts make retrieving the required information easier, and that they mainly depend on extracting features and classifying these features based on patterns, which allows the extraction of textual content for further analysis. Digitizing manuscripts and benefiting from AI applications will support the creation of a database for digitized manuscripts that includes all the physical and intellectual features of various forms of manuscripts and helps users to view digital copies without the need to refer to the original manuscripts except in special cases. In this regard, Rockenberger (2023) conducted two experiments at the level of automated text recognition in historical documents using ChatGPT 4 and found that small things such as letters were incorrectly recognized, but the results improved after several times of training and improving the HTR model. To provide access to manuscripts, Hassen and Khemakhem (2023) recommended creating a platform for digitizing Islamic manuscripts that facilitates storage, retrieval, and processing, as well as rapid access and unlimited storage capacity. This approach was agreed upon by Al-Hajji (2016), who recommended creating a website through which a single search strategy could be implemented, and all manuscript collections could be accessed. One of the distinguished experiences of applying AI technologies in smart processing of Arabic manuscripts and documents is Zenki Engine developed by Zenki Company for Arabic Optical Character Recognition. The forms of smart processing that can be performed on Arabic manuscripts include improving image quality, document segmentation, document layout detection, line and word division, letter optical character recognition, and image-to-text conversion. Also, the SEBR system also enables institutions, governments, and specialists to convert the archive of photographed or scanned Arabic documents into a searchable digital archive using the latest AI methods (SEBR system, 2024).
In terms of data accuracy, ChatGPT as a generative AI system is an important tool that helps extract and display information from different sources, especially in research environments that care about data accuracy and reliability. Although Kocoń et al. (2023) confirm ChatGPT's ability to provide detailed and accurate answers in various fields, in some cases ChatGPT did not provide an answer to the question, and when the question was rephrased, it gave an inaccurate answer. This is consistent with the study by Cheng et al. (2023), which indicates that ChatGPT cannot be fully relied upon because it sometimes “hallucinates” and makes errors in reasoning. Therefore, the researchers in this study believe that ChatGPT, as a model for generative AI applications, needs to be trained on models of Islamic manuscripts and their metadata to improve the accuracy of the information it provides about them.
However, the researchers believe that the difficulty ChatGPT faces in retrieving information about Islamic manuscripts can be attributed to several reasons, including:
the inadequacy of reliable sources provided to ChatGPT about Islamic manuscripts, which means that it needs to update its sources with more accurate and comprehensive information about Islamic manuscripts; digitization projects not covering a wide range of Islamic manuscripts, posing a challenge to leveraging AI applications to support access to these manuscripts; AI applications, including ChatGPT, needing to be trained on Islamic manuscripts, to perform OCR on handwritten manuscripts, make them searchable and facilitate access to content. This is also necessary for AI models trained on Arabic calligraphy to recognize the scripts of Islamic manuscripts.
Thus, the researchers believe that AI can play an important role in preserving, analyzing, and studying Islamic manuscripts. One area where AI can support Islamic manuscripts is digitization and preservation, making them accessible to a global audience. Handwritten texts in Arabic, Persian, and Ottoman can be converted into machine-readable formats. Artificial intelligence can also help restore damaged manuscripts by reconstructing lost or faded text, recognizing handwriting, and analyzing texts. Artificial intelligence models trained on historical Arabic script can recognize and copy different Islamic scripts such as Naskh, Kufic, Thuluth, and Diwani. In this regard, machine-learning algorithms can help read complex handwritten texts, even those written in ancient or rare styles. Topics for future studies that can focus on Islamic manuscripts include evaluating AI capabilities in translating and interpreting texts, automatically classifying and thematically analyzing manuscript texts using natural language processing tools. In addition, verifying the authenticity of manuscripts, using AI-powered semantic search engines will improve access to Islamic manuscripts by understanding context rather than just keywords, and using AI-powered chatbots and virtual assistants to help students study Islamic texts.
Study suggestions
The results of the study showed that ChatGPT provided responses when asked about the Islamic manuscripts included in the study (23 manuscripts) at a rate of 60.9%. At the level of identifying or descriptive data for the manuscript, ChatGPT was limited to descriptive data: title, subject, physical features, and author. Regarding providing a summary of the manuscript, ChatGPT was limited to providing an idea or general summary of the manuscript's subject in 84.6% of cases, and a brief evaluation of the manuscript in 76.9% of the manuscripts included in the study (14 manuscripts). Furthermore, ChatGPT did not provide any links through which a digital copy of the manuscript could be viewed or downloaded.
Considering these findings, the study proposes the following measures to improve the results provided by ChatGPT and other AI applications regarding Islamic manuscripts:
Encouraging the digitization of Islamic manuscripts can help store and analyze them using AI, providing an opportunity to extract new and useful information from these manuscripts. Training AI tools on manuscript databases can help improve the accuracy and effectiveness of ChatGPT, enhancing its ability to provide accurate and reliable information on Islamic manuscript heritage. Learning recurring patterns and characteristics in these manuscripts can be used to improve its ability to analyze and present information. Developing hybrid retrieval systems that combine AI and human intelligence can increase accuracy and enhance the capabilities of AI tools. Artificial intelligence can analyze data and provide accurate and reliable answers, while human intelligence can identify the need for additional information or clarifications, improving the accuracy of information and providing a more personalized and satisfying experience for users seeking to benefit from Islamic manuscripts. Arabic manuscript platforms should provide, in addition to the identification and descriptive data of the manuscript, several levels of manuscript summary (general summary of the manuscript, summary of the chapters of the manuscript, summary of each chapter, manuscript evaluation) to enhance and improve the results provided by AI technologies about manuscripts. Those responsible for digital platforms of Arabic and Islamic manuscripts should ensure that these manuscripts are available in digital form (both image and text) to support the capabilities of AI technologies in accessing the manuscript and its components. Those responsible for AI technologies should focus on feeding these technologies, including chat applications (such as ChatGPT), with prior text sequences about Arabic and Islamic manuscripts. Additionally, they should provide digital platforms that offer reference services about manuscripts, making them available and evaluating them to support research and studies focused on Islamic manuscripts, thereby maximizing the expected benefits from them. The AI/ML tools can be used to analyze handwritten Islamic manuscripts to convert them into an electronic format to support text analysis and manuscript content research. Investing in AI applications to enrich or correct the descriptive data of manuscripts and support the classification of manuscripts and the images and figures they contain.
Conclusion
This study supports the initiatives undertaken in several Islamic countries to establish digital platforms of Islamic manuscripts and can benefit from the capabilities of AI as a means of retrieving information about manuscripts for those searching for them. Therefore, if efforts are made to support the application AI in digital platforms for Islamic manuscripts, especially for their indexes, it will increase their capabilities in providing more accurate and complete information about Islamic manuscripts. By examining ChatGPT's capabilities in this area, the study aimed to highlight its effectiveness in increasing our knowledge of Islamic heritage. Although ChatGPT is widely recognized as providing many opportunities in various fields, users still need to use it correctly, review the data it provides, and ensure its accuracy and completeness, as it is a chatbot and not an AI tool specialized in Islamic manuscripts or Islamic heritage in general. Based on the results of the study, it can be concluded that ChatGPT showed a medium level of effectiveness in responding to inquiries related to Islamic manuscripts. However, its capabilities were primarily limited to providing descriptive data such as title, subject, physical features, and author. When it came to providing summaries and evaluations of manuscripts, ChatGPT showed a higher level of efficiency, as it was able to provide general summaries and brief evaluations of most of the manuscripts examined. However, a notable negative point is the lack of links to access digital copies of manuscripts, which would have enhanced the user's experience and provided more comprehensive information. Although ChatGPT has proved to be efficient in providing descriptive and evaluative information about Islamic manuscripts, there is still room for improvement, especially in facilitating access to digital copies for further exploration and research. Further improvements to its capabilities could increase its usefulness in assisting with manuscript studies inquiries. It can provide valuable information about Islamic manuscripts if it is fed with previous text sequences from trusted digital platforms. This will maximize the utility of ChatGPT for both researchers and readers.
Footnotes
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
