Abstract
Classification of research articles into different subject areas is an extremely important task in bibliometric analysis and information retrieval. There are primarily two kinds of subject classification approaches used in different academic databases: journal-based (aka source-level) and article-based (aka publication-level). The two popular academic databases- Web of Science and Scopus- use journal-based subject classification scheme for articles, which assigns articles into a subject based on the subject category assigned to the journal in which they are published. On the other hand, the recently introduced Dimensions database is the first large academic database that uses article-based subject classification scheme that assigns the article to a subject category based on its contents. Though the subject classification schemes of Web of Science have been compared in several studies, no research studies have been done on comparison of the article-based and journal-based subject classification systems in different academic databases. This paper aims to compare the accuracy of subject classification system of the three popular academic databases: Web of Science, Scopus and Dimensions through a large-scale user-based study. Results show that the commonly held belief of superiority of article-based subject classification over the journal-based subject classification scheme does not hold at least at the moment, as Web of Science appears to have the most accurate subject classification.
Introduction
Ever since the inception of the digital revolution, growth of electronic databases and Internet-based services, new subject classification systems for academic literature have been evolving and at the same time facing newer challenges due to emergence of interdisciplinary, transdisciplinary and multidisciplinary articles. Traditionally, libraries, encyclopedias and journal publishers have developed several subject classification systems for articles in different contexts. The two main academic databases in existence for some time- Web of Science and Scopus- work at the level of journals, i.e. an article published in a particular journal is assigned a subject category mapped to that journal. This journal-level subject classification of articles has often been criticized on the grounds of inaccuracies.
The recently introduced academic database- Dimensions- uses an article-level subject classification scheme. Unlike the other two existing popular databases, it is not based on journal classifications but uses a machine learning approach that works at article-level rather than its source of publication. Though some initial work has been done on evaluating reliability and validity of subject classification of articles in Dimensions database (see [2]), no work has been done on comparing it with the Web of Science and Scopus databases. Some studies [15–16] tried to compare the accuracy of subject classification of articles in Web of Science and Scopus, but the article-level subject classification schemes adopted by Dimensions database is not included in any such existing comparison. Therefore, it is still not known as to how does Dimensions database perform on subject classification vis-á-vis Web of Science and Scopus databases. A comparison of subject classification among these three databases will not only evaluate their relative accuracies but would also help us understand whether article-level subject classification scheme is actually better than journal-level subject classification.
This paper presents our research work towards this goal. We have compared accuracy of subject classification of the three databases using a large-scale user-based study involving more than 1000 papers annotated by three independent annotators. To the best of our knowledge, this is a first of its kind study on accuracy of subject classification scheme in the three databases. Suitable independent annotators are recruited to annotate accuracy of each paper’s subject classification on a scale of 1 to 3 (with 1 being least accurate and 3 being most accurate). The annotations are then subjected to inter annotator agreement computation. After computing inter-annotator agreement, the accuracy of subject classification in the three databases is evaluated. The results present interesting observations on subject classification in these three academic databases.
Research questions
The paper mainly aims to answer the following research questions about subject classification in academic databases:
Related work
There are several subject classification systems used in different countries for categorizing academic works. Some of these include the US National Science Foundation System (NSF), the Australian and New Zealand Standard Research Classification (ANZSRC) system and the Chinese Journal Classification System etc. However, these are not primarily dedicated to classifying research articles only. Among the research article classification systems, the journal-based subject classification systems of Web of Science and Scopus are the two most well-known and commonly used systems. However, other than these classifications there are several studies that either tried to develop their own schemes of subject classification or developed related algorithmic approaches (such as [1, 20]). The next three paragraph summarizes some of the important points of previous studies on construction of classification systems or related algorithmic approaches.
In an early paper, Gomez et al. [6] studied the delimitation of a research field in bibliometric studies and observed that classification of documents according to thematic codes or keywords was the most accurate method. Glanzel et al. [5] analyzed the subject classification of papers on the basis of assigning journals to subject categories (like those found in the various supplements of ISI databases) and found that it works well in case of highly specialized journals, but fails for multidisciplinary journals such as Nature, Science and PNAS. Glanzel & Schubert [3] proposed a two-level hierarchic system of fields and subfields of the sciences, social sciences and arts & humanities. The system was specifically designed for scientometric (evaluation) purposes with the ultimate goal of classifying every single document into a well-defined category. This goal was achieved using a three-step iterative process. Rafols & Leydesdorff [11] tested the results of content-based classifications of journals of ISI subject categories with the field/ subfield classification of Glanzel & Schubert [3]. They observed that content-based schemes allow for the attribution of more than a single category to a journal, whereas the algorithms maximize the ratio of within-category citations over between-category citations. They, however, found that the differences between the two content-based classification systems were marginal in statistical terms.
Zhang et al. [18] used a clustering algorithm based on journal cross-citation to validate and improve the journal-based subject classification schemes. They compared their approach with the 15-field subject classification proposed by Glanzel & Schubert [3]. They, however, concluded that for the final classification the clustering results should be combined with the intellectual approach. Waltman & van Eck [14] proposed a new methodology for a publication-level classification system of science. They observed that journal-level classification systems have two main limitations: they offer only a limited amount of detail and they have difficulties with multidisciplinary journals. The publication-level classification system proposed by them used clustering of publications into research areas based on citation relations. They proposed that it can deal with large number of publications and that it can be further improved if things like bibliographic coupling could be added to direct citation relations. Lee et al. [8] analyzed the mapping systems between classification systems in order to design a structure to connect a variety of classification systems used in the academic information database of the Korea Institute of Science and Technology Information.
Xu et al. [17] developed new journal classification methods based on the h-index and carried out some experiments for classification of operations research and management science (OR/MS) journals using this index, and compare it with other well-known journal rankings. Ruiz-Castelo & Waltman [12] constructed classification systems by performing a large-scale clustering of publications based on their citation relations. They designed 12 classification systems, each at a different granularity level. The number of fields in these systems ranged from 390 to 73,205 in granularity levels 1–12. Based on an investigation of some key characteristics of the 12 classification systems, they argued that working with a few thousand fields may be an optimal choice.
In addition to studies on developing subject classification systems or related algorithmic approaches, several studies tried to compare the different subject classification systems. Some of these studies are as follows: Wang & Waltman [16] systematically investigated the accuracy of the classification systems of Web of Science and Scopus. They defined two criteria on the basis of direct citation relations between journals and categories. They performed a more in-depth analysis for the field of Library and Information Science to assess whether the proposed criteria were appropriate and whether they yield meaningful results. They found that according to the citation-based criteria, Web of Science performed significantly better than Scopus in terms of the accuracy of its journal classification system. Shu et al. [13] carried out a study for comparing journal and paper level classification of science. They compared the journal- and paper-level classifications for the same set of journals and papers to isolate the effects of classification precision. For paper-level classification system, they used CLC codes of the Chinese Library Classification system. They concluded that almost half the papers could be misclassified in journal classification systems. None of these two studies, however, worked with Dimensions data.
Bornmann [2] was the first study to investigate the reliability and validity of field classification in Dimensions. He used his own set of papers and investigated whether they were reliably and validly assigned to fields. He concluded that the results obtained put in question the reliability and validity of the field classification scheme of Dimensions. Herzog & Lunn [7] responded to Bornmann’s analysis stating that “Dimensions opted for applying a categorization approach using machine learning and based on the content of the documents and well-established classification systems for which a training set was available. The implementation at launch was a first step and requires to be improved”. They further stated that “the implementation of a content-based classification will always be work in progress”.
However, to the best of knowledge there are no studies as on date that compare the accuracy of subject classification schemes of the three most prominent academic databases: Web of Science, Scopus and Dimensions. This paper aims to bridge this gap by comparing the subject classification of the three databases by taking a large sample of papers and adopting a user-based evaluation approach.
Data and methodology
The comparison of accuracy of the three academic databases required us to work on a set of papers that are indexed in all the three databases. Therefore, we obtained a random selection of 2000 research papers published in 2016, ensuring that papers selected are distributed among different disciplines. These papers were those indexed in all the three databases. Thereafter, we obtained the subject classification for each of the papers in all the three databases. Some of the papers did not have a clearly identified subject classification associated with them in all the three databases, therefore we dropped them from the data. Out of the remaining set of papers, we randomly selected 1100 research papers and recorded their subject classification information. Thereafter, relevant information for these 1100 papers from the three databases was obtained. This information included title, abstract, author keywords, journal keywords and cited references for the paper. All this information was put together in a systematic manner to design a user-based study.
For the user study, three independent annotators were recruited. The annotators were all doctoral research students working in the area of Scientometrics. The annotators were asked to rate the accuracy of subject assignment to a paper on a scale of 1 to 3, with 3 being most accurate and 1 being least accurate. Annotators had to thus annotate the accuracy level of each paper for all the three databases. This resulted into each annotator doing 3300 annotations. All three annotators taken together did 9900 annotations. It took about three months’ time for the annotators to complete the annotation process. After the annotations of all the three annotators were available, we computed inter-annotator agreements by computing Inter Indexer Consistency (IIC) and Cohen’s Kappa statistics. Out of the 1100 papers, we dropped the 100 papers that had largest annotation variations. Finally, we were left with 1000 annotated papers. Table 1 shows the details of data and inter-annotator agreement measures computed. The values obtained show that there was a reasonable degree of agreement among the three annotators and hence the annotations can be used for evaluation of accuracy of subject classification.
Inter Indexer Consistency of annotations
Inter Indexer Consistency of annotations
The data samples annotated by three independent annotators was analyzed to determine the ratings of accuracy. It may be noted again that while the first two databases use journal-based subject classification the third one uses paper-based subject classification. Annotations for the same set of papers were done with respect to classification accuracy in each of the three databases: Web of Science, Scopus and Dimensions. The accuracy ratings of the three annotators for each paper in a particular database were aggregated by computing mean of the three ratings. Since the ratings ranged between 1 to 3, the mean values computed also ranged between 1 to 3, with highest value of 3 (when all the three annotators rate the accuracy of the paper as 3) and the lowest value of 1 (when all the three annotators rate the accuracy of the paper as 1). Thereafter, the mean accuracy ratings of the 1000 papers for a particular database were aggregated into a single value summary by computing mean and median values. Table 2 shows the mean, median and variance and standard deviation values of aggregated accuracy ratings for all the three databases.
Statistical results of accuracy of the three databases
Statistical results of accuracy of the three databases
It can be observed that Web of Science shows the highest aggregated accuracy values (mean = 2.489 and median = 2.666) for the data sample. The Dimensions database is also quite close to the accuracy levels of Web of Science, with a mean value of 2.474 and median of 2.665. Scopus database has the lowest values of mean (2.342) and median (2.333). Thus, it can be clearly seen that Web of Science and Dimensions databases are quite close in accuracy levels, whereas Scopus is a little less accurate.
It may further be observed that Web of Science has lowest variance (0.228) in subject classification accuracy levels as compared to Scopus (0.239) and Dimensions (0.266). This indicates that Web of Science in general has more accurate and consistent subject classification as compared to the other two databases. Figures 1, 2 and 3 plot the accuracy levels of Web of Science, Scopus and Dimensions, respectively. The figures show more detailed results of accuracy of classification. It can be observed that Web of Science has more frequency of higher accuracy ratings as compared to the other two databases.

Histogram of average accuracy values in Web of Science.

Histogram of average accuracy values in Scopus.

Histogram of average accuracy values in Dimensions.
The results present an interesting observation, particularly taking into account the fact that Dimensions uses article-level classification whereas Web of Science uses journal-level classification. Thus, at the moment, Web of Science, with its journal-level classification, appears to be as good as (in fact a little better) than Dimensions database, which classifies each individual article into a subject category based on its content.
This may have two broader implications. Either the currently used approach in Dimensions for individual article’s subject classification is not fully mature and requires improvement or that the article-level subject classification of papers may not be necessarily superior to journal-level classification, as is usually believed. Herzog & Lunn’s [7] response to observations of Bornmann [2] about accuracy of subject classification of Dimensions appear to substantiate the first implication.
The paper presents a user-based study for evaluating the subject-classification accuracy of the three academic databases: Web of Science, Scopus and Dimensions. A large sample of 1000 research papers is used for analysis. Each paper’s subject classification in all the three databases is rated by three independent annotators on a scale of 1 to 3 for classification accuracy. The inter-indexer consistency values are found reasonable enough to use the annotations for evaluation of subject classification accuracy. Single value summaries of classification accuracy for the three databases is computed by computing mean and median. The standard deviation and variance values of the accuracy levels are also computed.
The results show that Web of Science has the most accurate subject classification followed by Dimensions and Scopus. Web of Science and Dimensions are in fact very close in accuracy levels. There are three main conclusions that can be drawn from these results, which also answer the research questions proposed. Firstly, out of the three databases, Web of Science still has the most accurate subject classification, followed by Dimensions and Scopus. Secondly, at the moment, article-level classification may not be necessarily better than journal-level classification, as is commonly believed. Thirdly, the new database Dimension is at least as accurate as Web of Science in subject category assignment, which is an impressive performance for a newly created academic database. It would be interesting to see how does Dimensions database further improves its subject classification accuracy, taking into account the advantage it may have in terms of using an article-level subject classification scheme.
Footnotes
Acknowledgments
The authors would like to acknowledge the support provided by the DST-NSTMIS funded project ‘Design of a Computational Framework for Discipline-wise and Thematic Mapping of Research Performance of Indian Higher Education Institutions (HEIs)’, bearing Grant No. DST/NSTMIS/05/04/2019-20, for this work.
