Abstract
To survive and prosper in a highly competitive environment where uncertainty and ambiguity are the norm, today’s firms are faced with the need for new information management methods and tools. Two of the most prominent strategies that take information and its treatment as a value-generating element in firms’ decision-making are Technology Watch and Competitive Intelligence. In addition, one of the fundamental components that a system based on these strategies must have is an efficient method of Information Retrieval. The present study describes a Competitive Intelligence–based decision-support system that uses a Genetic Algorithm. The system contributes to improving information retrieval through search optimisation, thus enhancing the performance of this knowledge-generating tool for organisations.
Keywords
1. Introduction
The processes of the globalisation of economic activities have given rise to a new scenario in which firms no longer compete only in local environments but in global environments as well. Knowledge of what is happening in other, sometimes very distant, countries is thus as vital for their development and survival as what happens in their immediate surroundings. Also, the fact that data are everywhere, easily accessible and inexpensive; that innovation cycle times are getting ever shorter; and that innovative firms demand ever more information of all kinds to increase their competitive advantage brings about the need to seek methods that can support complex decisions within organisations [1]. All of this implies, for example, the introduction of new technological products and services, and the use of information within the innovation value chain [2,3]. For organisations to make the right decisions requires them to use information that can help them establish, prioritise and develop strategies; anticipate threats; identify opportunities which generate change; develop new technologies and products; open up to new markets; and set up strategic alliances and thus improve their competitive position in a changing environment.
One consequence of this changing environment has been that Intelligence, as an integral process of (a) identification of information sources; (b) acquisition, treatment and processing of that information; (c) analysis; and (d) dissemination and decision-making, is applied to the business environment [4].
Prescott [5] points out that there are three currents in Intelligence: (a) the first is based on Sun Tzu’s work ‘The Art of War’ and is applied to the military field; (b) the second is focused on national security as a political issue which, although it had been developed over previous centuries, gained strength after World War II and is linked to Political Science; and (c) the third is Intelligence applied to the world of business, the economy and the firm. This third current can be viewed from two perspectives: ‘Economic Intelligence’, understood as that conducted by nation states to protect their economies against external threats, with its focus on protecting a country’s economic interests; and ‘Competitive Intelligence’ (CI), which is what firms implement to aid them in their decision-making. All three currents fundamentally imply the capture and management of external information, although the concept of Intelligence as such is polysemous and can also imply the management of data and information within organisations [6].
The focus of the present case is on CI. Although the appearance in 1986 of the Society of Competitive Intelligence Professionals (SCIP) in the United States was a catalyst for the discipline (in 2010, it changed its name to Strategic & Competitive Intelligence Professionals so as to include all professionals of competitive and strategic intelligence), there had been earlier studies and analyses of the importance of information in business management (such as its acquisition or protection). For instance, there has been literature about the need to protect information from competitors since the first third of the 20th century [7–11]. Businesses’ awareness of the need for detailed knowledge about their own environment was subsequently accelerated by such studies as those of Aguilar [12] or Porter [13].
The study by Madureira et al. [14] analyses how the term Competitive Intelligence has on many occasions been defined polysemously, and how the studies carried out have taken multiple approaches, treating it as a process, a product, a tool, an ability, knowledge, a discipline and so on. It has sometimes also been approached as an art [15] or a science [16]. With regard to which reports or environments are necessary to have knowledge of, Management manuals have been quick to respond and at the same time resolve a terminological problem, since it is usual for them to associate the context of the knowledge required with the term Intelligence and thus create multiple different intelligences. For example, in their ‘Management: Quality and Competitiveness’ manual, Ivancevich et al. [17] distinguish in the exterior of organisations (a) the external task environment (comprising suppliers, competitors, employees, shareholders, creditors and customers) and (b) the remote external environment (comprising the socio-cultural, economic, technological, political-legal and ecological environments). Rooted within these is the need to obtain information of very diverse types with which the Intelligence cycle in firms can begin.
Seeking a general definition, Madureira et al. [14] understand CI to be the process and prospective practices used to create knowledge about the competitive environment so as to improve an organisation’s performance. Bulger [1] points out that traditional CI has migrated towards an integrated intelligence formed by a ‘pool’ of intelligences.
Together with the concept of CI which assumes the integral Intelligence cycle, there is that of Surveillance. The Spanish standard on Intelligence and Surveillance ‘UNE 166006: 2018’ [18] clearly distinguishes between the two concepts. Intelligence is more focused on information analysis and generating reports, is more strategic in nature, and makes a greater contribution to value. Surveillance is more focused on collecting information, exploiting sources and generating alerts, is more operational in nature, and makes a lesser contribution to value. The concept of Surveillance has traditionally been linked to Technology Watch (TW), although nowadays it tends to be mentioned in an isolated manner – only as Monitoring – since it is understood that the decision-making process is influenced not only by technology but also by other elements that are necessary to stay informed about [18]. In this study, we shall be referring to the terms CI and TW. It is thus understood that within the Knowledge Economy – as a discipline that takes information and its treatment as a value-generating element in firms’ decision-making – CI and TW are two of the most outstanding strategies.
These systems are fed by many types of sources of information, among which are those of the so-called OSINT (Open Source Intelligence). As a product, this Intelligence is in turn fed by open information resources, that is, information legally obtained from public access sources whose main channel is the Internet. These sources conform the principal, and most accessible, resource for Intelligence that a firm can develop for its Intelligence strategy, and their use is not only justified, but imperative [19,20]. The development of TW/CI tools that use this type of information is aimed at providing new methods and techniques that can contribute to better and more effective management of these intangible resources [21,22].
A fundamental tool within TW/CI systems is that of the Decision-Support System (DSS). These systems represent the framework supporting the selection, collection and interpretation of external information, and their data are related to that generated within the organisation with the intention of obtaining knowledge about the opportunities and threats the organisation faces. One aim is to identify how economic, social, technological and political circumstances are affecting the present situation, and to try to anticipate future scenarios so as to obtain information that will lead to a competitive advantage.
Numerous research studies in this field have identified the knowledge base as being the most significant factor for achieving success in strategic decision-making, since decision quality increases when decision-makers have more detailed knowledge, and this in turn can reduce the complexity in anticipating future changes in the organisational environment. Ultimately, when there is access to more detailed knowledge, the quality and speed with which decisions are made increase [23].
Therefore, a fundamental component of a TW/CI system is an efficient Information Retrieval System (IRS), understanding such a system to be one which deals with document databases and that processes user queries to allow them to access relevant information in the shortest possible time. These systems also have to deal with incomplete, unstructured or heterogeneous information [24]. One of the main challenges faced by an IRS is to identify the true need for information that underlies the query made by the user, which generally barely manages to capture the complexity of the needs [25]. In the words of Büttcher et al. [26], Information Retrieval (IR) deals with the representation, search and manipulation of large collections of electronic texts, and the IRS must be able to handle this immense amount of information so that it can be efficiently retrieved when desired. One of the most widely used formal models of IR is the vector space model [27]. It is one of the simplest, as well as efficient, alternatives for processing text data [28] and is the one used in this study. In this model, each document (and query) is a vector in an n-dimensional document space, where n is the number of representative terms used to describe the content of the documents in a collection, and each term represents a dimension of the document space. The retrieval of documents is based on a similarity measure between the query and the documents, so that the documents with a greater similarity to the query are considered to be more relevant, and the retrieved documents are displayed to the user in an ordered manner based on their relevance regarding the query [29].
One of the main areas of research in the development of IRS in recent years has been the application of Artificial Intelligence (AI) techniques [30], and one of the most robust definitions of AI according to this study is provided by Rai et al. [31] who define it as ‘the ability of a machine to perform cognitive functions that are associated with human minds, such as perceiving, reasoning, learning, interacting with the environment, problem solving, decision-making, and even demonstrating creativity’.
It is not surprising that in recent years, there has been a considerable increase in the number of companies that use AI techniques to gain advantages in terms of added business value, such as higher revenues and cost reduction. Competitive pressure is a very important factor when incorporating AI techniques in organisations [32]. This is not only to improve decision-making by applying these techniques to the IRS of the TW/CI system to analyse the information collected from multiple sources, optimising decision-making processes, but these techniques are also increasingly being implemented in human resource management to support and speed up information processes that require a lot of work, such as the evaluation of CVs and conducting numerous interviews [33].
Among the AI techniques applied to IR, there stand out those encompassed within Evolutionary Computation [29,34], a branch focused on solving optimisation problems based on the natural processes of biological evolution. Genetic Algorithms (GAs) are one of the tools of Evolutionary Computation that can be used to develop methods of IR [35]. They have been applied to various of its aspects, with search optimisation being one of the objects of its most numerous and fruitful applications [36], mostly using techniques of relevance feedback (RF) [34,37–39].
RF is a particularly popular automatic search modification technique [27,40]. In it, once certain previously retrieved documents have been identified by the user as being relevant or irrelevant, that information provided is used to adjust the query and thus improve the next round of retrievals. It is an iterative process which is repeated as many times as the user wants [41–45].
The effect of this technique is to ‘shift’ the search in the direction of the relevant documents and away from the irrelevant ones, in the hope of retrieving more wanted documents and fewer unwanted documents in subsequent searches. To improve the decision-making process, it is very useful to have an efficient IRS that includes a search reformulation module that allows the initially formulated query to be improved [46].
The first contribution of this study is to present an overview of the literature about the different concepts and disciplines, such as CI, TW, IR, AI and Evolutionary Computing, as well as the different techniques and related tools that have been used in the development of the application. One of the main notable contributions of the proposed method is the innovative combination of tools or components, such as the DSS and IRS, from disciplines such as TW and CI, or the combination of techniques, such as GAs and RF for query optimisation. The second contribution is the proposal of a methodological framework for the implementation of a DSS based on this novel combination of techniques and tools: Research on GAs aimed at search optimisation has been applied to knowledge-generating tools for organisations. These tools include TW/CI Systems, and one result was the present DSS prototype, called Delphos, which uses a GA to optimise the retrieval of information destined for TW/CI processes. This prototype is detailed in this study.
The article is organised as follows. Section 2 introduces the main concepts and a review of the literature related to the field of this research study. Section 3 presents the method used in the development of the prototype. Section 4 presents and analyses the operation of the tool, which is the main result of this research, and Section 5 describes its evaluation and discusses the results of the analyses. Finally, Section 6 summarises the conclusions, the main contributions as well as the limitations of the study and future lines of research.
2. Background
Borges et al. [47] present a conceptual framework in which four sources of value creation are distinguished from the application of AI tools in business organisations: decision support, customer and employee engagement, automation, and new products and services. The authors note, similar to Enholm et al. [32], two broad categories of AI applications depending on their use: AI for automation and AI for augmentation. Automation refers to systems that are dedicated to replacing human work, while augmentation enhances human intelligence by providing information that can help the decision-making process, that is, AI integration with human experience to improve decisions and optimise actions. The present research would be included in this second category, as also would be the study by Shrestha et al. [48] which presents a novel framework outlining how human and AI-based decision-making can be combined to optimally benefit the quality of decision-making in organisations.
Another review of the literature about the subject can be found in Keding [23], where the sample of articles is classified into two areas of research: condition-oriented research, which explores the background to take advantage of the use of AI in strategic management, and results-oriented research, which studies the consequences of AI in strategic management, at both the individual and the organisational levels. There also stand out research studies applied to optimising decision-making. Duan et al. [49] present 12 research proposals focused on the use and impact of AI for decision-making in terms of theoretical development, technology–human interaction and the implementation of AI. Even large-scale decision-making has become a rapidly developing and emerging research field with extensive studies over the last decade [50], with the authors especially highlighting the potential rise of this complex type of AI-based decision-making. Feedback techniques also markedly improve the level of agreement among large-scale decision-makers [51,52].
Authors in this area of research agree that making better business decisions is often the key to gaining competitive advantage [53]. The DSSs play a key role in this. In terms of the main element with which they operate in TW/CI, five types can be identified. These correspond to the dominant architectural components which provide the functionality to support decision-making [54]: communications based (dedicated to group decision tasks); data based (e.g. the Data Warehouses); document based using Natural Language Processing techniques for content extraction (a good example of this is represented by the Document Warehouses); knowledge based, such as the Intelligent DSS, using AI techniques; and finally model based, which combine statistical, financial, mathematical, analytical, simulation, optimisation models and so on, in order to simulate scenarios and analyse the responses of the variables involved.
Intelligent DSSs are one of the tools used by organisations to implement AI techniques to optimise their performance. In the 2020s, the use of AI in decision-making will probably become widespread [55]. The prototype developed in this present research belongs to this type of DSS.
Depending on the type of treatment the data undergo, it is possible to differentiate between exploration and exploitation DSSs. The former serve to identify the environment, while the latter facilitate its understanding. The combination of the two actions on the data helps with the establishment of strategies and the move to action based on the reduction of uncertainty.
The DSS prototype developed in the present work corresponds to a Dashboard (taking the data exploitation perspective) that uses an application based on searches optimised by Evolutionary Computation techniques. Evolutionary Computation is one of the most promising branches of AI applied to IR. Numerous studies have been carried out using its tools, especially GAs. GAs, like the rest of the tools in this branch, belong to the family of problem-solving models based on Darwin’s principles of evolution and natural selection, and they work on problems in a specific domain using a model based on chromosomes [56,57].
The basis of a GA is dealing with a population of possible solutions to the problem at hand. A series of operations are carried out on this population that alter it with the intention of identifying those possible solutions which offer the best responses to the problem, determining which of them should be kept for future generations and which should be discarded.
The operations that a GA applies to a population are those that Darwin described in his theory of natural selection: mutation, crossover, selection and an adaptation function. Applied to IR, GAs interpret documents as chromosomes whose genes are the different terms that make up each document. The information contained in each gene can be represented in different ways. In the present case, we used binary coding, assigning the gene a value of 1 if the term corresponding to the query is found in the document or a value of 0 if it is not found.
To work with GAs, a series of parameters that must be controlled have to be taken into account before operating with them: the size of the population, the probability of crossover (Pc), the probability of mutation (Pm) and the number of generations. The fundamental thing in this type of algorithm is the adaptation function with which it operates, since this plays a basic role in the correct functioning of the algorithm.
Research on GA-based methods as a mechanism for the identification of optimal solutions in IR that has been developed intensively in recent years demonstrates their robustness as very efficient techniques for large search spaces. The indexing of texts and images, the classification and grouping of documents, and the improvement of user queries are some of the main fields [25,29,35]. The GA that we implemented applies precisely to this last field – the automatic optimisation of searches.
López-Pujalte et al. [37–39] propose an optimal GA based on RF techniques using the vector space model. The tests they carried out showed that the use of GA gave better results than the other RF techniques with which it was compared. These included the one that had proven to be the most efficient of all the classical techniques thus far – Ide’s dec-hi method [58]. In addition, the authors confirmed the importance of the adaptation function, since success or failure of the exploration is directly dependent on it. They concluded that, of all the adaptation functions they tested, the functions that not only take into account which documents are retrieved, but also the order of their retrieval give the best results, that is, these are adaptation functions that not only assess whether the potential solution retrieves many relevant documents and few irrelevant documents, but also whether the relevant documents are at the top or bottom of the list. Similar conclusions are drawn in the research study by Azmi and Kusumaningrum [59], where they point out that, due to the rapid growth of technological developments in Indonesia, it is urgent to develop a new, more efficient IR environment in organisations by implementing RF techniques. The authors conclude that GA is the best-fitting optimisation technique, outperforming the best of the standard optimisation techniques – the Ide dec-hi RF technique.
Gowthul Alam and Baulkani [60] implement an IRS based on clustering that offers the possibility of finding similar documents for a given user query using a GA-based kernel fuzzy c-means algorithm. Sharma et al. [61] propose a new query expansion approach using a hybrid evolutionary algorithm combining cuckoo search and an accelerated particle swarm optimisation technique. The latter’s performance was also improved with the application of fuzzy logic. The work of Gupta and Saini [62] is in the same line, also using particle swarm optimisation for the automatic expansion of queries and fuzzy logic to improve it.
RF techniques have been used successfully not only with text documents but also with images. Stejić et al. [63] propose a method to optimise the calculation of image similarity and, to improve it, a GA is used to perform the RF mechanism, allowing the user to automatically specify the queries, obtaining satisfactory results. Mahmood et al. [64] propose an approach to RF based on a hybrid evolutionary algorithm which is used for image recovery using the combination of particle swarm optimisation and a GA to achieve better results in the first iteration, and thus reduce the user’s interaction with the system.
In view of the benefits that GAs bring to the informational environment, an increasing volume of research is being directed at the development of TW/CI applications with a focus on these techniques, with excellent results. In recent times, GAs have been used for the search among alternatives, multi-objective and multi-criteria decision-making, constrained optimisation problems, job scheduling, selection of suppliers and project selection, among other applications [65].
The search for models predicting the prices of financial markets continues to be a topic of intense research, despite the great challenges that it entails. The prices of financial assets are non-linear, dynamic and chaotic, making financial time series difficult to predict. The most commonly used forecasting models involve support vector machines and neural networks [66] and other AI techniques such as GAs. There stands out the research by Kim and Han [67] which proposes a GA approach to determine the connection weights of artificial neural networks to predict the stock price index, although there is other relevant research on the subject [68–70]. Chou et al. [71] use a GA that employs an adaptation function based on fuzzy logic for a prediction system of financial crises in an organisation. In this same line is the research by Aruldoss et al. [72] in which a GA is applied to generate prediction rules until the most accurate is found.
Gupta et al. [73] propose a system for market analysis in which a GA is used to treat user opinions, obtaining a clear advantage in its optimisation. Anusha and Nallaperumal [74] implement an algorithm to find the optimal client from the immense quantity of clients in industry today by combining a GA and an artificial bee colony algorithm, an extension of Evolutionary Computation.
A GA can be used to improve the methods of displaying CI systems, like in Chung et al. [75] where the GA organises web pages into a tree structure, a format that is more efficient and useful for these systems than a bare list of results.
GAs have also been used in medical applications such as that of Rakhmetullina et al. [76] where the structure of the medical process for the diagnosis of anaemia is described and a mathematical model of GA-based decision rules is developed. Li et al. [77] develop a diagnostic system for heart diseases based on a large-scale multi-objective cooperative co-evolutionary algorithm that yields high precision. Other applications of this type focus on the treatment of medical images with very good results [78–80].
Patents have always been a primary source of information in CI systems. In this line, Galindo-Melero et al. [81] propose a global CI method for small- and medium-sized firms based on TW that is focused mainly on patent analysis. Here we would also include the work of Jürgens and Herrero-Solana [82]. James et al. designed a system to explore innovation by analysing patent applications using a GA to obtain groups of words that appear together in the titles [83].
Finally, directly related to the present study, many of the applications of these algorithms are framed in the field of information systems, focusing on operational and tactical decisions and strategic decisions at the functional, entrepreneurial and corporate levels in fields such as agriculture, urban planning, military, health, education, governance and other sectors of application [65].
Hernández-Julio et al. [84] develop a framework for the design of CI software, in which GAs are included in order to implement decision-support tools. In the same line, Ying and Liu [85] create a decision-making model based on big data analysis of business information that uses interactive GAs to obtain the optimal strategy through experimental tests.
Dawood et al. [86] apply a GA to solve a dynamic problem in the activities of industrial project networks, describing the result as being a tool for decision-support. In this framework, other evolutionary computing tools have also been applied, such as Artificial Swarm Intelligence, to address the limitations associated with group decision-making, amplifying the intelligence of human groups and facilitating the best business decision-making, since they include the perspectives of all group members [87]. These types of applications have also been implemented in the energy market field to improve decision-making and help construct expansion plans, investigating the different possible reactions of competitors [88,89].
Finally, in a recent bibliometric review of scientific literature on CI [6], it can be observed that both the application of AI techniques and the improvement in searches are the main thematic areas currently being investigated in research in this field. It is precisely in this line, among others, that we also have worked in this present study, applying a GA to the improvement of searches in a TW/CI system.
3. Methods
Delphos, the tool developed to help firms in their decision-making, is based on OSINT sources [19,20] and offers three main functionalities [90].
Search and analysis of information from open sources on the Internet.
Automatic surveillance of information, and generation of alerts for terms and/or sectors that it is wished to monitor. Direct access to said information.
Display and analysis of the trends of a term or a sector over a certain period of time, also allowing the generation of alerts when there is a significant change in a trend and the analysis of indicators that serve to identify the possible reasons that may have generated said change.
Functionalities that could be placed within the generic framework of a TW/CI application of this type consist of four layers: data acquisition and preparation, business knowledge learning, decision-making and recommendation [22]. In addition to the searches for information relevant to the company, a key functionality of the system is the automatic monitoring of information through the generation of alerts, and the visualisation and analysis of trends are fundamental elements in a decision-making support tool of this type [32,82,83,87].
In order to establish the requirements of the tool, in the conceptual design phase of the prototype, we had the collaboration of the technological observatory of one of the main Spanish firms dedicated to the sectors of renewable energy, water, environment and infrastructure. It is present in more than 40 countries and has extensive experience in the use of TW/CI techniques and applications. These were used to identify the requirements of the prototype tool and the principal information needs that had to be covered.
The programming language used was Java. The working method followed the Scrum model. This model is inspired by empirical inspection and the adaptation of feedback loops as tools to deal with complexity and risk. It gives more weight to the results of decision-making in the real world versus speculation, and divides time into short working cadences (‘sprints’) of generally between 1 and 2 weeks [91].
The intention was for the product to remain in a properly integrated and tested state at all times. At the end of each sprint block, the interested parties (the researchers and the collaborating firm’s users) met to test the new functionalities implemented in the tool and plan the next steps.
A final assessment of Delphos was carried out with the collaboration of a different firm, an important organisation specialised in the integral operation and maintenance of renewable energy projects. The intention was to achieve an objective assessment of the perception that real users, totally unrelated to the research, have of the system’s functioning.
The Delphos IRS operates with two kinds of documents: structured and unstructured. In this first phase of the prototype, the structured documents were patents, tenders and academic documents. The unstructured documents were web-pages.
To represent some of the essential types of documents to obtain knowledge about organisations’ external environment, our prototype uses external databases of patents, tenders and scientific documents. These sources configure the structured information with which the tool operates. Patents constitute a fundamental type of document in TW/CI tools to protect companies from unproductive or duplicative research (research in which new R&D + i projects give rise to existing technical solutions), which would cause the companies to waste resources and harm their ability to survive against their competitors [81,82].
For unstructured information, our prototype uses the vector space model as a basis [28,38,59], and the application’s IRS relies on the application programming interface (API) of the Microsoft search engine, Bing – the Bing Web Search API (BWSA). On it are configured the different options for filtering, expanding and improving the searches, while also using the GA for search optimisation by implementing the RF technique. To work with unstructured information, first the information needs that were included in the collaborating firm’s TW/CI strategy were determined. Then a search was made for all those web-pages that might contain quality information related to the recognised key intelligence topics (KITs). More than 4800 sources were identified as potential providers of relevant information for the organisation’s different operational sectors.
These sources were classed into six levels of content quality: excellent, very good, good, medium, poor and very poor. Those identified as ‘very poor’ were discarded and added to a list denominated ‘Unuseful sources’. The rest were categorised in accordance with a taxonomy developed to unambiguously classify the sources according to their geographical location, the sector with which they are related and the type of organisation that produces them. In selecting any of these filters, the search was restricted to the information contained on each of the web-pages (complete host) that coincided with the marked category or categories, instead of searching the entire network.
The Delphos design also includes a Web crawler that identifies new URLs corresponding to external links and adds them to the initial list. Thus, the list of sources can benefit from automatic feedback. It is based on the principle that these initial sources, whose quality is recognised, contain links to sources of also high informational quality. The sources extracted from this process inherit the categories assigned to their originators. The maximum depth of this process was set at a single level, that is, no further sources are extracted from those that were retrieved by the crawler. The intention was that the quality of the sources would not be distorted in this automatic process, as it might be as the new sources move away from the initial URLs that were identified, valued and categorised manually.
A parser was created to pick up coincidences between these new automatically retrieved sources and those in both the initial list and the list of automatically discarded ‘Unuseful sources’.
Both crawler and parser are elements started automatically when the Delphos system is initiated, and they extract URLs that point to new hosts which in turn inherit the classification (sector, location and type of organisation) from their originating source, providing automatic feedback for the host listing of the system’s database. Currently, it contains more than 48,500 different hosts.
Of course, the system allows manual source editing in addition to this automatic method.
With respect to searches, when the user enters a query, the first two processes that the search goes through are the elimination of stopwords and reduction through stemming [38,59,83]. Both operations use the Snowball library, an Apache Lucene project based on implementing the Porter algorithm [92].
Once the initial search has been carried out and delivered a series of results, the user can tag those they consider to be relevant to their need for information. This allows automatic execution of the RF technique of search improvement.
The new expanded query will be formed taking into account the terms contained in the title, URL and description fields of the feedback Web documents in addition to those from the initial search, in all cases after subjecting the query and the documents to stopword elimination and stemming processes. If new results are found (different from those already identified as relevant, since a residual collection method is used), they are presented to the user.
The process is iterative. As long as the RF operations continue to yield results that can be presented to the user, execution can continue as long as the user finds relevant documents from which successive searches can continue to be constructed. Then, no further new results can be found, automatically (and transparently for the user) the GA begins to be executed, thus reinforcing and optimising the RF technique.
3.1. Operation of the GA
The GA designed for Delphos uses binary coding, that is, vectors in which 1 appears if the term corresponding to each of the components is found in the representation of the query and 0 if it is not found. As a selection method, it uses global elitist selection by tournament in which the parent chromosomes are all those the user tagged as relevant in the previous RF search process. As operators, it uses simple single-point crossover (crossover probability of 0.7) and simple mutation (mutation probability of 0.05).
The adaptation function is constructed on a pure fitness Minimisation basis. Pure fitness is that measure of adaptation indicated in the natural terminology of the problem itself. An example is the problem put forward by Koza [93] in which a population of ants have to collect as much food as possible. In the pure fitness type of approach, the goodness of each ant is determined by the number of pieces of food it has collected and stored. In Koza’s example, the goodness values range from 0 (in the case when an ant collects no food units) and 89 (the most that can be collected being the total number of grains).
Pure fitness is determined by the following formula
where r(i, t) = goodness of individual i in generation t; s(i, j) = desired value for individual i in case j; c(i, j) = value obtained by individual i in case j; and Nc = number of cases.
In the case of the ants proposed by Koza, the goodness of an individual would be given by (substituting in the above formula)
In maximisation problems, where high values are sought as in the ant colony, individuals with high pure fitness will be the fittest. On the contrary, in minimisation problems such as the present case (and others, e.g. Mahmood et al. [64]), the best solution will be represented by individuals (query vectors or chromosomes) with low pure fitness.
Standardised fitness attempts to solve the duality of minimisation or maximisation in the face of problems, and reaffirms pure fitness so that a lower numerical value is always a better value.
If for a particular problem a lower value of pure fitness is better, the standardised fitness equals the pure fitness for that problem, and the best value desirable for standardised fitness would be equal to 0.
If the problem is one of maximisation, it is subtracted from an upper bound (rmax) that represents the maximum desired value the pure fitness value obtained by the individual, (r(i, t)).
So for minimisation problems, the standardised fitness would be represented as follows
For maximisation problems, the standardised fitness would be
Continuing with the example of the ants, if an individual collected five grains
The closer to 0 the value obtained, the better the goodness of the individual, so within a generation Ant i collects nine grains: Ant j collects five grains:
The kindest individual will then be the ant
In our GA, goodness is determined by the total number of cases in which the participation of the individual (query vector or chromosome) in combination with the rest of the individuals (chromosomes) of that same generation did NOT obtain new results, its optimal value being 0. That is, if the number of times in which a given vector does NOT achieve new results is 0, it means that on all the occasions that vector was tested, there WOULD have been retrieval of new potentially relevant documents not previously viewed by the user.
The procedure was as follows:
(a). In order to identify which descriptors (and by extension, which chromosomes) are preventing new results from being obtained, the combination used is of m elements taken n by n without repetition (where m is the number of chromosomes in the population and n the size of the different groups of chromosomes that can be formed, which in this case is always less than or equal to m1). For example, there are four vectors (m = 4) which represent the results of the operations of selection (four relevant documents that were identified by the user in the previous search process), crossover and mutation. To identify which descriptors of the contents in these four vectors are those preventing new results from being obtained, the combinations (sum) of four elements (vectors) are tested, taking them 1 by 1 (n = 1), 2 by 2 (n = 2) and 3 by 3 (n = 3). 1 The various combinations are tested through the BWSA, and the combinations that do NOT retrieve new results are registered. The goodness of a vector will be given by its non-occurrence in these combinations, that is, its optimal value will be 0. 2
(b). The chromosomes are ranked from lowest to highest goodness. Goodness is determined by the formula (a).
(c). The results obtained by the combinations are shown to the user in the following order: first, the documents retrieved by the combination in which the lowest ranked chromosome did not participate (if they exist), then the documents of the combination in which the two lowest ranked were missing, then the three lowest and so on.
In this way, the retrieval of new results is achieved with a maximum of precision (the minimum possible relative to the degree of precision for which the last search made with RF did not achieve any new results), reactivating the process by which the user can keep exploring new results and tag those which are relevant. In our GA process therefore, precision is prioritised over comprehensiveness in terms of document exploration.
4. Results
The main result of this research is the development of a prototype, called Delphos, of a CI-based DSS, improved with GA search optimisation incorporated in the IR module, and designed in such a way that the IRS is capable of better defining the real need for information and providing results in line with that need.
As commented on above, Delphos works with various types of documents: patents, tenders, academic documents and web-pages (Figure 1). The main functionalities of the DSS prototype are:
Information search and retrieval. Instant searches. Automatic information searches and generation of alerts.
Information analysis. Instant search for and display of trends. Programmed search for and display of trends. Trend analysis.

Delphos search panel.
4.1. Information search and retrieval
In addition to executing instant searches, the system allows automatic searches to be set up. At the top of the main Delphos panel (Figure 1) is the ‘Emerging technologies’ section giving access to the option of editing automatic searches and consulting the alerts that are generated. The matches that the system finds in the automatic search process are added to the alerts panel for later review by the user.
Figure 2 shows an example of a Web document search, in particular, a search on the Internet for documents related to wind energy projects and specifically of competing firms.

Example of web document search (‘wind energy’ and ‘projects’) using the option ‘Competing firms’ of the filter ‘Organisation type’.
Once this initial search has been carried out, the IRS returns a set of results, and the user will be able to start the RF process to refine their search, tagging those documents they consider relevant to their informational need (Figure 3).

Results obtained for the example of the previous web document search (‘wind energy’ and ‘projects’) using the option ‘Competing firms’ of the filter ‘Organisation type’.
When no new results can be found, the GA designed in accordance with the method described in the previous section is automatically executed, which helps to overcome the obstacles encountered by the RF technique in the search for new potentially relevant documents.
4.2. Information analysis
The analysis of the information in Delphos is carried out on tenders, patents and academic document types, that is, on information of a structured type. The options offered by this functionality are the search for and display of trends and their analysis.
Delphos allows both instant searches and display of trends, as well as programming the system to generate an alert if there is a change in the constant of a trend. Editing event alerts allows the user to keep abreast of significant variations in the timelines drawn by the trends, without the need for continuously consulting them manually.
Figure 4 shows an example of a patent search (‘liquefied natural gas’) when applying ‘Free Text’ like filter to visualise ROK patents that contain the term ‘liquefied natural gas’ relative to the total of ROK patents published from November 2018 to November 2021.

Example of a search in the Delphos ‘Trends’ panel.
Figure 5 shows the trend plot from the previous search.

ROK patents that contain the term ‘liquefied natural gas’ with respect to the total of ROK patents for the period 2 November 2018–8 November 2021.
With respect to trend analysis, whether the graph being viewed comes from an instant search for information about the trend or is the result of an alert generated by the system, the user has the option to analyse the documents in greater depth, concentrating on a specific period in the graph. This might be one that catches their attention as corresponding to a turning point, or a sudden rise or fall, for example.
Continuing with the search example ‘liquefied natural gas’ on ROK, exploring the plots the ROK patents that contain the term ‘liquefied natural gas’ relative to the total of world patents published from November 2018 to November 2021 that contain the term ‘liquefied natural gas’, one observes an increase from November 2018 until August 2019 and a significant decrease again until February 2020. Although a new rebound is observed in October 2020, the average values do not reach the percentages prior to the aforementioned decrease. Using the trend analysis functionality that is incorporated in Delphos, the user can, for example, select the period ranging from August 2019 to May 2020 to analyse the patents published in this period, and the periods corresponding to the same time interval (9 months) before and after to inquire the possible reasons for such an increase and decrease (Figure 6).

ROK patents that contain the term ‘liquefied natural gas’ with respect to the world total of patents that contain the same terms for the period 2 November 2018–8 November 2021.
At the bottom of the screen is the ‘Analyse’ button and two boxes for entering the start and end dates of the temporal section of the graph to analyse (Figure 7).

Table of the trend analysis results.
5. Assessment
The vast majority of publications in this field only use one case study or illustrative example to validate the performance of their model. This may mainly be due to the lack of sufficient datasets describing real decision problem cases (used for comparison and validation) and the lack of well-established evaluation metrics to objectively assess the performance of such models [50].
Thus, as mentioned in the ‘Methods’ section, the prototype was assessed from a cognitive rather than a traditional point of view, with the fundamental objective being to determine the degree of user satisfaction after having used the system’s functionalities.
In addition to the continuous intermediate assessments carried out between the researchers and the firm’s users following the Scrum method, a final assessment was made which included users of an outside firm in the same sector. After using Delphos, these new users responded to a questionnaire concerning their real information needs in their usual intelligence and surveillance tasks [94,95]. The responses to the questionnaire, based on the 68 real information needs that were finally distinguished for both our tool and for Google (their usual search engine) yielded the following results (Tables 1 and 2).
Comparison of results for initial queries.
The maximum result review limit is 100.
Comparison of results for improved queries.
For the improved queries, the relevant ones not previously displayed are counted.
As can be seen in Table 1 for the initial searches, when the queries were formulated in our tool without applying filters, Google performed better, with 2.94% more of the cases in which the searches returned relevant documents among the first 20 retrieved and 2.94% fewer searches that did not return a relevant document among the first 100 retrieved. When the queries in the Delphos system were formulated applying sector, location and/or type of organisation filters, the searches that returned relevant documents among the first 20 increased by 11.76% compared with the same searches without filters.
Nonetheless, the most significant difference was that which came from applying the search enhancement functionality which included, among other techniques, the search optimisation GA (Table 2). After a single interaction between user and system, the percentage of searches that managed to retrieve relevant documents among the first 20 results rose to 78.94%, representing an improvement of 21.59% compared with the initial search without filters.
The improvement effect translates into shifting the relevant documents to higher positions in the results ranking. For this reason, the percentage of searches whose relevant results were found after the first 20 decreased, with a cumulative 17.54% compared with the 24.99% and 32.35% of initial searches in Delphos with and without filters, respectively.
When the user searches manually again for information with the intention of improving (or expanding) the IR provided by the initial search, they manage to obtain new relevant information (other than what was found in the first search) among the first 20 results returned by the Web search engine in 55.17% of the cases, and the accumulated number of cases of finding relevant information after the first 20 results corresponds to 20.68%. The real difference lies in those cases in which the first attempt to manually improve the search failed to obtain relevant results − 24.13% versus a scant 3.5% failure rate obtained after the first search improvement interaction in Delphos.
Another datum that can be obtained from the information provided by the user is that of the number of attempts to improve the searches. In our Delphos application, the user makes a second improved query in 43.85% of the searches and a third or more in 5.26% of the cases. This contrasts with the Web search engine use in which the users manually reformulate their search a second time in 62.06% of the cases, and a third time or more in 41.37%. These data seem to indicate that, with Delphos, a second interaction to improve the results is usually quite effective, much more so than a second manual reformulation, in which a new construction is made in a large proportion of cases.
Finally, the user was asked about their general feeling after looking back over their use of all the search and information analysis functionalities of the tool. A representation of their responses is shown in Figure 8. The responses show a high level of general satisfaction with the prototype developed, highlighting among all its functions its usefulness in decision-making, which is the main objective of a DSS.

General user satisfaction.
6. Conclusion and discussion
In a scenario such as the current one affecting the world, globalisation has led to the internationalisation of the markets, and the new technologies are accelerating the rate at which technological advances occur, being increasingly explosive and with ever greater penetration in society. This means that they stay on the crest of the wave of innovation for ever shorter periods of time before being overtaken by others in this changing and fiercely competitive environment. Intelligence and surveillance tools, and the rest of the disciplines that aim to manage information with the common purpose of extracting knowledge from the immense amount of content generated daily worldwide, undoubtedly have a promising future, but the researchers who work in this field definitely face an immense challenge.
In this work, we have presented Delphos, a DSS based on TW/CI, that uses a GA and contributes to the improvement of IR through query optimisation using the RF technique.
One of the main contributions of this study in comparison with other studies is the innovative combination of tools and techniques from different disciplines such as Business Intelligence, IR and AI, which are put together in the development of the prototype to achieve its optimisation. Another contribution of this research is the method proposed as a framework for the implementation of a DSS and which can help other researchers in the development of tools of this type. In addition, the proposed framework can be used not only by companies in the energy sector (field for which it was developed) but can also be extrapolated to other sectors of activity.
The Delphos prototype is just a first development, and naturally can be consolidated and improved at the request of the firm to which it is directed. Nevertheless, a conclusion that can be drawn from the responses given by its users is that the designed DSS represents a useful tool with which to obtain information from the environment and then analyse it to generate knowledge that can support the firm’s decision-making, since its use has been very satisfactory.
In sum, after the design and development of our TW/CI DSS, and taking into account the considerations offered by the users who tested the system, dealing with real information needs in their usual intelligence and TW tasks, we can affirm that the implementation of an appropriate intelligence and surveillance tool, as simple as it may be, helps to manage effectively information about the external environment in which an organisation operates. Also, if the information the system works with is of quality, the analysis of that information will lead to useful knowledge for better decision-making, and this in turn will ultimately translate into the organisation’s better competitive positioning.
There was major improvement in the performance of the tool when applying the GA implemented in it. This could be seen in the differences between the results provided by the GA compared with the usual information search techniques of the users who put the system to the test. Using this Evolutionary Computation technique contributed to a substantial improvement in the retrieval of information that can serve as the basis for obtaining knowledge about a firm’s external environment.
One of the advantages of the prototype developed is that it is based on OSINT sources, free access information sources, in order to respond to a specific intelligence request such as decision-making. Another important advantage is its versatility. The prototype developed is focused on firms in the renewable energy sector. Implementing a DSS in an organisation dedicated to another sector should not make a difference in the tool’s usability and usefulness as long as one of the main steps for its correct implementation is met: identifying and adding quality sources (hosts) to the internal database through the editor generated for that purpose and establishing a classification adjusted to the needs of each firm and with which these are appropriately represented. The crawler and parser modules will then take over, continuing to feed the database from these initial sources.
There are some limitations in the study. In the first place, although the group of sources used is a fairly considerable and representative sample of the different types suitable for a TW/CI tool such as Delphos, it would be beneficial to increase their number, since the more information a DSS handles, the greater quantity and higher quality will be the knowledge obtained to support decision-making.
Another possible limitation of our study is that the tool has yet to have a collaborative approach to its use. This is because, before continuing to generate accesses to new databases, the system must provide a base on which to support the various KITs identified in preceding intelligence efforts. Thus, different users connected in a collaborative mode (or only one if the size and resources of the organisation thus dictates) can channel information to each of these KITs as well as knowing the state it is in individually. For this, the application must also allow a system to share content among users, to pose challenges and to know the current state of resolving those challenges collectively, distributing the tasks that need to be completed to address each of the problems raised in defining the strategy. The future of Delphos lies in it continuing to be built from a fundamental perspective to support the intelligence and surveillance strategy of any organisation and collaborative work.
Finally, it should be noted that the UNE 166002 and UNE 166006 standards argue in defence of these tools as aids for organisations’ effective R&D + i management through which they can achieve strategic positioning with respect to their competitors. Therefore, we can affirm that the use of TW/CI tools in an organisation facilitates its management of external information and contributes to better knowledge of its external environment, and this in turn affects the management of its R&D + i as a strategic resource with respect to its competitors. These tools, understood as being part of the framework that configures the implementation of a solid knowledge-based strategy, will underpin the efficient progress of the organisation of the future’s R&D + i.
As future lines of research, the next step needed to improve the Delphos prototype is, as already mentioned, to expand the information with which the system operates so as to get a broader, and at the same time more reliable, vision of the environment and hence to achieve better knowledge. This will begin by adding access to new resources from quality structured sources for the types already implemented and then continue by diversifying those types. Another future line of research is to broaden the focus for collaborative use. Also, there has to be further research into adaptation functions for the GA responsible for search optimisation using the RF technique, since these functions constitute a key element of the system.
Footnotes
Acknowledgements
This work was carried out within the framework of the collaboration agreement between the University of Extremadura and the Spanish company Acciona S.A. for the execution of the Delphos project for the development of tools and methods for the implementation of a Competitive and Prospective Intelligence system in post-normal environments.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship and/or publication of this article: The study was made possible by funding from the Regional R&Dþi Plan of the Government of Extremadura (Spain).
