Abstract
This paper introduces the concept of a Smart Metrology Campus, with its primary objective being to facilitate reliable data science in the realm of smart cities. This is achieved by connecting sensor and meter data with metadata, particularly from the field of metrology. The establishment of a robust data infrastructure, responsible for collecting this data, requires a strategic data engineering approach. The challenge lies in identifying a data engineering approach that is well-suited for design considerations characterized by high demands in scale, complexity, sensitivity, and reliability. To address this challenge, the paper conducts a comprehensive analysis of the applicability of existing guides for data engineering, alongside other guiding processes from related disciplines such as data mining, data science, and systems engineering. As a result, systems engineering is identified as a highly relevant approach for data engineering meeting the needs of smart cities. The subsequent sections explain how this approach can be utilized and adapted for the specific requirements of a data engineering process in smart cities.
Introduction
In a smart city, sensor networks play a crucial role in securing the sustainable provision of water, heat, and energy. Additionally, they contribute by offering essential data regarding light and noise pollution, air quality, and climate parameters. The effectiveness of these measurement networks hinges on the accessibility of the data and information quality. This involves addressing aspects such as uncertainty assessment, measurement frequencies, and calibration to ensure accurate and reliable information (Jung, 2023). Data management systems (DMSs) responsible for gathering and storing sensor data have been extensively studied in the context of smart city-related topics, specifically in areas such as the smart grid (Daki et al., 2017) and the internet of things (IoT) (Aiello, 2022; Fremantle, 2015; Weyrich and Ebert, 2016). However, there has been a notable gap in research regarding the integration of metadata necessary to ensure high trust and reliability in these DMSs.

The context of the Smart Metrology Campus data infrastructure which includes the domains of Smart City, Metrology and Data Science.
It is crucial for data scientists involved in smart city projects, as highlighted in Technikum Wien Academy (2023), to have access to metadata for assessing the unusual behavior of sensor data. This capability is essential to conduct trustworthy analyses and allow an effective and reliable application of methods supporting data science such as machine learning models.
The paper “Security and privacy concerns in assisted living environments” (Condado and Lobo, 2023) which is part of this journal, has shown that it is very difficult for IoT technologies to earn trust. Given the complexity and sensitivity inherent in the smart city context, it is important that the DMS gains the trust of the involved parties. For this, the structured data engineering process should not only account for the needs of smart city stakeholders and decision support but also address fundamental concepts such as the selection of data warehousing or data lake solutions, as well as the incorporation of relevant technologies such as Cloud Computing and Big Data. Moreover, it must ensure the long-term stability and maintainability of the system as well as the transferability into a subsequent system. It is important that the design and related decisions, which lead to the final DMS and data infrastructure, get comprehensively documented by a guiding process to gain the trust of all involved parties.
This paper attempts to complement the prior work (Condado and Lobo, 2023) by identifying and articulating such a guiding process. The initial segment will examine the current landscape of data engineering guides. Following this, the exploration will extend to guides from related disciplines, including data mining, data science, and systems engineering. A noteworthy aspect of this paper is its examination and subsequent discussion about which process guide is most suitable for a data engineering approach capable of meeting the challenges posed by the smart city context. The second significant contribution of this research will also be a detailed examination of how this chosen process can be effectively applied to the data engineering requirements. Finally, the findings will be summarized in a conclusion.
The work is part of the “Smart Metrology Campus (SMC)” initiative at Physikalisch-Technische Bundesanstalt, which is the National Metrology Institute of Germany (PTB-Homepage, 2024). The aim is to develop parts of a smart city infrastructure that help institutional goals such as energy saving, but also give a platform for research activities toward metrology in smart cities. The overarching goals of the SMC include promoting sustainability, facilitating data communication to the inhabitants, fostering interconnection, and encouraging participation. The key to these goals is a data infrastructure that collects and documents the data to enable trustworthy data science. The principals get developed by the electricity meter in corporation with the Elenia Institute of the Technical University of Braunschweig (Kurrat and Engel, 2024). With this, a link across the three domains, metrology, smart city, and data science, as shown in Figure 1 will be established.
The data infrastructure of the Smart Metrology Campus fuses the domains of Smart City, Metrology, and Data Science. Each of these domains imposes specific requirements on the development of the data infrastructure. However, it is the domains of Data Engineering and Smart City that establish the criteria for the structured data engineering process in the context of a smart city.
The challenges inherent in smart cities stem from the scale, complexity, and sensitivity of the data, primarily driven by the sheer volume and diversity of both data and stakeholders involved. Therefore, it becomes crucial for the process to thoroughly analyze the contextual nuances of the data and use cases, incorporating a method to address the diverse requirements posed by different stakeholders. Additionally, given that smart city systems often involve data spanning a significant period, it is imperative for the process to offer guidance for ongoing solutions in use.
The main task for data engineering is to provide a data infrastructure so that the data are ready for further analysis such as data science at the final state (Crickard, 2020). The data engineering life cycle involves key steps, including generation, storage, ingestion, transformation, and serving. At the core of this life cycle are concepts such as data warehouse, data lake, etc. (Reis and Housley, 2022). These concepts are collectively referred to here as DMSs. The implementation of DMS can take various forms using different technologies. Therefore, it is imperative for the process to offer decision support on the choice of DMS and how the process should be implemented, including the selection of technology. In summary, the requirements that need to be considered for the data engineering process include:
Smart city requirements
Context understanding Requirements engineering Long-term planning Data engineering requirements
Decision support DMS Decision support technology
State-of-the-art
This state-of-the-art exploration aims to summarize the unique contributions of existing data engineering approaches, data mining, data science, and systems engineering. By examining their strengths and intersections, the aim is to identify synergies that can enhance and optimize the data engineering processes crucial for the advancement of smart city initiatives.
Data engineering
Relevant literature with a focus on data engineering will be reviewed in the following section. The guide “Data Engineering with Alteryx” by Houghton (2022) focuses on data operation (DataOps). It covers the different life cycles of a data engineering process but this does not cover any requirements analysis of the stakeholder or decision support for what concept or technology to use but presents the platform Alteryx as the only solution.
The book “Fundamentals of data engineering: plan and build robust data systems” by Reis and Housley (2022) introduces the data engineering life cycle with the main steps of generation, storage, and analytics. The book introduces data engineering and its different components as well as the common concepts like a data warehouse, a data lake, etc., which are grouped here under the term data architecture. The book continues by explaining every step of the life cycle and finishes with security and privacy concerns.
The guide “Official google cloud certified professional data engineer study guide” by Sullivan (2020) gives advice about how to select an appropriate, which is described as storage technology. This is followed by the process of building and designing data pipelines. It then also considers data processing steps, operation, and characteristics of the solution. The guide finishes then how a system like this can be used for machine learning.
Data mining
In the realm of data mining, there are established methods that have become quasi-standards. Given the similar context and challenges shared by data mining and data engineering, this method might be also applicable for data engineering as well. The primary ones to be examined here include Knowledge Discovery in Databases (KDD), Cross-Industry Standard Process for Data Mining (CRISP-DM), and Sampling, Exploring, Modifying, Modeling, and Assessing (SEMMA).
An overview of the steps of the different processes can be seen in Table 1. The initial process among them is KDD, introduced by Fayyad in 1989 (Shafique and Qaiser, 2014). It serves as a comprehensive approach to extracting information from database technologies (Fayyad et al., 1996). Following this, CRISP-DM emerged in 1996, founded by three “veterans” in the nascent and evolving data mining market (Chapman, 2000). SEMMA was defined by the SAS Institute. This method is designed for handling substantial amounts of data to unveil previously unknown patterns that can be leveraged as a business advantage (SAS Help Center, 2023; Shafique and Qaiser, 2014).
Comparison of different data mining guides by their steps.
Comparison of different data mining guides by their steps.
KDD: Knowledge Discovery in Databases; CRISP-DM: Cross-Industry Standard Process for Data Mining; SEMMA: Sampling, Exploring, Modifying, Modeling, and Assessing.
All three processes share quite comparable steps, with the exception that CRISP-DM is extending these with two additional steps. The first extra step is business understanding which only CRISP-DM defines. This step focuses on understanding the project objectives and requirements from a business perspective, then converting this knowledge into a data mining problem definition and a preliminary plan designed to achieve the objectives (Chapman, 2000). Next, all three processes initiate the data selection phase, where data are chosen from the entirety of available datasets. In the KDD process, the defined product is termed as the target data (Fayyad et al., 1996). This selected data then undergoes preparation, a step grouped here under the term “data preparation.” Data preparation covers all activities to construct the final datasets from the initial raw data (Chapman, 2000). The KDD process defines the product as clean date (Fayyad et al., 1996). Subsequently, this clean data undergoes transformation in the next step. The objective of this phase is to format the data in a way that is advantageous for analysis. Typical transformation steps include converting nominal values to numeric values or grouping metric values into intervals (Fayyad et al., 1996). CRISP-DM does not have this modification step but it is also part of the data preparation step. The data, after being prepared and transformed, is then ready for the actual data mining step. In KDD, this is described as the search for patterns of interest. In the CRISP-DM process and SEMMA, this step is open to applying any modeling technique, which can also include neural networks, for example. Subsequently, these patterns and models undergo evaluation and interpretation in the final step. It is crucial in this phase that the knowledge gained is not only new but also useful (Chapman, 2000; Fayyad et al., 1996; SAS Help Center, 2023). The CRISP-DM process introduces an additional step to this process, known as the deployment step. Depending on the requirements, the deployment phase can be as simple as generating a report or as complex as implementing a repeatable data mining process across the enterprise (Chapman, 2000). If the acquired knowledge is not deemed new or useful, the steps can be iterated until the desired outcome is achieved (Chapman, 2000; Fayyad et al., 1996; SAS Help Center, 2023).
The National Institute of Standards and Technology defines data science as the methodology for the synthesis of useful knowledge directly from data. This definition also refers to the management and execution of the end-to-end data processes, including the behavior of the components of the data system (Chang et al., 2019).
The end-to-end data processes of data science also encompass data engineering, raising the prospect that common methods in data science can be applied to data engineering as well. Notable data science processes to be examined include the Microsoft Team Data Science Process (TDSP), The Six Phases of Data Science, Obtain, Scrub, Explore, Model, iNterpret (OSEMN), and Agile Data Science (ADS).
The TDSP, Six Phases of Data Science, and OSEMN processes exhibit comparable phases, as illustrated in Table 2.
Comparison of different data science guides by their steps.
Comparison of different data science guides by their steps.
TDSP: Team Data Science Process; OSMEN: Obtain, Scrub, Explore, Model, Interpret.

Steps of the technical process of systems engineering described in the ISO/IEC/IEE 12588.
The steps, understanding, data preparation, modeling, data evaluation, and deployment, are identical to the corresponding steps in data mining. The step that is switched is data selection, which is replaced by data collection. In this step, data are automatically moved from the source to the target location (Tabladillo et al., 2023; Mason and Wiggins, 2023).
ADS is a process designed to document, facilitate, and guide exploratory data analysis with the aim of discovering and progressing along the critical path toward creating a compelling analytics product (Jurney, 2017). According to Jurney (2017), ADS does not adhere to strict guidelines but is guided by principles, which include the following steps:
Iterate, iterate, iterate: tables, charts, reports, predictions Ship intermediate output. Even failed experiments have output. Prototype experiments over implementing tasks. Integrate the tyrannical opinion of data in product management. Climb up and down the data-value pyramid as we work. Discover and pursue the critical path to a killer product. Get meta. Describe the process, not just the end state.
ADS does not prescribe a specific method, but it references the OSEMN framework for use in conjunction with ADS (Trapani, 2023).
System engineering has the aim to be a guide to facilitate the development of complex systems (Kossiakoff et al., 2011). The approach tries to enable realization, use, and retirement of engineered systems using systems principles and concepts and scientific, technological, and management methods (15288-2023-ISO/IEC/IEEE, 2023). Both data and system engineering share the aim of developing a system that will possibly be used in production. In both cases, the methods are handling usually very complex use cases. This is why a combination of both looks promising to handle the challenges of the smart city context.
The guide that will be examined for systems engineering is the ISO/IEC/IEE 15288 systems and software engineering—system life cycle processes. According to 15288-2023-ISO/IEC/IEEE (2023), the scope of the standard is to establish a framework of process descriptions for describing the life cycle of systems created by humans, defining a set of processes and associated terminology from an engineering viewpoint.
The standard differentiates between different life cycle processes. The agreement processes, the organizational project-enabling processes, the technical management process, and the technical process. The process that is relevant for designing and realizing a system is the technical process. The technical processes describe the technical actions throughout the life cycle. In the standard, there are 14 technical processes that are split into four different phases which can also be seen in Figure 2. Each phase will now be explained in detail:
Concept definition
The concept definition phase starts with the business or mission analysis process. In this process, the overall strategic problem or opportunity gets defined as well as the solution space characterized. Parallel to this process is the stakeholder needs and requirements definition process which defines the stakeholder needs and requirements for a system that can provide the capabilities needed by users and other stakeholders in a defined environment. As a result, the operational context of the solutions should be defined and the operational concept should address what the system will do and why.
System definition
With the concept defined, the next step is the system definition. The system definition starts with the system requirements definition process in which the stakeholder/user-oriented view of desired capabilities, which have been found before, gets transformed into a technical view of a solution that meets the operational needs of the user. After this follows the system architecture definition process which generates system architecture alternatives to this technical view, selects one or more alternative(s) that address stakeholder concerns and system requirements, and expresses this in consistent views and models. At the same time, there is the system analysis process which provides a rigorous basis of data and information for technical understanding. This is needed to aid decision-making and technical assessments. When an appropriate system has been found, it is the role of the design definition process to provide sufficient detailed data and information about the system and its elements. Those data and information can then be used to realize the solution in accordance with the system requirements and architecture.
System realization
The system realization starts with the implementation process. At this step, the results from the system definition process get transformed into actions that create system elements according to the practices of the selected implementation technology. This system element then gets integrated into a realized system in the system integration process. The task is to ensure that the system, the system elements, or artifacts fulfil their specified requirements and characteristics task of the following verification process. This follows the validation process which extends the tasks to provide evidence that the system, when in use, fulfills its business or mission objectives and stakeholder needs and requirements.
System deployment and use
The realized system enters now the system deployment and uses phase. Two aspects that need to be considered when the system is in use are the maintenance processes and the operation processes.
The maintenance process sustains the capability of the system to provide a product or service which then gets used in the operation process. The maintenance process monitors the system’s capability to deliver products or services, records incidents for analysis, takes corrective, preventive, adaptive, additive, and perfectible actions, and confirms restored capability. The operation process uses the system to provide its products or services. This process establishes requirements for and assigns personnel to operate the system, and monitors the products or services and operator-system performance.
The other two aspects of the process describe methods with the system at the end of life. The two options are the transition process, which moves the system in an orderly, planned manner to be operable in the intended environment, which may be a new or changed environment. The other process is the disposal process which deactivates, disassembles, and removes the system or any of its system elements from the specific use.
The standard 15288-2023-ISO/IEC/IEEE (2023) is also mentioning that often these processes are performed at the same time and they iterate between each other. By doing this it can consider new requirements and findings.
Discussion
Now, as the different guides and processes have been introduced, a discussion will ensue to determine their appropriateness for the smart city use case. The requirements, applied to each domain and method, have been defined in Section 2. To facilitate clear decision-making, each method will be scored, with a 0 indicating that the method does not fulfill the requirement, 0.5 indicating partial fulfillment, and 1 indicating that the method fully satisfies the requirement. The result of the analysis can be seen in Table 3. The justification of each scoring will now follow:
Result of the analysis: Determine whether the methods fulfill the requirements from the domains of smart city and data engineering.
Result of the analysis: Determine whether the methods fulfill the requirements from the domains of smart city and data engineering.
DMS: data management system; KDD: Knowledge Discovery in Databases; CRISP-DM: Cross-Industry Standard Process for Data Mining; SEMMA: Sampling, Exploring, Modifying, Modeling, and Assessing; TDSP: Team Data Science Process; OSEMN: Obtain, Scrub, Explore, Model, Interpret.
The literature provided for data engineering lacks specific steps from the smart city domain, such as context understanding, requirements engineering, or long-term planning. In the case of data engineering with Alteryx, the introduction of Alteryx as a platform implies also a lack of suggestions for decision support regarding DMSs or technology. The literature titled “Fundamentals of Data Engineering,” does offer guidance on finding a DMS, referred to as data architectures, but the decision criteria are based on available technology rather than the requirements stemming from the use case. Additionally, there is no decision support provided on how to choose an implementation. The Data Engineer Study Guide from Google provides some guidance on selecting a technical solution, considering the requirements of the use case. However, the technical solution does not incorporate decision support for DMSs.
All the introduced data mining methods assume the availability of data. Consequently, there is no guidance provided for certain decision-support aspects of the data engineering requirements. Furthermore, only the CRISP-DM method supports the first step of the smart city requirements with its business understanding step.
The described data science processes include data collection as an explicit step. However, there is a lack of further definition guidance on how this data should be collected, the conceptual framework, or the preferred format. Consequently, none of the data engineering requirements are fulfilled in this aspect. Both Microsoft TDSP and the six phases of data science include business understanding, thus fulfilling one of the smart city requirements. This is not the case with the OSEMN methods, as well as ADS, which utilizes the OSEMN framework.
The technical process of systems engineering, as outlined in ISO/IEC/IEEE 12588, emerges as the most suitable framework for data engineering within the domain of smart cities. This conclusion is drawn from its comprehensive approach that fulfills all the requirements necessary for implementing a solution tailored to smart city challenges. Firstly, the systems engineering process addresses the essential steps of context and requirement analysis in its initial phase, concept definition. Through business or mission analysis, as well as stakeholder needs and requirements definition, it thoroughly assesses the smart city environment, ensuring a robust understanding of the context.
The engagement of the stakeholders is especially the strength of the standard. It starts with the holistic identification of stakeholders over the life phases of the system. With this approach, not only operators who are actively involved in the system development are in the focus of the system design but also operators from the middle and end of life. Then the stakeholders are actively involved during both the concept definition and system definition phases. The process documents during each step the path that leads to the decisions, ensuring that all decisions are grounded in stakeholder requirements. With this dual approach, the process provides a broad stakeholder inclusion and maximizes transparency. Such measures are particularly crucial in the context of smart city projects to foster widespread acceptance and trust in the system. An additional advantage is the systematic integration of new or modified requirements, allowing them to be evaluated and aligned with existing requirements.
Moreover, the process incorporates advice for long-term planning in the system deployment and use phase, a crucial aspect for sustainable development within the dynamic landscape of smart cities. While the process lacks specific guidance on selecting a DMS or technology, its concept definition phase inherently aims to identify an operational concept, here to a DMS. Similarly, the system definition phase aims to specify system elements, providing a foundation for choosing the appropriate technology for the DMS.
In summary, the technical process of systems engineering stands out as the most suitable framework for data engineering in the smart city domain due to its holistic approach, addressing context understanding, requirement analysis, and long-term planning, while also laying the groundwork for DMS and technology selection. In its phases and sub-phases, it provides enough metadata such as requirement analysis to make a design choice traceable and understandable.
This section will explain how the technical process from systems engineering can be used for a structured data engineering process for the smart city use case. The structure of the process is the same as the original and can be seen in Figure 3. Also as in the original, the execution of the process is meant to be not linear but iterative. However, the aims and especially the outcomes are more defined for data engineering in the context of a smart city. The aim is to realize a long-lasting system that collects, saves, and possibly distributes data. Systems which are here described under the group term DMSs.

Steps of the technical process of systems engineering described in the ISO/IEC/IEE 12588 with extensions for data engineering.
The first phase of concept definition has the aim to find the most fitting concept for the use case. The 1.1 Business and/or Mission Analysis and the 1.2 Stakeholder Needs and Requirements Definition process should analyze what kind of tasks and use cases the system should fulfill to define the system constraints. The processes are very much mutually dependent on each other. For data engineering, it is very important to include next to the business and mission analysis and also a data analysis part. Further at the end of the process 1.1 is an evaluation of different solution alternatives. In this case, it would be very beneficial to use a criteria catalog such as it is presented in former publications (Ulbig et al., 2023). The outcome of this process should be the decision of a concept for a DMS such as a data warehouse or a data lake. The aim of the system definition phase is to decide and plan how and with which technologies this concept can be realized.
The first step in this phase is the 2.1 Systems Requirements Definition process which sets and describes further requirements for the architecture and technologies. In addition, the 2.3 System Architecture Definition process looks at different options which technologies, and how they can be combined. Those technologies then get evaluated in the 2.2 System analysis process using the requirements that have been defined in process 2.1 and provide feedback back to the 2.3 System Architecture Definition process. The feedback is then used to define and finalize the architecture. This architecture will then be more specified in the 2.4 Design Definition process. The outcome of this process and also this phase should be the traceable decision for certain technology solutions and the resulting concepts and descriptions.
These descriptions can be used in the 3.1 Implementation and 3.2 Integration process. In both cases, this would include the necessary hardware and software solutions as well as their interaction with external systems. The results of both processes should then be verified and validated in the processes 3.3 Verification and 3.4 Validation. The end result of the system realization phase should be the realized system which is then ready to be used.
The last phase is then the System Deployment and Use phase. The 4.A Maintenance and 4.B process are handling the use phase of the DMS. This in the context of data engineering is the maintenance and interaction with the software. The 4.2A Transition and 4.2B Disposal process are both related to the end-of-life stage of the system. In this case, it is important to develop not only the related strategies about the potential hardware and software but about the data as well. All of these phases and their process will now be explained in the following in greater detail. For this part of the processes will be summarized and then be set in the context of data engineering in smart cities.
In addition, a minimal real-life example was added as well. A real-life example is the creation of a system that collects energy data of a city district for research purposes and data-driven automation in the district operation. The example does not list all steps that must be included but tries to give a view ideas how the norm could interpreted for a project.
The main aim of the concept definition is to find and select a concept for the system. This concept should be reasoned and described by the use case context and possibilities as well as the requirements of the stakeholders. The process 1.1 Business or mission analysis and 1.2 Stakeholder need and requirements definition are not carried out one after the other but are interlinked as can be seen in Figure 4.

The processes of the conception definition phase and their interactions including the outcome of each process.
The phase begins with the definition of the problem or opportunity space. As the standard describes, this include an analysis of the background to understand the problem scope, basis, or drivers as the first step. The analysis should then result in a definition of the problem or opportunity including, key parameters and critical business success measures. The problem will then be prioritized against other business needs.
In the context of data engineering for smart cities, this would mean first analyzing the problem in the city which should be solved by the system which collects and saves data. This would include the analysis of city needs or opportunity as well as specifically a rough look at what data the system needs to work with. Very important factors might be in this regard the data input and output. The common context might include an analysis of the data producer and consumer as well as questions about necessary hardware and software. Key parameters would be of course safety and security as well as the consistency and quality of the data collection and the usage of the collected data. Therefore the security and protection of residents is the most important and it must be checked whether a restriction in this respect justifies the use of a data infrastructure.
The problem and opportunity analysis in the given real-life example would analyze the problem background, the data input, and the data output. The analysis would be in this example a discussion with the stakeholders. The discussion could conclude that there is a need and wish to enable research and smart city technologies in the district. So far the energy provider has installed smart meters in the district without the development of a system to save the data. The data produces are the smart meters and belong to the energy provider and the data consumers are researchers and the public.
Operational and concepts in life cycle stages
The next step is then to define preliminary operational concepts and other life cycle concepts as the first part to characterize the solution space. The standard describes, in this task, it is important to identify and define the stakeholders and roles here listed as customers, users, administrations, regulators, and system owners as well as their connected operational concept. The standard lists here preliminary concepts for acquisition deployment, operation, support, and retirement. A key for this is the operational concept which furthermore should include high-level operational modes or states, operational scenarios, potential use cases, or usage within a proposed business strategy. It is important to check these concepts and operations of vulnerabilities and security issues. The interpretation for data engineering of the operational process would include the input and output of data to and from the system.
The next step in the characterization of the solutions space would be the identification of alternative solutions. However, a prerequisite for this is the definition of system constraints and stakeholder requirements. For this reason, the next step would be the further development of the operational concept and other life cycle concepts which belong to 1.2. Stakeholder needs and requirements definition process.
The real-life example separates the system into the three life cycle stages: beginning, middle, and end of life. The beginning of life includes the installation of the system which requires hardware and software and ends when the data can be collected and displayed. This involves system integrators as well as software developers as stakeholders. Operations of the system integrators are the installation and integration of servers and operations of the software developers are the installation and configuration of a database as well as the development of middleware. The middle of life describes how the system is used until it needs to be shut down or transitioned. This includes researchers and the public as consumers but also administrators who maintain the system. The operation of the researchers is accessing the data from the system and the public should get access to the data via a dashboard. Possible end scenarios for the systems’ end of life might be the transition into a different system that includes more districts. A disposal scenario might be the end of the research because of a lack of funding or end of research. In those cases, the stakeholders might be stakeholders of different systems but also the administrators who then need to migrate the data to a new system or delete data, uninstall software, and dispose of the hardware.
The standard describes that it is important to define the context of using the steps of the preliminary concepts specifically the concept of operations, preliminary life cycle concepts, a set of scenarios (or use cases), preferred solution class(es), and identify all required capabilities. The context of use for a set of scenarios or use cases should be used to identify all requirements that belong to the defined operational and other life cycle concepts. This step can in addition be used to characterize the operational environment and the intended users as well as the identification of interactions between users, the system, and the factors affecting the interactions and the interaction of all interface boundaries across which the system interacts with external systems.
The next step after the operations and concepts step are identified is the communication with the stakeholders to define the stakeholders’ needs and wants.
Stakeholder needs
Stakeholder needs are the wants, desires, and expectations of the stakeholders which are not further defined (29148-2018-ISO/IEC/IEEE, 2018). The needs often include measures of effectiveness and identification of critical operational issues and bring in their domain knowledge and context understanding. The next step is then to prioritize and down-select the needs. It is good practice also to record and track the needs along the process.
A stakeholder need in the real example could come from the administrator might be the wish of the researcher to access not only the sensor data but the system must also provide documents about the sensor to judge the quality of the data. The administrators want to keep the maintenance to a minimum and want the transparency of the system to a maximum.

The processes and preparation which are involved to define the system constraints.
The next step is the transformation of the stakeholder needs into stakeholder requirements. Stakeholder requirements are well-formed in the way that they include measurable conditions and bounded constraints. A stakeholder requirement for the given example would be a further definition of what kind of value should be displayed in which units. The first step of the transformation of needs into requirements regarding a system are functions that relate to critical quality characteristics. Such are assurance, safety, security, environment, or health. Defining stakeholder requirements can be developed by discussing need-related life cycle concepts, scenarios, interactions, constraints, critical quality characteristics, or system of system considerations.
An analysis for a requirement regarding the energy dashboard might raise security concerns that need to be considered.
The set of defined stakeholder requirements then needs to be analyzed. Single requirements need to be analyzed about their necessity, implementation interdependence, unambiguity, completeness, singularity, achievability, verifiability, and conformity. For sets of requirements apply the characteristics complete, consistent, feasible (or affordable), and bounded. The requirements furthermore need to define critical performance measures and quality characteristics that enable the assessment of technical achievement. Critical performance measures in the case of an energy dashboard for example might be the latency of the data being displayed. Those defined stakeholder needs and requirements then need to be validated by the stakeholder to resolve issues and finally ensure that they have been either adequately captured or expressed.
In the real-life example, the need of the researchers could lead to further specifies such as the wish to download the data to a csv file as an application programming interface. For a given time frame the data also could reference sensor-related documents such as a digital calibration certificate. The administrators want additional software to automate processes such as the addition, alteration, and removal of sensors. Another requirement is the use of logging and logging monitoring.
System constrains
These stakeholder needs and requirements then provide the necessary information terms and conditions that have a direct influence on the system constraints as well as the strategies than be developed for the system realization and system deployment and use. These strategies will be explained in the following. Outcomes will be the further system constraints. An outcome of the verification might be in the case of an energy dashboard that the data should be checked against other existing systems. This would lead to the constraint that the system also needs an interface to further systems.
A constraint outcome regarding the maintenance might be that the system will be handled by administrators without or low programming skills which would lead to the constraint that certain process must be automated or well documented.
As can be seen in Figure 5 the constraints come from the definition of the concepts, requirements, and strategies collect together in the general system constraints. The system constraints then build the space for the solution space.
Solution space
In data engineering, there are a lot of different options to store data. The solutions space can span minimal solutions such as saving data in a shared file to highly complex DMSs that automatically collect and process data. A comparison of stakeholder requirements and system limitations versus the possibilities of concepts and technologies should narrow down the selection. In the case of data engineering in smart cities it is more likely that the use cases require a DMS that can handle the large scale and variety of data.
The solution space in the real-life example would be a DMS since the system needs to handle different structures of ingestion methods of data.
Alternative solution classes and selected solution
The standard defines a solution class, which can be a new system, an adaptation or modification of an existing system, or a link between systems. Multiple alternative solutions need to be compared against defined criteria. Previous work (Ulbig et al., 2023) has shown how a criteria catalog can be used to compare different DMSs (there referred to as data storage systems). The catalog defines relevant criteria. Then it can be matched if the use case requires such a criteria and if the solution fulfills it. For this, it is important that the domain requirements get transformed into technical requirements. As an example, the use case has the domain requirement to collect sensor data. The resulting technical requirements for the DMSs would be the need for a streaming data ingest as well as the ability to save structured data. A classical data lake can ingest streaming data but a data warehouse can not for example. The result should be an overall concept. This can be a single concept as well as a combination of concepts such as a data lake with a data warehouse combined.
The solution classes in the real-life example are common DMSs such as data warehouse, data lake, lambda architecture, and kappa architecture. With the help of the criteria catalog from the previous work (Ulbig et al., 2023), the different DMSs can be matched with the requirements of the use case. In this example, the need to save structured and unstructured data leads to the conclusion that a data lake is a fitting concept.
If a stakeholder would set new requirements, such as the establishment of live monitoring of the energy data, the new requirement would be added to the other stakeholder requirements. The requirement would be transformed into a system constraint such as real-life data analysis. This would also change the requirements in the criteria catalog. The effect would be that the criteria catalog would change the final solution to a lambda architecture.
System definition
The system definition phase is defining how the concept can be realized via different technologies. This includes the systems requirement definition process, the systems analysis process, the system architecture definition process, and the design definition process. The outcome in the context of data engineering in smart cities should be a kind of blueprint that describes how a DMS should be realized for its use case. The example continues here with the data lake as the chosen concept.
System requirements definition process
By the standard, the purpose of the system requirements definition process is to convert stakeholder and user expectations into a technical framework that addresses operational needs. This process establishes measurable system requirements detailing the characteristics, attributes, and functional and performance criteria that the system must meet to fulfill stakeholder needs. The requirements are designed to specify what the system should do without dictating specific implementation methods, within the constraints of the project. The requirements are set by the stakeholders as well as the constraints that need to be applied. Compared to the concept definition the system definition requirements should be more viewed from a technical perspective.
The system requirements definition process is divided into the steps of system requirement definition and system requirement analysis and their sub-steps. For data engineering in smart cities, the outcome should be an overview of the requirements and constraints based on the knowledge from the concept definition. This overview then can be used in the decision-making process in the form of evaluation criteria in the system analysis process and the system architecture process.
A requirement could be from the administrators. For example, the requirement that databases that can save both structured and unstructured data are preferred to a solution with two databases. Also, the administrator would like to use dashboard technology which requires less maintenance.
System architecture definition process
The standard sets as best practice that the process also should involve defining a solution based on interconnected principles, concepts, and properties that align with each other. This process should be set to transform various related architectures, organizational policies, life cycle concepts, stakeholder requirements, and system constraints into the core concepts and governing principles that shape the system and guide its evolution throughout its life cycle. The standard divides the performance of this process further into the three steps of conceptualization, evaluation, and elaboration of the system architecture. In the interpretation of data engineering in smart cities, the outcome of this process should be a description or visualization of the system architecture which should include a rough overview of the selected technologies and their interaction with each other and external systems.
In the context of data engineering, this process would be used to define the tools and technologies that should be used for the system realization. As the standard points out it is also important to compare different system architecture alternatives. A data lake could be realized with Hadoop Distributed File System to store documents and an influxDB to store time series data or with MinIo and PostgresSQL. The decision-making for a selected solution should be defined with the help of the system analysis process and the system requirements and design constraints as can be seen in Figure 6.

The processes of the system definition phase and their interactions including the outcome of each process.
The system analysis process, in this context, would be an evaluation of the system elements that are specified in the system architecture definition process with the criteria given by the system requirements definition process. A tool to use this could also be an adaptation of the criteria catalog presented in the concept definition (Ulbig et al., 2023). The system analysis process goes even further and encompasses a variety of analytical functions and levels of complexity, tailored to the criticality of the information required. System analysis is used for technical assessments, including evaluating operational concepts, resolving requirement conflicts, assessing alternative architectures, and analyzing performance and risks. It often involves mathematical analysis, modeling, simulation, and experimentation to evaluate technical performance, feasibility, affordability, and life cycle costs.
In the context of data engineering in smart cities it also can be used for all of this as well but the important outcome would be the system analysis result which supports the decision-making in the system architecture process. In the data lake example, the PostgresSQL would be compared to MinIO and InfluxDB via a newly defined criteria catalog. Since one of the requirements is to use a database that can handle both structured and unstructured data the decision would fall here for PostgresSQL.
Design definition process
The design definition process applied in this example would be a furthermore detailed description of the system elements described in the system architecture. This could be for example the definition of internet protocol addresses and ports but also entity-relationship models for different databases. As described by the standard, this process transforms architecture and requirements into a realizable system design, providing detailed descriptions and drawings that align with architectural models and views, and conform to system requirements. The standard also advises to split the process into the creation of the system design and an evaluation step. In data engineering, the outcome should be a system design with possibly detailed descriptions of the system and system elements.
In the data lake example, the design definition process would be used to define the database schema for the sensor data. Another part might be the creation of a mockup of the dashboard.
System realization
The aim of the system realization phase is to set up a solution designed by the previous phases so that the end result can be used for operation. In this context, a system realization would be , for example, the implementation of a DMS which is already integrated into the data infrastructure. As shown in Figure 3, the system realization is divided into the processes of implementation, integration, verification, and validation. As can be seen in Figure 7, the system realization phase starts with the implementation process. The implementation result then gets integrated and both, the implementation and integration results, then need to be verified and validated. The verification or validation could lead to the necessity of further improvements which creates a feedback loop that should end when all the verification and validation requirements are fulfilled. All four processes are split into the tasks: prepare, perform, and manage. The prepare tasks in all four result in constraints for the system which have been mentioned before and the enable of systems and services for each process. Part of each preparation is the development of a strategy which will be explained for each process as well as the steps of the actual performance of each process. The outcome of the manage tasks is the traceability of the results and will here not be further discussed.

The processes of the system realization phase and their interactions including the outcome of each process.
The standard describes the implementation process as a transformation of the requirements, architecture, and design into actions to create the system elements. In the case of data engineering, the implementation is the installation of hardware and software which has been specified in the system definition process. Since data engineering is entirely software-based the implementation strategy would consider decisions such as agile project management, deployment, and documentation. The implementation process itself parts the standard into three steps. The first step is the realization of the system elements. In this use case that could be the purchase, construction, and installation of hardware if necessary. The implementation of software would be the installation of the defined system elements such as databases, IoT broker, access management, etc. The second step is the placement of the system elements for future use, which would be the placement of the databases for example on the defined hardware or virtual machines. The last step is then to record the evidence that the system elements meet the requirements for the verification and validation processes.
In a real-life example, this would be the implementation of the PostgresSQL database and the Grafana dashboard. The implementation would include the setup of the system but also of the hardware that might be necessary for it. Also, the configuration and implementation of the dashboard are part of this process.
Integration
The integration process is described by the standard as the synthesis of set elements into a realized system. The integration process for data engineering is the connection of the data producers as an input and the data consumer as an output. Common data producers might be sensors or other external databases. Tasks for the output might be the configuration of an access manager for example.
Part of the integration process is the planning of an integration strategy. The strategy should include the order for aggregating the evolving system elements based on the priorities of the system requirements and system architecture definition. The focus should be on the interfaces, while minimizing integration time and cost and providing appropriate risk treatments. An integration strategy for data engineering would need to specify and prioritize the interfaces to which the system needs to be connected. This also includes the definition of the interfaces themselves such as the specification of supported communication protocols. Furthermore, the data transfer to and from the system might require also configurations or the development of middleware that transfers data and act as a bridge between systems. It is also important that the necessary permissions are granted to access the interfaces. For the domain of smart cities, it is also very important, if and how the integration may interrupt city operations and what kind of risk it might bring. Also, it is very likely that the use case would include a high number and a big variety of connections to data producers and consumers which might make it necessary to include data brokers or other concepts.
In the real-life example, this would be the configuration of the server and proxy so that the meter data can be collected. Another part of this process might be the development of middleware or protocols to send data to the database and from there to the dashboard.
Verification
As outlined in the standard, the verification process provides objective evidence that a system, element, or artifact meets its specified requirements and characteristics. It identifies anomalies in system components or processes using appropriate methods and standards, providing information for resolving these issues. In the execution of data engineering, this would be testing the data flow including also the correct generation and distribution of metadata.
The verification strategy, as defined by the standard, involves balancing what will be in the scope against existing constraints as well as determining the necessary verification actions. The strategy prioritizes the most appropriate methods for each verification action, along with the necessary systems and resources such as simulators, test benches, and qualified personnel. In some regulatory cases, all system elements may require verification. The strategy and schedule are updated as the project progresses to address changes or unexpected events, with a focus on minimizing costs, schedule delays, and risks to ensure the system is built correctly. A verification strategy for data engineering would be the specification of automatic and manual testing methods and the aimed test coverage. The smart city context also makes it very important, that any safety and security risks are verified as well.
According to the standard, verification procedures are performed at the appropriate time in the system life cycle, using specified environments, systems, and resources. Results are captured and compared with expected outcomes to assess the correctness of the system element. If anomalies are identified, the need for repeating the verification process is determined as they are resolved. For this use case, the right time would be the test before every deployment.
In the real-life example, the verification would include the use of unit tests for the middleware as well as manual tests of the system.
Validation
In line with the standard, the validation process ensures that the system fulfills its business or mission objectives and satisfies stakeholder requirements in its intended operational environment. It wants to confirm that the system meets validation criteria, with stakeholders providing validation. Any identified anomalies are addressed through the relevant technical processes. For example, for data engineering for smart cities, this might be a survey from users or an evaluation with the stakeholders.
The definition of the validation strategy includes analyzing the tradeoffs between what will be validated and the existing constraints. The prioritized strategy involves selecting appropriate validation methods and necessary enablers, such as simulators, test benches, and qualified personnel. The strategy and schedule are updated based on project progress, with validation actions being redefined or rescheduled in response to unexpected events or changes in the system.
To perform validation procedures, capture results from executing the validation procedure and compare them with the expected outcomes defined by the success criteria. Determine the degree of compliance and decide on acceptability, addressing any remaining uncertainty if possible. Validation activities are conducted at the appropriate stage in the system life cycle, in an environment representative of the operational context, with intended users or suitable surrogates, and using defined enablers and resources. Review validation results to ensure that the system meets the required services for the stakeholders.
In the real-life example, the validation would be an ongoing process in the system development. This would include regular meetings with the stakeholders such as the administrators and researchers.
System deployment and use
If a system which supposed to operate over an extended period, the system deployment and use phase assumes significant importance. The processes within the system deployment and use phase not only ensure the sustained functionality of the system but also lay the groundwork for its adaptability and evolution. They define requirements from the outset of the technical process, emphasizing the necessity for simpler maintenance, seamless operation, and preparedness for potential transitions or disposal. This proactive approach, via preparing strategies, increases the probability of the long-term success and sustainability of the data engineering system in the changing landscape of a smart city. The standard splits the processes into the steps: preparing, performing, and managing, which is the same as in the system realization phase. The focus in the context of data engineering is on the preparing side since the performing and managing steps including the strategy for each process.
Maintenance
The maintenance applied to data engineering would be the maintenance of the software and hardware. The standard includes also the aspect of the logistics in the process. Since a system is mostly software-based this aspect will be left out of the focus for data engineering. By the standard, the maintenance process is designed to sustain the system’s capability to provide products or services. It involves monitoring the system’s performance, recording incidents for analysis, and taking corrective, preventive, adaptive, and perfective actions to ensure restored capability. Tasks of corrective maintenance applied to a DMS would be the fixing of potential software errors but also correcting false data if possible. Preventive maintenance would be the development of tests such as described in the verification process. Adaptive and additive maintenance would be not only the development of new features for data management but also the administration of the access of data producers and consumers. Examples of perfective maintenance might be the improvement of performance and usability. Furthermore, the standard points out that maintenance needs can arise not only from failures but also from changes in interfacing systems, evolving security threats, and the technical obsolescence of system elements over the system’s life cycle.
A main maintenance strategy for this context should be set next to the described actions and also who would perform these actions. In this case, it would mostly be administrators but also software developers are important. The standard summarizes the maintenance strategy by a collection of approaches, priorities, schedules, resources, and considerations necessary to perform maintenance in line with operational availability requirements. It should include strategies for corrective, preventive, adaptive, additive, and perfective maintenance to sustain products or services and ensure customer satisfaction. The strategy should also detail scheduled preventive maintenance actions to minimize system failures without significantly disrupting operations. It emphasizes preventing the introduction of substandard or counterfeit materials and system elements. Additionally, the strategy defines the required skill levels and personnel for repairs, considering relevant health, safety, security, and environmental legislation, and includes measures to evaluate maintenance performance, effectiveness, and efficiency.
In a real-life example, this would include the maintenance of the hardware and software. Software maintenance could be the fixing of bugs and the development of new features for the dashboard, for example. A strategic decision could also be to give the maintenance to a third party which would then lead to new requirements regarding the system documentation.
Operation
By the standard, the purpose of the operation process is to use the system to deliver its products or services. This involves establishing requirements, assigning personnel to operate the system, and monitoring product/service and operator performance. It also identifies and analyzes operational anomalies to ensure alignment with agreements, stakeholder requirements, and organizational constraints. Operation in the context of data engineering is strongly connected to the use cases and involves usually the roles of administrators and customers. For the energy dashboard use case for example the operations may span from installing new sensors or other devices as well as different use of the data.
The operation strategy defines the approaches, schedules, resources, and considerations required to effectively perform system operation, typically established early in the system’s life cycle. This strategy includes ensuring the capacity, availability, and security of products or services throughout their lifecycle, from introduction to disposal. It also outlines the human resources strategy and qualification requirements, criteria for system release as well as re-acceptance, and methods for implementing operational modes, including contingency plans and resilience against cybersecurity threats. Additionally, the strategy incorporates measures for assessing operational performance, safety strategies for operators and users, environmental protection and sustainability plans, and procedures for monitoring changes in external conditions and operational activities. An operation strategy for the energy dashboard might include an analysis about who should and who might have access to the dashboard, who is, and what is necessary for the operation of the dashboard. The criteria that should be applied and tested before a new release of the dashboard could be software tests as mentioned in the verification section. In addition, some manual tests with the operators could also be discussed.
In a real-life example, the operation strategy might decide that the role of an administrator includes also user management and sensor management. A strategic decision could also be to separate the user management role.
Transition
The transition process interpreted for data engineering in smart cities could be a transfer from the system to new hardware. It can also be interpreted as a partly or whole transition of the software or data as well. By the standard, the transition process aims to establish a system’s capability to deliver services as specified by stakeholder requirements in the operational environment. This process involves moving the system in a planned and orderly manner to its intended environment, which could be new or changed. The transition ensures that the system is functional and compatible with enabling, interfacing, and interoperating systems. It also involves installing a verified system along with necessary enabling systems, such as planning and training systems, as defined in agreements. This process can be applied whenever the system or its elements are transitioned between entities or environments. Notably, during system upgrades, the goal is to minimize disruption to ongoing operations.
In all three cases, hardware, software, and data, it is important that all interfaces are documented in the transition strategy. The standard points out that the transition strategy in general should involve planning all activities from site delivery and installation through to the deployment and commissioning of the system. This strategy should also ensure that the system’s integrity is maintained and involves all stakeholders, including human operators. It outlines roles and responsibilities, considers facilities, and addresses shipping, receiving, and contingency plans. The strategy also includes training, installation acceptance, operational readiness reviews, and the criteria for transition success. Additionally, it covers rights of access, data rights, and integration with other plans.
In the real-life example, a transition strategy might include that the selected technology needs to support a certain data standard that is compatible with other systems.
Disposal
The disposal process, as described by the standard, ends the existence of a system or its elements for a specific intended use, manages retired or replaced elements, and addresses waste products in compliance with environmental, legal, safety, and security requirements. It deactivates, disassembles, and removes the system or elements from use, while managing waste by restoring the environment to an acceptable condition. Waste is destroyed, stored, or reclaimed according to legislation and stakeholder requirements, preventing inadequate elements from re-entering the supply chain. The process also maintains records to monitor operator health and environmental safety and ensures proper handling when only parts of the system are disposed of. This process applies throughout the system life cycle, from prototype disposal to final system retirement. The disposal of a system might be in this case either the whole system but as well as the removal of a sensor or the user of the system. Furthermore could also data be disposed of via a data governance strategy.
The disposal strategy, as required by the standard, should define schedules, actions, and resources to manage each system element and any resulting waste products. This strategy must ensure the permanent termination of the system’s functions and services, transforming or retaining the system in a socially and physically acceptable state to prevent any adverse effects on stakeholders, society, and the environment. It should also address health, safety, security, and privacy concerns during disposal and in the long-term handling of physical materials and information. Furthermore, the strategy should consider the potential transition of the system for future use in a modified form, including legacy migration. Since a data engineering solution is entirely digital they are not much physical waste. However, a disposal strategy for a system should include what kind of data should be disposed of or saved in what way. Furthermore, also the disposal of the hardware and the included data should be specified.
In a real-life example, the disposal strategy might be a data governance concept. This would include policies when data needs to be deleted or archived. Also, policies regarding the uninstallment of the PostgresSQL database and dashboard could be included.
Conclusion
This paper began with the goal of identifying a process that can effectively support and guide the data engineering process in the context of applications requiring an interdisciplinary approach across reliable developments in smart cities, metrology, and data science. To attain a comprehensive understanding, the introduction pioneered a comparative analysis of diverse approaches encompassing data engineering, data mining, data science, and system engineering. The objective was to evaluate their suitability in guiding the data engineering process. After careful consideration, it was determined that the most suitable guidance comes from the technical process outlined in the ISO/IEC/IEEE 12588 standard, focusing on the systems engineering lifecycle. The approach provides a structured and traceable framework to develop a traceable and trustworthy DMS as a base for trustworthy data-driven smart cities.
The second distinctive feature of this paper is the application of systems engineering to data engineering. The key insight of the interpretation is the definition of data engineering products and their interaction for all phases in further detail. This kind of approach can be used for all kinds of data-driven smart city projects. This might be smaller and enclosed projects such as the collection of sensor and meter data similar to the example in Section 5. Other examples might be bigger such as data-driven traffic management. In this case, the process would also start with the analysis of the context and the stakeholder requirements to define the right data-saving concept to save traffic data. Then it would continue with the selection of technologies fitting to the concept and further requirements such as the use of maybe streaming, cloud, or big data technologies. The end result would be a blueprint for a system design where every decision can be tracked back and which also includes the whole life cycle of the system. The blueprint can then be used to realize the data-driven traffic management system.
A drawback of the obtained results lies in the intricate nature of the process, potentially leading to prolonged development times. Conventionally, it has been noted by the standard that certain stages of the process may be shortened through the application of model-based systems engineering. In the forthcoming research, the systems engineering-based data engineering approach will be applied to the electric metering use case to progress the planning, designing, and implementation of the data infrastructure and management system of the Smart Metrology Campus.
Footnotes
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
