Abstract
Cloud computing plays a predominant role in storage technologies. It enables the tenant user to deploy their infrastructure without any investment. Cloud storage offers flexibility with storage and sharing facilities using the Internet platform. Storing sensitive information such as clinical data requires high privacy preservation and is associated with serious concern over data privacy on the cloud platform. Privacy preservation becomes the most adherent issue when a large volume of data is stored in public clouds. Subtree anonymization using the bottom–up generalization (BUG) and top–down specialization (TDS) approaches has been widely adopted for anonymizing data sets. This ensures individual data privacy; however, it causes potential violations when the new update is received, and it suffers from valuing the k-anonymity parameter. In this proposed model, a pseudo-identity was anticipated to accomplish privacy preservation with maximum data utility on incremental data sets. Initially, the Data Set (DS) was partitioned in the preprocessing stage; subsequently, the processed data sets were clustered into groups. The genetic model was used for indexing and updating incremental data sets. This was consistent with repeatedly modified data sets. In the evaluation process, an incremental and distributed DS was deployed, and our model exhibited efficient and optimal performance for privacy preservation in comparison with existing models.
Introduction
Cloud computing along with big data structures are dual complex models that create impressive influences on IT services and research domains [1, 2]. Cloud computing elastically delivers significant computation energy and storage space capacity through a huge number of computer systems, allowing users to deploy big data solutions cost-effectively. Big data applications and cloud platforms offer significant benefits. Many public networks of infrastructure applications are moving to the cloud platforms. All these applications use big data models and have become increasingly privacy-sensitive. Therefore, our proposed model uses the healthcare data for data storing and processing. Subsequently, protection against privacy matters with big data applications is such a huge challenge. In fact, numerous cloud users are still hesitant to benefit from cloud computing owing to privacy and safety issues [3]. Confidentiality is one of the primary alarming disputes in the big data solutions that comprise numerous workgroups. Nevertheless, cloud privacy is a crucial issue. Data privacy and integrity must be notified and highlighted before uploading refined data sets on the cloud.
Cloud computing also delivers smart features for modern research applications, such as artificial intelligence (robotics), educational domains, and social networks [4]. Furthermore, cloud computing offers a multitenant environment similar to pay-as-you-go models; they are expedient for cloud consumers to stake their data and associate with each other. Consequently, several corporations and social welfare organizations have already migrated their IT works to the cloud platforms. Recently, numerous corporations and hospitals tend to position their healthcare data and service details on the cloud platform, e.g., Microsoft Health Vault (MHV). Individuals with data sets engaged in these cloud platforms are extremely concerned with privacy-sensitive matters. In case fraudulent practices are used to collect these data sets and compromise the privacy-sensitive data, this will cause considerable financial loss or severe impairment to the health of an individual. Apparently, in most of the cases, these data sets are stored in a cloud for sharing and utilization by multiple users, rather than just for storage purposes. Applying security features on all data sets for privacy protection is a forthright process [5]. Nevertheless, the effectual handling of encoded data sets on a multilevel cloud can be a fairly puzzling task, because most prevailing software executes unencrypted data sets.
Healthcare maneuver projects contain typical progressions of diagnosis, such as refined treatments and anticipation of syndromes, injuries, or other physical and psychological deficiencies in the human body. Significant advancements have been made in the services and drug production in the healthcare industry and other social welfare organizations. Hence, the healthcare industry may be the fastest emerging segment of the country’s growth.
It has been extensively recognized that the provision of highly excellent healthcare amenities is dependent on effective and competent health issue detection, pioneering elucidation, and medical resource provision [6], and it is purely dependent on the proper collection of drug information, utilization, and management [7]. Accordingly, many organizations collectively gather health information and share this information with each other. At present, complex situations exist in which numerous diseases can occur without any symptoms. In such cases, sharing and distributing medical data sets play a vital role in empowering the remedial evidence flow through these groups and then refining the superiority of healthcare amenities. The advent of cloud computing advancement and connectivity through the Internet enables cloud users to consume scalable resources with distributed platforms [8, 9].
Privacy preservation is widely achieved using data anonymization methods. The resultant models are more efficient and effective for distributed data sets [4]. Apparently, data anonymization consists of simply hiding an identity that could be the user’s private or sensitive information. Thus, the confidentiality of valid users is effectively preserved; however, some summarized collective information can be disclosed to the users for further analysis. The subtree anonymization scheme is a broadly accepted method for privacy preservation. In turn, this method produces tangible trade-off among data usage and privacy preservation. Subtree anonymization can be performed using two methods: bottom–up generalization (BUG) and top–down specialization (TDS) [13–15]. Nevertheless, the big data application models use a large volume of data sets on cloud platforms and are the most extensive challenge for the anonymization process. Once the data sets are scalable in nature, there is a lack of parallelization proficiency. However, the efficient model of BUG and TDS can solve and improve the enactment of the scalability factor. The k-anonymity parameter [18] is the most effective method for subtree anonymization. In this study, a hybrid model was adopted, which inculcates unique features and performs effective anonymization using k-anonymity.
Related works
There are utmost models investigated extensively by numerous people in privacy preservation on the cloud platform. The scope of our projected model is privacy preservation on incremental data sets [1]. The residual area that is one of the most usable platforms is cloud data nodes. In the following text, some of the identifiable privacy preservation techniques and prototypes are briefly discussed. In our anticipated model, privacy preservation methods were classified into three broad categories, as follows: privacy through statistical methods [2, 3], privacy through user-defined policy [4], and privacy through cryptanalysis [5]. Multiple privacy models are available for anonymization and have already been analyzed for privacy-sensitive information such as health care data sets and education data sets. Correspondingly, the I-diversity [10] and k-anonymity [11] are two extensively accepted privacy preservation models. The credibility of these two models is measured based on linkage and disclosure of sensitive information. In addition, two other privacy models, namely t-closeness [12] and m-invariance [13], have been widely adopted against privacy breaches. However, many types of anonymization approaches have been reported, such as generalization [14, 15], specialization [16], suppression [17], disassociation [18], and slicing [19].
In the following discussion, statistical analysis-based models for privacy preservation are classified. A MapReduce technique for privacy preservation is a widely implemented technique for incremental data sets. Author [23] proposed a Hadoop framework that integrates multiple methods. The hybrid model implies task partitioning initially and then scheduling through cache memory. This technique is named as Incoop Architecture. Subsequently, author [24] proposed a data management, architectural model through the stream approach. This model uses the stream-as-you-go tactic for accomplishing privacy for cloud storage. Another author [25] defined stateful and incremental data sets. This system is implemented as a layer on top of Hadoop/Pig. It implies continued process flow on upcoming data sets. Author [26] anticipated a dual and twin model using hashing and public keys. This model is also a Hadoop platform on incremental data sets that uses a one-pass analysis. However, all these statistical-based privacy models are not fitted for a large volume of data sets.
Typically, the generalization approach substitutes the domain values with root values in the taxonomy tree based on attribute values. The generalization methods are classified into four schemes, full-domain generalization, subtree generalization, multidimensional generalization, and cell generalization. Policy-driven privacy preservation is the most common for standard information processing. This is extensively rational for the large volume of data sets. Nevertheless, this is not beneficial for incremental data sets. Dual conventional access approaches, such as MAC and RBAC [27], are sufficient for the large volume of data sets. However, these dual models have some drawbacks; for example, they require a standard interface and acceptable privacy regulation by the user [28]. Another policy-driven privacy model is the UCON simulation model. This model supports basic control over contiguous updates and justified requirements. The user has adequate control over the data sets [29]. Nevertheless, these two models have some partial drawbacks that can be addressed using digital rights management (DRM), which supports inimitable authorization and perceptible management over control policies [30]. Subsequently, the privacy preference policy (P3P) is authorized by the W3C consortium [31]. This privacy model supports a web service that defines the privacy platform recognized by intelligent systems and retrieves and interprets the data inevitably used by the user agents. Finally, this privacy preservation model compares user preferences. Congruently, this will produce fair results. The subtree generalization method concentrates on the specialization of incremental data sets. This produces reliable and anticipated anonymization of data sets, which can be deployed using existing data mining approaches. It generates a reasonable balance between data utility and data consistency. In addition, this scheme comprehensively explores BUG and TDS.
In most cases, the exploitation of data indexing to support anonymization is crucial. The data sets are commonly stored in the taxonomy tree format. This produces an additional burden on generalization and specialization. Thus, the moderated Taxonomy Indexed Partition (TIPS) and Taxonomy Encoded Anonymity (TEA) are specifically needed. Although data anonymization is made faster through the indexing process, these approaches have often been failed to be adopted in parallel and distributed allocation systems. Cloud storage is complicated in terms of the data indexing structure. Author [32] proposed a distributed model that is concerned with privacy matters rather than handling the large volume of data sets. Nevertheless, the resulting approach achieves data utility. Another author [33] described the MapReduce model that is applicable to big data structures. This model reduces time complexity while handling the large volume of data sets. This model uses k-anonymity, and the k value is too small.
Numerous approaches have been proposed for handling privacy preservation over incremental data sets. Author [20] anticipated a trade-off technique using k-anonymity for incremental data sets. This model correlates the data sets using clusters. Correspondingly, this method requires less time for execution. Therefore, another author [21] analyzed the effects of the monotonic anonymity property on incremental data sets. The aforementioned approaches recommend multidimensional or cell generalization schemes [34]. Nevertheless, these models require data investigation.
In this study, the hybrid model for generalization and specialization was projected to anonymize the Data Set (DS). The proposed model was ordered as follows, the taxonomy tree-structured DS was partitioned and indexed accordingly. Subsequently, the genetic BUG and TDS algorithm was used for incremental DS, including data updates.
Fundamentals of subtree generalization
Information loss per privacy gain (I L P G ) and information gain per privacy loss (I G P L )
The presented study considered a medical diagnosis system such as electronic medical record (EMR), which includes more sensitive information, and ordered EMR as taxonomy tree format. The original records with multiple attributes are stored in the Data Set (DS). The main aim of this study is to produce a privacy model without loss of generality. Moreover, this study used k-anonymity for generating privacy models. The quasi-identifier is represented as the pseudo-identity P ID . P ID is equal to zero or at least k through which the pseudo-identity may not have dissimilarity with other identities [9].
In the fundamental subtree generalization structure, the domain hierarchy is generalized with all child nodes, which are nonleaf nodes or null. This study used dual optimal methods to accomplish generalization and specialization, namely a bottom–up method for generalization and a top-down method for specialization. The proposed model is demonstrated and implemented as follows. Moreover, data set partitioning is adopted for our EMR data set, which has highly sensitive values. Thereafter, the Modified Genetic Algorithm (MGA) is used for Data Set indexing and Data Set update.
In the generalization process, information for each domain is swapped with its root node in a taxonomy tree. By contrast, in the specialization process, the domain information is interchanged with its leaf node. Apparently, the generalization process is exemplified with the term G
N
G
N
: Leaf (P
t
) → P
t
. The specialization process is represented in the form of S
P
. S
P
: P
t
→ Leaf (P
t
), where the pointer P
t
is associated with all domain values. Intuitively, with generalization, the projected scheme causes minimal data utility loss, but the privacy factor is higher. By contrast, specialization focuses on more utilization of data; however, data privacy is disputed [22]. The notation Information Loss per Privacy Gain (I
L
P
G
) is used with generalization and is denoted as
To have control over the selection of the optimal anonymization procedure, the consent with generalization or specialization is computed through a separate search metric. It is the trade-off factor based on the information/privacy requirement; that is, BUG used I L P G and TDS used I G P L .
The term I
L
(G
N
) is used for generalization of information loss. DS
K
represents the set of records, K defines the attribute value, which is generalized with the Data Set (DS). The subscript notations R and L represent the root and leaf nodes, respectively. The term E (DS
K
) defines the entropy value used to compute the ordering with the taxonomy treestructure.
The entropy value can be measured based on the sensitive values S
V
. The anonymity values are accumulated to calculate the privacy gain of generalization. DS
K
, S
V
describes the number of sensitive values of DS
K
. The pseudo-identity is to define the anonymity of the domain value based on the datasets.
Est (P ID )← To determine the domain values of the pseudo-identity
I G L P for all domain sets are computed
Specialization S
P
mainly focuses on the individual selection process. This resembles the value for Information Gain per Privacy loss (I
G
P
L
).
The cluster group C g is labeled as valid.
Our projected model gave priority to healthcare data, which are more crucial currently. The medical care data are extracted as EMR in the conceptual model. These records are categorized into three separate tables: patient personal information, patient health profile, and treatment data. Presumably, in EMR, all data are stored on cloud storage in the form of a single entity taxonomy tree structure . The table is categorized into three disjoint unique identities: Categorized Identity (C ID ), Pseudo-Identity (P ID ), and Auxiliary Information (A I ). C ID is applied to identify the individual records using patient personal information. The P ID attributes are mixed, which helps to identify an individual record of a patient [9]. Finally, A I is an additional attribute for storing medical history, such as the type of health issues and the type of treatment given. The first two attributes pertain to structured data. The other attribute (i.e., A I ) might be structured or semi-structured information.
The input table is represented as the taxonomy tree format with the combination of the three basic attributes. Every information tuple value of I has set of attributes Attrib = {A1, A1, A1,. . . . . , A K }. Consequently, the attribute is formulated using these three identifiers. After partitioning, the user input is categorized as the actual table along with the auxiliary information table and the anonymized table with the pseudo-identity . Finally, the encrypted table is generated with the combination of C ID and P ID .
Algorithm: Partitioning of DS
Init: Input parameters: ,
For each
individual cluster with
form into g-groups; 1 ≤ q ≤ g
After computing the anonymity level (k) and its violation, the data set is merged into the single table I
Symmetric encryption is used in our model. A shared public key is given only to authorized users. The table A
I
is a combined attribute referred to as . The combined attribute retains the information tuples that are used to calculate the semantic distance between attributes. P
ID
is required to identify each individual Data Set (DS
i
) that can be partitioned with minimal information loss without the loss of generality. The first level of partitioning supports reduction of complexity with k-anonymization. The I-diversity retains privacy with all sensitive values. The pseudo-identifier attributes are formulated using range values (R
K
). The range values are also encrypted, offering assurance of uniqueness with k-1 records. The resultant values are partitioned into three individual tables and are stored on the cloud platform. The table containing patient personal information can be directly accessed by the user. However, patient clinical information and treatment-related information are stored in the other two tables. A single user may have multiple separate treatment details, all of which are stored without losing generality and incomplete privacy format. In the second segment of this data partitioning algorithm, the segregated tables are merged into a single table after generalization and specialization have been applied. This also ensures integrity and privacy. Authorization negotiation with the data owner ensures that the shared key is used to decrypt the medical data, because the key is not disclosed to any unauthorized persons.
Data set indexing and update
After the input dataset is partitioned into three different data sets, the dataset is indexed based on the domain values to enhance efficacy when the new dataset update occurs. The identical data sets cluster groups are ordered in same data nodes [15, 22]. The anonymized Data Set (DS) tends to produce the highest utilization with new updates. The Data Set is always represented by the taxonomy tree format .
Algorithm: Data set indexing
Init: Data Set (DS i )
Each DS having multiple nodes,
DS i = {DS1, DS2,. . . . . . DS n }
Initialize the control pointer P t
Form the DS into individual cluster group g
Whole group cluster C g
C g = {g1, g2, g3,. . . . . , g x }
g← Pseudo-identity for each group cluster
K← No. of connection links and domain values of each group cluster
Initialize the connection links
Con g ← Connection link within the groups
Con s ← Connection link between the groups
Con g ∈ Con s
Data partitioning involves pseudo-identity, which helps in acquiring maximum utilization when the new update occurs.
Each cluster group can be easily controlled by a unique pointer P t . The group cluster is arranged in a list C g = {g1, g2, g3,. . . . . , g x }. Subsequently, the group cluster contains all the three identities, namely C ID , P ID , AI. Each domain value is connected to the individual connection pointer Con g . All connection pointers Con s are arranged in ascending order; thus, the pointers are sorted when the new update occurs .
Algorithm: Bottom–up generalization and Top–down specialization
In the preconditioning process, the large volume of data set (i.e., clinical information [EMR]) is divided into relative tables. All these tables are directly stored on the cloud nodes. Correspondingly, the partitioned data set is generalized and specialized using pseudo-identifiers. The fundamental generalization G N and the specialization S P describe the skeleton map for engendering and diminishing the anonymity level. After partitioning is completed, the data set is indexed using their domain values based on their generalization level. Generalization and specialization are achieved without the new updates. Moreover, verification is performed to check either the k-anonymity violation or the overgeneralized data set groups.
Init: PG current - current generalization level,
ρ, ρ′ - Constant parameters to calculate the most recent updates.
Compute Data Set and Data Subset .
Map PG
current
with respective cluster groups C
g
= {g1, g2, g3,. . . . . , g
x
}
Algorithm: Data set update
After the new updates are received, the following anticipated approach is implemented. To fulfill the k-anonymity level with the new data set, two levels are defined, namely PG current and UG new . The primary purpose of generalization is to combine the existing cluster groups with new and upcoming cluster groups with the size of less than k. The non-k-anonymous cluster groups are again considered and combined with an existing group if their pseudo-identity is similar [15]. This will be extended when a slight difference is observed.
Generalization: Initialize the current generalization level PG current and upcoming generalization level UG new [0, 1]
Define current cluster group C g and new arrival cluster group
Initialize ρ, ρ′ pseudo-identity for current and new cluster groups
Every newly constructed group is connected with the existing group.
Determine k-Anonymity violation,
if|C g (P ID ) | ≥ k-Anonymity
then cluster group C g contain k-Anonymity.
elseif
|C g (P ID ) | < k-Anonymity
then cluster group C g is non _ k-anonymous
Specialization: Initialize if more than one domain value takes place.
DS R &DS L denote the root and leaf node of a data set, respectively.
Initialize the pointer P
t
The Privacy Loss per Information Gain P
L
(S
P
) is computed as
The specialization S P is more particular about the domain values. The input is given to the specialization from the generalized group incurring the highest I L P G (G N ). The parameter η is used to determine the violation with k-anonymity. Finally, the Privacy Loss per Information Gain is computed.
Experimental evaluation
Environmental setup
In our proposed model, all experiments were performed in the simulated cloud environment U-Cloud. U-Cloud was originated and developed by the University of Technology Sydney (UTS) [15]. This public cloud facility encourages several cloud platforms to compute the dynamics of novel simulated models. U-Cloud is implanted at the Faculty of Engineering, UTS [22]. The system architecture layer model is depicted in Fig. 1.
The high-end data systems involve three outmost configurations on top of hardware; the Linux operating system is installed. In the middle layer, the KVM virtualization software modules are installed, which helps in the virtualization of ground-level infrastructures and yields unified computing along with storage modules. To create virtualized data centers, Open_Stack is deployed for virtual machine management, resource provisioning, and finally scheduling allocation and task distribution [32, 33]. Moreover, at the user end, the Hadoop platform is installed for big data processing. Moreover, at the user end, the Hadoop platform is configured with the number of clusters, and each cluster consists of 30 virtual machines. In addition, each VM has two virtual CPUs with 4-GB memory.
Our evaluation process used healthcare data from EMR [9]. The healthcare data repository consists of many variations that are used for privacy preservation and distribution for next-level diagnosis and equivalent treatments. In the preprocessing process, the data are partitioned, and only basic attributes with respective identity are encrypted and stored in the cloud. For large data processing, the enciphering and deciphering technique is used for clinical information data sets with dissimilar sizes (50 k, 75 k, 100 k, 500 k, etc.). The partitioning process also enables three paradigms for handling healthcare information. First, the is defined to protect the individual’s information. Second, the categorized table with respect to the pseudo-identity and A I are preserved in anonymized individual information. Finally, the absolute table I is the combination of all the identities and A I . Data sharing of medical information is immune to accessibility for the cloud service provider. In addition, there are many prospects for privacy breach when data sets are stored on multiple data nodes using the connection pointer. The proposed model uses AES and the hashing technique using user private key for enciphering the final table before uploading. Therefore, the proposed model achieved first-level privacy preservation.
Initially, the data sets are described as . The control pointer is initiated, and the pseudo-identity is defined for each cluster group. Each data set is partitioned without loss of generality. The group cluster is encrypted into different data nodes. In our anticipated model, the data nodes are restricted to 10. The pointer is kept updated with the current generalization level when a new update occurs. The proposed model has credibility for handling large data sets compared with existing approaches [15, 22]. This study implemented the subtree generalization scheme for new updates [4]. The time complexity is highly consistent when compared with other models. Twin (BUGPr, TDSPr & BUGEx, TDSEx) notations are described to compute performance consistency with our model. The generated graph depicts (Fig. 2) the time variations, and both the proposed and existing models used the same data sets to compare the time complexity.
Experimental results and discussion
In this study, the experiments were conducted using healthcare information (EMR). The threshold value (k-anonymity) is constant for all data sets. However, the k-anonymity value is user defined. At the initial experimental level, the constant k-anonymity value is defined for both models. Subsequently, three ranges of values (2, 4, and 12) are defined for k-anonymity [22]. Moreover, the k value will not affect both models because of the dynamism for new and updated data sets. The k-anonymity leverages slightly on every new data set when the update occurs. The basic generalization and specialization levels are identical for both models, and slight difference is observed in their performance. The update of large data sets requires more time. Correspondingly, when large data sets are applied to existing models, exponential time is taken when updating the new and upcoming data set. More data sets entail the more domain values; subsequently, this will create the number of cluster groups. The data sets are processed several times to achieve the maximum results. However, more data sets require more connections between the nodes. This will result in more expensive handling of multiple data nodes. The advantage of our projected model is that it reduces multiple connection pointers when handling many data sets.
The projected model reduces time complexity and enables cost-effective handling of many data sets. In Fig. 2, the performance of the model is slightly higher than that of existing models. Our proposed model can reasonably increase the efficacy of privacy preservation on large incremental data sets.
Conclusion
This study investigated anonymization using the genetic approach to leverage Internet-dependent paradigms to optimize data utility. In this approach, the efficacy and tangible performance of a large volume of incremental and distributed DSs were briefly clarified. This model ensures privacy preservation for large-scale privacy-sensitive information, such as EMR. Our hybrid model shows the synergy of various privacy preservation techniques, which produces enough trade-off among data utility and privacy protection. The subtree anonymization scheme using BUG and TDS is a widely adopted model. Nevertheless, subtree anonymization is slightly difficult for incremental DSs and also suffers from the k-anonymity parameter. In our projected model, the DSs are clustered into groups using the pseudo-identifier. After data partitioning is completed, indexing is performed based on their domain values. This step provides the current generalization level. This genetic model is a repetitive process. To further improve performance, the Data Set is enciphered using AES and the hashing technique. The real-world medical Data Set from EMR is utilized in the evaluation process. This methodology is compared with existing models by fixing the constant k-parameter. Privacy preservation is crucial and has very interesting research scope. The future scope of this model is the competent scheduling of incremental anonymized data sets.
