What is big data desensitization?
Big data desensitization, also known as data bleaching, data deprivatization or data deformation, refers to the deformation of certain sensitive information through desensitization rules to achieve reliable protection of sensitive private data. to safely use desensitized real data sets in development, testing and other non-production and outsourced environments.
Privacy data desensitization technology
Usually in big data platforms, data is stored in a structured format. Each table consists of many rows, and each row of data has Composed of many columns. According to the data attributes of the column, data columns can usually be divided into the following types:
Columns that can accurately locate a person are called identifiable columns, such as ID number, address, and Name etc.
A single column cannot locate an individual, but multiple columns of information can be used to potentially identify a person. These columns are called semi-identifying columns, such as postal code, birthday and gender. A US research paper claims that 87% of Americans can be identified using only postal code, birthday and gender information.
Columns containing sensitive user information, such as transaction amounts, illnesses, and income.
Other columns that do not contain user sensitive information.
Privacy data leakage types
Privacy data leakage can be divided into many types. Depending on the type, different privacy data can usually be used Leakage risk model to measure the risk of preventing privacy data leakage and desensitize data corresponding to different data desensitization algorithms. Generally speaking, types of privacy data leaks include:
Personal identity leakage. When a data user confirms through any means that a piece of data in a data table belongs to a certain person, it is called a personal identity leak. Personal identity leakage is the most serious, because once personal identity leakage occurs, data users can obtain sensitive information about specific individuals.
Attribute leakage, when data users learn new attribute information about a person based on the data table they access, it is called attribute leakage. Personal identity leakage will certainly lead to attribute leakage, but attribute leakage can also occur independently.
Member relationship leaked. When a data user can confirm that a person's data exists in a data table, it is called membership disclosure. The risk of membership relationship leakage is relatively small. Personal identity leakage and attribute leakage definitely mean membership relationship leakage, but membership relationship leakage may also occur independently.
Recommended tutorial: "PHP"
The above is the detailed content of What is big data desensitization?. For more information, please follow other related articles on the PHP Chinese website!