Learn free · topic 141
Data Subsetting
Data Subsetting is as simple as taking a small piece of data, with referential integrity in place, for some activity. For example, there is a requirement to build a data science model and it requires data for users aged between 25-35 from X Location having certain buying/ spending patterns. For this requirement, a particular dataset will be extracted for the data science development team to build and train their data models.
Don't confuse Data Subsetting with Data Masking. Data Masking is still required on this smaller piece of the dataset from a security point of view. Having said that, conducting a Data Subsetting exercise is not as simple. One first must understand the functional requirements of the use case, then one must also understand what data is there in the database as whole. Wrongs or partial Data Subsetting will end up producing wrong data model results and will damage trust in data by decision-makers.
Finished reading? Test yourself with 10 questions on this topic.
Go to the questions →From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.