Learn free · topic 12
Data vs Metadata
If you have read Edition 2 you know what Data what Metadata is. Now the big question is, are both must-to-have or there can be a possibility to have one only.
Questions:
- Do we need Data when we have Metadata?
- My response is, No, we don’t.
- In other words, can there be Data with Metadata?
- Again, my response is, No.
- As one of friend Daniel Lundin said, ‘One person Metadata is Data for another person’.
Exciting topic right….. Yes, it is, as we know size of data is getting huge to humongous lolz 😊.
Let’s decode it…..
Here, we are only referring to Raw Data i.e., which is coming from source systems e.g., OLTP applications or social media or unstructured datasets like video, audio, PDF, email etc., into next datastores e.g., Data Lake or in Data Warehouse etc.
So, when it’s out in the next storage, do we need to keep raw data for life or for certain time e.g., to fulfill regulatory requirements or we can purge it once Metadata has been extracted?
As a decision maker, which kind of data do you need for your decision? Is it like what are the total sales of a product in a region or is it you want to see all the sale transactions of each product sold in a region? If I am a decision maker and if I want to decide to reinvest in a product, I would like to see TOTAL sales, TOTAL cost of the product and other expenses like TOTAL marketing cost etc. Decision makers normally look at the summarized data and summarized data is not actual data, it’s transformed data which is Metadata i.e., data about data right?
So, if one knows how far he/ she can go to generate Metadata then the need for actual raw data becomes very less. Now, someone can say, hey what about if decision maker wants to see transactional data, then? Can we allow them to hit the source system for every type of investigation? I would say, let’s do impact analysis, what is the percentage when decision makers decide to investigate. The reality is whenever there is occasion like something is wrong which are 5%-10% then investigation comes into picture. Now the next question is, do we need to store, process and service raw level data when it will have hit only 5%-10%.
Having said above, yes I agree, source systems normally don’t allow a lot of read hits, the question is why? The main reason is, source systems never want to compromise performance of their systems, which again is TRUE in case of on-premises infrastructure, but since the introduction of Cloud with auto-scaling capability, this challenge might not be there sooner or later. So, will there still be a need to pull and maintain raw transactional data from source systems into the next storage, that is a question mark. For me nah, no need. TIME WILL TELL 😊.
Finished reading? Test yourself with 10 questions on this topic.
Go to the questions →From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.