Learn free · topic 164
Data Locality
As per the name depiction, Data Locality looks like knowing the location of the data; it is partially relevant.
Traditionally, to use data for analysis, at the backend the data is fetched from storage/ data nodes into master nodes, for computation. There is a complete history of how initially whole data was pulled then blocks were pulled then the columnar pull concept came-in, etc. so we won't go into detail about how technology has evolved to process data.
Since semi and unstructured data has become a necessary part of the Data World, it comes with huge storage requirements i.e., in TB and even in PB. To host this humongous data, clusters of 100s of storage/ data nodes are set up, so it is not possible to pull this size of data for computation in master nodes which are normally few. It was a critical requirement to process queries where data is residing.
With the introduction of Big Data/ Hadoop/ HDFS where data is stored in the shape of blocks, the concept surfaced i.e., instead of fetching data for computation, let's take the query to the storage/ data nodes, process it there and fetch only processed data. This process of moving the computation close to where the actual data resides on the storage/ data node, instead of moving large data to computation, is called Data Locality.
In recent times, a new concept of separating computing and storage is surfacing.
In nutshell, when query runs on the same node where data is located, it is called Data Locality.
Finished reading? Test yourself with 10 questions on this topic.
Go to the questions →From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.