← All topics

Learn free Β· topic 259

Data Skew Issue

Data Skew Issue is one of the main hidden culprits for query performance 😊. It is normally referred to as Skewed data distribution which is used in Data Science domain.

Have you ever noticed, your query is correct, your joins are correct, your all indexes are enabled, you have high data quality, you have data standardization in place but still your query performance is not as expected? You noticed when you loaded new datasets, queries were flying but as time passed, the performance has gone down. You have added more CPU, more RAM, more nodes but NO performance is not like newly added data. It’s because of the Data Skew Issue.

Data Skew Issue is about when data is not distribution evenly. Uneven distribution of data at cluster level causes this issue. It heavily degrades the performance of the queries especially those with joins.

We all have faced a situation where query run exceptionally good till 99% and then hangs for 3-4 hours right lolz yeah, I can understand your pain 😊. Let me tell you the reason, it’s because of Data Skew Issue.

In the era of Big Data where we have huge data flooding in, if data is not distributed evenly across all the nodes in a cluster, then when we run a query, the MPP (Massive Parallel Processing) cannot be utilized properly, ultimately slowing down the query response.

One of the main reasons of Data Skew Issue is when we do Data Enrichment like when data is inserted and updated (it also includes delete tags). Data Enrichment moves data from one location to another creating imbalance of data across locations which are at backend, residing in difference nodes in clusters. For example, there are 10 nodes in a cluster and each node contain 10% of dataset evenly distributed across the cluster, now when we move data due to any reason and let’s say now 3 nodes contains 2 % data each and there are 3 other nodes now having additional 8% of data. So now when query runs, it will run in MPP mode using all 10 nodes till data is fetched from those 3 nodes with 2% data. For the data from those 3 nodes with 18% data will take longer time as now MPP is run on 7 nodes. This way, all 10 nodes compute power is not 100% utilized. I know above was confusing, read it few times and it will make sense 😊.

To cater Data Screw Issue challenge, we normally run Load Balancing exercises so after very Data Pipeline execution or Data Enrichment, data is evenly redistributed across the clusters.

In the era of Cloud and MPP, the first thumb rule is, when we run a query all possible compute power should be utilized for ultimate query performance. And this can only happen if your data is evenly distributed on possibly on all the nodes in a cluster.

In Data Science domain, there is a concept of Skewed Data Distribution which is related to Statistics.

Diagram

Description automatically generatedTypes of Skewed Data

  • Skewed Right (Positive)
  • Skewed Left (Negative)
  • Symmetric Skewed
  • Evenly Distributed Skewed
  • Two Mode Symmetric Skewed
  • Two Mode Non-Symmetric Skewed
  • Narrow Range Skewed
  • Wide Range Skewed
  • Outlier Skewed

Finished reading? Test yourself with 10 questions on this topic.

Go to the questions β†’

From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.