← All topics

Learn free · topic 203

Data Wrangling,Mangling orMunging

Data Wrangling or Data Mangling or Data Munging are used interchangeably and do the same function.

Table

Description automatically generatedData Wrangling is something that every ETL developer, Data Engineer, Data Analyst etc. do daily. It is something that is done before Data Mining, Data Science, Business Intelligence or Business Analytics.

In simple words, Data Wrangling is to prepare data for processing. Note, prepare data for processing, not process data. Yes, you heard it correctly, sometimes Data Wrangling techniques are used for permanent data fixes and sometimes to prepare data for specific use cases.

Let's take an example, refer to the tables above, fixing and streamlining data in fields i.e., Name, Phone, Birth Date and State will clean the data permanently which will assist in fast job processing, save analysts time for investigation etc. Secondly, you can see row#5 has been removed as, if the analysis is on Date then Year is missing from row#5 so this row will not assist in modelling.

Removing spaces from left and right of data in a field which we do using LTRIM or RTRIM, replacing empty date rows with 01/01/9999, removing special characters and many more are examples of Data Wrangling.

Every day, ETL developers and Data Engineers spend most of their time cleaning this kind of data. Policies and processes to do all the above from landing to staging reduce huge amounts of engineers’ effort. Not only human effort, the processing power of the CPU/ Memory will also be reduced heavily as it will not have to run checks here and there to make sure data quality is good. It also assists in creating the impression that Data is Good for processing.

In a nutshell, Data Wrangling is to transform 'Raw Data' into 'Quality & Useful Raw Data'.

General steps to perform Data Wrangling

  1. Data Discover: In this step, you collect data at one place if it’s coming from multiple sources then analysis it, understand it before taking any action.
  2. Data Structuring: In this step, you organize data in single format e.g., if data is coming in semi or unstructured format, convert it in structured formal.:
  3. Data Cleaning: In this step, you fix dirty data or remove unnecessary data, please refer to the above table in this topic. This can be done using automated algorithms or some third-party tools.
  4. Data Enriching (optional): This step is optional which is normally done as part of Data Self-Service. In this step, if data from another internal or third-party is required to add additional context to the first-hand then this can be done. This step normally is done as first-hand data is not enough to accurate decision making.
  5. Data Validating: This step is very important which can be repeated many a times to assure the following rules are followed on the data in-place. Validation rules can be quality, consistency, accuracy, security, and authenticity.
  6. Data Publishing: In this last step, you publish the data for stakeholder to consume.

Finished reading? Test yourself with 10 questions on this topic.

Go to the questions →

From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.