← All topics

Learn free · topic 199

Data Splitting forData Science

As the name says itself, Data Splitting is about dividing data into multiple datasets which can be for multiple departments, for multiple regions, for multiple time periods etc. but here we will discuss how it’s used for data science?

Data Splitting for data science activities is an important aspect of creating, processing, validating, and testing machine learning models.

In Data Science, mostly 2 approaches are adopted to split datasets.

  • Approach#1: Dataset are split in 2 subsets i.e., 1) Training 2) Testing
  • Approach#2: Dataset are split in 3 subsets i.e., 1) Training 2) Validating 3) Testing

Let’s understand with an example, consider a dataset contains 100,000 bank customers, now a created model’s aim is to tag each customer under three categories e.g., Silver, Gold or Premium based on their credit history, buying behavior, transaction values etc.

Now let’s see how both approaches on with above example.

A picture containing diagram

Description automatically generatedApproach#1 is widely used by data modelers or data scientists where datasets are divided into 2 subsets. Referring to the above-mentioned customers example, model will run on first dataset for Training purpose where each customer is tagged with relevant category. In next stage, model will be run on the second subset of data to Test whether model is working is fine or not. The model’s output will be judged based on the accuracy of results on the second subset of data.

Approach#2 is widely used where it is utmost important that model’s output accuracy is at the highest possible value. Referring to the same above-mentioned customers example, the dataset is divided into 3 subsets. Same as approach#1, first set is used for Training the model whereas third set is used for Testing the model but before Testing third set an additional Validation stage is added where model’s output is evaluated before its run on the third set of data. In this 3-stage approach, the expected accuracy is mostly at its best.

‘The percentages of subsets can vary based on the criticality of the output i.e., whether model is for marketing purpose or for cancer patients.’

Along with understanding the above-mentioned approaches, it is also very important to know how to split the datasets. Below are a few data sampling methods to split datasets.

  • Random Sampling is where data is divided randomly with a risk of uneven distribution of valuable data across subsets.
  • Graded Random Sampling is where again data is divided randomly but with some parameters e.g., if the dataset contains customers from 10 countries, then random data division in done within each country so yes risk of uneven data distribution will still be there but with-in each country, not across all.
  • Non-random Sampling is where specific dataset is required e.g., sometimes model requires latest data.

Finished reading? Test yourself with 10 questions on this topic.

Go to the questions →

From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.