← All topics

Learn free · topic 261

Data Imputation

Data Imputation is majorly used in Data Science domain.

‘Imputation means, fill missing data.’

Table

Description automatically generatedData Imputation normally comes into play so business rules can be imposed on all data rather on subsets of data. Before Machine Learning Model runs, data imputation is done by either missing data is filled up or the whole row is removed.

Missing value can create many kinds of issues e.g., during categorizing, classifying, clustering etc., if data is missing then results will not be 100%.

Generally, in numeric datasets, missing values are replaced with MEAN values. Having said that, there many ways to deal with missing data.

Few Imputations Techniques

  • The first technique is where data is missing at random basis but before using it, it is completely removed from the dataset.
  • The second technique is also where data is missing at random basis but here missing values are replaced referring to the other data points available with current dataset.
  • Third technique is where data is not missing at random basis and mainly replaced using Mean/Median/Mode values.

Finished reading? Test yourself with 10 questions on this topic.

Go to the questions →

From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.