← All topics

Learn free · topic 107

Data Classification,Categorization and Data Clustering

These two terms are used in Machine Learning for Data Mining and Data Science models. As we all know, within Machine Learning, there are Supervised and Unsupervised Learning methods.

Data Classification/Categorization=Supervised Learning

Data Clustering=Unsupervised Learnings

Both these terms are used for patterns identification before complex algorithms come into action.

Classification is done based on the uniqueness of the data where e.g., all animals are tagged separately whereas Clustering is done based on the characteristics e.g., 1) Animals with 2 legs 2) Animals with 4 legs.

  • Data Classification uses labeled data whereas Data Clustering uses unlabeled data.
  • The most popular Data Classification algorithms in data mining are the K-Nearest Neighbor and decision tree algorithms.
  • The two common Data Clustering algorithms in data mining are K-means clustering and hierarchical clustering.
  • Data Classification output is known.
  • Data Clustering out is unknown.
  • Training data is provided for Data Classification.
  • No training data is provided for Data Clustering.

These two 2 terms are heavily used in Data Mining and Data Science.

Finished reading? Test yourself with 10 questions on this topic.

Go to the questions →

From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.