Learn free · topic 107
Data Classification,Categorization and Data Clustering
These two terms are used in Machine Learning for Data Mining and Data Science models. As we all know, within Machine Learning, there are Supervised and Unsupervised Learning methods.
Data Classification/Categorization=Supervised Learning
Data Clustering=Unsupervised Learnings
Both these terms are used for patterns identification before complex algorithms come into action.
Classification is done based on the uniqueness of the data where e.g., all animals are tagged separately whereas Clustering is done based on the characteristics e.g., 1) Animals with 2 legs 2) Animals with 4 legs.
- Data Classification uses labeled data whereas Data Clustering uses unlabeled data.
- The most popular Data Classification algorithms in data mining are the K-Nearest Neighbor and decision tree algorithms.
- The two common Data Clustering algorithms in data mining are K-means clustering and hierarchical clustering.
- Data Classification output is known.
- Data Clustering out is unknown.
- Training data is provided for Data Classification.
- No training data is provided for Data Clustering.
These two 2 terms are heavily used in Data Mining and Data Science.
Finished reading? Test yourself with 10 questions on this topic.
Go to the questions →From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.