Learn free · topic 167
Hadoop, HDFS and Hive
Obvious question, difference between Hadoop, HDFS and Hive?
Hadoop is an Open-Source Framework to store structured, semi/unstructured data. Default block size of 128MB (changeable).
HDFS is a Hadoop Files System just like FAT, NTFS, EXT etc.
Hive is only a wrapper on top of files stored in HDFS. Hive behaves just like tables in traditional databases, but it doesn't store any data, it just manages table structures on top of HDFS files. There are 2 ways Hive works, either store files in HDFS directly or create tables which will create HDFS files at the back end. The difference in these approaches is 1) each file stored in HDFS creates its own block so if there are 2 small files e.g., 50 MB each, then there will be 2 blocks created. 2) on the other hand, Hive manages its own blocks so if there are 2 small files of e.g., 50 MB each, then Hive will store both in a single block.
Consider 1000s of 1MB files float in from social media means 1000s of 1MB small blocks created impacting MPP performance. The best practice is to have blocks of 128MB.
Finished reading? Test yourself with 10 questions on this topic.
Go to the questions →From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.