Learn free · topic 177
Big Data File Formats
This is again a very crucial topic but is heavily ignored in planning to host data in the big data ecosystem. It can impact the performance of analysis and even data pipelines execution.
Most of us know database tables are used to store data but at the backend where it's stored, we normally don’t bother. As far as databases are concerned, it is still fine to ignore the file format behind it but when it comes to Big Data it’s of utmost importance to plan the right File Formats for the right Data processing.
There are normally two kinds of data usage 1) to do ETL/ ELT 2) to extract data.
- Row File Format: In row format type, when the user requests to extract data from one column, the whole row is extracted in memory and a particular column is displayed on the screen.
- Columnar File Format: In columnar format type, when a user requests to extract data from one column, only that column data is extracted in memory and displayed on screen.
For example, there is a file that has 100 columns and 1 million rows. User runs a query for only 5 columns, now the behavior of Row file format type will be to extract all 100 columns with 1 million rows inside memory to process. This fully utilizes the memory whereas the requirement was only for 5 columns. Whereas the Columnar file format type will extract only 5 columns with 1 million rows inside memory to process.
- Row file format type is best for ETL/ ELT processing.
- Columnar file format type is best for data extraction for analysis.
Commonly used File Formats
- CSV: It’s a human-readable format as in ASCII or Unicode text format.
- JSON: It’s a human-readable format as in ASCII or Unicode text format.
- ORC: It’s not a human-readable format as in binary format.
- AVRO: It’s not a human-readable format as in binary format.
- Parquet: It’s not a human-readable format as in binary format.
Things to consider before choosing a file format.
- Structure of your data: If your dataset requires nested data then go for JSON, Avro, and Parquet.
- Performance: JSON is a very CPU-intensive file format. Considering nested data, Avro (file format type) can be used for ETL/ ETL, and Parquet (file format type) can be used to query/ extract data.
- Easy to Read: CSV and JSON are in ASCII or Unicode text format so is easy to read.
- Compression: Different file formats have different compression rates so based on storage limitation; the file format can be selected.
- Schema Evolution: ORC, Avro and Parquet provide some degree of schema evolution.
- Compatibility: CSV and JSON are the widely adopted file formats.
Finished reading? Test yourself with 10 questions on this topic.
Go to the questions →From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.