Learn free · topic 6
Types of Data
In data management, there are three fundamental data types: structured, semi-structured, and unstructured. These data types can originate from numerous sources, including databases, mobile applications, social media platforms, sensors, logs, CCTV footage, radio signals, and more. While structured and unstructured data are widely recognized, the classification of semi-structured data often leads to debate.
Structured Data
Structured data refers to information organized in a predefined format, typically consisting of rows and columns. For instance, imagine an “employee-contact-details” table containing fields such as employee code, email, mobile, LinkedIn URL, Instagram, and Twitter. Although not every employee will have complete details, the table structure remains intact when exported, with each row containing a set number of columns, even if some fields are empty. For example, a dataset with 6 columns and 10 rows will consistently display this format, regardless of missing data entries.
Common examples of structured data include tables, spreadsheets, CSV files, and text documents. Structured data is characterized by:
- A stable, predefined format.
- Known structure before data entry.
Easy and automated extraction of information.
An important attribute of structured data is its predictability. As an illustration, a facial image can be broken down into identifiable parts (e.g., eyes, nose, ears) within a structured 4x4 grid format. If all images are aligned consistently, coded algorithms can extract facial features, provided that metadata helps ensure consistency across the dataset.
Semi-Structured Data
Semi-structured data, unlike structured data, does not have a consistent set of fields across all entries. Formats like XML and JSON are often categorized as semi-structured because they allow for flexibility in the data structure. Referring back to the “employee-contact-details” example, semi-structured data might only include populated fields when exported, meaning rows with missing data will have variable numbers of columns. For instance, if a column in one row is empty, that row may export with only the columns containing data, resulting in varying column counts.
Semi-structured data is generally:
- Consistent within a flexible structure.
Not fully defined until data entry occurs.
Platforms like Facebook or Twitter exemplify semi-structured data, as user-generated content can vary widely in format.
Unstructured Data
Unstructured data lacks a predefined structure or consistent format. Traditional definitions describe it as information without a recognizable data model or structure. For instance, an image library with diverse subjects, such as faces, animals, or landscapes, may be classified as unstructured. However, metadata can sometimes help impose a semblance of structure on unstructured data.
Characteristics of unstructured data include:
- A lack of predefined structure.
- Unknown format until examined.
- Ongoing assumptions needed for interpretation.
- Continuous analysis required.
- Possibility for insights through machine learning models.
Typical examples of unstructured data include images, videos, audio files, emails, word documents, presentations, and PDFs.
Finished reading? Test yourself with 10 questions on this topic.
Go to the questions →From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.