← All topics

Learn free · topic 217

Data Discovery

Data discovery involves the collection and evaluation of data from various sources and is often used to understand trends and patterns in the data. But understanding trends and patterns cannot be simply done based on data coming from sources such as.

  • Data can have Null values.
  • Data can have different types of date formats.
  • Data can have spelling mistakes.
  • Data can have special characters.
  • Data can have different codes for the same product e.g., a Glass code can be GLS in one source system and GS in another. So, it's very important to streamline the data sets before presenting them to end users.
  • Data can have duplicates.
  • And many more....

Data Discovery is different for different roles i.e., for Data Scientist or Data Analyst or Business End User or for C-Level.

  • Data Scientists mostly prefer raw data without any data cleaning, cleansing, and scrubbing for Data Dredging, Snooping, p-hacking, and Fishing [there are separate topics on these terms].
  • Data analysts mostly prefer data after it’s at least cleaned and cleansed.
  • Business End Users and C-Level expect data to be cleaned, cleansed, scrubbed, wrangled, and munged [there are separate topics on these terms] before it reaches them for decision making. Well, a fair expectation.

Data discovery, usually associated with business intelligence (BI), helps inform business decisions by bringing together disparate, siloed data sources to be analyzed.

Finished reading? Test yourself with 10 questions on this topic.

Go to the questions →

From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.