← All topics

Learn free · topic 358

Textual orContextual ETL

The worlds of structured data and textual data are as far apart as the ocean is from the desert. These worlds are extremely different. But there is great value in being able to bring them together.

To that end there is textual ETL. Textual ETL reads text and disambiguates the text. As data is placed into a structured format from a textual format, the word that has been found and the context of the word that has been found are captured. That is the mechanics of how textual ETL operates.

How does textual ETL know what words to pluck out from text? Textual ETL knows by fiat of one or more taxonomies. The taxonomies alert textual ETL as to what words are important and need to be placed into a database for the purpose of analysis. In addition, the taxonomies provide the context for the word that has been selected.

Another feature of textual is inline contextualization. While many words can be contextualized by the usage of taxonomies, not all words can be treated this way. In addition to taxonomies, inline contextualization allows the word and its context to be selected by means of the positioning of the word in raw text and the proximity of the word or phrase to other words and phrases.

Once the text has passed through the disambiguation process, it is placed into a standard database.

By placing the extracted text in a standard database, the analyst can take advantage of the analytical software that depends on operation from a database.

Matching the data from the structured world and the textual world is still problematic even after the text has been converted to a standard database. In order to make a meaningful connection, there must be some connector between the two types of data. In some cases, there exists such a connector. In other cases, there is no such connector.

The lasting, a connector does not mean that analytical processing cannot be done.

In the case of no connection between the structured world and the textual world, analytics can still be done. Analytics against just structured data can be done. Analytics against just textual data can be done. But analytics against both textual and structured data cannot be done in the absence of some common connector of data.

One of the main differences between the worlds of textual data and structured data is that structured data is precise whereas textual data is probabilistic. For example, if in the structured world a deposit is made in a bank for $5000, this is an exact amount. The amount had better not show up as $4999.99. Precision is of the essence in structured data.

However, in textual data the interpretation is probabilistic. Suppose a customer says – “I like biscuits.” Does that mean that they really like biscuits? Do they like biscuits just a little? Are they ravenously hungry for biscuits? The truth is that the word like can be interpreted in ALL of these ways.

Text is especially useful because it can be used to examine documents that contain large amounts of text that have not been entered in a database. Especially useful is the notion of a card Catalog for documents. The document card Catalog is similar to the library’s card Catalog. The document card Catalog can be used to quickly and efficiently find documents in the corporation. Instead of the library using a card Catalog to find a book, the corporation uses the document card Catalog to find documents in the corporation.

Not only is the document card Catalog useful for finding documents, it is useful for finding documents that relate to each other.

In order to produce an effective document card Catalog, it is necessary to go into the contents of the document. Merely looking at the metadata attached to the document is hardly sufficient.

One of the features of textual ETL is its ease of use. The simple interface that is provided is like an ATM machine. If you can use an ATM machine you can use textual ETL. No programming. No consultants. No systems programmer.

SOME MORE INTERESTING FEATURES OF TEXTUAL ETL

Textual ETL is the technology that opens the door to analysis of text. Prior to textual ETL the world of analytics was strictly for structured data. But with the advent of textual ETL, textual data can now be analyzed.

THE WORLD OF TEXT

The world of text is huge. It is estimated that 85% to 95% of the data in the corporation is comprised of text. Furthermore, there is a tremendous amount of business value that is wrapped up in the form of text. So, it is very worthwhile paying attention to text in the corporation.

Some of the valuable data found in text includes

  • The voice of the customer: What are people saying about a product or a company
  • Medical records: In order to study a disease, it is necessary to look at the records of thousands of patients. Patient records are in the form of text.
  • Corporate contracts: Corporate obligations and opportunities are wrapped up in the form of text

And so forth

Nearly all of this data is untouched by today’s analysts in the corporation. The waste of text as a source of data on which to make decisions is like gold in the stream that you are too busy to pan for.

A screenshot of a computer

Description automatically generatedTEXTUAL ETL INFRASTRUCTURE

A close-up of a key

Description automatically generatedTextual ETL operates by reading raw text and turning that raw text into a standard database. Once in the form of a database, the text can be analyzed. Textual ETL makes use of taxonomies, ontologies and inline contextualization in order to process and transform text into a data base. Once the text is processed it can be placed in any database – DB2, Oracle, SQL Server, MySQL, etc.

The major business value of textual ETL is in unlocking text to analysis. And that is a big opportunity for many corporations.

But as text is unlocked, there are some interesting features of textual ETL that often go unnoticed. Let’s discuss some of those often-overlooked features.

TEXT IN ANY FORMAT

A group of squares on a black background

Description automatically generatedThe first feature is that textual ETL operates on text in any format. Stated differently, in reading and interpreting text, textual ETL does not require the raw text to be in any particular format.Early attempts at text analytics operated on only inputs of text documents that were in the same rigid format. The need for a homogeneous selection of input greatly restricts the ability of the analyst to gather and analyze textual data. Text is notoriously undisciplined and comes in many different formats.

Having to have text aligned in the same type of format is an unwanted restriction. And textual ETL does not have this requirement.

As an example of the freedom allowed by the formatting flexibility of textual ETL, consider some real estate documents – deeds of trust, quit claims, foreclosure, etc.

Textual ETL can read all of these different document types and take data from each of them and combine the data into a cohesive, unified understanding of the data.

For example, textual ETL sees “owner” in one document and records that information. Textual ETL sees “trustee” in another document and records that data as the same data as owner. The text is different, but their logical meaning is the same. It does not matter to textual ETL what the source of the document is or the format of the document. Textual ETL can read and interpret the document in any format and can tie the data together when appropriate. The data from the documents all end up in the same database with the same contextual interpretation of the data where data is grouped together by its logical meaning, not its physical manifestation.

DIFFERENT FILE TYPES

A black background with white text

Description automatically generatedAs important as is the freedom to read and interpret any kind of document, there are other aspects of document ingestion that are important. One of those aspects is the ability to read and process text that is housed in different file types. Textual ETL reads many different file types – JSON, txt, pdf, excel, and so forth.

A black background with white text

Description automatically generatedThe ability to read and interpret many different document types and file formats is one of the features that make textual ETL so easy to use.

THE TEXTUAL DATABASE

A second major feature of textual ETL is the database that is produced by textual ETL processing. Many different programs produce databases. But the database produced by textual ETL is fundamentally different than the databases produced elsewhere.

A screenshot of a computer

Description automatically generated

Consider a classical “normal” database.

The following figure shows that in a standard database that the data in a column is all the same type. For example, all the data found in the name column is a name. And all of the data in the sex column is a designation of gender. There is homogeneity of data within the column throughout.

This standard database design arrangement has been around for as long as there have been databases.

The standard database is shaped by the fact that metadata is used to define the structure that defines the database. In other words, when the database designer sits down and defines the structure of the database, the designer designates that one column is for name, another column is for sex, another column is for account, and so forth. The column name itself is a form of metadata and restricts the type of data that will be found in the column.

A black background with a black square

Description automatically generated with medium confidenceThe practice of defining the column types in a homogametic manner as part of the database design is a standard and well-known practice.

A black background with a black square

Description automatically generated with medium confidenceThe result is that metadata is used in the definition of the database structure for a “normal” database. That metadata then defines the placement of data in the database.

THE TEXTUAL ETL DATABASE

Now consider a database formed by textual ETL. The database formed by textual ETL has very different properties than the database that is normally created.

A black background with a black square

Description automatically generated with medium confidenceFrom an external perspective the database defined by textual ETL appears to be the same as a database formed in a standard fashion.

But that is hardly the case.

In a database formed by textual ETL, there is a column known as classification (or context).

A black background with a black square

Description automatically generated with medium confidenceThe classification column contains the classification of the “word” that is found in the same row as classification. For example, flu is a disease. Eliquis is a medication. Male is a gender, and so forth. The classification column gives meaning/context to the word in the same row.

The metadata (or the context) of the word is contained in the database structure itself. As such, the elements found in word are free to be any kind of word that is encountered in the document. The word column is heterogenous not homogenous. This is in stark contrast to the columns of data found in the standard database that are all homogeneous.

Contrast the word column in a textual ETL database with the name column found in a standard data base.

A white flag in the dark

Description automatically generated

A black background with red arrows

Description automatically generatedIn the name column of the standard database all elements are names. But in the word column of the textual ETL database there are many different kinds of words relating to many different things. The word column in the textual ETL database is free to contain any kind of data found on the document being processed. This freedom is necessary for document and textual processing.

The freedom and flexibility of the database structure found in a textual ETL database provides great freedom for the analyst. The analyst can get a truly multi-dimensional view of the data found in a document. This freedom is necessary because textual documents are undisciplined in their content.

FLEXIBLE ANALYTICS

In the example, the analyst can tell which disease the patient has, the medications taken, the gender of the patient, the marital status of the patient, and so forth. Given the free form nature of text and documents that contain text, this freedom is absolutely A screenshot of a computer

Description automatically generatednecessary.

The net result of the database structure associated with textual ETL provides a freedom that is absolutely essential for the analysis of data that comes from a document. Text is freeform and the database that represents text needs to be freeform as well.

Finished reading? Test yourself with 10 questions on this topic.

Go to the questions →

From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.