← All topics

Learn free · topic 349

ILM – InformationLifecycle Management

Let’s start by breaking down each part of this acronym based on the context of its usage.

  • Information – In this case the Information is related to the usage of the related data stored in a database or a data platform.
  • Lifecycle – The standard definition for life cycle is the different stages of life for a living thing. In this case the living thing is the Data. The Lifecycle for data is classified based on its usage as: - ‘HOT’, ‘WARM’, ‘COLD’ and that also could be extended to ‘DEAD’ for deleted.
  • Management – The process to facilitate the movement of the data related to the information over its Lifecycle.

So, using the context above, Information Lifecycle Management (ILM) is the process to optimize data storage by automatically tiering data to different locations based on its age, access frequency, and required read/write speeds. This approach minimizes storage costs by placing frequently accessed data in high-performance tiers and infrequently accessed or archived data in lower-cost tiers. ILM can also simplify data archiving and deletion according to data retention policies.

This was driven by the cost of Storage and scalability of Storage available for on premise databases and any Non-Functional Requirements from the business for the desired. A Database Appliance would normally have three tiers.

On Premise Storage Tiers

  • Tier 1: SSD (Solid State Drive)
    • Used to Store frequently accessed "hot" data for optimal performance, like frequently used tables or table partitions or indexes.
    • Speed: Extremely fast - Reads: up to 500 microseconds, Writes: up to 50,000 microseconds. SSDs access data electronically
    • Cost: More expensive per gigabyte compared to normal storage and external storage.
  • Tier 2: Normal Storage (HDD - Hard Disk Drive)
    • Used to store less frequently accessed "warm" data that still needs to be readily available, like historical data or backups.
    • Speed: Slower than SSDs - Reads: around 4-8 milliseconds, Writes: 5-15 milliseconds.
    • Cost: More affordable per gigabyte compared to SSDs, but more expensive than external storage.
  • Tier 3: External Storage
    • Used to support access to "cold" data for long-term retention, like old backups or compliance archives. This approach optimizes performance, manageability, and storage costs.
    • Speed: Varies depending on the technology used. External storage can connect via various interfaces like SAS, SATA, or USB, each with different speed capabilities.
    • Cost: Can be the most cost-effective option per gigabyte, especially for very large datasets.

Managing the Movement of Data Lifecycle

Data naturally transitions through different access frequencies over time. "Hot" data, which is frequently accessed, gradually becomes "warm" and eventually "cold" as its access needs decrease.

ILM facilitates this data movement transparently for users and systems. Data files referenced by tables or indexes are automatically tiered to different storage locations based on their access patterns. For very large tables, ILM can be applied at the partition level for even finger-grained control.

Example: Partitioning a Sales Transaction Table

Consider a 20 billion-row table storing sales transactions, partitioned by transaction year and further sub-partitioned by month. We can implement ILM with the following tiers:

  • Tier 1 (High-performance): Stores data files for the current year's transactions, further separated into sub-partitions for the most recent six months (Months 1-6) for optimal performance for active queries.
  • Tier 2 (Capacity): Holds data files for older sub-partitions (Months 7-12) of the current year and transactions from years 2 and 3. This tier offers a balance between cost and accessibility for data that may still be accessed occasionally.
  • Tier 3 (Archive): Stores data from years 4 to 7 along with their sub-partitions. This tier prioritizes lower costs for long-term archival purposes, with potentially longer retrieval times.
  • Data Deletion: Transactions older than 7 years can be deleted based on defined data retention policies.

By leveraging ILM, database appliances or storage systems can automatically migrate data between tiers based on access patterns. This approach ensures optimal performance for critical data while minimizing storage costs for less frequently accessed information.

But this management came with an overhead associated with the DBA’s time for planning, scripting, task management and the associated monitoring required.

So Traditional on-premises setups often required complex storage tiering for optimal performance and cost efficiency. But SaaS-based Cloud Data Platforms like Snowflake, Databricks, and BigQuery handle this complexity behind the scenes.

This is how Cloud Data Platforms (CDP) simplified storage management:

  • Data Loading: You simply load your data onto the platform.
  • Storage Management: The CDP provider automatically manages the underlying storage infrastructure, including cloud blob storage and compute resources.
  • Tiered Storage (Implicit): While you don't directly manage the tiers, CDPs often leverage tiered storage internally to optimize costs and performance based on data access patterns. This provides the benefits of tiering without the manual setup and maintenance overhead.

In essence, CDPs offer a managed service approach to data storage for analytics. They provide the benefits of tiered storage without the complexity, allowing you to focus on your data analysis tasks. But with the increase in data come the increase in costs and recently CDP platforms have started to support external tables on using Iceberg

Why is ILM relevant again?

Now Cloud Data platforms are starting to provide external table support that will provide the ability to tier your data based on the usage. For example, using Apache Iceberg file format that can enable the ability to move storage tiers. Apache Iceberg itself doesn't directly manage data storage tiers. However, it can work with underlying storage systems that offer tiering capabilities.

Here's how Iceberg works with tiered storage:

  • Leveraging Storage Tags: Iceberg can utilize tags supported by cloud storage solutions like Amazon S3. When writing data files, tags can be added specifying a preferred storage tier. These tags can then be used by the storage system to automatically transition data to different tiers based on access patterns or policies.
  • Integration with External Tools: Iceberg integrates with various data processing frameworks like Apache Spark. These frameworks might have their own mechanisms for managing tiered storage within the storage system. Iceberg can work alongside these tools to ensure data consistency even when files are moved between tiers.
  • While Iceberg doesn't handle tiering itself, it allows for smooth integration with existing tiered storage solutions. This enables data lakes to optimize storage costs by keeping actively used data readily accessible and archiving less frequently accessed data in lower-cost tiers.

Cloud Storage Tiers

Each cloud provider now provides tiering options for its storage. For example, Amazon S3 offers a variety of storage tiers designed for different data access needs and cost optimization.

  • S3 Standard: This is the go-to option for frequently accessed data. It provides high durability, availability, and performance, making it ideal for cloud applications, websites, and big data analytics.
  • S3 Intelligent-Tiering: This tier is designed for data with unknown or changing access patterns. It automatically monitors access and migrates data between frequent, infrequent, and rarely accessed tiers, optimizing costs without impacting performance.
  • S3 Standard-Infrequent Access (S3 Standard-IA) and S3 One Zone-IA: These tiers are suited for long-lived, less frequently accessed data. They offer lower storage costs compared to S3 Standard but have slightly higher retrieval times. S3 One Zone-IA stores data in a single Availability Zone for further cost savings.
  • S3 Glacier Instant Retrieval, S3 Glacier Flexible Retrieval and S3 Glacier Deep Archive: These tiers are targeted for long-term archive and data preservation. They provide the lowest storage costs but have retrieval times ranging from minutes (Instant Retrieval) to hours (Deep Archive).

Choosing the right tier depends on your data access requirements and budget. S3 Standard is ideal for frequently used data, while Intelligent-Tiering offers flexibility for data with unpredictable access patterns. Standard-IA and One Zone-IA are cost-effective for infrequently accessed archives, and Glacier tiers are perfect for long-term, rarely accessed data.

In summary, although the original concept for ILM that was created to manage the storage on your on-premises database appliances efficiently. The concept has now been adapted to cloud storage to help manage the costs based on the usage requirements of the data.

Finished reading? Test yourself with 10 questions on this topic.

Go to the questions →

From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.