Learn free · topic 36
Conceptual Data Modelling (CDM)
Conceptual Data Modelling (CDM) provides a high-level, business-oriented view of data by identifying core entities and the relationships between them without introducing technical implementation details. Its purpose is to describe how the business views its data, not how the database stores it. At this stage, the emphasis is on clarity and shared understanding rather than normalization, indexing, or storage optimization.
The primary goal of Conceptual Data Modelling is to capture the essence of business information in a way that non-technical stakeholders can understand and validate. Business leaders, analysts, and architects must first agree on what exists in the domain and how those things relate to each other before moving into more detailed design stages. CDM therefore serves as a bridge between business requirements and logical data modelling.
In a loan approval domain, for example, the conceptual model identifies high-level entities such as Customer, Loan, Application, Approval Decision, and Repayment Schedule. It then defines how these entities relate. A Customer applies for a Loan. A Loan results in an Approval Decision. A Loan has a Repayment Schedule. These relationships are represented visually as connected entities, as shown in the referenced diagram, where lines between boxes indicate how entities are associated. At this level, attributes such as interest rate, date fields, or data types are intentionally omitted. The focus remains on “what exists” and “how it connects.”
An important aspect of conceptual modelling is the definition of cardinality, which specifies how many instances of one entity may relate to another. For example, in the relationship between Customer and Loan, each Loan must be linked to exactly one
Customer. A loan cannot exist without being associated with a customer. However, a Customer may have zero, one, or many Loans. A customer may exist in the system without ever taking a loan, may have a single active loan, or may take multiple loans over time. This is expressed as a one-to-zero-or-many relationship.
Similarly, consider the relationship between Loan and Approval Decision. Each Approval Decision corresponds to exactly one Loan. However, a Loan may not yet have a decision if it is still under review. Once processed, it will have exactly one decision, such as approved or rejected. This represents a one-to-zero-or-one relationship.
In the relationship between Loan and Repayment Schedule, every approved Loan must have exactly one Repayment Schedule, and each Repayment Schedule belongs to exactly one Loan. This is a one-to-one relationship. The clarity provided by these cardinality definitions ensures that business rules are captured early and understood by all stakeholders.
Conceptual modelling also introduces the idea of many-to-many relationships. For example, in a university context, a Student can enroll in multiple Courses, and a Course can have multiple Students. At the conceptual level, this is represented as a many-to-many relationship. The resolution of such relationships into associative or junction entities typically occurs later in logical or physical modelling stages.
The strength of Conceptual Data Modelling lies in its simplicity and business alignment. It avoids technical complexity and focuses solely on identifying entities, defining relationships, and clarifying cardinality rules. By doing so, it ensures that everyone shares the same understanding of the domain before structural refinement begins.
Conceptual Data Modelling marks the transition from process thinking to structural thinking. Where process models describe how work flows, conceptual models describe what exists in the business and how those elements are related. It establishes a clean and stable foundation upon which logical and physical data models can be constructed.
From Conceptual to Logical Modelling
Conceptual Data Modelling defined what exists in the business and how entities relate at a high level. It gave us clarity about Customer, Loan, Approval Decision, and Repayment Schedule, and established cardinalities without entering into structural or technical detail. Logical Data Modelling (LDM) now takes the next step. It refines those business concepts into structurally precise representations that can be implemented in database systems, while still remaining independent of specific technologies.
It is important to understand that Logical Data Modelling is not a single modelling technique. It is a stage in the modelling journey. At this stage, different modelling families and design philosophies emerge, each optimized for particular needs such as transactional integrity, analytics performance, flexibility, scalability, or historical tracking. Logical modelling is where structural rules, normalization, constraints, and data integrity are formally defined before physical optimization begins.
One of the most fundamental transformations at the logical stage is normalization. Normalization organizes data into structured forms to reduce redundancy and improve integrity. In a simple loan example, a flattened structure might initially store customer details, loan details, and approval status in a single table. Logical refinement would separate these into distinct but related tables such as Customer, Loan, and Approval Decision. This progression from First Normal Form through Third Normal Form, and in advanced cases up to Sixth Normal Form, ensures consistency and eliminates duplication. Such normalized structures are typically favored in OLTP (Online Transaction Processing) systems, where data integrity and update consistency are critical.
However, logical modelling does not follow a single path. There is an important divergence between OLTP and OLAP (Online Analytical Processing) needs. OLTP environments prioritize normalization to maintain transactional accuracy. OLAP environments prioritize query performance and analytical simplicity. As a result, dimensional modelling techniques such as star schema, snowflake schema, or Unified Star Schema often appear at the logical stage for analytical systems. These approaches intentionally introduce controlled redundancy to enable faster reporting and simplified aggregation.
Beyond normalization and dimensional approaches, several specialized modelling families operate within the logical layer. Ensemble modelling techniques such as Data Vault, Anchor Modelling, Focal Point Modelling, and Hook structures focus on agility, auditability, and schema evolution. These techniques are particularly valuable in large, evolving enterprises where traceability and historical preservation are essential. Object-oriented modelling supports inheritance and complex structures, aligning logical design with software development paradigms. NoSQL modelling approaches introduce schema-on-read flexibility, document structures, key-value stores, wide-column patterns, or graph representations to support distributed and high-scale workloads. Temporal modelling techniques explicitly incorporate time dimensions into the logical structure, enabling historical reconstruction and time-travel analysis.
The logical modelling stage therefore represents the structural refinement of business concepts. It defines how entities are decomposed, constrained, and related in a way that preserves meaning while preparing for implementation. In many cases, the path follows Conceptual Data Modelling to Logical Data Modelling and then to Physical Data Modelling, where platform-specific optimization occurs.
The following table summarizes the major modelling families that commonly appear at the logical stage, along with their purpose, strengths, limitations, and how they might represent a loan domain.
Logical Modelling Families Overview
Technique | Purpose | Strength | Limitation | Loan Example |
Normal Forms (1NF–3NF+) | Ensure OLTP integrity | Reduces redundancy, enforces consistency | Complex joins, harder analytics | Separate Customer, Loan, and ApprovalDecision tables |
Dimensional Modelling | Optimize analytics (OLAP) | Fast queries, simple reporting | Controlled redundancy, limited transactional suitability | FactLoanApproval with Customer, Date, and Product dimensions |
Data Vault | Historical tracking & auditability | Full traceability, scalable integration | Complex for business users | Hubs: Customer, Loan; Link: Application; Satellites: LoanDetails |
Anchor / Focal / Hook | Extreme flexibility & schema evolution | Agile change handling | Steeper learning curve | Anchor for Loan; Hook linking Loan–Customer |
Object-Oriented Modelling | Align with OO software design | Natural mapping to code | Less mature in traditional DBMS | Loan as class inheriting from Product |
NoSQL Modelling | Scalability & workload flexibility | Schema-on-read, distributed design | Weaker schema discipline | Loan stored as JSON document |
Temporal Modelling | Track changes over time | Supports auditing and time-travel analysis | Storage overhead, added complexity | Loan with valid_from and valid_to timestamps |
Logical Data Modelling is therefore not a single method but a structural refinement stage where multiple modelling philosophies coexist. The choice of technique depends on workload, performance requirements, governance needs, and system architecture.
This bridge marks the transition from defining what exists to defining how it is structurally organized under formal rules. The next section will examine Logical Data Modelling in more detail, building on this foundation and preparing the path toward physical implementation.
Finished reading? Test yourself with 10 questions on this topic.
Go to the questions →From I Am Datapedia! by Mustafa Qizilbash, published here free by the author. Nothing about your reading is stored.