Insights

Master Data Management is All Around Us

December 6, 2018

Master Data Management (MDM) is not only addressed by large organizations handling vast amounts of data but also by individuals in their everyday lives. Let’s be convinced.

Numerous definitions of MDM exist. One such definition suggests that it involves an organization’s effort to establish a single, primary reference source for all critical business data, leading to fewer errors and redundancies in business processes. However, the practical application of this remains a point of interest.

The operation of MDM can be illustrated through examples from various areas, such as in our personal lives where we can solve certain problems using MDM on our mobile phones, or in the business realm, demonstrating what MDM can achieve in a bank. Through these insights, we can broadly discuss the areas of MDM and speculate on its future development.

MDM in Your Pocket

Ask yourself: what data do you have on your mobile phone? Or, perhaps more simply: what data don’t we have yet? Let’s examine the basics. Primarily, a phone is meant for making calls, so let’s focus only on telephone numbers and related information.

The evolution of phone contacts has been significant. In the era of landlines, an “external file system” was required, alphabetically organizing the number owners and the numbers themselves. However, times have changed, and fewer of us remember the paper notepad beside the phone or the directory in phone booths.

With the advent of mobile phones, this information began to be stored directly in the device, marking the beginning of electronic data management. Initially, the space for describing a phone number was severely limited to just a few characters, without distinguishing between individual attributes like names, surnames, addresses, emails, etc. Data were stored on the SIM card, with only a limited number of contacts possible, typically in the hundreds.

Later devices allowed contacts to be stored directly in the phone’s memory, displaying contacts from two data sources simultaneously. You might have expanded the original contact of your grandfather in the phone to include his first name, last name, and birth date. Some kept the original contact on the SIM card, while others did not. When your grandfather called, some phones showed “grandfather” calling, while others displayed “Novak Frantisek.” Only you knew that was your grandfather, not the device.

Later, your grandfather switched to a different telecom operator, changing his phone number. He forgot to tell you, but fortunately, your wife knew his new number and sent it to you as a contact named “Grandpa Franta.” You quickly imported it and called your grandfather, postponing contact consolidation for later.

As time went by, Android offered to back up your contacts to cloud storage efficiently, which you welcomed and regularly backed up. It also proved practical to use the same central storage when you got an additional SIM card for your car.

One day, your phone fell to the ground and would not turn on. You were probably relieved that your contacts were stored in the cloud, but less so when you found that each restoration from the storage duplicated your phone records.

For a moment, you thought about manually fixing the phone directory but spent more time looking for a suitable app for deduplicating contacts and found that there were two types of apps. One for a one-time data fix and another offering an online deduplicated directory. The former needs to be run at regular intervals, usually when the level of chaos becomes unbearable.

The latter solves everything online for you but is slower and different. The deduplication automation worked well for 90% of the records, but for the remaining 10%, you would have preferred to cancel or modify the merging process differently.

Aging and Enriching Contacts

The longer you have your phone and contacts, the more you’ll notice data aging over time. Some contacts are active, others less so, and some have long been inactive. It would be great if the phone directory could keep statistics on calls for each contact and indicate the aging of contacts itself.

Similarly, having a preference for a specific contact based on the availability of the person would be an interesting feature. The corporate phone number would be automatically selected based on standard working hours, while your grandmother’s landline number would not be chosen on Thursday afternoons when she goes out for coffee with a friend.

The same applies to automatically enriching contacts with new attributes, such as LinkedIn, Slack, Email, WhatsApp, opening hours, or addresses. Unfortunately, this is more a vision of the future than a current reality. The feature for automatically syncing the phone book among friends and family is imaginable even now. It’s about extending services for cloud storage.

However, the question with external services is their security and, importantly, how the data will be handled in the future.

MDM in Banking

While MDM issues for the average person can easily fit in a pocket, large organizations face these challenges on an entirely different scale. Instead of handling hundreds of contacts, they process tens of millions of contacts from dozens of internal systems and must focus on security both externally and internally.

Externally, the primary concern is data access. Internally, it’s about how data is managed. MDM issues are often addressed at the level of the parent company and its subsidiaries. The same data thus appear across many systems in various forms, leading to inconsistencies in content and timing due to redundancy.

Master Data Management Areas

Let’s explore the areas of MDM we have touched upon and how they can be addressed. The issue can be divided into four basic areas: mastering, quality, integration, and data discovery. Each of these can be described in several use cases. We will specifically return to the example of the phone directory.

To implement an MDM solution, we must first identify what data we have and in which systems it resides, a process known as data discovery. Once we know this, we aim to bring the data together in one place for analysis, which involves data integration. Then we begin to focus on data quality.

Initially, we monitor the quality based on specific criteria, then strive to improve it through manual or automated adjustments. Now, all the prerequisites are met to focus on mastering the data, which won’t succeed without high-quality data at a central location. Mastered data serves as a reference source for other systems.

To the mastered data, additional data, known as metadata, can be added, which dictates how to handle these data.

Discovery

Data discovery enables a deeper understanding of the data to be managed with MDM, aiming also to identify which systems contain which data. Contrary to the assumption that everyone knows their data well, the reality is often different. Automated data discovery allows for identifying not only basic data types like strings, numbers, booleans, dates, and times but also more complex types such as names, addresses, phone numbers, postal codes, secondary phones, and emails.

Through data profiling, based on frequency analysis and histograms, various insights about the data values can be obtained.

Integration

Implementing MDM processes requires centralizing all data. The discovery phase initially provides information about the systems where the data reside, which is then transferred through integration to a singular location for further use. This integration process merges incoming data from various systems across different platforms (Windows, Linux, iOS, Android), utilizing diverse technologies (web services, REST API, SQL, CSV, MS Excel) in various segments (incremental, full snapshot).

The entire data transfer process must be managed since the data are provided at different intervals and are interdependent. This management process needs to be monitored, and its results audited.

Quality

Quality is crucial for subsequent data mastering, and poor-quality data can significantly degrade the outcomes. The quality level must be assessed based on specific criteria and then improved, either manually or automatically:

  • Typos – Commonly, typos can be identified using dictionaries.
  • Diacritics – They are fundamental issues in data quality.
  • Numbers – Incorrectly written numerical codes can be checked using mathematical formulas.
  • Date and time – Date and time entries can be evaluated and then corrected based on their content (e.g., 24:15) or format (e.g., 23_15).
  • Enumeration – Enumerative values can be assessed using code lists.
  • Address – Complex data structures, such as address points, can be evaluated against external data sources, which act as a data registry.

Every quality assessment provides a value that indicates how quality the record is and how it can be managed in subsequent processes. A very poor-quality record can completely alter the perspective on the final mastering.

Mastering

Mastering data addresses how to consolidate data from various sources, such as SIM cards, phone memory, corporate phone directories, SMS business cards, backups, and others, to ultimately provide a unified view of the records for surrounding systems. Data consolidation always pertains to a specific domain (phone, contact, address, company, etc.).

In mastering, groups of closely related records are formed. The closeness of records is determined by the attributes defined for the domain. For instance, in the domain of phone contacts, attributes include phone number, first name, last name, title, social security number, or tax ID.

Each of these attributes carries information of a certain quality, and based on the assessment of their quality, a representative for the domain is created, although more than one representative may exist for the group.

The best representative is often referred to as the “golden record” or ideal record, which is used for propagation to other systems or processing. Thus, a mobile phone user should only work with golden records.

For golden records, the goal is to make them easily searchable and editable. Users should be able to modify not only the rules for domain creation but also the criteria for creating a golden record. Golden records can be organized in various hierarchies and shared with other users.

MDM in the Future

Looking into the future, where automation and service provision are commonplace, yet still designed by humans, we find ourselves between the present and the era of Skynet. How could MDM function in this landscape?

A vendor physically delivers an MDM device to the customer, placing it centrally in the room. Local IT allows the device to access the internal network, where it acts as a man-in-the-middle, listening on the network layer to corporate traffic and identifying physical infrastructure devices like servers and printers.

Subsequently, it determines the specific applications and their versions based on protocols. Eventually, the application system analyzes the structure and content of the data flowing between them.

The system automatically creates a data lake and business glossary, which can be utilized by other systems, such as GDPR compliance or data anonymization. Additionally, it can monitor the data quality, providing recommendations for enhancement, with the guidance generated by an AI module.

The subsequent data consolidation and mastering become relatively straightforward tasks. The system suggests merge proposals, and fine-tuning the output flow according to granted permissions becomes the icing on the cake. An element that automatically adjusts both data quality and mastering is inserted between the source and the target, eliminating the need to alter data in the source system, thereby bypassing modifications in both the primary and target systems.

Automatic integration at the network layer enables the system to easily add or remove components. Such an MDM implementation in an organization’s IT ecosystem could facilitate a non-intrusive service delivery.

If this scenario seems too futuristic, consider that IT attackers and antivirus systems already operate in a similar manner. However, both groups use network traffic monitoring to achieve different objectives, not for metadata acquisition and data exploration.

Join hundreds of professionals who enjoy regular updates by our experts. You can unsubscribe at any time.

SUBSCRIBE - Sidebar Newsletter

More Insights