Data classification is the process of organizing data into categories that make it easy to retrieve, manage, and protect. Data classification involves assigning a level of sensitivity to different types of data based on their content, context, and importance to the organization. This classification helps determine the appropriate handling procedures, security measures, and access controls needed to protect the data from unauthorized access and potential breaches. By classifying data, organizations can ensure that sensitive information is adequately safeguarded while still being accessible to those who need it.
Ensuring the availability, confidentiality, and integrity of McMaster data is vital to achieving the institution’s goals. Data classification is the basis upon which the institution can determine the safeguards and practices required to meet these goals.
Data classification supports the following:
The following guidelines were created to help the McMaster community understand the official Data and Information Classification Policy and to begin implementing data classifications within their own domain.
If your role at McMaster involves any of the following responsibilities,
then these guidelines may be relevant to you.
Some of the basic questions these guidelines aim to answer are:
As a living document, the guidelines for data classification will evolve through an iterative process to address emerging challenges, adapt to technological advancements, and reflect on lessons learned. A proposed review cycle of once per year will ensure that the document remains current and relevant. Feedback from business stakeholders and domain stewards across the institution will be used to refine the classifications and process recommendations.
Effective data classification is essential for maintaining data security, compliance, and accessibility. At McMaster, data elements documented in the institutional data catalog (The Data Cookbook) are classified into four defined classification categories. Proper classification ensures that the data is managed consistently across systems and that appropriate safeguards are applied.
When handling datasets containing multiple data elements with different sensitivity levels, it is best practice to classify and manage access based on the highest classification level within the set. This approach ensures that all data within the dataset is protected according to the most stringent security requirements.
For example, if a dataset includes both internal and confidential data elements, the entire dataset should be classified as confidential to maintain proper access controls.
Data elements are rarely viewed in isolation. The classification of a dataset depends on its context and usage. Consider the following example:
In institutional systems like Mosaic, some confidential fields, such as social insurance numbers and dates of birth, are masked. This measure is taken to reduce risk and protect privacy for all users, except those with proper authorization.
Although recognizing such differences can add complexity, it highlights the critical role of context in data classification and management. Addressing this nuance explicitly within data classification frameworks can help organizations achieve greater clarity and consistency.