Women in Big Data Global

Donate
×

Blogs

9 Pillars of Data Quality: The Foundation of Strategic Data Governance

FC

Key takeaways from “Data Governance Right Data to Right People at Right Timing to Make Right Decision” webinar with Prof. For Cheng by WiBD NW chapter

Data Governance: Building Trust in an AI-Driven World

In the modern data landscape, there is a common trap: focusing so much on the “sexy” output of a dashboard that we forget the foundation holding it up. Recently, the Women in Big Data, Pacific Northwest Chapter hosted a transformative session that shifted this perspective, focusing on the critical discipline of Data Governance.

The event was a centerpiece for the Pacific Phoenixes program, a cohort specifically designed to support women pivoting into data science or returning to the workforce after a career break. Led by Dr. For Cheng—an industry veteran with 36 years of experience at Intel and adjunct professor at George Fox University —the workshop moved beyond theory to provide what one participant called “what data governance really means” and its tangible benefits for the enterprise.

Breaking Silos and Defining Ownership

A recurring theme of the meeting was the human side of technology. One attendee reflected on how much she “liked the emphasis on breaking down silos between technical and business teams.” This resonance is vital; Dr. Cheng explained that data governance is not just a technical hurdle but “the exercise of authority and control over the management of data assets”. It is about moving away from the “siloed” approach where departments keep data in isolated spreadsheets. Instead, as the session highlighted, effective Data Governance requires practitioners to “go out and engage” with different departments to understand their unique requirements. By bridging this gap, technologists ensure that they are delivering the right data to the right people at the right time.

The Golden Rule: Avoiding “Garbage In, Garbage Out”

Dr. Cheng opened the discussion by framing data as a pipeline that begins with collection and ends with decision-making. He emphasized that while data analytics often assumes the data is already “good,” Data Governance is actually about guaranteeing both data quality and security from the very start. “Garbage in, garbage out,” he noted, reminding the audience that untrustworthy data leads to untrustworthy outputs. For those looking to grow in their data careers, understanding the authority and control over data assets—as defined by the DAMA Book of Knowledge—is essential for moving from a programmer to a strategic leader.

The Technical Foundation: Quality and Workflows

As one attendee put it, the deep dive into data cleaning and consolidation was a game-changer. Dr. Cheng walked us through the DAMA (Data Administration Management Association) standards, outlining the nine vectors of data quality that act as a safeguard against “garbage in, garbage out”. These include:

  1. Validity: The degree to which the data conforms to a specific format or rule. For example, ensuring that name fields do not contain numeric characters.
  2. Completeness: Refers to whether all required attributes and fields are present in the data collection. Missing fields, like a first name, can lead to duplicate or unclear records.
  3. Consistency: The requirement that different departments or systems use the same set of possible values for the same information, such as race or ethnicity categories.
  4. Integrity: The ability to reconcile and trust records across different departments. It ensures that if a person exists in two systems, their data (like gender) does not contradict each other
  5. Timeliness: The capacity to provide data immediately at the moment the user needs it for decision-making.
  6. Currency: Refers to the freshness of the data. For instance, yesterday’s data is more valuable and accurate for current reports than data that is three months old.
  7. Reasonableness: Whether the data makes logical sense. An example would be questioning a birth date that indicates a person is 150 years old.
  8. Uniqueness: The practice of removing duplicate records as much as possible to prevent inconsistencies within the dataset.
  9. Accuracy: The degree to which the data represents reality. This involves verifying if information, such as an email address, is real or fake through methods like two-way authentication.

Participants specifically noted the value of the “specific examples of vocabulary and the business situations” used to illustrate these points. For instance, Dr. Cheng’s practical case study showed how to use Python and SQL to automate the consolidation of messy city government records into a single, verified master list.

Securing the Pipeline: 10 Dimensions of Data Security

The conversation also turned to the critical nature of protection. Dr. Cheng detailed 10 dimensions of data security, ranging from confidentiality and auditability to incident response.

A highlight of this section was the “Least Privilege Principle”—the idea that organizations should minimize access to only those who truly need it to perform their roles. This balance between accessibility and security is a common pain point for organizations attempting to stay compliant with regulations like HIPAA. Dr. Cheng emphasized the “least privilege principle”—the idea that people should only have access to the data they absolutely need for their role.

One participant sparked a vital conversation by asking how organizations can balance accessibility for decision-making with these strict security controls. The answer, Dr. Cheng noted, lies in role-based access and a culture of constant “awareness and training”. He even shared a human moment, admitting that “even myself, many times [I] fell into the phishing email,” proving that vigilance is a shared journey for everyone in the Women in Big Data community.

The Future of the Data Pipeline

The session left attendees looking toward the horizon. While the overview of how data flows from sources to a final dashboard was clear, participants were eager to push further into governing data within ML pipelines. There is a growing desire to understand how data dictionaries and feature stores can manage shifting data meanings as they move through complex machine learning models.

This forward-thinking mindset is exactly what the Women in Big Data, Pacific Northwest Chapter aims to foster. As we move from “Model to Decision”—the topic of our next webinar on April 16th, 2026—we continue to build the technical and emotional confidence needed to lead.

Join the Women in Big Data

Are you ready to stop working in a silo and start building a foundation of data integrity? Whether you are a seasoned engineer or pivoting into a new field, your voice is needed.

Connect with the Women in Big Data community at https://www.womeninbigdata.org to access our mentoring programs, training tracks through DataCamp, and a global network of 30,000 professionals.

Recording and materials from this session are available here.

 

Related Posts