Two scientists working at lab benches and a data dashboard in a biotech laboratory
Back to Resources
Life Sciences
Data Governance

Stopping the Bleeding on an Unknown Data Liability

How a global biotech reduced Covered Data exposure before quantifying it. A DOJ Executive Order 14117 program proven inside the environment before the first dollar of a multi-year commitment.

Download Case Study

We didn't want another report telling us how exposed we were. We wanted to see our Covered Data, govern it automatically, and know it would keep pace as the rules change.

Director of Data and Analytics Governance
Key Takeaways

The Problem

The client faced a DOJ Executive Order 14117 compliance obligation described in its own SOW as of “unknown volume and severity.” Leadership could not see how much Covered Data the client held, or where it sat.

The Approach

Data Sentinel stood up its Trust Layer inside the client's own production environment and ran it continuously. Proven against three representative sources under proof gating before scaling. Risk reduction started on day one.

The Outcome

A self-learning classifier at 95%+ accuracy across all four Covered Data categories. A continuously shrinking DOJ exposure. A posture that absorbs new regulatory definitions without re-procuring software.

Results to Date

95%+

Classification accuracy across all four DOJ Covered Data categories.

3

Representative production sources cleared under proof gating before scale-up.

U.S. + China

Dual footprint governed under one continuously operated Trust Layer.

Adaptive

Posture that absorbs new DOJ definitions without re-procuring software.

An Unknown Liability, With a Deadline

The regulatory clock had already started. The client did not know how much Covered Data they held, or where.

Executive Order 14117 was signed in February 2024. The Department of Justice published the Final Rule implementing it in January 2025. Compliance phased in through the year.

The rule restricts covered transactions that give countries of concern access to bulk sensitive personal data on U.S. persons. Six countries are named. China is one of them.

Covered Data spans several categories: personal identifiers, precise geolocation, biometric identifiers, human genomic data and biospecimens, other human ‘omic data, personal health data, and personal financial data. Each has its own threshold. Human genomic data crosses into bulk at just 100 U.S. persons.

A global biotech operating commercial and research across both the United States and China is squarely in scope. When they engaged Data Sentinel, the compliance liability was, in the client's own words, of “unknown volume and severity.” They did not know how much Covered Data they held. They did not know where it sat. They did not know which cross-border flows carried it.

Why the Discovery Report Model Doesn't Reduce Risk

The default playbook for a compliance obligation of this shape is a discovery project. A consulting firm scopes the assessment, interviews stakeholders, samples systems, produces a classification report, delivers a remediation roadmap, and closes the engagement. Remediation gets scoped later, budgeted later, staffed later.

That model has two problems, and both hit exactly where this regulation lives.

The report is out of date the day it ships. Data keeps entering the environment after the sample window closes. Retention decisions get made or not. Regulatory definitions evolve. The classification snapshot describes an environment that no longer exists.

Worse, a classification report tells the company how much Covered Data was there. It does not reduce how much is there today. The exposure sits in a binder waiting for a project to be funded and staffed. Every day of that gap is a day the enforcement risk continues at its measured level, not below it.

Data Sentinel does not run discovery projects. The platform is deployed inside the client's own environment on day one and starts classifying immediately. Classification is continuous, not a snapshot. Risk reduction starts the moment the classifier is on.

Prove It Live, Then Scale

The client did not want a multi-year commitment before they saw the platform work inside their own environment. That constraint shaped how the engagement was designed.

Data Sentinel proposed a proof-gated deployment. The platform would run inside the client's actual production estate, not a sandbox. It would run against three sources chosen to span the environment's diversity: one structured enterprise data warehouse, one unstructured collaboration platform, and one analytics lakehouse. Each source type presents a different classification challenge. Getting all three right meant the platform could handle the rest.

Before any scan touched production data, the label taxonomy had to be aligned with the client's data classification governance and privacy office. Data-owner sign-off was required. Automated classification at scale, without governed labels, would have created expensive reclassification rework months later. Governance sequencing wasn't a design preference. It was the only way the automation would hold up.

The proof-gate had specific success criteria. Classification accuracy across the four DOJ Covered Data categories had to clear a threshold. The platform had to run inside production without disrupting normal operations. Data-owner sign-off had to be maintained through the scan cycle. Three sources. Three success criteria per source.

Every one cleared.

The proof also carried a second effect. While the POC ran, aggregate risk on net-new Covered Data started dropping, because the classifier was already flagging what was in-scope as it arrived. The reduction was real, not projected. Scaling was no longer a bet on whether the platform would work. It was extending something that already was.

1

2

3

Diagram of the Trust Layer pipeline: sources (Operations, Commercial, Clinical, Genomic) pass through Classify, Govern, Prove and Gate stages to downstream integrations, dashboards, analytics and AI initiatives

The Trust Layer pipeline. Data enters from operational systems, passes through Classify, Govern, Prove, and Gate stages, and exits ready for downstream consumption.

The Trust Layer Inside the Environment

The Trust Layer runs as a pipeline inside the client's environment. Data enters unclassified, passes through four stages, and exits ready for downstream use.

Classify.

The self-learning engine runs at 95%+ accuracy across all four DOJ Covered Data categories. Personal identifiers. Genomic and biospecimen data. Biometric and health data. Financial and location data. Each labeled as it moves.

Govern.

Retention, lifecycle, and exception handling. Policies aligned to the client's classification governance are enforced as data moves through the environment. Legal holds get labeled and preserved. Retention timers run.

Prove.

Audit-ready evidence produced continuously. Governance, compliance, and data quality dashboards report on what has been classified, enforced, and remediated. When the auditor asks, the evidence is already there.

Gate.

Every downstream request evaluated. Allow. Sanitize. Deny. Elevate for review. Analytics, dashboards, integrations, and AI initiatives only see what has cleared the pipeline.

What the POC proved.

The POC cleared every success criterion across every source type: 95%+ classification accuracy across all four DOJ Covered Data categories, three representative production sources cleared under proof gating, zero disruption to client operations, data-owner sign-off maintained throughout, and aggregate risk on net-new Covered Data dropping continuously as the classifier ran.

What that opens up now that the platform is scaling beyond proof:

What It Opens Up

  • A governed AI data pipeline. Analytics and AI running against classified, policy-eligible data with full auditability.
  • A continuously shrinking, defensible DOJ exposure. Aggregate risk drops as net-new data enters compliance, then drops further as historical data is remediated.
  • A posture that absorbs new regulatory definitions without re-procuring software.
  • Cross-border U.S. and China research that keeps moving while staying compliant.

From Triage to Managed Service

Phase One. Triage. Full production rollout beyond the three POC sources. Identity and single sign-on integration. Expanded connectivity to the next wave of sources across clinical trial systems and clinical operations platforms. ITSM integration goes live. Governance activated across the environment.

Phase Two. Assessment and Remediation. The bottoms-up historical scan begins. Historical Covered Data that has been accumulating for years gets sized, prioritized, and remediated. The Trust Layer that stopped the bleeding shrinks the accumulated liability.

Ongoing. Managed Service. Long-term managed-service operation with monthly strategic reviews. At that point, DOJ 14117 compliance is a state the environment maintains, not a project the client runs.

Prove it. Scale it. Manage it. Reduce risk at every step.

Download Case Study

Reduce Risk Before Measuring It.

The Trust Layer at the client did not wait for a discovery report to finish before it started reducing risk. It ran, it classified, it produced evidence, and the exposure started dropping. That is what a platform inside the environment does that a report cannot.

Download the Full Case Study

Take the client case study with you as a printable PDF. Also available in our resource library.

Download PDF

How trustworthy is your data?

Book a demo to see how a continuously operated Trust Layer becomes the foundation for compliance, quality, and AI.