DEV Community

Deb_Ghosh_Analsyt Layer
Deb_Ghosh_Analsyt Layer

Posted on

Finding Personal Data Across Old Systems Before a Data Law Deadline

A new data protection law does not care how old your systems are.
It asks a much simpler question:

Where does personal data live across every system the company has ever used — and can you prove it?
Most engineering and security teams can answer that question for the systems they actively manage today.

They know where the CRM lives. They know which cloud storage contains HR records. They can usually identify the databases behind current applications.

The harder problem is everything that came before.
Old file shares. Retired databases. Email archives. Backups. Spreadsheets sitting on departmental drives. Systems that were migrated years ago but never completely decommissioned.
And then there is the data that isn't inside your environment at all: personal data handed to contractors, processors, and other third parties.
That is where data discovery becomes more than a security exercise. It becomes a compliance engineering problem.
Kyndryl is one of the larger IT services companies now offering data-security capabilities built around Microsoft Purview for hybrid and multi-cloud environments.
The interesting buyer question isn't whether automated data discovery exists.
It's how far it actually reaches.
The three data-discovery problems

This is why "we have a data discovery tool" isn't the end of the discussion.
For buyers evaluating data-discovery platforms, the harder question is often not whether a tool can find and classify data, but whether it can provide enough coverage and evidence to support the actual compliance decision. Analyst Layer approaches these technology decisions from the buyer's perspective, looking beyond vendor claims to understand what a solution actually covers, where the gaps and trade-offs sit, and what buyers need to verify before relying on it.

The architecture of the discovery problem matters.
A scanner connected to your current cloud environment may give you excellent coverage of the systems you already know about while completely missing the systems nobody remembers.
1. Does the tool find personal data everywhere, or only in modern systems?
This is the first test I would run.
A tool that searches current cloud storage and well-known applications may work perfectly — while still failing to answer the actual compliance question.

The risk may be sitting in an old file share that hasn't been touched since the last migration.

Our take: Kyndryl's data-security offering, built on Microsoft Purview, is designed to work across hybrid and multi-cloud environments. That gives it a credible foundation for broad data discovery.

What is less clear from public material is how effectively that coverage extends into genuinely old environments: retired file shares, legacy systems that never moved to the cloud, or storage repositories that have effectively become invisible to current operations teams.
Ask for a real example of the tool finding personal data in an old, forgotten system — not a cloud environment it was already designed to search.

For an engineering team, this should be a live test rather than a slide.
Give the vendor a representative legacy repository and see what it actually finds.
2. Does it classify personal data using this law's categories, or generic security labels?
Finding data is only half the problem.
The next question is what the system thinks that data actually is.
A data-protection law may distinguish between personal data and specific categories of sensitive personal data. A security platform may use its own classification system based on broader concepts such as confidential, sensitive, or regulated information.
Those aren't automatically equivalent.
Our reading: we could not find Kyndryl's data-security work publicly tied to another country's personal-data law as a specific implementation example.

That doesn't mean the platform cannot support the required classification.
It means the mapping needs to be demonstrated rather than assumed.
Ask the vendor to show, in writing, how its data classifications map onto the law's specific definitions — not just onto generic labels such as "sensitive" or "confidential."

For developers and security engineers, the useful artifact here is a classification matrix.
You want to know exactly which detection rule maps to which legal category and what happens when the system isn't confident.
3. What happens when the data sits with an outside vendor?
This is where a purely technical view of data discovery starts to break down.
A company can have excellent visibility into its own databases and still have a blind spot if personal data has been copied into a contractor's environment.

The data may no longer be physically inside your infrastructure.
The responsibility may not have disappeared with it.
Our reading: Kyndryl's own data-processing terms describe its role as a processor of client data. That's a normal company-to-company processing relationship.

But that's different from answering whether the discovery tooling also provides visibility into your organization's third-party vendors and the personal data those vendors hold on your behalf.
Ask whether the coverage includes personal data held by your own vendors, or only data sitting inside your company's systems.
From an implementation perspective, this is not something a scanner alone can necessarily solve.
You may need a combination of data inventories, vendor registers, contracts, questionnaires, API integrations, and evidence from the third party itself.

4. What happens to backup copies when someone asks for deletion?
This is one of those problems that looks simple until you draw the actual data architecture.
A user record exists in the production database.
It gets copied into a data warehouse.
A backup job captures it.
A disaster-recovery system replicates it.
An archive retains another version.
Now someone requests deletion.
Deleting the production record is one operation.
Knowing where all the other copies are — and what your policy requires you to do with them — is another.

Our reading: public material describes capabilities around finding, classifying, and protecting personal data. What is less clear is a concrete walkthrough showing how backup and archive copies are handled following a deletion request.

Ask for a specific walkthrough of what happens to backup and archive copies after a deletion request — not just the copy in the live production system.

For engineering teams, this should be tested against the actual backup architecture.
If the answer depends on retention policies, immutable backups, restore procedures, or scheduled expiration, those dependencies should be documented before the compliance deadline arrives.

5. How fast can the system turn a request into an answer?
A compliance deadline changes the engineering requirement.
It's not enough to eventually find the data.
You need to find it within the time allowed to respond.
That means discovery latency matters.
So do query coverage, false positives, false negatives, indexing strategy, connector availability, and the amount of manual investigation required after the automated scan finishes.
Our reading: general capability claims are easy to find in public material.
A specific, timed example — from receiving an actual request to producing the final response — was not.

Ask for one real example: how long did it take, start to finish, to answer a comparable data request in a previous engagement?

And don't accept only the scan time.
Ask for the end-to-end time.
That includes identifying the subject, searching connected systems, validating results, checking third parties, handling exceptions, and producing the evidence needed for the final response.

Where a tool like this fits
A solution built around Microsoft Purview can be a strong fit for an organization that already runs heavily on Microsoft's ecosystem and has most of its data in modern, documented environments.
It becomes particularly useful when the goal is to create one connected view of data across hybrid and multi-cloud environments and improve automated discovery and classification.

Where it does not fit
The fit is less obvious for organizations with large amounts of:
old on-premises infrastructure
legacy or retired systems
non-Microsoft repositories
poorly documented data stores
personal data held primarily by third-party vendors
It is also not enough for an organization that needs to prove, with a real measured number, how quickly it can respond once a legal deadline starts running.
That requires testing the entire workflow, not just the discovery engine.

FAQs

Does this mean Kyndryl's tool doesn't work for this law?
No.
It means the public material does not yet show the capability mapped specifically to this law's rules. That's a diligence gap worth testing, not a verdict on the underlying technology.

Is this problem unique to Kyndryl?
No.
Any large IT services company selling data-discovery capabilities faces the same fundamental challenges: legacy systems, third-party data, and backup copies are where straightforward discovery approaches tend to become complicated.

What's the one thing most buyers forget to ask?
Whether "finding personal data" includes the data sitting with an outside vendor.
Most teams naturally think about their own databases, file systems, and cloud accounts first.
The harder question is what happens to the personal data that left those systems years ago — and whether anyone can still tell you where it went.
For a data-protection deadline, "we have a discovery tool" is not the finish line.
The real test is whether you can identify the data, classify it correctly, account for its copies and third parties, and produce an answer quickly enough to meet the law.

That distinction is particularly important when the deadline is measured in days rather than months, because a discovery platform can appear comprehensive while still leaving important parts of the data environment untested. This analysis examines the practical questions buyers should ask about finding personal data in legacy systems, mapping classifications to legal requirements, accounting for third-party and backup data, and measuring how quickly the full response process can actually run.

Top comments (0)