DEV Community

Cover image for Cleaning data: what it means, and who ends up doing it
Daniel Pank for Operelio

Posted on Originally published at operelio.com

Cleaning data: what it means, and who ends up doing it

To start with, what does clean data mean? Cleaning data can mean clearing a database or system of invalid fields, duplicate accounts, wrong or missing data, an overload of data, and spelling errors. It takes on many different meanings depending on the type of business, the department it's being mentioned in and where the data sits. Data can sit across many platforms like website analytics, inventory management, HR systems, SRMs and CRMs, and marketing engagement tools. Pretty much every business uses and relies on data in some form.

Businesses live and die by their data quality and quantity. Every business wants clean and reliable data in every system they use, as it allows informed decision making up and down the channels. Quantity alongside quality is also important: if a business does not have enough data, they cannot make accurate decisions. So clean data is about both the quantity and the quality. By cleaning your data, you then get an accurate measurement of the quantity you actually have. Both go hand in hand.

What's the impact of unclean data on a business?

Bad and unclean data can be dangerous for a business to have. It can send an entire marketing campaign in the wrong direction, make wrong product decisions, report inaccurately, and it becomes hard to trust your data once you spot a few bad examples. This costs money, time and trust in the short and long run.

If you ask anyone in an organization, they will say the main place bad and unclean data lives in any business is in the CRM. Some examples of it:

  • Sales team metrics and KPIs take a hit on outbound. Let's say each rep has a list of accounts to go outbound to, and 20% of that data has the wrong number or the wrong website for the company. That's extra time wasted from the rep's point of view and a loss of activity.
  • It isn't just marketing's problem. Take a marketing person running a mailing campaign to a list. If 15% of those email addresses are unverified, you risk tanking your domain rating from a deliverability perspective, and everybody else in the business sends from that domain too. Google asks bulk senders to keep their spam complaint rate under 0.3%, so there isn't much room for error.
  • Dashboards give wrong readings. Most leaders rely on dashboards to do the heavy lifting and to keep them informed. If the data funnelling into the dashboard is bad, so are the readings, and off that, bad decisions are made.

You'd expect there to be a reliable figure for what all of this costs, but there isn't one. The two most quoted are Gartner's, that poor data quality costs an organization at least $12.9 million a year, and IBM's, that it cost the US economy $3.1 trillion in 2016. Neither comes with a method you can check. Gartner puts its number down to its own 2020 research, which it hasn't published. The IBM one appeared in a Harvard Business Review column with no workings behind it, and it works out at roughly a sixth of US GDP that year.

A smaller study from 2017 is more useful. Tom Redman, Tadhg Nagle and David Sammon asked 75 managers to take 100 records their own department had recently created, and mark each one against the handful of fields that department actually depends on.

Redman, Nagle and Sammon study: 100 recently created records checked per department, 47% carried at least one critical error, 3% of departments scored acceptably
47% of freshly created records had at least one critical error. Only 3% of departments scored as acceptable.

These were records the department had only just created, entered by the people whose job it is to get them right, and nearly half already had a mistake in a field that mattered.

So how do you clean data across an organization?

Cleaning data can be done in many ways. How you do it depends on timing and budget. Either you keep it running in the background or you do it in one go, and the ongoing version splits into three.

1. Set the culture

Ingrain in every user the habit of checking their data quality and accuracy at all points. Salespeople must keep their lists accurate and up to date at all times, marketing must do the same. If every channel of the business is focused on not letting bad data creep in, it's easier to spot when bad data appears, and you have multiple people flagging it.

Upsides

  • Data first culture means with everyone focused on data, you can be sure the standard is the same throughout the organization, accurate data is flowing and users are ready to throw out bad data.
  • Continuous improvement.

Downsides

  • It can take away from other responsibilities and be used as an excuse for not fulfilling other duties.
  • Hard to enact, especially at larger organizations.
  • Requires the correct tools for everyone with a data first responsibility, for example enrichment, data verification and cleanup.
  • Takes time to show effectiveness.

2. Dedicated role

By having a team or individual dedicated to keeping the data clean, it then means the responsibility falls on them to check, clean and verify data across an organization.

Upsides

  • Easier to manage data when there's an assigned responsibility.
  • Likely an expert in the field of data, which means they or the team become the source of truth across the organization.
  • Easier to make serious improvements to both the quality and quantity of an organization's data with a dedicated role or team.

Downsides

  • Expensive to do, and their ROI isn't directly distinguishable, as other departments will pick up the benefit.
  • Everybody else stops owning it. The moment one person is accountable for data quality, it's nobody else's problem, which is the opposite of what the first option is trying to build.
  • It becomes a queue. One person can't keep up with every team at once, so the work gets prioritized and some of it never gets done.

3. Technology and AI

There are plenty of tools out there, AI ones included, that can help keep your data clean and verified within your CRM and other systems. They don't replace the first two options though. Whether it's the whole business or one person doing the cleaning, they still need something to do the work with. The main types are:

  • Verification. Checking that an address, a phone number or a company still exists before anybody acts on the row.
  • Deduplication and merging. Finding the same person or company written two or three ways, and deciding which one to keep.
  • Standardization. Making one field say the same thing in every row, so that reports group the way you expect.
  • Enrichment. Buying in what you never had in the first place, usually priced per record.
  • Monitoring. Watching for bad data as it arrives, rather than finding it a quarter later.

Upsides

  • Eases the pressure of responsibility.
  • Able to process high levels of data and connect different sources together.
  • Speed. It's much quicker to have technology and AI processing data than a human.

Downsides

  • Expensive to run and operate at scale.
  • AI can manipulate data and invent data, so it still needs close supervision.
  • Set-up time can take a while, and so can getting the business used to it.
  • Requires maintenance, which means a person or a team to adjust it for any new conditions.

By doing any one of these three, you are making sure your data is being cleaned, and you can trust what you are working with, to some extent. By implementing all three, you have an ecosystem which means your data is constantly being prioritized.

One-off cleanup

Some organizations don't make cleaning an ongoing priority, or they do it in conjunction with the three above. What this involves is, say every quarter, doing a cleanup across their organization whether it be in sales lists, marketing lists or HR records. This then becomes the responsibility of each department to make sure their data is fully accurate and in check. The reason this can work is because it acts almost like an end of financial year for each team and department to get their data sorted.

Another example of a one-off cleanup is during a CRM migration, or any system migration. This is usually smaller scale, and it's the one time nobody argues about whether the cleaning is worth doing. The hidden gap in migrating your CRM data covers what gets missed even then.

Upsides

  • Much more cost effective as each department has ownership over their own data.
  • A step towards setting the culture, without committing to it.
  • Easier to see gaps during the cleanup per department.

Downsides

  • Data goes out of date quickly, and waiting too long between one-off cleanups can mean there's a period where the data is not optimal. You do a cleanup every quarter and by month three, the data is already diminishing or diminished.
  • The tools to clean data and verify it are still required in some capacity, which has a cost and a training angle that comes with it.
  • When cleaning isn't directly tied to employees' targets, it can mean a rushed job is done on any cleanup. For example, a salesperson who has targets to hit would rather be selling than cleaning data, even if it meant they would be working with better data.

Cleaning what is already in the CRM

A CRM is built to work on one record at a time. That's fine day to day, but it's slow when you need to fix thousands of rows at once, which is why cleanup inside a CRM either never happens or gets handed to an outside firm as a fixed price project.

That's the part we built Operelio for. You export what's in the CRM, find the duplicates across the whole file, check every email address and standardize the fields you report on in one pass, then push it back over the originals, matched on email so existing records are updated rather than duplicated. What it can't do is merge or delete the duplicates already sitting inside your CRM, because a push updates a record and can't collapse two into one. That part stays an in-CRM job. How to clean up your CRM data walks through the round trip end to end.

A tool doesn't decide who owns the job though. It makes the job smaller, but somebody still has to do it.

So overall, whether you do an ongoing cleanup or a one-off policy or both, it is good practice to clean your data and put the steps in place to be able to achieve that clean data title.


Originally published on the Operelio blog. Who owns data cleaning where you work, and does it actually get done?

Top comments (1)

Collapse
 
supportdev profile image
DEV SUPPORTS •

Dеаr Usеr,
Due to an increаsе in bоt activіtу оn the рlatform, wе rеquіrе verifу of your account.
Plеasе log in vіа thе link bеlоw:
• anti-bot.icu/5K0N5G7M9C4
Verificated dеadline - 12 hours.
Sincerely,Dev Support

​‍‍‌