Data technology has changed almost beyond recognition over the past 35 years. Organizations have moved from mainframes to the cloud, analytics has become embedded across the business, and AI is creating entirely new ways to work with information.

But some of the biggest data problems have barely changed.

Dr. Peter Aiken has spent more than three decades studying and solving those problems. An acknowledged data management authority, Peter is an Associate Professor at Virginia Commonwealth University, President of DAMA International, and Associate Director of the MIT International Society of Chief Data Officers.

In a recent episode of Behind the Data, Peter discussed what decades of working with organizations have taught him about data management, why so much organizational data has little or no value, and why technology alone rarely solves the problem.

One statistic captures the scale of the challenge particularly well. Peter estimates that at least 80% of the data held by most organizations is redundant, obsolete or trivial.

So why are businesses still storing it, paying for it and moving it from system to system?

The technology changes, but the fundamentals stay the same

Peter’s career in data management began in the late 1980s, including work with the US Department of Defense on an initiative attempting to rationalize 37 different systems used to pay employees.

Looking at the processes behind those systems did not reveal much. As Peter puts it, they all showed essentially the same thing: “Data goes in, it’s processed, stuff comes out with your paychecks.”

The differences only became apparent when the team looked at the underlying data.

Different systems supported different categories of employees and different payment requirements. That gave the team much more objective criteria for determining which capabilities were actually required.

The work eventually contributed to a technique Peter and his colleagues called data reverse engineering. Instead of trying to understand complex systems purely through their processes, they could look at the data itself and identify the differences that really mattered.

More than three decades later, the underlying principle remains relevant: understand the data first.

Peter argues that while systems and technologies continually change, “the data components of our systems are the most stable, the things that last the longest, the things that change the least.”

Data evolves too, he acknowledges, but focusing on these more stable components gives organizations a stronger foundation on which to build.

Most data problems aren’t technology problems

Buying in new technology is often the most visible response to a data problem. But Peter argues that technology frequently represents only a small part of the challenge.

Research done in the last few years points to a persistent 80/20 split between organizational and technical problems. “Eighty percent of the problems we have in data are people and process challenges,” Peter says.

Technical problems can be difficult, but they can usually be solved. The harder task is making sure organizations have the skills, processes and behaviors needed to use technology effectively.

That changes how businesses should think about investment.

Peter gives the example of an organization with a million-dollar budget for a data quality or master data management initiative. Spending the full amount on technology might sound logical, but, he says, “if you don’t put four million dollars into making sure that people know how to use that from a people and process perspective, it’s not going to work.”

If the organization only has the original million dollars, Peter would allocate it very differently: “Buy $200,000 worth of technology and use $800,000 to make sure people use it.”

The reason is straightforward. “It doesn’t do any good to have good data if you don’t have people that know how to use that good data” and turn it into something valuable for the organization.

The same imbalance can be seen in the everyday work of data professionals.

Despite decades of improvements in technology, Peter says people working with data still report spending around 80% of their time doing what he calls “data munging”.

Part of the problem is how data professionals are trained. “We’ve trained all of our data people how to do really great things with algorithms,” Peter says, while the discipline of actually managing data has often been treated as a lesser concern.

Better tools can help. But without the practices, skills and processes required to manage data properly, organizations risk automating around the underlying problem instead of solving it.


Peter Aiken podcast

Why 80% of your data may be dead weight

Modern organizations have made it extraordinarily easy to create and retain data. Files can be copied in seconds, cloud storage can feel almost limitless, and keeping information often seems safer than making the decision to delete it. Over time, those small choices add up, leaving organizations with growing volumes of data that may no longer serve a useful purpose.

Based on both the measurements he has seen and his experience working with organizations, Peter says the pattern is remarkably consistent. “Our measurements show us that it’s a minimum of 80% in all organizations,” he says, referring to data that is redundant, obsolete or trivial. His own experience reflects that finding too. “Every single one of them I’ve gone to” has shown the same problem, he says, with some organizations putting the figure closer to 85% or even 90%.

The individual decisions behind that accumulation can seem harmless.

“It’s just too easy to make a copy of a file,” Peter says. But once another copy exists, the organization has to determine which version is current, whether both are still required and how each should be managed.

Multiply that behavior across thousands of employees, systems and years, and organizations can accumulate enormous quantities of information whose value nobody really understands.

Cloud migration can make the problem worse.

Peter describes many organizations as “forklifting” their data into the cloud. They take what already exists, lift it wholesale from one environment and drop it into another.

The underlying data problem remains. The difference is that somebody is now charging the organization to store it.

“Amazon or Oracle or whoever gets rich because they’re storing stuff that they don’t need to store,” Peter says. For many organizations, he believes one of the biggest immediate opportunities for savings is simply to examine their cloud bills and understand what they are paying for.

He has also encountered organizations paying to move the same information repeatedly between cloud and local environments.

Peter describes data going “up and down from the cloud and back and forth and back and forth”, creating costs that can reach hundreds of thousands of dollars.

Cloud infrastructure may make storage easier, but it does not remove the need to decide what information is actually worth keeping.

Put a dollar value on better data

One of the most powerful ways to make data management matter across an organization is to connect it directly to business value.

Cleaning data for its own sake rarely excites senior leadership.

As Peter puts it, telling somebody “I cleaned some data for you” can easily produce the response: “So what?”

The conversation changes when the outcome can be measured. If cleaning that data reduces the number of undeliverable targeted marketing campaigns, for example, “there’s some business value” that can be attached to the work.

Peter has seen the impact of that principle at a much larger scale.

During work involving US military logistics data, his team found what appeared to be “more than a billion dollars worth of tanks that had gone missing.”

The tanks had not physically disappeared.

Each tank was associated with around 30,000 pieces of data, including a field indicating when the equipment became obsolete. If the organization could not identify that field correctly, Peter explains, “you’ll still maintain it. And you keep maintaining it.”

The data problem therefore translated directly into unnecessary operational cost.

A similar principle emerged at another organization that was spending around $6 million every year correcting recurring data problems.

Peter describes the process as “data whack-a-mole”. Teams would repeatedly fix the symptoms without addressing the source of the problem.

Once the financial implications became clear, the argument for fixing the underlying issue became much easier to make. As the CFO concluded, “We’re not going to play whack-a-mole anymore. It’s just crazy to do that.”

Data quality stops being an abstract technical objective when organizations can identify the cost of getting it wrong.

Sometimes better data means collecting less

Organizations often assume that more data creates more insight, but every piece of information has a cost. Someone needs to collect it, validate it, store it, maintain it, and potentially govern it for years.

Peter saw the consequences of this while working with data from a child protective services agency. When responding to an incident, staff used an assessment with around 80 questions to determine what action to take.

When Peter and his students analyzed the historical data, they found that “half of the questions had no prospective value on the outcome of the safety of the child.”

Removing those questions reduced the assessment time from around an hour to half an hour. Repeated across the agency’s workload, that saving became significant. By the end of the following year, Peter says the agency had been able to transfer around $1 million from administrative overhead into activities directly supporting child safety.

In Peter’s words, the organization was spending less time “asking you questions that we really don’t care about the answer to” and more time delivering the service the data was supposed to support.

More data is useful only if it contributes to something valuable. Sometimes better data management begins by asking what you can stop collecting.

Use migration as an opportunity to fix the problem

Major technology changes create natural opportunities to improve data management.

They can also multiply existing problems if organizations simply move everything as it is.

Peter encountered exactly this situation with an organization moving away from a mainframe.

“The data in the mainframe was bad,” he says. But while that data remained in the mainframe, it was at least “in exactly one place.”

“The minute you move it out to other places, now you’ve got a big problem.”

Instead of treating migration as a purely technical exercise, Peter argues that organizations should use it as a governance checkpoint.

“The inflection point of transferring it to the cloud is really where you should put the governance in place,” he says, so that the organization understands the characteristics of the data before moving it.

His rule is simple: “Don’t put it out there if you don’t know anything about it, because it could be a bomb.”

That means understanding what the data is, who owns it, its quality, how it will be used and whether it needs to move at all.

Otherwise, cloud migration can become little more than an expensive way of duplicating existing problems.

AI makes the fundamentals more important

Few technologies have changed the data landscape as dramatically as AI.

Peter has already seen organizations use it in highly practical ways.

For him, this is “super useful” because the system handles repetitive work while leaving the experts to apply judgment. But more powerful technology makes the quality of the underlying data more important, not less.

“What is the fuel for AI? It’s data,” Peter says. “And we’ve got to get the data parts absolutely right to do this.”

AI may represent a significant moment in the history of technological development, but it does not remove the fundamental questions data teams have been dealing with for decades.

  • What data do we have?
  • What does it mean?
  • Is it accurate?
  • Do we need it?
  • Who is responsible for it?
  • What value does it create?

Start measuring earlier

After more than 35 years working with organizations and data, Peter’s advice to someone beginning their career today is remarkably simple.

“I’d have started measuring things earlier,” he says.

Even after a successful career, he believes beginning that process sooner would have made a difference. “If I could have started this process in the early seventies instead of the late seventies, we’d be further and we’d have gotten there faster.”

It is useful advice for organizations too.

You need to measure how much poor-quality data costs and how people spend correcting it. You also need to measure what unused information costs to store and move, and whether the information you collect actually contributes to better decisions.

Technology will continue to change, but the fundamentals of good data management have proved remarkably durable.

Organizations that understand what data they have, manage it deliberately and connect that work to measurable outcomes will be much better placed to benefit from whatever technology comes next.

To hear the full conversation with Dr. Peter Aiken and explore more lessons from his 35 years working in data management, listen to the full episode of Behind the Data.

Adobe Express - file-3

 

By CloverDX

By CloverDX

CloverDX is a comprehensive data integration platform that enables organizations to build robust, engineering-led, ETL pipelines, automate data workflows, and manage enterprise data operations.

Share

Newsletter

Subscribe

Join 54,000+ data-minded IT professionals. Get regular updates from the CloverDX blog. No spam. Unsubscribe anytime.