Listen to this episode:
About this episode
In this episode of Behind the Data, host Matthew Stibbe talks with Dr. Peter Aiken — associate professor at Virginia Commonwealth University, president of DAMA International, and a 35-year veteran of data management. Peter shares how his early work helping the Department of Defense consolidate 37 separate payroll systems led him to co-invent data reverse engineering, and why most data failures are people and process problems, not technology problems. The conversation covers the economics of data quality, why at least 80% of organizational data is redundant, obsolete, or trivial, and how to make the business case for data investment.
Meet the guest: Dr. Peter Aiken
Dr. Peter Aiken is an associate professor at Virginia Commonwealth University, president of DAMA International, and associate director of the MIT International Society of Chief Data Officers. He is the founder of Anything Awesome and has spent 35 years working hands-on with organizations including the US Department of Defense, Nokia, Deutsche Bank, Wells Fargo, and Walmart. He holds a PhD in Information Technology Engineering from George Mason University (1989) and is the author of 13 books, including Data Reverse Engineering, which grew out of his pioneering work on the integrated DoD process and data model — a model still in use today. Peter also hosts a long-running data management webinar series hosted by Dataversity.
Key takeaways
- 80% of data problems are people and process, not technology. There's no point having good data if nobody knows how to turn it into value. It's one thing to teach people to do great things with algorithms, but it's also important to train people on how to actually manage data.
- At least 80% of organizational data is redundant, obsolete, or trivial. And the convenience of the cloud has meant it's easy to just 'forklift' this data as-is. A lack of data management results in increased costs and data governance issues.
- Put a dollar value on data work, or the business won't act. 'Cleaning data' isn't what gives you business value - being able to say 'I've reduced the number of undeliverable targeted marketing ads' - that is a clear business saving.
- You need to speak the CFOs language in order to make the case for data investment. Currency, not concepts, is what gets people's attention.
Timestamps
- 01:04 Why one-third of the world's data centers are in Virginia — and whether we'll still need them
- 06:59 37 systems to pay people: how a data perspective solved what process models couldn't
- 11:29 What's changed in 35 years of data management — and what hasn't
- 15:33 The Lego kit problem: what happens when three teams build the same system with no shared guide
- 17:37 Putting a dollar value on data work — and a billion dollars in "missing" tanks
- 18:18 Tokenmaxxing and the danger of incentivizing the wrong behavior
- 23:39 The Data Doctrine: building an Agile Manifesto for data
- 27:15 80% of your data is redundant, obsolete, or trivial — so why are we paying to store it?
- 28:48 The biggest source of savings: auditing your cloud bill before you migrate
- 31:08 CDO tenure vs. CFO tenure: why data leaders need to speak in currency, not concepts
- 34:13 Career advice: start measuring earlier
Episode transcript
Matthew Stibbe (00:01) Welcome to Behind the Data with CloverDX. I'm your host, Matthew Stibbe. Today's guest is Dr. Peter Aiken. He has an illustrious resume and we've had already several interesting conversations. He is an associate professor at Virginia Commonwealth University, president of DAMA International, Associate Director of the MIT International Society of Chief Data Officers, founder of Anything Awesome and previously Data Blueprint. With thirty-five years in data management and thirteen books to his name — how did you find the time? — he's a data geek's data geek, and I am delighted to welcome him to the show. Welcome, Peter.
Peter Aiken (00:40) Thank you, Matthew. It's a pleasure to be called a data geek's data geek. I take it as a compliment.
Matthew Stibbe (00:44) It was so meant as a compliment. I was trying to find the nicest way of going, oh my God, I can't believe you're on my show. So, let me start with a question I ask everybody. What are you geeking out about at the moment?
Peter Aiken (01:04) So it's really interesting, but I've got a local issue. I live in rural Virginia, and that means I only got a cable connection to the internet in April of this year. This is 2026. And at the same time, as much as I love data and have lived my life in it, I'm really opposed to some of the economics around the data centers that they're putting in place. I think it's not a good idea for a federal government or a local government to be investing in this kind of stuff when there's so much volatility in the market. But as far as geeking out, it gets lots of data. So we've got, for example, in Virginia, we have one third of the entire data centers in the world here in Virginia. So what are we doing that's so special about this? Let's look at this and see if we can figure it out and maybe put some of these things on the moon or in space or in Texas or in Aberdeen, depending on where they want to be. But don't put them where they're not wanted.
Matthew Stibbe (01:59) One third of... in Virginia, that's an astonishing statistic. No wonder you're enthusiastic — for people listening and not watching, he just held up a badge saying "grow tomatoes, not data centers." And I can see if they're very thick on the ground where you live, that could be a big issue. So, do you think putting them in space is the answer?
Peter Aiken (02:23) Well, it would certainly answer some things in the sense that it needs to be cooled and all that, but there's a lot of bad chemicals that we'd be putting out in space as well. And of course the real challenge with the data centers, and this is where it gets to the economics, is that we just don't know how long we're going to need them. The devices are so good in today's environment that the new Apple chips that they're putting out, and Intel is working as hard as they can to catch up on this, are going to be able to run the LLMs on your device. You won't need to send them to a data center. And that from a privacy perspective is going to be much more appealing to the consumers.
Matthew Stibbe (02:58) It's astonishing. I did some calculations the other day, very back of the envelope sort of, but my M3 laptop has more compute power than all the computers in the world in 1985.
Peter Aiken (03:15) And that was late. I started back in the seventies.
Matthew Stibbe (03:16) Yeah. And all the computers — including four and a half million Commodore 64s. Also, the point at which the fastest computer in the world became faster than the laptop I use now was 2000. So in 2000, my laptop could beat in compute power the fastest mainframe anywhere. Which is astonishing. So as you say, the hardware is overtaking or playing catch-up with the software. We'll come on to later what that means for data, because I think this data proliferation is a huge issue. But you've been in this world of data for thirty-five years or more. Tell us how you started. Where did you pick up the first of the many hats that you wear?
Peter Aiken (04:13) I was very fortunate in that I was doing a master's program and my master's thesis advisor came by to tell me that I was done. You know, that's how they do it. They wait and see if your stuff's piled high enough and deep enough, and then they just come by and say, you're done, right? And they said also, by the way, I signed you up for a PhD program at George Mason University.
Matthew Stibbe (04:35) That came as a surprise to you, huh?
Peter Aiken (04:37) Quite a surprise to me. But I checked into it. It was a really neat program in Information Technology engineering, done by some people I really, really respected. And so I ended up at George Mason in 1985. Finished there in 1989, but stuck around the Washington, DC area for a couple more years to start plying my trade. And the nice thing about George Mason was that the faculty there taught me not just how to be good with the things you have to do to be a PhD, but they also taught me how to be a consultant, which was an added bonus. And getting consulting training is fabulous if you've got a really good teacher on it. I had several.
And so I finished that process and then went into the Pentagon. And I was hired by something called the Center for Information Management at the Corporate Information Management Initiative. So they used two different things with the same acronym in the same office I was in. So there's a good data management issue for starters. By the way, this is Percy the Cat. He likes to participate on these.
And the group was doing something called information engineering, which is what I had been studying. So it was very nice in that sense. They were creating... our main takeaway project was that we needed to have something called the integrated DOD process and data model. And we put together a data model that really is as complicated as the background that you see on me. And those of you that can't see — there's books and guitars and heaters and all sorts of other things back there — in order to do it. And it was a tremendously complicated situation. But the Department of Defense hiring me into this had a very practical first problem to start on. And that was that they had 37 systems that paid people in the Department of Defense.
They needed, of course, exactly one. Now, I say one in the sense that one is not always the best answer, and we'll come perhaps back to that if we get the conversation in that direction. But clearly 37 systems was obvious, right? They all pay people. So the reason they hired me was because they had detailed process models for each of these systems, all on the walls and looking at these things.
Peter Aiken (06:59) And they all told them exactly the same thing. Data goes in, it's processed, stuff comes out with your paychecks. Pretty straightforward. I'm not at all pooh-poohing people who do HR work and figure these things out. But it is not rocket science, and we've been doing it literally for millennia. So we understand this process. We look at all those process models, all 37 of them, they all looked exactly the same and didn't give us any differentiating characteristics.
So we decided to look at it from the data perspective just to see if there was a difference. And we found out specifically that different systems had different capabilities for different people. For example, there is a category in the Defense Department for paying one-legged engineers working in waist-deep water underneath rotating helicopter blades on overtime. No problem. We need people to do that and we need them to get paid. So we've got one of those set up. Super, right?
Well, when the people in New Orleans called us up and said, why didn't you pick our system as one of the winners? We said, you didn't have the bit about them in the waist-deep water. And they looked and they said, you're right. So what I didn't think I was inventing, but in fact turned out to invent with a couple of my colleagues, Russ Richards and Alice Munz, was a technique called data reverse engineering. And they finished the process at DOD and had a good outcome and nobody nonconcurred. That's the way you disagree with things in the military — you say you nonconcur on this.
Now I'm stretching my things… This is 1989 to 1993, I guess, the period we were looking at. But these things were just so complicated and so hard to make a gestalt type decision. But when you have specific things and say, you don't have the one-legged engineers, you don't have the rotating helicopter blades, you don't have this — this one does. That's why we're picking this one. Everybody said that's a good reason. And they were so pleased about the result of the outcome, they then told me, you need to write a book. And I said, what? I'm 30 years old. I don't know anything about writing a book. And they said, do you want to ask that question twice?
Matthew Stibbe (09:09) Mm.
Peter Aiken (09:21) Now, I was not in the military, I was civilian military, but it doesn't matter. You get somebody that gives you an order, you go off and do it. So I sat down and worked on it for about a year and came up with a book called Data Reverse Engineering. There's one that you're gonna go out and put on your shelf, right, Matthew? I guarantee anybody out there that's having trouble sleeping, send me a note. I will absolutely send you a PDF copy of the book. It's guaranteed to work better than the old Ambien prescriptions from the old days.
Matthew Stibbe (09:47) And less addictive, perhaps.
Peter Aiken (09:48) Quite a bit, yes. So doing this kind of work turned me on to the enterprise architecture field. And of course, enterprise architecture is largely about data. Think of all the things that are going digital today or even just in AI. What is the fuel for AI? It's data. And we've got to get the data parts absolutely right in order to do this. So I went out of the DOD and eventually went back into the university community, which is a lot of fun.
Peter Aiken (10:18) And instead of being just a university professor, I go and live with companies on things. I'm getting old now, so maybe I won't do quite as much of it, but I spent four years in Finland with Nokia. I spent five years with Deutsche Bank in and out of Germany, Wells Fargo, a little bit of work with the intelligence communities, all sorts of different — three years in Northwest Arkansas. Do you know what's in Northwest Arkansas?
Matthew Stibbe (10:46) I've never been there.
Peter Aiken (10:47) It's a beautiful place. So Walmart — actually there's three companies in this tiny little town of Bentonville that make more than a trillion dollars. So there's something in the water in Northwest Arkansas that clearly makes you smart and able to do trillions. I'm probably not there yet, but you get the picture. So instead of staying all the time in the university and teaching, which they would have been very happy to let me do, I go out, loan myself to these other organizations, learn what they do and how it works. And that part has been just fabulous. I have the opportunity to work with literally hundreds of practices and thousands of people and I learn from each one of them I go to.
Matthew Stibbe (11:29) And in that thirty-five-year arc, what hasn't changed and what's one thing that's changed beyond recognition?
Peter Aiken (11:38) So, of course, the AI is the beyond recognition piece. In the old days, we created the DOD data model. It's still in use today, by the way, 35 years later. It was really good work that the team did in order to come up with that. And the basics of DOD are still pretty much the same. Your MOD is very similar to our DOD. In fact, they've shared parts of the model back and forth and done collaboration with it and all that sort of thing. So that part has not changed. We need to understand the data that we're going to use.
For example, I'll just toss out the number 42. Any idea what that means?
Matthew Stibbe (12:13) Yes, the answer to life, the universe and everything.
Peter Aiken (12:16) Because I, knowing you, Matthew, knew you had probably read The Hitchhiker's Guide to the Galaxy, which is one of the books back there. And I've been able to do that joke, if you will, around the world for 35 years — there's always been at least one person in the audience who goes, are you talking about the meaning of life? Yes, I am. On the other hand, 42 could be my age 25 years ago. And you say, okay, that's great. So now we know whether Peter's old enough to buy an adult beverage in the Commonwealth of Virginia, right? So there's a different meaning for the number 42. All of these things just end up being the most specifiable, the most objective, the most testable types of requirements that we can do. By the way, that's not an earthquake — that's the cat rubbing the mirror.
Matthew Stibbe (13:05) As I said, Peter, this is a very pet-friendly podcast.
Peter Aiken (13:13) Anyway, it's just absolutely fascinating to be able to work with these different groups. I started getting involved in DAMA because after I wrote the material in a book, we actually published a paper in it. And there was a local DAMA conference in Washington, DC. And they were like, the Department of Defense is letting you write about its legacy systems? Oh my gosh. Now, there wasn't anything interesting about the DOD legacy systems and most of them were being decommissioned. But it was still interesting stuff.
And we started to put together some work around these areas. And the legacy part of the world is the part that's been really, really stagnant. No matter what we do, everybody who's working with data says they spend 80 percent of their time doing what we call data munging. It's the idea that although I — you'll have to tell me this is correct or not, Matthew — I saw somewhere somebody said munging is wiping somebody else's nose in the UK. Have you ever heard that?
Matthew Stibbe (14:11) I've never heard that expression. I have heard the expression munging in respect of data. So from my perspective, you're in the clear. You munge away.
Peter Aiken (14:19) So the key idea of course is that people don't understand that we've trained all of our data people how to do really great things with algorithms, but the focus of training people on how to actually manage data has been considered a lesser piece. It's not in most university curriculums. People come out of it, they're doing it in the real world and they're saying, gosh, are there places I can go to learn more about this stuff? And eventually somebody will point them to DAMA and say, hey, DAMA's got some resources such as its Data Management Body of Knowledge that's considered a Bible at this point and gives some guidance on what needs to be done. We don't tell them how to do it. There's still a very big need for that aspect of it, but at least we've got the what part down pretty well.
Matthew Stibbe (15:06) And the need for that data munging discipline, that data wrangling piece is constant. But can we discuss a project — of the 200 organizations you've worked with, was there one that you still think about, successful or otherwise, that sticks with you?
Peter Aiken (15:33) Well, another thing that I use — I try to tell people that I'm introducing myself as a professor with a positive cash flow. So instead of the money flowing to me from the university — and don't worry, they pay me — but I also push money back into the university as a result of these activities. And when you look at the overall types of things that have been going on, it really is the same sort of thing where somebody is working on something and they're trying to get it right, but they don't get the data right.
I'll give you an example, again from our US military. Not anything super bad, but still an oopsie. We designed a system that had three instances of SAP — one here, one here, one here. And we handed them to different companies. So this one gets to company A, this one gets to company B, and this one gets to company C. And we were in a position to say, why are you doing that? And they said, well, we wanted to make sure that everybody did some work. I said, yes, but you've handed them three Lego kits and they're putting the Lego parts together in completely different fashions. They don't have a guide like your rocket back there over your shoulder, right? In terms of how it should be done.
And so each of the systems ended up being different from the other three systems, and they couldn't talk to each other. They turned it on and they were all SAP systems, so they all should have worked, right? And of course we knew it wasn't going to work. The various consulting people didn't, but the people who were managing the contract knew it wouldn't work.
So if there's something I've been able to do, it's been capable of putting some sort of valuation on these pieces, because it's not really great to be able to go in and help people saying, hey, I cleaned some data for you. And they're like, so what? But if you say, by cleaning data, I've managed to reduce the number of undeliverable targeted marketing ads — there's some business value that we can work with. So I at least put dollars on mine.
And you asked about the one that worked — well, we found more than a billion dollars worth of tanks that had gone missing.
Matthew Stibbe (17:37) You'd have thought they were hard to mislay, right?
Peter Aiken (17:39) And of course it was an accounting error. Turns out when you're keeping track of tanks, there's one variable in the 30,000 variables that you get for each tank. So when you buy a tank, you get 30,000 pieces of data that come with it. And there's exactly one of those that tells when it's obsolete. If you don't know which one it is, you'll still maintain it. And you keep maintaining it. And this is a tank that we don't — we would give to our enemies now. We had to say it that way.
Matthew Stibbe (18:08) When data projects fail, do they still fail the same way? Or do they fail for new and more advanced reasons with the passage of time?
Peter Aiken (18:18) Most of the time things fail because of the incentives that we've provided. So let's jump into a couple of AI-related topics. Uber, again very well written up, just happened this last quarter, incentivized their people. They called it tokenmaxxing. So they really wanted their technical people to use a lot of AI. Well, if you tell technical people to use a lot of AI, we become really good at doing that. And Uber ran out of their entire token budget for the entire year in a single quarter. So they were using it four times as fast as they had planned to use it.
Matthew Stibbe (18:55) It's such an easily gameable number, right? Maximum model, maximum thought, maximum parallelization to answer random questions. Why?
Peter Aiken (19:05) Matthew, if I said I was gonna give you a million pounds for breaking the windows that are in your room there, a million pounds for each pane of glass, what would you do?
Matthew Stibbe (19:12) Smash them all and shake your hand off, right?
Peter Aiken (19:14) Assuming I actually had the million pounds on the other side of it.
Matthew Stibbe (19:17) So, incentives — we were discussing.
Peter Aiken (19:24) So you look at what's happening on these things and people say — I used to get called into conferences all the time because people would say you can't put a dollar value on quality. And certainly I understand Deming's quote, right? But if you put some dollar values on things that aren't going right, people will pay more attention to them. And so this was really the key — to say, what is it that we're doing?
In at least the example that I gave earlier about a million tanks — they weren't really missing, but they weren't on the right lists, and consequently it was costing people a lot more money. So we really just untangled some logistics and supply chain data, and they were able to come out and say, wow, that's like a billion dollars, right? With a B.
So it does go into that. Another part of this though is that these are, for the most part, people and process problems. There's a wonderful guy in our industry named Randy Bean, who was up at the conference last week. He's run the same survey year after year after year. And one of the things he does — he does a wide variety, it's Randy Bean Data, and he's a great person to interview. You would have a wonderful conversation with him as well, Matthew. But he's asked over the past 10 years, are your problems mainly people and process problems, or are they mainly technology problems? And it is just a clear 80-20 split. 80 percent of the problems we have in data are people and process challenges.
And that's the thing I think that I have learned. We can fix the technical challenges. I'm not gonna say they are not trivial — they are difficult problems. We've got some great solutions. But that's only getting this much of the problem. And where we see most organizations having trouble is that they'll come in and say, I've got a million dollars and I'm gonna buy a data quality tool. Whatever tool — master data management tool, anything that you put into those pieces. But we've also discovered that if you don't put four million dollars into making sure that people know how to use that from a people and process perspective, it's not going to work. So then they go, I've only got the million dollars. The answer is you buy $200,000 worth of technology and use $800,000 to make sure people use it. It doesn't do any good to have good data if you don't have people that know how to use that good data in order to turn it into something that's valuable for the organization.
Let me give you one more example, just to give you a range of things.
Peter Aiken (21:49) Always in class we like to bring in some real-world application for the students so that they get some practice. So here's an agency of the state government and it's concerned with child protective services. Basically, they get a call, they run to a house. At the house, they're gonna try and give somebody a questionnaire — it's about 80 questions — to see whether one of three outcomes occur. One, there's a problem with the kid and the kid needs to be taken away and put in protective custody. Two, there's a misunderstanding, and we're going to give everybody a ticket and tell them to go appear in front of the judge in the next week. Or three, there was no problem at all, and it was literally a swatting-type activity.
Well, the 80 questions that were there were put together when the program had been started at least 10 years ago. And we took that data into the classroom and analyzed it and found out that half of the questions had no predictive value on the outcome of the safety of the child, which is of course the paramount concern. My wife calls it the primary sort criteria — you want to make sure you get that piece right.
But by eliminating half the questions, the agency in question was able to reduce the amount of time to do this questionnaire from an hour to a half hour. Now that doesn't sound like a lot, but you repeat it enough times with enough interviews, and the agency was able to show at the end of the next year that they had been able to take a million dollars and transfer it from what had previously been administrative overhead — I'm asking you questions that we really don't care about the answer to — to actually delivering child safety activities. These are the things that make the real difference in the data community. And we've got thousands and thousands of people around the world that are dedicated to doing this, just as you're dedicated to helping share their stories out there, Matthew.
Matthew Stibbe (23:39) How — I'm really enjoying this 80/20 rule in terms of a resource allocation guideline. What other heuristics or guidelines or things have proven to work over time that people could apply?
Peter Aiken (23:58) Well, that's a really interesting question because when you are trying to make a fundamental change in things, you have to go to the general. And my thought on this — the inspiration was, I don't know how much you are in software, but somewhere many moons ago somebody put out something called agilemanifesto.org. You can go and look it up. It's four short sentences. And it says these are some principles that we've put together, and we think that by following these principles, we'll produce higher quality software, faster. And truly, in the last 50 years we've been doing this, there hasn't been a development that has made more of a difference.
So bring that to data. I said, I wonder what that would look like from a data perspective. And a couple of us put our heads together and came up with something called the Data Doctrine. And that's just that we value certain things more in the organizations. And it talks about what the organizational changes are that need to be made. Again, you said it yourself — it's a fairly simple change in behavior.
But if we do think — I'll give you one that I use all the time. I'm old. When I was taught to drive, I was taught to defensively drive. It's a type of driving that you're taught. And for some reason it's not taught anymore. When I talk to young people about it, about one in three gets a defensive driving course and the others just get it. But as part of my defensive driving course, I always put my seatbelt on when I get in the car. Just right. I sit down and do it. That type of automatic behavior is what we have to start working towards.
Letting people know that the data components of our systems are the most stable, the things that last the longest, the things that change the least, evolve the least. They all evolve and change, right? There's no question of that. But in order to focus on the most stable ones, that's where we're really going to be able to build things.
And I'm leading up to a topic here, Matthew, which I think is hilarious because our Office of Management and Budget at the federal level — we've had this thing called DOGE going on, you may have heard of it. I got "DOGEd" myself personally on the process.
Peter Aiken (26:15) They've put out an announcement last April that said they're gonna put together a gigantic system that is going to hold all of the HR systems for the federal government. More than a hundred of them are going to be consolidated into this. Now, go back to what I was just relating a little while ago. The reason we had 37 systems that grew up to pay people was because we needed a system somewhere in the world that paid people who were under rotating helicopter blades in waist-deep water, having only one leg and working overtime. Just weird combinations of that.
If you create one system to handle all of the HR government functions in the entire thing, you will have the biggest HR system you've ever thought about in the entire world. And the chances of it working correctly are zero. This is going to be a boondoggle that was given on a no-bid contract to Oracle. And Oracle, as we know, is in a little bit of trouble right at the moment as it is. So this may be one of the other things that kind of helps them circle the drain and get kicked out of the investment-grade behaviors. By the way, that's the economics of the data centers part that is really a frustrating piece.
Matthew Stibbe (27:15) It's at a very macro, gigantic level a problem that I think most organizations can relate to. One of the things you said earlier was that 80 percent of organizational data is dead, obsolete and trivial. To what extent is that part of the problem and to what extent is that a problem in itself?
Peter Aiken (27:39) That's a really great question. So if we go back to the data center piece, if we think that — and our measurements show us that it's a minimum of 80 in all organizations. Every single one of them I've gone to. And the only thing I get at all is that some people come to you and go, actually here at this organization, it's 85 percent or even — I had somebody tell me it was 90 percent redundant, obsolete, or trivial.
And that's, you know, why are we storing things? Back there on the shelf, it's kind of hard to see, but there's a bunch of old dead hard drives that I'm getting ready to have crushed up into little tiny pieces because it's got a bunch of data on it that nobody should have, but there's no real good way to get rid of it, even if you overwrite it 35 times. Nobody's got the connections to make these things, except the bad guys who can get anything they want.
Matthew Stibbe (28:29) What kind of strategies are available to us — society-wide and within an organization — to tackle that problem, to crunch those old hard drives that don't need to exist but nobody wants to get rid of?
Peter Aiken (28:48) Well, it's a cleanup problem, right? Look behind me, the mess I've got. Those of you that are listening unfortunately can't see it, but it's a bookshelf full of books. And literally my project this summer is to continue to try and find ones that I either need to give away to students or donate to a library or just simply throw away because there's just no need. At sixty-seven years old, I'm not gonna start rereading these books. And managing data is the same way.
It's just too easy to make a copy of a file because it's no problem. We can make it, but then we've got two copies of it on our hard drive. We're trying to figure out which one was the most recent version. And then we end up with versioning and things like that. Thank goodness Windows and the Mac have gone to versioning of the documents, so we don't actually have to maintain our own versions anymore.
But yes, there's so much absolute redundant stuff. And most organizations are doing what we call forklifting it up. So they grab their data, get it all in a forklift, they turn around and they drop it in the cloud, and Amazon or Oracle or whoever gets rich because they're storing stuff that they don't need to store.
I'm telling organizations the biggest source of savings for them right now is to go take a look at their cloud bills and see what sorts of things that they're paying for. I've found organizations that are paying to run the same data up and down from the cloud and back and forth and back and forth. It's like, if you're gonna do it, bring it down, work on it, and then put it back up at the end of it. That's literally hundreds of thousands of dollars.
Matthew Stibbe (30:20) Yeah, and I've been writing and marketing in IT for long enough to remember everyone saying, well, the cloud's going to reduce your costs and make everything cheaper. But I think it's allowed people to be a bit lazy as well. Just easier to shovel it or forklift it off into the cloud and leave it there and go, well, we've taken care of the archiving and it exists. Nobody's made the difficult decision to delete anything.
If someone is listening to this and they're a data engineer or a CDO in a company and they want to make the argument to go through and tidy up the data — crunch out that 80 percent — what's the best argument that lands with the CFO for doing that?
Peter Aiken (31:08) Dollars, right? Pounds, currency. Something that's in there. The CFO is a pretty sophisticated individual. Let me give you just a couple of statistics. The average CDO, chief data officer, is in the job about 18 months. The chief information officer is there about four years. The chief financial officer is there about twelve years. So they've got a lot more time in the job to be able to do things. They're very sophisticated.
I had a company I was working with in the Midwest at one point, and they were spending six million dollars a year fixing a bunch of data whack-a-mole things that they were going around to fix. And we showed them that by fixing it at the source, they would not have to spend the six million dollars a year, every year. And more importantly, they didn't really like that initial plan. So I went to the CFO, who's a very sophisticated lady, who said, well, look, if you can solve that problem and this, this, and this — and she was just able to reel the numbers off in her head — and said, we're doing it that way, right? We're not gonna play whack-a-mole anymore. It's just crazy to do that.
Now the really interesting part of that — where you get into what actually happened — they were moving off the mainframe. And the data in the mainframe was bad, right? Well, the data in the mainframe's in exactly one place. The minute you move it out to other places, now you've got a big problem. So the inflection point of transferring it to the cloud is really where you should put the governance in place to make sure that data that goes into the cloud is of known characteristics. That's my mantra. Don't put it out there if you don't know anything about it, because it could be a bomb.
Matthew Stibbe (32:48) And you've replicated, duplicated, and multiplied your problem rather than solving it. So, you were at a CDO conference, I think the last time we spoke, or just going to one, and you were making an observation about AI maturity and the sense of the industry. How mature is the industry on AI right now, in your view?
Peter Aiken (33:15) Unfortunately — well, let's start with the good news, right? There are some companies that are doing incredibly good things. I'll give you an example. A colleague of mine does business in all fifty states and a bunch of different places around the world. And each time he does business in them, he needs to do an NDA, a non-disclosure agreement. He had a lawyer who would sit down and put the standard one that they used and the other one up and they'd read back and forth and do redlining.
He wrote an AI to do this so that now all the guy has to do is drop it into this piece. And it comes up and it doesn't rewrite the agreement, but it comes up with the redlining and says, here's what's problematic in this state, here's how we've addressed it in the past. And the lawyer's like, bring it on. This is absolutely super useful stuff that we can put in place and be able to use.
Matthew Stibbe (34:13) If you were starting out in this data engineering world today, what advice would you give yourself? What does your career look like for the next 35 years? How do you make sure that you're building future-proof in yourself?
Peter Aiken (34:37) I'd have started measuring things earlier. Even though I've had a great career and feel fairly comfortable in terms of what we're doing, if I could have started this process in the early seventies instead of the late seventies, we'd be further and we'd have gotten there faster.
Matthew Stibbe (34:57) I think "start in the seventies and buy Apple stock" would be my time-traveling advice. Well, Peter, it's been an absolute delight talking to you and also meeting Percy the Data Cat. If you're listening to this and you'd like to find Peter online, anythingawesome.com is a good place to start, the Dataversity Data-Ed webinar series, of course his books available at all good bookshops, and DAMA.org — D-A-M-A dot org. Peter, thank you so much for being on the show today. It's been great talking to you.
Peter Aiken (35:33) Matthew, a pleasure. I looked forward to it.
Matthew Stibbe (35:36) And that brings this episode to a close. If you'd like to get more practical data insights or learn more about CloverDX, please visit cloverdx.com/behind-the-data. Thank you for listening and goodbye.
Resources
- DAMA International — the global professional association for data management, which publishes the DAMA Data Management Body of Knowledge (DMBOK)
- Anything Awesome — Peter Aiken's consultancy (previously Data Blueprint)
- Agile Manifesto — the four-sentence software development philosophy that inspired Peter's Data Doctrine
- Randy Bean's Executive Leadership data survey — the long-running annual survey on data and AI leadership that Peter references for the 80/20 people-vs-technology split
- Dataversity Data-Ed webinar series — where Peter regularly presents on data management topics
Related content
- Data quality with CloverDX
- Blog: How to establish a strong data governance framework
- Webinar: How to effectively migrate data from legacy systems
More from Behind the Data
- The vital importance of data governance - with James Courtney-Smith) Solutions Consultant at Lucid Data Services
- The human impact of data transformation - with Todd Schnirel and Jennifer McCalicher of Airline Hydraulics
