Generative AI promises to make it easier than ever to work with business data.
Instead of relying on predefined dashboards or asking a data team to build another report, users can increasingly ask questions in plain English and receive an answer almost immediately.
But there is a problem. AI can only work effectively with data if the organization understands what that data means in the first place.
James Serra has spent more than 40 years working with data, from programming and database administration to data warehousing and modern cloud architectures. Today, he is a Data & AI Architect at Microsoft and the author of Deciphering Data Architectures.
In a recent episode of Behind the Data, James explored what it really takes to make organizational data ready for AI, why the technology cannot bypass decades-old data management problems, and what happens when businesses mistake a convincing prototype for a production-ready solution.
AI exposes a problem businesses already had
Organizations rarely suffer from a lack of data. The harder problem is often agreeing on what that data actually means.
James regularly sees this emerge during architecture sessions with customers. A business may arrive wanting to build a data warehouse, dashboard or AI assistant, but the technical work quickly exposes disagreements that had previously remained hidden.
“They start arguing over what should be in the report, what the numbers mean, what’s most important,” James says. “They’ve sometimes never had that discussion before.”
A traditional report is usually built around questions the organization has already identified and metrics it has explicitly defined. A conversational AI interface allows users to ask almost anything, but that freedom depends on shared meaning.
Organizations first have to agree on questions such as who owns particular data, how it should be cleaned and what the correct answer actually is. “There’s a lot of stuff you’ve got to get through sometimes to get to that point where you’re even building a solution,” James says.
You need to make data AI-ready
Putting an AI interface over structured business data may look straightforward.
In practice, the model is being asked to interpret information that was often designed for applications and human analysts rather than large language models. That creates new opportunities for misunderstanding.
“The bot will assume a lot of things,” James says. Even something as simple as an acronym can cause problems if the model interprets it differently from the organization.
Preparing data for AI might therefore mean expanding acronyms, clarifying terminology or providing additional instructions that explain how concepts relate to one another.
“If somebody’s talking about a person, they’re really talking about a customer,” James gives as one example of the type of contextual information a model may need.
This is important because an incorrect AI-generated answer can be harder to identify than an incorrect dashboard.
With an established report, the numbers have normally been tested against expected outputs. With a conversational interface, the organization cannot predict every question a user might ask.
“You may not even know that the answer is not correct,” James says, “because nobody’s there to validate it like they would validate a report and a dashboard beforehand.”
The risk is particularly important for business-critical information.
James still sees an important role for traditional dashboards where accuracy cannot be compromised, whether that involves financial reporting or another high-stakes use case.
AI can complement these interfaces by helping users explore information and ask follow-up questions. But users need to understand that the answers require scrutiny.
Generative AI is, by its nature, non-deterministic. “You can ask the same question and get a different answer,” James says. “So you always gotta be aware of that.”
A better interface does not fix bad data
One of the biggest misconceptions around AI is that it can somehow compensate for poor underlying data.
James describes the problem using a familiar principle: “junk in, junk out.”
“If you have junk data and you put an AI bot in front of it, it’s gonna be even worse than reports and dashboards,” he says.
The model can make assumptions, incorrectly join information and still return an answer that sounds entirely plausible. That means organizations need to clean, master, join and aggregate their data before expecting AI to reason reliably over it.
“You have to get data AI-ready,” James says, “and it’s gonna take just as long, if not longer, than if you put reports and dashboards on that.”
New tools can accelerate elements of the work, but they cannot remove the underlying process.
“We have all these great tools and accelerators, but you don’t shortcut the process.”
The problem is that organizations frequently overestimate the quality of the data they already have.
James often hears customers say their data is “pretty clean”. Once the source data is examined in detail, another picture appears.
“You have birth dates that are in the future, or people are 200 years old,” he says. Often the problem originated years earlier when an application required a date and somebody entered anything that would allow them to continue.
The immediate error may look trivial. At scale, these inconsistencies create gaps that have to be understood and corrected before the data can reliably support analytics or AI.
Bad data destroys trust quickly
Data quality is not only a technical issue. It directly affects whether people trust the systems built for them.
James sees inadequate data cleaning as one of the most common reasons data projects fail. Teams underestimate the time required to prepare the information, only for users to discover errors once the finished solution reaches them.
A seemingly simple problem can be enough to undermine the entire project. If an end user opens a new report and immediately asks, “Why is this person here twice?”, the damage extends beyond that individual error.
“You lost their trust right from the get-go,” James says. “And to get it back is nearly impossible.”
The answer is to spend more time before launch cleaning the data and working with end users to validate it.
Involving those users also makes them part of the process. They can contribute their domain knowledge, help identify discrepancies and better understand why certain decisions were made.
AI makes this challenge even more acute. A report normally restricts users to predefined views. A conversational interface might allow them to ask hundreds of different questions.
An AI system can then deliver the wrong answer with complete confidence.
As James puts it, it can be “confidently wrong”.
The user may not even realize there is a problem. “It sounds right,” James says, while the person receiving the answer may have no obvious way to validate it.
That makes rigorous preparation and testing essential before an organization exposes business data through AI.
AI-ready data may require a new layer
Traditional data architectures often focus on moving raw information through a series of stages until it becomes ready for reporting and analysis. Yet James believes AI introduces an additional requirement.
He describes conventional architectures as moving toward a “gold” layer in which data becomes presentation-ready. That may be sufficient for reports and dashboards. But AI may require another step.
“I would go from presentation-ready to take that data and make it AI-ready,” James says.
That additional layer might expand acronyms, clarify terminology, add context or provide instructions about how different parts of the data should be interpreted.
Reports may not need these changes because their logic has already been explicitly defined. AI systems need enough context to interpret questions and work out what information the user is referring to.
For James, this is significant enough to think of AI readiness as another stage of data maturity.
“There’s a whole other layer,” he says. “Your solution’s going to take longer because now you have this fifth layer.”
AI may therefore increase the amount of data preparation organizations need to do, even while making the finished data easier for people to access.
Self-service still needs governance
Giving business users direct access to data has been a goal of analytics teams for decades.
Modern semantic layers and business intelligence tools have made genuine self-service much more achievable. IT teams can define models and relationships behind the scenes, allowing users to build reports without needing to understand every underlying table and join.
Generative AI pushes that idea further. Instead of dragging fields onto a dashboard, a user can simply ask a tool to create the report.
But the underlying question remains the same. Has the system interpreted the data correctly?
“It’ll build them, but does it understand correctly?” James asks.
A report can look convincing while being based on the wrong relationships between data.
That means the ultimate goal should still be self-service, but not self-service without oversight.
“IT’s gotta monitor what the end users are doing,” James says. If somebody creates a report or receives an AI-generated answer that is wrong, there need to be systems capable of identifying the problem.
“You have to have these checks and balances, this auditing in place,” he says.
Many organizations still fall short here. They either give users powerful tools and assume the output will be correct, or attempt to restrict the technology while employees find their own ways to use it anyway.
Effective self-service therefore depends on a combination of accessibility and governance.
The easier it becomes to generate analysis, the more important it becomes to understand how that analysis was produced.
There is still no shortcut
After more than four decades in technology, James has seen architectures, platforms and terminology change repeatedly. The underlying principles have proved much more persistent.
What he would most like organizations to stop doing is assuming that a new technology can bypass the work required to build reliable data systems.
“There’s no shortcut process,” James says.
AI has accelerated development, but “you still gotta go through all the steps.”
Cleaning is one of the clearest examples.
“Don’t go to AI and just say, ‘Clean all my data,’ and think it’s gonna be done with it.”
In some cases, modern data projects may actually involve more work because organizations are giving users more ways to interact with information. Reports, dashboards, self-service tools and conversational AI can all sit on top of the same underlying data. The end result can be far more useful, but only if the foundation is trustworthy.
To hear more from James Serra about AI-ready data, modern data architectures and the lessons he has learned across more than 40 years in the industry, listen to the full episode of Behind the Data.
By CloverDX
CloverDX is a comprehensive data integration platform that enables organizations to build robust, engineering-led, ETL pipelines, automate data workflows, and manage enterprise data operations.

