News
August 12, 2026

The Challenges of Data Cleansing for AI

Published By

Joe Franco

Continuing the theme of preparing and cleansing data before transferring it for AI purposes, organizations need to recognize the significant challenges involved. Data collected from multiple systems creates inherent complexities that must be addressed before the information can become a reliable foundation for machine learning and generative AI.

Some of the most common challenges include:

  • Inconsistent data formats across different systems
  • Missing or incomplete data
  • Incorrect or outdated information
  • Unclear or conflicting data definitions
  • Unstructured data contained in documents, notes, contracts, and other sources

Addressing these issues requires more than a technical data-cleansing process. Data governance becomes critical. Organizations must establish who owns the data, who is responsible for maintaining it, and—perhaps most importantly—who determines what data is correct when different systems contain conflicting information.

This becomes an even greater challenge when dealing with legacy systems that have been in use for decades. Over time, these systems accumulate historical data, workarounds, inconsistent definitions, duplicate records, and information that may no longer meet today's business or technology requirements.

For large finance companies, the challenge can be enormous. Portfolios containing billions of dollars in assets, millions of contracts, and multiple assets associated with individual transactions represent an extraordinary amount of data to evaluate, cleanse, reconcile, and prepare for AI applications.

Unfortunately, some companies may choose to ignore the problem because the effort and cost of addressing decades of legacy data can appear overwhelming. However, I believe there will come a point when these organizations can no longer avoid it.

As AI and modern digital technologies continue to transform the financial services marketplace, clean, accurate, accessible data will become a competitive necessity—not simply an IT initiative. Companies that invest in preparing their data today will be better positioned to take advantage of machine learning, generative AI, automation, and advanced analytics tomorrow.