Sprggun

Mobile Navigation Menu

There is no AI without data.

3 August 2026

Emma Di iorio | CEO & Co-founder, Spriggun

Between 80 and 95 percent of enterprise AI investments are producing no measurable return on investment. Gartner, RAND, and MIT have each measured this differently and converged on the same picture: the money is going in, and the returns are not following it. Boards have increased AI spend against those numbers for two years, and the pattern has not shifted.

The honest reading is not that the AI is bad. It is that AI is its data; the quality of what comes out directly reflects the quality of what went in, at training, at fine-tuning, at every point the system touches data. Most businesses are running AI on data that was never designed for the questions the AI is being asked, and the return on investment failure follows from that directly.

What businesses think they are dealing with

In a recent board conversation I was asked whether data even needs to be cleaned or governed, given AI. The question reveals the misunderstanding at the centre of much of the failed investment. The assumption is that AI is a smart layer, something the business puts on top of its data that works despite whatever is underneath. However, AI does not fix bad data; it scales it. Where the data does not support a confident answer, the model produces one anyway, and the hallucination toll, the cost of reviewing and correcting those false outputs, can wipe out the savings the AI was meant to deliver. In many businesses, because the organisation does not know what data it can lawfully use, a privacy and IP lockdown keeps high-value data away from AI systems entirely.

Training data determines what a model has seen and can produce. Fine-tuning data determines what it privileges when choosing between possible answers. Retrieval and agentic systems act on whatever data is made available at runtime. The model does not reason beyond what it has been given; it reflects it. A business running AI on undocumented data is running AI whose behaviour it cannot fully explain, even to itself.

Data sits in silos across legacy systems never designed to talk to each other, let alone feed an AI model. Different systems define the same thing differently, leading to duplicate records and unclear provenance. Data lineage is incomplete at best, meaning many organisations cannot trace a single data field from its source to the AI model it powers.

AI does not just need clean data; it needs data that carries the specific context of the business: how the business defines its own terms; what decisions created the gaps in the data; and what internal policies govern its use. Being able to say what the business holds, where it came from, what it means in the context of the business, and what may lawfully be done with it, on demand, is what Spriggun calls Know Your Data and Know Your AI: KYD and KYAI. The discipline is not new. I saw it working in banking a decade ago, advising on machine learning models at a time when few sectors were thinking about AI governance at all. Banking is not perfect; it has failures of its own. However, it does show the discipline is achievable when the stakes demand it. Most other sectors are being asked to build it in months, on data estates never designed for it.

The 2026 figures tell the story: only 5 to 7 percent of enterprises say their data is ready to support AI at scale, despite near-universal investment, and over half expect AI experiments to reach production within six months while acknowledging the underlying data work will take over a year. They know the foundations are not there and are pressing ahead regardless. The advice I have given most consistently to boards is this: go slow to go fast, though it seems counter-intuitive. Do the data work once, properly, and every use case that follows moves faster because the foundation is already there.

Three challenges, one root

The return on investment failure has three dimensions, and all three channel through the data.

The first is not knowing the data well enough. It extends further than most businesses have grasped. AI does not just consume data; it creates it. Every time someone uses an AI tool, the outputs become new data the organisation may not be governing, may not be able to trace, and may not know exists.

AI does not need to understand or intentionally construct protected characteristics to discriminate; it simply surfaces the mathematical patterns hidden deep within proxy data like postcodes, browsing behaviour, or device type. The Dutch child benefits scandal is the reference case: a system that used postal codes to flag fraud risk, producing outcomes that were discriminatory on racial lines through inference from the data it was given. The consequences land on individuals; discriminatory credit decisions, wrongful benefit denials, biased hiring screens. These are already actionable under Article 22 of the GDPR and under indirect discrimination provisions in equalities law, and the AI Act catches this directly. These obligations apply across jurisdictions and domains; the data-fitness discipline has to hold globally.

Agentic AI moves businesses from content risk to conduct risk. Autonomous agents consume and generate data at volumes no human team can audit, often using over-scoped, inherited credentials that give access to data broader and more sensitive than anything originally scoped. Knowing your data means knowing all of it, end to end: inputs, inferences, outputs, what happens after deployment, and whether the same can be said of vendors and partners in the chain. Getting this right requires effective top-down governance and genuine fluency across the organisation; not basic AI literacy, but the capability to recognise when outputs are not to be trusted. Data drift adds a further dimension: the data changes over time, the model's accuracy degrades, and bias testing done at launch does not catch what develops later. KYD and KYAI is an ongoing discipline, not a point-in-time exercise.

The second is not knowing whether the business challenge is one AI can address, given the data available. Some of what is being solved with AI is not an AI challenge; it is a management, product, or customer-experience challenge that has been sitting in the business for years, and reaching for AI often makes the underlying issue more visible rather than resolving it. Air Canada was held liable for a refund policy its chatbot invented; the technology did not create the problem, but it did make it harder to contain. Where processes require predictability and certainty of outcome, regulatory compliance, financial reporting, payroll, AI's variability is not a feature; it is a risk. There is a related fallacy in treating a given accuracy rate as good enough without asking what the error rate means in practice; in credit decisioning or benefits eligibility, even a small error rate means real harm at scale.

There is also a question most non-tech-forward businesses have not sat with honestly enough: whether they have the data maturity and capability to deploy AI at scale. Building in-house rarely works where the organisation was not built for the pace of change AI development demands; buying and partnering through providers set up for it is usually the more honest solution. The gap between a successful pilot and live deployment on real data at company scale is where much of the value is lost.

The third is not understanding the data-AI relationship well enough to lead. The misconception this piece opened on, that AI works despite its data, is not only a frontline issue; it sits at the top. Executives authorising AI spend without understanding the data underneath it are committing capital on a foundation the board has not assessed. When leadership does not understand what the data can support, the wrong challenges get picked for AI in the first place, and the right ones get overlooked. The work drifts to whoever will accept it, legal, IT, a head of AI not embedded in the business, and in the process moves away from the people who understand the business problem well enough to know whether AI is the right instrument for it. The other factors the research cites as causes of return on investment failure, workflow redesign not done, horizontal deployment rather than vertical, leadership disagreement on measurement, are downstream of this same gap. Until that changes, the pattern will repeat.

Questions the board needs answered before the next AI budget call

  • Is the data ready? Has it been demonstrated that the data underneath this investment is understood, fit for purpose, and lawfully usable? If the answer takes weeks to produce, neither is ready.

  • Is AI the right instrument for this challenge? Given the data the business has and the state it is in, is AI the right solution, or is this a broader business problem that AI will not resolve?

  • Is there effective AI governance in place, including for this use case? Is there governance with named executive ownership, covering what happens with the data now and what follows after deployment; new data the AI creates, what it may infer, and what happens across the vendor chain?

  • Is the data work costed and owned? Given the current state of the data, what realistically needs to happen before this investment can deliver, and is that work costed, scheduled, and owned at executive level with team buy-in?

  • Does the board have the fluency to assess the answers? Does the board understand the relationship between AI and data well enough to know whether the answers it is hearing are adequate? If not, what is the plan to get there?

Return on investment comes back to the data. It always does. The businesses that will see returns are the ones that have done the work: knowing what they hold, whether it supports what they are trying to do, and who owns the answer. That is not a governance platform or a policy library; it is knowing your own business well enough to know what the AI will actually do. The same questions apply to every AI investment the board considers, because the data underneath does not change between one proposal and the next.


Emma Di Iorio is Co-Founder and CEO of Spriggun, a UK-based RegTech and AI governance advisory firm. A qualified solicitor with senior in-house and advisory experience, she writes and speaks internationally on AI, data, privacy, and web compliance.