Product development organizations in industrial sectors are under more pressure than ever. Regulatory requirements keep mounting, sustainability mandates are reshaping materials decisions, and the pressure for speed is relentless.
In this environment, AI arrived with enormous promise. Yet for many organizations, it has not paid off. Pilots get stuck, models trained in one context fall apart in another, and scientists do not trust results that cannot be traced back to something they understand. External surveys have found that most companies struggle to achieve value from AI, especially in the lab or on the production line.
In most cases, the problem is not the algorithms. It is the data underneath them. Any serious AI initiative has to start with an honest look at data quality, structure, and accessibility.
Most organizations find significant gaps. Analysts have warned that a large percentage of AI projects will be abandoned because the underlying data was never made AI‑ready. The highest‑leverage investment is rarely in the latest model. It is in the experimental data that the model will be trained on.
The Data Problem AI Cannot Solve for You
Walk into the R&D or QC function of most industrial companies and you will find data captured in ways that made sense within each department, but do not connect into anything coherent.
Formulation records sit in spreadsheets. Test results live in a LIMS that was set up to track compliance, not to support analysis. Process parameters are buried in equipment logs that nobody ever thought to link to the experiments they were part of. Observations, photos, and customer notes sit on shared drives, inconsistently named and inconsistently filed.
That fragmentation creates a real obstacle for AI. Machine learning needs data in context: formulation compositions connected to process conditions, linked to measured properties, and tied to how the final product performed.
When those connections do not exist, or when someone must spend days rebuilding them by hand before any analysis can happen, the economics of AI stop making sense. These issues are a major reason AI projects in product development fail more often than other IT projects.
Data Structure Is a Strategic Decision
The gap between structured and unstructured data is both a software problem and an organizational problem.
Scientists working on their own, with the tools they have always used, naturally generate unstructured data as the path of least resistance. Spreadsheets, free‑text notebooks, and PDFs with no consistent naming or metadata can be produced quickly, but are nearly impossible to use at scale.
Structured data is different. It requires someone to sit down and design a shared model before data is ever captured: deciding what gets measured and how, which metadata travels with every observation, and how different data types connect to each other.
Take something as basic as Brookfield viscosity, a common property used to measure a fluid’s resistance to flow. Pull records from a typical unstructured lab system and you might see the same property written three different ways: “Viscosity, 7D = 3000,” “BV, ON = 1800,” or “Brookfield Visc. Sp #4 = 5500.”
Are those comparable. Maybe. It depends on multiple factors, including the spindle, the RPM, the temperature, how long the sample was aged, and which instrument was used. In many cases, none of that context was recorded.
A structured data model captures all of this information as separate, searchable fields attached to every measurement. The difference in what you can do with the data afterward is not trivial. It is the difference between data you can use and data you cannot.
Getting to that level of structure is harder than it sounds, for two reasons that often compound each other.
The first is technical. Existing systems such as LIMS, ELNs, and equipment logs were built for specific purposes and were not designed to feed AI. Stitching them together after the fact through data lake initiatives usually gives you marginally better access without real improvement in whether the data means the same thing across systems or locations.
Purpose‑built product development platforms take a different approach. They design the data model around scientific workflows from the ground up so that formulations, process conditions, and test results are linked by default rather than assembled after the fact.
The second problem is cultural, and it is often the more difficult one. Scientists have built their workflows around familiar tools, and when you try to change those workflows, you meet resistance even when the new approach is technically better. Plenty of sound data initiatives have failed in implementation because this piece was not taken seriously enough.
Failure Modes to Avoid
Several failure modes come up repeatedly when organizations try to apply AI without a strong data foundation.
Not enough usable data
A simple model trained on a large dataset almost always outperforms a sophisticated one trained on too little data. When AI goes live before enough useful experiments have been captured in a consistent format, models do not generalize. When they fail, blame lands on the AI rather than on the data shortage that caused the problem.
Applying AI where it does not fit
AI does not work equally well on every type of problem. Push predictive models onto projects with sparse data, poorly defined outcomes, or highly variable results and you will get noise, not answers. Choosing the right targets matters just as much as the modeling work itself. The strongest candidates are areas with a solid history of consistent experiments and measurable outcomes.
Setting the bar in the wrong place
It is tempting to point AI at the hardest problems first, the ones with hundreds of interacting variables and only a few dozen historical data points to draw from. When nothing useful comes out, everyone walks away thinking AI failed. In reality, the expectation was never realistic. Starting with grounded targets lets AI build a real track record before you ask it to do something more difficult.
A Maturity Model for Data‑Driven Product Development
Most organizations follow a recognizable path when it comes to data maturity. It usually starts with paper notebooks, moves to spreadsheets and Word documents, and eventually arrives at some kind of shared digital repository such as SharePoint or an ELN. At that point, the data is at least findable, but it is rarely usable across teams. Running any kind of cross‑functional analysis usually requires a lot of manual cleanup.
The real turning point is when an organization builds a unified data model with scientific context in mind. When formulations, process conditions, and measured properties are captured consistently across teams and sites, something genuinely useful starts to accumulate: institutional knowledge that does not evaporate when someone leaves.
Modern platforms that are purpose‑built for this inflection point give both R&D and QC teams a shared environment where experimental data accumulates in a consistent, connected structure.
With this system in place, old experiments become training data. Patterns that no single scientist could spot on their own begin to surface. Once that foundation is solid, AI stops being a high‑stakes gamble and becomes something more incremental and manageable.
You find the projects with enough historical depth. You validate what the models suggest against what your experienced scientists know. You build AI recommendations into existing workflows. You keep the feedback loops tight so that data quality and model performance improve together over time.
Institutional Knowledge as a Competitive Advantage
There is also an argument for data infrastructure that has nothing to do with AI itself. Industrial organizations hold decades of experimental knowledge. They lose enormous amounts of it to turnover, inconsistent documentation, and siloed files. A well‑designed data architecture changes that equation. It turns tacit knowledge into something an organization can build on.
Over a five‑ to ten‑year horizon, two organizations with equivalent AI tools but different approaches to data will end up in very different places. The one with structured, connected, historically rich data will see stronger model performance, spot innovation opportunities faster, catch problems before they become late‑stage failures, and bring new scientists up to speed more quickly.
The one that tried to run AI on top of fragmented spreadsheets will keep cycling through vendor pilots without getting the returns it expected.
This article was originally published on Technology Networks.

.png)
.png)
.png)