The Hidden Bottleneck of AI Isn’t Models — It’s Real-World Data

I Work FOR YOU, Not Factories.

I’m Leon Xu, based in Shenzhen, China — 15+ years across hardware engineering, embedded systems, product management, and supply chain execution for consumer electronics. I work for overseas teams who need a grounded China-side operator to help turn their technology visions into manufacturable, scalable products. This article is about something I have been watching quietly from the factory floor: a shift in where AI’s real constraints actually live.

The Hidden Bottleneck of AI Isn’t Models — It’s Real-World Data

By Leon Xu | Easelink Tech | Shenzhen, China | May 29, 2026

Everyone Is Looking At The Wrong Constraint

If you follow the AI discourse at all, you know the script. The next breakthrough is about model architecture. Or compute scaling. Or chip efficiency. Or inference optimization at the edge. These are real conversations and real constraints, and smart people are working on them.

But from where I sit — embedded in Shenzhen’s manufacturing ecosystem, dealing daily with engineering workflows, factory systems, supplier data, and production processes — I keep noticing something that almost nobody in the mainstream AI conversation is talking about directly: the actual next bottleneck for AI is not models. It is data. Specifically, it is the kind of data that cannot be scraped from the internet.

The first generation of large AI systems was trained on the internet. Text, images, code, structured documents, Wikipedia, GitHub, web crawls. This worked because the internet is a massive, dense, relatively accessible repository of human-generated knowledge. It also set a mental model in most people’s minds: AI scales by ingesting more data, and data is abundant.

That assumption is quietly breaking down. And the industries where it is breaking down fastest are exactly the ones that matter most for real-world AI capability — engineering, manufacturing, robotics, industrial operations, and physical-world workflow automation.

Why Web Data Is No Longer Enough

There is a category of knowledge that has never lived on the internet and probably never will. Not because it is secret — though some of it is — but because it was never created in a form that gets posted, indexed, and scraped.

Consider what a machinist knows. Not the textbook explanation of CNC operation, which is absolutely on the internet. The specific way this machine, in this factory, with this tooling setup, needs to be approached to hold tolerance on a particular aluminum alloy at a specific cut depth. The experienced judgment about which settings to adjust when a surface finish starts drifting. The workflow pattern — step one, check this, then adjust that, then verify the other thing — that took years to develop and lives entirely in operational practice.

Or consider what a mechanical design engineer knows after fifteen years of building products through full manufacturing cycles. Not the CAD manual, which is thoroughly documented. The intuition about which feature tree structure will survive downstream manufacturing changes. The knowledge of which geometric tolerances are achievable by which process families in which supplier ecosystems. The implicit design-for-manufacturing awareness that distinguishes a product that can be built in volume from one that works beautifully as a prototype and breaks the factory.

None of this is on the internet in any form that an AI model can currently consume. It exists in factory operating procedures, in engineering revision histories, in the embodied practice of skilled operators, and in the complex workflows of industrial production. Getting access to it requires being inside the ecosystem — not browsing it from the outside.

This is the constraint that AI’s second generation is going to run into head-first. And I think most of the people currently building AI systems significantly underestimate how hard this problem actually is.

The Rise of Physical-World AI

Let me map the territory, because the scope of this problem is larger than any single industry.

Look at where the most commercially significant AI applications are actually heading. Not chatbots — those are mostly solved, or at least commoditizing fast. The interesting applications are all fundamentally physical: AI systems that help engineers design products that can be manufactured. AI systems that operate robots on factory floors. AI systems that simulate garment fit on digital bodies before a single physical sample is cut. AI systems that predict machine failure before it happens. AI systems that optimize logistics through warehouses and supply chains. AI systems that assist surgeons or help develop medical devices.

Every single one of these categories is fundamentally dependent on data that describes how physical systems behave — data that mostly lives inside the operational environments where those physical systems exist. And access to that data is not a technical problem. It is a trust, relationship, and operational coordination problem.

This is a significant shift from the first generation of AI infrastructure. Building a web crawl is an engineering problem. Getting a tier-one automotive manufacturer to let you instrument their assembly line and capture operator workflow data is a completely different kind of challenge. One requires servers and bandwidth. The other requires years of relationship building, ironclad data governance agreements, organizational buy-in across multiple levels of a complex institution, and usually some form of direct operational value exchange.

Why Real-World Workflow Data Is Fundamentally Different

There are several dimensions to this difficulty that are worth unpacking carefully, because they interact in ways that make the problem harder than it first appears.

The data is not legible without domain context. A CAD file without understanding the engineering intent behind it is just a collection of geometric entities. A CNC machine log without knowledge of the workpiece material, tooling geometry, and target specification is just numbers. A sequence of robot arm movements without understanding the assembly workflow it is executing is meaningless. Real-world operational data almost always requires deep domain expertise to annotate, clean, and make usable for AI training. This is not a problem you solve by hiring more data labelers — it requires people with genuine operational expertise who are expensive, scarce, and not interested in doing annotation work for below-market wages.

The data is fragmented across incompatible systems. A typical manufacturing operation might have a different ERP system, MES, quality management platform, machine telemetry stack, CAD environment, and engineering document management system — all deployed at different times, from different vendors, with different data models, and with varying degrees of API accessibility. Assembling a coherent workflow dataset from these sources is an integration project that can easily take years and cost more than the AI system it is meant to enable.

The data contains proprietary operational knowledge that companies protect actively. Your manufacturing process data is your competitive advantage. The specific workflow optimizations you have developed over twenty years of production are not something you share with an AI company’s data acquisition team, regardless of the NDA in front of you. This is not paranoia — it is rational competitive behavior. The result is that the most valuable operational data is the hardest to access.

The data changes constantly. Unlike text on the internet, which is relatively static once published, operational workflow data is live. The optimal assembly sequence for a product changes when you get a new generation of components. The toolpath that worked last quarter needs adjustment when the tooling wears. The fitting behavior of a garment changes when the fabric supplier changes. Building AI systems on this kind of data requires not just initial acquisition but ongoing operational integration — a fundamentally different and more expensive infrastructure commitment.

Four Categories Where This Problem Is Already Critical

Engineering and CAD AI

The ambition of engineering AI is real and large. Generative CAD systems that can propose manufacturable geometries. Engineering copilots that understand design intent and flag manufacturability risks in real time. AI systems that can read a SolidWorks feature tree and suggest design-for-manufacturing improvements before the part ever goes to a factory. Intelligent drawing interpretation that converts legacy 2D documentation into parametric 3D models.

Every one of these applications requires training data that almost does not exist in public form. SolidWorks files are not shared publicly. Engineering drawings contain proprietary specifications that companies protect carefully. The feature tree choices that distinguish an expert designer from a novice are implicit knowledge that has never been systematically captured. The relationship between CAD geometry and manufacturing outcome — which features cause problems at which tolerances in which processes — is documented nowhere except in the institutional memory of experienced engineers and the informal post-mortems that follow failed prototype batches.

The AI companies trying to build engineering automation tools are discovering that the model architecture challenge is almost secondary to the data acquisition challenge. You can train a capable geometric reasoning system on public 3D datasets. Training a system that genuinely understands manufacturability requires partnering with engineering organizations willing to contribute real production data, which requires trust, value alignment, and usually a direct integration into their existing workflows.

Robotics and Factory-Floor AI

Humanoid robotics is one of the most heavily funded areas in AI right now. And the fundamental challenge that every team in the space is running into, despite significant hardware progress, is training data. Specifically: diverse, high-quality, annotated data of humans performing the physical manipulation tasks that robots need to learn.

Factory-floor motion capture data. Assembly workflow sequences showing skilled operators performing complex pick-and-place, fastening, and inspection tasks. Warehouse robotics operational logs that capture the full diversity of object types, placement configurations, and failure modes that real logistics operations encounter. These datasets do not exist in public repositories in any useful form. Building them requires getting access to factory floors and warehouses, instrumenting the environment, capturing operator behavior over long periods, and doing the difficult annotation work to make that captured behavior useful for training.

This is why you see so much emphasis on simulation in robotics AI — not because simulation is adequate, but because real-world data acquisition is so hard that simulation is used as a substitute until the real data access problem can be solved. The companies that solve the real-world data access problem at scale — that build genuine relationships with factory operators, logistics companies, and industrial manufacturers who allow systematic workflow capture — will have training advantages that cannot be replicated by better simulation alone.

Virtual Try-On and Fashion AI

This is a category that deserves more attention than it usually gets in the AI conversation, partly because fashion seems less serious than industrial robotics, and partly because the underlying data problem is genuinely underappreciated.

Building a realistic virtual try-on system — one that can accurately predict how a specific garment will drape, fit, and move on a specific body — requires training data that combines garment construction information (digital patterns, material properties, construction specifications) with body measurement data and real-world fitting outcomes. This data is distributed across fashion design software systems, garment manufacturers’ production databases, fitting room measurement systems, and the institutional knowledge of pattern makers and fit specialists.

Fashion CAD files from systems like Browzwear, CLO, or Optitex are not publicly available. Apparel brands’ fitting data — the measurements, the adjustment records, the outcome data from fit sessions — is proprietary and closely held. The knowledge of a skilled pattern maker about how different fabric compositions behave differently in the same cut is nowhere in any dataset. Getting access to this data requires deep integration into fashion industry workflows, which means understanding the industry’s specific data formats, operational rhythms, and organizational dynamics.

The virtual try-on AI companies that succeed will not be the ones with the best neural architecture. They will be the ones that figure out how to acquire high-quality garment fitting data at scale, which is fundamentally an operational and relationship problem.

Industrial AI and Manufacturing Process Intelligence

CNC machining workflows. Tooling process optimization. Factory quality assurance systems with decades of inspection data. DFM review processes and the institutional knowledge behind which design choices cause manufacturing failures. Maintenance procedure documentation and the failure pattern data that makes predictive maintenance actually work. Supply chain operational data capturing supplier reliability, lead time variation, and component quality across years of production history.

All of this data exists. None of it is publicly accessible in a form useful for AI training. It lives inside manufacturing organizations, protected by competitive concern, operational complexity, and the simple fact that nobody has built the infrastructure to make it available in a standardized, AI-consumable form.

From my position in Shenzhen, I deal with fragments of this data problem constantly. When a team asks me to help evaluate a new component supplier, I am drawing on a mental model built from years of supplier interaction that has never been formally captured anywhere. When we review a product design for manufacturability, the issues we flag are based on pattern recognition from hundreds of past projects — pattern recognition that lives in my head and in the heads of the engineers I work with, not in any training dataset.

The Hidden Trust Problem Behind Industrial Data

Here is the piece of this puzzle that I think gets the least attention from people building AI systems: even when industrial data technically exists and is technically accessible, the trust infrastructure required to actually use it is largely absent.

Consider what it would take to build a comprehensive AI training dataset from manufacturing operational data. You need factory operators to allow you to instrument their production environment. You need engineering organizations to share their CAD workflows and design history. You need suppliers to expose their process data. You need quality teams to share their inspection records and failure data.

Each of these relationships requires a different form of trust. Factory operators need to trust that workflow capture will not be used to automate them out of jobs. Engineering organizations need to trust that their proprietary design patterns will not end up in competitors’ hands. Suppliers need to trust that process data will not be used to squeeze them on price. Quality teams need to trust that failure data will not be used to assign blame.

These trust barriers are not theoretical. They are the reason that most AI companies trying to build industrial applications find that the data acquisition problem is ten times harder than they expected when they wrote their initial pitch deck. The technology for capturing and processing this data is largely solved. The human and organizational dynamics of actually getting access to it are not.

The organizations that will solve this — that will build the trust infrastructure enabling systematic real-world workflow data access — will likely not look like typical AI companies. They will look like operational partners who have been inside these ecosystems long enough to understand the implicit rules, the legitimate concerns, and the value exchanges that make data sharing feel worthwhile for all parties.

Why Data Sourcing Is Becoming Infrastructure

There is a useful analogy here to cloud infrastructure. In 2005, every company that wanted to run web services had to build its own server infrastructure. The capital requirements, operational complexity, and expertise demands were high enough that this was a significant barrier. Then AWS demonstrated that you could abstract the infrastructure layer and provide compute as a service, and the entire economics of building software products shifted dramatically.

Something similar is happening with real-world operational data, except the infrastructure problem is harder because data access is fundamentally a relationship problem, not just an engineering problem. You cannot provision real-world workflow data by spinning up instances. You have to build the operational relationships, the data governance frameworks, the annotation pipelines, and the quality validation systems that transform raw operational activity into AI-consumable training data.

This is becoming an infrastructure investment, not just a project. Companies that build genuine, sustained access to specific categories of real-world operational data will have training advantages that compound over time. A robotics company with five years of continuous access to a major logistics operator’s warehouse workflow data will train systems that a company without that access simply cannot replicate — regardless of how good their model architecture is.

This dynamic creates a new category of competitive moat that is genuinely different from the model-based moats that have defined the first generation of AI competition. Model weights can be distilled, fine-tuned, or eventually commoditized. A deep operational relationship with a major manufacturing ecosystem, built over years, with proprietary workflow capture infrastructure in place, is much harder to replicate quickly.

Why China’s Manufacturing Ecosystem Is Strategically Significant Here

I want to be direct about something that I think gets lost when this topic is discussed from a Western perspective: China’s manufacturing ecosystem — and Shenzhen’s specifically — is probably the largest and most accessible concentration of the exact kind of real-world operational data that AI systems need.

The density of manufacturing processes, engineering workflows, supplier relationships, production data, and operational expertise in Shenzhen is extraordinary. Multiple complete supply chains for virtually every consumer electronics category, operating continuously, generating operational data at massive scale. Engineering teams working on hundreds of product categories simultaneously. Suppliers with production experience across materials, processes, and tolerances that span decades of accumulated operational knowledge.

Most of this data has never been formalized, let alone made available for AI training. But the operational relationships required to access it exist — in the supplier ecosystems, in the engineering service companies, in the manufacturing consultants and product development firms that are embedded in this ecosystem as trusted operational partners.

I say this not as a promotional claim about China’s manufacturing position — that argument is well-established and does not need my contribution. I say it because I think the AI companies that are serious about building physical-world AI capabilities will eventually discover that their data acquisition strategy has a significant China component, and the teams that have navigated the access question most successfully will be the ones with genuine operational relationships inside these ecosystems, not the ones parachuting in for a factory tour.

The trust infrastructure that enables real-world data access is built over time, through operational value exchange, not through vendor agreements. From where I sit, helping overseas teams navigate the China-side execution environment, this dynamic is visible every day. The partners who get real operational visibility are the ones who have earned it through sustained, useful participation in the ecosystem — not the ones who showed up with a compelling deck and an NDA.

Why Workflow Networks May Become More Valuable Than Models

Let me push this argument to its logical endpoint, because I think it has implications that go beyond just data acquisition strategy.

If the next major AI bottleneck is access to real-world operational workflow data, then the organizations that hold structural positions inside operational workflow networks — not just access to data, but embedded, trusted, ongoing roles in how workflows actually function — will have a compounding advantage that is fundamentally different from model-based advantages.

A language model trained on internet text can be improved by training a better model on more internet text. The weights can be released, fine-tuned, distilled, and replicated. The competitive dynamic is real but ultimately addressable with compute and engineering talent.

An organization that is genuinely embedded inside a major manufacturing ecosystem — that participates in the engineering review process, that has ongoing access to production workflow data, that has built the trust relationships that enable data sharing — has something that cannot be replicated quickly by a competitor with better model architecture. The operational relationships are the moat. The workflow embeddedness is the moat. The data that flows from genuine operational participation is the moat.

This suggests a strategic posture that most AI companies are not currently taking seriously: investing in operational network position, not just model capability. Finding the workflows where AI can provide genuine value and embedding into those workflows in ways that create sustained data access, rather than treating data acquisition as a one-time project to be completed before training begins.

The internet trained the first generation of AI. Factories, engineering systems, and real-world workflows may train the next generation. But only for the organizations that have done the slow, difficult, relationship-intensive work of earning access to those environments — and only if they understand that the data problem is, at its root, a trust and operational coordination problem that no amount of technical sophistication can shortcut.

Final Thoughts

I want to be clear about what I am not arguing. I am not arguing that model architecture does not matter, or that compute scaling is irrelevant, or that the technical challenges of AI development are solved. They are not.

What I am arguing is that a significant and underappreciated constraint is emerging that is fundamentally different in character from the technical challenges the AI field is well-equipped to address. Real-world operational workflow data is hard to access not because the technology for capturing it is immature, but because accessing it requires operational credibility, trust relationships, and sustained value exchange with the organizations that hold it. These are slow, human-scale processes that do not compress with more investment or better engineering.

The AI companies and teams that recognize this early — that build their strategy around operational network position and workflow data access rather than treating data as a solved problem — will likely have advantages that compound quietly but significantly over the next decade.

From the factory floor, this shift is already visible. The question is whether the broader AI community will notice it before the bottleneck becomes undeniable.

I Work FOR YOU, Not Factories.

But being embedded inside the factory ecosystem is exactly why I can see what most people looking at AI from the outside are missing.

Working On AI Hardware Or Industrial AI Products?

If you are building AI systems that depend on real-world manufacturing, engineering, or physical-world workflow data — or if you are trying to deploy AI capabilities inside China’s manufacturing ecosystem — this is the operational environment I work in every day. I help overseas hardware teams navigate the China-side execution layer: chip selection, supplier relationships, production coordination, and the kind of embedded operational access that takes years to build from scratch.

Reach me at [email protected]. I am not selling data access — but I am happy to discuss what the operational reality of this ecosystem looks like from the inside, and whether it is relevant to what you are building.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top