Why AI in Construction Stalls at the Data Layer
CodeBranch Team
Eighty-seven percent of US contractors believe AI will significantly transform their business. Twenty-six percent rate the quality of their current data as high. That gap is not a gap in enthusiasm or in technology — it is the reason most construction AI projects stall between a promising pilot and anything in production.
Quick Summary
- In a late-2025 survey of 235 US general and trade contractors, 87% expected AI to transform their business while only 26% rated their data quality as high.
- Of 23 AI functions evaluated in that survey, 20 were in use by fewer than 15% of respondents — expectation is near-universal, adoption is not.
- A separate 2026 survey of 2,500 industry leaders found integration with legacy systems to be the number one barrier to AI, cited by 50%.
- The bottleneck is not the model. It is that the data a model would need is spread across systems that do not agree, recorded inconsistently, or not recorded at all.
- Which means most construction AI work is integration and data engineering work, and scoping it as anything else is how projects run long.
That framing — the interesting problem sits below the feature, in the layer nobody demos — runs through this whole vertical. More on it in what construction companies build when their platform runs out.
What Do the Numbers Actually Say?
The most specific recent data comes from a Dodge Construction Network and CMiC survey of 235 US general and trade contractors, conducted in late 2025. Three findings from it matter more than the headline.
Expectation is nearly universal. 87% believe AI will significantly transform their business; 85% expect reduced time on repetitive tasks. Among large contractors, 86% believe AI confers competitive advantage.
Adoption is not. Of 23 AI functions evaluated, 20 were used by fewer than 15% of respondents. Only 19% were adapting legacy workflows for AI — which is the step that turns a tool into a capability.
And the constraint is visible in the data. Only 26% rated the quality of their current data as high. Meanwhile 57% cited concerns about output accuracy and reliability as a barrier. Those two findings are the same finding: a model trained or grounded on inconsistent data produces unreliable output, and the people using it correctly stop trusting it.
A 2026 survey of 2,500 industry leaders points at the same place from a different angle: integration with legacy systems was the single most-cited barrier to AI adoption, named by 50% of respondents — ahead of talent, cost, regulation and ROI measurement.
Why Is Construction Data Harder Than It Looks?
Because the information exists; it is just not in a form anything can consume.
A contractor knows what a linear meter of conduit costs. That knowledge lives in an estimator’s judgment, in a spreadsheet from the last similar project, in the accounting system under a cost code that means something slightly different than it did two years ago, and in a supplier email. All four are correct. None of them is queryable, and they do not agree.
This is the structural version of a finding that has held for years: in the last systematic survey of construction technology, 62% of firms were still estimating in spreadsheets and 49% moved data between applications by hand. A spreadsheet is not a data quality problem in itself — plenty of good work happens in spreadsheets. It becomes one the moment you want a system to learn from four hundred of them.
The three failure modes, in increasing order of difficulty:
The data is not in a system. It is in spreadsheets, emails, or people. Fixable with a project, measured in months.
The data is in a system but recorded inconsistently. The same activity is coded differently across projects, or across teams, or before and after someone reorganized the cost code structure. Fixable, but it requires changing how people work, which is measured in quarters.
The data is in a system, consistent, and unreachable. It is in a platform whose export is a CSV a human has to download. This is the one that looks easiest and is often the one that decides whether the project is viable, because it determines whether anything can be automated or someone is running an export every week forever.
Which AI Work Does Not Depend on Clean Data?
This is the distinction that makes a project deliverable in weeks rather than quarters, and it gets collapsed constantly.
Language work tolerates mess. Reading a submittal, summarizing an RFI thread, extracting terms from a contract, classifying an incoming document — these operate on unstructured input because unstructured input is what they were built for. A contractor with genuinely chaotic data can still get value here, immediately.
Prediction and estimation do not. Forecasting a schedule, estimating a cost from history, optimizing a sequence, flagging an anomaly — all of these learn from your past, so they inherit whatever inconsistency your past contains. Run these on unready data and the output is confidently wrong, which is worse than no output.
| Language and document work | Prediction and estimation | |
|---|---|---|
| Input | Unstructured, messy | Historical, structured |
| Data readiness required | Low | High |
| Time to first value | Weeks | Quarters, after data work |
| Failure mode | Visible — a bad summary looks bad | Invisible — a bad forecast looks fine |
| Where it fits | Start here | Build the data layer first |
The failure mode row is the one worth sitting with. A bad summary is obviously bad, so a human catches it. A cost forecast that is 18% low looks exactly like a cost forecast that is right, and nobody catches it until the project is underway. CodeBranch treats that asymmetry as the reason to sequence these differently, not as a detail.
What Does a Readiness Assessment Look For?
Three questions, answered against real data rather than a description of the system.
Does the data exist in a system that can be queried without a person in the loop. Is the same real-world thing recorded the same way across projects, teams and time. And can it be retrieved on a schedule without a manual export.
The answers determine the shape of the project. If all three are yes, the AI work is the AI work. If the third is no, the first phase is integration. If the second is no, the first phase is a conversation about how the organization records things, which is not a software project at all and should not be priced as one.
CodeBranch runs this as a discrete engagement — an AI readiness assessment — precisely because the answer changes the scope by an order of magnitude, and discovering it after a proposal is signed is how projects go badly for both sides.
What Does This Look Like When It Works?
The pattern worth copying is that the model is the small part.
In a quoting platform CodeBranch built for an electrical contractor, an AI recommendation engine suggests components as the user designs a home automation project on a floor plan, and it works — the recommendations are complete and installable. The reason it works is not the model. It is that the dependencies between components were modeled explicitly first: which devices require which controllers, what a given configuration implies for wiring and labor. The AI operates on a structure that describes the domain correctly.
Strip that structure out and the same model would produce plausible suggestions that do not add up to a working installation. The intelligence in the feature comes mostly from the data layer beneath it, and that layer was engineering work, not prompting.
That is the general shape. The AI feature is the visible tenth of a project whose other nine tenths is making the data mean something consistent.
Where Should a Contractor Start?
Start with the thing that is true regardless of the AI decision: get your project data into a system that can be queried.
That is worth doing on its own merits. A contractor who can ask their own history what a square meter of a given assembly has actually cost across the last twenty projects has gained something real, with or without a model attached. It also happens to be the prerequisite for every prediction and estimation capability worth having.
Then run the language and document work in parallel, since it does not depend on that foundation and can deliver value while the foundation is being built.
What CodeBranch would avoid is the sequence that most stalled projects followed: pick an impressive AI use case, build a pilot on hand-curated data, demonstrate it successfully, and then discover that production data cannot support it. The pilot was not wrong. It just measured the wrong thing.
That order — data layer first, model second — is how CodeBranch approaches software for construction companies generally, because the integration work underneath is what makes anything above it possible.
Frequently Asked Questions
Why do construction AI pilots fail to scale?
What does data readiness actually mean for a contractor?
Can we use AI without cleaning up our data first?
How long does it take to get construction data ready for AI?
Is it worth building AI features if our data is not ready?
How do we evaluate a partner for an AI project in construction?
CodeBranch Team
CodeBranch is an agentic software development partner for U.S. companies, from building new products to scaling existing ones — senior-led teams in Colombia, working U.S. hours. We deliver with AI-native pipelines and our own Spec-Driven Development framework, with quality and security built into every line of code.
Related Articles
Construction Job Costing Starts at the Cost Code
Job costing is an attribution problem, not a reporting problem. The cost code structure decides what the reports will ever be able to tell you.
Unit Price Analysis and the Limits of Generic Estimating
A unit price is not a price. It is a calculation with yields, crews and equipment underneath — and that structure is what generic tools do not store.
When You Should Not Build Custom Construction Software
Every build-vs-buy guide in construction is published by someone selling a platform. Here is the case against building, from a company that builds.