Skip to content
Construction

Why AI in Construction Stalls at the Data Layer

CT

CodeBranch Team

Why AI in Construction Stalls at the Data Layer

Eighty-seven percent of US contractors believe AI will significantly transform their business. Twenty-six percent rate the quality of their current data as high. That gap is not a gap in enthusiasm or in technology — it is the reason most construction AI projects stall between a promising pilot and anything in production.

Quick Summary

  • In a late-2025 survey of 235 US general and trade contractors, 87% expected AI to transform their business while only 26% rated their data quality as high.
  • Of 23 AI functions evaluated in that survey, 20 were in use by fewer than 15% of respondents — expectation is near-universal, adoption is not.
  • A separate 2026 survey of 2,500 industry leaders found integration with legacy systems to be the number one barrier to AI, cited by 50%.
  • The bottleneck is not the model. It is that the data a model would need is spread across systems that do not agree, recorded inconsistently, or not recorded at all.
  • Which means most construction AI work is integration and data engineering work, and scoping it as anything else is how projects run long.

That framing — the interesting problem sits below the feature, in the layer nobody demos — runs through this whole vertical. More on it in what construction companies build when their platform runs out.

What Do the Numbers Actually Say?

The most specific recent data comes from a Dodge Construction Network and CMiC survey of 235 US general and trade contractors, conducted in late 2025. Three findings from it matter more than the headline.

Expectation is nearly universal. 87% believe AI will significantly transform their business; 85% expect reduced time on repetitive tasks. Among large contractors, 86% believe AI confers competitive advantage.

Adoption is not. Of 23 AI functions evaluated, 20 were used by fewer than 15% of respondents. Only 19% were adapting legacy workflows for AI — which is the step that turns a tool into a capability.

And the constraint is visible in the data. Only 26% rated the quality of their current data as high. Meanwhile 57% cited concerns about output accuracy and reliability as a barrier. Those two findings are the same finding: a model trained or grounded on inconsistent data produces unreliable output, and the people using it correctly stop trusting it.

A 2026 survey of 2,500 industry leaders points at the same place from a different angle: integration with legacy systems was the single most-cited barrier to AI adoption, named by 50% of respondents — ahead of talent, cost, regulation and ROI measurement.

Why Is Construction Data Harder Than It Looks?

Because the information exists; it is just not in a form anything can consume.

A contractor knows what a linear meter of conduit costs. That knowledge lives in an estimator’s judgment, in a spreadsheet from the last similar project, in the accounting system under a cost code that means something slightly different than it did two years ago, and in a supplier email. All four are correct. None of them is queryable, and they do not agree.

This is the structural version of a finding that has held for years: in the last systematic survey of construction technology, 62% of firms were still estimating in spreadsheets and 49% moved data between applications by hand. A spreadsheet is not a data quality problem in itself — plenty of good work happens in spreadsheets. It becomes one the moment you want a system to learn from four hundred of them.

The three failure modes, in increasing order of difficulty:

The data is not in a system. It is in spreadsheets, emails, or people. Fixable with a project, measured in months.

The data is in a system but recorded inconsistently. The same activity is coded differently across projects, or across teams, or before and after someone reorganized the cost code structure. Fixable, but it requires changing how people work, which is measured in quarters.

The data is in a system, consistent, and unreachable. It is in a platform whose export is a CSV a human has to download. This is the one that looks easiest and is often the one that decides whether the project is viable, because it determines whether anything can be automated or someone is running an export every week forever.

Which AI Work Does Not Depend on Clean Data?

This is the distinction that makes a project deliverable in weeks rather than quarters, and it gets collapsed constantly.

Language work tolerates mess. Reading a submittal, summarizing an RFI thread, extracting terms from a contract, classifying an incoming document — these operate on unstructured input because unstructured input is what they were built for. A contractor with genuinely chaotic data can still get value here, immediately.

Prediction and estimation do not. Forecasting a schedule, estimating a cost from history, optimizing a sequence, flagging an anomaly — all of these learn from your past, so they inherit whatever inconsistency your past contains. Run these on unready data and the output is confidently wrong, which is worse than no output.

Language and document workPrediction and estimation
InputUnstructured, messyHistorical, structured
Data readiness requiredLowHigh
Time to first valueWeeksQuarters, after data work
Failure modeVisible — a bad summary looks badInvisible — a bad forecast looks fine
Where it fitsStart hereBuild the data layer first

The failure mode row is the one worth sitting with. A bad summary is obviously bad, so a human catches it. A cost forecast that is 18% low looks exactly like a cost forecast that is right, and nobody catches it until the project is underway. CodeBranch treats that asymmetry as the reason to sequence these differently, not as a detail.

What Does a Readiness Assessment Look For?

Three questions, answered against real data rather than a description of the system.

Does the data exist in a system that can be queried without a person in the loop. Is the same real-world thing recorded the same way across projects, teams and time. And can it be retrieved on a schedule without a manual export.

The answers determine the shape of the project. If all three are yes, the AI work is the AI work. If the third is no, the first phase is integration. If the second is no, the first phase is a conversation about how the organization records things, which is not a software project at all and should not be priced as one.

CodeBranch runs this as a discrete engagement — an AI readiness assessment — precisely because the answer changes the scope by an order of magnitude, and discovering it after a proposal is signed is how projects go badly for both sides.

What Does This Look Like When It Works?

The pattern worth copying is that the model is the small part.

In a quoting platform CodeBranch built for an electrical contractor, an AI recommendation engine suggests components as the user designs a home automation project on a floor plan, and it works — the recommendations are complete and installable. The reason it works is not the model. It is that the dependencies between components were modeled explicitly first: which devices require which controllers, what a given configuration implies for wiring and labor. The AI operates on a structure that describes the domain correctly.

Strip that structure out and the same model would produce plausible suggestions that do not add up to a working installation. The intelligence in the feature comes mostly from the data layer beneath it, and that layer was engineering work, not prompting.

That is the general shape. The AI feature is the visible tenth of a project whose other nine tenths is making the data mean something consistent.

Where Should a Contractor Start?

Start with the thing that is true regardless of the AI decision: get your project data into a system that can be queried.

That is worth doing on its own merits. A contractor who can ask their own history what a square meter of a given assembly has actually cost across the last twenty projects has gained something real, with or without a model attached. It also happens to be the prerequisite for every prediction and estimation capability worth having.

Then run the language and document work in parallel, since it does not depend on that foundation and can deliver value while the foundation is being built.

What CodeBranch would avoid is the sequence that most stalled projects followed: pick an impressive AI use case, build a pilot on hand-curated data, demonstrate it successfully, and then discover that production data cannot support it. The pilot was not wrong. It just measured the wrong thing.

That order — data layer first, model second — is how CodeBranch approaches software for construction companies generally, because the integration work underneath is what makes anything above it possible.

Frequently Asked Questions

Why do construction AI pilots fail to scale?
Usually because the pilot ran on a curated dataset that does not exist in production. Someone cleaned a few projects worth of data by hand to prove the concept, and the model worked. Scaling means the model meets the data as it actually is — inconsistent, incomplete, spread across systems that do not agree. CodeBranch assesses that gap before a project starts, because it determines whether the work is an AI project or a data engineering project wearing an AI label.
What does data readiness actually mean for a contractor?
Three things: the data exists in a system rather than in someone head or a spreadsheet, it is structured consistently enough that the same thing is recorded the same way across projects, and it can be retrieved without a manual export. Most contractors have the first and lack the other two. CodeBranch runs an AI readiness assessment that maps which of the three is missing, because the answer changes the scope by an order of magnitude.
Can we use AI without cleaning up our data first?
For some things, yes. Document processing, drafting, summarization and classification work on messy inputs because that is what they are designed for. What does not work on messy data is anything that predicts, estimates or optimizes based on your history — those depend on the history being consistent. CodeBranch separates these two categories at the start of an engagement, because the first can deliver value in weeks and the second cannot.
How long does it take to get construction data ready for AI?
It depends entirely on which of the three readiness problems you have, which is why an honest answer requires looking first. Getting data out of spreadsheets into a system is a project measured in months. Standardizing how an existing system records things is measured in quarters and requires changing how people work. CodeBranch scopes this in the Product Definition phase rather than estimating it blind, because the range between the easy case and the hard one is too wide to guess.
Is it worth building AI features if our data is not ready?
It is worth building the data layer, which is usually the unglamorous part of the same project. In practice, the integration and structuring work is what makes the AI feature possible, and it delivers value on its own — a contractor who can query their own project history has gained something even before a model touches it. CodeBranch builds in that order deliberately.
How do we evaluate a partner for an AI project in construction?
Ask them what they would check before writing any model code, and how they would respond if the answer came back badly. A partner who says they would assess data readiness first, and who can describe what a failing assessment would change about the proposal, is telling you they have seen this fail. CodeBranch treats a poor readiness result as a reason to change the scope rather than a reason to proceed more optimistically.
CT

CodeBranch Team

CodeBranch is an agentic software development partner for U.S. companies, from building new products to scaling existing ones — senior-led teams in Colombia, working U.S. hours. We deliver with AI-native pipelines and our own Spec-Driven Development framework, with quality and security built into every line of code.

LinkedIn · codebranch.co

construction AI data integrations

Related Articles