And most firms don’t even know it.
Every day, your team opens an AI tool, pastes in a project specification, a client contract, a proprietary process document, and asks it a question. The AI answers. Everyone moves on.
What most people don’t ask is: where did that information just go?
The answer should concern every firm leader in the AEC industry.
The Consumer AI Tools Your Employees Are Already Using
The most immediate threat to your competitive advantage is not a sophisticated cyberattack. It’s an engineer pasting source code into ChatGPT at 11 PM to debug a problem. A project manager uploading a client proposal to Claude to improve the writing. A director feeding quarterly data into Gemini to build a forecast. This is shadow AI, and it’s already pervasive.
According to the LayerX Enterprise AI & SaaS Data Security Report, 77% of employees paste data into generative AI tools, with 82% of that activity occurring through unmanaged personal accounts that fall entirely outside corporate security controls. Cyberhaven’s analysis found that 39.7% of all AI interactions involve sensitive data.
For architecture and engineering firms, that means design files, client specifications, proprietary BIM standards, cost estimates, and engineering calculations flowing into systems owned by vendors who face no legal obligation to protect them as your intellectual property.
What the terms of service actually permit
OpenAI’s consumer terms permit data use to “provide, maintain, develop, and improve” their services, with training opt-out buried in settings most users never find. Anthropic’s consumer terms now default to a five-year data retention period for users who don’t actively opt out. Google’s own privacy documentation warns users directly: “Please don’t enter confidential information that you wouldn’t want a reviewer to see or Google to use to improve our services.”
The business-tier protections (the ones that actually prohibit training on your data) apply only to enterprise contracts. The consumer tools your employees are using right now? None of those protections apply.
Samsung learned this the hard way
On March 11, 2023, Samsung’s Device Solutions division authorized engineers to use ChatGPT. Within twenty days, three separate data compromises occurred. An engineer pasted buggy source code into ChatGPT to find a fix. A second employee submitted proprietary code for defective equipment identification and chip yield optimization. A third recorded an internal company meeting, transcribed it using an AI tool, then entered the full transcript into ChatGPT to generate meeting minutes.
Because ChatGPT’s terms at the time permitted training on user inputs, Samsung’s semiconductor trade secrets effectively became part of OpenAI’s training corpus with no recourse for retrieval or deletion. Samsung implemented emergency measures before issuing a company-wide ban on all generative AI tools.
Samsung was far from alone. Amazon’s corporate counsel warned employees not to share “any Amazon confidential information” with ChatGPT. JPMorgan Chase, Goldman Sachs, Bank of America, Citigroup, and Deutsche Bank all restricted or banned the tool. Apple blocked both ChatGPT and GitHub Copilot, citing concerns about confidential data and product roadmap leaks.
The Invisible Threat: AI Startups Built on Borrowed Models
The most insidious threat isn’t from established vendors. It’s from the wave of new AI startups flooding the AEC market. These companies are solving real problems with genuinely useful tools, but built on a fundamentally risky architecture for the businesses that use their services.
These startups don’t have proprietary AI models. Instead, AI companies call frontier models from OpenAI, Anthropic, or other providers through APIs, wrap them in purpose-built workflows, and charge you a subscription. The value proposition is real: faster proposals, smarter analysis, better workflows. The risk is structural and often invisible.
Here’s what you don’t know: what happens to your data once it lands on their servers?
Most startups are typically pre-revenue or barely profitable. They’re under intense pressure to find monetization paths. Your data (architectural specifications, engineering calculations, project details, client information, cost estimates) sits on their infrastructure with no real guarantee about how it will be used, retained, or shared.
The terms of service probably say something like: “We don’t use your data to train models.” But that’s a promise from a company that, honestly, may not exist in two years. If the company pivots, gets acquired, or faces financial pressure, those promises evaporate.
Worse: even if the startup is trustworthy, you have no visibility into their sub-processor chain. They’re calling OpenAI or Anthropic APIs. What data do those calls include? How long is it retained? Is it used for model improvement? The startup may not know either - they’re just passing your data through to get an answer back.
The Legal Ground Is Shifting
In February 2026, Judge Jed Rakoff of the Southern District of New York ruled that documents prepared using consumer AI tools were not protected by attorney-client privilege, because users “do not have substantial privacy interests” in communications with public AI platforms. The reasoning is directly applicable to trade secret law: if your employees input proprietary information into consumer AI tools, a court may find your firm failed the “reasonable measures” test required to maintain trade secret protection.
There is currently no comprehensive legal framework protecting corporate intellectual property in AI systems equivalent to how GDPR protects personal data. Trade secret law protects information kept secret through reasonable measures - but protection is lost upon disclosure.
The IBM 2025 Cost of a Data Breach Report put financial teeth on this: shadow AI is now one of the top three costliest breach factors, adding $670,000 in additional costs per incident. Yet only 37% of organizations have policies to manage or detect shadow AI at all.
The Real Cost: Your Expertise Is Being Commoditized
Think about what actually differentiates your firm from the competition.
It’s not your software subscriptions. It’s not your headcount. It’s the institutional knowledge that took decades to build - your design standards, your engineering judgment, your lessons learned, your client relationships, your proprietary workflows.
When your architects use cloud AI tools to analyze designs or generate specifications, the patterns embedded in those interactions contribute - marginally, but cumulatively - to AI systems that serve your competitors too. The aggregate effect across thousands of interactions creates a flywheel that systematically democratizes specialized knowledge that took your firm decades to develop.
McKinsey put it plainly: “If you have a generative model, your competitor probably has it as well. The likely moat will be customization.”
The firms that preserve competitive advantage will be those that keep their proprietary data within systems they control, using that data to build AI capabilities that are specifically and exclusively theirs.
The Path Forward: From Risk to Sovereignty
Step 1: Classify your data
Identify at least four tiers: public information that can be freely shared; internal information for approved enterprise AI services; confidential information (client data, financial projections, competitive strategy) for hyperscaler-hosted AI only; and restricted information (trade secrets, patented processes, unreleased designs) that should never leave infrastructure you control.
Step 2: Audit your current AI usage
For every AI tool your team is currently using, whether officially approved or shadow AI, understand where your data flows. This audit will likely reveal gaps between your stated policy and actual practice.
Step 3: Shift to hyperscaler-hosted models for frontier AI
For organizations that need frontier AI capabilities today, hyperscaler-hosted models (on Microsoft Azure, Google AI Studio, or AWS Bedrock) represent the best available balance of capability and protection. Microsoft’s official documentation for Azure AI Foundry states the distinction with unusual clarity: “Customer prompts and completions are NOT available to OpenAI or other Azure Direct Model providers, are NOT used by Azure Direct Model providers to improve their models or services, and are NOT used to train any generative AI foundation models without your permission or instruction.”
Step 4: Build toward private AI infrastructure
The strategic trajectory points unmistakably toward privately hosted AI as the sovereign standard. The open-source model landscape has reached a capability threshold that makes this practical. Lockheed Martin’s trajectory is instructive: the company built its AI Factory on NVIDIA DGX SuperPOD architecture with over 8,000 engineers using the platform.
How HallianAI Is Built Differently
HallianAI is a private, centralized AI engine for architecture and engineering firms.
Your data stays 100% behind your firewall. The platform deploys on your infrastructure, either on-premises or in your private cloud. Your documents, conversations, workflows, and knowledge bases never touch a shared server.
You own the AI supply chain. You manage your own cloud accounts (Azure AI Foundry, AWS Bedrock, or Google Vertex AI) for model access. You control the keys. You control the costs. There’s no third-party markup, no black-box processing, no vendor lock-in.
It’s a centralized engine, not a point solution. Instead of bolting together dozens of disconnected AI tools, HallianAI consolidates all AI capabilities into one unified platform.
What this means in practice
- Technical Agents that reference your firm’s codes, standards, specs, and internal processes
- Project Manager Assistants that handle repetitive admin work - status reports, meeting minutes, and change orders generated in minutes instead of hours
- Proposal Response Automation that identifies relevant past projects, generates tailored SOQs, and pulls team resumes
- Reporting Automation Workflows that draft technical and non-technical reports from your templates
- Executive Management Agents that pull real-time financials, strategic context, and institutional knowledge into one trusted partner
The compounding advantage
Every interaction your team has with HallianAI makes your firm’s AI smarter about your business specifically. Organizations using retrieval-augmented generation systems grounded in proprietary knowledge bases report three-to-five-times-faster information retrieval and 45-65% reduction in time searching for answers.
Measured results across multiple firms show HallianAI consistently reduces high-value staff admin time by 40-70% in targeted workflows. For a 200-person firm, that’s thousands of hours reclaimed annually.
The Choice Is Yours
Every firm deploying AI right now is making an architectural decision whether they realize it or not. Some are choosing convenience - and surrendering control over their most valuable assets. Others are choosing sovereignty - and building AI infrastructure where their proprietary knowledge compounds in value.
The difference is about ownership. It’s about whether your firm’s accumulated expertise becomes a shared commodity or a proprietary asset that only you can leverage.
HallianAI exists to help your firm build that infrastructure.