All posts
GDPRSovereign AISME

GDPR-compliant AI for companies: what really matters

10 min readThomas Stermole
Concept graphic: a densely meshed data cluster inside the EU on the left, a third-country zone with a few distant systems on the right; at the border between them, a glowing shield icon representing lawful transfer safeguards such as the DPF and Standard Contractual Clauses.

GDPR compliance is not a product attribute. EU hosting alone is not enough, and self-hosting is not automatically GDPR-compliant either. The earlier question is more useful: Which data does this AI use case actually need to see, which systems may process them, and how much control must the company retain?

That is an architecture decision before it is a provider decision. Start with the use case, data class, data flow and system boundaries; they narrow down the right operating model. The legal assessment follows from there, not the other way around. Whether the EU AI Act bans ChatGPT & Copilot from August 2026 is covered separately in ChatGPT & Copilot at work.

Short answer: A robust AI architecture does not start with the provider. It starts with the use case, data class and control boundary. It makes clear which data reach the application, retrieval layer, model, logs and external services, and who controls access, storage and deletion. That leads to an appropriate operating model; none is automatically legally approved.

Last substantively updated: 17 August 2026

The architecture decision starts with the data

The guiding question is not, “Which model is best?” It is: Which data does this task actually need to see? Broad classes can help with technical risk assessment:

  • Public or low-sensitivity content: publicly available text or ideas with no connection to internal operations.
  • Internal company knowledge: policies, process documents, product knowledge or non-public work in progress.
  • Confidential or business-critical data: contracts, proposals, strategy, source code, prices or trade secrets.
  • Personal data: customer, employee or contact data, as well as content that identifies a person.
  • Special-category or highly regulated data: for example health data or data from highly regulated specialist processes.

These classes do not replace legal classification. They make technical protection needs visible early. A model marketed “for business” says nothing conclusive about them.

Gate 1: Use case and data class

The architecture is determined by the combination of task and data, not by the model.

Drafting general text without internal content can be a sensible fit for centrally managed Business SaaS. An internal RAG system, by contrast, must see documents, permissions and retrieval context. A support assistant working with customer data introduces personal data; particularly sensitive or regulated data raise the bar for isolation, traceability and specialist review.

The specific use case therefore filters out two poor decisions: expensive isolation for harmless work, and convenient standard SaaS for data that should not cross that control boundary.

Gate 2: The actual data flow

“EU storage” or “EU hosting” describes only one part of the system. Architecture needs to account for the complete path of the data:

  • User and client: What do employees enter, what files do they upload, and what remains in the browser?
  • Application, auth and identity: Who may use the use case, and which roles are enforced?
  • Data source and retrieval: Which documents are indexed, which passages are selected for a request, and are source permissions retained?
  • Prompt, context and model inference: Which content reaches the model, and where is inference actually performed?
  • Logs, telemetry and monitoring: Which prompts, metadata or error data remain there, who can view them and for how long?
  • Support access, subprocessors and external integrations: Which additional recipients may receive data, even if they are not visible in the primary AI product?
  • Backups and retention: When do files, embeddings, chats and logs actually disappear from production and backup systems?

Storage and inference must be assessed separately. Data may be stored in the EU while model requests, telemetry or support access cross different boundaries. The visible AI provider is not necessarily the only recipient of data.

This applies to established business products as well. OpenAI's Data Processing Addendum describes transfer mechanisms for EEA data, and OpenAI states that ChatGPT Business and Enterprise data are not used for training by default. Its data and inference residency is available only to eligible Enterprise and Edu customers and does not necessarily cover every item of metadata, support processing or integration. This is not an assessment of OpenAI; it illustrates why the data flow matters more than a single product feature.

Gate 3: Where is the control boundary?

A control boundary defines which components the company controls itself and from which point processing is delegated to external providers. It is not an abstract debate about sovereignty; it is a concrete operating decision.

Key control points include:

  • identity, roles and administrative access,
  • data access and network paths,
  • retrieval data and model access,
  • logging without unnecessary sensitive content,
  • retention, deletion and backups,
  • support access and subprocessors,
  • provider dependency and a realistic exit or migration path.

The further these points sit outside the company’s control, the more closely contracts, configuration and actual operations need to be assessed. Taking on more control creates more room to design the system, but also more responsibility for security, updates, availability and recovery.

Narrowing the operating model from protection needs

  1. 01

    What use case should the AI support?

    • General productivity

      No internal or personal content is needed.

    • Knowledge, support or specialist workflow

      The use case needs company or customer data.

  2. 02

    Which data class must pass through the data flow?

    • Public or low-sensitivity

      Business SaaS can be a reasonable starting point.

    • Internal, confidential or personal

      Assess retrieval, logs, inference and recipients deliberately.

  3. 03

    Is external processing permissible within the required boundary?

    • Yes, with a controlled contract and setup

      Business SaaS or EU-/Enterprise SaaS may fit.

    • Only to a limited extent, or not at all

      Give Private Cloud or isolated self-operation stronger consideration.

  4. 04

    What level of control and operating responsibility is viable?

    • Fast start, less operational work

      Business SaaS or EU-hosted SaaS.

    • Higher control, more operating responsibility

      Assess Private Cloud, self-hosted or on-premise options more closely.

This decision tree narrows architecture paths; it does not replace legal approval.

Gate 4: The right operating model

Operating models differ by more than location. They shift data control, external data flows, operational effort and lock-in risk in different ways.

Operating modelData control and external data flowOperations and time to valueResponsibility and exit
Business SaaSContract and configuration shape control; external processing is part of the model.Fastest start, low in-house operational effort.Assess provider dependency and data paths actively.
EU-hosted / Enterprise SaaSCan improve data residency and administrative control; inference, support and subprocessors still need separate review.Usually quick to introduce.Responsibility remains shared between company and provider.
Private CloudMore isolated instances and more controlled network and data paths are possible.More architecture and operational work, but often a pragmatic middle ground.More control, with clear ownership for platform and operations still required.
Self-hosted / On-premiseHighest technical control when uncontrolled external components are excluded.Longer introduction and substantial work for operations, scaling and resilience.Responsibility sits largely with the company; model and infrastructure changes can be planned more freely.

More control is not automatically better. For a low-sensitivity productivity use case, SaaS can be the sensible decision. On-premise is not a shortcut to compliance; it adds operational responsibility. Private Cloud often fits when protection needs exceed standard SaaS but a fully self-operated platform does not fit the team. Private AI vs public cloud explores the broader trade-offs between public cloud, EU hosting, Private Cloud, on-premise and hybrid architectures.

If that decision leads to a custom, private or more deeply integrated architecture, review the wider system before build or go-live: data, identity, security, evaluation, failure modes and operations. The Enterprise AI Architecture Checklist provides that review.

Three decision scenarios

Scenario A: General productivity AI

For drafting, summarising or ideating without confidential or personal content, Business SaaS can be sufficient. The architecture decision is deliberately simple: use central accounts and administration, define clear usage boundaries, and review training settings and retention. A local platform would not automatically be proportionate here.

Scenario B: RAG over internal company knowledge

A RAG system over policies, documents or product knowledge changes the decision. The data source, index, retrieval layer and model are separate components; source permissions must remain intact during retrieval. That is not an implementation detail but part of the control boundary. Depending on the protection need, EU SaaS, Private Cloud or a controlled in-house architecture may fit. RAG implementation & AI integration covers the technical implementation; RAG permissions and security trimming explains why permissions cannot be lost during retrieval.

Scenario C: Sensitive or regulated data

For special-category personal data, sensitive customer data or highly regulated processes, isolation, access paths, retention and support access are central architecture questions. Private Cloud, self-hosting or on-premise deserve much closer consideration. Architecture can reduce risks, but a binding legal assessment still requires qualified specialist advice.

What architecture does not decide

A sound architecture creates a basis for decisions; it does not replace legal advice. In particular, it does not conclusively determine:

For Austrian companies, the GDPR applies directly and is supplemented by the Austrian Data Protection Act. The Austrian Data Protection Authority explains the relevant legal sources. An architecture decision can structure that assessment; it cannot replace it.

Technical proof: controlled architecture under real constraints

The references overview documents a fully self-hosted AI platform built for experdoo. The starting point was stringent regulatory requirements, sensitive company knowledge and the requirement not to use external public-cloud AI for this data.

The implementation included a self-hosted platform, RAG over documents, policies and specialist processes, multiple agents, and a multi-tenant architecture with clear separation of data and permissions. This does not prove “self-hosting = GDPR compliance.” It shows the technical value of a controlled architecture: data flows, system boundaries, access and knowledge sources can be designed deliberately instead of inherited from a standard SaaS setup.

Practical consequence

The biggest risk is often not the wrong platform but the absence of a deliberately provided, controlled alternative. Shadow AI in the workplace often appears because employees are trying to solve a real productivity problem and have no usable approved option.

The Sovereignty Check turns the use case, protection needs, current data flows and operational effort into a traceable architecture decision. If a problematic upload has already occurred, AI incident response helps with the initial technical assessment and documentation.

Note: This article provides technical orientation and does not replace legal advice. Obtain qualified legal counsel for a binding assessment of a specific case.

Frequently asked questions

Is ChatGPT GDPR-compliant? Whether ChatGPT can be used in compliance with the GDPR depends on the plan, contract, configuration and actual data flow. ChatGPT Business and Enterprise data are not used for training by default according to the provider; EU residency is limited to eligible Enterprise and Edu customers and does not cover every processing activity.

Which AI is GDPR-compliant? No architecture is automatically compliant. The use case, data class, data flow, control boundary and legal framework matter. EU-hosted or self-operated solutions can reduce control and transfer risks, but do not replace the assessment.

Do my data have to stay in the EU? Not necessarily. Transfers to certified US companies can currently rely on the DPF; depending on the recipient, Standard Contractual Clauses and supplementary measures may also be available. Processing in the EU reduces the transfer burden, but it does not automatically satisfy every requirement.

Is local AI automatically GDPR-compliant? No. Self-operation can reduce external data flows if APIs, telemetry, support access and connected services are controlled as well. Legal basis, purpose limitation, permissions, deletion and transparency still apply.


Want to know which control boundary your AI use case needs? The Sovereignty Check creates clarity within a week on data flows, operating model and the next architecture decisions.

Next step

Sounds relevant for your company?

In a no-obligation initial call, we clarify within 30 minutes whether and where getting started is worthwhile for you — honestly and without sales pressure.

Request an initial call