
Data Access Is Not Data Ownership: What AI SaaS Agreements Need to Define
Learn how to structure AI SaaS data rights effectively. Navigate GDPR and EU Data Act requirements for data access, use, model training, and offboarding.
Key Takeaways
- A standard 'customer owns all customer data' clause does not, by itself, define an AI SaaS vendor's rights to access, process, retain, reuse, or use that data for model training.
- AI SaaS contracts should distinguish customer inputs, prompts, outputs, logs, feedback, and embeddings, applying tailored use, retention, security, and confidentiality terms to each.
- For data processing services within Chapter VI of the EU Data Act, Article 23 requires providers to remove specified switching obstacles. Art. 25 generally limits the notice period to two months, provides for a transitional period of up to 30 days subject to a technical-feasibility exception, and requires at least 30 days for data retrieval after that period.
- Under Art. 13 Data Act, covered B2B terms that are unilaterally imposed are non-binding if they are unfair. The rule is limited to terms on data access and use, or related liability and remedies.
Introduction
When negotiating commercial contracts for generative AI platforms, enterprise customers and software vendors frequently clash over a straightforward question: who owns the data? In traditional cloud models, agreements often rely on a binary division where the customer owns its uploaded data and the vendor owns the software. For AI platforms, however, that distinction is often incomplete.
Legal ownership of data under intellectual property frameworks is narrow and does not govern operational data access, reuse, or model-training rights. Defining AI SaaS data rights requires moving beyond abstract ownership clauses to establish precise, purpose-driven rules. This guide outlines how to structure customer data ownership and SaaS data access in AI SaaS agreements to support enterprise procurement and compliance with applicable European Union rules.
1. Begin with the Data, Not the Ownership Clause
A standard clause stating that the customer owns its data does not determine the vendor's actual access, processing, retention, or reuse rights. EU law does not generally recognize a unitary ownership right in data as such, and raw facts are not protected by copyright merely because they are facts. Database rights, trade-secret protection, contractual rights, and other legal interests may nevertheless apply. The contract should therefore define operational permissions rather than rely on a general property claim.
1.1. Identify the relevant data categories
AI SaaS platforms process a range of information. Agreements should avoid relying only on a single, catch-all definition of 'Customer Data' and should identify the following categories where relevant:
- Customer Inputs and Prompts: Raw datasets, files, or natural language prompts entered into the system by users.
- AI-Generated Outputs: The text, code, structured files, or predictions generated by the models.
- System Logs and Metadata: Operational telemetry, system performance metrics, and audit trails.
- Embeddings: Vectorized, mathematical representations of customer inputs generated during real-time processing.
- User Feedback: Explicit corrections, ratings, or qualitative reviews provided by users regarding output quality.
Failing to separate these categories makes it difficult to define appropriate operational boundaries and security controls.
1.2. Separate ownership, control and access
Legal protection for data and databases is limited and fact-specific. The Database Directive (Directive 96/9/EC) protects a database by copyright where the selection or arrangement of its contents is the author's own intellectual creation and, separately, provides a sui generis right where obtaining, verifying, or presenting the contents entails substantial investment. For data obtained from or generated by connected products or related services within the scope of the Data Act, Art. 43 excludes that sui generis right.
Because exclusive property rights in data are limited, parties should use contract terms to define access and use. A customer may retain rights in its proprietary databases and other materials, while granting the AI SaaS vendor a license limited to the access, transformation, and analysis needed to provide the service.
1.3. Identify the legal and contractual basis for each use
Where personal data are involved, processing must comply with the principles in Art. 5 GDPR and rest on a valid lawful basis under Art. 6. Controller and processor roles are determined functionally for each processing operation, as explained in EDPB Guidelines 07/2020. The contract should reflect whether the vendor acts as a processor under Art. 28 GDPR or as an independent or joint controller for a separate purpose, but contractual labels do not determine the legal role by themselves.
2. Define Who May Use Data and for What Purpose
Enterprise customers often resist broad, open-ended data licenses. AI SaaS agreements should limit data use to specified purposes and distinguish uses needed to provide the service from optional reuse.
2.1. Core service delivery
The contract should state the purposes for which the vendor may access customer inputs and prompts, including delivery, maintenance, and security of the SaaS platform. Where the vendor acts as a processor, Art. 28(3)(a) GDPR requires it to process personal data only on the controller's documented instructions, unless Union or Member State law requires otherwise.
2.2. Model training and product improvement
The use of customer data for model training is an important negotiation point.
- Personal Data Restrictions: If model training involves personal data, the purpose-limitation and data-minimisation principles in Art. 5(1)(b) and 5(1)(c) GDPR apply. EDPB Opinion 28/2024 also addresses, on a case-by-case basis, when an AI model may be anonymous and when legitimate interests may provide a lawful basis; it does not create a separate legal basis for AI training.
- Technical Mitigation: The CNIL's January 2026 recommendations describe data minimisation, anonymisation or pseudonymisation, and risk-appropriate security measures as possible safeguards. The appropriate measures depend on the processing and do not, by themselves, establish compliance.
Enterprise agreements should state whether customer inputs and prompts may be used for model training and, where that use is optional, provide a clear opt-out or separate opt-in mechanism. A contractual customer choice does not replace any required lawful basis, transparency, or data-subject rights.
2.3. Third-party and sub-processor access
Where the AI SaaS platform integrates an upstream API provider, that provider is a sub-processor only if it processes personal data on behalf of the SaaS vendor in the vendor's capacity as a processor. In that case, Art. 28 GDPR requires the controller's specific or general written authorization and the flow-down of the same data-protection obligations. Art. 32 GDPR requires security measures appropriate to the risk. For transfers outside the EEA, the parties must use an applicable mechanism under Art. 44–49; standard contractual clauses are one option, not a universal requirement.
2.4. Retention, deletion and continuing use
Agreements should define retention periods for each data category. Raw prompts may be deleted quickly, while the influence of particular training data on model parameters may not be readily isolated. The right to erasure under Art. 17 GDPR applies only where its conditions are met, and rights may extend to a model if the model is not anonymous. The CNIL's January 2026 guidance calls for realistic and proportionate ways to facilitate those rights and treats machine unlearning as a developing technique, not a universal requirement. Contracts should define post-termination use and retention, legal-retention exceptions, and the treatment of model parameters.
3. Address Outputs, Derived Data and Business Knowledge
Outputs, technical representations, and aggregated information require separate treatment in the contract. Addressing these issues early can reduce intellectual property disputes.
3.1. AI-generated outputs
Contracts should allocate, as between the vendor and the customer, any intellectual property rights that arise in generated outputs and grant the customer the rights needed for its intended use. They should not assume that every output is protectable or that all rights can be assigned. The vendor should retain its rights in pre-existing model architectures, prompt templates, and software, subject to third-party rights and the agreed service terms.
3.2. Embeddings and technical representations
Vector embeddings are mathematical representations of input data used in vector databases and may preserve or reveal sensitive information about the source data. Under the Trade Secrets Directive (Directive (EU) 2016/943), embeddings or vector databases qualify as trade secrets only if the information is secret, has commercial value because it is secret, and is subject to reasonable steps to keep it secret. Agreements should apply confidentiality and security measures proportionate to the information and risk.
3.3. Usage analytics and aggregated information
SaaS providers may use telemetry to optimize performance and manage infrastructure costs. The agreement may permit aggregated or effectively anonymized telemetry, provided it does not reveal personal data or confidential customer information and cannot reasonably be used to identify a customer or individual. If personal data remain, the GDPR continues to apply.
3.4. Feedback and evaluation results
To reduce intellectual property disputes, the agreement may grant the vendor a non-exclusive license to use customer feedback or evaluation results for specified product-improvement purposes. The scope, duration, and any irrevocability should match the parties' negotiation and should not automatically extend to customer confidential information, personal data, or embedded third-party material.
3.5. Confidential information and trade secrets
To support trade-secret protection under the Trade Secrets Directive, agreements should establish clear confidentiality boundaries and require measures appropriate to the information and risk. Depending on the service, those measures may include identity and access management (IAM), tenant isolation, and encryption in transit and at rest. Where personal data are involved, Art. 32 GDPR independently requires risk-appropriate security.
4. Make Data Rights Work Across the Contract Lifecycle
Data rights should be operationalized throughout the contract lifecycle, from onboarding to termination and offboarding.
4.1. Access, export and portability
Standard contracts sometimes confuse individual data portability under GDPR Article 20 with B2B data portability under Data Act.
On one hand, Art. 20 of GDPR applies to personal data provided by the data subject where processing is based on consent or contract and is carried out by automated means.
On the other hand, under Chapter VI of the Data Act, providers of data processing services, including SaaS, must remove specified obstacles to switching. Art. 25 requires written terms that generally limit the notice period to two months, provide for a transitional period of up to 30 days unless a technical-feasibility extension applies, and allow at least 30 days to retrieve exportable data after that period. Art. 30 addresses functional equivalence and export formats. Art. 31 contains limited exceptions for certain custom-built services and non-production testing services.
4.2. Instructions, security and accountability
Art. 12–14 and 32 GDPR do not create a general duty for every SaaS vendor to maintain detailed logs of all AI processing. Depending on the parties' roles and the processing, Art. 28 and 30 GDPR may require documented instructions, assistance, and records. The agreement should identify which logs and compliance documentation the vendor will maintain or provide, with appropriate security and confidentiality limits.
4.3. Changes to data use
AI SaaS agreements should define how data-use terms may change during the contract. UnderArt. 13 Data Act, a covered B2B term on data access or use, or related liability and remedies, is non-binding if it was unilaterally imposed and is unfair. The Article does not prohibit every unilateral change; for indefinite contracts, certain reserved changes may be permissible where the contract states a valid reason and provides reasonable notice and a right to terminate without cost. In Germany, the DADG designates BNetzA as the competent authority. Remedies and penalties depend on the specific infringement.
4.4. Exit and post-termination arrangements
A common compliance gap is the misalignment of return, retrieval, and deletion timelines. Art. 28(3)(g) GDPR requires a processor, at the controller's choice, to delete or return personal data after the end of the services, unless applicable law requires storage. Art. 25 GDPR separately requires qualifying data processing service agreements to address transition, retrieval, and erasure. These rules do not necessarily conflict: the agreement should coordinate when services and transitions end, which data are returned or retrieved, and when active copies and backups are deleted.
Conclusion
Moving from a generic 'customer owns all data' clause to structured AI SaaS data rights gives enterprise customers and vendors a clearer basis for procurement and compliance. Defining rights and limits for prompts, outputs, logs, embeddings, training, and offboarding can reduce ambiguity and support alignment with the GDPR, where applicable, the EU Data Act, and the Trade Secrets Directive.
About the author
Junzhe Dai
Junzhe Dai is a PhD candidate at the Faculty of Law, Humboldt University of Berlin. His research focuses on data market regulation, data protection law, and AI governance, with particular interest in the GDPR, the AI Act, the Data Act, and comparative analyses of EU and Chinese digital regulatory frameworks.
Need help with compliance?
Book a free 30-minute call to review your GDPR and EU AI Act readiness.