
User Prompts and AI Interaction Data in Generative AI SaaS: GDPR Risks and Procurement Readiness
Learn to manage AI interaction data and GDPR risks for EU enterprise procurement. A guide for generative AI SaaS founders on compliance and data architecture.
Learn how to manage AI interaction data, user prompts and related GDPR, AI Act issues when preparing for EU enterprise procurement. A guide for generative AI SaaS founders.
Key Takeaways
- AI interaction data, including prompts, uploaded files, outputs, feedback and metadata , may contain personal data within the meaning of GDPR Article 4(1). Where it does, the relevant controller needs a defined purpose, an appropriate legal basis and processes for handling applicable data-subject requests, including erasure requests under Article 17.
- One recurring distinction in enterprise procurement is between service provision and model development. A 'no training by default' commitment can function as a contractual and technical safeguard that reduces confidentiality and trade-secret risks, but it is not itself a statutory requirement under Directive on the Protection of Trade Secrets (Directive (EU) 2016/943).
- AI Act Article 53(1)(d) requires providers of general-purpose AI models to publish a sufficiently detailed summary of the content used for training, using the AI Office template. This obligation applies only where the SaaS provider is also a provider of a general-purpose AI model; user-contributed interaction data is relevant if it is used as training or fine-tuning content.
- Technical founders should assess and document processing operations separately by purpose and legal role. Security and billing logs, prompt content used to provide the service, product analytics and data used for model development may require different retention periods, access controls and legal assessments; neither legitimate interests nor consent applies automatically.
Why AI Interaction Data Is Becoming Important and Sensitive
AI interaction data is a broad category covering digital interactions between a user and an AI system. This includes user prompts, uploaded files (such as PDFs, spreadsheets, or images), generated outputs, user feedback (such as 'thumbs up' or 'thumbs down' ratings) and associated metadata and logs. In the current generative AI SaaS landscape, AI interaction data has evolved from simple telemetry into an important compliance dataset. For technical founders, understanding this data is no longer only about optimizing a model; it can also help them respond to EU enterprise procurement reviews.From a regulatory perspective, this data can be sensitive because it acts as a container. A prompt may look like a simple query, but it may contain personal data relating to the customer's clients or confidential business information that could qualify as a trade secret. As EU companies integrate generative AI, they may ask how this data is stored, who can access it and whether it is reused to develop a provider's own or a third party's models without clear authorization.
Interaction Data as a Compliance Issue
For a B2B SaaS company, AI interaction data may become a focus during legal and security review. Typical lifecycle questions include: Is the data stored? Is it deleted according to a defined schedule? Are third parties, such as subprocessors or LLM API providers, able to access it? And is it used for model training? Under the GDPR, a prompt falls within personal-data processing where its content or associated metadata relates to an identified or identifiable natural person. An email address will generally meet this threshold, while a job title will do so only where the person is identifiable in context. If a provider cannot answer these questions with sufficient detail, the customer's legal or data-protection review may be affected.
Overlapping Legal Frameworks for Interaction Data
The legal treatment of a prompt and related Interaction Data may involve more than one framework. The GDPR applies where the prompt or associated data contains personal data. Directive on the Protection of Trade Secrets (Directive (EU) 2016/943) may also be relevant where the prompt contains information that meets all elements of the trade-secret definition in Article 2(1): the information is secret, has commercial value because it is secret and has been subject to reasonable steps to keep it secret. Broad contractual rights allowing a SaaS provider to reuse prompt data for 'model improvement' without adequate safeguards may make it harder for a customer to demonstrate that it took reasonable steps to preserve secrecy. This does not automatically eliminate trade-secret protection, but it can increase confidentiality and intellectual-property risk.
Core GDPR Concerns: Personal Data, Roles, Processing Purposes and Storage
Managing AI interaction data calls for careful application of the GDPR's core principles. The EDPB's May 2024 ChatGPT Taskforce report identified prompts, file uploads and user feedback as forms of input data and examined the use of such content for training. The report presented preliminary views in the context of ongoing investigations rather than establishing a general rule for all generative AI services. A prudent approach is therefore to avoid catch-all descriptions of processing and instead define purposes, data flows and roles with sufficient specificity.
Whether Interaction Data Constitutes Personal Data
Under GDPR Article 4(1), personal data is any information relating to an identified or identifiable natural person. User prompts may contain names, contact details, professional information, biographical details or other information that meets this definition, but they are not automatically personal data in every case. The Italian Garante’s 2023 action concerning ChatGPT focused, among other issues, on inadequate transparency, the absence of an appropriate legal basis for processing personal data to train the service, inaccurate personal data in outputs and insufficient age-related safeguards. It did not establish that every prompt is personal data. Providers should also consider GDPR Article 9: prompts may reveal special categories of personal data, such as health information or political opinions, in which case an Article 9(2) condition is required in addition to an Article 6 legal basis.
The Role of the SaaS Provider
The SaaS provider's legal role depends on who determines the purposes and essential means of each processing operation and must be assessed in light of the actual arrangement. Where a provider processes customer interaction data solely on the customer's documented instructions to deliver the service, it may act as a processor under GDPR Article 28, although the qualification remains fact-specific. Where it independently decides to reuse that data for its own analytics, product development or model training, it may act as a controller for that separate purpose. Enterprise customers may seek clear restrictions, transparency and approval rights for such secondary uses in the Data Processing Agreement (DPA) or related documentation.
Legal Basis and Processing Purposes
Each controller must identify a legal basis under GDPR Article 6 for its processing operations. Article 6(1)(b) may support service provision only where the processing is objectively necessary to perform a contract with the data subject; in a B2B setting, the end user is not always a party to the customer's contract with the SaaS provider. Legitimate interests under Article 6(1)(f) may be available for certain security, product-improvement or model-development activities, but only after assessing necessity, reasonable expectations, impacts and safeguards. Consent or an opt-in mechanism may be selected for some secondary uses, but consent must be freely given, specific, informed and unambiguous, and it is not automatically the most appropriate basis. Where the SaaS provider acts as a processor for a particular operation, it ordinarily processes the data on the controller's documented instructions under the Article 28 arrangement, while remaining subject to the processor obligations imposed directly by the GDPR.
Storage, Retention and Deletion
GDPR Article 5(1)(e), rather than Article 5(1)(c), establishes the storage-limitation principle: personal data should be kept in identifiable form no longer than necessary for the relevant purpose. Where an AI SaaS provider acts as a controller, it should define and justify appropriate retention periods for prompts, outputs, security logs, billing data and support records. Where it acts as a processor, the relevant retention and deletion arrangements should normally be reflected in the controller's documented instructions and the Article 28 agreement. Article 17 may require erasure in applicable circumstances, but the right is not absolute and the controller must also consider statutory exceptions and the scope of data held by its processors. Deleting records from source databases and vector stores is generally a different task from addressing data that may have influenced model parameters. Where personal data has been used for fine-tuning, the provider should be able to explain the technical architecture, identify what can be deleted or isolated and describe how rights requests are handled. Filtering personal data before training or avoiding the use of customer prompts for training can reduce this complexity, but these measures do not by themselves guarantee GDPR compliance.
Concerns from EU Clients: Is Interaction Data Used for Model Training?
One question likely to arise in EU AI procurement is: 'Is my data being used to train your model or your provider's model?' The answer should be specific and supported by the product architecture, contracts and actual settings.
No Training by Default
For enterprise SaaS, 'No Training by Default' can be adopted as a procurement and contractual position. In practice, it typically means that customer interaction data is used to provide the contracted service and is not reused to train or improve general-purpose or shared models unless the customer has separately agreed. When using a third-party LLM provider, founders should verify the provider's contractual terms, retention settings, abuse-monitoring practices and available zero-retention or equivalent controls. A feature labelled 'zero data retention' should not be treated as sufficient without confirming its scope and exceptions.
Product Improvement Without Model Training
Product improvement and model training should be distinguished by the actual processing performed, not only by the label used. Product improvement may include analysing latency, interface performance, failure patterns or security events without using prompt content to adjust model parameters. Model training or fine-tuning involves using data to change model behaviour or parameters. These purposes may need to be described separately in privacy and customer-facing documentation where they are materially distinct and where the applicable transparency framework requires that distinction. Whether customers must be offered an opt-out depends on the legal basis, contractual arrangement and specific processing; there is no general rule that all improvement logging must be optional while service access remains unchanged.
Fine-Tuning and Model Development
If you offer fine-tuning using a customer's data, it is prudent to define it as a separate scope of processing under a specific agreement, including the purpose, role allocation, access controls, retention, permitted reuse and deletion procedures. AI Act Article 53(1)(d) applies only to providers of general-purpose AI models and requires a public, sufficiently detailed summary of the content used for training in accordance with the AI Office template. The obligations for providers of general-purpose AI models have applied since 2 August 2025, while models placed on the market before that date benefit from the transition rule in Article 111(3). If user interactions are included in training or fine-tuning content, the provider may need to document the relevant data sources, categories, selection and preparation processes and safeguards, depending on its role and the applicable legal and contractual framework. The GDPR continues to apply independently where personal data is processed.
Preparation: Data Classification and Purpose Statements
To support procurement readiness, the technical architecture should be capable of implementing and evidencing the provider's legal and contractual commitments. One useful starting point is data classification.
Classification of Interaction Data
A data map may distinguish between:
- User Content: Prompts and uploaded files (may contain personal data, confidential information or trade secrets).
- System Outputs: The AI's response (may contain personal data, including inaccurate or generated statements about individuals).
- Feedback Data: User ratings or comments (may be personal data and require a defined purpose, retention period and legal assessment).
- Metadata/Logs: Timestamps, account identifiers, IP addresses, token counts and technical events (may be used for billing, security, support or service analytics, depending on the product).
Processing Purpose Statement
A formal 'Purpose Statement' can help explain the processing in customer-facing and internal documentation. It is preferable to avoid vague terms such as 'to improve our services' and instead use language that matches the actual architecture, for example: 'Interaction data is processed to generate AI responses and to monitor security threats. Customer prompts are not used to train or fine-tune general-purpose or shared models unless this has been separately agreed and configured.' The statement would need to be adjusted where prompt content is retained for support, evaluation, safety monitoring or other purposes.
Customer-Facing Documentation
Where the provider acts as a processor, the DPA must contain the terms required by GDPR Article 28(3). Depending on the product and procurement process, this may be supplemented by an AI data addendum, security schedule or product-specific data sheet. Customers may also expect the documentation to identify subprocessors, permitted purposes, retention, deletion, access controls and any use of customer data for model development. Confidentiality and trade-secret concerns can be addressed through clear contractual restrictions and technical safeguards rather than by assuming that a reference to Directive on the Protection of Trade Secrets alone creates the necessary protection.
Trust and Procurement Readiness
EU enterprise procurement readiness often depends on documented alignment between product behaviour, contractual commitments and operational controls. Measures that may support customer reviews include:
- An administrative control that can disable the use of customer data for model training or fine-tuning where such use is not part of the contracted service.
- Tested detection, redaction or filtering controls, where appropriate, to reduce the transmission of personal or confidential data, while accounting for false positives, false negatives and data that cannot be reliably detected.
- An audit trail showing how prompt data is received, routed, accessed, retained and deleted.
Treating AI interaction data as a governed processing activity, rather than as an unrestricted resource for reuse, can make legal and security reviews more predictable and may improve enterprise procurement readiness.
About the author
Junzhe Dai
Junzhe Dai is a PhD candidate at the Faculty of Law, Humboldt University of Berlin. His research focuses on data market regulation, data protection law, and AI governance, with particular interest in the GDPR, the AI Act, the Data Act, and comparative analyses of EU and Chinese digital regulatory frameworks.
Need help with compliance?
Book a free 30-minute call to review your GDPR and EU AI Act readiness.