PepoChat
GDPRPrivacyCompliance

GDPR and AI Chatbots: What Data a Support Bot Stores and How to Stay Compliant

A non-lawyer's guide to chatbot GDPR compliance: what a support bot stores, controller vs processor roles, retention, sub-processors and a vendor checklist.

PepoChat TeamPublished Last verified 11 min read
Brass combination padlock and gold payment cards resting on a white computer keyboard

Short answer

A support chatbot is GDPR compliant when you, the website owner, can answer six questions: what personal data it stores (transcripts, emails, page origin, order data), on which lawful basis, for how long, who else touches it (the vendor and its sub-processors, including the AI model provider), where it is processed, and how you honour access and deletion requests. Under GDPR you are the controller and the chatbot vendor is your processor, so you need a written data processing agreement with them and a current sub-processor list.

"Is my chatbot GDPR compliant?" is the wrong question; the right one is are we. The General Data Protection Regulation (GDPR) applies to the organisation that puts a bot on its website and to the vendors that run it on that organisation's behalf. This guide covers what a support bot actually collects, who is responsible for what, and which questions to put to any vendor before you paste their script tag.

It is written for founders, support leads and marketing managers who own a website with EU or UK visitors and do not have a lawyer on call. It is not legal advice; if you handle health, financial or children's data, talk to a data protection professional. The examples reference PepoChat, an AI support agent for websites, but the checklist applies to every vendor in the category.

What personal data does a support chatbot actually process?

Personal data is, in the GDPR's own words, "any information relating to an identified or identifiable natural person". That definition is deliberately wide, and a support bot collects more of it than most people assume. A typical deployment touches these categories:

  • Transcripts. Every message a visitor types is stored so the bot can keep context and your team can read the conversation later. Visitors paste order numbers, addresses, health details and, occasionally, card numbers into chat boxes. Whatever they type becomes data you hold.
  • Contact details. A name and email address, collected through a pre-chat form or when the visitor verifies their email to get order information.
  • Technical identifiers. The page the visitor was on, the origin domain the widget reports, a session identifier, and usually the IP address in the hosting provider's server logs. An IP address counts as personal data under GDPR.
  • Data pulled through actions. When the bot looks up an order in Shopify, a subscription in Stripe or a ticket in Zendesk, the response flows through the chatbot vendor and often through the AI model provider. Your customer's purchase history has just been copied into two more systems.

Is my chatbot GDPR compliant? First settle the roles

GDPR splits responsibility between two roles. The controller is the organisation that "determines the purposes and means of the processing of personal data", and the processor is the one that "processes personal data on behalf of the controller". The ICO's plain-English explanation uses a gym that hires a printing company: the gym decides why member addresses are used and is the controller, the printer follows instructions and is the processor.

For a support chatbot, you are the controller. You chose to put the widget on your site, you decided it should answer order questions, and you decide how long transcripts are kept. The chatbot vendor is your processor, and every service the vendor uses to deliver the product (its database host, its AI model provider, its email service) is a sub-processor, meaning a processor engaged by your processor.

That role assignment has a practical consequence: Article 28 GDPR requires a written contract between you and your processor. A controller must use only processors "providing sufficient guarantees to implement appropriate technical and organisational measures", and the article lists what the contract must contain. The processor must act "only on documented instructions from the controller", keep its staff bound to confidentiality, apply the security measures of Article 32, help you answer data subject requests, delete or return the data when the service ends, and make available "all information necessary to demonstrate compliance".

That contract is what vendors call a data processing agreement (DPA). If a vendor cannot show you one, you cannot lawfully use them for personal data. Article 28 also says a processor "shall not engage another processor without prior specific or general written authorisation of the controller" and must tell the controller about changes so it can object, which is why serious vendors publish a sub-processor list.

Which lawful basis covers a support chatbot?

Every processing activity needs one of the six lawful bases in Article 6. For a support bot, two matter.

Contract (Article 6(1)(b)) covers processing that is "necessary for the performance of a contract to which the data subject is party or in order to take steps at the request of the data subject prior to entering into a contract". A customer asking where their order is, or a prospect asking what your product costs, fits comfortably.

Legitimate interests (Article 6(1)(f)) covers processing you need for a business purpose that does not override the visitor's rights. Keeping transcripts to improve your knowledge base, or producing weekly analytics of what visitors ask about, usually rests here. You must document the balancing test: what you gain, what the visitor risks, and why the first outweighs the second.

Consent is rarely the right basis for the core chat. Consent must be freely given and withdrawable, which is awkward when the visitor has already typed their question. Where consent belongs is for anything optional layered on top: marketing follow-ups, a newsletter subscription action, or any tracking cookie the widget might set.

Transparency: tell people it is a bot and what you keep

Article 5(1)(a) requires that data be "processed lawfully, fairly and in a transparent manner", and Article 13 spells out what visitors must be told when you collect data from them directly: who you are, why you process the data and on which basis, who receives it, whether it leaves the EU, how long you keep it, and their rights. That information belongs in your privacy notice, and the widget should be one click away from it.

There is a second transparency rule that is specific to AI. Article 50 of the EU AI Act requires that AI systems "intended to interact directly with natural persons" are designed so that people "are informed that they are interacting with an AI system", unless that is obvious from the context. The Article 50 text applies from 2 August 2026, so as of September 2026 it is live. The obligation sits with the system's provider, but as the business putting the bot in front of customers you should make sure it is honoured: a greeting like "Hi, I'm an AI assistant. Ask me anything or type 'human' to reach the team" does the job.

Transparency also covers hand-offs: when a person takes over, the visitor should see that a human is now replying. The human handoff guide shows how the switch should look.

Data minimisation: collect less, and stop the bot collecting more

Article 5(1)(c) says data must be "adequate, relevant and limited to what is necessary in relation to the purposes for which they are processed". The full list of Article 5 principles is short and worth reading once; minimisation is the one most chatbot deployments fail.

Three places to apply it:

  1. No pre-chat form unless you need one. Asking for name, email and phone number before the visitor can type a question collects data you often never use. Let chats start anonymous and ask for an email only when the conversation needs it, such as to look up an order or send a follow-up.
  2. Limit what actions return. When the bot calls your store or CRM, the raw API response often contains far more than the answer requires. A good vendor lets you allow-list which response fields the AI may see, so an order-status lookup returns the status and delivery date, not the customer's full record. The Shopify order status guide shows how that looks in practice.
  3. Keep the knowledge base clean. The documents you upload are processed data too. Internal spreadsheets with customer names, or support tickets with real emails, do not belong in a knowledge base that a public bot answers from. If you want to train on past tickets, anonymise them first.

How long should chatbot transcripts be kept?

Article 5(1)(e), the storage limitation principle, says data may be "kept in a form which permits identification of data subjects for no longer than is necessary". GDPR does not give you a number; you set one and justify it.

A reasonable default for a support bot is to keep resolved transcripts for as long as a dispute could plausibly arise, often 6 to 24 months depending on your product, and to keep anonymised statistics indefinitely. Sessions that never became a conversation (the visitor opened the widget and left) should disappear far sooner, since there is no purpose in keeping them.

Check two things with your vendor: whether you can set a retention period yourself, and how long deleted data survives in backups. A 30-day backup window is normal, but it belongs in your records.

Retention also applies to the AI model provider. OpenAI's API documentation states that "data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us)", and that "abuse monitoring logs are generated for all API feature usage and retained for up to 30 days, unless longer retention is required by law". The same OpenAI data controls page describes a zero-data-retention option and regional storage for eligible customers. Whether your chatbot vendor has any of those enabled is a question only they can answer.

Subject access and deletion requests when the data sits in a vendor's system

Article 15 gives visitors the right to a copy of their data, and Article 17 the right to have it erased "without undue delay" when, among other grounds, the data is "no longer necessary in relation to the purposes for which they were collected" or consent has been withdrawn. Article 12 gives you one month to respond, extendable for complex requests.

The operational problem is that chatbot data is spread across systems. To honour a request you need to be able to:

  • find every conversation attached to a given email address or verified identity;
  • export the transcript in a readable form;
  • delete it, including from the vendor's system, and know when backups age out;
  • know whether copies exist downstream, such as a ticket the bot created in Zendesk or a contact it filed in HubSpot.

Anonymous chats are harder to link to a person, which cuts both ways: less risk, but a visitor who asks you to delete "everything I told your bot last Tuesday" may be impossible to identify. That is not a failure as long as you explain it; GDPR does not require you to collect more data just to be able to delete it later.

Ask your vendor how search-by-email works in their inbox, whether they can delete a single conversation, and whether deleting your account removes everything. Article 28 requires the processor to assist with these requests.

Close-up of a row of server hard drives in a rack, their green activity lights blinking
Every chat message passes through at least three companies: the chatbot vendor, its database host, and the AI model provider. Each one is a sub-processor you need to know about.

Sub-processors: the AI model provider is part of your chain

A modern support bot is not one company's software. The vendor writes the application, but it runs on a cloud host, stores data in a managed database, sends email through an API and, critically, sends every message to a large language model run by OpenAI or a similar provider. Each is a sub-processor, and under Article 28(4) "the initial processor shall remain fully liable to the controller for the performance of that other processor's obligations".

You do not sign contracts with the model provider yourself, but you are entitled to know what it receives and how long it keeps it. Three questions capture most of the risk:

  1. Is the model provider's no-training default in place? OpenAI's API does not train on customer data by default, as quoted above; consumer chat products often do. Ask which product your vendor is built on, and check the equivalent page for any other provider.
  2. What exactly is sent to the model? Usually the visitor's message, the retrieved knowledge-base passages and the recent conversation history. If actions are involved, the API response is sent too, which is why field allow-listing matters.
  3. How is the prompt protected from injection? A web page in your knowledge base could contain instructions that trick the bot into revealing another visitor's details. Ask how retrieved content is separated from instructions.

If the widget offers voice conversations through a voice-AI platform, audio and transcripts go to that platform under its own terms, so it belongs on the list as well.

International transfers: where does the data go?

GDPR restricts transfers of personal data outside the EU and EEA unless a safeguard applies. The simplest safeguard is an adequacy decision, a formal finding by the European Commission that a country's protection is adequate, after which "personal data can flow from the EU (and Norway, Liechtenstein and Iceland) to that third country without any further safeguard being necessary". The Commission's adequacy decisions page lists the countries covered, including the United Kingdom, Japan, Switzerland and, for "commercial organisations participating in the EU-US Data Privacy Framework", the United States.

Most chatbot vendors and model providers are American, so the question is whether each company on the sub-processor list is certified under the Data Privacy Framework or, failing that, whether the vendor's DPA includes the EU Standard Contractual Clauses. EU data residency for the database rarely covers the model calls, so ask about both separately.

A hand ticking boxes on a handwritten checklist on a tablet with a stylus
Vendor due diligence is a short list of documents to request and a shorter list of questions to ask before the widget goes live.

The buyer's checklist: questions to ask any chatbot vendor

Run every vendor, PepoChat included, through this table before you sign up. A good vendor answers each row in a sentence with a link.

QuestionWhy it mattersA good answer looks like
Will you sign a data processing agreement?Article 28 requires a written contract before any personal data is processedA standard DPA you can accept online, with the Article 28(3) terms and a transfer mechanism
Can I see your sub-processor list?You must authorise sub-processors and be told when they changeA public page naming host, database, model provider, email and payment vendors, with a change-notification option
What is sent to the AI model and is it used for training?The model provider is the most sensitive link in the chainNamed provider, API with no-training default, stated retention of prompts
Can I set retention for transcripts and sessions?Storage limitation is your obligation, and you need the controls to meet itConfigurable retention, plus a documented default and a clear rule for abandoned sessions
How is data encrypted?Article 32 security measuresTLS in transit, encryption at rest, and separately encrypted credentials for integrations
Is my workspace isolated from other customers?A multi-tenant bug can expose one company's chats to anotherPer-organisation isolation of knowledge base, conversations and credentials, enforced in the backend
Can I delete a single conversation, and my whole account?Erasure requests and offboardingDelete from the inbox for one, account deletion for all, with a stated backup window
Does the widget set cookies on my site?Cookie consent rules apply on top of GDPRNo tracking cookies on the host domain; identity kept inside the widget's own iframe
Where is data hosted, and where do the model calls go?International transfersA stated region for storage, plus the model provider's transfer basis
How do you limit what the bot sees from my integrations?Data minimisation for actionsA per-action allow-list of response fields

Keep the vendor's answers with your records of processing. Under Article 5(2) the controller must "be able to demonstrate compliance", and a dated email from the vendor is the cheapest evidence there is.

What PepoChat does and does not tell you today

Here is what can be stated about PepoChat from the product itself, and what you should still ask for.

Isolation. Each organisation's knowledge base, conversations and integration credentials are isolated from every other organisation's.

Credentials and actions. API keys for integrations such as Stripe, Shopify or Zendesk are encrypted per organisation with AES-256-GCM and cannot be read back from the dashboard once saved. Custom actions carry an allow-list of the response fields the AI may see, which is the minimisation control described above. Outbound calls are HTTPS-only and block private and internal hosts.

Visitor identity. Chats start anonymous; there is no pre-chat form. A visitor can choose to verify their email with a six-digit one-time code, which is what unlocks order lookups by email and, if you enable it, meeting memory for returning customers. Visitor sessions last 24 hours, and expired sessions that never became a conversation are deleted after a further 24 hours. The widget sets no tracking cookies on your site.

Prompt safety. Retrieved knowledge-base content is wrapped as untrusted data in the prompt to resist prompt injection from web pages, and messages are capped at 4,000 characters.

Sub-processors. The service runs on Convex (backend and database), Vercel (frontend hosting), OpenAI (the language models: gpt-4o-mini for replies and gpt-4o for document extraction), Resend (transactional email), Vapi (optional voice, using your own Vapi account) and Dodo Payments (billing).

What this page deliberately does not claim: a hosting region, a specific DPA, or any SOC 2, ISO or GDPR certification. Those are exactly the items you should request in writing from PepoChat, as from any vendor, before you go live. Use the contact page to ask for the DPA and the current sub-processor list, and evaluate the answers against the checklist above.

What to do next

Start with the paperwork, not the widget: pick your lawful bases, set a retention period, write both into your privacy notice, then request the DPA and sub-processor list from whichever vendor you are evaluating. After that the technical setup takes an afternoon; the one-script-tag install guide covers the widget, and the pricing page confirms that every feature, including email verification and the action allow-lists, is on the free plan. To see the controls first, create a free workspace and open the Widget Customization and Actions pages.

Frequently asked questions

Is my chatbot GDPR compliant?
A chatbot itself is not compliant or non-compliant; your deployment is. You are compliant when you have a lawful basis for the chat data, a privacy notice that explains what is stored and for how long, a signed data processing agreement with the vendor, an authorised sub-processor list, a retention period you actually enforce, and a way to answer access and deletion requests.
What personal data does a support chatbot store?
Typically the full transcript of every conversation, any name or email the visitor provides or verifies, the page and origin domain the widget was opened from, a session identifier, and any customer data the bot pulls from connected systems such as an order status or a subscription. Free-text messages can also contain sensitive details the visitor volunteers.
Am I the controller or the processor for my chatbot?
You are the controller, because you decide why the bot exists and what it does with the data. The chatbot vendor is your processor, and the services it relies on, including the AI model provider, are sub-processors. Article 28 GDPR requires a written contract, usually called a data processing agreement, between you and the vendor.
Do I need consent to run an AI chatbot on my website?
Usually not for the chat itself. Answering a customer's question rests on the contract basis, and keeping transcripts for quality or knowledge-base improvement typically rests on legitimate interests. Consent is needed for optional extras such as marketing follow-ups or tracking cookies. Whatever basis you choose must be stated in your privacy notice.
How long can I keep chatbot transcripts under GDPR?
GDPR gives no fixed number; the storage limitation principle says only as long as necessary for your stated purpose. Many businesses keep resolved transcripts for six to twenty-four months in case of disputes and keep anonymised statistics longer. Sessions that never became a conversation should be deleted quickly. Write the period down and make sure your vendor can enforce it.
Does the AI model provider count as a sub-processor?
Yes. Every message your bot handles is sent to a language model run by a third party such as OpenAI, so that provider processes personal data on your behalf through your vendor. It must appear on the vendor's sub-processor list, and you should know whether the provider trains on the data, how long it retains prompts, and on what basis it receives EU data.

Try this on your own site in ten minutes

PepoChat includes every feature on the free plan — 500 AI replies and 10 knowledge sources a month, no credit card.