Guides / Ownership

Your builder may train on your project. Your users' data is a different clause.

On 5 August 2026, Lovable gave notice that from 9 September 2026 it may use customer content from Free and Pro plans, meaning your prompts and the files you attach, your code, your project files and the outputs, to train its AI models unless you opt out. Two days later it renamed the setting that controls this, and the new label reads in the opposite direction to the old one. That is one platform in one fortnight, and every other builder has its own clause with its own default. This guide is the version we read for clients: what Lovable, Replit, v0, Bolt, Supabase and Firebase each actually say about your project, checked at primary sources this month, the exact setting to look at on each, and why the real exposure is not your database but the spreadsheet you pasted into a prompt at midnight.

Two clauses, two risks. Founders read the wrong one.

"Does the platform train on my data" is two questions wearing one coat, and they have different answers, different consequences and different laws behind them.

  1. Your customer content. The prompts you type, the images and files you attach, your code, your project files, your configurations and the generated output. This is a confidentiality and competitive question about your business, and on most platforms it is the clause that is actually changing.
  2. Your app's end-user data. The rows your customers create by using the thing you built: their names, orders, messages, documents. This is a POPIA and GDPR question in which you are the responsible party and the platform is your operator, and it is the clause founders assume they are reading when they panic about a headline.

Read at primary sources this month, every platform below says the same thing about the second clause: your app's end-user data is not used to train their models. Lovable states it three separate times across its documentation, its change notice and its FAQ. So the honest summary is not "the platforms are taking your users' data". It is narrower and more awkward: the platforms are increasingly training on your content, and your content is where founders keep putting their users' data by hand.

That is the whole risk model in one sentence. A training clause that covers "prompts, including images and files you attach" covers the support transcript you pasted to debug a parser, the CSV export you dropped in to ask why a total was wrong, and the screenshot of a live admin dashboard with fourteen real customers visible in it. Your database is protected by the clause. Your prompt box is not, and your prompt box is the one with a human in front of it at 23:00.

What each platform actually says, read in August 2026.

Quotes are from the vendors' own policies and documentation, with the date each document carries. Terms change: treat this as a worked example of what to look for, and re-read the source before you rely on it.

  1. Lovable: opt-out, and a dated switch. The change summary published 5 August 2026 says that from 9 September 2026 Lovable may use customer content from Free and Pro plans, defined as "your prompts (including images and files you attach), code, project files, and generated outputs", plus usage data, to train and improve its models, relying on legitimate interests. What it will not train on is stated as plainly: "We will not use your users' data ... to train our models", nor your account and billing details. Business and Enterprise workspace data is excluded by default under the organisation's data processing agreement. The same update adds sharing of pseudonymised identifiers (explicitly not your project content, prompts, code or end-user data) with advertising platforms including Meta and Google, by consent in the EEA, UK, Switzerland and Brazil and by opt-out in the United States. The responsible entity is named as Lovable Labs Sweden AB, with the Swedish authority IMY as lead supervisory authority.
  2. Replit: no dated switch, and one trap in the model picker. The privacy policy last updated 3 August 2026 says Replit collects "code, project files, text, commands and prompts, and other inputs that you enter", and claims a legitimate interest in using personal data "to improve the accuracy of our machine learning technologies such as code generation". There is no announced effective date and no toggle named in the policy. The sharper detail is in the AI Integrations documentation, last updated 1 August 2026, which lists the privacy settings Replit applies for individual and team users routing through OpenRouter: "Paid endpoints training disabled", but "Free endpoints training enabled: Free model providers may train on your prompts and completions" and "Free endpoints publishing enabled: Free providers may publish your prompts and completions to public datasets". Picking the free model to save credits is a disclosure decision, not a billing one.
  3. v0 and Vercel: the default flips with the plan. The Vercel AI Policy, last updated 17 March 2026 and effective 31 March 2026, sets out three defaults at once. On a Hobby plan or a trial Pro plan your content is used for model training unless you have opted out; on a paid Pro plan it is not used unless you have opted in; and "we do not use or share data from Enterprise plan customers for Model Training". The control lives in Team Settings, and the policy notes you can also change the outcome by changing subscription.
  4. Bolt and StackBlitz: conditional, undated, unnamed. The privacy policy last updated 12 May 2026 says Bolt may use AI Inputs and AI Outputs "to operate, maintain, and improve the Services, including improving AI performance, reliability, and safety", and that "Such use shall be on aggregated, anonymized, or de-identified data". The opt-out is written as a possibility rather than a right: "Depending on account type or plan, users may have the ability to limit or opt out of the use of AI Inputs and AI Outputs for model training or improvement". If a client contract requires an exclusion, that sentence is not an exclusion. Get it in writing on the account you actually use.
  5. Supabase: your schema travels, your rows do not. The privacy policy's clause on the Supabase AI tool says it collects the content of your query (your inputs or prompts) and the generated output, plus information about the structure of your databases and metadata "such as column and row headings or other information about how that content is organized", and then draws the line explicitly: "We will not access the content of the databases itself or the information you manage through the Service without your consent." Processing for that feature is on a legitimate-interests basis to assess the tool and improve the service, and the service providers named for it include OpenAI, LLC and its affiliates, Amazon Bedrock, Fly.io and AWS. Practical reading: assume table and column names are disclosed, so do not encode secrets in them, and stop assuming the dashboard assistant is a closed loop.
  6. Firebase and Google Cloud: contractual, not a toggle. Google Cloud's Service Specific Terms carry a Training Restriction clause that reads in full: "Google will not use Customer Data to train or fine-tune any AI/ML models without Customer's prior permission or instruction." Firebase sits under those terms. This is the strongest position of the six precisely because it is a contract term rather than a settings page, which is also why it is the easiest one to hand to a client's lawyer.

The pattern across all six is uncomfortable but consistent: enterprise agreements buy exclusion, and free tiers pay for the product with your content. If you build client work on a free or hobby tier, you are the person making a confidentiality decision on your client's behalf, usually without telling them. That is worth one sentence in your proposal and one line in your own notes.

The 9 September date, and the toggle that changed its name.

Lovable's notice was published 5 August 2026 and points opted-out users at Account Settings, then Privacy, then a control called Data collection opt out that you enable. On 7 August 2026 the product moved and renamed it: the account Privacy section became AI model training, and the toggle inside it became On Your Customer Content, where, in the changelog's own words, "enabled means your content may be used to train Lovable's AI models, disabled means it isn't". The workspace-level control moved out of Data protection into its own AI model training section in Settings, then Privacy and security, and is now called On Workspace Data. Previous selections carry over, and the documentation confirms an account-level opt-out applies regardless of the workspace setting.

Read that sequence again, because it is the practical lesson: the label and its polarity both changed inside 48 hours of the legal notice that referenced them. Anyone who remembers "I turned the opt-out on in July" and later sees a toggle described as on for training has no way to reason about it from memory. So:

  1. Open the setting and read the sentence beside it, not the name of the section you remember. Enabled-means-what is the only fact that matters.
  2. Screenshot it with the date visible and file it with your other compliance evidence. This costs ten seconds and is the difference between an answer and a shrug when a client asks in November.
  3. Check both controls if a workspace is involved. Account-level and workspace-level are separate; the account-level opt-out is the one that always applies to your own content.
  4. Diarise 9 September 2026 and re-verify after the change lands. Effective dates are when settings pages get rebuilt.
  5. Do the same pass on every platform you use. Team Settings on Vercel and v0; the model picker on Replit if you route through free endpoints; written confirmation on Bolt, since its policy only offers the possibility of an opt-out.

For client work with a confidentiality obligation, the cheaper answer than toggle-watching is a plan whose data is excluded by default and governed by a data processing agreement. That is what the enterprise tiers are actually selling, and it survives the next rename.

The leak path is your prompt box, not your database.

Lovable's definition of customer content includes "configurations". Replit's includes "commands". Both include attachments. Anything you paste to get unstuck is inside the clause, which makes prompt hygiene an engineering practice rather than a manners lecture. Five rules, in the order they pay off.

One: debug against the schema, not the rows. Almost every "why is this query wrong" question needs table shapes, not real records. Take a structure-only dump and work from that:

supabase db dump --schema-only -f schema.sql

pg_dump --schema-only --no-owner --no-privileges "$DATABASE_URL" > schema.sql

Two: when you genuinely need realistic data, generate it. Build a small masked sample once, keep it in a seed schema, and paste from that forever after. Dropping precision matters more than hashing, because a hashed name plus an exact timestamp plus a suburb is still a person:

create schema if not exists seed;

create table seed.customers as
select
  row_number() over (order by id)                        as id,
  'user' || row_number() over (order by id) || '@example.test' as email,
  'Customer ' || row_number() over (order by id)         as full_name,
  date_trunc('month', created_at)                        as created_at,
  country,
  round(order_total::numeric, -2)                        as order_total
from public.customers
limit 200;

Three: secrets belong in the platform's secret store. Every builder has one, whether it is Lovable's build secrets, Supabase's project secrets or environment variables on the host, and none of them is the chat window. If a key, connection string or service-role token has ever appeared in a prompt, treat it as leaked and rotate it, following the exposed API key runbook. Rotation is cheap; an unrotated key in a training corpus is a permanent unknown.

Four: attachments and screenshots count. A cropped screenshot of one row is a debugging aid. A full-screen capture of your admin table is a personal information transfer with no record, no purpose and no legal basis. Crop, redact, or reproduce the problem against the seed data from rule two.

Five: understand that opting out is forward-looking. Lovable's own wording is that opting out later excludes your data "from training going forward". There is no retraction button anywhere in this market. So the response to a bad paste is the same as the response to any disclosure: rotate what can be rotated, record what was exposed and when, and assess whether anyone needs to be told. If personal information was involved, our POPIA guide covers the notification test. Keeping the agent's blast radius small in the first place is the guardrails guide.

Write down what you checked. That is the compliance artefact.

Because no vendor here claims your end-user rows for training, building on Lovable or Replit does not by itself put an "we train AI on your data" line in your privacy notice. What it does put on you is the duty to know that, and to show how you know it. POPIA's framing is unchanged by any of this: the platforms are operators, you remain the responsible party, and the paperwork is yours.

One short record per platform is enough, and it takes about fifteen minutes for a typical stack:

  1. Which legal entity you contract with and where it sits. For Lovable, the change notice names Lovable Labs Sweden AB in Stockholm.
  2. Which document you relied on, and its date. A privacy policy "last updated 3 August 2026" is a citable fact; "their website says" is not.
  3. The state of each setting and the date you read it, with the screenshot from the previous section attached.
  4. Whether end-user data is excluded from training, quoted, since this is the question a client or a data subject will actually ask.
  5. Who re-checks it and when. Terms move under you: this fortnight alone produced a new Lovable privacy policy, a renamed control, a Replit policy update and a self-hosted Supabase gateway change. A calendar entry every quarter, plus one on any announced effective date, is the whole control.

If the answer you need is "none of this applies to my data", the route is ownership rather than settings: your app on infrastructure you control, and if control is a hard requirement, self-hosted Supabase with its costs stated honestly. Most businesses do not need that. They need to know which clause applies to whose data, to have checked it on a date they can name, and to stop pasting customer spreadsheets into a chat box. The rest of the guides library covers the neighbouring gaps.

Questions

Questions founders ask.

Does Lovable train on my app's users' data?

No, on its own account. Lovable's change summary published 5 August 2026 states that it will not use your users' data, meaning the information your apps' visitors or customers provide to the apps you build, to train its models, and that this data stays in your project's own database and storage. What may be used from 9 September 2026, unless you opt out, is customer content on Free and Pro plans: your prompts including images and files you attach, your code, project files, configurations and generated outputs, plus usage data. Business and Enterprise workspace data is excluded from training by default under the organisation's data processing agreement. The practical exposure is therefore anything you personally pasted into a prompt, which is where your users' data ends up if you debug with real records.

I opted out of Lovable's AI training in July. Am I still opted out?

Probably, but verify rather than assume, because the control was renamed and inverted two days after the legal notice referred to it. Lovable's 7 August 2026 changelog says the account-level Privacy section became AI model training and the Data collection opt out toggle became On Your Customer Content, where enabled means your content may be used for training and disabled means it is not. Previous selections carry over, but a founder who remembers enabling an opt-out will now see a toggle whose name means the opposite. Open Account settings, find the AI model training section, read the sentence next to the toggle, and screenshot it with the date. If your projects live in a workspace, check the workspace-level control too, noting that an account-level opt-out applies regardless of the workspace setting.

Is it safe to do client work on a free or hobby plan?

It is a decision you are making on your client's behalf, so make it deliberately. The defaults are not uniform: Vercel's AI Policy uses your content for model training on Hobby and trial Pro plans unless you opt out, while paid Pro requires opt-in and Enterprise data is never used; Lovable will train on Free and Pro customer content from 9 September 2026 unless you opt out, with Business and Enterprise excluded by default; and Replit's own documentation says free OpenRouter endpoints may train on your prompts and completions and may publish them to public datasets. If your engagement carries a confidentiality clause, the cheapest defensible answer is a paid tier whose exclusion is contractual and a one-page record of the setting states with dates, not a toggle you hope is still where you left it.

Does using Supabase AI expose my customers' data?

It exposes your database's shape, not its contents. Supabase's privacy policy says the AI tool collects the content of your query and the generated output, along with information about the structure of your databases and metadata such as column and row headings or other information about how that content is organized, and states that Supabase will not access the content of the databases itself or the information you manage through the service without your consent. The service providers named for that flow include OpenAI, LLC and its affiliates, Amazon Bedrock, Fly.io and AWS. Two practical consequences: do not encode anything sensitive in table or column names, and remember that whatever you type into the assistant, including a query with real values in the WHERE clause, is your input rather than protected database content.

Do I have to tell my users that I build on an AI platform?

You have to tell them who processes their personal information and where, which in practice means naming your hosting and database providers as operators and disclosing overseas processing, exactly as POPIA's section 72 conditions expect. You do not need to tell them their data trains an AI model if it does not, and on every builder we checked in August 2026 end-user data is excluded from model training. What you should be able to produce on request is the reasoning: which document you read, its date, and the state of the training settings on your account. That record is also the fastest way to answer the same question when it arrives from a client's procurement team rather than a user.

Ownership, not just settings

Know which clause applies to whose data, before a client asks.

A production-readiness review covers the ownership layer as well as the code: which platforms hold your data and under what terms, whether secrets have ever passed through a prompt or a repository, whether your users' personal information is where you think it is, and what moving onto infrastructure you own would actually take. Ranked by the engineers who would do the work.