The question comes up at the same point in every training session we run: someone connects their mail, sees that the AI assistant can read messages from the past week, and stops to ask whether they are allowed to work this way.
We can’t settle the legal side for you. We can show you the decisions you can make right now: how to narrow the AI assistant’s access to files, how long to keep your conversations, and when masking data is worth the effort. At the end, we list the points worth taking to a lawyer or a data protection officer.
What leaves your computer, and what stays on it
Claude Code runs on your machine, but it talks to the model over the network. Your prompts and the model’s replies go to the provider, and so does everything the AI assistant read along the way: part of an open file, the output of a command, an email you pasted in. Whether something came from a file or from the chat makes no difference to what gets transmitted. If Claude Code read it, treat it as having left the building, whether or not it turned up in the answer.
The other half of this is easier to miss: some of that data stays with you. Claude Code writes session transcripts to ~/.claude/projects as plain text. Fragments of files and command output that scrolled past in the conversation can end up in those transcripts. The default retention is 30 days, and cleanupPeriodDays changes it. Those files are not encrypted – your operating system’s permissions are the only protection, so check who else can reach that user account and your backups.
Two features can send more than an ordinary conversation does. The /feedback command sends Anthropic a copy of your session history along with your code, and those reports are kept for five years. You may also see the question “How is Claude doing this session?”. The rating itself sends no conversation content. A follow-up may then ask for permission to look at the transcript, and once you grant it, the whole session log is kept for up to six months.
You can also narrow what the AI assistant reads in the first place. Rather than watching for it in every conversation, ask once:
Deny yourself access to the folder with client documents in this project permanently, so that you can neither read nor change it. First tell me which rule you intend to write and in which file, and change nothing until I approve it.
Claude Code should suggest a rule that denies reads on the path you named, something like Read(./clients/**), and tell you where it will save it. Before you approve, check the scope of the path and the kind of rule.
That protection has a limit. The rule covers Claude Code’s own tools, including file reads and commands such as cat or head. It will not stop a script that opens the file itself while running. For a harder boundary there is the sandbox, where the operating system polices access to files rather than Claude Code’s configuration.
A rule like that is part of your permission settings. The mode decides when the AI assistant stops to ask you; the rules decide what goes through without asking. For help choosing a mode, read the tip on permission modes.
Set how long your conversations are kept
On the Free, Pro and Max plans, a single setting decides two things at once. It is called Help Improve our AI models and it sits in your account settings, under Privacy. The claude.ai interface changes from time to time, so look for the name of the setting rather than a particular spot on the page.
When it is on, your conversations and Claude Code sessions may be used to improve future models, and Anthropic keeps them for five years. Switch it off and that drops to 30 days. The change only applies going forward: data already in a training run or baked into a trained model cannot be pulled back this way. Anthropic does stop using your previously stored conversations in later rounds of model improvement.
There is one exception. Conversations flagged by Anthropic’s safety classifiers may still be used to develop safety models, detect harmful content and enforce the usage policies.
Masking: the model gets labels instead of data
Masking can cut down how much data reaches the model at all. A script on your own computer goes through the text and swaps chosen items for labels. A VAT number becomes [TAXID_2], a first name becomes [PERSON_1]. The substitution dictionary stays on your disk. The model receives the text with labels in it, and once the answer comes back the same script can put the originals in again.
The important part is that detection happens locally, in the script, and not in the model in the cloud. A model asked to find personal data would have to receive it first. For the same reason, the dictionary must never go into the conversation or into a file you send to the model.
You don’t have to write this yourself:
I want to work with you on texts where personal data has been masked. Write me two scripts: the first takes a file, swaps personal data in it for labels and saves the substitution dictionary on my disk; the second takes a finished text with labels and puts the originals back. Propose first which data types you will cover and where you will save the dictionary, and wait for my approval.
Claude Code should then give you a list of data types to detect and a location for the dictionary. It should create files only after you approve. In practice, you ask the first script to run on a given file, and only the labelled text goes into the conversation. The original and the dictionary stay on your disk.
Go through that list of data types carefully, because this is where the method reaches its limit.
It works best with data that has a recognisable format and can be checked: email addresses, phone numbers, tax and company registration numbers, bank account and card numbers. Even in those cases, the script needs testing on your own documents, because a pattern can miss an unusual way of writing something or flag a string of digits that is not personal data at all.
Names, company names, postal addresses and anything that only becomes clear from context are much harder. They have no single format that a simple script can reliably recognise. A company name can itself be personal data, particularly when it identifies a sole trader. If the sensitive part of a message is who wrote it and which company they work for, pattern-based masking alone will not cover it.
Masking of this kind is pseudonymisation, not anonymisation. It lowers the risk without relieving you of the judgement about which data may be sent. Labelled text has the same retention period as any other conversation. The difference is that the provider’s servers hold the labels, not the data in your dictionary.
A local model can widen what gets masked
A local model can take this further. It runs on your computer and can use context, so it may recognise the names, company names and addresses a simple pattern misses. It can prepare the substitution dictionary, hand Claude the labelled text, and put the originals back into the answer afterwards.
Local means your computer does the computing. The data does not go to an external model provider, as long as the tool does not also call a network service. A local model makes mistakes too. It can miss data, and it can label something as personal data when it isn’t. Try it on a representative, controlled sample of documents before you rely on it.
Which model you can run, and how fast it works, comes down to your hardware. The same text may take seconds or minutes depending on your available memory, graphics memory and the size of the model.
You can hand the hardware question over as well:
Check this computer's specifications and tell me whether it could run a local language model for recognising personal data in text. Say which models would run here. Install nothing until I approve it.
We haven’t found a ready-made tool that does all this reliably, especially for languages other than English. Models that recognise names exist, but building a safe workflow around them – detection, the dictionary, restoring the data and keeping the grammar right – means building and testing it yourself. Think of it as something to experiment with, not something to roll out next week.
What to check on the legal side
None of this is individual legal advice. In general, the business that decides what client data is used for and how it is used is the controller of that data. The AI provider may process it on that business’s behalf. That distinction affects your legal basis, your processing agreement and how you assess transfers abroad.
Anthropic offers a data processing addendum under its commercial terms, which cover the Team and Enterprise plans and API access among others. The Pro and Max plans do not include it. Before you send client data anywhere, check the current terms for your own plan, who holds which role, and where and under what rules the data is processed. If you are in the EU or the UK and data goes to the United States, the transfer needs a basis of its own – a provider’s current status page is not a substitute for that assessment.
The EU’s AI Act does not impose one uniform set of obligations on every small business using an off-the-shelf assistant. How the tool is used is what counts, and the cases that call for real care are the ones where a system feeds into decisions about people, hiring being the obvious one. If you want to work through your own case, the European Data Protection Board collects its guidance on AI in one place, and most national authorities publish their own checklists.
In short
Start with your account settings: check whether Help Improve our AI models is on, and choose how long to keep the conversations. Then point the AI assistant at the folders it should stay out of, and reach for the sandbox when you need a harder boundary.
You can mask numbers, email addresses and other fixed-format data with a local script – but test it on your own documents first. With names, company names and anything that follows from context, don’t assume the detection will be complete.
If client data is going to reach the AI assistant regularly, work out your legal basis, the processing terms and whether you need a processing agreement. If the data shouldn’t leave your own environment at all, don’t send it to a cloud service.