Before an AI tool touches a client matter, a firm needs clear answers to five questions: does the vendor train on your inputs, how long is data retained, where is it physically stored, who at the vendor can access it, and is it encrypted in transit and at rest. These map directly onto the duty to make reasonable efforts to protect confidential information under Rule 1.6 and FLSC Rule 3.3-1 โ the same analysis firms already apply to cloud storage, now applied to AI.
Data security is where AICPA-style due diligence meets professional-conduct rules. This guide gives you the vendor questions, the settings to check, and the jurisdiction-specific residency issues that catch Canadian and US firms. It is part of our AI in Legal Practice library.
Related: AI in Legal Practice ยท AI & Client Confidentiality ยท Writing an AI Use Policy ยท Evaluating AI Research Tools ยท Bar Rules on AI ยท Practice-Area AI
Training on inputs is the first question
The most important data-security fact about any AI tool is whether it uses your inputs to train its models. Consumer ChatGPT and consumer Gemini may, by default, retain conversations and use them to improve the underlying models โ which means anything you paste can influence outputs delivered to other users and sits in a system you do not control. That is a disclosure to a third party under Rule 1.6. Business and enterprise tiers change this: ChatGPT Team and Enterprise, Microsoft Copilot with enterprise data protection, Google Workspace Gemini, and the major legal research platforms contractually commit not to train on your inputs. Get that commitment in writing, and confirm the specific tier and account settings โ the commitment often depends on which plan you are on.
The practical rule most firms adopt: no client-related data in any tool without a written no-training commitment. Our companion guide to client confidentiality and AI tools covers the consent and engagement-letter side.
The vendor due-diligence checklist
Before onboarding any AI tool for matter work, get answers to these in writing:
- Training: Do you train on our inputs or outputs? Under what plan? Can training be disabled at the account level?
- Retention: How long is data retained, is it configurable, and is it deleted on request and on account termination?
- Data residency: In which countries is data stored and processed? Can we require Canadian, EU, or US-only residency?
- Access: Who at the vendor can access our data, under what controls, and are sub-processors disclosed?
- Encryption: Is data encrypted in transit and at rest, and to what standard?
- Certifications: Do you hold SOC 2 Type II, ISO 27001, or equivalent, and will you share the report under NDA?
- Breach: What is your breach-notification commitment and timeline?
- Litigation holds: Can a court order override your ordinary deletion โ as happened in at least one high-profile US matter where a vendor was directed to preserve data it would otherwise have deleted?
Data residency: the Canadian and US wrinkle
Where data physically lives is a live issue, and it differs by jurisdiction. Canadian law societies' cloud-computing guidance โ from the Law Society of British Columbia and the Law Society of Ontario among others โ directs lawyers to understand where client data is stored and whether foreign storage exposes it to another country's access laws, such as US government access under the CLOUD Act when data sits on US servers. Some Canadian firms, and clients in regulated sectors, require Canadian data residency for that reason. US firms face the mirror issue with state privacy statutes and, for firms serving EU or UK clients, GDPR and UK data-protection requirements. Confirm the vendor can meet your residency requirement before you commit, because it is often a plan-level or enterprise-only option. The regulator-side view is in our survey of bar and law society rules on AI.
Settings you actually have to change
A contractual no-training commitment does not configure itself. After onboarding, verify the account: training and model-improvement toggled off, retention set to the shortest period your workflow allows, chat history controls configured, and personal consumer logins blocked for matter work. Centralize on firm-managed accounts so an administrator can see usage and enforce settings โ you cannot answer a client's or regulator's question about what was shared when work happens on individual personal logins. Review the settings after vendor updates, which occasionally reset defaults.
Prompts and outputs are records too
Data security does not end at the vendor contract. Prompts and AI outputs stored in the tool, exported to a file, or pasted into an email are firm records subject to the same protection, retention, and potential discovery as any document. Treat them accordingly: store AI work product in your matter files under normal access controls, and do not leave privileged analysis sitting in a chat interface indefinitely. The habit of treating every prompt as a document that could one day be produced does more for security discipline than any single setting.
Built-in AI in tools you already use
The trickiest data-security exposure is the AI that arrives inside software you already run. Microsoft Copilot, Google Workspace Gemini, and AI features now embedded in practice-management platforms like Clio, document-management systems, and email clients can process client data by default once enabled โ and staff may turn them on without realizing they have introduced a new AI processor. Inventory where AI has appeared in your existing stack, confirm each one's data terms (enterprise tiers of Microsoft and Google generally offer enterprise data protection with no training on your content, but only on the right plan), and disable features you have not vetted. The confidentiality analysis is identical to a standalone tool; the risk is that no one ran it because no one chose to adopt the feature.
The same applies to browser extensions and free plug-ins that promise AI summaries or drafting. Many route your text through servers with consumer-grade or undisclosed data terms. Treat any such extension as an unvetted AI vendor until you have confirmed otherwise, and keep them off machines that touch client matters.
Write it into policy
The due-diligence answers only protect the firm if they are recorded and enforced. Keep a short vendor file documenting each tool's training, retention, residency, and certification answers for your regulator, and fold the data-security rules into your firm AI policy so they bind everyone, not just whoever ran the diligence. The policy walkthrough is in how to write a law firm AI use policy. For help building an AI stack that satisfies both security and growth needs, talk to LexScale.ai.
Frequently Asked Questions
Grow your AI in Legal Practice practice with AI
LexScale.ai builds AI search visibility, websites, and intake systems for ai in legal practice firms across North America. Book a free strategy call to see what would move the needle for your practice.
Book a Free Strategy Call →