I run a company that sends your file data to AI models. So read this as testimony from inside the machine: here’s what actually happens to a document when title software calls an AI, and what you should be demanding in writing from anyone who does it — including us.

The mechanics first. When an AI-powered title product summarizes an instrument or drafts a commitment, the document usually doesn’t stay inside your vendor’s walls. It gets transmitted to a model provider — Anthropic, OpenAI, Google, or a model the vendor hosts on cloud infrastructure — processed, and returned. Which means the interesting questions aren’t about your vendor’s marketing page. They’re about the chain of subprocessors underneath it, and the contract governing each link.

The Four Questions

Retention. When the model provider processes your document, how long do they keep it? The honest answer is almost never “zero” and almost never “forever.” Enterprise API terms typically include a bounded retention window — commonly on the order of thirty days — for abuse monitoring, after which inputs are deleted, and some providers offer zero-retention arrangements for qualifying customers. What you want in writing is the actual number, per subprocessor, and whether your vendor took the default or negotiated the stricter option. “We don’t store your data” from a vendor who hasn’t read their model provider’s retention terms isn’t an answer. It’s an aspiration.

Training carve-outs. The one everyone asks about, and the one with the cleanest answer available — if the vendor is on the right product. The major providers’ consumer products may use conversations to improve models, depending on settings. Their enterprise API terms generally exclude customer inputs from training by default. Those are different products under different contracts, and the question to put to your vendor is precise: are you accessing models through commercial agreements that contractually exclude our data from training, and will you attach that language to our contract? If the answer hedges about which tier they’re on, the hedge is your answer.

Subprocessors. The data processing agreement should enumerate everyone who touches file data — model providers, cloud hosts, the OCR service, all of it — and obligate the vendor to notify you before adding one. When a vendor can’t produce this list, it’s usually not because the list is secret. It’s because they never mapped it themselves, which tells you how much diligence went into the pipeline you’re about to feed nonpublic personal information.

SOC 2 scope. This is where diligence theater lives. Plenty of vendors hold a SOC 2 covering their core application while the AI pipeline — the new part, the part shipping your documents to third parties — sits outside the audited boundary. Don’t ask “do you have a SOC 2.” Ask “is the AI processing path inside the scope, and show me the section of the report that says so.” Their model providers all publish their own attestations; your vendor should know exactly where those are.

Why This Is a Title Problem, Not a Generic One

A title file is not ordinary business data. It carries nonpublic personal information under GLBA, which flows down through your underwriter agreements and your Pillar 3 obligations. When your vendor transmits a payoff letter to a model API, you have not outsourced your responsibility for that data — you’ve extended your trust boundary to include the vendor and everything underneath them, and your regulator and your underwriter will treat it exactly that way. Your contracts should too: subprocessor list, no-training clause, stated retention windows, breach notification timelines, data residency. As exhibits. Not as answers a salesperson gave on a call.

One more question worth adding, because almost nobody asks it: what does the vendor send at all? A well-built pipeline sends the model what the task requires — not the whole file, not the escrow ledger, not the buyer’s Social Security number when the job is summarizing a 1974 easement. Data minimization is an architecture decision made early or never, and vendors who made it can describe it in one sentence. Listen for whether you get the sentence or a paragraph.

Our Answers, Since I’m Asking You to Demand Them

It would be cheap to write this post and duck the questions, so: we publish our subprocessor list. We access models exclusively through enterprise agreements that exclude customer data from training. We’ll put retention windows and breach-notification terms in your contract, because a contract is where commitments live. I’d rather compete on being auditable than on being vague — and candidly, an informed buyer asking these four questions is the best sales channel we have, because most of the market can’t answer them.

Scribe reads your instruments and drafts the report with exactly this architecture underneath it. Put the four questions to us. I’d genuinely enjoy the conversation.

Put the four questions to Scribe →