A regulated company can use a provider's API under contract, its own servers, or a hybrid that splits documents by class. A hybrid needs a routing table with an owner and a log of each document's route.
A regulated company can run LLM document processing through a provider's API under contract, on its own servers, or as a hybrid of the two. Use the API if every class of document may leave under contract, and self-host if none may. A hybrid pays off only when many documents fall on each side, because you run two systems and own the routing table. To tell which firms have taken such a system to production, ask for the records it leaves behind.
The short answer: write down your document classes and where each may go. If you choose a hybrid, give its routing table an owner and a version number, and log the route of every document. To check a firm, ask to see its route log and to speak with whoever runs the system today.
Read next
The closed-perimeter article linked above compares them in depth. With a provider's API, your data is covered by the contract and by the data controls the provider agrees to. OpenAI's data controls page, read on 8 October 2026, says data sent to its API since 1 March 2023 is not used for training unless you opt in.
The same page says abuse monitoring logs, which OpenAI uses to detect misuse and which may contain prompts and responses, are kept for up to 30 days by default. They are kept longer if the law requires it or if that is reasonably necessary to protect OpenAI's services or others from harm. Leaving your content out of these logs needs OpenAI's prior approval.
In a cloud platform's managed model service, the cloud company runs the model for you, and some of its models are sold under a cloud contract you may already hold. Microsoft's documentation says that, for models sold by Azure, your prompts and the model's answers are not available to OpenAI or other providers of those models. Amazon Bedrock's documentation says model providers have no access to customers' prompts and the model's answers. The model still runs on the cloud's servers, and your risk team decides whether that counts as inside your perimeter, the network and systems your company controls.
A self-hosted open-weight model, one its maker publishes for anyone to download and run, works on servers you control, so no document goes to a model provider. In return, running the servers and the model is your team's job.
A hybrid puts an outside route (a provider's API or a cloud managed service) and an inside route into one pipeline. The logic that decides where each document goes is called the router. There are three basic designs, and they can be combined:
Risk or data protection decides, and engineering writes the decision into code. Put it on paper first as a short table, called the routing table here. For each document class, it lists the routes allowed, whether masking is required and how long the provider may keep the text.
The pipeline then has to work out each document's class. Signals you control are safer than a model's judgement: the channel it came through, the sender, the document type. When signals disagree, the stricter class wins, and a document nobody can classify stays inside.
This check runs inside the perimeter. A check that asks the outside API whether a document may leave has already sent it. A change to the routing table goes through review and gets a version number and a date, like any code change.
Sometimes, within limits. Open-source tools such as Presidio replace names, account numbers and other identifiers with placeholders or encrypt them with a key. If you keep the key inside, Presidio's decrypt step puts the real values back when the answer returns; with placeholders, you keep your own table of which placeholder stands for which value. Data masked this way is called pseudonymised.
The European Data Protection Board adopted Guidelines 01/2025 on pseudonymisation on 16 January 2025, as a version for public consultation. The guidelines say such data still counts as personal data if additional information could link it to a person. They add that this holds even when the masked text and that information sit with different parties, as when the provider holds the text and you hold the table or the key.
The EU Court of Justice took up the same question in September 2025, in case C-413/23 P under the data protection rules for EU bodies. It held that pseudonymised data is not personal data in every case and for every person: depending on the circumstances, masking can stop anyone other than the company that masked the data from identifying the people in it. For you, holding the table or the key, the data stays personal. Ask your data protection counsel what follows for your contracts.
Detection is the second limit. Presidio finds personal data with a model that spots names of people, places and organisations, and with rules that match known formats such as account numbers. Its documentation says that, because detection is automated, “there is no guarantee that Presidio will find all sensitive information”. Typical misses in document work:
Before anyone approves a masked route, have people mark every identifier by hand in a sample of your own documents, then count what the detector missed.
In a hybrid, a leak is a routing bug. A scanned passport attached to a routine notice, or a client's message quoted at the bottom of a forwarded email, can carry restricted content onto the outside route. Classify each attachment on its own, and treat the quoted history in a thread as part of the content.
Build the system so that a wrong decision stops the document instead of sending it out. The closed-perimeter article linked above covers the network side: the inside route has no keys or passwords for the outside provider and no network path to it. Add one rule: when the self-hosted model is down, its queue waits or goes to people, and never switches to the API.
If a document does go out by mistake, the route log shows which documents left, when and to which provider. The provider's contract shows how long it may keep them. Treat it as an incident with your data protection team, who decide whether it must be reported.
An acceptance set is a group of real documents with the correct answer for each, used to decide whether the system passes. The two models in a hybrid answer differently, so one score for the whole pipeline hides the weaker route.
You also cannot test the API model on restricted documents, because they may not go there. So the shared acceptance set that the closed-perimeter article recommends for every route can only hold documents allowed on both routes. Use it to compare the two models, and give each route its own larger set drawn from the classes it handles. When the provider retires the API model, that route is scored again before the switch.
A hybrid doubles the contracts: the provider's terms and data processing agreement, plus the hardware or cloud contract under the self-hosted model. The European Banking Authority notes that DORA, the EU's Digital Operational Resilience Act, has applied since 17 January 2025. EU financial entities in its scope must keep a register of their contractual arrangements with ICT (information and communication technology) third-party service providers. Adding an outside model provider is a question for that register, which your compliance team answers.
Operations double too. The provider caps how many requests you may send per minute and retires models on its own schedule, so someone has to watch both. Your own servers need capacity planning and someone on call at night. The router and the masking layer are yours alone to run.
Watch the share of documents on each route, by class, every day. When the self-hosted model sees every document first, a rising share sent to the API means more documents leave and the API bill grows. Often the cause is a new document template that the internal model cannot read.
Store its class and the signals that set it, the routing table version and the route taken. Add whether it was masked and by which detector version, and the provider and exact model version that answered. With that record, “show every document of this class that left the perimeter last quarter” takes one database query.
One route is enough when all your documents fall into one data class. With a small volume, a second route costs more than it saves, and a person can handle what the self-hosted model cannot read. A hybrid also inherits any gap in night cover for the self-hosted servers. And when a cloud managed service in your region meets the rules for every class, it gives you one contract and one set of operations.
This article ranks no firms, because any firm can write “production AI” on its website. A system in production leaves records that a pilot does not, so ask for these, with client details removed.
The one language model project amBrain describes in public is this: “We have taken an LLM integration to production inside a client's FinTech perimeter: extracting and normalising unstructured broker and venue notices — corporate actions, instrument and margin changes — into structured records the trading system consumes.”
This article is not a case study: the client is not named, and where the model ran in that project is not disclosed. It does not claim that amBrain has built a hybrid setup, a masking layer or a router for any client, and it gives no prices or timelines.
amBrain takes over projects that stalled with another team and brings them to production.
amBrain has been building software since 2019. It works in three formats: full delivery, a dedicated team, or engineers embedded in your team. The client keeps full ownership of the product and the code, except amBrain's reusable components.
If you are weighing these options, bring your document classes and the list of records above to every firm you talk to, amBrain included.
Bring your current architecture and the failure mode that worries you, and we will go through it together in half an hour.