amBrain
FinTechOct 1, 202610 min read

A Private ChatGPT for Your Company on Your Own Servers: What It Takes and Who Can Set It Up

Private ChatGPTAI on Your Own ServersFrom Pilot to ProductionWho Builds It
Error loading image

You can give staff a ChatGPT-style assistant without client data leaving the company. Before daily use, whoever sets it up should show you the assistant following each person's document permissions and keeping chats and logs for as long as your compliance team decides.

A private ChatGPT is a chat assistant that works like ChatGPT but runs on servers your company controls. Staff sign into a chat app that sends their questions to a language model on those servers, and a search over your own documents lets the model answer from them. Each of these parts exists today as open-source or licensed software.

The work is in connecting those parts to your company sign-in, document permissions and logs, and then in running the system. Your IT team can set up a trial. Before every employee relies on it, a regulated firm should hire a team that can show each requirement below working on the firm's own systems.

The short answer: your IT team can run a trial on documents everyone in the trial group may read. For daily use across the firm, hire a team only after it shows the assistant working with your own sign-in and documents. The assistant should hide documents a person cannot open and keep chats and logs on your servers for a period your compliance team sets.

Why do companies want a private ChatGPT?

Staff are not waiting for their employer to provide AI tools. In the 2024 Work Trend Index from Microsoft and LinkedIn, a survey of 31,000 knowledge workers in 31 countries, 78% of AI users said they were bringing their own AI tools to work.

In May 2023 Samsung temporarily restricted generative AI tools on company devices and internal networks, after internal data had been accidentally leaked to ChatGPT in April. TechCrunch reported Samsung's concern that data on outside servers is hard to “retrieve and delete”. The company said it was “reviewing measures to create a secure environment for safely using generative AI”.

In a survey of UK financial firms that the Bank of England and the FCA published in November 2024, data protection and privacy was the largest perceived regulatory constraint on the use of AI.

What is a private ChatGPT made of?

“Private ChatGPT” is not the name of an OpenAI product. It describes a system put together from separate parts, and a model from OpenAI, the company behind ChatGPT, can be one of them. The open-source projects named below are examples you may meet in proposals, and this article does not rank them. These are the parts:

  • The chat app is the web page where staff sign in, ask questions, upload files and find past chats. LibreChat is one open-source example, under the MIT licence. Its project page lists company sign-in through OAuth2 and LDAP, two common ways of connecting an app to the accounts staff already have. It also works with any model server that accepts requests in the format of OpenAI's own service, known as an OpenAI-compatible API
  • The model server runs the language model on your hardware and answers the chat app's requests. vLLM is an open-source example. Its documentation says it “implements OpenAI's Completions API, Chat API, and more”, so a chat app written for OpenAI's service can send its questions to your server instead
  • The model is a set of files, called weights, that the server loads. Anyone can download and run an open-weight model, on the terms of that model's licence. OpenAI's gpt-oss-120b, for example, is published under the Apache 2.0 licence, and OpenAI says it fits on a single 80 GB graphics processor (GPU) like an NVIDIA H100
  • Document search lets the assistant answer from your own policies and contracts. The documents are split into passages and kept in a search index, and the passages that match a question are passed to the model along with it. Engineers call this retrieval-augmented generation, or RAG
  • Sign-in should go through the company accounts staff already use, so that closing an employee's account also ends their access to the assistant
  • Chat history and logs record who asked what and when, and what the assistant answered. Both are copies of your data
  • Admin settings decide who may use which model, which document collections each group can search, whether files can be uploaded and how long chats are kept

Licences matter once the whole company uses the system. Open WebUI, another chat app whose code is public, comes with company sign-in, user groups and built-in document search. Since version 0.6.6 (April 2025), its licence requires the Open WebUI name and logo to stay visible once a deployment has more than 50 users in a 30-day period. An enterprise licence lifts that rule, and so does the project's written permission for a substantial contributor. The project says the clause means its licence would not meet the Open Source Initiative's criteria for open source.

Does “our own infrastructure” rule out ChatGPT Enterprise or Microsoft Copilot?

Yes, if “our own” means servers your company controls, because OpenAI runs the servers behind ChatGPT Enterprise and Microsoft runs those behind Copilot. If your compliance team accepts data held by a provider under contract, both can stay on the list. The model behind a company assistant can run in three places, and only the third keeps every prompt and file on servers you control:

  • With a provider's business plan, such as ChatGPT Enterprise or Microsoft Copilot, the provider runs the servers and your data is covered by its contract. OpenAI's ChatGPT documentation, read in October 2026, describes data residency and inference residency for Enterprise workspaces. They tie where data is stored and where the model runs to a chosen region, and the documentation says both “apply only to eligible content and supported workloads”. It adds that “some processing or synced indexes can follow separate location rules”
  • Cloud providers also sell model services, such as Azure OpenAI and Amazon Bedrock, which run the model in the provider's data centres. Your chat app can sit in your own cloud account, but each question, with the document passages attached to it, goes to that service
  • An open-weight model can run on servers you operate, in your own data centre or on virtual machines in your own cloud account. Prompts, uploaded files and logs then stay on systems you control, and the servers and the model's security become your job, or the job of the team you hire

The articles on adding a language model to your own systems and on running one in a closed perimeter, both linked above, go into more detail. They compare a provider's service, a cloud service and servers you run. The pages linked at the bottom of this page cover cost and the two cloud services.

What separates a trial from an assistant your staff use every day?

A trial shows whether people find the answers useful. Daily use in a regulated firm adds six requirements, and a builder should be able to show each one working before launch:

  • Each person sees only the documents they can already open. Microsoft sets this rule for its own assistant, writing that Copilot “only surfaces organizational data to which individual users have at least view permissions”. The OWASP Top 10 for LLM Applications, a list of the main security risks in language model software, recommends “permission-aware vector and embedding stores”. That means the search index must record who may read each passage
  • Documents are treated as untrusted input, because a file can carry hidden instructions. In one OWASP example, a CV has white text on a white background that tells the system to recommend the candidate, and the model obeys it when someone later asks about the candidate. OWASP adds that answering from your own documents is meant to make answers more accurate, and research shows it does not fully protect against such attacks. So the assistant should only write answers, and any action in another system, such as sending an email, should wait for a person's approval
  • Chats, uploads and logs stay on your servers for a period your compliance team sets. Microsoft lets administrators use its Purview compliance tools to “set retention policies for the data related to chat interactions with Copilot”. A private assistant needs the same control over the chat database, uploaded files, the search index and the model server's logs, which can hold the full text of questions. Your compliance team should be able to read who asked what without asking an engineer
  • A set of your own questions, each with its correct answer, is rerun after every change. The people who know the answers write the questions from their real work. A change here means a new model, new instructions to the model or a large batch of new documents. The pilot article linked above explains how to build such a set
  • Capacity is sized for the busiest hour, because a trial for a small group says little about a morning when much of the office asks at once. Ollama is a tool for running a model on one machine. In October 2026 its documentation gave a default of one request at a time for each model, with further requests waiting in a queue. The comparison of vLLM and Ollama, linked at the bottom of this page, explains how each handles many people asking at once
  • A named person owns the assistant and can switch it off, or back to the previous model and settings, without waiting for the supplier. Staff know where to report a wrong answer

Who can set up a private ChatGPT on our own servers?

Your own IT team, a vendor or an engineering team can set it up. Which one fits depends on how many people will use the assistant and what it has to connect to.

For a trial with a small group, an IT team that already runs Linux servers can install an open-source chat app and model server and load an open-weight model. A smaller model can be enough to start, and OpenAI says its gpt-oss-20b runs within 16 GB of memory. Read every licence first, the chat app's included, and keep the trial to documents everyone in the group may read, so that permissions do not matter yet.

A packaged product, sold by a vendor as software you install on your own servers, can suit plain chat over a shared document library. Ask which of your systems it connects to and who installs its updates, and put it through the five demonstrations in the next section, as you would a custom build.

An engineering team is needed when the assistant has to work with your company sign-in and the permissions of several systems. The same applies when it must connect to other systems, such as a document management system, or keep logs your compliance team can read. The team can be an outside company or engineers who join your IT team. Agree before the work starts who will run the system after launch.

How do we check a company that offers to set it up?

Before you sign, ask each supplier for five demonstrations. Where you can, use a test installation with your own company sign-in and documents you are allowed to share with the supplier:

  • Sign in as someone who may not open a particular document, and ask about it. The assistant should answer as if the document did not exist. Then remove a colleague's access to a file in the original system and time how long the assistant takes to stop using it for that colleague
  • Have the supplier show where chats, uploads and logs are stored and delete one person's history while you watch. Then find out what remains in backups and logs, and when it is removed
  • Look at the set of test questions the supplier reruns after a model update and the record of its last run, and find out who wrote the correct answers
  • Ask how many people at once the system was sized for, and see the load test behind that number. The test should use questions and documents like yours, at the number of users you expect in your busiest hour
  • Check that you could run the system without the supplier. The code, the configuration, the model files, the test questions and the running instructions should all be handed over, and nothing should depend on the supplier's servers or licence keys

Where does amBrain fit?

On its own work with language models, amBrain says: “We have taken an LLM integration to production inside a client's FinTech perimeter: extracting and normalising unstructured broker and venue notices — corporate actions, instrument and margin changes — into structured records the trading system consumes.” The client is not named, and no numbers about the project are published.

This article is not a case study. It does not say where the model ran in that project, and it does not claim that amBrain has built a private ChatGPT, a staff assistant or a document search for any client. It gives no prices or timelines.

amBrain has been building software since 2019. amBrain works in three formats: full delivery, a dedicated team, or engineers embedded in your team. The client keeps full ownership of the product and the code, except amBrain's reusable components.

Write one page that says which of the three cases above fits you. For each of the six requirements, add what it means in your firm, such as which systems hold the document permissions and how long chats must be kept. Send that page to amBrain, or to any other team on your list, and ask every team for the same five demonstrations.

Common questions

  • Will it be as good as ChatGPT? Nobody can say before it is tested on your own work. Run your set of questions through each model you are considering, and have the people who know the answers mark the results. To compare with a public service, use questions that contain no client data
  • Can we start with a provider's model and move to our own servers later? Yes, if your chat app reaches the model through the OpenAI-compatible API. Until the move, send the provider only data your compliance team allows to leave the company. vLLM implements the same chat API, so the app then needs only your server's address and model name, as long as it uses no feature that only the provider's service offers. Rerun your questions before staff switch over, because the answers change with the model, and ask the provider how the data already sent to it will be deleted
  • Can staff upload their own files? Yes. An upload is a copy of the file, kept by the chat app and sometimes added to the search index. It needs the same retention period as the chats and has to be deleted with them. Decide whether an upload stays in the uploader's own chats or can be shared with a group

Have a design like this on the table?

Bring your current architecture and the failure mode that worries you, and we will go through it together in half an hour.