amBrain
FinTechSep 28, 202610 min read

Hire an ML Team or Outsource AI Development: How to Decide and Who to Ask

ML TeamOutsourcing AIHire or PartnerWho Builds It
Error loading image

Hire an ML team if AI is the product you sell. Outsource if you need one or two AI systems and nobody in the company can lead ML hires.

Hire your own ML team if AI is what your customers pay for and the models will need work every week for years. Outsource if you need one or two AI systems inside operations you already run and nobody in the company can lead ML hires or judge the candidates. When you choose who builds it, look for a firm that can show you an AI system it took to production and that is still in daily use, and make the contract hand you the code, the model files and the evaluation set.

The short answer: hire when AI is your product, and outsource when it is a tool inside the business and nobody in-house can lead an ML team. If you outsource but will run the system for years, do both in order: an outside team builds the first version while you hire the two or three people who will own it.

Do we need ML researchers, or engineers who build with existing models?

Most companies asking this question do not need anyone to invent or train a new model. They need a system that takes an existing model, gives it their documents, tickets or transactions, checks what comes out and puts the result where staff already work. That is engineering around a model, and it is a different hire from research.

Training your own model pays off mainly when the model is what customers buy, and it takes a lot of labelled data. Building on an existing model, whether a provider's service or a model you download and run on your own servers, takes people who connect systems, check the quality of the answers and keep software running. Decide which of the two you need before you write a job description or call a partner.

What does an in-house ML team actually consist of?

One ML engineer is not a team. A model in production sits inside a lot of ordinary software, and the people who build and run that software are most of the team. Google's architecture guide on MLOps puts it this way: “Only a small fraction of a real-world ML system is composed of the ML code. The required surrounding elements are vast and complex.”

A team that can build and run one production system without outside help covers four roles:

  • A lead who can plan the work and judge ML candidates. Without this person the company cannot tell a good hire from a confident one
  • ML engineers who choose the model, prepare the data, write the evaluation and improve the results
  • A data engineer who gets clean, permitted data out of the systems you already run and keeps it flowing
  • An engineer who runs the infrastructure: servers or cloud accounts, deployment, monitoring, and the alert that fires when the model's answers get worse

In a small team one person can hold two of these roles, but nobody holds all four. One more role sits on the business side, and no hire replaces it: the person who decides what a correct answer is.

How much does an in-house ML team cost?

Salaries are usually the largest line, and public data gives a reference point. The US Bureau of Labor Statistics does not publish a separate wage figure for machine learning engineers. For the closest occupations it covers, the median annual wage in May 2025 was $120,230 for data scientists, $135,980 for software developers and $140,300 for computer and information research scientists.

These are medians for whole US occupations. People who have already put ML systems into production are a narrower group, and your location and the seniority you need move the figure either way.

Salary is not the whole cost of an employee. Across all US private-sector jobs in June 2026, wages and salaries were 70.0 percent of what employers spent on compensation, and benefits were the other 30.0 percent, according to the BLS.

Other costs come on top of salaries:

  • Recruiting for each role, and the delay an empty lead position puts on every hire after it
  • Computing capacity for experiments and for production, which grows with use
  • Tools for labelling data, tracking experiments and monitoring answers
  • Hours from your own staff, who write correct answers and check the model's output in the first months
  • A team sized for building, which is often larger than running the finished system requires

This article gives no price range for outsourcing: no firm can price the work honestly before it has seen the task, the data rules and the volume.

When does hiring our own ML team make sense?

Hiring pays off when the work never ends and the knowledge is worth keeping inside:

  • Customers pay for what the model does. Competitors can buy the same base models you can, so the work your team does on top of them is what you sell
  • The models need attention every week. Google's MLOps guide names two reasons a model's performance drops: suboptimal code and “constantly evolving data profiles”. Where the data shifts all the time, retraining and retesting is a permanent job
  • You hold data nobody else has. The people who learn its quirks become hard to replace, and they should work for you
  • You can hire an experienced lead first and keep them. The rest of the team is built around that person

If the first two are true, hire. Outside engineers can still shorten the start while your team forms.

When is outsourcing AI development the better choice?

Outsourcing fits when AI is a tool inside the business rather than the business itself:

  • You need one or two systems, such as reading incoming documents or sorting support tickets, rather than a steady stream of new models
  • Most of the work is connecting a model to systems you already run, such as the helpdesk, the document store or the core database. A firm that does this regularly has met your integration problems before
  • Nobody in the company can lead ML hires or judge the candidates. Hiring a team you cannot manage is an expensive way to find out you needed a partner
  • You want to learn whether the task works at all before committing to salaries. A first system built by an outside team answers that with your own data

One public data point leans the same way. In “The GenAI Divide”, a July 2025 report from MIT NANDA based on interviews at 52 organizations, generative AI tools bought from or developed with outside vendors reached deployment about 67% of the time, against about 33% for tools built fully in-house. The authors describe the figures as self-reported and warn that part of the gap may come from the organizations themselves.

Outsourcing has its own cost: the knowledge of how the system works sits outside your company until someone moves it in.

Can we outsource the first AI system and hire a team to own it later?

Yes, and for a company that expects to run the system for years it is often the safest order. An outside team builds the first version, and you hire two or three people during the build rather than after it. They review code, join design decisions and, near the end, run the system themselves while the builders watch.

This only works if the handover is written into the contract as a list of things you receive and can use without the builder:

  • The source code in your own repositories, with its full history
  • The model itself wherever one was trained or fine-tuned: the model files, the settings that produced them, and the data they came from or a description of it
  • The prompts and configuration, versioned together with the code that uses them
  • The evaluation set: real examples with agreed correct answers, and the script that scores the system against them
  • The data pipeline, and the written rules for which data may go where
  • Runbooks for deployment, for rollback and for the day the model starts giving bad answers

Check the licence of any base model in that package, because its conditions travel with what is built on it. Meta's Llama 3.3 licence, for example, says that if you use Llama to “create, train, fine tune, or otherwise improve an AI model, which is distributed or made available, you shall also include ‘Llama’ at the beginning of any such AI model name.”

A handover test that works with any builder: your new hires change a prompt or a setting, run the evaluation set, deploy the change and roll it back, while nobody from the builder touches the keyboard. Anything they have to ask the builder about is something you do not own yet.

What should we ask an AI development company before signing?

Ask every candidate the same questions, in writing:

  • Show us one AI system you took to production that is still in daily use. What does it do, who runs it today, and what happens when the model gets an answer wrong?
  • Where will our data go during the work, including logs, test copies and anything used to tune a model? Will any of it leave our servers or accounts, or improve tools you use for other clients?
  • How will you measure output quality? We expect an evaluation set built from our own examples, agreed before the build starts, with a pass mark
  • What exactly do we receive at the end, and can our own engineers run and change the system without you?
  • Which of your own components will stay in the system? Name each one and our terms for using it

Also note what each firm asks you. One that has built this before asks about your task and your data before it names a model.

What are the red flags when outsourcing AI development?

  • The proposal names a model before anyone has looked at your data
  • Accuracy figures come without a test set, or with a test set the firm chose alone
  • The contract allows your data to improve the firm's shared tools, or says nothing about it
  • Nobody asks who will run the system after launch
  • The ownership clause keeps “our platform” or “our components” without a list of what those are
  • A custom-trained model is the first answer to every task. A firm should explain why an existing model will not do before it charges to train one

Which companies build AI for other businesses, and who would you recommend?

This article does not name a best company. A recommendation that ignores your task, your data rules and who will run the system afterwards is a guess. Five kinds of firm do this work, and each fits a different situation:

  • The professional services arms of cloud providers, and their partner networks. A sensible choice when the system will live in that cloud anyway; expect a design built on the provider's own managed services
  • Large consultancies, when AI is one part of a wider change across several departments. Ask who will write the code and whether those people are their own staff
  • Specialised engineering firms, for one or two systems built and connected to what you already run. Ask for a production system close to your problem, not a list of technologies
  • Freelance ML engineers, for a bounded task that someone inside the company can judge. When that person leaves, the knowledge leaves too
  • Software vendors, if the task is common, such as a support chatbot or reading standard invoices. A finished product can beat both hiring and building, so check this before commissioning anything custom

To build a shortlist, write one page: the task in one sentence, the data the system may see and where it may go, real examples with correct answers, the daily volume, and who will own the system after launch. Send the same page to three firms of the kind that fits, and compare their questions as closely as their proposals. If you can, pay for a small first phase with written acceptance criteria before signing for the whole build.

Where does amBrain fit?

Of the kinds of firm described above, amBrain is an engineering firm.

On its own work with language models, amBrain says: “We have taken an LLM integration to production inside a client's FinTech perimeter: extracting and normalising unstructured broker and venue notices — corporate actions, instrument and margin changes — into structured records the trading system consumes.” The client is not named, and no numbers about the project are published.

amBrain describes how it works with clients in one line: “Three formats: full delivery, a dedicated team, or engineers embedded in your team.” On ownership, its line is: “The client keeps full ownership of the product and the code, except our reusable components.” Ask amBrain for the list of those components by name, as you would ask any other firm.

If you are deciding now, start with the one-page brief from the previous section. Send it to amBrain or to anyone else, and compare what comes back.

Common questions

  • Can we start by hiring one ML engineer? You can, but one person cannot cover leading, building, data and operations, and nobody inside can judge their work. If you hire one person first, hire someone who can lead and later bring in the rest
  • Does outsourcing mean our data leaves the company? Not necessarily. An outside team can work inside your servers or cloud accounts under your access rules, and the contract can say so. Ask specifically where logs and test copies go, because those are the copies people forget
  • Can we bring an outsourced system in-house later? Yes, if the handover list above is in the contract from the start. Adding it after the build is harder, because nobody wrote the evaluation set or the runbooks while the work was fresh

Have a design like this on the table?

Bring your current architecture and the failure mode that worries you, and we will go through it together in half an hour.