amBrain
FinTechOct 1, 20269 min read

Low-Latency Software Development Companies: Which Kind You Need and How to Test One on Your Own System

Low LatencyChecking a PartnerLatency MeasurementTail Latency
Error loading image

For a trading or ad tech system where a late answer counts as a wrong one, the right outside firm works in the same range as your deadline and can show how its speed figures were measured. The one you favour should measure your live system before it builds anything.

No list of low-latency software development companies fits every buyer, because the term covers deadlines that can be a million times apart. In trading hardware it can mean tens or hundreds of nanoseconds. In an app or a web page, an answer within about a tenth of a second already feels instant to the person using it. So the first question is where your own deadline falls.

The short answer: write down your deadline and the two points it is measured between, and ask each firm which of the systems it built already run in that range in production. Before anyone builds anything, pay the firm you favour to measure your live system, and keep its report whatever you decide.

What does “low latency” mean for our system?

For your system, low latency is a deadline measured between two points you can name. Deadlines fall into four broad ranges:

  • Trading hardware, in tens or hundreds of nanoseconds. STAC-T0, a benchmark developed in consultation with trading firms in the STAC Benchmark Council, times how fast a system's network hardware and software turn simulated market data into a simulated order, with no trading logic in between. STAC's overview of 5 November 2020 says the benchmark treats the system as a black box, “interacting with it solely via network packets, which it timestamps in hardware”, and that software timestamps can carry considerable error at tens or hundreds of nanoseconds. The stack under test, as STAC calls the system being measured, can be an FPGA card, which carries a chip whose circuits are programmed for one task
  • Trading venues, from microseconds up to about a millisecond. The EU rule on how accurate a trading venue's clocks must be ties that accuracy to how long the venue's trading system takes to process an order and send back a confirmation. Since 2 March 2026 that rule is Commission Delegated Regulation (EU) 2025/1155, which replaced the earlier rule known as RTS 25 and calls this time gateway-to-gateway latency, “the time measured from the moment a message is received by an outer gateway of the trading venue's system, sent through the order submission protocol, processed by the matching engine, and then sent back until an acknowledgement is sent from the gateway”. Where that takes 1 millisecond or less, the venue's clocks may be no more than 100 microseconds off Coordinated Universal Time (UTC), and timestamps must be 0.1 microseconds or finer
  • Ad tech, from tens of milliseconds up to a second. OpenRTB 2.6, the IAB Tech Lab standard for real-time bidding, lets an exchange set a deadline in each bid request, and the time spent crossing the Internet counts against it. Google's Authorized Buyers developer documentation, last updated on 17 September 2026, says the response deadline “ranges from 80 to 1000 ms, depending on format and auction type”
  • Apps and web pages that people use, around a tenth of a second. Jakob Nielsen wrote in 1993 that “0.1 second is about the limit for having the user feel that the system is reacting instantaneously”. For web pages, Google's web.dev guide, last updated in September 2025, rates responsiveness as good when Interaction to Next Paint (INP) is 200 milliseconds or less. INP measures how long a page takes to respond to clicks, taps and key presses

One product can have parts in different ranges. A trader reads prices on a screen, while the order that trader sends may have to meet a much tighter deadline on its way to the venue. Write down each deadline with the two points it is measured between. The range each one falls in tells you which kind of firm to call.

Which low-latency software development companies should we talk to?

This article does not rank firms. The right kind of firm depends on your range and on the work you need done:

  • Hardware and network specialists, for deadlines counted in nanoseconds or a few microseconds. They work on FPGA cards, network cards, switches and the links to an exchange. Ask which of their figures come from timestamps taken by hardware on the network cable, and on whose equipment the tests ran
  • Trading technology engineering firms, for deadlines in microseconds or milliseconds when the time is lost inside your own software or on the way to your broker or venue. They write the software that sends orders, handles market data, checks each order's risk before it leaves and, at a venue, matches buy and sell orders. Ask which systems they built are in production today, and which broker or venue connections they wrote
  • Ad tech engineering firms, when an exchange sets your deadline request by request and your server bill grows with your traffic. They build bidders and ad exchanges. Ask which bidder or exchange they built is still live, and whether its timeout figures come from the exchange's count or their own
  • Independent performance engineers, when you do not know the cause yet, or when your own engineers will make the changes. Usually working alone or in a small group, they measure a running system and report where the time goes. Ask to see an earlier report with the client's name and data taken out, and expect a diagnosis rather than a rebuild
  • Component vendors, when the part you need works the same way for everyone and your advantage lies elsewhere. They sell a ready-made part, such as a matching engine or an exchange connection, that you run instead of commissioning your own. Ask where their latency figures were measured from and to, and what you may change in the code you license
  • General outsourcing companies with a named low-latency team, when you need many engineers and the fast path is a small part of the work. Ask for the names of the people on that team and what each of them built in your range, and get it in writing that they will work on your project
  • Your own hires, when speed is how you compete and will stay that way, and you can recruit and keep engineers who have already shipped this kind of system. Acuiti, a research firm, asked 50 systematic hedge funds, funds that trade by computer models, how they build their trading technology. In January 2023 it reported that latency “is the key factor in determining attitudes to outsourcing front office technology, with firms for which latency is critical more likely to develop in-house”

Many firms fit more than one kind, so ask which kind of work the engineers assigned to you have done.

How do we check that a firm has built fast systems before?

The article on dedicated teams, linked above, lists what should come with every performance figure and gives a checklist you can paste into a request for proposal. The questions below find out how a firm's own figures were produced.

Ask where the clock starts and where it stops. OPRA, which publishes consolidated trades and quotes from the US options exchanges, starts its clock when an incoming message “arrives at the application entrance to the OPRA environment” and stops it when the outgoing message “arrives at the application exit from the OPRA environment”. The EU definition quoted above is as precise about both points. A figure without named end points cannot be compared with yours.

Look at the slowest answers as well as the typical one. In OPRA's published metrics, the median latency, the middle value, was 19.5 microseconds in January 2024 and 20.5 in February. Over the same two months the 99th percentile, the time that only the slowest 1 percent of messages exceeded, fell from 543.5 microseconds to 57.5. A report with only the median would have shown almost no change.

Averages hide the slow answers too. Google's Site Reliability Engineering book, published by O'Reilly in 2016, describes a web service with an average latency of 100 milliseconds at 1,000 requests per second, in which “1% of requests might easily take 5 seconds”. Ask for the 99th percentile of each path and, if your system handles enough requests to measure it, the 99.9th, which only 1 answer in 1,000 exceeds.

Check the load behind each figure. In STAC's 2020 overview, STAC-T0 sends its test traffic at three rates. The lowest is “designed to see how systems behave when they are mostly idle”, and the highest, typically near the most the tested system can handle, is there “to see how systems behave when they are very busy”. Ask for figures from your own busiest minute, or from a test that reproduces it.

Ask how the load test was run. Many load-testing tools send a request, wait for the answer and only then send the next one. When the system stalls, such a tool stops sending, so the requests that would have arrived during the stall are never timed.

Gil Tene, the author of the load-testing tool wrk2, calls this effect coordinated omission. In the tool's documentation, last changed in September 2019, he writes that “high latency responses result in the load generator coordinating with the server to avoid measurement during high latency periods”. His tool sends requests at a fixed rate and times each answer “from the time the transmission should have occurred”. Ask whether the firm's tool works this way, or how its results were corrected.

Ask where each timestamp was taken. For latencies in microseconds, STAC's overview calls a benchmark with software timestamps “the best option”. It also warns that the small, uneven delays in software timestamps “can represent considerable error when measuring latencies in tens or hundreds of nanoseconds”, and STAC-T0 takes its timestamps in hardware instead. A firm that quotes nanoseconds should be able to show where its hardware timestamps were taken.

What test should we run before we hire anyone?

If you do not know yet what is wrong, only that the system seems slower or costs more than it should, start here. Pay the firm you favour for a fixed-scope measurement of your live system, with a report that stays yours whatever you decide next. Agree in writing what the firm may install or change to take its measurements, and how its tools are removed afterwards.

The report should show:

  • Where the time goes in your busiest minute, step by step along the path you wrote down
  • For each path, the 99th and 99.9th percentile response times
  • The fixes, ranked by what each one costs and what it removes, whether that is time on the path or servers that were added to hide a slow one
  • What the firm would leave alone, and why

For trading systems, the article on slow order execution, linked above, lists the four timestamps to record on every order. In ad tech, the articles on traffic spikes and on DSP timeouts cover what to measure at a peak and for each exchange connection.

Judge the firm by its report. If your own engineers could act on it without the firm, you can choose who does the work next from real numbers, whether that is this firm or another.

What are the warning signs when choosing a low-latency firm?

  • Speed described in words, with no number and no end points
  • Averages only, with nothing about the slowest answers
  • A benchmark from the firm's own lab, run on its hardware with traffic it generated, offered as evidence about your system
  • A rewrite, or a move to a new programming language, proposed before anyone has measured your system
  • A latency figure promised in the first call
  • The same pitch for trading work measured in microseconds and for an ad bidder with a 100-millisecond deadline

If a firm proposes a new language, the article linked below shows how engineers check whether the language is the cause.

Where does amBrain fit?

amBrain diagnoses slow systems in trading and ad tech: the running platform is measured end to end and the report names where the time goes. Its trading work includes trading terminal development, order management systems, and FIX protocol exchange integration.

In AdTech, amBrain works on DSP development, real-time bidding platforms, and ad exchange engineering.

amBrain works in three formats: full delivery, a dedicated team, or engineers embedded in your team. The client keeps full ownership of the product and the code, except amBrain's reusable components.

amBrain has been building software since 2019.

This article is not a case study and describes no client work. It quotes no latency figure for any system amBrain has built, and no prices or timelines.

Ask amBrain, or any other firm on your list, what it would measure first on your system, and put every firm through the same checks.

Common questions

  • Should we build our own team instead? In Acuiti's 2023 study of 50 systematic hedge funds, those for which latency is critical were more likely to develop their trading technology in-house. Measure first either way, because the report tells you which kind of engineer to hire, or what to ask a firm to do
  • Can one firm cover both trading and ad tech? It can, if it has production systems in your range in both fields. Ask for one live system in each field, and check its figures with the questions above

Have a design like this on the table?

Bring your current architecture and the failure mode that worries you, and we will go through it together in half an hour.