For a trading or ad tech system where a late answer counts as a wrong one, the right outside firm works in the same range as your deadline and can show how its speed figures were measured. The one you favour should measure your live system before it builds anything.
No list of low-latency software development companies fits every buyer, because the term covers deadlines that can be a million times apart. In trading hardware it can mean tens or hundreds of nanoseconds. In an app or a web page, an answer within about a tenth of a second already feels instant to the person using it. So the first question is where your own deadline falls.
The short answer: write down your deadline and the two points it is measured between, and ask each firm which of the systems it built already run in that range in production. Before anyone builds anything, pay the firm you favour to measure your live system, and keep its report whatever you decide.
Read next
For your system, low latency is a deadline measured between two points you can name. Deadlines fall into four broad ranges:
One product can have parts in different ranges. A trader reads prices on a screen, while the order that trader sends may have to meet a much tighter deadline on its way to the venue. Write down each deadline with the two points it is measured between. The range each one falls in tells you which kind of firm to call.
This article does not rank firms. The right kind of firm depends on your range and on the work you need done:
Many firms fit more than one kind, so ask which kind of work the engineers assigned to you have done.
The article on dedicated teams, linked above, lists what should come with every performance figure and gives a checklist you can paste into a request for proposal. The questions below find out how a firm's own figures were produced.
Ask where the clock starts and where it stops. OPRA, which publishes consolidated trades and quotes from the US options exchanges, starts its clock when an incoming message “arrives at the application entrance to the OPRA environment” and stops it when the outgoing message “arrives at the application exit from the OPRA environment”. The EU definition quoted above is as precise about both points. A figure without named end points cannot be compared with yours.
Look at the slowest answers as well as the typical one. In OPRA's published metrics, the median latency, the middle value, was 19.5 microseconds in January 2024 and 20.5 in February. Over the same two months the 99th percentile, the time that only the slowest 1 percent of messages exceeded, fell from 543.5 microseconds to 57.5. A report with only the median would have shown almost no change.
Averages hide the slow answers too. Google's Site Reliability Engineering book, published by O'Reilly in 2016, describes a web service with an average latency of 100 milliseconds at 1,000 requests per second, in which “1% of requests might easily take 5 seconds”. Ask for the 99th percentile of each path and, if your system handles enough requests to measure it, the 99.9th, which only 1 answer in 1,000 exceeds.
Check the load behind each figure. In STAC's 2020 overview, STAC-T0 sends its test traffic at three rates. The lowest is “designed to see how systems behave when they are mostly idle”, and the highest, typically near the most the tested system can handle, is there “to see how systems behave when they are very busy”. Ask for figures from your own busiest minute, or from a test that reproduces it.
Ask how the load test was run. Many load-testing tools send a request, wait for the answer and only then send the next one. When the system stalls, such a tool stops sending, so the requests that would have arrived during the stall are never timed.
Gil Tene, the author of the load-testing tool wrk2, calls this effect coordinated omission. In the tool's documentation, last changed in September 2019, he writes that “high latency responses result in the load generator coordinating with the server to avoid measurement during high latency periods”. His tool sends requests at a fixed rate and times each answer “from the time the transmission should have occurred”. Ask whether the firm's tool works this way, or how its results were corrected.
Ask where each timestamp was taken. For latencies in microseconds, STAC's overview calls a benchmark with software timestamps “the best option”. It also warns that the small, uneven delays in software timestamps “can represent considerable error when measuring latencies in tens or hundreds of nanoseconds”, and STAC-T0 takes its timestamps in hardware instead. A firm that quotes nanoseconds should be able to show where its hardware timestamps were taken.
If you do not know yet what is wrong, only that the system seems slower or costs more than it should, start here. Pay the firm you favour for a fixed-scope measurement of your live system, with a report that stays yours whatever you decide next. Agree in writing what the firm may install or change to take its measurements, and how its tools are removed afterwards.
The report should show:
For trading systems, the article on slow order execution, linked above, lists the four timestamps to record on every order. In ad tech, the articles on traffic spikes and on DSP timeouts cover what to measure at a peak and for each exchange connection.
Judge the firm by its report. If your own engineers could act on it without the firm, you can choose who does the work next from real numbers, whether that is this firm or another.
If a firm proposes a new language, the article linked below shows how engineers check whether the language is the cause.
amBrain diagnoses slow systems in trading and ad tech: the running platform is measured end to end and the report names where the time goes. Its trading work includes trading terminal development, order management systems, and FIX protocol exchange integration.
In AdTech, amBrain works on DSP development, real-time bidding platforms, and ad exchange engineering.
amBrain works in three formats: full delivery, a dedicated team, or engineers embedded in your team. The client keeps full ownership of the product and the code, except amBrain's reusable components.
amBrain has been building software since 2019.
This article is not a case study and describes no client work. It quotes no latency figure for any system amBrain has built, and no prices or timelines.
Ask amBrain, or any other firm on your list, what it would measure first on your system, and put every firm through the same checks.
Bring your current architecture and the failure mode that worries you, and we will go through it together in half an hour.