Timeouts on two exchange connections and an infrastructure bill that grows faster than revenue can both be measured connection by connection before any code is rewritten. Those numbers also let you check any team that offers to fix the bid path.
If your DSP times out with two exchange connections and not with the others, look first at what those two have that the rest do not. Possible causes are a longer network route, a shorter deadline, heavier requests or network connections that are reopened too often. The same causes can push the bill up faster than revenue, because your servers still do the work for answers that arrive too late to count.
The short answer: nobody can name the right team for you without numbers from your two connections. Whether a team has actually fixed bid-path latency elsewhere is something only its clients can confirm, on a call you arrange. For your own problem, ask for a written plan built on your data. Send one page of per-connection numbers to two or three teams, and go further only with a team that says, for each connection, what it will measure and how you will both know the problem is fixed.
Why does our DSP time out on two exchange connections and not on the others?
In OpenRTB, the IAB Tech Lab protocol for real-time bidding (RTB), an exchange can put the deadline into each request through an optional field called tmax, and time spent on the Internet counts against it. When only two connections fail, start with what makes those two different:
- Distance uses up part of each deadline. Google's Authorized Buyers documentation lists four trading locations for its bid requests, in Northern Virginia, the San Francisco Bay Area, Amsterdam and Singapore, and advises bidders to put their servers close to them. For bidders that receive many requests, Google also recommends peering, a direct link between their network and Google's, to reduce latency and how much it varies
- Deadlines differ by exchange and by request. At Google, the deadline depends on the ad format and the auction type. An exchange that passes a request on can also keep part of the time for itself. Equativ's bid request specification, for one, says the tmax value sent to its bidders is always lower, to leave enough time to process the bid responses
- Some exchanges send heavier requests. OpenRTB lets each exchange add its own extra fields and offer several impressions in one request, and whether requests arrive as plain JSON, in a binary format or compressed is agreed with each exchange. A larger request takes longer to receive and decode, and each extra impression is one more round of campaigns to check
- New network connections start with less time. Google's best-practice guide for RTB applications says the first request on a new connection has a shorter effective deadline and is more likely to time out, and it recommends keeping idle connections open for 2.5 minutes. If your servers, or a load balancer or proxy in front of them, close idle connections sooner, some requests have to wait for a new connection to open inside their deadline
- Some steps run only for certain requests. When a lookup of user data, or a model used for one ad format, is slow, the delay lands only on the connections whose requests use it
- One exchange's traffic can land on busier servers, where requests wait in a queue before any work starts. Google notes that network connections made through a proxy can become unbalanced over time and leave the load on your servers uneven
A slowdown of the whole bidder process delays every connection at once. In a Go or Java bidder, garbage collection (GC), the runtime's work of recycling memory the program no longer needs, can cause one. The connections with the least time to spare, after network time and the work each request needs, miss their deadlines first. So compare each connection's deadline with its network time and request size before looking for a cause that only those two connections have. The article on Go GC pauses, linked above, shows how engineers separate these causes.
Why is our infrastructure bill growing faster than revenue?
The bill grows with every request your servers receive and answer, and revenue comes only from the auctions you win. The gap widens in several ways:
- A request for a format or country you have no campaigns for still has to be received and parsed, and it cannot win. Google's pretargeting lets a bidder receive only the requests that match its targeting criteria. Ask each exchange what filtering it offers
- Late answers can also shrink the traffic you are sent. Google's help page on RTB graphs says that when more than 15 percent of responses are invalid or timed out, Google sends fewer requests until the error rate drops below 15 percent or the requests fall to a minimum. If traffic is throttled often and for long periods, Google may adjust the bidder's quota, the most requests per second it will send, to a level the bidder can handle more consistently. Servers sized for the old quota keep costing money unless someone resizes them
- Servers added in the same data centre raise the bill but do not shorten a long route to the exchange or stop idle network connections from being closed too early
- The same impression can reach you through more than one exchange. OpenRTB 2.6 describes a transaction ID that must be common across all participants in a bid request, potentially across several exchanges, and a supply chain object that lists the companies involved in the direct flow of payment. Where exchanges fill them in, these fields can show that two connections offer you the same impression, and each copy costs you server time
- Some spare capacity is necessary. To absorb temporary shifts of traffic between regions, Google recommends a 15 percent cushion between the seven-day peak and the requests per second set for each trading location. The spare capacity to question first is capacity added after an incident without a measurement to justify it
To see where the gap opens, put cost and revenue per million requests side by side for each exchange connection. Look first at a connection that brings many requests and few wins, and check whether it is one of the two that time out.
What should we measure before hiring anyone?
Put these on one page, one row per exchange connection, for a normal week and its busiest hour:
- Requests per second by exchange and region, and the average request size
- The spread of deadlines in those requests, read from tmax where the exchange sends it and from its documentation where it does not
- On each exchange connection, the response time that only your slowest 1 percent of answers exceed (the 99th percentile), with network time in both directions separated from time inside your servers
- The timeouts each exchange reports, next to the count in your own logs. Google's RTB graphs, for example, count requests that matched your pretargeting, requests actually sent, valid responses within the timeout, bids and auctions won, and show latency percentiles for each endpoint, the address where your bidder receives requests. Find out what the other exchanges report
- New network connections opened per minute for each exchange, and where the servers that answer that exchange are located
- Bid rate, win rate, spend and revenue per exchange
- Infrastructure cost per exchange per million requests, servers and bandwidth included
Give this page to every team you talk to, and keep today's numbers, because each change a team makes will be judged against them. Several of the causes above can be confirmed or ruled out from these numbers before anyone opens the code. For peaks, the article on traffic spikes lists what else to measure.
Can you recommend a team that has actually fixed bid-path latency?
This article does not rank firms. Whether a team has actually fixed bid-path latency shows in evidence you can check yourself:
- A bidder or exchange that the team built or repaired and that is still in production, with the client's name or the reason it cannot be given
- An engineer at that client who will talk to you without the team on the call
- Before-and-after numbers for named exchange connections, confirmed by the client, such as the timeout rate as the exchange counted it and the cost per million requests
- A diagnosis report or plan from an earlier engagement, with the client's data removed
How do we check that a team is real?
Run these five checks in order, before any work starts on the bid path.
Send the team your page of per-connection numbers and ask what it would test first. A team that has done this work names likely causes for your two connections and the measurement that would confirm or rule out each one. If the first answer names a programming language or a price, the team has probably not read your numbers yet.
Before any rewrite, ask for a written plan. It should state:
- What the team will measure first, and what access it needs
- Which changes come first, starting with the cheapest, and how each can be reversed
- For each of the two exchange connections, the target timeout rate as that exchange reports it, the target cost per million requests, and the traffic level at which both will be checked
- How a rebuilt part runs beside the current one on a copy of live traffic, then takes over one connection at a time
- What the team will leave alone
Pay for the diagnosis and the plan as separate work, and make the report yours whether or not you continue.
Call a client whose bidder or exchange the team built or repaired and that still runs it. Speak to that client's engineers without the team on the call, and ask:
- What did the team measure before it changed anything?
- Which numbers moved, on which exchange connections, and who measured them?
- Who runs and changes the code today?
- What went wrong during the work, and what did the team do about it?
Meet the engineers who will do the work, and ask the lead about the last latency problem they fixed. Someone who did that work can name the exchange and the number that moved, and usually remembers what they tried first that did not help. Put their names in the contract.
Read the ownership terms last. The code should live in your repositories from the first day and be assigned to your company in writing. Anything the team keeps should be listed by name, with a licence to use and change it after the work ends, and the bidder must run without the team's servers or licence keys.
Specialist ad tech engineering firms build and repair bidders and exchanges as their main work. An independent performance engineer can run a diagnosis alone, but a rebuild takes a team. A broader software firm can do the same job if the people it assigns have worked on a bidder before, so ask for those people by name.
A brief that asks for low latency, no GC pauses and high QPS should say what each phrase means:
- “Low latency” means answers inside each exchange's deadline at the 99th percentile, on your own connections. An average across the whole bidder can hide the two connections that fail
- “No GC pauses” is about garbage collection, so it is a requirement on the language the bidder is written in and its runtime. The Rust book says Rust manages memory through “a system of ownership with a set of rules that the compiler checks”, so a Rust bidder has no collector to pause it. Go and Java bidders do have a collector, and its settings can be tuned. A long network route or a queue can still make any bidder late
- “High QPS”, many queries per second, means little without the request size and the number of live campaigns behind it. Ask whether a figure comes from production or from a test, and on whose traffic
A team inside your platform works in your repositories and cloud accounts, and its changes go through your review process. You grant its access and can remove it, and someone on your side sets its priorities. Agree who is on call for the bid path during and after the work, and pair your engineers with the team so the knowledge stays with them.
What are the warning signs when hiring for bid-path work?
- A rewrite or a new language is proposed before anyone has seen your per-connection numbers
- A latency figure, or a saving on your bill, is promised in the first call
- The proof on offer is a benchmark on the team's own hardware with its own requests
- The bidder would run on the team's servers or under its licence, although you asked for a team inside your platform
- No client will take a call, and no live bidder or exchange can be shown
- Every fix in the proposal adds servers
Where does amBrain fit?
In AdTech, amBrain works on DSP development, real-time bidding platforms, and ad exchange engineering.
amBrain diagnoses slow systems in trading and ad tech: the running platform is measured end to end and the report names where the time goes.
amBrain works in three formats: full delivery, a dedicated team, or engineers embedded in your team. The client keeps full ownership of the product and the code, except amBrain's reusable components.
amBrain has been building software since 2019. It built RTBBidder, a demand-side platform, for a client.
This article is not a case study. It does not describe that platform (its design, language or performance), and it does not claim that amBrain has diagnosed or fixed bid-path latency for any client. It gives no prices or timelines.
If two of your exchange connections time out, put their numbers on one page before you talk to anyone. Then ask amBrain, or any other team on your list, what it would measure first on those two connections, and put every team through the same five checks.
Common questions
- Do we need a bidder in every region an exchange sends from? Not always. Google tries to send each request to the trading location closest to the user but does not guarantee it, so receiving all of its impressions takes servers reachable from all four locations. Its testing guide adds that taking impressions from several trading locations usually means running bidding servers in each region. If you want only part of the traffic, Google says servers in some of the locations may be enough, so choose regions by where your campaigns buy
- Will limiting the requests an exchange sends us cost us wins? It depends on which requests the exchange holds back. At Google, when the requests matching a bidder's pretargeting exceed its quota, the excess is throttled, and requests the bidder is likely to respond to are sometimes prioritised based on its recent bidding history. Other exchanges may decide differently, so ask each one