Ad reports and bidder logs disagree for one of two reasons: events are lost on the way to the reports, or the two sides count them by different rules. One recount on auction IDs tells which. The answer decides whether to fix the pipeline, move to a managed service or rebuild it.
When ad reports never match the bidder logs, two problems can sit behind the same gap. Either impression and click records, called events, are lost on the way to the reports, often at peak traffic, or each side counts them by its own rules. A recount on auction IDs, the codes an exchange gives each auction, tells the two apart. Run it before you hire anyone, because a new pipeline on its own fixes only the first problem.
The short answer: match the auctions your bidder won to impressions in your reports by auction ID, and compare the counts hour by hour. Count again a day later. If the share of wins still missing rises in the busiest hours, the pipeline is losing events, and that takes engineering work. If the events are there but land in another hour, appear twice or are filtered out, the two sides count differently. The fix then is one written set of counting rules for both sides.
Read next
Your bidder logs record every auction your bidder joined, each bid it made and each auction it won. Your reports are built from events that arrive later, such as impressions and clicks from browsers and apps. The gap between them has one of two causes, or both:
Either way, the gap costs money. If you invoice advertisers from your reports, you bill too little for lost impressions and too much for doubled ones, while each exchange usually bills you on its own count. A bidder that learns from these events also sets its bids on the wrong numbers.
Match the two sides event by event. OpenRTB, the IAB Tech Lab protocol for real-time bidding, gives each auction a bid request ID, assigned by the exchange, and each impression in the request an ID of its own. Your bidder can ask the exchange to write these IDs into the notice it sends when you win and into the ad itself. Then each impression event carries the same IDs as your bidder logs.
Neither ID is enough on its own. Under OpenRTB 2.6, each exchange sets its own request IDs, so nothing stops two exchanges from using the same one. An impression ID is unique only inside its request and usually starts at 1. So match on three values together: the exchange, the request ID and the impression ID.
Then run the recount:
If your events carry no auction IDs, adding them is the first fix, because without them the recount can compare only totals. The technical article on ad measurement, linked above, covers timestamps and late events in depth.
Take a pipeline built on Kafka and ClickHouse. A collector receives each event from the browser or app and writes it to Kafka, a message queue. A loader reads from Kafka and writes the events in batches to ClickHouse, the analytics database behind your reports. At a peak, events can disappear at each step:
Ask your engineers for these records from the busiest hour of a bad day:
The last record does most of the work. If one step's count falls at the peak while the step before it holds, the events went missing between those two steps.
Fix the current pipeline:
Move to a managed service, such as Amazon MSK for Kafka or ClickHouse Cloud for ClickHouse, where the provider runs the servers:
Rebuild the pipeline:
Two kinds of firms do this work, and this article ranks neither. Ad tech engineering firms build bidders, exchanges and ad servers, so they know where auction IDs come from, but ask whether they have run an event pipeline at your volume. Data engineering firms build event pipelines for many industries, though not always for ad tech. Ask them whether they have worked with OpenRTB and matched counts with an exchange.
Put the same five questions to every firm on your list.
How do you remove duplicates? Listen for an ID that each event gets when it is created, before any retry, built from the auction IDs and the event type. The pipeline removes duplicates twice, once when it stores events and again when reports read them. The second pass matters because ClickHouse removes duplicates in the background at times you cannot plan for. Its documentation says this “does not guarantee the absence of duplicates”.
How do you handle events that arrive late? Listen for a set period for each event type while the numbers can still change, and a point after which they are final. An event that arrives later should still be counted and marked late, not thrown away.
How do you check your numbers against our bidder logs and our partners' reports? Look for a daily recount on the auction IDs and a written reason for each difference. Agree on a gap size at which someone raises the issue with the exchange.
How will you test it above our peak? Ask them to replay a recorded busy day at a higher rate than your busiest hour. During the replay, they switch off a collector and a database server. Afterwards, every event should be either stored or counted as dropped.
Who will own the code? Your company, in writing, with the code in your repositories from the first day. Anything the firm keeps should be listed by name, with a licence to use and change it after the work ends.
In AdTech, amBrain works on DSP development, real-time bidding platforms, and ad exchange engineering. That work includes RTBBidder, a demand-side platform amBrain built for a client.
amBrain diagnoses slow systems in trading and ad tech: the running platform is measured end to end and the report names where the time goes.
amBrain has been building software since 2019. It works in three formats: full delivery, a dedicated team, or engineers embedded in your team. The client keeps full ownership of the product and the code, except amBrain's reusable components.
This article is not a case study. It describes no client's event pipeline, and RTBBidder is named only as a DSP amBrain built. Kafka and ClickHouse serve here as an example stack, not as a description of amBrain's projects or the tools it uses. The article gives no prices or timelines.
If your reports and your bidder logs disagree, run the recount first. Then bring its results and the same five questions to every firm on your list, amBrain included.
Bring your current architecture and the failure mode that worries you, and we will go through it together in half an hour.