amBrain
AdTechOct 7, 20269 min read

Why Ad Reports Don't Match Bidder Logs, and Who Can Rebuild the Pipeline

Ad MeasurementEvent PipelinesBidder LogsWho Builds It
Error loading image

Ad reports and bidder logs disagree for one of two reasons: events are lost on the way to the reports, or the two sides count them by different rules. One recount on auction IDs tells which. The answer decides whether to fix the pipeline, move to a managed service or rebuild it.

When ad reports never match the bidder logs, two problems can sit behind the same gap. Either impression and click records, called events, are lost on the way to the reports, often at peak traffic, or each side counts them by its own rules. A recount on auction IDs, the codes an exchange gives each auction, tells the two apart. Run it before you hire anyone, because a new pipeline on its own fixes only the first problem.

The short answer: match the auctions your bidder won to impressions in your reports by auction ID, and compare the counts hour by hour. Count again a day later. If the share of wins still missing rises in the busiest hours, the pipeline is losing events, and that takes engineering work. If the events are there but land in another hour, appear twice or are filtered out, the two sides count differently. The fix then is one written set of counting rules for both sides.

What does it mean when ad reports don't match the bidder logs?

Your bidder logs record every auction your bidder joined, each bid it made and each auction it won. Your reports are built from events that arrive later, such as impressions and clicks from browsers and apps. The gap between them has one of two causes, or both:

  • Lost events. The impression or click happened, but its event never reached the reports. A gap that grows at peak traffic points to a part of the pipeline that drops what it cannot handle
  • Different rules. The two sides may close the day in different time zones or put a late event into another hour. One side may count an event twice after a resend, or filter out bot traffic the bidder still counted. Attribution adds rules of its own, such as how long after a click a conversion still counts

Either way, the gap costs money. If you invoice advertisers from your reports, you bill too little for lost impressions and too much for doubled ones, while each exchange usually bills you on its own count. A bidder that learns from these events also sets its bids on the wrong numbers.

How can we tell lost events from events counted by different rules?

Match the two sides event by event. OpenRTB, the IAB Tech Lab protocol for real-time bidding, gives each auction a bid request ID, assigned by the exchange, and each impression in the request an ID of its own. Your bidder can ask the exchange to write these IDs into the notice it sends when you win and into the ad itself. Then each impression event carries the same IDs as your bidder logs.

Neither ID is enough on its own. Under OpenRTB 2.6, each exchange sets its own request IDs, so nothing stops two exchanges from using the same one. An impression ID is unique only inside its request and usually starts at 1. So match on three values together: the exchange, the request ID and the impression ID.

Then run the recount:

  • For each hour of one busy day, count wins in the bidder logs and impressions in your reports, both in UTC, matched on that three-part key
  • Count again the next day. If the gap shrinks, part of it was late events
  • If the share of wins that are still missing rises in the busiest hours, the pipeline is losing events at peak. A share that stays similar from hour to hour is expected, because some wins never become impressions
  • Sort the rest. An event that lands in a different hour on each side means the two use different time zones or cut-off times. A double count usually comes from a retry. If an event is missing only from the final report, a filter such as bot filtering removed it

If your events carry no auction IDs, adding them is the first fix, because without them the recount can compare only totals. The technical article on ad measurement, linked above, covers timestamps and late events in depth.

Where do events go missing at peak traffic?

Take a pipeline built on Kafka and ClickHouse. A collector receives each event from the browser or app and writes it to Kafka, a message queue. A loader reads from Kafka and writes the events in batches to ClickHouse, the analytics database behind your reports. At a peak, events can disappear at each step:

  • The collector is overloaded. Requests it refuses or answers too late are gone unless the browser or app sends them again. A collector that answers “received” before the event is in Kafka also loses whatever it held when it crashes
  • Kafka cannot take events fast enough. Kafka's documentation describes what happens when events arrive faster than they can be passed on. The code that writes to Kafka waits for a set time, then gives up with an error. A collector that ignores that error loses the event without a trace
  • Retries create duplicates. Kafka has a setting that stops its own retries from writing a second copy. Kafka's Java client turns it on by default; client libraries in other languages set their own defaults, and some leave it off. It does not catch duplicates your code creates, such as a batch resent after a restart or a pixel that fires twice
  • Batch inserts fail as a whole. ClickHouse's documentation recommends loading events in large batches. In one of its loading modes, a single malformed row gets the whole batch rejected. A loader that gives up loses every event in it, and one that resends it unchanged hits the same bad row again, so bad rows have to be set aside and counted. When a write times out and nobody knows whether it landed, a resend of exactly the same batch is safe only if the table is set up to drop repeated batches, which a basic self-hosted ClickHouse table does not do by default

Ask your engineers for these records from the busiest hour of a bad day:

  • Errors and timeouts at the load balancer and the collector, and failed writes to Kafka in the collector's logs
  • Consumer lag, meaning how far the loaders fell behind, and any events Kafka deleted unread because they waited longer than it keeps data
  • Failed writes to ClickHouse, with their error messages
  • A count per hour at every step: received by the collector, written to Kafka, read by the loader, stored in ClickHouse

The last record does most of the work. If one step's count falls at the peak while the step before it holds, the events went missing between those two steps.

Should we fix the pipeline, move to a managed service or rebuild it?

Fix the current pipeline:

  • When it fits: the recount shows mostly different rules or a few leaks you can name, and after the fixes the pipeline keeps up with your busiest hour
  • What you pay: engineering time, and the risk that a bigger peak finds the next weak point
  • Who owns the code: you do, and the knowledge stays with your engineers

Move to a managed service, such as Amazon MSK for Kafka or ClickHouse Cloud for ClickHouse, where the provider runs the servers:

  • When it fits: your engineers spend more time keeping Kafka and ClickHouse servers running than on the counting logic, and the losses come from those servers running out of capacity at peak
  • What you pay: a monthly bill that grows with your traffic. Both services charge for compute and storage, and ClickHouse Cloud lists data transfer out and its own ingestion service separately, according to their pricing pages read on 7 October 2026. Estimate your busiest month's bill and the cost of moving away later
  • What it does not fix: the collector and the loader you still run, duplicates, late events and the reconciliation with your bidder logs. The service stores what you send it and knows nothing of what your bidder counted
  • Who owns the code: you do, and the service runs on the provider's terms

Rebuild the pipeline:

  • When it fits: the recount shows losses at several steps, or the design cannot grow with your traffic. Missing auction IDs push toward a rebuild only if adding them means changing every step
  • What you pay: the most engineering work of the three routes, plus a period when the old and the new pipelines run side by side

Which firms can rebuild an ad event pipeline on Kafka and ClickHouse?

Two kinds of firms do this work, and this article ranks neither. Ad tech engineering firms build bidders, exchanges and ad servers, so they know where auction IDs come from, but ask whether they have run an event pipeline at your volume. Data engineering firms build event pipelines for many industries, though not always for ad tech. Ask them whether they have worked with OpenRTB and matched counts with an exchange.

What should we ask before hiring a firm?

Put the same five questions to every firm on your list.

How do you remove duplicates? Listen for an ID that each event gets when it is created, before any retry, built from the auction IDs and the event type. The pipeline removes duplicates twice, once when it stores events and again when reports read them. The second pass matters because ClickHouse removes duplicates in the background at times you cannot plan for. Its documentation says this “does not guarantee the absence of duplicates”.

How do you handle events that arrive late? Listen for a set period for each event type while the numbers can still change, and a point after which they are final. An event that arrives later should still be counted and marked late, not thrown away.

How do you check your numbers against our bidder logs and our partners' reports? Look for a daily recount on the auction IDs and a written reason for each difference. Agree on a gap size at which someone raises the issue with the exchange.

How will you test it above our peak? Ask them to replay a recorded busy day at a higher rate than your busiest hour. During the replay, they switch off a collector and a database server. Afterwards, every event should be either stored or counted as dropped.

Who will own the code? Your company, in writing, with the code in your repositories from the first day. Anything the firm keeps should be listed by name, with a licence to use and change it after the work ends.

What are the warning signs?

  • A rebuild is proposed before anyone has run a recount or opened your bidder logs
  • The firm promises that the new reports will match the bidder exactly
  • The only answer on duplicates is a promise that each event is delivered “exactly once”, with nothing said about duplicates your own code creates, such as a pixel that fires twice
  • The only proof is a load test on made-up events, not a replay of your traffic

Where does amBrain fit?

In AdTech, amBrain works on DSP development, real-time bidding platforms, and ad exchange engineering. That work includes RTBBidder, a demand-side platform amBrain built for a client.

amBrain diagnoses slow systems in trading and ad tech: the running platform is measured end to end and the report names where the time goes.

amBrain has been building software since 2019. It works in three formats: full delivery, a dedicated team, or engineers embedded in your team. The client keeps full ownership of the product and the code, except amBrain's reusable components.

This article is not a case study. It describes no client's event pipeline, and RTBBidder is named only as a DSP amBrain built. Kafka and ClickHouse serve here as an example stack, not as a description of amBrain's projects or the tools it uses. The article gives no prices or timelines.

If your reports and your bidder logs disagree, run the recount first. Then bring its results and the same five questions to every firm on your list, amBrain included.

Common questions

  • Will our reports and the bidder logs ever match 100 percent? No. OpenRTB says the exchange's message that you won is “not necessarily indicative of a delivered, viewed, or billable ad”, and some events are filtered out as bot traffic. Aim for a stable gap with a written reason for each part of it, and agree with each partner which count you bill on
  • Do we have to use Kafka and ClickHouse? No. Other queues and analytics databases can do the same job, and the recount and the five questions apply to any of them
  • How do we keep the old reports running while a new pipeline is built? Feed both pipelines the same events and compare each with the bidder logs every day. Switch the reports you bill on last, after a full billing period with every difference explained

Have a design like this on the table?

Bring your current architecture and the failure mode that worries you, and we will go through it together in half an hour.