amBrain
FinTechOct 7, 20268 min read

Order Book Freezes When the Market Gets Busy? How to Fix the Market Data Feed and Who Can Do It

Market DataOrder BookTrading TerminalWho Builds It
Error loading image

When the order book on trading screens freezes or jumps in a busy market, the market data feed is usually at fault. It loses updates on the way in, builds the book wrong, or sends it too slowly to hundreds of screens. Measure your busiest hour first, then test every firm on a recording of that day.

If the order book on your trading screens freezes, jumps or shows impossible prices in a busy market, the fault usually sits in one of three places. Updates get lost where the exchange feed comes in, the book is put together wrong, or it reaches hundreds of screens too slowly. Measure your busiest hour before you hire anyone, then test every firm you consider on a recording of that day.

The short answer: an exchange numbers every update it sends. A well-built system spots a missing number at once, marks the book as out of date and rebuilds it. Trouble starts when the gap goes unnoticed or the rebuild takes seconds, or when one slow connection holds up every trader. Record the feed of your busiest day and make a replay of it the test every firm has to pass: before you buy a product, and as the pass mark of the first phase when a team builds for you.

What does a broken market data feed look like on the screen?

It shows up at the busiest moments, such as a central bank announcement or the market open. The book on the screen stops for a second or two, then jumps. Cancelled orders stay visible. Sometimes the highest price a buyer is bidding sits above the lowest price a seller is asking, which is called a crossed book. On a single exchange, outside the opening and closing auctions, such orders would match at once. So a crossed book for a single exchange on your screen means your copy of that book is wrong.

Support then gets screenshots from two traders who see different books for the same instrument. Or a trader disputes the price an order traded at, because the screen showed a different one.

Why do updates go missing or arrive out of order?

An exchange's order book feed is a stream of small changes: an order added, an order cancelled, a trade. Your system starts from a full copy of the book, called a snapshot, and applies the changes in order. Each change is numbered, so a missing one can be spotted. Nasdaq's specification for its TotalView-ITCH 5.0 feed says the feed “is made up of a series of sequenced messages”, that is, numbered in order.

In a busy market the flow of changes rises sharply. Some are lost on the way or inside your own servers, and some arrive out of order. If the system misses the gap, it applies whatever arrives and shows a book that no longer matches the exchange. If it notices the gap but takes seconds to recover, the screen stands still.

Exchanges expect their clients to lose updates. CME Group's documentation for its MDP 3.0 feed says that after a gap, “it should be assumed that all books maintained in the client system may no longer have the correct, latest state”.

Where in the system does the feed break?

There are three places, and each needs its own fix. Engineers call the third one fan-out, because one stream of updates fans out to many screens. The technical article linked above covers all three in detail.

The first place is the intake, where the exchange feed arrives. Some exchanges offer ways to get lost data back. One of Nasdaq's delivery protocols, MoldUDP64, lets receivers “detect and re-request missed packets”. CME sends its stream twice, on lines called A and B, and runs a separate feed of snapshots to bring books up to date. None of this helps if your system does not notice the gap.

Then the book is put together, and here the danger is a change applied twice, out of order or on top of the wrong snapshot. Binance, a crypto exchange, sets out in its guide the exact steps for joining a snapshot to the live stream. A book built without them still shows prices and looks fine, but the prices are wrong.

Last comes fan-out, where the book goes out to hundreds of trader sessions, one for each connected screen. A trader on a weak mobile connection, or a terminal that has stalled, reads updates slowly. If the server waits for that session, every other session waits too. Letting that session's backlog of updates grow without limit is no better, because the server runs out of memory and fails for everyone.

In a well-built system the server puts the updates in order once and sends the same result to every session. A session that falls behind either gets the latest picture of the book and skips the steps in between, or the server disconnects it with a reason and the screen reconnects with a fresh copy. A feed can break in more than one place at once.

What should we measure before we hire anyone?

Take your busiest hour of the last month and collect these numbers for it:

  • Gaps: how many times the system found a missing update number on each exchange feed. If the system does not count them, that is your first finding
  • Recovery time: how long each rebuild took, and what traders saw in the meantime
  • Delay: the time from the exchange's timestamp on a message to the moment the update leaves your server for the trader, and on a few test terminals to the screen. Take the median and the 99th percentile, the time that 99 of every 100 updates stay under. Keep your servers' clocks in sync with an accurate time source, or the figures mean nothing
  • Sessions: how many fell behind, by how much, and how many were disconnected, with the reason for each

Then record the raw feed of a busy day exactly as it arrived, with the time each packet came in. Replaying that recording at real speed and faster is the test you give every firm on your list and repeat after every fix.

Fix our own handler, buy one, or rebuild the fan-out?

The software that takes in an exchange feed and keeps the book is called a feed handler. You have three routes, and your measurements should point to one. The last two can be combined.

When does fixing our own feed handler make sense?

Fix the feed handler you have when the measurements point to one clear fault, such as gaps that go unnoticed or a slow rebuild, and the people who know the code are still around.

  • When it fits: a fault you can name and a design that otherwise works
  • What you pay: engineering time and a test bench that replays your recordings
  • Who owns the code: you do
  • The limit: a fix at the intake does not help if the trouble is in sending the book to the screens

When should we buy a feed handler or a managed feed?

A ready-made feed handler is licensed software that connects to an exchange, catches gaps and hands your system a correct, up-to-date book. A managed feed goes further: a market data vendor connects to the exchanges, and you take one stream in one format from it.

  • When it fits: many exchanges in standard formats, and no wish to follow every change each exchange makes to its feed
  • What you pay: the licence and the work of connecting the product to your system. The exchanges' own data fees and licence terms usually still apply, whoever delivers the data
  • What stays your job: getting the book to hundreds of trader screens and handling slow sessions
  • Who owns the code: the vendor owns the product. You own the integration and everything after it

When should we hire a team to rebuild the fan-out?

Rebuild the layer that sends data to traders when the intake works but the trouble stays. Traders still see different books, one slow session drags down the rest, or you plan to serve many more sessions than today.

  • When it fits: the delay grows between your servers and the screens, not between the exchange and your servers
  • What you pay: engineering and testing time, then the people who run the system after launch
  • Who owns the code: you, if the contract says so

How do we check a firm before we hire it?

Put the same five tests to every firm on your list:

  • A replay of your recording, at real speed and several times faster, through whatever the firm delivers or demonstrates. Compare its gaps, rebuilds, recovery times and delays with those of your current system
  • How it catches a gap. A good answer says the system applies a change only when its number is the next one expected. If a number is missing, the system waits briefly, then asks for the change again or rebuilds the book. Until then it marks the book as out of date. Ask what traders see meanwhile
  • What happens to one slow trader. Ask the firm to slow one session down on purpose during the replay. The delays of the other sessions should not change
  • How it measures delay: from which timestamp to which point, at what percentile, load and hardware, and how the clocks are kept in sync
  • Who owns the code, and on what terms you use any parts the firm keeps as its own

What are the warning signs?

  • “We will add more servers” as the first answer, before anyone has seen your measurements
  • “We use a reliable connection, so nothing gets lost.” Binance sends its updates over a WebSocket, a connection that does not lose data in transit, and its guide still tells clients what to do when events are skipped

Where does amBrain fit?

amBrain builds algorithmic trading infrastructure: order execution, market data and pre-trade risk controls.

One line on amBrain's website reads: “Trading terminal development, order management systems, and FIX protocol exchange integration.”

amBrain diagnoses slow systems in trading and ad tech: the running platform is measured end to end and the report names where the time goes.

amBrain has been building software since 2019. It describes its team in one line: “A team of up to 40 people, about 75% of them senior.” It works in three formats: full delivery, a dedicated team, or engineers embedded in your team. The client keeps full ownership of the product and the code, except amBrain's reusable components.

This article is not a case study and describes no client work. It quotes no latency figure for any system amBrain has built, and no prices or timelines.

If amBrain is on your shortlist, put the same five questions to it as to every other firm, and make your recording the pass mark for any work you agree.

Common questions

  • Will more servers fix it? Not on their own. If the system applies changes out of order or waits for its slowest session, more servers repeat the same fault. Add servers once the measurements show the current ones are running out of capacity
  • Can we merge updates and send fewer of them? For the order book, yes. Engineers call it conflation, and some exchanges do it themselves. Binance's documentation sets the update speed of its spot stream of order book changes at 1000 ms or 100 ms. Tell traders the stream shows the latest picture, not every step. Do not merge trades or order confirmations, because dropping one makes the trade history wrong
  • Do we have to rewrite it in Rust or C++? Not necessarily. Missed gaps and a server that waits for its slowest session are design faults that a new language does not remove. Rust and C++ have no garbage collector, the automatic memory clean-up that can pause a program written in a language such as Java or Go. That helps in the parts of the system that handle every update. Whatever the language, ask for the measured delay at peak
  • How long does a fix take? It depends on where the fault is and how many exchanges and sessions you have. Ask each firm to price and schedule a first phase that covers the measurements and a test setup that replays your recording. Use the five tests above as pass marks

Have a design like this on the table?

Bring your current architecture and the failure mode that worries you, and we will go through it together in half an hour.