<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: amBrain</title>
    <description>The latest articles on DEV Community by amBrain (@ambrain).</description>
    <link>https://dev.to/ambrain</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4109909%2Ffb4dd3ef-b04c-4772-a8f3-e72f7b194ece.png</url>
      <title>DEV Community: amBrain</title>
      <link>https://dev.to/ambrain</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ambrain"/>
    <language>en</language>
    <item>
      <title>Your Own DSP, Ad Exchange or Ad Network: Build, License or Rent?</title>
      <dc:creator>amBrain</dc:creator>
      <pubDate>Wed, 23 Sep 2026 10:01:29 +0000</pubDate>
      <link>https://dev.to/ambrainorg/your-own-dsp-ad-exchange-or-ad-network-build-license-or-rent-4ged</link>
      <guid>https://dev.to/ambrainorg/your-own-dsp-ad-exchange-or-ad-network-build-license-or-rent-4ged</guid>
      <description>&lt;p&gt;Originally published at &lt;a href="https://ambrain.org/blog/own-dsp-ad-exchange-build-license-rent/" rel="noopener noreferrer"&gt;https://ambrain.org/blog/own-dsp-ad-exchange-build-license-rent/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When someone says “we want our own DSP”, they can mean one of three different businesses. A DSP, or demand-side platform, buys ad space automatically on behalf of advertisers. An ad exchange runs the auction where publishers and apps sell that space.&lt;/p&gt;

&lt;p&gt;An ad network signs up publishers, packages their space and resells it to advertisers. A DSP, an ad exchange and an ad network share a vocabulary, but they need different software, different partners and a very different amount of work.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The short answer: first decide whether you are building a DSP that buys for advertisers, an ad exchange that runs the auction where publishers sell their space, or an ad network that resells space it has signed up. Then settle where your partners and traffic come from, what data you are allowed to use, and how money moves between everyone involved. Rent a platform to test demand, license one when your business model is standard, and build a custom system when the bidding logic, the data or the margin is what your company sells.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The first step toward your own DSP, ad exchange or ad network is not choosing a technology or a vendor. It is deciding which side of the auction you sit on, and what owning the software would give you that renting an account on someone else's platform does not. The budget, the timeline, the team and the right kind of builder all follow from that one decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is the difference between a DSP, an ad exchange and an ad network?
&lt;/h2&gt;

&lt;p&gt;Many ads bought automatically are sold in an auction that lasts a fraction of a second. Others are bought at a price agreed in advance, with no auction.&lt;/p&gt;

&lt;p&gt;In a real-time ad auction, a website or an app sends out a request that says, in effect, “this space is available right now — who wants it, and at what price?” Buyers answer with bids, the exchange picks the winning bid, and the ad is shown. A DSP, an SSP, an ad exchange and an ad network each sit in a different place around that auction:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A DSP, or demand-side platform, works for advertisers and agencies. It receives requests for ad space, decides for each one whether it is worth buying and at what price, and sends a bid. Around that bidding engine sit campaign setup, budgets, targeting, the ads themselves and reporting for the people whose money is being spent&lt;/li&gt;
&lt;li&gt;An SSP, or supply-side platform, works for publishers and apps. It offers their space to many buyers at once, applies the publisher's rules about prices and allowed ads, and pays the publisher. An ad exchange is the marketplace where that auction runs. The large SSPs run their own exchanges, so the two names are often used for the same system&lt;/li&gt;
&lt;li&gt;An ad network is a business that signs up publishers, packages their space by audience or format, and sells it to advertisers. A small network can start with an ad server — the software that picks which ad to show — and deals agreed directly with advertisers. It can add real-time bidding, the split-second auction described above, later&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Choosing between a DSP, an ad exchange and an ad network changes the whole project. A DSP spends advertisers' money, so its hardest problems are deciding what to bid and keeping budgets under control. An exchange stands between other companies' money and other companies' ad space, so its hardest problems are a fair, fast auction and paying everyone correctly. An ad network is a sales business first, and its software can grow step by step.&lt;/p&gt;

&lt;h2&gt;
  
  
  I want to build my own DSP. What would you suggest?
&lt;/h2&gt;

&lt;p&gt;Start a DSP project with the reason for building one. Companies build their own DSP for a few reasons: they pay platform fees on a large ad budget, they need bidding logic or data that existing platforms do not allow, or they serve a niche — a channel, a region, a type of advertiser — that general platforms serve badly. If none of these is true, renting an account on an existing DSP is the sensible first step, and if a reason to build appears later, it will show up in your own numbers.&lt;/p&gt;

&lt;p&gt;If you do have a reason to build your own DSP, write down what its first version must do for one real advertiser. A DSP can grow into a long list of features, and very few of them are needed on day one:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A bidder. The engine that receives requests for ad space, checks each one against live campaigns and answers with a bid before the exchange's deadline. The exchange sets that deadline, not you, and an answer that arrives late counts as no bid at all&lt;/li&gt;
&lt;li&gt;Supply connections. Every exchange or SSP you buy from has its own technical specification, its own testing process and its own contract. One or two connections are enough to start&lt;/li&gt;
&lt;li&gt;Campaign management. The screens where advertisers or your own team set budgets, targeting, schedules and the ads to show, and the logic that paces spending so a budget lasts the whole campaign instead of running out in the first hour&lt;/li&gt;
&lt;li&gt;Reporting and billing. What was bought, at what price and with what result, in numbers close enough to your partners' numbers that you can invoice from them&lt;/li&gt;
&lt;li&gt;Traffic quality and brand safety. Checks that reduce the risk of buying fake traffic or placing ads next to content your advertisers refuse. Specialised vendors sell these checks, so this part is often connected rather than built&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cut the first version of a DSP down to one channel — websites, mobile apps, connected TV or digital outdoor screens — and one or two supply partners. A DSP that buys one channel well gives you something to sell while the rest is still a plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  We want to launch an ad exchange. Should we build it, license it or rent it?
&lt;/h2&gt;

&lt;p&gt;There are three routes to a working ad exchange or DSP, and the right one depends on how much of your business lives in the software. Here is what each route gives you, what it takes away, and when it is the right call.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rent: an account on an existing platform, or a white-label version its provider runs under your brand. You get: the fastest start, no servers to run, and a known fee, often a percentage of the ad spend that passes through. You give up: control over the auction and the bidding logic, part of your margin on every ad shown, and some control over your data. Right when: you are testing demand and the software is not what makes you different&lt;/li&gt;
&lt;li&gt;License an existing ad exchange or ad server and run it yourself. You get: a working product with the standard features already built, and room to configure it. You give up: you follow the vendor's roadmap, licence fees often grow with the traffic you process, and connecting the product to your own data and billing is real work. Right when: your business model is standard for your market and your advantage is in sales, supply or service&lt;/li&gt;
&lt;li&gt;Build a custom platform. You get: exactly the auction or bidding logic you want, your data in your own systems, no platform fee taken from every ad shown, and ownership of what your contract assigns to you. You give up: the cost and the time of building the first version, and the duty to run the system, pay for its servers and keep up with industry standards afterwards. Right when: the bidding logic, the data or the margin is your product, or an existing platform blocks something central to your business&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Mixed answers to build, license or rent are common. An ad network rents an ad server to launch and builds its own once the rental fee starts to hurt. An exchange licenses reporting and billing and builds the auction itself, because the auction is where it has to behave differently from everyone else. Deciding component by component is often smarter than answering once for the whole ad tech project.&lt;/p&gt;

&lt;h2&gt;
  
  
  I want to build an ad network. Where do I start?
&lt;/h2&gt;

&lt;p&gt;Start an ad network with its two sides, not with software: the publishers and apps that will give you their ad space, and the advertisers who will pay for it. The first question is where each side comes from, and why publishers and advertisers would work with you instead of an existing network or exchange.&lt;/p&gt;

&lt;p&gt;The first version of ad network software can be modest: an ad server that decides which ad to show, a dashboard where publishers see what they earned, reporting for advertisers, and a way to pay everyone accurately. To sell to DSPs, the network either connects its ad space to an existing SSP or runs its own auction. To buy from exchanges, it needs a bidder, the same kind of software a DSP runs. Either way, the build gets bigger.&lt;/p&gt;

&lt;p&gt;Before any code, decide how an ad network will check traffic quality. A network that pays publishers for fake traffic loses its advertisers, and adding those checks after launch is harder than designing them in from the start.&lt;/p&gt;

&lt;h2&gt;
  
  
  What has to be decided before anyone writes code?
&lt;/h2&gt;

&lt;p&gt;Before anyone writes code for an ad platform, seven questions need written answers. None of them needs an engineering background to answer, and a company that starts building before it has the answers in writing is guessing with your budget.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which side of the auction are you on? Buying for advertisers, selling for publishers, or both. Doing both at once raises conflict-of-interest questions your partners will ask, so decide it openly&lt;/li&gt;
&lt;li&gt;Which channels and formats? Website banners, video, mobile apps, connected TV, digital outdoor screens, audio. Each channel has its own standards and its own buyers, and video on connected TV is a different build from banners on websites&lt;/li&gt;
&lt;li&gt;Where do the partners come from? Name the exchanges, SSPs, DSPs, publishers or advertisers you already have agreements with, or are negotiating with. Every integration has its own specification, testing and contract, and agreements move on their own calendar&lt;/li&gt;
&lt;li&gt;How much traffic, and how fast? The number of requests you receive, not the number of ads you win, decides how many servers you pay for. The auction deadline is set by your partners. Get both written down as expectations before design starts&lt;/li&gt;
&lt;li&gt;What data may you use, and where? Which user data you collect or buy, what consent you need, and in which countries. Privacy laws decide what the system may store and pass on to partners. Confirm the specifics with a lawyer in each market before committing to a design&lt;/li&gt;
&lt;li&gt;How does money move? Who pays whom, how invoices are produced, and how you settle the differences between your numbers and your partners' numbers&lt;/li&gt;
&lt;li&gt;Who runs it after launch? Partners change their specifications, traffic patterns shift, and a bidding system needs someone watching it every day. Decide now whether that is your team, the builder under a support agreement, or someone you have not hired yet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Written answers to these seven questions are your brief. Handed to three different builders, they produce three proposals you can compare. Without them, you receive three sales presentations that cannot be compared at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  How much does it cost to build a DSP or an ad exchange?
&lt;/h2&gt;

&lt;p&gt;This article gives no prices. The same phrase — “our own DSP” — covers projects that differ many times over in size, and a number quoted before anyone has heard your answers to the seven questions above is a sales number, not an estimate.&lt;/p&gt;

&lt;p&gt;What is useful is knowing which decisions move the price of an ad platform, because these are the levers you control:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which product it is. A bidder that buys through one exchange, a full DSP with campaign tools for outside advertisers, and an exchange that runs auctions for other companies' money are three different sizes of project&lt;/li&gt;
&lt;li&gt;How many integrations. Each exchange, SSP, DSP or data provider adds its own specification, testing and later changes. The second integration costs less than the first; the tenth still costs something&lt;/li&gt;
&lt;li&gt;Traffic volume. Server and network costs grow with the requests you answer, including the auctions you lose. Filtering out requests you will never bid on is one of the first savings to design in&lt;/li&gt;
&lt;li&gt;Channels and formats. Each new channel — video, mobile apps, connected TV, outdoor screens — brings its own standards, creative checks and reporting&lt;/li&gt;
&lt;li&gt;Data and targeting. Your own audience data, bought data and matching users across devices each add storage, processing and privacy work&lt;/li&gt;
&lt;li&gt;Reporting and billing. Invoices that advertisers and publishers accept need numbers that match your partners' numbers closely. This is a real part of the build, not paperwork at the end&lt;/li&gt;
&lt;li&gt;Traffic quality and brand safety. Checks you build, vendors you connect, or both&lt;/li&gt;
&lt;li&gt;Who runs it. The team that watches the system every day, and the support hours your partners expect&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The cheapest lever on the price of an ad platform is scope. Cutting the first version to one channel, one or two partners and one way of buying or selling saves more money than most technical choices made afterwards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who builds custom ad tech software for publishers and advertisers?
&lt;/h2&gt;

&lt;p&gt;Three kinds of suppliers answer when you say “we want our own DSP”, and they are easy to confuse because they use the same words. Platform vendors sell you accounts or licences on their own product. White-label providers run their platform for you under your brand.&lt;/p&gt;

&lt;p&gt;Engineering companies build a system to your specification, and the contract decides how much of it becomes yours. All three kinds of ad tech supplier can help; ask each of them what you will own at the end.&lt;/p&gt;

&lt;p&gt;If you want something custom built — a bidder, a DSP, an ad exchange or the auction side of an ad network — ask a supplier for these proofs before anything else:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A system they built that is live with real traffic, with the client named, or a clear reason why the client cannot be named&lt;/li&gt;
&lt;li&gt;Who operates that system today, and what happens when a partner's traffic suddenly jumps or a connection breaks&lt;/li&gt;
&lt;li&gt;A walk-through of something working: a campaign set up, a bid placed, a report produced. A live system tells you more than any slide&lt;/li&gt;
&lt;li&gt;What they will hand over: the code, the build and deployment instructions, the documentation — and whether anything in the system stays theirs&lt;/li&gt;
&lt;li&gt;How they handled a partner changing its specification or its auction deadline, with a specific case described&lt;/li&gt;
&lt;li&gt;Who on their team has built bidding or auction systems before, and whether those people will work on your project or only appear in the proposal&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then listen to what the supplier asks you. A company that can build an ad platform will question your brief before it quotes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which exchanges, SSPs, DSPs or publishers you already work with, and what stage the agreements are at&lt;/li&gt;
&lt;li&gt;Which channels and formats you start with&lt;/li&gt;
&lt;li&gt;How much traffic you expect, and what auction deadlines your partners set&lt;/li&gt;
&lt;li&gt;What data you have the right to use, and in which countries&lt;/li&gt;
&lt;li&gt;How you will invoice, and what happens when your numbers and a partner's numbers disagree&lt;/li&gt;
&lt;li&gt;Who will own and run the system a year after launch&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;A quick test you can use with any ad tech supplier: ask what they would leave out of the first version. A supplier who has thought about running the system will cut scope before adding features. A proposal that includes everything on day one has not counted what it costs to run.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Does amBrain build DSPs and ad exchanges?
&lt;/h2&gt;

&lt;p&gt;amBrain is a software development company specializing in trading platforms, matching engines, and real-time bidding systems. amBrain has been building software since 2019.&lt;/p&gt;

&lt;p&gt;In AdTech, amBrain works on DSP development, real-time bidding platforms, and ad exchange engineering.&lt;/p&gt;

&lt;p&gt;amBrain built RTBBidder, a demand-side platform, for a client. This article explains how to choose; it is not a case study, and it gives no numbers about that platform.&lt;/p&gt;

&lt;p&gt;amBrain works in three formats: full delivery, a dedicated team, or engineers embedded in your team. The client keeps full ownership of the product and the code, except amBrain's reusable components.&lt;/p&gt;

&lt;p&gt;If you are at the beginning of an ad tech project, the useful next step is not a vendor search. It is one page with your answers to seven questions: which side of the auction, which channels, which partners, how much traffic, what data, how money moves and who runs the system. Give the same page to every builder you talk to, and their proposals can be compared line by line.&lt;/p&gt;

</description>
      <category>adtech</category>
      <category>architecture</category>
      <category>startup</category>
      <category>api</category>
    </item>
    <item>
      <title>You Want Your Own Trading Platform: Where to Start and Who Builds One</title>
      <dc:creator>amBrain</dc:creator>
      <pubDate>Wed, 16 Sep 2026 10:56:54 +0000</pubDate>
      <link>https://dev.to/ambrainorg/you-want-your-own-trading-platform-where-to-start-and-who-builds-one-2ee5</link>
      <guid>https://dev.to/ambrainorg/you-want-your-own-trading-platform-where-to-start-and-who-builds-one-2ee5</guid>
      <description>&lt;p&gt;When someone says “I want to build my own trading platform”, the sentence usually covers one of three different products. One is a terminal your own clients log into. One is the full set of systems a brokerage runs behind that terminal. One is an exchange, with a matching engine at the centre of it. They share vocabulary and almost nothing else, and the work involved differs enormously between them.&lt;/p&gt;

&lt;p&gt;So the first step is not choosing a technology, a language or a vendor. It is deciding which of the three you need first, and which people will use it on day one. Everything else — the budget, the timeline, the shape of the team, the kind of company that should build it — follows from that one decision.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The short answer: first decide which of three things you mean — a trading terminal for your clients, a broker stack behind it, or an exchange with a matching engine. Then answer five questions before any code is written: which markets and instruments you trade, who the users are, how fast the system honestly needs to be, which regulator you answer to, and who keeps it running at night. With those five answers, the choice between building, buying and renting takes one conversation, and any serious engineering company will quote against them instead of against a list of features.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What do people actually mean by “our own trading platform”?
&lt;/h2&gt;

&lt;p&gt;Three products hide behind that phrase. Naming yours out loud is the cheapest decision you will make, because it changes the size of the project more than any other choice in it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A trading terminal. The screen your clients use: prices, charts, an order ticket, positions, balances, history. It connects to a broker, a venue or an exchange that already exists. You are building the experience, not the market&lt;/li&gt;
&lt;li&gt;A brokerage stack. Everything behind that screen: client accounts, money in and out, risk limits, routing orders to the venues you have access to, end-of-day reconciliation, the reports your regulator asks for. The terminal is one part of it&lt;/li&gt;
&lt;li&gt;An exchange. You are not sending orders somewhere else — you are the place where they meet. That means a matching engine, the software that pairs a buy order with a sell order. Around it: the live list of orders waiting to be filled, prices and trades sent out to the firms connected to you, an arrangement for moving the money and the assets once a trade is agreed, a way to spot abusive trading, and a rulebook you publish and enforce&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In the conversations amBrain has, “exchange” often turns out to mean the first or the second item. That is not a failure of vocabulary — the words are used loosely everywhere. But a terminal and an exchange are different businesses with different licences, and building the wrong one first is the most expensive mistake available at this stage.&lt;/p&gt;

&lt;p&gt;There is also a common in-between case: you already use a third-party platform such as MetaTrader and you want your clients on something of your own instead. That is usually the terminal case with a migration attached — existing accounts, existing habits, and a period where both systems are live at once.&lt;/p&gt;

&lt;p&gt;A crypto exchange is the third product with a different set of surrounding problems. The matching engine, the order book and the market data feed are the same kind of work. What differs is everything around them: holding client funds in wallets instead of at a bank, moving assets on and off chains that have their own outages and fees, and a licensing picture that changes by country and by year. Someone who has built a matching engine can build yours; ask them separately who handles custody and the chain side.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I know which of the three I need first?
&lt;/h2&gt;

&lt;p&gt;Answer it by user, not by feature. Write down the first person who will use the system, on a specific day, to do a specific thing. If that person is your client placing a trade, you need a terminal. If that person is your own operations team taking in money, checking limits and routing orders, you need the broker stack. If that person is another firm connecting to you to trade against other members, you need an exchange.&lt;/p&gt;

&lt;p&gt;Almost nobody needs all three at once, and almost everyone eventually builds outward from one of them. Starting with the piece that touches a real user first gives you something to test and sell while the rest is still a plan.&lt;/p&gt;

&lt;h2&gt;
  
  
  What has to be decided before anyone writes code?
&lt;/h2&gt;

&lt;p&gt;Five questions. They are not technical questions, and you do not need an engineering background to answer them. A company that starts building before it has them in writing is guessing, and you will pay for the guess later.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What is traded, and where does it come from? Stocks, futures, currencies, crypto, or several at once. Name the specific brokers, exchanges or liquidity providers you can already connect to, or the ones you are negotiating with. This decides more of the work than anything else on the list&lt;/li&gt;
&lt;li&gt;Who are the users, and how many? Retail clients on phones, professional traders at desks, or your own staff. A hundred people or a hundred thousand. Ten users who trade constantly are a different system from ten thousand who trade once a month&lt;/li&gt;
&lt;li&gt;How fast does it honestly need to be? Many platforms only need to feel instant to a human being, which is a comfortable target. A few need to compete with other machines, which is a much harder and more expensive target. Answer this honestly — paying for speed you do not need is a classic way to burn a budget&lt;/li&gt;
&lt;li&gt;Which regulator do you answer to, and where? The country you operate in decides what you must record, what you must report, how long you keep it, and what you are allowed to show to whom. Licensing and market data agreements run on their own calendar, and no engineering team can make them go faster&lt;/li&gt;
&lt;li&gt;Who keeps it running at night? A trading system is not delivered once and left alone. Markets open while you sleep, feeds break, venues change something without telling you. Decide now whether that is your team, the builder under a support agreement, or someone you have not hired yet&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These five answers are your brief. Handed to three different companies, they will produce three comparable proposals. Without them, you will receive three sales decks that cannot be compared at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should I build it, buy it, or rent it?
&lt;/h2&gt;

&lt;p&gt;There are three routes to a working platform, and the right one depends on how much of your business lives in the software. Here is what each one gives you, what it takes away, and when it is the right call.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rent a white-label platform. You get: the fastest start, someone else running the servers, a known monthly cost. You give up: the look and the workflow are largely fixed, your data sits with the provider, and moving away later is a project in itself. Right when: you are testing demand, you need to be live soon, and nothing about your offer depends on the software being different&lt;/li&gt;
&lt;li&gt;Buy or licence an existing platform and configure it. You get: a mature product with features you would take years to write, plus room to adjust it. You give up: you live inside someone else's roadmap, the configuration work is real work, and integration with your own systems is usually the hard part. Right when: your business is standard for your market and your difference is in pricing, service or reach&lt;/li&gt;
&lt;li&gt;Build a custom platform. You get: exactly the workflow you want, systems that fit how you actually operate, and ownership of the result. You give up: time before the first version, and the duty to keep it alive afterwards. Right when: the software is the product, an existing platform blocks something central to your business, or your speed, instruments or rules do not fit what is on the market&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Platforms often end up mixed. A broker rents to launch, then builds the one part that clients actually choose them for. An exchange licences the surrounding systems and builds the matching engine itself, because that is the part it cannot afford to have behave like everyone else's. Splitting the decision per component is usually smarter than answering it once for the whole project.&lt;/p&gt;

&lt;h2&gt;
  
  
  How much does it cost, and how long does it take?
&lt;/h2&gt;

&lt;p&gt;amBrain does not quote a price for this kind of work before the scope exists, and a price quoted before anyone has heard your answers to the five questions above is worth little. The same sentence — “a trading platform” — covers products that differ by an order of magnitude in size. A number quoted before the scope exists is a sales number, not an estimate.&lt;/p&gt;

&lt;p&gt;What is useful is knowing which decisions move the number, because these are the levers you actually control:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which of the three products it is. A terminal on top of an existing broker, a full brokerage stack, and an exchange with a matching engine are three different sizes of project&lt;/li&gt;
&lt;li&gt;How many external connections. Every broker, venue or liquidity provider you connect to has its own protocol, its own quirks and its own certification process. The second connection costs less than the first; the tenth still costs something&lt;/li&gt;
&lt;li&gt;How many asset classes. Adding a second one is rarely a small change. A different asset class comes with different contracts, different rules about how much money a client has to hold against a position, and a different process for completing the trade after it is agreed&lt;/li&gt;
&lt;li&gt;The speed target. Feeling instant to a person and competing with other machines are separated by a large amount of engineering&lt;/li&gt;
&lt;li&gt;Regulation and reporting. Audit trails, record keeping, client reporting and the evidence a regulator expects are a real part of the build, not paperwork at the end&lt;/li&gt;
&lt;li&gt;How many users and what support hours. Serving a desk of traders and serving a hundred thousand retail accounts are different systems and different running costs&lt;/li&gt;
&lt;li&gt;Which clients you serve. Web, mobile, desktop — each one is a separate surface to build and keep updated&lt;/li&gt;
&lt;li&gt;Migration. Moving live accounts, balances and history off an existing platform without a break is often larger than the new features everyone is excited about&lt;/li&gt;
&lt;li&gt;Market data rights. What you may display, to whom, and at what delay is a commercial agreement that shapes the product&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Timelines move on the same levers, plus two you do not own: licensing, and the certification windows your venues and brokers schedule. A proposal that promises a date without naming those dependencies has not looked at them.&lt;/p&gt;

&lt;p&gt;The cheapest lever is scope. Cutting the first version down to one market, one asset class and one group of users usually saves more money than any technical choice made afterwards.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which companies build matching engines and trading platforms, and how do I check one is real?
&lt;/h2&gt;

&lt;p&gt;Three kinds of suppliers exist, and they are easy to confuse because they use the same words. Platform vendors licence their own product to you. White-label providers run their platform for you under your brand. Engineering companies build a system that becomes yours. All three will answer the phone when you say “I want a trading platform”, and only one of them is answering the question you asked.&lt;/p&gt;

&lt;p&gt;If you want something custom built — a terminal, a broker stack or a matching engine — ask for these proofs before anything else:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A system they built that is in production with real money going through it, with the client named, or an explanation of why it cannot be named&lt;/li&gt;
&lt;li&gt;Who operates that system today, and what happens when it breaks at three in the morning&lt;/li&gt;
&lt;li&gt;A walk-through of something working. A live demo of a real system tells you more than any slide&lt;/li&gt;
&lt;li&gt;What they will hand over: the code, the build instructions, the deployment steps, the documentation — and whether anything in the system stays theirs&lt;/li&gt;
&lt;li&gt;How they have handled a venue or feed going down, with a specific incident described&lt;/li&gt;
&lt;li&gt;Who on their team has done this before, and whether those people will be on your project or only in the proposal&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then listen to what they ask you. A company that can build this will not accept your brief at face value — it will interrogate it before quoting.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which venues, brokers or liquidity providers you already have access to, and what stage the agreements are at&lt;/li&gt;
&lt;li&gt;What your users do all day, and which of the three products you actually mean&lt;/li&gt;
&lt;li&gt;What speed means for your business, stated as something measurable rather than “fast”&lt;/li&gt;
&lt;li&gt;Which regulator you report to and what evidence it expects&lt;/li&gt;
&lt;li&gt;What happens when part of the price feed goes missing, or a venue rejects an order — the unglamorous cases that decide whether a platform is trusted&lt;/li&gt;
&lt;li&gt;Who will own and run the system a year after launch&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;A quick test that works on any supplier: describe your idea in three sentences and see what comes back. A proposal and a price in the first reply means the questions were skipped. Six questions and no price yet means someone is trying to find out what you actually need. The second reply is the one worth continuing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Where amBrain fits into this
&lt;/h2&gt;

&lt;p&gt;amBrain is an engineering company in Yerevan, Armenia. amBrain has been building software since 2019.&lt;/p&gt;

&lt;p&gt;amBrain built the trading terminal for Spectre Trade.&lt;/p&gt;

&lt;p&gt;amBrain has also built a mini-exchange that runs in production on MOEX colocation.&lt;/p&gt;

&lt;p&gt;If your question is which companies build matching engines, this is the part of the list amBrain works in: amBrain is a software development company specializing in trading platforms, matching engines, and real-time bidding systems.&lt;/p&gt;

&lt;p&gt;The parts where speed matters are written in Rust. Two numbers here are measured on the paths amBrain builds. Market data latency under 5 ms: that is how long a price change takes to reach the screen or the system waiting for it. Risk latency under 1 ms: that is how long the check takes that decides whether an order is allowed through before it goes to the market. Both figures describe the paths we build, not any one system named above.&lt;/p&gt;

&lt;p&gt;We work in three formats: full delivery, where we build and hand over the finished system; a dedicated team working on your product and nothing else; or our engineers embedded in a team you already have. The product and the code stay with the client, except for amBrain's own reusable components.&lt;/p&gt;

&lt;p&gt;If you are at the beginning of this, the useful next step is not a vendor search. It is writing down the answers to the five questions above in one page. With that page, a conversation with any builder — us or anyone else — starts at what you need and how it can be built, instead of at a demo of something that was built for somebody else.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://ambrain.org/blog/own-trading-platform-where-to-start/" rel="noopener noreferrer"&gt;ambrain.org&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>fintech</category>
      <category>trading</category>
      <category>startup</category>
      <category>architecture</category>
    </item>
    <item>
      <title>L2 Market Data Under Bursts: Sequence Gaps, Recovery, and Fan-Out to Hundreds of Sessions</title>
      <dc:creator>amBrain</dc:creator>
      <pubDate>Mon, 14 Sep 2026 07:14:36 +0000</pubDate>
      <link>https://dev.to/ambrainorg/l2-market-data-under-bursts-sequence-gaps-recovery-and-fan-out-to-hundreds-of-sessions-20ni</link>
      <guid>https://dev.to/ambrainorg/l2-market-data-under-bursts-sequence-gaps-recovery-and-fan-out-to-hundreds-of-sessions-20ni</guid>
      <description>&lt;p&gt;A handler that drops and reorders L2 updates under bursts is usually two separate failures behind one symptom. One lives on the ingest side, where the feed sequence is tracked and a gap has to be detected rather than absorbed. The other lives on the distribution side, where one normalised book is fanned out to many sessions and a single slow reader changes what the others receive.&lt;/p&gt;

&lt;p&gt;They have different fixes, and applying the wrong one moves the symptom instead of removing it. What follows is the shape of the pipeline when the sequence has to survive a burst: what a gap actually is, how recovery joins a snapshot to a live stream, how fan-out decides ordering once, and where conflation is honest.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The short answer is structural. Order is decided in exactly one place, upstream of every session: a single writer per instrument folds the numbered feed into a book, and sessions receive views derived from that fold - they never reorder anything for themselves. What amBrain can substantiate publicly: a mini-exchange we built runs in production on MOEX colocation, we built the Spectre Trade trading terminal, and the market data latency we publish is measured - under 5 ms on the paths we build. That figure describes our paths, not a benchmark of the design below.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A sequence number promises order, not delivery
&lt;/h2&gt;

&lt;p&gt;Feeds number their updates, and that number is the only ordering authority you have. Arrival time is not one: multicast paths reorder, several channels carry one instrument, receive queues are spread across cores, and a burst stretches all of it. A handler that orders by arrival is correct only while the network is calm - the condition nobody was worried about.&lt;/p&gt;

&lt;p&gt;Six properties of the feed have to be known before recovery logic is written. Each changes what a gap means.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The unit the sequence covers - channel, instrument, or book. A per-channel number does not tell you which instrument lost an update, and a per-instrument number does not tell you that a channel stopped&lt;/li&gt;
&lt;li&gt;The increment rule: strictly consecutive within the unit, or increasing with permitted holes. Both exist, and reading the second as the first produces recoveries that were never needed&lt;/li&gt;
&lt;li&gt;Whether numbers restart at a session boundary and what marks it - a restart read as a gap sends every instrument into recovery at the same moment&lt;/li&gt;
&lt;li&gt;Whether heartbeats carry the current sequence. Without them, a dead connection and a quiet instrument look identical&lt;/li&gt;
&lt;li&gt;Whether retransmission exists and over what window. If it does not, snapshot recovery is the only path back, and it has to be cheap enough to use often&lt;/li&gt;
&lt;li&gt;Which sequence number a snapshot is aligned with. Without it, a snapshot cannot be joined to a live stream at all&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where a property is genuinely unknown, measure it rather than encode a guess. Each of the six becomes a branch in the recovery path, and a wrong assumption there is found later as a book that quietly disagrees with the venue.&lt;/p&gt;

&lt;h2&gt;
  
  
  A gap and a reorder look identical for a few milliseconds
&lt;/h2&gt;

&lt;p&gt;Both start the same way: the next update does not carry the number you expected. The difference is time, so the classification is not made on arrival but when a bounded wait expires.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Out of order: you expected N, received N+2, and N+1 arrives while the wait is still open. Nothing is missing, and the only cost is the wait&lt;/li&gt;
&lt;li&gt;Duplicate or retransmission: a number at or below the last applied one. Dropped without touching the book, and counted, because a rising duplicate rate says something about the path&lt;/li&gt;
&lt;li&gt;Gap: the wait expired and N+1 never arrived. The book cannot advance past the hole, and this instrument goes to recovery&lt;/li&gt;
&lt;li&gt;Stale: the right number, too late to be useful. The bytes arrived, and downstream it is a loss&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;One rule keeps the corruption from becoming silent: an update is applied only when its sequence is exactly the expected one. Everything else goes to the wait buffer or to recovery. A book that accepts an out-of-order delta keeps serving prices and looks healthy - the disagreement with the venue is found later, by a client, on a fill that did not make sense.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The wait is a bounded structure, not a queue that grows. It holds updates ahead of the expected number, keyed by sequence, so releasing them is a lookup rather than a sort.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Release is a loop: apply the expected number, then apply what is already buffered while the numbers stay consecutive&lt;/li&gt;
&lt;li&gt;The deadline is expressed in time, not only in a count of pending updates - a burst fills a count-based window far sooner than the design intended&lt;/li&gt;
&lt;li&gt;Whatever the buffer waits, every consumer waits. Size the deadline from the reordering measured on your own path, not from a number that felt safe&lt;/li&gt;
&lt;li&gt;Buffer overflow is itself a gap declaration: the wait is bounded in memory as well as in time&lt;/li&gt;
&lt;li&gt;The wait is per instrument or per channel, never global. One quiet instrument must not hold back everything around it&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Recovery is a snapshot joined to a stream you were already buffering
&lt;/h2&gt;

&lt;p&gt;The join is the part that goes wrong. A snapshot is a book as of some sequence number, and it is stale the moment it is produced; what makes it usable is the incremental stream buffered while it was being fetched.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Buffer the incremental stream before the snapshot is requested. A snapshot without a live stream behind it is already behind the market when it lands&lt;/li&gt;
&lt;li&gt;Read the sequence number the snapshot is consistent with. If the feed publishes none, the feed is snapshot-only in practice, and the design has to say so out loud&lt;/li&gt;
&lt;li&gt;Discard buffered updates at or below the snapshot sequence, then apply the rest in order. If the first of them is not the update immediately after the snapshot, the join failed and recovery restarts&lt;/li&gt;
&lt;li&gt;If the buffer fills before the snapshot arrives, restart the recovery instead of applying part of it - a partially applied recovery is indistinguishable from a healthy book&lt;/li&gt;
&lt;li&gt;Publish the instrument as degraded while it recovers, as an explicit state on the stream. A book with a hole in it, served as current, is worse than no book&lt;/li&gt;
&lt;li&gt;Verify after the join: the checksum the feed publishes, if it publishes one, or agreement between your folded book and the next snapshot&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Recovery is a normal event, not an incident, and its cost belongs in the capacity plan: how long a snapshot takes to fetch, how much stream is buffered meanwhile, and how many instruments can recover at once before the snapshot service becomes the bottleneck.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fan-out: normalise once, encode once, send many
&lt;/h2&gt;

&lt;p&gt;Hundreds of terminal sessions want the same book. The mistake that multiplies under a burst is doing per-session work that is not per-session in nature: rebuilding a book for each subscriber, or serialising the same update once per socket.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One writer per instrument shard owns the book. Readers never mutate it, which removes both the lock and the question of whose version is authoritative&lt;/li&gt;
&lt;li&gt;The writer publishes versioned updates into a ring buffer that readers follow at their own pace, so a reader falling behind slows nobody down&lt;/li&gt;
&lt;li&gt;Each update is encoded once per wire format and shared across sessions by reference. Only framing and flow control are per session&lt;/li&gt;
&lt;li&gt;Every session carries its own outbound sequence number, so a client can detect its own losses without knowing anything about the upstream feed&lt;/li&gt;
&lt;li&gt;Ordering is guaranteed per instrument, because that is the guarantee clients depend on. Ordering across instruments is either promised explicitly and implemented, or not promised at all&lt;/li&gt;
&lt;li&gt;Beyond one process, fan-out becomes a relay tier: each relay takes one subscription upstream and serves a share of sessions, so the work of the writer stays constant&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;Fan-out cost is decided by how many times an update is transformed, not by how many sockets receive it. Encoding once and passing a reference scales with sessions; rebuilding a book per session does not.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  A slow consumer is a policy you choose, not an accident that happens
&lt;/h2&gt;

&lt;p&gt;Somewhere there is a session on a bad network, or a terminal whose render loop stalled, and its outbound buffer fills. There are four possible behaviours, and two of them are chosen only by accident.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Block the writer until the slow session drains: never. It turns one bad connection into a latency event for everyone on the shard&lt;/li&gt;
&lt;li&gt;Grow the queue without a limit: a slow consumer becomes memory exhaustion, and then an outage unrelated to the original session&lt;/li&gt;
&lt;li&gt;Bounded queue with conflation: correct for book state, where a client wants the current picture rather than every intermediate step&lt;/li&gt;
&lt;li&gt;Bounded queue with disconnect at a high watermark: correct for streams that cannot be conflated, where dropping an item drops meaning&lt;/li&gt;
&lt;li&gt;Whatever the policy, the queue is per session, and lag is measured continuously - queue depth, and the distance between the sequence published and the sequence written to the socket&lt;/li&gt;
&lt;li&gt;A disconnect states its reason. An unexplained close is retried in a loop; an explained one is followed by a resubscribe&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Back pressure is where the two sides meet. If the outbound path can push back on the book writer, a slow terminal eventually delays the folding of the feed, and gap detection starts firing for reasons that have nothing to do with the venue. A bounded ring between the two stops that chain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conflation is honest for state and wrong for events
&lt;/h2&gt;

&lt;p&gt;A book is state: the client wants the current levels, and a value already replaced carries no meaning of its own. A trade tape is a log of events, where each item is a fact that happened and cannot be summarised away.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Conflate: price level updates, top of book, aggregated depth, and derived statistics such as last price or session volume&lt;/li&gt;
&lt;li&gt;Do not conflate: trades and prints, order and execution reports, auction and phase changes, and anything a client aggregates over time - a tape built from a conflated stream is a wrong number held confidently&lt;/li&gt;
&lt;li&gt;Conflate per key, not per stream. Keeping the latest update for each price level preserves the book; keeping the latest update overall throws away every level that did not change last&lt;/li&gt;
&lt;li&gt;A conflated update carries the sequence number of the state it represents, so a client knows the point it corresponds to&lt;/li&gt;
&lt;li&gt;The conflation interval is part of the latency you report. A stream conflated on an interval is not described by the latency measured on the unconflated one&lt;/li&gt;
&lt;li&gt;A client that needs every intermediate state - a backtest, a compliance record - takes the unconflated stream and pays in bandwidth&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;Conflation is a change of shape, not a compression setting. Once a stream is conflated, a client cannot reconstruct what happened between two updates, and it must not be told the stream is complete. Publishing both - a conflated book stream and an unconflated event stream - is what keeps both kinds of client correct.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Reconnect is a resynchronisation, and they all arrive at once
&lt;/h2&gt;

&lt;p&gt;When a session comes back, the book it holds is worthless unless the server can prove continuity. The default is a fresh snapshot per subscription, with its sequence number, applied to a client that dropped its local state first.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Resume from a sequence number is offered only where a bounded replay buffer exists. When the requested number has aged out, the server says so and falls back to a snapshot rather than sending a stream with a hole&lt;/li&gt;
&lt;li&gt;Session state across a reconnect is an explicit decision: either the server keeps subscriptions for a bounded time under a session token, or the client restates them on connect. Both work; an implicit mixture does not&lt;/li&gt;
&lt;li&gt;Duplicate delivery after a resume is expected, and the client discards by sequence. At-least-once plus sequence numbering is easier to implement correctly than exactly-once&lt;/li&gt;
&lt;li&gt;Reconnects arrive together, because whatever disconnected one session usually disconnected many. Jittered backoff on the client and admission control on the server keep the recovery from becoming the second outage&lt;/li&gt;
&lt;li&gt;Snapshots for that crowd come from a per-instrument cache refreshed on a cadence, so the writer serialises a snapshot on a schedule rather than once per reconnecting session&lt;/li&gt;
&lt;li&gt;The client-side book is rebuilt, never patched. A terminal that keeps its old levels and applies new deltas on top carries the pre-disconnect error into a book that now looks fresh&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The failure worth designing for is not a single reconnect. It is a network event that returns hundreds of sessions in the same second, each asking for a snapshot of every instrument it was watching, while the ingest side recovers from the gap that same event produced.&lt;/p&gt;

&lt;p&gt;What amBrain can substantiate publicly: we build low latency trading platforms, matching engines and real-time bidding systems in Rust from Yerevan, Armenia, and the market data latency we publish - under 5 ms - is measured on the paths we build. If your handler is losing sequence under bursts, the conversation worth having is the one that separates the ingest side from the distribution side before either is rewritten.&lt;/p&gt;

</description>
      <category>fintech</category>
      <category>distributedsystems</category>
      <category>architecture</category>
      <category>performance</category>
    </item>
    <item>
      <title>Ad Measurement Loses Events at Peak: Seams, Duplicate Keys, and the Bidder Join</title>
      <dc:creator>amBrain</dc:creator>
      <pubDate>Thu, 10 Sep 2026 16:59:01 +0000</pubDate>
      <link>https://dev.to/ambrainorg/ad-measurement-loses-events-at-peak-seams-duplicate-keys-and-the-bidder-join-cdg</link>
      <guid>https://dev.to/ambrainorg/ad-measurement-loses-events-at-peak-seams-duplicate-keys-and-the-bidder-join-cdg</guid>
      <description>&lt;p&gt;An event pipeline that loses events at peak traffic and never reconciles with the bidder logs can be two failures behind one symptom. One is transport: events created and never landed, at a seam nobody counted. The other is definitional: events that landed and were counted under a different rule than the auction record.&lt;/p&gt;

&lt;p&gt;The usual framing - the pipeline drops events, so rebuild the pipeline - fixes at most one of them. What follows separates the two and says what a reconciliation produces, which is not equality. The defaults quoted below are Kafka on the transport side and ClickHouse on the storage side.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The short answer is structural: one identifier per auction carried end to end, a counter on both sides of every hop, and a reconciliation over a window that has already closed. What amBrain can substantiate publicly is the AdTech work: DSP development, real-time bidding platforms, and ad exchange engineering. In any RTB stack that is the side the bid, win and impression records come from. The pipeline below is described from the mechanics of the problem, not from a case of ours.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  One query tells a lost event from a late one
&lt;/h2&gt;

&lt;p&gt;The first failure is loss: the event was created and never arrived, at a specific hop, for a specific reason. A beacon that never left the page, an edge restarted mid-deploy, a producer buffer that filled, a consumer that committed its offset before processing.&lt;/p&gt;

&lt;p&gt;The second is not a failure at all. The measurement side counts a client-initiated event; the bidder records a server-side auction outcome. One is an award, the other an observation of what happened to it. The MRC measurement guidelines treat pre-fetch, pre-render and auto-refresh as separate things to detect and disclose: a counting rule, not a transport fault. Telling the two apart costs one query and some patience.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Re-run the same event-time window one hour, six hours and a full day after the time it covers&lt;/li&gt;
&lt;li&gt;A deficit that shrinks with each run means the events were late, not lost, and the transport is fine&lt;/li&gt;
&lt;li&gt;Count distinct deduplication keys, not rows: at-least-once transport guarantees redeliveries, and a rows-based curve hides an overcount&lt;/li&gt;
&lt;li&gt;A deficit that stays flat means the events are gone, and the question is which seam&lt;/li&gt;
&lt;li&gt;The test needs an event-time timestamp, an identifier that survives retries, and retention long enough to re-run the window&lt;/li&gt;
&lt;li&gt;Publish the settling curve as a chart, per event type, next to the number it explains&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Until that curve exists, both sides of the argument are opinions. Afterwards its shape decides which half of this article applies, and the halves are not exclusive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Events go missing at named seams, and an uncounted seam cannot be blamed
&lt;/h2&gt;

&lt;p&gt;There are seven places where an ad event is created and then quietly stops existing, plus one setting that looks like a guarantee and is not.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Client collection: the beacon fires but the document unloads first, or the creative was cached, pre-fetched or auto-refreshed and counts what the auction record does not&lt;/li&gt;
&lt;li&gt;Edge ingest: connection limits, keep-alive exhaustion, restarts during deploys, and the dangerous variant - success returned before the event is durable&lt;/li&gt;
&lt;li&gt;Producer buffer: the client blocks for a bounded time, then raises, and code that catches the error and counts nothing is where data dies&lt;/li&gt;
&lt;li&gt;Broker durability: with no acknowledgement required nothing guarantees the record arrived; with the leader alone it is lost if that leader fails before followers replicate&lt;/li&gt;
&lt;li&gt;Acknowledgement from all replicas is not durability: it waits for the current in-sync set, whose minimum size defaults to one, so a peak that pushes followers behind commits on the leader alone&lt;/li&gt;
&lt;li&gt;Consumer: committing the offset before processing is at-most-once, and it is a default rather than a decision - the client commits on a timer unless that was switched off&lt;/li&gt;
&lt;li&gt;Retention overrun: a consumer that falls behind past the retention window finds its next offset deleted, and the default reset policy jumps it to the head of the log&lt;/li&gt;
&lt;li&gt;Load into the columnar store: fire-and-forget inserts acknowledge once buffered, and dependent materialised views deduplicate through a separate setting - where raw table and report part ways&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rule is a reporting rule, not an engineering one: a loss is attributed to a named seam, or not attributed at all. Two alarms keep the quiet seams visible - consumer lag measured in time against retention, and a reset policy that fails rather than jumps.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deduplication needs a key that exists before the first retry
&lt;/h2&gt;

&lt;p&gt;A deduplication key is not a convenient primary key chosen at the destination. It is assigned upstream of every retry, at auction time or when the event is created, and the receiver never invents one. Arrival timestamps stay out: a retry carries a new arrival time and a new key.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The key has to be identical across every attempt, which decides its parts: the exchange or seat, the auction identifier, the impression identifier and the event type. Partition the topic on that key, so retries land together and per-key ordering survives them. That instruction is for the transport only: the same word applied to the columnar store gives one partition per event, and the insert dies on the per-block limit. Partition storage by time, order by the key.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;ul&gt;
&lt;li&gt;The idempotent producer removes duplicates from producer retries inside one session, and Kafka enables it by default since 3.0 alongside acknowledgement from all replicas&lt;/li&gt;
&lt;li&gt;That default is conditional: a conflicting setting from an older configuration disables idempotence silently, so ask the running process what it has&lt;/li&gt;
&lt;li&gt;It cannot see an application-level duplicate: a crashed process that re-sent, a beacon fired twice, an operator re-running an ingest job&lt;/li&gt;
&lt;li&gt;ClickHouse insert deduplication hashes the block contents, so a consumer that re-batches after a rebalance sends the same rows in a new shape and the hash misses&lt;/li&gt;
&lt;li&gt;The window is bounded in blocks and in time, and on non-replicated tables it defaults to zero, meaning off; an insert token removes that dependence&lt;/li&gt;
&lt;li&gt;Exactly-once inside the log covers consume-transform-produce, and the hop into an analytical database sits outside that boundary whatever the transport promises&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the working shape is at-least-once transport with idempotent keys. Deduplication at write time keeps the storage bill sane; deduplication at read time is what makes the number correct. Merges never combine parts from different partitions, so a duplicate landing in the next partition is resolved only when a query asks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Back pressure decides whether a loss is a number or a rumour
&lt;/h2&gt;

&lt;p&gt;Under overload a system has three options: slow the producer down, shed with a counter, or lose quietly. Only the third is unacceptable, and it is the default behaviour of code that was never asked the question. Loss at peak is a queue that grew until memory ran out, or an acknowledgement issued before durability.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Bounded queues at every hop, with explicit rejection instead of growth. An unbounded queue relocates the loss into memory pressure and a restart&lt;/li&gt;
&lt;li&gt;Producer blocking time and buffer size are capacity decisions: size them from the peak you measured, and alarm on time spent blocked&lt;/li&gt;
&lt;li&gt;Shed by class rather than at random: impression and billable events survive, diagnostics go first, and every shed event increments a labelled counter&lt;/li&gt;
&lt;li&gt;Consumer lag is back pressure made visible. Alarm on the age of the oldest unprocessed event and on how fast the lag changes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A drop with a labelled counter is a known quantity that can be reconciled later. A drop without one is not lost data, it is a lost number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Late arrival is structural, and half of the mismatch is a calendar
&lt;/h2&gt;

&lt;p&gt;The OpenRTB implementation guidance states it directly: the sequence from ad request through the auction to rendering and billing is fundamentally not transactional. Too many parties sit between the two counts.&lt;/p&gt;

&lt;p&gt;Delay is expected rather than exceptional. The bid request can carry an impression expiry and the bid the delay the bidder tolerates, and the same guidance gives rules of thumb running from the order of a minute for web to far longer for cached in-app formats and stitched video.&lt;/p&gt;

&lt;p&gt;Neither field is a contract. The guidance says plainly that a billing notice arriving later than the expiry a bidder declared may still be billable - a policy discussion between bidder and exchange rather than something the protocol imposes.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Three timestamps exist per event and exactly one drives the window: the device clock, untrusted; the edge receipt time, late; the auction time, authoritative&lt;/li&gt;
&lt;li&gt;A fourth exists where the exchange supplies it, the macro carrying the moment the impression was fulfilled; where it is absent the specification assumes the notice followed by seconds&lt;/li&gt;
&lt;li&gt;A watermark declares that event time has reached a point and no earlier elements are expected, so allowed lateness is a parameter you choose&lt;/li&gt;
&lt;li&gt;Publish the late window per event type with the restatement policy: numbers move while it is open, then freeze, and movements are logged&lt;/li&gt;
&lt;li&gt;Count the late event and mark it late, because dropping it is dropping spend you were billed for&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;An event that arrives after the window is not a defect in the pipeline, it is a property of the medium. The only real choice is whether the number moves in public while the window is open, or moves in private afterwards.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Reconciliation joins on the identifiers the protocol already carries
&lt;/h2&gt;

&lt;p&gt;The identifiers the exchange substitutes into notice and tracking URLs are the auction identifier from the bid request, the impression identifier, and, where the bidder minted one, the bid identifier. None of the three is a key on its own.&lt;/p&gt;

&lt;p&gt;The specification calls the auction identifier exchange unique, not globally unique: two exchanges can hand you the same string on the same day. The impression identifier is unique only inside its own bid request, often the literal 1. The bid identifier is optional.&lt;/p&gt;

&lt;p&gt;The key that holds is the composite: the exchange or seat you transacted through, plus the auction identifier, plus the impression identifier. Mint it on the bidder side at auction time, and treat anything shorter as a prefix rather than a key.&lt;/p&gt;

&lt;p&gt;If those macros are absent from the beacons, event-level reconciliation is impossible, and what remains is matching on time, placement and creative. What you can build instead is a bridge of six counts, each naming its reason for differing from the step above.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Auction wins, from the bidder log - the only count that is entirely yours&lt;/li&gt;
&lt;li&gt;Win notices received by the exchange - the difference is notice loss and timeouts, and by the specification a win notice does not necessarily imply delivery&lt;/li&gt;
&lt;li&gt;Beacons received at your edge - the difference is client collection and every seam above it&lt;/li&gt;
&lt;li&gt;Events after deduplication - the difference is retries, and it should be stable week to week&lt;/li&gt;
&lt;li&gt;Events after invalid-traffic filtration - the difference is a filtration rate you publish rather than discover&lt;/li&gt;
&lt;li&gt;Billable events - the difference is the billing rule, and the notice belongs server-side, where the exchange books revenue&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A stable ratio between steps is the goal and an unexplained move is the alarm. Inside one system the hop ratios belong at one, and any excursion is the signal. Across the auction-to-measurement boundary, one is the suspicious reading.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The same counters, read as ratios, turn the question into arithmetic: accepted over sent, produced over accepted, consumed over produced, inserted over consumed. Four ratios on one chart say where the events went before anyone opens a log. Add a producer-side sequence number per source and partition, and a hole becomes evidence rather than a suspicion.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What a rebuilt pipeline does not give you
&lt;/h2&gt;

&lt;p&gt;One question the bridge does not answer and a contract does: which of these numbers you pay on. The seller books revenue on its own billable event, the buyer paces on its own, and the guidance treats a persistent gap as a support conversation between the parties.&lt;/p&gt;

&lt;p&gt;Decide in advance which count is the system of record for spend, and at what gap a report note becomes a ticket with the exchange. The work above buys attribution of loss, honest duplicates and a reconciliation explained line by line. It does not buy the following.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It does not make the two counts equal: both sides count different events on purpose, and the difference is explained, never removed&lt;/li&gt;
&lt;li&gt;It does not recover events dropped before the instrumentation existed - the settling curve starts on the day the counters do&lt;/li&gt;
&lt;li&gt;It does not remove restatement: yesterday moves while the late window is open, and a business that cannot tolerate it needs a later close&lt;/li&gt;
&lt;li&gt;It does not survive missing macros: without the auction identifiers in the beacons, no storage design produces an event-level join&lt;/li&gt;
&lt;li&gt;It does not make sampled data joinable afterwards, because sampling decides which questions stay answerable before the row is written&lt;/li&gt;
&lt;li&gt;It does not replace the disclosure list an MRC-style audit expects: capture point, logging frequency, latency estimates, rules for inconsistencies&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A pipeline that can answer those disclosure questions has an integrity story. One that cannot has an opinion, and an opinion is what gets argued about at the end of a quarter.&lt;/p&gt;

&lt;p&gt;The failure worth designing against is not the missing hour that starts an investigation. It is the quiet version: a seam that sheds without a counter, an insert acknowledged before it was durable, and a reconciliation over a window still open.&lt;/p&gt;

&lt;p&gt;What amBrain can substantiate publicly: amBrain is a Yerevan, Armenia software engineering company building low latency trading platforms, matching engines, and real-time bidding systems in Rust. amBrain has been building software since 2019. Rebuilding a measurement pipeline is not work described here. If the numbers stop reconciling on the bidder and exchange side - the bid, win and impression records themselves - that is the conversation worth having, and it starts with the settling curve rather than with a rebuild.&lt;/p&gt;

</description>
      <category>kafka</category>
      <category>clickhouse</category>
      <category>dataengineering</category>
      <category>adtech</category>
    </item>
    <item>
      <title>Go GC Pauses in an RTB Bidder: Mark Assist, Deadlines, and the Rust Decision</title>
      <dc:creator>amBrain</dc:creator>
      <pubDate>Thu, 10 Sep 2026 16:52:57 +0000</pubDate>
      <link>https://dev.to/ambrainorg/go-gc-pauses-in-an-rtb-bidder-mark-assist-deadlines-and-the-rust-decision-309e</link>
      <guid>https://dev.to/ambrainorg/go-gc-pauses-in-an-rtb-bidder-mark-assist-deadlines-and-the-rust-decision-309e</guid>
      <description>&lt;p&gt;A bidder that loses auctions on timeout while its average looks healthy is being described by the wrong number. The deadline belongs to the exchange and covers the network in both directions. A collector is a property of the process, so its pause is charged to every connection at once, and a spike on one connection alone needs another explanation.&lt;/p&gt;

&lt;p&gt;Rewrite in Rust or tune the runtime is a choice between two answers to a question nobody has asked: which part of the deadline is being spent, and on what. What follows separates the budget from the collector, the collector from the scheduler, and the rewrite from the component that deserves it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The short answer is structural. The exchange sets the deadline and it includes the network, so the first repair splits the p99 of one connection into handler time, time waiting for a processor, and time on the wire. What amBrain can substantiate publicly: we built RTBBidder, a demand-side platform delivered from scratch, where each bid decision evaluates dozens of targeting conditions per impression. The latency figures we publish as measured come from trading paths, not from an ad bidder, and no number below is measured on a bidder of ours.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The deadline belongs to the exchange, and it covers the round trip
&lt;/h2&gt;

&lt;p&gt;The OpenRTB specification defines tmax as the maximum time in milliseconds the exchange allows for bids to be received, including Internet latency, and says the value supersedes any prior guidance. The budget is a round trip that arrives inside each request.&lt;/p&gt;

&lt;p&gt;The threshold you are judged against is not the one on your dashboard. Google's Authorized Buyers documentation requires that 85 percent of responses arrive inside the deadline as measured at the trading location, and throttles bidders that miss it. Between its clock and yours sits everything that is not computation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The round trip to the exchange and back, which is geography and peering before it is engineering&lt;/li&gt;
&lt;li&gt;Connection setup when keepalive lapses, because a fresh TLS handshake inside an auction budget is a lost auction&lt;/li&gt;
&lt;li&gt;Time in the accept queue before your handler sees the request, which grows exactly when you are busiest&lt;/li&gt;
&lt;li&gt;Deserialising, at a cost set by how much of the request you turn into objects, not by its size&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the internal deadline sits below tmax by the amount your own histogram says an answer costs on that connection, re-derived per exchange rather than set once for the fleet. A deadline is not a capacity plan: work cancelled at the deadline has already spent its CPU.&lt;/p&gt;

&lt;p&gt;Under overload that means paying full price for responses nobody counts, so the missing repair is admission control: read tmax, compare it against the queue delay you measure, and answer no-bid when the arithmetic does not close. A fast no-bid counts toward the 85 percent; a late bid does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  One collector serves every connection, so a single connection is a different fault
&lt;/h2&gt;

&lt;p&gt;A garbage collector is a property of the process, so a cycle triggered on any connection charges every connection. The one that misses first is the one with the tightest tmax and the heaviest request. Rule out what produces the same picture without a collector:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Too few connections from one exchange, since HTTP/1.1 carries one request at a time: QPS per connection times handler time near one puts the queue on the connection&lt;/li&gt;
&lt;li&gt;A single HTTP/2 connection, where one lost packet stalls every stream sharing it - the head-of-line blocking RFC 9114 cites in 2022 as the reason HTTP/3 exists&lt;/li&gt;
&lt;li&gt;Keepalive lapsing, a default more often than a fault: http.Server.IdleTimeout falls back to ReadTimeout when it is zero, while Google asks for a 2.5-minute idle timeout and nginx closes at 75 seconds&lt;/li&gt;
&lt;li&gt;A synchronous feature lookup only some exchanges trigger, where the tail belongs to a remote store, not to you&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of the four is repaired by a collector setting, and the split has instruments. A CPU profile separates the collector bills by symbol, runtime.gcAssistAlloc for handler charges and runtime.gcBgMarkWorker for background marking. Stop-the-world time is /sched/pauses/total/gc:seconds, runnable-wait is /sched/latencies:seconds, and the accept queue is read outside the process through ListenOverflows.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;One distinction decides which Go repair you need. A stop-the-world pause is charged to every goroutine at once, so it appears as a flat spike on all connections. A mark assist is charged to the goroutine that allocated, so it lands on the requests that allocated most. Two things lower an assist: fewer bytes per bid request, or a longer cycle in which background workers cover more of the marking. Only the first survives a change in traffic mix.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Mark assist is the bill your bid request pays
&lt;/h2&gt;

&lt;p&gt;The Go collector is concurrent, and the official guide is explicit that pause length does not scale with heap size, so stop-the-world transitions are brief. Assists are the source that matters: goroutines assist the collector when allocation is fast, because background marking gets a fixed quarter of the processors and the shortfall is charged to the allocator.&lt;/p&gt;

&lt;p&gt;Rate makes this a threshold rather than a slope. Allocation rate is QPS multiplied by bytes per bid request against a fixed background share, so code that never assists at a fifth of your traffic can assist on nearly every request at 100K QPS.&lt;/p&gt;

&lt;p&gt;Marking cost is proportional to the live pointer graph, not to the garbage, and a bidder holds the wrong shape for it: campaign indexes, audience segments, frequency caches. Discord published the same finding in 2020, with a collector scanning an entire LRU cache to decide whether the memory was free. Five mechanisms hide under one phrase:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Stop-the-world transitions: a flat spike on every connection in the same instant, rarely long enough to lose an auction on its own&lt;/li&gt;
&lt;li&gt;Mark assist: no pause in the trace at all, only handler time that grew - read it from the assist share of GC CPU&lt;/li&gt;
&lt;li&gt;Marking cost: GC CPU that rises when the live heap grows even though allocation did not, which flat arrays move and less garbage does not&lt;/li&gt;
&lt;li&gt;Scheduler contention: a request that sits runnable without executing, which the scheduler latency metric shows and the pause metric will not&lt;/li&gt;
&lt;li&gt;The forced cycle: a spike on a roughly two-minute period on a quiet instance, pointing at the collection floor rather than your traffic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Read that as five separate bills. Exactly one is settled by a collector setting, and none by changing the language before the split is measured.&lt;/p&gt;

&lt;h2&gt;
  
  
  A closed loop deletes the evidence you needed
&lt;/h2&gt;

&lt;p&gt;Gil Tene named the failure coordinated omission: the measuring system coordinates with the system under test in a way that avoids measuring outliers, because a closed loop waits for a reply and stops sending during a stall. ScyllaDB published a comparison in 2021 where one workload reported a p99 of 249 microseconds closed-loop and 665 ms under open load with correction, off by about 2,700 times.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Open-loop load at a fixed rate, with latency counted from the intended send time rather than from the moment the request left&lt;/li&gt;
&lt;li&gt;Correction in the style of HdrHistogram whenever the generator does not queue the requests it failed to send&lt;/li&gt;
&lt;li&gt;p99 and p99.9 rather than an average, with your own fan-out counted: Dean and Barroso showed in 2013 that touching 100 servers with a one-second p99 leaves 63 percent of requests slow&lt;/li&gt;
&lt;li&gt;A traffic profile copied from the exchange that hurts, run longer than the forced collection interval and under the same cgroup quota&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;A load generator that waits for a reply stops sending during exactly the stall it was built to find, and then averages the silence into the result. The percentile it prints afterwards describes the generator, not your bidder.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Allocate less first, then turn three knobs
&lt;/h2&gt;

&lt;p&gt;Name the stopping number before you tune, because a rewrite decided by exhaustion is not a decision. Two figures, not one: heap allocations per request, which drives assist, and the live heap, which drives marking.&lt;/p&gt;

&lt;p&gt;The order is source before ceiling, not largest win first. Start where the assist is generated: escape analysis on the bid path, buffers reused instead of allocated, and a codec that reads the fields you need instead of materialising a fresh object graph. Keep its slices short-lived: a slice into the request buffer pins that whole buffer for the life of the bid.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;sync.Pool relieves pressure and promises nothing: an item may be removed at any time without notification, and a pooled object is still marked between its return and its eviction&lt;/li&gt;
&lt;li&gt;The live set is the bill: rebuild campaign and segment indexes into flat arrays off the hot path, so the marked graph stops growing with campaign count&lt;/li&gt;
&lt;li&gt;GOGC trades memory for collector CPU at a rate the guide states plainly: doubling it roughly halves GC CPU cost, and the assist share falls with it&lt;/li&gt;
&lt;li&gt;GOMEMLIMIT is that trade against a ceiling, soft by design, because a hard limit turns a heap spike into an indefinite stall&lt;/li&gt;
&lt;li&gt;GOMAXPROCS inside a container: Go 1.25 reads the cgroup CPU limit and explicitly not CPU requests, so a pod with requests and no limit keeps the old behaviour&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The approach has a ceiling: Uber reported in 2021 that tuning GOGC against the container memory limit recovered around 70,000 cores across its mission-critical services. That is a cost result, not a percentile.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Rust hot path gives you, and what you pay for it
&lt;/h2&gt;

&lt;p&gt;RTB House described a JVM bidding service in June 2025 where a split into microservices produced a high volume of small requests: the added latency had to stay inside 7 ms against an average request of about 2.5 ms, and the 98th and 99th percentiles broke under frequent G1 pauses. They moved to generational ZGC and paid in memory.&lt;/p&gt;

&lt;p&gt;Note what that repair required: a second collector to switch to. Go ships one and it is not pluggable, so the Go levers are allocation rate, the shape of the live set, and GOGC against GOMEMLIMIT. What a Rust hot path removes instead is specific: no assist, no background marking, no forced cycle. What does not go away is longer than most teams expect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The allocator, plus page faults and NUMA placement: the Rust standard library states that the default global allocator is unspecified, and malloc under many threads is a tail source&lt;/li&gt;
&lt;li&gt;Scheduling, because an async runtime with a worker pool reproduces the Go scheduler effects the moment a blocking task lands on a worker&lt;/li&gt;
&lt;li&gt;Memory accounting, since GOMEMLIMIT covers Go runtime memory only: a Rust allocator in the same process sits outside your ceiling and enforcement moves to the OOM killer&lt;/li&gt;
&lt;li&gt;People, meaning who is on call for the hot path at night and what two toolchains cost in one repository&lt;/li&gt;
&lt;li&gt;The boundary, if Go stays outside: Cockroach Labs measured a cgo call at 171 ns against 1.83 ns for a Go call in 2015, and the ratio is what survives&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the tail lives in parsing the request, in the connection to the exchange or in scheduler queueing, Rust returns none of those milliseconds, and rewriting the wrong component spends a quarter to keep the same timeout rate.&lt;/p&gt;

&lt;p&gt;So move the smallest piece that owns the allocations, not the service: the impression evaluation loop and its indexes, candidate selection, targeting, frequency and budget lookups, scoring. Price the boundary per crossing: once per bid request with a flat buffer, never once per targeting rule. Validate on separate instances fed a mirrored stream, never inside the process under test, where a shadow path doubles the two quantities you are measuring.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;An acceptance test that works on either path: run a short window with the collector off, under a memory ceiling you control, and record p99.9 inside the handler. Run it on one instance behind a fraction of traffic with an automatic revert, because with GOGC off a heap spike against the ceiling puts the runtime into back-to-back cycles, and the guide says that stall can be indefinite. Read the result only if the GC CPU limiter never engaged: once it does, the percentile describes the limiter.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The firm that fixes this asks for the timeout report first
&lt;/h2&gt;

&lt;p&gt;The second half of the question, who does this kind of work, has a test that does not need a vendor list. A firm that repairs GC-driven latency behaves differently from one that sells a rewrite:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Asks for the tmax distribution and the per-exchange timeout report before it asks for the repository&lt;/li&gt;
&lt;li&gt;States the split - handler, scheduler, wire - and which of the three it expects to own the milliseconds, before it proposes a language&lt;/li&gt;
&lt;li&gt;Brings its own open-loop generator and traffic profile, and refuses a closed-loop percentile as evidence&lt;/li&gt;
&lt;li&gt;Names the exit criterion in advance: which exchange, which percentile, what margin against its tmax, what memory per thousand requests&lt;/li&gt;
&lt;li&gt;Can staff the hot path afterwards, because a rewrite nobody is on call for at night is the second incident&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of them requires trust: each is a document you can ask for in the first conversation, and one that comes back vague says the diagnosis is being skipped.&lt;/p&gt;

&lt;p&gt;So the first question is not which language. It is which of the three sums - handler, scheduler or wire - owns the missing milliseconds on the connection that times out, and whether bytes allocated per bid request moves when you attack it.&lt;/p&gt;

&lt;p&gt;What amBrain can substantiate publicly: amBrain is a software development company specializing in trading platforms, matching engines, and real-time bidding systems, we have been building software since 2019, and in AdTech what we build is DSP development, real-time bidding platforms and ad exchange engineering. We work in three formats: full delivery, a dedicated team, or engineers embedded in yours. If you are weighing a rewrite against a tuning pass, the conversation worth having runs the split before it picks a language.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>backend</category>
      <category>performance</category>
      <category>rust</category>
    </item>
    <item>
      <title>Pre-Trade Risk Checks Inside the Order Path</title>
      <dc:creator>amBrain</dc:creator>
      <pubDate>Tue, 08 Sep 2026 11:37:00 +0000</pubDate>
      <link>https://dev.to/ambrainorg/pre-trade-risk-checks-inside-the-order-path-5850</link>
      <guid>https://dev.to/ambrainorg/pre-trade-risk-checks-inside-the-order-path-5850</guid>
      <description>&lt;p&gt;An order arrives at the gateway. Before it goes out to the venue, something has to decide whether the account is allowed to send it. That decision runs on every order, including the overwhelming majority that are perfectly fine, so its cost is paid by all normal traffic and not only by the rejects.&lt;/p&gt;

&lt;p&gt;That is the constraint to name first. A risk check placed in the order path is a tax on order entry. The engineering question is not how to make the check clever, it is how to make it small enough that traders do not feel it, and honest enough that it still refuses the orders it must refuse.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The short answer first. The checks that belong in the order path are position and exposure limits, margin, fat-finger bounds, instrument and account state, the kill switch, and duplicate guards. On the risk path amBrain builds for brokers and prop firms, the pre-trade decision on in-memory state completes in under 1 ms - a figure that covers the check itself, not the end-to-end journey of an order. The rest of this article is how that path is built and where it stops.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What actually belongs in the order path
&lt;/h2&gt;

&lt;p&gt;The hot path answers exactly one question: may this order be sent right now, given what we currently know about this account. Checks that answer that question stay in. Checks that answer a different question move out.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Position and exposure limits - the resulting position in the instrument, the group and the account, compared against configured bounds&lt;/li&gt;
&lt;li&gt;Margin or buying power - whether the account still has room for the order under the current margin model&lt;/li&gt;
&lt;li&gt;Fat-finger bounds - order size, notional and price distance from a reference, catching the typo before the venue does&lt;/li&gt;
&lt;li&gt;Instrument and account state - trading halted, account restricted, close-only, product not enabled for this account&lt;/li&gt;
&lt;li&gt;Kill switch state - a single flag that overrides everything above&lt;/li&gt;
&lt;li&gt;Duplicate and self-trade guards where the venue does not provide them&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else runs beside the path, on the same state, without holding the order. It informs the limits that the hot path enforces, but it does not sit between the trader and the venue.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Portfolio risk analytics - scenario runs, stress tests, correlated exposure across accounts&lt;/li&gt;
&lt;li&gt;Margin model re-rating when parameters change, and any recalculation that touches the whole book&lt;/li&gt;
&lt;li&gt;Surveillance and pattern detection, which needs history the hot path deliberately does not carry&lt;/li&gt;
&lt;li&gt;Reporting, reconciliation and anything that talks to a database or an external service&lt;/li&gt;
&lt;li&gt;Credit and counterparty review, which operates on a slower clock by nature&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;The dividing line is a question, not a category. In the path: may this order go out. Beside the path: what should the limits be. Anything that answers the second question and still blocks the order is a design mistake, however important the check is.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Where position state lives, and why the database is not on the path
&lt;/h2&gt;

&lt;p&gt;The state a pre-trade check needs - current positions, working orders, used and available margin, limit configuration - lives in the memory of the process that makes the decision. Not in a cache in front of a database, not behind a network call. In the process.&lt;/p&gt;

&lt;p&gt;The reason is not only speed, though a query is orders of magnitude more expensive than a lookup in a local array. The reason is correctness. A database holds the position as it was written. The risk check needs the position including orders sent a moment ago that have not been filled, acknowledged or persisted yet. If you read from storage, you check against a past that has already been overtaken by your own flow.&lt;/p&gt;

&lt;p&gt;Practically, that shapes the process the way any low-latency component gets shaped:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One writer per account. Accounts are sharded across risk instances so that an account's state is never contended, and no lock is taken on the order path&lt;/li&gt;
&lt;li&gt;Flat, pre-allocated structures - fixed-size arrays indexed by account and instrument id, resolved at session start, not hash lookups on strings built per order&lt;/li&gt;
&lt;li&gt;No allocation, no I/O and no logging that blocks on the decision path; the audit record is handed to another thread through a queue&lt;/li&gt;
&lt;li&gt;Configuration that changes without a restart is swapped in as a whole immutable snapshot, so the check never reads a half-updated limit&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;The database is where the position is recorded. It is not where the position is known.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Incremental recalculation, not a full pass
&lt;/h2&gt;

&lt;p&gt;A full recomputation of an account's exposure and margin walks every position and every working order. That cost grows with the size of the book, which means the risk check would get slower for exactly the clients who trade the most. So the hot path does not recompute. It applies a delta.&lt;/p&gt;

&lt;p&gt;The account carries running aggregates - net and gross exposure per instrument and per group, used margin, notional in flight. An incoming order produces a small change to those aggregates, the changed values are compared against the limits, and the order is accepted or rejected. The work is proportional to the order, not to the portfolio.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On send, the order's worst-case effect is reserved against the aggregates, so two orders in flight cannot both fit into the same remaining headroom&lt;/li&gt;
&lt;li&gt;On reject, cancel or expiry, the reservation is released; on fill, the reservation is replaced by the realised position change&lt;/li&gt;
&lt;li&gt;Partial fills adjust both sides of that in one step, which is where most bugs in this kind of engine actually live&lt;/li&gt;
&lt;li&gt;Netting and grouping rules are resolved when the instrument is loaded, not per order, so the delta is a handful of arithmetic operations&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Full recomputation still happens - on a schedule, when margin parameters change, and as a periodic self-check against the incremental result. It runs off the path, on a copy, and its result is either swapped in or raised as a discrepancy. Incremental state that silently drifts from the true state is worse than no check at all, so the comparison is not optional.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Measured on the risk path we build, the pre-trade check itself completes in &amp;lt;1 ms. That figure covers the decision on in-memory state, not the full journey of an order from the client to the venue and back.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The kill switch is a separate path
&lt;/h2&gt;

&lt;p&gt;A kill switch is used precisely when something is already wrong. That rules out building it on top of the machinery that may itself be the thing that is wrong. It is a separate path, with its own rules.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;It is a single atomic flag read at the top of the check, before position state, margin or instrument data is touched - so it works even when those are stale, missing or broken&lt;/li&gt;
&lt;li&gt;It is set from several independent triggers: an operator action, an automated condition, a loss of the market data or fill feed the risk state depends on&lt;/li&gt;
&lt;li&gt;It fails closed. If the risk process cannot establish that it has valid state, the gateway behaves as if the switch is engaged&lt;/li&gt;
&lt;li&gt;It has scopes - the whole firm, one desk, one account, one strategy - because a switch that can only stop everything gets used too late&lt;/li&gt;
&lt;li&gt;Engaging it is one action and one confirmation, not a configuration deploy; disengaging it is deliberate and always recorded&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Stopping new orders is the easy half. The harder half is what the switch does to orders already resting at the venue: pulling quotes and cancelling working orders has to be possible while the sending path is disabled. That cancel path deserves its own testing, because it is exercised on the worst day rather than on a normal one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Restart, and how state comes back
&lt;/h2&gt;

&lt;p&gt;In-memory state is a derived view of a durable record. That is what makes a restart survivable. Every event that changes risk state - an order accepted, a reservation released, a fill applied, a limit changed, the switch engaged - is appended to a journal on the local machine before it is acted upon downstream.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;On start, the process replays the journal to rebuild aggregates, then reconciles against the venue and clearing drop copy for positions and working orders&lt;/li&gt;
&lt;li&gt;Until reconciliation completes, the account is not open for trading. A risk engine that accepts orders while it is still figuring out the position is not a risk engine&lt;/li&gt;
&lt;li&gt;A mismatch between the replayed state and the venue's view stops that account and raises an alert; it is never resolved by quietly preferring one side&lt;/li&gt;
&lt;li&gt;A hot standby follows the same journal, so a failover restores a warm state instead of a cold replay, and the standby is verified by being promoted regularly rather than in theory&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Recovery time then depends on journal length and drop copy availability, not on the size of the book, and the failure mode of every unknown is the same: refuse to trade the account.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this design does not give you
&lt;/h2&gt;

&lt;p&gt;The honest limits are worth stating plainly, because they decide whether this architecture fits at all:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The check is only as correct as the fill feed. If drop copy or execution reports lag, exposure is understated, and the correct response is to degrade to conservative limits or engage the switch rather than to keep trading on stale state&lt;/li&gt;
&lt;li&gt;Portfolio margin models that are genuinely non-additive resist incremental evaluation. What works is a conservative incremental bound on the path plus a full model off the path; the price is that some orders are rejected which a full model would have allowed&lt;/li&gt;
&lt;li&gt;The switch protects against your own flow, not against the market. It cannot prevent a gap or a slippage on positions you already hold&lt;/li&gt;
&lt;li&gt;In-process state means the risk engine and the order gateway share a fate. That buys latency and costs you the ability to scale them independently&lt;/li&gt;
&lt;li&gt;Single-writer sharding by account makes cross-account limits harder, and firm-wide checks need a slower aggregation layer with its own staleness&lt;/li&gt;
&lt;li&gt;It is more operational work than a database-backed check: journals, reconciliation, standby promotion drills. If order entry is not latency-sensitive, this complexity is not worth buying&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;amBrain builds this kind of pre-trade risk path for brokers and prop firms, with the hot paths written in Rust. The team has worked on trading infrastructure from Yerevan, Armenia since 2019.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://ambrain.org/blog/pre-trade-risk-hot-path/" rel="noopener noreferrer"&gt;ambrain.org&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>rust</category>
      <category>architecture</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Designing a Matching Engine in Rust: Price-Time Priority Without GC Pauses</title>
      <dc:creator>amBrain</dc:creator>
      <pubDate>Tue, 08 Sep 2026 11:29:07 +0000</pubDate>
      <link>https://dev.to/ambrainorg/designing-a-matching-engine-in-rust-price-time-priority-without-gc-pausesr-2h56</link>
      <guid>https://dev.to/ambrainorg/designing-a-matching-engine-in-rust-price-time-priority-without-gc-pausesr-2h56</guid>
      <description>&lt;p&gt;A spot exchange matching engine has a narrow job. It takes an ordered stream of commands - new order, cancel, replace - applies them to an order book under price-time priority, and emits an ordered stream of events: trades, book updates, rejects, acknowledgements.&lt;/p&gt;

&lt;p&gt;Everything hard about it comes from three constraints stacked on top of that job: the result must be identical on every replay of the same input, the remainder of a partially filled order must keep its place in the queue, and the latency tail must not move when a burst arrives.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What amBrain can substantiate publicly: a mini-exchange we built runs in production on MOEX colocation, and the hot paths of our trading systems are written in Rust. The two figures we publish - market data delivered in under 5 ms and pre-trade risk checks completing in under 1 ms - are measured on those paths, not on the matching loop described below. The rest of this article is how we design the engine, not a benchmark of one.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The order book is a sorted index of price levels over FIFO queues
&lt;/h2&gt;

&lt;p&gt;The book is two sides, each a price-ordered collection of levels. A level is not a number - it is a queue of resting orders at that price, in arrival order. Matching touches the best level constantly and the deep levels rarely, so the structure is chosen for that access pattern rather than for elegance.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prices are integers in ticks, never floating point - a tick is the unit of the instrument, and comparison and arithmetic on integers are exact&lt;/li&gt;
&lt;li&gt;Each side keeps its levels in price order, with the best price reachable without a search - the top of book is read on every single command&lt;/li&gt;
&lt;li&gt;A level holds a FIFO queue of resting orders, plus the aggregated resting quantity, so the aggregate does not have to be recomputed by walking the queue&lt;/li&gt;
&lt;li&gt;Orders are held in a preallocated slab and referenced by index handles, and the queue is intrusive: the next and previous links live inside the order record itself&lt;/li&gt;
&lt;li&gt;A separate map from client order id to slab handle makes cancel and replace a direct lookup, so a cancel never scans the book&lt;/li&gt;
&lt;li&gt;Removal from a queue is by handle, not by search - a cancel of a deep resting order costs the same as a cancel at the top of book&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The consequence of intrusive queues and index handles is that a resting order never moves in memory while it lives. Its queue position is a property of its links, not of where it happens to sit, which is what makes partial fills cheap later on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Price-time priority is a loop over levels, then over a queue
&lt;/h2&gt;

&lt;p&gt;An incoming aggressive order walks the opposite side from the best price inward. At each level it walks the FIFO queue from the front. It stops when the level price is no longer acceptable to the incoming order or when the incoming quantity reaches zero.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Take the best opposite level; if its price does not cross the incoming limit price, stop&lt;/li&gt;
&lt;li&gt;Take the front order of that level queue - it is the oldest at that price, and time priority means it fills first&lt;/li&gt;
&lt;li&gt;The traded quantity is the smaller of the two remaining quantities; the trade price is the resting order price, because the resting order set the terms&lt;/li&gt;
&lt;li&gt;Decrement both remainders, emit the trade event, and decrement the level aggregate&lt;/li&gt;
&lt;li&gt;If the resting order remainder reaches zero, unlink it from the queue and return its slab slot; if the level becomes empty, remove the level&lt;/li&gt;
&lt;li&gt;Repeat until the incoming quantity is zero or no acceptable level remains&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;A partially filled resting order keeps its place. Its remainder stays at the front of its queue with its original arrival sequence, because a fill changes a quantity and nothing else. A partially filled aggressive order that is a plain limit becomes a resting order at the back of its own price level, with a new arrival sequence - it arrived now, not earlier.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Order type semantics are decisions taken at the boundary of this loop, not inside it. Immediate-or-cancel drops the remainder instead of resting it. Fill-or-kill runs a dry pass first and either executes whole or rejects. Post-only rejects if the order would cross on arrival. Keeping these outside the loop means the loop stays the only place where book state changes.&lt;/p&gt;

&lt;p&gt;Self-trade prevention, minimum quantities, and tick and lot validation belong before the loop as well. An order that reaches matching has already been proven well formed, so the loop has no error branches to slow it down or to disagree about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Determinism is what makes the tail predictable, and allocation is what breaks it
&lt;/h2&gt;

&lt;p&gt;Average latency is rarely the problem. The problem is the worst observation during a burst, which is when the engine matters most and when a stop-the-world pause is most likely to land. Under a managed runtime with a garbage collector, that pause is scheduled by the collector rather than by you, and it lands in the middle of the burst that produced the garbage.&lt;/p&gt;

&lt;p&gt;Manual allocation is a smaller version of the same problem. A general purpose allocator can walk free lists, take a lock, or ask the kernel for more memory, and the call that does so is the call that shows up in the tail. The fix is the same in either case: do not allocate on the hot path at all.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Order records, level records and event buffers come from arenas sized at startup - steady state allocation count on the matching path is zero&lt;/li&gt;
&lt;li&gt;Freed slots return to a free list inside the arena, so a busy instrument recycles the same memory all session&lt;/li&gt;
&lt;li&gt;Structures are fixed shape and fixed size, with capacity limits enforced as a rejection rather than as a growth event&lt;/li&gt;
&lt;li&gt;Outbound events are written into a preallocated ring buffer that another thread drains - the matching thread never blocks on a consumer&lt;/li&gt;
&lt;li&gt;The matching thread is a single writer over the book, so there is no lock on book state and no ordering ambiguity to resolve&lt;/li&gt;
&lt;li&gt;Inputs are sequenced before they reach the engine, and the sequence number, not the arrival time, decides order&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;An engine is deterministic when the same input sequence produces the same output sequence, byte for byte, on a different machine and a year later. Anything that reads wall clock time, thread scheduling or hash iteration order inside the matching path breaks that property.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Timestamps are therefore an input, not something the engine reads for itself. The sequencer stamps a command when it accepts it, and the matching loop treats the stamp as data. Randomness, if any is needed, comes from a seeded generator whose seed is part of the journal.&lt;/p&gt;

&lt;h2&gt;
  
  
  The journal is the engine state, and the book is a cache of it
&lt;/h2&gt;

&lt;p&gt;Recovery is not a feature bolted on after matching works. The engine writes an append-only journal of accepted commands in sequence order, and the in-memory book is nothing more than the result of folding that journal. Rebuilding after a crash means replaying it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A command is journaled and durable before it is matched, so an accepted order cannot be lost by a crash between acknowledgement and execution&lt;/li&gt;
&lt;li&gt;The journal is the input sequence, not the output - outputs are derived, and re-deriving them is exactly what replay does&lt;/li&gt;
&lt;li&gt;Periodic snapshots of the book carry the sequence number they were taken at, so recovery loads a snapshot and replays only the tail of the journal&lt;/li&gt;
&lt;li&gt;Snapshot and replay agreement is checked rather than assumed: replaying from the previous snapshot must reproduce the next one&lt;/li&gt;
&lt;li&gt;The event stream carries the same sequence numbers, so downstream consumers - risk, settlement, market data - can be resumed from a known point instead of resynchronised by hand&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The engine writes the journal, but durability is a property of the storage path and of how many machines have the record before the acknowledgement goes out. That is a replication and hardware decision, and it is where recovery time is actually won or lost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing is deterministic replay plus invariants that must always hold
&lt;/h2&gt;

&lt;p&gt;Determinism is what makes the engine testable. Because the same input gives the same output, a captured session is a regression test, and a failure found once can be reproduced exactly instead of chased.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Replay harness: feed a recorded command sequence, compare the emitted event stream against the stored one, and fail on the first divergence with the sequence number&lt;/li&gt;
&lt;li&gt;Invariant checks after every command in test builds - queues sorted by arrival, level aggregates equal to the sum of their queue, no crossed book, total quantity conserved across every trade&lt;/li&gt;
&lt;li&gt;Property based tests that generate random but well formed command sequences and assert the invariants rather than specific outcomes&lt;/li&gt;
&lt;li&gt;A differential reference model: a slow, obviously correct implementation with naive data structures, run against the same input, with any disagreement treated as a bug in the fast one&lt;/li&gt;
&lt;li&gt;Fuzzing at the decoder boundary, where malformed input arrives from outside and where a panic would take the matching thread down&lt;/li&gt;
&lt;li&gt;Latency measurement under the burst shape that worries you, recording the distribution rather than an average, since the tail is the number that decides the design&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;A reference model is worth more than it looks. Two implementations written from the same specification disagree in exactly the places where the specification was ambiguous, and matching rules are full of ambiguity at the edges - crossed limits, zero remainders, cancels racing fills.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What Rust gives you here, and what it does not
&lt;/h2&gt;

&lt;p&gt;Rust removes a category of problem rather than making the loop faster by itself. There is no garbage collector, so no pause is scheduled behind your back. Ownership makes the single-writer discipline something the compiler enforces instead of something a code review has to notice. Slab handles and intrusive links, which are error prone in a language without lifetimes, are checkable here. Panics on integer overflow in debug builds catch a class of bug that silently corrupts a book.&lt;/p&gt;

&lt;p&gt;The honest boundary is that most of what determines tail latency is not the language:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Kernel scheduling, interrupt handling, CPU pinning and power management move the tail more than the matching code does&lt;/li&gt;
&lt;li&gt;Network interface, kernel bypass or its absence, and the physical path to the venue set a floor the engine cannot go below&lt;/li&gt;
&lt;li&gt;Serialisation and the wire protocol at the boundary are frequently the dominant cost, not the match itself&lt;/li&gt;
&lt;li&gt;Matching semantics, order types, fee and rebate rules and market phases are business decisions - a wrong rule implemented quickly is still wrong&lt;/li&gt;
&lt;li&gt;Risk checks, position limits and the settlement path live outside the engine and have their own latency and their own failure modes&lt;/li&gt;
&lt;li&gt;Operations - deployment, monitoring, the runbook for a failed replay - decide whether the guarantees survive contact with a production incident&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Rust also costs something. The borrow checker slows down the first weeks of a design that is still moving, the ecosystem for exchange specific protocols is thinner than in older languages, and unsafe blocks around lock-free structures need the same review discipline as the equivalent code anywhere else. Choosing it is a decision about the latency tail and about memory safety in a single-writer core, not a decision about developer comfort.&lt;/p&gt;

&lt;p&gt;amBrain has built trading infrastructure in Yerevan, Armenia since 2019, with hot paths in Rust; a mini-exchange we built runs in production on MOEX colocation. If you are designing a matching engine and want to walk through the book structure, the journal format or the replay harness, that conversation is worth having before the first line of the hot path is written.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://ambrain.org/blog/matching-engine-rust-design/" rel="noopener noreferrer"&gt;ambrain.org&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>rust</category>
      <category>architecture</category>
      <category>systemdesign</category>
      <category>microservices</category>
    </item>
  </channel>
</rss>
