Journal

Journal

Scaling a Trading Platform: Moving from a Single Event Loop to Event-Driven Workers

February 21, 20267 min read
Scaling a Trading Platform: Moving from a Single Event Loop to Event-Driven Workers

Scaling a Trading Platform: Moving from a Single Event Loop to Event-Driven Workers

By Mohammed Sadiq Ali, Software Engineer at ThWorks


Disclaimer

This article is a purely technical discussion of trading system architecture and software engineering decisions. ThWorks does not provide investment advice, trading signals, strategy recommendations, or any form of financial guidance.

As per SEBI regulations (Investment Advisers Regulations, 2013 and Research Analysts Regulations, 2014), only registered and licensed entities are permitted to offer investment advice or research recommendations in India. ThWorks is a technology company. We build fintech infrastructure and trading automation solutions exclusively for trading firms, brokerages, prop desks, and other licensed financial entities — not for individual retail investors.

Nothing in this article constitutes financial or investment advice.


What Happened

For weeks, our trading platform ran without a major bug. Strategies fired on time, results matched expectations, and we believed we had solved the stability problem that trips up most trading systems.

Then we onboarded 8 users in our beta cohort, and it started breaking.

Strategies didn't fire at their scheduled minute. Background jobs were skipped. Stop-loss handling was delayed. Order placement, which normally took 7–8 seconds, went past a full minute whenever several users had strategies firing in the same minute. We halted all trading on the platform for a day.

The cause was not a bug in any single component. It was an architecture decision we made on day one, and it only showed itself under load.

This post walks through that decision, why it failed, and what we replaced it with. I was part of the architecture and design group that worked through the rebuild, deciding what to split, in what order, and what got priority.


1. The Original Decision: One Event Loop for Everything

Our platform is built in Python on asyncio. A trading platform has several distinct layers, and in our original design all of them shared a single event loop:

  • the scheduler that fires strategies at their configured time
  • the entry and exit layer that places and manages orders
  • the risk layer
  • the reconciliation layer that syncs our state with the broker
  • assorted background jobs
  • the API layer serving the web interface, in the same container

On paper this is clean. One process, one loop, no inter-service communication, no coordination problems. And with the 1 to 3 test accounts we used during development, it performed exactly as we expected.

That was the problem. We never ran a stress test. The assumption was that once we deployed to the cloud, scaling would be handled by infrastructure. It wouldn't have been. More CPU does nothing for an architecture where every task waits for the one before it.


2. Why It Broke: Starvation

An asyncio event loop runs one task at a time. Concurrency comes from tasks voluntarily yielding when they wait on I/O. If any task takes a long time between yields, every other task on the loop waits, regardless of how important it is.

That is starvation, and our loop had no notion of importance. A stop-loss evaluation and a routine background job had equal standing. Whichever got the loop first held it.

With 1 to 3 users, no task ran long enough for this to matter. With 8, work piled up:

  • Strategies missed their minute. The scheduler couldn't get loop time to fire them.
  • Jobs were skipped. By the time a job got its turn, its window had passed.
  • Stop-loss and risk handling slowed down. The layer that exists to protect positions was queued behind everything else.
  • The whole platform slowed. Since the API shared the container, even loading the dashboard competed with order placement.

The latency numbers show the pattern clearly. If one user had a strategy firing in a given minute, execution was quick. If several users had the same strategy firing in the same minute, the system processed them one user at a time, and the last one waited over a minute. Small-scale testing could never surface this, because small-scale testing never has contention.

For a trading system, this is the worst possible failure mode. Nothing crashes. Nothing throws an error. Things just happen late, and in trading, late is wrong.


3. The First Fix: The Database, Not the Architecture

Before touching the architecture, we halted trading, shipped a short-term workaround, and went through the system layer by layer. We kept asking three questions:

  1. What is actually slowing this down?
  2. What is breaking because of it?
  3. What must run on time no matter what, and what can wait?

The first answer surprised us. The biggest single bottleneck was our database access pattern.

Every cycle, we checked users one by one: whether they were active, whether they had trades scheduled, and so on. Every user was checked individually, including users who were inactive or had nothing scheduled. That came to 14 queries for what is logically one question: which users need work this cycle?

We rewrote it as 2 queries.

That change alone made the platform noticeably faster, before any architectural change. The lesson is uncomfortable but useful: per-user loops against the database are invisible at 3 users and dominant at scale. They are worth auditing before reaching for bigger infrastructure changes.


4. The Rebuild: Event-Driven Workers

The query fix bought time. It did not fix the underlying problem, which was a misunderstanding of what kind of software we were building.

Most applications are request-driven. A request arrives, work is done, the process returns. A trading system is not like that. It runs continuously: watching, waiting, acting, checking, and repeating until the goal is reached or the session ends. Treating it like a request-driven app was the original mistake.

We rebuilt the backend in one week around event-driven workers running in parallel, with each layer isolated in its own container:

The split is about blast radius as much as performance. A slow background job can now only slow down background jobs. A burst of dashboard traffic can't delay an order. Each layer can be scaled and resourced on its own.

Work moves between layers through a queue with workers, backed by Redis. (We are keeping the specific queue implementation to ourselves.)


5. Priority as a First-Class Property

Splitting containers solves isolation. It doesn't solve ordering within a layer. For that, every task in the system now carries a priority.

Time-critical work such as stop-loss handling and risk checks is never queued behind a reporting job or a routine sync. Each layer has its own priority rules and its own monitoring, so when something starts to slow down, we see it in our metrics before a user sees it in their positions.

The principle is simple: if everything has the same priority, nothing does. A trading system has to know, structurally, what matters most.


6. Connection Pooling Without Giving Up ACID

Moving to many parallel workers created a new problem: database connections. Each worker may need several connections, and PostgreSQL's connection limit is reached quickly once worker counts grow.

We put PgBouncer in front of Postgres to pool connections. For us this was not optional. Trading writes have to be atomic, isolated, consistent and durable. A partially written order state or a race between two workers updating the same trade is not an acceptable failure. PgBouncer let us scale the number of workers without hitting connection limits and without loosening transactional guarantees.


7. Results

After the rebuild:

Dashboard with performance metrics: 100% runs executed, 91.1% stop-loss protection, 100% exits, 99% speed. Failures by type and broker issues shown.
trading-engine-health-dashboard
  • Zero missed jobs and zero missed strategies over the last month of operation.
  • Order placement at the broker within 0–1 second, including when many users fire strategies in the same minute. Previously this exceeded a minute under the same conditions.
  • Database queries per cycle reduced from 14 to 2.
  • Load tested and verified to handle hundreds of thousands of users.

Closing

The question isn't "does it work?"

The question is: "Does it still work when everyone uses it at the same minute?"

A single event loop answered the first question perfectly and the second not at all. Our system ran cleanly for 3 users, and that told us almost nothing about 8, and nothing at all about a hundred thousand.

Two things we would tell any team building continuous, time-critical systems:

Don't put everything in one place. Decide early what must run on time, isolate it, and give it priority and resources ahead of everything else. Then build the rest around it.

Load test before your users do it for you. Working for you means working for you. It says nothing about working for everyone else.


Building trading infrastructure that has to hold up at scale? Talk to us at thworks.org.

ThWorks is a fintech technology company. We build execution engines, risk management systems, and trading automation infrastructure for trading firms and licensed financial entities. We do not provide investment advice, trading signals, or financial recommendations of any kind.

Next step

Tell us what is breaking, and what it costs you.

A scoping call is 30 minutes and gets you a fixed estimate. If we are the wrong people for it, we will say so.

See the work
  • Reply within one business day
  • Fixed price, fixed scope
  • You own the code