In the fast-paced world of financial trading, every millisecond counts. Gemini, a leading cryptocurrency exchange, has achieved a remarkable 80% reduction in trading latency, bringing round-trip times down to under 2 milliseconds. This significant leap is attributed to a strategic collaboration with AWS, deploying a sophisticated hybrid cloud architecture featuring AWS Outposts and Local Zones, fundamentally transforming how institutional traders interact with their platform.
“The third column is what we built. AWS local zone and outposts and VPC peering and all these technologies combined allow us to deliver sub 3 millisecond end-to-end latency with minutes of setup time. I think that's really the key unlock here.”
- Paget Stanco, Lead, Institutional team at Gemini
Discover how Gemini slashed trading latency by 80% using AWS Outposts and Local Zones. This session reveals the critical infrastructure decisions that transformed their institutional trading platform, ensuring speed, scalability, and deterministic performance in the heart of financial markets.
Hi everyone. So, today we are going to talk about how Gemini cut down our trading latency by 80% using AWS's Outposts and Local Zones. As you just heard, my name is Paget Stanco. I lead the Institutional team at Gemini. I will be joined by Joshua Smith, um who is a senior solutions architect at AWS, as well as my colleague Yu Kong Sun, who is our head of trading solutions at Gemini. Quick disclosure, you don't need to read this. this. >> [laughter] >> [laughter] >> So, just to run through a little bit of the agenda that we're going to speak to shortly here. I'm going to cover Gemini's vision for institutional grade trading. Joshua will then speak to AWS's services and the back ground that we have used to build together, and Yu Kong will speak to a deeper dive in more detail of the actual Gemini infrastructure that we've built. So, Gemini is evolving from a crypto
exchange into a broader markets company. We are focused on the future of money, markets, and AI-driven financial infrastructure. Our company's long-term vision is building a financial super app where users can trade, hedge, spend, and interact with multiple market products in one cohesive ecosystem. We spent more than a decade building the core exchange infrastructure to have crypto buy, sell, store capabilities. Internally, this has included matching engines, order books, custody systems, KYC and onboarding, as well as post-trade reporting systems. Um but it wasn't really until last year of 2025 that marked a big transition. Uh this was our transition into the public markets. So, we've just had a successful IPO in September. And as we see a lot in the crypto space, uh that period of time was also underscored by the core reality of crypto, which is cycles are unavoidable. And even through a low cycle of Q4, we still saw over $11 billion in trading volume. And I think this speaks through many things that we've learned in these cycles, both positive and negative in the past, but the only way to move beyond them and move through them is to actually continue to build. So, what we have continued to build in
2025, as of December when we got our DCM, is a predictions market. So, the prediction market launched not as a separate business line, but actually as an extension of Gemini's existing market infrastructure and expertise. We believe prediction markets represent the next major evolution of capital markets. Faster information discovery, real-time pricing of future events, and markets that can actually used as as truth mechanisms. Gemini is building and operating this prediction marketplace end-to-end rather than licensing any third part party technology. Making this decision has given us full control over the latency, the risk systems, execution quality, and the operational scalability. Since our December DCM launch, as I've mentioned, um things have been accelerating. We had over 15,000 customers engaged shortly after that launch. We've had many expansions and and categories of contracts. So, this has expanded rapidly from monthly to weekly, daily, hourly, 15-minute, 5-minutes as of last week contract intervals. So, it really has created significant higher demands on our infrastructure. We have more market events that we have
to support, more order activity that has to be supported, settlement cycles. And so a lot of this comes down to locality. The actual infrastructure locality matters. We need this in order to reduce network distance, which directly impacts obviously the speed of the execution and thus the quality of the trades that are getting put through. So especially for the high frequency markets and traders that are interacting on the predictions platform, the speed and quality of trading is top of mind. Um next up as you'll see here, we've mentioned our DCO license, which we actually just got 2 weeks ago, which is really exciting and now allows us to act as the clearing house as well.
So alongside the launch of predictions in 2025, we launched a multi-week institutional trading competition. We had real prize money and live trading conditions. This competition was targeted at university quant financial teams to engage the next generation of institutional grade customers and and traders. So we had 51 active teams participate over a 2-month period. They generated more than 20 million in trading volume on the platform in predictions alone. So the event really served as a real-world use case for us to test the trading infrastructure itself, execution quality, the API performance, and of course scalability of our platform. These teams traded in ways very similar that you would see to professional firms. So they had algorithmic rapid execution and overall active portfolio management, which really helped our team reinforce the fact that growing demand for low-latency institutional grade crypto markets is here. That infrastructure matters. So the live trading environment that we saw through this competition and obviously have built through our predictions marketplace today has been one of the best ways that we've been able to validate the platform performance at scale and all that we've built with AWS.
In conclusion, what these traders need from us, I will repeat this again, but it's connection timing, streamlined onboarding, APIs, infrastructure that's built for rapid institutional access makes all the difference. That faster execution of ultra-low latency architecture that is powered by AWS. The outposts and the locals zones that allow for this faster execution, consistency, and all in all trader confidence. It's what keeps them coming back. Capacity, we need to have resilient infrastructure that's institutional grade and systems that are designed for continuous uptime, particularly during volatility spikes. Privacy, this of course is very important to have secure, transparent, and compliant market operations. And finally, one venue. Expanding beyond just crypto, which we started out with, into predictions, perps, futures, equities, and next generation market products. It's really important for us to have this in one venue. So, I'm going to pass it off to Joshua shortly here, but all in all for Gemini, the institutional grade trading is no longer just about this exchange access. It's really about combining the low latency infrastructure, the AI acceleration, and market intelligence to have a vertically integrated one unified platform for our customers.
>> All right. Thank you, Pagid. So, we heard the what, right? What's the goal here? It's a single venue for spot, derivatives, and prediction markets with traders that expect institutional grade execution. So, today I'm going to discuss why AWS and what AWS services power this. So, Gemini ultimately wanted to bring a low latency experience to cloud. What does that mean, right? The old path where traders connect through US East 1, our region in Virginia, round tripped in around 10 to 12 milliseconds back to Gemini's matching engine in NY5. This new path is under 2 milliseconds. So, this is a four to five times multiple improvement, and that's the difference realistically between a cloud-native trader being competitive in this venue or not.
So, this comes down to two things. And there's really no clever engineering around it. Light only moves so fast. You're going to burn multiple milliseconds just going from New York to Virginia before any hops. Distance is latency. And so, this is why the answer was never optimize more in the US East 1 region. We can do that, but we have to move the compute. The second thing, and this is the thing that trading teams really obsess over, it's not the average, it's the determinism. Every hop, every gateway, NAT, firewall, these all don't just add latency, they add jitter. And so, a path that averages a faster 5 milliseconds, but spikes to 50 at the P99 latency is actually worse than a steady-state 8 millisecond. Cuz traders are looking for the tightest P99, 99.9, because that's the path that doesn't miss the market fill when the market moves. And this shows up downstream in the opportunity. A slow market path, or a non-deterministic market path, means that the quote you wanted is gone by the time your order arrives. It means your fills come at worse prices. Market making has to the spreads have to widen to compensate for this. Uh the order flow eventually over time will route somewhere else. So, when we talk about microseconds and milliseconds being money, we're being precise about that. >> [clears throat] >> So, the AWS primitive that lets you actually move the computer, as I was saying, is the local zone. A local zone is an extension of an AWS region that's placed physically inside of a major metro area with a parent region. So, the parent region in this case is Virginia, US East 1, and the child of it is New York City 2A. So, the compute ends up right in the New York City metro region. And you get low millisecond latency to everything in that metro region. So, Gemini's matching engine being in NY5, just across town, the round trip time between them is a millisecond on the line. So, the moment you put your compute in that local zone, you are nearly collocated with everyone else in the metro area. And second, this is the part that really is easy underestimate the importance of, it's managed like a region. And so, it's the same console. It's the same APIs, IAM security, deployments, security posture. A local zone shows up as another availability zone subnet in your VPC. And so, your team doesn't have to learn a new product. You just launch EC2, you launch EBS, EKS. You're using the same experience that you already operate, so it's very easy for institutional risk, audit teams, and for operations teams, because there's no change to the control plane. And thirdly, traffic traverses the AWS private backbone, not the public internet. So, when Gemini needs to reach back to a service or control plane call in US East 1, the traffic stays on the AWS fiber with no additional hops or jitter.
So, there's a few more services in play, and I'm going to walk through them in sort of a flow. We've already covered local zones, and that's the foundation. That's where the trading API runs for low latency. But, we have to build around build the network around it. There's two networking primitives that stitch this metro together. Direct Connect is a dedicated private fiber circuit, in this case from the local zone to Gemini's exchange. There's no public internet, there's no shared peering, there's predictable bandwidth, predictable latency, and for a trading workload, this is incredibly critical. And VPC peering And VPC peering is how clients connect into Gemini once they're in AWS. So, when they set up a VPC network inside the local zone, they can peer directly to the VPC that's Gemini operates in local zone. And this creates a direct private network link between the two networks with no firewalls in the middle, no transit gateway hops, nothing to slow down or make the traffic unpredictable. Just routed traffic on the AWS backbone. And we chose peering here specifically because it's the lowest overhead way to connect two networks. And lastly, we land on Gemini's compute
in the data center. Outposts is AWS infrastructure, but with the same control plane, APIs, and operating model. But the rack exists physically inside of Gemini's cage in NY5. We used second generation Outposts, which are have AMD Solarflare NIC cards, and these support things like kernel bypass to shave tens of microseconds off packets. The platform supports native layer two multicast. It supports precision time sync over PTP, equal cable lengths in the rack. These are the types of deterministic networking primitives that capital markets regulators care about when you talk about fairness and equal access for trading. They're purpose-built for customers like Gemini. This also allows the matching engine itself to run the same console as before, same security posture, without giving up a single microsecond of that cross connect determinism they had inside the data center. And so, what you get with all of this combined, this whole flow, is that end-to-end you get the same consistent experience, you get the ability to deploy same APIs, and you get the same security posture end-to-end across both sides of the path, whether you're operating in the data center or in the cloud. And then two more managed services to
help round out the story. EKS, which is our managed Kubernetes service, is where Gemini's web sockets APIs for their fast API actually operate inside the local zone. This gives them elastic scalability. When news drops, the contracts change rapidly, traders jump on the platform, the pods scale up and when things are quiet they scale back down. MSK, which is managed streaming for Kafka, is the streaming backbone for market data within the region. So, this actually doesn't run inside the local zone. In the local zone Gemini uses an in-memory messaging bus to control that latency. Um but in the region where the bulk of the market data, orders, execution reports are being accessed, all of that moves across the platform on Kafka. And so, MSK takes that operational burden off their plate. And with that said, you have a good foundation and I'm going to pass it over to Yukong to go deeper into how Gemini use these primitives.
>> Yeah, thank you Joshua for the detailed explanation. So, as as you all know, my name is Yukong. I currently lead the trading system team at Gemini. So, what I'm going to do in the next section is to walk you through the journey that we've been through the last year to set up to really leverage our posts and local zone to achieve what we just talked about. So, as Patrick mentioned mentioned, Gemini's the journey started about like 10 years ago. At the time at 10 years ago, Gemini's was very proud to be the first trading system that runs on bare metal offering access to the crypto crypto asset class in the data center in the heart of the financial world. So, so as as of this design, Gemini's trading system actually runs in 5 in in New Jersey and our software is written in C++ and rely on multicast network in the in the system. in the system. So, in order for trading system to trading clients to connect to our trading system, what do you typically do in the finance world is you have to do a physical fiber cross connect. As you can see, um the the process for connect getting connected is long and uh but at the time I was okay. As time moves on, like in the past 10 years, what we have been observed is uh more and more clients start to wanting to connect from AWS. This This is because they have moved their cloud their own trading infrastructure onto the cloud. And also Gemini is quickly outgrowing our own data center needs. We also need to run part of our website on the cloud. So, in order to do that, we have to change both the software and the hardware to uh to create this option. So, first uh we will talk about our post, right? So, remember what I said, Gemini's soft uh trading software was written in C++. We run the data center, rely on the special hardware, we rely on multicast. Uh but that hardware is quickly aging. Right? So, AWS Outposts allow us to install uh essentially the same AWS hardware into our data center and allow us to migrate our our software without change, still leveraging all these hardware uh features such as multicast, such as Broadcom uh SolarFlare cards for low latency access. And AWS Outposts allow us to manage that using the standard AWS uh EC2 APIs. So, to us, that's really a no-brainer. It help us modernize our tech uh technology stack, help us uh reduce our cost, and help us managing our software better. >> [snorts] >> [snorts] >> So, the second part is in order for uh
to allow user to connect to to our trading system through the AWS, we actually have to extend our trading system gateways to the cloud. Right? In order to do that, uh we had we actually had to do uh a software rewrite. Uh, instead of relying on multicast as you would do in the data center world, we actually end up rewriting our software to send the same multicast package through AWS MSK. And that AWS MSK will connect to it will be consumed by some gateway services that will end up providing access to the to the client on AWS through AWS. So, as you can see here, um there we not only we need a hardware changes, we need software changes, but there's there's still compromises, right? Like the user the user who connects through AWS, they will they will spend about 8 to 15 million milliseconds just on the network alone. Just because AWS data the closest AWS data center is actually almost 8 milliseconds away from our trading system. So, no matter how fast our software is, you still got to have to spend all that long. So, the year is the 2025. So, last year was a as Praveen mentioned was kind of a bad a pretty big moment for Gemini. We are working on defining a new Gemini 2.0 architecture and we are we want to launch prediction markets. So, there's a quite a few important decisions was made earlier on to make sure that what we set up to do will support Gemini for the next 10 years. All right, one of the decision we want is we want to launch prediction market as one venue. A lot of other exchanges have to choose to do the different thing. They we don't want to run the totally separate exchange just for prediction market because that wouldn't allow you to trade multiple asset on the same venue at the same time. So, that was very important for us. But, we also didn't want to you know put put users through the same compromise that we did before that to have them choose between the the faster cross connect versus the slow AWS connect. So we worked out a solution together with AWS team to allow us to have a um sub 3 millisecond latency end-to-end. Right? A little bit summary here.
So on the on the first column you have the uh normally what all financial services does. They offer you the fastest access to the exchange. They They but it takes the longest time to set up. We're talking about the months. We're talking about a huge capital commitment. You know, just a couple million dollars just to get in the door and then to start trade. On the on the middle side that's what most people does uh in the industry today. You can connect but you're slow. Right? And there's no guarantee. There's no performance uh improvement. You're going to be uh you're going to be slow. to the HFTs and all the institutions. The third column is what we built. AWS local zone and outposts and VPC peering and all these technologies combined allow us to deliver sub 3 millisecond end-to-end latency with minutes of setup time. I think that's really the key unlock here. It's like once you have that, onboarding a new client is no longer a issue. Like it it's no longer something that you have to do for close to months. Okay, so now we get into a little bit details on how exactly we piece all these technologies together. So in this graph we start from the the right side which is you know, on the top right you see our trading system in the box. So in order to connect the trading system in uh runs in the NY5 data center, two local zone. Uh you use the same technology that you use to connect to AWS in region. Except uh the local zone version of Direct Connect automatically does this in the backend to make sure that your data travels directly to the to the local zone data center, not the in-region data center. So, the difference between, you know, going to West Virginia versus uh just going across the street to the local zone. So, uh So, first you have to do that. Uh this is kind of standard for any sort of like a uh setup, uh you know, trading system setup. Then, in order to actually serve the user, what we work out with a best practice with the um AWS team is that you want to set up this DMZ concept. Instead of connecting you your core trading system directly to all the client, you want to connect them to a VPC tree-like structure where uh only the uh this will make sure only the gateways runs in the DMZ zone, in DMZ VPC, and DMZ VPC will talk to the uh trading trading system, and the DMZ VPC will talk to the client. As a final step, uh as a um as we mentioned before, the last thing that hook all this together is a VPC peering. So, VPC peering is a AWS feature that basically guarantee connect all the three VPCs runs on the AWS backbone together, like as a uh as a like a one network. So, the the benefit for doing that instead of going the public internet is obviously, you know, lower latency, higher performance, also allows you to uh reduce some of the security postures you you need because it guarantees uh the the data package is all encrypted, all um all private. So, like I said, you what you what do we end up building is uh on [clears throat] the on the right side, that's our trading system. In the middle, we have a bunch of DMZ VPCs. On the left, you know, even though we only show one, but you can imagine there's a thousands of VPC uh client own the VPCs that can appear with the DMZ VPC to connect to our trading system in the middle. Uh scalability is also a building feature. The reason we chose this tree-like structure will guarantee us for future expansion ability uh to connect as many users as possible. Even in like data center, you know, connecting a hundred a thousand clients will be a big deal. Here, we can easily do a thousands. >> [snorts] >> [snorts] >> All right. So, uh next [clears throat] I will go into the a little bit uh technical details around how do we set up the gateways to run inside the inside the uh the DMZ VPC. So, AWS uh local zone does have uh some limitations. For example, ALB is not a supported services inside inside a inside a AWS local zone. So, what we end up having to do is we have to sort of create our own in-house uh you know, load load balancing and gateway services. Uh but in order to do that, uh you do need to do some automation to make sure we you want to be able to launch a multiple instance of the the gateway, have them all register with DNS, and then automatically discovered by the client. Remember, we have like thousands of clients, and the solution we end up picking on is we we actually gave each individual user their special URLs, and those URLs are hosted on a public AWS Route 53 console and we end up using this feature called a um um uh ASG termination hook uh life cycle hook to automatically register the DNS IP to the uh to the DNS or automatically deregister the IP when when the machine terminates. So, this allows us to run as many as uh as we want those gateway services and we never have to tell user a single thing about oh, you have to connect with a different IP now. Which is actually a pretty big deal in the data center because we only have like two machines and every month we have to do a rotation. That was a a huge operational burden. And the second part uh for this gateway service to automatically uh discover the backend services. So, when when you run things in a data center, you actually still have this problem but uh much smaller scale because you maybe have like a couple machines that that runs your trading system. But on cloud, you're going to have a you know, a dynamic list of machines that need to be discovered by uh this gateway server. So, what we end up to do is we actually run uh you know, hook into this EKS uh pod life cycle policy. Whenever there's a a new pod being started, which we're talking about a backend service being started, we automatically register that into a uh into a DNS URL that's automatically configured by the gateway app. >> [snorts] >> So, all in all, I think uh um we I mean the we spent the last year prove that there is a path where you can uh have fully automated uh uh uh fully automated uh VPC peering and gateway setup using this uh using this setup. So, a little bit of callback to the earliest slide that you already seen. Uh we said uh you know, a few The The reason we set out to do all this stuff is because we want we don't want the user to choose anymore. We want to be able to provide the user easy connectivities which set up in minutes rather than weeks, uh but also results in better latency. Like you remember, in a data center you got 1 ms in uh if you do it on a database, you get 8 to 15 ms which is a big deal. But here, you get 2 ms, very stable, very uh very private access. And doing all this will allow us to run
all the So, really allow us to run all the products in one venue. And that will bring the that secure the future for the Gemini trading system. >> [snorts] >> [snorts] >> And to end the today's um talk you know, talk on, I think you know, the trading competition we did You also seen this slide earlier. It's a It's the best uh best example of this. You know, when Gemini started, we want if we want to run any sort of trading competition If you If you want to run an event that involves 50 people connecting to an API you can imagine the logistics required for that, right? You're going to worry about a compute, you're going to worry about the physical infrastructure, worrying about how people can get access to the trading system. Instead Instead uh what we end up doing uh in 2025 is we just spend a week um actually so we onboarded 51 trading system teams. These are not like a you know, uh HFTs. These are like a college kids. They want to trade the crypto, they want to learn how to trade in. We gave them access, we set up gave them some instruction, they finished set up themselves, and we end up onboarding 51 uh teams over a course of week. Uh just like that, right? No other exchange have done have been able to do anything like the you know the same the same. And uh the infrastructure we build allow us to host even bigger event in the future. And without any any compromise. Everybody get the same latency, everybody get the same API, same playground. And of course, for us, we we see more trading activities on the exchange. You know, win-win. Yeah. So, this is where I I will I will end the talk. Thank you so much.
No more market misses!
Unprecedented scalability!
Instant low-latency access!
Massive speed boost!
Edge computing advantage!
Scale client access!
Hybrid cloud unlocked!














