← All posts

Flash-Sale Ticket Booking System

It is 8:59 p.m., and Asha has been staring at a countdown for ten minutes. At 9:00 sharp, a hundred thousand seats for a two-night concert go on sale, and so does everyone she knows. I made Asha up, but every engineer who has run a sale like this will recognise the three things that happen next.

First, the stampede. A million people press refresh in the same minute. The database connection pool empties, the pages turn into error screens, and the retry logic in every browser and app turns a bad minute into a bad hour.

Second, the double sale. Two fans are told, within the same second, that seat 14F is theirs. One of them gets an apology email at midnight and a story they will tell for years.

Third, the phantom charge. Asha’s bank shows the money gone. Her screen shows a spinner. The tickets never arrive, because the call to the payment provider timed out after the provider had already said yes.

These are three different kinds of failure: overload, concurrency, and ambiguity. None of them is fixed by adding servers. Each needs a specific, named defense, and the defenses have to work together. This article is how I designed one system to have all three. It’s called Encore, a flash-sale platform I built as ten cooperating services, and Asha will come back as we go.

I approached it with a narrower question than “can I design it?”: which of the usual claims can I actually prove, and which am I just asserting? The classic brief reads like a wish list: never sell a seat twice, 100,000 reads a second, 10,000 writes a second, verify a token in under a millisecond, survive a payment provider that drops mid-charge, keep bots out, and let a fan through a stadium gate with no signal. Every line is easy to write and expensive to build. So this walkthrough is opinionated about why, not just what, and the table in the next section keeps me honest about what I can back up.

What we are building

Functional requirements

  • Waiting room. When checkout is at capacity, users queue first-in-first-out, see their position live, and are admitted in bounded batches with a short-lived signed pass.
  • Live seat map. Browse sections and seats, and watch other people’s holds and sales as they happen.
  • Atomic hold. Hold one to six seats all-or-nothing, for ten minutes.
  • Reclamation. A hold that isn’t paid for returns to inventory by itself.
  • Pricing and promotions. Scarcity-based pricing, capped promotion codes, and a price frozen at hold time.
  • Payment. Start a payment with an external provider, and finalize the order only from a signed webhook.
  • Fulfillment. After payment, issue tickets and notify, asynchronously.
  • Entry. Rotating QR codes, and gate validation that keeps working when the venue’s internet doesn’t.

Non-functional requirements. The middle column is what the classic brief asks for. The right column is what I can actually show.

RequirementTarget in the briefWhat I can show
Inventory integrityZero double-bookingEnforced by a database constraint; tested with 12 concurrent buyers
IdempotencyEvery state-changing endpointReserve, pay and cancel carry stored responses; webhooks are deduplicated by event ID
Read throughput100k RPSNot demonstrated. A smoke script exists; it is not a capacity test
Write throughput10k RPSNot demonstrated
Token verificationUnder 1 ms, no network I/ONot met. The gateway makes a private auth call, which is a network hop
Seat-update fan-outUnder 50 msNot measured
Idle connections10k–25k per instanceConnection ownership proven; memory per connection not measured
Seat map payloadAbout 12.5 KB for 50,000 seatsMet. Exactly 12,500 bytes, asserted in a test
Availability and durability99.99%, RPO 0Not demonstrated. One host, single instances

The rules that can never break

Before any boxes and arrows, I write down the properties that must hold regardless of load, retries, or which process crashed. Everything later in this article exists to protect one of these:

  1. A seat has at most one live owner. This is enforced by a unique index, not by application code.
  2. Fast layers may delay or reject, but never grant. Losing every cache key must not double-book a seat.
  3. A hold is all-or-nothing. All six seats, or none.
  4. Money moves only on a verified webhook. Signature, timestamp, provider reference, amount and currency must all match.
  5. A reservation ends exactly one way: paid, expired or cancelled. Every path that changes it locks the same row first.
  6. The price quoted at hold time is the price charged.
  7. Duplicate delivery is harmless. Idempotency keys, webhook fingerprints, an inbox table and a unique ticket per seat absorb every retry.
  8. An idle connection never holds a booking thread or a database connection.

The design on one page

Here is the whole system with no product names on it. Colors mark the plane a component belongs to; the moving dots are traffic.

Technology-agnostic high-level design Fans and the payment provider reach one edge gateway. Behind it are three planes: admit, browse and buy. Below them are three stores: a fast coordination store, the transactional system of record, and a durable event backbone. Fansbrowser · mobile app Venue gatesscan on the venue's own LAN Payment providerintent out · webhook in Edge gateway the only public door · routing · request limits · verifies the admission pass before anything reaches the core manifest sync webhook ADMIT · PROTECTS THE CORE Waiting roomFIFO queue · live positionsockets only, no database access Admission controllerhands out checkout capacityshort-lived signed passes BROWSE · SCALES WITH FANS Read model servicecached seat snapshotsread-only database login Live-update fan-outone upstream feed per sectionslow viewers get dropped BUY · OWNS MONEY AND SEATS Booking serviceholds · prices · paymentsshort transactions only Background workersexpiry · fulfillment · refundsnotifications · retries STATE & MESSAGING Coordination storefast, atomic, disposablequeue order · capacity leasesseat hints · caches · pub/submay lag, is never the authority System of recordauthoritative, transactionalseats · reservations · pricespayments · outboxthe only place ownership lives Event backbonedurable, ordered log plusa retrying work queuefinancial events · notificationsdead-letter for the unlucky read-only outbox relay consume
Three planes that share nothing except the state layer. Fans enter through one edge, split into admit, browse and buy, and every plane leans on the same three stores: a fast one that may be wrong (coordination), an authoritative one (record) and a durable pipe (events).

The most important decision in that picture isn’t any single box. I separate the system by what each part holds scarce, not by feature. The admit and browse planes hold connections: memory, sockets, file descriptors. Browse also absorbs read bursts. The buy plane holds locks and database connections, the one resource that can’t be scaled by adding instances. If those three share a process, a flood in one starves the others, and the failure you get is the stampede from the opening.

The separation buys three specific guarantees. An idle socket never occupies a booking thread or a database connection. The components that face the public and hold sockets have no database credentials at all. And each plane can fail, restart, and scale independently. I wanted to test that last claim rather than assert it, so there’s a drill in the repo that stops the buy plane while clients stay connected and checks that browsing and queueing keep working.

There’s also a single dial. The admit plane decides how many people are inside checkout at once, and that one number is the lever that protects everything downstream. We’ll come back to it twice: once in the waiting-room design, and once when I talk about operating the system.

Choosing the technology

The design above is deliberately technology-agnostic. Now I’ll pick technologies, and explain each pick by the role it has to play, because the decisive boundary is where scarce resources are held, not which language is fashionable. A language choice alone proves nothing about throughput. It only makes some designs easier to express.

RoleMy pickWhy this, for this roleWhat I gave up, and what would change my mind
Edge gatewayNGINXOne public port, cheap idle connections, request and time limits, WebSocket upgrade, unbuffered streaming, and a private auth call to verify the passAn in-process, sub-millisecond token check. An ingress JWT module or a mesh proxy would change my mind
Waiting room and admissionNode.jsThe work is I/O-bound: many idle sockets, tiny messages, plus signing tokens. The event loop is a good fit. No database credentials, everCPU-heavy work would stall the event loop, so none goes there. Go would be a fair swap; the boundary matters more than the language
Read model and live fan-outGoCheap concurrency for streaming and a bounded, per-client channel is idiomatic, which is the whole backpressure story. Small footprint per connectionAnother toolchain to build and patch
Booking service and workersJava 21, Spring Boot, plain JDBCMature transactions, bounded pools, graceful shutdown, health and metrics. I kept SQL explicit instead of using an ORM so the lock behavior is visible in code reviewHeavier memory. That’s fine, because it holds no idle sockets by design
System of recordPostgreSQLPartial unique indexes and check constraints turn invariants into constraints. It has row locks with NOWAIT and SKIP LOCKED, advisory locks, transactional outbox writes, and database timeA single writer per event. If the hottest event outgrows one writer, I’d move to one authoritative writer per section, not to a weaker consistency model
Coordination storeRedisAtomic scripts over several keys, sorted sets for order plus leases plus heartbeats, TTLs, pub/sub, and sub-millisecond latencyDurability. That’s why everything in it is treated as disposable and every consumer fails closed
Financial event logKafkaDurable, ordered per key, replayable; offsets advance only after processing succeedsOperational weight
Notification queueRabbitMQPer-message acknowledgement, delayed retry queues and a dead-letter queue. That’s a work queue semantic, which a log doesn’t give you nativelyA second broker to run. Different delivery semantics justify it
PaymentsHosted provider SDK, idempotency keys, signed webhooksNo card data ever touches my servers, which shrinks compliance scopeProvider lock-in
Venue ledgerSQLite in WAL mode behind a tiny LAN serviceA primary key on the ticket ID makes double-entry atomic with no cloud dependencyHigh availability. A consensus-backed ledger would be the upgrade

Two of those picks deserve a paragraph, because they’re the ones people argue about.

Why Redis is a bouncer and not the owner. The textbook flash-sale design puts the seat lock in memory: fast, simple, sub-10 ms. I deliberately did not. If the cache is the source of truth for a hold, then a key expiring at the wrong instant, a failover, or a crash between “write to the cache” and “write to the database” can leave the two disagreeing about who owns seat 14F, and one of them has already taken someone’s money. So ownership lives in PostgreSQL, and Redis only turns obvious contention into a cheap rejection. Every Redis failure then degrades into a slower or rejected request, never a wrong sale.

Why Kafka and RabbitMQ. They aren’t redundant. Financial events need an ordered, replayable log where a failed event stops progress and stays retryable rather than being parked in a queue nobody watches. Notifications need the opposite: independent messages, per-message retry with growing delays, and a dead-letter queue for the ones that keep failing. Using one broker for both means forcing one set of semantics onto a workload that wants the other.

The data model

PostgreSQL holds everything that must be right: events, seats, reservations, per-seat price snapshots, promotions, payments, tickets, and the plumbing for reliable messaging. Three constraints carry the whole “never double-book” promise, exactly as they appear in the schema:

CREATE UNIQUE INDEX one_active_checkout
    ON reservations (user_id, event_id)
    WHERE status = 'PENDING_PAYMENT';

ALTER TABLE seats
    ADD CONSTRAINT seats_owner_state
    CHECK (
        (status IN ('AVAILABLE', 'BLOCKED') AND reservation_id IS NULL)
        OR (status IN ('HELD', 'SOLD') AND reservation_id IS NOT NULL)
    );

CREATE UNIQUE INDEX one_live_owner_per_seat
    ON reservation_items (seat_id)
    WHERE status IN ('HELD', 'SOLD', 'CHECKED_IN');

one_live_owner_per_seat is invariant number one. It’s a partial unique index: released and expired items fall out of it, so a seat can be held again later, but two live items can never point at the same seat. seats_owner_state means a seat can’t be HELD without naming who holds it. one_active_checkout stops one user from hoarding by opening a dozen pending checkouts.

Principal-level talking point: constraints beat code. An application-level check (“is this seat free?”) is a promise that every code path, every future engineer and every retry will remember to make. A constraint is a promise the database makes on their behalf. Design so that the worst bug you can write fails loudly at the constraint instead of silently double-selling.

Prices are frozen per item (base_price, surge_bps, discount, price_at_lock), so a later price change can’t alter what someone already agreed to pay. The outbox table is the other important one, and it shows up again in the payments section:

CREATE TABLE outbox (
    id BIGSERIAL PRIMARY KEY,
    kind TEXT NOT NULL,
    aggregate_id TEXT NOT NULL,
    payload JSONB NOT NULL,
    created_at TIMESTAMPTZ NOT NULL DEFAULT clock_timestamp(),
    published_at TIMESTAMPTZ
);

CREATE INDEX outbox_unpublished
    ON outbox (id)
    WHERE published_at IS NULL;

The waiting room

Back to Asha. It’s 9:00:01 and a million people are on the site. This is where the stampede gets stopped, and the waiting room is the most misunderstood component in the design. It looks like a UX feature, a “you’re number 4,812” screen. It is actually a concurrency limiter for your database.

The problem it solves

A database can only do so much work at once, and the limit has nothing to do with how much traffic you receive. It is set by how many transactions can hold a connection, and how long each one holds it. That relationship is Little’s law: concurrent work equals arrival rate times time in system. In my booking service the transaction pool is 12 connections. Suppose a hold transaction keeps its connection for about 25 ms. I haven’t profiled that number, so treat it as an assumption. Then the pool can finish roughly 12 ÷ 0.025 = 480 holds a second, no matter how many servers sit in front of it.

Now suppose only one in ten of a million fans reaches a hold attempt in the first minute. That’s 100,000 attempts over 60 seconds, about 1,700 per second, against a system that can finish about 480. Nothing about the architecture is wrong. The load is simply 3.5 times what the scarce resource can carry, and what happens next is worse than “some requests are slow”.

Without a limiter the excess doesn’t wait politely. It piles up, requests time out, and clients retry, so the offered load rises just as the system is least able to cope. Goodput, meaning the purchases that actually complete, doesn’t plateau at capacity. It collapses below it:

Goodput with and without admission control Illustrative curves. Both rise together until the capacity knee. With a waiting room, successful purchases per second hold flat. Without one, they fall as queues, timeouts and retries take over. capacity knee: the pool and the lock queues saturate with a waiting room: the excess waits in line, goodput holds without one: queues, timeouts, retries, then goodput falls offered load (arrivals per second) → successful purchases per second
Illustrative shape, not a measurement. Both systems behave identically until the knee. What differs is what happens to the load beyond it: one queues it in memory in front of the database, the other lets it in.

The second thing the waiting room buys you is fairness: first in, first out is a promise you can explain to a customer, while “whoever’s retry logic happened to land on a healthy server” is not. The third is protecting everything else that shares the pool. In my design, the payment webhook is served by the same booking service as the seat holds. An unbounded stampede of hold attempts competes for the same request threads and database connections as the webhook that confirms a paying customer’s order. Ask yourself what your webhook latency looks like at 9:00:01 on a night like this.

Lock contention, up close

To see why the limit matters, look at what a hold does. Asha and 999 other fans all want seat 14F. Each one runs, in effect, a transaction that ends in SELECT ... FOR UPDATE on the seat’s row. The database grants that lock to one transaction and makes the rest wait in line.

That has consequences that aren’t obvious from a diagram:

  • Throughput on one hot row is capped by how long the lock is held. If a transaction holds the seat lock for 25 ms, that row can change hands about 40 times a second, whether you have 10 servers or 1,000.
  • Waiters queue behind each other. The k-th waiter waits about k × 25 ms. A thousand waiters means the last one waits 25 seconds, far longer than a client will tolerate.
  • Every waiter occupies a connection. A blocked transaction does no useful work but still pins one of the 12. A handful of hot rows can hold the whole pool hostage while colder seats, which could have been sold immediately, wait for a connection.
  • Deadlocks appear when carts overlap. Buyer A wants seats 5 and 9. Buyer B wants 9 and 5. If they lock in different orders, each waits for the other, and the database eventually notices and kills one of them, after its deadlock timeout, with the transaction’s work wasted.

My defense against that last one is boring and reliable: the service sorts the requested seat IDs before touching the database, so every buyer locks in the same order. Two buyers who both want 5 and 9 will both try 5 first, and neither can hold 9 while waiting for 5.

FOR UPDATE, NOWAIT and SKIP LOCKED at scale

The choice of how to wait for a row is where a lot of the real design lives. Here’s how I think about the options, and where each one breaks:

StrategyWhen the row is busy…Good forWhat breaks at flash-sale scaleWhere Encore uses it
FOR UPDATEWaits its turn until the holder commits or rolls backPaths where the caller must have that exact rowWaiters pin connections; the k-th waiter waits k × hold time; the tail becomes timeouts and retriesSeat locks in the hold transaction, bounded by the waiting room and a 2-second lock timeout
FOR UPDATE NOWAITFails immediately with a lock-not-available errorInteractive picks where a busy seat probably won’t stay free; protecting the poolRetry storms if clients fire again instantly; no queue fairness; needs jittered backoff and a UX for “someone is holding those seats”Not used. It’s the first lever I’d add at 10× scale
FOR UPDATE SKIP LOCKEDSilently skips locked rows and returns whatever is leftWork queues where any row will doWrong for “I want 14F”: you’d get a different seat, or fewer than you asked for, non-deterministicallyExpiry worker, where any expired reservation is fair game
lock_timeoutWaits, but gives up after a boundA safety net around every blocking pathBounds the damage but doesn’t reduce contention; set it too high and the pool still drainsSet on every connection: 2 s lock timeout, 5 s statement timeout
Advisory lockAn application-defined mutex on a hashSerializing something with no natural rowA single global hot key becomes a global bottleneckIdempotency check per user and operation; the outbox relay singleton
Optimistic update (UPDATE ... WHERE status='AVAILABLE')Under the default isolation level, the update still waits for a competing writer, then re-checks the conditionSingle-row state changesIt doesn’t remove the wait. A multi-seat cart still needs one transaction, so locks are held until commitNot used. The version column exists for it

The two subtle rows are the last one and SKIP LOCKED. People reach for optimistic updates hoping to escape locking, but in SQL the update itself takes a row lock and holds it to the end of the transaction, so under contention you still queue. And SKIP LOCKED feels like the answer to hot rows, but it changes the meaning of the query. That’s exactly what you want for “give me any 100 expired holds to clean up”, and exactly what you don’t want for “give me seat 14F”.

Principal-level talking point: how I’d scale the hottest event 10×. First, keep the waiting room. It bounds concurrency before any lock strategy matters. Second, change the seat locks to NOWAIT (or a much shorter lock_timeout), with jittered client retries and a clear “someone is holding those seats” message, so that a hot seat costs a fast failure rather than a pinned connection. Third, add a best-available purchase mode (“any two adjacent seats in section B”) and implement it with SKIP LOCKED, which spreads contention across all the free seats instead of piling every buyer onto the same one. I haven’t built that mode; it is where SKIP LOCKED earns its place. Fourth, shard the hottest event by section, with one authoritative writer per section. Fifth, give the payment webhook its own connection pool so a stampede can never starve a paying customer.

What the backend feels without a waiting room

Here is the cascade, using the real limits from my booking service configuration: 64 request threads, a queue of 64 more connections, a pool of 12 database connections with a 3-second wait for one, and a 2-second lock timeout.

  1. 9:00:01. Attempts arrive at roughly 1,700 a second. The 12 pooled connections are busy within milliseconds.
  2. Within a second. The rest wait for a connection. The 64 request threads fill up with requests parked on the pool.
  3. After 3 seconds. Requests that never got a connection fail with a timeout. The client sees a 503 and, in many clients, retries at once, so the arrival rate goes up.
  4. Meanwhile. Connections that did get through pile onto a few hot seats and sit blocked on row locks for up to 2 seconds each. They’re occupied and doing no useful work.
  5. Now the queue behind Tomcat overflows. New connections are refused outright. This isn’t a slow system anymore. It’s a system that answers most callers with nothing.
  6. The quiet disaster. A payment webhook from a customer who paid a minute ago is stuck in that same queue. The provider retries later. The hold may expire in between, so that customer’s payment lands as a refund. Asha would be charged and refunded in a single evening.

With the waiting room, the same one million fans see a position number. Admission caps concurrent checkouts, so the transaction pool only ever sees the load it can carry. The stampede still exists. It just lives in a lightweight, connection-only tier with no database access, where holding a million idle sockets is affordable.

How admission is designed

Every aspect of the waiting room is a small decision with a reason behind it. I’ll take them in order.

Ordering: a sequence number, not a timestamp. The queue is a sorted set per event, scored by a monotonic counter that’s incremented atomically at join time. Timestamps would tie whenever two fans join in the same millisecond, and clocks disagree between servers; a counter gives strict, tie-free order. Rejoining doesn’t reset your place, so refreshing the page doesn’t send you to the back of the line.

Capacity, not rate. I don’t admit “N users per second”. I admit until N are inside checkout. Because the scarce resource is concurrency, a concurrency limit self-corrects: if payments get slow and people linger, fewer slots free up and admission slows on its own; if people finish faster, it speeds up. A pure rate limit would keep admitting at a fixed pace into a system that was already backing up. There’s also a per-tick batch cap so a burst of freed capacity can’t release a thundering herd all at once.

Two different Little’s-law budgets. The database pool is sized by transaction time: how long a hold keeps a connection. Admission capacity is sized by time in checkout: how long a fan spends browsing seats and paying. These are different numbers, often minutes versus milliseconds, and mixing them up is a classic mistake. As a worked example, to sell 25,000 orders in 30 minutes I’d need about 14 orders a second; at a mean of 90 seconds in checkout, admission capacity should be around 1,250 concurrent checkouts, while the same 14 holds a second only keep a pool connection busy for a third of a second per second. My local defaults are 100 concurrent checkouts with a batch of 10 per tick. They’re demo values, not tuned ones.

A single elected admitter. If ten admission workers each admit a batch, you’ve admitted ten batches. So workers compete for a one-second lease per event, and only the winner runs that tick. The lease also removes the need for any coordination between the workers beyond one atomic key.

Abandonment: heartbeats. People close laptops. A queue entry that outlives its owner would eventually be admitted and burn a checkout slot. Each queued fan refreshes a heartbeat every 5 seconds over the open socket, and entries silent for 120 seconds are dropped. The cleanup is capped at 1,000 per tick so a giant abandoned queue can’t monopolize the store, which creates a subtle trap the scripts below handle explicitly.

The pass: asymmetric signature plus a server-side lease. Admission produces a signed pass, valid at most five minutes and bound to the user and event. It’s signed with a private key that only the issuer holds, so everything that verifies it needs only a public key and can’t mint passes. But a signature can’t be revoked, so the pass is not sufficient by itself: the booking service also requires an active lease on the server side. A perfectly valid but stale pass can’t start a checkout.

Transport: a WebSocket, not polling. A million fans polling every few seconds is a million requests a few times a minute. A held socket with a five-second server push is far cheaper for the server and gives the fan a live number. That only works because the tier that holds the sockets has no database behind it.

Failure: fail closed. If the coordination store is lost, queue, admission and rate limits all return a retryable error. Fans rejoin the line. Nobody is granted a seat, because the store never granted seats in the first place. The cost is an ugly minute; the alternative is a bypass.

Here is the whole flow. Watch where the fan waits and where the booking service finally sees them:

%%{init: {"themeVariables": {"actorBkg": "#FFFFFF", "actorBorder": "#0E1A2B", "signalColor": "#22324A", "noteBkgColor": "#FFFFFF", "noteBorderColor": "#0E1A2B", "labelBoxBkgColor": "#0B8580", "labelBoxBorderColor": "#0B8580", "labelTextColor": "#FFFFFF", "loopTextColor": "#0E1A2B", "sequenceNumberColor": "#FFFFFF"}}}%%
sequenceDiagram
    autonumber
    box rgb(232,238,253) Outside
        actor Fan
    end
    box rgb(224,244,242) Admit plane
        participant WR as Waiting room
        participant AC as Admission controller
    end
    box rgb(252,231,227) Fast, disposable
        participant CS as Coordination store
    end
    box rgb(251,240,217) Buy plane
        participant BK as Booking service
    end
    rect rgb(232,238,253)
        Note over Fan,BK: 1 · Join the line
        Fan->>WR: Join the queue with a signed session token
        WR->>CS: Atomic join, next sequence number, heartbeat
        WR-->>Fan: You are number 4,812
    end
    rect rgb(224,244,242)
        Note over Fan,BK: 2 · Wait cheaply
        loop every 5 seconds while queued
            Fan->>WR: Socket ping with token
            WR->>CS: Refresh heartbeat, read position
            WR-->>Fan: Position update
        end
    end
    rect rgb(252,231,227)
        Note over Fan,BK: 3 · Admission tick, one winner per event per second
        AC->>CS: Take the one-second tick lease
        AC->>CS: Drop expired leases and stale members
        AC->>CS: Move min(free capacity, batch) from the head of the line
    end
    rect rgb(251,240,217)
        Note over Fan,BK: 4 · Enter checkout
        WR->>CS: Is there an active lease for this fan?
        WR-->>Fan: Signed pass, five minutes at most, socket closes
        Fan->>BK: Reserve seats, presenting the pass
        BK->>CS: Verify the active lease, fail closed on error
        BK->>CS: After commit, extend the lease to the hold duration
    end

Joining and admitting are each a single Lua script, so they run atomically inside the store. Joining reads the store’s own clock (so replicas can’t disagree about “now”), keeps an existing place in line, and assigns a sequence number only once:

-- KEYS: queue, active leases, sequence, heartbeat; ARGV: subject.
local clock = redis.call('TIME')
local now = tonumber(clock[1]) + tonumber(clock[2]) / 1000000
redis.call('ZREMRANGEBYSCORE', KEYS[2], '-inf', now)
if redis.call('ZSCORE', KEYS[2], ARGV[1]) then return -1 end

if not redis.call('ZSCORE', KEYS[1], ARGV[1]) then
  local sequence = redis.call('INCR', KEYS[3])
  redis.call('ZADD', KEYS[1], sequence, ARGV[1])
end
-- Admission removes stale queue/heartbeat members together. Do not prune only one here.
redis.call('ZADD', KEYS[4], now + 120, ARGV[1])
return redis.call('ZRANK', KEYS[1], ARGV[1]) + 1

Admission is where capacity, batch limits and the abandonment trap live together:

-- KEYS: queue, active leases, sequence, heartbeat; ARGV: capacity, batch, lease TTL.
-- Redis time avoids disagreement between admission replicas.
local clock = redis.call('TIME')
local now = tonumber(clock[1]) + tonumber(clock[2]) / 1000000
redis.call('ZREMRANGEBYSCORE', KEYS[2], '-inf', now)

-- Keep each cleanup bounded so an abandoned queue cannot monopolize Redis.
local stale = redis.call('ZRANGEBYSCORE', KEYS[4], '-inf', now, 'LIMIT', 0, 1000)
for _, subject in ipairs(stale) do
  redis.call('ZREM', KEYS[1], subject)
  redis.call('ZREM', KEYS[4], subject)
end

local freeSlots = tonumber(ARGV[1]) - redis.call('ZCARD', KEYS[2])
local attempts = math.min(freeSlots, tonumber(ARGV[2]))
local admitted = {}
for index = 1, attempts do
  local entry = redis.call('ZPOPMIN', KEYS[1], 1)
  if #entry == 0 then break end
  local subject = entry[1]
  local heartbeatUntil = redis.call('ZSCORE', KEYS[4], subject)
  redis.call('ZREM', KEYS[4], subject)
  -- More than 1,000 stale members may remain after cleanup. Never reactivate them.
  if heartbeatUntil and tonumber(heartbeatUntil) > now then
    redis.call('ZADD', KEYS[2], now + tonumber(ARGV[3]), subject)
    table.insert(admitted, subject)
  end
end
return admitted

The comment near the end is the trap I mentioned. Cleanup is capped, so abandoned users can still be at the front of the line afterward. Each candidate’s heartbeat is therefore re-checked as it’s popped; without that check, a fan who closed their laptop an hour ago would be admitted and burn a slot for five minutes. There’s a regression test for exactly that case.

Principal-level talking points for the waiting room

  • It’s a concurrency limiter for the database, not a UX feature. Size it from the transaction budget and the checkout time, not from the traffic forecast.
  • Admission by capacity self-corrects; admission by rate doesn’t.
  • The fair alternative to first-in-first-out is a lottery for the first few minutes, which removes the advantage of fast connections and simple bots. I chose FIFO because it’s predictable and easy to explain, and I’d revisit that if bots became the problem.
  • The wait estimate shown to the fan is a rough function of the batch size. It can’t know how long people ahead of you will linger in checkout, so present it as an estimate, never a promise.
  • Fail closed. Every failure mode of the queue must end with “nobody got in”, never “everybody got in”.

Browsing: a million eyes, one source of truth

While Asha waits, and after she’s admitted, she watches the seat map. This is the browse plane, and its problem is the reverse of the buy plane’s: enormous read volume with very forgiving correctness. A seat map that’s a second stale is fine, because the database will refuse the hold if the seat is gone.

Reads: cache, and stop the stampede on the cache itself. Seat snapshots are served from a short-lived cache (a couple of seconds). On a miss, one request takes a fill lease, a random token with a 5-second expiry, and refills the cache while other readers wait briefly or get a retryable 503. So a thousand simultaneous misses collapse toward one database query. The lease is released only if its token still matches, so a slow request can’t delete another request’s lease. The read service also has its own database login that can only SELECT, with a pool of four connections, so even a bug there can’t write.

Payload: a bitmap. Full seat state for a section is available as two bits per seat, four seats to a byte:

// Four two-bit states fit in each byte, ordered most significant pair first.
func writeBitmap(w http.ResponseWriter, data []byte) {
	var seats []Seat
	if json.Unmarshal(data, &seats) != nil {
		httpkit.Error(w, 503, "Invalid snapshot")
		return
	}
	packed := make([]byte, (len(seats)+3)/4)
	states := map[string]byte{"AVAILABLE": 0, "HELD": 1, "SOLD": 2, "BLOCKED": 3}
	for i, seat := range seats {
		packed[i/4] |= states[seat.Status] << uint(6-2*(i%4))
	}
	w.Header().Set("Content-Type", "application/octet-stream")
	w.Header().Set("X-Seat-Count", fmt.Sprint(len(seats)))
	w.Header().Set("X-Seat-Order", "section,row_num,seat_num")
	_, _ = w.Write(packed)
}

Fifty thousand seats become exactly 12,500 bytes, and a test asserts that number. Seat order comes from a fixed layout, so the payload carries only state. It is also strictly advisory; nothing decides ownership from it.

Updates: push an invalidation, not a delta. When a seat changes, the browser doesn’t receive “14F is now HELD”. It receives “section B changed, go fetch the current picture”. That’s less clever and much more robust: a delta stream needs ordering and replay, while an invalidation can be lost and the next refresh repairs it. This matters because the transport is Pub/Sub, which is not durable. Every stream therefore opens with a reset telling the browser to fetch a fresh snapshot, and browsers also refresh periodically.

Fan-out: one upstream subscription per section, however many viewers. The naive design gives every browser its own subscription: 5,000 people looking at section B means 5,000 subscriptions. Mine inverts that. Each viewer is a small struct with a bounded queue, and each active section has exactly one subscription per fan-out instance:

func (h *Hub) Broadcast(topic, message string) {
	h.mu.Lock()
	defer h.mu.Unlock()
	s := h.topics[topic]
	if s == nil {
		return
	}
	for c := range s.clients {
		select {
		case <-c.Done:
			continue
		default:
		}
		select {
		case c.Messages <- message:
		default:
			c.stop()
			h.Slow.Add(1)
		}
	}
}

The inner select with a default is the entire backpressure policy. Each viewer’s channel holds eight messages. If a phone on a bad network can’t keep up and its channel is full, it’s disconnected, not waited on, so one slow client can’t delay the other 4,999. When the last viewer leaves a section, its subscription closes.

Principal-level talking points for browse

  • Push an invalidation, pull the truth. It turns a hard ordered-delivery problem into an easy eventually-correct one.
  • Bound every queue. An unbounded per-client buffer is a memory leak with a slow phone’s name on it.
  • Drop the slow consumer, not the fast ones. Backpressure is a policy decision; make it explicit.
  • The read side reaches the database with a credential that can’t write. Least privilege is also a blast-radius decision.

Reserving seats

Asha has been admitted. She picks seats 14F and 14G. So does someone else. This is the double sale, and everything in the rules section is about it.

%%{init: {"themeVariables": {"actorBkg": "#FFFFFF", "actorBorder": "#0E1A2B", "signalColor": "#22324A", "noteBkgColor": "#FFFFFF", "noteBorderColor": "#0E1A2B", "labelBoxBkgColor": "#A8721A", "labelBoxBorderColor": "#A8721A", "labelTextColor": "#FFFFFF", "loopTextColor": "#0E1A2B", "sequenceNumberColor": "#FFFFFF"}}}%%
sequenceDiagram
    autonumber
    box rgb(232,238,253) Outside
        actor Fan
    end
    box rgb(251,240,217) Buy plane
        participant BK as Booking service
    end
    box rgb(252,231,227) Fast, disposable
        participant CS as Coordination store
    end
    box rgb(238,232,250) Authoritative
        participant DB as System of record
    end
    Fan->>BK: Reserve 1 to 6 seats, pass, idempotency key
    BK->>CS: Verify the active lease and the rate limit
    BK->>DB: BEGIN, take the idempotency lock
    alt identical retry
        DB-->>BK: Stored response
        BK-->>Fan: Original response, nothing re-executed
    else new request
        rect rgb(252,231,227)
            Note over BK,DB: 1 · Cheap rejection first
            BK->>DB: Any pending checkout for this user?
            BK->>CS: Hint all the seats atomically, all or none
        end
        rect rgb(238,232,250)
            Note over BK,DB: 2 · Lock and verify
            BK->>DB: Lock the seat rows in sorted order
            BK->>DB: Is every seat AVAILABLE?
            BK->>DB: Lock the promotion row, check its caps
        end
        rect rgb(251,240,217)
            Note over BK,DB: 3 · Price and write
            BK->>DB: Freeze prices, insert reservation and items
            BK->>DB: Seats to HELD, versions up
            BK->>DB: Outbox row: inventory changed
        end
        rect rgb(226,244,234)
            Note over BK,DB: 4 · Commit
            BK->>DB: Store the response and COMMIT
            BK->>CS: Extend the admission lease
            BK-->>Fan: 201 with frozen total and expiry
        end
    end
    opt any step fails
        BK->>DB: ROLLBACK
        BK->>CS: Owner-checked release of seat hints
        BK-->>Fan: 409 conflict or a retryable 503
    end

The code, unabridged:

private HoldResult createHold(
    String userId, UUID reservationId, ReserveCommand command, String idempotencyKey) {
  HoldResult previous =
      idempotency.lock(userId, "reserve", idempotencyKey, command, HoldResult.class);
  if (previous != null) return previous;

  cache.requireActive(command.eventId().toString(), userId);
  requireNoPendingCheckout(userId, command.eventId());
  Event event = requireOpenEvent(command.eventId());
  cache.hold(
      command.eventId().toString(),
      command.seatIds().stream().map(UUID::toString).toList(),
      reservationId.toString());

  List<Seat> seats = lockAvailableSeats(command.eventId(), command.seatIds());
  int discountBps = lockPromotion(command.coupon(), userId);
  List<LockedPrice> prices = freezePrices(event, seats, discountBps);
  int total = prices.stream().mapToInt(LockedPrice::priceAtLock).reduce(0, Math::addExact);
  var expiresAt = bookings.databaseExpiry(settings.holdSeconds);

  bookings.insertReservation(
      reservationId,
      userId,
      command.eventId(),
      expiresAt,
      total,
      event.currency(),
      command.coupon());
  consumePromotion(command.coupon(), userId, reservationId);
  prices.forEach(item -> bookings.insertHeldItem(reservationId, item));
  publishInventoryChanges(command.eventId(), seats);

  HoldResult result =
      new HoldResult(
          reservationId,
          command.eventId(),
          ReservationStatus.PENDING_PAYMENT,
          expiresAt,
          total,
          event.currency(),
          prices);
  idempotency.save(userId, "reserve", idempotencyKey, command, result);
  return result;
}

Reading top to bottom, the decisions worth defending:

Idempotency comes first. Networks retry and people double-tap. The service takes an advisory lock scoped to the user and operation, so concurrent retries queue behind each other, then looks up the stored fingerprint for the key. An identical retry gets the original response back, and the same key with a different request is rejected as a conflict. The stored response is written in the same transaction as the reservation, so a crash can’t leave a reservation with no recorded answer, or an answer with no reservation.

Cheap checks before expensive ones. Rate limits, the admission lease, the “one active checkout per user” rule and the seat hints are all cheap, and they run before any seat row is locked. The cheapest way to handle a doomed request is to refuse it before it costs a lock.

Deterministic lock order. The sort I described earlier happens before this function runs, and this is the loop it enables:

public List<Seat> lockAvailableSeats(UUID eventId, List<UUID> seatIds) {
  return seatIds.stream()
      .map(
          seatId ->
              jdbc.query(
                  "SELECT * FROM seats WHERE id=? AND event_id=? FOR UPDATE",
                  rs -> rs.next() ? mapSeat(rs) : null,
                  seatId,
                  eventId))
      .toList();
}

Then the code checks that every seat is AVAILABLE and fails the whole cart if any isn’t, so a hold is genuinely all-or-nothing.

The fast layer is a bouncer. Note that the hint step sits before the locks. It’s one atomic script across the whole cart:

-- Check the entire cart before writing any hint. SQL remains the ownership authority.
-- KEYS: seat hints in one event hash slot; ARGV: reservation ID, TTL seconds.
for _, key in ipairs(KEYS) do
  local owner = redis.call('GET', key)
  if owner and owner ~= ARGV[1] then return 0 end
end
for _, key in ipairs(KEYS) do
  redis.call('SET', key, ARGV[1], 'EX', ARGV[2])
end
return 1

Its only job is to turn obvious contention into a cheap rejection: if someone has already hinted seat 14G, this buyer is bounced without taking a database lock. The release script is its mirror image, and it compares the reservation ID, not just the user:

-- An old release must never delete a hint acquired by a newer reservation.
for _, key in ipairs(KEYS) do
  if redis.call('GET', key) == ARGV[1] then
    redis.call('DEL', key)
  end
end
return 1

That matters because cleanup is asynchronous: a delayed release for an expired reservation must never delete the hint of the next buyer who took the same seat.

Proving the cache can’t cause a double sale. I tested it directly. A test holds a seat, then deletes its hint from the cache by hand, then lets a second buyer request that seat plus another:

cache.redis.delete('hold:{' + event_id + '}:' + seats[0])
_, other = admitted(client, event_id)
assert hold(client, event_id, seats[:2], other).status_code == 409

The database rejects the second buyer with a 409, and the other seat in their cart stays AVAILABLE, which shows the cart rolled back as a unit. A second test admits 12 buyers, has them all request the same two seats from 12 threads at once, and asserts exactly one 201 and eleven 409s, with exactly two HELD items in the table afterward. Twelve is a small number, so this shows correctness under a modest race rather than behavior at scale, but it’s the invariant working end to end.

The lock modes I chose, and why. Seat locks are plain blocking FOR UPDATE, because after admission and the fast-layer bounce, the number of true contenders for any one seat is small, and a buyer who needs that seat has nothing better to do than wait a few milliseconds. The blocking is bounded on every connection by a 2-second lock timeout and a 5-second statement timeout, so no transaction can wait forever. Everything I described in the lock trade-off table above is what I’d change first as contention grew.

Principal-level talking points for the hold

  • The database owns the invariant; everything else is advice. Then any failure of the advice degrades to “slower or rejected”, never “wrong”.
  • Lock in a global order. It’s the one-line fix to a whole class of deadlocks.
  • Order the checks by cost, cheapest first, so doomed requests are refused before they cost a lock.
  • Make retries safe by design. Store the response with the state change in the same transaction.
  • Timeouts are part of the design, not configuration. A lock timeout is what converts “the system is stuck” into “this request failed fast and can be retried”.

Pricing and promotions

Pricing is a small deterministic function, not a merchandising system:

public static int multiplier(long occupied, long total, long days) {
  long ratio = occupied * 100 / Math.max(1, total);
  return Math.min(
      15000, (ratio >= 80 ? 15000 : ratio >= 50 ? 12500 : 10000) + (days <= 7 ? 500 : 0));
}

public static Quote quote(int base, int multiplier, int discountBps) {
  long gross = (Math.multiplyExact((long) base, multiplier) + 5000) / 10000;
  long discount = gross * discountBps / 10000;
  return new Quote(Math.toIntExact(gross - discount), Math.toIntExact(discount));
}

Multipliers are basis points, so 10,000 is 1.0×. A section under half sold is 1.0×, half sold is 1.25×, 80% or more is 1.5×, with an extra 0.05× within a week of the event and a hard cap at 1.5×. Everything is integer arithmetic, and Math.multiplyExact throws on overflow instead of wrapping, with explicit half-up rounding, so no floating-point money ever appears.

The promotion cap (“first 500 orders”) is enforced by locking the promotion row, checking used < max_uses, and inserting a (code, user_id) row whose composite primary key makes the per-user limit a constraint and the global cap a lock. That’s the simple, honest answer, and it’s also a hotspot: one hot promotion serializes every checkout that uses it. Pre-allocating tokens would relieve that, but it needs its own reconciliation story, and I’d want measurements before building it.

Principal-level talking point: snapshots and freezing. The scarcity read (how full is this section?) uses a normal committed snapshot, not a serializable one, so two buyers in the same instant can see the same occupancy. What is serialized is ownership of each seat. That’s why browsing shows a base price, and the frozen quote returned by the hold is the only price that counts. Decide which numbers must be exact and which are allowed to be advisory, and be explicit about the difference.

Payments: the phantom charge

Now the third failure. Asha holds her seats and taps Pay. The payment provider call can’t be part of a database transaction, and it can time out after the provider has acted. The whole design here is about that ambiguity.

%%{init: {"themeVariables": {"actorBkg": "#FFFFFF", "actorBorder": "#0E1A2B", "signalColor": "#22324A", "noteBkgColor": "#FFFFFF", "noteBorderColor": "#0E1A2B", "labelBoxBkgColor": "#B02F7E", "labelBoxBorderColor": "#B02F7E", "labelTextColor": "#FFFFFF", "loopTextColor": "#0E1A2B", "sequenceNumberColor": "#FFFFFF"}}}%%
sequenceDiagram
    autonumber
    box rgb(232,238,253) Outside
        actor Fan
        participant PP as Payment provider
    end
    box rgb(251,240,217) Buy plane
        participant BK as Booking service
        participant FW as Fulfillment worker
    end
    box rgb(238,232,250) Authoritative
        participant DB as System of record
    end
    box rgb(250,229,240) Async
        participant EB as Event backbone
    end
    rect rgb(232,238,253)
        Note over Fan,EB: 1 · Create the intent without ever double-charging
        Fan->>BK: Pay, with an idempotency key
        BK->>DB: Persist a stable payment ID and an initialize event, commit
        BK->>PP: Create the intent, payment ID as the provider idempotency key
        PP-->>BK: Provider reference and client secret
        BK->>DB: Save the reference and the response
        BK-->>Fan: Client secret for the hosted checkout
    end
    rect rgb(251,240,217)
        Note over Fan,EB: 2 · Only the webhook moves money
        Fan->>PP: Complete the payment, 3-D Secure or OTP
        PP->>BK: Signed webhook
        BK->>BK: Verify signature and timestamp
        BK->>DB: Lock reservation and payment, verify reference, amount, currency
        alt still pending and before the deadline
            BK->>DB: PAID, seats SOLD, order paid event to the outbox, COMMIT
        else expired or cancelled
            BK->>DB: REFUND_REQUIRED, COMMIT
        end
        BK-->>PP: 200
    end
    rect rgb(250,229,240)
        Note over Fan,EB: 3 · Everything after the commit is retryable
        DB-->>EB: Relay publishes outbox rows in order
        EB-->>FW: order paid, or refund required
        FW->>DB: Inbox row and one ticket per seat, or a refund with a stable key
    end

The first step is a small transaction that persists a stable internal payment ID and an outbox event before the provider is contacted. The provider call then happens outside any transaction, using that ID as the provider-side idempotency key, and a second transaction stores the result. If the process dies in between, the durable event lets a background worker retry with the same key, so the provider hands back the same intent rather than creating a second charge. A webhook that arrives before the provider reference has been saved gets a retryable 503 and comes back later.

The webhook handler is where money moves, so it is deliberately paranoid:

private WebhookResult applySuccessfulPayment(SuccessfulPayment event) {
  bookings.lockWebhookEvent(event.eventId());
  var existing = bookings.findWebhookFingerprint(settings.provider, event.eventId());
  if (existing.isPresent()) {
    if (!existing.get().equals(event.fingerprint())) {
      throw Problem.conflict("Webhook event ID has different content");
    }
    return new WebhookResult("duplicate");
  }

  Reservation reservation =
      bookings
          .lockReservation(event.reservationId())
          .orElseThrow(() -> Problem.conflict("Unknown payment"));
  Payment payment =
      bookings
          .lockPaymentByReservation(event.reservationId())
          .orElseThrow(() -> Problem.conflict("Unknown payment"));
  verifySuccessfulPayment(event, reservation, payment);
  bookings.insertProcessedWebhook(settings.provider, event.eventId(), event.fingerprint());

  if (List.of(
          PaymentStatus.SUCCEEDED, PaymentStatus.REFUND_REQUIRED, PaymentStatus.REFUNDED)
      .contains(payment.status())) {
    return new WebhookResult("duplicate");
  }
  if (reservation.status() != ReservationStatus.PENDING_PAYMENT
      || !reservation.expiresAt().isAfter(Instant.now())) {
    return requireRefund(reservation, payment);
  }
  completeSale(reservation, payment);
  return new WebhookResult("paid");
}

In order: the signature and timestamp are verified before this function is reached. Then the event ID is deduplicated with a content fingerprint, so a redelivered event is a no-op but the same ID carrying different content is a conflict. Then the reservation and payment rows are locked, and provider reference, amount and currency must all match. Only then does it decide the outcome, and there are exactly two: completeSale marks the reservation PAID and the seats SOLD and writes an order.paid outbox row, or requireRefund records a durable refund obligation.

That second branch answers “what if Asha pays at minute 10:01?” A late success never resurrects the hold, because that seat may already belong to someone else. It becomes a payment.refund_required event, and a worker refunds through the provider with a stable idempotency key. The test suite covers a late payment, a wrong amount (which must never sell), and eight identical webhooks fired in parallel: exactly one PAID transition, one set of tickets, and one notification.

Principal-level talking points for payments

  • You can’t make a network call atomic with a database write. So write your intent first, make the call idempotent, and make recovery a background job that doesn’t depend on the customer’s browser surviving.
  • The webhook is the source of truth; the browser is a courtesy. The customer may close the tab, lose signal or never come back.
  • Design the failure branch as carefully as the happy one. “Paid too late” isn’t an edge case on a flash sale; it’s a Tuesday.
  • Deduplicate on content, not just on ID. The same ID with different content is a red flag, not a retry.

After the money moves

Nothing after the commit is allowed to be fire-and-forget. The rule: the domain change and its event commit in one transaction (the outbox table from earlier), and a separate relay process publishes the events afterwards. That’s the transactional outbox, and it exists because “write to the database and publish to a broker” is two systems that can’t share a transaction.

private int relayBatch(KafkaProducer<String, String> producer) {
  if (!events.tryAcquireOutboxRelayLock()) return 0;
  List<OutboxEvent> batch = events.lockUnpublishedEvents(BATCH_SIZE);
  for (OutboxEvent event : batch) {
    applyDerivedEffect(event);
    publishToKafka(producer, event);
    events.markPublished(event.id());
  }
  return batch.size();
}

The relay takes an advisory lock so only one instance runs at a time, reads unpublished rows in ID order, publishes each with idempotence on and full acknowledgement, waits for the acknowledgement, and only then marks the row published. A crash between publish and mark means the event goes out twice, and that’s by design: the guarantee is at-least-once transport plus idempotent effects, not exactly-once transport. The fulfillment consumer writes an inbox row and issues tickets with a unique constraint per reservation item, so a duplicate delivery is a no-op, and ticket creation and the dedupe record commit together.

The same relay also drives the live seat map. When it processes an inventory.changed event, it invalidates the cache and then publishes to the section’s channel:

private void invalidateSection(OutboxEvent event) {
  UUID eventId = requiredUuid(event, "event_id");
  String section = requiredText(event, "section");
  cache.redis.del("map:" + eventId + ":" + section, "map:" + eventId + ":*");
  cache.redis.publish(
      "inventory:" + eventId + ":" + section,
      Json.write(
          Map.of("type", "invalidate", "event_id", eventId, "section", section)));
}

The order is deliberate: delete the cache first, then publish. A browser that reacts to the invalidation instantly can’t read the entry that was just declared stale. It also means the entire live-map pipeline is driven by committed transactions, not by a side channel that could run ahead of the truth.

Notifications go through RabbitMQ with publisher confirms, manual acknowledgements, retry queues at 5, 10, 20 and 40 seconds, and then a dead-letter queue. Financial events get the opposite treatment: a failed one stops progress and stays retryable, rather than being parked in a dead-letter queue nobody watches.

Principal-level talking points for messaging

  • The outbox turns a distributed transaction into two local ones with a durable bridge between them.
  • At-least-once plus idempotent consumers is the honest guarantee. “Exactly-once” is a property of your effects, and you build it.
  • A single ordered relay is a deliberate trade-off. It gives strict ordering, and it’s also head-of-line blocking if a broker is down, so the age of the oldest unpublished event is one of the most important numbers in the system.
  • Different failures deserve different queues. A lost email is a retry. A lost payment event is an incident.

Expiry: two things race for the same row

A hold must return to inventory when it isn’t paid for, and that reclamation races against a payment that might be confirming at the same moment. The expiry worker runs once a second, using the one place I do use SKIP LOCKED:

"SELECT * FROM reservations WHERE status='PENDING_PAYMENT' AND expires_at <= "
    + "clock_timestamp() ORDER BY expires_at FOR UPDATE SKIP LOCKED LIMIT ?"

It’s the right use because any expired reservation is fair game: several workers can reclaim different batches without waiting on each other, or on a row that a webhook is holding. The rule that makes it safe is the same one the webhook follows: both paths lock the reservation row first. Whoever gets the lock decides the terminal state.

%%{init: {"themeVariables": {"actorBkg": "#FFFFFF", "actorBorder": "#0E1A2B", "signalColor": "#22324A", "noteBkgColor": "#FFFFFF", "noteBorderColor": "#0E1A2B", "labelBoxBkgColor": "#6B46C1", "labelBoxBorderColor": "#6B46C1", "labelTextColor": "#FFFFFF", "loopTextColor": "#0E1A2B", "sequenceNumberColor": "#FFFFFF"}}}%%
sequenceDiagram
    autonumber
    box rgb(251,240,217) Buy plane
        participant WH as Webhook handler
        participant EW as Expiry worker
    end
    box rgb(238,232,250) Authoritative
        participant DB as System of record
    end
    Note over WH,EW: Both paths must lock the reservation row first
    par
        WH->>DB: Lock the reservation row
    and
        EW->>DB: Select expired rows, skip any that are locked
    end
    alt the webhook got the lock first
        rect rgb(226,244,234)
            DB-->>WH: Locked
            WH->>DB: PAID, or REFUND_REQUIRED if past the deadline
            DB-->>EW: Row skipped, nothing to do
        end
    else expiry got the lock first
        rect rgb(252,231,227)
            DB-->>EW: Locked
            EW->>DB: EXPIRED, release the seats and the promotion use
            DB-->>WH: The webhook now sees EXPIRED
            WH->>DB: REFUND_REQUIRED
        end
    end

There is no order in which a seat is both sold and released. Releasing a reservation also gives back its promotion use and emits an event, after which the relay does owner-checked cleanup of the cache hints. A backed-up relay can leave a seat looking unavailable in the cache for a while. It can never leave one sold twice.

Principal-level talking point: lock the same thing first. Two concurrent paths that can each end a reservation are safe if and only if they contend on the same lock before deciding. That single discipline replaces a great deal of clever conflict resolution.

Getting in: tickets and venue entry

A ticket’s QR code isn’t a static image. Its payload is an HMAC over the ticket ID and a 15-second time window, keyed with a per-ticket secret stored encrypted:

public static String qr(byte[] secret, String ticket, long window) {
  return ticket
      + "."
      + window
      + "."
      + HexFormat.of()
          .formatHex(hmac(secret, (ticket + ":" + window).getBytes(StandardCharsets.UTF_8)));
}

public static boolean validQr(byte[] secret, String payload) {
  try {
    String[] p = payload.split("\\.");
    long w = Long.parseLong(p[1]);
    return p.length == 3
        && Math.abs(Instant.now().getEpochSecond() / 15 - w) <= 1
        && equal(qr(secret, p[0], w), payload);
  } catch (Exception e) {
    return false;
  }
}

A screenshot goes stale within seconds. Note the <= 1: the verifier accepts the current window and one either side to tolerate clock drift between phone and scanner, so a code is really good for up to 30 to 45 seconds, not 15.

For entry, the venue runs a small standalone service backed by SQLite in WAL mode. Before the event, an encrypted, expiring manifest of paid tickets is exported into it. A scan is then a single INSERT into an entries table keyed by ticket ID, and a replay at a second gate hits a uniqueness error and is rejected as already used. It’s the same idea as everywhere else in this system: a uniqueness constraint does the arbitration, not application logic.

Principal-level talking point: the CAP corner you’re actually in. Two fully disconnected gates can’t both stay available and globally reject the same ticket’s second use, because rejecting it requires knowing about the first, which requires talking. So I chose consistency over availability at the venue: every gate shares one authoritative ledger on the venue LAN, and if that ledger or the LAN is down, gates fail closed. The alternative, independent offline scanners, is more available and permits a double entry in the window before they sync. Which one is right is a business decision, and the honest thing is to make it explicitly and write it down rather than pretend the trade-off isn’t there.

Running it: SLIs, SLOs, SLAs and tags for SRE

Building it is half the job; the other half is being able to say, at 9:04 p.m., whether it’s healthy. Here is how I’d instrument and operate a system like this. The services already expose Prometheus-style metrics endpoints; the targets below are the ones I’d build on top of them. They’re starting points to tune against load tests and real sales, not measured results.

SLI, SLO, SLA: three different things

  • An SLI (indicator) is a measurement: the fraction of good events out of valid events.
  • An SLO (objective) is the internal target for that measurement, over a window.
  • An SLA (agreement) is the external promise, with consequences, and it should be weaker than the SLO so you have room to be wrong.

The subtle craft is defining good. For example, a hold that returns 409 Conflict because someone else got the seat first is a success, not an error. It’s the system doing exactly what it should. If you count conflicts as failures, your dashboards go red at precisely the moment the system is behaving perfectly, and you train your team to ignore red.

The SLIs and SLOs I’d start with

JourneySLI (good ÷ valid)Starting SLOWhy
Join the queueJoins answered with a position, not an error99.9% during sale windowsThe front door; a failure here is invisible to the database
Browse the seat mapSnapshot requests not answered 5xx99.9%Forgiving content; stale is fine, absent is not
Seat-map freshnessInventory changes visible to viewers within 2 s99%Measured by outbox age; the brief asked for far tighter, unmeasured
Reserve seatsRequests ending 201, 409 or 422 (not 5xx or timeout)99.9%Conflicts count as good
Reserve latencyp99 of successful holdsUnder 1 sA lock timeout of 2 s means a p99 above that is a smell
Start a paymentPayment initializations succeeding99.9%Idempotent, so retries are safe
Webhook handlingValid webhooks acknowledged within the provider’s timeout99.9%A slow webhook is how paid fans get refunded
Ticket fulfillmentPaid orders with tickets issued within 60 s99.9%The promise the fan actually cares about
Hold reclamationExpired holds released within 10 s of the deadline99%Cadence is a second; batch is 100
NotificationsDelivered within 5 minutes99%Retry queues absorb the rest

Correctness SLIs have a zero-tolerance target, not a percentage. These are checks I’d run continuously and alert on the first violation:

  • A seat with more than one live owner. The constraint should make this impossible, so any constraint violation in the logs is a page.
  • A PAID reservation with no tickets after five minutes.
  • A payment stuck in INITIALIZING for more than two minutes.
  • A REFUND_REQUIRED payment older than 15 minutes.
  • Anything in a dead-letter queue on the financial path.

The SLA I’d actually publish is simple and human: “Your tickets appear in your account within 15 minutes of payment, or your money is returned automatically.” That’s much weaker than the 60-second fulfillment SLO on purpose. The gap between them is the room to be wrong.

Error budgets, and why a flash sale is different

An SLO of 99.9% over 28 days gives an error budget of about 40 minutes (28 days is 40,320 minutes, and 0.1% of that is 40.3). At 99.95% it’s about 20 minutes; at 99.99% it’s about 4 minutes, which is what “four nines” really means in practice.

But a flash sale breaks the usual model. Almost all the risk is concentrated in a 30-minute window on a few days a year, so a 28-day rolling average smooths away exactly the period that matters. I’d therefore scope budgets per sale window and count them by requests, not by minutes: during the on-sale window, no more than 0.1% of requests may fail. And I’d tighten alerting windows during the sale, for example checking a 15-minute and a 2-minute window rather than the 6-hour and 30-minute pair a steady-state service would use.

For steady-state alerting I’d use the standard multi-window, multi-burn-rate pattern: page when the budget is burning at 14.4× over one hour (and 5 minutes), page again at 6× over six hours (and 30 minutes), and open a ticket at 1× over three days. Alerting on burn rate rather than raw error percentage means you’re paged for problems that threaten the objective and left alone for blips that don’t.

Tags that make the data usable

Metrics are only as good as their labels. The set I’d standardize on, and the one I’d refuse:

LabelValuesPurpose
planeadmit, browse, buy, stateWhich of the three planes owns this signal
journeyjoin, browse, reserve, pay, fulfil, enterTies every metric to a customer-visible flow
tiert0-money, t1-purchase, t2-browse, t3-commsDrives alert routing and severity
sloe.g. reserve-availabilityLets a dashboard and an alert refer to the same objective
sale_phasepre, on_sale, postApplies the right thresholds and budgets
outcomesuccess, conflict, rejected, errorKeeps conflicts out of the error rate
event_idOnly for the current top few eventsPer-sale comparison without exploding cardinality

Never label with a user ID, reservation ID or seat ID. Metrics are for aggregates; per-request detail belongs in traces and logs, joined to metrics by exemplars or trace IDs. Routing follows the tier: t0-money pages a human immediately at any hour, t1 pages during a sale, t2 opens a ticket, and t3 goes to a queue.

What I’d watch, per component

  • Waiting room: queue depth, admission rate, active-lease count against capacity, stale-heartbeat cleanups per tick, and time from join to admit.
  • Fan-out: viewer count, upstream subscription count, and the slow-client disconnect counter that the hub already keeps.
  • Read model: cache hit ratio, fill-lease contention, and pool utilization on the read-only connection.
  • Booking service: hold latency and outcome mix, lock-wait time, database pool active versus pending, and request-thread saturation.
  • Coordination store: memory, evictions, command latency, and subscriber counts.
  • Outbox and messaging: oldest unpublished outbox age, consumer lag, retry-queue depth, and dead-letter depth.

Load-shedding, in order

When the system is overloaded, what you refuse first should be a decision made in advance, not improvised at 9:03 p.m. Mine is a ladder, from least to most painful:

  1. Serve slightly staler seat snapshots by extending the cache lifetime.
  2. Pause live push updates and let browsers poll.
  3. Shrink the admission batch, then the capacity. This is the dial from the design overview: one change that slows everything downstream.
  4. Stop admitting entirely. The line holds, and fans just wait.
  5. Never shed payment webhooks, refunds or fulfillment. Those are the money path.

Before any sale, I’d run a readiness drill: kill the coordination store and confirm that the queue and holds fail closed; stop the booking service and confirm browsing and queueing carry on (the isolation script in the repo does exactly this); stop the outbox relay and confirm the freshness alert fires. An SLO you’ve never seen page is a guess.

Epilogue

At 9:00:01, Asha sees “You are number 4,812”, a live number that ticks down. She’s never in the stampede, because the stampede is happening in a lightweight tier with no database behind it. At 9:07 she’s admitted, picks 14F and 14G, and gets a price that won’t change under her. The database, not the cache, is what says yes. She pays; the provider times out mid-call; the same intent comes back on retry; the webhook arrives; her tickets appear. If someone else had beaten her to 14F, she’d have got a clean, instant “those seats are taken”, not a phantom charge.

Four ideas did most of the work. Let the database own the invariant and treat every faster layer as advice. Bound the concurrency in front of your scarcest resource, and let everything else wait cheaply. Make every side effect idempotent, because delivery is always at-least-once. And separate load by what it holds scarce, so a flood in one place can’t starve another. Everything else, the Lua scripts, the lock modes, the outbox, the burn-rate alerts, is those four ideas applied one layer at a time.