Flash-Sale Ticket Booking System
It is 8:59 p.m., and Asha has been staring at a countdown for ten minutes. At 9:00 sharp, a hundred thousand seats for a two-night concert go on sale, and so does everyone she knows. I made Asha up, but every engineer who has run a sale like this will recognise the three things that happen next.
First, the stampede. A million people press refresh in the same minute. The database connection pool empties, the pages turn into error screens, and the retry logic in every browser and app turns a bad minute into a bad hour.
Second, the double sale. Two fans are told, within the same second, that seat 14F is theirs. One of them gets an apology email at midnight and a story they will tell for years.
Third, the phantom charge. Asha’s bank shows the money gone. Her screen shows a spinner. The tickets never arrive, because the call to the payment provider timed out after the provider had already said yes.
These are three different kinds of failure: overload, concurrency, and ambiguity. None of them is fixed by adding servers. Each needs a specific, named defense, and the defenses have to work together. This article is how I designed one system to have all three. It’s called Encore, a flash-sale platform I built as ten cooperating services, and Asha will come back as we go.
I approached it with a narrower question than “can I design it?”: which of the usual claims can I actually prove, and which am I just asserting? The classic brief reads like a wish list: never sell a seat twice, 100,000 reads a second, 10,000 writes a second, verify a token in under a millisecond, survive a payment provider that drops mid-charge, keep bots out, and let a fan through a stadium gate with no signal. Every line is easy to write and expensive to build. So this walkthrough is opinionated about why, not just what, and the table in the next section keeps me honest about what I can back up.
What we are building
Functional requirements
- Waiting room. When checkout is at capacity, users queue first-in-first-out, see their position live, and are admitted in bounded batches with a short-lived signed pass.
- Live seat map. Browse sections and seats, and watch other people’s holds and sales as they happen.
- Atomic hold. Hold one to six seats all-or-nothing, for ten minutes.
- Reclamation. A hold that isn’t paid for returns to inventory by itself.
- Pricing and promotions. Scarcity-based pricing, capped promotion codes, and a price frozen at hold time.
- Payment. Start a payment with an external provider, and finalize the order only from a signed webhook.
- Fulfillment. After payment, issue tickets and notify, asynchronously.
- Entry. Rotating QR codes, and gate validation that keeps working when the venue’s internet doesn’t.
Non-functional requirements. The middle column is what the classic brief asks for. The right column is what I can actually show.
| Requirement | Target in the brief | What I can show |
|---|---|---|
| Inventory integrity | Zero double-booking | Enforced by a database constraint; tested with 12 concurrent buyers |
| Idempotency | Every state-changing endpoint | Reserve, pay and cancel carry stored responses; webhooks are deduplicated by event ID |
| Read throughput | 100k RPS | Not demonstrated. A smoke script exists; it is not a capacity test |
| Write throughput | 10k RPS | Not demonstrated |
| Token verification | Under 1 ms, no network I/O | Not met. The gateway makes a private auth call, which is a network hop |
| Seat-update fan-out | Under 50 ms | Not measured |
| Idle connections | 10k–25k per instance | Connection ownership proven; memory per connection not measured |
| Seat map payload | About 12.5 KB for 50,000 seats | Met. Exactly 12,500 bytes, asserted in a test |
| Availability and durability | 99.99%, RPO 0 | Not demonstrated. One host, single instances |
The rules that can never break
Before any boxes and arrows, I write down the properties that must hold regardless of load, retries, or which process crashed. Everything later in this article exists to protect one of these:
- A seat has at most one live owner. This is enforced by a unique index, not by application code.
- Fast layers may delay or reject, but never grant. Losing every cache key must not double-book a seat.
- A hold is all-or-nothing. All six seats, or none.
- Money moves only on a verified webhook. Signature, timestamp, provider reference, amount and currency must all match.
- A reservation ends exactly one way: paid, expired or cancelled. Every path that changes it locks the same row first.
- The price quoted at hold time is the price charged.
- Duplicate delivery is harmless. Idempotency keys, webhook fingerprints, an inbox table and a unique ticket per seat absorb every retry.
- An idle connection never holds a booking thread or a database connection.
The design on one page
Here is the whole system with no product names on it. Colors mark the plane a component belongs to; the moving dots are traffic.
The most important decision in that picture isn’t any single box. I separate the system by what each part holds scarce, not by feature. The admit and browse planes hold connections: memory, sockets, file descriptors. Browse also absorbs read bursts. The buy plane holds locks and database connections, the one resource that can’t be scaled by adding instances. If those three share a process, a flood in one starves the others, and the failure you get is the stampede from the opening.
The separation buys three specific guarantees. An idle socket never occupies a booking thread or a database connection. The components that face the public and hold sockets have no database credentials at all. And each plane can fail, restart, and scale independently. I wanted to test that last claim rather than assert it, so there’s a drill in the repo that stops the buy plane while clients stay connected and checks that browsing and queueing keep working.
There’s also a single dial. The admit plane decides how many people are inside checkout at once, and that one number is the lever that protects everything downstream. We’ll come back to it twice: once in the waiting-room design, and once when I talk about operating the system.
Choosing the technology
The design above is deliberately technology-agnostic. Now I’ll pick technologies, and explain each pick by the role it has to play, because the decisive boundary is where scarce resources are held, not which language is fashionable. A language choice alone proves nothing about throughput. It only makes some designs easier to express.
| Role | My pick | Why this, for this role | What I gave up, and what would change my mind |
|---|---|---|---|
| Edge gateway | NGINX | One public port, cheap idle connections, request and time limits, WebSocket upgrade, unbuffered streaming, and a private auth call to verify the pass | An in-process, sub-millisecond token check. An ingress JWT module or a mesh proxy would change my mind |
| Waiting room and admission | Node.js | The work is I/O-bound: many idle sockets, tiny messages, plus signing tokens. The event loop is a good fit. No database credentials, ever | CPU-heavy work would stall the event loop, so none goes there. Go would be a fair swap; the boundary matters more than the language |
| Read model and live fan-out | Go | Cheap concurrency for streaming and a bounded, per-client channel is idiomatic, which is the whole backpressure story. Small footprint per connection | Another toolchain to build and patch |
| Booking service and workers | Java 21, Spring Boot, plain JDBC | Mature transactions, bounded pools, graceful shutdown, health and metrics. I kept SQL explicit instead of using an ORM so the lock behavior is visible in code review | Heavier memory. That’s fine, because it holds no idle sockets by design |
| System of record | PostgreSQL | Partial unique indexes and check constraints turn invariants into constraints. It has row locks with NOWAIT and SKIP LOCKED, advisory locks, transactional outbox writes, and database time | A single writer per event. If the hottest event outgrows one writer, I’d move to one authoritative writer per section, not to a weaker consistency model |
| Coordination store | Redis | Atomic scripts over several keys, sorted sets for order plus leases plus heartbeats, TTLs, pub/sub, and sub-millisecond latency | Durability. That’s why everything in it is treated as disposable and every consumer fails closed |
| Financial event log | Kafka | Durable, ordered per key, replayable; offsets advance only after processing succeeds | Operational weight |
| Notification queue | RabbitMQ | Per-message acknowledgement, delayed retry queues and a dead-letter queue. That’s a work queue semantic, which a log doesn’t give you natively | A second broker to run. Different delivery semantics justify it |
| Payments | Hosted provider SDK, idempotency keys, signed webhooks | No card data ever touches my servers, which shrinks compliance scope | Provider lock-in |
| Venue ledger | SQLite in WAL mode behind a tiny LAN service | A primary key on the ticket ID makes double-entry atomic with no cloud dependency | High availability. A consensus-backed ledger would be the upgrade |
Two of those picks deserve a paragraph, because they’re the ones people argue about.
Why Redis is a bouncer and not the owner. The textbook flash-sale design puts the seat lock in memory: fast, simple, sub-10 ms. I deliberately did not. If the cache is the source of truth for a hold, then a key expiring at the wrong instant, a failover, or a crash between “write to the cache” and “write to the database” can leave the two disagreeing about who owns seat 14F, and one of them has already taken someone’s money. So ownership lives in PostgreSQL, and Redis only turns obvious contention into a cheap rejection. Every Redis failure then degrades into a slower or rejected request, never a wrong sale.
Why Kafka and RabbitMQ. They aren’t redundant. Financial events need an ordered, replayable log where a failed event stops progress and stays retryable rather than being parked in a queue nobody watches. Notifications need the opposite: independent messages, per-message retry with growing delays, and a dead-letter queue for the ones that keep failing. Using one broker for both means forcing one set of semantics onto a workload that wants the other.
The data model
PostgreSQL holds everything that must be right: events, seats, reservations, per-seat price snapshots, promotions, payments, tickets, and the plumbing for reliable messaging. Three constraints carry the whole “never double-book” promise, exactly as they appear in the schema:
CREATE UNIQUE INDEX one_active_checkout
ON reservations (user_id, event_id)
WHERE status = 'PENDING_PAYMENT';
ALTER TABLE seats
ADD CONSTRAINT seats_owner_state
CHECK (
(status IN ('AVAILABLE', 'BLOCKED') AND reservation_id IS NULL)
OR (status IN ('HELD', 'SOLD') AND reservation_id IS NOT NULL)
);
CREATE UNIQUE INDEX one_live_owner_per_seat
ON reservation_items (seat_id)
WHERE status IN ('HELD', 'SOLD', 'CHECKED_IN');
one_live_owner_per_seat is invariant number one. It’s a partial unique index: released and expired items fall out of it, so a seat can be held again later, but two live items can never point at the same seat. seats_owner_state means a seat can’t be HELD without naming who holds it. one_active_checkout stops one user from hoarding by opening a dozen pending checkouts.
Principal-level talking point: constraints beat code. An application-level check (“is this seat free?”) is a promise that every code path, every future engineer and every retry will remember to make. A constraint is a promise the database makes on their behalf. Design so that the worst bug you can write fails loudly at the constraint instead of silently double-selling.
Prices are frozen per item (base_price, surge_bps, discount, price_at_lock), so a later price change can’t alter what someone already agreed to pay. The outbox table is the other important one, and it shows up again in the payments section:
CREATE TABLE outbox (
id BIGSERIAL PRIMARY KEY,
kind TEXT NOT NULL,
aggregate_id TEXT NOT NULL,
payload JSONB NOT NULL,
created_at TIMESTAMPTZ NOT NULL DEFAULT clock_timestamp(),
published_at TIMESTAMPTZ
);
CREATE INDEX outbox_unpublished
ON outbox (id)
WHERE published_at IS NULL;
The waiting room
Back to Asha. It’s 9:00:01 and a million people are on the site. This is where the stampede gets stopped, and the waiting room is the most misunderstood component in the design. It looks like a UX feature, a “you’re number 4,812” screen. It is actually a concurrency limiter for your database.
The problem it solves
A database can only do so much work at once, and the limit has nothing to do with how much traffic you receive. It is set by how many transactions can hold a connection, and how long each one holds it. That relationship is Little’s law: concurrent work equals arrival rate times time in system. In my booking service the transaction pool is 12 connections. Suppose a hold transaction keeps its connection for about 25 ms. I haven’t profiled that number, so treat it as an assumption. Then the pool can finish roughly 12 ÷ 0.025 = 480 holds a second, no matter how many servers sit in front of it.
Now suppose only one in ten of a million fans reaches a hold attempt in the first minute. That’s 100,000 attempts over 60 seconds, about 1,700 per second, against a system that can finish about 480. Nothing about the architecture is wrong. The load is simply 3.5 times what the scarce resource can carry, and what happens next is worse than “some requests are slow”.
Without a limiter the excess doesn’t wait politely. It piles up, requests time out, and clients retry, so the offered load rises just as the system is least able to cope. Goodput, meaning the purchases that actually complete, doesn’t plateau at capacity. It collapses below it:
The second thing the waiting room buys you is fairness: first in, first out is a promise you can explain to a customer, while “whoever’s retry logic happened to land on a healthy server” is not. The third is protecting everything else that shares the pool. In my design, the payment webhook is served by the same booking service as the seat holds. An unbounded stampede of hold attempts competes for the same request threads and database connections as the webhook that confirms a paying customer’s order. Ask yourself what your webhook latency looks like at 9:00:01 on a night like this.
Lock contention, up close
To see why the limit matters, look at what a hold does. Asha and 999 other fans all want seat 14F. Each one runs, in effect, a transaction that ends in SELECT ... FOR UPDATE on the seat’s row. The database grants that lock to one transaction and makes the rest wait in line.
That has consequences that aren’t obvious from a diagram:
- Throughput on one hot row is capped by how long the lock is held. If a transaction holds the seat lock for 25 ms, that row can change hands about 40 times a second, whether you have 10 servers or 1,000.
- Waiters queue behind each other. The k-th waiter waits about k × 25 ms. A thousand waiters means the last one waits 25 seconds, far longer than a client will tolerate.
- Every waiter occupies a connection. A blocked transaction does no useful work but still pins one of the 12. A handful of hot rows can hold the whole pool hostage while colder seats, which could have been sold immediately, wait for a connection.
- Deadlocks appear when carts overlap. Buyer A wants seats 5 and 9. Buyer B wants 9 and 5. If they lock in different orders, each waits for the other, and the database eventually notices and kills one of them, after its deadlock timeout, with the transaction’s work wasted.
My defense against that last one is boring and reliable: the service sorts the requested seat IDs before touching the database, so every buyer locks in the same order. Two buyers who both want 5 and 9 will both try 5 first, and neither can hold 9 while waiting for 5.
FOR UPDATE, NOWAIT and SKIP LOCKED at scale
The choice of how to wait for a row is where a lot of the real design lives. Here’s how I think about the options, and where each one breaks:
| Strategy | When the row is busy… | Good for | What breaks at flash-sale scale | Where Encore uses it |
|---|---|---|---|---|
FOR UPDATE | Waits its turn until the holder commits or rolls back | Paths where the caller must have that exact row | Waiters pin connections; the k-th waiter waits k × hold time; the tail becomes timeouts and retries | Seat locks in the hold transaction, bounded by the waiting room and a 2-second lock timeout |
FOR UPDATE NOWAIT | Fails immediately with a lock-not-available error | Interactive picks where a busy seat probably won’t stay free; protecting the pool | Retry storms if clients fire again instantly; no queue fairness; needs jittered backoff and a UX for “someone is holding those seats” | Not used. It’s the first lever I’d add at 10× scale |
FOR UPDATE SKIP LOCKED | Silently skips locked rows and returns whatever is left | Work queues where any row will do | Wrong for “I want 14F”: you’d get a different seat, or fewer than you asked for, non-deterministically | Expiry worker, where any expired reservation is fair game |
lock_timeout | Waits, but gives up after a bound | A safety net around every blocking path | Bounds the damage but doesn’t reduce contention; set it too high and the pool still drains | Set on every connection: 2 s lock timeout, 5 s statement timeout |
| Advisory lock | An application-defined mutex on a hash | Serializing something with no natural row | A single global hot key becomes a global bottleneck | Idempotency check per user and operation; the outbox relay singleton |
Optimistic update (UPDATE ... WHERE status='AVAILABLE') | Under the default isolation level, the update still waits for a competing writer, then re-checks the condition | Single-row state changes | It doesn’t remove the wait. A multi-seat cart still needs one transaction, so locks are held until commit | Not used. The version column exists for it |
The two subtle rows are the last one and SKIP LOCKED. People reach for optimistic updates hoping to escape locking, but in SQL the update itself takes a row lock and holds it to the end of the transaction, so under contention you still queue. And SKIP LOCKED feels like the answer to hot rows, but it changes the meaning of the query. That’s exactly what you want for “give me any 100 expired holds to clean up”, and exactly what you don’t want for “give me seat 14F”.
Principal-level talking point: how I’d scale the hottest event 10×. First, keep the waiting room. It bounds concurrency before any lock strategy matters. Second, change the seat locks to
NOWAIT(or a much shorterlock_timeout), with jittered client retries and a clear “someone is holding those seats” message, so that a hot seat costs a fast failure rather than a pinned connection. Third, add a best-available purchase mode (“any two adjacent seats in section B”) and implement it withSKIP LOCKED, which spreads contention across all the free seats instead of piling every buyer onto the same one. I haven’t built that mode; it is whereSKIP LOCKEDearns its place. Fourth, shard the hottest event by section, with one authoritative writer per section. Fifth, give the payment webhook its own connection pool so a stampede can never starve a paying customer.
What the backend feels without a waiting room
Here is the cascade, using the real limits from my booking service configuration: 64 request threads, a queue of 64 more connections, a pool of 12 database connections with a 3-second wait for one, and a 2-second lock timeout.
- 9:00:01. Attempts arrive at roughly 1,700 a second. The 12 pooled connections are busy within milliseconds.
- Within a second. The rest wait for a connection. The 64 request threads fill up with requests parked on the pool.
- After 3 seconds. Requests that never got a connection fail with a timeout. The client sees a 503 and, in many clients, retries at once, so the arrival rate goes up.
- Meanwhile. Connections that did get through pile onto a few hot seats and sit blocked on row locks for up to 2 seconds each. They’re occupied and doing no useful work.
- Now the queue behind Tomcat overflows. New connections are refused outright. This isn’t a slow system anymore. It’s a system that answers most callers with nothing.
- The quiet disaster. A payment webhook from a customer who paid a minute ago is stuck in that same queue. The provider retries later. The hold may expire in between, so that customer’s payment lands as a refund. Asha would be charged and refunded in a single evening.
With the waiting room, the same one million fans see a position number. Admission caps concurrent checkouts, so the transaction pool only ever sees the load it can carry. The stampede still exists. It just lives in a lightweight, connection-only tier with no database access, where holding a million idle sockets is affordable.
How admission is designed
Every aspect of the waiting room is a small decision with a reason behind it. I’ll take them in order.
Ordering: a sequence number, not a timestamp. The queue is a sorted set per event, scored by a monotonic counter that’s incremented atomically at join time. Timestamps would tie whenever two fans join in the same millisecond, and clocks disagree between servers; a counter gives strict, tie-free order. Rejoining doesn’t reset your place, so refreshing the page doesn’t send you to the back of the line.
Capacity, not rate. I don’t admit “N users per second”. I admit until N are inside checkout. Because the scarce resource is concurrency, a concurrency limit self-corrects: if payments get slow and people linger, fewer slots free up and admission slows on its own; if people finish faster, it speeds up. A pure rate limit would keep admitting at a fixed pace into a system that was already backing up. There’s also a per-tick batch cap so a burst of freed capacity can’t release a thundering herd all at once.
Two different Little’s-law budgets. The database pool is sized by transaction time: how long a hold keeps a connection. Admission capacity is sized by time in checkout: how long a fan spends browsing seats and paying. These are different numbers, often minutes versus milliseconds, and mixing them up is a classic mistake. As a worked example, to sell 25,000 orders in 30 minutes I’d need about 14 orders a second; at a mean of 90 seconds in checkout, admission capacity should be around 1,250 concurrent checkouts, while the same 14 holds a second only keep a pool connection busy for a third of a second per second. My local defaults are 100 concurrent checkouts with a batch of 10 per tick. They’re demo values, not tuned ones.
A single elected admitter. If ten admission workers each admit a batch, you’ve admitted ten batches. So workers compete for a one-second lease per event, and only the winner runs that tick. The lease also removes the need for any coordination between the workers beyond one atomic key.
Abandonment: heartbeats. People close laptops. A queue entry that outlives its owner would eventually be admitted and burn a checkout slot. Each queued fan refreshes a heartbeat every 5 seconds over the open socket, and entries silent for 120 seconds are dropped. The cleanup is capped at 1,000 per tick so a giant abandoned queue can’t monopolize the store, which creates a subtle trap the scripts below handle explicitly.
The pass: asymmetric signature plus a server-side lease. Admission produces a signed pass, valid at most five minutes and bound to the user and event. It’s signed with a private key that only the issuer holds, so everything that verifies it needs only a public key and can’t mint passes. But a signature can’t be revoked, so the pass is not sufficient by itself: the booking service also requires an active lease on the server side. A perfectly valid but stale pass can’t start a checkout.
Transport: a WebSocket, not polling. A million fans polling every few seconds is a million requests a few times a minute. A held socket with a five-second server push is far cheaper for the server and gives the fan a live number. That only works because the tier that holds the sockets has no database behind it.
Failure: fail closed. If the coordination store is lost, queue, admission and rate limits all return a retryable error. Fans rejoin the line. Nobody is granted a seat, because the store never granted seats in the first place. The cost is an ugly minute; the alternative is a bypass.
Here is the whole flow. Watch where the fan waits and where the booking service finally sees them:
%%{init: {"themeVariables": {"actorBkg": "#FFFFFF", "actorBorder": "#0E1A2B", "signalColor": "#22324A", "noteBkgColor": "#FFFFFF", "noteBorderColor": "#0E1A2B", "labelBoxBkgColor": "#0B8580", "labelBoxBorderColor": "#0B8580", "labelTextColor": "#FFFFFF", "loopTextColor": "#0E1A2B", "sequenceNumberColor": "#FFFFFF"}}}%%
sequenceDiagram
autonumber
box rgb(232,238,253) Outside
actor Fan
end
box rgb(224,244,242) Admit plane
participant WR as Waiting room
participant AC as Admission controller
end
box rgb(252,231,227) Fast, disposable
participant CS as Coordination store
end
box rgb(251,240,217) Buy plane
participant BK as Booking service
end
rect rgb(232,238,253)
Note over Fan,BK: 1 · Join the line
Fan->>WR: Join the queue with a signed session token
WR->>CS: Atomic join, next sequence number, heartbeat
WR-->>Fan: You are number 4,812
end
rect rgb(224,244,242)
Note over Fan,BK: 2 · Wait cheaply
loop every 5 seconds while queued
Fan->>WR: Socket ping with token
WR->>CS: Refresh heartbeat, read position
WR-->>Fan: Position update
end
end
rect rgb(252,231,227)
Note over Fan,BK: 3 · Admission tick, one winner per event per second
AC->>CS: Take the one-second tick lease
AC->>CS: Drop expired leases and stale members
AC->>CS: Move min(free capacity, batch) from the head of the line
end
rect rgb(251,240,217)
Note over Fan,BK: 4 · Enter checkout
WR->>CS: Is there an active lease for this fan?
WR-->>Fan: Signed pass, five minutes at most, socket closes
Fan->>BK: Reserve seats, presenting the pass
BK->>CS: Verify the active lease, fail closed on error
BK->>CS: After commit, extend the lease to the hold duration
end
Joining and admitting are each a single Lua script, so they run atomically inside the store. Joining reads the store’s own clock (so replicas can’t disagree about “now”), keeps an existing place in line, and assigns a sequence number only once:
-- KEYS: queue, active leases, sequence, heartbeat; ARGV: subject.
local clock = redis.call('TIME')
local now = tonumber(clock[1]) + tonumber(clock[2]) / 1000000
redis.call('ZREMRANGEBYSCORE', KEYS[2], '-inf', now)
if redis.call('ZSCORE', KEYS[2], ARGV[1]) then return -1 end
if not redis.call('ZSCORE', KEYS[1], ARGV[1]) then
local sequence = redis.call('INCR', KEYS[3])
redis.call('ZADD', KEYS[1], sequence, ARGV[1])
end
-- Admission removes stale queue/heartbeat members together. Do not prune only one here.
redis.call('ZADD', KEYS[4], now + 120, ARGV[1])
return redis.call('ZRANK', KEYS[1], ARGV[1]) + 1
Admission is where capacity, batch limits and the abandonment trap live together:
-- KEYS: queue, active leases, sequence, heartbeat; ARGV: capacity, batch, lease TTL.
-- Redis time avoids disagreement between admission replicas.
local clock = redis.call('TIME')
local now = tonumber(clock[1]) + tonumber(clock[2]) / 1000000
redis.call('ZREMRANGEBYSCORE', KEYS[2], '-inf', now)
-- Keep each cleanup bounded so an abandoned queue cannot monopolize Redis.
local stale = redis.call('ZRANGEBYSCORE', KEYS[4], '-inf', now, 'LIMIT', 0, 1000)
for _, subject in ipairs(stale) do
redis.call('ZREM', KEYS[1], subject)
redis.call('ZREM', KEYS[4], subject)
end
local freeSlots = tonumber(ARGV[1]) - redis.call('ZCARD', KEYS[2])
local attempts = math.min(freeSlots, tonumber(ARGV[2]))
local admitted = {}
for index = 1, attempts do
local entry = redis.call('ZPOPMIN', KEYS[1], 1)
if #entry == 0 then break end
local subject = entry[1]
local heartbeatUntil = redis.call('ZSCORE', KEYS[4], subject)
redis.call('ZREM', KEYS[4], subject)
-- More than 1,000 stale members may remain after cleanup. Never reactivate them.
if heartbeatUntil and tonumber(heartbeatUntil) > now then
redis.call('ZADD', KEYS[2], now + tonumber(ARGV[3]), subject)
table.insert(admitted, subject)
end
end
return admitted
The comment near the end is the trap I mentioned. Cleanup is capped, so abandoned users can still be at the front of the line afterward. Each candidate’s heartbeat is therefore re-checked as it’s popped; without that check, a fan who closed their laptop an hour ago would be admitted and burn a slot for five minutes. There’s a regression test for exactly that case.
Principal-level talking points for the waiting room
- It’s a concurrency limiter for the database, not a UX feature. Size it from the transaction budget and the checkout time, not from the traffic forecast.
- Admission by capacity self-corrects; admission by rate doesn’t.
- The fair alternative to first-in-first-out is a lottery for the first few minutes, which removes the advantage of fast connections and simple bots. I chose FIFO because it’s predictable and easy to explain, and I’d revisit that if bots became the problem.
- The wait estimate shown to the fan is a rough function of the batch size. It can’t know how long people ahead of you will linger in checkout, so present it as an estimate, never a promise.
- Fail closed. Every failure mode of the queue must end with “nobody got in”, never “everybody got in”.
Browsing: a million eyes, one source of truth
While Asha waits, and after she’s admitted, she watches the seat map. This is the browse plane, and its problem is the reverse of the buy plane’s: enormous read volume with very forgiving correctness. A seat map that’s a second stale is fine, because the database will refuse the hold if the seat is gone.
Reads: cache, and stop the stampede on the cache itself. Seat snapshots are served from a short-lived cache (a couple of seconds). On a miss, one request takes a fill lease, a random token with a 5-second expiry, and refills the cache while other readers wait briefly or get a retryable 503. So a thousand simultaneous misses collapse toward one database query. The lease is released only if its token still matches, so a slow request can’t delete another request’s lease. The read service also has its own database login that can only SELECT, with a pool of four connections, so even a bug there can’t write.
Payload: a bitmap. Full seat state for a section is available as two bits per seat, four seats to a byte:
// Four two-bit states fit in each byte, ordered most significant pair first.
func writeBitmap(w http.ResponseWriter, data []byte) {
var seats []Seat
if json.Unmarshal(data, &seats) != nil {
httpkit.Error(w, 503, "Invalid snapshot")
return
}
packed := make([]byte, (len(seats)+3)/4)
states := map[string]byte{"AVAILABLE": 0, "HELD": 1, "SOLD": 2, "BLOCKED": 3}
for i, seat := range seats {
packed[i/4] |= states[seat.Status] << uint(6-2*(i%4))
}
w.Header().Set("Content-Type", "application/octet-stream")
w.Header().Set("X-Seat-Count", fmt.Sprint(len(seats)))
w.Header().Set("X-Seat-Order", "section,row_num,seat_num")
_, _ = w.Write(packed)
}
Fifty thousand seats become exactly 12,500 bytes, and a test asserts that number. Seat order comes from a fixed layout, so the payload carries only state. It is also strictly advisory; nothing decides ownership from it.
Updates: push an invalidation, not a delta. When a seat changes, the browser doesn’t receive “14F is now HELD”. It receives “section B changed, go fetch the current picture”. That’s less clever and much more robust: a delta stream needs ordering and replay, while an invalidation can be lost and the next refresh repairs it. This matters because the transport is Pub/Sub, which is not durable. Every stream therefore opens with a reset telling the browser to fetch a fresh snapshot, and browsers also refresh periodically.
Fan-out: one upstream subscription per section, however many viewers. The naive design gives every browser its own subscription: 5,000 people looking at section B means 5,000 subscriptions. Mine inverts that. Each viewer is a small struct with a bounded queue, and each active section has exactly one subscription per fan-out instance:
func (h *Hub) Broadcast(topic, message string) {
h.mu.Lock()
defer h.mu.Unlock()
s := h.topics[topic]
if s == nil {
return
}
for c := range s.clients {
select {
case <-c.Done:
continue
default:
}
select {
case c.Messages <- message:
default:
c.stop()
h.Slow.Add(1)
}
}
}
The inner select with a default is the entire backpressure policy. Each viewer’s channel holds eight messages. If a phone on a bad network can’t keep up and its channel is full, it’s disconnected, not waited on, so one slow client can’t delay the other 4,999. When the last viewer leaves a section, its subscription closes.
Principal-level talking points for browse
- Push an invalidation, pull the truth. It turns a hard ordered-delivery problem into an easy eventually-correct one.
- Bound every queue. An unbounded per-client buffer is a memory leak with a slow phone’s name on it.
- Drop the slow consumer, not the fast ones. Backpressure is a policy decision; make it explicit.
- The read side reaches the database with a credential that can’t write. Least privilege is also a blast-radius decision.
Reserving seats
Asha has been admitted. She picks seats 14F and 14G. So does someone else. This is the double sale, and everything in the rules section is about it.
%%{init: {"themeVariables": {"actorBkg": "#FFFFFF", "actorBorder": "#0E1A2B", "signalColor": "#22324A", "noteBkgColor": "#FFFFFF", "noteBorderColor": "#0E1A2B", "labelBoxBkgColor": "#A8721A", "labelBoxBorderColor": "#A8721A", "labelTextColor": "#FFFFFF", "loopTextColor": "#0E1A2B", "sequenceNumberColor": "#FFFFFF"}}}%%
sequenceDiagram
autonumber
box rgb(232,238,253) Outside
actor Fan
end
box rgb(251,240,217) Buy plane
participant BK as Booking service
end
box rgb(252,231,227) Fast, disposable
participant CS as Coordination store
end
box rgb(238,232,250) Authoritative
participant DB as System of record
end
Fan->>BK: Reserve 1 to 6 seats, pass, idempotency key
BK->>CS: Verify the active lease and the rate limit
BK->>DB: BEGIN, take the idempotency lock
alt identical retry
DB-->>BK: Stored response
BK-->>Fan: Original response, nothing re-executed
else new request
rect rgb(252,231,227)
Note over BK,DB: 1 · Cheap rejection first
BK->>DB: Any pending checkout for this user?
BK->>CS: Hint all the seats atomically, all or none
end
rect rgb(238,232,250)
Note over BK,DB: 2 · Lock and verify
BK->>DB: Lock the seat rows in sorted order
BK->>DB: Is every seat AVAILABLE?
BK->>DB: Lock the promotion row, check its caps
end
rect rgb(251,240,217)
Note over BK,DB: 3 · Price and write
BK->>DB: Freeze prices, insert reservation and items
BK->>DB: Seats to HELD, versions up
BK->>DB: Outbox row: inventory changed
end
rect rgb(226,244,234)
Note over BK,DB: 4 · Commit
BK->>DB: Store the response and COMMIT
BK->>CS: Extend the admission lease
BK-->>Fan: 201 with frozen total and expiry
end
end
opt any step fails
BK->>DB: ROLLBACK
BK->>CS: Owner-checked release of seat hints
BK-->>Fan: 409 conflict or a retryable 503
end
The code, unabridged:
private HoldResult createHold(
String userId, UUID reservationId, ReserveCommand command, String idempotencyKey) {
HoldResult previous =
idempotency.lock(userId, "reserve", idempotencyKey, command, HoldResult.class);
if (previous != null) return previous;
cache.requireActive(command.eventId().toString(), userId);
requireNoPendingCheckout(userId, command.eventId());
Event event = requireOpenEvent(command.eventId());
cache.hold(
command.eventId().toString(),
command.seatIds().stream().map(UUID::toString).toList(),
reservationId.toString());
List<Seat> seats = lockAvailableSeats(command.eventId(), command.seatIds());
int discountBps = lockPromotion(command.coupon(), userId);
List<LockedPrice> prices = freezePrices(event, seats, discountBps);
int total = prices.stream().mapToInt(LockedPrice::priceAtLock).reduce(0, Math::addExact);
var expiresAt = bookings.databaseExpiry(settings.holdSeconds);
bookings.insertReservation(
reservationId,
userId,
command.eventId(),
expiresAt,
total,
event.currency(),
command.coupon());
consumePromotion(command.coupon(), userId, reservationId);
prices.forEach(item -> bookings.insertHeldItem(reservationId, item));
publishInventoryChanges(command.eventId(), seats);
HoldResult result =
new HoldResult(
reservationId,
command.eventId(),
ReservationStatus.PENDING_PAYMENT,
expiresAt,
total,
event.currency(),
prices);
idempotency.save(userId, "reserve", idempotencyKey, command, result);
return result;
}
Reading top to bottom, the decisions worth defending:
Idempotency comes first. Networks retry and people double-tap. The service takes an advisory lock scoped to the user and operation, so concurrent retries queue behind each other, then looks up the stored fingerprint for the key. An identical retry gets the original response back, and the same key with a different request is rejected as a conflict. The stored response is written in the same transaction as the reservation, so a crash can’t leave a reservation with no recorded answer, or an answer with no reservation.
Cheap checks before expensive ones. Rate limits, the admission lease, the “one active checkout per user” rule and the seat hints are all cheap, and they run before any seat row is locked. The cheapest way to handle a doomed request is to refuse it before it costs a lock.
Deterministic lock order. The sort I described earlier happens before this function runs, and this is the loop it enables:
public List<Seat> lockAvailableSeats(UUID eventId, List<UUID> seatIds) {
return seatIds.stream()
.map(
seatId ->
jdbc.query(
"SELECT * FROM seats WHERE id=? AND event_id=? FOR UPDATE",
rs -> rs.next() ? mapSeat(rs) : null,
seatId,
eventId))
.toList();
}
Then the code checks that every seat is AVAILABLE and fails the whole cart if any isn’t, so a hold is genuinely all-or-nothing.
The fast layer is a bouncer. Note that the hint step sits before the locks. It’s one atomic script across the whole cart:
-- Check the entire cart before writing any hint. SQL remains the ownership authority.
-- KEYS: seat hints in one event hash slot; ARGV: reservation ID, TTL seconds.
for _, key in ipairs(KEYS) do
local owner = redis.call('GET', key)
if owner and owner ~= ARGV[1] then return 0 end
end
for _, key in ipairs(KEYS) do
redis.call('SET', key, ARGV[1], 'EX', ARGV[2])
end
return 1
Its only job is to turn obvious contention into a cheap rejection: if someone has already hinted seat 14G, this buyer is bounced without taking a database lock. The release script is its mirror image, and it compares the reservation ID, not just the user:
-- An old release must never delete a hint acquired by a newer reservation.
for _, key in ipairs(KEYS) do
if redis.call('GET', key) == ARGV[1] then
redis.call('DEL', key)
end
end
return 1
That matters because cleanup is asynchronous: a delayed release for an expired reservation must never delete the hint of the next buyer who took the same seat.
Proving the cache can’t cause a double sale. I tested it directly. A test holds a seat, then deletes its hint from the cache by hand, then lets a second buyer request that seat plus another:
cache.redis.delete('hold:{' + event_id + '}:' + seats[0])
_, other = admitted(client, event_id)
assert hold(client, event_id, seats[:2], other).status_code == 409
The database rejects the second buyer with a 409, and the other seat in their cart stays AVAILABLE, which shows the cart rolled back as a unit. A second test admits 12 buyers, has them all request the same two seats from 12 threads at once, and asserts exactly one 201 and eleven 409s, with exactly two HELD items in the table afterward. Twelve is a small number, so this shows correctness under a modest race rather than behavior at scale, but it’s the invariant working end to end.
The lock modes I chose, and why. Seat locks are plain blocking FOR UPDATE, because after admission and the fast-layer bounce, the number of true contenders for any one seat is small, and a buyer who needs that seat has nothing better to do than wait a few milliseconds. The blocking is bounded on every connection by a 2-second lock timeout and a 5-second statement timeout, so no transaction can wait forever. Everything I described in the lock trade-off table above is what I’d change first as contention grew.
Principal-level talking points for the hold
- The database owns the invariant; everything else is advice. Then any failure of the advice degrades to “slower or rejected”, never “wrong”.
- Lock in a global order. It’s the one-line fix to a whole class of deadlocks.
- Order the checks by cost, cheapest first, so doomed requests are refused before they cost a lock.
- Make retries safe by design. Store the response with the state change in the same transaction.
- Timeouts are part of the design, not configuration. A lock timeout is what converts “the system is stuck” into “this request failed fast and can be retried”.
Pricing and promotions
Pricing is a small deterministic function, not a merchandising system:
public static int multiplier(long occupied, long total, long days) {
long ratio = occupied * 100 / Math.max(1, total);
return Math.min(
15000, (ratio >= 80 ? 15000 : ratio >= 50 ? 12500 : 10000) + (days <= 7 ? 500 : 0));
}
public static Quote quote(int base, int multiplier, int discountBps) {
long gross = (Math.multiplyExact((long) base, multiplier) + 5000) / 10000;
long discount = gross * discountBps / 10000;
return new Quote(Math.toIntExact(gross - discount), Math.toIntExact(discount));
}
Multipliers are basis points, so 10,000 is 1.0×. A section under half sold is 1.0×, half sold is 1.25×, 80% or more is 1.5×, with an extra 0.05× within a week of the event and a hard cap at 1.5×. Everything is integer arithmetic, and Math.multiplyExact throws on overflow instead of wrapping, with explicit half-up rounding, so no floating-point money ever appears.
The promotion cap (“first 500 orders”) is enforced by locking the promotion row, checking used < max_uses, and inserting a (code, user_id) row whose composite primary key makes the per-user limit a constraint and the global cap a lock. That’s the simple, honest answer, and it’s also a hotspot: one hot promotion serializes every checkout that uses it. Pre-allocating tokens would relieve that, but it needs its own reconciliation story, and I’d want measurements before building it.
Principal-level talking point: snapshots and freezing. The scarcity read (how full is this section?) uses a normal committed snapshot, not a serializable one, so two buyers in the same instant can see the same occupancy. What is serialized is ownership of each seat. That’s why browsing shows a base price, and the frozen quote returned by the hold is the only price that counts. Decide which numbers must be exact and which are allowed to be advisory, and be explicit about the difference.
Payments: the phantom charge
Now the third failure. Asha holds her seats and taps Pay. The payment provider call can’t be part of a database transaction, and it can time out after the provider has acted. The whole design here is about that ambiguity.
%%{init: {"themeVariables": {"actorBkg": "#FFFFFF", "actorBorder": "#0E1A2B", "signalColor": "#22324A", "noteBkgColor": "#FFFFFF", "noteBorderColor": "#0E1A2B", "labelBoxBkgColor": "#B02F7E", "labelBoxBorderColor": "#B02F7E", "labelTextColor": "#FFFFFF", "loopTextColor": "#0E1A2B", "sequenceNumberColor": "#FFFFFF"}}}%%
sequenceDiagram
autonumber
box rgb(232,238,253) Outside
actor Fan
participant PP as Payment provider
end
box rgb(251,240,217) Buy plane
participant BK as Booking service
participant FW as Fulfillment worker
end
box rgb(238,232,250) Authoritative
participant DB as System of record
end
box rgb(250,229,240) Async
participant EB as Event backbone
end
rect rgb(232,238,253)
Note over Fan,EB: 1 · Create the intent without ever double-charging
Fan->>BK: Pay, with an idempotency key
BK->>DB: Persist a stable payment ID and an initialize event, commit
BK->>PP: Create the intent, payment ID as the provider idempotency key
PP-->>BK: Provider reference and client secret
BK->>DB: Save the reference and the response
BK-->>Fan: Client secret for the hosted checkout
end
rect rgb(251,240,217)
Note over Fan,EB: 2 · Only the webhook moves money
Fan->>PP: Complete the payment, 3-D Secure or OTP
PP->>BK: Signed webhook
BK->>BK: Verify signature and timestamp
BK->>DB: Lock reservation and payment, verify reference, amount, currency
alt still pending and before the deadline
BK->>DB: PAID, seats SOLD, order paid event to the outbox, COMMIT
else expired or cancelled
BK->>DB: REFUND_REQUIRED, COMMIT
end
BK-->>PP: 200
end
rect rgb(250,229,240)
Note over Fan,EB: 3 · Everything after the commit is retryable
DB-->>EB: Relay publishes outbox rows in order
EB-->>FW: order paid, or refund required
FW->>DB: Inbox row and one ticket per seat, or a refund with a stable key
end
The first step is a small transaction that persists a stable internal payment ID and an outbox event before the provider is contacted. The provider call then happens outside any transaction, using that ID as the provider-side idempotency key, and a second transaction stores the result. If the process dies in between, the durable event lets a background worker retry with the same key, so the provider hands back the same intent rather than creating a second charge. A webhook that arrives before the provider reference has been saved gets a retryable 503 and comes back later.
The webhook handler is where money moves, so it is deliberately paranoid:
private WebhookResult applySuccessfulPayment(SuccessfulPayment event) {
bookings.lockWebhookEvent(event.eventId());
var existing = bookings.findWebhookFingerprint(settings.provider, event.eventId());
if (existing.isPresent()) {
if (!existing.get().equals(event.fingerprint())) {
throw Problem.conflict("Webhook event ID has different content");
}
return new WebhookResult("duplicate");
}
Reservation reservation =
bookings
.lockReservation(event.reservationId())
.orElseThrow(() -> Problem.conflict("Unknown payment"));
Payment payment =
bookings
.lockPaymentByReservation(event.reservationId())
.orElseThrow(() -> Problem.conflict("Unknown payment"));
verifySuccessfulPayment(event, reservation, payment);
bookings.insertProcessedWebhook(settings.provider, event.eventId(), event.fingerprint());
if (List.of(
PaymentStatus.SUCCEEDED, PaymentStatus.REFUND_REQUIRED, PaymentStatus.REFUNDED)
.contains(payment.status())) {
return new WebhookResult("duplicate");
}
if (reservation.status() != ReservationStatus.PENDING_PAYMENT
|| !reservation.expiresAt().isAfter(Instant.now())) {
return requireRefund(reservation, payment);
}
completeSale(reservation, payment);
return new WebhookResult("paid");
}
In order: the signature and timestamp are verified before this function is reached. Then the event ID is deduplicated with a content fingerprint, so a redelivered event is a no-op but the same ID carrying different content is a conflict. Then the reservation and payment rows are locked, and provider reference, amount and currency must all match. Only then does it decide the outcome, and there are exactly two: completeSale marks the reservation PAID and the seats SOLD and writes an order.paid outbox row, or requireRefund records a durable refund obligation.
That second branch answers “what if Asha pays at minute 10:01?” A late success never resurrects the hold, because that seat may already belong to someone else. It becomes a payment.refund_required event, and a worker refunds through the provider with a stable idempotency key. The test suite covers a late payment, a wrong amount (which must never sell), and eight identical webhooks fired in parallel: exactly one PAID transition, one set of tickets, and one notification.
Principal-level talking points for payments
- You can’t make a network call atomic with a database write. So write your intent first, make the call idempotent, and make recovery a background job that doesn’t depend on the customer’s browser surviving.
- The webhook is the source of truth; the browser is a courtesy. The customer may close the tab, lose signal or never come back.
- Design the failure branch as carefully as the happy one. “Paid too late” isn’t an edge case on a flash sale; it’s a Tuesday.
- Deduplicate on content, not just on ID. The same ID with different content is a red flag, not a retry.
After the money moves
Nothing after the commit is allowed to be fire-and-forget. The rule: the domain change and its event commit in one transaction (the outbox table from earlier), and a separate relay process publishes the events afterwards. That’s the transactional outbox, and it exists because “write to the database and publish to a broker” is two systems that can’t share a transaction.
private int relayBatch(KafkaProducer<String, String> producer) {
if (!events.tryAcquireOutboxRelayLock()) return 0;
List<OutboxEvent> batch = events.lockUnpublishedEvents(BATCH_SIZE);
for (OutboxEvent event : batch) {
applyDerivedEffect(event);
publishToKafka(producer, event);
events.markPublished(event.id());
}
return batch.size();
}
The relay takes an advisory lock so only one instance runs at a time, reads unpublished rows in ID order, publishes each with idempotence on and full acknowledgement, waits for the acknowledgement, and only then marks the row published. A crash between publish and mark means the event goes out twice, and that’s by design: the guarantee is at-least-once transport plus idempotent effects, not exactly-once transport. The fulfillment consumer writes an inbox row and issues tickets with a unique constraint per reservation item, so a duplicate delivery is a no-op, and ticket creation and the dedupe record commit together.
The same relay also drives the live seat map. When it processes an inventory.changed event, it invalidates the cache and then publishes to the section’s channel:
private void invalidateSection(OutboxEvent event) {
UUID eventId = requiredUuid(event, "event_id");
String section = requiredText(event, "section");
cache.redis.del("map:" + eventId + ":" + section, "map:" + eventId + ":*");
cache.redis.publish(
"inventory:" + eventId + ":" + section,
Json.write(
Map.of("type", "invalidate", "event_id", eventId, "section", section)));
}
The order is deliberate: delete the cache first, then publish. A browser that reacts to the invalidation instantly can’t read the entry that was just declared stale. It also means the entire live-map pipeline is driven by committed transactions, not by a side channel that could run ahead of the truth.
Notifications go through RabbitMQ with publisher confirms, manual acknowledgements, retry queues at 5, 10, 20 and 40 seconds, and then a dead-letter queue. Financial events get the opposite treatment: a failed one stops progress and stays retryable, rather than being parked in a dead-letter queue nobody watches.
Principal-level talking points for messaging
- The outbox turns a distributed transaction into two local ones with a durable bridge between them.
- At-least-once plus idempotent consumers is the honest guarantee. “Exactly-once” is a property of your effects, and you build it.
- A single ordered relay is a deliberate trade-off. It gives strict ordering, and it’s also head-of-line blocking if a broker is down, so the age of the oldest unpublished event is one of the most important numbers in the system.
- Different failures deserve different queues. A lost email is a retry. A lost payment event is an incident.
Expiry: two things race for the same row
A hold must return to inventory when it isn’t paid for, and that reclamation races against a payment that might be confirming at the same moment. The expiry worker runs once a second, using the one place I do use SKIP LOCKED:
"SELECT * FROM reservations WHERE status='PENDING_PAYMENT' AND expires_at <= "
+ "clock_timestamp() ORDER BY expires_at FOR UPDATE SKIP LOCKED LIMIT ?"
It’s the right use because any expired reservation is fair game: several workers can reclaim different batches without waiting on each other, or on a row that a webhook is holding. The rule that makes it safe is the same one the webhook follows: both paths lock the reservation row first. Whoever gets the lock decides the terminal state.
%%{init: {"themeVariables": {"actorBkg": "#FFFFFF", "actorBorder": "#0E1A2B", "signalColor": "#22324A", "noteBkgColor": "#FFFFFF", "noteBorderColor": "#0E1A2B", "labelBoxBkgColor": "#6B46C1", "labelBoxBorderColor": "#6B46C1", "labelTextColor": "#FFFFFF", "loopTextColor": "#0E1A2B", "sequenceNumberColor": "#FFFFFF"}}}%%
sequenceDiagram
autonumber
box rgb(251,240,217) Buy plane
participant WH as Webhook handler
participant EW as Expiry worker
end
box rgb(238,232,250) Authoritative
participant DB as System of record
end
Note over WH,EW: Both paths must lock the reservation row first
par
WH->>DB: Lock the reservation row
and
EW->>DB: Select expired rows, skip any that are locked
end
alt the webhook got the lock first
rect rgb(226,244,234)
DB-->>WH: Locked
WH->>DB: PAID, or REFUND_REQUIRED if past the deadline
DB-->>EW: Row skipped, nothing to do
end
else expiry got the lock first
rect rgb(252,231,227)
DB-->>EW: Locked
EW->>DB: EXPIRED, release the seats and the promotion use
DB-->>WH: The webhook now sees EXPIRED
WH->>DB: REFUND_REQUIRED
end
end
There is no order in which a seat is both sold and released. Releasing a reservation also gives back its promotion use and emits an event, after which the relay does owner-checked cleanup of the cache hints. A backed-up relay can leave a seat looking unavailable in the cache for a while. It can never leave one sold twice.
Principal-level talking point: lock the same thing first. Two concurrent paths that can each end a reservation are safe if and only if they contend on the same lock before deciding. That single discipline replaces a great deal of clever conflict resolution.
Getting in: tickets and venue entry
A ticket’s QR code isn’t a static image. Its payload is an HMAC over the ticket ID and a 15-second time window, keyed with a per-ticket secret stored encrypted:
public static String qr(byte[] secret, String ticket, long window) {
return ticket
+ "."
+ window
+ "."
+ HexFormat.of()
.formatHex(hmac(secret, (ticket + ":" + window).getBytes(StandardCharsets.UTF_8)));
}
public static boolean validQr(byte[] secret, String payload) {
try {
String[] p = payload.split("\\.");
long w = Long.parseLong(p[1]);
return p.length == 3
&& Math.abs(Instant.now().getEpochSecond() / 15 - w) <= 1
&& equal(qr(secret, p[0], w), payload);
} catch (Exception e) {
return false;
}
}
A screenshot goes stale within seconds. Note the <= 1: the verifier accepts the current window and one either side to tolerate clock drift between phone and scanner, so a code is really good for up to 30 to 45 seconds, not 15.
For entry, the venue runs a small standalone service backed by SQLite in WAL mode. Before the event, an encrypted, expiring manifest of paid tickets is exported into it. A scan is then a single INSERT into an entries table keyed by ticket ID, and a replay at a second gate hits a uniqueness error and is rejected as already used. It’s the same idea as everywhere else in this system: a uniqueness constraint does the arbitration, not application logic.
Principal-level talking point: the CAP corner you’re actually in. Two fully disconnected gates can’t both stay available and globally reject the same ticket’s second use, because rejecting it requires knowing about the first, which requires talking. So I chose consistency over availability at the venue: every gate shares one authoritative ledger on the venue LAN, and if that ledger or the LAN is down, gates fail closed. The alternative, independent offline scanners, is more available and permits a double entry in the window before they sync. Which one is right is a business decision, and the honest thing is to make it explicitly and write it down rather than pretend the trade-off isn’t there.
Running it: SLIs, SLOs, SLAs and tags for SRE
Building it is half the job; the other half is being able to say, at 9:04 p.m., whether it’s healthy. Here is how I’d instrument and operate a system like this. The services already expose Prometheus-style metrics endpoints; the targets below are the ones I’d build on top of them. They’re starting points to tune against load tests and real sales, not measured results.
SLI, SLO, SLA: three different things
- An SLI (indicator) is a measurement: the fraction of good events out of valid events.
- An SLO (objective) is the internal target for that measurement, over a window.
- An SLA (agreement) is the external promise, with consequences, and it should be weaker than the SLO so you have room to be wrong.
The subtle craft is defining good. For example, a hold that returns 409 Conflict because someone else got the seat first is a success, not an error. It’s the system doing exactly what it should. If you count conflicts as failures, your dashboards go red at precisely the moment the system is behaving perfectly, and you train your team to ignore red.
The SLIs and SLOs I’d start with
| Journey | SLI (good ÷ valid) | Starting SLO | Why |
|---|---|---|---|
| Join the queue | Joins answered with a position, not an error | 99.9% during sale windows | The front door; a failure here is invisible to the database |
| Browse the seat map | Snapshot requests not answered 5xx | 99.9% | Forgiving content; stale is fine, absent is not |
| Seat-map freshness | Inventory changes visible to viewers within 2 s | 99% | Measured by outbox age; the brief asked for far tighter, unmeasured |
| Reserve seats | Requests ending 201, 409 or 422 (not 5xx or timeout) | 99.9% | Conflicts count as good |
| Reserve latency | p99 of successful holds | Under 1 s | A lock timeout of 2 s means a p99 above that is a smell |
| Start a payment | Payment initializations succeeding | 99.9% | Idempotent, so retries are safe |
| Webhook handling | Valid webhooks acknowledged within the provider’s timeout | 99.9% | A slow webhook is how paid fans get refunded |
| Ticket fulfillment | Paid orders with tickets issued within 60 s | 99.9% | The promise the fan actually cares about |
| Hold reclamation | Expired holds released within 10 s of the deadline | 99% | Cadence is a second; batch is 100 |
| Notifications | Delivered within 5 minutes | 99% | Retry queues absorb the rest |
Correctness SLIs have a zero-tolerance target, not a percentage. These are checks I’d run continuously and alert on the first violation:
- A seat with more than one live owner. The constraint should make this impossible, so any constraint violation in the logs is a page.
- A
PAIDreservation with no tickets after five minutes. - A payment stuck in
INITIALIZINGfor more than two minutes. - A
REFUND_REQUIREDpayment older than 15 minutes. - Anything in a dead-letter queue on the financial path.
The SLA I’d actually publish is simple and human: “Your tickets appear in your account within 15 minutes of payment, or your money is returned automatically.” That’s much weaker than the 60-second fulfillment SLO on purpose. The gap between them is the room to be wrong.
Error budgets, and why a flash sale is different
An SLO of 99.9% over 28 days gives an error budget of about 40 minutes (28 days is 40,320 minutes, and 0.1% of that is 40.3). At 99.95% it’s about 20 minutes; at 99.99% it’s about 4 minutes, which is what “four nines” really means in practice.
But a flash sale breaks the usual model. Almost all the risk is concentrated in a 30-minute window on a few days a year, so a 28-day rolling average smooths away exactly the period that matters. I’d therefore scope budgets per sale window and count them by requests, not by minutes: during the on-sale window, no more than 0.1% of requests may fail. And I’d tighten alerting windows during the sale, for example checking a 15-minute and a 2-minute window rather than the 6-hour and 30-minute pair a steady-state service would use.
For steady-state alerting I’d use the standard multi-window, multi-burn-rate pattern: page when the budget is burning at 14.4× over one hour (and 5 minutes), page again at 6× over six hours (and 30 minutes), and open a ticket at 1× over three days. Alerting on burn rate rather than raw error percentage means you’re paged for problems that threaten the objective and left alone for blips that don’t.
Tags that make the data usable
Metrics are only as good as their labels. The set I’d standardize on, and the one I’d refuse:
| Label | Values | Purpose |
|---|---|---|
plane | admit, browse, buy, state | Which of the three planes owns this signal |
journey | join, browse, reserve, pay, fulfil, enter | Ties every metric to a customer-visible flow |
tier | t0-money, t1-purchase, t2-browse, t3-comms | Drives alert routing and severity |
slo | e.g. reserve-availability | Lets a dashboard and an alert refer to the same objective |
sale_phase | pre, on_sale, post | Applies the right thresholds and budgets |
outcome | success, conflict, rejected, error | Keeps conflicts out of the error rate |
event_id | Only for the current top few events | Per-sale comparison without exploding cardinality |
Never label with a user ID, reservation ID or seat ID. Metrics are for aggregates; per-request detail belongs in traces and logs, joined to metrics by exemplars or trace IDs. Routing follows the tier: t0-money pages a human immediately at any hour, t1 pages during a sale, t2 opens a ticket, and t3 goes to a queue.
What I’d watch, per component
- Waiting room: queue depth, admission rate, active-lease count against capacity, stale-heartbeat cleanups per tick, and time from join to admit.
- Fan-out: viewer count, upstream subscription count, and the slow-client disconnect counter that the hub already keeps.
- Read model: cache hit ratio, fill-lease contention, and pool utilization on the read-only connection.
- Booking service: hold latency and outcome mix, lock-wait time, database pool active versus pending, and request-thread saturation.
- Coordination store: memory, evictions, command latency, and subscriber counts.
- Outbox and messaging: oldest unpublished outbox age, consumer lag, retry-queue depth, and dead-letter depth.
Load-shedding, in order
When the system is overloaded, what you refuse first should be a decision made in advance, not improvised at 9:03 p.m. Mine is a ladder, from least to most painful:
- Serve slightly staler seat snapshots by extending the cache lifetime.
- Pause live push updates and let browsers poll.
- Shrink the admission batch, then the capacity. This is the dial from the design overview: one change that slows everything downstream.
- Stop admitting entirely. The line holds, and fans just wait.
- Never shed payment webhooks, refunds or fulfillment. Those are the money path.
Before any sale, I’d run a readiness drill: kill the coordination store and confirm that the queue and holds fail closed; stop the booking service and confirm browsing and queueing carry on (the isolation script in the repo does exactly this); stop the outbox relay and confirm the freshness alert fires. An SLO you’ve never seen page is a guess.
Epilogue
At 9:00:01, Asha sees “You are number 4,812”, a live number that ticks down. She’s never in the stampede, because the stampede is happening in a lightweight tier with no database behind it. At 9:07 she’s admitted, picks 14F and 14G, and gets a price that won’t change under her. The database, not the cache, is what says yes. She pays; the provider times out mid-call; the same intent comes back on retry; the webhook arrives; her tickets appear. If someone else had beaten her to 14F, she’d have got a clean, instant “those seats are taken”, not a phantom charge.
Four ideas did most of the work. Let the database own the invariant and treat every faster layer as advice. Bound the concurrency in front of your scarcest resource, and let everything else wait cheaply. Make every side effect idempotent, because delivery is always at-least-once. And separate load by what it holds scarce, so a flood in one place can’t starve another. Everything else, the Lua scripts, the lock modes, the outbox, the burn-rate alerts, is those four ideas applied one layer at a time.