0001. Actors and a typed event bus in the backend
- Status: Accepted
- Date: 2026-09-30
Context
The backend holds several streams of state that change at the same time: order books for many tokens, game state, open orders, fills, positions, risk counters and alerts. Market, user and sports WebSocket feeds update this state while HTTP handlers read it and submit commands. Each feed can drop and reconnect independently. Mistakes here cost money: a data race on an order or a risk counter can send a wrong order.Decision
- Each piece of mutable state is owned by one actor: a goroutine that reads a mailbox channel. No other code reads or writes that state.
- Other code talks to an actor by message. Request/reply messages carry a reply channel
and a
context.Contextfor the timeout. - Actors publish events on a typed bus:
BookUpdate,Trade,GameState,OrderEvent,Fill,PositionChanged,RiskEvent,AlertFired,FeedStatus. - Bus subscribers have bounded buffers. A slow subscriber drops messages and resyncs from a snapshot; it never blocks the publisher.
thejerf/suture/v4supervises every long-lived actor and restarts a crashed one with backoff. Feed actors reconnect themselves, publishFeedStatus, and do not exit on a transient error.- Only
store.Writerwrites to SQLite. - Helpers live in
internal/actor(mailbox, request/reply, supervised run) andinternal/bus.
Consequences
- No locks are shared across components, and each actor can be tested by sending it messages.
- Each order passes through one actor (
engine/orders), which gives one place for idempotency keys, audit writes and the call toengine/risk. - A dropped feed shows up as a
FeedStatusevent, which the UI turns into greyed prices. - Reading state from another component takes a message round trip instead of a function call. Request/reply needs a timeout, which the caller’s context provides.
- A subscriber that drops messages must be able to rebuild from a snapshot, so each topic needs a snapshot path.
- Message types add some boilerplate compared with direct method calls.
Alternatives considered
- Shared structs guarded by mutexes. Less code at first, but lock order and hidden shared state become hard to review as feeds and actors grow.
- One event loop for the whole backend. Simple to reason about, but one slow handler delays every feed and every order.
- An external broker (NATS, Redis). Adds a process to run and secure, and gives nothing for a single-user, single-process app.
- Unbounded bus buffers. Never drop, but a stuck subscriber grows memory without limit and hides the problem.