Skip to main content

0001. Actors and a typed event bus in the backend

  • Status: Accepted
  • Date: 2026-09-30

Context

The backend holds several streams of state that change at the same time: order books for many tokens, game state, open orders, fills, positions, risk counters and alerts. Market, user and sports WebSocket feeds update this state while HTTP handlers read it and submit commands. Each feed can drop and reconnect independently. Mistakes here cost money: a data race on an order or a risk counter can send a wrong order.

Decision

  • Each piece of mutable state is owned by one actor: a goroutine that reads a mailbox channel. No other code reads or writes that state.
  • Other code talks to an actor by message. Request/reply messages carry a reply channel and a context.Context for the timeout.
  • Actors publish events on a typed bus: BookUpdate, Trade, GameState, OrderEvent, Fill, PositionChanged, RiskEvent, AlertFired, FeedStatus.
  • Bus subscribers have bounded buffers. A slow subscriber drops messages and resyncs from a snapshot; it never blocks the publisher.
  • thejerf/suture/v4 supervises every long-lived actor and restarts a crashed one with backoff. Feed actors reconnect themselves, publish FeedStatus, and do not exit on a transient error.
  • Only store.Writer writes to SQLite.
  • Helpers live in internal/actor (mailbox, request/reply, supervised run) and internal/bus.

Consequences

  • No locks are shared across components, and each actor can be tested by sending it messages.
  • Each order passes through one actor (engine/orders), which gives one place for idempotency keys, audit writes and the call to engine/risk.
  • A dropped feed shows up as a FeedStatus event, which the UI turns into greyed prices.
  • Reading state from another component takes a message round trip instead of a function call. Request/reply needs a timeout, which the caller’s context provides.
  • A subscriber that drops messages must be able to rebuild from a snapshot, so each topic needs a snapshot path.
  • Message types add some boilerplate compared with direct method calls.

Alternatives considered

  • Shared structs guarded by mutexes. Less code at first, but lock order and hidden shared state become hard to review as feeds and actors grow.
  • One event loop for the whole backend. Simple to reason about, but one slow handler delays every feed and every order.
  • An external broker (NATS, Redis). Adds a process to run and secure, and gives nothing for a single-user, single-process app.
  • Unbounded bus buffers. Never drop, but a stuck subscriber grows memory without limit and hides the problem.