Skip to main content
Fixed: fake books report the market’s own tick (37659ff). Fixed: fake fills carry their outcome (37659ff). Fixed: notifications carry client_order_id, and the tab that clicked leaves the toast to the click (contract bebb70f, BE 70a6f28, FE fe/phase-4-fixes). Fixed: shared formatCount caps the tab at 99+ (fe/phase-4-fixes). Fixed: the phone score takes its own line (fe/phase-4-fixes). Fixed: each delete has its own toast id (fe/phase-4-fixes). # End-to-end UI suite A Playwright suite in e2e/ (TypeScript, exact pinned versions, its own package.json and lockfile) that drives the real frontend against the real backend running the fake venue, including the Phase 3 trading engine and the Phase 4 alert and notification actors (engine/alerts, notify), which fire on the fake venue’s moving books and scores. It is not part of task lint or task test. It runs before every phase tag: the tech lead (or QA for the exit gate) runs task e2e on main and it must pass before phase-<n> is tagged. The manual review that goes with it is ui-checklist.md.

Running it

On this host the full run takes about 6 minutes with the default 6 workers. Most of it is the trading, alert and notification tests, each on its own backend, the idle-timeout test (about 70 s by design) and the two tests that wait for a real book move or score change (usually 5–20 s, at most 120 s). Reports: e2e/playwright-report/index.html (HTML, with traces for failures: npx playwright show-report in e2e/), e2e/reports/axe-report.md (every axe violation of the run), and backend logs per stack in e2e/.cache/logs/. All are git-ignored. Environment switches:

How it is built

  • Projects: chromium, firefox, webkit at 1440 × 900; mobile-chromium (Pixel 7 profile) and mobile-webkit (iPhone 14 profile) at 390 × 844 with touch; stress-chromium (1440) and stress-mobile-chromium (390), which run only stress.spec.ts, against backends started with SESAME_FAKE_VENUE_STRESS=1 (the stackEnv worker option).
  • Global setup (support/global-setup.ts) builds backend/cmd/sesame into e2e/.cache/bin, builds the frontend into e2e/.cache/web, makes a random throwaway password, hashes it with sesame hash-password, and starts the shared stack: the backend with SESAME_FAKE_VENUE=1, that hash, a random SESSION_SECRET, a temp DATA_DIR and free ports, plus Vite in front of it. It logs in once through the API and saves the storage state, sets theme_mode to system, and pins two markets so the watchlist has stable content. Teardown stops everything and writes the axe report.
  • Free ports without touching the frontend: frontend/vite.config.ts hard-codes the proxy target (127.0.0.1:8080) and port 5173. support/web-server.mjs loads that same config through Vite’s JS API with an inline override of the port and proxy target, so the suite runs beside a developer’s task dev and other agents’ servers.
  • Clean environment: backend and Vite get an allowlisted environment (PATH, temp dirs, proxy settings). Nothing from .env, which Task loads, reaches them, so real venue credentials never mix with the fake venue.
  • Three kinds of stack (support/fixtures.ts):
    • test: the shared stack and its one session, for read-only flows. It sets no alert and makes no notification (the fake account generates no activity of its own), so no toast or unread badge reaches the visual baselines or the hygiene checks.
    • privateTest: one backend and Vite per worker with its own session, for tests that change server state or stop the backend (login and logout, pins, theme, stale/restart, stress).
    • freshTest: one backend per test, for trading, alerts and notifications. Every test starts SAFE with the seeded account, no alert and no notification, and no order, fill, alert or ARMED state leaks into the next. It also gives the test an API client with its session and CSRF token, for setup the test does not cover (stake presets, risk limits) and for reading the backend’s view after a UI action.
  • Themes: the shared backend’s theme_mode is system, and each test picks dark or light with Playwright’s colorScheme, so parallel tests never fight over a server-side setting. The product default (dark) is covered on a private stack by theme.spec.ts.

Why 6 workers

Every shared-stack page uses the one saved session, and the hub allows 8 WebSockets per session (hub.DefaultMaxSessionSockets). With 16 workers the backend refused upgrades (ws upgrade refused reason="socket limit") and pages sat on “Loading orders”. Separate sessions per worker are not possible on one backend either, because logins are limited to 5 per IP per minute. Six workers leave headroom for sockets that are still closing.

What it covers

Trading tests

Each runs on a fresh stack (freshTest) against engine/orders, engine/risk and the fake venue’s rules (backend/internal/testutil/fakevenue/trading.go): 13 shares is rejected, a price ending in 7 ticks is “unknown” and reaches the venue 1 s later, and marketable orders on sports markets wait out a 3 s delay window before they fill.

Phase 4: alerts and notifications

Each runs on a fresh stack. The fake books move every 1–2 s; about one turn in ten moves the mid by a tick, and the book keeps one empty level between bid and ask (a 2¢ spread). The two live games (Celtics v Nuggets, Everton v Fulham) score about once every 15 s. A spread alert at or below 2¢ fires as soon as it is set, which is how the tests make many notifications quickly. Helpers are in support/alerts.ts: a recorder of the page’s /ws frames, a toast timer that runs in the page, the one-click ladder alert (the bell on the hovered row on desktop, a dispatched 500 ms touch long-press and Set alert at … on a phone), and a book replay that finds the delta that met an alert’s condition. Measured latency, book or score change to toast on screen, over five runs: 10–150 ms in every engine, once 840 ms on the phone project under full parallel load (the gate allows 2 s). Toasts come only from notifications that arrive as deltas, so tests that expect a toast first wait for the page’s notifications snapshot (subscribed); a notification made before the page subscribed arrives inside the snapshot and, correctly, raises no toast. Telegram is not covered end to end. notify.TelegramConfig.BaseURL exists, but internal/app never sets it from the environment, so the backend cannot be pointed at a local stub without a code change. The client is covered by the Go tests against httptest servers.

Visual baselines

Re-baselined with task e2e:update on lead/phase-4-kickoff (the Phase 3 polish-round baselines before that). Every logged-in screen changed, for Phase 4 shell work only: the Alerts item in the rail (desktop), the inbox button in the top bar, the refitted top bar and phone header (4a2af0b), and the trading-state note bar under the header (Adopted 3 orders open at first start), which pushes the page content down. The login screens did not change. alerts and notifications are new (empty states, as the shared stack has none). visual.spec.ts keeps a pendingRebaseline switch: set it to true while a UI round is in flight to hold the comparison as fixme (task e2e:update still runs the tests, because it passes --update-snapshots=all), and back to false with the new PNGs.

Counts on fe/phase-8-polish (Phase 8)

The full run on lead/phase-8 plus the Phase 8 polish (the quiet panel indicator, esports scores, the phone combo exit kept on screen): 419 tests pass, none fail, none are flaky and none are expected failures, 302 are skipped, in about 6 minutes with 6 workers. 721 tests in all; Phase 8 adds combos.spec.ts (9 tests, one of them phone only: the click budgets, quote states and exits) and combo screens in the visual, axe and stress checks. The run on lead/phase-8 before the polish passed 413 and skipped 296. Rebaselined with --update-snapshots=changed, so only the screens that changed were rewritten: live-combo-dark and live-combo-light in chromium and mobile-chromium. The setup clicks inside the board in Safe, which used to draw the heavy focus ring round the board; the quiet indicator does not show there. The new baselines passed 20 of 20 with --repeat-each=5. Stress mode adds two CS2 games carrying the esports scores captured from the sports channel (backend/internal/venue/polymarket/testdata/sports_ws_frames_esports.jsonl). Its long league name found a league header that truncated without a tooltip; the header now has one.

Counts on qa/phase-6-finish (final, before the phase 6 tag)

The full run on lead/phase-6 plus the last Phase 6 fixes (virtualised lists, the mobile pass): 361 tests pass, none fail, none are flaky and none are expected failures, 264 are skipped, in about 5.5 minutes with 6 workers. The three tests that assumed every row renders (history paging in chromium and mobile-chromium, stress orders on the phone) now read aria-setsize or scroll and count. The new and changed tests also passed 50 of 50 with --repeat-each=5. No visual baseline changed, so none was regenerated: the shared stack has no unread badge and short names, so at 390 px the status word, league headers and orders filter bar look as before, and the visual test empties the orders list. The fixes show in stress mode and at 360 px, which the stress checks and the Ghostframe pass cover. Known issues: none new. mobile-webkit still skips on this Windows host (Known gaps), and the status word depends on the host (this one is geoblocked, US-VA), which is why the pill is masked.

Counts on lead/phase-4-kickoff (final, before the phase 4 tag)

The full run after the hardening and the fixes below: 355 tests pass, none fail, none are flaky and none are expected failures (support/known-issues.ts lists no findings), 261 are skipped, in about 5 minutes. mobile-webkit still skips all 118 (the host crash). The earlier run on 7332fc4 had 7 expected failures (AL-01, AL-02, AL-03, NT-01); all are fixed. Skips are by design: axe, visual and hygiene checks run in Chromium only, and some flows are desktop-only (drag, right-click, hotkeys) or phone-only.

Flakiness policy

  • No fixed sleeps. Every wait is a web-first assertion (expect(locator).toBeVisible(), toHaveAccessibleName, expect.poll) or waits for a real event (a process printing ready, the backend’s /api/health). Locators use roles and accessible names, the same names screen readers get. The one timed loop samples the page for pictographs under reduced motion and stops at the first hit.
  • No retries (retries: 0). A flaky test is fixed at its cause or quarantined with test.fixme and a note here; a retry would hide exactly the kind of race a trading UI must not have.
  • Market movement is expected. The fake books random-walk, so a marketable order may fill or rest; trading tests accept either outcome where the venue’s behaviour allows both, and choose resting prices relative to the current best bid.
  • Stable visuals: numerals are made transparent by support/visual.css (the boxes keep their size), prices, clocks, charts, order dots, account figures and the host-dependent status pill are masked, the page clock is paused before the capture, and the orders list is emptied to its fixed-size box.
  • Time-dependent fake data: the books and games move on their own, so tests assert what is always there (the seeded orders and positions), never exact counts. The fake account generates no order or fill of its own in the app (that would be foreign activity).
  • Book readiness: a test that reads a ladder first waits for data-book="ready" on its level area. The event-view test that watched the book move was flaky while the fake books moved rarely; with a move every 1–2 s and that wait it passed 40 of 40 (--repeat-each=10, four projects, 6 workers in parallel) and in every full run since.
  • The baseline notice: the shared stack never arms, so the first start’s Adopted N orders open at first start bar stays under the header. It appears when the trading topic’s snapshot arrives and moves the page down, so the visual test waits for it before the capture.
  • Waiting for a real trigger: the fired-alert and score tests set alerts that any mid move or either game’s next score fires, so they usually finish in 5–20 s; their budget is 120 s.

Known findings

Measured on lead/phase-4-kickoff (7332fc4). Layout, content, trading and hygiene findings are test.fail entries in support/known-issues.ts: the test runs, is reported as an expected failure, and turns into an unexpected pass (which fails the run) once the UI is fixed, so the entry must then be deleted. Axe findings are listed in knownAxe in support/axe.ts: reported in reports/axe-report.md, not failing. E2E_IGNORE_KNOWN=1 shows the current state of all of them. No axe rule is known-failing: knownAxe is empty, so every serious or critical violation fails. Fixed and removed since the first runs: every earlier layout finding (L-01 to L-09, including the 1024 px top-bar overflow), trailing zeros, units in cells and 0/×0 for nothing (S-01 to S-03), all axe findings (A-01 to A-05), the default document title (H-01), the upper-case transforms and BUY/SELL/SAFE/ARMED capitals and the chart range ALL (H-02), the missing focus rings on the portfolio table and phone tabs (H-03), the reduced-motion ▲/▼ glyph (E-01), Cancel all under the phone ladder sheet (T-02) and the truncated tennis score on the phone board (C-03). Observed, not findings: stress mode publishes its 200 fills of the day as live fill events at start, so a stress stack begins with 200 Order filled notifications (fake data, not a replay by the app). The layout helper now measures a segmented control (a fieldset or role=group of buttons) as one control at its own height. The earlier “24 px segments beside 28 px chips” findings compared the buttons inside the padded group with their neighbours, which the rule does not ask for.

Known gaps

  • Mobile WebKit on Windows: WebKit for Windows (WinCairo) crashes the page at phone widths (390, and intermittently other widths) as soon as the app’s stylesheet applies; desktop WebKit at 1440 is fine, and a plain page at 390 is fine. The mobile-webkit project stays configured and runs on macOS and Linux; on Windows it is skipped unless E2E_MOBILE_WEBKIT=1. Because this may also be a real Safari crash, check 390 in Safari on a Mac or iPhone by hand each phase (checklist §13).
  • Visual baselines exist for Chromium only (1440 and 390) and for Windows. A Linux runner needs its own baselines (task e2e:update there); the path includes the platform.
  • Telegram is not reached end to end: the backend’s Telegram base URL cannot be set from the environment (Phase 4 section above). Covered by Go tests; the live check is the Telegram runbook.
  • Desktop notifications themselves (the OS notification from /sw.js, its click opening the target, one per tab set) are not observable from Playwright; the permission flow and the worker registration are. Checklist §23.
  • Sound is not checked (Web Audio output); checklist §23.
  • Never-seen placement (Order not placed, venue_not_found): the fake venue has no rule that loses an order; unit tests cover it.
  • feed_down, foreign_activity, credentials, game_start_orders, session_key_expiring notifications need a feed outage over 30 s, outside trading or venue state the fake venue does not produce on demand; checklist §23.
  • axe and the hygiene checks run in Chromium only; the DOM is the same in every engine.
  • Not automated: real iOS keyboard and safe areas, 200% zoom, screen-reader output, long sessions and performance with 50+ games, background tabs, the glossary and British spelling; these are in the checklist.
  • Venue error states (425, 429, cancel-only, post-only mode, no credentials) cannot be produced by the fake venue, so they are checklist-only.
  • Long-press cancel on phones is not automated; the desktop right-click path is. Playwright has no long-press gesture; support/alerts.ts longPress dispatches the touch pointer events and holds until the action sheet shows, and is used for the ladder’s Set alert at ….
  • Structure without test ids: the spots below are located by structure rather than role. They are the first to break when markup changes.

Test ids that would help

The suite uses roles and accessible names everywhere else. These are located by structure today: