client_order_id, and the tab that clicked leaves the toast to the click (contract bebb70f, BE 70a6f28, FE fe/phase-4-fixes). Fixed: shared formatCount caps the tab at 99+ (fe/phase-4-fixes). Fixed: the phone score takes its own line (fe/phase-4-fixes). Fixed: each delete has its own toast id (fe/phase-4-fixes). # End-to-end UI suite
A Playwright suite in e2e/ (TypeScript, exact pinned versions, its own package.json and
lockfile) that drives the real frontend against the real backend running the fake venue,
including the Phase 3 trading engine and the Phase 4 alert and notification actors
(engine/alerts, notify), which fire on the fake venue’s moving books and scores. It is not part of task lint or task test. It
runs before every phase tag: the tech lead (or QA for the exit gate) runs task e2e on
main and it must pass before phase-<n> is tagged. The manual review that goes with it is
ui-checklist.md.
Running it
On this host the full run takes about 6 minutes with the default 6 workers. Most of it is
the trading, alert and notification tests, each on its own backend, the idle-timeout test (about
70 s by design) and the two tests that wait for a real book move or score change (usually
5–20 s, at most 120 s).
Reports:
e2e/playwright-report/index.html (HTML, with traces for failures: npx playwright show-report in e2e/), e2e/reports/axe-report.md (every axe violation of the run), and backend
logs per stack in e2e/.cache/logs/. All are git-ignored.
Environment switches:
How it is built
- Projects:
chromium,firefox,webkitat 1440 × 900;mobile-chromium(Pixel 7 profile) andmobile-webkit(iPhone 14 profile) at 390 × 844 with touch;stress-chromium(1440) andstress-mobile-chromium(390), which run onlystress.spec.ts, against backends started withSESAME_FAKE_VENUE_STRESS=1(thestackEnvworker option). - Global setup (
support/global-setup.ts) buildsbackend/cmd/sesameintoe2e/.cache/bin, builds the frontend intoe2e/.cache/web, makes a random throwaway password, hashes it withsesame hash-password, and starts the shared stack: the backend withSESAME_FAKE_VENUE=1, that hash, a randomSESSION_SECRET, a tempDATA_DIRand free ports, plus Vite in front of it. It logs in once through the API and saves the storage state, setstheme_modetosystem, and pins two markets so the watchlist has stable content. Teardown stops everything and writes the axe report. - Free ports without touching the frontend:
frontend/vite.config.tshard-codes the proxy target (127.0.0.1:8080) and port 5173.support/web-server.mjsloads that same config through Vite’s JS API with an inline override of the port and proxy target, so the suite runs beside a developer’stask devand other agents’ servers. - Clean environment: backend and Vite get an allowlisted environment (PATH, temp dirs, proxy
settings). Nothing from
.env, which Task loads, reaches them, so real venue credentials never mix with the fake venue. - Three kinds of stack (
support/fixtures.ts):test: the shared stack and its one session, for read-only flows. It sets no alert and makes no notification (the fake account generates no activity of its own), so no toast or unread badge reaches the visual baselines or the hygiene checks.privateTest: one backend and Vite per worker with its own session, for tests that change server state or stop the backend (login and logout, pins, theme, stale/restart, stress).freshTest: one backend per test, for trading, alerts and notifications. Every test starts SAFE with the seeded account, no alert and no notification, and no order, fill, alert or ARMED state leaks into the next. It also gives the test an API client with its session and CSRF token, for setup the test does not cover (stake presets, risk limits) and for reading the backend’s view after a UI action.
- Themes: the shared backend’s
theme_modeissystem, and each test picks dark or light with Playwright’scolorScheme, so parallel tests never fight over a server-side setting. The product default (dark) is covered on a private stack bytheme.spec.ts.
Why 6 workers
Every shared-stack page uses the one saved session, and the hub allows 8 WebSockets per session (hub.DefaultMaxSessionSockets). With 16 workers the backend refused upgrades (ws upgrade refused reason="socket limit") and pages sat on “Loading orders”. Separate sessions per worker
are not possible on one backend either, because logins are limited to 5 per IP per minute. Six
workers leave headroom for sockets that are still closing.
What it covers
Trading tests
Each runs on a fresh stack (freshTest) against engine/orders, engine/risk and the fake
venue’s rules (backend/internal/testutil/fakevenue/trading.go): 13 shares is rejected, a price
ending in 7 ticks is “unknown” and reaches the venue 1 s later, and marketable orders on sports
markets wait out a 3 s delay window before they fill.
Phase 4: alerts and notifications
Each runs on a fresh stack. The fake books move every 1–2 s; about one turn in ten moves the mid by a tick, and the book keeps one empty level between bid and ask (a 2¢ spread). The two live games (Celtics v Nuggets, Everton v Fulham) score about once every 15 s. A spread alertat or below 2¢ fires as soon as it is set, which is how the tests make many notifications
quickly. Helpers are in support/alerts.ts: a recorder of the page’s /ws frames, a toast
timer that runs in the page, the one-click ladder alert (the bell on the hovered row on
desktop, a dispatched 500 ms touch long-press and Set alert at … on a phone), and a book
replay that finds the delta that met an alert’s condition.
Measured latency, book or score change to toast on screen, over five runs: 10–150 ms in every
engine, once 840 ms on the phone project under full parallel load (the gate allows 2 s). Toasts come only from notifications that arrive as deltas, so tests that expect a
toast first wait for the page’s
notifications snapshot (subscribed); a notification made
before the page subscribed arrives inside the snapshot and, correctly, raises no toast.
Telegram is not covered end to end. notify.TelegramConfig.BaseURL exists, but
internal/app never sets it from the environment, so the backend cannot be pointed at a local
stub without a code change. The client is covered by the Go tests against httptest servers.
Visual baselines
Re-baselined withtask e2e:update on lead/phase-4-kickoff (the Phase 3 polish-round
baselines before that). Every logged-in screen changed, for Phase 4 shell work only: the
Alerts item in the rail (desktop), the inbox button in the top bar, the refitted top bar and
phone header (4a2af0b), and the trading-state note bar under the header (Adopted 3 orders open at first start), which pushes the page content down. The login screens did not change. alerts and notifications are new
(empty states, as the shared stack has none). visual.spec.ts
keeps a pendingRebaseline switch: set it to true while a UI round is in flight to hold the
comparison as fixme (task e2e:update still runs the tests, because it passes
--update-snapshots=all), and back to false with the new PNGs.
Counts on fe/phase-8-polish (Phase 8)
The full run onlead/phase-8 plus the Phase 8 polish (the quiet panel indicator, esports
scores, the phone combo exit kept on screen): 419 tests pass, none fail, none are flaky and
none are expected failures, 302 are skipped, in about 6 minutes with 6 workers. 721 tests in
all; Phase 8 adds combos.spec.ts (9 tests, one of them phone only: the click budgets, quote
states and exits) and combo screens in the visual, axe and stress checks. The run on lead/phase-8
before the polish passed 413 and skipped 296.
Rebaselined with
--update-snapshots=changed, so only the screens that changed were
rewritten: live-combo-dark and live-combo-light in chromium and mobile-chromium. The
setup clicks inside the board in Safe, which used to draw the heavy focus ring round the
board; the quiet indicator does not show there. The new baselines passed 20 of 20 with
--repeat-each=5.
Stress mode adds two CS2 games carrying the esports scores captured from the sports channel
(backend/internal/venue/polymarket/testdata/sports_ws_frames_esports.jsonl). Its long
league name found a league header that truncated without a tooltip; the header now has one.
Counts on qa/phase-6-finish (final, before the phase 6 tag)
The full run onlead/phase-6 plus the last Phase 6 fixes (virtualised lists, the mobile pass):
361 tests pass, none fail, none are flaky and none are expected failures, 264 are skipped, in
about 5.5 minutes with 6 workers. The three tests that assumed every row renders (history paging
in chromium and mobile-chromium, stress orders on the phone) now read aria-setsize or scroll
and count. The new and changed tests also passed 50 of 50 with --repeat-each=5.
No visual baseline changed, so none was regenerated: the shared stack has no unread badge and
short names, so at 390 px the status word, league headers and orders filter bar look as before,
and the visual test empties the orders list. The fixes show in stress mode and at 360 px, which
the stress checks and the Ghostframe pass cover.
Known issues: none new. mobile-webkit still skips on this Windows host (Known gaps), and the
status word depends on the host (this one is geoblocked, US-VA), which is why the pill is masked.
Counts on lead/phase-4-kickoff (final, before the phase 4 tag)
The full run after the hardening and the fixes below: 355 tests pass, none fail, none are flaky and none are expected failures (support/known-issues.ts lists no findings), 261 are
skipped, in about 5 minutes. mobile-webkit still skips all 118 (the host crash).
The earlier run on 7332fc4 had 7 expected failures (AL-01, AL-02, AL-03, NT-01); all are fixed.
Skips are by design: axe, visual and hygiene checks run in Chromium only, and some flows are
desktop-only (drag, right-click, hotkeys) or phone-only.
Flakiness policy
- No fixed sleeps. Every wait is a web-first assertion (
expect(locator).toBeVisible(),toHaveAccessibleName,expect.poll) or waits for a real event (a process printingready, the backend’s/api/health). Locators use roles and accessible names, the same names screen readers get. The one timed loop samples the page for pictographs under reduced motion and stops at the first hit. - No retries (
retries: 0). A flaky test is fixed at its cause or quarantined withtest.fixmeand a note here; a retry would hide exactly the kind of race a trading UI must not have. - Market movement is expected. The fake books random-walk, so a marketable order may fill or rest; trading tests accept either outcome where the venue’s behaviour allows both, and choose resting prices relative to the current best bid.
- Stable visuals: numerals are made transparent by
support/visual.css(the boxes keep their size), prices, clocks, charts, order dots, account figures and the host-dependent status pill are masked, the page clock is paused before the capture, and the orders list is emptied to its fixed-size box. - Time-dependent fake data: the books and games move on their own, so tests assert what is always there (the seeded orders and positions), never exact counts. The fake account generates no order or fill of its own in the app (that would be foreign activity).
- Book readiness: a test that reads a ladder first waits for
data-book="ready"on its level area. The event-view test that watched the book move was flaky while the fake books moved rarely; with a move every 1–2 s and that wait it passed 40 of 40 (--repeat-each=10, four projects, 6 workers in parallel) and in every full run since. - The baseline notice: the shared stack never arms, so the first start’s
Adopted N orders open at first startbar stays under the header. It appears when the trading topic’s snapshot arrives and moves the page down, so the visual test waits for it before the capture. - Waiting for a real trigger: the fired-alert and score tests set alerts that any mid move or either game’s next score fires, so they usually finish in 5–20 s; their budget is 120 s.
Known findings
Measured onlead/phase-4-kickoff (7332fc4). Layout, content, trading and hygiene findings are
test.fail entries in support/known-issues.ts: the test runs, is reported as an expected
failure, and turns into an unexpected pass (which fails the run) once the UI is fixed, so the
entry must then be deleted. Axe findings are listed in knownAxe in support/axe.ts: reported
in reports/axe-report.md, not failing. E2E_IGNORE_KNOWN=1 shows the current state of all of
them.
No axe rule is known-failing:
knownAxe is empty, so every serious or critical violation fails.
Fixed and removed since the first runs: every earlier layout finding (L-01 to L-09, including the
1024 px top-bar overflow), trailing zeros, units in cells and 0/×0 for nothing (S-01 to S-03),
all axe findings (A-01 to A-05), the default document title (H-01), the upper-case transforms and
BUY/SELL/SAFE/ARMED capitals and the chart range ALL (H-02), the missing focus rings on the
portfolio table and phone tabs (H-03), the reduced-motion ▲/▼ glyph (E-01), Cancel all under
the phone ladder sheet (T-02) and the truncated tennis score on the phone board (C-03).
Observed, not findings: stress mode publishes its 200 fills of the day as live fill events at
start, so a stress stack begins with 200 Order filled notifications (fake data, not a replay
by the app).
The layout helper now measures a segmented control (a fieldset or role=group of buttons) as
one control at its own height. The earlier “24 px segments beside 28 px chips” findings compared
the buttons inside the padded group with their neighbours, which the rule does not ask for.
Known gaps
- Mobile WebKit on Windows: WebKit for Windows (WinCairo) crashes the page at phone widths
(390, and intermittently other widths) as soon as the app’s stylesheet applies; desktop WebKit
at 1440 is fine, and a plain page at 390 is fine. The
mobile-webkitproject stays configured and runs on macOS and Linux; on Windows it is skipped unlessE2E_MOBILE_WEBKIT=1. Because this may also be a real Safari crash, check 390 in Safari on a Mac or iPhone by hand each phase (checklist §13). - Visual baselines exist for Chromium only (1440 and 390) and for Windows. A Linux runner
needs its own baselines (
task e2e:updatethere); the path includes the platform. - Telegram is not reached end to end: the backend’s Telegram base URL cannot be set from the environment (Phase 4 section above). Covered by Go tests; the live check is the Telegram runbook.
- Desktop notifications themselves (the OS notification from
/sw.js, its click opening the target, one per tab set) are not observable from Playwright; the permission flow and the worker registration are. Checklist §23. - Sound is not checked (Web Audio output); checklist §23.
- Never-seen placement (
Order not placed,venue_not_found): the fake venue has no rule that loses an order; unit tests cover it. feed_down,foreign_activity,credentials,game_start_orders,session_key_expiringnotifications need a feed outage over 30 s, outside trading or venue state the fake venue does not produce on demand; checklist §23.- axe and the hygiene checks run in Chromium only; the DOM is the same in every engine.
- Not automated: real iOS keyboard and safe areas, 200% zoom, screen-reader output, long sessions and performance with 50+ games, background tabs, the glossary and British spelling; these are in the checklist.
- Venue error states (425, 429, cancel-only, post-only mode, no credentials) cannot be produced by the fake venue, so they are checklist-only.
- Long-press cancel on phones is not automated; the desktop right-click path is. Playwright
has no long-press gesture;
support/alerts.tslongPressdispatches the touch pointer events and holds until the action sheet shows, and is used for the ladder’sSet alert at …. - Structure without test ids: the spots below are located by structure rather than role. They are the first to break when markup changes.