# BuyerKiosk E2E — AI-driven testing with Midscene.js

AI-agent E2E testing for the three BuyerKiosk apps using [Midscene.js](https://midscenejs.com)
(v1.12.x). Midscene is **pure vision**: a vision-language model looks at screenshots, plans
actions in natural language, and drives the UI. No selectors, no DOM coupling — which is
exactly why it works on the Flutter apps (canvas rendering is irrelevant to it).

| App | Platform | Driver | Target |
|---|---|---|---|
| buyerkiosk-web | Web (PHP/Slim) | `@midscene/web` (Puppeteer/CLI) | `http://localhost:8080` (Docker) or dev2 |
| buyerkiosk-live-flutter | iOS | `@midscene/ios` + WebDriverAgent | Simulator, bundle `com.v2ts.BuyerKiosk` |
| buyerkiosk-team | iOS | `@midscene/ios` + WebDriverAgent | Simulator, bundle `com.buyerkiosk.buyerkioskTeam` |

> **Why not `lhuanyu/midscene-ios`?** That community repo (iPhone-Mirroring + PyAutoGUI) is
> unmaintained since Sept 2025, can't drive simulators, and can't launch apps by bundle ID.
> Its author's work was upstreamed into the official `@midscene/ios` package (actively
> maintained, WebDriverAgent-based, simulators + real devices), which is what we use.

## One-time setup

1. **Model API key** — paste your OpenRouter key into `MIDSCENE_MODEL_API_KEY` in `.env`
   (https://openrouter.ai/settings/keys). Default model is `z-ai/glm-4.6v`
   (`family: glm-v`) — best cost/performance for GUI grounding at ~$0.30/M input.
   Escalate accuracy-critical flows to `qwen/qwen3.8-max` (`family: qwen3`, ~7x cost);
   the full model menu with pricing is in `.env.example`. Midscene needs a
   *visual-grounding* VLM — a plain chat model won't work.
2. **Dependencies** — already done, but for a fresh clone: `npm install && npm rebuild`
   (approve install scripts if prompted) and `npx puppeteer browsers install chrome`.
3. **WebDriverAgent** — `npm run wda:setup` downloads Appium's prebuilt simulator WDA
   (no signing, no Apple account needed for simulators).

## Running web tests

```bash
npm run test:web            # headless
npm run test:web:headed     # watch it live
```

`BK_WEB_URL` defaults to **`https://dev2.buyerkiosk.com`** — which is an ngrok tunnel to
THIS machine's Apache serving the local working tree, backed by the local dev databases.
It's the same backend the mobile dev builds use, so web + iOS suites share one dataset
(enabling cross-app flows: create data on web, assert it in the app).

Use the Docker stack (`./docker-dev.sh up`, `BK_WEB_URL=http://localhost:8080`) only for
**destructive suites** that need an isolated, resettable database — Docker mounts the same
code but runs its own MySQL, so tests can trash and reset data without touching dev2's DBs.

Fill `BK_WEB_USER`/`BK_WEB_PASS` with a dev-store account. Authenticated suites
use the secure CDP login helper so credentials are not copied into Midscene
prompts or reports; the older `tests/web/login.yaml` is retained only as a
reference and is not part of the default scripts.

## Running iOS tests

```bash
# 1. Boot simulator + install/launch WDA (waits for http://localhost:8100/status)
npm run wda:start

# 2. Install the app(s) under test — debug simulator builds:
#    (cd ../buyerkiosk-live-flutter && flutter build ios --simulator --debug)
xcrun simctl install booted ../buyerkiosk-live-flutter/build/ios/iphonesimulator/Runner.app
xcrun simctl install booted ../buyerkiosk-team/build/ios/iphonesimulator/Runner.app

# 3. Run
npm run test:ios:live
npm run test:ios:team
```

**Real devices** work too but need more ceremony: WDA built + signed in Xcode with your
team, Developer Mode + UI Automation enabled on the phone, and `iproxy 8100 8100` for port
forwarding. Simulators need none of that — start there.

### ⚠️ Apps currently point at PRODUCTION

- Live: `lib/core/constants/app_config.dart` has `isDebug = false` → `buyerkiosk.com`
- Team: `lib/core/constants/api_constants.dart` has `baseUrl = prodBaseUrl` → `buyerkiosk.com`

Before writing login/data-mutating E2E flows, flip those to the dev backends
(`dev2.buyerkiosk.com` / `try.buyerkiosk.com`) and rebuild, so tests never touch prod data.
Launch-smoke tests (current suites) are safe — they don't log in.

## Scheduling integrity suite

The scheduling suite covers the failures hardened in the three application
repositories: explicit employee clearing, publication visibility, canonical
DST/overnight elapsed hours, split-shift totals, payroll approval/rounding/
idempotency, AI cancellation/atomic application/provenance, and terminal
pending-request pagination.

```bash
npm run test:scheduling:web
npm run test:scheduling:ios:live
npm run test:scheduling:ios:team
```

Safety requirements:

- Use an isolated Docker store for destructive payroll/AI runs. If Docker is
  unavailable, run only shift fixtures with a unique `E2E_RUN_ID`; those flows
  delete only rows containing that marker.
- Build both iOS apps with an explicit non-production endpoint before running:

  ```bash
  (cd ../buyerkiosk-live-flutter && flutter build ios --simulator --debug \
    --dart-define=BK_API_BASE_URL=https://dev2.buyerkiosk.com/api/)
  (cd ../buyerkiosk-team && flutter build ios --simulator --debug \
    --dart-define=BK_API_BASE_URL=https://dev2.buyerkiosk.com)
  ```

- Web scheduling runners create an authenticated CDP browser with
  `scripts/start-authenticated-browser.mjs`. Credentials are read in Node and
  never interpolated into Midscene prompts/reports/logs.
- iOS runners authenticate through WDA before Midscene starts. Configure
  `BK_LIVE_USER/BK_LIVE_PASS` and `BK_TEAM_USER/BK_TEAM_PASS` directly in
  `.env`; never paste them into YAML or chat.
- `scheduling-payroll.yaml` and `scheduling-ai.yaml` require dedicated seeded
  fixtures identified by `E2E_RUN_ID`; they intentionally fail closed when the
  fixtures are absent rather than operating on unrelated rows.
- Every mutating prompt is scoped to the marker, and lifecycle/boundary flows
  end with marker-based cleanup.

## Reports

Every run writes a self-contained HTML report (screenshots embedded, per-step AI reasoning)
to `./midscene_run/report/<id>.html` — just open it in a browser. `aiQuery` extractions land
in the JSON files configured per script.

## Letting AI agents drive tests (Claude Code)

Midscene **retired its MCP servers** (last MCP version 1.9.8 — don't build on it). The
current integration is **Midscene Skills**, installed via:

```bash
npx skills add web-infra-dev/midscene-skills -a claude-code
```

With skills installed, Claude Code drives the platform CLIs directly
(`npx @midscene/web`, `npx @midscene/ios`), reading screenshots from CLI output and
deciding the next action. Modes for web: default (own headless Chrome), `--bridge`
(your desktop Chrome via the [Midscene Chrome extension](https://chromewebstore.google.com/detail/midscene/gbldofcpkknbggpkmbdaefngejllnief)),
or `--cdp <ws-endpoint>` (attach to any CDP browser).

Agents can also simply run the YAML suites (`npm run test:web` etc.) and read the reports —
that's the most reproducible loop for CI-style testing.

## Writing new tests

YAML is the primary format (`tests/**/*.yaml`). Flow steps: `ai`/`aiAct` (free-form
action), `aiTap`, `aiInput`, `aiScroll`, `aiKeyboardPress`, `aiHover`, `aiAssert`,
`aiQuery` (structured extraction), `aiBoolean`, `aiWaitFor`, `sleep`, `javascript` (web
only). `${VAR}` interpolates from `.env`. iOS targets support `launch: <bundleId>`,
`terminate:`, and raw `runWdaRequest`. See `sdk-examples/` for the TypeScript SDK
equivalents (useful when you need loops/logic around AI steps).

Useful screen-flow references for authoring iOS tests:
- Live app routes/screens: `../buyerkiosk-live-flutter/CLAUDE.md` (Navigation Flow section)
- Team app: `../buyerkiosk-team/CLAUDE.md`

## Scheduling fixture cleanup

Scheduling E2E rows must use one exact unique `E2E-*` notes marker. Before cleanup,
run the fail-closed dry run:

```bash
php scripts/cleanup-e2e-shifts.php --marker=E2E-MY-UNIQUE-RUN
```

Apply only after the reported count matches the fixture count:

```bash
php scripts/cleanup-e2e-shifts.php --marker=E2E-MY-UNIQUE-RUN --apply
```

The helper is pinned to the non-production `ou00` store, rejects wildcard/short
markers, refuses more than one matching row by default, uses a transaction, and
reads back zero remaining rows before reporting success. Use `--max=N` only for
a deliberately multi-row fixture. Pass `--web-root=/path/to/userfrosting` when
the Web repository is not at the default sibling location.

See `SCHEDULING_E2E_EVIDENCE.md` for the current pass/fail ledger and deterministic
DB/iOS read-back evidence. A Midscene replanning-limit failure is never counted as
a passing E2E test.

## Troubleshooting

- **WDA won't come up**: `npm run wda:status`; relaunch with
  `xcrun simctl launch booted com.facebook.WebDriverAgentRunner.xctrunner`. First cold boot
  of a new simulator is slow — give it a couple of minutes.
- **Taps landing offset on iOS**: make sure you're on `@midscene/ios` ≥ 1.12 (older
  versions had a Dynamic Island offset bug, fixed via `/wda/screen` scale factor).
- **Keyboard won't dismiss in Flutter apps**: custom accessory bars aren't auto-detected;
  use `device.hideKeyboard(['Done', 'Close Keyboard'])` in SDK code.
- **Web report flashing/blank screenshots**: set `deviceScaleFactor` to match your display
  (2 on Retina) or leave the YAML value at 1 for headless runs.
- **Model errors**: run with `MIDSCENE_RECORD_MODEL_CALL=true` to dump raw model traffic;
  `MIDSCENE_MODEL_TIMEOUT` (default 180s) if a slow model times out.
