All checks were successful
Co-authored-by: Jeffrey Smith <jasafpro@gmail.com> Co-committed-by: Jeffrey Smith <jasafpro@gmail.com>
408 lines
14 KiB
Markdown
408 lines
14 KiB
Markdown
# DESIGN: Surface Runners — v0.7.1–v0.7.3
|
||
|
||
## Status: v0.7.1 Shipped (framework + migrations), v0.7.2–v0.7.3 Proposed
|
||
|
||
## Problem
|
||
|
||
Armature has two test tiers today:
|
||
|
||
1. **Go unit tests** — test store methods, handlers, middleware, sandbox.
|
||
Run in CI on every push. Coverage is good for kernel internals.
|
||
|
||
2. **ICD/SDK test runners** — browser-based test suites that validate API
|
||
endpoint contracts and SDK domain methods. Run manually by navigating
|
||
to `/s/icd-test-runner` or `/s/sdk-test-runner`.
|
||
|
||
Neither tier catches the class of bugs discovered during manual testing:
|
||
|
||
- **Cross-surface state:** Notification bell state not syncing, announcement
|
||
dismiss not persisting across surface navigations.
|
||
- **Package integration:** Workflow demo shows "not installed" because its
|
||
API call fails silently. The API works (unit tests pass); the surface's
|
||
integration with the API is broken.
|
||
- **Test side-effects:** ICD security tests install `evil.surface` with no
|
||
cleanup — it leaks into the menu.
|
||
- **Surface lifecycle:** Surfaces fail to load, mount into wrong containers,
|
||
miss SDK boot, or render without shell chrome.
|
||
|
||
These are integration bugs — they live at the boundary between kernel and
|
||
package, between surface and surface, between API and UI. They require a
|
||
new test tier.
|
||
|
||
## Solution Overview
|
||
|
||
Three deliverables across three versions:
|
||
|
||
| Version | Deliverable | What it catches |
|
||
|---------|-------------|-----------------|
|
||
| v0.7.1 | Runner framework (`sw.testing`) | Framework bugs, standardizes existing runners |
|
||
| v0.7.2 | Package runners + CI gate | Package integration bugs, regressions |
|
||
| v0.7.3 | Headless E2E (Playwright) | DOM rendering bugs, navigation flows, visual regressions |
|
||
|
||
## v0.7.1 — Runner Framework
|
||
|
||
### `sw.testing` SDK Module
|
||
|
||
New kernel SDK module at `src/js/sw/sdk/testing.js`. Provides structured
|
||
test authoring, lifecycle hooks, cleanup tracking, and machine-readable
|
||
results.
|
||
|
||
```js
|
||
// Extension runner registers suites during load
|
||
sw.testing.suite('notes-crud', async (s) => {
|
||
let folderId, noteId;
|
||
|
||
s.beforeAll(async () => {
|
||
// Setup: create a test folder
|
||
const r = await sw.api.post('/api/v1/ext/notes/folders', {
|
||
name: 'test-' + Date.now()
|
||
});
|
||
folderId = r.id;
|
||
s.track('folder', folderId); // auto-cleanup
|
||
});
|
||
|
||
s.test('create note', async (t) => {
|
||
const r = await sw.api.post('/api/v1/ext/notes/notes', {
|
||
title: 'Test Note', folder_id: folderId, content: '# Hello'
|
||
});
|
||
t.assert.ok(r.id, 'note has ID');
|
||
t.assert.eq(r.title, 'Test Note');
|
||
noteId = r.id;
|
||
t.track('note', noteId); // auto-cleanup
|
||
});
|
||
|
||
s.test('renderers fire', async (t) => {
|
||
// Test that mermaid block in note content triggers renderer
|
||
await sw.api.patch('/api/v1/ext/notes/notes/' + noteId, {
|
||
content: '```mermaid\ngraph LR; A-->B\n```'
|
||
});
|
||
// Renderer integration tested via DOM assertion
|
||
// (only meaningful in headless E2E — marked as browser-only)
|
||
t.browserOnly(() => {
|
||
const el = document.querySelector('.mermaid svg');
|
||
t.assert.ok(el, 'mermaid rendered to SVG');
|
||
});
|
||
});
|
||
|
||
s.afterAll(async () => {
|
||
// s.track() resources auto-cleaned here
|
||
// Manual cleanup for anything not tracked
|
||
});
|
||
});
|
||
```
|
||
|
||
### Core API
|
||
|
||
```js
|
||
sw.testing.suite(name, fn) // Register a test suite
|
||
sw.testing.run(name?) // Run one suite or all
|
||
sw.testing.results() // Get structured results (JSON)
|
||
sw.testing.on('complete', fn) // Event when run finishes
|
||
|
||
// Inside suite:
|
||
s.test(name, fn) // Register a test
|
||
s.beforeAll(fn) // Runs once before all tests
|
||
s.afterAll(fn) // Runs once after all tests (always, even on failure)
|
||
s.beforeEach(fn) // Runs before each test
|
||
s.afterEach(fn) // Runs after each test
|
||
s.track(type, id) // Register resource for auto-cleanup
|
||
s.skip(reason) // Skip entire suite
|
||
|
||
// Inside test:
|
||
t.assert.ok(val, msg) // Truthy
|
||
t.assert.eq(a, b, msg) // Deep equality
|
||
t.assert.neq(a, b, msg) // Not equal
|
||
t.assert.gt(a, b, msg) // Greater than
|
||
t.assert.match(str, re, msg) // Regex match
|
||
t.assert.throws(fn, msg) // Expects throw
|
||
t.assert.status(resp, code, msg) // HTTP status check
|
||
t.track(type, id) // Register resource for auto-cleanup
|
||
t.warn(msg) // Emit warning (non-fatal)
|
||
t.browserOnly(fn) // Only runs in headless E2E, skipped in API-only mode
|
||
t.skip(reason) // Skip this test
|
||
```
|
||
|
||
### `requires` Declarations
|
||
|
||
Runner packages declare dependencies in their manifest:
|
||
|
||
```json
|
||
{
|
||
"id": "chat-runner",
|
||
"type": "test-runner",
|
||
"title": "Chat Runner",
|
||
"requires": ["chat", "chat-core"],
|
||
"version": "0.1.0"
|
||
}
|
||
```
|
||
|
||
On load, the framework calls `GET /api/v1/surfaces` (or equivalent) to
|
||
check which packages are installed. If any `requires` entry is missing:
|
||
|
||
- Suite is marked `skipped` with reason: `"Missing required package: chat-core"`
|
||
- No tests execute — clean skip, not a failure
|
||
- The runner registry surface shows the skip reason prominently
|
||
|
||
This directly solves the "workflow demo shows not-installed" pattern:
|
||
the runner *knows* what should be installed and reports clearly when it isn't.
|
||
|
||
### Auto-Cleanup
|
||
|
||
The `track(type, id)` method registers resources for deletion in `afterAll`.
|
||
Supported resource types and their cleanup endpoints:
|
||
|
||
| Type | Cleanup Action |
|
||
|------|---------------|
|
||
| `channel` | `DELETE /api/v1/channels/:id` |
|
||
| `note` | `DELETE /api/v1/ext/notes/notes/:id` |
|
||
| `folder` | `DELETE /api/v1/ext/notes/folders/:id` |
|
||
| `workflow` | `DELETE /api/v1/workflows/:id` |
|
||
| `schedule` | `DELETE /api/v1/schedules/:id` |
|
||
| `package` | `DELETE /api/v1/admin/packages/:id` |
|
||
| `user` | `DELETE /api/v1/admin/users/:id` |
|
||
| `team` | `DELETE /api/v1/admin/teams/:id` |
|
||
|
||
Cleanup runs in reverse order (LIFO) in `afterAll`, regardless of
|
||
test pass/fail. Cleanup failures are reported as warnings, not failures.
|
||
The framework never swallows cleanup errors silently.
|
||
|
||
### Result Structure
|
||
|
||
```json
|
||
{
|
||
"runner": "notes-runner",
|
||
"timestamp": "2026-04-01T12:00:00Z",
|
||
"duration_ms": 1234,
|
||
"summary": { "total": 5, "passed": 4, "failed": 0, "warned": 1, "skipped": 0 },
|
||
"suites": [
|
||
{
|
||
"name": "notes-crud",
|
||
"status": "passed",
|
||
"duration_ms": 890,
|
||
"tests": [
|
||
{
|
||
"name": "create note",
|
||
"status": "passed",
|
||
"duration_ms": 120,
|
||
"warnings": [],
|
||
"cleanup": { "tracked": 1, "cleaned": 1, "failed": 0 }
|
||
}
|
||
]
|
||
}
|
||
],
|
||
"requires": { "met": ["notes"], "missing": [] }
|
||
}
|
||
```
|
||
|
||
### Warning Tier
|
||
|
||
Three result statuses:
|
||
|
||
- **`passed`** — assertions all passed, cleanup succeeded
|
||
- **`failed`** — at least one assertion failed
|
||
- **`warned`** — assertions passed but something non-fatal happened:
|
||
- API returned unexpected shape (extra/missing fields) but test doesn't depend on the exact field
|
||
- Cleanup failed for a tracked resource
|
||
- Timing exceeded a soft threshold
|
||
- `t.warn(msg)` called explicitly
|
||
|
||
Warnings are **never silent.** They appear in the UI and structured results.
|
||
The difference from the current `catch (e) { /* ignore */ }` pattern is
|
||
that warnings are *visible* — a human or CI system can decide whether
|
||
to investigate.
|
||
|
||
### ICD/SDK Runner Migration
|
||
|
||
The existing runners use a hand-rolled framework (`T.test()`, `T.assert()`,
|
||
`T.authFetch()`). Migration preserves all test logic:
|
||
|
||
| Current | New |
|
||
|---------|-----|
|
||
| `T.test(tier, group, name, fn)` | `s.test(name, fn)` inside `sw.testing.suite(tier + '/' + group, fn)` |
|
||
| `T.assert(cond, msg)` | `t.assert.ok(cond, msg)` |
|
||
| `T.authFetch(token, method, path, body)` | Kept as utility — not an assertion primitive |
|
||
| `T.apiPost(...)` | Kept as utility |
|
||
| Result rendering (`ui.js`) | Delegated to runner registry surface |
|
||
| No cleanup hooks | `s.track()` + `s.afterAll()` |
|
||
|
||
The ICD and SDK runners become packages with `"type": "test-runner"` in
|
||
their manifests. Their existing surfaces (`/s/icd-test-runner`,
|
||
`/s/sdk-test-runner`) are replaced by the unified runner registry at
|
||
`/s/test-runners`.
|
||
|
||
## v0.7.2 — Package Runners
|
||
|
||
### Runner Inventory
|
||
|
||
| Runner | `requires` | Key Assertions |
|
||
|--------|-----------|---------------|
|
||
| `notes-runner` | `["notes"]` | CRUD, folders, tags, backlinks, search, markdown rendering, SDK integration |
|
||
| `chat-runner` | `["chat", "chat-core"]` | Channel CRUD, messaging, participant display, renderer blocks in messages |
|
||
| `schedules-runner` | `["schedules"]` | Schedule CRUD, cron expression, toggle, Starlark exec |
|
||
| `workflow-runner` | `["content-approval"]` | Install detection, stage progression, form submission, signoff |
|
||
| `renderer-runner` | `["mermaid-renderer"]` | `sw.renderers.register` contract, post-render hooks, block rendering |
|
||
|
||
### Runner Result API
|
||
|
||
New kernel endpoints (no package required — kernel-provided):
|
||
|
||
```
|
||
POST /api/v1/test-runners/run → Run all installed runners
|
||
POST /api/v1/test-runners/run/:id → Run specific runner
|
||
GET /api/v1/test-runners/results → Last run results (JSON)
|
||
GET /api/v1/test-runners/results/:id → Last run results for specific runner
|
||
```
|
||
|
||
These endpoints enable CI to trigger and consume runner results via
|
||
`curl` without browser automation. The v0.7.3 Playwright harness is
|
||
additive — not required for CI gating.
|
||
|
||
**Auth:** Admin-only. Runners create/delete resources — they must run
|
||
with elevated permissions.
|
||
|
||
### CI Integration
|
||
|
||
New stage in `.gitea/workflows/ci.yaml`:
|
||
|
||
```yaml
|
||
test-runners:
|
||
needs: [unit-tests]
|
||
steps:
|
||
- name: Boot server
|
||
run: |
|
||
docker compose up -d
|
||
./ci/wait-for-healthy.sh
|
||
- name: Run surface runners
|
||
run: |
|
||
RESULT=$(curl -s -X POST http://localhost:8080/api/v1/test-runners/run \
|
||
-H "Authorization: Bearer $ADMIN_TOKEN")
|
||
FAILED=$(echo "$RESULT" | jq '.summary.failed')
|
||
if [ "$FAILED" != "0" ]; then
|
||
echo "$RESULT" | jq '.suites[] | select(.status == "failed")'
|
||
exit 1
|
||
fi
|
||
```
|
||
|
||
Runs in both PG and SQLite pipelines. Server boots with `BUNDLED_PACKAGES=*`
|
||
so all packages and their runners are installed.
|
||
|
||
## v0.7.3 — Headless E2E
|
||
|
||
### Playwright Harness
|
||
|
||
`ci/e2e-surface-test.sh`:
|
||
|
||
1. `docker compose up -d` (server + DB)
|
||
2. `npx playwright install chromium` (CI caches this)
|
||
3. Run `ci/e2e-surfaces.spec.ts`
|
||
4. Collect screenshots on failure
|
||
5. `docker compose down`
|
||
|
||
### Surface Navigation Smoke Test
|
||
|
||
```ts
|
||
test('all surfaces reachable', async ({ page }) => {
|
||
// Login
|
||
await page.goto('/');
|
||
await page.fill('#username', 'admin');
|
||
await page.fill('#password', 'admin');
|
||
await page.click('button[type="submit"]');
|
||
|
||
// Navigate through every installed surface
|
||
const surfaces = ['notes', 'chat', 'admin', 'settings', 'docs'];
|
||
for (const s of surfaces) {
|
||
await page.goto(`/s/${s}`);
|
||
// Assert: page loaded, no uncaught JS errors
|
||
await expect(page.locator('.sw-topbar')).toBeVisible();
|
||
// Assert: home link works
|
||
await page.click('.sw-topbar__home');
|
||
await expect(page).toHaveURL('/');
|
||
}
|
||
});
|
||
```
|
||
|
||
This is the automated version of "hello dashboard has no way out" — if
|
||
any surface fails to render a topbar or its home link doesn't work, CI
|
||
catches it.
|
||
|
||
### Screenshot on Failure
|
||
|
||
```ts
|
||
test.afterEach(async ({ page }, testInfo) => {
|
||
if (testInfo.status !== 'passed') {
|
||
await page.screenshot({
|
||
path: `ci/artifacts/failure-${testInfo.title}.png`,
|
||
fullPage: true
|
||
});
|
||
const logs = await page.evaluate(() =>
|
||
(window.__consoleErrors || []).join('\n')
|
||
);
|
||
fs.writeFileSync(
|
||
`ci/artifacts/console-${testInfo.title}.log`, logs
|
||
);
|
||
}
|
||
});
|
||
```
|
||
|
||
Artifacts saved to CI workspace. On failure, the developer gets a
|
||
screenshot + console log dump without needing to reproduce locally.
|
||
|
||
### Visual Regression (Optional)
|
||
|
||
Not a CI gate in v0.7.3 — produces a diff report for human review:
|
||
|
||
```ts
|
||
test('visual baseline - notes', async ({ page }) => {
|
||
await page.goto('/s/notes');
|
||
await expect(page.locator('.sw-topbar')).toBeVisible();
|
||
await expect(page).toHaveScreenshot('notes.png', {
|
||
maxDiffPixelRatio: 0.01
|
||
});
|
||
});
|
||
```
|
||
|
||
Playwright stores baseline screenshots in `ci/visual-baselines/`.
|
||
`toHaveScreenshot` compares against baseline and produces a diff image
|
||
on mismatch. Foundation for future visual regression gating.
|
||
|
||
## Sequencing
|
||
|
||
```
|
||
v0.7.0 Shell Contract ← prerequisite: surfaces need topbar before
|
||
│ runners can assert on it
|
||
↓
|
||
v0.7.1 Runner Framework ← standardize test authoring
|
||
│
|
||
↓
|
||
v0.7.2 Package Runners + CI ← write the actual tests, wire into CI
|
||
│
|
||
↓
|
||
v0.7.3 Headless E2E ← automate browser-based runner execution
|
||
```
|
||
|
||
Each version is independently shippable. v0.7.2's API-based CI gate
|
||
works without v0.7.3's Playwright. v0.7.3 adds coverage for DOM-level
|
||
bugs that API-only runners can't catch.
|
||
|
||
## Open Questions
|
||
|
||
1. **Runner package type.** Should `"type": "test-runner"` be a new
|
||
manifest type, or should runners be `"type": "surface"` with a
|
||
`"tags": ["test-runner"]` convention? New type is cleaner but requires
|
||
a `ValidateManifest()` update.
|
||
|
||
2. **Runner discovery.** The registry surface needs to find all installed
|
||
runners. Options: (a) scan installed packages for `type: "test-runner"`,
|
||
(b) runners register themselves via `sw.testing.register()` during SDK
|
||
boot. Option (a) is declarative and doesn't require runner JS to load
|
||
before discovery.
|
||
|
||
3. **Parallel vs sequential.** Should runners execute in parallel?
|
||
Probably not initially — shared DB state means test isolation is hard.
|
||
Sequential is safer. Parallel can be a future optimization.
|
||
|
||
4. **SQLite limitations.** Some runners (cluster, multi-node) are
|
||
PG-only. The `requires` mechanism should support
|
||
`"requires_db": "postgres"` for these cases, or runners should
|
||
self-skip when `sw.config.db_driver === 'sqlite'`.
|