Changeset 0.33.0 (#207)

Co-authored-by: Jeffrey Smith <jasafpro@gmail.com>
Co-committed-by: Jeffrey Smith <jasafpro@gmail.com>
This commit is contained in:
2026-03-19 21:37:32 +00:00
committed by xcaliber
parent b1266b0d7c
commit ed3e9363f2
42 changed files with 2527 additions and 129 deletions

238
docs/DESIGN-0.33.0.md Normal file
View File

@@ -0,0 +1,238 @@
# DESIGN-0.33.0 — Observability
Operate the platform without reading Go source code. Six changesets
covering structured logging, Prometheus metrics, OpenAPI docs, Grafana
dashboards, alerting rules, and a built-in admin dashboard.
Depends on: v0.32.0 (Multi-Replica HA, current HEAD).
**Design decision:** Zero new database migrations. All metrics are
in-memory per-pod, scraped by Prometheus. Existing tables
(`provider_health`, `tool_health_windows`, `usage_log`, `audit_log`)
continue serving the admin dashboard via stores.
**Library choices:**
- `log/slog` (Go 1.22 stdlib) — zero-dep structured logging
- `prometheus/client_golang` v1.20+ — canonical Go Prometheus client
- Hand-curated OpenAPI 3.0.3 YAML — avoids annotation sprawl across 97 handlers
- Swagger UI v5 via CDN — single HTML embed, zero build step
---
## What Already Existed
These required no structural changes, only instrumentation:
- **Health accumulator** — in-memory provider+tool metrics, 60s DB flush
- **`Hub.ConnCount()`** — active WebSocket connection count
- **`database.DB.Stats()`** — Go stdlib DB pool metrics
- **K8s probes** — `/healthz/ready` + `/healthz/live` in Helm chart
- **Middleware skip paths** — `/metrics` already in `SkipPaths` list
- **`google/uuid`** — already in go.mod for request ID generation
---
## CS0 — Structured Logging + Request ID
Foundation changeset. All subsequent work emits structured logs.
### `server/logging/logger.go` (new)
`Init(format, level)` configures the global `slog` logger:
- `format=json``slog.NewJSONHandler(os.Stdout, opts)`
- `format=text``slog.NewTextHandler(os.Stdout, opts)`
- Levels: debug, info, warn, error
### `server/middleware/request_id.go` (new)
Generates UUID per request. Reuses existing `google/uuid`.
Honors inbound `X-Request-Id` header for trace propagation.
Sets `request_id` in Gin context and echoes header on response.
### Middleware chain rewrite
Changed `gin.Default()``gin.New()` with explicit chain:
```
RequestID → Prometheus → Logger → Recovery → CORS
```
Logger middleware rewritten from `log.Printf` to `slog.Info("request", ...)`
with structured fields: method, path, status, latency_ms, client_ip,
request_id, user_id.
### Config additions
- `LogFormat` (env: `LOG_FORMAT`, default: `text`)
- `LogLevel` (env: `LOG_LEVEL`, default: `info`)
---
## CS1 — Prometheus `/metrics` Endpoint
### `server/metrics/metrics.go` (new)
All metric definitions in one file using `promauto`:
| Metric | Type | Labels |
|--------|------|--------|
| `switchboard_http_requests_total` | Counter | method, path_pattern, status |
| `switchboard_http_request_duration_seconds` | Histogram | method, path_pattern |
| `switchboard_websocket_connections` | Gauge | — |
| `switchboard_completion_tokens_total` | Counter | direction, model_id |
| `switchboard_completions_total` | Counter | provider_config_id, model_id, status |
| `switchboard_completion_duration_seconds` | Histogram | provider_config_id, model_id |
| `switchboard_provider_status` | Gauge | provider_config_id |
| `switchboard_db_open_connections` | Gauge | — |
| `switchboard_db_in_use_connections` | Gauge | — |
| `switchboard_db_idle_connections` | Gauge | — |
| `switchboard_db_wait_count` | Gauge | — |
| `switchboard_db_wait_duration_seconds` | Gauge | — |
| `switchboard_task_executions_total` | Counter | status |
| `switchboard_sandbox_executions_total` | Counter | entry_point, status |
### `server/metrics/db_collector.go` (new)
Background goroutine reads `database.DB.Stats()` every 15s, updates
Prometheus gauges.
### `server/middleware/prometheus.go` (new)
Gin middleware. Uses `c.FullPath()` for `path_pattern` label to avoid
cardinality explosion from path parameters (e.g. `/channels/:id`
not `/channels/abc123`).
### Instrumentation points
- `events/ws.go``WebSocketConnections.Inc()/Dec()` on connect/disconnect
- `handlers/completion.go` — token counters in `logUsage()`, completion
duration/status in `recordHealth()`
- `handlers/stream_loop.go` — same instrumentation in `recordHealthFn()`
- `health/accumulator.go``ProviderStatus` gauge updated in `flush()`
### Route
`/metrics` mounted via `promhttp.Handler()` (no auth, standard scraping).
---
## CS2 — OpenAPI Spec + Swagger UI
### `server/static/openapi.yaml` (new)
Hand-curated OpenAPI 3.0.3 spec covering core API groups:
Auth, Channels, Completions, Health, Metrics, WebSocket Tickets,
Admin (stats, health, usage, dashboard).
### `server/static/swagger.html` (new)
Minimal HTML loading Swagger UI v5 from unpkg CDN, pointing at
`openapi.yaml`. Zero build step.
### Routes
- `GET /api/docs` → serves embedded `swagger.html`
- `GET /api/docs/openapi.yaml` → serves embedded `openapi.yaml`
Both files embedded via `//go:embed` directives.
---
## CS3 — Grafana Dashboard + Alerting Rules + Helm
Static files + Helm templates. No Go code changes.
### `chart/dashboards/switchboard-overview.json` (new)
Grafana dashboard with 10 panels:
- Request rate, error rate, latency percentiles
- WebSocket connections, completion rate/latency by provider
- Tokens/min by model, provider status
- DB connection pool, task executions
Template variables: `$datasource`, `$namespace`, `$pod`.
### `chart/alerting/switchboard-rules.yaml` (new)
Source PrometheusRule with 6 alerts:
| Alert | Condition | Severity |
|-------|-----------|----------|
| SwitchboardPodRestart | Restarts in 1h | warning |
| SwitchboardProviderDown | Status > 2 for 5m | critical |
| SwitchboardDBPoolExhaustion | In-use/open > 80% for 5m | warning |
| SwitchboardHighErrorRate | 5xx rate > 5% for 5m | warning |
| SwitchboardTaskFailureRate | Error rate > 25% for 10m | warning |
| SwitchboardNoCompletions | Zero completions for 15m | critical |
### Helm templates (new, all gated `enabled: false` by default)
- `chart/templates/servicemonitor.yaml` — ServiceMonitor CRD
- `chart/templates/grafana-dashboard-configmap.yaml` — ConfigMap for
Grafana sidecar discovery
- `chart/templates/prometheusrule.yaml` — PrometheusRule CRD
### Helm modifications
- `chart/values.yaml``logging` section + `monitoring` section
- `chart/templates/configmap.yaml``LOG_FORMAT`, `LOG_LEVEL` env vars
- `chart/templates/services.yaml` — Prometheus scrape annotations
---
## CS4 — Admin Observability Dashboard
Built-in admin page for real-time health without Grafana dependency.
### `server/handlers/dashboard_admin.go` (new)
`GET /api/v1/admin/dashboard` aggregating:
- Provider health summaries (from `HealthStore`)
- 24h usage totals (from `UsageStore`)
- DB pool stats (`database.DB.Stats()`)
- WebSocket connection count (`Hub.ConnCount()`)
- Recent errors (from `AuditStore`)
- Process uptime
### Frontend
- `SCAFFOLDING.dashboard` — stat cards grid + content container
- `ADMIN_LOADERS.dashboard``UI.loadAdminDashboard()`
- `ADMIN_SECTIONS.monitoring``dashboard` added as first section
- Dashboard auto-refreshes every 30s
- CSS: provider grid, DB pool bar gauge, error list
### Admin template
Monitoring tab now lands on `/admin/dashboard` instead of `/admin/usage`.
---
## CS5 — Wiring, docker-compose, ICD Tests
### docker-compose.yml
Added `LOG_FORMAT` and `LOG_LEVEL` env vars with defaults.
### ICD tests (`crud/observability.js`)
| Test | Assertion |
|------|-----------|
| `GET /metrics` | 200, contains `switchboard_*` metrics |
| `GET /api/docs` | 200, contains `swagger-ui` |
| `GET /api/docs/openapi.yaml` | 200, contains `openapi:` and `paths:` |
| X-Request-Id generation | Response includes 36-char UUID header |
| X-Request-Id passthrough | Client-sent header echoed back |
| `GET /admin/dashboard` | 200, has `uptime`, `ws_connections`, `provider_health` |
---
## Verification
1. `LOG_FORMAT=json docker compose up` → JSON log lines on stdout
2. `curl localhost:8080/metrics | grep switchboard_` → all metrics present
3. `curl -I localhost:8080/api/v1/channels``X-Request-Id` header
4. `/api/docs` in browser → Swagger UI renders
5. `/admin` → Monitoring → Dashboard shows live stat cards
6. `helm lint chart/ --set monitoring.serviceMonitor.enabled=true` → passes
7. ICD test suite passes (existing + observability tests)