8.1 KiB
DESIGN-0.33.0 — Observability
Operate the platform without reading Go source code. Six changesets covering structured logging, Prometheus metrics, OpenAPI docs, Grafana dashboards, alerting rules, and a built-in admin dashboard.
Depends on: v0.32.0 (Multi-Replica HA, current HEAD).
Design decision: Zero new database migrations. All metrics are
in-memory per-pod, scraped by Prometheus. Existing tables
(provider_health, tool_health_windows, usage_log, audit_log)
continue serving the admin dashboard via stores.
Library choices:
log/slog(Go 1.22 stdlib) — zero-dep structured loggingprometheus/client_golangv1.20+ — canonical Go Prometheus client- Hand-curated OpenAPI 3.0.3 YAML — avoids annotation sprawl across 97 handlers
- Swagger UI v5 via CDN — single HTML embed, zero build step
What Already Existed
These required no structural changes, only instrumentation:
- Health accumulator — in-memory provider+tool metrics, 60s DB flush
Hub.ConnCount()— active WebSocket connection countdatabase.DB.Stats()— Go stdlib DB pool metrics- K8s probes —
/healthz/ready+/healthz/livein Helm chart - Middleware skip paths —
/metricsalready inSkipPathslist google/uuid— already in go.mod for request ID generation
CS0 — Structured Logging + Request ID
Foundation changeset. All subsequent work emits structured logs.
server/logging/logger.go (new)
Init(format, level) configures the global slog logger:
format=json→slog.NewJSONHandler(os.Stdout, opts)format=text→slog.NewTextHandler(os.Stdout, opts)- Levels: debug, info, warn, error
server/middleware/request_id.go (new)
Generates UUID per request. Reuses existing google/uuid.
Honors inbound X-Request-Id header for trace propagation.
Sets request_id in Gin context and echoes header on response.
Middleware chain rewrite
Changed gin.Default() → gin.New() with explicit chain:
RequestID → Prometheus → Logger → Recovery → CORS
Logger middleware rewritten from log.Printf to slog.Info("request", ...)
with structured fields: method, path, status, latency_ms, client_ip,
request_id, user_id.
Config additions
LogFormat(env:LOG_FORMAT, default:text)LogLevel(env:LOG_LEVEL, default:info)
CS1 — Prometheus /metrics Endpoint
server/metrics/metrics.go (new)
All metric definitions in one file using promauto:
| Metric | Type | Labels |
|---|---|---|
switchboard_http_requests_total |
Counter | method, path_pattern, status |
switchboard_http_request_duration_seconds |
Histogram | method, path_pattern |
switchboard_websocket_connections |
Gauge | — |
switchboard_completion_tokens_total |
Counter | direction, model_id |
switchboard_completions_total |
Counter | provider_config_id, model_id, status |
switchboard_completion_duration_seconds |
Histogram | provider_config_id, model_id |
switchboard_provider_status |
Gauge | provider_config_id |
switchboard_db_open_connections |
Gauge | — |
switchboard_db_in_use_connections |
Gauge | — |
switchboard_db_idle_connections |
Gauge | — |
switchboard_db_wait_count |
Gauge | — |
switchboard_db_wait_duration_seconds |
Gauge | — |
switchboard_task_executions_total |
Counter | status |
switchboard_sandbox_executions_total |
Counter | entry_point, status |
server/metrics/db_collector.go (new)
Background goroutine reads database.DB.Stats() every 15s, updates
Prometheus gauges.
server/middleware/prometheus.go (new)
Gin middleware. Uses c.FullPath() for path_pattern label to avoid
cardinality explosion from path parameters (e.g. /channels/:id
not /channels/abc123).
Instrumentation points
events/ws.go—WebSocketConnections.Inc()/Dec()on connect/disconnecthandlers/completion.go— token counters inlogUsage(), completion duration/status inrecordHealth()handlers/stream_loop.go— same instrumentation inrecordHealthFn()health/accumulator.go—ProviderStatusgauge updated inflush()
Route
/metrics mounted via promhttp.Handler() (no auth, standard scraping).
CS2 — OpenAPI Spec + Swagger UI
server/static/openapi.yaml (new)
Hand-curated OpenAPI 3.0.3 spec covering core API groups: Auth, Channels, Completions, Health, Metrics, WebSocket Tickets, Admin (stats, health, usage, dashboard).
server/static/swagger.html (new)
Minimal HTML loading Swagger UI v5 from unpkg CDN, pointing at
openapi.yaml. Zero build step.
Routes
GET /api/docs→ serves embeddedswagger.htmlGET /api/docs/openapi.yaml→ serves embeddedopenapi.yaml
Both files embedded via //go:embed directives.
CS3 — Grafana Dashboard + Alerting Rules + Helm
Static files + Helm templates. No Go code changes.
chart/dashboards/switchboard-overview.json (new)
Grafana dashboard with 10 panels:
- Request rate, error rate, latency percentiles
- WebSocket connections, completion rate/latency by provider
- Tokens/min by model, provider status
- DB connection pool, task executions
Template variables: $datasource, $namespace, $pod.
chart/alerting/switchboard-rules.yaml (new)
Source PrometheusRule with 6 alerts:
| Alert | Condition | Severity |
|---|---|---|
| SwitchboardPodRestart | Restarts in 1h | warning |
| SwitchboardProviderDown | Status > 2 for 5m | critical |
| SwitchboardDBPoolExhaustion | In-use/open > 80% for 5m | warning |
| SwitchboardHighErrorRate | 5xx rate > 5% for 5m | warning |
| SwitchboardTaskFailureRate | Error rate > 25% for 10m | warning |
| SwitchboardNoCompletions | Zero completions for 15m | critical |
Helm templates (new, all gated enabled: false by default)
chart/templates/servicemonitor.yaml— ServiceMonitor CRDchart/templates/grafana-dashboard-configmap.yaml— ConfigMap for Grafana sidecar discoverychart/templates/prometheusrule.yaml— PrometheusRule CRD
Helm modifications
chart/values.yaml—loggingsection +monitoringsectionchart/templates/configmap.yaml—LOG_FORMAT,LOG_LEVELenv varschart/templates/services.yaml— Prometheus scrape annotations
CS4 — Admin Observability Dashboard
Built-in admin page for real-time health without Grafana dependency.
server/handlers/dashboard_admin.go (new)
GET /api/v1/admin/dashboard aggregating:
- Provider health summaries (from
HealthStore) - 24h usage totals (from
UsageStore) - DB pool stats (
database.DB.Stats()) - WebSocket connection count (
Hub.ConnCount()) - Recent errors (from
AuditStore) - Process uptime
Frontend
SCAFFOLDING.dashboard— stat cards grid + content containerADMIN_LOADERS.dashboard→UI.loadAdminDashboard()ADMIN_SECTIONS.monitoring—dashboardadded as first section- Dashboard auto-refreshes every 30s
- CSS: provider grid, DB pool bar gauge, error list
Admin template
Monitoring tab now lands on /admin/dashboard instead of /admin/usage.
CS5 — Wiring, docker-compose, ICD Tests
docker-compose.yml
Added LOG_FORMAT and LOG_LEVEL env vars with defaults.
ICD tests (crud/observability.js)
| Test | Assertion |
|---|---|
GET /metrics |
200, contains switchboard_* metrics |
GET /api/docs |
200, contains swagger-ui |
GET /api/docs/openapi.yaml |
200, contains openapi: and paths: |
| X-Request-Id generation | Response includes 36-char UUID header |
| X-Request-Id passthrough | Client-sent header echoed back |
GET /admin/dashboard |
200, has uptime, ws_connections, provider_health |
Verification
LOG_FORMAT=json docker compose up→ JSON log lines on stdoutcurl localhost:8080/metrics | grep switchboard_→ all metrics presentcurl -I localhost:8080/api/v1/channels→X-Request-Idheader/api/docsin browser → Swagger UI renders/admin→ Monitoring → Dashboard shows live stat cardshelm lint chart/ --set monitoring.serviceMonitor.enabled=true→ passes- ICD test suite passes (existing + observability tests)