Changeset 0.6.0 (#36)

This commit is contained in:
2026-02-20 22:24:47 +00:00
parent 30d0c11219
commit 925b70f98c
34 changed files with 2575 additions and 637 deletions

View File

@@ -4,177 +4,376 @@
- [ARCHITECTURE.md](ARCHITECTURE.md) — Core services design (Notes, KBs, Tasks, Channels, Embeddings)
- [EXTENSIONS.md](EXTENSIONS.md) — Extension system spec (Browser/Starlark/Sidecar tiers, tools, surfaces)
## Current State: v0.5.4
**Versioning (pre-1.0):** `0.<major>.<minor>` — hotfixes use quad: `0.x.y.z`
No compatibility guarantees before 1.0. Post-1.0: major = long-lived (break nothing),
minor = add features (deprecate, don't remove), patch = fix something.
---
## Current State: v0.6.2
### ✅ Done
**Backend (Go)**
- [x] PostgreSQL schema + auto-migrations (go:embed, startup)
- [x] JWT auth with refresh token rotation
- [x] Chat CRUD + message persistence
- [x] **Unified channel model** — chats→channels, "everything is a channel"
- [x] Channel CRUD + message persistence (`/api/v1/channels`)
- [x] Message tree (`parent_id`) with linear backfill
- [x] Participant tracking (`participant_type`/`participant_id` on messages)
- [x] `channel_members` + `channel_models` tables (schema foundation)
- [x] `channel_cursors` for branch tracking (schema foundation)
- [x] Folders + Projects tables for organization
- [x] Environment banner system (global_settings, admin presets, position control)
- [x] Streaming completion proxy (OpenAI + Anthropic + Venice + OpenRouter)
- [x] Per-user + global API config management
- [x] Admin endpoints (users, providers, models, settings, stats)
- [x] Admin bootstrap from env vars (K8s secret → upsert on every restart)
- [x] Registration with pending state (admin approval workflow)
- [x] Registration default state setting (active / pending)
- [x] User providers toggle (admin can restrict to global-only providers)
- [x] Bulk model enable/disable
- [x] Public settings endpoint (non-admin users read safe subset)
- [x] Rate limiting, error middleware, request logging
- [x] Health check endpoint (includes schema_version)
- [x] EventBus with WebSocket hub
- [x] EventBus with WebSocket hub (rooms, JWT auth, heartbeat, reconnection)
- [x] Provider capability system (known models, heuristic detection, resolution chain)
- [x] Dynamic max_tokens resolution (no more hardcoded 4096)
- [x] Dynamic max_tokens resolution
- [x] User + admin provider model listing with capabilities
**Frontend (Vanilla JS)**
- [x] Professional splash page (split-panel hero + tabbed auth)
- [x] Split "New Chat" button with dropdown (Group Chat, Channel — coming soon)
- [x] Environment banner system (CSS custom props + JS init from settings)
- [x] Collapsible sidebar with time-grouped chat history
- [x] User menu flyout (Settings, Admin, Debug, Sign Out)
- [x] Model selector with capability badges (output, context, tools, vision, thinking)
- [x] Client-side known model table with backend overlay
- [x] Settings modal with tabs (General, Providers, Models)
- [x] Admin modal with tabs (Users, Providers, Models, Settings, Stats)
- [x] Admin settings: registration, user providers, banner config with presets
- [x] Pending user badge + approve workflow in admin Users tab
- [x] Streaming SSE display with smart scroll
- [x] Full markdown rendering (marked.js + DOMPurify)
- [x] Full markdown rendering (marked.js + DOMPurify, vendor + CDN fallback)
- [x] Thinking block display (`<think>`/`<thinking>` tags)
- [x] Settings modal (profile, chat prefs, provider management)
- [x] Admin modal (users, providers, models with cap badges, settings, stats)
- [x] Debug modal (console intercept, network log, state inspector)
- [x] EventBus client with exponential backoff + max retries
- [x] Export (Markdown, JSON, Text)
- [x] Vendor libs with CDN fallback (SCIF-safe)
- [x] 3-stage Dockerfile (Go build → vendor download → nginx)
**CI/CD (Gitea Actions)**
- [x] Three-env pipeline (dev/test/prod)
- [x] Dev: upgrade test → schema validate → wipe → fresh install test
- [x] Backend auto-migrates on startup (no CI migration step)
- [x] Admin secret sync (Gitea secrets → K8s secret → env vars)
- [x] Shared PG safety (_dev suffix guard on wipe)
- [x] Post-deploy schema verification via /api/v1/health
**Architecture Design (v0.5.4)**
**Architecture Design**
- [x] ARCHITECTURE.md: core services spec (13 sections)
- [x] EXTENSIONS.md: three-tier extension system spec
- [x] Terminology: Roles, Teams, Channels, Message Tree, etc.
- [x] "Everything is a channel" — chat is a channel with type:'direct'
- [x] Conversation forking (message tree via parent_id)
- [x] Auth strategy (builtin/mTLS/OIDC)
- [x] Classification banners (CSS custom props)
### Implementation Phasing
**See ARCHITECTURE.md §10 for the authoritative implementation sequence.**
Phases 18, from channel foundation through RBAC.
- [x] "Everything is a channel" — unified model with type:'direct'
- [x] Message tree (parent_id for conversation forking)
- [x] Auth strategy design (builtin/mTLS/OIDC)
---
## Phase 1: Real-Time Foundation
## 0.7.0 — Custom Models / Presets
### 1.1 WebSocket Hub
**Priority:** HIGH — blocks Channels, typing indicators, live updates
Named wrappers around base models with bundled configuration. Admins create
org-wide presets, users create personal ones (if user providers are enabled).
- [ ] WebSocket upgrade endpoint (`/ws`)
- [ ] Connection manager (register, unregister, broadcast)
- [ ] Room-based pub/sub (per-chat, per-channel, global)
- [ ] JWT auth on WebSocket handshake
- [ ] Heartbeat / ping-pong keepalive
- [ ] Reconnection handling (frontend)
- [ ] Message types: `chat.message`, `chat.typing`, `user.presence`, `system.notify`
- [ ] `model_presets` table: name, base_model_id, system_prompt, temperature,
max_tokens, tools_enabled (jsonb), created_by, scope (global/team/personal),
team_id (nullable, for future team scoping), is_shared
- [ ] Permission gating: personal presets require user_providers_enabled
- [ ] Admin preset management UI (create, edit, delete org-wide presets)
- [ ] User preset management UI in Settings (personal presets)
- [ ] Presets appear as first-class entries in model selector
("Code Reviewer (GPT-4o)", "Research Assistant (Claude Opus)")
- [ ] Completion handler unwraps preset → base model + config overrides
### 1.2 Live Chat Updates
- [ ] New messages pushed via WebSocket (not just SSE streaming)
- [ ] Typing indicators
- [ ] Chat list updates when other sessions create/delete chats
- [ ] Online user presence
## 0.7.x — Branding + Polish
White-label support and UX improvements that exploit existing infrastructure.
**Branding**
- [ ] Admin branding settings in global_settings: org name, tagline, accent color
- [ ] `/branding/` volume mount (K8s ConfigMap, optional) for favicon, logo
- [ ] Splash page reads branding config on load, falls back to Switchboard defaults
- [ ] CSS accent color override from branding settings
**Message Editing + Forking**
- [ ] Edit message → creates sibling (uses existing parent_id tree)
- [ ] Regenerate → creates sibling model response
- [ ] Branch indicator at fork points (← 1/2 →)
- [ ] Context assembly follows active path, not full channel history
**UX Polish**
- [ ] Chat search / filter in sidebar
- [ ] UI preferences (theme toggle, font size — stored in user settings)
- [ ] Keyboard shortcuts (Ctrl+K command palette)
- [ ] PWA manifest + offline shell + install prompt
---
## Phase 2: Channels
## 0.8.0 — Teams + Onboarding
### 2.1 Backend
- [ ] Channel model (name, description, type: public/private/DM, creator)
- [ ] Membership model (user_id, channel_id, role: owner/member)
- [ ] Channel messages (extends message model with channel_id)
- [ ] CRUD endpoints: `/api/v1/channels`, `/api/v1/channels/:id/messages`
- [ ] @mention parsing — users and AI models
- [ ] AI response triggered by @model-name mentions
- [ ] WebSocket broadcast per channel room
The missing middle tier: scoped administration without system-admin access.
### 2.2 Frontend
- [ ] Channels section in sidebar (below chats, collapsible)
- [ ] Channel creation modal (name, description, public/private)
- [ ] Member management (invite, remove, role change)
- [ ] Message display with user avatars and @mention highlighting
- [ ] AI responses inline with user messages
- [ ] Unread count badges
**Schema**
- [ ] `teams` table: id, name, description, created_by
- [ ] `team_members` table: team_id, user_id, role ('admin' | 'member')
- [ ] Add `team_id` (nullable) to: model_presets, channels, (future: notes, KBs)
- [ ] Team admin role: scoped per-team, not system-wide
(one person can be admin of Team A, member of Team B)
**Onboarding**
- [ ] Registration approval → assign role + team(s) during approval
(upgrade from current binary approve → assign-and-activate)
- [ ] System admin creates teams + assigns first team admin
- [ ] Team admin self-manages: add/remove members, create team presets,
manage team channels, view team usage
**Private Provider Policy**
- [ ] `is_private` flag on provider configs (marks local/self-hosted endpoints)
- [ ] `require_private_providers` policy per team
- [ ] Completion handler enforces policy — team members restricted to private providers
- [ ] Enables HIPAA/compliance posture: provable data boundary per team
## 0.8.x — Audit + Usage Tracking
Required for enterprise and compliance. Cheap to build, expensive to retrofit.
**Audit Log**
- [ ] `audit_log` table: actor_id, action, resource_type, resource_id,
metadata (jsonb), ip_address, timestamp
- [ ] Every mutating handler inserts audit entry
- [ ] Admin audit viewer (filter by actor, action, resource, date range)
- [ ] Team admin sees audit entries scoped to their team
**Usage / Cost Tracking**
- [ ] Capture from provider responses: prompt_tokens, completion_tokens,
cache_creation_tokens, cache_read_tokens
- [ ] `usage_log` table: channel_id, user_id, model_id, provider_id,
token counts, timestamp
- [ ] Model pricing fields on `model_configs`:
cost_input, cost_output, cost_cache_input, cost_cache_output
(per M tokens, nullable, decimal)
- [ ] `cost_source` field: 'provider' | 'admin' | null
— provider APIs populate on model fetch, admin can override,
source tracking prevents silent overwrites on re-fetch
- [ ] Cost calculated at query time (join usage × pricing)
— never store computed cost; pricing changes and historical
token counts should recalculate against current rates
- [ ] Per-user and per-team usage stats in admin panel
- [ ] Team admins see usage for their team scope
---
## Phase 3: Plugin System
## 0.9.0 — Tool Execution + Notes
### 3.1 Architecture (see discussion below)
- [ ] Plugin manifest format (`plugin.json`)
- [ ] Plugin lifecycle (install, enable, disable, uninstall)
- [ ] Plugin storage (per-plugin isolated DB namespace or key-value)
- [ ] Hook system (pre-completion, post-completion, on-message, on-channel-message)
- [ ] Admin UI for plugin management (install, configure, enable/disable)
- [ ] Plugin API (what plugins can access: messages, user context, settings)
The tool-calling infrastructure that everything downstream depends on.
### 3.2 Built-in Plugin: Web Search
- [ ] Tool-use integration in completion flow
- [ ] Search provider abstraction (DuckDuckGo, SearXNG, Brave)
- [ ] Results injected into context
**Tool Framework**
- [ ] Tool calling pipeline in completion handler
- [ ] Tool registry (built-in + future plugin tools)
- [ ] Tool execution loop: model requests tool → backend executes → result fed back
- [ ] Tool permission model (which tools enabled per preset/channel)
### 3.3 Built-in Plugin: Token Counter
- [ ] Per-message token estimation
- [ ] Per-chat cost tracking
- [ ] Provider-specific pricing data
---
## Phase 4: Notes & Knowledge Base
### 4.1 Notes
- [ ] Note model (title, content, folder, tags)
- [ ] CRUD endpoints
- [ ] Folder organization
**Notes**
- [ ] `notes` table: title, content, folder_id, tags, team_id (nullable), created_by
- [ ] Notes CRUD endpoints
- [ ] `note_create`, `note_update`, `note_search`, `note_list` tools
- [ ] Full-text search (PostgreSQL `tsvector`)
- [ ] Markdown editor in frontend
- [ ] Link notes to chats
- [ ] Team-scoped notes (schema ready from 0.8.0)
### 4.2 RAG (Retrieval Augmented Generation)
- [ ] Document upload (PDF, DOCX, TXT, MD)
- [ ] Chunking strategies (recursive, semantic)
- [ ] Embedding generation (OpenAI, local models)
- [ ] pgvector storage and similarity search
**Conversation Forking UI**
- [ ] Edit-and-resubmit creates siblings (tree structure from 0.6.0 schema)
- [ ] Branch indicator ← 1/2 → at fork points
- [ ] Context assembly follows active path
## 0.9.x — Context Management
Usability before compaction exists — long conversations shouldn't silently fail.
- [ ] Token counting on outbound requests (estimate before sending)
- [ ] Truncation strategy (sliding window or drop oldest, configurable)
- [ ] "Conversation is getting long" warning in UI
- [ ] Manual "summarize and continue" action (user-triggered pre-compaction)
- [ ] Groundwork for automated compaction in 0.14.0
---
## 0.10.0 — Web Search + URL Fetch
First real external tools, built on 0.9.0 framework.
- [ ] `web_search` tool: search provider abstraction (DuckDuckGo, SearXNG, Brave)
- [ ] `url_fetch` tool: retrieve and extract content from URLs
- [ ] Sidecar or direct HTTP from backend (configurable)
- [ ] Results injected into context
- [ ] Paired with Notes: "research X and save findings to my notes"
---
## 0.11.0 — File Handling + Vision
File input into chat — table stakes for serious use. Blobs live in object
storage, metadata and search indexes live in PostgreSQL.
**Storage Backend Abstraction**
- [ ] S3-compatible API as primary interface
(MinIO, Ceph RGW, AWS S3, GCS — anything with S3 API)
- [ ] Local PVC as zero-config fallback (single-node / dev)
- [ ] Admin config: storage backend selection, endpoint, credentials,
bucket/path, per-file and total size limits
- [ ] Reused by 0.13.0 (KB documents) and 0.14.0 (compaction snapshots)
**Metadata + Search Index (PostgreSQL)**
- [ ] `attachments` table (metadata only, never blobs):
id, message_id, channel_id, team_id, filename, mime_type,
size_bytes, storage_backend ('s3' | 'pvc'), storage_path,
checksum_sha256, uploaded_by, created_at,
search_text (tsvector, nullable)
- [ ] Text extraction on upload (PDF, DOCX, TXT, MD → tsvector)
- [ ] Full-text search across attachment content
- [ ] Access control via JOIN: attachments → channels → team_members
**Chat Integration**
- [ ] Image/file upload in chat messages
- [ ] Multimodal message assembly for vision-capable models
- [ ] Document preview in chat (images inline, files as download links)
- [ ] Paste-to-upload (clipboard image support)
*Note: media generation (image gen, video models) is a separate concern —
those are tool-use actions that depend on the tool framework (0.9.0) and
produce attachments as output. Tracked under Future.*
---
## 0.12.0 — @mention Routing + Multi-model
The channel schema exists from 0.6.0. This phase adds the routing logic.
- [ ] @mention parsing in messages (users and AI models)
- [ ] Resolve mentions against `channel_models`
- [ ] Multi-model channels: fire completions per mentioned model
- [ ] "Add a model" UI per channel
- [ ] Enables: editor mode, second opinions, cross-model conversations
---
## 0.13.0 — Embeddings + Knowledge Bases
- [ ] Embedding pipeline: chunking (recursive, semantic), generation (OpenAI, local)
- [ ] `pgvector` storage and similarity search
- [ ] `knowledge_bases` table: name, description, team_id (nullable), created_by
- [ ] KB document storage via 0.11.0 storage backend (same S3/PVC abstraction)
- [ ] KB CRUD endpoints + admin UI
- [ ] `kb_search` tool (built on 0.9.0 tool framework)
- [ ] Context injection in completion flow
- [ ] Per-KB toggle in chat settings
- [ ] Notes get embedded too (once pipeline exists)
- [ ] Team admins manage team KBs (permission layer from 0.8.0)
- [ ] Per-channel KB toggle
---
## Phase 5: Polish
## 0.14.0 — Compaction
- [ ] Chat search / filter in sidebar
- [ ] Message editing
- [ ] Chat folders / pinning
- [ ] Bulk chat operations
- [ ] Keyboard shortcuts (Ctrl+K command palette)
- [ ] PWA manifest + offline shell
- [ ] Performance: lazy load chat history, virtual scroll for long chats
Replaces the manual 0.9.x context management with automated background processing.
- [ ] Auto-compaction service: background job that calls an LLM to summarize
- [ ] Channel-scoped: triggers when channel exceeds context threshold
- [ ] Compaction summaries stored as system messages in the channel
- [ ] Configurable: per-channel opt-in/out, summary model selection
- [ ] Admin controls for resource limits
---
## Future
## 0.15.0 — Smart Model Routing
Rules-based routing engine — not ML, just policy.
- [ ] Admin defines routing policies:
"cheapest model with required capabilities",
"prefer private providers, fallback to cloud",
"for this team, use X; for that team, use Y"
- [ ] Model capability system already exists — routing is policy on top
- [ ] Fallback chains: primary provider down → next provider with same model
- [ ] Cost-aware: factor pricing into routing decisions
- [ ] Latency-aware: track response times per provider, prefer faster
---
## 0.16.0 — Tasks / Autonomous Agents
The capstone: everything below it combined into autonomous workflows.
- [ ] Scheduler + task runner
- [ ] `task_create` tool
- [ ] Creates `type: 'service'` channels with no human members
- [ ] Depends on: completion handler, tool execution, notes, web search
- [ ] Admin controls for resource limits, execution budgets
---
## 0.17.0 — Auth Strategy (mTLS/OIDC) + Full RBAC
Enterprise auth modes and fine-grained permissions on top of the teams foundation.
- [ ] `AUTH_MODE` env var: `builtin` | `mtls` | `oidc`
- [ ] All three resolve to the same internal user model
- [ ] `auth_source` + `external_id` columns on users table
- [ ] mTLS: header trust, auto-provision from cert DN
- [ ] OIDC: Keycloak/Okta token validation, claim extraction, role mapping
- [ ] WebSocket auth per mode
- [ ] Per-source auto-activate policy:
auto_activate (bool), default_team, default_role
- [ ] Fine-grained permissions: model access, KB write, task create,
admin delegation, token budgets per user/team
- [ ] SSO/SAML (may fold in or follow as 0.17.x)
---
## Future (post-1.0 candidates)
Items that are real but don't yet have a version assignment. Any of these
could pull left based on need.
**Desktop + Mobile**
- Desktop app (Tauri)
- Smart model routing (cost/quality/latency auto-selection)
- Workflow builder (visual DAG for chaining models + tools)
- Plugin marketplace
- SSO/SAML
- Audit logging
- Full PWA with offline capability
- Mobile-optimized layouts
**Generation + Media**
- Image generation tools (DALL-E, Stable Diffusion endpoints)
- Video model integration
- Audio/TTS tools
**Data + Portability**
- Bulk export/import (account data, conversations, settings)
- ChatGPT/other tool import
- GDPR-style "download my data"
- Backup/restore CronJob manifests + operational docs
**Platform**
- Rate limiting per user/team/tier (token budgets)
- Provider health monitoring + key rotation
- Multi-tenant SaaS mode
- Workflow builder (visual DAG for chaining models + tools)
- Plugin/extension marketplace
- Live collaboration (typing indicators, presence, co-editing)
- Virtual scroll for long conversations
---
## Plugin / Extension Architecture
## Extension / Plugin Architecture
**Moved to dedicated documents:**
- [EXTENSIONS.md](EXTENSIONS.md) — Full extension system spec with three tiers
(Browser JS, Starlark sandbox, Sidecar containers), manifest format,
browser tool bridge, surface/mode system, and implementation roadmap.
**Deferred to post-1.0.** The tool execution framework (0.9.0) provides the
internal hook points. A formal plugin API, manifest format, and marketplace
are tracked in the dedicated design documents:
- [EXTENSIONS.md](EXTENSIONS.md) — Three-tier extension system spec
(Browser JS, Starlark sandbox, Sidecar containers)
- [ARCHITECTURE.md](ARCHITECTURE.md) — Core backend services that extensions
build on: Notes, Knowledge Bases, Tasks, Channels, Embeddings, and
Folders/Projects. Includes data models, sequencing, and admin controls.
build on