Changeset 0.15.0 (#71)

This commit is contained in:
2026-02-26 21:19:55 +00:00
parent e2149e249d
commit 2d79ff593b
35 changed files with 3218 additions and 504 deletions

584
docs/DESIGN-0.15.0.md Normal file
View File

@@ -0,0 +1,584 @@
# DESIGN-0.15.0 — Compaction
## Overview
Automatic conversation compaction. A background service monitors channels
for context pressure and triggers summarization via the utility model role,
replacing the manual "Summarize & Continue" button as the primary
compaction path. Manual summarization remains available as an explicit
user action.
Depends on: utility model role (v0.10.0), summarize handler (v0.10.2).
**Design principle: don't reinvent the wheel.** The existing
`SummarizeHandler` already has the full pipeline — path loading, summary
boundary detection, prompt construction, role resolution, tree insertion,
cursor update, usage logging. This design extracts that core logic into
a shared `compaction` package and adds a background scanner on top.
---
## 1. Extract: `compaction` Package
Move the reusable summarization pipeline out of `handlers/summarize.go`
into `server/compaction/compaction.go`. The HTTP handler becomes a thin
wrapper.
### Core type
```go
package compaction
// Service provides conversation compaction (summarization).
// Used by both the HTTP handler (manual) and the background scanner (auto).
type Service struct {
stores store.Stores
resolver *roles.Resolver
}
func NewService(stores store.Stores, resolver *roles.Resolver) *Service {
return &Service{stores: stores, resolver: resolver}
}
```
### Extracted method: `Compact`
The current `SummarizeHandler.Summarize` body becomes `Service.Compact`:
```go
type CompactRequest struct {
ChannelID string
UserID string // channel owner
TeamID *string // for role resolution
Trigger string // "manual" | "auto"
}
type CompactResult struct {
SummaryID string
SummarizedCount int
Model string
UsedFallback bool
Content string
InputTokens int
OutputTokens int
}
func (s *Service) Compact(ctx context.Context, req CompactRequest) (*CompactResult, error)
```
The method performs the same steps as today's handler:
1. Load active path via `getActivePath` (already exported within `handlers`)
2. Find summary boundary, collect messages to summarize
3. Guard: minimum 4 messages since last summary
4. Build summarization prompt
5. Call `resolver.Complete(ctx, RoleUtility, ...)`
6. Insert summary tree node with metadata
7. Update cursor
8. Log usage
The only new metadata field is `"trigger"` (`"manual"` or `"auto"`) so
the frontend can distinguish how the summary was created.
### Refactored handler
`handlers/summarize.go` shrinks to:
```go
func (h *SummarizeHandler) Summarize(c *gin.Context) {
// ownership check, rate limit check (unchanged)
// ...
result, err := h.compaction.Compact(c.Request.Context(), compaction.CompactRequest{
ChannelID: channelID,
UserID: userID,
TeamID: teamID,
Trigger: "manual",
})
// return JSON response (unchanged)
}
```
### Tree helpers
`getActivePath`, `isSummaryMessage`, `nextSiblingIndex`, `updateCursor`
are currently unexported in `handlers/tree.go`. Two options:
**Option A** — Export them from `handlers` and import in `compaction`.
Clean but creates a dependency from `compaction``handlers`.
**Option B** — Move tree helpers into a `treepath` (or similar) package
imported by both `handlers` and `compaction`.
Recommend **Option B** to keep the dependency graph clean:
`handlers``compaction``treepath``handlers`. The `treepath`
package is pure data + SQL queries with no handler logic.
```
server/
treepath/ ← NEW: extracted from handlers/tree.go
path.go (getActivePath, getPathToLeaf, getActiveLeaf)
summary.go (isSummaryMessage, summary metadata helpers)
siblings.go (getSiblingCount, getSiblings, findLeafFromMessage)
cursor.go (updateCursor, nextSiblingIndex)
compaction/ ← NEW
compaction.go (Service, Compact)
scanner.go (Scanner, background loop)
estimator.go (token estimation)
handlers/
tree.go → thin wrappers or deleted, imports treepath
summarize.go → thin HTTP wrapper, delegates to compaction.Service
completion.go → imports treepath for path building
```
---
## 2. Server-Side Token Estimation
The background scanner needs to evaluate context pressure without making
an LLM call. The frontend already does this with a `chars / 4` heuristic
(`tokens.js`). Replicate server-side:
```go
// compaction/estimator.go
// EstimateTokens returns a rough token count using the ~4 chars/token
// heuristic. Matches the frontend Tokens.estimate() for consistency.
func EstimateTokens(text string) int {
return (len(text) + 3) / 4
}
// EstimatePath returns total estimated tokens for a message path,
// including system prompt overhead.
func EstimatePath(path []treepath.PathMessage, systemPromptLen int) int {
total := 0
if systemPromptLen > 0 {
total += EstimateTokens(string(make([]byte, systemPromptLen))) + 4
}
for _, m := range path {
total += EstimateTokens(m.Content) + 4 // +4 per-message overhead
}
return total
}
```
Not exact, but it's the same bar the user already sees in the UI. If we
need better accuracy later, swap in a tokenizer library without changing
the interface.
### Context budget resolution
The scanner needs `max_context` for whatever model the channel is using.
Resolution chain:
1. Channel has a `model` field → look up `model_catalog` for capabilities
2. Channel has a `provider_config_id` → use that provider's catalog entry
3. Fallback: use the utility role's model context window as a proxy (it's
the model that would be doing the summarization anyway)
4. Hard fallback: 128K tokens (conservative default)
```go
func (s *Scanner) getContextBudget(ctx context.Context, ch *models.Channel) int {
if ch.Model != "" {
if caps, err := s.stores.Catalog.GetCapabilities(ctx, ch.Model); err == nil {
if caps.MaxContext > 0 {
return caps.MaxContext
}
}
}
return DefaultContextBudget // 128000
}
```
---
## 3. Background Scanner
Follows the Ingester pattern from `knowledge/ingest.go`: goroutine pool
with semaphore, WaitGroup for graceful shutdown.
### Scanner type
```go
// compaction/scanner.go
type Scanner struct {
service *Service
stores store.Stores
sem chan struct{}
wg sync.WaitGroup
stopCh chan struct{}
ticker *time.Ticker
}
type ScannerConfig struct {
Enabled bool // global kill switch
Interval time.Duration // scan frequency (default: 5 minutes)
Concurrency int // max parallel compactions (default: 2)
}
func NewScanner(svc *Service, stores store.Stores, cfg ScannerConfig) *Scanner
func (sc *Scanner) Start()
func (sc *Scanner) Stop() // signals stop + drains in-flight
```
### Scan loop
```
every <interval>:
if !globalEnabled(): continue
candidates = findCandidates()
for each candidate:
sem.acquire()
go compact(candidate)
```
### Candidate query
A single SQL query finds channels that need compaction:
```sql
SELECT c.id, c.user_id, c.model, c.settings,
COUNT(m.id) as msg_count,
SUM(LENGTH(m.content)) as total_chars
FROM channels c
JOIN messages m ON m.channel_id = c.id
AND m.deleted_at IS NULL
WHERE c.deleted_at IS NULL
AND c.type = 'direct' -- only user conversations (not groups yet)
AND c.is_archived = false
-- Skip channels with very recent activity (let the user finish typing)
AND c.updated_at < NOW() - INTERVAL '2 minutes'
-- Only channels active in the last 7 days (don't compact stale channels)
AND c.updated_at > NOW() - INTERVAL '7 days'
GROUP BY c.id
HAVING COUNT(m.id) >= 10 -- minimum message count
AND SUM(LENGTH(m.content)) > 20000 -- rough: 20K chars ≈ 5K tokens minimum
ORDER BY c.updated_at DESC
LIMIT 50 -- batch size cap
```
This is a coarse filter. Each candidate then gets a precise check:
```go
func (sc *Scanner) shouldCompact(ctx context.Context, ch *models.Channel) bool {
// 1. Check channel-level opt-out
if ch.Settings != nil {
if v, ok := ch.Settings["auto_compaction"]; ok && v == false {
return false
}
}
// 2. Load active path for the channel owner
path, err := treepath.GetActivePath(ch.ID, ch.UserID)
if err != nil || len(path) < 10 {
return false
}
// 3. Find post-summary messages only
startIdx := 0
for i, m := range path {
if treepath.IsSummaryMessage(&m) {
startIdx = i + 1
}
}
postSummary := path[startIdx:]
if len(postSummary) < 8 {
return false // not enough new messages since last summary
}
// 4. Estimate tokens and check threshold
tokens := EstimatePath(postSummary, 0)
budget := sc.getContextBudget(ctx, ch)
threshold := sc.getThreshold(ch) // default 0.70
return float64(tokens) / float64(budget) >= threshold
}
```
### Threshold
Default: **0.70** (70% of context window). This is intentionally lower
than the frontend's 75% warning bar — auto-compaction should fire before
the user sees a warning, not after.
Configurable per-channel via `settings.compaction_threshold` (float) and
globally via `global_settings` key `auto_compaction_threshold`.
### Rate limiting
Auto-compaction calls use the utility role and are org-funded. They share
the existing `utility_rate_limit` budget but get a separate counter
prefix so admins can see auto vs manual usage:
```go
entry := &models.UsageEntry{
// ...
Role: &role, // "utility"
Metadata: models.JSONMap{
"trigger": "auto", // distinguishes from manual
},
}
```
If the global utility rate limit is hit, auto-compaction silently skips
the channel and retries next scan cycle. It never blocks or errors.
---
## 4. Per-Channel Configuration
Uses the existing `Channel.Settings` JSONB column. No migration needed.
| Key | Type | Default | Description |
|-----|------|---------|-------------|
| `auto_compaction` | `bool` | `true` (from global) | Enable/disable for this channel |
| `compaction_threshold` | `float` | `0.70` (from global) | Context usage ratio to trigger |
### Frontend: channel settings
Add a "Compaction" section to the channel settings panel (the gear icon
in the chat header). Two controls:
```
┌─ Compaction ──────────────────────────────────┐
│ Auto-compact [✓] │
│ Threshold [70]% of context window │
└───────────────────────────────────────────────┘
```
The manual "Summarize & Continue" button remains in the context warning
bar and works regardless of auto-compaction settings.
---
## 5. Admin Controls
New `global_settings` keys:
| Key | Type | Default | Description |
|-----|------|---------|-------------|
| `auto_compaction_enabled` | `bool` | `false` | Global kill switch. Ships **off** — admins opt in. |
| `auto_compaction_threshold` | `float` | `0.70` | Default threshold for channels without override |
| `auto_compaction_interval_minutes` | `int` | `5` | Scanner tick interval |
| `auto_compaction_concurrency` | `int` | `2` | Max parallel auto-compactions |
| `auto_compaction_cooldown_minutes` | `int` | `30` | Min time between auto-compactions of the same channel |
### Admin UI
Add an "Auto-Compaction" card to the existing admin AI settings panel
(alongside the utility rate limit and model roles):
```
┌─ Auto-Compaction ─────────────────────────────┐
│ Enabled [✓] │
│ Threshold [70]% │
│ Scan interval [5] minutes │
│ Concurrency [2] parallel │
│ Cooldown [30] minutes per channel │
└───────────────────────────────────────────────┘
```
---
## 6. Cooldown + Dedup
Prevent rapid re-compaction of the same channel:
1. **Cooldown** — after compacting channel X, don't re-evaluate it for
`cooldown_minutes`. Track in-memory:
```go
lastCompacted map[string]time.Time // channelID → timestamp
```
2. **In-flight dedup** — if channel X is currently being compacted (sem
acquired, goroutine running), skip it in the current scan. Track via:
```go
inFlight sync.Map // channelID → struct{}
```
3. **Stacking** — the existing summary boundary logic already handles
multiple summaries correctly. Auto-compaction produces the same tree
nodes as manual, so stacking "just works."
---
## 7. Lifecycle + Wiring
### main.go changes
```go
// After role resolver and stores setup:
compactionSvc := compaction.NewService(stores, roleResolver)
compactionScanner := compaction.NewScanner(compactionSvc, stores, compaction.ScannerConfig{
Enabled: getGlobalBool(stores, "auto_compaction_enabled", false),
Interval: time.Duration(getGlobalInt(stores, "auto_compaction_interval_minutes", 5)) * time.Minute,
Concurrency: getGlobalInt(stores, "auto_compaction_concurrency", 2),
})
compactionScanner.Start()
defer compactionScanner.Stop() // drain in-flight on shutdown
// Summarize handler now wraps compaction service:
summarize := handlers.NewSummarizeHandler(stores, roleResolver, compactionSvc)
```
Follows the same pattern as `kbIngester` with `defer kbIngester.Wait()`.
### Config reload
The scanner re-reads global settings at the start of each scan cycle.
No restart required to change thresholds, intervals, or the kill switch.
If `auto_compaction_enabled` flips to `false`, the scanner skips all
work on the next tick.
---
## 8. Observability
### Logging
```
🔍 compaction: scanning 50 candidates
🔍 compaction: channel abc123 qualifies (78% of 128K context, 42 messages post-summary)
📝 compaction: compacting channel abc123 for user xyz (trigger=auto)
✅ compaction: channel abc123 done (42 messages → 847 chars, model=claude-sonnet-4-5-20250929)
⏭ compaction: channel def456 skipped (cooldown, 12m remaining)
⏭ compaction: channel ghi789 skipped (rate limit)
```
### Audit log
Auto-compaction events are logged to `audit_log`:
```go
stores.Audit.Log(ctx, &models.AuditEntry{
Actor: "system:compaction",
Action: "compaction.auto",
Resource: "channel:" + channelID,
Details: models.JSONMap{
"summarized_count": result.SummarizedCount,
"model": result.Model,
"trigger": "auto",
},
})
```
### Usage tracking
All auto-compaction calls flow through the existing usage logging path
(same as manual summarization). The `metadata.trigger = "auto"` field
lets the admin usage dashboard distinguish auto from manual.
---
## 9. Frontend Changes
Minimal — the backend does the heavy lifting.
### Auto-compaction indicator
When loading a conversation that was auto-compacted, the frontend
already handles summary boundaries correctly (the collapsed history
toggle, the "X earlier messages summarized" bar). The only new UI is
a subtle indicator on the summary node:
```
▸ 42 earlier messages summarized (auto)
```
The `(auto)` label comes from `metadata.trigger` on the summary message.
### Channel settings
Add compaction controls to the channel settings panel (Section 4 above).
Saves to `PATCH /channels/:id` with `settings` field — no new endpoint.
### Admin panel
Add the auto-compaction card to the AI settings section (Section 5).
Saves via existing `PUT /admin/settings/:key` endpoints.
---
## 10. Migration
**No new migration required.**
- Channel compaction settings → existing `Channel.Settings` JSONB
- Admin settings → existing `global_settings` table
- Summary messages → existing `messages` table with metadata
- Usage tracking → existing `usage_log` table
- Audit logging → existing `audit_log` table
---
## 11. File Manifest
```
server/
treepath/ ← NEW package
path.go (extracted from handlers/tree.go)
summary.go (isSummaryMessage + helpers)
siblings.go (sibling queries)
cursor.go (cursor + sibling index)
compaction/ ← NEW package
compaction.go (Service, Compact, CompactRequest/Result)
scanner.go (Scanner, scan loop, candidate query)
estimator.go (EstimateTokens, EstimatePath)
handlers/
tree.go (updated: delegates to treepath)
summarize.go (updated: delegates to compaction.Service)
completion.go (updated: imports treepath)
main.go (updated: wire compaction scanner)
src/js/
ui-core.js (updated: show "(auto)" on auto-compacted summaries)
ui-settings.js (updated: channel compaction controls)
ui-admin.js (updated: auto-compaction admin card)
```
---
## 12. Implementation Order
**Phase 1: Extract** — Move tree helpers to `treepath`, core
summarization to `compaction.Service`. Refactor `SummarizeHandler` to
delegate. All existing tests must still pass. No new behavior.
**Phase 2: Estimator + Scanner** — Add `estimator.go`, `scanner.go`,
candidate query, `shouldCompact` logic. Wire into `main.go` with
`Start()`/`Stop()`. Global settings for admin controls. Ship with
`auto_compaction_enabled` defaulting to `false`.
**Phase 3: Frontend** — Channel compaction settings, admin panel card,
auto-compaction indicator on summary nodes.
**Phase 4: Soak + tune** — Enable on a test instance, watch logs, tune
default threshold and cooldown. Adjust candidate query filters if scan
is too expensive.
---
## 13. What This Intentionally Doesn't Do
- **Token-accurate estimation** — chars/4 is good enough. A real
tokenizer adds a dependency for marginal accuracy gain. Revisit if
users report summaries firing too early or too late.
- **Per-model prompt tuning** — the summarization prompt is one-size-fits
-all. Works fine across Claude/GPT/Llama. Revisit if quality varies
significantly across providers.
- **Group channel compaction** — auto-compaction targets `direct` channels
only (single-owner). Group/channel types need multi-cursor awareness
(which user's path to compact?). Deferred to v0.19.0 multi-participant.
- **Compaction snapshots for KB** — the roadmap mentions this. The
summary nodes we create here are already stored as messages with full
metadata. KB integration can query them later without changes to this
design.
- **Event-driven triggers** — a post-completion hook that checks
immediately after each message would reduce latency vs. the ticker.
Worth adding later but the ticker is simpler to ship and debug. The
2-minute activity gap in the candidate query achieves a similar effect
(don't compact mid-conversation).