This repository has been archived on 2026-04-03. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
core/docs/DESIGN-0.14.0.md
2026-02-26 21:19:55 +00:00

33 KiB

DESIGN-0.14.0 — Knowledge Bases

Overview

RAG (Retrieval-Augmented Generation) for Chat Switchboard. Users upload documents into named knowledge bases, the backend chunks and embeds them via the embedding model role (v0.10.0), stores vectors in pgvector, and a kb_search tool lets the LLM pull relevant context at completion time.

Depends on: embedding model role (v0.10.0), file storage (v0.12.0), tool framework (v0.11.0), admin panel (v0.13.0).


0. Pre-Req Fix: Embedding Model Selection

The embedding role dropdown is broken — filterModels() in ui-primitives.js filters by model_type === 'embedding', but many providers don't report a type for embedding models (defaults to 'chat' and gets filtered out). This blocks KB setup entirely.

Fix (ship before or with Phase 1):

  1. Manual model ID entry — add a text input fallback alongside the dropdown. If the dropdown shows no results for the selected provider, display an editable text field pre-populated with the model_id from the saved binding (if any). Users type the model ID directly (e.g. text-embedding-3-small). The save handler accepts either dropdown selection or manual text entry.

  2. Tolerant type filter — change filterModels() to include models where model_type is empty/null/undefined when the role's type filter is 'embedding'. Embedding models are rare enough that showing untyped models alongside typed ones is better than showing nothing.

// Before (strict — breaks when model_type missing)
(m.model_type || 'chat') === typeFilter

// After (tolerant for embedding role)
typeFilter === 'embedding'
    ? (!m.model_type || m.model_type === 'embedding')
    : (m.model_type || 'chat') === typeFilter
  1. Combo UI — the slot renderer gets a hybrid dropdown+input:
┌──────────────────────────────────────────────┐
│ Provider: [OpenAI           ▾]               │
│ Model:    [text-embedding-3-small    ▾] [✎]  │
│                                    ↑ toggle  │
│           [text-embedding-3-small       ]    │ ← manual entry (shown on ✎ click)
└──────────────────────────────────────────────┘

The pencil icon toggles between dropdown and text input. If the dropdown has zero options after provider change, auto-switch to manual entry. Saved bindings with model IDs not in the dropdown auto-show in manual mode on load.


1. Schema — Migration 008_knowledge_bases.sql

pgvector Extension

CREATE EXTENSION IF NOT EXISTS vector;

Requires pgvector installed on the PostgreSQL server. The migration checks for the extension and fails with a clear error if unavailable, rather than silently degrading. The README/deployment docs will note the dependency.

Dimension handling: Embedding dimensions vary by model (OpenAI ada-002 = 1536, text-embedding-3-small = 1536, text-embedding-3-large = 3072, many open-source = 768 or 1024). The dimension is stored per KB and the vector column uses the max supported (3072), with vectors zero-padded or truncated at insert time. Alternative: use halfvec for storage efficiency. We go with a single fixed column at 3072 and document the tradeoff.

Decision: 3072 is generous but future-proof. If storage becomes an issue, we add a dimensions column to knowledge_bases and migrate to per-dimension indexes later. For v0.14.0, KISS.

Tables

-- ── Knowledge Bases ──────────────────────────

CREATE TABLE knowledge_bases (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    name            TEXT NOT NULL,
    description     TEXT NOT NULL DEFAULT '',
    scope           TEXT NOT NULL DEFAULT 'global',  -- global, team, personal
    owner_id        UUID,                            -- user_id (personal) or NULL
    team_id         UUID REFERENCES teams(id) ON DELETE CASCADE,
    embedding_config JSONB NOT NULL DEFAULT '{}',    -- snapshot: { provider_config_id, model_id, dimensions }

    -- Stats (denormalized, updated on ingest)
    document_count  INT NOT NULL DEFAULT 0,
    chunk_count     INT NOT NULL DEFAULT 0,
    total_bytes     BIGINT NOT NULL DEFAULT 0,

    status          TEXT NOT NULL DEFAULT 'active',  -- active, processing, error
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
    updated_at      TIMESTAMPTZ NOT NULL DEFAULT now(),

    CONSTRAINT kb_scope_check CHECK (
        (scope = 'global' AND owner_id IS NULL) OR
        (scope = 'team' AND team_id IS NOT NULL) OR
        (scope = 'personal' AND owner_id IS NOT NULL)
    )
);

CREATE INDEX idx_kb_scope ON knowledge_bases(scope);
CREATE INDEX idx_kb_owner ON knowledge_bases(owner_id) WHERE owner_id IS NOT NULL;
CREATE INDEX idx_kb_team ON knowledge_bases(team_id) WHERE team_id IS NOT NULL;

-- ── KB Documents ─────────────────────────────
-- Metadata row per uploaded document. Blobs live in ObjectStore
-- under key: kb/{kb_id}/{doc_id}_{filename}

CREATE TABLE kb_documents (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    kb_id           UUID NOT NULL REFERENCES knowledge_bases(id) ON DELETE CASCADE,
    filename        TEXT NOT NULL,
    content_type    TEXT NOT NULL,
    size_bytes      BIGINT NOT NULL,
    storage_key     TEXT NOT NULL,            -- ObjectStore key
    extracted_text  TEXT,                     -- full extracted text (for re-chunking)
    chunk_count     INT NOT NULL DEFAULT 0,
    status          TEXT NOT NULL DEFAULT 'pending', -- pending, chunking, embedding, ready, error
    error           TEXT,
    uploaded_by     UUID NOT NULL,
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
    updated_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE INDEX idx_kbdoc_kb ON kb_documents(kb_id);

-- ── KB Chunks ────────────────────────────────
-- Chunked text + embedding vector for similarity search.

CREATE TABLE kb_chunks (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    kb_id           UUID NOT NULL REFERENCES knowledge_bases(id) ON DELETE CASCADE,
    document_id     UUID NOT NULL REFERENCES kb_documents(id) ON DELETE CASCADE,
    chunk_index     INT NOT NULL,              -- ordinal within document
    content         TEXT NOT NULL,              -- chunk text
    token_count     INT NOT NULL DEFAULT 0,
    embedding       vector(3072),              -- pgvector column
    metadata        JSONB NOT NULL DEFAULT '{}', -- { page, section, heading, ... }
    created_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE INDEX idx_kbchunk_kb ON kb_chunks(kb_id);
CREATE INDEX idx_kbchunk_doc ON kb_chunks(document_id);

-- IVFFlat index for similarity search
-- (created after initial data load; needs rows to train)
-- Deferred: created by ingest handler after first batch
-- CREATE INDEX idx_kbchunk_embedding ON kb_chunks
--     USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);

-- ── Channel-KB Links ─────────────────────────
-- Which KBs are active for a given channel.

CREATE TABLE channel_knowledge_bases (
    channel_id      UUID NOT NULL REFERENCES channels(id) ON DELETE CASCADE,
    kb_id           UUID NOT NULL REFERENCES knowledge_bases(id) ON DELETE CASCADE,
    enabled         BOOLEAN NOT NULL DEFAULT true,
    added_at        TIMESTAMPTZ NOT NULL DEFAULT now(),
    PRIMARY KEY (channel_id, kb_id)
);

Scope Rules

Scope Visible to Created by Embedding cost
global All users Admin Org provider
team Team members Team admin Team provider
personal Owner only BYOK user User's own key

Personal KBs follow the same BYOK pattern as personal providers — the user's UEK decrypts their embedding provider key. Embedding costs are logged to usage_log with role = 'embedding'.


2. Chunking Pipeline

Strategy: Recursive Character Splitter

Start simple, iterate. Recursive splitting with configurable chunk size and overlap. No semantic chunking in v0.14.0 (requires embedding every boundary candidate — expensive, diminishing returns for a first pass).

// server/knowledge/chunker.go

type ChunkConfig struct {
    ChunkSize    int      // target chars per chunk (default 1000)
    ChunkOverlap int      // overlap chars between chunks (default 200)
    Separators   []string // split hierarchy: ["\n\n", "\n", ". ", " "]
}

type Chunk struct {
    Content    string
    Index      int               // ordinal in document
    TokenCount int               // estimated token count
    Metadata   map[string]any    // page, heading context
}

func SplitText(text string, cfg ChunkConfig) []Chunk

Default config: 1000 chars, 200 overlap (~250 tokens per chunk). These defaults are stored in knowledge_bases.embedding_config and overridable per-KB via the admin UI. The chunker tries separators in order: paragraph breaks first, then line breaks, then sentences, then words.

Text Extraction Reuse

Documents uploaded to KBs use the same extraction pipeline as attachments (v0.12.0). For inline-extractable types (TXT, MD, CSV), text is read directly. For sidecar types (PDF, DOCX), the extraction queue processes them. The extracted text is stored in kb_documents.extracted_text for re-chunking if config changes.


3. Embedding Pipeline

Flow

Upload → Extract Text → Chunk → Embed (batch) → Store vectors

Batch Embedding

The embedding role's Embed() method already accepts []string input. Chunks are batched in groups of 100 (or fewer, respecting provider token limits). Each batch is a single API call.

// server/knowledge/embedder.go

type Embedder struct {
    roleResolver *roles.Resolver
}

// EmbedChunks generates embeddings for a slice of chunks.
// Uses the embedding role with the resolution chain:
// personal → team → global (same as other role consumers).
func (e *Embedder) EmbedChunks(ctx context.Context, userID string,
    teamID *string, chunks []Chunk) ([][]float64, error)

Dimension Normalization

The embedding response includes raw vectors. If the vector dimension is less than 3072 (the column width), zero-pad. If greater, truncate (unlikely but defensive). Store the actual dimension in knowledge_bases.embedding_config.dimensions so search queries can account for it.

Note: Mixing embedding models within a KB is invalid — vectors from different models aren't comparable. The embedding_config is snapshotted at KB creation. Changing the embedding model requires re-embedding all documents (explicit admin action).


4. Ingestion Flow

No background job system exists yet (that's v0.15.0 compaction). Ingestion is synchronous with progress polling — same pattern as the extraction queue (v0.12.0).

API Endpoints

POST   /api/v1/knowledge-bases                   — Create KB
GET    /api/v1/knowledge-bases                    — List (scoped)
GET    /api/v1/knowledge-bases/:id                — Get KB details
PUT    /api/v1/knowledge-bases/:id                — Update name/description
DELETE /api/v1/knowledge-bases/:id                — Delete KB + all docs/chunks

POST   /api/v1/knowledge-bases/:id/documents      — Upload document(s)
GET    /api/v1/knowledge-bases/:id/documents       — List documents
DELETE /api/v1/knowledge-bases/:id/documents/:docId — Delete document + chunks

GET    /api/v1/knowledge-bases/:id/documents/:docId/status — Poll ingestion status

POST   /api/v1/knowledge-bases/:id/rebuild         — Re-chunk + re-embed all docs

POST   /api/v1/knowledge-bases/:id/search          — Direct search (debug/admin)

Admin endpoints mirror the pattern under /api/v1/admin/knowledge-bases/ for global KBs. Team KBs under /api/v1/teams/:teamId/knowledge-bases/. Personal KBs under the base path (scoped by auth).

Upload + Ingest Sequence

Client                          Server
  │                               │
  ├── POST /kb/:id/documents ───► │ Save blob to ObjectStore
  │   (multipart/form-data)       │ Create kb_documents row (status=pending)
  │                               │ Return { document_id, status: "pending" }
  │                               │
  │                               │── goroutine: ingestDocument()
  │                               │   ├─ Extract text (inline or queue)
  │                               │   ├─ status → "chunking"
  │                               │   ├─ Chunk text
  │                               │   ├─ status → "embedding"
  │                               │   ├─ Batch embed chunks
  │                               │   ├─ INSERT kb_chunks with vectors
  │                               │   ├─ Update KB stats
  │                               │   └─ status → "ready"
  │                               │
  ├── GET /kb/:id/docs/:d/status► │ Return current status + progress
  │   (poll every 2s)             │   { status, chunk_count, error }
  ◄─────────────────────────────  │

The goroutine is fire-and-forget per document. The status column on kb_documents is the progress indicator. Frontend polls until status = 'ready' or status = 'error'.

Concurrency: A per-process semaphore (channel of size 3) limits concurrent ingestion goroutines. Additional uploads queue behind the semaphore. This prevents overwhelming the embedding provider with parallel batch requests.


5. Search — kb_search Tool

Tool Definition

// server/tools/kbsearch.go

var kbSearchDef = ToolDef{
    Name:        "kb_search",
    DisplayName: "Knowledge Base",
    Description: "Search knowledge bases for relevant information. " +
        "Returns text passages from uploaded documents that match the query.",
    Category: "knowledge",
    Parameters: JSONSchema(map[string]interface{}{
        "query": Prop("string", "Search query — use natural language"),
        "max_results": map[string]interface{}{
            "type":        "integer",
            "description": "Maximum results to return (1-20, default 5)",
        },
    }, []string{"query"}),
}

Execution

func (t *KBSearchTool) Execute(ctx context.Context,
    execCtx ExecutionContext, argsJSON string) (string, error) {

    // 1. Parse args
    // 2. Get active KBs for this channel
    //    (from channel_knowledge_bases where enabled = true)
    //    PLUS personal KBs owned by user (always available)
    // 3. Embed the query using the embedding role
    // 4. Similarity search across all active KB chunks
    // 5. Format results with source attribution
}

Similarity Query

SELECT c.content, c.metadata, d.filename, kb.name as kb_name,
       1 - (c.embedding <=> $1::vector) AS similarity
FROM kb_chunks c
JOIN kb_documents d ON c.document_id = d.id
JOIN knowledge_bases kb ON c.kb_id = kb.id
WHERE c.kb_id = ANY($2)                    -- active KB IDs
  AND 1 - (c.embedding <=> $1) > $3        -- similarity threshold (0.3 default)
ORDER BY c.embedding <=> $1
LIMIT $4;                                  -- max_results

Result Format

{
  "results": [
    {
      "content": "Chunk text here...",
      "source": "quarterly-report.pdf",
      "kb": "Q4 Reports",
      "similarity": 0.87,
      "metadata": { "page": 12 }
    }
  ],
  "query": "revenue growth Q4",
  "searched_kbs": ["Q4 Reports", "Financial Data"]
}

ExecutionContext Extension

The kb_search tool needs access to the store layer and role resolver, which the current ExecutionContext doesn't provide. Two options:

Option A: Expand ExecutionContext — add optional fields:

type ExecutionContext struct {
    UserID    string
    ChannelID string
    // v0.14.0 additions
    Stores       *store.Stores       // nil for tools that don't need it
    RoleResolver *roles.Resolver     // nil for tools that don't need it
    TeamID       *string             // user's team (for role resolution)
}

Option B: Closure injection — the tool struct captures dependencies at registration time:

type kbSearchTool struct {
    stores       store.Stores
    roleResolver *roles.Resolver
}

func init() {
    // Deferred registration — called from main.go after stores init
}

Decision: Option B (closure injection). It's the same pattern the notes tool already uses (it captures the NoteStore). The kb_search tool is registered after stores and role resolver init in main.go, not via init(). Add a RegisterLate() function or explicit call.


6. Context Injection

When kb_search returns results, they flow through the existing tool call loop in stream_loop.go — no special injection needed. The tool result becomes a role: "tool" message in the conversation, and the LLM sees the retrieved context naturally.

Automatic vs. Tool-Based

Two approaches to getting KB context into the conversation:

Approach Pros Cons
Tool-based (v0.14.0) LLM decides when to search; transparent; works with existing tool loop Extra round-trip; LLM might not search when it should
Auto-inject (future) Always available; no tool call overhead Wastes tokens if not needed; opaque to user

v0.14.0: Tool-based only. The LLM calls kb_search when it needs information. Auto-injection (pre-pend top-K results to every prompt) is a v0.15.x consideration when we have compaction to manage context budget.

System Prompt Hint

When KBs are active on a channel, append a hint to the system prompt:

You have access to the following knowledge bases:
- "Q4 Reports" (3 documents)
- "Engineering Wiki" (12 documents)
Use the kb_search tool to find relevant information from these sources.

This nudges the model to use the tool without auto-injecting content.


7. Notes Embedding

Notes already have full-text search via PostgreSQL tsvector. Adding vector search means notes can be found semantically, not just by keyword.

Approach: Treat notes as single-chunk documents. On note create/update, embed the full note text and store in a new column:

ALTER TABLE notes ADD COLUMN embedding vector(3072);

The note_search tool (already exists) gains an optional semantic: true parameter that switches from tsvector to vector similarity search. Both results can be merged and deduplicated.

Deferred to Phase 3 — notes embedding is additive and independent of the core KB pipeline. Ship KB search first, add note embeddings after the pipeline is proven.


8. Per-Channel KB Toggle

UI: Chat Sidebar

A "Knowledge" button in the chat header or input bar (next to the tools toggle). Click opens a popup listing available KBs with toggle switches.

┌─────────────────────────────────┐
│ 📚 Knowledge Bases              │
│                                 │
│ ☑ Q4 Reports        (3 docs)   │
│ ☑ Engineering Wiki   (12 docs)  │
│ ☐ HR Policies        (5 docs)   │
│                                 │
│ Manage KBs...                   │
└─────────────────────────────────┘

State

Active KBs per channel stored in channel_knowledge_bases. When a user toggles a KB on/off for a channel, the frontend calls:

PUT /api/v1/channels/:id/knowledge-bases
Body: { kb_ids: ["uuid1", "uuid2"] }

The completion handler reads active KBs to populate the system prompt hint and to scope kb_search results.

Visibility Rules

A user can toggle KBs they have access to:

  • Global KBs — all users
  • Team KBs — team members
  • Personal KBs — owner only

The list endpoint returns all KBs the user can see, with enabled status per channel.


9. Store Interface

// server/store/interfaces.go — additions

// =========================================
// KNOWLEDGE BASE STORE
// =========================================

type KnowledgeBaseStore interface {
    // KB CRUD
    Create(ctx context.Context, kb *models.KnowledgeBase) error
    GetByID(ctx context.Context, id string) (*models.KnowledgeBase, error)
    Update(ctx context.Context, id string, fields map[string]interface{}) error
    Delete(ctx context.Context, id string) error

    // Scoped listing
    ListForUser(ctx context.Context, userID string, teamIDs []string) ([]models.KnowledgeBase, error)
    ListGlobal(ctx context.Context) ([]models.KnowledgeBase, error)
    ListForTeam(ctx context.Context, teamID string) ([]models.KnowledgeBase, error)
    ListPersonal(ctx context.Context, userID string) ([]models.KnowledgeBase, error)

    // Documents
    CreateDocument(ctx context.Context, doc *models.KBDocument) error
    GetDocument(ctx context.Context, id string) (*models.KBDocument, error)
    ListDocuments(ctx context.Context, kbID string) ([]models.KBDocument, error)
    UpdateDocumentStatus(ctx context.Context, id string, status string, err *string) error
    DeleteDocument(ctx context.Context, id string) (*models.KBDocument, error) // returns for storage cleanup

    // Chunks
    InsertChunks(ctx context.Context, chunks []models.KBChunk) error
    DeleteChunksForDocument(ctx context.Context, documentID string) error
    SimilaritySearch(ctx context.Context, kbIDs []string, queryVec []float64,
        threshold float64, limit int) ([]models.KBSearchResult, error)

    // Channel links
    SetChannelKBs(ctx context.Context, channelID string, kbIDs []string) error
    GetChannelKBs(ctx context.Context, channelID string) ([]models.ChannelKB, error)
    GetActiveKBIDs(ctx context.Context, channelID string, userID string,
        teamIDs []string) ([]string, error) // enabled + user has access

    // Stats
    UpdateStats(ctx context.Context, kbID string) error // recount from chunks
}

Add KnowledgeBases KnowledgeBaseStore to the Stores struct.


10. Models

// server/models/models.go — additions

type KnowledgeBase struct {
    ID              string                 `json:"id" db:"id"`
    Name            string                 `json:"name" db:"name"`
    Description     string                 `json:"description" db:"description"`
    Scope           string                 `json:"scope" db:"scope"`
    OwnerID         *string                `json:"owner_id,omitempty" db:"owner_id"`
    TeamID          *string                `json:"team_id,omitempty" db:"team_id"`
    EmbeddingConfig map[string]interface{} `json:"embedding_config" db:"embedding_config"`
    DocumentCount   int                    `json:"document_count" db:"document_count"`
    ChunkCount      int                    `json:"chunk_count" db:"chunk_count"`
    TotalBytes      int64                  `json:"total_bytes" db:"total_bytes"`
    Status          string                 `json:"status" db:"status"`
    CreatedAt       time.Time              `json:"created_at" db:"created_at"`
    UpdatedAt       time.Time              `json:"updated_at" db:"updated_at"`
}

type KBDocument struct {
    ID            string     `json:"id" db:"id"`
    KBID          string     `json:"kb_id" db:"kb_id"`
    Filename      string     `json:"filename" db:"filename"`
    ContentType   string     `json:"content_type" db:"content_type"`
    SizeBytes     int64      `json:"size_bytes" db:"size_bytes"`
    StorageKey    string     `json:"storage_key" db:"storage_key"`
    ExtractedText *string    `json:"extracted_text,omitempty" db:"extracted_text"`
    ChunkCount    int        `json:"chunk_count" db:"chunk_count"`
    Status        string     `json:"status" db:"status"`
    Error         *string    `json:"error,omitempty" db:"error"`
    UploadedBy    string     `json:"uploaded_by" db:"uploaded_by"`
    CreatedAt     time.Time  `json:"created_at" db:"created_at"`
    UpdatedAt     time.Time  `json:"updated_at" db:"updated_at"`
}

type KBChunk struct {
    ID          string                 `json:"id" db:"id"`
    KBID        string                 `json:"kb_id" db:"kb_id"`
    DocumentID  string                 `json:"document_id" db:"document_id"`
    ChunkIndex  int                    `json:"chunk_index" db:"chunk_index"`
    Content     string                 `json:"content" db:"content"`
    TokenCount  int                    `json:"token_count" db:"token_count"`
    Embedding   []float64              `json:"-" db:"embedding"` // not serialized to API
    Metadata    map[string]interface{} `json:"metadata" db:"metadata"`
    CreatedAt   time.Time              `json:"created_at" db:"created_at"`
}

type KBSearchResult struct {
    Content    string                 `json:"content"`
    Filename   string                 `json:"source"`
    KBName     string                 `json:"kb"`
    Similarity float64                `json:"similarity"`
    Metadata   map[string]interface{} `json:"metadata,omitempty"`
}

type ChannelKB struct {
    KBID    string `json:"kb_id" db:"kb_id"`
    KBName  string `json:"kb_name" db:"name"`
    Enabled bool   `json:"enabled" db:"enabled"`
    DocCount int   `json:"document_count" db:"document_count"`
}

11. File Structure

New Go Packages

server/
├── knowledge/                    # NEW — chunking + embedding logic
│   ├── chunker.go                # recursive text splitter
│   ├── chunker_test.go
│   ├── embedder.go               # batch embed via role resolver
│   └── ingest.go                 # orchestrator: extract → chunk → embed → store
├── handlers/
│   ├── knowledge_bases.go        # NEW — KB CRUD + document upload + search
│   └── ... (existing)
├── store/
│   └── postgres/
│       ├── knowledge_bases.go    # NEW — KnowledgeBaseStore impl
│       └── ... (existing)
├── tools/
│   ├── kbsearch.go               # NEW — kb_search tool
│   └── ... (existing)
├── database/
│   └── migrations/
│       └── 008_knowledge_bases.sql  # NEW
└── models/
    └── models.go                 # additions (KB types)

New/Modified Frontend Files

src/js/
├── ui-admin.js          # MODIFIED — KB management section in admin panel
├── ui-primitives.js     # MODIFIED — embedding dropdown fix, manual entry
├── chat.js              # MODIFIED — KB toggle popup, system prompt hint
├── api.js               # MODIFIED — KB API methods
└── settings-handlers.js # MODIFIED — personal KB management

Go Dependencies

go get github.com/pgvector/pgvector-go  # pgvector type support for lib/pq

12. Admin Panel Integration

KBs live under AI → Knowledge Bases in the admin panel (v0.13.0).

Admin View

┌──────────────────────────────────────────────────────┐
│ AI → Knowledge Bases                                 │
│                                                      │
│ [+ New Knowledge Base]                               │
│                                                      │
│ ┌────────────────────────────────────────────────┐   │
│ │ Q4 Reports              global   3 docs  ready │   │
│ │ Engineering Wiki         global  12 docs  ready │   │
│ │ Onboarding Materials     team:Eng 5 docs  ready │   │
│ └────────────────────────────────────────────────┘   │
│                                                      │
│ ── Selected: Q4 Reports ──────────────────────────   │
│                                                      │
│ Name: [Q4 Reports                ]                   │
│ Description: [Quarterly financial reports     ]      │
│ Scope: [Global ▾]   Team: [— ▾]                     │
│                                                      │
│ Embedding: OpenAI / text-embedding-3-small           │
│ Chunk size: [1000]  Overlap: [200]                   │
│                                                      │
│ Documents:                                           │
│ ┌────────────────────────────────────────────────┐   │
│ │ quarterly-report-q4.pdf  2.1 MB  42 chunks ✓  │   │
│ │ revenue-breakdown.xlsx   340 KB  12 chunks ✓   │   │
│ │ market-analysis.docx     1.8 MB  processing... │   │
│ └────────────────────────────────────────────────┘   │
│ [📎 Upload Documents]  [🔄 Rebuild All]  [🗑 Delete] │
└──────────────────────────────────────────────────────┘

Personal KB (Settings)

BYOK users see a "Knowledge Bases" section in Settings → Model Roles (or a new Settings tab). Same UI, scoped to personal KBs using the user's embedding role override.

Team KB (Team Management)

Team admins see a "Knowledge Bases" tab in the team management modal. Same UI pattern, scoped to the team.


13. Phased Delivery

Phase 1: Schema + Chunker + Embed Fix

Goal: Database ready, chunking works, embedding dropdown fixed.

  • 008_knowledge_bases.sql migration (pgvector + tables)
  • go get pgvector-go, update go.mod
  • knowledge/chunker.go — recursive text splitter + tests
  • Embedding dropdown fix in ui-primitives.js (tolerant filter + manual entry)
  • KnowledgeBaseStore interface + postgres implementation (CRUD only, no vector ops)
  • KnowledgeBase, KBDocument, KBChunk model types
  • Wire Stores.KnowledgeBases in main.go

Phase 2: Ingestion Pipeline

Goal: Documents can be uploaded, chunked, and embedded end-to-end.

  • knowledge/embedder.go — batch embed via role resolver
  • knowledge/ingest.go — orchestrator (extract → chunk → embed → store)
  • handlers/knowledge_bases.go — KB CRUD endpoints
  • Document upload with async ingestion goroutine
  • Status polling endpoint
  • SimilaritySearch() in store (pgvector query)
  • Admin panel: KB management section (AI → Knowledge Bases)
  • IVFFlat index creation (manual or after first batch)

Phase 3: kb_search Tool + Channel Integration

Goal: LLMs can search KBs during conversations.

  • tools/kbsearch.go — kb_search tool (late registration pattern)
  • Per-channel KB toggle (channel_knowledge_bases)
  • System prompt KB hint injection in completion handler
  • KB toggle popup in chat UI
  • Frontend: channel KB management
  • Category "knowledge" in tools toggle menu

Phase 4: Polish + Notes

Goal: Team/personal scoping works, notes get vectors.

  • Team KB endpoints + team management UI tab
  • Personal KB endpoints + settings UI
  • KB rebuild endpoint (re-chunk + re-embed)
  • Notes embedding column + semantic note search
  • Audit logging for KB operations
  • Usage tracking for embedding calls (role = 'embedding')
  • Integration tests

14. Risks + Mitigations

Risk Impact Mitigation
pgvector not installed on PG host Migration fails, blocks deploy Clear error message in migration; add to deployment docs/Dockerfile
Embedding provider rate limits Ingestion stalls on large uploads Semaphore + exponential backoff; batch size tuning
Dimension mismatch (model change) Search returns garbage Snapshot model in embedding_config; warn if role model ≠ KB model
Large documents (100+ pages) Memory pressure during chunking Stream text extraction; chunk incrementally
Vector storage size DB bloat with many KBs Monitor with stats endpoint; future: HNSW index for better perf
Embedding dropdown broken Can't configure embedding role at all Phase 1 priority — ship fix before KB features

15. Future (Not v0.14.0)

  • Auto-injection: Pre-pend top-K results to system prompt (v0.15.x, needs compaction for context budget)
  • Hybrid search: Combine vector similarity with full-text tsvector search, re-rank
  • Semantic chunking: Use embedding distance to detect topic boundaries
  • HNSW index: Better query performance than IVFFlat for large datasets
  • Web scraping source: Ingest URLs as KB documents (natural extension of url_fetch)
  • Scheduled re-indexing: Periodic rebuild when source documents are updated
  • KB sharing: Cross-team KB access grants