846 lines
33 KiB
Markdown
846 lines
33 KiB
Markdown
# DESIGN-0.14.0 — Knowledge Bases
|
|
|
|
## Overview
|
|
|
|
RAG (Retrieval-Augmented Generation) for Chat Switchboard. Users upload
|
|
documents into named knowledge bases, the backend chunks and embeds them
|
|
via the embedding model role (v0.10.0), stores vectors in pgvector, and a
|
|
`kb_search` tool lets the LLM pull relevant context at completion time.
|
|
|
|
Depends on: embedding model role (v0.10.0), file storage (v0.12.0),
|
|
tool framework (v0.11.0), admin panel (v0.13.0).
|
|
|
|
---
|
|
|
|
## 0. Pre-Req Fix: Embedding Model Selection
|
|
|
|
The embedding role dropdown is broken — `filterModels()` in
|
|
`ui-primitives.js` filters by `model_type === 'embedding'`, but many
|
|
providers don't report a type for embedding models (defaults to `'chat'`
|
|
and gets filtered out). This blocks KB setup entirely.
|
|
|
|
**Fix (ship before or with Phase 1):**
|
|
|
|
1. **Manual model ID entry** — add a text input fallback alongside the
|
|
dropdown. If the dropdown shows no results for the selected provider,
|
|
display an editable text field pre-populated with the `model_id` from
|
|
the saved binding (if any). Users type the model ID directly
|
|
(e.g. `text-embedding-3-small`). The save handler accepts either
|
|
dropdown selection or manual text entry.
|
|
|
|
2. **Tolerant type filter** — change `filterModels()` to include models
|
|
where `model_type` is empty/null/undefined when the role's type filter
|
|
is `'embedding'`. Embedding models are rare enough that showing
|
|
untyped models alongside typed ones is better than showing nothing.
|
|
|
|
```javascript
|
|
// Before (strict — breaks when model_type missing)
|
|
(m.model_type || 'chat') === typeFilter
|
|
|
|
// After (tolerant for embedding role)
|
|
typeFilter === 'embedding'
|
|
? (!m.model_type || m.model_type === 'embedding')
|
|
: (m.model_type || 'chat') === typeFilter
|
|
```
|
|
|
|
3. **Combo UI** — the slot renderer gets a hybrid dropdown+input:
|
|
|
|
```
|
|
┌──────────────────────────────────────────────┐
|
|
│ Provider: [OpenAI ▾] │
|
|
│ Model: [text-embedding-3-small ▾] [✎] │
|
|
│ ↑ toggle │
|
|
│ [text-embedding-3-small ] │ ← manual entry (shown on ✎ click)
|
|
└──────────────────────────────────────────────┘
|
|
```
|
|
|
|
The pencil icon toggles between dropdown and text input. If the dropdown
|
|
has zero options after provider change, auto-switch to manual entry.
|
|
Saved bindings with model IDs not in the dropdown auto-show in manual
|
|
mode on load.
|
|
|
|
---
|
|
|
|
## 1. Schema — Migration `008_knowledge_bases.sql`
|
|
|
|
### pgvector Extension
|
|
|
|
```sql
|
|
CREATE EXTENSION IF NOT EXISTS vector;
|
|
```
|
|
|
|
Requires `pgvector` installed on the PostgreSQL server. The migration
|
|
checks for the extension and fails with a clear error if unavailable,
|
|
rather than silently degrading. The README/deployment docs will note the
|
|
dependency.
|
|
|
|
**Dimension handling:** Embedding dimensions vary by model
|
|
(OpenAI ada-002 = 1536, text-embedding-3-small = 1536,
|
|
text-embedding-3-large = 3072, many open-source = 768 or 1024).
|
|
The dimension is stored per KB and the vector column uses the max
|
|
supported (3072), with vectors zero-padded or truncated at insert time.
|
|
Alternative: use `halfvec` for storage efficiency. We go with a single
|
|
fixed column at 3072 and document the tradeoff.
|
|
|
|
> **Decision:** 3072 is generous but future-proof. If storage becomes
|
|
> an issue, we add a `dimensions` column to `knowledge_bases` and
|
|
> migrate to per-dimension indexes later. For v0.14.0, KISS.
|
|
|
|
### Tables
|
|
|
|
```sql
|
|
-- ── Knowledge Bases ──────────────────────────
|
|
|
|
CREATE TABLE knowledge_bases (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
name TEXT NOT NULL,
|
|
description TEXT NOT NULL DEFAULT '',
|
|
scope TEXT NOT NULL DEFAULT 'global', -- global, team, personal
|
|
owner_id UUID, -- user_id (personal) or NULL
|
|
team_id UUID REFERENCES teams(id) ON DELETE CASCADE,
|
|
embedding_config JSONB NOT NULL DEFAULT '{}', -- snapshot: { provider_config_id, model_id, dimensions }
|
|
|
|
-- Stats (denormalized, updated on ingest)
|
|
document_count INT NOT NULL DEFAULT 0,
|
|
chunk_count INT NOT NULL DEFAULT 0,
|
|
total_bytes BIGINT NOT NULL DEFAULT 0,
|
|
|
|
status TEXT NOT NULL DEFAULT 'active', -- active, processing, error
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
|
|
updated_at TIMESTAMPTZ NOT NULL DEFAULT now(),
|
|
|
|
CONSTRAINT kb_scope_check CHECK (
|
|
(scope = 'global' AND owner_id IS NULL) OR
|
|
(scope = 'team' AND team_id IS NOT NULL) OR
|
|
(scope = 'personal' AND owner_id IS NOT NULL)
|
|
)
|
|
);
|
|
|
|
CREATE INDEX idx_kb_scope ON knowledge_bases(scope);
|
|
CREATE INDEX idx_kb_owner ON knowledge_bases(owner_id) WHERE owner_id IS NOT NULL;
|
|
CREATE INDEX idx_kb_team ON knowledge_bases(team_id) WHERE team_id IS NOT NULL;
|
|
|
|
-- ── KB Documents ─────────────────────────────
|
|
-- Metadata row per uploaded document. Blobs live in ObjectStore
|
|
-- under key: kb/{kb_id}/{doc_id}_{filename}
|
|
|
|
CREATE TABLE kb_documents (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
kb_id UUID NOT NULL REFERENCES knowledge_bases(id) ON DELETE CASCADE,
|
|
filename TEXT NOT NULL,
|
|
content_type TEXT NOT NULL,
|
|
size_bytes BIGINT NOT NULL,
|
|
storage_key TEXT NOT NULL, -- ObjectStore key
|
|
extracted_text TEXT, -- full extracted text (for re-chunking)
|
|
chunk_count INT NOT NULL DEFAULT 0,
|
|
status TEXT NOT NULL DEFAULT 'pending', -- pending, chunking, embedding, ready, error
|
|
error TEXT,
|
|
uploaded_by UUID NOT NULL,
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
|
|
updated_at TIMESTAMPTZ NOT NULL DEFAULT now()
|
|
);
|
|
|
|
CREATE INDEX idx_kbdoc_kb ON kb_documents(kb_id);
|
|
|
|
-- ── KB Chunks ────────────────────────────────
|
|
-- Chunked text + embedding vector for similarity search.
|
|
|
|
CREATE TABLE kb_chunks (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
kb_id UUID NOT NULL REFERENCES knowledge_bases(id) ON DELETE CASCADE,
|
|
document_id UUID NOT NULL REFERENCES kb_documents(id) ON DELETE CASCADE,
|
|
chunk_index INT NOT NULL, -- ordinal within document
|
|
content TEXT NOT NULL, -- chunk text
|
|
token_count INT NOT NULL DEFAULT 0,
|
|
embedding vector(3072), -- pgvector column
|
|
metadata JSONB NOT NULL DEFAULT '{}', -- { page, section, heading, ... }
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
|
|
);
|
|
|
|
CREATE INDEX idx_kbchunk_kb ON kb_chunks(kb_id);
|
|
CREATE INDEX idx_kbchunk_doc ON kb_chunks(document_id);
|
|
|
|
-- IVFFlat index for similarity search
|
|
-- (created after initial data load; needs rows to train)
|
|
-- Deferred: created by ingest handler after first batch
|
|
-- CREATE INDEX idx_kbchunk_embedding ON kb_chunks
|
|
-- USING ivfflat (embedding vector_cosine_ops) WITH (lists = 100);
|
|
|
|
-- ── Channel-KB Links ─────────────────────────
|
|
-- Which KBs are active for a given channel.
|
|
|
|
CREATE TABLE channel_knowledge_bases (
|
|
channel_id UUID NOT NULL REFERENCES channels(id) ON DELETE CASCADE,
|
|
kb_id UUID NOT NULL REFERENCES knowledge_bases(id) ON DELETE CASCADE,
|
|
enabled BOOLEAN NOT NULL DEFAULT true,
|
|
added_at TIMESTAMPTZ NOT NULL DEFAULT now(),
|
|
PRIMARY KEY (channel_id, kb_id)
|
|
);
|
|
```
|
|
|
|
### Scope Rules
|
|
|
|
| Scope | Visible to | Created by | Embedding cost |
|
|
|----------|------------------------|-------------------|-----------------|
|
|
| global | All users | Admin | Org provider |
|
|
| team | Team members | Team admin | Team provider |
|
|
| personal | Owner only | BYOK user | User's own key |
|
|
|
|
Personal KBs follow the same BYOK pattern as personal providers — the
|
|
user's UEK decrypts their embedding provider key. Embedding costs are
|
|
logged to usage_log with `role = 'embedding'`.
|
|
|
|
---
|
|
|
|
## 2. Chunking Pipeline
|
|
|
|
### Strategy: Recursive Character Splitter
|
|
|
|
Start simple, iterate. Recursive splitting with configurable chunk size
|
|
and overlap. No semantic chunking in v0.14.0 (requires embedding every
|
|
boundary candidate — expensive, diminishing returns for a first pass).
|
|
|
|
```go
|
|
// server/knowledge/chunker.go
|
|
|
|
type ChunkConfig struct {
|
|
ChunkSize int // target chars per chunk (default 1000)
|
|
ChunkOverlap int // overlap chars between chunks (default 200)
|
|
Separators []string // split hierarchy: ["\n\n", "\n", ". ", " "]
|
|
}
|
|
|
|
type Chunk struct {
|
|
Content string
|
|
Index int // ordinal in document
|
|
TokenCount int // estimated token count
|
|
Metadata map[string]any // page, heading context
|
|
}
|
|
|
|
func SplitText(text string, cfg ChunkConfig) []Chunk
|
|
```
|
|
|
|
**Default config:** 1000 chars, 200 overlap (~250 tokens per chunk).
|
|
These defaults are stored in `knowledge_bases.embedding_config` and
|
|
overridable per-KB via the admin UI. The chunker tries separators in
|
|
order: paragraph breaks first, then line breaks, then sentences, then
|
|
words.
|
|
|
|
### Text Extraction Reuse
|
|
|
|
Documents uploaded to KBs use the same extraction pipeline as
|
|
attachments (v0.12.0). For inline-extractable types (TXT, MD, CSV),
|
|
text is read directly. For sidecar types (PDF, DOCX), the extraction
|
|
queue processes them. The extracted text is stored in
|
|
`kb_documents.extracted_text` for re-chunking if config changes.
|
|
|
|
---
|
|
|
|
## 3. Embedding Pipeline
|
|
|
|
### Flow
|
|
|
|
```
|
|
Upload → Extract Text → Chunk → Embed (batch) → Store vectors
|
|
```
|
|
|
|
### Batch Embedding
|
|
|
|
The embedding role's `Embed()` method already accepts `[]string` input.
|
|
Chunks are batched in groups of 100 (or fewer, respecting provider token
|
|
limits). Each batch is a single API call.
|
|
|
|
```go
|
|
// server/knowledge/embedder.go
|
|
|
|
type Embedder struct {
|
|
roleResolver *roles.Resolver
|
|
}
|
|
|
|
// EmbedChunks generates embeddings for a slice of chunks.
|
|
// Uses the embedding role with the resolution chain:
|
|
// personal → team → global (same as other role consumers).
|
|
func (e *Embedder) EmbedChunks(ctx context.Context, userID string,
|
|
teamID *string, chunks []Chunk) ([][]float64, error)
|
|
```
|
|
|
|
### Dimension Normalization
|
|
|
|
The embedding response includes raw vectors. If the vector dimension
|
|
is less than 3072 (the column width), zero-pad. If greater, truncate
|
|
(unlikely but defensive). Store the actual dimension in
|
|
`knowledge_bases.embedding_config.dimensions` so search queries can
|
|
account for it.
|
|
|
|
> **Note:** Mixing embedding models within a KB is invalid — vectors
|
|
> from different models aren't comparable. The `embedding_config` is
|
|
> snapshotted at KB creation. Changing the embedding model requires
|
|
> re-embedding all documents (explicit admin action).
|
|
|
|
---
|
|
|
|
## 4. Ingestion Flow
|
|
|
|
No background job system exists yet (that's v0.15.0 compaction).
|
|
Ingestion is **synchronous with progress polling** — same pattern as
|
|
the extraction queue (v0.12.0).
|
|
|
|
### API Endpoints
|
|
|
|
```
|
|
POST /api/v1/knowledge-bases — Create KB
|
|
GET /api/v1/knowledge-bases — List (scoped)
|
|
GET /api/v1/knowledge-bases/:id — Get KB details
|
|
PUT /api/v1/knowledge-bases/:id — Update name/description
|
|
DELETE /api/v1/knowledge-bases/:id — Delete KB + all docs/chunks
|
|
|
|
POST /api/v1/knowledge-bases/:id/documents — Upload document(s)
|
|
GET /api/v1/knowledge-bases/:id/documents — List documents
|
|
DELETE /api/v1/knowledge-bases/:id/documents/:docId — Delete document + chunks
|
|
|
|
GET /api/v1/knowledge-bases/:id/documents/:docId/status — Poll ingestion status
|
|
|
|
POST /api/v1/knowledge-bases/:id/rebuild — Re-chunk + re-embed all docs
|
|
|
|
POST /api/v1/knowledge-bases/:id/search — Direct search (debug/admin)
|
|
```
|
|
|
|
Admin endpoints mirror the pattern under `/api/v1/admin/knowledge-bases/`
|
|
for global KBs. Team KBs under `/api/v1/teams/:teamId/knowledge-bases/`.
|
|
Personal KBs under the base path (scoped by auth).
|
|
|
|
### Upload + Ingest Sequence
|
|
|
|
```
|
|
Client Server
|
|
│ │
|
|
├── POST /kb/:id/documents ───► │ Save blob to ObjectStore
|
|
│ (multipart/form-data) │ Create kb_documents row (status=pending)
|
|
│ │ Return { document_id, status: "pending" }
|
|
│ │
|
|
│ │── goroutine: ingestDocument()
|
|
│ │ ├─ Extract text (inline or queue)
|
|
│ │ ├─ status → "chunking"
|
|
│ │ ├─ Chunk text
|
|
│ │ ├─ status → "embedding"
|
|
│ │ ├─ Batch embed chunks
|
|
│ │ ├─ INSERT kb_chunks with vectors
|
|
│ │ ├─ Update KB stats
|
|
│ │ └─ status → "ready"
|
|
│ │
|
|
├── GET /kb/:id/docs/:d/status► │ Return current status + progress
|
|
│ (poll every 2s) │ { status, chunk_count, error }
|
|
◄───────────────────────────── │
|
|
```
|
|
|
|
The goroutine is fire-and-forget per document. The status column on
|
|
`kb_documents` is the progress indicator. Frontend polls until
|
|
`status = 'ready'` or `status = 'error'`.
|
|
|
|
**Concurrency:** A per-process semaphore (channel of size 3) limits
|
|
concurrent ingestion goroutines. Additional uploads queue behind the
|
|
semaphore. This prevents overwhelming the embedding provider with
|
|
parallel batch requests.
|
|
|
|
---
|
|
|
|
## 5. Search — `kb_search` Tool
|
|
|
|
### Tool Definition
|
|
|
|
```go
|
|
// server/tools/kbsearch.go
|
|
|
|
var kbSearchDef = ToolDef{
|
|
Name: "kb_search",
|
|
DisplayName: "Knowledge Base",
|
|
Description: "Search knowledge bases for relevant information. " +
|
|
"Returns text passages from uploaded documents that match the query.",
|
|
Category: "knowledge",
|
|
Parameters: JSONSchema(map[string]interface{}{
|
|
"query": Prop("string", "Search query — use natural language"),
|
|
"max_results": map[string]interface{}{
|
|
"type": "integer",
|
|
"description": "Maximum results to return (1-20, default 5)",
|
|
},
|
|
}, []string{"query"}),
|
|
}
|
|
```
|
|
|
|
### Execution
|
|
|
|
```go
|
|
func (t *KBSearchTool) Execute(ctx context.Context,
|
|
execCtx ExecutionContext, argsJSON string) (string, error) {
|
|
|
|
// 1. Parse args
|
|
// 2. Get active KBs for this channel
|
|
// (from channel_knowledge_bases where enabled = true)
|
|
// PLUS personal KBs owned by user (always available)
|
|
// 3. Embed the query using the embedding role
|
|
// 4. Similarity search across all active KB chunks
|
|
// 5. Format results with source attribution
|
|
}
|
|
```
|
|
|
|
### Similarity Query
|
|
|
|
```sql
|
|
SELECT c.content, c.metadata, d.filename, kb.name as kb_name,
|
|
1 - (c.embedding <=> $1::vector) AS similarity
|
|
FROM kb_chunks c
|
|
JOIN kb_documents d ON c.document_id = d.id
|
|
JOIN knowledge_bases kb ON c.kb_id = kb.id
|
|
WHERE c.kb_id = ANY($2) -- active KB IDs
|
|
AND 1 - (c.embedding <=> $1) > $3 -- similarity threshold (0.3 default)
|
|
ORDER BY c.embedding <=> $1
|
|
LIMIT $4; -- max_results
|
|
```
|
|
|
|
### Result Format
|
|
|
|
```json
|
|
{
|
|
"results": [
|
|
{
|
|
"content": "Chunk text here...",
|
|
"source": "quarterly-report.pdf",
|
|
"kb": "Q4 Reports",
|
|
"similarity": 0.87,
|
|
"metadata": { "page": 12 }
|
|
}
|
|
],
|
|
"query": "revenue growth Q4",
|
|
"searched_kbs": ["Q4 Reports", "Financial Data"]
|
|
}
|
|
```
|
|
|
|
### ExecutionContext Extension
|
|
|
|
The `kb_search` tool needs access to the store layer and role resolver,
|
|
which the current `ExecutionContext` doesn't provide. Two options:
|
|
|
|
**Option A: Expand ExecutionContext** — add optional fields:
|
|
|
|
```go
|
|
type ExecutionContext struct {
|
|
UserID string
|
|
ChannelID string
|
|
// v0.14.0 additions
|
|
Stores *store.Stores // nil for tools that don't need it
|
|
RoleResolver *roles.Resolver // nil for tools that don't need it
|
|
TeamID *string // user's team (for role resolution)
|
|
}
|
|
```
|
|
|
|
**Option B: Closure injection** — the tool struct captures dependencies
|
|
at registration time:
|
|
|
|
```go
|
|
type kbSearchTool struct {
|
|
stores store.Stores
|
|
roleResolver *roles.Resolver
|
|
}
|
|
|
|
func init() {
|
|
// Deferred registration — called from main.go after stores init
|
|
}
|
|
```
|
|
|
|
**Decision: Option B (closure injection).** It's the same pattern the
|
|
notes tool already uses (it captures the NoteStore). The `kb_search`
|
|
tool is registered after stores and role resolver init in `main.go`,
|
|
not via `init()`. Add a `RegisterLate()` function or explicit call.
|
|
|
|
---
|
|
|
|
## 6. Context Injection
|
|
|
|
When `kb_search` returns results, they flow through the existing tool
|
|
call loop in `stream_loop.go` — no special injection needed. The tool
|
|
result becomes a `role: "tool"` message in the conversation, and the
|
|
LLM sees the retrieved context naturally.
|
|
|
|
### Automatic vs. Tool-Based
|
|
|
|
Two approaches to getting KB context into the conversation:
|
|
|
|
| Approach | Pros | Cons |
|
|
|----------|------|------|
|
|
| **Tool-based** (v0.14.0) | LLM decides when to search; transparent; works with existing tool loop | Extra round-trip; LLM might not search when it should |
|
|
| **Auto-inject** (future) | Always available; no tool call overhead | Wastes tokens if not needed; opaque to user |
|
|
|
|
**v0.14.0: Tool-based only.** The LLM calls `kb_search` when it needs
|
|
information. Auto-injection (pre-pend top-K results to every prompt) is
|
|
a v0.15.x consideration when we have compaction to manage context
|
|
budget.
|
|
|
|
### System Prompt Hint
|
|
|
|
When KBs are active on a channel, append a hint to the system prompt:
|
|
|
|
```
|
|
You have access to the following knowledge bases:
|
|
- "Q4 Reports" (3 documents)
|
|
- "Engineering Wiki" (12 documents)
|
|
Use the kb_search tool to find relevant information from these sources.
|
|
```
|
|
|
|
This nudges the model to use the tool without auto-injecting content.
|
|
|
|
---
|
|
|
|
## 7. Notes Embedding
|
|
|
|
Notes already have full-text search via PostgreSQL `tsvector`. Adding
|
|
vector search means notes can be found semantically, not just by keyword.
|
|
|
|
**Approach:** Treat notes as single-chunk documents. On note
|
|
create/update, embed the full note text and store in a new column:
|
|
|
|
```sql
|
|
ALTER TABLE notes ADD COLUMN embedding vector(3072);
|
|
```
|
|
|
|
The `note_search` tool (already exists) gains an optional
|
|
`semantic: true` parameter that switches from `tsvector` to vector
|
|
similarity search. Both results can be merged and deduplicated.
|
|
|
|
**Deferred to Phase 3** — notes embedding is additive and independent
|
|
of the core KB pipeline. Ship KB search first, add note embeddings
|
|
after the pipeline is proven.
|
|
|
|
---
|
|
|
|
## 8. Per-Channel KB Toggle
|
|
|
|
### UI: Chat Sidebar
|
|
|
|
A "Knowledge" button in the chat header or input bar (next to the tools
|
|
toggle). Click opens a popup listing available KBs with toggle switches.
|
|
|
|
```
|
|
┌─────────────────────────────────┐
|
|
│ 📚 Knowledge Bases │
|
|
│ │
|
|
│ ☑ Q4 Reports (3 docs) │
|
|
│ ☑ Engineering Wiki (12 docs) │
|
|
│ ☐ HR Policies (5 docs) │
|
|
│ │
|
|
│ Manage KBs... │
|
|
└─────────────────────────────────┘
|
|
```
|
|
|
|
### State
|
|
|
|
Active KBs per channel stored in `channel_knowledge_bases`. When a user
|
|
toggles a KB on/off for a channel, the frontend calls:
|
|
|
|
```
|
|
PUT /api/v1/channels/:id/knowledge-bases
|
|
Body: { kb_ids: ["uuid1", "uuid2"] }
|
|
```
|
|
|
|
The completion handler reads active KBs to populate the system prompt
|
|
hint and to scope `kb_search` results.
|
|
|
|
### Visibility Rules
|
|
|
|
A user can toggle KBs they have access to:
|
|
- **Global KBs** — all users
|
|
- **Team KBs** — team members
|
|
- **Personal KBs** — owner only
|
|
|
|
The list endpoint returns all KBs the user can see, with `enabled`
|
|
status per channel.
|
|
|
|
---
|
|
|
|
## 9. Store Interface
|
|
|
|
```go
|
|
// server/store/interfaces.go — additions
|
|
|
|
// =========================================
|
|
// KNOWLEDGE BASE STORE
|
|
// =========================================
|
|
|
|
type KnowledgeBaseStore interface {
|
|
// KB CRUD
|
|
Create(ctx context.Context, kb *models.KnowledgeBase) error
|
|
GetByID(ctx context.Context, id string) (*models.KnowledgeBase, error)
|
|
Update(ctx context.Context, id string, fields map[string]interface{}) error
|
|
Delete(ctx context.Context, id string) error
|
|
|
|
// Scoped listing
|
|
ListForUser(ctx context.Context, userID string, teamIDs []string) ([]models.KnowledgeBase, error)
|
|
ListGlobal(ctx context.Context) ([]models.KnowledgeBase, error)
|
|
ListForTeam(ctx context.Context, teamID string) ([]models.KnowledgeBase, error)
|
|
ListPersonal(ctx context.Context, userID string) ([]models.KnowledgeBase, error)
|
|
|
|
// Documents
|
|
CreateDocument(ctx context.Context, doc *models.KBDocument) error
|
|
GetDocument(ctx context.Context, id string) (*models.KBDocument, error)
|
|
ListDocuments(ctx context.Context, kbID string) ([]models.KBDocument, error)
|
|
UpdateDocumentStatus(ctx context.Context, id string, status string, err *string) error
|
|
DeleteDocument(ctx context.Context, id string) (*models.KBDocument, error) // returns for storage cleanup
|
|
|
|
// Chunks
|
|
InsertChunks(ctx context.Context, chunks []models.KBChunk) error
|
|
DeleteChunksForDocument(ctx context.Context, documentID string) error
|
|
SimilaritySearch(ctx context.Context, kbIDs []string, queryVec []float64,
|
|
threshold float64, limit int) ([]models.KBSearchResult, error)
|
|
|
|
// Channel links
|
|
SetChannelKBs(ctx context.Context, channelID string, kbIDs []string) error
|
|
GetChannelKBs(ctx context.Context, channelID string) ([]models.ChannelKB, error)
|
|
GetActiveKBIDs(ctx context.Context, channelID string, userID string,
|
|
teamIDs []string) ([]string, error) // enabled + user has access
|
|
|
|
// Stats
|
|
UpdateStats(ctx context.Context, kbID string) error // recount from chunks
|
|
}
|
|
```
|
|
|
|
Add `KnowledgeBases KnowledgeBaseStore` to the `Stores` struct.
|
|
|
|
---
|
|
|
|
## 10. Models
|
|
|
|
```go
|
|
// server/models/models.go — additions
|
|
|
|
type KnowledgeBase struct {
|
|
ID string `json:"id" db:"id"`
|
|
Name string `json:"name" db:"name"`
|
|
Description string `json:"description" db:"description"`
|
|
Scope string `json:"scope" db:"scope"`
|
|
OwnerID *string `json:"owner_id,omitempty" db:"owner_id"`
|
|
TeamID *string `json:"team_id,omitempty" db:"team_id"`
|
|
EmbeddingConfig map[string]interface{} `json:"embedding_config" db:"embedding_config"`
|
|
DocumentCount int `json:"document_count" db:"document_count"`
|
|
ChunkCount int `json:"chunk_count" db:"chunk_count"`
|
|
TotalBytes int64 `json:"total_bytes" db:"total_bytes"`
|
|
Status string `json:"status" db:"status"`
|
|
CreatedAt time.Time `json:"created_at" db:"created_at"`
|
|
UpdatedAt time.Time `json:"updated_at" db:"updated_at"`
|
|
}
|
|
|
|
type KBDocument struct {
|
|
ID string `json:"id" db:"id"`
|
|
KBID string `json:"kb_id" db:"kb_id"`
|
|
Filename string `json:"filename" db:"filename"`
|
|
ContentType string `json:"content_type" db:"content_type"`
|
|
SizeBytes int64 `json:"size_bytes" db:"size_bytes"`
|
|
StorageKey string `json:"storage_key" db:"storage_key"`
|
|
ExtractedText *string `json:"extracted_text,omitempty" db:"extracted_text"`
|
|
ChunkCount int `json:"chunk_count" db:"chunk_count"`
|
|
Status string `json:"status" db:"status"`
|
|
Error *string `json:"error,omitempty" db:"error"`
|
|
UploadedBy string `json:"uploaded_by" db:"uploaded_by"`
|
|
CreatedAt time.Time `json:"created_at" db:"created_at"`
|
|
UpdatedAt time.Time `json:"updated_at" db:"updated_at"`
|
|
}
|
|
|
|
type KBChunk struct {
|
|
ID string `json:"id" db:"id"`
|
|
KBID string `json:"kb_id" db:"kb_id"`
|
|
DocumentID string `json:"document_id" db:"document_id"`
|
|
ChunkIndex int `json:"chunk_index" db:"chunk_index"`
|
|
Content string `json:"content" db:"content"`
|
|
TokenCount int `json:"token_count" db:"token_count"`
|
|
Embedding []float64 `json:"-" db:"embedding"` // not serialized to API
|
|
Metadata map[string]interface{} `json:"metadata" db:"metadata"`
|
|
CreatedAt time.Time `json:"created_at" db:"created_at"`
|
|
}
|
|
|
|
type KBSearchResult struct {
|
|
Content string `json:"content"`
|
|
Filename string `json:"source"`
|
|
KBName string `json:"kb"`
|
|
Similarity float64 `json:"similarity"`
|
|
Metadata map[string]interface{} `json:"metadata,omitempty"`
|
|
}
|
|
|
|
type ChannelKB struct {
|
|
KBID string `json:"kb_id" db:"kb_id"`
|
|
KBName string `json:"kb_name" db:"name"`
|
|
Enabled bool `json:"enabled" db:"enabled"`
|
|
DocCount int `json:"document_count" db:"document_count"`
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## 11. File Structure
|
|
|
|
### New Go Packages
|
|
|
|
```
|
|
server/
|
|
├── knowledge/ # NEW — chunking + embedding logic
|
|
│ ├── chunker.go # recursive text splitter
|
|
│ ├── chunker_test.go
|
|
│ ├── embedder.go # batch embed via role resolver
|
|
│ └── ingest.go # orchestrator: extract → chunk → embed → store
|
|
├── handlers/
|
|
│ ├── knowledge_bases.go # NEW — KB CRUD + document upload + search
|
|
│ └── ... (existing)
|
|
├── store/
|
|
│ └── postgres/
|
|
│ ├── knowledge_bases.go # NEW — KnowledgeBaseStore impl
|
|
│ └── ... (existing)
|
|
├── tools/
|
|
│ ├── kbsearch.go # NEW — kb_search tool
|
|
│ └── ... (existing)
|
|
├── database/
|
|
│ └── migrations/
|
|
│ └── 008_knowledge_bases.sql # NEW
|
|
└── models/
|
|
└── models.go # additions (KB types)
|
|
```
|
|
|
|
### New/Modified Frontend Files
|
|
|
|
```
|
|
src/js/
|
|
├── ui-admin.js # MODIFIED — KB management section in admin panel
|
|
├── ui-primitives.js # MODIFIED — embedding dropdown fix, manual entry
|
|
├── chat.js # MODIFIED — KB toggle popup, system prompt hint
|
|
├── api.js # MODIFIED — KB API methods
|
|
└── settings-handlers.js # MODIFIED — personal KB management
|
|
```
|
|
|
|
### Go Dependencies
|
|
|
|
```
|
|
go get github.com/pgvector/pgvector-go # pgvector type support for lib/pq
|
|
```
|
|
|
|
---
|
|
|
|
## 12. Admin Panel Integration
|
|
|
|
KBs live under **AI → Knowledge Bases** in the admin panel (v0.13.0).
|
|
|
|
### Admin View
|
|
|
|
```
|
|
┌──────────────────────────────────────────────────────┐
|
|
│ AI → Knowledge Bases │
|
|
│ │
|
|
│ [+ New Knowledge Base] │
|
|
│ │
|
|
│ ┌────────────────────────────────────────────────┐ │
|
|
│ │ Q4 Reports global 3 docs ready │ │
|
|
│ │ Engineering Wiki global 12 docs ready │ │
|
|
│ │ Onboarding Materials team:Eng 5 docs ready │ │
|
|
│ └────────────────────────────────────────────────┘ │
|
|
│ │
|
|
│ ── Selected: Q4 Reports ────────────────────────── │
|
|
│ │
|
|
│ Name: [Q4 Reports ] │
|
|
│ Description: [Quarterly financial reports ] │
|
|
│ Scope: [Global ▾] Team: [— ▾] │
|
|
│ │
|
|
│ Embedding: OpenAI / text-embedding-3-small │
|
|
│ Chunk size: [1000] Overlap: [200] │
|
|
│ │
|
|
│ Documents: │
|
|
│ ┌────────────────────────────────────────────────┐ │
|
|
│ │ quarterly-report-q4.pdf 2.1 MB 42 chunks ✓ │ │
|
|
│ │ revenue-breakdown.xlsx 340 KB 12 chunks ✓ │ │
|
|
│ │ market-analysis.docx 1.8 MB processing... │ │
|
|
│ └────────────────────────────────────────────────┘ │
|
|
│ [📎 Upload Documents] [🔄 Rebuild All] [🗑 Delete] │
|
|
└──────────────────────────────────────────────────────┘
|
|
```
|
|
|
|
### Personal KB (Settings)
|
|
|
|
BYOK users see a "Knowledge Bases" section in Settings → Model Roles
|
|
(or a new Settings tab). Same UI, scoped to personal KBs using the
|
|
user's embedding role override.
|
|
|
|
### Team KB (Team Management)
|
|
|
|
Team admins see a "Knowledge Bases" tab in the team management modal.
|
|
Same UI pattern, scoped to the team.
|
|
|
|
---
|
|
|
|
## 13. Phased Delivery
|
|
|
|
### Phase 1: Schema + Chunker + Embed Fix
|
|
|
|
**Goal:** Database ready, chunking works, embedding dropdown fixed.
|
|
|
|
- [ ] `008_knowledge_bases.sql` migration (pgvector + tables)
|
|
- [ ] `go get pgvector-go`, update `go.mod`
|
|
- [ ] `knowledge/chunker.go` — recursive text splitter + tests
|
|
- [ ] Embedding dropdown fix in `ui-primitives.js` (tolerant filter + manual entry)
|
|
- [ ] `KnowledgeBaseStore` interface + postgres implementation (CRUD only, no vector ops)
|
|
- [ ] `KnowledgeBase`, `KBDocument`, `KBChunk` model types
|
|
- [ ] Wire `Stores.KnowledgeBases` in main.go
|
|
|
|
### Phase 2: Ingestion Pipeline
|
|
|
|
**Goal:** Documents can be uploaded, chunked, and embedded end-to-end.
|
|
|
|
- [ ] `knowledge/embedder.go` — batch embed via role resolver
|
|
- [ ] `knowledge/ingest.go` — orchestrator (extract → chunk → embed → store)
|
|
- [ ] `handlers/knowledge_bases.go` — KB CRUD endpoints
|
|
- [ ] Document upload with async ingestion goroutine
|
|
- [ ] Status polling endpoint
|
|
- [ ] `SimilaritySearch()` in store (pgvector query)
|
|
- [ ] Admin panel: KB management section (AI → Knowledge Bases)
|
|
- [ ] IVFFlat index creation (manual or after first batch)
|
|
|
|
### Phase 3: kb_search Tool + Channel Integration
|
|
|
|
**Goal:** LLMs can search KBs during conversations.
|
|
|
|
- [ ] `tools/kbsearch.go` — kb_search tool (late registration pattern)
|
|
- [ ] Per-channel KB toggle (channel_knowledge_bases)
|
|
- [ ] System prompt KB hint injection in completion handler
|
|
- [ ] KB toggle popup in chat UI
|
|
- [ ] Frontend: channel KB management
|
|
- [ ] Category "knowledge" in tools toggle menu
|
|
|
|
### Phase 4: Polish + Notes
|
|
|
|
**Goal:** Team/personal scoping works, notes get vectors.
|
|
|
|
- [ ] Team KB endpoints + team management UI tab
|
|
- [ ] Personal KB endpoints + settings UI
|
|
- [ ] KB rebuild endpoint (re-chunk + re-embed)
|
|
- [ ] Notes embedding column + semantic note search
|
|
- [ ] Audit logging for KB operations
|
|
- [ ] Usage tracking for embedding calls (role = 'embedding')
|
|
- [ ] Integration tests
|
|
|
|
---
|
|
|
|
## 14. Risks + Mitigations
|
|
|
|
| Risk | Impact | Mitigation |
|
|
|------|--------|------------|
|
|
| pgvector not installed on PG host | Migration fails, blocks deploy | Clear error message in migration; add to deployment docs/Dockerfile |
|
|
| Embedding provider rate limits | Ingestion stalls on large uploads | Semaphore + exponential backoff; batch size tuning |
|
|
| Dimension mismatch (model change) | Search returns garbage | Snapshot model in `embedding_config`; warn if role model ≠ KB model |
|
|
| Large documents (100+ pages) | Memory pressure during chunking | Stream text extraction; chunk incrementally |
|
|
| Vector storage size | DB bloat with many KBs | Monitor with stats endpoint; future: HNSW index for better perf |
|
|
| Embedding dropdown broken | Can't configure embedding role at all | Phase 1 priority — ship fix before KB features |
|
|
|
|
---
|
|
|
|
## 15. Future (Not v0.14.0)
|
|
|
|
- **Auto-injection:** Pre-pend top-K results to system prompt (v0.15.x, needs compaction for context budget)
|
|
- **Hybrid search:** Combine vector similarity with full-text `tsvector` search, re-rank
|
|
- **Semantic chunking:** Use embedding distance to detect topic boundaries
|
|
- **HNSW index:** Better query performance than IVFFlat for large datasets
|
|
- **Web scraping source:** Ingest URLs as KB documents (natural extension of url_fetch)
|
|
- **Scheduled re-indexing:** Periodic rebuild when source documents are updated
|
|
- **KB sharing:** Cross-team KB access grants
|