47 KiB
v0.12.0 — File Handling + Vision (Design)
Overview
File input into chat — table stakes for serious use. Image uploads for vision-capable models, document uploads for text extraction into context, and the storage backend that v0.14.0 (Knowledge Bases) and v0.15.0 (Compaction snapshots) will reuse.
Also picks up stragglers deferred from earlier releases.
Depends on: v0.11.0 (extension foundation — complete). Reused by: v0.14.0 (KBs), v0.15.0 (compaction), v0.22.0 (exports).
Stragglers (from v0.9.4 deferred)
Items deferred during vault work that were never scheduled:
switchboard vault rekey CLI command
Re-encrypts all global/team API keys when ENCRYPTION_KEY is rotated.
Personal BYOK keys are unaffected (keyed to UEK, not env var).
Implementation: standalone CLI entrypoint that reads old + new env vars,
iterates api_configs and team_providers where key_scope IN ('global', 'team'), decrypts with old key, re-encrypts with new key, updates in a
transaction.
Admin UI: encryption status indicator
Badge in admin Settings showing whether ENCRYPTION_KEY is set and how
many keys are encrypted vs. plaintext (migration stragglers). Single
GET /admin/storage/status endpoint.
Per-chat model/preset persistence (server-side)
Currently localStorage bandaid from v0.10.2 — doesn't roam across devices.
Fix: Store last_selector_id in channels.settings JSONB on each
completion success. On chat selection, resolve:
channel.settings.last_selector_id → match against available
models/presets → fall back to channel.model → global default.
No migration needed (channels.settings JSONB column already exists).
Single PATCH /channels/:id with settings merge on completion success.
Track 1: Storage Backend
Environment Variables
| Variable | Default | Description |
|---|---|---|
STORAGE_BACKEND |
(auto) | pvc or s3. If not set: pvc when STORAGE_PATH is writable, disabled otherwise. |
STORAGE_PATH |
/data/storage |
Mount point for PVC backend. Also scratch dir for extraction with S3. |
STORAGE_CLASS |
— | K8s only. Gitea CI variable → PVC manifest. Value: cephfs (RWX for multi-pod). |
S3_ENDPOINT |
— | S3-compatible endpoint (e.g. http://minio:9000). Required when STORAGE_BACKEND=s3. |
S3_BUCKET |
— | Bucket name. Must exist before first start. |
S3_ACCESS_KEY |
— | S3 access key ID. |
S3_SECRET_KEY |
— | S3 secret access key. |
S3_REGION |
us-east-1 |
AWS region. |
S3_PREFIX |
— | Optional key prefix for shared buckets (e.g. switchboard/). |
S3_FORCE_PATH_STYLE |
true |
Path-style URLs. Required for MinIO, Ceph RGW, most self-hosted. |
EXTRACTION_MODE |
inline |
inline (in-process, unified image) or sidecar (K8s, shared PVC). |
EXTRACTION_CONCURRENCY |
3 |
Max concurrent extraction jobs. Caps LibreOffice memory usage. |
Auto-detection when STORAGE_BACKEND is not set:
STORAGE_PATH writable → pvc (implicit)
STORAGE_PATH missing → file features disabled
admin panel shows "Storage not configured"
When STORAGE_BACKEND=pvc is explicit, fail startup if path is not
writable (fail-safe, same pattern as ENCRYPTION_KEY enforcement).
S3 backend: reads S3_ENDPOINT, S3_BUCKET, S3_ACCESS_KEY, S3_SECRET_KEY,
S3_REGION, S3_PREFIX, S3_FORCE_PATH_STYLE from env vars. Uses minio-go v7.
Works with MinIO, Ceph RGW, AWS S3. Same interface as PVC — all handlers are
backend-agnostic. PVC still required as scratch dir for extraction queue.
Config Addition
// In config.go
StorageBackend string // "pvc", "s3", or "" (auto-detect)
StoragePath string // mount point for PVC backend
ExtractionMode string // "inline" or "sidecar"
ExtractionConcurrency int // max concurrent extraction jobs
Interface
// server/storage/storage.go
package storage
import (
"context"
"io"
)
// ObjectStore is the abstraction for blob storage.
// Implementations: PVC (filesystem), S3 (minio-go v7).
type ObjectStore interface {
// Put writes data to the given key. Creates parent dirs as needed.
Put(ctx context.Context, key string, r io.Reader, size int64, contentType string) error
// Get returns a reader for the given key. Caller must close.
// Returns ErrNotFound if key does not exist.
Get(ctx context.Context, key string) (io.ReadCloser, int64, string, error)
// Delete removes the object at key. No error if already gone.
Delete(ctx context.Context, key string) error
// DeletePrefix removes all objects under the given prefix.
// Used for channel deletion (bulk cleanup).
DeletePrefix(ctx context.Context, prefix string) error
// Exists checks if an object exists at key without reading it.
Exists(ctx context.Context, key string) (bool, error)
// Healthy returns nil if the backend is operational.
Healthy(ctx context.Context) error
}
var ErrNotFound = errors.New("storage: object not found")
PVC Implementation
server/storage/pvc.go — ~100 lines. Thin wrapper around os.*:
type PVCStore struct {
basePath string // e.g. "/data/storage"
}
func NewPVC(basePath string) (*PVCStore, error) {
// Validate basePath is writable (create test file, remove)
// Create top-level subdirs: attachments/
}
Key → filesystem path: filepath.Join(basePath, key). The key itself
provides all the directory structure.
Filesystem Layout
/data/storage/ ← STORAGE_PATH mount point
└── attachments/
└── {channel_id}/
└── {attachment_id}_{filename}
Channel-first because:
- Channel deletion →
DeletePrefix("attachments/{channel_id}/")— one call. - User deletion cascades to owned channels → same cleanup path.
- Group chats / multi-participant channels: all files in one place regardless of who uploaded them.
- Quota enforcement uses PG (
SUM(size_bytes) WHERE user_id = $1), not filesystem walks.
Future subdirs at the STORAGE_PATH level (not this release):
knowledge-bases/— v0.14.0compaction/— v0.15.0exports/— TBD
K8s Manifest (PVC)
# k8s/storage-pvc.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: switchboard-storage
namespace: switchboard
spec:
accessModes: [ReadWriteMany] # CephFS RWX — multi-pod safe
storageClassName: "${STORAGE_CLASS}"
resources:
requests:
storage: 10Gi # configurable per deployment
Backend deployment addition:
# In k8s/backend.yaml
volumes:
- name: storage
persistentVolumeClaim:
claimName: switchboard-storage
containers:
- name: backend
volumeMounts:
- name: storage
mountPath: /data/storage
env:
- name: STORAGE_BACKEND
value: "pvc"
- name: STORAGE_PATH
value: "/data/storage"
Docker Compose
# In docker-compose.yml
services:
switchboard:
volumes:
- ./data/storage:/data/storage
environment:
- STORAGE_PATH=/data/storage
Admin Status Endpoint
GET /admin/storage/status
{
"backend": "pvc",
"path": "/data/storage",
"healthy": true,
"total_files": 142,
"total_bytes": 52428800,
"configured": true
}
When storage is not configured, file upload endpoints return 503 with
{"error": "File storage not configured"}. Frontend hides upload UI
based on a capabilities flag in the boot payload.
Track 2: Attachments Schema + API
Migration 007_attachments.sql
-- ── Attachments ────────────────────────────
CREATE TABLE IF NOT EXISTS attachments (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
channel_id UUID NOT NULL REFERENCES channels(id) ON DELETE CASCADE,
user_id UUID NOT NULL REFERENCES users(id),
message_id UUID REFERENCES messages(id) ON DELETE SET NULL,
filename VARCHAR(255) NOT NULL,
content_type VARCHAR(127) NOT NULL,
size_bytes BIGINT NOT NULL,
storage_key TEXT NOT NULL,
extracted_text TEXT,
metadata JSONB DEFAULT '{}'::jsonb,
created_at TIMESTAMPTZ DEFAULT NOW()
);
-- Access pattern: always join through channel for auth
CREATE INDEX idx_attachments_channel ON attachments(channel_id);
-- Quota pattern: sum by user
CREATE INDEX idx_attachments_user_size ON attachments(user_id);
-- Message association: find attachments for a message
CREATE INDEX idx_attachments_message ON attachments(message_id);
Notes on columns:
message_idis nullable — attachment is uploaded before the message is persisted. Updated after message creation. If message is deleted, attachment persists (belongs to channel, not message).storage_keyis the relative path withinSTORAGE_PATH, e.g.attachments/{channel_id}/{attachment_id}_{filename}.extracted_textis populated asynchronously for PDFs/DOCX. NULL for images. Used for context injection into completions.metadataJSONB holds dimensions (images), page count (PDFs), extraction status, thumbnails, etc.
Store Layer
// server/store/attachments.go
type AttachmentStore interface {
Create(ctx context.Context, a *models.Attachment) error
GetByID(ctx context.Context, id string) (*models.Attachment, error)
GetByChannel(ctx context.Context, channelID string) ([]models.Attachment, error)
GetByMessage(ctx context.Context, messageID string) ([]models.Attachment, error)
SetMessageID(ctx context.Context, id, messageID string) error
Delete(ctx context.Context, id string) error
DeleteByChannel(ctx context.Context, channelID string) ([]string, error) // returns storage_keys for cleanup
UserUsageBytes(ctx context.Context, userID string) (int64, error)
}
API Endpoints
POST /api/v1/channels/:id/attachments Upload file
GET /api/v1/attachments/:id Attachment metadata
GET /api/v1/attachments/:id/download Download file content
DELETE /api/v1/attachments/:id Delete attachment
GET /api/v1/channels/:id/attachments List channel attachments
Upload Flow
Client Backend Storage
│ │ │
├── POST multipart ────────►│ │
│ (file + channel_id) │ │
│ ├── Auth check │
│ ├── Channel membership ─────│
│ │ verify │
│ ├── Size limit check │
│ ├── MIME type validation │
│ ├── Generate attachment_id │
│ ├── PUT ───────────────────►│
│ │ key: att/{ch}/{id}_{fn} │
│ ├── INSERT attachments row │
│ │ │
│◄── 201 { attachment } ────┤ │
The attachment is created before the message is sent. The frontend holds the attachment ID(s), includes them in the completion request, and the backend links them to the persisted message after creation.
Download Flow (Access Control)
GET /api/v1/attachments/:id/download
→ SELECT a.*, c.id, c.user_id
FROM attachments a
JOIN channels c ON a.channel_id = c.id
WHERE a.id = $1
→ Verify: requester is channel owner
(future RBAC: OR requester is channel participant
OR requester is team member with channel access
OR requester is admin with audit log)
→ Stream file from ObjectStore with:
Content-Type: {content_type}
Content-Disposition: attachment; filename="{filename}"
Content-Length: {size_bytes}
Critical: Files are NEVER served directly from nginx or any static file path. Every byte goes through the Go backend with auth checks. No pre-signed URLs, no direct filesystem exposure.
Size & Type Limits
Configured via global settings (see Global Settings section below). Enforced server-side on upload — frontend mirrors limits for UX but the backend is authoritative.
Validation order on upload:
storage_allowed_types— reject disallowed MIME typesstorage_max_file_size— reject oversized individual filesstorage_max_upload_size— reject if request total exceeds limitstorage_max_attachments_per_message— reject if too many filesstorage_user_quota_bytes— reject if user quota exceeded (future)
MIME type is detected server-side (http.DetectContentType + extension
validation), not trusted from the client Content-Type header.
Track 3: Multimodal Message Assembly
Provider Message Polymorphism
Currently providers.Message.Content is a plain string. Vision requires
content arrays with image parts. The change is additive — no existing
callsites break.
// In server/providers/provider.go
// ContentPart represents one part of a multimodal message.
type ContentPart struct {
Type string `json:"type"` // "text", "image_url"
Text string `json:"text,omitempty"` // for type="text"
ImageURL *ImageURL `json:"image_url,omitempty"` // for type="image_url"
}
type ImageURL struct {
URL string `json:"url"` // "data:image/jpeg;base64,..." or URL
Detail string `json:"detail,omitempty"` // "auto", "low", "high"
}
type Message struct {
Role string `json:"role"`
Content string `json:"content"`
// Multimodal content. When non-nil, providers use this instead of Content.
ContentParts []ContentPart `json:"content_parts,omitempty"`
// ... existing tool fields unchanged
}
Resolution rule: if ContentParts is non-nil, providers build their
wire format from it. If nil, they use Content (string) as today. This
means every existing codepath works unchanged until an attachment is
involved.
Provider Wire Format Mapping
OpenAI / OpenRouter / Venice (OpenAI-compatible):
// In openai.go — message conversion
if len(m.ContentParts) > 0 {
// Content becomes array of objects
oaiMsg.Content = nil // clear string content
oaiMsg.ContentArray = toOpenAIContentParts(m.ContentParts)
}
OpenAI wire format:
{
"role": "user",
"content": [
{"type": "text", "text": "What's in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
]
}
The openaiMessage struct needs a ContentArray field with custom JSON
marshaling — when ContentArray is set, marshal content as the array;
otherwise marshal as string. This is the standard OpenAI multimodal
format.
Anthropic:
// In anthropic.go — message conversion
if len(m.ContentParts) > 0 {
antMsg.Content = toAnthropicContentBlocks(m.ContentParts)
}
Anthropic wire format:
{
"role": "user",
"content": [
{"type": "text", "text": "What's in this image?"},
{"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": "..."}}
]
}
The anthropicContentBlock struct gets new fields:
type anthropicContentBlock struct {
// ... existing fields ...
// Image support
Source *anthropicImageSource `json:"source,omitempty"`
}
type anthropicImageSource struct {
Type string `json:"type"` // "base64"
MediaType string `json:"media_type"` // "image/jpeg", "image/png", etc.
Data string `json:"data"` // base64-encoded image
}
Image Delivery: Base64 Inline
For v0.12.0, all images are sent as base64-encoded data inline in the API request. No pre-signed URLs.
Rationale: The target audience is government/enterprise, often in air-gapped or restricted networks. Pre-signed URLs require the provider's API servers to reach the storage endpoint — impossible in disconnected environments. Base64 works everywhere.
Trade-off: Larger request payloads. Acceptable for v0.12.0 — images are typically <10MB each and context windows are large. URL-based delivery can be an optimization in a future release for cloud deployments.
Completion Handler Changes
// In completion.go — after loading conversation, before building provReq
// If message has attachments, build multimodal content
if len(req.AttachmentIDs) > 0 {
parts := []providers.ContentPart{
{Type: "text", Text: req.Content},
}
for _, attID := range req.AttachmentIDs {
att, err := h.store.Attachments().GetByID(c, attID)
// ... verify att.channel_id matches ...
if isImageType(att.ContentType) {
// Verify vision capability
if !caps.Vision {
c.JSON(400, gin.H{"error": "Selected model does not support image input"})
return
}
// Read from storage, base64 encode
reader, _, _, _ := h.objectStore.Get(c, att.StorageKey)
data, _ := io.ReadAll(reader)
reader.Close()
b64 := base64.StdEncoding.EncodeToString(data)
dataURL := fmt.Sprintf("data:%s;base64,%s", att.ContentType, b64)
parts = append(parts, providers.ContentPart{
Type: "image_url",
ImageURL: &providers.ImageURL{URL: dataURL, Detail: "auto"},
})
} else if att.ExtractedText != nil && *att.ExtractedText != "" {
// Document: inject extracted text as context
parts = append(parts, providers.ContentPart{
Type: "text",
Text: fmt.Sprintf("[Document: %s]\n%s", att.Filename, *att.ExtractedText),
})
} else {
// Extraction not yet complete or failed — include filename only
status := att.Metadata["extraction_status"]
parts = append(parts, providers.ContentPart{
Type: "text",
Text: fmt.Sprintf("[Attached file: %s (extraction %s)]", att.Filename, status),
})
}
}
// Use ContentParts instead of Content string
userMsg := providers.Message{
Role: "user",
ContentParts: parts,
}
messages = append(messages, userMsg)
} else {
// Existing text-only path — unchanged
messages = append(messages, providers.Message{
Role: "user",
Content: req.Content,
})
}
Frontend blocks send only during active upload — the send button
stays disabled while any staged attachment is in uploading state
(bytes still in flight). Once the file is on disk and the attachment
row exists in PG, send is enabled regardless of extraction status.
The completion handler checks extraction_status at request time:
complete→ inject extracted text as context- anything else → inject filename-only placeholder
This avoids blocking users on slow extractions while still providing extracted context when available. For images (no extraction), the attachment is immediately usable.
Capability Gating
Images are only assembled into content parts when the resolved model
has capabilities.vision == true. If a user attaches an image to a
non-vision model:
- Documents (PDF, DOCX, TXT): always allowed — extracted text is injected as a text content part. No vision needed.
- Images to non-vision model: backend returns 400 with
{"error": "Selected model does not support image input"}. - Frontend disables image upload button when selected model lacks vision.
Text Extraction (Documents)
License Constraint
The project is Apache 2.0. Libraries with AGPL/GPL copyleft (e.g.
unipdf) are incompatible — their license would infect the codebase.
Pure Go PDF libraries under permissive licenses (ledongthuc/pdf,
pdfcpu) handle simple cases but choke on complex layouts, scanned
PDFs, and encrypted documents.
Approach: Tiered Extraction
Tier 1 — Native Go (no external deps, handles ~80% of uploads):
| Format | Library | License |
|---|---|---|
| TXT, Markdown, CSV | io.ReadAll |
— |
| PDF (simple) | pdfcpu |
Apache 2.0 |
| DOCX | fumiama/go-docx |
MIT |
Tier 2 — LibreOffice Headless (handles everything else):
For PDF with complex layouts, ODT, ODS, ODP, RTF, DOC (legacy), XLS, and any format Tier 1 can't handle — shell out to LibreOffice headless:
libreoffice --headless --convert-to txt:Text /tmp/input.odt --outdir /tmp/
LibreOffice handles the full Office family:
- OpenDocument:
.odt,.ods,.odp - Microsoft:
.doc,.docx,.xls,.xlsx,.ppt,.pptx - PDF (including scanned with basic OCR via hunspell)
- RTF, HTML, and dozens more
Extraction Queue (Filesystem-Backed)
Extraction state lives on the PVC, not in memory. Pod restarts don't lose queue state. The queue is a directory convention:
/data/storage/
├── attachments/ ← completed uploads (served to clients)
└── processing/ ← extraction queue (internal only)
└── {attachment_id}/
├── input ← original file (copy or hardlink)
├── status.json ← queue state + metadata
├── output.txt ← extracted text (when complete)
└── thumb.jpg ← thumbnail (when complete)
status.json:
{
"state": "pending",
"queued_at": "2026-02-25T14:30:00Z",
"started_at": null,
"completed_at": null,
"content_type": "application/pdf",
"filename": "report.pdf",
"error": null
}
States: pending → processing → complete | failed
Worker loop (runs in both inline and sidecar mode):
1. Scan processing/ for dirs with status.state == "pending"
(ordered by queued_at — FIFO)
2. Respect concurrency cap (semaphore, default 3)
3. Pick job → flip state to "processing"
4. Tier 1: try native Go extraction
├── success → write output.txt, generate thumb, flip "complete"
└── fail or unsupported →
5. Tier 2: LibreOffice headless
├── success → write output.txt, generate thumb, flip "complete"
└── fail → flip "failed", write error
6. Copy results to PG:
├── UPDATE attachments SET extracted_text = ...,
│ metadata = jsonb_set(metadata, '{extraction_status}', '"complete"')
└── Thumbnail storage_key written to metadata
7. Clean up processing/{id}/ dir
In inline mode (Unified image): the worker is a goroutine pool in the Go backend, scanning the queue directory on a 2-second tick.
In sidecar mode (K8s): the sidecar runs the same loop as a
standalone process. When extraction completes, it writes output.txt
and thumb.jpg to the processing dir and flips status to complete.
The Go backend polls the status file and copies results to PG.
Crash recovery: On startup, scan for state == "processing" (stale
from a crash). Reset to pending for re-processing. This is why the
queue is on the filesystem — nothing is lost.
Concurrency: Capped by semaphore. Default 3 concurrent extractions.
Configurable via EXTRACTION_CONCURRENCY env var. LibreOffice is
memory-hungry on complex docs — the cap prevents OOM on burst uploads.
Frontend queue awareness:
Attachment status:
uploading → progress bar
pending → "Queued (2 ahead)" ← position from queue scan
processing → "Extracting..." ← spinner
complete → ✓ (ready)
failed → ⚠ (send still enabled, model sees filename only)
The frontend polls GET /api/v1/attachments/:id every 2s for
non-terminal states. The response includes extraction_status and
queue_position (0 = currently processing, N = N jobs ahead).
Send is blocked while any attachment is uploading. Once all are at
least pending, send is enabled — the extraction results will be
available by the time the model processes the request (optimistic), or
the model sees the filename-only fallback (acceptable degradation).
Correction to earlier design: Send is blocked only during uploading
phase. Once the file is on disk and queued, the user can send. The
completion handler checks extraction status at request time and uses
whatever is available.
Extraction Metadata in PG
Attachment metadata JSONB tracks status (mirrored from queue):
{
"extraction_status": "complete",
"extraction_error": null,
"queue_position": 0,
"page_count": 12,
"dimensions": {"width": 1920, "height": 1080},
"thumbnail_key": "attachments/{channel_id}/{id}_thumb.jpg",
"thumb_dimensions": {"width": 280, "height": 147}
}
States: pending → processing → complete | failed | unavailable
LibreOffice Deployment Strategy
Two deployment modes, matching existing Docker architecture:
| Deployment | LibreOffice location | Reason |
|---|---|---|
| Unified (docker-compose, dev) | Baked into image | Single container, simplicity wins |
| K8s (backend + frontend split) | Sidecar container | Separate scaling, lean backend image, independent version pinning |
Unified Dockerfile addition:
# In Dockerfile (unified)
RUN apt-get update && apt-get install -y --no-install-recommends \
libreoffice-core libreoffice-writer libreoffice-calc \
&& rm -rf /var/lib/apt/lists/*
K8s sidecar (backend.yaml addition):
# In k8s/backend.yaml
containers:
- name: backend
image: switchboard-backend:0.12.0
volumeMounts:
- name: storage
mountPath: /data/storage
env:
- name: STORAGE_BACKEND
value: "pvc"
- name: STORAGE_PATH
value: "/data/storage"
- name: EXTRACTION_MODE
value: "sidecar" # "inline" for unified, "sidecar" for k8s
- name: extractor
image: switchboard-extractor:0.12.0 # minimal image: alpine + libreoffice
volumeMounts:
- name: storage
mountPath: /data/storage
resources:
requests:
memory: "256Mi"
limits:
memory: "1Gi" # LibreOffice can spike on complex docs
volumes:
- name: storage
persistentVolumeClaim:
claimName: switchboard-storage
The sidecar image is small and purpose-built: Alpine + LibreOffice + a watcher script. No Go code, no API surface. Upgrades to LibreOffice are a tag bump on the sidecar image — backend image untouched.
Communication: Filesystem-only via the shared PVC. No gRPC, no HTTP between containers. See Extraction Queue below.
Allowed Types (Updated)
{
"storage_allowed_types": [
"image/jpeg", "image/png", "image/gif", "image/webp", "image/svg+xml",
"application/pdf",
"text/plain", "text/markdown", "text/csv",
"application/vnd.openxmlformats-officedocument.wordprocessingml.document",
"application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
"application/vnd.openxmlformats-officedocument.presentationml.presentation",
"application/vnd.oasis.opendocument.text",
"application/vnd.oasis.opendocument.spreadsheet",
"application/vnd.oasis.opendocument.presentation",
"application/rtf",
"application/msword",
"application/vnd.ms-excel"
]
}
Thumbnail Generation
Thumbnails are generated asynchronously alongside text extraction as part of the same post-upload processing pipeline.
Image Thumbnails
Use golang.org/x/image/draw (BSD, stdlib-adjacent) for resizing:
// Resize to max 280px on longest edge, preserve aspect ratio
// Output: JPEG at 80% quality
// Stored at: attachments/{channel_id}/{attachment_id}_thumb.jpg
No external dependency. Pure Go. Fast enough for on-upload generation.
Document Thumbnails
When LibreOffice is available, generate a first-page preview:
# Convert first page to PDF, then to image
libreoffice --headless --convert-to pdf /tmp/input.docx --outdir /tmp/
# Then use Go image libraries to render first page
# Or: convert to PNG directly via LibreOffice draw filter
libreoffice --headless --convert-to png /tmp/input.docx --outdir /tmp/
For PDFs specifically, pdfcpu can extract page images under Apache 2.0.
When LibreOffice is not available, documents get a generic file-type icon (handled in CSS, no thumbnail row needed).
Storage
Thumbnails stored alongside originals with _thumb suffix:
/data/storage/attachments/{channel_id}/
├── {id}_report.pdf ← original
├── {id}_report.pdf_thumb.jpg ← thumbnail
├── {id}_photo.jpg ← original
└── {id}_photo.jpg_thumb.jpg ← thumbnail
Thumbnail URL returned in attachment metadata:
{
"thumbnail_key": "attachments/{channel_id}/{id}_photo.jpg_thumb.jpg",
"dimensions": {"width": 1920, "height": 1080},
"thumb_dimensions": {"width": 280, "height": 147}
}
Download endpoint: GET /api/v1/attachments/:id/thumbnail — same
channel-scoped auth check as the full download.
Track 4: Frontend
Input Area Changes
The <div class="input-wrap"> gets an attachment button and a preview
strip:
┌──────────────────────────────────────────────┐
│ ┌──────────┐ ┌──────────┐ │ ← attachment preview strip
│ │ image.png│ │ report.pdf│ │ (visible only when files
│ │ ✕ │ │ ✕ │ │ are staged)
│ └──────────┘ └──────────┘ │
├──────────────────────────────────────────────┤
│ 📎 │ Send a message... │ ■ │ ▶ │ ← existing input + new clip btn
└──────────────────────────────────────────────┘
Upload Interactions
| Trigger | Behavior |
|---|---|
| 📎 button click | File picker dialog (filtered by allowed types) |
| Paste (Ctrl+V) binary | Auto-attach: screenshot/image → image attachment ({uuid}.{ext}) |
| Paste (Ctrl+V) large text | Auto-attach: text > threshold → text file ({uuid}.txt) |
| Paste (Ctrl+V) small text | Normal paste into textarea (below threshold) |
| Drag-and-drop on chat area | Auto-attach file(s) as attachments (max 5 per message) |
Smart Paste (Auto-Attach)
Paste events are intercepted and routed automatically — no prompts, no user decisions. The clipboard content determines the behavior:
Detection logic:
Clipboard paste event
│
├── Contains files/images (binary)?
│ └── YES → auto-attach each item
│ ├── Image (screenshot, copied image) → upload as image attachment
│ ├── File (dragged from OS) → upload as file attachment
│ └── Filename: {uuid}.{ext} (ext from MIME type)
│
├── Contains text?
│ ├── Length ≤ threshold → normal paste into textarea
│ └── Length > threshold → auto-attach as text file
│ ├── Upload as attachment (content_type: text/plain)
│ ├── Filename: {uuid}.txt
│ └── Textarea stays empty (or keeps existing draft)
│
└── Empty → no-op
Binary detection: The paste event's clipboardData.items array
carries type information. Items with type.startsWith('image/') or
any non-text type are binary — always auto-attached. Text items
(text/plain, text/html) are checked against the threshold.
UUID filenames: All auto-attached pastes get UUID-based filenames.
No timestamps, no pasted-text.txt collision issues. The original
content type and paste source are captured in attachment metadata:
{
"source": "clipboard",
"original_type": "image/png",
"paste_length": 4200
}
Threshold: storage_paste_to_file_chars global setting (default:
2000 characters, admin-configurable). Setting to 0 disables text
auto-attach (binary pastes still auto-attach regardless).
UX flow for a screenshot paste:
- User hits Ctrl+V with a screenshot on clipboard
- Image appears instantly in the attachment preview strip (uploading...)
- Upload completes → thumbnail shown with checkmark
- User types their question, hits send
- Image goes to vision model as base64
UX flow for a large log paste:
- User copies 5,000 chars of server logs, hits Ctrl+V
- Text file chip appears in attachment preview strip:
📄 {uuid}.txt (4.9 KB) - Extraction runs (trivial for plain text — near-instant)
- User types "what's wrong with these logs?", hits send
- Log content injected as extracted text context
No prompts, no dialogs, no decisions. Paste just works.
State Management
Staged attachments live in App.stagedAttachments[] — an array of
{id, filename, contentType, sizeBytes, previewUrl, extractionStatus}
populated after the upload API returns. On send, attachment IDs are
included in the completion request body. On cancel/clear, staged
attachments are deleted via the API.
Upload lifecycle per attachment:
uploading → uploaded/queued → extracting → ready | failed
uploading: progress bar visible, send disabled.uploaded/queued: file on disk, extraction queued. Queue position shown: "Queued (2 ahead)". Send enabled — extraction will complete before or during model processing, or the model gets the filename-only fallback.extracting: spinner on thumbnail. Send remains enabled.ready: extraction complete (or image — no extraction needed). Green checkmark.failed: extraction failed. Warning icon on thumbnail. Send still enabled — model sees[Attached file: X (extraction failed)].
Key principle: Send is blocked only during active upload (bytes in flight). Once the file is on disk and the attachment row exists, the user can send. This prevents the frustrating case where a user types their message while a 50-page PDF extracts for 30 seconds.
Polling: Frontend polls GET /api/v1/attachments/:id every 2s for
any attachment with extractionStatus == "pending" | "processing".
Stops when all settle. WebSocket notification is a future optimization.
Polling updates the queue position indicator and thumbnail status icon.
Message Rendering
Attachments on sent/received messages render inline:
- Images: Thumbnail (from
_thumb.jpg) with click-to-expand (lightbox or side panel). Rendered in the message bubble below/above the text. - Documents: File chip with icon, filename, size, and download link.
📄 report.pdf (2.3 MB) ⬇ - Pasted text files: Same as documents —
📄 {uuid}.txt (4.9 KB) ⬇. Metadatasource: "clipboard"available for future UX refinement.
Attachment data comes from messages.metadata.attachment_ids →
batch-fetched on conversation load via
GET /channels/:id/attachments.
Capability-Aware UI
The frontend already knows model capabilities from the model selector.
Images are always auto-attached regardless of current model — the
user might switch models before sending. When the selected model has
vision: false:
- Image attachments show a subtle indicator:
👁️🗨️icon dimmed with tooltip "Current model doesn't support images" - File picker filters to documents only, but paste/drop still accepts images (user can switch models)
- The 📎 button tooltip says "Attach document" instead of "Attach file"
- On send with image + non-vision model: backend returns 400 with
"Selected model does not support image input — switch to a vision model or remove image attachments" - Document uploads (PDF, DOCX, ODT, etc.) always work — text extraction doesn't require vision
When storage is not configured (boot payload flag), the 📎 button is hidden entirely and all paste/drop auto-attach handlers are bypassed.
Privacy & Access Control Model
Principle
Attachments have no access identity of their own. They inherit access from their parent channel. The channel membership check is the single gate for all attachment operations.
Access Matrix (v0.12.0 — Current Auth Model)
| Channel type | Who can upload | Who can download | Delete |
|---|---|---|---|
| Personal DM | Channel owner | Channel owner | Owner |
| (Future) Team channel | Team members with access | Team members with access | Uploader or team admin |
| (Future) Group chat | Channel participants | Channel participants | Uploader or channel owner |
| (Future) Workflow | Assigned participants | Assigned participants | Stage rules |
Enforcement Points
Every attachment endpoint performs channel-scoped authorization:
// Pseudocode — every handler follows this pattern
func (h *Handler) downloadAttachment(c *gin.Context) {
att := h.store.Attachments().GetByID(c, attID)
channel := h.store.Channels().GetByID(c, att.ChannelID)
// THE GATE: is the requester authorized for this channel?
if channel.UserID != requesterID {
c.JSON(403, gin.H{"error": "access denied"})
return
}
// Future RBAC (v0.20.0): replace the above with:
// if !h.rbac.CanAccess(requesterID, channel.ID, "attachment:read") {
// c.JSON(403, ...)
// }
h.objectStore.Get(c, att.StorageKey) // → stream to client
}
Threat Mitigations
| Threat | Mitigation |
|---|---|
| ID guessing / enumeration | Every download joins through channel + verifies requester membership. UUIDs are non-sequential. |
| Direct filesystem access | Files never served by nginx. All bytes go through Go backend with auth. No static file routes for storage path. |
| Cross-team data leak | Channel membership is the boundary. HR channel attachments are only accessible to HR channel participants. |
| Admin overreach | v0.12.0: admin can access (consistent with existing admin trust model). Audit-logged. v0.20.0: RBAC with team-scoped admin. |
| Metadata leak | Channel attachment list endpoint also requires channel membership. No global search across attachments. |
| Storage key prediction | Keys include UUID attachment_id — not guessable from channel_id + filename alone. |
| Upload as attack vector | MIME type validated against allowlist. File size enforced. Content-Type set from server-side detection, not client header. |
Consistency with Existing Scoped Resources
The attachment access pattern is identical to the pattern used for all scoped data in the system:
Resource Scope Access Gate
──────────────────────────────────────────────────
Notes (personal) user_id user_id = requester
Notes (team) team_id requester ∈ team members
Attachments channel_id requester ∈ channel (owner/participant)
Usage data user/team scoped queries
API keys user/org vault (UEK) / env-derived key
The row knows its scope. The handler enforces the boundary. The query filters by it.
Future RBAC Hooks (v0.20.0)
The v0.12.0 access checks are written as explicit if checks that
can be trivially replaced by a centralized RBAC call:
// v0.12.0
if channel.UserID != requesterID { deny }
// v0.20.0
if !rbac.Can(requesterID, channel.ID, permission) { deny }
The schema, storage layer, and API surface don't change. Only the authorization check function gets swapped.
Implementation Order
Phase 1: Storage Backend
├── Config additions (STORAGE_BACKEND, STORAGE_PATH, EXTRACTION_MODE)
├── ObjectStore interface
├── PVC implementation (~100 lines)
├── Startup validation (writable check, health)
├── K8s manifest (PVC + volume mount + STORAGE_CLASS)
├── Docker-compose volume mount
├── Admin status endpoint
└── Boot payload: storage_configured flag
Phase 2: Attachments CRUD + Extraction Pipeline
├── Migration 007_attachments.sql
├── models.Attachment struct
├── AttachmentStore (PG CRUD)
├── handlers/attachments.go (upload, download, delete, list, thumbnail)
├── Channel-scoped access checks on all endpoints
├── MIME validation + size enforcement (server-side detection)
├── Filesystem cleanup on channel delete (CASCADE + DeletePrefix)
├── Extraction queue (filesystem-backed):
│ ├── processing/ directory convention + status.json
│ ├── Worker loop with semaphore (default 3 concurrent)
│ ├── Tier 1: native Go (pdfcpu, go-docx, io.ReadAll)
│ ├── Tier 2: LibreOffice headless (inline or sidecar)
│ ├── Crash recovery (reset stale "processing" → "pending")
│ └── PG metadata mirror (extraction_status, queue_position)
├── Thumbnail generation (images: Go resize, docs: LibreOffice)
├── Unified Dockerfile: add libreoffice-core
├── Sidecar Dockerfile: alpine + libreoffice + watcher
└── K8s backend.yaml: sidecar container definition
Phase 3: Multimodal Message Assembly
├── ContentPart + ImageURL types in providers/provider.go
├── OpenAI message conversion (content array marshaling)
├── Anthropic message conversion (image source blocks)
├── OpenRouter / Venice (inherit OpenAI path)
├── Completion handler: attachment → base64 → content parts
├── Completion handler: extraction_status check at request time
├── Capability gating (vision check before image assembly)
├── Document text injection (extracted_text → text content part)
└── Usage logging for multimodal requests
Phase 4: Frontend
├── Attachment button (📎) in input area
├── File picker with type filtering (capability-aware)
├── Smart paste handler:
│ ├── Binary detection (images, files → auto-attach)
│ ├── Large text detection (> threshold → auto-attach as {uuid}.txt)
│ ├── Small text → normal textarea paste
│ └── UUID filename generation
├── Drag-and-drop on chat area
├── Staged attachment preview strip with status indicators:
│ ├── Upload progress bar
│ ├── Queue position ("Queued (2 ahead)")
│ ├── Extraction spinner
│ ├── Ready checkmark / failed warning
│ └── Send blocked only during active upload
├── Extraction status polling (2s interval, stops at terminal state)
├── Attachment rendering in messages (thumbnails + doc chips)
├── Image lightbox / side panel expand
├── Capability-aware UI (hide/filter by vision support)
├── Storage-not-configured state (hide upload UI)
└── Admin Storage panel (status, orphan count, cleanup button)
Phase 5: Stragglers
├── vault rekey CLI command
├── Admin encryption status indicator
└── Per-chat model/preset persistence (server-side)
Phase 6: Admin Cleanup
├── POST /admin/storage/cleanup endpoint
├── Orphan detection query (message_id IS NULL, age > 24h)
├── Admin Storage panel: orphan count + reclaimable space card
└── "Run Cleanup Now" button
Testing
- Integration tests: upload → send message with attachment → verify multimodal completion request → verify download auth
- Access control tests: upload to channel A, attempt download from user who owns channel B → 403
- Storage backend tests: put/get/delete/deletePrefix round-trip
- Extraction tests: PDF/DOCX/TXT → extracted_text populated
- Frontend tests: staged attachments state, capability gating
Migration Summary
| Migration | Tables/Columns |
|---|---|
| 007_attachments.sql | attachments table with indexes |
No changes to existing tables. The messages.metadata JSONB carries
attachment_ids without schema changes. The channels.settings JSONB
carries last_selector_id without schema changes.
Decisions Made
-
Text extraction — Tiered: native Go (pdfcpu Apache 2.0, go-docx MIT) for common cases, LibreOffice headless for everything else including ODF formats. No AGPL dependencies.
-
Thumbnails — Yes. Generated async post-upload. Images via Go stdlib (
golang.org/x/image/draw). Documents via LibreOffice first-page render. Stored alongside originals with_thumbsuffix. -
Max attachments per message — 5 files (admin-configurable).
-
Extraction timing — Async with filesystem-backed queue. Upload returns immediately, extraction runs in background via worker loop with concurrency cap. Frontend polls for status. Send blocked only during active upload (bytes in flight), not during extraction.
-
Smart paste — Automatic, no prompts. Binary clipboard items (screenshots, images) always auto-attach. Text exceeding threshold (default 2000 chars) auto-attaches as
{uuid}.txt. Below threshold pastes normally into textarea. UUID filenames, MIME detection from clipboard data. Zero user decisions.
Global Settings (New)
| Key | Default | Description |
|---|---|---|
storage_max_file_size |
10485760 (10MB) |
Per-file upload limit |
storage_max_upload_size |
52428800 (50MB) |
Per-request total limit |
storage_allowed_types |
(see Allowed Types) | MIME type allowlist |
storage_max_attachments_per_message |
5 |
Max files per message |
storage_paste_to_file_chars |
2000 |
Paste-to-file threshold (0=disabled) |
storage_user_quota_bytes |
0 |
Per-user quota (0=unlimited, future enforcement) |
Decisions Made (Continued)
-
Sidecar base image — Debian slim. Rock-solid LibreOffice packages, ~80MB larger than Alpine but no flaky package issues.
-
Queue persistence — Dual: filesystem
status.jsonis source of truth (crash recovery,lsdebuggability), mirrored to PG metadata for API serving. Worth the bookkeeping for operational transparency.
Open Questions
(None remaining — all design decisions resolved.)
Orphan Cleanup (Admin Action)
Attachments uploaded but never associated with a message (user closes tab, browser crash, abandoned upload) accumulate as orphans.
v0.12.0: manual admin action with "Run Now" button.
POST /admin/storage/cleanup
→ Scan attachments WHERE message_id IS NULL
AND created_at < NOW() - interval '24 hours'
→ Delete PG rows + storage files + processing dirs
→ Return { deleted: 3, freed_bytes: 15728640 }
Admin Storage panel card:
┌─ Orphan Files ──────────────────────────────┐
│ 3 orphaned attachments (15 MB reclaimable) │
│ │
│ Files not linked to any message for >24hrs │
│ │
│ [Run Cleanup Now] │
└─────────────────────────────────────────────┘
The 24-hour grace period prevents cleaning up files that are simply in the middle of a long compose session.
Future (v0.22.0 task runner): This becomes a scheduled task running daily. The admin button stays as an on-demand override. The endpoint is identical — the scheduler just calls it on a cron.