This repository has been archived on 2026-04-03. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
core/docs/DESIGN-0.12.0.md
2026-02-26 21:19:55 +00:00

47 KiB

v0.12.0 — File Handling + Vision (Design)

Overview

File input into chat — table stakes for serious use. Image uploads for vision-capable models, document uploads for text extraction into context, and the storage backend that v0.14.0 (Knowledge Bases) and v0.15.0 (Compaction snapshots) will reuse.

Also picks up stragglers deferred from earlier releases.

Depends on: v0.11.0 (extension foundation — complete). Reused by: v0.14.0 (KBs), v0.15.0 (compaction), v0.22.0 (exports).


Stragglers (from v0.9.4 deferred)

Items deferred during vault work that were never scheduled:

switchboard vault rekey CLI command

Re-encrypts all global/team API keys when ENCRYPTION_KEY is rotated. Personal BYOK keys are unaffected (keyed to UEK, not env var).

Implementation: standalone CLI entrypoint that reads old + new env vars, iterates api_configs and team_providers where key_scope IN ('global', 'team'), decrypts with old key, re-encrypts with new key, updates in a transaction.

Admin UI: encryption status indicator

Badge in admin Settings showing whether ENCRYPTION_KEY is set and how many keys are encrypted vs. plaintext (migration stragglers). Single GET /admin/storage/status endpoint.

Per-chat model/preset persistence (server-side)

Currently localStorage bandaid from v0.10.2 — doesn't roam across devices.

Fix: Store last_selector_id in channels.settings JSONB on each completion success. On chat selection, resolve: channel.settings.last_selector_id → match against available models/presets → fall back to channel.model → global default.

No migration needed (channels.settings JSONB column already exists). Single PATCH /channels/:id with settings merge on completion success.


Track 1: Storage Backend

Environment Variables

Variable Default Description
STORAGE_BACKEND (auto) pvc or s3. If not set: pvc when STORAGE_PATH is writable, disabled otherwise.
STORAGE_PATH /data/storage Mount point for PVC backend. Also scratch dir for extraction with S3.
STORAGE_CLASS K8s only. Gitea CI variable → PVC manifest. Value: cephfs (RWX for multi-pod).
S3_ENDPOINT S3-compatible endpoint (e.g. http://minio:9000). Required when STORAGE_BACKEND=s3.
S3_BUCKET Bucket name. Must exist before first start.
S3_ACCESS_KEY S3 access key ID.
S3_SECRET_KEY S3 secret access key.
S3_REGION us-east-1 AWS region.
S3_PREFIX Optional key prefix for shared buckets (e.g. switchboard/).
S3_FORCE_PATH_STYLE true Path-style URLs. Required for MinIO, Ceph RGW, most self-hosted.
EXTRACTION_MODE inline inline (in-process, unified image) or sidecar (K8s, shared PVC).
EXTRACTION_CONCURRENCY 3 Max concurrent extraction jobs. Caps LibreOffice memory usage.

Auto-detection when STORAGE_BACKEND is not set:

STORAGE_PATH writable → pvc (implicit)
STORAGE_PATH missing  → file features disabled
                        admin panel shows "Storage not configured"

When STORAGE_BACKEND=pvc is explicit, fail startup if path is not writable (fail-safe, same pattern as ENCRYPTION_KEY enforcement).

S3 backend: reads S3_ENDPOINT, S3_BUCKET, S3_ACCESS_KEY, S3_SECRET_KEY, S3_REGION, S3_PREFIX, S3_FORCE_PATH_STYLE from env vars. Uses minio-go v7. Works with MinIO, Ceph RGW, AWS S3. Same interface as PVC — all handlers are backend-agnostic. PVC still required as scratch dir for extraction queue.

Config Addition

// In config.go
StorageBackend         string // "pvc", "s3", or "" (auto-detect)
StoragePath            string // mount point for PVC backend
ExtractionMode         string // "inline" or "sidecar"
ExtractionConcurrency  int    // max concurrent extraction jobs

Interface

// server/storage/storage.go
package storage

import (
    "context"
    "io"
)

// ObjectStore is the abstraction for blob storage.
// Implementations: PVC (filesystem), S3 (minio-go v7).
type ObjectStore interface {
    // Put writes data to the given key. Creates parent dirs as needed.
    Put(ctx context.Context, key string, r io.Reader, size int64, contentType string) error

    // Get returns a reader for the given key. Caller must close.
    // Returns ErrNotFound if key does not exist.
    Get(ctx context.Context, key string) (io.ReadCloser, int64, string, error)

    // Delete removes the object at key. No error if already gone.
    Delete(ctx context.Context, key string) error

    // DeletePrefix removes all objects under the given prefix.
    // Used for channel deletion (bulk cleanup).
    DeletePrefix(ctx context.Context, prefix string) error

    // Exists checks if an object exists at key without reading it.
    Exists(ctx context.Context, key string) (bool, error)

    // Healthy returns nil if the backend is operational.
    Healthy(ctx context.Context) error
}

var ErrNotFound = errors.New("storage: object not found")

PVC Implementation

server/storage/pvc.go — ~100 lines. Thin wrapper around os.*:

type PVCStore struct {
    basePath string // e.g. "/data/storage"
}

func NewPVC(basePath string) (*PVCStore, error) {
    // Validate basePath is writable (create test file, remove)
    // Create top-level subdirs: attachments/
}

Key → filesystem path: filepath.Join(basePath, key). The key itself provides all the directory structure.

Filesystem Layout

/data/storage/                    ← STORAGE_PATH mount point
└── attachments/
    └── {channel_id}/
        └── {attachment_id}_{filename}

Channel-first because:

  • Channel deletion → DeletePrefix("attachments/{channel_id}/") — one call.
  • User deletion cascades to owned channels → same cleanup path.
  • Group chats / multi-participant channels: all files in one place regardless of who uploaded them.
  • Quota enforcement uses PG (SUM(size_bytes) WHERE user_id = $1), not filesystem walks.

Future subdirs at the STORAGE_PATH level (not this release):

  • knowledge-bases/ — v0.14.0
  • compaction/ — v0.15.0
  • exports/ — TBD

K8s Manifest (PVC)

# k8s/storage-pvc.yaml
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: switchboard-storage
  namespace: switchboard
spec:
  accessModes: [ReadWriteMany]     # CephFS RWX — multi-pod safe
  storageClassName: "${STORAGE_CLASS}"
  resources:
    requests:
      storage: 10Gi                # configurable per deployment

Backend deployment addition:

# In k8s/backend.yaml
volumes:
  - name: storage
    persistentVolumeClaim:
      claimName: switchboard-storage
containers:
  - name: backend
    volumeMounts:
      - name: storage
        mountPath: /data/storage
    env:
      - name: STORAGE_BACKEND
        value: "pvc"
      - name: STORAGE_PATH
        value: "/data/storage"

Docker Compose

# In docker-compose.yml
services:
  switchboard:
    volumes:
      - ./data/storage:/data/storage
    environment:
      - STORAGE_PATH=/data/storage

Admin Status Endpoint

GET /admin/storage/status

{
  "backend": "pvc",
  "path": "/data/storage",
  "healthy": true,
  "total_files": 142,
  "total_bytes": 52428800,
  "configured": true
}

When storage is not configured, file upload endpoints return 503 with {"error": "File storage not configured"}. Frontend hides upload UI based on a capabilities flag in the boot payload.


Track 2: Attachments Schema + API

Migration 007_attachments.sql

-- ── Attachments ────────────────────────────
CREATE TABLE IF NOT EXISTS attachments (
    id              UUID PRIMARY KEY DEFAULT gen_random_uuid(),
    channel_id      UUID NOT NULL REFERENCES channels(id) ON DELETE CASCADE,
    user_id         UUID NOT NULL REFERENCES users(id),
    message_id      UUID REFERENCES messages(id) ON DELETE SET NULL,
    filename        VARCHAR(255) NOT NULL,
    content_type    VARCHAR(127) NOT NULL,
    size_bytes      BIGINT NOT NULL,
    storage_key     TEXT NOT NULL,
    extracted_text  TEXT,
    metadata        JSONB DEFAULT '{}'::jsonb,
    created_at      TIMESTAMPTZ DEFAULT NOW()
);

-- Access pattern: always join through channel for auth
CREATE INDEX idx_attachments_channel ON attachments(channel_id);
-- Quota pattern: sum by user
CREATE INDEX idx_attachments_user_size ON attachments(user_id);
-- Message association: find attachments for a message
CREATE INDEX idx_attachments_message ON attachments(message_id);

Notes on columns:

  • message_id is nullable — attachment is uploaded before the message is persisted. Updated after message creation. If message is deleted, attachment persists (belongs to channel, not message).
  • storage_key is the relative path within STORAGE_PATH, e.g. attachments/{channel_id}/{attachment_id}_{filename}.
  • extracted_text is populated asynchronously for PDFs/DOCX. NULL for images. Used for context injection into completions.
  • metadata JSONB holds dimensions (images), page count (PDFs), extraction status, thumbnails, etc.

Store Layer

// server/store/attachments.go
type AttachmentStore interface {
    Create(ctx context.Context, a *models.Attachment) error
    GetByID(ctx context.Context, id string) (*models.Attachment, error)
    GetByChannel(ctx context.Context, channelID string) ([]models.Attachment, error)
    GetByMessage(ctx context.Context, messageID string) ([]models.Attachment, error)
    SetMessageID(ctx context.Context, id, messageID string) error
    Delete(ctx context.Context, id string) error
    DeleteByChannel(ctx context.Context, channelID string) ([]string, error)  // returns storage_keys for cleanup
    UserUsageBytes(ctx context.Context, userID string) (int64, error)
}

API Endpoints

POST   /api/v1/channels/:id/attachments     Upload file
GET    /api/v1/attachments/:id               Attachment metadata
GET    /api/v1/attachments/:id/download      Download file content
DELETE /api/v1/attachments/:id               Delete attachment
GET    /api/v1/channels/:id/attachments      List channel attachments

Upload Flow

Client                      Backend                     Storage
  │                           │                           │
  ├── POST multipart ────────►│                           │
  │   (file + channel_id)     │                           │
  │                           ├── Auth check              │
  │                           ├── Channel membership ─────│
  │                           │   verify                  │
  │                           ├── Size limit check        │
  │                           ├── MIME type validation    │
  │                           ├── Generate attachment_id  │
  │                           ├── PUT ───────────────────►│
  │                           │   key: att/{ch}/{id}_{fn} │
  │                           ├── INSERT attachments row  │
  │                           │                           │
  │◄── 201 { attachment } ────┤                           │

The attachment is created before the message is sent. The frontend holds the attachment ID(s), includes them in the completion request, and the backend links them to the persisted message after creation.

Download Flow (Access Control)

GET /api/v1/attachments/:id/download

  → SELECT a.*, c.id, c.user_id
    FROM attachments a
    JOIN channels c ON a.channel_id = c.id
    WHERE a.id = $1

  → Verify: requester is channel owner
    (future RBAC: OR requester is channel participant
     OR requester is team member with channel access
     OR requester is admin with audit log)

  → Stream file from ObjectStore with:
    Content-Type: {content_type}
    Content-Disposition: attachment; filename="{filename}"
    Content-Length: {size_bytes}

Critical: Files are NEVER served directly from nginx or any static file path. Every byte goes through the Go backend with auth checks. No pre-signed URLs, no direct filesystem exposure.

Size & Type Limits

Configured via global settings (see Global Settings section below). Enforced server-side on upload — frontend mirrors limits for UX but the backend is authoritative.

Validation order on upload:

  1. storage_allowed_types — reject disallowed MIME types
  2. storage_max_file_size — reject oversized individual files
  3. storage_max_upload_size — reject if request total exceeds limit
  4. storage_max_attachments_per_message — reject if too many files
  5. storage_user_quota_bytes — reject if user quota exceeded (future)

MIME type is detected server-side (http.DetectContentType + extension validation), not trusted from the client Content-Type header.


Track 3: Multimodal Message Assembly

Provider Message Polymorphism

Currently providers.Message.Content is a plain string. Vision requires content arrays with image parts. The change is additive — no existing callsites break.

// In server/providers/provider.go

// ContentPart represents one part of a multimodal message.
type ContentPart struct {
    Type     string `json:"type"`               // "text", "image_url"
    Text     string `json:"text,omitempty"`      // for type="text"
    ImageURL *ImageURL `json:"image_url,omitempty"` // for type="image_url"
}

type ImageURL struct {
    URL    string `json:"url"`              // "data:image/jpeg;base64,..." or URL
    Detail string `json:"detail,omitempty"` // "auto", "low", "high"
}

type Message struct {
    Role    string `json:"role"`
    Content string `json:"content"`

    // Multimodal content. When non-nil, providers use this instead of Content.
    ContentParts []ContentPart `json:"content_parts,omitempty"`

    // ... existing tool fields unchanged
}

Resolution rule: if ContentParts is non-nil, providers build their wire format from it. If nil, they use Content (string) as today. This means every existing codepath works unchanged until an attachment is involved.

Provider Wire Format Mapping

OpenAI / OpenRouter / Venice (OpenAI-compatible):

// In openai.go — message conversion
if len(m.ContentParts) > 0 {
    // Content becomes array of objects
    oaiMsg.Content = nil  // clear string content
    oaiMsg.ContentArray = toOpenAIContentParts(m.ContentParts)
}

OpenAI wire format:

{
  "role": "user",
  "content": [
    {"type": "text", "text": "What's in this image?"},
    {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
  ]
}

The openaiMessage struct needs a ContentArray field with custom JSON marshaling — when ContentArray is set, marshal content as the array; otherwise marshal as string. This is the standard OpenAI multimodal format.

Anthropic:

// In anthropic.go — message conversion
if len(m.ContentParts) > 0 {
    antMsg.Content = toAnthropicContentBlocks(m.ContentParts)
}

Anthropic wire format:

{
  "role": "user",
  "content": [
    {"type": "text", "text": "What's in this image?"},
    {"type": "image", "source": {"type": "base64", "media_type": "image/jpeg", "data": "..."}}
  ]
}

The anthropicContentBlock struct gets new fields:

type anthropicContentBlock struct {
    // ... existing fields ...
    // Image support
    Source *anthropicImageSource `json:"source,omitempty"`
}

type anthropicImageSource struct {
    Type      string `json:"type"`       // "base64"
    MediaType string `json:"media_type"` // "image/jpeg", "image/png", etc.
    Data      string `json:"data"`       // base64-encoded image
}

Image Delivery: Base64 Inline

For v0.12.0, all images are sent as base64-encoded data inline in the API request. No pre-signed URLs.

Rationale: The target audience is government/enterprise, often in air-gapped or restricted networks. Pre-signed URLs require the provider's API servers to reach the storage endpoint — impossible in disconnected environments. Base64 works everywhere.

Trade-off: Larger request payloads. Acceptable for v0.12.0 — images are typically <10MB each and context windows are large. URL-based delivery can be an optimization in a future release for cloud deployments.

Completion Handler Changes

// In completion.go — after loading conversation, before building provReq

// If message has attachments, build multimodal content
if len(req.AttachmentIDs) > 0 {
    parts := []providers.ContentPart{
        {Type: "text", Text: req.Content},
    }
    for _, attID := range req.AttachmentIDs {
        att, err := h.store.Attachments().GetByID(c, attID)
        // ... verify att.channel_id matches ...

        if isImageType(att.ContentType) {
            // Verify vision capability
            if !caps.Vision {
                c.JSON(400, gin.H{"error": "Selected model does not support image input"})
                return
            }

            // Read from storage, base64 encode
            reader, _, _, _ := h.objectStore.Get(c, att.StorageKey)
            data, _ := io.ReadAll(reader)
            reader.Close()
            b64 := base64.StdEncoding.EncodeToString(data)
            dataURL := fmt.Sprintf("data:%s;base64,%s", att.ContentType, b64)

            parts = append(parts, providers.ContentPart{
                Type:     "image_url",
                ImageURL: &providers.ImageURL{URL: dataURL, Detail: "auto"},
            })
        } else if att.ExtractedText != nil && *att.ExtractedText != "" {
            // Document: inject extracted text as context
            parts = append(parts, providers.ContentPart{
                Type: "text",
                Text: fmt.Sprintf("[Document: %s]\n%s", att.Filename, *att.ExtractedText),
            })
        } else {
            // Extraction not yet complete or failed — include filename only
            status := att.Metadata["extraction_status"]
            parts = append(parts, providers.ContentPart{
                Type: "text",
                Text: fmt.Sprintf("[Attached file: %s (extraction %s)]", att.Filename, status),
            })
        }
    }

    // Use ContentParts instead of Content string
    userMsg := providers.Message{
        Role:         "user",
        ContentParts: parts,
    }
    messages = append(messages, userMsg)
} else {
    // Existing text-only path — unchanged
    messages = append(messages, providers.Message{
        Role:    "user",
        Content: req.Content,
    })
}

Frontend blocks send only during active upload — the send button stays disabled while any staged attachment is in uploading state (bytes still in flight). Once the file is on disk and the attachment row exists in PG, send is enabled regardless of extraction status.

The completion handler checks extraction_status at request time:

  • complete → inject extracted text as context
  • anything else → inject filename-only placeholder

This avoids blocking users on slow extractions while still providing extracted context when available. For images (no extraction), the attachment is immediately usable.

Capability Gating

Images are only assembled into content parts when the resolved model has capabilities.vision == true. If a user attaches an image to a non-vision model:

  • Documents (PDF, DOCX, TXT): always allowed — extracted text is injected as a text content part. No vision needed.
  • Images to non-vision model: backend returns 400 with {"error": "Selected model does not support image input"}.
  • Frontend disables image upload button when selected model lacks vision.

Text Extraction (Documents)

License Constraint

The project is Apache 2.0. Libraries with AGPL/GPL copyleft (e.g. unipdf) are incompatible — their license would infect the codebase. Pure Go PDF libraries under permissive licenses (ledongthuc/pdf, pdfcpu) handle simple cases but choke on complex layouts, scanned PDFs, and encrypted documents.

Approach: Tiered Extraction

Tier 1 — Native Go (no external deps, handles ~80% of uploads):

Format Library License
TXT, Markdown, CSV io.ReadAll
PDF (simple) pdfcpu Apache 2.0
DOCX fumiama/go-docx MIT

Tier 2 — LibreOffice Headless (handles everything else):

For PDF with complex layouts, ODT, ODS, ODP, RTF, DOC (legacy), XLS, and any format Tier 1 can't handle — shell out to LibreOffice headless:

libreoffice --headless --convert-to txt:Text /tmp/input.odt --outdir /tmp/

LibreOffice handles the full Office family:

  • OpenDocument: .odt, .ods, .odp
  • Microsoft: .doc, .docx, .xls, .xlsx, .ppt, .pptx
  • PDF (including scanned with basic OCR via hunspell)
  • RTF, HTML, and dozens more

Extraction Queue (Filesystem-Backed)

Extraction state lives on the PVC, not in memory. Pod restarts don't lose queue state. The queue is a directory convention:

/data/storage/
├── attachments/              ← completed uploads (served to clients)
└── processing/               ← extraction queue (internal only)
    └── {attachment_id}/
        ├── input             ← original file (copy or hardlink)
        ├── status.json       ← queue state + metadata
        ├── output.txt        ← extracted text (when complete)
        └── thumb.jpg         ← thumbnail (when complete)

status.json:

{
  "state": "pending",
  "queued_at": "2026-02-25T14:30:00Z",
  "started_at": null,
  "completed_at": null,
  "content_type": "application/pdf",
  "filename": "report.pdf",
  "error": null
}

States: pendingprocessingcomplete | failed

Worker loop (runs in both inline and sidecar mode):

1. Scan processing/ for dirs with status.state == "pending"
   (ordered by queued_at — FIFO)
2. Respect concurrency cap (semaphore, default 3)
3. Pick job → flip state to "processing"
4. Tier 1: try native Go extraction
   ├── success → write output.txt, generate thumb, flip "complete"
   └── fail or unsupported →
5. Tier 2: LibreOffice headless
   ├── success → write output.txt, generate thumb, flip "complete"
   └── fail → flip "failed", write error
6. Copy results to PG:
   ├── UPDATE attachments SET extracted_text = ...,
   │   metadata = jsonb_set(metadata, '{extraction_status}', '"complete"')
   └── Thumbnail storage_key written to metadata
7. Clean up processing/{id}/ dir

In inline mode (Unified image): the worker is a goroutine pool in the Go backend, scanning the queue directory on a 2-second tick.

In sidecar mode (K8s): the sidecar runs the same loop as a standalone process. When extraction completes, it writes output.txt and thumb.jpg to the processing dir and flips status to complete. The Go backend polls the status file and copies results to PG.

Crash recovery: On startup, scan for state == "processing" (stale from a crash). Reset to pending for re-processing. This is why the queue is on the filesystem — nothing is lost.

Concurrency: Capped by semaphore. Default 3 concurrent extractions. Configurable via EXTRACTION_CONCURRENCY env var. LibreOffice is memory-hungry on complex docs — the cap prevents OOM on burst uploads.

Frontend queue awareness:

Attachment status:
  uploading       → progress bar
  pending         → "Queued (2 ahead)"     ← position from queue scan
  processing      → "Extracting..."        ← spinner
  complete        → ✓ (ready)
  failed          → ⚠ (send still enabled, model sees filename only)

The frontend polls GET /api/v1/attachments/:id every 2s for non-terminal states. The response includes extraction_status and queue_position (0 = currently processing, N = N jobs ahead).

Send is blocked while any attachment is uploading. Once all are at least pending, send is enabled — the extraction results will be available by the time the model processes the request (optimistic), or the model sees the filename-only fallback (acceptable degradation).

Correction to earlier design: Send is blocked only during uploading phase. Once the file is on disk and queued, the user can send. The completion handler checks extraction status at request time and uses whatever is available.

Extraction Metadata in PG

Attachment metadata JSONB tracks status (mirrored from queue):

{
  "extraction_status": "complete",
  "extraction_error": null,
  "queue_position": 0,
  "page_count": 12,
  "dimensions": {"width": 1920, "height": 1080},
  "thumbnail_key": "attachments/{channel_id}/{id}_thumb.jpg",
  "thumb_dimensions": {"width": 280, "height": 147}
}

States: pendingprocessingcomplete | failed | unavailable

LibreOffice Deployment Strategy

Two deployment modes, matching existing Docker architecture:

Deployment LibreOffice location Reason
Unified (docker-compose, dev) Baked into image Single container, simplicity wins
K8s (backend + frontend split) Sidecar container Separate scaling, lean backend image, independent version pinning

Unified Dockerfile addition:

# In Dockerfile (unified)
RUN apt-get update && apt-get install -y --no-install-recommends \
    libreoffice-core libreoffice-writer libreoffice-calc \
    && rm -rf /var/lib/apt/lists/*

K8s sidecar (backend.yaml addition):

# In k8s/backend.yaml
containers:
  - name: backend
    image: switchboard-backend:0.12.0
    volumeMounts:
      - name: storage
        mountPath: /data/storage
    env:
      - name: STORAGE_BACKEND
        value: "pvc"
      - name: STORAGE_PATH
        value: "/data/storage"
      - name: EXTRACTION_MODE
        value: "sidecar"    # "inline" for unified, "sidecar" for k8s

  - name: extractor
    image: switchboard-extractor:0.12.0   # minimal image: alpine + libreoffice
    volumeMounts:
      - name: storage
        mountPath: /data/storage
    resources:
      requests:
        memory: "256Mi"
      limits:
        memory: "1Gi"       # LibreOffice can spike on complex docs

volumes:
  - name: storage
    persistentVolumeClaim:
      claimName: switchboard-storage

The sidecar image is small and purpose-built: Alpine + LibreOffice + a watcher script. No Go code, no API surface. Upgrades to LibreOffice are a tag bump on the sidecar image — backend image untouched.

Communication: Filesystem-only via the shared PVC. No gRPC, no HTTP between containers. See Extraction Queue below.

Allowed Types (Updated)

{
  "storage_allowed_types": [
    "image/jpeg", "image/png", "image/gif", "image/webp", "image/svg+xml",
    "application/pdf",
    "text/plain", "text/markdown", "text/csv",
    "application/vnd.openxmlformats-officedocument.wordprocessingml.document",
    "application/vnd.openxmlformats-officedocument.spreadsheetml.sheet",
    "application/vnd.openxmlformats-officedocument.presentationml.presentation",
    "application/vnd.oasis.opendocument.text",
    "application/vnd.oasis.opendocument.spreadsheet",
    "application/vnd.oasis.opendocument.presentation",
    "application/rtf",
    "application/msword",
    "application/vnd.ms-excel"
  ]
}

Thumbnail Generation

Thumbnails are generated asynchronously alongside text extraction as part of the same post-upload processing pipeline.

Image Thumbnails

Use golang.org/x/image/draw (BSD, stdlib-adjacent) for resizing:

// Resize to max 280px on longest edge, preserve aspect ratio
// Output: JPEG at 80% quality
// Stored at: attachments/{channel_id}/{attachment_id}_thumb.jpg

No external dependency. Pure Go. Fast enough for on-upload generation.

Document Thumbnails

When LibreOffice is available, generate a first-page preview:

# Convert first page to PDF, then to image
libreoffice --headless --convert-to pdf /tmp/input.docx --outdir /tmp/
# Then use Go image libraries to render first page
# Or: convert to PNG directly via LibreOffice draw filter
libreoffice --headless --convert-to png /tmp/input.docx --outdir /tmp/

For PDFs specifically, pdfcpu can extract page images under Apache 2.0.

When LibreOffice is not available, documents get a generic file-type icon (handled in CSS, no thumbnail row needed).

Storage

Thumbnails stored alongside originals with _thumb suffix:

/data/storage/attachments/{channel_id}/
├── {id}_report.pdf            ← original
├── {id}_report.pdf_thumb.jpg  ← thumbnail
├── {id}_photo.jpg             ← original
└── {id}_photo.jpg_thumb.jpg   ← thumbnail

Thumbnail URL returned in attachment metadata:

{
  "thumbnail_key": "attachments/{channel_id}/{id}_photo.jpg_thumb.jpg",
  "dimensions": {"width": 1920, "height": 1080},
  "thumb_dimensions": {"width": 280, "height": 147}
}

Download endpoint: GET /api/v1/attachments/:id/thumbnail — same channel-scoped auth check as the full download.


Track 4: Frontend

Input Area Changes

The <div class="input-wrap"> gets an attachment button and a preview strip:

┌──────────────────────────────────────────────┐
│ ┌──────────┐ ┌──────────┐                    │  ← attachment preview strip
│ │ image.png│ │ report.pdf│                    │    (visible only when files
│ │    ✕     │ │    ✕     │                    │     are staged)
│ └──────────┘ └──────────┘                    │
├──────────────────────────────────────────────┤
│ 📎 │ Send a message...              │ ■ │ ▶ │  ← existing input + new clip btn
└──────────────────────────────────────────────┘

Upload Interactions

Trigger Behavior
📎 button click File picker dialog (filtered by allowed types)
Paste (Ctrl+V) binary Auto-attach: screenshot/image → image attachment ({uuid}.{ext})
Paste (Ctrl+V) large text Auto-attach: text > threshold → text file ({uuid}.txt)
Paste (Ctrl+V) small text Normal paste into textarea (below threshold)
Drag-and-drop on chat area Auto-attach file(s) as attachments (max 5 per message)

Smart Paste (Auto-Attach)

Paste events are intercepted and routed automatically — no prompts, no user decisions. The clipboard content determines the behavior:

Detection logic:

Clipboard paste event
  │
  ├── Contains files/images (binary)?
  │   └── YES → auto-attach each item
  │       ├── Image (screenshot, copied image) → upload as image attachment
  │       ├── File (dragged from OS) → upload as file attachment
  │       └── Filename: {uuid}.{ext} (ext from MIME type)
  │
  ├── Contains text?
  │   ├── Length ≤ threshold → normal paste into textarea
  │   └── Length > threshold → auto-attach as text file
  │       ├── Upload as attachment (content_type: text/plain)
  │       ├── Filename: {uuid}.txt
  │       └── Textarea stays empty (or keeps existing draft)
  │
  └── Empty → no-op

Binary detection: The paste event's clipboardData.items array carries type information. Items with type.startsWith('image/') or any non-text type are binary — always auto-attached. Text items (text/plain, text/html) are checked against the threshold.

UUID filenames: All auto-attached pastes get UUID-based filenames. No timestamps, no pasted-text.txt collision issues. The original content type and paste source are captured in attachment metadata:

{
  "source": "clipboard",
  "original_type": "image/png",
  "paste_length": 4200
}

Threshold: storage_paste_to_file_chars global setting (default: 2000 characters, admin-configurable). Setting to 0 disables text auto-attach (binary pastes still auto-attach regardless).

UX flow for a screenshot paste:

  1. User hits Ctrl+V with a screenshot on clipboard
  2. Image appears instantly in the attachment preview strip (uploading...)
  3. Upload completes → thumbnail shown with checkmark
  4. User types their question, hits send
  5. Image goes to vision model as base64

UX flow for a large log paste:

  1. User copies 5,000 chars of server logs, hits Ctrl+V
  2. Text file chip appears in attachment preview strip: 📄 {uuid}.txt (4.9 KB)
  3. Extraction runs (trivial for plain text — near-instant)
  4. User types "what's wrong with these logs?", hits send
  5. Log content injected as extracted text context

No prompts, no dialogs, no decisions. Paste just works.

State Management

Staged attachments live in App.stagedAttachments[] — an array of {id, filename, contentType, sizeBytes, previewUrl, extractionStatus} populated after the upload API returns. On send, attachment IDs are included in the completion request body. On cancel/clear, staged attachments are deleted via the API.

Upload lifecycle per attachment:

uploading → uploaded/queued → extracting → ready | failed
  • uploading: progress bar visible, send disabled.
  • uploaded/queued: file on disk, extraction queued. Queue position shown: "Queued (2 ahead)". Send enabled — extraction will complete before or during model processing, or the model gets the filename-only fallback.
  • extracting: spinner on thumbnail. Send remains enabled.
  • ready: extraction complete (or image — no extraction needed). Green checkmark.
  • failed: extraction failed. Warning icon on thumbnail. Send still enabled — model sees [Attached file: X (extraction failed)].

Key principle: Send is blocked only during active upload (bytes in flight). Once the file is on disk and the attachment row exists, the user can send. This prevents the frustrating case where a user types their message while a 50-page PDF extracts for 30 seconds.

Polling: Frontend polls GET /api/v1/attachments/:id every 2s for any attachment with extractionStatus == "pending" | "processing". Stops when all settle. WebSocket notification is a future optimization. Polling updates the queue position indicator and thumbnail status icon.

Message Rendering

Attachments on sent/received messages render inline:

  • Images: Thumbnail (from _thumb.jpg) with click-to-expand (lightbox or side panel). Rendered in the message bubble below/above the text.
  • Documents: File chip with icon, filename, size, and download link. 📄 report.pdf (2.3 MB) ⬇
  • Pasted text files: Same as documents — 📄 {uuid}.txt (4.9 KB) ⬇. Metadata source: "clipboard" available for future UX refinement.

Attachment data comes from messages.metadata.attachment_ids → batch-fetched on conversation load via GET /channels/:id/attachments.

Capability-Aware UI

The frontend already knows model capabilities from the model selector.

Images are always auto-attached regardless of current model — the user might switch models before sending. When the selected model has vision: false:

  • Image attachments show a subtle indicator: 👁️‍🗨️ icon dimmed with tooltip "Current model doesn't support images"
  • File picker filters to documents only, but paste/drop still accepts images (user can switch models)
  • The 📎 button tooltip says "Attach document" instead of "Attach file"
  • On send with image + non-vision model: backend returns 400 with "Selected model does not support image input — switch to a vision model or remove image attachments"
  • Document uploads (PDF, DOCX, ODT, etc.) always work — text extraction doesn't require vision

When storage is not configured (boot payload flag), the 📎 button is hidden entirely and all paste/drop auto-attach handlers are bypassed.


Privacy & Access Control Model

Principle

Attachments have no access identity of their own. They inherit access from their parent channel. The channel membership check is the single gate for all attachment operations.

Access Matrix (v0.12.0 — Current Auth Model)

Channel type Who can upload Who can download Delete
Personal DM Channel owner Channel owner Owner
(Future) Team channel Team members with access Team members with access Uploader or team admin
(Future) Group chat Channel participants Channel participants Uploader or channel owner
(Future) Workflow Assigned participants Assigned participants Stage rules

Enforcement Points

Every attachment endpoint performs channel-scoped authorization:

// Pseudocode — every handler follows this pattern
func (h *Handler) downloadAttachment(c *gin.Context) {
    att := h.store.Attachments().GetByID(c, attID)
    channel := h.store.Channels().GetByID(c, att.ChannelID)

    // THE GATE: is the requester authorized for this channel?
    if channel.UserID != requesterID {
        c.JSON(403, gin.H{"error": "access denied"})
        return
    }

    // Future RBAC (v0.20.0): replace the above with:
    // if !h.rbac.CanAccess(requesterID, channel.ID, "attachment:read") {
    //     c.JSON(403, ...)
    // }

    h.objectStore.Get(c, att.StorageKey) // → stream to client
}

Threat Mitigations

Threat Mitigation
ID guessing / enumeration Every download joins through channel + verifies requester membership. UUIDs are non-sequential.
Direct filesystem access Files never served by nginx. All bytes go through Go backend with auth. No static file routes for storage path.
Cross-team data leak Channel membership is the boundary. HR channel attachments are only accessible to HR channel participants.
Admin overreach v0.12.0: admin can access (consistent with existing admin trust model). Audit-logged. v0.20.0: RBAC with team-scoped admin.
Metadata leak Channel attachment list endpoint also requires channel membership. No global search across attachments.
Storage key prediction Keys include UUID attachment_id — not guessable from channel_id + filename alone.
Upload as attack vector MIME type validated against allowlist. File size enforced. Content-Type set from server-side detection, not client header.

Consistency with Existing Scoped Resources

The attachment access pattern is identical to the pattern used for all scoped data in the system:

Resource          Scope          Access Gate
──────────────────────────────────────────────────
Notes (personal)  user_id        user_id = requester
Notes (team)      team_id        requester ∈ team members
Attachments       channel_id     requester ∈ channel (owner/participant)
Usage data        user/team      scoped queries
API keys          user/org       vault (UEK) / env-derived key

The row knows its scope. The handler enforces the boundary. The query filters by it.

Future RBAC Hooks (v0.20.0)

The v0.12.0 access checks are written as explicit if checks that can be trivially replaced by a centralized RBAC call:

// v0.12.0
if channel.UserID != requesterID { deny }

// v0.20.0
if !rbac.Can(requesterID, channel.ID, permission) { deny }

The schema, storage layer, and API surface don't change. Only the authorization check function gets swapped.


Implementation Order

Phase 1: Storage Backend
  ├── Config additions (STORAGE_BACKEND, STORAGE_PATH, EXTRACTION_MODE)
  ├── ObjectStore interface
  ├── PVC implementation (~100 lines)
  ├── Startup validation (writable check, health)
  ├── K8s manifest (PVC + volume mount + STORAGE_CLASS)
  ├── Docker-compose volume mount
  ├── Admin status endpoint
  └── Boot payload: storage_configured flag

Phase 2: Attachments CRUD + Extraction Pipeline
  ├── Migration 007_attachments.sql
  ├── models.Attachment struct
  ├── AttachmentStore (PG CRUD)
  ├── handlers/attachments.go (upload, download, delete, list, thumbnail)
  ├── Channel-scoped access checks on all endpoints
  ├── MIME validation + size enforcement (server-side detection)
  ├── Filesystem cleanup on channel delete (CASCADE + DeletePrefix)
  ├── Extraction queue (filesystem-backed):
  │   ├── processing/ directory convention + status.json
  │   ├── Worker loop with semaphore (default 3 concurrent)
  │   ├── Tier 1: native Go (pdfcpu, go-docx, io.ReadAll)
  │   ├── Tier 2: LibreOffice headless (inline or sidecar)
  │   ├── Crash recovery (reset stale "processing" → "pending")
  │   └── PG metadata mirror (extraction_status, queue_position)
  ├── Thumbnail generation (images: Go resize, docs: LibreOffice)
  ├── Unified Dockerfile: add libreoffice-core
  ├── Sidecar Dockerfile: alpine + libreoffice + watcher
  └── K8s backend.yaml: sidecar container definition

Phase 3: Multimodal Message Assembly
  ├── ContentPart + ImageURL types in providers/provider.go
  ├── OpenAI message conversion (content array marshaling)
  ├── Anthropic message conversion (image source blocks)
  ├── OpenRouter / Venice (inherit OpenAI path)
  ├── Completion handler: attachment → base64 → content parts
  ├── Completion handler: extraction_status check at request time
  ├── Capability gating (vision check before image assembly)
  ├── Document text injection (extracted_text → text content part)
  └── Usage logging for multimodal requests

Phase 4: Frontend
  ├── Attachment button (📎) in input area
  ├── File picker with type filtering (capability-aware)
  ├── Smart paste handler:
  │   ├── Binary detection (images, files → auto-attach)
  │   ├── Large text detection (> threshold → auto-attach as {uuid}.txt)
  │   ├── Small text → normal textarea paste
  │   └── UUID filename generation
  ├── Drag-and-drop on chat area
  ├── Staged attachment preview strip with status indicators:
  │   ├── Upload progress bar
  │   ├── Queue position ("Queued (2 ahead)")
  │   ├── Extraction spinner
  │   ├── Ready checkmark / failed warning
  │   └── Send blocked only during active upload
  ├── Extraction status polling (2s interval, stops at terminal state)
  ├── Attachment rendering in messages (thumbnails + doc chips)
  ├── Image lightbox / side panel expand
  ├── Capability-aware UI (hide/filter by vision support)
  ├── Storage-not-configured state (hide upload UI)
  └── Admin Storage panel (status, orphan count, cleanup button)

Phase 5: Stragglers
  ├── vault rekey CLI command
  ├── Admin encryption status indicator
  └── Per-chat model/preset persistence (server-side)

Phase 6: Admin Cleanup
  ├── POST /admin/storage/cleanup endpoint
  ├── Orphan detection query (message_id IS NULL, age > 24h)
  ├── Admin Storage panel: orphan count + reclaimable space card
  └── "Run Cleanup Now" button

Testing

  • Integration tests: upload → send message with attachment → verify multimodal completion request → verify download auth
  • Access control tests: upload to channel A, attempt download from user who owns channel B → 403
  • Storage backend tests: put/get/delete/deletePrefix round-trip
  • Extraction tests: PDF/DOCX/TXT → extracted_text populated
  • Frontend tests: staged attachments state, capability gating

Migration Summary

Migration Tables/Columns
007_attachments.sql attachments table with indexes

No changes to existing tables. The messages.metadata JSONB carries attachment_ids without schema changes. The channels.settings JSONB carries last_selector_id without schema changes.


Decisions Made

  1. Text extraction — Tiered: native Go (pdfcpu Apache 2.0, go-docx MIT) for common cases, LibreOffice headless for everything else including ODF formats. No AGPL dependencies.

  2. Thumbnails — Yes. Generated async post-upload. Images via Go stdlib (golang.org/x/image/draw). Documents via LibreOffice first-page render. Stored alongside originals with _thumb suffix.

  3. Max attachments per message — 5 files (admin-configurable).

  4. Extraction timing — Async with filesystem-backed queue. Upload returns immediately, extraction runs in background via worker loop with concurrency cap. Frontend polls for status. Send blocked only during active upload (bytes in flight), not during extraction.

  5. Smart paste — Automatic, no prompts. Binary clipboard items (screenshots, images) always auto-attach. Text exceeding threshold (default 2000 chars) auto-attaches as {uuid}.txt. Below threshold pastes normally into textarea. UUID filenames, MIME detection from clipboard data. Zero user decisions.

Global Settings (New)

Key Default Description
storage_max_file_size 10485760 (10MB) Per-file upload limit
storage_max_upload_size 52428800 (50MB) Per-request total limit
storage_allowed_types (see Allowed Types) MIME type allowlist
storage_max_attachments_per_message 5 Max files per message
storage_paste_to_file_chars 2000 Paste-to-file threshold (0=disabled)
storage_user_quota_bytes 0 Per-user quota (0=unlimited, future enforcement)

Decisions Made (Continued)

  1. Sidecar base image — Debian slim. Rock-solid LibreOffice packages, ~80MB larger than Alpine but no flaky package issues.

  2. Queue persistence — Dual: filesystem status.json is source of truth (crash recovery, ls debuggability), mirrored to PG metadata for API serving. Worth the bookkeeping for operational transparency.

Open Questions

(None remaining — all design decisions resolved.)

Orphan Cleanup (Admin Action)

Attachments uploaded but never associated with a message (user closes tab, browser crash, abandoned upload) accumulate as orphans.

v0.12.0: manual admin action with "Run Now" button.

POST /admin/storage/cleanup
  → Scan attachments WHERE message_id IS NULL
    AND created_at < NOW() - interval '24 hours'
  → Delete PG rows + storage files + processing dirs
  → Return { deleted: 3, freed_bytes: 15728640 }

Admin Storage panel card:

┌─ Orphan Files ──────────────────────────────┐
│ 3 orphaned attachments (15 MB reclaimable)  │
│                                             │
│ Files not linked to any message for >24hrs  │
│                                             │
│                            [Run Cleanup Now] │
└─────────────────────────────────────────────┘

The 24-hour grace period prevents cleaning up files that are simply in the middle of a long compose session.

Future (v0.22.0 task runner): This becomes a scheduled task running daily. The admin button stays as an on-demand override. The endpoint is identical — the scheduler just calls it on a cron.