This repository has been archived on 2026-04-03. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
core/docs/DESIGN-storage-primitives.md
Jeffrey Smith 55b83baa4c Add v0.8.x design docs and expand roadmap through v1.0
Land two design documents from planning sessions:
- DESIGN-storage-primitives.md (files module, workspace, capability
  negotiation, vector columns)
- DESIGN-extension-composability.md (slots/contributes manifest fields,
  lib.require relaxation, SDK helpers)

Expand ROADMAP.md with detailed v0.8.x-v1.0 plan: storage primitives,
reference extensions, sidecar tier, stable release gate criteria.
Merge design decisions from both planning docs.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 15:18:56 +00:00

21 KiB

DESIGN — Storage Primitives

Version: v0.8.0 Status: Draft Author: Jeff / Claude session 2026-04-02


Problem

Extensions cannot access blob storage or managed disk paths. The existing ObjectStore interface (PVC + S3 backends) is fully implemented but only exposed to the admin status endpoint — no Starlark bridge exists. The db module provides structured data storage via extension-scoped tables, but there is no equivalent for binary files, no managed filesystem for tools like git, and no mechanism for extensions to declare environment requirements (pgvector, workspace root, S3) that the kernel validates at install time.

This blocks every future capability that depends on file handling: RAG/vector search, LLM bridge, code workspaces, file upload/sharing, media processing, and document indexing.


Non-Goals

  • Replacing the db module. Extension-scoped tables (ext_{pkg}_{table}) are the structured data primitive and remain unchanged.
  • Building a KV store. Extensions that need key-value semantics declare a table with key TEXT, value TEXT columns — the infrastructure exists.
  • Multi-tenant file isolation beyond package scoping. Team/user-level file ACLs are extension-layer concerns built on top of these primitives.
  • Streaming / chunked upload through Starlark. Large file ingestion goes through HTTP routes; the files module handles storage after receipt.

Architecture

Three new primitives plus one extension to the existing db module.

Primitive 1: files Starlark Module

Bridges the existing ObjectStore into the sandbox. Follows the same pattern as db_module.go: a Build*Module factory, permission-gated, package-scoped key namespacing.

Permissions: files.read, files.write

Starlark API:

# Write a file. content is string (UTF-8) or bytes.
# metadata is an optional dict stored alongside (JSON-serialized).
files.put(name, content, content_type="application/octet-stream", metadata={})

# Read a file. Returns dict: {"content": <bytes>, "content_type": "...", "size": N, "metadata": {...}}
# Returns None if not found.
result = files.get(name)

# Read metadata only (no content transfer). Returns dict or None.
meta = files.meta(name)

# List files by prefix. Returns list of dicts: [{"name": "...", "size": N, "content_type": "..."}]
entries = files.list(prefix="", limit=100)

# Delete a file. Idempotent — no error if missing.
files.delete(name)

# Delete all files under a prefix. Use with caution.
files.delete_prefix(prefix)

# Check existence without reading.
exists = files.exists(name)

Key namespacing:

All keys are automatically prefixed with ext/{packageID}/. An extension calling files.put("models/v1.bin", data) writes to the ObjectStore key ext/my-extension/models/v1.bin. Extensions cannot escape their namespace.

Name validation rejects .., absolute paths, and control characters — same sanitization rules as physicalTable() in db_module.go.

Implementation notes:

  • New file: sandbox/files_module.go
  • FilesModuleConfig struct mirrors DBModuleConfig:
    type FilesModuleConfig struct {
        PackageID    string
        CanWrite     bool
        Store        storage.ObjectStore
    }
    
  • Metadata is stored as a companion JSON object at key ext/{packageID}/_meta/{name}. This avoids schema changes — the ObjectStore interface is unchanged. The files module manages the companion transparently.
  • Size limit per files.put() call: 50MB (configurable via EXT_FILES_MAX_SIZE). Enforced in the builtin before calling Store.Put().
  • files.get() returns content as starlark.Bytes for binary safety. Starlark's Bytes type was added in go.starlark.net v0.0.0-20240725214946 and handles non-UTF-8 content correctly.
  • If ObjectStore is nil (storage not configured), the module is not injected — same pattern as db module when r.db == nil.

Runner wiring:

// In buildModulesWithLibCtx, add cases:
case models.ExtPermFilesRead:
    if filesLevel < 1 { filesLevel = 1 }
case models.ExtPermFilesWrite:
    filesLevel = 2

// After permission loop:
if filesLevel > 0 && r.objectStore != nil {
    modules["files"] = BuildFilesModule(ctx, FilesModuleConfig{
        PackageID: packageID,
        CanWrite:  filesLevel == 2,
        Store:     r.objectStore,
    })
}

Runner gains SetObjectStore(s storage.ObjectStore) setter, called from main.go after storage.Init().


Primitive 2: workspace Starlark Module

Managed disk directories for extensions that need a real filesystem — git repos, compilers, ffmpeg, pandoc, code analysis tools. These tools cannot operate through put/get blob semantics; they need paths.

Permission: workspace.manage

Starlark API:

# Create a named workspace directory. Returns the absolute path.
# Idempotent — returns existing path if already created.
path = workspace.create(name)

# Get the path for an existing workspace. Returns string or None.
path = workspace.path(name)

# List workspace names for this extension.
names = workspace.list()

# Delete a workspace and all its contents.
workspace.delete(name)

# Disk usage in bytes for a workspace.
size = workspace.usage(name)

Directory layout:

{WORKSPACE_ROOT}/
  {packageID}/
    {name}/
      ... (extension-managed contents)

WORKSPACE_ROOT defaults to /data/workspaces (configurable via WORKSPACE_ROOT env var). The kernel creates the package subdirectory on first workspace.create(). Extensions own everything below their directory — the kernel does not inspect contents.

Implementation notes:

  • New file: sandbox/workspace_module.go
  • WorkspaceModuleConfig:
    type WorkspaceModuleConfig struct {
        PackageID     string
        WorkspaceRoot string
    }
    
  • Name validation: same rules as table names (^[a-z][a-z0-9_]{0,62}$). No path separators, no .., no spaces.
  • workspace.create() calls os.MkdirAll for the scoped path.
  • workspace.delete() calls os.RemoveAll — destructive by design. Extensions must handle confirmation in their own UX.
  • workspace.usage() walks the directory tree and sums file sizes. Bounded by a 10-second context timeout to prevent hangs on huge trees.
  • Quota enforcement (optional): WORKSPACE_QUOTA_MB env var. When set, workspace.create() checks cumulative usage for the package before creating. Returns error if quota exceeded. Default: unlimited.
  • If WORKSPACE_ROOT is empty or not writable, the module is not injected. Extensions that declare workspace.manage without the root configured will have their permission granted but the module absent — same degradation pattern as db when no DB is available.

Runner wiring:

case models.ExtPermWorkspaceManage:
    if r.workspaceRoot != "" {
        modules["workspace"] = BuildWorkspaceModule(ctx, WorkspaceModuleConfig{
            PackageID:     packageID,
            WorkspaceRoot: r.workspaceRoot,
        })
    }

Runner gains SetWorkspaceRoot(path string) setter.

Security considerations:

  • The returned path is an absolute filesystem path. Extensions can pass this to http module calls (e.g., POST a file to an API) or use it in db records as a reference. They cannot execute arbitrary binaries — the Starlark sandbox has no os.exec. Execution requires a sidecar tier package or an api_route handler that shells out server-side.
  • For sidecar-tier packages that can execute binaries, the workspace path is the designated scratch space. The kernel ensures the path is within the scoped directory via filepath.Clean + prefix check.

Primitive 3: Capability Negotiation

Extensions declare environment requirements in their manifest. The kernel validates these at install time and reports failures with actionable messages. This is what makes progressive enhancement strategies (e.g., three-tier vector search) work.

Manifest field:

{
  "id": "vector-store",
  "capabilities": {
    "required": ["files.read", "files.write"],
    "optional": ["pgvector"]
  },
  "permissions": ["db.write", "files.write"],
  "db_tables": {
    "embeddings": {
      "columns": {
        "source_id": "text",
        "chunk_text": "text",
        "embedding": "vector(384)"
      },
      "indexes": [["source_id"]]
    }
  }
}

capabilities.required — install fails if any are missing. Admin gets a clear error: "Package 'vector-store' requires capability 'pgvector' which is not available. Run CREATE EXTENSION vector; in your PostgreSQL database to enable it."

capabilities.optional — install succeeds regardless. The capability state is queryable at runtime so extensions can degrade gracefully.

Runtime query (Starlark):

Settings module is the natural home since it's always available:

# Returns True/False for a capability name.
has_pgvector = settings.has_capability("pgvector")

Kernel capability registry:

A simple function in the handlers package that probes the environment:

func DetectCapabilities(db *sql.DB, isPostgres bool, workspaceRoot string, objStore storage.ObjectStore) map[string]bool {
    caps := make(map[string]bool)

    // pgvector: check pg_extension
    if isPostgres && db != nil {
        var exists bool
        row := db.QueryRow("SELECT EXISTS(SELECT 1 FROM pg_extension WHERE extname='vector')")
        if row.Scan(&exists) == nil && exists {
            caps["pgvector"] = true
        }
    }

    // workspace: check root is writable
    if workspaceRoot != "" && storage.IsPathWritable(workspaceRoot) {
        caps["workspace"] = true
    }

    // object storage: check configured and healthy
    if objStore != nil {
        if err := objStore.Healthy(context.Background()); err == nil {
            caps["object_storage"] = true
        }
    }

    // s3: specific backend check
    if objStore != nil && objStore.Backend() == "s3" {
        caps["s3"] = true
    }

    // postgres: dialect check
    if isPostgres {
        caps["postgres"] = true
    }

    return caps
}

Called once at startup, stored on the Runner (or a shared config struct). Re-probed on admin request for the capabilities endpoint.

Install-time validation:

In the package install handler (handlers/extensions.go), after ParseDBTables and before SyncManifestPermissions:

if caps, ok := ParseCapabilities(manifestMap); ok {
    missing := CheckRequiredCapabilities(caps.Required, detectedCaps)
    if len(missing) > 0 {
        // Roll back: delete the just-created package row
        h.stores.Packages.Delete(c.Request.Context(), pkg.ID)
        c.JSON(422, gin.H{
            "error":   "missing required capabilities",
            "missing": missing,
            "help":    capabilityHelpText(missing),
        })
        return
    }
}

Admin API:

GET /api/v1/admin/capabilities
→ {"pgvector": true, "workspace": true, "object_storage": true, "s3": false, "postgres": true}

Displayed in the Admin UI alongside storage status. Gives operators visibility into what their deployment supports.


Extension to db Module: Vector Column Type

The existing mapColType() in ext_db_schema.go gains a vector type with progressive enhancement across backends.

Manifest declaration:

"db_tables": {
  "embeddings": {
    "columns": {
      "embedding": "vector(384)"
    }
  }
}

Column type mapping:

Manifest type PG + pgvector PG without pgvector SQLite
vector(N) vector(N) JSONB TEXT

Implementation in mapColType:

case "vector":
    // vector or vector(384) — extract dimension if present
    dim := extractVectorDim(typStr) // returns "384" or ""
    if isPostgres && hasPgVector {
        if dim != "" {
            return fmt.Sprintf("vector(%s)", dim)
        }
        return "vector"
    }
    if isPostgres {
        return "JSONB" // store as JSON array, brute-force search
    }
    return "TEXT" // SQLite: JSON array as text

hasPgVector is passed through a new field on a SchemaConfig struct (or detected inline — the capability registry result is available at table creation time).

New db module function — db.query_similar():

# Find rows with embeddings closest to the query vector.
# Returns list of row dicts with _distance appended.
results = db.query_similar(
    table="embeddings",
    column="embedding",
    vector=[0.1, 0.2, ...],  # query vector (list of floats)
    limit=10,
    filters={"source_id": "doc-123"}  # optional WHERE clause
)

Backend dispatch:

  • PG + pgvector: Uses ORDER BY embedding <=> $1 LIMIT $2 with native vector distance operator.
  • PG without pgvector / SQLite: Loads candidate rows (respecting filters), deserializes JSON arrays, computes cosine similarity in Go, sorts, returns top-N. This is the "it works but slowly" fallback — acceptable for small corpora (<10k rows), documented as such.

Implementation: new function dbQuerySimilar() in db_module.go, gated on db.read permission (read-only operation). The function checks cfg.HasPgVector to choose the fast or slow path.

DBModuleConfig gains:

type DBModuleConfig struct {
    PackageID   string
    CanWrite    bool
    DB          *sql.DB
    IsPostgres  bool
    HasPgVector bool  // NEW: enables native vector ops
}

Wired from the capability registry at module build time.


New Permission Constants

// In models/models_extension_perm.go:
const (
    ExtPermFilesRead        = "files.read"
    ExtPermFilesWrite       = "files.write"
    ExtPermWorkspaceManage  = "workspace.manage"
)

Added to ValidExtensionPermissions map.


Schema Changes

Migration 015 — none required for kernel tables.

The files module uses the existing ObjectStore interface — no new kernel tables. Metadata companions are stored as ObjectStore objects.

The workspace module uses the filesystem — no tables.

Capabilities are detected at runtime — no tables.

The vector column type is handled by extension DDL generation in ext_db_schema.go — no kernel migration.

The new permission constants are code-only changes.


New Files

File Purpose
sandbox/files_module.go files Starlark module
sandbox/files_module_test.go Tests using mock ObjectStore
sandbox/workspace_module.go workspace Starlark module
sandbox/workspace_module_test.go Tests using temp directory
handlers/capabilities.go DetectCapabilities(), ParseCapabilities(), admin endpoint
handlers/capabilities_test.go Tests for capability detection and validation

Modified Files

File Change
models/models_extension_perm.go Add files.read, files.write, workspace.manage
sandbox/runner.go Add SetObjectStore(), SetWorkspaceRoot(), wire new modules
sandbox/db_module.go Add dbQuerySimilar(), HasPgVector field
handlers/ext_db_schema.go Extend mapColType() for vector(N), pass hasPgVector
handlers/extensions.go Add capability validation in install handler
sandbox/settings_module.go Add settings.has_capability()
server/main.go Call DetectCapabilities(), wire to runner
docs/PACKAGE-FORMAT.md Document capabilities manifest field
docs/EXTENSION-GUIDE.md Document new modules and vector column type

Changeset Plan

CS1 — files module + permissions

New files: files_module.go, files_module_test.go. Modified: models_extension_perm.go, runner.go, main.go. Scope: ObjectStore bridge, permission constants, runner wiring. CI-green independently — no schema changes, no existing behavior affected.

CS2 — workspace module

New files: workspace_module.go, workspace_module_test.go. Modified: models_extension_perm.go, runner.go, main.go. Scope: Managed disk directories, quota enforcement. CI-green independently.

CS3 — Capability negotiation

New files: capabilities.go, capabilities_test.go. Modified: extensions.go (install validation), settings_module.go (has_capability), main.go. Scope: Detection, install-time validation, runtime query, admin endpoint. CI-green independently.

CS4 — Vector column type + db.query_similar()

Modified: ext_db_schema.go, db_module.go, db_module_test.go. Scope: Column type mapping, similarity query with dual-path dispatch. Depends on CS3 (needs HasPgVector from capability registry).


Configuration Summary

Env Var Default Purpose
WORKSPACE_ROOT /data/workspaces Root directory for workspace module
WORKSPACE_QUOTA_MB 0 (unlimited) Per-extension disk quota
EXT_FILES_MAX_SIZE 52428800 (50MB) Max single file size via files.put()

Example: Vector Store Extension

A vector-store library extension consuming all four primitives:

manifest.json:

{
  "id": "vector-store",
  "title": "Vector Store",
  "type": "library",
  "tier": "starlark",
  "version": "0.1.0",
  "permissions": ["db.write", "files.read", "connections.read"],
  "capabilities": {
    "required": [],
    "optional": ["pgvector"]
  },
  "db_tables": {
    "documents": {
      "columns": {
        "source": "text",
        "chunk_text": "text",
        "page": "int",
        "metadata": "text"
      },
      "indexes": [["source"]]
    },
    "embeddings": {
      "columns": {
        "document_id": "text",
        "embedding": "vector(384)",
        "chunk_index": "int"
      },
      "indexes": [["document_id"]]
    }
  },
  "exports": ["ingest", "search", "search_clustered"]
}

script.star:

def ingest(source_name, chunks):
    """Store document chunks and generate embeddings."""
    for i, chunk in enumerate(chunks):
        row = db.insert("documents", {
            "source": source_name,
            "chunk_text": chunk["text"],
            "page": chunk.get("page", 0),
            "metadata": json.encode(chunk.get("metadata", {})),
        })
        # Embedding generation delegated to caller (llm-bridge)
        # Caller passes pre-computed vectors
        if "embedding" in chunk:
            db.insert("embeddings", {
                "document_id": row["id"],
                "embedding": json.encode(chunk["embedding"]),
                "chunk_index": i,
            })

def search(query_vector, limit=10, source_filter=None):
    """Semantic similarity search — dispatches to native or brute-force."""
    filters = {}
    if source_filter:
        filters["source"] = source_filter
    # query_similar handles pgvector vs fallback transparently
    results = db.query_similar(
        table="embeddings",
        column="embedding",
        vector=query_vector,
        limit=limit,
    )
    # Hydrate with document text
    enriched = []
    for r in results:
        docs = db.query("documents", filters={"id": r["document_id"]}, limit=1)
        if docs:
            enriched.append({
                "text": docs[0]["chunk_text"],
                "source": docs[0]["source"],
                "page": docs[0]["page"],
                "distance": r["_distance"],
            })
    return enriched

def search_clustered(embeddings_list, k):
    """K-means clustering for query-free thematic selection.
    Returns k representative chunks. Clustering runs in the
    extension — the kernel provides the data, not the algorithm."""
    # This would be implemented by a consumer extension with
    # http access to call a clustering API, or by a sidecar
    # that runs sklearn. The vector-store library provides
    # the data access pattern; clustering logic lives elsewhere.
    pass

Future Considerations

  • Content-addressed deduplication. If multiple extensions store the same PDF, the ObjectStore holds duplicate bytes. A SHA256-keyed blob layer with refcounting would deduplicate transparently. Deferred — adds complexity without clear need pre-1.0.

  • Extension-to-extension file sharing. The current design isolates file namespaces per extension. A files.grant(name, target_pkg_id) primitive could enable controlled sharing. Deferred — composition through API routes is sufficient initially.

  • Streaming upload for large files. Starlark files.put() buffers in memory. For files >50MB, extensions should use HTTP api_routes with Go handlers that stream directly to ObjectStore. The files module is for extension-internal storage, not user-facing upload.

  • Workspace snapshots. workspace.snapshot(name) → creates a tarball in the files store. Useful for backup/reproducibility. Deferred — extensions can implement this themselves.

  • Rate limiting / quota on db.query_similar() fallback. The brute-force path loads all candidate rows into Go memory. For large tables this is dangerous. A row-count guard (e.g., refuse if >50k candidates) with a clear error message pointing to pgvector is the right safety valve.