# DESIGN — Storage Primitives **Version:** v0.8.0 **Status:** Draft **Author:** Jeff / Claude session 2026-04-02 --- ## Problem Extensions cannot access blob storage or managed disk paths. The existing `ObjectStore` interface (PVC + S3 backends) is fully implemented but only exposed to the admin status endpoint — no Starlark bridge exists. The `db` module provides structured data storage via extension-scoped tables, but there is no equivalent for binary files, no managed filesystem for tools like `git`, and no mechanism for extensions to declare environment requirements (pgvector, workspace root, S3) that the kernel validates at install time. This blocks every future capability that depends on file handling: RAG/vector search, LLM bridge, code workspaces, file upload/sharing, media processing, and document indexing. --- ## Non-Goals - Replacing the `db` module. Extension-scoped tables (`ext_{pkg}_{table}`) are the structured data primitive and remain unchanged. - Building a KV store. Extensions that need key-value semantics declare a table with `key TEXT, value TEXT` columns — the infrastructure exists. - Multi-tenant file isolation beyond package scoping. Team/user-level file ACLs are extension-layer concerns built on top of these primitives. - Streaming / chunked upload through Starlark. Large file ingestion goes through HTTP routes; the `files` module handles storage after receipt. --- ## Architecture Three new primitives plus one extension to the existing `db` module. ### Primitive 1: `files` Starlark Module Bridges the existing `ObjectStore` into the sandbox. Follows the same pattern as `db_module.go`: a `Build*Module` factory, permission-gated, package-scoped key namespacing. **Permissions:** `files.read`, `files.write` **Starlark API:** ```python # Write a file. content is string (UTF-8) or bytes. # metadata is an optional dict stored alongside (JSON-serialized). files.put(name, content, content_type="application/octet-stream", metadata={}) # Read a file. Returns dict: {"content": , "content_type": "...", "size": N, "metadata": {...}} # Returns None if not found. result = files.get(name) # Read metadata only (no content transfer). Returns dict or None. meta = files.meta(name) # List files by prefix. Returns list of dicts: [{"name": "...", "size": N, "content_type": "..."}] entries = files.list(prefix="", limit=100) # Delete a file. Idempotent — no error if missing. files.delete(name) # Delete all files under a prefix. Use with caution. files.delete_prefix(prefix) # Check existence without reading. exists = files.exists(name) ``` **Key namespacing:** All keys are automatically prefixed with `ext/{packageID}/`. An extension calling `files.put("models/v1.bin", data)` writes to the ObjectStore key `ext/my-extension/models/v1.bin`. Extensions cannot escape their namespace. Name validation rejects `..`, absolute paths, and control characters — same sanitization rules as `physicalTable()` in `db_module.go`. **Implementation notes:** - New file: `sandbox/files_module.go` - `FilesModuleConfig` struct mirrors `DBModuleConfig`: ```go type FilesModuleConfig struct { PackageID string CanWrite bool Store storage.ObjectStore } ``` - Metadata is stored as a companion JSON object at key `ext/{packageID}/_meta/{name}`. This avoids schema changes — the ObjectStore interface is unchanged. The `files` module manages the companion transparently. - Size limit per `files.put()` call: 50MB (configurable via `EXT_FILES_MAX_SIZE`). Enforced in the builtin before calling `Store.Put()`. - `files.get()` returns content as `starlark.Bytes` for binary safety. Starlark's `Bytes` type was added in go.starlark.net v0.0.0-20240725214946 and handles non-UTF-8 content correctly. - If `ObjectStore` is nil (storage not configured), the module is not injected — same pattern as `db` module when `r.db == nil`. **Runner wiring:** ```go // In buildModulesWithLibCtx, add cases: case models.ExtPermFilesRead: if filesLevel < 1 { filesLevel = 1 } case models.ExtPermFilesWrite: filesLevel = 2 // After permission loop: if filesLevel > 0 && r.objectStore != nil { modules["files"] = BuildFilesModule(ctx, FilesModuleConfig{ PackageID: packageID, CanWrite: filesLevel == 2, Store: r.objectStore, }) } ``` Runner gains `SetObjectStore(s storage.ObjectStore)` setter, called from `main.go` after `storage.Init()`. --- ### Primitive 2: `workspace` Starlark Module Managed disk directories for extensions that need a real filesystem — git repos, compilers, ffmpeg, pandoc, code analysis tools. These tools cannot operate through put/get blob semantics; they need paths. **Permission:** `workspace.manage` **Starlark API:** ```python # Create a named workspace directory. Returns the absolute path. # Idempotent — returns existing path if already created. path = workspace.create(name) # Get the path for an existing workspace. Returns string or None. path = workspace.path(name) # List workspace names for this extension. names = workspace.list() # Delete a workspace and all its contents. workspace.delete(name) # Disk usage in bytes for a workspace. size = workspace.usage(name) ``` **Directory layout:** ``` {WORKSPACE_ROOT}/ {packageID}/ {name}/ ... (extension-managed contents) ``` `WORKSPACE_ROOT` defaults to `/data/workspaces` (configurable via `WORKSPACE_ROOT` env var). The kernel creates the package subdirectory on first `workspace.create()`. Extensions own everything below their directory — the kernel does not inspect contents. **Implementation notes:** - New file: `sandbox/workspace_module.go` - `WorkspaceModuleConfig`: ```go type WorkspaceModuleConfig struct { PackageID string WorkspaceRoot string } ``` - Name validation: same rules as table names (`^[a-z][a-z0-9_]{0,62}$`). No path separators, no `..`, no spaces. - `workspace.create()` calls `os.MkdirAll` for the scoped path. - `workspace.delete()` calls `os.RemoveAll` — destructive by design. Extensions must handle confirmation in their own UX. - `workspace.usage()` walks the directory tree and sums file sizes. Bounded by a 10-second context timeout to prevent hangs on huge trees. - **Quota enforcement** (optional): `WORKSPACE_QUOTA_MB` env var. When set, `workspace.create()` checks cumulative usage for the package before creating. Returns error if quota exceeded. Default: unlimited. - If `WORKSPACE_ROOT` is empty or not writable, the module is not injected. Extensions that declare `workspace.manage` without the root configured will have their permission granted but the module absent — same degradation pattern as `db` when no DB is available. **Runner wiring:** ```go case models.ExtPermWorkspaceManage: if r.workspaceRoot != "" { modules["workspace"] = BuildWorkspaceModule(ctx, WorkspaceModuleConfig{ PackageID: packageID, WorkspaceRoot: r.workspaceRoot, }) } ``` Runner gains `SetWorkspaceRoot(path string)` setter. **Security considerations:** - The returned path is an absolute filesystem path. Extensions can pass this to `http` module calls (e.g., POST a file to an API) or use it in `db` records as a reference. They cannot execute arbitrary binaries — the Starlark sandbox has no `os.exec`. Execution requires a sidecar tier package or an `api_route` handler that shells out server-side. - For sidecar-tier packages that _can_ execute binaries, the workspace path is the designated scratch space. The kernel ensures the path is within the scoped directory via `filepath.Clean` + prefix check. --- ### Primitive 3: Capability Negotiation Extensions declare environment requirements in their manifest. The kernel validates these at install time and reports failures with actionable messages. This is what makes progressive enhancement strategies (e.g., three-tier vector search) work. **Manifest field:** ```json { "id": "vector-store", "capabilities": { "required": ["files.read", "files.write"], "optional": ["pgvector"] }, "permissions": ["db.write", "files.write"], "db_tables": { "embeddings": { "columns": { "source_id": "text", "chunk_text": "text", "embedding": "vector(384)" }, "indexes": [["source_id"]] } } } ``` `capabilities.required` — install fails if any are missing. Admin gets a clear error: "Package 'vector-store' requires capability 'pgvector' which is not available. Run `CREATE EXTENSION vector;` in your PostgreSQL database to enable it." `capabilities.optional` — install succeeds regardless. The capability state is queryable at runtime so extensions can degrade gracefully. **Runtime query (Starlark):** Settings module is the natural home since it's always available: ```python # Returns True/False for a capability name. has_pgvector = settings.has_capability("pgvector") ``` **Kernel capability registry:** A simple function in the `handlers` package that probes the environment: ```go func DetectCapabilities(db *sql.DB, isPostgres bool, workspaceRoot string, objStore storage.ObjectStore) map[string]bool { caps := make(map[string]bool) // pgvector: check pg_extension if isPostgres && db != nil { var exists bool row := db.QueryRow("SELECT EXISTS(SELECT 1 FROM pg_extension WHERE extname='vector')") if row.Scan(&exists) == nil && exists { caps["pgvector"] = true } } // workspace: check root is writable if workspaceRoot != "" && storage.IsPathWritable(workspaceRoot) { caps["workspace"] = true } // object storage: check configured and healthy if objStore != nil { if err := objStore.Healthy(context.Background()); err == nil { caps["object_storage"] = true } } // s3: specific backend check if objStore != nil && objStore.Backend() == "s3" { caps["s3"] = true } // postgres: dialect check if isPostgres { caps["postgres"] = true } return caps } ``` Called once at startup, stored on the `Runner` (or a shared config struct). Re-probed on admin request for the capabilities endpoint. **Install-time validation:** In the package install handler (`handlers/extensions.go`), after `ParseDBTables` and before `SyncManifestPermissions`: ```go if caps, ok := ParseCapabilities(manifestMap); ok { missing := CheckRequiredCapabilities(caps.Required, detectedCaps) if len(missing) > 0 { // Roll back: delete the just-created package row h.stores.Packages.Delete(c.Request.Context(), pkg.ID) c.JSON(422, gin.H{ "error": "missing required capabilities", "missing": missing, "help": capabilityHelpText(missing), }) return } } ``` **Admin API:** ``` GET /api/v1/admin/capabilities → {"pgvector": true, "workspace": true, "object_storage": true, "s3": false, "postgres": true} ``` Displayed in the Admin UI alongside storage status. Gives operators visibility into what their deployment supports. --- ### Extension to `db` Module: Vector Column Type The existing `mapColType()` in `ext_db_schema.go` gains a `vector` type with progressive enhancement across backends. **Manifest declaration:** ```json "db_tables": { "embeddings": { "columns": { "embedding": "vector(384)" } } } ``` **Column type mapping:** | Manifest type | PG + pgvector | PG without pgvector | SQLite | |------------------|-------------------------|---------------------|--------------| | `vector(N)` | `vector(N)` | `JSONB` | `TEXT` | Implementation in `mapColType`: ```go case "vector": // vector or vector(384) — extract dimension if present dim := extractVectorDim(typStr) // returns "384" or "" if isPostgres && hasPgVector { if dim != "" { return fmt.Sprintf("vector(%s)", dim) } return "vector" } if isPostgres { return "JSONB" // store as JSON array, brute-force search } return "TEXT" // SQLite: JSON array as text ``` `hasPgVector` is passed through a new field on a `SchemaConfig` struct (or detected inline — the capability registry result is available at table creation time). **New `db` module function — `db.query_similar()`:** ```python # Find rows with embeddings closest to the query vector. # Returns list of row dicts with _distance appended. results = db.query_similar( table="embeddings", column="embedding", vector=[0.1, 0.2, ...], # query vector (list of floats) limit=10, filters={"source_id": "doc-123"} # optional WHERE clause ) ``` **Backend dispatch:** - **PG + pgvector:** Uses `ORDER BY embedding <=> $1 LIMIT $2` with native vector distance operator. - **PG without pgvector / SQLite:** Loads candidate rows (respecting `filters`), deserializes JSON arrays, computes cosine similarity in Go, sorts, returns top-N. This is the "it works but slowly" fallback — acceptable for small corpora (<10k rows), documented as such. Implementation: new function `dbQuerySimilar()` in `db_module.go`, gated on `db.read` permission (read-only operation). The function checks `cfg.HasPgVector` to choose the fast or slow path. `DBModuleConfig` gains: ```go type DBModuleConfig struct { PackageID string CanWrite bool DB *sql.DB IsPostgres bool HasPgVector bool // NEW: enables native vector ops } ``` Wired from the capability registry at module build time. --- ## New Permission Constants ```go // In models/models_extension_perm.go: const ( ExtPermFilesRead = "files.read" ExtPermFilesWrite = "files.write" ExtPermWorkspaceManage = "workspace.manage" ) ``` Added to `ValidExtensionPermissions` map. --- ## Schema Changes **Migration 015 — none required for kernel tables.** The `files` module uses the existing `ObjectStore` interface — no new kernel tables. Metadata companions are stored as ObjectStore objects. The `workspace` module uses the filesystem — no tables. Capabilities are detected at runtime — no tables. The `vector` column type is handled by extension DDL generation in `ext_db_schema.go` — no kernel migration. The new permission constants are code-only changes. --- ## New Files | File | Purpose | |------|---------| | `sandbox/files_module.go` | `files` Starlark module | | `sandbox/files_module_test.go` | Tests using mock ObjectStore | | `sandbox/workspace_module.go` | `workspace` Starlark module | | `sandbox/workspace_module_test.go` | Tests using temp directory | | `handlers/capabilities.go` | `DetectCapabilities()`, `ParseCapabilities()`, admin endpoint | | `handlers/capabilities_test.go` | Tests for capability detection and validation | ## Modified Files | File | Change | |------|--------| | `models/models_extension_perm.go` | Add `files.read`, `files.write`, `workspace.manage` | | `sandbox/runner.go` | Add `SetObjectStore()`, `SetWorkspaceRoot()`, wire new modules | | `sandbox/db_module.go` | Add `dbQuerySimilar()`, `HasPgVector` field | | `handlers/ext_db_schema.go` | Extend `mapColType()` for `vector(N)`, pass `hasPgVector` | | `handlers/extensions.go` | Add capability validation in install handler | | `sandbox/settings_module.go` | Add `settings.has_capability()` | | `server/main.go` | Call `DetectCapabilities()`, wire to runner | | `docs/PACKAGE-FORMAT.md` | Document `capabilities` manifest field | | `docs/EXTENSION-GUIDE.md` | Document new modules and vector column type | --- ## Changeset Plan **CS1 — `files` module + permissions** New files: `files_module.go`, `files_module_test.go`. Modified: `models_extension_perm.go`, `runner.go`, `main.go`. Scope: ObjectStore bridge, permission constants, runner wiring. CI-green independently — no schema changes, no existing behavior affected. **CS2 — `workspace` module** New files: `workspace_module.go`, `workspace_module_test.go`. Modified: `models_extension_perm.go`, `runner.go`, `main.go`. Scope: Managed disk directories, quota enforcement. CI-green independently. **CS3 — Capability negotiation** New files: `capabilities.go`, `capabilities_test.go`. Modified: `extensions.go` (install validation), `settings_module.go` (`has_capability`), `main.go`. Scope: Detection, install-time validation, runtime query, admin endpoint. CI-green independently. **CS4 — Vector column type + `db.query_similar()`** Modified: `ext_db_schema.go`, `db_module.go`, `db_module_test.go`. Scope: Column type mapping, similarity query with dual-path dispatch. Depends on CS3 (needs `HasPgVector` from capability registry). --- ## Configuration Summary | Env Var | Default | Purpose | |---------|---------|---------| | `WORKSPACE_ROOT` | `/data/workspaces` | Root directory for workspace module | | `WORKSPACE_QUOTA_MB` | `0` (unlimited) | Per-extension disk quota | | `EXT_FILES_MAX_SIZE` | `52428800` (50MB) | Max single file size via `files.put()` | --- ## Example: Vector Store Extension A `vector-store` library extension consuming all four primitives: **manifest.json:** ```json { "id": "vector-store", "title": "Vector Store", "type": "library", "tier": "starlark", "version": "0.1.0", "permissions": ["db.write", "files.read", "connections.read"], "capabilities": { "required": [], "optional": ["pgvector"] }, "db_tables": { "documents": { "columns": { "source": "text", "chunk_text": "text", "page": "int", "metadata": "text" }, "indexes": [["source"]] }, "embeddings": { "columns": { "document_id": "text", "embedding": "vector(384)", "chunk_index": "int" }, "indexes": [["document_id"]] } }, "exports": ["ingest", "search", "search_clustered"] } ``` **script.star:** ```python def ingest(source_name, chunks): """Store document chunks and generate embeddings.""" for i, chunk in enumerate(chunks): row = db.insert("documents", { "source": source_name, "chunk_text": chunk["text"], "page": chunk.get("page", 0), "metadata": json.encode(chunk.get("metadata", {})), }) # Embedding generation delegated to caller (llm-bridge) # Caller passes pre-computed vectors if "embedding" in chunk: db.insert("embeddings", { "document_id": row["id"], "embedding": json.encode(chunk["embedding"]), "chunk_index": i, }) def search(query_vector, limit=10, source_filter=None): """Semantic similarity search — dispatches to native or brute-force.""" filters = {} if source_filter: filters["source"] = source_filter # query_similar handles pgvector vs fallback transparently results = db.query_similar( table="embeddings", column="embedding", vector=query_vector, limit=limit, ) # Hydrate with document text enriched = [] for r in results: docs = db.query("documents", filters={"id": r["document_id"]}, limit=1) if docs: enriched.append({ "text": docs[0]["chunk_text"], "source": docs[0]["source"], "page": docs[0]["page"], "distance": r["_distance"], }) return enriched def search_clustered(embeddings_list, k): """K-means clustering for query-free thematic selection. Returns k representative chunks. Clustering runs in the extension — the kernel provides the data, not the algorithm.""" # This would be implemented by a consumer extension with # http access to call a clustering API, or by a sidecar # that runs sklearn. The vector-store library provides # the data access pattern; clustering logic lives elsewhere. pass ``` --- ## Future Considerations - **Content-addressed deduplication.** If multiple extensions store the same PDF, the ObjectStore holds duplicate bytes. A SHA256-keyed blob layer with refcounting would deduplicate transparently. Deferred — adds complexity without clear need pre-1.0. - **Extension-to-extension file sharing.** The current design isolates file namespaces per extension. A `files.grant(name, target_pkg_id)` primitive could enable controlled sharing. Deferred — composition through API routes is sufficient initially. - **Streaming upload for large files.** Starlark `files.put()` buffers in memory. For files >50MB, extensions should use HTTP `api_routes` with Go handlers that stream directly to ObjectStore. The `files` module is for extension-internal storage, not user-facing upload. - **Workspace snapshots.** `workspace.snapshot(name)` → creates a tarball in the `files` store. Useful for backup/reproducibility. Deferred — extensions can implement this themselves. - **Rate limiting / quota on `db.query_similar()` fallback.** The brute-force path loads all candidate rows into Go memory. For large tables this is dangerous. A row-count guard (e.g., refuse if >50k candidates) with a clear error message pointing to pgvector is the right safety valve.