Land two design documents from planning sessions: - DESIGN-storage-primitives.md (files module, workspace, capability negotiation, vector columns) - DESIGN-extension-composability.md (slots/contributes manifest fields, lib.require relaxation, SDK helpers) Expand ROADMAP.md with detailed v0.8.x-v1.0 plan: storage primitives, reference extensions, sidecar tier, stable release gate criteria. Merge design decisions from both planning docs. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
21 KiB
DESIGN — Storage Primitives
Version: v0.8.0 Status: Draft Author: Jeff / Claude session 2026-04-02
Problem
Extensions cannot access blob storage or managed disk paths. The existing
ObjectStore interface (PVC + S3 backends) is fully implemented but only
exposed to the admin status endpoint — no Starlark bridge exists. The db
module provides structured data storage via extension-scoped tables, but
there is no equivalent for binary files, no managed filesystem for tools
like git, and no mechanism for extensions to declare environment
requirements (pgvector, workspace root, S3) that the kernel validates at
install time.
This blocks every future capability that depends on file handling: RAG/vector search, LLM bridge, code workspaces, file upload/sharing, media processing, and document indexing.
Non-Goals
- Replacing the
dbmodule. Extension-scoped tables (ext_{pkg}_{table}) are the structured data primitive and remain unchanged. - Building a KV store. Extensions that need key-value semantics declare a
table with
key TEXT, value TEXTcolumns — the infrastructure exists. - Multi-tenant file isolation beyond package scoping. Team/user-level file ACLs are extension-layer concerns built on top of these primitives.
- Streaming / chunked upload through Starlark. Large file ingestion goes
through HTTP routes; the
filesmodule handles storage after receipt.
Architecture
Three new primitives plus one extension to the existing db module.
Primitive 1: files Starlark Module
Bridges the existing ObjectStore into the sandbox. Follows the same
pattern as db_module.go: a Build*Module factory, permission-gated,
package-scoped key namespacing.
Permissions: files.read, files.write
Starlark API:
# Write a file. content is string (UTF-8) or bytes.
# metadata is an optional dict stored alongside (JSON-serialized).
files.put(name, content, content_type="application/octet-stream", metadata={})
# Read a file. Returns dict: {"content": <bytes>, "content_type": "...", "size": N, "metadata": {...}}
# Returns None if not found.
result = files.get(name)
# Read metadata only (no content transfer). Returns dict or None.
meta = files.meta(name)
# List files by prefix. Returns list of dicts: [{"name": "...", "size": N, "content_type": "..."}]
entries = files.list(prefix="", limit=100)
# Delete a file. Idempotent — no error if missing.
files.delete(name)
# Delete all files under a prefix. Use with caution.
files.delete_prefix(prefix)
# Check existence without reading.
exists = files.exists(name)
Key namespacing:
All keys are automatically prefixed with ext/{packageID}/. An extension
calling files.put("models/v1.bin", data) writes to the ObjectStore key
ext/my-extension/models/v1.bin. Extensions cannot escape their namespace.
Name validation rejects .., absolute paths, and control characters —
same sanitization rules as physicalTable() in db_module.go.
Implementation notes:
- New file:
sandbox/files_module.go FilesModuleConfigstruct mirrorsDBModuleConfig:type FilesModuleConfig struct { PackageID string CanWrite bool Store storage.ObjectStore }- Metadata is stored as a companion JSON object at key
ext/{packageID}/_meta/{name}. This avoids schema changes — the ObjectStore interface is unchanged. Thefilesmodule manages the companion transparently. - Size limit per
files.put()call: 50MB (configurable viaEXT_FILES_MAX_SIZE). Enforced in the builtin before callingStore.Put(). files.get()returns content asstarlark.Bytesfor binary safety. Starlark'sBytestype was added in go.starlark.net v0.0.0-20240725214946 and handles non-UTF-8 content correctly.- If
ObjectStoreis nil (storage not configured), the module is not injected — same pattern asdbmodule whenr.db == nil.
Runner wiring:
// In buildModulesWithLibCtx, add cases:
case models.ExtPermFilesRead:
if filesLevel < 1 { filesLevel = 1 }
case models.ExtPermFilesWrite:
filesLevel = 2
// After permission loop:
if filesLevel > 0 && r.objectStore != nil {
modules["files"] = BuildFilesModule(ctx, FilesModuleConfig{
PackageID: packageID,
CanWrite: filesLevel == 2,
Store: r.objectStore,
})
}
Runner gains SetObjectStore(s storage.ObjectStore) setter, called from
main.go after storage.Init().
Primitive 2: workspace Starlark Module
Managed disk directories for extensions that need a real filesystem — git repos, compilers, ffmpeg, pandoc, code analysis tools. These tools cannot operate through put/get blob semantics; they need paths.
Permission: workspace.manage
Starlark API:
# Create a named workspace directory. Returns the absolute path.
# Idempotent — returns existing path if already created.
path = workspace.create(name)
# Get the path for an existing workspace. Returns string or None.
path = workspace.path(name)
# List workspace names for this extension.
names = workspace.list()
# Delete a workspace and all its contents.
workspace.delete(name)
# Disk usage in bytes for a workspace.
size = workspace.usage(name)
Directory layout:
{WORKSPACE_ROOT}/
{packageID}/
{name}/
... (extension-managed contents)
WORKSPACE_ROOT defaults to /data/workspaces (configurable via
WORKSPACE_ROOT env var). The kernel creates the package subdirectory
on first workspace.create(). Extensions own everything below their
directory — the kernel does not inspect contents.
Implementation notes:
- New file:
sandbox/workspace_module.go WorkspaceModuleConfig:type WorkspaceModuleConfig struct { PackageID string WorkspaceRoot string }- Name validation: same rules as table names (
^[a-z][a-z0-9_]{0,62}$). No path separators, no.., no spaces. workspace.create()callsos.MkdirAllfor the scoped path.workspace.delete()callsos.RemoveAll— destructive by design. Extensions must handle confirmation in their own UX.workspace.usage()walks the directory tree and sums file sizes. Bounded by a 10-second context timeout to prevent hangs on huge trees.- Quota enforcement (optional):
WORKSPACE_QUOTA_MBenv var. When set,workspace.create()checks cumulative usage for the package before creating. Returns error if quota exceeded. Default: unlimited. - If
WORKSPACE_ROOTis empty or not writable, the module is not injected. Extensions that declareworkspace.managewithout the root configured will have their permission granted but the module absent — same degradation pattern asdbwhen no DB is available.
Runner wiring:
case models.ExtPermWorkspaceManage:
if r.workspaceRoot != "" {
modules["workspace"] = BuildWorkspaceModule(ctx, WorkspaceModuleConfig{
PackageID: packageID,
WorkspaceRoot: r.workspaceRoot,
})
}
Runner gains SetWorkspaceRoot(path string) setter.
Security considerations:
- The returned path is an absolute filesystem path. Extensions can pass
this to
httpmodule calls (e.g., POST a file to an API) or use it indbrecords as a reference. They cannot execute arbitrary binaries — the Starlark sandbox has noos.exec. Execution requires a sidecar tier package or anapi_routehandler that shells out server-side. - For sidecar-tier packages that can execute binaries, the workspace
path is the designated scratch space. The kernel ensures the path is
within the scoped directory via
filepath.Clean+ prefix check.
Primitive 3: Capability Negotiation
Extensions declare environment requirements in their manifest. The kernel validates these at install time and reports failures with actionable messages. This is what makes progressive enhancement strategies (e.g., three-tier vector search) work.
Manifest field:
{
"id": "vector-store",
"capabilities": {
"required": ["files.read", "files.write"],
"optional": ["pgvector"]
},
"permissions": ["db.write", "files.write"],
"db_tables": {
"embeddings": {
"columns": {
"source_id": "text",
"chunk_text": "text",
"embedding": "vector(384)"
},
"indexes": [["source_id"]]
}
}
}
capabilities.required — install fails if any are missing. Admin gets
a clear error: "Package 'vector-store' requires capability 'pgvector'
which is not available. Run CREATE EXTENSION vector; in your PostgreSQL
database to enable it."
capabilities.optional — install succeeds regardless. The capability
state is queryable at runtime so extensions can degrade gracefully.
Runtime query (Starlark):
Settings module is the natural home since it's always available:
# Returns True/False for a capability name.
has_pgvector = settings.has_capability("pgvector")
Kernel capability registry:
A simple function in the handlers package that probes the environment:
func DetectCapabilities(db *sql.DB, isPostgres bool, workspaceRoot string, objStore storage.ObjectStore) map[string]bool {
caps := make(map[string]bool)
// pgvector: check pg_extension
if isPostgres && db != nil {
var exists bool
row := db.QueryRow("SELECT EXISTS(SELECT 1 FROM pg_extension WHERE extname='vector')")
if row.Scan(&exists) == nil && exists {
caps["pgvector"] = true
}
}
// workspace: check root is writable
if workspaceRoot != "" && storage.IsPathWritable(workspaceRoot) {
caps["workspace"] = true
}
// object storage: check configured and healthy
if objStore != nil {
if err := objStore.Healthy(context.Background()); err == nil {
caps["object_storage"] = true
}
}
// s3: specific backend check
if objStore != nil && objStore.Backend() == "s3" {
caps["s3"] = true
}
// postgres: dialect check
if isPostgres {
caps["postgres"] = true
}
return caps
}
Called once at startup, stored on the Runner (or a shared config
struct). Re-probed on admin request for the capabilities endpoint.
Install-time validation:
In the package install handler (handlers/extensions.go), after
ParseDBTables and before SyncManifestPermissions:
if caps, ok := ParseCapabilities(manifestMap); ok {
missing := CheckRequiredCapabilities(caps.Required, detectedCaps)
if len(missing) > 0 {
// Roll back: delete the just-created package row
h.stores.Packages.Delete(c.Request.Context(), pkg.ID)
c.JSON(422, gin.H{
"error": "missing required capabilities",
"missing": missing,
"help": capabilityHelpText(missing),
})
return
}
}
Admin API:
GET /api/v1/admin/capabilities
→ {"pgvector": true, "workspace": true, "object_storage": true, "s3": false, "postgres": true}
Displayed in the Admin UI alongside storage status. Gives operators visibility into what their deployment supports.
Extension to db Module: Vector Column Type
The existing mapColType() in ext_db_schema.go gains a vector type
with progressive enhancement across backends.
Manifest declaration:
"db_tables": {
"embeddings": {
"columns": {
"embedding": "vector(384)"
}
}
}
Column type mapping:
| Manifest type | PG + pgvector | PG without pgvector | SQLite |
|---|---|---|---|
vector(N) |
vector(N) |
JSONB |
TEXT |
Implementation in mapColType:
case "vector":
// vector or vector(384) — extract dimension if present
dim := extractVectorDim(typStr) // returns "384" or ""
if isPostgres && hasPgVector {
if dim != "" {
return fmt.Sprintf("vector(%s)", dim)
}
return "vector"
}
if isPostgres {
return "JSONB" // store as JSON array, brute-force search
}
return "TEXT" // SQLite: JSON array as text
hasPgVector is passed through a new field on a SchemaConfig struct
(or detected inline — the capability registry result is available at
table creation time).
New db module function — db.query_similar():
# Find rows with embeddings closest to the query vector.
# Returns list of row dicts with _distance appended.
results = db.query_similar(
table="embeddings",
column="embedding",
vector=[0.1, 0.2, ...], # query vector (list of floats)
limit=10,
filters={"source_id": "doc-123"} # optional WHERE clause
)
Backend dispatch:
- PG + pgvector: Uses
ORDER BY embedding <=> $1 LIMIT $2with native vector distance operator. - PG without pgvector / SQLite: Loads candidate rows (respecting
filters), deserializes JSON arrays, computes cosine similarity in Go, sorts, returns top-N. This is the "it works but slowly" fallback — acceptable for small corpora (<10k rows), documented as such.
Implementation: new function dbQuerySimilar() in db_module.go,
gated on db.read permission (read-only operation). The function
checks cfg.HasPgVector to choose the fast or slow path.
DBModuleConfig gains:
type DBModuleConfig struct {
PackageID string
CanWrite bool
DB *sql.DB
IsPostgres bool
HasPgVector bool // NEW: enables native vector ops
}
Wired from the capability registry at module build time.
New Permission Constants
// In models/models_extension_perm.go:
const (
ExtPermFilesRead = "files.read"
ExtPermFilesWrite = "files.write"
ExtPermWorkspaceManage = "workspace.manage"
)
Added to ValidExtensionPermissions map.
Schema Changes
Migration 015 — none required for kernel tables.
The files module uses the existing ObjectStore interface — no new
kernel tables. Metadata companions are stored as ObjectStore objects.
The workspace module uses the filesystem — no tables.
Capabilities are detected at runtime — no tables.
The vector column type is handled by extension DDL generation in
ext_db_schema.go — no kernel migration.
The new permission constants are code-only changes.
New Files
| File | Purpose |
|---|---|
sandbox/files_module.go |
files Starlark module |
sandbox/files_module_test.go |
Tests using mock ObjectStore |
sandbox/workspace_module.go |
workspace Starlark module |
sandbox/workspace_module_test.go |
Tests using temp directory |
handlers/capabilities.go |
DetectCapabilities(), ParseCapabilities(), admin endpoint |
handlers/capabilities_test.go |
Tests for capability detection and validation |
Modified Files
| File | Change |
|---|---|
models/models_extension_perm.go |
Add files.read, files.write, workspace.manage |
sandbox/runner.go |
Add SetObjectStore(), SetWorkspaceRoot(), wire new modules |
sandbox/db_module.go |
Add dbQuerySimilar(), HasPgVector field |
handlers/ext_db_schema.go |
Extend mapColType() for vector(N), pass hasPgVector |
handlers/extensions.go |
Add capability validation in install handler |
sandbox/settings_module.go |
Add settings.has_capability() |
server/main.go |
Call DetectCapabilities(), wire to runner |
docs/PACKAGE-FORMAT.md |
Document capabilities manifest field |
docs/EXTENSION-GUIDE.md |
Document new modules and vector column type |
Changeset Plan
CS1 — files module + permissions
New files: files_module.go, files_module_test.go.
Modified: models_extension_perm.go, runner.go, main.go.
Scope: ObjectStore bridge, permission constants, runner wiring.
CI-green independently — no schema changes, no existing behavior affected.
CS2 — workspace module
New files: workspace_module.go, workspace_module_test.go.
Modified: models_extension_perm.go, runner.go, main.go.
Scope: Managed disk directories, quota enforcement.
CI-green independently.
CS3 — Capability negotiation
New files: capabilities.go, capabilities_test.go.
Modified: extensions.go (install validation), settings_module.go
(has_capability), main.go.
Scope: Detection, install-time validation, runtime query, admin endpoint.
CI-green independently.
CS4 — Vector column type + db.query_similar()
Modified: ext_db_schema.go, db_module.go, db_module_test.go.
Scope: Column type mapping, similarity query with dual-path dispatch.
Depends on CS3 (needs HasPgVector from capability registry).
Configuration Summary
| Env Var | Default | Purpose |
|---|---|---|
WORKSPACE_ROOT |
/data/workspaces |
Root directory for workspace module |
WORKSPACE_QUOTA_MB |
0 (unlimited) |
Per-extension disk quota |
EXT_FILES_MAX_SIZE |
52428800 (50MB) |
Max single file size via files.put() |
Example: Vector Store Extension
A vector-store library extension consuming all four primitives:
manifest.json:
{
"id": "vector-store",
"title": "Vector Store",
"type": "library",
"tier": "starlark",
"version": "0.1.0",
"permissions": ["db.write", "files.read", "connections.read"],
"capabilities": {
"required": [],
"optional": ["pgvector"]
},
"db_tables": {
"documents": {
"columns": {
"source": "text",
"chunk_text": "text",
"page": "int",
"metadata": "text"
},
"indexes": [["source"]]
},
"embeddings": {
"columns": {
"document_id": "text",
"embedding": "vector(384)",
"chunk_index": "int"
},
"indexes": [["document_id"]]
}
},
"exports": ["ingest", "search", "search_clustered"]
}
script.star:
def ingest(source_name, chunks):
"""Store document chunks and generate embeddings."""
for i, chunk in enumerate(chunks):
row = db.insert("documents", {
"source": source_name,
"chunk_text": chunk["text"],
"page": chunk.get("page", 0),
"metadata": json.encode(chunk.get("metadata", {})),
})
# Embedding generation delegated to caller (llm-bridge)
# Caller passes pre-computed vectors
if "embedding" in chunk:
db.insert("embeddings", {
"document_id": row["id"],
"embedding": json.encode(chunk["embedding"]),
"chunk_index": i,
})
def search(query_vector, limit=10, source_filter=None):
"""Semantic similarity search — dispatches to native or brute-force."""
filters = {}
if source_filter:
filters["source"] = source_filter
# query_similar handles pgvector vs fallback transparently
results = db.query_similar(
table="embeddings",
column="embedding",
vector=query_vector,
limit=limit,
)
# Hydrate with document text
enriched = []
for r in results:
docs = db.query("documents", filters={"id": r["document_id"]}, limit=1)
if docs:
enriched.append({
"text": docs[0]["chunk_text"],
"source": docs[0]["source"],
"page": docs[0]["page"],
"distance": r["_distance"],
})
return enriched
def search_clustered(embeddings_list, k):
"""K-means clustering for query-free thematic selection.
Returns k representative chunks. Clustering runs in the
extension — the kernel provides the data, not the algorithm."""
# This would be implemented by a consumer extension with
# http access to call a clustering API, or by a sidecar
# that runs sklearn. The vector-store library provides
# the data access pattern; clustering logic lives elsewhere.
pass
Future Considerations
-
Content-addressed deduplication. If multiple extensions store the same PDF, the ObjectStore holds duplicate bytes. A SHA256-keyed blob layer with refcounting would deduplicate transparently. Deferred — adds complexity without clear need pre-1.0.
-
Extension-to-extension file sharing. The current design isolates file namespaces per extension. A
files.grant(name, target_pkg_id)primitive could enable controlled sharing. Deferred — composition through API routes is sufficient initially. -
Streaming upload for large files. Starlark
files.put()buffers in memory. For files >50MB, extensions should use HTTPapi_routeswith Go handlers that stream directly to ObjectStore. Thefilesmodule is for extension-internal storage, not user-facing upload. -
Workspace snapshots.
workspace.snapshot(name)→ creates a tarball in thefilesstore. Useful for backup/reproducibility. Deferred — extensions can implement this themselves. -
Rate limiting / quota on
db.query_similar()fallback. The brute-force path loads all candidate rows into Go memory. For large tables this is dangerous. A row-count guard (e.g., refuse if >50k candidates) with a clear error message pointing to pgvector is the right safety valve.