This repository has been archived on 2026-04-03. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
core/docs/DESIGN-storage-primitives.md
Jeffrey Smith 55b83baa4c Add v0.8.x design docs and expand roadmap through v1.0
Land two design documents from planning sessions:
- DESIGN-storage-primitives.md (files module, workspace, capability
  negotiation, vector columns)
- DESIGN-extension-composability.md (slots/contributes manifest fields,
  lib.require relaxation, SDK helpers)

Expand ROADMAP.md with detailed v0.8.x-v1.0 plan: storage primitives,
reference extensions, sidecar tier, stable release gate criteria.
Merge design decisions from both planning docs.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 15:18:56 +00:00

669 lines
21 KiB
Markdown

# DESIGN — Storage Primitives
**Version:** v0.8.0
**Status:** Draft
**Author:** Jeff / Claude session 2026-04-02
---
## Problem
Extensions cannot access blob storage or managed disk paths. The existing
`ObjectStore` interface (PVC + S3 backends) is fully implemented but only
exposed to the admin status endpoint — no Starlark bridge exists. The `db`
module provides structured data storage via extension-scoped tables, but
there is no equivalent for binary files, no managed filesystem for tools
like `git`, and no mechanism for extensions to declare environment
requirements (pgvector, workspace root, S3) that the kernel validates at
install time.
This blocks every future capability that depends on file handling:
RAG/vector search, LLM bridge, code workspaces, file upload/sharing,
media processing, and document indexing.
---
## Non-Goals
- Replacing the `db` module. Extension-scoped tables (`ext_{pkg}_{table}`)
are the structured data primitive and remain unchanged.
- Building a KV store. Extensions that need key-value semantics declare a
table with `key TEXT, value TEXT` columns — the infrastructure exists.
- Multi-tenant file isolation beyond package scoping. Team/user-level
file ACLs are extension-layer concerns built on top of these primitives.
- Streaming / chunked upload through Starlark. Large file ingestion goes
through HTTP routes; the `files` module handles storage after receipt.
---
## Architecture
Three new primitives plus one extension to the existing `db` module.
### Primitive 1: `files` Starlark Module
Bridges the existing `ObjectStore` into the sandbox. Follows the same
pattern as `db_module.go`: a `Build*Module` factory, permission-gated,
package-scoped key namespacing.
**Permissions:** `files.read`, `files.write`
**Starlark API:**
```python
# Write a file. content is string (UTF-8) or bytes.
# metadata is an optional dict stored alongside (JSON-serialized).
files.put(name, content, content_type="application/octet-stream", metadata={})
# Read a file. Returns dict: {"content": <bytes>, "content_type": "...", "size": N, "metadata": {...}}
# Returns None if not found.
result = files.get(name)
# Read metadata only (no content transfer). Returns dict or None.
meta = files.meta(name)
# List files by prefix. Returns list of dicts: [{"name": "...", "size": N, "content_type": "..."}]
entries = files.list(prefix="", limit=100)
# Delete a file. Idempotent — no error if missing.
files.delete(name)
# Delete all files under a prefix. Use with caution.
files.delete_prefix(prefix)
# Check existence without reading.
exists = files.exists(name)
```
**Key namespacing:**
All keys are automatically prefixed with `ext/{packageID}/`. An extension
calling `files.put("models/v1.bin", data)` writes to the ObjectStore key
`ext/my-extension/models/v1.bin`. Extensions cannot escape their namespace.
Name validation rejects `..`, absolute paths, and control characters —
same sanitization rules as `physicalTable()` in `db_module.go`.
**Implementation notes:**
- New file: `sandbox/files_module.go`
- `FilesModuleConfig` struct mirrors `DBModuleConfig`:
```go
type FilesModuleConfig struct {
PackageID string
CanWrite bool
Store storage.ObjectStore
}
```
- Metadata is stored as a companion JSON object at key
`ext/{packageID}/_meta/{name}`. This avoids schema changes — the
ObjectStore interface is unchanged. The `files` module manages the
companion transparently.
- Size limit per `files.put()` call: 50MB (configurable via
`EXT_FILES_MAX_SIZE`). Enforced in the builtin before calling
`Store.Put()`.
- `files.get()` returns content as `starlark.Bytes` for binary safety.
Starlark's `Bytes` type was added in go.starlark.net v0.0.0-20240725214946
and handles non-UTF-8 content correctly.
- If `ObjectStore` is nil (storage not configured), the module is not
injected — same pattern as `db` module when `r.db == nil`.
**Runner wiring:**
```go
// In buildModulesWithLibCtx, add cases:
case models.ExtPermFilesRead:
if filesLevel < 1 { filesLevel = 1 }
case models.ExtPermFilesWrite:
filesLevel = 2
// After permission loop:
if filesLevel > 0 && r.objectStore != nil {
modules["files"] = BuildFilesModule(ctx, FilesModuleConfig{
PackageID: packageID,
CanWrite: filesLevel == 2,
Store: r.objectStore,
})
}
```
Runner gains `SetObjectStore(s storage.ObjectStore)` setter, called from
`main.go` after `storage.Init()`.
---
### Primitive 2: `workspace` Starlark Module
Managed disk directories for extensions that need a real filesystem —
git repos, compilers, ffmpeg, pandoc, code analysis tools. These tools
cannot operate through put/get blob semantics; they need paths.
**Permission:** `workspace.manage`
**Starlark API:**
```python
# Create a named workspace directory. Returns the absolute path.
# Idempotent — returns existing path if already created.
path = workspace.create(name)
# Get the path for an existing workspace. Returns string or None.
path = workspace.path(name)
# List workspace names for this extension.
names = workspace.list()
# Delete a workspace and all its contents.
workspace.delete(name)
# Disk usage in bytes for a workspace.
size = workspace.usage(name)
```
**Directory layout:**
```
{WORKSPACE_ROOT}/
{packageID}/
{name}/
... (extension-managed contents)
```
`WORKSPACE_ROOT` defaults to `/data/workspaces` (configurable via
`WORKSPACE_ROOT` env var). The kernel creates the package subdirectory
on first `workspace.create()`. Extensions own everything below their
directory — the kernel does not inspect contents.
**Implementation notes:**
- New file: `sandbox/workspace_module.go`
- `WorkspaceModuleConfig`:
```go
type WorkspaceModuleConfig struct {
PackageID string
WorkspaceRoot string
}
```
- Name validation: same rules as table names (`^[a-z][a-z0-9_]{0,62}$`).
No path separators, no `..`, no spaces.
- `workspace.create()` calls `os.MkdirAll` for the scoped path.
- `workspace.delete()` calls `os.RemoveAll` — destructive by design.
Extensions must handle confirmation in their own UX.
- `workspace.usage()` walks the directory tree and sums file sizes.
Bounded by a 10-second context timeout to prevent hangs on huge trees.
- **Quota enforcement** (optional): `WORKSPACE_QUOTA_MB` env var. When
set, `workspace.create()` checks cumulative usage for the package
before creating. Returns error if quota exceeded. Default: unlimited.
- If `WORKSPACE_ROOT` is empty or not writable, the module is not
injected. Extensions that declare `workspace.manage` without the
root configured will have their permission granted but the module
absent — same degradation pattern as `db` when no DB is available.
**Runner wiring:**
```go
case models.ExtPermWorkspaceManage:
if r.workspaceRoot != "" {
modules["workspace"] = BuildWorkspaceModule(ctx, WorkspaceModuleConfig{
PackageID: packageID,
WorkspaceRoot: r.workspaceRoot,
})
}
```
Runner gains `SetWorkspaceRoot(path string)` setter.
**Security considerations:**
- The returned path is an absolute filesystem path. Extensions can pass
this to `http` module calls (e.g., POST a file to an API) or use it
in `db` records as a reference. They cannot execute arbitrary binaries —
the Starlark sandbox has no `os.exec`. Execution requires a sidecar
tier package or an `api_route` handler that shells out server-side.
- For sidecar-tier packages that _can_ execute binaries, the workspace
path is the designated scratch space. The kernel ensures the path is
within the scoped directory via `filepath.Clean` + prefix check.
---
### Primitive 3: Capability Negotiation
Extensions declare environment requirements in their manifest. The kernel
validates these at install time and reports failures with actionable
messages. This is what makes progressive enhancement strategies (e.g.,
three-tier vector search) work.
**Manifest field:**
```json
{
"id": "vector-store",
"capabilities": {
"required": ["files.read", "files.write"],
"optional": ["pgvector"]
},
"permissions": ["db.write", "files.write"],
"db_tables": {
"embeddings": {
"columns": {
"source_id": "text",
"chunk_text": "text",
"embedding": "vector(384)"
},
"indexes": [["source_id"]]
}
}
}
```
`capabilities.required` — install fails if any are missing. Admin gets
a clear error: "Package 'vector-store' requires capability 'pgvector'
which is not available. Run `CREATE EXTENSION vector;` in your PostgreSQL
database to enable it."
`capabilities.optional` — install succeeds regardless. The capability
state is queryable at runtime so extensions can degrade gracefully.
**Runtime query (Starlark):**
Settings module is the natural home since it's always available:
```python
# Returns True/False for a capability name.
has_pgvector = settings.has_capability("pgvector")
```
**Kernel capability registry:**
A simple function in the `handlers` package that probes the environment:
```go
func DetectCapabilities(db *sql.DB, isPostgres bool, workspaceRoot string, objStore storage.ObjectStore) map[string]bool {
caps := make(map[string]bool)
// pgvector: check pg_extension
if isPostgres && db != nil {
var exists bool
row := db.QueryRow("SELECT EXISTS(SELECT 1 FROM pg_extension WHERE extname='vector')")
if row.Scan(&exists) == nil && exists {
caps["pgvector"] = true
}
}
// workspace: check root is writable
if workspaceRoot != "" && storage.IsPathWritable(workspaceRoot) {
caps["workspace"] = true
}
// object storage: check configured and healthy
if objStore != nil {
if err := objStore.Healthy(context.Background()); err == nil {
caps["object_storage"] = true
}
}
// s3: specific backend check
if objStore != nil && objStore.Backend() == "s3" {
caps["s3"] = true
}
// postgres: dialect check
if isPostgres {
caps["postgres"] = true
}
return caps
}
```
Called once at startup, stored on the `Runner` (or a shared config
struct). Re-probed on admin request for the capabilities endpoint.
**Install-time validation:**
In the package install handler (`handlers/extensions.go`), after
`ParseDBTables` and before `SyncManifestPermissions`:
```go
if caps, ok := ParseCapabilities(manifestMap); ok {
missing := CheckRequiredCapabilities(caps.Required, detectedCaps)
if len(missing) > 0 {
// Roll back: delete the just-created package row
h.stores.Packages.Delete(c.Request.Context(), pkg.ID)
c.JSON(422, gin.H{
"error": "missing required capabilities",
"missing": missing,
"help": capabilityHelpText(missing),
})
return
}
}
```
**Admin API:**
```
GET /api/v1/admin/capabilities
→ {"pgvector": true, "workspace": true, "object_storage": true, "s3": false, "postgres": true}
```
Displayed in the Admin UI alongside storage status. Gives operators
visibility into what their deployment supports.
---
### Extension to `db` Module: Vector Column Type
The existing `mapColType()` in `ext_db_schema.go` gains a `vector` type
with progressive enhancement across backends.
**Manifest declaration:**
```json
"db_tables": {
"embeddings": {
"columns": {
"embedding": "vector(384)"
}
}
}
```
**Column type mapping:**
| Manifest type | PG + pgvector | PG without pgvector | SQLite |
|------------------|-------------------------|---------------------|--------------|
| `vector(N)` | `vector(N)` | `JSONB` | `TEXT` |
Implementation in `mapColType`:
```go
case "vector":
// vector or vector(384) — extract dimension if present
dim := extractVectorDim(typStr) // returns "384" or ""
if isPostgres && hasPgVector {
if dim != "" {
return fmt.Sprintf("vector(%s)", dim)
}
return "vector"
}
if isPostgres {
return "JSONB" // store as JSON array, brute-force search
}
return "TEXT" // SQLite: JSON array as text
```
`hasPgVector` is passed through a new field on a `SchemaConfig` struct
(or detected inline — the capability registry result is available at
table creation time).
**New `db` module function — `db.query_similar()`:**
```python
# Find rows with embeddings closest to the query vector.
# Returns list of row dicts with _distance appended.
results = db.query_similar(
table="embeddings",
column="embedding",
vector=[0.1, 0.2, ...], # query vector (list of floats)
limit=10,
filters={"source_id": "doc-123"} # optional WHERE clause
)
```
**Backend dispatch:**
- **PG + pgvector:** Uses `ORDER BY embedding <=> $1 LIMIT $2` with
native vector distance operator.
- **PG without pgvector / SQLite:** Loads candidate rows (respecting
`filters`), deserializes JSON arrays, computes cosine similarity in
Go, sorts, returns top-N. This is the "it works but slowly" fallback —
acceptable for small corpora (<10k rows), documented as such.
Implementation: new function `dbQuerySimilar()` in `db_module.go`,
gated on `db.read` permission (read-only operation). The function
checks `cfg.HasPgVector` to choose the fast or slow path.
`DBModuleConfig` gains:
```go
type DBModuleConfig struct {
PackageID string
CanWrite bool
DB *sql.DB
IsPostgres bool
HasPgVector bool // NEW: enables native vector ops
}
```
Wired from the capability registry at module build time.
---
## New Permission Constants
```go
// In models/models_extension_perm.go:
const (
ExtPermFilesRead = "files.read"
ExtPermFilesWrite = "files.write"
ExtPermWorkspaceManage = "workspace.manage"
)
```
Added to `ValidExtensionPermissions` map.
---
## Schema Changes
**Migration 015 — none required for kernel tables.**
The `files` module uses the existing `ObjectStore` interface — no new
kernel tables. Metadata companions are stored as ObjectStore objects.
The `workspace` module uses the filesystem — no tables.
Capabilities are detected at runtime — no tables.
The `vector` column type is handled by extension DDL generation in
`ext_db_schema.go` — no kernel migration.
The new permission constants are code-only changes.
---
## New Files
| File | Purpose |
|------|---------|
| `sandbox/files_module.go` | `files` Starlark module |
| `sandbox/files_module_test.go` | Tests using mock ObjectStore |
| `sandbox/workspace_module.go` | `workspace` Starlark module |
| `sandbox/workspace_module_test.go` | Tests using temp directory |
| `handlers/capabilities.go` | `DetectCapabilities()`, `ParseCapabilities()`, admin endpoint |
| `handlers/capabilities_test.go` | Tests for capability detection and validation |
## Modified Files
| File | Change |
|------|--------|
| `models/models_extension_perm.go` | Add `files.read`, `files.write`, `workspace.manage` |
| `sandbox/runner.go` | Add `SetObjectStore()`, `SetWorkspaceRoot()`, wire new modules |
| `sandbox/db_module.go` | Add `dbQuerySimilar()`, `HasPgVector` field |
| `handlers/ext_db_schema.go` | Extend `mapColType()` for `vector(N)`, pass `hasPgVector` |
| `handlers/extensions.go` | Add capability validation in install handler |
| `sandbox/settings_module.go` | Add `settings.has_capability()` |
| `server/main.go` | Call `DetectCapabilities()`, wire to runner |
| `docs/PACKAGE-FORMAT.md` | Document `capabilities` manifest field |
| `docs/EXTENSION-GUIDE.md` | Document new modules and vector column type |
---
## Changeset Plan
**CS1 — `files` module + permissions**
New files: `files_module.go`, `files_module_test.go`.
Modified: `models_extension_perm.go`, `runner.go`, `main.go`.
Scope: ObjectStore bridge, permission constants, runner wiring.
CI-green independently — no schema changes, no existing behavior affected.
**CS2 — `workspace` module**
New files: `workspace_module.go`, `workspace_module_test.go`.
Modified: `models_extension_perm.go`, `runner.go`, `main.go`.
Scope: Managed disk directories, quota enforcement.
CI-green independently.
**CS3 — Capability negotiation**
New files: `capabilities.go`, `capabilities_test.go`.
Modified: `extensions.go` (install validation), `settings_module.go`
(`has_capability`), `main.go`.
Scope: Detection, install-time validation, runtime query, admin endpoint.
CI-green independently.
**CS4 — Vector column type + `db.query_similar()`**
Modified: `ext_db_schema.go`, `db_module.go`, `db_module_test.go`.
Scope: Column type mapping, similarity query with dual-path dispatch.
Depends on CS3 (needs `HasPgVector` from capability registry).
---
## Configuration Summary
| Env Var | Default | Purpose |
|---------|---------|---------|
| `WORKSPACE_ROOT` | `/data/workspaces` | Root directory for workspace module |
| `WORKSPACE_QUOTA_MB` | `0` (unlimited) | Per-extension disk quota |
| `EXT_FILES_MAX_SIZE` | `52428800` (50MB) | Max single file size via `files.put()` |
---
## Example: Vector Store Extension
A `vector-store` library extension consuming all four primitives:
**manifest.json:**
```json
{
"id": "vector-store",
"title": "Vector Store",
"type": "library",
"tier": "starlark",
"version": "0.1.0",
"permissions": ["db.write", "files.read", "connections.read"],
"capabilities": {
"required": [],
"optional": ["pgvector"]
},
"db_tables": {
"documents": {
"columns": {
"source": "text",
"chunk_text": "text",
"page": "int",
"metadata": "text"
},
"indexes": [["source"]]
},
"embeddings": {
"columns": {
"document_id": "text",
"embedding": "vector(384)",
"chunk_index": "int"
},
"indexes": [["document_id"]]
}
},
"exports": ["ingest", "search", "search_clustered"]
}
```
**script.star:**
```python
def ingest(source_name, chunks):
"""Store document chunks and generate embeddings."""
for i, chunk in enumerate(chunks):
row = db.insert("documents", {
"source": source_name,
"chunk_text": chunk["text"],
"page": chunk.get("page", 0),
"metadata": json.encode(chunk.get("metadata", {})),
})
# Embedding generation delegated to caller (llm-bridge)
# Caller passes pre-computed vectors
if "embedding" in chunk:
db.insert("embeddings", {
"document_id": row["id"],
"embedding": json.encode(chunk["embedding"]),
"chunk_index": i,
})
def search(query_vector, limit=10, source_filter=None):
"""Semantic similarity search — dispatches to native or brute-force."""
filters = {}
if source_filter:
filters["source"] = source_filter
# query_similar handles pgvector vs fallback transparently
results = db.query_similar(
table="embeddings",
column="embedding",
vector=query_vector,
limit=limit,
)
# Hydrate with document text
enriched = []
for r in results:
docs = db.query("documents", filters={"id": r["document_id"]}, limit=1)
if docs:
enriched.append({
"text": docs[0]["chunk_text"],
"source": docs[0]["source"],
"page": docs[0]["page"],
"distance": r["_distance"],
})
return enriched
def search_clustered(embeddings_list, k):
"""K-means clustering for query-free thematic selection.
Returns k representative chunks. Clustering runs in the
extension — the kernel provides the data, not the algorithm."""
# This would be implemented by a consumer extension with
# http access to call a clustering API, or by a sidecar
# that runs sklearn. The vector-store library provides
# the data access pattern; clustering logic lives elsewhere.
pass
```
---
## Future Considerations
- **Content-addressed deduplication.** If multiple extensions store the
same PDF, the ObjectStore holds duplicate bytes. A SHA256-keyed blob
layer with refcounting would deduplicate transparently. Deferred —
adds complexity without clear need pre-1.0.
- **Extension-to-extension file sharing.** The current design isolates
file namespaces per extension. A `files.grant(name, target_pkg_id)`
primitive could enable controlled sharing. Deferred — composition
through API routes is sufficient initially.
- **Streaming upload for large files.** Starlark `files.put()` buffers
in memory. For files >50MB, extensions should use HTTP `api_routes`
with Go handlers that stream directly to ObjectStore. The `files`
module is for extension-internal storage, not user-facing upload.
- **Workspace snapshots.** `workspace.snapshot(name)` → creates a
tarball in the `files` store. Useful for backup/reproducibility.
Deferred — extensions can implement this themselves.
- **Rate limiting / quota on `db.query_similar()` fallback.** The
brute-force path loads all candidate rows into Go memory. For large
tables this is dangerous. A row-count guard (e.g., refuse if >50k
candidates) with a clear error message pointing to pgvector is the
right safety valve.