This repository has been archived on 2026-04-03. You can view files and clone it. You cannot open issues or pull requests or push a commit.
Files
core/docs/DESIGN-batch-exec.md
Jeffrey Smith 8957e99610 Feat v0.7.12 batch.exec concurrent execution primitive
Add batch.exec(callables, timeout=10) — a general-purpose concurrent
execution primitive that runs arbitrary Starlark callables in parallel,
each in its own thread with independent step budget. Enables extensions
to fan out library function calls (frozen exports from lib.require)
without decomposing them back into raw http.post parameters.

New permission: batch.exec. Max 8 callables, timeout 1-30s.
Nesting prohibited via atomic flag. 12 new tests, all pass with -race.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 23:52:01 +00:00

467 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# DESIGN — Concurrent Execution Primitive (`batch.exec`)
**Version:** v0.7.12
**Status:** Implemented
**Author:** Jeff / Claude session 2026-04-02
---
## Problem
Starlark is single-threaded by design (`go.starlark.net` enforces one
thread per execution). Extensions that need to fan out — calling multiple
external APIs, invoking several library functions, or performing independent
I/O operations — must do so sequentially. For two `http.post()` calls
taking 200ms each, the extension blocks for 400ms regardless of whether
the calls are independent.
The v0.7.10 `http.batch()` primitive solves the narrow case of parallel
HTTP dispatch. But it doesn't help when the work is wrapped in library
functions. If a `jira-client` library exposes `create_issue()` and a
`confluence-client` library exposes `create_page()`, the extension author
must either:
1. Call them sequentially (slow), or
2. Decompose the library calls back into raw `http.post()` parameters
to use `http.batch()` (defeats the purpose of having libraries).
The platform needs a general-purpose concurrent execution primitive that
works with arbitrary Starlark callables — including library exports.
---
## Non-Goals
- **Shared mutable state between branches.** Each concurrent branch is
fully isolated. No channels, no mutexes, no shared dicts. If branches
need to coordinate, they don't belong in `batch.exec`.
- **Unlimited concurrency.** A hard cap prevents extensions from spawning
unbounded goroutines. This is a fan-out primitive, not a thread pool.
- **Automatic retry or circuit breaking.** Error handling is the caller's
responsibility. The kernel reports per-branch results and errors.
- **Nested `batch.exec()`.** A callable inside `batch.exec` cannot itself
call `batch.exec`. This prevents exponential goroutine growth and keeps
the concurrency model flat.
---
## Key Insight: Frozen Libraries Are Thread-Safe
The reason this works without exotic machinery is `lib.require()`.
When a library is loaded via `lib.require()`, its exports are wrapped in
a `starlarkstruct.FromStringDict()` — which produces a **frozen** struct.
Frozen Starlark values are immutable and safe to read from any number of
goroutines concurrently. This is a property of `go.starlark.net`, not
something we enforce.
The only mutable state in a Starlark execution is:
1. **The `starlark.Thread` itself** — step counter, print buffer, cancel
channel. Each branch gets its own thread.
2. **Module instances**`db`, `http`, `settings`, etc. contain
configuration and hold references to shared Go objects (`*sql.DB`,
`http.Client`). Each branch gets fresh module instances, but the
underlying Go resources (`*sql.DB` connection pool, etc.) are already
designed for concurrent access.
3. **Local variables** — thread-local by definition in Starlark.
So the construction is: one new `starlark.Thread` + one new module set
per branch, with frozen library structs shared read-only across all
branches. This is exactly what `triggers/schedule.go` already does for
scheduled task execution — `buildRestrictedModules` creates a fresh
module set for each cron fire. `batch.exec` generalizes that pattern.
---
## API
```python
results, errors = batch.exec([
lambda: jira.create_issue(issue_data),
lambda: confluence.create_page(page_data),
lambda: slack.post_message(channel, msg),
])
# results[0] = return value of jira.create_issue(), or None on error
# errors[0] = None on success, or error string on failure
# All three ran concurrently.
```
### Signature
```
batch.exec(callables, timeout=10) → (results: list, errors: list)
```
**Parameters:**
| Param | Type | Description |
|-------|------|-------------|
| `callables` | `list[callable]` | Starlark callables (lambdas, named functions, bound methods). Max length: 8. |
| `timeout` | `int` (optional) | Per-branch timeout in seconds. Default 10. Max 30. Inherits parent context deadline if shorter. |
**Returns:** A 2-tuple of equal-length lists.
- `results[i]` — the return value of `callables[i]`, or `None` if it
errored.
- `errors[i]``None` if `callables[i]` succeeded, or a string error
message if it failed (timeout, step limit, runtime error).
**Errors (whole-call):**
- `callables` is empty → error
- `callables` length > 8 → error
- Any element is not callable → error
- Permission `batch.exec` not granted → error
### Permission
New extension permission: `batch.exec`. Declared in manifest:
```json
{
"permissions": ["batch.exec"]
}
```
This is a separate permission because concurrent execution has resource
implications (goroutines, module construction overhead). Extensions that
don't need it shouldn't pay for it. The permission doesn't imply any
other permissions — the branch inherits whatever modules the calling
package already has.
---
## Execution Model
```
batch.exec([fn_a, fn_b, fn_c])
├─── goroutine 1: Thread₁ + Modules₁ → fn_a() → result[0]
├─── goroutine 2: Thread₂ + Modules₂ → fn_b() → result[1]
└─── goroutine 3: Thread₃ + Modules₃ → fn_c() → result[2]
sync.WaitGroup.Wait()
return (results, errors)
```
### Per-Branch Construction
For each callable in the input list, the kernel:
1. Creates a new `sandbox.Sandbox` with the same `Config` as the parent
(same `MaxSteps` limit — each branch gets its own step budget, not a
shared one).
2. Calls `runner.buildModulesWithLibCtx()` with the **same** `packageID`,
`manifest`, and `RunContext` as the parent invocation. This produces
a fresh module set — new `DBModuleConfig`, new `HTTPModuleConfig`, etc.
— pointing at the same underlying Go resources (`*sql.DB`, etc.).
3. The `libContext` is **shared** (read path only — cached frozen exports).
Library exports are immutable. The `loading` map (cycle detection) is
not relevant because libraries are already loaded before `batch.exec`
runs. If a branch triggers a new `lib.require()`, it would need its
own `libContext` — see Open Questions.
4. Creates a new `starlark.Thread` with the branch's print handler,
step limit, and context-based cancellation.
5. Calls `starlark.Call(thread, callable, nil, nil)` — the callable
is a zero-arg lambda that closes over its arguments.
### Context & Cancellation
Each branch gets a child context derived from the parent with the
per-branch timeout applied:
```go
branchCtx, cancel := context.WithTimeout(parentCtx, branchTimeout)
defer cancel()
```
If the parent context is cancelled (e.g., HTTP request timeout), all
branches are cancelled. If one branch exceeds its timeout, only that
branch is cancelled — others continue.
### Goroutine Cap
Hard limit: **8 concurrent branches.** This is enforced at the API
boundary (list length check), not via a semaphore. Rationale:
- 8 covers the real-world fan-out patterns (2-5 API calls, small batch
operations). Nobody needs 50 concurrent Starlark branches.
- Each branch allocates a `starlark.Thread` + module instances. At 8
branches, overhead is bounded at ~8KB per thread + module construction
time (~50μs per module set).
- No semaphore means no queuing surprises. You get 8, period.
---
## Implementation
### New File: `sandbox/batch_module.go`
```go
// BuildBatchModule creates the "batch" module.
// Requires the Runner reference for per-branch module construction.
func BuildBatchModule(
ctx context.Context,
runner *Runner,
packageID string,
manifest map[string]any,
rc *RunContext,
lc *libContext,
) *starlarkstruct.Module
```
The module holds a reference to the `Runner` — same pattern as
`BuildLibModule`. It needs the runner to call `buildModulesWithLibCtx`
for each branch.
### Core Implementation Sketch
```go
func batchExec(ctx context.Context, runner *Runner, packageID string,
manifest map[string]any, rc *RunContext, parentLC *libContext,
) func(*starlark.Thread, *starlark.Builtin, starlark.Tuple, []starlark.Tuple) (starlark.Value, error) {
return func(thread *starlark.Thread, b *starlark.Builtin,
args starlark.Tuple, kwargs []starlark.Tuple,
) (starlark.Value, error) {
var callableList *starlark.List
var timeout int = 10
if err := starlark.UnpackArgs(b.Name(), args, kwargs,
"callables", &callableList,
"timeout?", &timeout,
); err != nil {
return nil, err
}
n := callableList.Len()
if n == 0 {
return nil, fmt.Errorf("batch.exec: callables list is empty")
}
if n > 8 {
return nil, fmt.Errorf("batch.exec: max 8 callables, got %d", n)
}
if timeout < 1 || timeout > 30 {
timeout = 10
}
// Validate all elements are callable.
callables := make([]starlark.Callable, n)
for i := 0; i < n; i++ {
c, ok := callableList.Index(i).(starlark.Callable)
if !ok {
return nil, fmt.Errorf("batch.exec: element %d is %s, not callable",
i, callableList.Index(i).Type())
}
callables[i] = c
}
// Execute concurrently.
type branchResult struct {
index int
value starlark.Value
err error
}
results := make([]starlark.Value, n)
errors := make([]starlark.Value, n)
var wg sync.WaitGroup
for i, callable := range callables {
wg.Add(1)
go func(idx int, fn starlark.Callable) {
defer wg.Done()
// Per-branch context with timeout.
branchCtx, cancel := context.WithTimeout(ctx,
time.Duration(timeout)*time.Second)
defer cancel()
// Fresh module set for this branch.
modules, err := runner.buildModulesWithLibCtx(
branchCtx, packageID, manifest, rc, parentLC)
if err != nil {
results[idx] = starlark.None
errors[idx] = starlark.String(err.Error())
return
}
// Fresh sandbox + thread.
sb := New(DefaultConfig())
val, _, callErr := sb.Call(branchCtx, fn, nil, nil)
if callErr != nil {
results[idx] = starlark.None
errors[idx] = starlark.String(callErr.Error())
} else {
results[idx] = val
errors[idx] = starlark.None
}
}(i, callable)
}
wg.Wait()
return starlark.Tuple{
starlark.NewList(results),
starlark.NewList(errors),
}, nil
}
}
```
### Runner Wiring
In `buildModulesWithLibCtx`, add the permission case:
```go
case models.ExtPermBatchExec:
// Deferred — wired after module map is complete (needs runner ref).
hasBatchExec = true
```
After the module map is assembled:
```go
if hasBatchExec {
modules["batch"] = BuildBatchModule(ctx, r, packageID, manifest, rc, lc)
}
```
### Permission Constant
In `models/permissions.go`:
```go
ExtPermBatchExec = "batch.exec"
```
Add to `AllExtensionPermissions` slice.
---
## Callable Closure Semantics
The callables passed to `batch.exec` are typically lambdas that close
over variables from the calling scope:
```python
issue_data = {"summary": "Review Q3 report"}
page_data = {"title": "Q3 Report", "body": content}
results, errors = batch.exec([
lambda: jira.create_issue(issue_data),
lambda: confluence.create_page(page_data),
])
```
The closed-over values (`issue_data`, `page_data`, `jira`, `confluence`)
are references to Starlark values in the calling thread's scope. Two
safety properties make this work:
1. **Library exports (`jira`, `confluence`) are frozen.** They were
returned by `lib.require()` as `starlarkstruct.FromStringDict()`
deeply immutable. Safe to read from any goroutine.
2. **Dict/list arguments may be mutable**, but Starlark's execution
model means the calling thread is **blocked** waiting for
`batch.exec` to return. No concurrent mutation is possible because
the caller can't execute while the branches are running.
This is the same safety model as Go's `sync.WaitGroup` pattern: the
goroutine that calls `wg.Wait()` cannot proceed until all goroutines
complete, so values passed to goroutines before `wg.Add` are safe to
read without locks.
---
## Open Questions
### 1. `lib.require()` Inside Branches
If a callable triggers a `lib.require()` that hasn't been cached yet,
the shared `libContext.cache` would be written from a goroutine. Options:
**A. Prohibit: branches cannot call `lib.require()`.** The branch gets
a nil `libContext`, so `lib` module is unavailable inside `batch.exec`.
Libraries must be loaded before the batch call. Simplest, safest.
**B. Per-branch `libContext` with shared read cache.** Each branch gets
its own `libContext` whose `cache` is pre-populated from the parent's
cache (snapshot). New loads go into the branch's local cache only.
Slightly wasteful if two branches load the same library (loaded twice),
but safe.
**C. Mutex-protected shared `libContext`.** Add a `sync.RWMutex` to
`libContext`. Reads use `RLock`, writes use `Lock`. Minimal overhead,
but makes `libContext` aware of concurrency — violates its current
assumptions.
**Recommendation: Option A for v0.7.11, Option B as follow-up if needed.**
In practice, extensions call `lib.require()` at module scope (top of
script), not inside request handlers. The lambdas passed to `batch.exec`
call methods on already-loaded library structs. Option A covers all
real-world patterns.
### 2. Print Output
Each branch has its own print buffer (the `output strings.Builder` in
`Sandbox.Call`). Options:
**A. Discard.** Branch print output is lost. Simple, avoids interleaving.
**B. Collect per-branch.** Return a third list: `(results, errors, outputs)`.
Useful for debugging but clutters the API.
**C. Merge into parent.** Append all branch output to the parent thread's
print buffer, prefixed with branch index. Requires passing the parent's
`outputMu` and `output` builder — invasive.
**Recommendation: Option A for v0.7.11.** `print()` in Starlark is a
debugging tool, not a production logging facility. Branch callables that
need to report status should return structured data. If debugging demand
emerges, Option B is a backward-compatible addition.
### 3. Step Limit Scope
Each branch gets its own `MaxSteps` budget (default 1M). Should the
total across all branches be capped?
**No.** The per-branch cap is sufficient. 8 branches × 1M steps = 8M
total, which completes in under a second on any modern hardware. The
wall-clock timeout (per-branch, max 30s) is the real resource guard.
Adding a cross-branch step budget creates coupling between independent
execution paths — branch A's step count shouldn't affect branch B's
ability to complete.
---
## Testing
| Test | Description |
|------|-------------|
| Parallel ordering | 3 callables with different sleep durations. Results in input order, not completion order. |
| Partial failure | 3 callables, middle one raises error. results = [val, None, val], errors = [None, "err msg", None]. |
| Timeout per-branch | One callable sleeps beyond timeout. Others succeed. Timed-out branch returns error. |
| Parent cancellation | Cancel parent context mid-execution. All branches cancelled. |
| Cap enforcement | List of 9 callables → immediate error, nothing executed. |
| Empty list | `batch.exec([])` → error. |
| Non-callable element | `batch.exec([1, 2])` → error, nothing executed. |
| Permission gating | Package without `batch.exec` permission → module not available. |
| Frozen library sharing | Two branches call same frozen library function concurrently. No race. |
| Nested batch.exec | Callable inside batch.exec attempts batch.exec → error (module not injected in branch). |
| db module isolation | Two branches insert into same table concurrently. Both succeed, no corruption. |
| http module isolation | Two branches make HTTP calls with different headers. No cross-contamination. |
---
## Migration
No schema changes. No new tables. No new migrations.
New permission constant `batch.exec` added to `models/permissions.go`.
Extensions must declare the permission in their manifest to access the
`batch` module.