Define generalized workload SDK contract

This commit is contained in:
Emil
2026-08-01 16:30:25 +03:00
parent 0a759a3f01
commit 11e9333033
3 changed files with 698 additions and 0 deletions
+2
View File
@@ -498,6 +498,8 @@ reducer behavior, worker allowlist/capabilities, input and output artifacts,
UI/API parameters, reproducible execution environment, result verification,
golden cross-worker fixtures, and applicable resource limits. The future public
contract is described in [`docs/scimesh-sdk-roadmap.md`](docs/scimesh-sdk-roadmap.md).
Its normative future interfaces and execution semantics are defined in the
design-draft [`docs/scimesh-sdk-contract.md`](docs/scimesh-sdk-contract.md).
Every workload declaration must classify its task decomposition and input/output
artifact shapes, determinism, reduction semantics, verifier mode, supported
+533
View File
@@ -0,0 +1,533 @@
# SciMesh Workload SDK contract
**Status:** design draft `0.1`; not implemented. Normative words **MUST**,
**MUST NOT**, **SHOULD**, and **MAY** describe the intended future contract,
not capabilities of the current release.
This document defines the compatibility boundary for approved SciMesh workload
packages. The sequencing and unresolved product decisions remain in
[`scimesh-sdk-roadmap.md`](scimesh-sdk-roadmap.md). The currently implemented
protocol is still [`ctx-07-distributed-workload-protocol.md`](ctx-07-distributed-workload-protocol.md).
## 1. Scope and invariants
The SDK must express current molecular workloads and future batch, iterative,
streaming, optimization, simulation, ML, image, engineering, and accelerator
workloads through installed, allowlisted packages. "Universal" means that a
workload can declare its dataflow, resources, execution semantics, and result
validation; it does not mean users may submit arbitrary executable code.
Every implementation MUST preserve these invariants:
1. The coordinator owns Jobs, workflow state, Tasks, Attempts, Leases, resource
assignments, and durable Artifact metadata.
2. A Worker Agent executes only an installed package and entry point whose
digest is enabled by an administrator.
3. Scientific code never receives PostgreSQL credentials, worker identity
credentials, or unrestricted coordinator credentials.
4. Every input and output crosses a typed artifact port. Worker-local paths are
attempt-scoped implementation details and never become durable references.
5. A Task reports success only after all declared outputs are durable and its
verifier policy has produced an admissible decision.
6. Resource allocation is explicit and lease-fenced. A worker MUST NOT execute
a Task without reserving its complete declared resource set.
7. Unknown fields, unsupported versions, undeclared outputs, non-finite numeric
values, and limit violations fail closed.
8. No verifier, reducer, or trust policy may silently downgrade to a weaker
mode.
## 2. Version and identity model
The following versions are independent and MUST be recorded in each Job and
output provenance manifest:
| Identifier | Meaning |
| --- | --- |
| `sdk_api_version` | Python authoring API compatibility. |
| `protocol_version` | Coordinator/Worker wire contract. |
| `workload.name` + `workload.version` | Scientific behavior and planner contract. |
| `manifest_schema_version` | Shape of the workload manifest. |
| `workflow_schema_version` | Shape of workflow/stage/task plans. |
| `artifact_schema.name` + `version` | Logical data format. |
| `verifier.name` + `version` | Result-acceptance semantics. |
| `environment.digest` | Exact executable environment or image. |
Versions use explicit compatibility ranges. The coordinator enables a workload
only when there is a non-empty intersection among coordinator protocol, Worker
runtime, SDK API, workload package, and verifier versions. A missing or unknown
version is never interpreted as "latest". Jobs remain pinned to the resolved
versions even after an administrator upgrades the installed package.
### 2.1 Feature negotiation and conformance profiles
Universality does not require every deployment to enable every execution mode.
The protocol negotiates explicit feature IDs; a workload declares required and
optional features, and coordinator, Worker, and verifier runtimes advertise
supported versions. Planning fails before Job creation when a required feature
is absent.
| Profile | Required feature set |
| --- | --- |
| `core-batch-v1` | Typed parameters/artifacts, static map/reduce DAG, subprocess runner, exact verifier, CPU/memory/scratch reservations. |
| `dynamic-v1` | Transactional expansion manifests and bounded loop controllers. |
| `stream-v1` | Partition offsets, windows, checkpoints, backpressure, and declared delivery guarantees. |
| `accelerator-v1` | Generic device inventory, fenced allocation, visibility/isolation, and device failure codes. |
| `gang-v1` | Atomic multi-Agent reservation, rendezvous, group lease, and fail-all semantics. |
| `side-effect-v1` | Idempotency/audit/compensation and scoped external credentials. |
Example feature IDs include `artifact-collections@1`, `dynamic-expansion@1`,
`bounded-loops@1`, `stream-checkpoints@1`, `gpu-exclusive@1`, `gpu-mig@1`,
`gang-leases@1`, and `numeric-verifier@1`. Optional features may select a
manifest-declared fallback such as CPU execution; they MUST NOT change output
schema or verification semantics unless the fallback is a separately versioned
workflow variant.
## 3. Workload package and manifest
An SDK workload is an installed Python distribution containing:
- one or more versioned workload manifests;
- explicit Python entry points for planner, runners, reducers, and verifiers;
- artifact and parameter schemas;
- pinned environment/image metadata;
- golden fixtures and conformance tests;
- provenance and license metadata.
Discovery MUST use configured package entry points and an administrator
allowlist. It MUST NOT import modules from job parameters, uploaded archives, or
user-provided filesystem paths.
Illustrative manifest shape:
```yaml
manifest_schema_version: 1
sdk_api: ">=1.0,<2.0"
protocol: ">=2,<3"
workload:
name: descriptor-batch
version: 1.0.0
description: Pinned RDKit 2D descriptors
package:
distribution: scimesh-descriptors
digest: sha256:...
signature: cosign-or-project-signature-reference
environment:
kind: oci
digest: sha256:...
parameters_schema: schemas/parameters-v1.json
workflow: workflows/default-v1.yaml
inputs:
molecules: {schema: molecule-table@1, cardinality: one}
outputs:
descriptors: {schema: descriptor-table@1, cardinality: one}
determinism: byte_exact
trust_modes: [trusted, verified, untrusted_quorum]
verifier: {name: exact-artifact, version: 1}
limits:
max_input_bytes: 10737418240
max_tasks: 10000
max_output_bytes: 10737418240
capabilities: [descriptor-batch]
```
The manifest MUST declare canonical hyphenated workload names, parameter
schema, external ports, workflow, determinism, supported trust modes, verifier,
resource bounds, output-growth bounds, environment, and capabilities.
## 4. Workflow model
### 4.1 Workflow graph
`WorkflowSpec` defines versioned stages and artifact edges. Its persisted form
is a DAG. Iteration is represented by a bounded controller that materializes a
new DAG segment for each iteration; persisted task dependencies never contain a
cycle.
```yaml
workflow_schema_version: 1
id: default
inputs: [molecules]
stages:
- id: partition
kind: plan
runner: descriptor.partition:v1
- id: calculate
kind: map
needs: [partition]
runner: descriptor.calculate:v1
- id: combine
kind: reduce
needs: [calculate]
reducer: descriptor.combine:v1
- id: verify
kind: verify
needs: [combine]
outputs: [descriptors]
failure_policy: fail_fast
```
Stage IDs are stable within a workflow version. Valid stage kinds are initially
`plan`, `map`, `reduce`, `verify`, `loop-controller`, `stream`, `service`, and
`side-effect`. Runtimes MAY add kinds only through a new workflow schema
version.
### 4.2 Stage contract
A `StageSpec` MUST declare:
- stable ID, kind, entry-point identity, and dependencies;
- named input/output ports and their artifact schemas/cardinality;
- parameter projection from immutable Job parameters;
- resource and execution profiles;
- retry, timeout, checkpoint, cancellation, and failure policies;
- fan-out/fan-in bounds and ordering semantics;
- verifier and trust requirements where stage outputs affect acceptance;
- cacheability, side effects, and network/secrets policy.
Dynamic fan-out requires an accepted `ExpansionManifest` containing stable child
keys, bounded child count, TaskSpecs, artifact bindings, and a digest. The
coordinator validates and persists the entire expansion transactionally. A
retry producing a different expansion digest is a conflict, not a replacement.
### 4.3 Bounded iteration
`LoopSpec` supports training epochs, optimization, adaptive sampling, MD
segments, and convergence algorithms:
```yaml
loop:
state_schema: optimizer-state@1
max_iterations: 100
max_wall_seconds: 86400
body_workflow: optimize-step@2
continue_when: verifier-entry-point-reference
checkpoint_every: 5
on_limit: fail # fail | accept-best | return-inconclusive
```
The loop controller is trusted orchestration code from the workload package.
It consumes immutable prior-state and evaluation artifacts and emits an
immutable next-iteration expansion. It MUST NOT mutate completed Tasks or reuse
an Attempt directory. Termination is bounded by iterations, wall time, cost,
and output growth. The condition and best-result selection are versioned and
auditable.
### 4.4 Streaming profile
Streaming is an explicit profile rather than an indefinitely running batch
Task. A `StreamSpec` declares source identity, partitioning, offset/checkpoint
schema, event-time or processing-time windows, watermark behavior, backpressure
limit, delivery guarantee, idle/terminal condition, and output compaction.
Supported guarantees are `at_least_once` initially and, only where the source
and sink support transactional offsets, `exactly_once`. A checkpoint commits
source offsets only after corresponding output artifacts are durable. Stream
processors MUST be restartable from a sealed checkpoint artifact. An unbounded
stream produces versioned window/result artifacts and does not hold one Task
lease forever.
### 4.5 Side effects and human interaction
Stages controlling instruments or writing external systems are trusted-only.
They require an idempotency key, declared target, credential scope, audit event,
timeout, and compensation/manual-recovery policy. Retries are disabled unless
the stage proves idempotency. Human approval is modeled as a coordinator state
transition with an authenticated decision record, never as a worker waiting
indefinitely while holding resources.
## 5. Task and artifact contracts
`TaskSpec` is the concrete unit leased to a Worker Agent:
```yaml
task_schema_version: 1
task_key: calculate/shard-000042
stage_id: calculate
parameters: {...validated JSON...}
inputs:
molecules:
collection: ordered
artifacts: [{artifact_id: uuid, sha256: ..., schema: molecule-table@1}]
expected_outputs:
descriptors: {schema: descriptor-table@1, cardinality: one}
resources: {profile: cpu-medium@1}
execution: {profile: python-process@1}
```
`task_key` is deterministic within a workflow expansion. The coordinator adds
its durable Task ID, Attempt number, Lease, and generated download/upload URLs.
Plans contain artifact IDs and checksums, never external credentials or local
paths.
### 5.1 Artifact schemas and collections
Each port references an `ArtifactSchema` defining logical type, schema version,
media type, encoding, cardinality, maximum bytes/records/dimensions, streaming
support, canonicalizer, validation entry point, and privacy/retention class.
Collections are `single`, `ordered`, `keyed`, or `set`:
- `ordered` preserves declared order and is included in the collection digest;
- `keyed` requires unique canonical string keys;
- `set` canonicalizes by artifact identity and forbids duplicates;
- nested collections require explicit schema permission and depth limits.
Protocol v1 compatibility uses one immutable composite-manifest artifact to
represent a collection. A later protocol may persist collection edges directly.
### 5.2 Output and provenance manifest
Before completion, a runner uploads an `OutputManifest` listing every declared
output artifact, checksum, schema, size, record/dimension summary, metrics, and
provenance. Provenance includes resolved versions, package/environment digest,
Worker runtime, allocated resource IDs, parameters digest, input collection
digest, timestamps, random seed where applicable, and checkpoint lineage.
Unexpected ports, missing required outputs, extra artifacts, schema failures,
or limit violations reject the Attempt. Logs and checkpoints are separate
artifact kinds and never satisfy scientific output ports.
### 5.3 Locality and cache
Artifacts remain coordinator-owned even when cached. Worker Agents MAY maintain
a content-addressed read cache verified by checksum. Cache entries carry size,
last-use, schema, environment sensitivity, and retention class; eviction never
deletes the durable coordinator copy.
Task matching MAY score data locality after eligibility and fairness checks.
It MUST NOT weaken trust, resource, lease, or ownership constraints. Large input
staging occurs before execution timeout starts, with a bounded staging lease.
Cache hits are verified before use; private artifacts are isolated by tenant or
encrypted policy.
## 6. Resource and execution contracts
### 6.1 Resource inventory and requests
A Worker Agent advertises versioned, periodically refreshed inventory:
```yaml
agent:
cpu:
logical_cores: 32
allocatable_cores: 28
architecture: x86_64
memory_mb: 131072
scratch_mb: 1000000
accelerators:
- kind: gpu
vendor: nvidia
device_uuid: GPU-...
model: A100
memory_mb: 81920
compute_capability: "8.0"
partitioning: [exclusive, mig]
topology_group: nvlink-0
runtime:
os: linux
container_runtime: ...
driver_versions: {...}
environment_digests: [sha256:...]
```
A `ResourceRequirements` request distinguishes minimums, preferred values, and
hard constraints. Core fields are CPU cores, memory, scratch, accelerator count
and kind, device memory, architecture/capability, exclusivity, topology,
network/interconnect class, environment digest, estimated input/output bytes,
and maximum duration.
Allocation is atomic and lease-fenced. A Task cannot start until the Agent has
confirmed the reservation token. Resources are released only after the process
group exits and attempt cleanup completes. Device IDs and secrets are not part
of scientific parameters.
### 6.2 CPU concurrency
One machine runs one Worker Agent with `max_concurrency` execution slots. Each
Task separately requests `cpu_cores`; the sum of reservations cannot exceed
allocatable capacity. `ExecutionProfile` declares:
- `process_model`: `single`, `process_pool`, `thread_pool`, or `external_runtime`;
- maximum worker processes and threads per process;
- OpenMP/BLAS/native-library thread limits;
- affinity/NUMA preference when required;
- whether nested parallelism is prohibited (default) or explicitly bounded.
CPU-bound Python SHOULD use isolated processes. Threads remain valid for I/O or
native extensions that release the GIL. Independent task heartbeat supervisors
remain outside scientific subprocesses. Draining stops claims first, maintains
active leases, then checkpoints/cancels at the declared deadline.
### 6.3 GPU and accelerator allocation
GPU requests can specify `exclusive_device`, `fractional`, or `partition`
(including MIG-like partitions); runtimes MUST advertise which modes they can
enforce. Requests may require multiple devices in one topology group. The
coordinator performs generic eligibility and gang selection; the Agent owns
device isolation and sets backend-specific visibility variables.
The Agent MUST fence allocation by device UUID/partition ID, validate driver,
runtime and environment compatibility, prevent incompatible sharing, monitor
device health and memory, and terminate the whole process group on lease loss.
OOM, device reset, ECC failure, and unavailable-device errors have distinct
sanitized codes and workload-declared retry policies. Preemptible GPU Tasks need
an explicit compatible checkpoint contract.
The workload owns CPU/GPU implementation, batching, mixed precision, seeds,
algorithm determinism, memory strategy, and scientific parity tests. The
coordinator contains no CUDA calls or domain formulas. A GPU result follows the
same output schema and verifier semantics as any CPU result.
### 6.4 Multi-node and gang execution
`GangSpec` expresses MPI, multi-node training, and tightly coupled simulations:
```yaml
gang:
replicas: 4
per_replica_resources: {cpu_cores: 8, gpu_count: 1, memory_mb: 32768}
topology: {same_fabric: true, min_bandwidth_class: infiniband}
rendezvous: worker-agent-managed
failure_mode: fail_all
```
The coordinator atomically reserves all replicas or none, then issues one gang
lease and per-replica fenced assignments. Worker Agents establish a scoped
rendezvous channel without exposing general coordinator credentials. A failed,
expired, or cancelled replica invalidates the gang according to `failure_mode`;
partial success cannot complete the stage. Gang retries use a new Attempt and
new rendezvous credentials.
## 7. Verification and trust
Verifier decisions are `accepted`, `rejected`, or `inconclusive`. Only
`accepted` can satisfy a stage. Evidence is a bounded, sanitized artifact tied
to input/output and verifier digests.
| Verifier | Intended semantics |
| --- | --- |
| `ExactArtifactVerifier` | Whole-file or declared collection digest equality. |
| `CanonicalRecordVerifier` | Parse, validate, normalize, order, and serialize through a versioned canonicalizer. |
| `NumericToleranceVerifier` | Structured element/aggregate comparison with declared absolute, relative, ULP, NaN, and shape policies. |
| `StatisticalVerifier` | Repeated seeded evidence and versioned statistical acceptance criteria. |
| `DomainSpecificVerifier` | Workload-owned invariants, constraints, objective bounds, or reference checks. |
| `TrustedWorkerPolicy` | Accept only from allowed trust/environment attestations; still validate schema and bounds. |
The current whole-artifact SHA quorum maps only to `ExactArtifactVerifier`.
Canonical, numeric, stochastic, optimization, and side-effecting workloads MUST
declare an appropriate verifier/trust combination. Reducers consume only
accepted partial outputs and MUST detect missing, duplicate, conflicting, or
inconclusive inputs.
## 8. Failure, retry, cancellation, and checkpoint semantics
Every failure has a stable sanitized code, category (`input`, `scientific`,
`resource`, `infrastructure`, `lease`, `verification`, or `policy`), retryability,
and optional bounded evidence reference. Raw tracebacks and private paths remain
local.
- Retries create a new Attempt directory and resource lease; they never mutate
prior artifacts.
- Idempotent completion accepts the identical output manifest for the same
Attempt; conflicting manifests are rejected.
- Speculative execution, when enabled, creates multiple Attempts but commits at
most one accepted result and cancels the rest.
- Lease loss immediately fences upload/completion and terminates execution.
- Cancellation propagates to process groups/gangs and disposes uncommitted
staging artifacts.
- Retry budgets may be per Task, Stage, Loop, and Job; the strictest exhausted
budget wins.
- Checkpoint resume requires matching workload, schema, environment, and
checkpoint compatibility versions. Otherwise execution restarts cleanly.
- Side-effecting retries require an idempotency record or explicit operator
recovery.
Workflow failure policies are `fail_fast`, `continue_independent`,
`allow_partial` (only with an output schema/verifier that defines partial
results), and `compensate`. A failed reducer/verifier never leaves a Job marked
completed.
## 9. Package security and permissions
Installation and job submission are separate authorities. Only administrators
or managed policy may install, sign, approve, enable, upgrade, or revoke a
workload package. A job references an enabled immutable package digest.
Package policy MUST cover signature trust roots, dependency/image scanning,
license/provenance records, supported platforms, vulnerability/revocation
status, and reproducible build evidence. Upgrade does not rewrite running or
historical Jobs.
Each stage declares least-privilege permissions:
- network: none, coordinator-artifacts-only, allowlisted egress, or trusted;
- filesystem: read-only inputs, attempt scratch, declared outputs;
- secrets: named scoped handles, never raw values in Task parameters;
- subprocess: denied by default except the installed runner/runtime contract;
- devices and host interfaces: only allocated resources.
Worker Agents enforce permissions through the strongest available OS/container
sandbox and report the enforcement profile. A workload requiring unavailable
isolation is ineligible rather than silently unsandboxed.
## 10. SDK interfaces and conformance
The future Python API SHOULD expose protocols equivalent to:
```python
class Planner(Protocol):
def validate(self, request: JobRequest) -> ValidatedJob: ...
def plan(self, job: ValidatedJob, artifacts: ArtifactCatalog) -> WorkflowPlan: ...
class Runner(Protocol):
def run(self, context: TaskContext) -> OutputManifest: ...
class Reducer(Protocol):
def reduce(self, context: ReduceContext, inputs: AcceptedOutputs) -> OutputManifest: ...
class Verifier(Protocol):
def verify(self, context: VerifyContext, candidates: CandidateOutputs) -> VerificationDecision: ...
```
Concrete public value objects are immutable, typed, JSON-safe, schema-versioned,
and reject unknown fields. Scientific cores SHOULD remain callable without a
coordinator so the same implementation powers local and distributed adapters.
An SDK conformance suite MUST test manifest/schema validation, deterministic
planning, no local-path/URI leakage, output bounds, local/distributed parity,
retry and completion-order invariance, lease-loss cleanup, verifier behavior,
resource eligibility, and package permission declarations. Profile-specific
suites add two-worker byte equality, numeric tolerance, stochastic evidence,
stream recovery, loop limits, gang failure, or GPU parity as applicable.
## 11. Existing workload compatibility
- Local and distributed `similarity-search` map to a bounded shard-map and
top-k reducer workflow without duplicating the scientific algorithm.
- Local `similarity-graph` and future CTX-10 distribution map to triangular
block-pair expansion plus duplicate-safe edge reduction and pair-coverage
verification.
- Existing `DistributedPlan`, single input artifact, single partial result, and
`chunk_index` become the SDK compatibility profile `map-reduce-v1`.
- `descriptor-batch` is the first new reference implementation for the full
manifest, exact verifier, golden fixtures, and local/distributed parity.
No current workload is removed or renamed by adopting the SDK. Migration is an
adapter and manifest exercise first; protocol/database generalization occurs in
versioned later phases.
## 12. Deferred decisions
The roadmap must resolve these before implementation reaches the affected
phase:
1. SDK package ownership and independent release cadence.
2. Go-to-Python planner/reducer/verifier bridge and isolation boundary.
3. Verifier execution placement and environment attestation.
4. Streaming source/sink integrations and exactly-once scope.
5. Multi-node rendezvous, network identity, and gang scheduling persistence.
6. Accelerator sharing/MIG portability and accounting.
7. Workload signing technology, trust roots, and revocation distribution.
8. Tenant quotas, costs, priorities, fairness, and data-retention policy.
9. First-class artifact collections versus composite manifests.
10. Compatibility negotiation and deprecation support windows.
+163
View File
@@ -3,6 +3,11 @@
**Status:** future design and sequencing document. No SDK package, commands, or
general verifier abstraction described here is implemented yet.
The normative future API, workflow, execution, resource, security, and failure
semantics are specified in the design-draft
[`scimesh-sdk-contract.md`](scimesh-sdk-contract.md). This roadmap controls
delivery order and does not override that contract.
## Purpose and boundaries
The SDK should let a scientific developer add an allowlisted workload without
@@ -25,6 +30,9 @@ coordinator API contract. See [CTX-07](ctx-07-distributed-workload-protocol.md),
## Proposed public concepts
The detailed schemas and invariants are defined in the SDK contract; this table
is the roadmap-level responsibility map.
| Concept | Responsibility |
| --- | --- |
| `WorkloadDefinition` / `WorkloadManifest` | Name, versions, schemas, execution and verification metadata. |
@@ -91,6 +99,161 @@ Discovery should use an installed Python package, manifest, pinned environment
metadata, explicit entry points, and golden fixtures. It must be allowlisted;
never scan or execute user-provided module paths.
## General workload model: a versioned artifact workflow
Map/reduce is the first execution shape, not the limit of the SDK. The target
abstraction is an acyclic **workflow graph**: typed artifact ports connect
versioned stages, and a stage may fan out, fan in, or run once per job. This
allows the same SDK to express scientific ETL, simulations, parameter sweeps,
multi-step pipelines, model inference, image/video analysis, and the current
molecular workloads without placing scientific logic in the coordinator.
```text
Job inputs -> validate -> plan -> [map/partition stages] -> [join/reduce stages]
| |
accepted artifacts -----------+-> verify -> final manifest
```
The coordinator persists the graph, task attempts, leases, and artifact
ownership. The SDK declares stage behavior; it does not receive database access
or arbitrary commands. The initial `DistributedWorkload` protocol maps to a
single input, many map tasks, one reducer, and one final artifact. It remains a
supported compatibility profile rather than being replaced abruptly.
### Workflow and stage contracts
| Concept | Target responsibility |
| --- | --- |
| `WorkflowSpec` | Versioned DAG, external input ports, terminal outputs, global limits, and failure policy. |
| `StageSpec` | Stable stage ID, kind (`map`, `reduce`, `service`, `verify`), input/output port schemas, retry and resource policy. |
| `TaskSpec` | One concrete deterministic unit: stage ID, ordered artifact bindings, parameters, execution profile, and expected output manifest. |
| `ArtifactSchema` | Logical media type, schema version, cardinality, size bound, canonicalization rules, and privacy/retention class. |
| `ArtifactCollection` | Ordered, named, or keyed artifact set; used for shards, paired inputs, model bundles, and multiple outputs. |
| `OutputManifest` | Every output's artifact reference, schema/version/digest, metrics, provenance, and verifier evidence. |
| `FailurePolicy` | Retryable versus terminal errors, timeout, cancellation, partial-output disposal, and compensating cleanup rules. |
A stage is a pure artifact transformation wherever possible. Interactive,
long-running, or external-side-effect stages must declare that fact explicitly
and are initially trusted-only. A workflow cannot form cycles, read a
worker-local path from another stage, mutate a sealed input artifact, or produce
undeclared output ports. Dynamic fan-out is permitted only through a bounded,
versioned manifest emitted by an accepted planning stage; the coordinator must
enforce declared task, artifact, scratch, and output limits.
### Artifact and data-shape generality
The SDK must support more than CSV while retaining streamability and audit
trails. An `ArtifactSchema` can describe tabular records, scientific arrays,
images, meshes, molecular structures, model weights, archives, JSON manifests,
binary checkpoints, or opaque domain formats. It always declares how a consumer
validates structure and bounds bytes/records/dimensions before loading it.
Collections solve multi-input/multi-output work without an immediate database
rewrite. A task can initially receive one composite manifest artifact whose
entries name ordered or keyed logical inputs; it can return a composite output
manifest. Later protocol versions may persist first-class collection edges. The
collection manifest itself is immutable, coordinator-owned, schema-versioned,
and hash-addressed, so the old one-input/one-result API remains compatible.
## Execution model: Worker Agent, slots, and isolation
The Worker Agent is a resource manager, not a scientific runtime. One physical
machine registers one Agent. The Agent advertises a finite inventory and creates
isolated **execution slots**; each leased Task owns exactly one slot until it
finishes, loses its lease, or is cancelled.
```text
machine -> Worker Agent -> CPU / GPU / memory / scratch slot -> task subprocess -> attempt directory
```
`ExecutionProfile` declares whether a task uses a single process, a bounded
process pool, a distributed runtime, or an accelerator backend. It also carries
environment image/digest, entry-point identity, timeout, network policy,
scratch/output bounds, checkpoint policy, and determinism declaration. The
Agent—not a workload—sets environment variables, process groups, filesystem
roots, credentials, resource limits, and lifecycle signals.
### CPU parallelism
`cpu_cores` is a reservation, while `max_concurrency` is the number of slots;
neither is inferred from the other. CPU-bound Python work normally uses a
process pool constrained to the task's allocated cores. A workload must declare
its own internal parallelism and thread-library limits (for example OpenMP,
BLAS, Torch, or RDKit-related native code) so nested pools cannot oversubscribe
the host. The Agent starts independent heartbeat supervision per task and never
claims a task if it cannot reserve all declared resources.
Graceful draining means: stop new claims, continue heartbeat for active
attempts, request checkpoint/cancellation at deadline, then clean only that
attempt directory. Checkpoints are immutable artifacts and may be resumed only
when the workload's manifest explicitly supports checkpoint compatibility; they
are never treated as a completed result.
### Accelerator support
GPU/accelerator capability is generic inventory, not a coordinator-specific
CUDA feature. A future Agent reports device kind/vendor, UUID, compute
capability, memory, driver/runtime/image digest, supported backends, and
allocatable slot count. The coordinator only matches `ResourceRequirements` to
this inventory. The Agent assigns exclusive or shareable devices, sets device
visibility (for example `CUDA_VISIBLE_DEVICES`), reserves memory where the
platform supports it, starts the subprocess, measures usage, and releases the
slot.
The workload implementation chooses CUDA, ROCm, Metal, TPU, FPGA, SIMD, or a
CPU fallback; batches work and manages model/device memory; and declares the
scientific equivalence policy. Reducers and verifiers compare domain outputs,
not device-specific logs or floating-point text. GPU work cannot be enabled for
untrusted quorum merely because it runs: it additionally needs pinned images,
appropriate verifier/trust mode, and CPU/GPU or domain-valid parity evidence.
## Determinism, verification, and scientific validity
`DeterminismProfile` separates reproducibility from correctness:
| Profile | Examples | Minimum acceptance route |
| --- | --- | --- |
| `byte_exact` | canonical descriptors, fingerprints, sorted ETL | Exact artifact SHA-256 from independent owners. |
| `canonical_exact` | format-normalized records, deterministic structures | Versioned parser/canonicalizer then exact records. |
| `numeric_tolerance` | numerical solvers, GPU linear algebra | Structured comparison with absolute/relative/ULP tolerances and invariants. |
| `seeded_stochastic` | conformers, randomized search | Recorded seed, repeated-run policy, statistical/domain verifier. |
| `search_or_optimization` | routing, docking, retrosynthesis | Objective/constraint/domain evidence; often trusted execution. |
| `side_effecting` | instrument control, external database writes | Trusted-only, idempotency key, audit/compensation policy. |
The verifier consumes `OutputManifest` values, declared schemas, and bounded
streams; it returns accept/reject/inconclusive plus evidence. `inconclusive`
must never become success by a reducer default. Verification may be run by a
coordinator adapter, a pinned Python verifier subprocess, or a separate trusted
service—selection remains an open architectural decision. The manifest versions
the verifier configuration, tolerance values, canonicalizer, reference data,
and environment assumptions so historical results remain interpretable.
## Existing-workload migration matrix
| Existing capability | SDK workflow profile | Future adapter path |
| --- | --- | --- |
| Local `similarity-search` | single-process map + bounded top-k reduce | Keep local algorithm; expose a manifest and use current distributed planner/reducer. |
| Distributed `similarity-search` | deterministic shard map -> ordered reduce | Compatibility workload v1; later attach exact verifier and provenance manifest. |
| Local `similarity-graph` | triangular pair-partition map -> edge-set reduce | Preserve pair-coverage invariant as stage verifier. |
| Distributed `similarity-graph` | planned block-pair DAG -> duplicate-safe reduce | First major non-linear partition reference; implement before SDK generalization. |
| `descriptor-batch` | row-partition map -> ordered concatenation | First SDK reference workload and byte-exact quorum candidate. |
| Future ML/docking/QM/MD | parameter sweep, ensemble, or iterative workflow | Use numeric/domain/trusted verifier profile and explicit resource/environment contracts. |
## Authoring and operational lifecycle
An installed workload package should contain a signed or administrator-approved
manifest, Python entry points, schema migrations where required, pinned
environment metadata, golden fixtures, test vectors, and documentation. An
administrator controls enablement; users choose only among enabled manifests and
validated parameter ranges. Workload installation is separate from job
submission, preventing a user from sending code through the normal API.
The eventual author workflow remains deliberate: initialize a template, define
schemas and bounds, implement local scientific core, add planner/runner/reducer/
verifier adapters, generate fixtures, test local parity, test two-worker and
retry behavior, package, review, and enable. Any future CLI names are examples,
not implemented commands.
## Delivery sequence
1. Finish distributed `similarity-graph` and reliability/cross-language CI.