Define generalized workload SDK contract
This commit is contained in:
@@ -3,6 +3,11 @@
|
||||
**Status:** future design and sequencing document. No SDK package, commands, or
|
||||
general verifier abstraction described here is implemented yet.
|
||||
|
||||
The normative future API, workflow, execution, resource, security, and failure
|
||||
semantics are specified in the design-draft
|
||||
[`scimesh-sdk-contract.md`](scimesh-sdk-contract.md). This roadmap controls
|
||||
delivery order and does not override that contract.
|
||||
|
||||
## Purpose and boundaries
|
||||
|
||||
The SDK should let a scientific developer add an allowlisted workload without
|
||||
@@ -25,6 +30,9 @@ coordinator API contract. See [CTX-07](ctx-07-distributed-workload-protocol.md),
|
||||
|
||||
## Proposed public concepts
|
||||
|
||||
The detailed schemas and invariants are defined in the SDK contract; this table
|
||||
is the roadmap-level responsibility map.
|
||||
|
||||
| Concept | Responsibility |
|
||||
| --- | --- |
|
||||
| `WorkloadDefinition` / `WorkloadManifest` | Name, versions, schemas, execution and verification metadata. |
|
||||
@@ -91,6 +99,161 @@ Discovery should use an installed Python package, manifest, pinned environment
|
||||
metadata, explicit entry points, and golden fixtures. It must be allowlisted;
|
||||
never scan or execute user-provided module paths.
|
||||
|
||||
## General workload model: a versioned artifact workflow
|
||||
|
||||
Map/reduce is the first execution shape, not the limit of the SDK. The target
|
||||
abstraction is an acyclic **workflow graph**: typed artifact ports connect
|
||||
versioned stages, and a stage may fan out, fan in, or run once per job. This
|
||||
allows the same SDK to express scientific ETL, simulations, parameter sweeps,
|
||||
multi-step pipelines, model inference, image/video analysis, and the current
|
||||
molecular workloads without placing scientific logic in the coordinator.
|
||||
|
||||
```text
|
||||
Job inputs -> validate -> plan -> [map/partition stages] -> [join/reduce stages]
|
||||
| |
|
||||
accepted artifacts -----------+-> verify -> final manifest
|
||||
```
|
||||
|
||||
The coordinator persists the graph, task attempts, leases, and artifact
|
||||
ownership. The SDK declares stage behavior; it does not receive database access
|
||||
or arbitrary commands. The initial `DistributedWorkload` protocol maps to a
|
||||
single input, many map tasks, one reducer, and one final artifact. It remains a
|
||||
supported compatibility profile rather than being replaced abruptly.
|
||||
|
||||
### Workflow and stage contracts
|
||||
|
||||
| Concept | Target responsibility |
|
||||
| --- | --- |
|
||||
| `WorkflowSpec` | Versioned DAG, external input ports, terminal outputs, global limits, and failure policy. |
|
||||
| `StageSpec` | Stable stage ID, kind (`map`, `reduce`, `service`, `verify`), input/output port schemas, retry and resource policy. |
|
||||
| `TaskSpec` | One concrete deterministic unit: stage ID, ordered artifact bindings, parameters, execution profile, and expected output manifest. |
|
||||
| `ArtifactSchema` | Logical media type, schema version, cardinality, size bound, canonicalization rules, and privacy/retention class. |
|
||||
| `ArtifactCollection` | Ordered, named, or keyed artifact set; used for shards, paired inputs, model bundles, and multiple outputs. |
|
||||
| `OutputManifest` | Every output's artifact reference, schema/version/digest, metrics, provenance, and verifier evidence. |
|
||||
| `FailurePolicy` | Retryable versus terminal errors, timeout, cancellation, partial-output disposal, and compensating cleanup rules. |
|
||||
|
||||
A stage is a pure artifact transformation wherever possible. Interactive,
|
||||
long-running, or external-side-effect stages must declare that fact explicitly
|
||||
and are initially trusted-only. A workflow cannot form cycles, read a
|
||||
worker-local path from another stage, mutate a sealed input artifact, or produce
|
||||
undeclared output ports. Dynamic fan-out is permitted only through a bounded,
|
||||
versioned manifest emitted by an accepted planning stage; the coordinator must
|
||||
enforce declared task, artifact, scratch, and output limits.
|
||||
|
||||
### Artifact and data-shape generality
|
||||
|
||||
The SDK must support more than CSV while retaining streamability and audit
|
||||
trails. An `ArtifactSchema` can describe tabular records, scientific arrays,
|
||||
images, meshes, molecular structures, model weights, archives, JSON manifests,
|
||||
binary checkpoints, or opaque domain formats. It always declares how a consumer
|
||||
validates structure and bounds bytes/records/dimensions before loading it.
|
||||
|
||||
Collections solve multi-input/multi-output work without an immediate database
|
||||
rewrite. A task can initially receive one composite manifest artifact whose
|
||||
entries name ordered or keyed logical inputs; it can return a composite output
|
||||
manifest. Later protocol versions may persist first-class collection edges. The
|
||||
collection manifest itself is immutable, coordinator-owned, schema-versioned,
|
||||
and hash-addressed, so the old one-input/one-result API remains compatible.
|
||||
|
||||
## Execution model: Worker Agent, slots, and isolation
|
||||
|
||||
The Worker Agent is a resource manager, not a scientific runtime. One physical
|
||||
machine registers one Agent. The Agent advertises a finite inventory and creates
|
||||
isolated **execution slots**; each leased Task owns exactly one slot until it
|
||||
finishes, loses its lease, or is cancelled.
|
||||
|
||||
```text
|
||||
machine -> Worker Agent -> CPU / GPU / memory / scratch slot -> task subprocess -> attempt directory
|
||||
```
|
||||
|
||||
`ExecutionProfile` declares whether a task uses a single process, a bounded
|
||||
process pool, a distributed runtime, or an accelerator backend. It also carries
|
||||
environment image/digest, entry-point identity, timeout, network policy,
|
||||
scratch/output bounds, checkpoint policy, and determinism declaration. The
|
||||
Agent—not a workload—sets environment variables, process groups, filesystem
|
||||
roots, credentials, resource limits, and lifecycle signals.
|
||||
|
||||
### CPU parallelism
|
||||
|
||||
`cpu_cores` is a reservation, while `max_concurrency` is the number of slots;
|
||||
neither is inferred from the other. CPU-bound Python work normally uses a
|
||||
process pool constrained to the task's allocated cores. A workload must declare
|
||||
its own internal parallelism and thread-library limits (for example OpenMP,
|
||||
BLAS, Torch, or RDKit-related native code) so nested pools cannot oversubscribe
|
||||
the host. The Agent starts independent heartbeat supervision per task and never
|
||||
claims a task if it cannot reserve all declared resources.
|
||||
|
||||
Graceful draining means: stop new claims, continue heartbeat for active
|
||||
attempts, request checkpoint/cancellation at deadline, then clean only that
|
||||
attempt directory. Checkpoints are immutable artifacts and may be resumed only
|
||||
when the workload's manifest explicitly supports checkpoint compatibility; they
|
||||
are never treated as a completed result.
|
||||
|
||||
### Accelerator support
|
||||
|
||||
GPU/accelerator capability is generic inventory, not a coordinator-specific
|
||||
CUDA feature. A future Agent reports device kind/vendor, UUID, compute
|
||||
capability, memory, driver/runtime/image digest, supported backends, and
|
||||
allocatable slot count. The coordinator only matches `ResourceRequirements` to
|
||||
this inventory. The Agent assigns exclusive or shareable devices, sets device
|
||||
visibility (for example `CUDA_VISIBLE_DEVICES`), reserves memory where the
|
||||
platform supports it, starts the subprocess, measures usage, and releases the
|
||||
slot.
|
||||
|
||||
The workload implementation chooses CUDA, ROCm, Metal, TPU, FPGA, SIMD, or a
|
||||
CPU fallback; batches work and manages model/device memory; and declares the
|
||||
scientific equivalence policy. Reducers and verifiers compare domain outputs,
|
||||
not device-specific logs or floating-point text. GPU work cannot be enabled for
|
||||
untrusted quorum merely because it runs: it additionally needs pinned images,
|
||||
appropriate verifier/trust mode, and CPU/GPU or domain-valid parity evidence.
|
||||
|
||||
## Determinism, verification, and scientific validity
|
||||
|
||||
`DeterminismProfile` separates reproducibility from correctness:
|
||||
|
||||
| Profile | Examples | Minimum acceptance route |
|
||||
| --- | --- | --- |
|
||||
| `byte_exact` | canonical descriptors, fingerprints, sorted ETL | Exact artifact SHA-256 from independent owners. |
|
||||
| `canonical_exact` | format-normalized records, deterministic structures | Versioned parser/canonicalizer then exact records. |
|
||||
| `numeric_tolerance` | numerical solvers, GPU linear algebra | Structured comparison with absolute/relative/ULP tolerances and invariants. |
|
||||
| `seeded_stochastic` | conformers, randomized search | Recorded seed, repeated-run policy, statistical/domain verifier. |
|
||||
| `search_or_optimization` | routing, docking, retrosynthesis | Objective/constraint/domain evidence; often trusted execution. |
|
||||
| `side_effecting` | instrument control, external database writes | Trusted-only, idempotency key, audit/compensation policy. |
|
||||
|
||||
The verifier consumes `OutputManifest` values, declared schemas, and bounded
|
||||
streams; it returns accept/reject/inconclusive plus evidence. `inconclusive`
|
||||
must never become success by a reducer default. Verification may be run by a
|
||||
coordinator adapter, a pinned Python verifier subprocess, or a separate trusted
|
||||
service—selection remains an open architectural decision. The manifest versions
|
||||
the verifier configuration, tolerance values, canonicalizer, reference data,
|
||||
and environment assumptions so historical results remain interpretable.
|
||||
|
||||
## Existing-workload migration matrix
|
||||
|
||||
| Existing capability | SDK workflow profile | Future adapter path |
|
||||
| --- | --- | --- |
|
||||
| Local `similarity-search` | single-process map + bounded top-k reduce | Keep local algorithm; expose a manifest and use current distributed planner/reducer. |
|
||||
| Distributed `similarity-search` | deterministic shard map -> ordered reduce | Compatibility workload v1; later attach exact verifier and provenance manifest. |
|
||||
| Local `similarity-graph` | triangular pair-partition map -> edge-set reduce | Preserve pair-coverage invariant as stage verifier. |
|
||||
| Distributed `similarity-graph` | planned block-pair DAG -> duplicate-safe reduce | First major non-linear partition reference; implement before SDK generalization. |
|
||||
| `descriptor-batch` | row-partition map -> ordered concatenation | First SDK reference workload and byte-exact quorum candidate. |
|
||||
| Future ML/docking/QM/MD | parameter sweep, ensemble, or iterative workflow | Use numeric/domain/trusted verifier profile and explicit resource/environment contracts. |
|
||||
|
||||
## Authoring and operational lifecycle
|
||||
|
||||
An installed workload package should contain a signed or administrator-approved
|
||||
manifest, Python entry points, schema migrations where required, pinned
|
||||
environment metadata, golden fixtures, test vectors, and documentation. An
|
||||
administrator controls enablement; users choose only among enabled manifests and
|
||||
validated parameter ranges. Workload installation is separate from job
|
||||
submission, preventing a user from sending code through the normal API.
|
||||
|
||||
The eventual author workflow remains deliberate: initialize a template, define
|
||||
schemas and bounds, implement local scientific core, add planner/runner/reducer/
|
||||
verifier adapters, generate fixtures, test local parity, test two-worker and
|
||||
retry behavior, package, review, and enable. Any future CLI names are examples,
|
||||
not implemented commands.
|
||||
|
||||
## Delivery sequence
|
||||
|
||||
1. Finish distributed `similarity-graph` and reliability/cross-language CI.
|
||||
|
||||
Reference in New Issue
Block a user