diff --git a/docs/specs/drafts/VASTAI_E2E_FUZZ_PLAN.md b/docs/specs/drafts/VASTAI_E2E_FUZZ_PLAN.md new file mode 100644 index 0000000..ff4a614 --- /dev/null +++ b/docs/specs/drafts/VASTAI_E2E_FUZZ_PLAN.md @@ -0,0 +1,931 @@ +# VastAI Real-Node E2E Fuzz Plan + +## Start here + +**Status as of 2026-09-14:** the latest retained run passed five-node Gate A +and the complete local campaign, then exceeded the complete warm-workflow +ceiling before remaining Gate B checks. Its source identity differs from the +reviewed tree. Fresh ordered qualification and paid acceptance remain pending. +Paid execution is blocked. + +Checkpoint handoffs: [remote-ready work](VASTAI_REMOTE_READY.md) and +[local blockers, in priority order](VASTAI_LOCAL_BLOCKERS.md). The internal +workload/cleanup budget split is not a separate blocker for this checkpoint +review; the handoffs leave runtime limits and paid-access guards unchanged. + +This is the behavioral acceptance contract for the existing real-binary fuzz +harness. It defines guarantees, scenarios, gate order, safety limits, and +required evidence—not file layout, internal APIs, identifiers, serialization, +or process-supervision mechanics. Follow existing repository conventions for +those implementation choices. + +### Agent execution rules + +1. Start at the earliest gate without fresh passing evidence: + **Gate A → Gate B → paid canary**. Gate definitions are in Section 18; + workflow order and failure transitions are in Sections 4 and 17. +2. Fix the behavior blocking that gate. Focused tests and diagnostic runs may + explain a failure; they never substitute for the complete gate. +3. Preserve the case counts, coverage, independent oracle, isolation, deadlines, + acquisition bounds, and cleanup proofs below. A partial campaign is a failed + or diagnostic run, not reduced acceptance. +4. After source, executable, runtime-image, or fixture/deployment-identity + changes, obtain fresh ordered evidence. Do not reuse an old pass or combine + campaign coverage across deployment identities. +5. Paid campaign invocations require `--gate-attestation` referencing fresh, + successful `ordered_acceptance.py` evidence. Cleanup-only requires no + attestation and must remain available. +6. Report the gate, artifact path, exercised scope, PASS/FAIL result, and next + blocker. Use Section 21 for readiness review. Do not carry forward old resume + commands as authorization to skip gates. + +### Evidence checkpoint (historical, not authorization) + +| Evidence | Established scope | Limit | +|---|---|---| +| `target/vastai-plan-acceptance-20260910-3/warm-1/gate-a/deployment-e2e/gate-a-evidence.json` | Untraced five-node Gate A: all 12 rounds and complete cleanup | Subsequent source changes require ordered revalidation | +| `target/vastai-plan-campaign-diagnostic-2/vastai-e2e-000000000135282e/` | 26 normal cases passed before shared workload allocation exhausted at `normal-026-blob-4-5`; admission accounts for 208 cases / 224 segments | Not Gate B acceptance | +| `target/vastai-binding-shutdown-smoke.json` | Direct-only child bindings avoid unrelated default public relays; main-to-return teardown fell from 1.001 s to 0.002 s without shorter timeouts | Not complete campaign timing | +| `target/vastai-gate-b-scripted.json` and `target/vastai-gate-b-cleanup-safety.json` | Focused scripted-provider admission and cleanup safety checks passed | Not the complete ordered Gate B | + +Ordered source attestation includes Cargo configuration and actor-control-flow +lint tooling alongside runtime and harness source. Historical evidence above +is a debugging reference, not a claim that current artifacts pass. + +--- + +## 1. Objective + +One manually invoked workflow MUST: + +1. select five cheap, verified VastAI offers on distinct hosts; +2. provision and bootstrap those five nodes once; +3. prove that the cluster converged and is usable within ten minutes; +4. retain that cluster while generated E2E cases use fresh processes, namespace + state, blobs, streams, descriptors, observations, and model state; +5. replay and shrink a failure without provisioning more infrastructure; +6. restart and recover the orchestrator without re-provisioning or + re-bootstrapping nodes; +7. finish with five-to-four and four-to-three destructive phases; and +8. destroy and account for every contract created by the run. + +Node acquisition and bootstrap are the expensive operations. The fixture is +therefore reused, but no generated case may consume state or observations from +a previous case. + +Model-based generation is the center of the suite. Fixed examples are limited +to readiness checks, persisted minimized regressions, and scenarios whose exact +ordering is itself the contract. + +--- + +## 2. Fixed campaign decisions + +The fixed campaign shape and limits are: + +- initial topology: exactly five worker nodes; +- normal campaign: 128 generated cases on all five nodes; +- recovery campaign: 16 cases that cross an orchestrator restart; +- first destructive phase: remove one node, then run 32 cases on the four exact + survivors; +- second destructive phase: remove another node, then run 32 cases on the three + exact survivors; +- paid preparation deadline: ten minutes for provisioning, bootstrap, convergence, + readiness proof, readiness-proof cleanup, and prepared-fixture commit; +- complete local campaign deadline: five minutes, including preparation, + all cases, recovery, destructive transitions, cleanup and evidence; +- prepared-fixture workload allocation: at most four minutes, shared by all + 208 logical cases / 224 pre/post workload segments and auxiliary probes; +- diagnosis allowance: at most 15 seconds, subordinate to remaining execution + and fixture lifetime; exhaustion cannot discard the original failure; +- fixture cleanup reserve: 15 seconds locally; five minutes for paid cleanup, + separate from the failing workload; +- per-case admission ceilings: 8 MiB payload and 64 MiB allocation; +- campaign admission ceilings: 512 MiB payload and 8 GiB cumulative allocation, + accounting separately for endpoint rings and public binding copies; and +- invocation model: manual only, with no CI or recurring scheduler in scope. + +Offer selection MUST use eligible verified offers, choose distinct hosts, and +prefer the cheapest eligible set. If five suitable distinct hosts are not +available, the run fails before provisioning. + +The operator MUST provide explicit fixture lifetime and cost ceilings. The +workflow MUST reject a run before provisioning when its planned campaign cannot +fit those limits. + +These are monotonic budget and admission ceilings, not measured speed or +resident-memory claims. The [speed migration](VASTAI_E2E_FUZZ_SPEED_MIGRATION.md) +requires three consecutive complete warm ordered workflows within ten minutes +each, with Gate A before all Gate B checks, plus separately qualified cold +timing. Paid planning reserves up to 20 minutes for preparation, campaign, and +cleanup; offer search and actual provider capacity are reported separately. +Foreground cleanup expiry fails acceptance; the durable paid cleanup duty +continues. + +The run may acquire at most the five initially selected contracts and may start +at most one initial logical bootstrap session for each selected node. Provider +or initial-bootstrap failure during preparation fails the fixture; it does not +trigger replacement acquisition. Explicit retained-node binary redeployment is +a separately accounted operation: it may reuse those contracts, but it may +never acquire, replace, or remap infrastructure. + +--- + +## 3. Existing harness changes required + +Extend the existing case model, generated public-binding programs, independent +oracle, shrinker, regression corpus, real-binary process control, and fixture +lifecycle. Do not create a parallel VastAI-only harness. Required changes: + +- select real-provider execution throughout fixture startup; +- separate fixture preparation from per-case execution and cleanup; +- generate cases over exact, possibly non-contiguous survivor node sets; +- support more processes and richer multi-node topologies; +- collect observations fairly from concurrent executions; +- replay and shrink against an already prepared fixture; +- restart the orchestrator while retaining prepared nodes; and +- use provider-neutral health and teardown checks. + +Control requests, replies, provisioning facts, and artifacts MUST use shared, +versioned contracts rather than separate lookalike representations. + +--- + +### 3.1 Blocking retained-node binary-redeployment gate + +Retained-node binary redeployment MUST be implemented and proven before any +remaining paid-campaign work. It must provide: + +```text +established provider resource with reachable root SSH ++ arbitrary prior Myelin/Swactor execution state ++ a new deployment bundle and deployment identity +→ the same provider resource running exactly the new binaries ++ fresh Iroh/Swactor membership, routing, readiness, and telemetry +``` + +The node's storage contents, caches, model state, and GPU state are outside this +contract. The redeployment may preserve or destroy them. It must not depend on +their contents. Its substrate preconditions are limited to reachable root SSH +and a functional node capable of receiving, writing, and executing the +deployment. Failure of those substrate capabilities is reported explicitly; +it must not be disguised as an indefinitely bootstrapping worker. + +The local gate MUST use a configurable fleet of `N` pre-created Docker nodes +that behave as remote machines. Before paid execution, the gate MUST pass with +the five-node campaign shape. Each node MUST: + +- use the production remote-node base environment; +- have an independently masked network, with no shared Docker LAN that can + accidentally satisfy cluster communication; +- be reachable for deployment only through its published SSH endpoint; and +- retain the same container and network identity through every redeployment + round. + +Docker control may construct adversarial starting states and census fixture +resources. It may not install or launch the tested binaries. Artifact delivery, +reset, launch, and recovery MUST use the same orchestrator/provider path that a +retained VastAI node uses. + +The gate uses at least two distinguishable deployment artifacts and identities. +It first deploys generation A, proves cluster behavior, damages the nodes into +different prior states, and then deploys generation B without recreating the +fixture: + +```text +create N remote-shaped nodes once +→ deploy generation A through the orchestrator +→ prove membership, routing, telemetry, and an all-pairs behavioral exchange +→ quiesce generated work +→ construct adversarial and mutually different node states +→ deploy generation B to the same N nodes +→ prove generation-B membership, routing, telemetry, and all-pairs behavior +→ repeat with transport and process crashes at every deployment boundary +``` + +Across deterministic rounds, the adversarial states MUST include: + +- no worker and no trustworthy deployment metadata; +- one healthy stale worker; +- a dead worker with stale pid, lock, socket, or completion metadata; +- multiple stale workers or process descendants; +- an interrupted artifact transfer or partial installation; +- a corrupt active binary, activation pointer, or deployment descriptor; +- a deployment transaction interrupted before launch, after launch, and before + receipt observation; +- loss of the SSH client or orchestrator at each observable transaction + boundary; and +- a temporarily unreachable node that later becomes reachable with its prior + state intact. + +The same node may be used for multiple states; every required state must be +covered before the gate passes. Fault construction must not teach the +deployment path which state was injected. + +For every redeployment round, the harness MUST prove: + +1. the exact provider resources, Docker containers, networks, logical node + identities, and node-to-resource mapping did not change; +2. all stale Myelin/Swactor process incarnations and descendants are gone + before a new worker is accepted; +3. exactly one worker incarnation per logical node reports the requested + artifact digest and deployment generation; +4. the active executable bytes match the requested artifact; +5. a successful deployment requires a valid, matching launch receipt rather + than merely a zero SSH exit status; +6. SSH closure or failure after remote launch is not interpreted as worker + exit and does not create duplicate workers; +7. repeating an interrupted deployment is idempotent and converges without + manual node repair; +8. stale runtime-ready, rejoin, membership, route, or telemetry observations + cannot make an older generation live; +9. each new worker establishes communication and telemetry independently of + the SSH session; and +10. cluster operations resume only after every required node has passed fresh + membership, route-ownership, runtime-ready acknowledgement, telemetry, and + all-pairs behavioral gates. + +Deployment convergence is level-triggered. Transient SSH, process, and +transport failures retry with bounded backoff and no elapsed-time success or +failure criterion. Authentication rejection, an invalid artifact, or loss of +the substrate preconditions is a typed terminal failure. + +An explicit retained-node redeployment on a paid development fixture follows +the same contract. It MUST begin from a quiescent fixture, retain the exact +contracts and logical topology, prohibit acquisition and replacement, assign a +fresh deployment identity, and rerun readiness and behavioral gates before +cases resume. A redeployment invalidates prior campaign coverage and fixture +baselines: coverage-bearing execution restarts from the beginning on the latest +deployment identity. This permits iterative remote debugging without claiming +that results collected across different binaries form one accepted campaign. + +The redeployment gate is complete only when its fault-injection and cleanup +checks pass. The remaining local real-binary E2E suite runs only afterward. + +--- + +## 4. Required execution order + +Advancement to paid execution is strictly gated: + +```text +implement retained-node binary redeployment +→ pass the local N-node remote-shaped redeployment fault suite +→ pass the complete remaining local real-binary E2E suite +→ pass the remaining mock, scripted-provider, safety, and cleanup checks +→ unlock real-VastAI development and canary execution +``` + +A failure at any local gate blocks every later gate. Passing an isolated test +or manually demonstrating a remote deployment cannot skip this order. + +The paid workflow then runs in this order: + +```text +preflight and generate/validate the complete campaign +→ start the cleanup owner and the real orchestrator +→ search for and select five offers +→ authorize the bounded paid operation +→ provision and converge five nodes +→ run the all-pairs readiness gate and clean its state +→ commit the prepared fixture and prohibit further acquisition +→ run compatible persisted regressions +→ run 128 five-node generated cases +→ run 16 orchestrator-recovery cases +→ remove one node and verify the exact survivor set +→ run 32 four-node generated cases +→ remove one node and verify the exact survivor set +→ run 32 three-node generated cases +→ stop the orchestrator and destroy remaining contracts +→ prove every contract created or discovered for this run is absent +``` + +The complete generated campaign and its resource bounds MUST be validated +before provider access. A failure before the prepared fixture is committed in a +formal acceptance run is cleanup-only: the workflow does not resume initial +bootstrap or restart preparation. + +Once preparation succeeds, every path used by normal execution, recovery, +replay, shrinking, and regressions MUST be unable to acquire more +infrastructure. Retained-node binary redeployment is an explicit, +quiescent-fixture transition, never an automatic response to health failure. A +successful redeployment returns the workflow to the readiness gate and resets +campaign coverage; a failed redeployment permits only another explicit +redeployment attempt, evidence collection, or provider cleanup. + +--- + +## 5. Fixture safety and lifecycle + +### 5.1 Preparation + +Preparation succeeds only when: + +- all five selected nodes correspond to distinct expected contracts and hosts; +- all five nodes are running and contextual control is reachable; +- required provisioning and bootstrap milestones are present; +- no unexpected provider acquisition or bootstrap occurred; +- the all-pairs readiness gate passed; +- readiness workloads and resources were cleaned up; and +- the reusable fixture baseline was durably recorded before the deadline. + +Repeated connection attempts while a selected node becomes reachable are +allowed. They remain part of that node's single logical bootstrap session. + +Any selected-node failure, vanished contract, failed readiness proof, resource +leak, or preparation timeout immediately ends preparation and enters cleanup. +No generated case starts from a partially prepared fixture. + +### 5.2 Paid-operation boundary + +Every path capable of creating a contract or starting bootstrap MUST share one +run-scoped acquisition bound. The bound MUST: + +- authorize only the five selected nodes; +- conservatively count ambiguous provider outcomes; +- survive runner or orchestrator failure; +- prevent retries from becoming replacement acquisitions; and +- become permanently deny-only when preparation succeeds. + +The durable run record MUST contain enough expected provider identity to find +and clean contracts even if a process fails between provider acceptance and +normal telemetry. + +No crash or retry may exceed the acquisition bound or leave a possibly live +contract unaccounted for. Representation, key derivation, persistence, and +locking remain implementation choices. + +### 5.3 Per-case isolation + +Infrastructure is reused; case state is not. + +Every attempt MUST have a fresh isolated ownership scope covering all writable +namespace paths, processes, streams, descriptors, observations, and model state. +Before execution, the harness proves that owned state is absent and records the +fixture's resource baseline. + +Every terminal path—success, expected process failure, launch failure, oracle +failure, or timeout—MUST: + +1. stop remaining generated processes; +2. release or remove all case-owned data-plane resources; +3. prove owned state is absent; +4. prove all generated processes are terminal; +5. restore execution, actor, and live resource gauges to baseline; and +6. prove provider contracts and bootstrap counts did not change. + +Execution and cleanup use separate deadlines. An execution timeout never skips +cleanup. + +Failure to restore the baseline quarantines the fixture. No later case, replay, +or shrink candidate may run against it; only evidence collection and provider +cleanup remain. + +--- + +## 6. Readiness gate + +The prepared fixture MUST pass a fresh directed all-pairs data proof across all +five nodes: + +- every ordered node pair transfers and validates a blob; +- every ordered node pair transfers and validates a framed stream; +- payloads include small, boundary-sized, and multi-chunk values; +- stream endpoints open concurrently so the proof does not depend on a + writer-first or reader-first schedule; +- identity, order, length, content digest, and terminal stream state are + checked; and +- every proof process and namespace resource is removed afterward. + +The proof shares the ten-minute preparation deadline. It is a gate for the +entire campaign, not an ordinary generated case. + +--- + +## 7. Generated case IR + +A persisted generated case MUST describe behavior rather than concrete harness +implementation. It includes: + +- schema and generator versions; +- seed and stable case identity; +- the exact live logical-node set required by the case; +- one to twenty process programs; +- process-to-node assignments; +- public-binding actions; +- acyclic process-completion dependencies; +- blob and stream routes, including branch and join structure; +- expected or admissible outcomes; and +- modeled non-node-loss faults. + +Cases record exact node membership and never assume contiguous node numbering. +Replay requires the same live set. Shrinking may remove an unused node from a +case but may not silently substitute a different fixture node. + +The IR distinguishes writable case-owned data from explicit read-only fixture +data. Concrete names, paths, request identities, and other execution-local +values are produced per attempt and are not part of the behavioral contract. + +Node loss is not a reusable generated-case fault. It is a controlled campaign +phase transition. + +--- + +## 8. Reference model and oracle + +The independent model tracks the minimum state needed to judge behavior: + +- namespace entry kind, revision, ownership, and legal mutation outcomes; +- blob length and content digest; +- stream incarnation, participants, frame sequence, clean close or abort, and + replacement isolation; +- process lifecycle, dependencies, expected terminal class, and owned + resources; and +- authorization boundaries and expected errors. + +The oracle MUST derive expected payload and state transitions independently of +the system under test. + +Concurrent exclusive publish, rename, unlink, replacement, and active-stream +mutation may have multiple legal outcomes. The oracle MUST retain the bounded +set of legal states and use causal observations to eliminate impossible states. +Polling order or local receipt time MUST NOT choose a race winner. + +The case generator MUST reject a case before provisioning if its modeled race +space or planned resource usage exceeds configured bounds. Runtime model +explosion is a harness failure, not permission to choose an arbitrary outcome. + +Case success requires both behavioral-oracle success and successful cleanup and +baseline restoration. + +--- + +## 9. Topology and scenario generation + +Healthy generated cases contain 16 to 64 actions over 2 to 20 processes. The +campaign uses these topology families: + +| Weight | Family | Required shape | +|---:|---|---| +| 20% | Chain | 3 to 5 distinct nodes | +| 15% | Ring or random walk | repeated traversal without reusing a data edge | +| 15% | Fan-out | one source and 2 to 4 independent branches | +| 15% | Fan-in | 2 to 4 producers and one concurrent sink | +| 10% | Diamond | fork, independent transformed branches, and join | +| 25% | Random DAG | 3 to 20 vertices with bounded fan-in and fan-out | + +A route may revisit a node, but each directed data edge has its own stream or +blob path and process role. Process-completion dependencies remain acyclic even +when the data route is a ring or walk. + +Generated cases compose these behaviors: + +- multi-hop blob and framed-stream relay; +- concurrent fan-in and fan-out with complete branch accounting; +- branch joins that depend on every expected input; +- multiple routes sharing nodes but not writable paths; +- descriptor access and namespace mutation; +- active and quiescent stream mutation with modeled outcomes; +- concurrent exclusive operations and legal loser errors; +- authorized and denied access-prefix operations; +- process churn while unrelated routes continue; +- writer abort, reader stop, and stream replacement; +- one expected-failed or slow branch without cancellation of healthy siblings; +- hot high-volume work while cold flows make observable progress; and +- expected process failure isolated from healthy processes and routes. + +Payload selection MUST cover empty, small, framing/chunk boundaries, large +multi-chunk values, and randomized sizes within the campaign's byte budget. +Streams MUST test framing independent of transport chunk boundaries, multiple +logical frames, clean EOF, abort propagation, and stale-incarnation rejection. + +The precise payload generator, frame binary layout, and buffering strategy are +implementation details. They must be deterministic across generated programs +and the independent oracle, bounded in memory, and capable of detecting loss, +duplication, reordering, corruption, and cross-attempt contamination. + +--- + +## 10. Required ordering scenarios + +Most concurrency is judged by partial order, not a total schedule. The following +scenarios require explicit ordering guarantees. + +### 10.1 Concurrent stream startup + +For rings, fan-in, fan-out, diamonds, and all-pairs readiness, participants open +required endpoints concurrently before waiting for peer completion. The test +MUST prove progress without relying on endpoint creation order. + +### 10.2 Ring completion + +Ring participants start concurrently. Tokens traverse the required laps, each +edge validates and forwards them, and completion propagates only after all +injected tokens are accounted for. The origin MUST continue receiving while +injection or forwarding can block, and the ring MUST terminate with clean EOF +rather than deadlock. + +### 10.3 Hot/cold fairness + +The generated scenario establishes this causal order: + +```text +hot flow is active and parked +→ its first frame is observed +→ each cold flow starts useful work +→ every cold flow terminates +→ the hot flow receives release +→ the hot flow terminates +``` + +The hot stream remains active until release. The guarantee is observable +progress of cold work under concurrent load, not a latency threshold or a +particular scheduler behavior. + +### 10.4 Failure isolation + +An expected failure in one branch or process MUST NOT implicitly cancel healthy +siblings. The model and observations must account independently for every +branch's outcome and cleanup. + +--- + +## 11. Coverage requirements + +The five-node normal campaign MUST close both a planned coverage ledger and an +observed completed ledger. It requires: + +- every topology family at least eight times; +- every logical node as source, sink, and interior relay for blobs and streams; +- every ordered node pair as a blob edge and stream edge; +- every existing primitive action-category adjacency; +- all required payload boundary classes; +- every descriptor read/write mode and terminal mode; +- successful operations, modeled errors, and contextual-process failures; +- concurrent starts for chain, ring/walk, fan-in/out, and diamond families; and +- the actor-style progress, churn, mutation, and failure-isolation scenarios in + this plan. + +Generation fills required coverage slots before applying random weights. A plan +that cannot meet coverage or resource bounds fails before provider access. +Observed incomplete coverage fails the campaign rather than being reported as a +reduced successful run. + +Each survivor campaign MUST cover every remaining node as source, sink, and +relay, and every ordered survivor pair as both a blob and stream edge. + +--- + +## 12. Observation requirements + +Generated workloads use only the public application binding. Observations MUST +carry enough typed information for the independent oracle to establish: + +- attempt and process-local identity; +- process lifecycle and local action order; +- route, token, hop, and stream-incarnation relationships; +- source and destination ownership where relevant; +- input/output lengths and digests; +- progress barriers and fault triggers; and +- typed outcomes and errors. + +There is no assumed global event order across processes. The oracle constructs a +partial order from process-local order, dependencies, data-route edges, explicit +barriers, fault triggers, and namespace revisions. + +Observation transport MUST be bounded, untorn, loss-detecting, and drained +fairly across live executions. The implementation may choose the encoding and +collection mechanics, but it MUST prove before a paid run that concurrent +records can be recovered completely without truncating required evidence. + +Timeout artifacts include enough model, process, fleet, and recent telemetry +state to explain what remained pending. + +--- + +## 13. Recovery campaign + +Each of the 16 recovery cases has a pre-restart segment and a post-restart +segment within one case attempt. + +Before restart: + +- all pre-restart contextual processes are terminal; +- no stream is active; +- live resource gauges are at baseline; and +- the model explicitly identifies the blobs and quiescent namespace entries + intended to persist. + +The orchestrator is then stopped and replaced using the same prepared fixture +and persistent state. Recovery MUST: + +- perform no offer search, contract creation, node bootstrap, or topology repair; +- adopt the exact existing contracts and logical nodes; +- restore all five nodes to running and reachable; +- begin with no live contextual executions; +- preserve persisted blob kind, revision, length, and digest; +- preserve quiescent stream namespace identity and inactive state, without + treating consumed bytes as persistent stream contents; and +- allow fresh post-restart processes to complete the remaining segment. + +Active contextual processes and active streams are outside the recovery +contract. Normal per-attempt cleanup occurs after the post-restart segment. + +--- + +## 14. Destructive tail + +After the normal and recovery campaigns, the workflow removes nodes through +normal orchestrator control. + +For each removal, command acceptance alone is insufficient. The harness MUST +prove that: + +- the selected target reached the stopped state; +- every non-target survivor remained running and reachable; +- the provider contract set lost exactly the target contract; +- the running logical-node set lost exactly the target node; and +- no replacement intent, node, contract, acquisition, or bootstrap appeared. + +The four-node campaign runs on the exact first survivor set. The three-node +campaign runs on the exact second survivor set. Neither phase claims fresh +four-node or three-node bootstrap coverage. + +A failing workload on a degraded live set may be replayed and shrunk before the +next node is removed. The node-removal transition itself is not part of a +shrinkable case. + +--- + +## 15. Replay, shrinking, and regressions + +The first unexpected behavioral failure stops ordinary campaign execution. +After the failed attempt has cleaned up and restored the fixture baseline, the +runner replays and shrinks it within the same fixture, lifetime, cost, and +resource bounds. + +Replay MUST: + +- require the recorded exact live-node set; +- use fresh per-attempt state; +- perform no infrastructure acquisition, repair, remapping, or node mutation; +- reproduce the same typed behavioral failure; and +- pass the same cleanup and health checks as an ordinary case. + +Shrinking may reduce faults, processes, dependencies, actions, topology edges, +laps, tokens, payloads, and unused node participation. It MUST preserve the +failure's behavioral signature rather than incidental identities, timestamps, +or artifact positions. + +Cleanup failure stops shrinking. The unshrunk failure remains available when the +fixture cannot safely run candidates. + +Minimized compatible regressions run before random cases. Incompatible artifacts +are rejected clearly rather than interpreted under a different schema or +silently remapped. + +--- + +## 16. Health and teardown + +Provider-neutral health checks MUST verify: + +- the orchestrator is live when expected; +- the exact expected logical nodes are in the expected phases; +- contextual control is reachable on live nodes; +- no unexpected execution, panic, poison, or transient actor remains; +- live resource gauges return to baseline; and +- contract and bootstrap accounting remains unchanged during case execution. + +Health behavior MUST reflect the selected provider. Real-provider health and +cleanup cannot rely on local-container census or local mock-resource behavior. +The concrete provider abstraction is an implementation decision. + +Teardown MUST continue despite individual errors and must: + +1. stop all still-managed nodes; +2. stop and reap the orchestrator and its provider-facing work; +3. discover every contract attributable to the run, including contracts missing + from ordinary telemetry after an ambiguous create; +4. destroy every discovered contract; +5. distinguish confirmed absence from provider query failure; +6. prove every exact contract absent before declaring cleanup successful; and +7. persist final accounting and all cleanup errors. + +A successful destroy request is not proof of absence. Continued typed provider +absence is the spending-stop condition. + +Cleanup discovery MUST be constrained to this run's expected labels and known +contract identities. It must never destroy or persist unrelated account +resources. Duplicate or ambiguous matches fail accounting but are still cleaned +to stop spend. + +Credentials MUST never appear in arguments, URLs, manifests, telemetry, +artifacts, or durable diagnostics. Outside an explicitly retained development +fixture, cleanup-only recovery after runner, supervisor, or workstation failure +MUST be able to finish exact discovery, destruction, and absence proof, but +cannot search offers, acquire nodes, bootstrap, or run cases. + +An explicitly retained development fixture may survive those failures only +while its cleanup owner, fixture-lifetime ceiling, and cost ceiling remain +active. Recovery may perform an explicit in-place binary redeployment or +provider cleanup; it may not acquire infrastructure or resume cases until the +full readiness and behavioral gates establish a new baseline. + +--- + +## 17. Failure handling summary + +| Failure point | Required result | +|---|---| +| Preflight or campaign validation | Zero provider acquisition; no cleanup needed | +| Cleanup-owner or orchestrator startup | Zero provider acquisition; stop started processes | +| Offer selection or admission | Zero provider acquisition | +| Provisioning, bootstrap, or readiness | Stop preparation; clean every possible contract | +| Case execution or oracle | Clean attempt; replay/shrink only from restored baseline | +| Local retained-node redeployment gate | Stop the local phase, clean its Docker fixture, and keep paid execution blocked | +| Attempt cleanup or health restoration | Quarantine fixture; skip further cases; clean provider | +| Formal orchestrator-recovery phase failure | Enter cleanup-only; never reacquire or repair | +| Explicit paid-fixture binary redeployment | Retain the exact contracts, pause cases, retry only in place, and reacquire nothing | +| Destructive transition | Do not continue to the next survivor campaign unless exact state is proven | +| Normal teardown | Continue best-effort cleanup and fail the run on any unresolved contract | +| Runner/supervisor/workstation loss outside an explicitly retained development fixture | Resume cleanup-only from durable run accounting | + +--- + +## 18. Pre-paid verification + +Pre-paid verification has two ordered gates. Gate B MUST NOT begin until Gate A +passes, and paid execution MUST NOT begin until both pass. + +### 18.1 Gate A: retained-node redeployment + +The local remote-shaped fleet and fault suite in Section 3.1 MUST pass first. +Its persisted evidence MUST identify every injected starting state and +interruption boundary and prove, for every redeployment round: + +- unchanged fixture resources and logical topology; +- matching installed and runtime deployment identity; +- absence of stale and duplicate workers; +- independence of worker lifetime from SSH lifetime; +- restored Iroh/Swactor membership, routing, readiness, and telemetry; +- successful post-redeployment all-pairs behavior; and +- complete local fixture cleanup. + +### 18.2 Gate B: remaining local and scripted verification + +Only after Gate A passes, local real-binary, mock-provider, and +scripted-provider checks MUST prove these scenarios: + +1. the complete five-node local real-binary campaign shape passes after the + redeployment suite, not instead of it; +2. the generator, independent model, oracle, replay, and shrinker agree on + fixed vectors and representative generated cases; +3. every control and telemetry contract round-trips through the shared + versioned representation; +4. offer filtering selects five verified distinct hosts by the required price + policy and rejects insufficient or malformed results before acquisition; +5. cost, lifetime, image, credential, and campaign-resource admission failures + produce zero provider acquisition; +6. every provider-capable path shares the same five-node acquisition bound, and + retries, replacement attempts, recovery, replay, shrinking, destructive + phases, and retained-node redeployment cannot exceed it; +7. connection retries do not become extra initial logical bootstrap sessions + or overlapping binary-redeployment transactions; +8. failures and injected crashes at each preparation boundary preserve enough + accounting to clean every possible contract; +9. outside an explicitly retained development fixture, runner, orchestrator, + cleanup-owner, and workstation failure paths enter cleanup-only behavior and + do not resume acquisition or initial bootstrap; +10. recovery adopts the prepared fixture with zero acquisition and preserves + the required namespace state while excluding active execution recovery; +11. all-pairs streams, rings, fan-in/out, joins, hot/cold fairness, and expected + branch failures terminate with complete observations; +12. concurrent observation collection recovers all required records without + using polling order as causal order; +13. every case failure path cleans before replay or shrinking; +14. real-provider health uses no local mock/container assumptions; +15. concurrent node and contract cleanup proves typed absence within the + cleanup deadline while preserving unrelated account resources; and +16. injected credential values are absent from all diagnostics and artifacts. + +A paid canary is allowed only after both gates pass in order. + +--- + +## 19. Paid canary acceptance + +The first paid run is accepted only when: + +- five verified distinct hosts are provisioned once; +- all five nodes converge, pass typed readiness, complete the all-pairs proof, + clean proof state, and commit the prepared fixture within ten minutes; +- acquisition accounting shows exactly five contract creations and five + initial logical bootstrap sessions, with no later acquisition; +- every retained-node binary redeployment, if exercised during development, is + explicit, preserves those exact contracts, has complete per-node transaction + accounting, and is followed by fresh readiness and all-pairs proof; +- the accepted coverage ledger contains results from one final deployment + identity only; +- selected worst-case cost remains within the operator ceiling; +- planned and observed coverage ledgers close; +- every attempt proves fresh preconditions and clean postconditions; +- recovery performs no acquisition or implicit binary redeployment and + preserves the specified state; +- destructive phases run against exact survivor sets and create no replacement; +- replay and shrinking, if exercised, remain inside the prepared fixture; +- cleanup proves every contract attributable to the run absent; and +- paid evidence is persisted as typed, versioned artifacts. + +A cleanup error fails the run even if all behavioral cases passed. + +--- + +## 20. Non-goals + +This work does not: + +- create a scheduler or recurring test service; +- prescribe source files, module boundaries, internal APIs, or implementation + steps beyond the required gate order; +- standardize identifier generation, concrete containers, frame layout, or + persistence internals beyond the observable guarantees in this plan; +- test unverified hosts; +- automatically replace failed nodes; +- promise survival of active contextual processes or active streams across an + orchestrator restart; +- treat degraded campaigns as fresh bootstrap coverage; +- bypass public application bindings in generated workloads; +- allow a later case to consume earlier case state; +- require storage, cache, model, dataset, or GPU state to survive a retained-node + binary redeployment; or +- recover a node whose root SSH or basic write-and-execute substrate is broken. + +--- + +## 21. Design-review checklist + +A readiness review MUST return PASS or FAIL, with behavioral evidence, for each +item: + +1. **Retained-node redeployment gate:** the local remote-shaped fleet converges + from every required corrupt state and interruption boundary without changing + fixture resources, duplicating workers, or depending on SSH after launch. +2. **Ordered advancement:** the redeployment gate passes before the remaining + local E2E suite, and every local and scripted gate passes before paid access. +3. **Bounded paid work:** no execution or failure path can acquire more than the + five selected nodes or leave a possible contract outside cleanup accounting. +4. **Prepared-fixture isolation:** every case begins fresh, ends at the recorded + baseline, and quarantines the fixture on restoration failure. +5. **Topology executability:** all required topology families, stream startup, + ring completion, joins, faults, and progress scenarios can terminate without + relying on a favorable schedule. +6. **Independent oracle:** expected data and legal concurrent outcomes are + derived independently and never selected by observation polling order. +7. **Zero-acquisition reuse:** recovery, replay, shrinking, regressions, + survivor campaigns, and explicit binary redeployment cannot provision, + repair, replace, or remap topology. +8. **Exact cleanup:** normal and crash-recovery paths discover every attributable + contract, preserve unrelated resources, and prove typed absence. +9. **Integration evidence:** the complete local redeployment suite, local + real-binary campaign, mock campaign, scripted failure and crash scenarios, + credential-leak checks, and cleanup gates pass in the required order before + paid execution. + +--- + +## 22. Execution discipline + +This plan is a verification checklist, not an implementation backlog. + +- Do not edit without a concrete, focused reproduction of a violated + requirement. +- Do not invent goals from unchecked requirements, speculative reviews, or + agent suggestions. +- Keep exactly one active blocker. Defer everything unrelated. +- The full campaign is final qualification, never the debugging loop. +- After a campaign failure, extract and run the exact failing case directly. Do + not rerun the campaign until that focused case fails before the fix and passes + after it. +- Freeze source, binary, and image identities before running one complete + ordered qualification. + +### Current state — 2026-09-14 + +- `target/ordered-final-3/ordered-acceptance.json` records successful five-node, + 12-round Gate A (331.13 s) and complete local campaign (274.08 s). +- The complete warm workflow failed at 605.23 s against its 600 s ceiling, + before the remaining Gate B checks. This is the active demonstrated blocker. +- A separate failure-case run passed in 14.39 s; it is not ordered acceptance. +- Review-time current-source checks passed: 169 harness library tests, + 5 shared-contract tests, 4 contextual-process tests, 4 telemetry transport + tests, and 4 attestation-guard tests. +- The old active-stream replacement failure is historical; do not carry it + forward as a current blocker without a new reproduction. +- Release executable hashes matched the retained ordered run at review time, + but the source digest did not. No current identity has successful complete + ordered qualification. +- No remote or paid execution was performed during the checkpoint review. +- Continue work using the remote and local handoffs linked above. Local + artifact paths are not part of the committed evidence and must be retained + separately. \ No newline at end of file diff --git a/docs/specs/drafts/VASTAI_LOCAL_BLOCKERS.md b/docs/specs/drafts/VASTAI_LOCAL_BLOCKERS.md new file mode 100644 index 0000000..384acfa --- /dev/null +++ b/docs/specs/drafts/VASTAI_LOCAL_BLOCKERS.md @@ -0,0 +1,137 @@ +# VastAI checkpoint: local blockers + +Reviewed 2026-09-14. This is the local-work handoff, in priority order. +The [remote-ready outline](VASTAI_REMOTE_READY.md) describes the work that can +proceed independently on a frozen checkpoint. The [E2E plan](VASTAI_E2E_FUZZ_PLAN.md) +remains the acceptance contract; this report does not change runtime limits or +paid-access guards. + +## Decision + +The main demonstrated local blocker is end-to-end qualification time, not an +unfinished deployment system or a campaign that cannot complete. Finish the +remaining safety verification and qualify a frozen identity after addressing +that blocker. Do not reopen historical failures without a current reproduction. + +## 1. Make the complete ordered workflow fit its time envelope + +**Status: demonstrated failure; highest priority.** + +`target/ordered-final-3/ordered-acceptance.json` records: + +| Stage | Result | Elapsed | +|---|---|---:| +| Five-node, 12-round Gate A | PASS | 331.13 s | +| Complete local campaign | PASS | 274.08 s | +| Combined warm workflow at that point | FAIL | 605.23 s | + +The workflow limit is 600 seconds. It expired before failure-case, contract/model, +and scripted-provider checks ran. Saving only 5.23 seconds is therefore not +sufficient: the remaining required stages also need time within that envelope. + +### Next work + +1. Use the retained stage/round timing evidence to identify the dominant work. + Keep build/cache preparation separate from measured warm execution. +2. Reproduce the slow stage or operation directly. Make changes only against a + measured bottleneck, not a speculative broad runtime rewrite. +3. Verify the focused path after each fix. Preserve campaign shape, real public + bindings, independent observations, fault coverage, and cleanup guarantees. +4. Measure the remaining Gate B stages to establish the headroom actually needed. + Full ordered execution is final qualification, not the profiling loop. + +### Completion condition + +The complete warm workflow, including every required Gate B stage, passes in +600 seconds; the campaign remains within its 300-second envelope. Obtain the +required three consecutive warm passes rather than accepting a shortened run. + +## 2. Establish the remaining safety and failure-path evidence + +**Status: verification gap, not a demonstrated current functional failure.** + +The newest ordered run stopped before its failure-case, contract/model, and +scripted-provider stages. There is a separate successful failure-case run and +passing focused tests, but these do not establish complete ordered Gate B. + +### Next work + +Use the existing checks to verify: + +- admission rejection before acquisition, exact five-node acquisition bounds, + and offer/host selection; +- preparation failures and injected crashes with durable contract accounting; +- cleanup-owner recovery, typed contract absence, unrelated-resource + preservation, and credential redaction; +- case failure cleanup, replay/shrink isolation, and recovery without acquisition. + +Run the scripted provider only through its existing loopback-restricted path; +no paid resources are needed for this local work. Reuse +`tools/myelin-e2e-fuzz/scripted_safety_gate.sh` and the current contract tests. +Focused safety checks are diagnostics until included in ordered qualification. + +If a check fails, capture its exact reproduction and make that the active local +bug. Do not label unexecuted scenarios as known broken behavior. + +### Completion condition + +The full existing failure/safety checks pass with retained artifacts and are +included after Gate A in the final ordered run. A successful destroy request +alone is not cleanup proof. + +## 3. Qualify a frozen checkpoint and hand it to paid testing + +**Status: integration prerequisite after the first two items.** + +At review time the release binaries matched the retained ordered evidence, but +the source digest did not. The ordered record itself is failed, not an +attestation that can authorize paid work. + +### Next work + +- Freeze the final source checkout, deployment artifacts, and runtime images. +- Use `tools/myelin-e2e-fuzz/ordered_acceptance.py` for the full ordered workflow; + keep its identity/provenance checks and required stage sequence intact. +- Record all three successful warm runs and separately handle the plan's cold + timing requirement. Do not report a warm-cache build as cold evidence. +- If source, binaries, or images change, create new qualification evidence. + Independent remote diagnostics must not mutate this frozen checkout. +- Hand the remote track the exact qualified checkpoint, image identity, and + fresh attestation. Keep real-provider execution subject to the existing + admission and cleanup requirements. + +### Completion condition + +A successful, fresh, identity-matching ordered attestation is available for the +paid path. The current CLI requires at least three warm runs and an attestation +valid for no more than 24 hours. Qualification completion and paid launch must +be coordinated; a stored historical pass is not permanent authorization. + +## What not to work on without new evidence + +- The old active-stream replacement failure is not the current blocker. A newer + complete local campaign passed. Reopen it only on a current failing case. +- Do not start another broad runtime, transport, or harness redesign merely + because the final qualification is incomplete. +- Do not remove safety gates or reduce scenario coverage to obtain a pass. +- Keep unrelated tooling and UI polish outside this critical path. + +## Verification already performed for this checkpoint + +All of these current-source checks passed during the review: + +```text +cargo test --locked -p myelin-e2e-fuzz -p myelin-control-contract --lib + 169 harness tests + 5 shared-contract tests +cargo test --locked -p myelin --lib contextual_process_guarantees + 4 tests +cargo test --locked -p iroh-driver --test telemetry_transport + 4 tests +cargo test --locked -p myelin-e2e-fuzz --bin myelin-e2e-fuzz ordered_gates_ + 4 tests +``` + +Total: 186 passing tests. These are focused checks, not the complete workspace +suite or new ordered acceptance. No remote resources or paid execution were +started during review. Historical `target/` evidence is local to the development +machine and must be preserved separately from these committed reports. diff --git a/docs/specs/drafts/VASTAI_REMOTE_READY.md b/docs/specs/drafts/VASTAI_REMOTE_READY.md new file mode 100644 index 0000000..0e9984d --- /dev/null +++ b/docs/specs/drafts/VASTAI_REMOTE_READY.md @@ -0,0 +1,112 @@ +# VastAI checkpoint: remote-ready work + +Reviewed 2026-09-14. This is the remote-work handoff, not a paid-run authorization. +See [local blockers](VASTAI_LOCAL_BLOCKERS.md) for the parallel local track and +[the E2E plan](VASTAI_E2E_FUZZ_PLAN.md) for the acceptance contract. + +## Decision + +The deployment, runtime, and campaign machinery is implemented far enough to +move a frozen checkpoint onto an existing remote development host and test it +there. Do not wait for all local qualification work before collecting useful +remote evidence. + +There are two different milestones: + +- **Ready now: remote-host diagnostics.** Run the existing Docker/SSH fixture and + real-binary workloads on infrastructure already available for development. +- **Not yet authorized: paid VastAI execution.** This still requires successful, + fresh ordered qualification and the existing paid admission/cleanup guards. + Moving a run to another machine does not bypass those requirements. + +## What is ready to exercise + +| Area | Remote work | Existing evidence | +|---|---|---| +| Deployment and retained-node redeployment | Install new bundles, replace prior worker state, interrupt SSH/orchestrator boundaries, and verify fresh membership without replacing nodes. | Five-node Gate A passed all 12 rounds, all-pairs behavior, and fixture cleanup. | +| Contextual processes and data paths | Exercise public Python bindings, process launch/stop, namespace/blob/stream behavior, and resource reclamation. | Complete local campaign plus current focused process-lifecycle tests. | +| Prepared-fixture reuse | Keep the fixture while running fresh cases, orchestrator recovery, and exact four-node/three-node survivor phases. | 128 normal, 16 recovery, 32 four-node, and 32 three-node cases completed: 208 logical cases / 224 workload segments. | +| Telemetry transport | Observe startup records, reconnect/replay, cancellation, and consistency between worker identity and collected observations. | Current transport tests passed catalog/frame delivery, compressed-frame validation, cancellation, and replay without duplicates. | +| Failure handling | Exercise the existing failure-case workload and inspect terminal observations and cleanup. | A separate local failure-case run records success; it is not complete ordered Gate B evidence. | + +These are working test candidates, not claims of stability on real VastAI hosts. +Provider offer selection, real network conditions, and provider-side cleanup +still need live-provider evidence after admission is authorized. + +## Remote track: execution outline + +### 1. Freeze and prepare + +- Use the committed checkpoint in a separate checkout from ongoing local edits. +- Build and record exact source, executable, deployment-bundle, and image + identities. Do not assume the development machine's cached binaries belong to + the new checkout. +- Use an existing development host with the required Docker, SSH, image, and + networking capabilities. The static-SSH provider manages pre-created Docker + nodes; it is not a generic adapter for arbitrary SSH machines. +- Keep the existing fixture ownership and cleanup machinery. Do not introduce a + second deployment harness or manually install binaries inside tested nodes. + +### 2. Exercise the established paths + +Start with the five-node retained-node redeployment gate. After it passes, +exercise the complete campaign and focused failure cases. Preserve exact node +sets, public-binding workloads, and cleanup checks. + +Measure preparation, redeployment rounds, campaign phases, and cleanup +separately. Useful findings include remote-only connectivity failures, stale +membership after redeployment, missing telemetry, process-lifetime problems, +and stages that dominate elapsed time. + +An isolated diagnostic run is useful even when it is not ordered acceptance. +Label it as diagnostic; do not merge its coverage into a later accepted run. +Do not use another full campaign as the debugging loop after a failure: retain +and reproduce the exact failing case or deployment round. + +### 3. Hand off failures without moving the baseline + +For each result, record: + +- checkpoint/source, executable, image, and deployment identities; +- host environment, gate or case, exact node set, and artifact location; +- PASS/FAIL, stage timings, and cleanup outcome; +- for failure, the first failing observation and smallest known reproduction. + +Continue local fixes in the separate local checkout. Promote them to the remote +track only as a new checkpoint, with fresh identities and evidence. Do not +silently update binaries in the middle of an accepted campaign. + +### 4. Advance to real VastAI only after qualification + +The paid path requires fresh successful `ordered_acceptance.py` evidence, +matching artifacts and immutable image provenance, eligible distinct-host +offers, explicit cost/lifetime limits, credentials supplied through the existing +private configuration, and active durable cleanup ownership. + +The current checkpoint has no such successful attestation. Leave the gate +checks intact. Do not substitute the static-SSH fixture, a scripted-provider +pass, or the historical local campaign for paid authorization. + +## Evidence and limits + +Development-machine artifacts inspected for this checkpoint: + +- `target/ordered-final-3/ordered-acceptance.json`: Gate A passed in 331.13 s; + campaign passed in 274.08 s; the complete workflow failed at 605.23 s before + later Gate B checks. +- `target/ordered-final-3/warm-1/gate-a/deployment-e2e/gate-a-evidence.json`: + five nodes, 12 rounds, successful behavior and complete cleanup. +- `target/ordered-final-3/warm-1/campaign/vastai-e2e-000000000135282e/`: + coverage ledgers and checkpoint account for all campaign phases. +- `target/gate-b-failure-cases/timing-runs/1789241279248-532.json`: + separate failure-case run passed in 14.39 s. + +At review time all three release executable hashes matched the ordered-run +record, but the current source digest did not. These are historical execution +results, not current-tree qualification. The artifact paths are local build +outputs and are not included in Git; retain or transfer the evidence explicitly +when handing off a run. + +Current-source verification during review: 169 harness library tests, 5 shared +contract tests, 4 contextual-process tests, 4 telemetry transport tests, and 4 +attestation-guard tests passed. No new remote or paid run was performed.