swactor/examples/pipeline-parallel-inference/N3_COVERAGE_EXTENSION_SPEC.md

521 lines
24 KiB
Markdown
Raw Normal View History

# N=3 collection coverage extension — behavioral spec
Companion to `N3_POSTMORTEM_2026-05-25_1779733878.md`,
`N3_SIM_TEST_BATTERY_SPEC.md`, and the simulator's `SIM_SPEC.md`.
This document is the contract for a separate coding agent to extend
diagnostic collection coverage along three layers — **production
diagnostics**, **simulator emit/model**, and **simulator test
verification** — for the gaps the `1779733878` run surfaced.
This is a *behavioral* spec. It names the gap, the contract the
collected data must satisfy, and the layer(s) the contract threads
through. It does not prescribe field names, file layout, or
implementation choices.
---
## 0. Motivation
The `1779733878` run validated the prior observability upgrade —
A+B+C tiers were load-bearing, the bundle attributed the failure to
"dials to orchestrator fail by timeout 7/11 while every inter-stage
dial succeeds 19/19" in one table — and surfaced five residual
collection gaps the upgrade either left as carry-forward (gaps 1, 5,
8 from the original scorecard) or that this run exposed for the
first time (response-leg event absence, bundle-serve behavior under
run-id reuse).
A gap whose collection landed only in production but not in the sim
is a gap that the sim test battery can never guard — the next regression
in that field will be caught only by another live deploy. A gap whose
sim model exists but is not exercised by a test is dead code. The
coverage in this spec is required to thread through every layer
where it can — and the spec is explicit when a layer does not
apply.
Five coverages, each threaded through up to three layers. Each
coverage may close one of {Pass, Mixed, Fail} against the current
source, and each names the close criterion.
---
## 1. Cross-cutting requirements
These hold for every coverage in §2.
### 1.1 Three-layer threading
For each coverage, the spec names which of the three layers it
threads through:
- **D — Diagnostics**: the production bundle gains a field, event,
or section that closes the gap the postmortem named.
- **S — Sim**: the simulator's relevant component (host kind,
network, relay vertex, bundle writer) emits the same field /
event / section under the same conditions, with bundle-shape
parity per `SIM_SPEC.md §5` (cross-cutting "Bundle-shape parity
with prod") and §9 (bundle schema).
- **T — Tests**: the sim test battery gains a scenario or property
test asserting the bundle carries the new data when the
triggering condition holds, and gains a discriminator assertion
when absence is meaningful (per the honesty-under-absence pattern
the prior upgrade established).
A coverage that threads through fewer than three layers is honest
about which it skips and why. Skipping S because "the simulator
does not model this surface" is acceptable; skipping it because
"this is not interesting" is not.
### 1.2 Honesty-under-absence carries forward
The prior upgrade's `status_source` discriminator pattern (a status
field always paired with a field naming how that status was derived
— `"iroh"` for native, `"derived"` for inferred) is the model. Any
new field whose value might be absent or derived must carry an
adjacent discriminator. A bundle reader must never be left guessing
"unknown means the thing is unknown" vs "we couldn't ask."
### 1.3 Additive evolution
Every new field on `SnapshotBody`, every new event variant, every
new section in the post-processor output is additive. An old bundle
reader on a new bundle still parses; a new bundle reader on an old
bundle reports the new field absent rather than erroring. The
prior upgrade established this contract; coverage 2.x preserves it.
### 1.4 Sim/prod schema parity
Per `SIM_SPEC.md §9.2`: the event payload schema is exactly the
production diagnostics schema for that kind. The sim invents no new
event kinds. A coverage that lands an event in prod and in sim
**uses the same schema in both**, verified by the existing parity
tests under `crates/simulation/tests/sim_cross_pollination.rs`. A
schema added to sim ahead of prod is a deliberate amendment and
declares so explicitly.
### 1.5 Verdict-first per layer
Every coverage in §2 declares, per layer, its expected status on
the current source: **landed** (the layer satisfies the contract;
the work is verification / regression-guard), **partial** (the
layer has structure but not data flow), **absent** (the layer has
nothing today). The implementing agent's work is to bring each
layer to "landed" against this spec or to file a structural reason
why a layer cannot land.
---
## 2. The coverages
Five, ordered by the postmortem's own ranking of residual gaps.
### 2.1 Orchestrator-side host provider metadata forwarding
**Source**: postmortem §"Observability upgrade scorecard" row "gap
5 host metadata" (◐); postmortem §"Data-collection / deployment
gaps surfaced by this run" item 1.
**Gap**: The orchestrator has each rental's public IP, datacenter,
country, and contract id at `lease_chain` return time. The
container can read these from `SWACTOR_DIAG_*` env vars. The
container env is never set. The boot record's
`host_ip_public`/`datacenter_id`/`host_country`/`vastai_contract_id`/
`home_relay_url_at_boot` fields are still null in every bundle.
The contract id arrives but in `container_id`, not
`vastai_contract_id` — the naming is currently load-bearing-but-wrong.
**Contract — what closing the gap looks like**:
- D: the orchestrator's per-rental env payload, at the point it
creates each container, carries every field the boot record can
consume — public IP, datacenter id, host country, vast.ai
contract id, the home relay URL the container will use. The boot
record reflects every field as a concrete value, not `null`,
whenever the orchestrator had the data. The fields that name
cloud-provider state stay absent only on hosts where they
genuinely do not apply (e.g., local development), and the
bundle's `## Hosts` section renders `?` for absent fields
(already implemented per `S-A2`).
- S: scenarios declare per-peer host context as part of the peer's
`kind_config`. The sim's stage host populates its boot record /
`HostContext` from the scenario declaration the same way prod
populates from env. A scenario without declared host context
produces a bundle whose `## Hosts` section is all-`?` for that
peer — same absence shape as a local-dev prod bundle.
- T: a scenario declaring heterogeneous host context across three
peers (e.g., two datacenters, two countries) produces a bundle
whose `## Hosts` section renders the declared fields verbatim.
A scenario that declares no context for one peer and full context
for the others produces a bundle distinguishable from "no context
declared for any peer" by the `?` placement.
**Expected status**:
- D: partial. The container reads the env vars (per `S-A2`); the
orchestrator does not set them. The misnaming of contract id →
`container_id` is a separate cleanup.
- S: absent. The sim's stage host carries no host context in its
current scenario schema.
- T: absent. No test exercises this discriminator.
**Close criterion**: a deployed bundle's `## Hosts` section names
the datacenter, country, public IP, and contract id of every
vast.ai rental, and the docker `container_id` field carries the
docker container id, not the vast.ai contract id. A sim bundle
with declared host context produces the matching shape.
### 2.2 Relay-port reachability probe
**Source**: postmortem §"Observability upgrade scorecard" row "gap
8 relay-port probe" (✗); postmortem §"Data-collection / deployment
gaps surfaced by this run" item 3.
**Gap**: Stage probe arrays carry only `collector_udp_echo`
(:9081). No probe targets the relay's actual port (:7843).
Whether a stage retained transport-level reachability to the relay
at the moment its peer-connection died is currently inferable only
from a *different* port on the same host. The `S-E1` work is
documented as landed (per the prior iteration log) but the
`1779733878` bundle shows no relay-port probe records. The wiring
is in place; the data is not.
**Contract — what closing the gap looks like**:
- D: every stage's snapshot carries a probe outcome for the
relay's UDP listener (host + port resolved from the home relay
URL). The outcome is one of the five-discriminator set the
prior upgrade established: `ok` / `timeout` / `refused` /
`unresolved` / `error`. A snapshot taken when the relay is
reachable carries `ok` with an RTT; a snapshot taken when the
relay is unreachable carries the appropriate failure
discriminator with no silent fallback to "absent."
- S: the sim's stage host emits the same probe record on every
snapshot, sourced from a query the network answers about the
stage→relay edge. The relay vertex's `RelayKill` /
`RelayCapacityChange` mutations are reflected in the probe's
outcome distribution.
- T: a scenario that issues a `RelayKill` mutation mid-run
produces a bundle whose every stage's relay-port probe outcome
flips from `ok` to `unresolved` (or `timeout`, per the
network's policy) at the mutation's `at_ns` and remains there
through `RelayBoot`. The probe-outcome timeline is the test's
discriminator between "tunnel down" and "tunnel up but peer
conn down" — coverage 2.x.A from the battery spec consumes
this signal.
**Expected status**:
- D: partial. Probe scheduler wires the target; emission to the
bundle is unverified by this run's evidence.
- S: absent. The sim's network has no probe-query surface today.
- T: absent.
**Close criterion**: the next deployment's bundle has a
relay-port probe outcome on every stage's snapshots. A sim
scenario with `RelayKill` produces the probe-outcome flip in the
bundle.
### 2.3 Relay session lifecycle on the relay side
**Source**: postmortem §"Observability upgrade scorecard" row "gap
1 relay observability" (◐); postmortem §"Data-collection /
deployment gaps surfaced by this run" item 4.
**Gap**: The relay reports identity and 186 snapshots into the
bundle but cannot answer "who closed session X and why" — the
per-session lifecycle hooks are the documented skeleton with
`active=0 opens=0 closes=0`. `iroh_relay::server` exposes no
session hooks. Until it does, a relay-side eviction is
unanswerable from the relay's own data; the postmortem fell back
to node-side dial outcomes.
**Contract — what closing the gap looks like**:
- D: the relay's bundle contribution names, per peer session, the
open time, close time, close-initiator discriminator
(`relay` / `peer` / `transport` / `unknown`), close reason
string (relay-specific or transport-specific), bytes
transferred per direction, and duration. The mechanism is
free — middleware around the relay binary, kernel-layer
observation, a forked relay, or upstream hooks when iroh
exposes them. The contract is the *shape*, not the source.
When the source is unavailable, the relay's bundle
contribution still emits the gap-1 absence-line the prior
upgrade introduced in `summary.md` (the post-processor's
acceptance branch for "no relay-role node has session data").
- S: the sim's relay vertex emits `RelaySessionOpened` /
`RelaySessionClosed` records when it accepts and releases
per-peer queues. The records carry the same shape D
requires. A `RelayKill` mutation produces a
`RelaySessionClosed { initiator: "relay", reason: "killed",
... }` for every session active at the mutation time.
- T: a scenario where the relay accepts three peer sessions, runs
to steady state, then receives a `RelayKill` mutation,
produces a bundle whose relay contribution names three
`RelaySessionOpened` events at the convergence boundary and
three `RelaySessionClosed { initiator: "relay" }` events at
the mutation time. A scenario where a peer voluntarily
disconnects produces a session-closed event with
`initiator: "peer"`. The discriminator must hold.
**Expected status**:
- D: skeleton — wired call sites, no data flow. Whether the
unblock path is upstream hooks, middleware, or kernel
observation is implementer's call.
- S: partial. `RelayObservability` exists on the host side per the
prior upgrade (`S-B1`); the sim's relay vertex itself does not
emit lifecycle events as engine-synthesized records.
- T: absent.
**Close criterion**: a deployed bundle from a run that included a
peer dial failure attributable to a relay-side close names the
close-initiator and reason in the relay's bundle contribution. A
sim `RelayKill` scenario produces the matching event stream.
### 2.4 Inference response-leg instrumentation
**Source**: postmortem §"Data-collection / deployment gaps
surfaced by this run" item 5.
**Gap**: The `1779733878` postmortem's conclusion — "last stage
could not deliver the response" — was inferred from dial timeouts
plus the absence of an inbound `InferenceResponse`, not from a
typed event on the last stage saying "I tried to send the response
and the send outcome was X." The chain `stage-(N-1)
→ InferenceResponse → orchestrator's inbox` has no event on the
sending side. A typed event makes attribution a one-line read
rather than a triangulation.
**Contract — what closing the gap looks like**:
- D: the production stage actor, on attempting to send an
`InferenceResponse` upstream, emits a typed event naming the
target peer, the request id the response corresponds to, the
byte size, and the send outcome. The outcome discriminator is
the iroh-level result the transport returns (succeed / timeout
/ connection-closed / refused / unresolved / queued-but-not-
acked-in-budget). The post-processor surfaces these in
`summary.md` under a section that names which inference
request was answered by which stage's send and how that send
resolved.
- S: the sim's stage host kind grows a minimal inference
message surface (`InferenceRequest` inbound to stage-0,
`InferenceResponse` outbound from stage-(N-1), forwarded
between adjacent stages as opaque payload in the MVP). The
stage host emits the same typed response-send event when it
attempts the outbound to the orchestrator. The codec contract
(`SIM_SPEC.md §3.3`) carries the inference messages with
byte-equality between sim and prod encoding.
- T: a scenario where the orchestrator's inbound path is broken
via `RelayPeerConnDown` on the last leg (stage-(N-1) → orch)
while every other leg works produces a bundle whose last
stage emits exactly one `InferenceResponseSent` event with
`send_outcome` in the failure-discriminator set. A scenario
where every leg works produces an `InferenceResponseSent`
with `send_outcome=success` and a matching
`InferenceResponseReceived` (or equivalent) on the
orchestrator's side.
**Expected status**:
- D: absent. The current stage actor's send call is not wrapped
in a typed diagnostic event for the response leg.
- S: absent. The sim's stage host kind today produces no `Send`
actions during its lifecycle (`SIM_SPEC.md §6A.5` notes this
explicitly and defers inter-stage traffic to a later revision).
Closing this coverage moves that deferral forward.
- T: absent.
**Close criterion**: a deployed bundle from any run where the
response did not return names the send outcome of the last
stage's response attempt in a single event. A sim scenario
modeling the same failure produces the same shape.
### 2.5 Bundle serve hardening under run-id reuse
**Source**: postmortem §"Bundle recovery" caveat; postmortem
§"Data-collection / deployment gaps surfaced by this run" item 2.
**Gap**: When a run id is reused across the failed-first-lease /
successful-second-lease shape the `1779733878` run exhibited, a
finalize record from the first phase pins a stale canonical
bundle in the collector's cache. A subsequent `GET` serves the
stale 5.3 KB bundle instead of synthesizing the rich 9.3 MB one
from current staging. Two adjacent quirks: `finalize_received`
stays `true` after the on-disk `finalize-*.json` is deleted, and
the synthesized manifest still lists a removed node directory.
**Contract — what closing the gap looks like**:
- D: the collector's `download_bundle` handler prefers the
*richer* of {canonical-cached, synthesized-from-current-staging}
by a size or node-count heuristic, or rebuilds canonical when
staging has grown past the cached bundle's manifest. Deleting a
node directory from staging clears the corresponding finalize
record from in-memory state. The synthesized manifest reflects
the current on-disk state, never a stale in-memory record. The
`finalize_received` boolean is sourced from the same place the
serve decision is sourced from — a single source of truth, not
two diverging caches.
- S: not applicable. The sim writes bundles directly to a
destination directory; there is no serve logic, no finalize
cache, no run-id reuse semantics. The coverage threads through
D only.
- T: not applicable as a *sim test*. The discriminator (stale vs
fresh serve on a finalize-then-staging-growth sequence) is a
collector unit-test concern living under
`crates/distribution/tests/`, not a scenario the sim engine
can express. The implementing agent should land the collector
test alongside the D-layer change; it is named here so that the
coverage's verification surface is honest about where it lives.
**Expected status**:
- D: absent. Current serve logic prefers cached canonical
unconditionally when `finalize_received` is true.
- S: not applicable.
- T: collector unit test absent.
**Close criterion**: a collector unit test writes two phases of
staging with an intervening finalize, deletes the first-phase
node, and verifies the second `GET` serves the richer bundle and
that the cleared node does not appear in the manifest.
---
### 2.6 Per-SWIM-probe RTT and observed latency distribution
**Source**: postmortem §"SWIM churn and relay events" (1701
SwimTransitions over ~7 min, all with `conn_type=Relay`); postmortem
§"UDP echo probes" (tier-2 RTTs spread 181–405 ms; SWIM probes
ride a relay-mediated path on top of these); `SWIM_TUNING_REPORT.md`
§6 limit 3 ("SWIM host adapter does not emit `probe_sent` /
`probe_received` / `probe_timed_out` events").
**Gap**: The bundle has tier-2 UDP-echo RTT to docean:9081 — a
host-level surface that does not represent the latency SWIM
actually sees. SWIM rides a relay-mediated peer connection whose
RTT is at least one extra hop and is subject to relay-side HOL
queueing under load. The bundle currently exposes:
- per-snapshot iroh counters (cumulative `MessageSent` /
`MessageReceived`),
- aggregate `SwimTransition` counts,
- per-peer dial outcomes (`Timeout` / `Success` rollup),
but it does not expose per-probe RTT, per-peer RTT distribution
over the run window, or correlation between
`probe_timed_out`-class outcomes and observed RTT spikes. Without
this surface, SWIM tuning is a guess against the deploy's actual
latency distribution rather than a measurement.
This gap also mirrors the simulator's own limit per
`SWIM_TUNING_REPORT.md` §6.3: the SWIM host adapter does not emit
the probe lifecycle events, so the §10 evaluator's
`no_flap_while_probes_ok` is structurally `Inconclusive`. Closing
the gap on both sides closes the assertion's precondition.
**Contract — what closing the gap looks like**:
- D: each SWIM ping/ack pair emits a typed event naming the
observer, target, virtual-or-wall send time, virtual-or-wall
receive time, the resulting RTT, and the discriminator
(`success` / `timeout` / `connection-closed` / etc.). The
post-processor surfaces a `## Probe RTT distribution` section
with median, p95, p99 per (observer, target) pair, plus per
five-second bucket so degradation over time is visible. A
`probe_timed_out` outcome carries the configured timeout
budget alongside the observed RTT (where one exists) so a
reader sees "probe missed a 3 s budget by 200 ms" vs "no
response within 3 s, never arrived."
- S: the simulator's SWIM host adapter emits the same probe
lifecycle events. Per `SIM_SPEC.md §9.2` parity, the schema is
identical to D's. This is the §6.3 limit from
`SWIM_TUNING_REPORT.md` closing simultaneously with D — the
bundle reader cannot tell a sim run from a prod run by this
surface.
- T: a scenario with a declared per-link latency distribution
(heavy-tailed, peer-symmetric) produces a bundle whose
postproc RTT section's median, p95, p99 fall within stated
tolerance of the scenario's declared distribution. A scenario
with a `LatencySpike` mutation produces a bundle whose RTT
section shows the spike at the mutation time. The
precondition for `no_flap_while_probes_ok` is now satisfied;
the assertion moves off `Inconclusive` for every scenario
using a SWIM-host kind.
**Expected status**:
- D: absent. No per-probe event today.
- S: absent. `SWIM_TUNING_REPORT.md` §6.3 names this explicitly.
- T: absent.
**Close criterion**: a deployed bundle's postproc summary names
the median / p99 RTT per (observer, target) and a sim bundle
produces the matching surface. `no_flap_while_probes_ok` resolves
to `Pass` or `Fail` (not `Inconclusive`) on every SWIM scenario in
the calibration library.
**Downstream**: this coverage is the data surface
`N3_SWIM_TUNING_SPEC.md` consumes. SWIM tuning itself is
downstream of collection and lives in that sibling document.
---
## 3. Out of scope
- **Inference protocol surface beyond the response leg.** Coverage
2.4 instruments the response-send event. A full inference-
protocol event stream (microbatch routing, KV cache, per-stage
worker activity) is broader than what the `1779733878` postmortem
could not answer; it belongs in a separate spec when a
postmortem demands it.
- **Post-processor summary enhancements.** SWIM transition
distributions, per-(observer, target, reason) breakdowns,
cross-node temporal alignment around the moment of failure —
these are renderer concerns, not collection concerns. They
presuppose the data is in the bundle; this spec is about the
data.
- **Orchestrator-topology fixes.** The `1779733878` postmortem's
item 6 names the root cause as a NAT'd local orchestrator with
no reachable port. That is a deployment-shape question for the
runbook, not a collection-coverage question.
- **Runbook fixes.** The `--gpu RTX_4090` vs `RTX 4090` line in
`DEPLOYMENT_TEST.md` (postmortem item 7) is a runbook bug, not a
collection gap.
- **Sim coverage of upstream-blocked surfaces.** If
`iroh_relay::server` continues to expose no session hooks, the
sim's relay vertex can model the lifecycle events the contract
requires, but the production D layer of coverage 2.3 may remain
partial. That partiality is a structural blind spot to file per
the established blind-spot discipline; this spec does not
resolve it.
---
## 4. References
- `N3_POSTMORTEM_2026-05-25_1779733878.md` — the second
2026-05-25 deployment's postmortem. §"Observability upgrade
scorecard" is the source for coverages 2.1, 2.2, 2.3; §"Data-
collection / deployment gaps surfaced by this run" items 1–5 map
to coverages 2.1, 2.5, 2.2, 2.3, 2.4 respectively.
- `N3_SIM_TEST_BATTERY_SPEC.md` — the sim-test battery spec. The
battery's families A (relay peer-conn down) and the discriminator
it builds against the relay-port probe (coverage 2.2) and the
relay session lifecycle (coverage 2.3) consume the data this
spec lands.
- `crates/simulation/SIM_SPEC.md` — the simulator's behavioral
surface. §3.3 codec contract, §5A relay vertex, §6A stage host
kind, §9 bundle layout are the load-bearing references for the
S-layer contracts.
- `crates/simulation/SWIM_TUNING_REPORT.md` — the prior tuning
pass against simulated 60 ms latency. §6 limits (especially
§6.3 "SWIM host adapter does not emit `probe_sent` /
`probe_received` / `probe_timed_out` events") are the source
for the S-layer of coverage 2.6.
- `N3_SWIM_TUNING_SPEC.md` — the downstream spec that consumes
coverage 2.6's data surface to retune SWIM against the
observed `1779733878` latency distribution. Sibling document.