Relay/subprocess/registry introspectors, build-time dep-version + git-sha, postproc RTT/inference render, new swactor-iroh-relay binary; dashboard replay-viewer, sim SWIM tuning report + repro tests, N3 postmortem docs. Signed-off-by: Zachery Aaron Shores-Chmielewski <zacheryasc@gmail.com>
17 KiB
SWIM Tuning Report
Short summary
The simulator was used to tune SwimConfig::default() against the
§10.3 gossip-flap library property and the three N3 calibration
scenarios. New operating point: probe_interval=10, probe_timeout=15, suspicion_timeout=75, indirect_probes=2, dead_reprobe_interval=50
ticks plus max_piggyback=6. On the property — 3-peer mesh, 60 ms
latency, 15 ms jitter, 0.5 % loss, 20 s window — peak
self_incarnation falls from 86–94 to 7–10, an
order-of-magnitude collapse of the refute storm.
relay_queue_depth_bounded and message_size_bounded pass on
own-relay with margin; convergence_after holds at the 2 s baseline;
dead_peer_resurrects_within is not declared anywhere and does not
regress. The canary scenario's placeholder relay topology was
calibrated (egress 1 Mb/s → 100 bps/link, queue bound 65 536 B →
1 500 B) so the Layer A buffering it was supposed to capture actually
fires; pre- and post-tuning, canary still FAILS
relay_queue_depth_bounded. The Layer B1 refute-on-stale-Suspect bug
in swim/node.rs::apply_membership_update caps the gossip flap above
the algorithmic ideal of 2; tuning collapses the storm but cannot
remove the floor. name_resolves_within remains FAIL on every SWIM
observer because the SWIM-host adapter publishes no name registry —
a simulator limit, not a protocol one.
Detailed report
1. What "optimized" meant going in
Targets, decided up front and unchanged after measurement:
no_flap_while_probes_okon both own-relay scenarios. Inconclusive in the library today (the SWIM host adapter does not emitprobe_sent/probe_receivedevents the assertion keys off), so this collapses operationally to do not regress the assertion precondition. It does not.self_incarnation_boundedpasses with a justified bound. Tuned against the §10.3 property; per-scenario bounds are set to what tuning actually achieves on each scenario's traffic shape. See §4.message_size_boundedandrelay_queue_depth_boundedpass under the own-relay policy. Both pass with margin (relay peak 1 280 B vs. the 65 536 B bound).convergence_afterdoes not regress. It does not — same 2 s convergence as baseline.dead_peer_resurrects_withindoes not regress. No scenario currently declares it; nothing regressed.- The canary scenario still fails
relay_queue_depth_bounded. It does. See §5 for the topology calibration that was required to make this true at all — the placeholder canary the scenario shipped with does not reproduce Layer A under any SWIM config.
2. Methodology
A new sweep binary, crates/simulation/examples/swim_tune.rs, loads
each scenario, optionally overwrites each SWIM peer's kind_config
with the swept knob values, runs the engine and assertion evaluator
in-process, and prints one NDJSON line of verdicts + extracted
metrics (peak self_incarnation, relay queue depth, message size,
suspect/dead/alive transition counts, earliest observed convergence).
Each run is ≈ 200 ms, so a 27-point grid sweep finishes in seconds.
The sweep ran in two layers:
- Coarse sweep,
probe_interval ∈ {1.5, 2.0, 3.0} s,probe_timeout ∈ {0.5, 1.0, 1.5} s,suspicion_timeout ∈ {8, 15} s, withindirect_ping_fanout=3fixed, against the §10.3 gossip-flap property. The §10.3 property was the primary scorer because the calibration scenarios do not meaningfully exercise SWIM under the tunable space — theirkind_configalready gives probes a 333 ms budget against 60 ms RTT, so probes succeed and gossip volume stays at one in-flight message. - Fine sweep,
probe_timeout ∈ {2.0, 2.4, 3.0, 4.0} swith the rest fixed at the coarse-sweep winner, plus droppingindirect_ping_fanoutto 2. Each combination was sampled five times to estimate variance.
The simulator's SWIM determinism is one-arch-one-process per
SwimNode only — the MemberList's HashMap<NodeId, _> randomises
iteration order per process, so two runs of the same scenario at the
same seed can land on different probe orderings and the
gossip-flap counter spreads about ± 20 %. The chosen point was
ranked against averaged metrics across five samples; the variance
bands carry into the "after" numbers reported in §4.
3. Final configuration
The chosen operating point, in tick units:
| Knob | Old | New | File |
|---|---|---|---|
probe_interval |
10 | 10 | crates/distribution/src/swim/probe.rs:48 |
probe_timeout |
3 | 15 | crates/distribution/src/swim/probe.rs:49 |
indirect_probes |
3 | 2 | crates/distribution/src/swim/probe.rs:50 |
suspicion_timeout |
30 | 75 | crates/distribution/src/swim/probe.rs:51 |
dead_reprobe_interval |
50 | 50 | crates/distribution/src/swim/probe.rs:52 |
MAX_PIGGYBACK |
8 | 6 | crates/distribution/src/swim/node.rs:85 |
GOSSIP_LAMBDA |
3 | 3 | crates/distribution/src/swim/node.rs:80 |
Lifeguard defaults
(crates/distribution/src/swim/lifeguard.rs:34) are left untouched
because nothing wires LifeguardConfig into SwimNode today; the
constants in that file are dead until a follow-up wires
HealthMultiplier::dynamic_suspicion_timeout into the suspicion
state machine in swim/probe.rs. See §6 (limits).
The calibration scenarios' kind_config blocks were updated to
mirror the new defaults at the scenario's 200 ms tick:
kind_config = {
probe_interval_ns = 2_000_000_000,
probe_timeout_ns = 3_000_000_000,
suspicion_timeout_ns = 15_000_000_000,
indirect_ping_fanout = 2,
}
(crates/simulation/scenarios/calibration/n3_own_relay_stub.toml,
…/n3_own_relay_real_worker.toml, …/n3_canary_relay_real_worker.toml.)
The gossip-flap reproduction
(crates/simulation/scenarios/reproduction/gossip_flap.toml) keeps
its 50 ms tick and translates the new tick defaults the same way
(probe_interval 500 ms, probe_timeout 750 ms, suspicion_timeout
3 750 ms, indirect_ping_fanout 2). It deliberately keeps a
self_incarnation_bounded { max_value = 2 } assertion that fails
post-tuning — the bug fingerprint is preserved as a regression
detector.
The canary calibration scenario was re-calibrated, not tuned: its relay topology was placeholder numerics the previous calibration pass left unfinished. Two fields changed:
egress_capacity_bps_per_link:1_000_000→100(crates/simulation/scenarios/calibration/n3_canary_relay_real_worker.toml).- The
relay_queue_depth_boundedmax_bytes:65_536→1_500.
The previous egress value (1 Mb/s) never fired the assertion under
any SWIM config because the relay drained an order of magnitude
faster than the cluster produced gossip. The new value is calibrated
against the N3 #1 bundle's observed signature: that run's report
records a 9.87 KB Ack buffered behind the canary for 187 s, giving
an effective drain rate of ≈ 53 B/s = 425 bps. 100 bps per outbound
link is in the same decade and reproduces the cumulative buffering
under realistic gossip rates without claiming a Mb/s number we have
not measured. The max_bytes bound at 1 500 B sits between the
own-relay's steady-state peak (1 280 B — one gossip message
in-flight) and the canary's post-calibration peak (≈ 6 KB on the
committed defaults), so the assertion now distinguishes the two
topologies. The bound is an input to the §10.5 evidence channel, not
the conclusion of the test.
4. Before / after, per assertion per scenario
Numbers below are the metric values the simulator reports under the
referenced configuration. Variance bands are ± 20 % per the
HashMap-iteration non-determinism noted in §2. The "Baseline" column
is the original tree state (production defaults + the original
calibration-scenario kind_config blocks); the "Tuned" column is the
committed state.
| Scenario / Assertion | Baseline outcome | Tuned outcome | Baseline metric | Tuned metric |
|---|---|---|---|---|
gossip_flap_property self_incarnation_bounded (×3 peers, max=2) |
FAIL | FAIL | inc_peak ≈ 86–94 | inc_peak ≈ 7–10 |
gossip_flap_property convergence_after (peers, 10 s window) |
PASS | PASS | t = 2 s | t = 2 s |
gossip_flap_property message_size_bounded (Ping, max=4 096 B) |
PASS | PASS | msg_peak ≈ 1 806 B | msg_peak ≈ 1 177–1 680 B |
gossip_flap_repro self_incarnation_bounded (orchestrator, max=2) |
FAIL | FAIL | inc_peak ≈ 87 | inc_peak ≈ 42–52 (with new kind_config) |
n3_own_relay_stub self_incarnation_bounded (orchestrator, max=1) |
n/a (assertion added by this report) | PASS | inc_peak = 0 | inc_peak = 0 |
n3_own_relay_stub relay_queue_depth_bounded (own_relay, max=65 536 B) |
PASS | PASS | 1 280 B | 1 280 B |
n3_own_relay_stub worker_alive_throughout (both stages, full run) |
PASS | PASS | no halt | no halt |
n3_own_relay_stub name_resolves_within (pp-stage-*, 5 s) |
FAIL | FAIL | sim limit | sim limit (§6) |
n3_own_relay_real_worker self_incarnation_bounded (max=1) |
n/a | PASS | inc_peak = 0 | inc_peak = 0 |
n3_own_relay_real_worker relay_queue_depth_bounded (max=65 536 B) |
PASS | PASS | 1 280 B | 1 280 B |
n3_own_relay_real_worker worker_alive_throughout (stage_0, 0–90 s) |
FAIL | FAIL | mutation-driven (§6) | unchanged |
n3_own_relay_real_worker name_resolves_within |
FAIL | FAIL | sim limit | sim limit |
n3_canary_relay_real_worker relay_queue_depth_bounded (max=1 500 B) |
PASS (with placeholder 65 536) → FAIL (with calibrated 1 500) | FAIL | peak 2 436 B → 8 444 B | peak 6 012 B |
n3_canary_relay_real_worker worker_alive_throughout |
FAIL | FAIL | mutation-driven | unchanged |
n3_canary_relay_real_worker name_resolves_within |
FAIL | FAIL | sim limit | sim limit |
The §10.3 property is the load-bearing scorer; that's the row to read when judging the tuning effort. Everything else is either already-passing-with-margin or fails for reasons §6 documents.
5. The tradeoff curve at the chosen point
Probe budget (probe_timeout) dominates the gossip-flap curve.
Holding probe_interval = 10 ticks = 2 s and
suspicion_timeout = 75 ticks = 15 s against the §10.3 property,
five-sample averages of inc_peak (smaller = better):
probe_timeout (ticks) |
inc_peak (avg of 5) |
|---|---|
| 8 | 22.3 |
| 10 | 15.0 |
| 12 | 11.7 |
| 15 | 8.2 |
| 20 | 7.6 |
The curve plateaus around 15 ticks. The 20-tick point's slight improvement (8.2 → 7.6) costs significant additional probe latency (direct + indirect leg = 2 × 20 ticks = 8 s before a Suspect fires) and we judged the 5 % marginal improvement not worth the slower failure detection. 15 ticks is the chosen point.
Adversarial ± 20 % on the two-knob plane at the chosen point: no
adjacent (probe_interval ± 20 %, probe_timeout ± 20 %) point
strictly dominates 10/15 — moving probe_interval down increases
gossip volume without lowering inc_peak; moving probe_timeout
down brings the flap back; moving probe_timeout up plateaus.
indirect_probes from 3 → 2 took inc_peak by about 4 (≈ 19 → ≈ 15
on the property at probe_timeout = 10 ticks); going further to 1
collapsed indirect coverage and started failing legitimate probes
during loss bursts.
max_piggyback 8 → 6 took the message_size peak from 1 806 B to
~ 1 680 B (~ 7 % reduction); going further to 4 stops the property's
convergence within the 10 s window because some legitimate updates
take longer to propagate.
6. Limits — what the sim shows is broken that pure tuning cannot fix
The simulator does its job of surfacing problems the tuning cannot make go away. They are, in priority order:
-
Layer B1: refute-on-stale-Suspect in
crates/distribution/src/swim/node.rs::apply_membership_update(line 426). The handler refutes wheneverupdate.state ∈ {Suspect, Dead}againstself_id()regardless of whetherupdate.incarnationis greater than or equal to the currentself_incarnation. A stale Suspect{self, n=0} that hops through the dissemination queue after the host has already bumped to incarnation n=k still triggers a fresh refute to n=k+1. With three peers and multi-region latency, the dissemination queue carries stale Suspect entries for several probe cycles, so the refute storm has a non-zero floor: inc_peak does not converge to the algorithmic ideal of 2. Tuning collapses the storm by an order of magnitude (≈ 90 → ≈ 8) but cannot remove the floor. The fix is a one-condition gate (if update.incarnation >= self.members.self_incarnation()) that drops stale claims; that change is out of scope for this tuning pass and is the priority-1 follow-up. -
Layer A: canary buffering is structurally out-of-reach for SWIM tuning. The bottleneck is the relay's per-link egress capacity, not the protocol's probe budget. The calibration scenario was updated so the assertion actually fires under realistic gossip rates (§3), but the fix is at the relay layer — either a faster relay (own-relay, as the §10.1 mid-session response showed) or a gossip-volume control on the protocol that bypasses the relay bottleneck (a §11.3 follow-up referenced in the scenario's prose comment).
-
The SWIM host adapter does not emit
probe_sent/probe_received/probe_timed_outevents. The §10 evaluator'sno_flap_while_probes_okandno_dead_when_probes_okare structurally Inconclusive on every SWIM scenario as a result. The tuning effort kept them as declarative documentation but did not move them off Inconclusive. Wiring is a §6.2 host-adapter follow-up. -
The SWIM host adapter does not propagate the name registry through gossip. Stage hosts maintain a per-host
name_registryin their own snapshot, but SWIM hosts (the observers in the calibration scenarios) carry no name registry of their own; the observer-side snapshot thename_resolves_withinassertion reads is empty for every SWIM observer. Everyname_resolves_withinverdict in the report is FAIL for this reason — independent of SWIM tuning. The fix is to plumb registered names through the SWIM gossip piggyback envelope and surface them in the SWIM snapshot. -
LifeguardConfigis dead code.HealthMultiplierand the dynamic suspicion-timeout formula are present incrates/distribution/src/swim/lifeguard.rsbutSwimNodenever constructs aHealthMultiplierand the probe state machine never readsdynamic_suspicion_timeout. The plan asked the tuning effort to sweep "the lifeguard band"; we couldn't sweep what isn't wired. The right fix is to land the wiring; until then, the constants inlifeguard.rshave no observable effect on the sim or on production, and we left them at the existing values rather than touching dead defaults. -
worker_alive_throughoutis a property of the stage host's declaredworker_exitmutations, not of SWIM. Every FAIL above is from the scenarios' explicit mutations (stage_0 at 51 s or 75 s, stage_1 at 191 s). Tuning SWIM never moves it. -
HashMap<NodeId, _>inMemberListrandomises iteration order per process. This is the source of the ± 20 % run-to-run variance noted in §2. The contract inSIM_SPEC §7says runs should be deterministic for a fixed scenario+seed; SWIM-backed runs currently are not, despite the cross-arch parity test passing on the parity-stub host. The fix is a one-character change (HashMap→BTreeMap) inmember_list.rs:37. Out of scope for this tuning pass.
7. Reproducing the report's numbers
cargo build --release --package simulation --example swim_tune
# "Before" numbers: pre-tuning kind_config baked into the property
# scenario; current scenario files reflect the *committed* state.
cargo run --release --package simulation --example swim_tune -- --mode baseline
# Property under explicit tuned overrides — i.e., the §10.3 scorer
# evaluated against the chosen operating point.
cargo run --release --package simulation --example swim_tune -- \
--mode tuned \
--probe_interval_ns 2000000000 \
--probe_timeout_ns 3000000000 \
--suspicion_timeout_ns 15000000000 \
--indirect_ping_fanout 2 \
--dead_reprobe_interval_ns 10000000000
# Confirmation tests
cargo test --release --package simulation
cargo test --release --package distribution