swactor/examples/pipeline-parallel-inference
Zachery Aaron Shores-Chmielewski f1f98e9f7d spec: add behavioral test subsections to SIM_SPEC.md
Six per-component property-test subsections (§4.10 engine, §5.8 network,
§6.4 hosting, §8.4 scenario loader, §9.6 bundle writer, §10.5 assertion
evaluator) written in property-language form. Each names the property the
component must have; how the property is verified is the test author's
call.

These were drafted in conversation but never committed, causing the
boundary check to flag SIM_SPEC.md on iteration 1 of the implementation
loop. Committing now puts them at HEAD so future loop iterations diff
cleanly.


Signed-off-by: Zachery Aaron Shores-Chmielewski <zacheryasc@gmail.com>
2026-05-22 08:08:01 +00:00
..
scripts feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
src feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
tests feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
Cargo.lock feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
Cargo.toml feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
DIAGNOSTICS_PLAN.md feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
docker-compose.diag.yml feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
Dockerfile feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
Dockerfile.diag feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
Dockerfile.stub feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
N3_DEPLOYMENT_REPORT.md feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
pp_tinygrad_worker.py feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
README.md feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
SIM_SPEC.md spec: add behavioral test subsections to SIM_SPEC.md 2026-05-22 08:08:01 +00:00
smoke.log feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
SPEC.md feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
TEST_SPEC.md feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
VASTAI_RUNBOOK.md feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00
VASTAI_STATUS.md feat(distribution): add diagnostics subsystem 2026-05-20 11:41:30 +04:00

pipeline-parallel-inference

A two-node pipeline-parallel inference example over swactor + iroh. Stage 0 holds the embedding and the first half of the transformer blocks; stage 1 holds the second half plus the output norm and LM head. The two stages exchange StageActivation (forward) and NextToken (feedback) messages over real iroh QUIC; the orchestrator submits one InferenceRequest to pp-entry and waits for the InferenceResponse.

The model is llama3.2:1b. See SPEC.md for the wire protocol and TEST_SPEC.md for the authoritative list of test names.


Layout

examples/pipeline-parallel-inference/
├── Cargo.toml              # standalone crate; two binary targets
├── pp_tinygrad_worker.py   # stage worker (stub + real modes)
├── src/
│   ├── lib.rs              # public modules
│   ├── messages.rs         # InferenceRequest/Response, StageActivation, NextToken + codecs
│   ├── stage_actor.rs      # Stage0Actor, Stage1Actor, ProcessBridge plumbing
│   ├── topology.rs         # pp-entry / pp-exit / pp-stage-{i} name helpers
│   ├── iroh_transport.rs   # actor message pump over iroh
│   ├── vastai.rs           # rent N / destroy-all helpers
│   └── bin/
│       ├── pp_gpu_node.rs  # one stage per process, env-configured
│       └── pp_smoke_run.rs # orchestrator (--seed | --vastai)
└── tests/
    ├── t_codec.rs          # §1
    ├── test_worker.py      # §2, §3
    ├── t_actor.rs          # §4
    ├── t_topology.rs       # §5
    ├── t_cluster.rs        # §6
    ├── t_vastai.rs         # §7
    ├── t_integration.rs    # §8 (stub) + §9 equivalence (#[ignore])
    ├── t_binary.rs         # §10
    └── t_docker.rs         # §13b (Docker-coordinated E2E, #[ignore])

Running the tests

All paths below assume the repository root as the working directory.

Rust fast tier

Everything that does not need a real model:

cargo test --manifest-path examples/pipeline-parallel-inference/Cargo.toml

This covers codecs, stage actors against the stub worker, topology / SWIM name resolution, the in-process iroh cluster, the wiremock-backed vast.ai client, the stub end-to-end integration tests, and the two stub binary E2E tests in t_binary.rs.

Rust slow tier (#[ignore])

The slow tier needs:

  • clang (tinygrad's CLANG backend) or an NVIDIA GPU with the CUDA toolchain (set CUDA=1), with tinygrad importable on python3.
  • The llama3.2:1b GGUF cached in tinygrad's standard model location.
  • ~30s startup per worker plus per-token decode time; budget at least a few minutes per test.
cargo test --manifest-path examples/pipeline-parallel-inference/Cargo.toml -- --ignored

Slow Rust tests:

  • t_integration.rs::sliced_pipeline_matches_full_generate_* — drives the real-mode pipeline against a single-process reference and asserts token-for-token equality (max_tokens=8).
  • t_binary.rs::binary_e2e_two_tinygrad_workers_returns_response — real-mode variant of the binary smoke run.

Python tests

pytest examples/pipeline-parallel-inference/tests/

Fast pytest run covers the stub-mode protocol: env handling, op dispatch, stage-0 / stage-1 boundary cases, error handling, malformed JSON, EOF behavior. The real-worker class is gated behind an env var:

PP_REAL_WORKER_TESTS=1 pytest examples/pipeline-parallel-inference/tests/test_worker.py::TestRealTinygradWorker

The gated class loads llama3.2:1b and exercises one prefill + decode round-trip per stage. Same prerequisites as the Rust slow tier.


Localhost smoke run

cargo build --manifest-path examples/pipeline-parallel-inference/Cargo.toml --release

Then:

PP_WORKER_STUB=1 ./examples/pipeline-parallel-inference/target/release/pp-smoke-run \
    --seed \
    --gpu-node ./examples/pipeline-parallel-inference/target/release/pp-gpu-node \
    --worker  ./examples/pipeline-parallel-inference/pp_tinygrad_worker.py \
    --prompt "Say hello" \
    --max-tokens 16

PP_WORKER_STUB=1 selects the deterministic stub worker — no GPU, no GGUF, no tinygrad import. The orchestrator forwards PP_WORKER_STUB, MODEL, PYTHON, CUDA, and WORKER_CMD from its own environment to each child pp-gpu-node, so changing modes is a single env-var flip.

For real-mode local inference, drop PP_WORKER_STUB and set the appropriate backend env:

CUDA=1 MODEL=llama3.2:1b ./examples/pipeline-parallel-inference/target/release/pp-smoke-run \
    --seed \
    --gpu-node ./examples/pipeline-parallel-inference/target/release/pp-gpu-node \
    --worker  ./examples/pipeline-parallel-inference/pp_tinygrad_worker.py \
    --prompt "Say hello" \
    --max-tokens 64

Use PYTHON=1 for the CPU fallback. On success the orchestrator prints the response between === pipeline-parallel Inference Response === banners and exits 0. It always kills both pp-gpu-node children before returning.

Flags

Flag Default Purpose
--seed / --vastai (required, exactly one) mode selector
--prompt <text> Say hello request prompt
--max-tokens <n> 64 decode-loop cap
--gpu-node <path> sibling of pp-smoke-run pp-gpu-node binary path
--worker <path> pp_tinygrad_worker.py next to Cargo.toml worker script path
--api-key <key> — vast.ai API key (--vastai only)
--gpu <name> RTX 4090 GPU constraint (--vastai only)
--image <name> swactor-pp-gpu:latest container image (--vastai only)

Docker-coordinated localhost run (the pre-deploy gate)

scripts/docker-e2e.sh N is the one-line pre-deploy gate (TEST_SPEC §13b). It builds the release binaries, packages them into a CPU-only stub-mode container image, brings up N containers on the host network, drives one inference request through them, prints the response, and tears everything down — failing fast if any container exits prematurely or any container survives the run.

examples/pipeline-parallel-inference/scripts/docker-e2e.sh 3

Useful environment overrides:

Variable Default Purpose
PP_IMAGE pp-gpu-node-stub:latest image tag to build/use
PP_CONTAINER_PREFIX pp-stage container name prefix (each container is <prefix>-<stage>)
PP_MAX_TOKENS 4 per-request decode cap (stub workers are deterministic, 4 keeps it fast)
PP_PROMPT Say hello inference prompt
PP_SKIP_BUILD unset skip cargo build --release when artefacts are already present
PP_SKIP_IMAGE_BUILD unset skip the docker build step when the image is already current

What the script asserts:

  • pp-smoke-run exits 0.
  • The response banner (=== pipeline-parallel Inference Response ===) appears on stdout with a non-empty body.
  • No container whose name matches ^<prefix>-[0-9]+$ survives the run.

The script can be rerun safely — TEST_SPEC §13b includes an idempotency test that runs the harness twice back-to-back. If you need to reproduce the failure path that drives the script's fail-fast behaviour, launch pp-smoke-run with the docker shim manually and docker kill one stage mid-run; the orchestrator exits non-zero within seconds.

Same machinery is exercised under cargo test as t_docker.rs:

cargo test --manifest-path examples/pipeline-parallel-inference/Cargo.toml \
    --test t_docker -- --ignored

vast.ai mode (documented, not executed here)

pp-smoke-run --vastai rents two GPU instances, deploys the pp-gpu-node image to each (one with STAGE=0, one with STAGE=1, both with SEED_ADDR=<local node> and an iroh relay URL), waits for SWIM convergence over WAN, sends one InferenceRequest, prints the response, and destroys both instances on the way out — including on partial-success rentals or destroy failures.

./examples/pipeline-parallel-inference/target/release/pp-smoke-run \
    --vastai \
    --api-key "$VASTAI_API_KEY" \
    --gpu "RTX 4090" \
    --image swactor-pp-gpu:latest \
    --prompt "Say hello" \
    --max-tokens 64

This path is exercised end-to-end against a wiremock server in t_vastai.rs (rental, partial-success rollback, destroy-all error fan-out). It is not run against real vast.ai by the test suite; the real run is the human-driven step described in TEST_SPEC §11.


Stub vs real mode

The Python worker is a single file (pp_tinygrad_worker.py) that branches on --stub / PP_WORKER_STUB=1:

  • Stub mode. No tinygrad import, no GGUF load. Deterministic byte-shaped activations, a small synthetic vocab, and an EOS rule that lets the integration tests cover the EOS termination path. The whole test suite except the slow tiers runs in this mode.
  • Real mode. Loads llama3.2:1b via Transformer.from_gguf, slices model.blk to its stage's range, runs the real forward and argmax sampler. Stage 0 owns tokenization + model.token_embd; stage 1 owns model.output_norm + model.output and detokenization via SimpleTokenizer.from_gguf_kv.

Switching modes from Rust: pp-smoke-run reads PP_WORKER_STUB from its environment and forwards it to each child. pp-gpu-node and the stage actors toggle their tokenization path on the same variable, so a real-mode run does not feed synthetic whitespace-split ids into a real embed lookup.