swactor/CLAUDE/notes/progress.md
Developer 771c38c863 feat: OneForAll and RestForOne supervisor strategies (Cycle 18)
Add coordinated restart strategies to the Supervisor actor. OneForAll
restarts all children when one fails; RestForOne restarts the failed
child and all children after it in spec order. Uses a SupervisorPhase
state machine to coordinate stop signals and Down confirmations.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-12 14:12:28 +00:00

455 lines
33 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Progress Log
## Current Stage: Phase 1 — Research + First Improvement Cycle
### Status: Cycle 18 COMPLETE
## Plan Overview
1. **Phase 0**: Codebase audit — understand current swactor architecture, existing tests, benchmarks ✅
2. **Phase 1**: Broad survey + interleaved improvements
3. **Phase 2**: Deeper improvements based on findings
4. **Phase 3**: Testing methodology improvements
5. **Phase 4**: Final evaluation & documentation
## Completed This Session
### Cycle 1: Fairness (Message Budget)
- **Research**: Studied ractor, tokio, Erlang/OTP BEAM, Linux CFS/EEVDF, libuv
- **Finding**: `tick_all` drained ENTIRE mailbox per actor per tick — critical fairness bug
- BEAM uses 4000 reduction budget, tokio uses 128-op cooperative budget
- Swactor had zero budget — one hot actor could starve all others on same worker
- **Implementation**: Added `actor_message_budget` to `RuntimeConfig` (default: 64)
- Modified `tick_all` to break after `budget` messages per actor
- `budget=0` means unlimited (backward compatible)
- **Tests**: 3 new fairness tests (hot_actor_does_not_starve_cold_actor, unlimited_budget_drains_all, budget_messages_drain_across_multiple_ticks)
- **Benchmarks**: Added fairness benchmark group (cold_latency_under_pressure, throughput_by_budget)
- **Fixes**: Updated RuntimeConfig struct literals across crates (python, runtime-dashboard, mt_benchmarks)
- **Result**: 45 tests pass (42 original + 3 new), all workspace crates compile
### Cycle 2: Stress Tests, Benchmarks, Research Expansion
- **Research**: Added Kameo and Actix analysis to synthesis
- Actix uses custom Vyukov lock-free MPSC queue (why it's fastest)
- Kameo has dual bounded/unbounded mailbox, default capacity 64
- Both use vtable dispatch (not Box<dyn Any> downcast)
- Actix has 256-message assertion guard (validates our budget approach)
- **Stress tests**: 6 new tests
- `message_ordering_preserved_under_budget` — FIFO order with budget=8
- `mt_stress_many_senders_one_receiver` — 50 senders × 100 msgs, 4 threads
- `mt_stress_concurrent_spawn_and_send` — 200 concurrent spawn+send, 4 threads
- `mt_chain_spawning_under_load` — 50-level chain across 2 workers
- `mt_panic_isolation_under_load` — 10 panicking + 10 healthy actors, 4 threads
- `sustained_throughput_does_not_drop_messages` — 10 batches × 100 msgs
- **Benchmarks**: 2 new benchmark groups
- `msg_size`: throughput and send_latency by message size (8B-4KB)
- `contention`: fanin (1-100 senders to 1 sink), cross_worker (1-4 threads)
- **Result**: 51 tests pass (42 original + 3 fairness + 6 stress), all workspace compiles
### Cycle 3: Thread Parking (Adaptive Backoff)
- **Implementation**: Replaced `thread::sleep` with `thread::park_timeout` in worker run loop
- Workers register `thread::current()` via `OnceLock<Thread>` on startup
- `send_to` and `spawn` call `Thread::unpark()` on target worker
- Cross-worker sends from `WorkerContext` also unpark target
- Zero new dependencies (uses `std::sync::OnceLock` + `std::thread::park_timeout`)
- **Design source**: Tokio's parker state machine, Linux NO_HZ adaptive ticks
- **Benefits**: Parked workers wake instantly when work arrives (vs waiting for sleep timer)
- Reduces idle-to-active latency from up to 1ms to near-zero
- No overhead on hot path — `unpark()` is no-op if thread isn't parked
- **Tests**: 1 new test (`mt_parked_worker_wakes_on_send`)
- **Result**: 52 tests pass (51 + 1 new), all workspace compiles
### Cycle 4: Shutdown Fix + Bug-Inspired Tests
- **Shutdown improvement**: `shutdown()` now unparks all workers for immediate exit
- Previously, parked workers wouldn't notice shutdown until park_timeout expired
- **Bug-inspired tests** (5 new, from competitor bug reports):
- `stats_snapshot_is_read_only` — from ractor #310 (destructive get_children)
- `stats_under_load_do_not_interfere_with_processing` — stats don't affect msg processing
- `shutdown_wakes_parked_workers_immediately` — validates fast shutdown with parking
- `mt_send_after_run_delivers_to_running_actors` — from kameo #185 (startup delivery)
- `budget_respected_even_with_self_sends` — from actix #515 (mailbox bypass)
- **Result**: 57 tests pass, all workspace compiles
### Cycle 5: Work Stealing Research + Load-Aware Placement
- **Research**: Deep analysis of work stealing in Tokio, Go, BEAM, ForkJoinPool
- Tokio: fixed 256-slot ring, steal-half, LIFO slot (3-use starvation cap), N/2 searcher limit
- Go: M:N scheduler, runnext + 256-slot local queue, steal-half, 4 tries with random permutation
- BEAM: unique dual approach — reactive stealing + proactive migration via check_balance()
- ForkJoinPool: owner LIFO / thief FIFO deque, even/odd queue indexing
- **Feasibility analysis**: Full actor migration IS mechanically possible (ActorSlot is Send), but:
- Requires push-based donation (ActorPool not Sync → no pull stealing)
- 1-tick message loss window during migration
- Significant complexity for uncertain benefit
- **Implementation**: Load-aware placement replaces blind round-robin
- `Placement::next_worker()` now reads per-worker stats (num_actors + mailbox_depth)
- Scan starts from rotating position → round-robin when all stats equal (initial burst)
- O(N) relaxed atomic loads per spawn, trivial for N≤8 workers
- **Tests**: 3 new tests
- `load_aware_placement_prefers_lighter_worker` — imbalanced load biases toward lighter worker
- `load_aware_placement_single_worker_degrades_gracefully` — single-thread works correctly
- `load_aware_placement_falls_back_to_round_robin_on_fresh_runtime` — even distribution before ticks
- **Benchmarks**: 1 new group — `placement/spawn_under_load` (2t, 4t)
- **Result**: 60 tests pass, all workspace compiles
### Cycle 9: Lifecycle Hooks + Graceful Stop
- **Research**: Cross-framework lifecycle analysis (Erlang init/terminate, Akka preStart/postStop,
Actix started/stopping/stopped, Kameo on_start/on_stop/on_panic, Ractor pre_start/post_stop,
Stakker state-based Prep/Ready/Zombie, CAF on_exit)
- Also researched graceful stop across Erlang (gen_server:stop, exit, kill), Akka (stop, PoisonPill,
Kill, gracefulStop), Actix (ctx.stop, Running::Stop), Kameo (stop_gracefully, kill), Go (context.Done)
- Key finding: most frameworks have on_stop NOT called on panic (state may be corrupt)
- Key finding: self-stop should be immediate (after current message), external stop is queued
- **Implementation**: Lifecycle hooks + dual-mode graceful stop
- `ActorInterface::on_start()` and `on_stop()` — default no-ops, backward compatible
- `AnyActor::on_start()`/`on_stop()` forwarded from `Actor<A>` impl
- `ctx.stop_self()` — immediate stop after current message via `request_stop` buffer
- `runtime.stop_actor(addr)` — external stop via StopSignal message (PoisonPill semantics)
- `ActorSlot` gains `started: bool` and `stopping: bool` flags
- `on_start` called in tick_all before first message; panic in on_start → immediate poison
- `on_stop` called in cleanup_dead for stopping (not poisoned) actors, wrapped in catch_unwind
- Restarted actors get `started=false` so on_start fires again on fresh instance
- `stops: AtomicU64` added to WorkerStats and WorkerInfo
- `ContextInner::request_stop()` method for same-worker immediate stop
- Phase 7 cleanup_dead now handles both poisoned AND stopping actors, with on_stop context
- **Tests**: 12 new tests
- `on_start_called_before_first_message` — on_start fires on first tick, before messages
- `on_start_called_per_actor` — 5 actors each get one on_start call
- `on_start_panic_poisons_actor` — panic in on_start → poisoned, no messages processed
- `actor_can_stop_self` — 5 msgs sent, stops after 3, only 3 processed, on_stop called
- `runtime_can_stop_actor` — external stop via runtime, on_stop called, actor removed
- `send_to_stopped_actor_returns_error` — stopped actor gone from address map
- `stop_vs_panic_tracked_separately_in_stats` — stops and panics counted independently
- `on_stop_can_send_messages` — farewell message sent during on_stop is delivered
- `on_start_called_again_after_restart` — restartable actor gets on_start on fresh instance
- `external_stop_is_queued_after_pending_messages` — PoisonPill semantics for external stop
- `external_stop_before_new_messages_prevents_processing` — stop before send blocks msgs
- `stop_nonexistent_actor_returns_error` — stop on bad address returns Err
- **Result**: 82 tests pass, all workspace compiles
### Cycle 17: Supervision Trees (handle_down + Supervisor Actor)
- **Research**: Cross-framework supervision analysis — Erlang (one_for_one/all/rest, child specs, intensity/period),
Akka (SupervisorStrategy, Resume/Restart/Stop/Escalate, BackoffSupervisor), Ractor (SupervisionEvent,
ractor-supervisor crate), Bastion (hierarchy, redundancy groups), CAF (no built-in supervisor, monitor-based)
- Key finding: swactor has all building blocks (monitor, spawn_restartable, lifecycle hooks, Down messages)
- Decision: Supervisor as a user-space actor built on existing primitives (like Ractor base crate)
- handle_down callback enables any actor to react to monitored deaths without making Down the Incoming type
- **Implementation**: Three features added to `src/actor.rs`:
1. **`handle_down` callback on ActorInterface** — default no-op, called when monitored actor dies
and actor's Incoming type is NOT Down. Implemented via second downcast attempt in `handle_any`.
Fully backward-compatible: actors with `Incoming = Down` still receive via `handle()`.
2. **`ctx.stop_actor(addr)`** — send graceful stop to another actor from handler context.
Uses StopSignal through normal message routing (PoisonPill semantics).
3. **`Supervisor` actor** — manages child actors with configurable restart policies:
- `SupervisorStrategy::OneForOne` — only failed child is restarted
- `RestartPolicy::Permanent` — always restart
- `RestartPolicy::Transient` — restart only on Panicked, not Normal
- `RestartPolicy::Temporary` — never restart
- `ChildSpec` with id, restart policy, and factory closure `Fn(&Ctx) -> Result<ActorAddress>`
- Meltdown detection: stops itself when `total_restarts > max_restarts`
- Cascading shutdown: on_stop sends stop signals to all living children
- **Tests**: 10 new behavioral tests
- `handle_down_receives_death_notification` — handle_down callback fires on monitored death
- `handle_down_skipped_when_incoming_is_down` — backward compat: Incoming=Down uses handle()
- `ctx_stop_actor_stops_target` — one actor can stop another via ctx.stop_actor()
- `supervisor_restarts_permanent_child_on_panic` — panic → restart (OneForOne + Permanent)
- `supervisor_does_not_restart_transient_child_on_normal_stop` — Normal stop → no restart
- `supervisor_restarts_transient_child_on_panic` — Panicked → restart (Transient)
- `supervisor_never_restarts_temporary_child` — Temporary → never restart
- `supervisor_meltdown_after_max_restarts` — exceeding max_restarts stops supervisor
- `supervisor_one_for_one_only_restarts_failed_child` — multi-child, only crashed child affected
- `supervisor_on_stop_kills_children` — supervisor shutdown cascades to children
- **Result**: 138 tests pass (130 behavioral + 7 proptest + 1 doctest), zero warnings, full workspace compiles
### Cycle 18: OneForAll + RestForOne Supervisor Strategies
- **Research**: Investigated SmallBox/InlineAny optimization (44% queue throughput improvement)
but deferred due to unsafe code risk and 32+ call-site changes violating structural constraints.
Chose to extend Supervisor with remaining Erlang-style strategies instead.
- **Implementation**: Extended `Supervisor` in `src/actor.rs` with coordinated restart strategies:
- `SupervisorStrategy::OneForAll` — all children restarted when one fails
- `SupervisorStrategy::RestForOne` — failed child + all children after it (in spec order) restarted
- `SupervisorPhase` state machine: `Normal` (steady state) | `Stopping { awaiting, restart_set }` (coordinated)
- During coordinated restart: supervisor stops living siblings, waits for all Down confirmations,
then restarts the full restart set in spec order
- `begin_coordinated_restart(ctx, indices)` — sends stop signals, transitions to Stopping phase
- `finish_restart(ctx)` — called when all awaiting Downs received, restarts from spec order
- `check_intensity()` factored out for restart budget checking
- Already-dead children are handled: if all targets are already dead, immediate restart (no Stopping phase)
- **Tests**: 3 new behavioral tests
- `supervisor_one_for_all_restarts_all_on_single_failure` — one child panics, all 3 get new addresses
- `supervisor_rest_for_one_restarts_rest_after_failed` — child_b panics, child_a unchanged, child_c restarted
- `supervisor_one_for_all_waits_for_all_downs_before_restart` — verifies coordinated shutdown completes before restart
- **Result**: 141 tests pass (133 behavioral + 7 proptest + 1 doctest), zero warnings, full workspace compiles
### Cycle 16: Benchmark New Features
- **Scope**: Added benchmark group for features from Cycles 12-15 (named actors, groups, monitors, ask)
- **New benchmarks** (5 total in `registry` group):
- `named_spawn_lookup` — spawn_named + where_is roundtrip: **~2.4µs** (vs bare spawn 1.9µs → +0.5µs overhead for name registration)
- `where_is_100_names` — lookup in 100-name registry: **~9.0µs** (includes setup overhead)
- `group_publish/{10,50,100}` — broadcast to N members: 4.8µs/15.5µs/60µs (linear with O(N) clones)
- `monitor_setup` — monitor + stop + cleanup: **~13.4µs**
- `ask_roundtrip` — ask + recv_ticking: **~4.5µs** (vs manual roundtrip 3.0µs → +1.5µs for inbox creation)
- **Analysis**: All registry operations are efficient. Named lookup adds <1µs over bare spawn.
Ask adds ~50% overhead vs manual inbox pattern (acceptable for convenience). Group publish
scales linearly — expected for O(N) message cloning. No optimization needed.
- **Result**: All benchmarks run cleanly, 127 tests pass, zero warnings
### Cycle 15: Ask Pattern (Request-Response)
- **Research**: Studied ask/call/request-response patterns across Erlang gen_server:call (From + reply),
Akka ask (temporary actor + Future), Ractor call (RpcReplyPort), Kameo ask (async + Reply trait),
xactor Handler (return value auto-routing)
- Key finding: swactor's synchronous tick model requires explicit reply_to, not implicit routing
- Decision: convenience wrapper over existing inbox pattern, not implicit auto-reply
- **Implementation**: `Ask<R>` struct + `Runtime::ask()` method
- `Ask<R>`: wraps `Inbox<R>` with `try_recv()` and `recv_ticking(rt, max_ticks)`
- `rt.ask(addr, |reply_to| Msg { reply_to })` — creates inbox, builds message, sends, returns Ask<R>
- `ask.recv_ticking(&rt, max_ticks)` — ticks until response or timeout (single-threaded only)
- `ask.try_recv()` — poll without ticking (works in both modes)
- `ask.reply_addr()` — access inbox address for manual use
- Purely sugar over `new_inbox → send_to → tick → try_recv` pattern
- Zero changes to ContextInner or ActorInterface — no implicit auto-reply magic
- **Tests**: 5 new behavioral tests
- `ask_recv_ticking_returns_response` — basic PingPong ask roundtrip
- `ask_multiple_times_tracks_state` — 3 sequential asks to CounterActor
- `ask_timeout_when_no_response` — ask dead actor → timeout error
- `ask_try_recv_returns_none_before_tick` — poll before tick → None, after tick → Some
- `ask_reply_addr_is_accessible` — reply address is valid
- **Result**: 127 tests pass (120 behavioral + 7 proptest), all workspace compiles, zero warnings
### Cycle 14: Actor Groups (Pub-Sub)
- **Research**: Studied group/pub-sub patterns across Erlang pg (scopes, join/leave/get_members),
Akka DistributedPubSub (mediator, topics), Ractor pg (join/leave/broadcast), Bastion (Dispatcher),
Redis pub/sub (channels, patterns)
- Common patterns: auto-cleanup on death, at-most-once delivery, string-based naming,
flat groups (not hierarchical), lazy creation/deletion
- Decision: Erlang pg-style flat groups, string keys, auto-cleanup, RwLock<HashMap> pattern
- **Implementation**: `GroupRegistry` in delivery.rs with forward + reverse maps
- `groups: RwLock<HashMap<String, HashSet<ActorAddress>>>` — group→members
- `memberships: RwLock<HashMap<ActorAddress, HashSet<String>>>` — actor→groups (reverse for cleanup)
- Groups auto-create on first join, auto-delete when empty
- Runtime API: `join_group(addr, name)`, `leave_group(addr, name)`, `publish_to(group, msg)`,
`group_members(group)`, `groups()`
- Ctx API: `join_group(name)`, `leave_group(name)`, `publish(group, msg)`, `group_members(group)`
- `publish` clones at the typed level (Message: Clone), sends to each member via normal routing
- Auto-cleanup: `group_registry.cleanup(&addr)` in cleanup_dead phase removes dead actor from all groups
- ContextInner extended: `join_group()`, `leave_group()`, `group_members()` (publish is Ctx-level only)
- **Tests**: 9 new behavioral tests
- `group_members_returns_joined_actors` — join + query
- `empty_group_returns_no_members` — nonexistent group → empty
- `publish_broadcasts_to_all_members` — 2 members, both receive
- `leave_group_stops_receiving_publishes` — leave → excluded from broadcast
- `dead_actor_auto_removed_from_group` — stop → removed from group
- `actor_removed_from_all_groups_on_death` — multi-group membership cleanup
- `empty_group_auto_deleted` — last member leaves → group removed from groups()
- `ctx_join_group_from_handler` — join via on_start
- `ctx_publish_broadcasts_from_handler` — publish via handler
- **Result**: 122 tests pass (115 behavioral + 7 proptest), all workspace compiles, zero warnings
### Cycle 13: Actor Monitoring / Death Watch
- **Research**: Studied monitoring across Erlang (monitor/2, DOWN messages), Akka (watch/Terminated),
Ractor (link, SupervisionEvent), Actix (none), Kameo (link, on_link_died callback)
- Key finding: Erlang's unidirectional monitor + message delivery is the best fit for swactor
(reuses existing type-erased handler, zero trait changes, composable)
- Callbacks (Ractor/Kameo style) rejected: would require adding to AnyActor/ActorInterface traits
- Bidirectional links deferred: can layer on top of monitors later
- **Implementation**: `MonitorRegistry` in delivery.rs + `Down`/`StopReason`/`MonitorRef` in actor.rs
- `MonitorRegistry`: `RwLock<HashMap<ActorAddress, Vec<(MonitorRef, ActorAddress)>>>` (watched→watchers)
+ reverse `RwLock<HashMap<MonitorRef, ActorAddress>>` for O(1) demonitor
- `MonitorRef(u64)`: unique token from `AtomicU64` counter
- `Down { addr: ActorAddress, reason: StopReason }`: delivered as normal mailbox message
- `StopReason`: `Normal` (graceful stop) | `Panicked` (panic, not restartable)
- `ctx.monitor(target)` → `MonitorRef` — subscribe to death notifications
- `ctx.demonitor(mref)` — cancel a subscription
- `cleanup_dead` now returns `Vec<(ActorAddress, StopReason)>` instead of `Vec<ActorAddress>`
- After cleanup_dead: iterate dead actors, take_monitors from registry, route Down through normal
delivery (pool.deliver for same-worker, transfer_txs for cross-worker, inbox_registry for inboxes)
- Dead watcher cleanup: `remove_watcher()` strips monitor subscriptions for dead watchers
- Multiple monitors of same target produce independent notifications (stacking, like Erlang)
- **Tests**: 7 new behavioral tests
- `monitor_notifies_on_graceful_stop` — Down{reason: Normal} on stop
- `monitor_notifies_on_panic` — Down{reason: Panicked} on panic
- `multiple_watchers_all_notified` — two watchers both get Down
- `demonitor_cancels_notification` — demonitor → no Down delivered
- `dead_watcher_does_not_receive_down` — dead watcher's monitors cleaned up
- `down_delivered_to_external_inbox` — Down forwarded through inbox
- `stacked_monitors_produce_multiple_notifications` — two monitors on same target → two Downs
- **Result**: 113 tests pass (106 behavioral + 7 proptest), all workspace compiles, zero warnings
### Cycle 12: Named Actor Registry
- **Research**: Studied named actor/service discovery across Erlang (register/2, whereis/1, global, pg),
Actix (Registry, SystemRegistry — TypeId keys), Bastion (hierarchy-based), Ractor (String keys, DashMap,
global static), xactor (TypeId singleton), Akka (Receptionist, ServiceKey[T])
- Key findings: TypeId keys (Actix/xactor) don't fit swactor's type-erased model; global static
(Ractor) breaks multi-runtime scenarios; Erlang's register/whereis is the gold standard
- Decision: String keys, RwLock<HashMap> (matches existing AddressMap/InboxRegistry pattern),
per-runtime scope, error on collision, auto-unregister on death
- **Implementation**: `NameRegistry` in delivery.rs with forward + reverse maps
- `NameRegistry`: `RwLock<HashMap<String, ActorAddress>>` + `RwLock<HashMap<ActorAddress, String>>`
- Forward map for O(1) name→addr lookup, reverse map for O(1) addr→name cleanup
- Added to `Runtime` as `Arc<NameRegistry>`, threaded through `TickContext`
- Runtime API: `spawn_named(name, actor)`, `where_is(name)`, `unregister(name)`, `registered_names()`
- Ctx API: `spawn_named(name, actor)`, `where_is(name)` — usable from inside handlers
- `ContextInner` trait extended: `where_is()` + `register_name()` (private, supports both Runtime and WorkerContext)
- Auto-unregister on death: `cleanup_dead` phase calls `name_registry.unregister_by_addr()` for each dead actor
- Name reservation is immediate (before spawn queue push) — prevents TOCTOU race
- Collision returns `Err("Name already registered")` — original binding preserved
- **Tests**: 11 new behavioral tests
- `named_actor_lookup_returns_spawn_address` — spawn_named → where_is roundtrip
- `named_actor_receives_messages_via_lookup` — send to looked-up address works
- `duplicate_name_returns_error` — collision error, original preserved
- `where_is_returns_none_for_unknown_name` — nonexistent name → None
- `name_auto_unregistered_on_actor_death` — stop_actor → name freed
- `name_can_be_reused_after_actor_death` — death → respawn with same name
- `name_auto_unregistered_on_panic` — panic → name freed
- `registered_names_lists_all` — all registered names returned
- `manual_unregister_frees_name_but_actor_lives` — unregister doesn't kill actor
- `ctx_where_is_resolves_inside_handler` — where_is from handler context
- `ctx_spawn_named_registers_from_handler` — spawn_named from handler context
- **Result**: 106 tests pass (99 behavioral + 7 proptest), all workspace compiles, zero warnings
### Cycle 11: Property-Based Testing (proptest + fuzz extension)
- **Research**: Studied testing approaches across tokio (loom), Erlang (PropEr, QuickCheck, Concuerror),
Rust property-based testing (proptest vs quickcheck), cargo-fuzz, and actor-specific testing patterns.
- Ranked approaches: #1 proptest-state-machine (perfect fit for deterministic ticks),
#2 extend cargo-fuzz, #3 simple proptest, #4 shuttle, #5 loom, #6 DST
- Also researched remaining feature gaps: named actors, monitoring/death watch, groups, ask pattern
- **Implementation**: Property-based testing suite with proptest-state-machine
- Added `proptest` and `proptest-state-machine` to dev-dependencies
- New test file: `tests/proptest_runtime.rs` with 7 tests:
- `fifo_ordering_for_any_message_sequence` — FIFO preserved for 1-100 random messages
- `budget_limits_per_actor_processing` — budget caps per-tick processing for 2-10 actors
- `one_shot_timer_fires_at_correct_tick` — timer with delay 1-20 fires at exact right tick
- `interval_timer_fires_at_correct_period` — period 1-10, verifies 3 consecutive fires
- `bounded_mailbox_never_exceeds_capacity` — capacity 1-20, 1-200 messages, never exceeds
- `spawn_n_actors_all_tracked` — 1-50 actors, all unique, all in stats
- `swactor_state_machine` — stateful property test: random Spawn/Send/Tick/Stop/CheckStats
sequences (up to 40 transitions, 128 cases), verifies runtime invariants after each step
- State machine test defines SwactorModel (reference) vs SwactorTest (SUT) with:
- Reference model: HashMap<id, alive> tracking expected actor lifecycle
- Invariants checked after every transition: worker count, actor placement, mailbox safety
- Automatic shrinking finds minimal failing sequences
- Extended fuzz targets (fuzz_runtime.rs) with 4 new RawAction variants:
- `StopActor` — graceful stop via runtime.stop_actor
- `SpawnRestartable` — spawn_restartable with configurable max_restarts
- `ScheduleTimer` — one-shot timer via TimerSchedulerActor
- `ScheduleInterval` — interval timer via IntervalSchedulerActor
- Added 3 new actor types to fuzz: TimerSchedulerActor, IntervalSchedulerActor, RestartableEchoActor
- **Bug found**: State machine test immediately caught invariant mismatch: address map tracks spawned
actors immediately, but per-worker num_actors lags until first tick. Fixed invariant to use <= check.
- **Result**: 95 tests pass (88 behavioral + 7 proptest), fuzz targets compile, zero warnings
### Cycle 10: Actor Timers (Tick-Counting)
- **Research**: Studied timer/scheduling patterns across Erlang (timer:send_after, erlang:start_timer),
Akka (scheduleOnce, scheduler), Actix (ctx.run_later, ctx.run_interval), Kameo (tokio::time::sleep),
Tokio (tokio::time), Go (time.After, time.NewTicker)
- Also researched priority messages (REJECTED: lifecycle hooks cover 95% of use cases)
- Also researched SmallBox optimization (DEFERRED: measure allocation cost first)
- Key finding: per-worker tick-counting is ideal for swactor's synchronous model (deterministic)
- **Implementation**: Per-worker `TimerWheel` with deterministic tick-based scheduling
- `OnceTimer`: fire once at `fire_at` tick, consumed after firing
- `IntervalTimer`: fire every `period` ticks, message cloned via `CloneMsg` trait
- `CloneMsg` trait: type-erased clone for interval timer messages (blanket impl for `Message`)
- `TimerRequest` enum: `Once { dest, msg, ticks }` | `Interval { dest, msg, period }`
- `ctx.send_after_ticks(addr, msg, ticks)` — one-shot timer API
- `ctx.send_interval_ticks(addr, msg, period)` — interval timer API
- Phase 2.5 in tick_once: fire due timers, route through full delivery system
(pool.deliver for local actors, transfer_txs for cross-worker, inbox_registry for inboxes)
- Phase 5.5: drain timer requests from handler buffer into TimerWheel
- GC: interval timers for removed actors cleaned up after cleanup_dead
- `schedule_timer` on Runtime's ContextInner: no-op with warning (timers are per-worker only)
- **Bug fixed**: `gc_dead_intervals` was over-aggressive — removed timers for ANY address not
in the local pool, including inboxes and cross-worker actors. Fixed to only GC timers for
addresses in the `dead` set from cleanup_dead.
- **Tests**: 6 new tests
- `one_shot_timer_fires_after_n_ticks` — timer with delay=3 fires on tick 4
- `handler_can_schedule_one_shot_timer` — timer scheduled from handler, fires correctly
- `one_shot_timer_fires_only_once` — consumed after firing, doesn't repeat
- `interval_timer_fires_repeatedly` — period=2, fires every 2 ticks (3 fires verified)
- `interval_timer_cleaned_up_when_actor_dies` — GC removes orphaned timers
- `timer_with_zero_delay_fires_next_tick` — delay=0 fires on next tick
- **Result**: 88 tests pass, all workspace compiles, zero warnings
### Cycle 8: Dead Actor Cleanup (Memory Leak Fix)
- **Research**: Audited remaining improvement gaps, found ActorPool and AddressMap both leak
permanently poisoned actors. Known class of bug in Akka (#22990), CAF (#420).
- **Implementation**: Automatic cleanup of poisoned actors after tick_all
- `AddressMap::remove()` added to delivery.rs
- `ActorPool::cleanup_dead()` collects and removes poisoned actors, returns their addresses
- Phase 7 in tick_once: cleanup_dead → remove from address_map → update num_actors stat
- Re-publish num_actors after cleanup so stats immediately reflect removal
- **Behavior change**: Sends to poisoned actors now return Err (address not found) instead of
silently discarding. This is better — callers learn the actor is gone.
- **Tests**: 2 new tests + 2 existing tests updated
- `dead_actor_cleaned_up_from_stats_and_address_map` — good actor persists, bad actor removed
- `bulk_dead_actor_cleanup` — 20 panicked actors all cleaned up
- Updated `send_to_poisoned_actor_is_a_silent_black_hole` — now asserts send returns Err
- Updated `poisoned_actor_messages_not_counted_as_processed` — sends fail to cleaned-up actor
- **Result**: 70 tests pass, all workspace compiles
### Cycle 7: Actor Recovery (Factory Restart)
- **Research**: Deep analysis of supervision/recovery across Erlang (supervision trees, restart intensity),
Akka (Resume/Restart/Stop/Escalate), Kameo (on_panic hook), Actix (Supervised trait), Ractor (SupervisionEvent)
- Erlang: fresh process via factory (MFA tuple), mailbox lost, PID changes
- Akka: replace internals but keep ActorRef stable, mailbox preserved (docs say this is usually wrong)
- Kameo: on_panic(&mut self) — risky with corrupt state after panic
- Decision: factory-based restart (Erlang-style), safest approach
- **Implementation**: `spawn_restartable(actor, factory, max_restarts)` on Runtime and Ctx
- `Actor<A>` expanded from tuple struct to named fields: inner, restart_factory, max_restarts, restart_count
- `AnyActor::try_restart(&self)` trait method (default None, backward compatible)
- Factory stored as `Arc<dyn Fn() -> A + Send + Sync>` — cloned into fresh Actor on restart
- `tick_all` panic handler: try_restart before poisoning, clear mailbox, fresh state
- `restarts` counter added to `WorkerStats` and `WorkerInfo`
- **Safety**: Factory fields are "cold" (never touched by handle_any), safe to read after catch_unwind
- **Tests**: 4 new tests
- `restartable_actor_recovers_after_panic` — basic restart works
- `restartable_actor_resets_state_on_restart` — fresh state post-restart
- `restartable_actor_respects_max_restarts` — 2 restarts then permanent poison
- `non_restartable_actor_still_poisons_on_panic` — backward compatibility
- **Result**: 68 tests pass, all workspace compiles
### Cycle 6: Mailbox Backpressure
- **Research**: Compared backpressure across Erlang (unbounded, pobox), Actix (cap 16, do_send bypass),
Kameo (cap 64, bounded), Tokio mpsc (bounded, permit pattern), Go channels (blocking)
- Consensus: bounded by default, configurable overflow policy
- **Implementation**: Per-actor bounded mailboxes with configurable overflow
- Added `MailboxOverflow` enum: `DropNewest` (discard incoming) and `DropOldest` (evict oldest)
- Added `default_mailbox_capacity` and `mailbox_overflow` to `RuntimeConfig`
- Default: capacity=0 (unbounded) — 100% backward compatible
- `ActorSlot` stores per-actor capacity and policy (from runtime defaults)
- `deliver()` enforces bounds; dropped messages tracked via `drops_this_tick` counter
- `messages_dropped: AtomicU64` added to `WorkerStats` and `WorkerInfo`
- **Tests**: 4 new tests
- `bounded_mailbox_drop_newest_caps_at_capacity` — 50 msgs, cap 10 → only 10 delivered
- `bounded_mailbox_drop_oldest_keeps_newest` — 10 msgs, cap 5 → newest 5 kept
- `unbounded_mailbox_delivers_all_messages` — backward compatibility check
- `bounded_mailbox_refills_after_processing` — cap 5, process, refill works
- **Result**: 64 tests pass, all workspace compiles
### Research Notes
- Full analysis in `CLAUDE/notes/research_synthesis.md`
- Baseline benchmarks in `CLAUDE/notes/baseline_benchmarks.md`
- Constraints in `CLAUDE/notes/constraints.md`
## Next Steps
- [x] **Cycle 2: Stress testing + property-based tests** ✅
- [x] **Cycle 3: Adaptive backoff with thread parking** ✅
- [x] **Cycle 4: Enhanced benchmarks + bug-inspired tests** ✅
- [x] **Cycle 5: Work stealing research + load-aware placement** ✅
- [x] **Cycle 6: Mailbox backpressure** ✅
- [x] **Cycle 7: Actor recovery (factory restart)** ✅
- [x] **Cycle 8: Dead actor cleanup** ✅
- [x] **Cycle 9: Lifecycle hooks + graceful stop** ✅
- [x] **Cycle 10: Actor timers (tick-counting)** ✅
- [x] **Cycle 11: Property-based testing (proptest-state-machine + fuzz extension)** ✅
- [ ] **Cycle 12: Next improvement**
- Candidates: named actors/registry (small effort, high value), actor monitoring/death watch,
actor groups/pub-sub, SmallBox optimization
- Priority messages REJECTED (lifecycle hooks cover 95% of cases)
- LIFO slot rejected (0-5% benefit for typical workloads, not worth complexity)
## Open Questions
- Should budget be configurable per-actor (not just per-runtime)?
- Is 64 the right default budget? Benchmarks show budget=32 slightly faster for throughput
- ~~Thread parking: notification mechanism~~ RESOLVED: OnceLock<Thread> + unpark()
- Should load-aware placement weight mailbox depth more than actor count?
- LIFO slot for same-worker sends: worth the complexity?
## Blockers
- (none)