swactor/WORKER.md
Zachery Aaron Shores-Chmielewski ed26f0c080 feat(wip): rewrite
Our first implementation is incredibly slow, and failed to scale with increased threads. In order to fix this, we
are implementing cache-aware batching of message processing, and an epoch-based actor rebalancing using shard pools.
2026-01-30 22:52:38 +07:00

4.7 KiB

struct MessageRing: buffer: [u8; 4096] head: AtomicUsize # Sender writes (Release) tail: AtomicUsize # Worker reads (Acquire)

struct LocalArena: buffer: [u8; 262144] # Raw message bytes from rings bump: usize

Key addition: Bucket buffer (indices into arena, not copies)

struct BucketBuffer: # Pre-allocated array of slices. Max 64K actors, resize if needed. # bucket[i] contains indices of messages for actor i. buckets: Vec<Vec> # Or flat Vec with head/tail if arena-allocated actor_order: Vec # Which actors have messages (for iteration)

struct Worker: worker_id: ID inbox_rings: Vec # Per-sender rings arena: LocalArena # Contiguous message storage buckets: BucketBuffer # Grouped by actor actor_table: Vec # ActorId -> Actor

Sender side (unchanged, 30 cycles)

function send_message(sender, target_worker, actor_id, payload): ring = sender.rings[target_worker] offset = reserve_in_ring(ring, 8 + len(payload)) serialize(ring.buffer[offset:], actor_id, len(payload), payload) ring.head.store(offset + 8 + len(payload), Release)

Worker side: Three-phase pipeline

function worker_run(worker): while true: # PHASE 1: DRAIN (same as before, ~5 cycles per message) # ----------------------------------------------- for ring in worker.inbox_rings: head = ring.head.load(Acquire) tail = ring.tail.load(Relaxed) if head == tail: continue

        # Copy sequential chunk from ring -> arena (hardware prefetch)
        size = head - tail
        memcpy(worker.arena.buffer[worker.arena.bump:], 
               ring.buffer[tail:], size)
        
        # Parse boundaries while copying to avoid second pass
        parse_and_bucket(worker.arena, worker.arena.bump, size, worker.buckets)
        
        worker.arena.bump += size
        ring.tail.store(head, Relaxed)
    
    if worker.arena.bump == 0: 
        cpu_relax()
        continue
    
    # PHASE 2: RADIX BUCKET (O(N), deterministic ~300 cycles)
    # ------------------------------------------------------
    # We already built buckets during parse_and_bucket above, 
    # but if we deferred parsing, do it now:
    
    # Option A: If parsed during drain (optimal)
    # Buckets already filled with (offset, len) pairs pointing into arena
    
    # Option B: Linear scan to build buckets (if raw bytes in arena)
    offset = 0
    while offset < worker.arena.bump:
        actor_id = read_u32(arena[offset:])
        msg_len = read_u32(arena[offset+4:])
        
        # Append to actor's bucket (Vec push, amortized O(1))
        # Each bucket entry: (offset, msg_len) = 16 bytes
        worker.buckets.buckets[actor_id].append((offset+8, msg_len))
        
        # Track unique actors (optional, avoids empty bucket scans)
        if worker.buckets.buckets[actor_id].len() == 1:
            worker.buckets.actor_order.append(actor_id)
        
        offset += 8 + msg_len
    
    # PHASE 3: PROCESS BY ACTOR (hidden message fetch, hot actor state)
    # ----------------------------------------------------------------
    for actor_id in worker.buckets.actor_order:
        actor = worker.actor_table[actor_id]  # First access: L3 miss (250 cycles)
        
        # Prefetch next actor's state while processing current (optional)
        prefetch_actor(worker.buckets.actor_order, worker.actor_table)
        
        # Process all messages for this actor
        # arena[slice] is L1 hit (12 cycles) - scanned sequentially within actor
        for (msg_offset, msg_len) in worker.buckets.buckets[actor_id]:
            msg_data = worker.arena.buffer[msg_offset : msg_offset+msg_len]
            actor.process(msg_data)  # 100 cycles work
        
        # actor state stays in L1 for entire inner loop
        
    # PHASE 4: RESET (zero cost)
    worker.arena.bump = 0
    clear_buckets(worker.buckets)  # Just reset lengths, don't free

Helper: Parse during drain to avoid touching bytes twice

function parse_and_bucket(arena, base_offset, size, buckets): ptr = 0 while ptr < size: actor_id = read_u32(arena[base_offset + ptr:]) msg_len = read_u32(arena[base_offset + ptr + 4:])

    # Append metadata to bucket (16 bytes: offset, len)
    buckets.buckets[actor_id].append((base_offset + ptr + 8, msg_len))
    
    if buckets.buckets[actor_id].len() == 1:
        buckets.actor_order.append(actor_id)
    
    ptr += 8 + msg_len