swactor/CACHE.md
Zachery Aaron Shores-Chmielewski ed26f0c080 feat(wip): rewrite
Our first implementation is incredibly slow, and failed to scale with increased threads. In order to fix this, we
are implementing cache-aware batching of message processing, and an epoch-based actor rebalancing using shard pools.
2026-01-30 22:52:38 +07:00

1050 lines
No EOL
33 KiB
Markdown

# Cache Locality Design Guidelines for Actor Runtimes
## Table of Contents
1. [Foundational Concepts](#foundational-concepts)
2. [Actor Data Structure Layout](#actor-data-structure-layout)
3. [Mailbox Design for Cache Efficiency](#mailbox-design-for-cache-efficiency)
4. [Memory Allocation Strategies](#memory-allocation-strategies)
5. [Scheduler Design for Locality](#scheduler-design-for-locality)
6. [NUMA-Aware Considerations](#numa-aware-considerations)
7. [Post-Design Tuning Strategies](#post-design-tuning-strategies)
8. [Framework-Specific Examples](#framework-specific-examples)
---
## Foundational Concepts
### Cache Hierarchy Understanding
Before designing your actor runtime, you must internalize how modern CPUs access memory:
```
┌─────────────────────────────────────────────────────────────┐
│ CPU Core │
│ ┌─────────┐ │
│ │ Registers│ ~0.5 cycles │
│ └────┬────┘ │
│ ▼ │
│ ┌─────────┐ │
│ │ L1 Cache│ 32-64KB, ~4 cycles, per-core │
│ └────┬────┘ │
│ ▼ │
│ ┌─────────┐ │
│ │ L2 Cache│ 256KB-1MB, ~12 cycles, per-core │
│ └────┬────┘ │
└───────┼─────────────────────────────────────────────────────┘
▼
┌─────────┐
│ L3 Cache│ 8-64MB, ~40 cycles, shared across cores
└────┬────┘
▼
┌─────────┐
│ RAM │ ~100-300 cycles
└─────────┘
```
**Key insight**: A cache miss at L1 that propagates to RAM can be 50-100x slower than a cache hit. Your actor runtime design should minimize these misses.
### Cache Lines: The Atomic Unit of Memory Transfer
Modern CPUs don't transfer individual bytes—they transfer **cache lines**, typically 64 bytes on x86-64 and ARM (though 128 bytes on some ARM implementations like Apple M-series for certain operations).
```cpp
// On most x86-64 systems:
constexpr size_t CACHE_LINE_SIZE = 64;
// C++17 provides portable constants:
#include <new>
// std::hardware_destructive_interference_size (typically 64)
// std::hardware_constructive_interference_size (typically 64)
```
This has two critical implications:
1. **Spatial Locality**: When you access one byte, the CPU fetches 63 more. Design your data structures so related data lives nearby.
2. **False Sharing**: When two threads write to different variables on the same cache line, the CPU invalidates that line for all cores, causing severe performance degradation.
---
## Actor Data Structure Layout
### Principle 1: Keep Hot Data Together
The most frequently accessed data during message processing should fit within one or two cache lines. Here's how CAF (C++ Actor Framework) approaches this:
```cpp
// Conceptual layout inspired by CAF's actor design
// CAF actors are designed to be only a few hundred bytes
struct alignas(64) Actor {
// === CACHE LINE 1: Critical scheduling data ===
std::atomic<uint32_t> state; // 4 bytes: running/waiting/done
uint32_t flags; // 4 bytes: various flags
MailboxPtr mailbox; // 8 bytes: pointer to mailbox
SchedulerPtr home_scheduler; // 8 bytes: affinity hint
ActorId id; // 8 bytes: unique identifier
RefCount ref_count; // 8 bytes: reference counting
BehaviorPtr current_behavior; // 8 bytes: message handler
// Remaining: 16 bytes for future use
char padding1[16];
// === CACHE LINE 2: Less frequently accessed ===
ActorPtr parent; // 8 bytes: supervision
ChildList children; // 16 bytes: supervised actors
ErrorHandler on_error; // 8 bytes: error handling
// ... additional metadata
};
static_assert(offsetof(Actor, parent) == 64,
"Second cache line should start at offset 64");
```
**Why this matters**: When the scheduler decides whether to run an actor, it only needs the first cache line. The message handler only needs the first two cache lines for most operations.
### Principle 2: Separate Hot and Cold Data
Actors often have data that's accessed frequently (hot) and data accessed rarely (cold). Separating these prevents cold data from evicting hot data from cache:
```cpp
// HOT: Accessed on every message dispatch
struct ActorHotData {
alignas(64) std::atomic<int> mailbox_count;
BehaviorFunction* behavior;
void* user_state; // Pointer to actual actor state
SchedulerWorker* affinity;
};
// COLD: Accessed only during lifecycle events or errors
struct ActorColdData {
std::string name;
ActorAddress parent;
std::vector<ActorAddress> children;
std::function<void(Error)> error_handler;
MonitorList monitors;
LinkList links;
CreationTimestamp created_at;
};
struct Actor {
ActorHotData hot; // First cache line(s)
ActorColdData* cold; // Pointer to cold data, allocated separately
};
```
### Principle 3: Prevent False Sharing in Scheduler Structures
When you have per-worker data structures, false sharing is your enemy. Here's the pattern used in work-stealing schedulers:
```cpp
// BAD: Dense packing causes false sharing
struct BadWorkerArray {
std::atomic<int> work_count[NUM_WORKERS]; // All on adjacent cache lines!
};
// GOOD: Cache-line padding prevents false sharing
struct alignas(64) PaddedCounter {
std::atomic<int> count;
char padding[64 - sizeof(std::atomic<int>)];
};
struct GoodWorkerArray {
PaddedCounter work_count[NUM_WORKERS]; // Each on its own cache line
};
```
The Rust ecosystem provides `cache-padded` (now deprecated but instructive):
```rust
// From the cache-padded crate (conceptually)
use std::sync::atomic::AtomicUsize;
// Concurrent queue with cache-padded indices
struct Queue<T> {
head: CachePadded<AtomicUsize>, // Own cache line
tail: CachePadded<AtomicUsize>, // Own cache line
buffer: *mut T,
}
```
CAF's work-stealing scheduler applies this pattern to its worker queues, ensuring that each worker's deque head/tail pointers don't share cache lines with other workers.
---
## Mailbox Design for Cache Efficiency
### MPSC Queue Optimizations (Multiple Producers, Single Consumer)
Actor mailboxes are typically MPSC queues. The design significantly impacts cache performance:
```cpp
// Intrusive linked-list approach (good for variable message sizes)
// Used conceptually by many actor frameworks
struct Message {
Message* next; // 8 bytes
MessageType type; // 4 bytes
uint32_t payload_size; // 4 bytes
alignas(16) char payload[]; // Flexible array member
};
class IntrusiveMailbox {
// Producer side (multiple threads write here)
alignas(64) std::atomic<Message*> tail;
// Consumer side (single thread reads here)
alignas(64) Message* head;
// Sentinel node to simplify empty-check
Message stub;
public:
void push(Message* msg) {
msg->next = nullptr;
Message* prev = tail.exchange(msg, std::memory_order_acq_rel);
prev->next = msg; // Linearization point
}
Message* pop() {
Message* h = head;
Message* next = h->next;
if (next) {
head = next;
// Return the message that was after stub, or actual message
return (h == &stub) ? pop() : h;
}
return nullptr;
}
};
```
**Cache analysis**:
- `tail` is on its own cache line (written by producers)
- `head` is on its own cache line (written by consumer)
- This prevents false sharing between producers and consumer
### Bounded Ring Buffer Approach (better cache locality for small messages)
```cpp
template<typename T, size_t Capacity>
class alignas(64) BoundedMailbox {
static_assert((Capacity & (Capacity - 1)) == 0,
"Capacity must be power of 2");
// Producer state - own cache line
alignas(64) std::atomic<size_t> write_pos{0};
// Consumer state - own cache line
alignas(64) size_t read_pos{0};
size_t cached_write{0}; // Cached write_pos to reduce atomic reads
// Buffer - contiguous for spatial locality
alignas(64) T buffer[Capacity];
public:
bool try_push(const T& item) {
size_t wp = write_pos.load(std::memory_order_relaxed);
size_t next = (wp + 1) & (Capacity - 1);
// Check if full (would need to load read_pos from consumer)
if (next == read_pos) return false;
buffer[wp] = item;
write_pos.store(next, std::memory_order_release);
return true;
}
bool try_pop(T& item) {
if (read_pos == cached_write) {
cached_write = write_pos.load(std::memory_order_acquire);
if (read_pos == cached_write) return false;
}
item = buffer[read_pos];
read_pos = (read_pos + 1) & (Capacity - 1);
return true;
}
};
```
**Cache benefits**:
- Sequential access patterns maximize hardware prefetching
- Bounded size means working set fits in cache
- The `cached_write` optimization reduces atomic operations
### CAF's Dual-Queue Strategy
CAF uses a double-ended queue with interesting cache locality properties:
```cpp
// Simplified version of CAF's mailbox concept
class CafStyleMailbox {
// Internal jobs (from same worker) - LIFO for cache locality
alignas(64) std::atomic<Job*> internal_head;
spinlock internal_lock;
// External jobs (from other workers) - FIFO for fairness
alignas(64) std::atomic<Job*> external_tail;
spinlock external_lock;
public:
// Called by the worker that owns this actor
void internal_enqueue(Job* job) {
// LIFO: new jobs go to front
// This maximizes temporal locality - recently touched
// actor state is still warm in cache
std::lock_guard<spinlock> lock(internal_lock);
job->next = internal_head.load(std::memory_order_relaxed);
internal_head.store(job, std::memory_order_release);
}
// Called by other workers sending messages
void external_enqueue(Job* job) {
// FIFO: maintains message ordering guarantees
std::lock_guard<spinlock> lock(external_lock);
job->next = nullptr;
Job* prev = external_tail.exchange(job, std::memory_order_acq_rel);
if (prev) prev->next = job;
}
};
```
The LIFO internal queue is crucial: when an actor sends a message to another actor and that message creates a response, processing the response immediately means the original actor's state is still cache-hot.
---
## Memory Allocation Strategies
### Arena Allocators: The Foundation of Cache-Friendly Allocation
Standard `malloc` has several cache-hostile properties:
1. Metadata overhead (typically 8-16 bytes per allocation)
2. Memory fragmentation scatters related objects
3. System call overhead for large allocations
Arena allocators solve these problems:
```cpp
class Arena {
char* buffer;
size_t capacity;
size_t offset;
public:
Arena(size_t size)
: buffer(static_cast<char*>(aligned_alloc(64, size)))
, capacity(size)
, offset(0) {}
void* allocate(size_t size, size_t alignment = alignof(std::max_align_t)) {
// Align the offset
size_t aligned_offset = (offset + alignment - 1) & ~(alignment - 1);
if (aligned_offset + size > capacity) {
return nullptr; // Or grow/chain arenas
}
void* ptr = buffer + aligned_offset;
offset = aligned_offset + size;
return ptr;
}
void reset() { offset = 0; } // "Free" everything at once
~Arena() { free(buffer); }
};
```
### Per-Actor Arenas
Each actor can have its own arena, ensuring that:
1. Actor state and its heap-allocated data are contiguous
2. GC is simplified (just reset the arena when actor dies)
3. No cross-thread allocation/deallocation (thread-local arenas)
```cpp
class ActorArena {
static constexpr size_t INITIAL_SIZE = 4096; // One page
struct Block {
Block* next;
size_t size;
size_t used;
alignas(16) char data[];
};
Block* current;
Block* blocks; // Linked list for cleanup
public:
void* allocate(size_t size) {
size = (size + 15) & ~15; // Align to 16 bytes
if (current->used + size > current->size) {
grow(size);
}
void* ptr = current->data + current->used;
current->used += size;
return ptr;
}
// For actor-local allocations that follow message processing
template<typename T, typename... Args>
T* create(Args&&... args) {
void* mem = allocate(sizeof(T));
return new (mem) T(std::forward<Args>(args)...);
}
};
```
### Message Pool Allocators
Messages often have predictable sizes. Pool allocators exploit this:
```cpp
template<size_t SlotSize, size_t SlotsPerBlock = 64>
class MessagePool {
static_assert(SlotSize >= sizeof(void*), "Slot must fit a pointer");
struct alignas(64) Block {
Block* next;
char slots[SlotSize * SlotsPerBlock];
};
Block* blocks = nullptr;
void* free_list = nullptr;
public:
void* allocate() {
if (!free_list) {
grow();
}
void* slot = free_list;
free_list = *static_cast<void**>(slot);
return slot;
}
void deallocate(void* ptr) {
*static_cast<void**>(ptr) = free_list;
free_list = ptr;
}
private:
void grow() {
Block* b = new Block;
b->next = blocks;
blocks = b;
// Thread all slots through free list
for (size_t i = 0; i < SlotsPerBlock; ++i) {
void* slot = b->slots + i * SlotSize;
*static_cast<void**>(slot) = free_list;
free_list = slot;
}
}
};
// Usage: Pool for common message sizes
MessagePool<64> small_messages; // ≤48 byte payload
MessagePool<256> medium_messages; // ≤240 byte payload
MessagePool<1024> large_messages; // ≤1008 byte payload
```
### Pony's Per-Actor Heap Approach
Pony takes an extreme approach: each actor has its own heap, and the garbage collector (ORCA) runs per-actor without stop-the-world pauses:
```
// Pony's conceptual memory model
Actor {
local_heap: Heap, // Only this actor allocates here
reference_counts: Map, // Track foreign references
message_queue: Queue, // Incoming messages
}
// Key insight: GC only runs when actor is not executing a behavior
// This means:
// 1. No stack to scan (no stack map needed)
// 2. No safepoints required
// 3. No synchronization with other actors during GC
```
---
## Scheduler Design for Locality
### Work-Stealing with Locality Awareness
The basic work-stealing algorithm is cache-oblivious. Here's how to make it cache-aware:
```cpp
class LocalityAwareScheduler {
struct alignas(64) WorkerState {
std::deque<Actor*> local_queue;
std::atomic<bool> is_active{true};
uint32_t numa_node;
uint32_t core_id;
// Steal victims ordered by locality (nearest first)
std::vector<uint32_t> steal_order;
};
std::vector<WorkerState> workers;
public:
void initialize_steal_order(uint32_t worker_id) {
WorkerState& w = workers[worker_id];
// Build steal order: same NUMA node first, then others
for (uint32_t i = 0; i < workers.size(); ++i) {
if (i == worker_id) continue;
if (workers[i].numa_node == w.numa_node) {
// Same NUMA node - insert at front
w.steal_order.insert(w.steal_order.begin(), i);
} else {
// Different NUMA node - append at end
w.steal_order.push_back(i);
}
}
}
Actor* try_steal(uint32_t worker_id) {
WorkerState& w = workers[worker_id];
// Try stealing in locality order
for (uint32_t victim_id : w.steal_order) {
WorkerState& victim = workers[victim_id];
// Steal from back (FIFO) to maintain parent-child locality
if (!victim.local_queue.empty()) {
Actor* stolen = victim.local_queue.back();
victim.local_queue.pop_back();
return stolen;
}
}
return nullptr;
}
};
```
### CAF's Scheduler Design
CAF's scheduler uses work-stealing with specific optimizations for cache locality:
```cpp
// From CAF's conceptual design
class CafScheduler {
// Each worker has its own deque
// Internal enqueue: LIFO (front) - maximizes cache reuse
// External enqueue: FIFO (back) - fair ordering
// Stealing: FIFO (back) - steals "cold" work
void worker_loop(WorkerId id) {
while (running) {
Actor* actor = try_get_local(id);
if (!actor) {
actor = try_steal_from_random(id);
}
if (actor) {
// Process messages until actor blocks or quota exhausted
run_actor(actor);
// If actor has more work, re-enqueue locally (LIFO)
// This keeps hot actors on the same core
if (actor->has_pending_messages()) {
internal_enqueue(id, actor);
}
} else {
// No work available - sleep briefly
sleep_or_poll(id);
}
}
}
};
```
### Affinity-Based Scheduling
For maximum cache efficiency, actors can have "home" workers:
```cpp
struct Actor {
std::atomic<WorkerId> home_worker{NO_AFFINITY};
// ... other fields
};
class AffinityScheduler {
void enqueue(Actor* actor, WorkerId sender) {
WorkerId home = actor->home_worker.load(std::memory_order_relaxed);
if (home == NO_AFFINITY) {
// First run: assign to sender's worker for initial locality
actor->home_worker.store(sender, std::memory_order_relaxed);
workers[sender].enqueue(actor);
} else {
// Subsequent runs: prefer home worker
workers[home].enqueue(actor);
}
}
// Periodically rebalance to prevent hot spots
void rebalance() {
// Move actors with high message rates to dedicated workers
// Consider communication patterns: actors that communicate
// frequently should be on the same core/NUMA node
}
};
```
---
## NUMA-Aware Considerations
### Understanding NUMA Topology
On multi-socket systems, memory access time depends on which CPU socket allocated the memory:
```
┌─────────────────┐ QPI/UPI ┌─────────────────┐
│ Socket 0 │◄────────────────►│ Socket 1 │
│ ┌───────────┐ │ │ ┌───────────┐ │
│ │ 8 Cores │ │ │ │ 8 Cores │ │
│ └───────────┘ │ │ └───────────┘ │
│ ┌───────────┐ │ │ ┌───────────┐ │
│ │ Local RAM │ │ │ │ Local RAM │ │
│ │ 32GB │ │ │ │ 32GB │ │
│ └───────────┘ │ │ └───────────┘ │
└─────────────────┘ └─────────────────┘
Local memory access: ~100ns
Remote memory access: ~300ns (3x slower!)
```
### NUMA-Aware Actor Placement
```cpp
class NumaAwareScheduler {
struct NumaNode {
std::vector<WorkerId> workers;
Arena* local_arena; // Memory allocated on this node
std::atomic<size_t> actor_count{0};
};
std::vector<NumaNode> nodes;
public:
void spawn_actor(ActorFactory factory, NumaHint hint = HINT_LOCAL) {
int target_node;
switch (hint) {
case HINT_LOCAL:
target_node = current_numa_node();
break;
case HINT_NEAR_PARENT:
target_node = parent_actor->numa_node;
break;
case HINT_LEAST_LOADED:
target_node = find_least_loaded_node();
break;
}
// Allocate actor memory on target NUMA node
NumaNode& node = nodes[target_node];
void* mem = node.local_arena->allocate(sizeof(Actor));
Actor* actor = new (mem) Actor(factory);
actor->numa_node = target_node;
// Assign to a worker on the same node
WorkerId worker = node.workers[round_robin++ % node.workers.size()];
actor->home_worker = worker;
}
// Work stealing prefers same-node victims
Actor* steal(WorkerId thief) {
int thief_node = worker_to_node[thief];
// First: try workers on same NUMA node
for (WorkerId victim : nodes[thief_node].workers) {
if (Actor* a = try_steal_from(victim)) {
return a;
}
}
// Then: try other nodes (expensive - remote memory!)
for (int n = 0; n < nodes.size(); ++n) {
if (n == thief_node) continue;
for (WorkerId victim : nodes[n].workers) {
if (Actor* a = try_steal_from(victim)) {
return a;
}
}
}
return nullptr;
}
};
```
### Research-Backed NUMA Strategies
The paper "Work-Stealing, Locality-Aware Actor Scheduling" by Barghi et al. identifies key strategies:
1. **Hub Actor Identification**: Actors that spawn many short-lived children should keep those children on the same NUMA node.
2. **Communication Locality**: Actors that frequently message each other should be co-located.
3. **Hierarchical Stealing**: Steal from local workers first, then socket-local, then remote.
```cpp
// Implementing hierarchical stealing based on the research
class HierarchicalStealer {
// Topology levels: Core → L3 Cache Group → Socket → System
struct Level {
std::vector<WorkerId> peers;
double steal_probability; // Higher for closer levels
};
std::vector<std::vector<Level>> topology; // [worker][level]
Actor* steal(WorkerId thief) {
for (const Level& level : topology[thief]) {
// Probabilistic stealing at each level
if (random() < level.steal_probability) {
WorkerId victim = level.peers[random() % level.peers.size()];
if (Actor* a = try_steal_from(victim)) {
return a;
}
}
}
return nullptr;
}
};
```
---
## Post-Design Tuning Strategies
### Profiling for Cache Misses
Once your runtime is working, these tools help identify cache problems:
```bash
# Linux perf for cache miss analysis
perf stat -e cache-references,cache-misses,L1-dcache-load-misses \
./your_actor_runtime
# Detailed cache analysis
perf record -e cache-misses ./your_actor_runtime
perf report
# NUMA statistics
numastat -p $(pgrep your_actor_runtime)
```
### Tunable Parameters
Design your runtime with these tunable knobs:
```cpp
struct RuntimeConfig {
// Scheduler tuning
size_t messages_per_actor_run = 10; // Batch size before yielding
size_t steal_batch_size = 1; // How many actors to steal
double steal_probability = 0.5; // Work-stealing aggressiveness
// Memory tuning
size_t actor_arena_initial = 4096; // Initial arena size
size_t actor_arena_max = 1024 * 1024; // Maximum before GC
size_t message_pool_slots = 1024; // Pool allocator size
// Cache tuning
size_t prefetch_distance = 3; // Mailbox prefetch depth
bool enable_numa_awareness = true; // NUMA-local allocation
bool enable_affinity = true; // Worker-actor affinity
};
```
### Prefetching Strategies
When you know you'll access data soon, tell the CPU:
```cpp
void process_mailbox(Actor* actor) {
Message* current = actor->mailbox.head;
while (current) {
Message* next = current->next;
// Prefetch next message while processing current
if (next) {
__builtin_prefetch(next, 0, 3); // Read, high temporal locality
__builtin_prefetch(next->payload, 0, 3);
}
process_message(actor, current);
current = next;
}
}
```
### Batch Processing for Cache Warmth
Process multiple messages while actor state is cache-hot:
```cpp
void run_actor(Actor* actor, size_t max_messages = 10) {
auto& behavior = actor->current_behavior;
// Keep processing while mailbox non-empty and quota not exhausted
size_t processed = 0;
while (processed < max_messages) {
Message* msg = actor->mailbox.try_pop();
if (!msg) break;
// Actor state is cache-hot from previous iteration
behavior->handle(actor, msg);
++processed;
}
// Record for tuning: if we hit quota, actor is "hot"
if (processed == max_messages) {
actor->hot_count++;
}
}
```
### Adaptive Granularity (ElasticActor Approach)
Research shows that optimal batch size depends on workload:
```cpp
class AdaptiveScheduler {
struct ActorStation {
std::vector<Actor*> actors;
Mailbox shared_mailbox;
// Actors in same station share a mailbox
// Coarser granularity = better cache utilization
// Finer granularity = better parallelism
};
std::vector<ActorStation> stations;
std::atomic<size_t> station_count;
void adjust_granularity() {
double avg_latency = measure_message_latency();
double target_latency = 1000; // 1 microsecond target
if (avg_latency > target_latency * 1.5) {
// Latency too high: split stations for more parallelism
split_hottest_station();
} else if (avg_latency < target_latency * 0.5) {
// Latency low: merge stations for better cache use
merge_coldest_stations();
}
}
};
```
---
## Framework-Specific Examples
### CAF (C++ Actor Framework)
CAF's key cache optimizations:
```cpp
// 1. Minimal actor footprint (~240 bytes)
// From CAF's design philosophy:
// "CAF actors consist of only a few hundred bytes"
// 2. Work-stealing with LIFO local queuing
// internal_enqueue adds to front (LIFO) for cache locality
// external_enqueue adds to back (FIFO) for fairness
// 3. Type-safe messaging reduces runtime overhead
template<typename... Args>
void typed_send(typed_actor<Args...>& receiver, Args... args) {
// Compile-time type checking eliminates runtime type dispatch
// Smaller message headers, better cache utilization
}
// 4. Configurable scheduler
auto cfg = actor_system_config{}
.set("caf.scheduler.max-threads", 8)
.set("caf.scheduler.policy", "stealing"); // or "sharing"
```
### Actix (Rust)
Actix leverages Rust's ownership for cache-friendly design:
```rust
use actix::prelude::*;
// Actors own their state - no shared mutable state
struct MyActor {
// State is exclusively owned, always cache-local to processing
counter: u64,
buffer: Vec<u8>, // Can pre-allocate for known workloads
}
impl Actor for MyActor {
type Context = Context<Self>;
fn started(&mut self, ctx: &mut Self::Context) {
// Set mailbox capacity to control memory usage
ctx.set_mailbox_capacity(16); // Default is 16
}
}
// Messages are moved, not copied - zero-copy when possible
struct DataMessage {
payload: Vec<u8>, // Ownership transferred, not cloned
}
impl Handler<DataMessage> for MyActor {
type Result = ();
fn handle(&mut self, msg: DataMessage, _ctx: &mut Context<Self>) {
// msg.payload is now owned by this actor
// No cache invalidation from other threads
self.buffer = msg.payload;
}
}
```
### Pony
Pony's extreme approach to cache locality:
```pony
// Each actor has isolated heap - perfect cache locality for GC
actor Counter
var _count: U64 = 0
// All state is actor-local, always cache-hot when processing
be increment() =>
_count = _count + 1
be get(main: Main) =>
// Sending is zero-copy for immutable data
main.print(_count)
// Reference capabilities ensure no data races
// This allows GC without read/write barriers
// GC runs only between behaviors - no stack scanning needed
```
Key Pony cache insights:
- Per-actor heaps mean allocation is always thread-local
- No stop-the-world GC means no cache thrashing from GC threads
- Zero-copy messaging through capability-based types
---
## Benchmarking Your Optimizations
### Micro-benchmarks
```cpp
// Message passing latency
void benchmark_mailbox_latency() {
Actor sender, receiver;
auto start = high_resolution_clock::now();
for (int i = 0; i < 1'000'000; ++i) {
sender.send(receiver, PingMessage{});
receiver.process_one();
}
auto end = high_resolution_clock::now();
auto ns_per_message = duration_cast<nanoseconds>(end - start).count() / 1'000'000;
std::cout << "Latency: " << ns_per_message << " ns/msg\n";
}
// Actor creation throughput
void benchmark_actor_creation() {
auto start = high_resolution_clock::now();
for (int i = 0; i < 100'000; ++i) {
auto actor = spawn<SimpleActor>();
// Actor immediately becomes garbage
}
auto end = high_resolution_clock::now();
auto actors_per_sec = 100'000.0 / duration_cast<seconds>(end - start).count();
std::cout << "Creation: " << actors_per_sec << " actors/sec\n";
}
```
### Monitoring Production Performance
```cpp
class SchedulerMetrics {
// Per-worker metrics (cache-padded to avoid false sharing)
struct alignas(64) WorkerMetrics {
std::atomic<uint64_t> messages_processed{0};
std::atomic<uint64_t> actors_run{0};
std::atomic<uint64_t> steal_attempts{0};
std::atomic<uint64_t> steal_successes{0};
std::atomic<uint64_t> cache_misses_estimate{0}; // From perf counters
};
std::vector<WorkerMetrics> workers;
public:
void report() {
double steal_success_rate = total_steal_successes / total_steal_attempts;
// Low steal success rate + high CPU might indicate poor locality
double messages_per_actor = total_messages / total_actors_run;
// High value indicates good batching
}
};
```
---
## Summary: Cache Locality Checklist
### Data Structure Design
- [ ] Actor hot data fits in 1-2 cache lines
- [ ] Scheduler per-worker state is cache-line padded
- [ ] Mailbox head/tail are on separate cache lines
- [ ] Cold data is allocated separately from hot data
### Memory Allocation
- [ ] Per-actor or per-worker arenas for locality
- [ ] Pool allocators for common message sizes
- [ ] NUMA-aware allocation on multi-socket systems
### Scheduling
- [ ] LIFO local scheduling for cache reuse
- [ ] Batch processing (multiple messages per actor run)
- [ ] Locality-aware work stealing (same NUMA node first)
- [ ] Actor affinity to workers
### Tuning
- [ ] Profile with `perf` for cache miss rates
- [ ] Configurable batch sizes and steal policies
- [ ] Prefetching for predictable access patterns
- [ ] Adaptive granularity based on latency metrics
---
## References
1. CAF: The C++ Actor Framework - https://actor-framework.org/
2. Barghi et al. "Work-Stealing, Locality-Aware Actor Scheduling" (IPDPS 2018)
3. Clebsch et al. "Orca: GC and Type System Co-Design for Actor Languages" (OOPSLA 2017)
4. Actix - Rust Actor Framework - https://actix.rs/
5. Pony Language - https://www.ponylang.io/
6. "Cache Line Alignment in C++" - Various Medium articles
7. Protocol Buffers Arena Allocation - https://protobuf.dev/reference/cpp/arenas/