NOTICE: AI GENERATED SLOP. KNOWN TO WORK, BUT BARELY REVIEWED. TAKE APPROPRIATE PERCAUTIONS IN YOUR DOWNSTREAM AI GENERATED SLOP.
Zero-serialization, zero-copy shared-memory RPC for .NET.
A sluice is a channel with a gate that controls fast flow. Sluice is a memory-mapped ring (the channel) with a wait-handle doorbell (the gate): messages flow between processes on one machine and are read in place — the bytes sitting in shared memory are the object. No serialize on send, no deserialize on receive, no allocation on the hot path.
It does two things:
- Point-to-point RPC — the daemon + thin-CLI pattern: a long-lived process holds expensive in-memory state, and short-lived command-line invocations discover it and exchange request/response messages — far cheaper than re-spinning that state per call, and far faster than HTTP, gRPC, or named pipes.
- A multi-way fabric — broadcast (1→N), a many-to-many bus (N↔N), and an addressed, full-duplex peer mesh, all over a Disruptor-style multicast ring with selectable reliable or lossy delivery.
Status: working, tested (42 tests green), benchmarked. Cross-platform — verified on both: the full suite passes on Windows and in a .NET 10 Linux container (
/dev/shmfile-backed maps + polling doorbell), and all three libraries build clean fornet8.0(the ACA LTS runtime) on Linux. On Windows the shared region is a named (page-file-backed) map with a wait-handle doorbell; on Linux it is a/dev/shm(tmpfs) file-backed map with a polling doorbell — so it runs in Azure Container Apps (Linux) as well as on Windows. The framing, cursor, and discovery (namedMutex+ lease file) logic is shared across both.
Sluice publishes to GitHub Packages (Sluice, Sluice.Gossip, Sluice.Fusion, Sluice.Supergraph). Each push to master
ships a CI prerelease (0.1.0-ci.<n>); tagging vX.Y.Z ships the stable X.Y.Z. Both target net8.0 and
net10.0.
GitHub Packages requires authentication even for reads, so add the feed once with a
personal access token that has the read:packages scope:
dotnet nuget add source "https://nuget.pkg.github.com/erinloy/index.json" \
--name sluice-github \
--username <your-github-username> \
--password <YOUR_PAT_WITH_read:packages> \
--store-password-in-clear-textThen reference it from any project:
dotnet add package Sluice # core RPC + multi-way fabric
dotnet add package Sluice.Gossip # epidemic dissemination (optional)
dotnet add package Sluice.Fusion # attach/mirror/cache overlay (optional)
dotnet add package Sluice.Supergraph # federated reactive graph (optional)To pin the latest CI prerelease, add --prerelease. Or declare the source in a repo-local nuget.config
(commit this; tokens stay in your environment, not the file):
<?xml version="1.0" encoding="utf-8"?>
<configuration>
<packageSources>
<add key="sluice-github" value="https://nuget.pkg.github.com/erinloy/index.json" />
</packageSources>
<packageSourceCredentials>
<sluice-github>
<add key="Username" value="%GITHUB_ACTOR%" />
<add key="ClearTextPassword" value="%GITHUB_TOKEN%" />
</sluice-github>
</packageSourceCredentials>
</configuration>Inside GitHub Actions the built-in GITHUB_TOKEN already authorizes the owner's feed — no PAT needed.
Every other local-IPC option in .NET copies and serializes:
- gRPC / MagicOnion — HTTP/2 framing + protobuf/MessagePack encode/decode, even on loopback.
- Named pipes — a kernel stream you must frame and serialize over.
- Cloudtoid.Interprocess — a genuine shared-memory queue, but
Dequeuecopies the message body out of the ring into a pooled/allocated buffer (it clears the slot and advances the read cursor immediately, so it cannot hand back a stable view). You still serialize your object on top.
Sluice keeps the slot alive until you call AdvanceRead(), so the consumer gets a ReadOnlySpan<byte>
pointing directly into the shared pages. Combined with blittable message headers read via
MemoryMarshal.AsRef, the read path is a pointer reinterpret — there is nothing to deserialize.
// ---- Owner (the daemon that holds state) ----
using var server = new SluiceServer("my-endpoint", (in RpcContext ctx) =>
{
// ctx.Request is a ReadOnlySpan<byte> pointing straight into shared memory — no copy, no decode.
switch (ctx.Kind)
{
case 1: ctx.Reply(ctx.Request); break; // echo
default: ctx.Reply("unknown"u8, ok: false); break;
}
});
server.Start();
SluiceDiscovery.Heartbeat("my-endpoint", 1 << 20); // publish a lease so clients can find us
// ---- Client (a short-lived CLI invocation) ----
if (!SluiceDiscovery.IsAlive("my-endpoint", TimeSpan.FromSeconds(10)))
return; // no daemon running
using var client = new SluiceClient("my-endpoint", exclusiveProducer: true);
RpcResponse resp = client.Send(kind: 1, "ping"u8);
Console.WriteLine(resp.Text); // "ping"Streaming responses:
foreach (var item in client.SendStream(kind: 3, request))
Process(item); // owner pushed N frames, ended with Complete()// Broadcast (1→N) or many-to-many bus (N↔N) — any process opens the same topic by name.
using var topic = new ShmTopic("prices", maxPayload: 256, mode: DeliveryMode.Lossy);
var sub = topic.Subscribe(); // each subscriber sees every message
topic.Publish(quoteBytes);
if (sub.TryRead(out var span)) { Use(span); sub.Advance(); } // read in place
// Addressed, full-duplex peer mesh — every peer both sends and receives, no client/server roles.
using var alice = new ShmPeer("alice");
using var bob = new ShmPeer("bob");
alice.Send("bob", "hello"u8);
bob.Receive(out var msg); // -> "hello"
bob.Send("alice", "hi back"u8); // full duplexDeliveryMode.Reliable never drops (the producer backpressures on the slowest subscriber);
DeliveryMode.Lossy never blocks (a lapped subscriber resyncs forward and reports Dropped).
Sluice.Gossip — epidemic dissemination: a replicated last-writer-wins key/value store that converges
across all processes. A local Set broadcasts the entry (O(1) on shared memory); periodic anti-entropy
re-gossips a random subset to repair lossy drops and bring late joiners up to date.
using var node = new GossipNode("cluster", "node-a");
node.Start();
node.Set("x", "1"u8.ToArray()); // converges to every node; LWW by Lamport version
node.TryGet("x", out var value);Sluice.Fusion — a Fusion-style attach/mirror/cache overlay (faithful to ActualLab/Fusion).
A ModelHost owns values and publishes a tiny invalidation (key + version) on a change — never the value.
A ModelMirror caches what it reads and refetches lazily only when the origin invalidates it: the Fusion
"signal stale, recompute on access" lifecycle, peer-to-peer over shared memory.
using var host = new ModelHost("prices");
host.Set("AAPL", quoteBytes);
using var mirror = new ModelMirror("prices");
SluiceComputed c = mirror.Get("AAPL"); // fetched once, then cached
// ... host.Set("AAPL", ...) → c.State becomes Invalidated → next Get refetches
using var state = new MirroredState(mirror, "AAPL");
state.Updated += v => Render(v); // push feed, like Fusion's ComputedState<T>Sluice.Supergraph — a federated reactive graph: Sluice.Fusion's invalidate→refetch model generalized
over the fabric so it spans hosts. Vertices are owned by participants across shared memory
and the network; a read of a vertex you own is a direct local read, a remote one routes a fetch over the fabric,
and a change broadcasts a tiny invalidation across the whole federation. A ComputedVertex whose dependencies live
on other participants recomputes and re-publishes when they change — so a change propagates transitively and many
processes behave as one logical reactive graph.
using var prices = new GraphPeer(new ShmTransport("prices"), "book");
using var risk = new GraphPeer(new ShmTransport("risk"), "book");
using var desk = new GraphPeer(new ShmTransport("desk"), "book");
var spot = prices.Define("spot", codec);
using var exposure = risk.Computed("exposure", codec, // depends on a vertex owned elsewhere
() => risk.Get(spot.Id, codec).Value * 100.0, spot.Id);
using var view = desk.Observe(exposure.Id, codec); // reactive view of a third participant's vertex
view.Updated += v => Render(v);
spot.Set(43.25); // ripples prices → risk → desk across two participant boundariesGive a GraphPeer a FederatedTransport instead of a bare ShmTransport and the same graph spans hosts. See
the supergraph guide.
A tiny in-memory key-value store served by a daemon and driven by a thin CLI (kvd):
kvd serve # start the daemon (holds the key-value state in memory)
kvd set foo bar # a separate process: discovers the daemon, sets a key
kvd get foo # -> bar (read from the running daemon, not reloaded)
kvd status # -> alive; keys=1; pid=12345
set/get/status never load the store — they reach into the already-running serve process over shared
memory. That is the efficiency win the whole library exists for.
┌─────────────────────────── shared memory (memory-mapped file) ───────────────────────────┐
│ header: [writeCursor | readCursor | capacity | magic] data: [len|payload][len|payload]… │
└────────────────────────────────────────────────────────────────────────────────────────────┘
producer ──write payload, publish writeCursor (release)──▶ ◀──read in place, advance readCursor (release)── consumer
doorbell: named EventWaitHandle (spin → block)
ShmRing— an SPSC lock-free ring in a memory-mapped file. Monotonic 64-bit cursors; frames are 4-byte aligned and never straddle the wrap (a skip-marker pads the tail). The consumer reads in place and only frees the slot onAdvanceRead().- RPC — one MPSC request ring (many clients → one owner; concurrent producers serialized by a named
mutex, or lock-free in
exclusiveProducermode) plus a per-client SPSC response ring, correlated by a 16-byte id in the blittableRpcHeader. Unary and streaming. ShmMulticast— the multi-way core: a Disruptor-style sequenced ring. Producers claim a monotonic sequence with oneInterlocked.Increment, write their slot, then stamp the slot's sequence as the publish barrier. Each subscriber holds an independent cursor and reads every message in place. The per-slot stamp is also the lap detector that makes lossy delivery (and torn-read re-validation) possible.ShmTopicandShmPeerare thin façades over it for pub/sub, bus, and addressed full-duplex mesh.- Discovery — a system-global named mutex elects the single owner; a lease file (pid + heartbeat +
capacity) lets a fresh CLI find the live instance and tell a crash from a running process by staleness.
Peers announce presence the same way (
ShmPeer.IsPeerAlive).
The ring protocol is checked with Microsoft Coyote: its systematic scheduler explores the producer/consumer interleavings exhaustively and proves no schedule loses, reorders, or tears a message (see delivery & safety).
Deeper guides live in docs/:
- Architecture — the layers, the zero-copy ring protocol, the sequenced multicast ring, the OS floor, and the correctness model.
- Choosing a topology — a decision table and a copy-paste cookbook for RPC, frame channels, topics, peers, gossip, and Fusion.
- The fabric — the transport-agnostic seam: one channel spanning local shared-memory participants (zero-copy) and remote participants over TCP, indistinguishably.
- The supergraph — a federated reactive graph over the fabric: Fusion's invalidate→refetch, generalized so many processes form one logical reactive graph across shared memory and the network.
- Delivery & safety — reliable vs lossy, torn-read handling, the memory model, crash robustness, and the systematic concurrency proof, failure mode by failure mode.
- Deployment & cross-platform — Windows vs Linux internals, container
/dev/shmsizing, and Azure Container Apps. - Performance — the numbers, where they come from, and the knobs that move them.
Contributions: see CONTRIBUTING.md. Security: see SECURITY.md. Release history: CHANGELOG.md.
- Cross-platform, with one asymmetry: the doorbell. Windows backs the region with a named
(page-file) map and signals readers via a named
EventWaitHandle(sub-µs wakeups). Linux — where named maps and named wait handles aren't supported in .NET — backs it with a/dev/shm(tmpfs) file map and has no OS doorbell, so a blocked reader polls with an escalating backoff capped at ~1 ms. Throughput and the in-place read path are identical; only the idle→first-message wakeup latency differs (spin-hot for the first ~64 iterations, then ≤1 ms). NamedMutex(discovery + the MPSC producer lock) works on both. The owner unlinks its tmpfs file on dispose. This is what lets Sluice run in Azure Container Apps (Linux). exclusiveProducerskips the cross-process producer mutex (a kernel transition per send) for the common single-CLI case. It still syncs the shared cursor — it only drops the mutex, never correctness — but you must not have two "exclusive" producers writing concurrently.- Crash robustness. A client that dies mid-request can't crash the daemon (each request is isolated).
The named mutex surfaces
AbandonedMutexExceptionas inherited ownership, so a crashed owner's endpoint is reclaimable. - Response-ring cache is bounded (FIFO, 256) so a stream of short-lived clients can't leak handles; an evicted persistent client just pays one re-open.
- Reliable multicast evicts a dead subscriber. A registered reliable subscriber that stalls would
backpressure every producer on the topic — so each subscriber carries a heartbeat lease (
LeaseMs, default 30 s), and a producer reclaims the cell of any subscriber whose lease has lapsed rather than wait on it forever. A subscriber that may pause longer than the lease between reads should callHeartbeat();Lossyis crash-robust regardless (a dead subscriber is simply lapped). - Lossy in-place reads can tear if a producer laps you mid-read; the
TryReadCopy/ShmPeer/ShmTopicconvenience paths re-validate the sequence stamp after copying, so use those for lossy unless you validate yourself. - Container
/dev/shmsizing (Linux). A ring's backing file lives in tmpfs, which counts against the container's shared-memory limit (often a 64 MB default). Size the total of your rings under that limit, or raise it — in Azure Container Apps via anEmptyDirvolume withstorageType: Memorymounted at/dev/shm, or in Docker via--shm-size. If/dev/shmis absent, Sluice falls back to the temp dir (real disk). - Not a network transport, and not a durable queue — it's the fastest path between processes on one box. Pair it with a durable store (FASTER, a log, etc.) if you need persistence.
src/Sluice the library
ShmRing.cs SPSC zero-copy ring (RPC transport)
ShmMulticast.cs Disruptor-style multi-way ring (broadcast / bus / mesh core)
ShmMap.cs the one OS split: named map (Windows) vs /dev/shm file map (Linux)
Rpc/ SluiceServer, SluiceClient, discovery
Multiway/ ShmTopic (pub/sub + bus), ShmPeer (addressed full-duplex mesh)
Fabric/ transport-agnostic seam: ITransport / IChannel over local-shm + remote
IFrameChannel.cs frame transport seam (drop-in for a duplex Stream, zero-copy on read)
ShmFrameChannel.cs bidirectional duplex channel = two rings
ShmFrameListener.cs Accept() — multiplex many client connections onto one daemon
src/Sluice.Gossip epidemic dissemination (GossipNode: LWW store, rumor + anti-entropy)
src/Sluice.Fusion Fusion-style overlay (ModelHost/ModelMirror/MirroredState, lazy invalidation)
src/Sluice.Supergraph federated reactive graph (GraphPeer: source/observed/computed vertices over the fabric)
tests/Sluice.Tests xUnit suite (ring, RPC unary/stream/MPSC, multicast MP-claim/lossy/reliable,
lease eviction, broadcast/bus/mesh/duplex, fabric, supergraph, discovery)
tests/Sluice.Concurrency Microsoft Coyote systematic-concurrency tests for the ring protocol
bench/Sluice.Benchmarks BenchmarkDotNet vs Cloudtoid + named pipes, serialization microbench, multicast
per-op, and the `throughput` aggregate harness
samples/Sluice.Demo the `kvd` daemon + thin-CLI demo
samples/Sluice.SupergraphDemo the federated reactive graph demo (prices → risk → desk)
Measured with BenchmarkDotNet v0.14.0 on Windows 11, .NET 10.0.9 (X64 RyuJIT
AVX-512), Concurrent Server GC. The serialization and multicast tables are the full default job (stable,
in-process); allocation figures are deterministic everywhere. See docs/performance.md for
the full methodology and the load-sensitivity caveat on cross-process latency. Reproduce with dotnet run -c Release --project bench/Sluice.Benchmarks -- --filter '*'.
This is the cost Sluice's hot path eliminates — reading a message is a pointer reinterpret
(MemoryMarshal.AsRef), not a decode. Measured against decoding the same message with
MemoryPack,
MessagePack, and
System.Text.Json:
| Reading one message | Payload | Mean | Allocated |
|---|---|---|---|
| Blittable (in place) | 64 B | 6.5 ns | 0 B |
| MemoryPack deserialize | 64 B | 107.2 ns | 272 B |
| MessagePack deserialize | 64 B | 313.3 ns | 280 B |
| Json deserialize | 64 B | 667.1 ns | 336 B |
| Blittable (in place) | 1024 B | 15.6 ns | 0 B |
| MemoryPack deserialize | 1024 B | 392.1 ns | 2192 B |
| MessagePack deserialize | 1024 B | 615.3 ns | 2200 B |
| Json deserialize | 1024 B | 1,182.9 ns | 2576 B |
Reading a message in place is ~16× faster than MemoryPack, ~100× faster than JSON, and allocates nothing (vs 270 B – 2.5 KB).
A full round trip through real shared memory / pipes (responder on a background thread; the payload still crosses the OS boundary). The alternatives must serialize a message to carry the payload — Sluice sends it raw.
Allocation — exact and load-independent:
| Transport | Allocated | vs Sluice |
|---|---|---|
| Sluice (zero-alloc receive) | 0 B | 0× |
Sluice (convenience Send) |
280 B | 1.0× |
| Cloudtoid + MemoryPack | 2008 B | 7.2× |
| Cloudtoid + JSON | 2520 B | 9.0× |
| Named pipe + MemoryPack | 2072 B | 7.4× |
| Named pipe + JSON | 2586 B | 9.2× |
Latency (clean run — two transport classes): Sluice 471 ns (zero-alloc 439 ns) vs Cloudtoid — also a shared-memory queue — 1065 ns (~2.3×) vs named pipes, a syscall transport, ~32 µs (~68×). So against a serialized shared-memory queue Sluice is ~2–4× faster and 7–9× lighter; against a kernel transport it is ~70× faster. Under machine load the syscall gap widens (OS scheduling enters the pipe's critical path, not the shared-memory read), so the named-pipe ratio grows on a busy server. Latency is hardware-sensitive — see performance.md for the full method and how to reproduce.
The 280 B is the response copied out across the API boundary by the convenience
Sendoverload. The span-callback overloadSend<TState>(kind, payload, state, reader)hands you the response in place and allocates 0 B/op at near-identical latency — the win is GC-pressure elimination. Use it on hot request paths to remove client-side allocation entirely.
The multicast ring's per-message cost — what sets the throughput ceiling. Read-in-place, zero allocation.
| Operation | 32 B | 256 B | Allocated |
|---|---|---|---|
| Publish (fan-out write) | 8.1 ns | 13.8 ns | 0 B |
| Publish + consume in place | 12.6 ns | 17.9 ns | 0 B |
dotnet run -c Release --project bench/Sluice.Benchmarks -- throughput
| Producers | Consumers | Messages/sec | GB/sec |
|---|---|---|---|
| 1 | 0 | 77.6 M | 5.9 |
| 1 | 1 | 135.6 M | 10.3 |
| 1 | 4 | 114.3 M | 8.7 |
| 2 | 2 | 26.4 M | 2.0 |
| 4 | 4 | 16.9 M | 1.3 |
| 4 | 1 | 18.1 M | 1.4 |
Single-producer fan-out runs at 75–135M msgs/sec; multiple producers contend on the one interlocked claim counter (the natural cost of lock-free multi-producer ordering) and still sustain 16–26M msgs/sec aggregate. Unlike the syscall-based transports above, shared-memory throughput holds up under machine load.
dotnet run -c Release --project bench/Sluice.Benchmarks -- lifecycle
| Subsystem | Stage | Median |
|---|---|---|
| RPC | init → bootstrapped (daemon) | 0.9 ms |
| RPC | bootstrapped → discovered | 6.6 ms |
| RPC | discovered → active (1st call) | 0.13 ms |
| Fusion | host bootstrapped | 4.0 ms |
| Fusion | mirror attach | 0.4 ms |
| Fusion | active (1st fetch) | 0.1 ms |
| Fusion | invalidation propagation | 9.5 ms |
| Gossip | init → all discovered (N=3..8) | ~31 ms |
| Gossip | seed → fully converged (N=3..8) | ~15 ms |
Gossip timings are flat from N=3 to N=8 — shared-memory multicast reaches every node at once (O(1)), not the O(log N) rounds a network gossip protocol needs.
MIT — see LICENSE.