Fault-Injection Harness

Production incidents in a mesh are ORDERINGS: a write lands while its owner's pod lingers in a stop, a move reads storage before a flush, a first query frame arrives after a grace, a route answers NotFound for a node that was just created, a webhook meets a 500 mid-roll. Reproducing them by timing — a sleep, a load generator, a lucky run — produces tests that pass alone and fail in the suite, or the reverse. The harness makes each ordering a switch: the fault is in force from the moment the switch is created until it is released, and nothing in between depends on a clock.

Two projects under test/:

Project What it holds
MeshWeaver.Testing.FaultInjection the single-process injectors: FaultSwitch, FaultInjectingStorageAdapter, HeldFirstFrameQueryProvider, CrossProcessChangeRelay, FaultInjectingHttpHandler, FaultInjectingInbox<TMessage,TResult>, and AddFaultInjectingStorage() / AddHeldFirstFrame() for a monolith mesh. No Orleans dependency.
MeshWeaver.Hosting.Orleans.TestBase → FaultInjectionCluster the multi-silo harness: N in-process silos over ONE store of record, each silo's store under its own injector, the LISTEN model between them, and Kill / Drain / Linger.
MeshWeaver.FaultInjection.Test one isolated regression per incident, plus one worked example per injector (InjectorExamplesTest).

A satellite reaches them the way it reaches MeshWeaver.Fixture: by ProjectReference through its platform checkout ($(MeshWeaverRoot)/test/…). Nothing here is packed.

The rules the harness keeps

The injectors, one example each

Multi-silo: start two processes and transfer from one to the other

FaultInjectionCluster (derive a Cluster class per test class — a case that removes a silo leaves the cluster changed, so it is never pooled; override SiloCount for three):

Call Production shape
Kill(i) SIGKILL / SIGSEGV — no graceful stop, no deactivation
Drain(i) a rolling restart — SIGTERM, then a graceful stop
await using (Linger(i)) the roll at its worst moment — told to stop, still in the cluster; every grain on it refuses as ShuttingDown and hands its address off, until the handle is disposed
Relay.Hold() / Relay.Drop() a slow / reconnected-without-replay PostgreSQL LISTEN channel between silos
HandOffTarget(path, leaving, active) which silo a lingering grain hands path to — so a case can CHOOSE a path whose hand-off lands where it needs it

Silo indices are stable across kills. Death detection and the held-stream heartbeat run at test cadence (FaultInjectionCluster.Heartbeat, 1 s). The cross-process relay is ON by default, so a case measures the mesh production runs — before it, a multi-silo test had a shared store and no cross-process invalidation at all, and published it by hand.

public class AWriteDuringAPodRollIsReDrivenTest(AWriteDuringAPodRollIsReDrivenTest.Cluster mesh)
    : IClassFixture<AWriteDuringAPodRollIsReDrivenTest.Cluster>
{
    [Fact(Timeout = 180_000)]
    public async Task AWriteToAnOwnerWhoseSiloIsStopping_IsReDriven_AndTheHeldReadDeliversIt()
    {
        var ct = TestContext.Current.CancellationToken;
        using var held = await HeldRead.Arrange(mesh, owner: 1, holder: 0, "roll-write", ct);
        await using (mesh.Linger(1))
        {
            await held.Write("v2").Should().Within(TestTimeouts.Convergence).Emit("…re-driven…", ct);
            await held.Delivers("v2").Should().Within(TestTimeouts.Convergence).Emit("…delivered…", ct);
        }
    }
    public class Cluster : FaultInjectionCluster;
}

Storage flush hold

FaultInjectingStorageAdapter.HoldWrites(path): every write of the path reaches the INNERMOST adapter — under every guard the platform stacks on top — and waits. The owner has committed in memory; storage still holds the previous state. HeldWrites(path) replays each held write, and EnumeratedWhileHeld(path) reports a subtree enumeration (only a copy or a move does that) made while the flush is held. This is the general form of MeshWeaver.Plugins' SubmitFlushHoldingStorageAdapter (#5670, Plugins#2536).

protected override MeshBuilder ConfigureMesh(MeshBuilder builder)
    => base.ConfigureMesh(builder.AddFaultInjectingStorage());   // BEFORE the base adds persistence

using (var flush = Storage.HoldWrites(path))
{
    await Mesh.GetMeshNodeStream(path).Update(n => n with { Name = "committed" }).Should().Emit(…);
    await Storage.HeldWrites(path).Where(n => n.Name == "committed").Should().Emit(…);
    // Storage.Inner still reads the previous state here.
}

🚨 Register it BEFORE the base configuration's AddInMemoryPersistence — persistence TryAdds its adapter, so registering first is what puts the injector under the guards. On a cluster the fixture does this for every silo.

Route / NotFound window

HidePath(path): reads, existence probes, route resolution and the parent's child listing answer as if the path did not exist yet; writes still land. The shape of a resolver that has not caught up with a create another process acknowledged. Resolution stops at the parent, so the router answers No node found at '…' exactly as production logs it.

Change-feed hold

HoldChangeFeed(): the silo's change notifications — its own commits and relayed ones — queue in order and are delivered on release. Relay.Hold() does the same for the cross-process leg only.

First-frame delay

AddHeldFirstFrame(partition) adds a HeldFirstFrameQueryProvider to the query fan-in. Hold() keeps it silent for queries into that partition, and the fan-in waits for every matching provider's first frame (bounded by QueryInitialBudget), so the query's Initial is held until release.

using (var hold = FirstFrame.Hold())
{
    var frames = MeshQuery.Query<MeshNode>(query).Replay();
    using var subscribed = frames.Connect();
    await hold.Arrivals.Should().Emit(…);
    await frames.Should().NotEmit(TimeSpan.FromMilliseconds(500), "no first frame while held");
    hold.Release();
    await frames.Should().Emit("released, the fan-in emits its first frame");
}

Webhook loss / 500 during a roll

FaultInjectingHttpHandler sits in front of an HTTP inbox: Refuse(status) answers without forwarding (the pod-roll 500 GitHub never retries), FailAfterDelivery(status) forwards and then answers the error anyway. FaultInjectingInbox<TMessage,TResult> injects the same two faults into an inbox that is a function. Faulted replays every delivery a fault met, so a case can assert that a sweep re-fed exactly what the sender lost.

The cases

# Incident Test (MeshWeaver.FaultInjection.Test) Injector Negative control
1 Write during a pod roll (#5873) AWriteDuringAPodRollIsReDrivenTest Linger #5873's two ShuttingDown arms reverted → red: MeshNode Unknown … is shutting down … Rejecting now
4 Steward's first write after its create (Plugins#2530) ARoutedWriteRightAfterItsCreateTest (5 cases) HidePath, Relay.Hold, a real pre-create probe MeshNodeStreamCache.ResetFailureState made a no-op → the probe case red: No node found at '…/Item'
6 Fleet watch freeze (#5011) AHeldReadSurvivesItsSourceSilo{BeingKilled,Draining}Test, AHeldReadOnAThirdSilo…, AHeldReadFollowsItsOwnersHandOffTest, AHeldReadFreezesWhenItsOwnersHandOffIsNotNotifiedTest Kill, Drain, Linger, HoldChangeFeed, HandOffTarget held-stream heartbeat pushed beyond the budget → all four kill/drain cases red (the owner never re-activates)

Cases 2, 3, 5, 7 and 8 exercise code that lives in MeshWeaver.Plugins (the Feedback submit, the Hosting record resolution, the plan-approve page, the running-action resumer, the PR steward's sweep). They move onto this harness there, by ProjectReference through the platform checkout, once the platform pin carries it.

What the cases found

#5011 — a held read's liveness across a roll is only as good as its process's change feed. Held reads survive a killed, drained or lingering owner in every topology modelled, with notifications flowing. But with the holder's change feed withheld and the owner handed off to a THIRD silo, the held read freezes: no value, no error, no completion — the #5011 symptom, "reads as holding while it holds nothing". It delivers the moment the one withheld notification is released. The reason is structural: the sync stream's heartbeat is fire-and-forget (JsonSynchronizationStream: the change-feed resubscribe is "the sole recycled-grain detector"), and a lingering owner's goodbye rides the refused router. AHeldReadFreezesWhenItsOwnersHandOffIsNotNotifiedTest PINS that gap, so a fix (a heartbeat that detects an owner which no longer knows the subscriber) flips it on purpose. The case was first written with the hand-off left to chance and passed alone while freezing in the suite: the hand-off target is a per-process ordinal string hash, which is why HandOffTarget exists. Not modelled: an in-process kill still writes Dead to the membership table, so a SIGSEGV whose row stays Active until the survivors vote it out is not exercised.

Plugins#2530 — a pre-create probe's window, closed only by a notification. A point read of a not-yet-existing path opens the storm breaker's window on the reading process, and that window fast-fails WRITES too. What closes it early is the create's change event reaching that process. When the probe, the create and the write all run in one process — the steward's shape — the harness shows the first write lands (APreCreateProbe_DoesNotPoisonTheIssuersFirstWrite_WhenNotificationsAreLate). When ANOTHER replica creates and its notification is late, the prober's first write fails with the steward's exact line, No node found at '…/Item', until the notification arrives (ACrossReplicaCreate_WithALateNotification_LeavesTheProbersWriteShut_UntilTheNotificationArrives, pinned). Whether the steward met that second shape is not established. Separately, a write routed into a transient NotFound window fails loudly naming the path, and the write-side breaker then holds the path shut for one base cooldown (2 s, measured) because no change event follows to clear it; the next natural write after that lands.

See also