Fault-Injection Harness
Production incidents in a mesh are ORDERINGS: a write lands while its owner's pod lingers in a stop, a move reads storage before a flush, a first query frame arrives after a grace, a route answers NotFound for a node that was just created, a webhook meets a 500 mid-roll. Reproducing them by timing — a sleep, a load generator, a lucky run — produces tests that pass alone and fail in the suite, or the reverse. The harness makes each ordering a switch: the fault is in force from the moment the switch is created until it is released, and nothing in between depends on a clock.
Two projects under test/:
| Project | What it holds |
|---|---|
MeshWeaver.Testing.FaultInjection |
the single-process injectors: FaultSwitch, FaultInjectingStorageAdapter, HeldFirstFrameQueryProvider, CrossProcessChangeRelay, FaultInjectingHttpHandler, FaultInjectingInbox<TMessage,TResult>, and AddFaultInjectingStorage() / AddHeldFirstFrame() for a monolith mesh. No Orleans dependency. |
MeshWeaver.Hosting.Orleans.TestBase → FaultInjectionCluster |
the multi-silo harness: N in-process silos over ONE store of record, each silo's store under its own injector, the LISTEN model between them, and Kill / Drain / Linger. |
MeshWeaver.FaultInjection.Test |
one isolated regression per incident, plus one worked example per injector (InjectorExamplesTest). |
A satellite reaches them the way it reaches MeshWeaver.Fixture: by ProjectReference through its
platform checkout ($(MeshWeaverRoot)/test/…). Nothing here is packed.
The rules the harness keeps
- A fault is a
FaultSwitch. Created closed, released exactly once,Dispose= release, sousing var hold = …(ortry/finally) releases it on every exit path — a failing assertion never strands the work it holds, and the fixture's teardown never waits on it. - Nothing parks a thread. A held operation is an observable that has not emitted yet
(
FaultSwitch.Gatecomposes it behind anAsyncSubjectthe release completes). That is what makes a switch legal inside a hub's action block. NoSemaphoreSlim, noTask.Delay, no sleep. - Every fault reports what it met.
FaultSwitch.Arrivalsreplays each operation the fault answered, noted when it ARRIVES. A case waits on that positive signal ("the write reached the hold") before it asserts, and asserts it again at the end — so a case whose fault never fired fails instead of passing having injected nothing. - The fixture proves its own wiring.
FaultInjectionClusterasserts, per silo, that the writable store is served through that silo's injector before any case runs. - Instance state only. Every injector is a singleton of one mesh or one silo; the cluster owns the relay. No statics.
The injectors, one example each
Multi-silo: start two processes and transfer from one to the other
FaultInjectionCluster (derive a Cluster class per test class — a case that removes a silo leaves
the cluster changed, so it is never pooled; override SiloCount for three):
| Call | Production shape |
|---|---|
Kill(i) |
SIGKILL / SIGSEGV — no graceful stop, no deactivation |
Drain(i) |
a rolling restart — SIGTERM, then a graceful stop |
await using (Linger(i)) |
the roll at its worst moment — told to stop, still in the cluster; every grain on it refuses as ShuttingDown and hands its address off, until the handle is disposed |
Relay.Hold() / Relay.Drop() |
a slow / reconnected-without-replay PostgreSQL LISTEN channel between silos |
HandOffTarget(path, leaving, active) |
which silo a lingering grain hands path to — so a case can CHOOSE a path whose hand-off lands where it needs it |
Silo indices are stable across kills. Death detection and the held-stream heartbeat run at test
cadence (FaultInjectionCluster.Heartbeat, 1 s). The cross-process relay is ON by default, so a
case measures the mesh production runs — before it, a multi-silo test had a shared store and no
cross-process invalidation at all, and published it by hand.
public class AWriteDuringAPodRollIsReDrivenTest(AWriteDuringAPodRollIsReDrivenTest.Cluster mesh)
: IClassFixture<AWriteDuringAPodRollIsReDrivenTest.Cluster>
{
[Fact(Timeout = 180_000)]
public async Task AWriteToAnOwnerWhoseSiloIsStopping_IsReDriven_AndTheHeldReadDeliversIt()
{
var ct = TestContext.Current.CancellationToken;
using var held = await HeldRead.Arrange(mesh, owner: 1, holder: 0, "roll-write", ct);
await using (mesh.Linger(1))
{
await held.Write("v2").Should().Within(TestTimeouts.Convergence).Emit("…re-driven…", ct);
await held.Delivers("v2").Should().Within(TestTimeouts.Convergence).Emit("…delivered…", ct);
}
}
public class Cluster : FaultInjectionCluster;
}
Storage flush hold
FaultInjectingStorageAdapter.HoldWrites(path): every write of the path reaches the INNERMOST
adapter — under every guard the platform stacks on top — and waits. The owner has committed in
memory; storage still holds the previous state. HeldWrites(path) replays each held write, and
EnumeratedWhileHeld(path) reports a subtree enumeration (only a copy or a move does that) made
while the flush is held. This is the general form of MeshWeaver.Plugins' SubmitFlushHoldingStorageAdapter
(#5670, Plugins#2536).
protected override MeshBuilder ConfigureMesh(MeshBuilder builder)
=> base.ConfigureMesh(builder.AddFaultInjectingStorage()); // BEFORE the base adds persistence
using (var flush = Storage.HoldWrites(path))
{
await Mesh.GetMeshNodeStream(path).Update(n => n with { Name = "committed" }).Should().Emit(…);
await Storage.HeldWrites(path).Where(n => n.Name == "committed").Should().Emit(…);
// Storage.Inner still reads the previous state here.
}
🚨 Register it BEFORE the base configuration's AddInMemoryPersistence — persistence TryAdds its
adapter, so registering first is what puts the injector under the guards. On a cluster the fixture
does this for every silo.
Route / NotFound window
HidePath(path): reads, existence probes, route resolution and the parent's child listing answer as
if the path did not exist yet; writes still land. The shape of a resolver that has not caught up
with a create another process acknowledged. Resolution stops at the parent, so the router answers
No node found at '…' exactly as production logs it.
Change-feed hold
HoldChangeFeed(): the silo's change notifications — its own commits and relayed ones — queue in
order and are delivered on release. Relay.Hold() does the same for the cross-process leg only.
First-frame delay
AddHeldFirstFrame(partition) adds a HeldFirstFrameQueryProvider to the query fan-in. Hold()
keeps it silent for queries into that partition, and the fan-in waits for every matching provider's
first frame (bounded by QueryInitialBudget), so the query's Initial is held until release.
using (var hold = FirstFrame.Hold())
{
var frames = MeshQuery.Query<MeshNode>(query).Replay();
using var subscribed = frames.Connect();
await hold.Arrivals.Should().Emit(…);
await frames.Should().NotEmit(TimeSpan.FromMilliseconds(500), "no first frame while held");
hold.Release();
await frames.Should().Emit("released, the fan-in emits its first frame");
}
Webhook loss / 500 during a roll
FaultInjectingHttpHandler sits in front of an HTTP inbox: Refuse(status) answers without
forwarding (the pod-roll 500 GitHub never retries), FailAfterDelivery(status) forwards and then
answers the error anyway. FaultInjectingInbox<TMessage,TResult> injects the same two faults into an
inbox that is a function. Faulted replays every delivery a fault met, so a case can assert that a
sweep re-fed exactly what the sender lost.
The cases
| # | Incident | Test (MeshWeaver.FaultInjection.Test) |
Injector | Negative control |
|---|---|---|---|---|
| 1 | Write during a pod roll (#5873) | AWriteDuringAPodRollIsReDrivenTest |
Linger |
#5873's two ShuttingDown arms reverted → red: MeshNode Unknown … is shutting down … Rejecting now |
| 4 | Steward's first write after its create (Plugins#2530) | ARoutedWriteRightAfterItsCreateTest (5 cases) |
HidePath, Relay.Hold, a real pre-create probe |
MeshNodeStreamCache.ResetFailureState made a no-op → the probe case red: No node found at '…/Item' |
| 6 | Fleet watch freeze (#5011) | AHeldReadSurvivesItsSourceSilo{BeingKilled,Draining}Test, AHeldReadOnAThirdSilo…, AHeldReadFollowsItsOwnersHandOffTest, AHeldReadFreezesWhenItsOwnersHandOffIsNotNotifiedTest |
Kill, Drain, Linger, HoldChangeFeed, HandOffTarget |
held-stream heartbeat pushed beyond the budget → all four kill/drain cases red (the owner never re-activates) |
Cases 2, 3, 5, 7 and 8 exercise code that lives in MeshWeaver.Plugins (the Feedback submit, the
Hosting record resolution, the plan-approve page, the running-action resumer, the PR steward's
sweep). They move onto this harness there, by ProjectReference through the platform checkout,
once the platform pin carries it.
What the cases found
#5011 — a held read's liveness across a roll is only as good as its process's change feed. Held
reads survive a killed, drained or lingering owner in every topology modelled, with notifications
flowing. But with the holder's change feed withheld and the owner handed off to a THIRD silo, the
held read freezes: no value, no error, no completion — the #5011 symptom, "reads as holding while
it holds nothing". It delivers the moment the one withheld notification is released. The reason is
structural: the sync stream's heartbeat is fire-and-forget (JsonSynchronizationStream: the
change-feed resubscribe is "the sole recycled-grain detector"), and a lingering owner's goodbye
rides the refused router. AHeldReadFreezesWhenItsOwnersHandOffIsNotNotifiedTest PINS that gap,
so a fix (a heartbeat that detects an owner which no longer knows the subscriber) flips it on
purpose. The case was first written with the hand-off left to chance and passed alone while
freezing in the suite: the hand-off target is a per-process ordinal string hash, which is why
HandOffTarget exists. Not modelled: an in-process kill still writes Dead to the membership
table, so a SIGSEGV whose row stays Active until the survivors vote it out is not exercised.
Plugins#2530 — a pre-create probe's window, closed only by a notification. A point read of a
not-yet-existing path opens the storm breaker's window on the reading process, and that window
fast-fails WRITES too. What closes it early is the create's change event reaching that process.
When the probe, the create and the write all run in one process — the steward's shape — the
harness shows the first write lands (APreCreateProbe_DoesNotPoisonTheIssuersFirstWrite_WhenNotificationsAreLate).
When ANOTHER replica creates and its notification is late, the prober's first write fails with the
steward's exact line, No node found at '…/Item', until the notification arrives
(ACrossReplicaCreate_WithALateNotification_LeavesTheProbersWriteShut_UntilTheNotificationArrives,
pinned). Whether the steward met that second shape is not established. Separately, a write
routed into a transient NotFound window fails loudly naming the path, and the write-side breaker
then holds the path shut for one base cooldown (2 s, measured) because no change event follows to
clear it; the next natural write after that lands.
See also
- Writing Tests · Negative Controls · Reactive Test Assertions
- Riding Out a ShuttingDown Address — the refusal
Lingerproduces, and the riders that must survive it - Debugging Message Flow —
MESHWEAVER_MSG_TRACE=1shows a routed delivery's resolution, which is how case 4's NotFound window was read