expect_not_called could pass on a step that made the forbidden call. callPaths
concatenates per server and callCount was a total, so slicing the concatenated
list by the total examined the wrong window. With praxis on three requests and
nexus on one, a fourth praxis call landed at index three and paths[4:] never
saw it, while the stale nexus call was reported as new. The mark is now
per server and the paths are taken per server from it.
Neither scenario ever produced an act, so all three fakes saw zero requests and
the fault lever changed no outcome. The two headline capabilities of the
harness had no coverage. act_degraded scripts an act against an enabled
allowlist row and runs it healthy, at 503 and healthy again, asserting the
reply, the call, the absence of a call on a tick, and that nothing was pushed
at him either way. That needed two seams the world did not have: allowlist rows
from the scenario, and a matcher on the real store rather than a nil API, which
would have panicked the moment any scenario produced an act.
expect_no_events compared bus.Len(), which stops growing at the ring capacity,
so a scenario long enough to fill the ring made every later expect_no_events
pass unconditionally. It counts publishes through a subscriber now.
A scenario could not express a fact below full confidence, because write
hardcoded 1.0, and morning_missed annotated its ambient step as if it could.
factPriority branches on exactly that, so no replay could reach the low branch.
signalStep takes a confidence, the ambient step sets the 0.6 the ambient path
writes, and the event line carries the priority so a scenario can assert it.
TestSimulatorIsDeterministic compared the transcript against time.Now, which
fails for the half hour a day the scenario itself covers. It checks that every
stamped line falls inside the scenario span instead. TestSimulatorRefusesBackwardsSteps
tested the forwards case, because reaching the backwards branch ended the test.
A fatalf seam makes the refusal observable.
Smaller notes: the step doc comment now states which assertions are run scoped
and which are step scoped, audioText parses the golden manifest once per world
rather than once per step, and the feminine checks list the masculine form with
its following character, since the earlier check on a comma alone passed on
"записал что ты выпил воды".
Found in review of #79.
A scenario is a JSON file under cmd/mavend/testdata/scenarios: a start
instant, a script of canned model answers, and a list of steps at "HH:MM".
Each step does one thing — say, audio, signal, arrive, tick, fault — and
then asserts on what she said, what was sent, which ecosystem services were
called, and what landed in the intake journal.
Between those boundaries the real components run: the real router cascade
(stage0, the LLM router over a scripted completer, the classifier
underneath it), the real store, the real reactive handler, the real tick
loop, and the same intake-decorated ipc.CoreAPI the daemon wires. What is
faked is only what a test cannot have: the model, the microphone, the
speaker, the delivery sink, and the ecosystem HTTP services.
Time is a single fakeClock threaded into every reader — the handler, the
intake publish stamp and tick(ctx, now) — so there is no time.Now() on the
replay path and a scenario is reproducible. TestSimulatorIsDeterministic
enforces that by replaying twice and diffing the transcripts byte for byte;
advanceTo refuses a step that goes backwards.
Two scenarios ship. morning_missed replays #284's own description: he
appears at the desk, a feed item, a mail candidate and a relayed
notification arrive through the morning, two ticks pass, and the assertions
are as much about nothing being sent at him unprompted as about what she
said. evening_degraded picks up the tier-2 pipeline case #288 deferred
here — a golden WAV through the STT seam to a written fact — and then puts
the ecosystem into 503 and checks that the proactive loop stays quiet and
that intake keeps working without it.
This is test-only code. Nothing in the production binaries changed, so the
daemon behaves identically when no scenario is running.
`make simulate` runs them verbose so the transcript is readable; `make
test` runs them with everything else.
Vikunja #284
The review's verification plan needs an instrument to settle whether the
classifier or the LLM router handles RU queries better, but that comparison
only means something if the daemon's own safety invariants are pinned
independently first.
These scenarios are deliberately narrow. They consume already-normalized
router decisions and assert what the post-router daemon owns: that a decision
requiring confirmation cannot execute before it is confirmed, that an
unresolved entity is never guessed at, and that named capabilities stay
unexecuted. Model routing quality is a separate question, evaluated against a
held-out contract fixture — mixing the two would produce a suite that fails for
two unrelated reasons.
The fixture is versioned (schema_version) so scenarios can be added without
rewriting the loader.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X5JApcrCRVGmqrxnhynSik