Score the chat, query and knowledge phrasing paths (#395)
The nudge fixture only covered nudges. The shared persona block now goes into five prompts, and the three conversational ones were unmeasured — those are the long free-form replies where a persona break is most likely. Adds talk_v1.json (27 Russian cases, 9 per path) and ScoreTalk, reporting per-path as well as per-check so a chat regression can be told apart from a knowledge one. Reuses the persona checks; the nudge-only ones (length, mood, no questions) are left out, since a chat reply is allowed 1-3 sentences and a follow-up question. The LLM run is opt-in on MAVEN_LLM_URL as before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CGeSZxh1DCtRxmFVSYVGvJ
This commit is contained in:
@@ -103,14 +103,17 @@ eval-router:
|
||||
eval-recall:
|
||||
MAVEN_ONNX_LIB="$(MAVEN_ONNX_LIB)" $(GO) test -v -count=1 ./internal/memory/recalleval/
|
||||
|
||||
# eval-phrasing -- score nudge phrasing (internal/phraser/eval). Verbose so the
|
||||
# eval-phrasing -- score nudge phrasing AND the conversational paths (chat,
|
||||
# query, general knowledge) in internal/phraser/eval. Verbose so the
|
||||
# report and every generated message land in the terminal. With no environment
|
||||
# it scores the deterministic Stub only, which is what CI runs. Set
|
||||
# MAVEN_LLM_URL to add the resident model:
|
||||
# MAVEN_LLM_URL=http://127.0.0.1:18099 make eval-phrasing
|
||||
# The model run is slow (minutes) -- the timeout is raised to match.
|
||||
# The model run is slow (minutes) -- the timeout is raised to match. It covers
|
||||
# two fixtures now (15 nudges + 27 conversational cases, and the chat replies are
|
||||
# the long ones), hence 90m rather than 40m.
|
||||
eval-phrasing:
|
||||
$(GO) test -v -count=1 -timeout 40m ./internal/phraser/eval/
|
||||
$(GO) test -v -count=1 -timeout 90m ./internal/phraser/eval/
|
||||
|
||||
# eval-models — score ONE llama-server against the same fixture, for the
|
||||
# resident-model bake-off (#278, #250). Start a server with the gguf you want,
|
||||
|
||||
Reference in New Issue
Block a user