fix: workflow inference chain — surface llama-server errors, user-turn-last, bundle-relative prompts

- LlamaCppInferenceProvider: check HTTP status and surface the real error body instead of masking it as a ChatCompletionResponse deserialization failure
- PromptRenderer: render the live conversation layer (L1) last so the user turn is the final message (fixes Qwen 'No user query found' 500 when L3 recall is present)
- TomlWorkflowLoader: resolve stage prompts relative to the workflow file's own dir (bundle), CWD-independent; config-dir fallback for globals
- SessionOrchestrator: hard-fail a stage whose declared prompt can't resolve (was a silent user-less request)
- move healthcheck prompts next to the example workflow (examples/workflows/prompts/)
This commit is contained in:
2026-06-01 23:23:29 +04:00
parent 55a18b4b3a
commit 94f7ad0ee9
6 changed files with 50 additions and 23 deletions
-4
View File
@@ -1,4 +0,0 @@
# healthcheck_execute.md
Run `scripts/healthcheck.sh` and report the results.
If the script fails, report the error and exit code.
-8
View File
@@ -1,8 +0,0 @@
Write a bash healthcheck script that reports:
- RAM usage %
- CPU usage %
- GPU usage % and VRAM usage % if an NVIDIA or AMD GPU is detected
Use standard tools only (free, top/vmstat, nvidia-smi, rocm-smi).
If no GPU is detected, skip that section gracefully.
Output path: scripts/healthcheck.sh