fix(context,tools): context hygiene + tool ergonomics from freestyle QA
Found while live-QAing freestyle_planning on a 12B local model: - list_dir tool: recursive, .gitignore-aware listing so weak models stop flooding context with `ls -R` over node_modules/build/dist. Wired into the fileRead toggle + advertised to the planner (architect_freestyle). - ContextClassifier: assistantToolCall turns are STRUCTURED, so the token pruner never shreds the model's own tool-call history — that was causing amnesia loops (re-issuing calls it had already made). - Retire instruction-doc LLMLingua pruning (DOC_SOURCE_TYPES emptied): it fused load-bearing procedural text into unparseable soup. The static block stays small by dropping CLAUDE.md at the loader instead. - AgentInstructionsLoader: load only AGENTS.md, not CLAUDE.md — the latter targets the outer assistant and polluted the agent's stage context. - DefaultSessionReducer: WorkflowFailed flips session status to FAILED (was stuck ACTIVE forever, so clients/approval loops never saw a terminal). - ShellTool: run shell command lines (cd/&&/pipes) via `sh -c`; unrunnable program is recoverable instead of an uncaught IOException killing the stage; malformed argv (non-string/collapsed-array) rejected with guidance. - llmlingua sidecar: cap force_tokens to max_force_token (big docs blew the assert and 500'd, so doc pruning silently failed open). Tests added/updated across all of the above.
This commit is contained in:
@@ -55,11 +55,18 @@ def prune(req: PruneRequest):
|
||||
text = req.text.strip()
|
||||
if not text:
|
||||
return PruneResponse(compressed=req.text)
|
||||
compressor = _get_compressor()
|
||||
# LLMLingua-2 asserts len(force_tokens) <= max_force_token (default 100). Large static docs
|
||||
# (CLAUDE.md/AGENTS.md) yield hundreds of protected spans, which used to blow the assert and
|
||||
# 500 -> the doc came back uncompressed. Dedup and keep the longest spans (most load-bearing:
|
||||
# full paths, hashes, code fences beat bare numbers) up to the cap.
|
||||
cap = getattr(compressor, "max_force_token", 100)
|
||||
forced = sorted(set(req.protected), key=len, reverse=True)[:cap] or None
|
||||
# rate is fraction to keep; LLMLingua-2 force_tokens keeps the protected spans verbatim.
|
||||
result = _get_compressor().compress_prompt(
|
||||
result = compressor.compress_prompt(
|
||||
text,
|
||||
rate=max(0.1, min(1.0, req.rate)),
|
||||
force_tokens=req.protected or None,
|
||||
force_tokens=forced,
|
||||
drop_consecutive=True,
|
||||
)
|
||||
return PruneResponse(compressed=result["compressed_prompt"])
|
||||
|
||||
Reference in New Issue
Block a user