047e2a4070
- build() is now suspend: pipeline runs all stages in fixed order, each gated by level - TOKEN_PRUNE (level 3): TokenPruner interface + LLMLingua-2 sidecar (sidecars/llmlingua) + HttpTokenPruner adapter (fails open if sidecar down); prunes freeform, preserves protected spans, skips tier-0 turns when TIER_SPLIT on - TOME_MERGE (level 8): ToMeMerger collapses near-duplicate freeform turns (Jaccard) - Stage 5 selection: RelevanceScorer + EmbeddingRelevanceScorer (cosine over Embedder); query-conditioned reorder so least-relevant freeform drops first under budget - [orchestration] compression_level + token_pruner_url config, wired in Main - suspend ripple fixed across builder callers/stubs
38 lines
1.3 KiB
Markdown
38 lines
1.3 KiB
Markdown
# LLMLingua-2 token-pruning sidecar
|
|
|
|
Prunes low-perplexity tokens from freeform prose before it hits the local LLM, so more usable
|
|
context fits a bounded window. Implements pipeline stage 3 (`TOKEN_PRUNE`, level 3+) — see
|
|
`docs/plans/correx-compression-pipeline.md` §4.
|
|
|
|
Python-only because LLMLingua-2 is a torch/BERT classifier with no JVM equivalent. correx calls
|
|
it over localhost HTTP via `HttpTokenPruner`, which **fails open**: if this sidecar is down, the
|
|
kernel passes context through uncompressed. Nothing breaks; you just don't get token pruning.
|
|
|
|
## Run
|
|
|
|
```bash
|
|
cd sidecars/llmlingua
|
|
python -m venv .venv && source .venv/bin/activate
|
|
pip install -r requirements.txt
|
|
uvicorn server:app --host 127.0.0.1 --port 8199
|
|
```
|
|
|
|
First `/prune` call downloads the model (~1-2 GB) and loads torch; `/health` responds immediately.
|
|
|
|
## Wire into correx
|
|
|
|
Set compression level ≥ 3 for the workflow and point the kernel at the sidecar:
|
|
|
|
```toml
|
|
[compression]
|
|
level = 4
|
|
token_pruner_url = "http://127.0.0.1:8199"
|
|
```
|
|
|
|
## API
|
|
|
|
- `GET /health` → `{"status":"ok"}`
|
|
- `POST /prune` `{"text": str, "protected": [str], "rate": 0.55}` → `{"compressed": str}`
|
|
- `rate` = fraction of tokens to **keep** (0.55 ≈ 45% compression)
|
|
- `protected` substrings (IDs, numbers, paths, code) are kept verbatim
|