# LLMLingua-2 token-pruning sidecar Prunes low-perplexity tokens from freeform prose before it hits the local LLM, so more usable context fits a bounded window. Implements pipeline stage 3 (`TOKEN_PRUNE`, level 3+) — see `docs/plans/correx-compression-pipeline.md` §4. Python-only because LLMLingua-2 is a torch/BERT classifier with no JVM equivalent. correx calls it over localhost HTTP via `HttpTokenPruner`, which **fails open**: if this sidecar is down, the kernel passes context through uncompressed. Nothing breaks; you just don't get token pruning. ## Run ```bash cd sidecars/llmlingua python -m venv .venv && source .venv/bin/activate pip install -r requirements.txt uvicorn server:app --host 127.0.0.1 --port 8199 ``` First `/prune` call downloads the model (~1-2 GB) and loads torch; `/health` responds immediately. ## Wire into correx Set compression level ≥ 3 for the workflow and point the kernel at the sidecar: ```toml [compression] level = 4 token_pruner_url = "http://127.0.0.1:8199" ``` ## API - `GET /health` → `{"status":"ok"}` - `POST /prune` `{"text": str, "protected": [str], "rate": 0.55}` → `{"compressed": str}` - `rate` = fraction of tokens to **keep** (0.55 ≈ 45% compression) - `protected` substrings (IDs, numbers, paths, code) are kept verbatim