Files
Maven/docs/evals/2026-08-03-llama-prompt-cache.md
T

3.6 KiB

Where the resident model's 7.9GB of RSS goes (2026-08-03, homesrv)

Measured for Vikunja #499. The deployed llama-server held 7.9GB RSS for a 1.1GB model file. Half a gigabyte of it was in swap, on a box that also runs whisper.cpp, piper and the embedder.

Method

maven-mavend-1 was stopped for the measurement, with the owner's approval. Its own binary then ran on the host with the exact deployed command line. That binary is /opt/maven/bin/llama-server, version 1 (4c65955), a Vulkan build.

llama-server -m /mnt/hdd1/llms/qwen3/Qwen3-1.7B-UD-Q4_K_XL.gguf \
  --host 127.0.0.1 --port 18099 -c 4096 -ngl 99 --no-webui

RSS was read from /proc/<pid>/status after load and after each of 8 distinct 1521-token prompts. smaps of the deployed process was read first, from inside the container, since the host user cannot read another user's maps.

The cause: the prompt cache, not the weights and not the offload

The startup log says it outright:

srv load_model: prompt cache is enabled, size limit: 8192 MiB
srv llama_server: n_parallel is set to auto, using n_parallel = 4 and kv_unified = true

The server saves the full KV state of every idle slot it evicts. It keeps up to 8GiB of those states in host RAM (llama.cpp PR 16391). One saved prompt of 1521 tokens costs 166.377 MiB. That is 112 kiB per token, exactly Qwen3-1.7B's KV footprint (28 layers x 2 x 1024 dims x 2 bytes).

RSS at rest, and per distinct prompt:

Prompts served RSS, default RSS, --cache-ram 512
0 (just loaded) 443 MB 411 MB
1 445 MB 411 MB
4 958 MB 929 MB
8 1641 MB 932 MB

Uncapped, RSS climbs about 170MB per distinct prompt and does not stop until the 8GiB limit. Capped at 512 MiB it plateaus at 932MB from the fourth prompt on, with the cache holding steady at 3 prompts, 499.132 MiB and evicting.

The 7.9GB on the running daemon was that climb, weeks of it. Its smaps showed one 6.03GB anonymous mapping at 5.32GB resident plus a 1.45GB mapping at 1.27GB resident, and only 30MB of file-backed RSS.

The task's leading guess was wrong

-ngl 99 on the Vega iGPU costs almost no process RSS. A freshly loaded server has 95MB of anonymous RSS in total. RADV allocates device memory through the kernel, outside the process, and the log sees 8202 MiB free on Vulkan0. The weights are mmapped and file-backed, so they are evictable and do not pin RSS. The logit buffer is not visible in the numbers above at all.

Decision

--cache-ram 512 is now the default, wired as phraser.cache_ram_mib and set in deploy/mavend.json. 512 MiB caps total RSS near 1GB, an eighth of what the box carried. It still holds three of the 1521-token probes above. Maven's real routing and phrasing prompts are much shorter, so it holds more of those than the table suggests. -c 4096 is untouched, as #499 required. A negative cache_ram_mib passes no flag, for a llama-server too old to know it.

Not changed: n_parallel = 4. With kv_unified = true the four slots share one 4096-token KV cache, so they do not multiply it.

The other half of #499 was that none of these lines were reachable. mavend scraped llama-server's stderr for the listen line and discarded it, and never piped stdout at all. Both streams now go to mavend's log with a llama: prefix. The last 12 startup lines go into the error when the server dies before it listens.

Deployed

The mavenai:latest image was rebuilt and maven-mavend-1 recreated the same day. The daemon's own log now carries the child's startup, it reads prompt cache is enabled, size limit: 512 MiB, and the resident server sat at 439MB RSS after load and 613MB after one served turn.