docs: record that contention is KFD presence, not a VRAM threshold (V-489)
This commit is contained in:
@@ -71,6 +71,22 @@ It arbitrates nothing between callers. It reports whether it can take work and
|
|||||||
manages one process to back that answer. Maven never asks it to start anything
|
manages one process to back that answer. Maven never asks it to start anything
|
||||||
and never learns that it did.
|
and never learns that it did.
|
||||||
|
|
||||||
|
Contention is decided by presence under `/sys/class/kfd/kfd/proc`, not by a VRAM
|
||||||
|
threshold. A ROCm process registers there when it initialises HIP, before it
|
||||||
|
allocates anything. So the supervisor sees a contender during that job's startup,
|
||||||
|
and yields before the job loses the memory it asked for. A
|
||||||
|
threshold reads the card too late. By the time free VRAM has dropped, the other
|
||||||
|
job has already lost the allocation race. Free VRAM is still read, but only as a
|
||||||
|
precondition for loading, never as the eviction signal. One blind spot is known.
|
||||||
|
A job can take the card without registering on the KFD, as a Vulkan or a
|
||||||
|
video-decode job would. `describe()` logs every contender's comm, and that log is
|
||||||
|
how we find out whether the blind spot is real.
|
||||||
|
|
||||||
|
`mavgpud` runs from a systemd unit on the workstation with
|
||||||
|
`deploy/mavgpud.json` as its config, and `llama_args` is passed to llama-server
|
||||||
|
untouched. The model, the context size, the layer count and the MTP flags are the
|
||||||
|
owner's business and not this daemon's schema.
|
||||||
|
|
||||||
## What stays on homesrv, permanently
|
## What stays on homesrv, permanently
|
||||||
|
|
||||||
The **embedder** (multilingual-e5-small, ONNX, CPU). It backs the classifier, which
|
The **embedder** (multilingual-e5-small, ONNX, CPU). It backs the classifier, which
|
||||||
|
|||||||
Reference in New Issue
Block a user