feat(inference): operator-tunable sampling knobs (top_k/min_p/repeat_penalty) for stage requests
GenerationConfig only carried temperature/top_p/max_tokens/stop/seed. Added nullable topK/minP/repeatPenalty, serialized to the llama.cpp and OpenAI-compat request bodies via @EncodeDefault(NEVER) so an unset knob is omitted (the model keeps its own default) and behavior is unchanged unless the operator opts in. Surfaced as a new [sampling] config section feeding the default stage GenerationConfig (the main agentic loop) through TomlWorkflowLoader + ExecutionPlanCompiler; the former hardcoded temperature=0.7/topP=1.0 stage defaults now come from config. Talkie chat/narration keep their own generation settings. Vikunja #46 (task 76) — sampling half.
This commit is contained in:
@@ -13,4 +13,9 @@ data class GenerationConfig(
|
||||
val maxTokens: Int,
|
||||
val stopSequences: List<String> = emptyList(),
|
||||
val seed: Long? = null, // null = non-deterministic; set for replay
|
||||
// Sampling knobs. null = omit from the request so the provider/model keeps its own default,
|
||||
// preserving prior behavior. Serialized only when set (top_k / min_p / repeat_penalty).
|
||||
val topK: Int? = null,
|
||||
val minP: Double? = null,
|
||||
val repeatPenalty: Double? = null,
|
||||
)
|
||||
|
||||
Reference in New Issue
Block a user