--- name: "Spec V0.1" description: "Early system specification (Harness v0.1)" depth: 1 links: ["../index.md", "../architecture/overview.md"] --- harness architecture specification version: 0.1-draft 1. purpose Harness is a config-driven orchestration runtime for local and remote LLM workflows. The system treats models as interchangeable execution engines while the Harness owns: - lifecycle - memory - orchestration - permissions - validation - event sourcing - workflow transitions - context synthesis Primary goals: - deterministic-enough execution on nondeterministic models - local-first operation - replayable execution - bounded autonomy - strong observability - minimal human steering - model/provider agnosticism Non-goals: - AGI simulation - unconstrained autonomous agents - hidden implicit memory - permanently accumulating context --- 2. architecture overview ┌──────────────────────────┐ │ User │ └────────────┬─────────────┘ │ ▼ ┌──────────────────────────┐ │ Router │ │ conversational interface │ └────────────┬─────────────┘ │ summaries ▼ ┌──────────────────────────┐ │ Harness │ │ orchestration kernel │ │ event bus │ │ approval engine │ │ lifecycle owner │ └────────────┬─────────────┘ │ ┌───────┴────────┐ ▼ ▼ ┌───────────┐ ┌────────────┐ │ Stage │ │ Context │ │ Runtime │ │ Processor │ └─────┬─────┘ └─────┬──────┘ │ │ ▼ ▼ ┌──────────────────────────┐ │ Agent Runtime │ └────────────┬─────────────┘ │ ▼ ┌──────────────────────────┐ │ Inference Layer │ │ local / remote models │ └────────────┬─────────────┘ │ ▼ ┌──────────────────────────┐ │ Tools │ └──────────────────────────┘ --- 3. core principles 3.1 event sourced All system state is reconstructable from events. No mutable hidden memory exists outside event projections. Benefits: - replayability - deterministic debugging - auditability - recovery - analytics - synthetic dataset generation --- 3.2 models are stateless Models do not own: - memory - workflow state - permissions - lifecycle Models receive synthesized context only. --- 3.3 validation-first execution Every agent output is validated at multiple levels: 1. routing validation 2. payload schema validation 3. semantic validation 4. approval validation Invalid artifacts cannot progress workflow state. --- 3.4 context is synthesized Raw accumulation is forbidden. All context is: - filtered - deduplicated - compressed - summarized - relevance-ranked --- 4. terminology term| meaning Harness| orchestration kernel Router| conversational interface layer Stage| workflow state Agent| execution worker Role| semantic responsibility Artifact| validated structured output Transition| movement between stages Event| immutable state mutation Projection| derived state view Context Pack| synthesized model input Approval Gate| human/system checkpoint --- 5. configuration system 5.1 config locations Global: ~/.config/harness/ Project-local: ./.harness/ Precedence: project > user > defaults --- 5.2 config files harness.yaml models.yaml tools.yaml stages.yaml policies.yaml compression.yaml --- 6. model registry 6.1 model definition models: qwen_coder_14b: provider: local path: /models/qwen-coder.gguf capabilities: - coding - tool_calling - reasoning context_size: 32768 kv_cache: quantization: q8_0 gpu: layers: 48 residency: dynamic swap_timeout_sec: 300 inference: temperature: 0.2 top_p: 0.9 repeat_penalty: 1.1 limits: max_parallel_sessions: 2 --- 7. providers 7.1 local provider Responsibilities: - llama.cpp process management - GPU residency - model loading/unloading - warm pools - swap scheduling - health monitoring --- 7.2 remote provider providers: openai_compatible: base_url: https://api.example.com/v1 api_key_env: HARNESS_API_KEY extra: timeout: 120 retries: 3 --- 8. stages Stages define execution states. Example: stages: implementation: role: coder requirements: - coding - tool_calling allowed_tools: - filesystem - git - shell approvals: tool_tier_3: required transitions: on_success: - validation on_failure: - retry --- 9. agents Agents are ephemeral execution workers. Agents: - consume Context Packs - emit Artifacts - do not persist memory --- 10. router Router responsibilities: - user interaction - conversational continuity - summarizing system state - steering interpretation Router is always available but not persistent in execution context. Execution agents unload router context during active work. After completion: - outputs are summarized - summaries are injected into router memory --- 11. artifacts All stage outputs MUST emit structured artifacts. Example: class ImplementationArtifact(BaseModel): summary: str modified_files: list[str] risks: list[str] commands_executed: list[str] next_recommendation: str --- 12. validation pipeline 12.1 routing validation Checks: - valid stage transitions - capability compatibility - policy compliance --- 12.2 payload validation Pydantic validation: - schema correctness - required fields - typing --- 12.3 semantic validation Checks: - contradictory outputs - hallucinated files - invalid commands - policy violations - unsafe operations --- 13. approval system 13.1 approval tiers tier| meaning T0| inference only T1| read-only T2| reversible mutation T3| external/network T4| destructive --- 13.2 approval modes mode| meaning prompt| require confirmation auto| auto approve deny| reject automatically yolo| bypass safeguards --- 13.3 approval actions User may: - approve - reject - auto-approve for session - steer execution Example: approved, but verify auth edge cases first --- 14. event system 14.1 event categories DomainEvents SystemEvents InferenceEvents ToolEvents ApprovalEvents CompressionEvents LifecycleEvents --- 14.2 event structure class Event(BaseModel): id: UUID session_id: UUID timestamp: datetime type: str payload: dict causation_id: UUID | None correlation_id: UUID | None --- 15. replay system Replay modes: - full replay - partial replay - replay from cursor - inference-skipping replay - deterministic simulation Replay reconstructs projections and workflow state. --- 16. context processor The ContextProcessor synthesizes minimal relevant context. Responsibilities: - deduplication - summarization - ranking - token budgeting - artifact extraction - tool compression --- 16.1 context layers L0 live execution L1 stage-local context L2 compressed session memory L3 project memory L4 archival history --- 16.2 compression strategies compression: tool_logs: mode: summarize artifacts: mode: latest_only events: mode: deduplicate conversations: mode: semantic_summary --- 17. tool system Tools are config-driven and capability-scoped. Default tools may be: - disabled - replaced - overridden Example: tools: shell: enabled: true tier: T2 git: enabled: true tier: T2 curl: enabled: true tier: T3 --- 18. transitions Transitions are rule-based. No hardcoded workflow graphs exist in code. Example: transitions: - when: artifact.status == "success" goto: validation - when: retries > 3 goto: failed --- 19. retry policies 19.1 strategies strategy| behavior retry| retry execution fail_safe| skip and continue fail_fast| terminate session --- 19.2 retry configuration retry: strategy: corrective max_attempts: 3 inject_failure_reason: true temperature_backoff: true --- 20. persistence Default persistence: SQLite Future: - PostgreSQL - event stores - distributed backends Persisted: - events - projections - artifacts - approvals - transitions - summaries --- 21. session lifecycle Harness owns session lifecycle. States: created active paused awaiting_approval failed completed cancelled --- 22. observability Required: - event tracing - transition tracing - inference timing - token accounting - tool execution logs - replay diagnostics Recommended: - DAG visualization - live stage graph - approval history --- 23. security model Principles: - least privilege - explicit approvals - isolated tools - auditability - bounded execution Recommendations: - sandbox shell tools - filesystem allowlists - network policy control - secret isolation - execution timeouts --- 24. future extensions Potential: - distributed agents - evaluator models - speculative execution - long-term semantic memory - automatic fine-tuning corpus extraction - capability benchmarking - adaptive routing --- 25. anti-goals Avoid: - hidden prompts - invisible memory mutation - unrestricted recursion - self-modifying workflows - implicit approvals - context accumulation without compression - permanent agent processes --- 26. philosophy summary Harness is not an “AI agent framework”. It is: - an orchestration kernel - an event-sourced execution runtime - a bounded autonomy system - a deterministic shell around probabilistic cognition