Files

672 lines
10 KiB
Markdown

---
name: "Spec V0.1"
description: "Early system specification (Harness v0.1)"
depth: 1
links: ["../index.md", "../architecture/overview.md"]
---
harness architecture specification
version: 0.1-draft
1. purpose
Harness is a config-driven orchestration runtime for local and remote LLM workflows.
The system treats models as interchangeable execution engines while the Harness owns:
- lifecycle
- memory
- orchestration
- permissions
- validation
- event sourcing
- workflow transitions
- context synthesis
Primary goals:
- deterministic-enough execution on nondeterministic models
- local-first operation
- replayable execution
- bounded autonomy
- strong observability
- minimal human steering
- model/provider agnosticism
Non-goals:
- AGI simulation
- unconstrained autonomous agents
- hidden implicit memory
- permanently accumulating context
---
2. architecture overview
┌──────────────────────────┐
│ User │
└────────────┬─────────────┘
┌──────────────────────────┐
│ Router │
│ conversational interface │
└────────────┬─────────────┘
│ summaries
┌──────────────────────────┐
│ Harness │
│ orchestration kernel │
│ event bus │
│ approval engine │
│ lifecycle owner │
└────────────┬─────────────┘
┌───────┴────────┐
▼ ▼
┌───────────┐ ┌────────────┐
│ Stage │ │ Context │
│ Runtime │ │ Processor │
└─────┬─────┘ └─────┬──────┘
│ │
▼ ▼
┌──────────────────────────┐
│ Agent Runtime │
└────────────┬─────────────┘
┌──────────────────────────┐
│ Inference Layer │
│ local / remote models │
└────────────┬─────────────┘
┌──────────────────────────┐
│ Tools │
└──────────────────────────┘
---
3. core principles
3.1 event sourced
All system state is reconstructable from events.
No mutable hidden memory exists outside event projections.
Benefits:
- replayability
- deterministic debugging
- auditability
- recovery
- analytics
- synthetic dataset generation
---
3.2 models are stateless
Models do not own:
- memory
- workflow state
- permissions
- lifecycle
Models receive synthesized context only.
---
3.3 validation-first execution
Every agent output is validated at multiple levels:
1. routing validation
2. payload schema validation
3. semantic validation
4. approval validation
Invalid artifacts cannot progress workflow state.
---
3.4 context is synthesized
Raw accumulation is forbidden.
All context is:
- filtered
- deduplicated
- compressed
- summarized
- relevance-ranked
---
4. terminology
term| meaning
Harness| orchestration kernel
Router| conversational interface layer
Stage| workflow state
Agent| execution worker
Role| semantic responsibility
Artifact| validated structured output
Transition| movement between stages
Event| immutable state mutation
Projection| derived state view
Context Pack| synthesized model input
Approval Gate| human/system checkpoint
---
5. configuration system
5.1 config locations
Global:
~/.config/harness/
Project-local:
./.harness/
Precedence:
project > user > defaults
---
5.2 config files
harness.yaml
models.yaml
tools.yaml
stages.yaml
policies.yaml
compression.yaml
---
6. model registry
6.1 model definition
models:
qwen_coder_14b:
provider: local
path: /models/qwen-coder.gguf
capabilities:
- coding
- tool_calling
- reasoning
context_size: 32768
kv_cache:
quantization: q8_0
gpu:
layers: 48
residency: dynamic
swap_timeout_sec: 300
inference:
temperature: 0.2
top_p: 0.9
repeat_penalty: 1.1
limits:
max_parallel_sessions: 2
---
7. providers
7.1 local provider
Responsibilities:
- llama.cpp process management
- GPU residency
- model loading/unloading
- warm pools
- swap scheduling
- health monitoring
---
7.2 remote provider
providers:
openai_compatible:
base_url: https://api.example.com/v1
api_key_env: HARNESS_API_KEY
extra:
timeout: 120
retries: 3
---
8. stages
Stages define execution states.
Example:
stages:
implementation:
role: coder
requirements:
- coding
- tool_calling
allowed_tools:
- filesystem
- git
- shell
approvals:
tool_tier_3: required
transitions:
on_success:
- validation
on_failure:
- retry
---
9. agents
Agents are ephemeral execution workers.
Agents:
- consume Context Packs
- emit Artifacts
- do not persist memory
---
10. router
Router responsibilities:
- user interaction
- conversational continuity
- summarizing system state
- steering interpretation
Router is always available but not persistent in execution context.
Execution agents unload router context during active work.
After completion:
- outputs are summarized
- summaries are injected into router memory
---
11. artifacts
All stage outputs MUST emit structured artifacts.
Example:
class ImplementationArtifact(BaseModel):
summary: str
modified_files: list[str]
risks: list[str]
commands_executed: list[str]
next_recommendation: str
---
12. validation pipeline
12.1 routing validation
Checks:
- valid stage transitions
- capability compatibility
- policy compliance
---
12.2 payload validation
Pydantic validation:
- schema correctness
- required fields
- typing
---
12.3 semantic validation
Checks:
- contradictory outputs
- hallucinated files
- invalid commands
- policy violations
- unsafe operations
---
13. approval system
13.1 approval tiers
tier| meaning
T0| inference only
T1| read-only
T2| reversible mutation
T3| external/network
T4| destructive
---
13.2 approval modes
mode| meaning
prompt| require confirmation
auto| auto approve
deny| reject automatically
yolo| bypass safeguards
---
13.3 approval actions
User may:
- approve
- reject
- auto-approve for session
- steer execution
Example:
approved, but verify auth edge cases first
---
14. event system
14.1 event categories
DomainEvents
SystemEvents
InferenceEvents
ToolEvents
ApprovalEvents
CompressionEvents
LifecycleEvents
---
14.2 event structure
class Event(BaseModel):
id: UUID
session_id: UUID
timestamp: datetime
type: str
payload: dict
causation_id: UUID | None
correlation_id: UUID | None
---
15. replay system
Replay modes:
- full replay
- partial replay
- replay from cursor
- inference-skipping replay
- deterministic simulation
Replay reconstructs projections and workflow state.
---
16. context processor
The ContextProcessor synthesizes minimal relevant context.
Responsibilities:
- deduplication
- summarization
- ranking
- token budgeting
- artifact extraction
- tool compression
---
16.1 context layers
L0 live execution
L1 stage-local context
L2 compressed session memory
L3 project memory
L4 archival history
---
16.2 compression strategies
compression:
tool_logs:
mode: summarize
artifacts:
mode: latest_only
events:
mode: deduplicate
conversations:
mode: semantic_summary
---
17. tool system
Tools are config-driven and capability-scoped.
Default tools may be:
- disabled
- replaced
- overridden
Example:
tools:
shell:
enabled: true
tier: T2
git:
enabled: true
tier: T2
curl:
enabled: true
tier: T3
---
18. transitions
Transitions are rule-based.
No hardcoded workflow graphs exist in code.
Example:
transitions:
- when:
artifact.status == "success"
goto: validation
- when:
retries > 3
goto: failed
---
19. retry policies
19.1 strategies
strategy| behavior
retry| retry execution
fail_safe| skip and continue
fail_fast| terminate session
---
19.2 retry configuration
retry:
strategy: corrective
max_attempts: 3
inject_failure_reason: true
temperature_backoff: true
---
20. persistence
Default persistence:
SQLite
Future:
- PostgreSQL
- event stores
- distributed backends
Persisted:
- events
- projections
- artifacts
- approvals
- transitions
- summaries
---
21. session lifecycle
Harness owns session lifecycle.
States:
created
active
paused
awaiting_approval
failed
completed
cancelled
---
22. observability
Required:
- event tracing
- transition tracing
- inference timing
- token accounting
- tool execution logs
- replay diagnostics
Recommended:
- DAG visualization
- live stage graph
- approval history
---
23. security model
Principles:
- least privilege
- explicit approvals
- isolated tools
- auditability
- bounded execution
Recommendations:
- sandbox shell tools
- filesystem allowlists
- network policy control
- secret isolation
- execution timeouts
---
24. future extensions
Potential:
- distributed agents
- evaluator models
- speculative execution
- long-term semantic memory
- automatic fine-tuning corpus extraction
- capability benchmarking
- adaptive routing
---
25. anti-goals
Avoid:
- hidden prompts
- invisible memory mutation
- unrestricted recursion
- self-modifying workflows
- implicit approvals
- context accumulation without compression
- permanent agent processes
---
26. philosophy summary
Harness is not an “AI agent framework”.
It is:
- an orchestration kernel
- an event-sourced execution runtime
- a bounded autonomy system
- a deterministic shell around probabilistic cognition