Files
correx/docs/design/des-doc-v0.1.md

662 lines
13 KiB
Markdown

---
name: "Des Doc V0.1"
description: "Design document with full layer diagrams"
depth: 1
links: ["../index.md", "./spec-v0.1.md"]
---
harness design document
version: 0.1
status: architectural draft
1. overview
Harness is a local-first orchestration runtime for structured LLM execution.
The system is designed around a strict separation of concerns:
layer| responsibility
Router| conversational UX
Harness| orchestration + policy
Agent Runtime| task execution
Model Manager| inference lifecycle
Context Processor| memory synthesis
Event Store| persistence + replay
Tool Runtime| external actions
The architecture assumes:
- LLMs are probabilistic
- context is expensive
- memory must be externalized
- workflows require validation
- inference is transient
- orchestration owns reliability
The system intentionally resembles operating system architecture more than chatbot architecture.
---
2. architectural goals
primary goals
- replayable execution
- bounded autonomy
- local-first operation
- model/provider abstraction
- deterministic-enough workflows
- strict validation
- structured artifacts
- observability
- hardware-aware scheduling
---
secondary goals
- distributed execution
- remote provider failover
- semantic memory
- automatic evaluation
- synthetic training data generation
---
3. system architecture
┌────────────────┐
│ User │
└───────┬────────┘
┌────────────────────────┐
│ Router │
│ conversational facade │
└──────────┬─────────────┘
┌──────────────────────────────────────────────────────────┐
│ Harness │
│ │
│ ┌────────────┐ ┌────────────┐ ┌──────────────────┐ │
│ │ Session │ │ Transition │ │ Approval Engine │ │
│ │ Lifecycle │ │ Engine │ │ │ │
│ └─────┬──────┘ └─────┬──────┘ └────────┬─────────┘ │
│ │ │ │ │
│ ▼ ▼ ▼ │
│ ┌───────────────────────────────────────────────────┐ │
│ │ Event Bus │ │
│ └───────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────┘
┌──────────────────────────────────────────────────────────┐
│ Execution Layer │
│ │
│ ┌──────────────┐ ┌──────────────────────────────┐ │
│ │ StageRuntime │ │ ContextProcessor │ │
│ └──────┬───────┘ └──────────────┬───────────────┘ │
│ │ │ │
│ ▼ ▼ │
│ ┌──────────────────────────────────────────────┐ │
│ │ Agent Runtime │ │
│ └──────────────────────┬───────────────────────┘ │
└─────────────────────────┼────────────────────────────────┘
┌──────────────────────────────────────────────────────────┐
│ Inference Layer │
│ │
│ ┌───────────────┐ ┌────────────────────────────┐ │
│ │ Model Manager │────▶│ Inference Providers │ │
│ └───────────────┘ │ local / remote / hybrid │ │
│ └────────────────────────────┘ │
└──────────────────────────────────────────────────────────┘
┌────────────────┐
│ Tool Runtime │
└────────────────┘
---
4. execution lifecycle
step 1 — user interaction
User communicates only with Router.
Router responsibilities:
- intent interpretation
- steering ingestion
- summarization
- conversational continuity
Router never owns workflow state.
---
step 2 — session initialization
Harness:
- creates session
- initializes projections
- emits SessionCreated event
- resolves workflow entry stage
---
step 3 — context synthesis
ContextProcessor:
- retrieves relevant events
- compresses history
- filters artifacts
- injects policies
- builds token budget
Output:
ContextPack
---
step 4 — stage execution
StageRuntime:
- resolves agent role
- resolves required capabilities
- acquires suitable model
- executes inference
---
step 5 — artifact validation
Pipeline:
1. routing validation
2. payload validation
3. semantic validation
4. approval validation
Only validated artifacts emit success transitions.
---
step 6 — transition evaluation
TransitionEngine:
- evaluates rule graph
- emits transition event
- schedules next stage
---
step 7 — replay + persistence
Every mutation:
- emits immutable event
- updates projections
- persists session state
---
5. event sourcing architecture
rationale
Event sourcing is mandatory because:
- LLM execution is nondeterministic
- debugging requires replayability
- context reconstruction must be deterministic
- memory requires compression pipelines
---
event categories
DomainEvents
InferenceEvents
ToolEvents
ApprovalEvents
CompressionEvents
LifecycleEvents
SystemEvents
---
event flow
UserInputReceived
StageScheduled
ContextBuilt
InferenceStarted
ArtifactProduced
ArtifactValidated
ApprovalRequested
TransitionExecuted
---
6. context processing design
core assumption
Raw conversational accumulation destroys smaller models.
Context must be:
- hierarchical
- compressed
- relevance-ranked
- bounded
---
context layers
layer| meaning
L0| live execution
L1| active stage
L2| compressed session memory
L3| durable project memory
L4| archival history
---
context synthesis pipeline
events
deduplication
relevance ranking
semantic compression
artifact extraction
policy injection
token budgeting
ContextPack
---
compression rules
tool logs
Raw:
4000 lines shell output
Compressed:
pytest failed:
- auth_test.py
- timeout in token refresh
---
retention policy
data| strategy
artifacts| latest valid
tool logs| summarize
conversation| semantic
transitions| retain
approvals| retain
---
7. model management
design goals
- limited hardware optimization
- hot-swapping
- GPU residency control
- provider abstraction
- inference isolation
---
model manager responsibilities
- model lifecycle
- loading/unloading
- GPU scheduling
- swap timeout enforcement
- health monitoring
- concurrency limits
---
scheduling strategies
residency modes
mode| behavior
persistent| never unload
dynamic| unload after timeout
ephemeral| unload immediately
---
capability routing
Stages do not request model names.
Stages request capabilities:
requirements:
- coding
- reasoning
- tool_calling
Registry resolves:
- best local model
- available GPU budget
- fallback providers
---
8. agent runtime
design assumptions
Agents are:
- stateless
- ephemeral
- replaceable
Agents never:
- persist memory
- mutate workflow state directly
- own permissions
---
agent lifecycle
spawn
receive ContextPack
execute
emit Artifact
terminate
---
9. approval system
rationale
Autonomous systems require bounded risk.
Approval gates prevent:
- destructive execution
- runaway automation
- hidden escalation
---
approval tiers
tier| meaning
T0| inference only
T1| read-only
T2| reversible mutation
T3| external/network
T4| destructive
---
steering-aware approvals
Approvals may inject corrective context.
Example:
approved, but verify migrations against staging schema first
Approval events become part of future context synthesis.
---
10. transition engine
rule-based execution
Transitions are declarative.
No workflow logic is hardcoded.
---
example
transitions:
- when:
artifact.status == "success"
goto: validation
- when:
retries > 3
goto: failed
---
safeguards
Required:
- cycle detection
- deadlock detection
- transition tracing
- graph visualization
---
11. validation system
layered validation
layer 1 — routing
Checks:
- stage compatibility
- capability availability
- policy alignment
---
layer 2 — schema
Pydantic validation:
- structure
- typing
- required fields
---
layer 3 — semantic
Checks:
- hallucinated paths
- invalid references
- unsafe commands
- contradictory artifacts
---
layer 4 — approval
Checks:
- policy thresholds
- user permissions
- escalation rules
---
12. persistence model
storage backend
Default:
SQLite
Future:
- PostgreSQL
- distributed event stores
---
persisted entities
entity| purpose
events| source of truth
projections| fast reads
artifacts| outputs
approvals| audit
transitions| replay
summaries| context synthesis
---
13. observability
required telemetry
- token usage
- inference latency
- stage duration
- retries
- approval frequency
- tool failures
- transition graphs
---
debugging features
- replay from cursor
- event inspection
- context inspection
- transition trace
- artifact lineage
---
14. security model
principles
- least privilege
- explicit escalation
- isolated execution
- auditable actions
---
recommendations
- sandbox shell tools
- filesystem allowlists
- network policies
- process isolation
- execution timeouts
- secret vault integration
---
15. scalability roadmap
v1
single-node:
- sqlite
- local inference
- sequential execution
---
v2
multi-provider:
- distributed workers
- remote execution
- shared event store
---
v3
adaptive orchestration:
- evaluator models
- speculative execution
- automatic routing optimization
---
16. architectural risks
risk| mitigation
context entropy| aggressive compression
workflow spaghetti| transition tracing
infinite retries| bounded retry policies
GPU thrashing| residency scheduler
hallucinated execution| semantic validators
hidden state| event sourcing
---
17. philosophy
Harness treats LLMs as bounded semantic processors embedded inside deterministic orchestration.
Reliability emerges from:
- validation
- event sourcing
- constrained execution
- context synthesis
- approval systems
- replayability
not from trusting model intelligence alone.