epic-12: after epic audit and init commit
This commit is contained in:
@@ -0,0 +1,672 @@
|
||||
---
|
||||
name: "Spec V0.1"
|
||||
description: "Early system specification (Harness v0.1)"
|
||||
depth: 1
|
||||
links: ["../index.md", "../architecture/overview.md"]
|
||||
---
|
||||
|
||||
harness architecture specification
|
||||
|
||||
version: 0.1-draft
|
||||
|
||||
1. purpose
|
||||
|
||||
Harness is a config-driven orchestration runtime for local and remote LLM workflows.
|
||||
|
||||
The system treats models as interchangeable execution engines while the Harness owns:
|
||||
|
||||
- lifecycle
|
||||
- memory
|
||||
- orchestration
|
||||
- permissions
|
||||
- validation
|
||||
- event sourcing
|
||||
- workflow transitions
|
||||
- context synthesis
|
||||
|
||||
Primary goals:
|
||||
|
||||
- deterministic-enough execution on nondeterministic models
|
||||
- local-first operation
|
||||
- replayable execution
|
||||
- bounded autonomy
|
||||
- strong observability
|
||||
- minimal human steering
|
||||
- model/provider agnosticism
|
||||
|
||||
Non-goals:
|
||||
|
||||
- AGI simulation
|
||||
- unconstrained autonomous agents
|
||||
- hidden implicit memory
|
||||
- permanently accumulating context
|
||||
|
||||
---
|
||||
|
||||
2. architecture overview
|
||||
|
||||
┌──────────────────────────┐
|
||||
│ User │
|
||||
└────────────┬─────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────┐
|
||||
│ Router │
|
||||
│ conversational interface │
|
||||
└────────────┬─────────────┘
|
||||
│ summaries
|
||||
▼
|
||||
┌──────────────────────────┐
|
||||
│ Harness │
|
||||
│ orchestration kernel │
|
||||
│ event bus │
|
||||
│ approval engine │
|
||||
│ lifecycle owner │
|
||||
└────────────┬─────────────┘
|
||||
│
|
||||
┌───────┴────────┐
|
||||
▼ ▼
|
||||
┌───────────┐ ┌────────────┐
|
||||
│ Stage │ │ Context │
|
||||
│ Runtime │ │ Processor │
|
||||
└─────┬─────┘ └─────┬──────┘
|
||||
│ │
|
||||
▼ ▼
|
||||
┌──────────────────────────┐
|
||||
│ Agent Runtime │
|
||||
└────────────┬─────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────┐
|
||||
│ Inference Layer │
|
||||
│ local / remote models │
|
||||
└────────────┬─────────────┘
|
||||
│
|
||||
▼
|
||||
┌──────────────────────────┐
|
||||
│ Tools │
|
||||
└──────────────────────────┘
|
||||
|
||||
---
|
||||
|
||||
3. core principles
|
||||
|
||||
3.1 event sourced
|
||||
|
||||
All system state is reconstructable from events.
|
||||
|
||||
No mutable hidden memory exists outside event projections.
|
||||
|
||||
Benefits:
|
||||
|
||||
- replayability
|
||||
- deterministic debugging
|
||||
- auditability
|
||||
- recovery
|
||||
- analytics
|
||||
- synthetic dataset generation
|
||||
|
||||
---
|
||||
|
||||
3.2 models are stateless
|
||||
|
||||
Models do not own:
|
||||
|
||||
- memory
|
||||
- workflow state
|
||||
- permissions
|
||||
- lifecycle
|
||||
|
||||
Models receive synthesized context only.
|
||||
|
||||
---
|
||||
|
||||
3.3 validation-first execution
|
||||
|
||||
Every agent output is validated at multiple levels:
|
||||
|
||||
1. routing validation
|
||||
2. payload schema validation
|
||||
3. semantic validation
|
||||
4. approval validation
|
||||
|
||||
Invalid artifacts cannot progress workflow state.
|
||||
|
||||
---
|
||||
|
||||
3.4 context is synthesized
|
||||
|
||||
Raw accumulation is forbidden.
|
||||
|
||||
All context is:
|
||||
|
||||
- filtered
|
||||
- deduplicated
|
||||
- compressed
|
||||
- summarized
|
||||
- relevance-ranked
|
||||
|
||||
---
|
||||
|
||||
4. terminology
|
||||
|
||||
term| meaning
|
||||
Harness| orchestration kernel
|
||||
Router| conversational interface layer
|
||||
Stage| workflow state
|
||||
Agent| execution worker
|
||||
Role| semantic responsibility
|
||||
Artifact| validated structured output
|
||||
Transition| movement between stages
|
||||
Event| immutable state mutation
|
||||
Projection| derived state view
|
||||
Context Pack| synthesized model input
|
||||
Approval Gate| human/system checkpoint
|
||||
|
||||
---
|
||||
|
||||
5. configuration system
|
||||
|
||||
5.1 config locations
|
||||
|
||||
Global:
|
||||
|
||||
~/.config/harness/
|
||||
|
||||
Project-local:
|
||||
|
||||
./.harness/
|
||||
|
||||
Precedence:
|
||||
|
||||
project > user > defaults
|
||||
|
||||
---
|
||||
|
||||
5.2 config files
|
||||
|
||||
harness.yaml
|
||||
models.yaml
|
||||
tools.yaml
|
||||
stages.yaml
|
||||
policies.yaml
|
||||
compression.yaml
|
||||
|
||||
---
|
||||
|
||||
6. model registry
|
||||
|
||||
6.1 model definition
|
||||
|
||||
models:
|
||||
qwen_coder_14b:
|
||||
provider: local
|
||||
path: /models/qwen-coder.gguf
|
||||
|
||||
capabilities:
|
||||
- coding
|
||||
- tool_calling
|
||||
- reasoning
|
||||
|
||||
context_size: 32768
|
||||
|
||||
kv_cache:
|
||||
quantization: q8_0
|
||||
|
||||
gpu:
|
||||
layers: 48
|
||||
residency: dynamic
|
||||
swap_timeout_sec: 300
|
||||
|
||||
inference:
|
||||
temperature: 0.2
|
||||
top_p: 0.9
|
||||
repeat_penalty: 1.1
|
||||
|
||||
limits:
|
||||
max_parallel_sessions: 2
|
||||
|
||||
---
|
||||
|
||||
7. providers
|
||||
|
||||
7.1 local provider
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- llama.cpp process management
|
||||
- GPU residency
|
||||
- model loading/unloading
|
||||
- warm pools
|
||||
- swap scheduling
|
||||
- health monitoring
|
||||
|
||||
---
|
||||
|
||||
7.2 remote provider
|
||||
|
||||
providers:
|
||||
openai_compatible:
|
||||
base_url: https://api.example.com/v1
|
||||
api_key_env: HARNESS_API_KEY
|
||||
|
||||
extra:
|
||||
timeout: 120
|
||||
retries: 3
|
||||
|
||||
---
|
||||
|
||||
8. stages
|
||||
|
||||
Stages define execution states.
|
||||
|
||||
Example:
|
||||
|
||||
stages:
|
||||
implementation:
|
||||
role: coder
|
||||
|
||||
requirements:
|
||||
- coding
|
||||
- tool_calling
|
||||
|
||||
allowed_tools:
|
||||
- filesystem
|
||||
- git
|
||||
- shell
|
||||
|
||||
approvals:
|
||||
tool_tier_3: required
|
||||
|
||||
transitions:
|
||||
on_success:
|
||||
- validation
|
||||
on_failure:
|
||||
- retry
|
||||
|
||||
---
|
||||
|
||||
9. agents
|
||||
|
||||
Agents are ephemeral execution workers.
|
||||
|
||||
Agents:
|
||||
|
||||
- consume Context Packs
|
||||
- emit Artifacts
|
||||
- do not persist memory
|
||||
|
||||
---
|
||||
|
||||
10. router
|
||||
|
||||
Router responsibilities:
|
||||
|
||||
- user interaction
|
||||
- conversational continuity
|
||||
- summarizing system state
|
||||
- steering interpretation
|
||||
|
||||
Router is always available but not persistent in execution context.
|
||||
|
||||
Execution agents unload router context during active work.
|
||||
|
||||
After completion:
|
||||
|
||||
- outputs are summarized
|
||||
- summaries are injected into router memory
|
||||
|
||||
---
|
||||
|
||||
11. artifacts
|
||||
|
||||
All stage outputs MUST emit structured artifacts.
|
||||
|
||||
Example:
|
||||
|
||||
class ImplementationArtifact(BaseModel):
|
||||
summary: str
|
||||
modified_files: list[str]
|
||||
risks: list[str]
|
||||
commands_executed: list[str]
|
||||
next_recommendation: str
|
||||
|
||||
---
|
||||
|
||||
12. validation pipeline
|
||||
|
||||
12.1 routing validation
|
||||
|
||||
Checks:
|
||||
|
||||
- valid stage transitions
|
||||
- capability compatibility
|
||||
- policy compliance
|
||||
|
||||
---
|
||||
|
||||
12.2 payload validation
|
||||
|
||||
Pydantic validation:
|
||||
|
||||
- schema correctness
|
||||
- required fields
|
||||
- typing
|
||||
|
||||
---
|
||||
|
||||
12.3 semantic validation
|
||||
|
||||
Checks:
|
||||
|
||||
- contradictory outputs
|
||||
- hallucinated files
|
||||
- invalid commands
|
||||
- policy violations
|
||||
- unsafe operations
|
||||
|
||||
---
|
||||
|
||||
13. approval system
|
||||
|
||||
13.1 approval tiers
|
||||
|
||||
tier| meaning
|
||||
T0| inference only
|
||||
T1| read-only
|
||||
T2| reversible mutation
|
||||
T3| external/network
|
||||
T4| destructive
|
||||
|
||||
---
|
||||
|
||||
13.2 approval modes
|
||||
|
||||
mode| meaning
|
||||
prompt| require confirmation
|
||||
auto| auto approve
|
||||
deny| reject automatically
|
||||
yolo| bypass safeguards
|
||||
|
||||
---
|
||||
|
||||
13.3 approval actions
|
||||
|
||||
User may:
|
||||
|
||||
- approve
|
||||
- reject
|
||||
- auto-approve for session
|
||||
- steer execution
|
||||
|
||||
Example:
|
||||
|
||||
approved, but verify auth edge cases first
|
||||
|
||||
---
|
||||
|
||||
14. event system
|
||||
|
||||
14.1 event categories
|
||||
|
||||
DomainEvents
|
||||
SystemEvents
|
||||
InferenceEvents
|
||||
ToolEvents
|
||||
ApprovalEvents
|
||||
CompressionEvents
|
||||
LifecycleEvents
|
||||
|
||||
---
|
||||
|
||||
14.2 event structure
|
||||
|
||||
class Event(BaseModel):
|
||||
id: UUID
|
||||
session_id: UUID
|
||||
timestamp: datetime
|
||||
type: str
|
||||
payload: dict
|
||||
causation_id: UUID | None
|
||||
correlation_id: UUID | None
|
||||
|
||||
---
|
||||
|
||||
15. replay system
|
||||
|
||||
Replay modes:
|
||||
|
||||
- full replay
|
||||
- partial replay
|
||||
- replay from cursor
|
||||
- inference-skipping replay
|
||||
- deterministic simulation
|
||||
|
||||
Replay reconstructs projections and workflow state.
|
||||
|
||||
---
|
||||
|
||||
16. context processor
|
||||
|
||||
The ContextProcessor synthesizes minimal relevant context.
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- deduplication
|
||||
- summarization
|
||||
- ranking
|
||||
- token budgeting
|
||||
- artifact extraction
|
||||
- tool compression
|
||||
|
||||
---
|
||||
|
||||
16.1 context layers
|
||||
|
||||
L0 live execution
|
||||
L1 stage-local context
|
||||
L2 compressed session memory
|
||||
L3 project memory
|
||||
L4 archival history
|
||||
|
||||
---
|
||||
|
||||
16.2 compression strategies
|
||||
|
||||
compression:
|
||||
tool_logs:
|
||||
mode: summarize
|
||||
|
||||
artifacts:
|
||||
mode: latest_only
|
||||
|
||||
events:
|
||||
mode: deduplicate
|
||||
|
||||
conversations:
|
||||
mode: semantic_summary
|
||||
|
||||
---
|
||||
|
||||
17. tool system
|
||||
|
||||
Tools are config-driven and capability-scoped.
|
||||
|
||||
Default tools may be:
|
||||
|
||||
- disabled
|
||||
- replaced
|
||||
- overridden
|
||||
|
||||
Example:
|
||||
|
||||
tools:
|
||||
shell:
|
||||
enabled: true
|
||||
tier: T2
|
||||
|
||||
git:
|
||||
enabled: true
|
||||
tier: T2
|
||||
|
||||
curl:
|
||||
enabled: true
|
||||
tier: T3
|
||||
|
||||
---
|
||||
|
||||
18. transitions
|
||||
|
||||
Transitions are rule-based.
|
||||
|
||||
No hardcoded workflow graphs exist in code.
|
||||
|
||||
Example:
|
||||
|
||||
transitions:
|
||||
- when:
|
||||
artifact.status == "success"
|
||||
goto: validation
|
||||
|
||||
- when:
|
||||
retries > 3
|
||||
goto: failed
|
||||
|
||||
---
|
||||
|
||||
19. retry policies
|
||||
|
||||
19.1 strategies
|
||||
|
||||
strategy| behavior
|
||||
retry| retry execution
|
||||
fail_safe| skip and continue
|
||||
fail_fast| terminate session
|
||||
|
||||
---
|
||||
|
||||
19.2 retry configuration
|
||||
|
||||
retry:
|
||||
strategy: corrective
|
||||
max_attempts: 3
|
||||
inject_failure_reason: true
|
||||
temperature_backoff: true
|
||||
|
||||
---
|
||||
|
||||
20. persistence
|
||||
|
||||
Default persistence:
|
||||
|
||||
SQLite
|
||||
|
||||
Future:
|
||||
|
||||
- PostgreSQL
|
||||
- event stores
|
||||
- distributed backends
|
||||
|
||||
Persisted:
|
||||
|
||||
- events
|
||||
- projections
|
||||
- artifacts
|
||||
- approvals
|
||||
- transitions
|
||||
- summaries
|
||||
|
||||
---
|
||||
|
||||
21. session lifecycle
|
||||
|
||||
Harness owns session lifecycle.
|
||||
|
||||
States:
|
||||
created
|
||||
active
|
||||
paused
|
||||
awaiting_approval
|
||||
failed
|
||||
completed
|
||||
cancelled
|
||||
|
||||
---
|
||||
|
||||
22. observability
|
||||
|
||||
Required:
|
||||
|
||||
- event tracing
|
||||
- transition tracing
|
||||
- inference timing
|
||||
- token accounting
|
||||
- tool execution logs
|
||||
- replay diagnostics
|
||||
|
||||
Recommended:
|
||||
|
||||
- DAG visualization
|
||||
- live stage graph
|
||||
- approval history
|
||||
|
||||
---
|
||||
|
||||
23. security model
|
||||
|
||||
Principles:
|
||||
|
||||
- least privilege
|
||||
- explicit approvals
|
||||
- isolated tools
|
||||
- auditability
|
||||
- bounded execution
|
||||
|
||||
Recommendations:
|
||||
|
||||
- sandbox shell tools
|
||||
- filesystem allowlists
|
||||
- network policy control
|
||||
- secret isolation
|
||||
- execution timeouts
|
||||
|
||||
---
|
||||
|
||||
24. future extensions
|
||||
|
||||
Potential:
|
||||
|
||||
- distributed agents
|
||||
- evaluator models
|
||||
- speculative execution
|
||||
- long-term semantic memory
|
||||
- automatic fine-tuning corpus extraction
|
||||
- capability benchmarking
|
||||
- adaptive routing
|
||||
|
||||
---
|
||||
|
||||
25. anti-goals
|
||||
|
||||
Avoid:
|
||||
|
||||
- hidden prompts
|
||||
- invisible memory mutation
|
||||
- unrestricted recursion
|
||||
- self-modifying workflows
|
||||
- implicit approvals
|
||||
- context accumulation without compression
|
||||
- permanent agent processes
|
||||
|
||||
---
|
||||
|
||||
26. philosophy summary
|
||||
|
||||
Harness is not an “AI agent framework”.
|
||||
|
||||
It is:
|
||||
|
||||
- an orchestration kernel
|
||||
- an event-sourced execution runtime
|
||||
- a bounded autonomy system
|
||||
- a deterministic shell around probabilistic cognition
|
||||
Reference in New Issue
Block a user