Compare commits
3
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
fdb22948d9 | ||
|
|
44ba53cf71 | ||
|
|
b96b334283 |
@@ -66,6 +66,26 @@ Delegation is **mandatory** when any of these apply:
|
||||
`curl`, `hermes tools list` — these are decision-making tools. The manager
|
||||
reads them directly.
|
||||
|
||||
## Data Source Integrity (CRITICAL)
|
||||
|
||||
**Workers MUST use the data provided in their task context. They MUST NOT
|
||||
fetch their own data from external sources unless explicitly told to.**
|
||||
|
||||
When a task says "Read file X and format it", the worker reads file X. It does
|
||||
not query a separate API, run its own diagnostics, or pull data from a different
|
||||
system. This is the #1 source of cross-worker inconsistency: one worker gathers
|
||||
SSH data, another queries the Proxmox API, and the report merges two incompatible
|
||||
datasets.
|
||||
|
||||
**Rule:** If a worker needs additional data beyond what's in its task description,
|
||||
it asks the manager (via relay) — it doesn't go find it on its own.
|
||||
|
||||
**This is a hard rule, not a recommendation.** Violating it produces the exact
|
||||
type of discrepancy the kanban pipeline exists to prevent: a review worker finds
|
||||
"5 nodes present" in the raw data but "5/5 online" in the report — even though
|
||||
one of those nodes was unreachable. The report lied because it used data the
|
||||
raw data never provided.
|
||||
|
||||
## Worker Selection Matrix
|
||||
|
||||
| Worker | Model | Toolsets | Role | Use When |
|
||||
@@ -100,6 +120,19 @@ dispatch sequentially.
|
||||
|
||||
Fire workers via `delegate_task`:
|
||||
|
||||
**Critical: Pass the data, not just the goal.** When dispatching a worker that
|
||||
processes output from another worker, include the file path AND explicit
|
||||
instructions to use ONLY that source. Example:
|
||||
|
||||
```
|
||||
delegate_task(
|
||||
goal="Format the cluster check into a clean report",
|
||||
context="Source data is at /tmp/proxmox-check-raw.md. Format ONLY the data
|
||||
in that file. Do NOT query the Proxmox API or any other data source. Use the
|
||||
file as your sole source of truth."
|
||||
)
|
||||
```
|
||||
|
||||
**Parallel (independent lanes):**
|
||||
```
|
||||
delegate_task(
|
||||
@@ -203,6 +236,7 @@ Only verified results reach Kwame. Format per channel:
|
||||
- ❌ Skipping verification → raw worker output never reaches the user
|
||||
- ❌ Delegating single tool calls → keep quick reads/writes at manager level
|
||||
- ❌ Firing more than 3 workers in parallel → hard limit
|
||||
- ❌ **Workers fetching their own data sources** → a writer worker that queries the Proxmox API when told to "format the raw file" is fabricating data. Use the input given, not external sources
|
||||
|
||||
## Emergency Exception
|
||||
|
||||
|
||||
@@ -0,0 +1,210 @@
|
||||
---
|
||||
kind: pattern
|
||||
name: ra-h-os-custodianship-contract
|
||||
description: >
|
||||
Operational standards for maintaining RA-H OS as a high-functioning shared
|
||||
memory system across all agents. Defines rules for embedding consistency,
|
||||
namespace discipline, staleness management, orphan prevention, agent
|
||||
custodianship, and recovery protocols. Prevents knowledge graph degradation
|
||||
and ensures reliable semantic search capabilities.
|
||||
---
|
||||
|
||||
# RA-H OS Custodianship Contract — Shared Memory Protocol
|
||||
|
||||
## Purpose
|
||||
This contract establishes the operational standards for maintaining RA-H OS as a high-functioning shared memory system across all agents. It defines the rules, responsibilities, and recovery protocols that prevent knowledge graph degradation and ensure consistent, reliable semantic search capabilities.
|
||||
|
||||
## Core Principles
|
||||
|
||||
### 1. Embedding Consistency
|
||||
**Rule:** All nodes MUST use the RA-H OS built-in embedding pipeline via `http://192.168.68.65:8080/v1/embeddings` for consistent vector representation.
|
||||
|
||||
**Requirements:**
|
||||
- The embedding service at `192.168.68.65:8080` is the single source of truth for vector generation
|
||||
- No external embedding APIs (OpenAI, Cohere, etc.) may be used for RA-H OS nodes
|
||||
- All nodes must have `embedding_status: "chunked"` and `chunk_status: "chunked"` upon creation
|
||||
- If embedding fails, the node must be immediately flagged as `state: "unsearchable"` and logged
|
||||
|
||||
### 2. Namespace Discipline
|
||||
**Rule:** All nodes MUST be categorized into one of three namespaces with strict metadata requirements.
|
||||
|
||||
**Namespace Structure:**
|
||||
- `agent-private`: Isolated working notes for individual agents (Mumuni, Tanko, Okyeame, etc.)
|
||||
- `shared`: Policy files, registry nodes, collective knowledge
|
||||
- `syslogsolution`: Business nodes, client data, operational context
|
||||
|
||||
**Metadata Requirements:**
|
||||
Every node MUST have these metadata fields:
|
||||
```json
|
||||
{
|
||||
"agent_id": "<agent_name>",
|
||||
"namespace": "<namespace>",
|
||||
"state": "<state>",
|
||||
"type": "<type>",
|
||||
"owner": "<owner>",
|
||||
"tenant": "<tenant>",
|
||||
"visibility": "<visibility>",
|
||||
"source": "<source>",
|
||||
"captured_by": "<captured_by>"
|
||||
}
|
||||
```
|
||||
|
||||
### 3. Staleness Management
|
||||
**Rule:** Nodes are categorized by type and have specific staleness thresholds.
|
||||
|
||||
**Staleness Thresholds:**
|
||||
- `policy` type: 30 days (e.g., Shared Memory Policy, Agent Registry)
|
||||
- `registry` type: 30 days (e.g., Node State Management, Health Dashboard)
|
||||
- `template` type: 30 days (e.g., Agent SOUL.md)
|
||||
- `note` type: 60 days (e.g., business context, project notes)
|
||||
- `information` type: 90 days (e.g., reference documentation)
|
||||
- `idea` type: 30 days (e.g., brainstorming, experimental notes)
|
||||
|
||||
**Transition Protocol:**
|
||||
- `active` → `stale`: After days since `updated_at` exceeds threshold
|
||||
- `stale` → `review_pending`: Daily cron check flags for human review
|
||||
- `review_pending` → `active`: Human reviews and updates content
|
||||
- `review_pending` → `archived`: After 15 days in review_pending state (automatic purge)
|
||||
|
||||
### 4. Orphan Prevention
|
||||
**Rule:** No node should remain orphaned (0 edges) for more than 7 days.
|
||||
|
||||
**Orphan Detection:**
|
||||
- Daily cron job identifies nodes with 0 incoming/outgoing edges
|
||||
- Orphans are flagged with `state: "review_pending"` and `metadata.orphan_detected: true`
|
||||
- 7-day grace period: Orphans must be either:
|
||||
- Connected to relevant nodes via edges
|
||||
- Merged into existing parent nodes
|
||||
- Archived after 15 days in review_pending state
|
||||
|
||||
### 5. Agent Custodianship
|
||||
**Rule:** Each agent is responsible for maintaining their own created nodes.
|
||||
|
||||
**Responsibilities:**
|
||||
- **Mumuni**: Primary custodian for `syslogsolution` namespace, health dashboard, and all business nodes
|
||||
- **Tanko**: Custodian for `agent-private` namespace, personal notes, and fitness/creative work
|
||||
- **Okyeame**: Custodian for `shared` namespace, policy files, and collective knowledge nodes
|
||||
- **All Agents**: Must verify embedding status before reporting node creation as complete
|
||||
|
||||
### 6. Embedding Failure Recovery
|
||||
**Rule:** When embedding pipeline is unavailable, nodes are blocked from creation or immediately flagged.
|
||||
|
||||
**Recovery Protocol:**
|
||||
1. **Detection**: Daily health check monitors `http://192.168.68.65:8080/health` endpoint
|
||||
2. **Alert**: If embedding service is down, alert is sent to all agents via relay
|
||||
3. **Mitigation**:
|
||||
- New node creation is suspended
|
||||
- Existing nodes with `chunk_status: "not_chunked"` are flagged as `unsearchable`
|
||||
4. **Recovery**: When service returns:
|
||||
- All `not_chunked` nodes are queued for re-embedding
|
||||
- Re-embedding status logged in `embedding_retry_log` table
|
||||
- 5-minute retry window with exponential backoff (1s, 2s, 4s, 8s, 16s)
|
||||
|
||||
### 7. Chunking Verification
|
||||
**Rule:** All node modifications must verify chunking status within 5 minutes.
|
||||
|
||||
**Verification Protocol:**
|
||||
- After `updateNode` or `createNode`, query `chunks` table for `node_id`
|
||||
- If `chunk_status != "chunked"` after 5 minutes, log failure and flag node
|
||||
- Automated cleanup script runs every 6 hours to retry failed chunks
|
||||
- Maximum 3 retry attempts per node before escalation to human
|
||||
|
||||
### 8. Graph Health Monitoring
|
||||
**Rule:** Health dashboard (Node #74) must reflect real-time graph status.
|
||||
|
||||
**Dashboard Requirements:**
|
||||
- Total nodes, edges, orphans, stale nodes, chunk errors
|
||||
- Embedding service health status
|
||||
- Last successful embedding attempt
|
||||
- Alert thresholds:
|
||||
- Orphan rate > 40%: Warning
|
||||
- Stale node rate > 50%: Critical
|
||||
- Chunk error rate > 30%: Critical
|
||||
- Embedding service down: Critical (immediate alert)
|
||||
|
||||
## Operational Procedures
|
||||
|
||||
### Daily Maintenance (Automated)
|
||||
1. **Staleness Check**: Identify nodes past their threshold
|
||||
2. **Orphan Detection**: Flag nodes with 0 edges
|
||||
3. **Chunk Verification**: Retry failed chunks, log errors
|
||||
4. **Health Update**: Refresh Node #74 dashboard
|
||||
|
||||
### Weekly Maintenance (Human Review)
|
||||
1. **Review Pending Nodes**: Examine flagged nodes
|
||||
2. **Edge Optimization**: Connect related orphaned nodes
|
||||
3. **Archive Cleanup**: Remove nodes past 15-day review period
|
||||
4. **Embedding Audit**: Verify all nodes have valid embeddings
|
||||
|
||||
### Emergency Recovery
|
||||
1. **Embedding Service Down**:
|
||||
- Suspend node creation
|
||||
- Alert all agents
|
||||
- Investigate root cause (check `192.168.68.65:8080`)
|
||||
2. **Database Corruption**:
|
||||
- Restore from latest PBS backup
|
||||
- Verify chunk integrity
|
||||
- Re-run embedding for affected nodes
|
||||
3. **Mass Orphan Creation**:
|
||||
- Identify source agent/namespace
|
||||
- Review recent changes
|
||||
- Reconnect or archive affected nodes
|
||||
|
||||
## Compliance & Enforcement
|
||||
|
||||
### Audit Schedule
|
||||
- **Daily**: Automated health checks
|
||||
- **Weekly**: Human review of review_pending nodes
|
||||
- **Monthly**: Comprehensive graph audit (full schema validation)
|
||||
- **Quarterly**: Contract review and threshold adjustment
|
||||
|
||||
### Violation Consequences
|
||||
- **First**: Alert to responsible agent
|
||||
- **Second**: Node placed in `review_pending` state
|
||||
- **Third**: Agent access suspended until remediation
|
||||
- **Fourth**: Escalation to human (Kwame)
|
||||
|
||||
### Metrics for Success
|
||||
- Orphan rate: < 10%
|
||||
- Stale node rate: < 20%
|
||||
- Chunk error rate: < 5%
|
||||
- Embedding service uptime: > 99.9%
|
||||
- Node creation verification: 100%
|
||||
|
||||
## Implementation Notes
|
||||
|
||||
### Database Schema Requirements
|
||||
- `nodes` table: Standard fields plus `embedding_status`, `chunk_status`, `last_embedding_attempt`
|
||||
- `chunks` table: Must track `chunk_status` and `embedding_status`
|
||||
- `embedding_retry_log` table: Track retry attempts for failed chunks
|
||||
- `node_state_history` table: Log state transitions for audit trail
|
||||
|
||||
### Cron Job Requirements
|
||||
- `ra-h-health-check`: Daily (4h interval)
|
||||
- `ra-h-staleness-detection`: Daily
|
||||
- `ra-h-orphan-detection`: Daily
|
||||
- `ra-h-chunk-retry`: Every 6 hours
|
||||
- `ra-h-review-cleanup`: Weekly (15-day archive)
|
||||
|
||||
### Agent Onboarding
|
||||
All new agents must:
|
||||
1. Read this contract
|
||||
2. Understand namespace responsibilities
|
||||
3. Know how to verify embedding status
|
||||
4. Know the alert escalation path
|
||||
|
||||
---
|
||||
|
||||
## Approval
|
||||
This contract is effective immediately upon approval by the primary custodian (Mumuni) and system owner (Kwame).
|
||||
|
||||
**Effective Date:** July 10, 2026
|
||||
**Review Date:** October 10, 2026
|
||||
**Primary Custodian:** Mumuni 🦅
|
||||
**System Owner:** Jerome Tabiri
|
||||
|
||||
---
|
||||
|
||||
*Contract Version: 1.0*
|
||||
*Last Updated: July 10, 2026*
|
||||
*Storage: `/root/.hermes/skills/ra-h-os-custodianship-contract.prose.md`*
|
||||
Reference in New Issue
Block a user