Files
prose-contracts/contract-registry.yaml
root de32f54337
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
feat(daily-digest): deliver via Zulip DM as an HTML attachment; drop mail entirely
Captain's decision 2026-09-26, clarified the same day: the digest is delivered to
his Zulip DM (user id 9) from abiba-bot as an HTML FILE - an attachment, not HTML
rendered in the message body and not a Markdown translation of it. Closes
daily-digest-mail-transport-20260921; the Google dependency is gone (no SMTP, no
EMAIL_PASSWORD, no app password, nothing to rotate).

WHAT CHANGES
* scripts/daily-infra-report.py: send_email() is replaced by send_zulip(), which
  writes the styled dashboard to /var/log/daily-infra-report/infra-report-<ts>.html,
  uploads it via POST /api/v1/user_uploads, then posts a SHORT Markdown pointer to
  user 9. The message body carries subject, top-line status and the attachment
  link; it does not reproduce the report.
* the 10,000-character cap is irrelevant here - it bounds message TEXT only, and
  the report travels as a file, so nothing is shrunk to fit.
* the key is abiba-bot's, already on the execution host at
  /root/.pi/agent/extensions/zulip/.env (mode 600). No vault entry was added:
  under the auth-keys charter that is a captain decision.
* daily-health-digest.prose.md -> v2.0.0 and contract-registry.yaml updated:
  transport, healthy/degraded definitions, and exit codes now match observed
  behaviour. There is NO degraded delivery leg any more - delivery is the only
  output path, so a missing or rejected key is a real failure (exit 1).
* queued defect folded in: a failed delivery used to print only the transport
  error while the report body never surfaced. Now the HTML is printed to stdout
  AND persisted on every failure, and the message names which step failed.

EVIDENCE (all against the live stack)
* real send: message id 86221 to user 9, attachment 16208 bytes at
  /user_uploads/2/45/m1cQesBFV78BGeNY2lN8xkN5/infra-report-20260926-153406.html
* the message is type=private, sender abiba-bot@chat.sysloggh.net, recipients
  [9, 21], body carries the top-line status and the attachment link, and does NOT
  contain a <table> - i.e. it does not reproduce the report
* the attachment fetches HTTP 200, 16208 bytes, content-type text/html, starts
  with <!DOCTYPE html>, and contains <style>, <table> and 16 class="card" blocks -
  it opens as a standalone styled document
* failure path: a bad key gives 'Delivery FAILED at upload: Malformed API key',
  EXIT=1, the HTML is printed to stdout and persisted to disk
* scheduled path: the run's own output is pasted in the PR

prose-lint: PASSED (19 warnings); secret scan clean.
2026-09-26 15:35:36 +00:00

2056 lines
50 KiB
YAML

registry_version: 0.1.0
last_updated: '2026-09-11T00:00:00Z'
updated_by: mumuni
categories:
- compliance
- monitoring
- remediation
- build_deploy
- maintenance
- reference
domains:
- hermes-agent
- zulip
- infrastructure
- gpu
- proxmox
- litellm
- memory
- relay
- shared-patterns
owners:
- abiba
- mumuni
- kwame
- ops
trigger_types:
- scheduled
- event_driven
- passive
- on_demand
- none
sensitivity_levels:
- critical
- high
- normal
- retired
index:
by_category:
compliance:
- hermes-key-enforcement
- litellm-api-keys
- hermes-config-template
- hermes-agent-baseline
monitoring:
- proxmox-monitor
- gpu-monitor
- infrastructure-monitoring
- zulip-health
- litellm-health
- daily-health-digest
remediation:
- litellm-self-heal
- pm2-self-heal
- disk-gc-threat-response
- hermes-zulip-restore
build_deploy:
- hermes-zulip-plugin
- build-zulip-plugin
- stirling-pdf-agent-access
maintenance:
- memory-audit-maintenance
- gpu-fleet
- infrastructure-update
- infrastructure-maintenance
reference:
- infrastructure-control
- ra-h-os-custodianship-contract
- mumuni-delegation
- zulip-approval-fix
- zulip-oidc-redirect-fix
- hello-world
- zulip-adapter-lessons
- pi-approval-architecture
retired:
- zulip-mention-reliability
- zulip-self-heal
by_domain:
hermes-agent:
- hermes-key-enforcement
- hermes-config-template
- hermes-agent-baseline
- mumuni-delegation
zulip:
- zulip-health
- hermes-zulip-restore
- hermes-zulip-plugin
- build-zulip-plugin
- zulip-approval-fix
- zulip-oidc-redirect-fix
- zulip-adapter-lessons
- zulip-mention-reliability
- zulip-self-heal
infrastructure:
- infrastructure-control
- infrastructure-monitoring
- infrastructure-update
- infrastructure-maintenance
- pm2-self-heal
- disk-gc-threat-response
- daily-health-digest
gpu:
- gpu-monitor
- gpu-fleet
proxmox:
- proxmox-monitor
litellm:
- litellm-api-keys
- litellm-health
- litellm-self-heal
memory:
- memory-audit-maintenance
shared-patterns:
- ra-h-os-custodianship-contract
- pi-approval-architecture
- stirling-pdf-agent-access
- hello-world
retired:
- zulip-mention-reliability
- zulip-self-heal
- zulip-adapter-lessons
- pi-approval-architecture
by_owner:
abiba:
- hermes-key-enforcement
- hermes-config-template
- hermes-agent-baseline
- proxmox-monitor
- gpu-monitor
- infrastructure-monitoring
- zulip-health
- litellm-health
- litellm-self-heal
- pm2-self-heal
- disk-gc-threat-response
- hermes-zulip-restore
- hermes-zulip-plugin
- build-zulip-plugin
- stirling-pdf-agent-access
- gpu-fleet
- infrastructure-control
- zulip-adapter-lessons
- pi-approval-architecture
- zulip-approval-fix
- zulip-oidc-redirect-fix
- zulip-mention-reliability
- zulip-self-heal
mumuni:
- memory-audit-maintenance
- ra-h-os-custodianship-contract
- mumuni-delegation
kwame:
- hello-world
ops:
- infrastructure-maintenance
- infrastructure-update
by_trigger:
scheduled:
- hermes-key-enforcement
- hermes-config-template
- hermes-agent-baseline
- proxmox-monitor
- gpu-monitor
- infrastructure-monitoring
- zulip-health
- litellm-health
- memory-audit-maintenance
- infrastructure-update
- infrastructure-maintenance
event_driven:
- litellm-self-heal
- pm2-self-heal
- disk-gc-threat-response
- hermes-zulip-restore
- gpu-fleet
- zulip-oidc-redirect-fix
passive:
- infrastructure-control
- ra-h-os-custodianship-contract
- mumuni-delegation
- zulip-adapter-lessons
- pi-approval-architecture
- zulip-approval-fix
on_demand:
- hermes-zulip-plugin
- build-zulip-plugin
- stirling-pdf-agent-access
- hello-world
none:
- zulip-mention-reliability
- zulip-self-heal
by_sensitivity:
critical:
- hermes-key-enforcement
- proxmox-monitor
- gpu-monitor
- litellm-health
- litellm-self-heal
- disk-gc-threat-response
- gpu-fleet
- infrastructure-control
high:
- hermes-config-template
- infrastructure-monitoring
- zulip-health
- pm2-self-heal
- hermes-zulip-restore
- hermes-zulip-plugin
- build-zulip-plugin
- infrastructure-update
- infrastructure-maintenance
- ra-h-os-custodianship-contract
- mumuni-delegation
normal:
- hermes-agent-baseline
- stirling-pdf-agent-access
- memory-audit-maintenance
- zulip-approval-fix
- zulip-oidc-redirect-fix
- hello-world
retired:
- zulip-adapter-lessons
- pi-approval-architecture
- zulip-mention-reliability
- zulip-self-heal
contracts:
- name: hermes-key-enforcement
file: hermes-key-enforcement.prose.md
kind: enforcement
category: compliance
sensitivity: critical
status: active
owner: abiba
version: 1.0.0
trigger:
type: scheduled
cadence: 0 6 * * *
description: Daily compliance scan at 6am ET
cron_job_id: null
execution:
agent: abiba
timeout: 300
requires: []
protocol:
- Load contract from prose-contracts/main
- 'Verify prerequisites: config.yaml readable, harness/litellm sections present'
- 'Execute enforcement: grep config for plaintext keys'
- 'On failure: stop, don''t hallucinate forward'
- Log actions to ~/.hermes/runs/hermes-key-enforcement/
verification:
postconditions:
- check: no plaintext API keys in config
verify: 'grep -rc ''api_key: sk-'' /root/.hermes/config.yaml'
expect: 0 matches
- check: api_key_env used for harness/litellm providers
verify: grep -c 'api_key_env.*LITELLM_API_KEY' /root/.hermes/config.yaml
expect: count > 0
artifact: before/after verification report
verify_commands:
- 'grep -rc ''api_key: sk-'' /root/.hermes/config.yaml'
- grep -c 'api_key_env.*LITELLM_API_KEY' /root/.hermes/config.yaml
receipt:
format: json
storage: ~/.hermes/runs/hermes-key-enforcement/
graph_node: true
schema:
contract: string
run_id: string
timestamp: ISO 8601
agent: string
status: pass|fail|escalated
actions_taken: array
postconditions: array
drift_alerts: array
evidence_path: string
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
critical:
action: relay_alert + pause
notify:
- abiba
- mumuni
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: hermes-config-template
file: hermes-config-template.prose.md
kind: template
category: compliance
sensitivity: high
status: active
owner: abiba
version: 2.1.0
trigger:
type: scheduled
cadence: 0 4 * * 1
description: Weekly config drift check Monday at 4am ET
cron_job_id: null
execution:
agent: abiba
timeout: 300
requires: []
verification:
postconditions:
- check: agent config template_version matches template file
verify: 'grep -q ''template_version'' /root/.hermes/config.yaml && diff <(grep
''template_version'' /root/.hermes/config.yaml | cut -d: -f2 | xargs) <(grep
''template_version'' /root/prose-contracts/hermes-config-template.prose.md
| cut -d: -f2 | xargs) && echo match || echo mismatch'
expect: match
- check: config file is valid YAML
verify: python3 -c 'import yaml; yaml.safe_load(open("/root/.hermes/config.yaml"))'
&& echo valid || echo invalid
expect: valid
artifact: config drift report
receipt:
format: json
storage: ~/.hermes/runs/hermes-config-template/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: hermes-agent-baseline
file: hermes-agent-baseline.prose.md
kind: template
category: compliance
sensitivity: normal
status: active
owner: abiba
version: 1.0.0
trigger:
type: scheduled
cadence: 0 5 * * 1
description: Weekly baseline verification Monday at 5am ET
cron_job_id: null
execution:
agent: abiba
timeout: 300
requires: []
verification:
postconditions:
- check: Hermes agent process running
verify: pgrep -f 'hermes' > /dev/null && echo running || echo stopped
expect: running
- check: agent config file exists and valid YAML
verify: test -f /root/.hermes/config.yaml && python3 -c 'import yaml; yaml.safe_load(open("/root/.hermes/config.yaml"))'
&& echo valid || echo invalid
expect: valid
- check: no uncommitted changes in hermes directory
verify: cd /root/.hermes && git status --porcelain | wc -l
expect: '0'
artifact: baseline drift report
receipt:
format: json
storage: ~/.hermes/runs/hermes-agent-baseline/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: proxmox-monitor
file: proxmox-monitor.prose.md
kind: responsibility
category: monitoring
sensitivity: critical
status: active
owner: abiba
version: 1.0.0
trigger:
type: scheduled
cadence: '*/15 * * * *'
description: Every 15 minutes
cron_job_id: null
execution:
agent: abiba
timeout: 120
requires: []
protocol:
- Load contract from prose-contracts/main
- Verify connectivity to all Proxmox nodes
- Execute monitoring checks per contract
- Log intermediate state
- 'On failure: stop, don''t hallucinate'
verification:
postconditions:
- check: all Proxmox nodes reachable
verify: curl -sf http://192.168.68.10:8006/api2/json/status | jq '.status'
expect: healthy
- check: no VMs in crashed state
verify: pvesh get /nodes -output-format=json | jq '.[] | select(.status=="Crashed")'
expect: empty
- check: backups running on schedule
verify: pbs-info --check
expect: last_backup < 24h ago
artifact: cluster health snapshot
verify_commands:
- curl -sf http://<pve-api>/status | jq '.status'
- pvesh get /nodes -output-format=json
receipt:
format: json
storage: ~/.hermes/runs/proxmox-monitor/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert + trigger_remediation
notify:
- abiba
- mumuni
triggers:
- disk-gc-threat-response
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: gpu-monitor
file: gpu-monitor.prose.md
kind: responsibility
category: monitoring
sensitivity: critical
status: active
owner: abiba
version: 1.0.0
trigger:
type: scheduled
cadence: '*/15 * * * *'
description: "Every 15 minutes \u2014 polls all GPU subsystems"
cron_job_id: null
execution:
agent: abiba
timeout: 120
requires: []
verification:
postconditions:
- check: GPU metrics accessible
verify: curl -sf http://localhost:9100/gpu-data
expect: 200 OK, populated data
- check: dashboard serving
verify: curl -sf http://localhost:9100/gpu-fleet.html
expect: 200 OK, HTML returned
- check: health endpoint responsive
verify: curl -sf http://localhost:9100/health
expect: 200 OK
artifact: GPU fleet snapshot
receipt:
format: json
storage: ~/.hermes/runs/gpu-monitor/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert + trigger_remediation
notify:
- abiba
- mumuni
triggers:
- gpu-fleet
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-monitoring
file: infrastructure-monitoring.prose.md
kind: function
category: monitoring
sensitivity: high
status: active
owner: abiba
version: 1.0.0
trigger:
type: scheduled
cadence: '*/30 * * * *'
description: Every 30 minutes
cron_job_id: null
execution:
agent: abiba
timeout: 180
requires: []
verification:
postconditions:
- check: Proxmox API reachable
verify: curl -sf http://192.168.68.10:8006/api2/json
expect: 200 OK
- check: Zulip API reachable
verify: curl -sf https://chat.sysloggh.net/api/v1/me
expect: 200 OK
- check: LiteLLM proxy reachable
verify: curl -sf http://192.168.68.116/litellm/v1/models
expect: 200 OK
- check: Gitea API reachable
verify: curl -sf https://git.sysloggh.net/api/v1/version
expect: 200 OK
- check: SearXNG reachable
verify: curl -sf http://192.168.68.7:8888
expect: 200 OK
artifact: infrastructure health report
receipt:
format: json
storage: ~/.hermes/runs/infrastructure-monitoring/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert + trigger_remediation
notify:
- abiba
- mumuni
triggers:
- litellm-health
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: zulip-health
file: zulip-health.prose.md
kind: responsibility
category: monitoring
sensitivity: high
status: active
owner: abiba
version: 3.4.0
trigger:
type: scheduled
cadence: '*/15 * * * *'
description: "Every 15 minutes \u2014 monitors the Zulip-connected agents under this host's control (pi, DSH, Agent Zero)"
cron_job_id: null
execution:
agent: abiba
timeout: 120
requires:
- Zulip API key for abiba-bot@chat.sysloggh.net
- SSH access to amdpve (192.168.68.15) for Tanko (CT 112) and the Agent Zero Docker host (.14)
verification:
postconditions:
- check: bot registration active
verify: curl -sf https://chat.sysloggh.net/api/v1/me | jq '.user_id'
expect: bot_id present
- check: DM delivery working
verify: curl -sf https://chat.sysloggh.net/api/v1/users/me/is-online
expect: 'online: true'
artifact: Zulip mesh health report
receipt:
format: json
storage: ~/.hermes/runs/zulip-health/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert + trigger_remediation
notify:
- abiba
- mumuni
triggers:
- hermes-zulip-restore
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: litellm-health
file: litellm-health.prose.md
kind: function
category: monitoring
sensitivity: critical
status: active
owner: abiba
version: 1.0.0
trigger:
type: scheduled
cadence: '5 3,7,11,15,19,23 * * *'
description: "4-hourly staggered dispatch via fm-send (run contract litellm-health)"
cron_job_id: null
execution:
agent: abiba
timeout: 60
requires: []
verification:
postconditions:
- check: LiteLLM proxy reachable
verify: curl -sf http://192.168.68.116/litellm/v1/models
expect: 200 OK, models returned
- check: router deprecated, nginx routes work
verify: curl -sf https://litellm.sysloggh.net/v1/models
expect: 200 OK (via nginx)
artifact: LiteLLM health snapshot
receipt:
format: json
storage: ~/.hermes/runs/litellm-health/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert + trigger_remediation
notify:
- abiba
- mumuni
triggers:
- litellm-self-heal
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: litellm-self-heal
file: litellm-self-heal.prose.md
kind: responsibility
category: remediation
sensitivity: critical
status: active
owner: abiba
version: 1.1.0
trigger:
type: event_driven
description: Triggered by relay message from litellm-health or infrastructure-monitoring
relay_message_type: remediation:litellm-self-heal
cron_job_id: null
execution:
agent: abiba
timeout: 600
requires:
- Relay message with failure details from monitoring contract
verification:
postconditions:
- check: LiteLLM proxy responding
verify: curl -sf http://192.168.68.116/litellm/v1/models
expect: 200 OK
artifact: self-heal run report
receipt:
format: json
storage: ~/.hermes/runs/litellm-self-heal/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert + pause
notify:
- abiba
- mumuni
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: pm2-self-heal
file: pm2-self-heal.prose.md
kind: responsibility
category: remediation
sensitivity: high
status: active
owner: abiba
version: 1.0.0
trigger:
type: event_driven
description: Triggered by relay message when PM2 process crashes
relay_message_type: remediation:pm2-self-heal
cron_job_id: null
execution:
agent: abiba
timeout: 300
requires:
- Relay message with process name and failure details
verification:
postconditions:
- check: PM2 process running
verify: pm2 list | grep <process_name>
expect: online
artifact: self-heal run report
receipt:
format: json
storage: ~/.hermes/runs/pm2-self-heal/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: disk-gc-threat-response
file: disk-gc-threat-response.prose.md
kind: responsibility
category: remediation
sensitivity: critical
status: active
owner: abiba
version: 1.0.0
trigger:
type: event_driven
description: Triggered by relay message from proxmox-monitor when disk >80%
relay_message_type: remediation:disk-gc-threat-response
cron_job_id: null
execution:
agent: abiba
timeout: 900
requires:
- Relay message with target CT/host and current disk usage
protocol:
- Load contract from prose-contracts/main
- Verify target node reachable
- Run garbage collection per contract SOP
- Verify disk usage reduced
verification:
postconditions:
- check: disk usage below 80%
verify: df -h <target>
expect: <80%
- check: no critical services stopped
verify: ps aux | grep <critical_services>
expect: all running
artifact: GC run report with reclaimed space
receipt:
format: json
storage: ~/.hermes/runs/disk-gc-threat-response/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert + pause
notify:
- abiba
- mumuni
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: hermes-zulip-restore
file: hermes-zulip-restore.prose.md
kind: function
category: remediation
sensitivity: high
status: active
owner: abiba
version: 1.0.0
trigger:
type: event_driven
description: Triggered by relay message from zulip-health when bot registration
fails
relay_message_type: remediation:hermes-zulip-restore
cron_job_id: null
execution:
agent: abiba
timeout: 600
requires:
- Relay message with failure details
verification:
postconditions:
- check: Zulip bot registration active
verify: curl -sf https://chat.sysloggh.net/api/v1/me | jq '.user_id'
expect: bot_id present
- check: DM delivery working
verify: send test DM
expect: message delivered
artifact: restore run report
receipt:
format: json
storage: ~/.hermes/runs/hermes-zulip-restore/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert + pause
notify:
- abiba
- mumuni
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: hermes-zulip-plugin
file: hermes-zulip-plugin.prose.md
kind: function
category: build_deploy
sensitivity: high
status: active
owner: abiba
version: 1.0.0
trigger:
type: on_demand
description: Manual or change request to update Zulip plugin
cron_job_id: null
execution:
agent: abiba
timeout: 1800
requires: []
verification:
postconditions:
- check: 3 adapter files present
verify: ls -la /usr/local/lib/hermes-agent/plugins/platforms/zulip/
expect: 3 files
- check: gateway connected
verify: ps aux | grep hermes
expect: zulip connected
artifact: deploy verification report
receipt:
format: json
storage: ~/.hermes/runs/hermes-zulip-plugin/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 1
window: 7200
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: build-zulip-plugin
file: build-zulip-plugin.prose.md
kind: responsibility
category: build_deploy
sensitivity: high
status: active
owner: abiba
version: 1.0.0
trigger:
type: on_demand
description: Generate or improve Zulip plugin via iterative refinement
cron_job_id: null
execution:
agent: abiba
timeout: 3600
requires: []
verification:
postconditions:
- check: plugin builds without errors
verify: npm install && npm run build
expect: 0 errors
- check: connectivity tests pass
verify: run plugin selftest
expect: all checks green
artifact: plugin version artifact
receipt:
format: json
storage: ~/.hermes/runs/build-zulip-plugin/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 1
window: 7200
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: stirling-pdf-agent-access
file: stirling-pdf-agent-access.prose.md
kind: function
category: build_deploy
sensitivity: normal
status: active
owner: abiba
version: 1.0.0
trigger:
type: on_demand
description: Grant agent access to Stirling PDF
cron_job_id: null
execution:
agent: abiba
timeout: 300
requires: []
verification:
postconditions:
- check: agent can access Stirling PDF
verify: curl -sf https://stirling.sysloggh.net
expect: 200 OK
artifact: access verification report
receipt:
format: json
storage: ~/.hermes/runs/stirling-pdf-agent-access/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 1
window: 7200
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: memory-audit-maintenance
file: memory-audit-maintenance.prose.md
kind: responsibility
category: maintenance
sensitivity: normal
status: active
owner: mumuni
version: 1.0.0
trigger:
type: scheduled
cadence: 0 3 * * *
description: Daily at 3am ET
cron_job_id: b59f3cc21f4c # provisioned on kagentz 2026-09-08 (okyeame-memory-audit, glm-5.3-flash)
execution:
agent: mumuni
timeout: 600
requires: []
protocol:
- Load contract from prose-contracts/main
- 'Verify prerequisites: memory files accessible'
- Execute memory audit per contract SOP
- Log actions to ~/.hermes/runs/memory-audit-maintenance/
verification:
postconditions:
- check: memory files below 80% capacity
verify: wc -l ~/.hermes/memories/*.md
expect: total lines < threshold
- check: no stale entries
verify: grep -r 'STALE' ~/.hermes/memories/
expect: 0 matches
artifact: memory audit report
receipt:
format: json
storage: ~/.hermes/runs/memory-audit-maintenance/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- mumuni
critical:
action: relay_alert
notify:
- mumuni
- abiba
fatal:
action: relay_alert + human_required
notify:
- mumuni
- abiba
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: gpu-fleet
file: gpu-fleet.prose.md
kind: responsibility
category: maintenance
sensitivity: critical
status: active
owner: abiba
version: 1.0.0
trigger:
type: event_driven
description: Triggered on model add/remove, GPU health degradation, agent key
rotation
cron_job_id: null
execution:
agent: abiba
timeout: 1800
requires: []
verification:
postconditions:
- check: all GPUs reported to LiteLLM
verify: curl -sf http://192.168.68.116/litellm/v1/models | jq '.data | length'
expect: count matches expected
- check: router deprecated
verify: check nginx routes
expect: no /v1/ prefix routes
artifact: GPU fleet status report
receipt:
format: json
storage: ~/.hermes/runs/gpu-fleet/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-maintenance
file: infrastructure-maintenance.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: ops
version: 1.0.0
trigger:
type: scheduled
cadence: 0 2 * * 0
description: Weekly host-level maintenance Sunday at 2am ET (replaces infrastructure-update
build-phase role; infra-update moves to ops)
cron_job_id: null
execution:
agent: ops
timeout: 3600
requires:
- infrastructure-monitoring run within last 30 minutes (pre-update health baseline)
- Proxmox snapshot of primary host OR /tmp backup dir created this run
- LiteLLM master key from Infisical vault for health verification
protocol:
- Load contract from prose-contracts/main
- Phase 0 preflight — capture health baseline, backup check, record image baseline
- Phase 1 apt update && apt upgrade -y on primary host
- Phase 2 docker compose pull for LiteLLM, SearXNG, and other running containers
- Phase 3 restart stacks one at a time with per-stack health verification
- Phase 4 verify every critical service (LiteLLM, SearXNG, Zulip, Gitea, PM2, Hermes gateways)
- On failure — rollback per protocol, escalate, do not loop beyond circuit breaker
- Log actions to ~/.hermes/runs/infrastructure-maintenance/
verification:
postconditions:
- check: all critical services running after update
verify: 'curl -sf http://192.168.68.116/litellm/v1/models && curl -sf https://chat.sysloggh.net/api/v1/server_settings && curl -sf https://git.sysloggh.net/api/v1/version && curl -sf http://192.168.68.7:8888 && pm2 jlist'
expect: all probes 200 OK / processes online
- check: no regressions from pre-update health baseline
verify: diff Phase 0 health-baseline against Phase 4 results
expect: no GREEN service turned RED
- check: docker containers on latest stable tags
verify: docker inspect --format '{{.Config.Image}}' <container> per service matches image-baseline.pulled_tag
expect: all containers running pulled tags
- check: APT packages up to date with no held broken packages
verify: apt list --upgradable 2>/dev/null | wc -l and apt-get -s upgrade | grep -ci broken
expect: upgradable == 0, broken == 0
artifact: maintenance run report with phase results and any rollback/escalation
verify_commands:
- curl -sf http://192.168.68.116/litellm/v1/models
- curl -sf http://192.168.68.7:8888
- curl -sf https://chat.sysloggh.net/api/v1/server_settings
- curl -sf https://git.sysloggh.net/api/v1/version
- pm2 jlist
receipt:
format: json
storage: ~/.hermes/runs/infrastructure-maintenance/
graph_node: true
schema:
contract: string
run_id: string
timestamp: ISO 8601
agent: string
status: pass|fail|escalated
phase: preflight|apt|images|restarts|verify|rollback|done|failed
actions_taken: array
postconditions: array
drift_alerts: array
evidence_path: string
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + pause + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 2
window: 7200
trip_action: escalate_to_fatal
depends_on:
- infrastructure-monitoring
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-update
file: infrastructure-update.prose.md
kind: responsibility
category: maintenance
sensitivity: high
status: active
owner: ops
version: 1.1.0
trigger:
type: scheduled
cadence: 0 2 * * 0
description: Weekly system updates Sunday at 2am ET
cron_job_id: null
execution:
agent: abiba
timeout: 3600
requires: []
verification:
postconditions:
- check: all services running after update
verify: systemctl list-units --state=running
expect: all critical services
artifact: update report with changelog
receipt:
format: json
storage: ~/.hermes/runs/infrastructure-update/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 1
window: 7200
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: infrastructure-control
file: infrastructure-control.prose.md
kind: pattern
category: reference
sensitivity: critical
status: active
owner: abiba
version: 1.0.0
trigger:
type: passive
description: Loaded on-demand by executing agents as source of truth
cron_job_id: null
execution:
agent: none
timeout: 0
requires: []
verification:
postconditions: []
artifact: none
receipt:
format: none
storage: none
graph_node: false
escalation:
info:
action: none
notify: []
warning:
action: none
notify: []
critical:
action: none
notify: []
fatal:
action: none
notify: []
circuit_breaker:
max_retries: 0
window: 0
trip_action: none
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: ra-h-os-custodianship-contract
file: ra-h-os-custodianship-contract.prose.md
kind: pattern
category: reference
sensitivity: high
status: active
owner: mumuni
version: 1.0.0
trigger:
type: passive
description: Loaded on-demand by agents needing RA-H OS guidance
cron_job_id: null
execution:
agent: none
timeout: 0
requires: []
verification:
postconditions: []
artifact: none
receipt:
format: none
storage: none
graph_node: false
escalation:
info:
action: none
notify: []
warning:
action: none
notify: []
critical:
action: none
notify: []
fatal:
action: none
notify: []
circuit_breaker:
max_retries: 0
window: 0
trip_action: none
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: mumuni-delegation
file: mumuni-delegation-prose-contract.prose.md
kind: pattern
category: reference
sensitivity: high
status: active
owner: mumuni
version: 1.0.0
trigger:
type: passive
description: Loaded by Mumuni when delegating tasks
cron_job_id: null
execution:
agent: none
timeout: 0
requires: []
verification:
postconditions: []
artifact: none
receipt:
format: none
storage: none
graph_node: false
escalation:
info:
action: none
notify: []
warning:
action: none
notify: []
critical:
action: none
notify: []
fatal:
action: none
notify: []
circuit_breaker:
max_retries: 0
window: 0
trip_action: none
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: zulip-adapter-lessons
file: zulip-adapter-lessons.prose.md
kind: pattern
category: reference
sensitivity: retired
status: historical
owner: abiba
version: 1.0.0
trigger:
type: passive
description: "Historical reference \u2014 pi Zulip extension retired 2026-07-04"
cron_job_id: null
execution:
agent: none
timeout: 0
requires: []
verification:
postconditions: []
artifact: none
receipt:
format: none
storage: none
graph_node: false
escalation:
info:
action: none
notify: []
warning:
action: none
notify: []
critical:
action: none
notify: []
fatal:
action: none
notify: []
circuit_breaker:
max_retries: 0
window: 0
trip_action: none
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: pi-approval-architecture
file: pi-approval-architecture.prose.md
kind: architecture
category: reference
sensitivity: retired
status: historical
owner: abiba
version: 1.0.0
trigger:
type: passive
description: "Historical reference \u2014 pi Zulip extension retired"
cron_job_id: null
execution:
agent: none
timeout: 0
requires: []
verification:
postconditions: []
artifact: none
receipt:
format: none
storage: none
graph_node: false
escalation:
info:
action: none
notify: []
warning:
action: none
notify: []
critical:
action: none
notify: []
fatal:
action: none
notify: []
circuit_breaker:
max_retries: 0
window: 0
trip_action: none
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: zulip-approval-fix
file: zulip-approval-fix.prose.md
kind: responsibility
category: reference
sensitivity: normal
status: active
owner: abiba
version: 2.0.0
trigger:
type: passive
description: Loaded on-demand when /approve or /deny commands fail
cron_job_id: null
execution:
agent: none
timeout: 0
requires: []
verification:
postconditions: []
artifact: none
receipt:
format: none
storage: none
graph_node: false
escalation:
info:
action: none
notify: []
warning:
action: none
notify: []
critical:
action: none
notify: []
fatal:
action: none
notify: []
circuit_breaker:
max_retries: 0
window: 0
trip_action: none
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: zulip-oidc-redirect-fix
file: zulip-oidc-redirect-fix.prose.md
kind: responsibility
category: reference
sensitivity: normal
status: active
owner: abiba
version: 1.0.0
trigger:
type: event_driven
description: Triggered when Zulip OIDC login returns redirect URI error
cron_job_id: null
execution:
agent: abiba
timeout: 300
requires: []
verification:
postconditions:
- check: OIDC redirect uses public domain
verify: grep redirect_uri /opt/zulip/zulip_env.py
expect: chat.sysloggh.net
artifact: fix verification report
receipt:
format: json
storage: ~/.hermes/runs/zulip-oidc-redirect-fix/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
critical:
action: relay_alert
notify:
- abiba
- mumuni
fatal:
action: relay_alert + human_required
notify:
- abiba
- mumuni
- kwame
circuit_breaker:
max_retries: 1
window: 7200
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: zulip-mention-reliability
file: zulip-mention-reliability.prose.md
kind: responsibility
category: retired
sensitivity: retired
status: retired
owner: abiba
version: 1.0.0
trigger:
type: none
description: "Retired 2026-07-04 \u2014 pi Zulip extension decommissioned"
cron_job_id: null
execution:
agent: none
timeout: 0
requires: []
verification:
postconditions: []
artifact: none
receipt:
format: none
storage: none
graph_node: false
escalation:
info:
action: none
notify: []
warning:
action: none
notify: []
critical:
action: none
notify: []
fatal:
action: none
notify: []
circuit_breaker:
max_retries: 0
window: 0
trip_action: none
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: zulip-self-heal
file: zulip-self-heal.prose.md
kind: responsibility
category: retired
sensitivity: retired
status: retired
owner: abiba
version: 1.0.0
trigger:
type: none
description: "Retired 2026-07-04 \u2014 pi Zulip extension decommissioned"
cron_job_id: null
execution:
agent: none
timeout: 0
requires: []
verification:
postconditions: []
artifact: none
receipt:
format: none
storage: none
graph_node: false
escalation:
info:
action: none
notify: []
warning:
action: none
notify: []
critical:
action: none
notify: []
fatal:
action: none
notify: []
circuit_breaker:
max_retries: 0
window: 0
trip_action: none
depends_on: []
last_run: null
last_status: null
drift_alerts: []
- name: hello-world
file: hello-world.prose.md
kind: function
category: reference
sensitivity: normal
status: active
owner: kwame
version: 1.0.0
trigger:
type: on_demand
description: Test contract for OpenProse on pi
cron_job_id: null
execution:
agent: kwame
timeout: 60
requires: []
verification:
postconditions:
- check: greeting returned
verify: contract output contains 'Hello'
expect: pass
artifact: greeting output
receipt:
format: json
storage: ~/.hermes/runs/hello-world/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify: []
critical:
action: relay_alert
notify:
- mumuni
fatal:
action: relay_alert + human_required
notify:
- mumuni
- kwame
circuit_breaker:
max_retries: 3
window: 3600
trip_action: escalate_to_fatal
depends_on: []
last_run: null
last_status: null
drift_alerts: []
# Koby Report-Only Registry (2026-08-17 — Captain)
# ⛔ KOBY IS NEVER REPAIRED — detect + report, never fix on .129
- name: litellm-api-keys
file: litellm-api-keys.prose.md
kind: function
category: compliance
sensitivity: critical
status: active
owner: abiba
version: 1.1.0
trigger:
type: on_demand
cadence: null
description: "Manual invocation when creating/rotating/verifying agent LiteLLM keys"
cron_job_id: null
execution:
agent: abiba
timeout: 120
requires: []
protocol:
- Load contract from prose-contracts/main
- Retrieve master key from Infisical (project=infrastructure env=production)
- Read live key-scoped model roster from CT 116 /v1/models
- Create/rotate/verify the requested agent key with an EXPLICIT models list
- 'Never create a key with an empty models list or all-proxy-models (Cloud leak)'
verification:
postconditions:
- check: standard agent key is local-only
verify: 'curl -s -H "Authorization: Bearer <KEY>" http://192.168.68.116/litellm/v1/models | jq -r ''.data[].id'' | grep -c /'
expect: 0 cloud models
- check: key exists with correct alias
verify: 'curl -s -H "Authorization: Bearer <MASTER>" http://192.168.68.116/litellm/v1/key/info?key_alias=<AGENT>'
expect: 200 with matching alias
artifact: key creation/rotation report
receipt:
format: json
storage: ~/.hermes/runs/litellm-api-keys/
graph_node: true
escalation:
info:
action: log_to_receipt
notify: []
warning:
action: relay_alert
notify:
- abiba
- mumuni
critical:
action: relay_alert
notify:
- abiba
- mumuni
- ops
- name: daily-health-digest
file: daily-health-digest.prose.md
kind: function
category: monitoring
sensitivity: normal
status: active
owner: abiba
version: 1.0.0
trigger:
type: scheduled
cadence: 30 10 * * *
description: Daily at 10:30 UTC, dispatched on CT 100 as a firstmate message
to the ops lane, which executes the pinned producer
cron_job_id: null
execution:
agent: abiba
timeout: 300
requires:
- infisical (vault credentials injected at run time)
protocol:
- 'Execute from the PINNED clone only: /root/abiba-workspace/projects/prose-contracts'
- cd /root/abiba-workspace/projects/prose-contracts
- infisical run --env=prod -- python3 scripts/daily-infra-report.py
- 'Never execute from a per-agent working copy (treehouse) - it drifts onto feature branches'
- 'On failure: do not treat a 0/0 Proxmox section as evidence about the estate - it means could not look'
verification:
postconditions:
- check: every Proxmox probe is reachable
verify: >-
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
| grep -c 'pve_probe_status'
expect: '1'
- check: all nodes reported online
verify: >-
infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
| grep 'nodes_online'
expect: nodes_online == node_count
artifact: timestamped HTML dashboard emailed to jerome@sysloggh.com
verify_commands:
- infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email
- python3 -m pytest tests/test_daily_infra_report.py -q
delivery:
transport: zulip-dm-attachment
recipient_user_id: 9
sender: abiba-bot@chat.sysloggh.net
key_source: abiba-bot Zulip key already on the execution host, read from the
600-mode env file /root/.pi/agent/extensions/zulip/.env
key_policy: do NOT add a vault entry - that is a captain decision under the auth-keys charter
body: short Markdown pointer; the HTML attachment IS the report
artifact: /var/log/daily-infra-report/infra-report-<UTCstamp>.html
note: Replaced SMTP/mail on 2026-09-26 by captain decision. Removes the Google
dependency entirely; closes daily-digest-mail-transport-20260921.
exit_semantics:
'1': missing PVE_TOKEN, unreachable Proxmox probe, missing/rejected Zulip
credential, or a failed upload/post - raises an alert
'0': healthy delivery only - there is no degraded delivery leg any more
depends_on: []
last_run: null
last_status: null
drift_alerts: []
koby_report_only: true
koby_host: "CT 111 (tdunna)"
koby_ip: ".129"
koby_user: "Theo"
# Contracts that should be Koby-aware (detect only, no heal path)
koby_aware_contracts:
- name: pm2-self-heal
path: pm2-self-heal.prose.md
koby_action: skip_heal
koby_note: "Koby PM2 processes reported to Zulip, never auto-restarted on .129"
- name: zulip-health
path: zulip-health.prose.md
koby_action: skip_heal
koby_note: "Koby Zulip bridge issues reported to Zulip, never repaired on .129"
- name: hermes-zulip-restore
path: hermes-zulip-restore.prose.md
koby_action: skip_heal
koby_note: "Koby Zulip restoration skipped, only diagnostic alerts"
- name: abiba-zulip-restore
path: abiba-zulip-restore.prose.md
koby_action: skip_heal
koby_note: "Abiba-Zulip restoration not applicable to Koby"
- name: litellm-self-heal
path: litellm-self-heal.prose.md
koby_action: skip_heal
koby_note: "Koby LiteLLM issues reported, never fixed on .129"
- name: disk-gc-threat-response
path: disk-gc-threat-response.prose.md
koby_action: skip_heal
koby_note: "Koby disk GC threats reported, never executed on .129"
- name: memory-fixer
path: memory-fixer.prose.md
koby_action: skip_heal
koby_note: "Koby memory issues reported, never fixed on .129"
- name: memory-audit-maintenance
path: memory-audit-maintenance.prose.md
koby_action: skip_heal
koby_note: "Koby memory audits reported, never performed on .129"
- name: gpu-self-heal
path: gpu-self-heal.prose.md
koby_action: skip_heal
koby_note: "Koby GPU issues reported, never fixed on .129"
- name: gpu-monitor
path: gpu-monitor.prose.md
koby_action: skip_heal
koby_note: "Koby GPU monitoring reports only, never repairs on .129"
- name: agent-health-check
path: agent-health-check.prose.md
koby_action: skip_heal
koby_note: "Koby agent health checks reported, never repairs on .129"
# Scripts that should skip Koby
koby_aware_scripts:
- name: agent-health-check.py
path: scripts/agent-health-check.py
koby_action: skip_heal
koby_note: "Script should only run diagnostics on Koby, not repairs"