ci.yml has been invalid YAML since introduction: the 'Config validation'
and old inline checks dedented out of their run:| block scalar, so Gitea
could never parse the workflow — CI never ran on any PR despite
CI_STATUS.md claiming 'Active'. The old 'No secrets check' also always
passed (|| echo swallows the grep hit) and never scanned *.yml — where
six embedded credentials were living.
- validation logic moved to ci_check.py (testable locally: python3 ci_check.py all)
- secrets check now FAILS on embedded http-basic URLs and long api_keys,
across .py/.ts/.yaml/.yml/.cjs/.sh, with placeholder allowlist
- added workflow-YAML parse gate so this class of breakage can't recur
- py_compile steps no longer swallow errors with '|| echo skipped'
- removed ci.yml's duplicate deploy job: deploy.yml is the sole deploy
pipeline (rc tags → Tanko canary only; stable → all agents). The ci.yml
copy would have deployed Mumuni on rc tags too, breaking canary policy,
and never ran anyway.
- CI_STATUS.md rewritten with the real state + caveats (history still
contains the old creds — rotation is a server-side task)
Dynamic resolution (ADR-006) overrides this on connect, but when the
/api/v1/users call fails the adapter fell back to 1, silently dropping
every @all-bots mention. The realm's all-bots user is 20 (verified
2026-09-25: 'Resolved @all-bots user_id=20 from all-bots@chat.sysloggh.net';
CONTRACT_VERIFICATION_2026-06-29 fixed the Pi config to 20 for the same
reason). Align the Hermes-side fallback with the verified realm value.
Completes aeb79c6 (main): four more clone steps in deploy.yml still
embedded abiba-bot HTTP Basic credentials in plaintext. Runner already
auto-checkouts the repo, so the manual clone was redundant — replaced
with actions/checkout@v4, same pattern as ci.yml.
No secrets remain in tracked workflow files after this change.
The truncation notice '[...truncated at Zulip limit]' was appended AFTER
slicing at MAX_ZULIP_MESSAGE (10000), causing the final message to exceed
Zulip's API limit. This fix subtracts the notice length from the slice so
the total stays within bounds.
(cherry picked from commit 19c52a9425)
Replaced manual 'git clone' with password in URL with actions/checkout@v4.
Runner already auto-checkouts the repo - manual clone was redundant.
Also fixed YAML syntax issues in Config validation and No secrets check steps.
Credentials were exposed in git history since initial commit.
(cherry picked from commit aeb79c6286)
Root cause: Workers hung because default model was DeepSeek v4-pro (reasoning
model that produces empty content with normal token limits). pi RPC workers
waited forever for content that never arrived.
Changes:
- Circuit breaker (CLOSED→OPEN→HALF_OPEN) around Zulip API calls
- Retry with exponential backoff + jitter for transient errors
- Queue lifecycle management (idle_queue_timeout, auto re-register on BAD_EVENT_QUEUE_ID)
- External supervisor watchdog (restarts router after 3 health check failures)
- Busy worker timeout (SIGKILL after 5 min + error DM)
- PM2 hardening (max_restarts=100, max_memory_restart=500M)
- Crash handlers (uncaughtException + unhandledRejection → reconnect, not die)
- Fixed default model: deepseek-v4-pro → syslog-harness/syslog-auto
- Enhanced health endpoint with circuit breaker stats
Architecture: Research-backed from Zulip event system docs + Node.js resilience
patterns. Inline circuit breaker (no dependency). Separate watchdog process (Hermes pattern).
Verified: Circuit breaker trips on outage, recovers gracefully. End-to-end DM
processed in 4 seconds. Watchdog monitoring every 30s.
Zulip delivers message content as HTML (<p>/approve</p>).
The gateway slash command parser expects plain text, so HTML
tags prevent command matching. This helper strips HTML tags
and decodes common entities (&, <, etc.).
Per contract: zulip-approval-fix.prose.md
Applied to: Mumuni (CT114), Tanko (CT112), Koby (CT111)
- Added Host column (Abiba: 192.168.68.24, Hermes agents: 192.168.68.123)
- Changed Health Port → Health URL with full http://host:port/health URLs
- Updated Check 1 to reference {{health_url}} instead of {{health_port}}
- Added .gitignore for .agents/ (OpenProse run state)
Adds two improvements from pi Zulip extension lessons:
1. _ensure_stream_subscriptions() — checks at connect time if the bot
is subscribed to the primary stream and general-chat; auto-subscribes
if missing. Without this, bots with 0 subscriptions can't receive
stream events (DMs only).
2. selftest check #9 — verifies stream subscriptions and reports
count + names. Catches the 'zero subscriptions' failure mode that
was a major debugging bottleneck in the pi extension.
Same fix as applied to the pi Zulip extension: last_error_msg and
last_error_at are now cleared on every successful poll cycle, not
just on reconnect. Health monitors no longer show stale errors.
- Full contract verification for pi (TypeScript) and Hermes (Python) Zulip plugins
- Fixed @all-bots user_id (1→20) — now dynamically resolved from Zulip API
- Fixed last_event_id stale display in health endpoint (B2)
- Fixed display_recipient logging (B3)
- Added bot_user_id and all_bots_user_id to health endpoint
- Created OpenProse cross-platform verification contract (docs/contracts/)
- Created Hermes plugin deployment script (scripts/deploy-hermes-zulip.py)
Pi extension: 8/9 checks pass (LLM relay untested — no external DMs)
Hermes adapter: 9/10 checks pass (deployment pending)
Bots with no incoming messages were reconnecting every ~60 seconds
(20 polls × 3s interval). Raised to 500 polls (~25min) so idle bots
don't cycle. Connection recovers immediately on BAD_EVENT_QUEUE_ID
regardless of this counter.
Zulip event queues expire after ~600s of inactivity. The adapter now
detects 400/BAD_EVENT_QUEUE_ID responses and auto-reconnects instead of
silently returning empty results.
Complete rewrite of the pi Zulip extension as a standalone Node.js service
with direct harness API calls. No longer depends on pi's session — survives
pi shutdown and restart.
Changes:
- src/index.ts: Complete rewrite — standalone process, not a pi extension
- Direct harness API calls (POST /v1/chat/completions) with conversation memory
- Per-sender conversation history (up to 50 messages)
- Placeholder->edit streaming for Zulip DMs
- Typing indicators via Zulip API
- Health endpoint on :9200
- Exponential backoff reconnection
- Graceful shutdown on SIGINT/SIGTERM
- abiba-zulip.service: systemd service unit file
- README.md: Updated deployment instructions
- package.json/tsconfig.json: Updated for standalone app
Deploy: npm install -> npx tsc -> systemctl enable abiba-zulip
Closes issue #24 (Hermes parity for pi)
A2A server uses Agent Zero's AgentContext.communicate() directly
to process Zulip messages inside the container.
Full chain verified: Zulip DM -> adapter -> A2A -> Agent Zero -> response -> Zulip
Zulip API edits messages via PATCH /api/v1/messages/{message_id}
(POST returns 405 on this endpoint). Also:
- edit_message() now returns SendResult (not bool) for Gateway compatibility
- Clean up duplicate gateway instances before restart
The Gateway streaming interface calls edit_message() with finalize=True
and expects SendResult. Fixed:
1. Added **kwargs to edit_message() signature — silently ignores finalize
2. Changed return type from bool to SendResult for Gateway compatibility
3. Switched from PATCH to POST /api/v1/messages/{id} — Zulip docs use
POST for message editing (PATCH returns 405 Method Not Allowed)
4. Same fixes applied to delete_message() and send()
Discovered during Tanko Gen 4 testing.
Fixes 4 critical reliability issues discovered during Zulip server outage recovery:
1. CLOSE-WAIT socket leak: _reconnect() now creates a fresh httpx.AsyncClient()
when the queue expires, preventing stuck connection pool. Old client is
explicitly aclose()'d.
2. Silent poll loop death: After sustained 502 errors, the poll loop was
returning [] without logging — transport errors are swallowed by _api_call
returning (None, 0). Added heartbeat logging every 5 minutes and silence
detection warnings after 60s of no events.
3. HTTP client stuck pool: Added proactive client recreation after 20 consecutive
empty polls (connection pool exhaustion detection) and periodic refresh every
500 polls (connection reuse limit).
4. Log spam: 502 error responses from Netbird contain full HTML pages — now
truncated to 80 chars with newlines stripped.
Also:
- Support ZULIP_URL env var (in addition to ZULIP_SITE)
- expose silence_seconds, consecutive_empty_polls, client_pool_resets in health stats
- Reduced health callback interval from 600s to 300s