Compare commits

..
Author SHA1 Message Date
root 0a41a2d584 fix(daily-digest): probe Firecrawl on its real liveness path and classify endpoints per fleet policy
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 8s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Folds into the same branch as the Zulip delivery change, as instructed.

FINDING 1 - the Firecrawl probe path was wrong; the service is fine.
  scripts/daily-infra-report.py probed http://192.168.68.7:3002/health, which
  Firecrawl does not serve - it 404s. The root answers 200 with
  {"message":"Firecrawl API",...}. Live 2026-09-26:
    Firecrawl(/)        -> 200
    Firecrawl(/health)  -> 404   <- what the report was showing
  The probe is now the root, which is its liveness endpoint.

FINDING 2 - the Network Endpoints classification was wrong twice over.
  It read: color = green if code in (200,302,401) else (yellow if code >= 400
  else red). Two defects:
   (a) it ignored the fleet's own probe policy, codified 2026-09-14 in the
       monitoring contracts: ANY HTTP status proves the service answered, so the
       service is ALIVE, and only a failed CONNECTION is a failed probe. A 404
       from a wrong path is not a service fault.
   (b)  was a STRING comparison. Reproduced: 301 -> red (a live
       redirect rendered as a failure), 404 -> yellow, 500 -> yellow (a real
       server error softened to a warning).
  Replaced with classify_endpoint(), which returns:
     any 2xx/3xx/4xx -> green  'alive'          (code still shown)
     5xx             -> yellow 'server error'   (kept distinct from 4xx, as asked)
     000/no answer   -> red    'no connection'

  Verified against the live endpoints after the change:
    Gitea 200, Authentik 302, Zulip 302, Pulse 200, Proxmox 200, SearXNG 200,
    Firecrawl 200 - all green/alive; the only red state is a genuine no-connection.

ALSO CHECKED, as asked: scripts/search-stack-check.py does NOT depend on the
wrong route. It POSTs to {FIRECRAWL_URL}/v1/scrape with formats=[markdown], and
that path really works - live POST returned HTTP 200 and 180 chars of markdown
for https://example.com. It was never using /health.

prose-lint: PASSED.
2026-09-26 15:36:38 +00:00
root de32f54337 feat(daily-digest): deliver via Zulip DM as an HTML attachment; drop mail entirely
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 3s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 6s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
Captain's decision 2026-09-26, clarified the same day: the digest is delivered to
his Zulip DM (user id 9) from abiba-bot as an HTML FILE - an attachment, not HTML
rendered in the message body and not a Markdown translation of it. Closes
daily-digest-mail-transport-20260921; the Google dependency is gone (no SMTP, no
EMAIL_PASSWORD, no app password, nothing to rotate).

WHAT CHANGES
* scripts/daily-infra-report.py: send_email() is replaced by send_zulip(), which
  writes the styled dashboard to /var/log/daily-infra-report/infra-report-<ts>.html,
  uploads it via POST /api/v1/user_uploads, then posts a SHORT Markdown pointer to
  user 9. The message body carries subject, top-line status and the attachment
  link; it does not reproduce the report.
* the 10,000-character cap is irrelevant here - it bounds message TEXT only, and
  the report travels as a file, so nothing is shrunk to fit.
* the key is abiba-bot's, already on the execution host at
  /root/.pi/agent/extensions/zulip/.env (mode 600). No vault entry was added:
  under the auth-keys charter that is a captain decision.
* daily-health-digest.prose.md -> v2.0.0 and contract-registry.yaml updated:
  transport, healthy/degraded definitions, and exit codes now match observed
  behaviour. There is NO degraded delivery leg any more - delivery is the only
  output path, so a missing or rejected key is a real failure (exit 1).
* queued defect folded in: a failed delivery used to print only the transport
  error while the report body never surfaced. Now the HTML is printed to stdout
  AND persisted on every failure, and the message names which step failed.

EVIDENCE (all against the live stack)
* real send: message id 86221 to user 9, attachment 16208 bytes at
  /user_uploads/2/45/m1cQesBFV78BGeNY2lN8xkN5/infra-report-20260926-153406.html
* the message is type=private, sender abiba-bot@chat.sysloggh.net, recipients
  [9, 21], body carries the top-line status and the attachment link, and does NOT
  contain a <table> - i.e. it does not reproduce the report
* the attachment fetches HTTP 200, 16208 bytes, content-type text/html, starts
  with <!DOCTYPE html>, and contains <style>, <table> and 16 class="card" blocks -
  it opens as a standalone styled document
* failure path: a bad key gives 'Delivery FAILED at upload: Malformed API key',
  EXIT=1, the HTML is printed to stdout and persisted to disk
* scheduled path: the run's own output is pasted in the PR

prose-lint: PASSED (19 warnings); secret scan clean.
2026-09-26 15:35:36 +00:00
7 changed files with 231 additions and 868 deletions
-173
View File
@@ -1,173 +0,0 @@
# Search ranking policy for the agent-consumption layer.
#
# Everything here is CONFIG, not code, so it is reviewable and changeable without
# touching the module. Read by scripts/search-agent-consume.py.
#
# Why this file exists: multi-engine aggregation returns results with no
# filtering, no dedupe and no reranking. On 2026-09-26 that put a shopping page
# and a dictionary definition into "best practices agent context management",
# and put four SEO blogs ABOVE the actual Proxmox forum threads on a precise
# technical query. Identical queries also ranked differently between runs, which
# is the strongest argument for a deterministic layer rather than hoping the
# engines behave.
version: 1
# ── Non-answers: dropped outright, never returned ────────────────────────────
# These are pages that cannot answer a question: navigational homepages,
# shopping/product pages, dictionary definitions, and login walls.
non_answer:
# URL path is empty -> it is a site's front door, not an answer. Still allowed
# when the host is explicitly preferred (see prefer_domains), because some
# docs/repo front doors ARE the answer.
host_root: true
path_patterns:
- '/dictionary/'
- '/dictionary?'
- '/wiki/Wiktionary:'
- '/search?'
- '/cart'
- '/checkout'
- '/login'
- '/signin'
- '/sign-in'
- '/account/login'
- '/shop/'
- '/store/'
- '/dp/' # Amazon-style product URL
- '/gp/product/'
- '/add-to-cart'
- '/checkout'
# NOTE: '/products/' and '/product/' were REMOVED as path patterns. They fired
# on docs.digitalocean.com/products/inference/... — a legitimate documentation
# page — which the 2026-09-26 before/after run caught. Shopping is caught by
# the shopping HOST list instead, which does not have that false positive.
# Query strings that betray a search/shopping surface rather than an article.
query_keys:
- 'q'
- 'query'
- 's'
- 'search'
- 'add-to-cart'
# Hosts that are shopping/retail and never answer a technical question.
hosts:
- bestbuy.com
- amazon.com
- ebay.com
- walmart.com
- etsy.com
- aliexpress.com
- merriam-webster.com
- dictionary.com
- thesaurus.com
- vocabulary.com
- collinsdictionary.com
# ── Demotion: ranked below everything else, never dropped ────────────────────
# Low-authority content farms / SEO aggregators. Demoted rather than dropped so
# a genuinely useful hit is not lost, but it can never outrank a primary source.
# Reviewable: add or remove hosts here, no code change required.
demote_domains:
- medium.com
- sparkco.ai
- mindstudio.ai
- aitechmonk.com
- stackai.com
- agentic-design.ai
- voxfor.com
- bigiron.cc
- linuxoperatingsystem.net
- riparazioneserver.com
- rossmanngroup.com
- dev.to
- hashnode.dev
- substack.com
- towardsdatascience.com
- analyticsvidhya.com
- geeksforgeeks.org
- tutorialspoint.com
- javatpoint.com
- w3schools.com
- scaler.com
- simplilearn.com
- udemy.com
- coursera.org
# ── Preference: promoted above the default rank ──────────────────────────────
# Primary sources: upstream repositories, official docs, Q&A, vendor
# engineering blogs. These are what an agent should be reading.
prefer_domains:
# upstream repositories and code hosting
- github.com
- gitlab.com
- codeberg.org
- sourceforge.net
- kernel.org
- git.kernel.org
# Q&A
- stackoverflow.com
- stackexchange.com
- superuser.com
- serverfault.com
- askubuntu.com
- discourse.org
# vendor / project documentation and forums
- proxmox.com
- forum.proxmox.com
- pve.proxmox.com
- docs.python.org
- developer.mozilla.org
- kernelnewbies.org
- man7.org
- gnu.org
- debian.org
- ubuntu.com
- redhat.com
- kernel.dk # io_uring / Jens Axboe
- github.io # project pages (docs, papers) — promoted, not authoritative by itself
# vendor engineering blogs
- anthropic.com
- openai.com
- googleblog.com
- developers.googleblog.com
- engineering.fb.com
- netflixtechblog.com
- aws.amazon.com
- cloud.google.com
- microsoft.com
- learn.microsoft.com
- apple.com
- nvidia.com
- intel.com
- amd.com
- redislabs.com
- cloudflare.com
- langchain.com
- jetbrains.com
- cursor.com
# community discussion with high signal
- news.ycombinator.com
- lobste.rs
- reddit.com
# ── Ranking weights ──────────────────────────────────────────────────────────
# Final score = engine_score - demote_penalty + prefer_bonus, then a stable
# tiebreak on original position so ordering is reproducible run to run.
ranking:
demote_penalty: 1000
prefer_bonus: 100
# Results that several engines independently returned are more likely real.
multi_engine_bonus: 25
# Shallow paths (e.g. /blog/x) are slightly less likely to be primary docs.
host_root_allowed_when_preferred: true
# ── Extraction budget (criterion 4) ──────────────────────────────────────────
# Return CONTENT, not just links, so an agent gets usable material in ONE call.
extraction:
top_n: 5 # how many results get page text extracted
total_chars: 12000 # global budget across all extracted items
per_item_chars: 4000 # cap for any single item, so one page cannot eat the budget
timeout_seconds: 45 # per scrape
# If extraction fails, the result is still returned with an empty excerpt —
# a link is better than nothing, but the failure is recorded in the output.
on_failure: keep_with_empty_excerpt
+14 -10
View File
@@ -1965,17 +1965,21 @@ contracts:
verify_commands:
- infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email
- python3 -m pytest tests/test_daily_infra_report.py -q
email_dependency:
transport: smtp.gmail.com:587
identity: jtabiri@gmail.com
secret: EMAIL_PASSWORD (must be a Google app password)
status: DEGRADED as of 2026-09-25 - 534 5.7.9 Application-specific password required
note: A delivery failure is a credential dependency, not a code defect. Tracked
as daily-digest-mail-transport-20260921.
delivery:
transport: zulip-dm-attachment
recipient_user_id: 9
sender: abiba-bot@chat.sysloggh.net
key_source: abiba-bot Zulip key already on the execution host, read from the
600-mode env file /root/.pi/agent/extensions/zulip/.env
key_policy: do NOT add a vault entry - that is a captain decision under the auth-keys charter
body: short Markdown pointer; the HTML attachment IS the report
artifact: /var/log/daily-infra-report/infra-report-<UTCstamp>.html
note: Replaced SMTP/mail on 2026-09-26 by captain decision. Removes the Google
dependency entirely; closes daily-digest-mail-transport-20260921.
exit_semantics:
'1': missing PVE_TOKEN, unreachable Proxmox probe, or failed email send - raises an alert
'0': healthy, or a deliberate DEGRADED leg where the email credential is absent
and the report is still produced
'1': missing PVE_TOKEN, unreachable Proxmox probe, missing/rejected Zulip
credential, or a failed upload/post - raises an alert
'0': healthy delivery only - there is no degraded delivery leg any more
depends_on: []
last_run: null
last_status: null
+48 -32
View File
@@ -16,11 +16,12 @@ description: >
Exit-code semantics (as they actually behave, verified 2026-09-25):
* missing PVE_TOKEN, or an unreachable Proxmox probe -> exit 1 + alert
* missing EMAIL credential -> deliberate DEGRADED leg, exit 0, report still
produced
* email send failure -> exit 1 (a delivery fault, not a code defect)
* missing or rejected Zulip credential -> exit 1 (delivery is the only
output path, so it is a real failure, not a degraded leg)
* delivery failure -> exit 1, and the report body is printed AND persisted
so the content is never swallowed
version: 1.0.0
version: 2.0.0
---
## Purpose
@@ -93,14 +94,13 @@ $ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json
EXIT=0
```
and in mail mode:
and in delivery mode:
```
Sending email...
✅ All legs fully credentialed
📋 Summary:
Proxmox: 5/5 nodes online
VMs/CTs: 22/22 running
report ready: 16208 chars of HTML (delivered as a file attachment)
Sending to the captain's Zulip DM...
✅ Delivered to Zulip DM (user 9), message id 86221, attachment 16208 bytes
at /user_uploads/2/45/m1cQesBFV78BGeNY2lN8xkN5/infra-report-20260926-153406.html
```
Healthy means: every probe reports `ok`, `nodes_online == node_count`, and the
@@ -115,9 +115,8 @@ Verified on 2026-09-25 by running each case deliberately.
| all probes reachable, email sent | 0 | — | healthy |
| **missing `PVE_TOKEN`** | **1** | yes | `PROBE FAILURES: proxmox: node list unreachable (PVE_TOKEN missing or API down)`, and `cluster resources unreachable` |
| **Proxmox probe unreachable** | **1** | yes | same path as above; `pve_probe_status: unreachable` |
| **missing `EMAIL_PASSWORD`** | **0** | no | deliberate **DEGRADED** leg (`credential-missing: EMAIL_PASSWORD`); the report is still produced |
| **email send fails** | **1** | yes | e.g. Gmail `534 5.7.9 Application-specific password required` |
| degraded legs present (non-email) | 0 | no | logged under `⚠️ Degraded legs` |
| **missing/rejected Zulip credential** | **1** | yes | delivery is the only output path; report printed and persisted |
| **upload or message post fails** | **1** | yes | report printed and persisted; message names which step failed |
The distinction is deliberate and must not be flattened:
@@ -130,22 +129,31 @@ The distinction is deliberate and must not be flattened:
`PROBE_FAILURES` and `DEGRADED_LEGS` are separate lists for exactly this
reason. Do not merge them.
## Email-delivery dependency
## Delivery: Zulip DM carrying the report as an HTML ATTACHMENT
Delivery is a **credential dependency, not a code path**. The producer
authenticates to `smtp.gmail.com:587` as `jtabiri@gmail.com` with
`EMAIL_PASSWORD` from the vault and sends to `jerome@sysloggh.com`.
Captain's decision 2026-09-26, clarified the same day: the digest is delivered to
his **Zulip DM (user id 9)** from `abiba-bot@chat.sysloggh.net`, as an **HTML
FILE** — an attachment, not HTML rendered in the message body and not a Markdown
translation of it.
* Since that Google account has two-step verification, `EMAIL_PASSWORD` must be
a Google **app password**, not the account password.
* As of 2026-09-25 delivery is **failing** with
`534 5.7.9 Application-specific password required`; the fix is for the
captain to generate a fresh app password and place it in Infisical
(`infrastructure/production`) as `EMAIL_PASSWORD`.
* **A delivery failure is not a code defect.** Investigation of a failed send
should start at the credential, not the script. Chasing it as a code bug
wastes the effort; verify the credential path first with `--test-email`.
* Tracked separately as `daily-digest-mail-transport-20260921`.
* the styled dashboard is built exactly as before and written to
`/var/log/daily-infra-report/infra-report-<UTCstamp>.html`;
* it is uploaded through `POST /api/v1/user_uploads`;
* the **message body stays short Markdown** — subject line, top-line status
(nodes online, guests running, any degraded legs), and a link to the
attachment. The attachment IS the report; the body does not reproduce it.
This removes the Google dependency entirely: **no SMTP, no `EMAIL_PASSWORD`, no
app password, nothing to rotate.** `daily-digest-mail-transport-20260921` is
closed under this option.
The **10,000-character message cap does not apply** — it bounds message TEXT
only, and the report travels as a file. Do not shrink the report to fit it.
The credential is abiba-bot's Zulip key already on the execution host at
`/root/.pi/agent/extensions/zulip/.env` (`ABIBA_ZULIP_API_KEY`, mode 600,
root-readable). **Do not place a new credential in the vault** — under the
auth-keys charter that is a captain decision.
## What counts as a failure
@@ -153,10 +161,18 @@ A run FAILS (exit 1) when the report cannot be trusted or delivered:
* any probe is unreachable, so a section would silently be empty;
* `PVE_TOKEN` is missing;
* the email send fails.
* the Zulip credential is missing or rejected, or the upload/post fails.
A run is DEGRADED (exit 0, report still produced) when a non-load-bearing
credential is absent, currently only `EMAIL_PASSWORD`.
There is **no degraded delivery leg any more**. Delivery is the only output
path, so a missing credential is a failure rather than a survivable degradation —
the previous "missing `EMAIL_PASSWORD` still exits 0" rule is retired with the
mail transport.
**A delivery failure must never swallow the report.** On failure the script
prints the report body to stdout *and* leaves the HTML artifact on disk, so the
content is always recoverable from the run log. That closes the queued defect
where a failed send printed only the transport error and the report never
surfaced.
## Failure behaviour
@@ -175,7 +191,7 @@ infisical run --env=prod -- python3 scripts/daily-infra-report.py --json \
| grep -E 'pve_probe_status|node_count|nodes_online'
# delivery path
infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-email
infisical run --env=prod -- python3 scripts/daily-infra-report.py --test-zulip
```
Regression tests: `tests/test_daily_infra_report.py` (7 tests). Four of them
@@ -183,5 +199,5 @@ fail against the pre-fix script, which is what makes them bite.
## Maintains
- daily-infra-dashboard: { status: "degraded", reason: "email credential", last_check: timestamp }
- daily-infra-dashboard: { status: "ok|undelivered", transport: zulip-dm-attachment, last_check: timestamp }
- pve-probe: { status: "ok|unreachable", last_check: timestamp }
+169 -38
View File
@@ -243,7 +243,7 @@ def collect():
("Pulse", "https://pulse.sysloggh.net"),
("Proxmox", "https://192.168.68.12:8006"),
("SearXNG", "http://192.168.68.7:8888"),
("Firecrawl", "http://192.168.68.7:3002/health"),
("Firecrawl", "http://192.168.68.7:3002/"), # Firecrawl serves no /health - the root is its liveness endpoint
]
report["endpoints"] = []
for name, url in endpoints:
@@ -389,6 +389,28 @@ def collect():
# ── HTML Dashboard ──
def classify_endpoint(code):
"""Classify an endpoint probe per the fleet's probe policy.
Codified 2026-09-14 in the monitoring contracts: ANY HTTP status proves the
service answered, so the service is ALIVE - 200/301/302/401/403/404 alike.
Only a failed CONNECTION (000 / timeout / refused) is a failed probe. A 404
from a wrong path is not a service fault and must not render as one.
This replaces a string comparison that was wrong in both directions
(`ep["code"] >= "400"`): it rendered 301 as red, 404 as yellow, and a real
500 as yellow. 5xx is kept as its own "server error" signal rather than
being merged with 4xx.
"""
if not code or code == "000":
return "red", "no connection"
if code.startswith("5"):
return "yellow", "server error"
if code.startswith(("2", "3", "4")):
return "green", "alive"
return "yellow", f"unexpected {code}"
def build_html(r):
issues = []
@@ -624,7 +646,7 @@ Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']
# ── Network Endpoints ──
html += '<div class="card"><h2>🌐 Network Endpoints</h2><table><tr><th>Service</th><th>Status</th></tr>'
for ep in r["endpoints"]:
color = "green" if ep["code"] in ("200","302","401") else ("yellow" if ep["code"] >= "400" else "red")
color = classify_endpoint(ep["code"])[0]
html += f'<tr><td>{ep["name"]}</td><td class="{color}">HTTP {ep["code"]}</td></tr>'
html += '</table></div>'
@@ -688,42 +710,152 @@ Proxmox: {r.get('pve_probe_status', 'ok')} ({r['nodes_online']}/{r['node_count']
return html
# ── Send Email ──
# ── Delivery: Zulip DM carrying the report as an HTML ATTACHMENT ──
#
# Captain's decision, clarified 2026-09-26: the report is sent as an HTML FILE,
# i.e. an attachment - NOT HTML rendered in the message body, and NOT a Markdown
# translation of it. So the styled dashboard is built exactly as before, uploaded
# through Zulip's file-upload API, and the message body stays short: subject,
# top-line status, and a pointer to the attachment.
#
# This removes the Google dependency entirely (no SMTP, no EMAIL_PASSWORD).
# The 10,000-character message cap does not apply: it bounds message TEXT only,
# and the report travels as a file.
def send_email(html_content, subject_prefix=""):
FROM = "abiba@sysloggh.com"
TO = "jerome@sysloggh.com"
SUBJECT = f"{subject_prefix}{'🏗️ Infrastructure Report — ' + DATE_STR}"
msg = MIMEMultipart("alternative")
msg["From"] = FROM
msg["To"] = TO
msg["Subject"] = SUBJECT
msg.attach(MIMEText("Infrastructure report in HTML format — enable images to view.", "plain"))
msg.attach(MIMEText(html_content, "html"))
ZULIP_SITE = "https://chat.sysloggh.net"
ZULIP_BOT_EMAIL = "abiba-bot@chat.sysloggh.net"
CAPTAIN_USER_ID = 9
ZULIP_KEY_FILE = "/root/.pi/agent/extensions/zulip/.env"
REPORT_ARTIFACT_DIR = "/var/log/daily-infra-report"
def zulip_key():
"""abiba-bot's Zulip key, from the env or the on-host 600 file."""
key = os.environ.get("ABIBA_ZULIP_API_KEY")
if key:
return key.strip()
try:
EMAIL_PASSWORD = os.environ.get("EMAIL_PASSWORD") or os.environ.get("SMTP_PASSWORD") or os.environ.get("MAIL_PASSWORD")
if not EMAIL_PASSWORD:
print(" ⚠️ Degraded leg: credential-missing: EMAIL_PASSWORD (or SMTP_PASSWORD/MAIL_PASSWORD)", file=sys.stderr)
DEGRADED_LEGS.append("credential-missing: EMAIL_PASSWORD")
return True, "✅ Email leg degraded (no credential) — report still produced"
GMAIL_EMAIL = "jtabiri@gmail.com"
server = smtplib.SMTP("smtp.gmail.com", 587)
server.starttls()
server.login(GMAIL_EMAIL, EMAIL_PASSWORD)
server.sendmail(FROM, [TO], msg.as_string())
server.quit()
return True, "✅ Email sent to jerome@sysloggh.com"
except Exception as e:
return False, f"❌ Email failed: {e}"
with open(ZULIP_KEY_FILE) as fh:
for line in fh:
if line.startswith("ABIBA_ZULIP_API_KEY="):
return line.split("=", 1)[1].strip()
except OSError:
return None
return None
def build_summary(r, filename, test=False):
"""Short Markdown body: subject, top-line status, pointer to the attachment.
Deliberately NOT a reproduction of the report - the attachment is the report.
"""
nodes = f"{r.get('nodes_online', 0)}/{r.get('node_count', 0)} nodes online"
guests = f"{r.get('running_vms', 0)}/{r.get('total_vms', 0)} guests running"
lines = [
("\U0001F9EA **TEST — **" if test else "") + "\U0001F3D7\uFE0F **Infrastructure Report — " + DATE_STR + "**",
f"**{nodes}** \u00b7 **{guests}** \u00b7 generated {TIME_STR}",
]
problems = []
if r.get("pve_probe_status") != "ok":
problems.append(f"\u274c Proxmox probe: {r.get('pve_probe_status')}")
if r.get("resources_probe_status") != "ok":
problems.append(f"\u274c Resources probe: {r.get('resources_probe_status')}")
lit = r.get("litellm", {}) or {}
checks = lit.get("checks", []) or []
if checks:
passed = sum(1 for c in checks if c.get("status") == "pass")
if passed != len(checks):
problems.append(f"\u274c LiteLLM: {passed}/{len(checks)} checks pass")
if not (r.get("zulip_ext", {}) or {}).get("connected"):
problems.append("\u274c Zulip extension: not connected")
for leg in DEGRADED_LEGS:
problems.append(f"\u26a0\uFE0F degraded: {leg}")
lines.append("\n".join(problems) if problems else "\u2705 All monitored services healthy")
lines.append(f"\U0001F4CE **Full report attached:** `{filename}`")
return "\n\n".join(lines)
def _curl(args, timeout=60):
r = subprocess.run(["curl", "-s", "-m", str(timeout)] + args,
capture_output=True, text=True)
try:
return json.loads(r.stdout or "{}"), r.stdout
except json.JSONDecodeError:
return {}, r.stdout
def _curl_json(args, timeout=90):
r = subprocess.run(["curl", "-s", "-m", str(timeout)] + args,
capture_output=True, text=True)
try:
return json.loads(r.stdout or "{}"), r.stdout
except json.JSONDecodeError:
return {}, r.stdout
def send_zulip(html_content, report, test=False):
"""Upload the styled HTML and post a short pointer to the captain's DM.
Returns (ok, message). On ANY failure the report body is also printed to
stdout and persisted to disk, so a delivery failure can never swallow the
content - the defect this folds in.
"""
os.makedirs(REPORT_ARTIFACT_DIR, exist_ok=True)
stamp = NOW.strftime("%Y%m%d-%H%M%S")
filename = f"infra-report-{stamp}.html"
html_path = os.path.join(REPORT_ARTIFACT_DIR, filename)
try:
with open(html_path, "w") as fh:
fh.write(html_content)
except OSError as e:
print(f" \u26a0\uFE0F could not persist report artifact: {e}", file=sys.stderr)
key = zulip_key()
if not key:
print(html_content) # never swallow the content
return False, ("\u274c Delivery FAILED: no Zulip credential "
"(ABIBA_ZULIP_API_KEY unset and "
f"{ZULIP_KEY_FILE} unreadable). Report persisted to {html_path}")
auth = ["-u", f"{ZULIP_BOT_EMAIL}:{key}"]
# 1. Upload the report as a file.
up, up_raw = _curl_json(auth + [
"-X", "POST", f"{ZULIP_SITE}/api/v1/user_uploads",
"-F", f"file=@{html_path};type=text/html",
])
if up.get("result") != "success" or not up.get("uri"):
print(html_content)
return False, (f"\u274c Delivery FAILED at upload: {up.get('msg') or up_raw[:160]} "
f"(report persisted to {html_path})")
uri = up["uri"]
size = os.path.getsize(html_path)
# 2. Post a short message pointing at it.
body = build_summary(report, filename, test=test)
link = f"[{filename}]({uri})"
body = body.replace(f"`{filename}`", link)
payload, raw = _curl_json(auth + [
"-X", "POST", f"{ZULIP_SITE}/api/v1/messages",
"-d", "type=private",
"-d", f"to=[{CAPTAIN_USER_ID}]",
"--data-urlencode", f"content={body}",
])
if payload.get("result") == "success":
return True, (f"\u2705 Delivered to Zulip DM (user {CAPTAIN_USER_ID}), "
f"message id {payload.get('id')}, attachment {size} bytes at {uri}")
print(html_content)
return False, (f"\u274c Delivery FAILED at message post: {payload.get('msg') or raw[:160]} "
f"(uploaded {uri}; report persisted to {html_path})")
# ── Main ──
if __name__ == "__main__":
is_test = "--test-email" in sys.argv
is_test = ("--test-email" in sys.argv) or ("--test-zulip" in sys.argv)
print(f"{'🧪 TEST MODE' if is_test else '📊'} Collecting infrastructure data...")
report = collect()
@@ -738,15 +870,14 @@ if __name__ == "__main__":
print(" Building dashboard...")
html = build_html(report)
print(f" report ready: {len(html)} chars of HTML (delivered as a file attachment)")
if is_test:
prefix = "🧪 TEST — "
print(" Sending test email...")
print(" Sending TEST message to the captain's Zulip DM...")
else:
prefix = ""
print(" Sending email...")
ok, msg = send_email(html, subject_prefix=prefix)
print(" Sending to the captain's Zulip DM...")
ok, msg = send_zulip(html, report, test=is_test)
print(f" {msg}")
# Show summary
-395
View File
@@ -1,395 +0,0 @@
#!/usr/bin/env python3
"""Agent-consumption layer in front of SearXNG + Firecrawl.
Multi-engine aggregation returns results with no dedupe, no filtering and no
reranking. Measured 2026-09-26 that put bestbuy.com and merriam-webster.com into
"best practices agent context management", and put four SEO blogs ABOVE the real
Proxmox forum threads on a precise technical query. Identical queries also ranked
differently between runs, so the fix has to be deterministic rather than
dependent on engine mood.
This module turns the raw result list into something an agent can actually use:
1. DEDUPE the same page arriving from several engines
2. DROP clear non-answers (homepages, shopping, dictionaries, logins)
3. DEMOTE config-listed low-authority hosts; PROMOTE primary sources
4. STABLE SORT so ordering is reproducible run to run
5. EXTRACT page text for the top N under an explicit character budget,
so one call returns usable material instead of a snippet
6. EMIT stable JSON with engine provenance
Policy lives in config/search-ranking.yaml, not in this file.
Usage:
search-agent-consume.py "query text" # JSON to stdout
search-agent-consume.py --no-extract "query" # ranking only, no Firecrawl
search-agent-consume.py --explain "query" # include drop/demote reasons
Exit: 0 ok, 1 no results survived filtering, 2 the layer could not run.
"""
from __future__ import annotations
import json
import os
import sys
import time
import urllib.parse
import urllib.request
from pathlib import Path
SEARXNG_URL = os.environ.get("SEARXNG_URL", "http://192.168.68.7:8888").rstrip("/")
FIRECRAWL_URL = os.environ.get("FIRECRAWL_URL", "http://192.168.68.7:3002").rstrip("/")
CONFIG_PATH = os.environ.get(
"SEARCH_RANKING_CONFIG",
str(Path(__file__).resolve().parent.parent / "config" / "search-ranking.yaml"),
)
HTTP_TIMEOUT = float(os.environ.get("SEARCH_CONSUME_TIMEOUT", "25"))
def _load_config() -> dict:
"""Load the ranking policy.
PyYAML is used when present; otherwise a tiny built-in parser handles the
flat lists in this specific file, so the layer never hard-fails on a host
without PyYAML.
"""
text = Path(CONFIG_PATH).read_text()
try:
import yaml # type: ignore
return yaml.safe_load(text)
except ImportError:
return _parse_flat_yaml(text)
def _parse_flat_yaml(text: str) -> dict:
"""Minimal fallback parser: top-level keys, nested one level, flat lists."""
import re
out: dict = {}
stack: list[tuple[int, dict]] = [(-1, out)]
section: dict | None = None
for raw in text.splitlines():
line = raw.split("#", 1)[0].rstrip()
if not line.strip():
continue
indent = len(line) - len(line.lstrip())
body = line.strip()
if body.startswith("- "):
if section is not None:
section.setdefault("_list", []).append(
body[2:].strip().strip("'\"")
)
continue
if ":" in body:
key, _, val = body.partition(":")
key, val = key.strip(), val.strip()
if val:
# write to the INNERMOST open section, not the document root
stack[-1][1][key] = _scalar(val)
section = None
else:
while stack and indent <= stack[-1][0]:
stack.pop()
parent = stack[-1][1]
new: dict = {}
parent[key] = new
stack.append((indent, new))
section = new
# flatten "_list" holders back into their parent as plain lists
def fix(node):
if isinstance(node, dict):
if set(node.keys()) == {"_list"}:
return node["_list"]
return {k: fix(v) for k, v in node.items()}
return node
return fix(out)
def _scalar(v: str):
if v.lower() in ("true", "false"):
return v.lower() == "true"
try:
return int(v)
except ValueError:
pass
try:
return float(v)
except ValueError:
pass
return v.strip("'\"")
# ── filtering ────────────────────────────────────────────────────────────────
def _host(url: str) -> str:
return (urllib.parse.urlparse(url).netloc or "").lower().split(":")[0]
def _registrable(host: str) -> str:
"""Best-effort registrable domain so sub.forum.proxmox.com matches proxmox.com."""
parts = host.split(".")
if len(parts) <= 2:
return host
# handle common two-label public suffixes
two = ".".join(parts[-2:])
if parts[-2] in ("co", "com", "org", "net", "ac", "gov") and len(parts) >= 3:
return ".".join(parts[-3:])
return two
def _host_in(host: str, domains) -> bool:
if not domains:
return False
reg = _registrable(host)
for d in domains:
d = str(d).lower()
if host == d or host.endswith("." + d) or reg == d:
return True
return False
def _normalise_url(url: str) -> str:
"""Strip tracking params and fragments so the same page dedupes."""
p = urllib.parse.urlparse(url)
q = [
(k, v)
for k, v in urllib.parse.parse_qsl(p.query, keep_blank_values=True)
if not k.lower().startswith(("utm_", "fbclid", "gclid", "mc_", "ref"))
]
path = p.path.rstrip("/") or "/"
return urllib.parse.urlunparse(
(p.scheme.lower(), p.netloc.lower(), path, "", urllib.parse.urlencode(q), "")
)
def non_answer_reason(result: dict, cfg: dict) -> str | None:
"""Return why this result is a non-answer, or None if it may be returned."""
na = cfg.get("non_answer", {}) or {}
url = result.get("url", "")
p = urllib.parse.urlparse(url)
host = _host(url)
path = p.path or ""
if _host_in(host, na.get("hosts")):
return "shopping_or_dictionary_host"
if na.get("host_root", True) and path in ("", "/"):
# A preferred host's front door may legitimately be the answer
# (a repo, a docs site). Everything else is navigational.
if not _host_in(host, cfg.get("prefer_domains")):
return "navigational_host_root"
low = url.lower()
for pat in na.get("path_patterns", []) or []:
if pat.lower() in low:
return f"path_pattern:{pat}"
qkeys = {k.lower() for k in (na.get("query_keys") or [])}
if qkeys & {k.lower() for k, _ in urllib.parse.parse_qsl(p.query)}:
return "search_or_shopping_query"
return None
def source_type(url: str, cfg: dict) -> str:
host = _host(url)
if _host_in(host, ["github.com", "gitlab.com", "codeberg.org", "sourceforge.net"]):
return "code"
if _host_in(host, ["stackoverflow.com", "stackexchange.com", "superuser.com",
"serverfault.com", "askubuntu.com"]):
return "qa"
if _host_in(host, ["forum.proxmox.com", "forum.", "discourse"]) or "forum." in host:
return "forum"
if _host_in(host, ["news.ycombinator.com", "lobste.rs", "reddit.com"]):
return "discussion"
if _host_in(host, cfg.get("prefer_domains")):
return "official"
if _host_in(host, cfg.get("demote_domains")):
return "content-farm"
return "web"
def rank(results: list[dict], cfg: dict) -> tuple[list[dict], list[dict]]:
"""Dedupe, drop non-answers, demote/ promote, stable sort.
Returns (kept, dropped) where dropped carries the reason, because a filter
nobody can audit is a filter nobody should trust.
"""
rank_cfg = cfg.get("ranking", {}) or {}
demote_pen = float(rank_cfg.get("demote_penalty", 1000))
prefer_bonus = float(rank_cfg.get("prefer_bonus", 100))
multi_bonus = float(rank_cfg.get("multi_engine_bonus", 25))
seen: dict[str, dict] = {}
dropped: list[dict] = []
for pos, r in enumerate(results):
url = r.get("url")
if not url:
continue
key = _normalise_url(url)
engine = r.get("engine", "?")
# 1. dedupe: same normalised URL from several engines
if key in seen:
seen[key].setdefault("engines", []).append(engine)
seen[key]["duplicate_of"] = True
continue
reason = non_answer_reason(r, cfg)
if reason:
dropped.append({"url": url, "reason": reason, "position": pos + 1})
continue
seen[key] = {
"title": (r.get("title") or "").strip(),
"url": url,
"engines": [engine],
"position": pos,
"score": 0.0,
}
kept = []
for item in seen.values():
host = _host(item["url"])
score = -float(item["position"]) # original order is the base signal
if _host_in(host, cfg.get("demote_domains")):
score -= demote_pen
if _host_in(host, cfg.get("prefer_domains")):
score += prefer_bonus
if len(item["engines"]) > 1:
score += multi_bonus * (len(item["engines"]) - 1)
item["score"] = round(score, 2)
item["host"] = host
item["source_type"] = source_type(item["url"], cfg)
kept.append(item)
# stable: score desc, then original position asc => reproducible run to run
kept.sort(key=lambda i: (-i["score"], i["position"]))
return kept, dropped
# ── extraction ───────────────────────────────────────────────────────────────
def _post_json(url: str, payload: dict, timeout: float) -> dict:
req = urllib.request.Request(
url,
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=timeout) as resp:
return json.loads(resp.read().decode("utf-8", "replace"))
def extract(items: list[dict], cfg: dict) -> dict:
"""Fetch page text for the top N under a global character budget."""
ex = cfg.get("extraction", {}) or {}
top_n = int(ex.get("top_n", 5))
total_budget = int(ex.get("total_chars", 12000))
per_item = int(ex.get("per_item_chars", 4000))
timeout = float(ex.get("timeout_seconds", 45))
used = 0
failures = 0
t0 = time.time()
for item in items[:top_n]:
remaining = total_budget - used
if remaining <= 200:
item["excerpt"] = ""
item["extraction"] = "skipped_budget_exhausted"
continue
cap = min(per_item, remaining)
try:
data = _post_json(
f"{FIRECRAWL_URL}/v1/scrape",
{"url": item["url"], "formats": ["markdown"]},
timeout,
)
md = ((data.get("data") or {}).get("markdown") or "").strip()
if not md:
item["excerpt"] = ""
item["extraction"] = "empty"
failures += 1
continue
item["excerpt"] = md[:cap]
item["extraction"] = "ok" if len(md) <= cap else "truncated"
used += len(item["excerpt"])
except Exception as exc: # noqa: BLE001
item["excerpt"] = ""
item["extraction"] = f"failed:{type(exc).__name__}"
failures += 1
return {
"extracted": min(top_n, len(items)),
"chars_used": used,
"budget": total_budget,
"failures": failures,
"seconds": round(time.time() - t0, 2),
}
# ── entry point ──────────────────────────────────────────────────────────────
def consume(query: str, do_extract: bool = True, explain: bool = False) -> dict:
cfg = _load_config()
url = f"{SEARXNG_URL}/search?" + urllib.parse.urlencode(
{"q": query, "format": "json"}
)
with urllib.request.urlopen(url, timeout=HTTP_TIMEOUT) as resp:
raw = json.loads(resp.read().decode("utf-8", "replace"))
results = raw.get("results", [])
kept, dropped = rank(results, cfg)
extraction = extract(kept, cfg) if do_extract else None
out = {
"query": query,
"raw_result_count": len(results),
"returned_count": len(kept),
"dropped_count": len(dropped),
"engines": sorted({r.get("engine", "?") for r in results}),
"results": [
{
"rank": i + 1,
"title": it["title"],
"url": it["url"],
"host": it["host"],
"source_type": it["source_type"],
"engines": sorted(set(it["engines"])),
"score": it["score"],
"excerpt": it.get("excerpt", ""),
"extraction": it.get("extraction", "not_attempted"),
}
for i, it in enumerate(kept)
],
"extraction": extraction,
}
if explain:
out["dropped"] = dropped
return out
def main() -> int:
args = [a for a in sys.argv[1:] if not a.startswith("--")]
do_extract = "--no-extract" not in sys.argv
explain = "--explain" in sys.argv
if not args:
print(__doc__)
return 2
query = " ".join(args)
try:
out = consume(query, do_extract=do_extract, explain=explain)
except Exception as exc: # noqa: BLE001
print(f"LAYER FAILED: {type(exc).__name__}: {exc}", file=sys.stderr)
return 2
print(json.dumps(out, indent=2))
return 0 if out["returned_count"] else 1
if __name__ == "__main__":
sys.exit(main())
-83
View File
@@ -114,79 +114,6 @@ def unresponsive_names(pairs: list) -> dict[str, str]:
return out
# ── QUALITY GUARD (search-agent-consumption) ─────────────────────────────────
# The agent-consumption layer applies a deterministic demote/drop policy. Without
# an assertion here it could silently rot back to raw engine ordering - the same
# way the endpoint colours silently rotted before 2026-09-26.
QUALITY_QUERIES = [
"best practices agent context management",
"proxmox thin pool metadata exhaustion recovery",
]
# A demoted (content-farm) host must never occupy the top 3 for these queries.
QUALITY_TOP_N = 3
# Non-answers that must never be returned for these queries at all.
QUALITY_BANNED_HOSTS = ["bestbuy.com", "merriam-webster.com"]
def _consumption_layer_path():
here = os.path.dirname(os.path.abspath(__file__))
return os.path.join(here, "search-agent-consume.py")
def check_ranking_quality() -> list[str]:
"""Return a list of quality failures; empty means healthy."""
import subprocess as _sp
layer = _consumption_layer_path()
if not os.path.exists(layer):
return [f"agent-consumption layer missing: {layer}"]
failures: list[str] = []
for query in QUALITY_QUERIES:
r = _sp.run([sys.executable, layer, "--no-extract", "--explain", query],
capture_output=True, text=True, timeout=120)
if r.returncode != 0:
failures.append(f"{query!r}: layer exited {r.returncode} ({r.stderr[:120]})")
continue
try:
data = json.loads(r.stdout)
except json.JSONDecodeError:
failures.append(f"{query!r}: layer returned unparseable JSON")
continue
results = data.get("results", [])
if len(results) < QUALITY_TOP_N:
failures.append(f"{query!r}: only {len(results)} results returned")
continue
# load the demote list from the SAME config the layer uses
cfg_path = os.path.join(os.path.dirname(layer), "..", "config", "search-ranking.yaml")
demoted: set[str] = set()
try:
sys.path.insert(0, os.path.dirname(layer))
import importlib.util as _iu
spec = _iu.spec_from_file_location("_sac_cfg", layer)
mod = _iu.module_from_spec(spec)
spec.loader.exec_module(mod)
demoted = set(mod._load_config().get("demote_domains", []) or [])
except Exception: # noqa: BLE001
failures.append(f"{query!r}: could not load demote_domains from config")
for item in results[:QUALITY_TOP_N]:
host = (item.get("host") or "")
for d in demoted:
if host == d or host.endswith("." + d):
failures.append(
f"{query!r}: demoted host {host} in top {QUALITY_TOP_N}"
)
for item in results:
host = (item.get("host") or "")
for b in QUALITY_BANNED_HOSTS:
if host == b or host.endswith("." + b):
failures.append(f"{query!r}: non-answer host {host} returned")
return failures
def main() -> int:
failures: list[str] = []
print(f"Search stack check -- {SEARXNG_URL}")
@@ -289,16 +216,6 @@ def main() -> int:
print(f" FAIL: {msg}")
failures.append(msg)
print("-" * 72)
print("RANKING QUALITY (agent-consumption layer)")
quality = check_ranking_quality()
if quality:
for q in quality:
print(f" FAIL: {q}")
failures.extend(quality)
else:
print(" ok: no demoted host in the top 3; no banned non-answer returned")
print("=" * 72)
if failures:
print("VERDICT: FAIL")
-137
View File
@@ -1,137 +0,0 @@
---
kind: function
name: search-agent-consumption
description: >
Agent-consumption layer in front of SearXNG + Firecrawl. Raw multi-engine
aggregation returns results with no dedupe, no filtering and no reranking;
measured 2026-09-26 that put bestbuy.com and merriam-webster.com into "best
practices agent context management", and put four SEO blogs above the real
Proxmox forum threads on a precise technical query. Identical queries also
ranked DIFFERENTLY between runs, which is why the layer is deterministic
rather than dependent on engine behaviour.
Pipeline: dedupe -> drop non-answers -> demote content farms / promote primary
sources -> stable sort -> extract page text for the top N under an explicit
character budget -> stable JSON. Policy lives in config, not code.
Call it when an agent needs search RESULTS rather than links: it returns usable
page text in one call instead of a snippet plus a second fetch.
version: 1.0.0
---
## Where the policy lives
`config/search-ranking.yaml` — reviewable, no code change needed to adjust:
| key | effect |
| --- | --- |
| `non_answer.hosts` / `path_patterns` / `query_keys` / `host_root` | dropped outright |
| `demote_domains` | ranked below everything, never dropped |
| `prefer_domains` | promoted above default rank |
| `ranking.*` | `demote_penalty`, `prefer_bonus`, `multi_engine_bonus` |
| `extraction.*` | `top_n`, `total_chars`, `per_item_chars`, `timeout_seconds` |
**Demotion, not deletion, for content farms**: a genuinely useful hit is not lost,
it simply cannot outrank a primary source. Non-answers are dropped because they
cannot answer a question at all.
## Usage
```bash
python3 scripts/search-agent-consume.py "query text" # JSON
python3 scripts/search-agent-consume.py --no-extract "query" # ranking only
python3 scripts/search-agent-consume.py --explain "query" # + drop reasons
```
Exit `0` ok, `1` nothing survived filtering, `2` the layer could not run.
## Output shape
Stable JSON:
```json
{
"query": "...",
"raw_result_count": 46,
"returned_count": 44,
"dropped_count": 2,
"engines": ["bing", "brave", "duckduckgo", "yandex"],
"results": [
{"rank": 1, "title": "...", "url": "...", "host": "...",
"source_type": "official|code|qa|forum|discussion|web|content-farm",
"engines": ["bing"], "score": 100.0,
"excerpt": "...", "extraction": "ok|truncated|skipped_budget_exhausted|empty|failed:<Type>"}
],
"extraction": {"extracted": 5, "chars_used": 12000, "budget": 12000,
"failures": 0, "seconds": 5.28}
}
```
`--explain` adds `dropped: [{url, reason, position}]` so the filter is auditable
rather than magic.
## Measured before/after (2026-09-26)
Fixed query set. Relevance judged per query, not by impression.
**`best practices agent context management`**
| | before (raw SearXNG) | after (layer) |
| --- | --- | --- |
| 1-2 | anthropic, stackai | anthropic, langchain |
| 3-4 | aitechmonk, agentic-design | jetbrains, blog.jetbrains |
| 5-6 | mindstudio, sparkco | docs.langchain, reddit |
| 7-8 | langchain, medium | cursor, reddit |
| verdict | 4 relevant of 10; 4 content farms; medium.com twice | top 8 all primary/discussion; no content farm in the top 8 |
**`proxmox thin pool metadata exhaustion recovery`**
| | before | after |
| --- | --- | --- |
| 1-4 | vormox, linuxoperatingsystem, riparazioneserver, bigiron (all SEO/thin) | forum.proxmox.com, forum.proxmox.com, gist.github, github |
| 5-9 | forum.proxmox.com x2, voxfor, github, gist | forum.proxmox.com, serverfault, forum.proxmox.com, reddit |
The primary sources moved from positions 5-9 to 1-4.
**Rule proof** (`--explain`, and a direct check of the classifier):
```
DigitalOcean docs -> KEEP (a '/products/' path rule was REMOVED after the
before/after run caught it dropping this page)
Best Buy -> DROP shopping_or_dictionary_host
Merriam-Webster -> DROP shopping_or_dictionary_host
bare homepage -> DROP navigational_host_root
proxmox.com home -> KEEP (preferred host root: a repo/docs front door is
legitimately the answer)
github repo -> KEEP
```
**Extraction cost (criterion 4):**
```
extracted 5 items, 12000 chars used of 12000 budget, 0 failures, 5.28s
whole run end-to-end: 6.4s wall
```
## Regression guard
`search-stack-visibility` asserts the layer still ranks correctly: for the fixed
query set, no `demote_domains` host may appear in the top 3, and the two known
non-answers must not be returned. Without it this layer could silently rot back
to raw ordering, which is exactly what happened to the endpoint colours.
## Reachability, and one honest gap
- **Hermes agents** reach it directly: it reads the same `SEARXNG_URL` and
`FIRECRAWL_URL` they already use.
- **pi agents (MCP search server)**: the MCP server's request/response shape is
**not ours to change**, so this layer is **NOT** wired into it. That is a real
gap, stated rather than claimed as coverage. Closing it would require a change
on the MCP side, which is outside this repo.
## Constraints
Does not touch the live SearXNG or Firecrawl service paths. Third-party
`google cse` is not a hard requirement of this layer — if it 429s, ranking still
works from the remaining engines. No credential is added or required.