Closes search-agent-consumption-20260926. Two commits: the layer, then the quality guard.
The finding
Raw multi-engine aggregation had no dedupe, no filtering and no reranking. Measured 2026-09-26:
best practices agent context management returned bestbuy.com and merriam-webster.com, four content farms, and medium.com twice;
proxmox thin pool metadata exhaustion recovery put four SEO blogs above the real Proxmox forum threads.
The strongest argument for a deterministic layer: the same query ranked differently between runs — one sample had the forum threads at the top, a later one had them at 5–9. Hoping the engines behave is not a strategy.
Policy is config, not code (config/search-ranking.yaml): demote_domains, prefer_domains, non_answer.*, ranking.*, extraction.*. Content farms are demoted, not dropped — a useful hit is not lost, it just cannot outrank a primary source.
Rules proven to fire — and one false positive the evidence caught
DigitalOcean docs -> KEEP <- a '/products/' path rule was REMOVED after the
before/after run caught it dropping
docs.digitalocean.com/products/inference/...
Best Buy -> DROP shopping_or_dictionary_host
Merriam-Webster -> DROP shopping_or_dictionary_host
bare homepage -> DROP navigational_host_root
proxmox.com home -> KEEP (preferred host root: a repo/docs front door IS the answer)
github repo -> KEEP
search URL -> DROP path_pattern:/search?
Extraction cost (criterion 4)
extracted 5 items, 12000 chars used of 12000 budget, 0 failures, 5.28s
whole run end-to-end: 6.4s wall
Results carry excerpt plus an extraction status (ok/truncated/skipped_budget_exhausted/empty/failed:<Type>), so a partial extraction is visible rather than silent.
Quality guard (criterion 7)
search-stack-visibility now asserts that for the fixed query set no demote_domains host appears in the top 3 and the two known non-answers are not returned. It reads the demote list from the same config the layer uses, so the guard cannot drift from the policy it guards. Without this the layer could rot back to raw ordering unnoticed — exactly what happened to the endpoint colours before 2026-09-26.
RANKING QUALITY (agent-consumption layer)
ok: no demoted host in the top 3; no banned non-answer returned
VERDICT: PASS -- multiple engines contributing, extraction healthy
EXIT=0
Reachability — one honest gap
Hermes agents reach it directly: same SEARXNG_URL / FIRECRAWL_URL.
pi agents (MCP search server): the MCP server's request/response shape is not ours to change, so this layer is NOT wired into it. That is a real gap, documented as a gap — not claimed as coverage. Closing it needs a change on the MCP side, outside this repo.
Constraints honoured
Live SearXNG/Firecrawl paths untouched; the hourly visibility contract still passes; third-party google cse is not a hard requirement (ranking works from the remaining engines if it 429s); no credential added or required.
Closes `search-agent-consumption-20260926`. Two commits: the layer, then the quality guard.
## The finding
Raw multi-engine aggregation had **no dedupe, no filtering and no reranking**. Measured 2026-09-26:
- `best practices agent context management` returned **bestbuy.com** and **merriam-webster.com**, four content farms, and **medium.com twice**;
- `proxmox thin pool metadata exhaustion recovery` put four SEO blogs **above** the real Proxmox forum threads.
**The strongest argument for a deterministic layer:** the same query ranked *differently between runs* — one sample had the forum threads at the top, a later one had them at 5–9. Hoping the engines behave is not a strategy.
## The layer
`scripts/search-agent-consume.py` — dedupe → drop non-answers → demote/promote → **stable sort** → extract → stable JSON.
Policy is **config, not code** (`config/search-ranking.yaml`): `demote_domains`, `prefer_domains`, `non_answer.*`, `ranking.*`, `extraction.*`. Content farms are **demoted, not dropped** — a useful hit is not lost, it just cannot outrank a primary source.
### Before / after on the fixed query set
**`best practices agent context management`**
| # | before | after |
| --- | --- | --- |
| 1–2 | anthropic, stackai | anthropic, langchain |
| 3–4 | aitechmonk, agentic-design | jetbrains, blog.jetbrains |
| 5–6 | mindstudio, sparkco | docs.langchain, reddit |
| 7–8 | langchain, medium | cursor, reddit |
**`proxmox thin pool metadata exhaustion recovery`**
| # | before | after |
| --- | --- | --- |
| 1–4 | vormox, linuxoperatingsystem, riparazioneserver, bigiron | forum.proxmox.com, forum.proxmox.com, gist.github, github |
| 5–9 | forum.proxmox.com ×2, voxfor, github, gist | forum.proxmox.com, serverfault, forum.proxmox.com, reddit |
Primary sources moved from **5–9 to 1–4**.
### Rules proven to fire — and one false positive the evidence caught
```
DigitalOcean docs -> KEEP <- a '/products/' path rule was REMOVED after the
before/after run caught it dropping
docs.digitalocean.com/products/inference/...
Best Buy -> DROP shopping_or_dictionary_host
Merriam-Webster -> DROP shopping_or_dictionary_host
bare homepage -> DROP navigational_host_root
proxmox.com home -> KEEP (preferred host root: a repo/docs front door IS the answer)
github repo -> KEEP
search URL -> DROP path_pattern:/search?
```
### Extraction cost (criterion 4)
```
extracted 5 items, 12000 chars used of 12000 budget, 0 failures, 5.28s
whole run end-to-end: 6.4s wall
```
Results carry `excerpt` plus an `extraction` status (`ok`/`truncated`/`skipped_budget_exhausted`/`empty`/`failed:<Type>`), so a partial extraction is visible rather than silent.
## Quality guard (criterion 7)
`search-stack-visibility` now asserts that for the fixed query set **no `demote_domains` host appears in the top 3** and the two known non-answers are not returned. It reads the demote list from the **same config** the layer uses, so the guard cannot drift from the policy it guards. Without this the layer could rot back to raw ordering unnoticed — exactly what happened to the endpoint colours before 2026-09-26.
```
RANKING QUALITY (agent-consumption layer)
ok: no demoted host in the top 3; no banned non-answer returned
VERDICT: PASS -- multiple engines contributing, extraction healthy
EXIT=0
```
## Reachability — one honest gap
- **Hermes agents** reach it directly: same `SEARXNG_URL` / `FIRECRAWL_URL`.
- **pi agents (MCP search server)**: the MCP server's request/response shape is **not ours to change**, so this layer is **NOT wired into it**. That is a real gap, documented as a gap — **not** claimed as coverage. Closing it needs a change on the MCP side, outside this repo.
## Constraints honoured
Live SearXNG/Firecrawl paths untouched; the hourly visibility contract still passes; third-party `google cse` is **not** a hard requirement (ranking works from the remaining engines if it 429s); **no credential added or required**.
```
prose-lint -> ✅ LINT PASSED (21 warning(s)) (secret scan clean)
```
No master push, no merge.
Raw multi-engine aggregation had no dedupe, no filtering and no reranking.
Measured 2026-09-26: 'best practices agent context management' returned
bestbuy.com and merriam-webster.com, plus 4 content farms, with medium.com twice;
'proxmox thin pool metadata exhaustion recovery' put four SEO blogs ABOVE the
real Proxmox forum threads. Identical queries also ranked DIFFERENTLY between
runs, which is why the fix is deterministic rather than trusting the engines.
scripts/search-agent-consume.py:
1. dedupe by normalised URL (tracking params and fragments stripped)
2. drop non-answers - shopping/dictionary hosts, navigational host roots,
search/shopping/cart/login paths and query keys
3. demote content farms and promote primary sources
4. STABLE sort (score desc, then original position) so runs are reproducible
5. extract page text for the top N via Firecrawl POST /v1/scrape under an
explicit character budget, so an agent gets usable material in ONE call
6. emit stable JSON with engine provenance and source_type
Policy is config, not code: config/search-ranking.yaml holds demote_domains,
prefer_domains, non_answer rules and the extraction budget, so it is reviewable
and changeable without touching the module. Content farms are DEMOTED rather
than dropped so a useful hit is not lost, it just cannot outrank a primary.
A '/products/' path rule was REMOVED after the before/after run caught it
dropping docs.digitalocean.com/products/inference/... - a legitimate docs page.
Shopping is caught by the host list instead, which has no such false positive.
Measured: 'best practices...' top 8 becomes anthropic, langchain, jetbrains,
blog.jetbrains, docs.langchain, reddit, cursor, reddit - no content farm.
'proxmox thin pool...' moves the forum threads from positions 5-9 to 1-4.
Extraction: 5 items, 12000 chars of 12000 budget, 0 failures, 5.28s; whole run
6.4s wall.
Contract: search-agent-consumption.prose.md, including the honest reachability
gap - the pi MCP search server's shape is not ours to change, so this layer is
NOT wired into it.
Extends search-stack-visibility with a ranking assertion: for the fixed query
set, no config demote_domains host may appear in the top 3, and the known
non-answers (bestbuy.com, merriam-webster.com) must not be returned at all.
Without this the layer could rot back to raw engine ordering unnoticed - the same
way the endpoint colours silently rotted before 2026-09-26. It reads the demote
list from the SAME config the layer uses, so the guard cannot drift from the
policy it is guarding.
Live: 'ok: no demoted host in the top 3; no banned non-answer returned';
visibility contract still PASSES end to end.
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Closes
search-agent-consumption-20260926. Two commits: the layer, then the quality guard.The finding
Raw multi-engine aggregation had no dedupe, no filtering and no reranking. Measured 2026-09-26:
best practices agent context managementreturned bestbuy.com and merriam-webster.com, four content farms, and medium.com twice;proxmox thin pool metadata exhaustion recoveryput four SEO blogs above the real Proxmox forum threads.The strongest argument for a deterministic layer: the same query ranked differently between runs — one sample had the forum threads at the top, a later one had them at 5–9. Hoping the engines behave is not a strategy.
The layer
scripts/search-agent-consume.py— dedupe → drop non-answers → demote/promote → stable sort → extract → stable JSON.Policy is config, not code (
config/search-ranking.yaml):demote_domains,prefer_domains,non_answer.*,ranking.*,extraction.*. Content farms are demoted, not dropped — a useful hit is not lost, it just cannot outrank a primary source.Before / after on the fixed query set
best practices agent context managementproxmox thin pool metadata exhaustion recoveryPrimary sources moved from 5–9 to 1–4.
Rules proven to fire — and one false positive the evidence caught
Extraction cost (criterion 4)
Results carry
excerptplus anextractionstatus (ok/truncated/skipped_budget_exhausted/empty/failed:<Type>), so a partial extraction is visible rather than silent.Quality guard (criterion 7)
search-stack-visibilitynow asserts that for the fixed query set nodemote_domainshost appears in the top 3 and the two known non-answers are not returned. It reads the demote list from the same config the layer uses, so the guard cannot drift from the policy it guards. Without this the layer could rot back to raw ordering unnoticed — exactly what happened to the endpoint colours before 2026-09-26.Reachability — one honest gap
SEARXNG_URL/FIRECRAWL_URL.Constraints honoured
Live SearXNG/Firecrawl paths untouched; the hourly visibility contract still passes; third-party
google cseis not a hard requirement (ranking works from the remaining engines if it 429s); no credential added or required.No master push, no merge.
Raw multi-engine aggregation had no dedupe, no filtering and no reranking. Measured 2026-09-26: 'best practices agent context management' returned bestbuy.com and merriam-webster.com, plus 4 content farms, with medium.com twice; 'proxmox thin pool metadata exhaustion recovery' put four SEO blogs ABOVE the real Proxmox forum threads. Identical queries also ranked DIFFERENTLY between runs, which is why the fix is deterministic rather than trusting the engines. scripts/search-agent-consume.py: 1. dedupe by normalised URL (tracking params and fragments stripped) 2. drop non-answers - shopping/dictionary hosts, navigational host roots, search/shopping/cart/login paths and query keys 3. demote content farms and promote primary sources 4. STABLE sort (score desc, then original position) so runs are reproducible 5. extract page text for the top N via Firecrawl POST /v1/scrape under an explicit character budget, so an agent gets usable material in ONE call 6. emit stable JSON with engine provenance and source_type Policy is config, not code: config/search-ranking.yaml holds demote_domains, prefer_domains, non_answer rules and the extraction budget, so it is reviewable and changeable without touching the module. Content farms are DEMOTED rather than dropped so a useful hit is not lost, it just cannot outrank a primary. A '/products/' path rule was REMOVED after the before/after run caught it dropping docs.digitalocean.com/products/inference/... - a legitimate docs page. Shopping is caught by the host list instead, which has no such false positive. Measured: 'best practices...' top 8 becomes anthropic, langchain, jetbrains, blog.jetbrains, docs.langchain, reddit, cursor, reddit - no content farm. 'proxmox thin pool...' moves the forum threads from positions 5-9 to 1-4. Extraction: 5 items, 12000 chars of 12000 budget, 0 failures, 5.28s; whole run 6.4s wall. Contract: search-agent-consumption.prose.md, including the honest reachability gap - the pi MCP search server's shape is not ours to change, so this layer is NOT wired into it.