--- kind: function name: search-agent-consumption description: > Agent-consumption layer in front of SearXNG + Firecrawl. Raw multi-engine aggregation returns results with no dedupe, no filtering and no reranking; measured 2026-09-26 that put bestbuy.com and merriam-webster.com into "best practices agent context management", and put four SEO blogs above the real Proxmox forum threads on a precise technical query. Identical queries also ranked DIFFERENTLY between runs, which is why the layer is deterministic rather than dependent on engine behaviour. Pipeline: dedupe -> drop non-answers -> demote content farms / promote primary sources -> stable sort -> extract page text for the top N under an explicit character budget -> stable JSON. Policy lives in config, not code. Call it when an agent needs search RESULTS rather than links: it returns usable page text in one call instead of a snippet plus a second fetch. version: 1.0.0 --- ## Where the policy lives `config/search-ranking.yaml` — reviewable, no code change needed to adjust: | key | effect | | --- | --- | | `non_answer.hosts` / `path_patterns` / `query_keys` / `host_root` | dropped outright | | `demote_domains` | ranked below everything, never dropped | | `prefer_domains` | promoted above default rank | | `ranking.*` | `demote_penalty`, `prefer_bonus`, `multi_engine_bonus` | | `extraction.*` | `top_n`, `total_chars`, `per_item_chars`, `timeout_seconds` | **Demotion, not deletion, for content farms**: a genuinely useful hit is not lost, it simply cannot outrank a primary source. Non-answers are dropped because they cannot answer a question at all. ## Usage ```bash python3 scripts/search-agent-consume.py "query text" # JSON python3 scripts/search-agent-consume.py --no-extract "query" # ranking only python3 scripts/search-agent-consume.py --explain "query" # + drop reasons ``` Exit `0` ok, `1` nothing survived filtering, `2` the layer could not run. ## Output shape Stable JSON: ```json { "query": "...", "raw_result_count": 46, "returned_count": 44, "dropped_count": 2, "engines": ["bing", "brave", "duckduckgo", "yandex"], "results": [ {"rank": 1, "title": "...", "url": "...", "host": "...", "source_type": "official|code|qa|forum|discussion|web|content-farm", "engines": ["bing"], "score": 100.0, "excerpt": "...", "extraction": "ok|truncated|skipped_budget_exhausted|empty|failed:"} ], "extraction": {"extracted": 5, "chars_used": 12000, "budget": 12000, "failures": 0, "seconds": 5.28} } ``` `--explain` adds `dropped: [{url, reason, position}]` so the filter is auditable rather than magic. ## Measured before/after (2026-09-26) Fixed query set. Relevance judged per query, not by impression. **`best practices agent context management`** | | before (raw SearXNG) | after (layer) | | --- | --- | --- | | 1-2 | anthropic, stackai | anthropic, langchain | | 3-4 | aitechmonk, agentic-design | jetbrains, blog.jetbrains | | 5-6 | mindstudio, sparkco | docs.langchain, reddit | | 7-8 | langchain, medium | cursor, reddit | | verdict | 4 relevant of 10; 4 content farms; medium.com twice | top 8 all primary/discussion; no content farm in the top 8 | **`proxmox thin pool metadata exhaustion recovery`** | | before | after | | --- | --- | --- | | 1-4 | vormox, linuxoperatingsystem, riparazioneserver, bigiron (all SEO/thin) | forum.proxmox.com, forum.proxmox.com, gist.github, github | | 5-9 | forum.proxmox.com x2, voxfor, github, gist | forum.proxmox.com, serverfault, forum.proxmox.com, reddit | The primary sources moved from positions 5-9 to 1-4. **Rule proof** (`--explain`, and a direct check of the classifier): ``` DigitalOcean docs -> KEEP (a '/products/' path rule was REMOVED after the before/after run caught it dropping this page) Best Buy -> DROP shopping_or_dictionary_host Merriam-Webster -> DROP shopping_or_dictionary_host bare homepage -> DROP navigational_host_root proxmox.com home -> KEEP (preferred host root: a repo/docs front door is legitimately the answer) github repo -> KEEP ``` **Extraction cost (criterion 4):** ``` extracted 5 items, 12000 chars used of 12000 budget, 0 failures, 5.28s whole run end-to-end: 6.4s wall ``` ## Regression guard `search-stack-visibility` asserts the layer still ranks correctly: for the fixed query set, no `demote_domains` host may appear in the top 3, and the two known non-answers must not be returned. Without it this layer could silently rot back to raw ordering, which is exactly what happened to the endpoint colours. ## Reachability, and one honest gap - **Hermes agents** reach it directly: it reads the same `SEARXNG_URL` and `FIRECRAWL_URL` they already use. - **pi agents (MCP search server)**: the MCP server's request/response shape is **not ours to change**, so this layer is **NOT** wired into it. That is a real gap, stated rather than claimed as coverage. Closing it would require a change on the MCP side, which is outside this repo. ## Constraints Does not touch the live SearXNG or Firecrawl service paths. Third-party `google cse` is not a hard requirement of this layer — if it 429s, ranking still works from the remaining engines. No credential is added or required.