Files
prose-contracts/search-agent-consumption.prose.md
root 0d30091f62 feat(search): agent-consumption layer - dedupe, filter, rerank, extract
Raw multi-engine aggregation had no dedupe, no filtering and no reranking.
Measured 2026-09-26: 'best practices agent context management' returned
bestbuy.com and merriam-webster.com, plus 4 content farms, with medium.com twice;
'proxmox thin pool metadata exhaustion recovery' put four SEO blogs ABOVE the
real Proxmox forum threads. Identical queries also ranked DIFFERENTLY between
runs, which is why the fix is deterministic rather than trusting the engines.

scripts/search-agent-consume.py:
  1. dedupe by normalised URL (tracking params and fragments stripped)
  2. drop non-answers - shopping/dictionary hosts, navigational host roots,
     search/shopping/cart/login paths and query keys
  3. demote content farms and promote primary sources
  4. STABLE sort (score desc, then original position) so runs are reproducible
  5. extract page text for the top N via Firecrawl POST /v1/scrape under an
     explicit character budget, so an agent gets usable material in ONE call
  6. emit stable JSON with engine provenance and source_type

Policy is config, not code: config/search-ranking.yaml holds demote_domains,
prefer_domains, non_answer rules and the extraction budget, so it is reviewable
and changeable without touching the module. Content farms are DEMOTED rather
than dropped so a useful hit is not lost, it just cannot outrank a primary.

A '/products/' path rule was REMOVED after the before/after run caught it
dropping docs.digitalocean.com/products/inference/... - a legitimate docs page.
Shopping is caught by the host list instead, which has no such false positive.

Measured: 'best practices...' top 8 becomes anthropic, langchain, jetbrains,
blog.jetbrains, docs.langchain, reddit, cursor, reddit - no content farm.
'proxmox thin pool...' moves the forum threads from positions 5-9 to 1-4.
Extraction: 5 items, 12000 chars of 12000 budget, 0 failures, 5.28s; whole run
6.4s wall.

Contract: search-agent-consumption.prose.md, including the honest reachability
gap - the pi MCP search server's shape is not ours to change, so this layer is
NOT wired into it.
2026-09-26 15:44:09 +00:00

138 lines
5.3 KiB
Markdown

---
kind: function
name: search-agent-consumption
description: >
Agent-consumption layer in front of SearXNG + Firecrawl. Raw multi-engine
aggregation returns results with no dedupe, no filtering and no reranking;
measured 2026-09-26 that put bestbuy.com and merriam-webster.com into "best
practices agent context management", and put four SEO blogs above the real
Proxmox forum threads on a precise technical query. Identical queries also
ranked DIFFERENTLY between runs, which is why the layer is deterministic
rather than dependent on engine behaviour.
Pipeline: dedupe -> drop non-answers -> demote content farms / promote primary
sources -> stable sort -> extract page text for the top N under an explicit
character budget -> stable JSON. Policy lives in config, not code.
Call it when an agent needs search RESULTS rather than links: it returns usable
page text in one call instead of a snippet plus a second fetch.
version: 1.0.0
---
## Where the policy lives
`config/search-ranking.yaml` — reviewable, no code change needed to adjust:
| key | effect |
| --- | --- |
| `non_answer.hosts` / `path_patterns` / `query_keys` / `host_root` | dropped outright |
| `demote_domains` | ranked below everything, never dropped |
| `prefer_domains` | promoted above default rank |
| `ranking.*` | `demote_penalty`, `prefer_bonus`, `multi_engine_bonus` |
| `extraction.*` | `top_n`, `total_chars`, `per_item_chars`, `timeout_seconds` |
**Demotion, not deletion, for content farms**: a genuinely useful hit is not lost,
it simply cannot outrank a primary source. Non-answers are dropped because they
cannot answer a question at all.
## Usage
```bash
python3 scripts/search-agent-consume.py "query text" # JSON
python3 scripts/search-agent-consume.py --no-extract "query" # ranking only
python3 scripts/search-agent-consume.py --explain "query" # + drop reasons
```
Exit `0` ok, `1` nothing survived filtering, `2` the layer could not run.
## Output shape
Stable JSON:
```json
{
"query": "...",
"raw_result_count": 46,
"returned_count": 44,
"dropped_count": 2,
"engines": ["bing", "brave", "duckduckgo", "yandex"],
"results": [
{"rank": 1, "title": "...", "url": "...", "host": "...",
"source_type": "official|code|qa|forum|discussion|web|content-farm",
"engines": ["bing"], "score": 100.0,
"excerpt": "...", "extraction": "ok|truncated|skipped_budget_exhausted|empty|failed:<Type>"}
],
"extraction": {"extracted": 5, "chars_used": 12000, "budget": 12000,
"failures": 0, "seconds": 5.28}
}
```
`--explain` adds `dropped: [{url, reason, position}]` so the filter is auditable
rather than magic.
## Measured before/after (2026-09-26)
Fixed query set. Relevance judged per query, not by impression.
**`best practices agent context management`**
| | before (raw SearXNG) | after (layer) |
| --- | --- | --- |
| 1-2 | anthropic, stackai | anthropic, langchain |
| 3-4 | aitechmonk, agentic-design | jetbrains, blog.jetbrains |
| 5-6 | mindstudio, sparkco | docs.langchain, reddit |
| 7-8 | langchain, medium | cursor, reddit |
| verdict | 4 relevant of 10; 4 content farms; medium.com twice | top 8 all primary/discussion; no content farm in the top 8 |
**`proxmox thin pool metadata exhaustion recovery`**
| | before | after |
| --- | --- | --- |
| 1-4 | vormox, linuxoperatingsystem, riparazioneserver, bigiron (all SEO/thin) | forum.proxmox.com, forum.proxmox.com, gist.github, github |
| 5-9 | forum.proxmox.com x2, voxfor, github, gist | forum.proxmox.com, serverfault, forum.proxmox.com, reddit |
The primary sources moved from positions 5-9 to 1-4.
**Rule proof** (`--explain`, and a direct check of the classifier):
```
DigitalOcean docs -> KEEP (a '/products/' path rule was REMOVED after the
before/after run caught it dropping this page)
Best Buy -> DROP shopping_or_dictionary_host
Merriam-Webster -> DROP shopping_or_dictionary_host
bare homepage -> DROP navigational_host_root
proxmox.com home -> KEEP (preferred host root: a repo/docs front door is
legitimately the answer)
github repo -> KEEP
```
**Extraction cost (criterion 4):**
```
extracted 5 items, 12000 chars used of 12000 budget, 0 failures, 5.28s
whole run end-to-end: 6.4s wall
```
## Regression guard
`search-stack-visibility` asserts the layer still ranks correctly: for the fixed
query set, no `demote_domains` host may appear in the top 3, and the two known
non-answers must not be returned. Without it this layer could silently rot back
to raw ordering, which is exactly what happened to the endpoint colours.
## Reachability, and one honest gap
- **Hermes agents** reach it directly: it reads the same `SEARXNG_URL` and
`FIRECRAWL_URL` they already use.
- **pi agents (MCP search server)**: the MCP server's request/response shape is
**not ours to change**, so this layer is **NOT** wired into it. That is a real
gap, stated rather than claimed as coverage. Closing it would require a change
on the MCP side, which is outside this repo.
## Constraints
Does not touch the live SearXNG or Firecrawl service paths. Third-party
`google cse` is not a hard requirement of this layer — if it 429s, ranking still
works from the remaining engines. No credential is added or required.