feat(search): agent-consumption layer — dedupe, filter, rerank, extract content #138

Merged
abiba-bot merged 2 commits from feat/search-agent-consumption-20260926 into master 2026-09-26 16:06:50 +00:00
Owner

Closes search-agent-consumption-20260926. Two commits: the layer, then the quality guard.

The finding

Raw multi-engine aggregation had no dedupe, no filtering and no reranking. Measured 2026-09-26:

  • best practices agent context management returned bestbuy.com and merriam-webster.com, four content farms, and medium.com twice;
  • proxmox thin pool metadata exhaustion recovery put four SEO blogs above the real Proxmox forum threads.

The strongest argument for a deterministic layer: the same query ranked differently between runs — one sample had the forum threads at the top, a later one had them at 5–9. Hoping the engines behave is not a strategy.

The layer

scripts/search-agent-consume.py — dedupe → drop non-answers → demote/promote → stable sort → extract → stable JSON.

Policy is config, not code (config/search-ranking.yaml): demote_domains, prefer_domains, non_answer.*, ranking.*, extraction.*. Content farms are demoted, not dropped — a useful hit is not lost, it just cannot outrank a primary source.

Before / after on the fixed query set

best practices agent context management

# before after
1–2 anthropic, stackai anthropic, langchain
3–4 aitechmonk, agentic-design jetbrains, blog.jetbrains
5–6 mindstudio, sparkco docs.langchain, reddit
7–8 langchain, medium cursor, reddit

proxmox thin pool metadata exhaustion recovery

# before after
1–4 vormox, linuxoperatingsystem, riparazioneserver, bigiron forum.proxmox.com, forum.proxmox.com, gist.github, github
5–9 forum.proxmox.com ×2, voxfor, github, gist forum.proxmox.com, serverfault, forum.proxmox.com, reddit

Primary sources moved from 5–9 to 1–4.

Rules proven to fire — and one false positive the evidence caught

DigitalOcean docs  -> KEEP   <- a '/products/' path rule was REMOVED after the
                              before/after run caught it dropping
                              docs.digitalocean.com/products/inference/...
Best Buy           -> DROP   shopping_or_dictionary_host
Merriam-Webster    -> DROP   shopping_or_dictionary_host
bare homepage      -> DROP   navigational_host_root
proxmox.com home   -> KEEP   (preferred host root: a repo/docs front door IS the answer)
github repo        -> KEEP
search URL         -> DROP   path_pattern:/search?

Extraction cost (criterion 4)

extracted 5 items, 12000 chars used of 12000 budget, 0 failures, 5.28s
whole run end-to-end: 6.4s wall

Results carry excerpt plus an extraction status (ok/truncated/skipped_budget_exhausted/empty/failed:<Type>), so a partial extraction is visible rather than silent.

Quality guard (criterion 7)

search-stack-visibility now asserts that for the fixed query set no demote_domains host appears in the top 3 and the two known non-answers are not returned. It reads the demote list from the same config the layer uses, so the guard cannot drift from the policy it guards. Without this the layer could rot back to raw ordering unnoticed — exactly what happened to the endpoint colours before 2026-09-26.

RANKING QUALITY (agent-consumption layer)
  ok: no demoted host in the top 3; no banned non-answer returned
VERDICT: PASS -- multiple engines contributing, extraction healthy
EXIT=0

Reachability — one honest gap

  • Hermes agents reach it directly: same SEARXNG_URL / FIRECRAWL_URL.
  • pi agents (MCP search server): the MCP server's request/response shape is not ours to change, so this layer is NOT wired into it. That is a real gap, documented as a gap — not claimed as coverage. Closing it needs a change on the MCP side, outside this repo.

Constraints honoured

Live SearXNG/Firecrawl paths untouched; the hourly visibility contract still passes; third-party google cse is not a hard requirement (ranking works from the remaining engines if it 429s); no credential added or required.

prose-lint -> ✅ LINT PASSED (21 warning(s))   (secret scan clean)

No master push, no merge.

Closes `search-agent-consumption-20260926`. Two commits: the layer, then the quality guard. ## The finding Raw multi-engine aggregation had **no dedupe, no filtering and no reranking**. Measured 2026-09-26: - `best practices agent context management` returned **bestbuy.com** and **merriam-webster.com**, four content farms, and **medium.com twice**; - `proxmox thin pool metadata exhaustion recovery` put four SEO blogs **above** the real Proxmox forum threads. **The strongest argument for a deterministic layer:** the same query ranked *differently between runs* — one sample had the forum threads at the top, a later one had them at 5–9. Hoping the engines behave is not a strategy. ## The layer `scripts/search-agent-consume.py` — dedupe → drop non-answers → demote/promote → **stable sort** → extract → stable JSON. Policy is **config, not code** (`config/search-ranking.yaml`): `demote_domains`, `prefer_domains`, `non_answer.*`, `ranking.*`, `extraction.*`. Content farms are **demoted, not dropped** — a useful hit is not lost, it just cannot outrank a primary source. ### Before / after on the fixed query set **`best practices agent context management`** | # | before | after | | --- | --- | --- | | 1–2 | anthropic, stackai | anthropic, langchain | | 3–4 | aitechmonk, agentic-design | jetbrains, blog.jetbrains | | 5–6 | mindstudio, sparkco | docs.langchain, reddit | | 7–8 | langchain, medium | cursor, reddit | **`proxmox thin pool metadata exhaustion recovery`** | # | before | after | | --- | --- | --- | | 1–4 | vormox, linuxoperatingsystem, riparazioneserver, bigiron | forum.proxmox.com, forum.proxmox.com, gist.github, github | | 5–9 | forum.proxmox.com ×2, voxfor, github, gist | forum.proxmox.com, serverfault, forum.proxmox.com, reddit | Primary sources moved from **5–9 to 1–4**. ### Rules proven to fire — and one false positive the evidence caught ``` DigitalOcean docs -> KEEP <- a '/products/' path rule was REMOVED after the before/after run caught it dropping docs.digitalocean.com/products/inference/... Best Buy -> DROP shopping_or_dictionary_host Merriam-Webster -> DROP shopping_or_dictionary_host bare homepage -> DROP navigational_host_root proxmox.com home -> KEEP (preferred host root: a repo/docs front door IS the answer) github repo -> KEEP search URL -> DROP path_pattern:/search? ``` ### Extraction cost (criterion 4) ``` extracted 5 items, 12000 chars used of 12000 budget, 0 failures, 5.28s whole run end-to-end: 6.4s wall ``` Results carry `excerpt` plus an `extraction` status (`ok`/`truncated`/`skipped_budget_exhausted`/`empty`/`failed:<Type>`), so a partial extraction is visible rather than silent. ## Quality guard (criterion 7) `search-stack-visibility` now asserts that for the fixed query set **no `demote_domains` host appears in the top 3** and the two known non-answers are not returned. It reads the demote list from the **same config** the layer uses, so the guard cannot drift from the policy it guards. Without this the layer could rot back to raw ordering unnoticed — exactly what happened to the endpoint colours before 2026-09-26. ``` RANKING QUALITY (agent-consumption layer) ok: no demoted host in the top 3; no banned non-answer returned VERDICT: PASS -- multiple engines contributing, extraction healthy EXIT=0 ``` ## Reachability — one honest gap - **Hermes agents** reach it directly: same `SEARXNG_URL` / `FIRECRAWL_URL`. - **pi agents (MCP search server)**: the MCP server's request/response shape is **not ours to change**, so this layer is **NOT wired into it**. That is a real gap, documented as a gap — **not** claimed as coverage. Closing it needs a change on the MCP side, outside this repo. ## Constraints honoured Live SearXNG/Firecrawl paths untouched; the hourly visibility contract still passes; third-party `google cse` is **not** a hard requirement (ranking works from the remaining engines if it 429s); **no credential added or required**. ``` prose-lint -> ✅ LINT PASSED (21 warning(s)) (secret scan clean) ``` No master push, no merge.
abiba-bot added 2 commits 2026-09-26 15:44:22 +00:00
Raw multi-engine aggregation had no dedupe, no filtering and no reranking.
Measured 2026-09-26: 'best practices agent context management' returned
bestbuy.com and merriam-webster.com, plus 4 content farms, with medium.com twice;
'proxmox thin pool metadata exhaustion recovery' put four SEO blogs ABOVE the
real Proxmox forum threads. Identical queries also ranked DIFFERENTLY between
runs, which is why the fix is deterministic rather than trusting the engines.

scripts/search-agent-consume.py:
  1. dedupe by normalised URL (tracking params and fragments stripped)
  2. drop non-answers - shopping/dictionary hosts, navigational host roots,
     search/shopping/cart/login paths and query keys
  3. demote content farms and promote primary sources
  4. STABLE sort (score desc, then original position) so runs are reproducible
  5. extract page text for the top N via Firecrawl POST /v1/scrape under an
     explicit character budget, so an agent gets usable material in ONE call
  6. emit stable JSON with engine provenance and source_type

Policy is config, not code: config/search-ranking.yaml holds demote_domains,
prefer_domains, non_answer rules and the extraction budget, so it is reviewable
and changeable without touching the module. Content farms are DEMOTED rather
than dropped so a useful hit is not lost, it just cannot outrank a primary.

A '/products/' path rule was REMOVED after the before/after run caught it
dropping docs.digitalocean.com/products/inference/... - a legitimate docs page.
Shopping is caught by the host list instead, which has no such false positive.

Measured: 'best practices...' top 8 becomes anthropic, langchain, jetbrains,
blog.jetbrains, docs.langchain, reddit, cursor, reddit - no content farm.
'proxmox thin pool...' moves the forum threads from positions 5-9 to 1-4.
Extraction: 5 items, 12000 chars of 12000 budget, 0 failures, 5.28s; whole run
6.4s wall.

Contract: search-agent-consumption.prose.md, including the honest reachability
gap - the pi MCP search server's shape is not ours to change, so this layer is
NOT wired into it.
test(search): quality guard so the ranking layer cannot silently rot
PR Pipeline — Authorize → Validate → Review → Merge / auth (pull_request) Successful in 5s
PR Pipeline — Authorize → Validate → Review → Merge / validate (pull_request) Successful in 4s
PR Pipeline — Authorize → Validate → Review → Merge / lint (pull_request) Successful in 13s
PR Pipeline — Authorize → Validate → Review → Merge / ai-review (pull_request) Successful in 11s
PR Pipeline — Authorize → Validate → Review → Merge / gate (pull_request) Successful in 1s
ba38efcd75
Extends search-stack-visibility with a ranking assertion: for the fixed query
set, no config demote_domains host may appear in the top 3, and the known
non-answers (bestbuy.com, merriam-webster.com) must not be returned at all.

Without this the layer could rot back to raw engine ordering unnoticed - the same
way the endpoint colours silently rotted before 2026-09-26. It reads the demote
list from the SAME config the layer uses, so the guard cannot drift from the
policy it is guarding.

Live: 'ok: no demoted host in the top 3; no banned non-answer returned';
visibility contract still PASSES end to end.
abiba-bot merged commit d99b552448 into master 2026-09-26 16:06:50 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: SyslogSolution/prose-contracts#138