add a query source that does not need a language model

Idea taken from TheNetsky/Microsoft-Rewards-Script, which builds search terms
from public feeds rather than a model. No code from it: that project is
GPL-3.0 and this one is MIT, so only the approach crosses over.

The LLM has exactly two call sites here, both producing a short string to type
into Bing. Everything the dependency costs, an Ollama account, cloud usage and
the provider work in #15, is paid for search strings. Three keyless sources
answer the same question:

  Google Trends RSS    queries people are actually typing right now
  Wikipedia most-read  topic seeds when trends is unavailable
  Bing autosuggest     expands a seed into related queries

Autosuggest is what makes the chaining work. Asking Bing what follows a term
returns queries Bing already expects, which is nearer to what the prompt in
llm_utils was reaching for than a model guessing unaided.

Selected with QUERY_SOURCE=trends. The default stays llm, so no existing setup
changes. stdlib only, no new dependencies.

Measured against the LLM on the same cards from a live account:

  card                     llm                                    trends
  airport parking          best rates airport parking reservations reserve airport parking best rates
  checking vs savings      compare checking vs savings accounts    compare checking savings account options
  cruise deals             best cruise deals and destinations      cruise deals destinations

Verified live with OLLAMA_HOST pointed at a dead port, so nothing could reach
a model: five queries generated from feeds and three typed into Bing, each
landing on a real results page.

Every source degrades to an empty list rather than raising, and both entry
points fall back, to the trimmed task description and to nouns.txt. A search
that does not happen costs points; a run that dies costs the rest of the day.
This commit is contained in:
Ethan Stoner
2026-08-26 15:05:09 -07:00
parent 6f6ffa3fa4
commit af03afcc6d
4 changed files with 281 additions and 5 deletions
+64
View File
@@ -0,0 +1,64 @@
"""Where search queries come from.
Two backends. `llm` is the default and is unchanged, so nothing about an
existing setup moves. `trends` uses public feeds and needs no account, no
model and no key, which is the difference between running this in five
minutes and installing Ollama first.
QUERY_SOURCE=trends python src/main.py
The LLM's whole job in this project is producing short strings to type into
Bing, and Bing's own autosuggest answers that question directly.
"""
import os
import llm_utils
import query_sources
LLM = "llm"
TRENDS = "trends"
DEFAULT_SOURCE = LLM
ENV_VAR = "QUERY_SOURCE"
def selected_source() -> str:
"""Read on each call so a test can change it without reimporting."""
choice = os.environ.get(ENV_VAR, DEFAULT_SOURCE).strip().lower()
return choice if choice in (LLM, TRENDS) else DEFAULT_SOURCE
def search_query_for_task(task_description: str) -> str:
"""A query for one "Search on Bing for X" card."""
if selected_source() == TRENDS:
query = query_sources.query_from_task_description(task_description)
if query:
return query
# Every feed was unreachable. The description still contains the topic,
# so a trimmed version beats skipping the card entirely.
print("[WARNING] No query source reachable, using the task description as written.")
return task_description.lower()
return llm_utils.get_search_query_from_task_description(task_description)
def related_queries(count: int):
"""`count` queries for the daily search quota."""
if selected_source() == TRENDS:
queries = query_sources.related_queries(count)
if queries:
return queries
print("[WARNING] No query source reachable, falling back to the wordlist.")
# nouns.txt is already in the repo for exactly this kind of seed.
return [llm_utils.get_random_noun() for _ in range(count)]
return llm_utils.get_related_search_queries(llm_utils.get_random_noun(), num_queries=count)