Skip to content

Make the quickstart finish, and stop eval scoring its own template - #4

Merged
sahilkalgutkar merged 1 commit into
mainfrom
fix/quickstart-feeds-and-eval
Sep 1, 2026
Merged

sahilkalgutkar merged 1 commit into
mainfrom
fix/quickstart-feeds-and-eval

Conversation

@sahilkalgutkar

Copy link
Copy Markdown
Owner

Part of the audit that began with sahilkalgutkar/modelforge#18 — looking for things CI never runs. I ran this repo's quickstart top to bottom on a clean clone. Four problems, none of which CI can see, because CI runs pytest, not the README.

1. One of the two default feeds is dead

https://www.anthropic.com/news/rss.xml returns 404, and Anthropic no longer publishes an RSS feed at any path I could find. feedparser reports that as a clean parse with zero entries, so nothing failed — half the corpus was just silently missing. Swapped for https://github.blog/changelog/feed/, which is a genuine product changelog and fits what config/feeds.yaml already recommends. poll now prints when a feed returns nothing, with its HTTP status.

2. The first documented command ran for seventeen minutes with no output

python -m ingest.poll took every entry a feed offered, and the OpenAI blog publishes its entire history in one document — 1161 entries, each costing a full trafilatura page fetch. The very first command in the quickstart therefore ran silently for about seventeen minutes before printing anything. It is not hung, but there is no way to tell that while it is happening.

Added max_articles_per_feed (default 25, newest first, 0 for no limit) and a per-feed line as it goes:

$ python -m ingest.poll
github-changelog: 10 new from 10 entries
openai-blog: 25 new from 25 entries
Ingested 35 new articles -> data/articles.jsonl

26.4s total

The cap is applied before fetching, so an entry past the limit is never downloaded at all — that is what the test asserts, rather than just checking the row count.

3. eval.run scored its own placeholder

eval/questions.yaml ships as a template with expected_article_id: "REPLACE_WITH_REAL_ARTICLE_ID". Running the documented command against it produced:

[MISS] What did <source> announce in <article title>?

recall@5: 0/1 (0%)

That is a real-looking number that measures nothing and reads as a broken retriever. eval.run now refuses to score an unfilled set, says how to fill it, and exits 1 without running a search.

The README also said "I hand-built a question/answer set to measure retrieval recall and citation correctness" — there is no such set in the repo, and there is a good reason there isn't: an expected_article_id is only meaningful against the feeds it came from, so shipping mine would leave every id dead for anyone pointing the bot at a different beat. The README now says that instead of claiming otherwise.

The guard does not get in the way of real use — with three questions filled in from an actual index, eval.run reports recall@5: 3/3 (100%) and exits 0.

4. A wrong comment in the quickstart

cp .env.example .env # add your feeds + API key — .env holds only ANTHROPIC_API_KEY; feeds live in config/feeds.yaml.

Verification

ruff check clean; 41 tests pass (8 new) at 96% coverage. Ran the full quickstart afterwards on a clean clone: poll, index build, and eval both refusing the template and scoring a filled-in set.

Three things the README described and the repo did not do.

The anthropic-news feed 404s and has for a while. feedparser reports that as a
clean parse with zero entries, so the only symptom was half the corpus quietly
missing. Swapped for the GitHub changelog, and poll now says out loud when a
feed returns nothing and what status it returned.

`python -m ingest.poll` had no bound on how much of a feed it took, and the
OpenAI blog publishes its whole history -- 1161 entries, one page fetch each.
The first documented command in the quickstart therefore ran for about
seventeen minutes with no output at all, which reads as a hang. Added
max_articles_per_feed (default 25, newest first, 0 for no limit) and a per-feed
count as it goes: the same first run is now 26 seconds.

eval/questions.yaml ships as a template with a REPLACE_WITH_REAL_ARTICLE_ID
placeholder, but `python -m eval.run` scored it anyway and printed
"recall@5: 0/1 (0%)" -- a real-looking number that says nothing about
retrieval and looks like the retriever is broken. It now says the set is
unfilled, explains how to fill it from data/articles.jsonl, and exits 1
without running a search. The README claimed a hand-built question set that
was never in the repo; it now describes the template and why an article id
only means something against the feeds it came from.

Also: the quickstart said `cp .env.example .env  # add your feeds + API key`,
but .env holds only ANTHROPIC_API_KEY -- feeds are in config/feeds.yaml.
@codecov

codecov Bot commented Sep 1, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.77419% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
eval/run.py 93.33% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@sahilkalgutkar
sahilkalgutkar merged commit d276413 into main Sep 1, 2026
2 checks passed
@sahilkalgutkar
sahilkalgutkar deleted the fix/quickstart-feeds-and-eval branch September 1, 2026 19:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant