Repository navigation
Make the quickstart finish, and stop eval scoring its own template - #4
Merged
Merged
Conversation
Three things the README described and the repo did not do. The anthropic-news feed 404s and has for a while. feedparser reports that as a clean parse with zero entries, so the only symptom was half the corpus quietly missing. Swapped for the GitHub changelog, and poll now says out loud when a feed returns nothing and what status it returned. `python -m ingest.poll` had no bound on how much of a feed it took, and the OpenAI blog publishes its whole history -- 1161 entries, one page fetch each. The first documented command in the quickstart therefore ran for about seventeen minutes with no output at all, which reads as a hang. Added max_articles_per_feed (default 25, newest first, 0 for no limit) and a per-feed count as it goes: the same first run is now 26 seconds. eval/questions.yaml ships as a template with a REPLACE_WITH_REAL_ARTICLE_ID placeholder, but `python -m eval.run` scored it anyway and printed "recall@5: 0/1 (0%)" -- a real-looking number that says nothing about retrieval and looks like the retriever is broken. It now says the set is unfilled, explains how to fill it from data/articles.jsonl, and exits 1 without running a search. The README claimed a hand-built question set that was never in the repo; it now describes the template and why an article id only means something against the feeds it came from. Also: the quickstart said `cp .env.example .env # add your feeds + API key`, but .env holds only ANTHROPIC_API_KEY -- feeds are in config/feeds.yaml.
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of the audit that began with sahilkalgutkar/modelforge#18 — looking for things CI never runs. I ran this repo's quickstart top to bottom on a clean clone. Four problems, none of which CI can see, because CI runs
pytest, not the README.1. One of the two default feeds is dead
https://www.anthropic.com/news/rss.xmlreturns 404, and Anthropic no longer publishes an RSS feed at any path I could find. feedparser reports that as a clean parse with zero entries, so nothing failed — half the corpus was just silently missing. Swapped forhttps://github.blog/changelog/feed/, which is a genuine product changelog and fits whatconfig/feeds.yamlalready recommends.pollnow prints when a feed returns nothing, with its HTTP status.2. The first documented command ran for seventeen minutes with no output
python -m ingest.polltook every entry a feed offered, and the OpenAI blog publishes its entire history in one document — 1161 entries, each costing a fulltrafilaturapage fetch. The very first command in the quickstart therefore ran silently for about seventeen minutes before printing anything. It is not hung, but there is no way to tell that while it is happening.Added
max_articles_per_feed(default 25, newest first,0for no limit) and a per-feed line as it goes:The cap is applied before fetching, so an entry past the limit is never downloaded at all — that is what the test asserts, rather than just checking the row count.
3.
eval.runscored its own placeholdereval/questions.yamlships as a template withexpected_article_id: "REPLACE_WITH_REAL_ARTICLE_ID". Running the documented command against it produced:That is a real-looking number that measures nothing and reads as a broken retriever.
eval.runnow refuses to score an unfilled set, says how to fill it, and exits 1 without running a search.The README also said "I hand-built a question/answer set to measure retrieval recall and citation correctness" — there is no such set in the repo, and there is a good reason there isn't: an
expected_article_idis only meaningful against the feeds it came from, so shipping mine would leave every id dead for anyone pointing the bot at a different beat. The README now says that instead of claiming otherwise.The guard does not get in the way of real use — with three questions filled in from an actual index,
eval.runreportsrecall@5: 3/3 (100%)and exits 0.4. A wrong comment in the quickstart
cp .env.example .env # add your feeds + API key—.envholds onlyANTHROPIC_API_KEY; feeds live inconfig/feeds.yaml.Verification
ruff checkclean; 41 tests pass (8 new) at 96% coverage. Ran the full quickstart afterwards on a clean clone: poll, index build, and eval both refusing the template and scoring a filled-in set.