Skip to content

Review: UC Berkeley mirror + task verifiers (site by @richard-peng-xia, verifiers by reviewer) - #116

Merged
Raibows merged 28 commits into
aiming-lab:mainfrom
evanz37:review/pr-11-berkeley-delivery
Sep 15, 2026
Merged

Raibows merged 28 commits into
aiming-lab:mainfrom
evanz37:review/pr-11-berkeley-delivery

Conversation

@evanz37

@evanz37 evanz37 commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

Review + deterministic grading contract for sites/berkeley

Reviewer PR for #11. Branch review/pr-11-berkeley-delivery. Author commit fa70f03 is unchanged; all commits below are the reviewer's, on top of it.

What this PR carries

Integrationc0a3fde register berkeley as site 26 / port 40026; 8e5bb7f merge upstream/main (Healthline #105 took index 26 / 40026, Kaggle #106 took index 27 / 40027) and re-slot berkeley to site 28 / port 40028 (29 sites, 40000-40028); docs sweep.
Seed + clock

  • 2f0c6c4 build-generated, byte-reproducible seed (3001bcf4… under PYTHONHASHSEED=0 and =1).
  • c1cdcb8 freeze the benchmark clock (BENCHMARK_NOW = 2026-05-12) so date filters are deterministic.
  • 1929b84 article detail no longer writes view_count (a read path was writing the DB).
    Verifiers + tasks
  • 1d17acb add 22 deterministic task verifiers + verify_lib (targets re-derived from the snapshot).
  • 0ff43bf task selection and verifier_path / judge_rubric backfill; 32 candidate rows triaged to 22.
  • 67ee9a5 add the run-signature writer and the browser-driven matrix harness.
  • eb19a2b site README and integration tests.
  • 50813e7 matrix writer drives the real recorder contract (grade from agent_demo/, pre-navigate, port guard).
  • 67cff3d negation before a titled name is no longer hidden by "Prof." (found by the mutation rows).
  • 5aca1de rubric polish — all 22 judge_rubric values rewritten (rubric-only diff).
    Robustness + UI
  • 66a0c8d app robustness: POST-only logout, per-process secret key, bounded inputs, FK enforcement.
  • 077f37f answer facts no longer render before their discovery route (+ the answer-leak sweep).
  • 0603a3d task 16 workflow pages to the catalogue page holding its target.
  • 678f703 neutral ORDER BY on list queries and /search field-set parity.
  • 57cb484 WCAG AA contrast, 320px layout, heading order.
  • 455e5f7 the Featured badge escapes its card instead of overlapping the top bar.
  • d5d42b9 seed generator self-locks its engine URI; scope a README claim.
  • ae2bdbd two-tone focus ring so the indicator clears 3:1 on the navy band.
  • a38df80 task 27 search gate accepts the route the ques names (resolved under the review's §0.4).
  • f237364 drop references to untracked review artifacts from two test docstrings.
    Imagery19f7322 synthetic imagery: 82 FLUX.1 [schnell] scenes + 82 deterministic Pillow avatars, the inventory and build gate, and the template wiring. New asset archive, uploaded to HF as PR Review: harden WebHarbor task validation (#45) #91. See Imagery below.
    Scripts exemption — droppable86ba06e fetch_assets.sh skips build-generated sites with no archive. Touches scripts/, not the site; maintainers may drop this commit if they prefer to own the exemption. Since 19f7322 the archive exists, so berkeley no longer takes that branch — the exemption is no longer exercised by this site, though it stays a valid fallback for any build-generated site without an archive at the pinned revision.

Imagery

The mirror shipped no images at all before this commit — static/ held two .gitkeep files and every card image was a flat colour plus a text label. That is a real fidelity gap for a mirror of berkeley.edu, where every section page leads with a photograph.

What. 164 generated files, 6,991,525 bytes (6.67 MiB): campus 8, college 14, research 25, news 21 (7 categories × 3 variants), event 14 (7 × 2), faculty 82 — 1024×768 WebP scenes and 256×256 PNG avatars. Variants bind to the row's primary key (id % 3 + 1, id % 2 + 1), never to a render index, so a listing card and a detail banner agree on the same file and pagination cannot reselect one. Every <img> carries explicit width/height inside an aspect-ratio or fixed-height box, and its alt is built from seed fields the page already renders — never a title, director, founding year or focus area.

How generated. Scenes: fal.ai FLUX.1 [schnell], 4 inference steps, one fixed seed per slot derived from its slug. Avatars: not model-generated — Pillow 11.0.0 draws monogram discs from each faculty row's initials and id, offline and byte-reproducibly. Prompt policy is one subject clause plus a fixed style suffix plus a fixed negative list: no faces, no portraits, no text, no lettering, no logos, no watermarks, no posters, no framed pictures, no screens with visible content, no recognizable landmarks or signage. Labs, research interiors, news and event venues are prompted unoccupied; campus and college-exterior slots permit at most two or three figures in the middle distance, backs turned, faces hidden. No image depicts a real person, a real face or a real landmark — the campus scenes are generic institutional architecture, deliberately not the Campanile, Sather Gate or the Golden Gate. The plan, exact prompts and every calibration live in sites/berkeley/scripts/IMAGE_PLAN.md; the per-file prompt is in generated_asset_inventory.json.

Gates. check_generated_assets.py runs in the Docker build (verified 164 ... assets) and enforces exact coverage, extension/kind agreement, per-file SHA-256, a full decode at the planned dimensions and a letterbox test. tests/test_generated_assets.py adds the OCR and face passes with positive controls, and fails if any row is left qa=flagged.

Honest limits. Neither detector is a guarantee and both were measured, not assumed. The OCR pass returned zero tokens for the word STCK rendered in large red capitals in colleges/chemistry.webp and missed a placard in research/bair.webp; it also flagged clean scenes on three-character junk until the floor was raised. The face pass is a frontal cascade, not a person detector — the OpenCV defaults gave 15 false positives on five clean pilot scenes, and the calibrated settings that score zero there also miss the small profile faces the rule cares about. What actually bounds the bundle is the by-eye contact sheet, which across five passes found the seated person, the crowds with faces to camera, the framed portrait, three letterboxed frames, STCK, a watermark-like label in the newsroom, a face to camera in journalism and faces in law — the last four regenerated on new seeds to the reviewer's instruction.

HF PR + re-pin. Archive berkeley.tar.gz, 6,951,483 bytes, sha256 ab9d2716ae8d06540a181b5e60c37f613d87b103864b467511da546b1b173789, 171 managed members, validated by scripts/validate_asset_archive.py. Uploaded as https://huggingface.co/datasets/ChilleD/WebHarbor/discussions/91 (head commit 4529b18c6fc23f9b203a281b2d1e1940ba93853b). .assets-revision pins revision: refs/pr/91 as an explicitly interim pin: it is a PR ref, so it moves if the PR branch is updated. Maintainers must re-pin revision: to the merged commit sha once PR #91 lands and delete the interim paragraph in that file; the "never a moving branch name" rule then applies again. fetch_assets.sh berkeley was verified against this exact ref: it downloads, validates 171 members, installs the roots, and the fetched sha256 equals the local tarball.

Grading contract

  1. Each of the 22 accepted rows has a deterministic verifier (verify/verify_N.py) — no LLM, no key.
  2. Targets are re-derived from the supplied initial_db, never frozen; ambiguity fails closed.
  3. Read-only tasks require every seeded table to be row-identical before and after execution.
  4. The stateful rows (30/31) require an exact bookmark delta for the demo account, an unchanged-everything-else check, and (31) the surviving row id.
  5. The LLM judge is secondary; judge_rubric carries scoring rules only, never an answer value.

Validation

  • Matrix: 112/112 cells correct, 0 mismatches — 22 tasks × {pass, no_op, shortcut, wrong_answer, collateral_write} + state_mismatch for 30/31; every pass PASSes, every variant FAILs on its intended first failing check.
  • Mutations: 176/176 rows FAIL, 0 infra errors — 22 tasks × 8 classes (1×1 PNG, catalog-wide search, truncated, another task's trajectory, missing after_db, read-only write, negated answer, answer-only).
  • Unit tests (separate pytest processes): berkeley 569 (site 35 + verify 534), walmart_careers 51, rotten_tomatoes 56 + 4129 subtests. Re-run on the imagery commit: 569 passed.
  • Imagery (19f7322): check_generated_assets.py verified 164 assets in the Docker build; OCR pass over all 82 scenes with a live positive control, 0 rows qa=flagged; broken-image sweep over 424 routes × 2 widths inside the image, 0 broken (the only 4xx is the intentional /nope-404).
  • Container: 29/29 sites HTTP 200; /health 29 alive+ready; POST /reset/berkeley byte-identical (f2f0187c… == instance_seed), still identical after a real authenticated write + reset and after docker restart; reset-all 27/27 in 2.6 s; 419-route smoke with the DB byte-identical afterwards; no-op matrix 22/22 FAIL, 0 errors (re-run on the imagery commit: 22 cells, 0 mismatches).
  • Real runs: stock agent_demo/agent.py 22 runs → 1 PASS / 21 FAIL, 0 verifier↔judge divergences. The endpoint returns done with a top-level text while the shipped parser reads params.text, so most stock runs recorded empty answers; a reviewer-side parser shim (recorder untouched, AST-verified byte-compatible) recovered 21 gradeable runs → verifier 10 PASS / 11 FAIL, judge agreement 20/21 (24/25 cumulative after re-runs).

Required before merge

Not tested

  • Browsers other than Chromium, and screen readers — the accessibility pass is automated only.
  • The full HF re-download of the other sites' asset archives in a cold checkout (they were already present). Berkeley's own archive was re-downloaded end to end at refs/pr/91 and its sha256 verified against the local tarball.
  • Per-site behaviour of the other 28 mirrors beyond the 29-site sweep.
  • The OCR and face passes are tripwires with measured blind spots, not guarantees — see Imagery. Small incidental fixture labels and small profile faces are known to be missed; the contact sheet is what covers that gap, and it is a manual step.

richard-peng-xia and others added 24 commits May 13, 2026 17:33
Adds a full Flask mirror of berkeley.edu as the 16th WebHarbor site.

**Site features:**
- 8 SQLAlchemy models: College, Department, Program, NewsArticle, Event,
  ResearchCenter, Faculty, Bookmark (+ User with auth)
- 20+ routes: homepage, news, programs, events, research centers,
  departments, faculty, admissions, about, unified search
- 23 Jinja2 templates styled with Berkeley Blue (#003262) / Gold (#FDB515)
- 30 benchmark tasks in tasks.jsonl (WebVoyager schema)

**Seed data (fully idempotent):**
- 14 UC Berkeley colleges/schools (real names)
- 83 degree programs (BA/BS/MA/MS/PhD/MBA/JD/MD/MEng)
- 121 news articles (2023–2025, 7 categories)
- 64 events (upcoming + past, 7 categories)
- 25 research centers (BAIR, QB3, MSRI, …)
- 82 faculty (Jennifer Doudna, Stuart Russell, Saul Perlmutter, …)
- 4 benchmark users: alice/bob/carol/dave (password: test1234)

**Infrastructure changes:**
- control_server.py: add 'berkeley' to SITES (port 40015)
- websyn_start.sh: add 'berkeley' to startup array
- Dockerfile: EXPOSE 40015, generate instance_seed DB at build time
  (no HF assets needed — all data is code-generated via seed_data.py)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…6 / port 40026, docs sweep

Merges upstream/main (26 sites @ f20b5ee) into the UC Berkeley PR branch. Berkeley is
appended as site index 26 rather than inserted at the PR's original index 15: appending
keeps merriam_webster..webmd_doctor on 40015-40025, so every merged site's tasks.jsonl
"web" field and the fixed-port assertions in walmart_careers/rotten_tomatoes tests stay
true. Inserting would have silently remapped 12 sites.

- websyn_start.sh: keep upstream's SITE_COUNT form, append berkeley last
- control_server.py: append 'berkeley' last (list == shell array, 27 entries)
- Dockerfile: keep every upstream RUN block; EXPOSE 8101 40000-40026; header 27 sites
- sites/berkeley/app.py: __main__ port from $PORT with 40026 default (was hard-coded 40015)
- sites/berkeley/tasks.jsonl: all 30 rows repointed to http://localhost:40026/
- docs: README/AGENTS/CONTRIBUTING/CLAUDE/agent_demo README + 5 skills swept
  40000-40025 -> 40000-40026 and "26 sites" -> "27 sites"; README mirror list gains
  UC Berkeley. agent_demo/README.md is included because walmart_careers' integration test
  asserts the current range appears in all five shared docs.
- .assets-revision: keep upstream's pinned sha ad6f424f (the PR branch had left it at "main")

Verification: step 1 (git status has no UU; author commit fa70f03 unchanged at the bottom of
upstream/main..HEAD), step 5 (walmart_careers test_shared_documentation_uses_the_current_site_range
now finds 40000-40026 in every doc it checks) and step 6 (no berkeley 40015 reference remains).

Co-Authored-By: Claude Code <noreply@anthropic.com>
- .build-generated-seed: declares instance_seed/berkeley.db a build artifact (this site
  has no HF archive), so check_assets.sh and build.sh stop requiring one.
- Dockerfile: generate the seed with the peers' exact shape —
  `rm -rf instance instance_seed && PYTHONHASHSEED=0 python seed_data.py && rm -rf instance` —
  instead of `python3 -c "from app import app"` plus a cp.
- seed_data.py: __main__ writes instance_seed/berkeley.db itself (build_seed_database), so
  the artifact comes from the documented command rather than an import side effect.
- seed_data.py: freeze the four benchmark password hashes (bcrypt salts are random, so
  set_password() gave every build a different DB) and pin created_at=datetime(2026, 5, 12)
  (the column default read the wall clock).
- app.py: guarded bootstrap_site() as in walmart_careers (WEBSYN_SKIP_BOOTSTRAP=1 suppresses
  it), still seeding on import so site_runner.py and the generator work unchanged.
- app.py: drop index=True from users.email / users.username. SQLAlchemy emits a table's
  named indexes in set-iteration order, so the two indexes were assigned different root
  pages from run to run (observed: pages 3/4 swapping) and the file was reproducible only
  by luck. unique=True keeps SQLite's implicit index, created in declaration order.

Verification: step 2 (md5 3001bcf4bcec169f4192c08609160ab6 across 5 scratch builds — 4x
PYTHONHASHSEED=0, 1x PYTHONHASHSEED=1 — and 3 further in-repo runs); step 3 bootstrap half
(empty instance/ seeds to the same md5; populated instance/ is a no-op; skip flag suppresses).

Co-Authored-By: Claude Code <noreply@anthropic.com>
sites/osu/app.py:27 pattern: BENCHMARK_NOW = datetime(2026, 5, 12), the same instant
seed_data.py pins the seeded event calendar to.

- app.py: BENCHMARK_NOW replaces every datetime.utcnow() on a request path — inject_globals,
  index()'s upcoming-events filter, events() (including its today_start/today_end window) and
  event_detail()'s related-events filter. Before this the seeded calendar (last event
  2026-07-16) was already behind the wall clock: / rendered an empty Upcoming Events section
  and /events answered "Showing 0 of 0 events", which made the event tasks unsolvable.
- app.py: the created_at / published_date column defaults now call a frozen utcnow() helper
  returning BENCHMARK_NOW, so no request path (register, bookmark_add) or seed path can stamp
  the wall clock into a row either.

Verified: seed md5 unchanged at 3001bcf4bcec169f4192c08609160ab6 (no seeded row ever relied
on the column default).

Verification: step 3 (all six endpoints 200; instance md5 still equals the seed md5 after the
request set; / renders 4 upcoming-event cards; /events default "Showing 20 of 52 events";
/events?category=Lecture "Showing 15 of 15").

Co-Authored-By: Claude Code <noreply@anthropic.com>
…rchive

A site whose seed is generated by the Dockerfile can legitimately have no archive at
all, not merely a media-only one — berkeley is the first such site. fetch_assets.sh was
the only release script still requiring an archive for every site directory (check_assets.sh,
build.sh and extract_assets.sh all honour .build-generated-seed), so a fresh clone died
before extracting anything:

    [fetch] scope: 27 registered site(s)
    fetch_assets: revision ad6f424f72cada9e6f5c09a58093d0ceeab9c52b has no archive for: berkeley
    exit 1

- single-site branch: after the download attempt, a missing archive on a
  .build-generated-seed site reports "nothing to fetch" and exits 0
- all-sites branch: such sites are skipped instead of added to `missing`
- an archive that does exist (fedex, webmd_doctor, ...) is still downloaded and validated;
  nothing changes for media-carrying build-generated sites

Verified: `./scripts/fetch_assets.sh berkeley` -> exit 0, no "expected archive";
`./scripts/fetch_assets.sh` -> exit 0, 26 sites extracted; `./scripts/check_assets.sh` -> exit 0.

Verification: step 4 (berkeley-only fetch exits 0 without the "expected archive" error, and
the full fetch + check_assets.sh succeed).

Co-Authored-By: Claude Code <noreply@anthropic.com>
GET /news/<slug> bumped news_articles.view_count and committed, so every article
visit wrote the DB. That broke two invariants: a read-only benchmark task could
never have an after-state equal to its initial snapshot, and instance/ diverged
from instance_seed as soon as an agent opened one article. The route is now a
pure read.

The column is kept and is still rendered as "N views" on /news and on the
article page from the frozen seed values; nothing orders or filters by it.

Verified: fresh-seed boot on :45011, three article detail pages all 200,
md5(instance/berkeley.db) == md5(instance_seed/berkeley.db) == 3001bcf4...,
news_articles rows (incl. view_count) identical before and after.

Co-Authored-By: Claude Code <noreply@anthropic.com>
sites/berkeley/verify/ now carries the grading contract for the 22 accepted
rows: the 20 kept ids (including the re-anchored --27) plus the new stateful
--30/--31.

- verify_lib.py: adapted from sites/webmd_doctor (newest merged), SITE=berkeley,
  no cross-site import, zero LLM calls on any verdict path. The snapshot
  contract is a pinned schema hash + nine-table set + seed counts + row-level
  catalog fingerprint; snapshots resolve from <run_dir>/initial.db|after.db,
  --initial_db/--after_db, or docker cp from $WH_CONTAINER, and missing or
  invalid input exits 1 with a structured infra_error instead of a traceback.
- ground_truth.py: every target is re-derived from the run's initial.db the way
  the app renders it (the BENCHMARK_NOW event filter, PER_PAGE, the app's
  ORDER BY clauses, the unordered LIMIT 3 related-centres query). The two
  source-rendered rows (--11, --17) parse tracked files and fail closed if the
  literals move. No answer constant is frozen anywhere.
- 22 verifiers: any-step URL gates plus the final action's declared target as
  an alternative satisfier; every multi-hop task gates each hop; set-valued
  answers use derived accepted sets and minimum counts; matching is
  negation-aware; --30/--31 bind to an exact bookmark row delta, and --31 also
  pins the surviving row id (2) as proof that both inserts and the removal
  happened.
- verify/tests: 516 tests. Stdlib-only fixture DBs reproduce the pinned
  fingerprint; trajectories follow the agent_demo/agent.py signature. Per task:
  genuine PASS, no-op, wrong task id, another task's trajectory, shortcuts,
  1-2 wrong answers, alternative phrasings, negated answers, truncated run,
  corrupt and 1x1 PNGs, missing after.db, catalog/schema drift, collateral
  writes and state mismatch, plus the About/Admissions source-fact assertions.

Run: .venv/bin/python -m pytest sites/berkeley/verify/tests -q  -> 516 passed

Co-Authored-By: Claude Code <noreply@anthropic.com>
- tasks.jsonl: 22 rows — the 20 kept ids (with --27 re-anchored onto the two
  exact programme durations) plus the new stateful --30/--31. The 19 unchanged
  rows keep the contributor's exact ques text; every row gains verifier_path
  and a rules-only judge_rubric (shared scoring preamble + per-task
  checkpoints). "web" stays http://localhost:40026/.
- verify/TASK_REVIEW.md: all 32 rows (30 contributor + 2 reviewer) with
  ACCEPT / DROP / ADDED and the workflow each accepted row must follow, plus
  the corrections found while deriving the targets (frozen-clock Lecture count,
  the AI-family allowlist for --7, BIDS' four focus areas and its rendered
  related-centre set, and the article-view write removal).
- verify/README.md: the contract per task, the snapshot rules, the matcher
  semantics, and how to run the verifiers and their tests.
- verify/tests/test_tasks_contract.py: validates the file (22 rows, seven keys,
  existing verifier paths, unique rubrics) and re-derives every target to prove
  no rubric carries a ground-truth value — with a self-test that the leak scan
  actually fires on a planted answer.

The image already excludes verify/tests/ via the existing .dockerignore pattern
(sites/*/verify/tests/), the same rule the merged peers rely on; no change was
needed there.

Run: .venv/bin/python -m pytest sites/berkeley/verify/tests -q  -> 524 passed

Co-Authored-By: Claude Code <noreply@anthropic.com>
verify/tests/run_matrix.py boots the mirror from a fresh seed per cell on an
alt port, drives a scripted Playwright workflow, snapshots the live database
and writes an agent_demo/agent.py-shaped run directory (trajectory.json with
url-before-action steps and one input step per filled field, screenshots/,
initial.db, after.db). It then grades every cell through
`uv run python agent_demo/eval_judge.py --run_dir <dir> --verifier True` and
compares the verdict with the cell's expectation.

Cells per task: pass (genuine walk), no_op, shortcut (catalog-wide search with
the correct answer), wrong_answer, collateral_write (one row injected straight
into the live DB after the walk) and, for the two stateful rows,
state_mismatch (the save skipped). Genuine answers are rendered from the
derived target, so the harness carries no second copy of the ground truth.
Artifacts land under sites/berkeley/scripts_dev/runs/matrix (gitignored, and
self-ignored by a generated .gitignore in the output root).

The matrix itself was NOT run in this window: that is the next one. What the
test suite runs is the browser-free replay contract
(verify/tests/test_run_matrix_contract.py), which replays each workflow as a
synthetic trajectory and asserts the genuine run passes every verifier, every
wrong answer is rejected and the injected collateral write fails.

Run: .venv/bin/python -m pytest sites/berkeley/verify/tests -q  -> 528 passed

Co-Authored-By: Claude Code <noreply@anthropic.com>
- sites/berkeley/README.md: scope and routes, seed generation (build-generated
  from tracked source, byte-reproducible, md5 recorded), the frozen
  BENCHMARK_NOW clock, why the mirror ships no imagery, the seeded row counts,
  the demo accounts, and a pointer to the grading contract.
- sites/berkeley/tests/test_integration.py (walmart_careers pattern): the
  registry is derived from control_server.SITES (cross-checked against
  websyn_start.sh) and berkeley is asserted at index 26 / port 40026; the
  Dockerfile exposes the current range and builds the seed from source; the
  .build-generated-seed marker is honoured by fetch_assets.sh; the seed's md5 is
  asserted when the build-generated DB is present; tasks.jsonl has 22 rows on
  the registered port with existing verifier paths and no answer key; app.py
  keeps the frozen clock and its article route neither writes view_count nor
  commits; the five shared docs carry the current port range.

Run: .venv/bin/python -m pytest sites/berkeley/tests -q  -> 7 passed

Co-Authored-By: Claude Code <noreply@anthropic.com>
…der contract

The first matrix run exposed four writer defects; each is fixed here, and the
fix was mutation-checked (the port guard was shown to fire on a stray server,
and the 30/31 pass/state_mismatch cells now grade correctly):

1. grade() launched eval_judge.py from the repo root, whose .venv has neither
   openai nor simpleArgParser -> ModuleNotFoundError and no eval.json, so every
   cell would have been reported as a mismatch with pass=None. It now runs from
   agent_demo/, the invocation AGENTS.md documents.
2. The page was never navigated before the first recorded step, so step 0's URL
   was about:blank and the verifier's all_urls_match_local_origin gate failed
   the genuine run of every task. It now pre-navigates to the start URL exactly
   as agent.py's browser.navigate_to(start_url) does.
3. Both login workflows clicked "button[type=submit]", which matches the navbar
   search button first: the login never happened, /account bounced to /login and
   the bookmark step timed out. Now form[action='/login'] button[type=submit].
4. boot() adopted any process already listening on the port. A leftover
   standalone server made cells grade against a foreign instance (the 30/31 pass
   cells failed with bookmarks_exact_delta). boot() now refuses that port.
   Also the state-mismatch cell (all writes skipped) clicked the bookmark-removal
   form that a fresh instance never renders; the skip rule now applies to any
   step marked skip_in_state_mismatch, and the removal click carries the mark.

Co-Authored-By: Claude Code <noreply@anthropic.com>
…by "Prof."

The C2 mutation rows found that a contradictory answer passed three verifiers:
"The chair of EECS is not Prof. James Demmel, ...", "BIDS is not directed by
Prof. David Culler, ..." and "The Economics department is not chaired by Prof.
Ulrike Malmendier, ..." were all graded as affirmative. Cause: the clause
splitter treated the period in an honorific as a sentence end, so the negation
that precedes the name landed in a previous clause and _match_is_affirmative
never saw it. The plain forms without the title ("James Demmel is not the
chair of EECS.") were already rejected, which is why the unit suite missed it.

_match_is_affirmative now computes clause boundaries over a length-preserving
mask of honorific / degree abbreviations (Prof. Dr. Mr. Mrs. Ms. Miss Mx Rev.
Fr. Sr. Jr. and "Ph.D."), so the two fixes compose: the name matcher sees the
negation in front of it wherever the rendering carries a title.

Mutation-checked: with the mask disabled the four new tests fail and only
those (4 failed, 103 passed); with it enabled the suite is 531 passed. Rows
re-run with the full variant + mutation sets; the matrix is re-run in the
same phase.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Rewrite judge_rubric for all 22 accepted tasks (rubric-only diff; every other key unchanged, no row added or removed). Rubrics open with the scoring rules (step list authoritative; a checkpoint is true unless positively contradicted; a detail page is not satisfied by a listing page; every step must be on the local mirror origin; the verifier owns exact values and DB state; an empty answer forces failure) and close with an explicit origin checkpoint. No rubric contains a derived answer value; the task-contract leak scan is clean.

Rebuilt from D3/D4 evidence: 16/21 -> 19/21 verifier-judge agreement on the shim arm; D5 (tasks 24, 27) run after this rewrite leaves one reproduced divergence (verify_27.py route gate, NOT-FIXED).

Co-Authored-By: Claude Code <noreply@anthropic.com>
… bounds, FK

Appendix A §6 merge blockers found by the Phase E executable probes:

- SECRET_KEY was the committed literal 'berkeley-mirror-secret-key-2024', so a
  cookie signed with it read /account without a password (probe: 200). It now
  comes from BERKELEY_SECRET_KEY or a per-process random key, and a wrong-key
  cookie is bounced to /login.
- GET/HEAD /logout returned 302 and cleared the session (prefetcher-reachable);
  the route is POST-only now (GET/HEAD -> 405) and the site chrome's Sign Out
  control is a CSRF-protected form instead of a link.
- A session cookie with a non-numeric _user_id raised int() into a 500;
  load_user now fails closed (anonymous).
- Huge query/form integers overflowed SQLite ('/news?page=9'*20 -> 500);
  bounded_int/page_arg and the event-id range check answer 200/404 instead.
- MAX_CONTENT_LENGTH (256 KB) and SESSION_COOKIE_HTTPONLY/SAMESITE are set.
- Empty or invalid /bookmark/add submissions silently redirected; they now
  answer 400 (closed item-type vocabulary) or 404 (missing row), and
  /bookmark/remove is owner-scoped (another user's id is a 404) with a bounded
  id.
- '?next=https://evil.example/' bounced /login and /bookmark/add off-mirror;
  safe_next keeps redirect targets same-origin.
- SQLite PRAGMA foreign_keys was off (orphan bookmark rows accepted); an
  Engine connect listener turns enforcement on.
- Duplicate registration now rolls back on IntegrityError.

Accessibility chrome on every page (the maintainers' b87db8f batch): skip link,
mirror/synthetic-data notice, focus-visible outlines, role=alert on flash
messages. sites/berkeley/tests/test_app_robustness.py pins all of it and was
mutation-checked (mutations archived under scripts_dev/logs/phase_e/).

Co-Authored-By: Claude Code <noreply@anthropic.com>
…ry route

The Phase E leak sweep (new sites/berkeley/tests/test_answer_leaks.py, adapted
from the maintainers' webmd_doctor sweep) found that listing cards rendered
the exact fields the tasks must discover on a detail page:

- /research, the home page and the /search results rendered "Director: X",
  "Founded: Y" and the first focus-area tags for every centre, so tasks 10,
  23, 30 and 31 had their answers visible before /research/<slug>.
- related-centre cards on another centre's page did the same for neighbouring
  centres.
- /departments cards rendered "Chair: X" and the location, so tasks 13 and 24
  had their answers visible before /departments/<slug>.

The listing cards now carry name/college/description only; the detail pages
still render director, founding year, focus areas, chair and location (pinned
by positive controls in the test). The test enumerates each task's answer
facts from verify/ground_truth.py, exempts only the routes its verifier
requires, and documents every shared-value co-occurrence with its reason; it
was mutation-verified by re-adding each removed field (5/5 mutations fail).

Co-Authored-By: Claude Code <noreply@anthropic.com>
…ng its target

/programs?page=3 lists programmes 41-60; the Data Science MS is the 24th row,
i.e. on page 2 (the run then opened the detail page by URL, so the listing hop
never showed the target). Replay fixtures are data-driven and unaffected.

Co-Authored-By: Claude Code <noreply@anthropic.com>
… parity

Appendix A §4:
- /search matched narrower field sets than the catalogue pages it duplicates:
  news content, event location/organizer and faculty title were reachable from
  the listing filters but not from the global search. All catalogue fields are
  now matched (superset direction only; no result set shrinks).
- The related-centre / related-programme / colleague / department-programme /
  home-page research queries had no ORDER BY and fell back to rowid insert
  order; each now has a neutral key (name) and ground_truth.related_centres
  mirrors the same ORDER BY. /search results are ordered by the same neutral
  keys as their listings.
- test_verify_23's fixture now names the first entry of the ORDER BY name list.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Appendix A §9 (4 widths x all routes, pixel-sampled contrast):

- Contrast tokens: --gray-mid #6c757d -> #64696f (4.38 -> 5.17 on the off-white
  card ground), --light-blue #3B7EA1 -> #2E6C8B (4.48 -> 5.78 on white),
  .badge-green #28a745 -> #1E7E34, gold-button hover #C4820A -> #E0A50C
  (blue text 4.01 -> 5.85), the 404 numeral gold -> --dark-gold (1.78 -> 3.56),
  the hero gradient's light stop darkened and its 0.9 opacity wash removed, the
  card-image label made solid white, and the event/news category palettes
  darkened (Sports #8A5A07, Career #1E7E34, Health #0F7A5A, Arts #5A3383,
  Virtual #55595E; the in-badge white 0.2 wash -> black 0.3).
  Measured by glyph-anchored pixel sampling (scripts_dev/phase_e_contrast.py):
  254 failing text elements before, 0 of 2725 after at 1440, 0 at 768/390/320.
- Layout: fixed 4/2-column and 3fr/2fr grids became responsive helpers that
  collapse at <=768px (about, academics, admissions, department/program detail,
  faculty profile, research, account, home hero); long unbreakable strings get
  `min-width: 0` / `overflow-wrap: anywhere`; the filter-bar controls shrink;
  the faculty profile header wraps. 320px horizontal scroll is gone.
- Heading order: footer and card headings no longer skip levels, listing pages
  carry an sr-only h2, detail-page section headings were re-levelled; the
  heading-order check is 0 jumps over 423 routes.
- run_matrix/site docs: the related-centres query is ORDER BY name now
  (mirrored in ground_truth and TASK_REVIEW).

Co-Authored-By: Claude Code <noreply@anthropic.com>
…op bar

The featured-article badge is position:absolute with no positioned ancestor:
.card-img was static, so the badge anchored to the page and painted over the
top-bar Sign In / Create Account links (measured overlaps at 768/390/320 px).
position: relative on .card-img keeps it inside the card image; verified in
Chromium at 4 widths (in_card 4/4, overlapped 0).

Co-Authored-By: Claude Code <noreply@anthropic.com>
…ADME claim

seed_data.py now refuses to build instance_seed/berkeley.db when the app
engine's URI is not the canonical BASE_DIR path, before deleting or creating
anything (the maintainers' §8 wound: a redirected URI silently writes the seed
elsewhere while the copy still reads DB_PATH). Baseline reproduces md5
3001bcf4...; the scratch-copy mutation with a redirected URI refuses and writes
nothing.

verify/README.md claimed 19/24/25/30/31 'gate each hop in order'; only 24/30/31
call check_paths_in_order — 19/25 gate both hops independently. Wording fixed.

Co-Authored-By: Claude Code <noreply@anthropic.com>
…he navy band

Measured: the #0b5cab focus outline is 6.7:1 on white but only 1.92:1 where it
crosses the navy top bar / header (skip link, header search input, search
button). A lone ring cannot contrast with both white and navy, so the ring is
now two-tone: the blue outline plus a white halo (12.86:1 on navy, 6.7:1 blue on
white, blue on gold 3.77:1). Verified by pixel measurement of the painted ring
around each of the first six tab stops.

Co-Authored-By: Claude Code <noreply@anthropic.com>
…es (§0.4)

verify_27.py gated "visited_program_search" on the programme catalogue only
(/programs?q~'master of engineering' OR /programs?degree~'MEng'), while the ques
says "Search the Berkeley site for 'Master of Engineering'". Both task-27 runs
with real answers (27shim, 27shimd5b) searched /search?q=Master+of+Engineering,
opened both detail pages and gave the right department and both durations; every
other check passed and only this gate failed (observed_urls=[]).

PIPELINE §0.4 conditions: (a) required to accept a genuinely correct run —
proved check by check; (b) the loosening keeps the anti-shortcut property —
re-grading task 27's full variant and mutation sets gives the identical verdict
and identical first failing check for all 13 rows (pass PASS; no_op
final_answer_nonempty; shortcut visited_program_search; wrong_answer
answer_has_meng_duration; collateral_write read_only_bookmarks_unchanged;
tiny_png screenshots_decode; catalog_search visited_program_search; truncated
trajectory_completed; other_task trajectory_task_matches; missing_after_db
database_unavailable; read_only_write read_only_bookmarks_unchanged;
negated_answer answer_has_department; answer_only visited_program_search);
(c) logged in DECISIONS.md.

The gate now also accepts /search?q~'master of engineering'. The query must name
the MEng term, so a catalog-wide search still fails. Both detail-page gates and
every answer check are unchanged and remain the binding anchors. 3 regression
tests added (site-search accepted; catalog-wide search still rejected; search
alone without both detail pages still rejected).

Suites: verify/tests 534 passed; site tests 27 passed; no-op container matrix 22
rows, 0 not FAILing, 0 infra errors.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Two test-file docstrings pointed readers at paths that are not part of this
branch: review-reports/berkeley/CHECKLIST_REPORT.md §6 and
scripts_dev/logs/phase_e/ (test_app_robustness.py), and
scripts_dev/REVIEW_STATUS.md §7.2 (test_integration.py). The prose keeps its
meaning — probes failed on the pre-fix tree, each detector was mutation-checked
— without naming artifacts that do not ship.

No functional change. verify/tests/run_matrix.py was deliberately left alone:
its four `scripts_dev`/`runs/` occurrences are the harness's own scratch
--out default (a path it creates on demand and self-ignores), not references
to review evidence.

Co-Authored-By: Claude Code <noreply@anthropic.com>
… into the UC Berkeley branch

Registry: append `berkeley` after `kaggle`, so Healthline keeps index 26 / port
40026 and Kaggle keeps index 27 / port 40027; UC Berkeley moves to index 28 /
port 40028 (29 sites, 40000-40028).

Conflict resolutions (append rule):
- `websyn_start.sh`, `control_server.py`: keep all of upstream's 28 entries in
  order and append `berkeley` last (both registries identical, 29 entries).
- `Dockerfile`: keep every upstream site block (Healthline's migrate + prune step
  included) and berkeley's build-generated seed block; header 28 -> 29 sites;
  `EXPOSE 8101 40000-40027` -> `40000-40028`.
- `README.md`, `AGENTS.md`, `CONTRIBUTING.md`, `CLAUDE.md`, `agent_demo/README.md`,
  `.claude/skills/*`: take upstream's text, then 29 sites, port range
  40000-40028, alt ports 41000-41028, and `UC Berkeley` appended to README's
  mirror list.
- `scripts/fetch_assets.sh`, `scripts/check_assets.sh`: take upstream's
  registry-scoped implementations and keep berkeley's `.build-generated-seed`
  exemption on top.

Follow-on work required by the port move:
- `sites/berkeley/tasks.jsonl` (22 `web` rows), `app.py` PORT default,
  `README.md`, `verify/TASK_REVIEW.md`, `verify/verify_lib.py` comment,
  `tests/test_integration.py` (SITE_INDEX 28 / SITE_PORT 40028) and
  `verify/tests/test_tasks_contract.py`: port 40028.
- `scripts/check_site_registry.py` (upstream's new gate) reports 29 sites
  consistent across both registries, `Dockerfile EXPOSE` and every
  `tasks.jsonl` port.

Verified after the merge: registry gate green; berkeley site 27, berkeley verify
534, healthline 27; walmart_careers 50 passed and rotten_tomatoes 55 passed
(+4129 subtests) each with one expected red that is pre-existing on upstream/main
and reproduced there from a pristine `git archive` tree (VERIFICATION.md
§re-slot); kaggle ships no pytest suite. Container rebuilt and re-run on
44000-44028: 29/29 sites 200, /health 29 alive+ready, `POST /reset/berkeley`
byte-identical (`f2f0187c...`) before and after `docker restart`, no-op matrix
22/22 FAIL with 0 infra errors.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@evanz37
evanz37 marked this pull request as ready for review September 13, 2026 15:58
…tic avatars), asset inventory and gate

The mirror shipped no images at all: static/ held two .gitkeep files, base.html
carried one inline stylesheet, and every card image was a flat colour plus a text
label. That is a real fidelity gap for a mirror of berkeley.edu, where every
section page leads with a photograph, and the README called it "by design".

164 files, 6,991,525 bytes (6.67 MiB) on disk, all declared in
generated_asset_inventory.json and gated by check_generated_assets.py:

  campus    8   1024x768 WebP   section heroes and page-header banners
  college  14   1024x768 WebP   one per college; odd ids exterior, even interior
  research 25   1024x768 WebP   one per research centre
  news     21   1024x768 WebP   7 categories x 3 variants, variant = id % 3 + 1
  event    14   1024x768 WebP   7 categories x 2 variants, variant = id % 2 + 1
  faculty  82    256x256 PNG    deterministic Pillow initials avatars

Scenes come from fal.ai FLUX.1 [schnell] (4 inference steps, one fixed seed per
slot derived from the slug); the avatars are drawn offline by Pillow 11.0.0 from
the faculty row's initials and id, and regenerate byte-identically. Variants bind
to the row's primary key, never to a render index, so a listing card and a detail
banner agree on the same file and pagination cannot reselect one. The plan and the
generators live in sites/berkeley/scripts/ (the tracked location
sites/compass/scripts/ uses), not in scripts_dev/, which stays a local scratch
directory.

Prompt policy: every prompt is one subject clause plus a fixed style suffix plus a
fixed negative list — no faces, no portraits, no text, no lettering, no logos, no
watermarks, no posters, no framed pictures, no screens with visible content, no
recognizable landmarks or signage, unlabeled containers. The exact prompt is
recorded per file. Slots are occupancy-classified: labs, research interiors, news
and event venues are prompted unoccupied; campus and college-exterior slots permit
at most two or three figures in the middle distance, backs turned, faces hidden.
No image depicts a real person, a real face or a real landmark — the campus scenes
are generic institutional architecture, deliberately not the Campanile, Sather
Gate or the Golden Gate, and the avatars are monogram discs.

Answer-leak safety: every alt is built from the same seed fields the page already
renders — college and centre names, category only for news and events — never a
title, director, founding year or focus area, which are the graded answers on the
detail pages. test_answer_leaks passes unchanged, including the two entity-bound
assertions that inspect a +-700-character window around every /research/<slug>
link, which now contains the new card image.

Verification:
- 569 tests pass: sites/berkeley/tests (answer leaks, app robustness, integration,
  generated assets) and sites/berkeley/verify/tests.
- check_generated_assets.py verified 164 assets at build time: exact coverage,
  per-file SHA-256, decode at the planned dimensions, letterbox test.
- OCR pass and manual face review. The OCR pass runs over every scene and has a
  positive control that proves it can fire on rendered text; it is NOT a text
  guarantee. It returned zero tokens for the word STCK rendered in large red
  capitals across a window in colleges/chemistry.webp and missed a placard in
  research/bair.webp entirely, and it was flagging clean scenes on three-character
  junk (aif, hea, saks) until the confidence floor was raised to 4 chars / 50.
  The face pass is a frontal-face cascade, not a person detector: the OpenCV
  defaults gave 15 false positives on five clean pilot scenes (foliage, mown
  grass) and were replaced with settings that score zero there and detect a
  generated face control — which is exactly why it also misses the small, profile
  faces the rule cares about. scripts/IMAGE_PLAN.md §9 records every measurement
  and every miss.
- The by-eye review is what actually bounds this, across five passes: a seated
  person in a research interior; crowds with faces to camera in the event family
  (which moved to unoccupied prompts — an empty hall still reads as a venue); a
  framed portrait in a faculty-office scene (fixed with anti-wall-art clauses on
  the 24 office and reading-room motifs that invite one); three letterboxed frames
  (now a checked defect class, self-healing via is_accepted); and four slots
  reported on the review contact sheet — chemistry (STCK), the newsroom
  (watermark-like label), journalism (a face to camera) and law (faces) — each
  regenerated on a new seed to its instruction: law unoccupied, journalism
  back-view figures only. The contact sheet is
  gen_images.py --contact-sheet, default scripts_dev/contact_sheet.png.
- §9 UI sweep at 1440/768/390/320 over 424 routes: no overflow, no broken images,
  no console errors, no aria or heading-order findings. The only findings are
  net::ERR_ABORTED on lazy images cancelled by navigation, with no 4xx and no
  external requests.
- Contrast is arithmetic, not measured. The pixel-sampling sweep reported no
  measured failure before or after this change, but every finding on both sides
  carries the sampler's "could not isolate a background" signature
  (bg_used == sampled, ratio 1.0) — it went blind exactly where the photos are, so
  it cannot corroborate photo-backed text. What the claim rests on is the computed
  worst case against a pure-white pixel behind: hero scrim white 8.6:1, the 18px
  paragraph 7.3:1, the gold eyebrow 4.8:1, page-header 8.3:1, and .card-img-label
  given its own scrim (>= 8:1). The gold eyebrow was already 3.25:1 on the old
  opaque gradient, so this improves it rather than fixing a regression.
- Container: 29/29 sites HTTP 200; /health 29 alive + ready; POST /reset/berkeley
  byte-identical (f2f0187c... == instance_seed) before, after reset and after
  docker restart; no-op matrix 22/22 FAIL, 0 mismatches; broken-image sweep over
  424 routes x 2 widths inside the image, 0 broken.

Assets. The bundle ships through the pinned Hugging Face tarball, not git:
berkeley.tar.gz, 6,951,483 bytes, sha256
ab9d2716ae8d06540a181b5e60c37f613d87b103864b467511da546b1b173789, 171 managed
members, validated by scripts/validate_asset_archive.py. Uploaded as
https://huggingface.co/datasets/ChilleD/WebHarbor/discussions/91 (head commit
4529b18c6fc23f9b203a281b2d1e1940ba93853b), and .assets-revision pins
`revision: refs/pr/91` as an explicitly INTERIM pin — a PR ref, so it moves if the
PR branch is updated. Maintainers should re-pin `revision:` to the merged commit
sha once PR aiming-lab#91 is merged, and delete the interim paragraph; the
"never a moving branch name" rule in that file then applies again.

86ba06e's fetch_assets.sh build-generated exemption is no longer exercised by
berkeley: the archive now exists, so the download branch is taken and the
exemption remains only as the general fallback for a build-generated site with no
archive at the pinned revision. Nothing needs removing; this records that berkeley
stopped being the case that motivated it.

run_matrix.py gains --cells so the no-op matrix can be run on its own; that is the
only non-imagery change here, and it is additive.

Co-Authored-By: Claude Code <noreply@anthropic.com>
@evanz37

evanz37 commented Sep 15, 2026

Copy link
Copy Markdown
Contributor Author

Imagery added in 19f7322: 82 synthetic FLUX.1 [schnell] scenes (no faces, no text, no logos, no real landmarks; back-view figures only) plus 82 deterministic Pillow avatars, wired into every card and detail banner under the webmd_doctor generated-asset contract (inventory, build-time gate, tests). Asset bundle is HF PR https://huggingface.co/datasets/ChilleD/WebHarbor/discussions/91; .assets-revision pins refs/pr/91 as an interim ref — please re-pin to the merged sha after merging that PR, as with webmd_doctor. Also rebased onto main after #105/#106 (berkeley now index 28 / port 40028).

…ey to index 29 / 40029 after aiming-lab#107 took index 28

Resolution of the conflicts between this branch and current main (aiming-lab#107 merged NVIDIA
as index 28):

- websyn_start.sh, control_server.py: keep main's 29-site order (nvidia at index 28)
  and append berkeley as index 29.
- Dockerfile: keep main's per-site build steps and add berkeley's generated-asset
  gate (check_generated_assets.py) plus its build-generated seed step; the site-count
  comment is now "30 Flask mirror sites" and EXPOSE 8101 40000-40029.
- .assets-revision: pin stays this branch's refs/pr/91 (berkeley.tar.gz's only
  revision). Checked while merging: for all 29 sites registered before berkeley, the
  archive size and LFS oid on refs/pr/91 are identical to the previous b7e605c0 pin,
  so the pin change alters no other site's assets. The comment block keeps both pins'
  documentation without contradicting itself.
- sites/berkeley/**: port contexts 40028 -> 40029 (tasks.jsonl web URLs, app.py
  default PORT, README, tests, verify comment).
- Documentation registries (README.md, AGENTS.md, CLAUDE.md, CONTRIBUTING.md,
  .claude/skills/**): main's side is taken verbatim here; the 30-site / 40000-40029 /
  41000-41029 statements follow in the next commit.
The registry grew from 29 to 30 sites when UC Berkeley was re-slotted behind
aiming-lab#107's NVIDIA (index 28). Current-state statements in AGENTS.md, CLAUDE.md,
CONTRIBUTING.md, README.md, agent_demo/README.md and the five .claude/skills
files now say 30 sites, container ports 40000-40029 and alt ports 41000-41029.

README.md also gains the UC Berkeley registry row (index 29, container port
40029, local review host port 48029) and an asset-delivery section that
describes the pin this branch actually ships (refs/pr/91) including the
berkeley.tar.gz artifact and the verified equivalence of every other site's
archive with the previous b7e605c0 pin.

review-reports/** is left untouched as historical record; templates and
placeholders are unchanged.
HF dataset PR aiming-lab#91 ("berkeley: synthetic imagery bundle (164 files)") is merged, so
the interim `refs/pr/91` pin is replaced by the immutable head commit of the
dataset's `main`, c32018ca3b3d67e7b858b1b85fb101aea5090cd7 (merged
2026-09-15T04:00:54Z). A `refs/pr/<n>` ref moves if its branch is updated; the
merged commit does not, so this removes the drift the PR ref carried.

Verified while changing the pin:

- `berkeley.tar.gz` on c32018ca is 6951483 bytes, sha256
  ab9d2716ae8d06540a181b5e60c37f613d87b103864b467511da546b1b173789, byte-identical
  (cmp) to the archive the `refs/pr/91` pin served, and passes
  `scripts/validate_asset_archive.py` (171 managed members, exit 0);
- all 30 registered site archives on c32018ca have the same size and LFS oid as on
  `refs/pr/91` (30/30 identical, 0 differences), and the same two unregistered
  archives (bandcamp.tar.gz, drugs_com.tar.gz) are present;
- a clean re-fetch from the new pin reproduces sites/berkeley/static/images (164
  files) and instance_seed/berkeley.db byte for byte.

`.assets-revision`'s comment block now records the merged pin, keeps the previous
pins as history, and no longer describes the revision as an interim PR ref.
README.md's asset-delivery section states the new pin and the merge.
@Raibows

Raibows commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Thanks for your contribution @evanz37 @richard-peng-xia

@Raibows
Raibows merged commit aa95bd8 into aiming-lab:main Sep 15, 2026
Raibows added a commit to jackjin1997/WebHarbor that referenced this pull request Sep 17, 2026
…ex 30 / 40030 after aiming-lab#116 took index 29

Conflict resolution against current main (aiming-lab#116 merged: nvidia index 28, UC Berkeley
index 29, 30 sites):

- websyn_start.sh, control_server.py: keep main's 30-site order and append imdb as
  index 30 (container port 40030).
- Dockerfile: keep main's build steps unchanged (the PR adds no build step for imdb,
  whose seed ships in the HF archive) and set the canonical header/EXPOSE lines:
  "31 Flask mirror sites" / "EXPOSE 8101 40000-40030". The merged file differs from
  main's by exactly those two lines.
- sites/imdb/tasks.jsonl: 20 web URLs 40024 -> 40030. No task semantics changed
  (ques/judge_rubric/verifier_path untouched); sites/imdb/app.py keeps its
  PORT-env default of 5000, so no other file in the site referenced 40024.
- nvidia (40028) and berkeley (40029) task URLs are unchanged from main.
- .assets-revision: kept as the PR's value (revision:
  f9ddfd2596229f2610418d57fc88c3051e1056bb, the head commit of HF dataset PR aiming-lab#57).
  This pin is NOT usable after this merge: refs/pr/57 carries 26 archives and has no
  archive for fedex, webmd_doctor, healthline, kaggle, nvidia or berkeley, and
  c32018ca (the pin main uses) has no imdb.tar.gz. The pin will be replaced with the
  merged HF dataset commit once HF PR aiming-lab#57 is merged; this commit does not change it.
- Documentation (README.md, AGENTS.md, CLAUDE.md, CONTRIBUTING.md) took main's side
  here; the 31-site / 40000-40030 statements follow in the next commit.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants