Skip to content

Stage 3W dataset - #35

Merged
oskrgab merged 12 commits into
mainfrom
stage
Jun 4, 2026
Merged

oskrgab merged 12 commits into
mainfrom
stage

Conversation

@oskrgab

@oskrgab oskrgab commented Jun 4, 2026

Copy link
Copy Markdown
Owner

No description provided.

oskrgab and others added 12 commits May 17, 2026 18:54
…pinning ADR

Captures the design decisions reached while grilling issue #17 / PRD #18:
- CONTEXT.md adds the 3W language (Well, Instance, Well kind, Event class,
  Transient label, Observation, Source file), cross-dataset ambiguity entries
  for Well and Event, and four 3W dataset sections (tables, operating
  principles, output layout, pre-publish validation).
- ADR-0001 records the observations-layout decision (one file per Instance,
  hive on event_class) with the four rejected alternatives.
- ADR-0002 records the upstream-tag-pinning policy (v2.0.0, event-driven
  refresh) and why tracking main would silently mutate already-published bytes.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Tracer-bullet slice that establishes the transform/export module layout
for the Petrobras 3W dataset and emits its smallest deliverable table.
Subsequent slices (#20, #21, #22) add the Instance catalog, real-Well
master, and Observations hive partition without re-touching this
scaffolding.

Key decisions:
- Source of truth for event_types is upstream `dataset.ini`. The
  pipeline parses it via configparser (with `optionxform=str` to
  preserve hyphenated sensor column names like `P-PDG` and `QGL`)
  rather than hard-coding canonical values — a future upstream rename
  surfaces as a parse-time change rather than silent drift.
- The shallow-clone is idempotent: presence of
  `<staging>/dataset/dataset.ini` short-circuits the `git clone`, so
  the smoke test can stand in a fixture upstream tree without
  network access.
- Pinned identity (git tag `v.1.70.0`, dataset version `2.0.0`) is
  emitted via the export's `petrobras_3w.export` logger before any
  parquet write, and is also recorded in `schema.json`, `schema.sql`,
  `README.md`, and the static-site tab. A parsed `dataset_version`
  that disagrees with the pin aborts publish (ADR-0002 event-driven
  refresh).
- `has_normal_prefix` is materialised equal to `has_transient` per
  CONTEXT.md (events 0, 3, 4 carry only the steady class — no NORMAL
  precursor).
- DDL identifier quoting in `schema.sql` is general (`_quote_identifier`)
  so the hyphenated sensor columns added in #22 round-trip without
  re-touching the writer.
- Website integration uses the same sentinel-comment idempotency
  pattern as the Argentina integrator.

Files added:
- `scripts/transform/petrobras_3w/` — constants, upstream stager + ini
  parser, event_types builder, orchestrator.
- `scripts/export/petrobras_3w/` — parquet writer, validator (logs the
  pin + asserts dataset version + event_types row count/PK/transient
  invariants), schema doc generator (md/json/sql + README + LICENSE),
  website integrator, orchestrator.
- `parquet/petrobras_3w/` — published deliverables: `event_types.parquet`
  (10 rows), `schema.md` (with 27-sensor glossary mirrored from
  upstream), `schema.json`, `schema.sql`, `README.md`,
  `LICENSE-3W-DATA.md` (CC BY 4.0 + attribution).
- `tests/petrobras_3w/test_smoke.py` — end-to-end coverage of the
  scaffolding (event_types contents, doc generation, website
  idempotency, pin logging, validator abort paths).
- `tests/fixtures/petrobras_3w/dataset/dataset.ini` — byte-identical
  mirror of upstream at the pinned tag, so tests run without a clone.

Files modified:
- Root `README.md` and `parquet/index.html` patched via the new
  website integrator (new Petrobras 3W entry / tab parallel to
  Argentina, Volve, FORCE 2020).
- `.gitignore` — exclude the `data/petrobras_3w/` staging dir and the
  intermediate `database/petrobras_3w.duckdb` file.

Notes for next iteration:
- #20 (instances.parquet), #21 (wells.parquet), #22 (observations/) all
  extend `scripts/transform/petrobras_3w/orchestrator.py` and
  `scripts/export/petrobras_3w/{parquet_writer,validator,schema_doc_generator}.py`.
- The validator's structural checks (#19 only covers event_types
  row-count, PK uniqueness, has_transient/transient_code invariants
  + pin assertion) need to be extended with the remaining seven rules
  from CONTEXT.md as later tables land.
- Pre-existing test failure in `scripts/export/test_integration.py`
  (Volve-era, missing `parquet/wells.parquet`) is unrelated to this
  slice and remains as-is.

Closes #19

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Tracer-bullet extension of the #19 scaffolding that emits the full
Instance catalog. One row per upstream per-Instance parquet, materialised
via pure DuckDB SQL: a single `read_parquet(..., filename=true)` pulls
every row, the filename is parsed into `instance_id` / `well_kind` /
`well_id`, and the per-file aggregates (`start_ts`, `end_ts`,
`duration_s`, `n_rows`, four `n_rows_*` counts, `source_url`) are
computed in the same query.

Key decisions:
- `n_rows_normal` is defined as `class = 0 AND event_class <> 0` (the
  NORMAL precursor before an anomaly). Strict literal reading of
  CONTEXT.md ("rows where `class = 0`") would double-count event 0,
  where `class = 0` and `class = event_class` overlap. The refined
  definition makes the four `n_rows_*` columns always partition `n_rows`
  without overlap, which is the operational requirement stated in
  acceptance criterion 5. CONTEXT.md, the column doc, and the validator
  rule are all aligned on this.
- `n_rows_transient` is materialised NULL exactly when the event's
  `has_transient = false` — joined from `event_types` rather than
  hardcoded to `{0, 3, 4}` so a future upstream toggle surfaces as a
  validation failure rather than a silent shift.
- `well_id` strips leading zeros from `WELL-NNNNN` and is NULL for
  simulated/drawn — enforced by the validator's well-kind contract.
- `source_url` is computed from the ADR-0001 URL pattern at catalog
  build time, before `observations/` exists. A new `PUBLIC_BASE_URL`
  constant centralises the base URL.
- Test fixtures use DuckDB to materialise five small instance parquets
  (one per representative event class, spanning real/simulated/drawn).
  The smoke test exercises the catalog contents, the well-kind/well-id
  contract, the four-bucket accounting, the source_url shape, and five
  validator abort paths.

Files added:
- `scripts/transform/petrobras_3w/instances_builder.py` — single-CTE
  SQL pipeline (raw → parsed → aggregated → joined-with-event_types).
- `tests/petrobras_3w/conftest.py` — `build_instance_parquets` helper
  that emits the deterministic per-test fixture tree.

Files modified:
- `scripts/transform/petrobras_3w/{orchestrator,constants}.py` — wire
  the new builder, add `PUBLIC_BASE_URL`.
- `scripts/export/petrobras_3w/parquet_writer.py` — add
  `write_instances`.
- `scripts/export/petrobras_3w/validator.py` — rule 1 (instance_id
  unique), rule 4 (event_class FK), well-kind contract, transient
  nullness vs event_types.has_transient, four-bucket accounting.
- `scripts/export/petrobras_3w/schema_doc_generator.py` — register
  `instances` in `TABLE_ORDER` + per-column descriptions; extend the
  README with corpus-balance and per-event-class-filter examples.
- `scripts/export/petrobras_3w/website_integrator.py` — bump tab count
  to 2 files, surface `instances.parquet` in the download grid, swap
  the README/page query example for the new catalog-balance one.
- `scripts/export/petrobras_3w/orchestrator.py` — call
  `write_instances` after `write_event_types`.
- `tests/petrobras_3w/test_smoke.py` — extend documentation/website
  tests + add seven `instances`-specific tests (contents, well_kind,
  row-count accounting, source_url, source_file, plus four validator
  abort paths).
- `CONTEXT.md` — clarify the `n_rows_normal` definition.

Notes for next iteration:
- The five remaining pre-publish validation rules from CONTEXT.md
  (rules 2, 3, 5, 6, 7, 8, 9) still need to land. Rule 7's expected
  real-Well count is **40** (resolved in triage on #21).
- Issue #21 (wells.parquet) builds on `instances` via
  `well_kind = 'real'` aggregation; #22 (observations) extends the
  validator with the per-Observations-file invariants (rule 5/6).
- Pre-existing test failure in `scripts/export/test_integration.py`
  (Volve-era, missing `parquet/wells.parquet`) remains as-is, same as
  #19 noted.

Closes #20

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Tracer-bullet extension of the #20 instances catalog that derives the
real-Well master. One row per distinct `well_id` from instances where
`well_kind = 'real'`, with per-Well aggregates (`n_instances`,
`first_ts`, `last_ts`, `n_observations`). Built via pure DuckDB SQL
against the existing intermediate DB — no additional upstream staging.

Key decisions:
- Rule 7's expected real-Well count is pinned to **40** in the
  validator (`EXPECTED_REAL_WELL_COUNT`), matching the locked decision
  from triage. Upstream's `dataset/README.md` claims "42 real wells
  covered" but only 40 IDs (`00001..00016`, `00019..00042`) actually
  appear in instance filenames at `v.1.70.0` / dataset version `2.0.0`
  — IDs 17 and 18 are absent. Pinning on the observed 40 preserves the
  fail-loud-on-upstream-drift property.
- Validator order is rule 3 (FK from instances to wells) → rule 7
  (rowcount) → rule 8 (wells limited to real well_ids). Rule 3 fires
  first because an orphan instance reference is the more fundamental
  invariant: a missing wells row breaks every downstream join, whereas
  rule 8 is the symmetric check that the master does not include
  synthetic IDs.
- Wells master deliberately excludes simulated and drawn instances by
  filtering at aggregation time. Upstream anonymises every physical
  attribute (no basin, field, depth, location), so the table is
  identity-plus-statistics only.
- Test fixtures gain 37 single-row event-0 padding instances so the
  total distinct real well count matches the pinned 40. The IDs leave
  the upstream gap at 17/18 intact, so happy-path tests exercise rule
  7 against a realistic catalog rather than a degenerate one. The
  existing 5 primary fixtures keep their event-class / well-kind mix
  for the catalog-shape assertions in #20's tests.

Files added:
- `scripts/transform/petrobras_3w/wells_builder.py` — pure-SQL
  aggregation from `instances WHERE well_kind = 'real'`.

Files modified:
- `scripts/transform/petrobras_3w/orchestrator.py` — wire the wells
  builder after the instances builder.
- `scripts/export/petrobras_3w/{parquet_writer,orchestrator}.py` — add
  `write_wells` and call it after `write_instances`.
- `scripts/export/petrobras_3w/validator.py` — rule 3
  (`WellsIdFkError`), rule 7 (`WellsRowCountError`,
  `EXPECTED_REAL_WELL_COUNT = 40`), rule 8 (`WellsKindError`).
- `scripts/export/petrobras_3w/schema_doc_generator.py` — register
  `wells` in `TABLE_ORDER`, add per-column descriptions, declare the
  `instances.well_id → wells.well_id` FK, add a wells/instances join
  query example to README.
- `scripts/export/petrobras_3w/website_integrator.py` — bump tab count
  to 3 files, surface `wells.parquet` in the download grid, refresh
  the description blurb.
- `tests/petrobras_3w/conftest.py` — extend `_INSTANCE_SPECS` with 37
  padding real-Well fixtures (IDs 4..16, 19..42).
- `tests/petrobras_3w/test_smoke.py` — relax the
  `test_pipeline_emits_instances_catalog` exact-row-count assertion;
  add 6 wells-specific tests (contents, exclusion of non-real,
  aggregates vs instances, rule 3 / 7 / 8 validator abort paths);
  extend the documentation + website-integration tests to cover the
  new table.

Notes for next iteration:
- The remaining pre-publish validation rules from CONTEXT.md (rules 2,
  5, 6, 9) cover the Observations files and land with #22.
- Pre-existing test failure in `scripts/export/test_integration.py`
  (Volve-era, missing `parquet/wells.parquet`) remains as-is, same as
  #19 and #20 noted.

Closes #21

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Tracer-bullet extension of #21 that emits the per-Instance Observations
time-series, completing the four-table Petrobras 3W shape. Each upstream
Instance parquet is republished to
`parquet/petrobras_3w/observations/event_class=N/<instance_id>.parquet`,
preserving the source columns verbatim (including hyphenated sensor
names) and adding three RLE-encoded constants: `instance_id`,
`well_id`, `well_kind`. `event_class` lives in the hive partition path
only — never in the file body. A `_files.json` manifest at the
partition root enumerates every published file.

Key decisions:
- `observations` is exposed as a DuckDB view over the staged sources
  rather than a materialised table. The full corpus (~2,228 files ×
  tens of thousands of rows × 30 columns) is too large for memory;
  the view's columns are derived on read and the validator pivots on
  them without materialising. `union_by_name=true` keeps the view
  tolerant of test fixtures that only carry a subset of columns.
- The writer iterates the catalog and runs one `COPY` per Instance
  against the matching staged source file. Going through the view's
  `WHERE instance_id = ?` would re-scan every staged file on every
  iteration (O(n²)); per-file COPY is O(n) total.
- Rule 9 (50 MB soft-warn) is emitted by `write_observations` after
  each file is written, not by `validator.validate`. The other
  Observations rules (2, 5, 6) run pre-write against the view because
  they catch data-semantic problems before any byte hits disk.
- `event_class` is documented as a logical column of `observations`
  with a new `hive_partition: true` flag in `schema.json`/`schema.md`,
  surfaced as a "Source" column in the table docs. The reflection
  passes `hive_partitioning=false` to read parquet metadata so DuckDB
  doesn't synthesize `event_class` from the parent dir name and
  double-count.
- The export orchestrator now takes a `staging_dir` parameter — the
  Observations writer needs direct access to the source files. The
  catalog tables (event_types/instances/wells) remain
  staging-independent (they live as materialised tables).
- Validator rule 5 is split into four pre-existing-style functions
  (FK, row count, timestamp monotonic, class domain) so each can fire
  with a precise exception class and the test suite can target them
  individually. The "single (instance_id, well_id, well_kind) triple
  per file" sub-clause is satisfied by construction (constants
  derived per filename) and is therefore not re-checked at runtime.
- Test fixtures gain a hyphenated `P-PDG` sensor column + `state` so
  the column-name fidelity policy is exercised end-to-end on the
  written parquets, not just in the docs.
- The validator-rejection tests for rules 5 (timestamp / class
  domain) and 6 (class) mutate the staged Instance file *after*
  transform builds the catalog, preserving the catalog's bucket
  accounting (rule from #20) so the targeted Observation-side rule
  is the failing one rather than the upstream catalog check.

Files added:
- `scripts/transform/petrobras_3w/observations_builder.py` — view
  with derived `event_class` / `instance_id` / `well_id` / `well_kind`.

Files modified:
- `scripts/transform/petrobras_3w/orchestrator.py` — wire the new
  builder after `wells_builder`.
- `scripts/export/petrobras_3w/parquet_writer.py` — add
  `write_observations` (per-Instance COPY + `_files.json` manifest +
  rule 9 soft-warn).
- `scripts/export/petrobras_3w/orchestrator.py` — pass `staging_dir`
  through; call the new writer after `write_wells`.
- `scripts/export/petrobras_3w/validator.py` — new exceptions
  (`ObservationsInstanceFkError`, `ObservationsRowCountError`,
  `ObservationsTimestampError`, `ObservationsClassDomainError`,
  `ObservationsNonTransientClassError`) and the five rule checks
  (rules 2, 5×3, 6).
- `scripts/export/petrobras_3w/schema_doc_generator.py` — register
  `observations` in `TABLE_ORDER`, add description / column docs /
  PK / FKs, prepend `event_class` as a hive-partition column, surface
  the four new query examples in `README.md` (event-class scan,
  single-Instance fetch, manifest enumeration, hive glob).
- `scripts/export/petrobras_3w/website_integrator.py` — bump tab
  count, surface `observations/_files.json`, swap the placeholder
  follow-up paragraph for the hive-glob query example.
- `tests/petrobras_3w/conftest.py` — add `state` and `P-PDG`
  columns to the per-Instance fixtures.
- `tests/petrobras_3w/test_smoke.py` — add 9 observations-specific
  tests (layout, hyphen fidelity, NULL well_id for non-real,
  manifest, README query, schema.sql round-trip, rules 2 / 5 / 5 /
  6 abort paths) + thread `staging_dir` through every export call.

Notes for next iteration:
- Issue #23 (parity test suite vs upstream) extends the validator
  with the nine parity queries described in PRD #18; the
  Observations view defined here is the SQL surface those parity
  queries will run against.
- Issue #24 (final docs + website polish) refines the README/schema
  examples and the index page beyond this slice's tracer-bullet
  level.
- Pre-existing test failure in `scripts/export/test_integration.py`
  (Volve-era, missing `parquet/wells.parquet`) remains as-is, same
  as #19/#20/#21 noted.

Closes #22

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds the nine-check parity suite from PRD #18, wired in as the
post-write correctness gate. Where the existing pre-write `validator`
asserts structural invariants of the intermediate DB, the new
`parity.check(staging_dir, output_dir)` proves the published bytes
round-trip the upstream bytes 1:1 — class, state, timestamp, and
every sensor column are preserved verbatim, so any divergence is a
writer bug rather than legitimate transformation.

Checks (PRD #18 order, all hard-fail on divergence):
1.  Per-event-class row count — upstream / catalog / pub.
2.  Per-instance row count — upstream / catalog / pub.
3.  Global `class` distribution — upstream vs pub.
4.  Global `state` distribution — upstream vs pub.
5.  Per-sensor SUM/AVG/MIN/MAX/COUNT/NULL — upstream vs pub.
6.  Per-sensor aggregates grouped by `event_class` — same metrics.
7.  Per-instance (start_ts, end_ts) — upstream / catalog / pub.
8.  Distinct real-Well count — upstream / wells.parquet / pinned 40.
9.  Per-event-class instance count — upstream vs catalog.

Key decisions:
- Parity runs *after* the parquets are written but *before* the
  static-site tab is patched, so a divergence keeps the existing
  published site intact. Implementation-wise, this means closing the
  validator's read-only DuckDB connection, calling `parity.check`
  with its own in-memory connection over the written parquets, then
  re-opening a read-only connection for the schema-doc generator.
- Sensor columns are discovered at runtime from the published
  `observations` schema (everything that is not an
  identifier/label/time/hive column). Keeps the check agnostic to
  upstream's exact 27-column set — a future column rename surfaces
  as a parity match on the renamed column instead of a hard-coded
  reference to a missing column name.
- NULL `class` and `state` values are legitimate on the warmup
  prefix of real-Well anomaly Instances, so the global-distribution
  checks (3 / 4) use `IS NOT DISTINCT FROM` for their join predicate.
  Plain `USING (class)` would split the NULL bucket into two phantom
  mismatches.
- Sensor columns are double-quoted via a `_quote` helper so hyphenated
  identifiers (`P-PDG`, `ESTADO-SDV-GL`) round-trip safely.
- The bit-for-bit comparison policy (no epsilon) is from PRD #18:
  every sensor float is copied unchanged, so identical SUM/AVG on
  identical multisets is the contract. If full-corpus runs ever
  surface non-associativity drift, we can switch to per-instance
  row-level EXCEPT — for now the simpler aggregate compare suffices.
- The real-Well count (check 8) pins on `EXPECTED_REAL_WELL_COUNT`
  imported from `validator` so the parity suite and the pre-write
  validator stay in lockstep on the pinned 40.
- The orchestrator-level integration test patches `parity.check` to
  raise, then asserts the website integrator never runs (the stub
  README + index.html stay unchanged). This proves the wiring
  without needing to artificially break the published bytes.
- The mutation-detection tests use a `_mutate_published_parquet`
  helper that re-reads a written file with `hive_partitioning=false`,
  applies a SQL mutation, and overwrites in place. Parallel pattern
  to `_rewrite_staged_instance` from #22, but operating against the
  published tree rather than the staged sources.

Files added:
- `scripts/export/petrobras_3w/parity.py` — module with `check()`
  entry point and nine `Parity*Error` subclasses, one per PRD check.

Files modified:
- `scripts/export/petrobras_3w/orchestrator.py` — split the existing
  single-connection block in two so `parity.check` can run between
  the writes and the schema-doc / website-integrator steps; updated
  the module docstring to describe the two correctness gates.
- `tests/petrobras_3w/test_smoke.py` — five parity tests: happy-path,
  sensor-value mutation, partition-internal per-instance drift, row
  drop (per-event-class), class-label flip, and orchestrator-aborts-
  on-parity-failure.

Notes for next iteration:
- Issue #24 (final docs + website polish) is now unblocked.
- The full-corpus parity run is part of the production publish path
  via the orchestrator hook; the smoke suite runs the 42-instance
  subset version under 20 s on the fixture, well within the 30 s cap.
- Pre-existing test failure in `scripts/export/test_integration.py`
  (Volve-era, missing `parquet/wells.parquet`) is unchanged.

Closes #23

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Refreshes the published documentation surface for the Petrobras 3W
dataset to the final four-table state. The committed `parquet/petrobras_3w/`
docs and the static-site entry were last touched at the #19 skeleton
slice and only described `event_types`; this PR brings them in lockstep
with the writers landed by #20 (`instances`), #21 (`wells`), #22
(`observations`), and #23 (parity).

Generator improvements:
- `schema.md` sensor-glossary intro drops the stale "once issue #22
  lands" forward-reference — those bytes have shipped, so the prose now
  describes the observations columns as present.
- The observations table's per-column docs now inline upstream's sensor
  descriptions from `dataset.ini` (e.g. `P-PDG` → "Downhole pressure at
  the PDG (permanent downhole gauge) [Pa]") instead of leaving the
  Description cell blank with the glossary as the only reference.
- README adds the per-Well cross-validation split query (acceptance
  criterion c from issue #24) — leave-one-Well-out via `well_id %
  n_folds`, with a follow-up note on handling simulated/drawn
  Instances. The four canonical examples (a load-by-event-class, b
  fetch-by-URL, c per-Well CV split, d corpus-balance from catalog
  only) are now all surfaced.

Test fixture upgrade:
- `tests/petrobras_3w/conftest.py` reads upstream's
  `PARQUET_FILE_PROPERTIES` and writes all 27 sensor columns into the
  per-Instance fixtures (placeholder values for the 26 non-P-PDG
  columns; `P-PDG` keeps its row-varying values so parity sensor
  aggregates still differentiate). This makes the fixture-derived
  schema docs structurally match production, which is what lets us
  commit them. The lone hyphenated `P-PDG` column was enough to
  exercise column-name fidelity end-to-end at #22's level, but it
  understated the column inventory in the reflected schema.
- `test_parity_detects_per_instance_only_drift`'s INSERT now uses
  `INSERT BY NAME SELECT *` so it doesn't need to enumerate every
  sensor column by hand.

Regenerated committed artifacts:
- `parquet/petrobras_3w/schema.md`, `schema.json`, `schema.sql`,
  `README.md` regenerated against the upgraded fixture — `schema.sql`
  now declares all 27 sensor columns with hyphen-quoted identifiers,
  `schema.json` carries the full column / FK / PK metadata for all
  four tables, `schema.md` and `README.md` carry every per-table
  description and every canonical query example.
- Root `README.md` and `parquet/index.html` regenerated via the
  website integrator — the tab count was stuck at "1 file" since
  #19 and the README blurb still claimed the catalog and Observations
  "ship in follow-up issues". Both now reflect the realised
  four-table state.

Key decisions:
- Committed schema docs are produced from the FIXTURE pipeline, not
  from a full upstream publish (3+ GB unstageable). The docs describe
  shape, not content — bumping fixture column inventory to match
  upstream's 27-column inventory makes the fixture-derived docs
  production-correct. The actual data parquets (`wells`, `instances`,
  `observations/`) still come from the upstream publish at deploy
  time; only `event_types.parquet` is byte-stable across the two
  paths because it is dataset.ini-driven.
- Per-Well CV split uses `well_id %% n_folds` as the fold key. Stable
  across refreshes and zero-state (no separate fold-assignment file
  needed); consumers picking another assignment can replace the
  `%% 5` clause without touching the rest of the query.
- `dataset.ini` parsing is reused in the test conftest via
  `parse_dataset_ini` rather than hard-coding the 27 column names —
  any future upstream column rename surfaces at fixture build time as
  a parse-time mismatch, the same fail-loud property the production
  pipeline has.

Files modified:
- `scripts/export/petrobras_3w/schema_doc_generator.py` — sensor-
  description inline lookup + glossary-intro prose update + new
  CV-split README section.
- `tests/petrobras_3w/conftest.py` — 27-sensor fixture columns from
  upstream `dataset.ini`.
- `tests/petrobras_3w/test_smoke.py` — assert no stale "#22 lands"
  prose; assert sensor descriptions are inlined in observations table;
  assert the four canonical queries (a/b/c/d) are present; switch the
  per-instance-drift parity test's INSERT to `BY NAME SELECT *`.
- `parquet/petrobras_3w/{schema.md,schema.json,schema.sql,README.md,LICENSE-3W-DATA.md}`
  — regenerated.
- `README.md`, `parquet/index.html` — refreshed via the website
  integrator.

Notes for next iteration:
- Pre-existing test failure in `scripts/export/test_integration.py`
  (Volve-era, missing `parquet/wells.parquet`) is unchanged, same as
  every #19..#23 noted.
- The deploy step that copies the full ~2,228-file Observations tree
  to `dev-petrodb.ocortez.com` is out of scope for this PR — the
  pipeline running in production is what materialises those bytes.
  The committed schema docs accurately describe what the deployed
  tree contains.

Closes #24

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Replaces the GitHub Pages workflow (which would fail on the 1 GB cap with
the 3W tree) with a single Cloudflare Pages workflow that handles both
environments via `wrangler --branch ${{ github.ref_name }}`: push to
`main` lands in production, push to `stage` in the preview deployment.

The 3W generated tree is gitignored — the workflow rebuilds it from the
pinned upstream tag (ADR-0002) and uploads the whole `parquet/` directory
so committed datasets ride along. Pipeline re-runs are cached on the 3W
sources + per-branch `BASE_URL`, so non-3W changes deploy without
rebuilding. `PUBLIC_BASE_URL` is now host-only via env, with the
`/petrobras_3w` segment derived from the dataset's directory.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Sync 3W dataset + Cloudflare Pages migration into dev
@oskrgab
oskrgab merged commit 67e0c45 into main Jun 4, 2026
1 check passed
@oskrgab
oskrgab deleted the stage branch June 14, 2026 05:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant