From 5e2a4dfa806fcf86f1f51ef3023f3c69605bbf21 Mon Sep 17 00:00:00 2001 From: oskrgab Date: Sun, 17 May 2026 18:54:12 +0000 Subject: [PATCH 01/11] Petrobras 3W domain context + observations-layout ADR + upstream-tag pinning ADR Captures the design decisions reached while grilling issue #17 / PRD #18: - CONTEXT.md adds the 3W language (Well, Instance, Well kind, Event class, Transient label, Observation, Source file), cross-dataset ambiguity entries for Well and Event, and four 3W dataset sections (tables, operating principles, output layout, pre-publish validation). - ADR-0001 records the observations-layout decision (one file per Instance, hive on event_class) with the four rejected alternatives. - ADR-0002 records the upstream-tag-pinning policy (v2.0.0, event-driven refresh) and why tracking main would silently mutate already-published bytes. Co-Authored-By: Claude Opus 4.7 --- CONTEXT.md | 91 ++++++++++++++++++- .../0001-petrobras-3w-observations-layout.md | 24 +++++ ...2-petrobras-3w-pin-upstream-release-tag.md | 22 +++++ 3 files changed, 136 insertions(+), 1 deletion(-) create mode 100644 docs/adr/0001-petrobras-3w-observations-layout.md create mode 100644 docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md diff --git a/CONTEXT.md b/CONTEXT.md index 0af8c66..4049db6 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,14 +16,43 @@ The human-readable code for a well (e.g. `YPF.BLO.x-8`). Treated as a label, not **Formprod** (producing formation): The geological formation a given `idpozo` produces from. Static per `idpozo` by construction (encoded into the ID itself), so it lives in the master table, never in the time-series. +### Petrobras 3W dataset + +**Well** (3W sense, `well_id`): +A physical offshore producing well operated by Petrobras, identified by an integer anonymized by upstream (e.g. `1`, `19`). 1-Hz sensor histories from the well are sliced into many **Instances**. The dataset covers ~40 distinct real wells. _Not the same as Argentina's Well_: there is no `formprod` split here, and no human-readable label. + +**Instance** (`instance_id`): +A single contiguous 1-Hz time-window of sensor data drawn from one source, framed around at most one labeled anomaly event. Identified by the upstream source filename without extension (e.g. `WELL-00019_20120601165020`). An instance is the natural unit of an experiment: one file = one window to train, evaluate, or visualize. Length varies from ~6 hours (~21k rows) to ~3 days (~243k rows). + +**Well kind** (`well_kind`): +The provenance of an Instance. One of `{real, simulated, drawn}` — `real` for instances sourced from a physical Well (~40 wells across the corpus), `simulated` for synthetic instances generated by upstream tooling, `drawn` for hand-drawn series. Simulated and drawn instances have **no** `well_id`. + +**Event class** (`event_class`): +An integer code in `0..9` identifying the operational regime an Instance is framed around. `0 = NORMAL`, `1..9 = anomaly categories` (Abrupt Increase of BSW, Spurious Closure of DHSV, Severe Slugging, Flow Instability, Rapid Productivity Loss, Quick Restriction in PCK, Scaling in PCK, Hydrate in Production Line, Hydrate in Service Line). The canonical names come from `dataset.ini` upstream. _Not the same as Argentina's "event"_, which means an operational state transition row. + +**Transient label**: +A per-observation `class` value equal to `event_class + 100`, marking the developing phase of an anomaly before it reaches steady state. Codes seen in the data: `101, 102, 105, 106, 107, 108, 109`. Defined only for the seven anomalies that upstream marks `TRANSIENT=True`. Events `3` (Severe Slugging) and `4` (Flow Instability) are `TRANSIENT=False` — no transient codes exist for them. + +**Observation**: +One row of an `observations` table — a single 1-Hz timestamped sample with 27 sensor floats, a `class` label (per-observation regime: NULL, 0, anomaly class, or transient code), and a `state` label (well operational status). + +**Source file**: +The upstream parquet file backing a single Instance. Filenames encode provenance: `WELL-_.parquet` for real, `SIMULATED_.parquet`, `DRAWN_.parquet`. In our published catalog the source filename is preserved as a column so consumers can cross-reference with upstream. + ## Relationships - A **Well** (`idpozo`) is the unit of identity for both the master table and the production time-series. - **Formprod** is a static attribute of a **Well** — recorded once in the master, never repeated in monthly rows. +- A **Well** (3W sense, `well_id`) is the parent of zero-or-more **Instances** (only when `well_kind = real`). +- An **Instance** is the parent of many **Observations** (1-Hz rows). +- An **Event class** is referenced by an **Instance** (the regime the window is framed around) and by an **Observation** (the per-row label, possibly via its transient offset). +- Simulated and drawn **Instances** have no parent **Well** — their `well_id` is NULL by design. ## Flagged ambiguities -- "well" was used to refer to both the physical wellbore and the producing-formation-specific record — resolved: in this dataset a **Well** = `idpozo` = wellbore × formprod. Use "physical wellbore" if the bore-only concept is needed. +- "well" was used to refer to both the physical wellbore and the producing-formation-specific record — resolved: in the Argentina dataset a **Well** = `idpozo` = wellbore × formprod. Use "physical wellbore" if the bore-only concept is needed. +- "Well" is overloaded across datasets — resolved by scope: Argentina **Well** = `idpozo`; Petrobras 3W **Well** = anonymized offshore wellbore `well_id`. Always qualify with the dataset name when the context is mixed. +- "Event" is overloaded — resolved: in the Argentina dataset an **Event** is a row in `well_events.parquet` marking an operational-state transition; in the Petrobras 3W dataset an **Event class** is an anomaly category referenced by an Instance and by every Observation. Different concepts, different tables. ## Argentina dataset — column buckets @@ -100,3 +129,63 @@ The export step asserts the following before writing Parquets; failure aborts pu 4. Geometry in `wells.parquet` is parseable WKB for every row that carries a `geom`. 5. Year-partition count is 22 and the sum of partition row counts equals the staged-source row count. 6. Soft-warn if any year-partition Parquet exceeds 50 MB (Cloudflare cache headroom). + +## Petrobras 3W dataset — tables + +A four-table relational shape, derived from upstream's per-instance Parquet files plus `dataset.ini`. Layouts and column names below are settled in the grill; pipeline implementation details (e.g. exact column-name policy for hyphenated source columns) are still open. + +**Static master** (`wells.parquet`, one row per real Well, ~40 rows): +`well_id`, `n_instances` (count of `instances` rows), `first_ts`, `last_ts`, `n_observations` (sum across instances). Upstream anonymizes physical well attributes (no basin, field, depth, or location is published), so this table is mostly an identity-plus-statistics master. Only `well_kind = real` instances have a parent row here. + +**Lookup** (`event_types.parquet`, exactly 10 rows): +`event_class` (PK, 0..9), `name` (canonical PascalCase from `dataset.ini`, e.g. `HYDRATE_IN_PRODUCTION_LINE`), `description` (human-readable, e.g. `Hydrate in Production Line`), `has_transient` (boolean, false for `{0, 3, 4}`), `transient_code` (`event_class + 100` when `has_transient`, NULL otherwise), `has_normal_prefix` (boolean, true for events that include a class=0 precursor in the data — correlates with `has_transient`). + +**Instance catalog** (`instances.parquet`, one row per Instance, ~2,228 rows): +`instance_id` (PK, source filename without extension), `well_kind` (enum: `real | simulated | drawn`), `well_id` (FK to `wells.parquet`, NULL when `well_kind != real`), `event_class` (FK to `event_types.parquet`), `start_ts`, `end_ts`, `duration_s` (derived), `n_rows`, `n_rows_warmup_null` (rows where `class IS NULL`), `n_rows_normal` (rows where `class = 0`), `n_rows_transient` (rows where `class = event_class + 100`; NULL when `has_transient = false`), `n_rows_steady` (rows where `class = event_class`), `source_file` (upstream filename for cross-reference), `source_url` (URL to the published Observations parquet). The four `n_rows_*` columns let corpus-wide balance and labeled-mass queries run purely against this catalog without scanning Observations. + +**Observations time-series** (`observations/event_class=N/.parquet`, ~2,228 files): +Hive-partitioned by `event_class` only. Each file is a single Instance's 1-Hz rows. Columns: the 27 source sensor columns + `class` (per-observation label, may include the +100 transient codes) + `state` (well operational status) + `timestamp` + three added constant columns: `instance_id`, `well_id`, `well_kind`. `event_class` is provided by the hive partition, not stored in the file body. + +## Petrobras 3W dataset — operating principles + +- **Source column names preserved verbatim**, including hyphens (`P-PDG`, `ABER-CKGL`, `ESTADO-SDV-GL`). Consumers must double-quote these identifiers in SQL. Same fidelity principle applied to Argentina, applied here to hyphenated forms instead of Spanish forms. +- **Source fidelity over smoothing**: the `class` column's NULL warmup prefix (~1 hour at the start of every real-Well instance) is preserved; the `+100` transient offset is preserved as a raw code (lookup table explains the convention); per-instance file boundaries match upstream's per-instance file boundaries. No row-level reinterpretation. +- **Per-row provenance**: `instance_id`, `well_id`, `well_kind` are duplicated into each Observations file as constant columns (RLE-encoded, negligible storage). The dataset stays self-describing without filename-parsing tribal knowledge. +- **TRANSIENT semantics**: instances of an event with `has_transient = true` carry a `NORMAL → TRANSIENT → STEADY` arc in their `class` column. Instances of an event with `has_transient = false` (events 3 and 4) carry **only** the steady class — no `NORMAL` precursor either. Consumers training early-detection models should filter on `has_transient`. +- **3W toolkit artefacts are excluded**: per-event detector hyperparameters (`WINDOW`, `STEP`) and fold-config knobs (`EXTRA_INSTANCES_TRAINING`) live in upstream's `dataset.ini` but belong to the *3W toolkit*, not to the dataset itself. Petrodb publishes the dataset; toolkit and folds remain the consumer's responsibility. +- **Pure DuckDB SQL transform**: pipeline is read-with-`filename=true` → parse filename into catalog columns → `COPY ... PARTITION_BY (event_class)`. Polars permitted only where SQL becomes unreadable; pandas is not used. Same rule as the Argentina pipeline. +- **Upstream pinned to a release tag** (currently `v2.0.0`). Refreshes are event-driven on new upstream tags, not calendar-driven. Past upstream releases have retroactively changed sensor values inside the same `instance_id`, so tracking `main` would silently mutate already-published parquet bytes. See [ADR-0002](docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md). + +## Petrobras 3W dataset — output layout + +``` +parquet/petrobras_3w/ +├── wells.parquet +├── event_types.parquet +├── instances.parquet +├── observations/ +│ ├── event_class=0/.parquet # 594 files +│ ├── event_class=1/.parquet # 128 files +│ ├── … +│ └── event_class=9/.parquet # 207 files +├── observations/_files.json # manifest of all Observations URLs for httpfs consumers +├── schema.md / schema.json / schema.sql +├── README.md +└── LICENSE-3W-DATA.md # CC BY 4.0 mirror with upstream attribution +``` + +Rationale: 2,228 instances × ~0.4–1 MB each keep every file well under Cloudflare's edge-cache size target. Hive partitioning by `event_class` prunes the dominant ML query pattern ("train on class N"). One-file-per-instance preserves the upstream conceptual unit and makes single-instance fetch a direct GET (the ML training-loop pattern). Catalog tables answer corpus-wide questions without touching Observations. + +## Petrobras 3W dataset — pre-publish validation + +The export step asserts the following before writing Parquets; failure aborts publish. + +1. `instances.instance_id` is unique. +2. Every `instance_id` referenced by an Observations file exists in `instances.parquet`. +3. Every non-NULL `instances.well_id` exists in `wells.parquet` (FK integrity); NULL only when `well_kind != real`. +4. Every `instances.event_class` exists in `event_types.parquet` (FK integrity). +5. Within each Observations file: only one distinct `(instance_id, well_id, well_kind)` triple; `class ∈ {NULL, 0, event_class, transient_code(event_class)}` for the file's `event_class`; row count equals `instances.n_rows`; `timestamp` is strictly monotonic at exactly 1-second cadence. +6. For Instances with `has_transient = false` event class (events 3, 4): no `class` value ≥ 100 appears; no `class = 0` appears. +7. Real-Well coverage equals upstream's stated count (currently 42 — fail-loud if our derived `wells.parquet` rowcount disagrees, to catch upstream version drift). +8. `wells.parquet` rows are limited to the union of `instances.well_id WHERE well_kind = 'real'`. +9. Soft-warn if any Observations Parquet exceeds 50 MB (Cloudflare cache headroom). diff --git a/docs/adr/0001-petrobras-3w-observations-layout.md b/docs/adr/0001-petrobras-3w-observations-layout.md new file mode 100644 index 0000000..a915042 --- /dev/null +++ b/docs/adr/0001-petrobras-3w-observations-layout.md @@ -0,0 +1,24 @@ +# Petrobras 3W observations layout: one file per Instance, hive on `event_class` + +**Status:** accepted + +## Context + +The Petrobras 3W dataset is a corpus of ~2,228 labeled 1-Hz sensor-data windows ("Instances"), totalling ~1.74 GB compressed. Instance length varies from ~21k rows (~6 h) to ~243k rows (~3 days). The ML workload that dominates is event-detector training, which has two distinct access patterns: (a) "load all instances of event class N" and (b) "load one specific Instance for visualization or single-window training." Petrodb is published as static Parquet files behind Cloudflare with a soft per-file size target of 50 MB (CONTEXT.md `Argentina dataset — pre-publish validation` rule 6). + +## Decision + +Publish `observations/` as `observations/event_class=N/.parquet` — hive-partitioned by `event_class` only, with one file per Instance preserving upstream's per-instance file boundaries. Inside each file, in addition to the 30 upstream columns, store `instance_id`, `well_id`, `well_kind` as constant columns (RLE-encoded, negligible cost) so the dataset stays self-describing under any future restructuring. + +## Considered alternatives + +- **Single coalesced `observations.parquet`** with row-group sort on `(event_class, instance_id, timestamp)`. Best for corpus-wide aggregate scans; rejected because the 1.74 GB single file busts the per-file cache target and a cold CDN miss serves the entire file even when DuckDB only needs a row-group range. +- **Hive on `event_class` only, coalesced within partition** (10 files, ~50–500 MB each). Rejected: partition 0 (NORMAL, 594 instances) alone is hundreds of MB; busts the cache target; also discards the per-Instance file boundary, which is the natural unit of the ML workload. +- **Hive on `(event_class × well_id)`** (~100–150 partitions of ~10–20 MB). Fits the cache target and is efficient for combined class+well filters, but the dominant pattern is single-Instance fetch — pulling a 15 MB partition to read a 0.5 MB Instance is wasted bandwidth and worse cold-start latency. +- **Thin mirror of upstream** (`dataset/N/*.parquet`). Rejected: leaves `well_id`, `event_class`, `well_kind`, `instance_id` encoded only in filenames and directory names, requiring tribal knowledge to query. Contradicts petrodb's mission of clean, non-redundant relational schemas (CONTEXT.md, line 3). + +## Consequences + +- Corpus-wide aggregate scans require 2,228 parallel HTTP requests rather than a single large GET. DuckDB httpfs handles this concurrently, and the catalog tables (`instances.parquet`, `wells.parquet`, `event_types.parquet`) answer most aggregate questions without touching `observations/` at all. +- The published URL space is committed: `observations/event_class=N/.parquet` becomes part of the public API. Reorganizing later breaks consumers. +- Adding `instance_id` / `well_id` / `well_kind` as constant columns means future coalescing (if ever needed) is a transparent operation — no schema migration for downstream consumers. diff --git a/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md b/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md new file mode 100644 index 0000000..465bb60 --- /dev/null +++ b/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md @@ -0,0 +1,22 @@ +# Pin Petrobras 3W upstream to a release tag, not `main` + +**Status:** accepted + +## Context + +Petrobras 3W uses semantic versioning (`VERSIONING.md` upstream). Past minor releases have reshaped instance contents in-place: 1.1.0 added/removed instances, adjusted expert labels, and corrected historian-tag misconfigurations that retroactively changed sensor values. The same `instance_id` can therefore have different bytes between upstream commits, with no rename or URL change to signal it. Petrodb publishes parquet at stable URLs, so silently-changing bytes would propagate as silently-changing training data to downstream consumers. + +## Decision + +The Petrobras 3W pipeline reads from a pinned upstream release tag (currently `v2.0.0`). Refreshes are event-driven (when upstream cuts a new tag and we've reviewed the release notes), not calendar-driven. The current pinned tag is recorded in `parquet/petrobras_3w/README.md` and emitted in the pipeline's validation logs. + +## Considered alternatives + +- **Track upstream `main`.** Rejected — re-pulling at any time risks silently mutating already-published `instance_id`s (sensor-value corrections, label adjustments). Consumers relying on bytes-stable URLs would observe non-reproducible behaviour. +- **Pin to a commit SHA.** Strongest reproducibility, but upstream cuts release tags deliberately at coherent dataset states; an arbitrary SHA boundary is harder to reason about and harder to compare against the release notes. The marginal reproducibility gain over a tag is small enough not to justify the extra friction. + +## Consequences + +- Pre-publish validation rule 7 (real-Well coverage equals upstream's stated count) implicitly enforces the pin — an accidental tag change shows up as a row-count mismatch, not a silent corruption. +- Petrodb's release cadence for this dataset is bounded above by upstream's release cadence. If upstream goes dormant we publish stale data; that's acceptable for the reproducibility guarantee it buys. +- The pinned tag is part of the dataset's published metadata; bumping it is a deliberate, reviewable action with release-note context, not an automated refresh. From 6dbe510ecc983c07f1073546b83208022c6e88a5 Mon Sep 17 00:00:00 2001 From: oskrgab Date: Sun, 17 May 2026 19:34:12 +0000 Subject: [PATCH 02/11] Covered additional context about the dataset tag --- CONTEXT.md | 4 ++-- docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md | 4 +++- 2 files changed, 5 insertions(+), 3 deletions(-) diff --git a/CONTEXT.md b/CONTEXT.md index 4049db6..41bc722 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -154,7 +154,7 @@ Hive-partitioned by `event_class` only. Each file is a single Instance's 1-Hz ro - **TRANSIENT semantics**: instances of an event with `has_transient = true` carry a `NORMAL → TRANSIENT → STEADY` arc in their `class` column. Instances of an event with `has_transient = false` (events 3 and 4) carry **only** the steady class — no `NORMAL` precursor either. Consumers training early-detection models should filter on `has_transient`. - **3W toolkit artefacts are excluded**: per-event detector hyperparameters (`WINDOW`, `STEP`) and fold-config knobs (`EXTRA_INSTANCES_TRAINING`) live in upstream's `dataset.ini` but belong to the *3W toolkit*, not to the dataset itself. Petrodb publishes the dataset; toolkit and folds remain the consumer's responsibility. - **Pure DuckDB SQL transform**: pipeline is read-with-`filename=true` → parse filename into catalog columns → `COPY ... PARTITION_BY (event_class)`. Polars permitted only where SQL becomes unreadable; pandas is not used. Same rule as the Argentina pipeline. -- **Upstream pinned to a release tag** (currently `v2.0.0`). Refreshes are event-driven on new upstream tags, not calendar-driven. Past upstream releases have retroactively changed sensor values inside the same `instance_id`, so tracking `main` would silently mutate already-published parquet bytes. See [ADR-0002](docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md). +- **Upstream pinned to a git tag that ships a specific dataset version.** Currently pinned to git tag `v.1.70.0`, which ships upstream dataset version `2.0.0`. Refreshes are event-driven on new upstream tags, not calendar-driven. Past upstream releases have retroactively changed sensor values inside the same `instance_id`, so tracking `main` would silently mutate already-published parquet bytes. Upstream uses two version namespaces — git tags (`v.1.NN.0`) and dataset semver in `dataset/README.md` (`1.0.0`, `1.1.0`, `1.1.1`, `2.0.0`, …); we pin the git tag because that is what `git clone --branch` accepts, and we record the dataset version because that is what identifies the data shape. See [ADR-0002](docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md). ## Petrobras 3W dataset — output layout @@ -186,6 +186,6 @@ The export step asserts the following before writing Parquets; failure aborts pu 4. Every `instances.event_class` exists in `event_types.parquet` (FK integrity). 5. Within each Observations file: only one distinct `(instance_id, well_id, well_kind)` triple; `class ∈ {NULL, 0, event_class, transient_code(event_class)}` for the file's `event_class`; row count equals `instances.n_rows`; `timestamp` is strictly monotonic at exactly 1-second cadence. 6. For Instances with `has_transient = false` event class (events 3, 4): no `class` value ≥ 100 appears; no `class = 0` appears. -7. Real-Well coverage equals upstream's stated count (currently 42 — fail-loud if our derived `wells.parquet` rowcount disagrees, to catch upstream version drift). +7. Real-Well coverage equals the count derived from upstream filename prefixes at the pinned git tag (currently 40 distinct `WELL-NNNNN` prefixes at `v.1.70.0` / dataset version 2.0.0 — fail-loud if our derived `wells.parquet` rowcount disagrees, to catch upstream version drift). Upstream's `dataset/README.md` states "42 real wells covered", but only 40 IDs (`00001..00016`, `00019..00042`) actually appear in instance filenames; IDs `00017` and `00018` are absent. The validator pins on the observed 40. 8. `wells.parquet` rows are limited to the union of `instances.well_id WHERE well_kind = 'real'`. 9. Soft-warn if any Observations Parquet exceeds 50 MB (Cloudflare cache headroom). diff --git a/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md b/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md index 465bb60..c44fde5 100644 --- a/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md +++ b/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md @@ -8,7 +8,9 @@ Petrobras 3W uses semantic versioning (`VERSIONING.md` upstream). Past minor rel ## Decision -The Petrobras 3W pipeline reads from a pinned upstream release tag (currently `v2.0.0`). Refreshes are event-driven (when upstream cuts a new tag and we've reviewed the release notes), not calendar-driven. The current pinned tag is recorded in `parquet/petrobras_3w/README.md` and emitted in the pipeline's validation logs. +The Petrobras 3W pipeline reads from a pinned upstream git tag that ships a specific upstream *dataset version*. The currently pinned git tag is `v.1.70.0`, which ships dataset version `2.0.0`. Refreshes are event-driven (when upstream cuts a new git tag — typically corresponding to a new dataset version — and we've reviewed the release notes), not calendar-driven. The current pinned git tag and dataset version are recorded in `parquet/petrobras_3w/README.md` and emitted in the pipeline's validation logs. + +Note on upstream versioning: upstream uses two distinct version namespaces. Git tags are formatted `v.1.NN.0` (with the dot after `v`); these are the only mechanism a clone can pin to. The *dataset* itself carries a separate semver in `dataset/README.md` (`1.0.0`, `1.1.0`, `1.1.1`, `2.0.0`, …) — this is the version that identifies the data shape and content. The dataset version is what consumers care about; the git tag is the pinning mechanism that delivers it byte-stably. A single dataset version is typically shipped by many consecutive git tags (toolkit/docs fixes don't bump the dataset); for stability we pin to the latest available git tag for the chosen dataset version. ## Considered alternatives From 88255cc1ae9694692471b6aada6ba3bcf68199e2 Mon Sep 17 00:00:00 2001 From: oskrgab Date: Sun, 17 May 2026 19:46:27 +0000 Subject: [PATCH 03/11] Petrobras 3W skeleton pipeline + event_types.parquet (#19) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Tracer-bullet slice that establishes the transform/export module layout for the Petrobras 3W dataset and emits its smallest deliverable table. Subsequent slices (#20, #21, #22) add the Instance catalog, real-Well master, and Observations hive partition without re-touching this scaffolding. Key decisions: - Source of truth for event_types is upstream `dataset.ini`. The pipeline parses it via configparser (with `optionxform=str` to preserve hyphenated sensor column names like `P-PDG` and `QGL`) rather than hard-coding canonical values — a future upstream rename surfaces as a parse-time change rather than silent drift. - The shallow-clone is idempotent: presence of `/dataset/dataset.ini` short-circuits the `git clone`, so the smoke test can stand in a fixture upstream tree without network access. - Pinned identity (git tag `v.1.70.0`, dataset version `2.0.0`) is emitted via the export's `petrobras_3w.export` logger before any parquet write, and is also recorded in `schema.json`, `schema.sql`, `README.md`, and the static-site tab. A parsed `dataset_version` that disagrees with the pin aborts publish (ADR-0002 event-driven refresh). - `has_normal_prefix` is materialised equal to `has_transient` per CONTEXT.md (events 0, 3, 4 carry only the steady class — no NORMAL precursor). - DDL identifier quoting in `schema.sql` is general (`_quote_identifier`) so the hyphenated sensor columns added in #22 round-trip without re-touching the writer. - Website integration uses the same sentinel-comment idempotency pattern as the Argentina integrator. Files added: - `scripts/transform/petrobras_3w/` — constants, upstream stager + ini parser, event_types builder, orchestrator. - `scripts/export/petrobras_3w/` — parquet writer, validator (logs the pin + asserts dataset version + event_types row count/PK/transient invariants), schema doc generator (md/json/sql + README + LICENSE), website integrator, orchestrator. - `parquet/petrobras_3w/` — published deliverables: `event_types.parquet` (10 rows), `schema.md` (with 27-sensor glossary mirrored from upstream), `schema.json`, `schema.sql`, `README.md`, `LICENSE-3W-DATA.md` (CC BY 4.0 + attribution). - `tests/petrobras_3w/test_smoke.py` — end-to-end coverage of the scaffolding (event_types contents, doc generation, website idempotency, pin logging, validator abort paths). - `tests/fixtures/petrobras_3w/dataset/dataset.ini` — byte-identical mirror of upstream at the pinned tag, so tests run without a clone. Files modified: - Root `README.md` and `parquet/index.html` patched via the new website integrator (new Petrobras 3W entry / tab parallel to Argentina, Volve, FORCE 2020). - `.gitignore` — exclude the `data/petrobras_3w/` staging dir and the intermediate `database/petrobras_3w.duckdb` file. Notes for next iteration: - #20 (instances.parquet), #21 (wells.parquet), #22 (observations/) all extend `scripts/transform/petrobras_3w/orchestrator.py` and `scripts/export/petrobras_3w/{parquet_writer,validator,schema_doc_generator}.py`. - The validator's structural checks (#19 only covers event_types row-count, PK uniqueness, has_transient/transient_code invariants + pin assertion) need to be extended with the remaining seven rules from CONTEXT.md as later tables land. - Pre-existing test failure in `scripts/export/test_integration.py` (Volve-era, missing `parquet/wells.parquet`) is unrelated to this slice and remains as-is. Closes #19 Co-Authored-By: Claude Opus 4.7 --- .gitignore | 5 + README.md | 30 ++ parquet/index.html | 159 +++++++ parquet/petrobras_3w/LICENSE-3W-DATA.md | 21 + parquet/petrobras_3w/README.md | 57 +++ parquet/petrobras_3w/event_types.parquet | Bin 0 -> 1464 bytes parquet/petrobras_3w/schema.json | 56 +++ parquet/petrobras_3w/schema.md | 59 +++ parquet/petrobras_3w/schema.sql | 15 + scripts/export/petrobras_3w/__init__.py | 0 scripts/export/petrobras_3w/orchestrator.py | 51 +++ scripts/export/petrobras_3w/parquet_writer.py | 22 + .../petrobras_3w/schema_doc_generator.py | 413 ++++++++++++++++++ scripts/export/petrobras_3w/validator.py | 145 ++++++ .../export/petrobras_3w/website_integrator.py | 309 +++++++++++++ scripts/transform/petrobras_3w/__init__.py | 0 scripts/transform/petrobras_3w/constants.py | 18 + .../petrobras_3w/event_types_builder.py | 51 +++ .../transform/petrobras_3w/orchestrator.py | 34 ++ .../transform/petrobras_3w/upstream_stager.py | 124 ++++++ .../fixtures/petrobras_3w/dataset/dataset.ini | 104 +++++ tests/petrobras_3w/__init__.py | 0 tests/petrobras_3w/test_smoke.py | 315 +++++++++++++ 23 files changed, 1988 insertions(+) create mode 100644 parquet/petrobras_3w/LICENSE-3W-DATA.md create mode 100644 parquet/petrobras_3w/README.md create mode 100644 parquet/petrobras_3w/event_types.parquet create mode 100644 parquet/petrobras_3w/schema.json create mode 100644 parquet/petrobras_3w/schema.md create mode 100644 parquet/petrobras_3w/schema.sql create mode 100644 scripts/export/petrobras_3w/__init__.py create mode 100644 scripts/export/petrobras_3w/orchestrator.py create mode 100644 scripts/export/petrobras_3w/parquet_writer.py create mode 100644 scripts/export/petrobras_3w/schema_doc_generator.py create mode 100644 scripts/export/petrobras_3w/validator.py create mode 100644 scripts/export/petrobras_3w/website_integrator.py create mode 100644 scripts/transform/petrobras_3w/__init__.py create mode 100644 scripts/transform/petrobras_3w/constants.py create mode 100644 scripts/transform/petrobras_3w/event_types_builder.py create mode 100644 scripts/transform/petrobras_3w/orchestrator.py create mode 100644 scripts/transform/petrobras_3w/upstream_stager.py create mode 100644 tests/fixtures/petrobras_3w/dataset/dataset.ini create mode 100644 tests/petrobras_3w/__init__.py create mode 100644 tests/petrobras_3w/test_smoke.py diff --git a/.gitignore b/.gitignore index b68e24c..6d0a066 100644 --- a/.gitignore +++ b/.gitignore @@ -210,6 +210,11 @@ __marimo__/ .duckdb/ database/argentina.duckdb database/argentina.duckdb.wal +database/petrobras_3w.duckdb +database/petrobras_3w.duckdb.wal + +# Petrobras 3W upstream shallow clone (see ADR-0002 for the pin policy) +data/petrobras_3w/ # Tool cache/config directories (from container) .config/ diff --git a/README.md b/README.md index bd8ce5b..552984a 100644 --- a/README.md +++ b/README.md @@ -53,6 +53,36 @@ four-bucket rationale, and three more canonical query patterns live in + + +### Petrobras 3W Dataset +Labelled 1-Hz sensor-data windows from the Petrobras 3W dataset, sliced +into per-Instance Parquet files. Pinned at upstream git tag `v.1.70.0` +(dataset version `2.0.0`). This initial release publishes +the event-class lookup and documentation scaffolding; the Instance catalog, +real-Well master, and Observations time-series ship in follow-up issues. + +List every event class (NORMAL plus the nine anomaly categories) with their +TRANSIENT-arc semantics: + +```python +import duckdb + +result = duckdb.sql(""" + SELECT event_class, name, description, + has_transient, transient_code + FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/event_types.parquet' + ORDER BY event_class +""").df() +``` + +Full per-column English docs (including the 27-sensor glossary mirrored +from upstream `dataset.ini`) live in +[`parquet/petrobras_3w/README.md`](parquet/petrobras_3w/README.md). Upstream +source: (CC BY 4.0). + + + ## Access Data Browse and download files at: **https://dev-petrodb.ocortez.com** diff --git a/parquet/index.html b/parquet/index.html index 1e60c7f..85a6a7d 100644 --- a/parquet/index.html +++ b/parquet/index.html @@ -794,6 +794,12 @@

PETRODATA REPOSITORY

4 files + + + @@ -2001,6 +2007,159 @@

Source & License

+ + +
+ +
+

Download Petrobras 3W Files

+

+ Labelled 1-Hz sensor-data windows from the Petrobras 3W dataset. + Pinned at upstream git tag v.1.70.0 + (dataset version 2.0.0). This initial + release publishes the event-class lookup and documentation + scaffolding; the Instance catalog, real-Well master, and + Observations time-series ship in follow-up issues. +

+ + + +

One lookup table · pinned upstream identity logged on every publish

+
+ + +
+

About This Dataset

+

+ The Petrobras 3W dataset is a corpus of + ~2,228 labelled 1-Hz sensor-data windows recorded on + Petrobras's offshore wells, framed around at most one + anomaly event per window. The full corpus covers ten + operational regimes (NORMAL plus nine anomaly categories + such as Hydrate in Production Line and + Severe Slugging) across ~40 distinct real wells, + supplemented by simulated and hand-drawn instances. +

+

+ Petrodb pins the upstream repository at git tag + v.1.70.0 (dataset version + 2.0.0) — refreshes are + event-driven on new upstream releases, never silent. +

+
+ + +
+

Quick Start with DuckDB

+

+ List every event class with its TRANSIENT-arc semantics: +

+
+
+ + + +
+
import duckdb
+
+# List the 10 event classes and their TRANSIENT-arc semantics
+result = duckdb.sql("""
+    SELECT event_class, name, description,
+           has_transient, transient_code
+    FROM 'petrobras_3w/event_types.parquet'
+    ORDER BY event_class
+""").df()
+
+

+ More canonical patterns (per-event-class filter, joins + against the Instance catalog, single-Instance fetch) + will land alongside the catalog and Observations files + in follow-up releases. +

+
+ + +
+

Schema Documents

+

+ Full per-column documentation is published alongside the parquets: +

+
+
+

README.md

+

Dataset overview, pinned upstream identity, query examples

+
+ → Open README.md +
+
+
+

schema.md

+

English column docs + 27-sensor glossary mirrored from upstream dataset.ini

+
+ → Open schema.md +
+
+
+

schema.json

+

Machine-readable column list, types, primary & foreign keys

+ +
+
+

schema.sql

+

DDL that mirrors the published structure in a fresh DuckDB

+
+ → Open schema.sql +
+
+
+
+ + +
+

Source & License

+

+ Upstream repository: https://github.com/petrobras/3W.git + (pinned at git tag v.1.70.0, dataset version + 2.0.0). +

+

+ Licensed under Creative Commons Attribution 4.0. + All credit for the underlying measurements, labelling, and dataset + design belongs to Petrobras and the upstream maintainers. +

+
+
+ + +