From 5e2a4dfa806fcf86f1f51ef3023f3c69605bbf21 Mon Sep 17 00:00:00 2001
From: oskrgab
Date: Sun, 17 May 2026 18:54:12 +0000
Subject: [PATCH 01/11] Petrobras 3W domain context + observations-layout ADR +
upstream-tag pinning ADR
Captures the design decisions reached while grilling issue #17 / PRD #18:
- CONTEXT.md adds the 3W language (Well, Instance, Well kind, Event class,
Transient label, Observation, Source file), cross-dataset ambiguity entries
for Well and Event, and four 3W dataset sections (tables, operating
principles, output layout, pre-publish validation).
- ADR-0001 records the observations-layout decision (one file per Instance,
hive on event_class) with the four rejected alternatives.
- ADR-0002 records the upstream-tag-pinning policy (v2.0.0, event-driven
refresh) and why tracking main would silently mutate already-published bytes.
Co-Authored-By: Claude Opus 4.7
---
CONTEXT.md | 91 ++++++++++++++++++-
.../0001-petrobras-3w-observations-layout.md | 24 +++++
...2-petrobras-3w-pin-upstream-release-tag.md | 22 +++++
3 files changed, 136 insertions(+), 1 deletion(-)
create mode 100644 docs/adr/0001-petrobras-3w-observations-layout.md
create mode 100644 docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md
diff --git a/CONTEXT.md b/CONTEXT.md
index 0af8c66..4049db6 100644
--- a/CONTEXT.md
+++ b/CONTEXT.md
@@ -16,14 +16,43 @@ The human-readable code for a well (e.g. `YPF.BLO.x-8`). Treated as a label, not
**Formprod** (producing formation):
The geological formation a given `idpozo` produces from. Static per `idpozo` by construction (encoded into the ID itself), so it lives in the master table, never in the time-series.
+### Petrobras 3W dataset
+
+**Well** (3W sense, `well_id`):
+A physical offshore producing well operated by Petrobras, identified by an integer anonymized by upstream (e.g. `1`, `19`). 1-Hz sensor histories from the well are sliced into many **Instances**. The dataset covers ~40 distinct real wells. _Not the same as Argentina's Well_: there is no `formprod` split here, and no human-readable label.
+
+**Instance** (`instance_id`):
+A single contiguous 1-Hz time-window of sensor data drawn from one source, framed around at most one labeled anomaly event. Identified by the upstream source filename without extension (e.g. `WELL-00019_20120601165020`). An instance is the natural unit of an experiment: one file = one window to train, evaluate, or visualize. Length varies from ~6 hours (~21k rows) to ~3 days (~243k rows).
+
+**Well kind** (`well_kind`):
+The provenance of an Instance. One of `{real, simulated, drawn}` — `real` for instances sourced from a physical Well (~40 wells across the corpus), `simulated` for synthetic instances generated by upstream tooling, `drawn` for hand-drawn series. Simulated and drawn instances have **no** `well_id`.
+
+**Event class** (`event_class`):
+An integer code in `0..9` identifying the operational regime an Instance is framed around. `0 = NORMAL`, `1..9 = anomaly categories` (Abrupt Increase of BSW, Spurious Closure of DHSV, Severe Slugging, Flow Instability, Rapid Productivity Loss, Quick Restriction in PCK, Scaling in PCK, Hydrate in Production Line, Hydrate in Service Line). The canonical names come from `dataset.ini` upstream. _Not the same as Argentina's "event"_, which means an operational state transition row.
+
+**Transient label**:
+A per-observation `class` value equal to `event_class + 100`, marking the developing phase of an anomaly before it reaches steady state. Codes seen in the data: `101, 102, 105, 106, 107, 108, 109`. Defined only for the seven anomalies that upstream marks `TRANSIENT=True`. Events `3` (Severe Slugging) and `4` (Flow Instability) are `TRANSIENT=False` — no transient codes exist for them.
+
+**Observation**:
+One row of an `observations` table — a single 1-Hz timestamped sample with 27 sensor floats, a `class` label (per-observation regime: NULL, 0, anomaly class, or transient code), and a `state` label (well operational status).
+
+**Source file**:
+The upstream parquet file backing a single Instance. Filenames encode provenance: `WELL-_.parquet` for real, `SIMULATED_.parquet`, `DRAWN_.parquet`. In our published catalog the source filename is preserved as a column so consumers can cross-reference with upstream.
+
## Relationships
- A **Well** (`idpozo`) is the unit of identity for both the master table and the production time-series.
- **Formprod** is a static attribute of a **Well** — recorded once in the master, never repeated in monthly rows.
+- A **Well** (3W sense, `well_id`) is the parent of zero-or-more **Instances** (only when `well_kind = real`).
+- An **Instance** is the parent of many **Observations** (1-Hz rows).
+- An **Event class** is referenced by an **Instance** (the regime the window is framed around) and by an **Observation** (the per-row label, possibly via its transient offset).
+- Simulated and drawn **Instances** have no parent **Well** — their `well_id` is NULL by design.
## Flagged ambiguities
-- "well" was used to refer to both the physical wellbore and the producing-formation-specific record — resolved: in this dataset a **Well** = `idpozo` = wellbore × formprod. Use "physical wellbore" if the bore-only concept is needed.
+- "well" was used to refer to both the physical wellbore and the producing-formation-specific record — resolved: in the Argentina dataset a **Well** = `idpozo` = wellbore × formprod. Use "physical wellbore" if the bore-only concept is needed.
+- "Well" is overloaded across datasets — resolved by scope: Argentina **Well** = `idpozo`; Petrobras 3W **Well** = anonymized offshore wellbore `well_id`. Always qualify with the dataset name when the context is mixed.
+- "Event" is overloaded — resolved: in the Argentina dataset an **Event** is a row in `well_events.parquet` marking an operational-state transition; in the Petrobras 3W dataset an **Event class** is an anomaly category referenced by an Instance and by every Observation. Different concepts, different tables.
## Argentina dataset — column buckets
@@ -100,3 +129,63 @@ The export step asserts the following before writing Parquets; failure aborts pu
4. Geometry in `wells.parquet` is parseable WKB for every row that carries a `geom`.
5. Year-partition count is 22 and the sum of partition row counts equals the staged-source row count.
6. Soft-warn if any year-partition Parquet exceeds 50 MB (Cloudflare cache headroom).
+
+## Petrobras 3W dataset — tables
+
+A four-table relational shape, derived from upstream's per-instance Parquet files plus `dataset.ini`. Layouts and column names below are settled in the grill; pipeline implementation details (e.g. exact column-name policy for hyphenated source columns) are still open.
+
+**Static master** (`wells.parquet`, one row per real Well, ~40 rows):
+`well_id`, `n_instances` (count of `instances` rows), `first_ts`, `last_ts`, `n_observations` (sum across instances). Upstream anonymizes physical well attributes (no basin, field, depth, or location is published), so this table is mostly an identity-plus-statistics master. Only `well_kind = real` instances have a parent row here.
+
+**Lookup** (`event_types.parquet`, exactly 10 rows):
+`event_class` (PK, 0..9), `name` (canonical PascalCase from `dataset.ini`, e.g. `HYDRATE_IN_PRODUCTION_LINE`), `description` (human-readable, e.g. `Hydrate in Production Line`), `has_transient` (boolean, false for `{0, 3, 4}`), `transient_code` (`event_class + 100` when `has_transient`, NULL otherwise), `has_normal_prefix` (boolean, true for events that include a class=0 precursor in the data — correlates with `has_transient`).
+
+**Instance catalog** (`instances.parquet`, one row per Instance, ~2,228 rows):
+`instance_id` (PK, source filename without extension), `well_kind` (enum: `real | simulated | drawn`), `well_id` (FK to `wells.parquet`, NULL when `well_kind != real`), `event_class` (FK to `event_types.parquet`), `start_ts`, `end_ts`, `duration_s` (derived), `n_rows`, `n_rows_warmup_null` (rows where `class IS NULL`), `n_rows_normal` (rows where `class = 0`), `n_rows_transient` (rows where `class = event_class + 100`; NULL when `has_transient = false`), `n_rows_steady` (rows where `class = event_class`), `source_file` (upstream filename for cross-reference), `source_url` (URL to the published Observations parquet). The four `n_rows_*` columns let corpus-wide balance and labeled-mass queries run purely against this catalog without scanning Observations.
+
+**Observations time-series** (`observations/event_class=N/.parquet`, ~2,228 files):
+Hive-partitioned by `event_class` only. Each file is a single Instance's 1-Hz rows. Columns: the 27 source sensor columns + `class` (per-observation label, may include the +100 transient codes) + `state` (well operational status) + `timestamp` + three added constant columns: `instance_id`, `well_id`, `well_kind`. `event_class` is provided by the hive partition, not stored in the file body.
+
+## Petrobras 3W dataset — operating principles
+
+- **Source column names preserved verbatim**, including hyphens (`P-PDG`, `ABER-CKGL`, `ESTADO-SDV-GL`). Consumers must double-quote these identifiers in SQL. Same fidelity principle applied to Argentina, applied here to hyphenated forms instead of Spanish forms.
+- **Source fidelity over smoothing**: the `class` column's NULL warmup prefix (~1 hour at the start of every real-Well instance) is preserved; the `+100` transient offset is preserved as a raw code (lookup table explains the convention); per-instance file boundaries match upstream's per-instance file boundaries. No row-level reinterpretation.
+- **Per-row provenance**: `instance_id`, `well_id`, `well_kind` are duplicated into each Observations file as constant columns (RLE-encoded, negligible storage). The dataset stays self-describing without filename-parsing tribal knowledge.
+- **TRANSIENT semantics**: instances of an event with `has_transient = true` carry a `NORMAL → TRANSIENT → STEADY` arc in their `class` column. Instances of an event with `has_transient = false` (events 3 and 4) carry **only** the steady class — no `NORMAL` precursor either. Consumers training early-detection models should filter on `has_transient`.
+- **3W toolkit artefacts are excluded**: per-event detector hyperparameters (`WINDOW`, `STEP`) and fold-config knobs (`EXTRA_INSTANCES_TRAINING`) live in upstream's `dataset.ini` but belong to the *3W toolkit*, not to the dataset itself. Petrodb publishes the dataset; toolkit and folds remain the consumer's responsibility.
+- **Pure DuckDB SQL transform**: pipeline is read-with-`filename=true` → parse filename into catalog columns → `COPY ... PARTITION_BY (event_class)`. Polars permitted only where SQL becomes unreadable; pandas is not used. Same rule as the Argentina pipeline.
+- **Upstream pinned to a release tag** (currently `v2.0.0`). Refreshes are event-driven on new upstream tags, not calendar-driven. Past upstream releases have retroactively changed sensor values inside the same `instance_id`, so tracking `main` would silently mutate already-published parquet bytes. See [ADR-0002](docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md).
+
+## Petrobras 3W dataset — output layout
+
+```
+parquet/petrobras_3w/
+├── wells.parquet
+├── event_types.parquet
+├── instances.parquet
+├── observations/
+│ ├── event_class=0/.parquet # 594 files
+│ ├── event_class=1/.parquet # 128 files
+│ ├── …
+│ └── event_class=9/.parquet # 207 files
+├── observations/_files.json # manifest of all Observations URLs for httpfs consumers
+├── schema.md / schema.json / schema.sql
+├── README.md
+└── LICENSE-3W-DATA.md # CC BY 4.0 mirror with upstream attribution
+```
+
+Rationale: 2,228 instances × ~0.4–1 MB each keep every file well under Cloudflare's edge-cache size target. Hive partitioning by `event_class` prunes the dominant ML query pattern ("train on class N"). One-file-per-instance preserves the upstream conceptual unit and makes single-instance fetch a direct GET (the ML training-loop pattern). Catalog tables answer corpus-wide questions without touching Observations.
+
+## Petrobras 3W dataset — pre-publish validation
+
+The export step asserts the following before writing Parquets; failure aborts publish.
+
+1. `instances.instance_id` is unique.
+2. Every `instance_id` referenced by an Observations file exists in `instances.parquet`.
+3. Every non-NULL `instances.well_id` exists in `wells.parquet` (FK integrity); NULL only when `well_kind != real`.
+4. Every `instances.event_class` exists in `event_types.parquet` (FK integrity).
+5. Within each Observations file: only one distinct `(instance_id, well_id, well_kind)` triple; `class ∈ {NULL, 0, event_class, transient_code(event_class)}` for the file's `event_class`; row count equals `instances.n_rows`; `timestamp` is strictly monotonic at exactly 1-second cadence.
+6. For Instances with `has_transient = false` event class (events 3, 4): no `class` value ≥ 100 appears; no `class = 0` appears.
+7. Real-Well coverage equals upstream's stated count (currently 42 — fail-loud if our derived `wells.parquet` rowcount disagrees, to catch upstream version drift).
+8. `wells.parquet` rows are limited to the union of `instances.well_id WHERE well_kind = 'real'`.
+9. Soft-warn if any Observations Parquet exceeds 50 MB (Cloudflare cache headroom).
diff --git a/docs/adr/0001-petrobras-3w-observations-layout.md b/docs/adr/0001-petrobras-3w-observations-layout.md
new file mode 100644
index 0000000..a915042
--- /dev/null
+++ b/docs/adr/0001-petrobras-3w-observations-layout.md
@@ -0,0 +1,24 @@
+# Petrobras 3W observations layout: one file per Instance, hive on `event_class`
+
+**Status:** accepted
+
+## Context
+
+The Petrobras 3W dataset is a corpus of ~2,228 labeled 1-Hz sensor-data windows ("Instances"), totalling ~1.74 GB compressed. Instance length varies from ~21k rows (~6 h) to ~243k rows (~3 days). The ML workload that dominates is event-detector training, which has two distinct access patterns: (a) "load all instances of event class N" and (b) "load one specific Instance for visualization or single-window training." Petrodb is published as static Parquet files behind Cloudflare with a soft per-file size target of 50 MB (CONTEXT.md `Argentina dataset — pre-publish validation` rule 6).
+
+## Decision
+
+Publish `observations/` as `observations/event_class=N/.parquet` — hive-partitioned by `event_class` only, with one file per Instance preserving upstream's per-instance file boundaries. Inside each file, in addition to the 30 upstream columns, store `instance_id`, `well_id`, `well_kind` as constant columns (RLE-encoded, negligible cost) so the dataset stays self-describing under any future restructuring.
+
+## Considered alternatives
+
+- **Single coalesced `observations.parquet`** with row-group sort on `(event_class, instance_id, timestamp)`. Best for corpus-wide aggregate scans; rejected because the 1.74 GB single file busts the per-file cache target and a cold CDN miss serves the entire file even when DuckDB only needs a row-group range.
+- **Hive on `event_class` only, coalesced within partition** (10 files, ~50–500 MB each). Rejected: partition 0 (NORMAL, 594 instances) alone is hundreds of MB; busts the cache target; also discards the per-Instance file boundary, which is the natural unit of the ML workload.
+- **Hive on `(event_class × well_id)`** (~100–150 partitions of ~10–20 MB). Fits the cache target and is efficient for combined class+well filters, but the dominant pattern is single-Instance fetch — pulling a 15 MB partition to read a 0.5 MB Instance is wasted bandwidth and worse cold-start latency.
+- **Thin mirror of upstream** (`dataset/N/*.parquet`). Rejected: leaves `well_id`, `event_class`, `well_kind`, `instance_id` encoded only in filenames and directory names, requiring tribal knowledge to query. Contradicts petrodb's mission of clean, non-redundant relational schemas (CONTEXT.md, line 3).
+
+## Consequences
+
+- Corpus-wide aggregate scans require 2,228 parallel HTTP requests rather than a single large GET. DuckDB httpfs handles this concurrently, and the catalog tables (`instances.parquet`, `wells.parquet`, `event_types.parquet`) answer most aggregate questions without touching `observations/` at all.
+- The published URL space is committed: `observations/event_class=N/.parquet` becomes part of the public API. Reorganizing later breaks consumers.
+- Adding `instance_id` / `well_id` / `well_kind` as constant columns means future coalescing (if ever needed) is a transparent operation — no schema migration for downstream consumers.
diff --git a/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md b/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md
new file mode 100644
index 0000000..465bb60
--- /dev/null
+++ b/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md
@@ -0,0 +1,22 @@
+# Pin Petrobras 3W upstream to a release tag, not `main`
+
+**Status:** accepted
+
+## Context
+
+Petrobras 3W uses semantic versioning (`VERSIONING.md` upstream). Past minor releases have reshaped instance contents in-place: 1.1.0 added/removed instances, adjusted expert labels, and corrected historian-tag misconfigurations that retroactively changed sensor values. The same `instance_id` can therefore have different bytes between upstream commits, with no rename or URL change to signal it. Petrodb publishes parquet at stable URLs, so silently-changing bytes would propagate as silently-changing training data to downstream consumers.
+
+## Decision
+
+The Petrobras 3W pipeline reads from a pinned upstream release tag (currently `v2.0.0`). Refreshes are event-driven (when upstream cuts a new tag and we've reviewed the release notes), not calendar-driven. The current pinned tag is recorded in `parquet/petrobras_3w/README.md` and emitted in the pipeline's validation logs.
+
+## Considered alternatives
+
+- **Track upstream `main`.** Rejected — re-pulling at any time risks silently mutating already-published `instance_id`s (sensor-value corrections, label adjustments). Consumers relying on bytes-stable URLs would observe non-reproducible behaviour.
+- **Pin to a commit SHA.** Strongest reproducibility, but upstream cuts release tags deliberately at coherent dataset states; an arbitrary SHA boundary is harder to reason about and harder to compare against the release notes. The marginal reproducibility gain over a tag is small enough not to justify the extra friction.
+
+## Consequences
+
+- Pre-publish validation rule 7 (real-Well coverage equals upstream's stated count) implicitly enforces the pin — an accidental tag change shows up as a row-count mismatch, not a silent corruption.
+- Petrodb's release cadence for this dataset is bounded above by upstream's release cadence. If upstream goes dormant we publish stale data; that's acceptable for the reproducibility guarantee it buys.
+- The pinned tag is part of the dataset's published metadata; bumping it is a deliberate, reviewable action with release-note context, not an automated refresh.
From 6dbe510ecc983c07f1073546b83208022c6e88a5 Mon Sep 17 00:00:00 2001
From: oskrgab
Date: Sun, 17 May 2026 19:34:12 +0000
Subject: [PATCH 02/11] Covered additional context about the dataset tag
---
CONTEXT.md | 4 ++--
docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md | 4 +++-
2 files changed, 5 insertions(+), 3 deletions(-)
diff --git a/CONTEXT.md b/CONTEXT.md
index 4049db6..41bc722 100644
--- a/CONTEXT.md
+++ b/CONTEXT.md
@@ -154,7 +154,7 @@ Hive-partitioned by `event_class` only. Each file is a single Instance's 1-Hz ro
- **TRANSIENT semantics**: instances of an event with `has_transient = true` carry a `NORMAL → TRANSIENT → STEADY` arc in their `class` column. Instances of an event with `has_transient = false` (events 3 and 4) carry **only** the steady class — no `NORMAL` precursor either. Consumers training early-detection models should filter on `has_transient`.
- **3W toolkit artefacts are excluded**: per-event detector hyperparameters (`WINDOW`, `STEP`) and fold-config knobs (`EXTRA_INSTANCES_TRAINING`) live in upstream's `dataset.ini` but belong to the *3W toolkit*, not to the dataset itself. Petrodb publishes the dataset; toolkit and folds remain the consumer's responsibility.
- **Pure DuckDB SQL transform**: pipeline is read-with-`filename=true` → parse filename into catalog columns → `COPY ... PARTITION_BY (event_class)`. Polars permitted only where SQL becomes unreadable; pandas is not used. Same rule as the Argentina pipeline.
-- **Upstream pinned to a release tag** (currently `v2.0.0`). Refreshes are event-driven on new upstream tags, not calendar-driven. Past upstream releases have retroactively changed sensor values inside the same `instance_id`, so tracking `main` would silently mutate already-published parquet bytes. See [ADR-0002](docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md).
+- **Upstream pinned to a git tag that ships a specific dataset version.** Currently pinned to git tag `v.1.70.0`, which ships upstream dataset version `2.0.0`. Refreshes are event-driven on new upstream tags, not calendar-driven. Past upstream releases have retroactively changed sensor values inside the same `instance_id`, so tracking `main` would silently mutate already-published parquet bytes. Upstream uses two version namespaces — git tags (`v.1.NN.0`) and dataset semver in `dataset/README.md` (`1.0.0`, `1.1.0`, `1.1.1`, `2.0.0`, …); we pin the git tag because that is what `git clone --branch` accepts, and we record the dataset version because that is what identifies the data shape. See [ADR-0002](docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md).
## Petrobras 3W dataset — output layout
@@ -186,6 +186,6 @@ The export step asserts the following before writing Parquets; failure aborts pu
4. Every `instances.event_class` exists in `event_types.parquet` (FK integrity).
5. Within each Observations file: only one distinct `(instance_id, well_id, well_kind)` triple; `class ∈ {NULL, 0, event_class, transient_code(event_class)}` for the file's `event_class`; row count equals `instances.n_rows`; `timestamp` is strictly monotonic at exactly 1-second cadence.
6. For Instances with `has_transient = false` event class (events 3, 4): no `class` value ≥ 100 appears; no `class = 0` appears.
-7. Real-Well coverage equals upstream's stated count (currently 42 — fail-loud if our derived `wells.parquet` rowcount disagrees, to catch upstream version drift).
+7. Real-Well coverage equals the count derived from upstream filename prefixes at the pinned git tag (currently 40 distinct `WELL-NNNNN` prefixes at `v.1.70.0` / dataset version 2.0.0 — fail-loud if our derived `wells.parquet` rowcount disagrees, to catch upstream version drift). Upstream's `dataset/README.md` states "42 real wells covered", but only 40 IDs (`00001..00016`, `00019..00042`) actually appear in instance filenames; IDs `00017` and `00018` are absent. The validator pins on the observed 40.
8. `wells.parquet` rows are limited to the union of `instances.well_id WHERE well_kind = 'real'`.
9. Soft-warn if any Observations Parquet exceeds 50 MB (Cloudflare cache headroom).
diff --git a/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md b/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md
index 465bb60..c44fde5 100644
--- a/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md
+++ b/docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md
@@ -8,7 +8,9 @@ Petrobras 3W uses semantic versioning (`VERSIONING.md` upstream). Past minor rel
## Decision
-The Petrobras 3W pipeline reads from a pinned upstream release tag (currently `v2.0.0`). Refreshes are event-driven (when upstream cuts a new tag and we've reviewed the release notes), not calendar-driven. The current pinned tag is recorded in `parquet/petrobras_3w/README.md` and emitted in the pipeline's validation logs.
+The Petrobras 3W pipeline reads from a pinned upstream git tag that ships a specific upstream *dataset version*. The currently pinned git tag is `v.1.70.0`, which ships dataset version `2.0.0`. Refreshes are event-driven (when upstream cuts a new git tag — typically corresponding to a new dataset version — and we've reviewed the release notes), not calendar-driven. The current pinned git tag and dataset version are recorded in `parquet/petrobras_3w/README.md` and emitted in the pipeline's validation logs.
+
+Note on upstream versioning: upstream uses two distinct version namespaces. Git tags are formatted `v.1.NN.0` (with the dot after `v`); these are the only mechanism a clone can pin to. The *dataset* itself carries a separate semver in `dataset/README.md` (`1.0.0`, `1.1.0`, `1.1.1`, `2.0.0`, …) — this is the version that identifies the data shape and content. The dataset version is what consumers care about; the git tag is the pinning mechanism that delivers it byte-stably. A single dataset version is typically shipped by many consecutive git tags (toolkit/docs fixes don't bump the dataset); for stability we pin to the latest available git tag for the chosen dataset version.
## Considered alternatives
From 88255cc1ae9694692471b6aada6ba3bcf68199e2 Mon Sep 17 00:00:00 2001
From: oskrgab
Date: Sun, 17 May 2026 19:46:27 +0000
Subject: [PATCH 03/11] Petrobras 3W skeleton pipeline + event_types.parquet
(#19)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Tracer-bullet slice that establishes the transform/export module layout
for the Petrobras 3W dataset and emits its smallest deliverable table.
Subsequent slices (#20, #21, #22) add the Instance catalog, real-Well
master, and Observations hive partition without re-touching this
scaffolding.
Key decisions:
- Source of truth for event_types is upstream `dataset.ini`. The
pipeline parses it via configparser (with `optionxform=str` to
preserve hyphenated sensor column names like `P-PDG` and `QGL`)
rather than hard-coding canonical values — a future upstream rename
surfaces as a parse-time change rather than silent drift.
- The shallow-clone is idempotent: presence of
`/dataset/dataset.ini` short-circuits the `git clone`, so
the smoke test can stand in a fixture upstream tree without
network access.
- Pinned identity (git tag `v.1.70.0`, dataset version `2.0.0`) is
emitted via the export's `petrobras_3w.export` logger before any
parquet write, and is also recorded in `schema.json`, `schema.sql`,
`README.md`, and the static-site tab. A parsed `dataset_version`
that disagrees with the pin aborts publish (ADR-0002 event-driven
refresh).
- `has_normal_prefix` is materialised equal to `has_transient` per
CONTEXT.md (events 0, 3, 4 carry only the steady class — no NORMAL
precursor).
- DDL identifier quoting in `schema.sql` is general (`_quote_identifier`)
so the hyphenated sensor columns added in #22 round-trip without
re-touching the writer.
- Website integration uses the same sentinel-comment idempotency
pattern as the Argentina integrator.
Files added:
- `scripts/transform/petrobras_3w/` — constants, upstream stager + ini
parser, event_types builder, orchestrator.
- `scripts/export/petrobras_3w/` — parquet writer, validator (logs the
pin + asserts dataset version + event_types row count/PK/transient
invariants), schema doc generator (md/json/sql + README + LICENSE),
website integrator, orchestrator.
- `parquet/petrobras_3w/` — published deliverables: `event_types.parquet`
(10 rows), `schema.md` (with 27-sensor glossary mirrored from
upstream), `schema.json`, `schema.sql`, `README.md`,
`LICENSE-3W-DATA.md` (CC BY 4.0 + attribution).
- `tests/petrobras_3w/test_smoke.py` — end-to-end coverage of the
scaffolding (event_types contents, doc generation, website
idempotency, pin logging, validator abort paths).
- `tests/fixtures/petrobras_3w/dataset/dataset.ini` — byte-identical
mirror of upstream at the pinned tag, so tests run without a clone.
Files modified:
- Root `README.md` and `parquet/index.html` patched via the new
website integrator (new Petrobras 3W entry / tab parallel to
Argentina, Volve, FORCE 2020).
- `.gitignore` — exclude the `data/petrobras_3w/` staging dir and the
intermediate `database/petrobras_3w.duckdb` file.
Notes for next iteration:
- #20 (instances.parquet), #21 (wells.parquet), #22 (observations/) all
extend `scripts/transform/petrobras_3w/orchestrator.py` and
`scripts/export/petrobras_3w/{parquet_writer,validator,schema_doc_generator}.py`.
- The validator's structural checks (#19 only covers event_types
row-count, PK uniqueness, has_transient/transient_code invariants
+ pin assertion) need to be extended with the remaining seven rules
from CONTEXT.md as later tables land.
- Pre-existing test failure in `scripts/export/test_integration.py`
(Volve-era, missing `parquet/wells.parquet`) is unrelated to this
slice and remains as-is.
Closes #19
Co-Authored-By: Claude Opus 4.7
---
.gitignore | 5 +
README.md | 30 ++
parquet/index.html | 159 +++++++
parquet/petrobras_3w/LICENSE-3W-DATA.md | 21 +
parquet/petrobras_3w/README.md | 57 +++
parquet/petrobras_3w/event_types.parquet | Bin 0 -> 1464 bytes
parquet/petrobras_3w/schema.json | 56 +++
parquet/petrobras_3w/schema.md | 59 +++
parquet/petrobras_3w/schema.sql | 15 +
scripts/export/petrobras_3w/__init__.py | 0
scripts/export/petrobras_3w/orchestrator.py | 51 +++
scripts/export/petrobras_3w/parquet_writer.py | 22 +
.../petrobras_3w/schema_doc_generator.py | 413 ++++++++++++++++++
scripts/export/petrobras_3w/validator.py | 145 ++++++
.../export/petrobras_3w/website_integrator.py | 309 +++++++++++++
scripts/transform/petrobras_3w/__init__.py | 0
scripts/transform/petrobras_3w/constants.py | 18 +
.../petrobras_3w/event_types_builder.py | 51 +++
.../transform/petrobras_3w/orchestrator.py | 34 ++
.../transform/petrobras_3w/upstream_stager.py | 124 ++++++
.../fixtures/petrobras_3w/dataset/dataset.ini | 104 +++++
tests/petrobras_3w/__init__.py | 0
tests/petrobras_3w/test_smoke.py | 315 +++++++++++++
23 files changed, 1988 insertions(+)
create mode 100644 parquet/petrobras_3w/LICENSE-3W-DATA.md
create mode 100644 parquet/petrobras_3w/README.md
create mode 100644 parquet/petrobras_3w/event_types.parquet
create mode 100644 parquet/petrobras_3w/schema.json
create mode 100644 parquet/petrobras_3w/schema.md
create mode 100644 parquet/petrobras_3w/schema.sql
create mode 100644 scripts/export/petrobras_3w/__init__.py
create mode 100644 scripts/export/petrobras_3w/orchestrator.py
create mode 100644 scripts/export/petrobras_3w/parquet_writer.py
create mode 100644 scripts/export/petrobras_3w/schema_doc_generator.py
create mode 100644 scripts/export/petrobras_3w/validator.py
create mode 100644 scripts/export/petrobras_3w/website_integrator.py
create mode 100644 scripts/transform/petrobras_3w/__init__.py
create mode 100644 scripts/transform/petrobras_3w/constants.py
create mode 100644 scripts/transform/petrobras_3w/event_types_builder.py
create mode 100644 scripts/transform/petrobras_3w/orchestrator.py
create mode 100644 scripts/transform/petrobras_3w/upstream_stager.py
create mode 100644 tests/fixtures/petrobras_3w/dataset/dataset.ini
create mode 100644 tests/petrobras_3w/__init__.py
create mode 100644 tests/petrobras_3w/test_smoke.py
diff --git a/.gitignore b/.gitignore
index b68e24c..6d0a066 100644
--- a/.gitignore
+++ b/.gitignore
@@ -210,6 +210,11 @@ __marimo__/
.duckdb/
database/argentina.duckdb
database/argentina.duckdb.wal
+database/petrobras_3w.duckdb
+database/petrobras_3w.duckdb.wal
+
+# Petrobras 3W upstream shallow clone (see ADR-0002 for the pin policy)
+data/petrobras_3w/
# Tool cache/config directories (from container)
.config/
diff --git a/README.md b/README.md
index bd8ce5b..552984a 100644
--- a/README.md
+++ b/README.md
@@ -53,6 +53,36 @@ four-bucket rationale, and three more canonical query patterns live in
+
+
+### Petrobras 3W Dataset
+Labelled 1-Hz sensor-data windows from the Petrobras 3W dataset, sliced
+into per-Instance Parquet files. Pinned at upstream git tag `v.1.70.0`
+(dataset version `2.0.0`). This initial release publishes
+the event-class lookup and documentation scaffolding; the Instance catalog,
+real-Well master, and Observations time-series ship in follow-up issues.
+
+List every event class (NORMAL plus the nine anomaly categories) with their
+TRANSIENT-arc semantics:
+
+```python
+import duckdb
+
+result = duckdb.sql("""
+ SELECT event_class, name, description,
+ has_transient, transient_code
+ FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/event_types.parquet'
+ ORDER BY event_class
+""").df()
+```
+
+Full per-column English docs (including the 27-sensor glossary mirrored
+from upstream `dataset.ini`) live in
+[`parquet/petrobras_3w/README.md`](parquet/petrobras_3w/README.md). Upstream
+source: (CC BY 4.0).
+
+
+
## Access Data
Browse and download files at: **https://dev-petrodb.ocortez.com**
diff --git a/parquet/index.html b/parquet/index.html
index 1e60c7f..85a6a7d 100644
--- a/parquet/index.html
+++ b/parquet/index.html
@@ -794,6 +794,12 @@
PETRODATA REPOSITORY
4 files
+
+
+
@@ -2001,6 +2007,159 @@
Source & License
+
+
+
+
+
+
Download Petrobras 3W Files
+
+ Labelled 1-Hz sensor-data windows from the Petrobras 3W dataset.
+ Pinned at upstream git tag v.1.70.0
+ (dataset version 2.0.0). This initial
+ release publishes the event-class lookup and documentation
+ scaffolding; the Instance catalog, real-Well master, and
+ Observations time-series ship in follow-up issues.
+
One lookup table · pinned upstream identity logged on every publish
+
+
+
+
+
About This Dataset
+
+ The Petrobras 3W dataset is a corpus of
+ ~2,228 labelled 1-Hz sensor-data windows recorded on
+ Petrobras's offshore wells, framed around at most one
+ anomaly event per window. The full corpus covers ten
+ operational regimes (NORMAL plus nine anomaly categories
+ such as Hydrate in Production Line and
+ Severe Slugging) across ~40 distinct real wells,
+ supplemented by simulated and hand-drawn instances.
+
+
+ Petrodb pins the upstream repository at git tag
+ v.1.70.0 (dataset version
+ 2.0.0) — refreshes are
+ event-driven on new upstream releases, never silent.
+
+
+
+
+
+
Quick Start with DuckDB
+
+ List every event class with its TRANSIENT-arc semantics:
+
+
+
+
+
+
+
+
import duckdb
+
+# List the 10 event classes and their TRANSIENT-arc semantics
+result = duckdb.sql("""
+ SELECT event_class, name, description,
+ has_transient, transient_code
+ FROM 'petrobras_3w/event_types.parquet'
+ ORDER BY event_class
+""").df()
+
+
+ More canonical patterns (per-event-class filter, joins
+ against the Instance catalog, single-Instance fetch)
+ will land alongside the catalog and Observations files
+ in follow-up releases.
+
+
+
+
+
+
Schema Documents
+
+ Full per-column documentation is published alongside the parquets:
+
+ Licensed under Creative Commons Attribution 4.0.
+ All credit for the underlying measurements, labelling, and dataset
+ design belongs to Petrobras and the upstream maintainers.
+
- More canonical patterns (single-Instance fetch, per-Well
- cross-validation splits) will land alongside the Observations
- files in follow-up releases.
+ The per-Instance Observations files are accessible via the
+ hive-partitioned URL pattern
+ observations/event_class=N/<instance_id>.parquet.
+ Each file embeds instance_id, well_id,
+ and well_kind as constant columns, so corpus-wide
+ queries against a single event class do not need to join the
+ catalog:
+
+
+
+
+
+
+
-- All real-Well Hydrate-in-Production-Line observations
+SELECT instance_id, well_id, "timestamp", "P-PDG", "T-PDG", class
+FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/observations/event_class=8/*.parquet'
+WHERE well_kind = 'real';
+
diff --git a/scripts/transform/petrobras_3w/observations_builder.py b/scripts/transform/petrobras_3w/observations_builder.py
new file mode 100644
index 0000000..611398c
--- /dev/null
+++ b/scripts/transform/petrobras_3w/observations_builder.py
@@ -0,0 +1,74 @@
+"""Build the `observations` view over the staged upstream tree.
+
+Defines a view rather than a materialised table because the full
+Observations corpus (2,228 files × tens of thousands of rows × 30
+columns) is too large to keep in RAM. The view scans the staged
+per-Instance parquets via a single `read_parquet(..., filename=true,
+union_by_name=true)`, then derives four columns from the source
+filename:
+
+- `instance_id` — upstream filename without `.parquet`
+- `event_class` — the integer carved out of `/dataset/N/`
+- `well_id` — leading-zero-stripped `WELL-NNNNN` prefix (NULL for
+ simulated / drawn instances)
+- `well_kind` — `real` / `simulated` / `drawn` keyed on the prefix
+
+The view exists for the validator (rules 2, 5, 6 in CONTEXT.md need
+per-Observation queries); the actual per-Instance parquets are written
+by `parquet_writer.write_observations`, which reads the staged sources
+directly rather than going through the view (one read per file vs. an
+O(n²) re-scan against the view).
+
+Per ADR-0001 the published layout hive-partitions by `event_class`
+only, so `event_class` is NOT stored in the per-file body — but it is
+present in the view so the validator can pivot on it.
+
+`union_by_name=true` keeps the view tolerant of the test fixtures,
+which only carry `timestamp` and `class`; columns missing from a file
+become NULL in the view.
+"""
+
+from __future__ import annotations
+
+from pathlib import Path
+
+import duckdb
+
+
+def build(con: duckdb.DuckDBPyConnection, staging_dir: Path) -> None:
+ """Create or replace the `observations` view over staged sources."""
+ staging_dir = Path(staging_dir)
+ glob_pattern = str(staging_dir / "dataset" / "*" / "*.parquet")
+
+ con.execute(
+ f"""
+ CREATE OR REPLACE VIEW observations AS
+ SELECT
+ * EXCLUDE (filename, _instance_id_, _event_class_),
+ _instance_id_ AS instance_id,
+ _event_class_ AS event_class,
+ CASE
+ WHEN starts_with(_instance_id_, 'WELL-')
+ THEN CAST(regexp_extract(_instance_id_, '^WELL-0*([0-9]+)_', 1)
+ AS INTEGER)
+ END AS well_id,
+ CASE
+ WHEN starts_with(_instance_id_, 'WELL-') THEN 'real'
+ WHEN starts_with(_instance_id_, 'SIMULATED_') THEN 'simulated'
+ WHEN starts_with(_instance_id_, 'DRAWN_') THEN 'drawn'
+ END AS well_kind
+ FROM (
+ SELECT
+ *,
+ regexp_extract(filename, '/([^/]+)\\.parquet$', 1)
+ AS _instance_id_,
+ CAST(regexp_extract(filename, '/dataset/([0-9]+)/', 1)
+ AS INTEGER) AS _event_class_
+ FROM read_parquet(
+ '{glob_pattern}',
+ filename=true,
+ union_by_name=true
+ )
+ )
+ """
+ )
diff --git a/scripts/transform/petrobras_3w/orchestrator.py b/scripts/transform/petrobras_3w/orchestrator.py
index ec6f71a..edbfad0 100644
--- a/scripts/transform/petrobras_3w/orchestrator.py
+++ b/scripts/transform/petrobras_3w/orchestrator.py
@@ -2,10 +2,11 @@
Stages the pinned upstream tag, parses `dataset.ini`, builds the
`event_types` lookup, aggregates every staged instance file into the
-`instances` catalog, and derives the `wells` master from those real-Well
-instances. The remaining slice (#22 — observations) extends this file
-with the per-Instance Observations writer; staging + ini-parse + the
-catalog tables are shared.
+`instances` catalog, derives the `wells` master from those real-Well
+instances, and exposes the `observations` view (a thin enrichment over
+the staged per-Instance parquets) for the validator. The catalog
+tables are materialised; `observations` stays as a view so the full
+corpus never has to live in RAM.
"""
from __future__ import annotations
@@ -17,6 +18,7 @@
from scripts.transform.petrobras_3w import (
event_types_builder,
instances_builder,
+ observations_builder,
upstream_stager,
wells_builder,
)
@@ -40,5 +42,6 @@ def run(db_path: Path, staging_dir: Path) -> DatasetIni:
event_types_builder.build(con, dataset_ini)
instances_builder.build(con, staging_dir)
wells_builder.build(con)
+ observations_builder.build(con, staging_dir)
return dataset_ini
diff --git a/tests/petrobras_3w/conftest.py b/tests/petrobras_3w/conftest.py
index 998e97b..7d88e3a 100644
--- a/tests/petrobras_3w/conftest.py
+++ b/tests/petrobras_3w/conftest.py
@@ -1,10 +1,10 @@
"""Shared fixture helpers for the Petrobras 3W pipeline tests.
The smoke test populates a fixture upstream tree (just `dataset.ini`)
-into a per-test `tmp_path` and short-circuits the shallow-clone. Once
-issue #20 lands, the transform pipeline also reads per-Instance parquet
-files from `/dataset/N/*.parquet`, so each test that exercises
-the orchestrator needs a small set of those files alongside the ini.
+into a per-test `tmp_path` and short-circuits the shallow-clone. The
+transform pipeline reads per-Instance parquet files from
+`/dataset/N/*.parquet`, so each test that exercises the
+orchestrator needs a small set of those files alongside the ini.
`build_instance_parquets` materializes them in a deterministic, minimal
shape:
@@ -27,9 +27,12 @@
pins on, and using the exact upstream gap means the happy-path tests
exercise rule 7 against a realistic catalog.
-Each file carries only the columns the instances builder needs
-(`timestamp`, `class`); the full 27-sensor schema lands with the
-Observations slice (#22).
+Each file carries `timestamp`, `class`, `state`, plus one hyphenated
+sensor column (`P-PDG`) so the Observations writer's column-name
+fidelity (rule from CONTEXT.md: hyphens preserved) is exercised end-
+to-end. The full 27-sensor production schema is broader; the writer
+preserves whatever the source carries, so a single representative
+hyphenated column is enough to assert the policy.
"""
from __future__ import annotations
@@ -126,6 +129,10 @@ def build_instance_parquets(staging_dir: Path) -> None:
Idempotent: re-running with an existing tree overwrites in place. Uses
DuckDB so the fixture format matches what `read_parquet` will see in
the pipeline (timestamp column written as a TIMESTAMP).
+
+ The fixture carries a hyphenated sensor column (`P-PDG`) so the
+ Observations writer's column-name fidelity policy is exercised in
+ the smoke test.
"""
dataset_root = Path(staging_dir) / "dataset"
con = duckdb.connect()
@@ -137,19 +144,27 @@ def build_instance_parquets(staging_dir: Path) -> None:
con.execute("DROP TABLE IF EXISTS staging_rows")
con.execute(
"CREATE TEMP TABLE staging_rows ("
- " timestamp TIMESTAMP,"
- " class INTEGER"
+ ' "timestamp" TIMESTAMP,'
+ ' "class" INTEGER,'
+ ' "state" INTEGER,'
+ ' "P-PDG" DOUBLE'
")"
)
rows = [
- (f"2012-01-01 00:00:{i:02d}", cls) for i, cls in enumerate(spec.classes)
+ (
+ f"2012-01-01 00:00:{i:02d}",
+ cls,
+ 0,
+ 1.0e7 + i,
+ )
+ for i, cls in enumerate(spec.classes)
]
con.executemany(
- "INSERT INTO staging_rows VALUES (?, ?)",
+ "INSERT INTO staging_rows VALUES (?, ?, ?, ?)",
rows,
)
con.execute(
- f"COPY (SELECT * FROM staging_rows ORDER BY timestamp) "
+ f'COPY (SELECT * FROM staging_rows ORDER BY "timestamp") '
f"TO '{target}' (FORMAT PARQUET)"
)
finally:
diff --git a/tests/petrobras_3w/test_smoke.py b/tests/petrobras_3w/test_smoke.py
index b1212f7..b34cc04 100644
--- a/tests/petrobras_3w/test_smoke.py
+++ b/tests/petrobras_3w/test_smoke.py
@@ -109,6 +109,7 @@ def test_pipeline_emits_event_types(tmp_path: Path) -> None:
db_path=db_path,
output_dir=out_dir,
dataset_ini=dataset_ini,
+ staging_dir=staging,
website_root=site_root,
)
@@ -178,7 +179,12 @@ def test_pipeline_emits_documentation(tmp_path: Path) -> None:
staging = _populated_staging(tmp_path)
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
for name in (
"schema.md",
@@ -191,7 +197,33 @@ def test_pipeline_emits_documentation(tmp_path: Path) -> None:
# schema.json reflects exactly the published parquet
schema_payload = json.loads((out_dir / "schema.json").read_text())
- assert set(schema_payload["tables"]) == {"event_types", "wells", "instances"}
+ assert set(schema_payload["tables"]) == {
+ "event_types",
+ "wells",
+ "instances",
+ "observations",
+ }
+ obs_cols = {c["name"] for c in schema_payload["tables"]["observations"]["columns"]}
+ # event_class lives in the hive partition, not in the file body — but
+ # we surface it as a logical column in the schema so consumers see
+ # the full data model. The constant columns instance_id / well_id /
+ # well_kind are present in every file body.
+ assert {
+ "event_class",
+ "instance_id",
+ "well_id",
+ "well_kind",
+ "timestamp",
+ "class",
+ }.issubset(obs_cols)
+ # Hyphenated source columns survive into the file body and the docs.
+ assert "P-PDG" in obs_cols
+ obs_hive_cols = {
+ c["name"]
+ for c in schema_payload["tables"]["observations"]["columns"]
+ if c.get("hive_partition")
+ }
+ assert obs_hive_cols == {"event_class"}
wells_cols = {c["name"] for c in schema_payload["tables"]["wells"]["columns"]}
assert wells_cols == {
"well_id",
@@ -285,6 +317,7 @@ def test_website_integration_is_idempotent(tmp_path: Path) -> None:
db_path=db_path,
output_dir=out_dir,
dataset_ini=dataset_ini,
+ staging_dir=staging,
website_root=site_root,
)
first_readme = (site_root / "README.md").read_text()
@@ -303,6 +336,10 @@ def test_website_integration_is_idempotent(tmp_path: Path) -> None:
assert "petrobras_3w/event_types.parquet" in first_index
assert "petrobras_3w/instances.parquet" in first_index
assert "petrobras_3w/wells.parquet" in first_index
+ # The Observations manifest is surfaced.
+ assert "petrobras_3w/observations/_files.json" in first_index
+ # The observations hive-glob query example is on the index page.
+ assert "observations/event_class=8/*.parquet" in first_index
# Pin metadata is surfaced on the site.
assert PIN_GIT_TAG in first_index
assert PIN_DATASET_VERSION in first_index
@@ -312,6 +349,7 @@ def test_website_integration_is_idempotent(tmp_path: Path) -> None:
db_path=db_path,
output_dir=out_dir,
dataset_ini=dataset_ini,
+ staging_dir=staging,
website_root=site_root,
)
assert (site_root / "README.md").read_text() == first_readme
@@ -327,7 +365,12 @@ def test_validator_logs_pinned_upstream(
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
with caplog.at_level(logging.INFO, logger="petrobras_3w.export"):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
messages = "\n".join(record.getMessage() for record in caplog.records)
assert PIN_GIT_TAG in messages
@@ -349,7 +392,12 @@ def test_validator_rejects_count_mismatch(
con.execute("DELETE FROM event_types WHERE event_class = 9")
with pytest.raises(validator.EventTypeCountError):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
assert not (out_dir / "event_types.parquet").exists()
@@ -367,7 +415,12 @@ def test_validator_rejects_dataset_version_drift(tmp_path: Path) -> None:
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
with pytest.raises(validator.UpstreamDatasetVersionError):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
assert not (out_dir / "event_types.parquet").exists()
@@ -398,7 +451,12 @@ def test_pipeline_emits_instances_catalog(tmp_path: Path) -> None:
staging = _populated_staging(tmp_path)
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
assert (out_dir / "instances.parquet").exists()
rows = _read_instances(out_dir)
@@ -424,7 +482,12 @@ def test_instances_well_kind_and_well_id(tmp_path: Path) -> None:
staging = _populated_staging(tmp_path)
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
by_id = {r["instance_id"]: r for r in _read_instances(out_dir)}
# well_kind values reflect the upstream filename prefix.
@@ -451,7 +514,12 @@ def test_instances_row_count_accounting(tmp_path: Path) -> None:
staging = _populated_staging(tmp_path)
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
rows = _read_instances(out_dir)
for row in rows:
@@ -499,7 +567,12 @@ def test_instances_source_url_matches_adr0001_pattern(tmp_path: Path) -> None:
staging = _populated_staging(tmp_path)
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
rows = _read_instances(out_dir)
for row in rows:
@@ -516,7 +589,12 @@ def test_instances_source_file_retains_extension(tmp_path: Path) -> None:
staging = _populated_staging(tmp_path)
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
rows = _read_instances(out_dir)
for row in rows:
@@ -539,7 +617,12 @@ def test_validator_rejects_duplicate_instance_id(tmp_path: Path) -> None:
)
with pytest.raises(validator.InstancePkError):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
assert not (out_dir / "instances.parquet").exists()
@@ -556,7 +639,12 @@ def test_validator_rejects_unknown_event_class(tmp_path: Path) -> None:
con.execute("UPDATE instances SET event_class = 42 WHERE event_class = 9")
with pytest.raises(validator.InstanceEventClassFkError):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
assert not (out_dir / "instances.parquet").exists()
@@ -575,7 +663,12 @@ def test_validator_rejects_well_id_violation(tmp_path: Path) -> None:
)
with pytest.raises(validator.InstanceWellKindError):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
def test_validator_rejects_transient_nullness_mismatch(tmp_path: Path) -> None:
@@ -592,7 +685,12 @@ def test_validator_rejects_transient_nullness_mismatch(tmp_path: Path) -> None:
con.execute("UPDATE instances SET n_rows_transient = 0 WHERE event_class = 0")
with pytest.raises(validator.InstanceTransientNullnessError):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
def test_validator_rejects_row_count_accounting_break(tmp_path: Path) -> None:
@@ -610,7 +708,12 @@ def test_validator_rejects_row_count_accounting_break(tmp_path: Path) -> None:
)
with pytest.raises(validator.InstanceRowCountAccountingError):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
# ---------------------------------------------------------------------------
@@ -642,7 +745,12 @@ def test_pipeline_emits_wells_master(tmp_path: Path) -> None:
staging = _populated_staging(tmp_path)
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
assert (out_dir / "wells.parquet").exists()
rows = _read_wells(out_dir)
@@ -686,7 +794,12 @@ def test_wells_excludes_simulated_and_drawn(tmp_path: Path) -> None:
staging = _populated_staging(tmp_path)
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
well_ids = {r["well_id"] for r in _read_wells(out_dir)}
# `wells.parquet` only has real-Well IDs; the simulated/drawn instances
@@ -703,7 +816,12 @@ def test_wells_aggregates_match_instances(tmp_path: Path) -> None:
staging = _populated_staging(tmp_path)
dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
con = duckdb.connect()
real_instance_count, real_row_total = con.execute(
@@ -736,7 +854,12 @@ def test_validator_rejects_well_count_mismatch(tmp_path: Path) -> None:
con.execute("DELETE FROM wells WHERE well_id = 42")
with pytest.raises(validator.WellsRowCountError):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
assert not (out_dir / "wells.parquet").exists()
@@ -760,7 +883,12 @@ def test_validator_rejects_well_id_orphan(tmp_path: Path) -> None:
)
with pytest.raises(validator.WellsIdFkError):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
def test_validator_rejects_non_real_well_in_wells(tmp_path: Path) -> None:
@@ -785,4 +913,360 @@ def test_validator_rejects_non_real_well_in_wells(tmp_path: Path) -> None:
)
with pytest.raises(validator.WellsKindError):
- export_orch.run(db_path=db_path, output_dir=out_dir, dataset_ini=dataset_ini)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
+
+
+# ---------------------------------------------------------------------------
+# observations time-series (issue #22)
+# ---------------------------------------------------------------------------
+
+
+def _run_pipeline(tmp_path: Path) -> tuple[Path, Path, Path]:
+ """Helper: run the full pipeline and return (db_path, out_dir, staging)."""
+ db_path = tmp_path / "petrobras_3w.duckdb"
+ out_dir = tmp_path / "parquet"
+ staging = _populated_staging(tmp_path)
+ dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
+ return db_path, out_dir, staging
+
+
+def test_observations_layout(tmp_path: Path) -> None:
+ """One parquet per Instance under `observations/event_class=N/...`."""
+ _, out_dir, _ = _run_pipeline(tmp_path)
+
+ obs_root = out_dir / "observations"
+ assert obs_root.is_dir()
+
+ # Primary fixtures land in their declared partitions.
+ primary = {
+ 0: "WELL-00001_20120101000000",
+ 1: "WELL-00002_20120102000000",
+ 3: "WELL-00003_20120103000000",
+ 8: "SIMULATED_00001",
+ 9: "DRAWN_00001",
+ }
+ for event_class, instance_id in primary.items():
+ target = obs_root / f"event_class={event_class}" / f"{instance_id}.parquet"
+ assert target.exists(), f"missing observations file: {target}"
+
+ # The padding event-0 fixtures are also published — total partition
+ # count of event_class=0 should match the count of event_0 instances
+ # in the catalog (1 primary + 37 padding = 38).
+ event_0_files = sorted((obs_root / "event_class=0").glob("*.parquet"))
+ assert len(event_0_files) == 38
+
+
+def test_observations_preserves_upstream_columns_and_adds_constants(
+ tmp_path: Path,
+) -> None:
+ """Body columns include hyphenated source columns + the three constants;
+ event_class is NOT stored in the file body (hive-only).
+ """
+ _, out_dir, _ = _run_pipeline(tmp_path)
+
+ target = (
+ out_dir / "observations" / "event_class=1" / "WELL-00002_20120102000000.parquet"
+ )
+ con = duckdb.connect()
+ # `hive_partitioning=false` because we are inspecting the file body
+ # specifically — DuckDB's hive autodetect would otherwise synthesize
+ # `event_class` from the parent directory name and mask the check.
+ described = con.execute(
+ f"DESCRIBE SELECT * FROM read_parquet('{target}', hive_partitioning=false)"
+ ).fetchall()
+ cols = {row[0] for row in described}
+
+ # Hyphenated source column survives the writer.
+ assert "P-PDG" in cols
+ # The body carries `class`, `state`, `timestamp` from upstream.
+ assert {"class", "state", "timestamp"}.issubset(cols)
+ # The three constant identifiers are added per row.
+ assert {"instance_id", "well_id", "well_kind"}.issubset(cols)
+ # `event_class` is NOT in the body — it lives in the hive partition.
+ assert "event_class" not in cols
+
+ # Constant columns are actually constant within the file.
+ rows = con.execute(
+ f"SELECT DISTINCT instance_id, well_id, well_kind "
+ f"FROM read_parquet('{target}', hive_partitioning=false)"
+ ).fetchall()
+ assert rows == [("WELL-00002_20120102000000", 2, "real")]
+
+
+def test_observations_simulated_well_id_null(tmp_path: Path) -> None:
+ """Simulated and drawn Instances carry NULL well_id in the body."""
+ _, out_dir, _ = _run_pipeline(tmp_path)
+
+ target = out_dir / "observations" / "event_class=8" / "SIMULATED_00001.parquet"
+ con = duckdb.connect()
+ rows = con.execute(
+ f"SELECT DISTINCT well_id, well_kind FROM read_parquet('{target}')"
+ ).fetchall()
+ assert rows == [(None, "simulated")]
+
+
+def test_observations_manifest_lists_every_file(tmp_path: Path) -> None:
+ """`observations/_files.json` enumerates every published file in
+ catalog order (sorted by event_class, then instance_id).
+ """
+ _, out_dir, _ = _run_pipeline(tmp_path)
+
+ manifest = json.loads((out_dir / "observations" / "_files.json").read_text())
+
+ # 5 primary + 37 padding = 42 published Observations files.
+ assert len(manifest) == 42
+ # Every entry resolves to an actual file.
+ for rel in manifest:
+ assert (out_dir / "observations" / rel).exists(), f"missing {rel}"
+ # Manifest contains relative paths only (no http URLs leaking in).
+ for rel in manifest:
+ assert not rel.startswith("http"), rel
+ assert rel.startswith("event_class="), rel
+ # Manifest is sorted by (event_class, instance_id) — matches catalog order.
+ assert manifest == sorted(manifest)
+
+
+def test_observations_query_pattern_in_readme(tmp_path: Path) -> None:
+ """The published README documents the hive-glob query pattern with
+ `well_kind = 'real'` filtering (parallel to the acceptance criterion).
+ """
+ _, out_dir, _ = _run_pipeline(tmp_path)
+
+ readme = (out_dir / "README.md").read_text()
+ assert "observations/event_class=8/*.parquet" in readme
+ assert "well_kind = 'real'" in readme
+
+
+def test_observations_event_class_in_schema_sql_is_quoted_safely(
+ tmp_path: Path,
+) -> None:
+ """The published schema.sql round-trips hyphenated identifiers."""
+ _, out_dir, _ = _run_pipeline(tmp_path)
+ schema_sql = (out_dir / "schema.sql").read_text()
+ # The hyphenated sensor columns must be quoted in the DDL.
+ assert '"P-PDG"' in schema_sql
+ # The observations table CREATE TABLE is emitted.
+ assert "CREATE TABLE observations" in schema_sql
+
+
+def test_validator_rejects_observations_orphan(tmp_path: Path) -> None:
+ """Rule 2: every observations `instance_id` must exist in `instances`."""
+ db_path = tmp_path / "petrobras_3w.duckdb"
+ out_dir = tmp_path / "parquet"
+ staging = _populated_staging(tmp_path)
+ dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
+
+ # Delete one Instance row but leave its staged parquet in place — the
+ # observations view still sees it, so rule 2 trips.
+ with duckdb.connect(str(db_path)) as con:
+ con.execute("DELETE FROM instances WHERE instance_id = 'SIMULATED_00001'")
+
+ with pytest.raises(validator.ObservationsInstanceFkError):
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
+
+
+def test_validator_rejects_observations_row_count_mismatch(tmp_path: Path) -> None:
+ """Rule 5: per-Instance row count must equal `instances.n_rows`."""
+ db_path = tmp_path / "petrobras_3w.duckdb"
+ out_dir = tmp_path / "parquet"
+ staging = _populated_staging(tmp_path)
+ dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
+
+ # Bump `n_rows` AND `n_rows_steady` by the same amount so the four
+ # buckets still sum to n_rows (rule from #20 stays happy) but the
+ # catalog claims one more row than the observations actually contain
+ # — exactly the divergence rule 5 watches for.
+ with duckdb.connect(str(db_path)) as con:
+ con.execute(
+ "UPDATE instances "
+ "SET n_rows = n_rows + 1, n_rows_steady = n_rows_steady + 1 "
+ "WHERE instance_id = 'WELL-00001_20120101000000'"
+ )
+
+ with pytest.raises(validator.ObservationsRowCountError):
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
+
+
+def _rewrite_staged_instance(staged_path: Path, mutate_sql: str) -> None:
+ """Read the staged instance parquet, apply `mutate_sql` against the
+ `staged` temp table, and overwrite the file in place.
+
+ Used to surgically break a single per-Observation invariant *after*
+ the transform pipeline has built the catalog, so the validator's
+ catalog-side checks (bucket accounting, FK, etc.) still pass and
+ the targeted Observation-side rule is the failing one.
+ """
+ con = duckdb.connect()
+ try:
+ con.execute(
+ f"CREATE TEMP TABLE staged AS "
+ f"SELECT * FROM read_parquet('{staged_path}', hive_partitioning=false)"
+ )
+ con.execute(mutate_sql)
+ con.execute(f"COPY staged TO '{staged_path}' (FORMAT PARQUET)")
+ finally:
+ con.close()
+
+
+def test_validator_rejects_observations_timestamp_gap(tmp_path: Path) -> None:
+ """Rule 5: timestamps must be strictly monotonic at 1-second cadence.
+
+ We run the transform pipeline first (so the catalog is clean), then
+ push the last timestamp on one Instance forward by one extra second
+ — this introduces a 2-second gap without changing the row count or
+ bucket distribution.
+ """
+ db_path = tmp_path / "petrobras_3w.duckdb"
+ out_dir = tmp_path / "parquet"
+ staging = _populated_staging(tmp_path)
+ dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
+
+ bad_path = staging / "dataset" / "1" / "WELL-00002_20120102000000.parquet"
+ _rewrite_staged_instance(
+ bad_path,
+ # Push the latest timestamp 1 second further out, creating a gap
+ # of 2 seconds between it and its predecessor.
+ "UPDATE staged "
+ 'SET "timestamp" = "timestamp" + INTERVAL \'1 second\' '
+ 'WHERE "timestamp" = (SELECT MAX("timestamp") FROM staged)',
+ )
+
+ with pytest.raises(validator.ObservationsTimestampError):
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
+
+
+def test_validator_rejects_observations_class_outside_domain(
+ tmp_path: Path,
+) -> None:
+ """Rule 5: per-row `class` must lie in {NULL, 0, event_class, transient_code}.
+
+ Replace one STEADY (class=1) row in an event-1 Instance with class=7
+ (a code that belongs to a different event). Row count and bucket
+ accounting on the catalog side are not affected by this in-place
+ swap; only the per-observation class domain trips.
+ """
+ db_path = tmp_path / "petrobras_3w.duckdb"
+ out_dir = tmp_path / "parquet"
+ staging = _populated_staging(tmp_path)
+ dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
+
+ bad_path = staging / "dataset" / "1" / "WELL-00002_20120102000000.parquet"
+ _rewrite_staged_instance(
+ bad_path,
+ # Replace exactly one of the STEADY (class=1) rows with class=7.
+ "UPDATE staged SET class = 7 "
+ 'WHERE "timestamp" = ('
+ ' SELECT MIN("timestamp") FROM staged WHERE class = 1'
+ ")",
+ )
+
+ with pytest.raises(validator.ObservationsClassDomainError):
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
+
+
+def test_validator_rejects_observations_non_transient_class_zero(
+ tmp_path: Path,
+) -> None:
+ """Rule 6: events 3 and 4 must carry no `class >= 100` and no `class = 0`.
+
+ Mutating one row of an event-3 Instance to `class = 0` keeps the
+ catalog's row count and bucket totals (the row was class = 3 = steady,
+ becomes class = 0 = invalid for event 3 by rule 6) — wait, this
+ changes the per-row class so the steady bucket count from the
+ catalog would no longer match observations. To trip rule 6
+ cleanly, we instead append a new event-3 Instance file by writing
+ a single-row file with class = 0 and re-running transform so the
+ catalog reflects the new file's 1-row count; rule 6 then fires.
+ """
+ db_path = tmp_path / "petrobras_3w.duckdb"
+ out_dir = tmp_path / "parquet"
+ staging = _populated_staging(tmp_path)
+
+ # Drop a new event-3 file with a single class=0 row, then build the
+ # catalog from the modified staging. n_rows=1, buckets are all 0
+ # except… well, none of the buckets count class=0 with event_class=3
+ # because n_rows_normal excludes event 0/3/4-class instances? Let me
+ # re-read CONTEXT.md: `n_rows_normal` is "rows where class=0 AND
+ # event_class<>0". For event 3, class=0 row counts to normal=1.
+ # Buckets: warmup=0, normal=1, transient=NULL (event 3 has_transient=
+ # false), steady=0. Sum = 1 == n_rows. Catalog accounting OK.
+ extra_path = staging / "dataset" / "3" / "WELL-00099_20990101000000.parquet"
+ con = duckdb.connect()
+ try:
+ con.execute(
+ "CREATE TEMP TABLE staged ("
+ ' "timestamp" TIMESTAMP,'
+ ' "class" INTEGER,'
+ ' "state" INTEGER,'
+ ' "P-PDG" DOUBLE'
+ ")"
+ )
+ con.execute(
+ "INSERT INTO staged VALUES (TIMESTAMP '2099-01-01 00:00:00', 0, 0, 1.0e7)"
+ )
+ con.execute(f"COPY staged TO '{extra_path}' (FORMAT PARQUET)")
+ finally:
+ con.close()
+
+ dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
+ # Adding a new well_id (99) bumps real-Well rowcount to 41 and trips
+ # rule 7 before rule 6 — drop the freshly added wells row so rule 6
+ # fires first. Also drop the corresponding instances row's well_id
+ # FK so rule 3 stays happy. Easiest: also drop the new instance from
+ # the wells master only; the instances table still references
+ # well_id 99 which would trip rule 3 (wells FK). So delete the new
+ # well_id 99 entirely from instances AND wells, leaving the staged
+ # file in place so the observations view still sees it as an
+ # orphan… that trips rule 2 first.
+ #
+ # Cleanest path: keep all the existing fixtures untouched and add
+ # the new well_id to wells too, so rule 7's count becomes 41. To
+ # keep rule 7 at 40, drop one padding well that has no instances
+ # in `instances` table beyond its single-row event-0 fixture, then
+ # delete that instance and its staged file.
+ with duckdb.connect(str(db_path)) as con:
+ # Remove well 42 (a single-instance padding fixture). Also
+ # remove its instance from the catalog and its staged file from
+ # disk, so the catalog stays consistent.
+ con.execute("DELETE FROM wells WHERE well_id = 42")
+ con.execute("DELETE FROM instances WHERE well_id = 42")
+ (staging / "dataset" / "0" / "WELL-00042_20120101000000.parquet").unlink()
+
+ with pytest.raises(validator.ObservationsNonTransientClassError):
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ )
From 8d3b33c5e537f82caea1a6245fc3daafcf988df5 Mon Sep 17 00:00:00 2001
From: oskrgab
Date: Sun, 17 May 2026 20:41:52 +0000
Subject: [PATCH 07/11] Petrobras 3W parity test suite vs upstream pinned tag
(#23)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Adds the nine-check parity suite from PRD #18, wired in as the
post-write correctness gate. Where the existing pre-write `validator`
asserts structural invariants of the intermediate DB, the new
`parity.check(staging_dir, output_dir)` proves the published bytes
round-trip the upstream bytes 1:1 — class, state, timestamp, and
every sensor column are preserved verbatim, so any divergence is a
writer bug rather than legitimate transformation.
Checks (PRD #18 order, all hard-fail on divergence):
1. Per-event-class row count — upstream / catalog / pub.
2. Per-instance row count — upstream / catalog / pub.
3. Global `class` distribution — upstream vs pub.
4. Global `state` distribution — upstream vs pub.
5. Per-sensor SUM/AVG/MIN/MAX/COUNT/NULL — upstream vs pub.
6. Per-sensor aggregates grouped by `event_class` — same metrics.
7. Per-instance (start_ts, end_ts) — upstream / catalog / pub.
8. Distinct real-Well count — upstream / wells.parquet / pinned 40.
9. Per-event-class instance count — upstream vs catalog.
Key decisions:
- Parity runs *after* the parquets are written but *before* the
static-site tab is patched, so a divergence keeps the existing
published site intact. Implementation-wise, this means closing the
validator's read-only DuckDB connection, calling `parity.check`
with its own in-memory connection over the written parquets, then
re-opening a read-only connection for the schema-doc generator.
- Sensor columns are discovered at runtime from the published
`observations` schema (everything that is not an
identifier/label/time/hive column). Keeps the check agnostic to
upstream's exact 27-column set — a future column rename surfaces
as a parity match on the renamed column instead of a hard-coded
reference to a missing column name.
- NULL `class` and `state` values are legitimate on the warmup
prefix of real-Well anomaly Instances, so the global-distribution
checks (3 / 4) use `IS NOT DISTINCT FROM` for their join predicate.
Plain `USING (class)` would split the NULL bucket into two phantom
mismatches.
- Sensor columns are double-quoted via a `_quote` helper so hyphenated
identifiers (`P-PDG`, `ESTADO-SDV-GL`) round-trip safely.
- The bit-for-bit comparison policy (no epsilon) is from PRD #18:
every sensor float is copied unchanged, so identical SUM/AVG on
identical multisets is the contract. If full-corpus runs ever
surface non-associativity drift, we can switch to per-instance
row-level EXCEPT — for now the simpler aggregate compare suffices.
- The real-Well count (check 8) pins on `EXPECTED_REAL_WELL_COUNT`
imported from `validator` so the parity suite and the pre-write
validator stay in lockstep on the pinned 40.
- The orchestrator-level integration test patches `parity.check` to
raise, then asserts the website integrator never runs (the stub
README + index.html stay unchanged). This proves the wiring
without needing to artificially break the published bytes.
- The mutation-detection tests use a `_mutate_published_parquet`
helper that re-reads a written file with `hive_partitioning=false`,
applies a SQL mutation, and overwrites in place. Parallel pattern
to `_rewrite_staged_instance` from #22, but operating against the
published tree rather than the staged sources.
Files added:
- `scripts/export/petrobras_3w/parity.py` — module with `check()`
entry point and nine `Parity*Error` subclasses, one per PRD check.
Files modified:
- `scripts/export/petrobras_3w/orchestrator.py` — split the existing
single-connection block in two so `parity.check` can run between
the writes and the schema-doc / website-integrator steps; updated
the module docstring to describe the two correctness gates.
- `tests/petrobras_3w/test_smoke.py` — five parity tests: happy-path,
sensor-value mutation, partition-internal per-instance drift, row
drop (per-event-class), class-label flip, and orchestrator-aborts-
on-parity-failure.
Notes for next iteration:
- Issue #24 (final docs + website polish) is now unblocked.
- The full-corpus parity run is part of the production publish path
via the orchestrator hook; the smoke suite runs the 42-instance
subset version under 20 s on the fixture, well within the 30 s cap.
- Pre-existing test failure in `scripts/export/test_integration.py`
(Volve-era, missing `parquet/wells.parquet`) is unchanged.
Closes #23
Co-Authored-By: Claude Opus 4.7
---
scripts/export/petrobras_3w/orchestrator.py | 21 +-
scripts/export/petrobras_3w/parity.py | 554 ++++++++++++++++++++
tests/petrobras_3w/test_smoke.py | 179 ++++++-
3 files changed, 750 insertions(+), 4 deletions(-)
create mode 100644 scripts/export/petrobras_3w/parity.py
diff --git a/scripts/export/petrobras_3w/orchestrator.py b/scripts/export/petrobras_3w/orchestrator.py
index 01016a1..e41ebf1 100644
--- a/scripts/export/petrobras_3w/orchestrator.py
+++ b/scripts/export/petrobras_3w/orchestrator.py
@@ -2,8 +2,15 @@
Validates the intermediate DB unconditionally (and emits the pinned
upstream identity to the validation log), then writes the published
-parquets, the schema documentation, and idempotently surfaces the
+parquets, runs the post-write parity suite against the upstream staged
+tree, writes the schema documentation, and idempotently surfaces the
dataset on the static site.
+
+Two correctness gates run in sequence: the pre-write ``validator`` covers
+structural invariants of the intermediate DB; the post-write ``parity``
+suite proves the written bytes round-trip the upstream bytes 1:1. A
+parity divergence aborts publish before the static-site tab is patched,
+so a broken export never becomes visible to consumers.
"""
from __future__ import annotations
@@ -13,6 +20,7 @@
import duckdb
from scripts.export.petrobras_3w import (
+ parity,
parquet_writer,
schema_doc_generator,
validator,
@@ -39,12 +47,15 @@ def run(
reads one parquet per Instance from `/dataset/N/` so
the catalog plus the source files together are sufficient to emit
`observations/event_class=N/.parquet` without holding
- the full corpus in RAM.
+ the full corpus in RAM. The same tree is the upstream source-of-
+ truth for the post-write ``parity`` suite.
``website_root`` is the repo root containing `README.md` and
`parquet/index.html`. When provided, the static site is patched in
place to surface the Petrobras 3W dataset; when None, the website
- integration step is skipped.
+ integration step is skipped. Both site-patching paths run only after
+ ``parity.check`` returns, so a divergence keeps the existing site
+ intact.
"""
db_path = Path(db_path)
output_dir = Path(output_dir)
@@ -56,6 +67,10 @@ def run(
parquet_writer.write_instances(con, output_dir)
parquet_writer.write_wells(con, output_dir)
parquet_writer.write_observations(con, output_dir, staging_dir)
+
+ parity.check(staging_dir, output_dir)
+
+ with duckdb.connect(str(db_path), read_only=True) as con:
schema_doc_generator.generate(con, output_dir, dataset_ini)
if website_root is not None:
diff --git a/scripts/export/petrobras_3w/parity.py b/scripts/export/petrobras_3w/parity.py
new file mode 100644
index 0000000..4951363
--- /dev/null
+++ b/scripts/export/petrobras_3w/parity.py
@@ -0,0 +1,554 @@
+"""Parity check suite for the Petrobras 3W publish.
+
+Runs the nine queries described in PRD #18 / issue #23 against three
+sources of truth — the staged upstream parquet tree pinned at git tag
+``v.1.70.0`` / dataset version ``2.0.0``, the published catalog
+(``instances.parquet``, ``wells.parquet``), and the published
+``observations/`` hive tree. Any divergence aborts publish.
+
+Where the validator (``validator.py``) enforces structural invariants
+of the *intermediate* DuckDB DB before writes, parity is the *post-write*
+correctness gate that proves the published bytes round-trip the upstream
+bytes 1:1. The bit-for-bit comparison policy (no epsilon tolerance) is
+set in PRD #18: the 27 sensor floats + ``class`` + ``state`` +
+``timestamp`` are preserved verbatim, so any discrepancy is a writer
+bug we want to catch.
+
+Each check raises a specific subclass of ``ParityError`` so the orchestrator
+can produce an informative abort message and the test suite can target
+single checks. The checks run in PRD-#18 order:
+
+1. Per-event-class total row count — upstream vs catalog vs published.
+2. Per-instance row count — upstream vs catalog vs published.
+3. Global ``class`` distribution — upstream vs published.
+4. Global ``state`` distribution — upstream vs published.
+5. Per-sensor global aggregates (SUM/AVG/MIN/MAX/COUNT/NULL) — upstream
+ vs published, for every non-identifier body column.
+6. Per-sensor aggregates grouped by ``event_class`` — same metrics.
+7. Per-instance ``(start_ts, end_ts)`` — upstream MIN/MAX vs catalog
+ ``start_ts``/``end_ts`` vs published MIN/MAX.
+8. Distinct real-Well count — upstream count vs ``wells.parquet``
+ rowcount vs the pinned ``EXPECTED_REAL_WELL_COUNT``.
+9. Per-event-class instance count — upstream distinct-instance count
+ vs catalog row count.
+
+Sensor columns are discovered at runtime from the published
+``observations`` schema (the upstream side carries the same body, minus
+the three RLE-encoded constants the writer adds). That keeps the check
+agnostic to upstream's exact 27-column schema — a future column add or
+rename surfaces as a parity match on the renamed column, not a hard-
+coded reference to a missing column name.
+"""
+
+from __future__ import annotations
+
+from pathlib import Path
+
+import duckdb
+
+from scripts.export.petrobras_3w.validator import EXPECTED_REAL_WELL_COUNT
+
+# Columns present in every published Observations file body but not part
+# of the upstream-sensor schema. The first three are upstream's own
+# label/operating-status/time columns; the next three are the RLE-encoded
+# constants the writer derives from the catalog. ``event_class`` is the
+# hive-partition column on the published side (not stored in the file
+# body) and the regex-derived column on the upstream side, so it is also
+# excluded from the "sensor" list.
+_NON_SENSOR_COLUMNS: frozenset[str] = frozenset(
+ {
+ "class",
+ "state",
+ "timestamp",
+ "instance_id",
+ "well_id",
+ "well_kind",
+ "event_class",
+ }
+)
+
+
+class ParityError(Exception):
+ """Base class for parity divergences between upstream and published."""
+
+
+class ParityRowCountPerEventClassError(ParityError):
+ """Check 1: per-event-class row count diverges (upstream / catalog / pub)."""
+
+
+class ParityRowCountPerInstanceError(ParityError):
+ """Check 2: per-instance row count diverges (upstream / catalog / pub)."""
+
+
+class ParityClassDistributionError(ParityError):
+ """Check 3: global `class`-value distribution diverges (upstream vs pub)."""
+
+
+class ParityStateDistributionError(ParityError):
+ """Check 4: global `state`-value distribution diverges (upstream vs pub)."""
+
+
+class ParitySensorAggregatesError(ParityError):
+ """Check 5: a per-sensor global aggregate diverges (upstream vs pub)."""
+
+
+class ParitySensorAggregatesByEventClassError(ParityError):
+ """Check 6: a per-sensor per-event-class aggregate diverges."""
+
+
+class ParityInstanceTimestampsError(ParityError):
+ """Check 7: per-instance (start_ts, end_ts) diverges (any of three sources)."""
+
+
+class ParityRealWellCountError(ParityError):
+ """Check 8: distinct real-Well count diverges from the pinned upstream."""
+
+
+class ParityInstanceCountPerEventClassError(ParityError):
+ """Check 9: per-event-class instance count diverges (upstream vs catalog)."""
+
+
+def check(staging_dir: Path, output_dir: Path) -> None:
+ """Run the full nine-check parity suite.
+
+ ``staging_dir`` is the staged upstream tree (a ``dataset/N/*.parquet``
+ layout). ``output_dir`` is the published-parquets directory written
+ by ``parquet_writer``. The function reads both via a fresh DuckDB
+ in-memory connection — it does not share state with the validator's
+ intermediate DB.
+ """
+ staging_dir = Path(staging_dir)
+ output_dir = Path(output_dir)
+
+ con = duckdb.connect()
+ try:
+ _register_views(con, staging_dir, output_dir)
+ sensors = _sensor_columns(con)
+ _check_row_count_per_event_class(con)
+ _check_row_count_per_instance(con)
+ _check_class_distribution(con)
+ _check_state_distribution(con)
+ _check_sensor_aggregates(con, sensors)
+ _check_sensor_aggregates_by_event_class(con, sensors)
+ _check_instance_timestamps(con)
+ _check_real_well_count(con)
+ _check_instance_count_per_event_class(con)
+ finally:
+ con.close()
+
+
+def _register_views(
+ con: duckdb.DuckDBPyConnection, staging_dir: Path, output_dir: Path
+) -> None:
+ """Wire the three sources into named views with consistent types.
+
+ Both views surface ``event_class`` as ``BIGINT``: the upstream view
+ derives it from the directory name (cast on the way out); the
+ published view reads it via ``hive_partitioning=true`` (DuckDB picks
+ BIGINT for the autodetected integer partition value). ``instance_id``
+ is regex-derived from the source filename on the upstream side and
+ materialised as a constant column on the published side — both are
+ VARCHAR.
+ """
+ staging_glob = str(staging_dir / "dataset" / "*" / "*.parquet")
+ pub_obs_glob = str(output_dir / "observations" / "event_class=*" / "*.parquet")
+ instances_path = output_dir / "instances.parquet"
+ wells_path = output_dir / "wells.parquet"
+
+ con.execute(
+ f"""
+ CREATE OR REPLACE VIEW upstream_obs AS
+ SELECT
+ * EXCLUDE (filename),
+ CAST(regexp_extract(filename, '/dataset/([0-9]+)/', 1)
+ AS BIGINT) AS event_class,
+ regexp_extract(filename, '/([^/]+)\\.parquet$', 1)
+ AS instance_id
+ FROM read_parquet(
+ '{staging_glob}',
+ filename = true,
+ union_by_name = true
+ )
+ """
+ )
+
+ con.execute(
+ f"""
+ CREATE OR REPLACE VIEW pub_obs AS
+ SELECT *
+ FROM read_parquet(
+ '{pub_obs_glob}',
+ hive_partitioning = true,
+ union_by_name = true
+ )
+ """
+ )
+
+ con.execute(
+ f"""
+ CREATE OR REPLACE VIEW pub_instances AS
+ SELECT * FROM read_parquet('{instances_path}')
+ """
+ )
+
+ con.execute(
+ f"""
+ CREATE OR REPLACE VIEW pub_wells AS
+ SELECT * FROM read_parquet('{wells_path}')
+ """
+ )
+
+
+def _sensor_columns(con: duckdb.DuckDBPyConnection) -> tuple[str, ...]:
+ """Discover sensor columns from the published `pub_obs` schema.
+
+ Everything that is not an identifier / label / time / hive column is
+ a sensor. Discovered at runtime so the check stays agnostic to the
+ exact 27-column upstream schema.
+ """
+ described = con.execute("DESCRIBE SELECT * FROM pub_obs").fetchall()
+ return tuple(row[0] for row in described if row[0] not in _NON_SENSOR_COLUMNS)
+
+
+def _check_row_count_per_event_class(con: duckdb.DuckDBPyConnection) -> None:
+ """Check 1: per-event-class row count matches across all three sources."""
+ rows = con.execute(
+ """
+ WITH
+ upstream AS (
+ SELECT event_class, COUNT(*) AS n
+ FROM upstream_obs GROUP BY event_class
+ ),
+ catalog AS (
+ SELECT event_class, SUM(n_rows) AS n
+ FROM pub_instances GROUP BY event_class
+ ),
+ pub AS (
+ SELECT event_class, COUNT(*) AS n
+ FROM pub_obs GROUP BY event_class
+ )
+ SELECT
+ COALESCE(u.event_class, c.event_class, p.event_class) AS event_class,
+ u.n AS upstream_n, c.n AS catalog_n, p.n AS pub_n
+ FROM upstream u
+ FULL OUTER JOIN catalog c USING (event_class)
+ FULL OUTER JOIN pub p USING (event_class)
+ WHERE u.n IS DISTINCT FROM c.n
+ OR c.n IS DISTINCT FROM p.n
+ ORDER BY event_class
+ """
+ ).fetchall()
+ if rows:
+ raise ParityRowCountPerEventClassError(
+ f"per-event-class row count diverges for {len(rows)} event class(es): "
+ f"{rows}"
+ )
+
+
+def _check_row_count_per_instance(con: duckdb.DuckDBPyConnection) -> None:
+ """Check 2: per-instance row count matches upstream / catalog / pub."""
+ rows = con.execute(
+ """
+ WITH
+ upstream AS (
+ SELECT instance_id, COUNT(*) AS n
+ FROM upstream_obs GROUP BY instance_id
+ ),
+ catalog AS (
+ SELECT instance_id, n_rows AS n FROM pub_instances
+ ),
+ pub AS (
+ SELECT instance_id, COUNT(*) AS n
+ FROM pub_obs GROUP BY instance_id
+ )
+ SELECT
+ COALESCE(u.instance_id, c.instance_id, p.instance_id) AS instance_id,
+ u.n AS upstream_n, c.n AS catalog_n, p.n AS pub_n
+ FROM upstream u
+ FULL OUTER JOIN catalog c USING (instance_id)
+ FULL OUTER JOIN pub p USING (instance_id)
+ WHERE u.n IS DISTINCT FROM c.n
+ OR c.n IS DISTINCT FROM p.n
+ ORDER BY instance_id
+ """
+ ).fetchall()
+ if rows:
+ raise ParityRowCountPerInstanceError(
+ f"per-instance row count diverges for {len(rows)} instance(s); "
+ f"first: {rows[0]}"
+ )
+
+
+def _check_class_distribution(con: duckdb.DuckDBPyConnection) -> None:
+ """Check 3: global ``class``-value distribution matches upstream exactly.
+
+ ``class`` is NULL on the warmup prefix of real-Well anomaly Instances,
+ and a NULL bucket on both sides is the expected normal — so the join
+ must treat NULL=NULL as a match. ``USING (class)`` would split it
+ into two phantom mismatches, so we use ``IS NOT DISTINCT FROM``.
+ """
+ rows = con.execute(
+ """
+ WITH
+ upstream AS (
+ SELECT class, COUNT(*) AS n FROM upstream_obs GROUP BY class
+ ),
+ pub AS (
+ SELECT class, COUNT(*) AS n FROM pub_obs GROUP BY class
+ )
+ SELECT COALESCE(u.class, p.class) AS class,
+ u.n AS upstream_n, p.n AS pub_n
+ FROM upstream u
+ FULL OUTER JOIN pub p ON u.class IS NOT DISTINCT FROM p.class
+ WHERE u.n IS DISTINCT FROM p.n
+ ORDER BY class
+ """
+ ).fetchall()
+ if rows:
+ raise ParityClassDistributionError(
+ f"`class` distribution diverges for {len(rows)} class value(s): {rows}"
+ )
+
+
+def _check_state_distribution(con: duckdb.DuckDBPyConnection) -> None:
+ """Check 4: global ``state``-value distribution matches upstream exactly.
+
+ NULL is a legitimate value here (same reasoning as the class check),
+ so ``IS NOT DISTINCT FROM`` is used for the join predicate.
+ """
+ rows = con.execute(
+ """
+ WITH
+ upstream AS (
+ SELECT state, COUNT(*) AS n FROM upstream_obs GROUP BY state
+ ),
+ pub AS (
+ SELECT state, COUNT(*) AS n FROM pub_obs GROUP BY state
+ )
+ SELECT COALESCE(u.state, p.state) AS state,
+ u.n AS upstream_n, p.n AS pub_n
+ FROM upstream u
+ FULL OUTER JOIN pub p ON u.state IS NOT DISTINCT FROM p.state
+ WHERE u.n IS DISTINCT FROM p.n
+ ORDER BY state
+ """
+ ).fetchall()
+ if rows:
+ raise ParityStateDistributionError(
+ f"`state` distribution diverges for {len(rows)} state value(s): {rows}"
+ )
+
+
+def _quote(identifier: str) -> str:
+ """Double-quote a SQL identifier, escaping embedded double quotes.
+
+ Sensor columns can contain hyphens (`P-PDG`, `ESTADO-SDV-GL`), so the
+ parity SQL must always quote them. Embedded ``"`` is doubled per the
+ SQL standard.
+ """
+ escaped = identifier.replace('"', '""')
+ return f'"{escaped}"'
+
+
+def _check_sensor_aggregates(
+ con: duckdb.DuckDBPyConnection, sensors: tuple[str, ...]
+) -> None:
+ """Check 5: per-sensor global SUM/AVG/MIN/MAX/COUNT/NULL match upstream."""
+ for col in sensors:
+ quoted = _quote(col)
+ result = con.execute(
+ f"""
+ WITH
+ u AS (
+ SELECT
+ SUM({quoted}) AS s, AVG({quoted}) AS a,
+ MIN({quoted}) AS mn, MAX({quoted}) AS mx,
+ COUNT({quoted}) AS c,
+ COUNT(*) - COUNT({quoted}) AS nl
+ FROM upstream_obs
+ ),
+ p AS (
+ SELECT
+ SUM({quoted}) AS s, AVG({quoted}) AS a,
+ MIN({quoted}) AS mn, MAX({quoted}) AS mx,
+ COUNT({quoted}) AS c,
+ COUNT(*) - COUNT({quoted}) AS nl
+ FROM pub_obs
+ )
+ SELECT u.s, u.a, u.mn, u.mx, u.c, u.nl,
+ p.s, p.a, p.mn, p.mx, p.c, p.nl
+ FROM u, p
+ WHERE u.s IS DISTINCT FROM p.s
+ OR u.a IS DISTINCT FROM p.a
+ OR u.mn IS DISTINCT FROM p.mn
+ OR u.mx IS DISTINCT FROM p.mx
+ OR u.c IS DISTINCT FROM p.c
+ OR u.nl IS DISTINCT FROM p.nl
+ """
+ ).fetchall()
+ if result:
+ upstream_agg = result[0][:6]
+ pub_agg = result[0][6:]
+ raise ParitySensorAggregatesError(
+ f"sensor `{col}` global aggregates diverge: "
+ f"upstream(s,a,mn,mx,c,nl)={upstream_agg}, "
+ f"published={pub_agg}"
+ )
+
+
+def _check_sensor_aggregates_by_event_class(
+ con: duckdb.DuckDBPyConnection, sensors: tuple[str, ...]
+) -> None:
+ """Check 6: per-sensor aggregates grouped by `event_class` match exactly."""
+ for col in sensors:
+ quoted = _quote(col)
+ rows = con.execute(
+ f"""
+ WITH
+ u AS (
+ SELECT
+ event_class,
+ SUM({quoted}) AS s, AVG({quoted}) AS a,
+ MIN({quoted}) AS mn, MAX({quoted}) AS mx,
+ COUNT({quoted}) AS c,
+ COUNT(*) - COUNT({quoted}) AS nl
+ FROM upstream_obs
+ GROUP BY event_class
+ ),
+ p AS (
+ SELECT
+ event_class,
+ SUM({quoted}) AS s, AVG({quoted}) AS a,
+ MIN({quoted}) AS mn, MAX({quoted}) AS mx,
+ COUNT({quoted}) AS c,
+ COUNT(*) - COUNT({quoted}) AS nl
+ FROM pub_obs
+ GROUP BY event_class
+ )
+ SELECT COALESCE(u.event_class, p.event_class) AS event_class,
+ u.s, u.a, u.mn, u.mx, u.c, u.nl,
+ p.s, p.a, p.mn, p.mx, p.c, p.nl
+ FROM u
+ FULL OUTER JOIN p USING (event_class)
+ WHERE u.s IS DISTINCT FROM p.s
+ OR u.a IS DISTINCT FROM p.a
+ OR u.mn IS DISTINCT FROM p.mn
+ OR u.mx IS DISTINCT FROM p.mx
+ OR u.c IS DISTINCT FROM p.c
+ OR u.nl IS DISTINCT FROM p.nl
+ ORDER BY event_class
+ """
+ ).fetchall()
+ if rows:
+ raise ParitySensorAggregatesByEventClassError(
+ f"sensor `{col}` per-event-class aggregates diverge for "
+ f"{len(rows)} event class(es); first: {rows[0]}"
+ )
+
+
+def _check_instance_timestamps(con: duckdb.DuckDBPyConnection) -> None:
+ """Check 7: per-instance MIN/MAX timestamp matches upstream / catalog / pub."""
+ rows = con.execute(
+ """
+ WITH
+ upstream AS (
+ SELECT instance_id,
+ MIN("timestamp") AS min_ts,
+ MAX("timestamp") AS max_ts
+ FROM upstream_obs
+ GROUP BY instance_id
+ ),
+ catalog AS (
+ SELECT instance_id, start_ts, end_ts FROM pub_instances
+ ),
+ pub AS (
+ SELECT instance_id,
+ MIN("timestamp") AS min_ts,
+ MAX("timestamp") AS max_ts
+ FROM pub_obs
+ GROUP BY instance_id
+ )
+ SELECT COALESCE(u.instance_id, c.instance_id, p.instance_id)
+ AS instance_id,
+ u.min_ts AS upstream_min, u.max_ts AS upstream_max,
+ c.start_ts AS catalog_start, c.end_ts AS catalog_end,
+ p.min_ts AS pub_min, p.max_ts AS pub_max
+ FROM upstream u
+ FULL OUTER JOIN catalog c USING (instance_id)
+ FULL OUTER JOIN pub p USING (instance_id)
+ WHERE u.min_ts IS DISTINCT FROM c.start_ts
+ OR u.max_ts IS DISTINCT FROM c.end_ts
+ OR u.min_ts IS DISTINCT FROM p.min_ts
+ OR u.max_ts IS DISTINCT FROM p.max_ts
+ ORDER BY instance_id
+ """
+ ).fetchall()
+ if rows:
+ raise ParityInstanceTimestampsError(
+ f"per-instance timestamps diverge for {len(rows)} instance(s); "
+ f"first: {rows[0]}"
+ )
+
+
+def _check_real_well_count(con: duckdb.DuckDBPyConnection) -> None:
+ """Check 8: distinct real-Well count matches the pinned upstream count.
+
+ Compares three counts: the count of distinct ``WELL-NNNNN`` prefixes
+ in the staged upstream tree, the row count of ``wells.parquet``, and
+ the pinned ``EXPECTED_REAL_WELL_COUNT`` (40 at git tag ``v.1.70.0``
+ / dataset version ``2.0.0``). Any disagreement is a hard fail.
+ """
+ upstream_count = con.execute(
+ """
+ SELECT COUNT(DISTINCT
+ CAST(regexp_extract(instance_id, '^WELL-0*([0-9]+)_', 1) AS INTEGER)
+ )
+ FROM upstream_obs
+ WHERE starts_with(instance_id, 'WELL-')
+ """
+ ).fetchone()[0]
+ pub_count = con.execute("SELECT COUNT(*) FROM pub_wells").fetchone()[0]
+ if (
+ upstream_count != EXPECTED_REAL_WELL_COUNT
+ or pub_count != EXPECTED_REAL_WELL_COUNT
+ ):
+ raise ParityRealWellCountError(
+ f"real-Well count diverges: upstream={upstream_count}, "
+ f"wells.parquet={pub_count}, pinned={EXPECTED_REAL_WELL_COUNT}"
+ )
+
+
+def _check_instance_count_per_event_class(con: duckdb.DuckDBPyConnection) -> None:
+ """Check 9: per-event-class instance count matches upstream's folder file count.
+
+ Each upstream file is one Instance, so ``COUNT(DISTINCT instance_id)``
+ grouped by ``event_class`` over ``upstream_obs`` is equivalent to the
+ file count of ``upstream/dataset/N/`` for each ``N`` (and earlier
+ checks would already have caught a missing file).
+ """
+ rows = con.execute(
+ """
+ WITH
+ upstream AS (
+ SELECT event_class, COUNT(DISTINCT instance_id) AS n
+ FROM upstream_obs
+ GROUP BY event_class
+ ),
+ catalog AS (
+ SELECT event_class, COUNT(*) AS n
+ FROM pub_instances
+ GROUP BY event_class
+ )
+ SELECT COALESCE(u.event_class, c.event_class) AS event_class,
+ u.n AS upstream_n, c.n AS catalog_n
+ FROM upstream u
+ FULL OUTER JOIN catalog c USING (event_class)
+ WHERE u.n IS DISTINCT FROM c.n
+ ORDER BY event_class
+ """
+ ).fetchall()
+ if rows:
+ raise ParityInstanceCountPerEventClassError(
+ f"per-event-class instance count diverges for {len(rows)} "
+ f"event class(es): {rows}"
+ )
diff --git a/tests/petrobras_3w/test_smoke.py b/tests/petrobras_3w/test_smoke.py
index b34cc04..c09b1ff 100644
--- a/tests/petrobras_3w/test_smoke.py
+++ b/tests/petrobras_3w/test_smoke.py
@@ -30,7 +30,7 @@
import pytest
from scripts.export.petrobras_3w import orchestrator as export_orch
-from scripts.export.petrobras_3w import validator
+from scripts.export.petrobras_3w import parity, validator
from scripts.transform.petrobras_3w import orchestrator as transform_orch
from scripts.transform.petrobras_3w.constants import (
PIN_DATASET_VERSION,
@@ -1270,3 +1270,180 @@ def test_validator_rejects_observations_non_transient_class_zero(
dataset_ini=dataset_ini,
staging_dir=staging,
)
+
+
+# ---------------------------------------------------------------------------
+# Parity check suite vs upstream pinned tag (issue #23)
+# ---------------------------------------------------------------------------
+
+
+def _mutate_published_parquet(target: Path, mutate_sql: str) -> None:
+ """Re-read a published Observations parquet, apply ``mutate_sql``
+ against the ``staged`` temp table, and overwrite the file in place.
+
+ Used to surgically break a single value or row in the published
+ bytes *after* the export pipeline has run, so the parity suite
+ operates against a real divergence rather than a synthetic input.
+ Hive partitioning is disabled on the read so the body columns line
+ up 1:1 with the source on disk.
+ """
+ con = duckdb.connect()
+ try:
+ con.execute(
+ f"CREATE TEMP TABLE staged AS SELECT * FROM "
+ f"read_parquet('{target}', hive_partitioning=false)"
+ )
+ con.execute(mutate_sql)
+ con.execute(f"COPY staged TO '{target}' (FORMAT PARQUET)")
+ finally:
+ con.close()
+
+
+def test_parity_passes_on_happy_path(tmp_path: Path) -> None:
+ """All nine parity checks succeed against the fixture upstream tree.
+
+ Asserts that the smoke-fixture run is end-to-end-consistent across
+ upstream / catalog / published Observations. The pipeline's own
+ success in `_run_pipeline` is *implicitly* this check, but exercising
+ `parity.check` directly here documents the public entry point.
+ """
+ _, out_dir, staging = _run_pipeline(tmp_path)
+ # Re-running parity outside the orchestrator must also succeed.
+ parity.check(staging, out_dir)
+
+
+def test_parity_detects_sensor_value_mutation(tmp_path: Path) -> None:
+ """A single-value edit to a sensor column in a published parquet trips
+ at least one parity check.
+
+ Bumps the `P-PDG` value in one row of one published Observations
+ file. The catalog and upstream stay untouched, so the catalog-side
+ structural checks (row count, FK, timestamps) all match — only the
+ sensor-aggregate checks see the divergence.
+ """
+ _, out_dir, staging = _run_pipeline(tmp_path)
+
+ target = (
+ out_dir / "observations" / "event_class=1" / "WELL-00002_20120102000000.parquet"
+ )
+ _mutate_published_parquet(
+ target,
+ # Bump the first row's `P-PDG` by 1.0 — keeps row count, NULL count,
+ # and class distribution intact; only the SUM/AVG/MIN/MAX trip.
+ 'UPDATE staged SET "P-PDG" = "P-PDG" + 1.0 '
+ 'WHERE "timestamp" = (SELECT MIN("timestamp") FROM staged)',
+ )
+
+ with pytest.raises(parity.ParitySensorAggregatesError):
+ parity.check(staging, out_dir)
+
+
+def test_parity_detects_row_drop_in_published(tmp_path: Path) -> None:
+ """Dropping a published Observations row trips a row-count parity check
+ (the catalog still says n_rows=12 for that Instance; the published file
+ now has 11 rows). The per-event-class check fires first in the suite
+ order, but the acceptance criterion only requires *some* check to fire.
+ """
+ _, out_dir, staging = _run_pipeline(tmp_path)
+
+ target = (
+ out_dir / "observations" / "event_class=1" / "WELL-00002_20120102000000.parquet"
+ )
+ _mutate_published_parquet(
+ target,
+ 'DELETE FROM staged WHERE "timestamp" = (SELECT MAX("timestamp") FROM staged)',
+ )
+
+ with pytest.raises(parity.ParityRowCountPerEventClassError):
+ parity.check(staging, out_dir)
+
+
+def test_parity_detects_per_instance_only_drift(tmp_path: Path) -> None:
+ """Reordering rows across two published files in the same event class
+ keeps the per-event-class row count intact (check 1 passes) but trips
+ the per-instance row count check (check 2). This isolates check 2
+ behind its dedicated exception class so the suite catches partition-
+ internal drift, not just cross-partition mismatches.
+ """
+ _, out_dir, staging = _run_pipeline(tmp_path)
+
+ # Both targets are event_class=0 padding fixtures (1 row each). We
+ # drop one row from `target_lose` and synthesise an extra (duplicate-
+ # timestamp-shifted) row in `target_gain` so the partition's total
+ # row count is unchanged.
+ base = out_dir / "observations" / "event_class=0"
+ target_lose = base / "WELL-00004_20120101000000.parquet"
+ target_gain = base / "WELL-00005_20120101000000.parquet"
+ _mutate_published_parquet(target_lose, "DELETE FROM staged")
+ _mutate_published_parquet(
+ target_gain,
+ # Append a synthetic second row 1 second later so the partition-
+ # level total stays the same but this Instance now has 2 rows
+ # instead of 1.
+ "INSERT INTO staged SELECT "
+ '"timestamp" + INTERVAL \'1 second\', class, state, "P-PDG", '
+ "instance_id, well_id, well_kind FROM staged",
+ )
+
+ with pytest.raises(parity.ParityRowCountPerInstanceError):
+ parity.check(staging, out_dir)
+
+
+def test_parity_detects_class_label_flip(tmp_path: Path) -> None:
+ """Flipping one `class` value in a published file trips the class
+ distribution check (the per-row class is preserved verbatim, so any
+ edit is a parity break).
+ """
+ _, out_dir, staging = _run_pipeline(tmp_path)
+
+ target = (
+ out_dir / "observations" / "event_class=1" / "WELL-00002_20120102000000.parquet"
+ )
+ _mutate_published_parquet(
+ target,
+ # Replace one STEADY row's class (1) with a value that already
+ # exists in the published distribution (101), so the global
+ # multiset changes its histogram without introducing a new value.
+ 'UPDATE staged SET "class" = 101 '
+ 'WHERE "timestamp" = (SELECT MIN("timestamp") FROM staged WHERE "class" = 1)',
+ )
+
+ with pytest.raises(parity.ParityClassDistributionError):
+ parity.check(staging, out_dir)
+
+
+def test_orchestrator_aborts_publish_on_parity_failure(
+ tmp_path: Path, monkeypatch: pytest.MonkeyPatch
+) -> None:
+ """A parity divergence raised from `parity.check` propagates out of
+ the orchestrator and skips the website-integration step.
+
+ Patches `parity.check` (as imported into the orchestrator module) so
+ it always raises. The website root is provided, so we can assert
+ the integrator did NOT run by checking the stub README is unchanged.
+ """
+ db_path = tmp_path / "petrobras_3w.duckdb"
+ out_dir = tmp_path / "parquet"
+ staging = _populated_staging(tmp_path)
+ site_root = _site_root(tmp_path)
+ pristine_readme = (site_root / "README.md").read_text()
+
+ def boom(*_args, **_kwargs) -> None:
+ raise parity.ParitySensorAggregatesError("synthetic divergence")
+
+ monkeypatch.setattr(export_orch.parity, "check", boom)
+
+ dataset_ini = transform_orch.run(db_path=db_path, staging_dir=staging)
+ with pytest.raises(parity.ParitySensorAggregatesError):
+ export_orch.run(
+ db_path=db_path,
+ output_dir=out_dir,
+ dataset_ini=dataset_ini,
+ staging_dir=staging,
+ website_root=site_root,
+ )
+
+ # The static site is untouched on a parity abort: the publish never
+ # made it past `parity.check`.
+ assert (site_root / "README.md").read_text() == pristine_readme
+ assert "petrobras_3w" not in (site_root / "parquet" / "index.html").read_text()
From bda0130f732e58794b334d85729ec64339d101ed Mon Sep 17 00:00:00 2001
From: oskrgab
Date: Sun, 17 May 2026 20:57:49 +0000
Subject: [PATCH 08/11] Petrobras 3W final docs + website polish (#24)
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Refreshes the published documentation surface for the Petrobras 3W
dataset to the final four-table state. The committed `parquet/petrobras_3w/`
docs and the static-site entry were last touched at the #19 skeleton
slice and only described `event_types`; this PR brings them in lockstep
with the writers landed by #20 (`instances`), #21 (`wells`), #22
(`observations`), and #23 (parity).
Generator improvements:
- `schema.md` sensor-glossary intro drops the stale "once issue #22
lands" forward-reference — those bytes have shipped, so the prose now
describes the observations columns as present.
- The observations table's per-column docs now inline upstream's sensor
descriptions from `dataset.ini` (e.g. `P-PDG` → "Downhole pressure at
the PDG (permanent downhole gauge) [Pa]") instead of leaving the
Description cell blank with the glossary as the only reference.
- README adds the per-Well cross-validation split query (acceptance
criterion c from issue #24) — leave-one-Well-out via `well_id %
n_folds`, with a follow-up note on handling simulated/drawn
Instances. The four canonical examples (a load-by-event-class, b
fetch-by-URL, c per-Well CV split, d corpus-balance from catalog
only) are now all surfaced.
Test fixture upgrade:
- `tests/petrobras_3w/conftest.py` reads upstream's
`PARQUET_FILE_PROPERTIES` and writes all 27 sensor columns into the
per-Instance fixtures (placeholder values for the 26 non-P-PDG
columns; `P-PDG` keeps its row-varying values so parity sensor
aggregates still differentiate). This makes the fixture-derived
schema docs structurally match production, which is what lets us
commit them. The lone hyphenated `P-PDG` column was enough to
exercise column-name fidelity end-to-end at #22's level, but it
understated the column inventory in the reflected schema.
- `test_parity_detects_per_instance_only_drift`'s INSERT now uses
`INSERT BY NAME SELECT *` so it doesn't need to enumerate every
sensor column by hand.
Regenerated committed artifacts:
- `parquet/petrobras_3w/schema.md`, `schema.json`, `schema.sql`,
`README.md` regenerated against the upgraded fixture — `schema.sql`
now declares all 27 sensor columns with hyphen-quoted identifiers,
`schema.json` carries the full column / FK / PK metadata for all
four tables, `schema.md` and `README.md` carry every per-table
description and every canonical query example.
- Root `README.md` and `parquet/index.html` regenerated via the
website integrator — the tab count was stuck at "1 file" since
#19 and the README blurb still claimed the catalog and Observations
"ship in follow-up issues". Both now reflect the realised
four-table state.
Key decisions:
- Committed schema docs are produced from the FIXTURE pipeline, not
from a full upstream publish (3+ GB unstageable). The docs describe
shape, not content — bumping fixture column inventory to match
upstream's 27-column inventory makes the fixture-derived docs
production-correct. The actual data parquets (`wells`, `instances`,
`observations/`) still come from the upstream publish at deploy
time; only `event_types.parquet` is byte-stable across the two
paths because it is dataset.ini-driven.
- Per-Well CV split uses `well_id %% n_folds` as the fold key. Stable
across refreshes and zero-state (no separate fold-assignment file
needed); consumers picking another assignment can replace the
`%% 5` clause without touching the rest of the query.
- `dataset.ini` parsing is reused in the test conftest via
`parse_dataset_ini` rather than hard-coding the 27 column names —
any future upstream column rename surfaces at fixture build time as
a parse-time mismatch, the same fail-loud property the production
pipeline has.
Files modified:
- `scripts/export/petrobras_3w/schema_doc_generator.py` — sensor-
description inline lookup + glossary-intro prose update + new
CV-split README section.
- `tests/petrobras_3w/conftest.py` — 27-sensor fixture columns from
upstream `dataset.ini`.
- `tests/petrobras_3w/test_smoke.py` — assert no stale "#22 lands"
prose; assert sensor descriptions are inlined in observations table;
assert the four canonical queries (a/b/c/d) are present; switch the
per-instance-drift parity test's INSERT to `BY NAME SELECT *`.
- `parquet/petrobras_3w/{schema.md,schema.json,schema.sql,README.md,LICENSE-3W-DATA.md}`
— regenerated.
- `README.md`, `parquet/index.html` — refreshed via the website
integrator.
Notes for next iteration:
- Pre-existing test failure in `scripts/export/test_integration.py`
(Volve-era, missing `parquet/wells.parquet`) is unchanged, same as
every #19..#23 noted.
- The deploy step that copies the full ~2,228-file Observations tree
to `dev-petrodb.ocortez.com` is out of scope for this PR — the
pipeline running in production is what materialises those bytes.
The committed schema docs accurately describe what the deployed
tree contains.
Closes #24
Co-Authored-By: Claude Opus 4.7
---
README.md | 27 +-
parquet/index.html | 67 ++-
parquet/petrobras_3w/README.md | 105 ++++-
parquet/petrobras_3w/schema.json | 444 +++++++++++++++++-
parquet/petrobras_3w/schema.md | 115 ++++-
parquet/petrobras_3w/schema.sql | 73 +++
.../petrobras_3w/schema_doc_generator.py | 41 +-
tests/petrobras_3w/conftest.py | 39 +-
tests/petrobras_3w/test_smoke.py | 34 +-
9 files changed, 886 insertions(+), 59 deletions(-)
diff --git a/README.md b/README.md
index 552984a..0932522 100644
--- a/README.md
+++ b/README.md
@@ -58,21 +58,28 @@ four-bucket rationale, and three more canonical query patterns live in
### Petrobras 3W Dataset
Labelled 1-Hz sensor-data windows from the Petrobras 3W dataset, sliced
into per-Instance Parquet files. Pinned at upstream git tag `v.1.70.0`
-(dataset version `2.0.0`). This initial release publishes
-the event-class lookup and documentation scaffolding; the Instance catalog,
-real-Well master, and Observations time-series ship in follow-up issues.
+(dataset version `2.0.0`). This release publishes the
+event-class lookup, the real-Well master, the full Instance catalog, and
+the per-Instance Observations time-series (hive-partitioned by event class).
-List every event class (NORMAL plus the nine anomaly categories) with their
-TRANSIENT-arc semantics:
+Measure the labelled-data balance across the corpus from the catalog alone
+(no Observations scan needed):
```python
import duckdb
-result = duckdb.sql("""
- SELECT event_class, name, description,
- has_transient, transient_code
- FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/event_types.parquet'
- ORDER BY event_class
+base = 'https://dev-petrodb.ocortez.com/petrobras_3w'
+result = duckdb.sql(f"""
+ SELECT
+ et.event_class,
+ et.description,
+ COUNT(*) AS n_instances,
+ SUM(i.n_rows) AS n_observations
+ FROM '{base}/instances.parquet' i
+ JOIN '{base}/event_types.parquet' et
+ ON et.event_class = i.event_class
+ GROUP BY et.event_class, et.description
+ ORDER BY et.event_class
""").df()
```
diff --git a/parquet/index.html b/parquet/index.html
index 85a6a7d..fef3c81 100644
--- a/parquet/index.html
+++ b/parquet/index.html
@@ -797,7 +797,7 @@
PETRODATA REPOSITORY
@@ -2016,16 +2016,28 @@
Download Petrobras 3W Files
Labelled 1-Hz sensor-data windows from the Petrobras 3W dataset.
Pinned at upstream git tag v.1.70.0
- (dataset version 2.0.0). This initial
- release publishes the event-class lookup and documentation
- scaffolding; the Instance catalog, real-Well master, and
- Observations time-series ship in follow-up issues.
+ (dataset version 2.0.0). This release
+ publishes the event-class lookup, the real-Well master, the full
+ Instance catalog, and the per-Instance Observations time-series
+ hive-partitioned by event_class.
One lookup table · pinned upstream identity logged on every publish
+
Lookup + Wells master + Instance catalog + per-Instance Observations (hive-partitioned by event_class) · pinned upstream identity logged on every publish
@@ -2078,7 +2090,8 @@
About This Dataset
Quick Start with DuckDB
- List every event class with its TRANSIENT-arc semantics:
+ Measure the labelled-data balance across the corpus from the
+ Instance catalog alone (no Observations scan needed):
@@ -2088,20 +2101,40 @@
Quick Start with DuckDB
import duckdb
-# List the 10 event classes and their TRANSIENT-arc semantics
-result = duckdb.sql("""
- SELECT event_class, name, description,
- has_transient, transient_code
- FROM 'petrobras_3w/event_types.parquet'
- ORDER BY event_class
+base = 'https://dev-petrodb.ocortez.com/petrobras_3w'
+result = duckdb.sql(f"""
+ SELECT
+ et.event_class,
+ et.description,
+ COUNT(*) AS n_instances,
+ SUM(i.n_rows) AS n_observations
+ FROM '{base}/instances.parquet' i
+ JOIN '{base}/event_types.parquet' et
+ ON et.event_class = i.event_class
+ GROUP BY et.event_class, et.description
+ ORDER BY et.event_class
""").df()
- More canonical patterns (per-event-class filter, joins
- against the Instance catalog, single-Instance fetch)
- will land alongside the catalog and Observations files
- in follow-up releases.
+ The per-Instance Observations files are accessible via the
+ hive-partitioned URL pattern
+ observations/event_class=N/<instance_id>.parquet.
+ Each file embeds instance_id, well_id,
+ and well_kind as constant columns, so corpus-wide
+ queries against a single event class do not need to join the
+ catalog:
+
+
+
+
+
+
+
-- All real-Well Hydrate-in-Production-Line observations
+SELECT instance_id, well_id, "timestamp", "P-PDG", "T-PDG", class
+FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/observations/event_class=8/*.parquet'
+WHERE well_kind = 'real';
+
diff --git a/parquet/petrobras_3w/README.md b/parquet/petrobras_3w/README.md
index ad4ec4e..ad225e9 100644
--- a/parquet/petrobras_3w/README.md
+++ b/parquet/petrobras_3w/README.md
@@ -1,6 +1,6 @@
# Petrobras 3W Dataset
-Labelled 1-Hz sensor-data windows for the Petrobras 3W dataset, republished as Parquet files. The full per-Instance time-series (`observations/`), the Instance catalog (`instances.parquet`), and the real-Well master (`wells.parquet`) ship in subsequent issues (#22, #20, #21); this initial release publishes only the event-class lookup (`event_types.parquet`) and the documentation scaffolding so consumers can preview the schema.
+Labelled 1-Hz sensor-data windows for the Petrobras 3W dataset, republished as Parquet files. This release publishes the event-class lookup (`event_types.parquet`), the full Instance catalog (`instances.parquet`), the real-Well master (`wells.parquet`), and the per-Instance Observations time-series (`observations/event_class=N/.parquet`).
## Upstream pin
@@ -10,11 +10,18 @@ Labelled 1-Hz sensor-data windows for the Petrobras 3W dataset, republished as P
Refreshes are event-driven on new upstream tags (see [ADR-0002](../../docs/adr/0002-petrobras-3w-pin-upstream-release-tag.md)). Both the git tag and the dataset version are emitted in the publish orchestrator's validation log.
-## Published files (this slice)
+## Published files
```
petrobras_3w/
├── event_types.parquet # 10 rows, one per upstream event class
+├── wells.parquet # 40 rows, one per real Well
+├── instances.parquet # one row per upstream Instance file
+├── observations/
+│ ├── event_class=0/.parquet # ~594 files
+│ ├── …
+│ ├── event_class=9/.parquet # ~207 files
+│ └── _files.json # manifest of every Observations file's relative path
├── schema.md
├── schema.json
├── schema.sql
@@ -48,6 +55,100 @@ WHERE event_class > 0
ORDER BY event_class;
```
+### Corpus balance from the Instance catalog
+
+The per-Instance `n_rows_*` counts let you measure the labelled data balance across the corpus without scanning the Observations time-series:
+
+```sql
+SELECT
+ et.event_class,
+ et.description,
+ COUNT(*) AS n_instances,
+ SUM(i.n_rows) AS n_observations,
+ SUM(i.n_rows_steady) AS n_observations_steady
+FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/instances.parquet' i
+JOIN 'https://dev-petrodb.ocortez.com/petrobras_3w/event_types.parquet' et
+ ON et.event_class = i.event_class
+GROUP BY et.event_class, et.description
+ORDER BY et.event_class;
+```
+
+### List Instances of a single event class (real wells only)
+
+```sql
+SELECT instance_id, well_id, start_ts, n_rows, source_url
+FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/instances.parquet'
+WHERE event_class = 8
+ AND well_kind = 'real'
+ORDER BY start_ts;
+```
+
+### Per-Well corpus footprint (join `wells` with `instances`)
+
+The `wells` master pre-aggregates each Well's Instance and Observation counts so coverage tables can be built without scanning Observations:
+
+```sql
+SELECT
+ w.well_id,
+ w.n_instances,
+ w.n_observations,
+ w.first_ts,
+ w.last_ts,
+ COUNT(DISTINCT i.event_class) AS distinct_event_classes
+FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/wells.parquet' w
+JOIN 'https://dev-petrodb.ocortez.com/petrobras_3w/instances.parquet' i USING (well_id)
+GROUP BY w.well_id, w.n_instances, w.n_observations,
+ w.first_ts, w.last_ts
+ORDER BY w.n_instances DESC;
+```
+
+### Per-Well cross-validation split (leave-one-Well-out)
+
+Training models on Petrobras 3W should split by `well_id`, not by Instance — Instances drawn from the same physical Well share operating conditions and would leak signal across a naive shuffle. Assign each real Well a stable fold index from the `well_id` modulo the desired fold count, then derive the test Instances for fold `k` directly from `instances.parquet`:
+
+```sql
+WITH folded AS (
+ SELECT well_id, well_id % 5 AS fold
+ FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/wells.parquet'
+)
+SELECT i.instance_id, i.event_class, i.n_rows, i.source_url
+FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/instances.parquet' i
+JOIN folded f USING (well_id)
+WHERE f.fold = 0 -- held-out test set; train on fold != 0
+ORDER BY i.event_class, i.instance_id;
+```
+
+The simulated and drawn Instances (`well_kind <> 'real'`, `well_id IS NULL`) are excluded from the join above and can be added to the training set independently — they have no physical Well to leak against.
+
+### Load all real-Well Observations of one event class
+
+The Observations tree is hive-partitioned by `event_class`, so a wildcard against one partition is a pruned scan — DuckDB only touches files under that path. Each file carries `instance_id`, `well_id`, `well_kind` as constant columns, so consumers can filter by Well or provenance without joining the catalog:
+
+```sql
+SELECT instance_id, well_id, "timestamp", "P-PDG", "T-PDG", class
+FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/observations/event_class=8/*.parquet'
+WHERE well_kind = 'real'
+ORDER BY instance_id, "timestamp";
+```
+
+### Fetch one specific Instance by `source_url`
+
+Each row of `instances.parquet` carries the published URL of its Observations file. Round-trip the catalog and the time-series in two queries:
+
+```sql
+-- 1. find the URL
+SELECT source_url
+FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/instances.parquet'
+WHERE instance_id = 'WELL-00019_20120601165020';
+
+-- 2. read the Observations
+SELECT * FROM 'https://dev-petrodb.ocortez.com/petrobras_3w/observations/event_class=8/WELL-00019_20120601165020.parquet';
+```
+
+### Enumerate every Observations file (`_files.json`)
+
+A JSON-array manifest at `observations/_files.json` lists every published file's path relative to the partition root, in catalog order. Useful for consumers that prefer enumeration over wildcard scans (e.g. ML training loops that iterate Instances one at a time).
+
## License
Upstream data is released under [Creative Commons Attribution 4.0](https://creativecommons.org/licenses/by/4.0/). See `LICENSE-3W-DATA.md` in this directory for the attribution text.
diff --git a/parquet/petrobras_3w/schema.json b/parquet/petrobras_3w/schema.json
index 86ee398..05a897c 100644
--- a/parquet/petrobras_3w/schema.json
+++ b/parquet/petrobras_3w/schema.json
@@ -14,43 +14,475 @@
"name": "event_class",
"type": "INTEGER",
"not_null": true,
- "primary_key": true
+ "primary_key": true,
+ "hive_partition": false
},
{
"name": "name",
"type": "VARCHAR",
"not_null": false,
- "primary_key": false
+ "primary_key": false,
+ "hive_partition": false
},
{
"name": "description",
"type": "VARCHAR",
"not_null": false,
- "primary_key": false
+ "primary_key": false,
+ "hive_partition": false
},
{
"name": "has_transient",
"type": "BOOLEAN",
"not_null": false,
- "primary_key": false
+ "primary_key": false,
+ "hive_partition": false
},
{
"name": "transient_code",
"type": "INTEGER",
"not_null": false,
- "primary_key": false
+ "primary_key": false,
+ "hive_partition": false
},
{
"name": "has_normal_prefix",
"type": "BOOLEAN",
"not_null": false,
- "primary_key": false
+ "primary_key": false,
+ "hive_partition": false
}
],
"primary_key": [
"event_class"
],
"foreign_keys": []
+ },
+ "wells": {
+ "description": "Real-Well master, one row per distinct `well_id` derived from Instances with `well_kind = 'real'` (40 rows at the current upstream pin). Upstream anonymises every physical-well attribute (no basin, field, depth, or location), so the master is an identity-plus-statistics table: count of Instances, total 1-Hz Observations, and the time span across which the Well appears in the corpus. Simulated and drawn Instances have NULL `well_id` and contribute nothing here.",
+ "columns": [
+ {
+ "name": "well_id",
+ "type": "INTEGER",
+ "not_null": true,
+ "primary_key": true,
+ "hive_partition": false
+ },
+ {
+ "name": "n_instances",
+ "type": "BIGINT",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "first_ts",
+ "type": "TIMESTAMP",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "last_ts",
+ "type": "TIMESTAMP",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "n_observations",
+ "type": "BIGINT",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ }
+ ],
+ "primary_key": [
+ "well_id"
+ ],
+ "foreign_keys": []
+ },
+ "instances": {
+ "description": "One row per upstream Instance file (~2,228 rows). Identifies the Instance (`instance_id`), its provenance (`well_kind`, `well_id`, `source_file`), the operational regime it is framed around (`event_class`), and pre-aggregated per-Instance statistics (`start_ts`, `end_ts`, `duration_s`, `n_rows`, plus four `n_rows_*` counts that partition `n_rows` by `class` value). Corpus-wide balance and labelled-mass queries can run purely against this catalog without scanning the Observations time-series. `source_url` points at the published Observations file for the Instance (the URL pattern is fixed by ADR-0001).",
+ "columns": [
+ {
+ "name": "instance_id",
+ "type": "VARCHAR",
+ "not_null": true,
+ "primary_key": true,
+ "hive_partition": false
+ },
+ {
+ "name": "well_kind",
+ "type": "VARCHAR",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "well_id",
+ "type": "INTEGER",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "event_class",
+ "type": "INTEGER",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "start_ts",
+ "type": "TIMESTAMP",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "end_ts",
+ "type": "TIMESTAMP",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "duration_s",
+ "type": "BIGINT",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "n_rows",
+ "type": "BIGINT",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "n_rows_warmup_null",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "n_rows_normal",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "n_rows_transient",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "n_rows_steady",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "source_file",
+ "type": "VARCHAR",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "source_url",
+ "type": "VARCHAR",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ }
+ ],
+ "primary_key": [
+ "instance_id"
+ ],
+ "foreign_keys": [
+ {
+ "column": "event_class",
+ "references_table": "event_types",
+ "references_column": "event_class"
+ },
+ {
+ "column": "well_id",
+ "references_table": "wells",
+ "references_column": "well_id"
+ }
+ ]
+ },
+ "observations": {
+ "description": "Per-Instance 1-Hz sensor time-series. Hive-partitioned by `event_class` into `observations/event_class=N/.parquet` — one file per Instance, ~2,228 files in total. Each file preserves the upstream sensor columns verbatim (including hyphens: `P-PDG`, `ABER-CKGL`, `ESTADO-SDV-GL`, …), plus `class`, `state`, and `timestamp`. Three constant columns identify provenance per row: `instance_id`, `well_id`, `well_kind` (RLE-encoded, negligible storage). `event_class` is provided by the hive partition and is NOT stored in the file body. A `_files.json` manifest at the partition root lists every published file's relative path for consumers that prefer enumeration over wildcards.",
+ "columns": [
+ {
+ "name": "event_class",
+ "type": "INTEGER",
+ "not_null": true,
+ "primary_key": false,
+ "hive_partition": true
+ },
+ {
+ "name": "timestamp",
+ "type": "TIMESTAMP",
+ "not_null": true,
+ "primary_key": true,
+ "hive_partition": false
+ },
+ {
+ "name": "class",
+ "type": "INTEGER",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "state",
+ "type": "INTEGER",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ABER-CKGL",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ABER-CKP",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ESTADO-DHSV",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ESTADO-M1",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ESTADO-M2",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ESTADO-PXO",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ESTADO-SDV-GL",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ESTADO-SDV-P",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ESTADO-W1",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ESTADO-W2",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "ESTADO-XO",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "P-ANULAR",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "P-JUS-BS",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "P-JUS-CKGL",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "P-JUS-CKP",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "P-MON-CKGL",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "P-MON-CKP",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "P-MON-SDV-P",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "P-PDG",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "PT-P",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "P-TPT",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "QBS",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "QGL",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "T-JUS-CKP",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "T-MON-CKP",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "T-PDG",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "T-TPT",
+ "type": "DOUBLE",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "instance_id",
+ "type": "VARCHAR",
+ "not_null": true,
+ "primary_key": true,
+ "hive_partition": false
+ },
+ {
+ "name": "well_id",
+ "type": "INTEGER",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ },
+ {
+ "name": "well_kind",
+ "type": "VARCHAR",
+ "not_null": false,
+ "primary_key": false,
+ "hive_partition": false
+ }
+ ],
+ "primary_key": [
+ "instance_id",
+ "timestamp"
+ ],
+ "foreign_keys": [
+ {
+ "column": "instance_id",
+ "references_table": "instances",
+ "references_column": "instance_id"
+ },
+ {
+ "column": "event_class",
+ "references_table": "event_types",
+ "references_column": "event_class"
+ },
+ {
+ "column": "well_id",
+ "references_table": "wells",
+ "references_column": "well_id"
+ }
+ ]
}
}
}
diff --git a/parquet/petrobras_3w/schema.md b/parquet/petrobras_3w/schema.md
index 210912f..b153403 100644
--- a/parquet/petrobras_3w/schema.md
+++ b/parquet/petrobras_3w/schema.md
@@ -10,20 +10,117 @@ Static lookup of upstream event classes (`0..9`). Mirrors the `[NAMES]` / per-cl
**Columns:**
-| Column | Type | Nullable | PK | Description |
-|--------|------|----------|----|-------------|
-| `event_class` | INTEGER | No | ✓ | Integer event class (0 = NORMAL; 1..9 = anomaly categories). Primary key. Matches upstream's `LABEL`. |
-| `name` | VARCHAR | Yes | | Internal name (PascalCase with underscores, e.g. `HYDRATE_IN_PRODUCTION_LINE`). Mirrors upstream's `NAMES` list. |
-| `description` | VARCHAR | Yes | | Human-readable description (e.g. `Hydrate in Production Line`). Mirrors upstream's per-class `DESCRIPTION`. |
-| `has_transient` | BOOLEAN | Yes | | `true` for classes that carry a `TRANSIENT` precursor phase in their `class` column (1, 2, 5, 6, 7, 8, 9). `false` for `NORMAL` (0) and the two events upstream marks `TRANSIENT=False` (3 = Severe Slugging, 4 = Flow Instability). |
-| `transient_code` | INTEGER | Yes | | Per-observation label seen during the transient phase: `event_class + 100` when `has_transient = true`, NULL otherwise. Decodes raw `class` codes such as 101, 105, 108 in the observations time-series. |
-| `has_normal_prefix` | BOOLEAN | Yes | | `true` when instances of this class include a `class = 0` (NORMAL) precursor before the labelled event. Correlates with `has_transient` — events 0, 3, 4 carry only the steady class. |
+| Column | Type | Nullable | PK | Source | Description |
+|--------|------|----------|----|--------|-------------|
+| `event_class` | INTEGER | No | ✓ | file body | Integer event class (0 = NORMAL; 1..9 = anomaly categories). On `event_types` this is the primary key; on `instances` it is a foreign key back to `event_types.event_class`. |
+| `name` | VARCHAR | Yes | | file body | Internal name (PascalCase with underscores, e.g. `HYDRATE_IN_PRODUCTION_LINE`). Mirrors upstream's `NAMES` list. |
+| `description` | VARCHAR | Yes | | file body | Human-readable description (e.g. `Hydrate in Production Line`). Mirrors upstream's per-class `DESCRIPTION`. |
+| `has_transient` | BOOLEAN | Yes | | file body | `true` for classes that carry a `TRANSIENT` precursor phase in their `class` column (1, 2, 5, 6, 7, 8, 9). `false` for `NORMAL` (0) and the two events upstream marks `TRANSIENT=False` (3 = Severe Slugging, 4 = Flow Instability). |
+| `transient_code` | INTEGER | Yes | | file body | Per-observation label seen during the transient phase: `event_class + 100` when `has_transient = true`, NULL otherwise. Decodes raw `class` codes such as 101, 105, 108 in the observations time-series. |
+| `has_normal_prefix` | BOOLEAN | Yes | | file body | `true` when instances of this class include a `class = 0` (NORMAL) precursor before the labelled event. Correlates with `has_transient` — events 0, 3, 4 carry only the steady class. |
+
+---
+
+### `wells`
+
+Real-Well master, one row per distinct `well_id` derived from Instances with `well_kind = 'real'` (40 rows at the current upstream pin). Upstream anonymises every physical-well attribute (no basin, field, depth, or location), so the master is an identity-plus-statistics table: count of Instances, total 1-Hz Observations, and the time span across which the Well appears in the corpus. Simulated and drawn Instances have NULL `well_id` and contribute nothing here.
+
+**Columns:**
+
+| Column | Type | Nullable | PK | Source | Description |
+|--------|------|----------|----|--------|-------------|
+| `well_id` | INTEGER | No | ✓ | file body | Anonymised physical-well integer ID. On `wells` this is the primary key (one row per real Well). On `instances` it is parsed from the `WELL-NNNNN` prefix of `instance_id` and is NULL when `well_kind != 'real'` — a foreign key back to `wells.well_id`. |
+| `n_instances` | BIGINT | Yes | | file body | Number of Instances in the corpus drawn from this real Well. Equal to `COUNT(*) FROM instances WHERE well_kind = 'real' AND well_id = wells.well_id`. |
+| `first_ts` | TIMESTAMP | Yes | | file body | Earliest Instance `start_ts` for the Well. The Well's first appearance in the corpus. |
+| `last_ts` | TIMESTAMP | Yes | | file body | Latest Instance `end_ts` for the Well. The Well's last appearance in the corpus. |
+| `n_observations` | BIGINT | Yes | | file body | Total count of 1-Hz Observations contributed by this Well across all of its Instances. Equal to `SUM(n_rows) FROM instances WHERE well_kind = 'real' AND well_id = wells.well_id`. |
+
+---
+
+### `instances`
+
+One row per upstream Instance file (~2,228 rows). Identifies the Instance (`instance_id`), its provenance (`well_kind`, `well_id`, `source_file`), the operational regime it is framed around (`event_class`), and pre-aggregated per-Instance statistics (`start_ts`, `end_ts`, `duration_s`, `n_rows`, plus four `n_rows_*` counts that partition `n_rows` by `class` value). Corpus-wide balance and labelled-mass queries can run purely against this catalog without scanning the Observations time-series. `source_url` points at the published Observations file for the Instance (the URL pattern is fixed by ADR-0001).
+
+**Columns:**
+
+| Column | Type | Nullable | PK | Source | Description |
+|--------|------|----------|----|--------|-------------|
+| `instance_id` | VARCHAR | No | ✓ | file body | Primary key. The upstream source filename without `.parquet` (e.g. `WELL-00019_20120601165020`, `SIMULATED_00012`, `DRAWN_00003`). Stable across refreshes. |
+| `well_kind` | VARCHAR | Yes | | file body | Provenance of the Instance: `real` (from a physical Petrobras well), `simulated` (synthetic, generated upstream), or `drawn` (hand-crafted series). `well_id` is non-NULL only when `well_kind = 'real'`. |
+| `well_id` | INTEGER | Yes | | file body | Anonymised physical-well integer ID. On `wells` this is the primary key (one row per real Well). On `instances` it is parsed from the `WELL-NNNNN` prefix of `instance_id` and is NULL when `well_kind != 'real'` — a foreign key back to `wells.well_id`. |
+| `event_class` | INTEGER | Yes | | file body | Integer event class (0 = NORMAL; 1..9 = anomaly categories). On `event_types` this is the primary key; on `instances` it is a foreign key back to `event_types.event_class`. |
+| `start_ts` | TIMESTAMP | Yes | | file body | First `timestamp` value in the upstream Instance file. |
+| `end_ts` | TIMESTAMP | Yes | | file body | Last `timestamp` value in the upstream Instance file. |
+| `duration_s` | BIGINT | Yes | | file body | `end_ts - start_ts` in seconds. Derived from the per-Instance aggregates so consumers do not have to recompute it. |
+| `n_rows` | BIGINT | Yes | | file body | Number of 1-Hz observations in the upstream Instance file. Range across the corpus: ~21k (~6h) to ~243k (~3 days). |
+| `n_rows_warmup_null` | DOUBLE | Yes | | file body | Count of rows where `class IS NULL` — the warmup prefix seen on real-Well Instances (typically ~1 hour) where upstream chose not to assign a label. |
+| `n_rows_normal` | DOUBLE | Yes | | file body | Count of rows where `class = 0` and the Instance's `event_class` is not itself 0 — i.e. the NORMAL precursor before an anomaly. Event 0's `class = 0` rows are its labelled regime itself and are counted under `n_rows_steady` instead, so the four `n_rows_*` columns always partition `n_rows` without overlap. |
+| `n_rows_transient` | DOUBLE | Yes | | file body | Count of rows where `class = event_class + 100` (the developing phase before steady state). NULL when the row's `event_class` has `has_transient = false` in `event_types` (events 0, 3, 4) — the transient phase does not exist by design, distinct from a zero-row count. |
+| `n_rows_steady` | DOUBLE | Yes | | file body | Count of rows where `class = event_class` (the labelled operational regime at steady state). |
+| `source_file` | VARCHAR | Yes | | file body | Upstream parquet filename including `.parquet` extension, for cross-reference with the upstream repository. |
+| `source_url` | VARCHAR | Yes | | file body | URL of the published Observations file for this Instance (`observations/event_class=N/.parquet`). The URL pattern is fixed by ADR-0001 so it can be materialised here before the Observations files exist. |
+
+**Foreign keys:**
+
+- `event_class` → `event_types.event_class`
+- `well_id` → `wells.well_id`
+
+---
+
+### `observations`
+
+Per-Instance 1-Hz sensor time-series. Hive-partitioned by `event_class` into `observations/event_class=N/.parquet` — one file per Instance, ~2,228 files in total. Each file preserves the upstream sensor columns verbatim (including hyphens: `P-PDG`, `ABER-CKGL`, `ESTADO-SDV-GL`, …), plus `class`, `state`, and `timestamp`. Three constant columns identify provenance per row: `instance_id`, `well_id`, `well_kind` (RLE-encoded, negligible storage). `event_class` is provided by the hive partition and is NOT stored in the file body. A `_files.json` manifest at the partition root lists every published file's relative path for consumers that prefer enumeration over wildcards.
+
+**Columns:**
+
+| Column | Type | Nullable | PK | Source | Description |
+|--------|------|----------|----|--------|-------------|
+| `event_class` | INTEGER | No | | hive partition | Integer event class (0 = NORMAL; 1..9 = anomaly categories). On `event_types` this is the primary key; on `instances` it is a foreign key back to `event_types.event_class`. |
+| `timestamp` | TIMESTAMP | No | ✓ | file body | Wall-clock timestamp of the 1-Hz observation. Strictly monotonic within an Instance file at exactly 1-second cadence. |
+| `class` | INTEGER | Yes | | file body | Per-observation regime label: `NULL` during the warmup prefix of a real-Well Instance, `0` for the NORMAL precursor before an anomaly, `event_class` for the labelled steady regime, or `event_class + 100` for the TRANSIENT phase (only on events where `event_types.has_transient = true`). Events 3 and 4 ship only the steady code — no transient and no NORMAL precursor. |
+| `state` | INTEGER | Yes | | file body | Upstream-provided well operational status. Preserved verbatim from the source file; semantics are documented in upstream's `dataset.ini`. |
+| `ABER-CKGL` | DOUBLE | Yes | | file body | Opening of the GLCK (gas lift choke) [%%] |
+| `ABER-CKP` | DOUBLE | Yes | | file body | Opening of the PCK (production choke) [%%] |
+| `ESTADO-DHSV` | DOUBLE | Yes | | file body | State of the DHSV (downhole safety valve) [0, 0.5, or 1] |
+| `ESTADO-M1` | DOUBLE | Yes | | file body | State of the PMV (production master valve) [0, 0.5, or 1] |
+| `ESTADO-M2` | DOUBLE | Yes | | file body | State of the AMV (annulus master valve) [0, 0.5, or 1] |
+| `ESTADO-PXO` | DOUBLE | Yes | | file body | State of the PXO (pig-crossover) valve [0, 0.5, or 1] |
+| `ESTADO-SDV-GL` | DOUBLE | Yes | | file body | State of the gas lift SDV (shutdown valve) [0, 0.5, or 1] |
+| `ESTADO-SDV-P` | DOUBLE | Yes | | file body | State of the production SDV (shutdown valve) [0, 0.5, or 1] |
+| `ESTADO-W1` | DOUBLE | Yes | | file body | State of the PWV (production wing valve) [0, 0.5, or 1] |
+| `ESTADO-W2` | DOUBLE | Yes | | file body | State of the AWV (annulus wing valve) [0, 0.5, or 1] |
+| `ESTADO-XO` | DOUBLE | Yes | | file body | State of the XO (crossover) valve [0, 0.5, or 1] |
+| `P-ANULAR` | DOUBLE | Yes | | file body | Pressure in the well annulus [Pa] |
+| `P-JUS-BS` | DOUBLE | Yes | | file body | Downstream pressure of the SP (service pump) [Pa] |
+| `P-JUS-CKGL` | DOUBLE | Yes | | file body | Downstream pressure of the GLCK (gas lift choke) [Pa] |
+| `P-JUS-CKP` | DOUBLE | Yes | | file body | Downstream pressure of the PCK (production choke) [Pa] |
+| `P-MON-CKGL` | DOUBLE | Yes | | file body | Upstream pressure of the GLCK (gas lift choke) [Pa] |
+| `P-MON-CKP` | DOUBLE | Yes | | file body | Upstream pressure of the PCK (production choke) [Pa] |
+| `P-MON-SDV-P` | DOUBLE | Yes | | file body | Upstream pressure of the production SDV (shutdown valve) [Pa] |
+| `P-PDG` | DOUBLE | Yes | | file body | Downhole pressure at the PDG (permanent downhole gauge) [Pa] |
+| `PT-P` | DOUBLE | Yes | | file body | Subsea Xmas-tree pressure downstream of the PWV (production wing valve) in the production line [Pa] |
+| `P-TPT` | DOUBLE | Yes | | file body | Subsea Xmas-tree pressure at the TPT (temperature and pressure transducer) [Pa] |
+| `QBS` | DOUBLE | Yes | | file body | Flow rate at the SP (service pump) [m3/s] |
+| `QGL` | DOUBLE | Yes | | file body | Gas lift flow rate [m3/s] |
+| `T-JUS-CKP` | DOUBLE | Yes | | file body | Downstream temperature of the PCK (production choke) [oC] |
+| `T-MON-CKP` | DOUBLE | Yes | | file body | Upstream temperature of the PCK (production choke) [oC] |
+| `T-PDG` | DOUBLE | Yes | | file body | Downhole temperature at the PDG (permanent downhole gauge) [oC] |
+| `T-TPT` | DOUBLE | Yes | | file body | Subsea Xmas-tree temperature at the TPT (temperature and pressure transducer) [oC] |
+| `instance_id` | VARCHAR | No | ✓ | file body | Primary key. The upstream source filename without `.parquet` (e.g. `WELL-00019_20120601165020`, `SIMULATED_00012`, `DRAWN_00003`). Stable across refreshes. |
+| `well_id` | INTEGER | Yes | | file body | Anonymised physical-well integer ID. On `wells` this is the primary key (one row per real Well). On `instances` it is parsed from the `WELL-NNNNN` prefix of `instance_id` and is NULL when `well_kind != 'real'` — a foreign key back to `wells.well_id`. |
+| `well_kind` | VARCHAR | Yes | | file body | Provenance of the Instance: `real` (from a physical Petrobras well), `simulated` (synthetic, generated upstream), or `drawn` (hand-crafted series). `well_id` is non-NULL only when `well_kind = 'real'`. |
+
+**Foreign keys:**
+
+- `instance_id` → `instances.instance_id`
+- `event_class` → `event_types.event_class`
+- `well_id` → `wells.well_id`
---
## Sensor-column glossary (Observations time-series)
-Mirrored verbatim from upstream `dataset.ini`'s `PARQUET_FILE_PROPERTIES` section. These columns will appear in `observations/event_class=N/.parquet` once issue #22 lands; the glossary is published here so consumers can plan queries against the table layout in advance.
+Mirrored verbatim from upstream `dataset.ini`'s `PARQUET_FILE_PROPERTIES` section. These columns appear in every `observations/event_class=N/.parquet` file body; the glossary is repeated here in a single table for quick reference.
| Column | Description |
|--------|-------------|
diff --git a/parquet/petrobras_3w/schema.sql b/parquet/petrobras_3w/schema.sql
index 290cadd..6020057 100644
--- a/parquet/petrobras_3w/schema.sql
+++ b/parquet/petrobras_3w/schema.sql
@@ -13,3 +13,76 @@ CREATE TABLE event_types (
has_normal_prefix BOOLEAN,
PRIMARY KEY (event_class)
);
+
+CREATE TABLE wells (
+ well_id INTEGER NOT NULL,
+ n_instances BIGINT,
+ first_ts TIMESTAMP,
+ last_ts TIMESTAMP,
+ n_observations BIGINT,
+ PRIMARY KEY (well_id)
+);
+
+CREATE TABLE instances (
+ instance_id VARCHAR NOT NULL,
+ well_kind VARCHAR,
+ well_id INTEGER,
+ event_class INTEGER,
+ start_ts TIMESTAMP,
+ end_ts TIMESTAMP,
+ duration_s BIGINT,
+ n_rows BIGINT,
+ n_rows_warmup_null DOUBLE,
+ n_rows_normal DOUBLE,
+ n_rows_transient DOUBLE,
+ n_rows_steady DOUBLE,
+ source_file VARCHAR,
+ source_url VARCHAR,
+ PRIMARY KEY (instance_id),
+ FOREIGN KEY (event_class) REFERENCES event_types (event_class),
+ FOREIGN KEY (well_id) REFERENCES wells (well_id)
+);
+
+-- `observations` is published as a hive-partitioned tree:
+-- observations/event_class=N/.parquet
+-- `event_class` lives in the partition path; every other column lives in the file body.
+CREATE TABLE observations (
+ event_class INTEGER NOT NULL,
+ timestamp TIMESTAMP NOT NULL,
+ class INTEGER,
+ state INTEGER,
+ "ABER-CKGL" DOUBLE,
+ "ABER-CKP" DOUBLE,
+ "ESTADO-DHSV" DOUBLE,
+ "ESTADO-M1" DOUBLE,
+ "ESTADO-M2" DOUBLE,
+ "ESTADO-PXO" DOUBLE,
+ "ESTADO-SDV-GL" DOUBLE,
+ "ESTADO-SDV-P" DOUBLE,
+ "ESTADO-W1" DOUBLE,
+ "ESTADO-W2" DOUBLE,
+ "ESTADO-XO" DOUBLE,
+ "P-ANULAR" DOUBLE,
+ "P-JUS-BS" DOUBLE,
+ "P-JUS-CKGL" DOUBLE,
+ "P-JUS-CKP" DOUBLE,
+ "P-MON-CKGL" DOUBLE,
+ "P-MON-CKP" DOUBLE,
+ "P-MON-SDV-P" DOUBLE,
+ "P-PDG" DOUBLE,
+ "PT-P" DOUBLE,
+ "P-TPT" DOUBLE,
+ "QBS" DOUBLE,
+ "QGL" DOUBLE,
+ "T-JUS-CKP" DOUBLE,
+ "T-MON-CKP" DOUBLE,
+ "T-PDG" DOUBLE,
+ "T-TPT" DOUBLE,
+ instance_id VARCHAR NOT NULL,
+ well_id INTEGER,
+ well_kind VARCHAR,
+ PRIMARY KEY (instance_id, timestamp),
+ FOREIGN KEY (instance_id) REFERENCES instances (instance_id),
+ FOREIGN KEY (event_class) REFERENCES event_types (event_class),
+ FOREIGN KEY (well_id) REFERENCES wells (well_id)
+);
diff --git a/scripts/export/petrobras_3w/schema_doc_generator.py b/scripts/export/petrobras_3w/schema_doc_generator.py
index f3383db..8a3309d 100644
--- a/scripts/export/petrobras_3w/schema_doc_generator.py
+++ b/scripts/export/petrobras_3w/schema_doc_generator.py
@@ -476,7 +476,9 @@ def _write_schema_md(
nullable = "No" if col["not_null"] else "Yes"
pk = "✓" if col["primary_key"] else ""
source = "hive partition" if col.get("hive_partition") else "file body"
- desc = COLUMN_DESCRIPTIONS.get(col["name"], "")
+ desc = COLUMN_DESCRIPTIONS.get(
+ col["name"]
+ ) or dataset_ini.sensor_descriptions.get(col["name"], "")
lines.append(
f"| `{col['name']}` | {col['type']} | {nullable} | {pk} | "
f"{source} | {desc} |"
@@ -497,10 +499,9 @@ def _write_schema_md(
lines.append("")
lines.append(
"Mirrored verbatim from upstream `dataset.ini`'s "
- "`PARQUET_FILE_PROPERTIES` section. These columns will appear in "
- "`observations/event_class=N/.parquet` once issue "
- "#22 lands; the glossary is published here so consumers can plan "
- "queries against the table layout in advance."
+ "`PARQUET_FILE_PROPERTIES` section. These columns appear in every "
+ "`observations/event_class=N/.parquet` file body; the "
+ "glossary is repeated here in a single table for quick reference."
)
lines.append("")
lines.append("| Column | Description |")
@@ -654,6 +655,36 @@ def _write_readme(schemas: dict[str, dict], path: Path) -> None:
lines.append("ORDER BY w.n_instances DESC;")
lines.append("```")
lines.append("")
+ lines.append("### Per-Well cross-validation split (leave-one-Well-out)")
+ lines.append("")
+ lines.append(
+ "Training models on Petrobras 3W should split by `well_id`, not by "
+ "Instance — Instances drawn from the same physical Well share "
+ "operating conditions and would leak signal across a naive shuffle. "
+ "Assign each real Well a stable fold index from the `well_id` "
+ "modulo the desired fold count, then derive the test Instances for "
+ "fold `k` directly from `instances.parquet`:"
+ )
+ lines.append("")
+ lines.append("```sql")
+ lines.append("WITH folded AS (")
+ lines.append(" SELECT well_id, well_id % 5 AS fold")
+ lines.append(f" FROM '{base_url}/wells.parquet'")
+ lines.append(")")
+ lines.append("SELECT i.instance_id, i.event_class, i.n_rows, i.source_url")
+ lines.append(f"FROM '{base_url}/instances.parquet' i")
+ lines.append("JOIN folded f USING (well_id)")
+ lines.append("WHERE f.fold = 0 -- held-out test set; train on fold != 0")
+ lines.append("ORDER BY i.event_class, i.instance_id;")
+ lines.append("```")
+ lines.append("")
+ lines.append(
+ "The simulated and drawn Instances (`well_kind <> 'real'`, "
+ "`well_id IS NULL`) are excluded from the join above and can be "
+ "added to the training set independently — they have no physical "
+ "Well to leak against."
+ )
+ lines.append("")
lines.append("### Load all real-Well Observations of one event class")
lines.append("")
lines.append(
diff --git a/tests/petrobras_3w/conftest.py b/tests/petrobras_3w/conftest.py
index 7d88e3a..19064ba 100644
--- a/tests/petrobras_3w/conftest.py
+++ b/tests/petrobras_3w/conftest.py
@@ -42,6 +42,8 @@
import duckdb
+from scripts.transform.petrobras_3w.upstream_stager import parse_dataset_ini
+
@dataclass(frozen=True)
class _InstanceSpec:
@@ -123,6 +125,19 @@ def _padding_specs() -> tuple[_InstanceSpec, ...]:
_INSTANCE_SPECS: tuple[_InstanceSpec, ...] = _PRIMARY_SPECS + _padding_specs()
+def _sensor_columns(staging_dir: Path) -> tuple[str, ...]:
+ """Return upstream's sensor columns in `PARQUET_FILE_PROPERTIES` order,
+ excluding the ones the fixture writes explicitly (`timestamp`, `class`,
+ `state`).
+ """
+ ini = parse_dataset_ini(staging_dir)
+ return tuple(
+ column
+ for column in ini.sensor_descriptions
+ if column not in {"timestamp", "class", "state"}
+ )
+
+
def build_instance_parquets(staging_dir: Path) -> None:
"""Write the minimal per-Instance parquet fixtures under `/dataset/N/`.
@@ -130,11 +145,17 @@ def build_instance_parquets(staging_dir: Path) -> None:
DuckDB so the fixture format matches what `read_parquet` will see in
the pipeline (timestamp column written as a TIMESTAMP).
- The fixture carries a hyphenated sensor column (`P-PDG`) so the
- Observations writer's column-name fidelity policy is exercised in
- the smoke test.
+ Carries all 27 upstream sensor columns from `dataset.ini`'s
+ `PARQUET_FILE_PROPERTIES`, so the reflected schema docs published from
+ a fixture-driven run match production's column inventory. Most sensor
+ values are constant placeholders; only `P-PDG` varies so the
+ column-name fidelity policy and parity sensor-aggregate checks have
+ distinguishable data to compare across upstream → catalog → published.
"""
dataset_root = Path(staging_dir) / "dataset"
+ sensor_cols = _sensor_columns(staging_dir)
+ sensor_col_ddl = ",".join(f' "{name}" DOUBLE' for name in sensor_cols)
+ placeholders = ",".join(["?"] * (3 + len(sensor_cols)))
con = duckdb.connect()
try:
for spec in _INSTANCE_SPECS:
@@ -147,7 +168,7 @@ def build_instance_parquets(staging_dir: Path) -> None:
' "timestamp" TIMESTAMP,'
' "class" INTEGER,'
' "state" INTEGER,'
- ' "P-PDG" DOUBLE'
+ f"{sensor_col_ddl}"
")"
)
rows = [
@@ -155,12 +176,18 @@ def build_instance_parquets(staging_dir: Path) -> None:
f"2012-01-01 00:00:{i:02d}",
cls,
0,
- 1.0e7 + i,
+ *(
+ # `P-PDG` varies row-to-row so sensor aggregates
+ # diverge under any byte-level mutation in the
+ # parity tests; other sensors stay constant.
+ 1.0e7 + i if name == "P-PDG" else 0.0
+ for name in sensor_cols
+ ),
)
for i, cls in enumerate(spec.classes)
]
con.executemany(
- "INSERT INTO staging_rows VALUES (?, ?, ?, ?)",
+ f"INSERT INTO staging_rows VALUES ({placeholders})",
rows,
)
con.execute(
diff --git a/tests/petrobras_3w/test_smoke.py b/tests/petrobras_3w/test_smoke.py
index c09b1ff..af7eb93 100644
--- a/tests/petrobras_3w/test_smoke.py
+++ b/tests/petrobras_3w/test_smoke.py
@@ -295,11 +295,36 @@ def test_pipeline_emits_documentation(tmp_path: Path) -> None:
):
assert f"`{column}`" in schema_md, f"schema.md missing {column}"
+ # The glossary intro no longer refers to issue #22 as future work; the
+ # Observations writer landed there. Stale forward-references would
+ # confuse consumers reading the published docs.
+ assert "once issue #22 lands" not in schema_md, (
+ "stale '#22 lands' forward-reference in schema.md"
+ )
+
+ # Sensor column descriptions are inlined in the observations table — not
+ # left blank with the glossary as the only reference.
+ obs_section, _, _ = schema_md.partition("## Sensor-column glossary")
+ _, _, observations_block = obs_section.partition("### `observations`")
+ assert "Downhole pressure at the PDG" in observations_block, (
+ "observations P-PDG row should inline upstream's sensor description"
+ )
+
# README records the pinned git tag + dataset version.
readme = (out_dir / "README.md").read_text()
assert PIN_GIT_TAG in readme
assert PIN_DATASET_VERSION in readme
+ # Four canonical query examples from issue #24 acceptance criteria must
+ # be present (a load-by-event-class, b fetch-by-URL, c per-Well CV
+ # split, d corpus-wide balance from catalog only).
+ assert "WHERE event_class = 8" in readme, "missing event-class load query"
+ assert "Fetch one specific Instance" in readme, "missing fetch-by-URL query"
+ assert "leave-one-Well-out" in readme, "missing per-well CV split query"
+ assert "Corpus balance from the Instance catalog" in readme, (
+ "missing corpus-balance catalog-only query"
+ )
+
# LICENSE-3W-DATA.md attributes CC BY 4.0 with upstream attribution.
license_text = (out_dir / "LICENSE-3W-DATA.md").read_text()
assert "CC BY 4.0" in license_text
@@ -1379,10 +1404,11 @@ def test_parity_detects_per_instance_only_drift(tmp_path: Path) -> None:
target_gain,
# Append a synthetic second row 1 second later so the partition-
# level total stays the same but this Instance now has 2 rows
- # instead of 1.
- "INSERT INTO staged SELECT "
- '"timestamp" + INTERVAL \'1 second\', class, state, "P-PDG", '
- "instance_id, well_id, well_kind FROM staged",
+ # instead of 1. `SELECT *` over the table re-uses the full column
+ # set unchanged; the column-list-by-name approach would have to be
+ # kept in lockstep with the fixture's sensor inventory.
+ 'INSERT INTO staged BY NAME SELECT * REPLACE ("timestamp" + '
+ "INTERVAL '1 second' AS \"timestamp\") FROM staged",
)
with pytest.raises(parity.ParityRowCountPerInstanceError):
From 7f91b19b7ac5a53010dd378f56b38b613e76c5ab Mon Sep 17 00:00:00 2001
From: oskrgab
Date: Mon, 18 May 2026 03:32:17 +0000
Subject: [PATCH 09/11] Corrected scripts and parity tests
---
.gitignore | 1 +
CONTEXT.md | 6 +-
parquet/petrobras_3w/instances.parquet | Bin 0 -> 100818 bytes
parquet/petrobras_3w/schema.json | 62 +++++-----
parquet/petrobras_3w/schema.md | 18 +--
parquet/petrobras_3w/schema.sql | 22 ++--
parquet/petrobras_3w/wells.parquet | Bin 0 -> 1932 bytes
scripts/export/petrobras_3w/parity.py | 109 +++++++++++-------
.../petrobras_3w/schema_doc_generator.py | 16 ++-
scripts/export/petrobras_3w/validator.py | 49 +++++---
scripts/run_pipeline/__init__.py | 0
scripts/run_pipeline/argentina/__init__.py | 0
scripts/run_pipeline/argentina/__main__.py | 29 +++++
scripts/run_pipeline/petrobras_3w/__init__.py | 0
scripts/run_pipeline/petrobras_3w/__main__.py | 34 ++++++
15 files changed, 229 insertions(+), 117 deletions(-)
create mode 100644 parquet/petrobras_3w/instances.parquet
create mode 100644 parquet/petrobras_3w/wells.parquet
create mode 100644 scripts/run_pipeline/__init__.py
create mode 100644 scripts/run_pipeline/argentina/__init__.py
create mode 100644 scripts/run_pipeline/argentina/__main__.py
create mode 100644 scripts/run_pipeline/petrobras_3w/__init__.py
create mode 100644 scripts/run_pipeline/petrobras_3w/__main__.py
diff --git a/.gitignore b/.gitignore
index 6d0a066..99ca9e1 100644
--- a/.gitignore
+++ b/.gitignore
@@ -225,3 +225,4 @@ data/petrobras_3w/
.DS_Store
.gemini/
data/Argentina
+parquet/petrobras_3w/observations
diff --git a/CONTEXT.md b/CONTEXT.md
index d0b8228..1204210 100644
--- a/CONTEXT.md
+++ b/CONTEXT.md
@@ -140,8 +140,8 @@ A four-table relational shape, derived from upstream's per-instance Parquet file
**Lookup** (`event_types.parquet`, exactly 10 rows):
`event_class` (PK, 0..9), `name` (canonical PascalCase from `dataset.ini`, e.g. `HYDRATE_IN_PRODUCTION_LINE`), `description` (human-readable, e.g. `Hydrate in Production Line`), `has_transient` (boolean, false for `{0, 3, 4}`), `transient_code` (`event_class + 100` when `has_transient`, NULL otherwise), `has_normal_prefix` (boolean, true for events that include a class=0 precursor in the data — correlates with `has_transient`).
-**Instance catalog** (`instances.parquet`, one row per Instance, ~2,228 rows):
-`instance_id` (PK, source filename without extension), `well_kind` (enum: `real | simulated | drawn`), `well_id` (FK to `wells.parquet`, NULL when `well_kind != real`), `event_class` (FK to `event_types.parquet`), `start_ts`, `end_ts`, `duration_s` (derived), `n_rows`, `n_rows_warmup_null` (rows where `class IS NULL`), `n_rows_normal` (rows where `class = 0` AND `event_class <> 0` — i.e. the NORMAL precursor before an anomaly; event 0's `class = 0` rows are its labelled regime and roll into `n_rows_steady` instead so the four buckets always partition `n_rows`), `n_rows_transient` (rows where `class = event_class + 100`; NULL when `has_transient = false`), `n_rows_steady` (rows where `class = event_class`), `source_file` (upstream filename for cross-reference), `source_url` (URL to the published Observations parquet). The four `n_rows_*` columns let corpus-wide balance and labeled-mass queries run purely against this catalog without scanning Observations.
+**Instance catalog** (`instances.parquet`, one row per `(instance_id, event_class)` pair):
+`instance_id` (source filename without extension; **not** unique on its own — upstream re-publishes each synthetic `SIMULATED_*` / `DRAWN_*` series under several event classes, so the composite `(instance_id, event_class)` is the primary key; the hive partition path on `observations/` encodes the same identity), `well_kind` (enum: `real | simulated | drawn`), `well_id` (FK to `wells.parquet`, NULL when `well_kind != real`), `event_class` (FK to `event_types.parquet`), `start_ts`, `end_ts`, `duration_s` (derived), `n_rows`, `n_rows_warmup_null` (rows where `class IS NULL`), `n_rows_normal` (rows where `class = 0` AND `event_class <> 0` — i.e. the NORMAL precursor before an anomaly; event 0's `class = 0` rows are its labelled regime and roll into `n_rows_steady` instead so the four buckets always partition `n_rows`), `n_rows_transient` (rows where `class = event_class + 100`; NULL when `has_transient = false`), `n_rows_steady` (rows where `class = event_class`), `source_file` (upstream filename for cross-reference), `source_url` (URL to the published Observations parquet). The four `n_rows_*` columns let corpus-wide balance and labeled-mass queries run purely against this catalog without scanning Observations.
**Observations time-series** (`observations/event_class=N/.parquet`, ~2,228 files):
Hive-partitioned by `event_class` only. Each file is a single Instance's 1-Hz rows. Columns: the 27 source sensor columns + `class` (per-observation label, may include the +100 transient codes) + `state` (well operational status) + `timestamp` + three added constant columns: `instance_id`, `well_id`, `well_kind`. `event_class` is provided by the hive partition, not stored in the file body.
@@ -180,7 +180,7 @@ Rationale: 2,228 instances × ~0.4–1 MB each keep every file well under Cloudf
The export step asserts the following before writing Parquets; failure aborts publish.
-1. `instances.instance_id` is unique.
+1. `instances` has unique `(instance_id, event_class)`. `instance_id` alone is not unique because upstream re-publishes each synthetic `SIMULATED_*` / `DRAWN_*` series under several event classes (~225 such pairs at the pinned tag).
2. Every `instance_id` referenced by an Observations file exists in `instances.parquet`.
3. Every non-NULL `instances.well_id` exists in `wells.parquet` (FK integrity); NULL only when `well_kind != real`.
4. Every `instances.event_class` exists in `event_types.parquet` (FK integrity).
diff --git a/parquet/petrobras_3w/instances.parquet b/parquet/petrobras_3w/instances.parquet
new file mode 100644
index 0000000000000000000000000000000000000000..0e0c9076c963814a9231eb8460348f79988730f5
GIT binary patch
literal 100818
zcmd4a3tSBC|3CiO+G;yZwe6-&n@S}lyR*Afjcm0lBuP{xsU)Y6gq%`@oRx&QNyxdJ
z;~plZEE%k>P%5I3itQC1AySGj
zcjrCeD_JZvQpp&3STbcKcbCg1V;z-Du26-AQ##DJB20%TkbB5=rnmnd
zmZ-4AQtp8tQYb=_DQhLZdB_-5NH}GxWbm7IKZBCVSEbx7s)FJx8J;I!t_q3dVVO#%
zi1O9qNh$ZJs`h!gA}W~@s2G`Jbx0A83P&AVfuq4rl~Iip*7TSlr6~Di4wDI(Fuc&H
zB1*u>RJ=|>vv}Bok@KvP2UtUYI!=I6_-vx&AzievSA50p$&z^Ngyj`SBvbm8$Q`~?
zrc;blq}<}eF++CmQ-b~alweOjCHTqo;D{p1uLRp6ZE}eg#s04o$C)@m%8exa!^H9T
zh)5pg?nAbnkg7$wlk~qP@nk=avnVEPm@fIkmop)Ib;wCZ_4;c8@|2?l?u;Vd0NKR%
zU?N`X{(Uy~BA@ThEczZci^r43X7Z7sY+nzD2r;W_pN;*8?|)BX<1g}g(lG;qo{;VS
znnU8A=OoYE;&16;sTl(K*dN`Qh)Up9=zF6FeN5iT9Y2oD(g
zNAh^#SRyu{n1DzgcH@p}!Co~KR|w>CMSKIKI6#j`HGB%56h9Eh7oHZ3s#gz0N3BhexTf7Nv?X(4tgWj5?x{qqw7#
z8>vvmV;f05VB%A?kP8YK*?>iv#J)($6R>jk_#%{HEUMzgoTijqwXtfJ$PIN*p^Avt
zHg%ILWlRvO6SGj@CijRp(WWmL@EJvfoZnvf=m92_Eb0SFqTSo5BRt^%lc*7{%|rYP#SGq!IS!*)0ar;dxD_N
z0$wtfi8)H;H#)>m@YgC5Zp`^8jtV%ZFXoUcmHK-cNu&WYjHJ?7j^eX&90m?P2Spj{ODPsdjtW5M8Hb@zU_E<8C)&Um2aW5m
ziRz+7dH9utxEZ31d8Aw5yeN+g*-^P5<{`=geb|_fYlSMn$Yky^h7=D*Mkg3V_-Y9~
zI*|1l`j|0t7T5Us?tPzTQb&ZF4vT7*%Oh)$dy5>i2=fz={|w7pjiV8;GB*TQ3#@=+
zXJcw8Mc8Hm#rUM*9Qp1tJFkdJUE0f7C=(-m@}zW*5Lc95Wb)VTVVeZS&XE-R3lSv4
zg!(8VF<)_+k%IMZ8Q~JpTWO>PJrKQ_jXgq%G~UV()v+1@D)xbhULdC2u|(i5b4P%s
zm=HsOacKKR!Yq#0FB;c?@(?aHz(K|@H4v82$m=3pTiJvXt^aH)kNFybi!8rTj66RO8Q}vFdsmB8Cpy!Py`uAPt07}$+;*`C>!qdQv(g+hLloh23v?x
zJ#lv-_a=ctCg;YBk)YBnw~$KdaVG2`A(o>B#I@QX(!4WAzF{>P=b|m;_$rzDo+IvX
zS*$N|fb!66G^bQ{k!W5i6!LUA*TDj3E$8EWVzeYu)Devi&sa(+9dpzM*r0u>ofqeX
zW5L}B8;lo=IeI`xO2$1UzrWd$-D)L~@?PR9*yUn^3ZXAyYqW8mI>xBBbb|#B$j6tn
z(UpV;R3dAzhN)9XBNcNRXDZ_|?Tk4RNi`)nfOucrXoM#u*nK_@AKAv^=te0!3|SRW
zXJ9O?Z?;pjI5R2jMjBA4S-fz6p;!`*NGR;#gbNnrhs$)K5(Q*(BQ*hm!aZb*NRo^h
zN`XntH(XA-vn7g>Kkj~6LArWUDXx)_lOl>!N~}T)#pFa(F)pa7!f+{mXEN?SJ;m7P
z1XOv|WVE1ESVTLEaH9~87b9ZQ-`Xn{ww~SrZN3TkvYl)|_UA0X{^*`o_)Tn2KWbl&
z9l0S97Y`u@H=X)h$R`}qR2R&=e9J5~kV=IqMyQMNIy|eDhG=8aE++Ic1NOVvI2lDM
z87)n;nlK?jk)F|$#|x^!GF)rYaJ1-w_KVe016C?U0rhoI?{`H3aMR4PK`9Bc<8pc;
z6Xy%q^$zWquQMza{GfWqOg8kcK?DVg%KATVsd^VB1zzG(ifsdG$jmYJ<3BTq?
z+WRCS%kQuC$YK0FA$G4$`)^jG~KN{f?-)KhxkB1(Tm6<;F
z*uD)zuQtQIme0uB*TfulVO$)B>Jmy4>W|}h5HiP#v5TNkv%#sze^=x<3A7AZ2UY45
zO`3`581T_}j}@R*^PLfg@9Yq!OIqM6`m-ovhz`YUW6|u;&A
zI^$}x^|w$)u*P~l`hlD`Jj7cMshDRB=+kn#)`g7_inAy;Y?okwH7G{o*;T?}_7a($
zkc@5BxxK{U*dYm8^|!d$dX$^R4I^@An$Qttp%tx0eU5f>rsSG+DwJhQYxQ9jf+L#Z
zTvJz+-+2=5cPyX-^w}z&L@EpvqiY&uX^oI}#h$cJEUnSnUW+y`9p{KJFOeGIrp|E1
zol)o&c3#L-w3bhG=RRm*ldj!KD5>j0YB@^LO`)206DMOqY%UK|6mOQ4M^ic41#;Y9
z(6v~=4lpqz5%8M}^0i%cfr}Z3Dj#!}biouP1<}OMDz<%YxgtK9@=4`tw2iZbWrlR$
zPE4gPsgP&_(G#udtyXNE0S$M0cpENKZIvC$0E<;y*;P4B0WxiJd#sTGD8Q}7M(2@e`Yztl&c%>?;{
z3D*PHfVDvcCQDc{aU&CVGVx$cC=3qZ8#Sc=z|cUSp?<#ldH&|uluQ(4q9hX)nXqKy
zMkelL;-N7|LmX`n|9@h__|HYKBU}9MP84M2e`=y6Tl`ZK71`pSny_Sxe`?}Jw)m$e
z?qrL9YT`k*_}`vr&XI=Y->0~c33~q@C+fwf!mvQnN$zMQb5|(1@20}>Jy53%d#gM7
zSoJjJyD4a)Dc?iE?l;nrq5;ebr`w<}U(EMUInmrRAOY3U~It5oX-zqso}Qt{BFX
zidZ?Ta$_(^W)UXl481X+(aV%RVv5#>2;*mRY(KFR!Z{=S+LLu~Mu7B?tB^Yb5mo&<
zut%)X=9SCn15T6-V>I;1p0tM;&2Yaid^Z`rQA!`RW-BN(IOOzwB{LWOJwF2;C(Kp{
zOf_gL(*FD3>>_P{sj>fV-L?%z|I}aM1|!!RU5kAEZJQ_=9qYtOI#Y7t{+_r8cnC*{
zIGMGpl+vt2mq7TTH=AdLIf_-k6XEbfd$Mu^9GVAxrXxFEA(rwws`Ks03M%!4Fs#ts
zjx{Hn(>DZ63Td1(^fPDXI66i2Q4Iz+8U*r4;nq-us^fC39wKEerDXNna)#V(9CdMr
zrRNRhGP+1PI~{~ifm|iyrbs26upGA+#@hgEYqa_zPHXigU-yKg<^B>41mPqpyPkyk
zaw%zprfSi;q$exDz$v|+MF8eWW8%49`uz=br2dCw!e#?p(=bnZu?1_QO%9>blF8P`
z7rPDYVTm-{KM^+rO%2*&e$KcL#h8#Ao7!Rp9U@plRn8P+}`?q<<-K
z8*+?9zlApJgLJgEC!-7O*9YAeyGS;}NK&yak=ziDbVBonUgeKIT)AcOQU{_q9VXRtltNW!lDsNZUz`e#ejm1gCyB57sqAg;21PpkKXGY
zdl$Def8;}gF)iXK_dPF|a1>~rYg`t3c-!3Y%l`GQQqF%i`a1s8(W1RD4DDloiH?{P
zn%E)Vt&TW`RiLRJOv%_L6A4-~c1Lj1!YLeDU9`DJh`(9A9XeGvgmN8=@rtNBL+O-m
zT!y}sqoqcl-|n~ZK@xXi*Y`R4ST~gQergRP6x*&v#LuPamj#OgaDr|XjuAm!QoRXv`zP{3@XvmPmbuK
zh3L@U;fPs?1_iyp7y1aouKg)_5V{`07WBbx>K^*?Tom5K&LJk+`W)pE+eJ&*QH1si
zGu2Q^MkPh|XcWdr@`N)-<0g$(rd?#vS#6yn@Ksx1Ispxo1`*e0E~&r}NeNc5*wh;X
zHA1t15cLpV7QHkUR~)0Hl$?n&mSzd3JD^?vvKd2u!f^p;iU!mM=w#sp(fj4yKws}}
z?@-_RX0)cBHwSGEy4{kZyI(~k7Cf%3NY2@H40lSQHwj8Ny6#c9_V|_>5jHnuBUpK4@!YYyp67~hd#dbXQC08$o&jy>M7|kx>{RQlXaZwFkmEHW=aO2Si4A0S5H`PN;7_bsThNnNbT%7GTz~nieYWB6bms_DWy}pvn5zc
z&oM?&kV<1}P(u*x#t4`>#u6^3hU-E`2o#zstc-DAMuFZh#s#hXg+X2%0$BLA#P;3V
zMWU1ElZxZRq>2}mswGbLL9c=m8E77>8-OW%tgm_q2~Z;}$x$(}>aW@o&gTOLCxs)1kj8kI2=!4t9QGSK!~tQd
zBdPEoNWwW^_0~{v^0vX0oJ%A>J;;N_c!`u==Fl!8H$7^I8fAgQS+=i{Xo`Tkb1#YD
zVm~!wfZYyOvZrX2mPn2~I1#ow2xKa`jug2fMf_y%VK`vt71Aq2*9F~iO#^<3Vm0lf
zMP!Ui__!lWMo(y&3450Oy0M7aBC@b?LhFO>(m%bG$>otCnIlBBM)Ed&wvPl0G&}j2
zrcY!_w8$JoZhN#y%(UW~la+=Vk_Gn!%p1~2J4Z%WgabPwk1Tjm40>MS7FC`Y&caHR
zhnf_+%RF>`g#AZwqcPM^*P3wEW5^z^8it|3Vb)=cD@Lar!jAAP@<)&n;4ew&P)Wmm
zy08n|Z;xsg`^6?2PjcTy2MI%%xUG?LfEfISXMu4Nj1O|c`2zBgNr`b2($o6kfe>NQ
zW*QZvMYoYLTp^|v}`+ms>X<)G+l^QEKUlT8AdoNJbFQA
zDWJ`id>X5F7$cw;&($r=eaRqUc?4jP&HJnP{8PUE!1Zk{*xRz-bfm=@!hU+f(!XuVz
zy_kCfNriBwH@3#uD7i(}nvgTjw2NJ-BgC|bnw+SES#@UzVVWy_coY^!%+|B`l7!|8
zSuKcHZA|Gw;B{|!&a{rh_Vrw8!wKNjTmuQ?RH+)b0N`49VQ@7VqOFa1w`
zW5`)E{tsv}TF#L%IL&EaESRtPuP12t{Iz2H=IxX51F#8B<-a=#Xz~7U@8e&ygMd^1
z|C$~BAGH?+xk&$Uf)P7Sr1Af4;s5091NwjR*Z!CD{|`U@f6}r1*GKwKq$p*#MAr2V$!OSa${;(X+GQvE?7~y8|zKB@qTBc%1
zUy`d`VJ4+N_7KQ1e*L;P`bH}p(BgTkZzNYLRcsMPZ)m9_pH(r9j`-@@net$-;)#&}
z?_zNocuIsm#Q`(YwVmMY@4&}Um;=9UPUzqNm^q>9RaR|dqWmyT}q5WuPATqJ1BnXY7D+`62$xCX&yU|`wI9jx|@g_wF}KPgCi7+HlA5Hz12&gQYo38c*ut$GgHG3i87N~
zd1!Dv!lTnD6~=4m%sI3$9D|W~R1{N+EZ#UsoQ!8A*3Kx*e~64L_#QGkZZn<-@ZHc=
z82cL@Ytg4fMy$d^!Q}Q9b20t-3bGY`LP6&*r|@DQ9`ln|C+s$&`=MV%YxSW6W^<`p
zW>OKu%2*ZaUtrxfQ2gKjyFO6t+FRFR+-zk3tp3P0+OpUmS)Ch1mwDsiMd@#NR;5{n
z?25-gEt8{*QDnS9NSk=!?IQI90j-UBnT&py%J)!mH_^BdE;pc`1k!7Vt4HafQ(HVw
zxILDei)KCwP1h+Jv~%C%wH2bzm<~;$
zd-cFW6TDicu;`8PS~)#z0d3UB6;F8-LTf|%O$L2m&21fvxnBmjcD;fGcq9|>r&HTT
zz5o8-^^JP!X;y6m=KMeXEgLXLJx#0ah4p{&H~EFNYowt~M4*rfraQZk%6PW}-(8`e
zGEI2IklwJ*2v3oeoL(Ht=Nm)3Ey)hXI4#})R4A3K)@X}m0kkaGR4G@fJlGgK=%#bT
zT!|5CSEtkT*9bLs=YqF`JcJ1%y2EZ}ho!}Xt{4+}E~XcpWv7Ytr96~#;hIy-2eNHI
zG>YRo3rBl;_(mhVPpaas97W#!Zovzu*9=DKd=`eaCW}xXlMz0JTY}h`f#}LOtzx@d
zqD~wTbM|=af;O+9Lupoo8@uqA2zA2F!U(UmD%r&eSh;y6tKR{;`ExLLH5O;(cagUC
zam9mgg>f=HI+}htnBiMsCq))?UI$8rzBi@l@rK^M49_UIVGtQp4y@lpMFruv)_BWFCg
z1`~t?m0d2Y?t$etrWpQov{v6m9ki%9%woV#IVPUIZ%K_E-7Z_WPN;@Hh$9d
zgqY+pad>xnoH(KVEo%N>Z&52N|Lq-WtsVN5h}Xc=5>U2UE>sRrt67I?*4470inQ^%
zoru?uD&T2ZQ~R~Fk|_LWsh0SvR-ibT{A@n&-Bupd!iVgR90WN4yroU);UmBYPdF)Z
zEC%?HXRAi|7~@0ZLmrr#;v>R`?9T$94)|E&L*DJS!lxrX*7%?RP&W85Pv9K92Y};nM{lXMDQi(+wY-`er?Gx%NMv{7vn5eMs64eO?QX67Xy(7hVpJYS!U3
z>rxiHB5lfuC!+fC3V4(?Z@(6m#M3>+qYAWnN-az60v`V5Sv?*s>?_g27
ztzhY4mkMKTkt|T#ebuudZHli=#bVc{e4jRLN}`=F;)Q7k@d9+>Yg1pKCh7ckwhw_nrg#it_p#VDIZrVSUfaPEF$^pycrYf|LsbV
ztAyO@Cl36_ymDCaZv};qf|91i3)H&9(=O8CXXxpi($YbHMv^<
z_^SS{<-AXYe*d;Tb?@kMDSxYoa`1Id6hY@#Q*JrHoQC`t_0UPeetwpXmLsZZU+1^K0II@^v!XazY}_=
zFAR$<6;aXFGmj;~v=b9V3t@gp)U>s*^v&Sv?XWJ>Pkt0q4jyaDpzV2=w|Ag-gl*7s
z82#D&z-O4YM$bt1f{3!5R6fZ9mX+wgKArMndoBUzC$!bXwEeRG8NB@rP9~
z|M;!JTVd(o%Lj{L-RlyQQb=X)pIilPJLli1gPU(mOU|izg|-)#%q@oAk#ioNhtXdrb5$@c
zJ9c9o%?H|s2jO|DFK5ho7Uox>hw9Ru!W}vsq^0o#>R~a%xVA{lAYevI-
z-H&gl!P0Fnf-+#8XU%~eNIke=v=iD+yD;$>^rnwpzYL@I?CyFGrulBpd;#;Htp4*0
z%zgA==?k?v<-V-G`y1$8B;EQErgfU%{X0CIYjBeCGN*L={zy0t1yAR`xd2xjPD!~6
zJzp#vyKx)#|2pa69@y;?&-oM#i2A(vI`oe^)8Qek8|D-K3hwP$wfzUYsVeDURDktt
zD&|;2?$Oa3E-)=xo9PW_n{Qq|5DMQ#ycr3b99M=-hDFn24yM9Ffzo9)w14lvW-E--
z-}a;!F6%K!UJ9v|y9%nHnlpY>2R(=74R{Y-I%NFD+m3pdSIrYbQz|gb9+nJGKFq+3
zd2+T7OqE?_nFlA+5x+jSZ6yL`u&9GFdy9<~#%xbWl9F_`x156g1c
zdi(X7WeEcos3V-=)2N@PN!(ag;%G_Y#^z7c
zBS_q^?lSmF+@%`DO=Ok`;k97$kaf9AdypP0<+9t-B#EmffSBHxBAWm6?
zSe}QzzhaVphnc^uA3lNH;R})PU|enPPHHdCcjhv_8QdXE9c>54MTeY~!-jri>@+Z}
z_qBN;WZv`ljFlGn+FN+G^|6q(kp5E$-Pc+T->50+@E^$(O@0f6DF9
zOR&`7%F)}fZs#fU1}He@_PL>4M7jR)-2XGoztMlY&K0aTuD6y2tkagqI6-QwLn#Zb
z4vlo_3pd+$Ulu_X$`9tKVbg+h1rHzMO5ORvBuWW*5q-x3yhxZ^`|F1yhUX`5WcWm
zl0Fi~8st5n1Sg;EsY-<{YA`!@@Gp)8`?5`MpUMT(Irh>(XFg*s1Ebt6&uSpOWtH9#n3nxv^k~>RTj%OD
zD6|jglmQDqoXXCDMLUC^?tsho`KBI&fo3J=E+SHM;lUmcE;lUO?u++qoZM?L1{R
zbsgpP(KxiF7)0%_S|z=p4qkA7g*iv{48C>=ho{nbr$gP
zC(Yn4F!*KR0XGn@(;&4g6cf#@aB_FMk>_0T&f$7VN-?wL8S1g2e|Q+pNW
zkD1cz4_Hc#&U*>#Hidor2Jtq`FoPOwU)}qJ74)7Yx9AR|1&-5u!L;p`cLumnl44Qn?DjvG@I6Y$q%)Wj1$^KfzN4o=k
z&O*VRpwl;C*yvZrkD>SdctfTEzn>H_K?Bo_gRX_Z{M|mzqhVWWg^-$U
z?Uw^>#il!Uz;3hN#lK%-PUV%4pju#R({D@K;Dq@+IFpz)MeX0j^>GrDZ@j)DC&}Q4WF#6@{`M+SW
zao@Klhj5&oX7;v)rKcurmci_?PlY~^dOtW~Fx*|-_f8D7Klsfi2?ky4nZ5wVzNo9s
zhO1gEytc!5*Z#YYz&h6-A1*=adh_5r(00tzgALG|x@Ynk25qWL(Lp?}Z0+y<9oH)|
z_?Lix9_eW%)yA;!`K!VXkh)Vl(+<{t?K`A9EbYRRs-XAoeGR6#E`KXGRO|0c$@Q#r
zjE!`KrGq}Xy2Jcgi+LLOw{jGHV{ak3zTRIwXN`ikcE1l852*n$Hq&5T^2c{`Vd<7d
zmlnhPs~vLJz_bs0#^u3iYtLRgptt``<05FA7+rS)Qh6T^Ux0O&G8R?9(l-{v?!q*m
zUOyXQ^u!y#e}LZWMisO|+w<>}bsu88mkaxv!aAXal@%=Y-2J*U%#Y{J_kd|@Dzn^R
z^r@&gAL!lijvWAP>4o~Cka8E-M!~w6T}9(y>2iLeu4Og!
zRxGI42yI7-wrq#glETCTuDSl5x@Pm=@?fcD;glUPU*oPR
zf@ulYMI~^Y&$7%%Pch&3r+EV`Es4r+f>gIsm-oD7KKe7w
z2k$?IyPt}bZ(%>#+ikyK*%4)<3EF#Kj(mu)h0cSX70Y0C_pN$9u&nxJ(qK5l_QJgw
zGG9EhYZBb?{!!Wjc<=qDkJ&J;X5YZ=(AU~!=Mng2ytaNBjQ+H6`;?tHp4Bt8(qMkK
zi7{(n>E)=>d{`G9(&-SSTK!TlK-={mk8VM4ribTK7=6Qb(+8L~PV`-?5Zh}DhMU9E
zEk90lhI!!^td)?u`*gZ56!QHPH%}u@@Ttv@p@aXzpUse;sU7+YR$p6u+~f?7JMm@*
zTbS25akUKY9=@^08_w8d?Kv2RJ=(lB24*LCzMBe@iq7<1K+;3`d$Qq&-W`m$L*X9=
z2}dB4+;;sE{Boh%;SS6-Igs4|8@BX+{t4dm%Lwi}U+FvCr#&D3{T)Aw1)UIKK-2
zXes<#IV|P5+Nd#Kx7lu95d2%Yv%M|oH3jqDlUp~=hkq+)X*-^U=VLzK>;CbBu(am7
z_yVk(elGbIq|6TAu7|e!cQPNKci)X!TD92z>56wkm^LqMP-mELnOvlRrANjZ`@*`Q
z5eY*f)f7|_2W=N=x=n}Pw(QbHF#2@Y#&s|)Oxm{)=D#;BJPu1&@O7@hx-KoF??dX+
ztBZd^+o&gYEztX0b((We|x~XGq6;x{!|I?1q6?O1RZK#pMC=^(%rhW!MvI=>*>e1t|x4s
z+rW@F0bEab&wtilFL*rj8!s5MmJY+r2XqZ{l+}f{!LxqRkXbl_wVJotB%0H**BhCB6e)3Q(5(@)|Ot|k4Bd31q6%J?EL@yf)
z_dED~nh8s`O`o3$@4e-AZGcY`CNg{AmWHs{lhCCpyBmqaHS#{|$3T9hW!-c*BY3VS
ziMyNs__{9_R=yvqvkOMmEe|7c*2vlO(iM0#J+9LOqWz+!jqq_`M0E>H$a4$ON4d0F
z=(9!w&kW4e>IRb?tcG)tdph@`2;D7FN<<1{0n4zf!!1t6rSBOMI|lzt#sR
z^Gk9iragZ-k!VwW!1p6umC!PS7#7s)1TmlL-at$lb>HR_ocwF#D&k@1wO5G|`&3_u
zH5Xsne}=AEAN{_-72SOn5OXum?jjByao_MK-0-xvg_zpbs`U$IKN-1;$i-53iFxzu
zji?b~YJdM(YGS|=(zw9yjL5v=B_mDOWl1%$bH1*6)-P%(WaHkF@}|F
zxiJj|0}nP6>3kij32c~S6G@z#9=@7bcJN3ku_baoWeRiF_st+SmCia!w48tT8PWTw
zzpWXJc>8T6vAnVAIPr0>-VcdM_tL)+U2Bh83Snx8QSL-*L)*E;JU(4PEU5YNiDSzwj^@nrB4Nn3V5Wk2j@`*Xy&y^EfE8o0sSAN<+gn8??-K&UoQS;V_VTt|MM&hR$
z(+(E!_{|D`;;L-FtwiUJ!ka`(?HTWhrHd>(b%3U)F5f0QpmfQ6`Ph@r{I&_4z=l)`%=l8ddiEVH8($=shB-NXE{KViS;_HJ0UJ{4-O>vOI
za;FQq#BKrCE)$bxIesOs;^`&Wkoo(c_YigO>%Ac6-ILgNg2($u>3Rm5Q?JzromyaA
z;h!UCA&xIw-1ODNM(
z-Xlh&H9wZXp&1D+RFRE0HMHg1Zm6z1wlZEaxW60Wfmbb4-n(iT>K7dP3#wndx2
zcZN@wb=CMO%&EGuJ<`nfm{ZYjzjsNyi~iQGq<0mjVmyD)Po1?7;|m`J$`GE4fTT=-R}$;3d_zreGP|#r%~m@;K2hmV@ANj
zfiZW9wuPSqB4M`>udH~oyvJCrv9Q`+C71-WPh8NR3BN40dJq6T*Zx-b9tgioc!bEAb&X2GnMZRlRdrCgQ1eX_P7^}cjI;Tgbhw|=WjU971JhX5wj0frv89O
z2ilC{A-&Y$=x%aeR&wi;UMp&clpgvM0V_W)O{D)!ZT;x19TuoCm+>=u$B(RT-?Yhw*Xg9XrFtAC@j5o++~X-2t}d+<58;
zmu-#M?gaO1_B3~g&NjD3xj<{#1s@hp@Y;944Mxw7`$Y_NIC-0#Uy;c&BXYj}v)7Cy
z2AP=2{h+AUIGf~4!PU;@{b0%t&k!P)k`=rT<*T&tQGK`PC~rQuyn8|ZP}$cYczm1Q
zh*2>2*^H%cpK)95QaCvt0SteeS*zm>GabEZLs-s{?5L+1s^G{9+p9t%^{r+q!N#86>Oo7qH*{!K0
zJ@sSeB9i`gsCWfQk2$_Ao1|BFaNS7K`^-AEm82i7@h^lzm-svTN&5QVMjavPf?uyr
zlJr^8DW%Z%)dl_)=$u@!?gmM}ks-Z9($z2a)WDh>63==VZMWy*pCmooU{Etj*G_o&
z3EED-G4==Cn^$EWj_cWQIZLnyM#s%5o_Yf7->vPH0_*(MS5u*F$=0EZV02er{c@6?
z7dt7NEMH;ZwHD)n(LX0vw2v1)F1-4^H=Z8{3qpD1Im(6c^?Kwv3R5wOhNatGzKb9w
zUbVO*NniX{Xiw6=25)sH%e6GVdU&4VJ#BlHG2|OD)bG%D-r5y+VMg;O%UamaGH%~f
zSlKv=eF@oqea^py5tmyAet~&xL`^Hq3>q+A+Z%CrO)O;qF9y{uF@-6qvpRNwZMOsk
zQdl)UZuFwQ=G6BocN&&Mu3j}Y8}88A+WiP@m_e1CguWZx-mL78=RX$<=jTA9SV{fb
z0CQ^G6{m@LFi(BBB_Hy)tV`bokNPIt7s14zS9TwTW$#ZqohIp{J07_JWrJT2yb5Ow
z5Zt&4QwD@jybHhdsCfPm8Xej@{V9yFeoDP0>86i1zlE;eGh|<&=+6O1T4BcTFI}|<
z;C>lWU8@gEPdte?g|hs4Pc2|Z&Xi~=%vqK5vop->={>(2q?*ee
zHROAi?Dd7#yIO_}AnDDrn}cD2aK@-%uw-56+bC$2Ax#(qgP!S|C&0M)mjy|1cE%6U
zENFk~+@AUHagtEB7`i;}e|iPvFYoWQ7V0kF@Mt4UY?>3F4_ABaMJSLuxLhttuLItzw5RDn4^+84~EfIr^>^j%xJ4R3Tnw3evgCY
z!-vF9gx+la^Qo|@`OEYa7$Otd&4E~-;
z)7Xc@>*2V7L3jRy!%k%GAGiehb@qqbP}t1xT@?ZMo?R0b1M`N7>&KGx1?l4_Lrr1%
z_hcx#Zn$C&Jhvd$dI2nzHf+v>E@Pa%R>N-J7hGLO=GXE2Z-$oLVrvWFou|iQ_rTOm
zMjwmeu!^lo$6>Y6Df6@Ng^}*=OEB#7_RiN~Z9m1{Dwy+Ql*fHIwA1c8kKntM5&arq
zQu(F(O|aT5!)f9!b1HWD#iLW9(?WUQ6xdQY<3TE{zNbHI5ft{D*svTL-QO}J8#+8P
z)7l8@){V*B3R^ZkFfW8>F83|m4_nzi%n|s#Li@}~_$k6Is1$B7U-h^g<``bQ_DpYC1qn0krlEO~juR==
z^TQv(SaoK}pBRIEk&^IA$@o()i?>wGg=fs^@
z2vfSA_g)6;bWCopfwnUfBi2KK?yfJt!R%2ZlXt+`&1+5fL3Y=lxrboLk#tE3bm3ku
zeBg-wOI)bY6G(k4UiTbY>jYT7hU#Av_kDm1W_)t}24$+Gu)45VGJ|R7Fz0$Ev)K-_Egr40r#Qx4!6sXpk_zfpaxhIS@W_9CLN5A
ze-FP0J2`Q+0HCUHsE9AW9s
zJwsjKgedAc16?zx%=Cc6Ry}F)f?9pqrT$QTIMHSxe4W_5a|j$dP_2l7-Dd7R9S!r!
z#Gzy1-rrvj{Y?Mj1DRv&vUfkDS+mal{8PUzvdPY3N**YKR)Mv$8G$4wzjo?E*#9F}|AkL!r~
z5V!pJM{D@_S-^NH%;?wUB=N&X^RYJ2CMTx1WDJC!R?faoQF(
zuiV(x9v(b+<}mR?Sk!xB-XBz_&d@*hdLLr6`@zw~#e=3@B-%Spc6NZRgO*e|!MQ&Y
zle@q>Q*S*a=7*?vJ43#Ay|62!ZobVX8mU#?yTPa2nK#|x@zj8KJ&3hd+g;(}p0nN(
z3qQW-)DzDBbBiDG`|{YRp(UjX}CiG3ui3e@`RZC
z*k_3wEb6x2+Z{Fz?);kvv|RBsl7o4j`dRgYzV$YHJYh6nLG^|=J2tVz**lJmAc{V(
zXEc!Ob*YJ{=HATc15-~|B>2Lj?uPG)nmHf!`@(qj@^a#`veF^_VcwN7Zv)_FL(@i$J(wU3w5v_ef+C@q<&~g28Y{Zhl}0?5D}7CDweaFA0N5Ezygp#%tq5p)d-c%~aK|YfQ6!uZk-eU{EHdVN6r?xpjgN*}CBM}Wm+j=_G0&9biHEx8;>E