Skip to content

Drop emails.source_account and emails.import_id columns (post-deprecation) #43

Description

@splaice

Problem

Since PR #39 / issue #36, the authoritative per-account attribution lives in the `email_accounts` join table. The scalar `emails.source_account` and `emails.import_id` columns are kept as read-only "first seen" markers, populated on initial insert but never used for querying (all `account=` filters go through `EXISTS(... email_accounts ...)`).

These columns have real downsides:

  • They imply a 1:1 account-per-email relationship that the DSL accidentally exposes (see issue DSL: expose true multi-account attribution (not just first-seen) #40), leading to silent wrong-answer scenarios.
  • They add ~50 bytes/row + B-tree index storage on every row, for no query benefit.
  • They make the `emails` schema look more account-aware than it actually is, which is confusing for contributors.
  • They exist only for a self-tightening `NOT NULL` invariant that serves no purpose once the column is gone.

The #36 issue explicitly said these would be removed "one release after" the join table lands. This issue is that follow-up.

Current state

  • `src/maildb/schema_tables.sql`: `emails.source_account TEXT`, `emails.import_id UUID REFERENCES imports(id)`.
  • `src/maildb/db.py::init_db`: self-tightens `source_account` to `NOT NULL` once all rows are tagged; backfills `email_accounts` from these columns.
  • `src/maildb/models.py::Email`: `source_account: str | None`, `import_id: UUID | None` dataclass fields.
  • `src/maildb/maildb.py::SELECT_COLS`: lists both columns.
  • `src/maildb/ingest/parse.py::INSERT_EMAIL_SQL`: writes both columns on initial insert.
  • `src/maildb/ingest/orchestrator.py::backfill_source_account`: writes to both columns then mirrors to `email_accounts`.
  • `src/maildb/dsl.py::_EMAILS_COLUMNS`: exposes both columns.
  • `src/maildb/schema_indexes.sql`: `idx_email_source_account`, `idx_email_import_id`.

Proposed solution

A single migration release that drops the columns. Prerequisite: issue #40 (DSL multi-account source) must ship first, so that DSL users have an alternative.

Steps

  1. Confirm no external consumers read `emails.source_account` or `emails.import_id` directly (grep the codebase; check any stored notebooks/scripts).
  2. Remove the columns from the Email dataclass. Deserialization in `Email.from_row` already uses `row.get(...)` with defaults, so it tolerates absent columns automatically.
  3. Update `SELECT_COLS` in `maildb.py` and the aliased version inside `unreplied()`.
  4. Update `INSERT_EMAIL_SQL` in `parse.py` to drop the two columns from the INSERT list.
  5. Update `backfill_source_account` to stop writing to them — it becomes simpler: insert an `imports` row, then populate `email_accounts` directly for every existing `emails` row that has no join entry yet.
  6. Remove both columns from `_EMAILS_COLUMNS` in the DSL.
  7. Drop the indexes: `idx_email_source_account`, `idx_email_import_id`.
  8. Remove `ALTER TABLE emails ADD COLUMN IF NOT EXISTS source_account`, `... import_id` lines from `schema_tables.sql`.
  9. Remove the NOT NULL self-tightening block from `init_db`.
  10. Schema migration: `ALTER TABLE emails DROP COLUMN source_account`, `ALTER TABLE emails DROP COLUMN import_id`. Do this in `init_db` as an idempotent `ALTER TABLE ... DROP COLUMN IF EXISTS` to smooth upgrades from the transitional release.
  11. Delete the two tests that specifically cover the scalar columns (`test_emails_has_source_account_and_import_id`, `test_init_db_tightens_source_account_when_no_nulls`, `test_init_db_leaves_nullable_when_some_nulls`).
  12. Update DESIGN.md §4.1 (emails schema) and §4.2 (remove the "first seen" design decision).

Acceptance criteria

  • Issue DSL: expose true multi-account attribution (not just first-seen) #40 shipped (DSL virtual source for true attribution)
  • Columns dropped from schema; `init_db` runs the DROP idempotently
  • `Email` dataclass no longer has `source_account` / `import_id` fields
  • No code path reads `emails.source_account` or `emails.import_id`
  • All queries against `account=` still work (via `email_accounts`)
  • Integration test: upgrade path — fresh DB created against previous schema (with columns) runs `init_db` from the new code and ends up with the columns dropped and `email_accounts` fully populated
  • `maildb ingest migrate` still works for DBs that predate the join table

Tradeoffs

  • Pro: Simpler schema. One place to look for account attribution. No more "first seen vs true" confusion.
  • Con: Migration is destructive — rolling back after the drop requires a restore from backup. Mitigated by shipping this as its own release with a clear upgrade note.

Do NOT do in this issue

  • Don't rename `email_accounts` or reshape the join table. This issue is strictly subtractive.
  • Don't try to preserve a "first seen" concept — it's not load-bearing; `email_accounts.first_seen_at` covers the use case if ever needed.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions