Skip to content

MySQL and MariaDB migrations are read, so every rule works on a MySQL source - #16

Merged
avison9 merged 3 commits into
mainfrom
feat/mysql-source
Sep 25, 2026
Merged

avison9 merged 3 commits into
mainfrom
feat/mysql-source

Conversation

@avison9

@avison9 avison9 commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

What it changes

cdclint reads MySQL and MariaDB migrations, so every existing rule works on a Debezium MySQL source. MySQL has no schemas: Debezium names a table database.table and a column database.table.column, so the new reader puts the database where the Postgres reader puts the schema, and the rules, include lists and topic names work unchanged.

  • Which reader. The connector's class picks it (MySqlConnector, MariaDbConnector); --migrations mysql:DIR or postgres:DIR says it outright. No new flag.
  • Which database. From a USE db; in the migrations, or from a connector whose database.include.list names one database (or whose table.include.list entries all start with the same one). When neither says, cdclint stops and says how to tell it; guessing would name the wrong table in every finding.
  • Database lists. database.include.list and database.exclude.list (and the old whitelist/blacklist spellings) are applied before the table lists, as Debezium does. sink-table-not-captured now names the list that actually keeps the table out, with the right entry: add billing to database.include.list, not billing.invoices.
  • The reader. CREATE TABLE (backticks, ENGINE options, KEY/INDEX/FULLTEXT/SPATIAL lines, LIKE, keywords glued to their column list), ALTER TABLE (ADD with FIRST/AFTER and several columns in parentheses, CHANGE, MODIFY, DROP, RENAME COLUMN, RENAME TO another database), RENAME TABLE, DROP TABLE, USE. Replica identity is left empty for MySQL tables (a Postgres setting).
  • The splitter's MySQL mode. # comments, backslash escapes in strings, no dollar quoting ($ is an identifier character), and DELIMITER for routine bodies. Postgres splitting is unchanged.
  • Prepared-statement DDL. MySQL has no ADD COLUMN IF NOT EXISTS, so re-runnable migrations write SET @s = (SELECT IF(<exists>, 'SELECT 1', 'ALTER TABLE t ADD c INT')); PREPARE ...; EXECUTE .... The reader applies the table statements in those strings. The IF only skips DDL whose effect is already there, and re-adding or re-dropping is a no-op in the reader, so this is the end state the migration guarantees.

Proven against a real repository

Mattermost v10.11.0 keeps the same schema as 140 MySQL and 140 Postgres migrations (golang-migrate .up.sql). A throwaway comparison, not committed, read both:

before reading prepared-statement strings after
tables found (MySQL / Postgres) 71 / 71 71 / 71
column differences of 610 97 12

All 12 remaining differences are on the Postgres reader's side, not this one: 9 are a glued UNIQUE(...) constraint read as a column, and 3 are columns Mattermost adds inside a PL/pgSQL DO $$ block. Both are existing Postgres-reader limitations, listed below as follow-ups.

The same run found a crash in the Postgres reader (a table name glued to a multi-line column list), fixed separately in #15; this reader has its own copy of the fix and a test for it.

Verified

  • Unit tests: internal/source/mysql (every construct above in one migration set, glued multi-line names, the missing-database error), internal/sqlsplit (MySQL mode: #, backslash, versioned comment, DELIMITER $$ trigger, $ in a name). The existing Postgres split test is unchanged.
  • Corpus: mysql-column-never-captured, mysql-change-column, mysql-database-not-included, mysql-diff-adds-column (the diff rule through the MySQL reader and the --base path). Honest note: this time the reader was written, with its unit tests, before these corpus entries; the expectations were still written by hand before running them, and they passed as written. Writing mysql-database-not-included's expectation is what exposed the wrong billing.invoices fix.
  • Every existing corpus entry unchanged; gofmt -l . clean, go vet, go test ./... pass.
  • RefuseRadar: 0 error(s), 0 warning(s), 159 info, unchanged, both detected and with postgres: given explicitly.

What was rejected

  • Guessing the database when neither the migrations nor the connector name one.
  • Creating a table for CREATE TABLE ... AS SELECT without a column list: its columns are the query's; left out rather than invented.
  • Treating MySQL generated columns differently. Whether Debezium emits them is not something I verified, so they are read as ordinary columns.

Follow-ups found on the way (not in this PR)

  1. Postgres: a glued UNIQUE(a, b) in CREATE TABLE is read as a column (9 phantom columns on Mattermost).
  2. Postgres: DDL inside DO $$ ... $$ blocks is not read (the Postgres twin of the prepared-statement idiom).
  3. Both readers apply every *.sql, including golang-migrate's .down.sql, which undoes each migration. Any golang-migrate repository reads wrong today.

Down migrations

Rebased after #17 merged: the MySQL reader leaves out down migrations through the same internal/migrate calls as the Postgres reader (migrate.Down in ReadFiles, migrate.Up before splitting), with TestDownMigrationsAreLeftOut (Mattermost's 000092 down file, and a goose down section).

Merging

Stacked on #14, which is stacked on #13, all rebased onto main after #15 and #17: until those merge, this PR's diff also shows their commits. The code conflicts with #13 and #14 in debezium.go and engine.go were resolved by hand in this rebase (both sides kept; #13's updated PatternLister comment kept). Merge #13, then #14, then this; each goes in without conflicts. On the three together: go test ./... passes, 33 corpus entries, RefuseRadar 0 error(s), 0 warning(s), 159 info.

…corrected entry when the mistake is a missing schema or a shell glob

Debezium logs a warning when a table.include.list entry matches nothing
and keeps running, so the topic is simply never produced. People ask it
to fail instead (debezium/dbz#872); a maintainer answered that a table
may be created later. cdclint raises it as a warning for the same
reason: when a sink actually reads the table, sink-table-not-captured
already fails the pull request.

The two shapes behind the most-viewed questions get a pointed fix,
because both follow from Debezium's documented matching (each entry is
a regular expression matched against the whole schema.table name):
an entry without its schema (Stack Overflow 74103659: ipaddrs for
myschema.ipaddrs) gets "write myschema\.ipaddrs", and a shell glob
(Stack Overflow 51345636: public.bg_* names no table as a regex) gets
"write public\.bg_.* to capture ..." with the tables it would capture.
An entry that already contains .* is read as the regex it is.

Exclude-list entries are not checked: one that matches nothing excludes
nothing, the same reasoning as captured-column-missing.

Three corpus entries, written before the code: include-table-no-schema,
include-table-glob, include-table-typo.
… connectors' transforms, and a table that misses a flatten is raised

A Flatten transform on the Debezium envelope names every field
after.<column>, and sinks match fields to columns by name. A ClickHouse
table that names its columns as Postgres does therefore receives a row
of zeros and empty strings for every change, with no error
(ClickHouse/clickhouse-kafka-connect discussion 182). cdclint said "ok"
on that exact configuration, because it took every sink column name as
a row column name. The same assumption made it raise twelve false
sink-column-unknown warnings on ClickHouse's own CDC example, whose
table names its columns before.<column> and after.<column>, and miss
that one of those columns was never captured.

A new internal/smt package works out the value's shape from each
connector's transforms, in order: the envelope for a Debezium connector
(or after.state.only false), the row after ExtractNewRecordState or the
Iceberg DebeziumTransform, and the flattened envelope after Flatten,
with its delimiter. The sink connector's transforms apply after the
source's. When the result is a flattened envelope, after.<c> and
before.<c> resolve to source column c, so the capture rules see through
them; op, ts_ms and source/transaction fields are envelope metadata; and
a column named as the row names it is raised by the new
sink-column-flattened rule, once per table.

Only documented transforms are modelled. Routers, key transforms and
Filter change no field name; any other value transform, or a modelled
one under a predicate, makes the shape unknown, and an unknown shape
keeps the old behaviour and produces no new finding.

Rejected: raising a bare envelope with no transform at all. Whether a
sink handles Debezium's envelope itself, or a converter reshapes it, is
not in the files, so that would be a guess.

Corpus first: flatten-bare-columns (discussion 182's configs),
flatten-envelope-columns (ClickHouse/examples cdc/postgresql at ae417ae,
with an include list that leaves postcode2 out), flatten-after-unwrap
(a guard). RefuseRadar, which unwraps, is unchanged at 0/0/159.
… source

Debezium's MySQL connector is where the most include-list questions come
from, and cdclint could only read Postgres. MySQL has no schemas:
Debezium names a table database.table and a column
database.table.column, so a new reader puts the database where the
Postgres reader puts the schema and the rules, include lists and topic
names work unchanged.

The connector's class picks the reader (MySqlConnector,
MariaDbConnector); --migrations mysql:DIR or postgres:DIR says it
outright. The database unqualified migrations run in comes from a USE
statement, or from a connector whose database.include.list names one
database (or whose table.include.list entries all start with the same
one); when neither says, cdclint stops and says how to tell it, since
guessing would name the wrong table in every finding.

database.include.list and database.exclude.list are applied before the
table lists, as Debezium does, and sink-table-not-captured names the
list that actually keeps a table out, with the entry that belongs in
it: "add billing to database.include.list", not billing.invoices.

The reader covers CREATE TABLE (backticks, ENGINE options, KEY, INDEX,
FULLTEXT and SPATIAL lines, LIKE, keywords glued to their column list),
ALTER TABLE (ADD with FIRST or AFTER and several columns in
parentheses, CHANGE, MODIFY, DROP, RENAME COLUMN, RENAME TO another
database), RENAME TABLE, DROP TABLE and USE. The splitter gains a MySQL
mode: # comments, backslash escapes, no dollar quoting, and DELIMITER.
MySQL has no ADD COLUMN IF NOT EXISTS, so re-runnable migrations put the
DDL in a string run through PREPARE; the reader applies the table
statements in those strings, which is the end state the migration
guarantees.

Proven against a real repository: Mattermost v10.11.0 keeps the same
schema as 140 MySQL and 140 Postgres migrations. Both readers find the
same 71 tables; of 610 columns, every difference left is on the
Postgres reader's side (a glued UNIQUE( read as a column, and columns
added inside DO blocks). Before reading the prepared-statement strings
the MySQL reader missed about 90 columns there.

Down migrations are left out here as in the Postgres reader (#17):
golang-migrate *.down.sql files are skipped and goose, sql-migrate and
dbmate down sections blanked, through the same internal/migrate calls.

Corpus: mysql-column-never-captured, mysql-change-column,
mysql-database-not-included, mysql-diff-adds-column. RefuseRadar is
unchanged at 0/0/159.
@avison9
avison9 merged commit bab645d into main Sep 25, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant