Skip to content

Commit 15bf186

Browse files
fix(driver-sql): aggregate count / count_distinct / sum / avg answer numbers on PostgreSQL and MySQL (#20372)
Fixes #20335 Clause-②: no ## What changed `SqlDriver.aggregate` handed the SQL client's answer straight through. node-postgres parses `bigint` (OID 20) and `numeric` (OID 1700) to strings, and mysql2 does the same for `DECIMAL`. So `count` / `count_distinct` / `sum` / `avg` reached the engine's `having` and the REST response as strings on the native path of PostgreSQL (all four) and MySQL (`sum` / `avg`), while SQLite and the rows path of every dialect answered numbers. The fix is one presentation step at the end of the native path, in `packages/drivers/driver-sql/src/sql-driver.ts`: - `AGGREGATE_ANSWER_KIND` is a `Record` over the spec's `AggregationFunction`. `count`, `count_distinct`, `sum` and `avg` are `'number'`; `min` and `max` are `'column'`. A function added to the enum without an entry fails `tsc`. - In `aggregate()`, an aliased aggregation whose function is `'number'` registers its alias with the existing `'number'` presenter (`presentReadValue`, the one `formatOutput` applies to a numeric field on a `find()` row). `min` / `max` keep the column's own presentation, unchanged. - It is keyed on the function the query asked for, never on what a value looks like. It is not gated by dialect: the presenter rewrites only a string, so SQLite's answers (and mysql2's `COUNT`) pass through untouched. - No connection-level type parser is touched. `find()`, `distinct()` and the `where` comparand path are unchanged. `packages/objectql` is untouched. ## Precision policy (H3) The answer is one JS number, an IEEE-754 double, on every dialect. A `sum` / `avg` whose exact value needs more than a double's 15 to 17 significant digits, or an integer total at or above 2^53, is rounded to the nearest double. The policy is stated on `AGGREGATE_ANSWER_KIND` and in the changeset. Options weighed: - A. A JS number, with the loss beyond double precision undeclared. Rejected, because the loss must be written down. - B. A string only when the value is unsafe as a double, a number otherwise. Rejected. The value's type would depend on its size, so the defect would come back for large totals only: a `having` `$in`, or a chart reading the column, would silently stop matching for exactly those groups. That is also the harder shape for an AI author to get right. - C. **Chosen:** always a number, with the loss declared. It is the same bound `find()` already puts on a read of the same exact-decimal column, since `formatOutput`'s numeric pass (`valueSchemaFor` gives the numeric class `z.number().finite()`, ADR-0104 D1). It is also the bound the rows path has always had: `in-memory-aggregation.ts` sums JS doubles. What the spec says an aggregate's value type is: `packages/spec/src/data/aggregation-conformance.ts` types `AggregationExpectation.value` as `number` and states it "stays a `number` for every case", across every enrolled face. `service-analytics`' `measure-result-type.ts` publishes `number` as the result type of `count` / `count_distinct` / `sum` / `avg` measures. Before this change, PostgreSQL delivered strings under that declaration. ## Measured (H1, H2) Base `26daf0b036` and head `f3b9e1390c`, on SQLite (better-sqlite3), live PostgreSQL 16.13 (server zone Asia/Shanghai) and live MySQL 8.0.46 (`+08:00`). Each cell went through `SqlDriver.aggregate`, `engine.aggregate` and `POST /api/v1/data/:object/query` (JSON round-trip). The fixture is `groupBy` customer, four groups, over `number` fields (one integer-valued, one decimal-valued), `currency`, `percent` and `rating`. The three native doors answered identically at base and at head, on every dialect. | face | base | head | |:--|:--|:--| | PostgreSQL native: `count`, `count_distinct` | string `"2"` | number `2` | | PostgreSQL native: `sum` / `avg` over number, currency, percent | string `"500.000…"` (30-digit scale) | number `500` | | PostgreSQL native: `sum` / `avg` over rating (int4) | string `"7"` / `"3.5000000000000000"` | number `7` / `3.5` | | MySQL native: `count`, `count_distinct` | number | number, unchanged | | MySQL native: `sum` / `avg` over number, currency, percent, rating | string (`DECIMAL`) | number | | `min` / `max` over any of the five, native, PG and MySQL | number | number, unchanged | | SQLite native, every cell | number | byte-identical | | rows path, every dialect, every cell | number | byte-identical | H2: a raw query through `pg` answers field OIDs `n:20 total:1700 mean:1700 mn:1700 mx:1700 ss:20 sa:1700 smin:23`. Every value is a string except `smin` (int4). The same query through `mysql2` answers column types `n:8` (LONGLONG, a number), `246` (NEWDECIMAL, a string) for `sum` / `avg` / `min` / `max` of the decimal column and for `sum` / `avg` of the int column, and `3` for `min` of the int column. So `min` / `max` over a `numeric` column is a string at the client too. It already left the driver as a number, because a declared numeric field takes the `'number'` column presentation since the numeric-representation change. The rows path yields numbers because `find()` presents the numeric column as a number and `in-memory-aggregation.ts` computes `count` as `rows.length` and `sum` / `avg` in JS arithmetic. The card's `having` table, engine and REST doors (`having` is evaluated by the engine on both paths): | `having` | base: PG native | base: MySQL native | head: every dialect × path | |:--|:--|:--|:--| | `{ n: { $in: [2] } }` | no group | c1, c2 | c1, c2 | | `{ total: { $in: [500, 20] } }` | no group | no group | c1, c4 | | `{ mean: { $in: [250, 600] } }` | no group | no group | c1, c2 | | `{ avg_rate: { $in: [0.375] } }` | no group | no group | c1 | | `{ n: { $lt: 'not-a-date' } }` | c1–c4 | no group | no group | | `{ total: { $lt: 'not-a-date' } }` | c1–c4 | c1–c4 | no group | | `{ n: { $gt: '+010000-01-01T00:00:00.000Z' } }` | c1–c4 | no group | no group | | `$eq` on count / sum / avg, `$gt` / `$gte` numbers | as elsewhere | as elsewhere | unchanged | At base, SQLite (both paths) and the rows path of PostgreSQL and MySQL already answered the head column. ## Collateral (H4) - `find()` and `distinct()` over the five numeric columns are byte-identical base to head, on all three dialects (compared as typed JSON). - SQLite: every aggregate cell is byte-identical base to head, on every door. - No connection-level parser changed. The temporal `min` / `max` presentation is untouched: the `'column'` arm is the pre-existing branch. `sql-driver-aggregate-temporal-output.test.ts` and `sql-driver-13973-canonical-iso-read-door.test.ts` pass in the live suite below. ## Consumers (H5) These read the changed values and now receive a number: - `@objectstack/objectql` `aggregateSummaryValue`: roll-up summaries write the value into the parent's `summary` field. That value is now a number where PostgreSQL (`count`, `sum`, `avg`) and MySQL (`sum`, `avg`) handed back a numeric string. `summary-backfill.ts` counts `nonEmpty` for a `count` / `sum` roll-up only when `typeof computed === 'number'`, so by reading, its report under-counted those roll-ups on PostgreSQL (`count`, `sum`) and MySQL (`sum`) before; it counts them now. - `service-analytics` `ObjectQLStrategy`: passes values through into responses that declare `number`, and now delivers one. `cross-object-rebucket.ts` and `dataset-executor.ts` coerce with `Number(...)`, which is the identity on a number. - `metadata-protocol` (the REST query route), `runtime` `action-execution` `aggregate`, `mcp` `stdio-data-bridge`, and the client SDK: pass-through, no coercion. - objectui, read by code search: `ObjectChart` `readValue`, `ObjectMetricWidget` and `MetricWidget` read these values through `Number(...)`, which is the identity on a number. None of the consumers above compares these values as strings or branches on `typeof value === 'string'`; the one `typeof` test (`summary-backfill.ts`) asks for `'number'`. ## Tests - New: `packages/drivers/driver-sql/src/sql-driver-20335-aggregate-numeric-presentation.test.ts`. Seven cases per dialect cell of the live matrix (`declareDialectCell`, so the PostgreSQL and MySQL cells run in `Temporal Conformance (live PG + MySQL)`): - `count` / `count_distinct` / `sum` / `avg` over number, currency, percent and rating, and `sum` / `avg` over a boolean, asserted with `toBe` against the rows' own JS arithmetic; - the all-NULL fold; - `min` / `max` presentation unchanged; - the precision policy: SQL-written `9007199254740993` and `12345678901234567.123456789` answer `Number(literal)`, never a string, and equal what `find()` reads for the same row. - New: `packages/rest/src/rest-aggregate-numeric-having.test.ts`. Engine and REST doors, native and rows paths: `having` `$in` / `$eq` on `count` / `count_distinct` / `sum` / `avg`, the string-comparand rows of the card, and the response's JSON numbers. The SQLite cell always runs. The PostgreSQL and MySQL cells run when `OS_TEST_POSTGRES_URL` / `OS_TEST_MYSQL_URL` are set and are otherwise a named skip. - Reverse verification: the fix was committed first, then the one `presentedOutput.set(agg.alias, 'number')` line was deleted through `scripts/ablation-replace.mjs` (anchor 1 → 0, blob `8dc157d535` → `c247bcff26`). driver-sql was rebuilt, and `ablation-dist-preflight --absent` confirmed the marker absent from all six built files. Result: - driver-sql test: 7 failed, 14 passed. The failures are every PG and MySQL string cell (for example `c1 count(*): expected '2' to be 2`). Every SQLite case and MySQL `count` stayed green. - REST test: 13 failed, 23 passed. The failures are the PG and MySQL cells the base table marks. SQLite stayed green. - The ablated native answers were byte-identical to the base on all three dialects. - Restore: blob equals HEAD, `git diff HEAD` empty. After a rebuild the marker is present in 2 built files, and `git status --porcelain` is empty. - Whole suites, at `06349dc973` (identical code to `f3b9e1390c`, which only corrects a docblock). PostgreSQL and MySQL were live, with `TZ=America/New_York` and `OS_EXPECT_LIVE_DIALECT_MATRIX=1`. - `@objectstack/driver-sql`: 204 files passed, 4690 tests passed, 1 skipped. The reporter confirmed that all 3 dialects were exercised. - `@objectstack/rest` `local` project, with no live URL, as in CI: 210 files passed, 3817 tests passed, 26 skipped (24 of them this file's PostgreSQL and MySQL cells). - `@objectstack/objectql` aggregate, `having`, in-memory aggregation and summary files: 19 files, 612 tests passed. - Typecheck: `@objectstack/driver-sql` exit 0 and `@objectstack/rest` exit 0 (including `check:test-typecheck`). Both new files are in a typecheck program (`--listFiles`). ## Gates At head `f3b9e1390c`, after a full workspace build (`turbo run build` over `./packages/*` and `./packages/*/*`, 71 of 71 tasks): - `node scripts/pm/dispatch-gates.mjs --commands --repo objectstack-ai/objectstack` derived 63 commands, the same 63 as the dispatch list. All 63 ran and exited 0. `--ran` reconciliation: 63 derived, 63 run, 0 NOT-MEASURED, 0 UNRUN. - `origin/main` moved during the run, so the three `--base` gates (`check-adr-0087-registration`, `check-changeset-no-major`, `check-empty-changeset`) and `check-issue-citations` were also run with `--base 26daf0b` (the merge base). Each exited 0. `check-changeset-no-major` reads the level from the PR, so its local run has no PR payload to judge. - Lint, a proved narrowing: `eslint --no-inline-config --format json` over the three touched TypeScript files reports 3 files, 0 errors, 0 warnings. `eslint --print-config` resolves a config for each, so none is ignored. The config enables no type-aware linting (`parserOptions` is `ecmaVersion` / `sourceType` only, with no `project`), so this diff cannot move the verdict on an untouched file. The repo-wide `pnpm lint` is CI's. - NOT MEASURED locally, and owned by CI: the type-check lanes over the whole workspace, `Test Core` shards, `Dogfood`, `Build Core`, and `Temporal Conformance` as CI spells it. The driver-sql suite was run locally against live PostgreSQL and MySQL as above. ## Acceptance notes - The rows path and the exact-decimal native path can differ in the last place of a double. A `number` column holding 0.1 and 0.2 sums to `0.3` on PostgreSQL and MySQL native (exact `numeric` / `DECIMAL` arithmetic) and to `0.30000000000000004` on the rows path and on SQLite (double arithmetic). So `having { s: { $eq: 0.3 } }` keeps that group on PG / MySQL native and on no other face. Measured identical at base and at head: this PR moves the type, not the arithmetic. Reported to the seat as a separate finding. - `readPresentationKind`'s `'number'` kind (used by `min` / `max` and `distinct()`) reads `numericFields`, which includes the driver-internal `integer` / `int` / `float` aliases, on every dialect. `formatOutput` narrows its row pass to `numericValueFields` on PostgreSQL and MySQL, to keep an introspected external `bigint` above 2^53 as a string. Not measured, not changed here. - The REST test's PostgreSQL and MySQL cells are provisioned by no CI job today; the live servers are attached to `driver-sql`'s suite. CI runs their SQLite cell. The PostgreSQL and MySQL value pins run in CI through the driver-sql file. Carrier: none. --- _Generated by [Claude Code](https://claude.ai/code/session_01Bvd69VPa6puiNzzPUroDBx)_ --------- Co-authored-by: Claude <noreply@anthropic.com>
1 parent 17e4f52 commit 15bf186

4 files changed

Lines changed: 529 additions & 5 deletions

File tree

Lines changed: 41 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,41 @@
1+
---
2+
'@objectstack/driver-sql': patch
3+
---
4+
5+
fix(driver-sql): `count` / `count_distinct` / `sum` / `avg` answer JS numbers on PostgreSQL and MySQL, as they do on SQLite and on the engine's rows path
6+
7+
Clause-②: no
8+
9+
`SqlDriver.aggregate` handed the SQL client's answer straight through. node-postgres parses
10+
`bigint` (`count`, and `sum` over an integer column) and `numeric` (`sum` / `avg` over the
11+
numeric family's exact-decimal column, `avg` over an integer column) to strings, and mysql2
12+
does the same for `DECIMAL` (`SUM` / `AVG`). So one grouped query answered
13+
14+
{ "n": "2", "total": "500.000000000000000000000000000000" }
15+
16+
on PostgreSQL's native path and `{ "n": 2, "total": 500 }` on SQLite and on the rows path of
17+
every dialect. The engine's `having` compares values as they
18+
arrive, so `having { n: { $in: [2] } }` kept no group on PostgreSQL alone, and
19+
`having { total: { $in: [500, 20] } }` kept no group on PostgreSQL or MySQL, while a string
20+
comparand such as `{ total: { $lt: 'not-a-date' } }` kept every group there and none anywhere
21+
else.
22+
23+
Those four functions now answer a JS number on every dialect, through `SqlDriver.aggregate`,
24+
`engine.aggregate` and `POST /api/v1/data/:object/query`. The presentation is keyed on the
25+
aggregate function the query asked for; it only rewrites a string, so SQLite's answers are
26+
byte-identical to before. `min` / `max` are unchanged: they answer a value of the column and
27+
keep that column's presentation (a declared numeric field was already a number).
28+
Non-aggregate reads (`find()`, `distinct()`) are unchanged, and no connection-level type
29+
parser is touched.
30+
31+
**Precision policy.** The answer is one JS number (an IEEE-754 double) on every dialect. A
32+
`sum` / `avg` whose exact value needs more than a double's 15 to 17 significant digits, or an
33+
integer total at or above 2^53, is rounded to the nearest double. That is the same bound
34+
`find()` already puts on a read of the same exact-decimal column, and the bound the rows path
35+
has always had. A total that fits keeps its exact value (`500`, `30.75`, `0.375`). The answer
36+
is never a string, including for large totals: an answer whose type depended on its size would
37+
break the same `having` or chart for exactly those totals.
38+
39+
A consumer that read these values through `Number(...)` gets the same number it computed
40+
before. A consumer that compared them as strings, or checked `typeof value === 'string'`,
41+
now receives a number.
Lines changed: 208 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,208 @@
1+
// Copyright (c) 2026 ObjectStack. Licensed under the Apache-2.0 license.
2+
3+
/**
4+
* [#20335] `count` / `count_distinct` / `sum` / `avg` answer a JS NUMBER from
5+
* `SqlDriver.aggregate`, on every dialect — the value the engine's rows path
6+
* and SQLite already answered.
7+
*
8+
* Measured on the base (`26daf0b036`) through this door, `engine.aggregate` and
9+
* `POST /api/v1/data/:object/query`, `groupBy` customer, four groups:
10+
*
11+
* | dialect | `count` | `count_distinct` | `sum` / `avg` over number, currency, percent | `sum` / `avg` over rating (integer column) |
12+
* |:--|:--|:--|:--|:--|
13+
* | SQLite | number | number | number | number |
14+
* | PostgreSQL 16.13 | `"2"` | `"2"` | `"500.000000000000000000000000000000"` | `"7"` / `"3.5000000000000000"` |
15+
* | MySQL 8.0.46 | number | number | `"500.000000000000000000000000000000"` | `"7"` / `"3.5000"` |
16+
*
17+
* node-postgres parses `bigint` (OID 20) and `numeric` (OID 1700) to strings,
18+
* and mysql2 does the same for `DECIMAL`; `min` / `max` over a declared numeric
19+
* field were already numbers (the column's own `'number'` presentation, #16318)
20+
* and are unchanged. The engine's `having` then compared `"2"` against `2`:
21+
* `having { n: { $in: [2] } }` kept no group on PostgreSQL's native path and
22+
* c1, c2 everywhere else (pinned at the engine and REST doors in
23+
* `@objectstack/rest`'s `rest-aggregate-numeric-having.test.ts`).
24+
*
25+
* The fixture's values are dyadic fractions on purpose, so a JS double holds
26+
* every sum and average EXACTLY: the expected values below are computed from the
27+
* rows with JS arithmetic — the rows path's own arithmetic — and asserted with
28+
* `toBe`, so a string, a boolean or a rounding difference each fail.
29+
*
30+
* The precision policy (one JS double, the loss beyond a double's precision
31+
* declared — `AGGREGATE_ANSWER_KIND` in `sql-driver.ts`) is pinned by the last
32+
* case of each cell: a total the exact-decimal column holds but a double cannot
33+
* answers the nearest double, which is also what `find()` reads for the row.
34+
*/
35+
36+
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
37+
import type { DriverQuery } from '@objectstack/spec/contracts';
38+
import { SqlDriver } from './sql-driver.js';
39+
import { DIALECT_CELLS, declareDialectCell, type DialectCell } from './live-dialect-matrix.testkit.js';
40+
41+
const TABLE = 'os20335_agg_numbers';
42+
43+
interface Row {
44+
id: string;
45+
customer_id: string;
46+
amount: number;
47+
price: number;
48+
rate: number;
49+
stars: number;
50+
flag: boolean;
51+
note: string;
52+
}
53+
54+
const ROWS: readonly Row[] = [
55+
{ id: 'o1', customer_id: 'c1', amount: 100, price: 10.25, rate: 0.25, stars: 3, flag: true, note: 'b' },
56+
{ id: 'o2', customer_id: 'c1', amount: 400, price: 20.5, rate: 0.5, stars: 4, flag: false, note: 'a' },
57+
{ id: 'o3', customer_id: 'c2', amount: 900, price: 30.75, rate: 0.75, stars: 5, flag: true, note: 'c' },
58+
{ id: 'o4', customer_id: 'c2', amount: 300, price: 40, rate: 0.125, stars: 2, flag: true, note: 'c' },
59+
{ id: 'o5', customer_id: 'c3', amount: 50, price: 5.5, rate: 0.5, stars: 1, flag: false, note: 'e' },
60+
{ id: 'o6', customer_id: 'c4', amount: 20, price: 1.25, rate: 0.25, stars: 5, flag: false, note: 'f' },
61+
];
62+
63+
/** number (integer-valued), currency, percent, and the one integer column of the family. */
64+
const MEASURED = ['amount', 'price', 'rate', 'stars'] as const;
65+
const GROUPS = ['c1', 'c2', 'c3', 'c4'] as const;
66+
67+
const byGroup = (g: string) => ROWS.filter((r) => r.customer_id === g);
68+
const sumOf = (rows: readonly Row[], f: keyof Row) => rows.reduce((a, r) => a + Number(r[f]), 0);
69+
70+
function grouped(): DriverQuery {
71+
const aggregations: Array<Record<string, unknown>> = [
72+
{ function: 'count', alias: 'n' },
73+
{ function: 'count_distinct', field: 'note', alias: 'nd' },
74+
{ function: 'sum', field: 'flag', alias: 'sum_flag' },
75+
{ function: 'avg', field: 'flag', alias: 'avg_flag' },
76+
{ function: 'sum', field: 'spare', alias: 'sum_spare' },
77+
{ function: 'avg', field: 'spare', alias: 'avg_spare' },
78+
{ function: 'min', field: 'note', alias: 'min_note' },
79+
];
80+
for (const f of MEASURED) {
81+
for (const fn of ['count', 'sum', 'avg', 'min', 'max']) aggregations.push({ function: fn, field: f, alias: `${fn}_${f}` });
82+
}
83+
return { groupBy: ['customer_id'], aggregations } as DriverQuery;
84+
}
85+
86+
function declareCell(cell: DialectCell): void {
87+
describe(`[#20335] driver-sql — aggregate counts and totals are numbers (${cell.label})`, () => {
88+
let driver: SqlDriver;
89+
let answers: Map<string, Record<string, unknown>>;
90+
91+
beforeAll(async () => {
92+
driver = new SqlDriver(cell.config());
93+
await driver.execute(`drop table if exists ${TABLE}`).catch(() => {});
94+
await driver.initObjects([
95+
{
96+
name: TABLE,
97+
fields: {
98+
customer_id: { type: 'text' },
99+
amount: { type: 'number' },
100+
price: { type: 'currency' },
101+
rate: { type: 'percent' },
102+
stars: { type: 'rating' },
103+
flag: { type: 'boolean' },
104+
note: { type: 'text' },
105+
// NULL in every row: `sum` folds to 0 (#15546) and `avg` stays null.
106+
spare: { type: 'number' },
107+
},
108+
},
109+
] as never);
110+
for (const row of ROWS) await driver.create(TABLE, { ...row }, { bypassTenantAudit: true });
111+
const rows = (await driver.aggregate(TABLE, grouped())) as Array<Record<string, unknown>>;
112+
answers = new Map(rows.map((r) => [String(r.customer_id), r]));
113+
});
114+
115+
afterAll(async () => {
116+
await driver?.execute(`drop table if exists ${TABLE}`).catch(() => {});
117+
await driver?.disconnect();
118+
});
119+
120+
it('answers the four groups', () => {
121+
expect([...answers.keys()].sort()).toEqual([...GROUPS]);
122+
});
123+
124+
it('count and count_distinct are numbers, equal to the rows', () => {
125+
for (const g of GROUPS) {
126+
const a = answers.get(g)!;
127+
const rows = byGroup(g);
128+
expect(a.n, `${g} count(*)`).toBe(rows.length);
129+
expect(a.nd, `${g} count_distinct(note)`).toBe(new Set(rows.map((r) => r.note)).size);
130+
for (const f of MEASURED) expect(a[`count_${f}`], `${g} count(${f})`).toBe(rows.length);
131+
}
132+
});
133+
134+
it('sum and avg over number, currency, percent and rating are numbers, equal to the rows', () => {
135+
for (const g of GROUPS) {
136+
const a = answers.get(g)!;
137+
const rows = byGroup(g);
138+
for (const f of MEASURED) {
139+
expect(a[`sum_${f}`], `${g} sum(${f})`).toBe(sumOf(rows, f));
140+
expect(a[`avg_${f}`], `${g} avg(${f})`).toBe(sumOf(rows, f) / rows.length);
141+
}
142+
}
143+
});
144+
145+
it('sum and avg over a boolean answer the #11152 numbers strictly', () => {
146+
for (const g of GROUPS) {
147+
const a = answers.get(g)!;
148+
const rows = byGroup(g);
149+
expect(a.sum_flag, `${g} sum(flag)`).toBe(sumOf(rows, 'flag'));
150+
expect(a.avg_flag, `${g} avg(flag)`).toBe(sumOf(rows, 'flag') / rows.length);
151+
}
152+
});
153+
154+
it('an all-NULL aggregand: sum folds to the number 0, avg stays null', () => {
155+
for (const g of GROUPS) {
156+
expect(answers.get(g)!.sum_spare, `${g} sum(spare)`).toBe(0);
157+
expect(answers.get(g)!.avg_spare, `${g} avg(spare)`).toBeNull();
158+
}
159+
});
160+
161+
it('min and max keep the column presentation — numbers for the numeric family, text for text', () => {
162+
for (const g of GROUPS) {
163+
const a = answers.get(g)!;
164+
const rows = byGroup(g);
165+
for (const f of MEASURED) {
166+
expect(a[`min_${f}`], `${g} min(${f})`).toBe(Math.min(...rows.map((r) => Number(r[f]))));
167+
expect(a[`max_${f}`], `${g} max(${f})`).toBe(Math.max(...rows.map((r) => Number(r[f]))));
168+
}
169+
expect(a.min_note, `${g} min(note)`).toBe([...rows.map((r) => r.note)].sort()[0]);
170+
}
171+
});
172+
173+
it('precision policy: a total a double cannot hold answers the nearest double, as find() does', async () => {
174+
// Written by SQL, not by the driver: a JS number could not carry these
175+
// values in the first place, which is the whole point of the case. The
176+
// literals are numeric, so the exact-decimal column stores them exactly
177+
// on PostgreSQL and MySQL (SQLite's REAL column rounds on write).
178+
const EXACT = ['9007199254740993', '12345678901234567.123456789'];
179+
for (const [i, literal] of EXACT.entries()) {
180+
await driver.execute(`insert into ${TABLE} (id, customer_id, amount) values ('p${i}', 'p${i}', ${literal})`);
181+
}
182+
const rows = (await driver.aggregate(TABLE, {
183+
where: { customer_id: { $in: ['p0', 'p1'] } },
184+
groupBy: ['customer_id'],
185+
aggregations: [
186+
{ function: 'sum', field: 'amount', alias: 'total' },
187+
{ function: 'avg', field: 'amount', alias: 'mean' },
188+
],
189+
} as DriverQuery)) as Array<Record<string, unknown>>;
190+
const found = (await driver.find(TABLE, { where: { customer_id: { $in: ['p0', 'p1'] } } })) as Array<
191+
Record<string, unknown>
192+
>;
193+
for (const [i, literal] of EXACT.entries()) {
194+
const row = rows.find((r) => r.customer_id === `p${i}`)!;
195+
expect(typeof row.total, `sum over ${literal} is a number, never a string`).toBe('number');
196+
expect(row.total, `sum over ${literal}`).toBe(Number(literal));
197+
expect(row.mean, `avg over ${literal}`).toBe(Number(literal));
198+
expect(row.total, `the same bound find() reads for ${literal}`).toBe(
199+
found.find((r) => r.customer_id === `p${i}`)!.amount,
200+
);
201+
}
202+
});
203+
});
204+
}
205+
206+
for (const cell of DIALECT_CELLS) {
207+
declareDialectCell(cell, 'aggregate numeric presentation (#20335)', declareCell);
208+
}

‎packages/drivers/driver-sql/src/sql-driver.ts‎

Lines changed: 80 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -1476,6 +1476,64 @@ const SQL_AGGREGATE_FUNCTIONS: ReadonlyMap<string, SqlAggregateLowering> = new M
14761476
['count_distinct', { sql: 'count', distinct: true }],
14771477
]);
14781478

1479+
/**
1480+
* [#20335] What each declared aggregate function ANSWERS — a derived `number`,
1481+
* or a value OF the aggregated column — and therefore which read presentation
1482+
* {@link SqlDriver.aggregate} gives its result column.
1483+
*
1484+
* - `'number'` — `count`, `count_distinct`, `sum`, `avg`. A count or a total is
1485+
* a number whatever the column held, and it is presented as one (`'number'`,
1486+
* the presenter `formatOutput` applies to a numeric field on a `find()` row).
1487+
* - `'column'` — `min`, `max`. The answer is one of the column's own values, so
1488+
* it takes that column's presentation ({@link SqlDriver.readPresentationKind}),
1489+
* exactly as before this table existed.
1490+
*
1491+
* Why the `'number'` half needs presenting at all: the SQL client hands a
1492+
* result back as the wire type of the SQL expression, not as the platform's
1493+
* value type. Measured on live PostgreSQL 16.13 and MySQL 8.0.46 through this
1494+
* driver's own connections: node-postgres parses `bigint` (OID 20 — `count`,
1495+
* and `sum` over an integer column) and `numeric` (OID 1700 — `sum` / `avg`
1496+
* over the exact-decimal numeric family, `avg` over an integer column) to
1497+
* STRINGS (`"2"`, `"500.000000000000000000000000000000"`), and mysql2 does the
1498+
* same for `DECIMAL` (`SUM` / `AVG`; its `COUNT` arrives as a number). The
1499+
* engine's rows path (`objectql`'s `in-memory-aggregation.ts`) and SQLite answer
1500+
* numbers for the same query, so `having { n: { $in: [2] } }` kept c1, c2 on
1501+
* those and no group on PostgreSQL's native path.
1502+
*
1503+
* Keyed on the function the query ASKED for, never on whether a value looks
1504+
* numeric, and deliberately not gated by dialect: the presenter only rewrites a
1505+
* STRING, so a client that already answers a number (better-sqlite3, mysql2's
1506+
* `COUNT`) passes through untouched — measured byte-identical on SQLite — and
1507+
* no list of "string-answering dialects" exists to drift.
1508+
*
1509+
* ## The precision policy — one JS number, the loss declared
1510+
*
1511+
* The answer is `Number(text)`: an IEEE-754 double, on every dialect. A `sum` /
1512+
* `avg` over the exact-decimal column (`numeric(65,30)` / `DECIMAL(65,30)`)
1513+
* whose value needs more than a double's ~15-17 significant digits, or an
1514+
* integer at or above 2^53, is ROUNDED to the nearest double — declared, not
1515+
* silent: it is the same bound `formatOutput` already puts on a `find()` read of
1516+
* that column (#16318, `valueSchemaFor`'s `z.number().finite()`, ADR-0104 D1),
1517+
* and the bound the rows path has always had (`toNumber` sums JS doubles). A
1518+
* value-dependent type — a number when it fits, a string when it does not — was
1519+
* rejected: it would reopen this defect for exactly the large totals, where a
1520+
* `having` `$in` or a chart silently stops matching. Only a string `Number()`
1521+
* reads as NaN (PostgreSQL's `numeric` `'NaN'`) is left as written, the
1522+
* presenter's existing rule.
1523+
*
1524+
* A `Record` over `AggregationFunction` on purpose: a function that joins the
1525+
* declared vocabulary without an answer here fails `tsc` rather than reaching a
1526+
* caller unpresented.
1527+
*/
1528+
const AGGREGATE_ANSWER_KIND: Readonly<Record<AggregationFunction, 'number' | 'column'>> = {
1529+
count: 'number',
1530+
count_distinct: 'number',
1531+
sum: 'number',
1532+
avg: 'number',
1533+
min: 'column',
1534+
max: 'column',
1535+
};
1536+
14791537
/**
14801538
* [#5907] The aggregate vocabulary the Query Protocol DECLARES, read from the
14811539
* spec rather than restated — `AggregationNodeSchema.function` is this enum, so
@@ -9858,8 +9916,10 @@ export class SqlDriver implements IDataDriver {
98589916
// GROUP BY bucket expression (#3773) and the result presentation (#3797).
98599917
const table = this.coercionKey(builder);
98609918

9861-
// Result columns that carry a raw column VALUE (rather than a count/total
9862-
// derived from one), keyed by the column name the caller will read.
9919+
// Result columns and the presentation each takes, keyed by the column name
9920+
// the caller will read: a raw column VALUE (a group key, `min`/`max`) takes
9921+
// its column's presentation, and [#20335] a count or total derived from one
9922+
// takes the `'number'` presentation (see `AGGREGATE_ANSWER_KIND`).
98639923
// Collected while the statement is built because that is the only point
98649924
// where a column name and its meaning are both known: a `min()` lands under
98659925
// its alias (never under the field name), and a date-BUCKETED column lands
@@ -10009,14 +10069,25 @@ export class SqlDriver implements IDataDriver {
1000910069
// their NULL passes through. See {@link foldEmptyAggregateAnswers}.
1001010070
const identity = emptyGroupValueFor(funcName);
1001110071
if (identity !== undefined) foldedOutput.set(agg.alias, identity);
10072+
// [#20335] A count or a total is presented as the number it is, on
10073+
// every dialect: node-postgres hands `bigint` / `numeric` back as a
10074+
// string and mysql2 `DECIMAL`, so without this the native path
10075+
// answered `"2"` where the rows path and SQLite answer `2`. Keyed on
10076+
// the function asked for; the precision policy (one JS double, the
10077+
// loss above a double's precision declared) is stated on
10078+
// `AGGREGATE_ANSWER_KIND`. The fold above runs first, so a folded
10079+
// `0` is already a number and passes through.
10080+
if (AGGREGATE_ANSWER_KIND[funcName] === 'number') {
10081+
presentedOutput.set(agg.alias, 'number');
10082+
}
1001210083
// `min`/`max` are the only supported functions that hand back a value
1001310084
// OF the column rather than a count/total derived from it, so they are
1001410085
// the only ones whose result still needs the column's presentation.
1001510086
// `alias` is required by `AggregationNodeSchema`; the unaliased branch
1001610087
// below lands under a dialect-dependent column name
1001710088
// (`max("closed_at")` on SQLite, `max` on Postgres) and is defensive
1001810089
// only, so it is deliberately not tracked.
10019-
if ((funcName === 'min' || funcName === 'max') && agg.field) {
10090+
if (AGGREGATE_ANSWER_KIND[funcName] === 'column' && agg.field) {
1002010091
// [#11152] A BOOLEAN aggregand is the ruled exception to "the
1002110092
// result still needs the column's presentation": the maintainer's
1002210093
// 2026-08-28 ruling (superseding #11249's `false`/`true`, which
@@ -15128,7 +15199,9 @@ export class SqlDriver implements IDataDriver {
1512815199
* the audit-stamp fold run everywhere (`datetime` and `audit_timestamp`
1512915200
* since #13973, [ADR-0053 D-F1] — the former SQLite-only and the latter
1513015201
* absent before, which handed the two live dialects' `Date` through), the
15131-
* numeric coercion is SQLite-only, and the boolean coercion runs on SQLite
15202+
* numeric coercion runs everywhere (#16318 for a declared numeric column;
15203+
* [#20335] for `aggregate()`'s counts and totals, `AGGREGATE_ANSWER_KIND`),
15204+
* and the boolean coercion runs on SQLite
1513215205
* and MySQL (#11782 — the two dialects whose stored boolean is a number).
1513315206
* {@link readPresentationKind} does the dialect gating for the scalar kinds,
1513415207
* so by the time one arrives here the dialect is settled.
@@ -15196,7 +15269,9 @@ export class SqlDriver implements IDataDriver {
1519615269

1519715270
/**
1519815271
* Apply {@link presentReadValue} to the result columns a caller of
15199-
* `aggregate()` will read as column VALUES — group keys, and `min`/`max`.
15272+
* `aggregate()` will read as column VALUES — group keys, and `min`/`max` —
15273+
* and [#20335] to its counts and totals (`count`, `count_distinct`, `sum`,
15274+
* `avg`), which take the `'number'` presenter (`AGGREGATE_ANSWER_KIND`).
1520015275
*
1520115276
* Which columns those are cannot be recovered from the rows — the driver has
1520215277
* to be told, because the mapping from column name to meaning is only

0 commit comments

Comments
 (0)