-
Notifications
You must be signed in to change notification settings - Fork 0
615 lines (586 loc) · 32.7 KB
/
Copy pathci.yml
File metadata and controls
615 lines (586 loc) · 32.7 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
name: CI
on:
push:
branches: [main]
pull_request:
branches: [main]
merge_group:
jobs:
# The four Node floor declarations (`engines.node` in the root, in `apps/docs`
# and in `tools/ci-scripts`, plus `.node-version`) are read by nothing in the
# install path: `.npmrc` sets no `engine-strict`, pnpm does not enforce
# `engines` by default, and every workflow here pins `node-version`
# explicitly instead of consulting them. This job is what makes them
# mechanically checkable. It needs no install — the script is zero-dependency
# and reads the lockfile as text — so it stays a seconds-long job that can
# run alongside `build`.
node-floor:
name: Node floor
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v7
- uses: actions/setup-node@v7
with:
node-version: 22
# 裁决 (PR #74): a validator observed only green is indistinguishable
# from one that cannot go red. The fixtures run before the real scan, so
# a rule that stopped being able to fail fails the job on its own.
- name: Self-test
shell: bash
run: node .github/scripts/check-node-floor.mjs --self-test
# `shell: bash` is load-bearing here, not tidiness. The DEFAULT shell for
# a `run:` step is `bash -e {0}` with no pipefail, so in `node ... | tee`
# the step takes tee's exit status and a gate that exits 1 passes the job
# silently. Naming the shell gets `bash --noprofile --norc -eo pipefail
# {0}`, which propagates it.
- name: Check
shell: bash
run: node .github/scripts/check-node-floor.mjs | tee -a "$GITHUB_STEP_SUMMARY"
build:
runs-on: ubuntu-latest
# `NEXT_PRIVATE_STANDALONE` is what `@opennextjs/aws` sets before it runs
# `next build` — its own comment reads "Equivalent to setting `output:
# "standalone"` in next.config.js". Without it a plain `next build`
# produces no `.next/standalone/`, and the packaging step below fails on a
# missing `pages-manifest.json` three directories inside it. Measured on
# this branch before it was set: `ENOENT ... .next/standalone/apps/docs/
# .next/server/pages-manifest.json`.
#
# Set for the whole job rather than for the deploy path only, so that what
# a pull request builds is the same shape as what gets published. A build
# that differs from the deploy build is a small instance of the defect this
# card is about.
#
# It has to be declared in `turbo.json` as well: turbo 2 runs tasks in
# strict env mode, so an undeclared variable never reaches `next build` —
# and declaring it is also what puts it in the cache key, so a `.next`
# cached from before this line cannot be replayed without the standalone
# tree the packaging step needs.
env:
NEXT_PRIVATE_STANDALONE: 'true'
steps:
- uses: actions/checkout@v7
- uses: pnpm/action-setup@v6
- uses: actions/setup-node@v7
with:
node-version: 22
cache: pnpm
- run: pnpm install --frozen-lockfile
# `scripts/pm/check-half-states.mjs` is a verbatim upstream copy (#237)
# whose 1551 cases were run by nothing here. It is a STEP and not a
# registry entry because `tools/ci-scripts/run-self-tests.mjs` scans
# `.github/scripts` top level only, and that is load-bearing — it is what
# lets `.github/scripts/lib/` exist without tripping the
# unregistered-self-test rule. Widening the scan to reach one script would
# change this repo's gate topology; a step changes nothing.
#
# What it buys is a drift detector, not a hash-pin: #237's ablation
# mutated `DEFAULT_SWEEP_REPO` in this copy and 2 of the 1551 went red, so
# an edit here that changes BEHAVIOUR fails. One that changes bytes
# without changing behaviour still passes — pinning this copy to upstream
# byte for byte needs a cross-repo credential and is a separate decision.
#
# `--self-test` is the whole of it. Without the flag the script runs a
# live sweep needing a transport prerequisite this job does not have; that
# caller is `.github/workflows/half-state-patrol.yml`, on its own
# schedule. Zero-dependency and about a second, so it runs before the
# build rather than behind it.
- name: Half-state sweeper self-test
shell: bash
run: node scripts/pm/check-half-states.mjs --self-test
# `content/docs/**/*.zh-Hant.mdx` and `meta.zh-Hant.json` are generated
# from the Simplified siblings by `apps/docs/scripts/gen-zh-hant.mjs` and
# committed, because `lib/seo.ts` tells a real translation from an English
# fallback by the presence of a locale-suffixed FILE — a conversion done
# while rendering would leave the locale out of every sitemap entry and
# hreflang cluster. Committed output needs a gate or it drifts: this
# regenerates in memory and compares bytes, so a hand edit, a stale file
# whose source was retired, and a converter upgrade nobody re-ran all fail
# here with the same one-line fix.
#
# Before `type-check` on purpose: it needs no build, and a content drift
# reported as a type error is a wrong first diagnosis.
- name: Generated zh-Hant is current
run: node apps/docs/scripts/gen-zh-hant.mjs --check
- run: pnpm turbo run type-check --continue
- run: pnpm turbo run build
# Reads the BUILT sitemap and asserts its locale composition against an
# oracle derived from `content/docs/`. It has to come after `build` — the
# measurement is taken off the artifact, because importing `sitemap.ts`
# pulls in the whole MDX collection and the `@/` alias.
#
# Not a turbo task on purpose: turbo would hash it against the ci-scripts
# package's own inputs, which do not include `apps/docs/.next/`, so a
# locale regression would replay a cached green. `pnpm turbo run test`
# below still runs this script's `--self-test`, which is what keeps its
# rules provably able to fail.
#
# #282 made the same step the guard for #197: no HTML numeric character
# reference and no malformed link target in the `llms-full.txt` body or
# in any per-page `llms.mdx` body. #197's fix is a patch pinned to one
# exact `fumadocs-core` version. On a bump, pnpm refuses the stale pin
# when the lockfile is regenerated, but the first remedy its error
# offers is deleting the pin, and after that install is green and the
# defect is back; this is where that turns red. Every run first feeds
# the #197 shape through the same scan and fails unless it goes red,
# and the log says so. It reads `.next`, not `.open-next`, so it sits
# here after `build` and not behind the artifact upload.
#
# `shell: bash` is load-bearing, not tidiness. The DEFAULT shell for a
# `run:` step is `bash -e {0}` with no pipefail, so in `node ... | tee`
# the step takes tee's exit status and a gate that exits 1 passes the job
# silently. Naming the shell gets `bash --noprofile --norc -eo pipefail
# {0}`, which propagates it.
- name: Locale surface
shell: bash
run: node .github/scripts/check-locale-surface.mjs | tee -a "$GITHUB_STEP_SUMMARY"
# #171: one positioning constant (`apps/docs/lib/positioning.ts`, quoting
# the objectstack README), one brand spelling, and no stale sentence
# coming back. The brand and the positioning copies are read off the
# BUILT pages and `llms` bodies — the 2026-09-08 mandate is about shipped
# output, and a source scan would flag the code comment in `lib/i18n.ts`
# that the ruling on #171 accepted — so this sits after `build`, on the
# same footing as `Locale surface` above: not a turbo task, its
# `--self-test` under `pnpm turbo run test` below, `shell: bash` for the
# same pipefail reason. The stale-sentence rules read the English sources
# under `content/docs/`, so a finding names the `path:line` to open.
- name: Positioning
shell: bash
run: node .github/scripts/check-positioning.mjs | tee -a "$GITHUB_STEP_SUMMARY"
- run: pnpm turbo run test
# Defect 2 of #269: the deploy used to run its own `pnpm install` and its
# own `opennextjs-cloudflare build`, so CI built the site, threw it away,
# and the deploy published a SECOND build that nothing here had checked.
# The published artifact was unverified by construction.
#
# `--skipNextBuild` packages the `.next` output `pnpm turbo run build`
# produced above — the same output `Locale surface` measured and the same
# tree every step in this job passed — and `deploy-docs.yml` uploads THIS
# bundle rather than making another one.
#
# Last in the job on purpose: the artifact then only exists for a commit
# that cleared every gate above it.
#
# #262 removed the `push` + `refs/heads/main` condition this step used to
# carry. It ran only where a deploy would follow, which meant a pull
# request never packaged a Worker and therefore could never be told its
# Worker was too big — the whole defect that card is about. It runs on
# every event now so that the size gate below has something to weigh, and
# the artifact upload stays `main`-only underneath it.
#
# This is NOT a second build, and that distinction is what makes the
# price acceptable: `--skipNextBuild` re-packages the `.next` tree
# `pnpm turbo run build` already produced. Measured on this branch:
# `turbo run build` 101s, this packaging step 24s, the dry-run weigh-in
# below 9s. A pull request pays ~33s more than before, not another 101s.
- name: Package the Worker from the build this job tested
working-directory: apps/docs
run: pnpm exec opennextjs-cloudflare build --skipNextBuild
# #262. Nothing in this repository ever weighed the Worker. The only
# thing that checked it was the Cloudflare API, at upload time, on
# `main`, AFTER merge — and the rejection lands on version CREATION, so
# nothing 500s, no page changes, and the site silently stops moving. That
# is how this repo ran 35 consecutive red deploys (runs #106-#140,
# 2026-08-25 to 09-02) with the `build` job green for every one of them:
# `build` compiles the Next app, it never bundled or weighed the Worker.
#
# ## Why `wrangler deploy --dry-run` and not `stat`
#
# The number Cloudflare enforces is wrangler's `Total Upload:` line, and
# `--dry-run` prints it from the same code path a real deploy uses,
# without calling the API and without credentials. Stat-ing files instead
# would mean re-deriving WHICH files count, and that guess is the trap
# #262 names: `handler.mjs` alone measures 48.48 MiB while the upload is
# 58553.98 KiB, so a budget stated against it tracks nothing.
#
# Verified on this branch, at `0e26657f`, from `--dry-run --outdir`:
# the uploaded set is `worker.js` (58383222 B) plus three sidecar
# modules — resvg.wasm (1378357 B), Geist-Regular.ttf.bin (125956 B),
# yoga.wasm (71736 B) = 59959271 B = 58553.98 KiB, which is the printed
# line to the hundredth. `worker.js.map` (85819146 B, larger than the
# whole budget) and the `.open-next/assets` + `.open-next/cache` trees
# are NOT in it; assets upload separately and do not count here.
#
# ## Calibration against a number Cloudflare actually accepted
#
# Same commit `0e26657f`, run 33891143864, Worker version
# 2170b929-5879-4b3f-b7a2-9eda750158dd, the upload Cloudflare ACCEPTED:
# Total Upload: 58555.94 KiB
# This step, on that commit:
# Total Upload: 58553.98 KiB
# 1.96 KiB low — 0.0033%. The reading tracks the enforced figure.
#
# ## Accurate against the limit; NOT repeatable to better than a few KiB
#
# #277. Accuracy is this gate's job and the figure above is the right
# claim for it. It says nothing about REPEATABILITY, and a reader who
# meets 0.0033% at the point of use will reasonably assume a small
# per-PR delta means something. Measured, it does not.
#
# The strongest control is ONE commit run twice, with no tree difference
# at all — not a workflow file, not a specifier line, nothing. `66cb0a4`
# is an empty commit whose tree hash is byte-identical to `main`'s at
# `50aacb9`; run 34252385192, attempts 1 and 2, landed on two different
# `ubuntu-latest` runners. The third reading is from `1bf8b70`, a
# DIFFERENT tree — it carries this comment — but the same bundle inputs,
# since a workflow file cannot enter the Worker. Bundle-identical, not
# tree-identical: the weaker of the two controls, and labelled as one.
#
# 66cb0a4 attempt 1, job 102149701548 58549.04 KiB
# 66cb0a4 attempt 2, job 102152795435 58553.46 KiB
# 1bf8b70 run 34277657373, job 102234545211 58549.06 KiB
#
# Largest gap: 4.42 KiB, between two runs of the SAME commit. That is
# larger than the 1.96 KiB agreement above, so the agreement cannot be
# read as sub-KiB precision either — it is one comparison, taken inside
# this much noise.
#
# The SHAPE is a tight cluster and one outlier, NOT a wander across a
# 4.42 KiB band: two of the three agree to 0.02 KiB and the third sits
# ~4.4 KiB above both. Three points cannot carry a distribution and none
# is modelled here. What is claimed is only that repeated readings over
# identical bundle inputs are not equal, and how far apart the measured
# ones landed.
#
# wrangler's gzip column moves too — 8800.26, 8800.65, 8797.64 KiB — so
# the artifact's bytes genuinely differ; this is not the same number
# being displayed differently. It does not track the raw figure (the
# third reading has the smallest compressed size and a middling raw
# one), so it confirms that much and nothing more.
#
# Holding the machine fixed shrinks the spread without closing it: four
# cold rebuilds of `66cb0a4`'s tree in one container gave 58548.93,
# 58548.93, 58548.94 and 58549.55 KiB — 0.62 KiB apart.
#
# So this number answers "is the bundle near the ceiling", not "did my
# PR grow the bundle". A single-digit-KiB move between two commits is
# inside the spread measured above and has not been shown to be a change
# in the bundle at all; a real growth of a few KiB is equally invisible.
# Attributing a small delta needs a same-tree control — the same commit
# run twice — not the previous commit's reading.
#
# n is small (3 CI runs, 4 local builds) and 4.42 KiB is the largest
# same-tree gap MEASURED, not a proven bound. Causes are NOT measured
# here and none is claimed.
#
# ⚠️ A SNAPSHOT at the commits named above, deliberately NOT maintained.
# Every run of this workflow produces another reading, so "fold in the
# latest" has no end; the population was frozen at these three instead
# of chased. Nobody is obliged to update this block, and a later reading
# that differs is the point being made, not a defect in it. If the
# question ever needs more than "a few KiB", take a fresh same-commit
# series rather than appending to this one.
#
# ## The budget
#
# ONE constant, below, with the limit written beside it; every
# percentage and headroom figure is COMPUTED from those two and never
# typed. #262's own banner is why: it quoted `89.3%` (against 65536) and
# `~5.3 MiB headroom` (against 64000) in the same paragraph, two ceilings
# in one card, ~1.5 MiB of phantom room. The limit is 65536 KiB, taken
# from run #140's own rejection text.
#
# 61440 KiB = 60 MiB = 93.75% of the limit. It has to clear two bars:
# - above today's 58553.98 KiB (89.35%), or it is red on arrival —
# 2886.02 KiB, 4.93% of growth room;
# - below run #105's 62747.87 KiB (95.75%), the LAST SUCCESSFUL deploy
# before the outage, or it would have watched that go by. The commit
# that finally crossed the line was 15 lines and made its own output
# smaller; at 95.7% anything landing that week would have done it.
# This budget goes red 1307.87 KiB before that point.
#
# No warn tier on purpose. Every band you could draw between today's
# 89.35% and this budget's 93.75% is under 4.5 points wide and the bundle
# is already inside it, so a warning would be lit from its first run and
# read as wallpaper. The reading is printed to the step summary on EVERY
# run instead — visible before it is a problem, which is what a warn tier
# was wanted for.
#
# `shell: bash` is load-bearing, not tidiness. The DEFAULT shell for a
# `run:` step is `bash -e {0}` with no pipefail, so in `awk ... | tee`
# the step takes tee's exit status and a gate that exits 1 passes the job
# silently. Naming the shell gets `bash --noprofile --norc -eo pipefail
# {0}`, which propagates it.
#
# Ablated before it was trusted, per the pattern
# `.github/scripts/smoke-docs.mjs` already sets here — a probe that
# cannot fail is indistinguishable from one that passed. Padding the
# bundle by 4 MiB took the reading to 62650.01 KiB (95.60%, within 98 KiB
# of run #105's pre-outage figure) and this step exited 1; restoring the
# bundle byte-for-byte returned it to exit 0. Deleting the `Total
# Upload:` line from wrangler's output exits 1 as well, rather than
# passing on a measurement it never took.
- name: Worker bundle fits the size budget
working-directory: apps/docs
shell: bash
env:
# The budget, and the limit it is set below. Nothing else in this
# repository states a Worker size; every other figure is derived.
# Cloudflare's limit 65536 KiB (64 MiB, from run #140's rejection)
# this budget 61440 KiB (60 MiB, 93.75% of the limit)
WORKER_BUDGET_KIB: '61440'
WORKER_LIMIT_KIB: '65536'
WRANGLER_SEND_METRICS: 'false'
run: |
set -euo pipefail
pnpm exec wrangler deploy --dry-run 2>&1 | tee "$RUNNER_TEMP/worker-size.log"
SIZE_KIB="$(sed -n 's/^.*Total Upload: \([0-9][0-9.]*\) KiB.*$/\1/p' "$RUNNER_TEMP/worker-size.log" | tail -n 1)"
# An unreadable measurement is a finding, never a skip. A size gate
# that silently weighs nothing passes forever, which is the failure
# mode this step exists to end rather than to reproduce.
if [ -z "$SIZE_KIB" ]; then
{
echo "### Worker bundle size — NOT MEASURED"
echo
echo "\`wrangler deploy --dry-run\` printed no \`Total Upload:\` line, so this step"
echo "weighed nothing. Failing rather than passing: an unmeasured bundle is the"
echo "state this gate exists to end."
} | tee -a "$GITHUB_STEP_SUMMARY"
echo "::error::Worker bundle NOT measured — no 'Total Upload:' line in the wrangler dry-run output."
exit 1
fi
awk -v size="$SIZE_KIB" -v budget="$WORKER_BUDGET_KIB" -v limit="$WORKER_LIMIT_KIB" '
BEGIN {
over = (size > budget)
printf "### Worker bundle size — %s\n\n", (over ? "OVER BUDGET" : "within budget")
printf "| | KiB | %% of limit |\n|:--|--:|--:|\n"
printf "| measured | %.2f | %.2f %% |\n", size, size / limit * 100
printf "| budget | %d | %.2f %% |\n", budget, budget / limit * 100
printf "| Cloudflare limit | %d | 100 %% |\n", limit
printf "\nHeadroom to budget: **%.2f KiB** · headroom to the limit: %.2f KiB\n", budget - size, limit - size
if (over) {
printf "\nThis change takes the Worker past the declared budget.\n\n"
printf "Cloudflare rejects an over-limit upload at VERSION CREATION, on `main`, after\n"
printf "merge: nothing 500s, no page changes, the site simply stops being updated.\n"
printf "This repo ran 35 consecutive red deploys that way with `build` green for every\n"
printf "one of them. Reduce the bundle, or raise the budget deliberately and say why.\n"
}
exit (over ? 1 : 0)
}' | tee -a "$GITHUB_STEP_SUMMARY"
# `include-hidden-files` is load-bearing, not tidiness: the compiled
# OpenNext config the deploy reads lives at `.open-next/.build/`, and
# upload-artifact excludes dotted paths by default. Without it the
# download succeeds, the deploy exits 1 on a missing config, and the
# cause is three directories away from the message.
#
# Still `main`-only: a pull request now packages and weighs a Worker, but
# it has nothing to deploy, so it uploads nothing. An `if:` carrying no
# status function implies `success()`, so the size gate above also gates
# this — an over-budget Worker never becomes an artifact and never
# reaches `deploy-docs.yml`.
- name: Upload the Worker bundle
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
uses: actions/upload-artifact@v7
with:
name: docs-worker
path: apps/docs/.open-next
include-hidden-files: true
if-no-files-found: error
retention-days: 3
# #274. The 09-04 outage was found by a human looking at the live site.
# #269 answered that with detection and recovery — the deploy is gated on
# CI, the published artifact is the one CI tested, the live site is
# smoke-checked after deploying, and a bad deploy auto-rolls back — but
# the only environment in which a rendering defect is DETECTED is still
# production, and the chain `deploy succeeds -> smoke fails -> rollback
# fires` has never once executed end to end. A check that fires before
# merge costs a red pull request; the same check firing after merge costs
# a live outage plus a recovery path nobody has ever seen run.
#
# ## The same script, pointed at a different base
#
# `.github/scripts/smoke-docs.mjs` is the post-deploy check
# `deploy-docs.yml` runs against `https://docs.objectos.ai`. It is
# invoked here unmodified, with `--base` pointing at a local preview.
# NOT a second implementation of "does the site render": two copies of
# those rules drift, and the copy that drifts is the one nobody watches.
# Its live negative control — a `/docs/` slug no page claims, which must
# produce findings or the run fails on `negative-control-passed` — comes
# along with it, which is what makes a green here worth reading.
#
# ## This is not a second build
#
# `opennextjs-cloudflare preview` does not build. It populates the
# incremental cache (for this app, a copy of `.open-next/cache` into the
# Workers static assets) and then runs `wrangler dev` on the `.open-next`
# package the step above produced with `--skipNextBuild`. Since #262
# removed the `main`-only condition from that packaging step, that
# package exists on every pull request, so this step adds a preview boot
# and four fetches and nothing else. Measured in this repo's container,
# against the package already sitting in the tree: `Ready on` at 41 s,
# the smoke run itself 1 s. No `opennextjs-cloudflare build`, no
# `next build`, no Cloudflare credentials — `wrangler dev` serves the
# Worker locally under real workerd.
#
# ## Why it runs LAST, after the artifact upload
#
# `preview` copies `.open-next/cache` into `.open-next/assets/cdn-cgi`,
# which for this app is 268 MB: measured, `.open-next` goes from 386 MB
# to 653 MB the moment the preview boots. `opennextjs-cloudflare deploy`
# makes that same copy in the deploy job from `.open-next/cache`, which
# the artifact already carries — so running this before the upload would
# add 268 MB to every `main` artifact, both ways across the wire, to
# ship a copy the deploy remakes anyway. Placed here, the uploaded bundle
# is byte-for-byte what it was before this step existed, and the gating
# is unchanged: `deploy-docs` needs the whole `build` job, so a red here
# keeps a bad render off production just as a red anywhere above it does.
#
# ## A dead server must not read as a pass
#
# "No findings" from a preview that never started is indistinguishable
# from "the site renders", and this lane logged four probes of exactly
# that shape in a single day. So readiness is asserted from wrangler's
# own `Ready on` line before anything is judged, with the preview log
# printed into the step summary when it does not arrive, and the step
# exits 1 rather than reporting a measurement it never took.
#
# `setsid` is load-bearing, not tidiness. The preview is a chain of six
# processes — pnpm, node, sh, pnpm, wrangler, workerd — and a SIGTERM to
# the pnpm wrapper at the top leaves workerd running and holding the
# port. Measured here: the cleanup `wait` never returned and the whole
# thing hung. Starting it in its own process group lets the trap signal
# the GROUP and take the chain with it; the `$$` comparison is there so
# that a `setsid` which did not take effect can never turn that into the
# step killing itself.
- name: The docs site renders — smoke-check a local preview
working-directory: apps/docs
shell: bash
env:
PREVIEW_PORT: '8792'
# Boot budget for the preview. Measured at 41 s in this repo's
# container; the margin is for a cold runner, and overrunning it is
# a finding (NOT MEASURED), never a skip.
PREVIEW_READY_TIMEOUT_S: '180'
WRANGLER_SEND_METRICS: 'false'
run: |
set -euo pipefail
BASE="http://127.0.0.1:${PREVIEW_PORT}"
PREVIEW_LOG="$RUNNER_TEMP/preview.log"
SMOKE_LOG="$RUNNER_TEMP/smoke.log"
: > "$PREVIEW_LOG"
setsid pnpm exec opennextjs-cloudflare preview -- \
--port "$PREVIEW_PORT" --ip 127.0.0.1 > "$PREVIEW_LOG" 2>&1 &
PREVIEW_PID=$!
SELF_PGID="$(ps -o pgid= -p $$ | tr -d ' ')"
# The preview's process group is read HERE, at kill time, and never
# cached at launch. `setsid` only changes the group once the forked
# child has exec'd it, so a `ps` issued straight after `&` is a race
# that can still see the STEP's own group. Measured on a GitHub
# runner, run 34250422860: the cached read lost that race, the guard
# below fell back to signalling the pnpm wrapper alone, and the
# runner's own orphan sweeper had to terminate esbuild and two
# workerd processes after the job. The identical code cleaned up
# correctly in this repo's container every time — which is exactly
# how a race presents, and why the group is resolved at use.
cleanup() {
PREVIEW_PGID="$(ps -o pgid= -p "$PREVIEW_PID" 2>/dev/null | tr -d ' ' || true)"
if [ -z "$PREVIEW_PGID" ]; then
: # already gone — nothing to signal
elif [ "$PREVIEW_PGID" != "$SELF_PGID" ]; then
kill -TERM "-$PREVIEW_PGID" 2>/dev/null || true
for _ in 1 2 3 4 5; do
pgrep -g "$PREVIEW_PGID" >/dev/null 2>&1 || break
sleep 1
done
kill -KILL "-$PREVIEW_PGID" 2>/dev/null || true
else
# `setsid` did not take effect and the preview is sharing this
# step's group, which must NEVER be signalled as a group or the
# step kills itself. Signal the pid and say so out loud rather
# than leaving a silent orphan for the runner to sweep.
kill -TERM "$PREVIEW_PID" 2>/dev/null || true
echo "::warning::preview pid $PREVIEW_PID is in this step's own process group — signalled the pid alone, its descendants may survive."
fi
}
trap cleanup EXIT
READY=0
for _ in $(seq 1 "$PREVIEW_READY_TIMEOUT_S"); do
if grep -q 'Ready on http' "$PREVIEW_LOG"; then READY=1; break; fi
if ! kill -0 "$PREVIEW_PID" 2>/dev/null; then break; fi
sleep 1
done
if [ "$READY" -ne 1 ]; then
# Which of the two shapes it was. They call for different fixes —
# a crashed preview is a broken bundle, an exhausted budget is a
# slow runner — and the log below is the same either way, so the
# sentence has to say which one the reader is looking at.
if kill -0 "$PREVIEW_PID" 2>/dev/null; then
WHY="the ${PREVIEW_READY_TIMEOUT_S}s boot budget ran out with the preview still starting"
else
WHY="the preview process exited before it was ready"
fi
{
echo "### Pre-merge render check — NOT MEASURED"
echo
echo "The local preview never printed \`Ready on\`: ${WHY}. Nothing was checked."
echo "Failing rather than passing: \"no findings\" from a server that never started"
echo "is indistinguishable from a rendered site."
echo
echo '```'
tail -n 40 "$PREVIEW_LOG"
echo '```'
} | tee -a "$GITHUB_STEP_SUMMARY"
echo "::error::Local preview never became ready (${WHY}) — the render check measured nothing."
exit 1
fi
echo "preview ready: $(grep -m1 'Ready on http' "$PREVIEW_LOG" || true)"
echo "preview pid $PREVIEW_PID in process group $(ps -o pgid= -p "$PREVIEW_PID" | tr -d ' '), step in $SELF_PGID"
# Exit code captured before anything pipes it. `cmd | tee` hands back
# tee's status, and this step's verdict is the script's.
set +e
node "$GITHUB_WORKSPACE/.github/scripts/smoke-docs.mjs" --base "$BASE" \
> "$SMOKE_LOG" 2>&1
SMOKE_EXIT=$?
set -e
cat "$SMOKE_LOG"
{
if [ "$SMOKE_EXIT" -eq 0 ]; then
echo "### Pre-merge render check — the site renders"
else
echo "### Pre-merge render check — FINDINGS"
fi
echo
echo "Ran \`.github/scripts/smoke-docs.mjs\` — the same script \`deploy-docs.yml\` runs"
echo "against the live site — against a local \`opennextjs-cloudflare preview\` of the"
echo "Worker this job packaged, at \`$BASE\`."
echo
echo '```'
cat "$SMOKE_LOG"
echo '```'
} >> "$GITHUB_STEP_SUMMARY"
exit "$SMOKE_EXIT"
# Defect 1 of #269: `deploy-docs.yml` used to hang off `push: branches:
# [main]` exactly as this workflow does, so the two ran in PARALLEL and a
# commit that failed any gate above still deployed. There was no `needs:` and
# no `workflow_run` anywhere.
#
# As a job here it cannot start until `node-floor` and `build` are green, and
# the `if:` keeps it off pull requests and merge groups. `workflow_run` would
# also have gated it, but it fires on a FAILED run too — the conclusion has
# to be re-checked by hand inside the workflow — and it runs detached from
# the run whose artifact it publishes, which is what the `with:` line here
# depends on.
#
# ⚠️ This deploy is expected to FAIL while #261 is open: `main` builds a
# Worker over Cloudflare's 64 MiB limit, so the upload is rejected at version
# creation and the serving version cannot be displaced. That is the current
# deliberate steady state, not a regression from this wiring.
deploy-docs:
name: Deploy docs
needs: [node-floor, build]
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
# Preserves what `deploy-docs.yml` declared for itself before it became a
# called workflow: one deploy at a time, and never cancel one in flight.
concurrency:
group: deploy-docs
cancel-in-progress: false
permissions:
contents: read
actions: write # dispatch rollback-docs.yml when the smoke check fails
issues: write # file or update the one deploy-failure card
uses: ./.github/workflows/deploy-docs.yml
with:
artifact_name: docs-worker
secrets: inherit