Skip to content

knowledge: daily intake batch 15 - #161

Open
tzhouam wants to merge 1 commit into
mainfrom
knowledge/intake-2026-09-15
Open

tzhouam wants to merge 1 commit into
mainfrom
knowledge/intake-2026-09-15

Conversation

@tzhouam

@tzhouam tzhouam commented Sep 15, 2026

Copy link
Copy Markdown
Collaborator

Daily knowledge-intake batch from omni-reviewbot: executable rules distilled from merged PRs and copilot bugfix runs, append-only per the knowledge contract. Merging this PR is the human promotion gate — review the rules like any other knowledge edit.

Proposed rules

  • knowledge/repos/vllm-omni/components/diffusion/rules-platform-kernels.mdDIFF-1ai (sources: PR #6982)
  • knowledge/repos/vllm-omni/components/serving/rules-request-input.mdSERV-4r (sources: PR #7447)
  • knowledge/repos/vllm-omni/components/distributed/rules.mdDIST-1l (sources: PR #6093)
  • knowledge/repos/vllm-omni/models/lingbot-video/rules.mdLBV-3a (sources: PR #6533)
  • knowledge/repos/vllm-omni/components/serving/rules-request-input.mdSERV-4s (sources: PR #7459)
  • knowledge/repos/vllm-omni/components/diffusion/rules-attention.mdDIFF-1aj (sources: PR #7324)
  • knowledge/repos/vllm-omni/ci/rules-amd.mdOMNI-CI-2i (sources: PR #6978)
  • knowledge/repos/vllm-omni/models/minicpm-o-4-5/rules.mdMCPMO-3f (sources: PR #7517)

Source events

  • merged_pr PR #6982: [Kernel][Boogu-Image] Fuse Q/K RMSNorm + interleaved RoPE via fused_qk_norm_rope
  • merged_pr PR #7447: [Bugfix][Frontend] Honor output_compression on the image generations route
  • merged_pr PR #6093: [Core] Add NIXL omni connector
  • merged_pr PR #6533: [AR-Diffusion] Add session-owned streaming VAE decode
  • merged_pr PR #7459: [Perf][Frontend] Chunk in-memory image file responses
  • merged_pr PR #7444: [Doc][NPU] Keep AllGather for MiniMax-H3 INT8 DLO
  • merged_pr PR #4770: [P0][RFC #3747] RL rollout serving: session management + world_model_env
  • merged_pr PR #7324: [Bugfix][NPU] Materialize causal mask in FlashAttention dense path
  • merged_pr PR #7445: [CI/Build] Align CUDA release image with vLLM 0.29
  • merged_pr PR #7490: [Doc] Update vLLM-Omni WeChat QR code
  • merged_pr PR #7456: [Frontend] Add ComfyUI FastH3 node, fix t2va aspect ratio and dropped audio
  • merged_pr PR #6978: [CI/Build] Add experimental AMD MI300 nightly lane
  • merged_pr PR #7028: [CI/Build] Share nightly CUDA YAML between H100/L4 and B200
  • merged_pr PR #7436: [Performance][MammothModa2] Enable and validate FP8 KV cache for AR stage
  • merged_pr PR #7325: [Bugfix] Forward Qwen3-Omni streaming video sampling parameters
  • merged_pr PR #7517: [Bugfix][MiniCPMO45] Return bare tensor in Thinker forward to fix repetition loop (#7497)
  • merged_pr PR #7461: fix(qwen-image): drop the shadowed dead _get_qwen_prompt_embeds duplicate
  • merged_pr PR #7316: [Bugfix][Frontend] Preserve duplex chat audio format and duration
  • merged_pr PR #7084: [Model] Add Breeze-TTS-2 two-stage AR TTS support (talker + codec, streaming)
  • merged_pr PR #7043: [Bugfix] Allow Fish Speech phoneme-control tokens through normalize_fish_speech_text

Public SDK curation result

{
  "batch_id": "sha256:f8bee4a81616084b0e657df3bfe4f2fd109070302b0dac3e912538d624cf8f41",
  "success": true,
  "attempted": 8,
  "applied": 8,
  "accepted_indexes": [
    0,
    1,
    2,
    3,
    4,
    5,
    6,
    7
  ],
  "rejected_indexes": [],
  "updated_document_ids": [
    "knowledge/repos/vllm-omni/ci/rules-amd.md",
    "knowledge/repos/vllm-omni/components/diffusion/rules-attention.md",
    "knowledge/repos/vllm-omni/components/diffusion/rules-platform-kernels.md",
    "knowledge/repos/vllm-omni/components/distributed/rules.md",
    "knowledge/repos/vllm-omni/components/serving/rules-request-input.md",
    "knowledge/repos/vllm-omni/models/lingbot-video/rules.md",
    "knowledge/repos/vllm-omni/models/minicpm-o-4-5/rules.md"
  ],
  "updated_on": "2026-09-15",
  "validators": [
    {
      "validator_id": "knowledge/tools/check_knowledge_tree.py",
      "passed": true,
      "status": "passed",
      "returncode": 0,
      "output": "提醒:文件接近拆分线:general/review/guides/review-execution-contract.md (183 个非空行,21548 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/ci/rules.md (243 个非空行,27775 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/components/configuration/rules.md (174 个非空行,30403 bytes)\n提醒:目录有 15 个普通页面,应该考虑分类:repos/vllm-omni/components/configuration\n提醒:文件接近拆分线:repos/vllm-omni/components/diffusion/rules-attention.md (195 个非空行,25964 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/components/diffusion/rules-checkpoint-loading.md (211 个非空行,26702 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/components/diffusion/rules-output-lifecycle.md (147 个非空行,21047 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/components/diffusion/rules-system-runtime.md (189 个非空行,27138 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/components/diffusion/rules.md (172 个非空行,27612 bytes)\n提醒:目录有 22 个普通页面,应该考虑分类:repos/vllm-omni/components/diffusion\n提醒:文件接近拆分线:repos/vllm-omni/components/model-executor/rules-bridge-batch.md (134 个非空行,18619 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/components/model-executor/rules.md (193 个非空行,31439 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/components/scheduler/rules.md (299 个非空行,32496 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/components/serving/rules-engine-lifecycle.md (246 个非空行,32271 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/components/serving/rules.md (196 个非空行,31197 bytes)\n提醒:目录有 12 个普通页面,应该考虑分类:repos/vllm-omni/components/serving\n提醒:文件接近拆分线:repos/vllm-omni/models/hunyuan-image3/rules.md (137 个非空行,25387 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/models/minicpm-o-4-5/rules.md (263 个非空行,31243 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/models/minimax-h3/rules-cache-task.md (204 个非空行,28481 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/models/minimax-h3/rules-deployment.md (236 个非空行,27707 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/models/minimax-h3/rules.md (183 个非空行,32421 bytes)\n提醒:文件接近拆分线:repos/vllm-omni/models/qwen3-tts/rules.md (190 个非空行,20981 bytes)\n知识目录检查通过:0 个错误,22 个提醒"
    },
    {
      "validator_id": "knowledge/tools/check_wiki_lint.py",
      "passed": true,
      "status": "passed",
      "returncode": 0,
      "output": "wiki lint 通过:0 个错误,0 个提醒(共 388 页)"
    }
  ],
  "rolled_back": false
}

The vllm-omni release audit runs in CI on this PR as usual.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 15, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review Completed 2026-09-15T07:35:26.688934Z 3685632 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 3685632afe

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".


## LBV-3a — AR 流式 VAE decode 状态必须按 session 持有并计入 admission

- 触发:修改 LingBot/Wan 系 AR-Diffusion `post_decode`、streaming VAE decode、`SupportsStreamingDecode`,或 `model_owned_state_bytes_per_session`。

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Move the cross-model streaming contract out of LingBot

This trigger explicitly covers Wan-family pipelines and shared SupportsStreamingDecode/admission hooks, even though this page declares itself limited to LingBot-specific semantics. Keeping the common lifecycle and accounting contract here means reviews routed to Wan or the shared AR-diffusion runtime will not load it; move the shared portion to the diffusion owner and retain only LingBot-specific requirements here.

AGENTS.md reference: AGENTS.md:L12-L17

Useful? React with 👍 / 👎.

- 验收:CPU/mock 覆盖 CUDA、NPU、XPU、MUSA、unsupported dtype/shape、空行和 contiguous/strided 输入,断言非 CUDA 路径不启动 Triton 且 NPU 保留 native RMSNorm;CUDA 目标硬件以 FP32 reference 核对值、dtype 和容差,并单独记录 kernel 数量/延迟。当前 PR 未新增该模块测试且 NPU CI 不覆盖 H3,在补齐前不得称为生产级跨平台支持。^[PR #6281]


## DIFF-1ai — fused_qk_norm_rope 的 interleaved 模式与 token gate 必须分模式闭合

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Update every affected owner's index

All seven rule pages gain rules, but this commit updates no nearest _index.md. For example, the diffusion index still advertises rules-platform-kernels.md as containing only DIFF-1t through DIFF-1w, and the LingBot Direct map has no LBV-3a route; consumers following the repository's index-first routing can therefore miss the new contracts. Update each affected owner index with the new route/summary and metadata.

AGENTS.md reference: AGENTS.md:L23-L24

Useful? React with 👍 / 👎.

- 禁止:让 `cards_2+` 测试在单卡 worker 上 spawn rank1→GPU1 导致 `invalid device ordinal`;用邻近 green shard 宣称 multi-GPU routing 已修好。
- 验收:pipeline argv/collection 断言单卡 job 排除 multi-card markers、双卡 job 收集目标文件;硬件 marker helper 覆盖 ROCm 声明。^[PR #7234]

## OMNI-CI-2i — AMD nightly suite 必须可显式选中且不受 L2/L3 skip-ci 误杀

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Register the new PRs in page-level sources

This rule cites PR #6978 only in the paragraph while leaving the page's frontmatter sources unchanged; the same omission affects all other rules added by this batch. The knowledge schema requires PR-learning pages to carry both page-level sources: provenance and paragraph-level references, so metadata consumers cannot discover the evidence behind these additions. Add each originating PR, and any relevant upstream source paths, to its affected page's sources list.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants