From 5a57938d7e9ed93f064d504bc66323aebcb045f1 Mon Sep 17 00:00:00 2001 From: kakyungkim Date: Mon, 27 Jul 2026 12:01:56 +0000 Subject: [PATCH 1/4] =?UTF-8?q?docs:=20BIOP01-45=20=EC=84=A4=EA=B3=84?= =?UTF-8?q?=EC=B4=88=EC=95=88=20=EC=84=A0=EA=B2=B0=20=EA=B0=B1=EC=8B=A0=20?= =?UTF-8?q?=E2=80=94=20=E2=91=A0=EB=9D=BC=EC=9A=B0=ED=84=B0=20blocker=20?= =?UTF-8?q?=ED=95=B4=EC=86=8C(BIOP01-71=20A=EB=B3=B5=EC=9B=90=C2=B7PR#5=20?= =?UTF-8?q?=EB=A8=B8=EC=A7=80=20cb886ee,=20doctor=20PASS=20=EC=8B=A4?= =?UTF-8?q?=EC=B8=A1)=20=E2=91=A1=EC=8B=A0=EA=B7=9C=20=ED=8A=B8=EB=A6=AC?= =?UTF-8?q?=EA=B1=B0=20=EA=B2=B0=EC=A0=95(openclaw=20CLI=20=ED=8C=80=200/6?= =?UTF-8?q?=20=EB=B6=80=EC=9E=AC=E2=86=92codex=20=EB=8C=80=EC=B2=B4,=20?= =?UTF-8?q?=ED=9A=8C=EC=9D=98=20=EC=83=81=EC=A0=95).=20manifest=20?= =?UTF-8?q?=EC=9B=8C=EC=BB=A4=EC=B8=B5=EC=9D=80=20=EB=91=90=20=EA=B2=B0?= =?UTF-8?q?=EC=A0=95=20=EB=B6=88=EB=B3=80,=20CPU=EA=B2=80=EC=A6=9D=20?= =?UTF-8?q?=EC=84=A0=ED=96=89=20=EA=B0=80=EB=8A=A5=20=EC=9E=AC=ED=99=95?= =?UTF-8?q?=EC=9D=B8?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- .../ORCHESTRATION-WIRING-DESIGN.md | 25 ++++++++++++------- 1 file changed, 16 insertions(+), 9 deletions(-) diff --git a/pipeline/hspc-velocity-benchmark/ORCHESTRATION-WIRING-DESIGN.md b/pipeline/hspc-velocity-benchmark/ORCHESTRATION-WIRING-DESIGN.md index d5d1800..ea7e355 100644 --- a/pipeline/hspc-velocity-benchmark/ORCHESTRATION-WIRING-DESIGN.md +++ b/pipeline/hspc-velocity-benchmark/ORCHESTRATION-WIRING-DESIGN.md @@ -4,15 +4,22 @@ > 이 문서는 코드가 아니라 **호출 규약의 제안**이다. 합의 전 실행 배선을 만들지 않는다. > 근거: BIOP01-45 본문 4개 할 일 + BIOP01-22 오케스트레이션 검토(계획 계층은 정합, 계획→실행 연결만 없음). -> 🚧 **선결 미해결 (2026-07-26 추가, BIOP01-71 [P0]) — 착수 보류.** -> 이건규 님 2차 전수조사에서, 이 초안 §3 route가 위임하는 **`skills/ROUTES.md` 라우터가 유실**된 것이 확인됐다. -> git 실측(지용기, 07-27): `bc7f824`(6/14)가 ROUTES.md를 **추가** → 5주 유지(내 env-repro `7f38b23` 7/15 포함) -> → **`275def2`(7/20 PR #4 충돌 해소)에서 41파일 전부 소실**(부모 41 → 0). 의도적 폐기가 아니라 머지 사고. -> 복원본은 PR #5(`7039bc4` 기준 41파일)에 대기 중. -> **→ 이 초안 §3-2·3-4의 route 의사코드는 그 라우터 위에 서 있으므로, BIOP01-71에서 A(복원)/B(문서정정)이 -> 확정되기 전에는 실행 배선을 만들지 않는다.** 단 §2 `runner_manifest.yaml`(워커 계층 계약)은 라우터 -> 아래층이라 A/B 어느 쪽이든 그대로 유효하다 — manifest는 "어느 runner를 어떤 env로"만 정의하고, -> "누가 트리거하나"(라우터)는 상위 결정이다. (근거: BIOP01-45 코멘트 11397·11439.) +> ✅ **선결 ① 해소 (2026-07-27, BIOP01-71 A(복원) 확정·머지).** +> 이 초안 §3 route가 위임하는 `skills/ROUTES.md` 라우터가 `275def2`(7/20 PR #4 머지)에서 41파일 사고 유실됐던 건, +> kkkim A(복원) 확정 → **PR #5 머지(`cb886ee`)로 복원 완료**. 실측: `skills/ROUTES.md`(69줄) 등 41파일 복귀, +> `harness_doctor` **RESULT PASS(phantom-path 0)**. → 라우터 blocker 제거, 초안 §3의 dataset→task 2단 라우팅이 실체 위에 섬. +> +> 🚧 **선결 ② 신규 (2026-07-27, 이건규 BIOP01-71 코멘트 11506) — 트리거 계층 결정 대기.** +> **`openclaw` CLI가 팀 컨테이너 6개 전부에 없다(0/6, 지용기 이 박스도 실측 부재).** 이 초안 제목·§3이 전제한 +> "OpenClaw route로 트리거"가 **실행 도구 부재**로 지금 성립하지 않는다. 대체 경로는 존재한다 — +> `skills/OPENCLAW-RUN.md`가 같은 `agents/openai.yaml`을 **codex로 실행**하는 길을 기록하고 있고 codex는 인증돼 동작한다. +> → **트리거를 openclaw로 갈지, codex 기준으로 문서·설계를 정정할지는 회의 결정**(BIOP01-45와 묶어 상정, 이건규 제안). +> **이 결정이 나기 전 §3-2·3-4 route 의사코드는 확정하지 않는다**(openclaw 전제라 codex면 표기가 달라짐). +> +> ✅ **단 §2 `runner_manifest.yaml`(워커 계층)은 두 결정 모두에 불변이다.** manifest는 "어느 runner를 어떤 env로"만 +> 정의하고, "누가 트리거하나"(openclaw/codex/사람)는 상위 계층이다. 라우터 복원으로 하위 배선 근거는 이미 실체가 됐고, +> 트리거가 openclaw든 codex든 manifest는 그대로 재사용된다. → **manifest 스키마 확정 + §6 CPU 검증은 트리거 결정과 무관하게 선행 가능.** +> (근거: 코멘트 11397·11439의 계층 분리 논지 그대로.) --- From 3531d1ede2c36f699233893b229165f5be2f3be9 Mon Sep 17 00:00:00 2001 From: kakyungkim Date: Mon, 27 Jul 2026 22:05:51 +0900 Subject: [PATCH 2/4] =?UTF-8?q?paper=5Fanalysis:=20Wu=202026=20GB=20veloci?= =?UTF-8?q?ty=20benchmark=20=EB=B6=84=EC=84=9D=20=EC=B6=94=EA=B0=80=20(?= =?UTF-8?q?=EC=8A=A4=EC=BF=B1=20=EC=A0=90=EA=B2=80)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 오늘(2026-07-27) Genome Biology 27:242에 게재된 세 번째 종합 velocity 벤치마크. 우리 원고 타깃 저널이라 스쿱 점검이 필요해 전문(36p) 정독 후 kkkim paper_analysis 구조로 저장. paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/ - paper-info.yaml (scoop_check 블록 포함) - _abstract / _core / _lens-academic / _lens-industry / _methodology-brief - sources/ (.bib·.url·abstract.txt·fulltext 추출본; PDF는 gitignore 관례) 판정: 스쿱 아님. Luo[12]·Huang[13]과 같은 계열(임베딩 CBDir/ICVCoh 순위). 우리 6개 차별점을 전문 grep으로 검증 — 행렬 재현성(층②)·ATAC-shuffle 인과· chromatin-lag 감사·외부 rate 앵커·사전등록·MoFlow/CRAK 포함 전부 Wu에 없음 (shuffle 0·permutation 0·preregist 0·MoFlow 0·CRAK 0). ★ Wu의 Data 6 = GSE209878 + GSE284047 = 우리 HSPC 1차 데이터. 미인용 불가. 전략: 반드시 인용 + '무엇이 다른가' 문단 사활 + GB desk-reject 위험 상향 → BIOP01-75 저널 결정에 반영. BIOP01-77에 요지 코멘트 게시. _index는 build_index.py(외부 하네스) 소관이라 직접 편집 안 함 (paper-info.yaml이 정본). manuscript에 임시로 뒀던 SCOOP_CHECK 노트는 이 폴더로 이관하며 삭제. --- .../paper-info.yaml | 126 ++ .../sources/abstract.txt | 1 + .../sources/fulltext_extracted.txt | 1546 +++++++++++++++++ .../sources/paper-doi.url | 2 + .../sources/paper-publisher.url | 2 + .../sources/wu-2026-velocity-benchmark.bib | 10 + .../wu-2026-velocity-benchmark_abstract.md | 18 + .../wu-2026-velocity-benchmark_core.md | 71 + ...u-2026-velocity-benchmark_lens-academic.md | 20 + ...u-2026-velocity-benchmark_lens-industry.md | 12 + ...26-velocity-benchmark_methodology-brief.md | 15 + 11 files changed, 1823 insertions(+) create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/paper-info.yaml create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/abstract.txt create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/fulltext_extracted.txt create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/paper-doi.url create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/paper-publisher.url create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/wu-2026-velocity-benchmark.bib create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_abstract.md create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_core.md create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_lens-academic.md create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_lens-industry.md create mode 100644 paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_methodology-brief.md diff --git a/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/paper-info.yaml b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/paper-info.yaml new file mode 100644 index 0000000..63565ad --- /dev/null +++ b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/paper-info.yaml @@ -0,0 +1,126 @@ +# Wu, Yida, 2026 — velocity-benchmark — Genome Biology +# DOI: 10.1186/s13059-026-04182-z +# Topics: epigenomic-lag (primary), RNA-velocity, single-cell-genomics | Importance: 상 + +# ─── end of header (auto-managed by build_index.py; do not edit this block manually) ─── + +# Identity +document_type: "paper" +title: "Comprehensive benchmarking of RNA velocity methods across single-cell datasets" +authors: + - "Wu, Yida" + - "Kong, Chuihan" + - "Liao, Xu" + - "Lin, Zhixiang" + - "Sun, Xiaobo" + - "Liu, Jin" +year: 2026 +venue: "Genome Biology" +doi: "10.1186/s13059-026-04182-z" +keywords: + - "RNA velocity" + - "benchmark" + - "single-cell" + - "multimodal" + +version: + type: "published" + peer_reviewed: true + published: + date: "2026-07-27" + citation: "Genome Biology 27:242" + preprint: + server: "Research Square" + doi: "10.21203/rs.3.rs-8708834/v1" + +topics: + - "epigenomic-lag" # primary + - "RNA-velocity" + - "single-cell-genomics" + +citation: + key: "wu2026velocitybenchmark" + short_id: "velocity-benchmark-wu" + bibtex_file: "sources/wu-2026-velocity-benchmark.bib" + bibtex_type: "article" + +sources: + paper: + url: "https://genomebiology.biomedcentral.com/articles/10.1186/s13059-026-04182-z" + url_alt: + - "https://link.springer.com/article/10.1186/s13059-026-04182-z" + local: "sources/wu-2026-velocity-benchmark.pdf" + status: "downloaded" + note: "GB open access. 본문 36p. 자동 fetch는 idp.springer.com 쿠키·인증 리다이렉트로 차단 → kkkim이 collab_workspace/references/papers에 업로드한 PDF를 sources/로 반입. fulltext_extracted.txt는 pypdf 추출본(grep용)." + supplementary: + - local: "sources/Additional_file_2.pdf" + type: "pdf" + status: "downloaded" + note: "MOESM2 — supplementary figures(Fig. S1–S17 등). 6.4MB." + - local: "sources/Additional_file_3.pdf" + type: "pdf" + status: "downloaded" + note: "MOESM3 — supplementary tables 추정. 0.2MB." + code: + - url: "<검토필요: 본문 Availability에서 확인>" + type: "github" + status: "url-only" + note: "분석 코드 repo — 전문 Availability of data and materials 절에서 확정 필요." + data: + - url: "https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE209878" + type: "geo-accession" + status: "url-only" + note: "★ 이들 Data 6 = GSE209878 + GSE284047 = 우리 HSPC 1차 데이터와 동일 accession. (Luo 2026이 Dataset12로 쓴 것과 같은 데이터.)" + abstract_text: + source: "publisher-metadata" + local: "sources/abstract.txt" + status: "downloaded" + +categorization: + domain: + - "RNA velocity" + - "single-cell-genomics" + - "benchmark" + use_case: + - "methodology-reference" + - "competitive-landscape" # 우리 원고의 직접 선행 벤치마크(스쿱 점검 대상) + importance: + level: "상" + perspective: "우리 원고(HSPC multiome velocity 출력 신뢰성 감사)의 타깃 저널(Genome Biology)에 2026-07-27 게재된 세 번째 종합 velocity 벤치마크. 우리 1차 데이터(GSE209878)를 Data 6으로 사용. 반드시 인용 + '무엇이 다른가' 문단 필수 + GB desk-reject 위험 상향 → BIOP01-75 저널 결정에 직접 영향. 단 축이 명확히 갈려(임베딩 순위 vs 우리 행렬 재현성·ATAC-shuffle 인과·외부 rate 앵커) 스쿱은 아님." + +related: + - key: "luo2026velocitybenchmark" + relation: "sibling-benchmark" + note: "같은 계열 벤치마크(CRM 2026). 둘 다 CBDir/ICVCoh 임베딩 지표. 우리 원고 refs [12]." + - key: "li2023multivelo" + relation: "cited" + note: "MultiVelo 원본. Wu가 multimodal task에서 평가(GraphVelo(ATAC)에 뒤짐)." + +# 스쿱 판정 (전문 근거) +scoop_check: + verdict: "not-a-scoop" + our_manuscript: "HSPC 10x multiome velocity 출력 신뢰성 감사 (BIOP01)" + differentiators_verified_in_fulltext: + - "cell×gene velocity 행렬의 method 간 재현성(층②) — Wu는 임베딩(CBDir/ICVCoh)만" + - "ATAC-shuffle 인과 대조 — Wu의 negative control은 self-transition/entropy(출력 robustness), 인과 아님. shuffle 0·permutation 0" + - "chromatin→transcription lag 신뢰성·부호 편향 — Wu 주제 아님(lag 1회, 무관)" + - "외부 측정 rate 앵커(TT-seq·half-life) — Wu 없음(metabolic labeling은 method 입력으로만)" + - "사전등록·permutation null·자기철회 방법론 — Wu 없음" + - "MoFlow·CRAK-Velo 포함 — Wu 0회(우리 4 arm 중 둘 미포함)" + strategic: + - "반드시 인용(타깃 저널 게재 + 우리 데이터 사용)" + - "'이전 벤치마크와 무엇이 다른가' 문단 사활" + - "GB desk-reject 위험 상향 → BIOP01-75 저널 결정에 반영" + +workflow: + last_updated: '2026-07-27' + analysis_status: + abstract: "done" + core: "done" + lens_academic: "done" + lens_industry: "done" + methodology_brief: "done" + slides: "skipped" + +# lit-search 발견 맥락 +# 역할: competitive-benchmark — 우리 원고 타깃 저널에 게재된 직접 선행 벤치마크. 스쿱 점검 + 인용·차별화 근거. diff --git a/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/abstract.txt b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/abstract.txt new file mode 100644 index 0000000..c563f52 --- /dev/null +++ b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/abstract.txt @@ -0,0 +1 @@ +RNA velocity provides a powerful framework for inferring cellular dynamics from single-cell RNA sequencing data. The rapid proliferation of compu- tational methods within this field has prompted a need for systematic evaluation. However, existing comparisons often suffer from limited scope or incomplete task design, leaving users without clear guidance. Consequently, there is a lack of a compre- hensive and standardized benchmark that evaluates methods across diverse biological and technical scenarios using appropriate, context-specific metrics. Results: In this study, we present a comprehensive benchmark of 19 computa- tional RNA velocity tools covering 30 distinct methods. We systematically evaluate 25 RNA-only methods across eight evaluation tasks, designating directional con- sistency, temporal precision, negative control robustness, and sequencing depth stability as core tasks, while assessing five multimodal-enhanced methods specifi- cally on the multimodal integration task. These assessments utilize 34 datasets span- ning 26 real-world and eight simulated scenarios. Our results reveal a clear trade-off between directional consistency and negative control robustness, distinct group-wise behaviors across temporal modeling strategies, and variability driven by sequencing depth and quantification choices. This study also identifies several methodological gaps, including the need for improved modeling of gene dependence, more accurate temporal inference strategies, and better-designed multimodal architectures. Conclusions: This benchmark establishes a unified framework for evaluating RNA velocity methods. Crucially, we provide task-aware guidance to facilitate method selec- tion based on specific biological contexts and technical constraints, rather than relying on a single overall ranking. \ No newline at end of file diff --git a/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/fulltext_extracted.txt b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/fulltext_extracted.txt new file mode 100644 index 0000000..1a782d2 --- /dev/null +++ b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/sources/fulltext_extracted.txt @@ -0,0 +1,1546 @@ +Open Access +© The Author(s) 2026. Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 Inter- +national License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you +give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified +the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The +images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a +credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by +statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of +this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/. +RESEARCH +Wu et al. Genome Biology (2026) 27:242 +https://doi.org/10.1186/s13059-026-04182-z +Genome Biology +Comprehensive benchmarking of RNA +velocity methods across single-cell datasets +Yida Wu1†, Chuihan Kong1†, Xu Liao1, Zhixiang Lin2,3,4, Xiaobo Sun5* and Jin Liu1* +Abstract +Background: RNA velocity provides a powerful framework for inferring cellular +dynamics from single-cell RNA sequencing data. The rapid proliferation of compu- +tational methods within this field has prompted a need for systematic evaluation. +However, existing comparisons often suffer from limited scope or incomplete task +design, leaving users without clear guidance. Consequently, there is a lack of a compre- +hensive and standardized benchmark that evaluates methods across diverse biological +and technical scenarios using appropriate, context-specific metrics. +Results: In this study, we present a comprehensive benchmark of 19 computa- +tional RNA velocity tools covering 30 distinct methods. We systematically evaluate +25 RNA-only methods across eight evaluation tasks, designating directional con- +sistency, temporal precision, negative control robustness, and sequencing depth +stability as core tasks, while assessing five multimodal-enhanced methods specifi- +cally on the multimodal integration task. These assessments utilize 34 datasets span- +ning 26 real-world and eight simulated scenarios. Our results reveal a clear trade-off +between directional consistency and negative control robustness, distinct group-wise +behaviors across temporal modeling strategies, and variability driven by sequencing +depth and quantification choices. This study also identifies several methodological +gaps, including the need for improved modeling of gene dependence, more accurate +temporal inference strategies, and better-designed multimodal architectures. +Conclusions: This benchmark establishes a unified framework for evaluating RNA +velocity methods. Crucially, we provide task-aware guidance to facilitate method selec- +tion based on specific biological contexts and technical constraints, rather than relying +on a single overall ranking. +Keywords: Benchmarking, RNA velocity, Single-cell RNA sequencing, Splicing +dynamics, Multimodal +Background +The advent of single-cell RNA sequencing (scRNA-seq) has revolutionized the study +of cellular differentiation and development at single-cell resolution [1 , 2]. Despite the +fact that scRNA-seq yields only static snapshots of cellular states  [3 , 4], trajectory +†Yida Wu and Chuihan Kong +contributed equally to this work. +*Correspondence: +xiaobo.sun@emory.edu; +liujinlab@cuhk.edu.cn +1 School of Data Science, The +Chinese University of Hong +Kong, Shenzhen, Shenzhen, +China +2 Department of Statistics +and Data Science, The Chinese +University of Hong Kong, Shatin, +Hong Kong SAR, China +3 Shenzhen Loop Area Institute, +Shenzhen, Guangdong, China +4 CUHK Shenzhen Research +Institute, Shenzhen, Guangdong, +China +5 Department of Human +Genetics, School of Medicine, +Emory University, Atlanta, GA, +USA +Page 2 of 36Wu et al. Genome Biology (2026) 27:242 +inference (TI) methods were developed to reconstruct the temporal ordering of cells +based on these gene expression profiles  [5 –7]. However, TI methods are subject to +two primary limitations. First, they rely on dense sampling of all intermediate states +to reconstruct continuous trajectories, limiting their applicability when data cover - +age is incomplete [8 ]. Second, the dependence on prior information, such as defined +starting cells, introduces subjectivity that can lead to multiple equally plausible solu - +tions  [9]. In contrast to traditional TI methods that infer global ordering between +cells, RNA velocity represents a breakthrough in predicting the future state of indi - +vidual cells [10]. By modeling the kinetics of unspliced and spliced mRNAs, it cap - +tures intrinsic cellular dynamics without relying on external prior knowledge [11]. +Practically, a standard RNA velocity computation pipeline consists of three distinct +stages: pre-processing, velocity estimation, and post-processing (Fig.  1a, left panel). +In the pre-processing stage, unspliced and spliced mRNA abundances are quanti - +fied from sequencing reads using established tools such as velocyto [ 10], dropEst [ 12], +and alevin [13]. Following quality control, selection of highly variable genes (HVGs), +and optional normalization and smoothing, the final spliced and unspliced expres - +sion matrices are generated. Next, RNA velocity is inferred by modeling the dynam - +ics between unspliced and spliced mRNAs. Specifically, these methods capture the +kinetics of the mRNA lifecycle, encompassing transcription, splicing, and even - +tual degradation. This innovation has triggered a surge in methodological develop - +ment, leading to the emergence of 17 RNA-only tools since the pioneering work of +La Manno et  al.  [10] (Fig.  1a, right panel). In addition, nine multimodal-enhanced +tools have been developed that incorporate auxiliary information, such as chromatin +accessibility [14] and metabolic labeling [15]. In the post-processing stage, a cell-to- +cell transition matrix is computed from the velocity matrix using a cosine similarity +kernel [10, 16]. Finally, velocities in the original gene space are projected into lower- +dimensional embeddings, such as principal component analysis (PCA) or uniform +Fig. 1 Overview of RNA velocity pipeline, benchmarking tasks and datasets. a Left: The standard RNA velocity +computation pipeline consists of three stages: pre-processing, velocity estimation, and post-processing. +During pre-processing, unspliced and spliced mRNA abundances are quantified from raw FASTQ files. +Following filtering, highly variable gene (HVG) selection, and optional normalization and smoothing, +unspliced and spliced count matrices ( n × p ) are generated, where n denotes the number of cells and p the +number of genes. Steps shown in parentheses indicate optional operations. Velocity estimation is based +on a continuous model of transcription, splicing, and degradation with rates α , β , and γ , respectively, from +which an n × p velocity matrix ( ds /dt ) is derived. Post-processing converts velocities into an n × n cell-to-cell +transition matrix and projects them onto low-dimensional embeddings for visualization or evaluation. Right: +Timeline of 17 RNA-only and nine multimodal-enhanced RNA velocity tools developed between 2018 +and 2025. b Eight benchmarking tasks are defined, comprising four core tasks: (I) directional consistency, +evaluated by cross-boundary directional correctness (CBDir) and in-cluster coherence (ICVCoh); (II) temporal +precision, quantified by cluster temporal ordering accuracy (CTO) and temporal Spearman correlation (TSC); +(III) negative control robustness, measured by self-transition score (STS) and effective entropy score (EES); and +(IV) sequencing depth stability, assessed by the stability score for direction ( SCBDir ) and cell-specific latent +time ( STSC ). Four additional tasks are also included: (V) quantification stability; (VI) simulation experiments; +(VII) multimodal integration; and (VIII) computational scalability. c Summary of the 26 real-world and eight +simulated datasets, including dataset identifiers, names, species, associated benchmarking tasks, and cell/ +gene counts. Bar plots show the numbers of cells and genes for each dataset. Gene counts for Data 31 and +Data 32 are averages across random-seed replicates (217–221 and 231–235 genes, respectively) +(See figure on next page.) +Page 3 of 36 +Wu et al. Genome Biology (2026) 27:242 + +manifold approximation and projection (UMAP) [17, 18], for visualization and down - +stream evaluation. +The rapid development of RNA velocity methods has prompted numerous compara - +tive studies; however, existing reviews are largely confined to theoretical discussions, +focusing primarily on model assumptions and limitations [19, 20]. Workflow-level stud- +ies supported by empirical data analyses have further revealed failure modes in RNA +velocity inference, including false-positive velocity fields, technical variation, hyper - +parameter sensitivity, k-nearest-neighbor (kNN) smoothing, and low-dimensional +Fig. 1 (See legend on previous page.) +Page 4 of 36Wu et al. Genome Biology (2026) 27:242 +mapping [21–23]. Nevertheless, these discussions have mainly focused on classical RNA +velocity methods, namely Velocyto  [10] and scVelo  [11]. Benchmarking studies have +expanded this perspective by systematically comparing a broader range of RNA velocity +methods across practical application scenarios. Although Ancheta et  al.  [24] assessed +RNA velocity methods across four tasks, including local consistency and driver-gene +identification, their study was limited in scale, evaluating only five methods across +three datasets. More recent comprehensive benchmarks by Luo et  al.  [25] and Huang +et  al.  [26] evaluated 15 and 29 RNA velocity methods, respectively, but did not sepa - +rately assess method variants, such as those of UniTVelo  [27] and Pyro-Velocity  [28]. +Since such variants rely on different model assumptions and show distinct task-specific +performance, using a single default variant to represent an entire method may obscure +practical differences and mislead method selection. In addition, these benchmarks pri - +marily assessed accuracy through directional correctness and velocity coherence, while +overlooking robustness in negative-control settings and key technical factors such as +variation in sequencing depth and pipelines for spliced/unspliced mRNA quantification, +all of which can substantially influence velocity estimates [21, 22, 24, 29]. +To address this gap, we established a comprehensive benchmark of 19 computational +RNA velocity tools covering 30 methods. Specifically, we systematically evaluated 25 +RNA-only methods across eight benchmarking tasks, categorized into four core (Fig.  1b, +top panels) and four additional assessments (Fig.  1b, bottom panels), while conducting a +targeted assessment of five multimodal-enhanced methods in the multimodal integra - +tion task. These evaluations utilized 34 datasets encompassing 26 real-world and eight +simulated scenarios (Methods ; Fig.  1c; Additional file 1: Tables S1–S2). Because meth - +ods exhibit heterogeneous performance across core tasks, selecting methods based +solely on global rankings can be misleading. We therefore offer task-aware guidance to +support informed method choices tailored to specific biological contexts and technical +constraints. +Results +Overall benchmarking of 25 RNA‑only methods +Before benchmarking, we systematically summarized the 25 RNA-only methods, which +included the 22 splicing dynamics–based methods and three scTour [30] variants that +uniquely leverage gene expression instead of spliced/unspliced transcript counts. First, +with respect to model architecture (Fig.  2, Col.  2), we categorized the methods into +11 deep learning (DL) and 14 non-DL approaches. The DL category is dominated by +variational autoencoder (VAE)-based generative frameworks, including veloVI  [31], +scTour (MSE/NB/ZINB), SvelvetVAE [32], LatentVelo (std) [33], and VeloVAE (std/Full +VB) [34]. By contrast, VeloAE [35], DeepVelo [36], and cellDancer [37] utilize non-gen - +erative architectures. The remaining non-DL approaches include Velocyto, scVelo (dyn/ +stc), UniTVelo (uni/ind), κ-velo  [38], Dynamo (m1)  [15], SDEvelo  [39], Pyro-Velocity +(m1/m2), cell2fate [40], TIVelo (std/simple) [41], and GraphVelo (std) [42]. Second, we +categorized output capabilities (Fig. 2, Col. 3), distinguishing 16 methods that infer both +velocity and cell-specific latent time (CLT) from nine methods that output only veloc - +ity vectors. In terms of implementation (Fig.  2, Col.  4), all methods are Python-based, +although Velocyto also provides an R package. Within this Python ecosystem, PyT orch +Page 5 of 36 +Wu et al. Genome Biology (2026) 27:242 + +predominates, with 14 direct implementations and three additional methods using its +Pyro extension. The remaining methods are implemented using either TensorFlow or +standard Python. Lastly, we stratified the methods into five tiers based on their scalabil - +ity with respect to cell and gene counts, where stronger signal intensity denotes superior +scalability (Fig. 2, Cols. 5–6). Detailed descriptions and implementations of each method +are provided in Methods. +Following this methodological overview, we defined four core benchmarking tasks. +(I) Directional consistency evaluates biological plausibility using cross-boundary direc - +tional correctness (CBDir) and in-cluster coherence (ICVCoh). (II) Temporal precision +assesses the accuracy of inferred CLT using cluster temporal ordering accuracy (CTO) +Fig. 2 Comprehensive benchmarking of 25 RNA-only methods. Method: The first column lists method +names. The second and third columns indicate the use of deep learning (DL) and support for cell-specific +latent time (CLT) output, respectively; for both columns, a black star denotes presence and a white star +denotes absence. The fourth column displays icons identifying the implementation platform (PyTorch, +TensorFlow, Pyro, Python, or R). Scalability: Cell scalability and gene scalability were measured on two +large-scale simulated datasets with up to 1,000,000 cells and 20,000 genes. For cell scalability, the five tiers +of signal intensity indicate successful completion by each method on datasets with 10,000, 50,000, 100,000, +500,000, and 1,000,000 cells, with the gene count fixed at p = 1, 000 . For gene scalability, the five tiers of +signal intensity indicate the methods completed at 2,000, 5,000, 10,000, 15,000, and 20,000 genes with +fixed cell counts ( n = 10, 000 ). Directionality: Performance in the directional consistency task (Task I), CBDir +and ICVCoh. Temporality: Performance in the temporal precision task (Task II), evaluated using CTO and +TSC. Neg. Ctrl.: Performance in the negative control robustness task (Task III), evaluated using STS and EES. +Stability: Performance in the sequencing depth stability task (Task IV), evaluated using SCBDir and STSC . For +each task, bars represent mean metric values across datasets, and error bars indicate standard deviations. Rk +( k ∈{ 1,... ,4} ) denotes the rank for Tasks I to IV, respectively (Methods). Rcore denotes the overall rank of the +four core tasks (Methods). Methods are listed in ascending order of Rcore +Page 6 of 36Wu et al. Genome Biology (2026) 27:242 +and temporal Spearman correlation (TSC). (III) Negative control robustness evaluates +the ability to avoid spurious dynamics in mature cell populations using the self-transi - +tion score (STS) and the effective entropy score (EES). (IV) Sequencing depth stability +quantifies performance robustness under binomial subsampling noise using stability +scores for CBDir ( SCBDir ) and TSC ( STSC ). Detailed definitions of all metrics are pro - +vided in Methods . T o aggregate performance across tasks, we averaged the two metric- +level ranks to obtain a single rank per task and defined the final ranking for each method +as the average of these four task-level ranks (Methods). +We applied this evaluation framework to the 25 RNA-only methods to obtain com - +prehensive, task-specific performance rankings (Fig.  2, Cols. 7–18). We found clear dif - +ferences in performance profiles among the top-ranked methods. Specifically, SDEvelo +(2nd) and Velocyto (3rd) performed well on Tasks II–IV but showed comparatively +weaker performance on Task I. By contrast, UniTVelo (uni) (1st) exhibited more bal - +anced performance, ranking within the top 15 across all core tasks. Regarding task- +specific strengths, LatentVelo (std), UniTVelo (uni), SDEvelo, and Pyro-Velocity (m2) +ranked first in Task I–Task IV , respectively. We further examined correlations between +method rankings across tasks (Additional file 2: Fig. S1) and found that a significant cor - +relation exists only between directional consistency and negative control robustness +(Spearman’s ρ =− 0.572 , P = 0.003 ). This negative correlation suggests that approaches +optimized for strong directional signals may compromise the ability to suppress false +positives. LatentVelo (std) exemplifies this trade-off: although it ranked first in direc - +tional consistency, it dropped to 23rd in negative control robustness. This pattern sug - +gests that an emphasis on dynamic inference can, in some cases, lead to the generation +of spurious velocity signals. +The overall rankings provide a high-level summary but do not capture the full com - +plexity of method performance. T o address this, we examine each evaluation task indi - +vidually in the following sections, clarifying the benchmarking metrics, identifying +patterns that recur across datasets, and linking specific modeling choices to their cor - +responding strengths and weaknesses. +Directional consistency performance comparison +A primary goal of RNA velocity methods is to reconstruct biologically meaningful tran - +sitions between cell types or developmental stages  [10]. To achieve this, ideal velocity +vectors are expected to satisfy two key criteria. First, they should point in the correct +biological direction, particularly at the boundaries between adjacent cell states. Second, +they should exhibit consistent behavior among phenotypically similar cells. To capture +these complementary aspects of velocity quality, we focused on directional correctness +at cell-type boundaries and velocity coherence within clusters [35], quantified using the +widely adopted CBDir and ICVCoh metrics, respectively. In this task, both metrics were +computed using UMAP-projected velocity vectors according to the rationale detailed in +Additional file 3: Note S1 and Additional file 2: Fig. S2. Our evaluation included ten data- +sets (Data 1–10) across different species, including human and mouse, as well as diverse +tissues, developmental settings, sequencing technologies, and batch-effect scenarios, all +with well-established developmental trajectories (Fig. 3a; Additional file 1: Table S3). +Page 7 of 36 +Wu et al. Genome Biology (2026) 27:242 + +Fig. 3 Directional consistency performance comparison. a Performance evaluation using CBDir and ICVCoh +on Data 1–10. Methods are ordered by mean metric values across datasets, such that stronger performers +appear in the top rows for each metric. GD denotes modeling of gene dependence, and a black star indicates +that the method incorporates gene dependence. Bars and the color scale represent the mean metric values +across cross-validation folds for each dataset, with darker colors indicating larger values. Error bars indicate +standard deviations across cross-validation folds. b Streamline visualizations for UniTVelo (ind), veloVI, VeloAE +and UniTVelo (uni) on Data 1 and Data 3. The leftmost column shows the reference cell-state transitions. +Each of the remaining four columns corresponds to one method. c Streamline visualizations of Velocyto and +scVelo (stc) on Data 7 (scRNA-seq), with the reference cell-state transitions shown below. Cell types include +cycling progenitors (Cyc. Prog.), early radial glia (Early RG), late radial glia (Late RG), neurogenic intermediate +progenitor cells (nIPC), and glutamatergic neurons (GluN). d Streamline visualizations of Velocyto and scVelo +(stc) on Data 8 (snRNA-seq), with the reference cell-state transitions shown below. For panels (b)–(d), average +values across cross-validation folds for CBDir (left) and ICVCoh (right) are shown below each streamline plot +Page 8 of 36Wu et al. Genome Biology (2026) 27:242 +Modeling gene dependence is crucial for RNA velocity inference, as it aims to cap - +ture cell-state transitions together with their underlying regulatory interactions  [21]. +Accordingly, many methods model genes jointly using either low-dimensional embed - +dings or a unified latent time (Fig.  3a, left panel, Col. 2). However, our analysis shows +that extending models to the multivariate setting does not necessarily result in improved +directional accuracy (Additional file 2: Fig. S3a, left panel). Consistent with this observa - +tion, enrichment analysis indicates that methods incorporating gene dependence are not +significantly overrepresented among the top 10 (Additional file  3: Note  S2; Additional +file 2: Fig. S3a, right panel). By contrast, results based on velocity coherence measured by +ICVCoh reveal a clear advantage for gene-dependent approaches (Fig.  3a, right panel). +Methods incorporating gene dependence significantly outperformed gene-independent +methods (Additional file 2: Fig. S3b, left panel) and were enriched among the top 10 per- +formers in ICVCoh, whereas gene-independent methods were overrepresented in the +bottom 10 (Additional file  2: Fig.  S3b, right panel). This trend was further supported +by pairwise comparisons between related methods. For example, UniTVelo (uni) (3rd) +clearly outperformed its independent counterpart (10th), and veloVI (8th) surpassed the +gene-independent scVelo (dyn) (14th), despite sharing the same underlying dynamical +model. Collectively, these results indicate that while gene dependence does not guaran - +tee superior directional accuracy, it confers a substantial advantage in generating coher - +ent RNA velocity fields. +In addition to assessing mean performance, we evaluated variation in method perfor - +mance across datasets. When analyzing CBDir, each method failed on at least one data - +set, and instability was evident even among top-performing approaches, as illustrated +by streamline visualizations of selected methods (Fig.  3b). For example, UniTVelo (ind), +veloVI, and VeloAE predicted reversed dynamics between hematopoietic stem cells and +erythroids on Data 1 (Fig.  3b, top panel), whereas they correctly inferred erythroid tra - +jectories on Data 3 (Fig.  3b, bottom panel). Conversely, UniTVelo (uni) captured the +correct transitions on Data 1 but performed poorly on Data 3. Cross-dataset variation +was more pronounced in technically challenging settings, including batch effects (Data +5 and 6; Additional file 2: Fig. S4a) and altered spliced-to-unspliced transcript ratios in +single-nucleus RNA sequencing (snRNA-seq; Data 8). LatentVelo (std) and all scTour +variants were the only methods whose original studies claimed batch-effect mitigation +during dynamic inference [30, 33, 43, 44] (Additional file 3: Note S3; Additional file 2: +Fig. S4b). To assess how batch effects influence velocity inference, we split Data 5 and +Data 6 into six single-batch subsets and compared the resulting method rankings with +those obtained from the original batch-confounded datasets (Additional file 2: Fig. S5, +left panel). LatentVelo (std), which explicitly accounts for batch effects, remained con - +sistently strong across both settings. Among methods without explicit batch-effect +modeling, cellDancer and TIVelo (simple) showed clear rank declines under batch-con - +founded settings, whereas cell2fate and SDEvelo showed substantial rank improvements +(Additional file 2: Fig. S5, right panel). Beyond batch effects, Data 8, the only snRNA- +seq dataset in Task  I, presented an additional technical challenge in which unspliced +counts exceeded spliced counts; the opposite pattern was observed in all non-snRNA- +seq datasets (Additional file 2: Fig. S6). This difference was particularly evident between +Data 7 and Data 8, two human brain datasets from the same species and tissue type but +Page 9 of 36 +Wu et al. Genome Biology (2026) 27:242 + +generated using different sequencing strategies. Compared with performance on Data 7, +Velocyto and scVelo (stc) showed markedly lower directional-consistency ranks on Data +8 (Additional file 2: Fig. S7a) and produced largely incorrect or reversed flows (Fig.  3c +and d), likely because the higher fraction of unspliced counts in nuclei affects gene-spe - +cific steady-state relationships (Additional file 2: Fig. S7b). Although scTour variants do +not use spliced or unspliced counts, their inferred trajectories can still be completely +reversed in practice, an issue also noted in the original study [30]. Together, these results +indicate that directional consistency is strongly influenced by dataset-specific technical +factors, including batch effects and sequencing modality. +Evaluation for temporal precision +Cell-specific latent time (CLT) is a critical output of RNA velocity methods, providing +the temporal ordering essential for downstream biological interpretation  [11]. Mean - +while, biologically measured ground-truth temporal labels across diverse settings, +including cell development (Data 3, 13, and 14), metabolic labeling (Data 10), the cell +cycle (Data 11), lineage tracing (Data 12), and cellular reprogramming (Data 15 and +16), enable direct comparison of CLT with ground-truth time and strengthen evalua - +tion reliability. In this section, we categorized the 25 RNA velocity methods into four +groups based on their temporal modeling strategies (Fig.  4a). Specifically, the first two +groups perform simultaneous inference, estimating velocity alongside either a unified +CLT (Group 1) or gene-specific latent times (Group 2). The remaining groups rely on +post-hoc reconstruction, utilizing either velocity graph–based (Group 3) or dynamical +model–based (Group 4) approaches. Details of CLT implementation for each group are +provided in Additional file 3: Note S4, including how outputs from Group 2 were stand - +ardized to a unified CLT to enable fair comparison. Performance was quantified using +CTO and TSC metrics to assess temporal accuracy at both cluster and single-cell levels +(Fig. 4b; Additional file 1: Table S4). +We first evaluated the precision of inferred temporal ordering at the cluster level using +CTO across all eight datasets. Based on this analysis, UniTVelo (uni) emerged as the top +method on average, ranking first in Data 13 and Data 14 (Fig.  4c, top panel; Additional +(See figure on next page.) +Fig. 4 Evaluation for temporal precision. a Classification of 25 RNA velocity methods into four groups +based on their temporal modeling strategies: (1) simultaneous modeling of velocity and a unified CLT; (2) +simultaneous modeling of velocity and per-gene, per-cell latent time; (3) velocity graph–based inference +following velocity estimation; and (4) model-based inference following velocity estimation. Methods +enclosed in the dashed box (Group 3) do not natively output CLT; therefore, a standardized post-hoc strategy +was applied (Additional file 3: Note S4). b Schematic illustrating temporal precision evaluation by comparing +inferred continuous CLT with discrete ground-truth temporal labels (e.g., reprogramming days) using CTO +and TSC. c Performance evaluation using CTO (top) and TSC (bottom) across eight real-world datasets (Data 3 +and Data 10–16). Boxplots are colored according to the groups defined in (a) and ordered by the average +rank across datasets (y-axis). Individual points represent scores for specific cross-validation folds in each +dataset. The x-axes indicate metric scores, ranging from 0 to 1 for CTO and −1 to 1 for TSC. d Correlation +between the average CTO rank and TSC rank for all 25 velocity methods. e Comparison of directional +consistency rank versus temporal precision rank. For panels (d) and (e), Spearman’s ρ and the associated +P-value between the two rankings are shown above each panel. f Method rank differences between Task I +and II. Positive rank differences indicate that a method achieved a better rank in temporal precision than in +directional consistency, whereas negative values indicated the opposite. For panels (d)–(f), method colors are +the same as in c +Page 10 of 36Wu et al. Genome Biology (2026) 27:242 +file 2: Fig. S8). To complement the cluster-level metric, we assessed CLT accuracy at the +single-cell level using TSC (Fig.  4c, bottom panel; Additional file 2: Fig. S9). Although +rankings differed between the two metrics across individual datasets, the average rank - +ings of the methods were highly consistent (Spearman’s ρ = 0.925 , P = 3.603 × 10 −11 ; +Fig.  4d, left panel). Several methods, including UniTVelo (uni), DeepVelo, SDEvelo, +Velocyto, Dynamo (m1), and SvelvetVAE, showed strong performance across both CTO +and TSC (Fig.  4d, right panel). By contrast, methods with joint temporal modeling did +not consistently achieve strong temporal precision. Among the eight Group 1 methods +that jointly infer both velocity and a unified CLT, only UniTVelo (uni) ranked within the +Fig. 4 (See legend on previous page.) +Page 11 of 36 +Wu et al. Genome Biology (2026) 27:242 + +top 10 for either CTO or TSC. None of the Group 2 methods, which jointly infer velocity +and gene-specific latent times, ranked within the top 10 for either metric. +We next integrated CTO and TSC into a unified temporal precision rank and com - +pared it with the directional consistency rank obtained in Task  I (Fig.  4e). UniTVelo +(uni) was the only method that ranked within the top five for both directional consist - +ency and temporal precision. Temporal precision showed no significant correlation with +directional consistency, suggesting that high directional-consistency performance in +Task I does not necessarily imply high temporal-precision performance in Task II. This +decoupling was particularly evident among five of the Group 3 methods that infer CLT +using post-hoc reconstruction based on velocity graphs (Fig.  4e). Although these meth - +ods showed similar directional consistency ranks (10th–15th), their temporal precision +ranks varied widely, ranging from 2nd to 25th. We further compared method rank differ- +ences between Task I and Task II (Fig.  4f). Notably, the seven methods with the largest +positive rank differences, indicating better ranks in temporal precision than in direc - +tional consistency, were all post-hoc methods from Groups 3 and 4. By contrast, several +of the Group 1 methods that jointly infer velocity and CLT showed negative rank differ - +ences, thus indicating the opposite pattern. These results suggest that, in terms of rela - +tive method rankings, several post-hoc methods can achieve higher temporal precision +than directional-consistency ranks after CLT reconstruction. We further examined this +pattern using Data 10, a dataset included in both Task I and Task II, to avoid potential +dataset-level confounding (Additional file 2: Fig. S10). Similarly, several Group 3 meth - +ods achieved substantially better temporal-precision ranks than directional-consistency +ranks, whereas several Group 1 methods showed the opposite trend. +Assessment of negative control robustness +In contrast to developing tissues, mature cell populations predominantly consist of +genes in a steady state and therefore provide little to no information about ongoing cel - +lular dynamics [21]. In such stable systems, an ideal RNA velocity method should avoid +inferring spurious directional flow and instead produce weak or near-random velocity +Fig. 5 Assessment of negative control robustness and sequencing depth stability. a Illustration of negative +control robustness. Fully mature cells are predominantly in a steady state and should not exhibit meaningful +velocity signals. The streamline plot in the green square shows the behavior of a robust RNA velocity method, +where no clear velocity pattern is present. The plot in the red square shows an example of erroneous velocity +estimation. b Streamline visualizations of cell2fate and Pyro-Velocity (m1) on Data 17 (negative control) +and Data 4 (positive control). The average EES values across cross-validation folds are shown below each +streamline plot. c Scatter plot of each method’s EES rank versus STS rank. Each point represents one RNA +velocity method listed in the legend. Spearman’s ρ and the associated P-value for the two rankings are +shown above the panel. d Low-dimensional embeddings colored by uncertainty score for VeloVAE (std) +and cell2fate. Each row corresponds to one method. Uncertainty scores are min-max scaled independently +for each method across six datasets (Data 4 and Data 17–21). PBMC datasets are marked with green or red +rectangles, where green indicates robust uncertainty estimation and red indicates problematic estimation. +The superscript asterisk (*) indicates that only one fold was used for the dataset. e Illustration of sequencing +depth simulation. For both spliced and unspliced abundance matrices, each count is subsampled using +a binomial distribution with sampling rates of 1.0, 0.8, 0.6, 0.4, and 0.2. Smaller sampling rates represent +shallower sequencing depth. f Scatter plot of average SCBDir versus STSC for each method across datasets to +assess sequencing depth stability. Spearman’s ρ and the associated P-value for the two stability scores are +shown above the panel. The legend is shared with panel (c) +(See figure on next page.) +Page 12 of 36Wu et al. Genome Biology (2026) 27:242 +patterns (Fig.  5a). However, existing methods are often prone to generating false-posi - +tive signals under these conditions [21, 22]. To benchmark this limitation, we evaluated +the negative control robustness of 25 methods using five datasets of peripheral blood +mononuclear cells (PBMCs; Data 17–21), as these fully differentiated cells lack meaning- +ful dynamic information. For this task, we used two complementary metrics (Additional +file 1: Table S5). First, we used STS to quantify the average probability of a cell transi - +tioning to itself. Because static cells should theoretically not transition to neighboring +states, a high STS indicates that a method correctly identifies steady-state behavior. Sec - +ond, in the absence of true biological dynamics, a robust method should yield diffuse, +near-uniform transition probabilities rather than confident but spurious trajectories. We +Fig. 5 (See legend on previous page.) +Page 13 of 36 +Wu et al. Genome Biology (2026) 27:242 + +therefore introduced EES to quantify the sparsity and uniformity of transition probabili - +ties, with higher values indicating stronger robustness under negative control conditions +(Additional file 2: Fig. S11). We validated EES using Data 4 as a positive-control data - +set and five datasets of PBMCs as negative controls. Most methods showed higher EES +scores on the negative controls than on Data 4, as expected (Additional file 2: Fig. S12). +For example, on Data 17, cell2fate produced nearly random transitions with a high EES +score (334.54), while Pyro-Velocity (m1) showed stronger but potentially spurious direc- +tional patterns with a lower EES score (9.51; Fig.  5b, top panel). By contrast, both meth - +ods showed low EES scores on Data 4, consistent with clear differentiation directions +(Fig. 5b, bottom panel). +We then integrated EES and STS to identify methods with strong performance across +both metrics (Fig. 5c). SDEvelo, cell2fate, κ-velo, and DeepVelo ranked within the top 10 +for both metrics, whereas LatentVelo (std), SvelvetVAE, GraphVelo (std), and the three +scT our variants consistently underperformed. Notably, four of these six low-ranking +methods utilize neural ordinary differential equations (ODEs) [45] to enforce temporal +ordering. This pattern suggests that imposing rigid temporal constraints during infer - +ence, without adequately accounting for stochastic variability, may undermine robust - +ness in steady-state settings. By contrast, methods that infer CLT only after estimating +kinetic parameters, such as SDEvelo, or that avoid explicit time parameterization alto - +gether, such as κ-velo, appear better equipped to mitigate spurious dynamics. +In addition to the metrics introduced here, several Bayesian methods use uncertainty +scores as a form of quality control, interpreting high uncertainty as indicative of weak or +inconsistent transcriptional activity [31]. These scores are quantified either by the coef - +ficient of variation of posterior cell times [34, 40], as implemented in cell2fate and Velo- +VAE (std/Full VB), or by variability in predicted future states [31], as in veloVI. T o assess +the reliability of these uncertainty measures, we compared a dynamic reference data - +set (Data 4) with steady-state PBMC datasets (Data 17–21) (Additional file 2: Fig. S13). +Although all methods assigned higher uncertainty to Data 17 relative to Data 4, each +failed to do so for at least one other PBMC dataset. For example, VeloVAE (std) failed on +Data 18, while cell2fate failed on Data 21 (Fig.  5d). T ogether, these results highlight the +limited reliability of current uncertainty-based measures in static tissue contexts. +Sequencing depth stability +An optimal RNA velocity method should be robust to technical and experimental noise +that is unrelated to true biological dynamics [40]. Because variation in sequencing depth +represents a major source of such noise [46], robustness to reduced sequencing depth +serves as a key benchmark of method reliability [24]. To evaluate this, we simulated pro - +gressively shallower datasets by subsampling spliced and unspliced mRNA counts using +a binomial distribution, with sampling rates ranging from 1.0 (original data) to 0.2 (shal - +lowest depth) (Fig. 5e) [47]. We then quantified the stability of directional accuracy using +SCBDir , and temporal precision using STSC . A high stability score suggests that a method +maintains consistently high performance with low variance on CBDir or TSC across +sequencing depths. In this section, we assessed directional stability using Data 2, 3, and +9, and temporal stability using Data 3, 11, and 13 (Additional file 1: Table S6). +Page 14 of 36Wu et al. Genome Biology (2026) 27:242 +As shown in Fig.  5f, a subset of methods exhibited strong robustness to sequencing- +depth variation. Specifically, Pyro-Velocity (m2), scVelo (stc), SDEvelo, and cellDancer +maintained high stability in both directional accuracy and temporal precision regard - +less of the sampling rates. Conversely, scTour (MSE) and GraphVelo (std) emerged as +clear outliers, showing highly unstable and weak performance in both dimensions. To +further examine the impact of sequencing depth on method performance, we analyzed +how CBDir and TSC varied across sampling rates for all methods. Unexpectedly, we did +not observe a monotonic decline in performance for either CBDir (Additional file  2: +Fig. S14a) or TSC (Additional file 2: Fig. S14b) as sequencing depth decreased. This find - +ing suggests that sequencing depth influences RNA velocity performance in more com - +plex ways than simple linear degradation. +Quantification stability +In addition to the core tasks, we extended our benchmarking framework to four com - +plex scenarios. (V) Quantification stability evaluates the robustness of RNA velocity +methods to variation in spliced and unspliced abundances arising from different quanti - +fication algorithms. (VI) Simulation experiments benchmark methods using diverse syn- +thetic datasets generated from dynamic models. (VII) Multimodal integration assesses +methods that incorporate additional biological modalities beyond spliced and unspliced +mRNA. (VIII) Computational scalability is evaluated by monitoring running time and +peak memory usage across varying numbers of genes and cells. Although these tasks +were excluded from the overall ranking (rationale detailed in Additional file 3: Note S5), +we conducted a comprehensive evaluation for each. In this section, we focus specifically +on quantification stability by examining five representative quantification algorithms, +namely alevin, dropEst, kallisto|bustools [48], STARsolo [49], and velocyto, which gener - +ate the initial spliced and unspliced abundance matrices [21, 29] (Fig. 6a). +Because quantification algorithms specifically affect spliced and unspliced matrices, +we evaluated the 22 splicing dynamics–based methods on Data 4 using outputs from +the five quantification algorithms [29]. Although these pipelines produced different pro - +portions of unspliced and spliced counts (Fig.  6b) and led to minor embedding shifts +among terminal cell types (alpha, beta, delta, and epsilon), the relative positioning of +pre-endocrine progenitors with respect to these endpoints was preserved (Fig.  6c). This +observation confirms that the underlying biological structure remains largely consistent +across quantification methods. A robust RNA velocity method should therefore toler - +ate quantification-induced noise and recover comparable trajectories irrespective of the +upstream quantification pipeline. +Using stability scores for CBDir ( SCBDir ) and ICVCoh ( SICVCoh ) (Fig. 6d; Additional +file 1: Table S7), we identified UniTVelo (ind), veloVI, and Pyro-Velocity (m2) as three +methods that demonstrated high stability across both metrics. T o contextualize these +results, we examined velocity streamline visualizations across quantification outputs +(Fig. 6e; Additional file 2: Fig. S15). Despite the noise introduced by different quantifica - +tion algorithms, UniTVelo (ind) consistently recovered differentiation trajectories from +pre-endocrine progenitors to terminal cell types. By contrast, many methods showed +robustness in either CBDir or ICVCoh alone, while cell2fate and cellDancer exhibited +Page 15 of 36 +Wu et al. Genome Biology (2026) 27:242 + +poor stability across both metrics. Collectively, these results indicate that robustness to +quantification-level variability remains an ongoing challenge for current RNA velocity +methods. +Fig. 6 Evaluation of quantification stability. a Illustration of abundance quantification from FASTQ files to +spliced and unspliced counts. b Proportions of unspliced and spliced counts for 100 common genes sampled +from Data 4 using different quantification algorithms: alevin, dropEst, kallisto|bustools, STARsolo, and velocyto. +Using velocyto as the baseline, genes are sorted in descending order of their unspliced proportions. c UMAPs +of spliced abundances quantified by the five algorithms. d Scatter plot of SICVCoh versus SCBDir across five +quantification algorithms for Data 4. Each point represents one RNA velocity method listed in the legend. +Spearman’s ρ and the associated P-value for the two stability scores are shown above the panel. e Streamline +visualizations of UniTVelo (ind) across five quantification algorithms for Data 4. The legend is shared with (c), +and the top-left panel shows the reference cell-state transitions. CBDir and ICVCoh values are provided below +each streamline plot on the left and right, respectively +Page 16 of 36Wu et al. Genome Biology (2026) 27:242 +Benchmarking analysis on simulation experiments +To complement real-world datasets, we evaluated 25 RNA velocity methods using six +simulated datasets generated from ODE and stochastic differential equation (SDE) mod- +els, as well as the multimodal simulator dyngen  [50] (Methods ). Because both RNA +velocity in the original gene space and CLT are known for these simulations, we could +directly compare estimated values with ground truth. Velocity estimation accuracy was +assessed using distance correlation [51, 52], which is well suited for measuring associa - +tions between vectors of unequal dimensionality and therefore accommodates methods +that predict low-dimensional velocity representations. Latent time estimation accuracy +was evaluated using Pearson’s correlation. Details of both metrics are described in Addi- +tional file 3: Note S6. +Analysis of overall performance on simulated datasets showed that the top five meth - +ods were exclusively non-DL approaches based on ODEs (scVelo (dyn/stc), Velocyto, +and κ-velo) or SDEs (SDEvelo) (Additional file 2: Fig. S16a; Additional file 1: Table S8). +Their strong performance is likely due to their close alignment between their modeling +assumptions and the generative processes underlying the simulations. By contrast, sev - +eral DL-based methods, including cellDancer, scT our (MSE), and VeloVAE (std/Full VB), +ranked near the bottom. Consistent with this observation, non-DL methods significantly +outperformed DL methods overall (Additional file  2: Fig.  S16b, left panel), with non- +DL methods enriched among the top 10 and DL methods over-represented among the +bottom 10 (Additional file  2: Fig.  S16b, middle panel). Importantly, no significant per - +formance difference between DL and non-DL methods was observed for the core tasks +based on real-world datasets (Additional file 2: Fig. S16b, right panel). This discrepancy +indicates a gap between simulated and biological settings, suggesting that current simu - +lation frameworks favor non-DL approaches. +T o explore this discrepancy, we compared method rankings derived from simulated +and real-world datasets. We observed that only four methods exhibited a substantial +rank divergence between the two settings: scVelo (dyn), GraphVelo (std), cellDancer, and +κ-velo (Additional file 2: Fig. S16c). After excluding these outliers, we identified a strong +positive correlation between rankings on simulated and real-world data (Spearman’s +ρ = 0.888 , P = 7.910 × 10−8 ; Additional file  2: Fig.  S16d). This concordance, observed +despite the differences between simulation and real-world datasets, indirectly validates +the robustness and reliability of our core task results. +Multimodal‑enhanced RNA velocity methods +Advanced RNA velocity methods extend beyond RNA-only approaches by incorporating +additional biological information (Fig. 7a). Several methods integrate single-cell assay for +transposase-accessible chromatin with sequencing (scATAC-seq) data, including Mul - +tiVelo [14], LatentVelo (ATAC) [33], and GraphVelo (ATAC) [42], while others, such as +Dynamo (m2) [15] and VelvetVAE [32] utilize single-cell metabolic labeling data (Meth - +ods). Moreover, velocity estimation has been extended to incorporate inputs such as +protein abundance [53], transcription factor activity [54], phylogenetic information [55], +gene regulatory networks  [56], and spatial transcriptomics  [57–59]. In this study, our +benchmarking focuses specifically on the contributions of chromatin accessibility and +metabolic labeling. +Page 17 of 36 +Wu et al. Genome Biology (2026) 27:242 + +We first benchmarked three chromatin accessibility–enhanced methods against 25 +RNA-only methods using paired scRNA-seq and scATAC-seq datasets (Data 22–24). +Performance was evaluated using CBDir and ICVCoh (Additional file 2: Fig. S17a; Addi - +tional file  1: Table  S9). Among the three enhanced methods, GraphVelo (ATAC) sub - +stantially outperformed MultiVelo and LatentVelo (ATAC) in CBDir, whereas MultiVelo +Fig. 7 Analysis of multimodal-enhanced RNA velocity methods. a Schematic overview of RNA velocity +methods incorporating multimodal information. Beyond using RNA-only, advanced frameworks leverage +data from chromatin accessibility (MultiVelo, LatentVelo (ATAC), GraphVelo (ATAC)) and metabolic labeling +(Dynamo (m2), VelvetVAE). Furthermore, velocity estimation can incorporate additional inputs such as protein +abundance, transcription factors, phylogenetic trees, gene regulatory networks, and spatial transcriptomics. +b Performance of chromatin accessibility–enhanced methods on Data 22–24, evaluated by CBDir (top) +and ICVCoh (bottom). Bar plots compare three multimodal methods against their corresponding RNA-only +counterparts. Ranks in parentheses denote the method’s position among all 28 evaluated methods (3 +multimodal and 25 RNA-only). Individual points represent scores for specific cross-validation folds in each +dataset. c Evaluation of metabolic labeling–enhanced methods. Top: CBDir scores on Data 9 and 26. Bottom: +TSC scores on Data 10 and 25. Comparisons include the corresponding RNA-only counterparts. Ranks +in parentheses denote the position among all 27 evaluated methods (2 multimodal and 25 RNA-only). +Individual points represent scores for specific cross-validation folds in each dataset +Page 18 of 36Wu et al. Genome Biology (2026) 27:242 +and LatentVelo (ATAC) performed relatively better in ICVCoh. Interestingly, none of +the enhanced methods achieved the top average rank for either metric, suggesting a +potential trade-off between the benefits of multimodal integration and increased model +complexity. To more directly assess the contribution of scATAC-seq data, we performed +pairwise comparisons between each enhanced method and its RNA-only baseline: Mul - +tiVelo versus scVelo (dyn), LatentVelo (ATAC) versus LatentVelo (std), and GraphVelo +(ATAC) versus GraphVelo (std) (Fig. 7b; Additional file 2: Fig. S17b). In these compari - +sons, GraphVelo (ATAC) was the only method that generally achieved performance +comparable to or better than its corresponding baselines across both metrics. +We next benchmarked two metabolic labeling–enhanced methods against 25 RNA- +only methods using four datasets containing both spliced/unspliced and newly synthe - +sized/total RNA counts. We evaluated performance using CBDir on Data 9 and 26 and +TSC on Data 10 and 25 (Additional file 2: Fig. S18a; Additional file 1: Table S10). Com - +paring the two enhanced methods, the results showed that Dynamo (m2) outperformed +VelvetVAE in both metrics. However, as observed for chromatin accessibility-based +methods, no metabolic labeling-enhanced method achieved the top rank for either +metric. We then prioritized comparing each enhanced method against its RNA-only +baselines, specifically Dynamo (m2) versus Dynamo (m1) and VelvetVAE versus Svel - +vetVAE (Fig. 7c; Additional file 2: Figs. S18b, c). These analyses showed that incorporat - +ing metabolic labeling data did not consistently improve performance over RNA-only +approaches. Both Dynamo (m2) and VelvetVAE failed to outperform their baselines in at +least one metric. For example, Dynamo (m2) ranked lower than Dynamo (m1) for TSC, +while VelvetVAE performed worse than SvelvetVAE for CBDir. +Computational scalability analysis +We assessed computational scalability by monitoring running time and peak memory +usage across synthetic datasets with varying cell counts (Data 33) and gene counts (Data +34). To evaluate scalability with respect to the number of cells, we fixed the gene count +( p = 1, 000 ) and generated seven datasets with cell counts ranging from n = 1, 000 to +1,  000,  000. Conversely, to assess scalability with respect to the number of genes, we +fixed the cell count ( n = 10, 000 ) and generated six datasets with gene counts ranging +from p = 1, 000 to 20, 000. T o quantify performance, we introduced an efficiency score +that integrates both running time and peak memory usage, and a scalability score that +considers both the maximum dataset size successfully processed at the cell or gene level +and efficiency (Methods ). Details of the computational environment are provided in +Additional file 3: Note S7. +We first evaluated scalability using the efficiency score with respect to gene counts and +found that seventeen methods successfully processed datasets with up to 20,000 genes +(Fig. 8a, top panel). Among these, the three scT our variants secured the top three rank - +ings, followed closely by veloVI, UniTVelo (uni), and SDEvelo. Conversely, cell2fate failed +even on the 5,000-gene datasets, while four others (Pyro-Velocity (m1/m2), GraphVelo +(std), cellDancer) could not handle the 10,000-gene datasets. Even among methods +that completed datasets at the same gene count, running time and peak memory usage +varied dramatically. For example, at the 20,000-gene level, running times ranged from +31.075 seconds (Velocyto) to 9.657 hours (VeloVAE (Full VB)), and peak memory usage +Page 19 of 36 +Wu et al. Genome Biology (2026) 27:242 + +Fig. 8 Computational scalability analysis for 25 RNA-only methods. a Top: Dot plot of scalability with +respect to gene number on simulated datasets, ranging from 1,000 to 20,000 genes with a fixed cell count +( n = 10, 000 ). In each row (corresponding to a gene count), methods are ranked by the efficiency score for +the corresponding data size in descending order (Methods). Methods are ordered along the x-axis based +on the scalability score Scalabilitygene in descending order (Methods). Bottom: Dot plot of scalability with +respect to cell number on simulated datasets, ranging from 1,000 to 1,000,000 cells with a fixed gene count +( p = 1, 000 ). In each row (corresponding to a cell count), methods are ranked by the efficiency score for +the corresponding data size in descending order (Methods). Methods are ordered along the x-axis based +on the scalability score Scalabilitycell in descending order (Methods). The top 10 methods for Scalabilitycell +are highlighted. b Line plot of accuracy versus downsampling rates for the top 10 methods in terms of +Scalabilitycell using two real-world datasets. The purple line indicates CBDir using Data 7, while the pink line +indicates TSC using Data 16. Sampling rates range from 0.2 to 1.0 (the original data). The dashed line at y = 0 +serves as a baseline +Page 20 of 36Wu et al. Genome Biology (2026) 27:242 +spanned from 0.848 GB (scTour (MSE)) to 61.790 GB (Dynamo (m1)) (Additional file 2: +Fig. S19a, top panel; Additional file 1: Table S11). +Next, we evaluated scalability with respect to cell counts (Fig.  8a, bottom panel). We +observed that all methods completed the task at 10,000 cells, and more than half handled +100,000 cells. However, only five methods (scTour (MSE/NB/ZINB), Velocyto, scVelo +(stc)) completed the task at 500,000 cells, with only the three variants of scTour complet- +ing it at 1,000,000 cells. Such scalability is largely attributable to scTour’s built-in sub - +sampling strategy, which by default uses 20% of cells for training when the total number +of cells exceeds 10,000 [30]. Notably, while scTour (MSE) demonstrated superior scal - +ability across both cell and gene dimensions, it ranked lowest in the core tasks. Regard - +ing resource efficiency, running time and peak memory usage also varied dramatically at +the same cell count. Among methods that successfully processed the 100,000-cell data - +set, running times ranged from 16.799 seconds (Velocyto) to 8.196 hours ( κ-velo), while +peak memory usage ranged from 0.181 GB (veloVI) to 60.232 GB (Dynamo (m1)) (Addi - +tional file 2: Fig. S19a, bottom panel; Additional file 1: Table S12). +Extending simulated analysis to real-world datasets, we selected the top ten meth - +ods according to the cell scalability score of (Methods ), including 7 DL methods and 3 +non-DL methods, and explored the relationship between accuracy and scalability. We +evaluated directional consistency using Data 7 ( n = 54, 542 cells) and assessed tempo - +ral precision using Data 16 ( n = 165, 522 cells), each with downsampling rates ranging +from 0.2 to 1.0 (original data). The results show that most methods, such as scT our (NB), +veloVI, and SDEvelo, displayed near-stable performance across varying sampling rates, +and were minimally affected by the difference between the smallest and full data sizes +(Fig. 8b; Additional file 2: Fig. S19b; Additional file 1: Table S13). This finding suggests +that downsampling serves as a viable strategy to enable RNA velocity estimation on +large-scale datasets by largely saving computational resources, provided that the chosen +method demonstrates robustness to varying sampling rates. +Guideline for method selection +RNA velocity methods are generally user-friendly, as they all provide Python packages +that operate directly on AnnData, a widely adopted data structure for single-cell analy - +sis. In this section, we offer practical recommendations based on directional consistency, +temporal precision, multimodal integration, computational scalability, and hyperparam - +eter tuning. +When exploring cell fate transitions, users should prioritize directional consistency +while carefully guarding against false positives. Although LatentVelo (std) ranked first +in Task I (Fig. 2, Col. 9), its tendency to infer spurious results in Task III (Fig.  2, Col. 15) +suggests that it is best suited for exploratory analyses and should be validated using com- +plementary approaches. For a more balanced trade-off between directional consistency +and negative control robustness, we recommend UniTVelo (uni), Pyro-Velocity (m2), +and veloVI. +For temporal analyses, UniTVelo (uni) and SDEvelo are recommended due to their +high temporal precision (Fig.  2, Col. 12). We also recommend four Group 3 methods, +DeepVelo, Velocyto, Dynamo (m1), and SvelvetVAE, although CLT must be derived +Page 21 of 36 +Wu et al. Genome Biology (2026) 27:242 + +through post-hoc processing (Additional file 3: Note S4). As a best practice, directional +and temporal results should be interpreted jointly, since inferred transition directions +and cell-state times are expected to be coherent. Users may compare direction-inference +and temporal-inference results across recommended methods, particularly prioritizing +UniTVelo (uni), which performs well in both analyses and estimates transition direction +and cell-state time. In cases where spliced/unspliced matrices are unavailable, scTour +(ZINB) can serve as a practical alternative, but its results may benefit from post-infer - +ence adjustments as suggested in the original study [30]. +When multimodal data are available, we recommend GraphVelo (ATAC) for analyses +incorporating chromatin accessibility (Fig.  7b), and Dynamo (m2) for metabolic labe - +ling data (Fig. 7c). However, given current architectural limitations of multimodal frame- +works, we advise users to prioritize RNA-only methods as the primary analysis strategy, +even when additional modalities are available. +From a practical perspective, computational scalability is often a key factor in method +adoption. Among the recommended methods, some showed limited scalability with +respect to cell number. Specifically, UniTVelo (uni) and SvelvetVAE could process only +up to 10,000 cells, whereas Pyro-Velocity (m2) could process only up to 50,000 cells +(Fig. 2, Col. 5). For these methods, we suggest users downsample the data to perform +estimation on large-scale datasets. For a balance between core task performance and +scalability, we recommend SDEvelo, Velocyto, scVelo (stc), DeepVelo, and veloVI as opti- +mal choices. Another practical consideration is the choice of method-specific hyperpa - +rameters. We performed a sensitivity analysis for the recommended methods by varying +key parameters discussed in the original papers or tutorials (Additional file 3: Note S8). +Although most methods showed relatively robust performance across different param - +eter settings, the default settings were not always optimal (Additional file  2: Fig.  S20). +We therefore suggest that users begin with the default settings and subsequently tune +method-specific hyperparameters according to their analysis goals. +Discussion +In this study, we established a comprehensive benchmarking framework for RNA veloc - +ity analysis using 34 datasets. We systematically evaluated 25 RNA-only methods across +eight evaluation tasks and assessed five multimodal-enhanced methods specifically for +multimodal integration. Compared with prior benchmarks [24–26], our study provides +a more comprehensive assessment of RNA velocity accuracy and robustness by match - +ing datasets to task-specific biological and technical questions, with a focus on the direct +outputs of RNA velocity methods (velocity vectors and latent time). We used cell differ - +entiation datasets to assess directional consistency, datasets with ground-truth temporal +labels to evaluate temporal precision, and mature cell population datasets as negative +controls for false-positive dynamic signals. By linking performance across different tasks, +our benchmarking framework revealed cross-task trade-offs and methodological gaps +that are difficult to capture using a single aggregate ranking. We were also able to directly +examine robustness under negative controls [21, 22], the role of gene dependence [21], +and stability across spliced/unspliced mRNA quantification pipelines  [20, 22, 29], key +issues raised in previous studies but rarely incorporated into existing benchmarks. Fur - +ther discussion of shared and disparate findings, as well as the complementary scopes +Page 22 of 36Wu et al. Genome Biology (2026) 27:242 +of existing benchmarks, is provided in Additional file 3: Note S9. Together, these design +choices enabled a systematic comparison and ranking of existing methods, offering clear, +evidence-based guidance for the community. +An ideal RNA velocity method should both recover biologically meaningful cell fate +transitions without introducing false positives and infer a CLT that accurately reflects +the underlying biological clock. However, our results show that no single method con - +sistently performs well across all core tasks. This limitation likely reflects several under - +lying factors. First, we observed a significant negative correlation between directional +consistency and negative control robustness. One contributing factor to this trade-off +is the use of constraints that enforce temporal progression in the velocity field; although +such constraints can improve directional inference, they limit the ability to correctly +characterize the absence of velocity in terminally differentiated states  [30]. Conse - +quently, existing methods often lack mechanisms to represent steady-state splicing +dynamics without over-inference, highlighting the need for more advanced modeling +strategies or the incorporation of appropriate prior information to suppress spurious +signals. Beyond controlling false positives in velocity vectors, existing RNA velocity +methods also struggle to accurately infer CLT simultaneously with velocity. Although +one simultaneous modeling method ranked within the top five for both directional and +temporal tasks, it exhibited suboptimal performance on the negative control task. These +findings suggest that future development should focus either on refining simultaneous +modeling architectures to mitigate spurious signals, or on improving CLT inference in +post-hoc strategies that already demonstrate robust directional consistency. +Complementing the core tasks, our analysis of additional tasks uncovers insights that +can guide the future development of RNA velocity methods. Quantification is a manda - +tory upstream step, yet it introduces technical noise that is independent of biological +context  [29]. Notably, our results show that only three methods remain robust across +different quantification outputs. This stability may be attributed to spliced RNA–ori - +ented designs  [27], which compute velocity using predicted counts, or to the use of +low-dimensional embeddings to mitigate technical noise [31]. In addition, model-based +simulation experiments show that non-DL methods, particularly ODE- and SDE-based +approaches, inherently outperform DL methods. Consequently, such simulations should +be used judiciously when evaluating newly developed methods to avoid potential bias. +Another challenge lies in integrating biological modalities beyond unspliced and spliced +RNA. In the multimodal integration task, neither chromatin accessibility–enhanced nor +metabolic labeling–enhanced methods achieved top rankings. This outcome suggests +that future efforts should prioritize advanced integration strategies and more refined +model architectures. Finally, this work assessed scalability on the large-scale datasets, +showing that existing top methods in core tasks have limited scalability with respect +to the number of cells. Although our experiments show that this can be handled by a +downsampling strategy when analyzing large datasets, future method development +should focus on more flexible and scalable frameworks. +Page 23 of 36 +Wu et al. Genome Biology (2026) 27:242 + +This benchmark offers key insights into RNA velocity methods, yet several impor - +tant directions warrant further investigation. With respect to gene selection, we fol - +lowed standard practice by using 2,000 HVGs; however, HVGs are not always the most +informative features for velocity inference [15]. Systematic evaluation of alternative gene +selection on velocity estimation represents a promising avenue for future work. Regard - +ing method implementation, while we applied default parameters to ensure a stand - +ardized and fair comparison, method-specific hyperparameter optimization remains a +critical factor that could further enhance performance in practical applications. Finally, +most existing RNA velocity methods rely on short-read sequencing, which may lack the +resolution needed to fully capture complex splicing dynamics. Emerging technologies +such as long-read sequencing [60, 61] offer the potential for more detailed RNA velocity +estimation. +Conclusions +In summary, we established a comprehensive benchmarking framework for RNA veloc - +ity, encompassing a systematic evaluation of 25 RNA-only methods across eight tasks, +together with a targeted assessment of five multimodal-enhanced methods. Our results +demonstrate that no single method performs consistently well across all tasks. Accord - +ingly, we provide practical guidelines to assist users in choosing tools that best align with +their specific biological questions and technical constraints. This study also identifies +several methodological gaps, including the need for improved modeling of gene depend- +ence, more accurate temporal inference strategies, and better-designed multimodal +architectures. We anticipate that this framework, together with our curated datasets and +evaluation metrics, will serve both as a practical guide for users and as a reference stand- +ard for method developers. Ultimately, these resources aim to facilitate more robust +modeling of transcriptional dynamics and improve the reliability of RNA velocity analy - +ses in single-cell studies. +Methods +Benchmarking datasets +This study comprises 26 real-world datasets and eight simulated datasets. The real-world +datasets span diverse biological contexts across mouse and human species and vary +widely in scale, with cell counts ranging from 753 to 165,522 and gene features from 936 +to 53,801. Complementing these are six full model-generated simulated datasets, includ- +ing two based on ODEs, two based on SDEs, and two generated using dyngen  [50]. +Finally, two additional ODE-based simulations were generated to systematically evalu - +ate method scalability across varying gene and cell counts, including large-scale settings +with up to 1,000,000 cells and 20,000 genes. Descriptions of each dataset are provided in +Additional file 3: Note S10. +Page 24 of 36Wu et al. Genome Biology (2026) 27:242 +Benchmarking methods +In this study, we present a comprehensive benchmark of 19 RNA velocity tools covering +30 distinct methods, including 25 RNA-only and five multimodal-enhanced approaches. +All methods were executed using their default parameter settings to ensure a fair base - +line comparison. For cases requiring dataset-specific configurations, we applied special - +ized model settings as detailed in Additional file 1: Table S14. +Velocyto +Velocyto [10] is a steady-state model based on an ODE assumption that estimates RNA +velocity via extreme quantile linear regression. Velocyto was applied using the Python +package scvelo (v0.3.3) through the function scvelo.tl.velocity() with the +parameter mode=‘deterministic’. +scVelo +scVelo  [11] provides three modes for RNA velocity estimation: deterministic (cor - +responding to Velocyto), dynamical, and stochastic. scVelo (dyn) solves a dynami - +cal ODE model using an expectation–maximization algorithm to infer kinetic rates. +scVelo (stc) generalizes the steady-state model by accounting for transcriptional sto - +chasticity via second-order moments, solving the steady-state ratio using general - +ized least squares. These modes were implemented in the Python package scvelo +(v0.3.3) using scvelo.tl.velocity() with parameters mode=‘dynamical’ +and mode=‘stochastic’, respectively. +VeloAE +VeloAE [35] estimates RNA velocity by using an autoencoder to project spliced and +unspliced expression into a shared low-dimensional latent space. The method was +applied using the Python module veloproj (v0.2.0). Analyses were run with the +function main_AE() using default parameters. VeloAE outputs velocity vectors only +in a latent space, which were used for velocity evaluation. +Dynamo +Dynamo  [15] is a computational framework for constructing transcriptomic vec - +tor fields by automatically detecting the modality of input data and selecting an +appropriate model. We benchmarked two modes implemented in the Python pack - +age dynamo-release (v1.4.2rc1) using the function dyn.tl.dynamics(). +Dynamo (m1) estimates splicing dynamics–based velocity using a negative bino - +mial (NB) distribution and the generalized method of moments, with parameters +model=‘stochastic’ and est_method= ‘negbin’. Dynamo (m2) models +labeled and total RNA using an ODE system with default settings. +Page 25 of 36 +Wu et al. Genome Biology (2026) 27:242 + +Pyro‑Velocity +Pyro-Velocity [28] is a probabilistic generative model for RNA dynamics that adapts +the Velocyto ODE framework by introducing a CLT shared across genes  [28]. We +benchmarked two methods implemented in the Python module pyrovelocity +(v0.3.0b2) using the function train_model(). Pyro-Velocity (m1) constrains the +CLT to be greater than the initial switching time by setting guide_type= ‘auto_ +t0_constraint’, whereas Pyro-Velocity (m2) removes this constraint to allow +gene-specific time lags using the parameter guide_type= ‘auto’. +UniTVelo +UniTVelo [27] is a computational framework that models gene-specific kinetics start - +ing from spliced RNA in a top-down manner. UniTVelo (uni) aggregates gene-spe - +cific temporal information into a single shared time per cell, whereas UniTVelo (ind) +retains a gene-specific time matrix during optimization. These modes were imple - +mented in the Python package unitvelo (v0.2.5) using unitvelo.run_model() +with parameters velo.FIT_OPTION= ‘1’ and velo.FIT_OPTION= ‘2’, +respectively. +VeloVAE +VeloVAE [34] is a Bayesian model for RNA velocity inference. We applied two methods +implemented in the Python package velovae (v0.1.3b0) using the function velovae. +VAE(). VeloVAE (std) estimates posterior distributions of latent time, cell states, and +kinetic rates using default settings. VeloVAE (Full VB) extends this framework by mod - +eling kinetic rate parameters as random variables, enabled by setting full_vb=True. +κ‑velo +κ-velo [38] is an ODE-based RNA velocity method that explicitly incorporates a time- +scale parameter. It was implemented in the Python package velocity-package +(v0.2) using the function velocity.tl.fit.get_velocity(). +MultiVelo +MultiVelo [14] extends scVelo (dyn) by incorporating paired scATAC-seq data. It intro - +duces an additional ODE to link transcription rates with promoter and enhancer accessi- +bility. MultiVelo was implemented in the Python module multivelo (v0.1.5) using the +function mv.recover_dynamics_chrom(). +cellDancer +cellDancer [37] is a deep neural network (DNN) model for RNA velocity estimation that +learns cell-specific kinetic parameters for each gene. It was implemented in the Python +package celldancer (v1.1.7) using the function cd.velocity(). +Page 26 of 36Wu et al. Genome Biology (2026) 27:242 +veloVI +veloVI [31] is a deep generative framework that reformulates scVelo’s dynamical model +using a VAE. We applied it using the Python package velovi (v0.3.1) via the function +VELOVI(). +LatentVelo +LatentVelo  [33] is a generative approach for inferring RNA velocity in a latent space. +Two methods were applied using the Python package latentvelo (v0.1). LatentVelo +(std) embeds cells using a VAE and models differentiation dynamics via a neural ODE, +implemented with ltv.models.VAE(). LatentVelo (ATAC) integrates scATAC- +seq data into the framework and was implemented using ltv.models.ATACReg - +Model(). For both methods, velocity evaluation was performed using the latent-space +velocity vectors. +scTour +scTour  [30] is a DNN-based method for estimating RNA velocity and CLT from gene +expression data rather than explicit spliced/unspliced counts. Three modes were +implemented in the Python package scTour (v1.0.0) via the function sct.train. +Trainer(). Specifically, scTour (MSE) uses mean squared error (MSE) with the +parameter loss_mode=‘mse’; scTour (NB) uses an NB-conditioned likelihood by +setting loss_mode=‘nb’; and scTour (ZINB) uses a zero-inflated negative binomial +(ZINB)-conditioned likelihood by setting loss_mode=‘zinb’. All three methods +output only latent-space velocity vectors, which were used for velocity evaluation. +DeepVelo +DeepVelo [36] is a DNN model that infers RNA velocity by learning kinetic parameters +for each cell and gene. It was implemented in the Python package deepvelo (v0.2.8) +using the function dv.train(). +SDEvelo +SDEvelo [39] is an adversarial learning approach that infers RNA velocity using multi - +variate SDEs. The method was applied using the Python module sdevelo (v0.2.12) via +the function sv.SDENN(). +VelvetVAE +VelvetVAE [32] introduces a VAE framework for RNA velocity inference based on meta- +bolic labeling data. SvelvetVAE  [32] is a variant designed for velocity estimation from +unspliced and spliced mRNA data. Both were applied using the Python package vel - +vetvae (v0.1.dev0) via vt.md.Velvet() and vt.md.Svelvet(), respectively. +cell2fate +cell2fate  [40] is a fully Bayesian RNA velocity method based on a linearized velocity +ODE. It was applied using the Python module cell2fate (v0.1a0) via the function +mod.compute_and_plot_total_velocity(). +Page 27 of 36 +Wu et al. Genome Biology (2026) 27:242 + +TIVelo +TIVelo  [41] is a model-free RNA velocity method that infers trajectory directions on +a cluster-level graph to supervise cell-level velocity. Two modes implemented in the +Python package tivelo (v0.1.4) were applied using tivelo(). TIVelo (std) includes +a cosine similarity regularization term in the objective function with the parameter +constrain=True, while TIVelo (simple) solves the objective function without this +regularization by setting constrain=False. +GraphVelo +GraphVelo [42] is a graph-based method that corrects RNA velocities by optimizing a +tangent space projection loss on the cell-state manifold. Because it requires an initial +velocity estimate, we benchmarked two variants implemented in the Python pack - +age graphvelo (v0.1.9) using GraphVelo(). GraphVelo (std) was initialized with +scVelo (dyn) output, whereas GraphVelo (ATAC) was initialized using MultiVelo output, +thereby incorporating scATAC-seq information. +Benchmarking metrics +To evaluate and rank the RNA velocity methods, we employed a suite of benchmarking +metrics implemented in our developed Python package VeloEV (v1.0) [62]. +CBDir +Cross-boundary directional correctness (CBDir) [ 35] quantifies the directional accu - +racy of inferred RNA velocities between predefined pairs of cell clusters. For consist - +ency with the original study, CBDir was computed using velocity vectors projected +into the UMAP space. Let T denote the set of predefined ground-truth transitions. +The metric is defined as: +Here, for each transition (A, B), CA and CB denote the source and target cell sets, respec - +tively; BA,B ={ c ∈ CA : CB ∩ N (c) �= ∅} denotes the subset of source cells with at least +one target-cell neighbor; N (c) denotes the expression neighborhood of cell c; x′ +c is the +UMAP coordinate of cell c; and v′ +c is the velocity vector projected into the same UMAP +space. CBDir ranges from −1 to 1, with higher values indicating stronger directional +agreement across cluster boundaries. +ICVCoh +In-cluster coherence (ICVCoh)  [ 35] measures the coherence of inferred velocity +directions by averaging cosine similarities between each cell’s velocity vector and +those of its expression neighbors. Like CBDir, ICVCoh was computed using UMAP- +projected velocities. Let G denote the set of cell types or clusters. ICVCoh is defined +as: +(1)CBDir = 1 +|T | (A,B)∈T +1 +|BA,B| c∈BA,B +1 +|CB ∩ N (c)| c′∈CB ∩N (c) +v′ +c · (x′ +c′ − x′ +c) +�v′c�� x′ +c′ − x′c� . +Page 28 of 36Wu et al. Genome Biology (2026) 27:242 +where CA is the set of cells in cluster A ; BA ={ c ∈ CA : CA ∩ N (c) �= ∅} denotes cells +with at least one same-cluster neighbor; N (c) is the expression neighborhood of cell c ; +and v′ +c is the velocity vector of cell c projected into the UMAP space. ICVCoh ranges +from −1 to 1, with higher values indicating greater local velocity coherence. +CTO +Cluster temporal ordering accuracy (CTO) is a metric designed to assess how well +inferred CLT preserves the correct temporal ordering between ground-truth develop - +mental time clusters (e.g., reprogramming days). CTO is defined as: +where {T1 ,T2 ,... ,TK } denotes the set of ground-truth time clusters in ascending order +and ti denotes the inferred CLT. CTO ranges from 0 to 1, with higher values indicating +more accurate temporal ordering. +TSC +Temporal Spearman correlation (TSC) measures the accuracy of inferred CLT by +comparing the rank-transformed inferred times with the true developmental times. +In contrast to CTO, which evaluates the ordering of discrete time clusters, TSC pro - +vides a continuous assessment of temporal concordance along the trajectory. TSC is +defined as follows: +where R t and R t∗ denote the ranks of inferred CLT and labeled developmental time, +respectively, and σ denotes the standard deviation. TSC ranges from −1 to 1, with higher +values indicating stronger agreement between the inferred CLT and labeled develop - +mental time. +STS +Self-transition score (STS) [11] serves as a partial indicator of negative control robust - +ness by quantifying the probability of a cell transitioning to itself. STS is computed based +on a velocity graph and defined as: +(2)ICVCoh = 1 +|G| +∑ +A∈G +1 +|BA | +∑ +c∈BA +1 +|CA ∩ N (c)| +∑ +c′∈CA ∩N (c) +v′ +c · v′ +c′ +�v′c�� v′ +c′ � , +(3)CTO = 1 +K +∑ +k +1 +|T k | +∑ +i∈T k +1 +|T k +1| +∑ +j∈T k +1 +1{ti 근거 자료: publisher metadata의 abstract 원문(`sources/abstract.txt`) + 본문 도입부(`sources/fulltext_extracted.txt`). abstract에 명시된 것만 근거로 한다. + +## Abstract Summary + +- **한 문장 요약**: RNA velocity method가 급증했으나 종합·표준 벤치마크가 없는 상황에서, 19개 도구(30 method)를 34개 데이터셋(26 real + 8 simulated)에서 8개 task로 평가하고, 단일 랭킹 대신 **task별·맥락별 선택 가이드**를 제시한 종합 벤치마크. +- **연구 목적**: diverse biological·technical scenario에서 맥락 특이적 metric으로 method를 평가하는 표준 벤치마크 수립. +- **문제/gap**: 기존 비교가 "limited scope or incomplete task design"이라 사용자에게 명확한 가이드가 없음. +- **핵심 방법**: 25개 RNA-only method를 8 task로 평가(핵심 4: directional consistency, temporal precision, negative control robustness, sequencing depth stability), multimodal-enhanced 5종은 multimodal integration task에서만 평가. +- **주요 결과**: directional consistency ↔ negative control robustness 사이 **명확한 trade-off**, temporal modeling 전략별 그룹 거동, sequencing depth·quantification 선택에 따른 변동. 개선 필요 gap 3종(gene dependence 모델링, temporal inference 정확도, multimodal 설계). +- **저자 주장 기여**: 단일 overall 랭킹이 아니라 생물·기술 맥락별 **task-aware 선택 지침**. + +## 우리 원고(BIOP01)와의 관계 — abstract 수준 + +- 같은 장르(velocity 벤치마크)이나 **RNA-only 중심 + 임베딩/전이벡터 지표**다. 우리는 multiome 행렬 재현성 + ATAC-shuffle 인과 + chromatin-lag 감사라 축이 다르다. +- abstract에는 우리 차별점 키워드(cell×gene 행렬 method 간 재현성, ATAC-shuffle, chromatin-lag 신뢰성, 외부 측정 rate 앵커, 사전등록)가 **없다**. 전문 검증은 `_core.md` §스쿱. +- **타깃 저널(Genome Biology) 게재**이므로 abstract만으로도 인용·차별화 필요가 확정된다. diff --git a/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_core.md b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_core.md new file mode 100644 index 0000000..f8c0018 --- /dev/null +++ b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_core.md @@ -0,0 +1,71 @@ +# Wu et al., 2026 — Comprehensive benchmarking of RNA velocity methods across single-cell datasets — core 분석 + +> 근거 자료: `sources/wu-2026-velocity-benchmark.pdf`(본문 36p, Genome Biology 27:242) + `sources/fulltext_extracted.txt`(pypdf 추출, grep 근거). 본문에 텍스트로 명시된 method 이름·task 정의·metric·수치·dataset accession만 단정한다. Figure에서 읽어야 하는 값은 `검토필요:`로 표시한다. +> +> 표기: `해석:` / `외부 맥락:` / `추정:` / `미제공:` / `검토필요:`. + +## Executive Summary + +- **무엇**: RNA velocity method가 급증(La Manno 2018 이후 RNA-only 17종 + multimodal-enhanced 9종)했으나 종합·표준 벤치마크가 없다는 문제의식에서, **19 도구 / 30 method를 34 데이터셋(26 real + 8 simulated)에서 8 task로 평가**하고 단일 랭킹 대신 **task-aware 선택 지침**을 제시한 resource/benchmark 논문. 새 method가 아니다. +- **평가 설계**: RNA-only 25종은 8 task 전부에서, multimodal-enhanced 5종은 **multimodal integration task 하나에서만** 평가. 핵심 4 task와 지표: + - (I) **directional consistency** — CBDir(cross-boundary direction correctness) · ICVCoh(in-cluster coherence) + - (II) **temporal precision** — CTO(cluster temporal ordering) · TSC(temporal Spearman correlation) + - (III) **negative control robustness** — STS(self-transition score) · EES(effective entropy score) + - (IV) **sequencing depth stability** — SCBDir · STSC + - (추가 4 task: quantification stability 등) +- **핵심 결과**: + - ① **directional consistency ↔ negative control robustness 사이 유일하게 유의한 음의 상관**(Spearman ρ = −0.572, P = 0.003). 방향 신호를 강하게 최적화한 method가 false positive 억제를 희생. **LatentVelo(std)가 directional 1위인데 negative control 23위**로 그 trade-off를 전형적으로 보여줌 — "dynamic inference 강조가 spurious velocity를 낳을 수 있다". + - ② method 간 방향 불일치가 top method에서도 큼. UniTVelo(ind)·veloVI·VeloAE가 Data 1에서 HSC↔erythroid 방향을 역전 추정(Fig 3b). 같은 method가 Data 3에서는 옳게 추정 → **dataset 의존**. + - ③ **multimodal integration task**: paired scRNA-seq+scATAC 3개(Data 22–24)에서 chromatin-enhanced 3종(MultiVelo, LatentVelo(ATAC), GraphVelo(ATAC))을 **CBDir·ICVCoh로 순위**. **GraphVelo(ATAC)가 CBDir에서 MultiVelo·LatentVelo를 크게 앞섬.** + - ④ 개선 필요 gap 3종: gene dependence 모델링, temporal inference 정확도, multimodal architecture 설계. +- **우리 적용 / 스쿱**: 아래 §스쿱. 요지 = **스쿱 아님**(축이 다름), **반드시 인용**(타깃 저널 게재 + 우리 데이터 Data 6 사용), **차별화 문단 사활**, **GB desk-reject 위험 상향**. + +## Identity + +- **Title**: Comprehensive benchmarking of RNA velocity methods across single-cell datasets +- **Authors**: Yida Wu, Chuihan Kong, Xu Liao, Zhixiang Lin, Xiaobo Sun, Jin Liu (corresponding — `검토필요:` 정확한 소속·교신 이메일은 전문 헤더에서 확정) +- **Venue**: Genome Biology **27:242** (2026-07-27 게재), open access +- **DOI**: 10.1186/s13059-026-04182-z (preprint Research Square 10.21203/rs.3.rs-8708834/v1, 2026-02) +- **Citation key**: `wu2026velocitybenchmark` + +## Background + +RNA velocity는 spliced/unspliced mRNA 비율로 cell state transition 방향을 추정한다(La Manno 2018). 본문은 파이프라인을 preprocessing → velocity estimation → postprocessing 3단계로 정리하고, postprocessing에서 velocity matrix로부터 cosine-kernel 전이행렬을 만들어 저차원 임베딩에 투영한다고 명시한다. method 계보를 DL 11종(veloVI·scTour·SvelvetVAE·LatentVelo·VeloVAE 등 VAE 계열 + VeloAE·DeepVelo·cellDancer 비생성형) vs non-DL 14종으로 나눈다. multimodal-enhanced 9종은 chromatin accessibility(MultiVelo 등) 또는 metabolic labeling(Dynamo·VelvetVAE) 등 auxiliary 정보를 입력으로 쓴다. + +- **해석**: 이 논문의 채점 지점은 [12] Luo·[13] Huang과 동일 계열이다 — velocity를 **임베딩/전이벡터 수준**(CBDir·ICVCoh)에서 평가하고 method를 순위 매긴다. multimodal은 8 task 중 1개(통합 task)에서만, 그것도 같은 CBDir·ICVCoh로 본다. + +## 핵심 방법·결과 (근거 표시) + +- **directional consistency**: CBDir(ground-truth transition으로 방향 정확도) + ICVCoh(cluster 내 cosine 일관성). — 본문 명시. +- **negative control robustness**: STS(self-transition score) + EES(effective entropy score). — 본문 명시. **해석: 이는 "정적/무-동역학 상황에서 velocity가 spurious하지 않은가"(출력 robustness)를 재는 것이다. 크로마틴을 파괴하는 우리 ATAC-shuffle 인과 대조와 목적이 다르다.** +- **trade-off 수치**: Spearman ρ = −0.572, P = 0.003 (directional vs negative control). LatentVelo(std) 1위 ↔ 23위. — 본문 텍스트 값. +- **multimodal**: Data 22–24(paired RNA+ATAC)에서 GraphVelo(ATAC) > MultiVelo·LatentVelo(ATAC) in CBDir. — 본문 명시. `검토필요:` 정확한 점수·ICVCoh 순위는 Fig 7 / Additional file 1 Table S9. +- **데이터셋**: 34개(26 real + 8 sim). **Data 6 = GSE209878 + GSE284047** — 본문 accession 명시. **외부 맥락: GSE209878은 우리 BIOP01 HSPC 1차 데이터. Luo 2026이 Dataset12로 쓴 것과 동일.** + +## ★ 스쿱 점검 — 우리 원고(BIOP01) 대비 (전문 근거) + +**판정: 스쿱 아님. [12][13]과 같은 계열의 세 번째 종합 벤치마크이며, 우리 6개 차별점이 전문 grep으로도 전부 유효.** + +| 우리 차별점 | Wu 2026 본문 (grep 실측) | +|---|---| +| cell×gene velocity **행렬의 method 간 재현성**(층②) | 없음 — 임베딩 CBDir/ICVCoh만 | +| **ATAC-shuffle 인과 대조** | 없음 — negative control은 STS/EES(출력 robustness). `shuffle` 0회·`permutation` 0회 | +| **chromatin→lag 신뢰성·부호 편향** | 주제 아님 — `lag` 1회(무관 맥락) | +| **외부 측정 rate 앵커**(TT-seq·half-life) | 없음 — metabolic labeling은 method **입력**으로만(Dynamo·VelvetVAE) | +| **사전등록·permutation null·자기철회** | 없음 (`preregist` 0·`permutation` 0) | +| **MoFlow·CRAK-Velo 포함** | `MoFlow` 0회·`CRAK` 0회 (우리 4 arm 중 둘 미포함) | +| HSPC 10x multiome 특이 | 일반 scRNA-seq 34종. multiome은 3개 데이터·1 task | + +- **negative control robustness 해소**: STS·EES 정의 확인 → "정적 상황 velocity가 spurious한가"이지 "크로마틴을 파괴하면 행렬이 움직이나"(우리 인과)가 아니다. 겹치지 않는다. +- **multimodal task 해소**: CBDir·ICVCoh로 MultiVelo 등을 **순위**한 것이지, 우리 **행렬의 method 간 재현성**도 chromatin-lag 신뢰성도 아니다. + +## 우리 적용 (BIOP01) + +1. **인용 필수** — 타깃 저널(GB) 게재 + 우리 데이터(GSE209878=Data 6) 사용. 미인용 시 리뷰어 첫 지적. refs에 [12][13] 옆 세 번째 벤치마크로. +2. **'무엇이 다른가' 문단 사활** — 같은 GB에 방금 종합 벤치마크. "우리는 임베딩이 아니라 **행렬**을, 순위가 아니라 **재현성·인과**를, RNA-only가 아니라 **multiome 4-arm**을, 그리고 **외부 측정 rate·사전등록**을 본다"를 앞세운다. +3. **저널 결정 영향(BIOP01-75)** — GB desk-reject 위험 상향("velocity 벤치마크를 방금 실었는데 또?"). CRM base-case 상대 안전성 ↑. 단 축이 갈려 '중복 아님' 방어 가능. +4. **부수 활용** — Wu의 GraphVelo(ATAC) > MultiVelo는 "새 multiome method가 계속 나온다 → 출력 신뢰성 감사가 필요"라는 우리 논지를 뒷받침. Discussion 한 줄. + +## 심층 + +한계·재현 ROI·산업 시선은 `..._lens-academic.md` / `..._lens-industry.md` / `..._methodology-brief.md` 참고. diff --git a/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_lens-academic.md b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_lens-academic.md new file mode 100644 index 0000000..03e400e --- /dev/null +++ b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_lens-academic.md @@ -0,0 +1,20 @@ +# Wu et al., 2026 (GB velocity benchmark) — lens: academic + +> 근거: `_core.md` + `sources/`. 학술 기여·한계·우리 원고 방어선 관점. + +## 학술적 기여 +- 규모: 19 도구/30 method × 34 데이터셋 × 8 task. RNA velocity 벤치마크 중 현재 최대 축(이전 Luo 14×17, Huang 29×176과 함께 3대 벤치마크). +- 개념 기여: **directional consistency ↔ negative control robustness trade-off**(ρ=−0.572, P=0.003)를 정량화. "방향 신호 최적화가 spurious velocity를 낳는다"를 순위 역전(LatentVelo 1위↔23위)으로 실증. +- 단일 랭킹을 거부하고 task-aware 선택 지침을 산출물로 삼음 — 사용성 측면 기여. + +## 한계 (우리 원고의 대비축) +- **채점이 임베딩/전이벡터 수준에 머문다**(CBDir·ICVCoh). velocity의 세 층(① 유전자별 모수 ② cell×gene 행렬 ③ 임베딩 화살표) 중 ③만 본다. ②(행렬)의 method 간 재현성은 다루지 않는다 — 우리 원고의 핵심 공백. +- **인과 대조가 없다**. negative control은 출력 robustness(STS/EES)이지 입력 교란(ATAC-shuffle)이 아니다. "무엇이 velocity를 만드는가"는 검정하지 않는다. +- **multimodal이 얕다**: 9개 multimodal 도구 중 5종, 그것도 1 task(통합), 3개 데이터셋. chromatin이 velocity에 기여하는지의 인과·재현성은 범위 밖. +- **외부 ground truth 부재**: 측정된 합성/분해 rate(TT-seq·SLAM) 앵커가 없다. 방향 정확도는 pseudo-ground-truth transition에 의존. +- **사전등록·자기철회 없음**: 결과를 본 뒤 지표를 고르는 것을 막는 장치가 명시되지 않음. + +## 우리 원고 방어선 (리뷰어 대응) +- "이미 벤치마크가 있는데?" → 세 벤치마크(Luo·Huang·Wu) 전부 **임베딩 순위**다. 우리는 행렬 재현성 + ATAC-shuffle 인과 + 외부 rate 앵커 + 사전등록. 축이 다르다. +- "Wu가 우리 데이터(GSE209878)를 썼는데?" → 그들은 Data 6을 임베딩 CBDir로 채점, 우리는 같은 데이터에서 cell×gene 행렬의 method 간 재현성과 chromatin-lag 신뢰성을 본다. 같은 데이터, 다른 질문. +- "GraphVelo(ATAC)가 MultiVelo보다 낫다던데?" → 새 multiome method가 계속 나온다는 사실 자체가 "출력이 신뢰 가능한지 먼저 감사하라"는 우리 논지를 강화한다. diff --git a/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_lens-industry.md b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_lens-industry.md new file mode 100644 index 0000000..3fde9e0 --- /dev/null +++ b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_lens-industry.md @@ -0,0 +1,12 @@ +# Wu et al., 2026 (GB velocity benchmark) — lens: industry / 실무 + +> 근거: `_core.md`. method 선택·파이프라인 적용 관점. + +## 실무 시사 +- **task-aware 선택**: 단일 최고 method가 없으므로, 목적(방향 정확도 vs false-positive 억제)에 따라 다른 method를 택해야 한다. 방향이 중요하면 directional 상위(예: LatentVelo std), 오탐 억제가 중요하면 negative-control 상위 — 둘은 trade-off라 동시 최적이 없다. +- **우리 파이프라인 적용**: 우리는 이미 MultiVelo·MoFlow·CRAK-Velo·MultiVeloVAE·scVelo 5-arm을 돌린다. Wu는 이 중 MultiVelo만(그것도 GraphVelo(ATAC)에 뒤짐) 평가 → **GraphVelo(ATAC)를 추가 arm 후보로 검토 가능**(향후 티켓, 필수 아님). +- **데이터 재사용**: Wu의 Data 6 = 우리 GSE209878. 전처리·transition 정의를 교차 참조 가능(단 Wu는 RNA-only/임베딩 관점). + +## 비용·재현 +- 코드·데이터 공개(open access, `검토필요:` repo URL). 재현 ROI는 낮음(우리 원고에 직접 재현할 이유는 없고, 인용·차별화가 목적). +- 우리 원고 관점의 실무 결론: 이 논문은 **경쟁 벤치마크**로 다루고(competitive-landscape), 채택할 도구나 지표는 없다. GraphVelo(ATAC)만 향후 arm 후보로 메모. diff --git a/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_methodology-brief.md b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_methodology-brief.md new file mode 100644 index 0000000..b03be2a --- /dev/null +++ b/paper_analysis/epigenomic-lag/wu-2026-velocity-benchmark/wu-2026-velocity-benchmark_methodology-brief.md @@ -0,0 +1,15 @@ +# Wu et al., 2026 (GB velocity benchmark) — methodology-brief + +> 우리 원고(BIOP01) 작업에 바로 쓰는 실행 지침. 근거 = `_core.md`. + +## 즉시 반영 (draft — 영/한 동시, 다른 창과 조율 후) +1. **참고문헌 추가**: Wu Y, Kong C, Liao X, Lin Z, Sun X, Liu J. Comprehensive benchmarking of RNA velocity methods across single-cell datasets. Genome Biology 27:242 (2026). doi:10.1186/s13059-026-04182-z. → [12][13] 옆 세 번째 벤치마크로. (p12/p15 재번호 스크립트로 편입, 앵커는 Background·Positioning의 벤치마크 언급 지점.) +2. **Positioning 문단 보강**: "임베딩 벡터를 채점하는 세 벤치마크(Luo·Huang·Wu)와 달리, 우리는 (a) cell×gene 행렬의 method 간 재현성 (b) ATAC-shuffle 인과 대조 (c) 외부 측정 rate 앵커 (d) 사전등록을 적용한다"로 명시. +3. **데이터 각주**: GSE209878이 Wu의 Data 6로도 쓰였음을 명기(중복이 아니라 다른 질문이라는 근거). + +## 저널 결정(BIOP01-75)에 넘길 판단 자료 +- GB에 종합 velocity 벤치마크가 방금 게재됨 → **GB desk-reject 위험 상향**. CRM base-case의 상대 안전성 ↑. 회의 안건(저널 결정)에 반영. + +## 하지 말 것 +- draft_v2/draft_v2_ko를 지금 임의 수정하지 않는다(다른 창과 draft 동시 편집 금지 규칙). 위 1~3은 **조율 후** 반영. +- Wu의 임베딩 지표(CBDir/ICVCoh)를 우리 지표로 도입하지 않는다 — 우리 축은 행렬 재현성·인과다. From 02b75a85ec268a1c4dc32f588db72e596d102261 Mon Sep 17 00:00:00 2001 From: Geon-Gyu LEE Date: Mon, 27 Jul 2026 22:06:07 +0900 Subject: [PATCH 3/4] =?UTF-8?q?BIOP02-100=20docs:=20README=20=EC=9E=AC?= =?UTF-8?q?=EC=9E=91=EC=84=B1=20=E2=80=94=20=EC=A0=95=EB=B3=B8(main)=20?= =?UTF-8?q?=EA=B8=B0=EC=A4=80=20=EB=A6=AC=ED=8F=AC=C2=B7=EB=B8=8C=EB=9E=9C?= =?UTF-8?q?=EC=B9=98=20=EC=A7=80=EB=8F=84=20+=20=ED=95=98=EB=84=A4?= =?UTF-8?q?=EC=8A=A4=20=EA=B2=8C=EC=9D=B4=ED=8A=B8=20=EB=B0=98=EC=98=81?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 기존 README는 2026-06-14(6792aa4)에서 멈춰 있었다. 그동안 바뀐 것이 README에 하나도 반영되지 않아, 리포의 첫 화면이 실제 상태와 어긋나 있었다. 정정한 것 - 제목이 "BioProject01 — kkkim-pipeline"이었다. 이 파일은 이제 main에 있다. - 하네스 정합성 게이트(harness.yaml + harness_doctor.py + 테스트 18종 + PR CI)가 통째로 빠져 있었다. PR #5→#7로 main에 들어간 것이 README에 없다. - "이 하네스는 OpenClaw로 실행한다"고 단언했다. 2026-07-27 실측 결과 팀 컨테이너 6개 중 openclaw CLI 설치는 0개다. 도입 미결 + 현재 실행수단(Claude Code/codex)으로 정정. - HANDOFF/TODO/SESSION-LOG를 "작업 기록"으로만 적어 새 clone에도 있는 것처럼 읽혔다. .gitignore 로컬 전용임을 명시. - 누락돼 있던 디렉터리 반영: .claude/, harness.yaml, scripts/, harness_after/, evals/, docs/, guide/, onboarding_gglee/, ai_scientist/, artifacts/, manuscript(심볼릭 링크). - paper_analysis 편수: dual-lens 14편 + 스쿱 점검 1편(todorovski)으로 구분. _index/papers.csv의 14는 dual-lens 기준이라 정합하다. 추가한 것 - 브랜치 지도 12개. main과의 차이는 git rev-list 실측(진행/archive/병합완료/미병합). - 연구 결론 한 문단. 단 수치는 싣지 않고 FINDINGS.md를 정본으로 가리킨다 (CLAUDE.md 방법론 주의 6 — 요약본의 숫자를 임계로 쓰지 않는다). - 담당표의 정본이 Confluence와 JIRA임을 명시. 개인별 배정 내역은 리포에 적지 않는다. 검증 - python scripts/harness_doctor.py --repo . --manifest harness.yaml → PASS, phantom_paths=0 - python harness_after/tests/test_harness_doctor.py → 18/18 OK - 작성 중 게이트가 이 README의 팬텀 경로 6건을 실제로 잡았다(브랜치 이름 3, 확장자 3). 게이트가 슬래시 든 백틱 토큰을 경로로 읽는 한계는 본문에 명시. --- README.md | 145 ++++++++++++++++++++++++++++++++++++++++++++++++------ 1 file changed, 129 insertions(+), 16 deletions(-) diff --git a/README.md b/README.md index ef70b65..d5128e0 100644 --- a/README.md +++ b/README.md @@ -1,27 +1,140 @@ -# BioProject01 — `kkkim-pipeline` +# BioProject01 — chromatin→transcription lag 와 RNA velocity 출력의 신뢰도 -**HSPC 연구 단일 작업 브랜치.** Human HSPC 10x Multiome(GSE209878)로 gene별 chromatin→transcription **lag**를 정량하고, baseline epigenomic feature로 drug response timing을 예측한다. 논문 *근거*와 그것을 돌리는 *코드*를 한 브랜치, 두 폴더로 관리한다. +**목표**: gene별 **chromatin→transcription lag**(activation/shutdown 시점차)을 정량해서, baseline +epigenomic feature로 **epigenetic drug response timing** 을 예측한다. +1차 데이터셋 = **Human HSPC 10x Multiome (GSE209878)**. -> 논문 분석 *하네스*(재사용 도구)는 외부 repo **`kakyungkim/paper-analysis-harness`** 에 있다. 새 분석은 거기서 돌리고 산출물만 `paper_analysis/`로 반입한다. 옛 `kkkim-paper-agent` 브랜치는 archive(보존만). +**현재까지의 결론** — 그 전제인 "lag이 method-robust한 양인가"(H1)를 먼저 검정했고, **아니었다.** +lag은 method 간 크기·방향 모두 재현되지 않고, ATAC 셔플 음성대조에서도 변하지 않아 chromatin 생물학이 +아니라 **모델 구조에서 나온 양**으로 판명됐다. 대신 **전사율 α는 method-robust하고 baseline ATAC으로 +예측 가능**하며, 이 순서는 외부 데이터셋 다섯 곳에서 보존된다(그중 하나는 fit 전 봉인한 사전등록 통과). -## 구조 -- `paper_analysis/` — paper 분석 산출물 14편(dual-lens) + `_index/`. 아래 파이프라인 method 선택의 근거. -- `pipeline/hspc-velocity-benchmark/` — 실제 실행 코드 + 논문 산출물 레이어: - - `scripts/`(P0–P5) · `env/` · `DESIGN.md` · `dataset/` · `download_manifest.tsv` — 코드·재현성·Methods. - - `results/` — 표·수치가 되는 요약(csv/md, tracked) · `figures/` — 그림 생성 스크립트(이미지는 ignore) · `manuscript/` — 원고(draft/refs/legends). -- `AGENTS.md` + `skills/` — 파이프라인 분석 하네스 (박상준 `Harness_Baseline` 반입, **OpenClaw/Codex 포맷**, Claude Code 호환). dataset 4종 × `download/preprocessing/model/visualization`. active = `human-hspc-10x-multiome`. -- `SESSION-LOG.md` / `HANDOFF.md` / `TODO.md` — 작업 기록·현황·할 일. +> ⚠️ **수치는 이 문서에서 인용하지 않는다.** 결과·해석의 정본은 +> `pipeline/hspc-velocity-benchmark/results/FINDINGS.md`(★통합 결론)이고, 원고는 `manuscript/draft_v2.md` +> (한국어 `draft_v2_ko.md`)다. 합격 기준·임계는 봉인 문서(`PREREGISTRATION_gse205117.md` 등)에서 +> `파일:줄`로 인용한다 — 발표자료·요약본의 숫자를 임계로 쓰지 않는다(`CLAUDE.md` 방법론 주의 6). + +- GitHub `biospin/BioProject01` · JIRA `BIOP01` · Confluence space `VC` > 프로젝트#01 (→ `Project-Info.md`) +- 정본 브랜치 = **`main`**. 팀 작업 브랜치 = `kkkim-pipeline`. (아래 [브랜치 지도](#브랜치-지도)) + +--- + +## 리포 지도 + +| 경로 | 무엇 | +| --- | --- | +| `pipeline/hspc-velocity-benchmark/` | **연구 본체.** 실행 코드(`scripts/` P0–P5)·격리 env(`env/`)·결과(`results/`)·그림(`figures/`)·원고(`manuscript/`)·설계(`DESIGN.md`)·데이터 출처(`dataset/`, `download_manifest.tsv`, `P0_provenance.md`) | +| `manuscript` | 위 `manuscript/` 로 가는 심볼릭 링크(단축 경로) | +| `paper_analysis/` | 선행연구 **dual-lens 분석 14편**(+ 스쿱 점검 1편) + 색인 `_index/`. 파이프라인 method 선택의 근거 | +| `AGENTS.md` + `skills/` | **데이터셋 분석 하네스** — dataset 4종 × task 4단계. 라우터 = `skills/ROUTES.md` | +| `.claude/` | **논문 생산 하네스** — agent 10종 + 오케스트레이터 Skill(`.claude/skills/paper-production-orchestrator/SKILL.md`) + 글쓰기 규율(`.claude/rules/writing-style.md`) | +| `harness.yaml` | 위 두 하네스의 **구성 SSOT(manifest)**. 문서·코드는 이 파일을 따른다 | +| `scripts/harness_doctor.py` | manifest ↔ 실물 대조 **검진기**(팬텀 역할·팬텀 경로) | +| `harness_after/` | 검진기 자체 테스트 18종(`tests/test_harness_doctor.py`) + 교체용 after 버전 | +| `evals/reproducibility_pilot/` | 사전등록 채점을 재현하는 eval + 회귀 케이스 코퍼스 | +| `docs/` | 랩 지도 `HARNESS.md` · 인프라 정본 `SHARED-INFRA-GUIDE.md` · 로드맵/정합성 보고 | +| `guide/` | 과제·가이드 원문(주차별 과제, 프로젝트 기획서) | +| `onboarding_gglee/` | 온보딩 1~3주차 산출물·회고 | +| `ai_scientist/` | "AI scientist" 구성 개념 문서(하네스 설계 배경) | +| `artifacts/` | 파이프라인 run 로그·리포트 요약 보관 규칙 | +| `CLAUDE.md` | 에이전트 운영 규칙 — 라우팅표·산출물 계약·**완료의 정의(DoD)**·commit 규칙 | +| `BIOP02_LINK.md` | BIOP02(SpatialPathoAgent)와의 cross-reference. **요약본, 정본 아님** | + +--- + +## 하네스와 그 정기검진 + +문서가 실재하지 않는 역할·경로를 가리키면 팀은 그 문서를 계속 믿는다. 2026-07 조사에서 실제로 +그런 결함 4종이 나왔다(팬텀 역할, cwd 의존 침묵 폴백, 게이트 순서 역전, 머지로 인한 `skills/` 41파일 +유실). 개별 수리 대신 **구성 자체를 검증하는 게이트**를 뒀다: + +``` +harness.yaml ← 단일 기준표(어떤 역할·산출물·게이트가 있어야 하는가) + └ scripts/harness_doctor.py ← 기준표 ↔ 실물 대조. 팬텀 역할·팬텀 경로를 FAIL + └ harness_after/tests/test_harness_doctor.py ← 검진기 자체 테스트 18종 + └ .github/workflows/harness-doctor.yml ← PR CI (blocking) +``` + +로컬에서 같은 검사를 돌린다: + +```bash +pip install pyyaml +python harness_after/tests/test_harness_doctor.py # 18/18 이어야 한다 +python scripts/harness_doctor.py --repo . --manifest harness.yaml # PASS 여야 한다 +``` + +`README.md`·`AGENTS.md`·`CLAUDE.md`·`docs/HARNESS.md`·오케스트레이터 SKILL 이 스캔 대상이다. +**이 문서에 백틱으로 쓴 경로도 검사 대상이므로**, 없는 경로를 적으면 PR이 막힌다. + +### 실행 도구 현황 (2026-07-27 실측) + +`skills/` 의 skill·agent 정의는 **OpenClaw/Codex 네이티브 포맷**을 유지한다. 다만 **팀 컨테이너 6개 중 +`openclaw` CLI가 설치된 곳은 0개**다. 현재 실제 실행은 Claude Code / codex 로 하며, OpenClaw 도입 +여부는 미결(회의 상정)이다. **문서는 openclaw 설치를 전제로 쓰지 않는다** — 전제로 쓰면 그 문서는 +아무도 실행할 수 없는 절차가 된다. + +--- + +## 브랜치 지도 + +2026-07-27 기준. `main` 과의 차이는 `git rev-list --count origin/main..` 실측. + +| 브랜치 | 상태 | +| --- | --- | +| main | **정본.** PR #5 → #7 머지로 하네스 게이트·`skills/` 복원 반영 완료 | +| kkkim-pipeline | 팀 작업 브랜치. main 으로 승격하는 경로 | +| gglee | 이건규 작업. PR #5 로 반영 완료 | +| feat/manuscript-condenser | 진행 중 (PR #6 열림) | +| jamie-paper-agent | 진행 중 (cross-paper insight 파이프라인) | +| kkkim-paper-agent · braveji-paper-agent · sezinie-paper-agent | **archive(보존만).** 새 작업은 하지 않는다 | +| epigenomics · braveji/team-owner-mousebrain · fix/BIOP01-22-braveji-env-repro | main 에 병합 완료 — 정리 가능 | +| team-table-update-20260709 | **미병합 1커밋**(팀 담당표 갱신). 당시 README 구조가 지금과 달라 그대로는 적용되지 않는다 | + +> 브랜치 이름은 백틱으로 감싸지 않는다 — 슬래시가 든 백틱 토큰을 정합성 게이트가 **리포 경로**로 +> 읽어 팬텀으로 잡기 때문이다. 게이트가 브랜치 이름과 경로를 구별하지 못하는 것은 알려진 한계다. + +--- ## 빠른 시작 + ```bash -# 1) env (miniforge/mamba) +# 1) 격리 conda env (framework 별로 분리 — CUDA 충돌 회피) bash pipeline/hspc-velocity-benchmark/env/setup_envs.sh -# 2) 데이터 +# 2) 데이터 (GSE209878; 체크섬은 download_manifest.tsv) bash pipeline/hspc-velocity-benchmark/scripts/download_data.sh -# 3) 전처리(P1) +# 3) 공통 전처리 (P1) conda run -n scv-preprocess python pipeline/hspc-velocity-benchmark/scripts/p1_build.py ``` -상세: `pipeline/hspc-velocity-benchmark/{P0_provenance,P1_README,DESIGN,env/README}.md`. -## OpenClaw -이 하네스는 OpenClaw로 실행하는 연습 대상이다. `skills///agents/openai.yaml`이 OpenClaw/Codex agent 정의이고, `AGENTS.md`+`skills/ROUTES.md`가 라우터다. 앞으로 분석을 OpenClaw 기반으로 돌리는 것을 감안해 이 포맷을 유지한다. +상세: `P0_provenance.md` · `P1_README.md` · `DESIGN.md` · `env/README.md`. +서버 접속·GPU 예절·env 위치는 `docs/SHARED-INFRA-GUIDE.md` 가 정본이다(`CLAUDE.md` 에 중복하지 않는다). + +**헤드라인 숫자를 커밋·공개하기 전에** 결정론적으로 재계산해 `FINDINGS.md` 와 대조한다 — +`p3_concordance.py` + `p3_crossdataset_concordance.py` + `p3_scrambled_null.py`. 전체 체크리스트는 +`CLAUDE.md` 의 **완료의 정의(DoD)**. + +--- + +## 이 리포에 없는 것 + +- **`HANDOFF.md` · `TODO.md` · `SESSION-LOG.md`** — 개인 작업기록. `.gitignore` 등재된 **로컬 전용**이라 + 새 clone 에는 없다. 문서가 이들을 필수 산출물로 부르지만 리포에서 찾지 말 것. +- **원본 데이터·대용량 binary**(`*.h5ad`/`*.h5mu`/`*.loom`/PDF) — tracked 는 `*.md`/`*.yaml`/요약 `*.tsv`/코드. +- **conda env 실체** — 팀 공유 서버에 있고 git 미추적. 위치는 `docs/SHARED-INFRA-GUIDE.md`. + +## 팀·추적 + +담당 데이터셋과 소유자의 **정본은 Confluence 프로젝트#01 페이지와 JIRA** 다. +⚠️ 리포 안 `Project-Info.md` 의 팀 표는 그 정본보다 뒤처져 있다 — 인용하지 말고 위를 볼 것. + +commit 메시지 규칙, 언어, 저자 표기는 `CLAUDE.md` 를 따른다. + +--- + +## 출처·라이선스 + +- 데이터셋 분석 하네스(`AGENTS.md` + `skills/`)는 **박상준(@poqopo) `Harness_Baseline`** 에서 반입해 + 이 프로젝트에 맞춘 것이다. 원저작자 박상준 — 원 repo LICENSE 미지정이므로 공유·수정은 동의 전제. +- 논문 생산 하네스(`.claude/` + `docs/HARNESS.md`)는 *Designed by Ka-Kyung Kim, 2026 — reusable + paper-production harness scaffold (CC BY 4.0)* 의 설치본이다. +- 리포 라이선스는 `LICENSE`. From 622da00a897639ba26d85eacddb0141e736ca554 Mon Sep 17 00:00:00 2001 From: kakyungkim Date: Mon, 27 Jul 2026 14:14:26 +0000 Subject: [PATCH 4/4] =?UTF-8?q?docs:=20ai=5Fscientist/=20=EC=A0=9C?= =?UTF-8?q?=EA=B1=B0=20=E2=80=94=20main=EC=9C=BC=EB=A1=9C=20=EC=9D=B4?= =?UTF-8?q?=EA=B4=80?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ai_scientist/ 문서·HTML 시각화를 main 브랜치로 이관(PR docs/ai-scientist-to-main). kkkim-pipeline은 HSPC 파이프라인 작업 브랜치로 유지하고 설계 문서는 main에 둔다. --- ai_scientist/01_overview.md | 37 -- ai_scientist/02_single_lab_harness.md | 107 ----- ai_scientist/03_multi_ai_collaboration.md | 118 ----- ai_scientist/04_design_principles.md | 59 --- ai_scientist/05_component_map.md | 51 --- ai_scientist/README.md | 48 -- ai_scientist/output_v01/README.md | 25 -- ai_scientist/output_v01/index.html | 507 ---------------------- 8 files changed, 952 deletions(-) delete mode 100644 ai_scientist/01_overview.md delete mode 100644 ai_scientist/02_single_lab_harness.md delete mode 100644 ai_scientist/03_multi_ai_collaboration.md delete mode 100644 ai_scientist/04_design_principles.md delete mode 100644 ai_scientist/05_component_map.md delete mode 100644 ai_scientist/README.md delete mode 100644 ai_scientist/output_v01/README.md delete mode 100644 ai_scientist/output_v01/index.html diff --git a/ai_scientist/01_overview.md b/ai_scientist/01_overview.md deleted file mode 100644 index 23c388a..0000000 --- a/ai_scientist/01_overview.md +++ /dev/null @@ -1,37 +0,0 @@ -# 01. 개요 — AI Scientist가 무엇을 자동화하는가 - -## 출발점 - -이 프로젝트가 풀려는 연구 문제는 gene별 chromatin→transcription **lag**(activation/shutdown)을 정량하고, baseline epigenomic feature로 epigenetic drug response timing을 예측하는 것이다. 1차 데이터셋은 Human HSPC 10x Multiome(GSE209878)이다. 이 도메인 문제 자체는 `pipeline/hspc-velocity-benchmark/`가 담당한다. - -AI Scientist 설계의 목표는 이 연구 문제를 푸는 **과정 전체**를 자동화하는 데 있다. 사람이 도구를 하나씩 손으로 돌리는 대신, AI 멤버들이 연구의 각 단계를 나눠 맡아 이어서 돌아가게 한다. - -## 연구 과정을 어떤 단계로 나눴나 - -전통적인 연구 흐름을 AI가 맡을 수 있는 단계로 나누면 다음과 같다. 괄호 안은 이 저장소에서 그 단계를 맡는 주체다. - -1. **논문 탐색·정리**: 선행연구를 찾아 정리하고, 우리 기여를 정직하게 위치시킨다. (`literature-scout`, 그리고 별도 하네스로 돌려 `paper_analysis/`에 반입한 dual-lens 분석 14편) -2. **가설 설정·차별화**: 무엇이 새로운지, 어떤 실험이 가장 싸게 그것을 입증하는지 정한다. (`novelty-strategist`, `research-methodologist`) -3. **실험 설계·감사**: 가설을 검증 가능한 실험으로 바꾸고 누수·통계 위험을 미리 잡는다. (`research-methodologist`) -4. **실험 수행·분석**: 파이프라인을 돌려 eval·통계·cross-dataset 재현을 계산한다. (`hspc-velocity-analyst` + `scripts/` P0–P5) -5. **집필·그림**: 결과 파일에서 원고와 그림을 만든다. (`manuscript-writer` + `figures/figNN_*.py`) -6. **검수·리뷰**: 제출 전 적대적 자체검토와 정식 venue 리뷰를 돌린다. (`paper-critic`, `reviewer`) -7. **발표**: 청중에 맞춰 슬라이드와 발제를 만든다. (`presenter`) - -이 일곱 단계를 사람이 매번 순서대로 부르지 않도록, 자연어 요청을 멤버에 배정하는 라우팅표(`CLAUDE.md`)와 여러 단계를 엮어 실행하는 오케스트레이터 Skill(`paper-production-orchestrator`)을 두었다. 자세한 구조는 [02_single_lab_harness.md](02_single_lab_harness.md)에서 다룬다. - -## 왜 한 명의 AI로 끝내지 않았나 - -연구 팀은 한 사람이 아니다. 이 프로젝트도 데이터셋별로 담당자가 다르고(mouse brain, SHARE-seq, human brain, HSPC 등), 팀원마다 쓰는 AI도 Claude, Codex, Gemini로 갈린다. 한 AI가 논문 한 편을 끝까지 끌고 가는 구조(레이어 A)만으로는 이 협업을 담을 수 없다. - -그래서 두 번째 레이어를 설계했다. 팀원 A의 AI가 끝낸 작업을 팀원 B의 AI가 자동으로 이어받는 인계 체계다. 신호를 JIRA 상태 전환 하나로 일원화하고, 인계 맥락을 정형화된 Handoff 코멘트로 강제하며, 모든 AI가 같은 MCP 설정으로 JIRA·GitHub를 읽고 쓰게 한다. 자세한 구조는 [03_multi_ai_collaboration.md](03_multi_ai_collaboration.md)에서 다룬다. - -## 두 레이어가 공유하는 발상 - -레이어 A와 B는 다른 문제를 풀지만 같은 원리 위에 서 있다. - -- **다음 주체가 하나만 읽어도 착수할 수 있게 한다.** A에서는 결과 파일(`results/FINDINGS.md`), B에서는 JIRA Handoff 코멘트가 그 역할을 한다. -- **자동화하되 사람 게이트를 남긴다.** A에서는 공개·main 병합, B에서는 초기 도입기의 Slack 승인이 사람 손을 거친다. -- **폭주와 비용을 구조로 막는다.** A에서는 검증 게이트가 근거 없는 주장을, B에서는 Hop Count 상한과 큐가 무한 인계와 토큰 낭비를 막는다. - -이 공통 원리는 [04_design_principles.md](04_design_principles.md)에 모아 두었다. diff --git a/ai_scientist/02_single_lab_harness.md b/ai_scientist/02_single_lab_harness.md deleted file mode 100644 index 6bbbd35..0000000 --- a/ai_scientist/02_single_lab_harness.md +++ /dev/null @@ -1,107 +0,0 @@ -# 02. 레이어 A — 단일 랩 자동화 (한 AI가 논문 한 편을 끝까지) - -한 연구자의 연구 과정 전체를 여러 agent 멤버가 나눠 맡아 자동으로 돌리는 구조다. 이 하네스를 "하나의 연구 랩"으로 보는 지도가 `docs/HARNESS.md`이고, 라우팅과 산출물 계약 요약은 `CLAUDE.md`의 *Agent routing & artifact contract* 절에 있다. - -핵심 발상은 이렇다. agent는 직원이 아니라 랩의 **멤버(연구원)**이고, 사람과 메인 루프가 랩을 이끄는 **PI**다. PI는 무엇을 할지 정하고 승인·공개를 책임지되, 실제 작업은 멤버가 파일로 주고받으며 이어서 한다. - -## 1. 멤버 명부 - -`.claude/agents/`에 정의된 멤버는 다음과 같다. 하나(`hspc-velocity-analyst`)만 이 프로젝트 도메인 전용이고, 나머지는 다른 논문에도 재사용할 수 있게 만들었다. - -| 멤버 | 벤치 | 역할 | -| --- | --- | --- | -| `hspc-velocity-analyst` | 분석실 | 도메인 슬롯. HSPC velocity-lag 파이프라인(P0–P5)·eval·통계·cross-dataset 실행/확장, 결과 파일 유지 | -| `literature-scout` | 문헌·기획 | 선행연구 탐색, 정직한 포지셔닝, related work | -| `novelty-strategist` | 문헌·기획 | 차별화 각도와 가장 값싼 입증 실험 제안 | -| `research-methodologist` | 문헌·기획 | 가설·기여문·실험설계, 누수·통계 감사 | -| `manuscript-writer` | 집필실 | 프리프린트·저널·블로그 본문 초안과 그림 연계 | -| `presenter` | 집필실 | 청중 맞춤 슬라이드·발제 | -| `paper-critic` | 심사·QA | 제출 전 적대적 자체검토와 그림 시각 QA | -| `reviewer` | 심사·QA | 정식 venue 스타일 공식 리뷰(선택) | -| `paper-orchestrator` | 코디네이션 | 멀티 agent 작업의 **계획**만 수립(실행은 PI) | -| `design` | 엔지니어링 | 로고·아이콘·브랜드·그림 미감 | -| 그림 생성 스크립트 | 엔지니어링 | `figures/figNN_*.py`. 결과 파일에서 그림 생성·번호 정합 | - -그림 생성을 agent가 아니라 결정론적 스크립트로 둔 점이 설계상의 선택이다. `manuscript-writer`가 스크립트를 실행해 결과 파일로부터 그림을 만들고, 단순 재생성이면 메인 루프가 직접 돌린다. 숫자를 손으로 하드코딩하지 않고 결과 파일에서만 뽑게 해 재현성을 지킨다. - -## 2. 자연어 라우팅 — 누가 시작할지 사람이 매번 안 정한다 - -요청에 agent 이름이 없어도 `CLAUDE.md`의 라우팅표가 자연어 요청을 멤버에 배정한다. 예를 들면 이렇게 나뉜다. - -- "분석 돌려줘 / 재실행 / eval·통계 / cross-dataset 재현" → `hspc-velocity-analyst` -- "프리프린트·섹션 써줘 / 그림 만들어줘" → `manuscript-writer` -- "선행연구 / 스쿱 확인" → `literature-scout` -- "차별화 각도 / 뭘 새로 해야 하나" → `novelty-strategist` -- "가설·실험설계 점검·감사" → `research-methodologist` -- "제출 전 자체검토 / 그림 QA" → `paper-critic` -- "발표자료 / 슬라이드" → `presenter` - -여러 단계를 엮는 요청("분석→집필→그림→검수까지", "critic 지적 반영해")은 단일 멤버가 아니라 오케스트레이터 Skill로 보낸다. - -## 3. 오케스트레이터 — 여러 단계를 정해진 순서로 - -`paper-production-orchestrator` Skill(`.claude/skills/paper-production-orchestrator/SKILL.md`)이 논문 생산 루프의 입구다. 메인 루프(PI)가 이 Skill을 실행하며 멤버를 순서대로 부른다. subagent는 subagent를 못 부르므로, "계획만 짜는" `paper-orchestrator` agent와 달리 실제 실행은 이 Skill이 맡는다. - -실행 흐름은 다음과 같다. - -``` -0. 단일 컨텍스트 로드 — manuscript/PAPER_DIRECTION.md 를 먼저 읽는다 - (현재 thesis · claim 등급표 · loop 규율 · 진행상태. 멤버 호출 전 이 문서를 넘긴다) -1. 모드 분기 — 풀 파이프라인 / 부분 재실행 / 하류만 다시 -2. (선택) 기획·근거 — research-methodologist / literature-scout / novelty-strategist -2.5 claim-defensibility 게이트 — headline·novelty claim이 본문에 들어가기 전 필수 -3. 분석·eval — hspc-velocity-analyst → results/FINDINGS.md -4. 집필 + 그림 — manuscript-writer → draft_v2.md + draft_v2_ko.md, figures/*.png -5. 검수 — paper-critic (적대적 + 그림 시각 QA) -6. 수정 — manuscript-writer 가 지적 반영 -7. (선택) 정식 리뷰 — reviewer → REVIEW--.md -8. 검증 게이트 — 헤드라인 숫자 결정론적 재계산 (실패하면 멈추고 사람에게 보고) -9. (선택) 발표 — presenter -``` - -핵심은 **부분 재실행**이다. 이미 만들어진 산출물이 있으면 요청한 단계만 다시 돌리고 나머지는 기존 파일을 재사용한다. "그림만 다시"면 4단계만, "최신 결과로 본문 갱신"이면 변경 지점의 하류 단계만 돌린다. - -## 4. 산출물 계약 — 대화가 아니라 파일로 넘긴다 - -멤버는 중간 결과를 대화에만 남기지 않고 정해진 파일로 넘긴다. 다음 멤버는 그 파일을 읽고 이어서 일한다. - -| 단계 | Writer | 산출물 | 다음이 읽음 | -| --- | --- | --- | --- | -| 분석·eval | hspc-velocity-analyst | `results/FINDINGS.md` + `results/*.csv` + `results/*.md` | 집필·검수 | -| 집필·그림 | manuscript-writer | `manuscript/draft_v2.md` + `draft_v2_ko.md`(영/한 동시), `figures/*.png` | 검수·리뷰·발표 | -| 검수 | paper-critic / reviewer | `manuscript/REVIEW--.md` | 집필(수정) | -| 발표 | presenter | 슬라이드·발제 | 사람 | -| 상태 핸드오프 | 전원 | `HANDOFF.md`, `TODO.md`, `SESSION-LOG.md` | 다음 세션 | - -이 계약 덕분에 멤버가 교체되거나 세션이 끊겨도 작업이 이어진다. 다음 주체가 산출물 파일 하나만 읽으면 착수할 수 있다는 기준을 지킨다. - -## 5. 실험 실행 엔진 — 파이프라인 P0–P5 - -분석 단계의 실제 계산은 `pipeline/hspc-velocity-benchmark/scripts/`가 담당한다. `hspc-velocity-analyst`가 이 스크립트들을 돌려 결과 파일을 만든다. 내부 단계 표기는 P0부터 P5까지다. - -- **P0: 다운로드·provenance.** `download_data.sh`로 GSE209878를 받고 `download_manifest.tsv`(sha256)와 `P0_provenance.md`를 남긴다. -- **P1: 통일 전처리.** `p1_build.py`가 공통 branch를 만든다. 여기서 preprocessing 차이와 method 차이를 분리한다. -- **P2: velocity method 실행.** `p2_multivelo.py`, `p2_moflow.py`, `p2_crakvelo_*`, `p2_multivelovae.py` 등으로 여러 method를 같은 전처리 위에서 돌린다. -- **P3: 재현성 검증.** `p3_concordance.py`, `p3_crossdataset_concordance.py`, `p3_scrambled_null.py`로 method 간·dataset 간 일치도와 null을 계산한다. -- **P4: permutation FDR.** gene 단위 다중검정을 통제한다. -- **P5: bootstrap 안정성.** shuffle/seed 변이 audit(`p10*`)까지 포함해 결과의 흔들림을 잰다. - -method 선택의 근거는 `DESIGN.md`와 `paper_analysis/`의 dual-lens 분석 14편에 있다. 프레임워크별로 conda env를 격리(`env/`)해 의존성 충돌을 막는다. - -## 6. 게이트 — 자동화가 넘지 못하는 선 - -이 랩은 전부를 자동으로 밀지 않는다. 두 종류의 게이트가 있다. - -**검증 게이트(커밋·공개 전).** 헤드라인 숫자를 결정론적으로 재계산해 결과 파일과 대조한다. - -```bash -cd pipeline/hspc-velocity-benchmark/scripts -conda run --no-capture-output -n scv-preprocess python p3_concordance.py -conda run --no-capture-output -n scv-preprocess python p3_crossdataset_concordance.py --dataset human_brain -conda run --no-capture-output -n scv-preprocess python p3_scrambled_null.py -# 출력 숫자를 results/FINDINGS.md 와 대조. 불일치면 멈추고 사람에게 보고. -``` - -**사람 승인 게이트.** 프리프린트·블로그 외부 공개와 main 병합은 사람이 승인한다. 저자·소속·IP·corresponding email이 확정되기 전에는 공개를 보류한다(원고에 ``로 표시). 작업 브랜치 `kkkim-pipeline`에 대한 커밋·push는 자동으로 수행하되, 이 검증·공개 게이트는 유지한다. - -claim 자체에도 게이트가 있다. headline claim은 반증기준, 가장 값싼 make-or-break 검정, advisor 확인을 통과하기 전에는 PROVISIONAL로 두고 본문에 넣지 않는다. within-method 적합 품질을 cross-method 재현성으로 승격하지 않는다는 규율(2층 융합 금지)도 여기에 든다. diff --git a/ai_scientist/03_multi_ai_collaboration.md b/ai_scientist/03_multi_ai_collaboration.md deleted file mode 100644 index f82cf55..0000000 --- a/ai_scientist/03_multi_ai_collaboration.md +++ /dev/null @@ -1,118 +0,0 @@ -# 03. 레이어 B — 멀티 AI 협업 인계 (여러 AI가 팀으로 이어달리기) - -팀원마다 다른 AI(Claude, Codex, Gemini)를 쓰고 결과물은 JIRA·Confluence·Git으로 공유한다. 문제는 한 작업이 끝나도 다음 담당자의 AI에 신호가 자동으로 가지 않는다는 점이다. 사람이 확인할 때까지 대기가 생기고 인계가 지연된다. - -이 레이어는 그 인계를 자동화한다. 설계 문서는 두 편이다. - -- `guide/ai-handoff-architecture-guide.md`: **무엇을** 인계하나. JIRA 상태 전환 신호에서 다음 AI 실행까지의 4계층 구조. -- `guide/openclaw-claude-guide.md`: **어떻게 싸고 안정적으로** 돌리나. OpenClaw와 메시지 큐로 그 구조를 실현하고 비용을 통제하는 방법. - -## 1. 설계 원칙 - -| 원칙 | 내용 | -| --- | --- | -| 단일 신호원 | 인계 신호는 JIRA 상태 전환만 쓴다. Git 머지 등은 JIRA 상태로 수렴시킨다 | -| 사람 승인 우선 | 초기엔 Slack 원클릭 승인 후 실행. 신뢰가 쌓이면 단계적으로 자동화 | -| 최소 권한 | AI별 서비스 계정 분리, 프로젝트 단위 권한, main 직접 push 금지 | -| 폭주 방지 | 티켓당 자동 인계 횟수 상한(기본 5회), 실패 시 즉시 사람 에스컬레이션 | - -## 2. 4계층 아키텍처 - -``` -① 이벤트 소스 (기존 스택) - JIRA 상태 전환(Ready for AI) / Git PR 머지 → JIRA 상태 자동 전환 - │ Webhook (JIRA Automation → HTTP POST) - ▼ -② 이벤트 허브 (신규) - Webhook 수신 → Next Agent 필드로 분기 → (선택) Slack 승인 → 워커 호출 - 실패 시 ai-failed 라벨 + Slack 알림 - │ Execute / SSH / HTTP - ▼ -③ AI 워커 (신규) - run_agent.sh - ├ claude -p ... (Claude Code headless) - ├ codex exec ... (Codex CLI 비대화) - └ gemini -p ... (Gemini CLI 비대화) - │ MCP (공통 mcp.json) - ▼ -④ MCP 공통 - Atlassian 원격 MCP → JIRA 이슈·코멘트·상태, Confluence - GitHub MCP → 저장소, PR, 이슈 - -작업 완료 → AI가 MCP로 JIRA 상태 전환 → 다시 ①의 신호 발생 → 체인 반복 -``` - -### 인계 루프 (티켓 생애주기) - -1. AI나 사람이 작업을 완료한다. 커밋·PR·문서와 함께 **Handoff 코멘트**를 남긴다. -2. JIRA 상태를 `Ready for AI`로 전환하고 `Next Agent` 필드를 지정한다. -3. JIRA Automation이 이벤트 허브로 웹훅을 보낸다. -4. 허브가 `Next Agent` 값으로 분기하고, 초기엔 Slack 승인을 거친다. -5. 해당 AI 워커가 실행되어 MCP로 티켓·코드 맥락을 읽고 작업한다. -6. 완료하면 1번으로 돌아간다. 체인이 이어진다. - -## 3. Handoff 코멘트 — 인계 맥락의 정형화 - -모든 AI의 규칙 파일(CLAUDE.md / AGENTS.md / GEMINI.md)에 같은 템플릿을 강제한다. - -```markdown -## Handoff -- 완료한 것: (요약 3줄 이내) -- 산출물: (커밋 해시 / PR 링크 / Confluence 페이지 링크) -- 다음 작업: (다음 AI가 해야 할 일, 구체적으로) -- 제약/주의: (건드리면 안 되는 것, 실패했던 접근) -- Next Agent: claude | codex | gemini | human -``` - -기준은 하나다. **다음 워커가 이 코멘트 하나만 읽어도 착수할 수 있어야 한다.** 이것이 레이어 A의 산출물 계약과 같은 발상이다. A는 파일로, B는 JIRA 코멘트로 맥락을 넘긴다. - -## 4. 공통 MCP — 모든 AI가 같은 방식으로 읽고 쓴다 - -`mcp.json` 하나를 설정 전용 저장소(`agent-config`)로 버전 관리하고, 세 AI에 같은 서버 정의를 물린다. 연결 대상은 Atlassian 원격 MCP(JIRA·Confluence)와 GitHub MCP다. 세 도구 모두 MCP 표준을 따르므로 서버 정의는 그대로 재사용하고 파일 형식만 각 도구에 맞게 바꾼다. 토큰은 파일에 직접 쓰지 않고 환경변수·시크릿 매니저로 주입한다. - -## 5. OpenClaw로 실현하기 — 허브와 워커를 대체 - -`ai-handoff-architecture-guide.md`는 이벤트 허브로 n8n을, 워커로 공용 서버의 `run_agent.sh`를 상정한다. `openclaw-claude-guide.md`는 그 ②+③(허브+워커)을 **OpenClaw와 메시지 큐로 대체**하는 경로를 제시한다. 별도 n8n·워커 서버를 세우지 않고 같은 인계 루프를 돌린다. - -| 인계 가이드 계층 | 원 구성 | OpenClaw로 실현 | -| --- | --- | --- | -| ① 이벤트 소스 | JIRA 상태 전환 / PR 머지 | 그대로 유지 | -| ② 이벤트 허브 | n8n | 메시지 큐 브리지 + OpenClaw Webhooks 플러그인 | -| ③ AI 워커 | 공용 서버 + `run_agent.sh` + `claude -p` | OpenClaw 세션(인증·모델선택·thinking 레벨을 OpenClaw가 관장) | -| ④ MCP 공통 | `mcp.json` | 동일. OpenClaw 세션에도 같은 MCP 서버를 물린다 | - -허브를 n8n으로 갈지 OpenClaw 웹훅+큐로 갈지는 팀 규모로 정한다. GUI 워크플로와 Slack 승인 버튼이 필요하면 n8n, 비용 통제를 한곳에서 하고 워커 서버 관리를 줄이고 싶으면 OpenClaw다. - -메시지 큐를 앞에 두는 이유는 안정성과 비용이다. 웹훅을 허브에 직결하면 허브가 재시작 중일 때 이벤트를 잃는다. 큐는 고속 이벤트 유입과 느린 Claude 처리를 분리한다. 브리지 컨슈머가 지켜야 할 네 가지는 다음과 같다. - -1. **ack는 Claude 처리 성공 이후에만.** 실패하면 ack하지 않고 데드레터큐로 격리한다. 이것이 인계 가이드의 "실패 시 상태 유지·자동 재시도 금지"를 자연히 만족한다. -2. **동시 처리 수 제한.** 레이트 리밋과 비용을 통제한다. -3. **멱등성·세션 키.** 재시도가 중복 인계를 만들지 않게 한다. -4. **병합(coalescing).** 같은 티켓의 연속 이벤트를 하나로 합쳐 Claude 호출 수 자체를 줄인다. - -## 6. 비용 — 인계 체인은 호출을 곱셈으로 늘린다 - -AI-to-AI 인계는 한 티켓이 여러 AI를 연쇄 호출하므로 단발 실행보다 토큰 지출이 배로 뛴다. 그래서 비용 레버가 인계 자동화에서 더 중요해진다. - -1. **병합**: 같은 티켓에 상태전환·코멘트 이벤트가 쏟아져도 브리지가 하나로 합쳐 호출 1회로. -2. **모델 티어링**: 저위험 작업(리뷰·테스트·문서화)은 Sonnet/Haiku로, 핵심 분석만 Opus로. `Next Agent`별로 모델을 다르게 물린다. -3. **프롬프트 캐싱**: 티켓 단위 sessionKey로 인계 맥락을 재사용해 입력 토큰을 줄인다. -4. **우선순위 큐**: 비싼 Opus 인계와 값싼 Sonnet 인계를 다른 큐로 분리 라우팅한다. -5. **DLQ**: poison 티켓이 무한 재인계로 과금되는 것을 막는다. - -## 7. 보안·승인 게이트 - -- **승인 게이트**: 도입 초기엔 Slack 승인 필수. OpenClaw 경로에서는 브리지가 큐→OpenClaw POST 직전에 Slack "Send-and-Wait"를 두거나, 웹훅을 수동 트리거로 둔다. -- **최소 권한**: AI별 서비스 계정 분리, JIRA는 해당 프로젝트만, main 직접 push 금지(브랜치+PR). -- **서명 검증 2구간**: JIRA Automation의 `X-Handoff-Token`과 OpenClaw webhook `secret`을 둘 다 건다. 이벤트 소스와 브리지 사이, 브리지와 OpenClaw 사이를 모두 검증한다. -- **토큰 비노출**: API 키·PAT는 환경변수·시크릿 매니저로만. `.mcp.json`에 토큰 직접 기입 금지. - -## 8. 도입 로드맵 - -| 주차 | 목표 | 산출물 | -| --- | --- | --- | -| 1주차 | JIRA 필드·워크플로·Automation + 허브 설치, Slack 알림까지만 | 인계 발생 즉시 알림(자동 실행 없음) | -| 2~3주차 | AI CLI·MCP 공통 설정·워커 구축, Slack 승인 후 반자동 | 첫 AI-to-AI 인계 파일럿 1건 | -| 4주차~ | 저위험 작업부터 승인 생략, 인계 상한·모니터링 정착 | 제한적 완전 자동 체인 + 비용 레버 계측 | - -설치 절차 전체는 `guide/ai-handoff-architecture-guide.md` §4에, OpenClaw판 세부는 `guide/openclaw-claude-guide.md` §3에 있다. diff --git a/ai_scientist/04_design_principles.md b/ai_scientist/04_design_principles.md deleted file mode 100644 index d73a7a2..0000000 --- a/ai_scientist/04_design_principles.md +++ /dev/null @@ -1,59 +0,0 @@ -# 04. 설계 원칙 — 두 레이어를 관통하는 것 - -레이어 A(단일 랩 자동화)와 레이어 B(멀티 AI 인계)는 다른 문제를 풀지만 같은 원리 위에 서 있다. 이 원리들이 AI Scientist 설계의 뼈대다. - -## 1. 다음 주체가 하나만 읽어도 착수할 수 있게 한다 - -두 레이어 모두 작업 맥락을 대화가 아니라 **정형화된 산출물**로 넘긴다. - -- 레이어 A: 결과 파일(`results/FINDINGS.md`), 원고(`draft_v2.md`), 상태 문서(`HANDOFF.md`). -- 레이어 B: JIRA Handoff 코멘트(완료한 것·산출물·다음 작업·제약·Next Agent). - -기준은 같다. 다음 멤버나 다음 AI가 그 산출물 하나만 읽으면 곧바로 일을 시작할 수 있어야 한다. 이 규율이 있어 멤버가 바뀌거나 세션이 끊겨도 작업이 이어지고, 매번 처음부터 다시 브리핑할 필요가 없다. - -## 2. 자동화하되 사람 게이트를 남긴다 - -전부를 자동으로 밀지 않는다. 되돌리기 어렵거나 외부로 나가는 지점에는 사람이 선다. - -- 레이어 A: 프리프린트·블로그 공개와 main 병합은 사람이 승인한다. 저자·소속·corresponding email이 확정되기 전에는 공개를 보류한다. -- 레이어 B: 도입 초기엔 모든 인계에 Slack 승인을 건다. 신뢰가 쌓이면 저위험 작업부터 단계적으로 승인을 생략한다. - -승인 게이트는 고정이 아니라 성숙도에 따라 옮긴다. 처음엔 촘촘하게, 검증되면 넓게. 이 단계적 완화가 두 레이어에 공통으로 들어 있다. - -## 3. 검증 게이트로 근거 없는 주장을 막는다 - -자동화가 그럴듯하지만 틀린 결과를 통과시키지 않도록, 커밋·공개 전에 숫자를 다시 계산해 대조한다. - -- 레이어 A: 헤드라인 숫자를 결정론적 스크립트로 재계산해 결과 파일과 대조하고, 불일치하면 멈추고 사람에게 보고한다. claim 자체도 반증기준·make-or-break 검정·advisor 확인을 통과하기 전에는 본문에 넣지 않는다. -- weak는 zero가 아니다. 통계가 뒷받침하지 않는 우월·재현 주장을 금지한다. - -## 4. 폭주와 비용을 구조로 막는다 - -AI 체인은 스스로를 무한히 호출하거나 토큰을 쏟아낼 수 있다. 이걸 사람의 주의가 아니라 구조로 막는다. - -- 레이어 B: 티켓당 자동 인계 상한(Hop Count < 5), 실패 메시지의 DLQ 격리, 동시 처리 수 제한. -- 비용 레버: 이벤트 병합, 모델 티어링(저위험은 Sonnet/Haiku, 핵심만 Opus), 프롬프트 캐싱, 우선순위 큐. - -인계 체인은 한 티켓이 여러 AI를 거치며 호출을 곱셈으로 늘리므로, 이 레버들이 단발 실행보다 더 중요해진다. - -## 5. 최소 권한과 격리 - -- AI별 서비스 계정을 분리해 JIRA·Git 이력을 추적한다. -- JIRA는 해당 프로젝트만, GitHub는 PR 권한만 준다. -- main 브랜치 직접 push를 금지하고 항상 브랜치와 PR로 간다. -- 자동 승인 옵션(`--permission-mode acceptEdits`, `--yolo`)은 격리된 작업 디렉토리와 최소 권한 계정을 전제로만 쓴다. -- 토큰·API 키는 코드나 코멘트에 남기지 않고 환경변수·시크릿 매니저로만 관리한다. - -## 6. 표준 포맷과 재사용 - -멤버 정의와 라우터를 특정 프로젝트에 묶지 않고 재사용할 수 있게 만들었다. - -- 논문 생산 하네스는 재사용 스캐폴드(CC BY 4.0)로 설계했고, 도메인 전용 슬롯 하나(`hspc-velocity-analyst`)만 이 프로젝트가 채웠다. -- 분석 하네스는 OpenClaw/Codex 네이티브 포맷(`AGENTS.md` + `skills/ROUTES.md` + `openai.yaml`)을 유지해 OpenClaw로 바로 실행하고 Claude Code에서도 동작한다. -- MCP 표준을 따르므로 서버 정의를 세 AI가 그대로 재사용한다. - -이 표준화가 있어 새 데이터셋·새 논문·다른 AI로 옮겨도 구조를 다시 짜지 않는다. - -## 7. 근거와 코드를 분리한다 - -이 프로젝트는 method 선택의 **근거**(`paper_analysis/`의 dual-lens 분석 14편)와 그 근거로 데이터를 돌리는 **코드**(`pipeline/`)를 한 브랜치, 두 폴더로 나눴다. 어떤 method와 confound를 쓸지의 판단(근거)과 실제 실행(코드)을 섞지 않아, 판단이 바뀌면 근거 레이어만, 실행이 바뀌면 코드 레이어만 고친다. diff --git a/ai_scientist/05_component_map.md b/ai_scientist/05_component_map.md deleted file mode 100644 index ef66b81..0000000 --- a/ai_scientist/05_component_map.md +++ /dev/null @@ -1,51 +0,0 @@ -# 05. 컴포넌트 매핑 — 설계 요소가 저장소 어디에 있나 - -AI Scientist 설계의 각 요소가 실제로 어느 파일에 구현·문서화돼 있는지 정리한 지도다. 이 폴더(`ai_scientist/`)는 설계를 **설명**하고, 아래 파일들이 그 설계를 **구현**한다. - -## 레이어 A — 단일 랩 자동화 - -| 설계 요소 | 저장소 위치 | -| --- | --- | -| 랩 구조 지도(멤버 명부·관계도·JD) | `docs/HARNESS.md` | -| 라우팅표 + 산출물 계약 요약 | `CLAUDE.md` (*Agent routing & artifact contract* 절) | -| 멤버 정의 8종 | `.claude/agents/{hspc-velocity-analyst,literature-scout,novelty-strategist,research-methodologist,manuscript-writer,presenter,paper-critic,paper-orchestrator,design}.md` | -| 오케스트레이터(실행 입구) | `.claude/skills/paper-production-orchestrator/SKILL.md` | -| 단일 컨텍스트(thesis·claim 등급표·loop 규율) | `pipeline/hspc-velocity-benchmark/manuscript/PAPER_DIRECTION.md` | -| 분석 실행 엔진(P0–P5) | `pipeline/hspc-velocity-benchmark/scripts/` (`download_data.sh`, `p1_build.py`, `p2_*.py`, `p3_*.py`, `p10*` 등) | -| method 선택 근거 | `pipeline/hspc-velocity-benchmark/DESIGN.md`, `paper_analysis/`(dual-lens 14편) | -| 실험 env 격리 | `pipeline/hspc-velocity-benchmark/env/` | -| 분석 산출물 계약 | `pipeline/hspc-velocity-benchmark/results/FINDINGS.md` + `results/*.csv` + `results/*.md` | -| 집필·그림 산출물 | `pipeline/hspc-velocity-benchmark/manuscript/draft_v2{,_ko}.md`, `figures/figNN_*.py` | -| 검수·리뷰 산출물 | `manuscript/REVIEW--.md` | -| 검증 게이트 스크립트 | `scripts/p3_concordance.py`, `p3_crossdataset_concordance.py`, `p3_scrambled_null.py` | -| 글쓰기 규율(한국어 윤문) | `.claude/rules/writing-style.md` | -| 상태 핸드오프 | `HANDOFF.md`, `TODO.md`, `SESSION-LOG.md` | - -## 레이어 B — 멀티 AI 협업 인계 - -| 설계 요소 | 저장소 위치 | -| --- | --- | -| 인계 아키텍처(4계층·인계 루프·설치 가이드) | `guide/ai-handoff-architecture-guide.md` | -| OpenClaw 실현(허브+워커 대체·메시지 큐·비용 레버) | `guide/openclaw-claude-guide.md` | -| 분석 하네스 project frame(OpenClaw/Codex 네이티브 포맷) | `AGENTS.md` (dataset 라우팅을 `skills/ROUTES.md`에 위임) | -| dataset→task 스킬 트리(`skills/ROUTES.md`, `skills///{SKILL.md,agents/openai.yaml}`) | 이 브랜치 체크아웃에는 없다. `AGENTS.md`·`README.md`가 규정하는 포맷이며, 실제 스킬 트리는 OpenClaw로 돌릴 때 채운다 | -| MCP 공통 설정 | `.mcp.json` (설계 목표는 `agent-config` 저장소로 버전 관리) | -| 팀·역할·AI 계정 매핑 | `Project-Info.md` (데이터셋 담당자 ↔ github·atlassian·slack·openclaw bot) | -| JIRA·Confluence 좌표 | `Project-Info.md` (JIRA space `BIOP01`, Confluence space `VC`) | - -## 두 레이어의 접점 - -| 공유 요소 | 레이어 A에서 | 레이어 B에서 | -| --- | --- | --- | -| 인계 계약 | 결과 파일(`results/FINDINGS.md`) | JIRA Handoff 코멘트 | -| 사람 게이트 | 공개·main 병합 승인 | 초기 Slack 승인 | -| 폭주·비용 방지 | 검증 게이트, claim 등급 | Hop Count 상한, 큐·DLQ, 모델 티어링 | -| 실행 도구 | Claude Code(agent·Skill) | OpenClaw 세션 또는 `run_agent.sh` | -| 라우터 포맷 | `CLAUDE.md` 라우팅표 | `Next Agent` 필드 → 브리지 분기 | - -## 읽는 순서 제안 - -1. 전체 그림만 빠르게: 이 폴더 [README.md](README.md)와 [01_overview.md](01_overview.md). -2. 단일 랩이 어떻게 도나: [02_single_lab_harness.md](02_single_lab_harness.md) → `docs/HARNESS.md` → `.claude/skills/paper-production-orchestrator/SKILL.md`. -3. 여러 AI가 어떻게 이어달리나: [03_multi_ai_collaboration.md](03_multi_ai_collaboration.md) → `guide/ai-handoff-architecture-guide.md` → `guide/openclaw-claude-guide.md`. -4. 왜 이렇게 설계했나: [04_design_principles.md](04_design_principles.md). diff --git a/ai_scientist/README.md b/ai_scientist/README.md deleted file mode 100644 index d1c4d1a..0000000 --- a/ai_scientist/README.md +++ /dev/null @@ -1,48 +0,0 @@ -# ai_scientist/ — AI Scientist 설계 정리 - -이 폴더는 이번 프로젝트에서 **AI Scientist**(인공지능이 연구 도구를 넘어 연구 과정 전반을 자동화하고, 여러 연구자가 함께 쓰도록 만든 구조)를 어떻게 설계했는지 한곳에 정리한 문서다. 새 코드를 만드는 것이 아니라, `kkkim-pipeline` 브랜치에 이미 흩어져 구현·기록된 설계를 확인해 하나의 지도로 묶었다. - -## 무엇을 다루나 - -목표는 두 가지였다. - -1. **연구 과정 전반의 자동화**: 논문 탐색과 정리, 가설 설정, 실험 수행, 그림·집필, 검수, 발표까지를 사람이 매 단계 손으로 잇지 않고 하나의 흐름으로 돌린다. -2. **여러 연구자와 협업하는 구조**: 팀원마다 다른 AI(Claude, Codex, Gemini)를 쓰더라도, 작업을 서로에게 자동으로 넘기고 이어받는 인계 체계를 표준화한다. - -이 두 목표는 각각 하나의 레이어로 설계했고, 서로 맞물린다. - -| 레이어 | 무엇인가 | 이 저장소의 구현·근거 | -| --- | --- | --- | -| **A. 단일 랩 자동화** | 한 연구자의 연구 과정 전체를 agent 멤버들이 나눠 맡아 자동으로 돌리는 "AI 연구 랩" | `.claude/agents/` 8종 + `paper-production-orchestrator` Skill + `AGENTS.md`/`skills/` 라우터 + 파이프라인 `scripts/`(P0–P5). 지도 = `docs/HARNESS.md` | -| **B. 멀티 AI 협업 인계** | 여러 연구자·여러 AI가 JIRA 상태 신호로 작업을 자동 인계하는 체계 | `guide/ai-handoff-architecture-guide.md` + `guide/openclaw-claude-guide.md` | - -레이어 A는 "AI 한 명이 논문 한 편을 어떻게 끝까지 끌고 가나"를, 레이어 B는 "그런 AI 여럿이 팀으로 어떻게 이어달리나"를 설계한다. A가 랩 안의 분업이라면 B는 랩과 랩, 사람과 사람 사이의 배턴 터치다. - -## 문서 구성 - -- [01_overview.md](01_overview.md) — AI Scientist가 무엇을 자동화하는지와 전체 그림 -- [02_single_lab_harness.md](02_single_lab_harness.md) — 레이어 A: 단일 랩 자동화(멤버 명부, 논문 생산 루프, 파이프라인, 게이트) -- [03_multi_ai_collaboration.md](03_multi_ai_collaboration.md) — 레이어 B: 멀티 AI 인계 자동화(JIRA→허브→워커→MCP, OpenClaw+큐) -- [04_design_principles.md](04_design_principles.md) — 두 레이어를 관통하는 설계 원칙 -- [05_component_map.md](05_component_map.md) — 설계 요소와 실제 저장소 파일의 매핑 - -## 한눈에 보는 전체 그림 - -``` - 사람 = PI (방향 설정 · 승인 · 공개 게이트) - │ - ┌───────────────────────────┴───────────────────────────┐ - │ │ - 레이어 A: 단일 랩 자동화 레이어 B: 멀티 AI 협업 인계 - (한 AI가 논문 한 편을 끝까지) (여러 AI가 팀으로 이어달리기) - │ │ - paper-production-orchestrator (Skill) JIRA 상태 전환(Ready for AI) - │ ↓ 멤버 호출 │ ↓ Automation 웹훅 - 기획 → 분석 → 집필·그림 → 검수 → 발표 이벤트 허브(n8n 또는 OpenClaw+큐) - │ ↓ 산출물 계약(파일로 인계) │ ↓ Next Agent 분기 - results/ · manuscript/ · figures/ AI 워커(claude/codex/gemini) - │ ↓ 검증 게이트(숫자 재계산) │ ↓ 공통 MCP(JIRA·GitHub) - 사람 승인 → 공개 Handoff 코멘트 → 다음 AI (체인) -``` - -두 레이어의 접점은 **산출물 계약**과 **Handoff 규율**이다. 레이어 A의 멤버가 결과를 파일로 남기는 규율(results/FINDINGS.md 등)과, 레이어 B의 AI가 JIRA에 Handoff 코멘트를 남기는 규율은 같은 발상이다. 다음에 일할 주체가 그 산출물 하나만 읽어도 곧바로 착수할 수 있게 만든다. diff --git a/ai_scientist/output_v01/README.md b/ai_scientist/output_v01/README.md deleted file mode 100644 index 4292a45..0000000 --- a/ai_scientist/output_v01/README.md +++ /dev/null @@ -1,25 +0,0 @@ -# output_v01 — AI Scientist 설계 시각화 (HTML + mermaid) - -`ai_scientist/`의 마크다운 6편(README, 01–05)을 하나의 인터랙티브 HTML 문서로 묶은 결과물이다. - -## 여는 법 - -`index.html`을 브라우저로 열면 된다. - -```bash -# 예: 로컬에서 바로 열기 -xdg-open ai_scientist/output_v01/index.html # Linux -open ai_scientist/output_v01/index.html # macOS -``` - -## 구성 - -- **단일 페이지**: 좌측 사이드바 목차 + 본문. 개요 → 레이어 A → 레이어 B → 설계 원칙 → 컴포넌트 맵. -- **mermaid 다이어그램 6종**: 전체 그림, 랩 조직도, 논문 생산 루프, 파이프라인 P0–P5, 4계층 아키텍처, 인계 루프. -- **라이트/다크 테마 토글**(좌측 하단 버튼). 시스템 설정도 자동 반영. - -## 알아둘 점 - -- mermaid 라이브러리를 CDN(jsdelivr)에서 불러온다. 따라서 **다이어그램 렌더에는 인터넷 연결이 필요**하다. 표·본문은 오프라인에서도 보인다. -- 완전 오프라인(자체 완결형)이 필요하면 mermaid를 파일에 인라인하는 버전으로 다시 만들 수 있다. -- 다이어그램 6종은 mermaid 파서로 문법 검증을 마쳤다. diff --git a/ai_scientist/output_v01/index.html b/ai_scientist/output_v01/index.html deleted file mode 100644 index 798e27b..0000000 --- a/ai_scientist/output_v01/index.html +++ /dev/null @@ -1,507 +0,0 @@ - - - - - -AI Scientist 설계 — BioProject01 / kkkim-pipeline - - - - -
- - -
-
- - -
-
설계 정리 · Design Overview
-

AI Scientist — 연구 과정 전반의 자동화와 멀티 연구자 협업 구조

-

인공지능이 연구 도구를 넘어, 논문 탐색·정리부터 가설 설정, 실험 수행, 논문 작성까지 연구 과정 전반을 자동화하고, 여러 연구자가 함께 쓰도록 만든 구조를 두 개의 레이어로 정리한다.

-

이 문서는 ai_scientist/ 안의 마크다운 6편을 mermaid 다이어그램과 함께 시각화한 것이다. 새 설계가 아니라, kkkim-pipeline 브랜치에 이미 구현·기록된 구조를 하나의 지도로 묶었다.

-
- - -
-

개요 — 두 개의 레이어

-

목표는 두 가지였고, 각각 하나의 레이어로 설계했다. 레이어 A는 "AI 한 명이 논문 한 편을 어떻게 끝까지 끌고 가나"를, 레이어 B는 "그런 AI 여럿이 팀으로 어떻게 이어달리나"를 설계한다. A가 랩 안의 분업이라면 B는 랩과 랩, 사람과 사람 사이의 배턴 터치다.

- -
- - - -
레이어무엇인가이 저장소의 구현·근거
A 단일 랩 자동화한 연구자의 연구 과정 전체를 agent 멤버들이 나눠 맡아 자동으로 돌리는 "AI 연구 랩".claude/agents/ + paper-production-orchestrator Skill + 파이프라인 scripts/. 지도 = docs/HARNESS.md
B 멀티 AI 협업 인계여러 연구자·여러 AI가 JIRA 상태 신호로 작업을 자동 인계하는 체계guide/ai-handoff-architecture-guide.md + guide/openclaw-claude-guide.md
- -
-

그림 1. 전체 그림 — 사람(PI) 아래 두 레이어가 산출물 계약과 Handoff 규율로 맞물린다

-
-flowchart TB
-  PI["사람 = PI
방향 설정 · 승인 · 공개 게이트"] - PI --> LA - PI --> LB - subgraph LA["레이어 A · 단일 랩 자동화"] - direction TB - A1["paper-production-orchestrator (Skill)"] - A2["기획 · 분석 · 집필/그림 · 검수 · 발표"] - A3["산출물 계약: results / manuscript / figures"] - A1 --> A2 --> A3 - end - subgraph LB["레이어 B · 멀티 AI 협업 인계"] - direction TB - B1["JIRA 상태 전환 (Ready for AI)"] - B2["이벤트 허브 (n8n 또는 OpenClaw+큐)"] - B3["AI 워커: claude / codex / gemini"] - B4["공통 MCP (JIRA · GitHub)"] - B1 --> B2 --> B3 --> B4 - end - A3 -->|"공유: 산출물 계약"| SH - B4 -->|"공유: Handoff 규율"| SH - SH["다음 주체가 산출물 하나만 읽어도 곧바로 착수"] - classDef pi fill:#0e7c86,stroke:#0e7c86,color:#fff; - classDef share fill:#b26a00,stroke:#b26a00,color:#fff; - class PI pi; - class SH share; -
-
-
- - -
-

01연구 과정을 어떤 단계로 나눴나

-

전통적인 연구 흐름을 AI가 맡을 수 있는 단계로 나누면 일곱 단계가 된다. 사람이 도구를 하나씩 손으로 돌리는 대신, AI 멤버들이 각 단계를 나눠 맡아 이어서 돌아가게 한다.

-
- - - - - - - - -
#단계담당 주체
1논문 탐색·정리 (정직한 포지셔닝)literature-scout, paper_analysis/ dual-lens 14편
2가설 설정·차별화 (가장 값싼 입증)novelty-strategist, research-methodologist
3실험 설계·감사 (누수·통계 위험 차단)research-methodologist
4실험 수행·분석 (eval·통계·cross-dataset)hspc-velocity-analyst + scripts/ P0–P5
5집필·그림manuscript-writer + figures/figNN_*.py
6검수·리뷰paper-critic, reviewer
7발표presenter
-

일곱 단계를 사람이 매번 순서대로 부르지 않도록, 자연어 요청을 멤버에 배정하는 라우팅표(CLAUDE.md)와 여러 단계를 엮어 실행하는 오케스트레이터 Skill을 두었다.

-
- - -
-

02레이어 A 단일 랩 자동화

-

agent는 직원이 아니라 랩의 멤버(연구원)이고, 사람과 메인 루프가 랩을 이끄는 PI다. PI는 무엇을 할지 정하고 승인·공개를 책임지되, 실제 작업은 멤버가 파일로 주고받으며 이어서 한다.

- -

멤버 명부

-

하나(hspc-velocity-analyst)만 이 프로젝트 도메인 전용이고, 나머지는 다른 논문에도 재사용할 수 있게 만들었다.

-
- - - - - - - - - - - - -
멤버벤치역할
hspc-velocity-analyst분석실도메인 슬롯. 파이프라인(P0–P5)·eval·통계·cross-dataset 실행, 결과 파일 유지
literature-scout문헌·기획선행연구 탐색, 정직한 포지셔닝, related work
novelty-strategist문헌·기획차별화 각도와 가장 값싼 입증 실험 제안
research-methodologist문헌·기획가설·기여문·실험설계, 누수·통계 감사
manuscript-writer집필실프리프린트·저널·블로그 본문 초안과 그림 연계
presenter집필실청중 맞춤 슬라이드·발제
paper-critic심사·QA제출 전 적대적 자체검토와 그림 시각 QA
reviewer심사·QA정식 venue 스타일 공식 리뷰 (선택)
paper-orchestrator코디네이션멀티 agent 작업의 계획만 수립 (실행은 PI)
design엔지니어링로고·아이콘·브랜드·그림 미감
그림 생성 스크립트엔지니어링figures/figNN_*.py — 결과 파일에서 그림 생성·번호 정합
-
그림 생성을 agent가 아니라 결정론적 스크립트로 둔 것이 설계상의 선택이다. 숫자를 손으로 하드코딩하지 않고 결과 파일에서만 뽑게 해 재현성을 지킨다.
- -
-

그림 2. 랩 조직도 — PI 아래 다섯 벤치에 멤버가 배치된다

-
-flowchart TB
-  PI["PI = 사람 + 메인 루프"]
-  ORC["paper-orchestrator
(계획만)"] - PI --> ORC - ORC --> G1 & G2 & G3 & G4 & G5 - subgraph G1["문헌·기획"] - m1["literature-scout"]; m2["novelty-strategist"]; m3["research-methodologist"] - end - subgraph G2["분석실"] - m4["hspc-velocity-analyst"] - end - subgraph G3["집필실"] - m5["manuscript-writer"]; m6["presenter"] - end - subgraph G4["심사·QA"] - m7["paper-critic"]; m8["reviewer (선택)"] - end - subgraph G5["엔지니어링"] - m9["design"]; m10["figNN_*.py 스크립트"] - end - classDef pi fill:#0e7c86,stroke:#0e7c86,color:#fff; - class PI pi; -
-
-
- -
-

자연어 라우팅과 오케스트레이터

-

요청에 agent 이름이 없어도 CLAUDE.md의 라우팅표가 자연어 요청을 멤버에 배정한다. 여러 단계를 엮는 요청("분석→집필→그림→검수까지", "critic 지적 반영해")은 단일 멤버가 아니라 paper-production-orchestrator Skill로 보낸다. 메인 루프(PI)가 이 Skill을 실행하며 멤버를 순서대로 부른다. subagent는 subagent를 못 부르므로, "계획만 짜는" paper-orchestrator agent와 달리 실제 실행은 이 Skill이 맡는다.

- -
-

그림 3. 논문 생산 루프 — 검증 게이트를 통과해야 발표·공개로 넘어간다

-
-flowchart LR
-  P["기획·근거
methodologist · scout · strategist"] --> AN["분석·eval
hspc-velocity-analyst"] - AN --> WR["집필+그림
manuscript-writer"] - WR --> CR["검수
paper-critic"] - CR -->|"블로킹 지적"| WR - CR --> VG{"검증 게이트
숫자 재계산"} - VG -->|"불일치"| STOP["멈춤 · 사람 보고"] - VG -->|"통과"| PR["발표
presenter"] - classDef gate fill:#b26a00,stroke:#b26a00,color:#fff; - classDef stop fill:#b3261e,stroke:#b3261e,color:#fff; - class VG gate; class STOP stop; -
-
-

핵심은 부분 재실행이다. 이미 만들어진 산출물이 있으면 요청한 단계만 다시 돌리고 나머지는 기존 파일을 재사용한다. "그림만 다시"면 집필+그림 단계만, "최신 결과로 본문 갱신"이면 변경 지점의 하류 단계만 돌린다.

-
- -
-

산출물 계약 — 대화가 아니라 파일로 넘긴다

-

멤버는 중간 결과를 대화에만 남기지 않고 정해진 파일로 넘긴다. 다음 멤버는 그 파일을 읽고 이어서 일한다. 이 계약 덕분에 멤버가 교체되거나 세션이 끊겨도 작업이 이어진다.

-
- - - - - - -
단계Writer산출물다음이 읽음
분석·evalhspc-velocity-analystresults/FINDINGS.md + results/*.csv + results/*.md집필·검수
집필·그림manuscript-writermanuscript/draft_v2.md + draft_v2_ko.md (영/한 동시), figures/*.png검수·리뷰·발표
검수·리뷰paper-critic / reviewermanuscript/REVIEW-<venue>-<date>.md집필(수정)
발표presenter슬라이드·발제사람
상태 핸드오프전원HANDOFF.md, TODO.md, SESSION-LOG.md다음 세션
-
- -
-

실험 실행 엔진 — 파이프라인 P0–P5

-

분석 단계의 실제 계산은 pipeline/hspc-velocity-benchmark/scripts/가 담당한다. hspc-velocity-analyst가 이 스크립트들을 돌려 결과 파일을 만든다.

-
-

그림 4. 파이프라인 단계 — 공통 전처리(P1) 위에서 method를 분기해 재현성을 검증한다

-
-flowchart LR
-  P0["P0
다운로드·provenance"] --> P1["P1
통일 전처리"] - P1 --> P2["P2
velocity method 실행"] - P2 --> P3["P3
재현성 검증"] - P3 --> P4["P4
permutation FDR"] - P4 --> P5["P5
bootstrap 안정성"] -
-
-
    -
  • P0download_data.sh로 GSE209878를 받고 download_manifest.tsv(sha256)와 P0_provenance.md를 남긴다.
  • -
  • P1p1_build.py가 공통 branch를 만든다. 여기서 preprocessing 차이와 method 차이를 분리한다.
  • -
  • P2p2_multivelo.py, p2_moflow.py, p2_crakvelo_*, p2_multivelovae.py 등으로 여러 method를 같은 전처리 위에서 돌린다.
  • -
  • P3p3_concordance.py, p3_crossdataset_concordance.py, p3_scrambled_null.py로 method 간·dataset 간 일치도와 null을 계산한다.
  • -
  • P4 — gene 단위 다중검정을 permutation FDR로 통제한다.
  • -
  • P5 — shuffle/seed 변이 audit(p10*)까지 포함해 결과의 흔들림을 잰다.
  • -
-
- -
-

게이트 — 자동화가 넘지 못하는 선

-

이 랩은 전부를 자동으로 밀지 않는다. 두 종류의 게이트가 있다.

-

검증 게이트 (커밋·공개 전). 헤드라인 숫자를 결정론적으로 재계산해 결과 파일과 대조한다.

-
cd pipeline/hspc-velocity-benchmark/scripts
-conda run --no-capture-output -n scv-preprocess python p3_concordance.py
-conda run --no-capture-output -n scv-preprocess python p3_crossdataset_concordance.py --dataset human_brain
-conda run --no-capture-output -n scv-preprocess python p3_scrambled_null.py
-# 출력 숫자를 results/FINDINGS.md 와 대조. 불일치면 멈추고 사람에게 보고.
-

사람 승인 게이트. 프리프린트·블로그 외부 공개와 main 병합은 사람이 승인한다. 저자·소속·IP·corresponding email이 확정되기 전에는 공개를 보류한다(<FILL>). claim 자체도 반증기준·make-or-break 검정·advisor 확인을 통과하기 전에는 PROVISIONAL로 두고 본문에 넣지 않는다.

-
- - -
-

03레이어 B 멀티 AI 협업 인계

-

팀원마다 다른 AI(Claude, Codex, Gemini)를 쓰고 결과물은 JIRA·Confluence·Git으로 공유한다. 문제는 한 작업이 끝나도 다음 담당자의 AI에 신호가 자동으로 가지 않아 인계가 지연된다는 점이다. 이 레이어는 그 인계를 자동화한다.

-
- - - - - -
원칙내용
단일 신호원인계 신호는 JIRA 상태 전환만 쓴다. Git 머지 등은 JIRA 상태로 수렴시킨다
사람 승인 우선초기엔 Slack 원클릭 승인 후 실행. 신뢰가 쌓이면 단계적으로 자동화
최소 권한AI별 서비스 계정 분리, 프로젝트 단위 권한, main 직접 push 금지
폭주 방지티켓당 자동 인계 상한(기본 5회), 실패 시 즉시 사람 에스컬레이션
- -
-

그림 5. 4계층 아키텍처 — 작업 완료가 다시 JIRA 상태 전환을 일으켜 체인이 반복된다

-
-flowchart TB
-  subgraph L1["① 이벤트 소스 (기존 스택)"]
-    S1["JIRA 상태 전환: Ready for AI"]
-    S2["Git PR 머지 → JIRA 상태 자동 전환"]
-  end
-  subgraph L2["② 이벤트 허브 (신규)"]
-    H1["Webhook 수신 · Next Agent 분기 · (선택) Slack 승인 · 워커 호출"]
-  end
-  subgraph L3["③ AI 워커 (신규)"]
-    W1["claude -p"]; W2["codex exec"]; W3["gemini -p"]
-  end
-  subgraph L4["④ MCP 공통"]
-    M1["Atlassian MCP: JIRA · Confluence"]
-    M2["GitHub MCP: 저장소 · PR"]
-  end
-  L1 -->|"Webhook (HTTP POST)"| L2
-  L2 -->|"Execute / SSH / HTTP"| L3
-  L3 -->|"공통 mcp.json"| L4
-  L4 -.->|"상태 전환 → 신호 재발생"| L1
-        
-
-
- -
-

인계 루프와 Handoff 코멘트

-
-

그림 6. 티켓 생애주기 — 후속 작업이 있으면 1번으로 돌아가 체인이 이어진다

-
-flowchart TB
-  T1["1. 작업 완료 + Handoff 코멘트"] --> T2["2. JIRA 상태 Ready for AI · Next Agent 지정"]
-  T2 --> T3["3. Automation 웹훅 발송"]
-  T3 --> T4["4. 허브가 Next Agent로 분기 (초기 Slack 승인)"]
-  T4 --> T5["5. AI 워커 실행: MCP로 맥락 로드 후 작업"]
-  T5 --> T6{"후속 작업?"}
-  T6 -->|"있음 → Ready for AI"| T1
-  T6 -->|"사람 검토 → In Review"| HU["사람"]
-        
-
-

모든 AI의 규칙 파일(CLAUDE.md / AGENTS.md / GEMINI.md)에 같은 Handoff 템플릿을 강제한다. 기준은 하나다. 다음 워커가 이 코멘트 하나만 읽어도 착수할 수 있어야 한다. 이것이 레이어 A의 산출물 계약과 같은 발상이다. A는 파일로, B는 JIRA 코멘트로 맥락을 넘긴다.

-
## Handoff
-- 완료한 것: (요약 3줄 이내)
-- 산출물: (커밋 해시 / PR 링크 / Confluence 페이지 링크)
-- 다음 작업: (다음 AI가 해야 할 일, 구체적으로)
-- 제약/주의: (건드리면 안 되는 것, 실패했던 접근)
-- Next Agent: claude | codex | gemini | human
-
- -
-

OpenClaw로 실현하기 — 허브와 워커를 대체

-

인계 가이드는 이벤트 허브로 n8n을, 워커로 공용 서버의 run_agent.sh를 상정한다. OpenClaw 가이드는 그 ②+③(허브+워커)을 OpenClaw와 메시지 큐로 대체하는 경로를 제시한다. 별도 서버를 세우지 않고 같은 인계 루프를 돌린다.

-
- - - - - -
인계 가이드 계층원 구성OpenClaw로 실현
① 이벤트 소스JIRA 상태 전환 / PR 머지그대로 유지
② 이벤트 허브n8n메시지 큐 브리지 + OpenClaw Webhooks 플러그인
③ AI 워커공용 서버 + run_agent.sh + claude -pOpenClaw 세션 (인증·모델선택·thinking 레벨 관장)
④ MCP 공통mcp.json동일. OpenClaw 세션에도 같은 MCP 서버를 물린다
-

메시지 큐를 앞에 두는 이유는 안정성과 비용이다. 브리지 컨슈머는 네 가지를 지킨다. (1) ack는 Claude 처리 성공 이후에만, (2) 동시 처리 수 제한, (3) 멱등성·세션 키, (4) 같은 티켓 연속 이벤트 병합.

- -
비용 — 인계 체인은 호출을 곱셈으로 늘린다. 한 티켓이 여러 AI를 연쇄 호출하므로 단발 실행보다 토큰 지출이 배로 뛴다. 그래서 비용 레버가 인계 자동화에서 더 중요해진다: 이벤트 병합, 모델 티어링(저위험은 Sonnet/Haiku, 핵심만 Opus), 티켓 단위 프롬프트 캐싱, 우선순위 큐, poison 티켓의 DLQ 격리.

-
- -
-

보안·승인 게이트와 도입 로드맵

-
    -
  • 서명 검증 2구간 — JIRA Automation의 X-Handoff-Token과 OpenClaw webhook secret을 둘 다 건다.
  • -
  • 최소 권한 — AI별 서비스 계정 분리, JIRA는 해당 프로젝트만, main 직접 push 금지(브랜치+PR).
  • -
  • 토큰 비노출 — API 키·PAT는 환경변수·시크릿 매니저로만. .mcp.json에 토큰 직접 기입 금지.
  • -
-
- - - - -
주차목표산출물
1주차JIRA 필드·워크플로·Automation + 허브 설치, Slack 알림까지만인계 발생 즉시 알림 (자동 실행 없음)
2~3주차AI CLI·MCP 공통 설정·워커 구축, Slack 승인 후 반자동첫 AI-to-AI 인계 파일럿 1건
4주차~저위험 작업부터 승인 생략, 인계 상한·모니터링 정착제한적 완전 자동 체인 + 비용 레버 계측
-
- - -
-

04설계 원칙 — 두 레이어를 관통하는 것

-

레이어 A와 B는 다른 문제를 풀지만 같은 원리 위에 서 있다. 이 원리들이 AI Scientist 설계의 뼈대다.

-
-
1하나만 읽어도 착수

작업 맥락을 대화가 아니라 정형화된 산출물로 넘긴다. A는 결과 파일, B는 JIRA Handoff 코멘트. 세션이 끊겨도 이어진다.

-
2사람 게이트를 남긴다

되돌리기 어렵거나 외부로 나가는 지점에는 사람이 선다. 승인 게이트는 성숙도에 따라 옮긴다. 처음엔 촘촘하게, 검증되면 넓게.

-
3검증 게이트

커밋·공개 전에 숫자를 다시 계산해 대조한다. weak는 zero가 아니다. 통계가 뒷받침하지 않는 우월·재현 주장을 금지한다.

-
4폭주·비용을 구조로

사람의 주의가 아니라 구조로 막는다. Hop Count 상한, DLQ, 동시성 제한, 모델 티어링, 캐싱, 우선순위 큐.

-
5최소 권한과 격리

AI별 서비스 계정 분리, 프로젝트 단위 권한, main 직접 push 금지. 자동 승인 옵션은 격리 환경 전제. 토큰은 시크릿으로만.

-
6표준 포맷과 재사용

멤버·라우터를 특정 프로젝트에 묶지 않는다. 재사용 스캐폴드(CC BY 4.0), OpenClaw/Codex 네이티브 포맷, MCP 표준.

-
7근거와 코드 분리

method 선택의 근거(paper_analysis/)와 그 근거로 돌리는 코드(pipeline/)를 두 폴더로 나눈다. 판단이 바뀌면 근거만, 실행이 바뀌면 코드만 고친다.

-
-
- - -
-

05컴포넌트 맵 — 설계 요소가 저장소 어디에 있나

-

이 폴더는 설계를 설명하고, 아래 파일들이 그 설계를 구현한다.

- -

레이어 A 단일 랩 자동화

-
- - - - - - - - - - -
설계 요소저장소 위치
랩 구조 지도docs/HARNESS.md
라우팅표 + 산출물 계약CLAUDE.md (Agent routing & artifact contract)
멤버 정의.claude/agents/*.md
오케스트레이터 (실행 입구).claude/skills/paper-production-orchestrator/SKILL.md
단일 컨텍스트 (thesis·claim 등급표)pipeline/hspc-velocity-benchmark/manuscript/PAPER_DIRECTION.md
분석 실행 엔진 (P0–P5)pipeline/hspc-velocity-benchmark/scripts/
method 선택 근거DESIGN.md, paper_analysis/ (dual-lens 14편)
검증 게이트 스크립트scripts/p3_concordance.py, p3_crossdataset_concordance.py, p3_scrambled_null.py
글쓰기 규율 (한국어 윤문).claude/rules/writing-style.md
- -

레이어 B 멀티 AI 협업 인계

-
- - - - - - -
설계 요소저장소 위치
인계 아키텍처 (4계층·설치 가이드)guide/ai-handoff-architecture-guide.md
OpenClaw 실현 (허브+워커·큐·비용)guide/openclaw-claude-guide.md
분석 하네스 project frameAGENTS.md (dataset 라우팅을 skills/ROUTES.md에 위임)
MCP 공통 설정.mcp.json (설계 목표는 agent-config 저장소)
팀·역할·AI 계정 매핑Project-Info.md
-
정확성 주의: AGENTS.md가 위임하는 skills/ROUTES.md·openai.yaml 스킬 트리는 이 브랜치 체크아웃에는 없다. 포맷만 규정되어 있고, 실제 스킬 트리는 OpenClaw로 돌릴 때 채운다.
- -

두 레이어의 접점

-
- - - - - - -
공유 요소레이어 A에서레이어 B에서
인계 계약결과 파일 (results/FINDINGS.md)JIRA Handoff 코멘트
사람 게이트공개·main 병합 승인초기 Slack 승인
폭주·비용 방지검증 게이트, claim 등급Hop Count 상한, 큐·DLQ, 모델 티어링
실행 도구Claude Code (agent·Skill)OpenClaw 세션 또는 run_agent.sh
라우터 포맷CLAUDE.md 라우팅표Next Agent 필드 → 브리지 분기
-
- - - -
-
-
- - - -