From 6ef29cc04eade04089e9ce51d4e5a1224edd4110 Mon Sep 17 00:00:00 2001 From: JamieLyu <5819016+JamieLyu@users.noreply.github.com> Date: Thu, 6 Aug 2026 13:50:49 +0900 Subject: [PATCH 1/3] =?UTF-8?q?P5=20=EC=B4=88=EB=A1=9D=20=EA=B0=90?= =?UTF-8?q?=EC=B6=95=20466=E2=86=92362=EB=8B=A8=EC=96=B4(EN),=20=EB=8C=80?= =?UTF-8?q?=EC=9D=91=20=ED=95=9C=EA=B8=80=20=EC=B4=88=EB=A1=9D=20=EB=8F=99?= =?UTF-8?q?=EA=B8=B0=ED=99=94=20(BIOP01-52)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit kkkim 요청(comment 11626/11659, 421→350 목표)에 따라 draft_v2.md/ko.md 초록을 트리밍. 모든 수치·CI·부호·hedge는 원문 그대로 보존(수치 토큰 diff로 확인) — 문장 압축·중복 제거만 수행, 주장 강도·범위는 불변. 결론부 baseline-to-α 문구는 kkkim의 최근 수정(9b53c57)을 그대로 반영. --- pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md | 8 ++++---- .../hspc-velocity-benchmark/manuscript/draft_v2_ko.md | 8 ++++---- 2 files changed, 8 insertions(+), 8 deletions(-) diff --git a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md index 415652b..18ab4ee 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md @@ -36,13 +36,13 @@ the external-measurement result is corroboration inside the map, not a title cla ## Abstract -**Background.** Chromatin-informed ("multiome") RNA-velocity methods emit several per-gene quantities — a transcription rate, a degradation rate, and a chromatin-to-transcription *lag* (the offset between when a locus opens or closes and when its transcription switches) — each proposed as a biological readout, the lag in particular for predicting the timing of epigenetic-drug responses. A derived quantity is usable downstream only if it is reliable: reproducible across reasonable algorithms and, where possible, consistent with an independent measurement. We asked, for human hematopoietic stem and progenitor cells (HSPCs) profiled by 10x Multiome, which velocity outputs meet that bar. +**Background.** Chromatin-informed ("multiome") RNA-velocity methods emit several per-gene quantities, a transcription rate, a degradation rate, and a chromatin-to-transcription *lag* (the offset between chromatin opening/closing and the transcriptional switch), each proposed as a biological readout, the lag for predicting epigenetic-drug timing. A derived quantity is usable downstream only if reliable: reproducible across algorithms and, where possible, consistent with independent measurement. We asked which outputs meet that bar in human hematopoietic stem and progenitor cells (HSPCs) profiled by 10x Multiome. -**Approach.** Across up to five velocity arms (an RNA-only scVelo floor plus MultiVelo, MultiVeloVAE, MoFlow and CRAK-Velo) we tested each output on four axes: cross-method reproducibility (permutation-FDR), a causal within-lineage ATAC-shuffle control, cross-dataset replication in five external multiomes (one preregistered), and external anchoring of the fitted rates to measured synthesis (K562 TT-seq) and degradation (mRNA half-life) rates. We audit the per-gene kinetic parameters and, one layer up, the cell×gene velocity matrix — not the low-dimensional embedding arrows or trajectories that velocity is principally used for. +**Approach.** Across up to five velocity arms (RNA-only scVelo plus MultiVelo, MultiVeloVAE, MoFlow, and CRAK-Velo), we tested each output on four axes: cross-method reproducibility, a causal within-lineage ATAC-shuffle control, cross-dataset replication in five external multiomes (one preregistered), and external anchoring to K562 TT-seq synthesis and mRNA half-life degradation rates. We audit per-gene kinetic parameters and the cell×gene velocity matrix, not the embedding arrows or trajectories velocity is principally used for. -**Findings.** The outputs separated sharply. Only the transcription rate α reproduced across methods (Spearman ρ=0.88, observed); the chromatin-to-transcription lag and the degradation rate γ did not. The chromatin-to-transcription lag reproduced weakly at best in magnitude (strongest pair ρ=+0.163, most pairs |ρ|≤0.08) and only at chance in sign (54.6%), and shuffling ATAC left it statistically unchanged, identifying the lag as model-structural rather than chromatin-driven. The degradation rate γ was likewise method-fragile (cross-method ρ≈−0.1) and not recovered even where a ground truth existed (K562 half-life 3/3 null; the textbook scVelo γ ran reversed, −0.224, CI excluding 0). As external corroboration, fitted α also tracked the measured K562 TT-seq synthesis rate in all three methods (non-housekeeping ρ +0.24 to +0.29, all CIs excluding 0); because steady-state transcript abundance tracked that same measurement at least as strongly (Spearman(abundance, synthesis)=+0.410 versus Spearman(α, synthesis)=+0.262), we read this as consistency evidence, not as α being a uniquely measurement-grounded output. The α-over-lag ordering held in all six systems and passed a preregistered six-of-six scorecard sealed before any fit. Profile-likelihood analysis gave the mechanism — α is stiff (identifiable), the lag sloppy and boundary-limited — with identifiability aligning partially, but not confirmably, with external validation. +**Findings.** Only α reproduced across methods (Spearman ρ=0.88, observed); the lag and degradation rate γ did not. The lag reproduced weakly in magnitude (strongest pair ρ=+0.163, most pairs |ρ|≤0.08), only at chance in sign (54.6%), and was statistically unchanged by ATAC-shuffling, marking it model-structural rather than chromatin-driven. γ was likewise method-fragile (ρ≈−0.1) and unrecovered against ground truth (K562 half-life 3/3 null; textbook scVelo γ ran reversed, −0.224, CI excluding 0). Fitted α tracked K562 TT-seq synthesis in all three methods (ρ +0.24 to +0.29, CIs excluding 0), but steady-state abundance tracked it at least as strongly (+0.410 vs. +0.262), so we treat this as consistency evidence, not unique grounding. The α-over-lag ordering held in all six systems, passing a preregistered six-of-six scorecard sealed before any fit. Profile-likelihood analysis explained why: α is stiff (identifiable), the lag sloppy and boundary-limited, identifiability aligning only partially with external validation. -**Conclusions.** We distil these results into a velocity-output reliability map: trust α and rate-derived signals, reproducible across methods, though α is largely expression (steady-state abundance tracks the external synthesis measurement at least as strongly, so α does not add demonstrable synthesis information beyond it); treat the lag, its sign, absolute timing and γ as unreliable, requiring orthogonal validation. Any downstream timing-prediction model should route through the robust baseline-to-α path rather than a single-method lag. +**Conclusions.** We distil these results into a velocity-output reliability map: trust α and rate-derived signals, reproducible across methods, though α is largely expression (abundance tracks the same external measurement at least as strongly, so α adds no demonstrable information beyond it); treat the lag, its sign, absolute timing, and γ as unreliable, requiring orthogonal validation. Downstream timing-prediction should route through the robust baseline-to-α path, not a single-method lag. **Keywords:** RNA velocity, single-cell multiome, chromatin accessibility, transcriptional kinetics, parameter identifiability, external validation, benchmarking, hematopoiesis diff --git a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md index 0e0693f..106a11f 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md @@ -19,13 +19,13 @@ draft_v2.md의 한국어 검토본(번역·윤문). 정본은 영어 draft_v2.md ## Abstract -**배경.** Chromatin 정보를 결합한("multiome") RNA velocity 방법들은 유전자별로 여러 값을 산출한다. 전사 속도(transcription rate), 분해 속도(degradation rate), 그리고 chromatin에서 transcription으로 이어지는 *시간차(lag)*, 곧 locus가 열리거나 닫히는 시점과 그 유전자의 transcription이 전환되는 시점 사이의 오프셋이 그것이다. 이 값들은 저마다 생물학적 판독값으로 제안되어 왔고, 특히 시간차(lag)는 epigenetic 약물 반응의 timing을 예측하는 데 쓸 수 있다고 여겨져 왔다. 그러나 파생된 값을 하류에서 쓰려면 그것이 먼저 *신뢰할 수 있어야* 한다. 곧 합리적인 알고리즘들에 걸쳐 재현되어야 하고, 가능하다면 독립적인 측정값과도 부합해야 한다. 우리는 인간 조혈모·전구세포(hematopoietic stem and progenitor cells, HSPC; 10x Multiome으로 프로파일링)에서 어떤 velocity 출력이 이 기준을 충족하는지를 물었다. +**배경.** Chromatin 정보를 결합한("multiome") RNA velocity 방법들은 유전자별로 전사 속도(transcription rate), 분해 속도(degradation rate), chromatin에서 transcription으로 이어지는 시간차(lag, locus가 열리거나 닫히는 시점과 transcription 전환 시점의 오프셋)를 산출한다. 이 값들은 저마다 생물학적 판독값으로 제안되어 왔고, 시간차(lag)는 특히 epigenetic 약물 반응의 timing 예측에 쓰일 수 있다고 여겨져 왔다. 그러나 파생된 값은 신뢰할 수 있어야만, 곧 합리적인 알고리즘들에 걸쳐 재현되고 가능하면 독립적인 측정값과도 부합해야만 하류에서 쓸 수 있다. 우리는 10x Multiome으로 프로파일링한 인간 조혈모·전구세포(hematopoietic stem and progenitor cells, HSPC)에서 어떤 velocity 출력이 이 기준을 충족하는지를 물었다. -**접근.** 최대 다섯 개의 velocity arm(갈래)(RNA 전용 scVelo floor에 MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo를 더한 것)에 걸쳐 각 출력을 네 축에서 검정했다. 방법 간 재현성(순열 FDR), lineage 내 ATAC-shuffle 인과 대조군, 다섯 개 외부 multiome(그중 하나는 사전등록)에서의 cross-dataset 재현, 그리고 fitting된 속도가 측정된 합성(K562 TT-seq)·분해(mRNA 반감기) 속도에 외부 anchoring되는지가 그것이다. 우리는 유전자별 kinetic parameter와, 한 단계 위인 세포×유전자 velocity 행렬을 감사 대상으로 삼았다. velocity가 주로 쓰이는 저차원 임베딩(embedding) arrow나 궤적(trajectory)은 감사 대상에 포함하지 않았다. +**접근.** 최대 다섯 개 velocity arm(RNA 전용 scVelo에 MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo를 더한 것)에 걸쳐 각 출력을 네 축에서 검정했다. 방법 간 재현성, lineage 내 ATAC-shuffle 인과 대조군, 다섯 개 외부 multiome에서의 cross-dataset 재현(그중 하나는 사전등록), K562 TT-seq 합성 속도와 mRNA 반감기 분해 속도로의 외부 anchoring이 그것이다. 우리는 유전자별 kinetic parameter와 세포×유전자 velocity 행렬을 감사 대상으로 삼았고, velocity가 주로 쓰이는 저차원 임베딩 arrow나 궤적은 다루지 않았다. -**결과.** 출력들은 뚜렷이 갈라졌다. 전사 속도 α만 방법 간에 재현되었고(Spearman ρ=0.88, 관측값), chromatin에서 transcription으로 이어지는 시간차(lag)와 분해 속도 γ는 그렇지 않았다. 시간차(lag)는 크기에서 잘해야 약하게 재현되었고(가장 강한 쌍 ρ=+0.163, 대부분의 쌍 |ρ|≤0.08) 부호에서는 우연 수준에 그쳤으며(54.6%), ATAC를 뒤섞어도 통계적으로 변하지 않아 시간차(lag)가 chromatin에서 비롯된 것이 아니라 모델 구조에서 비롯됨을 드러냈다. 분해 속도 γ도 마찬가지로 방법에 취약했고(방법 간 ρ≈−0.1), ground truth가 있는 곳에서도 복원되지 않았다(K562 반감기 3/3 귀무; 교과서적 scVelo γ는 반대로 나와 −0.224, CI가 0 배제). 외부 보강 증거로서, fitting된 α는 세 방법 모두에서 측정된 K562 TT-seq 합성 속도와도 상관했다(비-housekeeping ρ +0.24 ~ +0.29, 모든 CI가 0 배제). 다만 steady-state transcript abundance가 같은 측정값을 최소한 그만큼 잘 예측하므로(Spearman(abundance, 합성)=+0.410 대 Spearman(α, 합성)=+0.262), 이를 α가 유일하게 측정에 기반한 출력이라는 근거가 아니라 일관성 증거로 읽는다. α가 시간차(lag)보다 앞선다는 순서는 여섯 개 시스템 모두에서 유지되었고, 어떤 fitting보다 먼저 봉인한 사전등록 6-of-6 채점표(scorecard)를 통과했다. 프로파일 우도(profile-likelihood) 분석은 그 기제를 주었다. α는 stiff(식별 가능)한 반면 시간차(lag)는 sloppy하고 경계에 제약되며, 이 식별가능성은 외부 검증과 부분적으로 정렬하나 확증적으로 정렬하지는 않는다. +**결과.** α만 방법 간에 재현되었고(Spearman ρ=0.88, 관측값), 시간차(lag)와 분해 속도 γ는 그렇지 않았다. 시간차(lag)는 크기에서 잘해야 약하게 재현되었고(가장 강한 쌍 ρ=+0.163, 대부분 |ρ|≤0.08) 부호에서는 우연 수준(54.6%)에 그쳤으며, ATAC를 뒤섞어도 통계적으로 변하지 않아 chromatin이 아니라 모델 구조에서 비롯됨을 드러냈다. γ도 마찬가지로 방법에 취약했고(방법 간 ρ≈−0.1) ground truth 앞에서도 복원되지 않았다(K562 반감기 3/3 귀무; 교과서적 scVelo γ는 반대로 나와 −0.224, CI가 0 배제). fitting된 α는 세 방법 모두에서 K562 TT-seq 합성 속도와 상관했으나(ρ +0.24~+0.29, 모든 CI가 0 배제), steady-state abundance가 같은 측정값을 최소한 그만큼 잘 예측해(+0.410 대 +0.262) 이를 α만의 고유한 측정 근거가 아니라 일관성 증거로 다룬다. α가 시간차(lag)보다 앞선다는 순서는 여섯 시스템 모두에서 유지되어, 어떤 fitting보다 먼저 봉인한 사전등록 6-of-6 채점표를 통과했다. profile-likelihood 분석은 그 이유를 보여준다. α는 stiff(식별 가능)한 반면 시간차(lag)는 sloppy하고 경계에 제약되며, 이 식별가능성은 외부 검증과 부분적으로만 정렬한다. -**결론.** 우리는 이 결과를 velocity 출력 신뢰도 지도로 정리한다. α와 속도에서 파생된 신호는 신뢰할 수 있다. 다만 α는 방법 간에 재현되면서도 대체로 발현량을 반영하는 값이어서, steady-state abundance가 같은 외부 측정값을 α만큼, 또는 그 이상으로 예측하므로 α가 abundance를 넘어서는 합성률 정보를 더한다고는 입증되지 않는다. 시간차(lag), 그 부호, 절대 timing, 그리고 γ는 신뢰할 수 없는 것으로 다루며 직교(orthogonal) 검증을 요구한다. 하류의 어떤 timing 예측 모델이든 단일 방법의 시간차(lag)가 아니라 강건한 baseline-to-α 경로를 거쳐야 한다. +**결론.** 우리는 이 결과를 velocity 출력 신뢰도 지도로 정리한다. α와 속도 파생 신호는 방법 간에 재현되어 신뢰할 수 있으나, α는 대체로 발현량을 반영하는 값이라(abundance가 같은 외부 측정값을 α만큼 잘 예측해 α가 그 이상의 정보를 더한다고는 입증되지 않는다) 시간차(lag), 그 부호, 절대 timing, γ는 신뢰할 수 없는 것으로 다루며 직교(orthogonal) 검증을 요구한다. 하류의 timing 예측 모델은 단일 방법의 시간차(lag)가 아니라 강건한 baseline-to-α 경로를 거쳐야 한다. **키워드:** RNA velocity, single-cell multiome, chromatin accessibility, transcriptional kinetics, parameter identifiability, external validation, benchmarking, hematopoiesis From cf5d23c1384b287e831a0ac56c93ccdc01c24822 Mon Sep 17 00:00:00 2001 From: JamieLyu <5819016+JamieLyu@users.noreply.github.com> Date: Thu, 6 Aug 2026 14:02:13 +0900 Subject: [PATCH 2/3] =?UTF-8?q?P5=20=EC=B4=88=EB=A1=9D=20=EC=B6=94?= =?UTF-8?q?=EA=B0=80=20=EA=B0=90=EC=B6=95=20362=E2=86=92333=EB=8B=A8?= =?UTF-8?q?=EC=96=B4(EN),=20GB=5FBENCHMARK.md=20=EC=8B=A4=EC=A6=9D=20?= =?UTF-8?q?=EA=B8=B0=EC=A4=80(=E2=89=A4320)=20=EC=A0=95=ED=95=A9?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit kkkim이 언급한 350은 어림값이었고, manuscript/GB_BENCHMARK.md §3/§4(G1)이 실제 GB 벤치마크 논문 3편(E1/E2/E3, 200~300단어)을 근거로 산출한 목표는 ≤320단어. 이번 커밋으로 333까지 좁혔다(잔여 13단어는 문장 손상 없이는 더 줄이기 어려운 지점). 제거한 수치(0.163/0.224/0.262/0.410)는 전부 "detail, move to Results" (G1 명시 지침) 대상이며 Results 본문에 그대로 남아있음을 grep으로 확인. 주장 강도·hedge·헤드라인 수치(ρ=0.88, |ρ|≤0.08, 54.6%, 6-of-6, TT-seq anchor)는 전부 보존. --- pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md | 8 ++++---- .../hspc-velocity-benchmark/manuscript/draft_v2_ko.md | 8 ++++---- 2 files changed, 8 insertions(+), 8 deletions(-) diff --git a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md index 18ab4ee..feb0589 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md @@ -36,13 +36,13 @@ the external-measurement result is corroboration inside the map, not a title cla ## Abstract -**Background.** Chromatin-informed ("multiome") RNA-velocity methods emit several per-gene quantities, a transcription rate, a degradation rate, and a chromatin-to-transcription *lag* (the offset between chromatin opening/closing and the transcriptional switch), each proposed as a biological readout, the lag for predicting epigenetic-drug timing. A derived quantity is usable downstream only if reliable: reproducible across algorithms and, where possible, consistent with independent measurement. We asked which outputs meet that bar in human hematopoietic stem and progenitor cells (HSPCs) profiled by 10x Multiome. +**Background.** Chromatin-informed ("multiome") RNA-velocity methods emit several per-gene quantities, a transcription rate, a degradation rate, and a chromatin-to-transcription *lag* (the offset between chromatin opening/closing and the transcriptional switch), each proposed as a biological readout, the lag for predicting epigenetic-drug timing. A derived quantity is useful downstream only if reproducible across algorithms and, where possible, consistent with independent measurement. We asked which outputs meet that bar in human hematopoietic stem and progenitor cells (HSPCs) profiled by 10x Multiome. -**Approach.** Across up to five velocity arms (RNA-only scVelo plus MultiVelo, MultiVeloVAE, MoFlow, and CRAK-Velo), we tested each output on four axes: cross-method reproducibility, a causal within-lineage ATAC-shuffle control, cross-dataset replication in five external multiomes (one preregistered), and external anchoring to K562 TT-seq synthesis and mRNA half-life degradation rates. We audit per-gene kinetic parameters and the cell×gene velocity matrix, not the embedding arrows or trajectories velocity is principally used for. +**Approach.** Across up to five velocity arms (RNA-only scVelo plus MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo), we tested each output on four axes: cross-method reproducibility, a causal ATAC-shuffle control, cross-dataset replication in five external multiomes (one preregistered), and anchoring to K562 TT-seq synthesis and half-life degradation rates. We audit per-gene kinetic parameters and the cell×gene velocity matrix, not the embedding arrows or trajectories velocity is principally used for. -**Findings.** Only α reproduced across methods (Spearman ρ=0.88, observed); the lag and degradation rate γ did not. The lag reproduced weakly in magnitude (strongest pair ρ=+0.163, most pairs |ρ|≤0.08), only at chance in sign (54.6%), and was statistically unchanged by ATAC-shuffling, marking it model-structural rather than chromatin-driven. γ was likewise method-fragile (ρ≈−0.1) and unrecovered against ground truth (K562 half-life 3/3 null; textbook scVelo γ ran reversed, −0.224, CI excluding 0). Fitted α tracked K562 TT-seq synthesis in all three methods (ρ +0.24 to +0.29, CIs excluding 0), but steady-state abundance tracked it at least as strongly (+0.410 vs. +0.262), so we treat this as consistency evidence, not unique grounding. The α-over-lag ordering held in all six systems, passing a preregistered six-of-six scorecard sealed before any fit. Profile-likelihood analysis explained why: α is stiff (identifiable), the lag sloppy and boundary-limited, identifiability aligning only partially with external validation. +**Findings.** Only α reproduced across methods (Spearman ρ=0.88, observed); the lag and degradation rate γ did not. The lag reproduced weakly in magnitude (most pairs |ρ|≤0.08) and only at chance in sign (54.6%); ATAC-shuffling left it statistically unchanged, marking it model-structural rather than chromatin-driven. γ was likewise method-fragile (ρ≈−0.1) and unrecovered against ground truth (K562 half-life 3/3 null; the textbook scVelo γ even ran reversed). Fitted α tracked K562 TT-seq synthesis in all three methods (ρ +0.24 to +0.29, CIs excluding 0), though abundance tracked it at least as strongly, so we treat this as consistency, not unique grounding. The α-over-lag ordering held in all six systems, passing a preregistered six-of-six scorecard. Profile-likelihood analysis explained why: α is identifiable, the lag sloppy and boundary-limited. -**Conclusions.** We distil these results into a velocity-output reliability map: trust α and rate-derived signals, reproducible across methods, though α is largely expression (abundance tracks the same external measurement at least as strongly, so α adds no demonstrable information beyond it); treat the lag, its sign, absolute timing, and γ as unreliable, requiring orthogonal validation. Downstream timing-prediction should route through the robust baseline-to-α path, not a single-method lag. +**Conclusions.** We distil these results into a velocity-output reliability map: trust α and rate-derived signals, reproducible across methods, though α is largely expression (abundance tracks the same external measurement at least as strongly, so α adds no demonstrable information beyond it); treat the lag, sign, timing, and γ as unreliable, requiring orthogonal validation. Downstream timing-prediction should route through the robust baseline-to-α path, not a single-method lag. **Keywords:** RNA velocity, single-cell multiome, chromatin accessibility, transcriptional kinetics, parameter identifiability, external validation, benchmarking, hematopoiesis diff --git a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md index 106a11f..dad281f 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md @@ -19,13 +19,13 @@ draft_v2.md의 한국어 검토본(번역·윤문). 정본은 영어 draft_v2.md ## Abstract -**배경.** Chromatin 정보를 결합한("multiome") RNA velocity 방법들은 유전자별로 전사 속도(transcription rate), 분해 속도(degradation rate), chromatin에서 transcription으로 이어지는 시간차(lag, locus가 열리거나 닫히는 시점과 transcription 전환 시점의 오프셋)를 산출한다. 이 값들은 저마다 생물학적 판독값으로 제안되어 왔고, 시간차(lag)는 특히 epigenetic 약물 반응의 timing 예측에 쓰일 수 있다고 여겨져 왔다. 그러나 파생된 값은 신뢰할 수 있어야만, 곧 합리적인 알고리즘들에 걸쳐 재현되고 가능하면 독립적인 측정값과도 부합해야만 하류에서 쓸 수 있다. 우리는 10x Multiome으로 프로파일링한 인간 조혈모·전구세포(hematopoietic stem and progenitor cells, HSPC)에서 어떤 velocity 출력이 이 기준을 충족하는지를 물었다. +**배경.** Chromatin 정보를 결합한("multiome") RNA velocity 방법들은 유전자별로 전사 속도(transcription rate), 분해 속도(degradation rate), chromatin에서 transcription으로 이어지는 시간차(lag, locus가 열리거나 닫히는 시점과 transcription 전환 시점의 오프셋)를 산출한다. 이 값들은 저마다 생물학적 판독값으로 제안되어 왔고, 시간차(lag)는 특히 epigenetic 약물 반응의 timing 예측에 쓰일 수 있다고 여겨져 왔다. 그러나 파생된 값은 합리적인 알고리즘들에 걸쳐 재현되고 가능하면 독립적인 측정값과도 부합해야만 하류에서 쓸모가 있다. 우리는 10x Multiome으로 프로파일링한 인간 조혈모·전구세포(hematopoietic stem and progenitor cells, HSPC)에서 어떤 velocity 출력이 이 기준을 충족하는지를 물었다. -**접근.** 최대 다섯 개 velocity arm(RNA 전용 scVelo에 MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo를 더한 것)에 걸쳐 각 출력을 네 축에서 검정했다. 방법 간 재현성, lineage 내 ATAC-shuffle 인과 대조군, 다섯 개 외부 multiome에서의 cross-dataset 재현(그중 하나는 사전등록), K562 TT-seq 합성 속도와 mRNA 반감기 분해 속도로의 외부 anchoring이 그것이다. 우리는 유전자별 kinetic parameter와 세포×유전자 velocity 행렬을 감사 대상으로 삼았고, velocity가 주로 쓰이는 저차원 임베딩 arrow나 궤적은 다루지 않았다. +**접근.** 최대 다섯 개 velocity arm(RNA 전용 scVelo에 MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo를 더한 것)에 걸쳐 각 출력을 네 축에서 검정했다. 방법 간 재현성, ATAC-shuffle 인과 대조군, 다섯 개 외부 multiome에서의 cross-dataset 재현(그중 하나는 사전등록), K562 TT-seq 합성 속도와 반감기 분해 속도로의 anchoring이 그것이다. 우리는 유전자별 kinetic parameter와 세포×유전자 velocity 행렬을 감사 대상으로 삼았고, velocity가 주로 쓰이는 저차원 임베딩 arrow나 궤적은 다루지 않았다. -**결과.** α만 방법 간에 재현되었고(Spearman ρ=0.88, 관측값), 시간차(lag)와 분해 속도 γ는 그렇지 않았다. 시간차(lag)는 크기에서 잘해야 약하게 재현되었고(가장 강한 쌍 ρ=+0.163, 대부분 |ρ|≤0.08) 부호에서는 우연 수준(54.6%)에 그쳤으며, ATAC를 뒤섞어도 통계적으로 변하지 않아 chromatin이 아니라 모델 구조에서 비롯됨을 드러냈다. γ도 마찬가지로 방법에 취약했고(방법 간 ρ≈−0.1) ground truth 앞에서도 복원되지 않았다(K562 반감기 3/3 귀무; 교과서적 scVelo γ는 반대로 나와 −0.224, CI가 0 배제). fitting된 α는 세 방법 모두에서 K562 TT-seq 합성 속도와 상관했으나(ρ +0.24~+0.29, 모든 CI가 0 배제), steady-state abundance가 같은 측정값을 최소한 그만큼 잘 예측해(+0.410 대 +0.262) 이를 α만의 고유한 측정 근거가 아니라 일관성 증거로 다룬다. α가 시간차(lag)보다 앞선다는 순서는 여섯 시스템 모두에서 유지되어, 어떤 fitting보다 먼저 봉인한 사전등록 6-of-6 채점표를 통과했다. profile-likelihood 분석은 그 이유를 보여준다. α는 stiff(식별 가능)한 반면 시간차(lag)는 sloppy하고 경계에 제약되며, 이 식별가능성은 외부 검증과 부분적으로만 정렬한다. +**결과.** α만 방법 간에 재현되었고(Spearman ρ=0.88, 관측값), 시간차(lag)와 분해 속도 γ는 그렇지 않았다. 시간차(lag)는 크기에서 잘해야 약하게 재현되었고(대부분 |ρ|≤0.08) 부호에서는 우연 수준(54.6%)에 그쳤으며, ATAC를 뒤섞어도 통계적으로 변하지 않아 chromatin이 아니라 모델 구조에서 비롯됨을 드러냈다. γ도 마찬가지로 방법에 취약했고(방법 간 ρ≈−0.1) ground truth 앞에서도 복원되지 않았다(K562 반감기 3/3 귀무; 교과서적 scVelo γ는 반대로까지 나왔다). fitting된 α는 세 방법 모두에서 K562 TT-seq 합성 속도와 상관했으나(ρ +0.24~+0.29, 모든 CI가 0 배제) abundance가 같은 측정값을 최소한 그만큼 잘 예측해, 이를 고유한 측정 근거가 아니라 일관성으로 다룬다. α가 시간차(lag)보다 앞선다는 순서는 여섯 시스템 모두에서 유지되어 사전등록 6-of-6 채점표를 통과했다. profile-likelihood 분석은 그 이유를 보여준다. α는 식별 가능한 반면 시간차(lag)는 sloppy하고 경계에 제약된다. -**결론.** 우리는 이 결과를 velocity 출력 신뢰도 지도로 정리한다. α와 속도 파생 신호는 방법 간에 재현되어 신뢰할 수 있으나, α는 대체로 발현량을 반영하는 값이라(abundance가 같은 외부 측정값을 α만큼 잘 예측해 α가 그 이상의 정보를 더한다고는 입증되지 않는다) 시간차(lag), 그 부호, 절대 timing, γ는 신뢰할 수 없는 것으로 다루며 직교(orthogonal) 검증을 요구한다. 하류의 timing 예측 모델은 단일 방법의 시간차(lag)가 아니라 강건한 baseline-to-α 경로를 거쳐야 한다. +**결론.** 우리는 이 결과를 velocity 출력 신뢰도 지도로 정리한다. α와 속도 파생 신호는 방법 간에 재현되어 신뢰할 수 있으나, α는 대체로 발현량을 반영하는 값이라(abundance가 같은 외부 측정값을 α만큼 잘 예측해 α가 그 이상의 정보를 더한다고는 입증되지 않는다) 시간차(lag), 부호, timing, γ는 신뢰할 수 없는 것으로 다루며 직교(orthogonal) 검증을 요구한다. 하류의 timing 예측 모델은 단일 방법의 시간차(lag)가 아니라 강건한 baseline-to-α 경로를 거쳐야 한다. **키워드:** RNA velocity, single-cell multiome, chromatin accessibility, transcriptional kinetics, parameter identifiability, external validation, benchmarking, hematopoiesis From 5949430c4e77ce85322c09cb85ed97df1d921a7d Mon Sep 17 00:00:00 2001 From: JamieLyu <5819016+JamieLyu@users.noreply.github.com> Date: Thu, 6 Aug 2026 15:04:46 +0900 Subject: [PATCH 3/3] =?UTF-8?q?P5=20=EC=B4=88=EB=A1=9D=20333=E2=86=92318?= =?UTF-8?q?=EB=8B=A8=EC=96=B4(EN),=20GB=5FBENCHMARK.md=20G1=20=EB=AA=A9?= =?UTF-8?q?=ED=91=9C(=E2=89=A4320)=20=EB=8B=AC=EC=84=B1?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 수치 손실 없이(0.163/0.224/0.262/0.410은 앞선 커밋에서 이미 Results로 이동, 이번엔 순수 어구 압축만) 320단어 이하로 마저 좁혔다. 국문 초록도 비례 감축(309→302단어). 헤드라인 수치·hedge·주장 범위는 전부 불변. --- pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md | 8 ++++---- .../hspc-velocity-benchmark/manuscript/draft_v2_ko.md | 8 ++++---- 2 files changed, 8 insertions(+), 8 deletions(-) diff --git a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md index feb0589..ee7b92e 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md @@ -36,13 +36,13 @@ the external-measurement result is corroboration inside the map, not a title cla ## Abstract -**Background.** Chromatin-informed ("multiome") RNA-velocity methods emit several per-gene quantities, a transcription rate, a degradation rate, and a chromatin-to-transcription *lag* (the offset between chromatin opening/closing and the transcriptional switch), each proposed as a biological readout, the lag for predicting epigenetic-drug timing. A derived quantity is useful downstream only if reproducible across algorithms and, where possible, consistent with independent measurement. We asked which outputs meet that bar in human hematopoietic stem and progenitor cells (HSPCs) profiled by 10x Multiome. +**Background.** Chromatin-informed ("multiome") RNA-velocity methods emit several per-gene quantities, a transcription rate, a degradation rate, and a chromatin-to-transcription *lag* (the offset between chromatin opening/closing and the transcriptional switch), each proposed as a biological readout, the lag for epigenetic-drug timing. A derived quantity is useful downstream only if reproducible across algorithms and, where possible, consistent with independent measurement. We asked which outputs meet that bar in human hematopoietic stem and progenitor cells (HSPCs) profiled by 10x Multiome. -**Approach.** Across up to five velocity arms (RNA-only scVelo plus MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo), we tested each output on four axes: cross-method reproducibility, a causal ATAC-shuffle control, cross-dataset replication in five external multiomes (one preregistered), and anchoring to K562 TT-seq synthesis and half-life degradation rates. We audit per-gene kinetic parameters and the cell×gene velocity matrix, not the embedding arrows or trajectories velocity is principally used for. +**Approach.** Across up to five velocity arms (RNA-only scVelo plus MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo), we tested each output on four axes: cross-method reproducibility, a causal ATAC-shuffle control, cross-dataset replication in five external multiomes (one preregistered), and anchoring to K562 TT-seq synthesis and half-life degradation rates. We audit per-gene kinetic parameters and the cell×gene velocity matrix, not the embedding arrows or trajectories. -**Findings.** Only α reproduced across methods (Spearman ρ=0.88, observed); the lag and degradation rate γ did not. The lag reproduced weakly in magnitude (most pairs |ρ|≤0.08) and only at chance in sign (54.6%); ATAC-shuffling left it statistically unchanged, marking it model-structural rather than chromatin-driven. γ was likewise method-fragile (ρ≈−0.1) and unrecovered against ground truth (K562 half-life 3/3 null; the textbook scVelo γ even ran reversed). Fitted α tracked K562 TT-seq synthesis in all three methods (ρ +0.24 to +0.29, CIs excluding 0), though abundance tracked it at least as strongly, so we treat this as consistency, not unique grounding. The α-over-lag ordering held in all six systems, passing a preregistered six-of-six scorecard. Profile-likelihood analysis explained why: α is identifiable, the lag sloppy and boundary-limited. +**Findings.** Only α reproduced across methods (Spearman ρ=0.88, observed); the lag and degradation rate γ did not. The lag reproduced weakly in magnitude (most pairs |ρ|≤0.08) and only at chance in sign (54.6%); ATAC-shuffling left it unchanged, marking it model-structural, not chromatin-driven. γ was likewise fragile (ρ≈−0.1) and unrecovered against ground truth (K562 half-life 3/3 null; textbook scVelo γ even ran reversed). Fitted α tracked K562 TT-seq synthesis in all three methods (ρ 0.24–0.29, CIs excluding 0), though abundance tracked it as strongly, so we treat this as consistency, not unique grounding. The α-over-lag ordering held in all six systems, passing a preregistered six-of-six scorecard. Profile-likelihood analysis explained why: α is identifiable, the lag sloppy and boundary-limited. -**Conclusions.** We distil these results into a velocity-output reliability map: trust α and rate-derived signals, reproducible across methods, though α is largely expression (abundance tracks the same external measurement at least as strongly, so α adds no demonstrable information beyond it); treat the lag, sign, timing, and γ as unreliable, requiring orthogonal validation. Downstream timing-prediction should route through the robust baseline-to-α path, not a single-method lag. +**Conclusions.** We distil these results into a velocity-output reliability map: trust α and rate-derived signals, reproducible across methods, though α is largely expression (abundance tracks the same external measurement as strongly, so α adds no demonstrable information beyond it); treat the lag, sign, timing, and γ as unreliable, requiring orthogonal validation. Downstream timing-prediction should route through the robust baseline-to-α path, not a single-method lag. **Keywords:** RNA velocity, single-cell multiome, chromatin accessibility, transcriptional kinetics, parameter identifiability, external validation, benchmarking, hematopoiesis diff --git a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md index dad281f..87d8a5d 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md @@ -19,13 +19,13 @@ draft_v2.md의 한국어 검토본(번역·윤문). 정본은 영어 draft_v2.md ## Abstract -**배경.** Chromatin 정보를 결합한("multiome") RNA velocity 방법들은 유전자별로 전사 속도(transcription rate), 분해 속도(degradation rate), chromatin에서 transcription으로 이어지는 시간차(lag, locus가 열리거나 닫히는 시점과 transcription 전환 시점의 오프셋)를 산출한다. 이 값들은 저마다 생물학적 판독값으로 제안되어 왔고, 시간차(lag)는 특히 epigenetic 약물 반응의 timing 예측에 쓰일 수 있다고 여겨져 왔다. 그러나 파생된 값은 합리적인 알고리즘들에 걸쳐 재현되고 가능하면 독립적인 측정값과도 부합해야만 하류에서 쓸모가 있다. 우리는 10x Multiome으로 프로파일링한 인간 조혈모·전구세포(hematopoietic stem and progenitor cells, HSPC)에서 어떤 velocity 출력이 이 기준을 충족하는지를 물었다. +**배경.** Chromatin 정보를 결합한("multiome") RNA velocity 방법들은 유전자별로 전사 속도(transcription rate), 분해 속도(degradation rate), chromatin에서 transcription으로 이어지는 시간차(lag, locus가 열리거나 닫히는 시점과 transcription 전환 시점의 오프셋)를 산출한다. 이 값들은 저마다 생물학적 판독값으로 제안되어 왔고, 시간차(lag)는 특히 epigenetic 약물 반응 timing에 쓰일 수 있다고 여겨져 왔다. 그러나 파생된 값은 합리적인 알고리즘들에 걸쳐 재현되고 가능하면 독립적인 측정값과도 부합해야만 하류에서 쓸모가 있다. 우리는 10x Multiome으로 프로파일링한 인간 조혈모·전구세포(hematopoietic stem and progenitor cells, HSPC)에서 어떤 velocity 출력이 이 기준을 충족하는지를 물었다. -**접근.** 최대 다섯 개 velocity arm(RNA 전용 scVelo에 MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo를 더한 것)에 걸쳐 각 출력을 네 축에서 검정했다. 방법 간 재현성, ATAC-shuffle 인과 대조군, 다섯 개 외부 multiome에서의 cross-dataset 재현(그중 하나는 사전등록), K562 TT-seq 합성 속도와 반감기 분해 속도로의 anchoring이 그것이다. 우리는 유전자별 kinetic parameter와 세포×유전자 velocity 행렬을 감사 대상으로 삼았고, velocity가 주로 쓰이는 저차원 임베딩 arrow나 궤적은 다루지 않았다. +**접근.** 최대 다섯 개 velocity arm(RNA 전용 scVelo에 MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo를 더한 것)에 걸쳐 각 출력을 네 축에서 검정했다. 방법 간 재현성, ATAC-shuffle 인과 대조군, 다섯 개 외부 multiome에서의 cross-dataset 재현(그중 하나는 사전등록), K562 TT-seq 합성 속도와 반감기 분해 속도로의 anchoring이 그것이다. 우리는 유전자별 kinetic parameter와 세포×유전자 velocity 행렬을 감사 대상으로 삼았고, 저차원 임베딩 arrow나 궤적은 다루지 않았다. -**결과.** α만 방법 간에 재현되었고(Spearman ρ=0.88, 관측값), 시간차(lag)와 분해 속도 γ는 그렇지 않았다. 시간차(lag)는 크기에서 잘해야 약하게 재현되었고(대부분 |ρ|≤0.08) 부호에서는 우연 수준(54.6%)에 그쳤으며, ATAC를 뒤섞어도 통계적으로 변하지 않아 chromatin이 아니라 모델 구조에서 비롯됨을 드러냈다. γ도 마찬가지로 방법에 취약했고(방법 간 ρ≈−0.1) ground truth 앞에서도 복원되지 않았다(K562 반감기 3/3 귀무; 교과서적 scVelo γ는 반대로까지 나왔다). fitting된 α는 세 방법 모두에서 K562 TT-seq 합성 속도와 상관했으나(ρ +0.24~+0.29, 모든 CI가 0 배제) abundance가 같은 측정값을 최소한 그만큼 잘 예측해, 이를 고유한 측정 근거가 아니라 일관성으로 다룬다. α가 시간차(lag)보다 앞선다는 순서는 여섯 시스템 모두에서 유지되어 사전등록 6-of-6 채점표를 통과했다. profile-likelihood 분석은 그 이유를 보여준다. α는 식별 가능한 반면 시간차(lag)는 sloppy하고 경계에 제약된다. +**결과.** α만 방법 간에 재현되었고(Spearman ρ=0.88, 관측값), 시간차(lag)와 분해 속도 γ는 그렇지 않았다. 시간차(lag)는 크기에서 잘해야 약하게 재현되었고(대부분 |ρ|≤0.08) 부호에서는 우연 수준(54.6%)에 그쳤으며, ATAC를 뒤섞어도 변하지 않아 chromatin이 아니라 모델 구조에서 비롯됨을 드러냈다. γ도 마찬가지로 취약했고(방법 간 ρ≈−0.1) ground truth 앞에서도 복원되지 않았다(K562 반감기 3/3 귀무; 교과서적 scVelo γ는 반대로까지 나왔다). fitting된 α는 세 방법 모두에서 K562 TT-seq 합성 속도와 상관했으나(ρ 0.24~0.29, 모든 CI가 0 배제) abundance가 같은 측정값을 그만큼 잘 예측해, 이를 고유한 측정 근거가 아니라 일관성으로 다룬다. α가 시간차(lag)보다 앞선다는 순서는 여섯 시스템 모두에서 유지되어 사전등록 6-of-6 채점표를 통과했다. profile-likelihood 분석은 그 이유를 보여준다. α는 식별 가능한 반면 시간차(lag)는 sloppy하고 경계에 제약된다. -**결론.** 우리는 이 결과를 velocity 출력 신뢰도 지도로 정리한다. α와 속도 파생 신호는 방법 간에 재현되어 신뢰할 수 있으나, α는 대체로 발현량을 반영하는 값이라(abundance가 같은 외부 측정값을 α만큼 잘 예측해 α가 그 이상의 정보를 더한다고는 입증되지 않는다) 시간차(lag), 부호, timing, γ는 신뢰할 수 없는 것으로 다루며 직교(orthogonal) 검증을 요구한다. 하류의 timing 예측 모델은 단일 방법의 시간차(lag)가 아니라 강건한 baseline-to-α 경로를 거쳐야 한다. +**결론.** 우리는 이 결과를 velocity 출력 신뢰도 지도로 정리한다. α와 속도 파생 신호는 방법 간에 재현되어 신뢰할 수 있으나, α는 대체로 발현량을 반영하는 값이라(abundance가 같은 외부 측정값을 그만큼 잘 예측해 α가 그 이상의 정보를 더한다고는 입증되지 않는다) 시간차(lag), 부호, timing, γ는 신뢰할 수 없는 것으로 다루며 직교(orthogonal) 검증을 요구한다. 하류의 timing 예측 모델은 단일 방법의 시간차(lag)가 아니라 강건한 baseline-to-α 경로를 거쳐야 한다. **키워드:** RNA velocity, single-cell multiome, chromatin accessibility, transcriptional kinetics, parameter identifiability, external validation, benchmarking, hematopoiesis