diff --git a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md index 0991913..6b5f6a4 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md @@ -116,7 +116,7 @@ Per-gene kinetic parameters are not what velocity is mainly used for, so we exte Two controls frame these numbers, and both are MultiVelo's, because it is the only arm whose refits were retained. Refitting MultiVelo on resampled cells reproduces its own matrix at mean-centred cosine +0.872 (six refits, range +0.826 to +0.887; sign agreement 78.6%), so the measure does detect agreement where agreement exists. The second control destroys the chromatin channel by permuting the ATAC rows (cells) within each lineage, which leaves the RNA channel and the lineage-level chromatin structure intact. Our first version of this second comparison was not like-for-like, and we correct it here: the refit ceiling was computed on 15,315 resampled cells while the shuffled fit used all 21,878, so the two arms did not share a cell set, and the resulting shuffled value (+0.838) sat inside the ceiling range only under that mismatch. Repeating the shuffle on the *same* resampled cell sets S_b used by three of the refits — identical hyperparameters, gene set and shuffle protocol, with the two arms' cell-name vectors asserted equal at both fit and analysis time — puts the shuffled matrix below the intact refit range in all three pairs (+0.784, +0.813 and +0.810 against an intact range of +0.826 to +0.887), with non-overlapping cell-bootstrap intervals in three of three (`results/velocity_matrix_paired_shuffle.md`). We therefore withdraw our earlier statement that chromatin is inert in this matrix; it can no longer be asserted. -What replaces it is bounded rather than absent. The paired differences (Δ = A − B) are small and all positive, but their magnitude is not fixed to a single value: it varies with the shuffle draw, spanning roughly two- to three-fold within a single cell set, so no one draw is a representative size. Across draws Δ sits at a single-digit percentage of the +0.872 ceiling (of order 7%), well below the disagreement that method choice produces on the same measure (mean-centred cosine −0.530 to +0.131 across method pairs). Whether any of that residue is chromatin rather than refitting was initially unresolved, because no same-cell, same-configuration rerun null existed for MultiVelo; that gap has since been closed, and it closes towards a contribution rather than towards noise. Refitting on the same cell set S_b with chromatin intact, varying nothing but the worker count, reproduces the previous fit exactly — the two same-day fits are bit-identical and the run-to-run difference on the audit's own metric is Δ_rr = 0.000000, with a degenerate cell-bootstrap interval, for both cell sets measured (`results/velocity_matrix_runtorun_null.md`). Against the archived fits the only departure is a single gene (*LRIG1*): a 4×10⁻¹⁶ rounding difference in one cell set and, in the other, an alternative solution of near-equal loss (switch time 65.0 versus 66.9) that moves the per-cell mean-centred cosine by at most 2×10⁻⁶. Because the fit is deterministic under the conditions we use, the paired Δ carries no refit-noise component to subtract, and the preregistered condition for a positive statement — the upper interval of |Δ_rr| below the lower interval of Δ_paired — holds in both cases (0.000000 against 0.0792 and 0.0608). We therefore state, for MultiVelo and for this matrix, that destroying the chromatin channel does move the output and that the movement is not an artefact of refitting. Three limits keep the statement small. Only two of the three pairs have a matched rerun null, and the pair with the smallest original draw is not among them. Repeating the shuffle under four different draws per cell set (twelve fits in total) separates what survives resampling from what does not: the sign survives — all twelve draws give Δ>0 with intervals excluding zero — while the magnitude does not, since the spread across draws exceeds the median in one cell set (range 0.043 against a median of 0.031) and the single draw we had reported there (+0.063) turns out to be the largest of the four (`results/velocity_matrix_shuffle_seed_variability.md`). This dispersion verdict is sealed on the range, and a variance-based dispersion measure would instead pass; either way the substantive point — a two- to three-fold spread of magnitude across draws — is metric-independent. We therefore state the direction and decline to pin the size to a number; four draws per cell set is itself few, so the spread is coarsely estimated. And the effect stays an order of magnitude below what method choice produces on the same measure, so "chromatin contributes here" is not "chromatin makes this matrix reliable". This matrix-level movement is also not in tension with the lag result: the same class of ATAC shuffle does not perturb the priming-marker lags more than a bulk shuffle (Fig. 2, chromatin does not set the lag), yet here it moves the cell×gene velocity matrix; the two findings concern different targets and are consistent rather than contradictory. We report these comparisons on mean-centred cosine because the raw value is partly determined by a direction common to all cells (the mean vector accounts for 12.9–37.4% of squared row norm, depending on the arm); centring is a post-hoc diagnostic and is not part of the sealed metric list, and the preregistered verdict above rests on the raw metric. One further limit is load-bearing: reading the cross-method values as genuine disagreement rather than arm-internal instability is licensed only for MultiVelo, which has that control. The three pairs closest to zero all involve MoFlow, a stochastic deep model whose same-configuration rerun stability was never established (its own original-versus-shuffled value, +0.113, is uninterpretable for the same reason), so for those pairs disagreement and instability are not separable. +What replaces it is bounded rather than absent. The paired differences (Δ = A − B) are small and all positive, but their magnitude is not fixed to a single value: it varies with the shuffle draw, spanning roughly two- to three-fold within a single cell set, so no one draw is a representative size. Across draws Δ sits at a single-digit percentage of the +0.872 ceiling (of order 7%), well below the disagreement that method choice produces on the same measure (mean-centred cosine −0.530 to +0.131 across method pairs). Whether any of that residue is chromatin rather than refitting was initially unresolved, because no same-cell, same-configuration rerun null existed for MultiVelo; that gap has since been closed, and it closes towards a contribution rather than towards noise. Refitting on the same cell set S_b with chromatin intact, varying nothing but the worker count, reproduces the previous fit exactly — the two same-day fits are bit-identical and the run-to-run difference on the audit's own metric is Δ_rr = 0.000000, with a degenerate cell-bootstrap interval, for both cell sets measured (`results/velocity_matrix_runtorun_null.md`). Against the archived fits the only departure is a single gene (*LRIG1*): a 4×10⁻¹⁶ rounding difference in one cell set and, in the other, an alternative solution of near-equal loss (switch time 65.0 versus 66.9) that moves the per-cell mean-centred cosine by at most 2×10⁻⁶. Because the fit is deterministic under the conditions we use, the paired Δ carries no refit-noise component to subtract, and the preregistered condition for a positive statement — the upper interval of |Δ_rr| below the lower interval of Δ_paired — holds in both cases (0.000000 against 0.0792 and 0.0608). We therefore state, for MultiVelo and for this matrix, that destroying the chromatin channel does move the output and that the movement is not an artefact of refitting. Three limits keep the statement small. Only two of the three pairs have a matched rerun null, and the pair with the smallest original draw is not among them. Repeating the shuffle under four different draws per cell set (twelve fits in total) separates what survives resampling from what does not: the sign survives — all twelve draws give Δ>0 with intervals excluding zero — while the magnitude does not, since the spread across draws exceeds the median in one cell set (range 0.043 against a median of 0.031) and the single draw we had reported there (+0.063) turns out to be the largest of the four (`results/velocity_matrix_shuffle_seed_variability.md`). This dispersion verdict is sealed on the range, and a variance-based dispersion measure would instead pass; either way the substantive point — a two- to three-fold spread of magnitude across draws — is metric-independent. We therefore state the direction and decline to pin the size to a number; four draws per cell set is itself few, so the spread is coarsely estimated. And the effect stays an order of magnitude below what method choice produces on the same measure, so "chromatin contributes here" is not "chromatin makes this matrix reliable". This matrix-level movement is also not in tension with the lag result: the same class of ATAC shuffle does not perturb the priming-marker lags more than a bulk shuffle (Fig. 2, chromatin does not set the lag), yet here it moves the cell×gene velocity matrix; the two findings concern different targets and are consistent rather than contradictory. We report these comparisons on mean-centred cosine because the raw value is partly determined by a direction common to all cells (the mean vector accounts for 12.9–37.4% of squared row norm, depending on the arm); centring is a post-hoc diagnostic and is not part of the sealed metric list, and the preregistered verdict above rests on the raw metric. One further limit is load-bearing: reading the cross-method values as genuine disagreement rather than arm-internal instability is licensed only for MultiVelo, which has that control. That licence has since extended to MoFlow. A matched same-configuration rerun, two independent refits on the identical cell and gene set, reproduces MoFlow's original matrix at mean-centred cosine +0.9999 [+0.9999, +0.9999] (run-versus-run, +1.0000), a ceiling far above the 0.0118 maximum among the three near-zero pairs that involve it (`results/moflow_runtorun_null.md`). Those near-zero values therefore read as genuine near-zero cross-method agreement rather than MoFlow's own instability, and the same ceiling places MoFlow's original-versus-shuffled value (+0.113) well clear of run-to-run noise. Read within those limits, the largest mean-centred agreement anywhere in the comparison was between MultiVelo and the RNA-only scVelo floor (+0.583) rather than between two multiome methods — though this is not a family property, since the corresponding values for CRAK-Velo, MoFlow and MultiVeloVAE were +0.260, −0.004 and −0.292. MultiVelo and MultiVeloVAE assigned systematically opposite directions to the same cells (mean-centred −0.500); whether that is a substantive disagreement or an undocumented difference in sign or parameterisation convention cannot be settled by this design, but either way an analyst who swaps one output for the other without checking obtains opposing directions. Both contrasts failed on the metric and thresholds sealed before the fitted matrices were read (Additional file 12): multiome pairs did not agree more than the RNA-only baseline, and destroying chromatin did not collapse the matrix, although the paired comparison above shows that it does move it a little. The matrix therefore reproduces across methods no better than the per-gene parameters did. It also sits alongside the general benchmarks, which report low cross-method agreement of transition vectors in RNA-only settings (A1<0.3 across the twelve methods compared) [25]. @@ -204,7 +204,7 @@ All cross-method and cross-dataset concordances were Spearman rank correlations. ### Cell-level velocity-matrix audit -Definitions, exclusion rules and falsification criteria were sealed before the fitted matrices were read. Cells and genes were matched by name across arms; all five arms share the same 21,878 cells, and the 505 genes common to all five were reduced to 354 after excluding any gene with a missing (scVelo leaves 151 genes with unrecovered dynamics) or constant velocity in any arm. Only spliced velocity was compared (MultiVelo and MoFlow `velo_s`, CRAK-Velo and scVelo `velocity`, MultiVeloVAE `vae_velocity`); chromatin velocity layers were never mixed in. Agreement was the per-cell cosine similarity across genes, reported against a cell-shuffled null (method A's cell *i* against method B's randomly permuted cell *j*) because a globally shared velocity direction inflates the raw value; medians carry a cell bootstrap 95% CI (B=10³, seed 20260719). Because raw cosine is partly determined by a direction common to all cells (the mean vector accounts for 12.9–37.4% of squared row norm depending on the arm), we also report the mean-centred cosine, which measures agreement of the cell-specific deviations; centring is applied without scaling so that sign is preserved. Per-(cell,gene) sign agreement excludes entries that are exactly zero in either arm, following the same undetermined-direction convention used for the lag. Two reference points frame the cross-method values: a reproducibility ceiling from six MultiVelo bootstrap refits on resampled cells, and the ATAC-shuffle arms used for the lag control. Because those two were originally computed on different cell sets, the within-lineage ATAC shuffle was additionally refit on the same three resampled cell sets used by three of the refits, with identical hyperparameters, gene set and shuffle protocol and with the two arms' cell-name vectors asserted equal, giving a paired comparison and paired cell-bootstrap intervals (`results/velocity_matrix_paired_shuffle.md`). The causal reading is restricted to MultiVelo, the only arm with a refit control, and no same-cell same-configuration rerun null was available for either arm. +Definitions, exclusion rules and falsification criteria were sealed before the fitted matrices were read. Cells and genes were matched by name across arms; all five arms share the same 21,878 cells, and the 505 genes common to all five were reduced to 354 after excluding any gene with a missing (scVelo leaves 151 genes with unrecovered dynamics) or constant velocity in any arm. Only spliced velocity was compared (MultiVelo and MoFlow `velo_s`, CRAK-Velo and scVelo `velocity`, MultiVeloVAE `vae_velocity`); chromatin velocity layers were never mixed in. Agreement was the per-cell cosine similarity across genes, reported against a cell-shuffled null (method A's cell *i* against method B's randomly permuted cell *j*) because a globally shared velocity direction inflates the raw value; medians carry a cell bootstrap 95% CI (B=10³, seed 20260719). Because raw cosine is partly determined by a direction common to all cells (the mean vector accounts for 12.9–37.4% of squared row norm depending on the arm), we also report the mean-centred cosine, which measures agreement of the cell-specific deviations; centring is applied without scaling so that sign is preserved. Per-(cell,gene) sign agreement excludes entries that are exactly zero in either arm, following the same undetermined-direction convention used for the lag. Two reference points frame the cross-method values: a reproducibility ceiling from six MultiVelo bootstrap refits on resampled cells, and the ATAC-shuffle arms used for the lag control. Because those two were originally computed on different cell sets, the within-lineage ATAC shuffle was additionally refit on the same three resampled cell sets used by three of the refits, with identical hyperparameters, gene set and shuffle protocol and with the two arms' cell-name vectors asserted equal, giving a paired comparison and paired cell-bootstrap intervals (`results/velocity_matrix_paired_shuffle.md`). The causal reading is restricted to MultiVelo, the only arm with a cell-bootstrap refit control. A same-cell, same-configuration rerun null (fixed cells and genes, independent refits) was subsequently established for MultiVelo (`results/velocity_matrix_runtorun_null.md`) and, separately, for MoFlow (`results/moflow_runtorun_null.md`); see Results. ### Enrichment analysis diff --git a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md index 457a6d5..811947a 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2_ko.md @@ -99,7 +99,7 @@ fitting된 모수를 cross-method 재현성으로 순위 매기면 경험적 식 이 숫자들을 감싸는 대조군이 둘인데, 재적합 결과가 보존된 arm이 MultiVelo뿐이라 둘 다 MultiVelo 기준이다. MultiVelo를 재표집한 세포에 다시 fitting하면 자기 행렬을 평균 중심화 코사인 +0.872로 재현한다(재적합 6회, 범위 +0.826 ~ +0.887; 부호 일치 78.6%). 따라서 이 지표는 일치가 있는 곳에서 일치를 검출한다. 두 번째 대조군은 각 lineage 내부에서 ATAC 행(세포)을 permute해 크로마틴 채널을 파괴하며, RNA 채널과 lineage 수준의 크로마틴 구조는 그대로 남긴다. 이 두 번째 비교의 첫 판본은 같은 조건끼리의 비교가 아니었고, 여기서 바로잡는다. 재적합 천장은 재표집한 15,315개 세포에서, 셔플 적합은 21,878개 전량에서 계산해 두 arm이 세포집합을 공유하지 않았으며, 셔플 값(+0.838)이 천장 범위 안에 들어온 것도 그 불일치 아래에서만 성립했다. 재적합 3개가 쓴 *같은* 재표집 세포집합 S_b 위에서 셔플을 다시 적합하자(하이퍼파라미터·유전자 집합·셔플 규약 동일, 두 arm의 세포 이름 벡터가 같음을 적합 시점과 분석 시점 모두 assert로 확인) 셔플 행렬은 세 쌍 모두에서 온전한 재적합 범위 아래로 내려갔고(+0.784, +0.813, +0.810 대 온전한 범위 +0.826 ~ +0.887), 세포 부트스트랩 구간도 3개 중 3개가 겹치지 않았다(`results/velocity_matrix_paired_shuffle.md`). 따라서 크로마틴이 이 행렬에 무력하다는 앞선 진술을 철회한다. 더는 단정할 수 없다. -그 자리를 대신하는 진술은 없는 것이 아니라 경계가 그어진 것이다. 짝지은 차이(Δ = A − B)는 작고 모두 양수이지만, 그 크기는 하나의 값으로 고정되지 않는다. 셔플 draw에 따라 달라져 한 세포집합 안에서도 대략 두세 배로 흩어지므로, 어느 한 draw도 대표 크기가 아니다. draw 전반에서 Δ는 천장 +0.872의 한 자릿수 %(대략 7%) 수준으로, 같은 지표에서 method 선택이 만들어 내는 불일치(방법 쌍 간 평균 중심화 코사인 −0.530 ~ +0.131)보다 한참 아래다. 그 잔여분이 크로마틴인지 재적합인지는 처음에 가릴 수 없었다. MultiVelo에 대해 같은 세포·같은 설정의 재실행 귀무가 없었기 때문이다. 그 공백은 이후 메워졌고, 잡음이 아니라 기여 쪽으로 메워졌다. 같은 세포집합 S_b에서 크로마틴을 온전히 둔 채 worker 수만 바꿔 재적합하면 이전 적합이 그대로 재현된다. 같은 날 두 적합은 bit 단위로 동일하고, 감사 지표로 잰 실행 간 차이는 측정한 두 세포집합 모두에서 Δ_rr = 0.000000이며 세포 부트스트랩 구간도 퇴화한다(`results/velocity_matrix_runtorun_null.md`). 보관된 적합과 견주었을 때 어긋나는 것은 유전자 하나(*LRIG1*)뿐으로, 한 세포집합에서는 4×10⁻¹⁶의 반올림 차이이고 다른 하나에서는 loss가 거의 같은 대안 해(전환 시각 65.0 대 66.9)이며 세포별 평균 중심화 코사인을 최대 2×10⁻⁶ 움직인다. 우리가 쓰는 조건에서 적합이 결정론적이므로 짝지은 Δ에는 빼내야 할 재적합 잡음 성분이 없고, 양성 진술의 사전등록 조건, 곧 |Δ_rr|의 상한 구간이 Δ_paired의 하한 구간보다 낮다는 조건이 두 경우 모두에서 성립한다(0.000000 대 0.0792 및 0.0608). 따라서 MultiVelo에 대해, 그리고 이 행렬에 대해, 크로마틴 채널을 파괴하면 출력이 실제로 움직이며 그 움직임이 재적합의 인공물이 아니라고 진술한다. 이 진술을 작게 묶어 두는 한계가 셋이다. 세 쌍 중 둘만 짝맞춘 재실행 귀무를 가지며, 원래 draw가 가장 작았던 쌍은 거기 들어 있지 않다. 세포집합마다 서로 다른 셔플 추출 네 번(총 열두 번의 적합)으로 반복해 보면, 재표집을 견디는 것과 견디지 못하는 것이 나뉜다. **부호는 견딘다.** 열두 draw 전부 Δ>0이고 구간이 0을 배제한다. **크기는 견디지 못한다.** 한 세포집합에서 draw 간 산포가 중앙값을 넘고(범위 0.043 대 중앙값 0.031), 거기서 우리가 보고하던 단일 draw(+0.063)는 네 개 중 가장 큰 값이었다(`results/velocity_matrix_shuffle_seed_variability.md`). 이 산포 판정은 범위(range)를 기준으로 봉인했고, 분산 기반 산포 지표로 보면 오히려 통과한다. 어느 지표든 실질 결론, 곧 크기가 draw마다 두세 배로 흩어진다는 점은 지표와 무관하게 성립한다. 따라서 방향은 진술하되 크기는 특정 숫자로 못 박지 않는다. 세포집합당 네 draw도 그 자체로 적어 산포는 거칠게만 추정된다. 그리고 효과는 같은 지표에서 method 선택이 만드는 것보다 한 자릿수 아래에 머무르므로, "여기서 크로마틴이 기여한다"가 "크로마틴이 이 행렬을 신뢰할 만하게 만든다"는 뜻은 아니다. 이 행렬 수준의 움직임은 시간차 결과(Fig. 2)와 어긋나지 않는다. 같은 종류의 ATAC 셔플은 priming 마커 시간차를 bulk 셔플보다 더 흔들지 않지만(크로마틴이 시간차를 만들지 않는다), 여기서는 세포×유전자 velocity 행렬을 움직인다. 대상이 다르므로 모순이 아니다. 이 비교들을 평균 중심화 코사인으로 보고하는 이유는 원척도 값이 모든 세포에 공통인 방향에 부분적으로 좌우되기 때문이다(평균 벡터가 행 노름 제곱의 12.9~37.4%를 차지하며 arm마다 다르다). 중심화는 사후 진단이고 봉인된 지표 목록에 없으며, 위의 사전등록 판정은 원척도 지표에 기댄다. 무게가 실린 한계가 하나 더 있다. 방법 간 값을 arm 내부 불안정이 아니라 진짜 불일치로 읽는 것은 그 대조군을 가진 MultiVelo에서만 허용된다. 0에 가장 가까운 세 쌍은 모두 MoFlow가 낀 쌍인데, MoFlow는 확률적 심층 모형이고 동일 설정 재실행 안정성이 확립된 적이 없다(자기 원본 대 셔플 값 +0.113도 같은 이유로 해석 불가). 그 쌍들에서는 불일치와 불안정이 분리되지 않는다. +그 자리를 대신하는 진술은 없는 것이 아니라 경계가 그어진 것이다. 짝지은 차이(Δ = A − B)는 작고 모두 양수이지만, 그 크기는 하나의 값으로 고정되지 않는다. 셔플 draw에 따라 달라져 한 세포집합 안에서도 대략 두세 배로 흩어지므로, 어느 한 draw도 대표 크기가 아니다. draw 전반에서 Δ는 천장 +0.872의 한 자릿수 %(대략 7%) 수준으로, 같은 지표에서 method 선택이 만들어 내는 불일치(방법 쌍 간 평균 중심화 코사인 −0.530 ~ +0.131)보다 한참 아래다. 그 잔여분이 크로마틴인지 재적합인지는 처음에 가릴 수 없었다. MultiVelo에 대해 같은 세포·같은 설정의 재실행 귀무가 없었기 때문이다. 그 공백은 이후 메워졌고, 잡음이 아니라 기여 쪽으로 메워졌다. 같은 세포집합 S_b에서 크로마틴을 온전히 둔 채 worker 수만 바꿔 재적합하면 이전 적합이 그대로 재현된다. 같은 날 두 적합은 bit 단위로 동일하고, 감사 지표로 잰 실행 간 차이는 측정한 두 세포집합 모두에서 Δ_rr = 0.000000이며 세포 부트스트랩 구간도 퇴화한다(`results/velocity_matrix_runtorun_null.md`). 보관된 적합과 견주었을 때 어긋나는 것은 유전자 하나(*LRIG1*)뿐으로, 한 세포집합에서는 4×10⁻¹⁶의 반올림 차이이고 다른 하나에서는 loss가 거의 같은 대안 해(전환 시각 65.0 대 66.9)이며 세포별 평균 중심화 코사인을 최대 2×10⁻⁶ 움직인다. 우리가 쓰는 조건에서 적합이 결정론적이므로 짝지은 Δ에는 빼내야 할 재적합 잡음 성분이 없고, 양성 진술의 사전등록 조건, 곧 |Δ_rr|의 상한 구간이 Δ_paired의 하한 구간보다 낮다는 조건이 두 경우 모두에서 성립한다(0.000000 대 0.0792 및 0.0608). 따라서 MultiVelo에 대해, 그리고 이 행렬에 대해, 크로마틴 채널을 파괴하면 출력이 실제로 움직이며 그 움직임이 재적합의 인공물이 아니라고 진술한다. 이 진술을 작게 묶어 두는 한계가 셋이다. 세 쌍 중 둘만 짝맞춘 재실행 귀무를 가지며, 원래 draw가 가장 작았던 쌍은 거기 들어 있지 않다. 세포집합마다 서로 다른 셔플 추출 네 번(총 열두 번의 적합)으로 반복해 보면, 재표집을 견디는 것과 견디지 못하는 것이 나뉜다. **부호는 견딘다.** 열두 draw 전부 Δ>0이고 구간이 0을 배제한다. **크기는 견디지 못한다.** 한 세포집합에서 draw 간 산포가 중앙값을 넘고(범위 0.043 대 중앙값 0.031), 거기서 우리가 보고하던 단일 draw(+0.063)는 네 개 중 가장 큰 값이었다(`results/velocity_matrix_shuffle_seed_variability.md`). 이 산포 판정은 범위(range)를 기준으로 봉인했고, 분산 기반 산포 지표로 보면 오히려 통과한다. 어느 지표든 실질 결론, 곧 크기가 draw마다 두세 배로 흩어진다는 점은 지표와 무관하게 성립한다. 따라서 방향은 진술하되 크기는 특정 숫자로 못 박지 않는다. 세포집합당 네 draw도 그 자체로 적어 산포는 거칠게만 추정된다. 그리고 효과는 같은 지표에서 method 선택이 만드는 것보다 한 자릿수 아래에 머무르므로, "여기서 크로마틴이 기여한다"가 "크로마틴이 이 행렬을 신뢰할 만하게 만든다"는 뜻은 아니다. 이 행렬 수준의 움직임은 시간차 결과(Fig. 2)와 어긋나지 않는다. 같은 종류의 ATAC 셔플은 priming 마커 시간차를 bulk 셔플보다 더 흔들지 않지만(크로마틴이 시간차를 만들지 않는다), 여기서는 세포×유전자 velocity 행렬을 움직인다. 대상이 다르므로 모순이 아니다. 이 비교들을 평균 중심화 코사인으로 보고하는 이유는 원척도 값이 모든 세포에 공통인 방향에 부분적으로 좌우되기 때문이다(평균 벡터가 행 노름 제곱의 12.9~37.4%를 차지하며 arm마다 다르다). 중심화는 사후 진단이고 봉인된 지표 목록에 없으며, 위의 사전등록 판정은 원척도 지표에 기댄다. 무게가 실린 한계가 하나 더 있다. 방법 간 값을 arm 내부 불안정이 아니라 진짜 불일치로 읽는 것은 그 대조군을 가진 MultiVelo에서만 허용된다. 이 허용은 이후 MoFlow로도 넓어졌다. 같은 세포와 유전자 집합에서 독립적으로 두 번 재적합한 동일 설정 재실행 대조에서, MoFlow는 원래 행렬을 평균 중심화 코사인 +0.9999 [+0.9999, +0.9999]로 재현했다(실행 간 직접 비교는 +1.0000). 이 천장은 MoFlow가 낀 근접-0 세 쌍의 절댓값 최댓값인 0.0118을 훨씬 웃돈다(`results/moflow_runtorun_null.md`). 따라서 이 근접-0 값들은 MoFlow 자신의 불안정이 아니라 방법 간의 진짜 근접-0 일치로 읽을 수 있고, 같은 천장에 비추면 MoFlow의 원본 대 셔플 값(+0.113)도 재실행 잡음 수준을 훨씬 벗어난다. 그 한계 안에서 읽으면, 이 비교 전체에서 평균 중심화 일치가 가장 컸던 것은 두 multiome 방법 사이가 아니라 MultiVelo와 RNA 전용 scVelo floor 사이였다(+0.583). 다만 이는 계열 전체의 성질이 아니다. CRAK-Velo, MoFlow, MultiVeloVAE의 대응값은 각각 +0.260, −0.004, −0.292였다. MultiVelo와 MultiVeloVAE는 같은 세포에 체계적으로 반대 방향을 부여했다(평균 중심화 −0.500). 그것이 실질적 불일치인지 문서화되지 않은 부호·모수화 규약 차이인지는 이 설계로 가릴 수 없으나, 어느 쪽이든 확인 없이 한 출력을 다른 것으로 바꿔 쓰는 분석자는 반대 방향을 얻는다. 두 대조 모두 fitting된 행렬을 읽기 전에 봉인한 지표와 임계에서 실패했다(Additional file 12). multiome 쌍끼리의 일치가 RNA 전용 기준선을 넘지 못했고, 크로마틴을 파괴해도 행렬이 붕괴하지 않았다. 다만 위의 짝맞춘 비교가 보이듯 조금은 움직인다. 따라서 행렬의 방법 간 재현성도 유전자별 모수보다 나을 것이 없다. @@ -187,7 +187,7 @@ multiome velocity 방법들이 산출하는 여러 유전자별 값 가운데, ### 세포 수준 velocity 행렬 감사 -정의·제외 규칙·반증 기준은 fitting된 행렬을 읽기 전에 봉인했다. 세포와 유전자는 arm 간에 이름으로 대응시켰다. 다섯 arm은 같은 21,878개 세포를 공유하며, 다섯 arm 공통 505개 유전자는 어느 arm에서든 velocity가 결측이거나(scVelo는 dynamics를 복원하지 못한 151개 유전자를 남긴다) 상수인 유전자를 제외해 354개가 되었다. 비교는 spliced velocity로만 했고(MultiVelo·MoFlow `velo_s`, CRAK-Velo·scVelo `velocity`, MultiVeloVAE `vae_velocity`) chromatin velocity layer는 섞지 않았다. 일치도는 유전자 축에 대한 세포별 코사인 유사도이며, 세포 셔플 귀무(방법 A의 세포 *i* 대 방법 B의 무작위 치환된 세포 *j*)와 함께 보고한다. 모든 세포가 공유하는 전역 방향이 있으면 원척도 값이 부풀기 때문이다. 중앙값에는 세포 부트스트랩 95% CI를 붙인다(B=10³, seed 20260719). 원척도 코사인은 모든 세포에 공통인 방향에 부분적으로 좌우되므로(평균 벡터가 행 노름 제곱의 12.9~37.4%를 차지하며 arm마다 다르다), 세포별 편차의 일치를 재는 평균 중심화 코사인도 함께 보고한다. 중심화는 부호를 보존하도록 스케일링 없이 적용한다. (세포,유전자) 단위 부호 일치는 어느 한 arm에서든 정확히 0인 항목을 제외하며, 이는 시간차에 쓴 것과 같은 "방향 미정" 규약이다. 방법 간 값을 해석하는 기준자는 둘이다. 재표집한 세포에 대한 MultiVelo 부트스트랩 재적합 6회에서 얻은 재현성 천장, 그리고 시간차 대조에 쓴 ATAC-shuffle arm이다. 이 둘은 원래 서로 다른 세포집합에서 계산했으므로, lineage 내부 ATAC 셔플을 재적합 3개가 쓴 같은 재표집 세포집합에서 한 번 더 적합했다. 하이퍼파라미터·유전자 집합·셔플 규약을 동일하게 두고 두 arm의 세포 이름 벡터가 같음을 assert로 확인해 짝지은 비교와 짝지은 세포 부트스트랩 구간을 얻었다(`results/velocity_matrix_paired_shuffle.md`). 인과적 해석은 재적합 대조군을 가진 유일한 arm인 MultiVelo에 한정하며, 어느 arm에도 같은 세포·같은 설정의 재실행 귀무는 없었다. +정의·제외 규칙·반증 기준은 fitting된 행렬을 읽기 전에 봉인했다. 세포와 유전자는 arm 간에 이름으로 대응시켰다. 다섯 arm은 같은 21,878개 세포를 공유하며, 다섯 arm 공통 505개 유전자는 어느 arm에서든 velocity가 결측이거나(scVelo는 dynamics를 복원하지 못한 151개 유전자를 남긴다) 상수인 유전자를 제외해 354개가 되었다. 비교는 spliced velocity로만 했고(MultiVelo·MoFlow `velo_s`, CRAK-Velo·scVelo `velocity`, MultiVeloVAE `vae_velocity`) chromatin velocity layer는 섞지 않았다. 일치도는 유전자 축에 대한 세포별 코사인 유사도이며, 세포 셔플 귀무(방법 A의 세포 *i* 대 방법 B의 무작위 치환된 세포 *j*)와 함께 보고한다. 모든 세포가 공유하는 전역 방향이 있으면 원척도 값이 부풀기 때문이다. 중앙값에는 세포 부트스트랩 95% CI를 붙인다(B=10³, seed 20260719). 원척도 코사인은 모든 세포에 공통인 방향에 부분적으로 좌우되므로(평균 벡터가 행 노름 제곱의 12.9~37.4%를 차지하며 arm마다 다르다), 세포별 편차의 일치를 재는 평균 중심화 코사인도 함께 보고한다. 중심화는 부호를 보존하도록 스케일링 없이 적용한다. (세포,유전자) 단위 부호 일치는 어느 한 arm에서든 정확히 0인 항목을 제외하며, 이는 시간차에 쓴 것과 같은 "방향 미정" 규약이다. 방법 간 값을 해석하는 기준자는 둘이다. 재표집한 세포에 대한 MultiVelo 부트스트랩 재적합 6회에서 얻은 재현성 천장, 그리고 시간차 대조에 쓴 ATAC-shuffle arm이다. 이 둘은 원래 서로 다른 세포집합에서 계산했으므로, lineage 내부 ATAC 셔플을 재적합 3개가 쓴 같은 재표집 세포집합에서 한 번 더 적합했다. 하이퍼파라미터·유전자 집합·셔플 규약을 동일하게 두고 두 arm의 세포 이름 벡터가 같음을 assert로 확인해 짝지은 비교와 짝지은 세포 부트스트랩 구간을 얻었다(`results/velocity_matrix_paired_shuffle.md`). 인과적 해석은 세포 부트스트랩 재적합 대조군을 가진 유일한 arm인 MultiVelo에 한정한다. 이후 같은 세포와 유전자를 고정하고 독립적으로만 재적합하는 재실행 귀무를 MultiVelo(`results/velocity_matrix_runtorun_null.md`)와 MoFlow(`results/moflow_runtorun_null.md`) 각각에 대해 별도로 확보했다. 결과 절 참조. ### Enrichment 분석