diff --git a/docs/HARNESS.md b/docs/HARNESS.md index 0cd0a71..2c4f6d7 100644 --- a/docs/HARNESS.md +++ b/docs/HARNESS.md @@ -1,85 +1,52 @@ -# HARNESS.md — 랩 구조 (Agent 하네스 지도) — HSPC velocity-lag benchmark +# HARNESS.md — 랩 구조 (Agent 하네스 지도) — SpatialPathoAgent (BioProject02) *Designed by Ka-Kyung Kim, 2026 — a reusable paper-production harness, contributed as a scaffold (CC BY 4.0).* -이 문서는 이 프로젝트의 `.claude` 하네스를 **하나의 연구 랩**으로 본 지도다. -각 agent는 직원이 아니라 **랩의 멤버(연구원)** 이고, 사람(+메인 루프)이 랩을 이끄는 **PI**다. -운영 규칙·라우팅·산출물 계약 요약은 `CLAUDE.md`의 *Agent routing & artifact contract* 에 둔다. 이 파일은 그 확장판(멤버 명부 + 관계도 + JD)이다. +이 문서는 이 프로젝트의 `.claude` **논문 생산 하네스**를 하나의 연구 랩으로 본 지도다. 운영 규칙·라우팅·산출물 계약 요약은 `CLAUDE.md`의 *Agent routing & artifact contract*. 이 파일은 그 확장판(멤버 명부 + 관계도 + JD)이다. -- 멤버는 **누가 일을 시작할지** 사람이 매번 지정하지 않아도, CLAUDE.md 라우팅표로 자연어 요청에서 배정된다. -- 멤버는 결과를 **대화에만 남기지 않고** 산출물 계약(아래)에 따라 파일로 넘긴다. -- `paper-orchestrator`는 *계획만* 짠다. 실제 멤버 호출(실행)은 PI/메인 루프가 `paper-production-orchestrator` **Skill**로 한다 — subagent는 subagent를 못 부르기 때문. - ---- +> ⚠️ 이건 **논문 *생산* 하네스**(결과→논문·발표)다. BioProject02의 기존 **분석 파이프라인**(`agents//` 워크스페이스: data/embedding/modeling/therapeutic_evidence/critic)을 대체하지 않는다 — 그 위에 얹혀 결과를 논문으로 쓰는 레이어이고, 분석 레이어는 도메인 슬롯 `spatialpatho-analyst`가 대표한다. ## 1. 멤버 명부 (Roster) -| # | 멤버 | 소속 벤치 | 한 줄 역할 | 상태 | +| # | 멤버 | 벤치 | 한 줄 역할 | 상태 | | --- | --- | --- | --- | --- | -| D | `hspc-velocity-analyst` | 분석실 | **(도메인 슬롯)** HSPC velocity-lag 파이프라인(P0–P5)·eval·통계·cross-dataset 실행/확장, result 파일 유지 | ✅ 채움 | -| 1 | `literature-scout` | 문헌·기획 | 선행연구 탐색·정직한 포지셔닝·related work | 재사용 | -| 2 | `novelty-strategist` | 문헌·기획 | 차별화 각도 + 가장 싼 입증 실험 제안 | 재사용 | -| 3 | `research-methodologist` | 문헌·기획 | 가설·기여문·실험설계, 누수/통계 감사 | 재사용 | -| 4 | `manuscript-writer` | 집필실 | 프리프린트/저널/블로그 본문·초안 + 그림 연계 | ✅ 채움(저자/소속 FILL 잔여) | -| 5 | `presenter` | 집필실 | 슬라이드·발표·발제(청중 맞춤) | ✅ 채움(경로) | -| 6 | `paper-critic` | 심사·QA | 제출 전 적대적 자체검토 + 그림 시각 QA | 재사용 | -| 7 | `paper-orchestrator` | 코디네이션 | 멀티-agent 작업 **계획** 수립(실행은 PI) | 재사용 | -| 8 | `design` | 엔지니어링 | 로고·아이콘·브랜드·그림 미감(SVG/PNG) | 재사용 | -| 9 | `venue-reviewer` (프로젝트 로컬, 선택) | 심사·QA | 정식 venue 스타일 공식 리뷰 문서 | 선택 | -| S | 그림 생성 (스크립트) | 엔지니어링 | `figures/figNN_*.py` — 결과 파일에서 그림 생성·번호 정합 | ✅ (스크립트) | - -> ⚠️ 그림 생성은 스크립트로 둔다. `manuscript-writer`가 `pipeline/hspc-velocity-benchmark/figures/figNN_*.py`(예: `fig01_p2_concordance.py`)를 실행해 만든다. 단순 재생성은 메인 루프가 직접 돌려도 된다(결정론적). - ---- - -## 2. 관계도 (Org / collaboration chart) - +| D | `spatialpatho-analyst` | 분석실 | **(도메인 슬롯)** WSI→embedding→phenotype→therapeutic 파이프라인 대표·eval·통계, result 파일 유지 | ✅ 채움(경로 배선) | +| 1 | `literature-scout` | 문헌·기획 | 선행연구·포지셔닝·related work | 재사용 | +| 2 | `novelty-strategist` | 문헌·기획 | 차별화 각도 + 가장 싼 입증 실험 | 재사용 | +| 3 | `research-methodologist` | 문헌·기획 | 가설·기여문·설계, 누수/통계 감사 | 재사용 | +| 4 | `manuscript-writer` | 집필실 | 프리프린트/저널/블로그 본문·초안 + 그림 연계 | ✅ 채움(집필-단계 FILL 대기) | +| 5 | `presenter` | 집필실 | 슬라이드·발표(청중 맞춤) | ✅ 채움(경로 FILL 대기) | +| 6 | `paper-critic` | 심사·QA | 제출 전 적대적 자체검토 + 그림 QA (기존 `agents/critic/` 체크리스트와 병행) | 재사용 | +| 7 | `paper-orchestrator` | 코디네이션 | 멀티-agent 작업 **계획**(실행은 PI) | 재사용 | +| 8 | `design` | 엔지니어링 | 로고·아이콘·브랜드·그림 미감 | 재사용 | +| 9 | `venue-reviewer` (**프로젝트 로컬**, 선택) | 심사·QA | 정식 venue 스타일 리뷰 문서. 검증 게이트 ① 통과 후에만 호출 | ✅ 채움 (BIOP02-103) | +| S | 그림 생성 (스크립트) | 엔지니어링 | 결과 파일에서 그림 생성·번호 정합 | FILL(스크립트 지정) | + +## 2. 관계도 (일이 흐르는 표준 경로) ``` - PI = 사람 + 메인 루프 - (호출·승인·공개 게이트 책임) - │ 실행 입구 = paper-production-orchestrator (Skill) - ┌─────────┴─────────┐ - │ paper-orchestrator│ ← 계획만(실행 X) - └─────────┬─────────┘ - ┌──────────────┬──────────┼───────────────┬──────────────┐ - ▼ ▼ ▼ ▼ ▼ - 문헌·기획 분석실 집필실 심사·QA 엔지니어링 - ──────── ────── ────── ─────── ────────── - literature- hspc- manuscript- paper-critic design - scout velocity- writer venue-reviewer(선택) [그림 생성= - novelty- analyst presenter (그림 QA는 figNN_*.py, - strategist paper-critic) run by writer] - research- - methodologist +research-methodologist / literature-scout / novelty-strategist (기획·근거) + └─▶ spatialpatho-analyst ──▶ (분석·검증) + └─▶ manuscript-writer ──▶ (집필) + ║ ──▶ (그림) + └─▶ paper-critic (+ agents/critic/ 체크리스트) ──▶ reviewer (심사) + └─▶ (수정) manuscript-writer + └─▶ ──▶ presenter (검증→발표) ``` +실행 입구 = `paper-production-orchestrator` Skill(메인 루프가 실행). `paper-orchestrator`(agent)는 계획만. +- **검증 게이트**(헤드라인 숫자 재계산)와 **공개 게이트**(저자·소속·저자순서·IP·GPU 제공처)는 PI가 통과시킨다. -### 일이 흐르는 표준 경로 (논문 생산 루프) -``` -research-methodologist / literature-scout / novelty-strategist (기획·근거: paper_analysis/ 14편) - └─▶ hspc-velocity-analyst ──▶ results/FINDINGS.md + results/* (분석·검증) - └─▶ manuscript-writer ──▶ manuscript/draft_v2.md + draft_v2_ko.md (집필, 영/한 동시) - ║ figures/figNN_*.py ──▶ figures/*.png (그림) - └─▶ paper-critic ──▶ venue-reviewer ──▶ manuscript/REVIEW-*.md (심사) - └─▶ (수정 반영) manuscript-writer - └─▶ verify-gate(p3_concordance + p3_crossdataset_concordance + p3_scrambled_null) ──▶ presenter -``` -- **검증 게이트**(헤드라인 숫자 결정론적 재계산)와 **공개 게이트**(저자·소속·IP 검토)는 PI가 통과시킨다. - ---- - -## 3. 멤버별 JD 요약 -권위 있는 전체 정의는 각 `.claude/agents/.md` 본문. (분석=hspc-velocity-analyst; 집필=manuscript-writer; 발표=presenter; 검수=paper-critic; 기획=literature-scout/novelty-strategist/research-methodologist; 계획=paper-orchestrator; 디자인=design.) - ---- +## 3. 멤버별 JD +권위 있는 전체 정의는 각 `.claude/agents/.md` 본문. 분석=spatialpatho-analyst(기존 파이프라인 대표), 나머지는 재사용 연결조직. ## 4. 현재 하네스 상태 (성숙도) | 항목 | 상태 | | --- | --- | -| 멤버(agent) 정의 | ✅ 재사용 7 + 도메인 슬롯(hspc-velocity-analyst) 채움 | +| 멤버(agent) 정의 | ✅ 재사용 7 + 도메인 슬롯(spatialpatho-analyst) | | 자연어 라우팅 | ✅ CLAUDE.md 라우팅표 적용 | -| 산출물 계약 | ✅ 경로 검증(results/, manuscript/, figures/) | -| 입구(Orchestrator **Skill**) | ✅ `.claude/skills/paper-production-orchestrator/SKILL.md` | -| 검증 게이트 | ✅ `p3_concordance.py` + `p3_crossdataset_concordance.py` + `p3_scrambled_null.py` 재계산 | -| 개선 루프 | `SESSION-LOG.md`(세션별 회고 누적) | -| 미결(사람 확정) | 저자·소속·corresponding email·공개 정책 — manuscript-writer의 `` | +| 산출물 계약 | ⚠️ 경로 일부 FILL — 집필-단계 산출물(FINDINGS/manuscript/figures) 미존재(분석 진행 중) | +| 입구(Orchestrator Skill) | ✅ 설치 | +| 검증 게이트 | ⚠️ FILL — 헤드라인 AUC/AUPRC 재계산 스크립트 팀 확정 필요 | +| 미결(팀·사람 확정) | 결과요약 파일·verify-gate·headline 주장·manuscript 경로·저자순서·소속·corresponding email·GPU 제공처 | + +> **이유**: BioProject02는 분석 진행 단계(sprint 0/1). 연결 조직은 지금 설치했고, 집필-단계 FILL은 **첫 write-up-ready 결과가 나오면** 팀이 채운다. 과학적 주장·숫자는 지어내지 않는다(가정 금지). diff --git a/docs/adversarial_multi_llm_council_harness.md b/docs/adversarial_multi_llm_council_harness.md new file mode 100644 index 0000000..3c43e03 --- /dev/null +++ b/docs/adversarial_multi_llm_council_harness.md @@ -0,0 +1,437 @@ +# Adversarial Multi-LLM Council Harness +## 3 Models × 4 Independent Sessions + +## 0. 목적 + +첨부한 연구 아이디어 문서를 대상으로 다중 LLM 적대적 검토를 수행한다. + +이 작업의 목적은 아이디어를 친절하게 개선하거나 옹호하는 것이 아니다. + +목표는 다음과 같다. + +1. 아이디어의 논리적·수학적·통계적·생물학적 약점을 최대한 공격한다. +2. 서로 다른 모델과 독립 세션이 동일한 결론에 도달하는지 확인한다. +3. 각 모델이 자기 비판과 타 모델 비판을 모두 수행하게 한다. +4. 비판 자체의 오류, 과장, 문헌 누락, 환각도 다시 검증한다. +5. 최종적으로 아이디어가 폐기되어야 하는지, 수정 가능한지, 연구 가치가 남는지를 판정한다. + +각 세션은 가능한 한 독립적으로 실행한다. +사용 가능한 모델 중 각 플랫폼의 가장 높은 성능 모델을 선택한다. + +--- + +# 1. 전체 구조 + +사용할 모델은 다음 세 종류다. + +- Claude +- GPT +- Gemini + +각 모델은 서로 다른 네 개의 독립 세션을 사용한다. + +따라서 총 심사 세션 수는 다음과 같다. + +```text +Claude: 4 sessions +GPT: 4 sessions +Gemini: 4 sessions +-------------------- +Total: 12 sessions +``` + +이후 별도의 오케스트레이터 세션이 12개 결과를 종합한다. + +--- + +# 2. Stage 1 — 원본 독립 비판 + +먼저 세 모델이 원본 연구 아이디어만 받고 독립적으로 비판한다. + +```text +Claude Session 1 → Original Critique by Claude +GPT Session 1 → Original Critique by GPT +Gemini Session 1 → Original Critique by Gemini +``` + +이 단계에서는 다른 모델의 평가를 공유하지 않는다. + +각 모델은 원본 아이디어만 읽고 독립적인 적대적 리뷰를 작성한다. + +## Stage 1 산출물 + +1. `01_claude_original_critique.md` +2. `02_gpt_original_critique.md` +3. `03_gemini_original_critique.md` + +--- + +# 3. Stage 2 — 모델별 3방향 재비판 + +Stage 1의 세 원본 비판이 모두 완료되면 다음 자료를 하나의 검토 패키지로 구성한다. + +- 원본 연구 아이디어 +- Claude의 원본 비판 +- GPT의 원본 비판 +- Gemini의 원본 비판 + +이 동일한 패키지를 각 모델의 새로운 독립 세션 세 개에 제공한다. + +각 모델은 다음 세 역할을 각각 별도의 세션에서 수행한다. + +1. 자기 모델의 원본 비판을 비판하는 세션 +2. 다른 모델 A의 원본 비판을 비판하는 세션 +3. 다른 모델 B의 원본 비판을 비판하는 세션 + +같은 세션 안에서 세 리뷰를 한꺼번에 평가하지 않는다. + +--- + +## 3.1 Claude의 세 개 재비판 세션 + +```text +Claude Session 2 +→ Claude의 원본 비판을 재비판 +→ Self-Critique + +Claude Session 3 +→ GPT의 원본 비판을 비판 + +Claude Session 4 +→ Gemini의 원본 비판을 비판 +``` + +산출물: + +4. `04_claude_on_claude.md` +5. `05_claude_on_gpt.md` +6. `06_claude_on_gemini.md` + +--- + +## 3.2 GPT의 세 개 재비판 세션 + +```text +GPT Session 2 +→ GPT의 원본 비판을 재비판 +→ Self-Critique + +GPT Session 3 +→ Claude의 원본 비판을 비판 + +GPT Session 4 +→ Gemini의 원본 비판을 비판 +``` + +산출물: + +7. `07_gpt_on_gpt.md` +8. `08_gpt_on_claude.md` +9. `09_gpt_on_gemini.md` + +--- + +## 3.3 Gemini의 세 개 재비판 세션 + +```text +Gemini Session 2 +→ Gemini의 원본 비판을 재비판 +→ Self-Critique + +Gemini Session 3 +→ Claude의 원본 비판을 비판 + +Gemini Session 4 +→ GPT의 원본 비판을 비판 +``` + +산출물: + +10. `10_gemini_on_gemini.md` +11. `11_gemini_on_claude.md` +12. `12_gemini_on_gpt.md` + +--- + +# 4. 세션 독립성 규칙 + +모든 재비판은 별도의 새 세션에서 수행한다. + +각 세션은 자신에게 배정된 하나의 원본 비판만 직접 심사 대상으로 삼는다. + +다만 사실관계 확인과 문맥 파악을 위해 다음 전체 패키지는 함께 제공한다. + +- 원본 연구 아이디어 +- Claude 원본 비판 +- GPT 원본 비판 +- Gemini 원본 비판 + +각 세션에는 다음 역할을 명시한다. + +> 당신의 주된 심사 대상은 지정된 하나의 리뷰다. +> 다른 리뷰들은 비교 및 사실 검증을 위한 참고자료일 뿐이다. +> 세 리뷰를 종합하거나 최종 결론을 작성하지 말라. + +이 규칙을 통해 각 재비판의 초점이 흐려지는 것을 방지한다. + +--- + +# 5. Stage 1 원본 비판 프롬프트 + +각 모델의 첫 번째 세션에 다음 지시를 제공한다. + +```text +첨부된 연구 아이디어를 적대적 학술 심사자의 관점에서 검토하라. + +당신의 목적은 아이디어를 개선하거나 옹호하는 것이 아니라, +현재 형태의 아이디어가 왜 틀렸거나 불완전할 수 있는지 최대한 엄격하게 밝히는 것이다. + +ICML, NeurIPS, ICLR, Nature Machine Intelligence, +Bioinformatics 또는 유사 수준 학술지의 회의적인 리뷰어처럼 행동하라. + +다음을 반드시 검토하라. + +1. 핵심 주장과 실제로 검증 가능한 가설의 구분 +2. 숨겨진 전제 +3. 논리적 비약과 순환논증 +4. 수학적 정의의 부재 또는 오류 +5. uncertainty, noise, latent variable의 개념 혼동 +6. 식별 가능성 문제 +7. 통계적 confounding +8. 데이터 누수와 batch effect +9. 생물학적 타당성 +10. 알려진 pathway와 새로운 biology의 구분 가능성 +11. latent embedding을 biology로 오인할 위험 +12. 기존 연구 대비 novelty +13. 반증 가능한 실험 설계 +14. 실패할 가능성이 가장 높은 지점 +15. 논문으로서의 채택 가능성 + +모든 주장을 다음 중 하나로 분류하라. + +- Proven +- Plausible +- Speculative +- Unsupported +- Incorrect + +가능하면 관련 문헌을 확인하고 정확한 출처를 제시하라. +확인하지 못한 문헌이나 사실을 만들어내지 말라. + +친절한 표현보다 정확성을 우선하라. +비판의 강도를 낮추지 말라. +``` + +--- + +# 6. Stage 2 재비판 프롬프트 + +각 재비판 세션에는 원본 아이디어와 세 개의 Stage 1 리뷰를 제공하고, 심사할 리뷰 하나를 명시한다. + +```text +첨부 자료에는 다음이 포함되어 있다. + +1. 원본 연구 아이디어 +2. Claude의 원본 비판 +3. GPT의 원본 비판 +4. Gemini의 원본 비판 + +당신의 주된 심사 대상은 다음 리뷰다. + +[TARGET REVIEW] + +당신은 원본 연구 아이디어를 다시 심사하는 것이 아니라, +지정된 리뷰의 타당성과 품질을 적대적으로 검증해야 한다. + +다음을 확인하라. + +1. 리뷰가 원본 아이디어를 정확히 이해했는가? +2. 존재하지 않는 주장을 공격한 부분이 있는가? +3. 비판이 논리적으로 유효한가? +4. 수학적 또는 통계적 오류가 있는가? +5. uncertainty와 hidden biology를 잘못 동일시했는가? +6. 식별 가능성 문제를 정확히 설명했는가? +7. 생물학적 검증 가능성을 과소평가하거나 과대평가했는가? +8. 관련 문헌을 누락했는가? +9. 인용한 연구가 실제 주장을 지지하는가? +10. 비판 중 단순한 수사와 실제 치명적 결함을 구분했는가? +11. 지나치게 낙관적이거나 지나치게 비관적인 판단이 있는가? +12. 어떤 비판은 유효하고 어떤 비판은 폐기되어야 하는가? + +각 비판 항목을 다음 중 하나로 판정하라. + +- Valid and Fatal +- Valid but Fixable +- Partially Valid +- Weak +- Incorrect +- Hallucinated or Unsupported + +마지막에 반드시 다음을 작성하라. + +- 이 리뷰에서 유지해야 할 핵심 비판 +- 삭제해야 할 잘못된 비판 +- 추가해야 할 누락된 비판 +- 수정 후 리뷰의 최종 판정 +- 원본 아이디어에 대한 직접적인 새 종합은 하지 말 것 + +다른 두 리뷰는 비교와 사실 검증에만 사용하라. +세 리뷰 전체를 종합하지 말라. +``` + +--- + +# 7. Stage 3 — 최종 종합 + +12개의 결과가 모두 나온 후, 별도의 오케스트레이터 세션이 다음 자료를 모두 받는다. + +- 원본 연구 아이디어 +- Stage 1 원본 비판 3개 +- Stage 2 재비판 9개 + +총 13개 입력 문서를 기반으로 최종 종합을 수행한다. + +오케스트레이터는 단순 다수결을 해서는 안 된다. + +모델 수가 아니라 근거의 질, 논리적 유효성, 문헌 근거, 반증 가능성을 기준으로 판단한다. + +--- + +# 8. 최종 종합 프롬프트 + +```text +당신은 적대적 다중 LLM 학술심사 카운슬의 최종 메타 리뷰어다. + +입력에는 다음이 포함되어 있다. + +- 원본 연구 아이디어 1개 +- 독립 원본 비판 3개 +- 원본 비판에 대한 재비판 9개 + +총 12개의 심사 결과를 모두 검토하라. + +단순 요약이나 다수결을 하지 말라. + +각 주장과 비판을 증거의 질에 따라 재평가하라. + +다음을 수행하라. + +1. 모든 주요 비판 항목을 정규화하고 중복을 제거한다. +2. 세 모델이 합의한 비판을 식별한다. +3. 모델 간 의견이 갈린 비판을 식별한다. +4. 재비판 과정에서 폐기된 비판을 식별한다. +5. 자기비판에서만 발견된 오류를 식별한다. +6. 타 모델 비판에서만 발견된 오류를 식별한다. +7. 문헌으로 확인된 사실과 추측을 분리한다. +8. 아이디어의 치명적 결함과 수정 가능한 결함을 분리한다. +9. 현재 데이터로 검증 가능한 부분과 불가능한 부분을 분리한다. +10. 최소한의 반증 실험을 제안한다. +11. 연구를 계속할 가치가 있는지 냉정하게 판정한다. + +최종 판정은 다음 중 하나여야 한다. + +- Reject: Fundamentally Invalid +- Reject: Not Identifiable with Available Data +- Major Revision: Plausible but Severely Underspecified +- Conditional Go: Worth Testing under Strict Conditions +- Strong Go: Clear and Defensible Research Direction + +최종 문서에는 반드시 다음 섹션을 포함하라. + +- Executive Verdict +- Reconstructed Core Hypothesis +- Claims That Survived +- Claims That Failed +- Fatal Problems +- Fixable Problems +- Reviewer Disagreements +- Literature Gaps +- Minimum Falsification Experiments +- Required Mathematical Formalization +- Required Biological Validation +- Publication Potential +- Final Recommendation +``` + +--- + +# 9. 전체 산출물 + +## 원본 비판 3개 + +1. `01_claude_original_critique.md` +2. `02_gpt_original_critique.md` +3. `03_gemini_original_critique.md` + +## 재비판 9개 + +4. `04_claude_on_claude.md` +5. `05_claude_on_gpt.md` +6. `06_claude_on_gemini.md` +7. `07_gpt_on_gpt.md` +8. `08_gpt_on_claude.md` +9. `09_gpt_on_gemini.md` +10. `10_gemini_on_gemini.md` +11. `11_gemini_on_claude.md` +12. `12_gemini_on_gpt.md` + +## 최종 종합 + +13. `13_final_meta_review.md` + +## 선택 산출물 + +14. `14_revised_hypothesis.md` +15. `15_experiment_blueprint.md` +16. `16_claim_evidence_matrix.md` + +선택 산출물은 최종 메타 리뷰가 완료된 후에만 작성한다. +비판 단계에서 원본 아이디어를 임의로 개선하거나 다시 쓰지 않는다. + +--- + +# 10. 실행 원칙 + +- 각 세션에는 가능한 최고 성능 모델을 사용한다. +- 서로 다른 세션의 숨은 추론을 공유하지 않는다. +- 결과 문서만 다음 단계의 입력으로 전달한다. +- 모든 세션은 원본 아이디어를 직접 확인할 수 있어야 한다. +- 인용은 실제로 확인 가능한 문헌만 사용한다. +- 불확실한 내용은 불확실하다고 표시한다. +- 동일 모델의 자기비판과 타 모델 비판을 분리한다. +- 오케스트레이터는 모델 이름을 권위의 근거로 사용하지 않는다. +- 합의가 곧 진실이라는 가정을 금지한다. +- 비판의 강도보다 비판의 유효성을 우선한다. + +--- + +# 11. 구조 요약 + +```text +Original Research Idea + │ + ├── Claude Session 1 ── Claude Original Critique + ├── GPT Session 1 ───── GPT Original Critique + └── Gemini Session 1 ── Gemini Original Critique + │ + ▼ + Original + Three Original Critiques + │ + ┌─────────────────┼─────────────────┐ + │ │ │ + ▼ ▼ ▼ + Claude Sessions 2–4 GPT Sessions 2–4 Gemini Sessions 2–4 + │ │ │ + ├─ on Claude ├─ on GPT ├─ on Gemini + ├─ on GPT ├─ on Claude ├─ on Claude + └─ on Gemini └─ on Gemini └─ on GPT + │ + ▼ + 12 Review Documents + │ + ▼ + Final Meta Reviewer + │ + ▼ + 13_final_meta_review.md +``` diff --git a/docs/hyperresearchdeck.html b/docs/hyperresearchdeck.html new file mode 100644 index 0000000..3606f3f --- /dev/null +++ b/docs/hyperresearchdeck.html @@ -0,0 +1,536 @@ + + + + + +hyperresearch — 구조와 동작 원리 + + + +
+
+ + +
+
github.com/jordan-gibbs/hyperresearch · MIT · v0.9.1
+

hyperresearch
딥 리서치 하네스의 구조와 동작 원리

+
+ Claude Code를 "에이전트 파이프라인"으로 바꾸는 Python CLI + 스킬 팩.
+ 프롬프트 한 줄 → 16단계를 거쳐 출처가 검증된 리서치 리포트가 나옵니다. +
+
+
16
단계 파이프라인
(스킬 = 1단계)
+
16종
서브에이전트
(Sonnet / Opus 혼합)
+
55–450
1회 실행당 수집 소스
(기어 · 티어에 따라)
+
Vault에 영구 축적
(다음 실행이 재사용)
+
+
발표자료 · 2026-07 기준 main 브랜치(183443a) 코드/문서 직접 분석
+
+ + +
+
01 · 문제 정의
+

왜 또 하나의 딥 리서치 도구인가

+
기존 "Deep Research" 제품들이 공통적으로 놓치는 4가지를 정면으로 공격합니다.
+
+
+
+

일회성(one-shot)

+

리포트 하나 뽑고 읽은 자료는 전부 버림. 같은 주제를 또 물으면 처음부터 다시 크롤링.

+
+
+

Vault에 영구 저장

+

읽은 모든 소스가 마크다운 + SQLite 인덱스로 남고, 다음 세션은 fetch 전에 먼저 검색합니다.

+
+
+

인용 환각 / 유령 각주

+

그럴듯한 문장에 아무 링크나 붙는 문제. 실제로 그 소스가 그 문장을 뒷받침하는지 아무도 검증 안 함.

+
+
+

cite-check + quote-integrity

+

인용-문장 바인딩을 회의적으로 재검증하고, 인용부호 안 텍스트가 Vault 노트에 글자 그대로 없으면 출고 차단.

+
+
+

신디케이션 = 합의로 착각

+

보도자료 하나를 5개 매체가 받아쓴 걸 "5개 출처가 동의"로 계산.

+
+
+

독립성 감사

+

파생·복제본을 클러스터링해서 5부 = 1표로 눌러버립니다. 게다가 철회(retraction)된 논문은 품질 점수 0으로 바닥.

+
+
+
저자 주장: DeepResearch-Bench RACE 리더보드 1위 (내부 벤치마크, 제3자 검증 대기 중) — 발표 시 "자체 측정치"로 인용 권장.
+
+
+ + +
+
02 · 사용법
+

개발자 입장에서 실제로 뭘 하는가

+
pip 패키지 하나를 깔면 Claude Code 프로젝트에 스킬 · 서브에이전트 정의 · CLI가 주입됩니다.
+
+
# 1. 설치 — 벌트 초기화 + .claude/ 에 스킬/에이전트 주입 + CLAUDE.md 갱신
+cd your-project
+pip install hyperresearch && hyperresearch install      # --global 도 가능
+
+# 2. Claude Code 안에서 슬래시 커맨드 한 줄
+/hyperresearch 국내 배터리 3사의 전고체 로드맵과 실현 가능성을 비판적으로 분석해줘
+
+# 3. 실행 중/후 — 전부 CLI로 관측 가능 (모든 명령이 -j JSON 출력 지원)
+hyperresearch run status -j      # 단계별 진행 · 지출 · 에스컬레이션 큐 깊이
+hyperresearch run resume -j      # 죽은 지점의 정확한 다음 Skill 호출을 반환
+hyperresearch search "전고체" -j  # 벌트 전문 검색 (--semantic / --ranked)
+hyperresearch serve              # 벌트를 로컬 위키로 브라우징
+
+

런타임

Python 3.11–3.13 · typer/pydantic/jinja2 · Crawl4AI(크롤링) · pymupdf(PDF 본문 추출)

+

필수 전제

Claude Code. 모델 호출은 전부 Claude Code의 서브에이전트를 통해 나감 — 별도 LLM API 키 불필요

+

선택 확장

exa tavily 검색 프로바이더, mcp 서버 모드, watch, Voyage/OpenAI 임베딩

+
+
+
+ + +
+
03 · 아키텍처
+

3층 구조: CLI · 스킬 체인 · Vault

+
LLM이 하는 일과 결정론적 코드가 하는 일을 칼같이 분리한 게 이 프로젝트의 뼈대입니다.
+
+
+
+
LAYER 1
Python CLI
+
벌트 init · fetch · sync
SQLite 인덱스 · lint 게이트
run 매니페스트 · 프로파일 렌더
+
+
+
+
LAYER 2
스킬 체인 (17개 .md)
+
엔트리 스킬 = 라우터
단계별 절차서를 필요한 순간에만
컨텍스트로 로드
+
+
+
+
LAYER 3
서브에이전트 16종
+
fetcher · 분석가 · 드래프터
비평가 · 패처 · 인용검증기
모델은 프로파일 설정값
+
+
+
+
STATE
Vault + Run
+
research/notes/ 마크다운
research/runs/<tag>/ 작업공간
SQLite = 재생성 가능한 캐시
+
+
+
+
+

결정론적 코드가 담당

+

URL 중복 제거 · PDF 파싱 · 전문 검색 · PageRank · 인용수/철회 여부 조회(OpenAlex·Semantic Scholar) · 스키마 CHECK 제약 · 린트 규칙 · 예산 상한 · 재개 지점 계산. LLM에게 맡기면 흔들리는 것들을 전부 코드로 내림.

+
+
+

LLM이 담당

+

질의 분해 · 검색 관점 설계 · 소스 요약 · 모순 짝짓기 · 논점(loci) 선정 · 초안 집필 · 적대적 비평 · 외과적 패치. 판단이 필요한 곳에만 모델을 씁니다.

+
+
+
+
+ + +
+
04 · 파이프라인
+

16단계 — 폭 → 깊이 → 초안 → 적대적 감사

+
전 티어 full 티어 합성 감사/차단  각 단계는 독립된 스킬 파일이며 순서대로 Skill() 호출로 실행됩니다.
+
+
+
1
Decompose
질의→원자 항목·커버리지 행렬·티어 분류
+
1.5
Chapter partition
4–10개 챕터로 분할 (논문 티어)
+
2
Width sweep
다관점 검색 → 병렬 fetcher 웨이브
+
3
Contradiction graph
코퍼스 전반의 모순을 랭크된 클러스터로
+
4
Loci analysis
분석가 2인 병렬 → 논점 + 소스 예산
+
5
Depth investigation
K명 병렬 심층조사 → 입장을 확정한 중간노트
+
6
Cross-locus reconcile
확정 입장들 간 조정 → comparisons.md
+
7
Source tensions
전문가 간 이견 추출 → JSON
+
8
Corpus critic
"뭘 뒤집을 소스가 빠졌나" + 표적 보충
+
9
Evidence digest
핵심 주장 + 축자 인용 정리
+
10
Triple draft
관점별 소스 큐레이션 → 초안 3개 병렬
+
11
Synthesize
3초안 통합 → final_report.md (2패스 집필)
+
12
Critics ×4
변증·심도·폭·지시준수 비평가 병렬 공격
+
13
Gap-fetch
비평가가 지목한 공백만 표적 수집
+
14
Patcher
Read+Edit 락 · 외과적 hunk만 적용
+
14.5
Cite-check
인용-문장 결합 검증 + 2차 패치
+
15
Polish
군더더기·위생 누출 제거 (Edit 락)
+
16
Readability audit
가독성 제안 JSON → 선별 적용
+
+
1–2단계에서 을, 3–9단계에서 깊이와 대립을, 10–11에서 을, 12–16에서 공격과 수리를 합니다. 순서가 곧 방법론입니다.
+
+
+ + +
+
05 · 핵심 설계 ①
+

왜 스킬을 17개로 쪼갰나 — 컨텍스트 부패 방어

+
이 프로젝트에서 가장 배울 만한 부분. 저자가 겪은 실패에서 직접 나온 설계입니다.
+
+
+ "V7은 1200줄짜리 스킬 하나였다. Layer 4가 triple-draft 절차를 필요로 할 때쯤이면 그 부분은 이미 컴팩션으로 날아가 있었다. + 오케스트레이터는 절차를 잊었고, 초안을 하나만 썼고, 밋밋한 리포트가 나왔다." +
— hyperresearch.md (엔트리 스킬) 원문 요약
+
+
+
+

안티패턴: 모놀리식 프롬프트

+

긴 실행 = 긴 대화. 대화가 길어지면 앞부분이 압축·소실되고, 후반 단계의 절차가 조용히 사라집니다. 에러도 안 나고, 그냥 대충 합니다.

+
+
+

해법: 라우터 + 지연 로딩

+

엔트리 스킬은 순서만 압니다. 각 단계의 절차서는 그 단계를 실행하는 순간 Skill(...)으로 새로 로드됩니다. 축출 리스크 0.

+
+
+
+

내구성 있는 기억

단계 전환마다 run step N --status로 매니페스트에 기록. 대화가 아니라 파일이 상태입니다.

+

정본 질의(gospel)

사용자 원문을 query.md에 글자 그대로 1회 저장 → 모든 단계·모든 서브에이전트가 경로로 재읽기. 전언 게임 방지.

+

서브에이전트 = 컨텍스트 격리

fetcher 8–12개가 병렬로 도는 동안 오케스트레이터 컨텍스트는 깨끗하게 유지됩니다.

+
+
+
+ + +
+
06 · 스케일
+

티어 × 기어 — 직교하는 두 개의 레버

+
"어떤 단계를 도느냐(티어)"와 "얼마나 크게 도느냐(기어)"를 분리했습니다.
+
+
+
+

티어 — 질의별 자동 라우팅

+ + + + +
light범위가 닫힌 사실 질의·비교·서베이
1→2→10→15→16 · ~30–40분 · 15–25 소스
full기본값. 적대적 감사 포함 전 16단계
~1.5–2.5시간 · 55–80 소스
dissertation명시 요청 시에만. 4–10 챕터 루프
~4–8시간 · 300–450 소스 · 2.5–8만 단어
+

1단계가 분류하며, 스킬 문서는 "티어 게이트를 존중하라 — 꼼꼼히 하겠다고 건너뛴 단계를 되살리지 말라"고 명시적으로 금지합니다.

+
+
+

기어 — 프로젝트별 수동 설정

+ + + +
full기준선. 55–80 소스
premier100–130 소스, 심층 예산 2배, ~3–5시간
+
hyperresearch profile use premier
+# .hyperresearch/config.toml
+[profile.myown]
+extends = "full"
+source_min = 60
+models = { fetcher = "haiku" }
+

기어는 설치 시 스킬 프롬프트에 숫자로 렌더링(Jinja2)됩니다. 그래서 다음 실행부터 적용되고 실행 중간에는 절대 안 바뀝니다.

+
+
+
프로파일은 pydantic 모델로 엄격 검증(extra=forbid, frozen). 소스 목표·fetcher 팬아웃·논점 상한·초안 개수·단어 목표·인용 밀도 하한·에이전트별 모델까지 전부 하나의 설정 객체입니다.
+
+
+ + +
+
07 · 에이전트
+

서브에이전트 로스터 — 역할별 모델 배분

+
"싼 모델로 많이 읽고, 비싼 모델로 판단한다." 모델은 하드코딩이 아니라 프로파일 설정입니다.
+
+
+
+
SONNET — 수집 · 읽기 · 검증 (양)
+ + + + + + + + +
fetchercrawl4ai로 URL 수집. 웨이브당 8–12개 병렬
source-analyst5,000단어 넘는 장문 소스 1건 전담 다이제스트
loci-analyst폭 코퍼스를 읽고 1–8개 심층 논점 제안
depth-investigator논점 1개 조사 → 입장을 확정한 중간노트
corpus-critic초안 전 "무엇이 이 방향을 뒤집나" 공백 분석
cite-checker표본 인용-문장 결합을 회의적으로 재검증
browser-fetcher실제 로그인된 Chrome을 몰아 에스컬레이션 큐 처리
+
+
+
OPUS — 집필 · 비평 · 수리 (질)
+ + + + + + + + + +
draft-orchestrator관점 1개당 1명. 큐레이션된 소스로 초안 집필
synthesizer3개 초안을 읽고 최종본 집필 (Read+Write 락)
dialectic-critic초안이 놓친 반대 증거
depth-critic중간노트로 메울 수 있었던 얕은 지점
width-critic코퍼스는 받쳐주는데 초안이 무시한 주제 구석
instruction-critic원 프롬프트의 원자 항목 대비 구조적 불일치
patcherRead+Edit 락 비평 결과를 외과적 hunk로만 적용
polish-auditorRead+Edit 락 군더더기·위생 누출 제거
+
+
+
모델 배분은 [profile.x] models = {...} 한 줄로 통째 교체 가능 — 비용 실험이 설정 변경 수준으로 떨어집니다.
+
+
+ + +
+
08 · Vault
+

"마크다운이 진실, SQLite는 캐시"

+
읽은 것을 버리지 않는 구조. 세션을 거듭할수록 시작 지점이 앞당겨집니다.
+
+
+
+
research/
+├─ notes/        # 소스 1건 = 마크다운 1개 (YAML 프론트매터)
+├─ raw/          # 원본 PDF (note-id.pdf)
+├─ index/        # 자동 생성 인덱스
+└─ runs/<tag>/   # 실행별 격리 작업공간
+    ├─ query.md          # 정본 질의 (gospel)
+    ├─ run.json          # 매니페스트 = 내구성 메모리
+    └─ final_report.md
+.hyperresearch/index.db  # 지워도 sync 로 완전 재생성
+
+

왜 이게 중요한가

+

도구 없이도 내 리서치를 읽을 수 있습니다. 아무 에디터로 열고, git으로 버전 관리하고, 인덱스가 깨지면 hyperresearch sync 한 번. 락인 없음.

+
+
+
+

스키마가 강제된 노트

+

tier(ground_truth / institutional / practitioner / commentary)와 content_type(paper / docs / policy / dataset …)은 SQLite CHECK 제약으로 어휘가 고정. 프론트매터가 오염돼도 인덱스는 안 썩습니다.

+

출처의 출처(provenance)

+

모든 fetch는 --suggested-by로 "누가 이걸 물어왔는지"를 기록 → 시드에서 뻗은 루트 트리 형성. 끊긴 컴포넌트는 린트가 잡습니다.

+

검색 3종

+

전문 검색(FTS) · --include-body 본문 검색 · --semantic 임베딩 혼합(voyage/openai, 기본은 키 불필요한 none). 그래프 hubs/backlinks도 제공.

+
+
+
+
+ + +
+
09 · 소스 품질
+

품질을 "느낌"이 아니라 지속되는 점수

+
모든 소스는 4가지 신호가 합성된 quality_score를 갖고, 이 점수는 Vault에 남습니다.
+
+
+

소스 유형 티어

원 데이터 / 기관 / 실무자 / 논평 — 증거로서의 역할을 분류

+

수집 시점 유용성

fetcher가 실제로 읽고 매긴 utility 스코어

+

인용 권위

OpenAlex · Semantic Scholar에서 인용수·게재지·철회 플래그 조회

+

벌트 PageRank

링크 + provenance 그래프 상의 중심성

+
+
+
+

독립성 감사 — 5부 = 1표

+

sources independence가 신디케이트·파생 복제본을 클러스터링합니다. 같은 보도자료를 받아쓴 5개 기사는 증거 무게 1개분으로만 계산됩니다. 2단계의 "증거 중복 감사"에서 특정 항목의 독립 소스가 2개 미만으로 떨어지면 웨이브 3 표적 재수집이 발동합니다.

+
+
+

철회 논문 — 출고 차단 사유

+

철회된 소스는 품질 점수가 0에 수렴하도록 강제되고, 출고 직전 인용된 모든 DOI를 다시 조회합니다. 어제 철회된 논문이 오늘 잡힙니다. 옛날 실행에서 재사용한 벌트 소스도 예외 없이. 철회를 명시하지 않고 인용하면 하드 에러.

+
+
+
학술 주제는 웹 검색보다 학술 API를 먼저 칩니다 (Semantic Scholar · arXiv · OpenAlex · PubMed). 웹은 파생 논평을 주고, 학술 API는 인용 랭킹된 정본을 줍니다. 그 뒤에 맥락·뉴스·"X에 대한 비판"류 적대적 검색을 최소 1회 이상 반드시 수행.
+
+
+ + +
+
10 · 핵심 설계 ②
+

"고치되, 다시 쓰지 마라" — 도구 권한으로 강제

+
프롬프트로 부탁하는 대신, 물리적으로 불가능하게 만든 것이 이 프로젝트의 두 번째 큰 아이디어입니다.
+
+
+
STEP 11
최종본 생성
여기까지만 "쓰기" 허용
+
+
STEP 12
비평가 4인 병렬 공격
findings JSON 산출
+
+
STEP 13
공백 표적 수집
비평이 지목한 곳만
+
+
STEP 14 · 15 · 16
🔒 Read + Edit 만
Write 도구 자체가 없음
+
+
+
+

왜 락을 거나

+

LLM에게 "조금만 고쳐"라고 하면 결국 통째로 다시 씁니다. 그 과정에서 검증이 끝난 인용과 수치가 조용히 바뀝니다. Claude Code의 도구 허용목록 수준에서 Write를 제거하면 이 실패 모드가 원천 차단됩니다. hunk당 크기 상한까지 걸려 "그냥 다시 쓰기"가 기계적으로 불가능합니다.

+
+
+

부작용까지 설계됨

+

패처는 Write를 못 하므로 로그 파일을 만들 수도 없습니다. 그래서 오케스트레이터가 빈 스텁 파일을 먼저 써두고 패처가 Edit으로 채우게 합니다. 작은 hunk에 안 들어가는 비평은 "구조적 이슈"로 에스컬레이션되며, 미적용 CRITICAL은 patch-surgery 린트가 전부 드러냅니다.

+
+
+
+

quote-integrity

인용부호 안 문자열이 벌트 노트에 축자로 없으면 차단

+

retracted-citations

철회 소스 무고지 인용 = 하드 에러

+

numeric-consistency

증거로 추적 안 되는 숫자 플래그

+

scaffold-prompt · locus 커버리지

원문 프롬프트 누락, 조사 안 된 논점 에러

+
+
+
+ + +
+
11 · 운영
+

장시간 실행을 실무에서 굴리는 장치들

+
2–8시간짜리 에이전트 작업은 반드시 죽습니다. 죽는 걸 전제로 설계돼 있습니다.
+
+
+

재개 가능

실행마다 매니페스트를 남기고 run resume죽은 그 단계의 정확한 Skill 호출을 되돌려줍니다. 실행 작업공간은 태그로 격리돼 동시 실행끼리 충돌하지 않습니다.

+

예산 상한

run init --budget 50으로 추정 지출 한도를 걸면, 넘는 순간 조용히 불어나는 대신 실행이 막힙니다. run report는 단계별 벽시계·지출·소스 수율 텔레메트리.

+

출고 게이트

run verify가 제목 구조·분량·인용 밀도·cite-check 해소 여부를 최종 점검. 통과 못 하면 리포트는 나가지 않습니다.

+
+
+
+

브라우저 레인 — 막힌 fetch는 죽지 않고 줄을 섭니다

+

hyperresearch setup으로 한 번 로그인해두면 인증 크롤링이 가능합니다(LinkedIn·X·Facebook·Instagram·TikTok은 세션 킬 방지를 위해 보이는 브라우저 강제). 헤드리스가 로그인 벽/봇 벽에 막히면 해당 URL은 에스컬레이션 큐로 들어가고, Claude-in-Chrome 확장이 있으면 browser-fetcher 에이전트가 실제 로그인된 내 크롬을 몰아 큐를 비웁니다. 확장이 없으면 그냥 쌓이고 웨이브 요약에 개수만 보고 — 기존보다 나빠지지 않음.

+

단호한 경계선: CAPTCHA · 2FA · 로그인은 절대 자동으로 풀지 않습니다. 한 통의 메시지로 모아서 사람에게 넘깁니다.

+
+
+

수집 웨이브의 실제 동작

+

웨이브 1 — fetcher 8–12개를 한 메시지에 스폰(진짜 병렬), 겹치지 않는 배치 배분. 위키피디아는 인용 가능한 소스가 아니라 "소스 허브"로만 취급해 참고문헌을 뽑아 다음 웨이브로 넘깁니다.
+ 웨이브 2 — 커버리지 점검 후 얇거나 빈 항목만 표적 보충(20–40 URL).
+ 웨이브 3 — 독립 소스가 2개 미만인 항목에 대해서만 조건부 발동.

+
+
+
+
+ + +
+
12 · 정리
+

가져갈 것 · 조심할 것

+
+
+
+
에이전트를 만든다면 훔칠 만한 패턴
+
    +
  • 절차서를 지연 로딩하라. 긴 파이프라인의 진짜 적은 모델 성능이 아니라 컨텍스트 부패입니다.
  • +
  • 상태는 파일에, 대화에 두지 마라. 매니페스트 + 정본 질의 파일 = 재개 가능성과 일관성을 동시에.
  • +
  • 부탁하지 말고 도구를 빼앗아라. "다시 쓰지 마"는 프롬프트가 아니라 허용목록으로 강제.
  • +
  • 검증은 결정론적 코드로. 축자 인용 대조, 철회 조회, 스키마 CHECK — LLM 판단에 맡기지 않음.
  • +
  • 적대적으로 설계하라. 비평가를 병렬로 붙이고, 수집 단계에서도 "비판" 검색을 의무화.
  • +
  • 스케일 노브를 하나의 검증된 설정 객체로. 비용/품질 실험이 config 한 줄로 떨어집니다.
  • +
+
+
+
한계 (저자가 직접 명시)
+
    +
  • 사실 정확성은 보장 못 합니다. 린트 게이트가 잡는 건 구조적 실패(스캐폴드 누락, 끊긴 provenance, 미해소 CRITICAL)뿐.
  • +
  • 어떤 소스가 중요한지는 여전히 사람 판단. 에이전트가 고르고, 사람이 방향을 잡습니다.
  • +
  • 로그인 안 한 페이월은 못 뚫습니다.
  • +
  • Anthropic 모델 전용. Claude Code 서브에이전트 로스터에 묶여 있고, 사용량은 티어·기어·코퍼스 크기에 비례해 커집니다.
  • +
  • 벤치마크 1위 주장은 자체 측정. 제3자 검증 대기 중이고 README도 "전망치(projection)"라고 씀 — 인용 시 반드시 단서 필요.
  • +
  • v0.9.1 · Alpha. Python 3.14 미지원.
  • +
+
+
+
한 문장 요약 — hyperresearch는 "더 똑똑한 프롬프트"가 아니라, 에이전트가 실패하는 방식들을 하나씩 구조로 막아놓은 하네스다.
+
+
+ +
+ +
+ 1 / 13 + ← → 또는 Space 로 이동 · F 전체화면 +
+ + + + diff --git a/docs/pilot_external_skills_2026-07-27.md b/docs/pilot_external_skills_2026-07-27.md new file mode 100644 index 0000000..6005665 --- /dev/null +++ b/docs/pilot_external_skills_2026-07-27.md @@ -0,0 +1,187 @@ +# 파일럿: 외부 Agent Skill 재조사 — hyperresearch · K-Dense (2026-07-27) + +> **목적:** `jordan-gibbs/hyperresearch` 와 `K-Dense-AI/scientific-agent-skills` 두 저장소에 우리가 쓸 것이 있는지 판정한다. +> **판정 기준:** 새로 만들지 않았다. 팀이 2026-07-17~18에 정한 것을 그대로 쓴다 — `blog/2026-07-18_BIOP02_09_skill-benchmark-no-borrow.md`, `docs/HARNESS_REVIEW_2026-07-17.md` §4.5. +> **결론 선요약:** **기존 "차용 없음" 유지.** 지금 설치할 것은 없다. 다만 미조사 스코프 하나와, 어느 저장소도 안 잡는 새 실패 범주 하나가 드러났다. +> **작성:** 이건규. Claude·GPT·Gemini 3모델 교차검증(같은 사실 카드, 다른 질문). 합의·불일치를 §6에 그대로 적었다. + +--- + +## 0. 판정 기준 (인용, 새로 만들지 않음) + +> "별점도, 공식이라는 표시도, 도구가 스스로 매긴 점수도 아니다. **그 도구가 우리가 실제로 겪은 실패를 잡아내는지, 그리고 다시 돌려도 같은 답을 내는지. 이 두 가지뿐이다.**" +> +> "표준적인 기계 작업은 빌려 올 수 있지만, **'이게 진짜 맞는가'를 가리는 관문만큼은 우리 손으로** 지어야 한다." + +이 문서는 위 기준을 뒤집지 않는다. 적용할 뿐이다. + +--- + +## 1. 실패셋 (합격선) — 7건 → **11건으로 갱신** + +`docs/HARNESS_REVIEW_2026-07-17.md` §1.2 의 7건에, 2026-07-27 에 드러난 4건을 더했다. + +| # | 실패 | 성질 | +|---|---|---| +| 1 | csv `\r` 파일명 → openslide "missing" | 파이프라인·조용한 실패 | +| 2 | bash 자기참조 → `set -u` unbound → 다운로드 성공 후 임베딩 침묵 사망 | 파이프라인·조용한 실패 | +| 3 | detached shell 에 conda 없음 → 임베딩 실패 | 파이프라인 | +| 4 | n=187 vs n=85 혼동 | 문서·주장 | +| 5 | 인용 오류 5건 — 존재하지 않는 "Williams 2022" | 문서·주장 | +| 6 | "523 slides" vs "523 cases" 단위 혼동 | 문서·주장 | +| 7 | Virchow2 HF 캐시가 `.incomplete` 인데 `du` 크기만 보고 "완료" 보고 | 파이프라인 | +| **8** | **원고 R2 가 "공개 코호트는 원리적으로 검정력 부족"이라 주장하나, 코호트 전체 양성이 평가가 본 것의 3~5배**(위 MSI 82 vs 24). 단일 70/15/15 분할의 결과지 코호트 한계가 아님 | **상태 불일치** | +| **9** | **정본 스코어보드는 "위암 endpoint 전체 저신뢰", 원고 R4 는 "Lauren 만."** 정본이 요청한 진단이 나왔는데 정본 미갱신 | **상태 불일치** | +| **10** | **논문이 몇 편인지 리더 결정·`CLAUDE.md`·Jira 스프린트가 각각 다르게 말함**(2주째) | **상태 불일치** | +| **11** | **원고가 인용한 커밋(`a693984`)이 브랜치 HEAD 가 아니어서 stale 내용 기준으로 검토가 진행됨** | **상태 불일치** | + +--- + +## 2. 조사 범위 + +| 대상 | 규모 | 방법 | 기존 조사 여부 | +|---|---|---|---| +| `jordan-gibbs/hyperresearch` | 196 파일, `src/hyperresearch/` 12모듈 | 로컬 클론 + 소스 정독 | **신규** (2026-07-17 목록에 없음) | +| `K-Dense-AI/scientific-agent-skills` | **154** SKILL.md (당시 149) | 로컬 클론 + 스킬 목록 census | **기조사·기각** (병리 스코프) | + +클론만 했다. **설치·`pip install` 은 하지 않았다** — `/opt/envs/spatialpatho` 격리 규율 때문. + +--- + +## 3. K-Dense — 이미 기각됐다. 단 스코프가 한정적이었다 + +`docs/pilot_pathology_skills_2026-07-17.md` 판정을 재확인했고 **유효하다**: + +> "본진(병리)에 차용할 스킬은 없다. 후보 5종 전부 **실행코드 0줄**이고, 우리 실패 6건 중 **0건**을 잡는다." + +`skills/histolab`·`skills/pathml` 은 기존 pip 라이브러리 사용설명서다. 이 판정은 그대로 둔다. + +### 3.1 그런데 조사 스코프가 세 영역뿐이었다 + +기존 조사는 **(a) 병리·이미지 본진 (b) 논문 집필 (c) 단일세포** 만 봤다. BIOP02 의 실제 작업 층은 더 넓다. + +| 미조사 층 | 우리가 실제로 하는 일 | K-Dense 대응 후보 | +|---|---|---| +| **데이터 조달** | `guide/runbooks/download_cptac_from_idc.md`, TCGA/CPTAC WSI 다운로드 | `imaging-data-commons` | +| **치료증거 연결** | jhans Therapeutic Evidence Agent, DepMap/GDSC | `depmap` | +| 라벨·유전체 메타데이터 정합 | 라벨 추출·QC (jamie) | `genomic-coordinates`, `anndata` | + +**이 층들은 기각된 적이 없다. 조사된 적이 없다.** 154개로 5개 늘어난 것 때문이 아니라, 이 스코프 구멍 때문에 좁은 재조사가 정당하다. + +> ⚠️ 단, "미조사"는 "유망"이 아니다. 병리 스코프에서 나온 패턴(설명서만 있고 실행코드 0줄)이 여기서도 반복될 가능성이 높다. 실패셋 대조를 거치기 전에는 후보로도 세지 않는다. + +--- + +## 4. hyperresearch — 실행코드는 실재한다. 그러나 우리 원고엔 못 쓴다 + +MIT. Python 3.11–3.13. `src/hyperresearch/` 아래 12모듈(cli·core·graph·search·web·mcp·export 등). **실행코드가 실재한다** — 이 점에서 K-Dense 병리 스킬과 다르다. + +### 4.1 cite-check 를 코드로 확인했다 + +광고 문구가 아니라 구현을 읽었다. `src/hyperresearch/core/citecheck.py` (214줄): + +```python +# parse_sources_section: Map `[N]` -> note_id by matching Sources-section URLs/titles to the vault +row = conn.execute("SELECT note_id FROM sources WHERE url = ?", (url,)).fetchone() + +# triage_pairs 의 판정 범주 +# dangling — citation resolves to no vault note (finding) + +def sample_needs_llm(pairs, sample_rate: float = 0.6): ... +``` + +읽어낸 것: + +1. **대조 대상이 자기 vault SQLite 다.** 그 run 에서 hyperresearch 가 직접 받아온 소스하고만 맞춘다. Crossref·PubMed·DOI 해석이 아니다. +2. 그래서 **vault 가 없는 문서는 검사 자체가 불가능하다.** 우리 원고(`manuscript/sections/*.md`)는 우리 에이전트가 썼고 hyperresearch vault 가 없다. **실패 #5(존재하지 않는 Williams 2022)를 우리 원고에서 잡을 수 없다.** +3. LLM 스팟체크는 **60% 표본**(`sample_rate=0.6`)이다. 전수가 아니다. + +즉 이 기능이 유효한 범위는 "hyperresearch 가 처음부터 끝까지 생산한 리포트"뿐이다. 우리 워크플로 전체를 이 도구로 옮기지 않는 한 쓸 수 없다. + +### 4.2 벤치마크 주장은 근거로 쓰지 않는다 + +README 가 "DeepResearch-Bench RACE 리더보드 1위"라며 그래프를 싣는다. 그 그림의 캡션 원문: + +> "**Forward-looking projection** from a stratified pilot against the leaderboard snapshot. **Third party validation is pending.**" + +측정치가 아니라 투영이고 3자 검증 전이다. 본문도 "benchmarked internally"라고 쓴다. §0 기준에 따라 **판단 근거로 쓰지 않는다.** + +### 4.3 우리에게 없는 기능만 추리면 + +`.claude/agents/` 9종, `paper-production-orchestrator`, `auto_review_gate.py`, `verify_citations.py`, `evals/` 2종과 대조해 **중복을 뺀 신규**: + +| 기능 | 우리에게 없나 | 도입 판단 | +|---|---|---| +| **independence audit** — 파생 사본 클러스터링(재출판 5건이 합의 5표로 세어지지 않게) | 없음 | 발상은 참고할 만함. 우리 문헌 규모에서 실익은 미확인 | +| **run resume (manifest)** — 죽은 단계에서 재개 | 없음 | ⚠️ 우리 실패 #2·#3(침묵 사망)과 **인접**하나 겨냥이 다름. 재개는 사후 복구지 침묵 방지가 아님 | +| **persistent vault** (markdown+SQLite) | 없음 | 우리 `research/REFERENCE_LIST.md`(77편) 로 부분 대체 중 | +| 4 adversarial critics · tool-locked patcher | **있음** — `paper-critic`, `auto_review_gate` | 중복. 그리고 §5 참조 | +| cite-check | **있음** — `verify_citations.py` + `CITATION_AUDIT_2026-07-17.md` | §4.1 이유로 대체 불가 | + +--- + +## 5. 가져오면 안 되는 것 — 검증 관문 + +hyperresearch 의 4 critics + tool-locked patcher + cite-check 는 **검증·비평 층**이다. 이 층을 외부에 맡기는 것은 §0 기준의 후단("관문만큼은 우리 손으로")에 정면으로 걸린다. + +위험은 기술적인 것이 아니라 **규율 차원**이다. `cite-check`·`critic` 같은 이름이 붙어 있으면, 우리 내부 게이트가 잡아야 할 실패를 외부 도구가 잡아 줄 것처럼 착각하게 된다. 실제로 §4.1 에서 확인했듯 그 도구는 **우리 원고를 검사하지도 못한다.** + +--- + +## 6. 가장 중요한 발견 — 새 실패 범주를 아무도 안 잡는다 + +실패셋 11건을 성질별로 재분류하면 이렇다. + +| 성질 | 해당 | 잡는 수단 | 두 저장소가 잡나 | +|---|---|---|---| +| 파이프라인 조용한 실패 | #1 #2 #3 #7 | 스모크 회귀 | ❌ | +| 문서·주장 오류 | #4 #5 #6 | Critic eval | ❌ (§4.1) | +| **다중 소스 상태 불일치** | **#8 #9 #10 #11** | **현재 없음** | ❌ | + +#8~11 은 2026-07-27 하루에 나왔고, 앞의 두 범주와 성질이 다르다. 런타임 오류도 아니고 한 문서 안의 오기도 아니다. **git 커밋·정본 스코어보드·Jira·원고가 서로 다른 상태를 말하는 것**이다. + +`docs/HARNESS_REVIEW_2026-07-17.md` §1.2 의 경고를 여기에도 적용해야 한다: + +> "이 둘을 한 바구니에 넣으면 **scorer 가 못 잡는 걸 잡는 척하게 된다.**" + +세 번째 바구니가 생겼고, 그것을 겨냥한 도구는 조사한 두 저장소 어디에도 없다. **밖에서 사 올 자리가 아니라 우리가 만들 자리다** — BIOP02-107(원고↔결과문서 드리프트 체커)이 그 자리를 겨냥한다. + +--- + +## 7. 3모델 교차검증 기록 + +같은 사실 카드를 주고 질문을 갈랐다. Claude(오케스트레이션·실측) / GPT(적대적 심사) / Gemini(스코프 대조). + +**합의된 것** + +- 기존 no-borrow 결정을 뒤집을 근거는 없다 (GPT #1) +- README 벤치마크는 자기평가 점수로 취급, 근거로 쓰지 않는다 (GPT #4) +- 검증 관문 층은 외부에 맡기지 않는다 (GPT #2, Gemini #6) +- K-Dense 재조사는 **데이터 조달·치료증거 스코프에 한해** 정당하다 (Gemini #5) +- #8~11 은 기존과 다른 범주다 (Gemini #3 — "다중 소스 형상/상태 불일치") + +**갈린 것 — 기록해 둔다** + +hyperresearch cite-check 를 시범 도입할지에서 갈렸다. + +- **Gemini**: 도입 자체가 선을 넘는다. 우리에게 이미 `verify_citations.py`·eval 코퍼스가 있다. +- **GPT**: 도입하되 **독립 프로세스가 아니라 `evals/citation_verifier/` mutation fixture 에 붙인 비교군으로만**. 금지 조건 3개(PR 게이트 대체 금지 · 원고 자동수정 금지 · 실패셋 미통과 시 "도입" 표현 금지). + +작성자 판단은 GPT 쪽이다. 근거는 §4.1 을 코드로 확인하기 전까지 "우리보다 나은가"에 답할 수 없었다는 것이고, 그 확인 비용이 낮다는 점이다(`CITATION_AUDIT_2026-07-17.md` 77편 baseline 이 이미 있어 diff 파일럿이 싸다). **다만 §4.1 확인 결과 vault 결합이 드러났으므로, 이 비교군조차 실익이 크지 않을 수 있다.** 리뷰에서 판단을 구한다. + +--- + +## 8. 권고 (전부 조건부, 결정은 팀) + +1. **차용 없음 유지.** 두 저장소에서 지금 설치할 것은 없다. +2. **미조사 층만 좁게 재조사.** K-Dense `imaging-data-commons`·`depmap` 을 실패셋 11건 대조로. 클론만, 설치 없음. +3. **hyperresearch cite-check 비교군** (선택, §7 에서 갈린 항목). 채택 시 금지 조건 3개를 함께 명문화. + +--- + +## 9. 확인 못 한 것 (정직 기록) + +- **K-Dense `imaging-data-commons`·`depmap` 의 실행코드 유무를 아직 세지 않았다.** 권고 2가 그 작업이다. 이 문서는 "미조사"라고만 말하고 "유망"이라고 말하지 않는다. +- **hyperresearch 를 실제로 돌려 보지 않았다.** 설치가 env 격리 규율에 걸려 소스 정독으로만 판단했다. `sample_rate=0.6` 이 런타임에 어떻게 바뀌는지 등은 미확인. +- **independence audit·run resume 의 실효는 미측정이다.** 우리 문헌 규모(77편)에서 파생 사본 클러스터링이 실익이 있는지 재보지 않았다. +- K-Dense 154개 중 **스킬 이름만 census 했고 전수 정독은 하지 않았다.** 병리·집필·단일세포는 기조사분을 재확인했고, 나머지는 §3.1 표의 후보만 지목했다. diff --git a/pipeline/hspc-velocity-benchmark/ORCHESTRATION-WIRING-DESIGN.md b/pipeline/hspc-velocity-benchmark/ORCHESTRATION-WIRING-DESIGN.md index ea7e355..a3962ae 100644 --- a/pipeline/hspc-velocity-benchmark/ORCHESTRATION-WIRING-DESIGN.md +++ b/pipeline/hspc-velocity-benchmark/ORCHESTRATION-WIRING-DESIGN.md @@ -9,17 +9,20 @@ > kkkim A(복원) 확정 → **PR #5 머지(`cb886ee`)로 복원 완료**. 실측: `skills/ROUTES.md`(69줄) 등 41파일 복귀, > `harness_doctor` **RESULT PASS(phantom-path 0)**. → 라우터 blocker 제거, 초안 §3의 dataset→task 2단 라우팅이 실체 위에 섬. > -> 🚧 **선결 ② 신규 (2026-07-27, 이건규 BIOP01-71 코멘트 11506) — 트리거 계층 결정 대기.** -> **`openclaw` CLI가 팀 컨테이너 6개 전부에 없다(0/6, 지용기 이 박스도 실측 부재).** 이 초안 제목·§3이 전제한 -> "OpenClaw route로 트리거"가 **실행 도구 부재**로 지금 성립하지 않는다. 대체 경로는 존재한다 — -> `skills/OPENCLAW-RUN.md`가 같은 `agents/openai.yaml`을 **codex로 실행**하는 길을 기록하고 있고 codex는 인증돼 동작한다. -> → **트리거를 openclaw로 갈지, codex 기준으로 문서·설계를 정정할지는 회의 결정**(BIOP01-45와 묶어 상정, 이건규 제안). -> **이 결정이 나기 전 §3-2·3-4 route 의사코드는 확정하지 않는다**(openclaw 전제라 codex면 표기가 달라짐). +> ✅ **선결 ② 해소 (2026-08-04, 리더 결정) — 트리거 = codex, 형식 = YAML.** +> `openclaw` CLI가 팀 6대 중 실사용 불가(5대 바이너리 부재 + 6대 auth 부재, kkkim/이건규 실측)인 반면 **codex는 설치·인증 완료**. +> → **트리거는 codex**(같은 `agents/openai.yaml`을 codex로 실행, `skills/OPENCLAW-RUN.md` 3번 경로), **manifest 형식은 YAML**로 확정. +> 남은 서버측 선결 1건: "codex가 이 `openai.yaml`을 실제로 완주하는지" 1회 실증(codex 샌드박스 `bwrap` 이슈 포함) — kkkim 서버. > -> ✅ **단 §2 `runner_manifest.yaml`(워커 계층)은 두 결정 모두에 불변이다.** manifest는 "어느 runner를 어떤 env로"만 -> 정의하고, "누가 트리거하나"(openclaw/codex/사람)는 상위 계층이다. 라우터 복원으로 하위 배선 근거는 이미 실체가 됐고, -> 트리거가 openclaw든 codex든 manifest는 그대로 재사용된다. → **manifest 스키마 확정 + §6 CPU 검증은 트리거 결정과 무관하게 선행 가능.** -> (근거: 코멘트 11397·11439의 계층 분리 논지 그대로.) +> **결정에 따라 실체화된 것 (2026-08-04):** +> - `cross_dataset/runner_manifest.yaml` — 워커 계층 계약(YAML). gse205117 실산출물로 멱등·계약 인라인 검증 PASS. +> - `cross_dataset/verify_manifest_resolution.py` — 정식 검증기(멱등·누락감지·계약 3검사). 실행 PASS(exit 0). +> - `cross_dataset/run_from_manifest.sh` — codex가 트리거하는 실행 래퍼. manifest를 읽어 stage별 `conda run`, +> skip-if-output + required_cols 사후검증(BIOP01-41 게이트). `--dry-run`으로 gse205117 전 stage SKIP 확인. +> +> ✅ **§2 `runner_manifest.yaml`(워커 계층)은 트리거/형식 결정과 애초에 무관했다** — manifest는 "어느 runner를 어떤 env로"만 +> 정의하고, "누가 트리거하나"(codex)는 상위 계층이다. (근거: 코멘트 11397·11439의 계층 분리 논지.) +> **남은 것**: codex→래퍼 실증(서버) + numpy env byte-identical scorecard 재생성(kkkim) → DoD "end-to-end 1회" 마감. --- @@ -139,13 +142,16 @@ score: # P3 채점 (fit 위에서) --- -## 5. 협의가 필요한 열린 결정 (braveji ↔ kkkim) +## 5. 열린 결정 — 처리 상태 (리더 결정 2026-08-04) + +1. **manifest 형식** → ✅ **YAML 확정.** `cross_dataset/runner_manifest.yaml`. (PyYAML은 repo에 이미 있음 확인 6.0.1.) +2. **route 실행 주체** → ✅ **래퍼 확정.** `cross_dataset/run_from_manifest.sh`가 manifest를 읽어 실행; 트리거(codex)는 이 래퍼만 부른다. 기존 watchdog·flock 자산과 병행. +3. **GPU 스케줄링** → 래퍼는 manifest `gpu:true` stage에 `CUDA_VISIBLE_DEVICES=1` 고정(기존 규약). 여유 GPU 탐지는 후속 개선(비필수). +4. **dataset별 채점기 일반화** → **별건으로 분리 유지.** 지금은 `p3_prereg_.py`를 score stage에 직접 지정. parametric 통합은 후속 리팩터. +5. **적용 범위** → **P2(fit) + P3(prereg 채점)까지.** manifest `stages`(P2) + `score`(P3). P4/P5(bootstrap·FDR)는 후속. -1. **manifest 형식**: YAML 신규 파일 vs 기존 `p2_config.py`에 dict로 넣기. (초안은 YAML — 언어중립·OpenClaw가 읽기 쉬움. 단 repo에 PyYAML 의존 추가됨.) -2. **route 실행 주체**: OpenClaw가 직접 `conda run` vs 얇은 `run_from_manifest.sh` 래퍼를 호출. (초안은 래퍼 권장 — 기존 watchdog·flock 자산 재사용, OpenClaw는 트리거만.) -3. **GPU 스케줄링**: `CUDA_VISIBLE_DEVICES`를 manifest 고정 vs route가 여유 GPU 탐지 후 주입. (GPU 서버 Xid79 상황과 연동 — kkkim 판단.) -4. **dataset별 채점기 일반화**: 지금 `p3_prereg_.py`가 dataset마다 하나. manifest score stage를 dataset-parametric 단일 스크립트로 통합할지. (별건 리팩터로 분리 제안.) -5. **적용 범위**: P2(fit)만 자동화 vs P3~P5(채점·bootstrap·FDR)까지. (초안은 P2+P3까지, P4/P5는 후속.) +> **트리거 = codex 확정(2026-08-04):** openclaw 팀 전체 실사용 불가, codex 설치·인증됨. codex가 `agents/openai.yaml`을 실행 → `run_from_manifest.sh` 호출. +> 서버측 선결 1건: codex가 openai.yaml을 실제 완주하는지 실증(bwrap 샌드박스 결정 포함) — kkkim. --- diff --git a/pipeline/hspc-velocity-benchmark/cross_dataset/run_from_manifest.sh b/pipeline/hspc-velocity-benchmark/cross_dataset/run_from_manifest.sh new file mode 100755 index 0000000..879cff6 --- /dev/null +++ b/pipeline/hspc-velocity-benchmark/cross_dataset/run_from_manifest.sh @@ -0,0 +1,107 @@ +#!/usr/bin/env bash +# run_from_manifest.sh — BIOP01-45 워커 계층 실행 래퍼 (트리거=codex, 형식=YAML 확정 2026-08-04) +# +# codex가 openai.yaml → 이 래퍼를 트리거한다. 래퍼는 runner_manifest.yaml을 읽어 +# 각 stage를 올바른 env로 conda run 하고, 산출물이 있으면 skip, 완료 후 required_cols를 검증한다. +# "누가 트리거하나"(codex)와 "무엇을 어떻게 실행하나"(이 래퍼+manifest)를 분리 — 트리거가 바뀌어도 이 래퍼는 불변. +# +# 설계: ORCHESTRATION-WIRING-DESIGN.md §3·§5-2 (얇은 래퍼 — 기존 watchdog/flock 자산과 병행). +# 사용: +# run_from_manifest.sh --dataset gse205117 [--dry-run] [--conda /path/to/conda] +# --dry-run : 실행하지 않고 stage별 SKIP/RUN 결정과 명령만 출력 (CPU·GPU 불요, 어디서나 검증 가능) +set -euo pipefail + +HERE="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" # cross_dataset/ +ROOT="$(cd "$HERE/.." && pwd)" # pipeline/hspc-velocity-benchmark +SCRIPTS="$ROOT/scripts" +RESULTS="$ROOT/results" +MANIFEST="$HERE/runner_manifest.yaml" + +DATASET="" +DRYRUN=0 +CONDA="${CONDA:-/home/kkkim/miniconda3/bin/conda}" # 서버별 상이 → --conda로 override +while [ $# -gt 0 ]; do + case "$1" in + --dataset) DATASET="$2"; shift 2;; + --dry-run) DRYRUN=1; shift;; + --conda) CONDA="$2"; shift 2;; + *) echo "unknown arg: $1" >&2; exit 2;; + esac +done +[ -n "$DATASET" ] || { echo "--dataset 필요 (예: gse205117)" >&2; exit 2; } + +SUFFIX="_${DATASET}" +CONFIG="../cross_dataset/config_${DATASET}.py" # scripts/ 기준 상대경로 (기존 규약) + +log(){ echo "[$(date '+%F %T')] $*"; } + +# manifest에서 stage 메타를 파이프 구분 라인으로 방출 (id|env|gpu|runner|output|required_cols). +# score stage도 kind=score로 함께 (P3 채점까지 한 래퍼에서). +emit_stages() { +python3 - "$MANIFEST" <<'PY' +import sys, yaml +m = yaml.safe_load(open(sys.argv[1])) +def row(kind, s): + out = s.get("output","") or "" + req = ",".join(s.get("required_cols",[]) or []) + print("|".join([kind, s["id"], s["env"], str(s.get("gpu",False)).lower(), + s["runner"], out, req])) +for s in m["stages"]: row("fit", s) +for s in m["score"]: row("score", s) +PY +} + +verify_cols() { # $1=csv 경로, $2=쉼표구분 required_cols → 없으면 exit 1 + local csv="$1" req="$2" + [ -n "$req" ] || return 0 + python3 - "$csv" "$req" <<'PY' +import sys, csv +path, req = sys.argv[1], sys.argv[2].split(",") +with open(path, newline="") as f: hdr = next(csv.reader(f), []) +missing = [c for c in req if c not in hdr] +sys.exit(1 if missing else 0) +PY +} + +log "manifest=$MANIFEST dataset=$DATASET suffix=$SUFFIX dry_run=$DRYRUN" +[ "$DRYRUN" = 1 ] || { [ -x "$CONDA" ] || { echo "conda 실행 불가: $CONDA (--conda로 지정)" >&2; exit 3; }; } + +FAIL=0 +while IFS='|' read -r kind id env gpu runner output req; do + # 산출물 경로(있으면). {suffix} 치환. + outfile="" + if [ -n "$output" ]; then outfile="$RESULTS/${output/\{suffix\}/$SUFFIX}"; fi + + # skip 판정: 산출물이 있고 ≥2행이면 SKIP + if [ -n "$outfile" ] && [ -f "$outfile" ] && [ "$(wc -l <"$outfile")" -ge 2 ]; then + log "[$id] SKIP (산출물 존재: $(basename "$outfile"), $(wc -l <"$outfile")행)" + # 계약 검증은 skip이어도 수행 (BIOP01-41: 존재해도 컬럼 계약 위반이면 잡는다) + if [ -n "$req" ]; then + if verify_cols "$outfile" "$req"; then log " required_cols OK: $req" + else log " ❌ required_cols 위반: $req"; FAIL=1; fi + fi + continue + fi + + # 실행 명령 구성 + cuda=""; [ "$gpu" = "true" ] && cuda="CUDA_VISIBLE_DEVICES=1 " + cmd="cd $SCRIPTS && ${cuda}CROSS_DATASET_CONFIG=$CONFIG CROSS_DATASET_SUFFIX=$SUFFIX $CONDA run --no-capture-output -n $env python -u $runner" + + if [ "$DRYRUN" = 1 ]; then + log "[$id] RUN (dry) env=$env gpu=$gpu → ${output:-<산출물 없음>}" + echo " $cmd" + continue + fi + + log "[$id] RUN env=$env gpu=$gpu" + ( eval "$cmd" ) || { log "❌ [$id] 실행 실패"; FAIL=1; break; } + # 산출물·계약 사후 검증 + if [ -n "$outfile" ]; then + { [ -f "$outfile" ] && [ "$(wc -l <"$outfile")" -ge 2 ]; } || { log "❌ [$id] 산출물 미생성: $outfile"; FAIL=1; break; } + if [ -n "$req" ] && ! verify_cols "$outfile" "$req"; then log "❌ [$id] required_cols 위반: $req"; FAIL=1; break; fi + log " ✅ $(basename "$outfile") 생성·계약 충족" + fi +done < <(emit_stages) + +if [ "$FAIL" = 0 ]; then log "DONE — 전 stage 해소(SKIP/RUN) + 계약 충족"; exit 0 +else log "FAILED — 위 로그 참조"; exit 1; fi diff --git a/pipeline/hspc-velocity-benchmark/cross_dataset/runner_manifest.yaml b/pipeline/hspc-velocity-benchmark/cross_dataset/runner_manifest.yaml new file mode 100644 index 0000000..58c47b9 --- /dev/null +++ b/pipeline/hspc-velocity-benchmark/cross_dataset/runner_manifest.yaml @@ -0,0 +1,58 @@ +# runner_manifest.yaml — BIOP01-45 워커 계층 계약 (P2 fit → P3 채점) +# +# 이 파일은 흩어진 runner 실행 규약(env·GPU·출력·필수컬럼)을 한 곳에 선언한다. +# 트리거(openclaw/codex/사람)가 무엇이든 이 계약은 불변이다 — "누가 실행하나"가 아니라 +# "어느 runner를 어떤 env로, 무엇을 산출하는가"만 정의한다. (설계: ORCHESTRATION-WIRING-DESIGN.md §2) +# +# 상태: 내용 확정(설계 §2와 1:1), **직렬화 형식(YAML vs p2_config dict)은 BIOP01-45 §5-1 회의 결정 대기**. +# YAML을 기본 권장(언어중립·리뷰 용이). 형식이 dict로 확정되면 이 데이터를 그대로 옮기면 된다. +# 검증: gse205117 실 산출물로 stage 해소(멱등 SKIP)·required_cols 계약을 **인라인 대조 완료** +# (BIOP01-45 코멘트, commit 5a21945: 전 stage 산출물 존재 + required_cols 실체 일치, 계약 위반 0). +# 정식 검증기 스크립트화는 형식(YAML/dict) 확정 후 — 지금 미존재 경로를 가리키지 않도록 인라인으로 표기. + +version: 1 + +defaults: + cwd: scripts # 모든 runner는 scripts/에서 실행 + suffix_env: CROSS_DATASET_SUFFIX # 예: _gse205117 → 파일명 접미사 + config_env: CROSS_DATASET_CONFIG # 예: ../cross_dataset/config_gse205117.py + outputs_dir: results + runtime_log: results/runtime.csv + +stages: # P2 fit 산출 (선언 순서 = 의존성 순서) + - id: floor + runner: p2_rna_only.py + env: scv-preprocess + gpu: false + output: rna_only_dynamical_genes{suffix}.csv + required_cols: [gene, fit_alpha, fit_likelihood] + - id: multivelo + runner: p2_multivelo.py + env: velo-mv + gpu: false + output: multivelo_genes{suffix}.csv + required_cols: [gene, fit_alpha, fit_t_sw1, fit_t_sw2, fit_likelihood] + - id: dl_prep + runner: p2_dl_prep.py + env: velo-torch + gpu: true + produces_input_for: [multivelovae, moflow] # 산출 CSV 없음(중간 substrate) + - id: multivelovae + runner: p2_multivelovae.py + env: velo-torch + gpu: true + output: multivelovae_genes{suffix}.csv + required_cols: [gene, vae_alpha, vae_alpha_c] + - id: moflow # 사전등록 예측5(원정의 MV×MoFlow)에 필수 — BIOP01-41 교훈 + runner: p2_moflow.py + env: velo-torch + gpu: true + output: moflow_genes{suffix}.csv + required_cols: [gene, cs_lag_median] # cs_lag(오타) 아님 — BIOP01-41 조용한 폴백 사고 방지 + +score: # P3 채점 (fit 위에서, CPU) + - id: prereg + runner: cross_dataset/p3_prereg_gse205117.py # dataset별 채점기 (parametric 통합은 별건) + env: scv-preprocess + gpu: false + output: prereg{suffix}_scorecard.md diff --git a/pipeline/hspc-velocity-benchmark/cross_dataset/verify_manifest_resolution.py b/pipeline/hspc-velocity-benchmark/cross_dataset/verify_manifest_resolution.py new file mode 100644 index 0000000..93a2871 --- /dev/null +++ b/pipeline/hspc-velocity-benchmark/cross_dataset/verify_manifest_resolution.py @@ -0,0 +1,116 @@ +#!/usr/bin/env python +"""runner_manifest.yaml 검증기 — BIOP01-45 §6 CPU 검증 (트리거·GPU 무관). + +manifest의 stage 해소 로직과 required_cols 계약을, **이미 완주한 데이터셋 산출물**로 대조한다. +runner 실행·GPU 없음, stdlib + PyYAML만. (형식=YAML, 트리거=codex 확정 2026-08-04로 스크립트화.) + +세 검사: + (1) 멱등성 — 산출물이 다 있으면 모든 output stage가 SKIP으로 해소되나? + (2) 누락 감지 — 특정 산출물을 (가상) 부재로 두면 그 stage만 RUN으로 잡히나? + (3) 계약 검증 — 각 산출물이 manifest 선언 required_cols를 실제로 갖는가? + (BIOP01-41의 cs_lag vs cs_lag_median '조용한 폴백' 계열 사고를 manifest 층에서 차단) + +실행: python cross_dataset/verify_manifest_resolution.py [--suffix _gse205117] [--hide moflow] +종료코드: 0=PASS, 1=FAIL. +""" +from __future__ import annotations + +import argparse +import csv +import sys +from pathlib import Path + +import yaml + +ROOT = Path(__file__).resolve().parents[1] # pipeline/hspc-velocity-benchmark +MANIFEST = Path(__file__).with_name("runner_manifest.yaml") + + +def load_manifest(): + with open(MANIFEST) as fh: + return yaml.safe_load(fh) + + +def csv_header(path: Path): + with open(path, newline="") as fh: + return next(csv.reader(fh), []) + + +def resolve(tmpl: str, suffix: str) -> str: + return tmpl.replace("{suffix}", suffix) + + +def output_stages(m): + """output을 내는 stage만 — dl_prep 같은 중간 substrate(output 없음)는 제외.""" + return [s for s in m["stages"] if s.get("output")] + + +def check_idempotency(m, results: Path, suffix: str) -> bool: + print("\n(1) 멱등성 — 산출물이 있으면 SKIP") + ok = True + for s in output_stages(m): + out = results / resolve(s["output"], suffix) + rows = sum(1 for _ in open(out)) if out.exists() else 0 + state = "SKIP" if (out.exists() and rows >= 2) else "RUN" + print(f" {'✅' if state=='SKIP' else '❌'} {s['id']:<12} → {state:<4} ({out.name}, {rows}행)") + ok &= (state == "SKIP") + print(f" → {'전 stage SKIP(멱등)' if ok else '일부 RUN — 산출물 누락'}") + return ok + + +def check_missing_detection(m, results: Path, suffix: str, hide: str) -> bool: + print(f"\n(2) 누락 감지 — '{hide}' 산출물 가상 부재") + detected = False + for s in output_stages(m): + out = results / resolve(s["output"], suffix) + present = out.exists() and s["id"] != hide # hide stage만 부재로 시뮬레이션 + state = "SKIP" if present else "RUN" + if s["id"] == hide: + detected = (state == "RUN") + print(f" {'✅' if detected else '❌'} {s['id']:<12} → {state} (가상 부재 감지 {'OK' if detected else '실패'})") + else: + print(f" {s['id']:<12} → {state}") # 다른 stage는 SKIP 유지(오탐 없음) + print(f" → {'누락 stage만 RUN으로 감지' if detected else '감지 실패'}") + return detected + + +def check_contract(m, results: Path, suffix: str) -> bool: + print("\n(3) 계약 — 산출물이 required_cols를 실제로 갖는가 (BIOP01-41 방지)") + ok = True + for s in output_stages(m): + req = s.get("required_cols") or [] + out = results / resolve(s["output"], suffix) + if not out.exists(): + print(f" ⚠️ {s['id']:<12} 산출물 없음 — 계약 확인 skip") + continue + missing = [c for c in req if c not in csv_header(out)] + print(f" {'✅' if not missing else '❌'} {s['id']:<12} {req} {'전부 존재' if not missing else f'누락 {missing}'}") + ok &= not missing + print(f" → {'모든 계약 충족(컬럼명 실체 일치)' if ok else '계약 위반 — manifest/산출물 점검'}") + return ok + + +def main(argv=None): + ap = argparse.ArgumentParser() + ap.add_argument("--suffix", default="_gse205117") + ap.add_argument("--hide", default="moflow", help="누락 감지 시뮬레이션 대상 stage id") + args = ap.parse_args(argv) + + m = load_manifest() + results = ROOT / m["defaults"]["outputs_dir"] + print(f"manifest {MANIFEST.name} v{m['version']} | suffix={args.suffix} | results={results}") + print(f"stages={[s['id'] for s in m['stages']]} score={[s['id'] for s in m['score']]}") + + r1 = check_idempotency(m, results, args.suffix) + r2 = check_missing_detection(m, results, args.suffix, args.hide) + r3 = check_contract(m, results, args.suffix) + + passed = r1 and r2 and r3 + print(f"\n{'='*56}\n종합: 멱등 {'✅' if r1 else '❌'} · 누락감지 {'✅' if r2 else '❌'} · 계약 {'✅' if r3 else '❌'}" + f" → {'PASS' if passed else 'FAIL'}") + print("주의: byte-identical scorecard 재생성은 numpy env(scv-preprocess) 필요 → kkkim 서버.") + return 0 if passed else 1 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/pipeline/hspc-velocity-benchmark/manuscript/MEETING_BRIEF_2026-07-21.md b/pipeline/hspc-velocity-benchmark/manuscript/MEETING_BRIEF_2026-07-21.md index 249d360..f98332e 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/MEETING_BRIEF_2026-07-21.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/MEETING_BRIEF_2026-07-21.md @@ -134,3 +134,36 @@ HSPC를 포함해 human brain, E18 mouse brain, human BMMC, macrophage, mouse ga - BIOP01 쪽은 스터디 과제 티켓만 남아 있습니다(BIOP01-5, 12, 19). 이 트랙을 계속 갈지 정리할지 회의에서 정하겠습니다. 부담 없이 보실 수 있는 분량만 보고 오셔도 됩니다. 나머지는 회의에서 같이 정리하겠습니다. + +--- + +## 8. 저자·투고 메타데이터 확정 (추가 안건, 2026-08-04) + +저널은 **Genome Biology로 확정**됐다(단어 상한이 없어 본문 감축은 선택, 초록은 421→350단어로 반드시 줄여야 함). 남은 결정은 저자와 소속, 선언문이다. 근거는 BIOP01-75/79와 BIOP02-114에 정리돼 있고, 아래는 회의에서 정할 항목만 추린 것이다. + +**8-1. 저자 순서 (Paper A / HSPC).** git·JIRA 실기여 기준 초안이다. 커밋 수가 아니라 지적 기여의 크기로 정하며, ICMJE 네 항목을 모두 충족해야 저자다. + +| 순위 | 저자 | 근거 | +|---|---|---| +| 제1·교신 | 김가경 | 파이프라인 P0~P5, velocity 벤치마크, 층② 감사·외부재현, 원고 영/한 | +| 2 | 이건규 | 하네스 정합성 게이트, 라우터 복원, 저널 리스트업, 원고 간결화 | +| 3 | 지용기 | 오케스트레이션, Critic 총괄 | +| 4 | 박세진 | 독립 cross-review(박상준 승계) | +| Ack | 류재면, 박상준 | 8-4 참조 | + +정할 것은 이건규 님과 지용기 님의 순서, 박세진 님의 위치다. 코드에 안 남는 회의·설계 기여가 순서를 바꿀 수 있다. + +**8-2. 교신저자.** 가장 훌륭한 사람이 아니라 투고·심사·교정·출판 후 문의까지 끝까지 책임질 사람으로 정한다. 8-3(소속)과 한 세트다. + +**8-3. 소속 (핵심 맞교환: 게재료 대 IP).** GB는 완전 OA라 APC가 £3790(약 741만원)이고 무료 게재 경로가 없다. + +- 무소속(Independent Researcher)으로 가면 회사 IP 문제를 피하지만 741만원을 자비로 내야 하고 면제 지렛대가 없다. +- 협정 보유 기관 소속 저자를 교신저자로 세우면 Springer Read & Publish 협정으로 APC를 £0로 만들 수 있으나 IP가 노출된다. + +함정 하나. 다수 R&P 협정이 하이브리드 저널만 커버하고 BMC 완전-OA(GB 포함)는 제외한다. 기관별로 "Genome Biology 포함" 여부를 반드시 확인해야 한다. + +**8-4. 박상준 님 처리 (본인 의사 확인됨).** 저자가 아니라 Acknowledgements에 넣는다. ICMJE ①(주제 착상)만 충족하고 ②③④는 미충족이라, 이름만 넣으면 gift authorship으로 심사에서 흠이 된다. 본인도 "아니어도 괜찮다"고 했다. 하네스 출처(Harness_Baseline)는 저자 여부와 별개로 명기하고, 본인 동의 한 줄을 카톡으로 남겨 둔다. + +**8-5. 선언문·메타데이터 (사람만 채울 수 있음).** 저자 전원의 공식 영문명, competing interests, funding, author contributions(CRediT), acknowledgements, 코드·데이터 repository DOI 두 개. + +**8-6. 분량 목표 (BIOP01-74 착수 조건).** GB 확정이라 본문 감축은 선택이고 초록 감축만 필수다. 감축을 심사 인상용으로 돌릴지만 정하면 BIOP01-74가 풀린다. 같은 GB에 Wu 2026 종합 벤치마크가 실려 desk-reject 위험이 올라갔으므로, "이전 벤치마크와 무엇이 다른가" 문단을 방어선으로 유지한다. diff --git a/pipeline/hspc-velocity-benchmark/manuscript/README.md b/pipeline/hspc-velocity-benchmark/manuscript/README.md index f16fa44..064be08 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/README.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/README.md @@ -6,7 +6,7 @@ | 파일 | 내용 | |---|---| | `draft_v2.md` | **본문 정본**(영어). 한국어 검토본 = `draft_v2_ko.md`. 내용 수정은 항상 두 파일 동시 | -| `refs.bib` | 인용 — `paper_analysis/*//*.bib`에서 모음 | +| `refs.bib` | 인용 — 본문 참고문헌 목록[1]~[N]에서 `scripts/build_refs_bib.py draft_v2.md`로 재생성(진리원천=목록). 검증 `bib_to_cites.py \| verify_citations.py`. (구 "paper_analysis/*.bib 수집"은 method 논문 .bib 부재로 71개 못 만들어 폐기) | | `supplementary.md` | 보충 (추가 표·방법 상세) | | `figure-legends.md` | `../figures/figNN`과 1:1 대응하는 legend | diff --git a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md index febd57d..77b1b3f 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md +++ b/pipeline/hspc-velocity-benchmark/manuscript/draft_v2.md @@ -17,7 +17,7 @@ No arrow glyphs in title, body or tables. Research- and education-use draft. Written NEW alongside draft.md (draft.md preserved for comparison/gate). Do not overwrite draft.md. --> -# A reliability map for multiome RNA velocity outputs in single-cell kinetics +# A reliability map for per-gene multiome RNA velocity parameters in single-cell kinetics -# single-cell kinetics에서 multiome RNA velocity 출력의 신뢰도 지도 +# single-cell kinetics에서 유전자별 multiome RNA velocity 모수의 신뢰도 지도 **저자:** @@ -21,11 +21,11 @@ draft_v2.md의 한국어 검토본(번역·윤문). 정본은 영어 draft_v2.md **배경.** Chromatin 정보를 결합한("multiome") RNA velocity 방법들은 유전자별로 여러 값을 산출한다. 전사 속도(transcription rate), 분해 속도(degradation rate), 그리고 chromatin에서 transcription으로 이어지는 *시간차(lag)*, 곧 locus가 열리거나 닫히는 시점과 그 유전자의 transcription이 전환되는 시점 사이의 오프셋이 그것이다. 이 값들은 저마다 생물학적 판독값으로 제안되어 왔고, 특히 시간차(lag)는 epigenetic 약물 반응의 timing을 예측하는 데 쓸 수 있다고 여겨져 왔다. 그러나 파생된 값을 하류에서 쓰려면 그것이 먼저 *신뢰할 수 있어야* 한다. 곧 합리적인 알고리즘들에 걸쳐 재현되어야 하고, 가능하다면 독립적인 측정값과도 부합해야 한다. 우리는 인간 조혈모·전구세포(hematopoietic stem and progenitor cells, HSPC; 10x Multiome으로 프로파일링)에서 어떤 velocity 출력이 이 기준을 충족하는지를 물었다. -**접근.** 최대 다섯 개의 velocity arm(갈래)(RNA 전용 scVelo floor에 MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo를 더한 것)에 걸쳐 각 출력을 네 축에서 검정했다. 방법 간 재현성(순열 FDR), lineage 내 ATAC-shuffle 인과 대조군, 다섯 개 외부 multiome(그중 하나는 사전등록)에서의 cross-dataset 재현, 그리고 fitting된 속도가 측정된 합성(K562 TT-seq)·분해(mRNA 반감기) 속도에 외부 anchoring되는지가 그것이다. +**접근.** 최대 다섯 개의 velocity arm(갈래)(RNA 전용 scVelo floor에 MultiVelo, MultiVeloVAE, MoFlow, CRAK-Velo를 더한 것)에 걸쳐 각 출력을 네 축에서 검정했다. 방법 간 재현성(순열 FDR), lineage 내 ATAC-shuffle 인과 대조군, 다섯 개 외부 multiome(그중 하나는 사전등록)에서의 cross-dataset 재현, 그리고 fitting된 속도가 측정된 합성(K562 TT-seq)·분해(mRNA 반감기) 속도에 외부 anchoring되는지가 그것이다. 우리는 유전자별 kinetic parameter와, 한 단계 위인 세포×유전자 velocity 행렬을 감사 대상으로 삼았다. velocity가 주로 쓰이는 저차원 임베딩(embedding) arrow나 궤적(trajectory)은 감사 대상에 포함하지 않았다. **결과.** 출력들은 뚜렷이 갈라졌다. 전사 속도 α만 방법 간에 재현되었고(Spearman ρ=0.88, 관측값), chromatin에서 transcription으로 이어지는 시간차(lag)와 분해 속도 γ는 그렇지 않았다. 시간차(lag)는 크기에서 잘해야 약하게 재현되었고(가장 강한 쌍 ρ=+0.163, 대부분의 쌍 |ρ|≤0.08) 부호에서는 우연 수준에 그쳤으며(54.6%), ATAC를 뒤섞어도 통계적으로 변하지 않아 시간차(lag)가 chromatin에서 비롯된 것이 아니라 모델 구조에서 비롯됨을 드러냈다. 분해 속도 γ도 마찬가지로 방법에 취약했고(방법 간 ρ≈−0.1), ground truth가 있는 곳에서도 복원되지 않았다(K562 반감기 3/3 귀무; 교과서적 scVelo γ는 반대로 나와 −0.224, CI가 0 배제). 외부 보강 증거로서, fitting된 α는 세 방법 모두에서 측정된 K562 TT-seq 합성 속도와도 상관했다(비-housekeeping ρ +0.24 ~ +0.29, 모든 CI가 0 배제). 다만 steady-state transcript abundance가 같은 측정값을 최소한 그만큼 잘 예측하므로(Spearman(abundance, 합성)=+0.410 대 Spearman(α, 합성)=+0.262), 이를 α가 유일하게 측정에 기반한 출력이라는 근거가 아니라 일관성 증거로 읽는다. α가 시간차(lag)보다 앞선다는 순서는 여섯 개 시스템 모두에서 유지되었고, 어떤 fitting보다 먼저 봉인한 사전등록 6-of-6 채점표(scorecard)를 통과했다. 프로파일 우도(profile-likelihood) 분석은 그 기제를 주었다. α는 stiff(식별 가능)한 반면 시간차(lag)는 sloppy하고 경계에 제약되며, 이 식별가능성은 외부 검증과 부분적으로 정렬하나 확증적으로 정렬하지는 않는다. -**결론.** 우리는 이 결과를 velocity 출력 신뢰도 지도로 정리한다. α와 속도에서 파생된 신호는 신뢰한다(방법 간에 재현되고 외부 측정으로 보강되므로 직접 사용 가능). 시간차(lag), 그 부호, 절대 timing, 그리고 γ는 신뢰할 수 없는 것으로 다루며 직교(orthogonal) 검증을 요구한다. 하류의 어떤 timing 예측 모델이든 단일 방법의 시간차(lag)가 아니라 강건한 baseline-ATAC-to-α 경로를 거쳐야 한다. +**결론.** 우리는 이 결과를 velocity 출력 신뢰도 지도로 정리한다. α와 속도에서 파생된 신호는 신뢰할 수 있다. 다만 α는 방법 간에 재현되면서도 대체로 발현량을 반영하는 값이어서, steady-state abundance가 같은 외부 측정값을 α만큼, 또는 그 이상으로 예측하므로 α가 abundance를 넘어서는 합성률 정보를 더한다고는 입증되지 않는다. 시간차(lag), 그 부호, 절대 timing, 그리고 γ는 신뢰할 수 없는 것으로 다루며 직교(orthogonal) 검증을 요구한다. 하류의 어떤 timing 예측 모델이든 단일 방법의 시간차(lag)가 아니라 강건한 baseline-ATAC-to-α 경로를 거쳐야 한다. **키워드:** RNA velocity, single-cell multiome, chromatin accessibility, transcriptional kinetics, parameter identifiability, external validation, benchmarking, hematopoiesis @@ -307,7 +307,7 @@ cell-cycle, 전사 버스트(transcriptional burst), ambient/doublet 교란(conf [9] Gayoso A, Weiler P, Lotfollahi M, et al. Deep generative modeling of transcriptional dynamics for RNA velocity analysis in single cells. *Nature Methods* 21, 50–59 (2024). doi:10.1038/s41592-023-01994-w. -[10] Gu Y, et al. Bayesian inference of RNA velocity incorporating timepoints, lineage bifurcations, and count data (veloVAE). *PLOS Computational Biology* 22(3), e1014060 (2026). doi:10.1371/journal.pcbi.1014060. [Distinct from MultiVeloVAE [2].] +[10] Gu Y, et al. Bayesian inference of RNA velocity incorporating timepoints, lineage bifurcations, and count data (veloVAE). *PLOS Computational Biology* 22(3), e1014060 (2026). doi:10.1371/journal.pcbi.1014060. [Distinct from MultiVeloVAE [4].] [11] Qiao C, Huang Y. Representation learning of RNA velocity reveals robust cell transitions. *Proceedings of the National Academy of Sciences* 118(49), e2105859118 (2021). doi:10.1073/pnas.2105859118. @@ -397,7 +397,7 @@ cell-cycle, 전사 버스트(transcriptional burst), ambient/doublet 교란(conf [54] Lederer AR, Leonardi M, Talamanca L, et al. Statistical inference with a manifold-constrained RNA velocity model uncovers cell cycle speed modulations. *Nature Methods* 21(12), 2271–2286 (2024). doi:10.1038/s41592-024-02471-8. -[55] Gu et al. Profile-likelihood identifiability analysis of single-cell transcription (telegraph) kinetics. *Bioinformatics* 41(11), btaf581 (2025). doi:10.1093/bioinformatics/btaf581. [Distinct from [8].] +[55] Gu et al. Profile-likelihood identifiability analysis of single-cell transcription (telegraph) kinetics. *Bioinformatics* 41(11), btaf581 (2025). doi:10.1093/bioinformatics/btaf581. [56] Wang. Sloppiness and Action Constraint in Cell State Transitions: Are Single Cells Sloppy? bioRxiv 2025.12.31.697145 (v2, 2025). [Methodological analog on cell-state Gaussian coordinates.] diff --git a/pipeline/hspc-velocity-benchmark/manuscript/refs.bib b/pipeline/hspc-velocity-benchmark/manuscript/refs.bib index 2ced919..a2bbff9 100644 --- a/pipeline/hspc-velocity-benchmark/manuscript/refs.bib +++ b/pipeline/hspc-velocity-benchmark/manuscript/refs.bib @@ -1,248 +1,560 @@ -% refs.bib — HSPC velocity chromatin→transcription lag benchmark -% Extracted from manuscript/related_work.md (bibliographic detail verified vs CrossRef/PubMed 2026-07-05). -% Items flagged there as "verify author list / venue" are carried over with a note field. +# generated from draft_v2.md — 71 entries (refs [1]..[71]) +@article{ref01, + author = {La Manno, G and Soldatov, R and Zeisel, A and others}, + title = {RNA velocity of single cells}, + journal = {Nature}, + year = {2018}, + doi = {10.1038/s41586-018-0414-6} +} + +@article{ref02, + author = {Cao, J and Cusanovich, DA and Ramani, V and others}, + title = {Joint profiling of chromatin accessibility and gene expression in thousands of single cells}, + journal = {Science}, + year = {2018}, + doi = {10.1126/science.aau0730} +} -@article{li2023multivelo, - author = {Li, Chen and Virgilio, Maria C. and Collins, Kathleen L. and Welch, Joshua D.}, - title = {Multi-omic single-cell velocity models epigenome--transcriptome interactions and improves cell fate prediction}, +@article{ref03, + author = {Li, C and Virgilio, MC and Collins, KL and Welch, JD}, + title = {Multi-omic single-cell velocity models epigenome–transcriptome interactions and improves cell fate prediction}, journal = {Nature Biotechnology}, - volume = {41}, - pages = {387--398}, - year = {2023}, - doi = {10.1038/s41587-022-01476-y}, - note = {year=2023 is the issue year (Nat Biotechnol vol 41); online-first 2022 — issue year is correct for citation} + year = {2023}, + doi = {10.1038/s41587-022-01476-y} } -@article{li2025multivelovae, - author = {Li, Chen and Gu, Yichen and Virgilio, Maria C. and Lee, Kyung Hoi and Collins, Kathleen L. and Welch, Joshua D.}, - title = {Inferring differential dynamics from multi-lineage, multi-omic, and multi-sample single-cell data with MultiVeloVAE}, +@article{ref04, + author = {Li, C and Gu, Y and Virgilio, MC and Lee, KH and Collins, KL and Welch, JD}, + title = {Inferring differential dynamics from multi-lineage, multi-omic, and multi-sample single-cell data with MultiVeloVAE}, journal = {Nature Communications}, - volume = {16}, - pages = {11505}, - year = {2025}, - doi = {10.1038/s41467-025-66287-6} + year = {2025}, + doi = {10.1038/s41467-025-66287-6} } -@article{hong2025moflow, - author = {Hong, A. and Lee, S. and Kim, K.}, - title = {Multi-omic relay velocity modeling uncovers dynamic chromatin-transcription regulation across cell states}, +@article{ref05, + author = {Hong, A and Lee, S and Kim, K}, + title = {Multi-omic relay velocity modeling uncovers dynamic chromatin-transcription regulation across cell states}, journal = {Nature Communications}, - volume = {17}, - pages = {566}, - year = {2025}, - doi = {10.1038/s41467-025-67259-6}, - note = {PMID 41457082; early-access into the 2026 volume} + year = {2025}, + doi = {10.1038/s41467-025-67259-6} } -@article{elkazwini2026crakvelo, - author = {El Kazwini, N. and Gao, M. and Kouadri Boudjelthia, I. and Cai, F. and Huang, Y. and Sanguinetti, G.}, - title = {CRAK-Velo: chromatin accessibility kinetics integration improves RNA velocity estimation}, +@article{ref06, + author = {El Kazwini, N and Gao, M and Kouadri Boudjelthia, I and Cai, F and Huang, Y and Sanguinetti, G}, + title = {CRAK-Velo: chromatin accessibility kinetics integration improves RNA velocity estimation}, journal = {Genome Biology}, - volume = {27}, - number = {1}, - year = {2026}, - doi = {10.1186/s13059-026-04086-y}, - note = {PMID 42087173; bioRxiv 2024.09.12.612736} + year = {2026}, + doi = {10.1186/s13059-026-04086-y} } -@article{archvelo2026, - author = {Avdeeva, Maria and Walker, Sarah K. and van der Veeken, Joris and Rudensky, Alexander Y. and Pritykin, Yuri}, - title = {ArchVelo: archetypal velocity modeling for single-cell multi-omic trajectories}, +@article{ref07, + author = {ArchVelo: archetypal velocity modeling for single-cell multi-omic trajectories}, journal = {Nature Communications}, - year = {2026}, - doi = {10.1038/s41467-026-74000-4}, - note = {Author list confirmed via CrossRef (DOI 10.1038/s41467-026-74000-4)} + year = {2026}, + doi = {10.1038/s41467-026-74000-4} } -@article{su2023sckinetics, - author = {Burdziak, Cassandra and Zhao, Chujun Julia and Haviv, Doron and Alonso-Curbelo, Direna and Lowe, Scott W. and Pe'er, Dana}, - title = {scKINETICS: inference of regulatory velocity with single-cell transcriptomics data}, +@article{ref08, + author = {Su, M and others}, + title = {scKINETICS: inference of regulatory velocity with single-cell transcriptomics data}, journal = {Bioinformatics}, - volume = {39}, - number = {Suppl 1}, - pages = {i394--i403}, - year = {2023}, - note = {PMC10311321; author corrected via CrossRef (was 'Su' — first author is Burdziak). citekey su2023 kept as internal label} + year = {2023} } -@article{gayoso2024velovi, - author = {Gayoso, Adam and Weiler, Philipp and Lotfollahi, Mohammad and others}, - title = {Deep generative modeling of transcriptional dynamics for RNA velocity analysis in single cells}, +@article{ref09, + author = {Gayoso, A and Weiler, P and Lotfollahi, M and others}, + title = {Deep generative modeling of transcriptional dynamics for RNA velocity analysis in single cells}, journal = {Nature Methods}, - volume = {21}, - pages = {50--59}, - year = {2024}, - doi = {10.1038/s41592-023-01994-w}, - note = {year=2024 is the issue year (Nat Methods vol 21); online-first 2023 — issue year is correct for citation} + year = {2024}, + doi = {10.1038/s41592-023-01994-w} } -@article{gu2026velovae, - author = {Gu, Yichen and others}, - title = {Bayesian inference of RNA velocity incorporating timepoints, lineage bifurcations, and count data (veloVAE)}, +@article{ref10, + author = {Gu, Y and others}, + title = {Bayesian inference of RNA velocity incorporating timepoints, lineage bifurcations, and count data (veloVAE)}, journal = {PLOS Computational Biology}, - volume = {22}, - number = {3}, - pages = {e1014060}, - year = {2026}, - doi = {10.1371/journal.pcbi.1014060}, - note = {Full author list to confirm; distinct from MultiVeloVAE} + year = {2026}, + doi = {10.1371/journal.pcbi.1014060} +} + +@article{ref11, + author = {Qiao, C and Huang, Y}, + title = {Representation learning of RNA velocity reveals robust cell transitions}, + journal = {Proceedings of the National Academy of Sciences}, + year = {2021}, + doi = {10.1073/pnas.2105859118} +} + +@article{ref12, + author = {Gao, M and Qiao, C and Huang, Y}, + title = {UniTVelo: temporally unified RNA velocity reinforces single-cell trajectory inference}, + journal = {Nature Communications}, + year = {2022}, + doi = {10.1038/s41467-022-34188-7} +} + +@article{ref13, + author = {Li, S and Zhang, P and Chen, W and others}, + title = {A relay velocity model infers cell-dependent RNA velocity}, + journal = {Nature Biotechnology}, + year = {2023}, + doi = {10.1038/s41587-023-01728-5} +} + +@article{ref14, + author = {Qiu, X and Zhang, Y and Martin-Rufino, JD and others}, + title = {Mapping transcriptomic vector fields of single cells}, + journal = {Cell}, + year = {2022}, + doi = {10.1016/j.cell.2021.12.045} +} + +@article{ref15, + author = {Velten, L and Haas, SF and Raffel, S and others}, + title = {Human haematopoietic stem cell lineage commitment is a continuous process}, + journal = {Nature Cell Biology}, + year = {2017}, + doi = {10.1038/ncb3493} +} + +@article{ref16, + author = {Laurenti, E and Göttgens, B}, + title = {From haematopoietic stem cells to complex differentiation landscapes}, + journal = {Nature}, + year = {2018}, + doi = {10.1038/nature25022} +} + +@article{ref17, + author = {Buenrostro, JD and Corces, MR and Lareau, CA and others}, + title = {Integrated Single-Cell Analysis Maps the Continuous Regulatory Landscape of Human Hematopoietic Differentiation}, + journal = {Cell}, + year = {2018}, + doi = {10.1016/j.cell.2018.03.074} +} + +@article{ref18, + author = {Ma, S and Zhang, B and LaFave, LM and others}, + title = {Chromatin Potential Identified by Shared Single-Cell Profiling of RNA and Chromatin}, + journal = {Cell}, + year = {2020}, + doi = {10.1016/j.cell.2020.09.056} +} + +@article{ref19, + author = {Trevino, AE and Müller, F and Andersen, J and others}, + title = {Chromatin and gene-regulatory dynamics of the developing human cerebral cortex at single-cell resolution}, + journal = {Cell}, + year = {2021}, + doi = {10.1016/j.cell.2021.07.039} +} + +@article{ref20, + author = {Zheng, SC and Stein-O’Brien, G and Boukas, L and Goff, LA and Hansen, KD}, + title = {Pumping the brakes on RNA velocity by understanding and interpreting RNA velocity estimates}, + journal = {Genome Biology}, + year = {2023}, + doi = {10.1186/s13059-023-03065-x} } -@article{bergen2021challenges, - author = {Bergen, Volker and Soldatov, Ruslan A. and Kharchenko, Peter V. and Theis, Fabian J.}, - title = {RNA velocity --- current challenges and future perspectives}, +@article{ref21, + author = {Bergen, V and Soldatov, RA and Kharchenko, PV and Theis, FJ}, + title = {RNA velocity — current challenges and future perspectives}, journal = {Molecular Systems Biology}, - volume = {17}, - number = {8}, - pages = {e10282}, - year = {2021}, - doi = {10.15252/msb.202110282} + year = {2021}, + doi = {10.15252/msb.202110282} } -@article{gorin2022unraveled, - author = {Gorin, Gennady and Fang, Meichen and Chari, Tara and Pachter, Lior}, - title = {RNA velocity unraveled}, +@article{ref22, + author = {Gorin, G and Fang, M and Chari, T and Pachter, L}, + title = {RNA velocity unraveled}, journal = {PLOS Computational Biology}, - volume = {18}, - number = {9}, - pages = {e1010492}, - year = {2022}, - doi = {10.1371/journal.pcbi.1010492} + year = {2022}, + doi = {10.1371/journal.pcbi.1010492} } -@article{marotlassauzaie2022reliable, - author = {Marot-Lassauzaie, Valerie and Bouman, Brigitte J. and Donaghy, Fabian D. and Demerdash, Yasmin and Essers, Marieke A. G. and Haghverdi, Laleh}, - title = {Towards reliable quantification of cell state velocities}, +@article{ref23, + author = {Marot-Lassauzaie, V and Bouman, BJ and Donaghy, FD and Demerdash, Y and Essers, MAG and Haghverdi, L}, + title = {Towards reliable quantification of cell state velocities}, journal = {PLOS Computational Biology}, - volume = {18}, - number = {9}, - pages = {e1010031}, - year = {2022}, - doi = {10.1371/journal.pcbi.1010031}, - note = {PMC9550177; author order to confirm} + year = {2022}, + doi = {10.1371/journal.pcbi.1010031} } -@article{benchmark17studies2026, - author = {Luo, Ya and Ren, Jun and Yang, Qian and Zhou, Ying and You, Zhiyu and Li, Qiyuan}, - title = {Benchmarking RNA velocity methods across 17 independent studies}, - journal = {Cell Reports Methods}, - year = {2026}, - note = {S2667-2375(26)00067-6; bioRxiv 2025.08.02.668272; author list/DOI to confirm at proof} +@article{ref24, + author = {Sonrel, A and Luetge, A and Soneson, C and others}, + title = {Meta-analysis of (single-cell method) benchmarks reveals the need for extensibility and interoperability}, + journal = {Genome Biology}, + year = {2023}, + doi = {10.1186/s13059-023-02962-5} } -@misc{benchmark29algorithms2026, - author = {Huang, Kexin and Zhou, Yu and Wang, Tiangang and Li, Xiao and Zhao, Xinlong and Liu, Xi and Huang, Liyu and Zhou, Xiaobo and Liu, Jiajia}, - title = {Benchmarking algorithms for RNA velocity inference}, - howpublished = {Preprint, doi:10.64898/2026.01.03.697314}, - year = {2026}, - doi = {10.64898/2026.01.03.697314}, - note = {Author list + DOI confirmed via CrossRef — prefix is 10.64898 (not bioRxiv 10.1101, which 404s)} +@article{ref25, + author = {Luo, Y and Ren, J and Yang, Q and You, Z and Zhou, Y and Qin, Q and Li, Q}, + title = {Benchmarking RNA velocity methods across 17 independent studies}, + journal = {Cell Reports Methods}, + year = {2026}, + doi = {10.1016/j.crmeth.2026.101367} } -@article{ma2020chromatinpotential, - author = {Ma, Sai and Zhang, Bing and LaFave, Lindsay M. and Earl, Andrew S. and Chiang, Zachary and Hu, Yan and Ding, Jiarui and Brack, Andrew and Kartha, Vinay K. and Tay, Tristan and Law, Travis and Lareau, Caleb and Hsu, Ya-Chieh and Regev, Aviv and Buenrostro, Jason D.}, - title = {Chromatin Potential Identified by Shared Single-Cell Profiling of RNA and Chromatin}, - journal = {Cell}, - volume = {183}, - number = {4}, - pages = {1103--1116.e20}, - year = {2020}, - doi = {10.1016/j.cell.2020.09.056} +@article{ref26, + author = {Huang, K and Zhou, Y and Wang, T and Li, X and Zhao, X and Liu, X and Huang, L and Zhou, X and Liu, J}, + title = {Benchmarking algorithms for RNA velocity inference. bioRxiv 2026.01.03.697314}, + year = {2026}, + doi = {10.64898/2026.01.03.697314} } -@article{trevino2021cortex, - author = {Trevino, Alexandro E. and M{\"u}ller, Fabian and Andersen, Jimena and others}, - title = {Chromatin and gene-regulatory dynamics of the developing human cerebral cortex at single-cell resolution}, - journal = {Cell}, - volume = {184}, - number = {19}, - pages = {5053--5069.e23}, - year = {2021}, - doi = {10.1016/j.cell.2021.07.039}, - note = {GSE162170} +@article{ref27, + author = {Wu, Y and Kong, C and Liao, X and Lin, Z and Sun, X and Liu, J}, + title = {Comprehensive benchmarking of RNA velocity methods across single-cell datasets}, + journal = {Genome Biology}, + year = {2026}, + doi = {10.1186/s13059-026-04182-z} } -@article{zhang2024consensusvelo, - author = {Zhang, Huizi and Bochkina, Natalia and Wade, Sara}, - title = {Quantifying uncertainty in RNA velocity (ConsensusVelo)}, - journal = {Biometrics}, - year = {2026}, - note = {bioRxiv 2024.05.14.594102 (2024); Biometrics 82(1) ujag018, in press; doi:10.1101/2024.05.14.594102; full author list/final venue to confirm — closest prior art to the profile-likelihood section} +@article{ref28, + author = {Ancheta, S and Dorman, L and Le Treut, G and others}, + title = {Challenges and progress in RNA velocity: comparative analysis across multiple biological contexts}, + journal = {PLoS Computational Biology}, + year = {2026}, + doi = {10.1371/journal.pcbi.1014303} } -@article{gu2025profilelikelihood, - author = {Gu and others}, - title = {Profile-likelihood identifiability analysis of single-cell transcription (telegraph) kinetics}, - journal = {Bioinformatics}, - volume = {41}, - number = {11}, - pages = {btaf581}, - year = {2025}, - doi = {10.1093/bioinformatics/btaf581}, - note = {Exact title/author list to confirm; distinct from the veloVAE Gu et al.} +@article{ref29, + author = {Barile, M and Imaz-Rosshandler, I and Inzani, I and others}, + title = {Coordinated changes in gene expression kinetics underlie both mouse and human erythroid maturation}, + journal = {Genome Biology}, + year = {2021}, + doi = {10.1186/s13059-021-02414-y} } -@misc{wang2025sloppy, - author = {Wang, Yuxuan and Ying, Junda and Xiao, He and Huang, Miao and Zhang, Lei and Wang, Weikang}, - title = {Sloppiness and Action Constraint in Cell State Transitions: Are Single Cells Sloppy?}, - howpublished = {bioRxiv 2025.12.31.697145 (v2)}, - year = {2025}, - note = {Author list to confirm; Fisher analysis on cell-state Gaussian coordinates, methodological analog only} +@article{ref30, + author = {Schwalb, B and Michel, M and Zacher, B and others}, + title = {TT-seq maps the human transient transcriptome}, + journal = {Science}, + year = {2016}, + doi = {10.1126/science.aad9841} } -@misc{bayvel2025, - author = {Sabbioni, Elena and Bibbona, Enrico and Mastrantonio, Gianluca and Sanguinetti, Guido}, - title = {BayVel: A Bayesian Framework for RNA Velocity Estimation in Single-Cell Transcriptomics}, - howpublished = {arXiv:2505.03083}, - year = {2025}, - note = {Author list confirmed via arXiv 2505.03083} +@article{ref31, + author = {Herzog, VA and Reichholf, B and Neumann, T and others}, + title = {Thiol-linked alkylation of RNA to assess expression dynamics}, + journal = {Nature Methods}, + year = {2017}, + doi = {10.1038/nmeth.4435} } -@article{begley2012raise, - author = {Begley, C. Glenn and Ellis, Lee M.}, +@article{ref32, + author = {Begley, CG and Ellis, LM}, title = {Raise standards for preclinical cancer research}, - journal = {Nature}, volume = {483}, pages = {531--533}, year = {2012}, + journal = {Nature}, + year = {2012}, doi = {10.1038/483531a} } -@article{prinz2011believe, - author = {Prinz, Florian and Schlange, Thomas and Asadullah, Khusru}, +@article{ref33, + author = {Prinz, F and Schlange, T and Asadullah, K}, title = {Believe it or not: how much can we rely on published data on potential drug targets?}, - journal = {Nature Reviews Drug Discovery}, volume = {10}, pages = {712}, year = {2011}, + journal = {Nature Reviews Drug Discovery}, + year = {2011}, doi = {10.1038/nrd3439-c1} } -@article{osc2015estimating, - author = {{Open Science Collaboration}}, +@article{ref34, + author = {Open Science Collaboration}, title = {Estimating the reproducibility of psychological science}, - journal = {Science}, volume = {349}, pages = {aac4716}, year = {2015}, + journal = {Science}, + year = {2015}, doi = {10.1126/science.aac4716} } -@article{baker2016reproducibility, - author = {Baker, Monya}, +@article{ref35, + author = {Baker, M}, title = {1,500 scientists lift the lid on reproducibility}, - journal = {Nature}, volume = {533}, pages = {452--454}, year = {2016}, + journal = {Nature}, + year = {2016}, doi = {10.1038/533452a} } -@article{errington2021investigating, - author = {Errington, Timothy M. and Mathur, Maya and Denis, Matthias and others}, +@article{ref36, + author = {Errington, TM and Mathur, M and Denis, M and others}, title = {Investigating the replicability of preclinical cancer biology}, - journal = {eLife}, volume = {10}, pages = {e71601}, year = {2021}, + journal = {eLife}, + year = {2021}, doi = {10.7554/eLife.71601} } -@article{ioannidis2005why, - author = {Ioannidis, John P. A.}, +@article{ref37, + author = {Ioannidis, JPA}, title = {Why most published research findings are false}, - journal = {PLoS Medicine}, volume = {2}, number = {8}, pages = {e124}, year = {2005}, + journal = {PLoS Medicine}, + year = {2005}, doi = {10.1371/journal.pmed.0020124} } -@article{ioannidis2009repeatability, - author = {Ioannidis, John P. A. and Allison, David B. and Ball, Catherine A. and others}, +@article{ref38, + author = {Ioannidis, JPA and Allison, DB and Ball, CA and others}, title = {Repeatability of published microarray gene expression analyses}, - journal = {Nature Genetics}, volume = {41}, pages = {149--155}, year = {2009}, + journal = {Nature Genetics}, + year = {2009}, doi = {10.1038/ng.295} } + +@article{ref39, + author = {Weber, LM and Saelens, W and Cannoodt, R and others}, + title = {Essential guidelines for computational method benchmarking}, + journal = {Genome Biology}, + year = {2019}, + doi = {10.1186/s13059-019-1738-8} +} + +@article{ref40, + author = {Luecken, MD and Büttner, M and Chaichoompu, K and others}, + title = {Benchmarking atlas-level data integration in single-cell genomics}, + journal = {Nature Methods}, + year = {2021}, + doi = {10.1038/s41592-021-01336-8} +} + +@article{ref41, + author = {Zhang and others}, + title = {Quantifying uncertainty in RNA velocity (ConsensusVelo). bioRxiv 2024.05.14.594102 (2024);}, + journal = {Biometrics}, + year = {2024}, + doi = {10.1101/2024.05.14.594102} +} + +@article{ref42, + author = {Gutenkunst, RN and Waterfall, JJ and Casey, FP and Brown, KS and Myers, CR and Sethna, JP}, + title = {Universally Sloppy Parameter Sensitivities in Systems Biology Models}, + journal = {PLoS Computational Biology}, + year = {2007}, + doi = {10.1371/journal.pcbi.0030189} +} + +@article{ref43, + author = {Transtrum, MK and Machta, BB and Sethna, JP}, + title = {Geometry of nonlinear least squares with applications to sloppy models and optimization}, + journal = {Physical Review E}, + year = {2011}, + doi = {10.1103/physreve.83.036701} +} + +@article{ref44, + author = {Cowland, JB and Borregaard, N}, + title = {Granulopoiesis and granules of human neutrophils}, + journal = {Immunological Reviews}, + year = {2016}, + doi = {10.1111/imr.12440} +} + +@article{ref45, + author = {Borregaard, N and Sørensen, OE and Theilgaard-Mönch, K}, + title = {Neutrophil granules: a library of innate immunity proteins}, + journal = {Trends in Immunology}, + year = {2007}, + doi = {10.1016/j.it.2007.06.002} +} + +@article{ref46, + author = {Gillespie, M and Jassal, B and Stephan, R and others}, + title = {The reactome pathway knowledgebase 2022}, + journal = {Nucleic Acids Research}, + year = {2021}, + doi = {10.1093/nar/gkab1028} +} + +@article{ref47, + author = {Schep, AN and Wu, B and Buenrostro, JD and Greenleaf, WJ}, + title = {chromVAR: inferring transcription-factor-associated accessibility from single-cell epigenomic data}, + journal = {Nature Methods}, + year = {2017}, + doi = {10.1038/nmeth.4401} +} + +@article{ref48, + author = {Schuirmann, DJ}, + title = {A comparison of the Two One-Sided Tests Procedure and the Power Approach for assessing the equivalence of average bioavailability}, + journal = {Journal of Pharmacokinetics and Biopharmaceutics}, + year = {1987}, + doi = {10.1007/bf01068419} +} + +@article{ref49, + author = {Nosek, BA and Ebersole, CR and DeHaven, AC and Mellor, DT}, + title = {The preregistration revolution}, + journal = {Proceedings of the National Academy of Sciences}, + year = {2018}, + doi = {10.1073/pnas.1708274114} +} + +@article{ref50, + author = {Todorovski, I and Tsang, MJ and Feran, B and others}, + title = {RNA kinetics influence the response to transcriptional perturbation in leukaemia cell lines}, + journal = {NAR Cancer}, + year = {2024}, + doi = {10.1093/narcan/zcae039} +} + +@article{ref51, + author = {Raue, A and Kreutz, C and Maiwald, T and others}, + title = {Structural and practical identifiability analysis of partially observed dynamical models by exploiting the profile likelihood}, + journal = {Bioinformatics}, + year = {2009}, + doi = {10.1093/bioinformatics/btp358} +} + +@article{ref52, + author = {Kreutz, C and Raue, A and Kaschek, D and Timmer, J}, + title = {Profile likelihood in systems biology}, + journal = {The FEBS Journal}, + year = {2013}, + doi = {10.1111/febs.12276} +} + +@article{ref53, + author = {Villaverde, AF and Barreiro, A and Papachristodoulou, A}, + title = {Structural Identifiability of Dynamic Systems Biology Models}, + journal = {PLOS Computational Biology}, + year = {2016}, + doi = {10.1371/journal.pcbi.1005153} +} + +@article{ref54, + author = {Lederer, AR and Leonardi, M and Talamanca, L and others}, + title = {Statistical inference with a manifold-constrained RNA velocity model uncovers cell cycle speed modulations}, + journal = {Nature Methods}, + year = {2024}, + doi = {10.1038/s41592-024-02471-8} +} + +@article{ref55, + author = {Gu and others}, + title = {Profile-likelihood identifiability analysis of single-cell transcription (telegraph) kinetics}, + journal = {Bioinformatics}, + year = {2025}, + doi = {10.1093/bioinformatics/btaf581} +} + +@article{ref56, + author = {Wang}, + title = {Sloppiness and Action Constraint in Cell State Transitions: Are Single Cells Sloppy? bioRxiv 2025.12.31.697145 (v2, 2025). [Methodological analog on cell-state Gaussian coordinates.]} +} + +@article{ref57, + author = {BayVel: A Bayesian Framework for RNA Velocity Estimation in Single-Cell Transcriptomics}, + title = {arXiv:2505.03083}, + year = {2025} +} + +@article{ref58, + author = {Battich, N and Beumer, J and de Barbanson, B and others}, + title = {Sequencing metabolically labeled transcripts in single cells reveals mRNA turnover strategies}, + journal = {Science}, + year = {2020}, + doi = {10.1126/science.aax3072} +} + +@article{ref59, + author = {Cao, J and Zhou, W and Steemers, F and Trapnell, C and Shendure, J}, + title = {Sci-fate characterizes the dynamics of gene expression in single cells}, + journal = {Nature Biotechnology}, + year = {2020}, + doi = {10.1038/s41587-020-0480-9} +} + +@article{ref60, + author = {Lange, M and Bergen, V and Klein, M and others}, + title = {CellRank for directed single-cell fate mapping}, + journal = {Nature Methods}, + year = {2022}, + doi = {10.1038/s41592-021-01346-6} +} + +@article{ref61, + author = {Weinreb, C and Rodriguez-Fraticelli, A and Camargo, FD and Klein, AM}, + title = {Lineage tracing on transcriptional landscapes links state to fate during differentiation}, + journal = {Science}, + year = {2020}, + doi = {10.1126/science.aaw3381} +} + +@article{ref62, + author = {Kaminow, B and Yunusov, D and Dobin, A}, + title = {STARsolo: accurate, fast and versatile mapping/quantification of single-cell and single-nucleus RNA-seq data}, + year = {2021}, + doi = {10.1101/2021.05.05.442755} +} + +@article{ref63, + author = {Granja, JM and Corces, MR and Pierce, SE and others}, + title = {ArchR is a scalable software package for integrative single-cell chromatin accessibility analysis}, + journal = {Nature Genetics}, + year = {2021}, + doi = {10.1038/s41588-021-00790-6} +} + +@article{ref64, + author = {Stuart, T and Srivastava, A and Madad, S and Lareau, CA and Satija, R}, + title = {Single-cell chromatin state analysis with Signac}, + journal = {Nature Methods}, + year = {2021}, + doi = {10.1038/s41592-021-01282-5} +} + +@article{ref65, + author = {Hao, Y and Hao, S and Andersen-Nissen, E and others}, + title = {Integrated analysis of multimodal single-cell data}, + journal = {Cell}, + year = {2021}, + doi = {10.1016/j.cell.2021.04.048} +} + +@article{ref66, + author = {Bergen, V and Lange, M and Peidli, S and Wolf, FA and Theis, FJ}, + title = {Generalizing RNA velocity to transient cell states through dynamical modeling}, + journal = {Nature Biotechnology}, + year = {2020}, + doi = {10.1038/s41587-020-0591-3} +} + +@article{ref67, + author = {Benjamini, Y and Hochberg, Y}, + title = {Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing}, + journal = {Journal of the Royal Statistical Society Series B: Statistical Methodology}, + year = {1995}, + doi = {10.1111/j.2517-6161.1995.tb02031.x} +} + +@article{ref68, + author = {Chen, EY and Tan, CM and Kou, Y and others}, + title = {Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool}, + journal = {BMC Bioinformatics}, + year = {2013}, + doi = {10.1186/1471-2105-14-128} +} + +@article{ref69, + author = {Kuleshov, MV and Jones, MR and Rouillard, AD and others}, + title = {Enrichr: a comprehensive gene set enrichment analysis web server 2016 update}, + journal = {Nucleic Acids Research}, + year = {2016}, + doi = {10.1093/nar/gkw377} +} + +@article{ref70, + author = {Xie, Z and Bailey, A and Kuleshov, MV and others}, + title = {Gene Set Knowledge Discovery with Enrichr}, + journal = {Current Protocols}, + year = {2021}, + doi = {10.1002/cpz1.90} +} + +@article{ref71, + author = {Wolock, SL and Lopez, R and Klein, AM}, + title = {Scrublet: Computational Identification of Cell Doublets in Single-Cell Transcriptomic Data}, + journal = {Cell Systems}, + year = {2019}, + doi = {10.1016/j.cels.2018.11.005} +} + diff --git a/pipeline/hspc-velocity-benchmark/scripts/build_refs_bib.py b/pipeline/hspc-velocity-benchmark/scripts/build_refs_bib.py new file mode 100644 index 0000000..94bf68e --- /dev/null +++ b/pipeline/hspc-velocity-benchmark/scripts/build_refs_bib.py @@ -0,0 +1,111 @@ +#!/usr/bin/env python3 +"""build_refs_bib.py — draft_v2.md의 손수 작성 참고문헌 목록([1]~[N])을 refs.bib로 재생성. + +왜 필요한가 +----------- +manuscript/README.md는 refs.bib를 "paper_analysis/*//*.bib에서 모은다"고 +적었으나, method 논문(scVelo·MultiVelo·chromVAR…)은 paper_analysis .bib가 없어 그 과정은 +구조적으로 71개를 만들 수 없다(26에서 멈춤, BIOP01-52 jamie 지적). 진리원천은 본문 +참고문헌 목록이며 jamie가 in-text [1]~[N]과 1:1(고아·누락 0)임을 확인했다. 그래서 +목록을 파싱해 refs.bib를 만든다. 인용은 수동 번호 [27] 방식이라 cite key는 [N]에 +맞출 필요가 없고, verify_citations.py가 쓰는 first_author·year·doi만 정확하면 된다. + +형식 가정(목록 한 줄): + [N] AUTHORS. TITLE. *JOURNAL* VOL(ISSUE), PAGES (YEAR). doi:DOI. [optional note] + +사용: + python3 build_refs_bib.py draft_v2.md > refs.bib + # 검증: python3 bib_to_cites.py refs.bib | python3 verify_citations.py /dev/stdin +""" +import sys, re + + +def fmt_authors(s): + """프로즈 저자열('Wu Y, Kong C, et al')을 BibTeX('Wu, Y and Kong, C and others')로. + + verify_citations는 first author family만 비교하므로 마지막 공백토큰=이니셜로 보고 + 'Family, Initials'로 뒤집는다. 다어절 성(La Manno G)도 마지막 토큰만 이니셜로 처리. + """ + s = re.sub(r'\bet\s+al\.?', 'et al', s) + s = s.replace(' et al', ', et al') # 콤마 구분 보장 + out = [] + for p in [x.strip() for x in s.split(',') if x.strip()]: + if p.lower() == 'et al': + out.append('others'); continue + toks = p.split() + # 마지막 토큰이 이니셜형(짧은 대문자 시작)일 때만 'Family, Initials'로 분리. + # 'Open Science Collaboration' 같은 단체저자는 통째로 둔다. + if len(toks) >= 2 and re.match(r'^[A-Z][A-Za-z.]{0,3}$', toks[-1]): + out.append("%s, %s" % (' '.join(toks[:-1]), toks[-1])) + else: + out.append(p) # 성 하나 또는 단체저자(Zhang, Wang, Open Science Collaboration) + return ' and '.join(out) + + +def parse_reference_list(md_text): + entries = [] + for line in md_text.splitlines(): + m = re.match(r'^\[(\d+)\]\s+(.*)$', line.strip()) + if not m: + continue + num, rest = int(m.group(1)), m.group(2).strip() + # DOI: doi:10.xxx (공백·대괄호·괄호 앞에서 끊음 — 뒤따르는 (GSE…)·[note] 배제) + doi = '' + md = re.search(r'doi:\s*(10\.[^\s\[\]()]+)', rest) + if md: + doi = md.group(1).rstrip('.') + # year: 괄호 안 4자리 중 마지막(발행연도가 보통 doi 앞 괄호) + years = re.findall(r'\((\d{4})\)', rest) + year = years[-1] if years else '' + # authors: 첫 '. ' 앞까지 (제목 시작 전) + idx = rest.find('. ') + authors = rest[:idx].strip() if idx > 0 else rest + after = rest[idx + 2:] if idx > 0 else '' + # journal: *…* 이탤릭 + journal = '' + mj = re.search(r'\*([^*]+)\*', rest) + if mj: + journal = mj.group(1).strip() + # title: authors 뒤 ~ 저널(*) 또는 연도 괄호 앞 + title = after + if mj: + title = after[:after.find('*')] + else: + title = re.split(r'\s*\(\d{4}\)', after)[0] + title = title.rstrip(' .').strip() + # first author: 첫 콤마/ et al/ and 앞 + first_author = re.split(r',| et al| and ', authors)[0].strip() + entries.append(dict(num=num, authors=fmt_authors(authors), first_author=first_author, + title=title, journal=journal, year=year, doi=doi)) + return entries + + +def to_bibtex(e): + key = "ref%02d" % e['num'] + fields = [('author', e['authors']), ('title', e['title']), + ('journal', e['journal']), ('year', e['year']), ('doi', e['doi'])] + body = ",\n".join(" %s = {%s}" % (k, v) for k, v in fields if v) + return "@article{%s,\n%s\n}" % (key, body) + + +def main(): + if len(sys.argv) < 2: + sys.exit("usage: build_refs_bib.py ") + text = open(sys.argv[1], encoding='utf-8').read() + entries = parse_reference_list(text) + # 무결성: 번호가 1..N 연속인지, doi 결측이 몇 개인지 stderr로 보고 + nums = [e['num'] for e in entries] + expected = list(range(1, (max(nums) if nums else 0) + 1)) + missing = sorted(set(expected) - set(nums)) + no_doi = [e['num'] for e in entries if not e['doi']] + print("# generated from %s — %d entries (refs [1]..[%d])" + % (sys.argv[1].split('/')[-1], len(entries), max(nums) if nums else 0)) + for e in entries: + print(to_bibtex(e)) + print() + sys.stderr.write("parsed %d entries; missing nums=%s; entries without doi=%s\n" + % (len(entries), missing or "none", no_doi or "none")) + + +if __name__ == '__main__': + main()