Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions PERFORMANCE.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ This guide provides repeatable checks for local model latency, response quality,
```bash
cd sidecar
source .venv/bin/activate
pytest -q tests/test_v2_pipeline.py tests/test_local_inference_sanitize.py tests/test_memory_service_digest.py
pytest -q tests/v2/chat/test_v2_pipeline.py tests/test_local_inference_sanitize.py tests/test_memory_service_digest.py
```

## 2) End-to-End API Regression
Expand Down Expand Up @@ -47,7 +47,7 @@ Run week exact/filter tests:
```bash
cd sidecar
source .venv/bin/activate
pytest -q tests/test_v2_pipeline.py -k "week_exact_filter or requested_weeks"
pytest -q tests/v2/chat/test_v2_pipeline.py -k "week_exact_filter or requested_weeks"
```

## 5) Manual Comparison Matrix (Recommended)
Expand Down
11 changes: 9 additions & 2 deletions README.ja.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,11 @@ PLOSは SwiftUI アプリと Python サイドカー(FastAPI)を組み合わせ
- ハードウェアに応じたモデルカタログ
- 一般会話 Direct-First 応答ポリシー

## ランタイム状況
- ローカル推論バックエンド: `MLX`, `llama.cpp`(主要経路)
- 外部プロバイダ(任意): `OpenAI`, `Anthropic`
- Ollama: 現在の main ブランチでは一次ランタイムとして未統合

## アーキテクチャ
```mermaid
graph TD
Expand Down Expand Up @@ -67,6 +72,7 @@ brew install tesseract poppler
- Xcodeで `PLOS.xcodeproj` を開く
- `PLOS` ターゲットを実行
- サイドカーのライフサイクルはアプリが自動管理
- 初回確認: チャット画面が開き、簡単なプロンプトに応答すれば正常起動

## サイドカー単体起動(開発)
```bash
Expand All @@ -77,6 +83,8 @@ export LOCAL_AI_DATA_DIR="$(pwd)/data"
uvicorn local_ai_core.main:create_app --factory --host 127.0.0.1 --port 8787
```

`dev-token` はローカル開発用のサンプル値です。

## モデル推奨レンジ(実運用目安)
- 16GB: 7B/8B中心、12B/14Bは限定的
- 64GB+: 20B/70Bクラス
Expand All @@ -93,7 +101,7 @@ pytest -q

### 重点回帰
```bash
pytest -q tests/test_v2_pipeline.py tests/test_local_inference_sanitize.py tests/test_memory_service_digest.py
pytest -q tests/v2/chat/test_v2_pipeline.py tests/test_local_inference_sanitize.py tests/test_memory_service_digest.py
```

### アプリテスト
Expand All @@ -108,7 +116,6 @@ xcodebuild \
## ドキュメント
- [CONTRIBUTING.ja.md](CONTRIBUTING.ja.md)
- [PERFORMANCE.ja.md](PERFORMANCE.ja.md)
- [PLUGIN_RUNTIME_SPEC.md](PLUGIN_RUNTIME_SPEC.md)
- [CHANGELOG.ja.md](CHANGELOG.ja.md)
- [CHANGELOG.en.md](CHANGELOG.en.md)
- [CHANGELOG.ko.md](CHANGELOG.ko.md)
Expand Down
11 changes: 9 additions & 2 deletions README.ko.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,11 @@ PLOS는 SwiftUI 앱 + Python 사이드카(FastAPI) 구조로, 대화/검색/메
- 하드웨어 등급 기반 모델 카탈로그
- 일반 대화 Direct-First 응답 정책

## 런타임 상태
- 로컬 추론 백엔드: `MLX`, `llama.cpp` (기본 경로)
- 외부 프로바이더(선택): `OpenAI`, `Anthropic`
- Ollama: 현재 main 브랜치에서 1급 런타임으로 통합되어 있지 않음

## 아키텍처
```mermaid
graph TD
Expand Down Expand Up @@ -67,6 +72,7 @@ brew install tesseract poppler
- Xcode에서 `PLOS.xcodeproj` 열기
- `PLOS` 타깃 실행
- 앱이 사이드카 라이프사이클을 자동 관리
- 첫 실행 확인: 채팅 화면이 열리고 간단한 프롬프트에 응답하면 정상 기동

## 사이드카 단독 실행(개발)
```bash
Expand All @@ -77,6 +83,8 @@ export LOCAL_AI_DATA_DIR="$(pwd)/data"
uvicorn local_ai_core.main:create_app --factory --host 127.0.0.1 --port 8787
```

`dev-token`은 로컬 개발용 예시값입니다.

## 모델 권장 구간(실사용 기준)
- 16GB: 7B/8B 중심, 12B/14B 제한적 시도
- 64GB+: 20B/70B급
Expand All @@ -93,7 +101,7 @@ pytest -q

### 핵심 회귀
```bash
pytest -q tests/test_v2_pipeline.py tests/test_local_inference_sanitize.py tests/test_memory_service_digest.py
pytest -q tests/v2/chat/test_v2_pipeline.py tests/test_local_inference_sanitize.py tests/test_memory_service_digest.py
```

### 앱 테스트
Expand All @@ -108,7 +116,6 @@ xcodebuild \
## 문서
- [CONTRIBUTING.ko.md](CONTRIBUTING.ko.md)
- [PERFORMANCE.ko.md](PERFORMANCE.ko.md)
- [PLUGIN_RUNTIME_SPEC.md](PLUGIN_RUNTIME_SPEC.md)
- [CHANGELOG.ko.md](CHANGELOG.ko.md)
- [CHANGELOG.en.md](CHANGELOG.en.md)
- [CHANGELOG.ja.md](CHANGELOG.ja.md)
Expand Down
11 changes: 9 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,11 @@ PLOS combines a native SwiftUI app and a local Python sidecar to run chat, retri
- Hardware-aware local model catalog
- Conversational direct-first response policy

## Runtime Status
- Local inference backends: `MLX` and `llama.cpp` (primary path)
- External providers (optional): `OpenAI`, `Anthropic`
- Ollama: not integrated as a first-class runtime in current main branch

## Architecture
```mermaid
graph TD
Expand Down Expand Up @@ -67,6 +72,7 @@ brew install tesseract poppler
- Open `PLOS.xcodeproj` in Xcode
- Run `PLOS` target
- Sidecar lifecycle is managed by the app
- First-run check: if the chat view loads and answers a simple prompt, startup is healthy

## Sidecar Standalone (Dev)
```bash
Expand All @@ -77,6 +83,8 @@ export LOCAL_AI_DATA_DIR="$(pwd)/data"
uvicorn local_ai_core.main:create_app --factory --host 127.0.0.1 --port 8787
```

`dev-token` is an example value for local development only.

## Model Tiers (Practical)
- 16GB: 7B/8B class, limited 12B/14B attempts
- 64GB+: 20B/70B class
Expand All @@ -93,7 +101,7 @@ pytest -q

### Focused regressions
```bash
pytest -q tests/test_v2_pipeline.py tests/test_local_inference_sanitize.py tests/test_memory_service_digest.py
pytest -q tests/v2/chat/test_v2_pipeline.py tests/test_local_inference_sanitize.py tests/test_memory_service_digest.py
```

### App tests
Expand All @@ -108,7 +116,6 @@ xcodebuild \
## Docs
- [CONTRIBUTING.md](CONTRIBUTING.md)
- [PERFORMANCE.md](PERFORMANCE.md)
- [PLUGIN_RUNTIME_SPEC.md](PLUGIN_RUNTIME_SPEC.md)
- [CHANGELOG.en.md](CHANGELOG.en.md)
- [CHANGELOG.ko.md](CHANGELOG.ko.md)
- [CHANGELOG.ja.md](CHANGELOG.ja.md)
Expand Down
91 changes: 85 additions & 6 deletions sidecar/local_ai_core/inference/parsers/conversational_logic.py
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,25 @@

# Component: conversational_logic.py
class ConversationalLogic(BaseDelegate):
@staticmethod
def _looks_meta_only_promise_answer(text: str) -> bool:
value = re.sub(r"\s+", " ", str(text or "").strip())
if not value:
return False
trimmed = re.sub(r"^(?:그럼|그러면|좋아요|알겠습니다|네|자|이제)\s*[,.! ]*", "", value)
if re.search(r"[:\n]", trimmed):
return False
if len(re.findall(r"[.!?]", trimmed)) > 1:
return False
return bool(
re.match(
r"^(?:한\s*번에\s*)?(?:바로\s*)?(?:한\s*줄로\s*|한\s*문장으로만\s*|간단히\s*|짧게\s*|본문만\s*)?"
r"(?:정리|요약|설명|답변|말씀|말해|출력|안내|추천)(?:해\s*|해드리\s*|드리\s*)?"
r"(?:볼게요|할게요|드릴게요|해드릴게요)\.?\s*$",
trimmed,
)
)

@staticmethod
def _looks_user_role_confusion_answer(text: str) -> bool:
value = str(text or "").strip()
Expand Down Expand Up @@ -42,10 +61,11 @@ def _looks_truncated_conversation_answer(text: str) -> bool:
value = str(text or "").strip()
if not value:
return True
token_count = len(re.findall(r"\S+", value))
if token_count <= 2:
if re.search(r"(?m)(?:^|[\n ])(?:\*\*)?(?:\d+\.|[-*•])\s*$", value):
return True
if re.search(r"(?:\*\*)?(?:\d+\.)\s*$", value):
return True
if token_count <= 5 and re.search(r"(까요\?|게요\.?|습니다\.?|요\.?)$", value):
if value.count("**") % 2 == 1:
return True
if re.search(r"[:;,(\[{`\-]\s*$", value):
return True
Expand Down Expand Up @@ -89,7 +109,10 @@ def _continue_conversation_once(
).strip()
if not tail:
return None
merged = f"{str(draft_answer or '').rstrip()} {tail}".strip()
base = str(draft_answer or "").rstrip()
base = re.sub(r"(?:\s*)(?:\*\*)?(?:\d+\.)\s*$", "", base).rstrip()
base = re.sub(r"(?:\s*)[:;,(\[{`\-]\s*$", "", base).rstrip()
merged = f"{base} {tail}".strip()
merged = re.sub(r"\s{2,}", " ", merged).strip()
return merged

Expand Down Expand Up @@ -163,6 +186,19 @@ def _generate_conversation_candidate(
query=query,
response_language=response_language,
)
if answer and not self._looks_meta_only_promise_answer(answer) and self._looks_truncated_conversation_answer(answer):
continued = self._continue_conversation_once(
engine=engine,
query=query,
draft_answer=answer,
response_language=response_language,
profile=profile,
mlx_model_path=mlx_model_path,
llama_model_path=llama_model_path,
max_tokens=max_tokens,
)
if continued:
answer = continued
raw_preview = re.sub(r"\s+", " ", str(raw or "")).strip()[:140]
quality_issues = self._conversation_quality_issues(
query=query,
Expand Down Expand Up @@ -215,7 +251,7 @@ def _generate_conversation_candidate(
# Even when static fallbacks are disabled, allow a model-based rewrite
# if the primary output is invalid (empty/hard issues). This avoids
# failing fast into repeated runtime-error fallback text.
rewrite_enabled = str(os.getenv("LOCAL_AI_CONVERSATION_REWRITE_ENABLED", "0") or "0").strip().lower() in {
rewrite_enabled = str(os.getenv("LOCAL_AI_CONVERSATION_REWRITE_ENABLED", "1") or "1").strip().lower() in {
"1", "true", "yes", "on"
}
allow_model_rewrite = bool(
Expand Down Expand Up @@ -248,6 +284,19 @@ def _generate_conversation_candidate(
)

if repaired_answer:
if not self._looks_meta_only_promise_answer(repaired_answer) and self._looks_truncated_conversation_answer(repaired_answer):
continued = self._continue_conversation_once(
engine=engine,
query=query,
draft_answer=repaired_answer,
response_language=response_language,
profile=profile,
mlx_model_path=mlx_model_path,
llama_model_path=llama_model_path,
max_tokens=max_tokens,
)
if continued:
repaired_answer = continued
repaired_issues = self._conversation_quality_issues(
query=query,
answer=repaired_answer,
Expand Down Expand Up @@ -440,6 +489,8 @@ def _conversation_hard_issues(self, issues: list[str]) -> list[str]:
"avoidable_clarification",
"continuation_artifact",
"comma_loop_artifact",
"truncated_answer",
"intent_restatement",
}
return [item for item in issues if item in hard]

Expand Down Expand Up @@ -644,10 +695,14 @@ def _conversation_quality_issues(self, *, query: str, answer: str, response_lang
issues.append("clarification_template_leak")
if self._looks_avoidable_clarification_answer(query=query, answer=cleaned):
issues.append("avoidable_clarification")
if re.match(r"^\s*한\s*번에\s*(?:바로\s*)?(?:본문만\s*)?(?:출력|정리|답변)\s*(?:할게요|해볼게요|드릴게요|해드릴게요)\.?\s*$", cleaned):
if self._looks_meta_only_promise_answer(cleaned):
issues.append("meta_only_ack")
if self._looks_truncated_conversation_answer(cleaned):
issues.append("truncated_answer")
if self._looks_user_role_confusion_answer(cleaned):
issues.append("role_confusion")
if self._looks_intent_restatement_answer(query=query, answer=cleaned):
issues.append("intent_restatement")

if response_language == "ko":
ko_chars = len(re.findall(r"[가-힣]", cleaned))
Expand All @@ -664,6 +719,30 @@ def _conversation_quality_issues(self, *, query: str, answer: str, response_lang
issues.append("informal_tone")
return issues

@staticmethod
def _looks_intent_restatement_answer(*, query: str, answer: str) -> bool:
q = str(query or "").strip()
a = str(answer or "").strip()
if not q or not a:
return False
lowered = a.lower()
markers = (
"원하시는 것 같",
"알고 싶으신 것 같",
"추천받고 싶으신 것 같",
"고민 중이시군요",
"찾으시는군요",
"것 같습니다",
)
if not any(marker in lowered for marker in markers):
return False
if len(re.findall(r"[.!?]", a)) > 2:
return False
substantive_cues = ("1.", "2.", "-", "*", "예를", "굽", "방법", "먼저", "다음", "뒤집", "온도", "소금", "후추")
if any(token in a for token in substantive_cues):
return False
return True

@staticmethod
def _looks_avoidable_clarification_answer(*, query: str, answer: str) -> bool:
q = str(query or "").strip().lower()
Expand Down
48 changes: 14 additions & 34 deletions sidecar/local_ai_core/inference/parsers/prompt_constructor.py
Original file line number Diff line number Diff line change
Expand Up @@ -110,11 +110,9 @@ def _conversational_prompt(
context_line = f"참고: {session_summary}\n"
return (
"자연스러운 한국어로 답변하세요.\n"
"역할 라벨/메타 설명/내부 지시문은 출력하지 마세요.\n"
"문장을 '께서는', '당신은' 같은 어색한 호칭으로 시작하지 마세요.\n"
"사용자 메시지가 질문이면 답하고, 진술이면 맥락에 맞게 자연스럽게 반응하세요.\n"
"현재 사용자 메시지에 직접 답하세요. 과거 주제를 선택지처럼 다시 나열하지 마세요.\n"
"정보가 치명적으로 부족하지 않다면 되묻기보다 바로 실행 가능한 답을 먼저 제시하세요.\n"
"현재 사용자 메시지에 바로 답하세요.\n"
"메타 설명이나 역할 라벨 없이, 본문부터 시작하세요.\n"
"정보가 아주 부족한 경우가 아니면 되묻지 말고 답을 먼저 주세요.\n"
f"{context_line}"
f"{query}\n"
)
Expand All @@ -123,10 +121,9 @@ def _conversational_prompt(
context_line = f"Context: {session_summary}\n"
return (
"Reply naturally in English.\n"
"If the user message is a statement, respond contextually instead of forcing advice.\n"
"For statement-only messages, avoid unsolicited recommendations unless the user asks.\n"
"Answer the current user message directly and do not re-list older topics as options.\n"
"Unless critical information is missing, provide a directly usable answer first instead of asking follow-up questions.\n"
"Answer the current user message directly.\n"
"Start with the answer itself, without meta commentary or role labels.\n"
"Unless critical information is missing, answer first instead of asking follow-up questions.\n"
f"{context_line}"
f"{query}\n"
)
Expand All @@ -140,27 +137,17 @@ def _conversation_rewrite_prompt(self, *, query: str, draft_answer: str, respons
draft = (source[:900] if is_code_heavy else re.sub(r"\s+", " ", source)[:260])
if response_language == "ko":
return (
"다음 초안 답변을 자연스러운 한국어 존댓말로 다시 작성해 주세요.\n"
"규칙:\n"
"- 사용자 마지막 질문에 직접 답변\n"
"- 정책/지시문/메타 문장 금지\n"
"- 역할 라벨(User/Assistant/You/A) 금지\n"
"- 같은 문장 반복 금지\n"
"- 한국어 단어 사이 띄어쓰기를 자연스럽게 반드시 적용\n"
"- 목록/마크다운 헤더/과도한 이모지 사용 금지\n"
"- 핵심 의미는 유지하고 간결하게 작성\n"
"- 코드 블록이 있으면 코드 연산자와 줄바꿈을 그대로 보존\n"
"초안 답변을 자연스러운 한국어 존댓말로 고쳐 쓰세요.\n"
"사용자 마지막 메시지에 직접 답하고, 메타 문장 없이 답변 본문부터 바로 쓰세요.\n"
"예고문만 쓰지 말고, 끊긴 문장이 있으면 끝까지 완성하세요.\n"
f"사용자 질문: {query}\n"
f"초안 답변: {draft}\n"
"최종 답변:"
)
return (
"Rewrite the draft answer naturally.\n"
"Rules:\n"
"- Directly answer the user's latest message\n"
"- No policy text, no meta commentary, no role labels\n"
"- No repeated sentences\n"
"- If code is present, preserve operators and line breaks exactly\n"
"Answer the user's latest message directly and start with the answer itself.\n"
"Do not output meta commentary or an unfinished sentence.\n"
f"User message: {query}\n"
f"Draft answer: {draft}\n"
"Final answer:"
Expand All @@ -174,16 +161,9 @@ def _korean_rewrite_prompt(self, *, query: str, draft_answer: str) -> str:
)
draft = (source[:900] if is_code_heavy else re.sub(r"\s+", " ", source)[:240])
return (
"다음 초안 답변을 자연스러운 한국어 존댓말로 다시 작성해 주세요.\n"
"규칙:\n"
"- 사용자 마지막 질문에 직접 답변\n"
"- 정책/지시문/메타 문장 금지\n"
"- 역할 라벨(User/Assistant/You/A) 금지\n"
"- 같은 문장 반복 금지\n"
"- 한국어 단어 사이 띄어쓰기를 자연스럽게 반드시 적용\n"
"- 목록/마크다운 헤더/과도한 이모지 사용 금지\n"
"- 핵심 의미는 유지하고 간결하게 작성\n"
"- 코드 블록이 있으면 코드 연산자와 줄바꿈을 그대로 보존\n"
"초안 답변을 자연스러운 한국어 존댓말로 고쳐 쓰세요.\n"
"사용자 마지막 메시지에 직접 답하고, 메타 문장 없이 답변 본문부터 바로 쓰세요.\n"
"예고문만 쓰지 말고, 끊긴 문장이 있으면 끝까지 완성하세요.\n"
f"사용자 질문: {query}\n"
f"초안 답변: {draft}\n"
"최종 답변:"
Expand Down
Loading
Loading