Skip to content

feat(logs): track normalized API client identity without leaking raw User-Agent #425

Description

@LIghtJUNction

Upstream source

The upstream PR adds request-client identification from the original User-Agent, persists a snapshot into consume/error/task logs, adds client_family filtering/statistics, and exposes client details in the usage-log UI. It changes 32 files (+877/-39), including relay/plugin/WebSocket request capture, log schema/querying, async-task snapshots and frontend filters.

Current LMM gap

Checked against current main after 5431b981fac24dd7aa56dc19c1ee76a764dfe578.

  • apps/api-go/model/log.go has no normalized client identity fields; GetAllLogs(...) / user-log/stat queries have no client_family filter.
  • Repository search finds no ClientIdentity implementation or equivalent request-scoped normalized API-client snapshot.
  • Existing user_agent handling is for login/session/audit metadata, not per-API-request usage attribution.
  • LMM already has richer usage-log UX, source/attribution analytics and custom relay/task paths, so mechanically applying the upstream 32-file patch would risk coupling unrelated product analytics and privacy behavior.

This is useful for LMM because it can distinguish Codex / Claude Code / Pi / OpenCode / SDK / browser-style callers in request diagnostics and help validate real-client onboarding, but it should remain observability only: UA is spoofable and must never become an auth/trust/routing signal.

Proposed LMM adaptation

Implement this as a small, LMM-native observability slice after upstream #7488 finishes review:

  1. Capture the original inbound User-Agent before any relay/header rewriting on HTTP relay and Responses WebSocket paths.
  2. Normalize only known explicit prefixes to a stable {family, variant, version, confidence} structure. Unknown callers remain unknown; do not infer identity from model names or arbitrary substrings.
  3. Persist normalized identity with consume/error logs and async-task terminal logs so task settlement keeps the submission-time client identity.
  4. Add server-side client_family filtering before pagination/counting and use the same condition for usage statistics.
  5. Add a compact usage-log filter/detail UI using LMM's existing components and seven-language i18n; do not copy upstream branding/gradients/product-specific visual treatment.
  6. Treat raw UA as potentially sensitive/untrusted metadata. Prefer not to persist the full raw string by default; if retained for diagnostics, sanitize control characters, bound length, keep it out of normal user-facing exports/source-attribution analytics, and document retention/visibility explicitly.
  7. Do not let client identity affect pricing, groups, model price locks, /fast, OAuth policy, routing weights, trust level, rate limits or administrator-assistant authorization.

Acceptance criteria

  • Codex CLI/desktop, Claude Code, Pi, OpenCode and common OpenAI SDK test UAs map deterministically; malformed/control-character input cannot create a false known identity.
  • Unknown UA remains unknown and cannot be promoted by sanitization side effects.
  • HTTP relay and Responses WebSocket capture the pre-rewrite identity; two sequential WS turns cannot leak identity state between requests.
  • Async tasks retain the identity from submission through terminal logging without changing billing/refund/CAS behavior.
  • client_family filtering is applied before count/pagination and produces consistent admin/user statistics while preserving existing user-scope isolation.
  • Historical rows without identity continue to load and appear as unrecorded/unknown; migration is idempotent on SQLite/MySQL/PostgreSQL/ClickHouse paths used by LMM.
  • Raw UA, if stored at all, is bounded/sanitized and does not enter public/source-attribution exports unintentionally.
  • Regression tests cover known/unknown/malformed UA, old rows, filters, pagination/stat consistency, task settlement, and WebSocket request isolation.

Why issue instead of immediate backport

Upstream #7488 is still open with automated review in progress and introduces a database/logging/privacy surface across 32 files. The capability is applicable, but LMM should adopt the normalized observability contract rather than copy the full patch before upstream review settles or accidentally mix UA data into LMM's existing attribution/custom business logic.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions