Skip to content

design(billing): separate configurable relay pre-consume policy from fork balance gates #413

Description

@LIghtJUNction

Upstream source

QuantumNous/new-api commit: QuantumNous/new-api@9978ee1

Upstream replaces two hard-coded relay reservation rules with validated runtime settings:

  • quota_setting.trust_quota_usd (default 10 USD; 0 disables the trust bypass)
  • quota_setting.pre_consume_multiplier (default 1; finite positive number)

It also changes token-priced pre-consume from max(promptTokens, PreConsumedQuota) + max_tokens to estimated input cost × multiplier, and applies the same reservation policy to tiered_expr while keeping actual settlement unchanged. The billing snapshot stores the multiplier for compatibility/retry behavior.

Relevant upstream areas:

  • setting/operation_setting/quota_setting.go
  • relay/helper/price.go
  • service/billing_session.go
  • pkg/billingexpr/types.go
  • model/option.go
  • system billing/quota settings UI + tests

Why this is not a safe cherry-pick in LMM

LMM still has the upstream-old relay behavior:

  • apps/api-go/common/quota.go::GetTrustQuota() is fixed at 10 * QuotaPerUnit.
  • apps/api-go/relay/helper/price.go still reserves legacy token-priced requests using max(promptTokens, PreConsumedQuota) + meta.MaxTokens.
  • tiered billing still injects a default 8192 completion-token reservation when max_tokens is omitted.
  • apps/api-go/setting/operation_setting/quota_setting.go currently only contains EnableFreeModelPreConsume.
  • apps/api-go/service/billing_session.go uses common.GetTrustQuota() for the relay trust bypass.

But this fork also reuses common.GetTrustQuota() outside relay billing as a real product minimum-balance threshold, including Hero SMS purchase/reconciliation and drawing access/tests. Removing or globally redefining GetTrustQuota() as upstream now does would silently change unrelated LMM features.

Therefore the upstream idea is useful, but the fork needs a billing-specific adaptation rather than a global replacement.

Suggested adaptation

  1. Add billing-only settings such as quota_setting.trust_quota_usd and quota_setting.pre_consume_multiplier, with finite-number validation matching upstream semantics.
  2. Keep common.GetTrustQuota() (or replace it separately with an explicitly named product minimum-balance helper) for Hero SMS/drawing and other non-relay gates. Do not make those thresholds follow the relay trust-bypass setting by accident.
  3. Change only BillingSession trust-bypass logic to use the new billing threshold. Preserve the current requirements around wallet balance, limited token balance, subscriptions, forced pre-consume and assistants/playground.
  4. Treat the switch from output-inclusive/floor-based reservation to input-only × multiplier as a policy migration, not a mechanical refactor. Verify the effect on LMM's configured pricing, tiered expressions, dynamic pricing and high-output models before changing the default behavior.
  5. If input-only reservation is adopted, snapshot the multiplier with tiered billing so retries/settlement use a stable reservation policy; old snapshots must behave as multiplier=1.
  6. Fixed/request-priced models and asynchronous task billing should not inherit this multiplier unless explicitly justified and tested.
  7. Expose UI controls only after backend validation and persistence tests exist; price-lock behavior must remain unchanged.

Acceptance criteria

  • Relay trust bypass threshold is independently configurable without changing Hero SMS/drawing minimum-balance behavior.
  • trust_quota_usd=0 disables relay trust bypass; negative/NaN/Inf values are rejected.
  • pre_consume_multiplier rejects zero, negative, NaN/Inf and overflow; valid positive decimals persist/reload correctly.
  • Tests cover wallet above/equal/below threshold and limited-token above/equal/below threshold.
  • Unlimited tokens, subscriptions and ForcePreConsume preserve their intended behavior.
  • Legacy token-ratio reservation and tiered_expr reservation have explicit before/after tests for prompt tokens, client max_tokens, omitted max_tokens, multiplier <1 and >1.
  • Actual settlement remains based on actual usage and is unchanged by the reservation multiplier.
  • Tiered billing retries/group changes use a frozen/snapshot multiplier and old snapshots remain compatible.
  • Fixed-price/request-priced models and async task billing do not change unless separately documented.
  • Existing model price locks, dynamic pricing, group ratios, /fast, OAuth restrictions and administrator AI assistant behavior do not regress.
  • Full Go billing tests and release qualification pass.

Decision needed before implementation

The key behavioral choice is whether LMM wants to adopt upstream's input-only reservation semantics by default. That can substantially lower reservations for high-output requests compared with the current max_tokens/8192 fallback policy, so it should be validated against LMM's balance/settlement risk model rather than silently inherited.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions