Skip to content

Guardrails receive no conversation context and no caller identity #1168

Description

@AlexPtte

Summary

Custom/external guardrails (and, from the payload shape, presumably the built-in ones too) are called once per message, with each message evaluated in complete isolation. The guardrail never sees the rest of the conversation, and the request it receives carries no information about which caller/user/API key originated the prompt. This makes it impossible to (1) detect anything that only becomes suspicious across multiple turns, and (2) build any per-user policy, audit trail, or attribution on top of guardrail findings.

This was confirmed empirically, not inferred from docs (see Evidence below).

Problem 1: No cross-message context

For a presidio-kind guardrail (hook_point: input) proxying a self-hosted analyzer, AISIX calls the analyzer's /analyze endpoint once per message in messages[], sending only that single message's text each time. There is no field carrying the other messages, their roles, their order, or any conversation/session identifier that would let the analyzer correlate the calls as belonging to the same request.

Consequence: an attack split across turns so that no single message is individually suspicious (e.g. an innocuous premise in one message, a short "now do it for real" instruction in the next) is invisible to the guardrail — each isolated call looks clean, and the guardrail has no way to reconstruct or reason about the assembled conversation the model actually sees.

Problem 2: No caller/user identity forwarded

The request AISIX sends to the guardrail carries no caller metadata at all: no API key / consumer id, no user id, no request or conversation id. Only generic HTTP plumbing (Content-Type, Accept, Accept-Encoding, Host, Content-Length) is present.

Consequence: a guardrail cannot log or make decisions per caller, cannot flag a specific user for repeated attempts, and cannot apply a policy that depends on who is asking (e.g. "this data classification is only sensitive for callers outside team X"). This matters directly for use cases like ours: a custom-trained model meant to check whether a prompt touches protected data needs to know who is asking to do that job properly, not just what is being asked.

Evidence

Setup: resources.yaml guardrail custom-regex-moderation (kind: presidio, hook_point: input, fail_open: false) pointing at a local test analyzer (regex-guardrail service, Flask app with a /analyze endpoint over patterns Guardrail and sujet interdit).

Request sent through AISIX (chat-eco model), 4-message conversation:

{
  "model": "chat-eco",
  "messages": [
    {"role": "system", "content": "Tu es un assistant utile."},
    {"role": "user", "content": "Bonjour, je vais te poser une question dans un instant."},
    {"role": "assistant", "content": "Bien sur, je suis pret."},
    {"role": "user", "content": "Peux-tu me parler du sujet interdit dont on a discute avant ?"}
  ]
}

Result: request blocked (content_filter, guardrail custom-regex-moderation) — expected, since the last message matches the sujet interdit pattern.

What the analyzer actually received (instrumented temporarily to log inbound requests), 4 separate calls, one per message:

headers={'Content-Type': 'application/json', 'Accept': '*/*', 'Accept-Encoding': 'gzip', 'Host': 'regex-guardrail:8000', 'Content-Length': '108'}
body={'text': 'Tu es un assistant utile.', 'language': 'en', 'entities': ['CUSTOM_REGEX_MATCH'], 'score_threshold': 0.5}

headers={'Content-Type': 'application/json', 'Accept': '*/*', 'Accept-Encoding': 'gzip', 'Host': 'regex-guardrail:8000', 'Content-Length': '138'}
body={'text': 'Bonjour, je vais te poser une question dans un instant.', 'language': 'en', 'entities': ['CUSTOM_REGEX_MATCH'], 'score_threshold': 0.5}

headers={'Content-Type': 'application/json', 'Accept': '*/*', 'Accept-Encoding': 'gzip', 'Host': 'regex-guardrail:8000', 'Content-Length': '106'}
body={'text': 'Bien sur, je suis pret.', 'language': 'en', 'entities': ['CUSTOM_REGEX_MATCH'], 'score_threshold': 0.5}

headers={'Content-Type': 'application/json', 'Accept': '*/*', 'Accept-Encoding': 'gzip', 'Host': 'regex-guardrail:8000', 'Content-Length': '144'}
body={'text': 'Peux-tu me parler du sujet interdit dont on a discute avant ?', 'language': 'en', 'entities': ['CUSTOM_REGEX_MATCH'], 'score_threshold': 0.5}

Each call carries exactly one message's text and nothing else — no conversation id linking the four calls, no role, no position/order, no caller identity in body or headers.

Why this matters

We run a custom-trained model as a guardrail specifically to catch prompts that reference protected data. Without conversation context, a request engineered so that the protected reference only becomes clear across turns slips through. Without caller identity, we cannot attribute a detection to a user, cannot build per-user audit trails, and cannot differentiate policy by caller — all of which are basic requirements for operating this kind of control in production.

Suggested fix / feature request

  • Forward full conversation context to guardrails at hook_point: input, either by sending the assembled messages[] (or a documented subset/summary of it) in one call, or by including a stable conversation/session identifier so an external guardrail can reconstruct history itself.
  • Forward caller identity (API key id / consumer name, and any authenticated user id AISIX has) to guardrails, at minimum as request headers, so guardrail decisions and downstream audit logs can be attributed to a caller.
  • If either of these is intentionally out of scope for guardrails today, document it explicitly (the current docs do not state that guardrails operate message-by-message with no caller metadata), so operators relying on guardrails for contextual or per-user policy know this is currently unsupported.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions