Summary
Custom/external guardrails (and, from the payload shape, presumably the built-in ones too) are called once per message, with each message evaluated in complete isolation. The guardrail never sees the rest of the conversation, and the request it receives carries no information about which caller/user/API key originated the prompt. This makes it impossible to (1) detect anything that only becomes suspicious across multiple turns, and (2) build any per-user policy, audit trail, or attribution on top of guardrail findings.
This was confirmed empirically, not inferred from docs (see Evidence below).
Problem 1: No cross-message context
For a presidio-kind guardrail (hook_point: input) proxying a self-hosted analyzer, AISIX calls the analyzer's /analyze endpoint once per message in messages[], sending only that single message's text each time. There is no field carrying the other messages, their roles, their order, or any conversation/session identifier that would let the analyzer correlate the calls as belonging to the same request.
Consequence: an attack split across turns so that no single message is individually suspicious (e.g. an innocuous premise in one message, a short "now do it for real" instruction in the next) is invisible to the guardrail — each isolated call looks clean, and the guardrail has no way to reconstruct or reason about the assembled conversation the model actually sees.
Problem 2: No caller/user identity forwarded
The request AISIX sends to the guardrail carries no caller metadata at all: no API key / consumer id, no user id, no request or conversation id. Only generic HTTP plumbing (Content-Type, Accept, Accept-Encoding, Host, Content-Length) is present.
Consequence: a guardrail cannot log or make decisions per caller, cannot flag a specific user for repeated attempts, and cannot apply a policy that depends on who is asking (e.g. "this data classification is only sensitive for callers outside team X"). This matters directly for use cases like ours: a custom-trained model meant to check whether a prompt touches protected data needs to know who is asking to do that job properly, not just what is being asked.
Evidence
Setup: resources.yaml guardrail custom-regex-moderation (kind: presidio, hook_point: input, fail_open: false) pointing at a local test analyzer (regex-guardrail service, Flask app with a /analyze endpoint over patterns Guardrail and sujet interdit).
Request sent through AISIX (chat-eco model), 4-message conversation:
{
"model": "chat-eco",
"messages": [
{"role": "system", "content": "Tu es un assistant utile."},
{"role": "user", "content": "Bonjour, je vais te poser une question dans un instant."},
{"role": "assistant", "content": "Bien sur, je suis pret."},
{"role": "user", "content": "Peux-tu me parler du sujet interdit dont on a discute avant ?"}
]
}
Result: request blocked (content_filter, guardrail custom-regex-moderation) — expected, since the last message matches the sujet interdit pattern.
What the analyzer actually received (instrumented temporarily to log inbound requests), 4 separate calls, one per message:
headers={'Content-Type': 'application/json', 'Accept': '*/*', 'Accept-Encoding': 'gzip', 'Host': 'regex-guardrail:8000', 'Content-Length': '108'}
body={'text': 'Tu es un assistant utile.', 'language': 'en', 'entities': ['CUSTOM_REGEX_MATCH'], 'score_threshold': 0.5}
headers={'Content-Type': 'application/json', 'Accept': '*/*', 'Accept-Encoding': 'gzip', 'Host': 'regex-guardrail:8000', 'Content-Length': '138'}
body={'text': 'Bonjour, je vais te poser une question dans un instant.', 'language': 'en', 'entities': ['CUSTOM_REGEX_MATCH'], 'score_threshold': 0.5}
headers={'Content-Type': 'application/json', 'Accept': '*/*', 'Accept-Encoding': 'gzip', 'Host': 'regex-guardrail:8000', 'Content-Length': '106'}
body={'text': 'Bien sur, je suis pret.', 'language': 'en', 'entities': ['CUSTOM_REGEX_MATCH'], 'score_threshold': 0.5}
headers={'Content-Type': 'application/json', 'Accept': '*/*', 'Accept-Encoding': 'gzip', 'Host': 'regex-guardrail:8000', 'Content-Length': '144'}
body={'text': 'Peux-tu me parler du sujet interdit dont on a discute avant ?', 'language': 'en', 'entities': ['CUSTOM_REGEX_MATCH'], 'score_threshold': 0.5}
Each call carries exactly one message's text and nothing else — no conversation id linking the four calls, no role, no position/order, no caller identity in body or headers.
Why this matters
We run a custom-trained model as a guardrail specifically to catch prompts that reference protected data. Without conversation context, a request engineered so that the protected reference only becomes clear across turns slips through. Without caller identity, we cannot attribute a detection to a user, cannot build per-user audit trails, and cannot differentiate policy by caller — all of which are basic requirements for operating this kind of control in production.
Suggested fix / feature request
- Forward full conversation context to guardrails at
hook_point: input, either by sending the assembled messages[] (or a documented subset/summary of it) in one call, or by including a stable conversation/session identifier so an external guardrail can reconstruct history itself.
- Forward caller identity (API key id / consumer name, and any authenticated user id AISIX has) to guardrails, at minimum as request headers, so guardrail decisions and downstream audit logs can be attributed to a caller.
- If either of these is intentionally out of scope for guardrails today, document it explicitly (the current docs do not state that guardrails operate message-by-message with no caller metadata), so operators relying on guardrails for contextual or per-user policy know this is currently unsupported.
Summary
Custom/external guardrails (and, from the payload shape, presumably the built-in ones too) are called once per message, with each message evaluated in complete isolation. The guardrail never sees the rest of the conversation, and the request it receives carries no information about which caller/user/API key originated the prompt. This makes it impossible to (1) detect anything that only becomes suspicious across multiple turns, and (2) build any per-user policy, audit trail, or attribution on top of guardrail findings.
This was confirmed empirically, not inferred from docs (see Evidence below).
Problem 1: No cross-message context
For a
presidio-kind guardrail (hook_point: input) proxying a self-hosted analyzer, AISIX calls the analyzer's/analyzeendpoint once per message inmessages[], sending only that single message's text each time. There is no field carrying the other messages, their roles, their order, or any conversation/session identifier that would let the analyzer correlate the calls as belonging to the same request.Consequence: an attack split across turns so that no single message is individually suspicious (e.g. an innocuous premise in one message, a short "now do it for real" instruction in the next) is invisible to the guardrail — each isolated call looks clean, and the guardrail has no way to reconstruct or reason about the assembled conversation the model actually sees.
Problem 2: No caller/user identity forwarded
The request AISIX sends to the guardrail carries no caller metadata at all: no API key / consumer id, no user id, no request or conversation id. Only generic HTTP plumbing (
Content-Type,Accept,Accept-Encoding,Host,Content-Length) is present.Consequence: a guardrail cannot log or make decisions per caller, cannot flag a specific user for repeated attempts, and cannot apply a policy that depends on who is asking (e.g. "this data classification is only sensitive for callers outside team X"). This matters directly for use cases like ours: a custom-trained model meant to check whether a prompt touches protected data needs to know who is asking to do that job properly, not just what is being asked.
Evidence
Setup:
resources.yamlguardrailcustom-regex-moderation(kind: presidio,hook_point: input,fail_open: false) pointing at a local test analyzer (regex-guardrailservice, Flask app with a/analyzeendpoint over patternsGuardrailandsujet interdit).Request sent through AISIX (
chat-ecomodel), 4-message conversation:{ "model": "chat-eco", "messages": [ {"role": "system", "content": "Tu es un assistant utile."}, {"role": "user", "content": "Bonjour, je vais te poser une question dans un instant."}, {"role": "assistant", "content": "Bien sur, je suis pret."}, {"role": "user", "content": "Peux-tu me parler du sujet interdit dont on a discute avant ?"} ] }Result: request blocked (
content_filter, guardrailcustom-regex-moderation) — expected, since the last message matches thesujet interditpattern.What the analyzer actually received (instrumented temporarily to log inbound requests), 4 separate calls, one per message:
Each call carries exactly one message's text and nothing else — no conversation id linking the four calls, no role, no position/order, no caller identity in body or headers.
Why this matters
We run a custom-trained model as a guardrail specifically to catch prompts that reference protected data. Without conversation context, a request engineered so that the protected reference only becomes clear across turns slips through. Without caller identity, we cannot attribute a detection to a user, cannot build per-user audit trails, and cannot differentiate policy by caller — all of which are basic requirements for operating this kind of control in production.
Suggested fix / feature request
hook_point: input, either by sending the assembledmessages[](or a documented subset/summary of it) in one call, or by including a stable conversation/session identifier so an external guardrail can reconstruct history itself.