Skip to content

InvokeAgent spans emit non-spec gen_ai input/output message format #160

Description

@ninghu

Summary

After #149, passing agent_id/agent_name through instrumentation_options["langchain"] correctly puts gen_ai.agent.id on the emitted invoke_agent span. However, the same invoke_agent span appears to emit gen_ai.input.messages and gen_ai.output.messages in a non-OTel GenAI message format.

The OpenTelemetry semantic convention schema for gen_ai.input.messages is an array of ChatMessage objects with role and parts:

https://opentelemetry.io/docs/specs/semconv/gen-ai/gen-ai-input-messages.json

For a LangChain agent invocation, the parent invoke_agent span currently emits values shaped like this:

{
  "gen_ai.operation.name": "invoke_agent",
  "gen_ai.agent.id": "weather-agent-v1",
  "gen_ai.input.messages": "[\"System: You are a helpful weather assistant...\\nHuman: What should I wear in Seattle?\"]",
  "gen_ai.output.messages": "[\"chatcmpl-...\"]"
}

The child chat spans from the OpenAI instrumentation do emit the expected structured shape:

{
  "gen_ai.operation.name": "chat",
  "gen_ai.input.messages": "[{\"role\":\"user\",\"parts\":[{\"type\":\"text\",\"content\":\"...\"}]}]",
  "gen_ai.output.messages": "[{\"role\":\"constructor\",\"parts\":[{\"type\":\"text\",\"content\":\"...\"}]}]"
}

Request

Please emit gen_ai.input.messages and gen_ai.output.messages on invoke_agent spans using the OTel GenAI message schema, rather than JSON arrays of plain strings or response IDs.

For text-only content, that would be a JSON array of messages like:

[
  {
    "role": "user",
    "parts": [
      {
        "type": "text",
        "content": "What should I wear in Seattle?"
      }
    ]
  }
]

If the invoke_agent span cannot provide a spec-compliant output message, it may be better to omit gen_ai.output.messages from that parent span than to emit a non-message value such as chatcmpl-....

Why

  • Keeps invoke_agent spans aligned with the OpenTelemetry GenAI semantic convention.
  • Lets downstream consumers parse gen_ai.input.messages / gen_ai.output.messages uniformly across invoke_agent and child chat spans.
  • Avoids downstream trace-evaluation failures where the parent span looks like it has payload because the attributes are non-empty, but the payload is not parseable as OTel messages.
  • Prevents valid child chat span messages from being incorrectly treated as duplicate content of a non-spec parent invoke_agent span.

Context

This was observed with a LangChain external agent configured through:

use_microsoft_opentelemetry(
    enable_azure_monitor=True,
    instrumentation_options={
        "fastapi": {"enabled": False},
        "langchain": {
            "enabled": True,
            "agent_id": "weather-agent-v1",
            "agent_name": "weather-agent",
        },
    },
)

The installed distro code already has helpers in microsoft.opentelemetry.a365.core.message_utils that normalize strings into structured message wrappers. The issue appears to be that the LangChain invoke_agent span is still exporting legacy string-list payloads for gen_ai.input.messages / gen_ai.output.messages.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions