You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Telemetry always uses the S2S route with an app-only token, in every auth mode. The exporter posts to /observabilityService/ with a token for the exporting agent identity. OBO / Agentic User tokens are used only for workload (MCP / Graph) calls. Registered blueprint agent instances need no Agent365.Observability.OtelWrite permission or admin consent (subject to service policy).
instrument-observability
New app-only token resolver scaffolds: observability/app-token-resolver.ts, observability/app_token_resolver.py, and Observability/AgentAppTokenResolver.cs. Each reuses the hosting connection's blueprint credential: an FMI assertion for the turn's agent identity, then agent-identity client_credentials for the OBS scope.
The S2S route flag is set in every mode.
New "Migrating delegated telemetry" path for existing wiring, including a note on durable-delivery spools.
Idempotency no longer skips code that still wires delegated telemetry.
.NET:Microsoft.OpenTelemetry 1.0.3+ flattened o.Agent365.Exporter.* to o.Agent365.*. The snippets now target the flattened API, and the 1.0.2-and-earlier form is noted.
Stop hooks and a365-code-validator
With the distro, require the S2S route flag.
Flag delegated telemetry tokens, with new findings: *-obs-delegated-route, *-obs-delegated-token, python-obs-prefetch-connection-missing, and agent-registration-not-recorded.
The hook and the standalone scanner stay in parity.
A safe fix switches the route flag only when the token is already app-only.
Other skills, shared detection, docs, and evals: blueprint agents are authorized by registration (on a 403, a365 setup all --agent-registration-only). AI Teammates keep the OtelWrite application-role step that a365 setup all --aiteammate prints.
Each new detection branch was mutation-checked; disabling it fails a test.
The validate.yml job scripts pass on an LF checkout.
Reference snippets were compiled and run outside the repo:
Node.js resolver: tsc --strict against @microsoft/agents-hosting 1.8.1 and @microsoft/opentelemetry 1.4.0.
.NET resolver and handler: build with warnings-as-errors against Microsoft.OpenTelemetry 1.1.0 and Microsoft.Agents.* 1.4.83. The .NET S2S scaffold also builds.
Python resolver: behavior tests with stubbed MSAL.
E2E in a test tenant: a Node.js blueprint agent using this pattern exported to /observabilityService/ with an app-only token and no OtelWrite grant. It got HTTP 200, and activity appeared in MAC.
Notes
The plugin version is not bumped; that's left for release.
AI Teammate S2S authorization was not verified end to end (hiring failed in the test tenant). AI Teammate guidance therefore keeps the OtelWrite application-role step.
Telemetry now always uses the Observability API S2S route (/observabilityService/)
with an app-only token for the exporting agent identity. Registered blueprint agent
instances therefore need no Agent365.Observability.OtelWrite permission or admin
consent; OBO / Agentic User tokens are used only for workload (MCP / Graph) calls.
- instrument-observability references: app-only token resolver scaffolds for
Node.js, Python and .NET that reuse the hosting connection (FMI assertion for the
turn's agent identity, then agent-identity client_credentials); S2S route flag in
every auth mode; migration path for existing delegated wiring; .NET options moved
to the flattened Microsoft.OpenTelemetry 1.0.3+ API (o.Agent365.*).
- Stop hooks and a365-code-validator: require the S2S route flag and flag delegated
telemetry tokens (new findings, hook/standalone parity tests).
- Provisioning, test-local, shared detection, docs and evals: registration-based
authorization; AI Teammates keep the OtelWrite application-role step.
Related: microsoft/Agent365-nodejs#290, microsoft/Agent365-Samples#339,
microsoft/Agent365-devTools#501
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5cbf5f6b-cc40-4b7e-a591-65848db73a12
With the unified distro, the token resolver is the only export credential for
the S2S route: without one, .NET falls back to the delegated token cache, Node
has no token, and the Python exporter drops every span. The route flag alone
therefore passed validation for a configuration that cannot export.
- validate-instrument-observability: require o.Agent365.TokenResolver /
a365 tokenResolver / a365_token_resolver (or the contextual variants) when the
distro is used; legacy token-cache helpers stay accepted without the distro.
- a365-code-validator (stop hook and standalone scanner): new high findings
dotnet-, node-, and python-obs-token-resolver-missing, kept in parity.
- Tests for each language, the kwargs-dict resolver form, and scanner parity;
docs list the new check.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5cbf5f6b-cc40-4b7e-a591-65848db73a12
This mirrors the standalone false positive: aiTeammate: false is not sufficient to prove the project uses a Blueprint, so a non-blueprint app-registration agent with an agenticAppId but no agentRegistrationId is incorrectly reported as missing registration. Gate this finding on generated.agentBlueprintId (or another explicit blueprint signal) to keep the stop hook aligned with the actual registration flow.
Token anchors do not verify resolver wiring at distro call sites
plugins/agent365/shared/agent-detection.md:452
These token anchors still accept symbols/files anywhere in the project rather than proving that the distro call receives an app-only resolver. For example, ObservabilityTokenService.cs can exist without o.Agent365.TokenResolver, or tokenResolver can be an unused import while the Node/Python entry-point call omits it; with the other three anchors present, has_obs becomes true and the AI Teammate flow skips recovery even though export cannot authenticate. Tie each token signal to the corresponding distro call and resolver wiring, and keep the duplicated make-ai-teammate rules in sync.
Agent registration finding incorrectly applies to non-blueprint agents
This finding is emitted for every aiTeammate: false config with an agent identity, but that setting also covers non-blueprint app-registration agents that legitimately have no agentRegistrationId. Without requiring generated.agentBlueprintId (or an equivalent useBlueprint signal), the validator incorrectly recommends --agent-registration-only for agents where no blueprint instance exists.
obs_token is satisfied by the mere presence of AgentAppTokenResolver, ObservabilityTokenService, or AddAgent365Observability; it does not require the resolver to be assigned to o.Agent365.TokenResolver. A project can therefore have the S2S route, handler scopes, and an unregistered/unwired scaffold but be classified as has_obs_complete, causing Phase 9.5 to skip the instrumentation that the new stop validator would reject. Require the actual exporter option wiring (and likewise verify the Node/Python resolver is passed to the distro call), not just a scaffold symbol.
…a blueprint
Address the second Copilot review on #84:
- The resolver check now looks only at the distro call's own arguments
(useMicrosoftOpenTelemetry options, use_microsoft_opentelemetry keywords,
UseMicrosoftOpenTelemetry options callback). It accepts explicit and
shorthand properties and follows a variable passed to the call to its
initializer in the same file, so an unused import or a resolver defined
elsewhere no longer counts. One helper is shared verbatim by both stop hooks
and the standalone scanner, and a test keeps the copies identical.
- agent-registration-not-recorded now also requires a blueprint ID in
a365.generated.config.json.
- has_obs token anchors in shared/agent-detection.md and make-ai-teammate
require the resolver to be wired at the distro call.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5cbf5f6b-cc40-4b7e-a591-65848db73a12
Resolver wiring at the call site. Covers the three open threads; details are in the thread replies.
Previously missed:
agent-registration-not-recorded now also requires agentBlueprintId in a365.generated.config.json, in both the stop hook and the standalone scanner.
The has_obs token anchors in shared/agent-detection.md and make-ai-teammate now require the resolver to be passed in the distro call. A scaffold file or resolver symbol that isn't wired there doesn't count.
npm test: 180 passed (10 new). The validate.yml job scripts pass on an LF checkout.
The route check searches every TypeScript file rather than the useMicrosoftOpenTelemetry call. An unrelated object such as const workloadOptions = { useS2SEndpoint: true } makes this check pass even when the exporter call omits the flag, so the stop hook misses the high route finding and a delegated-route exporter can remain. Inspect the distro call's a365 options (including variable-passed options) instead of accepting any source-wide match.
Comments are falsely detected as delegated-token calls
The delegated-token scan passes raw source text to findCallBlocks, so a comment containing an example such as // Do NOT call refreshObservabilityToken(..., authorization) is classified as executable and emits a high finding. This can block a correctly migrated project whose comments document the forbidden legacy call; strip comments before call detection (while retaining the original file for the finding location).
The .NET route check is also file-wide: an unrelated UseS2SEndpoint = true assignment can suppress dotnet-obs-delegated-route even when the UseMicrosoftOpenTelemetry options omit it. Scope this predicate to the distro call/options callback; otherwise the stop hook can approve a non-S2S exporter.
This route check searches raw contents of every TS/JS file instead of the a365 options passed to useMicrosoftOpenTelemetry. A comment or unrelated object such as const workloadOptions = { useS2SEndpoint: true } can therefore satisfy the check while the exporter call still uses the delegated default route, allowing an app-only token to be sent to the wrong endpoint. Scope this check to the distro call (and its followed options variable) and ignore comments, as the resolver check already does.
Comments are falsely detected as delegated-token calls
The delegated-token scan passes raw source text to callBlocks, so a comment containing an example such as // Do NOT call refreshObservabilityToken(..., authorization) is classified as executable and emits a high finding. This can block a correctly migrated project whose comments document the forbidden legacy call; strip comments before call detection (while retaining the original file for the finding location).
pyS2SInCode scans all Python files, so an unrelated helper or comment with a365_use_s2s_endpoint=True suppresses the missing-route error even when use_microsoft_opentelemetry(...) omits the kwarg. Scope the route check to the distro call (while retaining the explicit env fallback) so this hook cannot accept a legacy-route exporter by coincidence.
This route check searches raw contents of every TS/JS file instead of the a365 options passed to useMicrosoftOpenTelemetry. A comment or unrelated object such as const workloadOptions = { useS2SEndpoint: true } can therefore satisfy the check while the exporter call still uses the delegated default route, allowing an app-only token to be sent to the wrong endpoint. Scope this check to the distro call (and its followed options variable) and ignore comments, as the resolver check already does.
This issue also appears on line 300 of the same file.
Address the "previously missed" items in the third Copilot review on #84:
- The S2S route flag now counts only when the distro call itself passes it,
using the same call-site matching as the resolver check. An unrelated
`useS2SEndpoint: true` object, a `UseS2SEndpoint = true` on other options,
or an `a365_use_s2s_endpoint=True` on another call no longer satisfies it.
The Python env fallback is unchanged.
- The delegated-token and prefetch scans ignore comments, so a comment that
documents the legacy call is not reported.
- One shared helper (DISTRO_CALLS / distroCallMatches / codeMatches) is kept
identical across both stop hooks and the standalone scanner.
- Detection docs: the route anchor must also be set at the distro call.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5cbf5f6b-cc40-4b7e-a591-65848db73a12
The seven "previously missed" items from the latest Copilot overview are addressed in b7e93c3:
Route checks are call-scoped (items 1, 3, 4, 6, 7).useS2SEndpoint: true, UseS2SEndpoint = true and a365_use_s2s_endpoint=True now count only when the distro call itself passes them. Matching uses the same helper as the resolver check: explicit or shorthand options, a variable passed to the call followed to its initializer, and comments ignored. An unrelated workloadOptions = { useS2SEndpoint: true }, UseS2SEndpoint = true on other options, or the kwarg on another helper call no longer passes. The Python A365_USE_S2S_ENDPOINT env fallback is unchanged.
Delegated-token scans ignore comments (items 2, 5, 7). A comment that documents the legacy refreshObservabilityToken(...authorization), RegisterObservability(..., new AgenticTokenStruct(...)) or observability exchange_token(...) call is no longer reported. The self.connection_manager prefetch check ignores comments too.
The helper is shared verbatim by both stop hooks and the standalone scanner, and a test keeps the three copies identical.
Verification:
npm test: 187 passed. New tests cover each case in the stop hook, plus a parity test across both scanners. Each change was mutation-checked.
The validate.yml job scripts pass on an LF checkout.
Address the fourth Copilot review on #84: options inside the distro call's
parentheses but outside the SDK's own options no longer count.
- Node.js: only the `a365` options object counts: inline, by variable,
shorthand, or a spread one level deep. `instrumentationOptions:
{ custom: { useS2SEndpoint: true } }` does not.
- .NET: only `.Agent365` (or `.Agent365.Exporter`) assignments count, so
`workloadOptions.UseS2SEndpoint = true` inside the callback does not.
- Python: only the call's own keyword arguments count (plus `**` spreads), not
keywords of a nested call. The call itself is masked when initializers are
looked up, because Python keyword arguments look like assignments.
- When the options cannot be read statically (a method group, a factory call,
or a callback that never touches Agent365), the whole file is checked
instead, to avoid false positives.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5cbf5f6b-cc40-4b7e-a591-65848db73a12
The reason will be displayed to describe this comment to others. Learn more.
Approve with nits. This is a solid, well-tested change. None of the notes below block it, but I'd like to see the five should-fix items addressed before merge, or as a quick follow-up.
Summary
This PR moves all observability guidance to the S2S route with an app-only token in every auth mode:
New app-only resolver scaffolds for obo / agentic-user: app-token-resolver.ts, app_token_resolver.py (sync resolve plus a per-turn prefetch) and AgentAppTokenResolver.cs. They reuse the hosting connection's blueprint credential: an FMI assertion, then agent-identity client_credentials.
The S2S route flag is set everywhere.
A migration path for existing delegated wiring.
Call-scoped route and resolver checks in the stop hooks and in a365-code-validator.
Merge gate
The CLI behavior described in make-a365-agent, AGENTS.md, copilot-instructions.md and CLAUDE.md comes from microsoft/Agent365-devTools#501, which is still open. That behavior is that setup all no longer requests OtelWrite for blueprint agents, and exits 1 when registration fails. Please merge after #501 is released, or pin the minimum CLI version in the wording. microsoft/Agent365-nodejs#290 and microsoft/Agent365-Samples#339 are also still open, and the "as the Agent 365 samples do" references point at unmerged sample code.
Findings on lines outside this diff
[Should-fix] plugins/agent365/skills/make-ai-teammate/references/nodejs-ai-teammate.md:744-745: this says agent365Observability__agentBlueprintId and __clientId/__clientSecret are "never written by the CLI… Stray. Omit." The CLI does write all three (ProjectSettingsSyncHelper.cs, around L659-671 on devTools main). This PR's new non-agentic fallback guard in nodejs-observability.md compares against agent365Observability__agentBlueprintId, so an agent following make-ai-teammate would drop the variable the guard needs. Suggest removing those two bullets, or saying the CLI writes them and the fallback guard uses __agentBlueprintId.
[Should-fix] plugins/agent365/skills/instrument-observability/references/nodejs-observability.md:687: agentId: recipient?.agenticAppId ?? process.env.agent365Observability__agentId ?? ''. For AI Teammates that env var holds the blueprint ID, as this PR notes. So non-agentic turns tag spans with the blueprint as gen_ai.agent.id, and the resolver then tries (and fails) to mint a token for it on every export batch. Suggest the same guard the .NET handler uses: a GUID, not equal to agent365Observability__agentBlueprintId, otherwise ''.
[Nit] Snippets that still leave out the route flag, which the new stop hook rejects:
Adding useS2SEndpoint: true / a365_use_s2s_endpoint=True (or a comment pointing to it) would keep copy-paste safe.
[Nit] Stale sample references:
python-observability.md:13: the sample-lag note says the AF sample imports get_observability_authentication_scope. Agent365-Samples#339 removes that import along with the delegated _setup_observability_token.
plugins/agent365/skills/purview-dlp-integration/assets/purview.py:230: the docstring says it "Uses the same exchange_token call the A365 host uses for observability", which is no longer true.
The other findings are inline.
Validation
Tests:npm test passes locally (194/194).
validate.yml steps: I ran them locally in Git Bash. JSON lint, JS hook syntax, version parity (1.0.2), cross-platform patterns, internal link integrity and hook coverage all pass. I couldn't run the kebab-case step locally because it needs python3; it passes in CI.
SKILL.md frontmatter: all 8 files parse as YAML.
Node.js compile check:observability/app-token-resolver.ts plus the OBO entry point compile with tsc --strict against @microsoft/agents-hosting 1.8.1 and @microsoft/opentelemetry 1.4.0.
API names: I checked the ones the skills rely on against the published packages and the CLI source, and they match:
Node distro:a365.useS2SEndpoint (defaults to false), the a365.tokenResolver signature, durableDelivery
Python:a365_use_s2s_endpoint, the sync token_resolver(agent_id, tenant_id) call, get_agentic_application_token, get_default_connection()
.NET: the o.Agent365.* flattening in 1.0.3, IAgenticTokenProvider.GetAgenticApplicationTokenAsync, and GetAgenticInstanceId() returning Recipient.AgenticAppId
CLI:Agent365Observability:AgentId holds the agent identity for blueprint agents and the blueprint ID for AI Teammates, and agentRegistrationId / agenticAppId are the right JSON names
Cross-PR notes
Agent365-nodejs#290 changes only the legacy @microsoft/agents-a365-observability* packages: the route flag becomes ignored, and the delegated RefreshObservabilityToken(..., authorization) overload now throws. The skills target the unified @microsoft/opentelemetry distro, where the flag is still required and defaults to false, so "set it in every mode" is correct. The validators rightly don't flag #290's new app-only RefreshObservabilityToken(agentId, tenantId, resolver) overload.
Agent365-Samples#339 takes a different resolver design (see the inline question) and keeps durable replay disabled for longer than this PR suggests.
Agent365-devTools#501: the claims here match its behavior. AI Teammate setup is unchanged there (see the inline note). #501 also says a365 setup permissions bot still configures the Observability API, which the CEA flow runs. That's harmless, but maybe worth a mention.
The reason will be displayed to describe this comment to others. Learn more.
[Should-fix] The spool-replay caveat applies to all three distros, not just Node. Microsoft.OpenTelemetry 1.1.0 has offline storage on by default (o.Agent365.DisableOfflineStorage is false) and replays with BuildEndpointPath(record.TenantId, record.AgentId, record.UseS2SEndpoint). Python microsoft-opentelemetry replays with use_s2s_endpoint=record.use_s2s_endpoint (disable_offline_storage also defaults to False).
Could we add o.Agent365.DisableOfflineStorage = true (.NET) and a365_exporter_disable_offline_storage=True (Python) alongside the Node durableDelivery: { enabled: false }?
Two smaller points:
durableDelivery first appears in @microsoft/opentelemetry 1.4.0 (it isn't in 1.2.0 or 1.3.0), so "1.4.0+" is more accurate than "1.4.x".
Agent365-Samples#339 keeps replay off "until the installed release enforces S2S for both live and replayed exports", which is stricter than "during the migration". It'd be good to align the wording.
The same applies to nodejs-observability.md:600-605 and to the .NET and Python migration sections.
It uses "1.4.0+" and the samples' wording: keep replay off until the installed release enforces S2S for both live and replayed exports. nodejs-observability.md and the .NET and Python migration sections say the same. The .NET note also says offline storage is on by default in 1.1.0+ (42f24b7).
- The `authHandlerName` should resolve to the agentic auth handler name (from config `AgentApplication:AgenticAuthHandlerName`) when `IsAgenticRequest()` is true, OBO handler name (from `AgentApplication:OboAuthHandlerName`) otherwise.
- **Keep the `Agent365Observability` section in `appsettings.json`** (`EnableAgent365Exporter` and base exporter settings are still required — Phase 6 handles these). For **OBO**, you do **not** need to hardcode per-agent IDs, tenant IDs, or S2S credentials in that section — the agent ID and tenant ID are resolved from the request at runtime on each turn.
- **Keep the `Agent365Observability` section in `appsettings.json`** (`EnableAgent365Exporter` and base exporter settings are still required — Phase 6 handles these). No S2S credentials are needed there: `AgentAppTokenResolver` uses the agent's `Connections` settings, and agentic turns resolve their agent ID at runtime.
- **The inline pattern shown above is preferred** for new code (mirrors PR #308 in `microsoft/Agent365-Samples`). The older `A365OtelWrapper.InvokeObservedAgentOperation(...)` static-wrapper pattern at `Agent365-samples/dotnet/agent-framework/sample-agent/telemetry/A365OtelWrapper.cs` is functionally equivalent but uses a separate helper class.
The reason will be displayed to describe this comment to others. Learn more.
[Should-fix] This still calls A365OtelWrapper.InvokeObservedAgentOperation(...) "functionally equivalent". But the migration note (L268) and both validators now treat A365OtelWrapper / new AgenticTokenStruct(...) as delegated wiring to remove, and Agent365-Samples#339 deletes the w365 sample's Telemetry/A365OtelWrapper.cs.
Maybe something like: "A365OtelWrapper is the legacy delegated-token wrapper. Remove it and migrate (see 'Migrating delegated telemetry' in Phase 3)." The "mirrors PR #308" note is also out of date now that the inline pattern has no RegisterObservability call.
The reason will be displayed to describe this comment to others. Learn more.
Fixed in bd73610. That bullet now says A365OtelWrapper is the legacy delegated-token wrapper: remove it and migrate the handler to the inline app-only S2S pattern (see "Migrating delegated telemetry" in Phase 3). The out-of-date "mirrors PR #308" note is gone.
1. **The Setup Summary table** from CLI output — verbatim.
2. **`Agent365.Observability.OtelWrite` is automatically granted** to the agent identity by `a365 setup all` — no GA consent step required for newly provisioned agents. If the CLI output includes a "Permission Grants" action item (upgrade scenario for pre-1.1 agents), display the PowerShell script verbatim so the user can hand it to a Global Admin.
2. **No Observability API permission is needed.** Recent `a365 setup all` versions no longer request `Agent365.Observability.OtelWrite` (or its admin consent) for blueprint agents. Telemetry is exported over the S2S route with an app-only token, and the route authorizes the registered agent instance. These versions also fail setup (exit code 1) when agent registration fails or cannot be verified. If that happens, show the error and have the user re-run `a365 setup all --agent-registration-only` after fixing it. Older CLI versions may still grant OtelWrite, which is harmless. If the CLI output includes an action item for other permissions (Graph, Bot API, custom resources), display the printed PowerShell script verbatim so the user can hand it to a Global Admin.
The reason will be displayed to describe this comment to others. Learn more.
[Should-fix] Both CLI behaviors described here (blueprint agents no longer request OtelWrite; setup exits 1 when registration fails or can't be verified) come from microsoft/Agent365-devTools#501, which is still open. On the released CLI, --agent-registration-only still exits 0 when registration fails.
Could we merge this after #501 ships, or name the minimum CLI version here instead of "recent versions"? The same wording appears in AGENTS.md:55, .github/copilot-instructions.md:167, CLAUDE.md rule 13, and the instrument-observability Known Issues section.
The reason will be displayed to describe this comment to others. Learn more.
There's no released CLI version to pin yet, so bd73610 makes the wording version-safe in make-a365-agent, AGENTS.md, .github/copilot-instructions.md, CLAUDE.md rule 13 and the Known Issues section:
Newer a365 setup all versions skip OtelWrite for blueprint agents and exit 1 when registration fails or can't be verified.
Older versions may still grant OtelWrite (harmless) and may exit 0 after a failed registration, so the skill checks the setup output for a registration error and reruns a365 setup all --agent-registration-only.
This PR will also wait for the release that contains microsoft/Agent365-devTools#501, and I can pin the minimum version once it's known.
The reason will be displayed to describe this comment to others. Learn more.
[Nit] This check only fires when the call arguments contain the literal authorization. In @microsoft/opentelemetry 1.4.0, refreshObservabilityToken(agentId, tenantId, turnContext, authorization, scopes?, authHandlerName?) has only the delegated signature, so every lowercase call is delegated. A fixture with AgenticTokenCacheInstance.refreshObservabilityToken('a', 't', ctx, app.auth) passes without a finding.
Suggested fix:
Flag any lowercase refreshObservabilityToken( call.
For PascalCase RefreshObservabilityToken, flag only calls with 4 or more arguments, so the new app-only (agentId, tenantId, resolver) overload from Agent365-nodejs#290 stays unflagged.
The same predicate is in validate-a365-code-validator.js and a365-code-validator.js:407.
The same predicate is in all three validators, and the parity test keeps the shared helper byte-identical. Tests: your app.auth fixture is now reported, and a new test checks that the app-only overload passes and a nested four-argument delegated call is reported.
The reason will be displayed to describe this comment to others. Learn more.
[Nit] As written, this snippet won't run. on_message is a module-level @AGENT_APP.activity function but calls self._setup_observability_token(...), and _setup_observability_token(self, ...) is also a module-level function, not a method.
The structure predates this PR, but since the body was rewritten here, could we show both as GenericAgentHost methods (as python-ai-teammate.md does), or drop self and pass the connection manager explicitly?
The reason will be displayed to describe this comment to others. Learn more.
Fixed in bd73610 and 42f24b7. Both are now GenericAgentHost methods, as in python-ai-teammate.md: _setup_observability_token is a method, and on_message is registered in _setup_handlers through self._adapter.on_activity(ActivityTypes.message). The handler also calls self._agent.process_user_message(...) with the AgentInterface arguments make-ai-teammate generates, so the snippet runs as written.
The reason will be displayed to describe this comment to others. Learn more.
[Nit] The Node and Python scaffolds also reject a token whose azp/appid isn't the exporting agent ID or whose tid isn't the tenant. Here only scp is checked, so it'd be good to add the same check.
Two smaller robustness points:
The single _refreshGate (L502) makes every agent identity wait behind one slow Entra call. A per-key gate would avoid that.
None of the three resolvers caches failures. An identity that can't mint a token (an unregistered instance, or a blueprint ID) calls Entra again on every export batch; a short backoff would help.
AgentAppTokenResolver now rejects a token that has scp, whose azp/appid isn't the exporting agent ID, whose tid isn't the tenant, or whose aud isn't the Observability API.
The single _refreshGate is now one gate per (tenant, agent) key.
All three resolvers back off a failed acquisition for 60 seconds, so an identity that can't mint a token no longer calls Entra on every export batch.
a single identity (cross-identity export is rejected)
stricter token checks: idtyp=app (or roles / oid == sub), aud equal to the OBS resource, and token_type=bearer
Reusing the hosting connection makes sense for multi-instance AI Teammates, and it covers cert/FIC/WID through getAgenticApplicationToken. A short note on why the skills differ would help anyone comparing against the samples. The aud check also looks cheap to add to all three scaffolds.
The reason will be displayed to describe this comment to others. Learn more.
Yes, it's intentional. Reusing the hosting connection (getAgenticApplicationToken) lets one deployment serve several agent instances and tenants, which multi-instance AI Teammates need, and it inherits certificate, federated-identity and managed-identity credentials. The samples use a dedicated single-identity OBS credential to keep telemetry auth isolated from business auth.
bd73610 adds a short note explaining this to the Node, Python and .NET scaffolds, and adds the aud check to all three. I didn't add idtyp or token_type checks: the resolver requests the token itself with agent-identity client_credentials, and the scp, azp/appid, tid and aud checks already reject a delegated or mis-targeted token.
- Agent (Non AI Teammate) → `obo` (On-Behalf-Of) or `s2s` (Service Principal, no user token)
All three `authMode` values use an auth handler reference in SDK code — the difference is Azure AD provisioning. For OBO paths: .NET reads `authHandlerName` from config (`AgentApplication:AgenticAuthHandlerName`); Node.js passes `agentApplication.authorization` (the auth object) to `RefreshObservabilityToken`; Python uses `auth_handler_id=self.auth_handler_name` (from config) in `exchange_token()` — never hardcode `"AGENTIC"`. Agent IDs are always resolved dynamically from TurnContext — .NET: `turnContext.Activity.GetAgenticInstanceId()` (service principal object ID); Node.js/Python: `recipient.agenticAppId` / `agentic_app_id`. Results are cached in `.a365-workspace-detection.local.json` under `agentType` and `authMode` fields so subsequent skill invocations skip re-questioning. If `authMode = s2s` and the skill is `add-workiq-tools`, the skill exits immediately — WorkIQ is not available for s2s agents (requires a delegated OBO token).
Telemetry never goes through the auth handler: every `authMode` exports over the S2S route with an app-only token for the exporting agent identity (the S2S route rejects delegated `scp` tokens). The auth handler reference is for workload calls (MCP / Graph): .NET reads `authHandlerName` from config (`AgentApplication:AgenticAuthHandlerName`); Node.js uses `agentApplication.authorization` (the auth object); Python uses `auth_handler_id=self.auth_handler_name` (from config) — never hardcode `"AGENTIC"`. Agent IDs are resolved dynamically from TurnContext — .NET: `turnContext.Activity.GetAgenticInstanceId()` (service principal object ID); Node.js/Python: `recipient.agenticAppId` / `agentic_app_id`. Only non-agentic turns of non-AI-Teammate agents fall back to the provisioned agent identity in config (never the blueprint). Results are cached in `.a365-workspace-detection.local.json` under `agentType` and `authMode` fields so subsequent skill invocations skip re-questioning. If `authMode = s2s` and the skill is `add-workiq-tools`, the skill exits immediately — WorkIQ is not available for s2s agents (requires a delegated OBO token).
The reason will be displayed to describe this comment to others. Learn more.
[Nit] This line still says GetAgenticInstanceId() returns the "service principal object ID". The SKILL now correctly says it returns Recipient.AgenticAppId, which matches ActivityExtensions.GetAgenticInstanceId in Agents-for-net. Since the resolver passes that value as client_id, it'd be good to fix it here too.
- For agents provisioned before CLI 1.1, this manual step is still required
- **403 `insufficient_scope` on export** → the agent instance is not registered and has no `OtelWrite` application role. Blueprint agents: run `a365 setup all --agent-registration-only` (idempotent), then retry. AI Teammates: that flag does not apply — complete the Observability API S2S app-role action item `a365 setup all --aiteammate` prints (the application role below).
- **Application role (AI Teammates, and fallback for blueprint agents):** a Global Administrator can grant the `Agent365.Observability.OtelWrite` **application** role on the Blueprint, which agent identities inherit. The S2S route accepts it. Use the PowerShell steps that `a365 setup all` prints, or the Entra portal: App registrations > Blueprint > API permissions > APIs my organization uses > `9b975845-388f-4429-889e-eab1ef63949c` > **Application** `Agent365.Observability.OtelWrite` > Grant admin consent.
- **Do not** add the delegated `OtelWrite` scope or run a delegated consent flow for telemetry. Delegated tokens are rejected by the S2S route. Only agents still on an old SDK that exports over the delegated route need that grant, and they should migrate instead (see "Migrating delegated telemetry" in Phase 3).
The reason will be displayed to describe this comment to others. Learn more.
[Nit] devTools#501 leaves AI Teammate setup unchanged, so a365 setup all --aiteammate still requests OtelWrite (both delegated and application), and the skill tells the agent to show that printed script verbatim. Could we add a sentence saying the delegated part of the CLI's AI Teammate grant is harmless and isn't used for telemetry? Otherwise an agent may read this line as "don't run the script".
The reason will be displayed to describe this comment to others. Learn more.
Fixed in bd73610. The Known Issues section now says a365 setup all --aiteammate still requests OtelWrite as both a delegated and an application permission, and tells the agent to run the printed script. The application role is what the S2S route accepts; the delegated part is harmless and isn't used for telemetry.
The reason will be displayed to describe this comment to others. Learn more.
[Question] With a sync resolver that only reads what prefetch cached: after a restart, the distro replays spooled records through token_resolver(record.agent_id, record.tenant_id). That returns None until the same identity handles a turn and prefetches, so records for idle identities can age out (the default maximum record age is 2 days). Is that acceptable? If so, a sentence noting it would help.
The reason will be displayed to describe this comment to others. Learn more.
Yes, that's acceptable, and 42f24b7 states it once under the resolver scaffold. After a restart, resolve returns None for an identity until it handles a turn and prefetches. If offline storage is enabled, records replayed for identities that stay idle can age out (default maximum record age: 2 days). The migration disables offline storage (a365_exporter_disable_offline_storage=True) until the installed release enforces S2S for replayed exports, so this matters only if someone turns it back on.
- The GenericAgentHost snippet now calls process_user_message with the
AgentInterface arguments make-ai-teammate generates.
- State the offline-storage replay caveat once, conditioned on storage being
enabled, and say .NET offline storage is on by default in 1.1.0+.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5cbf5f6b-cc40-4b7e-a591-65848db73a12
Thanks, Rick and Dominik. I pushed bd73610 and 42f24b7 and replied on each inline thread. For the findings outside the diff:
[Should-fix] nodejs-ai-teammate.md:744-745: fixed. agent365Observability__agentBlueprintId is now described as written by the CLI and used by the fallback guard. __clientId/__clientSecret are written by the CLI for compatibility: keep them if present, but don't hand-author them. make-ai-teammate/references/deploy-pipeline.md had the same bullets and is fixed too.
[Should-fix] nodejs-observability.md:687: fixed. On non-agentic turns, agent365Observability__agentId is used only if it is a GUID that differs from agent365Observability__agentBlueprintId; otherwise ''. This is the same guard as the .NET handler.
[Nit] Route flag in snippets: added useS2SEndpoint: true / a365_use_s2s_endpoint=True to the three nodejs-observability.md snippets, python-observability.md, the a365-code-validator SKILL.md examples and validation-checklist.md.
[Nit] Stale sample references: the python-observability.md sample-lag note no longer mentions get_observability_authentication_scope. The purview.py docstring now says its token is for Purview Graph calls only, and that observability uses an app-only S2S token.
Cross-PR:
AGENTS.md now notes that a365 setup permissions bot still configures the Observability API for the CEA flow, which is harmless because telemetry uses S2S.
The manual OtelWrite fallback here is application-only because it's for the S2S route. devTools#501's CHANGELOG note covers agents on the delegated route, which need both delegated and application permissions.
- require direct a365 route and resolver properties in scanner helper
- reject nested, sibling, and unresolvable observability options
- add regression coverage for direct and identifier-resolved a365 options
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5cbf5f6b-cc40-4b7e-a591-65848db73a12
This finding is gated on a365.config.json.aiTeammate === false, but the repository's non-AI path is identified by the detection cache's agentType: "system-agent"; the documented config lookup only guarantees blueprintId/agentBlueprintId. A normal make-a365-agent project without an aiTeammate key therefore skips this missing-registration warning even when agenticAppId is present and agentRegistrationId is absent. Use the authoritative agent-type signal and mirror the fix in the standalone scanner.
Use authoritative agent type for registration checks
This finding is gated on a365.config.json.aiTeammate === false, but the repository's non-AI path is identified by the detection cache's agentType: "system-agent"; the documented config lookup only guarantees blueprintId/agentBlueprintId. A normal make-a365-agent project without an aiTeammate key therefore skips this missing-registration warning even when agenticAppId is present and agentRegistrationId is absent. Use the authoritative agent-type signal and mirror the fix in the stop hook.
- evaluate Node A365 activation/exporter state per distro call
- keep route and resolver checks call-scoped with no file fallback
- use workspace detection cache for system-agent registration warnings
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5cbf5f6b-cc40-4b7e-a591-65848db73a12
nodeOptionsTexts read only the literal options object, so a route, resolver
or a365 block inside a same-file spread ({ ...baseOptions }) was missed, and
an unresolvable spread made the call look inactive. It now follows top-level
spreads to same-file object literals; any other spread keeps the call active.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5cbf5f6b-cc40-4b7e-a591-65848db73a12
When the Node checks run: both scanners now read enabled and enableObservabilityExporter from each useMicrosoftOpenTelemetry call's a365 options instead of searching for a literal enabled: true anywhere in the project. The skill's own enabled: A365_ENABLED snippet is now checked instead of skipped. Calls without a365 options stay quiet, as the distro never enables A365 for them, and options that can't be read are treated as active. Options spreads are followed.
Registration warning: it now uses .a365-workspace-detection.local.json (agentType: "system-agent") as the non-AI signal and falls back to a365.config.jsonaiTeammate: false only when the cache has no agent type. It never fires for AI Teammates. Copilot flagged this as "previously missed".
Helper comment: it no longer describes the removed whole-file fallback.
npm test: 205/205. The validate.yml checks pass locally.
a365Objects searches the entire options text with a regex, so it also discovers nested keys that are not the SDK's top-level a365 option. For example, useMicrosoftOpenTelemetry({ instrumentationOptions: { a365: { enabled: true, useS2SEndpoint: true, tokenResolver: r } } }) is treated as fully wired even though the distro ignores it, allowing the stop hook to pass a delegated/default-route configuration. Parse only top-level fields (the existing splitTopLevel helper is suitable here).
The standalone scanner has the same depth-insensitive a365Objects search. An a365 property nested under an unrelated option object is therefore accepted as the distro's A365 configuration, so both route and resolver findings can be suppressed while the SDK still uses its defaults. Restrict this helper to direct top-level fields and keep the parity test covering the nested-unrelated-object case.
Without the detection cache, the registration warning required an explicit
aiTeammate:false, so config-free blueprint setup (no a365.config.json) and
configs without the field never got it. The generated agenticAppId is written
only by blueprint-agent setup, so only an explicit aiTeammate:true now marks
an AI Teammate; the cached agentType stays authoritative.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 5cbf5f6b-cc40-4b7e-a591-65848db73a12
The reason will be displayed to describe this comment to others. Learn more.
Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
/observabilityService/with a token for the exporting agent identity. OBO / Agentic User tokens are used only for workload (MCP / Graph) calls. Registered blueprint agent instances need noAgent365.Observability.OtelWritepermission or admin consent (subject to service policy).observability/app-token-resolver.ts,observability/app_token_resolver.py, andObservability/AgentAppTokenResolver.cs. Each reuses the hosting connection's blueprint credential: an FMI assertion for the turn's agent identity, then agent-identityclient_credentialsfor the OBS scope.Microsoft.OpenTelemetry1.0.3+ flattenedo.Agent365.Exporter.*too.Agent365.*. The snippets now target the flattened API, and the 1.0.2-and-earlier form is noted.*-obs-delegated-route,*-obs-delegated-token,python-obs-prefetch-connection-missing, andagent-registration-not-recorded.a365 setup all --agent-registration-only). AI Teammates keep theOtelWriteapplication-role step thata365 setup all --aiteammateprints.Related
Testing
npm test: 165 tests pass (132 onmain, 33 new).validate.ymljob scripts pass on an LF checkout.tsc --strictagainst@microsoft/agents-hosting1.8.1 and@microsoft/opentelemetry1.4.0.Microsoft.OpenTelemetry1.1.0 andMicrosoft.Agents.*1.4.83. The .NET S2S scaffold also builds./observabilityService/with an app-only token and noOtelWritegrant. It got HTTP 200, and activity appeared in MAC.Notes
OtelWriteapplication-role step.