Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,7 @@
- Add GenAI v1.42 InvokeAgent request, response, cache-token, and provider attribute capture for manual A365 scopes. [#239](https://github.com/microsoft/opentelemetry-distro-javascript/pull/239)

### Other Changes
- Add offline SDK throughput and signed memory benchmarks with raw artifacts and explicitly configured named OTLP log result export.
- Consolidate Dependabot updates for Vitest 4.1.11, Hono 4.13.7, qs 6.16.0, fast-uri 3.1.7, actions/deploy-pages 5.0.1, and actions/checkout 7.0.1.
- Document local npm lockfile regeneration for contributors who cannot access the Microsoft package proxy, while retaining the proxy-generated lockfile. [#245](https://github.com/microsoft/opentelemetry-distro-javascript/pull/245)

Expand Down
85 changes: 85 additions & 0 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,3 +63,88 @@ the required proxy rather than using this workflow to bypass its restrictions.
- Link related issues when applicable.
- Update documentation when public behavior or setup changes.
- Keep the repository planning and README documents aligned with the implementation.

## SDK performance benchmarks

After building, run the standalone harness on Node.js 22 or later. It measures
recording span creation (with and without an attribute), counter aggregation,
and log emission through the built SDK. Non-exporting processors and a metric
reader keep serialization, network transport, and exporter batching out of the
measurement. Recording/aggregation probes fail rather than measuring no-op
providers. This is not an end-to-end exporter or application benchmark.

```sh
node --expose-gc perf/benchmark.mjs --output tmp/perf/raw.json --iterations 100000 --rounds 12 --memory-iterations 10000 --memory-trials 5
node perf/export-events.mjs --input tmp/perf/raw.json --output tmp/perf/events.json
```

Both commands are offline by default. The benchmark disables inherited
OpenTelemetry/Application Insights exporter and sampler configuration,
SDKStats, automatic instrumentation, and network resource discovery in its
dedicated processes. The VM resource detector is temporarily replaced during
SDK startup because the distro invokes it independently of the detector
environment setting. None of these changes affect normal library usage.
Use a clean Node process without instrumentation preloads.

`--package-root` defaults to the current directory and must contain the tested
`@microsoft/opentelemetry` manifest and `dist/esm/index.js`. It can also point to
an extracted npm package, with dependencies installed in that directory or an
ancestor. The manifest supplies the measured package name/version. Optional
`--revision` must identify that tested package's source revision, not the CI
orchestration repository; omit it if unknown. Optional `--run-id` supplies an
explicit execution correlation shared by the result events. Neither identifier
is inferred from a host, collector, timestamp, or session.

The options shown above are the defaults. Each throughput scenario has 20,000
warmup operations followed by the requested number of measured rounds, with GC
before each round. Raw output retains nanosecond durations, per-round operation
rates, and the existing `ns/op` samples/median consumed by `perf/compare.mjs`.
Only the original two span cases are regression-gating in that comparison.
Throughput telemetry is the **median of per-round rates**, in `operations/s`,
not the reciprocal of the median duration.

Each memory trial uses a fresh `--expose-gc` worker, warms up
`min(memory-iterations, 1000)` operations, settles GC three times, captures a
baseline, performs the requested operations, then captures immediate and
post-GC snapshots. The raw artifact retains all snapshots, counts, timestamps,
package identity, runtime/OS/architecture, and supplied provenance.
Memory results are medians of **signed deltas**, without clamping or noise
thresholds: immediate heap-used, post-GC retained heap, and immediate RSS minus
baseline, in `By`; heap and retained-heap per-operation deltas are in
`By/{operation}`. These are noisy process observations, not total allocated
bytes or a leak diagnosis. Negative and zero observations remain valid.

The second command validates the raw observations and writes native OTLP JSON
named log events (`microsoft.opentelemetry.benchmark.result`). Both throughput
and memory observations must match the canonical scenario names, test cases,
and categories defined in `perf/benchmark-sdk.mjs`, regardless of collection order.
Each event has `test.case.name`, `test.suite.name`, `benchmark.metric`, numeric `benchmark.value`,
`benchmark.unit`, `benchmark.statistic`, and actual iteration/round or
memory-trial counts. Case names are distinct from scenario labels. Metrics are
`microsoft.opentelemetry.benchmark.throughput` and
`microsoft.opentelemetry.benchmark.memory.{heap_used_delta,heap_used_delta_per_operation,retained_heap_delta,retained_heap_delta_per_operation,rss_delta}`.
The custom JSON reporting harness is identified by `telemetry.sdk.*` and
`service.name`, separately from the measured `package.name`/`package.version`.
Runtime/OS/architecture are measured resource attributes; optional provenance
uses `vcs.ref.head.revision` and `benchmark.run_id`.

Sending is a separate, explicit action for a trusted CI job or operator:

```sh
node perf/export-events.mjs --input tmp/perf/raw.json --output tmp/perf/events.json --endpoint https://YOUR-COLLECTOR/otlp/v1/logs
```

There is no default endpoint or environment-variable fallback. Keep the real
collector URL in private CI configuration, and gate submission separately from
offline benchmark execution (never submit untrusted PR results). HTTPS is
required except for loopback test servers. Credentials, query strings, and
redirects are not accepted. The request uses `Content-Type: application/json`;
timestamps/int64 attributes are strings and measurement `doubleValue` fields
are finite numbers. The default request timeout is 20,000 ms, configurable with
`--timeout-ms` up to 120,000 ms. Non-2xx responses, malformed success responses,
and partial success/error bodies fail the command. There are no automatic
retries because delivery may be uncertain. Both raw and generated payload files
remain available after export failure; archive them even on failed CI runs.
Output files are replaced atomically, and existing aliases of the raw input
(including hardlinks and symlinks) are rejected rather than overwritten.
HTTP success alone does not prove downstream ingestion or dashboard refresh.
146 changes: 146 additions & 0 deletions perf/benchmark-sdk.mjs
Original file line number Diff line number Diff line change
@@ -0,0 +1,146 @@
// Copyright (c) Microsoft Corporation.
// Licensed under the MIT License.

import assert from "node:assert/strict";
import { createRequire } from "node:module";
import { join } from "node:path";
import { pathToFileURL } from "node:url";

export const scenarios = [
{ name: "span", test: "span_creation", category: "span" },
{
name: "span_with_attribute",
test: "span_creation_with_attribute",
category: "span",
},
{ name: "counter_add", test: "metric_counter_add", category: "metric" },
{ name: "logger_emit", test: "log_emit", category: "log" },
];

export async function startBenchmarkSdk(packageRoot) {
// This helper runs only in dedicated benchmark processes, never in the library.
for (const key of Object.keys(process.env)) {
if (
/^(OTEL_|APPLICATIONINSIGHTS_|AZURE_MONITOR_|MICROSOFT_OTEL_|A365_|ENABLE_A365_)/i.test(key)
) {
delete process.env[key];
}
}
process.env.MICROSOFT_OTEL_SDKSTATS_DISABLED = "true";
process.env.APPLICATIONINSIGHTS_STATSBEAT_DISABLED_ALL = "true";
process.env.OTEL_NODE_RESOURCE_DETECTORS = "none";
process.env.OTEL_TRACES_SAMPLER = "always_on";

const requireFromPackage = createRequire(join(packageRoot, "package.json"));
const { metrics, trace } = requireFromPackage("@opentelemetry/api");
const { logs } = requireFromPackage("@opentelemetry/api-logs");
const { MetricReader } = requireFromPackage("@opentelemetry/sdk-metrics");
const { azureVmDetector } = requireFromPackage("@opentelemetry/resource-detector-azure");
// InternalConfig invokes this detector even with OTEL_NODE_RESOURCE_DETECTORS=none.
// Resource discovery/startup is not measured; suppress its metadata HTTP request.
const originalDetect = azureVmDetector.detect;
azureVmDetector.detect = () => ({ attributes: {} });

class NonExportingMetricReader extends MetricReader {
async onForceFlush() {}
async onShutdown() {}
}
Comment thread
JacksonWeber marked this conversation as resolved.
const metricReader = new NonExportingMetricReader();
const logRecordProcessor = {
enabled: () => true,
forceFlush: async () => {},
onEmit: () => {},
shutdown: async () => {},
};
const spanProcessor = {
forceFlush: async () => {},
onStart: () => {},
onEnd: () => {},
shutdown: async () => {},
};
let shutdown;
try {
const sdk = await import(pathToFileURL(join(packageRoot, "dist", "esm", "index.js")).href);
shutdown = sdk.shutdownMicrosoftOpenTelemetry;
sdk.useMicrosoftOpenTelemetry({
azureMonitor: { enabled: false },
a365: { enabled: false },
enableConsoleExporters: false,
instrumentationOptions: Object.fromEntries(
[
"azureSdk",
"bunyan",
"console",
"http",
"langchain",
"mongoDb",
"mySql",
"openaiAgents",
"postgreSql",
"redis",
"redis4",
"winston",
].map((name) => [name, { enabled: false }]),
),
logRecordProcessors: [logRecordProcessor],
metricReaders: [metricReader],
samplingRatio: 1,
spanProcessors: [spanProcessor],
tracesPerSecond: 0,
});
const tracer = trace.getTracer("performance-test");
const logger = logs.getLogger("performance-test");
const counter = metrics.getMeter("performance-test").createCounter("benchmark-counter");
const probe = tracer.startSpan("benchmark-probe");
assert(probe.isRecording(), "Benchmark requires a recording tracer");
probe.end();
let emitted = false;
logRecordProcessor.onEmit = () => {
emitted = true;
};
logger.emit({ body: "benchmark-probe" });
logRecordProcessor.onEmit = () => {};
assert(emitted, "Benchmark requires a recording logger");
counter.add(1);
const collected = await metricReader.collect();
assert.equal(collected.errors.length, 0, "Benchmark metric collection failed");
assert(
collected.resourceMetrics.scopeMetrics.some((scope) =>
scope.metrics.some(
(metric) =>
metric.descriptor.name === "benchmark-counter" &&
metric.dataPoints.some((point) => point.value === 1),
),
),
"Benchmark requires an aggregating counter",
);
const operations = [
() => {
tracer.startSpan("benchmark-span").end();
},
() => {
const span = tracer.startSpan("benchmark-span");
span.setAttribute("benchmark.attribute", 1);
span.end();
},
() => {
counter.add(1);
},
() => {
logger.emit({ body: "benchmark-log" });
},
];
return {
scenarios: scenarios.map((scenario, index) => ({
...scenario,
operation: operations[index],
})),
shutdown,
};
} catch (error) {
await shutdown?.();
throw error;
} finally {
azureVmDetector.detect = originalDetect;
}
}
Loading
Loading