Skip to content

About

Passively extract the reformulated web-search sub-queries AI assistants run behind the scenes, from your own browser session (read-only, local-only, via CDP).

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

ai-subquery-inspector

Modern AI assistants don't answer from memory — they rewrite your prompt into one or more web-search queries, fetch results, and synthesize. This tool passively observes your own authenticated browser session over the Chrome DevTools Protocol (CDP) and extracts those reformulated sub-queries as they stream to your client, storing them locally as JSONL and giving you commands to analyze how each engine decomposes intent. It attaches read-only to a Chrome you launch with remote debugging, tees the response bodies of the assistants' own streaming API calls, and decodes each engine's payload with a small per-engine parser. It does not submit prompts, run headless farms, modify requests, or touch any account but the one you are already using.


⚠️ Scope & ethics — read first

  • Your own browser, your own machine, your own sessions — only. This is a personal introspection tool for traffic already flowing to a session you are logged into. It is not built for, and must not be used for, capturing anyone else's activity.
  • No guidance is provided for pointing this at someone else's browser or a shared/managed machine, and you should not. Attaching to another party's browser session via CDP without their knowledge or consent is a serious invasion of privacy and is unlawful in many jurisdictions (wiretap/computer- misuse laws). Don't.
  • Local-only by design. Everything is written to a local JSONL file. There is no remote collection, no multi-user mode, no centralized upload, and none will be added. If you choose to export or share your capture store, that is an explicit, manual act on your part.
  • The capture store is sensitive. It contains everything you searched through these assistants. Treat it like your browser history. It is git-ignored by default (see Data handling & privacy).

What it captures

Six AI assistants are supported. Detection is two-tier: a broad domain filter picks which responses to read at all, then a precise per-engine parser decodes the transport and reports the JSON path each query was found at (it never asserts a fabricated path — a key rename surfaces loudly as drift, never as a silent empty).

Engine Endpoint / transport Where the sub-queries live Status
ChatGPT …/backend-api/f/conversation — SSE (delta-encoded) message.metadata.search_model_queries.queries (also reconstructs the safe_urls array) ✅ extracts
Gemini …/StreamGenerate — Google RPC envelope Not exposed client-side — only the answer + source citations are. Reported honestly as subqueriesExposed: false, not treated as an error. ⓘ sources only
Perplexity …/rest/sse/perplexity_ask — SSE blocks[].workflow_block.steps[].items[].payload.queries_payload.queries ✅ extracts
Claude …/chat_conversations/{uuid}/completion — SSE web_search tool-use block, reconstructed from streamed input_json_delta fragments ✅ extracts
DeepSeek …/api/v0/chat/completion — SSE SEARCH fragment response.fragments[].queries[].query (requires the "Search" toggle on) ✅ extracts
Grok …/load-responses — single JSON responses[].steps[].toolUsageCards[].webSearch.args.query ✅ extracts

Each engine encodes its queries differently (a full array, a queries_payload object, query text split across streaming JSON deltas, or tool-call cards), which is exactly why there is one parser module per engine. Full per-engine findings are in DISCOVERY.md. Gemini is a genuine finding, not a TODO: it does the grounding server-side and never ships the reformulated queries to the browser, so there is nothing to capture passively.

Selecting engines: capture all of them (default) or a subset with --engines chatgpt,perplexity (or AIQI_ENGINES).


Quick start

Full, zero-context walkthrough (prerequisites → first capture → troubleshooting) is in SETUP.md. The short version:

npm install          # 3 packages, no bundled browser
npm test             # offline: decode/extract plumbing + schema-drift check

# 1. Launch a dedicated debug Chrome and log into the AI apps once
node bin/cli.js launch

# 2. In another terminal, attach and capture
node bin/cli.js capture --no-launch
#    …now ask ChatGPT/Perplexity/etc. something that triggers a web search…

# 3. Read it back
node bin/cli.js groups

Prefer to explore first? Every analysis command runs against the bundled synthetic dataset — see examples/:

node bin/cli.js frequency --data examples/sample-captures.jsonl --app chatgpt

Why these design choices

Decision Choice Why
Runtime Node.js (≥18.3) CDP is JSON-over-WebSocket; Node is the natural, dependency-light fit. Uses the built-in parseArgs.
CDP client chrome-remote-interface The thin reference CDP client (deps: ws, commander). We deliberately avoid Puppeteer/Playwright: they bundle a browser and dozens of automation-oriented deps we don't need for passive eventing. Total install: 3 packages.
Attachment Flatten auto-attach (Target.setAutoAttach{flatten:true}) over one connection, with nested per-page auto-attach Covers pages and their Web/service workers as sessions on a single socket — required because some apps (Perplexity, Grok) stream from a worker a page-only attach never sees.
Streaming capture Network.streamResourceContent + flush on abort Passively tees the streamed body without pausing or altering what the page receives (unlike Fetch response-stage interception, which would stall SSE). Streams that the client cancels on completion end as loadingFailed(net::ERR_ABORTED), not loadingFinished, so we flush accumulated chunks on abort too. Failure mode: on a Chrome too old for streamResourceContent, we fall back to getResponseBody and surface a loud body_unavailable record — never a silent drop.
Storage JSONL Append-only durability (a crash costs ≤1 trailing line), zero dependencies, trivially re-parseable when a parser improves. No native SQLite module needed for this data volume (your own manual prompts).

Data model — exactly what is recorded

Each extracted sub-query is one JSONL line (type: "subquery"):

{
  "timestamp": "2026-08-01T10:00:00.000Z",
  "type": "subquery",
  "app": "chatgpt",
  "transport": "sse",
  "endpoint": "https://chatgpt.com/backend-api/f/conversation",
  "conversationId": "…|null",
  "prompt": "the original prompt, if observable in the request|null",
  "subQuery": "the reformulated search query",
  "seqIndex": 0,                              // position within this turn's fan-out
  "queryCount": 3,                            // total queries this turn
  "resultUrls": ["…"],                        // result/citation URLs, best-effort
  "resultTitles": ["…"],
  "patterns": ["year", "best_top"],           // fan-out pattern tags (see below)
  "sourceTypes": { "reddit": 2, "docs": 1 },  // source-type mix of resultUrls
  "safeUrls": ["https://…"],                  // ChatGPT only: "verified safe" URLs (turn-level)
  "jsonPath": "message.metadata.search_model_queries.queries",  // where it was found
  "shapeVersion": "2026-09-06-verified",
  "verified": true,                            // is this parser confirmed against live traffic?
  "tabUrl": "…",
  "rawFragment": "first 4 KB of the raw body, kept for re-parsing"
}

What is discarded: the full response body (only the first 4 KB is retained as rawFragment), all request/response headers, cookies, and auth tokens. The tool records the query text and its provenance, not the traffic.

Parser failures are never swallowed — they are written as type: "schema_drift" records (parse_error, body_unavailable, or zero_extracted_with_signal) so a payload change shows up as a loud record rather than a clean-looking empty result.


Analysis commands

All analysis is pure post-processing over the JSONL — read-only, nothing extra is fetched. Every command below is shown running against the bundled synthetic sample; see examples/README.md for full output.

  • query — flat list of captured sub-queries (or --drift for drift records).
    node bin/cli.js query --app chatgpt --contains desk
    # 2026-08-01T…  chatgpt   [1/3]  best budget standing desk 2026  [year,best_top,price]
  • groups — the cross-app comparison view: prompt → its sub-queries, grouped by turn, with result URLs.
    node bin/cli.js groups --app perplexity
  • frequency — fan-out is probabilistic (same prompt, different queries each run), so a single capture is noise. Run a prompt several times, then this aggregates by normalized prompt and reports how often each sub-query appears. It warns below 3 runs.
    node bin/cli.js frequency --app chatgpt
    #   100%  (3/3)  best budget standing desk 2026
    #    67%  (2/3)  standing desk reviews reddit
  • language-pack — per engine: top uni/bi/tri-grams (as distinct-queries/total-occurrences), the pattern-tag distribution, and modifiers (year values, site: targets, boolean-query count).
    node bin/cli.js language-pack --app chatgpt --top 15
    #   patterns:  price 55% · year 45% · best_top 27% · reviews 27% · reddit 18%
  • safe-urls — ChatGPT streams a safe_urls array ("URLs verified as safe to visit") per turn. This diffs it across captured runs and flags domains that appear in many distinct prompts while being topically unrelated to them — a signal a domain has been persisted across unrelated queries.
    node bin/cli.js safe-urls
    #   promo-tracker.example   2   2   0%   ⚠ persists across unrelated prompts
  • drift — re-parses the saved per-engine fixtures and exits non-zero if any engine's payload shape has shifted out from under its parser (npm run drift).

Each sub-query is tagged against a fan-out pattern taxonomy (src/patterns.js): year · best_top · reddit · reviews · comparison · site_operator · price · location · direct_official · question · boolean. Result URLs are classified into source types (src/sources.js): reddit · docs · review_platform · social · qa · video · wiki · news · web. Both are conservative by design — a missed tag is recoverable; a wrong tag pollutes the analysis.

"Pattern drift" vs "schema drift": frequency/language-pack reveal how an engine's query behavior changes over time; the drift command guards whether a parser still matches the payload shape. Different things.


Use cases

  • Personal research retrospectives — see the actual queries an assistant ran on your behalf, not just its prose answer.
  • Prompt-pattern analysis — study how your phrasing gets rewritten, and which modifiers (year, site:, boolean) each engine bolts on.
  • Tracking your own information-seeking over time — watch how your topics and an engine's fan-out behavior shift across weeks.
  • Comparing engines — put the same prompt to several assistants and compare how each decomposes it (the groups view).
  • Building a personal query corpus — a local, cross-engine store of the search intent behind your AI usage.
  • Generative-engine-optimization research — the pattern/frequency/language-pack analyses were built to study how engines rewrite intent (feed high-frequency terms into your own content). This is analysis of your own captures, offline.

Strengths

  • Passive — no manual logging; it records what actually happened.
  • Cross-engine, one store — six assistants normalized into one JSONL schema.
  • Local and private — nothing leaves your machine.
  • Honest under change — loud drift records + a fixture regression, so you know when an engine's payload moved instead of silently capturing nothing.

Limitations & failure modes

  • Breaks when an engine changes its DOM/network shape. These are private, undocumented APIs; a payload rename can zero out a parser. The generic-probe fallback and schema_drift records make this loud, not silent, but the parser still needs a refresh (see CONTRIBUTING.md).
  • CDP version drift. Very old Chrome lacks streamResourceContent; you'll get body_unavailable records until you update Chrome.
  • Headless vs headed differences. Designed for a headed session you log into and drive by hand; headless behavior is not a supported path.
  • Navigation races. A response that completes before Network.enable lands on a just-spawned worker can be missed; the tool pauses new targets on start to minimize this, but it isn't zero.
  • No non-browser clients. Desktop and mobile apps and API usage are invisible to it — it only sees a Chrome session.
  • Storage growth. The JSONL append-only store grows with use (see below).
  • Per-engine maintenance burden. Six hand-written parsers against moving targets; keeping them green is ongoing work.
  • Some engines don't expose the queries at all (Gemini today). That's a finding, reported as such, not something worked around.

Full detail in LIMITATIONS.md.


What to pay attention to (security)

  • A browser with remote debugging enabled is itself an exposure. Anything that can reach the debug port can drive the browser and read its traffic. The tool binds the port to 127.0.0.1 only — do not expose it, forward it, or bind it to a non-loopback interface.
  • The capture store contains everything you searched. Treat it as sensitive; back it up deliberately or exclude it deliberately.
  • Use a dedicated Chrome profile. Not your daily profile (recent Chrome blocks remote debugging on the default profile anyway; see LIMITATIONS.md).

Data handling & privacy

  • Captured records are written to data/captures.jsonl by default (override with --data / AIQI_DATA). The data/ directory and all *.jsonl files (except the synthetic examples/) are git-ignored so you cannot accidentally commit your captures.
  • Nothing is transmitted anywhere. There is no telemetry and no network egress beyond the local CDP connection to your own Chrome.
  • No credentials are ever read or stored — the tool records query text and provenance, and discards headers/cookies/tokens.

Cost

This is free and open-source, and runs entirely locally. There are no paid API calls anywhere in the tool — capture and all analysis are offline. The only real resource cost is local disk: the JSONL store grows roughly with the number of sub-queries you capture (each record is a few hundred bytes plus up to a 4 KB rawFragment). Prune or rotate the file yourself if it grows large — the tool intentionally does not delete your data for you (see the retention note in LIMITATIONS.md).


Configuration

Everything is a CLI flag, with optional environment-variable fallbacks (a flag always wins). See .env.example:

Flag Env Default
--port AIQI_PORT 9222
--profile AIQI_PROFILE ~/.ai-subquery-inspector/chrome-profile
--chrome-path AIQI_CHROME_PATH auto-detected
--data AIQI_DATA <repo>/data/captures.jsonl
--engines AIQI_ENGINES all registered
--top / --min-count AIQI_TOP / AIQI_MIN_COUNT 15 / 1
--min-prompts AIQI_MIN_PROMPTS 2

Adding an engine is a code change (one parser module), not a config entry, because each engine's payload shape is different — see CONTRIBUTING.md.

Project layout

bin/cli.js            # command dispatch
src/
  cdp.js              # CDP attach + passive network tee
  chrome.js           # launch a debug Chrome at a dedicated profile
  capture.js          # onBody pipeline: parse → enrich → store
  storage.js          # append-only JSONL read/write
  registry.js         # parser registry + domain matcher
  config.js           # env-var fallbacks (non-behavioral)
  parsers/            # one module per engine + shared probe + _template
  patterns.js sources.js aggregate.js langpack.js safeurls.js   # analysis
test/                 # offline smoke test + schema-drift regression + fixtures
tools/discovery/      # passive-capture harness + method playbook
examples/             # synthetic sample dataset + generator

License

MIT — see LICENSE.

About

Passively extract the reformulated web-search sub-queries AI assistants run behind the scenes, from your own browser session (read-only, local-only, via CDP).

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages