Skip to content

feat(enrichers): add org_to_news enricher backed by Serply - #227

Open
googio wants to merge 2 commits into
reconurge:mainfrom
googio:feat/enrichers/org-to-news
Open

googio wants to merge 2 commits into
reconurge:mainfrom
googio:feat/enrichers/org-to-news

Conversation

@googio

@googio googio commented Sep 19, 2026

Copy link
Copy Markdown

What

Adds org_to_news, an Organization to Website enricher that asks a search
index for news coverage of an organization. It runs the name as an
exact-phrase query against the news vertical of the Serply API and emits one
Website node per article, related back to the input Organization with
MENTIONED_IN.

Why

Third and last of the three enrichers from #220, after #223 (domain) and #225
(individual).

The three Organization enrichers today are org_to_domains (Whoxy),
org_to_infos (SIRENE) and org_to_asn (asnmap). All three are registry or
infrastructure pivots: they tell you what the organization filed and what it
runs. None of them tell you what has been said about it. Coverage is often
where the next pivot actually comes from, because an article names the people,
the partners, the acquisition, the breach and the locations that no registry
records, and it is the one source that exists for an organization too small or
too foreign to appear in SIRENE.

@dextmorgn this completes the set you green-lit on #220. All three are
independent PRs and can be merged, reordered or dropped one at a time.

Notes

  • Behavior is unchanged without a key. SERPLY_API_KEY is declared as a
    required vaultSecret, so the enricher is inert until a key exists in the
    Vault, same as domain_to_dehashed and ip_to_fraudscore.
  • No new HTTP client and no new vertical plumbing. News comes off the same
    endpoint as feat(enrichers): add domain_to_indexed_pages enricher backed by Serply #223 and feat(enrichers): add individual_to_mentions enricher backed by Serply #225, with the vertical selected by Google's own tbm
    parameter.
  • MENTIONED_IN is the same label feat(enrichers): add individual_to_mentions enricher backed by Serply #225 introduces, for the same reason: an
    article covering an organization mentions it, it is not a site the
    organization owns, so this is not the HAS_WEBSITE edge domain_to_website
    writes. That is the one cross-PR coupling here. If you would rather the label
    were named something else, say so on either PR and I will change both. If
    only one of the two merges the label still works, since neither reads the
    other's edges.
  • Language and market are pinned, by default en / us. This is not
    boilerplate: left unset, the vertical infers them from where the call
    originates, and a server is not a place. An English-language query for a
    well-known company came back as ten Japanese-language articles plus one row
    whose link was a thumbnail image rather than an article. With hl and gl
    set, all ten rows were articles. Both are exposed as params so a French
    investigation can ask for fr / fr.
  • time_range maps onto the index's own recency filter (tbs=qdr:d|w|m|y),
    which matters more here than on a web search: "the last week of coverage" is
    usually the question. An unrecognised value warns and searches all of time
    rather than failing the scan.
  • Publication date is deliberately not persisted. The vertical does return
    one, but Website has no date field, and I would rather not stuff it into
    description or invent a field in flowsint-types inside a PR that is about
    an enricher. Same for Website.domain, which I leave to
    website_to_domain. Both are a small follow-up if you want them, and adding
    a published_at to Website is the kind of change I would rather you
    decided than discovered.
  • Structure copies domain/to_dns.py and ip/to_fraudscore.py: registration
    via the decorator, no registry.py edit, per-item try/except, Logger.error
    on every failure path. The source name rides to postprocess on a temporary
    attribute and is deleted there, the way domain_to_dns threads its source
    domain, because one scan can cover several organizations and the results
    cannot be zipped back to their input.
  • Organization.name is typed Any, so it can arrive empty or unset from an
    upstream pivot. It is coerced and checked before anything is sent, and an
    organization with no usable name is skipped with a warning, so it costs no
    credit. Pinned by a test over "", " " and None.
  • Five params is more than most enrichers here carry; ip_to_ports carries
    six, so it is not unprecedented. Four of the five are optional with
    defaults.
  • Paging is the same as feat(enrichers): add domain_to_indexed_pages enricher backed by Serply #223 and feat(enrichers): add individual_to_mentions enricher backed by Serply #225: start is the only offset the API
    honours, num is an approximate cap of about ten rows per call, and a window
    that adds no new link ends the loop instead of spending a credit per
    duplicate window. Surplus rows are trimmed against max_results (default
    20).
  • Tests in flowsint-enrichers/tests/enrichers/test_org_to_news.py, HTTP layer
    mocked, no key and no network needed. Docs row added in
    docs/sources/available-enrichers.mdx, under ### Organization, so it does
    not collide with the ### Individual row feat(enrichers): add individual_to_mentions enricher backed by Serply #225 adds to the same file.

Testing

$ uv run ruff format --check $(PY_SRC)
328 files already formatted
$ uv run ruff check $(PY_SRC)
All checks passed!
$ make typecheck BASE_REF=upstream/main
== mypy: flowsint-enrichers ==
Success: no issues found in 2 source files
$ cd flowsint-enrichers && uv run pytest -q
55 passed, 2 warnings in 1.46s

55 is the whole flowsint-enrichers suite, 32 before this branch plus the 23
new ones. I did not run the frontend half of make lint, since no frontend
file changed here.

Live run, real API key against a local Neo4j 5, max_results 12 and
time_range "m" on "Mistral AI". 12 is over one window, so the second page was
fetched and merged:

scan() returned 12 Website objects
nodes in neo4j:  organization 1, website 12
edges:           MENTIONED_IN 12

(Mistral AI) -[MENTIONED_IN]-> https://www.reuters.com/world/europe/french-ai-...
(Mistral AI) -[MENTIONED_IN]-> https://www.wsj.com/tech/ai/mistral-ai-exceeds-...
(Mistral AI) -[MENTIONED_IN]-> https://techcrunch.com/2026/09/08/mistral-raise...
(Mistral AI) -[MENTIONED_IN]-> https://www.nytimes.com/2026/09/08/business/mis...
(Mistral AI) -[MENTIONED_IN]-> https://news.samsung.com/global/samsung-and-mis...
...

Every row was inside the one-month window the run asked for, and title and
snippet land on the node:

nodeProperties.title        French AI company Mistral hits $24 billion
                            valuation in funding round
nodeProperties.description  France's Mistral has raised EUR 3 billion at a
                            valuation of around EUR 21 billion ($24 billion),
                            in what the three-year-old AI company said [...]

Size

3 files, +598. This is the largest of the three and I would rather explain the
number than hide it. 291 lines are the test file and 3 are the docs row, so the
enricher itself is 304. It is longer than #225's 260 because the news vertical
has three knobs a plain web search does not (vertical selection, recency,
locale), and each one carries its schema entry, its mapping and its test.

For reference: #225 was +514, #223 +416, and the merged domain_to_dns (#182)
+336 and #183 +305 across the same three kinds of file.

Happy to cut the test file down if you would rather review it lighter. The
tests I would defend keeping are the locale one and the time_range ones,
since both pin behavior that is invisible in the output until it is wrong, and
the one that pins the source-name threading across several organizations in one
scan.

Disclosure

I work with Serply. Happy to adjust scope, naming, or drop this entirely if it
is not a direction you want for the project.

Pivots an Organization to the news coverage that mentions it, emitting one
Website per article and relating it back with MENTIONED_IN.

The existing Organization enrichers are registry and infrastructure pivots
(Whoxy, SIRENE, asnmap), so none of them surface what has been said about an
organization publicly. Coverage names the people, partners and incidents that
no registry records.

Language and market are pinned by default, since the vertical otherwise infers
them from where the call originates and answers an English query with articles
in another language. Recency is exposed as a time_range param.
The frontend needs a flag for whether an enricher takes params at all. Follows
the existing override in ip/to_asn.py, email/to_leaks.py and six other
enrichers, matching the fix already applied to reconurge#223.
@googio

googio commented Sep 20, 2026

Copy link
Copy Markdown
Author

Pulling your required_params note from #223 forward to this PR so you do not have to write it twice.

Added in 8c3b309, same shape as the override in ip/to_asn.py and email/to_leaks.py. This enricher declares a required SERPLY_API_KEY vault secret, so True is the right answer here. One test added alongside it; the file is now 24 passing, with ruff and mypy clean.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant