Skip to content

feat(enrichers): add individual_to_mentions enricher backed by Serply - #225

Open
googio wants to merge 2 commits into
reconurge:mainfrom
googio:feat/enrichers/individual-to-mentions
Open

googio wants to merge 2 commits into
reconurge:mainfrom
googio:feat/enrichers/individual-to-mentions

Conversation

@googio

@googio googio commented Sep 19, 2026

Copy link
Copy Markdown

What

Adds individual_to_mentions, an Individual to Website enricher that asks a
search index which public pages mention a person by name. It runs the full name
as an exact-phrase query through the Serply API and emits one Website node per
organic result, related back to the input Individual with MENTIONED_IN.

Why

Second of the three enrichers from #220. individual_to_domains and
individual_to_organization both start from a registry, so they only find a
person where a record already names them: a WHOIS registrant, a SIRENE leader.
This pivot finds the person where nobody registered anything, which is most of
the open web: a staff page, a conference bio, an article byline, a forum post.

@dextmorgn as promised. #223 is still the one that matters for the shape, and
one thing there is still open: whether discovered pages should reuse
HAS_WEBSITE or get their own label. This PR does not depend on that answer,
because a mention is a different relationship either way (see Notes), so it can
be reviewed on its own. If your answer on #223 changes anything here I will
follow it.

Notes

  • Behavior is unchanged without a key. SERPLY_API_KEY is declared as a
    required vaultSecret, so the enricher is inert until a key exists in the
    Vault, same as domain_to_dehashed and ip_to_fraudscore.
  • MENTIONED_IN is a new label, deliberately. I grepped the 34 labels in
    use first. HAS_WEBSITE is the closest, but it is what domain_to_website
    writes for a site a domain owns, and a name appearing on a page is evidence
    of a mention, not of ownership. Asserting ownership from a search hit would
    put an investigator's employer, a namesake and a news article on the same
    edge as their own site. If you would rather not grow the label set, I will
    switch it to whatever you prefer.
  • The name is quoted. "Jane Doe" matches the phrase; the unquoted form
    also matches a page carrying an unrelated Jane and an unrelated Doe, which is
    most of the noise on a name search. For the rest there is an optional
    additional_terms param, so a common name can be narrowed with an employer or
    a city.
  • Individual.compute_label back-fills full_name from first_name /
    last_name, so an individual carrying only the parts is still searchable and
    there is no need to reassemble the name here. An individual with no name at
    all (one pivoted from an email, say) is skipped with a warning before any
    request goes out, so it costs no credit. Both are pinned by a test.
  • Structure copies domain/to_dns.py and ip/to_fraudscore.py: registration
    via the decorator, no registry.py edit, per-item try/except, Logger.error
    on every failure path. The source name rides to postprocess on a temporary
    attribute and is deleted there, the way domain_to_dns threads its source
    domain, because one scan can cover several people and the results cannot be
    zipped back to their input.
  • Paging is the same as feat(enrichers): add domain_to_indexed_pages enricher backed by Serply #223: start is the only offset the API honours, num
    is an approximate cap of about ten rows per call, and a window that adds no
    new link ends the loop instead of spending a credit per duplicate window.
    Surplus rows are trimmed against max_results (default 20).
  • Tests in flowsint-enrichers/tests/enrichers/test_individual_to_mentions.py,
    HTTP layer mocked, no key and no network needed. Docs row added in
    docs/sources/available-enrichers.mdx.

Testing

$ uv run ruff format --check $(PY_SRC)
328 files already formatted
$ uv run ruff check $(PY_SRC)
All checks passed!
$ make typecheck BASE_REF=upstream/main
== mypy: flowsint-enrichers ==
Success: no issues found in 2 source files
$ cd flowsint-enrichers && uv run pytest -q
46 passed, 2 warnings in 1.33s

46 is the whole flowsint-enrichers suite, 32 before this branch plus the 14
new ones. I did not run the frontend half of make lint, since no frontend file
changed here.

Live run, real API key against a local Neo4j 5, max_results 12 on "Andrej
Karpathy". 12 is over one window, so the second page was fetched and merged:

scan() returned 12 Website objects
nodes in neo4j:  individual 1, website 12
edges:           MENTIONED_IN 12

(Andrej Karpathy) -[MENTIONED_IN]-> (https://karpathy.ai/)
(Andrej Karpathy) -[MENTIONED_IN]-> (https://en.wikipedia.org/wiki/Andrej_Karpathy)
(Andrej Karpathy) -[MENTIONED_IN]-> (https://www.linkedin.com/in/andrej-karpathy-9a650716)
(Andrej Karpathy) -[MENTIONED_IN]-> (https://www.youtube.com/andrejkarpathy)
(Andrej Karpathy) -[MENTIONED_IN]-> (https://x.com/karpathy?lang=en)
(Andrej Karpathy) -[MENTIONED_IN]-> (https://cs.stanford.edu/people/karpathy/)
(Andrej Karpathy) -[MENTIONED_IN]-> (http://karpathy.github.io/)
(Andrej Karpathy) -[MENTIONED_IN]-> (https://scholar.google.com/citations?user=l8WuQJgAAAAJ&hl=en)
...

Title and description land on the node, not just the URL:

nodeProperties.title        Andrej Karpathy - Wikipedia
nodeProperties.description  Andrej Karpathy (born 23 October 1986) is a
                            Slovak-Canadian AI researcher, who co-founded ...

Size

3 files, +514. Above the small-PR norm, and the split is the reason: 251 of
those lines are the test file and 3 are the docs row, so the enricher itself is
260. That is in line with #223 (+416 across the same three kinds of file) and
with the merged domain_to_dns (#182, +336) and #183 (+305).

Happy to cut the test file down if you would rather review it lighter. The two
tests I would defend keeping are the one that pins the source-name threading
across several people in one scan, and the one that pins the compute_label
back-fill, since both are behavior a future edit could break silently.

Disclosure

I work with Serply. Happy to adjust scope, naming, or drop this entirely if it
is not a direction you want for the project.

The frontend needs a flag for whether an enricher takes params at all. Follows
the existing override in ip/to_asn.py, email/to_leaks.py and six other
enrichers, matching the fix already applied to reconurge#223.
@googio

googio commented Sep 20, 2026

Copy link
Copy Markdown
Author

Pulling your required_params note from #223 forward to this PR so you do not have to write it twice.

Added in 700af02, same shape as the override in ip/to_asn.py and email/to_leaks.py. This enricher declares a required SERPLY_API_KEY vault secret, so True is the right answer here. One test added alongside it; the file is now 15 passing, with ruff and mypy clean.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant