Conversation
googio
force-pushed
the
feat/enrichers/individual-to-mentions
branch
from
September 19, 2026 13:33
d2045d7 to
a500ff7
Compare
This was referenced Sep 19, 2026
The frontend needs a flag for whether an enricher takes params at all. Follows the existing override in ip/to_asn.py, email/to_leaks.py and six other enrichers, matching the fix already applied to reconurge#223.
Author
|
Pulling your Added in |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds
individual_to_mentions, an Individual to Website enricher that asks asearch index which public pages mention a person by name. It runs the full name
as an exact-phrase query through the Serply API and emits one Website node per
organic result, related back to the input Individual with
MENTIONED_IN.Why
Second of the three enrichers from #220.
individual_to_domainsandindividual_to_organizationboth start from a registry, so they only find aperson where a record already names them: a WHOIS registrant, a SIRENE leader.
This pivot finds the person where nobody registered anything, which is most of
the open web: a staff page, a conference bio, an article byline, a forum post.
@dextmorgn as promised. #223 is still the one that matters for the shape, and
one thing there is still open: whether discovered pages should reuse
HAS_WEBSITEor get their own label. This PR does not depend on that answer,because a mention is a different relationship either way (see Notes), so it can
be reviewed on its own. If your answer on #223 changes anything here I will
follow it.
Notes
SERPLY_API_KEYis declared as arequired
vaultSecret, so the enricher is inert until a key exists in theVault, same as
domain_to_dehashedandip_to_fraudscore.MENTIONED_INis a new label, deliberately. I grepped the 34 labels inuse first.
HAS_WEBSITEis the closest, but it is whatdomain_to_websitewrites for a site a domain owns, and a name appearing on a page is evidence
of a mention, not of ownership. Asserting ownership from a search hit would
put an investigator's employer, a namesake and a news article on the same
edge as their own site. If you would rather not grow the label set, I will
switch it to whatever you prefer.
"Jane Doe"matches the phrase; the unquoted formalso matches a page carrying an unrelated Jane and an unrelated Doe, which is
most of the noise on a name search. For the rest there is an optional
additional_termsparam, so a common name can be narrowed with an employer ora city.
Individual.compute_labelback-fillsfull_namefromfirst_name/last_name, so an individual carrying only the parts is still searchable andthere is no need to reassemble the name here. An individual with no name at
all (one pivoted from an email, say) is skipped with a warning before any
request goes out, so it costs no credit. Both are pinned by a test.
domain/to_dns.pyandip/to_fraudscore.py: registrationvia the decorator, no
registry.pyedit, per-item try/except,Logger.erroron every failure path. The source name rides to
postprocesson a temporaryattribute and is deleted there, the way
domain_to_dnsthreads its sourcedomain, because one scan can cover several people and the results cannot be
zipped back to their input.
startis the only offset the API honours,numis an approximate cap of about ten rows per call, and a window that adds no
new link ends the loop instead of spending a credit per duplicate window.
Surplus rows are trimmed against
max_results(default 20).flowsint-enrichers/tests/enrichers/test_individual_to_mentions.py,HTTP layer mocked, no key and no network needed. Docs row added in
docs/sources/available-enrichers.mdx.Testing
46 is the whole
flowsint-enricherssuite, 32 before this branch plus the 14new ones. I did not run the frontend half of
make lint, since no frontend filechanged here.
Live run, real API key against a local Neo4j 5,
max_results12 on "AndrejKarpathy". 12 is over one window, so the second page was fetched and merged:
Title and description land on the node, not just the URL:
Size
3 files, +514. Above the small-PR norm, and the split is the reason: 251 of
those lines are the test file and 3 are the docs row, so the enricher itself is
260. That is in line with #223 (+416 across the same three kinds of file) and
with the merged
domain_to_dns(#182, +336) and #183 (+305).Happy to cut the test file down if you would rather review it lighter. The two
tests I would defend keeping are the one that pins the source-name threading
across several people in one scan, and the one that pins the
compute_labelback-fill, since both are behavior a future edit could break silently.
Disclosure
I work with Serply. Happy to adjust scope, naming, or drop this entirely if it
is not a direction you want for the project.